^sist NEWSLETTER vol.1, no. 2 Planned several projects: (a) Development of a list of recommendations addressed to potential researchers on study desiqn as it relates to data man- agement (the "do's and don'ts" list). Group members are to send in their suggestions to the DOM AG coordinators. A consolidated list will be produced in Toronto for publication in the Newsletter and elsewhere, (b) Collation of information on relevant monographs, technical reports, program writeups, etc., which the AG members would like to share. The intent of the collection is to publish this information in a "What's New In Data flanagement" section of the Newsletter . 5. Prepared an agenda for the Toronto meetings; items for that agenda have been referred to above. The DOM AG coordinators would like to suggest the following revised mandate for future discussion at international lASSIST meetings: "This group addresses the problems of data organization and management confronting those archiving or usinq social science data. The AG will investigate and evaluate software and procedures for data and documen- tation preparation and management; recommend nuidelines for preparation procedures and software development; and, sponsor workshops and seminars for professional training and the exchange of information in these areas." [The original mandate includes "hardware" as an area to be addressed by this AG; see Newsletter Volume 1, Number 1 for the full text of the mandate.] AN OVERVIEW OF PROBLEMS ASSOCIATED WITH PROC ESS - PROD UC ED DATA / paul miitler For the August 1976 lASSIST meetings, Paul Muller preoared a report which provided a broad overview of problems associated with process-produced data. What follows is an edited version of this report. [Future issues of the News-- letter will contain additional Action Group reports. The membership is encour- aged to begin a dialogue on this and other issues of concern to the data archive community.] "Administrative Bookkeeping as a Social Science Data Base" by,, Paul Muller Institute for Applied Social Research University of Cologne 1.0 [...] I will try to give a rather broad overview of problems associated with the "production, acguisition, preservation, processing, distribution, and utilization of machine-readable" process-produced data. This report is necessarily biased by my own viewpoints and experiences; other exper- iences may well be fundamentally different. But,' the function of this paper is to initiate discussion and later, joint actions. ^sist NEWSLETTER vol.1, no. 2 1 . 1 Process-produced data 1.1.1 Definitions Within this action group we should be concerned with "process-produced data" as defined by Rokkan, as well as with "official bookkeeping data" (e.g., administrative registers). A common problem with these data is that they are/were not originally collected for scientific purposes and/ or within explicit statistical routines, thus creating the research situ- ation that the data are to a great extent already "given". The generic term, "process-produced data," would encompass all data that are/were not collected for statistical or scientific research (e.g., censuses or surveys), but are instead by-products or traces of the daily routines of private or public organizations or persons. 1.1.2 Priorities As a first step, we should concentrate on already-created machine-readable data within public administration. This would imoly structured mass-data. 1 .2 Problem areas 1.2.1 Inventory of machine-readable administrative data bases The first task is to obtain an overview of the existing data bases within the restricted domain as defined in 1.1.2. In West Germany, QUANTUM plans a pilot study for Morth-Rhine-Westfalia , "A continuous inventory of admin- istrative machine-readable data bases." Prior documentation projects have identified some 185 machine-readable data bases at the state level (North- Rhine-Westfalia) , as well as some 655 files at the local government level (City of Cologne) [...] as by-products of EDP use in public administration. These inventories should be [...1 descriptions of the data bases: e.g., coverage, period of correction/update, variables, etc. The experiences with the project in Germany show that updating would not occur without the active participation of the administration. The first step towards this end should be en overview of existing EDP routines within public admini- stration. This can be rather easily achieved by using the information of the existing coordinating committees/institutions within public administra- tion and by jointly establishing a check list. 1.2.2 Documentation Existing documentation of machine-readable process-produced data e.g.. National Archives and Records Service, Catalog of Machine-Readable Records in the National Archives of the United States , (Washington, D.C., 1975) or Directory of Computerized Data Files & Related Software , (Na- tional Technical Information Service, 1974) are models far less ambi- tious than the envisaged codebook-1 ike documentation. Good examples for this would be: Department of Health, Education, and Welfare, 1973 Cur - rent Population Survey - Summary Earnings Record Exact Match File Code - book, Part I - Basic Information Studies from interagency data linkages , by F. Scheuren, D. Vaughan, and W. Alvey. Report No. 5 (1975) and De- i^^sist NEWSLETTER vol.1, no. 2 partment of Health, Education, and Welfare, 1973 Current Population Survey - Summary Earnings Record Exact Match File Codebook, Part II - Supplemental Information, Studies from interagency data linkages , by F. Scheuren, B. Kilss, and C. Cobleigh. Report No. 6 (1975). 1.2.3 Preservation It is time to follow up the initiatives that were made by the Ruggles-, Kaysen- and Dunn Report in the United States. Specifically, we have to define research needs vis-a-vis national, state, and local archiving in- stitutions. In so far as some 95% of data generated within public ad- ministrations are physically destroyed, joint actions must be launched to define worth-while material to be preserved. In Germany, we have an ongoing discussion concerning an archive law which should take into con- sideration the specific interests of social science research (criteria for preservation, the archiving of temporal and/or cross section samples of linked files). The German data law will have an effect on the qual- ity of the archived data, in so far as it allows for selective destroy- ing of individual records within a register. It is yet very unclear whether this contamination effect can be avoided in some way. I propose beginning with an overview of existing archiving criteria (Kassation, physically destroying of data) within governmental archives and investigating whether these criteria are compatible with the kinds of uses thai; are not case-studies or oriented towards identified persons (e.g., historical). 1.2.4 Data Laws The data laws will have an impact on accessibility to administrative data for research purposes, not only with reg.ird to access to identifiable in- dividual data (very often necessary in the data collection and management phases). As these data laws do not provide for exceptions, serious re- search projects will be effectively hampered when these projects are not in the interests of a data providing agency. Data laws (or their drafts) are increasingly used for the secretion of public administration data (even in the cases when anonymous information are required/souaht) . . . . 1.2.5 Uses already made of process-produced data To ensure response to research needs for process-produced data, we should survey the uses made of administrative registers (i.e., those uses that were beyond the simple drawing of samples). In Germany, we will take another look at around 6000 identified research projects within the INFORMATIONSZENTRUM social science research project. Similar endeavors should be made in other countries as far as there ex- ists information on research projects. 1-2.6 Quality and characteristics of administrative bookkeeping There has been very little attention paid to problems with the quality of process-produced data (e.g., how are these data created, in what ways are the collection or validation processes biased) or to the characteris- 19. sist NEWSLETTER vol.1, no. 2 tics of official bookkeeping. There is a plethora of "validity studies," which compare sample survey data with official statistics, but as a spe- cific kind of study are not very cumulative. The quality of process-produced data and characteristics of administra- tive registers are the objects of an ongoing research project at the In- stitute for Applied Social Research in Cologne (Wolfgang Bick and Paul J. Muller) in which we analyse the laterality of representation of the in- dividual's (client's) social context and the temporal changes in regis- tration by public administration of everyday activities. 1.2.7 Record linkage Because these data are very often "meager" due to the limited administra- tive purposes for which they are collected, linked registers or data sets should be of the greatest interest. To the extent that administrative registers are organized within different life sectors, the potential of record linkage, (or family reconstitution as it is called within histori- cal demography), for supplementing existing large scale, but limited reg- isters is promising. A codebook-1 ike documentation of linked files should by envisaged (cf . 1 .2.2) Research on record-linkage techniques increasingly concentrates on the problems associated with statistical versus exact links (cf. Department of Health, Education, and Welfare, Some Observations on Linkage of Survey and Administrative Record Data, Studies from interagency data linkages , by J. Steinberg. (1973). The 60-odd exact-matching studies done in the last decades showed that without a unique standard indentifier (SSN, PK, person identification number) exact record-linkage would continue to be a very expensive task achieving only an average of 85-90% matches. We need a methodological breakthrough for synthesizing different data bases ac- cording to configurations of socio-economic characteristics. 1.2.8 Instruments We will repeat the mistakes made in earlier phases of the archive "move- ment" if we neglect the need for specific computer programs to handle process-produced, often "ragged" data; therefore, we should consider the problems of analyzing strings of data (e.g. within CROSSTABS or TROLL, perhaps) hierarchies of relations (e.g. area-block-house-fanily-respon- dent), or "life histories". It is good to hear that the SPSS-survey has already brought these issues to the attention of the SPSS-people. [...] 1.2.9 Interdisciplinary communication There is a real opportunity for intensified communication between those researchers working in quantitative history (the analysis of tax regis- ters, birth certificates, marriage licences, etc.) and social scientists who have already worked with or are interested in using process-produced data for their research purposes. This communication, which is already institutionally organized within QUANTUM, should enable us to document, 20 sist NEWSLETTER vol.1, no. 2 store, and distribute process-produced data in such a manner that the kinds of uses have not to be invented after completing all these tasks. In particular, the development of a "source criticism" for mass data, analogous to the development of the methodology of survey data, can only be achieved in an interdisciplinary way. Such efforts are necessary for the envisaged descriptors for machine-readable process-produced data. Similarly, cooperation with those people working on record-linkage prob- lems (e.g., Oxford Record-Linkage Study on medical record linkage) or with large-scale process-produced data (e.g. criminal statistics utili- zing court-records) should be initiated.... 1 .3 Quantities It is not possible to deal with these problems solely within the existing social science archiving movement which has so heavily concentrated on survey data. Other institutes must be brouaht into a network of archives and information centers in which coordination and a division of labor must be planned. (Examples of these include the "Sozialdatenbank" [Social-In- formation-System of the Departiflent of Labour and Social Affairs, West Ger- many], the proposed "Zentrum fur Aggregatdaten" [Centre for aggregate data of the German National Science Foundation], and the National Archives.") In Germany, the prospects for coordination are a little bit better because of the pioneering work being done within the Information- and Documentation Program of the Federal Government. BOOK N OT I C E S / Kathleen m. heim Introduction This column is a preliminary step in defining the literature of data ar- chiving. Those of us who have tried to assess the state of the art in order to formulate annual reports, write articles, or keep professionally informed have been frustrated by the lack of bibl iograohic control over our area of concern. Indexing and abstracting services such as Social Science Citation Index , Infor- mation Science Abstracts , Library Literature , Social Science Index and Resources in Education are unsystematic in their assignation of subject headings to pieces of literature related to data archives. The problem is further confounded by the fact that seminal information concerning the establishment of data archiving has often been distributed informally at conferences or in unpublished papers. When our numbers were small we could depend upon an invisible college network to disseminate important information. However, as our numbers grow and as new ar- chivists enter the field without access to the established network, it becomes mandatory that we define and organize the literature of our profession.