NEWSLETTER vol. A no. 1 The Directory will be useAjl to all types of census data users: data archives, libraries, planning organizations, other Federal agencies, State and local goverments, and the private sector. By developnent of the Directory , the Ceta User Services Division is providing a new information service to the entire user comunity. STATE AND REGIONAL DATA ARCHIVES (SRDA) Mirray A. Straus Ihiversity of New Hanpshire Durham, NH 0382'J (603) 862-1888 A vast anomt of data is available on a state by state basis, or by region. Political scientists, demographers, economists, and geographers have used this valu^le resource. Other social scieni Ists have rarely used this data. Part of the reason is that the data is scattered over many sources. Social scientists are often not aware of its potential. Even those vho realize the potential do not know the full range of variables. They also do not have convenient access. In March 1979 we therefore began to compile a truly ccrprdTensive archive of data on Anerican states and regions. This archive (SRDA) is intended to be the equivalent, for states of the Lhited States, of the Unan Relations Area Files and other archives of cross-national data. Vfell over ?,000 variables have been identified in the preliminary search. Other variables are being added continuously. The data files will be available in machine readable form (SPSS file, card image tape, or cards) or as s printed listing. IWo overlapping archives will make up the SRDA. The State Arc^ve consists of data on each of the 50 states and the District of Cblunbia. TV.ere are now about 1,00C variables actually entered in the State Archive in the form of an SPSS system file. The Regional Archive consists of data on the nine divisions which the U.S. Census uses as its main regional classification. All of the variables in the State Archive will also be in the Regional Archive. However, the Regional Archive will contain additional variables that are not in the State Archive. These are variables for v*ilch state by state data could not be obtained. The initial Regional Archive file will be created sonetime in WHY USE STATE LEVEL DATA? There are many problans connected with the use of state level data for social science research. Anong the most obvious is the ambiguity inherent in a statistic which corisines, for exeinple. New York City and the Adirondack rtxntain region. These problems, and also vhat is to be gained fran using state level data, are analyzed in my forthcaning book. State and Regional Analysis in the NEWSLETTER vol. k no. 1 Social Sciences . Among the reasons for using state level data despite these problems are: 1. States are the theoretically appropriate uiit for issues vhlch involve many aspects of goverment, politics, taxes, schools, econcmics, adninistration, and legal issues, as for example in Hicks, FV-iedland and Jcinson.* However, the uses of Sffift data will include many issues for vhich the states are not the natural ixiit. In such cases the decision to use state level, as in many research decisions, involves a trade-off: one accepts the problematic aspects of state data in order to be able to do research v*iich VDuld otherwise not be possible or practical. For exanple: 2. Historical analysis is possible because sane state level data goes back to colonial times and the nunber of variables available has grown exponentially each generation. 3. Causal inferences can aanetimes be more clearly established becaijse the sane data is avail^le for two or more time periods. This permits the use of time series and cross-lagged correlation analysis. ^4. Variables can be linked , even thDugh they are located in different surveys and refer to different respondents. This is possible by first converting the individual level data to state level data: for exanple, the percent in each state \to agree that "Homosexuality is alveys wrong" fhcm one study, with the percent who oppose the Equal Rights AnaKinent from another study. 5. Contextual analysis is possible by oonbining variables fran the SRDA with individual level data fVon specifc surveys. For exanple, Kersti Yllo has constricted a Sexual Inequality Index for each state and used this to find out if the correlates of a male-ctxninant marriage are the same or different in states viiere women are generally disadvantaged versus those with gi-eater equality between the sexes. TYPES CF DftTA 1 . Published lists of state by state data . Many examples are to be foind in standard almanacs and the Statistical Abstract of the United States . The largest soiree of pi4)lished data is the U.S. Census. 2. Aggregated survey data. These surveys were designed for analysis on an individual by individual basis. For the SRDA. the results are tabulated by state. An exanple is the percent in each state agreeing with sane attitude question. 3. Ccnpilations from docunents . Many docunents give the state as one Itan of information. This can be the basis for conpiling a new variable. For example. Who's Wx) in America gives the state of birth. It is possible to obtain a measure of the extent to vhich each state has contributed eninent persons by tallying the nimber of eminent Anericans bom in each state and dividing that by the population of the state. ^4. Indexes fran other variables . ^ ccmbining existing variables it is possible to produce entirely new variables, or to produce an index vhich does a better job of measuring than any one of the variables vhich are conbined to form the index. An exanple would be a measure of the •Wioks, Alexander, Roger Friedland and Edwin Johnson 1978 "Qass power and state policy: The case of large business corporations, labor uiions and governmental redistribution in the Anerican states." Anerican Sociological Review 13 (3): 302-315 3lA/SSIST NEWSLETTER vol. 4 no. 1 socioeconanic status of each state's population. This could be made up by combining the median income, education, and occupational prestige values from each state. DATA SOURCES Another vay of classifying the data in the SRDA is according to the accessibility of the source. Sections A and E list readily available data sources, each of which contains many variables. The main value of including then in the SRDA is convenience. Researchers using the SRDA data do not have to pinch, verify, provide variable levels, etc. Instead they can acquire a proofed, clean, labeled, ready-to-TLn data set. The full value of the S?DA, however, will derive fron the inclusion of data from sources which are not readily available and which often will not even be known to researchers. These cone from dozens of books and research reports, reports of goverment agencies, special topic reports by the census, and reports of private organizations such as the Institute of Life Insurance, the Audit Bureau of Circulation, the Boy Scouts of America, the National Wanens Political Caucus, etc. Finally, the least accessible data of all are the state level statistics created by aggregating individual level surveys to provide rates and averages for the states and regions. This is a long and expensive process for which the procedures are now being developed. The initial archive will consist mainly of data given for the 50 states (and the District of Colurtiia) in the sources below. A^. Milti-topic Ccnpendia Bacheller, Martin A. (Ed.) 1979 The HanTTond Almanac of a ^tLlUon Facts, Records, Forecasts. Maplewood, N.J.: Almanac, Inc. Book of the States 1978-1979 1978 Cotncil of State Goverments. Lexington, Ky. County and City Ebta Book 1947- Washington, D.C. Bureau of the Census. Demographic, Social and Economic Profile of States: Spring, 1976 1979 Current Population Reports, Series P-20, No. 33*^. Washington, D.C: Bureau of the Census. Information Please 1979 Information Please Almanac, Atlas, f, Yearbook, 33rd Edition. New York: Information Please Rjbl. The World Almanac 1979 The World Almanac i Book of Facts 1979. New York: Newspaper Enterprise Association. Roswe, Arthur E. (Ed.) 1978 1978-79 Help: The Useful Almanac. Washington, D.C: Consuner News. Statistical Abstract of the Uhited States 1878- Washington D.C: Bureau of the Census, U.S. Department of Connsrce. B. Specific Topics (Printed Sources ) lAlmaiac of Anerican Rolitics 1978 New York: E. P. Cutton lELreau of Labor Statistics 1/ 1978 Handbook of Labor Statistics. Washington, D.C: U.S. Department of Labor. ! -9- NEWSLETTER vol. k no. 1 Gottfredacxi, Michael R., Michael J. Hindelang. and Nicolette Parisi (Eds.) 1978 Sourcebook of Q-iminal Justice Statistics, 1977. Washington, D.C.: CriMnal Justice Research Center, Law Ehforcement Assistance Mministration, U.S. Department of Justice. Grant, W. Vance and C. George lind 1979 Digest of Eaucation Statistics 1979. Washington, D.C.: U.S. Department of Health, Education, and Welfare Education Division. Rosten, Leo (Ed.) 1975 A Guide to the Religions of Anerica. New York: Simon and Schuster, Vital Statistics of the Ihited States 1978 Washir^ton, D.C.: National Center For Health Statistics, HEW. C. Aggregated Survey Data Riysical Violence in Anerican Fanilies 1976 Survey conducted by Response Analysis Corp. for Mrray A. Straus, principal investigator. reOGRAM FUELICAnOB IN FR0C3?ESS AT THE UNTVERSmr OF NEW HAMFJSHRE Straus, Mrray A. Codebock for the State and Regional Ifeta Ardiives. (Anticipated availability for Part I, Variables 1 to 1,iJ99 is March, 1980; for F^rt n. Variables 1,500 to 2,999, Septnber, 1980). State and Regional rteta Eljlletin This quarterly pufclication will replace the Codebook beginning with variable 3,000. The BLLLETIN will improve on the Codebook in three ways: (1) Quarterly piAlication will make materials availd^le more quickly. (2) In addition to docunentating the source and nature of the data, it will include a printed listing of the statistics for each state and region. (3) T>ie BULLETIN will include news items dxxit the SRDft and occasional ccnmentary and analyses of data included in that issue. The planned pLfclication date for the first issue is January 1981.