report of the gbif task group on biodiversity informatics, 7, 2010, pp.67 – 71 recommendations of the gbif task group on the global strategy and action plan for the mobilization of natural history collections data walter g. berendsohn(1), vishwas chavan(2) & james a. macklin(3) (1) dept. of biodiversity informatics and laboratories, botanic garden und botanical museum berlin-dahlem, freie universität berlin, königin-luise-straße 6-8, d-14195 berlin, germany (2) global biodiversity information facility secretariat, universitetsparken 15, 2100, copenhagen, denmark. (3) harvard university herbaria, 22 divinity avenue, cambridge, ma 02138, usa correspondence e-mail: w.berendsohn@bgbm.org abstract. – a task group to envision a global strategy and action plan for the mobilization of natural history collections data established by the global biodiversity information facility (gbif) has formulated three basic recommendations in order to increase the rate of mobilization of natural history collections data and improve the usage of this information resource: (i) gbif must facilitate access to information about non-digitized collection resources by publicizing the research potential of collections through metadata and assessing the number of non-digitized specimens; (ii) gbif must work with collections to continue to increase the efficiency of specimen data capture and to enhance data quality by means of technical measures, by means of ensuring attribution and professional credit and influencing institutional priorities, and by engaging with funding agencies; (iii) gbif must continue to improve and promote the global infrastructure used to mobilize digitized collection data through technical measures, outreach activities and political measures. key words. – natural history collections; collections; specimens; specimen data; metadata; digitization; gbif; biodiversity research. introduction mobilizing the biodiversity information intrinsic to the specimen holdings of natural history museums and herbaria of the world was one of the core aims of establishing the global biodiversity information facility (oecd 1999) and has been an integral part of its work program ever since. the value of specimen data for biodiversity research and natural resources management has been widely recognized (see references in scoble & bourgoin, 2010 and baird, 2010). gbif has created a functional technical infrastructure to discover and facilitate access to distributed data resources of primary biodiversity data (gbif, 2008a), including natural history collections data. as of july 2010, gbif facilitates access to more than 201 million such primary biodiversity data records, about 25.7% of which are specimen data (gbif, in prep) of which only 52.9% are geo-referenced. these are thought to represent a large proportion of the specimen data existing in digital form. effectively, the “low hanging fruit” have been picked and although the amount of specimen data increases steadily, the present growth rate is minimal when considering that estimates of the total numbers run in the range of 1.2 to 2.1 billion specimens (ariño, 2010) or even higher (see reference in vollmar & al., 2010). in 2008, the gbif secretariat concluded that a strategy and action plan is needed to incite further mobilization of specimen data. gbif constituted a task group of domain experts, with the aim of developing a draft action plan and strategy that is relevant at the global level and informs regional and national action plans for the digitization of natural history collections (gbif, 2008b). a work plan of the gbif task group for the global strategy and action plan for mobilization of natural history collections data was formally proposed to the gbif governing board in november 2008 (gbif, 2008c). mailto:w.berendsohn@bgbm.org berendsohn et al. – recommendations of the gbif task group the task group conducted a global survey to identify the barriers and challenges to digitization (vollmar et al., 2010) and undertook a prototype study to estimate the universe of the natural history collections data (ariño, 2010). it further held consultations with the natural history collections community and professional societies such as the society for preservation of natural history collections (berendsohn et al., 2009), and the consortium of european taxonomic facilities (cetaf). possible technical solutions were discussed in the biodiversity informatics communities, e.g. in tdwg (bourgoin et al., 2009). an interim report of the task group was presented and circulated at the 16th meeting of the gbif governing board held in october 2009 in copenhagen. the task group submitted its report to the gbif science committee in april 2010. the articles in this volume represent key parts of this report. in the following a summary of the conclusions the task group arrived at is given, they are based on the discussions during meetings and workshops and more details are provided in the individual articles cited. general conclusions the development of a global strategy and action plan for the mobilization of natural history collections data must be guided by a single basic strategic principle: user demand must be the driver of the detailed digitization of individual specimens. accordingly, priorities for digitization should be set either according to the demand from on-going or projected research, or in accordance with socio-political demands (conventions etc.). funding of digitization activities should be linked directly to these priorities, with the costs either incorporated into research proposals, or covered by (international) organizations, foundations, or governments. we posit that a demand-driven approach makes the size of the digitization task realistic and fundable, because the task can be focused and outside funding can be mobilized. however, a principal obstacle is the lack of access to potentially useful un-digitized collection holdings. the following three recommendations refer to helping users to address their demand for information, helping collections to answer that demand, and to the provision of the infrastructure necessary to transport the results to the user. recommendations 1: gbif must facilitate access to information about non-digitized collection resources 1.1. publicize the research potential of collections through metadata • metadata (i.e. higher-level information about the content of a collection) are an essential infrastructural component enabling demanddriven specimen-level digitisation (berents et al., 2010, scoble & bourgoin, 2010, berendsohn & seltmann, 2010). gbif participants and national funding agencies must provide funding for metadata capture, publication and discovery in order to build this essential research infrastructure. • collections must investigate the institutionspecific costs of metadata capture. • collections must expose their metadata holdings to promote the potential data availability to users and to expand the user base. • gbif must assist collections to use appropriate means to publish standard metadata to describe their holdings. o gbif and organizations of collection directors and custodians should work together to establish disciplinespecific best practice guidelines for metadata capture. o gbif must set up or identify data capture mechanisms allowing collection managers or metadata authors to report numbers of specimens that belong to certain metadata categories. o gbif must foster development of controlled vocabularies useful for authoring enriched metadata documents. o the metadata describing a collection represent publishable content. gbif should promote scholarly credit for metadata authoring, in order to 68 berendsohn et al. – recommendations of the gbif task group provide an incentive for quality metadata provision. • gbif and the collections should initially focus on metadata records documenting – for the entire collection – the number of un-digitized specimens of a particular (higher) taxon from a specific country or region of origin. large scale research questions as well as data repatriation efforts are centered upon these data areas. currently there is no straightforward way for interested parties to assess the amount of data that should by digitized for their purposes (berendsohn & seltmann, 2010). • gbif should create a mechanism to query the metadata (e.g. who has how many specimens of amphibia from tanzania) and to send requests initiating communications about demand for specimen-level digitization between a user and a collection manager (berendsohn & seltmann, 2010). • the gbif secretariat should set up a pilot metadata system (bourgoin et al., 2009), using the taxonomic groups suggested at the leiden workshop (berendsohn et al., 2010) as “doable targets”. 1.2. assessing the number of un-digitized specimens • in parallel to the metadata effort, gbif should use an analytical or statistical approach to determine the size of the holdings of undigitized specimens world-wide (ariño, 2010). 2: gbif must work with collections to continue to increase the efficiency of specimen data capture and to enhance data quality. 2.1. technical measures • gbif must investigate and document ways to industrialize specimen digitization. • gbif must foster the development and dissemination of collection type-specific procedures, best practice guidelines and tools aimed at streamlining workflows for data entry, imaging, text and feature recognition, georeferencing, and general quality enhancements. • gbif should develop and promote simple, easy-to-use, intuitive and efficient data capture tools that would be accessible to all user skill levels. • gbif should foster the creation of ‘data dictionaries’ in order to disperse the training load and reduce the number of input errors. • gbif and relevant societies/organizations must investigate innovative ways to use citizen science approaches for specimen-level digitization. 2.2. attribution, professional credit, institutional priorities • gbif must continue to foster digitization related capacity building and training activities. • natural history collection management should develop best practice and policy for data mobilization and curation accompanying their management of the physical collection. • gbif and other relevant stakeholder organizations must promote the recognition of natural history collection data publishing as a scholarly and scientifically useful exercise. • collection administrators and/or management must make efforts to allocate dedicated human resource for digitization activities. • gbif must work together with collection managers to help recognize digitization as a component of their formal job description. • digitization of type specimens (all type material) should be made a prerequisite for the publication of new names (berents et al., 2010). digitization and public data availability for all specimens used in a taxonomic revision should be recognized as best practice and made mandatory by collections. 2.3. funding • gbif must actively engage with funding agencies in convincing them of the significance of support for natural history collections digitization. • gbif must continue its seed money award scheme as it has acted as a catalytic agent and 69 berendsohn et al. – recommendations of the gbif task group created ripple effects in encouraging collections and many national and international donor agencies to support digitization. • science funding agencies and private donors must fund digitization where the research demand exists. • collections must investigate the institutionspecific costs of specimen data capture. 3: gbif must continue to improve and promote the global infrastructure used to mobilize digitized collection data 3.1. technical measures • gbif must encourage hosting environments to facilitate discovery and publishing of natural history collections data especially for small and mid-sized museums. • gbif must accelerate progress toward allocation and resolution of persistent identifiers at the dataset and data record level, including services and best practice guidelines. 3.2. outreach activities • gbif should make special efforts to include collections from the southern hemisphere. • gbif must pay special attention to small and medium sized collections based in research and academic institutions and their organizations, as they may be able to contribute with minimal investment and encouragement. 3.3. political measures • national and/or thematic digitization strategies must be developed to identify and draw from possible synergies with digitization efforts in other domains (e.g. primary data, archives, print publications). outlook the recommendations presented are intended to form the base for further discussion. some of the questions raised will only be answered when action is taken. it is clear, however, that the natural history collection community has to work closer together to address the question of how to mobilize their information resources in an efficient and usable way. further implementation of these recommendations calls for socio-political decisions, cultural changes, as well increased infrastructural, technological and financial investment by the stakeholder communities at all levels from local to global scale. references ariño, a. h. 2010. approaches to estimating the universe of natural history collection data. j. biodiversity informatics 7: 81-92. baird, r. 2010. leveraging the fullest potential of scientific collections through digitization. biodiversity informatics 7: 130-136. berendsohn, w. g., berents, p. and macklin, j.a. 2009. results of the leiden workshop on digitisation priorities. report from a workshop during the 2009 annual meeting of the society for the preservation of natural history collections (spnhc)1. [accessed september 1, 2010] berendsohn, w. g. and seltmann, p. 2010. using geographical and taxonomic metadata to set priorities in specimen digitization. j. biodiversity informatics 7: 120-129. berents, p., hamer, m. and chavan, v. 2010. towards demand-driven publishing: approaches to the prioritization of digitization of natural history collections data. j. biodiversity informatics, 7: 113-119. bourgoin, t., berendsohn, w. g. & macklin, j. a. 2009. the natural history collections geotaxonomic index. report from the working session during the 2009 meeting of the organisation for biodiversity information standards (tdwg) in montpelier2. [accessed september 1, 2010] gbif 2008a: gbif work programme 2009-20103. [accessed september 9, 2010] gbif 2008b: terms of reference for “task group on a global strategy and action plan for the mobilisation of natural history data”4. [accessed august 29, 2010] gbif 2008c: global strategy and action plan for the digitization of natural history collection data. 1 http://www2.gbif.org/gsap-nhc_spnhc09_report.pdf 2 http://www2.gbif.org/gsap-nhc_tdwg09_workingsession.pdf 3 http://www2.gbif.org/wp2009-10.pdf 4 http://tinyurl.com/gsaptg 70 http://www2.gbif.org/gsap-nhc_spnhc09_report.pdf http://www2.gbif.org/gsap-nhc_tdwg09_workingsession.pdf http://www2.gbif.org/wp2009-10.pdf http://tinyurl.com/gsaptg berendsohn et al. – recommendations of the gbif task group 71 t 29, 2010] preliminary report to the gbif governing board 15 in arusha5. [accessed september 12, 2010] gbif in prep.: discovery and publishing of primary biodiversity data through gbif network: the state-of-the-art and potentials (version 2.0), 52pp. copenhagen. oecd 1999. final report oecd megascience forum working group on biological informatics6. oecd, paris, jan. 1999. [accessed augus 5 http://www2.gbif.org/gsap-nhc-gb15reportwithannex.pdf 6 http://www.oecd.org/dataoecd/24/32/2105199.pdf scoble, m. j. and bourgoin, t. 2010. natural history collections digitization: rationale and value. j. biodiversity informatics 7: 77-80. vollmar, a., macklin, j. a., and ford, l.s. 2010. natural history specimen digitization: challenges and concerns, j. biodiversity informatics 7: 93112. http://www2.gbif.org/gsap-nhc-gb15reportwithannex.pdf http://www.oecd.org/dataoecd/24/32/2105199.pdf introduction general conclusions recommendations outlook references microsoft word biguentschtoqeproof1.doc biodiversity informatics, 6, 2009, pp. 53-58 53 effectively searching specimen and observation data with toqe, the thesaurus optimized query expander a. güntsch, n. hoffmann, p. kelbert and w. berendsohn freie universität berlin, botanic garden and botanical museum berlin-dahlem, königin-luise-str. 6-8, d-14195 berlin, germany abstract.⎯ today’s specimen and observation data portals lack a flexible search mechanism, able to link up thesaurus-enabled data sources such as taxonomic checklist databases and expand user queries to related terms, significantly enhancing result sets. the toqe system (thesaurus optimized query expander) is a rest-like xml web-service implemented in python and designed for this purpose. acting as an interface between portals and thesauri, toqe allows the implementation of specialized portal systems with a set of thesauri supporting its specific focus. it is both easy to use for portal programmers and easy to configure for thesaurus database holders who want to expose their system as a service for query expansions. currently, toqe is used in four specimen and observation data portals. the documentation is available from http://search.biocase.org/toqe/. key words.⎯ biocase, gbif, edit, query expansion, specimen portal, synthesys, taxonomy, thesaurus, xml. over the last decade, specimen and observational data portals have been developed with a wide variety of thematic scopes and technologies. however, all share the goal of offering unified access for distributed and often very heterogeneous data sources using standardized data schemas (e.g. abcd and darwincore) or common access protocols (e.g. biocase, digir, and tapir). early systems such as the species analyst network (vieglas, 2003) and the european natural history specimen information network (enhsin, n.d.; güntsch, 2002) used entirely distributed query mechanisms, instantly propagating user requests to the respective networks. the rapidly growing number of available data providers and data records led to serious performance and accessibility issues, and made this strategy infeasible. today, the vast majority of specimen and observation portals use index databases containing a projection of the networked concepts, which usually consist of a limited number of elements considered fundamental for typical queries. as a result, core data elements can be processed and retrieved with almost no delay. in a second query step, full data records can be accessed directly from the data provider as needed. the most complete index database, having a worldwide scope and used by several different portal systems, is the gbif index. presently it contains 180.000.000 data records (august 2009). _________________________ correspondence email: biodiversityinformatics@bgbm.org. apart from the obvious performance and stability benefit, the usage of an index database offers the opportunity to harmonize the highly heterogeneous vocabularies used by distributed and independent data providers. these can include: various naming conventions for scientific organism names; different taxonomic classifications for the same organism; misspelled country codes; varying representations of geographic coordinates; and different spellings of person names (e.g. collectors). these disparities can be partly resolved in the indexing process, for example: by mapping taxonomic names to standard taxonomies; by identifying and correcting potential misspellings; and by translating and converting data formats into standardized and common representations. however, even after the data stored in a common index database has been harmonized to a large extent, additional knowledge about the potential relations between terms can significantly broaden the result set for a given user query. a taxonomic checklist database might for example contain information about existing misapplications of scientific names. specimens or observations identified under the misapplied name will not be found if the query term is only the correct name. there are many specialized checklist or thesaurus databases which could potentially support portals by expanding user queries to related terms, such as taxonomic checklist databases, gazetteers, lists of person names, stratigraphic term lists, etc. it does biodiversity informatics, 6, 2009, pp. 53-58 54 not necessarily make sense to use all of them in a comprehensive system such as the gbif search portal. but specific regional or thematic networks can be improved significantly by using a selected set of thesaurus systems relevant for the given focus. the thesauri should therefore ideally be equipped with a service layer, one which is both easy to implement by the existing thesaurus database and easy to use by portals for query expansions. portal systems could then choose from the set of available thesauri and invoke them for their specific regional or thematic foci. the thesaurus optimized query expander toqe service was defined and implemented in the context of the eu 6th framework project synthesys (synthesis of taxonomic resources, synthesys, n.d.), and uses this approach to support queries to the biological collection access service for europe (biocase, 2005-) with high quality taxonomic data available from fauna europaea (fauna europaea, 2004), euro+med plantbase (euro+med, 2006-), and the european register of marine species (erms, 2004). we tried to make the toqe implementation as generic and configurable as possible with regard to the underlying thesaurus database, which led to many more provider databases and portals than originally planned using the services. in the following, we will describe the toqe service and its implementation, its primary deployment in the european biocase portal, and further implementation beyond the initial scope of the project. the toqe service a toqe service can deliver a set of concepts using a specific term, such as a set of taxa returned by a given scientific name in a floristic database. the service can then be used to retrieve related concepts (e.g. included taxa, misapplied names, synonyms) for the concepts that have been identified in the first step and for the terms being used (fig. 1). the resulting term-list is then used to expand the original user query. toqe itself is an xml web-service offering five methods supporting and partly simplifying the twostep query mechanism. a toqe-method is called with a get-request. the following example shows a call of the method getconceptsbyterm querying for all concepts in a given thesaurus using the term “calendula arvensis”: the syntax of the corresponding xml responsedocuments is defined in the toqe schema1. the following response belongs to the above function call: a full implementation of the toqe service offers the following five methods: getmethods() returns the list of methods implemented by a given toqe implementation. getmethodinfo(methodname: string) returns a description of a specific toqe method (required argument methodname). getconceptsbyterm(term: string, thesaurus: string) receives a search term and the name of a thesaurus connected to the given toqe instance and returns the list of concepts matching the search term in this thesaurus. the ‘%’ character can be used as a wildcard within the search term. both arguments are required. getrelatedconceptsbyterm(term: string, relation: string[], thesaurus: string) is more powerful than getconceptsbyterm and returns both matching concepts (ids, names, references, status values) for the search term and related concepts. 1 http://search.biocase.org/toqe/schema/ text box 1: http://search.biocase.org/toqe/toqe.py?term=cale ndula+arv%&thesaurus=standardliste&method= getconceptsbyterm text box 2: calendula arvensis @authorl. accepted r. wisskirchen et h. haeupler standardliste der farnund blütenpflanzen deutschlands 1998. biodiversity informatics, 6, 2009, pp. 53-58 55 figure 1. scheme of how queries are expanded using toqe. the relation types to be analyzed can be passed with the repeatable relation argument. again, the thesaurus to be used must be specified. all arguments are required. getallrelatedconcepts(conceptkey: string, relation: string[], recursive: boolean, thesaurus: string) receives a single concept key, a list of relation types to be analyzed in the given thesaurus and returns all related concepts (ids, names, references, status values). additionally, the boolean argument recursive is used to indicate whether the search should return concepts that are not directly related to the given concept and have to be searched recursively. the semantics of a recursive search depends on the thesaurus being queried, as well as on the available computing resources and the capabilities of the database management software in use. an example of a strategy for recursive searches of taxonomic concepts in the euro+med plantbase is here2. the recursive argument is optional, all other arguments are required. a more detailed description of the service, including the set of error-responses, is available here3. so far, a registry service for toqe has not been implemented, so client software systems are required to know the available and appropriate toqe services. the toqe software implemented in the context of the biocase portal development has two basic layers (fig. 2). the service layer processes the 2http://search.biocase.org/bgbm/static/extrasearch/extrasearch.htm. 3http://search.biocase.org/toqe/api.html#client. incoming http get-request and calls the associated database access method, hiding database specific functionalities in a generic way. the returned records are then wrapped up in xml response documents, following the toqe schema specification. the native database queries are generated in the toqe database layer. a database module has to be configured for each thesaurus database connected to the toqe service layer. all software components have been implemented using the python programming language. the source code is freely available and can be obtained from the bgbm subversion repository here4. the biocase portal and toqe in spring 2008, biocase released a new data portal which uses the toqe query expansion service. the biocase portal uses a subset of gbif specimen and observation data records referring to organisms that have been collected or observed in europe (holetschek et al. 2006). accordingly, the thesaurus systems connected with toqe are the three major european taxonomic checklist systems: euro+med plantbase, fauna europaea, and european register of marine species (erms). figure 3 shows the biocase response for a user query for calendula arvensis. the original query is expanded using 2 misapplied names, 32 synonyms, 4http://ww2.biocase.org/svn/synthesys/trunk/thesaurus. biodiversity informatics, 6, 2009, pp. 53-58 56 figure 2. schematic showing toqe implementation. and the higher taxon calendula,as well as a number of spelling variants. this leads to 3139 hits in the gbif index, compared to 2844 hits for the original query without query expansion. the portal returns both the expanded list of terms (scientific names) used in the query and the list of terms with actual hits in the gbif index database, together with the respective number of units which will be returned. users can now deselect terms which should not be contained in the specimen result set. other implementations because both biocase portal software and the toqe thesaurus service are highly generic and configurable, the implementation of additional portal systems with differing scopes can be accomplished with relatively little effort. up to now, the following additional toqe-enabled portals have been developed: checklist driven access to european biodiversity data (prototype)5: like the standard biocase portal, the system gives access to specimens and observations from europe with expanded queries using the major european floristic and faunistic taxonomic checklist databases. in contrast to the biocase portal, users 5http://search.biocase.org/toto/. have full control over the query expansion process and can freely select the thesaurus systems to be used and the relation types to be considered. edit specimen and observation explorer for taxonomists6: the portal offers an interface to all gbif data specifically tailored to serve the taxonomic work process. the toqe query functions are directly integrated into the specimen search interface so that taxonomists can easily pick from the list of available thesaurus systems. in a second query step, a clearly laid out form summarizes terms and relations retrieved from the selected thesauri for the given query. the system will be an important component of the edit platform for cybertaxonomy (edit, 2007-; berendsohn et al., 2007; döring, 2007). biocase portal for bgbm collections7: the botanic garden and botanical museum berlindahlem has linked almost all its collection databases to gbif, ranging from smaller collection databases for particular herbarium subcollections to the comprehensive accession management system of the botanical garden’s living collection. this opened up the opportunity to set up a portal specifically dedicated to bgbm collections. so far, the german standard lists for ferns and flowering plants and the reference list for german bryophytes 6http://search.biocase.org/edit/. 7http://search.biocase.org/bgbm/. biodiversity informatics, 6, 2009, pp. 53-58 57 figure 3. query expansion for the term calendula arvensis using toqe in biocase. the original result set is increased from 2844 to 3139 records. biodiversity informatics, 6, 2009, pp. 53-58 58 as well as euro+med plantbase and erms, have been added as thesaurus databases. the new portal offers a convenient new access point to bgbm collections which were previously lacking a common and unified representation. conclusions expanding queries to specimen and observational databases using the relevant thesaurus systems for a given portal scope can lead to richer query responses, with results which would have been ignored using traditional query mechanisms. we have demonstrated this approach with the definition and implementation of the toqe-service now in use by several specimen and observation portals together with a variety of taxonomic thesaurus systems. the approach should now be broadened to include non-taxonomic concepts such as place names, adding further value to specimen portals. we also believe that the query expansion process will require a sophisticated interactive explanation component monitoring the query mechanism transparently to the end-user. acknowledgements this work was funded by the network activity d of the european union 6th framework project synthesys (contract rii3-ct-2003-506117). the authors thank jörg holetschek, wolf-henning kusber and elke zippel of the synthesys project team for their critical comments and assistance during the implementation phase. we would also like to thank the fauna europaea, erms, and euro+med initiatives for the unbureaucratic provision of data. our special thanks go to pepe ciardelli for proofreading of the manuscript. literature cited berendsohn, w. g., m. döring, and m. c. ebach 2007. edit needs biodiversity information standards. p. 1 in: weitzman, a., and l. belbin ed. abstracts of the 2007 annual conference of the taxonomic databases working group in bratislava. biocase 2005-. biological collection access service. accessed 20 august 2008 from http://www.biocase.org. döring, m. 2007. a general concept for the design of the edit platform for cybertaxonomy. edit newsletter 3. muséum national d’histoire naturelle, paris. edit 2007-. edit platform for cybertaxonomy. accessed 20 august 2008 from http://wp5.etaxonomy.eu/blog/index.php. enhsin n.d. european natural history specimen information network. accessed 20 august 2008 from http://www.nhm.ac.uk/researchcuration/projects/enhsin/index.html. erms 2004. the european register of marine species. accessed 20 august 2008 from http://www.marbef.org/data/erms.php. euro+med 2006-. euro+med plantbase. accessed 20 august 2008 from http://ww2.bgbm.org/europlusmed/query.asp. fauna europaea 2004. fauna europaea. accessed 20 august 2008 from http://www.faunaeur.org/. güntsch, a. 2002. the enhsin pilot network. p.p. 3340 in scoble, m. j. ed. enhsin – the european natural history specimen information network. the natural history museum, london. holetschek, j., a. güntsch, c. oancea, m. döring, and w. g. berendsohn 2006. prototyping a generic slice generation system for the gbif index. pp. 51-52 in belbin, l., a. rissoné, and a. weitzman, a. eds. abstracts of the 2006 annual conference of the taxonomic databases working group in st louis, mi. synthesys n.d. synthesis of taxonomic resources. accessed 20 august 2008 from http://www.synthesys.info/index.htm. vieglas, d. 2003. species analyst. accessed 20 august 2008 from http://www.faunaeur.org/. rational and value of nhc digitisation biodiversity informatics, 7, 2010, pp. 77 – 80. natural history collections digitization: rationale and value malcolm j. scoble department of entomology, natural history museum, cromwell road, london sw7 5bd, uk, m.scoble@nhm.ac.uk abstract. – the value of digitizing natural history collections is well attested, although their rate of digitization should be increased. the task group of the global strategy and action plan for the digitization of natural history collections agreed that the only way of achieving a significant increase is by capturing metadata to encourage digitization at the specimen level. encouraging a metadata solution appears to be the best way of mobilizing the community responsible for caring for and providing access to such data. moreover, a user-driven approach is likely to offer the best means of prioritizing what should be digitized. key words. – natural history collections; digitization, museums, herbariums, metadata, gbif. questions about the natural world may be addressed by natural science collections, even if they comprise a small component of a more extensive source of data. although there may be as many as 2-3 billion specimens in natural science collections across the world (duckworth, genoways and rose 1993; ariño 2010), the amount of data is not large compared with the vast and increasing number of digital observations produced and used by monitoring and other projects. although data in collections are complementary to these other data, they offer, when digitized, an exceptional resource. within collections lies the most extensive dataset that exists of the planet’s biodiversity – a dataset that is relevant to research and decision-making. unlike observational data, which are restricted to relatively few of the estimated 2 million described species, collections hold a recoverable record of what species exist, or have existed over the past three hundred years (to a time even before linnaeus), and where they occur or occurred. this record is usually biased: so far, collections are anything but a representative sample of life on earth, although they provide, at least, minimum estimates. indeed, our entire knowledge of most species is based on just one or very few specimens housed in natural science collections. nevertheless, for most species, collections provide us with the best public record available. and, unlike observational data, the physical presence of specimens allows us to examine them many times using new techniques. examples include the extraction and study of molecular data, although many museum curators restrict destructive use of specimens. the recent push for stable isotopes, dna and fatty acid analysis presents a good example. collections are physical databases of the natural world. the specimens they house contain a wealth of taxonomic, spatial and temporal data, albeit with much variation in detail, quality and coverage. the problem is that most of this information is trapped in various museums, herbariums and private holdings. while it is accessible to bona fide researchers who have the means to visit collections, few outside the discipline of taxonomy have made much use of it. recently, modern methods have given us the capability to capture digital information from collections and expose it globally through computerized networks via a web interface either as data associated directly with individual specimens or as metadata describing collections. natural science collections have been used mostly for taxonomy (or systematics), the discipline associated with inventorying the planet’s biodiversity and describing its evolutionary relationships. these will remain scoble – natural history collections digitization: rationale and value primary tasks, although doubtless one that will occur increasingly online in a collaborative virtual environment. with the means of digitization at our disposal, however, we might not only help improve the species inventory but also realize a far wider potential of biological collections. such data could help researchers address questions on natural resource inventories, the effect of environmental change on biodiversity, on how to gain a better understanding of species distribution over time and why changes in species’ distributions have occurred. stated more broadly, specimen information in collections still has the great potential to be used in research, resource management, education and sustainability science. digitization priorities should, therefore, be set with the wider user community in mind. yet for the most part they have not, although there are notable exceptions where much taxonomy-related information is available (e.g. in fishbase, vertnet, gbif). while digitizing material in collections has a wider value, the effort expended will have to be proved cost-effective against other demands on those who digitize. this question is relevant at any time, but particularly so in a straightened economic environment. yet, there are well documented and compelling case studies of the use of information from natural history collections, particularly when integrated with data from research programs especially in the fields of biogeography, ecology and evolution (graham et al. 2004). clearly, both volume and quality of data are critical factors, which means that if digitization is to be achieved on a more comprehensive scale, a shift in the working patterns and current aims of curators and others managing collections will be required. in particular, effort will need to be prioritized, focused and sustained. the value of digitization locked up in collections, is information potentially relevant to broad questions that require information about species diversity and species distribution and their change through time. much information is available already from the global biodiversity data portal (http://data.gbif.org), but much more resides undigitized in collections. there is more to do in the mobilization of primary species data in collections in terms of encouraging further digitization and prioritizing what should be digitized. but collection managers should take heart from encouraging noises being made about the value of natural history collections to address a variety of questions. one of the most comprehensive accounts of the value of digitizing specimens in collections was written by chapman (2005), in a paper on the uses of species-occurrence data. this was preceded by a shorter review by graham et al. (2004). by providing a series of examples, chapman examined the uses both of specimen and observation data to a very wide range of fields from taxonomy (the traditional use), through biogeography, species diversity and invasive species, to education, and art and history. the value of specimen data in collections for addressing these questions varies considerably, with those from modern surveys having a greater variety of uses, primarily because of the higher quality of information associated with more recent specimens. it is crucial that appropriate metadata are collected to enable the widest use of this information. chapman also addressed the criticism that museum specimen data were outdated and unreliable, explaining that while some undoubtedly are, many are not only usable but can be improved by rendering them accessible across digital networks. a compelling case for the use of natural history collection data in modeling, alone or, particularly, combined with other types of data, was made by graham et al. (2004). these authors noted, for example, studies demonstrating the value of such data in predicting future distribution of invasive species. scaling up the response of curators to digitizing natural science collections has been rather more technology driven than strategically planned. but many curators and managers have adopted new technology to digitize specimen data opportunistically, often without any obvious purpose. where digitization has been more purposeful, it has been developed largely for the close community of taxonomists rather than the wider group of stakeholders and global users. while there are exceptions to this statement, largescale digitization is more likely to succeed if taxonomists are familiar with the major 78 http://data.gbif.org/ scoble – natural history collections digitization: rationale and value applications and if they forge partnerships with users and align their digital outputs with producers of other kinds of digital data. if a functional global infrastructure for biological collections is to be achieved, a responsiveness is needed that contributes to and helps those working in domains other than just in natural history collections. routine digitization can certainly contribute towards building capacity, even if in a limited way. for example, curators now frequently digitize label data and make images of specimens that can be sent to borrowers in lieu of the actual specimens, or they may lend the specimens while keeping the digital data as security. activities of this kind can form valuable contributions to mobilizing the digitization of natural history collections, but it is a relatively slow way of building a critical mass of digital objects. automated or semi-automated digitization is probably the only way in which digitization at the specimen level will be achieved at anything like the scale needed if data are to be useful for more quantitative studies. some kinds of specimens are intrinsically easier to subject to such an approach. herbarium sheets are probably the best example, having the advantage of being mounted flat with associated data written on labels attached to the sheet in the same plane. there are many examples of herbarium sheets being digitized as major scanning programs, in the botanischer garten und botanisches museum (berlin), kew gardens (london), and many other herbariums. specimens on microscope slides share with herbarium sheets similar characteristics. by contrast, digitization of label data on dried insect specimens poses an immense challenge given that labels are attached, often as a series, on the same pin as the specimen and underneath it. this arrangement renders it impossible to photograph drawers of specimens together with the associated specimen label data. solutions, which appear promising, are being sought for specimens of all kinds housed in drawers (blagoderov et al., 2010). digitization of natural science collections (images, information or digital surrogates) has often been undertaken because it is worthy and increasingly expected rather than overtly targeted to specific uses. although further digitization of collections (whether at specimen or metadata level) will almost certainly lead to uses as yet unanticipated, particularly when a critical mass of data becomes available, a more strategic approach to digitization will be developed better by partnering with users (particularly ecologists). creating partnerships is more beneficial than simply anticipating user needs. progress might be made by mobilizing the user community to fund digitization for specific purposes and to use offers of funding to prioritize, whether it be for a specific project or user community (e.g. specific taxa across many collections, or all collections from one area – for example for data repatriation). the metadata approach a metadata approach seems the most expeditious solution to the challenge of digitizing collections. berendsohn and seltmann (2010) state that capturing metadata is the only realistic solution to providing the scale of digitization across all kinds of collections that will mobilize the data gathering in a timely way. capturing specimen-level data from all collections is simply not possible in the short or even medium term with the kind of resources available to the collections community. this statement was confirmed by a survey of 228 respondents to the task group of the global strategy and action plan for the digitization of natural history collections. it is certainly not meant to suggest that the prioritized capture of specimen-level data should not be undertaken, for there are certain kinds of collections that are eminently capable of being digitized at scale (herbarium sheets being particularly tractable as already noted) or that have been digitized on a large scale already (e.g. the us vertebrate collections). but providing metadata is a means of providing the community with an understanding of what is potentially available at the specimen level across the universe of biological collections. such metadata might include the number of specimens an organization holds for a particular taxon and the country or countries of origin. this approach is demanding enough in its own right, but if collections-rich organizations are resolved to digitize their holdings, this method would provide a realistic and cost-effective initial solution to the problem of scaling up the digitization of collections. an important proviso is that the 79 scoble – natural history collections digitization: rationale and value 80 providers should aim for high quality metadata (see e.g. global biodiversity information facility 2008). the main conclusion of the task group was that metadata records describing a collection are an essential and achievable prerequisite for meeting user demands and best practices, but that this process could and should lead to prioritized specimen-level digitization. a metadata approach to the digitization of collections will in itself help achieve a number of ends. notably, it will facilitate the finding and using of data, identify gaps in coverage of our samples of biodiversity, help expose errors and other shortcomings of data quality, improve quality and peer-review, provide data on density of sampling to assess their suitability for analytical work, accelerate data capture, and improve the management and enhancement of collections. achieving even this metadata solution is an immense task and gbif has key roles to play as a facilitator for the international community of collection holders and as a data broker. it will require resolve and persistence to keep the international collections community engaged actively in digitization, a process that has hardly started in terms of what needs to be achieved to form a critical mass of information. it will also require leadership in acting as a forum for the debate about how to prioritize what to digitize. production of metadata about collections, however, will help inform the public and scientists alike. acknowledgements we thank three referees for thoughtful comments. references ariño, a. 2010. approaches to estimating the universe of natural history collections data. biodiversity informatics 7: 81-92. berendsohn, w.g. & seltmann, p. 2010. using geographical and taxonomic metadata to set priorities in specimen digitization. biodiversity informatics 7: 120-129. blagoderov, v., kitching, i., simonsen, t. and smith, v. report on trial of satscan tray scanner system by smartdrive ltd. available from nature precedings http://hdl.handle.net/10101/npre.2010.4486.1 (2010) chapman, a. d. 2005. uses of primary speciesoccurrence data, version 1.0. report for the global biodiversity information facility, copenhagen. duckworth, w.d., genoways, h.h. & rose, c.l. 1993. preserving natural science collections: chronicle of our environmental heritage. washington, d.c.: iii+140 pp. global biodiversity information facility. 2008. gbif training manual 1: digitisation of natural history collections data, version 1.0. copenhagen: global biodiversity information facility. graham, c. h., ferrier, s., huettmann, f., moritz, c. and peterson a.t. 2004. new developments in museum-based informatics and applications in biodiversity analysis. trends in ecology and evolution 19: 497-503. http://hdl.handle.net/10101/npre.2010.4486.1 the value of digitization scaling up the metadata approach acknowledgements references one of the incentives for scientists to publish research articles, conference papers, research monographs or patents is the explicit recognition that they receive by means of citations from fellow scholars biodiversity informatics, 7, 2010, pp.113 – 119. towards demand-driven publishing: approaches to the prioritization of digitization of natural history collection data penny berents (1), michelle hamer (2), and vishwas chavan (3)* (1)australian museum, 6 college street, sydney, nsw 2010, australia. email: penny.berents@austmus.gov.au (2) south african national biodiversity institute, 2 cussonia ave, brummeria, pretoria, south africa and school of biological & conservation sciences, university of kwazulu-natal. email: m.hamer@sanbi.org.za (3)global biodiversity information facility secretariat, universitetsparken 15, dk 2100, copenhagen, denmark. email: vchavan@gbif.org *corresponding author abstract – natural history collections represent a vast repository of biodiversity data of international significance. there is an imperative to capture the data through digitization projects in order to expose the data to new and established users of biodiversity data. on the basis of a review of the current state of digitization of natural history collections, a demand-driven approach is advocated through the use of metadata to promote and increase access to natural history collection data. key words. – natural history collections, publishing, demand-driven digitization, prioritization, metadata, biodiversity data natural history collection data is a critical component of biodiversity data which has widespread application in biodiversity research, natural resource management and biosecurity (chapman, 2005; tann, et al., 2008; pyke & ehrlich, 2010). the global biodiversity information facility (gbif) global strategy and action plan for mobilization of natural history collections data (gsap-nhc) task group was charged with examining priorities for the digitization of natural history collection data (gbif, 2008a). the drivers for specimen level digitization are many and varied, and often intrinsically linked (vollmer, et al., 2010), therefore it is not realistic for the task group to dictate priorities for specimen level digitization. the proportion of a collection which is digitized varies greatly from one collection to another with some collections fully digitized and others having made little progress. the opportunities and priorities for digitization vary from one institution and one country to another. there is no single answer to setting priorities for specimen level digitization and therefore, we argue that the use of metadata to describe a collection be a key component to achieve demand-driven data digitization and publishing. we are further of the opinion that: 1. metadata must be used to expose data to users and to expand the user base, 2. metadata and specimen level digitization should be considered as a part of the digitization process and be prioritized together, 3. metadata creation can be considered on a scale from local to global. we therefore recommend that the creation of metadata records for a collection is essential. this, however, does not replace the need for digitization itself but should be considered as a part of any digitization project. metadata creation will lead to digitization by increasing the exposure of the data. demand-driven prioritization of collection digitization the majority of digitization activities and initiatives are opportunistic in nature (vollmer, et al., 2010) but with such an opportunistic approach, digitization of the world’s natural history collections will not be achieved in the foreseeable future. further, as stated by scoble & bourgoin (2010), and berendsohn & seltmann (2010), with current resource allocation, and socio-political and scientific priorities, it may not be possible to achieve digitization of all the berents et al. – towards demand-driven publishing specimens housed in the world’s natural history collections, which comprise more than 3 billion specimens. the management of collections is greatly enhanced by digitization but despite this strong internal driver, institutions have been unable to generate adequate resources to achieve this goal. it is therefore necessary to mobilize external resources and expose the collections and data to a wider user group who will drive and resource digitization priorities (national science and technology council, committee on science, interagency working group on scientific collections, 2009). ‘demand-driven digitization’ will address the immediate needs of stakeholder communities and increase the potential for attracting financial and human resources and improved infrastructure. we therefore recommend that natural history collections adopt the approach of ‘demand-driven’ digitization of collections to address the requirements of stakeholder communities. we strongly advocate that the collection management and curatorial community develop an institutional or collection specific demand-driven prioritization plan for collection digitization on the basis of content needs of the major stakeholders and the user community. the integration of such institutional or collection specific plans would then help to develop a national or thematic collection digitization strategy and action plan. we further recommend that gbif works with major natural history collection stakeholder communities to develop best practice guidelines for developing (a) institution or collection specific, demand-driven plans for collection digitization, and (b) national or thematic collection digitization strategy and action plans. chris frazier, et al. (2008) has written a very useful guide on initiating a collection digitization project, as part of the gbif training manual on digitization of natural history collection data (gbif, 2008b). in our opinion, this will further help federal science funding agencies and private donors to cooperate in developing comprehensive funding strategies which will result in key scientific, ecological, and social issues being addressed. metrics for prioritizing the digitization we reviewed a variety of factors which influence decisions regarding digitization of natural history collections. these factors include (1) type specimens, (2) collections associated with projects, (3) historical significance, (4) taxonomic priorities, (5) ecosystem relevance, and (6) species of special concern. we suggest that the following criteria should be considered when setting priorities for digitization: 1. type specimens: type specimens are important not only for taxonomists, but also as a reference for accurate identification and naming of biological specimens, which is fundamental to all other biological and biology-related fields (agriculture, medicine, conservation, environmental management, etc.). type specimen records should allow links to the catalogue of life, the encyclopedia of life and other databases of spatial and temporal information from collections to provide ready access to information such as the type locality and the date of collection. type data must indicate the status of the type specimen, such as whether it refers to the type specimen of a currently valid species name. holotype (primary type) specimens should be a top priority, but other types such as paratypes should also be included in the databases. ideally, each type specimen (at least the holotype) should also be photographed or scanned so that access to information associated with type material can be made globally accessible in perpetuity. a global account of type specimens and the collections in which these specimens are housed will also allow some assessment of the value of different collections. the prioritization of type specimens in collections has three main benefits: a. profiles the jewels in the museum and herbarium collections and thus enhances their profile (prestige effect), b. creates an achievable target and gets institutions comfortable with what is involved in digitization and thus they may continue with these activities, and c. provides another register of all known species that can be compared to catalogue of life and other names received by gbif. gbif can then link species to their type specimen 114 berents et al. – towards demand-driven publishing locations, thus directing researchers to these essential resources. 2. digitization of collections associated with projects: e.g. museum exhibitions, biodiversity hotspots, biodiversity surveys, conservation questions such as important bird areas, or biological or other research projects. we strongly recommend that digitization of specimens associated with projects should be an integral part of a project and thus be achieved prior to the close of projects. in the case of completed projects an independent assessment needs to be carried out to decide the order of preference for digitization of specimens associated with such projects. 3. historical significance: the unique value of many natural history collections lies in the historical nature of the data. what constitutes an “historical” dataset may be debatable and be regionally variable depending on the time since the greatest change. for example, in developing countries, pre-1980 may be considered historical, while in industrialized countries pre-1900 may be of more value. specimens to be considered for these datasets should be identified to species level, and have locality data provided as accurately as possible, with some indication of uncertainty such as an uncertainty radius, and include at least the year of collection. 4. taxonomic priorities: digitization may be focused on taxonomic priorities as a result of taxonomists bringing resources for digitization driven by their need for specimens and data. taxonomists contribute the greatest amount of material and data to collections as specimens and specimen data are critical for taxonomists’ work. it is the taxonomic cadre that provides the value of the specimens and collections to society. ideally taxonomic priorities should be linked to global or national initiatives, involving several taxonomists and / or institutions working on a single taxon. for example, in south africa, the south african butterfly conservation assessment has driven the need for specimen data capture from all public institutions and private collectors and has resulted in more than 400,000 butterfly records being captured. 5. ecosystem relevance: this could involve prioritizing the collections which have resulted from biodiversity surveys or research projects on ecosystems or habitats that provide critical services (e.g. freshwater ecosystems, wetlands, coral reefs, forests, rangelands etc). 6. species of special concern: a. invasive alien species: invasive species are considered to be one of the greatest threats to biodiversity, and specifically to ecosystem functioning and resilience to change. information on the diversity and distribution of alien invasive species has global as well as regional and local relevance, and global data sets will be of value to a large number of biologists, environmental managers and conservationists. having large, global datasets of alien invasive species will also highlight the value of digitization of specimens or observations and the value of gbif activities to a wide range of stakeholders. b. species of direct relevance to people: i. harvested species (e.g. crops, medicinal plants, line fish) ii. pests iii. diseases or disease vectors (of humans, livestock or crops). the rationale for the selection of the species needs to be explicit and data must include species level identification, accurate locality data as well as the date (minimum of the year) of collection or observation. c. threatened, endangered, endemic species: data on these species are critical for natural resource management, conservation planning, and decisions about land use. the factors discussed in this section are considered useful for determining priorities which will generate demand-driven digitization. data capture initiatives based on these priorities will provide data of direct use to a wide range of stakeholders, perhaps expanding the traditional users of collection data, and thus increasing the value of gbif’s initiatives (and therefore the possibilities for funding). understanding current and past distributions of species of direct relevance to human survival is critical not only for human well-being, but also to illustrate the value of collections and taxonomy and of the biodiversity sciences in general. 115 berents et al. – towards demand-driven publishing 116 having recognized that metadata is essential for the demand-driven digitization of natural history collection data, in the remainder of this paper we discuss various aspects dealing with metadata authoring and publishing. metadata, a priority: why? there are several advantages to authoring and publishing enriched metadata documents, including: 1. increased discovery and visibility 2. stimulation of demand-driven digitization 3. increased usage and user base 4. comprehensive tracking of the progress of national to global scale digitization 5. early detection of collection risk assessment – identify collections at risk 6. improved capacity management – technical, infrastructure, human resources and finance 7. improved estimation of the scale of biological collections metadata: challenges or constraints? • what is metadata? discussions with curators and collection managers reveal that there is very poor understanding of what metadata is and its significance for improved discovery and visibility of the collections. this calls for increased awareness and outreach amongst the natural history collections community about the importance of metadata and how it can contribute towards the sustainability and increased use of collections. • metadata scale. one of the very critical decisions that influence the usefulness of a metadata document is the scope of the collection which the document describes. collections are arranged and organized on multiple bases, such as taxa, projects, collector, ecosystem etc. with such complexity, decisions about whether to describe collections on the basis of taxa, size, or any other criterion will determine the usefulness of the metadata document. further, inclusion of both digital and non-digital specimens in the same document adds to the challenge. metadata: criteria for determining scope of metadata documents authoring a metadata document that describes a collection adequately, resulting in sufficient exposure, visibility, renewed interest and increased support for digitization is a challenge. therefore, determining the scope of the metadata document is essential. some of the frequently asked questions are, (a) whether a single metadata document could be good enough to describe the collection, (b) how lengthy or detailed a metadata document should be, and (c) to what level of granularity / depth it should collate the details. we believe that answers to these questions largely depend on the answers to the following questions (1) how big is the collection? (2) what human resource capacity is available to do the metadata authoring? (3) how you would like to project the collection and (4) who is the target audience? the following criteria should be used to determine the scope of the metadata document, and for deciding whether single or multiple metadata documents will best describe the collection (table 1). table 1. criteria for determining scope of the metadata documents. aspects scale issues to consider taxon family or lower taxa size of the collection age of collection level of curation / digitization complexity and diversity of the level of metadata record geographic scope country or ocean / seas province / state biogeographic regions political boundaries v/s bioregions ease of management, and organization of collection projects individual project complexity in collection management (which is berents et al. – towards demand-driven publishing expeditions or cruise often taxon based, and projects / expeditions often cut across taxa) a collection from an expedition or cruise may be deposited across multiple institutions / countries temporal scale collector one record per collector or group of collectors homogeneity vs heterogeneity complexity in collection management (which is often taxon based, and projects / expeditions which encompass many taxa) a collection from an expedition or cruise may be deposited across multiple institutions / countries temporal scale size of the collection multiple metadata documents will describe extensive collections better <1000 specimens – single metadata document >1000 specimens – multiple metadata documents as a general principle large collections will require multiple metadata documents. furthermore, large collections may consider a taxon or region specific approach for collating metadata documents, whereas in small collections the author may employ a project or collector specific approach. however, the decision for adopting a specific criterion or combination of criteria is influenced by multiple factors. what metadata elements are essential? on the basis of our assessment of the central question as to what will describe the collections best, we suggest that details about the following elements must constitute the core component of the metadata document: 1. list of taxa – preferably to the level of family, but in the case of insects or invertebrates this could be to a higher taxon level (e.g. class or order) (low granularity) and where possible, also to a lower taxon level (family, subfamily, tribe) (higher granularity). 2. list of regions – preferably include biogeographical regions as it would enhance the use of metadata. 3. temporal scale – granularity depends on the size of the collection and temporal range of collection events (e.g. from 1990-2000). 4. an estimate of the size of the collection i.e. specify by order of magnitude of 100s, 1000s, or 10000s) (e.g. approx 1000-2000 specimens). 5. state of accession or curation e.g. state if the collection is sorted and pinned or not sorted yet, and whether the collection is accessioned into a catalogue book. 6. state of digitization – metadata, extent of digitization (e.g. %), detail of data captured (e.g., taxonomic details only, or locality data, collection data, imaging of each specimen or % of specimens). 7. type status – how many type specimens v/s non-type specimens. 8. persistent identifier (i.e. a unique number or code that unanimously identifies the record) for collection, curator and metadata record itself. interlinking between these persistent identifiers is crucial for easy and efficient discovery. 9. special significance – e.g. historical or social (productivity and public health), economic or environmental significance of collection. 10. collection risk assessment: level and description of the potential risk to the collection and reasons for such a risk. 117 berents et al. – towards demand-driven publishing metadata: implementation approach we recommend the following step-wise approach for constructing a metadata document: 1. work at curation or collection manager level 2. adopt a hierarchical approach: a. taxa (higher to lower taxonomic levels) b. bioregions (larger to smaller regions) 3. if tackling digitization from a low base (where few details are available) a few metadata records should be created to describe the collection on a large scale to achieve data exposure. the next priority would be digitization of the top priority elements of the collection. as digitization proceeds finer level metadata records should be created : e.g. (fictional example) metadata record 1: for the entire australian museum collection – we have a faunal collection > 16 million specimens from australia and the indo-pacific from 1800’s to present in various stages of curation and digitization. metadata record 2: mollusc collection (~50,000 specimens) from indo-pacific from 1850-2000. metadata record 3: create a metadata record for each of the 10 major families in the mollusc collection. metadata: exemplar use cases john studies the impact of climate change on amphibians in madagascar. he searches on the gbif data portal and other amphibian specific portals, which results in 2000 data records. a search on gbif data portal leads to 20 metadata records with 20000 specimens of which 18000 specimens are not digital. a scan through 20 metadata records reveals that there are an additional 4000 madagascan specimens in eight museums not digitized. john approaches the curators, and in the following month he has an additional 2500 records for analysis. conclusions and future work a demand-driven approach is considered to be the most successful approach to the daunting task of digitizing the data of the world’s natural history collections. however, we currently lack best practice guidelines on how to develop demand-driven strategies and action plans for digitization of natural history collections data. in the near future gbif, together with professional societies, needs to develop such guidelines. the use of metadata will expose the data to stakeholders and increase the resources available for data capture. a hierarchical and prioritized approach is recommended for the creation of metadata records. institutions and gbif should develop guidelines and plans for the digitization of collections including the use of metadata. competing interests the authors declare that they have no competing interests. acknowledgements penny berents is grateful to the australian museum for support to work on this project. michelle hamer is grateful to the south african national biodiversity institute. vishwas chavan is grateful to the global biodiversity information facility. references berendsohn, w. g., and p. seltmann. 2010. using geographical and taxonomic metadata to set priorities in specimen digitization. biodiversity informatics 7: 120-129. chapman, a.d. 2005. uses of primary species-occurrence data, version 1.0. copenhagen: global biodiversity information facility. 106 pp. isbn: 87-92020-01-1. accessible at http://www2.gbif.org/uses.pdf. (accessed september 20, 2010). frazier, c. k., wall, j., and grant, s. 2008. initiating a natural history collections digitization project, version 1.0. copenhagen: global biodiversity information facility. 75 pp. isbn: 87-92020-05-4 (pdf: http://www.gbif.org/) accessible at http://www2.gbif.org/digitization.pdf. (accessed september 20, 2010). gbif 2008a. terms of reference for “task group on a global strategy and action plan for the mobilisation of natural history data”1. accessible at http://tinyurl.com/gsaptg. (accessed september 20, 2010). gbif, 2008b. gbif training manual 1: digitization of natural history collections data, version 1.0. copenhagen: global biodiversity information facility. 1 http://tinyurl.com/gsaptg 118 http://www2.gbif.org/uses.pdf http://tinyurl.com/gsaptg http://tinyurl.com/gsaptg berents et al. – towards demand-driven publishing 119 isbn 87-92020-07-0. accessible at http://www.gbif.org. (accessed september 20, 2010). national science and technology council, committee on science, interagency working group on scientific collections, 2009. scientific collections: missioncritical infrastructure of federal science agencies. office of science and technology policy, washington, dc. accessible at http://www.whitehouse.gov/sites/default/files/scicollections-report-2009-rev2.pdf. (accessed september 20, 2010). pyke, g.h. and ehrlich, p.r. 2010. biological collections and ecological/environmental research: a review, some observations and a look to the future. biogical reviews., 85: 247 – 266. scoble, m. j., and t. bourgoin. 2010. natural history collections digitization: rationale and value. biodiversity informatics 7: 77-80. tann, j., kelly, l., and flemons, p. 2008. atlas of living australia – user needs analysis. (user needs analysis report | atlas of living australia). published electronically at: http://www.ala.org.au/documents/user-needs-analysisreport.html. (accessed september 20, 2010). vollmar, a., macklin, j. a., and ford, l.s. 2010. natural history specimen digitization: challenges and concerns, j. biodiversity informatics 7: 93-112. http://www.gbif.org/ http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf.%20(accessed%20september%2020 http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf.%20(accessed%20september%2020 http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf.%20(accessed%20september%2020 http://www.ala.org.au/documents/user-needs-analysis-report.html http://www.ala.org.au/documents/user-needs-analysis-report.html demand-driven prioritization of collection digitization metrics for prioritizing the digitization metadata, a priority: why? metadata: challenges or constraints? metadata: criteria for determining scope of metadata documents what metadata elements are essential? metadata: implementation approach metadata: exemplar use cases conclusions and future work competing interests acknowledgements references microsoft word jimenez-valverde_et_al_corrected.doc biodiversity informatics, 6, 2009, pp. 28-35 28 short communication environmental correlation structure and ecological niche model projections alberto jiménez-valverde, yoshinori nakazawa, andrés lira-noriega and a. townsend peterson biodiversity research center, the university of kansas, lawrence, kansas 66045 the environmental causation of species’ distributions depends on three general, interacting types of factors: the abiotic (or physical) environment, the biotic environment, and accessibility of areas across complex landscapes (pulliam 2000; soberón and peterson 2005; soberón 2007). indirect variables, such as elevation, are those associated with the presence of species owing to correlation with the actual variables that directly and causally affect the fitness of the species, such as temperature or precipitation (austin 2002). put another way, variables can be arranged along a gradient of proximal to distal, regarding the immediacy of causality regarding the fitness of the species: indirect variables are always distal variables (austin 2002). contrary to proximal variables, distal variables are often easy measurable, and thus available in georeferenced databases (fig. 1). many researchers now attempt to reconstruct these environmental dimensions as ecological niche models (also termed “bioclimatic envelopes,” “environmental niche models,” or even “species distribution models”), using a variety of inferential approaches. niche models have been used to predict geographic distributions of species (guisan et al. 2006), anticipate distributions of unknown species (raxworthy et al. 2003), estimate the invasive potential of species (peterson 2003; thuiller et al. 2005), and forecast climate change effects on species’ distributions (araújo et al. 2005). the predictive capacity of these approaches makes them particularly useful in applications involving “transferring” the niche model to make predictions regarding other landscapes or time periods (araújo and pearson 2005; peterson et al. 2007). such transferability applications, however, depend critically on the assumption that environmental variables relevant on one landscape or at one time will be relevant on another. niche models are probably never based directly on genuinely proximate variables, but rather rely on more easily measurable variables that are inevitably less directly related to the population biology of the species in question. as such, the correlation structure among environmental variables becomes key (morin and lechowicz 2008): if correlation structures are stable and consistent across different landscapes and time periods, then niche models may be transferable to those other situations; if, on the other hand, correlation structures are not consistent among situations, then models may not be transferable, or at least not as fully or as readily. as correlation methods, niche modeling techniques simply select the set of variables that is best to explain the largest part of the variation in the dependent variable. transferability exercises require the assumption that the variables selected (in isolation or in interaction terms, depending on the complexity of the technique) are those that have strongest influence on the real (unknown) causal variables. however, because of intercorrelations among variables, this relationship may not hold true, and other distal variables may be selected just because they are closely correlated with the key variables. in such situation, when transferring model predictions, if the correlation structure among distal variables is maintained, then model predictions will be robust; if not, then the model may not work properly. in this note, we present comparisons of correlation structures of suites of climatic, topographic, and surface-reflectance variables among continents and time periods. in general, jimenez-valverde et al. environmental correlation structure 29 table 1. summary of mantel tests used to evaluate similarity of correlation structure among environmental data sets for different continents and different time periods. the pearson product-moment correlation coefficient (r) was compared with similar calculations from 1000 randomized rearrangements of the original matrices to generate probability values (p). comparison r p worldclim climate data (19 bioclimatic variables) africa vs eurasia 0.712 <0.001 africa vs north america 0.673 <0.001 africa vs south america 0.734 <0.001 africa vs australia 0.933 <0.001 eurasia vs north america 0.909 <0.001 eurasia vs south america 0.895 <0.001 eurasia vs australia 0.758 <0.001 north america vs south america 0.825 <0.001 north america vs australia 0.734 <0.001 south america vs australia 0.767 <0.001 worldclim climate data (7 bioclimatic variables) africa vs eurasia 0.800 <0.001 africa vs north america 0.743 <0.001 africa vs south america 0.714 <0.001 africa vs australia 0.970 <0.001 eurasia vs north america 0.887 <0.005 eurasia vs south america 0.836 <0.001 eurasia vs australia 0.843 <0.001 north america vs south america 0.805 <0.005 north america vs australia 0.805 <0.005 south america vs australia 0.772 <0.001 ipcc mean monthly climate data (10 variables) africa vs eurasia 0.587 <0.001 africa vs north america 0.511 <0.001 africa vs south america 0.884 <0.001 africa vs australia 0.957 <0.001 eurasia vs north america 0.948 <0.001 eurasia vs south america 0.527 <0.001 eurasia vs australia 0.648 <0.001 north america vs south america 0.465 <0.001 north america vs australia 0.585 <0.001 south america vs australia 0.806 <0.001 normalized difference vegetation index (10 monthly composites) africa vs eurasia 0.431 <0.05 africa vs north america 0.541 <0.05 africa vs south america 0.932 <0.001 africa vs australia 0.968 <0.001 eurasia vs north america 0.948 <0.001 eurasia vs south america 0.530 <0.05 eurasia vs australia 0.574 <0.05 north america vs south america 0.613 <0.005 north america vs australia 0.665 <0.001 south america vs australia 0.939 <0.001 jimenez-valverde et al. environmental correlation structure 30 hydro-1k topographic variables africa vs eurasia 0.985 0.013 africa vs north america 0.987 0.013 africa vs south america 0.994 0.02 eurasia vs north america 0.983 0.008 eurasia vs south america 0.994 0.008 north america vs south america 0.987 0.007 pleistocene vs present africa 0.908 <0.001 australia 0.994 <0.001 eurasia 0.986 <0.001 north america 0.988 <0.001 south america 0.987 <0.001 pleistocene vs future africa 0.902 <0.001 australia 0.973 <0.001 eurasia 0.992 <0.001 north america 0.973 <0.001 south america 0.986 <0.001 present vs future africa 0.992 <0.001 australia 0.988 <0.001 eurasia 0.988 <0.005 north america 0.983 <0.001 south america 0.998 <0.001 correlation structures are conserved, which indicates that models based on distal variables can be transferred among regions and time periods. however, the conservative nature of the correlation structure is not absolute, indicating some degree of caution in interpretation, particularly when transferring model predictions between northern and southern hemispheres. methods we assembled sets of environmental data of global extent describing aspects of climate, topography, and surface reflectance. specifically, we used two climate data archives—worldclim (hijmans et al. 2005) and new et al. (1997), both widely used by the niche modeling community. from the former, we used the 19 “bioclimatic” variables (and in some analyses a subset of 7 of these variables: annual mean temperature, mean diurnal range, maximum temperature of warmest month, minimum temperature of coldest month, annual precipitation, precipitation of wettest month, and precipitation of driest month). pleistocene (last glacial maximum, 21,000 years bp) climate data were derived from the community climate system model (ccsm; collins et al. 2004), while future climate data (for 2100) were derived from the ccm3 climate model (govindasamy et al. 2003). these data sets were obtained, together with present-day climate data, from the worldclim website1 at a resolution of 2.5’. from the new et al. (1999) data set, we used 10 mean climate surfaces derived from the period 1961-1990 at a resolution of 0.5°, comprising precipitation; wet-day frequency; mean, maximum and minimum temperature; diurnal temperature range; vapor pressure; sunshine percent; cloud cover; and wind speed. for comparison, we also analyzed a global 8 km resolution dataset composed of 10 months’ (january, february, april, june, july, august, september, october, november, and december 1998) normalized difference vegetation index (ndvi) maximum value composites from the 1 http://www.worldclim.org/. jimenez-valverde et al. environmental correlation structure 31 figure 1. schematic diagram of immediately causal proximate variables (p) in affecting whether or not a site is suitable for a species, as well as the easily measurable distal variables (d, direct; i, indirect) that are correlated or associated to varying degrees of directness with the proximate variables. arrows indicate causal links—note the indirect nature of some of the causal links between easy-to-measure variables and proximate variables in this hypothetical case. advanced very high resolution radiometer (avhrr) sensor, and a global 1 km digital elevation model with layers describing elevation, slope, compound topographic index, flow direction, and flow accumulation (usgs 2001; note that this data set is incomplete for australia, so we omitted that continent from our analyses). we overlaid 10,000 random points on the extent of each of the 5 continents, and extracted grid values for each environmental data set at each point. for the ndvi data set, we developed analyses for the raw monthly data sets, and for a version in which the southern and northern hemispheres were offset by 6 months to reflect differences in seasonal timing. then, across each continent, for each data set, we calculated all pairwise correlation coefficients among environmental variables to produce a square correlation matrix. finally, we compared these sets of matrices using mantel tests in the vegan2 package of r, with 1000 permutations. to summarize patterns, we clustered continent matrices using the ward´s method as a linkage rule, based on similarity as measured by cell-by 2 http://r-forge.r-project.org/projects/vegan/. cell pearson product-moment correlation coefficients among continent matrices. results and discussion all pairwise matrix comparisons between variables from the same set (climate, topography, and surface reflectance) indicated a correlation structure statistically significantly more similar than null expectations (p < 0.05; table 1). correlation coefficients ranged 0.432-0.998, suggesting fairly-to-highly similar matrix structures; these numbers were generally higher for intertemporal comparisons, and lower for intercontinental comparisons. this result may be expected as past and future climate predictions are both derived from present-day models; for this reason, there is no guarantee that these measured correlations reflect the truth, instead of being affected to some degree by artifact. topographic variables also showed quite-high correlation values (>0.940), as would be expected considering that they are derived from the same single variable, elevation. distal variables (direct and indirect, easy-to-measure) proximal variables (direct) d1 d2 d3 i1 p1 p2 p3 i2 presence or absence of the species in question distal variables (direct and indirect, easy-to-measure) proximal variables (direct) d1 d2 d3 i1 p1 p2 p3 i2 presence or absence of the species in question jimenez-valverde et al. environmental correlation structure 32 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 0.16 0.18 linkage distance (1-r) south america north america eurasia australia africa 0.0 0.1 0.2 0.3 0.4 0.5 0.6 linkage distance (1-r) south america north america eurasia australia africa a 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 linkage distance (1-r) north america eurasia south america australia africa c 0.00 0.05 0.10 0.15 0.20 0.25 linkage distance (1-r) north america eurasia south america australia africa d b 0.004 0.006 0.008 0.010 0.012 0.014 0.016 0.018 linkage distance (1-r) north america eurasia south america africa e 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 0.16 0.18 linkage distance (1-r) south america north america eurasia australia africa 0.0 0.1 0.2 0.3 0.4 0.5 0.6 linkage distance (1-r) south america north america eurasia australia africa a 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 linkage distance (1-r) north america eurasia south america australia africa c 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 linkage distance (1-r) north america eurasia south america australia africa c 0.00 0.05 0.10 0.15 0.20 0.25 linkage distance (1-r) north america eurasia south america australia africa d 0.00 0.05 0.10 0.15 0.20 0.25 linkage distance (1-r) north america eurasia south america australia africa d b 0.004 0.006 0.008 0.010 0.012 0.014 0.016 0.018 linkage distance (1-r) north america eurasia south america africa e 0.004 0.006 0.008 0.010 0.012 0.014 0.016 0.018 linkage distance (1-r) north america eurasia south america africa e figure 2. summary of patterns of similarity of correlation structure among continents for different environmental data sets, using pearson product-moment correlation coefficients as similarity measures among matrices. (a) worldclim data set, all 19 present-day “bioclimatic” variables (hijmans et al. 2005); (b) worldclim data set, 7 present-day “bioclimatic” variables (see methods); (c) new et al. (1997), 10 mean montly climatic variables (see methods); (d) normalized difference vegetation index (ndvi) derived from the advanced very high resolution radiometer (avhrr) satellite, monthly measurements in 1998 (note northern and southern hemispheres offset by 6 months; results without offset were similar); (e) hydro-1k digital elevation model, 6 variables. jimenez-valverde et al. environmental correlation structure 33 clustering continents by the matrix correlation coefficients, for the ndvi data, topographic features and the new et al. (1999) data, southern hemisphere continents and northern hemisphere continents formed the two major branches of the dendrogram (fig. 2); analyses based on the worldclim data maintained the general northsouth division, but placed south america separately from the other southern continents. to analyze whether spatial resolution of data sets affected the outcome, we also performed the analyses with the worldclim dataset at 10’, 5’, and 2.5’ grid cell sizes: correlation structure is qualitatively identical (see appendix), except for some small decreases in correlations at highest resolutions. analyses with and without the 6month seasonal offset of the ndvi data both yielded the northern-versus-southern hemisphere dichotomy. transferability is a prerequisite for generalization of niche models, because they can otherwise be applied only locally and to a precise temporal “snapshot” (fielding and haworth 1995). local adaptations, biotic interactions, sink populations, and historical constraints all can reduce transferability of models (randin et al. 2006; staruss and biedermann 2007; vanreusel et al. 2007). however, besides these factors, when working with correlative models and indirect variables, a more basic consideration is needed: the maintenance of correlation structure of the set of factors. this point is even more important, given the current tendency to recommend use of complex modeling techniques (e.g., elith et al. 2006), which can potentially overfit the model to input data and reduce transferability (randin et al. 2006; peterson 2007). the results of this study are simultaneously encouraging and discouraging for broad-scale niche model projections across space and time. in general, the correlation structure of environmental data sets is conservative, and in that sense projections of model rule sets among continents should generally be robust. however, the relatively lower similarity of correlation structure among continents in the northern versus southern hemispheres could potentially produce less accurate or less complete projections among hemispheres. although our analyses were developed at an intercontinental scale, the problem of maintenance of the correlation structure among variables affects any transferability exercise at any spatial extent. thus, we recommend assessing the degree of maintenance of the correlation structure in any transferability study to assess this potential source of uncertainty. the biggest unknown surrounding these results is whether and to what degree observed similarities and differences in correlation structure will affect predictions of potential distributional areas of species. that is, all of these matrices for individual continents were more similar in structure than random expectations, but none had the exact same correlation structure: does this result mean that model transfers among continents will be less efficient than those within continents? similarly, to what degree could the inter-hemispheric reduction of similarity of correlation structure affect the predictive ability of the models when projected among hemispheres? these effects on model transferability, however, will depend on the correlation structure of the actual niches and the models we create thereof for the species, but this structure and its estimation are complex, and will require additional exploration. acknowledgements we are grateful to jorge soberón for his valuable suggestions. aj-v is supported by a mec (ministerio de educación y ciencia, spain) postdoctoral fellowship (ref.: ex-2007-0381). aln received financial support from the consejo nacional de ciencia y tecnología of mexico (189216). atp was supported by a grant from microsoft research. literature cited austin, m. p. 2002 spatial prediction of species distribution: an interface between ecological theory and statistical modeling. ecological modelling 157:101-118. araújo, m. b., and r. g. pearson. 2005. equilibrium of species' distributions with climate. ecography 28:693-695. araújo, m. b., r. g. pearson, w. thuiller, and m. erhard. 2005. validation of species-climate impact jimenez-valverde et al. environmental correlation structure 34 models under climate change. global change biolology 11:1504-1513. collins, w. d., c. m. bitz, m. l. blackmon, g. b. bonan, c. s. bretherton, j. a. carton, p. chang, s. c. doney, j. j. hack, t. b. henderson, j. t. kiehl, w. g. large, d. s. mckenna, b. d. santer, and r. d. smith. 2004. the community climate system model: ccsm3. journal of climate 19:2122-2143. elith, j., c. h. graham, r. p. anderson, m. dudik, s. ferrier, a. guisan, r. j. hijmans, f. huettman, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. m. overton, a. t. peterson, s. j. phillips, k. richardson, r. scachetti-pereira, r. e. schapire, j. soberón, s. e. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species' distributions from occurrence data. ecography 29:129-151. fielding, a. h., and p. f. haworth. 1995. testing the generality of bird-habitat models. conservation biology 9:1466-1481. govindasamy, b., p. b. duffy, and j. coquard. 2003. high-resolution simulations of global climate, part 2: effects of increased greenhouse cases. climate dynamics 21:391-404. guisan, a., o. broennimann, r. engler, m. vust, n. g. yoccoz, a. lehmann, and n. e. zimmermann. 2006. using niche-based models to improve the sampling of rare species. conservation biology 20:501-511. hijmans, r. j., s. e. cameron, j. l. parra, p. g. jones, and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25:1965-1978. morin, x., and m. j. lechowicz. 2008. contemporary perspectives on the niche that can improve models of species range shifts under climate change. biology letters 4:573-576. new, m., m. hulme, and p. jones. 1997. a 1961-1990 mean monthly climatology of global land areas. climatic research unit, university of east anglia, norwich, u.k. new, m., m. hulme, and p. d. jones. 1999. representing twentieth century space-time climate variability. part 1: development of a 1961-90 mean monthly terrestrial climatology. journal of climate 12:829-856. peterson, a. t. 2003. predicting the geography of species' invasions via ecological niche modeling. quarterly review of biology 78:419-433. peterson, a. t. 2007. why not whywhere: the need for more complex models of simpler environmental spaces. ecological modelling 203:527-530. peterson, a. t., m. papeş, and m. eaton. 2007. transferability and model evaluation in ecological niche modeling: a comparison of garp and maxent. ecography 30:550-560. pulliam, h. r. 2000. on the relationship between niche and distribution. ecology letters 3:349-361. randin, c. f., t. dirnbock, s. dullinger, n. e. zimmermann, m. zappa, and a. guisan. 2006. are niche-based species distribution models transferable in space? journal of biogeography 33:1689-1703. raxworthy, c. j., e. martínez-meyer, n. horning, r. a. nussbaum, g. e. schneider, m. a. ortegahuerta, and a. t. peterson. 2003. predicting distributions of known and unknown reptile species in madagascar. nature 426:837-841. soberón, j. 2007. grinnellian and eltonian niches and geographic distributions of species. ecology letters 10:1115-1123. soberón, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species' distributional areas. biodiversity informatics 2:1-10. staruss, b., and r. biedermann. 2007. evaluating temporal and spatial generality: how valid are species-habitat relationship models? ecological modelling 204:104-114. thuiller, w., d. m. richardson, p. pysek, g. f. midgley, g. o. hughes, and m. rouget. 2005. niche-based modelling as a tool for predicting the risk of alien plant invasions at a global scale. glob change biol 11:2234-2250. usgs. 2001. hydro1k elevation derivative database http://edcdaac.usgs.gov/gtopo30/hydro/. u.s. geological survey, washington, d.c. vanreusel, w., d. maes, and h. van dyck. 2007. transferability of species distribution models: a functional habitat approach for two regionally threatened butterflies. conservation biology 21:201-212. jimenez-valverde et al. environmental correlation structure 35 appendix i correlation structure of bioclimatic variables among continents at different spatial resolutions. tabular summary of mantel tests used to evaluate similarity of correlation structure among environmental data of different spatial resolutions among continents. p-values were below 0.001 in all cases. r worldclim climate data (19 bioclimatic variables) 10’ 5’ 2.5’ australia vs africa 0.932 0.9327 0.9324 australia vs eurasia 0.7574 0.7585 0.6672 australia vs north america 0.7332 0.7343 0.7348 australia vs south america 0.767 0.7681 0.7672 africa vs eurasia 0.7113 0.7127 0.677 africa vs north america 0.6726 0.6738 0.6743 africa vs south america 0.7332 0.7345 0.7345 eurasia vs north america 0.9086 0.9089 0.7007 eurasia vs south america 0.8941 0.8947 0.7705 north america vs south america 0.8253 0.8253 0.8239 microsoft word chavan_bi_2005_20.doc biodiversity informatics, 2, 2005, pp. 70-78 70 resolving taxonomic discrepancies: role of electronic catalogues of known organisms vishwas chavan, nilesh rane, aparna watve information division, national chemical laboratory, pune 411008, india. and michael ruggiero integrated taxonomic information system, u.s. geological survey, smithsonian institution, washington d.c., usa abstract. —there is a disparity in availability of nomenclature change literature to the taxonomists of the developing world and availability of taxonomic papers published by developing world scientists to their counterparts in developed part of the globe. this has resulted in several discrepancies in the naming of organisms. development of electronic catalogues of names of known organisms would help in pointing out these issues. we have attempted to highlight a few such discrepancies found while developing indfauna, an electronic catalogue of known indian fauna, and comparing it with existing global and regional databases. key words.—biodiversity informatics, electronic catalogue, taxonomic discrepancies, indian fauna identification of organisms is fundamental to biodiversity studies. owing to this, the discipline of taxonomy, especially scientific nomenclature, has gained immense importance. taxonomy provides a vocabulary to discuss the world (knapp et al., 2002). each name is unique and its representative organism is precisely described. it is estimated that about 1.8 million species of organisms have been formally named from the world (may, 1999) and each is recognized by a unique binomial. more than 2000 new generic names and 15000 new specific names alone are added to the zoological literature every year, and with such a multiplicity of names, problems are bound to occur. international mechanisms such as the international code of zoological nomenclature (iczn, 1999) and the international code of botanical nomenclature (icbn) are rulebooks that govern how organisms are named and they provide clear instructions on how to go about the process (knapp et al., 2002). international codes of nomenclature require taxonomic actions to be published and the data thus made available (agosti and johnson, 2002). however, nomenclatural additions or changes have to be conveyed to the iczn or icbn by the authors and is usually done when ratification is needed from the international authority. similarly, discrepancies in the nomenclature are brought to the notice of the iczn and icbn by scientists, which are later reviewed. this process requires a long time and the availability of a large amount of literature to the scientists discussing nomenclature. in several cases, especially for taxonomists in developing countries, recent taxonomic literature including the codes themselves are unavailable. very few libraries around the world have the financial capacity to carry the full range of literature in which systematic results are published (agosti and johnson, 2002). hence, nomenclature changes are in many cases unavailable or become available much later to the developing country scientists than to their counterparts in developed world. this leads to use of old or outdated nomenclature. on the other hand taxonomic papers by developing country scientists published in journals with regional scope, which are not scientifically abstracted, remain isolated and unnoticed by the wider scientific audience and taxonomic changes proposed or used in such papers are often neglected. this obviously leads to many discrepancies in the information available, especially about the current or correct taxonomic chavan et al. – resolving taxonomic discrepancies 71 hierarchy of organisms. it is thus necessary to create a system, which will lead to rapid identification of taxonomic discrepancies and their resolution. in addition, a permanent mechanism for registering and validating scientific names of organisms needs to be created at national as well as global levels. web-based electronic catalogues can be effective in creating such a central repository of taxonomic information. in this paper, we demonstrate the use of web-based electronic catalogues (elecats) in identifying taxonomic discrepancies in an indian context. the ncl center for biodiversity informatics (ncbi) is developing an electronic catalogue of known indian fauna (indfauna) (ncl, 2005). so far it has documented baseline information on more than 91,000 scientific names of the known indian faunal species. the data incorporated in the database is collected from multiple sources. the main focus is on published literature including research papers, faunas, and monographs as sources of authentic and reviewed information. for those taxa especially invertebrates, on which published literature is not readily available, preserved collections from natural history museums, web-based databases and checklists are also being referenced. thus, when collecting information, highest importance is given to “faunas” and “monographs” followed by “published research papers”, then “online and offline databases” followed by “regionand taxonspecific web sites” followed by “personal communications with experts”, and finally to “nontaxonomic publications”. the information is carefully scrutinized for validity and accepted only if it is from reputed taxonomic institutions or experts. for each species, the taxonomic hierarchy used by indian faunas is crosschecked diligently with that used in global taxonomic inventories such as integrated taxonomic information system (itis, 2005), species2000 (species2000, 2005), catalogue of life: 2005 annual checklist (bisby et al., 2005), index to organism names (ion, 2005), european register for marine species (erms, 2005), systema naturae 2000 (brands, s. j., 1989-2005) etc. in case of any problems regarding taxonomic placement of the species concerned taxonomy experts are contacted and as per their suggestions the species are being entered in the database. taxonomic discrepancies during this project we have noticed several taxonomic discrepancies, which need to be resolved by application of nomenclatural rules. these discrepancies can be grouped in 3 categories: hierarchical differences, spelling differences and homonymies. difference in taxonomic hierarchies several examples were found where the taxonomic hierarchy of organisms followed in india did not match that used by itis. this is especially true in case of some fishes, nematodes and insects. this is the result of differences in taxonomic opinions or the provisional nature of certain data in itis and a consensus is often difficult. in this case, the information managers can display the placement of the taxon according to alternative schemes. in spite of this option, it is necessary to conform to the international taxonomic opinion, to make the datasets interoperable with those developed in other parts of the world. this is an issue that needs to be discussed and resolved by taxonomists working in india. although making changes in taxonomic hierarchy is technically possible in case of the electronic datasets, each change needs to be validated by the taxonomic community as some taxa may or may not be conforming to that change. some of the examples where taxonomic hierarchy is different are given in table 1. as per itis and other taxonomic resources sub class elasmobranchii is placed under class chondricthyes, while systema naturae 2000 (brands, s. j., 1989-2005) still recognizes it as class elasmobranchii. differences in spelling the most common problem faced while digitizing the data was different spellings of organisms’ names. some examples are given in table 2. order cheilostomata as per itis is named differently as order cheilostomida by erms (2005). even the hierarchy under this order is not same for many species given in these two databases. chavan et al. – resolving taxonomic discrepancies 72 table 1. difference in hierarchies used in various sources. sr. no. taxon or scientific name indian sources integrated taxonomic information system 1 appendicularia histnae kingdom: animalia phylum: chordata subphylum:urochordata class: larvacea order: oikopleurida family: appendicularidae genus: appendicularia species: histnae (das, 2003; dhandapani, 1977) kingdom: animalia phylum: chordata subphylum:tunicata class: appendicularia order: copelata itis does not include genus appendicularia 2 pillaia indica pillaia khajuriai genus pillaia is placed under different order of class actinopterygii kingdom: animalia phylum: chordata class: actinopterygii order: perciformes family: chaudhuriidae genus pillaia (rao, 2000) kingdom: animalia phylum: chordata class: actinopterygii order: synbranchiformes family: chaudhuriidae genus pillaia 3 zenarchopterus ectuntio and zenarchopterus striga genus zenarchopterus is placed under different order kingdom: animalia phylum: chordata class: actinopterygii order: atheriniformes sub order: exocoetoidei family: hemiramphidae genus zenarchopterus (rao, 2000) kingdom: animalia phylum: chordata class: actinopterygii order: beloniformes sub order: belonoidei family: hemiramphidae genus zenarchopterus in some cases these were misspellings, especially typographical errors. however, to follow the taxonomic norms, each misspelling needs to be reported along with the valid scientific name for avoiding future problems. usually such wrongly spelled scientific names are reported as synonyms with a prefix “sic” as per the international code of zoological nomenclature (iczn, 1999). in some case the difference in spelling was due to different taxonomic opinions, and choosing which to use requires discussion with taxonomists. for example the blue whale shark rhincodon which is spelled rhiniodon (talwar and kacker, 1984; talwar, p. k., 1991) in indian literature. itis shows rhiniodon as a synonym of rhincodon that was suppressed by a ruling (itis, 2005). homonyms a few homonyms were identified in the cataloguing process. for instance, genus chaunoproctus and genus microcosmus were observed to be used in different regions for different taxa. this is a direct contradiction to nomenclature rules. closer scrutiny of literature and taxonomic opinion revealed interesting facts of uses of these two genera. according to itis (itis, 2005) and the taxonomicon of systema naturae 2000 (brands, s. j., 1989-2005), genus chaunoproctus bonaparte, 1850 (aves: passeriformes: fringillidae) is a bird, chaunoproctus ferreorostris vigors, 1829, which is now extinct as per the iucn redlist of threatened species (iucn, 2005). chavan et al. – resolving taxonomic discrepancies 73 table 2. misspellings and difference in hierarchies in various sources. sr. no. scientific names as per: integrated taxonomic information system indian sources 1 kingdom: animalia phylum: acanthocephala class: palaeacanthocephala order: polymorphida family: polymorphidae genus: corynosoma species: strumosum corynosoma streemosum (bhattacharya, 1998) 2 kingdom: animalia phylum: chordata class: mammalia order: cetacea family: delphinidae genus: pseudorca species: crassidens class mammalia order: cetacea genus: psudorca species: crassidens (agarwal v.c., 1998) 3 genus: amblypharyngodon genus previously present in itis but not found as on date. phylum: chordata class: actinopterygii family: cyprinidae genus: ambylopharyngodon (aditya and raut, 2001) 4 kingdom: animalia phylum: chordata class: actinopterygii order: beloniformes family: adrianichthyidae genus: oryzias species: melastigma [family: cyprinodontidae is present under order cyprinodontiformes] kingdom: animalia phylum: chordata class: actinopterygii order: cyprinodontiformes family: cyprinodontidae genus: oryzias species: melanostigma (nandi, 1993) chavan et al. – resolving taxonomic discrepancies 74 chaunoproctus pearce, 1906 (arachnida: acari) is a group of orabitid mites as per the indian faunas (sanyal and bhaduri, 1986; sanyal et al., 2003). the type species designated for this genus is chaunoproctus cancellatus pearce, 1906 collected from sikkim, india. interestingly, genus zetorchella berlese, 1916 reported from somaliland, and genus callopia balogh, 1958 reported from angola, africa were regarded as synonyms of the genus chaunoproctus by balogh (1965) (sanyal et.al., 2003). later in 1972 he again considered them as distinct genera. recently, j. balogh and p. balogh (1992) again concluded that genus zetorchella and genus caloppia as synonyms of genus chaunoproctus. in 2003, sanyal and team described three new species chaunoproctus orientalis, c. sisiri, and c. amarpurensis belonging to genus chaunoproctus from tripura, india (sanyal et.al., 2003). genus chaunoproctus is also reported from other parts of the globe as valid genus of order acari (mahunka, s., 1987). as per the international code of zoological nomenclature (iczn, 1999), generic name chaunoproctus pearce, 1906 (arachnida: acari) can be a junior homonym of chaunoproctus bonaparte, 1850 (aves: passeriformes: fringillidae). similarly the generic name microcosmus heller, 1878 (itis, 2005) [1877 (ion, 2005)] (chordata, ascidiacea) is a homonym of microcosmus chaudoir 1879 (insecta, coleoptera, carabidae) (saha, et.al., 1992). both names are in use and refer, respectively, to ascidiacea as per the itis (2005) and ion (2005), and for a beetle as per indian literature (saha, et al., 1992). opinion was sought from carabidologists and finally wolfgang schiller resolved this issue. microcosmus has been described by chaudoir (chaudoir, 1879) under tribe panageini within family carabidae (order coleoptera: class insecta: phylum arthropoda) for the species m. flavopilosus. however, heller already designated the same genus microcosmus in 1877 under family pyuridae, (class ascidiacea, phylum chordata). hence, emberik strand proposed new name microcosmodes (strand, 1936). later in 1940, andrewes replaced it to genus microschemus (yves, b., 2002). so until today the carabid species is cited as microcosmodes flavopilosus. thus, indian records (saha, et al., 1992) which show presence of microcosmus instead of microschemus needs to be updated. polypodium hydriforme, which is under phylum cnidaria, is still considered as a valid scientific name under phylum pisces along with phylum cnidaria by ion (2005). genus doto described by oken, 1815 is considered valid for phylum mollusca while; u.s. national museum of natural history database (2005) displays it under phylum arthropoda as well as phylum mollusca. in these cases it is really difficult to place the organism in a specific hierarchy. another example is of the genus cyaniris, which is present under family lycaenidae, order lepidoptera, class insecta (bingham, 1907). the same genus name cyaniris is also given to insects belonging to family chrysomelidae, order coleoptera, class insecta as per the database of the holotype collections from india present in the belgium museum (drugmand, 2002). the genus is not included in the itis database while the other web sites and literature sources shows presence of genus cyaniris in order lepidoptera. while checking hierarchy for a species, namely idia pristis, which is placed under class hydrozoa (ritchie, 1910), surprisingly we came across genus idia which is also present under class insecta, order lepidoptera, family noctuidae as per the checklist of noctuoidea of ontario present on the website hosted by canadian biodiversity information facility (troubridge and lafontaine, 2003). there is also no ready reference available for this generic designation anywhere. discussion a consensus on these issues is a matter of taxonomic discussion. they need to be resolved by using nomenclatural rules, which requires further detailed research. however these examples effectively demonstrate the potential of electronic catalogues (elecats) in bringing issues or discrepancies to the attention of the taxonomic community, starting a dialogue between taxonomists across the globe and identifying issues of common concern. in order chavan et al. – resolving taxonomic discrepancies 75 to notice such discrepancies and resolve them quickly, it is essential that a wrapper be developed which traverses through various electronic catalogues searching for taxonomic anomalies. this calls for increasing collaboration among the various elecats. the information available so far on the internet is largely limited to names and citations. it is understood, however, that taxonomic literature generated out of two centuries’ work cannot be put on the internet overnight. it is a formidable taxk, yet one that must be accomplished as promptly as possible. with the growing use of information and communication technologies in biodiversity research, it should be possible to make the taxonomic literature itself available on the internet, which can be used for checking inconsistency in taxonomy, used worldwide. although taxonomists from around the world have been dealing with these discrepancies, it is time-consuming and tedious to identify, check, and correct them using the traditional media such as published literature. modern information and communications tools can be of immense help in identifying taxonomic discrepancies quickly and resolving them in a collaborative manner leading to globally acceptable standardized inventories. with the use of the internet, there can be a truly two-way exchange of information between taxonomists from developed and developing countries. active collaboration and commitment of taxonomists and information managers are required to work towards the goal of developing information systems to bring uniformity and precision to taxonomic inventories across the world. many of the discrepancies arise because taxonomists are unable or find it difficult to check up on taxon names especially for taxa outside their field of expertise. hence, to build up easy communication pathways and reduce the time input, it would be extremely helpful to have a web-based central registry system for taxonomic names. checking of names being used in publication with the central registry would definitely eliminate many of the commonly encountered discrepancies described above. thorne (2003) also proposed the need for registration of new taxa names in a central registry of names. the journal nature has already taken a step towards registering of names (anonymous, 2002) by requiring the authors of papers featuring new taxonomy to file the information with a recognized institute such as linnaean society of london. central registry will be a repository of scientific name information or an index for scientific names in use, along with their history. it will be a dynamic register for proposed scientific names (which will be provisionally accepted, noted), which can later be added to the repository after annotation. these will serve as references for scientists describing new taxa to check if the name has been used before, and in which context. this will eliminate generation of homonyms. it can also provide a point of “normalization” for data. elecats offer an effective method of creating unique electronic registers. owing to the rules of acceptance of scientific names, names cannot be registered as valid before the publication of taxon description in a journal. to solve this, a precedent can be set that in case of each new description, together with the type specimen deposition number, a provisional registration number in the global, regional or national web based elecat should be quoted. there will be two-way information exchanges with other elecats. the registry will compare between elecats information, find out any points of mismatches or conflicting data, and also pick up new information automatically from the elecats. using this, a single number reference system for each scientific name can be developed. the central registry can provide a minimum standard and starting point for use in other databases. the global biodiversity information facility (gbif), together with the taxonomic database working group (tdwg), is currently seeking requirements for globally unique identifiers (guids) for biodiversity informatics and to establish infrastructure to support their use. guids once developed can overcome most of the current problems, such as (a) identification of same data records served from multiple locations, (b) referring to data from outside network, irrespective of frequent change of urls, and (c) referring to taxon concepts in reliable and consistent way. page (2005) chavan et al. – resolving taxonomic discrepancies 76 suggests a system of life science identifiers (lsids) as unique numerical identifiers for scientific names and itis currently employs a system of unique taxonomic serial numbers (tsns). databases could be mapped to tsns or some parallel concept. many organizations are working towards building up registers of published scientific names of taxa such as for beetles (vratislav, 2005). plant names can be checked using international plant names index (ipni, 2005). index to organism names (ion, 2005) database can be used to check zoological names. the most holistic efforts are those of species2000 and itis catalogue of life (leslie, 2005) and gbif, which aim to create an index of at least 95% of the known species by 2011 (gbif, 2005). this is a major step towards developing a central register of names, and increasing collaborations between similar efforts worldwide should shorten the time required. to complement these initiatives, international commission of zoological nomenclature announced its intention of setting up of webbased, open access mandatory registration system called “zoobank” to register descriptions of all new taxa and nomenclatural acts in animal taxonomy. (polaszek et al., 2005). while similar mandatory registration mechanism exists for descriptions of new bacteria, it needs to be extended for other kingdoms viz., plantae, archaea, chromista, fungi, protozoa, and viruses. similar to “zoobank” these mandatory registration mechanisms would facilitate retrospective registration of existing names, and of all nomenclatural acts in respective kingdom. this can be achieved through active linkages and collaborations with existing projects, rather than replacing them. in addition, links to other databases like image, dna sequences, protein sequences, lipid sequences, and collection locality maps, etc. will be a major step forward. further, linking valid scientific names to collection accession numbers of type specimens will help scientists track all collections quickly and know where they are deposited. applications could be built such as those on the itis canada website or ubio1 that display multiple classifications. in addition, a 1 http://www.ubio.org/. taxonomist’s time would be saved by having a tool that can readily compare data with that in the central file (such as itis’ taxcompare tool). the ability to check quickly for homonymies will also save time. in addition, development and use of national elecats should be encouraged to collate information at the national levels and make it available to the global users. this is especially important, as these national registers will be able to easily access locally available primary taxonomic information. it is also necessary to track changes in species concepts over time. itis has developed a capability for change tracking, but has not yet implemented it. availability of specimens, images, protologues, and classifying characters in use in different countries, comparing between specimens of a species with wide distribution transcending political boundaries and building biogeographic distribution maps, language barrier translating latin diagnoses, and picture data are some of the capabilities required to ensure accurate results in biodiversity research projects. these advances suggest that, in the future, the taxonomic discipline will make broad use of the web-based information and benefit greatly from it. therefore, it is crucial at the moment to build up or improve the collaborative activities among domain experts, information managers and users of taxonomic information. this would ultimately help in strengthening the biodiversity research necessary for conservation and management of global natural resources. conclusions the discrepancies found while developing indfauna, an electronic catalogue of known indian fauna and comparing it with existing databases can help to solve several issues like taxonomic ambiguities, inadequate documentation and incorrect placements of species. development of electronic catalogues of names of known organisms (elecats) will help in pointing out these issues. international organizations like gbif are trying to make all biodiversity data accessible to the largest possible section of the human population. recently gbif, species 2000, itis and ubio (gbif, 2005) have decided to cooperate on chavan et al. – resolving taxonomic discrepancies 77 compiling and utilizing taxonomic information resources. national and regional resources such as indfauna, after solving the types of discrepancies described here, can make valuable contributions to preparing of global taxonomic databases and standards. references aditya, g. and s. k. raut. 2001. food of the snail, pomacea bridgesi, introduced in india. current science. 80:919-920. agarwal, v. c. 1998. mammalia. pages 459-469 in j.r.b. alfred, a. k. das and a. k. sanyal, editors. faunal diversity in india. zoological survey of india, kolkata, india. agosti, d. and n. f. johnson. 2002. taxonomists need better access to published data. nature 417:222. anonymous. 2002. genomics and taxonomy for all. nature 417:573. bhattacharya, s. b. 1998. acanthocephala. pages 9398 in j.r.b. alfred, a. k. das and a. k. sanyal, editors. faunal diversity in india. zoological survey of india, kolkata, india. bingham, c. t. 1907. the fauna of british india including ceylon and burma. butterflies vol. ii, taylor and francis, london. bisby, f. a., ruggiero, m. a., wilson, k. l., cachuela-palacio, m., kimani, s. w., roskov, y. r., soulier-perkins a. and j van hertum, editors. 2005. species 2000 & itis catalogue of life: 2005 annual checklist. species 2000, reading, united kingdom. brands, s. j. (comp.). 1989-2005. systema naturae 2000. amsterdam, the netherlands.2 chaudoir, m. de. 1879. essai monographique sur les panageides. annales de la société entomologique de belgique 21: 83-186 (pp. 8384 published in 1878). das, a. 2003. a catalogue of new taxa described by the scientists of the zoological survey of india during 1916-1991. records of the zoological survey of india, occasional paper 208:1-530. dhandapani, p. 1977. descriptions of two new species of larvaceae with a list of other species collected from the bay of bengal. pages 60-64 in proceedings of the symposium on warm water zooplankton. unesco/nio (symp. national institute of oceanography, dona paula, goa, india. 2 http://sn2000.taxonomy.nl/. drugmand d. 2002. entomological and arachnological collections. royal belgian institute of natural sciences, brussels.3 european register for marine species (erms). 2005. european register for marine species (erms2), flanders marine institute, oostende, belgium.4 global biodiversity informatics facility (gbif). 2005. gbif strategic plan 2007-2011, version 6.0 (1 september 2005). global biodiversity informatics facility, copenhagen, denmark. international commission on zoological nomenclature (iczn). 1999. international code of zoological nomenclature, 4th edition. international commission on zoological nomenclature, london, u.k.5 index to organism names (ion). 2005. biosis index to organism names. thomson biosis ltd., york science park, york, u.k.6 integrated taxonomic information system (itis). 2005. integrated taxonomic information system, washington, d.c.7 international plant names index (ipni). 2005. international plant names index. plant names project.8 international union for conservation of nature and natural resources (iucn). 2005. iucn redlist of threatened species. international union for conservation of nature and natural resources, cambridge, u.k.9 kalavati, c. 1998. mesozoa. pages 21-26 in j. r. b. alfred, a. k. das and a. k. sanyal, editors. faunal diversity in india. zoological survey of india, kolkata, india. knapp, s., bateman, r. m., chalmers, n. r., humphries, c. j., rainbow, p. s., smith, a. b., taylor, p. d., vane-wright, r. i. and m. wilkinson. 2002. taxonomy needs evolution, not revolution. nature 419:59 leslie, m. 2005. species master list hits milestone. science 308:609. mahunka, s. 1987. oribatids from africa: acari oribatida v. folia entomologica hungarica 48:105-128. may, r. m. 1999. the dimensions of life of earth. pages 30-45 in p. h. raven and t. williams, editors. nature and human society, the quest 3 http://www.natuurwentenschappen.be/collections/entomo/ types_form/chrysomelinae_clythrinae.htm. 4 http://www.marbef.org/data/erms.php. 5 http://www.iczn.org/iczn.htm. 6 http://www.biosis.org.uk/ion/search.htm. 7 http://www.itis.usda.gov/. 8 http://www.ipni.org/. 9 http://www.redlist.org/search/details.php?species=4591. chavan et al. – resolving taxonomic discrepancies 78 for sustainable worlds. national academy press, washington, d.c.. nandi, n. c., s. r. das, bhuinya and j. m. dasgupta. 1993. wetland faunal resources of west bengal, i., north and south 24-parganas districts. records of the zoological survey of india, occasional paper 150:1-50. national chemical laboratory (ncl). 2005. electronic catalogue of known indian fauna. national chemical laboratory, pune, india.10 page, r. d. m. 2005. a taxonomic search engine: federating taxonomic databases using web services. bmc bioinformatics 6:48. polaszek, a., d. agosti, m. alonso-zarazaga, g. beccaloni, p. de p. bjørn, p. bouchet, d. j. brothers, e. n. earl, h. c. j. godfray, n. f. johnson, f.-t. krell, d. lipscomb, c. h. c. lyal, g. m. mace, s. mawatari, s. e. miller, a. minelli, s. morris, p. k. l. ng, d. j. patterson, r. l. pyle, n. robinson, l. rogo, j. taverne, f. c. thompson, j. tol, q. d. wheeler and e. o. wilson. 2005. a universal register for animal names. nature 437:477. rao, d.v., kamla devi and p.t.rajan.. 2000. an account of ichthyofauna of andaman and nicobar islands, bay of bengal. records of the zoological survey of india, occasional paper 178:1-434. ritchie, j. 1910. the hydroids of the indian museum. i. the deep sea collection. records of the indian museum v:1-30. 10 http://www.ncbi.org.in/. saha s. k., mukherjee a. k. and t. sengupta. 1992. carabidae (coleoptera: insecta) of calcutta. records of zoological survey of india, occasional paper 144:1-63. sanyal, a. k. and a. k. bhaduri. 1986. check list of oribatid mites (acari) of india. records of the zoological survey of india, occasional paper 83:1-79. sanyal, a. k. saha, s. and s. chakraborty. 2003. three new species of the genus chaunoproctus pearce (1906) (acarina: oribatida) from india. records of the zoological survey of india, 101:57-66. species2000. 2005. species2000, school of plant sciences, the university of reading, reading, united kingdom.11 strand, e. 1936. miscellanea nomenclatorica zoologica et palaeontologica ix. folia zoologica hydrobiologia riga 9:167-170. talwar, p. k. 1991. pisces. pages 1-143 in faunal resources of ganga. zoological survey of india, kolkata, india. talwar, p. k. and r. k. kacker 1984. commercial sea fishes of india. zoological survey of india, kolkata, india. thorne j.. 2003. zoological record and registration of new names in zoology. bulletin of zoological nomenclature 60:7-11. 11 http://www.sp2000.org/. microsoft word s_chm_2004_21.doc biodiversity informatics, 1, 2004, pp. 23-29 23 bioinformatics, the clearing-house mechanism, and the convention on biological diversity marcos silva clearing-house mechanism, secretariat of the convention on biological diversity, united nations environment programme, 393 saint-jacques street, suite 300, montreal, quebec, h2y 1n9, canada abstract.—this paper discusses the relevance of bioinformatics to the convention on biological diversity. it also discusses the role of the convention’s clearing-house mechanism and the biosafety clearinghouse, and their role in assisting parties and other governments with issues pertaining to bioinformatics in general. finally, it reviews many work programs established under the convention and the potential role of bioinformatics in assisting with their development and implementation. key words.—biodiversity informatics, biodiversity convention, international cooperation the word bioinformatics does not have a uniform usage: it may refer primarily to genomics (sugden and pennisi 2000, ouzounis and valencia 2003), or may also include more general issues related to biology and computation (scott 2004). this paper defines bioinformatics as the application of tools of computation and analysis to capture and interpretation of biological data (bayat 2002). for many reasons, the conference of the parties (cop) to the convention on biological diversity is paying increased attention to issues related to computational biology and to bioinformatics in general. indeed, data collection and dissemination, data mining and modeling, development, interoperability and further enhancement of databases, and visualization of data were explicitly identified by the cop during its seventh meeting, and are found throughout the work programmes of many thematic areas and cross-cutting issues of the convention. this paper discusses the development and role of the convention’s clearing-house mechanism (chm), synergies with the biosafety clearing-house (bch), and chm’s support for initiatives related to bioinformatics in general. it also examines the significance of bioinformatics in assisting parties in implementation of obligations under the convention, and chm’s role in facilitating activities by parties and governments to exploit benefits arising from the evolving bioinformatics global infrastructure. the view is also presented that development of bioinformatics tools and resources is having positive impacts on the ability of parties to meet the three objectives of the convention more effectively and to manage their biodiversity resources better. this trend is reflected by the overall impact of bioinformatics in the programme areas and cross-cutting issues of the convention and its related activities, especially work related to the 2010 target. the clearing-house mechanism chm was created pursuant to article 18, paragraph 3, of the convention, to promote and facilitate technical and scientific cooperation between state parties. as identified in its strategic plan (cbd 1999), chm has three primary objectives: cooperation, promotion, and facilitation of technical and scientific cooperation; information exchange; and network development. chm’s original mandate has been broadened to include matters pertaining to information exchange (article 17 of the convention) and the biosafety clearinghouse (bch), pursuant to article 20, paragraph 1, of the cartagena protocol on biosafety. in general, then, issues related to bioinformatics fall within the purview of the chm. the biosafety clearing-house the bch was established as part of the chm pursuant to article 20.1 (cbd 2001) of the cartagena protocol on biosafety to: (a) facilitate the exchange of scientific, technical, environmental and legal information on, and experience with living modified organisms; (b) assist parties in the implementation of the protocol, taking into account the special needs of developing state parties, in particular the least developed and small island developing states among them, and countries with economies in transition as well as countries that are centres of origin and centres of genetic diversity. silva – the biodiversity convention and bioinformatics 24 in this broad context, chm recommended the technical and architectural framework that was eventually implemented in the pilot phase of the bch, and it continues to oversee the framework’s technical functioning and enhancements. in addition, chm was responsible for the technical architecture of the bch pilot phase; protocols and standards supporting interoperability among disparate and distributed databases; and design, modules, and text of the draft bch toolkit. the bch continues to evolve, particularly in the context of decisions made at the first meeting of the cop, serving as the meeting of the parties to the cartagena protocol on biosafety. indeed, at the meeting, it was decided to approve the transition of the bch pilot phase to its fully operational phase. of interest are decisions pertaining to modalities of the bch which include use of a decentralized internetbased system, adherence to common formats, use of controlled vocabularies and metadata, and providing means to facilitate interoperability of data among disparate and remote databases (cbd 2004b). synergies between chm and bch it should be noted that the bch may be the first instance in which parties have the opportunity to register information electronically to fulfill obligations under a legally binding international treaty. indeed, it would be difficult, if not impossible, to imagine an effective and functioning cartagena protocol without a web-based bch. in other words, parties have a legally binding obligation to register data under the biosafety protocol, and this fact alone differentiates the narrower, more concise bch mandate from the very broad chm mandate. in contrast, the very broad chm mandate was developed to promote and facilitate technical and scientific cooperation among parties and governments. although cop has repeatedly called for parties to establish chm national focal points, parties are not legally obliged to do so. in addition, chm operates in a much broader environment, in which information exchange is but one of its activities, albeit a singularly essential one. even more, although chm has assumed responsibility for information exchange under the convention, this activity falls under the prerogative of article 17 of the convention, and not article 18.3. however, the uniqueness of bch can be understood as but one of the functions of chm: providing a secure, stable, and authenticated data registration, searching, and retrieval mechanism in support of activities and obligations undertaken by parties. if bch is further understood as being part of chm (per article 20.1 of the cartagena protocol), then, in this narrow context, it becomes possible to delineate synergies between the two clearing-houses, especially relating to issues pertaining to bioinformatics in general. capacity building first, both clearing-houses implement activities to assist parties in raising national and regional capacities. with regard to biosafety, capacities required are clearly associated with obligations under the protocol, particularly the use and navigation of the bch. chm, again, operates within the entire framework of the convention, including technical issues associated with bch. the synergies, therefore, between these two clearing-houses deal primarily with technical issues, namely assisting countries and regions to gain adequate capacity in the acquisition, on-going support and use of new information technologies. formats, protocols, and standards work on recommendations pertaining to adoption and use of data formats, protocols, and standards is a concern held by chm in relation to its general work on information exchange and in its development of a bch that can function in and accommodate a distributed, interoperable data and information environment. chm’s work with formats, protocols, and standards was formalized by a decision at the fifth meeting of the cop, in which the executive secretary was requested to identify possible formats, protocols, and standards for improved exchange of biodiversity-related data, information, and knowledge, including national reports, biodiversity assessments, and global biodiversity outlook reports, and convene an informal meeting on this issue (cbd 2000a). with regard to the bch, the meeting of the technical experts on bch, held in montreal, canada (september 2000), recommended that the pilot phase of bch be developed “…using a combination of centralized/decentralized information systems to offer the biosafety clearing-house the necessary flexibility for better coordination of the submission of data while ensuring timeliness and links to complementary distributed information” (cbd 2000b). this recommendation entailed formulation silva – the biodiversity convention and bioinformatics 25 of data formats, protocols, and standards, for use with bch and for adoption by parties desiring to make their national biosafety clearing-houses interoperable with bch, ensuring that bch can function in a distributed information environment. metadata recognition of the significance of metadata to the development of chm was voiced by the chm informal advisory committee (november 2001), with the recommendation to use the dublin core as the metadata standard for the cbd web site, and to continue further development of controlled vocabularies and metadata standards. a core group was constituted to examine this issue and to report at an informal meeting on formats, protocols, and standards for improved exchange of biodiversityrelated information (montreal, february 2002). this meeting gave further support for use of the dublin core as the metadata standard for the cbd web site. this work on metadata was ported to the bch. controlled vocabularies use of controlled vocabularies by chm or bch national focal points becomes an essential tool in assisting parties and governments in making information interoperable. the secretariat has made controlled vocabularies developed by chm and bch available through its website and through the chm and bch toolkits. in many ways, these activities complement those of other organizations, such as the global biodiversity information facility (gbif) and the inter-american biodiversity information network (iabin), and reflect elements needed for effective use of bioinformatics as a support element throughout activities related to the program areas and cross-cutting issues of the convention. bioinformatics and the convention it is within this context of promotion of technical and scientific cooperation, capacity-building, adherence to common formats, protocols and standards, and controlled vocabularies, that the chm strives to assist parties and governments to make use of resources from bioinformatics initiatives and projects, gbif’s data portal being a case in point. indeed, the need to search, retrieve, and analyze information hosted in very large distributed databases of biological data is increasingly important to the work of the cop. note, for example, the increasing use of web services, particularly xml and soap (simple object access protocol), to disseminate software applications that cannot be distributed as source code (jamison 2003), to search and aggregate data from different databases (stein 2002), and to assist in issues related to invasive alien species (morris 2004). discussion of a few activities, program areas, and cross-cutting issues of the convention will illustrate these points. in the strategic plan of the convention, parties emphasized the need for more effective and coherent implementation of the three objectives of the convention, and agreed to reduce current rates of biodiversity loss significantly at global, regional, and national levels by 2010. parties also adopted a framework to facilitate assessment of progress towards 2010 and communication of this assessment, to promote coherence among convention work programs, and to provide a flexible framework within which national and regional targets may be set and indicators identified. the framework includes seven focal areas. the cop identified indicators for assessing progress towards the 2010 target at the global level, and goals and sub-targets for each focal area, as well as a general approach for integration of goals and sub-targets into convention work programs. arguably, the challenge before the parties regarding the 2010 target is how to quantify and measure existing biodiversity, and how to quantify its loss or conservation. indeed, regarding measurement and quantification, one result of the meeting on “2010 – the global biodiversity challenge” (london, may 2003) was: “there is a need to make the biodiversity data that exists more readily accessible and available in a timely manner. actions to achieve this would include: a) disseminating information in appropriate formats for potential users; b) using best-practice in information management and dissemination; c) supporting the development and implementation of tools, standards and protocols for data exchange that allow more effective sharing of information; d) establishing inter-operable electronic databases that allow for more effective integration of information from multiple sources in real time; e) improving use of the internet as a tool for access and dissemination of biodiversity data, including increasing access to the internet; silva – the biodiversity convention and bioinformatics 26 f) reviewing the adequacy of the existing data, assessing gaps and the action that needs to be taken to fill them.” (cbd 2004a) the meeting also emphasized that “assessment is necessary of the datasets currently available, either through compilation of national-level data or through remote sensing, and of the processes for maintaining these data, in order to determine their potential value in addressing monitoring and assessment of achievement of the 2010 target, and their ability to contribute indicators” (cbd 2004a). such an exercise would necessitate assessment of relevant large-scale databases and review of issues related to their access. in addition, it would also demand adherence to common formats, protocols, and standards, as well as access to taxonomic databases, specimen and observational data databases, and remotely sensed data, all integrated via geographic information systems, analytical tools, and data mining programs. in all, these activities point to schmidt et al.’s (2004) comment that the exponential growth of biological databases is establishing the need for high-performance computing (hpc) in bioinformatics to analyze distributed data more optimally. it may also illustrate the need for systems with search and analysis functions analagous to blast1 (basic local alignment search tool capable of searching databases for genes with similar nucleotide structure; bayat 2004). regarding issues related to taxonomy and the global taxonomy initiative (gti), the cop pointed to the role of organizations and initiatives concerned with bioinformatics and taxonomy. for example, in its work program, it calls for facilitation of improved effective infrastructure for access to taxonomic information, with priority on ensuring access to information concerning elements of native biodiversity to countries of origin. principal actors identified include ecoport, gbif, species2000, integrated taxonomic information system (itis), tree of life, isis, and bionet international, as well as large-scale biosystematics research institutions and other stakeholders of taxonomic information, in collaboration with chm. such partnerships take on added urgency in light of new bioinformatics 1 http://www.ncbi.nlm.nih.gov/blast/. projects attempting to use dna barcodes to identify specimens (herbert et al. 2004). arguably, establishment of an internet-based “…public library of dna barcodes linked to named specimens could provide a new master key for identifying species, one whose power will rise with increased taxon coverage and with faster, cheaper sequencing (herbert et al. 2004).” similar emphasis on need for taxonomic data resources for identification and monitoring is inherent in decisions regarding invasive alien species. for example, cop requested the executive secretary to support development and dissemination of technical tools and related information on prevention, early detection, monitoring, eradication, and/or control of invasive alien species. moreover, the same decision requests several activities that would profit from an efficient global bioinformatics infrastructure (table 1), particularly once it becomes table i: decision vi/23, paragraph 28: alien species that threaten ecosystems, habitats or species. 1. compilation and dissemination of case-studies submitted by parties, other governments and organizations, best practices and lessons learned, drawing upon, as appropriate, tools listed in information document unep/cbd/sbstta/6/inf/3 and the “toolkit” compiled by the global invasive species programme; 2. further compilation and preparation of anthologies of existing terminology used in international instruments relevant to invasive alien species, and to develop, and update as necessary, a non-legally binding list of terms most commonly used; 3. compilation and making available lists of procedures for risk assessment/analysis and pathway analysis which may be relevant in assessing the risks of invasive alien species to biodiversity, habitats and ecosystems; 4. identification and inventory of existing expertise relevant to the prevention, early detection and warning, eradication and/or control of invasive alien species, and restoration of invaded ecosystems and habitats, which may be made available to other countries, including the roster of experts for the convention on biological diversity; 5. development of databases and facilitated access to such information for all countries including repatriation of information to source countries, through, inter alia, the clearing-house mechanism; 6. development of systems for reporting new invasions of alien species and the spread of alien species into new areas. silva – the biodiversity convention and bioinformatics 27 possible to resolve synonyms of scientific names of species that refer to the same taxonomic concepts, an issue currently under discussion by gbif. the need for a well-functioning global bioinformatics infrastructure is explicit in other convention-related program areas and cross-cutting issues. for example, phase iii of the proposed process for periodic assessment of status and trends of biological diversity in dry and sub-humid lands calls for data collection, processing, and communication according to agreed guidelines and mechanisms (cbd 2004c). the work program on mountain biodiversity likewise points to the need “…[t]o improve the infrastructure for data and information management for accurate assessment and monitoring of mountain biological diversity and develop associated databases” and “…[e]nhance and improve the technical capacity at a national level to monitor mountain biological diversity, benefiting from the opportunities offered by the clearing-house mechanism of the convention on biological diversity, including the development of associated databases as required at the global scale to facilitate exchange” (cbd 2004c). an implicit recognition regarding the link between monitoring and need for better data access is the invitation to parties by the cop to “[i]mprove and update national and regional databases on protected areas and consolidate the world database on protected areas (wdpa) as key support mechanisms in the assessment and monitoring of protected area status and trends” (cbd 2004c). the cop also recognized the important links between the mandate of the chm and issues related to the transfer of technology. the work program on technology transfer is unequivocal on the role of the chm as the “…central mechanism for the exchange of information on and facilitation of technology transfer and technical and scientific cooperation relevant for the convention on biological diversity” (cbd 2004c). furthermore, cop requested the executive secretary, in collaboration with parties, the informal advisory committee of the chm, and relevant organizations and initiatives, to work on “development of advice and guidance on the use of new information exchange formats, protocols and standards to enable interoperability among relevant existing systems of national and international information exchange, including technology and patent databases” (cbd 2004c). working relationships with bioinformatics initiatives and organizations this concern for access to data in support of programme goals and activities of the convention is reflected throughout other decisions from the seventh meeting of the cop. in response to these decisions and to ones from previous cops, the secretariat established memoranda of understanding with a number of bioinformatic initiatives and organizations with a view to facilitate access to data resources by parties and governments, repatriate information and enhance national and regional capacities with regard to the adherence to and use of common formats, protocols, and standards. for example, a memorandum of understanding was completed with gbif to establish a framework of collaboration between the cbd and gbif secretariats to further common goals. these goals include facilitating development and implementation of approaches, technologies, and best practices necessary to access, share, and disseminate biodiversity data at the species, ecosystem, and genetic levels via the internet. furthermore, the cbd secretariat is an ex-officio member in gbif governing council, and participates in many of its working committees. the potential impact of this working relationship should not be underestimated, given gbif’s development of its data portal; the electronic catalogue of names of known organisms; common formats, protocols, and standards to facilitate interoperability; and its efforts to enhance national and regional capacities regarding bioinformatics. these projects are, in part, efforts to solve problems related to accessing and using data, or “the scientific and data management communities have expressed a number of concerns in recent years regarding the collection and dissemination of data. considerable amounts of data are still held by scientists or institutes that have not released their findings (reynolds et al. 1997). many data are generated to be analyzed once, published, and often not visited again (gosz 1994). data suppliers either are disconnected from wide-area networks, lack a standard mechanism for informing potential users (walker et al. 1992), or prefer to advertise and distribute their products through more traditional means. the bottleneck in sharing information may be not knowing where to find useful information. in other cases, though information may exist on the network, no systematic directories exist to guide a user through the thousands of data sources.” (xu et el. 2003). silva – the biodiversity convention and bioinformatics 28 a similar working relationship has been established with iabin to facilitate development and implementation of technologies and best practices necessary to share knowledge and information relevant to biodiversity conservation and sustainable management. in response to the above, the secretariat and iabin are working to promote adoption of interoperability standards, share expertise, collaborate in development and implementation of programs pertaining to use of biodiversity information and management tools, and collaborate in enhancing national and regional capacities. another venue for collaboration regarding data exchange has been development of international thematic focal points under the chm. to date, four international thematic focal points have been established: birdlife international, the global invasive species programme (gisp), gti, and natureserve. several actions have been initiated as a result of these international thematic focal points. for example, the secretariat is supportive of birdlife international’s efforts to develop a bi-national ecoregion-based clearing-house mechanism for the dry forests of peru and ecuador (the tumbesian endemic bird area). as well, the secretariat is working with gisp in development of a global network of information on invasive species. natureserve has made its databases of information available through the chm, and is cooperating in enhancing national and regional capacities. these few examples serve to illustrate the impact of bioinformatics in the work program of the secretariat, and the role of chm in facilitating and promoting bioinformatics at national and regional levels. inarguably, access to and exchange of data and information are fundamental to the convention’s work program, and are key elements in ensuring development and enhancement of national and regional capacities. conclusion the potential and impact of bioinformatics and the role of the chm to the convention’s work programs and activities should not be underestimated. bioinformatics ensures success of activities related to the three objectives of the convention, and offers parties and governments means by which to implement obligations under the convention more effectively, the need for accurate data to measure biodiversity loss in relation to the 2010 target being a case in point. as technologies evolve, and more data and information become available for mining and analysis, bioinformatics will most likely continue to gain prominence and receive increased investment. this increase may indicate an increased role for chm, including more active and collaborative work programs with bioinformatics organizations and initiatives, particularly initiatives to develop local access to regional and global information networks. put another way, as stated by laihonen et al. (2003), in their discussion of geospatially structured biodiversity information as a component of a regional chm: “the regional viewpoint combined with the exploitation of geo-referenced information through modern information technology can offer opportunities lacking in coarse-grained national level chms. in further development of mechanisms for biodiversity information sharing, this should be seen as a resource enhancing our knowledge and wisdom on biodiversity at the national and global levels.” after all, an important premise of biodiversityrelated research is the need to examine variables under a holistic approach. this goal can be achieved only through a common and agreed upon strategy of information sharing; adherence to common formats, protocols, and standards; and equitable access to technologies and expertise. such agreements are possible only when parties, governments, and regions have equitable access to technologies and knowledge. for these reasons, parties should continue to strengthen and support the chm as the global tool for biodiversity related technical and scientific cooperation, particularly in light of work requiring access to bioinformatics resources and expertise. references bayat, a. 2002. science, medicine, and the future: bioinformatics. british medical journal 324:1018-1022. hebert, p. d. n., m. y. stoeckle, t. s. zemlak, c. m. francis. 2004. identification of birds through dna barcodes. plos biology 2:e312. jamison, d. curtis. (2003). open bioinformatics. bioinformatics 19:679–680. laihonen, p., m. rönka, h. tolvanen and r. kalliola. 2003. geospatially structured biodiversity information as a component of a silva – the biodiversity convention and bioinformatics 29 regional biodiversity clearing house. biodiversity and conservation 12:103-120. morris, r. a. 2004. web services: who, what, why, where, and when? pp. 37-41 in meeting on implementation of a global invasive species information network: proceedings of a workshop (e. sellers, a. simpson, j. p. fisher and s. curd-hetrick, eds.). national biological information infrastructure, reston, virginia. ouzounis, c. a. and a. valencia. 2003. early bioinformatics: the birth of a discipline—a personal view. bioinformatics 19:2176-2190. schmidt, b., l. feng, a. laud and y. santoso. 2004. development of distributed bioinformatics applications with gmp. concurrency and computation: practice and experience 16:945959. scott, r. l. 2004. bioinformatics: essay review. perspectives in biology and medicine 47:135139. cbd. 1999. strategic plan of the clearing-house mechanism. (unep/cbd/cop/5/inf/32) cbd. 2000a. report of the fifth meeting of the conference of the parties to the convention on biological diversity (unep/cbd/cop/5/233). cbd. 2000b. report of the meeting of technical experts on the biosafety clearing-house (unep/cbd/bs/te-bch/1/54). cbd. 2001. report of the intergovernmental committee for the cartagena protocol on biosafety on the work of its first meeting (unep/cbd/iccp/1/95). cbd. 2004a. consideration of the results of the meeting on “2010–the global biodiversity challenge” (unep/cbd/cop/7/inf/226). cbd. 2004b. report of the first meeting of the conference of the parties serving as the meeting of the parties to the protocol on biosafety (unep/cbd/bs/cop-mop/1/157). 2 http://www.biodiv.org/doc/meetings/cop/cop-05/information/cop-05inf-03-en.pdf. 3 http://www.biodiv.org/doc/meetings/cop/cop-05/official/cop-05-23en.doc. 4 http://www.biodiv.org/doc/meetings/bch/tebch-01/official/tebch-01-05en.pdf. 5 http://www.biodiv.org/doc/meetings/bs/iccp-01/official/iccp-01-09en.doc. 6 http://www.biodiv.org/doc/meetings/cop/cop-07/information/cop-07inf-22-en.doc. 7 http://www.biodiv.org/doc/meetings/bs/mop-01/official/mop-01-15en.doc. cbd. 2004c. report of the seventh meeting of the conference of the parties (unep/cbd/cop/7/218) stein, l. 2002. creating a bioinformatics nation. nature 417:119-120. sugden, a. and e. pennisi. 2000. diversity digitized. science 289:2305. xu, h., d. wang x. sun. 1999. biodiversity clearing-house mechanism in china: present status and future needs. biodiversity and conservation 9:361–378. the views expressed in this paper are those of the author and do not reflect the official position of the convention on biological diversity. 8 http://www.biodiv.org/doc/meetings/cop/cop-07/official/cop-07-21part1-en.doc. microsoft word sw_manis_2004.doc biodiversity informatics, 1, 2004, pp. 14-22 14 mammals of the world: manis as an example of data integration in a distributed network environment barbara r. stein and john wieczorek museum of vertebrate zoology, university of california, berkeley, berkeley, ca 94720 abstract.—natural history collections are the authoritative source of knowledge about the identity, evolutionary relationships, and attributes of species with which we share this planet. as such, collections of research specimens play a central and critical role in the conservation and sustainable use of biodiversity. the potential contribution of specimen data to systematic, genomic, and ecological analyses is enormous, and will be orders of magnitude greater when information is made easily accessible via distributed networks compared with stand-alone database systems in use up to the present. the mammal networked information system (manis) is a distributed database network that permits participating institutions to provide web-based global access to their collections data for research, education and informed decisionmaking. the simplicity of the network’s design ensures that any institution wishing to join manis may do so at relatively little cost and with relatively little technical expertise. although development of manis and its underlying architecture relied on a number of key programming tasks and innovations, much of what the project can offer at this pivotal juncture is insight into its approach and a template by which other disciplines can engage in a similar process with equal success. key words.—mammals, natural history museums, biodiversity data, distributed data networks, manis, digir the desire to link and retrieve electronic data from geographically distributed natural history collections to increase their effective use in research, conservation, and education is not new, but successful attempts to achieve that objective have only emerged within the last decade. the first distributed query system (1993) for natural history collections was the national science foundation (nsf)-funded fishgopher project1. in contrast, a centralized data warehouse approach was taken by the neodat ii project in 19972. this was followed by implementation of the z39.50 profile for distributed natural history collections data (zbig), which debuted in 1998 with the distributed database of north american bird data (peterson et al., 2003). at approximately the same time, a consortium of mexican and foreign institutions made their combined specimen information available online through remib, the mexican network of biodiversity information3 using tcp/ip sockets to ensure safe and efficient data transmission. an effort to exploit the z39.50 technology in another discipline followed in 2000 with a joint collaboration between the ichthyology and marine biology communities in development of fishnet (vieglais et al., 2000; wiley et al., in prep.). use of the z39.50 1 http://www.as.ua.edu/biology/uaic/fishgopher.html. 2 http://www.neodat.org/. 3 http://www.conabio.gob.mx/remib_ingles/doctos/remib_ing.html. protocol and tcp/ip sockets represented an advance over the centralized data warehouse concept by using a standard for distributed information retrieval that simultaneously accessed data directly from institutional databases. a serious commitment to develop an interdisciplinary standard for unified interoperability among natural history databases, as opposed to continued insular parallel developments within taxonomic disciplines, was first voiced at the annual meeting of the taxonomic database working group (tdwg) in frankfurt, germany in november, 20004. a committee was formed to create guidelines that would support such a collaboration and it was agreed that interoperability would benefit from the adoption of a network architecture that not only used a mainstream transport protocol (http) and document structure (xml), but that also allowed user communities to define the structure of the data to be shared without affecting the protocol or existing software. what emerged from subsequent deliberations is a highly successful, cooperative, open source, international development effort. the resulting protocol, distributed generic information retrieval (digir)5, is designed to support unified queries to geographically-distributed providers of data via one or more portals—software installed in 4 http://www.tdwg.org/news2001.html. 5 http://digir.net. stein and wieczorek – the manis data network 15 conjunction with a web server that manages the connections to, constructs and sends requests for data to, and receives and processes responses from digir providers. the mammal networked information system (manis), an nsf-funded project6 consisting of a consortium of seventeen north american mammal collections, was developed simultaneously with digir and was the first functional implementation of a digir-based network. the readiness of the mammal community to share and benefit from the power of the combined data in their institutional collections became a major force driving development of digir. functional implementations of both provider and portal software were needed by the digir project to demonstrate the utility of the protocol. a set of natural history concepts was also needed so that collections could map information in their databases to a well-defined semantic standard. these same requirements were central to manis for the creation of a scalable network. from an informatics perspective, the mammal community was well suited to become involved in development of a distributed database network. the discipline encompasses a relatively stable taxonomy; the taxa are well-known at the species level and they follow a generally-accepted taxonomic authority that has been rendered into electronic form (wilson and reeder, 1993). this volume, mammal species of the world, is used as the basis of the taxonomic entries in the integrated taxonomic information system (itis7) and species 20008. in addition, mammal collections in north america have achieved a high degree of computerization of their specimen data (hafner et al., 1997). under the auspices of the committee on information retrieval of the american society of mammalogists (asm), documentation standards for data processing in mammalogy (mclaren, 1999) were developed more than two decades ago (williams et al., 1979) and by late in 2000 had been implemented by essentially all north american institutions that house mammal collections, regardless of size. early guidelines for the use of computer-based collection data (mclaren, 1988) also highlight the preparedness of this taxonomic discipline to meet the social and technical 6 https://www.fastlane.nsf.gov/servlet/showaward?award=0108161. 7 http://www.itis.usda.gov/. 8 http://www.sp2000.org/. challenges of establishing a distributed database network. development of digir, and in turn the manis network9, relied on a number of key programming tasks and innovations, including designing conceptual schemas (for specimen data as well as taxonomic and geographic indexes) to assist in data discovery, writing migration scripts to periodically update data in manis repositories, designing worldwide web interfaces and data portals, creating tools for providers to monitor usage, responding to diagnostic messages from providers, allowing for data caching, dealing with unresponsive providers, and creating a provider installation package for institutions that were not funded through the original nsf award. independent development of web-based tools and guidelines to facilitate efficient georeferencing of specimen localities10 was of equal import to make these data valuable to a much broader segment of the research community (ecologists, biogeographers, conservationists, educators, etc.). although a great deal of technical development was required to create the manis network, it can be strongly argued that the sociological issues involved in successfully developing a distributed database network within a community of scientists and institutions eclipsed any of the technological problems that had to be overcome. data and software aside, much of what manis can offer at this pivotal juncture is insight into its approach to the project and a template by which other disciplines can engage in a similar process with equal success. manis objectives genesis of the manis network can be traced directly to a symposium held at the asm annual meeting in seattle, washington in 1999. entitled, “emerging database technologies”, the symposium included a presentation by a. townsend peterson (university of kansas natural history museum and biodiversity research center) about the z39.50based avian network and the vast array of research and conservation questions and applications to which the power of such combined specimen data could be applied. by the close of the seattle meeting, representatives from 17 north american mammal collections agreed to collaborate on the development 9 http://elib.cs.berkeley.edu/manis. 10 http://elib.cs.berkeley.edu/manis/search.shtml stein and wieczorek – the manis data network 16 of such a distributed database network. in order to keep the project affordable and manageable in a three-year funding period, no further recruiting was attempted. the initial clarity of the objectives of manis and the concrete benefits articulated to and recognized by the participants were directly responsible for development of a successful proposal to nsf. it was noted from the outset that design of the network needed to benefit the participants as well as the larger user community. recent closures and significant reductions in programs and operating budgets at many university and free-standing museums was, and still is, a reality that needed to be addressed when asking curators and their staff to make a significant commitment of time and effort to the manis project. when crafting the project proposal, it was acknowledged that institutions would be unwilling and/or unable to relinquish their current in-house database management systems in order to participate in manis, that they would have to be able to document network use of their collections, and that collections support and new hardware would enhance their ability to participate in the project, as well as their standing with institutional administrators. equally important, the design of the network would have to ensure that institutions retain control over which data were accessible to address concerns about the security of rare species and the intellectual property rights of institutional data providers (brooke, 2000; graves, 2000). developing an approach to the project that addressed these issues was essential. moreover, it was generally felt that manis would afford each of the participating institutions and their collections heightened visibility within the larger research community. for many of the original participants, manis presented the first opportunity to make their specimen data directly accessible via the internet. increased visibility and ease of data access via the network would, in turn, validate the importance of the collections, with greater institutional support viewed as a potential outcome. from the participants’ viewpoint, the project simply reinforced the unambiguous value they placed on their own collections for research, conservation, and education. implementation issues in addition to the clarity of objectives noted above, it was recognized in the earliest planning stages of manis that the network architecture had to be simple, low cost, and require minimal maintenance. other constraints driving project design included no visible long-term support for the network or its participants, known opposition within the community to centralization of operations, and uncertain availability of in-house technical expertise to maintain institutional systems after the funding period ended. in addition, given an articulated commitment to develop an interdisciplinary standard for unified interoperability among natural history databases, every attempt was made to conceive a design that could be easily adopted by others with similar needs. the resulting manis network architecture is novel in two respects. first, each manis data provider automatically maintains summary data (counts of specimen records indexed by taxonomy and geography), in addition to specimen data from its institutional database. second, the configuration of the publicly accessible repositories is optimized for query performance rather than for data management, for which the institutions’ in-house databases are optimized. data are filtered and standardized at each institution in automated periodic migrations from curatorial databases to the manis data provider repositories. this design allows institutions to retain control over public access to their data without the need to create new structures, or change original data for public consumption, in their curatorial databases. use of replicated databases also protects curatorial databases from increased traffic and unsolicited intrusion. given the age of many curatorial databases, this design element was considered essential. in addition, the network will continue to function even if a curatorial database experiences problems or requires maintenance. to keep the broader natural history community apprised of our progress, a project web site was established and all relevant documents were posted11. the site has made a much more interdisciplinary group of researchers aware of the project than was initially envisioned. thus, an educated community of users and enthusiasts has been slowly gaining ground through workshops, training sessions and presentations at scientific conferences. a page on the manis web site summarizes the most notable of these events12, providing a limited chronology of the evolution of distributed data networks within the 11 http://elib.cs.berkeley.edu/manis/documents.html. 12 http://elib.cs.berkeley.edu/manis/events.html. stein and wieczorek – the manis data network 17 natural history community, as well as the manis project itself. standards are paramount the importance of standards in all phases of development of manis cannot be overstated. because of the informatics groundwork that had been laid down by the mammal community prior to the initiation of manis, both the semantics (the standards expressed in a conceptual schema or map of database concepts and their relationships13), and data standards, (the recommended standard vocabulary for data content (mclaren, 1999)) were adopted by the group with relative ease. such agreements are essential in any federation to be sure that fields are populated with consistent values having the same meaning across collections. the quick adoption within the context of manis does not trivialize the necessity for other communities to formulate comparable standards if they do not exist prior to embarking on a similar project. furthermore, there are benefits to cross-disciplinary standards that should also be considered. for example, a set of common core concepts between mammalogy and parasitology would allow information from both disciplines to be accessed simultaneously and shared without confusion. within the mammal community, the existence of a relatively stable and accepted taxonomy and the prior existence of data content standards meant that the specimen data housed in participant collections were already largely consistent across institutions and in suitable electronic form at the time of proposal preparation. of equal import, this meant that subsequent discussion of a data exchange standard (the federated conceptual schema) precipitated relatively little debate. only slight modifications to the initial draft were required, because the concepts agreed upon simply mirrored those in the participants’ institutional databases and had long-standing acceptance within the discipline as a whole. this standard is now recognized as an extension of darwin core version 2 (dwc2)14, a profile describing the minimum set of standards for search and retrieval of natural history collections and observation databases. the advantage of having standards established within the mammal community was also reflected in 13 http://elib.cs.berkeley.edu/manis/darwin2conceptinfo030315jrw.htm. 14 http://tsadev.speciesanalyst.net/documentation/ow.asp?darwincorev2. the ease with which the participants were able to agree upon and adopt standards for georeferencing the specimen localities in their collections, i.e., assigning geographic coordinates and maximum error distances for those coordinates to locality descriptions. as stated above, much of the value inherent in specimen data is dependent upon the presence of accurately georeferenced locality information. this is particularly true if the data are to be used for modeling and predictive analyses (e.g., in the fields of conservation biology, ecosystem monitoring, and disease tracking). less than 26% of the roughly 1.4 million mammal specimens housed in participating museum collections contained coordinate data in conjunction with specimen localities at the time of proposal submission. hence, coordinating georeferencing activities was a major focus of activity for the project. at the outset of the manis project, there were no established standards for georeferencing descriptive locality data. the development of georeferencing guidelines15 and tools was critical to the success of the collaborative approach that was adopted, and both have proven useful to other initiatives as well. collaboration georeferencing collaboration, in addition to standards, has also been key to the success of manis. on one hand, it has resulted in cost efficiencies and economies of scale that would not otherwise have been realized. on the other hand, the project has recognized tremendous benefit from having a community of individuals available to address and propose solutions to problems as they arise. both of these factors can be demonstrated most clearly with examination of the collaborative georeferencing activities that lie at the heart of the project. getting participants to recognize the value of collaborative georeferencing did not require much effort. yet, on a specimen level, there was immediate consensus that georeferencing the locality from which every organism in a medium-sized or large collection had been collected was a daunting and seemingly overwhelming proposition. with every institution facing similar data problems (outdated place names, vague or ambiguous localities, specific localities mismatched with higher level geographic attributes) and limited resources, even with nsf 15 http://elib.cs.berkeley.edu/manis/georefguide.html. stein and wieczorek – the manis data network 18 support, duplication of effort had to be precluded. although the primary goal of each institution was to make sure that all of the localities from its own collections were georeferenced at the conclusion of the project, a plan to share the georeferencing workload among participants was readily adopted. perhaps the single most important innovation in the project was the creation of the manis georeferencing gazetteer, which contained all unique specimen collecting localities (296,737) for the 1,367,627 specimens from the seventeen original institutions in the project (figure 1), including the 64,073 localities for which geographic coordinates were already provided. this gazetteer was created by combining locality data from all participating institutions at the outset of the project into a single gazetteer database with an online query interface for browsing and downloading tab-delimited locality records16. the gazetteer was used as the basis for collaborative georeferencing, whereby each 16 http://elib.cs.berkeley.edu/manis/search.shtml. institution reserved geographic areas in which they had particular interest, map resources, or expertise17. having claimed a geographic region for georeferencing, an institution would query the gazetteer for all of the localities from that region, download them, and georeference all of the localities for all of the participating institutions. by having one institution georeference the localities of all institutions for a given region, tremendous economy of scale, uniformity, and use of expertise were achieved. it also meant that each of the participating institutions did not need to possess or acquire geographic resources (e.g., maps, gazetteers, atlases) for the entire world. this proved to be both a cost savings and a practical solution to a complex problem. 17 http://elib.cs.berkeley.edu/manis/checklist.html. figure 1. the distribution of the origin of mammal specimens by country for the seventeen original participating institutions in manis. stein and wieczorek – the manis data network 19 when georeferencing of a geographic region was finished, the data set was uploaded, standardized for consistency of values (datums, units, etc.), checked for completeness, assessed for compliance with the georeferencing standards, and underwent geographic validation to ensure that the coordinates provided were in agreement with the large-scale geography describing the locality. the geographic validation process revealed that ca. 4% of georeferenced localities required further attention. of those requiring further attention, roughly half were found to be correct, but had made it onto the list for additional examination due to inaccuracies in the boundary layer data (e.g., county shape files). the high rate of accuracy achieved using the guidelines and collaborative georeferencing are a testimony to the merit of the approach. looking forward, georeferencing in the manis project has defined a baseline against which to compare the accuracy and speed of automated georeferencing techniques, and from which taxonomic specialists can make further refinements as time and money permit. where digital maps were available for a geographic region, the mean (+1 sd) georeferencing rate was 16.6 (+8.3) localities per hour (n = 14 data sets from 14 institutions). the mean georeferencing rate for regions where printed maps were used instead of digital media was 9.6 (+6.8) localities per hour (n = 39 data sets from four institutions). these rates include the determinations of both coordinates and uncertainties, with full documentation (wieczorek, et al., 2004). the metadata for uncertainties and documentation were found to take roughly 30%, on average, of the time required to georeference a given locality. yet, without these additional data, it would not be readily possible to filter georeferences for suitability for a given line of research. thus, it was deemed essential to provide these additional data. this view was later adopted by the tdwg spatial standards subgroup in its definition of an “acceptably georeferenced specimen record” (frazier, 2002). collaborative georeferencing was initially an act of faith on the part of the participants, faith that time spent georeferencing localities from another institution would be rewarded by similar behavior on the part of one’s colleagues. although unanticipated, this course of action was richly rewarded through the attraction of georeferencing partnerships and collaborations with organizations outside the 17 participating institutions. this produced immediate, tangible benefits for the project, as well as the larger community. shortly after manis activities commenced, the lead programmer was contacted by mexico's national commission for the understanding and use of biodiversity (conabio) with an offer to assist in georeferencing mexican localities, using their own resources, for specimens housed in participating collections. given that the 17 participating institutions housed 172,596 mexican specimens from 30,059 distinct localities, this offer was accepted with alacrity. by adhering to the manis georeferencing guidelines, complete georeferencing of those localities occurred in record time and with unparalleled expertise. in turn, conabio was able to greatly enhance its knowledge of historical distributions for mexican mammals, which enables the organization to address ongoing species conservation and land management issues more accurately now and in the future. creating the network software georeferencing locality data in participant collections was only one of three major aspects of development of the manis network. the other two aspects included creating the network software and connecting the participant institutional databases to the network. though connecting institutions could not be initiated until a functional network was in place, the creation of the network software was another example of a collaborative effort that has benefited both manis and the larger biodiversity informatics community. as noted above, development of a reference implementation of the digir software was, in part, driven and influenced by the requirements of the manis participants. yet, because of its international collaborative origins, the software’s diverse capabilities exceed what manis programmers would have been able to achieve on their own within the limited period of awarded funding. by distributing responsibility for development of various software modules and functions (provider, portal, and network registration) among a diverse group of “biologists cum software developers” (see acknowledgments), the project has been able to avail itself of distributed expertise in excess of funds available to any one interested party. moreover, the collaborative nature of digir development has resulted in widespread adoption of the protocol because its generic character is easily adapted to varying community standards stein and wieczorek – the manis data network 20 and disciplines and does not reflect the imprint of any one group. benefits of collaboration for collaboration to be successful among individuals at institutions with similar high-level goals but diverse circumstances and constraints, it is important to recognize what drives such parties to cooperate. from the perspective of the manis participants, we would argue that a heightened sense of community and recognition of the enormous value of their combined data – “the whole” – has emerged. simultaneously, both large and small collections are now recognized for their respective contributions to that whole. appreciation of what can be gained through collaboration has replaced a sense of competition that is seemingly pervasive in scientific communities. education has also emerged as a valuable product of this collaboration. understanding of technical issues has increased enormously among the participants. with respect to georeferencing, for example, the nature and value of the geodetic datum (the geometric description of a geodetic surface model such as nad27, nad83, and wgs84), errors, and extents has led to a desire among the participants to modify their current database management systems to accommodate these fundamental concepts. this will mean that new specimens added to their collections will have locality data of quality comparable to that which has been achieved in the georeferencing for the manis project. if one path to success is through collaboration, how has this been achieved in this project? foremost was the establishment of trust among the participants with the programmer and with principles involved. manis was extremely fortunate in having a lead programmer who possessed a keen understanding of and experience with mammalian biology, curatorial practices, specimen data, museum databases, and field work. meetings between the project participants and the programmer were held at the annual asm meetings throughout the grant period. those meetings were designed to excite participants about each upcoming phase of the project, demonstrate new tools that would be made available to the project, answer questions face-to-face, and allow participants to interact and develop a rapport with the individual to whom they addressed their technical questions or were forced to reveal the dark secrets of their collection data during the course of the project. it was also made clear to participants from the moment of project conception that institutions would retain control over their data and over the entire project process. accordingly, a listserv was set up to facilitate grant submission. ultimately, this tool proved most valuable in creating a sense of community. by establishing a subscription list rather than a public list, all interested parties could take part in discussions, whether or not they were part of the original proposal. however, funded participants felt comfortable in posting questions about their data to the list, knowing that those who would read their postings probably shared their particular problem, had expertise in their problem area, or at least had a genuine interest in helping to solve problems. because all manis business was conducted openly on the list, and all decisions among participants were reached by consensus, each of the institutions was an equal player, regardless of size or perceived import of their collections. the enthusiasm among the participants was instrumental in making management of the project tractable and should not be underestimated when undertaking a project of this nature. because of the cooperative fulfillment of shared goals, in the final year of the grant period, the project is on target for successful completion and there have been no unexpected or unwanted obstacles. impact as stated in the nsf project proposal, the goals of the manis project were to facilitate open access to combined specimen data from a web browser, enhance the value of specimen collections, conserve curatorial resources, and use a design paradigm that could be easily adopted by other disciplines, all while adopting mainstream transport protocols, avoiding the long-term, external maintenance of a network and centralized data management, and exercising fiscal economy. ironically, remaining focused on those goals has presented the greatest challenge to the project due to the unexpected “ripple effect” that manis has had. given the enormous interest that the project has generated, it is increasingly difficult for participants not to be distracted by the burgeoning number of projects related to development of digir and creation of the stein and wieczorek – the manis data network 21 manis network (herpnet18; obis19; specieslink20; gbif21; ornis; biogeomancer22). from a programmer’s perspective, the temptation to become involved in this flurry of activity is almost irresistible. from an institutional perspective, external pressures to divert limited resources or overextend collections staff in the name of ancillary projects that proffer immediate heightened visibility is an ongoing challenge. a functional manis network was first publicly demonstrated in june 2002 at the asm annual meeting in lake charles, louisiana. later that month, the network was again demonstrated at the global biodiversity informatics facility (gbif) data access and database interoperability scientific and technical advisory group in an invited presentation entitled, “the mammal networked information system (manis), powered by distributed generic information retrieval (digir).” a third presentation of manis was given in the biodiversity and biocomplexity informatics panel at the july 2002 joint conference on digital libraries in portland, oregon. at the 2003 annual meeting of the asm in lubbock, texas, a symposium was held featuring five presentations covering topics relevant to the symposium title “the manis project: the development and applicability of a mammal collection database network.” these demonstrations helped establish standards and allowed for community input throughout the development process. there has been an enthusiastic reception to manis well beyond the scope of its participating institutions. as a reference implementation of digir, manis has been a highly visible pilot project for a variety of international consortia, such as the species analyst (tsa23 peterson et al., 1998), australia’s virtual herbarium24, landcare research of new zealand25, the european natural history specimen information network26, and centro de referência em informação ambiental (cria27) in brazil. among north american distributed database 18 http://herpnet.org/. 19 http://www.iobis.org/. 20 http://splink.cria.org.br/. 21 http://www.gbif.net/portal/index.jsp. 22 http://www.biogeomancer.org/yu/. 23 http://speciesanalyst.net/. 24 http://www.anbg.gov.au/avh.html. 25 http://www.landcareresearch.co.nz/. 26 http://www.bgbm.fu-berlin.de/biodivinf/projects/enhsin/. 27 http://www.cria.org.br/. initiatives, herpnet and the forthcoming effort on the part of the avian community, ornis, have both cited manis as their model and have based their network architectures on digir. in addition, there has been keen interest in the use of digir as an elegant, inexpensive solution to the problem of web data access and database interoperability for local consortia of natural history data providers. among these are the university of california (uc) berkeley natural history museums and the association of biological collections at uc davis. in early february 2004, the global biodiversity information facility (gbif28) announced the public debut of its global digir-based network, including the manis participating institutions among the data providers of nearly 10 million specimen records. the manis project contributed both financially and academically not only to the development of the digir protocol but also to the darwin core version 2 conceptual schema for natural history collections data. the darwin core version 2, an elaboration and reworking of the darwin core developed for the species analyst in 1999, captures the roughly 50 most prevalent biodiversity information concepts in common among natural history collections. this core provides a shared, standard information domain so that biodiversity information can be queried across multiple disciplines while ensuring compatible results. many of the questions posed during the development of the manis network, and many requests for use of the data even before the network was functional, have shed light on new and potential research applications to which these data can be applied. although museum curators have tirelessly and cogently articulated the value of their collections and associated data for decades, creation of the network has given substance to those words in a manner that is immediately accessible and understandable to administrators, colleagues, and the general public as never before. this has led to a renewed and enhanced appreciation of these invaluable resources and bodes well for their future support and the conservation of biodiversity on earth. acknowledgments with support from the national science foundation, the following natural history museums 28 http://www.gbif.org. stein and wieczorek – the manis data network 22 collaborated to create the original manis network: bernice p. bishop museum, california academy of sciences, colección nacional de mamíferos (mexico city), field museum, los angeles county museum of natural history, louisiana state university museum of natural science, michigan state university museum, university of california berkeley museum of vertebrate zoology, royal ontario museum, texas tech university museum, university of alaska museum, university of kansas natural history museum, university of michigan museum of zoology, university of new mexico museum of southwestern biology, university of puget sound slater museum, university of washington burke museum, and the utah museum of natural history. the principle investigators and lead programmer of the project wish to express their sincere gratitude to the curators, staff and the support personnel at these institutions who have worked enthusiastically and collaboratively to create the manis network. we also wish to extend our sincere gratitude to the curators and staff of four additional institutions that have made major contributions to the project’s collaborative georeferencing activities: comisión nacional para el conocimiento y uso de la biodiversidad (conabio); sternberg museum, fort hays state university; university of kansas natural history museum, division of birds; and the university of minnesota bell museum. the manis project is equally indebted to the following individuals and institutions for their programming expertise and contributions to almost all aspects of the project: pj schwartz (digir portal developer), dave vieglais (digir provider developer, university of kansas biodiversity research center), reed beaman (yale university), stan blum (california academy of sciences), renato giovanni (cria), australia national botanical garden (anbg), berkeley digital library project (dlp), biological collection access service for europe (biocase), committee on data for science and technology (codata), centro de referência em informação ambiental (cria, brazil), global biodiversity information facility (gbif), taxonomic databases working group (tdwg), university of kansas biodiversity research center (kubrc). references brook, m. de l. 2000. why museums matter. trends ecol. evol. 15:136-137. frazier, c.k. 2002. definition of the minimal attributes of an acceptably georeferenced specimen record. http://georef.peabody.yale.edu/tdwg/lisbon2003/def georeferencedspecimen_2003-03-25.rtf. graves, g.r. 2000. costs and benefits of web access to museum data. trends ecol. evol. 15:374. hafner, m.s., w.l. gannon, j. salazar-bravo and s.t. alvarez-castañeda. 1997. mammal collections in the western hemisphere: a survey and directory of existing collections. american society of mammalogists. http://www.mammalsociety.org/committees/commsys collection/collsurvey.pdf. mclaren, s.b. 1988. guidelines for usage of computerbased collection data. j. mammal. 69:217-218. mclaren, s.b. 1999. documentation standards for automatic data processing in mammalogy. version 2.0. american society of mammalogists. peterson, a.t., d.a. vieglais, a.g. navarro-sigüenza and m. silva. 2003. a global distributed biodiversity information network: building the world museum. bull. brit. ornithol. club 123a:186-196.. vieglais d.a., e.o. wiley, c.r. robins and a.t. peterson. 2000. harnessing museum resources for the census of marine life: the fishnet project. oceanography 13:10-13. wieczorek, j., q. guo and r.j. hijmans. 2004. the pointradius method for georeferencing locality descriptions and calculating associated uncertainty. int. j. geogr. inf. sci. 18:745-767. wiley, e. and c.r. robins. in prep. fish of the world: fishnet as an example integration of ichthyological data. biodiversity informatics. williams, s.l., m.j. smolen and a.a. brigida. 1979. documentation standards for automatic data processing in mammalogy. museum of texas tech university, lubbock, tx. wilson, d.e. and d.m. reeder. 1993. mammal species of the world: a taxonomic and geographic reference. 2nd edition. association of systematics collections and the smithsonian institution press, washington, d.c. http://www.nmnh.si.edu/msw/. microsoft word segopt-echng-vfmstartlayoutf.doc biodiversity informatics, 6, 2009, pp. 5-17 5 an efficient segmentation algorithm for entity interaction eugene ch’ng1 school of computing and information technology, the university of wolverhampton wulfruna street, wv1 1sb, wolverhampton, united kingdom. abstract.—the inventorying of biological diversity and studies in biocomplexity require the management of large electronic datasets of organisms. while species inventory has adopted structured electronic databases for some time, the computer modelling of the functional interactions between biological entities at all levels of life is still in the stage of development. one of the challenges for this type of modelling is the biotic interactions that occur between large datasets of entities represented as computer algorithms. in real-time simulation that models the biotic interactions of large population datasets, the use of computational processing time could be extensive. one way of increasing the efficiency of such simulation is to partition the landscape so that entities need only traverse its local space for entities that falls within the interaction proximity. this article presents an efficient segmentation algorithm for biotic interactions for research related to the modelling and simulation of biological systems. key words.— artificial life, entity interaction, individual-based model, optimisation algorithms, segmentation. biodiversity research studies the “variation of life at all levels of biological organisation” (gaston and spicer 2004) and is used to refer to “the whole range of activities traditionally connected with inventorying and studying living resources” (lévêque and mounolou 2003). such studies frequently encounter large datasets of biological entities, especially in functional interactions between different levels of organisms and their effects on the modifying and shaping of the environment. as research in biodiversity becomes complex, computer modelling and simulation become a necessary means for conducting experiments so that questions that cannot be answered by traditional approaches may be resolved through the use of technology. the feasibility of computer modelling of nature is amplified in studies in biocompexity, which is defined by lévêque and mounolou as “the result of functional interactions between biological entities, at all levels of organization, and their biological, chemical, physical and social environments” where observational studies of complex entities above the correspondence e-mail: e.chng@wlv.ac.uk species level such as population, biocenosis ecosystems, and biospheres require controlled experiments of their environment. for example, one may wish to project the impacts of future increase in yearly mean temperature on an ecosystem, or to attempt to computationally introduce an invasive species to see its effects on a population. in other cases, one may wish to increase the biological time and processes of a virtual ecosystem in order to predict its outcome in a hundred years’ time. other important needs are to visualise, not by charts and graphs but by utilising interactive computer graphics, the real-time interaction of a biological community in order to observe the ecosystem dynamics. in any case, there is much to be explored in the mergence of computers and biology for biodiversity research. traditional modelling in ecology depended upon statistical and probabilistic procedures such as classification and regression trees, generalised linear models, multivariate adaptive regression splines, and artificial neural networks. according to some comparative studies (cairns 2001, miller and franklin 2002, muñoz and felicísimo 2004), such methods often produce inconsistent results. in ch’ng. – an efficient segmentation algorithm for entity interaction 6 the worst case, they are fraught with errors and inaccuracies as factors considered crucial are often not being taken into account. a number of critiques have recently questioned the validity of modelling strategies for predicting the natural distribution of species. for example, a research review (pearson and dawson 2003) showed that many factors other than climate determine species distributions and the dynamics of distribution changes. while traditional approaches have been focused on the identification of a species’ bioclimatic envelope, crucial factors such as biotic interactions (connell 1961, silander and antonovics 1982, davis et al. 1998) and evolutionary change (woodward 1990, davis and shaw 2001, thomas et al. 2001) are not taken into account and therefore making predictions erroneous and misleading. species dispersal is another factor that has been disregarded, pearson and dawson (pearson and dawson 2003) suggest that migration limitations such as landscape barriers, deforestation, and manmade habitats can obstruct species movement and should be taken into account in such models. as incorporating these factors into the procedures of traditional approaches is difficult, it becomes necessary to explore new approaches for ecological modelling (ch'ng 2009). nature has intrinsic problem-solving principles and nature-inspired design seems a potential area for discovering new methodology for research. there may not be a better way to model biological life than to simply mimic its designs by using the algorithmic format to simulate the various levels of complexities (molecular, species, ecosystem, and biosphere). this synthesis of nature formulates algorithms by extracting nature’s principle rules in order that biological organisms can be replicated within the confines of voltage and silicon. indeed, the aim of the experimental science of artificial life (langton 1995) attempts to understand the mechanisms behind living systems by synthesising them so that the structure and functions of these systems may be understood and applied. one of its founders (langton 1990) states that “by extending the horizons of empirical research in biology beyond the territory currently circumscribed by life-as-we-know-it, the study of artificial life gives us access to the domain of life-as-it-could-be, and it is within this vastly larger domain that we must ground general theories of biology and in which we will discover novel and practical applications of biology in our engineering endeavors.” most, if not all of artificial life-based modelling techniques are individual-based models (ibms) and an extension of it called evolutionary individual-based model (eibm) (bornhofen and lattaud 2006), or agent-based models. breckling et al. stated that ‘ibms contrasts with common ecological models which frequently operate on the population level and represent a population as an overall state, thereby specifying rules how the overall state changes (breckling et al. 2005). ibms was identified in a visionary article by huston et al. (huston et al. 1988) as a potential model for simulating the effects of individual variation, spatial processes, cumulative stress, and natural complexities that are difficult with classical approaches: “individual-based models allow ecological modellers to investigate types of questions that have been difficult or impossible to address using the state-variable approach.” individual-based models have been thoroughly reviewed up to 1999 (grimm 1999) and 2005 (grimm and railsback 2005) followed by a recent review of eibm (bornhofen and lattaud 2006). another review of the approach have shown that individual-based models can also benefit from concepts in complex adaptive systems (railsback 2001). these novel modelling techniques could potentially resolve issues outlined earlier, such as biotic interaction, evolutionary change, and species dispersal barriers. what distinguishes ibms from other “individual-oriented” models that acknowledge the individual level in some way but still adhere mainly to the classical modelling paradigm? there are four criterion: (1) the degree to which the complexity of the individual's life cycle is reflected in the model; (2) whether or not the dynamics of resources used by individuals are explicitly represented; (3)whether real or integer numbers are used to represent the size of a population; and (4) the extent to which variability among individuals of the same age is considered (uchmanski j and v. 1996, grimm and railsback 2005). since the emergence phenomenon observed in ibms reflects that of ecological systems, three properties characterises them (breckling et al. 2005): (1) they do not exist on the level of isolated subsystems; (2) they emerge on higher levels as a result of interactions of the subsystems; and (3) new properties appear at one level of a system and ch’ng. – an efficient segmentation algorithm for entity interaction 7 are not deducible from the observation of the lower levels units or compartments of the system. this requires the modelling of individuals rather than at the population level. such techniques require the modelling of individual biological life, including their genotype, phenotype and dynamics. this implies the requirements for large computing resources as hundreds of thousands of calculations are performed for each individual as they react and interact with their biotic and abiotic environment. if optimisation problems associated with these models are properly developed, new opportunities for resource conservation planning, habitat and biodiversity management, mitigation strategies and studies related to the understanding of demographic problems can be initiated. this research attempts to solve a critical problem associated with computational resources that is connected to artificial life, ibms or agent-based modelling of biological systems at the species level and the simulation of biotic interactions among large datasets of spatially-explicit sessile organisms and a suggested extension for vagile life forms. the following sections explore optimisation techniques and formulate logical methods for optimising entity interactions. the paper concludes with a discussion of the results and opportunities for future work. optimisation techniques unlike statistical methods, which uses pixels as patch size of n meter of space to describe large communities of organisms on a landscape, new approaches require the modelling of individual organisms and all its internal processes and external interactions. experience tells us that the modelling of biological life using computer algorithms requires large numbers of variables, algorithmic structures, and calculations. this is true even for a single entity. the availability of computing resources for simulation becomes an exponential challenge when biotic and abiotic interactions occur and entities reproduce. resource bottleneck is a problem that requires solving before new modelling methods can see fruition. unfortunately, a review of literature in the area yielded little evidence of such studies. schulz and reggia developed a method for predicting nearest agent distances in artificial worlds (schulz and reggia 2002). other methods are developed for representing complex outdoor scenes in computer graphics (snyder and barr 1987, deussen et al. 1998, soler et al. 2003) but algorithms have not been formalised for improving the efficiency of large entity interactions. data structures for managing and categorising large collections of data are available. the simplest perhaps is the array data structure. more advanced structures are hierarchical data structures, which is based on recursive decomposition (aho et al. 1974). well known hierarchical data structures (samet 1984) used in the science and engineering for efficient representation and improving execution times are binary trees, quadtrees, and octrees. a tree in computer science terminology emulates a tree structure. it has sets of link nodes starting from a root node and proceeds through the child nodes (branch) and finally to the leaf nodes (final nodes). a child node has one parent node and a parent node may have zero to many nodes. a binary tree contains only two child nodes and is not suitable for partitioning terrains into equal parts (where the ratio is 1). quadtree and octrees contain only four and eight nodes respectively. as we shall soon see in the discussions, hierarchical data structures presents some problems for managing large indexes that have frequent changes. the approach presented in this article uses a data structure that is non-hierarchical. it divides terrains based on the square root of the number of segments, s . the approach is flexible and is able to segment a landscape into an infinite number of equal parts. the next section explores some concepts related to the efficiency of segmentation logic. exploring the logic of segmentation methodology the segmentation algorithm presented here is an optimisation technique developed and applied for investigating a new approach of vegetation modelling for studying palaeoenvironments (ch'ng and stone 2006b, a, ch'ng 2009). the optimisation technique targets non cell-based ibms and agentbased models. in other words, it does not apply to discrete environments such as cellular automata (ca), e.g., (gutowitz 1991, ginot et al. 2002). the ibms must be spatially explicit and in a continuous environment. the optimisation technique targets sessile organisms (e.g., plants) but can be easily extended to include vagile life forms (e.g., fish, ch’ng. – an efficient segmentation algorithm for entity interaction 8 animals and insects) and is suggested in the discussions section. the following paragraphs explore segmentation logics. in the ecosystem model, there are bound to be large numbers of entity interactions, both accessing and competing for resources. the efficiency of an optimisation technique will determine the computational speed of the simulation. in an algorithm, the entities are usually stored in a single array, in more advanced simulation requiring constant insertion and deletion (births and deaths) entities are stored within a collection data structure. in the simulation cycle, each entity accesses every other entity in the same collection to determine their proximities for interaction or competition. if entities are at proximity, interaction occurs. therefore, if the collection contains ten simple two state entities, every entity will have to go through ten loops including itself for determining the proximity of other entities before computing the interaction. this amounts to 210 , and if there are 1000 entities, the amount increases exponentially at 000,000,110002 = . in large landscapes and actual models, vegetations could amount to thousands to hundreds of thousands of plants with computable variables, algorithms and calculations in each entity. this problem gives reason for developing optimisation techniques. in the initial stages of research, three theoretical concepts were compared for their efficiencies. memory indices (mi) the first concept requires that each entity remembers the indices of nearby entities (figure 1). in this technique, each entity accesses only the indices in its memory for interaction and competition. the technique however, has serious limitations and is restricted only to sessile ibms. we know that vegetation reproduces abundantly during the spring-summer seasons each year. in the simulation, when new plants are reproduced plant proximities in the entire collection will have to be traversed and computed to determine which plants should be in the memory of other plants. and each time a single plant dies the collection has to be traversed again to remove the index of the dead plant from the memory of nearby plants. this becomes very slow as the number of plants increases. figure 1. inefficient memory indices (mi). collection class indices (cci) the second concept segments the landscape into different collection classes storing the indices of entities in each class (figure 2). this is a better algorithm compared to the first since the segments in a terrain is fixed provided the landscape is not continually being re-segmented during the simulation. however, different collection classes increases memory and the need to traverse each collection during simulation wastes computational time. furthermore, in each reproductive lifecycle, each collection has to be checked against the new plants to determine which segment boundary it belongs to. cci differs from cellular automata but is very similar to the particle in cell (pic) method in (bithell and macmillan 2007). segmentation algorithm the third technique is more efficient. in a landscape, a plant can only compete with adjacent plants at a given time and all other plants should be discounted from the interaction. in theory, the simulation time should decrease with the increase of landscape segments. we shall soon see the performance of the technique. figure 3a shows a landscape with groups of plants where only intersecting plants are being competed against. in the real world, computation is not a problem since interaction between entities is parallel and occurs simultaneously. however, since every virtual plant is stored within an array in a computer program, traversal is required to ch’ng. – an efficient segmentation algorithm for entity interaction 9 figure 2. inefficient collection class indices (cci). determine which plant is ‘visible’ for competition. in a simulation, the ideal competition scenario for segmentation in figure 3a can be divided into 4 segments where each plant needs only access the other plants in its own segment space (shown in b). in a difficult segmentation condition at c where the source plant (large circle) is near to a boundary or overlaps the boundary of a segment space, the source plant is required to access all other plants in neighbouring segment spaces for thorough interaction. in this case, the benefits of segmenting the landscape are not apparent as it is the same as condition a since plants within the entire collection is accessed. the benefits become apparent when the segments are increased (in d). the figure illustrates that a higher segmentation increases the speed of the program since each plant needs only access the plants within its segment space and adjacent segment spaces. the only rule needed is that the size of a segment should not be smaller than the canopy of a tree with the largest diameter. segmenting a landscape in a standard directx or opengl 3d coordinates is different from the 2d screen coordinate systems. the method developed in this research targets the directx 3d coordinate space but can be easily extended to include opengl and 2d coordinate spaces. segmentation begins with divisions. the number of divisions can be from 1 to ∞ subject to the limits of computer memory. the segments are derived from the square of the divisions with 2d d s= = , where d is the number of divisions and s is the number of segments. figure 3a contains 1 division with 1 segment, 21 1 1= = . the landscape in b and c both contain 2 divisions with 4 segments, 22 2 4= = . figure 3d contains 3 divisions with 9 segments 23 3 9= = . the number of divisions has equal width and height. the value of the width and height of the landscape can be obtained with, terrain divs terrain divs w w s h h s = = (1) where divsw and divsh are respectively the number of divisions in the width and height of the terrain. terrainw and terrainh are the size of the width and height of the terrain. s is the number of divisions which divides the terrain into divsw and divsh . constructing the segments requires an understanding of the 3d coordinate system, shown in figure 4a. the origin of the axis lies at the centre of the plane with extensions of the axis in both the negative and positive directions. the segmentation however, cannot begin from the origin but is offset in the negative (-x,-y) direction starting from the ‘start’ position show in b. this means that segment 0 begins from the lower left corner (-x, -y) and ends at the top right corner (+x, +y) at segment 15. the units (-50, -25, 0, +25, +50) are given as an example and can be replaced with any other units with equal divisions. based on the divisions of width at b, it is observed that when j=0, x=-50, when j=1, x=25, when j=2, x=0, and so on. this generates a graph at c where an equation of the line is given (eq. 2), ( ) 2 terrain terrainw w f j j s = − (2) ch’ng. – an efficient segmentation algorithm for entity interaction 10 figure 3. efficient segmentation logic. a landscape with 1 segment (a), a landscape with 4 segments (b), entity accessing other segments for interaction (c), and entity interaction in a 9 segments scenario (d). where j is the segment junction, terrainw is the width of the terrain, and s is the number of divisions. the equation offsets the first segment (0) from the origin to the ‘start’ position. a construction of the height divisions uses the same equation replacing terrainw with terrainh . in the algorithm, each segment is a rectangle object storing its size and position. the algorithm in standard format is given in listing 1. divs is the number of divisions, i and j are each division’s index, wdiv and hdiv are the width and length of each segment divided from terrainw s and terrainh s . every plant in its own segment space accesses its adjacent segment spaces. figure 5a shows the index accessing pattern using a one dimensional array. t is the target segment where a plant resides. the plant accesses its adjacent segment spaces for competition. the segmentation in b shows black coloured segments within the safe frame (red). in c, the segment within the safe frame (blue arrows) safely accesses segments that exist whereas segments outside the safe frame (red arrows) accesses non-existent segments and will generate errors in the program. this problem can be solved by preventing the left edge and corner segments from accessing certain non-existing segments. figure 6 shows some instances of the patterns of non-existent segments where the target segment should not have access to. the pattern illustrates that depending on where the edge segment (orange) is located (top, bottom, left, right), it should not access not less or more than three non-existent segments. the corner segments should not access not less or more than 5 non-existent segments. the algorithm for segregating the edge and corner segments is given in listing 2. pdivisions is the number of divisions on the landscape, pnumofsegment is the number of segments on the landscape. the rest are selfexplanatory. a uml diagram showing the class relationship between the segment and vegetation is given in figure 7. every plant has a unique vegetationid. within the plant, segmentid defines the segment the plant is in and has a default value of -1 when it is in the dormant stage. it is assigned when the seed is germinated as that is the beginning of competition. the segment object contains a unique segmentid and an array of plantindices[ ] which stores the index of each plant within the boundary of the segment rectangle. the index of a plant can only be in the plantindices[ ] array of one segment at any given time. the pseudo code in listing 3 shows how entities are managed during simulation time. optimisation results results from the segmentation algorithm covered in the previous section were recorded on a 150m² terrain with five species of trees. due to the nature of the simulation (large number and fluctuations of population), it is difficult to generate a comparison chart of the increase of the segments (1, 2, …, n) and for each case, the number of organisms and the simulation speed. therefore a chart has to be shown for each individual segmentation simulation. ch’ng. – an efficient segmentation algorithm for entity interaction 11 figure 4. segmentation constructions in 3d coordinate space (a, b, and c). figure 5. the concept of safe frame in segmentation algorithm: index accessing pattern using one dimensional array (a), safe frame with black numbers (b), and frames accessing non-existent segments (c). at the start of the simulation, the initial number of trees is 147 distributed among the species. three segmentation experiments were conducted to compare the results using the same settings where vegetation reproduces up to a maximum of 162 on one of the experiment. the first segmentation test uses 1 segment for the landscape with a total of 162 at peak production over 258 years, the second and third partitioned the terrain into 64 and 144 segments respectively. the 64 segments version produces 147 plants at its peak over 263 years and the 144 segments version produces up to 151 during its peak over 258 years. even though the settings are the same for all three experiments, the number of vegetation produced is different. this is due to the interaction among vegetation and the environmental variations affecting them. this however, does not significantly affect the results as comparison is made between the speed and the number of plants produced (figure 8-11). a comparison of the three graphs showed that there is significant increase in speed if the segmentation algorithm is applied. in figure 9, when the number of vegetation reaches its maximum at 140, the speed reaches only 180 milliseconds using 64 segments as compared to an average of 400 milliseconds using only 1 segment (figure 8). if the segments are increased to 144, the speed decreased to only an average of 100 milliseconds with the same amount of vegetation (figure 10). figure 11 is a comparison of efficiency of segmentation between the three experiments. the graph showed that the computational speed increases significantly between 1 and 64 segments. however, the increase in performance is not apparent between 64 and 144 segments. this is due to the small number of vegetation present in the landscape. the performance becomes obvious when the number of vegetation is greatly increased near the end of the graph. figure 12-13 are graphs ch’ng. – an efficient segmentation algorithm for entity interaction 12 listing 1: 1. int n = 0; // segment number 2. for (int j = 0; j < divs; j++) // rows (divided on length) 3. { 4. for (int i = 0; i < divs; i++) // columns (divided on width) 5. { 6. x = (terrainwidth/divs)*i terrainwidth/2; // set the x position (columns) 7. y = (terrainheight/divs)*j terrainheight/2; // set the y position (rows) 8. 9. // create a new segment 10. segment[n] = new segment(x*2, y*2, wdiv*2, hdiv*2); 11. n += 1; // next segment 12. } 13.} listing 2: 1. // collect segment corners 2. pbottomleftcorner = 0; 3. pbottomrightcorner = pdivisions-1; 4. ptopleftcorner = pdivisions * (pdivisions-1); 5. ptoprightcorner = pnumofsegment-1; 6. 7. // collect segment edges 8. pleftedge = new int[pdivisions-2]; 9. prightedge = new int[pdivisions-2]; 10. 11.for (int i=0 ; i < pdivisions-2 ; i++) 12.{ 13. pleftedge[i] = (i+1) * pdivisions; 14. prightedge[i]= (pdivisions-1) + pdivisions * (i+1); 15.} listing 3: initialise simulation: generate segments test all entities if entity is within segment assign vegetationid to segment plantindices based on its location assign segmentid to each entity simulation cycle: remove dead entities and assign new entities to segment simulate entity interaction within proximity segments listing 4: initialise simulation: generate segments test all entities if entity is within segment assign vegetationid to segment plantindices based on its location assign segmentid to each entity simulation cycle: remove dead entities and assign new entities to segment assign vagile entities to segment and assign segmentid to vagile entities simulate entity interaction within proximity segments simulate interaction for vagile entities (viewable entities within segment) ch’ng. – an efficient segmentation algorithm for entity interaction 13 comparing segmentation results of an environment using the same landscape. vegetation is a mixture of trees and herbaceous plants totalling 203 in number at the start of the simulation. the herbaceous plants are used because the rate of reproduction is much faster than the trees. the 64 segments version produces 14,795 plants at its peak over 23 years and the 144 segments version figure 6. accessing non-existent segments. produces up to 12,393 during its peak over 33 years. the comparison shows a significant increase in performance using 144 segments. when the number of plants increases to around 12,000 (figure 13), the speed decreases to only 300,000 milliseconds as compared to 64 segments at 700,000 milliseconds. discussion the implementation of an efficient segmentation algorithm will greatly reduce simulation time in studies dealing with large realtime datasets, such as the modelling of biocompexity where biotic interaction plays a major role in the dynamics of the system. the results in the study confirm that the use of the segmentation algorithm to divide the landscape and the devising of a strategy for managing entity interaction greatly speeds up the simulation, with an increase of >40% in all cases. the finding suggests that increasing the segmentation decreases the speed of simulation in all studies. in the first test, (figure 8-10), three scenarios with segments 1, 64, and 144 were compared. the number of plants in all three scenarios averaged 153. the results showed that the speed of simulation doubles from 400ms to 180ms if 64 segments were applied. if 144 segments were applied, the speed decreases from 400ms to 100 ms. tests conducted in a larger landscape with thousands of plants averaging 13,594 also showed a significant increase in simulation speed (figure 12-13). when plants in the landscape reached a population of 12,000, the 144 segments yielded 300k milliseconds as compared to 700k milliseconds using the 64 segments version. the experiments were conducted on sessile entities. but an additional cycle in the simulation will enable highly vagile entities to be included in the system. the bold lines in listing 4 show the extension. vagile entities may require a separate collection class from the static entities and the interaction should be after the static entities have been processed. static objects are at priority. for example, the seeds of plants should be produced first before predators have access to them, trees should first appear before habitats are made available, and etc. hierarchical data structures such as quadtrees and octrees may have potential solution for such problems. but they will need to be tested. hierarchical data structures divide a space into equal parts. quadtrees divide a space or the subset of a space into four equal parts that are stored as nodes while octrees divide a space into eight equal parts. apart from that, the two structures are essentially the same and achieve the same objective. these structures are found to be very efficient in a space that contains entities that do not go through large and frequent changes such as thousands of addition, deletion, interaction and movement in one simulation cycle. they are frequently used in computer graphics and visualisation. when there are large changes, it is predicted that the restructuring and traversing of the nodes will increase the simulation time. the simulation cycle in listing 3 and 4 show a simple breadth traversal of the segments for both sessile and vagile entities. hierarchical data structures require additional steps in each cycle to rebuild the nodes when entities die and new entities are produced. furthermore, entities need to traverse both the breadth and depth of the nodes in order to determine their interaction proximity. in theory, such structures should not have significant increase in speed, but tests should be carried out in future work. there is much to be explored as far as efficient algorithms for entity interaction is concerned. the algorithms have produced significant results for potential interaction among ch’ng. – an efficient segmentation algorithm for entity interaction 14 figure 7. a uml diagram illustrating the relationship between the class vegetation and segment . figure 8. experimental results using 1 segment on the landscape. the image shows the speed of simulation (the graph with high fluctuations at bottom) and the number of plants (the graph on the top). figure 9. experimental results using 64 segments on the landscape. figure 10. experimental results using 144 segments on the landscape. ch’ng. – an efficient segmentation algorithm for entity interaction 15 figure 11. a comparison of the efficiency of segmentations between the experiments. solid lines show the average speed of simulation and dashed lines of a paler shade show its corresponding plant population. figure 12. the comparison of efficiency of segmentation between experiments (64 segments). figure 13. the comparison of efficiency of segmentation between experiments (144 segments). ch’ng. – an efficient segmentation algorithm for entity interaction 16 thousands of entities. when vagile life forms are introduced, more efficient techniques will be required. in future work, high performance computing with hundreds of nodes of computer clusters is essential to support the calculations involved in biotic and abiotic interactions. research is underway to make this a reality. references aho, a. v., j. e. hopcroft, and j. d. ullman. 1974. the design and analysis of computer algorithms. addison-wesley, reading, mass. bithell, m. and w. d. macmillan. 2007. escape from the cell: spatially explicit modelling with and without grids. ecological modelling 200:59-78. bornhofen, s. and c. lattaud. 2006. outlines of artificial life : a brief history of evolutionary individual based models. lecture notes in computer science, springer-verlag 3871:226-237. breckling, b., f. müller, h. reuter, f. hölker, and o. fränzle. 2005. emergent properties in individualbased ecological models – introducing case studies in an ecosystem research context. ecological modelling 186:376-388. cairns, d. m. 2001. a comparison of methods for predicting vegetation. plant ecology 156:3-18. ch'ng, e. 2009. an artificial life-based vegetation modelling approach for biodiversity research. in r. chiong, editor. to appear in nature-inspired informatics for intelligent applications and knowledge discovery: implications in business, science and engineering. igi global, hershey, pa. ch'ng, e. and r. j. stone. 2006a. 3d archaeological reconstruction and visualization: an artificial life model for determining vegetation dispersal patterns in ancient landscapes. pp. 112-118.in computer graphics, imaging and visualization (cgiv). ieee computer society, sydney, australia. ch'ng, e. and r. j. stone. 2006b. enhancing virtual reality with artificial life: reconstructing a flooded european mesolithic landscape. presence: teleoperators and virtual environments 15:341352. connell, j. h. 1961. the influence of interspecific competition and other factors on the distribution of the barnacle chthamalus stellatus. ecology 42:710723. davis, a. j., l. s. jenkinson, j. h. lawton, b. shorrocks, and s. wood. 1998. making mistakes when predicting shifts in species range in response to global warming. nature 391:783-786. davis, m. b. and r. g. shaw. 2001. range shifts and adaptive responses to quaternary climate change. science 292:673-679. deussen, o., p. hanrahan, b. lintermann, r. mech, m. pharr, and p. prusinkiewicz. 1998. realistic modeling and rendering of plant ecosystems.in proceedings of siggraph '98 annual conference series 1998. acm press. gaston, k. j. and j. i. spicer. 2004. biodiversity: an introduction. blackwell publishing (2nd ed.), oxford. ginot, v., c. le page, and s. souissi. 2002. a multiagents architecture to enhance end-user individualbased modelling. ecological modelling 157:23-41. grimm, v. 1999. ten years of individual-based modeling in ecology: what have we learned and what could we learn in the future? ecological modelling 56:221-224. grimm, v. and s. f. railsback. 2005. individual-based modeling and ecology. princeton university press, princeton, new jersey. gutowitz, h. 1991. cellular automata: theory and experiment. physica d 45:1-3. huston, m., d. deangelis, and w. post. 1988. new computer models unify ecological theory. bioscience 38:682-692. langton, c. g. 1990. artificial life. addison-wesley longman publishing co., inc., boston, ma. langton, c. g., editor. 1995. artificial life: an overview. mit press, cambridge. lévêque, c. and j. c. mounolou. 2003. biodiversity. john wiley & sons, ltd., chichester, west sussex. miller, j. and j. franklin. 2002. modeling the distribution of four vegetation alliances using generalized linear models and classification trees with spatial dependence. ecological modelling 157:227-247. muñoz, j. and á. m. b. felicísimo. 2004. comparison of statistical methods commonly used in predictive modelling. journal of vegetation science 15:285292. pearson, r. g. and t. p. dawson. 2003. predicting the impacts of climate change on the distribution of species: are bioclimate envelope models useful? global ecology & biogeography 12:361-371. railsback, s. f. 2001. concepts from complex adaptive systems as a framework for individual-based modelling. ecological modelling 139:47-62. samet, h. 1984. the quadtree and related hierarchical data structures. computer surveys, acm 16. schulz, r. and j. a. reggia. 2002. predicting nearest agent distances in artificial worlds. artificial life 8:247 264. silander, j. a. and j. antonovics. 1982. analysis of interspecific interactions in a coastal plant ch’ng. – an efficient segmentation algorithm for entity interaction 17 community a perturbation approach. nature 298:557-560. snyder, j. m. and a. h. barr. 1987. ray tracing complex models containing surface tessellations. pages 1-10 siggraph '87 proceedings. acm. soler, c., f. x. xsillion, f. blaise, and p. dereffye. 2003. an efficient instantiation algorithm for simulating radiant energy transfer in plant models. acm transactions on graphics 11:204233. thomas, c. d., e. j. bodsworth, r. j. wilson, a. d. simmons, z. g. davies, m. musche, and l. conradt. 2001. ecological and evolutionary processes at expanding range margins. nature 411:577-581. uchmanski j and g. v. 1996. individual-based modelling in ecology: what makes the difference? trends in ecology & evolution 11:437-441. woodward, f. i. 1990. the impact of low temperatures in controlling the geographical distribution of plants. philosophical transactions of the royal society of london, series b, biological sciences 326:585-593. microsoft word gn_gis_2005_20050805.doc biodiversity informatics, 2, 2005, pp. 56-69 56 challenges building online gis services to support global biodiversity mapping and analysis: lessons from the mountain and plains database and informatics project robert guralnick1,2 and david neufeld1 1university of colorado museum, 2department of ecology and evolutionary biology, university of colorado, boulder, co 80309 usa abstract.—we argue that distributed mapping and analysis of biodiversity information becoming available on global distributed networks is a lynchpin activity linking together research and development challenges in biodiversity informatics. online mapping is key because it allows users to explore the spatial context of biodiversity information visually and assemble quickly the datasets needed to ask and answer biodiversity research and management questions. we make the case that free, online, global biodiversity mapping tools utilizing distributed species’ occurrence records are now within reach, and discuss how such a system can be built using existing technology. we also discuss additional technological and sociological challenges and solutions, given experiences building a regional distributed gis tool called mapstedi (mountain and plains spatio-temporal database and informatics initiative). we focus on solutions to 3 technology challenges: returning result queries in a reasonable amount of time given network limitations; accessing multiple, heterogeneous data sources using different transmission mechanisms; and scaling from a solution for a handful of data providers to hundreds or thousands of providers. we also discuss future challenges and potential solutions for integrating analysis tools into online mapping applications. we close with a discussion of sociological impediments and potential community solutions for biodiversity mapping endeavors. key words.—heterogeneous geospatial data, global biodiversity mapping, mapstedi, online gis, distribution modeling, species’ occurrence datasets. biologists are increasingly approaching a set of very complex questions related to current, past, and future geographic distribution of genes, organisms and ecosystems from a spatial ecological and evolutionary perspective (crisci, 2001). to approach these questions, researchers need tools to assist in acquiring and synthesizing biodiversity and environmental data. the biodiversity informatics community has realized the importance of making data available, and the amount and diversity of datasets continues to increase and become more readily accessible over the internet. as these data have become available, it has become clear that an equally crucial challenge is that of building tools for data synthesis and analysis. ultimately, such work achieves a major biodiversity informatics goal: a standards-based global computing infrastructure to allow rapid, real-time discovery, access, visualization, interpretation, and analysis of biodiversity information (wilson 1992; brisby 2000; krishtalka and humphrey 2000; sugden and pennisi 2000). meeting the challenge is particularly crucial given the accumulating evidence of accelerating biodiversity and habitat loss1 caused by human impacts on the environment. much of the current work in the field of biodiversity informatics is geared towards overcoming one of the largest problems the biodiversity community faces – access to vast quantities of baseline biodiversity data locked away from the broader research and management communities. for example, natural history museums worldwide contain >3 x 109 records of life--mostly plants and animals--representing one of the best sources of information on past and present biodiversity (brisby 2000; krishtalka and humphrey 2000; suarez & tsutsui 2004). however, much of the information on species diversity and distributions has neither been digitized nor georeferenced, exists in different data formats, or cannot be aggregated or integrated with existing computer applications. for example, it is estimated that <5% of species’ occurrences are digitized, and far fewer have computer-readable geospatial coordinates associated with them (beaman and conn 2003). 1 http://www.biodiv.org/gbo/gbo-pdf.asp. guralnick and neufeld – distributed mapping solutions for biodiversity analysis 57 in the correct format, data from natural history museums and local, regional and continental surveys could provide immensely valuable new data resources to the broader biodiversity community. towards this end, projects such as the mammal networked information system (manis2, stein and wieczorek, 2004), herpnet3, ornis4, and the global biodiversity information facility (gbif) biodiversity data portal5 are providing access to large, distributed species’ occurrence datasets. for example, gbif, as of may 2005, provides access >70 x 106 specimen records through its data portal. the ultimate goal is to provide standardized, high quality, and easily usable data back to the community. although these taxonomically focused efforts are critical for repurposing natural history data for biodiversity analyses, there are some limitations to broader use by the diverse community of scientists and managers who could benefit from the data. one impediment is that users must first accumulate the data on a local machine, in most cases taxon by taxon, and then convert the data into more usable formats for further analysis. taxonomic foci also limit the likelihood that land resource managers and conservation planners, who are typically more interested in particular areas, will adopt these systems. at the same time, several ongoing development projects aim to build tools to facilitate answering research questions in environmental biology. a general overview of available approaches and early tools is available from stockwell6 (stockwell 1997). graham et al. (2004) have recently published on the subject of how biodiversity informatics tools are being applied to answer research questions. two areas of interest in the community have been ecological niche modeling (overviews in peterson 2001; soberón and peterson 2004) and species richness and abundances estimates (soberón et al. 2000; ponder et al. 2001; rahbek and graves 2001; petersen et al. 2003; meier and dikow 2004; guralnick and van cleve 2005). ecological niche modeling and species richness estimations both require coupling biodiversity 2 http://elib.cs.berkeley.edu/manis/. 3 http://www.herpnet.org/. 4 http://ornisnet.org/. 5 http://www.gbif.net/portal/index.jsp. 6 http://biodi.sdsc.edu/doc/bis/overview.html. information--named species occurrences--with geographic and potentially environmental data, geographic information systems (gis) and statistical approaches. concurrent with methodological development has been production of a set of desktop and web-based tools. desktopgarp7, web services like whywhere8, and hybrids currently under development like openmodeller9 are freely available applications that allow users to perform their own niche model experiments. although species richness estimations have yet to be built into web-based applications, desktop tools like estimates10, which provides biodiversity summary estimates, and diva-gis11, which provides a gis-based set of biodiversity analysis tools, are both already available. these tools provide the logic and some of the geographic data for performing analyses, but do not provide a means to accumulate and explore effortlessly existing, up-to-date, georeferenced biodiversity information available from computer networks. at present, the communities of biodiversity informatics developers and users are working from two ends that will ideally converge in the middle. at one end are spatial research endeavors like ecological niche modeling and species richness estimations, and applications like desktopgarp, that are built to help answer research or management questions. on the other end is infrastructure development to share biodiversity information, especially species’ occurrence datasets, over the internet. we argue that the obvious bridge between infrastructure and research for the biodiversity informatics field is a global online distributed biodiversity mapping application. such an application would utilize georeferenced data from multiple global distributed species’ occurrence databases using community data standards and mechanisms, and provide functionalities for exploring, exporting, and analyzing the data (all discussed below in more detail). this mapping tool would remove many impediments that currently limit the utility of species occurrence data to the research community by allowing workers to perform spatial or text 7 http://lifemapper.org/desktopgarp. 8 http://biodi.sdsc.edu/ww_home.html. 9 http://sourceforge.net/projects/openmodeller/. 10 http://viceroy.eeb.uconn.edu/estimates. 11 http://www.diva-gis.org/. guralnick and neufeld – distributed mapping solutions for biodiversity analysis 58 figure 1. steps necessary to have natural history and survey collections data automatically displayable and analyzable in online gis web services. in each case, data move from one structured format into another, represented by boxes during data translation and transmission. biodiversity informatics, 2, 2005, pp. 56-69 59 based searches for the most up-to-date biodiversity information available. a user can delimit the taxonomic and spatial scale of the question that she wants to ask, and either run appropriate tests online or export the data for later analysis. we believe that online mapping simplifies the process of data exploration and ultimately lowers the cost barrier to analyses that have not yet been attempted, leading to potential novel research findings and wider use of the data by land managers and conservation planners. building online gis web services to support biodiversity mapping is a major development challenge, and many steps are still in the process of being worked out now, or will need to be addressed in the future. we believe the following questions are the crucial ones that need to be answered in order to move forward on such an endeavor: 1. what methodology and tools should be used to georeference the data? 2. how should data access and transmission be addressed, so that an online, global-scale gis can access appropriate biodiversity data? 3. how can a system be built that is efficient and fast enough for users to sort through the large amounts of data potentially available? 4. how can a distributed gis be built that can handle heterogeneous data sources? 5. how can such a system present attribute data effectively, both in text form and on maps, with potentially billions of data points and thousands of repositories? 6. how can analysis tools be built into a web mapping application so that users can perform as many tasks as possible online and easily export datasets out of the online applications for further use on their desktops? 7. how can the community overcome potential sociological barriers to build such a tool most effectively? below, we organize our discussion of these questions in order of tasks that need to be completed (summarized graphically in figure 1). we first discuss the steps necessary to prepare data, including retrospectively georeferencing species records without explicit geospatial coordinates and ensuring that converted records are compliant with community data standards like darwincore2, access to biological collections data (abcd), and tapir (all discussed in more detail below). we then discuss the challenge of building online mapping applications that can access those distributed georeferenced records, and display them along with other heterogeneous data sources. lastly, we discuss ways to link web-based analysis tools to online mapping applications in order to process biodiversity and environmental data and return summary spatial data. we base much of our discussion on our experiences developing an online biodiversity mapping application (mapstedi; the mountain and plains spatio-temporal database and informatics project12) at the university of colorado museum. we believe it is appropriate to use lessons learned developing the regional mapstedi application to extrapolate towards a fully integrated global biodiversity mapping application, given that mapstedi was built to scale to larger endeavors. data preparation needed prior to biodiversity mapping what methodology and tools should be used to georeference the data? a major challenge for the natural history community has been to establish standards for the process of georeferencing collections data. of the roughly 3 x 109 specimens stored in the world’s natural history museums, <5% have been digitally catalogued or georeferenced (beaman and conn, 2003). fortunately, progress is being made through an increasing number of projects that rely on manual and semi-automated techniques to assign geospatial coordinates to collection records based on the locality descriptions stored with the record; a process known as retrospective georeferencing (murphey et al. 2004, wieczorek et al. 2004). current and past manual georeferencing projects include mapstedi, manis, and inram13. collaborations among biodiversity informaticians are leading to georeferencing protocols with standardized methods for determining both the spatial coordinates for a location and the error and uncertainty regarding the assigned points. one of the most complete guides to georeferencing natural history museum specimen occurrence data is available online from the manis project14. this guide establishes a standard methodology with which to assign geospatial coordinates to historical locality descriptions which 12 http://www.mapstedi.org. 13 http://www.inram.org/. 14 http://elib.cs.berkeley.edu/manis/georefguide.html. guralnick and neufeld – distributed mapping solutions for biodiversity analysis 60 figure 2. a client requests a map containing data local to the server, and from two remote distributed gis servers. image fusion takes the returned images from both remote servers, fuses them with the image on the local server, and returns the composite image to the client. mapstedi.dmns.org www.geomuse.org client request wms arcxml jai – image fusion response terraserver biodiversity informatics, 2, 2005, pp. 56-69 61 oftentimes present unique challenges. these challenges include references in locality descriptions to place names which have since been renamed or eliminated from current gazetteers and maps, and changes in the physical geographic extent of referenced place names over time, for example the increasing boundaries of an urban area (murphey et al. 2004). perhaps most importantly, existing guidelines also establish a standard means to assign an uncertainty value or maximum error distance associated with geospatial coordinates. recording uncertainty associated with the georeferencing process is critical if the data is to be effectively utilized in future spatially explicit analyses. as further evidence of continued collaboration on protocols for georeferencing collections based on specimen occurrence data, a new initiative called biogeomancer15 is currently in progress. the goal of this multi-institutional international collaborative project is to develop a nextgeneration web services-based georeferencing tool. the tool uses natural language processing techniques to perform locality text to spatial coordinate conversions for single or multiple records, calculate uncertainties, and provide visualization and automated analysis tools (e.g., outlier detection) for validation. the georeferences and original data will ultimately be returned in a format compliant with the darwin core 2 data standard discussed more fully below. how should data access and transmission issues be addressed for online mapping applications? a key challenge, as collections information continues to be digitized and georeferenced, is to make digital data more widely available by providing online access. a key component to sharing and eventually mapping biodiversity data is development of agreed upon standards (figure 1 – community data standards) for accessing widely varying database structures, and associated metadata (bowker, 2000). progress in the natural history museum database community has led to relatively wide-scale adoption of standards like darwin core, and darwin core version 2, which has recently been proposed to the taxonomic database working group16 for adoption in 2005. 15 http://128.32.146.140/bgdev. 16 http://darwincore.calacademy.org. darwin core is a specification of data and concepts to support access to biological collections data that encompasses relatively common concepts across biological databases, such as institutional metadata, and taxonomic, collecting event, and geospatial elements. while its relatively simple structure facilitates use in retrieving and combining data from multiple sources, it is not intended to serve as a data model for managing primary collections databases or specialized disciplinary data within the biodiversity informatics community. the next challenge is how to transmit the data and information using those agreed-upon standards (figure 1: shared registration and transmission protocols). the first major attempt at a distributed biodiversity network was the species analyst, which employed the searchable concepts found in darwin core and utilized the z39.50 protocol. the original z39.50 transmission protocol has been supplanted by a new open source application known as distributed generic information retrieval (digir). digir provides a standardized mechanism by which stewards of natural history collections can make collections information available over the internet. the digir software has two main components. the first is a provider package that allows an institution to link its data to a federated, xml-based natural history data schema (darwin core version 2). the provider software interprets digir requests in the form of xml documents sent to the provider, using http as the transport protocol. the software then makes a native database query, creates an xml result set document and returns it to the requestor via http. the second software package allows institutions to create a portal or central interface for querying a network of distributed digir data providers. manis and gbif are two exemplars that have used digir to establish portal access to collections of data providers. while digir implementations were coming online in north america, the european union began using the biocase to access data from independent heterogeneous collections databases. biocase, like digir, is a software application that uses an xml-based protocol and http to search and retrieve distributed datasets. biocase differs from digir in that it allows a provider to select a conceptual schema, most commonly the access to biological collections data (abcd) schema. because of these two different protocol guralnick and neufeld – distributed mapping solutions for biodiversity analysis 62 figure 3. mapstedi’s implementation of the mvc2 design pattern. the user sends an xml request to the mapstedi controller. the controller determines which objects are called to process requests from both local and distributed gis services. gis services respond with results that are then presented to the browser using jsp. biodiversity informatics, 2, 2005, pp. 56-69 63 implementations, gbif is sponsoring development of a new unified protocol (“tapir”), which when fully implemented will allow sharing distributed datasets across both the digir provider and biocase provider software implementations (doring and giovanni 2004). challenges in developing a global online biodiversity mapping application we have argued that a global online biodiversity mapping application is a lynchpin, but as-yet unrealized, endeavor in biodiversity informatics that links infrastructure development with research questions. the main goal of the mapstedi project was to develop a proof-ofconcept mapping application that could potentially link to other regional projects or itself scale to more global map applications. a secondary goal was to build the application to allow analysis tools to utilize the distributed datasets available in the application. the project ultimately provided us with a set of insights into how global biodiversity mapping applications should be built. we attempt to impart that hard-won knowledge gleaned from mapstedi development to the larger question of how to build effectively a global biodiversity mapping application. mapstedi is a collaborative research project between the university of colorado museum (ucm), denver museum of nature and science (dmns), and the denver botanic gardens (dbg). it facilitated linking separate natural history collections data sets into one distributed biodiversity database accessible through an online mapping application. the project, which currently covers a 6-state region (colorado, montana, nebraska, north dakota, south dakota, and wyoming) in the united states, provides users with access to biodiversity data collected over the last 150 yr. to build the toolkit, we completed three main activities: (1) adding geospatial coordinates assigned from informal place information to existing collection databases following georeferencing procedures discussed above and in more detail in murphey et al. (2004); (2) exporting data into a new spatial database based on the darwin core 2 data standard; and (3) linking distributed online databases to online mapping applications along with other distributed spatial reference layers. the first step in our process (after georeferencing) was to load the data into a geospatial database that could be accessed by our online map server application. the geospatial database used in mapstedi is based on the darwin core version 2 data model. a data conversion application was written that loaded the collections data from comma delimited text files, converted utm coordinates into geographic coordinates (1983 north american datum), and loaded the data into the geospatial database. we used geographic coordinates here because they are well supported in darwincore 2, because our study region crossed multiple utm zones, and because geographic coordinates are more appropriate if our regional focus eventually grows to a more global scale. lastly, in order to link the spatial databases across institutions, we used geospatial multidatabases or database federations (abel 1998), similar in concept to digir but explicitly for spatial databases. the main advantage of this approach is the ability to allow institutions to continue to maintain and update their data locally while providing a mechanism to share and distribute the data through our online mapping application. how to build an efficient web biodiversity mapping given very large underlying datasets? we developed the mapping service to provide both image and tabular data, thus allowing users to explore data before downloading potentially large datasets. it is common knowledge within the gis community (peters 2005) that transporting spatial data across the network in images, as opposed to vector format, greatly reduces the amount of data transferred across the network. in this way online mapping applications permit end-users to preview smaller datasets from remote distributed servers by retrieving remote images from distributed gis data providers, rather than requesting the full results in an uncompressed xml format. mapstedi relied on java’s advanced imaging (jai) libraries to request georeferenced images of partner institutions’ specimen occurrence data and then to fuse seamlessly the combined georeferenced images together before displaying the results in the user’s browser (figures 2-4). other online mapping software packages, such as the university of minnesota’s mapserver, also support access to distributed internet map servers by offering guralnick and neufeld – distributed mapping solutions for biodiversity analysis 64 figure 4. mapstedi interface showing a fused, composite map image for boulder county, colorado made up of the following other images: an image created locally for ucm bird collections (orange dots) along with road (dark red) and river (light blue) layers; an image retrieved from a remote arcims server at the dmns using arcxml showing all their collections (blue dots); and an image (topographic map) retrieved remotely from terraserver-usa.com using wms protocols. tabular data are shown below the map, and records can be tagged and retrieved using tools above the map. biodiversity informatics, 2, 2005, pp. 56-69 65 cascading map server support using the open geospatial consortium’s (ogc) web map service (wms) specification. what image fusion cannot do is to transmit efficiently descriptive attributes of geographic features across the network. to deal with this issue we limited users to previewing attributes in selected record sets to the first 50-300 records, with the option to load the next 50-300 once they had previewed the first set. once a user decides that the data are of interest, they can all be downloaded in a compressed spatial data format that includes both spatial and tabular data. we believe that this approach represents one of the most efficient ways to support full access to data sets without compromising performance within the online mapping application. alternative approaches to deal with the limitations of using uncompressed xml have been explored by the ogc and a report on a binary xml encoding specification has been drafted (bruce 2003). we view transmission of attribute data as an ongoing implementation challenge that could be initially addressed by adding compression filters for provider software packages and decompression filters for portal applications. this approach would allow the community to continue to reap the benefits of using xml as a data exchange format, while significantly improving performance of these systems with minimal application development efforts. how to build a biodiversity mapping application that handles heterogeneous data sources? the decision to use image fusion as the primary means of transporting spatial data between institutions had the added benefit of allowing access to other distributed geographic datasets returning spatially referenced images. such datasets are already available from multiple services (e.g., terraserverusa17), and include datasets such as the u.s. geological survey digital raster graphics (drgs) and digital ortho-photo quadrangles (doqs). both datasets are useful reference layers in biodiversity studies. the main challenge we then faced was how to design the online mapping application to support accessing remote data sources efficiently using different underlying protocols? for example, remote data 17 http://terraservice.net/default.aspx. from mapstedi partner institution dmns is passed using arcxml, while other data sources, such as drgs and doqs from microsoft’s terraserver, rely on ogc’s wms protocol. given the need to support distributed heterogeneous data sources, we designed the system architecture for mapstedi using the java server pages model-view-controller2 (mvc2) design pattern (seshadri 1999). in this design pattern, a browser makes requests to a controller servlet, and the controller then forwards the request to the appropriate model class containing the required application logic. after the model class processes the request, the results are then rendered by a jsp page and subsequently displayed in the browser (see figure 2). mapstedi currently provides access to one local gis service and two distributed gis service providers. the local map service consists of an arcims image service that pulls ucm and dbg collection data from an arcsde database running on top of microsoft’s sql server database. the first distributed gis service based at dmns runs a similar software configuration to the ucm and so is accessible using arcxml, while the second distributed gis service is hosted by microsoft’s terraserverusa.com and is accessible using ogc’s wms protocol (figure 3). in both instances, requests to distributed gis services through the remotehandler classes are made in separate threads. this multi-threaded approach is important for performance reasons so that the application does not have to wait on any one service before sending out additional requests to other distributed gis services. in addition to development challenges of bandwidth and heterogeneous data sources, there are other significant issues to be considered when using distributed gis services. most importantly, because there is no longer central control of system redundancy and network utilization, it is necessary to implement error handling routines that rely on minimum time-outs in the case where a distributed gis service goes down or is not able to respond within a specified time. we chose to include distributed services that could respond within a timeout setting of 20 seconds, and ignored those that could not. for our distributed reference layers from terraserver, we improved application performance by incorporating scale dependencies, such that the drg and doq layers are only guralnick and neufeld – distributed mapping solutions for biodiversity analysis 66 accessible when the user selects scales larger than 1:1,500,000. we believe if online biodiversity mapping applications are to become more widely adopted in the biodiversity informatics community they will likely need to be well integrated with the digir software package. a simple image request mechanism, implemented in a test bed fashion within digir, has been developed. however, full support for the web mapping service18 or web feature service19 specifications has not yet been incorporated (d.a. vieglais, pers. comm.), although such work is now being undertaken by one of the authors. an alternative approach, called digirmapper, is to integrate the university of minnesota’s (umn) map server software20 with digir (hijmans and deck, pers. comm.). we view these efforts as complimentary in that wms/wfsenabled digir providers could serve data to a digirmapper portal given that umn’s map server offers support for distributed wms/wfs layers. how to present attribute data effectively with billions of data points and thousands of repositories? we realize that as digir providers and portals are linked to online gis toolkits, the number of potential data providers will increase far beyond the 3 partnering institutions used in mapstedi. as the number of data sources continues to potentially grow, there are issues with scaling up from two data or three remote data sources to potentially tens to thousands of potential data sources. these issues include but are not limited to providing tools for users to select data layers and re-render them in ways that are most meaningful for visualization and analysis. in particular, we anticipate further application development in the community so that users can customize the visibility, ordering, and transparency settings of spatial data layers. how to build analysis capabilities into online gis toolkits a major next step with online biodiversity mapping applications will be to provide more robust tools for analyzing data. the argument has been made (krishtalka and humphrey 2000, 18 http://www.opengeospatial.org/docs/01-068r2.pdf. 19 http://cite.occamlab.com/test_engine/wfs_1_0_0/wfs_1_0_0.html. 20 http://mapserver.gis.umn.edu/. guralnick and van cleve in press) that natural history collections contain data critical to biodiversity conservation decision-making, and that by examining these patterns we may be able to discover underlying causes for biodiversity change. although providing tools to visualize raw museum collections data location on maps can be useful for heuristic examination of patterns, it is equally important to provide tools for modeling ecological niches or creating summary measures of diversity and to allow tests for differences between these measures across space and time. as more collections and survey data come online, we believe the sampling will be adequate to track species richness and niche change through space and time. researchers have pointed out that there are problems with assuming that charismatic, but less diverse groups such as butterflies, birds, or trees represent the overall species richness for a region (colwell and coddington 1994). a useful feature of building the analytical functions into an online gis application is that they will provide results on any unit of biodiversity selected, whether snails or rodents, flowering plants or ferns. allowing users to select the taxon and geographic area of interest provides the means to examine patterns for less well-studied groups that may have divergent life histories and habitat needs. as well, such tools allow examination of geographical areas that may be of conservation concern but have not been examined or examined using just one higher level taxon (e.g., birds). having tools in an online gis that accesses distributed data sources has some advantages over existing useful desktop packages. one advantage will be that the tools work across platforms and operating systems, allowing more users access to them. another advantage is that the online gis application can constantly access updated distributed repositories for specimen occurrence data and spatial data layers. as new distributed collections databases come online, the online gis will link to them, rather than making the end user collate and update data sets. we believe that continued increases in availability of data and tool development will generally lead to higher quality data to use for analyses, although users will still need to assess the utility and quality of data for their particular applications. guralnick and neufeld – distributed mapping solutions for biodiversity analysis 67 we are particularly excited about building analysis tools using distributed application environments like kepler21. kepler opens the door for statistical tools to be integrated easily into online mapping applications while allowing the code to be repurposed for other applications. kepler is an application for managing scientific workflows and in addition has the ability to be executed as a run-time engine. given kepler’s ability to access distributed grid computing technologies, it makes sense to leverage this technology for complex and computationally intensive statistical analyses. a user could submit such tests to the online mapping application that would run the analyses through a kepler run-time environment and return both statistical and map layer results. thus, a next-generation online gis analysis package could combine the ability to generate species diversity raster maps and species accumulation curves, as well as compare species diversity levels between different datasets through hypothesis testing. finally, internet and desktop mapping tools can interact in useful ways. users may query and explore the data online, as well as eventually download spatial data formats like esri’s shapefiles or other export formats, and continue analyses on their desktop computers. we also realize that there is still much work to be done purely on the gis end towards providing greater analytical capabilities for distributed gis services. two examples of common spatial analyses that we have not yet incorporated include performing spatial searches based on other geographic features: for example, finding all of the collection data points within a federal land unit’s polygon, or locating all the collections with a buffer distance of a user selected hydrologic feature. in these situations, a potentially large number of spatial coordinates will need to be sent to distributed gis services as the query operator. we envision that some control over the size of these coordinate-based query strings could be implemented by storing reference layers on our local server and preprocessing those data sets with appropriate feature generalization techniques. in spite of some of the upcoming challenges and in some cases inherent limitations, we remain excited at the prospect of ongoing advances in distributed 21 http://seek.ecoinformatics.org/wiki.jsp?page=kepler. gis services and their ability to contribute to knowledge synthesis in biodiversity informatics. how to overcome sociological/community impediments to developing online gis for biodiversity? developing a global online distributed biodiversity mapping application involves community challenges, as well as technological challenges. it is beyond the scope of this paper to discuss multifaceted sociological issues in community development of informatics tools fully. instead, we focus on tractable problems that are most closely related to biodiversity mapping, and that offer the beginnings of some solutions. as we have discussed, in order to share data among regional or global map applications that may each have different biodiversity analysis requirements, community agreed-upon data concept and transmission standards are essential to facilitate sharing. we have previously discussed standards and transmission mechanisms like darwincore2 and digir for natural history collections and gml/wfs/wms for geographic information. we also believe that all code developed for such applications should be accessible as open source, with the intent of allowing all interested parties to help continue development of new or improved map application features. constructing the most effective global biodiversity online gis tools will also require building collaborations that include developers and the researchers and managers who will use the applications. such interdisciplinary projects are difficult because many environmental biologists have not kept up with advances in computing, while many computer scientists do not understand the difficult conceptual problems in environmental biology. ultimately, the online gis tools developed should be designed using the best available practices in the biodiversity informatics development community, while being usable by the broadest range of environmental biology users. we believe involvement in the biodiversity informatics community via the taxonomic database working group (tdwg) and gbif is essential for staying abreast of best practices in development. the challenge of building tools most useful to the user community is potentially more difficult. one short-term solution is to perform usability assessments at multiple stages during the guralnick and neufeld – distributed mapping solutions for biodiversity analysis 68 development process. a longer term but equally important solution is to make sure there are people who are able to effectively translate the software development challenges to the research community and vice-versa. because such individuals are currently limited in number, we believe it will require environmental biology and computer science cross-training programs to have a body of workers capable of working from both ends towards the middle. such programs for crossdisciplinary training do not yet exist but will be fundamental for continued growth of biodiversity mapping projects in particular and for biodiversity informatics more generally. finally, the utility of biodiversity mapping tools will need to be made manifest to the research and management user community through tracking access to the tools, papers and talks using the tools, training workshops, and other mechanisms that show the value of the data and tools and teach users how to leverage the technology maximally for their work. acknowledgments we would like to thank paul murphey, and all the undergraduate workers and graduate assistants whom he and the co-authors oversaw. together, these individuals helped with the many tasks needed to make the mapstedi project work. support provided by nsf grant dbi-0110133 for the mapstedi project is gratefully acknowledged. our understanding of the many issues in georeferencing and online gis development for biodiversity data has also increased through our participation in nsf and gbif supported workshops, especially the 2003 yale georeferencing workshop and a 2003 ucsd digir workshop. jeremy mennis, robert hijmans and two anonymous reviewers greatly helped improve the quality of previous manuscript versions. we acknowledge continuing support from the gordon and betty moore foundation for the collaborative research project “biogeomancer” administered through university of california berkeley, and from the global biodiversity information facility for wfs integration into digir portal software. references abel, d. 1998. towards integrated geographical information processing. int. j. geogr. inf. sys. 12:353-371. beaman, r. and b. conn. 2003. automated geoparsing and georeferencing of malesian collection locality data. telopea 10:43-52. bowker, g. 2000. mapping biodiversity. int. j. geogr. inf. sys. 14:739-754. brisby, f. 2000. the quiet revolution: biodiversity informatics and the internet. science 289: 23092312. bruce, g. 2003. binary-xml encoding specification, open gis consortium discussion paper22. colwell, r. k. and j. a. coddington. 1994. estimating terrestrial biodiversity through extrapolation. phil. trans. roy. soc. b 345:101-118. crisci, j. v. 2001. the voice of historical biogeography. j. biogeogr. 28:157-168 doring, m. and r. giovanni. 2004. a unified protocol for search and retrieval of distributed data, gbif23. graham, c. h., s. ferrier, f. huettman, c. moritz, and a. t. peterson. 2004. new developments in museum-based informatics and application in biodiversity analysis. trends ecol. evol. 19:497503. guralnick, r. p. and j. van cleve. 2005. strengths and weaknesses of museum and national survey datasets for predicting regional species richness: comparative and combined approaches. div. and distr. 11:349-359. krishtalka, l. and p. s. humphrey. 2000. can natural history museum capture the future? bioscience 50:611-617. meier, r. and t. dikow. 2004. significance of specimen databases from taxonomic revisions for estimating and mapping the global species diversity of invertebrates and repatriating reliable specimen data. conserv. biol. 18:478-488. murphey, p.c., r. p. guralnick, r. glaubitz, d. neufeld, and j. allen ryan. 2004. georeferencing of museum collections: a review of the problems and automated tools, and the methodology developed by the mountain and plains spatialtemporal database-informatics initiative (mapstedi). phyloinformatics 1:1-29. peters, d. 2005. system design strategies, esri technical reference document24. petersen, t. p., r. meier, and m. n. larsen. 2003. testing species richness estimation methods using the museum label data on the danish asilidae. biodiv. conserv. 12:687-701. peterson, a. t. 2001. predicting species' geographic distributions based on ecological niche modeling. condor 103:599-605. ponder, w.f, g. a. carter, p. flemons and r. r. chapman. 2001. evaluation of museum collection 22 http://www.opengis.org/docs/03-002r8.pdf. 23 http://www.cria.org.br/protocols/newprotocol.pdf. 24 http://www.esri.com/library/whitepapers/pdfs/sysdesig.pdf. guralnick and neufeld – distributed mapping solutions for biodiversity analysis 69 data for use in biodiversity assessment. conserv. biol. 15:1-11. rahbek, c. and g. r. graves. 2001. multiscale assessment of patterns of avian species richness. proc. nat. acad. sci. usa 98:4534-4539 seshadri, g. 1999. understanding javaserver pages model 2 architecture25. soberón, j.m., j. b. llorente, and l. oñate. 2000. the use of specimen-label databases for conservation purposes: an example using mexican papilionid and pierid butterflies. biodiv. conserv. 9:1441-1466. soberón, j. m. and a. t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. phil. trans. roy. soc. b 359:689698. stein, b. and j. wieczorek. 2004. mammals of the world: manis as an example of data integration in a distributed network environment. biodiv. inf. 1:14-22. stockwell, d. 1997. overview of computational biodiversity research.26 suarez, a.v. and n. d. tsutsui. 2004. the value of museum collections for research and society. bioscience 54:66-74. sugden, a and e. pennisi. 2000. diversity digitized. science 289:2305. wieczorek, j., q. guo, and r. hijmans. 2004. the point-radius method for georeferencing locality descriptions and calculating associated uncertainty. int. j. geogr. inf. sci. 18:745-767. wilson, e. o. 2000. a global biodiversity map. science 289:2279. 25 http://www.javaworld.com/javaworld/jw-12-1999/jw-12-ssjjspmvc.html. 26 http://biodi.sdsc.edu/doc/bis/overview.html. microsoft word setal_placeprior_2005_rev.doc biodiversity informatics, 2, 2005, 11-23 11 place prioritization for biodiversity representation using species’ ecological niche modeling víctor sánchez-cordero1, verónica cirelli1, mariana munguía1 and sahotra sarkar2 1 departamento de zoología, instituto de biología, universidad nacional autónoma de méxico. apartado postal 70-153, méxico, d. f. 04510, méxico 2 section of integrative biology and department of philosophy, university of texas at austin, waggener 316, university of texas at austin, austin, tx 78712 -1180, usa abstract.—place prioritization for biodiversity representation is essential for conservation planning, particularly in megadiverse countries where high deforestation threatens biodiversity. given the collecting biases and uneven sampling of biological inventories, there is a need to develop robust models of species’ distributions. by modeling species’ ecological niches using point occurrence data and digitized environmental feature maps, we can predict potential and extant distributions of species in untransformed landscapes, as well as in those transformed by vegetation change (including deforestation). such distributional predictions provide a framework for use of species as biodiversity surrogates in place prioritization procedures such as those based on rarity and complementarity. beyond biodiversity conservation, these predictions can also be used for place prioritization for ecological restoration under current conditions and under future scenarios of habitat change (e.g., deforestation) scenarios. to illustrate these points, we (1) predict distributions under current and future deforestation scenarios for the mexican endemic mammal dipodomys phillipsii, and show how areas for restoration may be selected; and (2) propose conservation areas by combining nonvolant mammal distributional predictions as biodiversity surrogates with place prioritization procedures, to connect decreed natural protected areas in a region holding exceptional biodiversity: the transvolcanic belt in central mexico. key words: biodiversity content, conservation, deforestation, ecological niche, endemic mammals, mexico, place prioritization, restoration. resumen.—la selección de áreas prioritarias de conservación es fundamental en la planeación sistemática de la conservación, particularmente en países de mega-diversidad, en donde la alta deforestación es una de las amenazas a la biodiversidad. debido a los sesgos taxonómicos y geográficos de colecta de los inventarios biológicos, es indispensable generar modelos robustos de distribución de especies. al modelar el nicho ecológico de especies usando localidades de colecta, mapas digitales de variables ambientales y sistemas de información geográficos, se proyecta las distribuciones potencial y actual en hábitat transformados y no transformados por la deforestación. estas hipótesis de distribución proveen un marco teórico para predecir presencia y ausencia de especies, como indicadores de la biodiversidad existente en áreas prioritarias seleccionadas con base en los principios de rareza y complementariedad. para ilustrar esto, se muestran dos ejemplos; (1) se modeló el nicho ecológico de un roedor endémico dipodomys phillipsii, proyectando su distribución en escenarios de deforestación actuales y a futuro. la predicción de la distribución de especies puede ser útil en la selección de áreas prioritarias para la conservación y la restauración, bajo escenarios actuales y futuros de deforestación, permitiendo una planeación sistemática adecuada de la conservación de la biodiversidad, y (2) proponer áreas de conservación, usando predicciones de distribuciones de mamíferos no voladores y procedimientos de selección de áreas prioritarias, como corredores que conecten las áreas naturales prioritarias decretadas en el eje neovolcánico, una región de alta biodiversidad. palabras clave: áreas prioritarias, contenido de biodiversidad, deforestación, mamíferos endémicos, méxico, nicho ecológico, restauración. sánchez-cordero et al. – place prioritization and species’ distributions 12 an emerging goal of conservation biology is to prioritize places for conservation and management on the basis of biodiversity value (margules and pressey 2000; sarkar and margules 2002). traditionally, priority areas for conservation have involved natural protected areas (npas), which have usually been decreed in regions where no previous rigorous quantitative place prioritization was performed (pressey et al. 1996a; sarkar 1999). this is particularly the case for most countries holding exceptionally rich biodiversity, where use of criteria such as scenic value, wilderness quality, or mere availability leads to the practice of ad hoc reservation (alcérreca et al. 1989; pressey 1994). recent efforts involve adding new areas to npa systems to improve biodiversity conservation, including worldwide (rodriguez et al. 2004), continental (andelman and willig 2004), national and regional approaches (pressey et al. 1993, 1996b). selection of priority conservation areas uses large occurrence data sets for particular biological groups of conservation interest. for example, williams et al. (1996) identified priority areas for bird conservation in the united kingdom, based on an exhaustive database of species distribution of more than 170,000 records. they proposed areas holding high species richness (richness hotspots), endemism (rarity hotspots), and sets of areas showing maximal species representation for birds in the united kingdom (williams et al. 1996). further steps for selection of priority areas for biodiversity conservation require inclusion of as many floristic and faunistic groups as possible, given that richness and rarity hotspots do not necessarily coincide geographically between such groups (saetersdal et al. 1993; williams 1998; egbert et al. 1999; peterson et al. 2000). recent studies proposing networks of conservation areas for improving biodiversity conservation include plants and birds in norway (saetersdal et al. 1993), plants (willis et al. 1996), and birds (godown and peterson 2000; fairbanks et al. 2001) in south africa and the united states, and terrestrial vertebrates in the united states (csuti et al. 1997). more complex approaches include probabilistic methods to identify reserve networks representing greatest expected numbers of species (polasky et al. 2000; sarkar et al. 2004), and incorporation of optimization procedures such as flexibility (ability to incorporate crucial information on real conservation problems), efficiency (ability to maximize species richness at the minimum cost), and accountability (solution transparency), identified as key attributes for prioritizing areas for biodiversity conservation (nichols and margules 1993; pressey et al. 1996a; williams 1998; rodrigues et al. 2004; sarkar et al. 2004). further noteworthy efforts for multi-taxa selection of priority areas for biodiversity conservation involve both governmental and academic participation. for example, the mexican national commission for the conservation and use of biodiversity (conabio) has launched one of the most prominent efforts for prioritization and regionalization of areas for conservation in conjunction with decreed npas at the national level, where taxonomic experts play a major role in identifying richness and rarity hotspots for a wide range of biological groups for terrestrial and marine environments (arriaga et al. 2000a,b; challenger 1998; conabio 1998; conabio and semarnat 2000). recently, margules and pressey (2000) proposed a synthetic framework for systematic conservation planning which involves place prioritization as one of its central goals; this framework was extended by sarkar (2004). it consists of a number of stages starting from compilation of information about biodiversity distribution and explicit conservation goals for a region. to achieve this, four problems for biodiversity planning and management must be solved (sarkar et al. 2002; sarkar 2004): (i) surrogate selection: surrogates form the explicit focus of conservation (species, vegetation types, ecosystem types, etc.), and the task of finding adequate surrogates is referred to as the “surrogacy problem”. the most common surrogates are distributions of well-known species, and environmental parameters such as average rainfall, average temperature, and soil type, since these are often the only data available (margules et al. 1995; nix et al. 2000). however, if total species diversity is the explicit conservation goal, whether or not these surrogates are adequate indicators or good predictors has not been solved theoretically (sarkar et al. 2002; sarkar 2004). this is particularly relevant and challenging for regions holding exceptionally high diversity, such as the so-called megadiverse countries; (ii) place prioritization: once surrogates have been chosen, selected places are ordered hierarchically according to their biodiversity sánchez-cordero et al. – place prioritization and species’ distributions 13 content; this is referred to as the “place prioritization problem” and is discussed in more detail below; (iii) viability of biota at prioritized places: for each place selected, projected future scenarios for biological entities (populations, species, communities, ecosystems) must be taken into account. this is usually referred as the “viability problem,” and is perhaps the most difficult stage of the planning protocol to execute in practice. viable areas must be estimated for all surrogates selected and the broader biota for which they are surrogates (sarkar et al. 2002; sarkar 2004). such analyses will result in re-ordering of prioritized areas, taking into account degrees of future viability. selected areas with low viabilities are ranked lower in the place prioritization than those with high viability; place prioritization for restoration requires further analyses. it should be emphasized that such a priority ranking of selected places must reflect their biodiversity value, a primary goal for place prioritization. there are a number of methodologies for estimating viabilities of these biological entities, from conducting stochastic population viability analysis (pva) for small and restricted populations (boyce 1992; burgman et al. 1993), to estimating viability of selected places based on risk of habitat conversion particularly into agrosystems (pressey et al. 1996a); (iv) final selection of appropriate places for management plan implementation: selection of appropriate places for management practices and sustainable use of natural resources should presumably start with those areas with highest biodiversity value. subsequently, socioeconomic, social, and political factors central to launching adequate management practices must also be incorporated into a conservation plan. this is usually referred as the “feasibility problem.” it should be noted that the solution of the last two problems requires reliable and frequent feedback, given that management practices can alter significantly the viability of biotas at selected places. human-induced changes in land use and viability of biota at selected places must be considered in the context of overall conservation and management practices and policies for entire regions. deforestation threatens biodiversity conservation and increasing human pressure for land conversion to agrosystems and urban settlements results in fewer extensive areas potentially devoted to biodiversity conservation. consequently, urgent action for conservation planning based on systematic place prioritization criteria is urgently needed. in particular, deforestation can impact biodiversity distributions significantly (sánchez-cordero et al. 2004), requiring proper adjustment or re-analysis of management plans. such “adaptive management” involves recurrent evaluation of place prioritizations for conservation content (meffe and carroll 1997; sarkar et al. 2002; sarkar 2004). in this contribution, we will focus only on the surrogacy and place prioritization problems. we propose the use of ecological niche modeling, by which presences or absences of species are interpreted into potential distributional areas. these modeled distributions provide a theoretical framework for use of species as surrogates for overall biodiversity content. we describe a place prioritization protocol, emphasizing selection of places containing rare surrogates (the principle of “rarity”) and places that add as many underrepresented surrogates as possible to a set of selected places (the principle of “complementarity”), and discuss fruitful options not only for systematic conservation, but also for restoration planning. this protocol combines ecological niche modeling with place prioritization for biodiversity representation under current and future deforestation scenarios. modeling species’ ecological niches natural history museum collections store massive amounts of information on biodiversity, containing primary information on species’ geographic occurrences for documenting biodiversity worldwide. this information can be compiled into databases containing species’ records and georeferenced collecting localities gathered from national and international institutions (soberón 1999). however, given the uneven and biased taxonomical and geographical nature of museum collections, tools for extrapolating from what is known to a more general prediction of species’ distributions are necessary. several efforts have advocated modeling species’ ecological niches, based on a grinnellian geographic niche concept (grinnell 1917; macarthur 1972), which are then projected as potential distributional maps (peterson et al. 1999; sánchez-cordero et al. 2001, 2004). one such sánchez-cordero et al. – place prioritization and species’ distributions 14 approach is based on use of a genetic algorithm, in combination with occurrence data sets and digitized maps of environmental features. in particular, the genetic algorithm for rule-set prediction (garp1, stockwell and peters 1999) uses an evolutionary computing approach to compute niche models that can be projected as potential geographic distributions of species (peterson et al. 1999; stockwell and peters 1999). garp has proven a robust method for modeling species’ ecological niches for large numbers of taxa (peterson 2001; peterson et al. 1999; peterson and kluza 2003; illoldi-rangel et al. 2004; sánchez-cordero et al. 2004). while ecological niche modeling predicts potential geographic distributions of species, certain areas may not be occupied currently, given effects of other factors external to the model, such as historical constraints, species interactions, and changes in land use patterns (anderson et al. 2003; peterson et al. 1999; sánchez-cordero et al. 2001, 2004). one such factor is deforestation, which leads to reduction and fragmentation of natural habitats, limiting species’ realized distributions to subsets of their potential distributions. we can quantify reductions of species’ distributions by constructing potential niche models based on climate, topography, and reference vegetation maps, and then overlaying these maps on actual land use/land cover maps. extant distributions of species are taken as areas holding untransformed natural habitats within potential distributions, assuming that deforested areas probably constitute inviable ecological conditions (sánchez-cordero et al. 2004). our approach to modeling species’ extant distributions based on current land-use patterns provides a testable framework for predicting where species are present, as well as regions in which populations are reduced or extirpated by loss of natural habitats. predictions of potential and actual distributions of species serve to indicate effects of scenarios of habitat transformation on biodiversity; these distributions can then be used to plan for conservation and restoration. to illustrate this approach, we show potential scenarios of changes in species’ distributions due to deforestation. the heteromyid dipodomys phillipsii is an endemic rodent associated with arid and semi-arid habitats and occurring on the mexican plateau, the transvolcanic belt and 1 http://www.lifemapper.org/desktopgarp. oaxaca (hall 1981). we modeled the ecological niche of d. phillipsii using garp, point occurrence data georeferenced to the nearest 0.01° of longitude and latitude for each locality using 1:250,000 topographic maps (comisión nacional para el conocimiento y uso de la biodiversidad (conabio 19982), and 10 environmental data layers (0.04 × 0.04° pixel resolution), including potential vegetation type (rzedowski 1986); elevation, slope, and aspect (from the u.s. geological survey’s hydro-1k data set3); and climatic parameters including mean annual precipitation, mean daily precipitation, maximum daily precipitation, minimum and maximum daily temperature, and mean annual temperature (conabio 1998). we then overlayed transformed areas due to human-induced habitat conversion, based on satellite imagery resulting in a land use/land cover map for 1980 and 2000 (velázquez et al. 2001). all point occurrence data from collecting localities (n = 73) for this species are dated before 1970 (hall 1981); as such, specimens were collected in natural habitats prior to the 1980 land use habitat transformation within its range (hall 1981). transformed areas converted into agrosystems and urban areas are presumed to represent non-viable ecological conditions for this rodent (sánchez-cordero et al. 2004). this assumption based on hypotheses of general niche conservatism tested for diverse taxa in mexico (peterson et al. 1999; peterson and holt 2003) assumes that rapid adaptation to new environments produced by human-induced habitat transformation is unlikely, particularly without recurrent immigration from adjacent natural habitats (peterson and holt 2003). these hypotheses are particularly likely to hold for locally-adapted endemic species. we found extensive transformed areas, particularly in the transvolcanic belt (tvb), as well as in the northwestern and southern portions of the distribution of this species (figure 1). moreover, habitat conversion from untransformed to transformed uses, when the species’ potential and actual distributions are compared (based on the inventario nacional forestal 1980 and 2000, respectively), showed that areas already fragmented are more likely to suffer additional reduction due to human-induced habitat 2 http://www.conabio.gob.mx. 3 http://www.usgs.gov. sánchez-cordero et al. – place prioritization and species’ distributions 15 figure 1. ecological niche models projected as potential geographic distribution (green on top map), and actual distribution (blue and red on top and bottom maps) for the mexican endemic mammal dipodomys phillipsii. species presence in actual distributions is predicted only in untransformed habitats (blue and red areas) within the species’ potential distribution. note that areas highly fragmented in the 1980 (blue fragments) in the northeast, central, and southern regions of the species’ distribution were more likely to suffer further fragmentation and area reduction (red fragments depicted by arrows). if this trend continues, local extirpations are likely to occur in these regions. sánchez-cordero et al. – place prioritization and species’ distributions 16 degradation (figure 1). if such a trend continues, fragments remaining in the tvb, and in the northwestern and southern parts of the species’ range will suffer further severe reductions or will simply disappear. as a consequence, these areas are predicted at being of high risk of population extirpations, and should be selected for restoration. (obviously, such distributional scenarios should first be validated in the field.) place prioritization procedures we will briefly discuss the place prioritization procedures incorporated in the resnet software package4 (sarkar et al. 2002; aggarwal et al. 2000). algorithms in resnet belong to the family of algorithms introduced by margules et al. (1988; nicholls and margules 1993), but add a novel dynamic memory allocation scheme that results in no constraints on size of the data set. several recent studies of regional planning for conservation purposes use resnet (sarakinos et al. 2001; sarkar et al. 2004; tognelli, in press). the operational procedure starts when a region is divided into a set of places on the basis of geographical coordinates, ecoregions, or biogeographic regions, and the algorithm orders those places by their biodiversity content. the algorithm assumes either that an explicit target has been set for adequate representation of each surrogate (e.g., number of selected places at which a surrogate must be present), or that a maximum allowed area or a maximum allowed cost of a proposed set of priority places has been specified. the goal of all such algorithms is to achieve the target as economically as possible, by selecting as few places as possible for reaching the conservation goal (margules et al. 1988; sarkar et al. 2004). three rules are incorporated into the algorithms of resnet: (i) rarity: surrogates are first ordered inversely by their frequency of appearance in the data set. then, places are ordered according to whether they contain the rarest surrogate, the next rarest surrogate, and so on, iteratively. (ii) complementarity: places are ordered based on numbers of surrogates that have not met their targeted representation. (iii) richness: places are ordered based on number of surrogates present; however, richness is used only in initial step (selecting the first place), since it has been shown previously that reliance on richness results in 4 http://uts.cc.utexas.edu/~philsci/sarkar/main.html. inefficient place selection (williams et al. 1996; csuti et al. 1997). in resnet, three criteria may be used to initialize the prioritization procedure: rarity, richness, or from a set of pre-selected places (e.g., existing protected areas). for both initialization and iterative place selection, ties are broken arbitrarily by selecting the first place on the list, so a unique place is chosen. further refinement of place prioritization can be achieved by introducing adjacency considerations, by which areas neighboring already-selected areas are given preference over others, resulting in larger and more contiguous areas. iterations continue until the target is met—that is, that all surrogates are adequately represented or the maximum allowed area or cost is exceeded. if no explicit target is set, the procedure continues until all places are selected (sarkar et al. 2002). the order in which places are selected produces a ranking of places based on their biodiversity content. biodiversity content is thus implicitly defined by the algorithm, and the intuition behind this approach is that diversity is adequately captured by rarity and complementarity (sarkar 2002; sarkar and margules 2002). as expected, depending on initialization and iteration criteria chosen, a number of different solutions can be achieved (sarkar et al. 2002). current challenges and an example the two techniques—ecological niche modeling and place prioritization—merge to form a synthetic protocol in the emerging field of biodiversity informatics, with potentially extensive applicability to conservation planning. ecological niche modeling facilitates inclusion of more taxa than would otherwise be possible, providing a framework for incorporation of large numbers of species, including those with high conservation priority, as biodiversity surrogates (egbert et al. 1999; peterson et al. 2000; rojas-soto et al. 2004). this is particularly true in megadiverse countries, despite controversies about whether the surrogates commonly employed are true indicators of biodiversity (margules and pressey 2000; sarkar et al. 2002). as a consequence, place prioritization for biodiversity content using criteria of rarity and complementarity can be implemented readily based on robust models of species’ geographic distributions. we envision a challenging research program for applying these approaches to prioritization challenges in megadiverse countries worldwide (rodrigues et al. 2004). current sánchez-cordero et al. – place prioritization and species’ distributions 17 proposals from international conservation organizations such as iucn and world park commission5 encourage governments worldwide to include at least 10% of their land into reserves for launching conservation programs; these potential natural protected areas can be selected following methodologies described herein and elsewhere (margules and pressey 2000; sarkar et al. 2002). recent studies using ecological niche modeling of multiple taxa, and place prioritization procedures hold promising for identifying additional areas devoted for conservation (egbert et al. 1999; peterson et al. 2000). deforestation ranks among the major threats to biodiversity conservation worldwide, making place prioritization for biodiversity content urgent. in many countries, institutional and governmental efforts on bioinformatics are making available massive amounts of information from natural history museum specimens and digital environmental data on the internet (conabio, inbio6, manis7). modeling ecological niches projected as potential and actual distributional hypotheses provide a framework for understanding species’ distributions across current untransformed and transformed landscapes (sánchez-cordero et al. 2001, 2004), useful for assigning probabilities of presence of surrogates in place prioritization procedures for biodiversity content. species’ actual distributions based on ecological niche models where only untransformed habitats are included can be further used as baseline distribution hypotheses for inclusion in place prioritization procedures (figure 1) (munguía 2003; sánchezcordero et al. 2004). we illustrate this point by combining species’ actual distributions (e.g., figure 1, bottom panel) with place prioritization procedures using terrestrial nonvolant mammals in the tvb as biodiversity surrogates to identify priority areas for connecting decreed natural protected areas (npas; fig. 2; see munguía 2003; munguía et al. in prep.). this region holds exceptionally rich biodiversity, but rampant deforestation threatens its conservation. conversely, the tvb holds many decreed npas, including 39 npas of 1000 ha or more (munguía 2003; sánchez-cordero et al. 2004; fuller et al. submitted). 5 http://www.iucn.org. 6 http://www.inbio.ac.cr. 7 http//elib.cs.berkeley.edu/manis/. searching among already-existing areas, we selected 13 priority areas based on endemicity, species richness, and complementarity: reserva de la biósfera sierra de manantlán, parque nacional volcán nevado de colima, parque nacional la primavera, and parque nacional sierra de quila, in the western region; reserva de la biosfera corredor biológico chichinautzin, parque nacional el tepozteco-zempoala, reserva de la biosfera mariposa monarca, and parque nacional el cimatario, in the central region; and parque nacional la malinche, reserva de la biosfera valle de tehuacán-cuicatlán, and parque nacional cofre de perote, in the eastern region, of the tvb (figure 2). we then proposed connecting these npas by choosing remnant untransformed habitats, based on the 2000 land use/land cover map (conabio), lying along straight paths, as follows: for the western region, corridors connected sierra de manantlán with volcán de colima, sierra de quila, and la primavera; for the central region, corridors connected izta-popo with corredor biológico chichinautzin, el tepoztecozempoala, mariposa monarca, and el cimatario; for the eastern region, corridors connected la malinche with valle de tehuacán-cuicatlán, pico de orizaba, and cofre de perote (fig. 2) (munguía 2003; munguía et al. in prep.). a further refinement of this analysis using actual distributions of 99 terrestrial nonvolant mammals as biodiversity surrogates, place prioritization procedures, and graph theory for identifying priority areas in the tvb is presented elsewhere (fuller et al. submitted). future projections of deforestation can also be incorporated into the niche modeling framework to generate predictions of potential distributional areas under forecasts of habitat transformation scenarios. such distributional models can be used as surrogates in place prioritization procedures to identify priority areas under current and future scenarios of deforestation (sánchez-cordero et al. 2001, 2004). deforestation has major impacts on species’ distributions (sánchez-cordero et al. 2004), so efforts to combine niche modelling with place prioritization procedures offer an extremely useful tool for planning current and future conservation strategies in megadiverse countries. place prioritization for biodiversity content also enables inclusion of selection of areas for habitat restoration based on comparison of potential and actual distributions in transformed and sánchez-cordero et al. – place prioritization and species’ distributions 18 a figure 2. proposed areas connecting decreed priority natural protected areas (npas) in the transvolcanic belt of central mexico, a region of exceptional biodiversity. priority npas were selected based on richness of endemism, species richness, and complementarity (munguía 2004; munguía et al, in prep.). (a) overview of the tvb, depicting remnant untransformed natural habitats based on 2000 land use/land cover map (green area), and selected priority decreed npas (delineated polygons). (b) selected areas identified as corridors (see arrows) of remnant untransformed habitat connecting priority npas for the western, central, and eastern regions of the tvb. priority npas shown are: (1) reserva de la biósfera sierra de manantlán, (2) parque nacional nevado de colima, (3) parque nacional la primavera (4) parque nacional sierra de quila, (5) parque nacional izta-popo, (6) reserva de la biosfera corredore biológico chichinautzin, (7) parque nacional el tepozteco-zempoala, (8) reserva de la biosfera mariposa monarca, (9) parque nacional el cimatario, (10) parque nacional la malinche, (11) reserva de la biosfera valle de tehuacancuicatlán, (12) parque nacional pico de orizaba, and (13) parque nacional cofre de perote. (3 pages) western central eastern a sánchez-cordero et al. – place prioritization and species’ distributions 19 central 1 2 3 4 b 5 6 7 8 9 c sánchez-cordero et al. – place prioritization and species’ distributions 20 10 11 12 13 d sánchez-cordero et al. – place prioritization and species’ distributions 21 untransformed landscapes (fuller et al. submitted). our distributional models computing potential and actual distributions in untransformed and transformed landscapes can provide baseline information for selection of priority areas for restoration. for example, transformed areas within the potential distributions of priority species are potential areas for restoration from transformed habitats to the original untransformed natural habitat. the above approaches of ecological niche modelling, reconstructing actual distributions, and incorporation into place prioritization procedures results in robust analytical tools for improving systematic conservation planning protocols for conservation and restoration sites (margules and pressey 2000; sarkar 2004). acknowledgments we thank a. t. peterson for his invitation to participate in this electronic journal. a. t. peterson, m. f. figueroa, t. escalante, and two anonymous reviewers provided useful comments to an earlier version of the manuscript. environmental digital maps and specimen data were provided by conabio. the following museum collections were consulted to built the mammal database: colección nacional de mamíferos, universidad nacional autónoma de méxico; colección de mamíferos, universidad autónoma metropolitana-iztapalapa; centro interdisciplinario de investigación y desarrollo regional de oaxaca; university of kansas natural history museum; american museum of natural history, new york; national museum of natural history, washington, d.c.; field museum, chicago, illinois; museum of zoology, university of michigan, ann arbor; michigan state university museum, east lansing; museum of vertebrate zoology, university of california, berkley; texas tech university museum, lubbock; texas cooperative wildlife collections, texas a&m university, college station, and collections served by manis8, including museum of vertebrate zoology and university of kansas natural history museum. this work was partially funded by the secretaría del medio ambiente y recursos naturales and the consejo nacional de ciencia y tecnología (semarnat-conacyt project 2002-c01-314-a1 to vs-c). v. cirelli was supported by a packard foundation scholarship for the masters program on biología ambiental 8 http//elib.cs.berkeley.edu/manis. (restauración ecológica), and the graduate program at the national autonomous university of mexico (unam) to visit the university of texas at austin. m. munguía was supported by the consejo nacional de ciencia y tecnología (project conacyt 2001-35472-v to vs-c) for her bachelor’s thesis. references alcérreca, c., j. consejo, o. flores, d. gutiérrez, e. hentschel, m. herzig, r. pérez-gil, j. m. reyes, and v. sánchez-cordero. 1989. fauna silvestre y áreas naturales protegidas. universo veintiuno. méxico, d. f. andelman, s., and m. willig. 2003 present patterns and future prospects for biodiversity in the western hemisphere. ecology letters 6:1-7. anderson, r. p., d. lew, and a.t. peterson. 2003. evaluating predictive models of species’ distributions: criteria for selecting optimal models. ecological modelling 162:211-232. arriaga, l., j. m. espinosa-rodríguez, c. aguilarzúñiga, e. martínez-romero, l. gómez-mendoza, and e. loa loza. 2000a. regiones prioritarias terrestres de méxico. comisión nacional para el conocimiento y uso de la biodiversidad. méxico, d. f. arriaga, l., v. aguilar sierra, and j. alcocer durand. 2000b. aguas continentales y diversidad biológica de méxico. comisión nacional para el conocimiento y uso de la biodiversidad. mexico, d.f. aggarwal, a., j. garson, c. r. margules, a. o. nicholls, and s. sarkar. 2000. resnet ver 1.1 manual. report. biodiversity and biocultural conservation laboratory, university of texas. boyce, m. s. 1992. population viability analysis. annual review of ecology and systematics 23:481-506. burgman, m., s. ferson, and h. r. akçakaya. 1993. risk assessment in conservation biology. new york: chapman and hall. challenger, a. 1998. utilización y conservación de los ecosistemas terrestres de méxico: pasado, presente y futuro. comisión nacional para el conocimiento y uso de la biodiversidad, instituto de biología, universidad nacional autónoma de méxico, y sierra madre. méxico d.f. comisión nacional para el conocimiento y uso de la biodiversidad (conabio). 1998. la diversidad biológica de méxico: estudio de país. conabio, méxico comisión nacional para el conocimiento y uso de la biodiverisdad (c onabio) and secretaría del medio ambiente y recursos naturales sánchez-cordero et al. – place prioritization and species’ distributions 22 (semarnap). 2000. estrategia nacional sobre biodiversidad de méxico. conabio, mexico csuti, b., s. polasky, p. h. williams, r. l. pressey, j. d. camm, m. kershaw, a. r. kiester, b. downs, r. hamilton, m. huso, and k. sahr. 1997. a comparison of reserve selection algorithms using data on terrestrial vertebrates in oregon. biological conservation 80:83 -97. egbert, s. l., a. t. peterson, v. sánchez-cordero, & k. price. 1999. modeling conservation priorities in veracruz, mexico. pp. 141-150 in gis solutions in natural resource management. (s. morain, ed.). onword press, santa fe, new mexico. fuller, t., m. munguía, m. mayfield, v. sánchezcordero, and s. sarkar. submitted. using connectivity to integrate conservation and restoration planning: a case study from central mexico. conservation biology. grinnell, j. 1917. the niche-relationship of the california thrasher. auk 43:427-433. godown, m., and a. t. peterson. 2000. preliminary distributional analysis of ud endangered bird species. biodiversity and conservation 9:13131322. hall., e. r. 1981. the mammals of north america. vol. i & ii. ronald press, new york. illoldi-rangel, p., v. sánchez-cordero, and a. t. peterson. 2004. predicting distributions of mexican mammals using ecological niche modeling. journal of mammalogy 85:658-662. macarthur, r. h. 1972. geographical ecology. harper and row. princeton, new jersey. margules, c. r., a. o. nicholls, and r. l. pressey. 1988. selecting networks of reserves to maximize biological diversity. biological conservation 43:63-76. margules, c. r. and r. l. pressey. 2000. systematic conservation planning. nature 405:242-253. margules, c. r., t. d. redhead, d. p. faith, and m. f. hutchinson. 1995. guidelines for using the biorap methodology and tools. csiro, canberra. meffe, g. k., and c. r. carroll. 1997. principles of conservation biology. second edition. ed: sinauer associates. inc., publishers. sunderland, massachusetts. munguía, m. 2004. representatividad mastofaunística en areas naturales protegidas y regiones terrestres prioritarias en el eje neovolcánico: un modelo de conservación. tesis de licenciatura. facultad de ciencias, universidad nacional autónoma de méxico. nicholls, a. o., and c. r. margules. 1993. an upgraded reserve selection algorithm. biological conservation 64:165-169. nix, h. a., d. p. faith, m. f. hutchinson, c. r. margules, j. west, a. allison, j. l. kesteven, g. natera, w. slater, j. l. stein, and p. walker. 2000. the biorap toolbox: a national study of biodiversity assessment and planning for papua new guinea. canberra: centre for resource and environmental studies, australian national university. peterson, a. t. 2001. predicting species geographic distributions based on ecological niche modeling. condor 103:599-605. peterson, a. t., j. soberón, and v. sánchez-cordero. 1999. conservatism of ecological niches in evolutionary time. science 285:1265-1267. peterson, t., s. l. egbert, v. sánchez-cordero, & k. v. price. 2000. geographic analysis of conservation priorities for biodiversity: a case study of endemic birds and mammals in veracruz, mexico. biological conservation 93:85-94. peterson, a. t., and d. kluza. 2003. new distributional modeling approaches for gap analysis. animal conservation 6:47-54. peterson, a. t., r. d. holt. 2003. niche differentiation in mexican birds: using point occurrences to detect ecological innovation. ecology letters 6:774-782. polasky, s., j. d. camm, a. r. solow, b. csuti, d. white, and r. ding. 2000. choosing reserve networks with incomplete species information. biological conservation 94:1-10. pressey, r. l. 1994. ad hoc reservations: forward of backward steps in developing representative reserve systems. conservation biology 8:662-668. pressey, r. l., c. j. humphries, c. r. margules, r. i. vane-wright, and p. h. williams. 1993. beyond opportunism: key principles for systematic reserve selection. trends in ecology and evolution 8:124128. pressey, r. l. and a. o. nicholls. 1989. efficiency in conservation evaluation: scoring versus iterative approaches. biological conservation 50:199-218. pressey, r. l., h. p. possingham, and c. r. margules, 1996a. optimality in reserve selection algorithms: when does it matter and how much? biological conservation 76:259-267. pressey, r. l., s. ferrier, t. c. hager, c. a. woods, s. l. tully, and k. m. weinman.1996b. how well protected are the forests of north-eastern new south wales? analyses of forest environments in relation to tenure, formal protection measures and vulnerability to clearing. forest ecology and management 85:311-333. rodrigues, a., s. l., s, j. andeman, m. i. bakarr, l. boitani, t. m. brooks, r. m. cowling, l. d. c. fishpool, g. a. b. da fonseca, k. j. gaston, m. i. hoffmann, j. s. long, p. a. marquet, j. d. pilgrim, r. l. pressey, j. schipper, w. sechrest, s. n. stuart, l. g. underhill, r. w. waller, watts, e. j. matthew. 2004. effectiveness of the global protected area network in representing species diversity. nature 428:640-643. sánchez-cordero et al. – place prioritization and species’ distributions 23 rojas-soto, o. r., o. alcántara-ayala y a. g. navarro. 2003. regionalization of the avifauna of the baja california peninsula, mexico: a parsimony analysis of endemicity and distributional modeling approach. journal of biogeography 30:449-461. sánchez-cordero, v., a. t. peterson, and p. pliegoescalante. 2001. modelado de la distribución de especies y conservación de la diversidad biológica. pp. 359-379. in enfoques contemporáneos en el estudio de la diversidad biológica. instituto de biología, unam y academia mexicana de ciencias, a.c., mexico, d.f. sánchez-cordero, v., m. munguía, and a. t. peterson. 2004. gis-based predictive biogeography in the context of conservation. pp. 311-323 in frontiers in biogeography. m. lomolino and l heaney, eds. sinauer press, sunderland, mass. sarakinos, h., a. o. nicholls, a. tubert, a. aggarwal, c. r. margules, and s. sarkar. 2001. area prioritization for biodiversity conservation in québec on the basis of species distributions: a preliminary analysis. biodiversity and conservation 10:1419-1472. sarkar, s. 1999. wilderness preservation and biodiversity conservation-keeping divergent goals distinct. bioscience 49:405-412. sarkar, s. 2002. defining “biodiversity”: assessing biodiversity. monist. 85:131-155. sarkar, s. 2004. conservation biology. the stanford encyclopedia of philosophy (summer 2004 edition), e.n. zalta (ed.) http://plato.stanford. edu/archives/sum2004/entries/conservationbiology/. sarkar, s., a. aggarwal, j. garson, c. r. margules, and j. zeidler. 2002. place prioritization for biodiversity content. journal of biosciences 27(s2):339-346. sarkar, s., and c. r. margules. 2002. operationalizing biodiversity for conservation planning. journal of biosciences 27(s2):299-308. sarkar, s., c. pappas, j. garson, a. aggarwal, and s. cameron. 2004. place prioritization for biodiversity conservation using probabilistic surrogate distribution data. diversity and distributions 10:125-133. society for ecological restoration (ser). 2002. the ser primer on ecological restoration. science and policy group. april 2002:1-9. soberón, j. 1999. linking biodiversity information sources. trends in ecology and evolution 14:291. stockwell, d. r. b., d. peters. 1999. the garp modeling system: problems and solutions to automated spatial prediction. international journal of geographical information science 13:143-158. tognelli, m. 2004. assessing the utility of surrogate groups for the conservation of south american terrestrial mammals. biological conservation. in press. velázquez, a., j. f. mas, j. r. diáz-gallegos, r. mayorga-saucedo, p. c. alcantara, r. castro, t. fernández, g. bocco, e. escurra, and j. l. palacios. 2001. patrones y tasas de cambio de uso de suelo en méxico. gaceta ecológica nueva época no. 62. instituto nacional de ecología y secretaria del medio ambiente y recursos naturales, méxico d.f., méxico. williams, p., d. gibbons, c. r. margules, a. rebelo, c. humphries, and r. pressey. 1996. a comparison of richness hotspots, rarity hotspots, and complementary areas for conserving diversity of british birds. conservation biology 10:155-174. approaches to estimating the universe of natural history collections data biodiversity informatics, 7, 2010, pp. 81 – 92 approaches to estimating the universe of natural history collections data arturo h. ariño department of zoology and ecology university of navarra, pamplona, spain, artarip@unav.es abstract.— this contribution explores the problem of recognizing and measuring the universe of specimen-level data existing in natural history collections around the world, and in absence of a complete, world-wide census or register. estimates of size seem necessary to plan for resource allocation for digitization or data capture, and may help to represent how many vouchered primary biodiversity data (in terms of collections, specimens or curatorial units) might remain to be mobilized. it further helps to set priorities, and assess certainties. three general approaches are proposed for further development, and initial estimates are given. probabilistic models involve crossing data from a set of biodiversity datasets, finding commonalities and estimating the likelihood of totally obscure data from the fraction of known data missing from specific datasets in the set. distribution models aim to find the underlying distribution of collections’ compositions, estimating the hidden sector of the distributions. finally, case studies seek to compare digitized data from collections known to the world to the amount of data known to exist in the collection but not generally available or not digitized. preliminary estimates of size range from 1.2 to 2.1 gigaunits (109) of which a mere 3% at most is currently web-accessible through gbif’s mobilization efforts. however, further data and analyses, along with other approaches relying more heavily on surveys, might change the picture and possibly help to narrow the estimate further. in particular, unknown collections not having emerged through literature are the major source of uncertainty. key words.— natural history collections, size, estimates, primary biodiversity data the global biodiversity information facility (gbif) aims to make the world’s biodiversity information freely and openly available via the internet (gbif, 2003), as recommended in the oecd bioinformatics report (oecd, 1999). a significant proportion of this information comes from specimens in natural history collections, generally hosted in museums whose mission includes documenting and studying life on earth and therefore biodiversity (krishtalka & humphrey, 2000). gbif recently convened a task group to catalyze the development of a global strategy and action plan for further mobilization of natural history collections data worldwide (gsapnhc), which among other objectives seeks to tackle the task of providing metadata describing the scope of natural history collections and the current status of their digitization (gbif, 2010). in addition, gbif encourages a national, regional and thematic implementation and enrichment of the global biodiversity resources discovery system (gbrds) to facilitate the decentralized discovery of all biodiversity datasets worldwide (gbif, 2010) to make them generally available for research. unfortunately, no complete, global repository of metadata about the contents of such collections (indeed, a single inventory of the world’s collections of natural history) seems to exist at present. there are extensive lists and catalogues for selected biological groups, such as index herbariorum (thiers, 2010), or initiatives seeking to index existing collections, such as the biodiversity collections index (bci, 2010), but that have yet to encompass all known collections. in addition, biodiversity collections in private mailto:artarip@unav.es ariño – approaches to estimating the universe of natural history collections data box 1: definitions the “size” of a collection is highly dependent on the “units” of measure. in general, this chapter uses the following concepts: ‐ specimen: an individual organism (or part thereof if treated as an individual), which may be accessioned and stored either separately or together with other specimens, i.e. a herbarium sheet, a pinned moth, or each springtail in a single vial holding hundreds. ‐ lot: a specimen or a set of specimens that are stored and generally handled together, i.e. individuals from a single species collected at a single site in a single sampling event. ‐ unit: a specimen or set of specimens that has been registered as a single data record and treated as a whole. generally coincidental with a lot. ‐ collection: a set of units that are held or linked together according to some criteria, i.e. the coleoptera collected by an expedition or the lichens from a herbarium. ‐ primary biodiversity record (pbr): a single data record that points to a unit, and that is backed by an actual observation or vouchered specimen. generally include at least a taxonomical identification, a location, and a time of capture or observation. a museum record for a specimen or a field observation is a pbr, but a datapoint in a map from a distribution model not backed by an actual observation is not. hands or poorly documented public collections may not have emerged through literature or registers. at present, it is thus impossible to have a census of all biodiversity data backed by vouchered specimens, and therefore the magnitude of this mass of data cannot be known with certainty. however, estimates of size seem necessary to plan for resource allocation for digitization or data capture, and may help represent how many vouchered primary biodiversity data (in terms of collections, specimens or curatorial units) might remain to be mobilized. current published estimates have been derived mostly from curatorial units (lots, specimens, and collections) and yield and aggregate 2.5 3 g (billion) specimens in collections worldwide (duckworth et al., 1993; oecd, 1999, in chapman, 2005). between 2005 and 2007, about one third of gbif node managers had reported an estimated 407 m (million) to 737 m sharable specimen holdings in potential data providers in their countries only (gbif, 2008b, 2006); although these may seem low figures, quite a number of data-rich countries had not yet joined gbif by then so the actual estimate was expected to grow as data from these countries became accessible. these data, however, are very dependent on the composition of the curatorial units. whereas in some fields (i.e. herbaria) accessions and specimens can generally be used interchangeably, this is not so in many zoological sections, or when dealing with collected field samples. therefore, as it respects to the unit of interest for gbif (the “primary biodiversity record” or pbr), it is relevant to establish the meaning of “size”; to try to translate “specimen” into pbr; and to try to refine current size estimates by figuring out what part of this “size” may have not been accounted for yet. this “size” marks the outer boundary within which useful, quality, “fit-for-use” data (hill et al., 2010) can be found to answer specific questions in, e.g. biodiversity, sustainability, or conservation. recognizing that surveys are incomplete (for example, not all nodes in gbif reported size estimates in their countries) and that a mandatory register of biodiversity collections does not exist yet, other ways to independently estimate this size (or size range) should be explored for any efficient planning of digitization efforts. in this paper, i propose several approaches and apply them with 82 ariño – approaches to estimating the universe of natural history collections data box 2: seber probabilistic model analogue seber (1982) applied probability theory to the problem of recognizing how many tagged animals had lost their marks in a recapture experiment, and therefore would have shown up in a sample as “untagged”. we may analogously define a collection as “tagged” if such collection is included within a repository, which is itself actually a sample of the universe of existing collections that, if fully known, should have been all included in the repository. if two independent repositories, a and b, exist which contain a fraction of all collections, a collection may belong to both repositories (rab), to one (ra), the other (rb), or none (r0). ideally, the total number of collections would be 0rrrrr abba +++= but since we do not know r0, we can only estimate r. seber (1982) shows from probability theory that this estimate can be derived as ( )abba rrr k r ++ − = 1 1ˆ where k could be interpreted, in our context, as the product of the probabilities for a collection of not belonging to one repository when belonging to the other: ))(( abbaba ba rrrr rrk ++ = . seber and felton (1981) noted that there is a certain bias in some estimates of the multinomial function used to derive k, and this bias, being negligible for sample sizes approaching the universe being sampled, may grow significantly for small sample sizes. sample data to derive initial estimates of size for first overview, planning and strategic use. further development and refinement of these approaches, and their application to a larger sample of data, could result in more precise estimates. for digitization purposes, and towards contributing to gbif’s stated goal of mobilizing about one billion specimen primary records (gbif, 2006, 2009), one may try to give (i) estimates of the number of accessions (“units”, with one or more specimens in each), as there will generally be a one-to-one correspondence between collection unit and occurrence record in gbif’s indexes, and (ii) the number of collections to be indexed. whereas the first set of estimates should itself serve as proxy for the amount of person-time work ahead, the second one would probably best related to the number of digitizing teams in the digitizing effort. i will first deal with collections, as the unit numbers may depend on this. collections several approaches can be followed in order to estimate the number of collections to be digitized, and with different results. i outline here some of these, adapted or derived from ecological and sampling theory. data crosscheck the biodiversity collections index (bci, 2010) inherits biocase metadata (berendsohn et al., 2000, 2002), and lists index herbariorum herbaria (thiers, 2010), invertebrate codens (samuelson and evenhuis, 1998), and other collections. as of october, 2008, bci included data about 4,265 primary collections (plus 13,865 subcollections embedded within the primary 83 ariño – approaches to estimating the universe of natural history collections data box 3: generalizing the seber model collections in the universe may have been unlisted (r0), listed in repository a (ra), in repository b (rb), in c (rc), etc., of l repositories, and therefore may have been also simultaneously listed in more than one repository (for example, rab or rabc). assuming that listings are independent from each other; that there are m(x: 0, a, b, c, ..., ab, ac, ..., abc, ...) possible outcomes; and that a collection has a probability θx of belonging to the outcome x, all possible outcomes follow a multinomial distribution ... m m r x θ . !!...! !)|,...,,( 0 0 0 0 x rr mx mx rrr rrrrrf θθ= which is the generalized version of the particular case for four outcomes proposed by seber and felton (1981, eq. 11). therefore, similar substitutions can be made as in that case, resulting in a generalized analogue where the correction factor is the product of the probabilities for a collection of not belonging to one repository when belonging to any other. thus, the case with l=3 has m=8 outcomes and is ct abba bt acca at bccb rr rrr rr rrr rr rrr k − ++ − ++ − ++ = where rt is the number of collections belonging to at least one repository. the general case is { } { } { }( ) { } ∏ ⎥ ⎦ ⎤ ⎢ ⎣ ⎡ − +++− = ¬¬¬ lbas st labcslcjslbsislasit rr rrrr k ,..., ...,...,...,... ... . collections, and other data sources), of which 58.6% declared specimen holdings, or whose holdings numbers could be obtained from other sources, accounting for 640 million specimens. biocase and other initiatives sought to discover how many collections (or units) existed already in digital form. for the purposes here, it was interesting to know how many collections had already reported metadata in gbif. a simple crosscheck of tables was not possible, as the common data field did not abide to a common structure. however, by taking a sample of bci records one could estimate the number of collections already present at least in part in gbif indexes, and therefore having been digitized to some extent. thus, a random subsample of 197 bci collections was manually searched in the entire gbif database 19 collections of the sample contributed data to gbif, or a 9.64% rate of findings in the sample. conversely, about 90% of the sample consisted of collections not contributing data to gbif. therefore, if the sample was representative, 90% of bci, or 3.8 k (thousand) already known collections would be either not contributing records, or not exist in a digitized form. conversely, one may ask how many of gbif’s datasets pertaining to collections were not in bci. thus, one may derive the degree of completeness of bci: this would show collections already in digital form but unknown to bci. about 320 gbif datasets were examined, of which only 16.5% included actual collection data. 41.5% of gbif collection datasets in the examined subsample corresponded to bci records. from this degree of incompleteness the number of completely “obscure” collections (neither in bci nor gbif) could in turn be estimated. seber’s (1982, in krebs, 1999) 84 ariño – approaches to estimating the universe of natural history collections data probabilistic model for mark loss in capturerecapture methods provided an interesting analog that could be readily adapted. collections would belong to one or the other set (bci or gbif), both, or none. the three first classes, known from the sample, could be used to estimate the fourth (unknown) case, as shown in box 2, therefore obtaining an estimate of the total number of collections (both known and unknown) in the subsample, from which the total number of collections in the sample (in our case, the universe of collections) can be extrapolated. the resulting number of extant collections is 8.5 k, of which all of gbif’s and at least 10% of bci’s would have digitized, contributed data. this would leave out some 6 k collections to be potentially contributed or digitized, including 1.9 k totally obscure collections known to neither bci nor gbif. seber and felton (1981) discussed the confidence intervals and bias of their estimate, that can be large for small overlaps. a possible extension of this probabilistic model could involve further datasets, thus increasing precision. this might require trying to generalize the seber model to more than two overlaps (box 3). as a test of concept, marginal data from the survey performed by the gsap-nhc tg (see macklin et al., this volume) were sampled for known collections. a 14.2% subsample of the survey’s results was thus searched both in bci and gbif, and conversely, the bci and gbif samples were searched in the survey. a number of candidate collections emerging in the survey sample could be found in neither previous sample, whereas many were common. applying the generalized model, the estimated number of extant collections remained at 8.4 k. assuming that the survey was also random and independent from bci and gbif entries, this close result seems to suggest that the crossover of data from independent listings may provide an adequate estimate of the size of the universe of collections. if these results hold through refinements of the model, the use of further independent listings, larger samples, and (possibly) a specificallydesigned survey, the crossover of listings could yield more precise figures but would nonetheless point to the existence “in the wild” of a significant number of yet unknown collections. distribution model declared holdings in bci and known records in gbif can be used to determine the distribution model of the records. from this model, one can try to derive an estimate of the total number of elements (collections) in the set. the approach i have followed here implies looking at the collections as categories having a given frequency (the number of records). this is akin to considering collections as operative taxonomical units (otus), the “sample” being the full inventory of known collections having at least one unit. one may explore the whittaker diagram (in krebs, 1999) of the log number of records in each collection for bci against their rank by magnitude (fig. 1). the plot very strongly resembles a broken-stick distribution (macarthur, 1957). however, it should be noted that the whittaker plot drops significantly at e8.5, which corresponds to collections having less than 5,000 specimens. berendsohn (pers. comm.) pointed out that data in index herbariorum (a main source for bci) has a cutoff at that size. therefore, the actual plot could also be suggestive of a lognormal model (inverted s-curve) if the sharp drop is an artifact of the bci selection. data from the gsap-nhc survey (see appendix, fig. 1) may support this hypothesis, as the sizes of collections were reported by their curators without limitations. in addition, the steep curve at the left of the distribution may represent the relative effort in bci to capture the metadata of the largest collections, which would not have been sampled but censused. if the lognormal distribution does indeed underlie the distribution of collection sizes (which could eventually be tested by having an unbiased sample if collections of all sizes were present), one could predict the number of unknown collections (i.e. missing from bci, belonging to the veiled sector of the distribution) by estimating the corresponding lognormal distribution parameters. the best fit would be achieved with a spread of 0,2295, which in turn yields a 39% veiled sector for the distribution. this sector should now be added to the full dataset for an upper limit of the complete universe, resulting 5.9 k extant 85 ariño – approaches to estimating the universe of natural history collections data y = -0,0022x + 13,01 r2 = 0,9198 y = -0,0017x + 12,146 r2 = 0,9946 0 2 4 6 8 10 12 14 16 18 0 1000 2000 3000 4000 5000 6000 7000 order of the collection ln (s pe ci m en s) collections. as we know from gbif data that digitization rate is 9.64%, according to this model 5.4 k potential collections would be not digitized. an alternate, simpler, but perhaps more robust approach (especially when taking into account the exponential effects of small differences when using logs), could take the central interquartile range of the distribution data, fitting it a lineal regressive model (fig. 1). this should reduce the bias caused both by the size cutoff in bci due to index herbariorum, and the weight from the largest collections, admittedly not sampled but censused. the intercept with the x axis, representing the order of the collection by number of specimens, should encompass all existing collections following a similar distribution. to these, one should add the zero-unit collections, but the model should not be applied to them. thus, from the regression coefficients one can calculate 8.6 k extant collections, of which 7.8 k potential collections would not be digitized. note the similarity of these estimates to the ones coming from the probabilistic model above. fig. 1: plot of log of number of specimens from bci collections against their rank when ordered by size. blue dots suggest a truncated lognormal distribution with a cutoff at units=5000. red dots are the interquartile range. the upper regression line (r2=0.99) corresponds to this interquartile range, and the lower (r2=0.92) to the full range. cases a third approach should derive from known, specialized cases where the actual number of collections is known through specialized literature, as well as the number of these collections that have emerged in general repositories or listings. a previous work (ariño, unpublished) dealing with invertebrate collections showed that 16 collections of collembola could be accessed through webbased searches, biocase, gbif or other databases, whereas a specialized publication listed 86 ariño – approaches to estimating the universe of natural history collections data 185 collections not known elsewhere. this 1115% “specialty” rate, if applied to the then known (in general repositories, such as samuelson & evenhuis’ codens) invertebrate collections, should yield a worst-case of 6.8 k hidden invertebrate collections. other groups could be treated similarly, i.e. identifying specialized publications citing collections not known elsewhere. in this case, this 6.8 k figure should be complemented with similar data for entomological, plant, and other general natural history collections, yielding a larger (albeit unknown) figure for which 6.8k would simply represent the lower bound. paradoxically, this approach seems the least robust one given that it depends heavily on the ratio between “sampled” data in general repositories and full datasets in comprehensive repositories for specific groups. however, it should be noted that, theoretically, a complete register of all existing collections should bring this ratio to 1, and then all uncertainty should disappear, therefore having the potential to become the most robust possible estimate (in fact, not an estimate but a census). specimens published estimates about the number of extant specimens are in the 2 3 g range (duckworth et al., 1993; oecd, 1999, in chapman, 2005). one can independently try to refine and contrast this against actual and modeled data. currently, bci lists collections that, either as declared by bci contributors or obtained elsewhere, amount to 640.5 m (million) specimens. gbif’s gsap-nhc survey returned 833.7 m specimens reported by respondents after removing duplicated data. from these data, their distribution, gbif distribution, number of collections, and other experimental data one may try to derive the total amount of data in collections, and what fraction has not been digitized or is not accessible through gbif. distribution data only 58.6% of bci collections declare holdings. assuming that the underlying distribution data of the holdings across the undeclared collections was similar, one would then have 1.09 g specimens from the known bci collections alone. however, as i have shown before, there may also be a number of “obscure” collections. for instance, the probabilistic model returns a 23% to 29% of unknown collections. 0% 20% 40% 60% 80% 100% google msn yahoo ask/teoma exalead gigablast natural history collections invertebrate collections entomology collections vertebrate collections herbarium fig. 2. distribution of number of results returned by some search engines when queried for certain keywords (legend: keywords used) in 2006. 87 ariño – approaches to estimating the universe of natural history collections data taking these into account, and assuming that their size distribution would be similar to the known distribution, a bracket may result from 1.40 g for the number of extant collections calculated by the lognormal model, to 2.01 g specimens if using the probabilistic model. the number of occurrences involving specimens from actual collections already in gbif can be estimated between 59 m and 96.1 m, depending upon how certain values given as basis of record are interpreted. if occurrences in gbif can be equated to specimens in bci (which is not always true) the remaining digitizable mass would bracket between 1.27 g and 1.97 g specimens. it should be noted that the assumption of similar underlying distribution of specimens across collections, which is required in this approach, has been regarded as a weak assumption (berendsohn, pers. comm.). besides the fact that certain requirements apply for insertion into one main source of bci (namely, active curation and minimal size), berendsohn points out that “it is correct to assume that most of the big collections declare at least an estimate of total numbers, while small collections either hide their numbers or give very exact figures.” fully testing this would require knowing the actual numbers of records in every collection (as opposed to their published numbers), which does not seem practical. however, an indirect approach is possible by looking at the relative precision with which numbers were reported. the implicit precision (the relative number of significant figures when reporting a collection size) tends to veer in the way predicted by berendsohn, although the effect seems too shallow to affect the estimates greatly (see appendix, figure 2). rate of digitization in gbif-mobilized data a similar approach but using gbif-mobilized data can be derived from digitized specimen counts. a random sample of 132 bci collections having declared numbers of specimens was searched in gbif indexes, and the amount of corresponding records in gbif tabulated. for all paired matches, 4.67% of the declared records did exist in gbif. if this is the rate of digitization in data received and distributed by gbif, then this rate, applied to the actual number of specimens s/l lots arthropods 4 0,08% bacteria 1 0,02% data aggregator/indexer 1 0,02% dna bank 1 0,02% entomology (insects/spiders) 1303 26,10% 257483000 40,20% 1,0 257483000 herbarium 3006 60,20% 342725710 53,51% 1,0 342725710 herpetology (amphibians and reptiles) 53 1,06% 2662300 0,42% 2,3 1158181 ichthyology (fishes) 10 0,20% 6861238 1,07% 1,9 3579002 images 1 0,02% 2000 0,00% 1,0 2000 invertebrates 1 0,02% 550000 0,09% 9,0 60956 library/archives 1 0,02% 500000 0,08% 1,0 500000 mammals 12 0,24% 315641 0,05% 1,5 212899 natural history museum (diverse collections) 69 1,38% 9802000 1,53% 7,4 1331793 not specified 486 9,73% 1379481 0,22% 1,0 1379481 ornithology (birds) 31 0,62% 4181408 0,65% 1,0 4043915 paleontology 1 0,02% 300000 0,05% 1,0 300000 sound/film 1 0,02% vertebrates 3 0,06% 1754000 0,27% 1,0 1754000 zoology 8 0,16% 12000000 1,87% 1 12000000 total 640516778 626530938 collections specimens table 1. estimated number of curatorial units from collections in bci assuming that the mean lot size is as in the example case (mzna, fig. 3). 88 ariño – approaches to estimating the universe of natural history collections data from collections in gbif (est. 59 – 96.1 m), should estimate the amount of unacquired specimens “in the wild”, either declared in bci or not. the resulting figures are 1.26 g – 2.06 g specimens. units the most relevant figure for the digitization goal would possibly be that of the digitizing unit. generally, this will coincide with the curatorial unit, although it is not always the case (i.e. multiple data from a single curatorial unit, such as successive exiccata or multi-taxon slides). it may also coincide with specimen data, especially for plants or pinned insects, but much less so for 1 mzna is the coden for the museum of zoology of the university of navarra, pamplona, spain. url: http://www.unav.es/unzyec/eng/ certain animal groups such as wet collections of small invertebrates or uncountable specimens. the expected number of units, thus, can be either sampled from known cases, or derived from the expected number of specimens or collections based on their types thereof. for that, a distribution of collections and specimens in units may be necessary. data from biocase, bci or gbif can be used to estimate the approximate composition of the types of units to be expected. also, the composition for data not yet covered in these repositories can be indirectly estimated by using search engines. 60.2% of bci collections are fig. 3. distribution of the composition of lots in a case study of a zoological collection (mzna1, approx. 200 k accessions, 2 m specimens). this refers to already digitized data only. pinned entomological collections were not included. <2 <8 <32 <128<512<2048<8192 <32768 mammals birds herps fish insects (w et) mites, ticks, s piders springtails, thrips shells & snails worms roundworms 0% 20% 40% 60% 80% 100% specimens per lot (binary log scale) 89 ariño – approaches to estimating the universe of natural history collections data table 2 summary of various size estimates of extant collections along with known data, ordered by estimate. concept source estimates collections bci (known) herbaria and 26% insect collections, reflecting the initial population from index herbariorum and bishop’s museum codens. search engines yield similar compositions, with a dominant number of references to herbaria (fig. 2). however, entomological collections amount for 40.2% of the specimens. as these do not distinguish pinned collections (mono-specimen) from wet collections (generally multi-specimen), it is therefore difficult to estimate the correlation for this quarter of data. we may derive an estimate from the analysis of an example of such a collection. reporting collections may have used interchangeably specimen and unit, and thus the number of specimens per lot of this analysis may be used as a lower bound, while a 1:1 relation would represent the upper bound. table 1 shows bci and analytical data from a case study collection (fig. 3, mzna collection). entomological collections have been assumed to be pinned (1:1), although it should be noted that a typical lot composition for entomological wet collections at mzna is about 16 specimens per lot. after applying both the distribution of lots and its composition, we obtain a correction factor of 1,022. therefore, the number of units in collections can be estimated to range, depending on the type of estimate selected, from 1.23 g to 1.97 g units. strategy it has been argued that the rate of accrual currently exceeds the available capacity for digitizing (beaman, unpublished), although this clearly depends on the amount of people and resources appropriated for the task. these numbers should provide a baseline for gauging the effort needed in order to both reach the 2 billion record goal, and to have most of the metadata from natural history collections mobilized and readily accessible. the highest estimate (2,06 g 4265 gbif (known) 2346 log-normal model 5941 case study (invertebrate only) 6779 probabilistic model from metadata 8372 8528 regressive model 8625 specimens bci (known) 640.5 million (m) gbif (known) 59 96.1 m gsap-nhc survey (known) 833.7 m gbif nodes reports (known) 407 – 737 m distribution data 1.61 – 2.04 billion (g) digitisation rate 1.26 – 2.06 g units composition of collections 1.23 – 1.97 g 90 ariño – approaches to estimating the universe of natural history collections data specimens) yields more than twenty times the amount of collection records already in gbif, and it should be noted that of the 20 largest known collections (amounting to one third of the total amount of specimens), almost one third had no data in gbif as of october 2008, while the remaining contributed less than one-tenth of their holdings (appendix, table 1). also, the large proportion of herbaria in the known lists of collections may not be mirrored in reality, as zoological collections (except vertebrates) have more slowly started to share data or to be digitized at all: zoological material lends itself less easily to automated digitization procedures, due to very varied physical layouts. a careful identification of the most critically-needed metadata should translate into an adjusted estimate of the human resources to be allocated for this task. while the above preliminary estimates decrease significantly the previously published estimated size of the digitizable universe, it should be taken into account that these figures are initial estimates, whose precision could be increased through more refined modeling and more extensive data collection. this paper suggests ways to estimate a more reliable range of size, better suited for planning and allocation of resources, and for obtaining certainties. as more records become available, and in particular as all available records are processed, narrower estimates could be obtained. for example, the following activities could all help to refine the current estimates: execute a wider survey among respondents to the gsap survey and bci curators with the specific purpose of getting the approximate numbers of digitized data; derive from that survey data to help estimating not only the universe of collections but also the depth of ongoing digitization; expand the analyses to use the fullest available datasets, including subcollections within bci, through allocation of resources for manual checking; reanalyze all data independently by types of collections. with better estimates, perhaps a way to allocate efforts for digitisation can be found, for the goal of having quality biodiversity data available for research. acknowledgements.the author is indebted to walter berendsohn, james macklin and two anonymous referees that provided very valuable criticism and comments, and to gbif for setting up the gsap-nhc task group that prompted part of this research. references berendsohn, w.g., m.j. costello, c. emblow, a. güntsch, a. hahn, j. koenemann, c. thomas, n. thomson and r. white, 2000. concepts for a european portal to biological collections. in: berendsohn, w.g. (ed.): resource identification for a biological collection information service in europe, botanic garden and botanical museum berlin dahlem, isbn 3-921800-44-7 2 berendsohn, w.g., m. döring, m. gebhardt and a. güntsch, 2002. biocase a biological collection access service for europe. trends and developments in biodiversity informatics. symposium: key innovations in biodiversity informatics, indaiatuba, sp, brasil 2002. 3 [bci] biodiversity collection index [continuously updated]. http://www.biodiversitycollectionsindex.org/static/i ndex.html (accessed july 30, 2010) chapman, a., 2005. uses of primary species occurrence data, version 1.0. copenhagen: global biodiversity information facility. 106 pp. isbn: 87-92020-01-14 [gbif] global biodiversity information facility, 2003. global biodiversity information facility annual report 2001-2002. copenhagen.5 [gbif] global biodiversity information facility, 2006. gbif strategic and operational plans 2007-2011: from prototype towards full operation. copenhagen.6[gbif] global biodiversity 2 http://www.bgbm.org/biocise/publications/results/11.htm 3 http://www.cria.org.br/eventos/tdbi/bis/biocase 4 http://www2.gbif.org/usesprimarydata.pdf 5 http://www2.gbif.org/annual_report_2001_2002.pdf 6 http://www2.gbif.org/strategic_plans.pdf 91 http://www.bgbm.org/biocise/publications/results/11.htm http://www.cria.org.br/eventos/tdbi/bis/biocase http://www2.gbif.org/usesprimarydata.pdf http://www2.gbif.org/annual_report_2001_2002.pdf http://www2.gbif.org/strategic_plans.pdf ariño – approaches to estimating the universe of natural history collections data 92 information facility, 2008. gbif work programme 2009-2010. copenhagen.7 [gbif] global biodiversity information facility, 2008b. gbif annual report 2008. copenhagen.8 [gbif] global biodiversity information facility [continuously updated]. data discovery and mobilisation. http://www.gbif.org/informatics/primary-data/datadiscovery-and-mobilisation/ (accessed july 30, 2010). [gbif] global biodiversity information facility [continuously updated]. global plan for natural history collections data. http://www.gbif.org/informatics/primary-data/taskgroups/gsap-nhc/ (accessed july 30, 2010). güntsch, a., a. hahn and w.g. berendsohn, 2001: biological collection databases in europe. pp. 917 in: riede, k. (ed.), new perspectives for monitoring migratory animals improving knowledge for conservation. bfn, zef, bonn. hahn, a. 2000. information resources the biocise survey. resource identification for a biological collection information service in europe. results of the concerted action project. http://www.bgbm.org/biocise/publications/results/ 7.htm (accessed july 30, 2010). hill, a., j. otegui, a.h. ariño and r. guralnick, 2010: gbif position paper on future directions and recommendations for enhancing fitness-for-use across the gbif network, version 1.0. copenhagen: global biodiversity information facility, 25 pp. isbn: 87-92020-11-9. http://www.gbif.org. 7 http://www2.gbif.org/wp2009-10.pdf 8 http://www2.gbif.org/annual_report_2008.pdf krebs c.j., 1999: ecological methodology, 2nd ed. addison-wesley, krishtalka, l., and p.s. humphrey, 2000. can natural history museums capture the future? bioscience, 50 (7): 611-617. mac arthur, r.h., 1957. on the relative abundance of bird species. proc. natl. acad. sci. 43, 293-295. [oecd] organisation for economic co-operation and development, 1999. final report of the oecd megascience forum working group on biological informatics. paris: oecd publications. 9 samuelson, a. and n. evenhuis, 1998+: abbreviations for insect and spider collections of the world. bishop museum, honolulu. seber g.a.f. and r. felton, 1981. tag loss and the petersen mark-recapture experiment. biometrika, 68 (1): 211-219. seber g.a.f., 1982: the estimation of animal abundance and related parameters. the blackburn press, new york. thiers, b. [continuously updated]. index herbariorum: a global directory of public herbaria and associated staff. new york botanical garden's virtual herbarium. 10 (accessed july 30, 2010). 9 http://www.oecd.org/dataoecd/24/32/2105199.pdf 10 http://sweetgum.nybg.org/ih/ http://www2.gbif.org/wp2009-10.pdf http://www2.gbif.org/annual_report_2008.pdf http://www.oecd.org/dataoecd/24/32/2105199.pdf http://sweetgum.nybg.org/ih/ microsoft word ketal_taxongrab_2005_final.doc biodiversity informatics, 2, 2005, pp. 79-82 79 taxongrab: extracting taxonomic names from text drew koning1, indra neil sarkar1,2, and thomas moritz1,3 1divisions of library services and 2invertebrate zoology, american museum of natural history, new york, ny 10024 usa 3e-mail: tmoritz@amnh.org abstract.––identification of organism names in biological texts is essential for the management of archival resources to facilitate comparative biological investigation. because organism nomenclature conforms closely to prescribed rules, automated techniques may be useful for identifying organism names from existing documents, and may also support the completion of comprehensive indices of taxonomic names; such comprehensive lists are not yet available. using a combination of contextual rules and a language lexicon, we have developed a set of simple computational techniques for extracting taxonomic names from biological text. our proposed method consistently performs at greater than 96% precision and 94% recall, and at a much higher speed than manual extraction techniques. an implementation of the described method is available as a web based tool written in php. additionally, the php source code is available from sourceforge: http://sourceforge.net/projects/taxongrab, and the project website is http://research.amnh.org/informatics/taxlit/apps/. key words.––named entity recognition; taxonomic name extraction introduction new and revised biological names are often embedded within the conventional biological and medical literature. while some taxonomic names are available via popular indexing services (e.g., zoological record), complete names data are only available in the form of printed articles. as the full legacy of this printed literature is digitized, automated techniques will be in demand to assist with the extraction and indexing of these, named entities. a range of computational techniques, categorized as named entity recognition (ner) techniques, exist for identifying named entities (cunningham et al. 2002). ner seeks to locate and classify atomic elements in text into predefined categories. for example, ner techniques have shown utility in identifying gene names from biomedical articles (krauthammer et al. 2004). organism names represent another common named entity that may be usefully extracted, particularly in cases where the literature consists of information pertaining to organism biology or biodiversity (soberón et al. 2004). organism names within taxonomy generally occur in natural language text as sequences of two or three words, called “binomen” and “trinomen” taxonomic names, respectively. taxonomic names also follow a prescribed set of linguistic and contextual rules (linnaeus 1753): the scientific name of an organism is written in either latin or greek. genus name precedes species name, and is capitalized. species name, and any subsequent subspecies, variant, or strain names are written in lower case. as organisms are discovered and described in scientific literature, these rules are mandated and prescribed by international commissions (e.g., iczn and icbn for zoological or botanical organisms, respectively). a commonly used approach in ner is to create dictionaries of terms that can later be referenced to identify known named entities (petasis et al. 2000). this approach has limited success, as it requires the term to pre-exist in the dictionary of terms. in order to create a comprehensive dictionary of taxonomic names for currently recognized species, one would currently need to craft a dictionary consisting of 1.5-1.8 million items (wilson 2003). with the number of new organisms identified increasing, in pace with advances in collection and description methods, this number could, ultimately, range from 30 to 100 million (wilson 2003). as a result, a comprehensive and current catalogue of taxonomic names will be needed to create reliable look-up and indexing algorithms. the linguistic and contextual nature of taxonomic names, as dictated by linnaean rules, enables the development of computational tools that can extract names from natural language text. since taxonomic names are most typically – though not exclusively – derived from greek or latin, filters can be used to separate taxonomic name candidates using a language-specific lexicon (e.g., a lexicon of english words). finally, after taxonomic name candidates have been identified, the linnaean conventions for capitalization (i.e., genus name is capitalized, followed by lowercase species and subspecies names) can be used to identify taxonomic name entities. here, we explore the applicability of ner methods for identifying taxonomic names from digitized texts. using a lexicon of words that we have compiled from existing english lexicons, we assess the efficacy of our method for extracting taxonomic names. we evaluate the proposed technique with respect to manually identified taxonomic names using a web-based interface. we conclude with some discussion of how this approach, which we term “taxongrab,” can be used in the design and development of future taxonomic literature organization systems. koning et al. taxongrab 80 building taxongrab the premise of the taxongrab approach is that taxonomic names can be identified from natural language text using a combination of taxonomic nomenclature rules and a lexicon of non-taxonomic terms. we composed a lexicon of english words by combining the terms from the wordnet® (fellbaum 1998) and specialist (mccray et al. 1993) lexicons. the wordnet lexicon contains common words that are associated with many facets of the english language. the specialist lexicon, which is part of the united states national library of medicine’s unified medical language system® (lindberg et al. 1993), consists of both common english and biomedical vocabulary words, including spelling variants and inflected forms (e.g., plural). an important consideration in building the language lexicon from existing resources was that many parts of taxonomic names have become part of the common language, and are therefore included in a complete lexicon; for example, coli (e.g., the second lexical unit of the binomen escherichia coli) is included in comprehensive language lexicons. therefore, in order to create a lexicon that did not contain any taxonomic terms (i.e., words used to describe genus, species, or subspecies), we manually culled words from a list of taxonomic terms drawn from popular taxonomic resources. specifically, we removed terms from our lexicon that were associated with 362,430 taxonomic names in either the national center for biotechnology information (ncbi) taxonomy, the integrated taxonomic information system (itis), or the german collection of microorganisms and cell cultures (deutsche sammlung von mikroorganismen und zellkulturen [dsmz]). the resulting lexicon consisted of 258,783 words. a script was written in php that used the resulting lexicon. the script isolates sets of two or three consecutive words that are not in the lexicon, as candidate names. these taxonomic name candidates are then validated according to the capitalization rules of linnaean nomenclature. additionally, strings involving a capital letter followed by a period and then at least one word that is not in the lexicon – a conventional form of abbreviation for binomens or trinomens that have been previously cited in a work (e.g., escherichia coli is often abbreviated as e. coli) – are also reported as taxonomic name candidates. the script also looks for more complex taxonomic formatting rules – such as variants preempted with var., subspecies preempted with subsp., or parentheses that are often used to indicate subgenus or author names. all of these rules were implemented in the script using regular expressions. a web interface was designed whereby users can enter text or upload text files. the interface then returns the list of taxonomic name candidates that are found using the taxongrab method. the interface allows one to enter text either by entering it directly, through uploading a text file, or specifying a web location. evaluation the web interface was used to examine a number of documents, consisting of archived publications, which were digitized using standard optical scanning and optical character recognition techniques (ocr. using an off-theshelf software package; abbyy finereader© that purports 97% accuracy. as a test corpus, we used the volume 1 of “the birds of the belgian congo” by james paul chapin (published in four parts in the series: bulletin of the american museum of natural history, 1932-1954). this corpus consists of 5000 pages that contain over 8000 taxonomic names. taxongrab was used via the web interface to extract names from the corpus. the extracted taxonomic names were then compared to a list of taxonomic names that had been manually identified by a team of experts. the results were then assessed using precision and recall values. precision, or the correctness of the reported taxonomic names, is defined as ratio of correct taxonomic names (tp) to the sum of correct and false taxonomic names (tp+fp): tp/(tp+fp). recall, or the ability to retrieve taxonomic names, is defined as the ratio of the sum of correct taxonomic names (tp) to the sum of correct and missed taxonomic names (tp+fn): tp/(tp+fn). compared to the manually extracted taxonomic names, taxongrab consistently identified taxonomic names with greater than 96% precision and 94% recall from the documents examined. errors arose mainly from ocr errors, manuscript typos, and the few common english words that are also used as scientific names, which had not been addressed in the lexicon creation. with respect to the speed of extraction, the manual extraction was reported to have taken 80 hours, while the automated method took approximately 330 seconds. discussion and future directions with the many advances in biological collection and description techniques, comprehensive and current catalogues of taxonomic names will increasingly be needed to support automated lexical lookup systems that are designed for organizing and aggregating textual data. currently, with only a small fraction of known organisms named, our taxonomic name catalogues are incomplete. testing just the three resources for taxonomic names described in this study, we found that there was only partial overlap between existing resources. this is probably attributable to differences in foci of the different catalogues – for example, while ncbi taxonomy is mostly concerned with organisms that are described in medline and have some biomedical significance, itis is more concerned with describing organisms in the context of governmental regulation and biodiversity information. at present, there seems insufficient investment in reconciliation of compiled names between biodiversity and biomedical resources. however, there are some links between resources, for koning et al. taxongrab 81 example, for itis taxonomic names there are links to appropriate ncbi taxonomy entries. however, there is not yet a centralized list containing all taxonomic names as they are identified and described. because taxonomic names generally (with the exception of many virus names) follow a set of prescribed syntax and linguistic rules, it is possible to create automated techniques to extract taxonomic names from literature resources. here, we have developed the taxongrab method, which leverages the linguistic and syntactic properties of taxonomic names. taxonomic names can be useful as index terms when organizing large sets of literature. to that end, the taxongrab method can be used to extract all the taxonomic names associated with documents, both legacy and prospective, which are subsequently used to organize them. because taxongrab does not rely on a particular taxonomic name catalogue, it seems an efficient tool in extracting and compiling new organism names for inclusion in suitable resources. in this way, taxongrab may be a tool that can be used for curating and updating taxonomic name catalogues. it is also possible that taxongrab can assist the editing of digitally captured taxonomic literature. because of their idiosyncratic properties, taxonomic names may not be easily treatable by normal corrective measures (standard dictionaries and spell-checkers) in ocr or other capture technologies. taxongrab may prove to be instrumental in rapidly identifying names for automated or semi-automated methods of proof reading and editing. we implemented taxongrab as a web-based interface written in php. the web interface enables one to search for taxonomic names that may exist in one of three forms (text, text file, web site). the source code and associated files are available for free download and can be modified for particular needs. using the taxongrab principle, larger queue application (for example, if implemented as a perl script), one could search through and organize a large set of documents for taxonomic names. this could be a useful utility that can address a number of important questions, such as “what taxa are described or mentioned in a particular corpus?” subsequent tools could be designed that organize the taxonomic names identified into an ontology, whereby one could track and register organism name changes. while there are implications of a taxonomic name ontology within existing resources (e.g., ncbi taxonomy is organized into a hierarchy of terms, that are also linked into larger upper-level ontologies), to date there are no specific projects focused on organizing taxonomic names and name changes according to an ontological framework. to that end, we are exploring the automated creation of hierarchic taxonomic name lists. as a list of taxonomic names is updated, we will be able to track the types and numbers of various organisms that are described in various different types of corpora. for example, we could consider the difference between what types of organisms are discussed in biodiversity resources versus biomedical resources. identifying organisms that are described in both can be used as a means to unify knowledge in some areas (e.g., medical entomology). in contrast, identifying organisms that are different between biodiversity resources and biomedical resources will highlight key differences between research foci, but may underscore the importance of studying other related organisms of the same genus for clues (e.g., drosophila studies may want to consider species besides just melanogaster). taxongrab was designed to extract taxonomic names from english texts, as reflected by the large english lexicon that was constructed. however, our web interface currently enables one to search for taxonomic names in, spanish, german, and french documents using very basic dictionaries (compiled from the winedit text editor language dictionaries). we anticipate development of nonenglish versions of taxongrab using available non-english lexicons resources (e.g., eurowordnet (vossen 1998)). conclusion the identification of taxonomic names within published literature can help guide comparative biological investigations. here, we have proposed a named entity recognition technique, taxongrab, which is based upon the systematic nomenclature rules conventionally used for taxonomic nomenclature in scientific publications. we believe that this method may show general utility for indexing documents containing embedded taxonomic names. acknowledgments the authors thank members of the project team (national science foundation: iis-0241229) for their comments towards the work described. the authors especially thank norm johnson, donat agosti, and liz nichols for their assistance with the manual extraction of taxonomic names. ins is partially supported by national science foundation: dbi-0421604 and the lewis b. & dorothy cullman program for molecular systematics. we also thank researchers from the marine biological laboratory at woods hole, massachusetts for their comments regarding the development of other extraction tools based on taxongrab. literature cited chapin, j. 1932. the birds of the belgian congo: part 1. bulletin of the american museum of natural history: 65. american museum of natural history, new york. chapin, j. 1939. the birds of the belgian congo: part 2. bulletin of the american museum of natural history: 75. american museum of natural history, new york. chapin, j. 1953. the birds of the belgian congo: part 3. bulletin of the american museum of natural history: 75a. american museum of natural history, new york. koning et al. taxongrab 82 chapin, j. 1954. the birds of the belgian congo: part 4. bulletin of the american museum of natural history: 75b. american museum of natural history, new york. cunningham, h., d. maynard, k. bontcheva, and v. tablan. 2002. gate: a framework and graphical development environment for robust nlp tools and applications. association for computational linguists, philadelphia. fellbaum, c. 1998. wordnet: an electronic lexical database. cambridge, mit press. krauthammer, m. and g. nenadic. 2004. term identification in the biomedical literature. journal of biomedical informatics 37:512-26. lindberg, d. a., b. l. humphreys, and a. t. mccray. 1993. the unified medical language system. methods of information in medicine 32:281-91. linnaeus, c. 1753. species plantarum, stockholm. sweden. mccray, a. t., a. r. aronson, a. c. browne, t. c. rindflesch, a. razi, and s. srinivasan. 1993. umls knowledge for biomedical language processing. bulletin medical library association 81:184-94. morgan, a. a., l. hirschman, m. colosimo, a. s. yeh, and j. b. colombe. 2004. gene name identification and normalization using a model organism database. journal of biomedical informatics 37:396-410. petasis, g., a. cucchiarelli, p. velardi, g. paliouras, v. karkaletsis, and c. d. spyropoulos. 2000. automatic adaptation of proper noun dictionaries through cooperation of machine learning and probabilistic methods. proceedings of the association for computing machinery: special interest group on information retrieval, athens, greece. soberon, j. and a. t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data philosophical transactions of the royal society of london b, 359:689-698 vossen, p. 1998. euro wordnet: a multilingual database with lexical semantic networks. norwell, kluwer academic publishers. wilson, e. o. 2003. the encyclopedia of life. trends in ecology and evolution 18:77-80. microsoft word downey_finallayout1_200907151-lld biodiversity informatics, 6, 2009, pp. 18-27 bridging the gap between technology and science with examples from ecology and biodiversity laura l. downey and deana pennington university of �ew mexico, lter �etwork office department of biology, msc03 2020 university of �ew mexico albuquerque, �m 87131-0001, laura.downey@dhs.gov, dpennington@lternet.edu abstract. — early informatics initiatives focused primarily on the application of technology and computer science to a specific domain; modern informatics has broadened to encompass human and knowledge dimensions. application of technology is but one aspect of informatics. understanding domain members’ issues, priorities, knowledge, abilities, interactions, tasks, and work environments is another aspect, one that directly impacts application success. involving domain members in the design and development of technology in their domain is a key factor in bridging the gap between technology and science. this usercentered design (ucd) approach in informatics is presented via an ecoinformatics case study in three areas: collaboration, usability, and education and training. keywords. — collaboration, ecoinformatics, science and technology, training, usability, user-centered design. the french term “informatique” first appeared in the 1960s, and referred to “the application of computing to the communication processes used by scientists in exchanging information and data among themselves.” 1 computing applications, adoption and usage remained the primary focus of informatics programs for many years. in the last decade, the focus has expanded to include cognitive, social and human dimensions. for example, at the school of informatics at the university of edinburgh, informatics is defined as “the study of how natural and artificial systems store, process and communicate information.” 2 the indiana university south bend informatics program defines informatics along three dimensions: 3 • understanding the impact that technology has on people • the development of new uses for technology • the application of information technology in the context of another field 1 http://informatics.buffalo.edu/school/informatics.asp. 2 http://www.inf.ed.ac.uk/about/. 3 http://www.informatics.iusb.edu/. within this broad perspective of informatics that encompasses humans, technology and knowledge, the science environment for ecological knowledge (seek) project (michener et al. 2007) bridged the gap between technology and science by employing a user-centered design approach. informatics began with scientists, and the usercentered design approach brings scientists to the forefront of the application, design, and development of technology in science. user-centered design user-centered design (stone et al. 2005) is a design approach to creating systems, software, and technology that actively seeks to understand and involve the target users in the design and development process. it includes understanding their abilities, knowledge, needs, and concerns, as well as their interactions, tasks, and work environments. example methods in ucd are surveys, field studies, structured feedback sessions, task analysis, and usability testing of software with target users. many ucd techniques have been applied in the process of developing the seek project. the benefits of ucd are well documented in the human factors and usability literature, in downey & pennington bridging the technology & science gap with ecology examples 19 books, and in many consultants’ white papers. as one example, bias and mayhew (2005) contains a nice cross-section of case studies and research that detail returns on usability investments. many corporations have come to realize the return on investment that the ucd process offers. the fields of human factors and industrial engineering and usability engineering have increased significantly in the past 20 years. there is an international standard (iso 13407) 4 which outlines the general process of user-centered design. basic business benefits, as summarized by the usability professionals association, are: • increased productivity • increased sales and revenues • decreased training and support costs • reduced development time and costs • reduced maintenance costs • increase customer satisfaction research teams that include ucd when developing software for scientists benefit from less resources being required for development and maintenance; and scientists who are the users will benefit from increased productivity. however, we believe strongly that the importance of ucd in science goes beyond these basics. the ucd process is important to science because: • technology should make a scientist’s job easier, not get in the way • technology should help scientists work faster, better and smarter • technology should be easily exploited by scientists to enable new analyses and discoveries there are various models of ucd, some with specific details, and some providing frameworks. figure 1 offers an example ucd process followed by the usability and accessibility center of michigan state university 5 . it has five phases and also demonstrates the inherent iterative nature of ucd. a variety of activities can be selected for each phase. 4 iso 13407 is an international standard produced by the international organization for standardization. the subject of the standard 13407 is “human-centred design processes for interactive systems. (see: http://www.iso.org/iso/iso_catalogue/catalogue_tc/catalogue_detail.ht m?csnumber=21197) 5 http://usability.msu.edu/approach.asp the seek project the seek project is a multi-institutional collaboration sponsored by the national science foundation that is (1) creating cyberinfrastructure and applications for ecological, environmental, and biodiversity research, and (2) educating the ecological community about ecoinformatics. figure 2 provides an overview of the seek architecture which illustrates knowledge, technology and community all working together to enable science and promote collaboration. a major focus of seek is the enhancement of a scientific workflow application called kepler (altintas et al. 2004, ludäscher 2006). kepler is an open-source modeling and analysis tool for creating, visualizing, executing, and documenting scientific workflows. a scientific workflow is a collection of data flow and analytical steps that formalizes the research process. two real-world research problems, one from ecology and one from biodiversity science are being pursued to demonstrate how technology can enhance and enable scientific research. semantic mediation is a primary infrastructural research area, with the goal of providing enhanced machine operations, relieving the manual computational analysis and data integration burden from scientists. the purpose of the semantic mediation system (sms) (bowers et al. 2004, bowers and ludäscher 2003, 2004) in kepler is to (1) help scientists discover relevant data and processing components for use in constructing scientific workflows, (2) automate or semi-automate the merging of heterogeneous data sets, and (3) perform automatic transformation in a scientific workflow. this technology application and usage is not done in a vacuum. scientists interact with the technology to achieve their objectives and therefore should be part of the design and development process so that the technology can be fully and usefully exploited. this paper uses the seek project as a case study to report on bridging the gap between technology and science with a focus on the user centered design approach. downey & pennington bridging the technology & science gap with ecology examples 20 figure 1. user-centered design process followed by the usability and accessibility center, michigan state university. building bridges one of the initial steps that the seek project took to ensure that scientists’ concerns and issues would be integral to the project was to put some of those scientists on the seek team. their perspective is an integral part of application design, helping to shape and formulate current and future directions. the cross-disciplinary team forms the foundation of our bridge between science and technology. the seek project built upon that foundation by pursuing additional collaborations with the scientific community, applying usability engineering techniques to its technology products, and conducting education and training in ecoinformatics. collaboration early in the project, seek initiated collaborative activities with two groups of scientists. each group had real-world research problems of interest to the project because their basic needs were in technical research areas targeted by seek. interactions with these communities allowed developers to delineate needs more clearly, focus development on specific articulated problems, and test solutions against real problems. the test beds were absolutely critical for grounding theoretical solutions devised by computer science researchers within the more comprehensive solutions developed to meet the full set of requirements. ecological niche modeling research group ecological niche modeling is an approach for understanding and predicting species’ present geographic distributions (nix 1986, carpenter et al. 1993) and distributions under scenarios of change, including transplantation to another continent as invasive species (beerling et al. 1995, peterson and vieglais 2001) or under changed climatic conditions (martínez-meyer et al. 2004, araújo et al. 2005). numerous conceptual approaches and software tools can be used in ecological niche modeling (nix 1986, walker and cocks 1991, carpenter et al. 1993). seek selected this community for its prototype application because the global-scale analyses in which they are engaged were complex and required substantial manual data discovery and processing. as such, clear gains could be made through applying cutting-edge technology. the group recruited for collaboration included ~20 leading researchers from around the world (united states, mexico, europe, south africa, australia, new zealand, etc.). seek held three working meetings with this group: (1) a full group meeting where analysis and modeling tasks were discussed and important datasets, algorithms, and downey & pennington bridging the technology & science gap with ecology examples 21 analytical environments were identified, (2) a small working group meeting where a specific research problem was fully specified from a workflow perspective, and (3) a full group meeting where prototypes were tested for appropriate functionality and usability (downey 2007). through the first two meetings, technical problems and tasks associated with ecological niche modeling were identified as: • labor-intensive data preprocessing and preparation • distributed data, archived in dozens of museum collections and environmental repositories • heterogeneous computing environments required in the workflow (e.g., c and java programs, scripts in matlab, sas and r, and gis analyses. • compute-intensive algorithm execution for multiple species, under multiple parameter sweeps, and when the algorithm is stochastic many iterations to construct a distribution of results for statistical comparison. these needs of the ecological niche modeling community drove development of the kepler workflow system (pennington et al. 2007), a problem-solving environment making use of visual modeling to construct workflows that access distributed data within a grid system and which allows integration of heterogeneous computing environments. biodiversity analysis research group biodiversity is a measure of the numbers and kinds of species occurring at a particular place and time. patterns of biodiversity are known to occur correlated with climate, productivity, evolutionary history, glacial history, and a host of other environmental characteristics. biodiversity is in documented decline worldwide (barbaut and sastrapadja 1995, pimm and raven 2000), but understanding the causes and consequences of this depends on integrating all of the relevant ecological characteristics with all available species information over broad areas (waide et al. 1999). most of the relevant data was collected by independent researchers working at small field sites over decades. hence, most biodiversity researchers invest a good deal of effort in simply gathering and integrating the available data to assess where gaps exist that could be addressed through collecting additional data. this group recruited to work with seek had a very different flavor than the previous group. development of solutions to help resolve semantic problems requires a truly interdisciplinary approach to understanding and representing domain concepts in ways that are formal and computationally tractable. the necessary level of engagement with technical personnel is fairly high; therefore, domain experts with at least some demonstrated technical skill were recruited. we held two small group working meetings with these scientists, where we meticulously examined multiple datasets and discussed the meaning of every attribute in those datasets, where they were derived from, and how they would need to be modified for integration. the group provided scripts that were used to manipulate data in prior analyses, and designed a new analysis from which information was gathered. technical problems associated with regional and global scale biodiversity analyses are: • distributed datasets • datasets that are heterogeneous at the physical, logical, and semantic levels • labor-intensive dataset integration • undocumented and non-repeatable manual integration steps the needs of the biodiversity analysis research group have driven development of seek’s observation ontologies (madin et al. submitted) and the sms. they also helped inform design of semi-automated integration tools within kepler and emphasized the need to capture provenance for derived datasets. usability usability is concerned with three major dimensions: effectiveness, efficiency, and user satisfaction. two primary methods to achieve usability are to apply research-based human factors design principles and to employ a usercentered design approach in which users are actively involved in the design and development process. both methods were applied on the seek project. in kepler, design principles were applied and two rounds of usability testing were conducted. two groups of scientists (total of 33) participated in user profiling, usability testing, and a facilitated downey & pennington bridging the technol figure 2. a schematic group discussion on the kepler application as part of two workshops: the first a meeting with the ecological niche modeling community disc above, the second an ecoinformatics training workshop for early career faculty discussed below. the usability tests uncovered usability issues that were translated to 16 design recommendations ( high priority). the follow-on facilitated discussions produced a list of seven new features for consideration. table 1 provides an example of the kind of data collected from users. understanding the characteristics of scientists, the tools they use, and their experience level in various areas, informs design and leads to more useful technology to support their scientific efforts. to date, we have profiled 74 scientists including 39 early career faculty. table 2 gives a small sample of the kinds of user profile data we collected on seek. the sample deals with technical and technology experience - structured query language programming experience and data format with (e.g., spreadsheets, relational databases) well who uses models and amount of experience. data was self-reported via a survey on bridging the technology & science gap with ecology 22 . a schematic presentation of the architecture of the seek project. group discussion on the kepler application as part of two workshops: the first a meeting with the ecological niche modeling community discussed above, the second an ecoinformatics training workshop for early career faculty discussed below. the usability tests uncovered usability issues that were translated to 16 design recommendations (9 on facilitated produced a list of seven new features provides an example of the kind of data collected from users. acteristics of scientists, the tools they use, and their experience level in various areas, informs design and leads to more useful technology to support their scientific efforts. to date, we have profiled 74 scientists gives a small sample of the kinds of user profile data we collected on seek. the and technology structured query language and data formats worked with (e.g., spreadsheets, relational databases)-as amount of modeling reported via a survey on a five-point scale (1 low and 5 high). user profiling, we also mapped the tasks kepler supports to two skills dimensions of our target users: information technology (it) skills and quantitative skills. figure 3 shows the basic task mapping along those dimensions. table 1. features for consideration identified by user scientists during usability activities. future features to consider workshop 1 natural language summary of workflows -summarization of workflow (in a publishable format) ability to assign check-points at various points in the workflow so that the user can check progress and make decisions on whether to modify or continue etc. ability to visualize data at points in the workflow guided analysis (wizard functionality for constructing workflows ecology examples point scale (1 low and 5 high). in addition to d the tasks kepler supports to two skills dimensions of our target users: information technology (it) skills and shows the basic task pping along those dimensions. . features for consideration identified by user future features to consider strongly workshop 2 provide browsing and filtering mechanisms especially for data and data nodes on the ecogrid. implement the “most recently used” concept for workflows and actors. downey & pennington bridging the technol a listing of the major tasks is listed on the right side of figure 3. in each quadrant associated tasks are displayed. running workflows can be done by those with any combination of it and quantitative skills. those with higher quantitative skills can create simple or complex workflows. technology implementation tasks (components and workflows) are done by those with high it skills. kepler promotes a cycle of knowledge sharing and collaboration in science. data, table # sql exp. overall 74 2.08 /5.0 early career faculty 39 1.40 /5.0 figure 3. task/skills mapping quantitative skills. bridging the technology & science gap with ecology 23 of the major tasks is listed on the . in each quadrant, the associated tasks are displayed. running workflows can be done by those with any combination of it and quantitative skills. those with higher ate simple or complex workflows. technology implementation tasks (components and workflows) are done by those kepler promotes a cycle of knowledge sharing data, components, workflows, and knowledge about w created and used by one group can be utilized by other groups. figure 4 shows some possible ways in which kepler components can be shared between groups. along with sharing, kepler promotes and supports collaboration among scientists. for example a scientist with strong quantitative skills can team with a scientist with strong it skills to produce sophisticated models and analyses. table 2. profile of seek users (selected results). sql exp. prog exp. data format sprdsht rel db models 2.08 /5.0 2.97 /5.0 49 (66%) 20 (27%) 50 (68%) 1.40 /5.0 2.64 /5.0 30 (77%) 8 (21%) 30 (77%) apping of kepler users along dimensions of information technology and ecology examples and knowledge about workflows created and used by one group can be utilized by shows some possible ways in which kepler components can be shared along with sharing, kepler promotes and supports collaboration among scientists. for example a scientist with strong quantitative skills can team with a scientist with strong it skills to produce sophisticated models model exp. 3.03 /5.0 2.56 /5.0 users along dimensions of information technology and downey & pennington bridging the technol figure 4 – illustration of how workflows, and knowledge. as mentioned earlier, the sms is a key feature of kepler, aimed at making computation and data integration much easier and quicker for scientists. ontologies, formal representations of related concepts, are being used to inform the sms are exposed to users through the kepler interface. given the newness of ontology development, the paucity of systems that make use of ontologies, and the lack of empirical studies to inform interface design, we constructed some early user interface prototypes with input from domain scientists on the team and by applying standa design principles. these interfaces allowed users to select terms from an ontology and assign them to workflow components (the process of annotation). however, our team members were somewhat knowledgeable about ontologies having participated in ontology design, and we were concerned that scientists with no knowledge engineering experience might have a different user interface expectation. subsequently, we conducted a design exercise (downey 2006) with a small number of scientists (3) as a starting point understanding the expectations of this user group. this type of ucd activity is formative, conducted bridging the technology & science gap with ecology 24 illustration of how kepler users of varying skills combinations can share data, as mentioned earlier, the sms is a key feature aimed at making computation and data integration much easier and quicker for scientists. ontologies, formal representations of related concepts, are being used to inform the sms, and are exposed to users through the kepler interface. ontology development, the paucity of systems that make use of ontologies, and the lack of empirical studies to inform interface design, we constructed some early user interface prototypes with input from domain scientists on the team and by applying standard design principles. these interfaces allowed users to select terms from an ontology and assign them to workflow components (the process of annotation). however, our team members were somewhat knowledgeable about ontologies, y design, and we were concerned that scientists with no knowledge engineering experience might have a different user interface expectation. subsequently, we conducted a design exercise (downey 2006) with a small number of scientists (3) as a starting point in understanding the expectations of this user group. this type of ucd activity is formative, conducted during the development process and is often exploratory (stone et al. 2005). results indicated a strong preference for a simple design rather than a feature-rich presentation. but the most interesting results were (1) that users expected the system to offer them contextually terms from the ontology depending on what component was selected for annotation (i.e., to filter out any noise or irrelevant terms based on the selection), and among terms selected from the ontology. this ucd exercise not only revealed user interface issues and user expectations but also highlighted other areas needing research like variance among term selection and how that might affect the efficacy of the sms. while kepler is a major focus, we have also applied ucd methods to other seek technology. formative design and evaluation activities with small numbers of representative users provide rich design information and reveal major usability issues. to inform the design of one of our taxonomic products, we conducted three activities: 1. user analysis and profiling with four taxonomists 2. task analysis (hackos and redish 1998) ecology examples ata, components, during the development process and is often exploratory (stone et al. 2005). results indicated a strong preference for a simple design rather than a rich presentation. but the most interesting 1) that users expected the system to terms from the ontology, depending on what component was selected for to filter out any noise or irrelevant and (2) the variance among terms selected from the ontology. this ucd exercise not only revealed user interface issues and user expectations but also highlighted other areas needing research like variance among how that might affect the while kepler is a major focus, we have also applied ucd methods to other seek technology. formative design and evaluation activities with small numbers of representative users provide rich mation and reveal major usability issues. to inform the design of one of our taxonomic products, we conducted three activities: user analysis and profiling with four taxonomists task analysis (hackos and redish 1998) downey & pennington bridging the technology & science gap with ecology examples 25 3. a paper prototyping (snyder 2003) session with an ecologist having taxonomic knowledge. for another taxonomic product that visualizes data, we conducted an analysis with two different user groups: four taxonomists and five museum collections managers. our purpose was to gauge whether our software included the right functionality needed by two groups. whenever we had the chance to apply usability engineering techniques on seek, we took the opportunity, whether it was a planned study with a small group or leveraging a large gathering of scientists at a workshop. future usability activities will include remote evaluations that can take advantage of geographically dispersed users/scientists. this will provide richer and more diversified scientist input into our tools as well as the opportunity to collect more data. education and training many of the approaches under research and development in seek represent completely new ways of problem solving for scientists, and depend on development of entirely new skills and new ways of thinking. ecological studies in the late twentieth century were characterized by single investigators making original observations in field notebooks, then analyzing their data with calculators or personal computers. by the end of the century, efforts to solve global environmental issues made clear that the old ways would not suffice. calls for cross-disciplinary, synthetic analysis over larger geographic and temporal scales enabled by advanced technologies are now pervasive throughout science, and especially within the natural sciences (cottingham 2002, kostoff 2002, cash et al. 2003, pickett et al. 2005). a new generation of scientists must emerge; scientists who are computationally savvy, can easily use technology to discover and integrate voluminous and/or heterogeneous data into their research, and who can collaborate effectively with computational scientists to construct innovative solutions to complex scientific problems. seek initiated an innovative training program for early career scientists: new faculty and postdoctoral researchers who are most likely to be the early adopters of new approaches. each year, 20 early-career scientists were chosen competitively to attend a one-week training workshop on ecoinformatics covering a broad range of cutting-edge technical concepts as well as more focused training on specific solutions being implemented by seek. participants were chosen in part because they were engaged actively in research that could clearly benefit from better technical approaches. most participants had little prior experience developing generic technical solutions to their research problems, as indicated by the user profiles mentioned above. many were involved in modeling activities that made use of scripts within scientific modeling software such as matlab, with little attention to useful information technology practices such as versioning, reusability, and program documentation. the goals of the workshop were far-reaching: to provide the necessary concepts and training that would allow participants to cross the barriers and begin incorporating new technical approaches into their ongoing research. participants were introduced to ecoinformatics and the seek project, and then exposed to a number of informatics topics relevant to conducting technology-empowered research (pennington et al., 2008). topics were covered in the order they would logically arise during a typical research project and included: • research design for enabling technical solutions • databases, metadata, data management and sharing • distributed data grids • analytical and visualization tools • scientific workflows • knowledge representation and ontologies. each workshop included pre-training and post-training surveys. an overwhelming percentage of participants agreed or strongly agreed that the workshops were useful (84.9%). most indicated that they learned a great deal during the workshops and that it met their expectations (90.4%). however, follow-up contact has indicated that few have been able to incorporate these approaches into their research as effectively as they would like. commonly stated reasons are that the week was simply not enough time, that they need additional training, and that they need mentoring as they actually start trying to use these new techniques. additional resources will be needed to develop more strategies for downey & pennington bridging the technology & science gap with ecology examples 26 bridging the gap, which is clearly broader than one technology research project can span. research implications one of the interesting benefits of applying ucd techniques on seek was that other areas of ucd research emerged as a result, within and between disciplines. within usability, a new testing method is being developed: group usability testing. this came about as a result of a pragmatic decision to take advantage of groups of scientists gathered for workshops, and because of limited time. instead of conducting traditional individual testing, we conducted group usability testing (downey 2007). future plans include conducting the same test both individually and as a group and comparing results. this research has the potential to improve the efficiency of usability testing. several research implications arose during the design activity on annotating workflow components from formal ontologies, described in detail in (downey 2006). they can be categorized into two broad areas: user interface design issues and annotation issues. in the short term, we plan to research the variance of term selection by users in the annotation process. from a broader perspective, our experience on the seek project has highlighted the need for theories and models of collaboration between scientists and technologists. we have initiated new research efforts towards this end 6 , searching for better ways to enable the incorporation of new technologies into scientists’ work processes (pennington, 2008). conversely, we also need better ways to determine user needs for innovative technologies, when those technologies are designed to fundamentally change the work tasks that normally inform design. lastly, we have begun to investigate mechanisms for designing technologies that simultaneously meet the needs of multiple collaborating user groups (pennington et al. 2007). summary ucd enhances and exploits the broadening scope of informatics that includes humans, technology and knowledge. the benefits of ucd on the seek project are clear. we believe that the usercentered design approach is essential for 6 http://www.scidesign.org. successfully bridging the gap between technology and science. individuals with this perspective and training have also been identified by the national science foundation (atkins et al.) 7 “the need for a new workforce – a new flavor of mixed science and technology professional – is emerging. these individuals have expertise in a particular domain science area, as well as considerable expertise in computer science and mathematics. also needed in this interdisciplinary mix are professionals who are trained to understand and address the human factors dimensions of working across disciplines, cultures, and institutions using technology-mediated collaborative tools. prior work on computer-supported collaborative work and social dimensions of collaboratories needs to be better codified, disseminated, and applied in the design and refinement of new knowledge environments for science based on cyberinfrastructure.” in order to build the best scientific tools possible, to promote and enhance collaboration and to enable new analyses and discoveries - involving scientists in the design and development of scientific tools is essential. the user-centered design approach is a proven methodology for achieving user involvement and is the bridge between technology and science. acknowledgements we acknowledge all the scientists who have participated in the user-centered design activities in seek, and all the members of the seek team for their support, enthusiasm, and contributions. finally, we thank the seek executive staff for their support and encouragement in our work. references altintas i., berkley c., jaeger e., jones m., ludäscher b., mock s. 2004. kepler: an extensible system for design and execution of scientific workflows. proceedings of 16th international conference on scientific and statistical database management (ssdbm). santorini island, greece. june 2004. 7 http://www.nsf.gov/cise/sci/reports/atkins.pdf. downey & pennington bridging the technology & science gap with ecology examples 27 araújo, m. b., pearson r. g., thuiller w., erhard m. 2005. validation of species-climate impact models under climate change. global change biology 11:1504-1513. atkins, d. (chair). 2003. revolutionizing science and engineering through cyberinfrastructure: report of the national science foundation blue-ribbon advisory panel on cyberinfrastructure.8 barbaut r., sastrapadja s. 1995. generation maintenance and loss of biodiversity. heywood, v. h., editor. global biodiversity assessment. cambridge university press. cambridge. pp. 193-274. beerling d. j., huntley b,, bailey j. p. 1995. climate and the distribution of fallopia japonica: use of an introduced species to test the predictive capacity of response surfaces. journal of vegetation science 6:269-282. bias r., mayhew d. 2005. cost-justifying usability, second edition. morgan kaufmann, san francisco, ca.. bowers s., lin k., ludäscher b. 2004. on integrating scientific resources through semantic registration. proceedings of the 16th international conference on scientific and statistical database management (ssdbm). santorini island, greece. june 2004. bowers s., ludäscher b. 2004. an ontology-driven framework for data transformation in scientific workflows. proceedings of the international workshop on data integration in the life sciences (dils). springer, lncs, volume 2994. bowers s., ludäscher b. 2003. towards a generic framework for semantic registration of scientific data. proceedings of the workshop on semantic web technologies for searching and retrieving scientific data, ceur workshop proceedings, volume 83, issn 16130073. carpenter g., gillison a. n., winter j. 1993. domain: a flexible modeling procedure for mapping potential distributions of animals and plants. biodiversity and conservation 2:667-680. cash d. w., clark w. c., alcock f., dickson n. m., eckley n., guston d. h., jager j., mitchell r. b. 2003. knowledge systems for sustainable development. pnas 100:8086-8091. cottingham, k. l. 2002. tackling biocomplexity: the role of people, tools, and scale. bioscience 52:793-799. downey l. l. 2006, designing annotation mechanisms with users in mind: a paper prototyping case study from the scientific environment for ecological knowledge (seek) project, semantic web user interaction workshop, fifth international semantic web conference. athens, ga. downey l. l. 2007. group usability testing: evolution in usability techniques. journal of usability studies [internet]. [cited 2007]. 2:133-144. hackos j. t., redish j. c. 1998. user and task analysis for interface design. john wiley & sons, inc., new york, ny. pp 71-72. kostoff r. n. 2003. overcoming specialization. bioscience 52:937-941. ludäscher b., altintas i., berkley c., higgins d., jaegerfrank e., jones m., lee e., tao j., zhao y. 2006. scientific workflow management and the kepler system. 8 http://www.nsf.gov/od/oci/reports/atkins.pdf. concurrency and computation: practice & experience 18:1039-1065. madin j., bowers s., krivov s., pennington d., schildhauer m., villa f. (submitted) an ontology for describing and synthesizing ecological observation data. international journal of ecological informatics, special issue on data management. martínez-meyer e., peterson a. t., hargrove w. w. 2004. ecological niches as stable distributional constraints on mammal species, with implications for pleistocene extinctions and climate change projections for biodiversity. global ecology and biogeography 13:305314. michener w. k., beach j. h., jones m. b., ludaescher b., pennington d. d., pereira r. s., rajasekar a., schildhauer m., 2007. a knowledge environment for the biodiversity and ecological sciences, journal of intelligent information systems. 29:111-126 nix h. a. 1986. a biogeographic analysis of australian elapid snakes. longmore r, editor. atlas of elapid snakes of australia. australian flora and fauna series 8:4-15. pennington, d. 2008. cross-disciplinary collaboration and learning. ecology and society. 13:8. pennington, d. 2008. a view from the trenches. computing in science and engineering, 10:28-33. pennington d., higgins d., peterson a. t., jones m. b., ludaescher b., bowers s. 2007. ecological niche modeling using the kepler workflow system. taylor i, gannon d, deelman e, shields m, editors, workflows for e-science. springer-verlag. pennington d. d., madin j., villa f., athanasiadis i. n. 2007. computer-supported collaborative knowledge modeling in ecology. proceedings of the workshop on social and collaborative construction of structured knowledge, 16th international world wide web conference (www2007). banff, canada, may 8, 2007. peterson a. t., vieglais d. a. 2001. predicting species invasions using ecological niche modeling. bioscience 51:363-371. pickett s. t. a., cadenasso m. l., grove j. m. 2005. biocomplexity in coupled natural-human systems: a multidimensional framework. ecosystems 8:225-232. pimm s. l., raven p. 2000. extinction of numbers. nature 403:843-845. snyder c. 2003. paper prototyping: the fast and easy way to design and refine user interfaces, elsevier, san francisco, ca. stone d., jarrett c., woodroffe m., shailey m. 2005. user interface design and evaluation, morgan kaufmann, san francisco, ca. pp 552. waide r. b., willig m. r., steiner c. f., mittelbach g., gough l., dodson s. i., juday g. p., parmenter r. 1999. the relationship between productivity and species richness. annual review of ecology and systematics 30:257-300. walker p. a., cocks k. d. 1991. habitat: a procedure for modelling a disjoint environmental envelope for a plant or animal species. global ecology and biogeography letters 1:108-118. microsoft word sautteretal07v2.doc biodiversity informatics, 4, 2007, pp. 1-13 a quantitative comparison of xml schemas for taxonomic publications guido sautter1,3, klemens böhm1, and donat agosti2 1department of computer science, universität karlsruhe (th), 76128 karlsruhe, germany; 2 division of invertebrate zoology, american museum of natural history, new york ny 10024-5192, and naturmuseum der burgergemeinde bern, 3005 bern switzerland; 3 sautter@ipd.uka.de abstract.— large numbers of legacy taxonomic publications are currently being digitized to make them online available and ready for full text search. the documents are being marked up with xml for two purposes: to preserve the document structure, and to facilitate access via standard query languages like xquery. with regard to the second aspect, the choice of an appropriate xml schema is crucial. it affects both query performance and the correctness of query results. over the last few years, several different xml schemas have been proposed as markup standards for taxonomic publications. in this paper, we report on a thorough evaluation and comparison of these schemas. we have examined if they facilitate formulation and correct processing of queries that are common when it comes to taxonomic literature. we also compare the performance of these queries on documents that are marked up with the different schemas. finally, we propose extensions to the schemas that enhance correctness of query results. key words. — heritage literature, quantiative analysis, systematics, taxonomy, taxonx, xml schema at present, legacy taxonomic publications are being digitized in large numbers (e.g. biodiversity heritage library1). the intention is to store these documents in digital archives and to make them available online. the documents are marked up with xml for two purposes. the first one is to preserve the original document structure and publication-related in-formation like publisher, title and issue. the sec-ond one is to facilitate deployment of standard query languages like xpath (xpath) to access the document collection. in the recent past, several institutions and projects have proposed a variety of xml schemas for this purpose, such as abcd (abcd), sdd (sdd), taxonx (taxonx), and taxmlit (weitzman & lyall). in this paper, we compare these schemas and include some additional ones. our comparison focuses on the second aspect mentioned above, relevant to the work of biologists. the queries typical for this domain are fine-grained (at the level of individual characters or distribution records), and their results are individual treatments, i.e., descriptions of a particular taxon. we have identified three basic types of criteria on which queries can be based: taxonomic names, the collection locations, i.e., the locations 1 http://www.bhl.si.edu/ where specimens of a particular taxon have been collected, and the morphological feature concepts as the selection criterion. we investigate both the ease of formulation and the execution performance of these queries. our results show that the four schemas mentioned above support queries over taxonomic names very well. the same is true for the collection locations. however, sdd is the only schema to allow formulating queries over morphological features at least to a certain degree, and they execute slowly in the environment investigated here. the other schemas do not provide any markup to identify individual concepts within morphologic descriptions. it is not possible to represent the relationship of a concept name and the associated descriptive terms. to overcome this problem, we propose some detail-level extensions to the different schemas. we show that the extended schemas better support queries that use morphological concepts as selection criteria. in our experiments, we have observed a moderate performance decrease due to the more complex markup. the issues investigated here are orthogonal to the question how the markup should actually be created, be it automated, semi-automated or completely manual. clearly, this question is important as well, and much research has sautter et al. comparison of xml schemas for taxonomic publications addressed it, including our own (sautter et al. 2006, sautter et al. 2007). in this current study, we assume that the documents are already marked up. we assume that the reader has some familiarity with xml (markup language for semi-structured data) (xml) and respective query languages (xpath, xquery in particular, (xpath)). queries in order to make maximum use of digital biosystematics archives, the information from the documents needs to be accessible on the treatment level. figure 1 displays such a treatment, an excerpt from wheeler (1922), which is also the basis for all examples throughout this paper. in the marked-up examples to follow, we will only use the passages printed in bold in figure 1. they will be sufficient for our purposes. pheidole lacerta, new species [...] mandibles and clypeus smooth and shining, the former with small scattered, elongate punctures. head and thorax subopaque, the head transversely rugose above, more reticulate-rugose laterally and in the occipital region, the scrobes finely and densely punctate. the gula is also reticulate but more loosely and finely and its sides are smooth and shining. thorax and petiole very finely and densely punctate, the pronotum also transversely rugulose above. postpetiole, gaster and legs smooth and shining, with fine, sparse, piliferous punctures. [...] a single specimen swept from foliage near port of spain by prof. roland thaxter. [...] figure 1: the example treatment (from wheeler 1922) there are three important information needs in biosystematics, or, in other words, three different ways of searching for information. they differ in the search criteria, which may be taxonomic names, collecting locations, or observable morphological features. in this section, we describe these three basic types of queries in more detail. the query types are also the basis for our evaluation experiments. these queries are fine-grained, i.e., they return individual treatments rather than entire publications. after the definitions, we give examples in natural language. the search criteria of the example queries are in italics. we will refer to these examples throughout this paper. name queries name queries find all information available on a given taxon. the selection criterion is the name of the taxon. the result is the set of all treatments on the particular taxon, i.e., descriptions as well as any collection event available. the latter serve to compute the dispersal of the taxon. the taxon range considered here is from genus down to variety. the natural language formulation of a typical name query would be: “find all treatments of a taxon with the given name. in particular, find the description and the dispersal”. example nq: “find the treatments on pheidole lacerta.” location queries location queries retrieve information on the fauna of a given location or area. the selection criterion is the name of a location or, ideally, longitude and latitude degrees. the result is the set of descriptions of all taxa that have ever been collected at this location, i.e., all taxa with a collection event which refers to it. in natural language, a typical location query could be: “find all taxa with a specimen having been collected at the given location”. example lq: “find all treatments on taxa endemic in port of spain”. concept queries concept queries identify a taxon based on its morphological feature concepts. the results sought are treatments on taxa with these concepts, in particular the treatments with a description. a concept query is as follows in natural language: “find all treatments describing a taxon showing a given morphological feature concept”. – note that our example treatment (see figure 1) will match both example cq1 and example cq2, but it does not match example cq3. the latter is because in the example text, shining refers to mandibles and clypheus, but not to head. example cq1: “find the treatments on all taxa whose head is subopaque”. example cq2: “find the treatments on all taxa whose thorax is subopaque”. example cq3: “find the treatments on all taxa whose head is shining”. sautter et al. comparison of xml schemas for taxonomic publications xml schemas in this section, we briefly introduce and discuss the different existing xml schemas that have been proposed for taxonomic publications. to our knowledge, no other relevant schemas exist. the second topic of this section is how different types of queries can be formulated (see previous section) based on the various schemas, and its performance estimations regarding query execution. due to space limitations, we restrict the examples to the treatment-internal markup, i.e., we do not provide the structural markup that delimits the treatments in the document. consequently, we also omit the markup of document meta-data, e.g., title, author(s), etc. in the mods elements of taxonx (taxonx). this applies to the markup examples as well as to the xpath expressions. in addition, we only include those xml elements in the markup examples that are relevant for the processing of the queries defined in the previous section. the markup examples are all based on the treatment shown in figure 1. abcd a design objective behind the abcd schema (abcd) is the preservation of the original document structure. abcd includes very detailed markup of publication-related data such as authors or publishers. treatments are marked up as units in this schema. figure 2 displays the example treatment marked up according to abcd. pheidole lacerta, new species mandibles and clypeus smooth and shining. head and thorax sub-opaque. a single specimen swept from foliage near port of spain by prof. roland thaxter. figure 2: example treatment marked up in abcd. for a document marked up in abcd, the three types of queries can be formulated as the following xpath expressions: name query (example nq): unit[contains(./identifications/identification/taxonidentified / scientificnamestring, “pheidole lactera”)] location query (example lq): unit[./gathring/gatheringsite/localitytext = “port of spain”] concept query (example cq1): unit[contains(./unitdescription, “head”) and contains(./unitdescription, “sub-opaque”)] the abcd schema supports the name queries and location queries very well. according to altinel and franklin (2000), the markup depth of a document affects query performance. therefore, we expect the name queries to execute quite slowly because of the deep nesting of the taxon name – there are five hierarchy levels unit, identifications, identification, taxonidentified, and scientificnamestring. the concept queries cannot be formulated precise enough – we would obtain incorrect results. the reason is that abcd markup does not make the relationship of the name of a morphological feature and its description explicit. for instance, the markup example would also match the query below, which formulates example cq3 as exact as possible, even though ‘shining’ is not part of the description of the morphological feature ‘head’. unit[contains(./unitdescription, “head”) and contains(./unitdescription, “shining”)] sdd / ubif the sdd schema (sdd) provides detailed markup for textual descriptions. it imports the ubif schema (ubif). the combination of both provides detailed markup for document-related as well as for datacentric aspects. figure 3 shows the example treatment marked up according to sdd. the key idea of this schema is that data on different aspects (taxon name, locations, morphologic feature concepts, textual description) is represented separately and linked by ref attributes: the taxonomic name is contained in a classname element. a geography element marks up the collection locations. the description is wrapped in a descriptivedata element, which has two children among others: a terminology element (transitively) contains concept elements, which sautter et al. comparison of xml schemas for taxonomic publications represent the individual morphological features the description refers to. they have an id attribute identifying them. the treatment text itself is enclosed in a naturallanguagedescription element. concept elements enclose the individual sentences of the description. these concept elements are linked to the ones in terminology with id-ref references. in particular, a concept in the naturallanguagedescription having a certain value as its ref attribute refers to the concept in the terminology whose id attribute has the same value.
mandibles and clypeus smooth & shining head and thorax subopaque. a single specimen swept from foliage near port of spain by prof. r. thaxter.
figure 3: example treatment marked up in sdd / ubif for a document marked up in sdd / ubif, the three types of queries can be formulated as the following xpath expressions: name query (example nq): dataset[./externaldatainterface/classnames/ classname/label/representation/text = “pheidole lactera”] location query (example lq): dataset[./externaldatainterface/geography/locality/ label/representation/text = “port of spain”] concept query (example cq1): dataset[descriptivedata/terminology/concepttrees/ concepttree//concept[.//label/representation/text = "head" and ./@id = string(//naturallanguagedata/ concept[contains(./text, "subopaque")]/@ref)]] the sdd / ubif schema supports the name queries and location queries. but according to altinel & franklin (2000), we can expect both to execute slowly because of the deep nesting of the taxon name and the locations. the concept queries can be formulated in a way that we expect to yield correct results. this is an advantage over the other schemas considered. but we expect these queries to execute very slowly due to the id references. dereferencing them results in a join, which is very complex and time-intensive to evaluate (wu et al. 2003). for large documents, in-memory query processing is not feasible, and the query engine has to perform file-based query evaluation. if it builds an index over the ids while performing the scan, all the elements have to remain in sautter et al. comparison of xml schemas for taxonomic publications memory. this is likely to result in out-ofmemory problems (ives et al. 2000). to avoid this, the query engine has to scan the document twice: once for collecting the referenced ids, and a second time to find the corresponding elements. this additional scan roughly doubles query-execution time. scanning the file takes the largest share of time in file-based query evaluation (böhm 2000). an additional problem is that a concept element in the natural language description can refer to only one concept in terminology. our fragment of the sample document matches example cq2 (“find the treatments on all taxa whose thorax is subopaque.”), but the concept query from above formulated against the sdd / ubif schema, with 'head' replaced with 'thorax', will not return a hit. an alternative formulation of example cq2 can overcome this problem, but would reduce the explicit modeling of the terminology element to absurdity: dataset[descriptivedata/naturallanguagedescriptions/ naturallanguagedescription/naturallanguagedata/ concept/text[contains(., "thorax") and contains(., “subopaque”]] finally, the original document is hard to reproduce once it has been transferred to an sdd / ubif representation. taxonx the taxonx schema (taxonx) preserves part of the original document structure. in particular, it provides paragraphs and different types of divisions. the individual treatments are enclosed in a tag. figure 4 displays the example treatment marked up according to taxonx. pheidole lacerta new species

mandibles and clypeus smooth and shining. head and thorax sub-opaque.

a single specimen swept from foliage near port of spain by prof. roland thaxter.

figure 4: example treatment marked up in taxonx for a document marked up in taxonx, one can formulate the three types of queries as the following xpath expressions: name query (example nq): treatment[./nomenclature/name = “pheidole lactera”] location query (example lq): treatment[./div[./@type = “materials_examined”] /p/seg/collection_event/locality = “port of spain”] concept query (example cq1): treatment[./div[./@type = “description”] /p[contains(./text(), “head”) and contains(./text(), “subopaque”)]] we found the name queries and location queries easy to formulate on taxonx. we expect the name queries to execute fast. the location queries, on the other hand, are likely to execute slowly because of the deep nesting in div, p, seg, collection_event, and locality tags. the concept queries can be formulated, but we expect incorrect results, for the same reason as with abcd. the example query (producing incorrect results) is as follows for taxonx: treatment[./div[./@type = “description”] /p[contains(./text(), “head”) and contains(./text(), “shining”)]] taxmlit the taxmlit schema (weitzman et al.) provides very detailed markup on different levels, document-centric as well as data-centric. its nesting is very deep: all textual content is wrapped in at least five hierarchy elements. it is somewhat cumbersome to formulate the long queries necessary for drilling down into the hierarchy. in addition, the documents are not exactly human-readable. the markup is relatively heavy, compared to the other schemas. for a document marked up in taxmlit, the three types of queries can be formulated as the following xpath expressions: name query (example nq): taxontreatment[./taxonheading/taxonheadingname/ taxonname/taxonnametext = “pheidole lactera”] location query (example lq): taxontreatment [./distributionandorspecimencitations/ individuallocalities/locality/detailedlocation/ detailedlocationtext = “port of spain”] concept query (example cq1): taxontreatment[./descriptions/ samelanguagedescription/ samelanguagedescriptonparagraphs/ samelanguagedescriptionparagraph/ samelanguagedescriptiontext[contains(./text(), “head”) and contains(./text(), “sub-opaque”)] sautter et al. comparison of xml schemas for taxonomic publications the taxmlit schema formally supports the name queries and location queries. but we again expect both to execute slowly because of the deep nesting. the concept queries can be formulated, but we expect incorrect results, for the same reasons as with abcd and taxonx. with taxmlit, example cq3 in xpath looks like this: taxontreatment[./descriptions/ samelanguagedescription/ samelanguagedescriptonparagraphs/ samelanguagedescriptionparagraph/ samelanguagedescriptiontext[contains(./text(), “head”) and contains(./text(), “shining”)] ... pheidole lacerta, new species species pheidole lacerta mandibles and clypeus smooth and shining. head and thorax sub-opaque. a single specimen swept from foliage near port of spain by prof. roland thaxter. port of spain figure 5: example treatment marked up in taxmlit tcs the taxonomic concept schema (tcs) is intended for the transfer of taxonomic data rather than for document markup. this means that the tags provided by tcs are not designed for sophisticated search and querying of treatments, but for transferring individual concepts between applications or machines. consequently, we do not include it in our evaluation of xml schemas for the markup of taxonomic publications. linneancore the linneancore schema (linneancore) provides very detailed markup for taxonomic names. it is intended for representing these names rather than for the markup of entire documents. thus, there are no elements for geographical information or textual descriptions of a taxon. hence, we do not include the linnaeancore schema in our evaluation either. natural collections descriptions the natural collections description schema (ncd) is currently being developed. its intention is the description of specimen collections rather than the markup of taxonomic publications. consequently, we do not include the natural collections descriptions schema in our evaluation. darwincore 2 the darwincore 2 schema (darwincore2) provides detailed markup for the key parts of taxonomic names and collection events. all these elements reside in a simple list element, thus the nesting of this schema is very flat. it may well be applicable for the representation of individual collection events. but it is not applicable for the markup of publications, for two reasons: first, it does not provide elements for the markup of descriptions. second, the schema provides no root element to enclose the entire document, which may well contain more than one treatment. consequently, we do not include the darwincore 2 schema in our evaluation of xml schemas for the markup of taxonomic publications. conclusion of schema analysis four of the eight schemas (abcd, taxmlit, taxonx and sdd / ubif) presented in this section are intended for the markup of taxonomic publications. to keep this paper focused, we will restrict further considerations sautter et al. comparison of xml schemas for taxonomic publications and our evaluation to these four schemas. all the four schemas support the name queries and location queries well, but we expect different query execution costs for the various schemas because of different nesting depths. none of them supports the concept queries in a way that all morphological features can be queried, and that the results are correct. abcd, taxonx and taxmlit are rather document-centric (nambiar et al. 2000). they focus on the structure and text of the document and do not rearrange the content. the three schemas do not provide markup for individual morphological concepts. sdd / ubif in turn clearly is data-centric, i.e., it organizes the content of a document with the focus on its semantics, and the original structure is (more or less) lost. on the other hand, sdd / ubif is the only schema that, at least to a certain degree, supports concept queries in a way so that we can expect correct results. we expect the name queries and location queries to execute fast with abcd and taxonx due to the relative simplicity of these two schemas. the taxmlit schema in turn provides a very complex element nesting. we expect this to result in a decrease of query performance. that complexity seems to be unnecessary from our specific perspective because it provides no advantages over the first two schemas from the querying point of view. all three schemas share the problem that concept queries cannot be formulated sufficiently exactly. the sdd / ubif schema supports all three types of queries, but we expect all of them to execute rather slowly. this is because resolving the id references results in higher query complexity. this in turn decreases query performance. possible extensions as the previous section has shown, only the sdd / ubif schema allows formulating the concept queries in a way so that we can expect correct results. this is because the other schemas are too document-centric. in particular, there is no way of expressing the relation of a morphological feature name and the associated description. even the level sdd / ubif provides is insufficient: morphological features can be queried exactly only if the terminology part provides a corresponding concept. otherwise, the problems are the same as those of the other schemas. to support the concept queries better, a marked-up document should allow querying individual morphological feature concepts. in this section, we propose and discuss three options to mark up the descriptions on a lower level in a more data-centric fashion to achieve this goal. these different types of detail-level markup can be added seamlessly to the schemas as children of the paragraph or treatment elements. the latter are part of all of the schemas. in other words, our proposed extensions are independent of a particular schema. morphologic indexing extending the concept element of sdd / ubif, we create a morphologic index. this index contains all morphological feature names contained in the description, and all descriptive terms for each of them. a morphological feature concept (concept) contains the feature name (label) and a list of the descriptive terms assigned to the name (description). figure 6 displays the morphologic index for our example treatment. due to space limitations, we restrict the figure to the terminology part of the dataset. in particular, we omit the externaldatainterface and naturallanguagedescription parts and the representation tags. – this highly data-centric idea results in the following dilemma: if the original document structure is to be preserved, the morphologic index contains much redundant data. if redundancy is not desirable, on the other hand, the original document structure cannot be preserved. subopaque subopaque smooth shining smooth shining figure 6: morphologic index sautter et al. comparison of xml schemas for taxonomic publications as a consequence of this extension, our example concept query for sdd is now significantly less complex than the original one. this is because the extension allows formulating the query without a join. for the other three schemas, only this extension enables formulating the concept queries in a way similar to sdd, and thus obtaining correct results. example cq1 now formulates like this: dataset[descriptivedata/terminology/concepttrees/ concepttree//concept[./label = "head" and contains(./description, "subopaque")]] we expect this query to execute significantly faster than the original one. on the other hand, we also expect the name queries and location queries to perform a little worse. the reason is the increased document size and the additional xml elements induced by the redundant data. aspect markup if the original document structure is important, and redundancy is not desirable, we have to use a more document-centric approach to mark up the descriptions. the idea is to use in-line tags for partitioning the description paragraphs into individual description aspects. such an aspect consists of a (set of) descriptive term(s) and the name(s) of the morphological features they refer to. in our example document, each sentence of the description is an aspect. the feature elements identify the names of the morphological features that are subject to the description. this finegrained markup allows a sufficiently exact formulation of the concept queries. using the taxonx example as the basis, (a fragment of) our example document becomes the one in figure 7.

mandibles and clypeus smooth and shining , head and thorax sub-opaque .

figure 7: morphologic description with individual aspects marked up. individually marking up the feature names is not necessary in this example. this is because the description applies to all feature names in a sentence. in more complex documents with longer sentences, however, such markup is necessary in order to distinguish the described features from the ones that are part of descriptive terms. consider the aspect in figure 8. the descriptive terms refer only to “mandibles” and “clypheus”. “hairs” is not a feature described. without the features marked up explicitly, such differentiations would not be possible in a query. mandibles and clypeus smooth and shining, with fine hairs on them figure 8: a more complex morphologic description because the additional markup is added only on levels beneath the most fine-grained element of the original schema, we can easily apply the approach to abcd and taxmlit as well. due to space constraints, we only display the div element containing the description. the markup is similar to the character and state elements provided by taxonx. the difference is that the latter mark up the description terms (state) instead of the morphological feature names. as a consequence of this extension, our example cq1 can now be formulated sufficiently exactly: treatment[./div[./@type = “description”] /p/aspect[./feature[./text() = “head”] and contains(./text(), “sub-opaque”)]] example cq3 now would formulate as below, and this query would not match the markup example in figure 7. this is because the aspect elements now delimit the individual concepts in the description. treatment[./div[./@type = “description”] /p/aspect[./feature[./text() = “head”] and contains(./text(), “shining”)]] consequently, we can express the relation between the name of a morphological feature and its associated description. again, this is not for free: we expect the performance of the name queries and location queries to decrease due to larger document size and the increased number of xml elements. normalized aspect markup a slight problem of the aspect markup approach is that we use the original text for the identification of the feature names: we only enclose the respective words in feature markup. this can result in spelling-related errors due to singular/plural or capitalization differences. by adding a normalized form of the feature name (e.g., all lower case) to the feature tag as an attribute value, we can overcome this problem. sautter et al. comparison of xml schemas for taxonomic publications the example from above would now look like the fragment in fig. 9.

mandibles and clypeus smooth and shining , head and thorax sub-opaque .

figure 9: morphologic description with individual aspects marked up, feature names normalized the resulting changes to the query formulating example cq1 are minimal: treatment[./div[./@type = “description”] /p/aspect[./feature[./@name = “head”] and contains(./text(), “sub-opaque”)]] the advantages are the same as those of the aspect markup, plus the spelling insensitivity. the latter comes at the cost of a little data redundancy and some more xml elements. consequently, we expect this extension to affect the performance of name queries and location queries a little more than the aspect markup. evaluation in this section, we report on our evaluation of the different schemas and the extensions proposed in the previous section. before we report on results, we briefly describe the experimental setup. experimental setup we have used altova xmlspy 2005, a widely used up-to-date xml editor, to execute the queries. the experiments have been run on a machine equipped with an intel pentium iv dual core and 1024 mb of ram. while we are aware of the fact that a native xml database or an sql database with xml extensions might provide better performance, our setup corresponds to the way biosystematicists work with documents today. in addition, processing queries on files is an approach that the databaseresearch community has paid much attention to in the recent past (abiteboul et al. 1993, 1995). test data because the digitization of biosystematics publications has just started, real documents marked in the different schemas are not available in numbers sufficiently large for large-scale performance experiments. hence, we have generated artificial documents based on data taken from wheeler (1922). the important parts for our experiments are the taxonomic names, collection events and textual descriptions of the taxa. we have generated these parts in the following way: the taxonomic names were randomly assembled from genus, subgenus, species, subspecies and variety names we have extracted from (wheeler 1922). genus and species are always given, the remaining parts were added with the following probabilities: subgenus: 30% subspecies: 60% variety: 30% although the resulting taxonomic names do not exist in the real world, they are syntactically identical to real ones. this is sufficient for performance measurements. the collection events mainly consist of a location, often accompanied by the name of the biologist who collected the specimen. we have synthesized these parts by inserting a location name into a sentence pattern, which lets it look more natural. the location was randomly picked from a given list. again, identity on the syntactic level is sufficient for performance experiments. the textual descriptions were generated by randomly lining up descriptive sentences. we have extracted the sentences from the textual descriptions in wheeler (1922). out of 70 different sentences, we have used 2 to 6 for each description paragraph. 2 to 7 paragraphs form a complete description. the intention was to produce descriptions of varying length, in order to arrive at a realistic distribution of the sizes of the documents. by inserting these three parts into a pattern, we have obtained artificial treatments. for our experiments, we have generated documents containing 1,000 of such treatments. the plain text has a size of about 2 mb. the markup in the various schemas results in the document sizes listed below: abcd: 2.5 mb taxonx: 2.6 mb taxmlit: 4.6 mb sdd / ubif: 9.6 mb the difference between the abcd and taxonx is minimal. it results from the slightly higher level of layout details in taxonx. but the sautter et al. comparison of xml schemas for taxonomic publications difference of these two schemas to taxmlit is significant. the reason is the large number of tags used in the latter schema. finally, sdd / ubif requires almost four times the storage space of abcd and taxonx. but in contrast to taxmlit, this comparison in isolation is not very significant. this is because sdd / ubif is the only data-centric schema in this evaluation. results with plain schemas in this subsection, we present the results of our performance experiments with the original schemas. in particular, we have run the name queries, location queries and concept queries on the documents. table 1 lists the execution times. as expected, abcd and taxonx provide almost equal performance. the more complex nesting in taxmlit results in the duplication of execution time. the processing time for the queries on the sdd / ubif document is significantly higher. the reason is that this particular document contains more than twice the number of xml elements of the others. schema name queries locatio n queries conce pt queries abcd 2.25 2 2.25 ir taxonx 2 3 3.25 ir taxmlit 4 5 5.5 ir sdd/ubif 12.5 22.5 oom table 1: results with unmodified schemas (query execution time in seconds; oom out of memory error). sdd / ubif is the only schema that allows formulating the concept queries such that results are always correct. however, resolving the id references (i.e., computing the join) results in an out of memory (oom) error for large documents, as we had hypothesized. as expected, all the other schemas produce incorrect results for the concept queries (‘ir’ in the table). in particular, the queries return treatments where the queried attribute does not describe the morphologic feature concept queried. results with extended schemas the experimental results presented in the last section have substantiated our expectation that none of the schemas supports the concept queries properly, except for sdd / ubif. we have proposed three possible extensions to overcome this problem. in this section, we present the results of our experiments with the extended schemas. the creation of a morphologic index for each treatment is the most data-centric extension. after generating the index, our test documents have the following sizes: abcd: 5.2 mb taxonx: 5.3 mb taxmlit: 7.3 mb sdd / ubif: 12.4 mb as expected, the size of the documents marked up in document-centric schemas has almost doubled. this is due to the redundancy caused by the index. only sdd / ubif is less affected (30% larger). this is because we could simply attach a description element to the existing concept elements instead of creating a complete index. nevertheless, the sdd / ubif document is still 2.5 times as big as the ones marked in abcd and taxonx, respectively. table 2 lists the execution times of the different queries: schema name queries location queries concept queries abcd 4.25 7 9 taxonx 4.5 8 9 taxmlit 6.5 11.5 14 sdd / ubif 14 30.75 34.25 table 2: results with morphologic index (query execution time in seconds) with the morphologic index, all schemas allow formulating the concept queries sufficiently exactly so that they return correct results. avoiding the join also overcomes the out-of-memory problems with sdd. the decreased performance of the name queries and location queries results from the increased number of xml elements. the fine-grained markup of description aspects preserves the original document structure and produces no redundant data. the document size increases only by the additional tags. in particular, our test documents have the following sizes with this extension: abcd: 3.2 mb taxonx: 3.3 mb taxmlit: 5.3 mb sdd / ubif: 10.3 mb this extension increases the document size by about 25% for abcd and taxonx. the increase is less for the other documents. this is due to their larger original size because the additional tags are the same for all schemas. table 3 lists the resulting query-execution times. sautter et al. comparison of xml schemas for taxonomic publications schema name queries location queries concept queries abcd 4.75 7 8.75 taxonx 5 8 10.25 taxmlit 7 10.5 13 sdd/ubif 15 30 oom / 37 table 3: results with aspect markup (query-execution time in seconds; oom out of memory error). the impact on the performance of the name queries and location queries is slightly higher than that of the morphologic index. this is because the special markup of all features and description aspects introduces more additional xml elements. but the concept queries all produce correct results. for sdd / ubif, we report two results in the concept queries column because the original query using the id references did not execute successfully, but produced an out of memory (oom) again. in order to avoid the join, we then reformulated the query for example cq1 so that it does not involve the concept elements in terminology, but only the aspects (see below). the new query executed successfully and produced correct results. dataset[descriptivedata/naturallanguagedescriptions/ naturallanguagedescription/naturallanguagedata/ concept/tex/aspect[./feature[./text() = “head”] and contains(./text(), “sub-opaque”)]] finally, the extension of the aspect markup with normalized feature names enlarges the documents as much as the markup of the aspects. the document sizes now are as follows: abcd: 3.5 mb taxonx: 3.6 mb taxmlit: 5.6 mb sdd / ubif: 10.7 mb this extension increases the document size a little more than the sole markup of the aspects and features. this is due to the additional name attribute in the feature tags. the execution times of the different queries are listed in table 4. name queries and location queries execute slower than with the plain aspect markup and with the morphologic index. this is because the additional name attribute of the feature element introduces an additional xml node that has to be processed. on the other hand, the normalized aspect extension provides the best support for the concept queries. this is because it abstracts from singular/plural and other spelling-related differences between different instances of the same feature name. the reason that we list two results for the sdd / ubif document is the same as in the previous section: the query is only processed successfully if we ignore the concept elements. otherwise, it produces an out of memory error. schema name queries location queries concept queries abcd 9.75 12.25 13 taxonx 6 11 14 taxmlit 8 14.5 18 sdd/ubif 18 35 oom/45 table 4: results with normalized aspect markup (query execution time in seconds; oom out of memory error) discussion our evaluation points out several similarities and differences between the schemas. with regard to query formulation, taxonx and abcd are almost equivalent. the latter produces slightly smaller documents, while the former preserves the original document structure better. the query performance is approximately equal, and both have the semantic problem with the concept queries. but these in turn are easy to solve with one of the extensions proposed in this paper. despite its complexity, taxmlit offers little semantic or structural advantages with regard to the aspects investigated here. it enlarges the documents by about 80%, compared to the first two schemas. finally, sdd / ubif is the only schema that supports the concept queries without schema modifications, at least in theory. the id references result in out of memory errors if the document size exceeds a certain limit. in addition, this advantage goes along with an increased document size, almost four times the one of documents marked in abcd or taxonx. finally, the advantage becomes less significant if we use one of the extensions with the latter two schemas. even with the aspect extension, the size of the abcd and taxonx documents is about a third of the size of a document marked up in original sdd / ubif. consequently, query execution is about twice as fast with the former two schemas. the three extensions we have proposed serve their intended purpose well: they all facilitate correct results for the concept queries with each of the schemas. but they also have some side effects: the data-centric approach of adding a morphologic index to the documents induces the most redundancy, and the documents become significantly larger. nevertheless, it does not offer any advantages sautter et al. comparison of xml schemas for taxonomic publications over the detailed markup of aspects in the textual description of the taxon. this applies to query formulation as well as to performance. the two flavors of the aspect markup only have slight differences with regard to the document size. regarding query execution, however, they differ significantly: while the normalized version provides slightly better support for formulating the concept queries, it also has a non-negligible impact on query performance. though the redundancy induced is very small, the number of xml elements increase significantly because attributes are represented as extra elements. the unnormalized version yields better querying performance for the name queries and the location queries, comparable to the morphological index, while avoiding the redundancy. thus it does not enlarge documents very much. the only non-negligible drawback in comparison to the other two versions is that formulations of concept queries have to pay attention to different possible spellings of morphological feature names. conclusion in this paper, we have presented and compared the xml schemas that are being used or developed as standards for the markup of taxonomic publications. we have considered both the size of the marked-up documents and the performance of queries against documents marked up using these standards. in particular, we have used three types of queries, which cover the three basic information needs in biosystematics: 1. finding the description and dispersal of a taxon with a given name. 2. finding all taxa ever reported to appear in a given area. 3. finding all taxa that have a given morphological feature. although there are several more schemas, we have restricted our comparison to the four which allow formulating these three types of queries. three of the schemas (abcd, taxonx and taxmlit) are document-centric, i.e., they intend to preserve the original structure of the publication. the fourth schema (sdd / ubif) is more data-centric, i.e., it focuses on representing the data in a form that better supports query formulation and execution. because none of the document-centric schemas properly supports the third type of queries, we have proposed and evaluated three possible extensions to overcome this problem. our evaluation has shown that they are all feasible for this purpose. they differ in the amount of redundant data introduced to the documents: 1. the creation of a morphologic index for each treatment is the most data-centric approach. unfortunately, it induces a redundant representation of almost the entire textual description. 2. the markup of description aspects works in-line. it adds data-centric fine-grained markup to the leaves of the document-centric schema. 3. the normalization of the feature names in the aspects slightly accelerates query execution, at the cost of a little redundancy. from the querying point of view, we deem abcd and taxonx most feasible for the markup of taxonomic publications. with the aspect markup extension, both support all queries and only slightly enlarge the documents. they provide acceptable query performance. taxmlit introduces very complex markup. it enlarges the documents to almost twice the size of abcd and taxonx, but offers no advantages over the latter two schemas, at least not with regard to the aspects covered by this evaluation. the number of xml elements and the nesting complexity also significantly decrease query performance. finally, sdd / ubif is the only schema that natively supports the third type of queries, but at a high price: the documents are almost four times the size of abcd or taxonx documents. in addition, this type of queries only executes for small documents. larger ones produce errors because of insufficient computation resources. finally, sdd / ubif does not preserve the original structure of the document. even with our description aspect extension, the size of an abcd or taxonx document is little more than a third of the one of an sdd / ubif document. on the other hand, the aspect markup compensates all querying advantages that sdd / ubif has over native abcd and taxonx. this applies to both correctness and performance. acknowledgments the authors thank the members of the project team (supported by awards from the us national science foundation iis-0241229 and deutsche forschungsgemeinschaft bib47) for their comments. sautter et al. comparison of xml schemas for taxonomic publications references abcd. access to biological collection data2 abiteboul, s., s. cluet, and t. milo, 1993. querying and updating the file. proceedings of the 19th international conference on very large data bases, 73 – 84, dublin, ireland abiteboul, s., s. cluet, and t. milo, 1995. a database interface for file update. acm sigmod record 24(2): 386-397 altinel, m., and m. j. franklin, 2000. efficient filtering of xml documents for selective dissemination of information. proceedings of the 26th international conference on very large data bases, 53 –64, cairo, egypt böhm, k., 2000. on extending the xml engine with query-processing capabilities. in proceedings of ieee advances in digital libraries, 127-138, washington, dc, usa darwincore23 ives, z.,, a. levy, d. weld, 2000. efficient evaluation of regular path expressions on streaming xml data, technical report, university of washington, seattle, wa, usa linneancore4 ncd: natural collections description5 nambiar, u., z. lacroix, s. bressan, l. l. mong, and l. yingguang, 2000. current approaches to xml management. ieee internet computing 6(4): 4351. sautter, g., d. agosti, k. böhm, 2006. a combining approach to find all taxon names (fat) in legacy biosystematics literature, artikel, biodiversity informatics 3: 41-53 sautter, g., d. agosti, k. böhm, 2007. semiautomated xml markup of biosystematics legacy literature with the goldengate editor, in proceedings of psb 2007, weilea, hi, usa sdd: structure of descriptive data6 taxonx7 tcs: taxonomic concept transfer schema8 ubif: unified biosciences information framework.9 weitzman, a. l., c. h. c. lyal. an xml schema for taxonomic literature – taxmlit10 wheeler, w. m., 1992. the ants of trinidad. american museum novitates 45: 1-16 wu, y. j.m. patel, and h. jagadish, 2003. structural join order selection for xml query optimization, in proceedings of icde, 443-454, bangalore, india 2 http://www.bgbm.org/tdwg/codata/schema/ 3 http://darwincore.calacademy.org/ 4 http://wiki.cs.umb.edu/twiki/bin/view/ubif/linneancore 5 http://www.tdwg.org/ncd/tdwg_ncd_subgroup.htm 6 http://wiki.cs.umb.edu/twiki/bin/view/sdd 7 http://sourceforge.net/projects/taxonx 8 http://tdwg.napier.ac.uk 9 http://wiki.cs.umb.edu/twiki/bin/view/ubif 10 http://www.sil.si.edu/digitalcollections/bca/status.cfm xml: extensible markup language11 xpath12 11 http://www.w3.org/xml/ 12 http://www.w3.org/tr/xpath the digit (originally: digitisation *** ) work programme of gbif was originally conceive to mainly focus on mobilising the biodiversity information enclosed in the specimen holdings of natural history museums and herbaria of the world biodiversity informatics, 7, 2010, pp. 130 – 136. leveraging the fullest potential of scientific collections through digitization. roger baird collection services division, canadian museum of nature, ottawa, ontario canada rbaird@mus-nature.ca abstract -access to digitized specimen data is a vital means to distribute information and in turn create knowledge. pooling the accessibility of specimen and observation data under common standards and harnessing the power of distributed datasets places more and more information and the disposal of a globally dispersed work force which would otherwise carry on its work in relative isolation, and with limited profile and impact. citing a number of higher profile national and international projects, it is argued that a globally coordinated approach to the digitization of a critical mass of scientific specimens and specimen-related data is highly desirable. an action plan of this scale is required to maximize the value of these collections to civil society and to support the advancement of our scientific knowledge globally. key words. biodiversity; metadata; digitization; science infrastructure natural and human history collections are a part of the wider scientific collections infrastructure. currently, the earth is estimated to be home to approximately 11.3 million species, but less than 2 million have been formally described by science (chapman, 2009). whether these species are exploited for commercial gain, or conserved for ethical, aesthetic and scientific reasons, the limited knowledge that we have about the biological diversity of the planet is of serious concern internationally. scientific collections can be characterized as assemblages of natural history specimens as well as human history artifacts that have been sufficiently documented, at the time of acquisition or through the course of analysis, as to have lasting value as part of a broad research infrastructure. how are scientific collections and their data used? cultural collections: scientific collections of material culture have been important resources for archaeologists, anthropologists and ethnologists, facilitating studies of ancient and living cultures as well as comparative analysis between cultures. the ability to supplement personal networks of colleagues and, in part, to overcome their ephemeral nature, is a key benefit from the creation of datasets that serve a global audience. the data allows institutions to perpetuate the knowledge created through the life work of a researcher, and to broaden the reach of its holdings for the benefit of others. equally, data on these human history collections is of interest to other individuals who may be related to, or representing, the very people studied. interest in traditional techniques, a desire to connect with the past, an objective to undertake a physical or “virtual” repatriation of one’s culture are all potential motivators for the broader audience for digitized material culture collections. natural history collections: the traditional users of natural history scientific collections are invariably taxonomists identifying, naming and classifying speciesand systematists who study the diversity of life on the planet’s past and present as well as the relationships among living things through time. these specialists make it possible for the comparative science of biology to flourish. without this essential work, a large portion of anatomy, physiology, biochemistry and microbiology, and ecology could not be realized. this work is essential to ensure economic wellbeing, to preserve natural resources, to maintain health, and to guard against invasive species. mailto:rbaird@mus-nature.ca baird – leveraging the fullest potential of collections whether they were amassed historically or compiled recently in response to pressures from development or exploitation of a site, scientific collections are invaluable resources to answer science-based questions far beyond the reach of a single individual. the collections themselves are sub samples of the world as it once was and as we know it today, and can also provide critical predictive modeling data on what the future could be. providing digitized access to the information that is inherent in these collections and having this inherent information converted into insight and knowledge by appropriate specialists allows a researcher or policy maker to verify a range of questions related to a wide variety of subjects: biodiversity and environmental change collections offer evidentiary value for documenting the biological diversity of life and in doing so, can also demonstrate changes in the environment that have taken place through time. the presence or absence of a species in a geographic region, as well as extensions and retractions in the distributional range of a species over the course of time are documented by examining the records associated with these collections. a habitat recovery program can be demonstrated to be successful if species once thought to be extirpated from the area or endangered are documented as being reestablished, through observation records as well as physical specimens. scientific collections that are well documented and deposited can offer proof that mitigation measures were successful, can provide evidence by proxy that climate change has taken place, and can help to effectively monitor rare, threatened and endangered species. invasive alien species examining the causes of biodiversity loss, digital mapping techniques have revealed that invasive alien species are second only to the threat posed by habitat destruction. in a 1993 report, the u.s. office of technology assessment1 estimated cumulative economic losses of $100 billion in the us due to noxious weeds, invasive insect pests, introduced aquatic species from ballast water, and other non-indigenous species (simberloff, 1996). 1 http://www.fas.org/ota/reports/9325.pdf the ability to differentiate native from non-native species is necessary and made possible by analyzing the specimen holdings in natural history collections, which provide that baseline data on which these analyses are built. collection-based science is essential to prevent, detect, and to rapidly respond to and manage these threats. public health and wildlife disease important scientific collections are also managed outside of museum environments. viruses, cultures, tissues and pathology samples are an important subset of this scientific infrastructure to be found in research laboratories and biological resource centers or brcs. infections from h1n1 influenza, west nile virus, lyme disease, tuberculosis, chronic wasting disease, and sars are all medical conditions that manifest themselves in human and non-human populations. for these and other zoonoses, approximately 70% of new or newly important diseases affecting human health are believed to have a wild animal source (blancou, 2005), and can have profound impacts on urban, rural, and human health, culture, and global economy. collections data over time can reveal biodiversity threats through knockout effects on ecosystems e.g. high mortality of uk rabbits after introduction of myxomatosis led to declines in predators such as stoats, buzzards, and owls (sumption, 1985, 2008); the reduced grazing pressure by rabbits on heath lands in turn removed the habitat for an ant species that assists developing butterfly larvae, leading to extirpation of populations of the endangered large blue butterfly. economics, biosecurity and regulatory frameworks the predominance of global trade and marketing in our modern world requires an internationally coordinated infrastructure to share expertise derived from scientific collections. individual customs or border security officials who exercise levels of control on the movement and transport of imported goods and products have a reliance on knowledge derived from these collections on a daily basis. for example, the invasive pest agrilus planipennis or emerald ash 131 http://www.fas.org/ota/reports/9325.pdf baird – leveraging the fullest potential of collections borer is considered to have been introduced to north america through infested wood crating materials in 2002, and has come to rival dutch elm disease in its impact on tree populations2. timely access to information can counter distribution of fungal or insect infections through horticultural trade, and can reduce economic loss from unwarranted delays in customs inspections and quarantines (renaud, 2008). as well, genetic species barcodes linked to taxonomic collections have been demonstrated to be of great significance in detecting market substitutions involving overfished species or market fraud such as lutjanus campechanus or red snapper (wong, 2008). access and benefit sharing material culture collections of ancient or indigenous cultures are predominantly held by developed countries, while parties related to the cultures studied are most prevalent in underdeveloped countries. equally, biological collections and related scientific expertise are held disproportionately within developed countries, although the greatest portion of the world’s biological diversity is found in countries which are presently not as economically advantaged. under the convention for biological diversity3 sharing information and assisting international development through sustainable development practices is considered a global imperative, and facilitating access to data and knowledge on material collections supports the perpetuation of knowledge from traditional cultures. documentation from archaeological sites has also been relevant to the resolution of land claims by aboriginal groups, by evidencing traditional use and occupation of geographic areas. scientific collections both historical and modern also have the potential to assist in the isolation and identification of biopharmaceuticals. for example, taxol as a treatment for ovarian cancer has been derived from taxus brevifolia (pacific yew) (stierle, 1994) and such ethno botanical sources or “traditional” medicines have become the subject 2 http://www.emeraldashborer.info/ 3 http://www.cbd.int/abs/ of discussion for recognition and possible compensation under patent regimes. these kinds of advances in information technologies and tissue sampling techniques for molecular biology and genomics are creating new applications for traditional collections beyond their original intents as objects of study and as vouchers that allow for the verification of research results. it is unfortunate therefore that the preservation of, and providing access to, scientific collections is not fully seen as the “big science” that it truly is on an international scale. researchers are spread out across the globe, searching for the new and unexpected, rather than working together on a single project or at a major new facility. the resulting lack of basic knowledge puts at risk other research and development investments in areas such as biotechnology, genomics, agriculture, forestry, fisheries and aquaculture, and public health. the promise of new information technologies, from genomics to geographic information systems, is that research results can be captured in a more systematic way, and will contribute to a greater understanding of the whole. significance of scientific collections the importance of scientific collections is exemplified to great effect in the united states of america, where the office of science and technology policy (within the office of the president) has recognized scientific collections as critical scientific infrastructure since 2005. an interagency working group on scientific collections was created to conduct a survey of scientific collections held by federal agencies and collaborated with the national science foundation to document collections not owned by the federal government. among its findings the working group has identified the need for: a) a comprehensive, government-wide mechanism for the responsible management of federal collections that addresses shortand long-term preservation issues both within and between agencies, and, b) the establishment of an information clearinghouse, for agencies to share policies, procedures, and other information4 the digitization of this 4 http://www.whitehouse.gov/sites/default/files/sci-collectionsreport-2009-rev2.pdf 132 http://www.emeraldashborer.info/ http://www.cbd.int/abs/ http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf baird – leveraging the fullest potential of collections infrastructure is a key element in making these specimens accessible, searchable and distributable to present and future researchers wherever they may be physically located. what is currently being digitized and to what effect? access to digitized data at the specimen-level has been a vital means to distribute information and to create the potential for that information to generate knowledge. the global biodiversity information facility5 is widely recognized as a leading example of the multiplier effect that is created by pooling the accessibility of specimen and observation data under common standards and harnessing the power of distributed datasets. and yet, the very strength of such a model is almost equally its weakness. the task at hand is made so daunting because of several factors: the sheer number of specimens, the variety of analogue and digital formats in which associated data are held, the fact they are dispersed in repositories large and small in all corners of the world, and the reality that they are amassed over decades and even centuries. the development and application of data standards which provide common descriptive language and common binary formats are critical tools to facilitate this access. metadata can be used to capture the essential information describing the general nature, as well as the how and when and by whom, of a particular set of data, as well its format. this has the effect of characterizing larger groupings of information into formats that are discoverable by users and are sufficiently described to stimulate interest in the grouping by the end user. the metadata is in effect the finding aid or catalogue that can lead the researcher onto further discovery. metadata standards and their application are therefore critical to rapidly characterize and quantify specimen information, but impose limits on what can be discovered by the data user. metadata can lead to promising paths of discovery but will rarely reach beyond the function of way finder. the scientific collections that can be accessed through metadata standards are not limited to biological or natural history holdings. a number of 5 http://www.gbif.org/ national and international initiatives have also been examining and highlighting the significance of material culture and their related data, in addition to biological material, and these are enumerated below: 1) the australian research council’s network for early european research (neer) is one of many digitization initiatives (burrows, 2006) which have made use of the power of metadata –data about data. europa inventa (the australian collections service) was envisaged as a gateway to all early european items in australian collections (burrows 2008), including manuscripts, papers, artworks, maps, furniture, fabrics, scientific instruments, and other material culture, with a focus on unique items and those with specific associations rather than on massproduced items such as early printed books. links from descriptive catalogue records to digitized images held on the distributed servers of the holding institutions provide this virtual access, and commissioning of digitization work in these repositories was also envisaged as part of the project. 2) colombia has engaged in a large scale systematic program of digitization under policy initiatives established by the ministry of culture. at an international workshop on digital preservation and copyright organized by the world intellectual property organization (wipo) in geneva in july 2008, participants heard general details of two important national initiatives. in recognition of 200 years of independence in 2010, the bicentennial digitization program will digitize and publish on the web the most important documents of the time from colombia’s national library, national archives, the national museum and the central bank’s library. the resulting standards and best practices will be leveraged to give form to a national digitization plan. as well, the colombian digital library6 has been established as a joint program between thirteen universities. working with colciencias (the national science committee) and renata7 (the national academic web of high technology), they will define standards and mechanisms for 6 http://www.bdcol.org:8080/ 7 http://www.renata.edu.co/ 133 http://www.gbif.org/ http://www.renata.edu.co/ baird – leveraging the fullest potential of collections digitization. this collaboration may overcome economic and legal issues that impede preservation of the works and access by researchers and the general public. 3) in the realm of biodiversity-related collections, twelve federal departments and agencies within the government of canada have been collaborating in a horizontal initiative known as the federal biodiversity information partnership (fbip)8. the ‘natural capital’ of the country’s biological resources from molecules and genes to organisms and ecosystems – are being positioned as resources that have the potential to yield benefits ranging from bio economic and environmental to social and human health. the key to unlocking these benefits is the creation of integrated bio-information systems of scientific information with interoperability of databases and national specimen repositories. 4) the european community demonstrated its recognition of collection institutions as important research infrastructures by its support and funding of the european distributed institute of taxonomy (edit)9. this “network of excellence” project is to integrate institutional policies and infrastructures of the participating 25 collection institutions. the continuing funding for the synthesis of systematics resources (synthesys)10 supports an initiative comprised of 20 european natural history museums and botanic gardens. its objective is to create an integrated european infrastructure for researchers in the natural sciences through access and networking activities. additionally, a preparatory project was granted for the preparation of a large scale and long term research facility (“lifewatch”)11, uniting these networks with those in the areas of marine and terrestrial eco technology committee released its follow up logy. 5) as recently as august 2008 in the united kingdom, the house of lords science and div unique global research infrastru 8 http://www.cbif.gc.ca/fbip/fbip_e.php 9 http://www.e-taxonomy.eu/ 10 http://www.synthesys.info/ 11 http://www.lifewatch.eu/ report on systematics and taxonomy12 (see para.2.13), measuring progress towards halting the decline in biodiversity is a key international obligation which cannot be achieved without baseline knowledge of biodiversity. creating baselines and monitoring change is dependent upon the availability of taxonomic expertise across the range of living organisms. systematic biology underpins our understanding of the natural world. a decline in taxonomy and systematics in the uk would directly and indirectly impact on the government's ability to deliver across a wide range of policy goals. the conclusions of the report also emphasize the need for financial and infrastructure support in digitizing biological collections to facilitate the aggregation of collections data. 6) on the global scale, the need for large scale collaboration is equally reflected in the reports of the megascience forum of the organization for economic cooperation and development (oecd) that led to the foundation of gbif13, and the global taxonomy initiative (gti) of the convention on biological ersity14. 7) an initiative by the netherlands in 2006 under the global science forum of the oecd15 led to workshops convened in leiden (june 2007), and washington d.c. (july 2008) with a steering committee guiding a proposal to establish an international coordinating mechanism supporting scientific collection-based institutions, in their collective role as part of a cture (see pp.1-2), scientific collections are essential parts of the research infrastructure of all countries with scientific enterprises, and they are critical to many areas of science, from microbiology to space science. 12 http://www.publications.parliament.uk/pa/ld200708/ldselect/ldscte ch/162/162.pdf 13 www.gbif.org 14 http://www.cbd.int/gti/ 15 http://www.oecd.org/topic/0,3373,en_2649_34319_1_1_1_1_3741 7,00.html 134 http://www.cbif.gc.ca/fbip/fbip_e.php http://www.e-taxonomy.eu/ http://www.synthesys.info/ http://www.lifewatch.eu/ http://www.publications.parliament.uk/pa/ld200708/ldselect/ldsctech/162/162.pdf http://www.publications.parliament.uk/pa/ld200708/ldselect/ldsctech/162/162.pdf http://www.gbif.org/ http://www.cbd.int/gti/ http://www.oecd.org/topic/0,3373,en_2649_34319_1_1_1_1_37417,00.html http://www.oecd.org/topic/0,3373,en_2649_34319_1_1_1_1_37417,00.html baird – leveraging the fullest potential of collections national governments share an interest in finding answers to basic research questions and many applied research challenges, and no one nation has all the assets to pursue major research challenges independently…the mission of an international coordinating mechanism for scientific collections would include bal-scale research ure e ated foster capacitynal standards deemed necessary ld in with the user ssessment tools for fic ce broader societal concerns/policies the following: • enable glo activities • promote an international cult of collections as large-scal distributed infrastructure • improve access to and mobility of collection objects and associated data, and the people associ with them; building • identify and integrate existing standards of community practice, and develop additio in order to fulfill this mission, this coordinating mechanism shou undertake the following actions: • create a research roadmap coordination community • create self-a collections • set standards of practice • promote research on scienti collections and collections management • provide opportunities for the global collections workfor • provide a clearinghouse mechanism/interface between collection-based science and the report16, submitted to the gsf in krakow, poland (oecd, 2008) elaborates an implementation plan for such a coordinating mechanism. strong satisfaction was expressed by the gsf delegates for the work done to date, with germany, france, australia, canada, belgium, holland, and the united kingdom registering formal comments of support. scientific collections international or scicoll17 has emerged as a nascent program, advancing on the strength of a well elaborated strategic plan, governance model and a program of outreach activities. in conclusion, a globally coordinated approach to the digitization of a critical mass of scientific specimens and specimen-related data is highly desirable and required, to maximize the value of these collections to civil society and to support the advancement of our scientific knowledge globally. a more cohesive and inclusive approach to the digitization of scientific collections is highly desirable for all of these reasons and more. references blancou j, chomel bb, belotto a, meslin fx. 2005. emerging or re-emerging bacterial zoonoses: factors of emergence, surveillance and control. vet res. may-jun;36(3):507-22. 18 burrows, t. 2006. network for early european research (neer)digital services-background, perth, australia. 19 burrows, t. 2008. europa inventa (australian collections services), perth, australia. 20 chapman, a.d. 2009. numbers of living species in australia and the world. report for the australian biological resources study, 2nd ed. 80 pp. canberra, australia. 21 ______, 2009 . oecd global science forum second activity on policy issues related to scientific research collections: final report on findings and recommendations submitted october 2010 to 16 http://www.oecd.org/dataoecd/7/58/42237442.pdf 17 www.scicoll.org 18 http://www.ncbi.nlm.nih.gov/pubmed/15845237 19 http://confluence.arts.uwa.edu.au/display/digital/neer+digita l+services+-+background 20 http://confluence.arts.uwa.edu.au/display/digital/europa+inven ta 21 http://www.environment.gov.au/biodiversity/abrs/publications/oth er/species-numbers/2009/pubs/nlsaw-2nd-complete.pdf 135 http://www.ncbi.nlm.nih.gov/pubmed?term=%22blancou%20j%22%5bauthor%5d http://www.ncbi.nlm.nih.gov/pubmed?term=%22chomel%20bb%22%5bauthor%5d http://www.ncbi.nlm.nih.gov/pubmed?term=%22belotto%20a%22%5bauthor%5d http://www.ncbi.nlm.nih.gov/pubmed?term=%22meslin%20fx%22%5bauthor%5d javascript:al_get(this,%20'jour',%20'vet%20res.'); javascript:al_get(this,%20'jour',%20'vet%20res.'); http://www.oecd.org/dataoecd/7/58/42237442.pdf http://www.scicoll.org/ http://www.ncbi.nlm.nih.gov/pubmed/15845237 http://confluence.arts.uwa.edu.au/display/digital/neer+digital+services+-+background http://confluence.arts.uwa.edu.au/display/digital/neer+digital+services+-+background http://confluence.arts.uwa.edu.au/display/digital/europa+inventa http://confluence.arts.uwa.edu.au/display/digital/europa+inventa http://www.environment.gov.au/biodiversity/abrs/publications/other/species-numbers/2009/pubs/nlsaw-2nd-complete.pdf http://www.environment.gov.au/biodiversity/abrs/publications/other/species-numbers/2009/pubs/nlsaw-2nd-complete.pdf baird – leveraging the fullest potential of collections 136 the oecd global science forum, krakow, poland. 22 martin, a. 2008 international workshop on digital preservation and copyright. world intellectual property organization, geneva. 23 national science and technology council, committee on science, interagency working group on scientific collections, 2009. scientific collections: mission-critical infrastructure of federal science agencies. office of science and technology policy, washington, d.c. 24 u.s. congress, office of technology assessment, 1993. harmful non-indigenous species in the united states, ota-f-565 u.s. government printing office, washington, d.c. 25 renaud, m.-a. 2008. dna barcode, trade, plant health and quarantine in the canadian ornamental industry. presented at canadian barcode of life network, 2nd scientific symposium, royal ontario museum toronto. 22 http://www.oecd.org/dataoecd/7/58/42237442.pdf 23 http://www.wipo.int/edocs/mdocs/copyright/en/wipo_cr_wk_ge_0 8/wipo_cr_wk_ge_08_www_105896.pdf 24 http://www.whitehouse.gov/sites/default/files/sci-collectionsreport-2009-rev2.pdf 25 http://www.fas.org/ota/reports/9325.pdf stierle, a. et al. 1994. endophytic fungi of pacific yew (taxus brevifolia) as a source of taxol, taxanes, and other pharmacophores. bioregulators for crop protection and pest control chapter 6, pp 64–77 chapter doi: 10.1021/bk-1994-0557.ch006 acs symposium series, vol. 557 . sumption k.j. and flowerdew, j.r., 1985. the ecological effects of the decline in rabbits (oryctolagus cuniculus l.) due to myxomatosis. mammal review,15: 151–186. re-published online 2008 at 26 sutherland of houndwood et al. 2008. systematics and taxonomy: follow-up. report of the house of lords science and technology committee, london 27 simberloff, d. 1996 impacts of introduced species in the united states consequences vol.2 no.2 united states global research information office, washington 28 wong, e. h.-k. and hanner, r.h. 2008. dna barcoding detects market substitution in north american seafood. food research international,volume 41, issue 8, october 2008, p.p. 828-837 29 26 http://onlinelibrary.wiley.com/doi/10.1111/j.13652907.1985.tb00396.x/references 27 http://www.publications.parliament.uk/pa/ld200708/ldselect/ldscte ch/162/162.pdf 28 http://www.gcrio.org/consequences/vol2no2/article2.html 29 http://www.sciencedirect.com/science?_ob=articleurl&_udi=b6 t6v-4syjs3m2&_user=8844459&_coverdate=10%2f31%2f2008&_rdoc=1&_ fmt=high&_orig=search&_sort=d&_docanchor=&view=c&_searc hstrid=1439577567&_rerunorigin=scholar.google&_acct=c0001 09520&_version=1&_urlversion=0&_userid=8844459&md5=ca1 f16fbe2a05e7b9fc1916102e565b3 http://www.oecd.org/dataoecd/7/58/42237442.pdf http://www.wipo.int/edocs/mdocs/copyright/en/wipo_cr_wk_ge_08/wipo_cr_wk_ge_08_www_105896.pdf http://www.wipo.int/edocs/mdocs/copyright/en/wipo_cr_wk_ge_08/wipo_cr_wk_ge_08_www_105896.pdf http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf http://www.fas.org/ota/reports/9325.pdf http://onlinelibrary.wiley.com/doi/10.1111/j.1365-2907.1985.tb00396.x/references http://onlinelibrary.wiley.com/doi/10.1111/j.1365-2907.1985.tb00396.x/references http://www.publications.parliament.uk/pa/ld200708/ldselect/ldsctech/162/162.pdf http://www.publications.parliament.uk/pa/ld200708/ldselect/ldsctech/162/162.pdf http://www.gcrio.org/consequences/vol2no2/article2.html http://www.sciencedirect.com/science?_ob=articleurl&_udi=b6t6v-4syjs3m-2&_user=8844459&_coverdate=10%2f31%2f2008&_rdoc=1&_fmt=high&_orig=search&_sort=d&_docanchor=&view=c&_searchstrid=1439577567&_rerunorigin=scholar.google&_acct=c000109520&_version=1&_urlversion=0&_userid=8844459&md5=ca1f16fbe2a05e7b9fc1916102e565b3 http://www.sciencedirect.com/science?_ob=articleurl&_udi=b6t6v-4syjs3m-2&_user=8844459&_coverdate=10%2f31%2f2008&_rdoc=1&_fmt=high&_orig=search&_sort=d&_docanchor=&view=c&_searchstrid=1439577567&_rerunorigin=scholar.google&_acct=c000109520&_version=1&_urlversion=0&_userid=8844459&md5=ca1f16fbe2a05e7b9fc1916102e565b3 http://www.sciencedirect.com/science?_ob=articleurl&_udi=b6t6v-4syjs3m-2&_user=8844459&_coverdate=10%2f31%2f2008&_rdoc=1&_fmt=high&_orig=search&_sort=d&_docanchor=&view=c&_searchstrid=1439577567&_rerunorigin=scholar.google&_acct=c000109520&_version=1&_urlversion=0&_userid=8844459&md5=ca1f16fbe2a05e7b9fc1916102e565b3 http://www.sciencedirect.com/science?_ob=articleurl&_udi=b6t6v-4syjs3m-2&_user=8844459&_coverdate=10%2f31%2f2008&_rdoc=1&_fmt=high&_orig=search&_sort=d&_docanchor=&view=c&_searchstrid=1439577567&_rerunorigin=scholar.google&_acct=c000109520&_version=1&_urlversion=0&_userid=8844459&md5=ca1f16fbe2a05e7b9fc1916102e565b3 http://www.sciencedirect.com/science?_ob=articleurl&_udi=b6t6v-4syjs3m-2&_user=8844459&_coverdate=10%2f31%2f2008&_rdoc=1&_fmt=high&_orig=search&_sort=d&_docanchor=&view=c&_searchstrid=1439577567&_rerunorigin=scholar.google&_acct=c000109520&_version=1&_urlversion=0&_userid=8844459&md5=ca1f16fbe2a05e7b9fc1916102e565b3 http://www.sciencedirect.com/science?_ob=articleurl&_udi=b6t6v-4syjs3m-2&_user=8844459&_coverdate=10%2f31%2f2008&_rdoc=1&_fmt=high&_orig=search&_sort=d&_docanchor=&view=c&_searchstrid=1439577567&_rerunorigin=scholar.google&_acct=c000109520&_version=1&_urlversion=0&_userid=8844459&md5=ca1f16fbe2a05e7b9fc1916102e565b3 http://www.sciencedirect.com/science?_ob=articleurl&_udi=b6t6v-4syjs3m-2&_user=8844459&_coverdate=10%2f31%2f2008&_rdoc=1&_fmt=high&_orig=search&_sort=d&_docanchor=&view=c&_searchstrid=1439577567&_rerunorigin=scholar.google&_acct=c000109520&_version=1&_urlversion=0&_userid=8844459&md5=ca1f16fbe2a05e7b9fc1916102e565b3 how are scientific collections and their data used? significance of scientific collections what is currently being digitized and to what effect? references microsoft word fernandezetallayoutproof4 biodiversity informatics, 6, 2009, pp. 36-52 36 locality uncertainty and the differential performance of four common niche-based modeling techniques miguel a. fernandez 1,5 , stanley d. blum 2 , steffen reichle 3 , qinghua guo 1 , barbara holzman 4 , and healy hamilton 5 1 sierra �evada research institute, university of california merced 2 research informatics, california academy of sciences 3 the �ature conservancy 4 department of geography & human environmental studies, san francisco state university 5 center for biodiversity research, california academy of sciences abstract. we address a poorly understood aspect of ecological niche modeling: its sensitivity to different levels of geographic uncertainty in organism occurrence data. our primary interest was to assess how accuracy degrades under increasing uncertainty, with performance measured indirectly through model consistency. we used monte carlo simulations and a similarity measure to assess model sensitivity across three variables: locality accuracy, niche modeling method, and species. randomly generated data sets with known levels of locality uncertainty were compared to an original prediction using fuzzy kappa. data sets where locality uncertainty is low were expected to produce similar distribution maps to the original. in contrast, data sets where locality uncertainty is high were expected to produce less similar maps. bioclim, domain, maxent and garp were used to predict the distributions for 1200 simulated datasets (3 species x 4 buffer sizes x 100 randomized data sets). thus, our experimental design produced a total of 4800 similarity measures, with each of the simulated distributions compared to the prediction of the original data set and corresponding modeling method. a general linear model (glm) analysis was performed which enables us to simultaneously measure the effect of buffer size, modeling method, and species, as well as interactions among all variables. our results show that modeling method has the largest effect on similarity scores and uniquely accounts for 40% of the total variance in the model. the second most important factor was buffer size, but it uniquely accounts for only 3% of the variation in the model. the newer and currently more popular methods, garp and maxent, were shown to produce more inconsistent predictions than the earlier and simpler methods, bioclim and domain. understanding the performance of different niche modeling methods under varying levels of geographic uncertainty is an important step toward more productive applications of historical biodiversity collections. key words. georeferencing; spatial uncertainty; ecological niche modeling; comparative performance; fuzzy kappa our maps of species’ distributions ultimately derive from primary observations of their occurrence in nature. because these data are typically sparse in comparison to the complete range of a species, biologists have devised a variety of methods to visualize and analyze species ranges based on field samples. these range from simply plotting occurrence points on maps, to drawing a free-form line around peripheral locality records. recently, researchers interested in species’ distributions have been able to integrate spatial tools and environmental data to produce probability distribution maps that indicate variation in habitat suitability, maps which convey more _________________________ correspondence email: mfernandez5@ucmerced.edu. information than either point locality maps or outline maps. known also as ecological niche models (enm; sensu grinnell 1917), these maps are the result of integrative algorithms embedded in a gis framework that use the taxonomic and geographic data associated with specimens and/or observations and fine scale environmental data to produce a set of rules that identify the environmental space where the species was collected or observed (peterson and vieglais 2001). this environmental space can be projected onto geographic space to identify appropriate conditions where the species may occur, resulting in a modeled distribution. despite the fact that these presence-only inferential maps are abstract representation of species ranges, they are still valuable summaries of fernandez et al. – uncertainty in locality data and ecological niche modeling 37 biogeographic information, and have been applied to a broad range of topics, from theoretical ecology and evolution (leathwick and whitehead 2001; hugall et al. 2002; graham et al. 2004), to practical uses in conservation (bustamante 1997; corsi et al. 1999; anderson et al. 2002; raxworthy et al. 2003; araújo et al. 2004) agriculture, invasive species (higgins et al. 2000; welk et al. 2002; underwood et al. 2004), and human health (mills and childs 1998; peterson and shaw 2003). while tremendous progress has been achieved on many aspects of building and evaluating enm (guisan and zimmermann 2000; pearce and ferrier 2000; williams and hero 2001; hirzel et al. 2002; stockwell and peterson 2002; brotons et al. 2004; reese et al. 2005; barry and elith 2006; pearson et al. 2006; wisz et al. 2008), enhanced frameworks for assessing errors and uncertainties have not been fully developed. specifically, uncertainty in the organism occurrence data (see fig. 1) has not been fully explored (graham et al. 2004; murphy et al. 2004; soberon and peterson 2004; wieczorek et al. 2004; rowe 2005; guo et al. 2008). understanding the susceptibility of enm methods to the positional error associated with a collection event becomes a critical factor in selecting a method to use in a particular case. comparing e�m performance against uncertainty a predicted species distribution is generally determined by three elements: the algorithm or modeling method, the environmental layers upon which it is based, and the occurrence data. although researchers have explored how each of these elements contributes separately or together to the overall performance of the technique, as yet, there is no agreement on the influence of uncertainties on enm. some studies show that different methods perform surprisingly similarly (peterson et al. 2007; peterson et al. 2008), while others studies show that alternative enm produce highly distinct outputs when predicting species’ geographic ranges (manel et al. 1999; elith et al. 2006; pearson et al. 2006; phillips et al. 2006; kelly et al. 2007; peterson et al. 2007; tsoar et al. 2007; ortega-huerta and peterson 2008). further research is required to address these discrepancies in model performance. specifically, standardized and improved parameterization and enhanced evaluation tools are needed to tease apart these differences in modeling outputs (araújo and guisan 2006; peterson et al. 2008). while few studies have measured the sensitivity of distribution models to grid cell size in the environmental layers (guisan et al. 2007), others have addressed the effect of remote sensing derived products as alternative environmental layers in enm (parra et al. 2004; roura-pascual et al. 2004; peterson et al. 2006; zimmermann et al. 2007; bradley and fleishman 2008; buermann et al. 2008). limited studies attempted to incorporate true absence and more meaningful pseudo-absence data in enm (manel et al. 2001; brotons et al. 2004; engler et al. 2004; chefaoui and lobo 2007; phillips 2008). numerous tests have also addressed the effect of occurrence data quantity on enm (peterson and cohoon 1999; stockwell and peterson 2002; kadmon et al. 2003; hernandez et al. 2006; pearson et al. 2007; wisz et al. 2008). however quality in species occurrence data can also have profound consequences in enm. localities may be geographically biased, for example, highly correlated with rivers and access roads (reddy and davalos 2003), or collected using different sampling intensity and sampling methods (anderson 2003). localities that have been retrospectively georeferenced have uncertainty associated with the lack of geographic details in the textual descriptions (beaman et al. 2004; rowe 2005; chapman and wieczorek 2006). more standardized techniques have been developed that allow a better quantification of the positional error of occurrence data (murphy et al. 2004; wieczorek et al. 2004; guralnick et al. 2007; guo et al. 2008). the effect of the positional error on resulting enm output using different methodologies has been underexplored. recently, graham et al. (2008) evaluated how locality uncertainty affects the performance of ten common niche modeling techniques by comparing a control model calibrated using the original accurate data to an error treatment where the positional accuracy of the data was degraded randomly in a radius of 5 km. even though they demonstrated that model performance can change markedly with increased locality uncertainty. their single randomization treatment is not sufficient to establish the relationship between the magnitude of locality uncertainty and enm performance. close to 2.5 billion specimens (duckworth et al. 1993) have been collected and housed in natural fernandez et al. – uncertainty in locality data and ecological niche modeling 38 figure 1. differences in geographic uncertainty between two labels of specimens housed in natural history museum collections. left, a vague description of the collection location; right, a much more precise description of the collecting location. history museums by different collectors, at different times, with different sampling techniques (see fig. 1). as a consequence, the geographic information associated with specimen collections has very different levels of geographic uncertainty. many historical localities were recorded only as textual descriptions, without geographic coordinates, which effectively makes them unavailable to gis-based analyses. as discussed above, the subsequent interpretation of textual localities as geocoordinates, known as retrospective georeferencing can introduce still greater spatial uncertainty (proctor 2004; wieczorek et al. 2004; rowe 2005). in order to make more effective use of the wealth of biodiversity information stored in natural history museums, it is critical to fully explore the sensitivity of enm techniques to different levels of geographic uncertainty in the organism occurrence data. only by quantifying how uncertainty interacts with modeling methods and landscape variability will we be able to understand the reliability of predicted distributions or their suitability to particular uses. methods because distribution modeling outputs differ, no simple statistic is available to measure intermodel performance across all approaches (phillips et al. 2006; lobo et al. 2008; peterson et al. 2008). many commonly used methods give results as probability surfaces, rather than binary distributions, in which the species is predicted to occur or not occur in a particular grid cell. evaluating these models directly requires selecting an arbitrary threshold value to create the binary prediction, which might then result in under and over prediction. given that our primary interest was to assess how accuracy degrades under increasing uncertainty, we chose to measure model performance indirectly, through the consistency of repeated simulations. in this study we used monte carlo simulations and a similarity measure to assess the consistency of predictions across three variables: locality uncertainty, niche modeling method, and species. we created data sets with known levels of locality uncertainty and compared them to an original prediction using a similarity measure, fuzzy kappa (discussed further below). we expected data sets where locality uncertainty is low to produce distribution maps that are similar to the original. in contrast, we expected data sets where locality uncertainty is high to produce maps that are less similar or more inconsistent. in addition, we wanted to examine whether the response to uncertainty would differ across modeling methods and whether taxonomic or landscape variability would also influence this sensitivity. ultimately, we would like to know what degree of data quality, estimated by maximum error distance we can tolerate to produce predictions that are sufficiently accurate. fernandez et al. – uncertainty in locality data and ecological niche modeling 39 species selection criteria three species of bolivian frogs were selected for this analysis: oreobates cruralis, leptodactylus elenae, and pleurodema marmoratum. these species were selected for the following reasons: (a) their geographic ranges are comparable in area; (b) none is narrowly endemic or broadly distributed; and (c) they represent each of the main geographic areas in bolivia, regions that are expected to have very different landscape characteristics: o. cruralis is widely distributed in the yungas region of bolivia, l. elenae is distributed in the lowlands of bolivia, commonly associated with savannas; and p. marmoratum is restricted to the highlands. simulating locality uncertainty by random displacement in our experiments, an “original data set” is the group of non-repeated collecting localities for each species, expressed as latitude and longitude, either taken by one of us using a gps (sr), or georeferenced by one of us (mf) (table 1 and fig. 2). even though we only use occurrences with positional uncertainty represented by maximum error estimates of less than 1 km, we note the goal of this exercise was not to evaluate how well the models fit the real distribution of the species, but to test what is the effect of degrading the localities across a broad range of positional accuracies. an “original ecological niche model” is the output produced by one of the modeling techniques using the original dataset and a set of 19 standard bioclimatic variables derived from worldclim 1.4 (hijmans et al. 2005) at a spatial resolution of ~1 km2. from this “original dataset” we generated 100 different “new data sets” that simulate an increased level of locality uncertainty using the random point generator arcview 3.x extension (jenness enterprises, 2005), which produces a random selection of points approaching a uniform distribution (see fig. 3). we used this tool to randomly displace every point in each of the original datasets to a new position within a selected buffer distance. each of these new 100 points per buffer size and per locality is combined randomly with other generated points for other localities to form a “new dataset”. this new dataset is composed of the same number of point localities as the “original dataset” but located at different distances within the selected buffer. therefore, buffer size represents our experimental model of locality uncertainty. this simulation is similar to the point-radius method of retrospective georeferencing described by wieczorek et al. (2004). this method encompasses a wide variety of processes that contribute to different degrees of uncertainty (see guo et al. 2008). although it is possible to derive a probability density function for each locality (guo et al. 2008), something other than equi-probable or even, these functions entail prior knowledge of the processes that produced the data, such as assumptions on referenced objects that are used to georeference species localities, and assumptions on spatial relationships that describe the species localities.. while these may be reasonable assumptions, their purpose is to minimize the effect of uncertainty and extract better information from occurrence data. that is not our purpose here. in this study we are measuring the effect of uncertainty, so our goal is to incorporate uncertainty in a reasonable and easily understood way. we chose to represent uncertainty as a circle around the original point, where any point in that area has an equal probability of selection. this is currently a common practice in estimating locality uncertainty in occurrence data derived from retrospective georeferencing (wieczorek et al. 2004; guo et al. 2008). modeling methods four distribution modeling techniques were used to predict the distributions for 1200 simulated datasets (3 species x 4 buffer sizes x 100 randomized data sets). thus, our experimental design produced a total of 4800 similarity measures, as each of these predicted distributions was compared to the prediction produced from the original data set and corresponding modeling method. two of the methods we used are based on a climatic envelope concept and presence only localities, bioclim (busby 1991) and domain (carpenter et al. 1993). the other two methods use both presence and pseudo-absence localities, garp (stockwell and peters 1999), which is based on a genetic algorithm, and maxent (phillips et al. 2006) which is based on the maximum entropy concept. fernandez et al. – uncertainty in locality data and ecological niche modeling 40 table 1. list of number of non-repeated localities for the original datasets per species species name number of localities oreobates cruralis 38 leptodactylus elenae 39 pleurodema marmoratum 29 bioclim relates occurrence localities to climatic conditions, and produces a single rule that identifies all areas with a similar climate to the locations of the species within a minimal rectilinear “climatic envelope”. in bioclim, user-specified thresholds for each environmental predictor are identified to define the multidimensional environmental space. this can be projected onto landscapes producing a model of appropriate climactic conditions for the species. however, the assumption that species’ distributions are controlled by a defined climatic envelope is largely simplistic. species ranges in nature are controlled by a complex combinations of factors and unlikely to be a box shape in the environmental space. domain is a tool based on a point-to-point similarity metric (gower metric). similarity between the site of interest and each of the recorded present occurrence locations is calculated by summing the standardized distance between the two points for each predictor variable. the standardization is achieved by dividing the distance by the predictor variable range for the presence sites, equalizing the contribution from each predictor variable. the standardized distance is subtracted from 1 to obtain the complementary similarity (carpenter et al. 1993). predictions are not to be interpreted as maps of probability of occurrence, but as a measure of classification confidence. neither of these two methods provides explanatory power of the relevant factors controlling the species’ distributions, nor statistically quantifies the variance, thus, the accuracy of the predictions is unknown (stockwell 2006). garp is a non-deterministic model that uses a machine learning approach to test several inferential algorithms (e.g. atomic, logistic regression, range rules, and negated range) in an iterative manner to develop multiple sets of rules that will provide multiple solutions given the same input. for each new iteration, garp divides the occurrences in: (a) training data, which is used to produce the rules that will define the model, and (b) testing data, which is used to internally evaluate the model based on omission and commission errors. in the next iteration the data is resampled, a new training and testing data set is produced, and the process starts over again. this process is repeated until the program can not create an improved model (stockwell and peters, 1999). since garp doesn’t produce a single probabilistic output, to deal with this stochasticity, multiple runs can be performed within the same garp session, producing a chosen number of output prediction maps. garp reports measures of omission and commission errors for each generated model, and provides the option to select a ‘best subset’ based on these accuracy measures. the predictions for the ‘best subset’ models can be arithmetically combined to produce a final predicted distribution map (anderson et al. 2003). maxent (phillips et al. 2004), estimates a species niche by finding the probability distribution of maximum entropy, subject to the constraint that the expected value of each environmental variable under this estimated distribution matches its empirical average. continuous environmental data can also be entered as both quadratic and product features, thereby adding further constraints to the estimation of the probability distribution by restricting it to be within the variance for each environmental predictor and covariance for each pair of environmental predictors. the program starts with a uniform probability distribution, and iteratively alters one weight at a time to maximize the likelihood of reaching an optimum probability distribution. the algorithm is guaranteed to converge, and therefore the outputs are deterministic. since the traditional implementation of maximum entropy is prone to over-fitting the probability distribution, maxent actually employs a relaxation method. it does not constrain the estimated distribution to the exact empirical fernandez et al. – uncertainty in locality data and ecological niche modeling 41 figure 2. original species occurrence datasets for the three species included in this study. average, but to within the empirical error bounds of the average value for a given predictor, in a procedure called ‘regularization’ (phillips et al. 2004). fuzzy kappa every predicted distribution was standardized (rescaled) into an idrisi andes compatible grid format in which cell values range from 0 to 100 (see below): ( ) ( ) a b x min x max min − = − where xb is the rescaled value of each cell in the raster layer, xa is the original value from the model output, and min and max are the minimum and maximum values from the model output, respectively. to compare the predicted distributions of simulated data sets against the original, we used the similarity measure called fuzzy kappa (hagen 2003), implemented in the map comparison kit 3.0 (visser and de nijs 2006). fuzzy kappa is based on the simple kappa algorithm, however, it enables the comparison of two maps (both categorical and noncategorical data), and produces a similarity statistic that represents the average similarity of the entire map. the principal benefit of using fuzzy kappa over kappa is that kappa is based on binary logic, where the result of comparing the values of two corresponding cells is either “equal” or “different.” in contrast, fuzzy kappa uses a fuzzy logic where fernandez et al. – uncertainty in locality data and ecological niche modeling 42 the measure of similarity is continuous and based on the values of corresponding cells, as well as the distance to similar cells within a buffer defined by the user. this is based on the notion that the fuzzy representation of a cell depends on the cell itself and its neighboring cells with correspondingly lesser weight. this key distinction allows fuzzy kappa not only to evaluate differences but actual levels of difference, and models a human assessment of similarity more closely than simple kappa (visser and de nijs 2006) (see fig. 5). fuzzy kappa is calculated in a similar manner as the traditional kappa: ( ) (1 ) fuzzy s e k e − = − where s is the average similarity over all cells based on fuzzy memberships, and e is the expected similarity. the fuzzy membership is used to account for the location error as shown in fig. 4. in this study, we used the gaussian distance decay functions to define the fuzzy membership (visser and de nijs 2006). detailed discussion regarding fuzzy kappa can be found in pontius (2000) and hagen-zanker et al. (2005). figure 3. random localities selected from a buffer zone, emulating different degrees of uncertainty in locality description. as shown in figure 4, the top a0 to a4 maps portray the ecological niche models based on the bioclim algorithm and increasingly degraded localities from left (original localities) to right (localities degraded in a buffer of 50 km). the second row of maps portrays the kappa map comparison based on consecutive comparison of the original ecological niche model (a0) to each map resulting from increasingly degraded localities (a1, a2, a3 and a4). the bottom row of maps represents the numerical fuzzy kappa map comparison based on consecutive comparison of the original ecological niche model (a2) to each of the maps created with degraded localities. even though both indexes show a decrease in similarity with increasing buffer size, the value of the kappa is too sensitive to small differences, and misses some of the basic similarity between the two maps. on the other hand, fuzzy kappa is a more conservative index that varies less dramatically when the position of a multi-pixel “object” shifts slightly, which makes it a better tool for measuring the similarity between two maps. experimental design we measured how the similarity of predicted distributions changes in response to buffer size, an experimentally controlled continuous variable, as well as two categorical variables, species and modeling method. the similarity measure, fuzzy kappa, varies between zero and one. our intention was to perform a general linear model (glm) analysis, which would enable us to measure simultaneously the effect of buffer size, modeling method, and species using a two-way analysis of variance with an ordinary least squares regression, as well as test for interactions among all variables. the full factorial model was specified as: sp mm bfr sp mm sp bfr mm bfr sp mm bfr+ + + × + × + × + × × where sp is the categorical effect for species, mm is the categorical effect for modeling method, bfr is the covariate, buffer, and interaction terms are specified with a multiplication symbol between the codes for the primary effects. the sample sizes were balanced, with every permutation of treatments evaluated with 100 simulated data sets. results the results of this study can be understood most directly through visualization. figure 6 shows box plots of similarity measures for the series of buffer sizes within each modeling method and species combination. several things are evident from this figure. first, large differences exist among the modeling methods; bioclim scores were highest, while maxent scores were lowest. second, very fernandez et al. – uncertainty in locality data and ecological niche modeling 43 figure 4. the top (a0 to a4) maps portray the enm based on bioclim. the second row portrays the kappa comparison. the bottom row represents the numerical fuzzy kappa map comparison. large differences also exist among the variances across treatment combinations; the largest variance is more than 700 times larger than the smallest. third, within most combinations of species and modeling method, the mean similarity score tends to decrease with increasing buffer size (i.e., locality uncertainty). fourth, the variance in similarity tends to increase with buffer size. fifth, the relationships between similarity and buffer size are not the same across combinations of species and modeling methods; i.e., there appear to be interaction effects between the categorical variables and the covariate. among the bioclim analyses for example, o. cruralis shows a strong relationship between similarity and buffer size, whereas the relationship is weaker in p. marmoratum. in contrast, this comparison is reversed in the domain analyses; o. cruralis shows a weaker relationship, while p. marmoratum shows a stronger one. the glm analysis assumes that deviations from expected are effectively summarized by a normally distributed random variant with equal variance across all treatment levels. because some cases show an increase in variance with a decrease in mean similarity, we tested for a correlation between mean similarity and its variance. the pearson correlation coefficient (r) between mean similarity and variance was -0.278, which has a probability of 0.028 in a one-tailed test. (we used a one-tailed test because we expected the variance to increase as the mean decreased.) we applied an arcsin transformation in an attempt to reduce this correlation; this transformation is commonly used with measures that range between zero and one. in the transformed data, the correlation (r) was reduced to -0.096, which has a one-tailed probability of 0.26. because the transformed data show a reduced and insignificant correlation, we used the transformed data in our primary analysis. the comparable box plots for the transformed data are shown in figure 7. while the correlation is reduced, the variances are still strongly heterogeneous across treatment combinations. the largest variance is still more than 130 times larger than the smallest. consequently, the probability values obtained in the primary analysis below can only be taken as broadly indicative. fernandez et al. – uncertainty in locality data and ecological niche modeling 44 the results of our glm analysis are shown in table 2. every term in the model is significant well beyond the commonly used 0.05 level. the fact that the interaction terms are significant means that the primary terms are not additive; the effect of any particular value depends on the values of the other variables. in particular, the rate at which consistency declines with uncertainty (the slope) depends on both the modeling method and the species. a more detailed view of our results can be seen in figure 8. these histograms show the distributions of similarity scores for each of the 48 permutations of the primary parameters. we include these graphs because the assumptions of normally distributed error terms and homogenous variances within groups are violated. these histograms show how the distributions of similarity scores change across the experimental variables. in 9 of the 12 combinations of species by modeling-method (columns of histograms in fig. 8a, b, and c) the distributions are close to normal and have similar variance across buffer-size. in the other three cases, the distributions change markedly with buffer size. the scores for l. elenae modeled with domain are skewed to the left at 5 and 10 km, become flatter at 25 km, and become skewed to the left again at 50 km. at the smallest buffer size, the scores for p. marmoratum and domain cluster toward the upper range with a sparse tail to the left. the maximum and minimum scores, and hence the range, do not change much between 5 and 50 km, but the distribution goes from skewed to flat and the variance gets 100 times larger from the smallest buffer size to the largest. in the p. marmoratum and maxent analyses, similarity scores cluster tightly in the 5 km simulations, while the distribution flattens and the mode decreases at the larger buffer sizes. discussion the range of uncertainty used in this study, 5 to 50 km, is realistic and meaningful in comparison to both the degree of uncertainty that exists in real data and the resolution or scale of various gridded environmental surfaces that are routinely employed in distribution modeling: 1 km to 1/2° cell sizes (hijmans et al. 2005; mitchell and jones 2005). furthermore, the variable specificity of historical localities introduces geographic uncertainty well within the range of the buffer sizes tested here. thus these results should help inform users of retrospectively georeferenced data regarding the distribution modeling methods that are most and least sensitive to degree of specimen locality uncertainty. we expect the difference between environmental space at a given point a and b to be inversely proportional to the distance that separates these two points; in other words, the closer the points in geographic space, the more similar they should be in terms of environmental space (tobler 1970). as a consequence, points selected from a 50 km buffer should be more different from the original point and from each other than points selected from the 5 km buffer. this environmental space translated into geographic space can have profound consequences in the modeling outputs. one possible outcome is that the area of the predicted distribution will be proportional to the differences among the points used to train the models, in other words, the model will become more general (see fig. 10). however, comparing predicted areas of suitability has one major difficulty that forms the basis of our choice to use fuzzy kappa: the issue of threshold selection. to measure the relationship between predicted area and buffer size, a threshold must be selected and binary outputs must be compared. the relationship between predicted area and buffer size and the issue of threshold selection are two very important elements deserving of further attention that we did not explicitly evaluate in this paper. in this study, we did not address the issue of spatial autocorrelation explicitly. there are two types of spatial autocorrelation that will influence the effective sample size of localities: 1) the spatial autocorrelation among species occurrence localities, and 2) the spatial autocorrelation within the buffer. although dormann et. al. (2007) suggest that differences in parameter estimates and inference between spatial and non-spatial models are small, i.e., (the spatial models accounted for spatial autocorrelation, while the non-spatial models did not), this problem may also depend on the degree of environmental heterogeneity across sampled environmental space. fernandez et al. – uncertainty in locality data and ecological niche modeling 45 table 2. summary of ecological niche modeling parameters under the four methods used in this study. bioclim domain garp maxent software used diva gis 5.4 (hijmans et al. 2001). diva gis 5.4 (hijmans et al. 2001). desktop garp 1.1.3 (kansas university) maxent 2.3 (phillips et al. 2004) removal of duplicated localities yes yes yes yes outlier detection no no ---- parameters details percentile used: 0.025 --atomic, range, negated range, and logit rules. regularization multiplier = 1 random test % = 0 internal evaluation --training: 50% localities 20 best-subset models training: 50% localities 20 best-subset models yes outputs rescaled from 0 to 100 yes yes yes yes table 3. anova table for the general linear model analysis of transformed similarity scores. source partial ss df ms f prob > f model 306.346 23 13.31940 3188.9 <0.001 buffer 9.343 1 9.34283 2236.9 <0.001 species 1.364 2 0.68186 163.3 <0.001 modeling method 122.772 3 40.92412 9798.0 <0.001 buffer*species 0.063 2 0.03171 7.6 <0.001 buffer* modeling method 3.099 3 1.03310 247.3 <0.001 species* modeling method 5.228 6 0.87128 208.6 <0.001 buffer*species* modeling method 2.752 6 0.45871 109.8 <0.001 residual 19.948 4776 0.00418 total 326.295 4799 0.06799 the glm analysis shows that modeling method has the largest effect on similarity scores and uniquely accounts for 40% of the total variance in the model. the second most important factor was buffer size, but it uniquely accounts for only 3% of the model. this may seem like a small percentage, but buffer size interacts with the categorical variables, which obscures the effect of buffer size alone. further interpretation of the proportion of variance is also ill-advised because buffer size is an experimenter-controlled variable, so the proportion of variance it explains is determined by the range of values we chose as inputs to the simulation. among the modeling methods, bioclim and domain produced distribution maps that were overall more similar to the original maps than either garp or maxent. given the popularity of garp and maxent, we were surprised to see these methods produce maps that were significantly less consistent than the two older and simpler methods. bioclim and domain behaved as expected in showing a decline in consistency with increasing locality uncertainty, but the other two methods were less sensitive to uncertainty, either consistently (garp), or in two of three cases (maxent). garp showed the least sensitivity to locality uncertainty; distributions generated from fernandez et al. – uncertainty in locality data and ecological niche modeling 46 figure 5. (a) two sets of maps, the first set with a slight difference in the position of the red cells; the second set with a more perceivable difference in the position of the red cells, but identical results for the kappa statistic. (b) the comparison of the same two sets of maps by a numerical fuzzy kappa algorithm. grayscales in the comparison map indicate the level of similarity, darker gray indicates less similarity, and lighter gray indicates more similarity. numerical fuzzy kappa is capable of discriminate differences between two maps based on distance decay function with constant value set by the user. points with up to 50 km of uncertainty were only moderately less similar to the originals than those generated from points with only a maximum of 5 km uncertainty. maxent distributions showed the lowest consistency and moderate sensitivity to locality uncertainty. several reasons may contribute to the differences: 1. the bioclim model identifies locations where all environmental factors fall within certain percentiles (e.g., 95%) of the observation records (busby 1986). therefore, unless a significant number of extreme large or small values are changed when increasing the buffer size, the locality uncertainty will have relatively little effect on the modeling results. 2. the domain method assigns a classification value to an unknown site based on the distance of its closest similar site in environmental space. the effort on the locality variation is local, and even extreme values are found, they will only influence some nearby points in environmental space. 3. garp is based on genetic algorithms, which aim to find exact or approximate solutions to an optimization or search problem. garp can be considered a non-parametric machine learning algorithm which normally makes few assumptions about the data distribution, and is more robust to data outliers. however, variation of the fuzzy kappa values is greater than that of bioclim and domain methods. this is due to the fact that variation also comes from the stochastic generation of rule sets for the garp method and the random sampling of the background area, which will generate slightly different results in each iteration of the garp model. 4. maxent is a general-purpose machine learning method. similar to generalized linear model (glm) and generalized additive models (gam), maxent needs to make certain assumptions on the probability distributions. exponential models are normally used (phillips et al. 2006), which could be more sensitive to variation of the training data compared to nonparametric approaches. finally, we would like to emphasize that the variable we labeled “species” in these experiments is not actually a simple repetition of the experiment with another taxon, with all other factors equal. the three species selected in this study are all allopatric and come from regions where environmental parameters are expected to change very differently with comparable horizontal displacement or uncertainty. we expected similarity scores based on o. cruralis to decline sharply with increasing buffer size, because it is found in the yungas or eastern andes where the elevation gradients are steep. we expected p. marmoratum from the andean highlands to exhibit intermediate sensitivity to buffer size, and l. elenae from the amazonian lowlands to show the least sensitivity. our expectations were never fully born out. in comparison to the other species, l. elenae produced the highest scores in the garp and maxent analyses, but p. mamorata produced the lowest scores in three out of four cases. conclusions in several respects the results of our simulations were very different from what we expected. modeling method produced the largest effect; more than the primary experimental treatment of displacing original localities by up to 50 km, more than species differences, and more than topographic heterogeneity. fernandez et al. – uncertainty in locality data and ecological niche modeling 47 figure 6. summary distributions of similarity scores from 48 experiments, each made of 100 simulated data sets. standard tukey’s box-plots represent the fuzzy kappa similarity scores for the series of buffer sizes within each modeling method and species combination. p. marmoratum (mar); o. cruralis (cru); and l. elenae (ele). figure 7. summary distributions of similarity scores from 48 experiments, each made of 100 simulated data sets. standard tukey’s box-plots represent the transformed similarity score arcsin of fuzzy kappa for the series of buffer sizes within each modeling method and species combination. p. marmoratum (mar); o. cruralis (cru); and l. elenae (ele). an inescapable observation is that the newer and currently more popular methods, garp and maxent, were shown to produce more inconsistent predictions than the earlier and simpler methods, bioclim and domain. we do not necessarily interpret this to mean that bioclim and domain predict distributions more accurately than garp or maxent. a method that predicts with higher consistency may not be closer to the true distribution because it could be biased. for example, it might consistently over-predict the true distribution. on the other hand, a single prediction may not be very close to the true distribution if the method is relatively inconsistent. it is worth investigating further why the garp and maxent analyses, as we performed them here, gave inconsistent predictions. graham et al (2008) conclude similarly that not all modeling techniques are equally influenced by positional error. they suggested that some modeling techniques (maxent and boosted regression trees) are particularly “robust” to moderate levels of uncertainty in locality data. on the contrary, our research finds that garp is the most robust technique to positional error and domain the most sensitive of the four techniques we evaluated. this contradictory finding may be explained in that graham et al. (2008) addressed a slightly different but complementary problem. they evaluated the effect of degrading positional accuracy on the capacity of the model to predict accurately an independent dataset, using auc as a metric. in contrast, our goal was to evaluate how different modeling methods respond to varying levels of degraded positional accuracy. moreover, graham et al. (2008) used a single error treatment (5 km), while our study addressed multiple levels of locality uncertainty. our finding that domain is the most sensitive method and garp is the more robust method of the four enm tested here doesn’t imply that one method is better over others. we aim to provide information model performance relative to one additional source of uncertainty that will assist the user in model selection. finally, we sampled only four points along the potentially larger domain of uncertainty values. fernandez et al. – uncertainty in locality data and ecological niche modeling 48 a b c figure 8. histograms of similarity scores for the 48 permutations of experimentally controlled primary variables, buffer-size, modeling-method, and species. arcsin transformed similarity is along the x-axis, frequency is on the yaxis, and the histograms are grouped by buffer-size (rows), modeling-method (columns) and species across pages (a. o. cruralis, b. l. elenae, and c. p. marmoratum). the scaling and range of the axes are the same across all histograms. fernandez et al. – uncertainty in locality data and ecological niche modeling 49 consequently, we cannot evaluate whether the response of consistency to uncertainty is linear or curvilinear. it also remains to be determined what might happen beyond the limits we sampled. figure 9. hypothetical relationship between the buffer size and environmental space. left figure: increasing uncertainty buffer size, and right figure: the possible change of its environmental space (using temperature and precipitation as example environmental features) due to the increasing uncertainty. note that the actual shape in the feature space may not be the ellipse shape, and there are situations that don’t follow the same trend (e.g. environmental space may not be so homogeneous). acknowledgments we would like to thank jerry davis, matthew merrifield, michel koo, kristin byrd, simon ferrier, catherine graham, juan parra, robert hijmans, and kazuya naoki for their input on our analytical methods. we also would like to thank town peterson and the anonymous reviewers who help to improve this manuscript. this project was funded by the california academy of sciences lakeside fund for international students and the wwf russell e. train education for nature program. literature cited anderson, r. 2003. real vs. artefactual absences in species distributions: tests for oryzomys albigularis (rodentia: muridae) in venezuela. journal of biogeography 30:591-605. anderson, r., m. gomez-laverde, and a. peterson. 2002. geographical distributions of spiny pocket mice in south america: insights from predictive models. global ecology and biogeography 11:131141. anderson, r., d. lew, and a. peterson. 2003. evaluating predictive models of species’ distributions: criteria for selecting optimal models. ecological modelling 162:211-232. araújo, m., m. cabeza, w. thuiller, l. hannah, and p. williams. 2004. would climate change drive species out of reserves? an assessment of existing reserveselection methods. global change biology 10:16181626. araújo, m., and a. guisan. 2006. five (or so) challenges for species distribution modelling. journal of biogeography 33:1677-1688. barry, s., and j. elith. 2006. error and uncertainty in habitat models. ecology 43:413-423. beaman, r., p. museum, j. wieczorek, and s. blum. 2004. determining space from place for natural history collections. d-lib magazine 10:1082-9873. bradley, b., and e. fleishman. 2008. can remote sensing of land cover improve species distribution modelling? journal of biogeography 35:1158-1159. brotons, l., w. thuiller, m. araújo, and a. hirzel. 2004. presence-absence versus presence-only modelling methods for predicting bird habitat suitability. ecography 27:437-448. buermann, w., s. saatchi, t. smith, b. zutta, j. chaves, b. milá, and c. graham. 2008. predicting species distributions across the amazonian and andean regions using remote sensing data. journal of biogeography 35:1160-1176. busby, j. 1986. a biogeoclimatic analysis of �othofagus cunninghamii (hook.) oerst. in southeastern australia. austral ecology 11:1-7. busby, j. 1991. bioclim-a bioclimate analysis and prediction system. pp. 64–68 in c. margules, and m. austin, eds. nature conservation: cost effective biological surveys and data analysis. csiro, camberra. bustamante, j. 1997. predictive models for lesser kestrel falco naumanni distribution, abundance and extinction in southern spain. biological conservation 80:153-160. carpenter, g., a. gillison, and j. winter. 1993. domain: a flexible modelling procedure for mapping potential distributions of plants and animals. biodiversity and conservation 2:667-680. chapman, a., and j. wieczorek. 2006. guide to best practices for georeferencing. global biodiversity information facility. fernandez et al. – uncertainty in locality data and ecological niche modeling 50 chefaoui, r., and j. lobo. 2007. assessing the conservation status of an iberian moth using pseudoabsences. journal of wildlife management 71:25072516. corsi, f., e. dupre, and l. boitani. 1999. a large-scale model of wolf distribution in italy for conservation planning. conservation biology 13:150-159. dormann, c., j. mcpherson, m. araújo, r. bivand, j. bolliger, g. carl, r. davies, a. hirzel, w. jetz, and w. kissling. 2007. methods to account for spatial autocorrelation in the analysis of species distributional data: a review. ecography 30:609-628. duckworth, w. d., h. h. genoways, and c. l. rose. 1993. preserving natural science collections: chronicle of our environmental heritage. national institute for the conservation of cultural property, washington, dc. elith, j., c. graham, r. anderson, m. dudík, s. ferrier, a. guisan, r. hijmans, f. huettmann, j. leathwick, and a. lehmann. 2006. novel methods improve prediction of species' distributions from occurrence data. ecography 29:129. engler, r., a. guisan, and l. rechsteiner. 2004. an improved approach for predicting the distribution of rare and endangered species from occurrence and pseudo-absence data. ecology 41:263-274. graham, c., j. elith, r. hijmans, a. guisan, a. t. peterson, and b. loiselle. 2008. the influence of spatial errors in species occurrence data used in distribution models. journal of applied ecology 45:239-247. graham, c., s. ron, j. santos, c. schneider, and c. moritz. 2004. integrating phylogenetics and environmental niche models to explore speciation mechanisms in dendrobatid frogs. evolution 58:1781-1793. grinnell, j. 1917. field tests of theories concerning distributional control. american naturalist 51:115. guisan, a., c. graham, j. elith, and f. huettmann. 2007. sensitivity of predictive species distribution models to change in grain size. diversity and distributions 13:332-340. guisan, a., and n. zimmermann. 2000. predictive habitat distribution models in ecology. ecological modelling 135:147-186. guo, q., y. liu, and j. wieczorek. 2008. georeferencing locality descriptions and computing associated uncertainty using a probabilistic approach. international journal of geographical information science 22:1067-1090. guralnick, r., a. hill, and m. lane. 2007. towards a collaborative, global infrastructure for biodiversity assessment. ecology letters 10:663-672. hagen-zanker, a., b. straatman, and i. uljee. 2005. further developments of a fuzzy set map comparison approach. international journal of geographical information science 19:769-785. hagen, a. 2003. multi-method assessment of map similarity. international journal of geographical information science 17:235-249. hernandez, p., c. graham, l. master, and d. albert. 2006. the effect of sample size and species characteristics on performance of different species distribution modeling methods. ecography 29:773785. higgins, s., d. richardson, and r. cowling. 2000. using a dynamic landscape model for planning the management of alien plant invasions. ecological applications 10:1833-1848. hijmans, r., s. cameron, j. parra, p. jones, and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25:1965-1978. hijmans, r., l. guarino, m. cruz, and e. rojas. 2001. computer tools for spatial analysis of plant genetic resources data: 1. diva-gis. plant genetic resources newsletter 127:15-19. hirzel, a., j. hausser, d. chessel, and n. perrin. 2002. ecological-niche factor analysis: how to compute habitat-suitability maps without absence data? ecology 83:2027-2036. hugall, a., c. moritz, a. moussalli, and j. stanisic. 2002. reconciling paleodistribution models and comparative phylogeography in the wet tropics rainforest land snail gnarosophia bellendenkerensis (brazier 1875). proceedings of the national academy of sciences usa 99:6112-6117. kadmon, r., o. farber, and a. danin. 2003. a systematic analysis of factors affecting the performance of climatic envelope models. ecological applications 13:853-867. kelly, m., q. guo, d. liu, and d. shaari. 2007. modeling the risk for a new invasive forest disease in the united states: an evaluation of five environmental niche models. computers, environment and urban systems 31:689-710. leathwick, j. r., and d. whitehead. 2001. soil and atmospheric water deficits and the distribution of new zealand's indigenous tree species. functional ecology 15:233-242. lobo, j., a. jimenez-valverde, and r. real. 2008. auc: a misleading measure of the performance of predictive distribution models. global ecology & biogeography 17:145-151. manel, s., j. dias, and s. ormerod. 1999. comparing discriminant analysis, neural networks and logistic regression for predicting species distributions: a case study with a himalayan river bird. ecological modelling 120:337-347. manel, s., h. williams, and s. ormerod. 2001. evaluating presence-absence models in ecology: the fernandez et al. – uncertainty in locality data and ecological niche modeling 51 need to account for prevalence. journal of applied ecology:921-931. mills, j., and j. childs. 1998. ecologic studies of rodent reservoirs: their relevance for human health. emerging infectious diseases 4:529-537. mitchell, t., and p. jones. 2005. an improved method of constructing a database of monthly climate observations and associated high-resolution grids. international journal of climatology 25:693–712. murphy, j., d. sexton, d. barnett, g. jones, m. webb, m. collins, and d. stainforth. 2004. quantification of modelling uncertainties in a large ensemble of climate change simulations. nature 430:768-772. ortega-huerta, m., and a. peterson. 2008. modeling ecological niches and predicting geographic distributions: a test of six presence-only methods. revista mexicana de biodiversidad 79:205-216. parra, j., c. graham, and j. freile. 2004. evaluating alternative data sets for ecological niche models of birds in the andes. ecography 27:350-360. pearce, j., and s. ferrier. 2000. evaluating the predictive performance of habitat models developed using logistic regression. ecological modelling 133:225245. pearson, r., c. raxworthy, m. nakamura, and a. peterson. 2007. predicting species distributions from small numbers of occurrence records: a test case using cryptic geckos in madagascar. journal of biogeography 34:102-117. pearson, r., w. thuiller, m. araújo, e. martinez-meyer, l. brotons, c. mcclean, l. miles, p. segurado, t. dawson, and d. lees. 2006. model-based uncertainty in species range prediction. journal of biogeography 33:1704-1711. peterson, a., and k. cohoon. 1999. sensitivity of distributional prediction algorithms to geographic data completeness. ecological modelling 117:159164. peterson, a., m. papes, and j. soberón. 2008. rethinking receiver operating characteristic analysis applications in ecological niche modeling. ecological modelling 213:63-72. peterson, a., v. sanchez-cordero, e. martínez-meyer, and a. navarro-sigüenza. 2006. tracking population extirpations via melding ecological niche modeling with land-cover information. ecological modelling 195:229-236. peterson, a., and j. shaw. 2003. lutzomyia vectors for cutaneous leishmaniasis in southern brazil: ecological niche models, predicted geographic distributions, and climate change effects. international journal of parasitology 33:919–931. peterson, a., and d. vieglais. 2001. predicting species invasions using ecological niche modeling: new approaches from bioinformatics attack a pressing problem. bioscience 51:363-371. peterson, t., m. papeş, and m. eaton. 2007. transferability and model evaluation in ecological niche modeling: a comparison of garp and maxent. ecography 30:550-560. phillips, s. 2008. transferability, sample selection bias and background data in presence-only modelling: a response to peterson et al.(2007). ecography 31:272-278. phillips, s., m. dudík, and r. schapire. 2004. a maximum entropy approach to species distribution modeling. pp. 83-84 in a. i. c. p. series, ed. proceedings of the twenty-first international conference on machine learning. acm press new york, ny, usa, banff, alberta, canada phillips, s. j., r. p. anderson, and r. e. schapire. 2006. maximum entropy modeling of species geographic distributions. ecological modelling 190: 231-259. pontius, r. 2000. quantification error versus location error in comparison of categorical maps. photogrammetric engineering and remote sensing 66:1011-1016. proctor, e. 2004. reducing variation in georeferenced locality descriptions. pp. 191. geography & human environmental studies. san francisco state university, san francisco. raxworthy, c., e. martinez-meyer, n. horning, r. nussbaum, g. schneider, m. ortega-huerta, and a. townsend peterson. 2003. predicting distributions of known and unknown reptile species in madagascar. nature 426:837-841. reddy, s., and l. davalos. 2003. geographical sampling bias and its implications for conservation priorities in africa. journal of biogeography 30:1719-1727. reese, g., k. wilson, j. hoeting, and c. flather. 2005. factors affecting species distribution predictions: a simulation modeling experiment. ecological applications 15:554-564. roura-pascual, n., a. suarez, c. gómez, p. pons, y. touyama, a. wild, and a. peterson. 2004. geographical potential of argentine ants (linepithema humile mayr) in the face of global climate change. proceedings of the royal society b: 271:2527-2535. rowe, r. 2005. elevational gradient analyses and the use of historical museum specimens: a cautionary tale. journal of biogeography 32:1883-1897. soberon, j., and t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. philosophical transactions of the royal society b: biological sciences 359:689-698. stockwell, d. 2006. niche modeling: predictions from statistical distributions. chapman & hall/crc, boca raton fl. stockwell, d., and d. peters. 1999. the garp modelling system: problems and solutions to fernandez et al. – uncertainty in locality data and ecological niche modeling 52 automated spatial prediction. international journal of geographical information science 13:143-158. stockwell, d., and a. peterson. 2002. effects of sample size on accuracy of species distribution models. ecological modelling 148:1-13. tobler, w. 1970. a computer movie simulating urban growth in the detroit region. economic geography:234-240. tsoar, a., o. allouche, o. steinitz, d. rotem, and r. kadmon. 2007. a comparative evaluation of presence-only methods for modelling species distribution. diversity & distributions 13:397-405. underwood, e., r. klinger, and p. moore. 2004. biodiversity research predicting patterns of nonnative plant invasions in yosemite national park, california, usa. diversity & distributions 10:447. visser, h., and t. de nijs. 2006. the map comparison kit. environmental modelling and software 21:346358. welk, e., k. schubert, and m. hoffmann. 2002. present and potential distribution of invasive garlic mustard (alliaria petiolata) in north america. diversity and distributions 8:219-233. wieczorek, j., q. guo, and r. hijmans. 2004. the pointradius method for georeferencing locality descriptions and calculating associated uncertainty. international journal of geographical information science 18:745-767. williams, s., and j. hero. 2001. multiple determinants of australian tropical frog biodiversity. biological conservation 98:1-10. wisz, m., r. hijmans, j. li, a. peterson, c. graham, and a. guisan. 2008. effects of sample size on the performance of species distribution models. diversity and distributions 14:763-773. zimmermann, n., t. edwards, g. moisen, t. frescino, and j. blackard. 2007. remote sensing-based predictors improve distribution models of rare, early successional and broadleaf tree species in utah. journal of applied ecology 44:1057-1067. microsoft word lobo_final_20080201_new.doc biodiversity informatics, 5, 2008, pp. 14-19 14 more complex distribution models or more representative data? jorge m. lobo dpto. de biodiversidad y biología evolutiva, museo nacional de ciencias naturales (csic), c/josé gutiérrez abascal, 2 – 28006, madrid, spain. e-mail: mcnj117@mncn.csic.es abstract.⎯ distribution models for species are increasingly used to summarize species’ geography in conservation analyses. these models use increasingly sophisticated modeling techniques, but often lack detailed examination of the quality of the biological occurrence data on which they are based. i analyze the results of the best comparative study of the performance of different modeling techniques, which used pseudo-absence data selected at random. i provide an example of variation in model accuracy depending on the type of absence information used, showing that good model predictions depend most critically on better biological data. key words.⎯ distribution models, model reliability, pseudo-absences, conservation usefulness. recently, many efforts have focused on creation of models able to predict species’ distributions from partial data. these distributional models use known distribution records of a species, as well as environmental and spatial explanatory variables, to build statistical functions for interpolating species’ distributions across the environmental spectrum (guisan & zimmermann 2000). models may also extrapolate species’ distributions to sets of environmental conditions outside those used to build the models (peterson 2003). the reliability of these predictions depends on many factors, but the three main ones are (1) quality of data used for model calibration, (2) predictive power of the explanatory variables, and (3) the modeling technique chosen to produce predictions from such variables. in spite of recent theoretical opinions on the errors that may result from the first two sources (soberón & peterson 2005, barry & elith 2006, araújo & guisan 2006), little effort has been devoted to testing experimentally the effects of these factors on the reliability of model outputs. instead, much effort has been devoted to comparisons of the different available modeling techniques (brotons et al. 2004, segurado & araújo 2004, pearson et al. 2006, and references therein). some recent papers (drake et al. 2006, elith et al. 2006) suggest that certain newly developed modeling techniques are better able to parametrize complex relationships, producing better distributional hypotheses for conservation purposes. the study by elith and collaborators (2006) drew especially interesting conclusions. this research paper is undoubtedly the best comparative study of the relative performance of different modeling techniques. comparing the reliability of 16 techniques, and modeling 226 species from six world regions, the researchers validated the predicted distributions with “independent” and reliable species presence/absence data that were withheld from model building. as distribution models were derived from both presence and presence-pseudo absence data, with absences randomly distributed throughout the territory considered, results illustrate the potential of existing techniques applied to widely available information. this short paper is designed to illustrate the shortcomings of the model comparison approach in improving the results of predictive models of species’ distributions: improvements in the biological occurrence data may provide more important advances than a more complex modeling approach. enough accuracy for conservation? unfortunately, in the study by elith and collaborators (elith et al. 2006), maximum mean scores of the area under the receiver operating characteristic curve (auc), a measure of predictive accuracy, do not surpass 0.82 (mean auc score around 0.70 for most species and types of models). average auc score for the modeling technique with the best predictions for all regions lobo – more complex models or representative data? 15 was 0.73. the average auc score in the study of drake and collaborators is similar (0.79). let us suppose that we obtain an auc score of 0.82 for a model accomplished at a 100 x 100 m resolution in switzerland (41.290 km2 or 4.129.000 pixels), using for that 6000 presence points (similar conditions to those of the best model in the elith et al. study). for an auc score like this it is exceptional to obtain an outstanding specificity score of 0.99 (99% of absences correctly predicted), but even in this case 41,230 pixels were erroneously ascribed as presences (4.123.000 absence pixels x 0.01); a predicted area almost 7fold larger (413 km2) than the observed one (60 km2). hence, the usefulness for conservation of even the best models identified by these studies is questionable. should we prioritize the use of such techniques, or search for others that are still more sophisticated? an example although biologists may know the places in which a species is unlikely to be observed (e.g., species not detected at a locality after intense sampling), such data are not usually published. thus, despite the potential usefulness of relatively reliable absence data, such information is generally not available. random selection of absences, a crude approach, may introduce an indeterminate number of false absences into models owing to the all-too-frequent sampling biases in biological information (see dennis et al. 1999, dennis & thomas 2000, zaniewski et al. 2002, reutter et al. 2003, graham et al. 2004, martínez-meyer 2005). the influence of random selection of absences on distribution models was illustrated for a large iberian dung beetle species (copris hispanus) that occurs mainly in the southern half of the iberian peninsula. using an exhaustive compilation of all available occurrence information regarding iberian dung beetles (54 species, 15,924 database records), i first used accumulation curves to select reliably inventoried 50 x 50 km utm cells. for each cell, i examined the number of species accumulated with the increase in the number of database records (an effective surrogate of the sampling effort carried out in each cell, see hortal et al. 2006). each curve was estimated 100 times, randomizing the entry order of the database records to smooth the curve, and subsequently fitted to the clench function (colwell & coddington 1994, soberón & llorente 1993) to estimate the asymptotic value (i.e. the estimated total richness score for an unlimited number of samples). the adequately-inventoried utm cells were defined as those with observed species richness of >80% of the asymptotic predicted scores. all of the 100 km2 utm cells belonging to the 2500 km2 well-surveyed cells at which c. hispanus had not been detected were considered as true absences. forty-seven presence points and an equal number of absences were selected: (1) at random from all cells lacking presence or (2) from the cells considered to be true absences, and were modeled via a widely-accepted prediction technique (gams). the model used 9 climatic and lithological variables as predictors (total annual precipitation, rainfall during summer months, yearly mean temperature, minimum annual temperature, an aridity index, area with stony siliceous soils, calcareous soils, siliceous sediments and calcareous sediments). all models were repeated 10 times, and predictions were validated using information from 205 cells (158 presences and 47 sound absences) not used in model calibration. beside the auc scores for the validation data, percentages of presences and absences correctly predicted (sensitivity and specificity scores) were also calculated after applying to model probabilities the threshold that minimizes the difference between sensitivity and specificity (see jiménez-valverde & lobo 2006 and 2007). model predictions were obtained for both the entire iberian peninsula and only the southern half of the peninsula to illustrate the effects of the geographic extent at which these types of models are developed on output probabilities and validation scores. predicting with sound absence data inclusion of reliable absence data significantly improved model predictions, especially for smaller territories with a less variable environment (fig. 1). as anticipated, average auc scores for the entire iberian peninsula (± 95% confidence intervals) using random absences are significantly lower (0.951 ± 0.011) than in the case of absences selected among well-surveyed cells (0.977 ± 0.007; f(2, 27) = 42.14, p< 0.0001). interestingly, this difference is still greater when only the southern lobo – more complex models or representative data? 16 random well surveyed 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 a u c random well surveyed 0.75 0.80 0.85 0.90 0.95 1.00 a u c figure 1. mean auc scores (± 95% confidence intervals) with respect to method of selection of absence information (at random or from among well-surveyed cells). left: model based on southern half of the iberian peninsula; right: model based on the entire iberian peninsula. half of the iberian territory is considered: auc scores derived from random absences are even lower and more variable (0.723 ± 0.090) than with well-surveyed cells (0.907 ± 0.027; f(2, 27) = 10.87, p= 0.0003). thus, the inclusion of reliable absence data significantly improves model predictions, especially for smaller ranges with less variable environments (see fig. 2). put another way, absences randomly distributed in a larger area lead to better predictions through reduced possibility of including false absences. the study by elith and collaborators (elith et al. 2006) is, without a doubt, the most comprehensive of all to date. unlike preceding studies, the authors validated model predictions with independent data and good absence information. what would have been the result if the reliable absence information used to validate their models had been used for calibration? interestingly, when good data and predictors were used by elith and collaborators (elith et al. 2006; see, e.g., the case of switzerland), accuracy differences among modeling methods seemed to diminish. unfortunately, most such modeling exercises do not use reliable absence information either to calibrate models or to validate them. this undesirable practice highly compromises the conservation usefulness of distribution model results. if one wants to generate a distributional simulation able to reflect the realized distribution of the species, good absence data need to be incorporated. these data should be located in climatically suitable localities in which the species does not occur due to historical factors, biotic interactions or dispersal limitation processes (pulliam 1988, ricklefts & schluter 1993, hanski 1998, pulliam 2000). including absences from a priori favorable environmental localities will inevitably diminish the predicted range size, so that the modeled distribution approaches the realized one (see chefaoui & lobo 2008). on the other hand, including absences from environmentally unsuitable places generates simulations which approach the potential distribution (all the environmentally suitable places in which a species could occur according to a group of environmental variables; see soberón & peterson 2005, peterson 2006). the probability of including false absences when absences are selected at random increases at lobo – more complex models or representative data? 17 figure 2. left: gam-derived distributional prediction for the iberian dung beetle species copris hispanus based on absence information derived from adequately-inventoried 50x50 km utm cells (yellow squares). white dots represent known presence localities (100 km2 utm cells). the three figures at right are distributional predictions based on absences randomly selected from the whole iberian peninsula (a); analyzing only the southern half of the iberian peninsula based on reliable (b) and randomly selected (c) absence information. the different shades represent probabilities from 0 (white) to 1 (dark red) that are averages of 10 replicate model predictions. note that use of randomly selected absences generates higher probability scores in regions of absence. smaller extents; at larger extents, it is more likely that random absence data are environmentally distant from the presence domain. thus, the drawback of selecting random absences is higher when the ratio between the extent of species occurrence and the extent of the entire studied territory increases (the relative occurrence area). this point is exemplified by the model based on the southern half of the iberian peninsula (fig. 2). for the same species, a model built at a smaller extent will produce inferior results if the absence data used are not reliable. as species modeled across the same region will frequently differ in relative occurrence area, the accuracy of models results cannot be compared among species (lobo et al. 2008), particularly when random selection of absences implies the choice of a high number of false absences. while recognizing the relevance of the search for improved modeling techniques, researchers must not forget that model prediction quality depends on data quality, and that species’ absences input into such models should be as reliable as species’ presences. among other steps, collaboration between modelers and taxonomists in designing data selection, and databases compiling all available information can allow assessment of inventory completeness; both of these points offer strategies towards better distribution hypotheses for conservation purposes. acknowledgments this paper was supported by a fundación bbva project (diseño de una red de reservas para la protección de la biodiversidad en américa austral), and two mec projects (cgl2004-04309 and cgl2006-09567/bos). special thanks to alberto jiménez valverde for his valuable comments references araújo, m. b. and a. guisan. 2006. five (or so) challenges for species distribution modelling. journal of biogeography 33:1677-1688. a b c lobo – more complex models or representative data? 18 barry, s. and j. elith. 2006. error and uncertainty in habitat models. journal of applied ecology 43:413-423. brotons, l., w. thuiller, m. b. araújo, and a. h. hirzel. 2004. presence-absence versus presence-only modelling methods for predicting bird habitat suitability. ecography 27:437-448. chefaoui, r. m. and j. m. lobo. 2008. assessing the effects of pseudo-absences on predictive distribution model performance. ecological modelling 210: 478-486. colwell, r. k. and j. a. coddington. 1994. estimating terrestrial biodiversity through extrapolation. philosophical transactions of the royal society of london b 345:101-118. dennis, r. l. h., t. h. sparks, and p. hardy. 1999. bias in butterfly distribution maps: the effects of sampling effort. journal of insect conservation 3:33-42. dennis, r. l. h. and c. d. thomas. 2000. bias in butterfly distribution maps: the influence of hot spots and recorder's home range. journal of insect conservation 4:73-77. drake, j. m., c. randin, and a. guisan. 2006. modelling ecological niches with support vector machines. journal of applied ecology 43:424-432. elith, j., c. h. graham, r. p. anderson, m. dudík, s. ferrier, a. guisan, r. j. hijmans, f. huettmann, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. m. overton, a. t. peterson, s. j. phillips, k. richardson, r. scachetti-pereira, r. e. schapire, j. soberón, s. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species’ distributions from occurrence data. ecography 29:129-151. engler, r., a. guisan, and l. rechsteiner. 2004. an improved approach for predicting the distribution of rare and endangered species from occurrence and pseudo-absence data. journal of applied ecology 41:263-274. graham, c. h., s. ferrier, f. huettman, c. moritz, and a. t. peterson. 2004. new developments in museum-based informatics and applications in biodiversity analysis. trends in ecology and evolution 19:497-503. guisan, a. and n. e. zimmermann. 2000. predictive habitat distribution models in ecology. ecological modelling 135:147-186. hanski, i., 1998. metapopulation dynamics. nature 396:41-49. hortal, j., borges, p. a., and c. gaspar. 2006. evaluating the performance of species richness estimators: sensitivity to sample grain size. journal of animal ecology 75:274-287. jiménez-valverde, a. and j. m. lobo. 2006. the ghost of unbalanced species distribution data in geographic model predictions. diversity and distributions 12:521-524. jiménez-valverde, a. and j. m. lobo. 2007. threshold criteria for conversion of probability of species presence to either-or presenceabsence. acta oecologica 31:361-369. lobo, j. m., a. jiménez-valverde, and r. real. 2008. auc: a misleading measure of the performance of predictive distribution models. global ecology and biogeography 17: 145151. martínez-meyer, e. 2005. climate change and biodiversity: some considerations in forecasting shifts in species potential distributions. biodiversity informatics 2:42-55. pearson, r. g., w. thuiller, m. b. araújo, e. martinez-meyer, l. brotons, c. mcclean, l. miles, p. segurado, t. p. dawson, and d. c. lees. 2006. model-based uncertainty in species range prediction. journal of biogeography 33:1704-1711. peterson, a. t. 2003. predicting the geography of species’ invasions vias ecological niche modelling. quarterly review of biology 78:419-433. peterson, a. t. 2006. uses and requirements of ecological niche models and related distributional models. biodiversity informatics 3:59-72. pulliam, h. r. 1988. sources, sinks and population regulation. american naturalist 132:652-661. pulliam, h. r. 2000. on the relationship between niche and distribution. ecology letters 3:349361. reutter, b. a., v. helfer, a. h. hirzel, and p. vogel. 2003. modelling habitat-suitability using museum collections: an example with three sympatric apodemus species from the alps. journal of biogeography 30:581–590. lobo – more complex models or representative data? 19 ricklefs, r. e. and d. schluter. 1993. species diversity in ecological communities. historical and geographical perspectives. university chicago press, chicago. segurado, p. and m. b.araújo. 2004. an evaluation of methods for modelling species distributions. journal of biogeography 31:1555-1568. soberón, j. and j. llorente. 1993. the use of species accumulation functions for the prediction of species richness. conservation biology 7:480-488. soberón, j. and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodiversity informatics 2:1-10. zaniewski, a. e., a. lehmann, and j. m. overton. 2002. predicting species spatial distributions using presence-only data: a case study of native new zealand ferns. ecological modelling 157:261-280. microsoft word teefmolv4.doc biodiversity informatics, 5, 2008, pp. 1-13 1 where and how to manage: optimal selection of conservation actions for multiple species astrid j.a.van teeffelen 1,2 & atte moilanen1 1 metapopulation research group, department of biological and environmental sciences, university of helsinki, p.o. box 65 (viikinkaari 1), fin-00014 helsinki, finland; 2 land use planning group, department of environmental sciences, wageningen university, p.o. box 47, nl-6700 aa wageningen, the netherlands. email: astrid.vanteeffelen@wur.nl abstract.— multiple alternative options are frequently available for the protection, maintenance or restoration of conservation areas. the choice of a particular management action can have large effects on the species occurring in the area, because different actions have different effects on different species. together with the fact that conservation funds are limited and particular management actions are costly, it would be desirable to be able to identify where, and what kind of management should be applied to maximize conservation benefits. currently available site-selection algorithms can identify the optimal set of sites for a reserve network. however, these algorithms have not been designed to answer what kind of action would be most beneficial at these sites when multiple alternative actions are available. we describe an algorithm capable of solving multi-species planning problems with multiple management options per site. the algorithm is based on benefit functions, which translate the effect of a management action on species representation levels into a value, in order to identify the most beneficial option. we test the performance of this algorithm with simulated data for different types of benefit functions and show that the algorithm’s solutions are optimal, or very near globally optimal, partially depending on the type of benefit function used. the good performance of the proposed algorithm suggests that it could be profitably used for large multi-action multi-species conservation planning problems. key words. — benefit function, conservation planning, habitat restoration, optimisation, reserve selection, site selection algorithm. biodiversity conservation is not just a matter of reservation of natural habitats. conservation encompasses a range of protection and management options that could be used to achieve conservation goals. numerous examples exist where conservation planning involves decision making between these alternative options: when defining reserves, each reserve area can be assigned to one of several protection levels, such as strict nature reserve, national park or protected landscape (iucn 1994). once protected, sites oftentimes require management to maintain their quality, which can change for example through succession. again, a number of alternative management options are available to maintain site quality. for example, to conserve semi-natural meadows overgrowth can be prevented by grazing with various herbivore species, or by mowing management, each of which can take place at various intensity levels (see for examples e.g. köhler et al. 2005; mitchley and xofis 2005; woodcock et al. 2005). also in the field of habitat restoration decisions are required between multiple alternative management options: the planning of agri-environment schemes and biodiversity offsets—i.e. areas where habitat is restored in order to compensate for habitat loss elsewhere due to development activities—are examples where multiple management actions are available to improve habitat quality of degraded sites (cuperus et al. 1999; manchester et al. 1999; ten kate et al. 2004; donald and evans 2006). each of the management options available for conservation can have different effects on different species. for example, in the case of the grazing of a meadow, extensive grazing could be most beneficial to particular species, whereas other species benefit more from an intensive grazing regime (pykälä 2003; pöyry et al. 2004; köhler et al. 2005). since each site can only be managed in one particular way, careful planning of site management is important for reaching conservation aims. in addition, conservation budgets are typically limited and particular management actions can be costly, and hence only a fraction of sites can be protected or managed. planning of conservation management van teeffelen and moilanen — selecting optimal conservation actions 2 actions therefore requires decisions concerning a) which sites will be chosen for conservation action and b) which actions will be applied to each of these sites. the field of reserve design has a tradition in the use of optimization tools to select optimal reserve networks for a large number of species in a cost-effective manner (margules and pressey 2000; cabeza and moilanen 2001; williams et al. 2004; sarkar et al. 2006). these optimization methods account for a single value per species per site, which typically is information either about the density, probability of occurrence or presence-absence of the species at the site. the optimization methods, which are often called site-selection or reserve selection algorithms, select a set of sites based on species occurrences at sites and complementarity in order to achieve a conservation objective, such as maximizing species representation for a given budget. we point out however, that in many real-world conservation planning problems sites could potentially be managed in several different ways and that each of these management options can have a different effect on each species. this results in a different value for each species in each site under each management option. as a consequence, the site selection problem extends from selecting the optimal set of sites for a particular option, to selecting the optimal set of sites and the optimal management option for each site. one could think of this as an optimal conservation action allocation problem. currently available site selection algorithms were not designed to deal with multiple alternative options per site. compared to the protect-or-not scenario, a multi-action site selection problem is computationally more complex for two reasons: first, the number of different solutions and consequently the search space size increases enormously. second, complex trade-offs between species can occur, as the best action for one species could be suboptimal for another species (van teeffelen 2007). the data demands for a multi-action reserve selection problem are large as well, and these factors together may be the reason why little has been published concerning optimal multi-action conservation planning (but see hof et al. (1994) and bevers et al. (1995) for timing of forest harvest and holzkämper et al (2006) for optimal land use planning). in this paper we first illustrate with a small example the characteristics of a planning problem with multiple options per site and multiple species when using benefit functions to value species representation. next, we present a formulation for the multi-species multi-action planning problem, in which the effect of an action can differ per species and per location, and costs can vary between actions and locations. the aim of optimization is to maximize the value of the chosen set of sites and actions, given that there is a cost constraint and that only one action can be chosen per site. we tested the algorithm with simulated data to demonstrate the optimality characteristics of the presented problem formulation and optimization algorithm. in addition to testing for the effects of different benefit functions on algorithm performance, we discuss the consequences of choosing a particular benefit function type for conservation, and give guidelines how our method can be used in realworld planning problems. methods consider a situation in which multiple management actions could be applied to each site in a landscape, but only a single action can be executed per site. due to budget limitations only a limited number of sites can be managed. in this case decision is required on which action to apply to which site, such that the benefit for all species is maximized for the given budget. to quantify the effect that each of the restoration actions would have on the species, we use the concept of a benefit function (hof and raphael 1993; bevers et al. 1995; arponen et al. 2005, 2007; cabeza and moilanen 2006). a benefit function fj[rj(x)] is an increasing function of the representation of species j rj(x), which specifies how the conservation value of a network (set of sites x) changes when the species’ representation level changes (figure 1; see table 1 for a list of symbols used). species representation can be understood as the expected abundance, number of occurrences, individuals or populations of a species in the set of selected sites, following chosen conservation actions. such estimates can for example be obtained from predictive species distribution models (guisan and thuiller 2005; elith et al. 2006). the actual value of species j in the network, vj(x), is determined by the benefit function fj for the species and a species-specific weight wj. this weight can be used to set different conservation priorities for different species; for example, one can set higher weights for endangered species (arponen et al. 2005; holzkämper et al. 2006; moilanen 2007; van teeffelen et al. in review). the following van teeffelen and moilanen — selecting optimal conservation actions 3 equation gives the value of one species in the selected set of sites: ( ) )]([ xx jjj rfwv = (1) the total value of the network is then simply a sum over species-specific values: ( ) ( )∑= j jvf xx (2) in contrast to target based planning, benefit functions account for how much the current level of species representation in the network is below tj species representation (rj (x)) v al ue (v j ( x ) ) ramp linear sigmoid concave wj – 0 – 0 figure 1. graphical representation of the type of benefit functions used. the benefit functions translate the representation of species j in the reserve network into a value, scaled by the weight given to the species (wj). at a chosen target level of representation (tj) the value will be equal to the species’ weight wj. or above a given nominal target level of representation. the assumption is that increasing species representation is always better for species of conservation concern. we next illustrate with a toy example how selection among multiple actions per site for multiple species can be handled with benefit functions. assume a planning problem concerning a single site. this site could for example be a semi-natural meadow that requires management to avoid overgrowth. we assume two alternative management options to be available for this site: option 1 concerns annual grazing of the site and option 2 concerns biennial grazing (every second year) of the site, both with cattle. each of these options has an associated cost, which we assume to be 1.5 and 1 for options 1 and 2, respectively (table 2). assume that two species of conservation concern occur at the site. species a has a population size of a individuals, whereas the population size of species b is b individuals, and we assume a < b. we assume that option 1 (annual grazing) is beneficial to species a, increasing its population size by x individuals, while it does not influence species b. likewise, we assume that option 2 (biennial grazing) is beneficial to species b, increasing its population size by x individuals, while it does not affect species a. we used single-species effects for each option for the sake of simplicity; note however that each option could influence multiple species simultaneously, positively as well as negatively. since only one of these management options can be applied to the site, a decision is required about which option is preferred. the representation of each species is translated into a value by the benefit function, for the current representation levels, as well as for the expected representation levels under each of the conservation options. we demonstrate two different benefit functions: concave and sigmoid (figure 2). furthermore, we assume species to have equal conservation priority; hence we set weights for both species to 1. although the species have an equal absolute increase in representation (x) under their respective beneficial actions, the change in value that each of the options generates depends on the benefit function used. the concave benefit function returns a larger increase in value for option 1 than for option 2 (f1 > f2), and although this difference is partly counterbalanced by the higher cost of option 1, the difference in f relative to option cost (marginal gain, or efficiency) is larger for option 1 than for option 2 (δ1 > δ2, table 2). option 1 is therefore preferred using the concave benefit function type. the sigmoid function returns a larger increase in value for option 2 than for option 1 (f1 < f2), which is emphasised by the lower cost of option 2 (δ1 < δ2), and hence option 2 is preferred when using the sigmoid benefit function type. it follows that the optimal decision depends on which benefit function is used to value species representation. note that this example corresponds to the cost-effective addition of one site to an existing set of sites. the problem extends when the number of sites, options per sites and species increases. in van teeffelen and moilanen — selecting optimal conservation actions 4 addition, effects of actions on species as well as action cost can vary between sites, and species can be prioritized by incorporating different weights between them. consequently, solving this kind of selection problem is no longer straightforward. we next describe an algorithm that is applicable to selection problems with multiple actions per site, and multiple species. the algorithm requires two main inputs: (i) a three-dimensional table giving the (estimated) representation of species j at site i given action k and (ii) a table giving cost of action k at site i. our problem formulation, unlike target-based planning formulations, is mathematically table 1. symbols used for multi-action conservation planning. symbol explanation a set of available sites i index for sites, i = 1, 2, ... ns j index for species, j = 1, 2, ... np k index for different conservation actions, k = 1, 2, ... na rij(k) representation of species j in site i under action k xi indicator variable, xi = 1 if site i has been selected for action, else xi = 0 ai action selected for site i x index set of sites that have been selected for a conservation action, i.e. the network rj(x) = ∑i xirij(ai), representation of species j in network x given actions ai fj[rj(x)] benefit function: an increasing function of species representation rj(x), which expresses how the conservation value of x, f(x), changes when the representation level of species j changes wj the weight of species j, to indicate a species’ conservation priority relative to other species under consideration vj(x) = wj fj[rj(x)], value of representation of species j f(x) = ∑j vj(x), value of network x cik cost of action k at site i cmax the total budget available cused the part of the budget allocated δik = [f(x + i) f(x)] / cik, the marginal gain: the rate of increase of network value, relative to the cost of the site-action pair, when site i is added to the network assuming action k at site i table 2. parameters and results for an example problem of multi-action selection. see figure 2 for a graphical display of species representation and value. current option 1 option 2 representation species a a a + x a representation species b b b b + x option cost ccurrent = 0 c1 = 1.5 c2 = 1 value fcurrent = va + vb f1 = va+x + vb f2 = va + vb+x marginal gain -δ1 = (f1 fcurrent) / c1 δ2 = (f2 fcurrent) / c2 concave function, value marginal gain fcurrent = 1.50 f1 = 1.71 δ1 = 0.14 f2 = 1.62 δ2 = 0.12 sigmoid function, value marginal gain fcurrent = 1.03 f1 = 1.22 δ1 = 0.13 f2 = 1.35 δ2 = 0.32 van teeffelen and moilanen — selecting optimal conservation actions 5 concave, which suggests that a stepwise search strategy could be used to solve it in a globally optimal manner. essentially, the proposed algorithm is a forward stepwise heuristic. it iteratively adds a site to the set of selected sites (x), in combination with a particular management action for that site. the site-action pair that returns the highest marginal increase in conservation value f to the set of selected sites, while accounting for action cost, is chosen to be added. after a site-action pair is added to the set of selected sites, no further actions are allowed for that site: the site is removed from the list of available sites. see box 1 for a description of the algorithm, and table 1 for explanation of symbols used. stepwise heuristic algorithms are not in general guaranteed to find globally optimal solutions. the present formulation has however mathematical properties that suggest that an iterative heuristic might do very well in optimizing it. these properties include linear additivity over values for species, and a potentially linear or concave form for the benefit function for a species. we tested for the optimality of the algorithm’s solutions by using simulated data sets, for which we could find the optimal solution by full enumerative search over all possible solutions. full enumerative search is only possible for small problems due to computational limitations, whereas real-world problems may be large with many sites, species and options per site. therefore heuristic algorithms are developed and used to provide (near-) optimal solutions for large problems. in order to verify whether certain problem characteristics affect the optimality, we defined a range of problem classes, each specified by a given number of sites, species, actions per site and available resource units (table 3). the value of the solution found by the algorithm, falg(x), was compared to the value of the optimal solution fopt(x) that was found by full enumerative search over all possible solutions. the average value of all possible solutions, fmean(x), corresponds to the mean expected value when conservation actions are randomly assigned to locations, under the budget constraint. to check to what extent the algorithm’s solutions are better than random allocation of conservation actions, we compared the difference between fopt(x) and falg(x) to the difference between fopt(x) and fmean(x): )()( )( xx x(x) meanopt algopt alg ff ff y − − = (3) where yalg is the relative performance of the algorithm. yalg < 1 indicates that the algorithm performs better than random, and yalg = 0 indicates that algorithm’s solution is equal to the globally optimal solution. the performance of the algorithm might also depend on the type of benefit function, for which we tested four different types: a linear, ramp, concave and sigmoid function (figure 1). each problem class was tested in combination with each benefit function in turn and 1000 replicates were created for each class-function pair by randomizing: species representation v al ue b a va a+x va+x vb vb+x b+x a va a+x b b+x species representation va lu e va+x vb vb+x concave sigmoid figure 2. translation of representation into value by a concave benefit function (left) and a sigmoid benefit function (right). a = current representation of species a, b = current representation of species b, x = increase in representation due to management action. v represents a value corresponding to each representation level, according to the benefit function used. see table 2 for an example. van teeffelen and moilanen — selecting optimal conservation actions 6 box 1. the algorithm for selecting sites and conservation actions at these sites. the pseudocode is given on the left hand side, with explanations on the right hand side. initialisation set a = {1,2, ..., ns} define the set of available sites. set x = ∅ no sites have been selected for conservation. set cused = 0 no resource has been used. set xi = ai = 0 for all i none of the sites and actions are selected. repeat site-action pair selection set max_improvement = 0 for all sites i∈a loop through all available sites and actions. for all actions k = 1, 2, ..., na if cik ≤ (cmax cused) then if the site-action pair can be afforded, then calculate the marginal gain of the site-action pair. calculate δik = (f(x+i)-f(x))/ cik, assuming ai = k if δik > max_improvement then if this is the largest gain found thus far, then set max_improvement=δik set the current improvement as maximum improvement. set best_site=i set the current site as best site. set best_action=k set the current action as best action. endif endif endfor endfor if max_improvement > 0 then best_site ∈ x include the best site in the set of selected sites. best_site ∉ a exclude the best site from the set of available sites. xbest site = 1 indicate that this site is now selected. abest site = best_action indicate which action is now selected for this site. cused = cused + cbest_site,best_action add the costs of the site-action pair to the used costs. endif until max_improvement = 0 quit when no improvement can be afforded. • rij(k), the representation of each species j in each site i under management action k. within each replicate, each site i obtained a different random base-level of representation for each species j (uniformly distributed between 0.0-1.0), on top of which each action k added an extra contribution (uniformly distributed between 0.0-1.0). • the weight of each species wj ∈ [1, 5]. for simplicity, cost levels of various management actions (cik) were kept equal at 1.0. the algorithm evaluates site-action pairs by their conservation value relative to cost (marginal gain), and since we varied conservation value already widely across sites, having equal cost levels per site-action pair should therefore not influence algorithm testing. the way the data was randomized, all of the problems had only a single unique globally optimal solution. we also tested the sensitivity of the algorithm to variation in problem characteristics. to do so, we varied the number of actions, the amount of resource available and the number of species of one problem class (d, table 3), and created 1000 replicates of each modified problem. only one factor was changed at a time, and solutions found by the algorithm were again compared to the globally optimal solution, relative to the mean of all solutions for that replicate. van teeffelen and moilanen — selecting optimal conservation actions 7 table 3. specifications of the problem classes tested. problem class specifications a b c d e # sites 6 8 10 12 14 # species 50 50 50 50 50 # actions per site 3 3 3 3 3 # site-action pairs to select 3 4 5 6 7 search space size* 540 5670 61,236 673,596 7,505,784 # randomized replicates 1000 1000 1000 1000 1000 *= ,as s n ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ where n = number of sites, s = number of selected sites and a = the number of possible actions per site. results we introduced an algorithm capable of optimizing conservation planning problems that have multiple actions per site. we applied the algorithm to simulated data and tested for the optimality of the algorithm’s solutions. with linear and ramp type benefit functions, the algorithm always found the globally optimal solution. with concave and sigmoid benefit functions the proportion of successful globally optimal replicates varied depending on the problem class (figure 3), where the concave function always returned a higher proportion of successful replicates than the sigmoid function. overall, the success rates where remarkably high: for the concave function >85% across all problem classes and for the sigmoid function >50% (figure 3a). even though not all replicates returned a value equal to the value of the globally optimal solution, the algorithm’s solutions were always close to optimal, proportional to the difference between the optimal solution and the mean value of all possible solutions. the average relative error for failed replicates across all tested problem classes was less than 0.08 from the global optimum for the sigmoid function and less than 0.04 for the concave function, compared to the mean value of all solutions (figure 4a and e). taken over all replicates, both failed and successful, these numbers are less than 0.005 and 0.02 for the concave and sigmoid benefit function respectively. such errors are likely to be much smaller than observation and prediction errors for biodiversity distribution data. with increasing search space size (problem class a → e, table 3) the number of globally optimal replicates decreased for both the sigmoid and concave function (figure 3a), but so did the average relative error (figure 4a and e). algorithm sensitivity to problem characteristics (number of species, actions and resource units) was tested by comparing variants of problem class d, varying one variable at a time. ramp and linear benefit functions again always returned the globally optimal solutions, and the following results therefore concern the concave and sigmoid benefit functions only: with an increasing number of actions, the proportion of globally optimal replicates decreased, mainly for the sigmoid function (figure 3b), with little or no effect on the relative error (figure 4b and f). the decrease in success rate is likely to be due to a major problem class a b c d e p ro po rti on g lo ba lly o pt im al 0.0 0.2 0.4 0.6 0.8 1.0 number of actions 2 3 4 5 number of resource units 4 6 8 10 number of species 25 50 100 200 a b d csigmoid function concave function figure 3. proportion of replicates where the algorithm returned an objective function value equal to that of the globally optimal solution, for a concave and a sigmoid benefit function. panel a displays success rates for problem definitions a-e (table 3). panels b-d display variants of problem d, in which problem characteristics were varied one at a time. the characteristic modified is given on the x-axis. van teeffelen and moilanen — selecting optimal conservation actions 8 increase in search space size. with an increasing resource level (more sites could be selected per replicate, because restoration cost was kept equal), the number of globally optimal replicates decreased (figure 3c), again with little effect on the relative error of replicates where the algorithm failed to find the optimum (figure 4c and g). changing the number of species hardly influenced the proportion of optimal replicates (figure 3d) and the relative error (figure 4d and h). discussion although planning problems concerning multiple alternative management actions per site are common in practice, few approaches to solve these problems in an optimal manner have been published. the algorithm we introduced is capable of solving planning problems with multiple potential actions per site for both single and multiple species in an optimal, or nearoptimal manner. potential applications of this type of decision support tool are for example found in the planning of mitigation and compensation measures (race and fonseca 1996; cuperus et al. 2001); the targeting of management within protected area networks to maintain or enhance quality, or outside protected areas, in order to complement or buffer protected area networks through reserve selection; restoration (hof et al. 2002; pieterse et al. 2002; cipollini et al. 2005) and agri-environment schemes (kleijn et al. 2001; donald and evans 2006). optimization tools such as the one presented in this paper can aid in more efficient and effective allocation of the available budgets for conservation (pressey et al. 1997; cowling et al. 2003). earlier studies on multi-action planning (hof et al. 1994; bevers et al. 1995; holzkämper et al. 2006) did not consider action cost in the optimization. cost of actions can however differ widely, and conservation budgets are typically tight. the budget available is therefore expected to influence which set of sites and actions is optimal for a given set of species and given budget (moilanen and cabeza 2002). the algorithm we used for multi-action planning is a stepwise heuristic, which have been criticized for returning sub-optimal results in conservation planning (underhill 1994; önal 2003). optimization methods that guarantee global optimality, such as integer programming or stochastic dynamic programming, are generally favourable over methods that do not guarantee global optimality. however, global s ig m oi d r el at iv e er ro r 0.00 0.02 0.04 0.06 0.08 c on ca ve problem class a b c d e r el at iv e er ro r 0.00 0.02 0.04 0.06 0.08 number of actions 2 3 4 5 number of resource units 4 6 8 10 number of species 25 50 100 200 f h a e b c d g all replicates failed replicates figure 4. average relative error (± se) for solutions of the algorithm, compared to the objective function value f(x) of the globally optimal solution and relative to the mean objective function value of all possible solutions, for a concave (panels a-d) and a sigmoid (panels e-h) benefit function. panels a and e display errors for problem classes a-e (table 3). panels b-d and f-h display variants of problem d, where problem characteristics were varied one at a time. the characteristic modified is given on the x-axis. van teeffelen and moilanen — selecting optimal conservation actions 9 optimality comes with a price: both integer programming (see williams et al. 2004) and stochastic dynamic programming (westphal et al. 2003) have limitations to the sizes of data sets that can be analysed. the problem formulation presented here, has the interesting property that it can be solved optimally or almost optimally using a stepwise search strategy due to the mathematical characteristics of the benefit function. this allows for convenient formulation and solution of conservation planning problems, and results in relatively rapid computations for large conservation planning problems, which is advantageous for the use of conservation planning tools in interactive planning (pressey et al. 1996; sarkar et al. 2004). the performance of the algorithm we propose is influenced by the type of benefit function employed to value species representations: namely whether the function is mathematically convex or concave, or not. the solutions for the ramp and linear benefit functions (which can be considered mathematically concave) were always globally optimal. also, the success rates for the concave benefit function were very high and relative errors low, while results using a sigmoid (neither solely convex nor concave) benefit function were suboptimal. this supports our assumption that a stepwise search strategy should perform optimally or almost optimally on problems that are formulated with either convex or concave functions. the mathematical properties of the concave benefit function allow the algorithm to find the global optimum with any given accuracy, if the landscape was divided into arbitrarily small land parcels. this follows because a concave or convex function can be optimised to an arbitrary precision using a gradient-ascent type optimisation method (rockafellar 1970; bazaraa et al. 1993). in our test with the concave benefit function formulation the global optimum was most often, though not always, found. sometimes the global optimum was missed because the selection was done in discrete steps (by selecting sites-action pairs with discrete cost), which may cause slightly suboptimal behaviour due to the curvature of the benefit function. if the problem has more sites, then the optimal solution will consist of a larger number of selection units each of which make up a relatively smaller contribution of the total solution. as a consequence, the optimal solution should be relatively closer to the globally optimal one. this kind of behaviour is evident in figures 3a and 4e, where the relative error of the stepwise search algorithm went down quickly as the problem sized increased. this demonstrates that the simple stepwise iterative algorithm can find practically optimal solutions for problems with many sites, when using problem formulation with a concave benefit function. the sigmoid function however, is neither convex nor concave. it therefore is relatively difficult to optimise, which is also demonstrated by suboptimal results in figures 3 and 4 (average relative errors of up to 8% for failed replicates). optimality should however not be the key factor for choosing one or the other benefit function type to value species representation, but rather, biological reasons should determine the choice. we therefore next take a closer look at different benefit functions tested and outline how they affect the selection of conservation actions, to inform decision making with respect to valuing species representation. the use of linear and ramp benefit functions to value species representation always returned globally optimal solutions for our simulated data sets. both functions value species representation linearly below a certain target level of representation. above the target level, the ramp function is horizontal and does not value overrepresentation, whereas the linear function values over-representation equal to underrepresentation. a drawback of the linear formulation is the possibility of species to fully compensate each other: increasing the representation of a better-represented species will generate the same increase in value as increasing the representation of a poorly represented species by an equal amount (see also holzkämper et al. 2006). as a result, using linear functions to value species representations does not explicitly promote the selection of sites most beneficial to the least-represented species, which is a property we do not find appropriate for conservation planning. the concave benefit function does value representation of species in a way that puts most emphasis on increasing the representation of the least-represented species, as we showed in fig 2. in other words, there are decreasing marginal gains with increasing representation levels, which is a plausible way of translating representation to conservation value. arponen et al. (2005) show that the use of a concave function benefits the representation of rare species, compared to using, for example, a step function. the concave formulation may be van teeffelen and moilanen — selecting optimal conservation actions 10 particularly suitable for plants or other species that are able to persist already in small areas (arponen et al. 2005). the sigmoid function is qualitatively very different. if there are not sufficient resources to cover all species at high representation levels, the sigmoid function has the tendency of obtaining high representation for some species while leaving others close to zero representation. due to the function’s s-shape, the total value of the network will increase more by adding a site that benefits a species with a representation level around the centre of the curve, than by selecting a site that benefits a species whose representation is at low levels close to the left part of the curve (figure 2). when the aim is to maximize the representation of as many species as possible, this effect of the benefit function’s shape may be undesirable. however, the sigmoid may be appropriate when there is reason to believe that the species can only persist at relatively high (meta)population sizes, in which case the sigmoid can be used to try to force the representation level of the species to an acceptably high level (arponen et al. 2005). the step function, which is implied by target-based planning, essentially counts the number of species that have reached a target level of representation, without valuing underand overrepresentation (arponen et al. 2005). this step function is the extreme limit of the sigmoid function, in which the increase in representation levels of under-represented species is not valued at all, which may lead to solutions where some species achieve targets whereas the representation of other species remains close to zero. how to use this method in practice? as mentioned in the introduction, the method requires estimated representations of each species at each site under each potential action, as well as costs of these actions. there are few examples where data are collected before and after a change in management, hence in most cases one needs to rely on estimates from e.g. species distribution models. these models can be used to link species occurrence to environmental variables such as climatic, soil and vegetation variables (guisan and thuiller 2005; elith et al. 2006), but can also include conservation actions as predictor variables. for example, it could be considered to increase the amount of dead wood in particular forest patches as a management action, to benefit species depending on dead wood. different actions under consideration could represent different amounts of dead wood. species distribution models with dead wood as a predictor variable could next be used to obtain estimates of the effects of the different actions at each site. van teeffelen et al. (in review) provide a real-world example of the use of this method in the context of planning grassland management for a set of species with contrasting management requirements with respect to grazing intensity. if no quantitative data is available on the effects of different actions on a species, qualitative data such as expert opinion can also be used. evidently there is uncertainty associated with such estimates, especially when management impacts are large, such as in the case of intensive restoration (suding et al. 2004; hilderbrand et al. 2005). nevertheless, estimates from well calibrated and evaluated distribution models provide a sound scientific basis to justify allocating—oftentimes costly—management actions (burgman et al. 2005). the method presented allows for a sensitivity analysis, such that the robustness of the results (selected sites and actions) to uncertainties in estimates of species responses and costs is analysed. with respect to the parameterisation of a benefit function, it is possible to scale the function between a minimum and a maximum level of representation for each species. the minimum level of representation would be obtained by managing all sites in the least beneficial way for that species, and could be set as the null-representation in figure 1. the maximum representation level is obtained by managing all sites in the most beneficial way for that species, and can be set equal to the target level of representation tj (figure 1). all other solutions in terms of sites and actions selected, will then obtain values between 0 and tj. the species-specific weight wj can be used to increase the slope of the benefit function, in order to prioritise particular species over other species. again, a sensitivity analysis can help gaining insight in the effects of setting different priorities for different species (van teeffelen et al. in review). since habitats become increasingly fragmented due to human action, the spatial configuration of conservation networks is considered important for species persistence (cabeza 2003; opdam et al. 2003; williams et al. 2005; nicholson et al. 2006; van teeffelen et al. 2006). the algorithm we presented is non-spatial, but it could account for the spatial configuration of selected sites and actions implicitly by implementing for example a van teeffelen and moilanen — selecting optimal conservation actions 11 penalty term for boundary length or for distance to existing reserve networks or habitat remnants to induce qualitative clustering (see, e.g., possingham et al. 2000; cabeza et al. 2004; crossman and bryan 2006). another way to implicitly account for connectivity is through the input data: as each action can have different effects on a species depending on the location of the action, one can adjust the estimated effect of a restoration action on a species, for example by connectivity to current populations of that species. multiple-action conservation problems where site value explicitly depends on the spatial configuration of selected sites and actions (cabeza 2003; moilanen 2005; westphal et al. 2007), are categorically more complicated than any of the problem types analyzed in this paper. when multiple management actions are possible for each site, and the value of a site is influenced by the quality of neighbouring sites, the problem becomes spatially non-linear (the ultimate conservation value of a site depends on habitat quality and conservation actions taken at nearby sites). the spatial pattern of restoration actions influences occurrence levels of species in potentially conflicting ways: an action that is good for one species may be bad for another, and effects of connectivity influence the distribution of species. finding the optimal set of sites and actions is therefore no longer straightforward. the problem may become even more complicated when species have diverse habitat requirements (for example different requirements in different stages of their life cycle). it is evident that further work is required to formulate general solutions for these more complicated, but nevertheless common types of multi-action conservation planning problems. acknowledgements we thank m. cabeza, a. arponen, o. ovaskainen and e. meyke for helpful comments on an earlier version of the manuscript. this work was financially supported by the academy of finland, project #202870. references arponen, a., r. k. heikkinen, c. d. thomas, and a. moilanen. 2005. the value of biodiversity in reserve selection: representation, species weighting, and benefit functions. conservation biology 19:2009-2014. arponen, a., h. kondelin, and a. moilanen. 2007. area-based refinement for selection of reserve sites with the benefit-function approach. conservation biology 21:527–533. bazaraa, m. s., h. d. sherali, and c. m. shetty. 1993. nonlinear programming: theory and algorithms. john wiley & sons, new york. bevers, m., j. hof, b. kent, and m. g. raphael. 1995. sustainable forest management for optimizing multispecies wildlife habitat: a coastal douglas-fir example. natural resource modeling 9:1-23. burgman, m. a., d. b. lindenmayer, and j. elith. 2005. managing landscapes for conservation under uncertainty. ecology 86:2007-2017. cabeza, m. 2003. habitat loss and connectivity of reserve networks in probability approaches to reserve design. ecology letters 6:665–672. cabeza, m., m. b. araújo, r. j. wilson, c. d. thomas, m. j. r. cowley, and a. moilanen. 2004. combining probabilities of occurrence with spatial reserve design. journal of applied ecology 41:252-262. cabeza, m., and a. moilanen. 2001. design of reserve networks and the persistence of biodiversity. trends in ecology and evolution 16:242-248. cabeza, m., and a. moilanen. 2006. replacement cost: a practical measure of site value for costeffective reserve planning. biological conservation 132:336-342. cipollini, k. a., a. l. maruyama, and c. l. zimmerman. 2005. planning for restoration: a decision analysis approach to prioritization. restoration ecology 13:460-470. cowling, r. m., r. l. pressey, r. sims castley, a. le roux, e. baard, c. j. burgers, and g. palmer. 2003. the expert or the algorithm? comparison of priority conservation areas in the cape floristic region identified by park managers and reserve selection software. biological conservation 112:147-167. crossman, n. d., and b. a. bryan. 2006. systematic landscape restoration using integer programming. biological conservation 128:369-383. cuperus, r., m. bakermans, h. a. u. de haes, and k. j. canters. 2001. ecological compensation in dutch highway planning. environmental management 27:75-89. cuperus, r., k. j. canters, h. a. udo de haes, and d. s. friedman. 1999. guidelines for ecological compensation associated with highways. biological conservation 90:41-51. donald, p. f., and a. d. evans. 2006. habitat connectivity and matrix restoration: the wider implications of agri-environment schemes. journal of applied ecology 43:209-218. elith, j., c. h. graham, r. p. anderson, m. dudik, s. ferrier, a. guisan, r. j. hijmans, f. huettmann, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. m. overton, a. t. peterson, s. j. phillips, k. richardson, r. van teeffelen and moilanen — selecting optimal conservation actions 12 scachetti-pereira, r. e. schapire, j. soberon, s. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species' distributions from occurrence data. ecography 29:129-151. guisan, a., and w. thuiller. 2005. predicting species distribution: offering more than simple habitat models. ecology letters 8:993-1009. hilderbrand, r. h., a. c. watts, and a. m. randle. 2005. the myths of restoration ecology. ecology and society 10:19. hof, j., m. bevers, l. joyce, and b. kent. 1994. an integer programming approach for spatially and temporally optimizing wildlife populations. forest science 40:177-191. hof, j., m. bevers, d. w. uresk, and g. l. schenbeck. 2002. optimizing habitat location for black-tailed prairie dogs in southwestern south dakota. ecological modelling 147:11-21. hof, j., and m. g. raphael. 1993. some mathematical programming approaches for optimizing timber age-class distributions to meet multispecies wildlife population objectives. canadian journal of forest research 23:828-834. holzkämper, a., a. lausch, and r. seppelt. 2006. optimizing landscape configuration to enhance habitat suitability for species with contrasting habitat requirements. ecological modelling 198:277-292. iucn. 1994. guidelines for protected area management categories. cnppa with the assistance of wcmc. iucn, gland, switserland an cambridge, uk. kleijn, d., f. berendse, r. smit, and n. gilissen. 2001. agri-environment schemes do not effectively protect biodiversity in dutch agricultural landschapes. nature 413:723-725. köhler, b., a. gigon, p. j. edwards, b. krusi, r. langenauer, a. luscher, and p. ryser. 2005. changes in the species composition and conservation value of limestone grasslands in northern switzerland after 22 years of contrasting managements. perspectives in plant ecology, evolution and systematics 7:51-67. manchester, s. j., s. mcnally, j. r. treweek, t. h. sparks, and j. o. mountford. 1999. the cost and practicality of techniques for the reversion of arable land to lowland wet grassland an experimental study and review. journal of environmental management 55:91-109. margules, c. r., and r. l. pressey. 2000. systematic conservation planning. nature 405:243-253. mitchley, j., and p. xofis. 2005. landscape structure and management regime as indicators of calcareous grassland habitat condition and species diversity. journal of nature conservation 13:171183. moilanen, a. 2005. reserve selection using nonlinear species distribution models. american naturalist 165:695-706. moilanen, a. 2007. landscape zonation, benefit functions and target-based planning: unifying reserve selection strategies. biological conservation 134:571-579. moilanen, a., and m. cabeza. 2002. single-species dynamic site selection. ecological applications 12:913-926. nicholson, e., m. i. westphal, k. frank, w. a. rochester, r. l. pressey, d. b. lindenmayer, and h. p. possingham. 2006. a new method for conservation planning for the persistence of multiple species. ecology letters 9:1049-1060. önal, h. 2003. first-best, second-best, and heuristic solutions in conservation reserve site selection. biological conservation 115:55-62. opdam, p., j. verboom, and r. pouwels. 2003. landscape cohesion: an index for the conservation potential of landscapes for biodiversity. landscape ecology 18:113-126. pieterse, n., a. verkroost, m. wassen, h. olde venterink, and c. kwakernaak. 2002. a decision support system for restoration planning of stream valley ecosystems. landscape ecology 17:69-81. possingham, h. p., i. ball, and s. andelman. 2000. mathematical methods for identifying representative reserve network. pp. 291-306 in s. ferson, and m. burgman, eds. quantitative methods for conservation biology. springer, new york. pöyry, j., s. lindgren, j. salminen, and m. kuussaari. 2004. restoration of butterfly and moth communities in semi-natural grasslands by cattle grazing. ecological applications 14:1656-1670. pressey, r. l., h. p. possingham, and j. r. day. 1997. effectiveness of alternative heuristic algorithms for identifying indicative minimum requirements for conservation reserves. biological conservation 80:207-219. pressey, r. l., h. p. possingham, and c. r. margules. 1996. optimality in reserve selection algorithms: when does it matter and how much? biological conservation 76:259-267. pykälä, j. 2003. effects of restoration with cattle grazing on plant species composition and richness of semi-natural grasslands. biodiversity and conservation 12:2211-2226. race, m. s., and m. s. fonseca. 1996. fixing compensatory mitigation: what will it take? ecological applications 6:94-101. rockafellar, r. t. 1970. convex analysis. princeton university press, new jersey. sarkar, s., c. pappas, j. garson, a. aggarwal, and s. cameron. 2004. place prioritization for biodiversity conservation using probabilistic surrogate distribution data. diversity and distributions 10:125-133. sarkar, s., r. l. pressey, d. p. faith, c. r. margules, t. fuller, d. m. stoms, a. moffett, k. a. wilson, k. j. williams, p. h. williams, and s. andelman. 2006. biodiversity conservation planning tools: van teeffelen and moilanen — selecting optimal conservation actions 13 present status and challenges for the future. annual review of environment and resources 31:123-159. suding, k. n., k. l. gross, and g. r. houseman. 2004. alternative states and positive feedbacks in restoration ecology. trends in ecology and evolution 19:46-53. ten kate, k., j. bishop, and r. bayon. 2004. biodiversity offsets: views, experience, and the business case. pp. 95. iucn, gland, switzerland and cambridge, uk and insight investment, london, uk. underhill, l. g. 1994. optimal and suboptimal reserve selection algorithms. biological conservation 70:85-87. van teeffelen, a. j. a. 2007. where and how to conserve: extending the scope of spatial reserve network design. phd-thesis. department of biological and environmental sciences. university of helsinki, finland. van teeffelen, a. j. a., m. cabeza, and a. moilanen. 2006. connectivity, probabilities and persistence: comparing reserve selection strategies. biodiversity and conservation 15:899-919. van teeffelen, a. j. a., m. cabeza, j. pöyry, k. m. raatikainen, and m. kuussaari. in review. maximizing conservation benefit for grassland species with contrasting management requirements. westphal, m. i., s. a. field, and h. possingham. 2007. optimizing landscape configuration: a case study of woodland birds in the mount lofty ranges, south australia. landscape and urban planning 81:56-66. westphal, m. i., m. pickett, w. m. getz, and h. p. possingham. 2003. the use of stochastic dynamic programming in optimal landscape reconstruction for metapopulations. ecological applications 13:543-555. williams, j., c. s. revelle, and s. a. levin. 2004. using mathematical optimization models to design nature reserves. frontiers in ecology and the environment 2:98-105. williams, j. c., c. s. revelle, and s. a. levin. 2005. spatial attributes and reserve design models: a review. environmental modeling and assessment 10:163-181. woodcock, b. a., r. f. pywell, d. b. roy, r. j. rose, and d. bell. 2005. grazing management of calcareous grasslands and its implications for the conservation of beetle communities. biological conservation 125:193-202. microsoft word cetal_env_2004b-ac_ik_2.doc biodiversity informatics, 2, 2005, pp. 24-41 24 environmental information: placing biodiversity phenomena in an ecological and environmental context arthur d. chapman, australian biodiversity information services, po box 7491, toowoomba south, qld 4352, australia mauro e. s. muñoz and ingrid koch centro de referência em informação ambiental (cria), av. romeu tórtima, 388, barão geraldo, 13083-885 campinas, sp, brazil abstract.—environmental niche models are increasingly being used to outline species’ distributions for a range of uses. this use has become an important component of the recent science known as biodiversity informatics. because of the nature of species’ occurrence data, considerable effort has often been spent in assessing their quality, but less attention has been paid to determining the quality of environmental data used to model species’ distributions. this paper examines a range of environmental data, and evaluates how they are prepared, their quality and use, and some commonly encountered pitfalls and problems in using environmental data in species’ distribution modeling. key words.—species modeling, environmental data, environmental modeling, climate data, data quality, scale. the world faces a challenge to manage its biodiversity resources in a sustainable manner, while conserving as much of it as possible for future generations. the study and conservation of biodiversity are not easy tasks, and the past 300 years of scientific endeavor has only just scraped the surface as far as knowing what biodiversity exists on earth, and how it functions. no amount of biological survey can adequately sample the whole earth, one country, or even a part of one country. so, in order to gain some understanding of which species occur in areas not yet surveyed or are under-surveyed, smart technologies need to be employed. there are a number of technologies available that allow the estimation of spatial distribution patterns of species (nix 1986, austin et al. 1990, margules and redhead 1995). by using environmental parameters such as climate, soils and vegetation, and knowledge of where species have been found to occur in the past, likely occurrences of species can be modeled, both now and into the future. environmental data and their use data are the essential starting point for all environmental management processes. this paper will concentrate on the non-biotic or environmental data used in biodiversity informatics along with ecological data. species’ occurrence data will not be covered here. non-biotic environmental data are increasingly being used in analyses aimed at estimating biodiversity and modeling distribution patterns of species or populations using point records obtained from collection data (faith and walker 1996, ferrier and watson 1997, williams et al. 2002). often, biological data have been collected opportunistically and thus, for large areas, it is often difficult to determine whether particular species actually occur there or not. the interrelationship between biological data and environmental data, and the knowledge of where those environments occur, can be used to fill in the gaps in the biological data. environmental data play an important role in biodiversity informatics, as all biological events are directly or indirectly related to environmental conditions. the theory behind ecological niche models is that species have certain habitat preferences that have an environmental basis (nix 1986). many models use climatological information such as temperature, rainfall, radiation, evaporation, soil moisture, etc. as the basis on which to define the habitat or ecological niche. environmental data are also generally more widely available, and generally exist in a more consistent form, than most biological data (williams et al. 2002). some ecological niche models also use classified vegetation maps, detailed habitat information, ranges of interacting species, and soil types. these data are often less accessible, in inappropriate formats, or at inappropriate scales chapman et al. – environmental data 25 for use in many biodiversity informatics studies. all too often, the data are of a categorical nature, making use in statistical models where continuous data are required difficult. these data may include both polygon-based vegetation information and pre-classified remotely sensed (rs) raster data. data exist in two basic formats: (1) primary data, such as individual point-referenced meteorological data, and (2) secondary or derived data such as climate surfaces and vegetation classifications. primary data, which are collected and referenced to individual points, largely eliminate problems of scale and categorization. categorized information commonly used to produce natural resource maps (e.g., soil types, vegetation categories, tree height classes and species – i.e. a collection of individual specimens – as well as climate layers) is problematic in a number of ways, but is essential for information presentation. one problem with pre-classified data occurs when the concepts underpinning the classification change, and thus the underlying data may become unusable. data stored as primary attributes (e.g., individual specimens, actual tree heights, etc.), in contrast, can be used to produce classified entities for display and communication while remaining available for use in alternative classifications and for use as individual data points (chapman and busby 1994). environmental data terrestrial environmental data fall into three basic categories: terrain, climate, and substrate. too often, these data are used in environmental modeling uncritically and without consideration of the error contained within, leading to erroneous results, misleading information, and even unwise decisions. terrain refers to surface morphology, and includes parameters such as elevation, slope, relief, and aspect. digital elevation models (dems) are representations of surface morphology, and can be developed at varying scales. the development of dems allows for consistent and repeatable interpolation across whole regions and constitutes the necessary first step in generating many climate surfaces (hutchinson 1991). construction of a dem is time-consuming and technically demanding (hutchinson 1991), but, once created, it doesn’t generally need to be developed again for a long time. errors in this type of data can arise in many ways, and the method used to create the dem can be important in determining both the type and dimension of likely errors. climate data are generally available from national meteorological agencies, but may have to be digitized and interpolated spatially for use in biodiversity modeling programs. spatial interpolation of climate data can be carried out with the aid of dems by fitting surfaces as smooth tri-variate functions of latitude, longitude and elevation (hutchinson 1995). these interpolations are usually developed at the scale of the dem, and, when done correctly, involve a lot of data cleaning and quality control. development of appropriately scaled climate surfaces is essential for modeling species’ distributions if models are to have any environmental meaning at scales required for management or decision-making. substrate data, both physical and chemical, can be the most difficult to obtain, and quality from one layer to another can be variable. mapping has been done in most regions of the world, but at varying scales and levels of completeness. substrate layers include soils, lithology, surficial geology, hydrology, and landform. these data are generally in the form of polygons, and are usually of a categorical nature; although in some cases continuous data may be derivable (e.g., soil texture). preparation of environmental layers is one of the most time-consuming, and computerintensive areas of modeling. fortunately, it only has to be done occasionally; once surfaces are prepared, they can be used for many models. climate surfaces, for example, have been prepared for much of the world’s land surface, and are available for use by researchers either for free or at nominal charge. these data sets, however, are at varying scales, and surfaces at suitable scales have not been available for some parts of the world until recently (e.g., south america). recent work has lead to release of globally-consistent climate layers at 30" (c. 1 km) resolution, with derived layers at 2.5', 5', and 10' resolution released in early 2004 (hijmans et al. 2004a, b). these layers are at ideal scales for modeling for both local area (30”) and continental (2.5' and 5') analysis, and provide a major advance over layers previously available. release of these layers allows for consistent modeling to be carried out across and between continents. ecological data ecological data can be of many forms: from point-based biological data and polygon-based chapman et al. – environmental data 26 vegetation data through to rs raster data. there are a number of issues associated with accuracy and error with each type of data, and types of error may vary depending on whether data are raster, polygon, or point. ecological data, such as vegetation and soils, are often categorical, and boundaries, although appearing discrete in the database, are seldom discrete in nature. for example, vegetation is usually stored as vegetation classes, and when mapped it is shown as polygons with distinct lines between one class and the next. in reality, distinctions between classes are not always clear, and mapped boundaries are usually subjective and quite artificial. in most ecological niche models, differences between classes have to be regarded as equal, whereas in reality some classes may be very close ecologically and others quite distant. this variability can lead to distinctions in the model output that may not exist in nature, and thus use of categorical data requires considerable precautions. as a result, it is often better to use categorical data as overlays in a geographic information system (gis) to refine the modeled distribution, after modeling is completed, rather than as a layer within the model itself. for example, if a species’ distribution model is obtained using climate, it can then be overlayed on vegetation types to exclude areas on unlikely vegetation types, and thus define the niche of the species more finely. ecological data are not always categorical, and continuous layers such as soil texture, ph, and water-holding capacity can be derived from them, and used in models as continuous data. managing environmental data layers an environmental data layer refers to a data set describing a characteristic of the environment that varies over a particular geographical region. environmental data, although sharing common attributes such as georeferencing, can be categorized into different data types. each data type has its own method of storage, with some data types having more than one file format. a brief introduction to the various environmental data types is given below, with information on how they are georeferenced, stored, and used. data types environmental data layers must be composed of compatible elements if they are to be manipulated individually. geographic information systems include two layer types: vector and raster. in shape files, all layer components are described geometrically, including points, lines, and polygons. in raster files, each grid square stores all of the available information for that square. point.—the point element is composed of a pair of georeferencing codes (e.g., longitude and latitude, utm), along with additional optional attributes associated with the locality the point represents. examples are gazetteer locations, specimen location data, and meteorological stations. for the latter, each station has its own georeferencing information (longitude, latitude) and additional attributes such as measured precipitation and temperature, the station’s name, the station’s responsibility, a textual description of location, etc. line.—a line data element is a set of connected linear segments. each segment has a pair of georeferencing codes (e.g., longitude, latitude) representing the beginning and the end of the segment, plus the line’s attributes. examples are roads, rivers, transect survey data, etc. in the river example, sequential georeferenced line segments describe its location; additional attributes may include name, flow direction, etc. polygon.—a polygon element is composed of a set of ordered georeferenced points such that the first and the last points are coincident (thus the shape is closed and defined) (noonan 2003). the points represent the polygons’ vertices, which can be ordered clockwise or counterclockwise. polygons are used to delimit geographical regions and each also has its own attributes. examples are cities, conservation areas, vegetation and soil classes, and rivers and roads (when their widths are relevant). in the city example, the polygons can delineate city limits and the attributes can be the city’s name, population size, etc. grid.—the grid element is a georeferenced matrix of cells. usually the grid represents a rectangular region, as does each cell. a rectangular grid is defined by the four georeferenced points that represent the corners of the rectangle, cell width and height, attributes of the grid, and individual cell values. examples of grid include climate grids, dems, and land use summaries. grids are used to represent phenomena that vary continuously over a geographic area. the phenomena can be discrete like soil type and land use, or continuous like temperature and elevation. when representing continuous phenomena, the grid stores samples of the chapman et al. – environmental data 27 phenomenon and not the phenomenon function itself. the information accuracy therefore depends on cell dimensions, because each cell holds one value to represent the phenomenon over the total of its region. some phenomena cannot be measured, are not important, or do not make sense for certain areas. for example: soil type in water, water ph in land, political divisions, etc. to represent the information in cells where the phenomenon value is not known the grid element is given a special value called “novalue“. the value actually used to represent “novalue“ can be different from one grid to another, and is defined in the grid metadata stored in the file’s header. georeferencing georeferencing is a simple concept, but is a difficult task. the concept of georeferencing is related to locating or positioning something on the earth or relating it to the “real world.” problems arise in trying to define the earth’s surface mathematically, because the earth has a highly irregular surface. the solution is to represent the surface by its ellipsoid of revolution, and to use a geographic coordinate system (gcs) to locate points on the surface. the gcs is a spherical coordinate system aligned with the spin axis of the earth. it defines two angles measured from the center of the earth. one angle (latitude) measures the angle between any point and the equator line. the other angle, called longitude, measures the angle along the equator from an arbitrary point on the earth. greenwich, england, is the accepted zerolongitude point (sobel 1995, wikipedia 2004). a problem arises when trying to fit the gcs, which is spherical, to an ellipsoid. to solve this problem, the concept of a ‘datum’ was created. a datum is a set of points used to reference the gcs position in the sphere to the ellipsoid of revolution. depending on the region of the earth being georeferenced, different datums are used so that the gcs better approaches the ellipsoid in the target region. to define a coordinate system one thus needs to know: • ellipsoid of revolution • datum • gcs spheroid radius (tied to datum definition) • origin meridian (usually greenwich) examples of ellipsoids are: • intl – international 1909 (hayford) • new intl – new international 1967 • wgs84 world geodetic system 1984 • evrst69 – everest 1969 examples of datums are: • wgs84 – world geodetic system 19841 • gda94 – geocentric datum of australia 1994 • nad83 – north american datum 1983 • sad69 – south american datum 1969 usually the datum name is sufficient to define the coordinate system (e.g., wgs84). projection is another aspect that influences the way maps are georeferenced. the earth is approximately a 3d sphere, but the usual way to communicate visual or graphical information is as a bi-dimensional or flat surface, which has lead to a proliferation of different projections. many projections are regional, while others are historical. projections or representations that function well at equatorial latitudes do not always function well at high latitudes, and viceversa. examples of projections are: • universal traverse mercator (utm) • albers equal area • azimuthal equidistant • equidistant cylindrical projection • hammer-aitoff equal area projection • geographic coordinates some projections preserve distances between points, so that one can measure the distance and multiply it by the map scale to obtain the real distance (e.g., lamberts conical projection). other projections preserve areas, but not distances (e.g., albers equal area). choice of the gcs and/or projection is determined by the location and extent of the region to be mapped and the way the map is to be presented and used. poor choices will lead to inaccurate or distorted data. a consequence is that when different sources of data are used, coordinate system transformations can become an important task for the environmental data user, and especially when working over large regions or areas. in general, analyses over large regions or globally use wgs84, as it is generally regarded as the best datum for relating one continent with another. working with grid layers grid layers are the most commonly used data 1 http://www.wgs84.com. chapman et al. – environmental data 28 type for ecological niche modeling. they also generally cause the greatest difficulties for users. different geographic information systems (gis) use different approaches to deal with grid layers. file formats.—grid layers are stored in many different formats. while some formats are open, others are restricted. open file formats have an internal organization that is publicly available, and relevant software applications know how to read and write them. on the other hand, restricted grid file formats are generally not known, and thus cannot be read or written by different software applications. examples of open grid file formats are: geotiff (geographic tagged image file2) and arc/info ascii grid (from esri®). fortunately for software developers, many computer libraries (e.g., gdal) provide implementers with uniform ways to access different formats. data related to georeferencing are stored in a grid’s header file. the header information varies with each file format. despite differences between formats, there is common information used by almost all formats: region extent, number of cell rows and columns, cell dimensions, georeferencing system, projection, content unit (e.g., km, m, kg), and “no data” value. some grid formats include additional software-specific information to speed up or simplify specific software tasks, such as maximum, minimum and average cell values; how cell values should be translated into colors; and cell value storage type (e.g., byte, word, floating point, etc). the user of a single application does not have to worry about file formats, as the application will handle reading and writing. on the other hand, if the user wants to use the grid in different applications, it is necessary to ensure that all applications can correctly handle the chosen file format(s). generating grid information.—another important issue with grid layers is the way information is read and written. as mentioned above, each cell in a grid layer stores a value associated with the region it represents and the grid is used to store a phenomenon over a whole region. due to the discrete nature of the cell and cell region of the grid, the phenomenon information needs to be discretized to fit in the grid. if the phenomenon is continuous (e.g., temperature), the cell can store a sample (e.g., the value at the center of the region) or the average value for the cell region. if the 2 http://remotesensing.org/geotiff/geotiff.html. phenomenon is discrete (e.g., soil category), the cell can store the most important, most abundant, or even a random category found within the region. choice of what is stored in a cell is determined during grid creation, and can depend on the way the phenomenon is captured. reading grid layer information.—the reading process presents the inverse problem. the grid is a discrete matrix of samples (or averages) and a map of continuous values (e.g., temperature) and/or continuous regions (e.g., temperature or soil categories on a shore) may be required. for example, figure 1 shows four cells (a, b, c, and d) of a grid that stores a phenomenon (annual average temperature). each cell region encompasses 1 x 1° of longitude and latitude, and has an associated temperature value. according to knowledge of the way the grid was originated, the value at (x, y) can be estimated using one of several methods. resampling of the grid values can assist in making this estimation, and is used when one wants to read a value of a grid layer assuming that the information is continuous within the region. figure 1. a 4-cell grid with a representative point (x,y). resampling.—the first step in resampling a grid layer is to find a function that represents its phenomenon. this function is usually piecewise, using a combination of values corresponding to cells in the neighborhood of the point being resampled. the phenomenon value is then calculated as the function at the desired coordinate. the three most common methods used are nearest neighbor, bilinear interpolation, and cubic convolution. with nearest neighbor, the value returned is the value of the cell at the given chapman et al. – environmental data 29 point (x, y). in other words, for all points (x, y) in a cell region, the resulting value is the value of that cell. this method is the fastest, and is the best for use with categorical data, because it assures that the returned value is present in the original grid. for example, in figure 1, the returned value for (x, y) is the value of cell a, 30. bilinear interpolation is a weighted average of the values of four nearest cells (cells a, b, c and d in figure 1). this method should not be used for categorical grids, as the result is not guaranteed to be a valid category. this method guarantees that the resulting value is always within the range of the 4 nearest cell values. in cubic convolution the 16 nearest cells values are used to find a smooth surface (usually with cubic splines) that passes through all its centers. the resultant value is not guaranteed to be within the values of the cells used. this method should not be used with categorical grids, as the result can be different from those defined in the grid. although computationally intensive, it gives good results when grid phenomena are continuous. digital elevation models use in biodiversity informatics digital elevation models (dem) or digital terrain models (dtm) form one important element in ecological niche modeling. they are used directly in providing data on elevation, slope, and aspect, and indirectly in development of climate layers. elevation can be an important environmental variable in determining niches of species. it is well known that some species grow at high altitudes, and others at low altitudes. the reasons are often climatic (temperature, occurrence of frost and snow, etc.), and for this reason elevation is important in derivation of climate layers. likewise, slope and aspect are important driving characteristics for species’ niches. some species grow preferentially on northern (southern hemisphere) or southern (northern hemisphere) aspects to obtain more of the sun’s warmth during the day. other species are the opposite, and prefer cooler daytime temperatures. some species prefer to grow in areas of high slope where water may not accumulate and pool and frost slides off, while others prefer flat areas where soil water content may be more consistent and less variable. methods of development stereo photogrammetry.—dems have been around for many decades. traditionally, they have been created using a combination of spot height information determined through field surveys and stereo-photogrammetry using aerial photographs. these methods are timeconsuming, laborious, and costly, and are really only suitable over small areas. data from these methods, especially the spot-height information, provide valuable information for use with the other techniques mentioned below. anudem.—the most common method for developing dems in recent years has been that used in the anudem program developed at the centre for resource and environmental studies in canberra (hutchinson 1989), a version of which is now included in arcinfo as topogrid. the method iteratively applies a spline interpolation algorithm to calculate values on a regular grid using irregularly spaced elevation data points from contour line data, streamline data and individual spot heights. the strength of the anudem method over most other methods is that it imposes a global drainage condition through an approach known as drainage enforcement, to produce elevation models that represent more closely the actual terrain surface and which contain fewer artifacts than those produced with more general-purpose surface interpolation routines (usgs 2003). remotely sensed imagery.—several remote-sensing (rs) techniques can also be used to create dems, such as use of radarsat or jers images like the synthetic aperture radar (sar). these methods can be used directly to create a dem using interferogram (phase difference) techniques that use two (or more) images of the same area taken from different angles (rao and rao 1999). it is predicted that these methods, possibly in conjunction with some others, could lead to accuracies as small as 1 mm (arora et al. 2002) but that is only likely to be over very small areas, and would be costly in both resources and computing power. combination of methods.—more recently, techniques have been developed that use the results derived from radarsat imagery in conjunction with the anudem software to provide a dem of much greater accuracy. this approach has been especially valuable in areas with little terrain variability, and has recently been used to create a dem of the antarctic (liu et al. 2001). accuracy and sources of error dems can be quite variable in accuracy, chapman et al. – environmental data 30 depending on their development method and the availability of suitable data from which to derive them. the elevation error at a single point in a dem depends on the resolution (cell size) and the roughness of the surface being modeled (hutchinson 1996, 2003b). imposed global drainage as used in anudem, has been found to increase dem accuracy significantly, especially in terms of drainage properties (hutchinson 2004). the usgs 1 km dem version 1.0, for example, is variable in accuracy across different areas, and is known to have considerable error in some areas of south america (ngdc 2000). recent use of anudem has led to development of dems with much less error than previous methods (hutchinson 1996) and these methods are now being used across much of the world to develop accurate, continental-scale dems at scales as fine as 80 m or 2” of latitudelongitude. an example is the global 30” dem-gtopo30--produced through a collaborative effort coordinated by the usgs’s eros data centre (usgs 2003). although output grids are improved over previous versions, their accuracy is still limited by the accuracy of the source materials used to create them (olsen and bliss 1997). ideally, using anudem, accuracy of the interpolation should approach one-half of the contour intervals of the source data. climate data climate data, as used in ecological niche modeling, are interpolated surfaces developed from information such as temperature and rainfall from weather stations at known locations, and integrated with terrain data (usually a dem) to form a smooth coverage over the earth’s surface. these data allow a user to estimate climate conditions at any point on the surface. this information is important in modeling because weather stations are not sufficiently abundant to permit accurate determination of climate at points where weather stations do not exist, and thus it is essential to use such derived surfaces. use in biodiversity informatics climate information has formed the basis of most ecological niche modeling applications over the past 20 years. it has been used in models associated with biogeographic studies (longmore 1986, peterson et al. 1999), conservation planning (faith et al. 2001), reserve selection (margules and pressey 2000), development of environmental regionalizations (thackway and cresswell 1995), climate change studies (chapman and milne 1998, peterson et al. 2002), agriculture and forestry production (booth 1996, nicholls 1997, cunningham et al. 2001), species translocation studies (mackey 1996, soberón et al. 2000, peterson and vieglais 2001), etc. in many cases, climate layers have been used as fairly raw layers, such as maximum and minimum temperatures in certain months of the year. one of the most important factors in determining where a plant may grow, however, at least at the macro-level, is the relationship between rainfall and temperature. agronomists have long relied on this knowledge to plan summer and winter plantings of different crops. some species respond to rainfall at certain times of the year, and others at other times. rainfall in the middle of winter, for example, may have little or no effect on species that are dormant during that period. in more tropical areas, however, this season may be the key growing period, as many plants reduce transpiration over summer when it is too hot, and most of the growing is done during the cooler period of the year. alternatively, a long dry period in the middle of a hot summer may have more detrimental effects on a plant than the same long dry period during the middle of winter. long experience of modeling in australia has determined that this combination of layers has more environmental relevance, and generally produces better models, than just raw monthly values (nix 1986). layers such as mean temperature of the wettest and driest, warmest and coolest quarters, and mean precipitation of the warmest and coolest, wettest, and driest quarters relate well to the environmental conditions that determine where a plant or animal is likely to occur. we recommend that consideration be given to greater use of such synthetic layers in future modeling efforts. methods of development the distribution of meteorological stations around the world is very uneven. traditionally, stations have been established in urban areas and areas of high agricultural production, with predominantly natural areas having very poor coverage. as such, important areas such as mountaintops, and wilderness areas have very little detailed climate information. because of the sparseness of this information, algorithms have had to be developed to fill in gaps and to produce chapman et al. – environmental data 31 a ‘blanket’ of climate information to cover all areas. rainfall and temperature patterns are heavily reliant upon altitude, slope, aspect, positioning of hills and mountains, and relationship to large water bodies. several algorithms have been developed to fit climate surfaces using known meteorological data points in conjunction with surface topology to create climate layers for use in ecological niche modeling. the underlying dem and the meteorological station data, therefore, form the basis of most (if not all) of these surface-fitting algorithms. different methods and algorithms used for interpolation produce layers with important differences. a few of the more important differences are discussed briefly below. all of these methods rely on the underlying dem to create the surfaces, and the finer and more accurate the dem, the better are the resultant surfaces. some algorithms (e.g., anusplin) place greater reliance on the dem than other methods and are thus more likely to result in more robust surfaces, but again that is dependent on the accuracy of the dem, and many have problems where relief is subtle. centro internacional de agricultura tropical (ciat, columbia).—the ciat method uses a simple interpolation algorithm based on the inverse square of the distance between the station and the interpolated point of the nearest five stations (ciat n.d.). this method has the advantage of speed and ease of use for large data sets where computational capacity is limited (booth and jones 1996). the influence of a bad data point can be significant and can cause significant circling in the resultant surface. because it is using only five data points, it relies less on the underlying dem than other methods. anusplin (australian national university).—anusplin (hutchinson 2001) is a technique that uses partial thin-plate splines to interpolate multivariate data. it is made up of nine programs (kesterven and hutchinson 1996), was developed in the 1980s, and has been refined extensively since. it is a proven methodology, with wide acceptance across the world. recent modifications now allow for simultaneous analysis of several surfaces (the concept of “surface independent variables”) (kesterven and hutchinson 1996). cost to the user is not high, and most of the work can be carried out ”inhouse“. the advantage in using anusplin is its heavy reliance on the underlying dem, and thus its tendency to produce more finely resolved surfaces. because most major climate-summary efforts around the world are using anusplin, surfaces created are likely to be consistent with surfaces in other parts of the world. this method has recently been used to develop globally consistent (as far as the meteorological data permits) 30" climate surfaces covering most areas of the earth’s terrestrial surface (hijmans et al. 2004a, b). parameter-evaluation regressions on independent slopes model (prism).—prism is a knowledge-based interpolation system developed at oregon state university in the 1990s. it has been used principally for developing climate layers for the united states (daly et al. 1994; gibson et al. 2004). the method incorporates a number of spatial interpolation quality control measures in similar ways to anusplin. it is being used in conjunction with automated data collection in the usa (daly et al. 2004). other methods.—several other methods have been used recently for estimating climates. these methods include regression (zheng and basher 1996), inverse distance (goovaerts 2000), first detrend for elevation, exposure, orographic influences, then kriging (holdaway 1996), cokriging with elevation (phillips et al. 1992), and gradient plus inverse-distance-squared (gids; nalder and wein 1998) summary several studies have compared methods for fitting climate surfaces. in nearly all cases, the conclusions have favored partial thin-plate spline techniques over others. for example, one study compared methods for interpolating climate surfaces using mexican data (hartkamp et al. 1999); the authors concluded: “taking in account error prediction, data assumptions, and computational simplicity, we would recommend use of thin-plate smoothing splines for interpolating climate variables.” another study (booth and jones 1996) suggested that more complex interpolation algorithms such as laplacian splines are better interpolators than most of the simpler methods, but need much more computing power. the same study states “the major difference between the techniques used by ciat (jones et al. 1990) and cres [=anusplin, above] (hutchinson 1989) is that the ciat method uses a standard lapse rate applied over the whole dataset; the cres method uses a 3-dimensional spline algorithm to determine a local lapse rate from the data.” they concluded that anusplin provides a powerful chapman et al. – environmental data 32 set of programs for climate analysis. a comparison between gids and anusplin (price et al. 2000) concluded that thin-plate splines generally produced better smoothing and better gradients at high elevations and in areas where climate station coverage was poor, and in predicting climate variables at points withheld at random from the source datasets. a study in new zealand (barringer and lilburne 2000) looked at methods for determining soil surface temperatures and also concluded that partial thinplate spline surfaces had the lowest residuals for long-term mean monthly and specific month/year soil temperature surfaces, and that multiple linear regression provided a simple and robust method for soil temperature interpolation where the data were not strongly spatially dependent. time period one issue often not taken into account is the time period from which the climate surfaces being used were developed. for many climate stations, it is difficult to obtain runs of consistent climate data for periods greater than 10 years. to prepare robust climate surfaces and ease out short-term variation, one should seek to obtain at least 30-year runs of data. wherever possible, for modeling with older species’ occurrence information, climate layers should be prepared for runs prior to at least 1990, and preferably as far back as 1970 (i.e., before recent climate change effects became noticeable). accuracy and sources of error similar to species’ occurrence data, mislocation of weather stations, or errors in readings at those stations, can create errors in resultant climate surfaces. georeferencing weather stations can be just as tedious and error-prone as georeferencing species’ occurrence data, but at least there are fewer weather stations than species collections! one common problem is that the geocode refers to the center of the town, when the meteorological station may be kilometers away (busby 1991). when prepared critically, and once individual meteorological data records have been cleaned, climate layers will have significantly lower levels of error. errors are of two major types--positional and attribute. positional error depends largely on the accuracy of the underlying dem. the dem for south africa, for example, at 10' has a standard error of between 20-150 m (hutchinson 2003a). the attribute error for the climate data, on the other hand, is different for temperature and rainfall. because anusplin depends on elevation, it is significantly more accurate than methods that use bivariate functions of longitude and latitude only (margules and redhead 1995). standard errors of temperature data are of the order of 0.5° and of rainfall about 5-15%, depending on data density and spatial variability of the actual monthly mean rainfall (margules and redhead 1995; hutchinson 1996, 2003a). one of the greatest sources of error in climate surface development is lack of reliable meteorological stations across large areas of the world. the overall accuracy of climate surfaces derived, therefore, depends largely on the ability of different methods to handle this scarcity of data, and to interpolate best into areas where meteorological data are scarce. evapotranspiration evapotranspiration (et) is the term used for the transfer of water, as water vapor, from land surfaces (both vegetated and non-vegetated) to the atmosphere by evaporation and transpiration. et depends on the energy supply (mainly direct solar radiation), vapor pressure, and winds, and is affected by climate, water availability, and vegetation cover. it is difficult to obtain accurate field measurements of physical parameters necessary to measure evapotranspiration, so procedures have been developed to assess et from meteorological data to produce continuous maps of et or potential evapotranspiration (pet). wang et al. (2004) adapted the complementary relationship areal et model (morton 1983) to estimate and map evapotranspiration as ‘areal actual,’ ‘areal potential,’ and ‘point potential’ using modified priestley-taylor and energy transfer and balance equations. the food and agriculture organization (fao), on the other hand, uses the penman-monteith equation, which they found best for use at global scales (allen et al. 1998). a third method often used is that of hargreaves and samani (1982), which estimates potential evapotranspiration as a function of solar radiation and air temperature. accuracy and sources of error mapping et is affected by the spatial coverage of available climate stations, accuracy of interpretation of vegetation cover, and by interpolation and mapping techniques used. errors appear greater at high latitudes owing to unreliability of solar radiation estimates in those chapman et al. – environmental data 33 areas. in a comparison of the penman-monteith and hargreaves methods, reynolds et al. (2000) recommended that the penman-monteith method was more data intensive and the most reliable method where accurate data was available, but that the hargreaves method was a preferred alternative where accurate data collection is less certain. et can be valuable in ecological niche modeling, however, it “is almost impossible to measure or observe directly at a meaningful scale in space or time” (wang 2004). topography topography is another set of layers derived from dems. the two most common derived layers are slope and aspect, although specific catchment area and contour curvature are also sometimes used for determining et and soil moisture (gallant and hutchinson 1996). the fineness of the underlying dem is key in determining the accuracy of the derived slope and aspect layers (see, e.g., figure 5). secondary layers such as solar radiation are based on slope and aspect and modified by topographic shadowing (moore et al. 1993; gessler et al. 1995). use in biodiversity informatics several process-based landscape-scale ecological niche models have included topographic data in their development (gallant and hutchinson 1996). accuracy and sources of error topographic layers are sensitive to the resolution of the source from which they were derived, the underlying dem. surface reconstruction using contouring from dems has been shown to be largely dependent on the scale of the dem, which can lead to large variation and error in derived layers such as slope, aspect, solar radiation, catchment area, soil moisture, and contour curvature (gallant and hutchinson 1996). soils discrete soil measurements can be converted to continuous soil layers by environmental correlation with continuous spatial data sets (e.g., dems, rs imagery) using statistical, geostatistical, or numerical models. the results of these models produce soil attributes necessary to generate continuous data sets for use in niche modeling. gessler et al. (1995) used the compound topographic index (cti) for allocating field sample locations and exploratory data analysis to search for useful relationships between modeled soil attributes and environmental attributes in a spatially continuous manner. these relationships were then confirmed and defined by statistical models, and improved by field verification. the methodology was improved on by gessler and chadwick (1997). continuous soil data such as ph, soil water holding capacity, texture, chemical composition, etc., are layers that may be used to advantage in niche modeling. remotely sensed data remotely sensed (rs) data represent a powerful resource where information about large geographic areas is required. rs data cover the energy captured from a sensor distant from the object or radiating phenomenon. data captured from airplanes and satellites are the most common type of rs data for use in biodiversity informatics and are what most people understand by the term “remotely-sensed” data. the two most common types of remote sensors used for the study of the earth’s surface are optical and radar. optical sensors use the visible, near-infrared, and short-wave infrared parts of the spectrum to form images of the earth’s surface by detecting solar radiation reflected from objects on the ground. different materials reflect and absorb radiation differently at different wavelengths, so objects such as bare rock, vegetation types, etc., can be differentiated by their spectral reflectance signatures. radar sensors use radar or sonar to detect variations in terrain, including ocean depths, elevation, etc. different types of optical remote sensors include (liew 2004): panchromatic – a single channel sensor that covers one broad wavelength, and measures the apparent brightness of the object; the resulting image resembling a black-and-white photograph (e.g., spot hrv-pan) multispectral – a multi-channel sensor covers just a few spectral bands (usually 3-7), with each channel recording reflectance in a narrow spectral band; results in a multi-layer image that can be colored in various ways to emphasize particular characteristics (e.g., avhrr, landsat tm) superspectral – has many more channels (usually >10), with each channel covering much narrower bandwidths, enabling finer characteristics of the environment to be detected (e.g., modis) hyperspectral – instruments known as “imaging spectrometers” acquire information in 100+ chapman et al. – environmental data 34 contiguous spectral bands. these data are used for monitoring phytoplankton, pollution, etc., but have high potential for use in biodiversity informatics. given the extraordinary amounts of data being reported, however, this data resource is beyond most existing project specifications (e.g., hyperion on eo1). raw rs data, the radiometric information data resulting directly from the sensors, needs to be preprocessed to generate useable thematic information, such as sea surface temperature or vegetation types. these exercises can be timeconsuming and computer-intensive, and are usually best carried out by specialists. use in biodiversity informatics a main use for rs data in biodiversity informatics has been the derived normalized difference vegetation index (ndvi), or greenness index. it approaches a representation of ‘greenness’ or amount of photosynthetic mass. it is calculated from reflected solar radiation in the near infrared (nir) and red (red) wavelengths as: ( ) ( )rednir rednirndvi + − = ndvi is correlated with photosynthesis through absorption of red light by plant chlorophyll and reflection of infrared radiation by water-filled cells in leaves. it is thus commonly used as an estimate or surrogate of green vegetation (deh 2004). in biodiversity informatics, it has largely been used as an overlay in a gis as a surrogate for vegetation following modeling (e.g., chapman and milne 1998). because it is a continuous (non-linear) function that varies between –1 and +1, it could easily be used as a layer in niche modeling (pereira 2002). its biggest problem may be when red and nir are zero, as then ndvi is undefined. the scale at which rs data are used determines their use and accuracy. the most commonly used images are from avhrr images at about 1 km pixel size. other commonly used images are from landsat (25100 m pixel size) and low-altitude spot images (about 10 m pixel size). for fine-scale regional studies, images from sensors similar to those used in satellites can be acquired from airplanes, bringing pixel sizes down to <1 m. such finescale data are expensive, however, and unlikely to be suitable in many modeling studies, other than as gis overlays. accuracy and sources of error many sources of error enter into rs data. these sources include variation between sensors on different satellites, signal decay over time, angle of incidence between images, angle of the sun’s rays when the image was taken, rectification of the image relative to the earth’s surface, and variation in screening procedures for cloud and pollutants in the atmosphere. several algorithms exist for calculating the magnitude of these errors. for example, calibration equations exist for rectifying data from various satellites (rao and chen 1995, 1996). these calibrations are usually calculated using reflectance in areas of low vegetation cover (e.g. deserts, water bodies) where vegetation cover (and hence reflectance) doesn’t vary over time. low sun angle can lead to poor data quality owing to large angle corrections--shadows in steep terrain, increased reflectance off atmospheric pollutants, etc. in australia, between april and september 1994, for example, noaa 11 data deteriorated considerably as to be almost unusable owing to low sun angles (deh 2004). one key aspect of error in rs data is the accuracy with which a pixel can be located on the ground. the margin of error, depending upon the method of geometric registration, is generally accepted to be ~1 pixel. thus, avhrr, which has a pixel size of ~1 km, has an accuracy of about +/1 km (mao et al. 1999). in many cases, however, especially in areas of the earth with few identifiable registration points (e.g., australian deserts, marine areas, the amazon), pixel accuracy cannot be relied on at better than +/2 pixels (s. cridland pers. comm. 2000). to account for cloud cover, cloud screening masks have been developed (deh 2004). these tools seek pixels that do not appear biologically consistent with images before and after the image of interest. cloud screening usually includes subjective steps, and thus are a possible source of error. vector-based ecological data not all environmental data are raster-based. terrain-based polygon data, for example, form the basis of many traditional paper-based maps, as well as many digital maps used in gis. these maps can have quite varying levels of accuracy. for example, topo-250k for australia chapman et al. – environmental data 35 (1:250,000 topographic data) is a well-researched data set and its accuracy is described as “not more than 10% of well-defined points being in error by more than 160 meters; and in the worst case, a well defined point is out of position by 300 meters” (geoscience australia 2003). this accuracy can be quite important, for example, if one is trying to determine if a species (with accuracy ~1 km) occurs in a national park (with an accuracy of ~160 m). with paper topographic maps, drawing constraints may restrict accuracy with which lines are placed. a 1 mm wide line depicting a road on a 1:250,000 map represents 250 meters on the ground. to depict a railway running beside the road, a separation of 1-2 mm (250-500 meters) is needed, and then the line for the railway (another 1 mm or 250 meters) makes a total of 750-1000 m as a minimum representation. if one is using such features to determine an occurrence locality, for example, then maximum precision would be ~1 km. accurate coastline representation (figure 2) can also be a nightmare--does the map use the highwater mark or mean sea level, and have neap, spring, or king tides been taken into account (bannerman 1999)? when it crosses the mouth of a river, does it take a direct line across, or does it follow the river upstream for a distance? are rivers depicted by two lines (one for each bank), or by just a centerline? similarly, are towns shown as points, and if so what part of the town does that point represent? if towns are represented by polygons, is it the municipal boundary or the boundary of outer development (wieczorek 2001)? how is terrain represented-just by contours, thus excluding highest and lowest points, or does it show spot heights? how accurate are they? does it also show low points? many features (e.g., coastlines and rivers) change over time. in many tropical areas of the world, seasonal billabongs are formed--how are they represented (bannerman 1999)? again, the scale of the map can make a significant difference. the depiction of phenomena that don’t have discrete boundaries in nature (vegetation, soils, geology, etc.) as polygons is also a source of error (burrough and mcdonnell 1998; discussion under ecological data, above). where lines are drawn can lead to major errors, both geographic and attribute-related. in some cases, a vegetation description may be a mosaic of vegetation types and not a single discrete type (sattler and williams 1999). if one is attempting to map a particular community (community ‘a’, for example), one group of polygons may contain 545% community ‘a’, one group may contain 5095%, and one group of polygons may contain 100%. the polygon with 5-45% of community ‘a’, doesn’t mean that it has a cover of 5-45% of that community, but that there is 5-45% chance of finding it somewhere within that polygon. depending on the scale of the mapping, there may be a small area of 100% of that community in one small area, and none in the rest. so, if you are talking about community ‘a’, where do you make the cut off? instead of using raw data, classification might be desirable, but can cause problems in interpretation and thus error. figure 2. example of coastlines at two different scales: 1:5,000,000 and 1:500,000. scale one very important consideration in choosing environmental layers is that of scale. too fine a scale can often lead to errors due to mismatching with biological data. for example, historical biological data often do not have accuracy of better than 5-6 km, but are often used in models as single latitude/longitude points. if the climate grid is at 1 km resolution, the actual point represented by those coordinates could in reality be in any one of 25 grid squares, so just taking the one where the point appears to be located may lead to wrong assumptions in regard to the climate where the species actually occurs; such errors will be propagated throughout the model. too coarse a scale may not adequately delineate niche dimensions. the type of biological data available should determine the scale of climate layers used in modeling; only where occurrence data are obtained using geographic positioning systems (gps) should fine environmental layers be used. too often, modelers give little consideration to chapman et al. – environmental data 36 scale in their selection of environmental layers, although these choices are determined more by availability of layers. another consideration is the computing power available. if one is modeling at a continental, or broad regional scale, it may not be practical to use environmental layers at 1 km resolution that may be very large and cumbersome for computing. a grid at 5 km (2.5') occupies 1/25 the storage space of a 1 km grid, and thus requires considerably less computing power to run. if such is the case, then it is important that it be noted, and assumptions arising elaborated. use of poor environmental layers can produce models that are of very little practical value in understanding environmental factors that drive species’ niche preferences. table 1. results from modeling raulvolfia nitida using climate data from distinct sources and the garp algorithm. pixel sizes; ciat=10 minutes, worldclim=30 seconds. prediction presence (pixels) absence (pixels) ciat (3-10 models) 466506 2012289 worldclim (3-10 models) 433070 2045842 (3-10 models) common to both 226119 1806438 an example was developed to demonstrate the effects of modeling using two different sets of environmental layers--one the 10' (~18 km) grid-based climate layers derived from ciat (jones 1991), and the other the 30" worldclim climate layers (hijmans 2004a, b). we used rauvolfia nitida (apocynaceae), a shrub species that occurs in the islands of central america. model results from the genetic algorithm for rule-set prediction (pereira 2002) showed little overall difference in the total extent of the modeled distribution. however, the finer scale climate layers allowed for improved delineation of the species’ distribution (figure 3c and d). using the 10' grid, the total area identified was 466,506 km2, compared with 433,070 km2 using the 0.5' grid (table 1). models based on the finer grid tended to exclude areas where the species is unlikely to occur given the climate profile at the points of known occurrence, and identified other areas that the coarser layers missed. the scale of environmental layers can also be a problem in assembling occurrence data before modeling. as an example, see the differences that can arise from the size of the grids alone in the coverage of the caribbean islands (figure 4a-c). parts of some islands, and even some entire islands, have no climatic information at either 30' or 10' grid resolution. this effect is most evident in coastal areas, because the definition of the grids does not allow faithful overlapping of environmental information with boundaries of islands. these effects in coastal areas will strongly influence the results of modeling species occurring there, because the only environmental data used are those coinciding with occurrence points. figure 4. annual precipitation information in the caribbean, with occurrence points of species of rauvolfia (apocynaceae) overlain. arrows indicate some points that fall outside of the layers. (a) ipcc dataset (30' resolution); (b) ciat dataset (10' resolution); (c) worldclim dataset (30" resolution). in general, modeling uses mathematical procedures to analyze environmental parameters related to species’ occurrence points, and the modeled output is produced when environmental parameters are projected onto a map showing where similar conditions are found--in other words, the potential occurrence areas of a species. thus, to produce a good model, it is desirable to have a good sample of environmental data related with the occurrence points; in some situations, e.g., in coastal areas, a low-resolution environmental dataset will not allow large enough samples. this effect occurs chapman et al. – environmental data 37 not only when some occurrence points have no environmental information, but also when occurrence points are highly clumped spatially. in that situation, common in the islands, many points will fall in the same grid squares, reducing the number of spatially unique points for use in analyses. dem – slope and aspect the scale of the dem used in a model is also very important, especially when using derived layers such as slope, aspect, and landform. if one has a dem at a scale of 9" (~250 m), one can derive quite-valuable slope and aspect layers. if one then reclassifies upwards to a scale of, say, 1' (~2 km), then one has already reduced the meaningfulness of those two criteria—slope, for example, is unlikely to remain consistent over the 2 km. if resolution is then further eroded to 10' (~18 km), then these layers can become meaningless. for example, in figure 5, the slope taken over 1 km is -10%, whereas over 10 km it is +2.5% and ignores major differences in terrain. because dems are derived using complicated algorithms, they should be used only at the scale at which they were derived, rather than reclassified in any way. a b figure 5. graph showing problems with determining slope using different scales. a. shows slope across 1 km (between 10 and 11 km from coast) of –10%. b. shows slope over 10 km (between 5 and 15 km from coast) of +2.5%. issues, pitfalls and challenges in using environment data many issues and pitfalls occur in using environmental data to model species’ ecological and geographic distributions. it is just as important to evaluate the modeling approach applied as the characteristics of the data being used--both the occurrence information and the environmental data. the following are a few additional considerations of common problems and pitfalls related to environmental data that often occur in niche modeling. just because a model produces a map doesn’t mean it is a good model. “one can lie with maps just as easily as one can lie with statistics” (wein 2002), and probably have them believed easier. many species (especially rare species) occur along transition zones between vegetation types. often these transition zones are not very well delineated in environmental data layers, particularly if the scale is coarse. scale is often thought of as a step-wise process, but it is continuous. one needs also to be careful mixing environmental layers at different scales within a single model. some modeling methods require continuous data as the input layers, and use of categorical data can cause problems. this requirement can be especially problematic where error inherent in the occurrence data may make the category into which a locality falls uncertain. by restricting the geographic boundaries of a model, one is also restricting the possible environment available for the species’ modeled niche. for example, if one is modeling a broadly distributed species within just one part of its range, the climatic range of the species may be underestimated, and the resulting modeled distribution under predicted. most environmental layers used for terrestrial species modeling are not necessarily suitable for modeling aquatic species. the habitat requirements of a fish, for example, may be more related to water temperature, ph and stream flow than it is to the air temperature or precipitation in a local area. indeed, precipitation up-stream may be more important to the species’ distribution than precipitation at the actual location. modeling marine organisms presents yet a different challenge. most data available for marine ecosystems (other than depth and bathymetry) are measured in the top 30 cm of the ocean, whereas many marine organisms move through a range of environments and depths. conclusion selection of environmental data layers is among the most important aspects of ecological niche modeling. a model is, at best, only as good as the input data, be they the occurrence data or the environmental layers. too often, models are run after hours and hours of data cleaning of the occurrence data, but with blind acceptance of the environmental layers used. the occurrence data may only be used once, but the environmental data layers are used over and over in model after model, and their accuracy and nature will chapman et al. – environmental data 38 determine the reliability of all resulting models. it is important that these be evaluated critically for quality before use; otherwise, resulting models will not adequately reflect the true ecological niche or potential distribution of the species being modeled. acknowledgments the staff of the centro de referência em informação ambiental, campinas, brazil, and the environmental resources information network, australia, have contributed ideas that form the basis of this paper. their discussions of error and accuracy in environmental information over the years, and the pioneering work done by them, as well as that of conabio (mexico), inbio (costa rica), and cres (australia) have helped to bring us to the stage of modeling development that we have reached today. we thank them for their ideas and constructive criticism. in addition, discussions with a. townsend peterson and others at the university of kansas, barry chernoff at wesleyan university in connecticut, read beaman at yale university, and john wieczorek and robert hijmans at the university of california, berkeley, have presented us with ideas and challenges that have lead to some of the ideas expressed in this paper. any errors, omissions or controversies are, however, the responsibility of the authors. references allen, r., g., l. s. pereira, d. raes, and m. smith 1998. crop evapotransporation – guidelines for computing crop water requirements. fao irrigation and drainage paper 56. fao, rome. [cited 6 august 2003]3. arora, m. k., v. patel, and m. l. sharma. 2002. sar interferometry for dem generation [online]. gis development, technology, remote sensing. [cited 13 january 2004].4 austin, m. p., a. o. nicholls, and c. r. margules. 1990. measurement of the realised qualitative niche of plant species: examples of the environmental niches of five eucalyptus species. ecol. monogr. 60:161-177. bannerman, b. s. 1999. positional accuracy, error and uncertainty in spatial information. geoinnovations, howard springs, nt, australia. [cited 13 january 2004].5 barringer, j. r. f., and l. r. lilburne. 2000. developing fundamental data layers to support 3 http://www.fao.org/docrep/x0490e/x0490e00.htm#contents. 4 http://www.gisdevelopment.net/technology/rs/techrs0021.htm. 5 http://www.geoinnovations.com.au/posacc/patoc.htm. environmental modeling in new zealand: progress and problems in 4th international conference on integrating gis and environmental modelling (gis/em4): problems, prospects and research needs. september 2-8, in banff, alberta, canada 2000. [cited 26 may 2004].6 booth, t. h. 1996. matching trees and sites in proceedings of an international workshop held in bangkok, thailand, 27-30 march 1995, aciar proceedings no. 63. booth, t. h., and p. g. jones. 1996. climatic databases for use in agricultural management and research in proceedings of the arendal ii workshop on unep/grid and cgiar cooperation to meet requirements for the use of digital data in agricultural management and research. arendal, norway 9-11 may 1995. [cited 13 january 2004].7 burrough, p. a., and r. a. mcdonnell. 1998. principles of geographical information systems. oxford university press, oxford. busby, j. r. 1991. bioclim – a bioclimatic analysis and prediction system. pp. 4-68 in c. r. margules, and m. p. austin (eds) nature conservation: cost effective biological surveys and data analysis. csiro, melbourne. chapman, a. d., and j. r. busby. 1994. linking plant species information to continental biodiversity inventory, climate and environmental monitoring. pp. 177-195 in r. i. miller (ed.). mapping the diversity of nature. chapman and hall, london. chapman, a. d., and d. j. milne. 1998. the impact of global warming on the distribution of selected australian plant and animal species in relation to soils and vegetation. environment australia, canberra. ciat n.d. floramap users manual. centro internacional de agricultura tropical (ciat), cali, colombia. cunningham, d., k. walsh, and e. anderson. 2001. potential for seed gum production from cassia brewsteri. rirdc project no. ucq-12a. rural industries research and development corporation, kingston, act. [cited 10 february 2004].8 daly, c., r. p. nielson, and d. l. phillips. 1994. a statistical-topographic model for mapping climatological precipitation over mountainous terrain. j. appl. meteor. 33:140-158. daly, c., w. gibson, m. doggett, j. smith, and g. taylor. 2004. a probabilistic-spatial approach to the quality control of climate observations. in proceedings of the 14th ams conference on applied climatology, amer. meteorological soc., seattle, wa. january 13-16, 2004. [cited 26 6 http://www.colorado.edu/research/cires/banff/pubpapers/221/. 7 http://www.grida.no/cgiar/htmls/climpres.htm. 8 http://www.rirdc.gov.au/reports/npp/ucq-12a.pdf. chapman et al. – environmental data 39 may 2004].9 department of environment and heritage (deh). 2004. normalized difference vegetation index. [cited 26 february 2004].10 faith, d. p., and p. a. walker. 1996. environmental diversity: on the best possible use of surrogate data for assessing the relative biodiversity of sets of areas. biodiv. and conserv. 5:399-415. faith, d. p., p. a. walker, c. r. margules, j. stein, and g. natera. 2001. practical application of biodiversity surrogates and percentage targets for conservation in papua new guinea. pacific conservation biology 6:289-303. [cited 10 february 2004].11 ferrier, s., and g. watson. 1997. an evaluation of the effectiveness of environmental surrogates and modelling techniques in predicting the distribution of biological diversity. department of environment, sport and territories, canberra. gallant, j. c., and m. f. hutchinson. 1996. towards an understanding of landscape scale and structure in proceedings of the third international conference integrating gis and environmental modeling. university of california, santa barbara, national center for geographic information and analysis: cd-rom [cited 6 august 2003].12 geoscience australia. 2003. geodata topo-250k (series 1) topographic data [online]. geoscience australia, canberra. [cited 14 january 2004].13 gessler, p. e., i. d. moore, n. j. mckenzie, and p. j. ryan. 1995. soil-landscape modelling in southeastern australia in m. goodchild et al. (eds) gis and environmental modeling: progress and research issues. gis world books. gessler, p. e. and o. a. chadwick. 1997. quantitative soil-landscape modeling: a key to linking ecosystem processes on hillslopes. pedometrics 97. [cited 19 september 2004].14 gibson, w., c. daly, t. kittel, d. nychka, c. johns, n. rosenbloom, a. mcnab, and g. taylor. 2004. development of a 103-year high-resolution climate data set for the conterminous united states in proceedings 13th ams conference on applied climatology, amer. meteorological soc., portland, or, may 13-16, 2002. pp. 181183. [cited 26 may 2004].15 goovaerts p. 2000. geostatistical approaches for incorporating elevation into the spatial interpolation of rainfall. j. hydrol. 228:113-129. 9 http://www.ocs.orst.edu/pub/prism/docs/applclim04probalistic_spatial_approach-daly.pdf. 10 http://www.deh.gov.au/erin/ndvi/ndvi.html. 11 http://wwwscience.murdoch.edu.au/centres/others/pcb/toc/ pcb_contents_v6.html. 12 http://www.ncgia.ucsb.edu/conf/santa_fe_cdrom/sf_papers/gallant_john/paper.html. 13 http://www.ga.gov.au/nmd/products/digidat/250k.htm. 14 http://www.essc.psu.edu/pedometrics/abstracts/html/ gessler.html. 15 http://culter.colorado.edu/~kittel/pdf_gibson02_ams.pdf. hargreaves, g. h., and z. a. samani. 1982. estimating potential evapotranspiration. tech. note, j. irrig. and drain. engrg., asce. 108:225230. hartkamp, a. d., k. de beurs, a. stein, and j. w. white. 1999. interpolation techniques for climate variables. nrg-gis series 99-01. cimmyt, mexico, d.f. [cited 17 october 2003].16 hijmans, r., j., s. cameron, and j. para. 2004a. worldclim, version 1.2. a square kilometer resolution database of global terrestrial surface climate. [cited 26 may 2004].17 hijmans, r., j., s. cameron, and j. parra. 2004b. worldclim, a new high-resolution global climate database. abstract: inter-american workshop on environmental data access. campinas, brazil. mar 2004. [cited 26 may 2004].18 holdaway, m. r. 1996. spatial modelling and interpolation of monthly temperature using kriging, climate research 6:215-225. hutchinson, m. f. 1989. a new procedure for gridding elevation and stream line data with automatic removal of spurious pits. j. hydrol. 106:211-232. hutchinson, m. f. 1991. the application of thin plate smoothing splines to continent-wide data assimilation. pp. 104-113 in j. d. jasper (ed.), data assimilation systems. bureau of meteorology research report no. 27, bureau of meteorology, melbourne. hutchinson, m. f. 1995. interpolating mean rainfall using thin plate smoothing splines. int. j. gis 9:305-403. hutchinson, m. f. 1996. a locally adaptive approach to the interpolation of digital elevation models in ncgia (ed.), proceedings of the third international conference integrating gis and environmental modeling. santa fe, new mexico, 21-25 january, 1996. university of california, santa barbara, national center for geographic information and analysis: cd-rom [cited 16 october 2003].19 hutchinson, m. f. 2001. anusplin version 4.3. [online]. centre for resource and environmental studies, australian national university, canberra. [cited 6 august 2004].20 hutchinson, m. f. 2003a. topographic and climate database for africa. version 1.0. [online]. centre for resource and environmental studies, australian national university, canberra. [cited 13 january 2004].21 16 http://www.cimmyt.org/research/nrg/pdf/ nrggis%2099_01.pdf. 17 http://biogeo.berkeley.edu/. 18 http://www.cria.org.br/eventos/iaed/rhijmans_pre.html. 19 http://www.ncgia.ucsb.edu/conf/santa_fe_cdrom/sf_papers/hutchinson_michael_dem/local.html and 5 of the 10 models were classified as present. i set this threshold to a more conservative value (as compared to that for present day models) due to the coarse scale of the climate change modeling and the reduced number of environmental layers used. finally, under an assumption of no dispersal ability (peterson et al. 2001), which is likely the most appropriate assumption for the present study, which focuses on rare, endangered, or declining species, i intersected present and future model predictions to identify portions of the present-day range that will remain habitable. place prioritization analysis the areas identified by the enms as suitable under present and future conditions were used to identify concentrations of bird species richness, rarity, and highest threat. a heuristic approach (pressey et al. 1996) was used, seeking highest numbers of new species or of rare species that can be added to the system with each new area. although this approach can lead to decisions that do not lead to the globally optimal solution for representing species in a network of places, the simplicity of patterns treated in this study made more complex approaches unnecessary (church et al. 1996, pressey et al. 1996, csuti et al. 1997), and i thus could take advantage of the computational simplicity of heuristic algorithms (pressey et al. 1996). my heuristic complementarity approach was thus as follows (after peterson et al. 2000): areas holding highest numbers of species (i.e., sum of equally weighted distributional predictions for species), or of rarest species (i.e., sum of distributional predictions for species weighted by the multiplicative inverse of range size; williams et al. 1996) were identified. if multiple areas of equal richness were identified in the first step, the largest and most entire area was chosen; next, areas holding maximum numbers of species (or maximally rare suites of species) not represented in the initial area were added; the process was repeated until all species were included or until the remaining species did not overlap distributionally. “rarity” in this study thus refers to species of restricted range, and not to abundance of individuals. to take into account existing protected areas, i counted species as present in a given protected area if >10% of the protected area was papeş – bird distributions in central and eastern europe 19 figure 1. comparison between present-day (orange) and future-climate (black) predicted distributional areas for 6 bird species. because future-climate predictions are shown on top of present-day predictions, visible orange areas are predicted to be uninhabitable under future climate conditions. papeş – bird distributions in central and eastern europe 20 predicted present for the particular species, an approach that probably overstates the importance of the current reserve system to the conservation of threatened birds. results distributional predictions all enms (on which all subsequent analyses were based) were statistically significantly better than random predictions for all species (all p < 0.02). species’ distributions reconstructed using this procedure ranged from broadly distributed across cee (e.g., anser erythropus) to narrowly restricted to small parts of the region (e.g., falco eleonorae; fig. 1). likely climate change effects on species’ potential distributions, as predicted by my enms, were variable. one species (falco eleonorae) was predicted to lose all of its potential distributional area in cee, whereas others (e.g., caprimulgus europaeus) were predicted to see only minor losses (0.06%). indeed, predicted distributions postclimate change in cee decreased appreciably only for 6 species: alectoris graeca, anser erythropus, falco eleonorae, aquila clanga, accipiter brevipes, and emberiza melanocephala (fig. 1). such variable effects of climate change agree well with results of parallel studies in other regions (peterson 2003a; peterson 2003b; peterson et al. 2002); in general, previous studies also showed low levels of extinction of bird species in europe and mexico under both assumptions of dispersal and no dispersal possible (thomas el al. 2004). these findings are in contrast with the severe losses predicted for more diverse environments, such as tropical rainforest (williams et al. 2003). complementarity analysis because preliminary analyses indicated that 34 of the 36 species were represented in at least one protected area (fig. 2), i did not explore the trivial case of simple species representation. hence, i developed 4 separate place-prioritization analyses based on species richness versus rarity, and including distributions under future versus presentday climates. prioritizations based on present-day distributional models generally identified a single area that included almost all of the species. the species richness/present day prioritization identified a single area (fig. 3a) holding 34 species, with aegypius monachus, falco eleonorae, and caprimulgus europaeus only marginally represented (just a few pixels) in the area; no areas of overlap were found between hippolais olivetorum, phoenicurus phoenicurus, and the rest of the species. prioritization with weighted representation by rarity (present-day climates) was swayed by representation of the rarest species, hippolais olivetorum, so the first area represented only 25 species, and a second area added 8 more (fig. 3b); finally, the predicted distribution of aegypius monachus did not overlap with any of the first 2 areas. figure 2. frequency of representation of species in protected areas in central and eastern european countries. when climate change effects on species’ distributions were considered, effectively rendering the models into prioritizations of portions of species’ distributional areas likely to remain habitable over the next 50 yr, results were somewhat different, in that prioritizations required more areas to represent most species, suggesting that climate change processes may act to reduce distributional coincidence among species. hence, a prioritization based on maximizing species richness in future potential distributional areas identified a single area holding 33 of the 36 species papeş – bird distributions in central and eastern europe 21 figure 3. summary of results of place-prioritization analyses: (a) species richness in present-day climates, (b) rarity representation in present-day climates, (c) species richness under future-climate conditions, and (d) rarity representation under future-climate conditions. first priority areas are outlined in blue, and subsequent ones in purple; representation of species diversity or rarity are depicted as color ramps from white (none) to orange (high); sums of species not included in prioritization areas identified first are shown in white-to-green color ramps; also, in (a), the predicted distribution of falco eleonorae is shown in light green. (fig. 3c), although 2 (caprimulgus europaeus and aegypius monachus) were only represented marginally. maximizing representation of rarity among future-climate potential distributions (fig. 3d), a first area included 25 species and the second added 7 more. most species’ modeled geographic distributions overlapped at least one existing protected area in the cee region. indeed, the intersection of either present-day and futureclimate maps indicated that only hippolais olivetorum and falco eleonorae do not intersect any of the protected areas; all other species are expected to be represented in one or more protected areas (fig. 2). discussion the relative lack of detailed, modern, point-based distributional data represented the main limitation in the development of this project. observational data sets were extensive for western europe, but sources for cee were very few. a reliable source of occurrence localities is the large base of natural history museum specimens (peterson et al. 1998, papeş – bird distributions in central and eastern europe 22 figure 4. location of reserves (purple) in cee with respect to the distributional predictions of branta ruficollis (yellow), falco eleonorae (green), and hippolais olivetorum (red), and priority areas (blue) under present-day and future-climate conditions. papeş – bird distributions in central and eastern europe 23 collar et al. 2003, peterson at al. 2005); nonetheless, because little recent specimen acquisition (e.g., salvage for threatened species) has occurred, it proved difficult to assemble large suites of point-occurrence information. because of these limitations, full validation of models (e.g., peterson 2001) based on predictions into unsampled regions was not always possible, so some models may not prove to be as robust as would be desired. the enm approach used herein has been tested in numerous studies (e.g. peterson & cohoon 1999; peterson et al. 1999; peterson 2001; peterson & vieglais 2001; peterson et al. 2001; stockwell & peterson 2002; anderson et al. 2003), and has been shown to produce robust distributional predictions under most conditions. the predictions are based on models of ecological niches and as such--as with any model--are subject to the assumptions outlined earlier in this paper. it must be borne in mind that i focused on identifying potential areas where threatened or/and rare species may occur in cee. because models were based on occurrences across europe and western asia, i avoided problems with incomplete representation of species’ ecological potential as much as possible (pearson & dawson 2003). predicted distributions obtained from enm approaches offer several clear advantages over raw occurrence information when used in synthetic analyses such as the place-prioritization analyses herein (rojas-soto et al. 2003, sánchez-cordero et al. 2005). whereas raw occurrence data have biases of detectability and sampling effort, and may focus attention on historically surveyed areas that no longer hold populations of species, modeled distributions as input information for these models can overcome these biases to an impressive degree (soberón & peterson 2004). the cost, of course, is that any results from this modelbased prioritization exercise should be validated and supplemented by targeted field validation (pressey & cowling 2001). as such, this approach helps to compensate for the lack of comprehensive distributional data in regional conservation planning. in general, overlap between species’ modeled distributions and existing protected areas was extensive (fig. 2), suggesting that most species already see some degree of protection. exceptions were falco eleonorae and hippolais olivetorum, which were not predicted to be present in any reserve, and branta ruficollis, which was predicted to occur in only one. hence, if the existing reserve system were a reliable protector of species distributed in each reserve, much of the challenge would have been met. over half of the 36 species were predicted to be present in reserves in >9 countries. of course, given the need for on-ground model validations, and given variable levels of intersection between reserve boundaries and species’ distributions, actual protection afforded may be less. broadening the summary of protected areas i used, which omitted reserves in albania, bulgaria, bosnia and herzegovina, and croatia, would improve this picture somewhat. the results of this study identified focal areas for threatened birds in cee. all were situated along the lower danube river, in areas little covered by existing protected areas (fig. 4). basically, foci were identified that included the bulk of the species; special additional areas were necessary to include the problematic branta ruficollis, falco eleonorae, and hippolais olivetorum. judging by the results obtained, and particularly comparing potential distributions under present-day and future-climate conditions, the whole lower danube river basin is identified as ‘of special interest’ for a regional conservation scheme. targeted data collection in the areas identified herein would be key in validating model predictions for each species, and (more importantly) in verifying the importance of focal areas identified in this study. another important issue is that of land use changes that have occurred in the past couple of decades in this region due to important political and social changes. to my knowledge, a comprehensive analysis of these transformations is not yet available; however, other studies assessing the possible future land use change in western and central europe (rounsevell et al. 2006; verburg et al. 2006) based on climate change economic scenarios show large reduction in cropland and grassland areas due to extensive technological development. at a more regional (western and eastern carpathian mountains) and shorter time (last decade) scale, it was observed that both forest and cropland areas decreased while the non-forest natural vegetation and cropland/natural vegetation mosaic areas increased (dezso et al. 2005).these studies also identified land abandonment to be a papeş – bird distributions in central and eastern europe 24 common phenomenon. all these findings emphasize the need for conservation planning along the lower danube river. acknowledgments i thank a. townsend peterson for discussions and useful comments regarding this manuscript. gis-related questions were answered with the help of e. martinez meyer, r. houser, and m. ortegahuerta. museum specimen data were gathered from the american museum of natural history, burke museum, institut royal des sciences naturelles de belgique, national museum of natural history (lieden, the netherlands), national museum of natural history (washington, d.c.), national museum of scotland, natural history museum (london, u.k.), polish academy of sciences (museum and institute of zoology), swedish museum of natural history, and zoologie institut und museum (hamburg, germany). i am grateful to the following curators, collection managers and ornithologists who provided distributional data taken from museum specimens and bird surveys: c. bracker, r. dekker, r. faucett, g. frisk, d. georgiev, m. grell, p. kaòuch, g. lenglet, b. mcgowan, and l. vergeichyk. i thank m. janaus and d. stanescu for sharing distributional data gathered for their own research programs. i am also grateful to a. nyári for his constant encouragement and support. this study was partially funded by the university of kansas natural history museum panorama grant. references anderson, r. p., d. lew, and a. t. peterson. 2003. evaluating predictive models of species’ distributions: criteria for selecting optimal models. ecological modelling 162:211-232. austin, m. p., a. o. nicholls, and c. r. margules. 1990. measurement of the realized qualitative niche: environmental niches of five eucalyptus species. ecological monographs. 60:161-177. birdlife international. 2000. threatened birds of the world. lynx edicions and birdlife international, barcelona and cambridge. birdlife international/european bird census council. 2000. european bird populations: estimates and trends. cambridge. boston, t. and d. r. b. stockwell. 1994. interactive species distribution reporting, mapping and modeling using the world wide web. computer networks and isdn systems 28:231-238. brotherton, i. 1996. protected area theory at the system level. journal of environmental management 47:369-379. carpenter, g., a. n. gillison, and j. winter. 1993. domain: a flexible modeling procedure for mapping potential distributions of animals and plants. biodiversity conservation 2:667-680. chapman, a. d., m. e. s. muñoz, and i. koch. 2005. environmental information: placing biodiversity phenomena in an ecological and environmental context. biodiversity informatics 2:24-41. church, r. l., d. m. stoms, and f. w. davis. 1996. reserve selection as a maximal covering location problem. biological conservation 76:105-112. collar, n., c. fisher, and c. feare (eds). 2003. why museums matter: avian archives in an age of extinction. bulletin of the british ornithologists’ club supplement 123a:1-360. cramp, s. and k. e. l. simmons. 1977. handbook of the birds of europe, the middle east, and north africa: the birds of western palearctic. vol. 1: ostrich to ducks. oxford university press, new york. cramp, s. and k. e. l. simmons. 1982. handbook of the birds of europe, the middle east, and north africa: the birds of western palearctic. vol. 2: hawks to bustards. oxford university press, new york. cramp, s. and k. e. l. simmons. 1985a. handbook of the birds of europe, the middle east, and north africa: the birds of western palearctic. vol. 3: waders to gulls. oxford university press, new york. cramp, s. and k. e. l. simmons. 1985b. handbook of the birds of europe, the middle east, and north africa: the birds of western palearctic. vol. 4: terns to woodpeckers. oxford university press, new york. cramp, s. 1988. handbook of the birds of europe, the middle east, and north africa: the birds of western palearctic. vol. 5: tyrant flycatchers to thrushes. oxford university press, new york. cramp, s. and c. m. perrins. 1993. handbook of the birds of europe, the middle east, and north africa: the birds of western palearctic. vol. 7: flycatchers to shrikes. oxford university press, new york. csuti, b., s. polasky, p. h. williams, r. l. pressey, j. d. camm, m. kershaw, a. r. kiester, b. downs, r. hamilton, m. huso, and k. sahr. 1997. a comparison of reserve selection algorithm using data on terrestrial vertebrates in oregon. biological conservation 80:83-97. dezso, s., j. bartholy, r. pongracz, and z. barcza. 2006. analysis of land-use/land-cover change in the carpathian region based on remote sensing papeş – bird distributions in central and eastern europe 25 techniques. physics and chemistry of the earth 30:109-115. gambino, r. 2002. park policies: a european perspective. environments 30:1-14. grinnell, j. 1917: field tests of theories concerning distributional control. american naturalist 51:115128. hagemeijer, e. j. m. and m. j. blair. 1997. the ebcc atlas of european breeding birds: their distribution and abundance. t & a.d. poyser, london. haslett, j. r. 2002. european protected areas go upmarket. trends in ecology and evolution 17:541-542. houghton, j. t., y. ding, d. j. griggs, m. noguer, p. j. van der linden, x. dai, k. maskell, and c. a. johnson. 2001: climate change 2001: the scientific basis. contribution of working group i to the third assessment report of the intergovernmental panel on climate change. cambridge university press, cambridge. kelley, c., j. garson, a. aggarwal, and s. sarkar. 2002. place prioritization for biodiversity reserve network design: a comparison of the sites and resnet software packages for coverage and efficiency. diversity and distributions 8:297-306. midgley, g. f., l. hannah, d. millar, w. thuiller, and a. booth. 2003. developing regional and specieslevel assessments of climate change impacts on biodiversity in the cape floristic region. biological conservation 112:87-97. moore, j. l., m. folkmann, a. balmford, t. brooks, n. burgess, c. rahbek, p. h. williams, and j. krarup. 2003. heuristic and optimal solutions for setcovering problems in conservation biology. ecography 26:595-601. pearson, r. g. and t. p. dawson. 2003. predicting the impacts of climate change on the distribution of species: are bioclimate envelope models useful? global ecology and biogeography 12:361-371. peterson, a. t. 2001. predicting species’ geographic distributions based on ecological niche modeling. condor 103:599-605. peterson, a. t. 2003a: projected climate change effects on rocky mountain and great plains birds: generalities of biodiversity consequences. global change biology 9:647-655. peterson, a. t. 2003b: subtle recent distributional shifts in great plains bird species. southwestern naturalist 48:289-292. peterson, a. t. and k. p. cohoon. 1999. sensitivity of distributional prediction algorithms to geographic data completeness. ecological modelling 117:159164. peterson, a. t. and d. a. vieglais. 2001: predicting species invasions using ecological niche modeling: new approaches from bioinformatics attack a pressing problem. bioscience 51:363-371. peterson, a. t. and d. kluza. 2003. new distributional modeling approaches for gap analysis. animal conservation 6:47-54. peterson, a. t., a. g. navarro-sigüenza, and h. benítez-díaz. 1998. the need for continued scientific collecting: a geographic analysis of mexican bird specimens. ibis 140:288-294. peterson, a. t., j. soberón, and v. sánchez-cordero. 1999. conservatism of ecological niches in evolutionary time. science 285:1265-1267. peterson, a. t., c. cicero, and j. wieczorek. 2006. free and open access to distributed data from bird specimens: why? auk 122:987-990. peterson, a. t., s. l. egbert, v.sánchez-cordero, and k. price. 2000. geographic analysis of conservation priority: endemic birds and mammals in veracruz, mexico. biological conservation 93:85-94. peterson, a. t., v. sánchez-cordero, j. soberón, j. bartley, r. h. buddemeier, and a. g. navarrosigüenza. 2001. effects of global climate change on geographic distributions of mexican cracidae. ecological modelling 144:21-30. peterson, a. t., m. a. ortega-huerta, j. bartley, v. sánchez-cordero, j. soberón, r. h. buddemeier, and d. r. b. stockwell. 2002. future projections for mexican faunas under global climate change scenarios. nature 416:626-629. pimm, s. l., g. j. russell, j. l. gittleman, and t. m. brooks. 1995. the future of biodiversity. science 269:347-350. prendergast, j. r., r. m. quinn, and j. h. lawton. 1999. the gaps between theory and practice in selecting nature reserves. conservation biology 13:484-492. pressey, r. l. 1994. ad hoc reservations: forward or backward steps in developing representative reserve systems? conservation biology 8:662-668. pressey, r. l., h. p. possingham, and c. r. margules. 1996. optimality in reserve selection algorithms: when does it matter and how much? biological conservation 76:259-267. pressey, r. l. and r. m. cowling. 2001. reserve selection algorithms and the real world. conservation biology 15:275-277. raxworthy, c. j., e. martínez-meyer, n. horning, r. a. nussbaum, g. e. schneider, m. a. ortegahuerta, and a. t. peterson. 2003. predicting distributions of known and unknown reptile species in madagascar. nature 426:837-841. rojas-soto, o. r., o. alcantara-ayala, and a. g navarro. 2003. regionalization of the avifauna of the baja california peninsula, mexico: a parsimony analysis of endemicity and distributional papeş – bird distributions in central and eastern europe 26 modelling approach. journal of biogeography 30:449-461. rounsevell, m. d. a., i. reginster, m. b. araújo, t. r. carter, n. dendoncker, f. ewert, j. i. house, s. kankaanpää, r. leemans, m. j. metzger, c. schmit, p. smith, and g. tuck. 2006. a coherent set of future land use change scenarios for europe. agriculture, ecosystems and environment 114:5768. sánchez-cordero, v., v. cirelli, m. munguia, and s. sarkar. 2005. place prioritization for biodiversity representation using species' ecological niche modeling. biodiversity informatics. 2:11-23. soberón, j. and a. t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. philosophical transactions of the royal society of london series b, biological sciences 359:689-698. stockwell, d. r. b. and d. peters. 1999. the garp modeling system: problems and solutions to automated spatial prediction. international journal of geographical information science 13:143-158. stockwell, d. r. b. and a. t. peterson. 2002. effects of sample size on accuracy of species’ distribution models. ecological modelling 148:1-13. thomas, c.d., a. cameron, r. e. green, m. bakkenes, l. j. beaumont, y. c. collingham, f. n. erasmus, m. ferreira de siqueira, a. grainger, l. hannah, l. hughes, b. huntley, a. s. van jaarsveld, g. f. midgley, l. miles, m. a. ortega-huerta, a. t. peterson, o. l. phillips, and s. e. williams. 2004. extinction risk from climate change. nature 427:145-148. tucker, g. m. and m. i. evans. 1997. habitats for birds in europe: a conservation strategy for the wider environment. birdlife international, cambridge. tucker, g. m. and m. f. heath. 1994. birds in europe: their conservation status. birdllife international, cambridge. unep-wcmc. 2002. protected areas and other associated areas of the world. version 5.0. unep, geneva. verburg, p.h., c. j. e. schulp, n. witte, and a. veldkamp. 2006. downscalling of land use change scenarios to assess the dynamics of european landscapes. agriculture, ecosystems and environment 114:39-56. walker, p. a. and k. d. cocks. 1991. habitat: a procedure for modelling a disjoint environmental envelope for a plant or animal species. global ecology and biogeography letters 1:108-118. williams, p., d. gibbons, c. margules, a. rebelo, c. humphries, and r. pressey. 1996. a comparison of richness hotspots, rarity hotspots, and complementary areas for conserving diversity of british birds. conservation biology 10:155-174. williams, s. e., e. e. bolitho, and s. fox. 2003. climate change in australian tropical rainforests: an impeding environmental catastrophe. proceedings of the royal society of london series b, biological sciences 270:1887-1892. a simple ms excel mapper for the quality control and publication of location data points biodiversity informatics, 7, no 3, 2010, pp. 137 – 142. a spreadsheet mapping approach for error checking and sharing collection point data desmond h. foley division of entomology, walter reed army institute of research, 503 robert grant avenue, silver spring, md, 20746, usa, foleydes@si.edu abstract. – the ready availability of online maps of plant and animal collection locations has drawn attention to the need for georeference accuracy. many obvious georeference errors, for example, that map land animals over sea, the wrong hemisphere, or the wrong country, may be avoided if collectors and data providers could easily map their data points prior to publication. various tools are available for quality control of georeference data, but many involve an investment of time to learn the software involved. this paper presents a method for the rapid map display of longitude and latitude data using the chart function in microsoft office excel®, arguably the most ubiquitous spreadsheet software. advantages of this method include: use of software that will be familiar to many data providers, immediate visual feedback to assess data point accuracy; and results that can be easily shared with others. methods for making custom excel chart maps are given, and we provide free charts for the world and a selection of countries at http://www.vectormap.org/resources.htm. key words. – map, excel; chart; scatter plot; georeference; collection data the past decade has seen an increasing number of online databases designed to host collection records of a wide variety of organisms (e.g. global biodiversity information facility1). these databases require georeferenced data, to map and compare collection points to other spatial information, sometimes in a geographical information system (gis) setting (e.g.2). guidelines for reporting georeference information for collection records have been developed (e.g. chapman and wieczorek 2006, manis/herpnet/ornis3). georeference errors can occur, for example, when converting degrees minutes seconds to a decimal degree format, omitting the negative sign for the hemisphere, or typographical errors such as omitting, repeating, or transposing numbers. various methods are available to automate the process of detecting errors in georeference information, such as using diva-gis4, geolocate5, and tools available through gbif6. for the former method, longitude and latitude point data are converted to shapefiles 1 http://www.gbif.org/ 2 http://www.mosquitomap.org/ 3 http://manisnet.org/ 4 http://www.diva-gis.org/ 5 http://www.museum.tulane.edu/geolocate/ 6 http://sourceforge.net/projects/gbif/files/data%20tester/ and data cleaning is then undertaken by the ‘check coordinates’ option of diva-gis, a “point-inpolygon” method (chapman, 2005), which identifies points located outside all polygons (i.e. over the sea), and points that do not match relations for an administrative region (e.g. fall in another country, or state). however, this process can be labor-intensive and requires a certain familiarity with the software packages involved. some databases (e.g. gbif) screen data that is submitted to them, and will report back to the data provider a list of questionable georeferences before data is accepted. however, no such process exists, for example, for individuals who want to send location data as part of a journal article, or gene sequence metadata to genbank, resulting in the frequent publication of imprecise or erroneous collection location details. if collectors and data providers could easily produce an orientation map for error checking, many georeference errors could be avoided at an early stage, when corrections tend to be easiest. this paper presents a simple method for the rapid map display of longitude and latitude data using the chart function in microsoft office excel®, arguably the most ubiquitous spreadsheet software. mailto:foleydes@si.edu http://www.vectormap.org/resources.htm http://www.gbif.org/ http://www.mosquitomap.org/ http://manisnet.org/ http://www.diva-gis.org/ http://www.museum.tulane.edu/geolocate/ http://sourceforge.net/projects/gbif/files/data%20tester/ foley – spreadsheet mapping of collection point data figure 1. microsoft excel chart (xy scatter plot) with background picture of the world and test locations in costa rica shown as red points. note the erroneous points over sea; these errors can easily be visually detected and corrected. methods and results converting the curved surface of the earth to a flat map requires a projection, and the geographic projection (also called the equidirectional, equidistant cylindrical, simple cylindrical, equirectangular, plate carrée or carte parallelogrammatique projection) maps longitude to equally spaced vertical straight lines, and latitude to evenly spread horizontal straight lines. data defined on a geographic coordinate system is displayed as if a degree is a linear unit of measure. latitude values are measured relative to the equator and range from –90° at the south pole to +90° at the north pole. longitude values are measured relative to the prime meridian at greenwich, united kingdom, and ranges from – 180° when traveling west to 180° when traveling east. for example, argentina, which is south of the equator and west of greenwich, has negative longitude values and negative latitude values. the geographic projection is the standard for computer applications that process global maps, because of the simple relationship between the position of an image pixel on the map and its corresponding geographic location on earth. at low resolution, such as at the level of countries, maps based on the earth as a sphere (e.g. google earth) display the lines of longitude and latitude as curvilinear rather than linear. however, at high resolution, these map lines approach linearity, and are therefore suitable for our purposes (see below). the excel chart option ‘scatter’ allows the graphical display of cartesian coordinates (xy) of 138 foley – spreadsheet mapping of collection point data figure 2. microsoft excel chart (xy scatter plot) with background picture from mosquitomap and test locations shown as red points. any parameter, including longitude (x-axis) and latitude (y-axis). if the chart contains a background image of a map in a geographic projection, cropped to known longitude and latitude, and the axes of the chart are constrained to these longitudes and latitudes, then geographic correspondence between xy data and geographic location is established. the resulting thematic map can be used for basic orientation, and for identifying some georeference errors, by visual comparison of data point locations with coast and political boundaries. by mousing over points, information about these coordinates is displayed in excel, allowing identification of particular coordinates that need checking. in the following i describe a method and provide examples for obtaining excel chart maps from a shapefile, and from online maps. possible, the gadm database of global creating maps from a shapefile while various options for obtaining maps are administrative areas7 provides free country administrative region shapefiles. these shapefiles can be loaded in a gis program such as the freely available diva-gis or the commercially available arcgis®. once loaded into a program capable of displaying a shapefile, an image file of the country or area of interest needs to be constructed, and this image is then cropped to the area of interest in a program such as ms paint or adobe photoshop. a bounding box or corner points can be drawn in the gis program to demarcate the area of interest, and the image exported as a file such as jpeg. for ease of presentation in excel, we define the bounding box or corner points at positions that are easy to display. for example, a bounding box from 1012.5 degrees latitude and 35-40 degrees longitude displays more easily in an excel chart than one from 10.12445-12.45773 degrees latitude and 34.55689-39.66554 degrees longitude. we used the tool ‘draw point’ in arcview gis 3.3, and 7 http://www.gadm.org/country/ 139 http://www.gadm.org/country/ foley – spreadsheet mapping of collection point data then edited the shape properties (right click) to achieve the desired longitude and latitude. an upper left and lower right point was drawn to demarcate the area of interest. graphics are selected, then graphics properties edited to make points small enough for accuracy and large enough to be visible (e.g. size 2). the zoom tool was then used to frame the view to just encompass the 2 points. this view was then exported as a jpeg format image file (200 dpi; 100% quality). figure 1 shows an excel scatter plot with a background picture of the world and test locations in costa rica shown as red dots. note the erroneous points appearing over sea, which can be visually detected and corrected. creating maps from an online map we used mosquitomap and google earth to create the custom maps in figure 2 and 3, respectively. in the mosquitomap viewer, we zoomed to the area of interest and demarcate this with tools>markup>create new item>rectangle. the settings were for a fine solid line, with no fill (line width = 1, line transparency = 0, fill transparency = 100%). the corner overview map was hidden and a rectangle was drawn between longitude and latitudes that are simple to chart. for example, figure 2 was derived from a rectangle with an upper left corner of latitude = 0 degrees, longitude = –64 degrees; and a lower right corner of latitude = –5 degrees, longitude = –59 degrees. the map was then exported as an image file. the image was cropped according to the rectangle in adobe photoshop, and then loaded into a chart as a background image, and the axis’ minimum and maximum adjusted to match the latitude and longitude. in the case of figure 2, the excel chart y-axis minimum was set to -5, maximum set to 0, and crossing the x-axis at –5; the x-axis was set to minimum –64, maximum –59, and crossing the yaxis at –64. the major unit was set to 1. to create excel charts using google earth images as backgrounds, the location in google earth must first be set to decimal degrees (set tools>options>3d view>show lat/long to decimal degrees), and the area of interest must be at such a resolution that the lines of longitude and latitude are linear. these lines can be visualized by turning on grid lines (view>grid). a new placemark is activated for the upper left corner of the area of interest. in the new placemark dialog figure 3. microsoft excel chart (xy scatter plot) with background picture from google earth and test locations shown as red points. 140 foley – spreadsheet mapping of collection point data box, the latitude and longitude can be adjusted to remove unnecessary precision and simplify the coordinates for charting, for example, the corners of figure 3 were limited to three decimal places. within the new placemark dialog box, under the view tab, the heading and tilt are changed to zero. the placemark icon in the dialog box is clicked to change the symbology to a radio button (circle), of scale 0.5, and the name of the placeholder is removed. this ensures that the placeholder is unobtrusive in the final image. this process is repeated for the lower right corner of the area of interest. under view, the grid lines, overview map, and sidebar are removed, and the area of interest positioned to allow an image to be taken that is clear of text. the image is then taken, either by edit>copy image or file>save>save image, and the image file cropped and loaded into an excel chart, as detailed above. as a final step, locations in the excel chart can be compared to the same coordinates in google earth to test accuracy. once a map is constructed in excel, there is no setting that forces equally spaced horizontal and vertical gridlines and axis ticks. an excel vba project is available for making square gridlines8, but the simplest solution is to use a ruler and manually adjust the chart dimensions on the screen until the units are equal. although points will appear at their correct longitude and latitude, printing from excel may distort the map, unless the margins in print preview are adjusted so that the degrees longitude and latitude are the same size. alternatively, the chart can be copied and pasted into another program, such as photoshop or ms powerpoint, with no distortion. discussion ideally, once a suitable map is located, the area of interest in longitude and latitude would be defined by the data provider, and an image file ready for excel would be produced in one action. screenshots can be obtained in ms windows xp and images pasted in ms paint, then cropped and saved as image files. windows vista and windows 7 has the screenshot function ‘snipping tool’, which in the latter program can be found by selecting start menu>all programs>accessories. freeware tools are available that have enhanced 8 http://peltiertech.com/excel/charts/squaregrid.html screenshot functionality9, as well as various commercial programs. with mac os x, a snapshot of a selected area can be created as a file on the desktop by command-shift-4, then dragging and dropping the mouse over the selected area. some of these screenshot options may replace reporting of the longitude and latitude of the underlying map application with screen coordinates of the mouse position. although, we concentrate on the charting function of excel because of the relative ubiquity of this program, enhanced mapping functionality for pdfs is now available in adobe reader 9 and adobe acrobat 9. as these and other product develop and become more widely available they may play a greater role in the cleaning and display of collection data. one limitation of the excel chart approach is that erroneous points that map beyond the boundaries of the map will not display on the chart, and could go undetected. the use of country and global excel chart maps allow the detection of some but not all of these possible georeference errors. although the size of charts can be expanded in excel, zoom capability is limited, so points that fall just off the coast or in the wrong country may not be detected. the excel charting step should not be seen as a replacement for other more detailed georeference error detection methods. there are alternative and more developed tools for data cleaning, like the previously mentioned diva-gis, and these will be indispensable when attempting to increase data quality for large collections. during the writing of this paper i became aware that others have promoted mapping using the chart function of excel, both for free (e.g.10, 11,12), and as part of software packages (e.g.13, 14, 15). however, i believe that this paper is the first to 9 http://www.inspire-soft.net/?nav=soft_screenshoter 10 http://www.excelhero.com/blog/2010/04/excel-locationmapping.html 11 http://forums.techguy.org/business-applications/917391-plot-latlon-points-world.html 12 http://www.excelcharts.com/blog/how-to-create-thematic-mapexcel/ 13 http://www.excelhero.com/blog/2010/04/excel-locationmapping.html 14 http://www.mapcrossing.com/ 15 http://office.microsoft.com/en-us/excel-help/create-a-map-withmappoint-and-excel-ha010277101.aspx 141 http://peltiertech.com/excel/charts/squaregrid.html http://www.inspire-soft.net/?nav=soft_screenshoter http://www.excelhero.com/blog/2010/04/excel-location-mapping.html http://www.excelhero.com/blog/2010/04/excel-location-mapping.html http://forums.techguy.org/business-applications/917391-plot-lat-lon-points-world.html http://forums.techguy.org/business-applications/917391-plot-lat-lon-points-world.html http://www.excelcharts.com/blog/how-to-create-thematic-map-excel/ http://www.excelcharts.com/blog/how-to-create-thematic-map-excel/ http://www.excelhero.com/blog/2010/04/excel-location-mapping.html http://www.excelhero.com/blog/2010/04/excel-location-mapping.html http://www.mapcrossing.com/ http://office.microsoft.com/en-us/excel-help/create-a-map-with-mappoint-and-excel-ha010277101.aspx http://office.microsoft.com/en-us/excel-help/create-a-map-with-mappoint-and-excel-ha010277101.aspx foley – spreadsheet mapping of collection point data 142 describe in detail the excel charting method specifically for those interested in error checking and sharing specimen location data. in addition, we provide excel charts for a variety of countries for free download (see16). users can cut and paste their collection coordinates into these files and immediately see global and country maps of their data. we have incorporated excel charts for mapping collection points in our spreadsheet collection forms for arthropod disease vectors (e.g.17), and we encourage other groups to include a mapping function in their collection spreadsheets. providing these simple tools is done in the hope that this will encourage mapping as a routine step before reporting geolocation data, and increase awareness of the importance of geolocation data quality. acknowledgements funding for this work was provided by the global emerging infections surveillance and response system, a division of the armed forces 16 http://www.vectormap.org/resources.htm 17 http://www.mosquitomap.org/contribute.htm health surveillance center. thanks go to l. m. rueda, j. f. ruiz, and r. c. wilkerson for comments and review of this manuscript. this research was performed under a memorandum of understanding between the walter reed army institute of research and the smithsonian institution, with institutional support provided by both organizations. the opinions and assertions contained herein are those of the authors and are not to be construed as official or reflecting the views of the department of the army or the department of defense. literature cited chapman, a.d. (2005) principles and methods of data cleaning – primary species and speciesoccurrence data, version 1.0. report for the global biodiversity information facility, copenhagen. chapman, a.d. and j. wieczorek (eds), 2006. guide to best practices for georeferencing. copenhagen: global biodiversity information facility. (accessed 2010-01-15) http://www.vectormap.org/resources.htm http://www.mosquitomap.org/contribute.htm methods and results creating maps from a shapefile creating maps from an online map discussion acknowledgements literature cited biodiversity informatics, 7, 2010, pp. 93 – 112 natural history specimen digitization: challenges and concerns ana vollmar (formerly 1), james a. macklin (1), linda s. ford (2) (1) harvard university herbaria, 22 divinity avenue, cambridge, ma 01238 (2) harvard university museum of comparative zoology, cambridge, ma 01238 correspondence e-mail: jmacklin@oeb.harvard.edu abstract. – a survey on the challenges and concerns involved with digitizing natural history specimens was circulated to curators, collections managers, and administrators in the natural history community in the spring of 2009, with over 200 responses received. the overwhelming barrier to digitizing collections was a lack of funding or issues directly related to funding, leaving institutions mostly responsible for providing the necessary support. the uneven digitization landscape leads to a patchy accumulation of records at varying qualities, and based on different priorities, ultimately influencing the data's fitness for use. the survey results also indicated that although the kind of specimens found in collections and their storage can be quite variable, there are many similar challenges across disciplines when digitizing including imaging, automated text scanning and parsing, geo-referencing, etc. thus, better communication between domains could foster knowledge on digitization leading to efficiencies that could be disseminated through documentation of best practices and training. key words. – natural history collections, survey, collections, specimens, specimen data, metadata, digitization, gbif, biodiversity research. introduction natural history collections are recognized as keepers of the primary information for the flora and fauna for both the present-day and historical record of the planet. a projected three billion or more specimens are estimated to be held in the biological and paleobiological collections of the world (butler et al. 1998; lane, 1999). these specimens have associated core data that is recognized as fundamental to discipline-specific research, as well as broader global issues such as invasive species, ecological/conservation issues, climate change, and emerging diseases (araújo et al., 2005; loarie et al., 2008; peterson and vieglais, 2001; pinto, 2010; saurez and tsutsui, 2004; shaffer et al., 1998; winker, 2004). currently, even with demand expanding rapidly beyond traditional taxonomic/systematic research to include other academic researchers, ngos, resource managers, governmental agencies, and citizen scientists, most of the world’s holdings has not been digitized and/or made available on-line. of the potential three billion specimens, only a small fraction have been digitized, which is evidenced by the approximately 50 million specimen records currently available through the gbif portal (gbif, 2010). in order to begin prioritizing data capture at a national or global level it is imperative to understand the current state of affairs of digitization initiatives at institutions housing collections. there have been other recent surveys taken on the state of natural history collections but none have specifically focused on digitization challenges (national science and technology council, 2009; synthesis, 2010). a workshop on identifying and addressing digitization bottlenecks was held at harvard university in 2006, which brought together experts in a diverse array of natural history collections, biodiversity informaticians, and various stakeholders. the conclusions and recommendations of this report helped to focus on both the need for the survey and the questions involved (beaman et al., 2007). the purpose of this survey was to assess what types of digitization work were ongoing in collections, to get a sense of how resources were handled and distributed for digitization projects, and to identify the nature of impediments to digitization work and how they have been addressed in diverse collections. the data are intended to provide practical information to vollmar et al. –specimen digitization: challenges and concerns curators, collection managers, and administrators at all levels about the challenges presented by digitization, and the ways in which they have been addressed, or not, in diverse collections. methods from may 20 to june 20, 2009, the survey “natural history specimen digitization: challenges and concerns” was circulated throughout the global natural history community to curators, collections managers, and administrators. the survey was initiated by the global biodiversity information facility (gbif) global strategy and action plan for digitization of natural history collections (gsap-nhc) task group, in collaboration with the national science foundation research coordination network, collectionsweb (www.collectionsweb.org), and the society for the preservation of natural history collections (www.spnhc.org). in this paper, the findings from the survey will be discussed in detail, beginning with an overview of the respondents. the second section turns to broad trends emerging from open-ended questions in the survey regarding barriers to digitization, along with specific ways in which those barriers have been addressed. it discusses issues associated with having collections data available on-line, including how various collections have addressed the difficulty of ensuring feedback about their on-line data. the third section examines how funding has been distributed in collections to support digitization projects and associated staff, and presents a comprehensive list of funding sources survey respondents actively use to support digitization in their collections (see appendix 1). section four tackles the technological logistics of digitization, examining data entry and imaging. it also outlines the specifics of a range of equipment employed in various digitization projects and associated costs. finally, drawing on survey and complementary interview material, the paper discusses specific digitization issues faced by different disciplines. interviews were conducted from march to may, 2009 with staff at harvard university museum of comparative zoology (mcz), harvard university herbarium (huh), and yale university, the natural history museum (nhm), uk, university of navarra, spain, university of kwazulu-natal, south africa, and the canadian museum of nature (cmn). the raw survey data is available in three files (gsap_questions.csv; gsap_var.csv; gsap_db_anon.csv) linked to on the gsapnhc home page on gbif's website ( http://www.gbif.org/informatics/primarydata/task-groups/gsap-nhc/) section 1: overview of survey respondents in total, 201 respondents completed the survey. reflecting a heavy bias in the geographic distribution of the survey, 62% were from north america, 22% from europe, 9% from south and central america, 3% from asia, 3% from the australasian region, 1% from the middle east, and none from africa. more than half (62%) of respondents were from institutions/museums affiliated with a university, 23% were from national or government-affiliated museums/institutions, and 8% were from freestanding museums. other respondents had private collections (2%) or were affiliated with non-profit, non-governmental organizations that maintained collections (2%). collections represented ranged greatly in size, from just 17 specimens up to 21 million specimens. the median collection size was 200,000 specimens. when asked to describe their position within the collections where they worked, nearly half of respondents selected more than one title to characterize the work they did, suggesting that respondents often played many roles within their collections and did diverse kinds of work. respondents who identified themselves as curators or assistant curators made up 47%, 39% were collections managers or collections assistants, 28% identified as researchers, and 18% were data managers, database technicians, or biodiversity informaticians. another 12% had more administrative positions as directors or assistant directors of museums/institutions, and as program directors. as with the diversity of roles respondents played in their collections, 43% of respondents worked with collections in multiple disciplines; 55% of respondents worked with botany collections; 26% with entomology collections; 18% with invertebrate zoology collections; 1416% with each of herpetology, ichthyology, mammalogy, mycology, and ornithology; and 10% 94 http://www.collectionsweb.org/ http://www.spnhc.org/ http://www.gbif.org/informatics/primary-data/task-groups/gsap-nhc/ http://www.gbif.org/informatics/primary-data/task-groups/gsap-nhc/ vollmar et al. –specimen digitization: challenges and concerns with each of vertebrate and invertebrate paleontology. other disciplines represented by survey respondents include geology (7%), mineralogy, anthropology, and archaeology (5% each), and ethnobotany, paleobotany, historical scientific instruments, and scientific photograph archives (less than 1% each). since almost half of the respondents worked either in multiple collections or collections encompassing multiple disciplines, drawing correlations between particular disciplines and digitization barriers, practices, or the like will not always be straightforward. section 2: barriers to digitization overall 94% of respondents reported digitization and/or imaging ongoing in their collections in the past two years, but the results also highlight impediments. when asked whether or not digitization work was currently ongoing in their collections, only 79% of respondents answered affirmatively, the remainder citing largely similar reasons for why digitization work was at a standstill. respondents (n=44) ranked the reasons why digitization was not ongoing in a collection; funding was the primary reason, followed in descending order of importance by time, staff, lack of institutional support, infrastructure/technology, and curation practices. when respondents were sorted into categories based on job responsibilities, this ranking remained largely stable. similarly, the least important reasons among respondents also were consistent, and included issues of data sharing, lack of collecting permits, sensitive species information, and indigenous rights (see table 1). table 1. what are the primary reasons digitization work is not ongoing in your collection? ranking is on a scale of 1 (most important) to 5 (least important). all respondents (n=44) answering for institution (n=7) answering for specific collection (n=28) directors (n=7) curators/collection managers (n=39) funding (1.3) staff (1.2) funding (1.1) funding (1.0) funding (1.3) time (1.7) funding (1.3) time (1.5) staff (1.2) staff (1.6) staff (1.7) lack of institutional support (1.8) staff (1.9) time (1.3) time (1.6) lack of institutional support (2.1) time (2.2) lack of institutional support (2.5) lack of institutional support (1.7) infrastructure/ technology (2.3) infrastructure/ technology (2.4) curation practices (2.4) infrastructure/ technology (2.5) curation practices (3.0) project complete (2.3) data sharing (3.9) data sharing (3.6) data sharing (4.2) project complete (4.0) data sharing (3.7) collecting permits (4.2) collecting permits (3.6) collecting permits (4.2) storage practices (4.3) sensitive species data (4.0) sensitive species data (4.2) sensitive species data (3.6) sensitive species data (4.4) collecting practices (4.3) collecting permits (4.1) indigenous rights (4.4) indigenous rights (4.8) indigenous rights (4.5) indigenous rights (5.0) indigenous rights (4.5) 95 vollmar et al. –specimen digitization: challenges and concerns among 171 respondents from collections in which digitization was either ongoing or had occurred within the past two years, the major impediments to digitization identified were similar to the reasons named by other respondents as to why digitization was not ongoing. funding, time, and staff were consistently the top three challenges faced, followed by issues of data entry, data quality, and georeferencing. these latter three in particular suggest a need for accessible guidelines and suggestions about how to structure digitization projects. collecting practices, sensitive species information, collecting permits, and indigenous rights were the least important barriers to digitizing collections (see table 2). their status as “least important” suggest these integral issues for collections in general take a backseat to more immediate logistical constraints and concerns of digitization. in an open-ended question about the impediments to digitization work, respondents elaborated on their initial rankings, giving more detailed descriptions of these identified barriers facing their digitization projects. funding other survey respondents pointed to budgetary issues and the myriad of problems related to the lack of funding. one respondent from israel described this succinctly: “[a lack of funding] did not allow good professional programmers. the process was long and not successful [so] we returned to excel and access data entry. [further, the] budget enables student work only [and] they table 2. if digitization is ongoing in your collections or has happened within the past two years, in your experience have any of the following issues impeded digitization work in your collection? ranking is on a scale of 1 (most important) to 5 (least important). all respondents (n=171) answering for institution (n=21) answering for specific collection (n=104) directors (n=33) curators/ collection managers (n=132) data managers (n=32) funding (1.5) funding (1.5) time (1.6) time (1.4) funding (1.5) funding (1.5) time (1.6) staff (1.7) funding (1.6) funding (1.6) time (1.5) time (1.7) staff (1.8) time (1.9) staff (1.8) staff (1.7) staff (1.7) staff (1.8) data entry (2.5) georeferencing (2.2) data quality (2.6) data entry (2.0) data entry (2.6) infrastructure/ technology (2.3) data quality (2.6) data quality (2.4) data entry (2.6) georeferencing (2.2) infrastructure/ technology (2.6) georeferencing (2.5) collecting practices (3.8) storage practices (3.3) collecting practices (3.9) curation practices (3.8) collecting practices (4.0) collecting practices (3.7) sensitive species data (3.9) indigenous rights (3.8) sensitive species data (4.0) collecting practices (3.9) sensitive species data (4.0) storage practices (3.9) collecting permits (4.2) collecting permits (3.8) collecting permits (4.3) collecting permits (4.1) collecting permits (4.3) collecting permits (4.1) indigenous rights (4.4) collecting practices (4.1) indigenous rights (4.5) indigenous rights (4.1) indigenous rights (4.4) indigenous rights (4.3) 96 vollmar et al. –specimen digitization: challenges and concerns tend to leave after a while. teaching them each time the job is very time consuming. they need closer supervising and [the] data requires more attention for quality.” respondents also identified obtaining funding sources as barriers to digitization, writing that “digitisation [is] seen as impossible to apply [for] funding” (denmark) and that “it is quite hard to find funds that pay just for the digitization of natural history collections” (netherlands). in an effort to get around funding crunches and incorporate digitization into daily curatorial tasks, one respondent described a system in which “data entry and imaging are now project-based, so that herbarium data and loan requests are answered simultaneously with the advancement of digitization. in other words, digitization has become an opportunistic activity that piggybacks on more pressing day-to-day herbarium service to clients” (canada). funding will be further addressed in section 3. staff closely related to funding, if not inseparable, is the importance of staff to digitization efforts, particularly well-trained staff. as respondents emphasized repeatedly in their responses to the survey, people are key to successful digitization projects. a respondent from canada wrote, “staff is the greatest limitation for digitizing the collection. there is a steep learning curve for accuracy and speed of data entry.” other respondents alluded to social difficulties, writing that impediments included “attitudes (digitization being seen as scientifically unproductive)” (denmark), and “cultural entrenchment” (usa). getting staff (at all position levels) on board with digitization projects was seen as a key element, and this was accomplished by “training and helping the head curators and staff understand why certain methods are in place…and moving everyone (including the ‘old guard’) forward” (usa). respondents also mentioned changes such as “hiring a curator concerned about digitization” (usa) and making a “change in the management view to [the] value of digitization” (malaysia) as important in increasing success. still others cited the importance of “workflow” or “staff organization and work planning” as significant to ensuring efficient digitization processes. for instance, thoughtful placement of disciplinary experts and non-experts throughout the process can be exceedingly helpful, as non-experts who know little about a given discipline can frequently do the first level of data entry, marking sites of doubt and uncertainty for experts to look over at a later time. in many contexts, data entry staffing can prove a huge bottleneck in the digitization process but, even this fundamental step, may be conditional. michelle hamer (pers. comm. 2009) of the south africa national biodiversity institute and the south africa gbif node made note of the low volume of applications for digitization funding coming to sagbif. she identified this particular bottleneck for digitization not as a lack of funding or available staff, but rather that receiving funding would necessitate fulfilling the terms of a grant, and this was rendered difficult in her region given significant barriers in terms of a lack of resources even for basic curation needs. to get a digitization project off the ground, involved staff must be well trained in the use and possibilities of the data capture client and/or whatever other technologies are being used in the digitization process. training requires time and money, both of which are limiting factors along with the question of who, in a given institution, is able and available to train staff. as indicated in interviews and by survey responses, this person is often the data manager or biodiversity informatician, which can divert effort from other necessary informatics work. creating local tutorials that explain, for example, the purpose of each field in a database and “data dictionaries,” which address common uncertainties encountered during data entry, can be very helpful in dispersing the training load and reducing the number of input errors. finally, databases in which staff must log in to modify or create new entries allow collection managers to track how long certain tasks take and how efficiently particular people are working. digitization of specimens, thus, not only expands possibilities for managing collections, but also the possibilities for managing the people who work in collections. because digitization of collections expands the access to collections, there are often increases in the number of visitors and more requests for information and loans. this can become an issue when a collection is understaffed, and highlights the importance of good planning in any digitization project; departments and 97 vollmar et al. –specimen digitization: challenges and concerns institutions must be aware and plan for the changes in the collections’ use once databases are available to a broader forum. space and curation respondents reporting on their institutions at large emphasized issues of physical space and money. on the difficulties presented by physical spaces, a respondent from austria cited as a hurdle the “size of collections [and the] complex structure of the very heterogeneously historically grown collections dating back to the beginning of the 19th century.” while many older institutions face this same challenge and are engaged in renovating buildings and internal spaces, the spatial organization of any collection impacts not only ease of use but also the process of digitizing collection data. collections are frequently organized by mixtures of taxonomic and geographic hierarchies, while at the same time also shaped by the constraints of particular physical spaces, by the ways in which people like to use collections, and by the storage practices demanded by the nature of the specimens themselves. ideally, the spatial and curation needs of a collection can to some extent be assisted during the digitization process. this can happen in a variety of ways, with parts of collections being reorganized and updated or, as practiced in an ornithology collection in the usa, collections are closely monitored for pests as part of digitizing specimens from their labels. curation of collections involves navigating the space in which a collection physically resides and, frequently, the present-day management is dependent on its historical curation. current work is always constrained by past decisions about how to gather and record information, and the spatial decisions of how to store and organize specimens. in collections that have had many curators with different practices, specialties, and goals, the levels of curation throughout a collection can be diverse, making digitization a challenge. curators might employ different numbering systems or organizational practices, and almost always have different research specialties that in turn shape where resources are focused. collections with a greater continuity of curation throughout their history, with fewer staff changes (at all levels), or with curators and collection managers who are in agreement about the goals and management of a collection, can be far more effectively and efficiently digitized. data entry and data quality digitization (as distinct from imaging) frequently begins with recording specimen data. this basic data entry involves interactions between the person doing the data entry and with both the specimens and the data-capture client. often catalogues and ledgers are targeted first for data entry because the data are more accessible, with a format essentially that of a basic spreadsheet. as one respondent (usa) noted, the “absence of bulk data sources such as ledgers limit[s] data capture to the handling of individual collections objects,” which is significantly more time consuming. bulk data sources, however, are not without problems. in older collections or collections that have had many curators, there are frequently multiple catalogues with overlapping or duplicate numbers for different parts of a single collection. specimens can be entered into the database using prefixes or suffixes attached to catalogue numbers in order to differentiate objects with the same number, allowing each specimen to be renumbered uniquely. in addition, if data are entered directly from ledgers and catalogues rather than from the specimens themselves, once the collection has been databased, it is often necessary to go back into the collection and see which specimens listed in the ledger/catalogue are actually present in the collection. even if a specimen is missing, however, its data should be recorded since the specimen may be recovered and the data, itself, has scientific merit although diminished without the voucher. for some collection types (e.g., botanical, entomological), there may only be specimen information on the objects themselves, or information may be located in multiple places. finally, decisions must be made about whether or not to capture data from all possible places for individual specimens, weighing the benefits of more complete data capture against constraints of time, resources, and efficiency. a usable and efficient interface for data entry that limits keystrokes and streamlines the process is to customize the appropriate fields required for data entry. frequent dialogue between the people doing the data capture and those doing the programming is key to crafting data-entry clients that are smooth functioning and tailored to the 98 vollmar et al. –specimen digitization: challenges and concerns needs of individual disciplines and digitization projects. however, not all collections are able to work with programmers or have access to more complex clients and databases, and in such cases two survey respondents cited using microsoft excel and/or access as clients that “increased efficiency of data entry” (israel) and proved “easey [sic] for everyone – even non-educated people – to use and understand. later other programs can be in use” (denmark). the quality of the specimen data as recorded on original sources, such as specimen labels, ledgers, card catalogues, and field notes, can present a variety of stumbling blocks to digitization projects. transferring data from specimens, ledgers, and catalogues often necessitates deciphering poor handwriting, translating labels in foreign languages, making inferences about dates (e.g., day or month first), and decoding place names (e.g., historic locations no longer exist, place names change). in short, data that are not precise at the point of capture can be tricky to fit into more determined database structures. however, uncertainties in collection data must be digitized as they are originally recorded. for instance, updating names of places or taxonomic determinations without including the original designations can result in data being tied anachronistically to places and things, or to a loss of geopolitical information. coming up with ways of recording uncertainties in the original data, and subsequent inferences or interpretations made by people doing data entry, are key for maintaining high quality data. tagging uncertainties and problems for later review by qualified staff is a common practice in collections with staff entirely dedicated to data entry, but who may not be familiar with the discipline itself. streamlining workflow for data entry and imaging, solving data problems, and reducing the number of steps in handling and databasing specimens was another important issue for respondents. one respondent from australia “[made] label generation for specimens a product of databasing, not an additional task prior to databasing.” a tactic employed by a respondent from the usa was to “database 1 specimen per species as a ‘first pass’ to obtain a complete taxon checklist for the collection before returning to database the remaining…90% of specimens.” another respondent (usa) reported doing a “predigitization critical assessment of specimens, and elimination of low/no-data specimens.” others “standardized and documented the digitization protocol” (spain) and “wrote [a] protocol manual before data entry began” (usa). backlog, which can present difficulties in data capture, was addressed by a respondent in the netherlands by “divid[ing] backlog digitisation of collection labels in two phases, i.e. initial image plus basic data[,] and second other data and field notes.” technology funding can significantly limit options for using various technologies for digitization, and outdated technology and equipment can significantly slow if not halt projects. addressing technology was an important issue for a respondent in nicaragua, who “changed the old computer. the significant barrier is the low network we have.” other issues included the “need to create workflow to handle issues during data capture” (usa) and software being diverse (difficult to integrate data entered in different systems/projects) and/or underdeveloped. for collections with access to funds, improvements to the technologies used were key to addressing digitization bottlenecks, including the implementation of a system of unique identifiers like barcodes, purchasing digital cameras and scanners, using voice recognition software, using optical character recognition (ocr) software to enter data, joining multiuser collection management systems, using international standards like darwincore, and improving cataloguing software. collections data on-line: benefits, challenges, and feedback mechanisms in total, 60% of all survey respondents reported that at least part of their collections data were available on-line (see table 3). 99 vollmar et al. –specimen digitization: challenges and concerns out of the respondents with at least some on-line data, 55% were affiliated with government institutions, 62% were affiliated with university institutions, and 80% were from free-standing institutions reported that some of their collections’ data were available on-line. forty-six respondents elaborated on why their data were not yet available on-line, and the reasons largely fell into two categories. first, respondents expressed the plan to put all of their data on-line at once, and thus needed to digitize more data and check its quality before making it available (13 respondents from various countries including, in alphabetical order, brazil, canada, malaysia, netherlands, new zealand, spain, sweden, and usa). second, respondents cited a lack of reliable internet service, web servers, and website/software support, or that the necessary it infrastructure did not exist at their institutions (14 respondents from canada, india, jamaica, malaysia, netherlands, spain, usa, and uruguay). various reasons were cited for why collection data were not available on-line. for respondents affiliated with government institutions, the reasons included the lack of web servers, support, and it resources, and digitization was not yet complete. for respondents from university-affiliated institutions, data were not on-line because digitization and data clean-up was not yet complete, there was not enough staff to tackle projects, and web servers and it support were not available. also mentioned were the lack of time and/or money, both of which might be translated into general issues of staffing and equipment. only 3 respondents from free-standing institutions answered the open-ended question, mentioning issues of accuracy of data, institutional policies, and lack of a suitable website. table 3. what are the reasons your collections data are not available online? responses were open-ended. respondents from university institutions (n=42) respondents from government institutions (n=19) staffing issues 21% not reported lack of it resources 26% 58% digitization and data cleaning incomplete 43% 32% in reflecting on the benefits to their institutions of having their collections data available on-line, respondents cited an increase in use of collections and requests for data (31 out of 70 respondents, 44%), and a general heightened visibility of collections. data availability to the general public and to remote researchers were also considered assets. increasingly, preliminary research could be conducted on-line, and so respondents reported that questions and loan requests were more specific, resulting in less physical handling of the collection, thereby extending the longevity of the collection. respondents also noted the benefit of data correction, and receiving feedback on errors and/or misidentifications from outside users. additionally, one data manager from austria mentioned the opportunity for “virtual repatriation of material to countries of origin.” some respondents, answering the survey with respect to their institutions as a whole, reported that they had not experienced any problems since making their collection information available online, however, another reported: “the only downside is a lot more work coming in: editing the data, answering questions, checking data entry, revisiting determinations, etc” (usa). many respondents reported experiencing an increase in workload due to heightened visibility of collections, and did not have adequate staff to handle queries. finding the funding, staff expertise, and support to keep databases up and running was also noted as a challenge. many curators and collections managers mentioned errors and data quality as primary issues, citing difficulties in checking data prior to posting on-line, and expressing concern at the public accessing potentially flawed data. others (16 out of 57) cited technological problems as their top issues, ranging from a lack of it support and training for staff, to minimal server functionality, a poor institutional network system (colombia), and limited storage facilities for highresolution digital images (austria). exposure of sensitive species data was also an issue, which 100 vollmar et al. –specimen digitization: challenges and concerns respondents reported on occasion was accidentally revealed and, at other times, was requested by researchers, necessitating an evaluation of the research. at the same time, revealing data to politically charged entities was also a concern. one respondent from australia described their policy: “registration is required to access specimen details; this sometimes requires arbitration to fairly address and assign access for some applicants (especially from the community or mining industries).” one of the most significant challenges presented by data sharing is how to navigate/ensure returns on projects that share information for free. curators and managers of collections put a great deal of time, energy, and resources into digitizing data and making it available to larger and, significantly, more remote audiences. collection managers frequently give voice to the idea that collections are “alive” because researchers and specialists physically work with the collections, redetermining taxonomic identifications, rearranging parts of the collection, etc. while increasing remote access and on-line use of collections clearly has many benefits, at the same time, it presents difficulties in ensuring that feedback about collections is received, in particular that the work researchers do with collections data on-line make it back to the collection itself. another challenge expressed was keeping track of who uses the data and attribution for its use. as one respondent (usa) wrote, “[i] saw a paper published referencing only data from “ornis” [an ornithology resource pooling specimen data from many institutions] but not specifying which collections actually contributed data. we had a dozen relevant specimens to that study but were not able to determine if they were used.” another respondent (sweden) also mentioned, “people may think what is on-line is all we have, although only a small percentage is databased.” respondents reflecting on how they received feedback about their individual collections broadly cited two mechanisms: voluntary feedback links/forms, and restricting access by log in. more specific iterations of these tactics included conspicuously located contact information, user surveys, collection agreement pages that users must navigate through in order to access the data, requests for reprints of publications drawing on data, and required membership in order to access data. others reported offering limited or basic specimen data, thus requiring users to contact collection managers for further detail. nineteen out of 57 respondents (33%) noted that there was no mechanism in place to solicit feedback from on-line users. section 3: funding the primary impediment to digitization was reported as the lack of funding which is intimately linked to the people, who do digitization, in the form of salaries, and to the technology and infrastructure of digitization through the necessary purchase of equipment. in this section, we examine where different kinds of institutions look for funding, how respondents prioritized spending those funds, and how staff work time was utilized for digitization projects. we also include a list of the specific funding sources cited by respondents as providing significant resources for digitization projects. in reporting sources of funding used for digitization projects within the last two years (since 2007), 69% (136) of respondents received internal institutional funding, 54% (107) received public funding, and 30% (59) received private funding. no official funding was received by 4% of the respondents; 3% used either personal income or pursued digitization projects in free time and 1% drew on volunteer efforts. on average, respondents received 53% of their funding from internal institutional sources, 49% from public sources, and 23% from private funding sources. respondents’ answers did not add up to 100% in this question since monies were received from multiple sources, and so the percentages of each funding category do not add up to accordingly. while the limitations of these proportions as exact measures must be recognized, they do give a general sense of where collections are finding support for digitization projects. when applying for funding, 72% of all respondents reported that they explicitly requested funds for digitization projects, equipment, or people. among respondents identified as directors of institutions/programs, 86% requested funds specifically for digitization projects, as did 88% of data managers, 78% of researchers and faculty, and 69% of curator/collections managers. among respondents requesting support for collections 101 vollmar et al. –specimen digitization: challenges and concerns digitization projects, staff were ranked as their top need, followed by money, and then equipment. it should be noted that “money” in this context can not practically be distinguished from either staff or equipment needs. when there is no funding for digitization projects, 48% of respondents reported reallocating resources from other jobs or projects to support digitization work. other sources of resources used to cover the costs of digitization projects included annual budgets, internal and departmental funding, salaries of staff, funds available to hire students, endowments, and volunteers. many respondents described including digitization as a routine part of collection maintenance activities, or even collecting expeditions. finally, three very dedicated respondents reported using their own personal funds! for a list of funding sources respondents actively used within the past two years to support digitization projects in their collections, please see appendix a. in assessing how they would prioritize managing their collections in general if given more money, respondents identified specimen curation as their top priority (average ranking of 1.7, with 1 being most important). second most important were research (2.2) and collection storage/equipment (2.2). education was a distant priority at 3.2. additional priorities mentioned included digitization (30 respondents), specimen acquisition (5 respondents), increased staff (4 respondents), and space (4 respondents). when asked to reflect on how they would use additional funding for digitization, respondents commonly voiced the need to “simply digitize collection data,” and this was often paired with an increase in staff as key to enabling the success of data capture projects. seventy-nine out of 180 respondents said they would direct funds toward hiring more staff to work on digitization and data entry and, as per one respondent (usa), “staff hours – time is of the essence.” respondents (40) also said they would allocate funds for improved equipment and technology, purchase more equipment, better software, and more digital storage space, and one (usa) wanted to “develop methods for the curation of digital media associated with specimens (e.g., field notes, digital images, radiographs, etc.).” despite the broad consensus among respondents that staff were essential, only 55% of respondents indicated that staff currently performing digitization work had such tasks specified in their job descriptions. of staff who performed digitization work that was not specified in their job descriptions, 72% reported doing so as an extension of their daily tasks. thus, although staff are widely recognized as perhaps the most important aspect of digitization efforts, for many this work is not formally recognized as a part of their job. section 4: technology and data entry in this section, we look at the ways in which collections are stored which affects access for digitization, the sources from which data are being entered, how long data entry takes from the various sources, and what kinds of collections are not using unique identifiers. we briefly look at georeferencing and imaging, and then close with an overview of various technologies and their costs being used for digitization and imaging as reported by survey respondents. on average, of the collections reported, 56% of specimens were stored as dried and pressed on sheets, 33% of specimens were pinned, 26% were fluid specimens in vials or bottles, 16% were dried and in packets, 13% were skeletons and bones, 11% were skins and hides, 10% were slides, 7% were dried in vials or bottles, 7% were fluid in tanks, 6% were taxidermy mounts, 5% were tissues, and 4% were cleared and stained (wet skeletal preparation). most respondents (74%) reported that their collections utilized two or more different ways of recording information about table 4. if digitization work has been ongoing in your collection within the past two years, from what sources are/were you entering data? respondents entering data from each source (n=185) specimen labels 90% catalogues 41% ledgers 28% literature 24% cards 21% 102 vollmar et al. –specimen digitization: challenges and concerns specimens in their collection. it should be noted that respondents’ answers did not have to add up to 100% in this question (e.g., one specimen with multiple preparations), and so the percentages of each storage/prep type do not add up accordingly. while the limitations of these proportions as exact measures must be recognized, they do give a general sense of the physical nature of the specimens in collections of respondents surveyed. it should be noted that this does not reflect a global proportioning of storage or prep types. the majority of respondents (95%) reported that they used a database to record data about specimens in their collection, and information was recorded from various sources (see table 4). for the average time and number of data fields entered from each of these sources, see table 5. additionally, 63% of all respondents reported georeferencing specimen data, although only 50% of those were doing so according to best practice guidelines or standards. apart from data entry, 64% of respondents surveyed reported that they were doing some imaging of specimens, which for most took anywhere from less than 5 minutes to 10 minutes (see table 6). most (91.5%) of all respondents said they assigned unique identifiers to specimens in their collection. of those specimens associated with unique identifiers, on average 70% were associated with catalogue numbers and 22% with barcodes. in the final section, we will touch on some of the challenges faced by various disciplines in assigning unique identifiers. digitization and imaging technology when reflecting on the primary factors they considered when purchasing digitization equipment/technology, on a scale from 1 (most table 5. on average, how long does it take you to enter data from each of the following sources? how would you characterize the number of data fields you are entering per specimen? low (1-10 data fields) medium (11-20 data fields) high (20+ data fields) respondents entering data from each source time to enter data respondents entering data from each source time to enter data respondents entering data from each source time to enter data specimen labels (n=161) 15% 5-9 minutes 59% 5-9 minutes 27% 5-9 minutes catalogues (n=71) 15% 5-9 minutes 57% 5-9 minutes 28% 5-9 minutes ledgers (n=49) 16% 5-9 minutes 55% 5-9 minutes 29% 5-9 minutes literature (n=44) 16% 10-14 minutes 55% 10-14 minutes 30% 20+ minutes cards (n=36) 8% 1-4 minutes 66% 5-9 minutes 26% 5-9 minutes table 6. if you are imaging specimens, how long does it take to image a specimen? respondents reporting the length of time to image specimens less than 5 minutes (n=37) 30% 6-10 minutes (n=38) 30% 11-15 minutes (n=16) 13% 16-20 minutes (n=9) 7% 20+ minutes (n=15) 12% 103 vollmar et al. –specimen digitization: challenges and concerns table 7. what equipment are you using for digitization projects in your collections and, if known, approximately how much did each piece of equipment cost? if you are imaging, what technology are you using to image and, if known, approximately how much did it cost? scanners price software price herbscan -geo locate - indus book scanner $25,000 imagemagick - epson expression 10000xl $3,000 nikon capture nx - microtek scanmaker 9800xl $2,000 phase capture one - nikon super coolscan 5000 $1,200 robogeo gps - digital cameras price sinarcapture shop - fuji s2 pro -leica automontage $3,000 nikon coolpix 995 -adobe design suite/ photoshop $400 nikon d90 -other equipment price nikon d100 plus 100mm macro lens -barcode printer - sony dh-5 -barcode scanner - tethered sony a900 -camera stand - digital photomicroscope -lighting equipment - x-ray imaging -sound recording devices - jenoptik eyelike m22 camera back and schneider apo-digitar 90mm/f 4.5 lens with tti camera stand/ lighting $60,000 file server space (3 terabytes) $3,600 syncroscopy automontage $60,000 apple mac pro computer, 2x2.8ghz $5,312 sinar evolution 75h multi shots digital back system 33 mp $39,794 datamax printer for archival quality specimen labels $4,000 leica mz16a stereomicroscope with leica motor focus and a leica dfc 420 $30,000 microchip labels and microchip scanner €3,000 tti-repro-graphic workstation 3040/digiflex 67ei/ sinar 75h $28,117 external hard drives for backup $1,000 leica mz10 $14,000 portable terabyte storage drives for archiving specimen scans au$600 large format camera with betterlight digital back $10,000 beseler copystand $500 coloreal ebox $3,000 canon rebel xti $1,000 usb microscope camera €70 104 vollmar et al. –specimen digitization: challenges and concerns botany/mycology storage types: important) to 3 (least important), respondents’ top concern was quality (1.4), followed by suitability to project (1.5), compatibility with existing/future technology (1.66), ease of use (1.7), cost (1.8), and durability (1.9). other factors reported included speed and efficiency (4% of respondents), availability of support/staff expertise (2% of respondents), and compliance with funding sources’ requirements (2% of respondents). the proximity of importance of these factors suggests that none of them takes full precedence, and that they are significantly dependent on the needs of any given digitization project. in table 7, the technologies used by respondents in both digitization/data entry and imaging projects are summarized. all costs were reported by respondents in us dollars unless otherwise specified, and have not been confirmed by any further external research. the equipment used reflects a wide variety of uses and also resource availability, with both $400 digital cameras and $60,000 imaging set-ups included in the survey results. section 5: disciplinary concerns in this section, particularities of collections in various disciplines are addressed. because many survey respondents worked in multiple collections and answered the survey with respect to more than one scientific discipline, survey data are less useful for addressing discipline-specific trends. to supplement the survey data, interviews with a variety of staff at the harvard university museum of comparative zoology, harvard university herbarium, and the yale university peabody museum were conducted to help form the foundation of this section. this section is not intended to establish the definitive characterizations of specific disciplines, since some observations from the interviews may be more reflective of challenges faced by individual collections and institutions. nevertheless, many of the issues noted cross institutional boundaries and give a general sense of the problems facing specific types of collections. n sheets, dried a total of 119 respondents indicated that they wor f herbarium spe dried and pressed o in packets, dried fruit and seed collections ethnographic artifacts, petrified wood, fluid specimens, fossil specimens. ked with either botanical or mycological collections, or both. many of these respondents also worked with other disciplines, so observations based strictly on discipline are somewhat general. on average, these respondents reported that 77% of their specimens were dried and pressed on sheets, 18% were dried and in packets, 6% in fluid-filled bottles/vials, and 3% slides. respondents also mentioned live plant collections, wood collections, specimens on rocks, fossils, and photographs. the majority of botany/mycology respondents (96%) said they used unique identifiers in their collection, and 84% reported there was digitization work ongoing. barriers to the digitization o cimens begin with the number of specimens: material frequently comes in faster than can be databased, leading to a significant backlog that is compounded in larger collections by the volume of specimens already contained within a particular collection. because herbarium sheets are relatively large pieces of paper, this means there is often abundant data to capture, making the process of digitization time consuming, and the possibility likely of encountering information that does not fit particular fields in a given data capture client. accordingly, 94% of respondents were entering data from specimen labels, 25% from catalogues, 21% from ledgers/accession records, and 21% from literature. most respondents (62%) were georeferencing specimen data, and of those, 48% were doing so according to best practice guidelines/standards. a majority of the respondents (59%) were imaging specimens, although the relative importance/abundance of imaging is unknown. see table 8 for data entry statistics. 105 vollmar et al. –specimen digitization: challenges and concerns additional barriers—none of which are unique to botanical specimens—include illegible handwriting, poor documentation, synonymy, space, and multiple sheets of a single specimen. ethnobotanical collections—particularly historic ones—present a unique set of challenges as ethnobotanical collecting methods and associated documentation have been largely unstandardized, and such collections encompass a wide range of storage types (herbarium sheets, raw materials, artifacts, fruits and seeds). the ethical dimension of ethnobotanical collections involves questions of repatriation and indigenous rights that are faced by anthropological and archaeological collections. entomology storage types: pinned insects, riker mounts, papered specimens (dried and folded in envelopes), fluid specimens, frozen tissues. a total of 55 respondents indicated that they worked with entomological collections. some of these respondents also worked with other disciplines, so observations based strictly on disciplines are somewhat generalized and exact proportions of collection storage/prep types could not be calculated. in general, entomological specimens are mounted on pins and organized in drawers, which may not always be organized in lots (i.e., many individuals of the same species from the same collecting event) and may contain multiple species since groups of insects are frequently clustered together in smaller boxes called unit trays. information about specimens can thus be associated with a specimen itself, with a unit tray, or with a drawer as a whole; parsing these different levels of information into a database can be challenging. entomological collections rarely have ledgers or card files; instead, nearly all data are physically associated with the specimen itself, usually on small pieces of paper pinned below the specimen. specimen data are often recorded in nonstandard shorthand that abbreviates the location and date of collection, and species identification for each specimen. in some cases, this information is recorded in a purposefully cryptic manner to hide collecting locations, and so deciphering these labels can be quite difficult, often requiring specialized knowledge accumulated over many years of involvement in the field. getting the necessary species data off the specimen and into a database is, thus, a time-consuming process; pins have to be removed, and specimens handled. for the same reason, adding items like barcodes to individual specimens is physically challenging. thus, it is not unsurprising that of the survey’s 18 respondents who said they did not assign unique identifiers, 14 were from entomological collections (housed in university-affiliated institutions). eight respondents reported that table 8. data entry statistics for botany/mycology collections. on average, how long does data entry take per specimen from each data source? (n=96) on average, how would you characterize the number of data fields per specimen you enter? approximately how long does it take to image a specimen? data source time data fields percent of respondents (n=100) time percent of respondents (n=66) specimen labels 6.1 minutes low 1-10 fields 12% 5 minutes or less 41% catalogues 7.5 minutes medium 11-20 fields 63% 6-10 minutes 32% ledgers/ accession records 6.2 minutes high 20+ fields 25% 11-15 minutes 9% cards 6.6 minutes 16-20 minutes 6% literature 9.8 minutes 20+ minutes 3% 106 vollmar et al. –specimen digitization: challenges and concerns digitization work was not ongoing in their collection, with the top reasons cited as lack of funding, time, institutional support, staff, and infrastructure and technology. eight respondents also did not have data from their collections available on-line, largely because data had not yet been digitized, and web servers/internet service were unreliable or absent. many of the respondents in entomological collections (69%) reported that there was some kind of digitization work ongoing in their collection. an ongoing lepidoptera imaging project at the museum of comparative zoology, harvard university uses barcodes that are double sided so that if they are not readable from above, the insect can be removed, turned over, and the barcode read without any further manipulation of the specimen. additionally, the barcodes used are readable even if punctured with a pin. because the specimen data is in shorthand and data entry personnel generally will be unable to interpret the meaning of what is written on each label, the project plans to utilize crowd-sourcing (i.e., upload images of specimen labels to a wiki-style website) entomologists all over the world— particularly amateurs—can then submit interpretations of labels and specimen information to the project. mammalogy storage types: study skins, cased skins (preparations without cuts to the abdomen), skeletons, fluid specimens, taxidermy mounts, histological slides, embryos, frozen tissues, observation data and measurements of each mammal. a total of 38 respondents worked with both mammalogy and ornithology collections, and 33 with only mammalogy collections. some respondents also worked with other disciplines (especially herpetology and ichthyology). because of the nearly complete overlap between respondents working with mammalogy and ornithology collections, discipline-specific observations from the survey are problematic. with mammalogy collections, cataloguing and databasing specimen information is not the most time consuming part of incorporating mammals into the collection; rather, the limiting step is the time it takes to prepare mammals. specimen preparation can take a relatively long time for large mammals, so databasing is only a brief moment in a much longer process. an abundance of diverse kinds of data about each specimen must be entered into a database, presenting challenges in crafting appropriate data entry interfaces. from the collecting end of the process, this can be streamlined by clear communication between collections managers and collectors in the field about what kinds of information are necessary to gather; one individual gives field-kits with data checklists to collectors so as to make data entry and cataloguing a smoother process. ornithology storage types: skins, skeletons, fluid specimens, egg and nest lots, taxidermy mounts, frozen tissues, histological slides. a total of 38 respondents worked with both ornithology and mammalogy collections, and 36 with only ornithology collections. problems with discipline-specific observations are as noted previously. one element of digitizing ornithology collections that requires attention is the need to associate different parts of particular specimens that are stored scattered throughout collections (e.g., skeletons with skins, or photographs of birds prior to collection and the resulting skins). fully digitizing information about particular specimens necessitates physical sleuthing and crossreferencing within collections. accommodating the different kinds of data that must be entered depending on the kind of specimen can also be a challenge, such as efforts to database eggs and nests, which incorporate observation data. a nest and its eggs are considered a “lot,” and the information recorded includes the number of eggs per nest, the location of the nest (e.g., ground or tree and, if tree, then height), how the species was identified, if the adult was seen while collecting, and how long the eggs were incubated prior to collection. in addition, the ability for databases to incorporate auditory data of bird sounds adds another layer of complexity and possibility to digitizing ornithology collections. 107 vollmar et al. –specimen digitization: challenges and concerns ichthyology and herpetology storage types: fluid specimens (including in metal tanks), skeletons, cleared and stained specimens, frozen tissues, histological slides, taxidermy mounts. specific to herpetology are skins/hides, and turtle shells. a total of 44 respondents worked with both herpetology and ichthyology collections, 35 with only herpetology collections and 32 with only ichthyology. respondents also worked with other disciplines (especially ornithology and mammalogy). the majority of herpetology/ichthyology collections (85%) had digitization ongoing. while many specimens are individuals, both ichthyology and herpetology collections contain lots, multiple individuals (ranging from two to thousands) gathered together in a single collecting event and catalogued as a single specimen. each individual organism does not have a unique identifier, unless the specimen is examined for research purposes, in which case ideally each is identified individually. digitizing lots presents challenges in terms of the number of subdivisions of the organism, and the levels of part enumeration supported by various databases. most of the respondents (68%) in herpetology/ichthyology collections were imaging specimens. significant challenges for both ichthyology and herpetology collections are the issues faced in imaging fluid specimens, which can be time consuming. some of the respondents (19%) reported that it took them less than 5 minutes to image a specimen, 41% said it took 610 minutes, 22% took 11-15 minutes, 4% took 1620 minutes, and 11% took more than 20 minutes. for example, some specimens must be removed from their jars and entirely submerged in fluid to create a single focal plane. large and oddly shaped specimens like snakes present still further difficulties. the majority of the respondents (71%) working with herpetology/ichthyology collections reported georeferencing specimen data. of those, who were georeferencing, however, only 53% were doing so according to best practice guidelines/standards. georeferencing presents a number of issues for collections with localities tied to water. historically, fresh water collections were often not identified with latitude and longitude but rather descriptions of particular places, while oceanographic collections are identified with longitude and latitude, but precision was less important. precision becomes of utmost importance in cases like collections tied to rivers. rivers can be thousands of miles long so, without further identifying information, it can be impossible to discern where along a river a specimen was collected. different styles by collectors of describing localities create challenges for digitization efforts attempting to enrich specimen data by assigning precise locations to specimens. georeferencing is further complicated because the error radius frequently employed to designate uncertainties in location demarcates a circular area. many of the geographic water features referenced in fish collections are not perfectly round and are near land. if the error radius is applied without consideration to the water boundaries, land will be included within the error radius. collections at rivers frequently occur at bridges and other sites of crossing; thus, a more useful designation of uncertainty would only extend upstream and downstream rather than in a circle radiating from the assumed point of collection. specimens collected in coastal locations can likewise be confusing since an error radius might encompass both fresh and salt water habitats, and land. georeferencing tools that are linked to data about terrain and the presence or absence of water are more useful for ichthyology specimens. invertebrate zoology storage types: fluid specimens, dried sponges and corals, shells, specimen slides, and frozen tissues. a total of 21 respondents worked with invertebrate zoology collections. these respondents also were heavily involved with other collections, particularly ornithology and herpetology (18 respondents) and mammalogy and ichthyology (16 respondents). responses by discipline, then, are difficult to gauge in this case for reasons discussed earlier. one of the challenges in digitizing invertebrate collections is numbers: there are so many species and specimens that it is hard to database them all. invertebrates are also often collected in large quantities and grouped in lots. 108 vollmar et al. –specimen digitization: challenges and concerns lots in invertebrate collections can contain thousands of organisms, which presents difficulties when loaning such specimens, as typically each organism in a lot is counted before being loaned. as a result, many lots have only approximations of the number of organisms they encompass. in databases, multiple levels of enumeration in lots (i.e., if they are subdivided) can be hard to accommodate and track. vertebrate and invertebrate paleontology storage types: fossil specimens, microfossils on slides, slabs and oversized specimens on carts. specific to vp are trackways and fossil skeletons. a total of 29 respondents worked with vertebrate and invertebrate paleontology collections. these respondents were also involved in other collections, particularly herpetology (16), mammalogy (15), and ornithology (14). again, responses by discipline are difficult to gauge. locality data are of utmost importance in digitizing paleontology specimens; without locality data, a specimen is nearly worthless because location is as important as a specimen’s taxonomic identification. however, digitizing paleontological specimens requires the inclusion not only of information about collection locality, but also geologic information like unit, age, series (upper, middle, lower), formation, and beds/members/zones. in particular, invertebrate zoology stratigraphic collections—which give data about specimens through time and in various locations—are difficult to database because of their more temporal orientation. specific to vertebrate paleontology is the challenge of describing what part of a specimen a given object might be. because vertebrate paleontology specimens are composed of diverse and complex objects, over time layers of narrative attempting to describe objects can become an incomprehensible description. attaining consistency in description, which is key to describing and identifying specimens, is thus critical. section 6: conclusions this paper serves to highlight the many challenges and concerns relevant to digitizing natural history collections. the detailed findings provide both a status quo for how specimens are being digitized presently, and a window into the issues that need to be resolved in order to break down barriers and make the digitization process more efficient. the survey strongly suggests that the greatest barrier to digitization is the cost of doing the work, from hiring staff to purchasing technology. the cost barrier is not a simple one to overcome as many respondents noted that there are very few funding sources available to tap. the burden of digitization, thus, generally falls to the resources at hand and institutional and/or collection priorities, making the process slow and very uneven across the collection landscape. this survey also highlighted that although each kind of collection has some important domain-specific issues, there are many challenges that are common to all. thus, communication of knowledge about digitization among disciplines could have great benefit toward overcoming common challenges, such as imaging, automated text scanning and parsing, georeferencing, etc. indeed, the survey respondents suggested that there was a great need for more communication of knowledge on digitization through training and documentation of best practices. in addition, many respondents, most likely from smaller institutions, reported already working with multiple collections, especially in the vertebrate and paleontological disciplines, which would facilitate the spreading of information across collections. community standards and collaborative efforts need to be further developed and embraced by those who create the software that the collection community relies on. the added benefit would be that newly created components or modules could then be linked together into workflows, which could address the specific needs of a particular collection. the barriers to digitization ultimately lead to a patchy set of digitized records being available to potential users, which can clearly be demonstrated by searching the gbif portal. this reality reduces the availability of specimens for research and broader purposes. this patchiness also extends to the amount and quality of the data records being captured, which has a direct impact on the data's fitness for use. reduction of fitness can negatively impact the data’s application to broader issues and the power of the conclusions on which any analyses are based. 109 vollmar et al. –specimen digitization: challenges and concerns although the social barriers to digitization were not specifically addressed in the survey, it became quickly apparent that these issues were prevalent. in no small part because of the social nature of natural history collections, we hope that some elements of this survey and paper can serve as practical tools for collections embarking on digitization projects, such as approximating how long it will take to digitize a ledger based on the number of data fields. we also hope that the survey results illuminate the paucity at the level of implementation and will help show the need for, and serve as a guide to, funding sources for digitization projects. given the abundance of natural history specimens and the potentially broad importance of their associated data ranging from disciplinespecific research, to public initiatives, to global issues, digitization must be undertaken as strategically as possible. a recurring theme in the survey was that digitization efforts must have the most impact in the shortest amount of time, and for a reasonable cost; this appeal, by its very design, requires coordination within and among collections and institutions. much as collecting and specimen acquisition has always been a social endeavor, so must be the effort to render those same collections digital. acknowledgements the authors would like to acknowledge funding for the society of preservation of natural history (spnhc) best practices intern, ana vollmar (first author), from alan prather, pi of the national science foundation research coordinating network:collectionsweb (# 0639214) and support from the global biodiversity information facility (gbif). we would especially like to thank the following collection professionals who were interviewed, including (in alphabetical order): arturo ariño (spain), roger baird (cmn), adam baldinger (mcz), susan butts (yale), judith chupasko (mcz), jessica cundiff (mcz), rod eastwood (mcz), brendan haley (mcz), michelle hamer (south africa), karsten e. hartel (mcz), shusheng hu (yale), eric lazo-wasem (yale), paul morris (huh and mcz), chris norris (yale), malcolm scoble (nhm), patrick sweeney (yale), jeremiah trimble (mcz), gregory watkins-colwell (yale), andrew williston (mcz), jonathan woodward (mcz), and krzysztof zyskowski (yale). references cited araújo, m.b., pearson, r.g., thuiller, w. and erhard, m. 2005. validation of species-climate impact models under climate change. global change biology, 11, 1504-1513. beaman, r., macklin, j.a., donoghue, m.j., and hanken, j. 2007. overcoming the digitization bottleneck in natural history collections: a summary report on a workshop held 7-9 september 2006 at harvard university. http://www.etaxonomy.org/wiki/images/b/b3/harv ard_data_capture_wkshp_rpt_2006.pdf [accessed september 11, 2010]. butler, d., gee, h. and macilwain, c. 1998. museum research comes off list of endangered species. nature: briefings, 394: 115-117. [doi:10.1038/28009]. gbif. 2010. global biodiversity information repository portal. http://data.gbif.org/welcome.htm [accessed september 11, 2010]. lane, m.a. 1999. weaving a web of wealth: biological informatics for industry, science and health. australian academy of science, canberra, australia. 40 pp. loarie, s.r., carter, b.e., hayhoe, k., mcmahon, s., moe, r., knight, c.a., and ackerly, d.d. 2008 climate change and the future of california's endemic flora. plos one 3(6): e2502. [doi:10.1371/journal.pone.0002502]. national science and technology council, committee on science, interagency working group on scientific collections. 2009. scientific collections: mission-critical infrastructure of federal science agencies. office of science and technology policy, washington, dc, 2009. http://www.whitehouse.gov/sites/default/files/scicollections-report-2009-rev2.pdf [accessed september 11, 2010]. peterson, a.t. and vieglais, d.a. 2001. predicting species invasions using ecological niche modeling: new approaches from bioinformatics attack a pressing problem. bioscience 51(5): 363371 [doi:10.1641/00063568(2001)051[0363:psiuen]2.0.co;2]. 110 http://www.etaxonomy.org/wiki/images/b/b3/harvard_data_capture_wkshp_rpt_2006.pdf http://www.etaxonomy.org/wiki/images/b/b3/harvard_data_capture_wkshp_rpt_2006.pdf http://data.gbif.org/welcome.htm http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf http://www.whitehouse.gov/sites/default/files/sci-collections-report-2009-rev2.pdf vollmar et al. –specimen digitization: challenges and concerns pinto, c.m., baxter, b.d., hanson, j.d., méndezharclerode, f.m., suchecki, j.r., grijalva, m.j., fulhorst, c.f., and bradley, r.d. 2010. using museum collections to detect pathogens [letter]. emerg. infect. dis. [serial on the internet]. http://www.cdc.gov/eid/content/16/2/356.htm [accessed september 11, 2010], [doi:10.3201/eid1602.090998]. shaffer, h. b., fisher, r. n. and davidson, c. 1998. the role of natural history collections in documenting species declines. trends in ecology and evolution 13: 27-30. suarez, a.v., and tsutsui, n. d. 2004. the value of museum collections for research and society. bioscience 54: 66-74. synthesis. 2010. synthesis of systematic resources. http://www.synthesys.info/ii_na.htm [accessed september 11, 2010]. winker, k. 2004. natural history museums in a postbiodiversity era. bioscience 54(5): 455-459. appendix a funding sources used within the past two years to support digitization projects in the collections of the survey respondents. • accessprojektet, sweden • african plant initiative/mellon foundation • amherst college, usa • artstor • atlantic canada conservation data center • australia virtual herbarium project • australian government • barcelona university, spain • batson endowment for the a. c. moore herbarium, usa • biology department, texas a&m • biomap project • boeing, usa • canadian foundation for innovation • canadian museums association • census of marine life • concejo distrito capital, colombia • concejo nacional de ciencia y tecnología de guatemala (concyt), guatemala • consejo nacional de investigaciones científicas y técnicas (conicet), argentina • conselho nacional de desenvolvimento científico e tecnológico (cnpq), brazil • darwin initiative, uk • department of conservation, terrestrial and freshwater biodiversity information system (tfbis) fund, new zealand • dirección general de investigación (digi), universidad de san carlos, guatemala • dutch scientific foundation • earthwatch institute • environment canada • environmental foundation of jamaica, virtual herbarium project • european distributed institute of taxonomy (edit) • european union funds • finnish ministry of education • fisheries and oceans canada • friends of the university of alberta museums, canada • fundação de amparo a pesquisa do estado de são paulo (fapesp), brazil • fundação para a ciência e tecnologia, portugal • global environment facility (gef)andes project • global biodiversity information facility (gbif) • government of jamaica • government of newfoundland and labrador, canada • government of spain • government of the netherlands • häagen-dazs • hanes trust • harvard university, usa • hearst scholarship foundation, california state university, usa • institute of museum and library services (imls), usa 111 http://www.cdc.gov/eid/content/16/2/356.htm http://www.synthesys.info/ii_na.htm vollmar et al. –specimen digitization: challenges and concerns 112 • inter-american biodiversity information network (iabin) • israel academy of sciences • kansas state university, usa • latin american plant initiative (lapi)/ mellon foundation • louisiana board of regents, usa • max planck gesellschaft, germany • mellon foundation • ministerie van onderwijs, cultuur en wetenschap, netherlands • ministerio de ciencia e innovación, spain • ministerio de educación y ciencia, spain • ministry of science and education, spain • mondriaan foundation, netherlands • museum assistance program, canada • museum of comparative zoology, harvard university, usa • nagisa program and goma program (alfred p. sloan foundation) • national cancer institute, usa • national history museum, stockholm, sweden • national science foundationbiological research collections (nsfbrc) grant, usa • national science foundation, usa • netherlands organization of scientific research (nederlandse organisatie voor wetenschappelijk onderzoek, nwo) • new brunswick wildlife trust fund, canada • new mexico state department of fish and game, usa • norwegian development agency (norad) • norwegian ministry of foreign affairs • overseas territories environment programme, uk • penn state, usa • plant health australia • pollinators thematic network (ptn) initiative, iabin • red nacional de información académica, colombia • renata, colombian ministry of education • riksbankens jubileumsfond formas, sweden • secretaría de estado de medio ambiente y recursos naturales (semarena), dominican republic • servizio civile nazionale, italy • smithsonian trust funds, usa • southeast regional network of expertise and collections (sernec), usa • svenska artprojektet, sweden • swedish species information center, artdatabanken • swedish university of agricultural sciences • the swedish taxonomy initiative • unesco-l'oréal for women in science fellowship • united states bureau of land management • united states department of agriculture (usda) • usda current research information system (cris) project funds, usa • united states fish and wildlife service • united states forest service • united states national park service • universidad de león, spain • universidad de san carlos de guatemala • university of florida, usa • university of iowa, usa • university of malay, malaysia • university of new brunswick department of biology, canada • university of wisconsin natural history museums council, usa • van eeden, a. mennege, h. de vries, van leersum foundations, sweden • young canada works in heritage institutions introduction methods section 1: overview of survey respondents section 2: barriers to digitization funding staff space and curation data entry and data quality technology collections data on-line: benefits, challenges, and feedback mechanisms section 3: funding section 4: technology and data entry digitization and imaging technology section 5: disciplinary concerns botany/mycology entomology mammalogy ornithology ichthyology and herpetology invertebrate zoology vertebrate and invertebrate paleontology section 6: conclusions acknowledgements references cited appendix a microsoft word bi terrestrial enm 20061130.doc biodiversity informatics, 3, 2006, pp. 59-72 59 uses and requirements of ecological niche models and related distributional models a. townsend peterson natural history museum, the university of kansas, lawrence, kansas 66045 usa e-mail town@ku.edu abstract.—modeling approaches that relate known occurrences of species to landscape features to discover ecological properties and predict geographic occurrences have seen extensive recent application in ecology, systematics, and conservation. a key component in this process is estimation or characterization of species’ distributions in ecological space, which can then be useful in understanding their potential distributions in geographic space. hence, this process is often termed ecological niche modeling or (less boldly) species distribution modeling. applications of this approach vary widely in their aims, products, and requirements; this variety is reviewed herein, examples are provided, and differences in data needs and possible interpretations are discussed. key words.− ecological niche, ecological niche modeling, distribution modeling, geographic distributions. recent years have seen impressive growth in use of modeling approaches based on relationships between known occurrences of species and features of the ecological and environmental landscape (guisan and zimmermann 2000; pearson and dawson 2003; peterson 2003a; soberón and peterson 2004). these models are often termed ‘distribution models,’ ‘climatic envelope models,’ or (most generally) ‘ecological niche models.’ the aim of these studies is generally to reconstruct species’ ecological requirements and/or predict geographic distributions of species. the earliest applications in this realm were undoubtedly those of joseph grinnell, in the 1910s and 1920s (grinnell 1917; grinnell 1924), who used the spatial distribution of occurrences of species to infer factors limiting their distributions, and laid a firm foundation for subsequent work in this field. the diversity of such applications, however, has now grown considerably. distributional models and ecological niche models are being used not just to understand species’ ecological requirements, but also to understand aspects of biogeography, predict existence of unknown populations and species, identify sites for translocations and reintroductions, plan area selection for conservation, forecast effects of environmental change, etc. (table 1). a basic dichotomy that pervades both the list of uses to which these methods are put and even the terminology used to refer to them is that of ecological niche modeling (enm) versus distributional modeling (dm). a recent paper (soberón and peterson 2005) formalized the idea of niche modeling, and clarified the differences between these two views. niches and distributions of species were visualized as a set of 3 intersecting circles, representing diagrammatically 3 classes of determinants: physical conditions necessary for a species’ survival and reproduction (“abiotic niche”; e.g., correct combinations of humidity, temperature, other biophysical variables, substrate types, disturbance regimes), biotic conditions necessary for a species’ survival and reproduction (“biotic niche”; e.g., presences of mutualists, absences of diseases and predators), and accessibility (i.e., within the dispersal capabilities of the species, either historically or at present) (figure 1, top). this latter set of factors is not a niche dimension, but rather is a set of nonecological factors that constrain the species to inhabit less than its full distributional potential, and may indeed not be permanent—as shown in the case of invasive species, dispersal limitations often with time are overcome. in this framework, where abiotic conditions are appropriate can be compared with the fundamental pe te r so n – u se s a n d r eq u ir em en ts o f ec o lo g ic a l n ic h e m o d el s 60 ta bl e 1. s um m ar y of u se s t o w hi ch e co lo gi ca l n ic he m od el s h av e be en p ut , a nd th e re qu ire m en ts th at th es e us es h av e in te rm s o f o ut pu t a nd in fo rm at io n co nt en t. q ua lit y of in te re st u nd er st an d ec ol og ic al re qu ire m en ts of sp ec ie s u nd er st an d bi og eo gr ap hy an d di sp er sa l ba rr ie rs fi nd u nk no w n po pu la tio ns fi nd n ew sp ec ie s id en tif y si te s f or tra ns lo ca tio ns a nd re in tro du ct io ns c on se rv at io n pl an ni ng a nd re se rv e sy st em de si gn pr ed ic t ef fe ct s o f ha bi ta t lo ss pr ed ic t sp ec ie s’ in va si on s pr ed ic t cl im at e ch an ge ef fe ct s fo rm o f pr ed ic tio n (e .g ., bi na ry , r an ke d, ab so lu te ) a ny a ny a ny a ny a bs ol ut e a bs ol ut e a ny a ny a ny g ra in re qu ire d a ny a ny po pu la tio n a ny in di vi du al po pu la tio n in di vi du al po pu la tio n a ny a ny a ny , b ut m ay be li m ite d by re so lu tio n of cl im at e ch an ge d at a se ts c au sa l v ar ia bl es ne ed ed , o r su rr og at es o k ? c au sa l c au sa l su rr og at es su rr og at es o k if w ith in ra ng e; ca us al n ec es sa ry if ex tra po la tin g su rr og at es o k if w ith in ra ng e; ca us al n ec es sa ry if ex tra po la tin g su rr og at es c au sa l c au sa l c au sa l n ee d m od el re sp on se c ur ve or p ar am et er re tri ev al y es y es n o n o u se fu l, to a vo id er ro r u se fu l, to av oi d er ro r n o n o n o er ro r n ee ds o ve ra ll lo w er ro r n ee de d o ve ra ll lo w er ro r n ee de d lo w o m is si on (d on ’t m in d se ar ch in g ex tra lo ca lit ie s, bu t do n’ t w an t t o le av e an yt hi ng ou t) lo w o m is si on (d on ’t m in d se ar ch in g ex tra lo ca lit ie s, bu t do n’ t w an t t o le av e an yt hi ng ou t) lo w c om m is si on (v er y hi gh c os t o f er ro rs ) lo w co m m is si on (v er y hi gh c os t of e rr or s) o ve ra ll o ve ra ll o ve ra ll po te nt ia l di st rib ut io na l m od el v er su s re al iz ed di st rib ut io n? po te nt ia l po te nt ia l r ea liz ed po te nt ia l r ea liz ed r ea liz ed r ea liz ed po te nt ia l b ot h u nc er ta in ty es tim at es ne ed ed ? n o n o n o n o y es y es n o n o n o 60 peterson – uses and requirements of ecological niche models 61 ecological niche of the species, and where abiotic and biotic conditions are fulfilled can be compared with the realized ecological niche of the species (hutchinson 1957), although hutchinson focused mostly on competition among the broader suite of potential biotic intereractions; these interactions could potentially be integrated more intimately into the niche modeling framework (araújo and guisan 2006). the geographic projection of these conditions (i.e., where both abiotic and biotic requirements are fulfilled) represents the potential distribution of the species (figure 1, top, blue area)—areas where the species could survive if introduced. finally, those areas where the potential distribution is accessible to the species is likely to approximate the actual distribution of the species (figure 1, top, black area). other authors (araújo and guisan 2006) have further distinguished between my ‘actual’ distribution and the area actually occupied at any point in time, taking into account stochasticity, metapopulation dynamics, etc. enm proponents are interested in using distributional information (i.e., known occurrences sampled from the actual distribution) to estimate ecological niches and potential distributions of species, which then provides a means of understanding and anticipating ecological and geographic features of species’ distibutional biology (soberón and peterson 2005). this approach has the advantage of allowing the effects of the three components listed above to be distinguished, which offers greater interpretability as to causation of phenomena, and permits predictability of phenomena that depend on the differences between components—e.g., the invasive potential of a species depends on the difference between potential and actual distributional areas. this approach, nonetheless, requires additional complexities of interpretation to produce estimates of actual geographic distributions, given that the data on which the niche models are based are not drawn from the entire abiotic niche or even from the potential distribution, but from the actual distributional area (araújo and pearson 2005; pearson and dawson 2003; soberón and peterson 2005; svenning and skov 2004). dm proponents, on the other hand, include effects of abiotic, biotic, and accessibility considerations in their models from the outset. they argue that because distributional information is an expression of a realized ecological niche, as such the realized niche (only) is the target of modeling. as such, dm proponents would often include in the modeling approach independent variables that summarize biotic considerations (e.g., distributions of other species in the region) and that bring in spatial considerations that may be relevant to dispersal ability and accessibility (leathwick 1998; latimer et al. 2006). whereas dm is simpler in producing estimates of species’ actual geographic distributions directly, predictivity across scenarios of change is largely lost, and assumptions regarding accessibility of areas may still required (soberón and peterson 2005). a further discussion of the differences between enm and dm is provided below. this diversity of ideas can place different demands on features of algorithms and approaches used to develop models—and clearly has the potential to lead to debate and perhaps misunderstanding between workers with distinct needs and interests. nonetheless, contrasts between different conceptualizations of the process (e.g., enm versus dm) have seen little direct discussion in the literature (araújo and guisan 2006; soberón and peterson 2005). such is the purpose of this contribution: to survey the diverse uses to which these approaches have been put, and discuss differences in data needs and interpretation that this diversity demands. ecological niches and evolutionary conservatism of niches in general, readers are referred to recent conceptual reviews (chase and leibold 2003; pulliam 2000; soberón and peterson 2005) of relationships among autecology, synecology (i.e., species interactions), and history and accessibility. once again, this general approach was pioneered by joseph grinnell (grinnell 1917; grinnell 1924), who was directing detailed series of biological inventories, and was thinking about why species are where they are, and why they are not where they are not. grinnell’s approach, obviously peterson – uses and requirements of ecological niche models 62 figure 1. illustration of the scale-dependent nature of the differences between ecological niche modeling (enm) and distribution modeling (dm). the 3 interacting sets of factors described in an earlier paper (soberón and peterson 2005) are shown as the two ‘worlds’ would presuppose—enm anticipates a broad biogeographic extent, presenting a highly structured landscape with extensive effects of accessibility, whereas dm focuses on a much finer scale, and as such is less concerned with issues of accessibility. in the enm world diagram, the blue area represents the potential distribution of the species, and the black area the hypothesized actual distribution. abiotic nicheabiotic niche biotic interactionsbiotic interactionsaccessibility enm world dm world abiotic accessibility biotic abiotic accessibility biotic abiotic nicheabiotic niche biotic interactionsbiotic interactionsaccessibility peterson – uses and requirements of ecological niche models 63 developed without benefit of computer aids for analysis, was to compare environments inside and outside of species’ distributions, asking what is different. what is more, grinnell fully appreciated the independent nature of ecological needs, versus barriers that can interrupt or truncate distributions (grinnell 1914). the concept of an ecological niche has— obviously—evolved quite a bit since grinnell. it was generalized to include biotic as well as abiotic dimensions, evolving to fit into modern community ecology and associated bodies of theory via treatments by several key workers (hutchinson 1957; macarthur 1972). curiously, though, more modern conceptualizations of ecological niches have become successively less useful for geographic, continental-scale views of species’ ecological requirements. as such, in enm applications, a grinnellian view of niches has generally been adopted—a species’ ecological niche can be defined as the set of conditions that permits it to maintain populations without immigrational subsidy. the idea that ecological niches place constraints on species’ geographic distributions hearkens directly back to the very definition of ecological niches by grinnell himself. however, this ‘constraint’ requires some clarification— particularly in light of modern population biological theory regarding metapopulations (pulliam 1988; pulliam 2000), in which some suitable areas are expected to be uninhabited, and source-sink dynamics may at times place populations in unsuitable conditions. certainly, an appreciation of the basic tenets of historical biogeography would suggest that species will not inhabit all areas that meet their niche requirements—rather, barriers to dispersal will often restrict species to a subset of these areas (peterson 2003a; peterson et al. 1999). moreover, the question remains as to whether species are restricted to this particular set of conditions only in the here and now, or are ecological niches evolved characteristics of species that would show inertia (conservatism) over evolutionary time periods? if the latter were the case, enms could offer considerable predictive ability for understanding the distribution of that species and perhaps related species. the roots of the answer to this question came from a paper on beech (fagus spp.) distributions in europe and north america, in which distributions of north american and european beech species were compared with respect to climatic dimensions; the result was that considerable coincidence exists between the two species (huntley et al. 1989). a more recent paper revisited this idea in a broader suite of species (peterson et al. 1999), testing coincidence of ecological niche dimensions in 37 sister species pairs separated across the isthmus of tehuantepec, in southern mexico—once again, each species’ ecological characteristics were highly predictive of those of its sister species. significantly, this interpredictivity among sister species broke down when confamilial species pairs were compared (peterson et al. 1999), suggesting the obvious— that evolutionary conservatism of ecological niches is not absolute, and that they do evolve on broader time scales (martínez-meyer 2002; wiens and graham 2005). further evidence for the stable nature of the constraint on geographic potential by ecological niches can be drawn from two additional types of studies. first, under rare circumstances, data availability permits direct before-and-after characterization of distributions for single species; take, for example, recent papers showing significant predictivity of species’ distributions across major events of change, such as the end of the pleistocene (martínez-meyer and peterson 2006; martínez-meyer et al. 2004a). finally, additional evidence comes from studies of invasive species, in which species are transplanted to a distinct geographic and community context. although it has been suggested based on theoretical musing and limited laboratory experiments that shifting species’ interactions would confound any possible predictivity (in this case in the context of anticipating climate change effects on species’ distributions) (davis et al. 1998), numerous studies have successfully predicted the invasive distributional potential of species based on native-range ecological characteristics (beerling et al. 1995; higgins et al. 1999; honig et al. 1992; iguchi et al. 2004; panetta and dodd 1987; papes and peterson 2003; peterson 2003a; peterson et al. 2003a; peterson and robins 2003; peterson et al. 2003b; peterson and vieglais 2001; richardson and mcmahon 1992; scott and panetta 1993; skov 2000; sutherst et al. 1999; zalba et al. 2000). peterson – uses and requirements of ecological niche models 64 hence, a diverse and growing body of evidence supports the idea that ecological niche evolution is conservative over short-to-medium periods of evolutionary time, and that models of ecological niches of species can hold significant predictive power for a variety of geographic and ecological phenomena related to biodiversity. recent theoretical treatments suggest that such should be the case—that, under many circumstances, ecological niche characteristics should not prove particularly labile in their evolution (brown and pavlovic 1992; holt 1996; holt and gaines 1992; holt and gomulkiewicz 1996; kawecki 1995). hence, in practice as well as in theory, ecological niches appear to represent long-term stable constraints on the geographic potential of species. functionalities and possibilities the following is a set of examples of applications to which enm approaches have been put. these uses are summarized in table 1, examples and brief discussion are provided in the text that follows. • understand ecological requirements of species too often, elements of biodiversity are so poorly known that a key first step is simply that of understanding the basic ecological dimensions that are relevant to a species’ geographic distribution. that is to say, for the vast majority of species, nothing more is known than a few geographic occurrences— effectively ‘dots on maps.’ all of the rich detail of natural history, ecology, and behavior can be essentially unknown, but some information can be inferred from ecological niche models. several examples of this sort of study have been published (austin and meyers 1996; costa et al. 2002; guisan and hofer 2003; hirzel et al. 2002; luoto et al. 2006; peterson et al. 2004a; ron 2005). • understand distributions, biogeography and dispersal barriers enm techniques also have considerable potential in identifying geographic phenomena that limit species’ distributional potential. as such, enm provides a tool by which the biogeography of species can be illuminated, providing information about species’ distributions that is otherwise basically unavailable. this use of enm hearkens directly back to grinnell’s original efforts. examples of this sort of enm application are numerous (anderson et al. 2002a; graham et al. 2004; manel et al. 1999; pearce et al. 2001; robertson et al. 2004; rojas-soto et al. 2003; skidmore et al. 1996; svenning and skov 2004). further extensions of these applications has addressed the seasonal distributions of migratory species (joseph 2003; joseph and stockwell 2000; martínezmeyer et al. 2004b; nakazawa et al. 2004), detection of species’ interactions (anderson et al. 2002b), and fine-scale temporal distributions of ephemeral species (peterson et al. 2005a). • find unknown populations and species enm, in its simplest manifestations, provides a framework by which one can interpolate between known populations of a species to anticipate existence of other, unknown populations. some species are sufficiently poorly known, or are sufficiently endangered, that encountering new populations can make a clear difference in understanding their distributions and in planning their conservation. including dimensions of conservatism of ecological niche evolution, this same reasoning can be used to predict the geographic distributions of unknown species closely related to known species. studies of this sort are growing in number (bourg et al. 2005; raxworthy et al. 2003). • identify sites for translocations and reintroductions recent discussions have noted that translocations and reintroductions of species closely resemble species’ invasions (bright and smithson 2001). that is, these deliberate introductions will work only to the extent that the species encounters appropriate conditions in the new landscape, and to the extent that all of the other factors affecting invasion success also coincide (e.g., demographic effects, biotic interactions). nonetheless, previous studies have focused mainly on factors affecting success of ‘establishable’ populations (armstrong and ewen 2002; carroll et al. 2003; howells and edwards-jones 1997; mccallum et al. 1995; nolet and baveco 1996; schadt et al. 2002; south et al. 2000; south et al. 2001; southgate and possingham 1995)—few have asked the question of what parts of the landscape are suitable for establishment. as such, enm provides a framework within which areas may be evaluated for their potential suitability for establishment of populations of species under intensive conservation management. examples of this sort of enm application are now beginning to appear (danks and klein 2002; mladenoff et al. 1995; peterson et al. 2006a). • conservation planning and reserve system design many exciting advances have been developed recently for prioritizing areas based on patterns of species’ occurrences (prendergast et al. 1999; pressey 1994; williams et al. 1996). these approaches peterson – uses and requirements of ecological niche models 65 generally focus on the challenge of assembling optimal suites of areas for priority conservation action, and as such represent important new tools in the conservation realm. nonetheless, the quality of the results of these analyses depends critically on the quality of the distributional information that is fed into them. given the existence of biases and gaps in the existing sampling, a modeling step to improve the picture of species’ distributional areas is warranted, and several efforts have now taken this step (araújo et al. 2005b; godown and peterson 2000; loiselle et al. 2003; peterson et al. 2000; sánchez-cordero et al. 2005a; wilson et al. 2005). further advances in this realm can be derived from use of niche models to identify areas of high probability of population persistence (araújo et al. 2002) and identifying optimal dispersal corridor locations (williams et al. 2005); indeed now these place-prioritization algorithms are able to base analyses on probability data (rather than just presence-absence data) (araújo and williams 2000; cabeza et al. 2004; williams and araújo 2002). an interesting discussion of issues related to these applications has recently been published (rondinini et al. 2006). • predict effects of habitat loss species appear often to obey different suites of environmental factors at different spatial scales (ortega-huerta and peterson 2003). that is to say, they may seek optimal suites of climatic conditions at relatively coarse conditions, but may respond to land cover type or soil type at finer scales (coudin et al. 2006; midgley et al. 2003), and to food distributions at micro-scales. as such, at times, investigators may be able to model species’ niches at coarse scales, but then use additional information (e.g., land cover type) to refine the model’s predictions. to the extent that species’ responses to these finer-scale phenomena remain constant over time, then these models can be used to anticipate future distributional potential in the face of changing patterns of habitat distribution and land use. explorations of applying these ideas in an enm framework have been developed and applied now in several situations (peterson et al. 2006b; sánchez-cordero et al. 2005a; sánchez-cordero et al. 2005b; thuiller et al. 2004). • predict potential for species’ invasions this enm application is perhaps that which has seen the most intensive exploration by many laboratories, with examples developed for numerous taxa worldwide. here, the idea is that—given apparently widespread evolutionary conservatism in ecological niche characteristics—species will often ‘obey’ the same set of ecological rules on invaded distributional areas as they do on their native distributional areas. as such, the geographic potential of invasive species is often quite predictable, based on their geographic and ecological distributions on their native distributional areas (beerling et al. 1995; higgins et al. 1999; hinojosa-díaz et al. 2005; hoffmann 2001; honig et al. 1992; iguchi et al. 2004; panetta and dodd 1987; papes and peterson 2003; peterson 2003a; peterson et al. 2003a; peterson and robins 2003; peterson et al. 2003b; peterson and vieglais 2001; podger et al. 1990; richardson and mcmahon 1992; robertson et al. 2004; sindel and michael 1992; skov 2000; sutherst et al. 1999; welk et al. 2002; zalba et al. 2000), although the factors that make a species invasive are clearly more complex than just niche considerations (thuiller et al. 2005b). • predict climate change effects to the extent that species’ ecological niches remain fairly constant, and do not evolve to meet changing conditions, it is possible to project present-day niche models onto future conditions of climate, as represented in large-scale climate models (flato et al. 1999; mcfarlane et al. 1992; pope et al. 2002). these projections, under assumptions made fairly explicit, provide hypotheses of species’ potential geographic distributions and how they will change over the next few decades of evolving world climates, and this field has now seen extensive activity (araújo et al. 2005a; araújo et al. 2006; berry et al. 2002; carey and brown 1994; erasmus et al. 2002; gottfried et al. 1999; huntley et al. 1995; kadmon and heller 1998; malanson et al. 1992; pearson and dawson 2003; pearson et al. 2002; peterson 2003b; peterson et al. 2004b; peterson et al. 2002; peterson et al. 2001; peterson and shaw 2003; peterson et al. 2005b; porter et al. 2000; price 2000; roura-pascual et al. 2005; sykes et al. 1996; thuiller et al. 2005a), including retro-projections aimed at reconstructing distributions in the pleistocene (hilbert et al. 2004; hugall et al. 2002; martínez-meyer and peterson 2006; martínezmeyer et al. 2004a). the complexities of these projections, however, are only beginning to be appreciated, given species’ responses to other factors such as atmospheric gas composition (thuiller et al. 2006). distributional modeling distribution modeling, as it is usually depicted by its proponents, appears to constitute an effort not to overinterpret conceptually the models resulting in what would otherwise be enm. that is, the usual argument goes, because the occurrence data on which models are trained are sampled from the actual distribution of the species, it is improbable that the model can say anything about the potential distribution or fundamental peterson – uses and requirements of ecological niche models 66 ecological niche of the species. while these arguments have some merit if interactions were to occur universally and in ecological dimensions (rather than in geographic dimensions—‘who gets there first’), i argue that such situations are rare, although this is clearly a topic for future detailed analyses. that is, i argue that many species interactions take place in geographic space, and that they are far from universal: in this way, ecological potential of species is manifest in some portion of the range, and good (dense, representative) sampling will usually reveal that potential (araújo and guisan 2006): a worked example is provided in a recent publication (anderson et al. 2002b), and the theoretical issues have also been reviewed recently (soberón and peterson 2005). another interpretation of the enm-dm debate focuses on the scale of the inquiry (see figure 1). here, whereas enm proponents would anticipate a full-species’-range scale for a study, dm proponents may be willing to examine a much smaller portion of species’ distributions. as such, accessibility considerations become unimportant in dm (as the entire study area tends to be available to the species), whereas enm—which is often developed at broader scales (such as that of an entire species’ distribution) must take them into account more explicitly. dm, given that it is able to rely on assumptions such as that dispersal limitation is relatively immaterial in sculpting species’ distributions, is often able to produce more accurate local predictions within regions. however, dm results are often limited by spatial references in environmental data sets used to build models, and the limited spatial scales of inquiry may often provide incomplete characterization of niches of species. enm, on the other hand, uses the conservative nature of ecological niche characteristics of species to open additional suites of predictive capabilities, and provides a clearer interpretation of causal forces affecting species’ distributions. however, enm models may be difficult to test and validate because special assumptions are required to turn potential distribution estimates into actual distribution estimates. discussion and conclusions the diversity of applications discussed above suggests that important methodological requirements may also differ among applications. that is, different applications may require distinct assumptions and interpretations to make possible a particular result. table 1 details a number of considerations that may be relevant in this respect. these points are relevant also to the challenge of choosing among the many options for creating ecological niche models in the first place. the community of investigators that use these approaches use many different inferential techniques. these techniques include very simple range-rule approaches that detect limits along independent environmental axes, multivariate statistical approaches that fit response curves in environmental space, and evolutionary computing approaches that explore solution space randomly to produce ‘best’ solutions to the challenge. these different techniques interact with the uses listed in table 1 as well—some techniques may be better suited to some uses. qualities of interest regarding uses and a brief discussion of their interaction with techniques and uses, are as follows: • form of prediction (e.g., binary, ranked, absolute) – some uses (e.g., identification of sites for translocations and reintroductions, conservation planning and reserve system design) require an absolute, rather than a relative, answer to the question of suitability. that is to say, it does no good to reintroduce a species to the best site in a landscape if that site is not highly suitable for the species to become established. these uses will thus require inferential approaches that have some mode of calibration to indicate which sites are equally suitable as those where the species does maintain or has maintained populations. • grain required – grain required for usable predictions varies from a grain relevant to individual or at least population distributions (e.g., identification of sites for translocations and reintroductions) up to broader-scale views (e.g., prediction of climate change effects). although the details depend on the extent of the area under consideration, this factor will clearly demand that inferential algorithms be able to deal with large data sets in developing models. • causal variables needed, or surrogates ok? – the issue of what is a causal variable is not simple. peterson – uses and requirements of ecological niche models 67 causation is not a yes-no issue, but rather is an issue of relative immediacy. for example, humidity may be directly related to survival, but the most proximate manifestation of ‘humidity’ may be whether a species’ eggs desiccate before they can hatch. nonetheless, clearly, some variables are likely to be surrogates for ‘real’ causal variables. similarly, many models use elevation as an independent variable—elevation nonetheless is really an excellent surrogate for temperature, but cannot be used in place of temperature, e.g., when climates change, because elevation would have a changing meaning in different climate regimes. the different uses break down about evenly as to whether surrogates are acceptable, or whether causal variables are needed (table 1). • need model response curve or parameter retrieval – some modeling approaches (e.g., multivariate statistical approaches) are able to reconstruct the roles of individual independent variables in model predictions, whereas others (e.g., evolutionary computing approaches) must reconstruct this information in a more post hoc manner. although both approaches can potentially ‘get at’ the issue of the shape of the response curve to particular environmental variables, the former are clearly more convenient. • error needs – two general types of error are possible in modeling species’ niches and distributions—omission (leaving out of the prediction areas that are within the species’ ecological potential), and commission (including in the prediction areas outside of the species’ ecological potential). different uses of modeling emphasize different error components as more or less important—for example, understanding ecological requirements of species would emphasize minimizing both error components simultaneously, whereas identification of sites for reintroductions would require low commission error (high cost of error as to which sites are suitable). • potential distributional model versus realized distribution? – several of the uses detailed in table 1 are clearly functions of potential distributions, rather than actual distributions of species. for example, all applications to predicting species’ invasions would perforce have to be based on estimates of potential distributions, whereas other applications (e.g., conservation planning and reserve design) would either demand actual distributional estimates or potential distributional estimates refined by specific assumptions regarding interactions with other species and dispersal abilities. • uncertainty estimates needed? – finally, because of high costs of being wrong in bases for decisions, applications such as identification of sites for translocations and reintroductions and conservation planning and reserve system design require careful estimates of uncertainty associated with predictions. such estimates are most easily drawn from multivariate statistical approaches, although they are possible using other approaches as well. in sum, the world of modeling ecological and geographic distributions of species is simultaneously complex (see the preceding list of considerations in choosing modeling methods) and promising (see the exciting list of uses farther above). this field is clearly just recently achieving much of the breadth of its potential, and as a consequence is seeing increasing interest and application. this contribution is intended principally as a platform for discussion. that is, much of what is said regarding particular applications and uses and their requirements for inferential algorithms is opinion, and is intended to spark discussion and debate, laying out one point of view. my hope is that this contribution can serve to initiate such a debate, and by this means improve the conceptual platform on which an even-more-vibrant field of inquiry can be based. acknowledgments this work was supported generously by the national center for ecological analysis and synthesis, university of california, santa barbara, and benefitted enormously from discussions and debates with members of the working group on predicting species’ distributions, particularly with simon ferrier. guy midgely and miguel b. araújo provided very helpful comments on an earlier version of this paper. references anderson, r. p., m. gómez-laverde, and a. t. peterson. 2002a. geographical distributions of spiny pocket mice in south america: insights from predictive models. global ecology and biogeography 11:131-141. anderson, r. p., a. t. peterson, and m. gómezlaverde. 2002b. using niche-based gis modeling to test geographic predictions of competitive exclusion and competitive release in south american pocket mice. oikos 93:3-16. peterson – uses and requirements of ecological niche models 68 araújo, m. b., and a. guisan. 2006. five (or so) challenges for species distribution modelling. journal of biogeography 33:1677-1688. araújo, m. b., and r. g. pearson. 2005. equilibrium of species' distributions with climate. ecography 28:693-695. araújo, m. b., r. g. pearson, w. thuiller, and m. erhard. 2005a. validation of species-climate impact models under climate change. global change biology 11:1504-1513. araújo, m. b., w. thuiller, and r. g. pearson. 2006. climate warming and the decline of amphibians and reptiles in europe. journal of biogeography 33:17121728. araújo, m. b., w. thuiller, p. h. williams, and i. reginster. 2005b. downscaling european species atlas distributions to a finer resolution: implications for conservation planning. global ecology and biogeography 14:17-30. araújo, m. b., and p. h. williams. 2000. selected areas for species persistence using occurrence data. biological conservation 96:331-345. araújo, m. b., p. h. williams, and r. j. fuller. 2002. dynamics of extinction and the selection of nature reserves. proceedings of the royal society of london b 269:1971-1980. armstrong, d. p., and j. g. ewen. 2002. dynamics and viability of a new zealand robin population reintroduced to regenerating fragmented habitat. conservation biology 16:1074-1085. austin, m. p., and j. a. meyers. 1996. current approaches to modelling the environmental niche of eucalypts: implications for management of forest biodiversity. forest ecology and management 85:95106. beerling, d. j., b. huntley, and j. p. bailey. 1995. climate and the distribution of fallopia japonica: use of an introduced species to test the predictive capacity of response surfaces. journal of vegetation science 6:269-282. berry, p. m., t. p. dawson, p. a. harrison, and r. g. pearson. 2002. modelling potential impacts of climate change on the bioclimatic envelope of species in britain and ireland. global ecology and biogeography 11:453-462. bourg, n. a., w. j. mcshea, and d. e. gill. 2005. putting a cart before the search: successful habitat prediction for a rare forest herb. ecology 86:27932804. bright, p. w., and t. j. smithson. 2001. biological invasions provide a framework for reintroductions: selecting areas in england for pine marten releases. biodiversity and conservation 10:1247-1265. brown, j. s., and n. b. pavlovic. 1992. evolution in heterogeneous environments: effects of migration on habitat specialization. evolutionary ecology 6:360382. cabeza, m., m. b. araújo, r. j. wilson, c. d. thomas, m. j. r. cowley, and a. moilanen. 2004. combining probabilities of occurrence with spatial reserve design. journal of applied ecology 41:252262. carey, p. d., and n. j. brown. 1994. the use of gis to identify sites that will become suitable for a rare orchid, himantoglossum hircinum l., in a future changed climate. biodiversity letters 2:117-123. carroll, c., m. k. phillips, n. h. schumaker, and d. w. smith. 2003. impacts of landscape change on wolf restoration success: planning a reintroduction program based on static and dynamic spatial models. conservation biology 17:536-548. chase, j. m., and m. a. leibold. 2003. ecological niches: linking classical and contemporary approaches. university of chicago press, chicago. costa, j., a. t. peterson, and c. b. beard. 2002. ecological niche modeling and differentiation of populations of triatoma brasiliensis neiva, 1911, the most important chagas disease vector in northeastern brazil (hemiptera, reduviidae, triatominae). american journal of tropical medicine & hygiene 67:516-520. coudin, c., j. c. gégout, c. piedallu, and j. c. rameau. 2006. soil nutritional factors improve plant species distribution models: an illustration with acer campestre (l.) in france. journal of biogeography, in press. danks, f. s., and d. r. klein. 2002. using gis to predict potential wildlife habitat: a case study of muskoxen in northern alaska. international journal of remote sensing 23:4611-4632. davis, a. j., l. s. jenkinson, j. h. lawton, b. shorrocks, and s. wood. 1998. making mistakes when predicting shifts in species range in response to global warming. nature 391:783-786. erasmus, b. f. n., a. s. van jaarsveld, s. l. chown, m. kshatriya, and k. j. wessels. 2002. vulnerability of south african animal taxa to climate change. global change biology 8:679-693. flato, g. m., g. j. boer, w. g. lee, n. a. mcfarlane, d. ramsden, m. c. reader, and a. j. weaver. 1999. the canadian center for climate modelling and analysis global coupled model and its climate. climate dynamics 16:451-467. godown, m. e., and a. t. peterson. 2000. preliminary distributional analysis of u.s. endangered bird species. biodiversity and conservation 9:1313-1322. gottfried, m., h. pauli, k. reiter, and g. grabherr. 1999. a fine-scaled predictive model for changes in species distribution patterns of high mountain plants induced by climate warming. diversity and distributions 5:241-251. peterson – uses and requirements of ecological niche models 69 graham, c. h., s. r. ron, j. c. santos, c. j. schneider, and c. moritz. 2004. integrating phylogenetics and environmental niche models to explore speciation mechanisms in dendrobatid frogs. evolution 58:17811793. grinnell, j. 1914. barriers to distribution as regards birds and mammals. american naturalist 48:248-254. grinnell, j. 1917. field tests of theories concerning distributional control. american naturalist 51:115128. grinnell, j. 1924. geography and evolution. ecology 5:225-229. guisan, a., and u. hofer. 2003. predicting reptile distributions at the mesoscale: relation to climate and topography. journal of biogeography 30:1233-1243. guisan, a., and n. e. zimmermann. 2000. predictive habitat distribution models in ecology. ecological modelling 135:147-186. higgins, s. i., d. m. richardson, r. m. cowling, and t. h. trinder-smith. 1999. predicting the landscapescale distribution of alien plants and their threat to plant diversity. conservation biology 13:303-313. hilbert, d. w., m. bradford, t. parker, and d. a. westcott. 2004. golden bowerbird (prionodura newtonia) habitat in past, present and future climates: predicted extinction of a vertebrate in tropical highlands due to global warming. biological conservation 116:367-377. hinojosa-díaz, i. a., o. yáñez-ordóñez, g. chen, and a. t. peterson. 2005. the north american invasion of the giant resin bee (hymenoptera: megachilidae). journal of hymenoptera research 14:69-77. hirzel, a. h., j. hausser, d. chessel, and n. perrin. 2002. ecological-niche factor analysis: how to compute habitat-suitability maps without absence data? ecology 83:2027-2036. hoffmann, m. h. 2001. the distribution of senecio vulgaris: capacity of climatic range models for predicting adventitious ranges. flora 196/5:395-403. holt, r. d. 1996. adaptive evolution in source-sink environments: direct and indirect effects of densitydependence on niche evolution. oikos 75:182-192. holt, r. d., and m. s. gaines. 1992. analysis of adaptation in heterogeneous landscapes: implications for the evolution of fundamental niches. evolutionary ecology 6:433-447. holt, r. d., and r. gomulkiewicz. 1996. the evolution of species' niches: a population dynamic perspective. pp. 25-50 in h. g. othmer, f. r. adler, m. a. lewis and j. c. dallon, eds. case studies in mathematical modeling: ecology, physiology and cell biology. prentice-hall, saddle river, n.j. honig, m. a., r. m. cowling, and d. m. richardson. 1992. the invasive potential of australian banksias in south-african fynbos--a comparison of the reproductive potential of banksia ericifolia and leucadendron laureolum. australian journal of ecology 17:305-314. howells, o., and g. edwards-jones. 1997. a feasibility study of reintroducing wild boar sus scrofa to scotland: are existing woodlands large enough to support minimum viable populations. biological conservation 81:77-89. hugall, a., c. moritz, a. moussalli, and j. stanisic. 2002. reconciling paleodistribution models and comparative phylogeography in the wet tropics rainforest land snail gnarosophia bellendenkerensis (brazier 1875). proceedings of the national academy of sciences usa 99:6112-6117. huntley, b., p. j. bartlein, and i. c. prentice. 1989. climatic control of the distribution and abundance of beech (fagus l.) in europe and north america. journal of biogeography 16:551-560. huntley, b., p. m. berry, w. cramer, and a. p. mcdonald. 1995. modelling present and potential future ranges of some european higher plants using climate response surfaces. journal of biogeography 22:967-1001. hutchinson, g. e. 1957. concluding remarks. cold spring harbor symposia on quantitative biology 22:415-427. iguchi, k., k. matsuura, k. mcnyset, a. t. peterson, r. scachetti-pereira, k. a. powers, d. a. vieglais, e. o. wiley, and t. yodo. 2004. predicting invasions of north american basses in japan using native range data and a genetic algorithm. transactions of the american fisheries society 133:845-854. joseph, l. 2003. predicting distributions of south american migrant birds in fragmented environments: a possible approach based on climate. pp. 263-284 in g. a. bradshaw and p. marquet, eds. how landscapes change: human disturbance and ecosystem fragmentation in the americas. springerverlag, inc., berlin. joseph, l., and d. r. b. stockwell. 2000. temperature-based models of the migration of swainson's flycatcher (myiarchus swainsoni) across south america: a new use for museum specimens of migratory birds. proceedings of the academy of natural sciences of philadelphia 150:293-300. kadmon, r., and j. heller. 1998. modelling faunal responses to climatic gradients with gis: land snails as a case study. journal of biogeography 25:527-539. kawecki, t. j. 1995. demography of source-sink populations and the evolution of ecological niches. evolutionary ecology 9:38-44. latimer, a. m., s. wu, a. e. gelfand, and j. a. silander jr. 2006. building statistical models to analyze species distributions. ecological applications 16:33-50. loiselle, b. a., c. a. howell, c. h. graham, j. m. goerck, t. m. brooks, k. g. smith, and p. h. peterson – uses and requirements of ecological niche models 70 williams. 2003. avoiding pitfalls of using species distribution models in conservation planning. conservation biology 17:1591-1600. luoto, m., r. k. heikkinen, j. pöyry, and k. saarinen. 2006. determinants of biogeographical distribution of butterflies in boreal regions. journal of biogeography 33:1764-1778. macarthur, r. 1972. geographical ecology. princeton university press, princeton, n.j. malanson, g. p., w. e. westman, and y.-l. yan. 1992. realized versus fundamental niche functions in a model of chaparral response to climatic change. ecological modelling 64:261-277. manel, s., j. m. dias, s. t. buckton, and s. j. ormerod. 1999. alternative methods for predicting species distribution: an illustration with himalayan river birds. journal of applied ecology 36:734-747. martínez-meyer, e. 2002. evolutionary trends in ecological niches of species. ph.d. thesis, department of geography, university of kansas, lawrence, kansas. martínez-meyer, e., and a. t. peterson. 2006. conservatism of ecological niche characteristics in north american plant species over the pleistocene-torecent transition. journal of biogeography 33:17791789. martínez-meyer, e., a. t. peterson, and w. w. hargrove. 2004a. ecological niches as stable distributional constraints on mammal species, with implications for pleistocene extinctions and climate change projections for biodiversity. global ecology and biogeography 13:305-314. martínez-meyer, e., a. t. peterson, and a. g. navarrosigüenza. 2004b. evolution of seasonal ecological niches in the passerina buntings (aves: cardinalidae). proceedings of the royal society b 271:1151-1157. mccallum, h., p. timmers, and s. hoyle. 1995. modeling the impact of predation on reintroductions of bridled nailtail wallabies. wildlife research 22:163-171. mcfarlane, n. a., g. j. boer, j.-p. blanchet, and m. lazare. 1992. the canadian climate centre secondgeneration general circulation model and its equilibrium climate. journal of climate 5:1013-1044. midgley, g. f., l. hannah, d. millar, w. thuiller, and a. booth. 2003. developing regional and specieslevel assessments of climate change impacts on biodiversity in the cape floristic region. biological conservation 112:87-97. mladenoff, d. j., t. a. sickley, r. g. haight, and a. p. wydeven. 1995. a regional landscape analysis and prediction of favorable gray wolf habitat in the northern great lakes region. conservation biology 9:279-294. nakazawa, y., a. t. peterson, e. martínez-meyer, and a. g. navarro-sigüenza. 2004. seasonal niches of nearctic-neotropical migratory birds: implications for the evolution of migration. auk 121:610-618. nolet, b. a., and j. m. baveco. 1996. development and viability of a translocated beaver castor fiber population in the netherlands. biological conservation 75:125-137. ortega-huerta, m. a., and a. t. peterson. 2003. effects of geographic scale on analyzing associations between regional habitats and distribution patterns of mexican birds. anales del instituto de biologia, u.n.a.m. 74:203-210. panetta, f. d., and j. dodd. 1987. bioclimatic prediction of the potential distribution of skeleton weed chondrilla juncea l. in western australia. journal of the australian institute of agricultural science 53:11-16. papes, m., and a. t. peterson. 2003. predicting the potential invasive distribution for eupatorium adenophorum spreng. in china. journal of wuhan botanical research 21:137-142. pearce, j., s. ferrier, and d. scotts. 2001. an evaluation of the predictive performance of distributional models for flora and fauna in north-east new south wales. journal of environmental management 62:171-184. pearson, r. g., and t. p. dawson. 2003. predicting the impacts of climate change on the distribution of species: are bioclimate envelope models useful? global ecology and biogeography 12:361-371. pearson, r. g., t. p. dawson, p. m. berry, and p. a. harrison. 2002. species: a spatial evaluation of climate impact on the envelope of species. ecological modelling 154:289-300. peterson, a. t. 2003a. predicting the geography of species' invasions via ecological niche modeling. quarterly review of biology 78:419-433. peterson, a. t. 2003b. projected climate change effects on rocky mountain and great plains birds: generalities of biodiversity consequences. global change biology 9:647-655. peterson, a. t., j. t. bauer, and j. n. mills. 2004a. ecological and geographic distribution of filovirus disease. emerging infectious diseases 10:40-47. peterson, a. t., s. l. egbert, v. sánchez-cordero, and k. p. price. 2000. geographic analysis of conservation priority: endemic birds and mammals in veracruz, mexico. biological conservation 93:85-94. peterson, a. t., c. martínez-campos, y. nakazawa, and e. martínez-meyer. 2005a. time-specific ecological niche modeling predicts spatial dynamics of vector insects and human dengue cases. transactions of the royal society of tropical medicine and hygiene 99:647-655. peterson – uses and requirements of ecological niche models 71 peterson, a. t., e. martínez-meyer, c. gonzálezsalazar, and p. hall. 2004b. modeled climate change effects on distributions of canadian butterfly species. canadian journal of zoology 82:851-858. peterson, a. t., e. martínez-meyer, j. servin, and l. f. kiff. 2006a. ecological niche modelling and strategizing for species reintroductions. oryx in press peterson, a. t., m. a. ortega-huerta, j. bartley, v. sanchez-cordero, j. soberon, r. h. buddemeier, and d. r. b. stockwell. 2002. future projections for mexican faunas under global climate change scenarios. nature 416:626-629. peterson, a. t., m. papes, and d. a. kluza. 2003a. predicting the potential invasive distributions of four alien plant species in north america. weed science 51:863-868. peterson, a. t., and c. r. robins. 2003. using ecological-niche modeling to predict barred owl invasions with implications for spotted owl conservation. conservation biology 17:1161-1165. peterson, a. t., v. sánchez-cordero, e. martínezmeyer, and a. g. navarro-sigüenza. 2006b. tracking population extirpations via melding ecological niche modeling with land-cover information. ecological modelling 195:229-236. peterson, a. t., v. sanchez-cordero, j. soberon, j. bartley, r. h. buddemeier, and a. g. navarrosiguenza. 2001. effects of global climate change on geographic distributions of mexican cracidae. ecological modelling 144:21-30. peterson, a. t., r. scachetti-pereira, and d. a. kluza. 2003b. assessment of invasive potential of homalodisca coagulata in western north america and south america. biota neotropica 3:online journal. peterson, a. t., and j. j. shaw. 2003. lutzomyia vectors for cutaneous leishmaniasis in southern brazil: ecological niche models, predicted geographic distributions, and climate change effects. international journal of parasitology 33:919-931. peterson, a. t., j. soberón, and v. sánchez-cordero. 1999. conservatism of ecological niches in evolutionary time. science 285:1265-1267. peterson, a. t., h. tian, e. martínez-meyer, j. soberón, v. sánchez-cordero, and b. huntley. 2005b. modeling distributional shifts of individual species and biomes. pp. 211-228 in t. e. lovejoy and l. hannah, eds. climate change and biodiversity. yale university press, new haven, conn. peterson, a. t., and d. a. vieglais. 2001. predicting species invasions using ecological niche modeling. bioscience 51:363-371. podger, f. d., d. c. mummery, c. r. palzer, and m. j. brown. 1990. bioclimatic analysis of the distribution of damage to native plants in tasmania by phytophthora cinnamomi. australian journal of ecology 15:281-290. pope, v. d., m. l. gallani, v. j. rowntree, and r. a. stratton. 2002. the impact of new physical parametrizations in the hadley centre climate model hadam3. hadley centre for climate prediction and research, bracknell, berks, uk. porter, w. p., s. budaraju, w. e. stewart, and n. ramankutty. 2000. calculating climate effects on birds and mammals: impacts on biodiversity, conservation, population parameters, and global community structure. american zoologist 40:597630. prendergast, j. r., r. m. quinn, and j. h. lawton. 1999. the gaps between theory and practice in selecting nature reserves. conservation biology 13:484-492. pressey, r. 1994. ad hoc reservations: forward or backward steps in developing representative reserve systems. conservation biology 8:662-668. price, j. 2000. modeling the potential impacts of climate change on the summer distributions of massachusetts passerines. bird observer 28:224-230. pulliam, h. r. 1988. sources, sinks and population regulation. american naturalist 132:652-661. pulliam, h. r. 2000. on the relationship between niche and distribution. ecology letters 3:349. raxworthy, c. j., e. martínez-meyer, n. horning, r. a. nussbaum, g. e. schneider, m. a. ortega-huerta, and a. t. peterson. 2003. predicting distributions of known and unknown reptile species in madagascar. nature 426:837-841. richardson, d. m., and j. p. mcmahon. 1992. a bioclimatic analysis of eucalyptus nintens to identify potential planting regions in southern africa. south african journal of science 88:380-387. robertson, m. p., m. h. villet, and a. r. palmer. 2004. a fuzzy classification technique for predicting species' distributions: applications using invasive alien plants and indigenous insects. diversity and distributions 10:461-474. rojas-soto, o. r., o. alcantara-ayala, and a. g. navarro. 2003. regionalization of the avifauna of the baja california peninsula, mexico: a parsimony analysis of endemicity and distributional modelling approach. journal of biogeography 30:449-461. ron, s. r. 2005. predicting the distribution of the amphibian pathogen batrachochytrium dendrobatidis in the new world. biotropica 37:209-221. roura-pascual, n., a. suarez, c. gómez, p. pons, y. touyama, a. l. wild, and a. t. peterson. 2005. geographic potential of argentine ants (linepithema humile mayr) in the face of global climate change. proceedings of the royal society of london b 271:2527-2535. sánchez-cordero, v., v. cirelli, m. munguia, and s. sarkar. 2005a. place prioritization for biodiversity peterson – uses and requirements of ecological niche models 72 representation using species' ecological niche modeling. biodiversity informatics 2:11-23. sánchez-cordero, v., p. illoldi-rangel, m. linaje, s. sarkar, and a. t. peterson. 2005b. deforestation and extant distributions of mexican endemic mammals. biological conservation 126:465-473. schadt, s., f. knauer, p. kaczensky, e. revilla, t. wiegand, and l. trepl. 2002. rule-based assessment of suitable habitat and patch connectivity for the eurasian lynx. ecological applications 12:1469-1483. scott, j. k., and f. d. panetta. 1993. predicting the australian weed status of southern african plants. journal of biogeography 20:87-93. sindel, b. m., and p. w. michael. 1992. spread and potential distribution of senecio madagascariensis pior (fireweed) in australia. australian journal of ecology 17:21-26. skidmore, a. k., a. gauld, and p. walker. 1996. classification of kangaroo habitat distribution using three gis models. international journal of geographical information systems 10:441-454. skov, f. 2000. potential plant distribution mapping based on climatic similarity. taxon 49:503-515. soberón, j., and a. t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. philosophical transactions of the royal society of london b 359:689-698. soberón, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species' distributional areas. biodiversity informatics 2:1-10. south, a., s. rushton, and d. macdonald. 2000. simulating the proposed reintroduction of the european beaver (castor fiber) to scotland. biological conservation 93:103-116. south, a. b., s. p. rushton, d. w. macdonald, and r. fuller. 2001. reintroduction of the european beaver (castor fiber) to norfolk, uk: a preliminary modelling analysis. journal of zoology 254:473-479. southgate, r., and h. possingham. 1995. modeling the reintroduction of the greater bilby macrotis lagotis: using the metapopulation model analysis of the likelihood of extinction (alex). biological conservation 73:151-160. sutherst, r. w., g. f. maywald, t. yonow, and p. m. stevens. 1999. climex user guide--predicting the effects of climate on plants and animals. csiro publishing, collingwood, victoria, australia. svenning, j.-c., and f. skov. 2004. limited filling of the potential range in european tree species. ecology letters 7:565-573. sykes, m. t., i. c. prentice, and w. cramer. 1996. a bioclimatic model for the potential distributions of north european tree species under present and future climates. journal of biogeography 23:203-233. thuiller, w., m. b. araújo, and s. lavorel. 2004. do we need land-cover data to model species distributions in europe? journal of biogeography 31:353-361. thuiller, w., s. lavorel, m. b. araújo, m. t. sykes, and i. c. prentice. 2005a. climate change threats to plant diversity in europe. proceedings of the national academy of sciences usa 102:8245-8250. thuiller, w., g. f. midgely, g. o. hughes, b. bomhard, g. drew, m. c. rutherford, and f. i. woodward. 2006. endemic species and ecosystem sensitivity to climate change in namibia. global change biology 12:759-776. thuiller, w., d. m. richardson, p. pysek, g. f. midgely, g. o. hughes, and m. rouget. 2005b. global risk assessment for plant invasions--the role of climatic suitability and propagule pressure. global change biology 11:2234-2259. welk, e., k. schubert, and m. h. hoffmann. 2002. present and potential distribution of invasive garlic mustard (alliaria petiolata) in north america. diversity and distributions 8:219-233. wiens, j. j., and c. h. graham. 2005. niche conservatism: integrating evolution, ecology, and conservation biology. annual review of ecology and systematics 36:519-539. williams, p., d. gibbons, c. r. margules, a. rebelo, c. humphries, and r. pressey. 1996. a comparison of richness hotspots, rarity hotspots, and complementary areas for conserving diversity of british birds. conservation biology 10:155-174. williams, p. h., and m. b. araújo. 2002. apples, oranges, and probabilities: integrating multiple factors into biodiversity conservation with consistency. environmental modeling and assessment 7:139-151. williams, p. h., l. hannah, s. andelman, g. f. midgely, m. b. araújo, g. o. hughes, l. l. manne, e. martínez-meyer, and r. g. pearson. 2005. planning for climate change: identifying minimumdispersal corridors for the cape proteaceae. conservation biology 19:1063-1074. wilson, k. a., m. i. westphal, h. p. possingham, and j. elith. 2005. sensitivity of conservation planning to different approaches to using predicted species distribution data. biological conservation 122:99112. zalba, s. m., m. i. sonaglioni, and c. j. belenguer. 2000. using a habitat model to assess the risk of invasion by an exotic plant. biological conservation 93:203-208. microsoft word franzthau-taxonomyontology-resubmission1-nofigs _2_-nmfcorrections.doc biodiversity informatics, 7, 2010, pp. 45-66   45   biological taxonomy and ontology development: scope and limitations nico m. franz 1 and david thau 2 1department of biology, call box 9000, university of puerto rico, mayagüez, pr 00681-9000 usa, nico.franz@upr.edu; 2 department of computer science, university of california at davis, kemper hall, one shield avenue, davis, ca 9561 usa, thau@learningsite.com abstract.  the prospects of integrating full-blown biological taxonomies into an ontological reasoning framework are reviewed. traditionally ontological representations of taxonomy have adopted the model of a single and static hierarchy. this model is contrasted with a more realistic situation involving dynamic revisions of particular groups and alignments among alternative taxonomic perspectives. taxonomic practice is bound by a range of epistemological constraints and linguistic conventions that run orthogonal to the logical background from which ontological entities and relationships originate, resulting in severe challenges for ontological representation and reasoning. in particular, the purported existence of a single hierarchy in nature forces taxonomists to gradually approximate this hierarchy and make frequent rearrangements in light of new evidence. the evolvability of taxa implies that taxon-defining features may be lost in subordinate members or independently gained across multiple sections of the tree of life. as a result, many terms for phenotypic properties are phylogenetically underdetermined and have limited hierarchical transitivity. the standard approach of defining taxa both in reference to properties (intensional) and members (ostensive) undermines the individual/class dichotomy that sustains conventional ontologies, and compromises the reasoning capabilities associated with this distinction. neither the use of linnaean ranks in taxonomy nor the 250-year legacy of nomenclatural adjustments have obvious analogues in the ontological realm. lastly, the piece-meal appearance of full-blown taxonomic information makes it necessary to use expert alignments to obtain a comprehensive static perspective. in light of these limitations, research along the taxonomy/ontology interface should focus either on strictly nomenclatural entities and relationships or on ontology-driven strategies for aligning multiple taxonomies, but not on building static networks for large portions of the tree of life. the prospects of using ontology-based services in taxonomy will largely depend on the ability of the taxonomic expert community to present its products in ways that are more compatible with ontological principles than concurrent practice. key words.  classes, evolvability, individuals, intension, legacy integration, nomenclature, ontology alignment, ostension, phylogeny, taxonomic concepts. biologically oriented ontologies are regarded as essential for integrating a vast range of biological data (goble and stevens, 2008). having originated in the biomedical sector (smith et al., 2007), the development of such ontologies is now branching out into other disciplines including behavior (midford, 2004), ecology (jones et al., 2006; madin et al., 2008), evolution (mabee et al., 2007), anatomy (ramírez et al., 2007; mikó and deans, 2009; vogt et al., 2009; balhoff et al., 2010; dahdul et al., 2010; mungall et al., 2010), and taxonomy (schulz et al., 2008). nevertheless, the impact of ontology-based information services has been uneven, in part because of differences in needs and resources among disciplines. it may seem particularly intriguing that the field of systematics1 was not among the first to shift towards using ontologies, even though species names are prime vehicles for transmitting biological data and taxonomic hierarchies appear very similar to ontology networks. in spite of such an apparent match, actual representations of taxonomy in information services such as genbank are rather simplistic (cf. wheeler et al., 2008); and therefore are unable to support                                                              1 the terms systematics and taxonomy are used interchangeably herein, and each encompasses phylogenetics as a subdiscipline that informs the establishment of classifications (wheeler, 2004). biodiversity informatics, 7, 2010, pp. 45-66   46   semantically complex searches or inferences across alternative taxonomic perspectives (page 2004, 2007; thau and ludäscher, 2007; dahdul et al., 2010). this is unfortunate because a thoroughgoing ontology for biological taxonomy could play the same overarching role in the bioontological domain that taxonomy plays in biology (schulz et al., 2008). this paper examines to what extent a fullblown representation of biological taxonomy in the ontological domain is possible. using schulz et al.'s (2008) pioneering work as a point of departure, we first observe that there are two fundamentally different models to create ontological representations of taxonomy; viz. strictly nomenclatural and full-blown taxonomic representations. each model serves a different purpose and user domain. we furthermore distinguish between static and dynamic perspectives of taxonomic classifications, arguing that only the latter is well suited to represent the reasoning needs of taxonomists. we then describe a number of ontological representation challenges posed by biological taxonomy under the full-blown dynamic model; including epistemological limitations, linguistic conventions, and alignment challenges. our review concludes with a series of practical recommendations for advancing the taxonomy/ontology interface, with an emphasis on research tasks for the taxonomic expert community. nomenclatural versus full-blown taxonomic representations biological taxonomy integrates a wide variety of heterogeneous data, including information on physical specimens and voucher samples, taxonomic names and nomenclatural relationships, phenotypic and genotypic properties of taxa, and hierarchical phylogenetic groupings and classifications. some user communities are well served by focusing on a subset of this immense data pool. for instance, taxonomic experts might have great use for an ontology of strictly nomenclatural relationships (homonymy, synonymy, typification, etc.) among accepted and rejected linnaean names (scoble, 2004; kennedy et al., 2005; page, 2006; franz and peet, 2009; huber and klump, 2009). such nomenclatural entities and relationships have a tractable history, and representations of this history in an ontological network are useful – up to a point – for identifying and integrating taxonomic legacy information (berendsohn et al., 2003). nevertheless, nomenclatural relationships such as "is a misapplied name for" or "is a pro parte synonym for" are often not sufficient to permit reasoning among of alternative and succeeding taxonomic perspectives (franz, 2005a). to our knowledge no comprehensive ontology of nomenclatural entities and relationships exists, making this a worthwhile if semantically limited pursuit. other examples of ontologies with a strong taxonomic component include phenotype ontologies for model organisms (smith et al., 2007). by design, these phenotype networks fall short of representing the depth and breadth of taxonomic information across large sections of the tree of life. in what follows, we will deliberately concentrate on ontologies that aim to represent authoritative, expert-produced phylogenies and classifications, as opposed to just nomenclatural relationships. we refer to such ontologies as fullblown if they permit inferences ranging from queries about the taxonomic identity of specimens to reasoning about evolutionary properties across multiple lineages (see dahdul et al., 2010, for a rationale of this approach). in this context, the ontological framework developed by schulz et al. (2008) is remarkable and pioneering effort. the authors (pp. i314-i315) intend "to provide classes and classificatory criteria to categorize the foundational kinds of biology, without any restriction to granularity, species, developmental stages or states of structural wellor illformedness". to this end they utilize a set of complementary ontological relationships (cf. fig. 2 on p. i318); including is_a (example: a taxon quality is a quality), part_of (example: the mammalia class region is part of the animalia kingdom region), and instance_of (example: the individual "clyde" is an instance of the species elephas maximus ); as well as more taxonomicspecific relationships such as inheres_in (example: the quality of having an elephant heart inheres in the species elephas maximus ); derives_from (example: the hela cell line derives from human cells [but exists separately from the human population]), has_granular_part (example: the population of thai elephants has a granular part "clyde"), and located_in (example: a species biodiversity informatics, 7, 2010, pp. 45-66   47   quality is located in the elephas maximus species region). suitable combinations of these relationships allow schulz et al. (2008) to connect individual organisms and populations (particulars) to their respective species (universals), link their properties (qualities) to particular species, and integrate these qualities up to higher levels (regions) in the taxonomic hierarchy (see fig. 1). schulz et al.'s (2008) ontology is therefore a first approximation of the kind of full-blown ontological representation of biological taxonomy that is our focus hereinafter. static versus dynamic perspectives ontologies are necessarily biased to suit particular representation and reasoning demands (mccray, 2006). accordingly, ontologies for biological taxonomy may focus on truly ontological aspects of taxa in nature (example question: what are species?); or alternatively, on the epistemological challenges of representing taxa in systematic practice (example question: how did rivera delimit mus musculus in her 1980 publication?).2 on one side of the spectrum, ontology creators must address such contentious issues as whether taxa are real in a suitable sense (ereshefsky, 2002; lee, 2003; johansson, 2006; merrill, 2010), what natural boundaries exist between species and populations (hey, 2006; schulz and hahn, 2007), and whether taxa are historical individuals or natural kinds (ghiselin, 1997; boyd, 1999). different perspectives on these theory-laden issues will lead to different ontological implementations, as is reflected in schulz et al.'s (2008) elaborate attempts to represent taxa as meta-properties, class hierarchies, populations, qualities, or qualia (regions of qualities). at the same time the authors place little emphasis on how actual taxonomies come about, relying instead on use cases that consider taxonomic hierarchies as a relatively simple and stable (fig. 1). ontology development on the other side of the spectrum focuses on representing concepts and relationships inherent in taxonomic publications (e.g., yoon and rose, 2001; pullan et al., 2005; thau and ludäscher, 2007; franz et al., 2008;                                                              2  this distinction permits, but does not presuppose, that concepts used in systematic practice may correspond closely to 'true' and causally sustained entities and relationships in nature. see merrill (2010) for further discussion.  thau, 2008; thau et al., 2008, 2009; franz and peet, 2009). these efforts are motivated by the need to integrate alternative taxonomic classifications and biological linked to them. the goal is therefore to extract whatever ontologycompatible information is encrypted in a published taxonomy. unfortunately most taxonomic works display some form of logical inconsistency (thau and ludäscher, 2007; franz, 2009; vogt et al., 2009); including regional and taxonomic sampling biases, incomplete semantic treatments, implicit references to other works, and mosaic (as opposed to strictly hierarchical) distributions of phenotypic features (fig. 2). ontological representations must capture the underlying similarities and differences between such multiple perspectives. in short, they are more concerned with representing what taxonomists mean than with what taxa are. the two disparate views of the taxonomy/ontology interface have strong implications for ontology development and usage. specifically, the various solutions presented in schulz et al. (2008) all adhere to a static perspective of taxonomy: ontological entities and relationships are regarded as correct and stable enough to permit reasoning across a single taxonomic hierarchy. the static perspective assumes that at any given time there is a community-wide agreement as to how taxonomic entities are constituted and connected in nature. schulz et al. (2008) are not alone in employing this view, which also underlies proposals to establish a unitary taxonomy (cf. scoble, 2004) and is commonplace in on-line taxonomic information services including genbank and the catalogue of life (page, 2004; berendsohn and geoffroy, 2007; franz et al., 2008). schulz et al.'s (2008) model use case (fig. 1) shows a hierarchy passing from the kingdom animalia through six intermediate ranks to the species elephas maximus. yet in context of taxonomy's century-long history, this classification may only represent a momentary snapshot, ignoring previous disagreements or ongoing debates about the identity of taxonomic entities. when instead classifications are modeled so as to accurately reflect their nature as intermittent summaries of ongoing research, the result is a complex network of competing hypotheses whose relative support changes over time. for instance, vane-wright (2003) reviews a series of alternative   48   fig. 1. idealized taxonomic hierarchy as depicted in figure 1 of schulz et al. (2008). taxon qualities inhere in individual organisms (or parts thereof), and at the same time represent instances of speciesand higher-level taxon qualities. biodiversity informatics, 7, 2010, pp. 45-66   49   perspectives on species identities in the amauris butterfly complex.3 he concludes (p. 10): [i]f we are to move to any system approaching godfray's unitary ideal, then we will have to come to terms with either making arbitrary choices among competing classifications […], or accept the additional burden of arriving at some compromise or consensus at the outset. [ …] in my experience, the taxonomic community does not seem to be ready to accept the idea of compromise or consensus. the extent to which taxonomies evolve over time is easily underestimated. geoffroy and berendsohn (2003) reanalyzed 12 succeeding classifications of german mosses from 1927 to 2000 (koperski et al., 2000). of the 1548 taxonomic entities recognized in 2000, 44.5% were unstable, 20% were potentially unstable, and an additional 22.2% had undergone nomenclatural changes in comparison to previous treatments; leaving only 13.3% of the entities consistent in name and circumscription over a 73-year period. franz et al. (2008) compared eight succeeding classifications of north american vascular plants from 1933 to 2006. only 55% of the currently valid concepts remained stable over the examined time period. these numbers are expected to decrease further when analyzing more diverse and understudied lineages such as invertebrates (cf. iise, 2009). in other words, the static model used by schulz et al. (2008) assumes a level of stability that has so far been unattainable for many taxonomic groups. the dynamic perspective relaxes the frequently unjustified requirement of taxonomic stability. under this view, names, specimens, descriptions, illustrations, proposed homologies, molecular evidence, and other kinds of core taxonomic information are potentially varying elements of taxonomic concepts that evolve and replace each other over time (berendsohn, 1995). a particular perspective, published by an author or group of authors at a specific time and place, makes explicit or implicit reference to some or all of these elements. alternative perspectives are either congruent or (partially) incongruent with each other. the ontological challenge, then, is to                                                              3 clark et al.'s (2009) model of taxonomy as an escience similarly allows multiple perspectives to be represented and updated dynamically following peer review. accurately represent the elements of alternative taxonomies and connect them in ways that are logically consistent and permit reasoning across these taxonomies (thau and ludäscher, 2007; thau, 2008; thau et al., 2008, 2009). to achieve this goal, taxonomic experts must provide an initial, non-exhaustive set of concept relationship mappings or alignments (koperski et al., 2000; franz and peet, 2009). given two taxonomies and the initial expert alignments, ontological reasoning can perform a series of useful inferences; including the implementation of global taxonomic constraints (non-emptiness, etc.), checks for logical consistency, logical repairs, inferences of implied relations, removal of redundant relations, and the merging of aligned concepts (thau et al., 2008, 2009). however, as currently implemented the dynamic perspective focuses solely on the issue of concept congruence and otherwise treats the contents of taxonomic concepts as black boxes. this narrow approach is not sufficient to facilitate a full-blown ontological representation of taxonomy. ontological representation challenges in allowing taxonomies to evolve over time, the dynamic perspective is best suited to represent the products and interests of the taxonomic expert community. the static approach, on the other hand, is effective when taxonomy is virtually a constant, which is only a reasonable assumption in well circumscribed lineages such as model organisms. as illustrated above, each view has its inherent strengths and limitations. moreover, both dynamic and static ontological representations of taxonomy have so far focused on use cases that are strongly simplified in comparison to full-blown taxonomic publications. they have been developed by following the semantic restrictions of ontological languages used in computer science (e.g., owl web ontology language4). because of their origin in description logic (baader et al., 2004), these languages do not necessarily respond to special representation demands posed by biological taxonomies. they were certainly not conceived to address occurrences of varying species concepts (wheeler and meier, 2000) or evolving perspectives of higher-level taxa.                                                              4 http://www.w3.org/tr/owl-features/.  biodiversity informatics, 7, 2010, pp. 45-66   50   while it is necessary to explore the taxonomy/ontology interface with existing ontology languages like owl, we also need to understand which phenomena they cannot represent, even though these phenomena may play an important role in taxonomy. by reviewing the full range of representation demands from the viewpoint of taxonomic practice, we may thus identify features that prominent ontological languages can, or cannot, handle. clearly, both sorts of features must be understood to advance the development of ontologies for taxonomy. in the following sections we describe a noncomprehensive set of full-blown taxonomic phenomena that constitute severe representation challenges in the ontological realm. these phenomena are often interconnected. they are organized here starting with three kinds of epistemological constraints (sections i to iii), continuing with special linguistic conventions (sections iv and v), and ending with challenges related to the alignment of taxonomies (sections vi and vii). throughout, the example of a phylogeny and classification of the weevil genus apodrosus marshall (girón and franz, 2010) is used to illustrate the challenges being discussed (fig. 2; table 1). sections i to iii – epistemological constraints i. nature has the last long word biological taxonomies are proposed and used by humans and in that sense they are socially mediated constructs (bowker and star, 1999). however, taxonomists working under the widely accepted paradigm of evolution have no intention to construct wholly artificial systems. instead, they strive to discover natural phylogenetic relationships among the taxa under study (hennig, 1966). most will agree that evolution on earth has unfolded along a linear time scale, resulting in a unique sequence of events and causally sustained relationships – a single, true phylogeny.5 this phylogeny is what systematists aim to reconstruct and represent in their classifications. taxonomists often argue about the quantity and quality of character evidence and the validity                                                              5 this is an oversimplification, particularly when applied to microbial lineages where hereditary information is often passed on horizontally (cf. bapteste et al., 2009; boto, 2009; fraser et al., 2009). the main point nevertheless remains valid enough for the present context. of inference methods placing a certain group close to another one. for instance, the twisted-wing parasites (strepsiptera) are notoriously difficult to place among the insect orders; with beetles and flies being the two most favored possibilities (whiting, 1998; bonneton et al., 2006). despite such disagreements, it is understood that there are no multiple correct answers, and no answer can prevail mainly for conventional reasons. indeed, the naturalness of taxa – their unique origin and coherence through time – is thought to exist independently of human classifications. biological taxonomy differs in this sense from more conventional ontologies of (e.g.) pizzas or human diseases (horridge et al., 2004; du et al., 2009). in many research disciplines outside of taxonomy, it is epistemologically permissible to maintain multiple alternative hierarchies. as a result, the process of generating standard ontologies more readily leads to agreements about the context-specific utility of entities and their relationships. for instance, it is more straightforward to stipulate the truth content of the relationship "cheesetopping is_a pizzatopping" or "carcinoma of the large intestine is_a carcinoma", than it is to assert that "strepsiptera is_a halteria" (= a member of the haltere-bearing lineage that presumably includes flies; cf. whiting, 1998). unlike the former two assertions, which may be useful or not depending on the inferential context, the validity of the latter assertion depends on the final outcome of complex investigations of homology – a special and highly theory-laden kind of similarity resulting from common ancestry (e.g., wagner, 2001; franz, 2005b; assis and brigandt, 2009). while systematists are allowed at any time to posit such a relationship about the strepsiptera, its truth content is subject to continuous inductive testing of phylogeny and its inferential reliability is an a posteriori phenomenon. boyd (2000) uses the term bicamericalism to describe the process of delimiting biological taxa. this means that both our linguistic conventions and the causal structure of the world make up the content of taxa – therein understood as natural kinds – but the latter has the "final word". indeed, the impact of nature's causal relationships on taxonomy is so strong that classifications require continuous revision in light of new evidence. taxonomic research is an ongoing quest to approximate phylogeny where new insights may   51   fig. 2 . phylogeny and classification of the weevi l genus apodrosus marshall (girón and franz, 2010). species that are not part of apodrosus represent three tribes: anypotactini ( a. bi caudatus and p. scans orius); polydrusini (c. ni grocinctus, p. coni cus, p. mutabilis and p. peninsularis); and sitonini (s. californicus), the latter chosen as the "root". the full scientific names are provided and select phylogenetically informative characters are mapped onto the internal branches. grey rectangles represent synapomorphic transformations and white rectangles represent homoplasious transformations. the numbers above and below each rectangle correspond to character numbers and states, respectively. for explanations of letters a to n see table 1.     t ab le 1 . e xa m pl es a nd e xp la na ti on s of li ng ui st ic p he no m en a (l et te rs a to n in f ig . 2 ) w it h re le va nc e to th e ta xo no m y/ on to lo gy in te rf ac e. l et te r t ax on om ic la ng ua ge p he no m en a a pe ri ph er al ta xo no m ic in fo rm at io n (" ou tg ro up ") . t ri be s an d ge ne ra a re n ot f ul ly r ep re se nt ed w it h re ga rd s to th ei r pr op er ti es o r co ns ti tu en t m em be rs . t he n am es a nd cl as si fi ca ti on o f sp ec ie s ar e in a cc or da nc e w it h o 'b ri en a nd w ib m er 's ( 19 82 ) ch ec kl is t. t he s ou rc es o f th e sp ec ie s co nc ep ts , w he th er o ri gi na l o r re vi si on al , r em ai n un sp ec if ie d. b c or e ta xo no m ic in fo rm at io n (" in gr ou p" ). f ea tu re -b as ed d ia gn os es a nd li st s of a ll s pe ci es -l ev el m em be rs ( fo r th e ge nu s) a nd s pe ci m en s (f or th e sp ec ie s) a re p ro vi de d. c u ns pe ci fi ed r oo t, pr es um ab ly c on ne ct in g to th e su bf am il y e nt im in ae s ec . a lo ns oz ar az ag a an d l ya l ( 19 99 ) (t ho ug h se e a ). d t he n am e p. sc an so ri us h as tw o sy no ny m s: ( 1) p . l uc tu os us d ej ea n (1 83 7: 2 78 ), a n om en n ud um ( in va li d na m e no t b ea ri ng d es cr ip ti on ); a nd ( 2) p . m od es tu s g yl le nh al (1 83 4: 1 31 ), a h et er ot yp ic ju ni or s yn on ym . t he s yn on ym ie s ar e li st ed in o 'b ri en a nd w ib m er ( 19 82 ), h ow ev er th is in fo rm at io n is n ot r ef er en ce d in g ir ón a nd f ra nz ( 20 10 ). n or is it c le ar f ro m e it he r pu bl ic at io n w ho e st ab li sh ed th es e. p ol yd ac ry s s ca ns or iu s w as o ri gi na ll y de sc ri be d in th e ge nu s si to na g er m ar , t hu s (k lu g) is in p ar en th es es . t hi s is a ls o no t m en ti on ed in th e 20 10 p ub li ca ti on . e t he tw o po ly dr us us s pe ci es a re n on -m on op hy le ti c in th e tr ee . y et d ue to th e fr ag m en ta ry n at ur e of th e pe ri ph er al ta xo no m ic in fo rm at io n, n o cl as si fi ca to ry c ha ng es a re un de rt ak en . f g ir ón a nd f ra nz ' ( 20 10 ) co nc ep t o f ap od ro su s; in cl ud in g 13 m em be r sp ec ie s (o st en si ve ) an d th e fo ll ow in g di ag no si s (i nt en si on al ; i n pa rt ): " ap od ro su s i s a ge nu s of re la ti ve ly s m al l s iz ed ( 37 m m ), o ft en m et al li c co lo re d, e xc lu si ve ly ( w es te rn ) c ar ib be an e nt im in e w ee vi ls w it h ph an er og na th ou s m ou th pa rt s, w it ho ut a p os to cu la r lo be a nd vi br is sa e, a nd w it h th e hu m er i a nd w in gs b ei ng w el l d ev el op ed . a cc or di ng to m ar sh al l ( 19 22 : 5 9) , t he g en us s ha re s w it h th e st ri ct ly c on ti ne nt al p ol yd ru su s ' it s m or e sa li en t ch ar ac te ri st ic s' , i nc lu di ng a la te ra ll y si tu at ed a nt en na l s cr ob e an d co nn at e cl aw s. h ow ev er , a po dr os us c an b e di st in gu is he d fr om p ol yd ru su s a nd o th er p ol yd ru si ne g en er a by a p ar tic ul ar c om bi na ti on o f ch ar ac te rs in cl ud in g a m ed ia n fu rr ow o n th e he ad ; a la rg e, b ar e, a nd s m oo th tr ia ng ul ar a re a fo rm ed b y th e ep is to m e on th e ro st ru m ; t he pr es en ce o f pr em uc ro ; t he p re se nc e of a m ed ia n fo ve a on th e ve nt ra l s te rn um v ii ; a nd a n ei th er j or y -s ha pe d fe m al e sp er m at he ca [ … ]. " t he d ia gn os is in cl ud es a m ix o f fe at ur es w it h a st ro ng p hy lo ge ne ti c si gn al ( e. g. " sp er m at he ca j or y -s ha pe d" ) or ju st m ai nl y fo r ge ne ra l r ec og ni ti on ( e. g. " si ze s m al l" ). t he 2 01 0 de fi ni ti on d if fe rs dr am at ic al ly f ro m it s 19 22 p re de ce ss or : i t i s m or e re st ri ct iv e in te ns io na ll y ye t i nc lu de s m or e sp ec ie s os te ns iv el y. t he r es ul tin g al ig nm en t i s ap od ro su s s ec . g ir ón a nd f ra nz (2 01 0) i n t a n d o s t a po dr os us s ec . m ar sh al l ( 19 22 ). g fe at ur es th at d ef in e ap od ro su s a cc or di ng to g ir ón a nd f ra nz ( 20 10 ). o nl y on e ch ar ac te r (2 1) r ep re se nt s a un iq ue a nd u nr ev er se d sy na po m or ph y, a t l ea st w ith in th e na rr ow co nt ex t o f th is r ev is io n. h e xa m pl e of a r ev er se d sy na po m or ph y fo r th e ge nu s ap od ro su s; c ha ra ct er 1 3, s ta te 1 : p re se nc e of a m ed ia n po st er io r fo ve a on s te rn um v ii , s ec on da ri ly lo st ( "a bs en t" ) in a . ep ip ol ev at us . i e xa m pl e of a h om op la si ou s ch ar ac te r; c ha ra ct er 1 7, s ta te 1 : t eg m in al p la te o f m al e te rm in al ia c om pl et e, c on ve rg en tl y pr es en t i n th e ou tg ro up ta xa p . c on ic us a nd p . pe ni ns ul ar is , a s w el l a s in a po dr os us w he re it is u nr ev er se d. j e xa m pl e of r ep ea te dl y re ve rs ed c ha ra ct er ; c ha ra ct er 2 2, s ta te 1 : p re se nc e of a p ro je ct io n on th e co rn u of th e sp er m at he ca , w ith a n in fe rr ed p hy lo ge ne ti c se qu en ce ( ro ot ) 0  1  0  1 ( a. w ol co tti ). k e xa m pl e of a h om op la si ou s ch ar ac te r, c ha ra ct er 1 2: e ly tr al s tr ia l i nt er va l x s li gh tl y (1 ) or s tr on gl y (2 ) pr od uc ed , c on ve rg en tl y pr es en t i n th re e cl ad es w it hi n ap od ro su s, w it h a se co nd ar y tr an sf or m at io n in th e a. e pi po le va tu sa. w ol co tti c la de . l e xa m pl e of a m on op hy le ti c, u nr an ke d an d un na m ed s pe ci es g ro up w it hi n ap od ro su s ( "a . e xi m iu sa. e m ph er ef as ci at us c la de ") . m o nl y a. a rg en ta tu s a nd a . w ol co tti w er e kn ow n pr io r to th e 20 10 r ev is io n. t he n um be r of s pe ci es h as in cr ea se d m or e th an f iv efo ld . t he tw o pr ev io us ly k no w n sp ec ie s ar e no lo ng er c on si de re d si st er ta xa , b ut in st ea d ar e ne st ed w it hi n di ff er en t c la de s. t he se tw o sp ec ie s ar e m ai nt ai ne d as v al id ; t he ir d ia gn os es ( in te ns io na l) a nd s pe ci m en m at er ia l ( os te ns iv e; in cl ud in g di st ri bu ti on r an ge s) h av e be en r ef in ed a nd u pd at ed . t he ty pe s pe ci m en s w er e no t s ee n, n or a re th ey e xp li ci tl y ci te d in th e 20 10 p ub li ca ti on . n ap od ro su s q ui sq ue ya nu s i s on e of 1 1 ne w s pe ci es a dd ed to th e ge nu s. t hi s sp ec ie s is a n ew o st en si ve e le m en t o f th e up da te d ge nu s ci rc um sc ri pt io n, w it h pr op er ty -b as ed an d ty pe s pe ci m en -b as ed d ef in it io ns o f it s ow n. atp typewritten text atp typewritten text atp typewritten text 52 biodiversity informatics, 7, 2010, pp. 45-66   53   lead to radical realignments even of higher-level categories (e.g. cavalier-smith, 2004). this iterative inferential process differs dramatically from the logic-driven, stipulative ontology building that is prevalent in the information sciences (baader et al., 2004). the epistemological challenge of generating reliable statements of homology often compromises our ability to produce stable taxonomies and develop ontologies based on them. the sheer magnitude of the task of building a full-blown ontology differentiates taxonomy from other domains. large portions of the tree of life remain insufficiently explored, particularly in megadiverse lineages such as arthropods (5-10 million species estimated; ødegaard, 2000), fungi (up to 1.5 million species estimated; schmit and mueller, 2007), and microbial organisms (up to 10 million species estimated; sogin et al., 2006; though see fraser et al., 2009, for a discussion of problems related to species delimitation in bacteria). some 18,500 new species were added to the global count in the year 2007 alone (iise, 2009). with only 1.8 million species named – likely less than 10% of the total species richness on earth – the task of completing the global inventory will take several additional centuries (wheeler, 2004). each species contains vast amounts of phylogenetic information that will require expert analysis to achieve a reliable classification. discoveries of "missing link" fossils may lead to rearrangements of the sequence of deeper splits in the hierarchy (e.g. franzen et al., 2009). on the other hand, working solutions for smaller and well known groups such as mammals appear feasible (~ 5420 species and 37,400 synonyms; wilson and reeder, 2005). ii. entities with evolving properties because biological taxonomies organize the products of evolutionary history, they have to accommodate certain phenomena that characterize species and higher-level taxa, such as the evolvability of traits (boyd, 1999). two of the most critical concepts related to the generation of phylogenies are (1) homology, the similarity of traits resulting from common ancestry (see above); and (2) homoplasy, the similarity of traits resulting from convergence or reversal (schuh and brower, 2009). according to hennig (1966), natural taxa are characterized by synapomorphies, i.e. derived homologous features whose unique origin in the tree of life has been inductively corroborated. the silk-spinning organs ("spinnerets") that characterize all spiders are a textbook example of a synapomorphy. synapomorphies may be unreversed or reversed within the taxon they define. in the latter case, a feature that defines the taxon as a whole, and presumably evolved in its ancestor, has subsequently been modified at the genetic level and is phenotypically absent ("lost") in one or more of the taxon's younger subgroups. in other words, the defining feature is not obviously present in all descendants of the ancestor. instead, its presence – manifested in a modified state at the molecular level – must be inferred on the basis of the overall tree structure (see also fig. 2; table 1). this is how we can accurately classify the phenotypically limband digit-less snakes as members of the natural taxon tetrapoda. while this sort of classificatory practice poses no deep problems for human recognition and communication, a feature that is phenotypically absent (though inferred as present in an altered genotypic state) in a subordinate member of the superordinate class defined by it, may be difficult to represent in the language of description logics. at the very least, the condition of transitivity of properties from higher to lower level members in the hierarchy is violated.6 convergent properties present additional problems. a shared property that is phenotypically the "same", or even rooted in the "same" transformations at the molecular level, is not considered homologous if phylogeny indicates an independent origin in two taxa with no recent common ancestor. to provide an example, both flies and twisted-wing parasites have one pair of their thoracic wings modified into halteres, which are stalked and terminally knob-like structures that function as balancing organs during flight. in flies these structures are attached to the third thoracic segment, whereas in the twisted-wing parasites they occur on the second thoracic segment. whether or not these structures are considered homologous depends in part on the outcome of future phylogenetic studies. if it turns out that flies                                                              6 molecular phylogeneticists frequently represent homology statements in a probabilistic framework (nielsen, 2002; sober, 2002), which constitutes an additional representation challenge (j. felsenstein, pers. comm.). biodiversity informatics, 7, 2010, pp. 45-66   54   and twisted-wing parasites are sister taxa (whiting, 1998), then they may jointly be named "halteria" and defined by the homologous presence of halteres. in the opposite case (bonneton et al., 2006), "halteres" are no longer a synapomorphy of a natural taxon. instead we would have to take into account the phylogenetic contextuality of this feature – (1) "halteres of flies" and (2) "halteres of twisted-wing parasites" – and use each descriptor separately to characterize the respective groups. the degree of homoplasy of a property across the tree of life frequently depends on its descriptive precision (wenzel and carpenter, 1994; proctor, 1996; rieppel, 2007; franz and engel, 2010). in the aforementioned example, "halteres" is a rather specific term that may ultimately refer to a homologous structure with a single evolutionary origin. the term is at the lower end of the range of levels of homoplasy commonly used in taxonomy. at the other end, we observe terms referring to broadly delimited properties, for instance "petal color red" or "ventral side of prothoracic tibia with a row of triangular teeth". such widely circumscribed properties may have hundreds of independent origins in the evolution of angiosperms and insects, respectively. within a particular lower-level group, "petal color red" or "presence of a protibial row of teeth" might well map onto a homologous state. when used at more inclusive levels, however, these descriptors will become increasingly homoplasious and phylogenetically uninformative, picking out sets of taxa that have no recent common ancestor. thus the referential validity of broadly defined phenotypic terms fundamentally depends on a precise specification of the phylogenetic context. in spite of the above, it is common practice to reuse broadly defined phenotypic terms without making the phylogenetic context explicit enough ("eyes globular", "hind legs saltatorial", "wingless", etc.; see also fig. 3). the result is a phylogenetic underdetermined and multireferential terminology. this problem also affects the utility of ontologies being defined for morphological structures of plants (avraham et al., 2008), fish (dahdul et al., 2010), spiders (ramírez et al., 2007), wasps (mikó and deans, 2009), and other taxonomic groups (smith et al., 2007). even though each of these controlled vocabularies includes many taxon-specific terms, the terms are not consistently embedded in a phylogenetic context that recognizes convergent occurrences in the tree of life. they are perhaps best thought of as a powerful set of phenotype-categorizing metadata (vogt et al., 2009). as such, they stand to make great contributions to the standardization of taxonomic descriptions (vogt, 2009), but will not automatically facilitate ontological reasoning within a phylogenetic framework. as argued above, any morphological term that is not mapped to a unique and unreversed property in the tree of life must be further annotated with a taxonomic delimiter ("halteres of flies ").7 only then does the term pass on from the strictly diagnostic to the phylogenetic language realm. the need for additional, context-specifying qualifiers for properties is much less prevalent in other ontological networks. iii. individualand class-like components the philosophical literature is testimony to a longstanding debate as to whether species and higher-level taxa are individuals or classes (= natural kinds). this discussion permeates schulz et al.'s (2008) ontologies and motivates their introduction of inherent taxon qualities (see above). but beyond the development of ontologies, the individuals-versus-kinds debate has had little impact on the practice of classifying taxa (keller et al., 2003; though see de queiroz and gauthier, 1992). as summarized by brigandt (2009: 78, 95): […] a species or a higher taxon can be construed both as an individual and a natural kind, i.e. both views are metaphysically compatible. yet one conceptualization can be pragmatically preferable depending on the epistemic considerations that are in play in a certain scientific context. taxa are best construed as natural kinds when they are viewed as taxonomic units, while it is preferable to view taxa as individuals when they are conceived of as units of evolutionary change. […] the upshot of my discussion for t he individualism vs. kinds debate is that the relevant question is not so much into which metaphysical category species and higher taxa fall, but how biological accounts of taxa (such as species concepts) underwrite                                                              7 one reviewer stated that owl allows such context specification through addition of a sufficient condition; and furthermore, that morphological ontologies are singled out unfairly here for not performing a service they were never designed for. both objections are valid to a degree, though neither refutes the point that standard taxonomic practice and its ontological implementation are poorly matched up with the need to support phylogenetic inferences.  5 55   fig. 3. example of a full-blown taxonomic representation of a species of damselfish, chromis circumaurea sec. pyle et al . (2008: 15), including (a ) intensional components (diagnosis, description), ostensive components (type specimens), as well as (b) links to images, dna data, and globally unique identifiers for individual specimens (cf. hyam, 2009; page, 2009). reproduced with permission of the authors and journal. biodiversity informatics, 7, 2010, pp. 45-66   56   classifications and generalizations, shed light on the unity of taxa across time, and permit explaining their ability to undergo change a s a unit – all of which are epistemic issues. taxonomists have adopted a hybrid a pproach to circumscribe taxa, i.e. one that uses both intensional (property-based) and ostensive (member-based) components (fig. 2; table 1). this practice is ingrained in the linnaean tradition (farber, 1976; stevens, 1984), where species definitions are fixated by the combination of a verbal diagnosis (intensional, often accompanied by illustrations) and the designation of a type specimen (ostensive). the practice of pointing to a type is especially critical at the species level (fig. 3). taxonomic disagreements and nomenclatural synonymies are often resolved in direct reference to the identity of type specimens (e.g. gardner and hayssen, 2004). typification is also mandatory at the genus level, but becomes increasingly less prevalent as one climbs up the hierarchy to the level of family, order, class, etc. while it is common to list all examined specimens when defining a new species (e.g., franz, 2010), there is virtually no use in designating a specimen to typify megadiverse lineages such as the class insecta. in practice the latter is sufficiently well defined by listing a set of diagnostic properties or synapomorphies (grimaldi and engel, 2005). thus, we observe a gradual shift from ostensive to intensional components in accordance with the inclusiveness of the taxon being defined. this convention matches up well with human cognitive abilities and inference needs (brigandt, 2009; franz, 2009). for instance, it is not necessary for humans to examine specimens of every species of the scarab beetle superfamily scarabaeoidea in order to reliably recognize them as such. a generic illustration of the synapomorphic, asymmetrically lamellate antennal club is sufficient for this purpose. similarly, two experts talking about the weevil genus perelleschus wibmer & o'brien may understand each other even though each has only seen specimens from central and south america, respectively, which share no common species (franz and o'brien, 2001). our cognitive tendency to shift towards "loosely typified" intensional definitions explains why regional and taxonomic sampling biases are acceptable in higher-level phylogenetic analyses. these definitions have greater predictive value and are especially useful for making wide-reaching inferences about the identity of a taxonomic group – past, present, and future, yet their referential precision is compromised by increasing levels of homoplasy and evolutionary transformation (section ii). ostensive definitions are more accurate but offer few inferential benefits beyond specific identification of a set of specimens or taxa.8 schulz et al. (2008) are highly responsive to the individuals-versus-kinds problem (see also gangemi et al., 2001). while recognizing the complex interaction of ostension and intension in taxonomy, the authors discard several of their initial proposals; viz. taxa as meta-properties, hierarchies, and populations. instead, they propose to represent taxa as qualities; in the sense that an individual organism or part thereof has an inherent quality of pertaining to a taxon (individual  species). that taxon, in turn, has the quality of pertaining to a higher-level taxon (species  genus, etc.). the addition of quality regions seemingly offers a workable transition from the realm of particulars to universals. however, this solution leads to a sort of "molecular essentialism" where taxon-level qualities can inhere in an isolated sequence of nucleic acids (p. e. midford, pers. comm.). moreover, the approach ignores why taxonomists single out specific properties to characterize taxa within a larger lineage (hennig, 1966; wheeler and meier, 2000), and fails to recognize that taxonomists use utilize ostensive and intensional elements flexibly and inconsistently at varying levels (e.g., sereno, 2005; franz and peet, 2009; schuh and brower, 2009). therefore the taxa-as-qualities proposal remains too simplistic in comparison to actual practice. in the end, any approach that decides one way or the other with regards to the purported individual/class dichotomy falls short of the reasoning powers that humans derive from hybrid definitions of taxa. sections iv and v – special linguistic conventions iv. linnaean ranks linnaeus (1758) advocated the strict use of ranks for taxonomic names – a convention that has                                                              8 from an ontological viewpoint it is even arguable whether preserved specimens count as members of their respective taxa, given that they have lost key molecular or behavioral (p. e. midford, pers. comm.). biodiversity informatics, 7, 2010, pp. 45-66   57   now persevered for more than 250 years (schuh, 2003). particularly names for midto higher-level taxa incorporate standardized terminations indicating their rank. in zoology, for instance, the names of tribes end with -ini, subfamilies end with -inae, families end with -idae, and so on. the botanical nomenclature has at least 13 fixed terminations for ranks up to the level of division (phyta, -mycota ) (icbn, 2006). although the number of ranks is in principle indeterminate, the nomenclatural codes only regulate the application of endings for a relatively small number of ranks, not reaching beyond the level of family in zoology (iczn, 1999). but the use of more finely tiered ranks is pervasive in taxonomy, where prefixes such as super-, sub-, and infraprovide additional levels of resolution to accommodate increasingly more bifurcated trees and refined phylogenetic classifications (e.g. mckenna and bell, 1997). a single classification may include (1) ranked names (i) with or (ii) without standardized endings; (2) informal names that map onto unranked sections of the hierarchy (e.g. "paleoherbs"; cf. nixon and carpenter, 2000); and (3) sections that are not named at all (unnamed clades, taxa of uncertain position, etc.). adding a rank ending to a taxonomic name might seem nothing more than an arbitrary convention with limited significance for ontological reasoning. yet it is precisely this convention which allows humans to make countless implicit and inter-subjectively reliable inferences about taxonomic relationships without necessarily having to visualize a reference tree. many midto higher-level taxa have conspicuous and well conserved synapomorphies. the coupling of the features with a ranked name and standardized ending reinforces a mental association between them. the cognitive pay-offs are aptly characterized in this example by platnick (2001: 8-9; see also platnick, 2009): i was wandering around john murphy's garden out in hampton, and came across a nice jumping spider. now, jumping spiders, the family salticidae, are probably the easiest of all spider families to recognize. with their large anterior median eyes, their excellent vision, the often highly exuberant and ornamented morphology that males use in their elaborate courtship displays, and their pr owess at jumping on prey several body-lengths away, salticids are quite distinctive. [… ] if you visit the [world spider catalog] site, you'll find a summary table that shows, for each of the 109 currently recognized spider families, the numbers of currently valid genera and species, including, at the very end of the list, the salticids, with 4,834 species. using the linnaean hierarchy, when i identified the spider in john's garden as a salticid, i was asserting that john's spider is more closely related to any single species currently included within the salticidae than it is to any single species that is currently excluded from that family. in other words, if my identification, and the current classification, are both correct, then john's spider is more closely related to salticid species #1 than it is to any of the 32,752 spider species currently excluded from the salticidae. it is also more closely related to salticid species #2 than it is to any non-salticid spider. so, assuming that the spider from john's garden belongs to one of the currently known 4,834 salticid species (and this being england, that's certainly a f air assumption), then my identification enables 4,833 (other salticids) times 32,752 (non-salticids) three-taxon statements. so by placing the animal as a salticid, the current linnaean hierarchy allows me to make 158,290,416 three-taxon statements about it, within spiders alone. if i were to expand the arena to include all arthropods, or all life, the number of implied three-taxon statements would, for all practical purposes, approach one third of infinity – the other two-thirds would be prohibited. that's none too shabby, for a single word – sa lticidae (admittedly, in a context provided – solely – by the linnaean hierarchy, and the mutual exclusivity of equally ranked names it requires). even if we concede that linnaean ranks alone are not sufficient to convey the level of speaker expertise portrayed in this example, it is clear that using ranks has immense inferential advantages. put simply, two taxa that have the same rank ending within a single coherent classification cannot share any subordinate members (see also thau et al., 2009). this perfect nestedness allows humans to make countless accurate inferences about the placement of subordinate members into these non-overlapping taxa – platnick's (2001) three-taxon statements – without having to memorize their diagnostic feature or exact position in the overall hierarchy. in addition, and despite the fact that rank assignments are not evolutionarily comparable across the tree of life (cf. avise and johns, 1999), many taxon/rank biodiversity informatics, 7, 2010, pp. 45-66   58   couplings do remain stable enough so that relevant biological phenomena become tied them over time (e.g., salticidae ↔ large anterior median eyes, prowess at jumping, elaborate courtship). students of biology often begin by recognizing orders, then families, and later on lower ranked entities such as genera and species. at each level they establish meaningful cognitive links between ranked names and biological traits; e.g. members of the formicidae (ants) are highly eusocial, the workers are female and wingless, etc. whenever (minor) rank changes occur, it is relatively easy for humans to adjust by making a limited number of rank reconnections (e.g. subfamily [-inae]  tribe [ini]) while keeping much of the learned name/information association intact. linnaean ranks and standardized endings have no obvious match in conventional ontologies. to address this issue, schulz et al. (2008) created a secondary hierarchy of rank classes that interfaces with the primary taxon quality hierarchy. however, as the authors concede themselves (p. i320): "the meaning of the taxonomic rank classes […] is somewhat counterintuitive, since every instance of speciesquality is also an instance of genusquality and so on. they are, therefore, not suited to comprehensively represent the meaning of species as disjoint from genus, kingdom, etc." other attempts to incorporate ranks into ontologies have opted to render them "ontologically weak"; treating a taxon as an instance_of its rank as opposed to the stronger is_a relationship which would give ranks the status of meta-classes and associated reasoning powers (dahdul et al., 2010; p. e. midford, pers. comm.). in short, translating the inferential benefits of ranks into the ontological realm has so far remained elusive. v. nomenclatural and taxonomic legacy taxonomy is bound in its use of names by a legacy that reaches back some 250 years to linnaeus' (1758) systema naturae. many taxon names and definitions were first published in that work. linnaeus (1758) also laid out a set of rules for naming taxa, including his binomial system and advocacy of ranks. in zoological taxonomy in particular, these rules were expanded by strickland et al. (1843) who introduced the law of priority and other requirements regarding the typification of names (farber, 1976). these efforts have gradually evolved into the current nomenclatural codes (e.g., iczn, 1999). the process of amending the codes continues in response to new threats to the stability of names and associated information (e.g., iczn, 2008). application of the rules of nomenclature for nearly 250 years has transformed the taxonomic literature into a continuous chain of publications with quasi-legal status (minelli, 2003). any new publication must recognize relevant nomenclatural and taxonomic precedents. indeed, new taxonomies are mostly communicated through existing names whose meanings were established in earlier publications. often these meanings are revised, expanded, or contracted in complex ways. as taxonomic revisions accumulate and supersede each other over time, they tend to create a network of many-to-many relationships among valid names, invalid synonyms, and past and present meanings (koperski et al., 2000; geoffroy and berendsohn, 2003; franz, 2005a; kennedy et al., 2005; franz et al., 2008). the situation is compounded by the fact that nomenclatural and full-blown taxonomic relationships are established in different and frequently non-congruent ways; the former being determined strictly on the basis of the identity of type specim ens, whereas the latter involve comparison of diagnostic features and other kinds of phylogenetic information. the trajectories of nomenclatural and taxonomic relationships among names are therefore semi-independent and must be modeled separately to record partial name/meaning disjunctions over time (koperski et al., 2000; kennedy et al., 2005; franz and peet, 2009). the history of taxonomy is littered not only with an indelible record of nomenclatural relationships but with thousands of published taxonomies that are partially incomplete, outdated, or of questionable quality and validity. nevertheless, some cross-section of this body of work represents our best present-day knowledge of nature's hierarchy (cf. maddison et al., 2007). many old classifications have been fully replaced by more recent revisions. yet even the least regarded publications (cf. jäch, 2006) are part of taxonomy's enduring ledger and must be linked in some ways to a more widely accepted view. it is generally not permissible in taxonomy to purge low quality work from the record or to ignore names coined in obscure treatments (cf. godfray, 2002). biodiversity informatics, 7, 2010, pp. 45-66   59   in this sense, taxonomy presents legacy integration difficulties that go beyond those of more conventional ontologies. while it is considered best practice in any domain to link a new ontology to a relevant predecessor (euzenat and shvaiko, 2007; smith et al., 2007; sahoo et al., 2008), there are no quasi-legal requirements to do so. no laws of priority or rules for typifying classes exist that, if violated, would render the new ontology invalid. there are also no clear analogies to the consistent separation of type-based nomenclatural versus full-blown taxonomic relationships among the elements of succeeding ontologies. lastly, it is permissible in most domains to ignore a previous ontology if it no longer has any utility. the result is a more uniform framework for reasoning that is not overly compromised by idiosyncratic entities and relationships of the past. sections vi and vii – alignment challenges vi. static alignments – core versus peripheral information for reasons given in section i, no team of authors is capable of publishing a full-blown taxonomy that spans across the entire tree of life. to the extent that such "complete" taxonomies are available, they represent compilations of multiple treatments published on subsections of the tree (scoble, 2004). they tend to have a limited amount of information associated with each taxon name (e.g.,9), and typically provide no information on types and synapomorphies, or even links to sources in the primary taxonomic literature. this means that all published classifications are somehow are incomplete, due either to insufficient breadth (not all taxa covered) or depth (not all information provided of the covered taxa), or both. new taxonomic contributions must adopt a piece-meal approach, focusing on select taxa and sources of data; e.g., morphological traits and/or molecular sequence information. beyond this core taxonomic focus, each new treatment usually connects to peripheral taxonomic information stemming from relevant predecessors. to provide an example, the most recent revision of the weevil genus cotithene voss (franz, 2008) contains a near-complete account of core taxonomic information on the genus itself and all eight                                                              9 http://www.catalogueoflife.org/.  constituent species; including differential diagnoses and detailed lists of specimens designated to represent each taxon. on the other hand, this work makes only peripheral and implicit statements about the taxonomy of related genera, without listing specimens or even naming all species per genus. mostly there are pointers to other treatments which contain this information. in contrast, a separately published genus-level reclassification of the tribe which contains cotithene offers much less taxonomic information on this genus in particular (franz, 2006). instead the focus is on inferring synapomorphies that define monophyletic groups within the tribe. beyond this narrow focus there is only minimal information on other tribes from which or into which certain genera are transferred (franz, 2009). thus the "semantic joints" to non-focal data remain vague (see also fig. 2; table 1). the piece-meal nature of taxonomic publications and incurred differential focus on core versus peripheral information result in severe challenges for ontology building. specifically, if the goal is to represent the entire tree of life based on both intensional and ostensive data, then information from many independently published sources must be linked together to obtain just a single static perspective. primarily ostensive classifications of catalogues (e.g. alonso-zarazaga and lyal, 1999) would have to be integrated with exemplary phylogenetic studies (cf. prendini, 2001) so that names and concepts listed in the former may be further defined through synapomorphies. phylogenetic mid-level analyses must be linked to species descriptions and specimen data provided in taxonomic revisions, and so on. considering the semantic complexity of the information to be integrated, the goal of a single full-blown ontology of the tree of life becomes almost as difficult as the underlying research itself. most critically, the integration process must involve third-party expert asses sments in order to connect core results from multiple publications at their peripheral "edges" and thereby construct a contiguous information-rich network. in light of the inherent vagueness of these "edges", the assessments require expertise of the historical and regional context in which each publication was produced. different expert teams may propose different ways in which to integrate such works. biodiversity informatics, 7, 2010, pp. 45-66   60   therefore the challenge of developing a single tree of life ontology rapidly devolves into making intersubjective assessments of concept equivalence among partial ontologies. in practice this translates into matching taxonomies at their edges using taxonomic concept relationships. this challenge is further discussed in the next section. vii. dynamic alignments as seen above, the assembly of a full-blown static ontology for taxonomy and dynamic integration of alternative taxonomies both require some measure of ontology alignment. thau et al. (2009) provide an overview of the logical and computational challenges involved in optimizing this process. here we will concentrate on the semantic challenges posed by the initial expert alignments. koperski et al. (2000) were the first to use a vocabulary of five terms derived from set theory in order to align two separately published taxonomic concepts c1 and c2 with each other; as follows (symbols according to thau et al., 2009): (1) congruence (c1  c2), (2) proper inclusion (c1 c2), (3) proper inverse inclusion (c1 c2), (4) partial overlap (c1  c2), and (5) exclusion (c1 ! c2). the terms have since been employed sporadically by other authors (e.g., gradstein et al., 2001; güntsch et al., 2003; kennedy et al., 2006; weakley, 2006; graham and kennedy, 2007; craig and kennedy, 2008; franz et al., 2008; krings, 2008). franz and peet (2009) subsequently showed that this vocabulary must be expanded to account for different relationships based on whether the intensional or ostensive subcomponents of two concepts are compared. an example for such a scenario is given in section iii: two experts working in central america (c1) and south america (c2) agree on the diagnosis of perelleschus based on jointly recognized synapomorphic features, but each lists a mutually exclusive set of species in the corresponding regional treatment of the genus. accordingly, the intensional alignment of the two concepts is c1  c2 whereas the ostensive alignment is c1 ! c2. the intensional congruence indicates that the concepts are in agreement as to what species generally belong to perelleschus – past, present, and future. the ostensive exclusion, in turn, reflects the fact that the two experts happen to work with a regionally biased and non-overlapping set of species. only the combined alignments capture how the two concepts relate to each other with in terms of their predictive content and explicitly included members. franz and peet (2009) showed how the expanded vocabulary can be applied to express (1) how taxonomic (c) and nomenclatural (n) relationships are interconnected (e.g. c1  c2 and n1 is a heterotypic synonym for n2); (2) whether there is uncertainty in an alignment (e.g. c1  c2 or c1 c2); (3) how to negate a relationship (e.g. c1 not  c2); (4) if there is a way to reconcile two classifications through addition or subtraction of concepts on one side (e.g. c1  c2 + c3); or (5) whether two authors stipulate different featurebased definitions (int) of a taxon even though they examined the same set (ost) of subordinate members (e.g. c1 int  c2 and c1 ost  c2). depending on the richness of the source data, such combined alignments can represent partial concept matches and implicit judgments of errors in the aligned taxonomies. these capabilities are not yet available in an ontological environment. even though alignments are essential for connecting concepts occurring in alternative taxonomies, they cannot replace expert judgment as to whether certain instances of congruence or non-congruence are significant with respect to a particular integration task. one might ask, for instance, whether two feature-based definitions of a taxon pick out a sufficiently similar set of subordinate members when each is applied outside its geographic context. the answer will vary according to case-specific standards of "sufficiently similar": what is similar enough for an ecological study may not suffice for an analysis of adaptive radiation, and so on (cf. peterson and navarro-sigüenza, 1999). in this regard it is difficult to image that taxonomic concept alignments will ever become fully automated (see also geoffroy and güntsch, 2003; thau and ludäscher, 2007; thau et al., 2009). summary and outlook as reviewed above, taxonomic practice is bound by a range of epistemological constraints and linguistic conventions that run orthogonal to the logical background from which ontological entities and relationships originate (baader et al., 2004). the enormous challenge of reconstructing biodiversity informatics, 7, 2010, pp. 45-66   61   the tree of life compromises the goal of creating an ontology that is comprehensive and reliable enough to permit reasoning about taxa and their properties. the inherent evolvability of taxa poses genuine ontological challenges in the philosophical sense of the term, because it undermines the individual/class dichotomy and associated reasoning capabilities that sustain conventional ontologies (cf. schulz et al. 2008). in addition, the idiosyncratic yet indelible 250-year legacy of nomenclatural and taxonomic changes has resulted in an immense network of names and concepts that can only be aligned with an expanded vocabulary whose application requires expert input. for better or worse, we should recognize that much of taxonomy's cumulative body of work is not well aligned with the requirements for ontological representation and reasoning. our analysis could explain why the development of ontologies for taxonomy has lagged behind in comparison to other disciplines. to the extent that taxonomic research is still focused on acquiring and interpreting primary data to generate a natural classification, the prospects of reliably applying these data in a metadata-driven ontological framework will remain limited. many ongoing projects in taxonomy are motivated precisely by the insight that the existing "ontology" for a particular lineage – i.e., the previously established classification – is incomplete or at least partly wrong (vane-wright, 2003). this is not an optimal foundation for positing stable, logic-based definitions and relationships among taxa. in light of these limitations, we suggest that a full-blown and static representation of taxonomic information for large portions of the tree of life is not the most fruitful path to advance research along the taxonomy/ontology interface. instead, our efforts should concentrate (1) on representing strictly nomenclatural relationships (huber and klump, 2009), and (2) on improving ontologydriven vocabularies and algorithms for producing alignments between multiple taxonomies (franz et al., 2009; thau et al., 2009). taxonomic experts stand to benefit from each of these developments because they will facilitate the identification and integration of taxonomic legacy information. both types of services may become stepping stones towards a dynamic ontological network representing the products of taxonomic research. conversely, full-blown and static ontologies for major organismal lineages will neither serve taxonomy nor its users in the long term. the utility of such ontologies is limited to smaller groups where the assumption of taxonomic stability is reasonable (cf. smith et al., 2007; dahdul et al., 2010; mungall et al., 2010). in either case, researchers should for the most part refrain from resolving "deep questions" about the ontological nature of taxa because the answers will vary according to the preferred inferential context (brigandt, 2009). we furthermore suggest that the prospects of utilizing ontological reasoning in taxonomy will largely depend on the ability of the expert community to present phylogenies and classifications in ways that are more compatible with ontological principles than concurrent practice. minimally, this means: (1) adopting strict conventions for linking new core taxonomic information to (provisionally accepted) peripheral information so that the relevant context of the new contribution is fully defined; (2) using lineagespecific phenotype ontologies for taxonomic descriptions while specifying the phylogenetic context of the descriptive terms in use (cf. ramírez et al., 2007; mikó and deans, 2009; dahdul et al., 2010); (3) presenting all nomenclatural and taxonomic novelties in an ontology-compatible format, including intensional and ostensive definitions (see also sereno, 2009); and (4) providing intensional and ostensive alignments to entities in relevant preceding taxonomies (franz and peet, 2009; thau et al., 2009). the implementation of these practices will require a wider acceptance of the taxonomic concept approach (berendsohn, 1995; koperski et al., 2000; franz et al., 2008). in particular, taxonomists will have to become more disciplined in recognizing acts of authoring or citing concepts as well as identifying specimens or subordinate concepts to them (franz and peet, 2009). adopting this approach may also force taxonomists to have more control over the economics of maintaining a cyberinfrastructure for the publication, continuous versioning, and cross-linking of such concepts. while these goals are worthy of pursuit, they will likely remain elusive in the short term. neither taxonomists nor developers of ontologies should be under the illusion that fullblown ontologies for taxonomy will soon be biodiversity informatics, 7, 2010, pp. 45-66   62   dynamically assembled and updated without also putting in place robust mechanisms for recognizing individual expert contributions (cf. clark et al., 2009). although phylogenies and classifications may represent no more than a means to an end for user communities (dahdul et al., 2010), they constitute the primary intellectual products of taxonomists. the latter rely on signaling their authorship of these products in order to advance their academic careers. schulz et al.'s (2008) ideal of an overarching ontology-based framework for organizing all organismal information implicitly requires that the field of taxonomy regains a more powerful role among the biological disciplines. acknowledgments the authors wish to thank members of the seek project (http://seek.ecoinformatics.org/) for critical input shaping the content of this paper; in particular james beach, shawn bowers, jessie kennedy, sergeui krivov, xianhua liu, bertram ludäscher, robert peet, and aimee stewart. peter midford and two anonymous reviewers provided insightful comments on an earlier version of this manuscript. the senior author's research on biodiversity informatics and entimine weevils was supported through a postdoctoral fellowship by the andrew w. mellon foundation and by the national science foundation, award deb-641231, respectively. the junior author's research on reasoning about taxonomies was supported by the national science foundation, awards iis-0630033, dbi-0743429, and dbi-0753144. references alonso-zarazaga, m. a., lyal, c. h. c. 1999. a world catalogue of families and genera of curculionoidea (insecta: coleoptera) excluding scolytidae and platypodidae. entomopraxis, barcelona. assis, l. c. s., brigandt, i. 2009. homology: homeostatic property cluster kinds in systematics and evolution. evol. biol. 36: 248-255. avise, j. c., johns, g. c. 1999. proposal for a standardized temporal scheme of biological classification for extant species. proc. natl. acad. sci. u.s.a. 96, pp. 7358-7363. avraham, s., tung, c. w., ilic, k., jaiswal, p., kellogg, e. a., mccouch, s., pujar, a., reiser, l., rhee, s. y., sachs, m. m., schaeffer, m., stein, l., stevens, p., vincent, l., zapata, f., ware, d. 2008. the plant ontology database: a community resource for plant structure and developmental stages controlled vocabulary and annotations. nucleic acids res. 36, d449-d454. baader, f., horrocks, i., sattler, u. 2004. description logics. in: staab, s., studer, r. (eds.), handbook on ontologies. springer, berlin, pp. 3-28. balhoff, j. p., dahdul, w. m., kothari, c. r., lapp, h., lundberg, j. g., mabee, p. m., midford, p. e., westerfield, m., vision, t. j. 2010. phenex: ontological annotation of phenotypic diversity. plos one. (in press). bapteste, e., o'malley, m. a., beiko, r. g., ereshefsky, m., gogarten, j. p., franklin-hall, l., lapointe, f. j., dupré, j., dagan, t., boucher, y., martin, w. 2009. prokaryotic evolution and the tree of life are two different things. biol. direct. 2009, 4:34, 1-20. berendsohn, w. g. 1995. the concept of “potential taxa” in databases. taxon 44, 207-212. berendsohn, w. g., döring, m., geoffroy, m., glück, k., güntsch, a., hahn, a., kusber, w.-h., li, j.-j., röpert, d., specht, f. 2003. the berlin taxonomic information model. schrift. veget. 39, 15-42. berendsohn, w. g., geoffroy, m. 2007. networking taxonomic concepts – uniting without "unitary-ism". in: curry, g., humphries, c. (eds.), biodiversity databases – techniques, politics, and applications. crc taylor & francis, baton rouge, pp. 13-22. bonneton, f., brunet, f. g., kathirithamby, j., laudet, v. 2006. the rapid divergence of the ecdysone receptor is a synapomorphy for mecopterida that clarifies the strepsiptera problem. insect mol. biol. 15, 351-362. boto, l. 2009. horizontal gene transfer in evolution: facts and challenges. proc. royal soc. b, 1-9. bowker, g. c., star, s. l. 1999. sorting things out: classification and its consequences. mit press, cambridge, ma. boyd, r. 1999. homeostasis, species, and higher taxa. in: wilson, r. a. (ed.), species: new interdisciplinary essays. bradford books, mit press, cambridge, pp. 141-185. boyd, r. 2000. kinds as the "workmanship of men": realism, constructivism, and natural kinds. in: nidarümelin, j. (ed.), rationality, realism, revision: proceedings of the third international congress of the society for analytical philosophy, september 15-18, 1997 in munich. walter de gruyter, berlin, pp. 52-89. brigandt, i. 2009. natural kinds in evolution and systematics: metaphysical and epistemological considerations. acta biotheor. 57, 77-97. cavalier-smith, t. 2004. only six kingdoms of life. proc. royal soc. lond. b 271, 1251-1262.clark, b. r., godfray, h. c. j., kitching, i. j., mayo, s. j., scoble, m. j. 2009. taxonomy as an escience. phil. trans. r. soc. a 367, 953-966. biodiversity informatics, 7, 2010, pp. 45-66   63   craig, p., kennedy, j. 2008. concept relationship editor: a visual interface to support the assertion of synonymy relationships between taxonomic classifications. in: börner, k., gröhn, m., park, j., roberts, j. (eds.), visualization and data analysis 2008, proceedings of the society of photo-optical instrumentation engineers (spie), vol. 6809, january 28-29, 2008, san jose, ca; pp. 6809.066809.12. dahdul, w. m., lundberg, j. g., midford, p. e., balhoff, j. p., lapp, h., vision, t. j., haendel, m. a., westerfield, m., mabee, p. m. 2010. the teleost anatomy ontology: anatomical representation for the genomic age. syst. biol. 59. (in press). du, p., feng, g., flatow, j., song, j., holko, m., kibbe, w. a., lin, s. m. 2009. from disease ontology to disease-ontology lite: statistical methods to adapt a general-purpose ontology for the test of geneontology associations. bioinformatics 25, i63-i68. ereshefsky, m. 2002. linnaean hierarchy: vestiges of a bygone era. phil. sci. 69, s305-s315. euzenat, j., shvaiko, p. 2007. ontology matching. springer, berlin. farber, p. l. 1976. the type concept in zoology in the first half of the nineteenth century. j. hist. biol. 9, 93-119. franz, n. m. 2005a. on the lack of good scientific reasons for the growing phylogeny/classification gap. cladistics 21, 495-500. franz, n. m. 2005b. outline of an explanatory account of cladistic practice. biol. phil. 20: 489-515. franz, n. m. 2006. towards a phylogenetic system of derelomine flower weevils (coleoptera: curculionidae). syst. entomol. 31, 220-287. franz, n. m. 2008. revision, phylogeny, and natural history of cotithene voss (coleoptera: curculionidae). zootaxa 1782 1-33. franz, n. m. 2009. letter to linnaeus. in: knapp, s., wheeler, q. d. (eds.), letters to linnaeus. linnean society of london, london, pp. 63-74. franz, n. m. 2010. revision and phylogeny of the genus apotomoderes dejean (coleoptera: curculionidae: entiminae). zookeys 49: 33-75. franz, n. m., engel, m. s. 2010. can higher-level phylogenies of weevils explain their evolutionary success? a critical reassessment. syst. entomol. 35. (in press). franz, n. m., o'brien, c. w. 2001. revision and phylogeny of perelleschus (coleoptera: curculionidae), with notes on its association with carludovica (cyclanthaceae). trans. am. entom. soc. 127, 255-287. franz, n. m., peet, r. k. 2009. towards a language for mapping relationships among taxonomic concepts. syst. biodiv. 7, 5-20. franz, n. m., peet, r. k., weakley, a. s. 2008. on the use of taxonomic concepts in support of biodiversity research and taxonomy. in: wheeler, q. d. (ed.), the new taxonomy. systematics association special volume series 74. taylor & francis, boca raton, pp. 63-86. franzen, j. l., gingerich, p. d., habersetzer, j., hurum, j. h., von koenigswald, w., smith, b. h. 2009. complete primate skeleton from the middle eocene of messel in germany: morphology and paleobiology. plos one 4(5), e5723, 1-27. fraser, c., alm, e. j., polz, m. f., spratt, b. g., hanage, w. p. 2009. the bacterial species challenge: making sense of genetic and ecological diversity. science 323, 741-746. gangemi, a., guarino, n., masolo, c., oltramari, a. 2001. understanding top-level ontological distinctions. in: gómez pérez, a., gruninger, m., stuckenschmidt, h., uschold, m. (eds.), ontologies and information sharing. proceedings of the ijcai01 workshop on ontologies and information sharing. morgan kaufmann, san francisco, pp. 2633. gardner, a. l., hayssen, v. 2004. a guide to constructing and understanding synonymies for mammalian species. mamm. spec. 739, 1-17. geoffroy, m., berendsohn, w.g. 2003. the concept problem in taxonomy: importance, components, approaches. schrift. veget. 39, 5-14. geoffroy, m., güntsch, a. 2003. assembling and navigating the potential taxon graph. schrift. veget. 39, 71-82. ghiselin, m. t. 1997. metaphysics and the origin of species. state university of new york press, albany. girón, j. c., n. m. franz, n. m. 2010. revision and phylogeny of the genus apodrosus marshall 1922 (coleoptera: curculionidae: entiminae). insect syst. evol. 41. (in press). goble, c., stevens, r. 2008. state of the nation in data integration for bioinformatics. j. biomed. inf. 41: 687-693. godfray, h. c. j. 2002. challenges for taxonomy. nature 417, 17-19. gradstein, s. r., sauer, m., koperski, m., braun, w., ludwig, g. 2001. taxlink, a program for computerassisted documentation of different circumscriptions of biological taxa. taxon 50, 1075-1084. graham, m., kennedy, j. 2007. visual exploration of alternative taxonomies through concepts. ecol. inform. 2, 248-261. grimaldi, d., engel, m. s. 2005. evolution of the insects. cambridge university press, cambridge. güntsch, a., döring, m., geoffroy, m., glück, k., li, j., röpert, d., specht, f., berendsohn w. 2003. the taxonomic editor. schrift. veget. 39, 43-56. biodiversity informatics, 7, 2010, pp. 45-66   64   hennig, w. 1966. phylogenetic systematics. university of illinois press, urbana, il. hey, j. 2006. on the failure of modern species concepts. trends ecol. evol. 21, 447-450. horridge, m., knublauch, h., rector, a., stevens, r., wroe, c. 2004. a practical guide to building owl ontologies using the protégé-owl plugin and coode tools edition 1.0. university of manchester10 [accessed xii-20-2009]. huber, r., klump, j. 2009. charting taxonomic knowledge through ontologies and ranking algorithms. comp. geosci. 35, 862-868. hyam, r. 2009. taxa, taxon names and globally unique identifiers in perspective. to appear in: descriptive taxonomy: the foundation of biodiversity research. systematics association, cambridge university press, 11 pp.11 [accessed xii-20-2009]. icbn – international code of botanical nomenclature (vienna code). regnum vegetabile 146, 1-568. iczn – international commission on zoological nomenclature. 1999. international code of zoological nomenclature. international trust for zoological nomenclature, london. iczn – international commission on zoological nomenclature. 2008. proposed amendment of the international code of zoological nomenclature to expand and refine methods of publication. zootaxa 1908, 57-67. iise – international institute for species exploration. 2009. sos – state of observed species12 [accessed xii-20-2009]. jäch, m. a. 2006. taxonomy and nomenclature threatened by d. makhan. koleop. rund. 76, 360. johansson, i. 2006. bioinformatics and biological reality. j. biomed. inf. 39, 274-287. jones, m. b., schildhauer, m. p., reichman, j., bowers, s. 2006. the new bioinformatics: integrating ecological data from the gene to the biosphere. ann. rev. ecol. evol. syst. 37, 519-544. keller, r. a., boyd, r. n., wheeler, q. d. 2003. the illogical basis of phylogenetic nomenclature. bot. rev. 69, 93-110. kennedy, j., hyam, r., kukla, r., paterson, t. 2006. a standard data model representation for taxonomic information. omics 10, 220-230. kennedy, j., kukla, r., raschid, l., paterson, t. 2005. scientific names are ambiguous as identifiers for biological taxa: their context and definition are required for accurate data integration. in: ludäscher,b., raschid, l. (eds.), data integration in the life sciences: proceedings of the second                                                              10 http://www.co-ode.org/resources/tutorials/protegeowltutorial.pdf.  11 http://www.hyam.net/blog/wp-content/uploads/2009/08/perspectives_09.pdf.  12 http://species.asu.edu/files/iise_sos_2009.pdf.  international workshop, san diego, ca, usa, july 20-22. dils 2005, lnbi 3615, pp. 80-95. koperski, m., sauer, m., braun, w., gradstein, s. r. 2000. referenzliste der moose deutschlands. schrift. veget. 34, 1-519. krings, a. 2008. revision of gonolobus s.s . (apocynaceae: asclepiadoideae) in the west indies. j. bot. res. inst. texas 2: 95-138. lee, m. s. y. 2003. species concepts and species reality: salvaging a linnaean rank. j. evol. biol. 16, 179-188. linnaeus, c. 1758. systema naturae p er regn a tria naturae, secundum c lasses, or dines, gene ra, species, cum caracteribus, differen tiis, synonymis, locis. vol. 1. 10th ed. l. salvii, holmiae (sweden). mabee, p. m., ashburner, m., cronk, q., gkoutos, g. v., haendel, m., segerdell, e., mungall, c., westerfield, m. 2007. phenotype ontologies: the bridge between genomics and evolution. trends ecol. evol. 22: 345-350. maddison, d. r., schulz, k.-s., maddison, w. p. 2007. the tree of life web project. zootaxa 1668: 19-40. madin, j. s., bowers, s., schildhauer, m. p., jones, m. b. 2008. advancing ecological research with ontologies. trends ecol. evol. 23: 159-168. marshall, g. a. k. 1922. some injurious neotropical weevils (curculionidae). bull. entomol. res. 13, 5971. mccray, a. t. 2006. conceptualizing the world: lessons from history. j. biomed. inf. 39, 267-273. mckenna, m. c., bell, s. k. 1997. classification of mammals above the species level. columbia university press, new york, ny. merrill, g. h. 2010. ontological realism: methodology or misdirection? appl. ontology 5, 79-108. midford, p. e. 2004. ontologies for behavior. bioinformatics 20, 3700-3701. mikó, i., deans, a. r. 2009. masner, a new genus of ceraphronidae (hymenoptera, ceraphronoidea) described using controlled vocabularies. zookeys 20, 127-153. minelli, a. 2003. the status of taxonomic literature. trends ecol. evol. 18, 75-76. mungall, c. j., gkoutos, g. v., smith, c. l., haendel, m. a., lewis, s. w., ashburner, m. 2010. integrating phenotype ontologies across multiple species. genome biol. 11: r2: 1-16. nielsen, r. 2002. mapping mutations on phylogenies. syst. biol. 51: 729-739. nixon, k. c., carpenter, j. m. 2000. on the other "phylogenetic systematics". cladistics 16, 298-318. o'brien, c. w., wibmer, g. j. 1982. annotated checklist of the weevils (curculionidae sensu lat o) of north america, centralamerica, and the west indies (coleoptera: curculionoidea). mem. am. entomol. inst. 34, 1-382. biodiversity informatics, 7, 2010, pp. 45-66   65   ødegaard, f. 2000. how many species of arthropods? erwin's estimate revised. biol. j. linn. soc. 71, 583597. page, r. d. m. 2004. phyloinformatics: towards a phylogenetic database. in: wang, j. t. l., mohammed, j. z., toivone, h. t. t., sasha, d. (eds.), data mining in bioinformatics. springer, london, pp. 219-241. page, r. d. m. 2006. taxonomic names, metadata, and the semantic web. biodiv. inf. 3, 1-15. page, r. d. m. 2007. tbmap: a taxonomic perspective on the phylogenetic database treebase. bmc bioinformatics 8, 158, 1-11. page, r. d. m. 2009. bioguid: resolving, discovering, and minting identifiers for biodiversity informatics. bmc bioinformatics 10 (suppl. 14): s5: 1-10. peterson, a. t., navarro-sigüenza, a. g. 1999. alternate species concepts as bases for determining priority conservation areas. conserv. biol. 13, 427431. platnick, n. i. 2001. from cladograms to classifications: the road to dephylocode.13 [accessed xii-20-2009]. platnick, n. i. 20099. letter to linnaeus. in: knapp, s., wheeler, q. d. (eds.), letters to linnaeus. linnean society of london, london, pp. 171-184. prendini, l. 2001. species or supraspecific taxa as terminals in cladistic analysis? groundplans versus exemplars revisited. syst. biol. 50, 290-300. proctor, h. c. 1996. behavioral characters and homoplasy: perception versus practice. in: sanderson, m. j., hufford, l. (eds.), homoplasy: the recurrence of similarity in evolution. academic press, san diego, ca, pp. 131-149. pullan, m. r., armstrong, k. e., paterson, t., cannon, a., kennedy, j. b., watson, m. f., mcdonald, s., raguenaud, c. 2005. the prometheus description model: an examination of the taxonomic description-building process and its representation. taxon 54, 751-765. pyle, r. l., earle, j. l., greene, b. d. 2008. five new species of the damselfish genus chromis (perciformes: labroidei: pomacentridae) from deep coral reefs in the tropical western pacific. zootaxa 1671: 3-31. de queiroz, k., gauthier, j. 1992. phylogenetic taxonomy. ann. rev. ecol. syst. 23, 449-480. ramírez, m. j., coddington, j. a., maddison, w. p., midford, p. e., prendini, l., miller, j., griswold, c. e., hormiga, g., sierwald, p., scharff, n., benjamin, s. p., wheeler, w. c. 2007. text linking of digital images to phylogenetic data matrices using a morphological ontology. syst. biol. 56, 283-294.                                                              13 http://www.systass.org/archive/events-archive/2001/platnick.pdf.  rieppel, o. 2007. the performance of morphological characters in broad-scale phylogenetic analyses. biol. j. linn. soc. 92, 297-308. sahoo, s. s., sheth, a., henson, c. 2008. semantic provenance for escience: managing the deluge of scientific data, ieee internet computing 12 (july/august 2009): 46-54. schmit, j. p., mueller, g. m. 2007. an estimate of the lower limit of global fungal diversity. biodivers. conserv. 16, 99-111. schuh, r. t. 2003. the linnaean system and its 250year persistence. bot. rev. 69, 59-78. schuh, r. t., brower, a. v. z. 2009. biological systematics: principles and applications, 2nd edition. cornell university press, ithaca, ny. schulz, s., hahn, u. 2007. towards the ontological foundations of symbolic biological theories. artif. intell. med. 39, 237-250. schulz, s., johansson, i. 2007. continua in biological systems. the monist 90, 499-522. schulz, s., stenzhorn, h., boeber, m. 2008. the ontology of biological taxa. bioinformatics 24, i313i321. scoble, m. j. 2004. unitary or unified taxonomy? phil. trans. r. soc. lond., ser. b 359, 699-711. sereno, p. c. 2005. the logical basis of phylogenetic taxonomy. syst. biol. 54: 595-619. sereno, p. c. 2009. comparative cladistics. cladistics 25: 624-659. smith, b., ashburner, m., rosse, c., bard, j., bug, w., ceusters, w., goldberg, l. j., eilbeck, k., ireland, a., mungall, c. j., the obi consortium, leontis, n., rocca-serra, p., ruttenberg, a., sansone, s.-a., scheuermann, r. h., shah, n., whetzel, p. l., lewis, s. 2007. the obo foundry: coordinated evolution of ontologies to support biomedical data integration. nature biotech. 25, 1251-1255. sober, e. 2002. reconstructing the character states of ancestors: a likelihood perspective on cladistic parsimony. monist 85: 156-176. sogin, m. l., morrison, h. g., huber, j. a., welch, d. m., huse, s. m., neal, p. r., arrieta, j. m., herndl, g. j. 2006. microbial diversity in the deep sea and the underexplored "rare biosphere". proc. natl. acad. sci. u.s.a. 103, 12115-12120. stevens, p. f. 1984. metaphors and typology in the development of botanical systematics 1690-1960, or the art of putting new wine in old bottles. taxon 33, 169-211. strickland, h. e., phillips, j., richardson, j., owen, r., jenyns, l., broderip, w. j., henslow, j. s., shuckard, w. e., waterhouse, g. r., yarrell, w., darwin, c., westwood, j. o. 1843. series of propositions rendering the nomenclature of zoology uniform and permanent. in: report of the twelfth meeting of the british association for the biodiversity informatics, 7, 2010, pp. 45-66   66   advancement of science held at manchester in june 1842; 105-121. thau, d. 2008. reasoning about taxonomies and articulations. in: ph.d. '08: proceedings of the 2008 edbt ph.d. workshop. acm, new york, ny, pp. 11-19. thau, d., bowers, s, ludäscher, b. 2008. merging taxonomies under rcc-5 algebraic articulations. in: conference on information and knowledge management. proceeding of the 2nd international workshop on ontologies and information systems for the semantic web. acm, pp. 47-54. thau, d., bowers, s, ludäscher, b. 2009. merging sets of taxonomically organized data using concept mappings under uncertainty. in: proceedings of the 8th international conference on ontologies, databases, and the applications of semantics (odbase 2009). lecture notes in computer science. otm 2009, pp. 1103-1120. thau, d., ludäscher, b. 2007. reasoning about taxonomies in first-order logic. ecol. inform. 2, 195209. vogt, l. 2009. the future role of bio-ontologies for developing a general data standard in biology: chance and challenge for zoo-morphology. zoomorphology 128, 201-217. vogt, l., bartolomaeus, t., giribet, g. 2009. the linguistic problem of morphology: structure versus homology and the standardization of morphological data. cladistics 25, 1-25. vane-wright, r. i. 2003. indifferent philosophy versus almighty authority: on consistency, consensus and unitary taxonomy. syst. biodiv. 1, 3-11. wagner, g. p. (ed). 2001. the character concept in evolutionary biology. academic press, san diego, ca. weakley, a. s. 2006. flora of the carolinas, virginia, and georgia. working draft of january 17, 200614 [accessed xii-20-2009]. wenzel, j. w., carpenter, j. m. 1994. comparing methods: adaptive traits and tests of adaptation. in: eggleton, p., vane-wright, r. i. (eds.), phylogenetics and ecology. academic press, london, pp. 79-101. wheeler, d. l., barrett, t., benson, d. a., bryant, s. h., canese, k., chetvernin, v., church, d. m., dicuccio, m., edgar, r., federhen, s., feolo, m., geer, l. y., helmberg, w., kapustin, y., khovayko, o., landsman, d., lipman, d. j., madden, t. l., maglott, d. r., miller, v., ostell, j., pruitt, k. d., schuler, g. d., shumway, m., sequeira, e., sherry, s. t., sirotkin, k., souvorov, a., starchenko, g., tatusov, r., tatusova, t. a., wagner, l., yaschenko, e. 2008. database resources of the                                                              14 http://www.herbarium.unc.edu/flora.htm.  national center for biotechnology information. nucleic acids res. 36, d13-d21. wheeler, q. d. 2004. taxonomic triage and the poverty of phylogeny. phil. trans. r. soc. lond. b 359, 571583. wheeler, q. d., meier, r. 2000. species concepts and phylogenetic theory: a debate. columbia university press, new york, ny. whiting, m. f. 1998. long-branch distraction and the strepsiptera. syst. biol. 47, 134-138. wilson, d. e., reeder, d. e. (eds). 2005. mammal species of the world: a taxonomic and geographic reference, 3rd edition. johns hopkins university press, baltimore, md. yoon, n., rose, j. 2001. an information model for the representation of multiple biological classifications. in: alexandrov, v. n., dongarra, j. j., juliano, b. a., renner, r. s., tan, c. j. k. (eds.), computational science – iccs 2001: international conference, san francisco, ca, usa. lecture notes in computer science. springer-verlag, berlin, pp. 937-946. microsoft word page_lsidsd _3_.doc biodiversity informatics, 3, 2006, pp. 1-15 1 taxonomic names, metadata, and the semantic web roderic d. m. page division of environmental and evolutional biology institute of biomedical and life sciences university of glasgow, glasgow g12 8qq, scotland email: r.page@bio.gla.ac.uk abstract.–– life science identifiers (lsids) offer an attractive solution to the problem of globally unique identifiers for digital objects in biology. however, i suggest that in the context of taxonomic names, the most compelling benefit of adopting these identifiers comes from the metadata associated with each lsid. by using existing vocabularies wherever possible, and using a simple vocabulary for taxonomy-specific concepts we can quickly capture the essential information about a taxonomic name in the resource description framework (rdf) format. this opens up the prospect of using technologies developed for the semantic web to add “taxonomic intelligence” to biodiversity databases. this essay explores some of these ideas in the context of providing a taxonomic framework for the phylogenetic database treebase. key words.–– life science identifiers, metadata, taxonomic names, semantic web, rdf, triple stores. integrating diverse sources of digital information is a major challenge facing biodiversity informatics. not only are we faced with numerous, disparate data providers, each with their own specific user communities, but the information in which we are interested is diverse, and includes taxonomic names and concepts, specimens in museum collections, scientific publications, genomic and phenotypic data, and images. of course, this problem is not unique to biodiversity informatics — the wider bioinformatics community is keenly aware of this problem (stein 2003) and indeed it is major topic of discussion concerning the future direction of the world wide web (“web 2.0”). my goal in this paper is to sketch some ideas on how we could create the infrastructure for constructing a distributed system for querying information on biodiversity. my contention is that, thanks to efforts by the semantic web community1 the elements we need are mostly already in place. the two key technologies i will advocate are the resource description format (rdf2) developed by the w3c, and the life science identifier (lsid) technology developed by ibm3. it is easy to enthuse about a technology and contribute to the “hype” that surrounds it, so i will try and keep my feet on the ground by providing some background on 1 http://www.w3.org/2001/sw/. 2 http://www.w3.org/rdf/. 3 http://lsid.sourceforge.net/. the problem that lead me to this conclusion, and by presenting working implementations wherever possible. motivation reconstructing the history of life on earth (the “tree of life”) is the holy grail of phylogenetics, yet we lack a comprehensive phylogenetic database that stores our efforts at reconstructing this tree. the most comprehensive phylogenetic database we currently have is treebase4 (piel et al. 2002). as i've outlined elsewhere (page 2004) a major limitation of this database is that it has no taxonomic “intelligence.” taxonomic names are entered into treebase without being validated against any external database of names, hence many of the names are not proper scientific names. efforts to map names in treebase to external databases rapidly run into problems. around half the names in treebase do not have an exact match in the ncbi's taxonomy database. using data cleaning tools (herbert et al. 2004) or a combination of approximate string matching, regular expressions, and manual matching5 can improve on this, but a significant fraction of names in treebase still have no obvious counterpart in the ncbi's database. in some cases this is because no dna 4 http://www.treebase.org/. 5 http://darwin.zoology.gla.ac.uk/~rpage/treebase. page – the semantic web 2 sequences have been (or indeed, can be) obtained from those taxa, in which case those names will not be in the ncbi database and hence matches may be sought in other taxonomic databases, such as the integrated taxonomic information system (itis6), the international plant names index (ipni7), indexfungorum8, and the universal biological indexer and organizer (ubio9). whereas it is relatively easy to search ncbi’s taxonomy because the entire database can be downloaded, this is not the case for most other taxonomic databases. a taxonomic search engine in 2004 i started to map treebase names onto various taxonomic databases (results can be viewed10). querying these source manually using their web interfaces is slow and tedious, so i developed a simple federated search engine that queries multiple taxonomic databases for information about a name (page 2005). the taxonomic search engine supports two basic queries, namesearch and getdataforid. the first query (namesearch) searches a database for a name, and if the name is found returns the name and its identifier in that database. the second query (getdataforid) “drills down” to get details about a single record in the source database. leaving aside the technical details of talking to databases that support very different query interfaces, i had two problems to deal with. the first was how to generate unique identifiers for names from each database. given that most of the databases use integers as their primary keys, in many cases the same identifier will be used by different databases. as one of many possible examples, 101593 is the identifier for odonata in ncbi's genbank and dahlia australis in itis. hence, if i were to store just the identifier it would not be clear what name that identifier referred to. an obvious solution is the idea of a “namespace” that specifies the context for a given identifier. in this case, the identifier from ncbi could be distinguished from that in itis by adding prefixes corresponding to the domain name address of the two databases, i.e., adding “ncbi.nlm.nih.gov” and “itis.usda.gov” to the respective identifiers. 6 http://www.itis.usda.gov/. 7 http://www.ipni.org/. 8 http://www.indexfungorum.org/. 9 http://wwww.ubio.org/. 10 http://darwin.zoology.gla.ac.uk/~rpage/treebase. the second problem is how to return information about a specific record in a database (e.g., the name, any synonyms, etc.). given that each database has its own format for returning information (ranging from delimited text, html, xml, and soap data structures), i transformed the result returned by each database into a common xml format that in turn could be transformed into html output for display in a web browser. so, to facilitate mapping names in treebase onto names in external databases we need (1) a mechanism for generating globally unique identifiers, and (2) standard format for providing information about the object the identifier refers to. before introducing one possible solution, let us first consider why names themselves are not enough. why taxonomic names aren't enough the taxonomic name of an organism is a key link between different databases that store information on that organism. however, taxonomic names themselves have serious limitations as identifiers in databases (kennedy 2003; kennedy et al. 2005) due to the existence of multiple names (synonyms) for the same taxon, and the use of the same name to refer to different taxa. for example, the genus morus applies to both an animal (the gannet) and a plant (the mulberry tree). even species names can be identical — a species of wasp and a species of conifer both share the name agathis montana. furthermore, there may be multiple names for the same taxon. hence, using names alone to link different data sources can be prone to error. as an example, at the time of writing ncbi's linkout feature mistakenly links the catfish genus loricaria (ncbi tax_id = 52085) to the treebase taxon loricaria (treebase taxonid = 1305), which is a plant genus (family compositae). this lack of uniqueness of names raises the issue of how to store taxonomic information in databases. uris, urls, and urns life science identifiers (lsid) are one solution to the problem of globally unique identifiers (clark et al. 2004). at the risk of drowning the reader in alphabet soup, it is useful to distinguish between two different types of identifiers in use in the internet, the uniform resource locator (url), and the uniform resource name (urn). urns and urls are two possible kinds of uniform resource identifier (uri). page – the semantic web 3 figure 1. the components of a life science identifier (lsid). most readers will be familiar with urls, which specify the location of a resource the internet (e.g., http://www.ubio.org/soapbrowser/index.php?func=n ame_detail&ubioid=454488), that is, they “point” to it. they can, in principle serve as a unique identifier, however they are prone to breakage — if the resource being pointed to moves, the url no longer points to the resource, leading to the dreaded “404 page not found” problem (dellavalle et al. 2003). urns, in contrast, provide a persistent name for a resource, but typically do not provide any information on how to access that resource. a lsid is a uniform resource name (urn). digital object identifiers (dois) are another example of a urn, and are widely used in the publishing industry to identify electronic publications. if the resource moves (e.g., one publishing house acquires another, and moves the acquired company’s digital resources to a new server) the resource still retains the original doi. the utility of a urn is somewhat limited, unless there is a mechanism to resolve the urn, that is, to retrieve the named resource. in the case of dois, the simplest way to see this mechanism in action is to append a doi, such as 10.1145/1024694.1024703, to the url11 giving in this instance12, and open the resulting url in a web browser. in this example the doi resolves to the electronic version of herbert et al. (2004). life science identifiers figure 1 shows an example lsid. each lsid is prefixed by ‘urn’ indicating that the lsid is a urn, ‘lsid’ indicates that the identifier is a lsid, then follow the authority, namespace, and identifier components. there may also be an optional revision component to indicate the version of the resource. the authority is a domain name that can be resolved by the 11 http://dx.doi.org/. 12 http://dx.doi.org/10.1145/1024694.1024703. internet dns (typically a domain name owned by the data provider), the namespace and identifier are specific to the data source which provides the resource. in this case the lsid is a taxonomic name in the ubio database. the authority ‘ubio.org.lsid. zoology.gla.ac.uk’ is a domain name of a server at the university of glasgow that serves lsids for ubio records. if ubio itself served lsids, the domain name could be ubio.org. note that the uniqueness of the lsid is in part guaranteed by the use of internet domain names, which are globally unique. providing that the data source ensures that each combination of namespace and identifier is unique within the data source, the lsid itself will be a globally unique identifier. a lsid is intended to refer to one unchanging digital object. hence, if two users retrieve data with the same lsid, they will have the exactly the same data. this contrasts with urls, where the content may change at any time (for example, if the author of the web page changes the layout). different versions of a digital object can be identified using the revision part of the lsid. in addition to data there may be metadata associated with a lsid. the lsid standard doesn’t require that the metadata remain unchanging. the life science identifier (lsid) standard specifies a mechanism for resolving a lsid and retrieving the data and/or metadata associated with that lsid. because a lsid is not a url, you can't simply paste a lsid into a web browser unless you have additional software installed, such as ibm's lsid launchpad for internet explorer 13 (figure 2) or the lsid extension for firefox14. the biopathways consortium provides a web-based lsid resolver15. the lsid shown in (figure 2) resolves to the ipni 13 http://lsid.sourceforge.net/ 14 http://lsid.mozdev.org/ 15 http://lsid.biopathways.org page – the semantic web 4 figure 2. ibm's lsid launchpad displaying metadata associated with a lsid, this case urn:lsid:ipni.org.zoology.gla.ac.uk:id:20012728-1. record for the name poissionia heterantha (a plant species). being a urn, the lsid is a name, not an address. currently few data sources serve their own lsids, the north temperate lakes long term ecological research project16 being the notable exception. third parties, such as the biopathways consortium, provide most working lsid authorities. search engine revisited the taxonomic search engine (page 2005) provides lsids for each source database that it queries. in order to provide metadata for an lsid, i needed to extract information from each source. one thing which struck me during the development of this search engine was that much of the code i used for the getdataforid method was being reused when implementing lsid authorities (i use the term “reused” loosely, the code had to be ported from the php scripting language to perl). hence, if a taxonomic database serves lsids, all a taxonomic search engine needs to provide is a means to search those databases (i.e., the namesearch method). 16 http://lsid.limnology.wisc.edu/ metadata having recounted how i was drawn to lsids in the context of trying to make sense of the taxonomic names in treebase, i will now turn to metadata. the lsid standard doesn't specify that metadata be in any particular format, but the resource description format (rdf17) is the most widely used. rdf is a framework for describing relationships between resources, where a resource is connected to another resource by a property. a resource must have a uri. the basic unit in rdf is the subject, property, object triple (figure 3). figure 3. a rdf subject, property, object triple. 17 http://www.w3.org/rdf/ page – the semantic web 5 the object is either a literal string, or another resource. rdf can be represented in a number of ways, most commonly using xml. this is a simple rdf document stating that the world wide web consortium is the publisher of the resource18. world wide web consortium rdf can be represented in various formats, but xml is most commonly used. the first tag in the document lists the namespaces being used in the document, which ensure that there is no “collision” between terms from different vocabularies, and that there is minimal ambiguity in interpreting a term. for example, dc:publisher tells us that the term publisher is part of the dublin core vocabulary, figure 4. an example rdf statement. which has the uri19. note that unlike other xml documents, the namespace uri must be retrievable. in the case of dublin core, opening the uri in a web browser will return an rdf document describing all the terms in that vocabulary. namespaces also help make the document more human-readable by defining prefixes that can be used in the document: rdf:rdf is more readable than http://www.w3.org/1999/ 02/22-rdf-syntax-ns#rdf. the same information shown in this example can be represented as a graph (figure 4). rather than attempt a tutorial on rdf i will illustrate its use with examples. for more background on rdf i recommend powers (2003). 18 http://www.w3.org/ 19 http://purl.org/dc/elements/1.1/ why rdf? it might be tempting to think of rdf as just another variant of xml, and given the substantial investment the taxonomic community has made in developing xml schema for specimens (e.g., abcd access to biological collections data20, darwin core21, and taxa22 (taxonomic concept schema) (kennedy et al. 2005), one could ask why should we contemplate an alternative format? wang et al. (2005) provide an insightful comparison of xml schema and rdf in the context of ‘omic’ data standards (e.g., genomics, proteomics. etc.). they conclude that xml schema solve the problem of standardising communication between different data sources at the level of messages, but is not equipped to convey semantics (i.e., meaning). rdf is explicitly designed to model semantics, and i suggest that this is where the real power of lsids resides. storing and querying rdf rdf triples can be stored in a database called a “triple store.” a triple store can be thought of as a database with a single table containing multiple rows, each row containing the subject, property, object triple. for this work i use 3store version 2.2.1823 from the university of southampton (harris and gibbins 2003). 3store uses a mysql database to physically store the triples, and supports the rdf data query language (rdql). a more recent version of this software adds support for the sparql language. taxonomic metadata this section explores the notion of taxonomic names as metadata. i am not going to attempt to develop an explicit “standard” for taxonomic metadata at this point. rather, i want to sketch a possible form the metadata could take, then explore what we can do with such metadata. hence, in the interests of illustrating the ideas i’m going to skirt around the complexities of taxonomic names and concepts (berendsohn et al. 2003; kennedy et al. 2005). figure 5 shows a graph representing the rdf metadata for a single record in the itis database for the taxon morus bassanus. 20 http://ww3.bgbm.org/abcddocs/ 21 http://darwincore.calacademy.org/ 22 http://www.soc.napier.ac.uk/tdwg/index.php 23 http://www.aktors.org/technologies/3store/ page – the semantic web 6 figure 5. graph representing the rdf metadata associated with the lsid for the itis record for the northern gannet morus bassanus. boxes represent literals, ovals represent resources (typically identified by their lsid, which to avoid clutter are abbreviated tsn:##### the labelled edges of the graph represent properties, such as the source of the information, the taxonomic name, rank, status, authority, and lineage (represented as a rdf:seq collection). source basic information on the source of the data, such as the name of the source database, the original source of the data (e.g., a contributing taxonomic database or specialist), and the date the record was created, can be represented using terms from the dublin core. copyright information on what users can and cannot do with the data can be expressed (typically in free format text) using dublin core, or perhaps more efficiently using one of the licenses developed by the creative commons24. the advantage of the latter is that these licenses are computer-readable, and hence software can discover what it can do with the data without requiring human intervention. bibliographic data some taxonomic databases provide bibliographic information. there are various ways this can be represented. the dublin core vocabulary provides dcterms:bibliographiccitation which can contain the bibliographic information in any suitable format. this is useful where the database stores the bibliographic information as unstructured text (e.g., a single field in the database contains the bibliographic record). if the database stores the bibliographic information in a more structured form, then a more expressive vocabulary such as publishing requirements for industry standard metadata (prism25) could be used. rss feeds provided by journal publishers make extensive use of this vocabulary (hammond et al. 2004). it would also be highly desirable to have identifiers for the publication, such as a doi26 or a pubmed number. person although i have not made use of it in this context, the friend of a friend project (foaf27) has developed a useful vocabulary for describing people, which could be used to describe authors of scientific names. taxonomic name a simple approach to representing the taxonomic name itself is through use of the dc:title and rdfs:label properties. one reason for adopting these tags is for consistency with tools such as lsid launchpad (figure 2), which use these tags to display a title for lsid metadata. for some information about taxonomic names, such as the authors of the name, the taxonomic rank, 24 http://creativecommons.org/ 25 http://www.prismstandard.org/ 26 http://www.doi.org/ 27 http://www.doi.org/ page – the semantic web 7 and synonymy, we need a vocabulary that is specific to biological taxonomy. i will use a simple vocabulary that describes the bare minimum of terms need. in the examples below i will use the namespace prefix gla, which is short for urn:lsid:lsid.zoology. gla.ac.uk:predicates, such that each property is defined by a lsid. this means that for any property there is associated metadata that describes that property. relationships among names given that each name has an lsid, we can describe the relationship between two names using the appropriate property. there are various kinds of relationships among taxonomic names that we might wish to model, of which synonymy is perhaps the most important. there are various kinds of synonyms, which we can loosely characterise as “objective” and “subjective.” two names are objective synonyms if there is no doubt that the two names refer to the same taxon. for example, if a species moves from one genus to another a new name combination results. these two names clearly refer to the same entity — they have the same type specimen. in other cases whether the name refers to the same entity may be contentious. consider the case where there are two species names, each with different holotypes. one taxonomic authority may regard these species as distinct, a second authority may regard them as the same species, and hence treat the names as synonyms. in the later case, this is an inference based (ideally) on data, not a simple consequence of the appropriate rules of nomenclature. as an example of objective synonymy, the plant originally described as tephrosia heterantha griseb. has been placed in the genus poissonia heterantha by lavin et al. (2002). hence, tephrosia heterantha is the basionym of poissonia heterantha. we can represent this relationship in rdf like this: poissonia heterantha where the names poissonia heterantha and tephrosia heterantha have the lsids urn:lsid:ipni.org.lsid.zoology.gla.ac.uk :id:20012728-1 and urn:lsid:ipni.org.lsid.zoology.gla.ac.uk :id:520610-1, respectively. these lsids are generated by the taxonomic search engine (page 2005) for data in the ipni database. the inverse relationship is gla:isbasionymof. according to the ipni database, tephrosia heterantha is also the basionym for coursetia heterantha (figure 6), which we can represent like this: tephrosia heterantha figure 6. two names linked to their basionym. poissonia heterantha and coursetia heterantha are synonyms because they share the same basionym, tephrosia heterantha. kinds of synonyms the depth of information on synonymy varies across taxonomic databases, and even within databases. for example, ipni comprises three source databases: index kewensis, the gray card index, and the australian plant names index. the gray card index specifies whether two names are “nomenclatural synonyms,” but does not for example specify the exact nature of the relationship (i.e., whether one name is a basionym of the other). based on the kinds of synonyms reported by the data sources queried by the taxonomic source engine, we can construct a graph depicting the relationships between synonym types (figure 7). these relationships are modelled by a page – the semantic web 8 simple rdf schema (rdfs28). the property gla:synonym has the sub-properties gla:objectivesynonym and gla:hasacceptedname. the later is based on the term used in itis to specify which name that database regards as the correct name to use for a taxon. the accepted name may in fact be an objective synonym, but because itis does not provide sufficient information to determine that, i've chosen not to regard this as an objective synonym. if a database merely specifies that a name is a synonym then the gla:synonym property can be used. the following is an abbreviated version of this schema: synonym objective synonym basionym of 28 http://www.w3.org/tr/rdf-mt/ in this schema synonyms are modelled using the rdf:property property, and if one kind of synonym is a refinement of another we use the rdsf:subpropertyof property. hence, the isbasionymof property is a refinement of the objectivesynonym property. if we ask for the objective synonyms of a name, we should recover any synonym that either has the property objectivesynonym, or a property that is a subproperty of objectivesynonym. as an example, consider the taxon names shown in figure 6 and the corresponding rdf: poissonia h eterantha coursetia heterantha tephrosia heterantha page – the semantic web 9 the following rdql query asks for the objective synonyms of tephrosia heterantha (urn:lsid:ipni.org.lsid.zoology.gla.ac.u k:id:520610-1): select ?synonym where (, , ?id), (?id, , ?synonym) using gla for , dc for this query returns the result ?synonym "poissonia heterantha" "coursetia heterantha" note that our rdql query doesn't mention the property isbasionymof by name, but it returns the two names for which tephrosia heterantha is the basionym because isbasionymof property is a refinement of the objectivesynonym property. put another way, if a name is an isbasionymof of another name, this entails that the name is also an objectivesynonym. unfortunately, support for entailment in rdf query languages is currently patchy (broekstra 2005). given the rdf schema above we can see that if a name a is a basionym of name b, then this entails the following statements: 1. a is an objective synonym of b 2. a is a synonym of b the rdql engine that comes with version 2.2.18 of 3store will infer the first statement, but not the second. as a (i hope) temporary fix until support for entailment improves, we can add additional edges to the graph to ensure that rdql can also infer statement 2. inferring synonymy once we start modelling relationships between names using graphs (which comes naturally if we adopt rdf) then we can quickly discover limitations in the information provided by taxonomic databases. for example, in the case of poissonia heterantha and coursetia heterantha shown in figure 6, there is no direct link between these two names (in graph terminology, there are no edges connecting the two nodes). this is because the ipni records for poissonia heterantha and coursetia heterantha do not make explicit that these two names are synonyms — the user has to work this out from the fact that both names share the same basionym. we can make this relationship explicit by computing the transitive closure of the graph. the transitive closure of a graph is obtained by adding an edge between nodes i and j if node j is reachable from node i, that is, there is a path from node i and node j. figure 8a shows the original graph representing the relationships between poissonia heterantha, coursetia heterantha, and tephrosia heterantha, and the transitive closure of that graph (figure 8b), showing that poissonia heterantha and coursetia heterantha are synonyms. vernacular names vernacular or common names exist for many taxa, and in many languages. we can represent these using a vernacularname tag, and use the xml:lang attribute to specify the language. common swift parts of names for names that have more than one part there is a logical relationship between those parts. for example, the trinomial coursetia caribaea caribaea is a variety of the species coursetia caribaea, which in turn is a species of the genus coursetia. some database models (pyle 2004) store the individual components of a multinomial name separately to facilitate tracking name changes, and searches for epithets. under such a model, the variety caribaea has as its parent the species caribaea, which in turn is a child of the genus coursetia. although this relationship also expresses a classification (i.e., that this species belongs in the genus coursetia), it is useful to separate the notion of a logical relationship between the parts of a name (a species name must include both genus name and specific epithet) and a classification. typically in a classification only the accepted names are part of the parent-child hierarchy. the basionym for coursetia caribaea is galega caribaea, hence in this instance caribaea is a child of galega. even though galega caribaea is not the accepted name for this plant, the page – the semantic web 10 figure 7. examples of different types of synonym. the property rdfs:subpropertyof links a more restricted kind of synonym to a more general kind of synonym. dotted lines indicate edges that need to ensure that rdql infers all the relationships entailed by the property rdfs:subpropertyof (see the text). logical relationship between the two components of the name remains. we can model this relationship using the dcterms:ispartof and dcterms: haspart properties from the dublin core terms (figure 9). classification a classification of a set of taxa can be regarded as a rooted tree with nodes are taxa and are labelled with the accepted names of those taxa. this, for example, is the model used by itis. in the case of taxa below the rank of genus, for the accepted name the classification will mirror the dcterms:ispartof and dcterms:haspart relationships among the components of the name. it is tempting to model this taxonomic hierarchy using the rdfs:subclassof property, particularly as rdfs:subclassof is transitive. if each taxon in a classification knows its parent taxon, then we could traverse the path from the tip to the root of a classification by using rdfs:subclassof. however, consider what happens if we want to go in the other direction, for example, if we want to find all birds. we would need to go through each node in the classification and go down the path to the root until we either (a) find “aves” or (b) reach the root of the tree. we can speed things up by storing for each node the path from that node to the root of the tree. this path is the node’s lineage. one natural way to do this would be to use the rdf:seq class, which stores an ordered sequence of uris. for example, the lineage for morus bassanus in itis is: page – the semantic web 11 figure 8. (a) the relationships between three names for the same plant (figure 7), and (b) the transitive closure of the graph. an edge between two nodes indicates that the corresponding names are synonyms. we use a sequence because order is important (in this example the order is from higher to lower taxon). by storing the lineage for a node, we can find all members of a given higher taxon by finding those that have the corresponding uri in their lineage. given than both approaches to representing a classification have their merits, here i use both: rdfs:subclassof to specify the immediate parent, and rdf:seq to specify the complete lineage. links between data sources some databases store links to information associated with the same name in other databases. examples include ncbi linkout, which links a taxon in genbank to various external databases, and ubio, which records the source of names it stores in its data warehouse. these links can be modelled using the rdfs:sameas property. this (and other) properties can be used to represent relationship between names in different databases, but also other data, such as specimens, sequences, and publications. for example, metadata for a specimen could link the name assigned to that specimen to the record for that name in ubio. treebase (again) federated search engines, globally unique identifiers, metadata, and the semantic web seem somewhat removed from the original problem, namely making sense of treebase. let me work through an example of how i think these technologies can help. the ultimate goal is to be able to query the treebase database with a taxonomic name, and recover all the studies that contain that taxon. there are some significant obstacles in our way. in many cases, the names in treebase (stored in a field called taxonname) are not names of taxa but names of operational taxonomic units (otus). that is, the names might be a combination of taxon name and some other identifier, such as a specimen or voucher code, or a genbank accession number. for example, in treebase study s813 (lavin et al. 2002) we have taxonname records such as “poissonia heterantha 5832”, “poissonia heterantha 5862”, and so on. another obstacle is that treebase does not have any notion of synonymy, so that if two studies refer to the same taxon using different names, treebase cannot tell the user that the two taxa are, in fact, the same. one solution would be to map each taxon name in treebase to one or more external data sources, and to use information from that external data source — such as lists of synonyms — to improve the performance of our query. name changes in legumes as our phylogenetic knowledge of a group of organisms grows it is not uncommon for this new understanding to be reflected in taxonomic changes. consequently, names used for the same taxon in successive studies submitted to treebase may have changed. as a concrete example, different names have been used for same legume taxa in successive papers published by matt lavin and collaborators. hence, treebase study s754 (lavin et al. 2001) includes names such as coursetia heterantha and c. weberbaueri. however, in study s813 (lavin et al. 2002) these taxa were moved to the genus poissonia. ideally, if we search treebase for “coursetia heterantha” we should find both study s754 and s813, because both studies contain this taxon, albeit under different names. one way to add this “taxonomic intelligence” to treebase would be to map names in that database to external name sources that knew the relationship between the different names. page – the semantic web 12 figure 9. an example of the relationships between the component parts of the name coursetia caribaea caribaea, modelled using the dcterms:ispartof and dcterms:haspart properties. mapping names to create the mapping i did the following. each treebase taxonname for study s813 was “cleaned” to remove extraneous suffixes indicating voucher codes, etc., using ubio's findit web service. the resulting cleaned names were then submitted to ipni using to the taxonomic search engine web service (page 2005). for each ipni name the corresponding lsid was resolved and the metadata obtained for that lsid (and for any lsid referred to by the metadata). at each step we incrementally gain some knowledge, which can be expressed in rdf and added to a triple store. for, example, we can represent what treebase knows about a taxonname in rdf: poissonia heterantha 5832 t30615 this simply asserts that there is a taxon with the identifier “t30615” and the name “poissonia heterantha 5832”. once we have cleaned the name and mapped it to ipni we can add a statement asserting that treebase taxonid t30615 is the same as ipni record 20012728-1: the attribute rdf:id="t83" is explained below in the section on reification. the final information we can add to our triple store is the metadata for the ipni record (shown here in abbreviated form): poissonia heterantha this states that the taxon’s name is “poissonia heterantha”, and its basionym is urn:lsid:ipni.org.lsid.zoology.gla.ac.uk :id:520610-1. queries once we have mapped each treebase taxonname, we can now search for treebase taxa using a proper scientific name, e.g.: select ?taxonid, ?taxonname where (?ipni, , "poissonia heterantha") (?tb, , ?ipni ) (?tb, , ?taxonname ) (?tb, , ?taxonid ) using dc for , rdfs for this query finds the ipni record for the name “poissonia heterantha”, then finds all treebase taxa that are mapped to that ipni record and lists their taxonid and taxonname. for study s813 this query yields: ?taxonid ?taxonname "t30610" "poissonia heterantha" page – the semantic web 13 "t30615" "poissonia heterantha 5832" "t30616" "poissonia heterantha 5862" "t30714" "poissonia heterantha 5856" "t30590" "poissonia heterantha 5843" "t30715" "poissonia heterantha 5800" "t30713" "poissonia heterantha 5860" "t30589" "poissonia heterantha 5785" query expansion and taxonomic intelligence simply finding taxa by name doesn't add much to treebase, and indeed we could retrieve all variants of “poissonia heterantha” directly via the web interface to treebase by using the search term “poissonia heterantha@”. however, the metadata for the ipni records enables us to query by synonyms. for example, we know that coursetia heterantha is a synonym of poissonia heterantha (figure 6), hence we should be able to use “coursetia heterantha” as a search term and retrieve records for poissonia heterantha, as in the following rdql query: select ?taxonid, ?taxonname where (?ipni, , "coursetia heterantha") (?ipni,,?basionym) (?basionym,,?synonym) (?tb, , ?synonym ) (?tb, , ?taxonname ) (?tb, , ?taxonid ) using dc for gla for rdfs for for study s813 this query yields: ?taxonid ?taxonname "t30610" "poissonia heterantha" "t30615" "poissonia heterantha 5832" "t30616" "poissonia heterantha 5862" "t30714" "poissonia heterantha 5856" "t30590" "poissonia heterantha 5843" "t30715" "poissonia heterantha 5800" "t30713" "poissonia heterantha 5860" "t30589" "poissonia heterantha 5785" this result is the same as if we had searched for “poissonia heterantha”. this query is more complex than it needs to be because, as discussed above (see figure 8), ipni does not explicitly state that poissonia heterantha and coursetia heterantha are synonyms. consequently, in the metadata retrieved from ipni there is no triple where the subject is poissonia heterantha and the object is coursetia heterantha (or visa versa). however, the key point is that by combining metadata from treebase and ipni, we have added the ability to handle synonyms to treebase. reification although sometimes called the “big ugly of rdf (powers 2003), reification is a useful technique which enables us to make statements about statements. for example, when mapping names in treebase onto names in an external database, some names have an exact match (such as taxonid t30610 “poissonia heterantha”), and other names match once they have been cleaned (for example, taxonid t30615 “poissonia heterantha 5832”). it would be useful to be able to state what type of match was obtained for each name, and reification provides a mechanism for doing so. consider this snippet of rdf: which states that treebase taxonid t30610 is the same as ipni record 20012728-1. the rdfs:sameas statement has a rdf:id attribute with the value “t83”. this enables us to refer to this statement and hence provide details what we mean by “same as”, for example: approximate this rdf triple informs us that statement t83 is an approximate match (i.e., the string “poissonia heterantha 5832” is not an exact match for the string “poissonia heterantha”). note that we could add additional statements, for example describing the degree to which the two strings differ, a date and time stamp for when the link between the two names was made, and any comments on the link. it might seem simpler to represent the kind of match like this (i.e., no reification): page – the semantic web 14 approximate however, in this case the dc:description property refers to the treebase taxon t30615, and not the statement that links this taxon with the ipni record. discussion my intention here has been to sketch out the case for using rdf and lsids to model taxonomic names. in one sense these two technologies are complementary, although strictly speaking neither requires the other. rdf requires uris for resources, and lsids are a natural candidate for these uris, but some other uri could be used. lsids have associated metadata for which rdf is the obvious candidate, but other formats could be used. however, i think these two technologies have the advantage of being available, relatively well understood, and in the case of rdf, part of a much broader effort (the semantic web). rdf is a relatively simple format, but the existence of numerous vocabularies that are relevant to biodiversity informatics, and its support for inference makes it potentially very powerful. it also introduces the notion of modelling relationships between taxonomic names using graphs, which can yield new insights (such as computing all synonyms using transitive closure). i am fully aware that the relationships between taxonomic names are more complicated that i've sketched here (berendsohn et al. 2003; kennedy et al. 2005). i have deliberately tried to keep the rdf described here as simple as possible, in the belief that keeping things simple and maximising the use of existing vocabularies is more productive than trying to design complex, domain specific schema. however, one area that could be usefully explored is the use of ontologies to explicitly model relationships between properties and to provide better support for inference. we have also seen that playing with even simple examples illustrates the potential to make inferences from very simple metadata, and also that existing taxonomic databases lack taxonomic intelligence, that is, they don't always know what they know. future if we adopt rdf for modelling taxonomic names and other objects in biodiversity informatics, then this leads naturally to rethinking the way we might develop biodiversity databases. most work in this area concentrates on using relational databases to store data (morris 2005) and xml schema for exchanging data (e.g., abcd, darwin core, and taxonomic concept schema) (kennedy et al. 2005). both these technologies have a role to play. relational databases support data integrity and a sophisticated query language (sql), however they have limitations — database schema can rapidly become large, complex, and domain specific. furthermore, the emphasis in designing such schema is on internal data integrity, rather than relationships with external data sources. this is a major limitation in an environment where most data is stored elsewhere. xml schema are good at describing messages, but poor at communicating meaning (wang et al. 2005). like relational database schema, xml schema can rapidly become large and unwieldy. a different (but complementary) vision is of a triple store containing rdf triples, where a globally unique identifier explicitly identifies every resource. where possible, identifiers can be resolved to a source of metadata, itself in rdf format. as shown in the treebase example above, one can rapidly construct and populate a triple store which supports inference of the sort that is relevant to the question at hand. simply specifying a relationship between two names, and adding the metadata for those names enables us to infer relationships between different data elements in that database. i am not suggesting that rdf and triple stores are the only technology we should use — relational databases and xml schema have important roles to play. rather, i suggest that rdf and triple stores are ideally suited to support a distributed query system of the kind that biodiversity informatics aspires to provide. resources the rdf examples presented here (often somewhat abbreviated) are available online29. 29 http://darwin.zoology.gla.ac.uk/~rpage/lsid/examples page – the semantic web 15 acknowledgements part of this work was funded by bbsrc grant bb/c004310/1. i thank the reviewers (walter berendsohn, anton güntsch, markus döring, and one who remained anonymous) for their helpful comments. literature cited berendsohn, w. g., m. döring, m. geoffroy, k. glück, a. güntsch, a. hahn, r. jahn, w.-h. kusper, j. li, d. röpert, and f. specht. 2003. moretax: handling factual information linked to taxonomic concepts in biology. bundesamt für naturshutz, bonn. broekstra, j. 2005. storage, querying and inferencing for semantic web languages. phd thesis. vrije universiteit. clark, t., s. martin, and t. liefeld. 2004. globally distributed object identification for biological knowledgebases. briefings in bioinformatics 50:59-70. dellavalle, r. p., e. j. hester, l. f. heilig, a. l. drake, j. w. kuntzman, m. graber, and l. m. schilling. 2003. going, going, gone: lost internet references. science 302:787-788. hammond, t., t. hannay, and b. lu. 2004. the role of rss in science publishing: syndication and annotation on the web. sigmod rec. 10. harris, s., and n. gibbins. 2003. 3store: efficient bulk rdf storage. proceedings of the 1st international workshop on practical and scalable semantic systems (psss'03), sanibel island, florida, 1-15. herbert, k. g., n. h. gehani, w. h. piel, j. t. wang, and c. h. wu. 2004. bio-ajax: an extensible framework for biological data cleaning. sigmod rec. 33:51-57. kennedy, j. 2003. supporting taxonomic names in cell and molecular biology databases. omics: a journal of integrative biology 7:13-16. kennedy, j., r. kukla, and t. paterson. 2005. scientific names are ambiguous as identifiers for biological taxa: their context and definition are required for accurate data integration. pp. 80-95 in b. ludäscher and l. raschid, eds. data integration in the life sciences: second international workshop, dils 2005, san diego, ca, usa, july 20-22, 2005 (lecture notes in computer science 3615). springer verlag. lavin, m., m. f. wojciechowski, p. gasson, c. e. hughes, and e. wheeler. 2002. phylogeny of robinioid legumes (fabaceae) revisited: coursetia and gliricidia recircumscribed, and a biogeographical appraisal of the caribbean endemics. systematic botany 28:387-409. lavin, m., m. f. wojciechowski, a. richman, j. rotella, m. j. sanderson, and a. beyra-matos. 2001. identifying tertiary radiations of fabaceae in the greater antilles: alternatives to cladistic vicariance analysis. international journal of plant science 162:s53-s76. morris, p. j. 2005. relational database design and implementation for biodiversity informatics. phyloinformatics 7:2. page, r. d. m. 2004. phyloinformatics: towards a phylogenetic database. pp. 219-241 in j. t. l. wang, m. j. zaki, h. t. t. toivonen and d. shasha, eds. data mining in bioinformatics. springer verlag. page, r. d. m. 2005. a taxonomic search engine: federating taxonomic databases using web services. bmc bioinformatics 6:48. piel, w. h., m. j. donoghue, and m. j. sanderson. 2002. treebase: a database of phylogenetic knowledge. pp. 41-47 in j. shimura, k. l. wilson and d. gordon, eds. to the interoperable "catalog of life" — with partners species 2000 asia oceanea. research report from the national institute for environmental studies no. 171, tsukuba. powers, s. 2003. practical rdf. o'reilly, sebastopol, california. pyle, r. l. 2004. taxonomer: a relational data model for managing information relevant to taxonomic research. phyloinformatics 1. stein, l. 2003. integrating biological databases. nature reviews: genetics 4:337-345. wang, x., r. gorlitsky, and j. s. almeida. 2005. from xml to rdf: how semantic web technologies will change the design of 'omic' standards. nature biotechnology 23:1099-1103. microsoft word final2.doc biodiversity informatics, 4, 2007, pp. 27-36 online biodiversity resources – principles for usability s. h. neale, m. r. pullan and m. f. watson royal botanic garden edinburgh, 20a inverleith row, edinburgh, eh3 5lr u.k. s.neale@rbge.org.uk abstract.—online biodiversity portals and databases enabling access to large volumes of biological information represent a potentially extensive set of resources for a variety of user groups. however, in order for these resources to live up to their promise they need to be demonstrably useful to the communities they are intended to serve. we discuss a number of principles that can be applied to portal development that assist in defining the scope of user communities, determining their requirements within the context of the data available and establishing realistic goals for a portal or portal development tools. we highlight a lack of user involvement and formalised requirements analysis in biodiversity portal projects to date, and compare this with a similar project in the astrophysics community. it is concluded that the poor understanding of both the users and their tasks that arises from this lack of analysis makes it difficult to assess the success of a portal and increases the risk of the portal being judged to have failed. we suggest that a change in the way large biodiversity portal projects are managed, presented and funded could lead to an increased perception of success with minimal change in the underlying infrastructure, yet enhancing the life expectancy of such projects. key words.—portal, user interface, requirements analysis, end-users online portals enabling access to biodiversity information are intended to increase both the amount of, and ease of access to data, and are claimed to have potential value for a great variety of users. they are intended to facilitate the investigation of complex biological questions and to allow researchers to develop new insights through the agglomeration of large data sets and provide access to appropriate analytical tools. for example, it has been reported that studies of the effect of climate change on flora and fauna (e.g., root et al. 2003, parmesan and yohe 2003) and the identification of priority areas for conservation (e.g. williams et al. 1996) will be enhanced through access to larger integrated datasets. biodiversity portals may also hold the potential for enabling better informed planning and policy decisions (sánchez-cordero and martínez-meyer 2000, jones and thornton 2003, peterson and shaw 2003). however, it is easy to overlook the risks involved with the development of such resources. a major study of over 8000 software development projects found that only 16% were considered successful in that they were completed on time, within budget and with all the features and functions as originally specified (standish group 1994). a more recent study indicates that this situation is improving but still only 29% of projects are considered successful (standish group 2006). the major factors responsible for this were found to be inadequate capture of user requirements and insufficient user involvement. there is no reason to believe that the development of biodiversity portals should be considered as being different from other software development projects. poorly designed biodiversity portals are just as likely to be perceived as having failed and to remain under-utilised. in this paper we therefore aim to investigate the extent to which a set of well known user–oriented principles for designing successful systems have been applied to the development of online biodiversity information resources, and suggest ways in which such projects could be better managed and presented so as to enhance their chances of success. user-oriented software design principles a large body of research exists within a set of overlapping disciplines (e.g. human computer interaction, requirements engineering and usability engineering) which discusses principles and methods of elucidating end-user requirements and incorporating them into design. although portal development is a relatively new phenomenon, many of these ideas and techniques still apply (waloszek 2001, zazelenchuk and boling 2003). the ultimate aim of understanding end-users and their requirements is to develop a product that is both useful and usable, the combination of which determine the usability of a system (löwgren 1995). the international organisation neale et al. – online biodiversity resources – principles for usability 28 for standardisation (iso) has provided the following definition of usability: “the extent to which a product can be used by specified users to achieve specified goals with effectiveness, efficiency and satisfaction in a specified context of use.” (iso 1998). the basic concept of usability has been used as the basis for establishing a set of principles that can guide a user-oriented design process (gould and lewis 1985, hewett and meadow 1986, gould et al. 1991, mulligan et al. 1991). whilst these principles, outlined below, appear rather intuitive they nevertheless provide a logical path through the design process. unfortunately they are not frequently applied to product development (gould et al. 1991, kujala et al. 2001). early and continued focus on users and tasks the involvement of users in portal development, including obtaining an understanding the tasks that they undertake and how these can be supported, is clearly important for the development of usable resources. indeed, users have been shown to be more satisfied with software design arising from thorough requirements analysis (kujala and mäntylä 2000; kujala 2002), and lack of user input has been cited as one of the major reasons for the failure of software development (standish group 1994). various techniques for acquiring this type of information have been discussed in goguen and linde (1993). empirical design early on in the development process potential users should be exposed to simulations and prototypes and asked to carry out tasks using the prototypes. the reactions, comments and performance should be observed and measured in order to identify problems and errors. feedback can then be incorporated into the design relatively easily, and improvements made before a project is committed to a full implementation. iterative design once user testing has been carried out using a simulation or prototype, improvements should be made and the product re-tested on the prototypes before full implementation begins. as much as possible the iterative phase of design should be confined to prototypes. measuring the success of a project ascertaining a measure of the success or failure of a project is not an exact science and depends largely upon the subjective assessment by its users. nevertheless, once a project design has been implemented it is important that the degree of success is measured since this will help to inform future developments. the success or effectiveness of commercial portals is usually measured in terms of user acquisition and retention (clarke and flaherty 2003). there are, however, a number of alternative ways in which success can be measured: achievement of requirements/goals set by project this involves identifying and describing user tasks that will be supported by a portal and using the results to establish a set of system requirements and deliverables prior to undertaking any development. the progress and success of a project can then be measured against this list of targets. usability tests these measure the usability of a portal by asking a number of users to complete a particular task and measuring their response. this is distinct from the iteration process, as described above in that it is carried out on the final product and should not be considered as part of the design process. nevertheless the results of such tests can be used to inform the design of future versions of the software. extent of uptake and use by potential end-user groups although this is the most commonly utilised method of measuring success (usually because it is easy to gather the necessary data) establishing a meaningful measure of success using this method is not straightforward. firstly, the number of potential end-users is clearly variable between portals, depending on the number and size of different user groups to be served. to overcome this it is possible to consider the number of users of a portal as a proportion of the potential number of users. this of course requires that a sensible assessment of the potential end-user base has been established during the design phase. feedback this is a passive and therefore easy to implement mechanism for obtaining information regarding the success or failure of the project. reliance on this mechanism is, however, likely to be less informative than the other three mechanisms. users are only likely to comment on negative aspects of a system although failure to neale et al. – online biodiversity resources – principles for usability 29 receive any feedback is as likely to indicate user dissatisfaction as it is to indicate success. an important point to note here is that the ability to measure and demonstrate the success of a project is highly dependent upon a thorough and appropriate initial design and the documentation thereof. in effect appropriate early design provides the context within which an entire project can be managed. astrogrid: a portal for the astrophysics community – a ‘model’ example of portal development we consider this project to represent a very good model of the way the development of an information portal can be managed. we will therefore use it to set the baseline against which biodiversity projects can be measured. astrogrid is a £10m project funded by uk's particle physics & astronomy research council (pparc) and the european commission. it is aimed at building a working virtual observatory for uk and international astronomers, and is one of three large world-wide virtual observatory projects (the astrogrid consortium 2002). the following sections examine the development of astrogrid in the context of the user-oriented software development principles detailed in the previous section. the development of astrogrid early and continuous focus on tasks and users at the start of this project, prior to any technical development, an extensive consultation exercise was undertaken. contributions to this were made through a number of routes; from consortium members, from research carried out by the project scientists, from consultation of documents detailing the results of requirements surveys from other projects, calls for input from the wider community at academic meetings, online and in journals, as well as from visits to laboratories, all of which are clearly documented in publicly available project resources (the astrogrid consortium 2002). through this exercise science problems that were of interest to the uk astronomical research community were identified. ten of these problems were then selected as priority objectives for the project based on the breadth of end-users showing interest in each issue. each science problem was then formally constructed as a use case, broken down to show a flow of events decomposed into tasks for each process and sub-process. the results of this analysis were made publicly available as a project resource. user input was maintained throughout the project through online testing, development workshops and an online forum. the use cases were used to define the project deliverables both generating measurable targets and maintaining a focus on the user tasks throughout development. in the astrogrid2 project, user involvement has been taken a step further with the development of an innovative process whereby users suggest functions that they would like to see developed. a number of projects are selected on a quarterly basis and then delivered during the astrogrid2 development cycle. the proposer of a new function becomes part of the development workgroup. the user community is therefore actively involved in dictating the direction of the development of the project thus creating a sense of community ownership. empirical measurement and iterative design user testing of the portal was carried out online as new versions and functions became available. potential end-users were actively signed up to conduct the testing and a three month cycle of development and testing was established for new functions. user workshops were held at the end of each iteration cycle. measuring the success of astrogrid a good indication that astrogrid has been perceived as a success is that it has received £4m for a second phase. this perception of success is no doubt attributable to the fact that a set of clearly defined project goals and concomitant functional deliverables was established and documentary evidence of the accomplishment of those goals is readily available. furthermore, as of 2005 the number of active users of astrogrid was measured at 1000. although this is not a large number it is nevertheless far in excess of their prediction (the astrogrid consortium 2002). comparing biodiversity portal projects with astrogrid the astrogrid project arose from a drive by the astrophysics community for resources to utilise and explore rapidly expanding astronomical datasets and the astrogrid project advertises itself as a virtual astronomical observatory. at the broadest level this makes astrogrid and many biodiversity portals directly comparable. indeed it might even be helpful to regard biodiversity portals as virtual biodiversity observatories. neale et al. – online biodiversity resources – principles for usability 30 when comparing biodiversity data portals with astrogrid we have opted to take a broad brush approach. as such, this paper is not intending to provide a direct critique of individual portal developments which we would consider to be unfair. rather it is aiming to draw out the common features and failings of biodiversity portal development in order to highlight where improvements could be made when future biodiversity portal developments are being initiated. given the current existence of large biodiversity data portals, it could of course be considered that putting these arguments forward at this point would be a case of closing the stable door after the horse has bolted. this would, however, be a short-sighted position to take. gbif is still under development and even now the data access portal is only being described as a prototype (gbif 2006a). when moving towards full implementation it would be highly advantageous for them to take a closer look at their end-users and their requirements. there is also a continuous stream of new projects very similar in nature to gbif being proposed and seeking funding. the eu funded edit programme is an example of a newly funded portal style project1. we also know of two other major biodiversity project proposals that are currently seeking funding, one concerned with the development of an archive for dna barcode data (m. watson pers. comm.) and the other to create an on-line encyclopaedia of life (r. hyam pers. comm.) all of which could no doubt benefit from undertaking a user-oriented design analysis. during the course of our research for this paper we have in large part concentrated on studying gbif and biocase. however in order to more fully develop our arguments we have studied the available documentation from a wide range of biodiversity portal projects (algaebase, biocase, biocise, cabi species fungorum, embl reptile database, enhsin, enbi, erms, eti-wbd, eti-wtd, fauna europaea, fishbase, gbif, speciesbank, ildis, linne, natureserve, natuurloket, nbn, species 2000 europa and wadsis). we have also been able to garner considerable information from two recent reports on user requirements analysis in biodiversity portal projects in which the design processes for a wide range of projects were obtained directly from the portal developers (neale et al. 2005, schalk and heitmans 2005). 1 http://www.e-taxonomy.eu/. it was during this process of information gathering that the first obvious contrast with astrogrid was observed. for most of the biodiversity projects the level of publicly available documentation regarding requirements analysis and design strategies was so low as to be almost non-existent despite the fact that many of them had substantial online documentation. in contrast, from the outset the astrogrid project appears to have been driven by well defined, clearly documented and publicly available user needs established early in the course of its development. early and continuous focus on tasks and users the users of the larger biodiversity portals are potentially very diverse. this can make it difficult to accurately predict exactly who will use a service before it is operational and its use can be monitored (kujala and kauppinen 2004). as stated above the astrogrid project overcame this difficulty by identifying the scientific questions to be addressed by the grid and thus the end-user community and the functional requirements become self-defining. no such specific context exists for biodiversity portals and therefore it is interesting to examine how the scope of these projects has been determined. for most projects it appears that a list of potential end-users was generated by brainstorming during preliminary design meetings (neale et al. 2005). often the participants in such meetings were either limited to the site developers or were expanded to include potential data providers. as a result, the lists of potential users generated often appear over-optimistic and represent more of a wish list than a realistic assessment of the potential user base. the brainstorming process of identifying potential users should only be considered a preliminary step towards a more fully fledged requirements analysis. unfortunately it appears that the majority of biodiversity portals have stopped their initial requirements analysis at this point and moved straight onto technical design and implementation. where evidence of requirements analysis has been published, the resulting analyses appear to have been rather unfocussed – providing no mapping between the tasks that the users wish to perform and the data and functions required to support those tasks (calabuig et al. 2001; larsen et al. 2004; robinson 2005). a recent study for the biocase project served to highlight this problem (neale and pullan, 2005). one of the preliminary development tasks neale et al. – online biodiversity resources – principles for usability 31 for biocase project was to establish a list of potential end-users. this was generated by a combination of brainstorming and by consolidating similar lists produced by other related projects (neale et al. 2005). this final list of potential end-users covered 43 diverse user groups which were split into the following five categories: 1. biological systematics and university research 2. private research and industry 3. public services and administration bodies 4. commercial services 5. general public immediately clear from this list was a probable decrease of interest in unit level information in favour of an increased interest in complex synthesised data from category one to category five of these user groups. the biocase portal provides access to unit level data, and is therefore clearly less likely to be of value to users as you move from category one to five. this was tested by attempting to match the data requirements of the potential users from each group with the data available and was achieved by conducting a series of face to face interviews with representatives of the groups. without reference to the biocase portal, users were asked to identify the tasks and data required to support those tasks that they perform as part of their day to day work. each task and its associated data were then ranked by the interviewee in order of importance. a mapping was then made between the data required for a task and the availability of that data in the portal. the results of this study reinforced the initial suspicion that the data accessible through the biocase portal was unlikely to directly satisfy the requirements of the non-taxonomic groups of potential users and that for future development of the biocase portal the interface design should concentrate on supporting the tasks most commonly undertaken by taxonomists and a more realistic set of objectives for the portal was established (neale et al. 2005). the effectiveness of this effort was however compromised by the fact that the requirements analysis was conducted in parallel with the portal development and so by the time the analysis was complete a large proportion of the portal had already been built. empirical measurement and iterative design biodiversity projects appear to fair better in this area than they did with regard to their initial requirements analysis. most of the biodiversity portal projects are iterative in the sense that new versions are released periodically, presumably incorporating elements of user feedback as and when deemed appropriate. however, compared with astrogrid, the feedback mechanisms appear to be somewhat unmanaged. the most frequent mechanism for gathering this feedback was through the rather passive mechanisms of online discussion and comment (e.g. fishbase). in some cases more active means of obtaining feedback were established such as online surveys of invited participants (e.g., gbif 2006b), conferences (e.g., enbi 2003, gbif 2005) and face-to-face user testing (e.g., pendry 2004). measuring the success of biodiversity portal projects in terms of the mechanisms for measuring success discussed earlier there was a lack evidence of either clear requirements goals or usability testing in the documentation of the projects examined. this lack of a structured framework and clear goals makes it very difficult to measure the success or progress of these projects. although some do monitor the number of visitors and associated statistics, meaning that they could obtain an assessment of who is accessing their services and for what. it was, however, noted in an enbi report that many online biodiversity resources do not take advantage of the ability to do this (schalk and heitmans 2005). discussion in this paper we have highlighted that during the initial development of many electronic biodiversity resources there appears to be a lack of adequate research into potential users and the tasks they are likely to want to perform. while there is undoubtedly a huge potential in these resources this lack of understanding represents a risk to their potential productivity, usefulness and perceived success. one possible reason for the lack of appropriate requirements analysis, aside from the fact that such analysis can be somewhat tedious and time consuming to undertake, is that when a project is funded it is taken as read that a definite requirement for the project has been established in the application for funding. if this were to be done properly, rather than simply relying on assertions, a full user requirements analysis would have had to be undertaken by the applicants prior to the application and at their own expense. projects that incorporate elements of user analysis into the application risk being rejected on the basis that neale et al. – online biodiversity resources – principles for usability 32 funding will not be granted to projects that can’t already demonstrate their usefulness, a classic catch 22 situation. we would like to suggest a change in funding attitudes such that it is possible to apply for research money to properly establish the appropriate scope and probable usefulness of potentially interesting projects prior to application for full funding. this would not only reduce the risk of money being wasted on ‘white elephants’ but would provide the framework within which the success of a project could be measured were it to eventually receive full funding. in this vein a great deal can be learnt from the success of the astrogrid project. moreover, because funding bodies do not appear to value requirements analysis as part of the grant assessment process a perhaps unexpected consequence is that once a grant has been awarded there is no external pressure to perform such a study. we would therefore also suggest that funding bodies should start to consider user-oriented requirements analysis as being a fundamental part of their pre and post award assessment processes. this will not only ensure a greater level of attention to design during a project but will give the awarding body a far more objective framework within which to assess the outcomes of a project. when we have found evidence of user requirements analysis this process appears to have been somewhat limited. little attempt has been made either to define the high level tasks to be supported by the resource, or to map the data available to the tasks that can be supported by that data, or even to identify the tasks that will need to be supported by the resource for individual user groups. the design and implementation of the resource cannot therefore be informed by an understanding of the relationships between users, tasks, content and functions/capabilities, even though such an understanding would be likely to increase the probability of developing a truly useful resource. in addition to the finding arguments, it is possible to identify a number of other possible causes for the lack of focus on users. firstly, at least in the case of larger portal projects such a gbif, there appears to be a perception that the portals will themselves generate requirements. this will, however, defer the acquisition of essential requirements data until well down the development path and is therefore a risky strategy to follow. the lack of focus on users may also be related to the fact that the larger portals have such a potentially wide-ranging group of end-users that to look at the needs and tasks of each group is considered too complex. one way to approach this may be to attempt to break down the complexity of the user base by changing the design emphasis of the projects. if portal providers focus on developing flexible infrastructure architecture, the plumbing so to speak, they could then support user-groups in the development of their own specific interfaces, tailored to support the tasks of these groups. the primary output of a project like gbif would then be a set of tools allowing portal interface developers to hook into the gbif infrastructure. this allows the interface development paradigm for the portals to be closer to that for the smaller more specific resources that already exist as stand alone data resources. many of the successful smaller sites are in fact developed by their end-user community (e.g. etiwbd). their success is undoubtedly linked to several factors: 1. the interfaces are developed by the end-user community, the developers therefore have an inherent understanding of the tasks they need to perform, and the interface functions required to support those tasks. 2. such interfaces will only be developed when the end-user community positively identifies a requirement for such an interface onto the resource. 3. because the user community obtains clear ownership of the portal, the development and enhancement cycle of such interfaces is therefore likely to be enduring. in addition to providing a clearer set of design goals, adopting the approach to design suggested here would also have the advantage of clearly separating the design pathway of the portal infrastructure from the design pathway of the interfaces. figure 1 illustrates the process by which a set of use cases can be developed by considering the relationships between users, tasks and the available data. figure 2 shows where the design process illustrated in figure 1 fits into the overall design process that could be applied to a portal and its interfaces. in figure 2 it can be seen that the design goals for the portal and its interfaces are clearly separated. this leads to a different selection of end-users when starting the user requirements analyses, thereby keeping the end-user groups smaller, more tightly defined and therefore less complex. interface developers focus their requirements analysis on the end-user community they wish to serve and the infrastructure developers focus their requirements analysis on the interface developers they intend to support. rather than a single set of use cases for the entire system this will lead to the development neale et al. – online biodiversity resources – principles for usability 33 aims / scope of project available data potential user groups tasks that can or will be supported user tasks use cases system design once potential user groups have been identified the tasks they perform and which could be supported can be identified – this should involve input from actual users the potential users of a portal will be influenced by both the aims of the project and the available data identify which tasks can be supported through the portal deconstruct the individual tasks requirements can be determined from the use cases the scope of the project is in part determined by the data which are available and vice versa. the data which are available will impact on the tasks which can be supported by the portal. figure 1. a process by which use cases can be developed considering users, tasks and available data. of a set of use cases for the infrastructure and a separate set of use cases for each interface to be developed on the infrastructure. having said this, the use cases for the interfaces are likely to be somewhat dependent on those developed for the portal infrastructure. it may at first glance appear that this is indeed the approach that gbif has taken. the project has been broken down into a number of operational categories each with a controlling committee. there are three operational categories that are directly relevant to the accumulation, indexing and access to data – dadi (data access and database interoperability) and ecat (electronic catalogue of names of known organisms), digit (digitisation of natural history collections; gbif 2003). however, examination of the work programmes associated with these efforts reveals little concern for the ultimate endusers and an overriding emphasis on developing tools and mechanisms for the accumulation and indexing of data at the expense of providing tools to access the data. while the data accumulation and indexing tools are undoubtedly essential elements of a portal project, because there is no clear understanding of the ultimate use of the accumulated data there is no guarantee that the data indexing and exchange systems developed will satisfy the user requirements as they emerge. indeed only now, seven years into the gbif development programme are we starting to see evidence that users are beginning to show an interest and looking for services, and consequently it is only at this point that the system can be tested (gbif 2006a). it is, therefore, very likely that, as intimated at the start of the paper, a considerable amount of the infrastructure will have to be redeveloped (“enhanced”) due to lack of timely focus on end-user requirements. moreover, although there are two demonstrator projects designed to show the utility of gbif the primary public output from the project currently appears to be the gbif data access portal. this approach is again problematic since in essence it creates a single point of failure for the system, i.e. regardless of how well the infrastructure has been developed gbif will almost certainly be judged on the success or failure of this interface. altering the design model such that the interfaces and the infrastructure are conceptually separated, means that the success of the portal infrastructure can then be measured independently of the success or failure of any of the interfaces developed on it. responsibility for design and implementation of the interfaces can be delegated to the end-user communities giving all the advantages of neale et al. – online biodiversity resources – principles for usability 34 aims / scope of portal infrastructure use cases use cases infrastructure implementation interface implementation aims / scope of portal interface system design system design . the system for determining users, tasks and use cases shown in figure 1 can be applied at this point. in the portal interface design the actors will mainly be end users, and the tasks high level e.g. identify pest species in garden the requirements for the interface are specified by the use cases and form a measurable set of deliverables for the system design the design of the interface will be in part determined by the design of the portal infrastructure some use cases will have overlapping sub elements the requirements for the interface are specified by the use cases and form a measurable set of deliverables for the system design the system for determining users, tasks and use cases shown in figure 1 can be applied at this point. in portal infrastructure design the actors are likely to be machines ‘talking’ to one another, and the tasks lower level ones for example ‘retrieve data’ figure 2. an illustration of the design process for portal infrastructure and interfaces showing how, although the two processes are interlinked and interdependent, they can also be separated so that end-user groups could generate individual interfaces on top of a single portal infrastructure. community-lead interface development discussed earlier. the only potential argument against this is that it is politically more difficult to sell a project intended only to develop infrastructure rather than high visibility public interfaces. it would appear that within both gbif and biocase there is indeed a move towards this approach. gbif provides a portal development tool based on the software it uses to drive its web site although this software does not provide tools for accessing gbif data. they also provide a uddi registry of web services which may be useful in data portal construction. gbif are also beginning internal discussions on the development of a universal data bus intended to provide just the kind of interface independent infrastructure we have been describing (r. hyam pers. comm.). biocase now provides the biocase and tapir protocols plus the pywrapper data provider, and ‘unit loader’ software2 which allow third parties to develop their own collaborative networks using a common architecture and a couple of prototype demonstration interfaces are available3. however, this approach is only just starting to emerge after 2 http://www.biocase.org/. 3 http://www.biocase.org/products/portals/. five to seven years of development in which, as far as can be judged from the project documentation, the primary focus appears to have been providing a single access point to distributed data. the developments described at the beginning of the paragraph therefore appear to have arisen as useful by-products rather than as the products of primary design goals. they have therefore involved adapting their software to the broader aim of infrastructure provision rather than designing for this from the outset. it can only be concluded that a move towards this truly useful kind of tools development could have been achieved much more quickly and efficiently had the appropriate user-oriented design analyses been undertaken from the beginning. abbreviations biocase – biological collection access service for europe biocise – resource identification for a biological collection information service in europe embl – european molecular biology laboratory enbi – european network for biodiversity information enhsin – european natural history specimen information network neale et al. – online biodiversity resources – principles for usability 35 erms (marbef) – european register of marine species (marine biodiversity and ecosystem functioning) eti-wbd – world biodiversity database eti-wtd – world taxonomist database gbif – global biodiversity information facility ildis – international legume database & information service linne – legacy infrastructure network for natural environments nbn – national biodiversity network wadsis – transnational wadden sea information service acknowledgments this work was funded by the network activity d of the eu synthesys project contract number r113/ct/2003/506117. we thank prof. david mann and dr. colin pendry for valuable comments on the manuscript. literature cited astrogrid consortium, the. 2002. astrogrid phase a report.4 bias, r. g., and d. j. mayhew, eds. 1994. costjustifying usability. harcourt brace & co., boston. calabuig, i., c. dieguez, i. izquierdo, n. scharff, and h. enghoff. 2001. enhsin – european network of biodiversity information – report on results from questionnaire to assess user needs. annual report 2001. zoological museum, university of copenhagen. clarke, i., and t.b. flaherty. 2003. web-based b2b portals. industrial marketing management 32:1523. enbi. 2003. first enbi e-conference: open access for biodiversity information. zoological museum of amsterdam, amsterdam.5 gbif. 2003. global biodiversity information facility strategic plan october 2003. global biodiversity information facility, copenhagen. gbif. 2005. building speciesbanks: how shall we shape the future? global biodiversity information facility, copenhagen. gbif. 2006a. gbif strategic and operational plans 2007-2011: from prototype towards full operation. m. lane, ed. global biodiversity information facility, copenhagen. gbif. 2006b. survey result summary: feedback and requirements for use of gbif data portal. global biodiversity information facility, copenhagen. goguen, j.a., and c. linde. 1993. techniques for requirements elicitation. pp.140-152 in thayer, r.h. and m. dorfman, eds., software requirements engineering (2nd ed.). ieee computer society press, los alamitos, california. 4 http://wiki.astrogrid.org/pub/astrogrid/phaseareport/redbook.pdf. 5 http://www.enbi.info/. gould, j. d., s.j. boies, and c.h. lewis. 1991. making usable, useful, productivity-enhancing computer applications. communications of the acm 34:7585. gould, j.d., and c. lewis. 1985. designing for usability: key principles and what designers think. communications of the acm 28:300-311. hewitt, t.t., and c.t. meadow. 1986. on designing for usability: an application of four key principles. pp. 247-252 in human factors in computing systems conference proceedings. acm press, new york. iso. 1998. 9241-11: ergonomic requirements for office work with visual display terminals (vdts), part 11: guidance on usability, 1st ed., 1998-03-15, international organization for standardization, geneva. jones, p.g., and p.k. thornton. 2003. the potential impacts of climate change on maize production in africa and latin america in 2055. global environmental change 13:51-59. kujala, s., and m. kauppinen. 2004. identifying and selecting users for user-centered design. p.p. 297303 in proceedings of nordic conference on computer-human interaction, tampere, finland, 25-27 october 2004. acm press, new york. kujala, s., m. kauppinen, and s. rekola. 2001. bridging the gap between user needs and user requirements. p.p. 45-50 in avouris, n., and n. fakotakis, eds. advances in human-computer interaction i. proceedings of the panhellenic conference with international participation in human-computer interaction (pc-hci 2001). typorama publications. kujala, s, and m. mäntylä. 2000. is user involvement harmful or useful in the early stages of product development? pp. 285-286 in computer human interaction 2000. the hague, netherlands. kujala, s. 2002. user studies: a practical approach to user involvement for gathering user needs and requirements. acta polytechnica scandinavica, mathmathics and computing series no. 116. finnish academies of technology, espoo. larsen, f.w., i. calabuig, and h. enghoff. 2004. report on non-european user needs for biodiversity data. enbi report. wp13_d13.3_9/ 2004. zoological museum, university of copenhagen, copenhagen. löwgren, l. 1995. perspectives on usability. institutionen för datavetenskap ida technical report lith-ida-r-95-23. university of linköping, sweden. mulligan, r. m., m.w. altom, and d.k. simkin. 1991. user interface design in the trenches: some tips on shooting from the hip. pp. 232-236 in robertson, s.p., g.m. olson, and j. s. olson eds. proceedings of the computer human interaction ‘91 human factors in computing systems conference. acm press, new york. neale, s., m. pullan, and m. watson. 2005. user requirements and interface testing documents of natural history collection and biodiversity online neale et al. – online biodiversity resources – principles for usability 36 portal projects. biocase/ synthesys report. royal botanic garden edinburgh, edinburgh. neale, s., and m. pullan. 2005. user interface requirements analysis for the biocase/synthesys biological information portal. synthesys networking activity d deliverable d 3.4.1. royal botanic garden edinburgh, edinburgh. parmesan, c., and g. yohe. 2003. a globally coherent fingerprint of climate change impacts across natural systems. nature 421:37-42. pendry, c. 2004. final report: testing the biocase user interface at rbge, april to july 2004. royal botanic garden edinburgh, edinburgh. rauterberg, m., and o. strohm. 1992. work organization and software development. annual review of automatic programming 16:121-128. sánchez-cordero, v., and e. martínez-meyer. 2000. museum specimen data predict crop damage by tropical rodents. proceedings of the national. academies of science usa 9:7074-7077. schalk, p., and w. heitmans. 2005. conclusions from an inventory of end-user forums and user feedback of biodiversity projects. enbi report wp12_d12.2b_12/2005. university of amsterdam, amsterdam. standish group. 1994. the chaos report. west yarmouth, ma: the standish group international, inc.6 6 http://www.standishgroup.com/sample_research/chaos_1994_1.php peterson, a. t., and j. shaw. 2003. lutzomyia vectors for cutaneous leishmaniasis in southern brazil: ecological niche models, predicted geographic distributions, and climate change effects. international journal for parasitology 33:919-931. robinson, n. 2005. user aspects breakout group, gbif species bank workshop, amsterdam 2-4 march 2005. global biodiversity information facility, copenhagen. root, t. l., j.t. price, k.r. hall, s.h. schneider, c. rosenzweigk, and j.a. pounds. 2003. fingerprints of global warming on wild animals and plants. nature 421:57-60. roth, c. 2002. top 10 portal pitfalls and how to avoid them. special report on portals, meta group.7 standish group. 2006. the chaos report. west yarmouth, ma: the standish group international, inc. waloszek, g. 2001. portal usability – is there such a thing? sap design guild editions. edition 3: portals. sap, walldorf, germany. 8 williams, p., d. gibbons, c. r. margules, a. rebelo, c. humphries, and r. pressey. 1996. a comparison of richness hotspots, rarity hotspots, and complementary areas for conserving diversity of british birds, conservation biology 10:155-174. zazelenchuk, t. w., and e. boling. 2003. considering user satisfaction in designing web-based portals, educause quarterly 26: 35-40. 7 http://techupdate.zdnet.com/techupdate/stories/main/ 0,14179,2867254,00.html. 8 http://www.sapdesignguild.org/editions/edition3/portal_usab.asp. microsoft word bi2009kakodkatdigim_final.doc biodiversity informatics, 6, 2009, pp. 1-4 darwin core based data streamlining with digimus 2.0 kakodkar a. p., kerkar s. s., varghese n. s., kavlekar d. p. & c.t. achuthankutty bioinformatics centre, national institute of oceanography, council of scientific and industrial research (csir), dona paula, goa 403 004, india abstract. cataloguing biological specimen is an important activity of biological museums world over. software developed especially for this purpose has evolved over time to achieve more accuracy in retrieving data from large and diverse datasets. combining smaller datasets into a larger information system requires uniformity of data based on a single data standard. in the developing world smaller datasets are maintained by individual researchers or small college and university groups. to standardize data from such datasets software needs to be developed, requiring expertise and sufficient funds which are often unavailable. we present a simple open source web based tool developed using php to enable an individual, with little or no knowledge of information systems or databases, to effectively streamline specimen data with data standard darwin core 1.2 ( dwc 1.2). such data can then be shared and easily provided to data aggregators like ocean biogeographic information systems (obis http://www.iobis.org ) and global biodiversity information facility (gbif http://www.gbif.org). this tool can be accessed at http://www.niobioinformatics.in/digimus.php and its source code is freely available at http://www.niobioinformatics.in/digimus_source.php. key words. biological specimen, darwin core, data standards, data streamlining, information exchange. darwin core data standards facilitate robust information exchange between distributed datasets (costello & berghe 2006; sautter et. al., 2007). data retrieval tools such as digir server have been meticulously designed for compatibility with darwin core data standards (hobern, 2002; greene, 2007; digir, 2008). integration of digir server with darwin core enables querying diverse databases with the same darwin core schema, but independent of their location, data content or methods of development (wieczorek, 2007). however, there aren’t any accurate estimates for the number of collections adopting darwin core data standards. a current estimate, by gbif enumerates 1683 collections from 226 providers (gbif, 2008), which may not include darwin core datasets which are not data providers to gbif. current scope of data standards (e.g. gbif); do not consider small datasets which are with individuals and institutions. most of these efforts take place in developed countries. individual researchers from developing countries possess valuable data which is mostly not shared with global datasets (schnase et. al., 2003). this data is unavailable to the global research community correspondence email: anaikkakodkar@nio.org (edwards, 2000). until data standards such as darwin core are implemented by researchers from developing world the information in their possession may prove difficult to access. funding for biological data management in developing countries is much less or entirely nonexistent (thomson, 2005). to solve this problem, a software tool needs to be developed that requires the user to possess little or no knowledge of web programming and database design. in this paper we present a web based open source software tool named “digimus 2.0” which enables individuals to manage biological specimen data from their collection and effectively structure it according to darwin core data standards v 1.2 ( dwc v 1.2). digimus 2.0 development digimus 2.0 presents a web interface allowing a user to enter biological specimen data into a database. the database is structured to comply with darwin core 1.2 ( dwc 1.2) data standards (wouter, 2008). data entered by individual users can be downloaded, transferred to another database, or shared with larger data aggregators (e.g. obis and gbif) that employ a digir server to query darwin core compliant datasets. as digimus 2.0 kakodka – darwin core data streamlining with digimus 2.0 2 figure 1. interface for digimus 2.0. page displayslinks to various data management modules. was developed mainly for marine biological specimens, we have used obis version of darwin core 1.2, to enable better information exchange between an individual dataset and obis. a php 1 script written for digimus 2.0 connects to a mysql 2 database and inserts user entered data. the script arranges the data into the tables designed to comply with dwc 1.2 schema. when the database is queried through the web interface, the php script runs a mysql query and fetches results to the browser window in html. in addition to dwc 1.2 compliant dataset, digimus 2.0 also allows users to enter additional specimen metadata (e.g., specimen description, ecological details and commercial importance), these tables are non dwc 1.2 compliant (see figure 1). html data is presented in the form of tables. a data download link below each table allows the users to download their data to a comma separated values (csv) file. data from this csv file can then be imported into a ms excel worksheet or a local mysql database. digimus 2.0 also proves a good choice for online 1http://www.php.net. 2http://dev.mysql.com. user data backup. a user cannot directly delete any data from the main database server but can request the system administrator to do it instead, if data privacy needs to be protected. custom unique record identifiers can be added by users in the “accession number” field. this is particularly useful in cases where users have predefined unique record identifiers. digimus 2.0 has two web interfaces (i) data management interface (dmi), it consists of a html data submission form, that enables a user to enter text and upload multimedia files. various modules are available to the user to manage various types of information related to a single species (e.g., biogeography, taxonomy, ecology, description, commercial importance etc.). (ii) record display interface (rdi), this is a output from the server in the form of html, it presents data in the web browser which enables users to view or download data entered by them through the data management interface, through the browser window (see figure 2). a minimum set of data fields are required to successfully complete a data record. these are programmatically implemented and are prompted within each module if minimum data fields are not met. these include scientific name (bionomial name, author and year) and basic taxonomy (from kingdom to family). evaluation of digimus 2.0 was carried out using data from seaweed herbarium collection of national institute of oceanography, goa. a total of 729 seaweed herbarium records and data from other sources were manually entered using the web interface of data management interface. a demo web interface using html and css was developed to display records. the demo can be accessed here 3 discussion the objective behind developing digimus 2.0 was to create an easy to use interface for marine biologists to manage their biological specimen data and streamline it with current data standards such as darwin core 1.2 and we found that digimus 2.0 can create a downloadable data file concurrent with dwc 1.2 data standards, which can be uploaded to larger aggregators such as obis (grassle, 2000) and gbif (lane, 2003). digimus 2.0 can prove an important tool for the marine biologists in the 3 http://www.niobioinformatics.in/digimus_demo.php kakodka – darwin core data streamlining with digimus 2.0 3 figure 2. representation of data flow through digimus 2.0. data management interface (dmi) is a html form through which an user enters data. data is processed by a sql query and stored in the mysql database. record display interface (rdi), displays query response in a html table, which is downloadable. a user can share download data with obis as a text file (csv) or through digir server. developing countries, as it will enable them to manage data according to current data standards and without spending more time in developing software tools for data management or learning web programming. as a web based tool digimus 2.0, does not require specific hardware or software other than a personal computer with internet connectivity and a web browser. such minimal requirements make it an ideal tool for individual researchers and small research groups in developing countries with limited data management budgets. web based biological specimen digitization software has always been developed to target a small number of data managers who have been specifically trained to have good knowledge of data management and information system concept. digimus 2.0 does not require the user to undergo any specific training. in fact it has been specifically designed to be user friendly. an undergraduate student or a researcher with basic knowledge of using a web browser can use digimus 2.0 with ease. being in the development phase digimus 2.0 has some weaknesses. on the user end the functioning of digimus 2.0 depends upon the internet connection speed of the client computer. low internet connection speed results in decreased number of managed data records. on the lower internet speeds uploading large image or multimedia files takes more time. at the server side, data storage can prove challenging with an increase in number of users. managing separate databases for individual user at the server side is another concern. these are some challenges which need to be addressed in future versions. digimus 2.0 exists as an open source tool that will ease marine biological data sharing. it prepares unorganized data into a standard format. being an open source software and licensed under gpl, its source code can be modified for specific needs of a particular research group. hopefully, digimus 2.0 will encourage marine biologists from developing countries to be a part of the larger worldwide data management culture and it will also motivate them to share biological specimen data. online availability of data will increase its usability at wider levels which will be in the interest of the global biological community. conclusion streamlining biological specimen data, with current data standards increases the potential of data in terms of information exchange from distributed datasets to global initiatives such as obis and gbif. digimus 2.0 is an effort to present a web based open source software tool for compilation of data to allow compatibility with darwin core data standards. we hope that this kakodka – darwin core data streamlining with digimus 2.0 4 tool shall be a handy utility to individual researchers and small research groups from developing countries for managing their biological specimen data. acknowledgements the authors thank department of biotechnology (govt. of india), new delhi for financial support through the btisnet programme. this is contribution no4458 of nio. references chapman, a. d. 2005. principles of data quality. global biodiversity information facility. 1-58. costello, m. j. and e. v. berghe. 2006. ‘ocean biodiversity informatics’: a new era in marinebiology research and management. marine ecological progress series. 316: 203-214. digir, 2008. distributed generic information retrieval. accessed. june 2, 2008. 4 edwards, j. l., lane, m. a., and e. s. nielsen. 2000. interoperability of biodiversity databases: biodiversity information on every desktop. science. 289: 2312-2314. gbif. 2008. global biodiversity information facility. accessed. june 19, 2008. 5 grassle, j. f. 2000 the ocean biogeographic information system (obis):an on-line, worldwide atlas foraccessing, modeling andmapping marine biological datain a multidimensionalgeographic context. oceanography. 13: 5-7. 4 http://digir.net/. 5 http://www.gbif.org/. greene, s. l., minoura, t. steiner, j. j. and g. pentacost. 2007. webgrms: prototype software for web-based mapping of biological collections. biodiversity conservation 16: 2611–2625 hobern, d. 2002. integrating biodiversity data standards and interoperability . accessed. march 4, 2008. 6 lane, m. a. 2003 the global biodiversity information facility. bulletin of the american society for information science and technology. oct/nov: 22-24. wieczorek, j. 2007. darwin core wiki site. accessed. march 3, 2008. 7 wouter, a. 2008. darwin core versions. accessed. march 3, 2008. 8 schnase, j. l., cushing, j., frame, m., frondorf, a., landis, e., maier, d., and a. silberschatz. 2003. information technology challenges of biodiversity and ecosystem informatics. information systems 28: 339-345. sautter, g., bohm, k. and d. agosti. 2007. a quantitative comparison of xml schemas for taxonomic publications. biodiversity informatics. 4: 1-13. thomson, k. s. 2005. natural history museum collections in the 21st century. accessed march 4, 2008. 9 6 http://www.cria.org.br/eventos/tdbi/bis/presentations/ bis_dhobern.ppt. 7 http://wiki.tdwg.org/twiki/bin/view/darwincore/webhome. 8 http://wiki.tdwg.org/twiki/bin/view/darwincore/ darwincoreversions. 9 http://www.actionbioscience.org/evolution/thomson.html. microsoft word 4853-9203-1-ed_accepted_js.docx biodiversity informatics, 10, 2015, 35-44 omws: a web service interface for ecological niche modelling renato de giovanni1, erik torres2, rafael b. amaral3, ignacio blanquer2, vinod rebello3, vanderlei p. canhos1 1cria centro de referência em informação ambiental, av. dr. romeu tórtima, 388, campinas, sp, brazil; 2institute of instrumentation for molecular imaging (i3m), universitat politècnica de valència, camino de vera s/n, valencia, spain; 3instituto de computação, universidade federal fluminense, niterói, rj, brazil abstract.—ecological niche modelling (enm) experiments often involve a high number of tasks to be performed. such tasks may consume a significant amount of computing resources and take a long time to complete, especially when using personal computers. omws is a web service interface that allows more powerful computing back-ends to be remotely exploited by other applications to carry out enm tasks. its latest version includes a new operation that can be used to specify complex workflows in a single request, adding the possibility of using workflow management systems on parallel computing back-end. in this paper we describe the omws protocol and compare its most recent version with the previous one by running the same enm experiment using two functionally equivalent clients, each designed for one of the omws interface versions. different back-end configurations were used to investigate how the performance scales for each protocol version when more processing power is made available. results show that the new version outperforms (by a factor of two) the previous one when more computing resources are used. key words.—workflow, high-throughput computing, openmodeller studies involving ecological niche modelling (enm) sometimes require hundreds (see segurado & araújo 2004, marmion et al. 2009, feeley & silman 2010, lorena et al. 2011), thousands (see farber & kadmon 2003, elith et al. 2006, wisz et al. 2008) or even millions (see diniz-filho et al. 2009) of models and related procedures to be carried out. the reason is that models can be created for multiple species using different algorithms with different sets of parameter values. there can also be many replicates for model tests, as well as many model projections based on different environmental scenarios. depending on the experiment, performing such tasks on personal computers can be a troublesome experience, if not prohibitive. even relatively simple experiments can take significant time to complete or, in some cases, exhaust machine resources when computingintensive algorithms and high-resolution environmental layers are used. for this reason, enm is a typical field where software applications can greatly benefit from parallelization techniques and high-throughput computing by outsourcing the execution of tasks to more a powerful pool of computing resources. over the internet, different strategies can be used to remotely exploit such computational infrastructures (kai et al. 2011). in all scenarios, an entry point, which can be an application directly accessible to end users or an intermediate web service, is used to communicate with the larger system. web services are web-based applications that support dynamic interactions with other software applications using open standards that include data formats like extensible markup language (xml) (world wide web consortium, 2006) to transmit data over a network via internetbased protocols. web services provide a standard means of interoperating between different software applications running on a variety of platforms and/or frameworks (world wide web consortium, 2004a). for this reason, web services currently permeate web development, as can be seen by the fact that most major web sites provide some mechanism for other programs to access their data or use their functionality. this allows all kinds of townpeterson typewritten text 35 biodiversity informatics, 10, 2015, 35-44 user interfaces and other applications to be developed on top of standard application programming interfaces (apis) created for remote calls, making web services important components of complex cyberinfrastructures (see stein 2008 and amaral et al. 2014 for examples). moreover, service-based applications can free users from the burden of configuring more specific computing resources. over the years, many tools were developed for enm, mostly as desktop applications. to our knowledge, the first web application for enm was released in 1994 (boston & stockwell 1995), while the first web service for enm appeared ten years later as part of openmodeller (muñoz et al. 2011), and it was recently named the openmodeller web service (omws). the first omws prototype was created for the biodiversity world project (pahwa et al. 2006) so that scientific workflows for enm could be built using the triana workflow management system (taylor et al. 2003). this prototype was improved and soon became officially part of the openmodeller toolbox. more recently, other initiatives such as the lifemapper project1 and ehabitat (skøien et al. 2013) also created web services for enm. although the first version of omws managed to cover the most important enm tasks by means of individual operations, any complex experiment would require a client program to send a large number of requests and manage any dependencies between them. this could lead to inefficient bandwidth use, preventing back-end implementtations to fully exploit more sophisticated tools and parallelization techniques. in this paper we present the latest version of omws that includes new operations and additional changes, allowing complex enm experiments to be fully specified in a single request without losing the possibility of using the previous individual calls that remained in the protocol. this paper presents an overview of the protocol and its general design, as more specific details of all operations and the corresponding input/output parameters can be found in the official documentation2. the capabilities of omws are then demonstrated with a typical enm experiment implemented for both protocol versions. the experiment is performed 1 http://lifemapper.org. 2 http://openmodeller.sf.net/web_service_2.html. 2 http://openmodeller.sf.net/web_service_2.html. against the same server hosting both versions of the service. different back-end configurations were used to measure speedup and saturation point (performance limit) when more processing power was made available. omws scope and general design omws was designed to cover essential enm tasks, such as model creation, testing and projection, without including other pre or post processing operations like data cleaning, species occurrence data retrieval from other sources or raster aggregation. although frequently used together with enm procedures, such additional tasks would make the protocol significantly more complex and would produce overlaps with other more specific protocols. due to the potentially long duration of the enm tasks, omws was deliberately designed to be a processing protocol with asynchronous operations, generating outputs that are temporarily stored on the server side to be retrieved later by clients. there are no operations for explicitly storing or manipulating objects on the server, such as occurrence points, algorithms, models or environmental layers. omws should therefore only be seen as a remote niche modelling engine service, not a repository. model repositories, such as the eubrazilopenbio niche modelling application (amaral et al. 2014) or biogeo3, can use omws behind the scenes, but they need to provide additional functionality related to authentication, storage and search capabilities not covered by omws. technically, omws is currently built on top of the simple object access protocol (soap) (world wide web consortium, 2007): a generic messaging framework created to facilitate communication between applications running on different platforms with different technologies. by using soap, the whole protocol is programmaticcally defined in a web service definition language (wsdl) (world wide web consortium, 2001) document specifying all operations, inputs and outputs. through wsdl, existing soap frameworks for different programming languages can automatically generate proxy code to interact with the service. for maximum interoperability and better performance, omws follows the soap document/literal style, which means that 3 http://biogeo.inct.florabrasil.net. townpeterson typewritten text 36 biodiversity informatics, 10, 2015, 35-44 messages are exchanged between client and server as plain xml documents fully encoded according to xml schema (world wide web consortium, 2004b) serialization rules. since omws was originally created as a web service for the openmodeller toolbox, its main xml definitions are all based on the openmodeller xml schema4. although this clearly facilitates using openmodeller tools in the background, by no means it prevents other niche modelling tools to be used. new server implementations could be developed, or the existing standard server implementation could use a plugin approach, translating inputs and outputs of different enm tools to the expected data structures. main data types like openmodeller, omws follows the enm correlative approach (soberón & peterson 2005) by relating species occurrence points with spatially explicit environmental variables so that the corresponding environmental data can be used by an algorithm to generate a niche model. therefore, environmental layers, occurrence points (presence, absence or background), algorithms, models and model projections are the fundamental data types used by the protocol. since environmental layers can frequently be very large raster files, they are referenced in the protocol by an identifier. there are no restrictions or assumptions for such identifiers other than being unique and resolvable by the service. they can refer to files stored on the server side or even point to remote resources. environmental scenarios are specified as a sequence of environmental layers followed by an optional mask also referenced by an identifier. with this approach, rasters that are frequently used can be previously stored on the server and advertised through the getlayers operation. this way, clients can browse the available options and use appropriate identifiers for each selected layer when building requests. remote raster sources can also be used, either as files directly available through http or ftp, or as more formal raster repositories exposed through services like the web coverage service (wcs)5. any of these options enables different mechanisms to be set up so that users can provide their own 4 http://openmodeller.cria.org.br/xml/2.0/openmodeller.xsd. 5 http://www.opengeospatial.org/standards/wcs. layers if necessary (e.g., by granting upload access to a specific directory on the server, or by configuring the server to access a remote resource where users have full control over the available rasters). however, for flexibility and simplicity, such mechanisms are not covered by the protocol, which also does not put any constraints on the content of layer identifiers. this approach makes the protocol more generic, but at the same time requires any additional mechanisms, capabilities or restrictions to be documented and communicated by service providers. occurrence points are often used in low numbers in enm, although there can be situations where thousands of points are available for the species. additionally, the same set of occurrence points is seldom reused by other enm experiments. for this reason, occurrence points are always completely included in requests. each point contains two mandatory attributes for the coordinates and an optional attribute for the corresponding environmental values, in which case the service is relieved of the task of reading environmental data from the corresponding layers. each point in a request must have its coordinates expressed in the same spatial reference system specified in well-known text (herring 2011) for the whole set of points. omws advertises available algorithms through the getalgorithms operation. there is no fixed or standardised set of algorithms, so each service is free to decide which algorithms can be used. algorithm metadata includes name, description, bibliography, authors, developers and parameter metadata (name, data type, domain and description), besides algorithm and parameter identifiers that are used by the other operations when specifying algorithm and parameter values. each different algorithm in enm produces a completely different kind of model. for instance, the result of bioclim (nix 1986) is a series of envelopes comprised by minimum, maximum, mean and standard deviation values for each variable, while random forests (breiman 2001) produces a set of decision trees, and artificial neural networks (tarassenko 1998) produces a system of interconnected neurons with activation functions and weights for each interconnection. omws deals with such diversity by allowing each algorithm to have its own particular xml representation for models, without imposing any townpeterson typewritten text 37 biodiversity informatics, 10, 2015, 35-44 kind of validation. regardless its representation, models can be reused in subsequent calls for testing or projection purposes. model projections are rasters produced in a specific format (e.g., geotiff) based on a specific template raster that indicates the resolution and spatial reference system to be used. each projection is based on a given environmental scenario. projections can be retrieved from a url obtained by calling the getlayerasurl operation when the procedure is finished. figure 1: paired individual asynchronous operations with their main inputs and outputs. each asynchronous operation generates a ticket that is used as input by its counterpart operation to retrieve results later. note: the first column shows only the main types of input for each operation (additional parameters are available). operations omws includes different kinds of operations, starting with a simple ping operation that can be used to monitor service status. two other operations, getalgorithms and getlayers are used to advertise algorithms and environmental layers available, respectively. the remaining operations can be divided into three groups. the first one is a set of individual asynchronous operations related to the main enm tasks (figure 1): createmodel, testmodel, projectmodel and, in the most recent version of the protocol, samplepoints and evaluatemodel. all these operations return a ticket that can be used later to call the corresponding operation for retrieving results: getmodel, gettestresult, getlayerasurl, getsamplingresult and getmodelevaluation. for model projections, an additional operation called getprojectionmetadata can be used to retrieve more information about a projection, such as the number of cells predicted present for a given threshold. despite the similar names, model testing and model evaluation are different operations in omws. the former is used for typical threshold-dependent (confusion matrix) or threshold-independent (roc curve) calculations, while the later is used to calculate raw model values for each given occurrence point in a given environmental scenario. the second group of operations was included in the latest version of the protocol, allowing complex experiments to be specified in a single call and then processed in an optimized way on the server. a runexperiment request may contain any number of enm jobs with or without dependencies between them (figure 2). each job contains its own set of parameters where each parameter either points to a fixed value specified in the first section of the request or to the output of another job. frequent situations such as generating models for multiple species using multiple algorithms and then testing or projecting results into different environmental scenarios can be expressed with this new kind of request. the runexperiment operation is also asynchronous, returning a set of individual tickets for each job. the corresponding operation townpeterson typewritten text 38 biodiversity informatics, 10, 2015, 35-44 getresults can be used to fetch sets of results given one or more tickets. finally, the last group of operations is used for job management after any asynchronous call: getprogress returns the status of one or more jobs, getlog returns the job log and cancel can be used to abort one or more jobs. figure 2: types of jobs and their possible dependencies in a runexperiment call. enm experiment to illustrate how the service can be used in a real world situation, a typical enm experiment involving multiple steps was created. the experiment was tested against two services hosted on the same server, each one compatible with one of the protocol versions (hereafter referred to as omws1 and omws2) for performance comparison. additionally, different back-end configuretions were used, each time adding more processing power on the server side. two equivalent client programs were developed in python, one for omws1, where the client has to be responsible for managing the whole workflow sending individual requests for each task, and the other for omws2 with all tasks specified in the new runexperiment operation where the server is responsible for managing the whole workflow. client programs were executed at the internet data center from the brazilian national research and educational network, while the server was located at the universitat politècnica de valència in spain. average connection speed between client and server was measured as 9.7mbits/s. on the server side, tests started with a single machine with 16gb of ram and 4 cores running the service initially configured to process a maximum of 3 parallel jobs. next, the use of htcondor (thain et al. 2005) was enabled on the server side, with the master node running on the same machine as the service. working nodes had the same computing resources (16gb of ram and 4 cores) and were gradually added to the pool until reaching a maximum of 8 nodes (128 gb of ram and 32 cores in total). the omws2 server implementation for htcondor used dagman (couvares et al. 2007) to handle complex experiment requests. the enm experiment consisted of generating individual models for several species using different algorithms. there was no specific concern about comparing the relative performance of each algorithm or even assessing model quality for each species for any particular use, although these would be obvious follow ups in a real use case. our sole interest was to demonstrate the service with a typical enm experiment, compare the two protocol versions and show how different townpeterson typewritten text 39 biodiversity informatics, 10, 2015, 35-44 back-ends can be used and how they influence the overall processing time and computing resource usage efficiency. five arbitrary species of passifloraceae from the brazilian flora having a minimum set of twenty occurrence points were selected. all points were downloaded from biogeo, where they were previously filtered and cleaned. also five algorithms were used to generate models: enfa (hirzel et al. 2002), garp best subsets (anderson et al. 2003), mahalanobis distance (farber & kadmon 2003), maxent (phillips et al. 2006) and one-class support vector machines (schölkopf et al. 2001). since some of these algorithms rely on background or pseudo-absence points and it is known that the area from where such points are sampled can influence model results (barve et al. 2011), individual masks were created for each species. the idea is that each mask approximates the area that has been historically accessible to the species, ensuring that background or pseudo-absence points are only sampled from environments where the species had the opportunity to colonize. our approximation was done by buffering each set of presence points by 500km and then merging the circles into a single polygon that was finally transformed into a raster. all masks were uploaded to a server where they became accessible to the service. each mask was used to sample 10k geographically unique background points for algorithms that required them, and also to delimit model projections for each species. for simplicity, and since all species are plants with similar requirements, the same set of high-resolution environmental layers currently used in biogeo was used for all species (seven bioclimatic variables and altitude). all layers, including masks, had the same resolution of 30 arc-seconds and were locally available on the server (masks were previously cached by running a preparatory experiment just to force mask download). to match the environmental data precision and avoid redundancies, data cleaning filters in biogeo selected points with a maximum location uncertainty of 500m and removed duplicate points for the same pixel. the actual experiment executed against the service included running extrinsic tests and generating final models for each pair speciesalgorithm. extrinsic tests used 5-fold crossvalidation, averaging the partial auc (peterson et al. 2008) for a maximum omission of 20%. final models were created using all points and were followed by an internal test using the same measurement of the extrinsic test and by a native projection. therefore, the whole experiment contained 125 model creations followed by 125 model tests for the extrinsic tests (5 folds * 5 algorithms * 5 species), and 25 model creations (5 algorithms * 5 species) followed by 25 internal model tests and 25 model projections for the final models, totalizing 325 steps. the experiment was repeated 3 times for each protocol and back-end configuration, using the average as the final measurement. the service code is open source and part of the openmodeller toolbox (we used revision #6045 from the openmodeller repository6). both clients used on the tests and all input data (masks, points and layer references) are publicly available7. experiment results the two protocols performed similarly from the initial single-machine configuration until an htcondor set up with 3 working nodes (12 cores), from where omws2 started to perform increasingly better than omws1 (figure 3). the initial duration for the experiment was 2h44min for both protocols. omws1 reached saturation point with 4 working nodes (duration time of 40min) with a 4.0 speedup compared with the single-machine configuration, which means that adding more computing resources to the backend did not improve efficiency. saturation point for omws2 could not be detected, as its performance continued to improve until the server infrastructure was saturated, reaching a speedup of 8.6 (19min) with 8 working nodes (32 cores). discussion and conclusions regardless the omws protocol version, when a service implementation is capable of exploiting more powerful computing resources there can be significant performance improvement in enm experiments, as clearly demonstrated by the results. omws2 performed better than omws1 when more computing resources became available. this can probably be explained by the fact that omws1 clients need to manage workflows on 6 http://sourceforge.net/p/openmodeller/svn/head/tree/trunk/ openmodeller/. 7 http://dx.doi.org/10.6084/m9.figshare.1301521. townpeterson typewritten text 40 biodiversity informatics, 10, 2015, 35-44 complex experiments without any clue about or control over server resources, while omws2 clients can completely delegate workflow management to the service, where more specialized tools can make use of additional information to optimize resources usage. additionally, by being able to specify complex experiments with a single request, omws2 requires fewer interactions between client and server, also simplifying client code. the task of developing new server software, however, gets more challenging with omws2 to handle workflows, although this also opens the possibility of using existing workflow management tools, such as htcondor dagman in our case. another example is a new omws server implementation under development using comp super scalar with cloud resources (lezzi et al. 2013), which was used to test a prototype protocol that was later improved and became omws2. figure 3: average performance after three repetitions for the different back-end configurations using both versions of omws. standard deviation was low for the graph scale (55s in average) so it is not being represented here. the example tested here also shows that a real world enm experiment likely requires additional tasks not covered by omws, such as preprocessing or post-processing data. this does not mean that such tasks can only be performed as unconnected individual steps, since there are many tools that can be used to integrate and orchestrate tasks performed by different services or software, such as kepler (altintas et al. 2004) and taverna (wolstencroft et al. 2013). omws was created before some of the existing geospatial standards, in particular those defined by the open geospatial consortium8. since omws is mostly based on geospatial data and operations, basic data types such as points could now be expressed according to ogc geography markup 8 http://opengeospatial.org. language (gml)9, or even according to the darwincore biodiversity data standard (wieczorek et al. 2012). in fact, the whole omws protocol could be encapsulated as an ogc web processing service (wps)10. however, in the particular situation of omws, the benefits of adhering to such standards are still unclear or premature in terms of improving interoperability with other software. similar enm initiatives already started to explore the use of wps with interesting results (cavner et al. 2011, skøien et al. 2013), although the incipient set of specific software libraries to interact with wps services and the lack of standard strategies to facilitate wps service chaining are still issues to be addressed. nonetheless, future 9 http://opengeospatial.org/standards/gml. 10 http://opengeospatial.org/standards/wps. townpeterson typewritten text 41 biodiversity informatics, 10, 2015, 35-44 versions of omws could be adjusted or even wrapped to become compatible with other standards. another possible future improvement for omws given the asynchronous nature of most of its operations is to become compatible with the websockets protocol11, which provides bidirectional, full-duplex tcp connections. this could reduce network traffic currently associated with omws getprogress calls. the number of enm applications interested in outsourcing most of the processing tasks to specialized web services is increasing over time. besides other recently emerged protocol initiatives for enm, the use of omws itself also increased over the years. since its first prototype version used by the biodiversity world project, omws was used for many years by the global biodiversity information facility12 data portal, and is still being used by openmodeller desktop users. more recently, biogeo, through the brazilian virtual herbarium, is using a separate omws server to process all enm requests. the eubrazilopenbio project created a web interface for enm where users can create complex experiments involving multiple species, algorithms and environmental scenarios to be processed by an omws2 service through the new runexperiment operation (amaral et al. 2014). all enm workflows created as part of the biovel project13 also interact with an omws2 service. by offering a standard interface for the most important enm tasks with an open source server implementation that can be deployed on more powerful computational infrastructures, omws can be a relevant tool to address some of the challenges of enm research. acknowledgments the latest version of omws contains improvements coming from different sets of requirements originated from two projects that funded their corresponding implementation: eubrazilopenbio14, with grants from the european commission and the national council for scientific and technological development of brazil (cnpq) of the brazilian ministry of science and technology (mct), and biovel, with grants 11 http://www.websocket.org. 12 http://gbif.org. 13 http://biovel.eu. 14 http://eubrazilopenbio.eu from the european commission. server infrastructure was operated through a provisioning system developed in the frame of the spanish project cluviem (tin2013-44390-r) funded by the "ministerio de economía y competitividad". literature cited altintas, i., c. berkley, e. jaeger, m. jones, b. ludäscher, and s. mock. 2004. kepler: an extensible system for design and execution of scientific workflows, in: proc 16th international conference on scientific and statistical database management, pp.423-424. amaral, r., r.m. badia, i. blanquer, r. braga-neto, l. candela, d. castelli, c. flann, r. giovanni, w.a. gray, a. jones, d. lezzi, p. pagano, v.p. canhos, f. quevedo, r. rafanell, v. rebello, m.s. sousabaena, and e. torres. 2014. supporting biodiversity studies with the eubrazilopenbio hybrid data infrastructure. concurr comp-pract e doi:10.1002/cpe.3238. anderson r.p., d. lew, and a.t. peterson. 2003. evaluating predictive models of species’ distributions: criteria for selecting optimal models. ecol model 162:211-232. barve, n., v. barve, a. jimenez-valverde, a. liranoriega, s.p. maher, a.t. peterson, j. soberón, and f. villalobos. 2011. the crucial role of the accessible area in ecological niche modeling and species distribution modeling. ecol model 222:1810-1819. boston, a.n., and d.r.b. stockwell. 1995. interactive species distribution reporting, mapping and modelling using the world wide web. comput networks isdn 28(1-2):231-238. breiman, l. 2001. random forests. mach learn 45(1):5-32. cavner, j.a., a.m. stewart, c.j. grady, and j.h. beach. 2011. an innovative web processing services based gis architecture for global biogeographic analyses of species distributions. in: foss4g 2011 proceedings, 10:15-25. couvares, p., t. kosar, a. roy, j. weber, and k. wenger. 2007. workflow in condor. workflows for e-science (eds: i. taylor, e. deelman, d. gannon, m. shields). springer press. isbn: 184628-519-4. diniz-filho, j.a.f., l.m. bini, t.f. rangel, r.d. loyola, c. hof, d. nogués-bravo, and m.b. araújo. 2009. partitioning and mapping uncertainties in ensembles of forecasts of species turnover under climate change. ecography 32:897-906. elith, j., c.h. graham, r.p. anderson, m. dudik, s. ferrier, a. guisan, r.j. hijmans, f. huettmann, townpeterson typewritten text 42 biodiversity informatics, 10, 2015, 35-44 j.r. leathwick, a. lehmann, j. li, l.g. lohmann, b.a. loiselle, g. manion, g. moritz, m. nakamura, y. nakazawa, j.mcc. overton, a.t. peterson, s.j. phillips, k. richardson, r. scachettipereira, r.e. schapire, j. soberón, s. williams, m.s. wisz, and n.e. zimmermann. 2006. novel methods improve prediction of species’ distributions from occurrence data. ecography 29:129-151. farber, o., and r. kadmon. 2003. assessment of alternative approaches for bioclimatic modeling with special emphasis on the mahalanobis distance. ecol model 160:115-130. feeley, k.j., and m.r. silman. 2010. modelling the responses of andean and amazonian plant species to climate change: the effects of georeferencing errors and the importance of data filtering. j biogeogr 37:733-740. herring, j.r., ed. opengis implementation standard for geographic information simple feature access part 1: common architecture, version 1.2.1, section 9. accessed september 2, 2014. http://portal.opengeospatial.org/files/?artifact_id=2 5355. hirzel, a.h., j. hausser, d. chessel, and n. perrin. 2002. ecological-niche factor analysis: how to compute habitat-suitability maps without absence data? ecology 83(7):2027-2036. kai, h., j. dongarra, and g.c. fox. 2011. distributed and cloud computing: from parallel processing to the internet of things. morgan kaufmann publishers inc., san francisco, ca, usa. lezzi, d., r. rafanell, e. torres, r. giovanni, i. blanquer, and r.m. badia. 2013. programming ecological niche modeling workflows in the cloud, in: 27th international conference on advanced information networking and applications workshops (waina), barcelona. pp. 1223-1228. lorena, a.c., l.f.o. jacintho, m.f. siqueira, r. giovanni, l.g. lohmann, a.c.p.l.f. carvalho, and m. yamamoto. 2011. comparing machine learning classifiers in potential distribution modelling. expert syst appl 38(5):5268-5275. nix, h.a. 1986. a biogeographic analysis of australian elapid snakes. atlas of australian elapid snakes (ed. by r. longmore), pp. 4-15. australian flora and fauna series 7, australian government publishing service, canberra. marmion, m., m. parviainen, m. luoto, r.k. heikkinen, and w. thuiller. 2009. evaluation of consensus methods in predictive species distribution modelling. divers distrib 15(1):59-69. muñoz, m.e.s., r. giovanni, m.f. siqueira, t. sutton, p. brewer, r.s. pereira, d.a.l. canhos, and v.p. canhos. 2011. openmodeller: a generic approach to species’ potential distribution modelling. geoinformatica 15:111-135. pahwa, j.s., p. brewer, t. sutton, c. yesson, m. burgess, x. xu, a.c. jones, r.j. white, w.a. gray, n.j. fiddian, f.a. bisby, a. culham, n. caithness, m. scoble, p. williams, and s. bhagwat. 2006. biodiversity world: a problem-solving environment for analysing biodiversity patterns, in: proc. 6th ieee international symposium on cluster computing and the grid (ccgrid 2006), singapore. peterson, a.t., m. papeş, and j. soberón. 2008. rethinking receiver operating characteristic analysis applications in ecological niche modeling. ecol model 213(1):63-72. phillips, s.j., r.p. anderson, and r.e. schapire. 2006. maximum entropy modelling of species geographic distributions. ecol model 190:231-259. schölkopf, b., j. platt, j. shawe-taylor, a.j. smola, and r.c. williamson. 2001. estimating the support of a high-dimensional distribution. neural comput 13:1443-1471. segurado, p. and m.b. araújo. 2004. an evaluation of methods for modelling species distributions. j biogeogr 31:1555-1568. skøien, j.o., m. schulz, g. dubois, i. fisher, m. balman, i. may, and é.ó. tuama. 2013. a model web approach to modelling climate change in biomes of important bird areas. ecol inform 14:38-43. soberón, j. and a.t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodiversity informatics 2:1-10. stein, l.d. 2008. towards a cyberinfrastructure for the biological sciences: progress, visions and challenges. nat rev genet 9:678-688. tarassenko, l. 1998. a guide to neural computing applications. arnold, london, 139 pp. taylor, i., m. shields, i. wang, and o. rana. 2003. triana applications within grid computing and peer to peer environments. j grid comput 1(2):199-217. thain, d., t. tannenbaum and m. livny. 2005. distributed computing in practice: the condor experience. concurr comp-pract e 17(2-4):323356. wieczorek, j., d. bloom, r. guralnick, s. blum, m. döring, r. giovanni, t. robertson, and d. vieglais. 2012. darwin core: an evolving community-developed biodiversity data standard. plos one, 7:e29715. wisz, m.s., r.j. hijmans, j. li, a.t. peterson, c.h. graham, and a. guisan. 2008. effects of sample size on the performance of species distribution models. divers distrib 14:763-773. townpeterson typewritten text 43 biodiversity informatics, 10, 2015, 35-44 wolstencroft, k., r. haines, d. fellows, a. williams, d. withers, s. owen, s. soiland-reyes, i. dunlop, a. nenadic, p. fisher, j. bhagat, k. belhajjame, f. bacall, a. hardisty, a.n. de la hidalga, m.p.b. vargas, s. sufi, and c. goble. 2013. the taverna workflow suite: designing and executing workflows of web services on the desktop, web or in the cloud. nucleic acids res 41(1):557-561. world wide web consortium, 2001. web services description language (wsdl) 1.1. w3c note 15 march 2001. http://www.w3.org/tr/wsdl. world wide web consortium, 2004a. web services architecture, w3c working group note 11 february 2004. http://www.w3.org/tr/ws-arch/. world wide web consortium, 2004b. xml schema part 0: primer second edition. w3c recommendation 28 october 2004. http://www.w3.org/tr/xmlschema-0/. world wide web consortium, 2006. extensible markup language (xml) 1.1 (second edition). w3c recommendation 16 august 2006. http://www.w3.org/tr/2006/rec-xml1120060816/. world wide web consortium, 2007. soap version 1.2 part 0: primer (second edition). w3c recommendation 27 april 2007. http://www.w3.org/tr/2007/rec-soap12-part020070427/. townpeterson typewritten text 44 townpeterson typewritten text titel biodiversity informatics, 7, 2010, pp. 1-16 1 correspondence e-mail: viktor.nilsson@emg.umu.se using taxonomic revision data to estimate the global species richness and characteristics of undescribed species of diving beetles (coleoptera: dytiscidae) viktor nilsson-örtman and anders n. nilsson dept. of ecology and environmental science, umeå university, se-90187, umeå, sweden abstract.— many methods used for estimating species richness are either difficult to use on poorly known taxa or require input data that are laborious and expensive to collect. in this paper we apply a method which takes advantage of the carefully conducted tests of how the described diversity compares to real species richness that are inherent in taxonomic revisions. we analyze the quantitative outcome from such revisions with respect to body size, zoogeographical region and phylogenetic relationship. the best fitting model is used to predict the diversity of unrevised groups if these would have been subject to as rigorous species level hypothesis-testing as the revised groups. the sensitivity of the predictive model to single observations is estimated by bootstrapping over resampled subsets of the original data. the dytiscidae is with its 4080 described species (end of may 2009) the most diverse group of aquatic beetles and have a world-wide distribution. extensive taxonomic work has been carried out on the family but still the number of described species increases exponentially in most zoogeographical regions making many commonly used methods of estimation difficult to apply. we provide independent species richness estimates of subsamples for which species richness estimates can be reached through extrapolation and compare these to the species richness estimates obtained through the method using revision data. we estimate there to be 5405 species of dytiscids, a 1.32-fold increase over the present number of described species. the undescribed diversity is likely to be biased towards species with small body size from tropical regions outside of africa. key words. — biodiversity, taxonomic bias, estimation, species richness, species description, dytiscidae knowledge on the magnitude of global species richness or the relative species richness of a certain taxa in different parts of the world are essential for understanding the impact of human activities on the world´s biodiversity (purvis and hector 2000), the factors responsible for them (gaston 2000) and assessing the role of biodiversity in maintaining ecosystem functioning (hooper et al. 2005). most researchers would agree that the 1.9 million species described today (chapman 2009) constitute a minor fraction of all the species that are out there (gaston 1991). with the current levels of taxonomic study, certain hyperdiverse taxa might require several hundred, if not thousands of years before they are completely described (gaston and may 1992). thus, developing and employing efficient and imaginative analytical methods that makes the best use of the information available is essential for efficiently planning taxonomic and conservation efforts. extrapolating from what we know in this paper we will focus on the concept of species richness (gaston 1996). estimating species richness at relevant geographic levels has been tackled in three major ways: 1) by extrapolation using the rate of species accumulation over some measure of sampling effort or in the proportion of rare species in a sample; 2) using known richness ratios of better studied taxa to infer the unknown richness of less well studied taxa, and 3) by assessing the level of underdescription based on expert opinion (colwell and coddington 1994). the first method is the most frequently used method to estimate species richness at the level of a single habitat or of smaller regions (colwell and coddington 1994). in an analogous manner, species richness of larger areas has often been inferred by extrapolating from the rate of species description using time as a crude estimator of sampling effort. often, more sophisticated measure of taxonomic effort is used (dolphin and quicke 2001). unfortunately the assumptions of this method is violated when the number of species increases too rapidly, which is the case for many arthropod taxa. in this paper we have applied this method to a subset of species for which the assumptions are fulfilled. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 2 applications of the second major strategy have proved informative in many instances (dolphin and quicke 2001; adamowicz and purvis 2005; jones et al. 2009), but the results are sometimes highly influenced by the assumptions made by the researcher and this method has in some cases produced highly controversial results. erwin (1982) caused much debate when he published estimates of a mindboggling 30 million arthropod species worldwide by taking the number of beetle species found in a field survey, estimating their host specificity and used this to extrapolate the total number of beetles using the much better known number of tree species in the tropics. a reanalysis of the same model using estimates of host specificity more in line with recent empirical findings (novotny et al. 2007), gave more moderate estimates of 4.8 million arthropod species (ødegaard 2000), in the present paper we use taxonomic ratios in a novel way by using revised groups of species as indicators of the yet undiscovered diversity of unrevised taxa. the third approach has largely been neglected (but see gaston 1990), but is certainly worth mentioning, since the experience of trained taxonomists acquired by working through enormous numbers of specimens must not be ignored when we assess the reliability of the ‗guesstimates‘ reached through more or less esoteric statistical artistry. taxonomic biases the data we use to estimate species richness at any level inevitably contain multiple taxonomic biases. any outcome of field surveys or regional species lists relies on and reproduces biases found in the taxonomic literature. if the taxonomic description level differs between two taxa, these differences will likely be reflected in species counts from surveys, even if the true number of species present in a sample is the same. unfortunately, taxonomical, geographical and trait-specific biases are common, strong and diverging (blackburn and gaston 1994; colleen et al. 2004, reed and boback 2002). taxonomic biases may arise through several mechanisms: some taxonomic groups have caught much more interest than others and the numbers of taxonomic workers differ between geographic regions (gaston and may 1992) and the effort spent in different regions has changed through time (allsopp 1997). methodological biases such as the efficiency in which different collecting methods capture different kinds of figure 1. delimitation of zoogeographical regions and species richness estimates. bars show the described and predicted number of diving beetle species in each region. the left-hand bar in each pair shows the number of species considered distinct up to and including may 2009 and the right-hand bar shows the number of species predicted by the method using taxonomic revision data. the total number of species is given on top of each bar. the number of unrevised endemic, revised endemic and multiregional are indicated by different shades, see the text for details. note that the estimation method only affects the number of species in endemic taxa which have not yet been revised. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 3 insects also have strong effects (king and porter 2005). one of the stronger relationships is that smaller and less conspicuous species generally are described later. this has been shown in for example american butterflies (gaston et al. 1995), british beetles (gaston 1991), western palaearctic dung beetles (cabrero-sanudo and lobo 2003), american oscine passerines (blackburn and gaston 1994) and carnivorous mammals (collen et al. 2004). in some instances however, size has proved to be a poor predictor of description date. this is the case in, for example, primates (collen et al. 2004), australian scarabaeid beetles (allsopp 1997), north american amphibians and reptiles (reed and boback 2002) and marine holozooplankton (gibbons et al. 2005). elucidating the effects of taxonomic biases will improve our understanding of the true patterns of diversity and help us interpret the bits and pieces of information that we have at hand. in this paper we use data from taxonomic revisions to study the effects of taxonomic biases on the species richness of predacious diving beetles. revisions are most often intended to counteract the disorder which builds up with time in the literature and museum collections, but we argue that they also constitute quantitative tests of how the previous level of description compares to the true species richness of a taxon. fundamental to this approach is the hypothesis-testing nature of ‗descriptive‘ taxonomy, which is rarely fully appreciated. every species name is an explicitly formulated hypothesis regarding the distribution of genetic and morphological characteristics among populations of organisms (wheeler 2004). every time we set out to identify a specimen we test such hypotheses. taxonomic revisions contain the summed effects of a large number of independent tests of all proposed species names applied to all available material in major museum collection. to use the outcome from taxonomic revisions for the purpose of species richness estimation was first proposed in a bachelor level thesis (nilsson 2006), which however only dealt with a subset of the revisions analyzed here and employed rather crude statistical methods. recently a very similar approach has successfully been used to estimate the number and distribution of undescribed species of braconid parasitic wasps (jones et al. 2009). however, revision data is not devoid of biases. some revisions are carried out when a large amount of new material has been collected, but a majority of revisions mostly deal with previously collected museum material (dikow et al. 2009), so it is likely that the effect from recent collecting events will have less impact on the results, making the estimates somewhat conservative in this respect. with 4080 described species, the dytiscidae is the largest aquatic beetle family. both adults and larvae of almost all species are aquatic, but pupation takes place on land. only five species are known to be fully terrestrial. dytiscids inhabit a wide range of both lotic and lentic freshwater habitats from 30 m below ground to 4,700 m above sea level (jäch and balke 2008) and no strictly marine species are known. adults of most species are carnivorous, but may also be scavengers (kristensen and beutel 2005). methods data collection to collect data on the outcome of revisionary work carried out on the dytiscidae so far, we made an exhaustive search in the zoological record for all publications published between 1978 and may 25 th 2009 containing the word dytiscidae together with the word revision, monograph or review in the title. this search returned 101 papers. we also searched the reference collection in the dytiscidae database (see below for details) for all papers published during this time span where at least 3 species where described or synonymized, expanding the list to include 129 papers which were considered in detail. for a complete list see the appendix i. we a priori decided on the following criteria for which papers to be classified as revisions: 1) the taxonomic scope of the revision must be explicitly stated at generic or subgeneric level, 2) it must not be a mere faunistic review and 3) must not have supra-specific taxonomy as its main focus. after close scrutiny 88 revisions fulfilled our criteria, collectively describing 499 new species and proposing 241 new synonyms. these publications accounts for 50% of all species described and 30% of all synonyms proposed during 1978-2009. based on the distribution records given in the revisions, we extracted the effect the proposed taxonomic changes had on the number of species in each zoogeographical region separately. the delimitation of zoogeographical regions (or simply ‗regions‘) used follows the most recent published version of the world catalogue of the nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 4 dytiscidae (nilsson 2001). these are the afrotropical (af), australian (au), nearctic (ne), neotropical (nt), oriental (or), pacific (pc) and palaearctic (pl) regions (fig. 1). from each revision we collected region-specific information on 1) the taxonomic coverage of the revision; 2) the number of separate species recognised after the revision; 3) the number of new species described and named; and 4) the number of species names synonymized. we calculated the number of names considered valid prior to the revision by subtracting the number of newly described species and adding the new synonymies to the number of species recognized after the revision. most revisions treated entire genera or tribes, while some dealt with subgenera or species groups. in a few cases there were considerable overlap in the species that where revised (see for example nilsson [1998], štastný [2003] and hendrich and balke [2000], which all dealt with oriental platynectes in the subgenus gueorguievtes). in such cases, the results from two or three revisions were treated as a single revision and the exact number of unique species treated where carefully extracted from the original publications. from the 88 revisions, 115 observations with regional data at the genus or species group level were extracted and used for analyzing taxonomic bias in the description process. a simplified version of the dataset can be found in appendix ii. from this data we calculated the number of revised species, n r i,k. a second dataset was compiled which contained data on the date of description, known global distribution divided into zoogeographical regions and mean body length (calculated as the average of minimum and maximum body length) for all 4080 species of diving beetles recognized as of may 25 th 2009. body length data was only missing for four species which were excluded from the analysis. this dataset was used for summarizing the final species richness estimated using revision data, to extrapolate from the rates of taxonomic description and to explore patterns in the current and historical knowledge of the dytiscidae. information for this dataset was collected from the dytiscidae database assembled by a.n. while working with the world catalogue of insects volume on dytiscidae (nilsson 2001). its information is drawn directly from original sources and studies of type material and includes information on type locality, type depository, global distribution, notes on synonymy and body length of all dytiscid taxa. the database is continuously kept updated and information from all taxonomic publications published before the end of may 2009 has been included in this analysis. method i: using taxonomic revision data this method assumes that the taxonomic revisionary process is a random process insofar as the groups of species subject to revisions are selected effectively at random. we further assume that revised groups approach their true diversity after a revision has been carried out. the second assumption is likely to be frequently violated, causing the method to underestimate the true diversity. neither of these assumptions has to our knowledge been subject to any scientific study but certainly merits closer examination. we used gamma generalized linear models (glm) with log link function (faraway 2004) to analyze the outcome of the taxonomic revisions. we used the loge-transformed ratio between the number of species considered valid after and before the revision as our response variable and zoogeographical region, body length and taxonomic group as additive predictor variables. the non-normal distribution of the response variable and body length data motivated the use of glm, which does not assume that the variables are normally distributed (mccullagh and nelder 1989). ratio data, being one random value divided by another random value, have rather unique statistical properties. assuming that data from taxonomic revisions represent randomly drawn observations on the level of underdescription of a taxon, the outcome of revisions form a rather distinct class of data where we do not expect that large values of the denominator (the number of species before) necessarily dictate large values of the nominator (the number of species after). such data can be modelled as two independent gamma random variables with the ratio of these being non-normal and the relationship between nominator and denominator being weak (liermann et al. 2004), an approach which we have adopted here. all statistical analyses were made using r version 2.9.0 with the lme4 and sampling packages. another bias which may be found is that the revision data that we use to reach our estimates mostly reflect the large amount of specimens nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 5 that are already collected, but hidden among poorly examined museum material, and only to a lesser extent reflect what happens when new material is collected from previously undersampled areas. as more and more remote areas are sampled, the pattern, and especially the magnitude of the undescribed diversity may change although the 88 revisions considered in this paper offer a window into this great unknown. for the taxonomic groupings used to investigate taxonomic biases, we decided not to use the ten dytiscid subfamilies as this would fail to capture any information on taxonomic biases found within the largest subfamily, hydroporinae, which encompass well over half of the world´s diving beetle species. instead, we constructed a taxonomic framework based on the following criteria: 1) the chosen groups should have support from molecular data while not severely violating the traditionally used classification, 2) include the vast majority of described species, and 3) preferably be treated in a fairly equal number of revisionary works. using these criteria we decided on a phylogenetic framework containing 7 groups corresponding to either subfamilies or tribes, collectively containing more than 95% of all dytiscid species and all revised taxa. these groups were: 1) bidessini+pachydrini; 2) colymbetinae sl., including agabinae, colymbetinae and dytiscini, chiefly corresponding to lineage 2 in the pr alignment of ribera et al. (2008); 3) copelatinae sl. (sensu ribera et al. [2008]); 4) hydaticini sl. (sensu ribera et al. [2008], eg. dytiscinae excluding dytiscini); 5) hydroporini sl. (sensu ribera et al. (2008); 6) hydrovatini+vatellini and 7) laccophilini. these groupings generally had strong molecular support from ribera et al.‘s (2008) phylogenic analysis. in their analysis they used four gene fragments with a combined length of about 4000 aligned base pairs and performed separate analyses using three different sequence alignments. in their bayesian analyses, almost all the groupings listed above had posterior probabilities (pp) above 0.90 in two or more of these alignments. there are two exceptions to this, however. one is the colymbetinae sl., which mostly contain large, bulky species. a monophyletic origin was strongly supported by just one of the alignments. however, internal relationships between these taxa were always poorly resolved, and they were consistently placed basal to other taxa. we feel that morphology and molecular data taken together lends sufficient support for treating them together in this context. the second exception is the very diverse assemblage hydroporini sl. this grouping had moderate support, with pp of 0.90 and 0.50 from two alignments, respectively, but this uncertainty is mainly caused by difficulties resolving the placement of a few associated species-poor genera. these taken together, hydroporini.sl form a sister group to the two well supported groups bidessini+pachydrini and hydrovatini+vatellini, effectively splitting the huge subfamily hydroporinae into two subequal halves, and the second and third criterion is thus well fulfilled. the regional coverage among the 115 outcomes of revisions on the regional level where as follows: afrotropical 19 (16.67%), australian 11 (9.65%), nearctic 26 (22.81%), neotropical 14 (12.28%), oriental 15 (13.16%) and palaearctic 28 (24.56%). the taxonomic coverage was: bidessini+pachydrini 21 (18.42%), colymbetinae sl. 26 (22.81%), copelatinae sl. 6 (5.26%), hydaticini sl. 5 (4.38%), hydroporini sl. 42 (36.84%), hydrovatini+vatellini 9 (7.89%) and laccophilini 5 (4.39%). predicting the unrevised diversity the glm model fitted to the revision data was used to predict the hypothetical loge(after/before) outcome of future revisions of all genera. the back-transformed after/before ratio was used to correct regional species richness for biases attributable to differences in body size, taxonomic group and zoogeographical region. to calculate the absolute effects the predictive model had on regional species richness, we first had to address two questions: the species richness of genera which have already been revised must not be corrected one more time and the number of species occurring on multiple continents should not be counted in each region separately, which would inflate the number of species predicted by this approach. to deal with these issues we partitioned the regional species richness of each genus into three components: revised endemic species, unrevised endemic species and multiregional species. from the revision dataset described above we counted the number of species in each genus which had been subject to nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 6 at least one taxonomic revision within the defined time span, n r i,k. then we calculated the number of unrevised endemic species, n u i,k, by assuming that species occurring in one or more zoogeographical regions are equally likely to be included in the taxonomic revisions, so that we for each genera i and region k get: n u i,k = n e i,k – ( ∙ n m i,k), (eq. 1) where n e i,k is the number of endemic species, is the total number of species revised, is the total number of species and n m i,k is the number of multiregional species. this way our total diversity estimate was given by: n total i,k = n m i,k + + n u i,k ∙ multiplier(predicted) (eq.2) where is the number of revised endemic species. to produce error estimates of the predicted outcomes and avoid biases from outlier observations, we carried out these predictions using a stratified resampling approach. in each resampling round, 80% of the observations from each of the seven taxonomic groups were sampled without replacement from the original revision dataset. this subset of observations was used to construct a new glm model which predicted the loge(after/before) ratio of unrevised groups, given the known region, body length and taxonomic group. the predicted logarithmic outcome ratios of all genera were backtransformed and multiplied with the number of unrevised species in each genus. the total predicted diversity was then calculated using equation 2. in each round we calculated the following summary statistics: 1) species richness of each genus, 2) species richness of each subfamily, 3) predicted global diversity of diving beetles, and 4) glm model summary statistics. the stratified resampling routines were written in the statistical computer language r (r development core team 2007). the three species poor tribes hydrodytini (4 species) lancetini (22 species) and matini (8 species) form a sister group to the rest of the colymbetinae sl. as defined above (ribera et al. 2008). they were excluded from the estimation model based on their remarkably low speciation rate compared to the other taxa in this group. seven genera with uncertain positions within the hydroporinae were also excluded from the analysis: hydrodessus (17), kuschelydrus (1), morimotoa (3), pachydrus (9), phreatodessus (2), terradessus (2), typhlodessus (1), and agabetini (2). no revisions dealt with any group of species within cybistrini. these have traditionally been placed in the dytiscinae, with which they share the generally large body size, but ribera et al. (2008) showed that they are closer to the generally much smaller hydroporinae. treating this group with either dytiscinae or hydroporinae both seemed rather spurious. thus, the 134 cybistrine species were excluded from the estimation procedure. uncorrected figures of the number of species from all excluded groups are presented in the total species richness summaries. when data from balke´s (1998) revision of new guinean exocelina, where the number of species increased by a factor of 16.5 from 2 to 33 species, was included in the analysis, the bootstrap did not converge. this observation was therefore excluded from the analysis. method ii: rates of taxonomic description in many instances, the process of taxonomic description proceeds in a manner very similar to the accumulation of new species caught in local biodiversity surveys with constant sampling effort over time. the similarity between these two classes of data is frequently used to justify the estimation of the magnitude of global species richness using data from the species description process in the same manner as estimations of local species richness (colwell and coddington 1994). however, the two types of data differ in several aspects. to be able to use species description data for this purpose, we must have considered that: 1) the effort spent on taxonomic discovery and description varies greatly over time and between regions; 2) curves showing the rate of description plotted against time often displays a distinct ―lag phase‖ of low description rates in the initial stage; 3) the precise relationship between taxonomic effort and description rate is poorly understood, making it difficult to justify which model to fit to the data; 4) the description process frequently proceed in pronounced leaps marking the publication of important monographic works and revision which are impossible to predict using extrapolation techniques; 5) if there is no sign of decrease in the rate of description over nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 7 time, extrapolating techniques cannot be applied (white 1979). this restricts the usage of such techniques to better studied taxa which likely are not very good representatives of the great majority of the world´s taxa. we explored for which groups of diving beetles this approach can be applied and used the species richness estimates of those groups to corroborate the results from the estimates reached using the method with revision data. we dealt with the uneven taxonomic effort by considering the varying number of taxonomic workers active at any point in time and standardizing the description rate against this index. following dolphin and quicke (2001), we defined the term ―prospective description rate‖ as the number of species described for every 20 taxonomist-years. the ―lag phase‖ of low description rates in the initial stage of species description is usually ascribed to the process of taxonomic organization, when much effort is required for discovering morphological features useful for species delimitation and developing a systematic framework above the species level in which to incorporate species (o'brien and wibmer 1979). sometime after this step the rate of description usually reaches its maximum and the effort required for the discovery of new species is at its lowest. provided that this is an adequate description of the lag-phase, we can assume that it contains little information about the magnitude of global species richness and avoid the problems inflicted by it by excluding data from the lag phase from our analysis and consider only the data from after the time when the description rate reached its highest level. plotting the prospective description rate against the number of species described, we expect a negative, linear relationship where the amount of work required for the discovery of new species steadily increases, and the prospective description rate decreases, as more and more species are described. fitting a linear model to such data, the x-intercept of the regression will provide us with an estimate of the number of species described when the description rate is zero. using the data from the dytiscidae database, we compiled a list of all taxonomists that have described at least one dytiscid species, now considered distinct. for each author we noted to which zoogeographical regions and dytiscid subfamilies the described species belong to. we also noted the publication date of the first and last species described, taking the period between these dates as the active period of each author. from this information we calculated the cumulative taxonomic effort that had been invested for each taxon and region and the number of species described in each 20 taxonomist-year. we also tested whether there were a relationship between body length and the date of description in diving beetles. neither body length nor date of description was normally distributed (shapiro-wilk test, data not shown), therefore kruskal-wallis rank sum test was used. results outcome of taxonomic revisions of the 4080 dytiscid species described today, 1644 have been critically examined as part of a taxonomic revision during the last 31 years. in the 88 revisions analyzed, the ratio of the number of species recognized after and before the revision increased on average by a factor of 1.65. if we would apply this correction factor directly to the 2352 endemic species that were not covered by these revisions, we would reach an estimate of the global diversity of diving beetles at about 5822 species, constituting an 1.43-fold increase. however, modelling the outcome of taxonomic revisions as a function of size, distribution and taxonomic group allowed us to correct for these biases, providing us with much more precise estimates. the full glm model analyzing the outcome of taxonomic revisions using the full revision dataset had an nagelkerke pseudo-r 2 (nagelkerke 1991) of 0.30 and both region (χ 2 =19.96, d.f. = 5, p = 0.001) and body length (χ 2 =6.988793, d.f.= 1, p<0.001) were significant while the effect of taxonomic groups was not (χ 2 = 10.81001, d.f = 6, p = 0.0944). nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 8 table 1. species richness of diving beetle subfamilies divided into zoogeographical regions. shown are the number of species which occur in more than one zoogeographical region, the numbers of named, distinct species known only from a given region, the estimated number of endemic species when we correct for taxonomic biases using the outcome of taxonomic revisions and the relative increase of species richness this constitutes. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 9 our best estimate of the global species richness of diving beetles is 5405 species (95% assymetrical quartiles = [5086-5661]), which we reached by carrying out 10.000 bootstrap replicates using a random subset of the revision data (table 1). this constitutes a 1.25 to 1.39fold increase in the number of species compared to the number of species known today. the largest relative regional increases are found in the neotropic (1.61-fold increase, 95% quantiles [1.42-1.81]), australian (1.56-fold increase, 95% quantiles [1.34-1.83] and oriental (1.49-fold increase, 95% quantiles [1.36-1.62]) regions, while the increase is less pronounced in the palaearctic (1.22-fold increase, 95% quantiles [1.15-1.29]), nearctic (1.18-fold increase, 95% quantiles [1.15-1.23]) and african (1.12-fold increase, 95% quantiles [1.02-1.21]) regions (fig 1). this has some minor effects on our understanding on the distribution of diversity among the zoogeographical regions. the neotropical region surpasses the palaearctic region as the second most species rich region following the afrotropical and it suggest that the least diversity is actually found in the nearctic region, not in the australian as the distribution of the currently described diversity suggest. among the subfamilies of dytiscidae, the largest relative increases are expected among the copelatinae (1.63-fold increase, 95% quantiles [1.32-1.91]) and laccophilinae (1.49-fold increase, 95% quantiles [1.13-1.85]), while for the agabinae (1.31-fold increase, 95% quantiles [1.25-1.37]), hydroporinae (1.24-fold increase, 95% quantiles [1.21-1.30]), colymbetinae (1.23fold increase, 95% quantiles [1.05-1.25]) and dytiscinae (1.06-fold increase, 95% quantiles [1.00-1.19]) the increases are less dramatic (table 1). these changes have no effect on the species richness rank-order between subfamilies. rates of description since 1758, when linnaeus described the first diving beetle, 308 taxonomists have spent 2491 taxonomist-years describing the 4080 currently recognized dytiscid species. the description rate reached its maximum around the end of the 19 th century for most taxa and regions, to a great extent due to the work of two individual taxonomists, maurice régimbart and david sharp, who described a most remarkable number of species from all continents between 1870 and 1910, of which 754 are still considered distinct today. for the family as a whole, there is a significant relationship between body length and date of description (r 2 = 0.057, p<0.0001) (fig 2). this decline in body size of described species over time is consistent across different time periods, with the correlation being equally strong between the years 1758-1865 (r 2 = 0.047, n = 540, p<0.0001) and 1910-2009 (r 2 = 0.044, n = 1756, p<0.0001). four zoogeographical regions and two subfamilies show signs of approaching saturation (fig 3). the oriental region showed no signs of saturation and the australian region even display a significant, positive trend in the prospective description rate (r 2 = 0.242, p=0.005) suggesting that this region is in a highly intensive phase of the descriptive process. by fitting a linear model to these groups of taxa and calculating the number of species described when the linear predictor of description rate was zero we calculated the final species richness of these groups. the estimate reached this way were generally directly comparable with the results gained from the method using revision data and the slope of the regression (1.15±0.4) was very close to 1 (fig 4). the major difference is that the estimates from the revision data method predict distinctly more undescribed species in the neotropical region, while the figure 2. relationship between a species‘ date of description and its body length for diving beetles. note that the y-axis is log transformed to show smaller species better. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 10 description rate method predicted larger increases in the palaearctic. discussion we have shown that even for a comparatively well studied arthropod taxon such as the dytiscidae, there is a considerable number of species still undescribed. we estimate there to be 5404 dytiscid species in the world, compared to the 4080 known today. this estimate compares rather well with previous estimates of 5000 species based on expert opinion (jäch and balke 2008) and constitutes a rather modest 1.32-fold increase. there were clear patterns regarding the characteristics of the 1324 species predicted to remain undescribed. the undescribed diversity is likely to be biased toward smaller species from tropical regions outside of africa. compared to the only other study that have used a comparable method, which applied it to braconid parasitic wasps (jones et al. 2009), taxonomic biases in the dytiscidae differs in one striking aspect: while the greatest number of undescribed species was predicted to be found in the afrotropical region for braconid wasps, our study suggest that for diving beetles the afrotropical region will experience the smallest relative increase. it is possible that these diverging patterns are artefactual, demonstrating biases in which groups are chosen for revision in different regions. the alternative explanation is that this reflects true differences between taxa caused by historical and biological factors. in dytiscidae, figure 3. rate of description of new dytiscid species (number of species per 20 taxonomist-year) plotted against the number of species already described at the same time for the four regions and two subfamilies showing signs of decreased description rates. the lines show linear models fitted to the data and the associated r 2 and p-values of the models are shown in the upper right corner. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 11 afrotropical taxa were represented by 19 revisions, covering 46% of the region´s fauna. all these cases displayed uniformly weak responses to revisions and the number of species increased on average by a factor of 1.075. in only a single case (clypeodytes: biström 1988) a taxon increased as much as 1.6-fold. this is further supported by the correlation between estimates for the afrotropical fauna reached by the two independent methods (fig. 3 and fig. 4). taken together, the dytiscid fauna of the afrotropical region indeed appears to be well studied and suggests that we should be cautious not to make simplistic statements about the level of unknown diversity in the tropics. compared to the methods used in the jones et al. (2009) paper, the present study most importantly differ in that we avoid applying the correction factors multiple times to already revised taxa by making a distinction between revised and unrevised taxa as well as between endemic and multiregional species when calculating the effects the predictive model has on the global species richness estimates. we feel that these distinctions are important, especially with regard to taxa where a non-negligible portion of the species pool has been subject to revision. the exclusion of multiregional species constitutes a potential pitfall with this method. doing this does not take into account that many multiregional species may constitute several, cryptic, sibling species, which are sometimes identified in taxonomic revision. as an example, adamowicz and purvis (2005) found that 64.3% of branchiopod species believed to have multiregional distributions de facto represented two or more genetically well separated species when studied in more detail. but since multiregional species constitute a small proportion of the total diversity of diving beetles, and in several cases merely represent the artificial nature of delimitating zoogeographical regions, we argue that treating the number of multiregional species as fixed is unlikely to have a major impact on the final results. the approach we have adopted here offers the potential to correct for a range of biases influencing our knowledge on the distribution of biodiversity. to give an idea of the extent of the taxonomic literature which could be utilized for this purpose, a literature search by meier and dikow (2002) found that more than 2300 zoological revisions were published between 1990 and 2002. given the success of the two attempts to utilize this information carried out so far, we believe that if this method is applied to a broader range of taxa, possibly incorporating additional explanatory variables, we will gain much insight into the magnitude and distribution of species richness which will help us focus taxonomic expertise and funding into areas where they are most needed. references allsopp, p.g. 1997. probability of describing an australian scarab beetle: influence of body size and distribution. journal of biogeography, 24(6), 717-724. adamowicz, s.j. and purvis, a. 2005. how many branchiopod crustacean species are there? quantifying the component of underestimation. global ecology and biogeography 14: 455-468. balke, m. 1998. revision of new guinea copelatus erichson, 1832 (insecta: coleoptera: dytiscidae): the running water species, part i. annalen des naturhistorischen museum in wien (b) 100: 301341. biström, o. 1988. revision of the genus clypeodytes régimbart in africa (coleoptera: dytiscidae). entomologica scandinavica 19: 199238. blackburn, t. and gaston, k. 1994. animal bodysize distributions change as more species are figure 4. comparison between the number of species predicted by the method using taxonomic revision data and the results using the rate of description of groups showing signs of saturation. regression over the observed data (solid line) falls very close to a constant ratio of 1 (dashed line). nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 12 described. proceedings of the royal society of london series b 257: 293–297. cabrero-sanudo, f. and lobo, j. 2003. estimating the number of species not yet described and their characteristics: the case of western palaearctic dung beetle species (coleoptera, scarabaeoidea). biodiversity and conservation, 12, 147–166. chapman, a.d. 2009. numbers of living species in australia and the world. report for the australian biological resources study. canberra, australia. collen, b., purvis, a. and gittleman, j. 2004. biological correlates of description date in carnivores and primates. global ecology and biogeography, 13, 459–467. colwell, r.k. and coddington, j.a. 1994. estimating terrestrial biodiversity through extrapolation. philosophical transactions of the royal society (series b). 345:101-118. dikow, t., meier, r., vaidya, g.g. and londt, j.g.h. 2009. biodiversity research based on taxonomic revisions—a tale of unrealized opportunities. in: diptera diversity: status, challenges and tools. (eds pape, t., bickel, d. and meier, r.). koninklijke brill nv, leiden. dolphin, k. and quicke, d. 2001. estimating the global species richness of an incompletely described taxon: an example using parasitoid wasps (hymenoptera : braconidae). biological journal of the linnean society 73(3): 279-286. erwin, t.l. 1982. tropical forests: their richness in coleoptera and other arthropod species. the coleopterists bulletin 36: 74-75. faraway, julian. 2005. extending the linear model with r: generalized linear, mixed effects and nonparametric regression models. chapman and hall. boca raton, usa. gaston, k.j. 1991. the magnitude of global insect species richness. conservation biology. 5(3): 283-296. gaston, k.j. 1996. species richness, measure and measurement. in: biodiversity (ed. by gaston, k.j.). 77–113. blackwell science, oxford, u.k. gaston, k.j., 2000. global patterns in biodiversity. nature, 405(6783): 220-227. gaston, k.j., blackburn, t. and loder, n. 1995. which species are described first the case of north-american butterflies. biodiversity and conservation 4(2): 119-127. gaston, k.j. and may, r.m. 1992. taxonomy of taxonomists. nature 356(6367): 281-282. gibbons, m. j., richardson, a. j., angel, m. v., buecher, e., esnal, g., fernandez alamo, m. a. gibson, r., itoh, h., pugh, p., boettger-schnack, r. and thuesen, e. 2005. what determines the likelihood of species discovery in marine holozooplankton: is size, range or depth important? oikos 109: 567-576. hendrich, l. and balke, m. 2000. the genus platynectes régimbart in the moluccas (indonesia): taxonomy, faunistics and zoogeography (coleoptera: dytiscidae). koleopterologische rundschau 70: 37-52. hooper, d.u., chapin, f.s., ewel, j.j., hector, a., inchausti, p., lavorel, s. et al. 2005. effects of biodiversity on ecosystem functioning: a consensus of current knowledge. ecological monographs 75(1), 3-35. jones, o.r., purvis, a., baumgart, e. and quicke, d.l.j. 2009. using taxonomic revision data to estimate the geographic and taxonomic distribution of undescribed species richness in the braconidae (hymenoptera: ichneumonoidea). insect conservation and diversity 2(3): 204-212. jäch, m.a. and balke, m. 2008. global diversity of water beetles (coleoptera) in freshwater. hydrobiologia 595: 419-442. king, j. and porter, s., 2005. evaluation of sampling methods and species richness estimators for ants in upland ecosystems in florida. environmental entomology 34(6), 1566-1578. kristensen, n.p. and beutel, r.g. 2005. handbook of zoology volume iv: arthropoda: insecta, volume 1: morphology and systematics (archostemate, adephaga, myxophaga, polyphaga partim) part 38. new york: walter de gruyter. liermann, m., steel, a., rosing, m. and guttorp, p. 2004. random denominators and the analysis of ratio data. environmental and ecological statistics 11: 55–71. mccullagh, p. and nelder, j. a. 1989. generalized linear models (2nd ed.). chapman and hall. london, uk. meier, r., and dikow, t. 2002. significance of specimen databases from taxonomic revisions for estimating and mapping the global species diversity of invertebrates and repatriating reliable specimen data. conservation biology 18(2): 478488. nagelkerke, n.j.d. 1991. a note on a general definition of the coefficient of determination. biometrika, 691–692. nilsson, a.n. 1998. dytiscidae: v. the genus platynectes regimbart in china, with a revision of the dissimilis-complex (coleoptera). in: water beetles of china iii (eds. jäch, m.a. and li, j.). 107-121. nilsson, a.n. 2001. dytiscidae – in: world catalogue of insects 3: 1-395. apollo books, stenstrup, denmark. nilsson, v.j. 2006: using taxonomic data to estimate species diversity: a multivariate approach applied to diving beetles (coleoptera: dytiscidae). bachelor´s thesis. umeå university. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 13 novotny, v., miller, s.e., hulcr, j. drew, r.a.i., basset, y. janda, m. et al. 2007. low beta diversity of herbivorous insects in tropical forests. nature 448(7154): 692-695. o'brien, c.w., and wibmer, g.j. 1979. the use of trend curves of rates of species descriptions: examples from the curculionidae (coleoptera). the coleopterists' bulletin: 151–166. purvis, a. and hector, a. 2000. getting the measure of biodiversity. nature 405(6783): 212-219. r development core team. 2007. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. ribera, i., vogler, a.p., and balke, m. 2008. phylogeny and diversification of diving beetles (coleoptera: dytiscidae). cladistics 24(4): 563590 reed, r.n. and boback, s.m. 2002. does body size predict dates of species description among north american and australian reptiles and amphibians? global ecology and biogeography 11(1): 41-47. štastný, j. 2003: dytiscidae: x. review of platynectes subgen. gueorguievtes vazirani from southeast asia (coleoptera). in: water beetles of china iii (eds. jäch, m.a. and li, j.). 217-259 white, r.e. 1979. response to the use of trend curves by erwin, frank and curtis, and o'brien and wibmer. the coleopterists bulletin 33(2): 167-168. ødegaard, f. 2000. how many species of arthropods? erwin‘s estimate revised. biological journal of the linnean society 71: 583-597. revisions analyzed anderson, r.d. 1971. a revision of the nearctic representatives of hygrotus (coleoptera: dytiscidae). annals of the entomological society of america 64: 503-512. anderson, r.d. 1983. revision of the nearctic species of hygrotus groups 4, 5, and 6 (coleoptera: dytiscidae), annals of the entomological society of america 76(2) 173196. angus, r.b., fresneda, j. and fery, h. 1992. a revision of the nebrioporus carinatus species complex (coleoptera, dytiscidae). nouvelle revue d'entomologie 9(4) 287-303. balke, m. 1995. revision of the afrotropicaloriental rhantus rugulosus-clade (coleoptera: dytiscidae). entomologica scandinavica 26(2) 229-239. balke, m. 1998. revision of new guinea copelatus erichson, 1832 (insecta: coleoptera: dytiscidae): the running water species, part i. annalen des naturhistorischen museum in wien (b) 100: 301341. balke, m. 2001. biogeography and classification of new guinean colymbetini (coleoptera: dytiscidae: colymbetinae). invertebrate taxonomy 15: 259-275. balke, m., larson, d.j., and hendrich, l. 1997. a review of the new guinea species of laccophilus leach 1815 with notes on regional melanism (coleoptera dytiscidae). tropical zoology 10: 295-320. balke, m., larson, d.j., hendrich, l., and konyorah, e. 2000. a revision of the new guinea water beetle genus of philaccolilus guignot, stat. n. (coleoptera dytiscidae). mitteilungen aus dem museum für naturkunde berlin, deutsche entomologische zeitung 47: 29-50. bergsten, j., and miller, k.b. 2006. taxonomic revision of the holarctic diving beetle genus acilius leach (coleoptera: dytiscidae). systematic entomology 31: 145-197. biström, o. 1979. a revision of the genus derovatellus sharp (coleoptera, dytiscidae) in africa. acta entomologica fennica 35: 128. biström, o. 1982. a revision of the genus hyphydrus illiger (coleoptera, dytiscidae). acta zoologica fennica 165: 1-121. biström, o. 1983. revision of the genera yola des gozis and yolina guignot (coleoptera, dytiscidae). acta zoologica fennica 176: 1-67. biström, o. 1985. a revision of the species group b. sharpi in the genus bidessus (coleoptera, dytiscidae). acta zoologica fennica 178: 1-40. biström, o. 1986. review of the genus hydroglyphus motschulsky (= guignotus houlbert) in africa (coleoptera, dytiscidae). acta zoologica fennica 182: 1-56. biström, o. 1987a. review of the genus leiodytes in africa (coleoptera, dytiscidae). annales entomologici fennici 53(3): 91-101. biström, o. 1987b. revision of the genus pachynectes regimbart (coleoptera, dytiscidae). annales entomologici fennici 53(2) 48-52. biström, o. 1988a. revision of the genus clypeodytes régimbart in africa (coleoptera: dytiscidae). entomologica scandinavica 19: 199 238. biström, o. 1988b. review of the genus liodessus in africa (coleoptera, dytiscidae). annales entomologici fennici 54: 21-28. biström, o. 1988c. review of the genus uvarus guignot in africa (coleoptera, dytiscidae). acta entomologica fennica 51: 1-38. biström, o. 1990. revision of the genus queda sharp (coleoptera: dytiscidae). quaestiones entomologicae 26(2): 211-220. biström, o. 1997. taxonomic revision of the genus hydrovatus motschulsky (coleoptera, dytiscidae). entomologica basiliensia 19: 57584. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 14 biström, o., and nilsson, a.n. 2002. herophydrus sharp: cladistic analysis, taxonomic revision of the african species, and world check list (coleoptera: dytiscidae). koleopterologische rundschau 72: 15-111. biström, o. and nilsson, a.n. 2003. taxonomic revision and cladistic analysis of the genus peschetius guignot (coleoptera: dytiscidae). aquatic insects 25(2): 125-155. biström, o., and nilsson, a.n. 2006. taxonomic revision of the ethiopian genus canthyporus (coleoptera dytiscidae). memorie della societa entomologica italiana 85: 207-304. brancucci, m. 1983. révision des espèces estpaléarctiques, orientales et australiennes du genre laccophilus (col. dytiscidae). entomologische arbeiten aus dem museum g. frey 31-32: 241426. brancucci, m. 1986. revision of the genus lacconectus motschulsky (coleoptera, dytiscidae). entomologica basiliensia 11: 81202. brancucci, m. 1988. a revision of the genus platambus thomson (coleoptera, dytiscidae). entomologica basiliensia 12: 165-239. brancucci, m. 2003. a review of the genus lacconectus motschulsky, 1855 from the indian subcontinent (coleoptera, dytiscidae). entomologica basiliensia 25: 3-39. brancucci, m. and hendrich, l. 2005. a review of indomalayan lacconectus motschulsky, 1855 (coleoptera: dytiscidae: copelatinae). mitteilungen der schweizerischen entomologischen gesellschaft 78(3-4) 265-297. fery, h. 1991. revision der 'minutissimus-gruppe' der gattung bidessus sharp (coleoptera: dytiscidae). entomologica basiliensia 14: 57-91. fery, h. 1992. revision der saginatus-gruppe der gattung coelambus thomson (coleoptera: dytiscidae). linzer biologische beitraege 24(1): 339-358. fery, h. 1999. revision of a part of the memnoniusgroup of hydroporus clairville, 1806 (insecta: coleoptera: dytiscidae) with the description of nine new taxa, and notes on other species of the genus. annalen des naturhistorischen museums in wien (b) 101: 217-269. fery, h. and brancucci, m. 1997. a taxonomic revision of deronectes sharp, 1882 (insecta: coleoptera: dytiscidae) (part i). annalen des naturhistorischen museums in wien (b) 99: 217302. fery, h. and hosseinie, s.o. 1998. a taxonomic revision of deronectes sharp, 1882 (insecta: coleoptera: dytiscidae) (part ii). annalen des naturhistorischen museums in wien (b) 100: 219-290. fery, h. and nilsson, a.n. 1993. a revision of the agabus chalconatusand erichsoni-groups (coleoptera: dytiscidae), with a proposed phylogeny. entomologica scandinavica 24: 79108. hendrich, l. and balke, m. 1997. taxonomische revision der südostasiatischen arten der gattung neptosternus sharp, 1882 (coleoptera: dytiscidae: laccophilinae). koleopterologische rundschau 67: 53-97. hendrich, l. and balke, m. 2000. the genus platynectes régimbart in the moluccas (indonesia): taxonomy, faunistics and zoogeography (coleoptera: dytiscidae). koleopterologische rundschau 70: 37-52. hendrich, l. and wang, l. 2006. taxonomic revision of australian clypeodytes (coleoptera: dytiscidae, bidessini). entomological problems 36 (1): 1-11. hendrich, l. and watts, c.h.s. 2004. taxonomic revision of the australian genus sternopriscus sharp, 1882 (coleoptera: dytiscidae: hydroporinae). koleopterologische rundschau 74: 75-142. hendrich, l. and watts, c.h.s. 2009. taxonomic revision of the australian predaceous water beetle genus carabhydrus watts, 1978 (col. dytiscidae, hydroporinae, hydroporini). zootaxa 2048: 1-30. larson, d.j. 1987. revision of north american species of ilybius erichson (coleoptera: dytiscidae), with systematic notes on palearctic species. journal of the new york entomological society 95: 341-413. larson, d.j. 1989. revision of north american agabus leach (coleoptera: dytiscidae): introduction, key to species groups, and classification of the ambiguus-, tristis-, and arcticus-groups. the canadian entomologist 121: 861-919. larson, d.j. 1991. revision of north american agabus leach (coleptera: dytiscidae): elongatus, zetterstedti-, and confinis-groups. the canadian entomologist 123: 1239-1317. larson, d.j. 1994. revision of north american agabus leach (coleoptera: dytiscidae): lutosus-, obsoletus-, and fuscipennis-groups. the canadian entomologist 126: 135-181. larson, d.j. 1996. revision of north american agabus leach (coleoptera: dytiscidae): the opacus-group. the canadian entomologist 128: 613-665. larson, d.j. 1997. revision of north american agabus leach (coleoptera: dytiscidae): the seriatus-group. the canadian entomologist 129: 105-149. larson, d.j. and roughley, r.e. 1990. a review of the species of liodessus guignot of north america north of mexico with the description of a new species (coleoptera: dytiscidae). journal of the new york entomological society 98: 233245. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 15 larson, d.j. and wolfe, g.w. 1998. revision of north american agabus leach (coleoptera: dytiscidae): the semivittatus-group. the canadian entomologist 130: 27-54. matta, j.f. and wolfe, g.w. 1981. a revision of the subgenus heterosternuta strand of hydroporus clairville (coleoptera: dytiscidae). the panpacific entomologist 57: 176-219. miller, k.b. 1998. revision of the nearctic liodessus affinis (say 1823) species group (coleoptera: dytiscidae, hydroporinae, bidessini). entomologica scandinavica 29(3): 281-314. miller, k.b. 2001a. revision of the neotropical genus hemibidessus zimmermann (coleoptera: dytiscidae: hydroporinae: bidessini). aquatic insects 23(4): 253-275. miller, k.b. 2001b. revision of the genus agaporomorphus zimmermann (coleoptera: dytiscidae). annals of the entomological society of america 94: 520-529. miller, k.b. 2001c. revision and phylogenetic analysis of the new world genus neoclypeodytes young (coleoptera: dytiscidae: hydroporinae: bidessini). systematic entomology 26: 87-123. miller, k.b. 2005. revision of the new world and south-east asian vatellini (coleoptera: dytiscidae: hydroporinae) and phylogenetic analysis of the tribe. zoological journal of the linnean society 144: 415-510. nilsson, a.n. 1992a. a revision of afrotropical agabus leach (coleoptera, dytiscidae), and the evolution of tropicoalpine super specialists. systematic entomology 17: 155-179. nilsson, a.n. 1992b. a revision of the east african nebrioporus abyssinicus group (coleoptera, dytiscidae). entomologica fennica 3(2) 81-93. nilsson, a.n. 1994a. revision of the hydroporus nigellus complex (coleoptera: dytiscidae) including multivariate species separation. entomologica scandinavica 25(1): 89-104. nilsson, a.n. 1994b. a revision of the palearctic ilybius crassus-complex (coleoptera, dytiscidae). entomologisk tidskrift 115(1-2): 55-61. nilsson, a.n. 1997. a redefinition and revision of the agabus optatus-group (coleoptera, dytiscidae); an example of pacific intercontinental disjunction. entomologica basiliensia 19: 621-651. nilsson, a.n. 1998. dytiscidae: v. the genus platynectes regimbart in china, with a revision of the dissimilis-complex (coleoptera). in: water beetles of china iii (eds. jäch, m.a. and li, j.). 107-121. nilsson, a.n. and larson, d.j. 1990. a review of the agabus affinis group (coleoptera: dytiscidae), with the description of a new species from siberia and a proposed phylogeny. systematic entomology 15(2): 227-239. nilsson, a.n. and nakane, t. 1993. a revision of the hydroporus species (coleoptera: dytiscidae) of japan, the kuril islands, and sakhalin. entomologica scandinavica 23(4): 419-428. roughley, r.e. 1990. a systematic revision of species of dytiscus linnaeus (coleoptera: dytiscidae). part 1. classification based on adult stage. quaestiones entomologicae 26: 383-557. satô, m. 1985. the genus copelatus of japan (coleoptera: dytiscidae). transactions of the shikoku entomological society 17: 57-67. štastný, j. 2003: dytiscidae: x. review of platynectes subgen. gueorguievtes vazirani from southeast asia (coleoptera). in: water beetles of china iii (eds. jäch, m.a. and li, j.). 217-259 shaverdo, h.v. 2004. revision of the nigrita-group of hydroporus clairville, 1806 (insecta: coleoptera: dytiscidae). annalen des naturhistorischen museums in wien (b) 105: 217-263. shaverdo, h.v. 2006. revision of the longiusculusgroup of the genus hydroporus clairville, 1806 (coleoptera: dytiscidae). zootaxa 1170: 27-56. shirt, d.b. and angus, r.b. 1992. a revision of the nearctic water beetles related to potamonectes depressus (fabricius) (coleoptera: dytiscidae). coleopterists bulletin 46(2): 109-141. toledo, m. 2009. revision in part of the genus nebrioporus regimbart, 1906, with emphasis on the n. laeviventris-group (coleoptera: dytiscidae). zootaxa 2040: 1-111. trémouilles, e.r. 1996. revisión del género hydaticus leach en américa del sur, con descripción de tres nuevas especies (coleoptera, dytiscidae). physis, secc. b52: 15-32. watts, c.h.s. 2000. three new species of tiporus watts (coleoptera: dytiscidae) with redescriptions of the other species in the genus. records of the south australian museum 33: 8999. watts, c.h.s. and leys, r. 2008. review of the epigean species of paroster sharp, 1882, with descriptions of three new species, and phylogeny based on dna sequence data of two mitochondrial genes (coleoptera: dytiscidae: hydroporinae). koleopterologische rundschau 78: 9-36. wewalka, g. 1979. revision der artengruppe des hydaticus (guignotites) fabricii (mac leay), (col., dytiscidae). koleopterologische rundschau 54: 119-139. wewalka, g. 1980. revision der afrikanischen gattung heterhydrus fairm. (coleoptera, dytiscidae). annales historico-naturales musei nationalis hungarici 72: 97-101. wewalka, g. 1997. taxonomic revision of microdytes balfour-browne (coleoptera: dytiscidae). koleopterologische rundschau 67: 13-51. nilsson-örtman & nilsson using taxonomic revision data to estimate species richness 16 wewalka, g. 2000. taxonomic revision of allopachria (coleoptera: dytiscidae). entomological problems 31: 97-128. wolfe, g.w. 1984. a revision of the vittatipennis species group of hydroporus clairville, subgenus neoporus guignot (coleoptera: dytiscidae). transactions of the american entomological society 110(3): 389-433. wolfe, g.w.and roughley r.e. 1990. a taxonomic phylogenetic and zoogeographic analysis of laccornis gozis (coleoptera: dytiscidae) with the description of laccornini, a new tribe of hydroporinae. quaestiones entomologicae 26: 273-354. young, f.n. 1981a. predaceous water beetles of the genus desmopachria babington: the leechiglabricula group (coleoptera: dytiscidae). the pan-pacific entomologist 57(1): 57-64. young, f.n. 1986. review of the predaceous water beetles of the genus bidessodes régimbart (coleoptera, dytiscidae). entomologica basiliensia 11: 203-220. young, f.n. 1981c. predaceous water beetles of the genus neobidessus from south america (coleoptera: dytiscidae). the coleopterists bulletin 35(3): 317-340. young, f.n. 1981b. predaceous water beetles of the genus desmopachria: the convexa-grana group (coleoptera: dytiscidae). occasional papers of the florida state collection of arthropods 2(iiiiv): 1-11. young, f.n. 1990. predaceous water beetles of the genus desmopachria babington: the subgenus pachriostrix guignot (coleoptera: dytiscidae). the coleopterists bulletin 44(2): 224-228. young, f.n. 1995. the genus desmopachria babington, subgenus portmannia young (coleoptera: dytiscidae). insecta mundi 9: 37-44. zimmerman, j.r. 1981. a revision of the colymbetes of north america (dytiscidae). the coleopterists bulletin 35(1): 1-52. zimmerman, j.r. 1985. a revision of the genus oreodytes in north america (coleoptera: dytiscidae). proceedings of the academy of natural sciences of philadelphia 137(1): 99-127. microsoft word honcuibilayoutapr08draft3.doc biodiversity informatics, 5, 2008, pp. 20-40 converting taxonomic descriptions to new digital formats hong cui 1 1 school of information resources and library science, university of arizona, 1515 e. first street, tucson, az, 85742. e-mail: mhongcui@email.arizona.edu abstract. the majority of taxonomic descriptions are currently in print format. the majority of digital descriptions are in a format, such as doc, html, or pdf, for human readers. these formats do not convey rich semantics in taxonomic descriptions for computer-aided processing. newer digital formats, such as xml and rdf, accommodate semantic annotations that allow a computer to process the rich semantics on human's behalf, opening up opportunities for a wide range of innovative usages of taxonomic descriptions, including searching in more precise and flexible ways, integrating morphological, genomic, georeference, or other information, automatically generating taxonomic keys, and knowledge mining and visualizing taxonomic data etc. this paper reports our experience with the development of an automated semantic markup system named martt and discusses challenging issues involved. to address these challenging issues, a number of utilities were implemented to make martt a more operable system. the utilities can be used to speed up the preparation of training examples for martt, to facilitate the creation of more comprehensive annotation schemas, and to predict system performance on a new collection of descriptions. martt has been tested on several plant and alga taxonomic publications including flora of china, flora of north america, and flora of north central texas. key words. digital formats, morphological descriptions, semantic markup, supervised machine learning, system evaluation , taxonomic descriptions, unsupervised machine learning, xml. taxonomic descriptions of living organisms are a major information resource used by systematists and evolutionary biologists. the majority of such information is in a print or digital format for human readers. on-going and planned digitalization projects such as those initiated by the global biodiversity information facility (gbif, 2007) and the biodiversity heritage library (bhl, 2007) will likely increase the volumes of taxonomic descriptions in legacy formats (e.g., doc, html, or pdf). these documents will have to be converted to a new digital format such as xml or rdf to allow for any innovative usages beyond keyword-based search. due to the scale of the problem, automated means for the conversion must be sought. large volumes of taxonomic descriptions, print or digital, have been produced over the past two hundred years. while descriptions created by trained taxonomists are of high quality and provide consistent information in general, there is not a well-defined and well-accepted standard to regulate the content of a description. a manual comparison among the descriptions of five plant species, found in six well-known floras, revealed surprisingly large variations in terms of description content and style (lydon et al, 2003). lydon and colleagues found that only 9% of information was exactly the same in six sources, over 55% of information was from a single source, and around 1% of information contradicted information from another source. besides the large variation, these findings also suggest that descriptions from different collections are mostly complementary to one another. as lydon et al. (2003) concluded, any automatic markup software program must take the variation into account to avoid an overly-tailored system that works only on one or a few description collections. in other words, it is highly desirable for a system to be easily portable to a different description collection. keeping this in mind, we designed and implemented a portable java application called martt (markuper for taxonomic treatments), which has marked-up >15,000 descriptions from three floras (i.e. flora of north america (fna, 1993 onwards), flora of china (foc, 1994 onwards), and flora of north central texas (diggs, lipscomb, & o’kennon, 1999)) into a cui – converting taxonomic descriptions to new digital formats 21 predefined xml format quite successfully without reconfiguring the system. this paper reports our experience with the development and evaluation of martt and discusses a number of challenging issues identified alone the way. the paper is organized as the following: starting with the design rationale of martt, we go on to report a series of experiments involving the aforementioned floras (readers not caring about technical details can safely skip this section without loss of continuity) and summarize the experimental results. the identified challenging issues are then discussed in detail and the utilities implemented as solutions are examined. after a review of relevant research, we conclude the paper with a plan for future research. system design rationale the design goal of martt was a highly portable system that would work with all professionally prepared taxonomic descriptions in english without having to re-adjust the system on a collection by collection basis. we also designed the system to learn from its experience with wellprepared descriptions, with the hope that it would become capable of tagging less-well-prepared ones (e.g. those created by amateur taxonomists) in the future. more specifically, the system should be able to mark up a plain-text description into an xml document like the one shown in figure 1. note the design goal emphasizes more the system’s ability of making the semantics of descriptions explicit by inserting appropriate tags than the resultant documents’ compliance to an encoding standard. this is because once a description is in xml format, it is easy to convert it to a standard format such as rdf or sdd (structure of descriptive data, an xml standard issued by the biodiversity information standards 1 ). the high portability may be achieved by employing an approach called “supervised machine learning”. in this approach, markup rules used to tag description sentences are not hardcoded but learned from examples of descriptions themselves. these examples are called training examples, which are selected descriptions tagged in a desired xml format by human experts according to an xml schema/dtd. a supervised machine learning algorithm examines/learns from 1 http://www.tdwg.org/standards/. training examples to come up with rules that may be used to tag unseen descriptions. learning from examples affords a flexible system that automatically adjusts its behavior according to the task on hand. for example, if a flora focuses entirely on flowering plants, then the system will not concern itself with tagging seed cones or pollen cones; on the other hand, if only main organ level annotations (i.e. flower, leaf, etc.) are desired and included in the training examples, then the algorithm will gracefully produce markup at that level and not try to insert bract or stamen tags. since the machine learning approach automatically learns markup rules from training examples, it does not require users to supply any rules. to taxonomists, preparing training examples is much easier than providing markup rules. on the other hand, we do realize that preparing training examples is time-consuming. this is one of the issues we shall address in later sections. for markup rules to be reusable across collections, they should not be based on text format cues. for example, a rule “the first bold words represent an organ name” is unlikely reusable, as not all collections use bold face for organ names. instead, the rules should be semantically rich and convey domain knowledge and/or convention, for example, “a berry is a type of fruit”. this type of semantic association rules is likely reusable across collections. based on all these considerations, martt was implemented with three main components. the first component is a machine learning component, which learns markup rules from training examples and applies the rules to tag new descriptions. the second component is a knowledge induction component, which takes a tagged collection to induce semantic association rules from it. the third component is a storage component for the association rules learned over time and is named “the markup rule bank”. when enabled, the markup rule bank answers queries initiated by the learning component. an example query may be “(according to the rule bank’s knowledge), what could be a good tag for ‘berries fleshy to somewhat leathery’”, and the rule bank would likely respond “fruit”. the learning component grows a learning hierarchy on the fly from the given training examples so the hierarchy is always the best fit for the markup task on hand. to illustrate this process, cui – converting taxonomic descriptions to new digital formats 22 let us use the xml description shown in figure 1 as an example. initially, the learning hierarchy has one root node “description”. when the xml description is read into the root node, the root node sees six elements (i.e. “taxon”, “plant-habit-andlife-style”, “leaves”, “flowers”, “fruit”, and “seeds”) in the description element. the root node thus creates six child nodes, one for each element, and dispatches the content of each element to its corresponding child node. for example the newly created child node “taxon” gets the family and genus names. each child node then reads the content received and if needed, creates its child nodes to accommodate any new elements, for example, the node “taxon” creates its two child nodes (“family” and “genus”), one for the family element and the other for the genus element. the process continues until a terminal element is reached in each branch. in the process each node saves the content of its corresponding element as part of its training data to be used later. by the end of reading the xml description into the learning hierarchy, a simple learning hierarchy is created and this hierarchy corresponds exactly to the xml structure of the description. each node in the hierarchy has one piece of training data: the “description” node has the entire description, the “taxon” node has the family and genus names, and the “family” node has the family name, etc. when another training example is read in, the learning hierarchy expands itself to accommodate any new elements not previously seen. suppose the second training example has a stems element in its description element. when the “description” node checks and sees that it does not have a child node for “stems”, it creates one to save the description of the stems there. if there are elements nested in the stems element, the newly created “stems” node creates its child nodes to accommodate those elements. by the time all training examples are read in the learning hierarchy, every element seen in the training examples will have a corresponding node in the hierarchy and the node will have its set of training data. a portion of the learning hierarchy is illustrated in figure 2. in addition to its training data, each node in the learning hierarchy is also equipped with a number of learning/markup algorithms. each node learns how to tag its corresponding segments in a description. when a new description comes, the root node (“description”) tags it into segments, such as plant-habit-and-life-style, leaves and stems, and then sends the segments to their corresponding child nodes, where the segments are further tagged. for example, the “leaves” node further tags its segment into pedicel, petiole, stipule, etc. segments. to see if new descriptions are tagged correctly at each node, the hierarchy also reads in and holds answer keys. in other words, each node is capable of calculating its performance scores. note the disadvantage of this top-down markup strategy is that if an error is made at an upper node, the error is passed down to lower levels. the current implementation of martt does not support back tracking of errors. bromeliaceae guzmania herbs, usually epiphytic, stemless to rarely caulescent. leaves many-ranked, usually ligulate; blade, margins entire. inflorescences 5many-flowered, many-ranked, mostly 2-pinnate to less commonly single spike, flowers laxly to densely arranged; floral bracts broad, conspicuous, mostly obscuring rachis. flowers bisexual; sepals distinct to connate over 1/2 length,usually symmetric; petals with claws adherent to subconnate petal, forming short tube, blade distinct; stamens usually included, adherent to adnate with petal claws; ovary superior. capsules cylindric, dehiscent. seeds with basal, usually tan-brown plumose appendage. figure 1. an example taxonomic description tagged in xml. cui – converting taxonomic descriptions to new digital formats 23 figure 2. a portion of a learning hierarchy in the learning component. illustration from cui & heidorn (2007) with permission. several markup algorithms are available at each node, including a naïve bayesian (nb) classifier, support vector machine (svm) classifier, and a number of homemade algorithms, in order to compare their performance. once descriptions are segmented into sentences, the task of semantic markup essentially becomes the task of text classification, hence nb and svms may be used to assign class labels (i.e. tags) to text segments. for svms, we used the implementation in the bow toolkit (mccallum, 1996). for nb, we implemented a version based on the algorithm described in mitchell (1997). experiments showed that nb and svms did not perform as well as some of our homemade algorithms, especially on elements with little training data. the lack of training data makes it difficult for nb to accurately estimate probabilities and for svms to identify good support vectors. details of the learning algorithms and their performance comparison can be found in cui (2005b) or cui & heidorn (2007). the following section describes the best homemade algorithm, sccp (semantic classes and character patterns), and reports the performance of martt/sccp on the three floras. readers not caring about technical details can safely skip the machine learning algorithm and the experiments with martt system without loss of continuity. the machine learning algorithm sccp markup algorithm first segments descriptions into sentences and then learns to tag the segments. sccp segments descriptions by periods (.) and semicolons (;), which are the typical punctuation marks used in taxonomic descriptions to set off semantic units. sccp uses a set of heuristics to avoid false segmentations at the periods used as a decimal point (e.g., 2.5) or in an abbreviation (such as var., subsp., h. l. james, diam. etc.) or at the semicolons that are part of html entities (e.g.,  ). sccp does not perform any text normalization procedures such as stemming or converting text to lower case. sccp does not use a part of speech (pos) tagger to identify nouns or noun phrases because available pos taggers are typically for the general domain and do not work well with taxonomic descriptions due to differences in grammar and lexicons. instead, sccp uses a frequent pattern and association rule learning method, originated from data mining research, to learn rules of the form: ngram → element (confidence, support), which reads “the n-gram is associated with the element description training examples learned model marked examples answer keys ... flowers training examples learned model marked examples answer keys fruits training examples learned model marked examples answer keys leaves training examples learned model marked examples answer keys pedicel training examples learned model marked examples answer keys petiole training examples learned model marked examples answer keys stipule training examples learned model marked examples answer keys plant-habit-and-life-style training examples learned model marked examples answer keys stems training examples learned model marked examples answer keys bark training examples learned model marked examples answer keys scale training examples learned model marked examples answer keys stele training examples learned model marked examples answer keys ... ... ... ... ... ... cui – converting taxonomic descriptions to new digital formats 24 with confidence (a numerical value) and support (a numerical value)”. in association rule learning, confidence and support are a pair of scores measuring the strength of an association. rules scored higher than a pair of user-defined thresholds areassumed to be good (han & kamber, 2000). adapting from the standard definitions, we define confidence as the ratio of the occurrence of an ngram in an element and the total occurrence of the n-gram, and support as the ratio of the occurrence of the n-gram in the element and the number of segments (i.e. sentences) belonging to the element. sccp learns the association rules from training examples by first generating sets of n-grams and then calculating the confidence and support scores for the candidate rules based on the occurrences of the n-grams in different elements. the leading l (a user defined variable) words in the sentences are used to generate ∑ ≤≤ +− mn nl 1 1 n-grams, where m < l is another user defined variable that defines the length of the longest n-grams. for example, a word sequence “a b c d” with m = 4, l = 4 generates four unigrams: a, b, c, and d; three bigrams: a b, b c, and c d; two 3-grams: a b c and b c d; and one 4gram: a b c d; totally ten n-grams, 41 ≤≤ n . we call the m-grams the “sub-grams” of an n-gram when they are generated from the same n-word sequence and m < n. the generation of n-grams of varied sizes creates a pool of noun phrase candidates. these noun phrases and all possible elements form candidate association rules. the strength of the association between an n-gram and an element is evaluated by the confidence and support scores, calculated from the occurrences of the n-gram in different elements in the training examples. note under this scheme, sub-grams inherit the occurrence counts of their n-grams. this causes undesirable consequences in some cases. suppose the n-gram “seed cones” occurs very frequently in the “seed cones” element and is recognized as a significant indicator of the element, the counting method automatically makes all its sub-grams (i.e. “seed” and “cones”) good indicators of the element as well, while in fact they are not (e.g. “seed” should be an indicator of the “seeds” element). to avoid this problem, the subgrams are not allowed to inherit its n-gram’s occurrence count when the confidence and support scores of the n-gram are greater than a pair of preset values (meaning the n-gram is likely a phrase and should be treated as one semantic unit). the pair of pre-set values should not be confused with the confidence/support thresholds for the association rules. the former values are set lower than the latter and they serve different purposes as described above. in the experiments reported below, we empirically set l = m = 3, the pre-set value pair was set to 0.7 for confidence and 0 for support, and the confidence threshold was set to 0.8 and support threshold was set to 0.035. settings close to these seemed to produce very similar performance. to mark up a new example, sccp segments the text and takes the first l words of the segments to generate n-grams, 1 < n < l. for each segment, by looking up the n-grams in the list of association rules learned earlier, sccp obtains a number of matching rules with confidence and support scores above the thresholds. the matching rules are ranked according to the following criteria applied in this order: the length of the n-gram (i.e., n), the location of the n-gram in the segment, the support score, and the confidence score. rules containing longer n-grams are ranked higher. rules matching n-grams closer to the beginning of the segment are ranked higher. the support score takes priority over the confidence score to favor the rules with more frequent n-grams. the top ranked rule determines the tag for the segment. sccp is also designed to recognize simple character patterns of the elements containing no words. the current version has only one such pattern for recognizing chromosome counts which take a form like “2n = 24” or “x = 12” in description text. experiments with martt the data sets for the experiments were taken from the published volumes of flora of china (foc), flora of north america (fna), and the monograph of flora of north central texas (fnct) with permission. three sets of training examples were manually prepared, including 378 examples selected from 12,000 foc descriptions, 310 from 1300 fna descriptions, and 378 from 1200 fnct descriptions.the tags used in the training examples and the resultant xml documents, such as “plant habit and life style”, were defined in an xml schema (cui, 2005a). the schema was a result of consulting a number of sources, including a plant systematics textbook cui – converting taxonomic descriptions to new digital formats 25 (radford, 1986), the delta format (dallwitz, 1980), and a plant taxonomist. the standard 10-fold cross-validation protocol routinely used to evaluate performance of a machine learning system was used to obtain the performance scores of martt. according to this protocol, each set of training examples was divided into ten equal-sized subsets. in each run, martt used nine subsets to learn markup rules and then tested the markup rules on the tenth subset. the ten subsets allowed for ten such runs, each with a different test set. the average performance over the ten runs was recorded as the final performance score on a collection. the soundness and completeness of the markup produced by martt were measured element by element (i.e., node by node). the soundness was measured by precision (p), which was defined as the ratio of the text segments tagged as an element e correctly and the total segments tagged as e by the algorithm. the completeness was measured by recall (r), which was defined as the ratio of the text segments tagged as e by the algorithm correctly and the total e segments in the collection. the harmonic mean of recall and precision, f-measure = 2pr / (p + r), was then calculated. precision, recall and f-measure are standard measures routinely used to evaluate performance of information retrieval systems. these measures were borrowed to measure the soundness and completeness of tag assignments. the performance of martt on the main organ level markup on each training set using sccp learning and markup algorithm is shown in table 1. the performance on each flora is displayed element by element with four columns: the number of examples (n), precision (p), recall (r), and fmeasure (f). note the “taxon” element shown in figure 1 was a result of a straightforward parsing of the text and was not involved in the learning process. blanks (i.e. no data) in table 1 were due to the variations in the descriptions, for example, fnct descriptions include discussions about the taxa, while fna and foc do not. the overall performance across all elements is a weighted average of recalls on n, indicating the percentage of correctly tagged segments. without any reconfiguration but relying solely on training examples, martt marked 94-98% of segments correctly on different collections (table 1). martt then used sccp and its learned rules to tag the entire collections of fna and foc to build the markup rule bank. finally, martt performance on fnct using the rule bank in different ways was compared with the performance without using the rule bank. these results are shown in table 2. table 2 shows the performance of martt on fnct with three different settings: the first was the normal training and learning process done by sccp, the second used the rule bank alone without sccp learning from the training examples, and the third used both—martt first queried the rule bank, if no good rule was returned, it used the rules sccp learned from the training examples. in other words, in this setting, the rule bank was used as the primary knowledge source while the training data was secondary. the results show higher precision scores when the rule bank alone is used, suggesting the rules learned from fna and foc are in general highly reliable and applicable on fnct. one exception here is the discussion element. this is due to the fact that fna and foc do not include any discussions in descriptions (see table 1, n column), so nothing about discussion can be learned from fna or foc. martt assumed that segments that did not belong to any other elements were discussion, resulting in a high recall (98%) yet a low precision (58%). the other exception is on phenology element. fna contains little information on phenology. in foc, all phenology elements start either with “fl.” for flowering time or “fr.” for fruiting time, while fnct uses normal english to describe when a plant gives flowers or fruits. thus the rules learned from foc do not apply to fnct. the lower recall scores (especially on flowers, only 0.34) are due to the limited coverage of the rules—which were learned from only two other floras (the published volumes only). overall, the rule bank alone tagged 69% of all segments from fnct correctly. the correct ratio of using training examples alone was 94%. when the rule bank and the training are combined, the overall performance is improved from 94% to 95%—the rule bank helped to correct 1/6 of the errors made by sccp. more interestingly, when martt used the training examples as the primary knowledge source and the rule bank secondary, the performance improvement was not that obvious, suggesting the rule bank was a more reliable source than the training examples, even though the rule bank was created from other collections. cui – converting taxonomic descriptions to new digital formats 26 table 1: martt performance in precision, recall, and f-measure on fna, foc, and fnct using sccp fnan p r f foc-n p r f fnctn p r f plant habit and life style 202 0.98 1.00 0.99 241 0.99 0.99 0.99 298 0.94 0 . 9 0 0.92 roots 28 1.00 0.94 0.97 30 0.95 0.90 0.92 6 0.83 0 . 7 2 0.77 buds 21 0.97 0.93 0.95 11 0.87 0.95 0.91 4 0.50 0 . 5 0 0.50 stems 230 0.92 0.98 0.95 278 0.92 0.97 0.94 111 0.92 0 . 9 1 0.91 leaves 296 0.99 0.98 0.98 343 0.97 0.98 0.98 270 0.93 0 . 9 4 0.93 flowers 198 1.00 0.99 0.99 345 0.99 0.99 0.99 307 0.94 0 . 9 4 0.94 fruit 192 0.98 0.96 0.97 233 0.98 0.96 0.97 178 0.94 0 . 8 8 0.91 cones 20 0.98 0.96 0.97 14 0.97 0.95 0.96 3 0.89 0 . 7 8 0.83 seeds 119 1.00 0.98 0.99 115 0.98 0.98 0.98 31 0.99 0 . 9 7 0.98 spore-related structures 68 0.97 0.96 0.96 7 0.57 0 . 5 0 0.53 gametophyte 19 1.00 0.96 0.98 chromosomes 191 0.97 0.89 0.93 53 1.00 1.00 1.00 3 1.00 1 . 0 0 1.00 phenology 269 1.00 1.00 1.00 234 0.97 0 . 9 8 0.97 discussion 638 0.95 0 . 9 7 0.96 total 1584 1932 2090 overall 0.97 0.98 0.94 further markup to the sub-organ level involves more than 240 elements. the element-by-element performance scores are shown in the appendix. in the appendix the hierarchical relations between elements are denoted by “/”. “phls/leaves” may seem strange, but this was how some descriptions had been written. martt made no attempt to rearrange original descriptions. the results suggest that at this markup granularity, there are more cases of other features element to accommodate sub-organs not covered by the xml schema. further, variations in element distributions across collections and within collections are more evident. the data also show that many elements have only one training example, which inevitably results in zero performance, because in a 10-fold cross-validation, the training example is either placed in the training set, leaving no test data, or in the test set, leaving no training data. excluding these elements, the overall markup performance at this level is 91% for fna, 94% for foc, and 89% for fnct (this figure drops to 87% if discussion element is excluded). the overall performance is 1% lower if these elements are counted. the calculation of the overall performance only involves the terminal elements and not their parent organ elements. we evaluated the reusability of the rule bank at sub-organ level markup as well, but limited the evaluation in stems, leaves, flowers, and fruit four elements since other main organ elements in fnct either do not have enough examples (e.g., roots, buds, cones, spore-related structures, and chromosomes), or do not have a counterpart in fna or foc (e.g., discussion), or do not have a good number of sub-elements (e.g., plant habit and life style, seeds, and phenology) to make the evaluation interesting (see the appendix). the results of the evaluation in stems, leaves, fruits, and flowers elements are shown in table 3-6 respectively. improved performance scores (compared to “training alone”) are highlighted in the tables. the results show that the sub-organ level markup in stems, leaves, and fruit elements benefits from the rule bank—using rule bank alone achieved about the same level of performance as that using hundreds of training examples. combining the rule bank and the training, the performance was further improved. however, for flowers element, the rule bank alone only marked 29% of the segments correctly. this is not entirely supervising because 1) the flower is the most complex organ of a flowering plant. 2) fnct contained descriptions of grass families and hence had specific sub-organs of grass cui – converting taxonomic descriptions to new digital formats 27 table 2: martt performance in precision, recall, and f-measure on fnct w / w/o the rule bank fnct training alone rule bank alone rule bank + training n p r f p r f p r f plant habit and life style 298 0.94 0.90 0.92 0.96 0.63 0.76 0.93 0.91 0.92 roots 6 0.83 0.72 0.77 0.83 0.89 0.86 0.83 0.89 0.86 buds 4 0.50 0.50 0.50 0.75 0.75 0.75 0.75 0.75 0.75 stems 111 0.92 0.91 0.91 0.94 0.88 0.91 0.89 0.97 0.93 leaves 270 0.93 0.94 0.93 0.98 0.84 0.90 0.94 0.95 0.94 flowers 307 0.94 0.94 0.94 0.99 0.34 0.51 0.96 0.92 0.94 fruit 178 0.94 0.88 0.91 0.98 0.83 0.90 0.93 0.90 0.92 cones 3 0.89 0.78 0.83 0.92 0.83 0.87 0.93 0.89 0.91 seeds 31 0.99 0.97 0.98 0.95 1.00 0.98 0.95 1.00 0.98 spore-related structures 7 0.57 0.50 0.53 0.86 0.74 0.79 0.86 0.74 0.79 chromosomes 3 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 phenology 234 0.97 0.98 0.97 0.00 0.00 0.00 0.97 0.98 0.98 discussion 638 0.95 0.97 0.96 0.58 0.98 0.73 0.95 0.97 0.96 total 2090 overall 0.94 0.69 0.95 flowers, such as pappus, ligule, glume, lemma, and palea etc, while fna and foc collections did not. 3) the recall on the flowers element was as low as 34% (see table 2). if a segment is not correctly identified as flowers, the further markup of its suborgans cannot be correct because of the hierarchical markup strategy. despite the overall low performance in flowers, the rule bank did help to improve the performance on some of its subelements (table 5). summary of martt experiments the experiments with martt show that the machine learning approach is highly portable: on all three floras martt achieved very good performance (in the range of 87% to 98%, depending on the markup granularity and data collection). biodiversity and other factors contribute to the rather skewed distributions of elements in description collections (see the appendix). martt fails at many elements with no or few training data. on the other hand, the results suggest that the induced knowledge (i.e. the rule bank) is reliable and reusable, in some circumstances, the rule bank provides more reliable rules than the training examples do. the rule bank is shown to help to improve the markup performance on elements with good coverage. continuing to enrich the rule bank with the markup rules learned from other description collections is likely to improve its coverage and make the rule bank more powerful. overall, the experiments showed that martt achieved its goal on portability and performance. using the learned rules, martt tagged all the 15,000 descriptions into xml format and turned them into three greenstone collections which can be searched by element 2 (witten et. al. 2000) is an open source digital library software which supports search in specified elements, such as in leaves element. if the collections are tagged according to one schema, like what we have done with foc, fna, and fnct, greenstone also supports crosscollection search. the experiments with martt and the three floras also identified a number of issues calling for further research, including the issues surrounding training examples, schema coverage, and performance variations. we shall discuss these issues and our current solutions in detail next. 2 http://research.sbs.arizona.edu/gs/cgi-bin/library.greenstone. cui – converting taxonomic descriptions to new digital formats 28 table 3: martt performance in precision, recall, and f-measure in stems with and without the rule bank. stems training alone rule bank alone rule bank+training n p r f p r f p r f stem-general 97 0.88 0.89 0.89 0.96 0.94 0.95 0.90 0.94 0.92 bark 3 0.67 0.67 0.67 0.33 0.33 0.33 0.33 0.33 0.33 node 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 culm 7 0.80 0.80 0.80 0.20 0.20 0.20 1.00 1.00 1.00 twig 2 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00 branch 2 0.00 0.00 0.00 0.67 1.00 0.80 0.67 1.00 0.80 branchlet 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 compound 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 overall 114 0.84 0.84 0.91 table 4: martt performance in precision, recall, and f-measure in leaves with and without the rule bank. leaves training alone rule bank alone rule bank+training n p r f p r f p r f leaf-general 206 0.92 0.96 0.94 0.97 0.96 0.97 0.95 0.97 0.96 petiole 18 0.72 0.72 0.72 0.89 0.83 0.86 0.83 0.83 0.83 stipule 10 1.00 0.94 0.97 0.00 0.00 0.00 1.00 0.94 0.97 sheath 9 0.79 0.71 0.75 0.00 0.00 0.00 0.79 0.79 0.79 leaf-blade 77 0.95 0.75 0.83 0.95 0.73 0.83 0.90 0.73 0.81 leaflet-general 32 0.91 0.80 0.85 1.00 0.93 0.96 0.97 0.94 0.95 spine 9 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00 tendril 3 1.00 1.00 1.00 0.33 0.33 0.33 1.00 1.00 1.00 ligule 11 0.71 0.79 0.75 0.00 0.00 0.00 0.79 0.86 0.82 gland 3 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 compound 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 overall 379 0.87 0.79 0.90 the training example issue the training example problem has two aspects: one has to do with the effort required to prepare training examples and the second is about the skewed distribution of training data in different elements. manually inserting tags in hundreds of documents is time-consuming and error-prone. to alleviate this problem, we developed a userfriendly utility that makes use of the rule bank induced from the foc, fna, and fnct collections to automate the training example preparation process. some screenshots of the interface are shown in figure 3. figure 3a shows a text description in the editing area. a click on the “mark up” button on the tool bar invokes martt to tag the description using the rule bank, which essentially tags every clause in the description as shown in figure 3b. in figure 3b, the hierarchy in the left pane displays the element structure of the tagged description. if a wrong tag is inserted by martt, the user can easily correct the error by bringing up the tag menu with a right-click on the mouse. the identified errors are saved automatically by the utility for further analyses. because of the shared domain knowledge across plant taxonomic descriptions, the rule bank can mark a large portion of a description with good tags, saving a significant amount of manual effort. cui – converting taxonomic descriptions to new digital formats 29 table 5: martt performance in precision, recall, and f-measure in fruits with and without the rule bank. fruits training alone rule bank alone rule bank+training n p r f p r f p r f fruit-general 176 0.94 0.94 0.94 0.98 0.94 0.96 0.97 0.97 0.97 infructescence-general 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 pedicel 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 mericarp 2 0.50 0.50 0.50 0.00 0.00 0.00 1.00 1.00 1.00 beak 4 0.00 0.00 0.00 0.33 0.33 0.33 1.00 1.00 1.00 wing 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 pappus 12 0.22 0.28 0.25 0.00 0.00 0.00 0.00 0.00 0.00 style 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 other-features 4 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 overall 202 0.84 0.82 0.87 for taxonomic descriptions that do not have a corresponding rule bank in martt (e.g., ant descriptions or alga descriptions), the utility has another feature to help with the manual markup as shown in figure 3c, where a selected text segment can be tagged with a tag chosen from the pop-up tag menu, which is populated from a specified xml schema (cui, 2005a). this interface ensures a tagged example is valid or at least well-formed. the second issue related to training examples has to do with the unbalanced distribution of elements. in description collections, due to the diversities in living organisms, authorship, and editorial policies, the coverage of different organs are quite uneven, resulting in a very skewed distribution of training data for individual elements: for example, in the 310 fna training examples, there were more than two hundred examples for “leaf blade” but zero for “tendril”. the training data distribution (see the appendix, n column) shows many sub-organs with zero or one examples. there were 42 elements from fna training examples, 34 from foc, and 20 from fnct with only one example, making learning impossible for sccp. this problem is somewhat alleviated by the induced knowledge from other collections (i.e. the rule bank), for example the markup rules learned from the several examples of “tendril” in foc and fnct training examples can be applied to fna descriptions. but we also investigated an unsupervised approach that would address this issue in a more direct manner, since no training examples are required at all. because this approach also helps to make rare organs more visible in the xml schema, we shall explain the unsupervised learning approach in detail in the next section. the schema coverage issue even though the xml schema (cui, 2005a) we created for the martt experiments was quite comprehensive to start with, there were occasions when we had to edit the schema to include new (sub)organs discovered from the training examples. we also had to use the other-features elements to accommodate any uncovered organs remaining in the collections (see the appendix for the occurrences of other-features elements). since a standard list covering all organs of living organisms does not exist, it is often difficult to enumerate in an xml schema all organs described in a sizeable collection. it is more difficult to create a comprehensive xml schema for multiple description collections. although it is not always necessary to formalize organ names at the schema level (e.g., sdd does not), from the application’s perspective, the need to tag all organs described in a collection and the need to search across collections basing on a common schema call for explicit declaration of all organ names. in absence of a comprehensive dictionary covering all organs, a simple way to discover them from collections of descriptions is needed in order to build a complete schema incrementally. in addition, the method can be used by martt to address the lack of training examples problem, because it can identify organ names without any training examples. cui – converting taxonomic descriptions to new digital formats 30 table 6: martt performance in precision, recall, and f-measure in flowers with and without the rule bank. flowers training alone rule bank alone rule bank+training n p r f p r f p r f inflorescence-general 187 0.84 0.82 0.83 1.00 0.13 0.23 0.89 0.65 0.75 bract 35 0.81 0.73 0.77 0.20 0.05 0.08 0.90 0.79 0.84 peduncle 4 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00 scape 3 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00 pedicel 16 0.89 0.92 0.90 0.00 0.00 0.00 0.89 0.92 0.90 rachis 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 rachilla 2 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00 branch 5 0.75 0.63 0.68 0.00 0.00 0.00 0.00 0.00 0.00 involucre 9 1.00 0.92 0.96 0.00 0.00 0.00 1.00 0.92 0.96 flower-general 132 0.86 0.90 0.88 0.76 0.91 0.83 0.73 0.94 0.82 perianth 24 0.81 0.82 0.81 0.10 0.05 0.07 0.90 0.88 0.89 corolla 96 0.95 0.93 0.94 0.30 0.05 0.08 0.98 0.93 0.95 corona 2 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 pappus 15 0.24 0.28 0.26 0.00 0.00 0.00 0.24 0.28 0.26 ligule 7 0.70 0.60 0.65 0.00 0.00 0.00 0.80 0.60 0.69 calyx 40 0.85 0.86 0.86 0.80 0.28 0.41 0.90 0.86 0.88 glume 11 1.00 0.86 0.92 0.00 0.00 0.00 1.00 0.86 0.92 lemma 24 0.88 0.93 0.91 0.00 0.00 0.00 0.88 0.93 0.91 palea 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 sepal 19 0.80 0.77 0.78 0.70 0.43 0.54 0.93 1.00 0.96 petal 59 0.94 0.93 0.94 0.90 0.46 0.61 0.97 0.94 0.96 tepal 2 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00 lip 2 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 hood 2 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 carpel 6 1.00 1.00 1.00 0.80 0.80 0.80 1.00 1.00 1.00 anther 9 1.00 0.83 0.91 1.00 0.92 0.96 1.00 0.92 0.96 style 15 0.96 1.00 0.98 0.14 0.14 0.14 0.96 1.00 0.98 stamen 38 0.97 0.98 0.98 1.00 0.72 0.84 0.97 0.98 0.98 pistil 6 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00 stigma 10 0.94 0.92 0.93 0.00 0.00 0.00 0.94 0.92 0.93 filament 6 0.60 0.50 0.55 0.40 0.40 0.40 0.40 0.40 0.40 ovary 12 1.00 0.86 0.93 0.00 0.00 0.00 1.00 0.86 0.93 placenta 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 receptacle 3 0.67 0.67 0.67 0.00 0.00 0.00 0.67 0.67 0.67 gynostegium 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 hypanthium 3 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00 keel 2 0.00 0.00 0.00 0.00 0.00 0.00 0.50 0.50 0.50 pollen 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 nectary 1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 gland 3 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 compound 10 0.67 0.50 0.57 0.00 0.00 0.00 0.67 0.50 0.57 other-features 10 0.29 0.19 0.23 0.00 0.00 0.00 0.29 0.19 0.23 overall 835 0.83 0.29 0.81 cui – converting taxonomic descriptions to new digital formats 31 to this end we developed a utility that simply take a collection of descriptions to generate a draft xml schema, which contains the names of the organs described in the collection. the utility employs an unsupervised machine learning algorithm that takes advantage of the formality in professionally prepared descriptions. in particular, we notice that. collectively, subjects of sentences in descriptions likely represent the complete set of organs described. the algorithm tries to separate subjects from remaining parts of sentences, and then collects organ names from the subjects and organ characters from the remaining parts for future markup at a finer granularity. being an unsupervised algorithm, this algorithm does not need any training examples. making use the organ names and the regularity in punctuation usage in the descriptions, the utility generates a raw but rather comprehensive xml schema that can be easily refined by a domain expert. here is how the unsupervised algorithm works on a collection: plain-text descriptions in the collection are segmented into sentences at full stops or semicolons. the algorithm makes the first three leading words of the sentences candidate subjects so no potential organ names is left out. next it finds nouns from the description collection in question by using the following heuristic rule: a word w is noun, iff the collection contains singular and plural forms of w, but no past, past participle, or present participle forms. seed nouns (nouns given to the algorithm are called seed nouns) may also be provided by the user directly or collected from a glossary. with the list of nouns, the algorithm marks the words in the candidate subjects as either noun or unknown. then, all the sentences in the collection are sorted according to the number of known nouns in their candidate subjects. next, the algorithm uses the following bootstrap procedure to infer the roles of the unknown words. the bootstrap procedure works in iterations and stops when no new discoveries are made in an iteration. new discoveries are used immediately in the next iteration to make other discoveries. a discovery is an identification of either a modifier – the word before a head noun (e.g. “basal” in “basal leaf”), a boundary word – the word following a head noun (e.g 2 in “cells 2”), or a noun. when the bootstrap procedure terminates, the algorithm uses the discovered modifiers, nouns, and boundary words to verify the candidate subjects: a verified subject is a noun with or without modifiers and is followed by a boundary word. if a subject can not be verified, the algorithm takes all the words up to the first known noun (inclusive) in the sentence as the subject. when the roles of the words in the subjects are known, it is straightforward to group different subjects to their head nouns, for example, “pistillate flowers” and “staminate flowers” are “flowers”. this in effect identifies an “is type of” relationship between the three concepts: pistillate flowers and staminate flowers are types of flowers. the relationship “is part of” may also be discovered by looking at the punctuation marks. many floras adopt the convention to “place each major structure in a separate sentence and separate subparts by semicolons” (fna editorial committee, 2006). this convention can be used to identify relationships such as sepals are a part of a flower. these relationships are integrated in the resultant raw schema, which is a good start for a domain expert to make refinements. the organ names and relationships will also be useful for a semantically richer ontology to be developed in the future. in addition, the subjects and their head nouns can be used as xml tags to tag the descriptions into well-formed xml documents. the well-formed xml documents may be imported to the training example preparation utility (figure 3(c)) to generate training examples for martt at a much reduced cost. martt may also directly use the tags to mark up elements with few training examples. hence the simple unsupervised learning algorithm addresses the schema coverage problem and the lack of training example problem at the same time. the bootstrap algorithm was tested, without being given any seed nouns, on three collections: one contained 120 algae descriptions extracted from feist, et. al, (2005), another contained 200 fna descriptions, and the third contained 2367 fna descriptions. table 7 shows the evaluation results. from the 538 sentences of the algae descriptions, the algorithm learned 37 good singular nouns (correct rate = 95%) and 13 good plural nouns (correct rate = 87%), and tagged 476 sentences correctly (correct rate = 88%). from the cui – converting taxonomic descriptions to new digital formats 32 (a) the composition area (b) oneclick markup and editing (c) manual markup figure 3. training example preparation and verification utility interface. cui – converting taxonomic descriptions to new digital formats 33 3195 sentences of the fna descriptions (labeled as fna-1 in the table), the algorithm learned 152 good singular nouns (correct rate = 99%) and 90 good plural nouns (correct rate = 100%), and tagged 3140 sentences correctly (correct rate = 98%). an example correct tag is “inner petals” while an incorrect one may be “petals generally” or “in some species base” (figure 4a). note the number of unique tags does not grow linearly with the size of description collections. this ensures that a visual display of the learned tags and their structure will not get overly crowded with larger collections. as table 7 shows, while the number of sentences in fna-1 is 6 times of that in algae collection, the number of tags learned from fna-1 is only 2 times of that from algae. to confirm this observation, a larger fna collection (labeled fna-2 in table 7) with 31387 sentences was processed and the result shows that, comparing fna-2 with fna-1, while the number of sentences increases 9-fold, the number of tags only increases 2-fold. the number of unique tags increases at a much lower rate than the number of sentences and is expected to reach a plateau. the diagram in figure 4 visualizes the resultant xml schema, including the discovered tags and their structural relationships. figure 4a and 4b shows the interactive diagrams generated from the algae and fna-1 descriptions respectively. the “is part of” relationships are displayed in the diagrams by connecting sub-organs to their parent organs. the visualization readily shows the organs and how consistently periods and semicolons were used in the text. fna descriptions often use periods and semicolons to set off major structure descriptions and subpart descriptions respectively, hence we see rather clearly the main organ elements such as leaves, inflorescences, flowers, fruits, and seeds as the first level elements and their subparts as the second level elements (figure 4b). in contrast, the algae descriptions do not follow the same convention in using periods and semicolons; instead, they use mostly semicolons to separate different descriptive segments. therefore in the diagram there is no clear-cut main organs level or subparts level (figure 4a). the diagram may be further explored; for example, when a tag is selected, the interface displays the original descriptions on which the tag is applied. a visual interface like this assists the human expert in refining the raw schema to make it fit the descriptions better. the performance variation issue the results from the martt experiments show that the system performed better on the fna and foc collections than on the fnct collection. performance differences were also seen among different elements, for example flowers elements were more difficult than others. what characteristics of data sets cause the performance difference? can these characteristics be measured and used to predict martt performance? a performance prediction model helps to answer questions such as “how well will this system work on this description collection?” instead of asking the user to prepare hundreds of training examples to test the system out, we developed a prototype utility that has the potential to predict the performance with just a few dozens of examples. at the center of the utility are two modules: one module measures characteristics of a set of examples, and the other uses the prediction model to make the prediction basing on the measurements. the prediction model was established and tested on fna, foc, and fnct descriptions using the following procedure: 1. a set of 11 corpus characteristic measures were derived. 2. 177 collections of description segments (5-56 files per collection) were created from fna, foc, and fnct training examples. 3. the characteristics of each collection were measured. 4. martt performance on these collections was evaluated. 5. statistical analyses were carried out to find correlations between the characteristic measures and system performance. steps 1 and 3: characteristic measurements we derived the following 11 corpus characteristic measures that can potentially have an impact on system performance. the statistical analyses carried out in step 5 will reveal the ones with statistically significant impact. cui – converting taxonomic descriptions to new digital formats 34 table 7: performance of the unsupervised algorithm on alga and two fna collections. alga fna-1 fna-2 descriptions 120 200 2367 sentences 538 3195 31387 sentences correctly tagged (%) 476(88) 3140(98) * unique tags 61 143 444 singular nouns learned 39 154 504 correct singular nouns(%) 37(95) 152(99) 490 plural nouns learned 15 90 297 correct plural nouns(%) 13(87) 90(100) 295 boundary words learned 44 317 932 correct boundary words 44 317 931 process time 1 minute 1 minute 15 minute 1. instance count is the number of examples (i.e. documents) in a collection. 2. class count is the number of unique terminal elements in a collection. 3-5. n-gram count (n ∈ [1,2,3]) is the number of unique n-grams in a collection. 6-8. n-gram distribution score (n ∈ [1,2,3]) gauges the distinctiveness of n-gram distributions in terminal elements in a collection. if an n-gram occurs m (m>1) times in a collection and all occurrences are in one terminal element e, then we say the distribution of the n-gram is very distinctive in that the presence of the ngram itself suggests the element. if all ngrams have such a distinctive distribution, the markup task would be trivial. at the other extreme, if the m occurrences are uniformly distributed in the elements, then the presence of the n-gram is of little value to the markup task. the final n-gram distribution score is the mean distinctiveness scores of all n-grams counted in a collection. the formula for the measure is || || 1|| || ||max . g g g g g scoredistri gig i i i i classes∑ ∈         − = where g = (g1, …, gn). in the formula, an ngram gi’s maximum occurrence in all terminal elements is divided by gi’s total occurrence in the collection. this simple division roughly measures the distinctiveness of gi’s distribution. the factor         − || 1|| i i g g is used to discount the effect of rare n-grams. the final score is obtained by taking the average over all ngrams to remove possible sample size effect. the score is a value between 0 and close to 1. 9. delimiter score measures the consistency of delimiters. here a delimiter is a textual pattern that separates a previous element from the current one and the current one from the next one. for example, a delimiter pattern “. /fruit berry/. /” indicates that following a period, a fruit type description starts with the words “fruit berry” and ends with another period. the delimiter score uses information entropy (ie) to measure the distribution of delimiting patterns in a collection. the lower the ie score, the more distinctive the distribution. if all examples in a collection shares one delimiting pattern, the markup task would be much easier than in a collection where each element has a unique pattern. the formula for this measure is iemax ie scoredelimiter dd iemax d d d d ie dd i dd i i i . 1 || 1 log || 1 . || || log || || −=       −=       −= ∑ ∑ ∈ ∈ cui – converting taxonomic descriptions to new digital formats 35 (a) visualization of the alga collection (b) visualization of the fna-1 collection figure 4. visualizations of learned tags and their structural relationships. cui – converting taxonomic descriptions to new digital formats 36 where d =(d1,…,dn). to find the delimiter score for a collection, the delimiters d1,….,dn of terminal elements are gathered. the standard ie is calculated using the occurrences of different patterns in the collection. the ie reaches its maximum when each pattern occurs only once. the maximum ie is used to make the delimiter score a positive measure of the distinctiveness of a distribution (i.e. the higher the score, the more distinctive the distribution). the score is a value between 0 and 1. 10. class order score and the next measure evaluate the consistency of the element sequences in a collection. class order score deals with the order of the terminal elements. an example of an order may be “inflorescence, sepal, petal, style” in a flower description. descriptions with some or all of these four terminal elements presented in that order are said to “fit” that sequence. consistent sequences are useful for a markup algorithm to make sound decisions on some otherwise difficult cases. similar to the delimiter score, the maximum ie is used here. this score is calculated as the following: iemax ie scoreorderclass ordersallofordersallof iemax examplestotal ofitexamplesof examplestotal ofitexamplesof ie ordersall i ooo i ni . 1 # 1 log # 1 . # log # ]...[ 1 −=       −=       −= ∑ ∑ ∈ to get the score for a collection, the sequences of terminal elements are collected and the examples fitting a sequence are counted. the maximum ie is calculated based on the number of all possible sequences, which is either the number of the total examples in the collection, or the number of all permutations of terminal elements, whichever is smaller. similar to the delimiter score, the class order score is a positive measure with a value between 0 and 1. 11. class presence score considers the presence/absence patterns of terminal elements regardless of their order. the score is calculated in a rather similar way as the class order score. for maximum ie, the number of all possible patterns is either the number of the total examples in the collection, or the number of all combinations of terminal elements, whichever is smaller. the class presence score is a positive measure with a value between 0 and 1. iemax ie scorepresenceclass patternsallofpatternsallof iemax examplestotal pwithexampleof examplestotal pwithexampleof ie patternspresentall i ppp i ni . 1 # 1 log # 1 . # log # ]...[ 1 −=       −=       −= ∑ ∑ ∈ step 2: creation of 177 collections the 177 collections were created using the following procedure. first, 1500 descriptions from the three floras (633 from fna, 492 from foc, and 378 from fnct) were manually marked-up in the xml format. the sample sizes were increased from those used in the martt experiments to generate enough collections for statistical analyses. these descriptions were then randomly divided into 30 sets of 50 descriptions. then each description was split into several parts, each of which contained a text segment describing a main organ (e.g. flowers, fruit, etc). from this point on, each part was treated as an individual document. the documents that were in the same set and contained the same main organ element formed a collection. of the resultant 200 collections, 23 collections had fewer than 5 documents and were removed because they were too small to measure martt performance using a 5-fold crossvalidation routine. each remaining collection consisted of 5 to 54 (mean = 23) documents. among the 177 remaining collections, 135 random collections were used in the statistical analyses to derive the performance prediction model, and the remaining 42 collections were reserved to test the prediction model. the collections produced provide a reasonable representation of the taxonomic description population, as the documents were drawn from three different sources. they also preserve the element distribution variations seen in the original descriptions. in the end, each document contained a 2-level, flat xml structure. this simple model cui – converting taxonomic descriptions to new digital formats 37 allowed us to focus on the effect of characteristic measures on system performance. the more involved multi-level hierarchical structures will be examined in the future. step 4: performance measurement instead of precision/recall, we used a single-valued cosine similarity-like measure to evaluate markup accuracy, which is essentially a normalized value characterizing the proportion of the words tagged correctly in a description. step 5: statistical analyses the spss linear regression analysis on 135 of the 177 collections between characteristic measurements (the independent variables) and system performance (the dependent variable) constructed the following model: )(*002.0 )(*372.0 )(*176.0725.0 countclass ondistributi unigram presenceclasseperformanc − + += this model explained 64% of the original variance in performance and the residual of the model is normally distributed, as figure 5 shown (the closer the plot of the residual to the diagonal line, the closer the distribution to the normal distribution), indicating the linear model is a good fit. the model shows that among the eleven characteristics measured, class presence, unigram distribution, and class count are the statistically significant factors for determining the performance score. the prediction model was tested on the reserved 42 of the 177 collections. the performances of martt on the 42 collections ranged from 60% to 100%. the differences between observed performance and predicted performance are plotted in figure 6, which shows the residual distribution is quite close to normal. the residual of over 50% of the cases is in a0.03 range (meaning the predicted value is 0.03 less or more than the observed value). this result seems very promising. the reader should keep in mind that the prediction model was derived basing on the data from the three floras. at this time, the coefficients should not be interpreted literally. we will continue to test and refine the prediction model with more data from other sources. literature review the majority of studies on structuring plaintext taxonomic descriptions have relied on handcrafted rules which make heavy use of formatting and textual cues. organism nomenclature, for example, conforms closely to prescribed rules and can be reliably extracted by software programs using a combination of contextual rules and a language lexicon (kirkup et.al., 2005; koning et al., 2005). sautter et al. (2006) built on top of koning et al.’s system a named entity recognition system for taxonomic names, using both hand-crafted rules and some learning components. fewer studies have focused on cue-poor yet semantic-rich sections (e.g. morphological descriptions) largely due to the lack of consistency in description contents. lydon et al. (2003)’s manual comparison revealed surprisingly large inter-collection variations among descriptions of the same species. earlier studies using syntactic parsing methods to extract information to populate relational databases or to mark up plant descriptions in xml have focused on a single collection (taylor, 1995; abascal et.al., 1999; vanel, 2004). recently, wood et. al.(2004) extracted plant features from the descriptions of five species found in six floras, using a hand-made gazetteer as a lookup list to link extracted terms with their tags. they also showed that features extracted from different sources were complementary to each other. the research reported in this paper involves multiple description collections and multiple user-friendly approaches, minimizing manual work as much as possible. goldengate (sautter et al., 2007) is an xml editor that facilitates the markup of plain-text taxonomic descriptions in xml. it works with complete documents and the user can invoke different functions to paginate documents and to tag taxonomic names and taxon names, in other words, to tag a document to taxonx level 1. taxonx is an xml schema that defines five levels of markup. the sentence level markup described in this paper is between taxonx level 2 and 3. goldengate relies on regular expressions and pre-compiled dictionaries to tag description text. this approach can be sensitive to text variations cui – converting taxonomic descriptions to new digital formats 38 0.0 0.2 0.4 0.6 0.8 1.0 observed cumulative probability 0.0 0.2 0.4 0.6 0.8 1.0 e x p e c t e d figure 5. normal distribution of the residual of the linear regression model. and is limited by the availability of the dictionaries and user skills in constructing regular expression patterns. goldengate supports manual editing of tagged text in a similar way as martt's training example preparation utility. others, such as cui et. al (2002), used text classification algorithms such as svms to mark up description paragraphs as nomenclature, description, distribution, discussion, and reference, etc. with good accuracy. few studies linked characteristic measures of text corpora to system performance statistically. an exception is bagga & biermann (1997) who proposed a measure called “fact level” to evaluate the complexity of a text corpus in the context of information extraction, basing on the observation that it is more difficult to extract a fact when its components are scattered around in the text. the study showed that higher fact levels are associated with lower performances in information extraction systems, indicating that fact level may be an appropriate measure for extraction difficulty. however, fact level is not applicable to the semantic markup scenario discussed here. conclusions and future work our experience with taxonomic descriptions confirms lydon et.al (2003)'s conclusion that large variations exist among collections of descriptions. domain practices (e.g., use of punctuation marks) are not adopted uniformly across collections. these variations demand any automated semantic markup systems to enhance not only its accuracy but also its portability. the uniqueness of martt lies in its ability to store and reuse markup rules learned over time from different description collections. this makes it highly portable across collections as demonstrated in the experiments with fna, foc, and fnct. because the learned markup rules are collection-independent, we hope that these rules accumulated over time will be also useful for tagging more free-style text related to taxonomy. as a machine learning system, martt compares candidate markup rules learned from training examples to select the rule with the lowest expected error rate and the highest expected correct rate. this feature releases the user from the difficult and time consuming task of crafting markup rules. to make the system more efficient and user-friendly, a number of utilities are also being developed. the training example preparation utility can significantly reduce the cost of training examples. the unsupervised learning utility identifies main concepts (organ names) from a description collection without any training example and helps the user to create a more comprehensive xml schema and training examples at low cost. lastly, the performance prediction utility shows the potential of predicting martt performance on a collection with only a few dozens of tagged examples. -0.10 -0.05 0.00 0.05 0.10 0.15 unstandardized residual 0 2 4 6 8 10 12 14 f re q u e n c y n= 42 mean = -.0025 median = -.0087 std. dev. = .04459 m inimum = -.09 maximum = .11 percentiles 25 = -.0326 50 = -.0087 75 = .0311 figure 6. the difference between observed performance and predicted performance on 42 test collections. cui – converting taxonomic descriptions to new digital formats 39 in the course of developing the martt system and its utilities, we essentially have tested two machine learning approaches: the first is a supervised learning approach where an xml schema and a set of training examples guide the markup decisions of the algorithm; the second is an unsupervised learning approach where the algorithm exploits implicit regularities in the description text without any training examples or a schema. although both approaches are capable of producing well-formed xml documents from plain-text taxonomic descriptions, the latter is more efficient but the former integrates more domain knowledge. for example, “is part of” and “is type of” relationships are more accurately represented in the supervised approach. it is important to note, however, the two approaches are mutually beneficial in that the unsupervised approach helps to create a comprehensive xml schema and training examples that the supervised approach needs, while the schema and the rules learned by the supervised approach can help to improve the performance of unsupervised approach (e.g., by providing good seed nouns). while the markup at the sentence level can benefit information retrieval by supporting fielded searches, in the immediate future we will further develop martt to tag at an even finer granularity; that is, to tag characters and character states in descriptions. the character level markups will prove more useful and powerful: they can be used to support database-like queries, to merge descriptions from multiple collections, to generate taxonomic keys either in a semi-automated or automated manner, and to compare descriptions along multiple dimensions, to name just a few possibilities. we will format the tagged description in standard formats such as sdd to share them with the community. sdd does not prescribe a standard set of characters to be included, but leaves the decision to individual applications. to ensure our sdd documents interoperate with others, a conceptual model (i.e., ontoglogy) with broad coverage is indispensable. we will look into the issues on ontology construction and how to use the ontology to guide the markup practice. as well, we will conduct further evaluation of the entire system from a more user-centered perspective. we will examine in a systematic manner the effort required on the user’s side to mark up a sizeable collection using martt and its utilities. to provide a comprehensive and useful evaluation, the author is more than willing to collaborate with contributors and rights-holders of any taxonomic collection. acknowledgments this work was in part supported by an internal research grant of faculty of information and media studies, university of western ontario. the author thanks the editorial committees and authors of the taxonomic works for the permissions to use their text in this project. the author thanks dr. richard mccourt, dr. p. bryan heidorn, dr. linda smith, and others for their constructive comments on the design and evaluation of martt and its utilities. the author also thanks the bi editor and reviewers for their valuable paper revision suggestions. references abascal, r. & sánchez. 1999. x-tract: structure extraction from botanical textual descriptions. proceedings of the string processing & information retrieval symposium and international workshop on groupware, spire/criwg, pp. 2-7. bagga, a., & biermann a.w. 1997. analyzing the complexity of a domain with respect to an information extraction task. proceedings of the tenth international conference on research on computational linguistics (rocling x), pp. 175194. bhl 2007. biodiversity heritage library. accessed 10 july 2007 from http://www.bhl.si.edu/. cui, h. 2005a. the xml schema for martt, accessed 10 july 2007 from http://publish.uwo.ca/~hcui7/research/xmlschema.xsd. cui, h. 2005b. martt: using knowledge based approach to automatically mark up plant taxonomic descriptions with xml. proceedings of the annual meeting of american association of information and technology. oct 28-nov 2. 2005 charlotte, north carolina, usa. cui, h. 2005c. automating semantic markup of semistructured text via an induced knowledge base: a case-study using floras. doctoral dissertation. the university of illinois at urbana-champaign. cui, h., heidorn, p.b., & zhang, h. 2002. an approach to automatic classification for information retrieval. proceedings of the joint conference of digital libraries 2002, 96-97. cui, h., & heidorn, p.b. 2007. the reusability of induced knowledge for the automatic semantic markup of taxonomic descriptions. j. am. soc. inf. sci. technol. 58:133-149. cui – converting taxonomic descriptions to new digital formats 40 dallwitz, m.j. 1980. a general system for coding taxonomic descriptions. taxon 29:41-46 diggs, g.m, lipscomb, b.l., & o’kennon r.j. 1999. shinners & mahler’s illustrated flora of north central texas. center for environmental studies and department of biology, austin college, sherman, texas, and botanical research institute of texas (brit), fort worth, texas. feist, m., crambast-fessard, n., guerlesquin, m., karol, k., huinan, l., & mccourt, r. m. et al. 2005. treaties on invertebrate paleontology, part b: protoctista 1 volume 1: charophyta. boulder, colorado: geological society of america, inc. & lawrence, kansas: university of kansas. fna editorial committee 2006. flora of north america north of mexico guide for contributors. accessed july 10, 2007 from http://www.fna.org/fna/guide/guide_2006.pdf. fna flora of north america editorial committee. (eds.). 1993 onwards. flora of north america north of mexico. accessed july 10, 2007 from http://www.fna.org/. foc flora of china editorial committee. (eds.). 1994 onwards. flora of china. beijing/st. louis: science press/missouri botanical garden press. accessed july 10, 2007 from http://flora.huh.harvard.edu/china/. gbif. 2007. global biodiversity information facility accessed july 10, 2007 from http://www.gbif.org/. han, j. & kamber, m. 2000. data mining: concepts and techniques. morgan kaufmann publishers. kirkup, d., malcolm, p., christian, g., & paton, a. 2005. towards a digital african flora. taxon 54:457-466. koning, d., sarkar, i.n., & moritz, t 2005. taxongrab: extracting taxonomic names from text. biodiv. inf. 2:79-82. lydon, s.j., wood, m. m., huxley, r., & sutton, d. 2003. data patterns in multiple botanical descriptions: implications for automatic processing of legacy data. syst. biodiv. 1:151-157. mccallum, a.k. 1996. bow: a toolkit for statistical language modeling, text retrieval, classification and clustering. accessed april 20, 2003. http://www.cs.cmu.edu/~mccallum/bow. mitchell, t. 1997. machine learning. mcgraw hill: new york, ny. radford, a.e. 1986. fundamentals of plant systematics. harper & row, publishers, inc.: new york, ny. sautter, g., agosti, d., & böhm, k. 2007. semiautomated xml markup of biosystematics legacy literature with the goldengate editor, proceedings of psb 2007, wailea, hi, usa. accessed july 10, 2007 from http://psb.stanford.edu/psbonline/proceedings/psb07/sautter.pdf. sautter, g., agosti, d., & böhm, k. 2006. a combining approach to find all taxon names (fat). biodiv. inf. 3:46-58. taylor, a.1995. extracting knowledge from biological descriptions. proceedings of 2nd international conference on building and sharing very largescale knowledge bases. pp114-119, enschede, the netherlands, april. vanel, j.-m. 2004. worldwide botanical knowledge base. accessed july 5, 2007 from http://wwbota.free.fr/. witten, i.h., mcnab, r.j., boddie, s.j. & bainbridge, d. 2000. greenstone: a comprehensive opensource digital library software system. proceedings of digital libraries 2000, pp. 113-121, san antonio, texas, june. wood, m., lydon, s., tablan, v., maynard, d. & cunningham, h. 2004. populating a database from parallel texts using ontology-based information extraction. proceedings of natural language processing and information systems, 9th international conference on applications of natural languages to information systems. pp.254264, salford, uk, june. biodiversity informatics, 8, 2013, pp. 185-197 discovery and publishing of primary biodiversity data associated with multimedia resources: the audubon core strategies and approaches robert a. morris(1)*, vijay barve(2), mihail carausu(3), vishwas chavan(4)*, josé cuadra(4), chris freeland(5), gregor hagedorn(6)*, patrick leary(7), dimitry mozzherin(7), annette olson(8), gregory riccardi(9), ivan teage(10), and greg whitbread(11) (1) university of massachusetts at boston, ma, usa, email: ram@cs.umb.edu (2) foundation for revitalisation of local health traditions, bangalore, india (3) danish biodiversity information facility (danbif), copenhagen, denmark (4) global biodiversity information facility secretariat, universitetsparken 15, dk 2100, copenhagen, denmark, email: vchavan@gbif.org (5) washington university in st. louis, usa (6) museum für naturkunde berlin, germany, email: g.m.hagedorn@gmail.com (7) encyclopedia of life, woods hole, ma, usa (8) us geological survey, reston, va, usa (under contract via information international associates) (9) florida state university, tallahassee, usa (10) natural history museum, united kingdom (11) australian national botanical garden, australia *corresponding authors abstract.—the audubon core multimedia resource metadata schema (simply “audubon core” or “ac”) is a representation-free vocabulary for the description of biodiversity multimedia resources and collections, now in the final stages as a proposed standard under tdwg biodiversity information standards. by defining only four terms as mandatory, it seeks to lighten the burden for providing or using multimedia useful for biodiversity science. at the same time it offers rich optional metadata terms that can help curators of multimedia collections provide authoritative media that document species occurrence, ecosystems, identification tools, ontologies, and many other kinds of biodiversity documents or data. about half of the vocabulary is re-used from other relevant controlled vocabularies that are often already in use for multimedia metadata, thereby reducing the mapping burden on existing repositories. a central design goal is to allow consuming applications to have a high likelihood of discovering suitable resources, reducing the human examination effort that might be required to decide if the resource is fit for the purpose of the application. introduction discovery and access to primary biodiversity data, as defined by the global biodiversity information facility (gbif, 2007) are critical components in ensuring informed decision-making on the sustainable use of biological resources and on the conservation of biodiversity at all levels. with an increasing need for a high volume of credible, quality data for research, instruction, and decision support, biodiversity information systems and networks must mobilize primary data associated with non-traditional sources including multimedia resources and their metadata. multimedia resources are digital or physical artifacts that normally comprise more than text. these include photographs, artwork, drawings, sound, video, animations, and presentation materials, as well as interactive online media such as species identification tools. a multimedia collection is an assemblage of such objects, whether curated or not and whether digitally accessible or not. collections are included under the umbrella of resources, though they sometimes mailto:ram@cs.umb.edu mailto:vchavan@gbif.org mailto:g.m.hagedorn@gmail.com morris, et al. – the audubon core strategies and approaches 186 need different kinds of treatment. multi-media resources can provide reliable evidence for the occurrence of a taxon in a particular place and time, and there is a growing recognition that a biodiversity-related multimedia object could be used as a ‘primary biodiversity record’ if the metadata associated with the object is available and of high quality. as such, mobilizing such metadata for network access is an extension of one of gbif’s central activities, the marshaling of occurrence data from its data contributors. metadata on multi-media resources, and those resources themselves, can also enhance other biodiversity informatics applications such as species and specimen descriptions, glossaries, and image processing. because the potential quantity and quality of biodiversity multimedia resources are at least as great as that represented by observational data and have widespread potential uses, multimedia data merit special consideration. as depicted in figure 1, applications exploiting a wide range of digital and physical biodiversity objects sometimes require the use of multimedia resources to document the objects. there is vast potential to channel the heterogeneous and distributed biodiversity-related multimedia resources through data publishers and partners. unlike observation or specimen data, however, the network loads and figure 1: relationships of multimedia resources to primary types of biodiversity resources, including some well-known example systems. the figure is adapted from tdwg ncd interest group (2009). morris, et al. – the audubon core strategies and approaches 187 latency for serving or acquiring multimedia resources may be so high that resource producers and consumers alike need mechanisms to determine the fitness-for-use of media upon discovery, before the media are fetched. to meet these and other goals described below, we describe the audubon core multimedia resource metadata schema proposal, now in the final stages of approval under the mechanisms of biodiversity information standards (tdwg), http://www.tdwg.org. gbif multimedia resources task group recognizing the need for primary biodiversity data and information to extend beyond its current focus of specimenand observation-based data records, in march 2008 gbif asked members of the tdwg image working group, and others whose work is related to images, audio, and video, to serve in the multimedia resources task group (mrtg), in order to suggest strategies to expand the types of primary biodiversity data that the gbif network can discover and publish through the mobilization of multimedia resources (gbif, 2008). mrtg was specifically asked to provide recommendations on (a) criteria for multimedia data sharing infrastructures, (b) best practices for multimedia resources metadata exchange/sharing, (c) estimation of the scale of multimedia resources in biodiversity, (d) metadata schema(s) for multimedia data management, and data exchange and/ sharing, (e) whether existing protocols for biodiversity data publishing services, such as digir, tapir, or biocase will need to be altered, or new tools developed to handle these data types, (f) ways to encourage potential data providers to participate in the gbif network for discovery of and access to multimedia resources, (g) ways to increase involvement of industry leaders, and (h) use of gps-enabled mobile devices and other recording tools. the gbif mrtg survey the mrtg conducted an online survey of multimedia resources in may 2008, with the objective to understand the extent of potentially useful, sharable biodiversity multimedia resources and the repositories that hold them. the survey revealed that a large quantity of biodiversity related multimedia objects are held in repositories with definite metadata recorded (such as scientific names and geo-references) indicating a huge potential for such resources to carry scientifically useful data. many of the reported resources are managed at general-purpose repositories like flickr (http://www.flickr.com) and picasaweb (http://picasaweb.google.com), and special purpose biodiversity image repositories such morphbank (http://www.morphbank.net), wildscreen (http://www.wildscreen.org.uk/).their diversity highlighted the need for an infrastructure that can (1) leverage such collections for scientific analysis and (2) assist in the better management of these vast biotic resources. the survey further highlighted the need for annotation and attribution services to enhance the usability of objects and to recognize the efforts towards mobilization of such resources. the gbif mrtg recommendations mrtg dealt with both social and technical issues related to the discovery, mobilization, and use of biodiversity-related multimedia resources. the principal recommendation of mrtg was that gbif should facilitate the discovery and publishing of multimedia resources as primary biodiversity data (morris et al., 2008). in particular, as a global information infrastructure, gbif must reduce burdens on its stakeholders as a strategy for increasing access to high-quality resources. recommendations about social issues called on gbif to (1) recognize the breadth and depth of information technology resources available to publishers of biodiversity media, (2) facilitate the publication of metadata with tools and training, (3) encourage free and open access and use of metadata, while increasing the ability to license resources, (4) support discovery and access of, at a minimum, thumbnails or other preview representation of resources, (5) encourage cultural change towards routine georeferencing of multimedia resources, and (6) encourage creation of national, regional, and thematic multimedia repositories across the gbif network. recommendations on technical issues focused on the development of georeferencing, annotation, and attribution services. morris et al. (2008) listed 28 recommendations about social and technical issues with a rationale and with the possible http://www.tdwg.org/ http://www.flickr.com/ http://picasaweb.google.com/ http://www.morphbank.net/ http://www.wildscreen.org.uk/ morris, et al. – the audubon core strategies and approaches 188 burdens they may impose on gbif or multimedia metadata publishers. the report further concluded that social and technical issues hampered the progress toward facilitating efficient discovery, publishing, and the use of biodiversity related multimedia objects or collections. many valuable multimedia resources exist that have no documentation information stored in databases. some may have a web presence and others not. even those available on line may not be adequately discovered by search engines, or may be lost in the noise of images, audio and videos from unreliable sources. a brief descriptive record can act as the ‘business card’ for researchers, aggregators, decision makers, educators, or the general public to discover these resources. the development of a multimedia metadata schema for easy discovery, publishing and use of biodiversity-related multimedia resources was deemed helpful to address these issues. gbif mrtg recommendation (morris et.al., 2008) facilitation through the audubon core r#3: metadata about media resources is provided either without any restriction on its use or reproduction, or under a suitable open-content license. provide for copyright attribution and terms of use. r#4: publishers will be able to license their resources. specific terms can reference various versions of a multimedia resource including license and other attributes. r#5: gbif metadata and data sharing agreements should give the gbif network the right to cache and display previews (e. g. thumbnails) if publisher grants the access. metadata identifies such resources. r#12: metadata should promote the ability of users of gbif services to determine fitness-for-use without requiring the users to acquire underlying resources. ability to signal biologically relevant content metadata, such as taxonomic and geographic coverage. r#13: ability to treat resource collections and objects uniformly. both resource collections and objects are described through a single schema. r#14: controlled vocabularies for metadata values should be encouraged and supported technically. specific values are suggested or required, particularly where arising from other vocabularies. r#15: specify that the copyright owner or available licenses are unknown when this is the case. 'unknown' is an accepted value for the terms specifying these. morris, et al. – the audubon core strategies and approaches 189 table 1: recommendations of the mrtg met through the audubon core. “r#” designates the recommendation addressed (morris et al., 2008, and morris et al., 2009). table 1 provides a list of 14 of the 28 issues from morris et al. (2008) that mrtg sought to address through the development of the multimedia resources metadata schema (morris et.al., 2009), now designated as the audubon core multimedia resources metadata schema (“audubon core”). development of the audubon core a subset of mrtg began development of the audubon core in august 2008, and a slightly different subset continued in september 2009, developing the key terms for a new metadata schema for multimedia resources. development of the schema included the participation of key stakeholders such as gbif, key to nature (http://www.keytonature.eu), the u s geological survey, morphbank (http://www.morphbank.net), and the encyclopedia of life (http://eol.org), as well as expressions of interest and inputs from the biodiversity heritage library (http://biodiversitylibrary.org), the university of massachusetts at boston electronic field guide r#16: support the identification of resources with publisherdefined guid schemes in resource or collection level metadata. an identifier is required for collections (strongly recommended for media), but the scheme for such identifiers is up to the provider, or to implementers of the representationneutral form of the specification. r#17: support the ability to specify relations among described objects. a generic relation 'relatedresourceid' is provided with no specified semantics. a small number of relations are provided for provenance, and a few for relations between different renderings of the same resource. r#18: services for georeferencing and scientific name recognition. all the georeferencing predicates of the darwin core are accepted by inclusion. a collection of terms designated as the 'taxonomic coverage vocabulary’ supports use of several darwin core nomenclatural predicates. r#19: allow support for the ‘documents’ relation, which asserts that a multimedia object provides evidence for an assertion that something else (e. g. an observation) is a primary biodiversity datum. subsets of the terms facilitate this, including the taxonomic, geographic, and temporal coverage vocabularies. r#20: lightweight metadata schema by combining existing schemata. accomplished by use of existing namespaces from other vocabularies where semantically reasonable. r#21: ability to specify media formats. service access points for different formats can be separately specified. r#22: allow specification of media manipulation by the publisher after acquisition. service access points for variants are supported, along with limited terminology for provenance description. http://www.keytonature.eu/ http://www.morphbank.net/ http://eol.org/ http://biodiversitylibrary.org/ morris, et al. – the audubon core strategies and approaches 190 project (http://efg.cs.umb.edu) , and the atlas of living australia (http://www.ala.org.au/). further work has been conducted on the schema since that meeting, and in february 2010, still known as the “mrtg schema”, version 0.9 was submitted for internal review to the biodiversity information standards (tdwg). as the schema progressed to the final stages of approval by the standards body, the name “audubon core” was proposed for the schema in honor of the great natural history illustrator, john james audubon. in november 2010, v1.0 the schema was submitted to the tdwg executive committee (ec). the submission included the response to an internal review and the proposal to officially name it “audubon core”. a second review was completed and substantial changes made based on it. responses to these and two more reviews have been completed and addressed, with further detailed changes. based on those, the executive committee permitted a period of public review as required by the tdwg rules, which is now complete. responses to that review will be submitted to the ec, including any changes arising from the response, with a request to accept the audubon core as a tdwg standard as may be revised based on the responses to the public review. several projects have been exploring the use of ac for their image management metadata in the form proposed for public review. of these, the most central to gbif's goals is a draft produced by the idigbio (2011) project of an audubon core ipt darwin core extension 1 now under testing. ipt denotes the gbif integrated publishing toolkit (gbif 2011), the recommended tool for publishing biodiversity data for harvesting by gbif and exposure through its portal. an ipt extension is an xml file that allows ipt to drive user interfaces and map the data publisher's data to easily harvested data using the domain vocabulary, in this case a subset of audubon core. a recently commissioned indo-norwegian ipbes capacity building pilot project aims to implement audubon core based dataflow to collate and publish the camera trap data through the gbif network. it is planned to use ms excel based ‘ac data 1 available at this writing at http://rs.gbif.org/sandbox/extension/audubon.xml templates’ to collate the multimedia data captured through camera traps in key protected areas. capabilities it is expected that audubon core will facilitate (1) the enhanced discovery of multimedia resources, (2) the evaluation of fitness-for-use prior to fetching a resource, (3) the use of metadata records as potential taxon occurrence evidence, or for other biological inferences such as evidence for species interactions, habitats, and phenotypic variations, (4) identification aids, and (5) the ability of multimedia resource producers and publishers to gather and serve resources contributed by a wide variety of producers and custodians, particularly those with little or no information technology expertise or support. the audubon core facilitates the above by describing with consistent metadata either media resources themselves or a collection of them. other existing standards present very little opportunity to provide media resource metadata that are specifically biologically relevant. for instance, although it can describe multimedia, the use of dublin core (dc); http://dublincore.org/) alone would not ease the discovery of media resources that require precision with respect to geolocation and identification. similarly, darwin core (dwc; http://rs.tdwg.org/dwc/ supports some biological semantics (e. g. taxonomy) but offers little about important intellectual property rights issues, or ways to express relations between alternate versions of media resources (e. g. services for different pixel resolution). the natural collections description (tdwg ncd interest group, 2009) provides useful metadata on object collections, but is missing some aspects relevant to biological media collections. metadata compliant with technical schemes, such as exif (http://www.exif.org/specifications.html), are frequently embedded directly in the media files by the imaging systems themselves. such embedded data often can be managed by tools such as adobe photoshop™ and the gimp open source image editor (http://www.gimp.org/). however technical metadata typically describe only the acquisition parameters of the media (e.g. pixel size, exposure data, etc.). they present little opportunity to embed biologically relevant information. furthermore, the combination of all of these standards still does not, http://efg.cs.umb.edu/ http://www.ala.org.au/ http://rs.gbif.org/sandbox/extension/audubon.xml http://dublincore.org/ http://rs.tdwg.org/dwc/ http://www.exif.org/specifications.html http://www.gimp.org/ morris, et al. – the audubon core strategies and approaches 191 or does only in a limited fashion; address the concerns of a wide variety of multimedia contributors, especially those with limited access to software engineers and digital librarians. among such concerns are various aspects of multimedia object provenance, intellectual property rights and attribution, access services, and the impact on service quality of large multimedia resources. below we discuss four examples: transfer cost, discovery of fitness-for-use, intellectual property rights, and provenance. transfer cost: individual digital multimedia resources such as images, video and sound may have very large file sizes. as a result, multimedia metadata must support use cases where humans or software agents fetch the resource in a reduced size (e.g., for images, small thumbnails or screen-sized resolutions). the management of multiple access points returning the resource in different forms and resolutions is therefore essential. fitness-for-use discovery. without specific examination of possibly many thousands of images, it can be difficult to determine whether a media resource carries all the biological context and technical properties required for the intended use. for example, it may be difficult to determine whether the resource depicts an organism in its natural habitat, a specific behavior, or particular morphological characters. furthermore, the resolution of an image may be too low, or it may contain labeling in an unsuitable language. the audubon core combines metadata terms representing these things (as well as several others from other widely used vocabularies) into one schema. it does so in a standardized way that makes it unambiguous what is being described and how it is made available by the provider. intellectual property rights: ownership of physical objects (e.g., specimens) is generally governed by property laws, while text and media resources are often subject to intellectual property rights (ipr). however, factual descriptions of objects are usually not subject to ipr (agosti and egloff, 2009). although similar considerations may apply to factual media representations of organisms, media have a history of being treated as creative works of art, not as expressions of facts of nature. consequently, the audubon core provides attributes to describe ipr, including ownership and license restrictions (such as reproduction permission and attribution requirements). provenance: for any scientific data, it is important to know the methodology used as well as how and when the data may have been changed from its original gathering. this is particularly important for media, which are commonly edited for a variety of purposes. if carelessly done, this may destroy some of the modified object's utility, or provide false impressions of data and thus influence research results. no current or proposed tdwg standard provides much provenance information, in part because widely accepted standards for specimen provenance and governance already exist. however, the creative aspects of media resources result in conflicting goals. the audubon core records object derivation (one media item is the source of another) and introduces a term called resource creation technique, for information about the technical aspects of the creation, digitization, and postprocessing (like background blurring, background elimination, color adjustment, etc.). relation to other standards a number of organizations concerned with addressing biodiversity multi-media in particular have informally or formally published specification for describing their resources. representatives of, or consultants to, several of these organizations are among the authors of this paper and architects of the audubon core. much of those organizations’ published metadata terminology has in one way or another been folded into ac (see http://terms.gbif.org/wiki/audubon_core_term_li st_(1.0_normative)#references.) most of the more general well-known multi-media metadata vocabularies focus on technical metadata of the image acquisition, or on curatorial, provenance, and intellectual property attributes. (niso 2008, iptc 2010, dcmi 2011, xmp 2010). they have limited expressivity about content, but we adopt their terms where we can. two crowd sourcing biodiversity media collections are worth mentioning, in part because they illustrate some of the problems of insufficiently formal or too dynamic metadata. the first of these, wikispecies, documents its image http://terms.gbif.org/wiki/audubon_core_term_list_(1.0_normative)#references http://terms.gbif.org/wiki/audubon_core_term_list_(1.0_normative)#references morris, et al. – the audubon core strategies and approaches 192 requirements at http://species.wikimedia.org/wiki/help:image_gui delines. most of the guidance is dedicated to licensing (wikispecies requires open access to material on its pages) and layout. however, wikispecies images are actually uploaded to the wikimedia commons (http://commons.wikimedia.org/wiki/main_page). as is generally the case for images supported in the wikimedia commons, image metadata per-se is limited to three sorts, mostly optional: a text caption, some specific image provenance and licensing text and the assignment of new or existing mediawiki “categories”. the last of these can be considered as lightly structured attributes (or rather “classes”) of the images, but at this writing, the overwhelming fraction of those are the names of geographically constrained taxa, e.g. (“australia arthropoda”). all that said, images on wikispecies are associated with a taxon page, and that has somewhat more information about the taxon, principally its taxonomy and nomenclature. the fact that contributors to wikispecies can add mediawiki categories at will could hold some promise for its contributor community to provide more organization to the website in ways that would provide more metadata to the embedded images. however, mediawiki categories are a typing mechanism and do not provide simple ways to place attributes of objects on wiki pages (as evidenced by the 330 categories of geographically constrained taxa such as mentioned above, and which reference fewer than 20 georegions.) wikispecies could be augmented by the semantic mediawiki extensions (http://semanticmediawiki.org/). note that the design of the audubon core puts emphasis on attributes rather than categories. ac only models as a class the access mechanism for retrieving media, because such mechanisms are highly variable and with many attributes. finally, we note that all mediawiki installations provide a permanent url for each version of a page. by the association of the image with a page version, this “permalink” can serve as a globally unique, persistent, dereferenceable uri for the image. a second crowdsourced biodiversity image repository may be seen in the encyclopedia of life image flickr group (http://www.flickr.com/groups/encyclopedia_of_lif e/) with metadata provided by a small set of flickr “machine tags” (http://www.flickr.com/groups/encyclopedia_of_lif e/discuss/72157612488733900). these are limited to taxonomy, georeference, and licensing information, but ownership, license metadata, and some technical metadata is available by flickr apis (http://www.flickr.com/services/api/). about 88,000 images are served this way by flickr, of which about 78,000 are harvested and associated with eol pages. eol itself offers similar metadata for all of its images (http://wiki.eol.org/display/dev/data_objects) the documentation supporting the submission to tdwg for ratification includes a normative specification of the audubon core as a set of multimedia resource metadata terms independent of any digital representation (http://terms.gbif.org/wiki/audubon_core_term_ list_(1.0_normative) ). that document will be updated to reflect any changes accepted for the standard after the period of public comment. the normative document provides metadata specifications describing biodiversity-related multimedia resources or collections. while focused on biodiversity-related multimedia resources, the audubon core addresses some of the same concepts as the dublin core, darwin core and other standards that describe access to resources. these standards include the adobe extensible metadata platform (xmp 2010), the international press and telecommunications council (iptc 2010) the metadata working group (mwg 2010) schema, the tdwg natural collections descriptions (tdwg-ncd 2009) schema, and others. where a particular term meets the same need met by the terminology within another standard, mrtg adopted that standard’s globally unique identifiers and definitions. where this is unsuitable, mrtg defined new. the design intends to ease the burden of holders using descriptions already specified either by dwc or dc, to allow them to use existing descriptions where appropriate. in other words, much of the audubon core may be viewed as a standard profile that defines the best practice use of certain terms from other metadata vocabularies, and provides further vocabulary for metadata that improves the http://species.wikimedia.org/wiki/help:image_guidelines http://species.wikimedia.org/wiki/help:image_guidelines http://commons.wikimedia.org/wiki/main_page http://semantic-mediawiki.org/ http://semantic-mediawiki.org/ http://www.flickr.com/groups/encyclopedia_of_life/ http://www.flickr.com/groups/encyclopedia_of_life/ http://www.flickr.com/groups/encyclopedia_of_life/discuss/72157612488733900 http://www.flickr.com/groups/encyclopedia_of_life/discuss/72157612488733900 http://www.flickr.com/services/api/ http://wiki.eol.org/display/dev/data_objects http://terms.gbif.org/wiki/audubon_core_term_list_(1.0_normative) http://terms.gbif.org/wiki/audubon_core_term_list_(1.0_normative) morris, et al. – the audubon core strategies and approaches 193 ability to utilize multimedia resources for scientific research. audubon core records an audubon core metadata record is a set of terms and term values that describe an underlying multimedia resource. each term is identified by a uniform resource identifier (uri). each uri refers to the attribute, not the underlying resource; it simply specifies which term is being provided. there are many uri schemes, some of which have been registered with the internet assigned names authority (iana). all audubon core uris conform to the widely used http uri scheme. mrtg chose this scheme because it uses the familiar internet url syntax. however, this familiarity gives rise to a common misconception that pasting the uri into a browser url line, or providing it to some other application that understands the http protocol, should result in the application returning some information about the object identified by the uri. such dereferencing 2 of the uri is in no way guaranteed for all audubon core terms. where possible, audubon core terms are dereferenceable, with information returned for how the metadata attribute identified by that uri is defined or used. human-centric audubon core applications, however, should not present the uris to users, nor use them as linking mechanisms. one possible exception is a selfdocumenting application that assigns metadata to multimedia resources. in that case the application might dereference the uri to provide a glossary entry aiding the user in the semantics of the metadata term. however, since dereferencing is not required for other functionality and may not be guaranteed in the long term, we suggest caution using it. in fact, all the “native” ac terms (as distinguished from those borrowed from other vocabularies) do have dereferenceable uris, presently pointing to the normative documentation. where borrowed terms have dereferencable uris links to documentation are provided. finally, note that some controlled vocabularies are defined in pdfs or other documents that do not have url links directly to each defined term. in 2 commonly called “resolution”, but the two terms are importantly distinguished in the ietf specification http://www.rfceditor.org/rfc/rfc3986.txt these cases, any dereference may only link to the beginning of the document, leaving it necessary to search in the document for the referenced definition. the proposed audubon core schema consists of 77 terms (plus the darwin core georeferencing terms by inclusion). every term has a plain text name, a normative uri, and a plain text normative definition. in addition, terms have a recommended english label for use in applications, the aforementioned details, some non-normative commentary on usage, and a non-normative, somewhat spare, set of usage notes. the final normative definitions of the standard, with full uris, will be found on the audubon core wiki http://terms.gbif.org/wiki/audubon_core_term_li st_(1.0_normative). it is expected that “best practices” documents will be developed by various user communities. to ensure that the barriers to use are as low as possible, only four terms of an audubon core record are considered to be mandatory. these are summarized in table 2, with abbreviated uris in parentheses. associated with each audubon core term is its value, whose data type is also specified. when the audubon core (or any vocabulary it references) uses literals, it is important that any metadata interchange use the literals verbatim, even if the record is declared to be in a different natural language. an example is the “type” metadata term, which is required to come from the corresponding vocabulary, dublin core. agents answering audubon core metadata queries must be able to consume and respond to queries framed with that controlled vocabulary. the normative document does not prevent a metadata publisher from asserting it has no records with a given controlled term, nor from internally mapping between a controlled vocabulary and its internal attributes, whose names may well be in a language other than english. only a small number of terms take values in a specific, english-based controlled vocabulary. of the mandatory audubon core terms, only type has any such requirements. http://www.rfc-editor.org/rfc/rfc3986.txt#_blank http://www.rfc-editor.org/rfc/rfc3986.txt#_blank http://terms.gbif.org/wiki/audubon_core_term_list_(1.0_normative) http://terms.gbif.org/wiki/audubon_core_term_list_(1.0_normative) morris, et al. – the audubon core strategies and approaches 194 it may seem odd that so few terms are mandatory. one reviewer suggested that there is no use for a metadata record that contains only the mandatory terms, because such a record would not assist in discovery or fitness-for-use evaluation. but this is definitely not the case in circumstances where the resource metadata and/or the resource data are themselves available from several disparate sources. the simplest example might be the case in which an extensive ac metadata service is offered by one server without any resource service, but with a reference to a service that holds the resource. in this case, the resource service is likely to need only the ac identifier value and might well be motivated to hold and serve only the mandatory metadata. a related scenario is one for which a user or software agent desires to formulate an ac-based query to a distant server as to whether that server holds any resources meeting a specific set of criteria relative to a given image for which only the mandatory data (including the identifier) is met. for example, entirely with mandatory data and a sufficiently expressive query language, a remote service can be asked for a list of resources it holds (or even simply knows about) that are known to have the same taxonomic coverage as the one in hand, even though the user doesn’t know what that coverage is. the reviewer suggested that mrtg could propose one or more standard subsets of ac to provide for various communities of practice, e.g. taxonomists, ecologists, etc., but the authors feel that such “profiles” are best organized by the communities themselves. thus, the architecture is meant to enable, rather than define such profiles. indeed, doing so will likely involve social and organizational considerations, e.g. the it resources available to organizations holding the media and metadata, and no single group is likely competent to provide several different profiles. instead, at the final adoption or soon thereafter, the tdwg annotation interest group (aig), of which mrtg is a task group, will probably also recommend mechanisms by which self-organizing communities can build such profiles and choose among the several tdwg mechanisms for recognizing applicability and use of standards. (see the section “sustainability” below.) in some cases, metadata terms are necessarily related to others. for example, an image might have several variants that must remain related even if they have their own metadata in another audubon core record. however, many image contributors are constrained to record their metadata in spreadsheets or other flat structures, use of which makes it difficult to represent such structural relationships. consequently the audubon core itself is primarily flat, the exception being a few structures designated as members of a serviceaccesspoint class, which describe various ways to access the media resource and related resources. one consequence of the flat structure is that a metadata publisher might have to make term definition identifier (dcterms:identifier) an arbitrary code that is unique for the resource, with the resource being a collection or a media item. the draft requires an identifier for collections and strongly recommends but does not require an identifier for media items. type (dcterms: type) any dcmi type term from http://dublincore.org/documents/dcmi-typevocabulary/ may be used. recommended terms are collection, stillimage, sound, movingimage, interactiveresource, text, panandzoomimage , 3dstillimage, and 3dmovingimage. metadata language (ac:metadatalanguage) language of description and other meta data (but not necessarily of the image itself) represented in iso639-1 or -3. copyright statement (dcterms:rights) information about rights held in and over the resource. a full-text, readable copyright statement, as required by the national legislation of the copyright holder. on collections, this applies to all contained objects, unless the object itself has a different statement. table 2: the four mandatory terms of audubon core http://dublincore.org/documents/dcmi-type-vocabulary/ http://dublincore.org/documents/dcmi-type-vocabulary/ morris, et al. – the audubon core strategies and approaches 195 several metadata records available about the same underlying resource. an important case surrounds multilingual metadata. because each metadata record is in a fixed language specified by the metadata language term (this is the language of the metadata record, not of any language featured within the multimedia resource itself), a provider might have to offer one metadata record about the same multimedia resource for each available language. the mandatory terms must be provided in every metadata record, even if repeated in other metadata records. this and other cases requiring multiple metadata carrying the same mandatory terms and only a little more, provide a huge number of combinations wherein extremely minimal metadata is in play. at the date of this writing, the normative document does not provide a mechanism for singling out a metadata record that might be overarching, the optional terms of which may be regarded as defaults for other records about the same resource. finally, many terms may be repeated in a record, but some may not. for example the “modified” term corresponds to a date at which the media resource was modified and may be repeated to reflect the history of the resource. by contrast, “date available” is a single date, or a single range of dates, at which the underlying resource became, or will become, available. audubon core designates terms that may be repeated. for use by software, two recommended serializations for the digital representation of the audubon core metadata are under development and will be submitted to the tdwg standardization process via the sustainability model described below. one is based on the world wide web consortium (w3c) resource description framework (rdf), a model for data interchange on the semantic web (http://www.w3.org/rdf/). the other is based on the extensible markup language (xml; http://www.w3.org/standards/xml/), a standard language for constraining the form of markup permitted in data interchange. another serialization, yet to be specified, will be based on delimited text files, such as commaor tabseparated text. one such serialization is implicitly provided by the audubon core ipt extension mentioned earlier. also of note, the language of the normative specification is english, but this in no way constrains applications from using labels or content of the metadata in other languages. a term is provided to denote in which language the metadata is recorded. the question of how much structure to provide in a metadata specification for science application is complex (beard, 1996). our choice was generally to avoid the issue and leave it to specific implementations, particularly as media providers may wish to have service or exchange profiles specialized to more than one purpose. we expect that the (nearly flat) normative schema enables the specification of a variety of profiles aimed at such different applications as metadata exchange, intelligent image discovery, and services providing quality control on evidence for species occurrence. sustainability and future developments in its submission to the tdwg executive committee, mrtg included a sustainability plan based on procedures similar to those of the darwin core namespace policy document (dwc 2011). this provides procedures for the introduction of new namespaces and new terms, for dealing with editorial errata, and for introducing semantic changes to terms. in addition, mrtg manages audubon core issue tracking with a google code project similar to that of the darwin core. an initial implementation is at http://code.google.com/p/auduboncore/ although google code provides a wiki, it is deliberately minimalistic and the normative documentation on terms.gbif.org will remain on that platform. more importantly, the gbif terminology platform implementation uses the semantic mediawiki extension (smw 2011) that supports reasoning, rdf export, and semantically enhanced data. several of the authors plan to explore the use of these facilities to support images with audubon core metadata published on the semantic web. summary the audubon core is a representation-neutral metadata vocabulary for the description of biodiversity multimedia resources. it is capable of implementation in various constraint languages, and with profiles that specify further constraints, best practices, or term subsets. because it is http://www.w3.org/rdf/ http://www.w3.org/standards/xml/ http://code.google.com/p/auduboncore/ morris, et al. – the audubon core strategies and approaches 196 representation-neutral, its uris may be used across a number of technologies, such as namespaces in xml schema-validated documents, rdf, and column headings in comma-delimited text files. its use of existing namespaces and vocabularies for a number of terms eases mappings from existing metadata to audubon core compliant metadata. the breadth and diversity of participation in the development of this schema, by multiple organizations, causes us to expect that a ratified audubon core will become a de facto standard for exchanging multimedia data that describe biodiversity multimedia resources. the gbif secretariat has begun to implement the audubon core schema in its next version of integrated publishing tools (ipt) as an ipt extension of darwin core. some tools have already explored use of the audubon core in xml-based implementations of the early normative document (e.g. saraiva and catalano, 2010).there is huge potential in the discovery, publishing, and usage of multimedia resources in scientific analysis, in addition to the interpretation that leads to informed decisions in the sustainable use of biodiversity resources. the need for such resources calls for service and access of biodiversity multimedia records/resources as robust and simple as for other primary biodiversity data. an early uptake of the audubon core by the stakeholder communities would not only ensure the mainstreaming of multimedia resources into biodiversity research, but would help engage citizen scientists and professional naturalists in creating, and sharing scientifically useful primary biodiversity data. references agosti, d. and egloff, w. 2009. “taxonomic information exchange and copyright: the plazi approach”, bmc research notes 2009, 2:53doi:10.1186/1756-0500-2-53 beard, k. 1996. “a structure for organizing metadata collection” in proceedings, third international conference/workshop on integrating gis and environmental modeling, santa fe, nm, january 21-26, 1996. santa barbara, ca: national center for geographic information and analysis. paper at http://www.ncgia.ucsb.edu/conf/santa_fe_cdrom/sf_papers/beard_kate/metadatapaper.html dcmi. 2011. dublin core metadata initiative, dcmi metadata terms. http://dublincore.org/documents/dcmi-terms/ dwc. 2011. darwin core namespace policy, http://rs.tdwg.org/ac/terms/serviceaccesspoint. gbif. 2007. gbif work programme 2009-2010. 54pp. http://imsgbif.gbif.org/cms_new/get_file.php?fi le=976c6c3015d7d2505d5ac1d45357c8 gbif. 2008. http://www.gbif.org/communications/news-andevents/showsingle/article/multimedia-resourcestask-group/. gbif. 2011. the integrated publishing toolkit, http://www.gbif.org/informatics/infrastructure/publi shing/ idigbio. 2011. national resource for advancing digitization of biological collections. https://www.idigbio.org/ iptc. 2010. iptc standard, photo metadata (july 2010), document revision 1, international press telecommunications council, http://www.iptc.org/std/photometadata/specificatio n/iptc-photometadata-201007_1.pdf morris r, olson a, freeland c, hagedorn g, riccardi g, carausu m, o’tauma e, and v chavan. 2009. mobilising multimedia resources in biodiversity: 2 nd report of the gbif multimedia resources task group (mrtg), march 2009. copenhagen: global biodiversity information facility. 22 pp. http://imsgbif.gbif.org/cms_new/get_file.php?fi le=a9369cf39af07c00d0891f34fe667a morris, r., olson, a., o’tuama, e., riccardi, g., whitbread, g., hagedorn, g., teage, i., heikkinen, m., leary, p., barve, v., and v. s. chavan. 2008. recommendations of the gbif multimedia resources task group. september 2008. copenhagen: global biodiversity information facility. 18pp. http://imsgbif.gbif.org/cms_new/get_file.php?fi le=e61a4e0dde908609320b1fb7dfcf3b mwg. 2010. metadata working group guidelines for handling image metadata, version 2.0, november 2010. http://www.metadataworkinggroup.org/pdf/mwg_g uidance.pdf niso. 2008. niso metadata for images in xml schema (mix) official web site. http://www.loc.gov/standards/mix/. http://www.ncgia.ucsb.edu/conf/santa_fe_cd-rom/sf_papers/beard_kate/metadatapaper.html http://www.ncgia.ucsb.edu/conf/santa_fe_cd-rom/sf_papers/beard_kate/metadatapaper.html http://dublincore.org/documents/dcmi-terms/ http://rs.tdwg.org/ac/terms/serviceaccesspoint http://imsgbif.gbif.org/cms_new/get_file.php?file=976c6c3015d7d2505d5ac1d45357c8 http://imsgbif.gbif.org/cms_new/get_file.php?file=976c6c3015d7d2505d5ac1d45357c8 http://www.gbif.org/communications/news-and-events/showsingle/article/multimedia-resources-task-group/ http://www.gbif.org/communications/news-and-events/showsingle/article/multimedia-resources-task-group/ http://www.gbif.org/communications/news-and-events/showsingle/article/multimedia-resources-task-group/ http://www.gbif.org/informatics/infrastructure/publishing/ http://www.gbif.org/informatics/infrastructure/publishing/ https://www.idigbio.org/ http://www.iptc.org/std/photometadata/specification/iptc-photometadata-201007_1.pdf http://www.iptc.org/std/photometadata/specification/iptc-photometadata-201007_1.pdf http://imsgbif.gbif.org/cms_new/get_file.php?file=a9369cf39af07c00d0891f34fe667a http://imsgbif.gbif.org/cms_new/get_file.php?file=a9369cf39af07c00d0891f34fe667a http://imsgbif.gbif.org/cms_new/get_file.php?file=e61a4e0dde908609320b1fb7dfcf3b http://imsgbif.gbif.org/cms_new/get_file.php?file=e61a4e0dde908609320b1fb7dfcf3b http://www.metadataworkinggroup.org/pdf/mwg_guidance.pdf#_blank http://www.metadataworkinggroup.org/pdf/mwg_guidance.pdf#_blank http://www.loc.gov/standards/mix/ morris, et al. – the audubon core strategies and approaches 197 saraiva, a. and cartolano, e. a. jr. 2010. biodiversity data digitizer. tdwg annual meeting, woods hole ma, 2010 http://www.tdwg.org/fileadmin/2010conference/sli des/saraiva_biodiversity_data_digitizer.pdf, smw. 2011. semantic media wiki. http://semanticmediawiki.org/ tdwg ncd interest group. 2009. natural collections descriptions: a data standard for exchanging data describing natural history collections. http://www.tdwg.org/standards/312/. xmp. 2010. xmp specification part 1; data model, serialization, and core properties, adobe systems, july 2010. http://www.adobe.com/content/dam/adobe/en/devn et/xmp/pdfs/xmpspecificationpart1.pdf http://www.tdwg.org/fileadmin/2010conference/slides/saraiva_biodiversity_data_digitizer.pdf http://www.tdwg.org/fileadmin/2010conference/slides/saraiva_biodiversity_data_digitizer.pdf http://semantic-mediawiki.org/ http://semantic-mediawiki.org/ http://www.tdwg.org/standards/312/ http://www.adobe.com/content/dam/adobe/en/devnet/xmp/pdfs/xmpspecificationpart1.pdf http://www.adobe.com/content/dam/adobe/en/devnet/xmp/pdfs/xmpspecificationpart1.pdf biodiversity informatics, 9, 2014, pp. 18-29 digiweb – a workflow environment for quality assurance of transcription in digitization of natural history collections tero mononen, riitta tegelberg, mira sääskilahti, markku a. huttunen, marko tähtinen and hannu saarenmaa* digitarium, sib labs, university of eastern finland, länsikatu 15, fi-80101 joensuu, finland *corresponding author: hannu.saarenmaa@uef.fi abstract – data produced by digitization increases the scientific use of natural history collections. however, in mass digitization, attention must be paid to the flawless management of the workflows, and high quantities of end results should not be compromised by a low standard of quality. a web-based environment digiweb was created for controlling the workflow of transcribing data from images of natural history specimens. using digiweb, it was possible to manage the workflow of transcription and data proofing, include all participants to the workflow, allow collaboration and training, and also to provide useful processing features. the data emerging from this process pass quality control standards which are supported by digiweb and based on the strict requirements of the iso 2859 standard. keywords – data entry, mass digitization introduction collections conserved by natural history museums are an important source of information on taxonomy, biodiversity and environmental change. digitization of these collections and the practices of providing open access data are expected to improve the world-wide utilization of museum data. recent advances in technology (e.g. schmidt et al. 2012; heerlien et al. 2013; tegelberg et al. 2014) offer solutions for increased efficiency in the imaging of natural history specimens. these automated imaging pipelines are now producing huge amounts of digital content in many projects, and thus there is an urgent need to develop new solutions for streamlining the transcribing of data from images. such data entry is still mostly based on manual, labor-intensive work. however many approaches are being tried to modernize transcription. optical character recognition (ocr) is a promising method for typewritten material (e.g. haston et al. 2012; tulig et al. 2012). crowdsourcing (flemons and berents 2012; hill et al. 2012; herbaria united 2014; les herbonautes 2014) has lately gained much attention, but is not an appropriate solution for inhouse and time-bound project work. the digitization of biological and geological collections can be described as a process or workflow, containing steps such as transportation, tagging, imaging, transcription and archiving (e.g. dou et al. 2011, 2012; lehtonen et al. 2011; nelson et al. 2012; tegelberg et al. 2012). the methods used in some of these steps may vary depending on specimen type. for example, the physical dimensions of the objects strongly affect the imaging phase (tegelberg et al. 2014). the management of mass digitization is however particularly sensitive to abnormalities in workflows and when digitizing different specimen types, the differing steps of alternate solutions may increase the risk of error if not carefully controlled. efficient management of the digitization process can be achieved by use of an information system that is designed not only for controlling the workflows, but also to promote the quality assurance of results. for example, based on a survey among digitizers, automatic filtering of data by country and collector is expected to reduce mistakes made during data entry (drinkwater et al. 2014). such a gentle change in digitization workflow leads to familiarity with the geography or handwriting of a collector. this allows the 18 mononen et al – digiweb digitizer to concentrate on specific parts of data entry, with good results. during the process, the collaboration of partners and specialists is required. when using a well-established imaging method for specimens, scientific expertise is particularly important for transcribing specimen data. fast feedback and the formation of a community of specialists around a digitization project may enhance the quality of the outcomes. the transcription would also benefit from a knowledge base of the results of earlier digitization projects, available both in-house and on the web. for example, the sometimes rather cryptic handwriting of certain collectors might already have been cracked by a devoted scientist, so providing helpful material for transcribers. digitarium is a digitization center providing services to museums of natural history (tegelberg et al. 2012). the basic idea is to outsource digitization from museums to a factory-like setting, where entire collections are processed, and all steps are automated as much as possible. this idea was first implemented by the commercial company océ for the herbarium of the natural history museum in paris (pignal & michiels 2012). this was a real break-through, the birth of mass-digitization, which led to increased focus and funding for digitization worldwide. mass-digitization of herbaria is now underway also in leiden (heerlien et al. 2013), helsinki (tegelberg et al. 2014), and oslo. digitarium has expanded the concept by not only providing a commercial mass-digitization service, but also research and development on industrial engineering and biodiversity informatics, tackling the bottlenecks of the whole digitization process in a real production environment. this paper discusses the management of the steps required for selecting the specimens for data entry, transcribing the labels from the images, and validating (quality control) the result. the aim was to develop a web-based environment that allows a gradual streamlining of the whole digitization process. especially, the process is intended to be open to all participants of the digitization projects concerned, promote collaboration, and include the use of helpful features which lead to a high-quality result. methods preparatory steps the first steps in digitarium's digitization process are specimen labeling and imaging (lehtonen et al. 2011). for each specimen, a unique identifier (id) is generated, and a physical label with the id is attached to the specimen. the labeling process enables different types and sizes of labels, with variable contents to be generated. all of the information concerning the generated ids and their hierarchy are stored in a mysql database, which contains all the necessary data to control the process. the labeling and imaging methodology used at digitarium in the massdigitization of herbarium sheets and insect specimens has been explained in more detail (tegelberg et al. 2014). the actual digitized data is not stored in the mysql database, but as image files and xml documents. this allows version control through the figure 1. files and folders – the digiweb data store. 19 mononen et al – digiweb various phases of the digitization process better than a relational database (fig. 1). the xml file has a unique id and also a corresponding web address created for the digital object corresponding to the specimen, to which data can be later uploaded. the folder contains all the images which relate to the same id, and the metadata of the specimen is contained in the xml file according to the darwin core standard. the digiweb environment digiweb is a web application created for the browsing of produced images and data, and for transcribing and verifying the images metadata. it relies on the digitarium file system hierarchy of imaged material, and also on a database which contains the hierarchy of labels and multiple assistant data sets. digiweb is a java enterprise edition (ee) application, and the user interface utilizes java serverfaces 2 (jsf2) and primefaces frameworks to implement rich standalone software-like features in a browser. digiweb has a built-in version control system of saved metadata where the newest version of the metadata is considered as the primary one. therefore, all versions are available to trace the accumulation of data, and a rollback function to a specific version is available. digiweb is divided into three main views: browsing, transcribing and administration. on the browsing page, a user can view the collections and specimens. with permission, the user may reserve specimens to form a queue of data to be transcribed. on the transcribing page, the user is able to transcribe the data from the queue of specimens. the administration page contains the administrative functions of digiweb, e.g. user management, user monitoring and a function that enables all the reports generated by the system to be viewed. in user management, users are allocated a role which defines the views they can access and the functions they can use. the roles are: figure 2. the user interface of digiweb. 20 mononen et al – digiweb administrator, in-house participant and client. users with an administrator role have access to all collections and all functionalities. the in-house role is for those users who are involved in the production and entering of metadata. each inhouse user has customized access rights based on their skills and experience. as a result, users can be divided into groups of standard staff and experienced staff. the client role is for those not participating in everyday data entry work but who need to monitor the process. on the transcribing page, the view is divided into two parts (fig. 2). first are the specimen images which can be zoomed and panned. the image viewer uses standard jpeg images as an input, which are served by the digitarium api (application programming interface). the second part of the view contains all of the input fields grouped into tabs. the users required that the image and the input text fields must be visible simultaneously. the data entry region is dynamically generated according to the structure and components defined in the separate xml file. it supports a full list of darwin core terms (taxonomic databases working group 2014) but can also be customized case specifically by hiding terms that are not needed. external lookups and web services in digiweb, for every darwin core term there is a configured input component in the user interface. supported components are the input text, input text area, auto-complete input text and a drop-down menu. for example, the entry of plant species names is implemented by an auto-complete input text component, including a suitable name suggestion feature. finalization of the name can be done by choosing the right name from the list, which is then automatically added to the text area. the taxonomic list of species names is taken from the plant list (http://www.theplantlist.org/) and contains genera that represent families which are currently being digitized. in future, also authority lists of names of other organism groups will be included. in a similar way, the country (and in some cases a smaller geographical area) representing the location of the collection effort can be chosen from a selection list, after inputting the first letters of the location. this look-up list originates from the university of oslo. a look-up function for a collector’s name from a list showing historical and present plant collectors around the world is also supported by digiweb. the name list has been created by harvard university, and in digiweb it aims to help the transcriber when the handwriting of the collector is difficult to translate. digiweb uses the collecting locality service (tähtinen et al. 2014) developed by the biovel project (http://www.biovel.eu). this is a restful web service using a json data transfer format with fields based on the darwin core standard. this service has been published in the biodiversity catalogue (tähtinen 2014) for general use. the idea of this service is to provide an easy way to access and reuse existing locality data, which may also include geo-reference information. as a source of information, it can serve the local database of already digitized specimens, and also gbif’s global data portal where there are currently over 500 million records. algorithms such as fuzzy text search can be used to find the most probable collecting locality of a specimen, by giving the collector’s name and additional information, in particular, the collection date. possible collecting localities are shown on the digiweb interface and the user may then choose from the options displayed. localities cannot be saved manually in separation of the specimens, but once the specimen data has been saved by digiweb, any new collecting locality information becomes available for transcribing labels of further specimens. test case oslo in 2013-2014, digiweb was used in a project to transcribe data from the herbarium of the university of oslo. in this project, the management of workflows by digiweb was developed for a large number of specimens and additionally, methods for the quality assurance of data were created. during the project, a total of 168,527 herbarium specimens were imaged. transcription began during the imaging process and it will continue until the end of 2014. in practice, an image became available for transcription within 21 http://www.theplantlist.org/ http://www.biovel.eu/ mononen et al – digiweb five minutes of its creation. the contracting company digforsk as performs transcription in five digitization centers in northern norway. up to 30 distance workers simultaneously use digiweb to transcribe the data from the images produced at digitarium. the transcribing process began by conducting a training session at the digforsk premises. aimed at the leaders of data entry centers, the training was conducted by the project manager, a botanical expert and a software developer working at digitarium. in addition, additional training was offered by taxonomists from the university of oslo. the training covered the use of the webbased transcription tool, reading and understanding old labels of botanical samples, and the accurate and uniform way to transcribe the data. after training, there was a practice period of two-four weeks, based at the digitization centers. during that time, transcribers in norway worked in a test environment of digiweb, and specialists at digitarium offered feedback on the quality of their work through digiweb and by e-mails. a person was only allowed access to the production environment of digiweb after approval by specialists based at digitarium. figure 3. the workflow developed for quality control. 22 mononen et al – digiweb results the workflow the functionality of the program in managing workflows was not found to be dependent on the type of specimens or the imaging methods used. thus, the program can use any data that follow the file system structure and format used at digitarium. the developed workflow for transcription and verification using digiweb is shown in fig. 3. the workflow contains altogether seven specimen states in two lines. the state of the specimen contains two types of information: difficulty and step. difficulty is indicated by color where green (g) represents an easy specimen and red (r) represents a complex specimen. the step in the transcribing process is indicated by the number of spots in the symbol for example 1g means one green spot, i.e., initial transcribing. the state transitions of a specimen are finalized when no modifications need to be made and data is ready to be delivered to the customer. the default route in transcribing workflow is 1g, 2g, 3g, and then finalized. in specimen state 1g, the initial transcription is done by nonprofessional, possibly inexperienced staff, who pass the specimen to the next level. in specimen state 2g, more experienced staff proof the specimens. they pass the specimen to state 3g, which indicates that transcription is finished and ready for the final quality check. in the initial transcription (1g), a specimen may be forwarded to the red line due to a complex label. then more experienced staff will check the specimen indicated as 1r. from that state, a resolved specimen may be forwarded back to the green line but if it stays unsolved, it will be passed to 2r. samples in state 2r are checked by digitarium's experts. if it seems almost impossible to transcribe, it will continue to state 3r and await inspection by specialists of the client organization. in a large collection, there may be tens, hundreds or occasionally thousands of specimens collected by the same collector. as a result, the same information will be transcribed several times, by persons with different skill levels. this may create different spellings of the same locality, unnecessary duplication of records, and results in poorly usable data. this can be avoided in digiweb by phasing the transcription of different fields. that is, the user does not need to transcribe all fields at once. instead, the locality data of specimens can be processed in a later phase as a bulk operation after the collector’s name has already been entered. in this case, the transcriber already knows the specific handwriting style and is familiar with the visited locations. consequently the produced data will be more uniform and of higher quality. this idea was preliminarily tested by the scientific curator of joe (herbarium of the university of eastern finland). during the first round of data entry, the scientific name of the taxon, the collector’s name, collector’s record number, and the date of collection were transcribed, but locality was skipped. during the second round, the locality information was transcribed. between the two rounds, the data were re-organized according to the collector and date to facilitate bulk operations. both the duration of the data entry and quality of the end results were assessed. results showed that doubling the sample size from 374 to 748 items reduced the portion of collectors’ first samples from 54 % to 41 %. it was also found that when the collector’s name was successfully transcribed, the average total time spent on the transcription of a specimen using the re-organized data dropped from original 5 minutes to 2 minutes. thus, the presented workflow could in future be changed to having two 1g states 1g and 1gfinal. from a quality assurance point of view, concentrating on the repeating localities increased the validity and uniformity of the metadata produced. acceptance procedure the procedure for quality assurance was developed at the beginning of the test project. in practice, before the data is delivered to the customer, a final quality control check takes place (fig. 3). when at least 1000 (test batch size) specimens have been marked as finished (state 3g), they are subjected to quality control. the quality check and reports are created by using the digiweb acceptance test tool which is based on the statistical iso 2859 standard. in the check, a batch 23 mononen et al – digiweb of 1000 samples is created and 80 randomly selected samples are checked by an expert. if there are ≤3 rejected samples, the whole batch of 1000 samples is accepted. however, in the case of >3 rejected samples, the whole batch (excluding the accepted samples) is returned for proofing (state 2g). the acceptance criteria may vary, depending on the customer demands. typically, however, at least the country, collector name, scientific name, and catalog number must be free of error, because without them, the specimen cannot be found in a search. in general, small errors in other fields may be tolerated as they can be checked directly from the image. the sets of specimens accepted in the qualitycontrol process are considered as finalized and can be delivered to the collection owner. quality and quantity of transcription during the test project, excluding the initial practice weeks, the data of around 5000 specimens were transcribed every week. there were on the average 24 different people at work each week who transcribed on the average 42 samples in a 3hour part-time working day. as transcribing progressed, this allowed proofing of the metadata to start. on average, about 8000 specimens were proofed per month. this resulted in some degree of backlog; however, this will be cleared at the end of the project by people gathering experience on transcription. on average, the transcription and proofing of the label data of a specimen lasted about five minutes, however this varied significantly between specimens. typed and printed labels were relatively fast to transcribe but some of the hand written labels went through the whole of the ‘red line’ of the workflow (fig. 3), thus expanding the time spend on that particular specimen. at digitarium, the quality control of data entry was centralized to a specialist to ensure the uniformity of decisions of acceptance. the requirements in quality control were created in cooperation with the experts in the customer organization. reports made by the quality controller by using digiweb showed the main errors that led to rejection of the data were missing information (especially some or all of the locality figure 4. errors found in transcribed data during the first four months of the test case project. in total, there were 187 rejected specimens. 24 mononen et al – digiweb information) and incorrect collector name (fig. 4). problems with taxa were also found, often connected to markings expressing hybrids and determiner’s, or other expert’s, doubts (e.g. “cf.”) about the identification of the taxon. the quality of transcription was followed from the beginning of the project, and problems were identified in applying dwc terms and syntax correctly. therefore, a new feature was introduced into digiweb. any deviation in the presentation style of year, month, day, or collector’s name automatically caused the input text to turn red. the color did not change to black until the term was presented according to the style defined by dwc. this feature was especially helpful in producing uniform data concerning the dates of records. in general, of the batches of 1000 specimens, about 50% were accepted during the first round of quality control. this percentage increased gradually over the first six months. fig. 5 presents the amounts of rejected samples in all batches checked by the quality controller during the first months of the test case project. the clear decrease in errors found can be explained by the increasing experience of data entry personnel, as well as thanks to use of the help features provided in the system. communication in order to enhance rapid communication between digitization centers which were spatially distributed in northern norway, digitarium (joensuu, finland) and the customer (oslo, norway), and to lower the transcriber’s threshold to ask questions from the experts, a built-in messaging system was implemented in digiweb. such an integrated communication system was found to be an easy way to send questions and answers, without knowing individual contact information like e-mail or skype addresses. general information was delivered to all digiweb users by using the news section in the main page, and included information and guidance on new digiweb features. to ensure that all transcribers had easy access to the most recent versions of all documents (e.g., data entry manual, guidelines and learning material), a document directory was also integrated into the system. figure 5. numbers of specimens rejected by the quality controller during the first four months of the test case project. 25 mononen et al – digiweb discussion in digiweb, attention has been paid to the support of managing the digitization of large collections. based on the test project, the management of large scale data entry work is possible using digiweb, and is not dependent on the whereabouts of the participants of the project. in addition, digiweb can be used as a training tool and an information source when pursuing a high level of quality in produced metadata. recent improvements in digiweb, combined with temporary spreadsheet assistance, made it possible to divide the current transcription workflow into two partial rounds. this enabled sequential entry of repeating information, with the aim of increasing the quality of locality data and improving the results of geo-referencing for end users. it is acknowledged, however, that further testing of the proposed workflow is still needed, and the ability to use digiweb efficiently in phased transcription needs still some refinement. the workflow developed here has special characteristics which derive from the requirement of the customer to support iso 2859 based quality control, and from the use of large number of (initially) inexperienced transcribers. this necessitates two validation steps, first by the team leader at digforsk as and then at digitarium. this is relatively expensive, and can only be defended by the fact that the basic transcribers’ work cost is subsidized by the employment office. compared with workflows developed elsewhere, there are two human data curation steps. the kurator workflow (dou et al. 2011, 2012) only contains one, and includes more automation. in crowdsourcing projects, the human curation steps have entirely been replaced by repeated transcription of the same samples (flemons et al. 2012; hill et al. 2012): when there are enough repetitions that match, the result is automatically accepted. what is “enough” repetition has been studied by shah (2014) who developed a consensus model for aligning the differing transcriptions. a digitization workflow tool such as digiweb would ideally support several quality assurance methods. the labels of specimens stored at natural history museums contain a set of basic information. however, the specimen data may be presented in a personalized way. for example, old handwriting, invalid locality names, and the variable uses of latin will cause problems to the transcribers and translators of the label information. therefore it is important that the tool used for transcription is easy to learn, easy to use, and provides help when possible to the user. in the case of digiweb, the functionality of the system has been tested with both academics and also those persons without any academic education but with reasonable it-skills. this work is on-going, with new ideas constantly emerging. a fundamental question for quality assurance is whether we want a literally accurate transcription of what is written in the label, or its interpretation, where differing spellings, ancient locality and taxon names are harmonized. if we can afford it, both would be nice. literally accurate transcription would ideally need to be saved because the interpretation can go wrong, and because of its cultural history value. on the other hand, when images are available, they serve as the literally accurate information to fall back when in doubt. multiple, slightly differing transcriptions showed to be problematic in this project, and our opinion now is to avoid them, and only save modern interpretations. however, such interpretations can only be made by relatively experienced staff, and not inexperienced workers or volunteers. in conclusion, aligning repeated transcriptions needs to be studied more. digiweb uses darwin core as the standard and basis of data terms. the aim of dwc is to facilitate the sharing and alignment (integration across records) of information found in the specimen labels. dwc still has some shortcomings in support of digitization. there are separate fields for verbatim data which facilitates data entry of the labels literally. however, the verbatim fields do not cover everything. therefore, a new field for all label information, e.g., “verbatimlabel”, might be necessary. in digiweb, the dwc-standard is being followed precisely and instructions for each term are easily available, with different languages 26 mononen et al – digiweb available if needed. thus, digiweb can be used as a training tool and it gently guides the users towards conformance in data entry work. tools such as taxonomic lists, geographic hierarchies, and search facilities for collectors’ names helped data entry workers to produce data efficiently. in addition, in long series of the same taxon, the ability to select quickly the previously written taxon name with a mouse click made the workflow easier and faster. however, for example, the taxon name lists accepted by the scientific communities are not available for all taxon groups. indeed, with the exception of the plant list, they appear frequently not to be publicly available: open access to such lists should be promoted. one possibility is that they should be available from the major nomenclators involved in their compilation, which would provide high quality and flawless information. in practice, incorporating tools such as lists in digiweb is quick and easy: it is possible to use such tools through the servers at digitarium or remotely through available apis. the interpretation of locality names (especially when presented in latin) was proven to be a difficult task. based on a request by the customer, the country was considered as the most important locality information. however, in the labels, country is not always mentioned, and specifying the country was demanding when for example, the only information given was a name of a mountain. according to the results, the most common errors found by quality control often concerned actions for which help was not readily available in the digiweb environment. therefore, the features aimed at helping the transcribers were deemed to be found useful. for geographic locations, gazetteers are publicly available. for georeferencing, efficient services have been created by geolocate (rios and bart 2008). however, we must point out that collecting localities are not just any localities, but rather place-collector combinations, repeatedly visited by the same collectors. therefore, we developed a collection locality service, which also takes into account the collector’s name and the time of collecting. by tracing the movements of collectors, we can get more accurate information for geo-referencing. furthermore, the geographic hierarchies obtained from gazetteers reflect the situation today, and not that of the past. thus there is a need for establishing historical names and also their periods of validity. if data is pooled and made available, with each new collection that is digitized, the supporting tools become more efficient. our experience in streamlining transcription is that much can be gained by using shared lookup services and big pools of data that are already available. samples must be distributed to the best available agent with regard to language, handwriting, taxonomy, and geography. this distribution should happen automatically in the contemporary electronic marketplace of digitization services. transcription in isolation is waste of time, and thus more web services are needed. the pooling of data is making this degree of distribution and access possible. the possibility to share knowledge and solve problems in a community is a feature expected to enhance the levels of motivation and skills of those involved. on the other hand, spreading important information to all users at the same time may encroach upon working time. for the “digiweb community,” releasing news for example about new features was important. digiweb was also used for showing the results of final validation, and for commenting on mistakes in data entry. this allowed all partners to recognize the specific problems in transcription and work together to find solutions to them. the data delivered by such workflows needs to be reliable enough to meet the standards of science. in the validation of data, human resources are needed. a person may have specialized in the taxonomy of certain families or in the handwriting of a collector from the 19th century; however, whether such persons exist in every organization is questionable. therefore access to the databases of earlier digitized collections might prove helpful. in this project, embedding access to other digital contents was worthwhile, especially when the handwriting of the collectors was poor. in these cases, after resolving the collector’s name, locations could often be discerned by tracking down the collector’s path by following the dates of 27 mononen et al – digiweb the collecting events. such “natural history intelligence,” where new facts are derived from examining pooled information, is actually common practice in museums. when we understand that transcription of the entire specimen data in one pass may not be the right thing to do, we are close to making a fundamental conclusion: digitization is annotation. the digitization of scientific objects will never be finished, since new facts and measurements which relate to them are always emerging. an information system for digitization must therefore support annotations. ideally, annotations can be inserted into distributed databases, wherever the specimen data may be found. these issues have been explored by the annosys (tschöpe et al. 2013) and filtered-push (morris et al. 2009) projects. digitization can also be seen as an asynchronous workflow that can span over decades. for instance, when important details such as geographic coordinates or a new identification have been annotated to a specimen, this may trigger workflows that push related data to other related specimens. another workflow can then notify the curators to validate such annotations. the development of digiweb will continue in co-operation with its users. as more lookup services emerge, these will be included as new features. finally, automatic data entry and georeferencing based on previously resolved label information will also be introduced as part of the workflows. acknowledgements the development of digiweb has been financed by the european social fund (grant no. 703990) and european regional development fund (grant no. 806527), and by the biovel project (the european union's seventh framework programme for research, technological development and demonstration, grant no. 283359). we thank the staff of the university of oslo and digforsk as for fruitful collaboration. references dou, l., d. zinn, t. mcphillips, s. köhler, s. riddle, s. bowers and b. ludäscher. 2011. scientific workflow design 2.0: demonstrating streaming data collections in kepler. international conference on data engineering 2011. doi.ieeecomputersociety.org/10.1109/icde.2011.5 767938 dou, l., g. cao, p.j. morris, r.a. morris, b. ludäscher, j.a. macklin and j. hanken. 2012. kurator: a kepler package for data curation workflows. procedia computer science 9: 16141619. doi: 10.1016/j.procs.2012.04.177 drinkwater, r.e., r.w.n. cubey and e.m. haston. 2014. the use of optical character recognition (ocr) in the digitization of herbarium specimen labels. phytokeys 38: 15-30. doi: 10.3897/phytokeys.38.7168 flemons, p. and p. berents. 2012. image based digitization of entomology collections: leveraging volunteers to increase digitization capacity. zookeys 209: 203-217. doi: 10.3897/zookeys.209.3146 haston, e., r. cubey, m. pullan, h. atkins and d.j. harris. 2012. developing integrated workflows for the digitization of herbarium specimens using a modular and scalable approach. zookeys 209: 93102. doi: 10.3897/zookeys.209.3121 heerlien, m., j. van leusen, s. schnörr and k. van hulsen. 2013. the natural history production line. in: digital heritage international congress (digitalheritage), oct. 28 – nov. 1. vol. 2, pp. 289-294. marseille, france. doi: 10.1109/digitalheritage.2013.6744766 herbaria united. 2014. accessed at http://herbariaunited.org/athome 2.6.2014. hill, a., r. guralnick., a. smith, a. sallans, r. gillespie, m. denslow, j. gross, z. murrell, t. conyers, p. oboyski, j. ball, a. thomer, r. prysjones, j. de la torre, p. kociolek and l. fortson. 2012. the notes from nature tool for unlocking biodiversity records from museum records through citizen science. zookeys 209: 219-233. doi: 10.3897/zookeys.209.3472 lehtonen, j., s. heiska, m. pajari, r. tegelberg and h. saarenmaa. 2011. the process of digitizing natural history collection specimens at digitarium. in: jones m. b. and c. gries (eds) proceedings of the environmental information management conference 2011 (eim 2011). september 28-29, 2011. santa barbara, ca. university of california, pp. 87-91. 28 mononen et al – digiweb https://eim.ecoinformatics.org/eim2011/eimprodeedings-2011. doi: 10.5060/d2nc5z4x les herbonautes. 2014. accessed at http://lesherbonautes.mnhn.fr/ 2.6.2014. pignal, m. and h. michiels. 2012. switching to the fast track: rapid digitization of the world's largest herbarium. botany 2011 – colombus, ohio marc pignal, henri michiels 11th of july, 2012. http://collections.mnhn.fr/wiki/attach/visit_october 2012/paris-herbarium-digitization_2012-07-12.pdf morris, p.j., m.a. kelly, d.b. lowery, j.a. macklin, r.a. morris, d. tremonte and z. wang. 2009. filtered push: annotating distributed data for quality control and fitness for use analysis. american geophysical union, fall meeting 2009, abstract #in34b-08. http://adsabs.harvard.edu/abs/2009agufmin34b.. 08m nelson, g., d. paul, g. riccardi and a.r. mast. 2012. five task clusters that enable efficient and effective digitization of biological collections. zookeys 209: 19-45. doi: 10.3897/zookeys.209.3135 rios, n. and h.l. bart. 2008. community building and collaborative georeferencing using geolocate. in: weitzman, a.l., and l. belbin, (eds). proceedings of tdwg (2008), fremantle, australia. http://www.tdwg.org/fileadmin/2008conference/do cuments/proceedings2008.pdf#page=46 schmidt, s., m. balke and s. lafogler. 2012. dscan – a high-performance digital scanning system for entomological collections. zookeys 209: 183-191. doi: 10.3897/zookeys.209.3115 shah, m. 2014. accuracy assessment of crowdsourced data in biological specimen transcription. master’s thesis 58 p., 2 appendix (13 p.), university of eastern finland, school of computing, joensuu. http://epublications.uef.fi/pub/urn_nbn_fi_uef20140793/urn_nbn_fi_uef-20140793.pdf tähtinen, m. 2014. collecting locality web service in biodiversitycatalogue. accessed at https://www.biodiversitycatalogue.org/services/62 16.5.2014. tähtinen, m., t. mononen and h. saarenmaa. 2014. workflows for automation of digitisation of biological collections. 57-59. in: saarenmaa h (ed.). use cases, workflows, benchmarking, and related sprints in biovel june 2013 february 2014. capacities programme of framework 7: ec e-infrastructure programme, e-science environments infra-2011-1.2.1 biovel biodiversity virtual e-laboratory. deliverable report d2.4. 78 p. european commission. http://www.biovel.eu/images/publications/internal documents/d2.4reportanddocumentationofsprints-final28february2014.pdf taxonomic databases working group. darwin core. 2014. accessed at http://rs.tdwg.org/dwc/ 17.4.2014. tegelberg, r., j. haapala, t. mononen, m. pajari and h. saarenmaa. 2012. the development of a digitising service centre for natural history collections. zookeys 209: 75-86. doi: 10.3897/zookeys.209.3119 tegelberg, r., t. mononen and h. saarenmaa. 2014. high performance digitization of natural history collections: automated imaging lines for herbarium and insect specimens. taxon. tschöpe, o., l. suhrbier, a. güntsch and w.g. berendsohn. 2013. annotating biodiversity data via the internet. taxon 62: 1248-1258. doi: http://dx.doi.org/10.12705/626.4 tulig, m., n. tarnowsky, m. bevans, a. kirchgessner and b.m. thiers. 2012. increasing the efficiency of digitization workflows for herbarium specimens. zookeys 209: 103-113. doi: 10.3897/zookeys.209.3125 29 introduction methods preparatory steps the digiweb environment external lookups and web services test case oslo results the workflow acceptance procedure quality and quantity of transcription communication discussion acknowledgements references microsoft word koffi_proofs.docx biodiversity informatics, 10, 2015, 56-64   56   the present state of botanical knowledge in côte d’ivoire kouao jean koffi1*, akossoua faustine kouassi2, constant yves adou yao3, adama bakayoko1, ipou joseph ipou2, jan bogaert4 1university nangui abrogoua, ufr-sn, 02 b. p. 801 abidjan 02, côte d’ivoire. 2centre national de floristique (ufhb), 22 b. p. 582 abidjan 22, côte d’ivoire. 3university félix houphouët boigny, ufr-biosciences, 22 b. p. 1682 abidjan 22, côte d’ivoire. 4university of liège, biodiversity and landscape unit, passage des déportés 2, b 5030 gembloux-belgique. *corresponding author: koffi kouao jean: kouaojean@yahoo.fr abstract.—the aim of this present study is to summarize the current state of research on the flora of the côte d’ivoire from the sig ivoire database to better direct future collection efforts. herbarium specimen data used for this study covered the period from 1894 to 2000, and were assembled by 226 collectors. this database comprises 15,228 samples, grouped in 3621 species, 1371 genera, and 198 families. a grid system was used to cover the ivorian territory at spatial resolution of 0.75° x 0.75°. indices of evenness and completeness were calculated to characterize sampling and identify floristically well-known regions. the exploration of the ivorian territory is far from uniform, such that some areas were more densely surveyed, but others partially or not at all. the regions of grands ponts, agnébytiassa, loh-djiboua, part of gbèkè, boukani, san pedro and cavally were floristically well known; environmentally, the largest gaps in coverge were in the mountains in western côte d'ivoire. key words.—côte d’ivoire, flora, biodiversity, sig ivoire, gis, completeness. the forests of the tropical world are characterized by particular community structures and specific floristic compositions. floristic richness of tropical forest have been highlighted by numerous studies (gentry 1988; wright 2002). for leigh et al. (2004), tropical forests are clearly museums of species diversity. the flora of tropical africa has attracted the interest of many botanists, such as auguste chevalierie, andré aubréville, george mangenot, and ake assi laurent. these botanists, via several summaries (de gouvenain and silander 2003; kelatwang and garzuglia 2005; lehmann and kioko 2005; jørgensen 2006; nair 2006; kowero et al. 2006; glenday 2008) have highlighted many of the details of the distribution and diversity of the african flora (aké-assi 2001, 2002). in côte d'ivoire, many studies (kouakou 1989; corthay 1996; kouamé 1998; nusbaumer 2003; adou yao and n'guessan 2005) have focused on the flora in general (aké-assi 2001, 2002). work on botany in côte d’ivoire began early in the twentieth century, and is active continuing today, with many detailed studies. the flora of côte d’ivoire is rich; according to recent estimates, it includes 3677 species of vascular plants, divided between forests and savannas (ake-assi 1961, 1962, 1976, 1984, 1998, 2001, 2002). several young researchers have launched botanical inventory studies in key areas of côte d'ivoire as part of their theses or other research. still, despite all these efforts, many areas are yet to be inventoried in detail and remain poorly known. these gaps underline the importance of increasing geographic knowledge of the flora of country to provide a solid base for development of future strategies for documenttation, understanding, conservation, and use of the biodiversity of the country. it is in this perspective that, for many years, the conservatory and botanical garden of geneva (cjbg) and the national floristic center of the university of abidjan, with the help of the swiss center for scientific research in côte d’ivoire (csrs), have conducted research on the flora and vegetation of the country (bänninger 1995; corthay 1996; kouamé 1993, 1998; dotty 1999; bakayoko 1999; menzies 2000). cjbg implemented a geographic information system (gis) for côte d'ivoire, which includes rich botanical data, with diverse map information on the physical environments of the country. it combines a relational botanical database and a mapping environment (gauthier et al. 1999); this system provides the base data for this study. biogeographic studies aim to understand how living organisms are distributed spatially, which factors influence those distributions, and how these patterns change over time (brown and lomolino 1998). biodiversity databases are the biodiversity informatics, 10, 2015, 56-64   57   main source of information for such studies, specifically data that place particular species at georeferenced locations at specific points in time. based on this information, spatial patterns can be investigated at scales ranging from local to global (brown and maurer 1989). hence, biogeographic understanding depends heavily on the quality and completeness of information in biodiversity databases. the present study aims to summarize the current state of knowledge of the flora of the côte d’ivoire as summarized in the sig ivoire database. specifically, study goals are (1) determine the frequency and distribution of botanical investigations across côte d’ivoire, (2) assess the intensity and completeness of these explorations, and (3) identify areas that remain poorly known in the country. methods study data the republic of côte d'ivoire covers an area of 322,462 km2, limited to the north by mali and burkina faso, to the west by liberia and guinea, to the east by ghana, and to the south by the atlantic ocean. the country has 22 million inhabitants, according to general census of population and housing of 2014. côte d'ivoire covers parts of two climatic zones: equatorial and tropical (eldin 1971), and as such has two rainy seasons and two dry seasons, although northern côte d'ivoire has only one rainy and one dry season. the natural vegetation types across the country include dense humid evergreen forest, semi-deciduous rain forest, montane rainforest, and savannah (guillaumet and adjanohoun 1971). the terrain consists of plains, plateaus, and mountains, with mount nimba (1750 m) representing the high point in southwestern côte d'ivoire. côte d'ivoire is characterized by desaturated lateritic soils, ferruginous soils, eutrophic brown soils, hydromorphic soils, and pseudo-podzols. (aubert and segalen 1996). plant occurrence database the database of herbarium specimens for côte d'ivoire has been developed at the conservatory and botanical gardens of geneva, and summarizes the ivoirian specimen holdings of herbaria in abidjan (cnf), geneva (cjgb), natural history museum of paris, and university of wageningen (wur). as each record in the database contains geographic coordinates, it is related to maps through a gis called sig ivoire (gautier et al. 1999). the data are managed in an access database, and related to ecological and spatial data in raster formats using idrisi (version 2.0) and arcview (version 3.2). sig ivore thus integrates botanical information with environmental data through a geographic information system (gautier et al. 1999), and aims marshal the best information to address conservation needs of overlooked species in the face of environmental degradation (chatelain 2002). database completeness data were cleaned via an iterative series of inspections and explorations designed to detect and document inconsistencies. (1) we created lists of unique names in each dataset in excel, and inspected them for repeated versions of the same taxonomic concepts: misspellings, name variants, different versions of authority information, etc. such repeated name variants were flagged, checked via independent sources, and corrected to produce single scientific names that correctly referred to single taxa. (2) we checked for geographic coordinates that fell outside of the country, but that were referred to côte d’ivoire. (3) within the country, we checked for consistency between textual descriptions of region and geographic coordinates. in each case, where possible, we corrected the data record; where no correction was clear, we discarded data, recording data losses at each step in the cleaning process. (4) we discarded data records for which information on year, month, or day of collection was lacking; we created a unique ‘stamp’ of time as year_month_day. we then aggregated point-based occurrence data to 0.75° spatial resolution across côte d’ivoire. this spatial resolution was the product of a detailed analysis of balancing the benefits of aggregating data (i.e., larger sample sizes), versus the negative of loss of spatial resolution that can make important geographic features imperceptible. details of this procedure are provided in ariño et al. (in prep.). we produced shapefiles representing the 0.75° grid system in the vector grid module of qgis, version 2.4. we added grid system identification codes to each individual occurrence datum, and aggregated each datum to the coarseresolution aggregation squares. in excel, we explored relationships between data on species identity, time, and aggregation square. we calculated total numbers of records available from each grid square (termed n); we reprojected biodiversity informatics, 10, 2015, 56-64   58   the aggregation grid to an africa albers equalarea conic projection, measured the area of each grid square (taking into account the area of grid squares that overlapped the edge of the country), and calculated numbers of data records per unit area. because data for plants of côte d’ivoire were not massively abundant, full development of completeness indices (sousa-baena et al. 2013) proved rather fruitless: very few aggregation aquares could be termed well-known; as a consequence, we used a more crude criterion of a density of 15.2 occurrence data points/1000 km2 to establish which squares should be considered well-known. once we had established these criteria in qgis, we linked the table of grid square statistics to the aggregation grid, and saved this file as a shapefile. we created a shapefile of well-sampled aggregation squares, which we in turn converted to raster (geotiff) format using custom scripts in r version 3.1.2 (venables and smith 2014). this raster coverage was the basis for our identification of coverage gaps, as follows. we used the proximity (raster distance) function in qgis to summarize geographic distances across the country to any well-sampled aggregation square. to create a parallel view of environmental difference from environments represented in well-sampled areas, we plotted 5000 random points across côte d’ivoire, and used the point sampling tool in qgis to extract values of each point to the geographic distance raster, and to raster coverages (2.5’ spatial resolution) summarizing annual mean temperature and annual precipitation drawn from the worldclim climate data archive (hijmans et al. 2005). we exported the attributes table associated with the random points, and analyzed further in excel, as follows. we first standardized the values of each environmental variable to the overall range of the variable as : (xi – xmin) / (xmax xmin), where xi is the particular observed value in question. we then created a matrix of euclidean distances in the two-dimensional climate space, calculating distances for each of the points with a geographic distance >0 to all of the points with geographic distance of zero; the latter represent points falling in well-sampled regions, whereas the former are scattered across the entire country; points in well-sampled regions were assigned (by definition) environmental distances of zero. finally, the environmental distances were imported back into qgis, and linked to the random points shapefile. then, a series of analyses was developed to explore patterns further. first, the whole country was considered; then it was partitioned into four cells, and then 16 cells (table 1). at each resolution, samples were counted within the cells. all of this information was submitted to the estimates software version 9.1.0. (colwell, 2013) to calculate a series of indices. table 1. sizes and number of cells covering côte d’ivoire at different spatial resolutions. scales cell area (km2) number of cells with ≤1 record c1 322,462 1 c4 80,615.5 4 c16 20,153.875 16 ivorian territory was divided into 64 finer resolution (0.75°) cells. we counted numbers of data points in each cell using the “count points in polygon” function in qgis. we used the index of evenness or regularity to describe the distribution of numbers of individuals between different cells (wala et al. 2005). this index of evenness (pielou 1975) is: 𝐸 = − 𝑝𝑖log(𝑝𝑖 ! !!! ) log(𝑛) where pi is the proportional frequency of cell of each number and n is the number of subdivisions. the index ranges from 0 to 1, tending to 0 when one value dominates, and to 1 when values are, evenly distributed. a quadrat method was used to characterize the spatial pattern with the index of dispersion 𝑥! =   𝑛𝑖 − 𝜇 2! !!! /𝜇, with 𝜇 =   𝑛𝑖 /𝑚, where ni is the proportional frequency of cells of each number and n is the number of subdivisions, m is the number of cells contained entirely within the territory and 𝜇 the mean. the test is two tailed: large values of 𝑥! indicate aggregation or heterogeneity, whereas small values of 𝑥! indicate regularity. results this study is based on herbarium data that cover the period 1894-2000, including work by 226 collectors. oldest specimens were collected by pobégui and jolly, and the most recent were from 2000 (jongkind, hawthorne, assi jean, aman kadjo and others). in general, we observed three major periods (figure 1). during 18941905, data accumulation was slow; it took off during 1906-1994, when it reached a plateau that lasted to the end of the database period in 2000. biodiversity informatics, 10, 2015, 56-64   59     figure 1. frequency of botanical investigations in côte d’ivoire over more than a century shown as the cumulative number of specimens in the sig ivoire database.           figure 2. frequency of collection of the plants of côte d’ivoire by month of collection as represented in the sig ivoire database. 0   2000   4000   6000   8000   10000   12000   14000   16000   18000   20000   22000   18 94   18 96   18 98   19 00   19 02   19 04   19 06   19 09   19 17   19 25   19 30   19 32   19 34   19 36   19 38   19 42   19 45   19 47   19 49   19 51   19 53   19 55   19 57   19 59   19 61   19 63   19 65   19 67   19 69   19 71   19 73   19 75   19 77   19 79   19 81   19 83   19 85   19 87   19 89   19 91   19 93   19 95   19 97   20 00   c um ul at iv e nu m be r of sa m pl es year 0   1000   2000   3000   1   2   3   4   5   6   7   8   9   10   11   12   n um be r of sa m pl es month of the year 0-1000 1000-2000 2000-3000 biodiversity informatics, 10, 2015, 56-64   60   months when collection was most intense were october and november (figure 2). the database for this study comprised 15.2 records of 3621 species in 1371 genera and 198 families. ten species were represented by 20-35 records, 31 species by 16-20 records, 197 species by 11-15 records, 718 species 6-10 records, and 2665 species by 1-5 records. the best represented species were andropogon gayanus kunth (poaceae), culcasia scandens p.beauv. (araceae), rauvolfia vomitoria afzel (apocynaceae) and voacanga africana stapf ex scott-elliot (apocynaceae). the best represented genera were cyperus l. (cyperaceae), ficus l. (moraceae) and indigofera l. (leguminosae); the best represented families were leguminosae with 1610 records (10.6%), poaceae with 1531 records (10.1%), rubiaceae with 1166 records (7.7%), cyperaceae with 771 records (5.1%) and apocynaceae with 710 records (4.7%). the sampling distribution map (figure 3) shows botanical exploration across côte d’ivoire. exploration of the ivorian territory is not uniform. some areas were more densely surveyed, while others have seen partial surveys only or none at all. the equitability index (e) had a value of 0.73, calculated across cells, confirming the irregularity of botanical exploration across côte d’ivoire. also, the value of the index of dispersion (x2 = 41419; x2 m-1 = 44.99) shows considerable aggregation.       figure 3. intensity of botanical exploration across côte d’ivoire, showing points from which data records exist in the sig ivoire database and the aggregation squares that cover the country.   figure 4. summary of floristically well-known areas across côte d'ivoire (in gray). because c values showed that few or no cells were completely inventoried, we used density of records instead. here, the general pattern of only 10% of the cells could be called floristically wellknown (figure 4). these cells correspond to the regions of the grands ponts, agnéby-tiassa, loh-djiboua, gbèkè and parts of boukani, san pedro, and cavally. the best-known regions floristically were grands ponts and agnébytiassa. on the other hand, 90% of cells were not densely sampled. figure 5 shows the relationship between sample size and number of species, showing that most grid squares hold only single records of a species. figure 6 shows that ice varies with spatial resolution, whereas the chao index is constant regardless of spatial resolution and proves relatively invariant. geographically, the well-sampled regions were concentrated in the south-east. geographic distance from these regions accumulated over space, and was highest in the tomkpi region. environmentally, difference from environments of well-sampled areas (figure 7), was relatively low in the south-east and north-east, but the montane region in western côte d'ivoire, was quite different from well-sampled environments. discussion the flora of côte d’ivoire ranks among the best-known floras of west africa (aké-assi 2001). several studies in the country allow an idea of the degree of exploration achieved. botanical investigations in côte d'ivoire began in 1882 (aké-assi 2001), and accumulated slowly biodiversity informatics, 10, 2015, 56-64   61   in the early years. twentieth century botanical explorations reflect the beginning of penetration of europeans into african forests (schenel 1950).     figure 5. relationship between between sample size (n) and observed number of species in a grid square (sobs). the work of one of the explorers, auguste chevalier, provided the foundations of botanical knowledge of the country (schenel 1950; akéassi 2001). he worked in the country in 19061907 (chevalier 1908). in 1932, aubréville made important collections and new discoveries by visiting tai at tabou. in 1950, the team of mangenot began important investigations in the south-west parts of the country, surveys that continued over succeding decades. adjanohoun and guillaumet realized important work on the flora and vegetation of the area (adjanohoun and guillaumet 1961; guillaumet 1967; guillaumet and adjanohoun 1968, 1971), and in 1975, ake assi dreaw up a list of plants in taï national park (aké-assi and pfeffer 1975). except for work by van rompaey (1993), botanical work in the country became less frequent after that point; a few inventories focused on forest fragmentation problems in the taï area (bakayoko 2005; chatelain et al. 2010; martin 2010). botanical exploration is not uniform across ivorian territory, as some regions have been explored more than others. regions where investigations were intense were concentrated along the coast, near activity centers, research institutions and universities, classified forests, natural parks, and reserves. outside these areas, however, large and diverse areas of natural vegetation remain unexplored or are only partially documented. these findings are consistent with those of hepper (1979) and koffi (2008), who indicated that well-known areas are few across africa, and that moderately and poorly known areas cover broad area. the most unique and unsampled environmental were concentrated in western côte d'ivoire. this area comprises one of the few true montane areas of west africa, with mount nimba, which rises to an elevation of 1752 m above a panorama of undulating forested plains. in sum, after more than a century of botanical research in côte d'ivoire, much effort is still needed. certainly, the work of several prominent researchers has produced a view of the basic dimensions and characteristics of the flora and vegetation of côte d’ivoire. however, many regions of côte d'ivoire are not known floristically or remain only partially documented. this study provides detailed analyses and mapping efforts that should guide new botanical investigations in côte d'ivoire. this study also underlines the importance of continued systematic study of plant diversity as a priority in the botanical gardens, universities, and other research organizations. figure 6. summary of diverse inventory statistics and their relationship to spatial resolution of the aggregation grid used across côte d’ivoire. y  =  1.2274x  -­‐  19.377   r²  =  0.99657   0   200   400   600   800   1000   1200   1400   1600   1800   2000   0   500   1000   1500   2000   n   sobs   0   5000   10000   15000   20000   individuals   s  (est)  analyacal   s  mean  (runs)   ace  mean   ice  mean   chao  1  mean   chao  2  mean   jack  1  mean   jack  2  mean   bootstrap  mean   simpson  mean   c16   c4   c1   biodiversity informatics, 10, 2015, 56-64   62     figure 7. map of distances in environmental space to well-inventoried grid cells across côte d'ivoire. acknowledgments we would like to thank the swiss center for scientific research in côte d'ivoire (csrs) for the availability of the database for this study. our thanks also go to the jrs biodiversity foundation for sponsoring our participation in biodiversity informatics training curriculum in entebbe (uganda), in january 2015. references adjanohoun, e. and j. l. guillaumet. 1961. etude botanique entre bas-sassandra et bas-cavally. orstom, adiopodoumé. côte d’ivoire. adou yao, c.y. and e. k. n'guessan. 2005. diversité botanique dans le sud du parc national de taï, côte d'ivoire. afrique science, 1: 295 313. aké-assi l. 1961. contribution à l’étude floristique de la côte d’ivoire et des territoires limitrophes. thèse unique. paris. 205 pp. aké-assi l. 1962. contribution à l’étude floristique de la côte d’ivoire et des territoires limitrophes. vol ii : les monocotylédones et ptéridophytes. thèse unique. paris. 147 pp. aké-assi l. 1976. esquisse de la flore générale de côte d’ivoire. boissiera 24: 543-549. aké-assi l. 1984. flore de la côte d’ivoire: étude descriptive et biogéographique avec quelques notes ethnobotaniques. thèse de doctorat d’état, faculté de sciences et techniques, université de cocody, abidjan (côte d’ivoire), 1206 p. aké-assi l. 1998. impact de l'exploitation forestière et du développement agricole sur la conservation de la biodiversité biologique en côte d'ivoire. le flamboyant, 46: 20-21. aké-assi l. 2001. flore de la côte d’ivoire 1, catalogue, systématique, biogéographie et écologie. genève, suisse: conservatoire et jardin botanique de genève (suisse); boisseria 57, 396 p. aké-assi l. 2002. flore de la côte d’ivoire 2, catalogue, systématique, biogéographie et écologie. genève, suisse: conservatoire et jardin biodiversity informatics, 10, 2015, 56-64   63   botanique de genève (suisse); boisseria 58, 441 p. aké assi, l. and p. pfeffer. 1975. etude d'aménagement touristique du parc national de taï. tome 2: inventaire de la flore et de la faune. bdpa, paris. france. aubert, g. and p. segalen. 1966. projet de classification des sols ferrallitiques. cahiers orstom-pédologie, 4: 97-112. bakayoko, a. 1999. comparaison de la composition floristique et de la structure forestière de parcelles de la forêt classée de bossématié, dans l’est de la côte d’ivoire. mémoire dea., ufr biosciences, université de cocody-abidjan, 72 p. bakayoko, a. 2005. influence de la fragmentation forestière sur la composition floristique et la structure végétale dans le sud-ouest de la côte d'ivoire. thèse de doctorat, université d’abidjan. bänninger, v. 1995. inventaire floristique des dicotylédones de la réserve de lamto (v baoulé) en côte d’ivoire central. diplôme, université de genève, 110 p. brown, j. h. and b. a. maurer. 1989. mac-roecology: the division of food and space among species on continents. science 243: 1145–1150. brown, j. h. and m. v. lomolino. 1998. biogeography. 2nd ed. sunderland: sinauer. massachusetts (sinauer associates, inc. publishers). chatelain, c., kadjo, b., koné, i. and j. refisch. 2000. relations faune flore dans le parc national de taï: une étude bibliographique. tropenbos côte d’ivoire. chatelain, c., bakayoko, a., martin, p. and l. gautier. 2010. monitoring tropical forest fragmentation in the zagne-tai area (west of tai national park, cote d'ivoire). biodiversity and conservation. 19: 2405–2420. chevalier, a. 1908. la forêt vierge de la côte d'ivoire. la géographie 17: 201-210. colwell, r. k. 2013. estimates: statistical estimation of species richness and shared species from samples. version 9. persistent url . corthay, r. 1996. analyse floristique de la forêt sempervirente de yapo (côte d'ivoire). université de genève. de gouvenain, r. c. and j. a. j. r. silander. 2003. do tropical storm regimes influence the structure of tropical lowland rain forests? biotropica, 35:166 180. dotia, y. p. 1999. arbres, arbustes et lianes ligneuses de la commune de kouto, nord de la côte d’ivoire. dea, université de cocody-abidjan. eldin, m. 1971. le climat de la côte d’ivoire. pp. 73108 in avenard j.m., eldin e., girard g., sircoulon j., touchebeuf p., guillaumet j.l., adjanohoun e. and a. perraud. ed. 1971. le milieu naturel de côte d’ivoire. mémoires orstom n° 50, paris (france). gautier, l., aké assi, l., chatelain, c. and r. spichiger. 1999. ivoire: a geographic information system for biodiversity managment in ivory coast. pp. 183-194 in timberlake, j. and s. kativus. ed. african plants: biodiversity taxonomy and uses. royal botanic gardens, kew. gentry, a. h. 1988. tree species richness of upper amazonian forests. proceeding of the national academy of sciences, usa, 85: 156-159. glenday, j. 2008. carbon storage and carbon emission offset potential in an african riverine forest, the lower tana river forests, kenya. journal of east african natural history, 97: 207 223. guillaumet, j. l. 1967. recherches sur la flore et la végétation du bas-cavally (côte d'ivoire). mémoires orstom n°20. guillaumet, j. l. and e. adjanohoun. 1968. carte de la végétation de la côte d'ivoire au 1/500.000. orstom. guillaumet, j. l. and e. adjanohoun. 1971. la végétation de la cote d’ivoire. pp. 161-263 in avenard j.m., eldin e., girard g., sircoulon j., touchebeuf p., guillaumet j.l., adjanohoun e. and a. perraud. ed. le milieu naturel de côte d’ivoire. mémoires orstom n° 50, paris (france). hepper, n. f. 1979. deuxième édition de la carte du degré d’exploration floristique de l’afrique au sud du sahara. pp. 157-162 in g. kunkel ed. proceedings of the ixth plenary meeting of the association pour l'etude taxonomique de la flore d'afrique tropicale (aetfat). las palmas de gran canaria. hijmans, r. j., cameron, s. e., parra, j. l., jones, p. g. and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25: 19651978. jørgensen, i. 2006. forestry in africa: a role for donors? international forestry review, 8: 178 182. kelatwang, s. and m. garzuglia. 2006. changes in forest area in africa 1990 2005. international forestry review, 8: 21 30. koffi, k. j., d. champluvier, e. robbrecht, m. el bana, r. rousseau and j. bogaert. 2008. acanthaceae species as potential indicators of phytogeographic territories in central africa. in: a. dupont & h. jacobs (eds.) landscape ecology research trends: 165-182. nova science publishers, isbn 978-1-60456-672-7. kouakou, n. 1989. contribution à l'étude de la régénération naturelle dans les trouées de l'exploitation en forêt de taï (côte d'ivoire): approches écologiques et phytosociologiques. faculté des sciences et technique. université d’abidjan. kouamé, n. f. 1993. contribution au recensement des monocotylédones de la réserve de lamto (côte d’ivoire centrale) et à la connaissance de leur biodiversity informatics, 10, 2015, 56-64   64   place dans les différents faciès savaniens. mémoire de dea. université de cocody-abidjan, 128 p. kouamé, n. f. 1998, influence de l'exploitation forestière sur la végétation et la flore de la forêt classée du haut sassandra (centre-ouest de la côte d'ivoire). ufr biosc., univ. cocody abidjan. kowero, g., kufakwandi, f. & chipeta, m. 2006. africa's capacity to manage its forests: an overview. international forestry review, 8: 110117. lehmann, i. and e. kioko. 2005. lepidoptera diversity, floristic composition and structure of three kaya forests on the south coast of kenya. journal of east african natural history, 94: 121 163. mangenot, g. 1956. etude sur la forêt des plaines et plateaux de côte d'ivoire. etudes éburnéennes. ifan, dakar, 4: 55 67. martin, p. j. a. 2010. influence de la fragmentation forestière sur la régénération des espèces arborées dans le sud-ouest de la côte d'ivoire. thèse de doctorat, université de genève. 306 p. menzies, a. 2000. structure et composition floristique de la forêt de la zone ouest du parc national de taï (côte d’ivoire). diplôme, université de génèse, 124 p. nair, c. t. s. 2006. what is the future for african forests and forestry? international forestry review, 8: 4 13. nusbaumer, l. 2003. structure et composition floristique de la forêt classée du scio (côte d'ivoire): étude descriptive et comparative. université de genève, suisse. 150 p. pielou, e. c. 1975. ecological diversity. new york, john wiley & sons. 165 p. mariane, s. s. b., letıcia, c. g. and t. p. andrew. 2013. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. diversity and distributions, 20: 1-13. schnell, r. 1950. etat actuel des recherches sur la végétation de l'afrique intertropicale française. vegetation, 1.2: 331-340. singh, k. d. 1993. l’évaluation des ressources forestières en 1990. unsaylva 174 (44): 10-20. van rompaey, r. s. a. r. 1993. forest gradients in west africa: a spatial gradient analysis. thesis wageningen. 142 p. venables, w. n. and d. m. smith. 2014. an introduction to r: a programming environment for data analysis and graphics version 3.2.2. the r development core team. r foundation for statistical computing, vienna, au. u.r.l.: http://www.r-project.org/. wala, k., sinsin, b., guelly, k. a., kokou, k. and k. akpagana. 2005. typologie et structure des parcs agroforestiers dans la préfecture de doufelegou (togo). sécheresse, 16 (3): 209-216. whittaker, r. j., araùjo, m. b., paul, j., ladle, r. j., watson, j. e. m. and k. j. willis. 2005. conservation biogeography: assessment and prospect. diversity and distributions 11: 3–23. biodiversity informatics, training modules, 13, 2018, pp. 49-50 49 sample data and training modules for cleaning biodiversity information marlon e. cobos, laura jiménez, claudia nuñez-penichet, daniel romero-alvarez, marianna simões department of ecology and evolutionary biology and biodiversity institute, university of kansas, lawrence, ks, usa abstract.—large-scale biodiversity databases have become crucial information sources in many analyses in biogeography, macroecology, and conservation biology, often involving development of empirical models of species’ ecological niches and predictions of their geographic distributions. these analyses, however, can be impaired by the presence of errors, particularly as regards taxonomic identifications and accurate geographic coordinates. here, we present an introductory data-cleaning exercise based on two contrasting datasets; we link these example data with a step-by-step guide to overcoming these problems and improving data quality for analyses based on these data. key words: accuracy, data cleaning, error, redundancy, precision, primary biodiversity data the availability and management of large online biodiversity databases has become an exciting, although challenging, step in the development of an authoritative and comprehensive basis for biodiversity knowledge. thanks to improvements in technology and data collection, steps that previously took years can now be accomplished in seconds (robin 2012; musa et al. 2013). the temptation to use this information and interpret analyses quickly, however, often undermines the fundamental necessity of checking data and addressing data quality carefully. this tension is a perpetual caveat to use of modern, open datasets, and has been particularly problematic as regards use of georeferenced, primary biodiversity data. still, these data are the building blocks for many interesting analyses, perhaps most prominently, species distribution and ecological niche models (peterson et al. 2011; anderson 2015). primary biodiversity data have become much more accessible in recent decades, with a major transition from recalcitrance (graves 2000) to enthusiasm. many institutions (e.g., museums, herbaria, observational data initiatives) have digitized data associated with their work, and increasing numbers now make these data available via the internet (soberón and peterson 2004). the global biodiversity information facility (gbif), the botanical information and ecology network (bien), and the distributed information system for biological collections (specieslink), are a few examples of repositories that provide online open access biodiversity information, now providing access to over a billion individual records. the benefit of these initiatives is clear; for instance, in 2016, 438 articles were published using data from gbif alone in multiple fields including data management, evolution, biogeography, biodiversity, and public health (gbif 2018). however, data quantity is compromised by frequent low data quality; occurrence data from these sources suffer from diverse errors that should be identified, assessed, and minimized before performing any analysis. low accuracy, low precision, occurrences from outside species’ ranges, abundant duplicate records, and taxonomic misidentifications, are just a few of the common errors found in these data (rahm and do 2000). to present a means of building and assessing capacity to handle such problems, we provide a handson exercise for data cleaning, with two worked examples, coupled with recommendations and suggestions on how to identify and overcome the most common errors. the goal, of course, is to obtain a cleaned database that will be robust and reliable in different downstream analyses (peterson et al. 2011). we have assembled two example datasets, one small (960 records) and one large (36,574 records), using records from gbif—in each case, we have cleaned the data extensively, and then re-populated the dataset with typical classes of errors—we then provide a step-bystep guide to the process of cleaning and improving them. the databases and the data-cleaning manual are available at: http://hdl.handle.net/1808/26512. these data files and the associated manual are not intended as a detailed treatment of biodiversity data quality, or to offer automated methods with which to solve these problems (chapman 2005; hijmans and http://hdl.handle.net/1808/26512 marlon e. cobos et al. – training modules 50 elith 2013; maldonado et al. 2015). rather, we focus on individualized, hands-on, user assessment of occurrence datasets, and providing detailed training materials with which to build capacity to make such assessments possible. this publication can be used as a tool for educational purposes, as an exercise for a broader audience that is starting to explore the field, and/or as a reminder of this often-overlooked step (anderson 2015). we are aware of the time-consuming nature of manual data cleaning, but, to avoid a ‘garbage in, garbage out’ situation, high-quality data should always be the goal, to allow researchers to develop informative experiments and operational models. references anderson, r. p. 2015. el modelado de nichos y distribuciones: no es simplemente clic, clic, clic. biogeografía 8:4–27. chapman, a. d. 2005. principles and methods of data cleaning. gbif, copenhagen. graves, g. r. 2000. costs and benefits of web access to museum data. trends ecol. evol. 15, 374. hijmans, r. j., and j. elith. 2017. species distribution modeling with r. available at cran.r-project. org/web/packages/dismo/vignettes/sdm.pdf. accessed 8 march 2018. maldonado, c., c. i. molina, a. zizka, c. persson, c. m. taylor, j. albán, e. chilquillo, n. rønsted, and a. antonelli. 2015. estimating species diversity and distribution in the era of big data: to what extent can we trust public databases? glob. ecol. biogeogr. 24:973–984. musa, g. j., p.-h. chiang, t. sylk, r. bavley, w. keating, b. lakew, h.-c. tsou, and c. w. hoven. 2013. use of gis mapping as a public health tool—from cholera to cancer. heal. serv. insights. 6:111–116. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. rahm, e., and h. h. do. 2000. data cleaning: problems and current approaches. ieee data eng. bull., 23(4): 3–13. robin w. 2012. robin’s blog. john snow’s famous cholera analysis data in modern gis formats. available at http://www.blog.rtwilson.com/johnsnows-famous-cholera-analysis-data-in-moderngis-formats. accessed march 8 2018. soberón, j., and a. t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. philos. trans. r. soc. lond. 359:689–698. http://www.blog.rtwilson.com/john-snows-famous-cholera-analysis-data-in-modern-gis-formats http://www.blog.rtwilson.com/john-snows-famous-cholera-analysis-data-in-modern-gis-formats http://www.blog.rtwilson.com/john-snows-famous-cholera-analysis-data-in-modern-gis-formats microsoft word m_cc_2005_final.doc biodiversity informatics, 2, 2005, pp. 42-55 climate change and biodiversity: some considerations in forecasting shifts in species’ potential distributions enrique martínez-meyer instituto de biología, universidad nacional autónoma de méxico tercer circuito exterior s/n, ciudad universitaria, méxico city 04510 emm@ibiologia.unam.mx global climate change and its broad spectrum of effects on human and natural systems has become a central research topic in recent years; biodiversity informatics tools—particularly ecological niche modeling (enm)—have been used extensively to anticipate potential effects on geographic distributions of species. misuse of these tools, however, is counterproductive, as biased conclusions might be reached. in this paper, i discuss some issues related to niche theory, geographic distributions, data quality, and algorithms, all of which are relevant when using enm in climate change projections for biodiversity. this assortment of opinions and ideas is presented in the hope that enm applications to climate change questions can be made more realistic and more predictive. key words: climate change, ecological niche, niche modeling, geographic distribution, bioclamatic envelope the large-scale climatic changes observed and documented since mid-twentieth century represents a major and growing concern for the academic community. considerable human and economical resources have been directed at understanding these phenomena and their possible consequences for humans and the natural world (houghton et al. 1990; ipcc 1995; watson 2001). one dimension in which climate change effects on natural systems has been studied is in understanding its implications for potential geographic distributions of species. a common rule of thumb has been that species’ potential distributional areas will likely shift poleward, as well as upward in elevation in areas of topographic relief. although a useful generality, the predictive power of such sweeping statements is low; as such, the need for tools that produce species-specific, predictive tools regarding the details of climate change effects on species’ distributions became clear. ecological niche modeling (enm), also known as bioclimatic modeling or climate envelope modeling, has been applied increasingly to this task. this approach uses georeferenced primary occurrence data for species, in combination with digital maps representing environmental parameters, to build models of the ecological requirements of species—the set of conditions suitable and necessary for long-term survival of populations of the species without immigrational input. then, such conditions are located on landscapes, and maps created to indicate the distributional potential of the species (pearson and dawson 2003; peterson et al. 2001; thuiller 2003). with this approach, distributional shifts caused by climatic change, in both the past and the future, can be estimated based on the fact that the niche model is characterized in ecological space—conditions with which a species is associated at present can be sought on modeled future or past climate scenarios (fig. 1) (hugall et al. 2002; martinez-meyer et al. 2004; meynecke 2004). the simplicity of the approach, improved availability of relevant data and software, and the importance of the topic have increased considerably the number of studies aiming to estimate effects of future climatic regimes on species’ distributions. for instance, several studies have projected future potential distributions of species for conservation purposes (aspinall and matthews 1994; bakkenes et al. 2002; beaumont and hughes 2002; burns et al. 2003; dunbar 1998; erasmus et al. 2002; iverson and prasad 2001; meynecke 2004; midgley et al. 2002; midgley et al. 2003; ortega-huerta and peterson 2004; peterson et al. 2002; skov and svenning 2004; tellez-valdes and davila42 martínez-meyer – climate change and ecological niche modeling figure 1. diagrammatic summary of the ecological niche modeling process (see text for details). 43 martínez-meyer – climate change and ecological niche modeling aranda 2003; thomas et al. 2004; williams et al. 2003), resource management (bradshaw et al. 1992; clark et al. 2001; holden et al. 2003; loukos et al. 2003; mati 2000; rafoss and saethre 2003; schwartz et al. 2001; sykes 2001; van staden et al. 2004; williams and liebhold 2002), public health (ando 1994; craig et al. 1999; peterson and shaw 2003; wittmann et al. 2001), and invasive species (kriticos et al. 2003). in many cases, however, insufficient attention has been paid to the limitations of the approach and the data, and analyses rely on untenable assumptions that may bias conclusions (thomas et al. 2004; thuiller 2003; thuiller et al. 2004b). in this paper, i present and discuss what i consider the most critical issues—both conceptual and operational—that should be taken into account when using enm to anticipate climate change effects on geographic distributions of species. conceptual considerations enm depends conceptually on the theory of the niche, developed in the early 20th century (soberón and peterson 2005). the contribution of niche theory to biogeography is highly relevant, since it provides key conceptual elements to understand why species have limited ranges, and how abiotic and biotic factors interact to mold species’ geographic distributions (macarthur 1972). an understanding of these concepts is necessary for adequate and appropriate application of enm tools to produce future potential distributional maps, and for correct interpretation of the results. although several often disparate niche definitions have been proposed through the years (elton 1927; grinnell 1917; hutchinson 1957), in the context of biogeography, a generally accepted definition refers to the set of environmental conditions—biotic and abiotic—under which populations of a species can survive indefinitely without immigration (grinnell 1917; hutchinson 1957). according to hutchinson (1957), measurements of a population’s performance along an environmental gradient (e.g., temperature) can be used to define tolerance limits along that axis—i.e., the niche limits. extending this notion to two dimensions (e.g., temperature and humidity), the niche takes the form of a polygon; considering multiple (n) axes, we can envision an n-dimensional hypervolume encompassing the set of appropriate conditions. hutchinson termed this hypervolume the fundamental niche—the full range of possibilities -in ecological space where the population can persist. in most cases, though, negative biotic interactions (chiefly competition and predation) prevent species from occupying the entirety of the fundamental niche; according to the theory, the portion of the fundamental niche actually occupied by the species in geographic space was called the realized niche (hutchinson 1957). hence realized niches are subsets of fundamental niches. from these ideas, several issues relevant to enm and studies of climate change can be identified. soberón and peterson (2005) developed a conceptual framework to answer a fundamental question—what precisely are we modeling with enm tools? as they described, differences exist between fundamental niches, realized niches, and geographic distributions, with each one implying a distinct set of influencing factors. according to the soberón and peterson framework, in a modeling exercise, depending on the nature of the input occurrence data (presence-only versus presence/absence data; source [within-the-niche] versus sink [outsidethe-niche] populations), the model may reflect the fundamental niche, realized niche, or actual geographic distribution of the species. in general, these authors argue that in most cases enm summarizes something more towards the fundamental niche, which explains the fact that some (or a lot) overprediction is generally contained in resulting maps. clearly, previous manifestations of niche theory are insufficient to explain geographic distributions of species for two reasons. (1) niche theory simply does not account for influences of historical factors, such as biogeographic barriers and life history traits of species such as dispersal capacity, which are manifested exclusively in geographical, rather than ecological, space. (2) although interactions may take place in ecological space (e.g., species a is able to survive at some temperatures only because its competitor species b cannot), they may also exist exclusively in geographic domains (e.g., under identical environmental conditions, species a is absent from site 1 because 44 martínez-meyer – climate change and ecological niche modeling species b is present, but species a is present at site 2 because species b has never been there for historical reasons), or in both realms. therefore, finding in the field a stable population of a species likely means that that place presents suitable environmental conditions—biotic and abiotic—and thus lies within the species’ realized niche. hence, neglecting the possibility of sink populations biasing model results, occurrence points can be assumed to represent within-the-niche populations. in such cases, maps resulting that show true overprediction (i.e., not owing to a poorly specified model), would represent a projection of the realized ecological niche model of the species onto an unconstrained fundamental geographic space: a landscape in which historical factors and/or interactions acting in the geographical domain are not reflected. put another way, this map would identify the potential geographic distribution of the species. if the additional factors of history, limited dispersal ability, and biotic interactions were incorporated into the modeling process, then we would be truly modeling the realized ecological niche projected onto its realized geographical space—in other words, the actual geographic distribution of the species. currently, to approximate actual geographic distribution maps from potential distributional models, post hoc procedures have sometimes been used. for example, electronic maps of biogeographic provinces, ecoregions, or vegetation types can be used as “cookiecutters” to trim potential distribution maps under the assumption that the species’ distribution is well-sampled at the level of ecoregions, and that disjunctions among such regions represent distributional barriers for the target species (anderson and martínez-meyer 2004). in studies in which the goal is to anticipate geographic shifts in species’ distributions resulting from climate change, additional problems arise. in the best-case scenario, a species’ present-day (modeled) potential distribution map resembles fairly well its known geographic distribution; in this case, an investigator may assume that environmental factors (rather than biotic or historical factors) are most influential in molding the species’ distribution. when this model is projected onto future climate scenario, the resulting map identifies areas likely to become habitable or uninhabitable for that species. several authors have argued that this sort of result should be taken cautiously, since biotic interactions may shift in the face of changing conditions, and may thus affect the distributional potential of species (davis et al. 1998a; davis et al. 1998b; pearson and dawson 2003). in situations in which a potential distribution map for present-day conditions shows broad areas of overprediction, biogeographic barriers to dispersal and colonization, biotic interactions, or both have an important influence on the geographic distribution of the species. such potential distribution maps would need to be reduced to the species’ actual geographic distribution prior to interpretation, based on explicit assumptions regarding which sorts of areas are likely to represent actual distributional areas as opposed to potential distributional areas (fig. 2). projecting such models to future scenarios poses important complications. if potential distribution maps are not processed post hoc into actual distribution maps, then establishing which habitable areas are likely to be reachable by the species is difficult, and projections to future conditions may become totally unrealistic (fig. 2). if the same scheme as for present-day maps is followed, and distributions are adjusted via maps representing biogeographic barriers, the outcome may be misleading because the barriers may also shift under effects of climate change. in other words, what is today a barrier for a species may not be a barrier in the future, because ecological differences become less abrupt and make previously inhabitable areas habitable for the species (thomas et al. 2001). sadly, no straightforward solution to this problem exists. at present, projecting niche models onto changed-climate landscapes produces expected suitability maps. when the goal is to produce plausible future distributional scenarios for species, understanding and incorporating dispersal considerations becomes critical (morecroft et al. 2002; svenning and skov 2004; thuiller et al. 2004a; travis and dytham 2002). previous studies have shown that dispersal patterns are species-specific, and also depend on the geography of the area, which can make predictions difficult (clark et al. 2003; davis 45 martínez-meyer – climate change and ecological niche modeling figure 2. distribution model of the volcano rabbit (romerolagus diazi) in central mexico, showing raw predictions for (a) the present and (b) in 2050. overprediction (commission error) is reduced with a post hoc clipping process both in the present (c) and the future (d) using a map of the ecoregions (black lines), assumed to represent biogeographical barriers for the species. gray areas represent the prediction for the present and blue areas are the modeled projection to 2050 under a future-climate scenario. a b c d 46 martínez-meyer – climate change and ecological niche modeling et al. 1998b; dullinger et al. 2004; post and forchhammer 2002; travis and dytham 2002). as a result, some authors do not attempt to present definite future distributional predictions, but rather a set of possible scenarios based on different dispersal assumptions, ranging from no dispersal to universal dispersal (peterson et al. 2001). in this way, one can at least bracket possible effects of climate change on species’ distributions. recent findings and new modeling tools regarding dispersal of species, however, are opening opportunities to incorporate dispersal information into predictive modeling to produce more realistic future distributional scenarios (clark et al. 2003; collingham and huntley 2000; nagelkerke and alkemade 2003; nathan et al. 2002). finally, even resolving issues of interactions, barriers, and dispersal, some uncertainty remains in future projections as regards the ability of species to survive in novel environments. as climates change, environmental regimes novel to populations emerge; populations will be able to survive under these conditions if (1) those conditions are within their fundamental niches (vetaas 2002); or (when such is not the case) (2) populations are able to adapt to conditions outside their present niches (holt 1990; holt and gaines 1992). in the first case, projections to future scenarios may fail to predict areas within the fundamental niche but not represented in the modeled realized niche; as such, the models may underestimate potential future distributions of species. in the second situation, models will also underrepresent potential distributional areas, given the species’ ability to adapt to new conditions; however, theoretical and empirical evidence strongly suggests that adaptation is less frequent than migration (holt 1996; hughes 2000; martínez-meyer et al. 2004). operational considerations modeling ecological niches and predicting geographic distributions of species frequently becomes complicated because input data are frequently far from ideal. as mentioned above, two main data streams are employed: (1) locality data documenting known occurrences of the species (some algorithms incorporate absence data as well), and (2) environmental data in the form of raster gis maps, normally including some combination of climatic, topographic, and land-cover data. in climate change studies, parallel environmental variables must be available for the future (or the past; fig. 1). in the following section, i discuss some important issues regarding data quality, both biological and geographic, that should be taken into account in this sort of analysis. biological data studies of climate change and biodiversity are frequently based on incomplete biological data; indeed, enm emerged as a solution to this challenge. perhaps the most frequent question from new workers in this field refers to minimum numbers of occurrence points needed to generate robust models. sadly, no easy answer to this question exists. controlled experiments and analyses have shown that minimum numbers of points required for modeling depends on the species in question, as well as on the number of geographic variables in the analysis, and the algorithm used (kadmon et al. 2003; stockwell and peterson 2002). in general, a rule of thumb is that the more occurrence data are available, the better, but even this seemingly obvious point is not necessarily true. actually, the distribution of occurrence data in ecological space becomes much more important than the overall density of points, particularly for generalist, widespread species that likely have broad ecological niches (kadmon et al. 2003; thuiller et al. 2004b). when the distribution of occurrence points in geographic space is too biased (i.e., observations are clustered in small regions of the species’ range), biases appear in ecological space as well, and models can misrepresent the species’ ecological requirements (fig. 3). in this case, systematic removal of points to “balance” their spatial distribution have resulted effective (hidalgomihart et al. 2004). although methods for this manipulation need exploration and standardization, a useful approach has been to overlay a reticule on the study area, and randomly select one point per cell (hidalgomihart et al. 2004). cell size for the reticule has to be large enough to eliminate clustered points, but not as large as to eliminate points from the same cell that represent different environments. 47 martínez-meyer – climate change and ecological niche modeling figure 3. geographic bias in occurrence data available for the mexican wolf (canis lupus baileyi). (a) input points show strong geographic bias towards the center of the distribution of the species, causing a bias in the distribution model. (b) systematic removal of points reduces the bias and produces better distribution models. a b 48 martínez-meyer – climate change and ecological niche modeling not only are the amount and distribution of points important to producing good models, but some species are more difficult to model than others. in general, models for species with narrow ecological niches (e.g., habitat specialists) result better and more predictive than models for species with broad ecological niches (thuiller et al. 2004b). of course, it is always highly recommended to validate model predictions prior to any extrapolation or interpretation (oreskes et al. 1994). validation usually takes the form of challenging models to predict the distribution of a suite of points that were not included in model development—a general, predictive, and usable model will be able to anticipate the distributions of such sets of points. model predictions may also be validated via cross-time tests, which such is feasible (martínez-meyer et al. 2004). different statistics have been used for this purpose, including the receiver operating characteristic (roc) curve, cohen’s kappa statistic, and chi-square tests (anderson et al. 2003; elith and burgman 2002; fleishman et al. 2003; oreskes et al. 1994); detailed reviews of statistical methods can be found elsewhere (fielding and bell 1997; pearce and ferrier 2000; pearce et al. 2001), but a couple of comments will be provided here. first, roc curves and kappa require both presence and absence data, whereas chisquare not. of course, the former methods are more powerful, but interpretation of ‘absence’ information is complex—absences of species will frequently not coincide with absence of appropriate niche conditions, making the meaning of the absence data difficult to interpret (soberón and peterson 2005). in instances in which true absence data are lacking and sampling from non-presence areas is performed to obtain ‘pseudo-absence’ points (engler et al. 2004; hirzel et al. 2001; zaniewski et al. 2002), these same concerns are relevant. furthermore, the independence of data sets for developing models (‘calibration’) and for evaluating them (‘validation’) is critical. different strategies have been followed here, including data partitioning, resubstitution, and independent sampling. in general, obtaining new and independent data directly from the field is preferable. when such new sampling is not feasible, data partitioning seems to provide better results than resubstitution (fielding and bell 1997). when biases exist in particular sampling methods, calibration and validation data sets may both reflect them, and the model may not be generally representative of the distribution of the species. in any case, true independence may not exist in the biological data because of the very nature of species’ distributions. however, this situation overestimates model fit, necessitating development of validation techniques that incorporate effect of spatial autocorrelation (hampe 2004). projections onto scenarios of change over time bring additional complications, because frequently no data are available to validate future predictions. projections to past climates are often the only means of obtaining statistical validation for cross-time predictions (martínez-meyer et al. 2004). in future climate change studies, a minimal step is validation of model predictions under current conditions—because errors are propagated or even exacerbated in projections (thuiller 2003), species for which present-day models are poor should not be used for future projections. geographic data several issues related to environmental datasets used in enm affect model performance in both present and future predictions, including the amount, type, and quality of variables included in analyses. experimental studies have demonstrated that some environmental variables are more informative than others when modeling ecological niches. in general, climatic variables (e.g., maximum and minimum temperatures, precipitation) are particularly useful, as they coincide with physiological tolerances (parra et al. 2004; peterson and cohoon 1999). however, a combination of climatic and topographic features, like elevation, slope, and orientation of slopes, yields better results (parra et al. 2004) because topographic features modify how individual animals or plants experience a particular climate regime. in future projections, elevation certainly is not useful, as certain elevation-temperature associations break down under different climates. not only are the type and number of environmental variables important, but also their quality in terms of resolution (spatial, temporal, and metric) is crucial. today, 49 martínez-meyer – climate change and ecological niche modeling thanks to efforts by individuals and research groups, global, regional, and local climatic and topographic data sets with different spatial and temporal resolutions are freely available (e.g., hydro-1k1, worldclim2; chapman et al. 2005). nonetheless, at least until recently, many regions lacked adequate geographic information (lim et al. 2002). in any case, a thorough understanding of the methods used to derive geographic datasets is highly desirable, since all of them have limitations (chapman et al. 2005), and these limitations will affect modeling outcomes. spatial resolution (pixel size) is of particular importance when interest is focused on species with relatively small ranges or in complex landscapes (chapman et al. 2005). low-resolution maps do not capture such variability and models result too coarse or imprecise (lim et al. 2002). high-resolution datasets may be difficult to generate for some regions because climatic station data are too sparse or are lacking (new et al. 1999). even when high-resolution maps are derived (e.g., worldclim), it is very important to validate their reliability with independent field data or other information (magana et al. 1997) before using them for predicting species’ distributions. in climate change studies, simulated future climates are generally obtained for projections by one of two means: (1) increasing one or more present-day temperature variables (minimum, maximum or mean) by some quantity (e.g., tellezvaldes and davila-aranda 2003), or (2) from data output from general circulation models (gcm) (e.g., peterson et al. 2002). use of gcm results is clearly preferable since climate change involves complex rearrangements of numerous parameters; gcms are currently the best means to account for these complexities (murphy et al. 2004). however, inconsistencies among predictions from different gcms available pose problems for users who may not be able to decide among alternatives. in recent years, research efforts have focused on quantifying uncertainties among models, with the aim of producing more reliable estimations (allen et al. 2000; murphy et al. 2004). regional climate models (rcm) have been developed 1 http://lpdaac.usgs.gov/gtopo30/hydro/. 2 http://biogeo.berkeley.edu/worldclim/worldclim.htm. for some areas as well, which improve spatial resolution, but which bring an additional suite of assumptions and potential complications (maccracken et al. 2004). for now, statistical downscaling of gcm results remains the best option for producing high-resolution scenarios (giorgi et al. 2001). modeling algorithms as with biological and geographic datasets, enm algorithms have seen considerable improvement and development in recent years. currently, many enm algorithms are available as stand-alone software packages (e.g., biomapper, desktopgarp, floramap, bioclim), or are implemented in statistical or gis packages (e.g., grasp). this rapid development has led to a series of studies testing and comparing algorithm performance under diverse circumstances. although results suggest that no single algorithm can be identifying as performing better than all others under all circumstances (brotons et al. 2004; pearson et al. submitted; thuiller et al. 2003), some generalities can be drawn. first, as mentioned above, all approaches face particular problems in dealing with widespread species. this effect results from increased probability of data being lacking or bias in representation of the niche (brotons et al. 2004), or from reduced statistical power (stockwell and peterson 2002). in this case, algorithms able to handle “true” absence data (e.g., generalized linear models, glm; artificial neural networks, ann) seem to perform better than presence-only methods (e.g., ecological niche factor analysis, enfa; bioclim). however, most biological data sources provide only occurrence records; so methods that at least take advantage of “pseudo-absence” information (e.g., genetic algorithm for rule-set prediction, garp) may have an advantage. second, since model performance varies among species, combining results of different methods may be desirable in cross-taxon studies (thuiller 2003). currently, only a few algorithms incorporate multiple methods in predicting distributions, like biodiversity modeling, biomod (thuiller 2003) and garp (stockwell and peters 1999). another useful approach to this challenge may be development of models using several distinct 50 martínez-meyer – climate change and ecological niche modeling algorithms, and combining the results after modeling. finally, projections to future climate scenarios may produce very different results depending on the algorithm used, even when present-day results are very similar (thuiller 2003). these differences arise because different algorithms make different assumptions when extrapolating to future scenarios presenting environmental combinations not found in the present (pearson et al. submitted). theoretical studies have indicated that significant elements in species’ adapting to new conditions are the degree of dissimilarity between conditions within and outside the niche, genetic variation for key traits determining abundance and distribution, dispersal dynamics, etc. (holt and gaines 1992). a fundamental problem is that very little is known about species’ responses to novel environments to permit their incorporation into the modeling process (holt 1990). probably, the most appropriate recommendation here is to evaluate algorithm performance in present-day predictions, to select only those algorithms that perform best, and to follow a ‘consensus’ approach with algorithms used in future projections. this approach at least provides a notion of ranges of distributional possibilities of target species, and strengths and weaknesses of algorithms employed. clearly, more research is needed both in fundamental aspects of the ecology and evolution of species, and in implementation of these findings in enm. conclusions in recent years, climate change has come to rank among the most active research topics in science because of its immediacy, and the profound effects on natural systems and human welfare that are anticipated. rapid development of biodiversity informatics tools has stimulated research focused on forecasting climate change impacts on biodiversity. enm has become particularly important, because it provides one of few predictive approaches to understanding geographic dynamics of species. as more people turn to such tools, it becomes relevant to review key aspects of their use. poor understanding of the enm approach, both in terms of conceptual basis and practical implementation, might lead to inappropriate results or interpretations. the first elements for consideration are the biological, ecological, and geographic characteristics of the species in question. natural history aspects such as niche breath, ecological affinities, and dispersal capacity are crucial to understanding geographic distributions, and to interpreting model results. here, the theory of the niche plays a key role—based on the this conceptual framework, we can affirm that the enm approach does not produce a representation of the geographic distribution of species, but rather an unconstrained geographic projection of the realized niche, estimated in ecological space. other elements that modify species’ geographic distributions from their fundamental potential, e.g., biogeographical barriers and competitors, are not integrated into the enm process. post hoc procedures are generally implemented to convert potential distribution models into actual distribution models if the intention is to produce realistic scenarios of present and future distributions of species. in addition to the conceptual framework, the quality of input data, both biological and geographic, is fundamental. species’ distributions and their responses to climate change processes are idiosyncratic. spatial and metric resolutions of geographic data are decisive in the quality of modeling outcomes. in projections to future scenarios, notwithstanding several complications, gcms are the most reliable source of future-climate information, but downscaling is mandatory to permit regional and local studies. currently, it is impossible to identify any single modeling algorithm that performs better than all others for all types of species and data conditions, and i suspect that use of multiple methods may prove to be the most robust current option. despite all the limitations discussed herein, enm is the best instrument currently available for anticipating effects of climate change on distributions of species. as more workers get involved in this field, i anticipate rigorous and critical evaluation of enm tools, as well as filling key knowledge gaps and incorporating them into the modeling systems. in this way, this approach will become more reliably predictive, and results will be increasingly useful in preservation of both natural and human resources. 51 martínez-meyer – climate change and ecological niche modeling acknowledgments i thank a. t. peterson, l. hannah, r. pearson, m. araujo, p. williams, and s. andelman for useful discussions of climate change and ecological niche modeling, as well as. two anonymous reviewers for their helpful comments. the dirección genearal de asuntos del personal académico of the universidad nacional autónoma de méxico (dgapa-papiit in215102-3) has provided the financial support for my studies on climate change and ecological niche modeling. literature cited allen, m. r., p. a. stott, j. f. b. mitchell, r. schnur, and t. l. delworth. 2000. quantifying the uncertainty in forecasts of anthropogenic climate change. nature 407:617-620. anderson, r. p., d. lew, and a. t. peterson. 2003. evaluating predictive models of species' distributions: criteria for selecting optimal models. ecological modelling 162:211-232. anderson, r. p., and e. martinez-meyer. 2004. modeling species' geographic distributions for preliminary conservation assessments: an implementation with the spiny pocket mice (heteromys) of ecuador. biological conservation 116:167-179. ando, k. 1994. the influences of global environmental changes on nematodes. japanese journal of parasitology 43:477-482. aspinall, r., and k. matthews. 1994. climate change impact on distribution and abundance of wildlife species: an analytical approach using gis. environmental pollution 86:217223. bakkenes, m., j. r. alkemade, f. ihle, r. leemans, and j. b. latour. 2002. assessing effects of forecasted climate change on the diversity and distribution of european higher plants for 2050. global change biology 8:390-407. beaumont, l. j., and l. hughes. 2002. potential changes in the distributions of latitudinally restricted australian butterfly species in response to climate change. global change biology 8:954-971. bradshaw, r. h. w., b. h. holmqvist, s. a. cowling, and m. t. sykes. 1992. the effects of climate change on the distribution and management of picea abies in southern scandinavia. canadian journal of forest research 30:1992-1998. brotons, l., w. thuiller, m. b. araujo, and a. h. hirzel. 2004. presence-absence versus presence-only modelling methods for predicting bird habitat suitability. ecography 27:437-448. burns, c. e., k. m. johnston, and o. j. schmitz. 2003. global climate change and mammalian species diversity in u.s. national parks. proceedings of the national academy of sciences, usa 100:11474-11477. chapman, a. d., m. e. s. muñoz, and i. koch. 2005. environmental information: placing environmental phenomena in an ecological and environmental context. biodiversity informatics 2:24-41. clark, j. s., m. lewis, j. s. mclachlan, and j. hillerislambers. 2003. estimating population spread: what can we forecast and how well? ecology 84:1979-1988. clark, m. e., k. a. rose, d. a. levine, and w. w. hargrove. 2001. predicting climate change effects on appalachian trout: combining gis and individual-based modeling. ecological applications 11:161-178. collingham, y. c., and b. huntley. 2000. impacts of habitat fragmentation and patch size upon migration rates. ecological applications 10:131-144. craig, m. h., r. w. snow, and d. le sueur. 1999. a climate-based distribution model of malaria transmission in sub-saharan africa. parasitology today 15:105-110. davis, a. j., l. s. jenkinson, j. h. lawton, b. shorrocks, and s. wood. 1998a. making mistakes when predicting shifts in species range in response to global warming. nature 391:783-786. davis, a. j., j. h. lawton, b. shorrocks, and l. s. jenkinson. 1998b. individualistic species responses invalidate simple physiological models of community dynamics under global environmental change. journal of animal ecology 67:600-612. dullinger, s., t. dirnboeck, and g. grabherr. 2004. modelling climate change-driven treeline shifts: relative effects of temperature increase, dispersal and invasibility. journal of ecology 92:241-252. dunbar, r. i. m. 1998. impact of global warming on the distribution and survival of the gelada baboon: a modelling approach. global change biology 4:293-304. elith, j., and m. burgman. 2002. predictions and their validation: rare plants in the central highlands, victoria in j. m. scott, p. j. heglund and m. l. morrison, eds. predicting species occurrences: issues of scale and accuracy. island press, washington, d.c. elton, c. s. 1927. animal ecology. sidgwich and jackson, london. engler, r., a. guisan, and l. rechsteiner. 2004. an improved approach for predicting the distribution of rare and endangered species from occurrence and pseudo-absence data. journal of applied ecology 41:263-274. 52 martínez-meyer – climate change and ecological niche modeling erasmus, b. f., a. s. van jaarsveld, s. l. chown, m. kshatriya, and k. j. wessels. 2002. vulnerability of south african animal taxa to climate change. global change biology 8:679-693. fielding, a. h., and j. f. bell. 1997. a review of methods for the assessment of prediction errors in conservation presence/absence models. environmental conservation 24:3849. fleishman, e., r. m. nally, and j. p. fay. 2003. validation tests of predictive models of butterfly occurrence based on environmental variables. conservation biology 17:806-817. giorgi, f., b. hewitson, j. christensen, m. hulme, h. v. storch, p. whetton, r. jones, l. mearns, and c. fu. 2001. regional climate information – evaluation and projections. pp. 585-638 in c. a. johnson, ed. climate change 2001: the scientific basis. contributions of working group i to the third assessment report of the intergovernmental panel on climate change. cambridge university press, new york. grinnell, j. 1917. field tests of theories concerning distributional control. american naturalist 51:115-128. hampe, a. 2004. bioclimate envelope models: what they detect and what they hide. global ecology and biogeography 13:469-471. hidalgo-mihart, m. g., l. cantú-salazar, a. gonzález-romero, and c. a. lópezgonzález. 2004. historical and present distribution of coyote (canis latrans) in mexico and central america. journal of biogeography 31:2025-2038. hirzel, a..h., v. helfer, and f. métral. 2001. assessing habitat-suitability models with a virtual species. ecological modelling 145:111-121. holden, n. m., a. j. brereton, r. fealy, and j. sweeney. 2003. possible change in irish climate and its impact on barley and potato yields. agricultural and forest meteorology 116:181-196. holt, r. d. 1990. the microevolutionary consequences of climate change. trends in ecology and evolution 5:311-315. holt, r. d. 1996. demographic constraints in evolution: towards unifying the evolutionary theories of senescence and niche conservatism. evolutionary ecology 10:1-11. holt, r. d., and m. s. gaines. 1992. analysis of adaptation in heterogeneous landscapes: implications for the evolution of fundamental niches. evolutionary ecology 6:433-447. houghton, j. t., g. j. jenkins, and j. j. ephramus. 1990. scientific assessment of climate change. cambridge university press, cambridge. hugall, a., c. moritz, a. moussalli, and j. stanisic. 2002. reconciling paleodistribution models and comparative phylogeography in the wet tropics rainforest land snail gnarosophia bellendenkerensis (brazier 1875). proceedings of the national academy of sciences, usa 99:6112-6117. hughes, l. 2000. biological consequences of global warming: is the signal already; apparent? trends in ecology & evolution 15:56-61. hutchinson, g. e. 1957. concluding remarks. cold spring harbor symposia on quantitative biology 22:415-427. ipcc. 1995. ipcc second assessment. climate change 1995. ipcc secretariat, geneva, switzerland. iverson, l. r., and a. m. prasad. 2001. potential changes in tree species richness and forest community types following climate change. ecosystems 4:186-199. kadmon, r., o. farber, and a. danin. 2003. a systematic analysis of factors affecting the performance of climatic envelope models. ecological applications 13:853-867. kriticos, d. j., r. w. sutherst, j. r. brown, s. w. adkins, and g. f. maywald. 2003. climate change and the potential distribution of an invasive alien plant: acacia nilotica ssp. indica in australia. journal of applied ecology 40:111-124. lim, b. k., a. t. peterson, and m. d. engstrom. 2002. robustness of ecological niche modeling algorithms in guyana. diversity and distributions 11:1237-1246. loukos, h., p. monfray, l. bopp, and p. lehodey. 2003. potential changes in skipjack tuna (katsuwonus pelamis) habitat from a global warming scenario: modelling approach and preliminary results. fisheries oceanography 12:474-482. maccracken, m., j. smith, and a. c. janetos. 2004. reliable regional climate model not yet on horizon. nature 429:699. magana, v., c. conde, o. sanchez, and c. gay. 1997. assessment of current and future regional climate scenarios for mexico. climate research 9:107-114. martínez-meyer, e., a. t. peterson, and w. w. hargrove. 2004. ecological niches as stable distributional constraints on mammal species, with implications for pleistocene extinctions and climate change projections for biodiversity. global ecology and biogeography 13:305-314. mati, b. m. 2000. the influence of climate change on maize production in the semi-humid-semiarid areas of kenya. journal of arid environments 46:333-344. 53 martínez-meyer – climate change and ecological niche modeling meynecke, j.-o. 2004. effects of global climate change on geographic distributions of vertebrates in north queensland. ecological modelling 174:347-357. midgley, g., l. hannah, d. millar, m. rutherford, and l. powrie. 2002. assessing the vulnerability of species richness to anthropogenic climate change in a biodiversity hotspot. global ecology and biogeography 11:445-451. midgley, g. f., l. hannah, d. millar, w. thuiller, and a. booth. 2003. developing regional and species-level assessments of climate change impacts on biodiversity in the cape floristic region. biological conservation 112:1-2. morecroft, m. d., c. e. bealey, o. howells, s. rennie, and i. p. woiwod. 2002. effects of drought on contrasting insect and plant species in the uk in the mid-1990s. global ecology and biogeography 11:7-22. murphy, j. m., d. m. h. sexton, d. n. barnett, g. s. jones, m. j. webb, and m. collins. 2004. quantification of modelling uncertainties in a large ensemble of climate change simulations. nature 430:768-772. nagelkerke, k., and r. alkemade. 2003. modelling the effect of climate change on species' ranges. levende natuur 104:114-118. nathan, r., g. g. katul, h. s. horn, s. m. thomas, r. oren, r. avissar, s. w. pacala, and s. a. levin. 2002. mechanisms of longdistance dispersal of seeds by wind. nature 418:409-413. new, m., m. hulme, and p. jones. 1999. representing twentieth-century space-time climate variability. part i: development of a 1961-90 mean monthly terrestrial climatology. journal of climate 12:829-856. oreskes, n., k. shrader-frechette, and k. belitz. 1994. verification, validation and confirmation of numerical models in the earth sciences. science 263:641-646. ortega-huerta, m. a., and a. t. peterson. 2004. modelling spatial patterns of biodiversity for conservation prioritization in north-eastern mexico. diversity and distributions 10:39-54. parra, j. l., c. c. graham, and j. f. freile. 2004. evaluating alternative data sets for ecological niche models of birds in the andes. ecography 27:350-360. pearce, j., and s. ferrier. 2000. evaluating the predictive performance of habitat models developed using logistic regression. ecological modelling 133:225-245. pearce, j., s. ferrier, and d. scotts. 2001. an evaluation of the predictive performance of distributional models for flora and fauna in north-east new south wales. journal of environmental management 62:171-184. pearson, r. g., and t. p. dawson. 2003. predicting the impacts of climate change on the distribution of species: are bioclimate envelope models useful? global ecology and biogeography 12:361-371. pearson, r. g., w. thuiller, m. b. araújo, e. martínez-meyer, l. brotons, c. mcclean, l. miles, p. segurado, t. p. dawson, and d. c. lees. submitted. model-based uncertainty in species' range prediction. proceedings of the royal society london b. peterson, a. t., and k. c. cohoon. 1999. sensitivity of distributional prediction algorithms to geographic data completeness. ecological modelling 117:159-164. peterson, a. t., m. a. ortega-huerta, j. bartley, v. sanchez-cordero, j. soberon, r. h. buddemeier, and d. r. b. stockwell. 2002. future projections for mexican faunas under global climate change scenarios. nature 416:626-629. peterson, a. t., v. sanchez-cordero, j. soberon, j. bartley, r. w. buddemeier, and a. g. navarro-siguenza. 2001. effects of global climate change on geographic distributions of mexican cracidae. ecological modelling 144:21-30. peterson, a. t., and j. shaw. 2003. lutzomyia vectors for cutaneous leishmaniasis in southern brazil: ecological niche models, predicted geographic distributions, and climate change effects. international journal for parasitology 33:919-931. post, e., and m. c. forchhammer. 2002. synchronization of animal population dynamics of large-scale climate. nature 420:168-171. rafoss, t., and m. g. saethre. 2003. spatial and temporal distribution of bioclimatic potential for the codling moth and the colorado potato beetle in norway: model predictions versus climate and field data from the 1990s. agricultural and forest entomology 5:75-86. schwartz, m. w., l. r. iverson, and a. m. prasad. 2001. predicting the potential future distribution of four tree species in ohio using current habitat availability and climatic forcing. ecosystems 4:568-581. skov, f., and j. c. svenning. 2004. potential impact of climatic change on the distribution of forest herbs in europe. ecography 27:366380. soberón, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species' distributional areas. biodiversity informatics 2:1-10. stockwell, d. r. b., and d. p. peters. 1999. the garp modelling system: problems and solutions to automated spatial prediction. 54 martínez-meyer – climate change and ecological niche modeling international journal of geographic information systems 13:143-158. stockwell, d. r. b., and a. t. peterson. 2002. effects of sample size on accuracy of species distribution models. ecological modelling 148:1-13. svenning, j. c., and f. skov. 2004. limited filling of the potential range in european tree species. ecology letters 7:565-573. sykes, m. t. 2001. modelling the potential distribution and community dynamics of lodgepole pine (pinus contorta dougl. ex. loud.) in scandinavia. forest ecology and management 141:1-2. tellez-valdes, o., and p. davila-aranda. 2003. protected areas and climate change: a case study of the cacti in the tehuacan-cuicatlan biosphere reserve, mexico. conservation biology 17:846-853. thomas, c. d., e. j. bodsworth, r. j. wilson, a. d. simmons, z. g. davies, m. musche, and l. conradt. 2001. ecological and evolutionary processes at expanding range margins. nature 411:577-581. thomas, c. d., a. cameron, r. e. green, m. bakkenes, l. j. beaumont, y. c. collingham, b. f. n. erasmus, m. f. d. siqueira, a. grainger, l. hannah, l. hughes, b. huntley, a. s. v. jaarsveld, g. f. midgley, l. miles, m. a. ortega-huerta, a. t. peterson, o. l. phillips, and s. e. williams. 2004. extinction risk from climate change. nature 427:145148. thuiller, w. 2003. optimizing predictions of species distributions and projecting potential future shifts under global change. global change biology 9:1353-1362. thuiller, w., m. b. araujo, and s. lavorel. 2003. generalized models vs. classification tree analysis: predicting spatial distributions of plant species at different scales. journal of vegetation science 14:669-680. thuiller, w., m. b. araujo, r. g. pearson, r. j. whittaker, l. brotons, and s. lavorel. 2004a. biodiversity conservation: uncertainty in predictions of extinction risk. nature 430, july 1. thuiller, w., l. brotons, m. b. araujo, and s. lavorel. 2004b. effects of restricting environmental range of data to project current and future species distributions. ecography 27:165-172. travis, j. m. j., and c. dytham. 2002. dispersal evolution during invasions. evolutionary ecology research 4:1119-1129. van staden, v., b. f. erasmus, j. roux, m. j. wingfield, and a. s. van jaarsveld. 2004. modelling the spatial distribution of two important south african plantation forestry pathogens. forest ecology and management 187:61-73. vetaas, o. r. 2002. realized and potential climate niches: a comparison of four rhododendron tree species. journal of biogeography 29:545554. watson, r. t. a. t. c. w. t. 2001. climate change report 2001: synthesis report, geneva, switzerland. williams, d. w., and a. m. liebhold. 2002. climate change and the outbreak ranges of two north american bark beetles. agricultural and forest entomology 4:87-99. williams, s. e., e. e. bolitho, and s. fox. 2003. climate change in australian tropical rainforests: an impending environmental catastrophe. procceedings of the royal society london b 270:1887-1892. wittmann, e. j., p. s. mellor, and m. baylis. 2001. using climate data to map the potential distribution of culicoides imicola (diptera: ceratopogonidae) in europe. revue scientifique et technique office international des epizooties 20:731-740. zaniewski, a. e., a. lehmann, and j. mcc. overton. 2002. predicting species spatial distributions using presence-only data: a case study of native new zealand ferns. ecological modelling 157:261-280. 55 microsoft word 4717-8494-1-ed_v3_proofs.docx biodiversity informatics, 9, 2014, pp. 13-17 13 pointsampler: a gis tool for point intercept sampling of digital images david l. gobbett1 and andre zerger2 1 csiro sustainable agriculture flagship / csiro ecosystem sciences, pmb 2 glen osmond, sa 5064, australia; 2environmental information services, bureau of meteorology, gpo box 2334, act 2600, australia abstract.— close-range digital photography to assess vegetation cover is useful in disciplines ranging from ecological monitoring to agricultural research. an on-screen point intercept sampling method, which is analogous to the equivalent field based method, can be used to manually derive the percentage occurrence of multiple cover classes within an image. pointsampler is a gis embedded tool that provides a semi-automated approach for performing point intercept sampling of digital images, and which integrates with existing gis functionality and workflows. we describe and illustrate the two general applications of this tool, in efficiently deriving primary ecological data from digital photographs, and for the generation of validation data to complement automated image classification of a time series of groundcover images. the flexible design and gis integration of pointsampler allows it to be put to a wide range of similar uses. key words.— classification, groundcover, image assessment, photography, vegetation close range digital imagery is useful in research and monitoring where the proportions of different types of cover need to be assessed. its applications include arid vegetation assessment (laliberte et al. 2007), green vegetation cover (liu et al. 2012), flower number estimation (adamsen et al. 2000), crop canopy coverage (purcell 2000), forest understory foliar coverage (macfarlane and ogden 2012) and lichen cover and biomass estimation (bowker et al. 2008). the benefits of digital imagery over field based methods include reduced fieldwork time and data collection costs, which can enable data collection at a larger number of sites, plots or quadrats. automated photography can also enable the capture of data at higher temporal frequency (e.g. hourly, daily). the rapid increase in a use of unmanned aerial vehicles as platforms for image capture has expanded the use of digital imagery. a further benefit of digital images is that they can be retained for future reference or analysis for purposes other than they were originally collected. typically, close range images are used to determine attributes such as percentage vegetation cover within a number of classes, canopy cover or other presence-absence estimates. in ground cover vegetation analyses, such classes might include live vegetation, flowers, bare soil, rock or perhaps assign vegetation to particular types. the same techniques can be applied in a range of other research and monitoring domains. by classifying an image into discrete classes, the proportion of those classes can be used to assess change at a single point over time, or at points along a transect, or to facilitate comparisons between different sites. where large numbers of images are captured, either across a spatial extent, or at at intervals as a time-series, it may be appropriate to utilize automated image analysis methods such as supervised classification. supervised classification assigns the pixels of an image to classes by determining the fit with previously input signatures, and generally treat each pixel as a discrete entity. object oriented classification methods also incorporate analysis of shape and context in classifying an image. regardless of the automated method used, a classification process needs to be validated. one approach is to undertake actual field observations, or ground truthing, but another option is to manually generate calibration or validation datasets from a subset of images to be statistically compared with the results of automated classification methods. alternatively, manual image analysis may also be used as a primary data collection method. the point intercept method (pim; also known as point contact or point frequency method) is a wellestablished method for estimating vegetation cover proportions (e.g. whitman and siggeirsson 1954; heady et al. 1959; brun and box 1963). in fieldwork the pim involves counting the proportion of cover types under randomly or regularly placed pins. pim is most suitable for single-layer vegetation, and has a lower level of subjectivity than visual estimation methods (bråkenhielm and qinghong 1995), although is less suitable for recording rare species (moen et al. 2007). the method allows input from experts with specialized knowledge, or alternatively can be applied by an operator with a relatively low skill level, provided that appropriate classification gobbett and zerger – pointsampler 14 rules are supplied. applying a pim to assessing cover in digital imagery is very similar to its use in a field situation. developing workflows that incorporate software tools and scripting can be essential for efficiently dealing with larger numbers of images captured in such studies. automated image analysis methods can be used to determine percentage cover of different classes. geographic information systems (gis) include many features useful for the handling of imagery and combining them with related datasets. scripting tools such as python (python software foundation 2011) provide the ability to organize and manipulate images and can provide the ‘glue’ to link various processing steps into a reusable workflow. while automated processes facilitate the processing of large number of images, it can be necessary to perform manual assessments of images, using methods such as the pim, in order to validate automated processing. manual image assessment can be time consuming, so there is a key niche for a software tool to provide a semi-automated approach to image assessment. we evaluated two existing software tools for their suitability in assessing vegetation cover in digital imagery. vegmeasure (johnson et al. 2003) is a semi-automated image classification tool that provides a variety of spectral band algorithms which can be calibrated by an operator using a point intercept method. vegmeasure is most suited to performing binary classifications, such as calculating cover of crop and bare ground, and did not readily accommodate our need to classify images into four or five classes. samplepoint (booth et al. 2006) is a tool for image classification using a point intercept method and supports a large number of cover classes. cagney et al. (2011) estimated that on screen image sampling using samplepoint took about one third of the time taken by a field based point intercept method. however, samplepoint does not allow the re-use of the same set of randomly generated sample points to facilitate the comparison of changes in individual points over time. further, because neither of these tools readily integrated with our existing gis workflows, we set about developing our own tool which we named pointsampler. figure 1. pointsampler being used to classify images within arcgis. gobbett and zerger – pointsampler 15 pointsampler development pointsampler was developed in visual basic .net and is installed as an add-in for arcgis version 10 (esri, redmond, ca). pointsampler (fig. 1) makes use of inbuilt arcgis functionality, such as the ability to load and display images and overlay shapefiles, apply symbology, zoom and pan. consequently, the user interface of pointsampler is kept simple, and is straightforward to use for anyone familiar with arcgis. being embedded within arcgis also adds the ability to work with georeferenced images, work to fixed scale and potentially to calculate areas, overlay ancillary data or spatial datasets, and to compare multiple coincident images. when launched, the operator assesses the cover class at a point displayed centrally on the screen then types the corresponding character code. the image then zooms automatically to the next point to be classified. the character codes are defined in a text file and can easily be modified to different needs. for each set of images to be classified, a shapefile contains the sample points and also stores the coded values for each point. the points to be sampled can be generated manually, or with tools such as geospatial modelling environment (beyer 2011) and may form a regular grid or random arrangement. for each image to be processed a data field is needed in the shapefile to store the coded values. this field can be added manually or, for example, through an ancillary python script which could create the shapefile and add all necessary fields as part of the image processing workflow. after coding with pointsampler is completed, the shapefile dataset can be used within arcgis or exported for further analysis. applications of pointsampler there are many possible uses for pointsampler. two key areas in which its value has been demonstrated are in the generation of primary dataset (fig 2a), and to assist in the generation of a dataset used to validate an automated image classification process (fig 2b). an example of each of these cases is presented here. generation of a primary dataset pointsampler was as part of an assessment of floodplain woodland structure and condition (mcginness et al. 2013) to streamline the capture of groundcover data along transects within study sites. photographs were taken 1.5 m above the ground using 12 megapixel compact digital cameras. prior to analysis using pointsampler, the individual images were named so as to identify the site at which they were captured, but no other image preprocessing was performed. pointsampler was effectively used to capture a range of attributes from each image, figure 2: applications of pointsampler in (a) primary data capture and (b) generation of a dataset for validating an automated image classification process. gobbett and zerger – pointsampler 16 including the presence of three indicator plant species. the camera images did not need to be spatially referenced since there was no intention to compare repeat images of the same sites, and so they could be processed with pointsampler without additional processing. in this case, the use of pointsampler resulted in the efficient generation of primary data for each of the study sites. generation of a dataset for validation of automated classifications in a ground-cover monitoring project, daily images of vegetation plots were captured over a six month period using fixed, downward pointing, consumer grade, weatherproof digital cameras (zerger et al. 2012). to assess changes in groundcover we focused on five classes, live (l), attached litter (a), detached litter (d), bare ground (b) and other (o). automated processing was used to apply a maximum likelihood classifier method within envi (itt visual solutions, 2008) to detect trends in cover proportions over the entire time series. image preparation included georeferencing and cropping each image to the plot boundary, which enabled appraisal of change at individual points in sequential images. in a similar way to rotz et al. (2008) we required a validation dataset based on manual image assessment by a skilled operator. pointsampler was used to classify 100 random points from weekly images selected from the larger timeseries. in this case, pointsampler enabled the generation of a validation dataset used to assess the automated image classification process. while performing the manual assessments using pointsampler, some subjectivity became apparent in the differentiation of classes which tend to occur across a gradient, such as from a to d and d to b. it was helpful to elucidate clear rules to assist the operator in performing the classifications. for a detailed description of these rules see zerger et al. (2012). overall pointsampler enabled the operator to process the images far more rapidly than would otherwise have been possible with a more manual process. discussion the above examples illustrate the utility of pointsampler for the collection of primary data for field research, and for generating a validation dataset used in an automated image classification process. in both case studies, clearly defined rules were needed to assist the operator in the classification task. in both cases pointsampler enabled the processing of larger numbers of images than would otherwise have been feasible in the available time. importantly, the integration of pointsampler within a gis allows it to be incorporated into image processing workflows that can be automated with the python scripting language. for example, using python, cover class scores from a series of images using pointsampler could be summarized, or cross tabulated against reference images to calculate accuracy scores such as user’s accuracy, producer’s accuracy and kappa coefficients (congalton 1991). we intend to update pointsampler from time-totime to accommodate new versions of arcgis software. pointsampler benefits from tight integration within the arcgis user interface, which however precludes it from use with different gis software. while pointsampler is a useful tool in its current form, a number of enhancements are being considered for a future version. these include the addition of program features to simplify editing of the classification codes, shapefile management (such as simplifying the addition of a field for a new image), a training mode – similar to that offered in samplepoint, which allows inexperienced users to compare their classification choices against those of an expert, and a function to generate producer’s and user’s accuracy statistics (congalton 1991) against equivalent automatically classified images. the use of low cost digital photography for primary field data collection and assessment, such as in estimating proportional composition of groundcover classes, can be substantially streamlined using pointsampler. its use is particularly appropriate where gis workflows form part of the image analysis process, and enables large numbers of images to be manually assessed efficiently. the versatility of pointsampler allows it to be applied to a range of similar uses, in any domain in which classification and sampling of digital images is used. as well as applications of the two types illustrated here, it is conceivable that pointsampler could be used for the generation of training datasets for input into automated processing. the pointsampler arcgis add-in can be downloaded from the csiro website1. acknowledgements feedback provided about pointsampler by various colleagues, especially micah davies and heather mcginness, is appreciated. two anonymous reviewers provided valuable comments and suggestions.   1 http://www.csiro.au/pointsampler gobbett and zerger – pointsampler 17 references adamsen, f. j., t. a. coffelt, j. m. nelson, e. m. barnes, and r. c. rice. 2000. method for using images from a color digital camera to estimate flower number. crop sci 40:704-709. beyer, h. l. 2011. geospatial modelling environment (version 0.5.3 beta). booth, d., s. cox, and r. berryman. 2006. point sampling digital imagery with ‘samplepoint.’ environ monit assess 123:97-108. bowker, m. a., n. c. johnson, j. belnap, and g. w. koch. 2008. short-term monitoring of aridland lichen cover and biomass using photography and fatty acids. j arid environ 72:869-878. bråkenhielm, s. and l. qinghong. 1995. comparison of field methods in vegetation monitoring. water air soil poll 79:75-87. brun, j. m. and t. w. box. 1963. a comparison of line intercepts and random point frames for sampling desert shrub vegetation. j range manage 16:21-25. cagney, j., s. e. cox, and d. t. booth. 2011. comparison of point intercept and image analysis for monitoring rangeland transects. rangeland ecol manag 64:309315. congalton, r. g. 1991. a review of assessing the accuracy of classifications of remotely sensed data. remote sens environ 37:35-46. heady, h. f., r. p. gibbens, and r. w. powell. 1959. a comparison of the charting, line intercept, and line point methods of sampling shrub types of vegetation. j range manage 12:180-188. johnson, d., m. vulfson, m. louhaichi, and n. harris. 2003. vegmeasure v1.6 user's manual. department of rangeland resources, oregon state university, corvallis, oregon. laliberte, a. s., a. rango, j. e. herrick, e. l. fredrickson, and l. burkett. 2007. an object-based image analysis approach for determining fractional cover of senescent and green vegetation with digital plot photography. j arid environ 69:1-14. liu, y., x. mu, h. wang, and g. yan. 2012. a novel method for extracting green fractional vegetation cover from digital images. j veg sci 23:406-418. macfarlane, c. and g. n. ogden. 2012. automated estimation of foliage cover in forest understorey from digital nadir images. methods ecol evol 3:405-415. mcginness, h. m., a. d. arthur, m. davies, and s. mcintyre. 2013. floodplain woodland structure and condition: the relative influence of flood history and surrounding irrigation land use intensity in contrasting regions of a dryland river. ecohydrology 6:201-213. moen, j., ö. danell, and r. holt. 2007. non-destructive estimation of lichen biomass. rangifer 27:41-46. purcell, l. c. 2000. soybean canopy coverage and light interception measurements using digital imagery. crop sci 40:834-837. python software foundation. 2011. python programming language. rotz, j. d., a. o. abaye, r. h. wynne, e. b. rayburn, g. scaglia, and r. d. phillips. 2008. classification of digital photography for measuring productive ground cover. rangeland ecol manag 61:245-248. whitman, w. c. and e. i. siggeirsson. 1954. comparison of line interception and point contact methods in the analysis of mixed grass range vegetation. ecology 35:431-436. zerger, a., d. gobbett, c. crossman, p. valencia, t. wark, m. davies, r. n. handcock, and j. stol. 2012. temporal monitoring of groundcover change using digital cameras. int j appl earth obs geoinformation 19:266-275.   microsoft word coetzer_final4.docx biodiversity informatics, 8, 2012, pp.1-11. 1 a new era for specimen databases and biodiversity information management in south africa willem coetzer, ofer gon south african institute for aquatic biodiversity, somerset street, grahamstown, south africa. email for correspondence: w.coetzer@saiab.ac.za. michelle hamer south african national biodiversity institute, cussonia ave, pretoria, south africa fatima parker-allie south african national biodiversity institute, rhodes drive, cape town, south africa abstract. ‒ we present observations and a commentary on the inherited legacy and current state of biodiversity information management in south african natural history museums, and make recommendations for the future. we emphasize the importance of using a recognized database application, and training and capacity development to improve the quality and integration of biodiversity information for research. in the last decade, biodiversity information in specimen databases of natural history museums has seen renewed interest and much innovation and development (bisby 2000, soberón and peterson 2004, johnson 2007, peterson et al. 2010). biodiversity informatics has been defined as ‘application of informatics to recorded and yetto-be discovered information specifically about biodiversity, and the linking of this information with genomic, geospatial and other biological and non-biological datasets.’ the mission of the global biodiversity information facility (gbif1) is to ‘facilitate free and open access to biodiversity data worldwide via the internet to underpin sustainable development.’ as of january 2012 the gbif data portal provided access to >317 million primary biodiversity data records. by january 2011 ~7.1 million records on the gbif data portal were contributed by the south african biodiversity information facility (sabif), to which natural history museums in south africa contribute their biodiversity information. these records originated from 8 south african data providers (mostly natural history museums) and 14 collections. eighty per cent of the records were unvouchered occurrences, mostly observation records from the south african bird atlas project, contributed by the animal demography unit of the university of cape town. in january 2011 the digital records of 1http://www.gbif.org/. approximately 26% (hamer 2011) of vouchered specimens in south african zoological collections could be queried through the sabif data portal or the gbif data portal. the vast majority of information about south african biodiversity, which is relatively well sampled (figure 1), originates from south african natural history museums and the south african bird atlas project (table 1). biodiversity information, including specimen records from natural history collections, is used in: • biodiversity monitoring (e.g., reyers and mcgeoch 2007); • bioregional planning (e.g., smith and wolfson 2004); • identifying and categorizing threatened species (e.g., tweddle et al. 2009); • understanding the impacts of global change on biodiversity (e.g., skelton and coetzer 2011, cherry 2009, skelton et al. 1995) and developing mitigation strategies; • informing sustainable harvesting programs; • control of alien, invasive species (e.g., foxcroft et al. 2009) and disease vectors; • environmental impact assessments. ecological niche modeling (phillips et al. 2006) is a productive research area that relies heavily on high-quality biodiversity data, especially with respect to taxonomic precision and precision of georeferencing. maintaining such biodiversity informatics, 8, 2012, pp.1-11. 2 high-quality data requires a well-designed and well-managed relational database and application, tailor-made for biodiversity information. perhaps the most practical use of biodiversity information is in systematic conservation planning, to identify areas that need to be protected for the persistence and spatial continuity of genetic variability (populations), species, communities or ecological services (nel et al. 2011). the biodiversity community has recently drawn extensively on biodiversity information held by south africa’s natural history museums for important, national biodiversity projects. examples of these are the alien zonation project (nem:ba, 2004), the national freshwater protected areas project (nel et al. 2011) and the national spatial biodiversity assessment (reyers et al. 2007). figure 1. the density of occurrences of south african biodiversity as published by gbif in april 2012. table 1. numbers of specimen occurrences of south african biodiversity contributed to gbif as of april 2012. numbers in parentheses refer to vouchered specimens. country georeferenced occurrences georeferenced as % of total not georeferenced argentina 14 <0.05 0 australia 3,676 <0.05 2,395 austria 775 <0.05 4,352 belgium 4,003 <0.05 3,406 canada 1,501 <0.05 1,019 colombia 15 <0.05 0 denmark 1,284 <0.05 319 estonia 6 <0.05 139 finland 733 <0.05 1,179 france 612 <0.05 24,754 germany 35,604 0.05 21,403 hong kong 2 <0.05 0 india 236 <0.05 0 japan 6 <0.05 1,556 mexico 56 <0.05 176 netherlands 2,056 <0.05 6,201 new zealand 0 <0.05 105 norway 376 <0.05 0 poland 529 <0.05 764 portugal 0 <0.05 38 south africa 7,115,071 (2,295,813) 98.3 (94.8) 638,843 spain 132 <0.05 531 sweden 1,643 <0.05 36,594 switzerland 2 <0.05 16,014 taiwan 1 <0.05 0 united kingdom 2,166 <0.05 37,773 united states 69,496 1 71,476 total 7,239,995 (2,420,737) 869,037 as the ultimate, verifiable source of much biodiversity information, natural history museums are now not only responsible for the curation, preservation and management of collections of physical specimens, but also for capturing, managing and disseminating accurate, precise, biodiversity informatics, 8, 2012, pp.1-11. 3 and current biodiversity information in the form of up-to-date species names and specimen records. similarly the perception of a specimen database is changing from that of a pure collection management tool to include the concept of a ‘biodiversity specimen database’: a repository of useful research data relating to occurrences of biological species in spatial, environmental and temporal contexts. as the former, the database is used independently of other organizations’ specimen databases, in an inward-looking fashion (e.g., coetzer et al. 2009), but as the latter, the information from different databases and organizations needs to be integrated in interoperable systems, to discover and analyze patterns in biodiversity, and to detect changes in these patterns, including those caused by global change. in this article we consider some of the challenges faced by the south african biodiversity community with respect to biodiversity information management, and emphasize the importance of adopting a more systematic and rigorous approach to formal biodiversity information management in south african natural history museums. we believe that training and skillsdevelopment will be central to any strategy that seeks to improve biodiversity information management in south africa. in the last decade, we have witnessed the development of a staggering array of standards, recommendations, products, tools, initiatives, and collaborations focusing on biodiversity informatics. we should not postpone introducing young technical staff into this new world any longer. information management is not the only challenge faced by natural history collections in south africa. in 2010, a survey of zoological collections was undertaken to assess the state and sustainability of collections. the report representing the initial outcome of the survey (hamer 2011) found that collections could be consolidated to achieve a critical mass of staff and economies of scale. there can be no doubt that any investment in biodiversity information management in museums will have complementary positive spin-offs for sustainability in curation and collection management, especially at this critical time. coordinated biodiversity information management in south africa drinkrow et al. (1994) were among the first to highlight the need for greater recognition of the value of south african natural history collections in biodiversity research, specifically the need for coordinated biodiversity database management through a national biodiversity network. in the late 1990s, such an initiative was launched. named biomap (and later renamed sa-isis), it was the first coordinated program to capture, collate and disseminate south african biodiversity information, including specimen records from natural history collections. sabif was established as a program under the management of the national research foundation (nrf) in 2003, following the signing of a memorandum of understanding with gbif. in 2006, sabif was incorporated into the south african national biodiversity institute (sanbi). as a node of gbif, the objectives of sabif are similar to those of gbif: to mobilize biodiversity data, provide a data-sharing platform, promote data standards and tools, and develop capacity in, and raise awareness of, biodiversity informatics. progress in digitizing south african natural history collections can be traced back almost two decades (gon and wertlen 1996), but the recent development of information technology has far outpaced the development of skills in museums. south african custodians of natural history collections have been unable to address this disparity, particularly in the present context of underfunded and understaffed natural history museums (cherry 2009). this situation contrasts with the recognition of south africa as a megadiverse country with a tradition of excellence in collection management and research in biodiversity science. in the last five years, however, a user-community of people directly involved in biodiversity informatics has been formed under the auspices of sanbi to address this need. practitioners meet annually in the biodiversity information management forum (bimf 2 ) to exchange information on the acquisition, maintenance, sharing and use of biodiversity information. among the many subjects discussed at bimf meetings is the array 2 bimf: biodiversity information management forum. http://www.infoforum.org.za. biodiversity informatics, 8, 2012, pp.1-11. 4 of recently developed standards and protocols in biodiversity informatics. despite significant achievements, most notably the successful establishment of sabif, the state of digitization and web mobilization of south african biodiversity collections leaves much to be desired. at least 18 organizations and 63 collections exist that could potentially contribute ~6.5 million as-yet undigitized zoological specimen records to sabif/gbif. the quality of biodiversity information contributed to sabif is highly variable, and in some cases substandard due to a history of inadequate control over data quality by contributing organizations, a result of inadequate human capacity, infrastructure, training, and capacity development, specifically in museums. if we increase the magnification and examine what is happening at the individual workstation, we still see a lack of fundamental skills and inability to adopt new technology. we suggest that standardization and formalization of biodiversity information management are therefore needed in the museum and at the individual workstation. the proposed mechanism to achieve this objective should be based on the common use of a recognized database product, such as specify6, a new version of the specify biodiversity collections management software platform developed by the university of kansas biodiversity institute. of the ~2.8 million records in south african collections that are already digitized, only about a third are presently managed using specify6 databases, representing ~900,000 specimen records (hamer 2011; table 2). table 2. the south african natural history collections currently using specify6. organization collection no. of specimen records using specify since south african institute for aquatic biodiversity fish amphibians diatoms total 100,000 500 50,000 150,500 august 2001 january 2009 january 2011 ditsong national museum of natural history reptiles and amphibians birds mammals archaeozoology total 74,000 44,000 46,000 2,500 166,500 january 2008 january 2008 january 2008 january 2008 albany museum aquatic invertebrates fish terrestrial insects total 67,000 15,000 33,000 115,000 january 2009 january 2009 december 2009 kwazulu-natal museum diptera oligochaeta total 54,000 5,000 59,000 september 2010 september 2010 agricultural research council apoidea and chalcidoidea homoptera total 30,510 8,347 38,857 january 2011 april 2012 iziko museums of cape town paleontology and geology: plant fossils karoo vertebrate fossils invertebrate fossils rocks and minerals biodiversity: arthropods (incl. arachnida) invertebrates (incl. crustacea) fish reptiles and amphibians birds mammals total 2,295 7,748 16,694 3,592 245,890 48,230 19,826 13,700 17,048 13,590 388,532 june 2011 grand total 918,470 biodiversity informatics, 8, 2012, pp.1-11. 5 a platform for biodiversity information is a platform for biodiversity science berendsohn (2003) listed 26 software applications that could be used particularly for paleontological collections, entomological collections, botanical collections, or for many kinds of collections. an application qualified for the list if it could be used to manage specimens or observations, was available free or for purchase, was in use by at least one collection, did not require re-programming, and came with some support from the software developer. fourteen applications could handle many kinds of collections, a requirement of most south african natural history museums. current websites could be found for 11 of these applications, and only two could be downloaded and used free of charge. these were specify6 (university of kansas biodiversity institute) and biótica5 (comisión nacional para el conocimiento y uso de la biodiversidad, government of mexico). many of the world’s largest natural history museums, including those of the smithsonian institution, american museum of natural history and the natural history museum, london, use commercial software (knowledge enterprises’ electronic museum, or ke-emu, is popular) or custom software, developed in-house. as recently as 2009, naturalis, the central natural history museum of the netherlands, was finalizing a contract with a commercial company to develop a new system in-house. specify6 software together with its predecessor, muse, specify has been developed in lawrence, kansas, since the early 1990s. specify6 was released in april 2009. the specify software project is funded by the advances in biological informatics program of the u.s. national science foundation (nsf) and has received nsf support since 1987. the java application can be used on microsoft windows®, mac os x or linux operating systems, and future releases will be capable of connecting to any relational database management system. specify6 is free and open source software used by 404 collections in 43 us states and 26 countries worldwide (e.g. countries in south america and europe, as well as india and kenya). a complete list of the collections using specify software as of november 2011 may be found on the specify software website (http://www.specifysoftware.org). in addition to conventional functionality expected in a collection database, such as report design for loan invoices and labels, specify 6.4, released on 22 november 2011, included a new module specifically designed for the spatial visualization of specimen records and ecological niche modeling. through the use of state-of-theart web services, the lifemapper module allows the user to visualize not only data in the local database, but also data currently served by gbif. the need to search for, compile, collate and assemble static copies of datasets has therefore been eliminated. this illustrates that specify6 optimally marries the concepts of collection management and biodiversity analysis in a single repository-and-workbench. an upcoming web interface will closely reflect the functionality of the desktop application, and will greatly facilitate the management of off-site collections and the dissemination of information. the museum data migration project in early 2009 the first author began to clean the specimen data of selected collections in four of the country’s large natural history museums, and migrate the data to specify6 databases that were to be installed in the museums. funding was secured from sabif for a pilot project involving ditsong national museum of natural history, albany museum, kwazulu-natal museum and iziko south african museum. the objectives of this work were to: • clean legacy biodiversity data and migrate the data to specify6; • install the customized specify6 databases in the four museums; • train museum staff to use and manage the databases; • train museum staff to export formatted data to contribute to sabif; • stimulate interest in, and understanding of, and develop skills for, biodiversity information management in the south african biodiversity community, especially zoological museum collections. biodiversity informatics, 8, 2012, pp.1-11. 6 during the project every effort was made to communicate and demonstrate data cleaning and data-migration techniques, and to provide enduser training. it was during this data-cleaning and data-migration work, and during these interactions, that the insights expounded here were formed, and where a discussion and opinions among various role-players and organizations originated. while the learning curves at all the museums were steep, the specify 6 databases continue to be used successfully in all cases. the training workshops that were presented during the mdm project were, however, not comprehensive nor sufficient. no other training program exists, across the various museums, in the use of a particular biodiversity database application. the reason for this is that very few collections use a particular database product; many collections relying on flat microsoft access tables or microsoft excel spreadsheets exhibiting a wide range of expertise in design (hamer 2011). there is much enthusiasm in the south african biodiversity information community, including within sanbi and sabif, to design a sustainable biodiversity-database training program. at the very least there are now five museums using a particular database application that did not do so before the mdm project, and it can be argued that any database training that is conducted will therefore not be merely theoretical but will actually make a difference to existing collections and practices, and have a more meaningful and lasting effect. the south african biodiversity information legacy evidence suggests that the erosion of data quality and data integrity observed in some databases is directly attributable to the common practice of indiscriminately importing legacy disk operating system databases (e.g. dbase, pc file etc.) into microsoft access files (usually as single tables), and the use of these with no user-interface, or with a poorly designed user-interface. among many symptoms, arguably the worst symptom of this disease is the truncation of text fields, or perhaps the formatting of date fields as text, caused by a complete lack of design or planning, and possibly the result of inappropriate manipulation by unskilled people. the tendency of these microsoft access files to multiply is also well known, with current information becoming distributed among the copies of files. microsoft access files can easily become corrupted. the maximum file size of a microsoft access file is 2gb whereas that of a mysql file is 8tb. there is ample evidence of all these afflictions, and more, in the databases of our natural history museums. arguably the pretoria computerized information system (precis) of sanbi, designed in 1974 (morris 1974) and now considered technologically out of date, is nevertheless capable of far better data validation than the hundreds of duplicated access files, with scant validation rules, of our zoological museums. in 1974 information was managed according to strict system constraints that users accepted implicitly–they had no choice. the ill-considered and inappropriate use of microsoft access in south african natural history collections, however, probably since the early 1990s, has created a need for a large amount of data cleaning, validation and restoration of data integrity, even before migration to an appropriate platform can begin. this work will require highly skilled database analysts who understand biodiversity information–a domain that is attractive to researchers in applied computer science and ontology engineering due to its complexity. it is unclear who will do this work, how much it will cost, who will pay for it or how long it will take to complete. yet this work is urgently needed to bring south africa’s collections, biodiversity information, information systems, analysts and users up to date, and to allow researchers to more easily and effectively use the biodiversity information. there is a risk of losing biodiversity information, especially in south africa where there is a shortage of appropriate skills and the future of some natural history museums and collections is in question (hamer 2011). this risk has been recognized by gbif, which has initiated a global program (rebind 3 ) to rescue biodiversity data at risk of being lost because they are neglected in outdated systems and are not properly documented. 3biodiversity needs data. http://rebind.bgbm.org/. biodiversity informatics, 8, 2012, pp.1-11. 7 the scale and integration of biodiversity data presently the manual collation of data, usually in ad hoc spreadsheets, is accepted by many as the way to digitize specimen information. once the speciesor specimen list or analysis is published, however, the ‘flat’, static dataset resulting from the short-lived research project gets copied and distributed (unless it is lost), and the copies get edited in an uncoordinated fashion. this perpetuates the cycle of poor information management practices, and begets poor data quality, which makes these data inaccessible in the future. rather, data need to be captured in a central museum information system designed for managing natural history specimens and collections, which can accommodate the needs of a curator as well as the needs of a biodiversity scientist. thereafter, the data need to be managed on an ongoing basis in the museum. standardized biodiversity information for research can be simply extracted from a good collection management tool such as specify6. in other words, data integration is an automatic consequence of good database design and management, rather than a process itself requiring design. keeping a national system of integrated biodiversity databases up-to-date will be greatly facilitated by the standard implementation of a recognized product, such as specify6, in museums. gbif and the collaboration known as biodiversity information standards (bis 4 ), formerly the taxonomic databases working group (tdwg) of the international union of biological sciences, have developed data standards and data exchange protocols. if these protocols are followed properly there is no need to manipulate data at the record-level as a means to integrate and distribute data. for example, specify6 includes a field-mapping and dataexport utility. all that is needed is an initial mapping of fields from the specimen database schema to the set of standard terms chosen by the data analyst, such as the ‘simple darwin core’ (wieczorek et al. 2012). exporting updated data from the database in the future is then merely a matter of updating the data cache by clicking a 4bis: biodiversity information standards, http://www.tdwg.org/. button. in contrast to this, we presently see enormous investments of effort and time in manipulating individual rows and columns manually in huge datasets, even at a national scale. this practice is doomed from the start because it relies on the literal interpretation of data by humans rather than the computation, by computers, of the arbitrary unique identifiers that make a relational database useful and necessary. we are not harnessing the real power of the relational database if, in the final stage of the process, we are second-guessing it or circumventing it completely. if we fail to scale up our systems to match the overwhelming tide of incoming biodiversity information we shouldn’t be surprised when, in a few years, we are drowning in data and yet the data remain inaccessible. the usability of data like other information that is awash on the web, biodiversity information is distributed and heterogeneous, and is in need of improvement and integration. south african researchers in biodiversity science desperately need vouchered, improved, standardized and integrated, highquality biodiversity information that originates from across the board and through time. well-managed, and therefore clean, biodiversity specimen information has been easy to integrate into derived analyses in projects which have become important contributions to our knowledge of the current distribution of, and threats faced by, south african biodiversity. this readiness was demonstrated by the 2006 conservation assessments of freshwater fishes conducted by the south african institute for aquatic biodiversity (saiab) and iucn (tweddle et al. 2009). in contrast, data that have become corrupted due to poor information management practices need to be cleaned at great cost before being used or are not used at all. cherry (2009) lamented the lack of examples of the effect of global change on south african biodiversity (see foden et al. 2007 for an excellent exception). we believe that the inaccessibility of south african biodiversity information may partly explain such information shortfalls. inadequate information management systems and practices result in data of a poor biodiversity informatics, 8, 2012, pp.1-11. 8 quality. there is no doubt, at least in the case of the saiab/iucn project, that the use of specify software contributed to the quality and integrity, and therefore the usability, of the data. the perception of the museum specimen database and the responsibility for biodiversity information the museum specimen database has not received, and is still not receiving, due recognition as the origin of speciesand specimen-related biodiversity information. as individual researchers come and go, datasets tend to become more idiosyncratically designed, more fractured, less consistent, separated from the voucher specimens and more isolated from one another. for this reason, paradoxically, information management has become more difficult since the 1970s because there is now no control over how individuals manage or mismanage information. a useful thought experiment is to imagine leaving the responsibility for an organization’s financial data to the whims of individual employees. unlike budget projections and actual expenditure figures, however, biodiversity information is not only complex and interesting, but also remains useful after three years (provided that it is managed well), and its value increases with time, even after, and especially when, the voucher specimens have dried up or been reduced to dust by museum bugs despite our best efforts. biodiversity information management in south africa needs to be formalized and recognized as a profession, or we face a future of ever-more-haphazard datasets and practices, and our return on investment will be in glitz (a ‘facebook of biodiversity’?), and not in data, information, research, knowledge, conservation policy or conservation action. vouchered biodiversity information originates in the museum. as long as museums exist, biodiversity information will continue to originate in museums. as the origin of the information, museums are the only places where the information can be created and managed in parallel with the physical specimens that it describes and documents. anything that happens to the information downstream from the museum can only be to facilitate the discovery, transmission, and use of the information. the challenges we face in capacity development are not insurmountable if we recognize the role of the biodiversity information manager in the museum, a competency that is presently missing between users and senior management, or between users and the network administrator or it support staff. recommendations and conclusions south african natural history museums need to employ qualified biodiversity information managers. the responsibilities of the biodiversity information manager (a kind of a biologist) are to: • understand the particular meaning of this museum’s biodiversity information; • maintain oversight and control over data quality in the museum; • train museum staff to use the biodiversity database and information; • manage and administer the biodiversity database; • develop the museum’s broader biodiversity information management system (e.g. expand the scope to include new kinds of data); • curate and analyze the biodiversity information using appropriate technology; • communicate with users, management and it staff in the museum, and with the wider community. the way we work with biodiversity information has been revolutionized in the last decade and has continued to change significantly in the last few years. specimen records used to be manipulated exclusively by collection managers and systematists, who often became database designers out of necessity. information management is now a specialized discipline, requiring qualified, skilled and experienced designers and analysts to work together with collection managers and expert systematists to develop and maintain sophisticated, web-enabled biodiversity information management systems. our hope is that by publicizing the importance and vulnerability of biodiversity information, we will highlight the need to develop formal information technology skills among our young biologists, who already have some understanding of the biological and information domains. we need to train specimen cataloguers, georeferencers, biodiversity informatics, 8, 2012, pp.1-11. 9 taxonomic editors, information analysts and information managers. we propose not only that the adoption of a recognized biodiversity database product, such as specify6, will be a solution for information management in museums, but that the very use of a recognized product will facilitate the development of capacity to manage biodiversity information in museums. for example, the developers of specify6 have employed a language that can be used to communicate effectively about biodiversity specimens and information. words and phrases like ‘collecting event,’ ‘collection object,’ and ‘preparation’ are not only a part of conceptual data modeling but are also concise human expressions of the concepts that we need to represent in biodiversity information. the wider use of this ‘language’ on an everyday basis in museums will allow data capturers, specimen handlers, collection managers and information managers to understand their work and communicate with each other more effectively and easily. the subject of support, training and capacity development in collections information management and biodiversity informatics has never been more important. the emergence of relational databases in specimen collections in the 1980s was an opportunity that the south african collections and biodiversity community largely missed. today the ability to manipulate specimen records and species records is still seen as a mystifying trick on the side-line rather than the mainstream profession that it should be seen as. delaying the introduction of state-of-the-art tools will only delay efforts to support, train, and develop capacity among curators and collection managers. while we are not recommending blindly enforcing superficial software standardization, we believe that standardizing on an established product will greatly facilitate training, specifically through enhanced communication through the use of a ‘database language’ across different museums. it’s no wonder that the dina project, 5 a collaboration of museums in sweden, denmark and estonia, is choosing to stand on the shoulders 5 digital information system for natural history collections. http://www.dina-project.net. of giants by adopting the specify6 schema as a standard for developing a web-based information system to allow scientists and amateurs to manage collections and distribute biodiversity information. not only is the specify6 schema an integral part of the developing dina system, but during a transitional phase the specify6 application itself will be implemented wherever an adequate biodiversity database is lacking. south africa has the potential to continue this trend, and become one of the first countries to adopt a national database training, capacity development and database implementation program. smith et al. (2003) investigated the value of south african natural history collections and their associated information, a subject that has been studied in detail previously (pietsch and anderson 1990, allmon 1994, davison 1994). natural history collections and biodiversity information are highly valuable to science and society (scoble 2010). local practitioners and professionals ought to demonstrate this value by managing collections and biodiversity information using the right tools. certainly, nobody else is likely to fly the flag of natural history collections on their behalf. we owe it not only to the legacies of the thousands of naturalists and curators who built these priceless collections to look after the specimens and information, but to the research community as well. the present and future scientific research and conservation action that stand to benefit from a more rigorous approach to biodiversity information management are too important to sacrifice on the altar of absent or second-rate biodiversity database design and management. by stimulating interest in, and appreciation of, natural history collections, taxonomy, systematics, biodiversity informatics and biodiversity science through the use of sophisticated tools, especially among the younger generation, we will be able to rise to the challenge of a new era. in doing so we could simultaneously address the country’s shortage of skills, particularly by training and employing skilled technicians. acknowledgments through sabif, sanbi funded the museum data migration project. we wish to thank the following people, as well as two anonymous biodiversity informatics, 8, 2012, pp.1-11. 10 reviewers, for reading the manuscript and making suggestions for improvement: ferdy de moor, hamish robertson, billy de klerk; teresa kearney, sarah gess, ansie dippenaar-schoeman, charnie craemer and eddie ueckermann. literature cited allmon, w.d. 1994. the value of natural history collections. curator 37:83-89. berendsohn, w. 2003. survey of existing publicly distributed collection management and data capture software solutions used by the world’s natural history collections: global biodiversity information facility, copenhagen. bisby, f.a. 2000. the quiet revolution: biodiversity informatics and the internet. science 289: 23092312. cherry, m.i. 2009. what can museum and herbarium collections tell us about climate change? south african journal of science 105:87-88. coetzer, w., o. gon, and p.h. skelton. 2009. moving into the future: relocating the national fish collection to a new, dedicated collection facility. collection forum 23:1-10. davison, p. 1994. museum collections as cultural resources. south african journal of science 90:435-436. drinkrow, d.r., m.i. cherry, and w.r. siegfried. 1994. the role of natural history museums in preserving biodiversity in south africa. south african journal of science 90:470-479. foden, w., g.f. midgley, g. hughes, w.j. bond, w. thuiller, m.t. hoffman, p. kaleme, l.g. underhill, a. rebelo and l. hannah. 2007. a changing climate is eroding the geographical range of the namib desert tree aloe through population declines and dispersal lags. diversity and distributions 13:645-653. foxcroft, l.c., d.m. richardson, m. rouget, and s. macfadyen. 2009. patterns of alien plant distribution at multiple spatial scales in a large national park: implications for ecology, management and monitoring. diversity and distributions 15:367-378. gon, o., and r. wertlen. 1996. fishnet, a computerized database management system for the national fish collection at the j.l.b. smith institute of ichthyology. south african journal of science 92:117-121. hamer, m. 2011. an assessment of the zoological research collections in south africa: report to the national research foundation (unpublished). johnson, n.f. 2007. biodiversity informatics. annual review of entomology 52:421-38. morris, j.w. 1974. progress in the computerization of herbarium procedures. bothalia 11:349-354. nem:ba national environmental management: biodiversity act, 2004: draft alien and invasive species regulations, 2009. government gazette, south africa 32090:3-37. nel, j.l., k.m. murray, a.m. maherry, c.p. petersen, d.j. roux, a. driver, l. hill, h. van deventer, n. funke, e.r. swartz, l.b. smith-adao, n. mbona, l. downsborough, s. nienaber, s. 2011. technical report for the national freshwater ecosystem priority areas project: water research commission report no. 1801/2/11. water research commission, pretoria. peterson, a.t., s. knapp, r. guralnick, j. soberón and m.t. holder. 2010. the big questions for biodiversity informatics. systematics and biodiversity 8:159-168. phillips, s., r. anderson, and r. schapire. 2006. maximum entropy modeling of species geographic distributions. ecological modelling 190:231-259. pietsch, t.w. and anderson w.d. 1990. collection building in ichthyology and herpetology. bulletin of marine science 62:957-958. reyers, b., m. rouget, z. jonas, r.m. cowling, a. driver, k. maze, and p. desmet. 2007. developing products for conservation decisionmaking: lessons from a spatial biodiversity assessment for south africa. diversity and distributions 13:608-619. reyers, b. and m.a. mcgeoch. 2007. a biodiversity monitoring framework for south africa: progress and directions. south african journal of science 103:295-300. scoble, m. 2010. rationale and value of natural history collections digitisation. biodiversity informatics 7:77-80. skelton, p.h., j.a. cambray, a. lombard, and g.a. benn. 1995. patterns of distribution and conservation status of freshwater fishes in south africa. proceedings of the zoological society of southern africa 1994 30:71-81. skelton, p.h., and w. coetzer. 2011. changing patterns of freshwater fish diversity in south africa. pp. 190-192 in observations on environmental change in south africa (l. zietsman (ed.). sun press, stellenbosch. smith, g.f., y. steenkamp, r.r. klopper, s.j. siebert, and t.h. arnold. 2003. the price of collecting life. nature 422:375-376. smith, g.f. and m.m. wolfson. 2004. mainstreaming biodiversity: the role of taxonomy in bioregional planning activities in south africa. taxon 53:467. soberón, j. and a.t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity informatics, 8, 2012, pp.1-11. 11 biodiversity data. philosophical transactions of the royal society b 359:689-698. tweddle, d., r. bills, e. swartz, w. coetzer, l. da costa, j. engelbrecht, j. cambray, b. marshall, d. impson, p.h. skelton, w.r.t darwall, k.s. smith. 2009. the status and distribution of freshwater fishes. pp. 21-37 in the status and distribution of freshwater biodiversity in southern africa (w.r.t. darwall, k.g. smith, d. tweddle, and p.h. skelton, eds.). international union for the conservation of nature, gland. wieczorek, j., d. bloom, r. guralnick, s. blum, m. döring, r. giovanni, t. robertson, and d. vieglais. 2012. darwin core: an evolving community-developed biodiversity data standard. plos one 7:e29715. microsoft word 6975-13772-1-ce_1.docx biodiversity informatics, 13, 2018, pp. 1-10 an open-access platform for camera-trapping data eugenio padilla-gómez1, mario c. lavariega2*, pablo antonio garcía-santiago1, josé santiago-velasco1, and raúl oswaldo méndez-méndez1 1dirección sierra juárez-mixteca, comisión nacional de áreas naturales protegidas, av. independencia 709, centro, oaxaca de juárez, 68100, mexico. 2centro interdisciplinario de investigación para el desarrollo integral regional, unidad oaxaca, instituto politécnico nacional, hornos 1003, santa cruz xoxocotlán, oaxaca, 71230, mexico. *corresponding author: mariolavnol@yahoo.com.mx abstract.—in southern mexico, local communities have been playing important roles in the design and collection of wildlife data through camera-trapping in community-based monitoring of biodiversity projects. however, the methods used to store the data have limited their use in matters of decisionmaking and research. thus, we present the platform for community-based monitoring of biodiversity (pcmb), a repository, which allows storage, visualization, and downloading of photographs captured by community-based monitoring of biodiversity projects in protected areas of southern mexico. the platform was developed using agile software development with extensive interaction between computer scientists and biologists. system development included gathering data, design, built, database and attributes creation, and quality control. the pcmb currently contains 28,180 images of 6478 animals (69.4% mammals and 30.3% birds). of the 32 species of mammals recorded in 18 pa since 2012, approximately a quarter of all photographs were of white-tailed deer (odocoileus virginianus). platforms permitting access to camera-trapping data are a valuable step in opening access to data of biodiversity; the pcmb is a practical new tool for wildlife management and research with data generated through local participation. thus, this work encourages research on the data generated through the communitybased monitoring of biodiversity projects in protected areas, to provide an important information infrastructure for effective management and conservation of wildlife. key words: agile software development, community-based monitoring, community conservation areas, protected areas, oaxaca. community-based monitoring of biodiversity (cbm) requires participation of local people in study design and data collection (meffe et al. 2002, danielsen et al. 2007, conrad and hilchey 2011). this local intervention increases probability of success of conservation projects because it creates a sense of ownership among participants (danielsen et al. 2007, conrad and hilchey 2011, dickinson et al. 2012). cbm yields other benefits, such as creation of local employment, increased human capital, and increased tolerance of humanwildlife conflicts (treves et al. 2009, burton 2012). cbm is particularly necessary in areas with high biological and cultural diversity, as well as in areas where land tenure is communal, as in state of oaxaca in southern mexico. oaxaca holds some of the richest biodiversity in mexico (flores-villela and garcía-vázquez 2014, navarro-sigüenza et al. 2014, parra-olea et al. 2014, sánchez-cordero et al. 2014, villaseñor and ortiz 2014). it also has impressive human ethnic diversity: 16 of the 58 native groups of mexico, and 158 of the 291 known languages in the country (de ávila 2008). almost 70% of its territory is under communal land tenure, and community members carry out decisions about management of natural resources (bray et al. 2008, martin et al. 2011). communities of oaxaca have been pioneers in community conservation processes (de la maza 2010); currently, at least 880 community conservation areas and 12 governmental protected areas exist in the state, together protecting 12% of the state’s land area (brionessalas et al. 2016). to involve landowners in gathering data of wildlife populations, the mexican government has implemented the cbm projects (conanp 2016). the purpose of these projects is to provide equipment and training to local people on use of camera-traps and global positioning systems; padilla-gómez et al. – camera-trap monitoring of biodiversity count and photograph mammal signs; census birds; and update databases (conanp 2016). in the last three years, cbm projects in protected areas of oaxaca have generated ~150,000 photographs and videos. however, the resulting data are being stored in ways that restrict access to the information, thereby limiting their use in decision-making and research. since analyses of wildlife data are essential to conservation and management efforts (meffee et al. 2002), open access to biodiversity data becomes crucial (molloy 2011, thessen and patterson 2011, hanssen et al. 2014). because biodiversity processes are dynamic—and conservation and management efforts must be designed to incorporate this characteristic—it is crucial to implement tools that allow prompt distribution of information regarding occurrences of species, their abundance, their trends, and their ecosystem services (thessen and patterson 2011, nesshöver et al. 2016). as such, biodiversity data repositories have proven crucial in supporting management efforts, adding to scientific knowledge, and increasing citizen appreciation of biodiversity (conrad and hilchey 2011, thessen and patterson 2011). therefore, with the goals of processing large amounts of data, assembling geographic information with photographic records, and providing educational materials in accessible formats, a platform for serving data resulting from biodiversity monitoring was created. the aim of this work is to present the development of the platform for community-based monitoring of biodiversity (pcbm1). materials and methods study site community-based monitoring of biodiversity projects had been implemented in 18 protected areas in southern mexico, distributed in various sectors of the region: the western mountains and valleys (mixteca region), sierra madre de oaxaca, central valleys of oaxaca, central mountains and valleys, and balsas depression physiographic sub-provinces (figure 1). the western mountains and valleys are characterized by temperate climate, holding pine 1 http://dsjm-conanp-monitoreo.org/. forest and pine-oak forest, and at low elevations tropical deciduous forests. the sierra madre de oaxaca, in northern oaxaca, holds pine-oak forest, oak forest, and montane cloud forest. tropical perennial forest is found in the foothills. lowlands and knolls dominate the central valleys of oaxaca, where landscapes have been modified for agriculture, pastureland, and settlements, in a tropical climate setting (ortizpérez 2004, inegi 2013). community-based monitoring of biodiversity since 2011, the sierra juárez-mixteca office of the mexican national commission of natural protected areas began implementing programs aimed at raising awareness of the significance of conservation in protected areas located in oaxaca. in a series of meetings, we presented program objectives for wildlife monitoring to the authorities of communities and community assemblies in protected areas. to involve local people in monitoring, incentives provided included equipment, training, temporary employment, and field assistance, with regular reports. members of monitoring committees were selected by the community assemblies, with basis on experience in conservation projects and knowledge of the territory. in workshops, monitoring committees were trained in use of camera traps, geographic positioning systems (gps), cameras, and databases. numbers of camera traps in communities were a function of the annual budget of the sierra juárez-mixteca office and ranged 4–10 devices. during surveys, monitoring committees followed standardized protocols for camera-trap data collection (chávez et al. 2013, padilla-gómez et al. 2015). training protocols included details of distances between camera traps (1-3 km), placement of camera traps at sites (i.e., distance between targets and camera traps, height above the ground, and orientation with respect to sunlight), and several dry-run tests. an important element was recording geographic coordinates and elevation with a gps of each camera trap site. experienced team members accompanied monitoring committees during initial field surveys and at intervals thereafter. camera traps figure 1. map showing the study area. red dots show locations of camera traps; gray polygons show administrative boundaries of subprovinces; blue polygons show mexican government protected areas; and green polygons show areas voluntarily destined for conservation. figure 2. workflow in development of the platform for community-based monitoring of biodiversity. padilla-gómez et al. – camera-trap monitoring of biodiversity were checked every ~15 days to download the images and change batteries. camera traps remained in the field for 1–6 months whenever feasible. platform development the pcmb was developed using an agile software development approach. the process consisted of an initial work plan and several proofing cycles that included testing, assessment, and improvement until the platform was functioning according to desired objectives (pressman 2006). these iterations were carried out via multiple meetings between computer scientists and biologists, which improved the workflow greatly (figure 2). system development was conducted in three stages: (1) gathering data, in which photographs, videos, and data acquired were compiled; (2) development and design of the platform, in which the maps module was created to feature an interactive map showing the locations of the camera-trap stations via googleearthtm; and (3) designing database relationships using workbench software 2 (pressman 2006), and developing schemes to show relationships between data and attributes in the databases. the database was created in mysql. once diagrams were built, a database with all information on the biodiversity of each protected area was created. all interfaces were designed using hypertext markup language (html) and cascading style sheets (css) to create the appearance of a web page. the website was developed using the pre hypertext processor (php) programming language to link the site to the databases, organize the data, and return the content as html to the browser. to simplify interactions between html documents and to make pages more dynamic, the jquery library of the javascript programming language was used, which includes plug-ins like jquery.validate, jquery.ui, jquery.autosize and modernizr. finally, netbeans was used to integrate the website; for geographic information, the google earth and google maps apis were used. once completed, the platform was housed in a 2 https://www.mysql.com/products/workbench/ 3 http://rs.tdwg.org/dwc/ 4 http://www.naturalista.mx/ dreamhost server, which provided flexibility to the site in terms of storing capacity and speed connection. the platform allows access to two types of users: managers and users. the manager has control over stored information through the attribute options. this interface includes options to convert excel databases to tables in darwin core3 format, and to label photographs also in darwin core format (wieczorek et al. 2012). meanwhile, the user or visitor can access an interactive google earth map showing locations of camera trap stations. the visitor can perform a search on species or biodiversity measures, using filters such year, month, activity patterns, protected area, and vegetation type, and download a data table in excel. the platform also generates quick response (qr) codes for each camera trap station, so a qr reader can view species recorded in each of them. for each photograph, the platform can generate a standardized data card with appropriate metadata (botello et al. 2007, thessen and patterson 2011, wihtlock 2011; figure 3). it can also generate a fact sheet for display or download as a pdf with information on each species recorded in a protected area. interoperability in the pcmb, we added the capability to export data in darwin core format (tegelberg et al. 2012). it also can migrate data to naturalista4, the main citizen-science platform in mexico (koleff et al. 2014), and the global biodiversity information facility (gbif) platform5 (graham et al. 2004). to this end, a script was created in the pcmb that generates a file of metadata resources, a metafile describing the content and relationships of text files in the darwin core file, and a text data file. the information is then sent to the national biodiversity information system (snib), which is housed by the national commission for the knowledge and use of biodiversity (conabio6) and linked to gbif. although in the current version of the pcmb, the audubon core metadata schema (morris et al. 2013) was not considered, it could be integrated in the future. 5 http://www.gbif.org/ 6 https://www.gob.mx/conabio figure 3. example of a standardized data card generated by the platform for community-based monitoring of biodiversity, corresponding to a photographic record of a jaguar (panthera onca). figure 4. mammal species with highest numbers of independent photographic records (a) and the highest relative abundance index (b) deposited in the platform for community-based monitoring of biodiversity. padilla-gómez et al. – camera-trap monitoring of biodiversity data workflow during surveys, monitoring committees downloaded images, naming them according to the camera trap site. subsequently, we visited communities within the program to gather images and associated data. data were centralized and stored at the sierra juárezmixteca headquarters. before uploading photographs to the platform, a team of experts on mammals or birds identified the animals in each photograph. these specialists acted as a quality control (thessen and patterson 2011). photographs were considered as independent and uploaded to the platform when they met the following criteria (monroy-vilchis et al. 2011): (1) pertaining to different individuals, and (2) same individual or species taken at intervals of >24 hr. data analysis biodiversity measures implemented in pcmb included species richness, frequency index, relative-abundance index, and diversity index. we also included number of species in categories of threat and protection according to the mexican norma oficial 059 (semarnat 2010), number of species in the red list (iucn 2016 7 ), number of species listed in the appendices of the convention on international trade in endangered species of wild fauna and flora (cites8), and number of endemic species (briones-salas et al. 2015). results the pcmb was finished and launched in june 2015. to date, the data comprise 6478 independent photographs of 28,180 records obtained since 2012. photographs correspond mainly to mammals (4497 photographs; 69.4%) and birds (1962 photographs; 30.3%). the rest were images of reptiles (16 photographs; 0.2%). in all, 4000 independent photographs and associated metadata have been shared with conabio, the mexican node of gbif. mammal species diversity over the course of the project, 32 species of medium and large-sized mammals were recorded in the protected areas. the average number of 7 http://www.iucnredlist.org/ species documented in protected areas was 12.2 (range 4–21). eight species, including coati (nasua narica), white-tailed deer (odocoileus virginianus), and gray fox (urocyon cinereoargenteus) were found regularly in protected areas, being recorded in >10 areas. in contrast, the striped hog-nosed skunk (conepatus semistriatus), ocelot (leopardus pardalis), long-tailed weasel (mustela frenata), and red brocket deer (mazama temama) were recorded only in one pa each one. one quarter (25.2%) of all independent photographs were of white-tailed deer. the species with the second largest number of independent photographs was opossum (didelphis virginiana; 7.7%), followed by coati (6.3%; figure 4a). the relative abundance index was highest for white-tailed deer, followed by collared peccary (pecari tajacu) and eastern cottontail (sylvilagus floridanus; figure 3b). conservation status we found 7 mammal species listed in the mexican norma oficial: baird’s tapir (tapirella bairdii), jaguar, ocelot, margay (leopardus wiedii), and tayra (eira barbara) were listed as endangered; hog-nosed skunk is listed as subject to special protection; and yagouaroundi (herpailurus yagouaroundi) is threatened (semarnat 2010). worldwide, the mexican agouti (dasyprocta mexicana) is considered as critically endangered; baird’s tapir is considered as endangered; and jaguar and margay are near threatened (iucn 2016). seven species were listed in the appendices of the cites: baird’s tapir, jaguar ocelot, margay, and yagouaroundi are in appendix i, and bobcat and puma (puma concolor) are in appendix ii (cites 2015). discussion pcmb is an innovative platform of biodiversity data developed to centralize, standardize, and serve open-access biodiversity data and analyses from community-based monitoring (cbm) projects using camera traps in protected areas of southern mexico. pcmb is proposed to advance wildlife management and conservation efforts in protected areas. in addition, the platform fosters collaboration and 8 https://www.cites.org/ padilla-gómez et al. – camera-trap monitoring of biodiversity exchange of information among specialists, scholars, researchers, and the general public through consultation and decision making. four additional platforms now provide camera-trap data, including wildlife insights9 , emammal10, deskteam11, and the wii camera trap data web portal 12 . these platforms and pcmb share the objectives of data mobilization from camera-trapping projects (fegraus et al. 2011, hanssen et al. 2014). pcmb is comparable to deskteam and the wii portal. deskteam is a partnership among conservation international, missouri botanical garden, wildlife conservation society, and smithsonian institution (fegraus et al. 2011). the wii portal resulted from a collaboration between indian and norwegian researchers (hanssen et al. 2014). emphasis on local participation in deskteam and pcmb is an important difference from the wii portal (fegraus et al. 2011). for example, the wii portal includes only a very few cbm studies (wild mammal biodiversity in the pune district), with most data coming from specialized researchers. deskteam, emammal, and pcmb include training in use of camera-traps (in-person or online): emammal users pay subscription fees, whereas deskteam and pcmb are free of cost. emammal presently serves data from 64 projects, and deskteam from 16 projects, both worldwide. pcmb manages data from projects in 18 sites, all in southern mexico, although it can manage data from anywhere. a next step of these different camera-trap data repositories should be to integrate or share data between platforms, as proposed by forrester et al. (2016). thessen and patterson (2011) noted that a problem with repositories developed through projects is lack of long-term funding. emammal is a non-governmental initiative, supported by subscription fees. wildlife insights and deskteam are supported by non-governmental and private agencies. the wii portal is supported by the norwegian and indian governments, while pcmb is supported by the mexican government. despite the differences, all of these platforms fit one of the core goals of the intergovernmental platform on biodiversity and 9 https://www.wildlifeinsights.org 10 http://emammal.si.edu/ ecosystem services (ipbes) by “filling knowledge gaps; build local capacities; and assessing the state of the planet’s biodiversity” (kok et al. 2016, schmeller and bridgewater 2016). these platforms provide access to unique data resources, protect the integrity of the data, and incorporate normalization, standardization, automation, quality control, and analysis of the data (thessen and patterson 2011). through these platforms, decision makers and nonspecialists can learn about the presence of medium and large-sized mammals around the world (hanssen et al. 2014, fegraus and maccarthy 2016). projects implemented in southern mexico have sought to integrate local communities since their initiation (conanp 2016). we noted that, at the beginning, villagers regarded the project as just another task. however, after retrieving the first images captured with the camera-traps, persons involved in the monitoring committees were able to see the different species of mammals and birds that inhabit their forests. in many cases, local residents did not know that particular species inhabited their lands because of nocturnal or shy habits. photographs of jaguars, puma with cubs, margay, jaguarondi, and lynx (lynx rufus) were distributed among the local population through cell phones. most communities where biological monitoring has been implemented have received some level of economic support, which, although insufficient, allows them to carry out monitoring activities. interestingly, we noted that some communities have continued to monitor even without any financial support. pcmb facilitates evaluation of efforts by communities to protect wild species. the long-term continuity of biodiversity monitoring nonetheless will depend on the funding received. thus, funding for these programs should be considered an investment that will eventually yield earnings in research, evidence-based management, and wildlife conservation (molloy 2013, piwowar et al. 2011, fegraus and maccarthy 2016). as a whole, this project recorded 32 medium and large-sized mammal species, which documents in these protected areas almost 60% 11 http://www.teamnetwork.org/ 12 http://www.wii.gov.in/. padilla-gómez et al. – camera-trap monitoring of biodiversity of the medium and large mammals of oaxaca (briones-salas et al. 2015). through this paper, we seek to encourage research from communitybased monitoring of biodiversity projects in protected areas to improve the tools and knowledge available for effective wildlife management and conservation. pcmb has already made important contributions to the general knowledge and information regarding the conservation of several threatened and endangered mammals, and offers additional opportunities for projects related to biogeography and ecosystem services. pcmb provides open-access to camera-traps records gathered by local committees contributing to the dissemination of knowledge to inform biodiversity conservation efforts. acknowledgments we thank the monitoring committees of ejido donají, tlalixtac cabrera, san pablo etla, san andrés ixtlahuaca, santo domingo tonalá, agencia santa catarina, san marcos arteaga, santa maria tindú, san francisco yosocuta, santa catarina estancia, santiago asunción, yagul, villa díaz ordaz, ejido union zapata, villa de mitla, santa catarina yetzelalag, and san juan yetzecobi. anne thessen and joe figel reviewed drafts of this manuscript; m. garcía and a. polo improved the english. anonymous reviewers and editors for their input in helping to improve the manuscript. pavel palacios and the sierra norte mixteca office provided important support. this work was supported by proyecto mixteca, national commission of natural protected areas, global environmental fund, u.n. environment programme, and world wildlife fund. literature cited botello, f., g. monroy, p. lloldi-rangel, i. trujillobolio, and v. sánchez-cordero. 2007. sistematización de imágenes obtenidas por fototrampeo: una propuesta de ficha. revista mexicana de biodiversidad 78:207-210. bray, d.b., l. merino, p. negreros-castillo, g. segura-warnholtz, j.m. torres-rojo, and h.f.m. vester. 2003. mexico’s community-managed forests as a global model for sustainable landscapes. conservation biology 17:672-677. briones-salas, m., m. cortés-marcial, and m.c. lavariega. 2015. diversidad y distribución geográfica de los mamíferos terrestres del estado de oaxaca, méxico. revista mexicana de biodiversidad 86:685-710. briones-salas, m., m.c. lavariega, m. cortésmarcial, a.g. monroy-gamboa, and c.a. masesgarcía. 2016. iniciativas de conservación para los mamíferos de oaxaca, méxico. in: m. brionessalas, y. hortelano-moncada, g. magaña-cota, g. sánchez-rojas, & j.e. sosa-escalante (eds.), riqueza y conservación de los mamíferos en méxico a nivel estatal, instituto de biología, universidad nacional autónoma de méxico, mexico city. pp. 329-366. burton, a.c. 2012. critical evaluation of a long-term, locally-based monitoring program in west africa. biodiversity conservation 21:3079-3094. chávez, c., a. de la torre, h. bárcenas, r.a. medellín, h. zarza, and g. ceballos. 2013. manual de fototrampeo para estudio de la fauna silvestre. el jaguar en méxico como estudio de caso. alianza wwf-telcel, universidad nacional autónoma de méxico, mexico city. cites, convention on international trade in endangered species of wild fauna and flora. 2015. appendices i, ii, and iii. convention on international trade in endangered species of wild fauna and flora, geneva. conanp, comisión nacional de áreas naturales protegidas. 2016. programa de conservación para el desarrollo sostenible. comisión nacional de áreas naturales protegidas, mexico city. conrad, c.c., and k.g. hilchey. 2011. a review of citizen science and community-based environmental monitoring: issues and opportunities. environmental monitoring and assessment 176:273-291. danielsen, f., m.m. mendoza, a. tagtag, p.a. alviola, d.s. balete, a.e. jensen, m. enghoff, and m.k. poulsen. 2007. increasing conservation management action by involving local people in natural resource monitoring. ambio 36:566-570. de ávila, a. 2008. la diversidad lingüística y el conocimiento etnobiológico. in: conabio (ed.), capital natural de méxico, vol. i: conocimiento actual de la biodiversidad. comisión nacional para el conocimiento y uso de la biodiverssidad, mexico city. pp. 497-556. de la maza, j. 2010. áreas naturales certificadas. in: j. carabias, j. sarukhán, j. de la maza, & c. galindo (eds.), patrimonio natural de méxico cien casos de exito. comisión nacional para el conocimiento y uso de la biodiversidad, mexico city. pp. 18-19. dickinson, j.l., j. shirk, d. boner, r. bonney, r.l. crain, j. martin, t. phillips, and k. purcell. 2012. the current state of citizen science as tool for padilla-gómez et al. – camera-trap monitoring of biodiversity ecological research and public engagement. frontiers in ecology and the environment 10:291297. fegraus, e., k. lin, j.a. ahumada, c. baru, s. chancra, and c. youn. 2011. data acquisition and management software for camera trap data: a case study from the team network. ecological informatics 6:345-353. fegraus, e., and j. maccarthy. 2016. camera trap data management and interoperability. in: f. rovero, & f. zimmermann (eds.), camera trapping for wildlife research. pelagic publishing, exeter, united kingdom. pp. 33-42. flores-villela, o., and u. garcía-vázquez. 2014. biodiversidad de reptiles en méxico. revista mexicana de biodiversidad 85:s467-s475. graham, c.h., s. ferrier, f. huettman, c. moritz, and a.t. peterson. 2004. new developments in museum-based informatics and applications in biodiversity analysis. trends in ecology and evolution 19:497-503. hanssen, f., v.b. mathur, v. athreya, v. barve, r. bhardwaj, l. boumans, m. cadman, v. chavan, m. ghosh, a. lindgaard, ø. lofthus, b. mehlum, b. pandav, g.a. punjabi, a. gonzález, g. talukdar, n. valland, and r. vang, 2014. capacity building for intergovernmental platform for biodiversity and ecosystem services (ipbes). final report 2014: indo-norwegian pilot project on capacity building in biodiversity informatics for enhanced decision making, improved nature conservation and sustainable development. nina report 1079, norway. inegi. 2013. vectorial map of land use and vegetation, series v, scale 1:250,000. instituto nacional de estadística, geografía e informática, mexico. iucn. 2016. red list of threatened species. international union for conservancy of nature and natural resources, switzerland. jost, l. 2006. partitioning diversity into independent alpha and beta components. ecology 88:24272439. kok, m.t.j., k. kok, g.d. peterson, r. hill, j. agard, and s.r. carpenter. 2016. biodiversity and ecosystem services require ipbes to take novel approach to scenarios. sustainability science 12:1-5. koleff, p., t. urquiza-haas, and j. sarukhán. 2014. scientific evaluation of biological diversity: process, needs, challenges and perspectives. investigaciones ambientales 6:61-75. martin, g.j., c.i. camacho, c.a. del campo, s. anta, f. chapela, and m.a. gonzález. 2011. indigenous and community conserved areas in oaxaca, mexico. management of environmental quality 22:250-266. meffe, g., l. nielsen, r.l. knight, and d. schenborn. 2002. ecosystem management: adaptive, community-based conservation. island press, washington. molloy, j.c. 2011. the open knowledge foundation: open data means better science. plos biology 9:1-4. monroy-vilchis, o., m.m. zarco-gonzález, and c. rodríguez-soto. 2011. fototrampeo de mamíferos en la sierra nanchititla. revista de biología tropical 59:373-383. moreno, c.e., f. barragán, e. pineda, and n.p. pavón. 2011. reanálisis de la diversidad alfa: alternativas para interpretar y comparar información sobre comunidades ecológicas. revista mexicana de biodiversidad 82:12491261. morris, r.a., v. barve, m. carausu, v. chavan, j. cuadra, c. freeland, g. hagedorn, p. leary, m. mozzherin, a. olson, g. riccardi, i. teage, and g. whitbread. 2013. discovery and publishing of primary biodiversity data associated with multimedia resources: the audubon core strategies and approaches. biodiversity informatics 8:185-197. navarro-sigüenza, a.g., m.f. rebón-gallardo, a. gordillo-martínez, a.t. peterson, h. berlangagarcía, and l.a. sánchez-gonzález. 2014. biodiversidad de aves en méxico. revista mexicana de biodiversidad 85:s476-s495. nesshöver, c., b. livoreil, s. schindler, and m. vandewall. 2016. challenges and solutions for networking knowledge holders and better informing decision-making on biodiversity and ecosystem services. biodiversity conservation 25:1215-1233. ortiz-pérez, m.a., j.r. hernández, and j.m. figueroa. 2004. reconocimiento fisiográfico y geomorfológico. in: a.j. garcía mendoza, m.j. ordóñez, and m. briones-salas (eds.) biodiversidad de oaxaca. instituto de biología, universidad nacional autónoma de méxico, mexico city. pp. 43-54. padilla-gómez e., m.c. lavariega, and r. floresdiego. 2015. protocolo estandarizado para las anp que atiende las dirección sierra juárezmixteca. world wildlife fund, oaxaca, mexico. parra-olea, g., o. flores-villela, and c. mendozaalmerall. 2014. biodiversidad de anfibios en méxico. revista mexicana de biodiversidad 85:s460-s466. piwowar, h.a., t.j. vision, and m.c. whitlock. 2011. data archiving is a good investment. nature 473. padilla-gómez et al. – camera-trap monitoring of biodiversity pressman, r.s. 2006. ingeniería del software un enfoque práctico. mcgraw hill, mexico city. sánchez-cordero, v., f. botello, j.j. flores-martínez, r.a. gómez-rodríguez, l. guevara, g. gutiérrez-granados, and a. rodríguez-moreno. 2014. biodiversidad de chordata (mammalia) en méxico. revista mexicana de biodiversidad 85:s496-s504. schmeller, d.s., and p. bridgewater. 2016. the intergovernmental platform on biodiversity and ecosystem services (ipbes): progress and next steps. biodiversity conservation 25:801-805. semarnat, secretaría de medio ambiente y recursos naturales. 2010. norma oficial mexicana nom-059-semarnat-2010, protección ambiental—especies nativas de méxico de flora y fauna silvestres—categorías de riesgo y especificaciones para su inclusión, exclusión o cambio-lista de especies en riesgo. diario oficial de la federacion 2454:1-77. sergio, f., t. caro, d. brown, b. clucas, j. hunter, j. ketchum, k. mchugh, and f. hiraldo. 2008. top predators as conservation tools: ecological rationale, assumptions, and efficacy. annual review of ecology, evolution and systematics 39:1-19. tegelberg, r., j. haapala, t. mononen, m. pajari, and h. saarenmaa. 2012. the development of a digitizing service center for natural history collections. in: v. blagodev, & v.s. smith (eds.), no specimen left behind: mass digitization of natural history collections. zookeys, bulgaria. pp. 75-86. thessen, a.e., and d.j. patterson. 2011. data issues in the life sciences. in: v. blagodev, & v.s. smith (eds.), no specimen left behind: mass digitization of natural history collections. zookeys, bulgaria. pp. 15-51. treves, a., r.b. wallace, and s. white. 2009. participatory planning of interventions to mitigate human-wildlife conflicts. conservation biology 23:1523-1739. villaseñor, j.l., and e. ortiz. 2014. biodiversidad de las plantas con flores (división magnoliophyta) en méxico. revista mexicana de biodiversidad 85:s134-s142. wieczorek, j., d. bloom, r. guralnick, s. blum, m. döring, r. giovanni, t. robertson, and d. vieglais. 2012. darwin core: an evolving community-developed biodiversity data standard. plos one 7:e29715. whitlock, m.c. 2011. data archiving in ecology and evolution: best practices. trends in ecology and evolution 26:61-65. microsoft word 6507-12385-1-ed.docx biodiversity informatics, 12, 45-57 introduccion a los análisis espaciales con énfasis en modelos de nicho ecológico angela p. cuervo-robayo1,2, luis e. escobar3*, luis a. osorio-olvera1,4, javier nori5, sara varela6, enrique martínez-meyer1, jorge velásquez-tibatá7, clarita rodríguez-soto8, mariana munguía2, nora p. castañeda-àlvarez9, andrés liranoriega10, mariano soley-guardia11, josep m. serra-díaz12, a. townsend peterson13 1instituto de biología, universidad nacional autónoma de méxico, méxico distrito federal 04510, méxico. 2comisión nacional para el conocimiento y uso de la biodiversidad (conabio), insurgentes sur-periférico 4903, tlalpan, méxico distrito federal 14010, méxico. 3department of fisheries, wildlife and conservation biology, university of minnesota, estados unidos1. 4departamento de matemáticas de la facultad de ciencias, universidad nacional autónoma de méxico, ciudad de méxico, méxico. 5instituto de diversidad y ecología animal (idea-conicet), centro de zoología aplicada, universidad nacional de córdoba, rondeau 798, córdoba, argentina. 6museum fur naturkunde, leibniz institute for evolution and biodiversity science, berlin, germany. 7laboratorio de biogeografía aplicada, instituto alexander von humboldt, calle 28ª # 15-09, bogotá d.c., colombia. 8centro de estudios e investigación en desarrollo sustentable, universidad autónoma del estado de méxico. toluca, estado de méxico. 9global crop diversity trust, platz der vereinten nationen 7, 53113, bonn, germany. 10catedrático conacyt, red de estudios moleculares avanzados, instituto de ecología a.c., xalapa, veracruz. 11escuela de biología, universidad de costa rica, ciudad universitaria, 11501-2060, san pedro, costa rica. 12ecoinformatics and biodiversity, department of bioscience, aarhus university, ny munkegade 114, aarhus c dk-8000, denmark. 13biodiversity institute, university of kansas, lawrence, ks 66045, estados unidos. resumen.—en 2016 implementamos un sistema de seminarios de enseñanza, en formato de videos libres y accesibles desde internet, con la finalidad de dar a conocer de forma sencilla y en castellano, las bases conceptuales y aplicaciones de los modelos de nicho ecológico en estudios de ecología, conservación biológica, epidemiología y agrobiodviersidad, así como su implementación para el diseño de políticas públicas de los recursos naturales. cada seminario fue desarrollado por uno o varios expertos discutiendo conceptos, métodos y diferentes herramientas disponibles para elaborar modelos de distribución de especies. este manuscrito reúne los resúmenes de cada uno de los seminarios en línea, dando referencias clave para cada tema y el enlace al video correspondiente. los videos están disponibles de forma libre en youtube o en formato .mp4 bajo solicitud. palabras clave.—ecología espacial; biogeografía; conservación; distribución de especies; en línea; modelos de nicho ecológico; seminarios; video; enseñanza; capacitación en línea. 1enviar correspondencia a: e-mail: lescobar@umn.edu. 45 biodiversity informatics, 12, 45-57 abstract.—in 2016, we implemented a system of seminars for teaching, in open access video format available via the internet, aiming to show in spanish and in simple form the conceptual bases and applications of ecological niche modeling in studies of ecology, biological conservation, epidemiology, and agro-biodiversity, as well as its implementation for designing public policies for natural resources. each seminar was developed by one or several experts discussing concepts, methods, and tools available to construct species distribution models. this manuscript assembles the abstracts for each of the seminars online, providing key references for each topic and links to the corresponding videos. the videos are available freely via youtube or in .mp4 format by request. key words.—spatial analysis; biogeography; conservation; species distributions; online; ecological niche modeling; seminar; video; teaching, online training. el modelado de nicho ecológico permite estudiar la distribución geográfica de las especies e identificar aquellos factores ambientales que la limitan (peterson et al. 2011). en general, los modelos de nicho ecológico relacionan datos de presencia (y algunas veces también de ausencia) de las especies con una serie de parámetros ambientales para generar una aproximación de las condiciones que favorecen la presencia de las poblaciones de la especie (el nicho ecológico). este modelo se calcula en un espacio ambiental multidimensional para luego ser proyectado al espacio geográfico para generar un mapa que representa una distribución potencial (peterson et al. 2011). junto con otra serie de análisis espaciales, los modelos de nicho ecológico han permitido ampliar las preguntas que se abordan desde el campo de la biogeografía (peterson 2008). el modelado de nicho ecológico es un campo en constante evolución y adaptación a nuevas preguntas y métodos. por esto, los ecólogos, biólogos, epidemiólogos y biogeógrafos necesitan estar en constante actualización. en este sentido, las redes sociales son un medio eficiente para conectar y actualizar a la comunidad científica, diluyendo las barreras en espacio y tiempo, ya que permiten la comunicación en tiempo real en casi todo el mundo (kaplan & haenlein 2010); por ello, actualmente son un medio altamente utilizado entre jóvenes investigadores para el uso compartido de la información (tachibana 2014). este trabajo tiene como objetivo dar a conocer a científicos de habla hispana interesados en el modelado de nichos ecológicos y otros análisis espaciales una primera serie de seminarios en línea. se presentaron 13 seminarios por parte de 15 especialistas en el campo radicados en 6 países. según las estadísticas de la página de youtube (www.youtube.com; a enero 2017) se registraron al menos ~8500 visualizaciones a los videos, con un rango de edad de la audiencia de 25–34 años, seguido por un rango de 35–44 años. las visualizaciones se efectuaron desde 19 países, con mayor audiencia en méxico, colombia, perú, argentina y ecuador (cuadro 1), pero la presente publicación puede ampliar la audiencia de los seminarios (expandiendo el número de usuarios y la lista de países desde los cuales se efectúan las visualizaciones). los seminarios están organizados en tres grupos: (i) bases conceptuales de los modelos de nicho ecológico; (ii) aplicaciones de los modelos; y (iii) tutoriales de las herramientas disponibles. este artículo pretende estimular el interés en este campo de investigación presentando una síntesis de los seminarios ofrecidos y los enlaces a sus videos correspondientes en los que se explican las bases conceptuales y métodos modernos para generar modelos de nicho ecológico. bases conceptuales de los modelos de nicho ecológico el diagrama bam el diagrama bam (soberón & peterson 2005) es un marco conceptual de referencia para pensar en torno a los factores bióticos (b), abióticos (a) y de accesibilidad o movilidad (m) que son determinantes para explicar las áreas de distribución de las especies. cuando los conjuntos de áreas b, a y m de una especie coinciden geográficamente, se podrán encontrar poblaciones de dicha especie, ya que esas localidades serían accesibles, m, y presentarían las condiciones ambientales, a, y las interacciones interespecíficas suficientes, b, para mantener una tasa 46 biodiversity informatics, 12, 45-57 demográfica neta positiva. estas localidades constituyen el área ocupada (b ∩ a ∩ m). el diagrama bam puede usarse para entender por separado la influencia de cada uno de los tres conjuntos de factores, pero también para deducir cómo su variación espacial y temporal puede estar afectando la distribución de una especie (hutchinson 1978, soberón 2007). en este sentido b está compuesto por variables de tipo bionómico (relacionadas dinámicamente con la especie), que caracterizan interacciones interespecíficas (positivas o negativas), incluyendo recursos de los cuales depende la especie de interés. a corresponde a variables escenopoéticas o no dinámicas (no son modificadas, en un sentido amplio, por la especie), y que típicamente son más estables en el tiempo, tales como el clima de una región. finalmente, m corresponde a una hipótesis sobre el área sobre la cual la especie tiene, o ha tenido, acceso para dispersarse (barve et al. 2011; ver también anderson & raza 2010). cada uno de estos tres factores se cuantifica y está disponible de manera diferente. la información de b es, por lo general, muy escasa, ya que generalmente desconocemos cuáles son las interacciones con otras especies y cómo varían éstas a través del espacio y tiempo. gran parte de la información de a está disponible a diversas resoluciones espaciales y temporales, por ejemplo, en forma de coberturas bioclimáticas; y m debería calcularse a partir de la capacidad de dispersión de la especie a través del tiempo. en la práctica, debido al amplio desconocimiento de b, se suele operar con los factores a y m, aunque b pueda ser un factor determinante a escalas pequeñas (soberón & nakamura 2009) e incluso se ha sugerido un efecto importante de b a escalas continentales (e.g., gutiérrez et al. 2014; wisz et al. 2012). quizás lo más importante es que el diagrama bam es una referencia para pensar formalmente sobre distintas combinaciones de los tres conjuntos de factores (peterson et al. 2011) para preguntarse: ¿qué zonas son potencialmente habitables para una especie (área invadible)? ¿qué determina una localidad de presencia o ausencia de una especie? ¿cómo operan los algoritmos para hacer modelos de nicho ecológico dependiendo de las configuraciones del bam (por ej., bam clásico, mundos de wallace y hutchinson; saupe et al. 2012)? ¿los algoritmos analizan la periferialidad ambiental y riesgos de extrapolación (owens et al. 2013)? este enfoque también permite explorar una teoría más general 2https://youtu.be/px_bgt-neh8?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e de la biodiversidad que incorpore la demografía (soberón 2010). la sugerencia es que un buen ejercicio de modelado de nichos ecológicos, así como la edición posterior del modelo para aproximarlo al área de distribución deseada, debe ir acompañado de su propia configuración del diagrama bam.2 cuadro 1. lista de países en los que se ha reportado por lo menos una hora de visualización. país horas de visualización méxico 711 colombia 295 perú 102 argentina 78 ecuador 60 estados unidos 59 brasil 42 chile 29 españa 27 guatemala 15 costa rica 14 venezuela, reino unido 8 c/u bolivia 5 francia 3 nueva zelanda, canadá, uruguay, nicaragua 2 c/u aplicaciones de los modelos de nicho y modelos de distribución biogeografía de la conservación, anfibios y reptiles en la actualidad existe una crisis global de pérdida de biodiversidad. históricamente muchas decisiones para la conservación de la biodiversidad se han basado en disposiciones políticas sin dar suficiente importancia a la preservación de los ambientes naturales, su biodiversidad y los servicios ecosistémicos que proveen (margules & pressey 2000). no obstante, este tipo de decisiones son críticas en la conservación, y requieren un entendimiento de la situación al menos a escalas espaciales gruesas. en los últimos años, se ha consolidado la biogeografía de la conservación, sub-disciplina de la biología de la conservación enfocada en los procesos y patrones que operan a escalas espaciotemporales amplias y resoluciones gruesas (ladle & whittaker 2011). la biogeografía de la conservación utiliza principios, teorías y análisis de la biogeografía relacionados a la dinámica de las distribuciones de diversos taxa, para abordar 47 biodiversity informatics, 12, 45-57 problemas relacionados a la conservación de la biodiversidad (ladle & whittaker 2011). en este seminario se abordan estudios de biogeografía de la conservación que buscan generar información científica para la toma de decisiones para la conservación de los anfibios y reptiles de sudamérica. entre los ejes temáticos se destacan: (i) la determinación de zonas vulnerables a invasiones de la rana toro norteamericana (lithobates catesbianus) en argentina y, luego, a lo largo de sudamérica, considerando escenarios de cambio climático (nori et al. 2011) y zonas vulnerable a la invasión de las tortugas acuáticas más comercializadas en argentina (nori et al. 2016); (ii) el estudio de la eficiencia del sistema global de áreas protegidas para representar a los anfibios (nori et al. 2015; nori & loyola 2015); (iii) el estudio de la exposición de los reptiles de argentina al cambio climático (nori et al. 2016); y (iv) la determinación de áreas prioritarias para conservación en el gran chaco sudamericano (nori et al. 2016) y la provincia de córdoba en argentina (nori et al. 2013). estos ejemplos representan una muestra de los artículos científicos realizados exclusivamente para aportar información útil en la toma de decisiones para la conservación de la biodiversidad. algunos estudios están siendo considerados por los tomadores de decisiones o incluso han sido generados específicamente por requerimiento de éstos. sin embargo, muchas veces la valiosa información, luego de ser generada, no es utilizada o siquiera detectada por los tomadores de decisiones. en ese sentido, la interacción y articulación entre los científicos del área y las esferas políticas encargadas de la toma de decisiones resulta indispensable.3 territorios de oportunidad para mejorar la conservación considerando factores socioambientales la identificación y planeación de áreas prioritarias para la conservación, así como el uso eficiente de los recursos son aspectos fundamentales en el éxito de la conservación biológica (valenzuela & vázquez 2007). por ello, se ha desarrollado un enfoque de investigación cuyo objetivo es identificar las áreas que deben priorizarse para la distribución de los escasos recursos dedicados al manejo de la biodiversidad y desvincular estas áreas de los factores que amenazan su persistencia, por ejemplo, la 3 https://youtu.be/semcgnttld8?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_ planificación sistemática para la conservación (psc; margules & sarkar 2009). la psc debe tener en cuenta muchos criterios para garantizar un equilibrio entre la conservación de la biodiversidad y otros tipos de uso de la tierra (faleiro & loyola 2013). su eficiencia puede ser mayor si se tienen en cuenta las dimensiones sociales y humanas, incluyendo la gobernabilidad y la voluntad política (p.ej., faleiro & loyola 2013). debido a que la biodiversidad es difícil de estimar y conservar en su totalidad, se propone usar medidas parciales o subrogados que pueden ser determinadas más fácilmente (margules & sarkar 2009). en este sentido, existen especies de mayor interés para la conservación, como las especies bandera y las especies con ámbitos hogareños amplios (valenzuela & vázquez 2007). en este seminario se presenta la aplicación de la psc en la identificación de escenarios de priorización para la conservación de vertebrados terrestres en méxico, considerando factores socioambientales que pueden influir en la conservación. en este ejercicio se usó el software zonation (moilanen et al. 2011) para desarrollar diferentes escenarios de priorización a través de un análisis de costo-beneficio que reduce las limitaciones socioeconómicas y permite aprovechar las oportunidades políticas para la conservación (para un enfoque similar ver faleiro & loyola 2013). el algoritmo de priorización zonation calcula la contribución relativa de cada celda para lograr el objetivo de conservación, utilizando la regla de eliminación de superficie original “área núcleo” (moilanen et al. 2009 para más detalles). zonation ofrece la posibilidad de penalizar las zonas de acuerdo a la importancia de los factores, lo que permite un equilibrio entre los beneficios y costos para las acciones de conservación (moilanen et al. 2011). en este trabajo, a cada especie y variable se le asignó un valor de importancia diferente (ver faleiro & loyola 2013). los resultados mostraron que la mayor concentración de biodiversidad converge en regiones con gran persistencia de la cobertura natural y alta gobernabilidad local, definida como la voluntad de las personas de participar en acciones de conservación. los resultados del estudio de caso presentado en este seminario resaltan la relevancia de las variables socioeconómicas para futuros modelos de nicho ecológico, diseño de políticas ambientales e implicaciones para el cambio climático.4 4 https://youtu.be/dmykn5votc0?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e 48 biodiversity informatics, 12, 45-57 gradiente de impacto humano y biodiversidad las actividades humanas han causado cambios drásticos en los paisajes naturales; estos cambios han provocado la extinción local de algunas especies y por lo tanto pérdida de la biodiversidad. la extinción no es azarosa, si no depende de una serie de factores biológicos y ambientales. en particular, la sensibilidad de las especies a la degradación puede estar asociada a caracteres especie-específicos. por lo tanto, asociar los caracteres de las especies (p.ej., grupos tróficos, masa corporal) con atributos ambientales, permite detectar el potencial impacto ecológico de la pérdida de una especie en los hábitats, debido a la función que desempeña la especie en los ecosistemas estudiados (p.ej., los frugívoros son potenciales dispersores de semillas y los nectarívoros polinizadores). adicionalmente, esta relación permite la detección de sitios con riesgo latente para la biodiversidad mediante la identificación de los impactos que más afectan al grupo taxonómico de interés. en este seminario se presenta una propuesta para diseñar un gradiente de impacto humano de biodiversidad (gihb) basado en el análisis de ordenación rlq (dolédec et al. 1996) y "fourth corner" (dray & legendre 2008), el cual tiene como objetivo identificar especies y caracteres asociados a sitios con diferentes grados de degradación humana (munguía et al. 2016). se utilizaron datos de localidades de especies de mamíferos (conabio 2012) para seleccionar sitios de 50 km de radio con alta completitud de especies en las comunidades analizadas. se incluyeron 211 especies de mamíferos terrestres y nueve variables ambientales biofísicas, geofísicas y de impacto humano, las cuales fueron procesadas en un sistema de información geográfico (esri 2014) y analizadas con la librería ade4 (dray & dufor 2007) en el paquete estadístico r (r core team 2014). los caracteres de las especies evaluados fueron de tres tipos: modo de locomoción, hábito trófico y masa corporal. el gihb detectado para los mamíferos fue conformado principalmente por el porcentaje de cobertura de plántulas, la riqueza de plantas, el porcentaje de cobertura vegetal bajo dosel (conafor 2009) y el índice de densidad humana basado en imágenes espaciales de luces nocturnas (noaa/nesdis/ncei 2011). los resultados muestran que los caracteres asociados con sitios menos impactados por el hombre fueron los grupos tróficos carnívoros y los frugívoro 5 https://youtu.be/yxuyvgkm9g4?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e herbívoros, así como los mamíferos con tamaño corporal mayor a 17.8 kg, lo que sugiere que estos son caracteres sensibles a la degradación. por otro lado, los mamíferos con hábito trófico granívoro, y tipo de locomoción fosorial y semi-forsorial están asociados a los sitios más impactados. este marco puede ser replicado en otras áreas con menos información, en donde se necesite identificar especies en riesgo al cambio en el uso del suelo. las variables más importantes asociadas al impacto humano del gihb fueron mapeados en conjunto para resaltar los sitios con mayor riesgo para los mamíferos en el país. dada la crisis de biodiversidad y la acelerada tasa de extinción sin precedentes que actualmente enfrentamos (levin 2005; barnosky et al. 2011) existe la necesidad de continuar detectando diversos indicadores y especies sensibles a la degradación para el monitoreo de los ecosistemas.5 identificación de prioridades de conservación de los parientes silvestres de cultivos los sistemas de producción de alimentos enfrentan nuevos retos que impulsan el desarrollo de alternativas para la producción de más alimentos para una población humana en crecimiento, utilizando recursos naturales de manera eficiente, y reduciendo los impactos ambientales negativos asociados a la agricultura. con el mejoramiento genético de cultivos es posible obtener variedades para incrementar la producción por hectárea cultivada, reducir gases de efecto invernadero (subbarao et al. 2007), asimismo variedades más nutritivas y tolerantes a plagas y enfermedades (borlaug 1983). el mejoramiento genético de cultivos requiere como materia prima los recursos fitogenéticos, ya que estos ofrecen la variabilidad genética para obtener nuevas variedades (gepts 2006). los recursos fitogenéticos se pueden organizar de acuerdo a la variabilidad genética que ofrecen: variedades obtenidas a través del mejoramiento genético, variedades tradicionales y especies silvestres emparentadas con las especies cultivadas (también conocidas como parientes silvestres de cultivos). estas últimas son consideradas como portadoras de una alta diversidad genética ya que no han sido expuestas a procesos de selección propios de la domesticación (mariac et al. 2006). es así, como los parientes silvestres son utilizados para contribuir a que la agricultura sea una actividad más sostenible, debido a que son fuentes 49 biodiversity informatics, 12, 45-57 de genes en mejoramiento con el potencial para ayudar a adaptar cultivos a las condiciones ambientales extremas asociadas al cambio climático y para obtener variedades que usen el agua, la tierra y los fertilizantes de forma eficiente (guarino & lobell 2011; dempewolf et al. 2013). en este seminario se presenta una introducción sobre parientes silvestres, la cual incluye su importancia, usos en la agricultura, algunas de las amenazas que enfrentan en sus hábitats, y los objetivos globales de conservación y desarrollo que reconocen su importancia para la seguridad alimentaria global. adicionalmente, en el video se presenta en detalle la metodología utilizada por castañeda-álvarez et al. (2016) para estimar la representatividad de estos recursos fitogenéticos en bancos de germoplasma, y cómo los análisis obtenidos a través de esta metodología están siendo utilizados para establecer prioridades de conservación de las especies analizadas. el caso presentado en este video es un ejemplo práctico de cómo pueden utilizarse herramientas de modelación como maxent, y datos de acceso público, como los facilitados a través de la infraestructura mundial de información en biodiversidad (gbif6), para establecer prioridades de elementos de la naturaleza importantes para la alimentación.7 cambio global, plantas y modelos de distribución de especies el nicho ecológico de las especies vegetales ha sido ampliamente interpretado en términos de clima, puesto que el clima condiciona en gran medida las respuestas fisiológicas y la ecología de las especies. en consecuencia, los modelos de distribución de especies o modelos de nicho ecológico (franklin 2010, peterson et al. 2011) utilizan principalmente variables climáticas, ya que son generalmente accesibles, no obstante existen otras variables que también se relacionan con el desempeño fisiológico. en este seminario discutimos diferentes dimensiones de la distribución de especies comparando modelos eco-fisiológicos, modelos de nicho ecológico y modelos de interacciones. asimismo, analizamos las diferencias en las proyecciones de cambio climático bajo esos modelos. en una primera comparación entre modelos de nicho ecológico y modelos ecofisiológicos, observamos que diversas especies de árboles de la península ibérica podrían crecer en espacios climáticos que no ocupan actualmente (serra-diaz et al. 2013). 6 http://www.gbif.org/ menos frecuente fue el caso de encontrar espacios climáticos ocupados por la especie que no tuvieran un bajo rendimiento fisiológico (p.ej., crecimiento). en la comparación entre proyecciones de cambio climático realizadas con modelos ecofisiológicos y modelos de nicho correlativos, observamos que el efecto del co2 puede modificar la relación entre crecimiento y clima debido a un aumento en la eficiencia del uso de agua. por ello, incluso en el caso de que la distribución actual de especies estuviera en equilibrio climático, el no considerar las relaciones biogeoquímicas que alteran el uso de variables climáticas y su respuesta en plantas (p.ej., producción, esfuerzo reproductivo) puede influir las proyecciones realizadas a partir de los modelos correlativos (keenan et al. 2011). utilizando un modelo basado en individuos (landis-ii, scheller et al. 2007), proyectamos el cambio en la distribución de un conjunto de especies debido a un desplazamiento espacial del nicho. este experimento de modelado identificó potenciales expansiones y contracciones del área de distribución de muchas especies arbóreas debido a la interacción de características funcionales, competición-facilitación, heterogeneidad espacial y distribución de microrefugios. incluso, para especies cuyo nicho potencial se expandía con el cambio climático, el área de distribución se reducía debido a la combinación de perturbaciones y heterogeneidad espacial (serradiaz et al. 2016). como resultado de tales comparaciones, podemos entender que los modelos correlativos de nicho ecológico son en realidad una buena técnica para captar la exposición al cambio climático. así, se analizaron los cambios de distribución de la exposición al cambio climático para varias especies de árboles de california y se desarrollaron varias métricas para comparar la velocidad del cambio en la exposición de las especies. se observó que, para varias especies con distribución espacial relativamente similar, la exposición al cambio climático varía fuertemente entre ellas y entre mitad y final del siglo 21 (serra-diaz et al. 2014). en la actualidad existen varias aproximaciones y modelos que caracterizan bien diferentes dimensiones de la distribución de la especie (bam sensu soberón 2007). en el caso de plantas y cambio climático los modelos correlativos pueden caracterizar algunas facetas, pero deben ser analizados e interpretados conjuntamente con otras aproximaciones para poder proyectar 7 https://youtu.be/wifm1l_wq8q?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e 50 biodiversity informatics, 12, 45-57 cambios de distribución de plantas que sean plausibles en el curso de este siglo (franklin et al. 2016).8 mapeo de riesgo de transmisión de enfermedades los sistemas de transmisión de enfermedades representan un conjunto de especies interactuando: patógenos, vectores y hospederos. como tal, el modelado de las dimensiones geográficas del riesgo de transmisión requiere, en efecto, la integración de la ecología de la distribución de múltiples especies en un solo modelo compuesto. una aplicación útil de los modelos de nicho ecológico es la de evaluar la geografía del riesgo de transmisión de enfermedades; las diversas y complejas consideraciones involucradas en estas aplicaciones han sido revisadas en un reciente compendio que abarca un libro (peterson 2014). en este seminario, se revisan los diversos conceptos y consideraciones prácticas involucradas en el mapeo del riesgo de transmisión de enfermedades usando modelos de nicho ecológico. una bifurcación clave aparece entre los esfuerzos que buscan reconstruir el nicho ecológico y la distribución potencial de cada componente del ciclo de transmisión de la enfermedad, versus situaciones en las cuales solo está disponible la información sobre los casos de la enfermedad, lo que se define como aplicaciones de “caja negra”. se presenta una serie de ejemplos y casos de estudio para ilustrar diferentes tipos de aplica-ciones, así como algunos inconvenientes clave y problemas que se encuentran durante su desarrollo.9 herramientas disponibles nichetoolbox: de la obtención de datos de biodiversidad a la validación de los sdms el modelado de nicho es un campo de la ecología que ha permitido estimar partes del nicho ecológico y la distribución geográfica de las especies (elith & leathwick 2009). se utiliza un conjunto de herramientas estadísticas, matemáticas y computacionales para estimar la relación entre variables ambientales y la presencia de las especies (franklin 2010). el proceso de modelación de la distribución involucra por lo menos cuatro etapas: (i) obtención de datos georreferenciados de presencia (y algunas veces de ausencia) de especies, (ii) depuración de las bases de datos, 8 https://youtu.be/btby0sayuck?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e (iii) ajuste de un modelo de distribución utilizando variables de información ambiental y los registros depurados y (iv) validación del modelo. si bien existe un gran desarrollo de herramientas computacionales disponibles en r para modelar el nicho ecológico de las especies (p.ej., dismo, sdm), estas herramientas se encuentran dispersas y no cuentan con un flujo de trabajo que permita construir modelos de distribución de forma estandarizada. más aún, algunos de los programas de uso libre en cierto sentido son considerados cajas negras debido a que su código es cerrado y por lo tanto no se sabe con certeza lo que hacen y, por otro lado, aprender a usar programas de código abierto representa un gran reto para aquellas personas que no están familiarizadas con la programación. en este sentido, este seminario presenta el programa nichetoolbox (osorio-olvera et al. 2016), una plataforma y paquete de r con una interfaz de usuario amigable desarrollada en shiny (chang et al. 2016). el objetivo de nichetoolbox es facilitar el proceso de modelado de nicho ecológico y de las distribuciones de las especies. la plataforma incorpora funciones propias y otras disponibles en diferentes paquetes de r (p.ej., dismo, enmgadgets, spocc) para buscar datos de presencia, limpiar duplicados, seleccionar variables ambientales, calibrar algoritmos de nicho ecológico (p.ej., bioclim, maxent y modelos basados en elipsoides) y evaluar los modelos de distribución con las métricas dependientes de umbral, como la sensibilidad, o independientes, como el auc (area under the curve) de la curva roc (receiver operating characteristic) y la roc parcial (peterson et al. 2008). una de las características notables de nichetoolbox es que cuenta con funciones para descargar el flujo de trabajo (workflow) de lo que el usuario ha hecho dentro de la aplicación. este flujo de trabajo, además de contener los archivos de los análisis realizados en la sesión, guarda en un documento el código de r con el que los produjo. con lo anterior se pretende hacer que el proceso de modelado sea transparente, además de que los usuarios aprendan a programar en r mientras realizan su investigación.10 nichea los modelos de nicho ecológico buscan una aproximación de la distribución de las especies a través de la estimación sus nichos ecológicos (ver 9 https://youtu.be/ybzllisuufy?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e 10 https://youtu.be/cco3kfwzal4?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e 51 biodiversity informatics, 12, 45-57 el diagrama de bam). así, los modelos de nicho ecológico son análisis generados en espacios ambientales que pueden ser proyectados al espacio geográfico (warren 2012). no obstante, muchas veces los estudios son diseñados, evaluados, e interpretados considerando únicamente la geografía (p.ej., peterson et al. 2008; fischer et al. 2014; radosavljevic & anderson 2014). por otro lado, los avances en los métodos y las variables disponibles para la calibración de los modelos de nicho ecológico han permitido avanzar considerablemente en este campo (escobar & craft 2016). por ejemplo, cada vez se desarrollan y proponen nuevos algoritmos para el modelado de los nichos ecológicos (p.ej., blonder et al. 2014; drake 2015; qiao et al. 2015). estos nuevos algoritmos requieren de rigurosas evaluaciones para validar su aplicación en la práctica (peterson et al. 2008), y las evaluaciones a su vez requieren datos robustos y confiables para determinar si los modelos pueden o no reconstruir los nichos ecológicos y la distribución de las especies. sin embargo, los datos frecuentemente sufren de sesgo en el muestreo o no representan correctamente el nicho ecológico de las especies, lo que limita la correcta evaluación de los modelos (kadmon et al. 2004). en este sentido, han surgido las especies virtuales como alternativa que permite controlar la forma, posición y amplitud de un nicho ecológico. además, permiten manejar la respuesta de las especies a las variables ambientales y controlar la distribución espacial de las muestras para obtener muestras con o sin sesgo en el muestreo. en este seminario se presenta el programa nichea, que permite diseñar especies virtuales en espacios ambientales multidimensionales (qiao et al. 2016). nichea está construido en un ambiente java enlazado a r, permitiendo desarrollar nichos virtuales en una plataforma amigable (figura 1). una vez creada la especie virtual, nichea permite generar desde análisis descriptivos básicos hasta análisis predictivos multivariados complejos. la versatilidad y facilidad de uso de nichea hacen de este programa una herramienta ideal en docencia en cursos de biología básica o capacitaciones avanzadas para el modelado de nichos ecológicos. en este video, se explican las bases conceptuales y teorías ecológicas que respaldan el uso de nichea para crear especies virtuales. además, se muestran aspectos introductorios del funcionamiento de nichea, el 11 https://youtu.be/csbod4yqdtg?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e 12 https://www.cs.princeton.edu/~schapire/maxent/datasets/samples.zip 13 parte i: https://youtu.be/2pkc-xfmbo8?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e manejo de los datos, aplicaciones de las especies virtuales creadas y material de ayuda en línea para el proceso de instalación y utilización del programa. finalmente, se muestra cómo nichea puede ayudar a interpretar y comparar modelos en espacios ambientales, convirtiendo a nichea en una poderosa herramienta de última generación para el modelado e interpretación de nichos ecológicos en ecológica moderna.11 modelamiento de distribución de especies con maxent en r: tutorial y fundamentos el modelage de distribuciones de especies es un campo de rápido crecimiento en el que nuevos programas, métodos y recomenda-ciones de buenas prácticas son publicados continuamente. un elemento frecuente en los avances recientes en el campo es el uso del lenguaje estadístico r (r core team 2016), el cual permite el desarrollo de flujos completos de trabajo, facilita la replicación y permite el desarrollo de experimentos virtuales de alta complejidad. en este seminario se presenta una introducción al modelamiento de distribuciones de especies en r usando maxent (phillips et al. 2016), a través de dismo (hijmans et al. 2016) y un conjunto de datos para bradypus variegatus (phillips et al. 2016, disponible en12). inicialmente se facilita una “traducción” de las opciones de ingreso de datos y desarrollo de modelos con la configuración predeterminada de maxent. posteriormente se explica cómo (i) evaluar modelos y seleccionar umbrales usando validación cruzada, auc y tss; (ii) transferir modelos a otras épocas/regiones y (iii) configurar distintos argumentos para el desarrollo de modelos y su predicción en el espacio geográfico en maxent a través de r. finalmente se presentan dos implementaciones de buenas prácticas recomendadas en la literatura: (1) la optimización del multiplicador de regularización (warren & seifert 2011) y (2) el muestreo de datos de entorno del área accesible, m (anderson & raza 2010; barve et al. 2011). este tutorial está dirigido a usuarios con conocimientos mínimos del uso de maxent y de la sintaxis de r13,14. el código empleado en las demostraciones se encuentra disponible en: 15. enmeval: la teoría detrás del paquete en el modelaje existe un clásico balance entre generalidad y sobreajuste, en el cual la complejidad óptima se encuentra entre ambos 14 parte ii: https://youtu.be/q-x3vxppq5c?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e 15 https://github.com/jivelasquezt/courses 52 biodiversity informatics, 12, 45-57 figura 1. interface de usuario de nichea. los botones superiores (file, toolbox, workflow, view, niches, about nichea) despliegan herramientas para selección, análisis de las variables ambientales y nichos e información general sobre el programa. las herramientas laterales (izquierda) permiten diseñar el nicho ecológico de una especie virtual de forma manual (barras) o a través de valores de la dimensión y posición del nicho (cuadros blancos). derecha: las condiciones ambientales del área de estudio y los modelos de nicho son proyectados en un espacio ambiental de una (x), dos (y), o tres dimensiones (z). 53 biodiversity informatics, 12, 45-57 extremos y puede ser determinada mediante evaluaciones rigurosas (peterson et al. 2011). en este seminario se discuten los conceptos básicos detrás de la elaboración y evaluación de un modelo correlativo de nicho ecológico haciendo énfasis en el programa de modelación maxent (phillips et al. 2006). seguidamente, se exponen y explican las limitantes importantes que presentan las evaluaciones automatizadas de dicho programa, y cómo éstas pueden ser solucionadas mediante el paquete complementario de r (r core team 2016) enmeval (muscarella et al. 2014). en particular, maxent permite extraer relaciones altamente complejas entre la especie y el ambiente, las cuales pueden ser más representativas de lo que realmente ocurre en sistemas naturales, otorgando así una alta capacidad predictiva (olden et al. 2008; evans et al. 2013). no obstante, la complejidad potencial de maxent también puede resultar en un serio sobreajuste a la información existente. en este caso, los modelos pueden convertirse más bien en una representación de los problemas típicamente asociados a datos (p.ej., muestreo incompleto o sesgado) reduciendo su capacidad predictiva e incluso invalidando su utilidad en general (anderson 2012; merow et al. 2013). mientras maxent permite al usuario variar la complejidad potencial del modelo mediante el uso de distintas características de los argumentos (feature classes) y niveles de regularización (phillips & dudík 2008), las evaluaciones implementadas por este software no toman en cuenta la autocorrelación espacial presente entre los datos de calibración y evaluación. esto resulta en valores artificialmente inflados para las métricas de evaluación (veloz 2009), lo que no solo dificulta la determinación de una complejidad óptima, sino que conlleva a que los usuarios ni siquiera intenten calibrar modelos alternativos. para solventar este problema, enmeval permite al usuario calibrar múltiples modelos en maxent desde la plataforma de r y simultáneamente evaluarlos mediante diversas métricas y particiones de los datos (muscarella et al. 2014). estas particiones incluyen esquemas que reducen o eliminan la autocorrelación espacial (radosavljevic & anderson 2014) y se ajustan a distintos requerimientos según el tamaño de muestra y la configuración geográfica del sistema. por su lado, las diversas métricas permiten evaluar los modelos de acuerdo a distintos criterios, como lo son la tasa de omisión (omission rate: or), 16https://youtu.be/07vecd2os8u?list=plu_3tlnpcpdzd0vvvxvct9xhcnx7l9e_e discriminación del entorno (auc) y complejidad en general (aicc). de esta manera, enmeval permite al usuario escoger los modelos que mejor se ajusten a su determinado sistema y objetivos en general.16 consideraciones finales los videos en vivo permitieron la interacción de los expositores con la audiencia en tiempo real, mientras otros usuarios continúan interactuando con los expositores a través del sitio web de los videos a partir de comentarios o correos electrónicos. en conclusión, la constante evolución de modelado de nichos y distribución de especies hace necesario un constante desarrollo de plataformas de interacción y actualización para la comunidad científica. más esfuerzos son necesarios para que las nuevas generaciones de investigadores y revistas científicas tomen ventaja de las nuevas tecnológicas para la trasferencia de información. el uso de redes sociales y seminarios en líneas son una prometedora y sustentable plataforma para la interacción entre expertos y estudiantes. agradecimientos los seminarios en línea se realizaron con el apoyo de la dirección de análisis y prioridades de la comisión nacional para el conocimiento y uso de la biodiversidad (conabio), y el laboratorio de análisis espaciales del instituto de biología de la universidad nacional autónoma de méxico (unam). lee fue apoyado por el minnesota environment and natural resources trust fund, the minnesota aquatic invasive species research center, and the clean water land and legacy, y agradece a huijie qiao por su incansable esfuerzo para el desarrollo de nichea. aln agradece a jorge lobo, octavio rojas y juan l. parra por los comentarios. loo agradece al posgrado en ciencias biológicas de la unam, al papiit-unam in112715 (2015) y a gsoc 2016 por el apoyo parcial brindado para la realización de su investigación. agradecemos a jorge soberón y eliécier gutiérrez por sus comentarios, los cuales permitieron mejorar la versión final de este artículo. referencias anderson, r. p. 2012. harnessing the world’s biodiversity data: promise and peril in ecological niche modeling of species distributions. ann. n. y. acad. sci. 1260:66–80. 54 biodiversity informatics, 12, 45-57 anderson, r. p., & a. raza. 2010. the effect of the extent of the study region on gis models of species geographic distributions and estimates of niche evolution: preliminary tests with montane rodents (genus nephelomys) in venezuela. j. biogeogr. 37:1378–1393. barnosky, a. d., n. matzke, s. tomiya, g. o. u. wogan, b. swartz, t. b. quental, c. marshall, j. l. mcguire, e. l. lindsey, k. c. maguire, b. mersey, & e. a. ferrerer. 2011. has the earth’s sixth mass extinction already arrived? nature 471:51–57. barve, n., v. barve, a. jiménez-valverde, a. liranoriega, s. p. maher, a. t. peterson, j. soberón, & f. villalobos. 2011. the crucial role of the accessible area in ecological niche modeling and species distribution modeling. ecol. modell. 222:1810–1819. blonder, b, c. lamanna, c. violle, & b. j. enquist. 2014. the n-dimensional hypervolume. glob. ecol. biogeogr. 23:595–609. borlaug, n. e. 1983. contributions of conventional plant breeding to food production. sci. 219:689– 693. castañeda-álvarez, n. p., c. k. khoury, h. a. achicanoy, v. bernau, h. dempewolf, r. j. eastwood, l. guarino, r. h. harker, a. jarvis, n. maxted, j. v. müller, j. ramirez-villegas, c. c. sotelo, p. c. struik, h. vincent, j. toll. 2016. global conservation priorities for crop wild relatives. nat. plants 2:16022. chang w., j. cheng, j. j. allaire, y. xie, & j. mcpherson. 2016. shiny: web application framework for r. r package version 0.13.2. https://cran.r-project.org/package=shiny conabio, comisión nacional para el conocimiento y uso de la biodiversidad. 2012. base de datos de localidades de especies de mamíferos terrestres. sistema nacional de información sobre biodiversidad de méxico (snib). consultado en línea en el 2012: http://www.conabio.gob.mx/institucion/snib/docto s/acerca.html conafor, comisión nacional forestal. 2009. inventario nacional forestal y de suelos, méxico 2004–2009. dempewolf, h., r. j. eastwood, l. guarino, c. k. khoury, j. v müller, & j. toll. 2013. adapting agriculture to climate change: a global initiative to collect, conserve, and use crop wild relatives. agroecol. and sust. food. 38:369–77. dolédec, s., d. chessel, f. c. j. ter braak, & s. champely. 1996. matching species traits to environmental variables: a new three-table ordination method. environ. ecol. stat. 3:143–146. drake, j. m. 2015. range bagging: a new method for ecological niche modelling from presence-only data. j. r. soc. interface. 12:1–9. dray, s., & a. b. dufor. 2007. the ade4 package: implementing the duality diagram for ecologists. j. stat. softw. 22:1–20. dray, s., & p. legendre. 2008. testing the species traits-environment relationships: the fourth-corner problem revisited. ecology 89:3400–3412. elith, j., & j. leathwick. 2009. species distribution models: ecological explanation and prediction across space and time. annu. rev. ecol. evol. syst. 40:677–697. escobar, l. e., & m. e. craft. 2016. advances and limitations of disease biogeography using ecological niche modeling. front. microbiol. 7:1– 21. gutiérrez, e. e., boria r. a. & r. p. anderson. 2014. can biotic interactions cause allopatry? niche models, competition, and distributions of south american mouse opossums. ecography 37: 741753. evans, m. r., v. grimm, k. johst, t. knuuttila, r. de langhe, c. m. lessells, m. merz, m. a. o’malley, s. h. orzack, m. weisberg, d. j. wilkinson, o. wolkenhauer, & t. g. benton. 2013. do simple models lead to generality in ecology? trends ecol. evol. 28:578–583. faleiro, f. v., & r. d. loyola. 2013. socioeconomic and political trade-offs in biodiversity conservation: a case study of the cerrado biodiversity hotspot, brazil. divers. distrib. 19:977–987. fischer, d., s. m. thomas, m. neteler, n. b. tjaden, & c. beierkuhnlein. 2014. climatic suitability of aedes albopictus in europe referring to climate change projections: comparison of mechanistic and correlative niche modelling approaches. euro surveill. 19:1–13. franklin, j. 2010. mapping species distributions: spatial inference and prediction. cambridge: cambridge university press. franklin, j., j. m. serra-diaz, a. d. syphard, & h. m. regan. 2016. global change and terrestrial plant community dynamics. proc. natl. acad. sci. u.s.a. 113:3725–3734. gepts, p. 2006. plant genetic resources conservation and utilization: the accomplishments and future of a societal insurance policy. crop sci. 46: 2278– 2292. guarino, l. & d. b. lobell. 2011. a walk on the wild side. nat. clim. change. 1:374–375. hijmans, r. j., s. phillips, j. leathwick, & j. elith. 2016. dismo: species distribution modeling. r package version 1.0–15. https://cran.rproject.org/package=dismo hutchinson, g. e. 1957. concluding remarks. cold spring harb. symp. quant. biol. 22:415–427. hutchinson, g. e. 1978. an introduction to population ecology. new heaven: yale university press. kadmon, r., o. farber, & a. danin. 2004. effect of roadside bias on the accuracy of predictive maps produced by bioclimatic models. ecol appl. 14:401–413. 55 biodiversity informatics, 12, 45-57 kaplan, a. m., & haenlein, m. 2010. users of the world, unite! the challenges and opportunities of social media. bus horiz. 53:59–68. keenan, t., j. m. serra, f. lloret, m. ninyerola, & s. sabate. 2011. predicting the future of forests in the mediterranean under climate change, with niche and process-based models: co2 matters! glob chang biol. 17:565–579. ladle, r., & r. j. whittaker. editors. 2011. conservation biogeography. john wiley & sons. chichester, uk: wiley-blackwell. margules, c. r., & r. l. pressey. 2000. systematic conservation planning. nat. 405: 243–53. margules, c. r., & s. sarkar. 2009. planeación sistemática de la conservación. (trad. v. sánchezcordero & f. figueroa). universidad nacional autónoma de méxico, comisión nacional de áreas naturales protegidas y comisión nacional para el conocimiento y uso de la biodiversidad. 304 pp. méxico d.f. (original en inglés, 2007). mariac, c., v. luong, i. kapran, a. mamadou, f. sagnard, m. deu, j. chantereau, b. gerard, j. ndjeunga, g. bezançon, j. l. pham, & y. vigouroux. 2006. diversity of wild and cultivated pearl millet accessions (pennisetum glaucum [l.] r. br.) in niger assessed by microsatellite markers. theor. appl. genet. 114: 49–58. merow c., m. j. smirh, & j. a. silander. 2013. a practical guide to maxent for modeling species’ distributions: what it does, and why inputs and settings matter. ecography. 36:1058–1069. moilanen, a., h. p. possingham, & k. a. wilson. 2009. spatial conservation prioritization: past, present and future. en: a. moilanen, k. a. wilson, & h. p. possingham (eds.) spatial conservation prioritization: quantitative methods and computational tools. oxford university press, oxford, pp. 260–268. moilanen, a., b. j. anderson, f. eigenbrod, a. heinemeyer, d. b. roy, s. gillings, p. r. armsworth, k. j. gaston, & c. d. thomas. 2011. balancing alternative land uses in conservation prioritization. ecol appl. 21:1419–1426. munguía, m., i. trejo, c. gonzález-salazar, & o. pérez-maqueo. 2016. human impact gradient on mammalian biodiversity. glob. ecol. conserv. 6:79–92. muscarella, r., p. j. galante, m. soley-guardia, r. a boria, j. m. kass, m. uriarte, & r. p. anderson. 2014. enmeval: an r package for conducting spatially independent evaluations and estimating optimal model complexity for maxent ecological niche models. methods. ecol. evol. 5:1198–1205. noaa/nesdis/ncei. 2011. dmsp-ols nighttime light data. http://ngdc.noaa.gov/eog/dmsp/downloadv4comp osites.html nori, j., m. s. akmentins, r. ghirardi, n. frutos, & g. c. leynaud. 2011. american bullfrog invasion in argentina: where should we take urgent measures? biodivers. conserv. 20:1125–1132. nori, j. d. l. m. azócar, f. b. cruz, m. f. bonino, & g. c. leynaud. 2016. translating niche features: modelling differential exposure of argentine reptiles to global climate change. austral. ecol. 41:373–381. nori, j., p. lemes, n. urbina-cardona, d. baldo, j. lescano, & r. loyola. 2015. amphibian conservation, land-use changes and protected areas: a global overview. biol. conserv. 191:367–374. nori, j., j. n. lescano, p. illoldi-rangel, n. frutos, m. r. cabrera, & g. c. leynaud. 2013. the conflict between agricultural expansion and priority conservation areas: making the right decisions before it is too late. biol. conserv. 159:507–513. nori, j., & r. loyola. 2015. on the worrying fate of data deficient amphibians. plos one 10:1–8. nori, j. g. tessarolo, g. f. ficetola, v. di cola, r. loyola, & g. c. leynaud. 2016. buying environmental problems: the invasive potential of imported turtles in argentina. aquat. conserv. in press. nori, j., r. torres, j. n. lescano, j. m. cordier, m. e. periago, & d. baldo. 2016. protected areas and spatial conservation priorities for endemic vertebrates of the gran chaco, one of the most threatened ecoregions of the world. divers. distrib. 22:1212–1219. nori, j., j. n. urbina-cardona, r. d. loyola, j. n. lescano, & g. c. leynaud. 2011. climate change and american bullfrog invasion: what could we expect in south america? plos one 6: e25718. olden, j. d., j. j. lawler, & n. l. poff. 2008. machine learning methods without tears: a primer for ecologists. q. rev. biol. 83:171–193. osorio-olvera l., v. barve, n. barve, & j. soberón. 2016. nichetoolbox: from getting biodiversity data to evaluating species distribution models in a friendly gui environment. r package version 0.1.6.0. https://github.com/luismurao/nichetoolbox owens, h. l., l. p. campbell, l. l. dornak, e. e. saupe, n. barve, j. soberón, k. ingenloff, a. liranoriega, c. m. hensz, & c. e. myers. 2013. constraints on interpretation of ecological niche models by limited environmental ranges on calibration areas. ecol. modell. 263:10–18. peterson, a. t. 2014. mapping disease transmission risk in geographic and ecological contexts. baltimore: johns hopkins university press. peterson, a. t., m. papeş, & j. soberón. 2008. rethinking receiver operating characteristic analysis applications in ecological niche modeling. ecol. modell. 213:63–72. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, & m. b. araújo. 2011. ecological niches and geographic distributions. princeton: princeton university press. 56 biodiversity informatics, 12, 45-57 phillips, s. j., & m. dudík. 2008. modeling of species distributions with maxent: new extensions and a comprehensive evaluation. ecography 31:161–175. phillips, s. j., r. p. anderson, & r. e. shapire. 2006. maximum entropy modeling of species geographic distributions. ecol. modell. 190:231–259. qiao, h., c. lin, z. jiang, & l. ji. 2015. marble algorithm: a solution to estimating ecological niches from presence-only records. sci. rep. 5:14232. qiao, h., a. t. peterson, l. p. campbell, j. soberón, l. ji, & l. e. escobar. 2016. nichea: creating virtual species and ecological niches in multivariate environmental scenarios. ecography 39:805–813. r core team. 2016. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. url https://www.r-project.org/. radosavljevic, a., & r. p. anderson. 2014. making better maxent models of species distributions: complexity, overfitting and evaluation. j. biogeogr. 41:629–43. saupe, e. e., v. barve, c. e. myers, j. soberón, n. barve, c. m. hensz, a. t. peterson, h. l. owens, & a. lira-noriega. 2012. variation in niche and distribution model performance: the need for a priori assessment of key causal factors. ecol. modell. 237:11–-22. scheller, r. m., j. b. domingo, brian r. sturtevant, j. s. williams, a.d. rudy, e.j. gustafson, & d. j. mladenoff. 2007. design, development, and application of landis-ii, a spatial landscape simulation model with flexible temporal and spatial resolution. ecol. modell. 201:409–419. serra-diaz, j. m., j. franklin, miquel ninyerola, f.w. davis, a. d. syphard, h. m. regan, & m. ikegami. 2014. bioclimatic velocity: the pace of species exposure to climate change. divers. distrib. 20:169–180. serra-diaz, j. m., t. f. keenan, m. ninyerola, s. sabaté, c. gracia, & f. lloret. 2013. geographical patterns of congruence and incongruence between correlative species distribution models and a process-based ecophysiological growth model. j. biogeogr. 40:1928–1938. serra-diaz, j. m., r. m. scheller, a. d. syphard, & j. franklin. 2015. disturbance and climate microrefugia mediate tree range shifts during climate change. landsc. ecol. 30:1039–53. soberón, j. 2007. grinnellian and eltonian niches and geographic distributions of species. ecol. lett. 10:1115–1123. soberón, j. 2010 niche and area of distribution modeling: a population ecology perspective. ecography 33:159–167 soberón, j., & m. nakamura. 2009. niches and distributional areas: concepts, methods, and assumptions. proc. natl. acad. sci. u.s.a. 106:19644–19650. soberón, j., & a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species' distributional areas. biodivers. infor. 2:1– 10. subbarao, g. v., ban tomohiro, kishii masahiro, ito osamu, h. samejima, h. y. wang, s. j. pearse, s. gopalakrishnan, k. nakahara, a. k. m. zakir hossain, h. tsujimoto & w.l. berry. 2007. can biological nitrification inhibition (bni) genes from perennial leymus racemosus (triticeae) combat nitrification in wheat farming? plant soil. 299:55– 64. tachibana, c. 2014. a scientist's guide to social media. science. 343:1032–1035. valenzuela, d. g., & l. b. vázquez. 2007. consideraciones para priorizar la conservación de carnívoros mexicanos. en: g. sánchez-rojas, & a. rojas-martínez (eds) tópicos en sistemática, biogeografía ecología y conservación de mamíferos. universidad autónoma del estado de hidalgo, méxico. pp. 197–214. veloz, s. d. 2009. spatially autocorrelated sampling falsely inflates measures of accuracy for presence only niche models. j. biogeogr. 36:2290–2299. warren, d. l., & s. n. seifert. 2011. ecological niche modeling in maxent: the importance of model complexity and the performance of model selection criteria. ecol. appl. 21:335–342. warren, d. l. 2012. in defense of ‘niche modeling’. trends ecol. evol. 27:497–500. wisz, m. s. j. pottier, w. d. kissling, l. pellissier, j. lenoir, c.f. damgaard, m. f. forchhammer, j-a. grytnes, a. guisan, r. k. heikkinen, t. t. hoye, i. kuhn, m. luoto, l. maiorano, m-c. nilsson, s. normand, e. ockinger, n. m. schmidt, m. termansen, a. timmermann, d. a. wardle, p. aastrup, j-c. svenning. 2012. the role of biotic interactions in shaping distributions and realised assemblages of species: implications for species distribution modelling. biol. rev. 88: 15–30. 57 microsoft word enmevaluation_final.docx biodiversity  informatics,  8,  2012,  p.  41   biodiversity informatics training module niche modeling: model evaluation a. townsend peterson biodiversity institute, university of kansas, lawrence, kansas 66045 usa abstract.—ecological niche modeling has become a very popular tool in ecological and biogeographic studies across broad extents. the tool is used in hundreds of publications each year now, but some fundamental aspects of the approach have seen a fair amount of carelessness. among these aspects is that of model evaluation or validation. this unit provides a simple introduction to the basic concepts that are important in model validation, using some very simple examples. the focus is on two solutions to the challenge of model evaluation: a simply cumulative binomial approach that can be used with binary model outputs, and partial roc analysis, which can be used with continuous model outputs. exercises, illustrations, and reading materials relevant to this training module can be downloaded from http://hdl.handle.net/1808/10059. note that one file has the extension ".xyz". this extension must be changed to ".exe", as it is a program for partial roc analyses and should be installed on your computer. video  links  (via  youtube):   part  1  –  http://www.youtube.com/watch?v=ntlp8opcpl8   part  2  –  http://www.youtube.com/watch?v=scjcqj8m1ri   part  3  –  http://www.youtube.com/watch?v=xgajiidwqh0   part  4  –  http://www.youtube.com/watch?v=rvxnuq_yedy   part  5  –  http://www.youtube.com/watch?v=qnljkzra61q         for  more  information  and  more  training  modules:   http://biodiversity-­‐informatics-­‐training.org/     microsoft word hortal&lobo06v2 corrected proofs.doc biodiversity informatics, 3, 2006, pp. 16-45 16 towards a synecological framework for systematic conservation planning joaquín hortal dpto. biodiversidad y biología evolutiva, museo nacional de ciencias naturales (csic), 28006 madrid, spain, center for macroecology, institute of biology, university of copenhagen, 2100 copenhagen o, denmark, and depto. ciências agrárias – cita-a, universidade dos açores, 9700-851 angra do heroísmo, terceira, azores, portugal e-mail: jhortal@mncn.csic.es and jorge m. lobo dpto. biodiversidad y biología evolutiva, museo nacional de ciencias naturales (csic), madrid 28006, spain e-mail: mcnj117@mncn.csic.es abstract.− biodiversity conservation design, though difficult with fragmentary or insufficient biological data, can be planned and evaluated with several methods. one of them, the complementarity criterion, is commonly used to account for the distributions of a number of species (i.e., an autoecological approach). at the same time, the patchiness and spatial bias of available distribution data has also been dealt with through distribution modelling. however, both the uncertainty of the ranges estimated and the changes in species’ distributions in response to changing climates, limit the potential of single-species distributions as the biodiversity attribute to be used in complementarity strategies. several technical and theoretical advantages of composite biodiversity variables (i.e., a synecological approach) may, however, make them ideal biodiversity indicators for conservation area selection. the drawbacks associated with current biodiversity data are discussed herein, along with the possible advantages and disadvantages of conservation planning through a synecological or autoecological approach. key words.− biodiversity, complementarity, conservation biogeography, predictive modelling, climate change, surrogacy, emergent properties, systematic conservation planning conservation biogeography has recently been defined as “the application of biogeographical principles, theories, and analyses, being those concerned with the distributional dynamics of taxa individually and collectively, to problems concerning the conservation of biodiversity” (whittaker et al. 2005). here, the phenomenon of biodiversity at its whole width becomes the central target for conservation from a biogeographic perspective. however, current knowledge of biodiversity patterns and processes is not reliable enough for the models and scenarios needed for conservation policy decisions (see discussion in whittaker et al. 2005). such drawback has given rise to much debate as to what to protect, how to deal with available data insufficiency by means of surrogacy, and how to include ecosystem processes and natural services in systematic conservation planning processes (see, e.g., brooks et al. 2004a,b; molnar et al. 2004; higgins et al. 2004; cowling et al. 2004; pressey 2004, and references therein). conservation biology needs a solid framework from which to work with the current fragmentary state of biodiversity data, while at hortal and lobo – synecological conservation framework 17 the same time addressing the urgent need for a ‘evaluation scheme’, whereby conservation values can be input to the decision-making process (see green et al. 2005 for such a scheme). whittaker and collaborators (2005) identify four inter-related conservation biogeography weaknesses: i) inadequacies in taxonomic and distribution data; ii) spatial and temporal scale dependency; iii) effects of model structure and parameterization and iv) inadequacies of theory. this work aims to develop the basis for a framework to deal with the first of the above-mentioned weaknesses (and partly accounting for the latter) in the sub-discipline of systematic conservation planning (margules and pressey 2000). leaving apart the less effective expert assessment (although the gains it provides should be integrated in conservation planning; see cowling et al. 2003a), three approaches could be used as a basis for systematic conservation planning schemes: • the environment-based approach -use environmental variation as a biodiversity descriptor. • the autoecological approach (herein autoecology for short) -gather observed or predicted distribution information for a great number of species, and use them as surrogates of overall diversity. • the synecological approach (herein synecology) -use composite biodiversity variables (i.e. synecological variables; e.g., species richness, rarity, endemism, community composition, etc.), either observed or estimated, to describe the whole patterns of biodiversity variation. considerable effort would be required to gather the amount of information needed for a reliable use of the autoecology approach. therefore, as environmental information is more readily available, many authors favor the environment-based approach, which has been used during the development of several successful coarse-scale conservation schemes (e.g., the iucn-wpca reserve network approaches of the 1970s and 1980s; mackinnon and mackinnon 1986a,b; mackinnon et al. 1986). however, environmental variation is just a surrogate for biodiversity, rather than a conservation target. therefore, the drawbacks of the environmentbased approach should better be overcome using biodiversity data, which is a true conservation target itself. from the approaches based on biodiversity data, the autoecological has been the most used, in spite of presenting practical and theoretical problems that could hamper its reliability (see below). a mixed approach combining autoecology and synecology could be using together biological and environmental data to model the distribution of the species of interest, and then aggregate their predictions to construct biodiversity variables, or use them for complementarity analyses (see, e.g., rojassoto et al. 2003 or ferrier and guisan 2006; a. lira-noriega et al. unpubl.). such approach lies under the assumption that errors in the predictions of the distributions of a number of species will minimize if they are used altogether. we here propose that the effort required in the drawing up of biogeography conservation schemes could be minimized, and the process simplified, by the direct gathering (and use) of synecological information instead of autoecological one. we examine the drawbacks associated with the current state of biodiversity data and discuss the technical and theoretical advantages of using synecological variables in systematic conservation planning (see box 1). geographic representation of biodiversity biodiversity is a heterogeneous concept accounting for the multi-faceted, continuous variability of life, from genes to ecosystems. however, even though recent proposals hortal and lobo – synecological conservation framework 18 box 1. schematic swot analysis for the three systematic conservation planning approaches discussed: environment-based (e), autoecology (a) and synecology (s) (see introduction for definitions). tables on the right present the relevant characteristics of each approach. strengths (s) and weaknesses (w) are structural characteristics, and are intended as not to change in the short and/or medium-term. on the contrary, opportunities (o) and threats (t) are suited to the current status of the theoretical and practical knowledge on biodiversity and conservation fields, and could change in the near future. items are numbered according to the approach they pertain (e, a, s), and the kind of characteristic they are (s, w, o and t) (e.g., es1). swot analysis is based on studying the relevant combinations of s, w, o and t items inside of each approach through a twos matrix (matrices placed on the left), to identify its main drawbacks and advantages (see dyson and o’brien 1998 and dyson 2004 for further reviews on the method). in addition to the twos matrices, we identify i) the structural conflicts between strengths and weaknesses (represented over the twos matrices), and ii) current limitations to opportunities caused threats (i.e., the state-of-the-art) (represented below twos matrices). e. environment-based approach strengths weaknesses es1.rapid mapping techniques es2.information availability (gis and remote sensing) es3.fine-grain geographic resolution ew1.not a true conservation target (surrogate) ew2.land classifications�abstract human constructs opportunities threats eo1.ecosystem representation eo2.vegetation type representation eo3.ecological processes representation eo4.species distribution representation (relationship with environment) et1.lack of knowledge of the relationship between land classes and biodiversity variations et2.risk of exclusive use of surrogacy due to information easy availability (es1+es2+es3)xew1 (es1+es2)xew2 s w o (es1+es2)xeo1 (es1+es2)xeo2 es3xeo2 (es1+es2)xeo3 (es1+es2)xeo4 ew1x(eo1+eo2) ew1xeo3 ew1xeo4 ew2xeo3 ew2xeo4 t (es1+es2)xet1 (es1+es2)xet2 ew1xet2 ew2xet1 (eo1,eo2,eo3)xet1 a. autoecology approach strengths weaknesses as1.true conservation target as2.true biodiversity representation as3.representation of key species important for ecosystem processes aw1.incomplete information aw2.taxonomically biased information aw3.geographically biased information aw4.high amount of variables (species) aw5.represents only patterns, not processes opportunities threats ao1.development of gbif and other biodiversity databases ao2.predictive modelling ao3.good knowledge on population biology at1.lack of taxonomic work at2.model error aggregation at3.drawbacks for the use of rare species at4.fragmentary knowledge of the relationship between key species and ecosystem processes (as1+as2)x(aw1+aw2+aw3+aw4) as3x(aw1+aw2+aw3) s w o (as1+as2)xao1 (as1+as2)xao2 as3xao1 as3xao2 as3xao3 aw1x(ao1+ao2) aw2xao1 aw3x(ao1+ao2) aw5xao3 t as1xat1 (as1,as2,as3)xat3 as2xat2 as3xat4 (aw1,aw2)xat1 (aw1,aw3)xat2 aw4xat2 (aw1,aw2)xat3 ao1xat1 ao2x(at2,at3) ao3xat4 s. synecology approach strengths weaknesses ss1.true conservation target ss2.true biodiversity representation ss3.use of rare species ss4.reduced number of variables ss5.representation of emergent processes sw1.incomplete information sw2.taxonomically biased information sw3.geographically biased information opportunities threats so1.development of gbif and other biodiversity databases so2.variables can be estimated locally so3.predictive modelling st1.lack of taxonomic work st2.partial knowledge of the relationship between diversity and ecosystem processes (ss1+ss2)x(sw1+sw2+sw3) (ss3,ss4,ss5)x(sw1+sw2+sw3) s w o (ss1+ss2)xso1 (ss1+ss2)xso3 ss2xso2 ss3xso1 ss3xso2 ss4x(so2+so3) sw1x(so1+so2+so3) sw2x(so1+so2) sw3x(so1+so2+so3) t ss1xst1 ss5xst2 (sw1,sw2)xst1 so1xst1 hortal and lobo – synecological conservation framework 19 promise progress in assessing reductions in the rate of biodiversity loss (scholes and biggs 2005), information on genetic and population diversity is still grossly insufficient for conservation purposes, while that on the higher level (the species level) is fragmentary and biased (gewin 2002). biogeography-based conservation planning makes use of only those biodiversity features that can be mapped (brooks et al. 2004a). as mapping and measuring ecological processes are still in their infancy (cowling et al. 1999), most biogeography conservation efforts have targeted either (i) the observed or predicted pattern of species variation, or (ii) abiotic surrogates able to describe such a pattern. whilst the former makes use of either species distribution data (autoecology approach) or synecological variables (synecology approach) as targets, abiotic surrogacy has relied on environmental variation as a well-known surrogate of biodiversity variation (the environment-based approach), employing either continuous environmental variation data (see e.g. faith and walker 1996a; ferrier 2002; venevsky and venevskaia 2005) or land class data (land types sensu pressey 2004). however, most former systematic conservation planning applications have been academic rather than practical, being the results of area selection methods generally ignored by conservation planners (prendergast et al. 1999; cabeza and moilanen 2001; but see pressey and cowling 2001; cowling et al. 2003b; or airame et al. 2003). given the foregoing, conservation biogeography during the coming decades should strive to reconcile the need for a ‘great synthesis,’ to be presented as a unitary front to society and layman, with the possible scepticism that might arise from rapid science-based decision-taking in a field still in its infancy (whittaker et al. 2005). species or environment data? the long-latent disagreement over the superiority of either species or environment information, for both planning and for measuring the success of conservation efforts, has recently figured in debate (araújo et al. 2004a; brooks et al. 2004a,b; cowling et al. 2004; faith et al. 2004; higgins et al. 2004; molnar et al. 2004; pressey 2004), although its origin can be traced back to the approaches that preceded recent systematic conservation planning framework (see, e.g., dasmann 1972; mackinnon et al. 1986). those that favor the species-based approach claim that broad-scale biodiversity attributes such as land types or habitats are abstract and subjective human constructs; that environmental diversity does not necessarily reflect species diversity (araújo et al. 2001); and that species should be the main unit of biodiversity (wilson 2002). they argue that more taxonomic and distribution information should be gathered, that available information on species should be used (rodrigues et al. 2003), and that geographical data of species distribution should be improved through the use of environmental variables to generate coherent distribution hypothesis. they maintain that while environment data and land classifications are important, these should be used just to upgrade the reliability of species data, not as conservation targets themselves (brooks et al. 2004b). on the other hand, defenders of the environment-based approach assert that conservation planning must also represent such other biodiversity attributes as land types, habitats or ecological processes; that ecosystems within species-poor areas may be crucial regarding the goods and services needed for nature functioning (kareiva and marvier 2003); and that lack of reliability is the main impediment to the use of species data (pressey 2004). in addition, other authors suggest that the use of broad-scale biodiversity attributes can facilitate the study of the most poorlyunderstood species and assemblages, such as invertebrates (ward et al. 1999; mac nally et hortal and lobo – synecological conservation framework 20 al. 2002), and that strategy should concentrate not on pattern but on processes. more information should be gathered on the ecological role of species, and their relevance in nutrient cycles and energy flows (kareiva and marvier 2003) and population size should be taken into consideration in the design of conservation strategies. unfortunately for this argument, the relevance of species in ecological processes is no better understood than is species distribution. a reconciliation of these apparently opposing arguments would entail that both environment and species data are necessary components of conservation assessment (noss 1987; clark and slusher 2000; faith and ferrier 2002; ferrier 2002; ferrier et al. 2002b; cowling et al. 2004; higgins et al. 2004). conservation decisions still have to be made in spite of bias and gaps in species information. methods incorporating both types of data would establish promising conservation strategies (ward et al. 1999; cowling et al. 2003b; lombard et al. 2003), while using a complete and complementary set of biodiversity surrogates would guarantee a better selection (pressey 2004). therefore, available taxonomic and distribution information should be combined with environment, land type and vegetation data to take advantage of advances in predictive modelling. such a framework does not need to include the assumption that environment data by itself represents biodiversity variation (wessels et al. 1999; araújo et al. 2001; pharo and beattie 2001; cushman and mcgarigal 2002; mac nally et al. 2002; lombard et al. 2003; oliver et al. 2004; su et al. 2004), but rather that environmental characteristics, combined with available species distribution information, should improve such data. as both species and environment data are relevant, their respective contribution to conservation planning should depend on their quality and amount. however, in the case where both types of information are available, which should be considered more meaningful? unfortunately, the available information on the whole biodiversity “agents” (species) has often been ignored in the design of protection strategies (see, e.g., european nature 2000 network1). such a situation should be corrected, as the species are true conservation targets. as brooks et al. (2004b) argued, there is a risk that conservation policies may be based on rapidly-obtained remote-sensing data and computer models, and may not incorporate the species data that could ensure the success of conservation decisions. therefore, both species data and effective species-based methodologies should be put together on the table. could available species data account for biodiversity variations? if there is any basic unit of biogeography, it is the geographic range of species; the shapes of ranges and the dynamic changes in their boundaries reflect the interacting influences of limiting environmental conditions (niche variables), dispersal and extinction dynamics, and historic effects (brown et al. 1996; gaston 2003). thus, the present-day geographic information on species is able to reflect the environmental variation in nature and species data summarizes not only the entire spectrum of environmental conditions (perhaps unaccounted-for by the observer eyesight), but also the historic and demographic effects that produce dissimilar species assemblages in similar environments (hortal and lobo 2005). having said this, the lack of knowledge of the geographic distribution of organisms, their interactions, and their role in natural processes is the major obstacle to develop reliable strategies of biodiversity conservation. environment data, useful in the absence of detailed species data, should be used in conjunction with any newly-obtained species information, and the recording of taxonomic and distribution information should be 1 http://europa.eu.int/comm/environment/nature/home.htm. hortal and lobo – synecological conservation framework 21 encouraged. rather than considering only a limited number of organisms, or a limited, exclusive, environment database, as many species as possible, acceptably distributed throughout the evolutionary tree, should be taken into account. the exclusive use of distribution data is not advocated herein, but rather a framework based on biodiversity variables as conservation targets, that might reduce the need for environment surrogates, a complement to biodiversity distribution data. should the latter be lacking, the former would become a target. conservation priority selection based on information available on known taxa would at least guarantee their protection as conservation targets in their own right (brooks et al. 2004b). as detailed distribution and abundance data for most species in many regions is lacking (andelman and fagan 2000), former conservation approaches targeted just one or a few species (flagship, endangered, umbrella, and/or indicator species), assuming that regional biodiversity will be preserved by protecting these species and their habitats. however, this approach does not guarantee the protection of sites that encompass all regional biodiversity (see simberloff 1998; andelman and fagan 2000; williams et al. 2000; possingham et al. 2002; saetersdal et al. 2005), and can result in taking incomplete and/or erroneous decisions, so is currently abandoned. to avoid oversimplification of describing biodiversity through the distribution of a reduced number of species, as much species’ distribution information as possible should be compiled. the more detailed the information (both on spatial location and habitat description), the more useful for the monitoring and conservation of regional biodiversity (austin 1998). many initiatives are now devoted to gather extensive distribution data of organisms (among other information; see edwards et al. 2000; graham et al. 2004), making large amounts of information on a growing number of species and higher taxa available. the compilation of all the information on species now disseminated in the literature and natural history collections will make an enormous source of information available. although the analysis of this data will surely involve a number of problems (see, e.g., dennis et al. 1999; dennis and thomas 2000; dennis and shreeve 2003; gu and swihart 2004; molnar et al. 2004; higgins et al. 2004; cowling et al. 2004; pressey 2004; seoane et al. 2005), it is the only available basis to describe the spatial distribution of species (brooks et al. 2004b). interestingly, gaston and rodrigues (2003) show that even data gathered with relatively poor sampling effort can be highly effective for species representation. however, as higgins et al. (2004) point out, working in data-poor areas leads to the conclusion that such information is not sufficiently representative of biodiversity patterns. there are two shortfalls associated with the use of biological data for conservation biogeography (see whittaker et al. 2005). first, our knowledge of global biodiversity, the conservation target, is fragmentary and taxonomically and geographically biased (the linnean shortfall; brown and lomolino 1998). second, adequate distribution data for many of the known species and higher taxa is lacking (the wallacean shortfall; lomolino 2004), and prone to taxonomic and geographic bias. these shortfalls could partially be compensated by selecting as many species as possible from those well-distributed in the tree of life, and by forecasting species distributions with methodologies capable of coping with incomplete data. then, forecasted distributions could be used either for complementarity analyses (the autoecological approach) or to calculate biodiversity indices (a mixed autoecological-synecological approach) (see below). hortal and lobo – synecological conservation framework 22 obtaining reliable data for regional conservation assessment one of systematic conservation planning’s main objectives is to identify a set of representative areas (dasmann 1972; mackey et al. 1988; belbin 1993; church et al. 1996; vane-wright 1996; powell et al. 2000), in which all species may persist if included in a reserve network (araújo and williams 2000; araújo et al. 2004b). this is a key objective of several international conservation policy schemes, such as iucn (mackinnon and mackinnon 1986a,b, mackinnon et al. 1986) or wwf (dinerstein et al. 1995; olson and dinerstein 1998) frameworks. to develop processes for the selection and improvement of protected areas, objectives need to be specified (margules and pressey 2000). along this line, much attention has been directed to algorithms based on complementarity, a measure of the biodiversity attributes found in newly-selected areas that would be added to those in a pre-existing network (vane-wright et al. 1991; faith 1994; see also hingston 1932). complementarity was initially applied since the late 1980’s, using presence-absence data to select areas for protection with one or more presences of as many species of a given group as possible (margules et al. 1988; pressey et al. 1993; williams and humphries 1994; vane-wright 1996). other approaches have explored the power of composite variables to represent assemblage composition variability (e.g., ferrier 2002; ferrier et al. 2002b; araújo et al. 2004a); or that of mixed datasets including environment variables, distribution information and predicted species distributions (e.g., lombard et al. 2003; cowling et al. 2003b; sarkar et al. 2005). the three of them are enhanced by distribution information in biological databases, as well as by the latest methodologies now available (church et al. 1996; williams 1999; cabeza and moilanen 2001; rodrigues and gaston 2002), up-to-date but even so still having difficulties with complex conservation targets and datasets. therefore, in addition to using environmental surrogate data (see above), reserve selection procedures have also been developed from: i) single species distribution data (an attempt to represent all species found in the region; e.g. dobson et al. 1997; van jaarsveld et al. 1998; howard et al. 1998; araújo 1999; araújo and williams 2000; andriamampianina et al. 2000; polasky et al. 2000; martín-piera 2001; raxworthy et al. 2003); ii) composite synecological attributes of biodiversity (aimed at protecting sites of great richness, rarity or endemism, or representing as much variability in community composition as possible), an approach less-used for conservation purposes (but see margules et al. 1987; bojórquez-tapia et al. 1996; iverson and prasad 1998a,b; zimmermann and kienast 1999; ferrier et al. 2002b; gladstone 2002 or araújo et al. 2004a). the effectiveness of these selection algorithms depends on the accuracy of the input data. therefore, biases in the geographical and/or ecological space covered by distribution information obtained from nonsystematic sampling may in turn bias descriptions of species geographic ranges and lead to major errors in the distribution of endangered or target species (dennis 2001). this is probably a consequence of present atlas data on the spatial distribution of species and biodiversity measures being far from accurate (van jaarsveld et al. 1998; dennis et al. 1999; dennis and thomas 2000), one of the most important weaknesses in current planning techniques for conservation area selection. unavoidably, data obtained from sampling a given region constitutes just a group of samples, not a complete inventory (nicholls and margules 1993), and further sampling is needed to improve its quality, which usually hortal and lobo – synecological conservation framework 23 would involve high costs, both in time and money (prendergast et al. 1999). atlases developed from additional sampling may yield diversity measures that vary greatly on broader scales and extents, while old and newly observed scores may not correlate (e.g. european butterflies: dennis 1997; dennis and shreeve 2003). thus estimates of synecological variables from well-sampled areas (e.g., colwell and coddington 1994 for species richness, or chao et al. 2005 for compositional differences) would be of pragmatic value (see discussion below). modelling techniques can reduce the costs mentioned above by the extrapolation of biodiversity data to unexplored or poorlysampled areas, through the use of available environmental information with several methodologies of bioclimatic and geostatistic modelling (nicholls 1989; austin 1998; ferrier 2002; ferrier et al. 2002a,b; lehmann et al. 2002a). for example, as atlas data has improved, many predictions developed from old atlases agree with the data presented in new ones (dennis and shreeve 2003). biodiversity is now modelled mainly from: i) single species distribution data, one-by-one (autoecology, see guisan and zimmermann 2000; scott et al. 2002; ferrier et al. 2002a; pearson and dawson 2003; peterson et al. 2004; soberón and peterson 2005; araújo and guisan 2006); or ii) composite biodiversity variables, such as species richness, assemblage composition, endemism, rarity, and others, assumed to be biodiversity surrogates for monophyletic groups (synecology; see lobo and martín-piera 2002; hortal et al. 2001, 2003, 2004; ferrier 2002; ferrier et al. 2002b; ferrier and guisan 2006). if spatial distribution ranges of a number of species are modelled one-by-one, the complementarity criteria can be used to select nature reserves from such predictions. nowadays, single-species distribution modelling predictions are considered reliable indicators for area selection (see, e.g., williams et al. 2000; andriamampianina et al. 2000; araújo and williams 2000; peterson et al. 2000; polasky et al. 2000; araújo et al. 2002a; lehmann et al. 2002b; cabeza et al. 2004). synecology, on the other hand, used infrequently for area selection (but see, e.g., araújo et al. 2004a), has commonly been used just to identify biodiversity hotspots. however, several of these synecological variables (especially those describing community composition), may be of great utility for conservation (ferrier 2002; ferrier et al. 2002b). whilst autoecology has received more attention and support, modelling synecological variables can be advantageous for area selection. the advantages of the synecological approach in practical terms, compensating for data bias and model unreliability could be easier with the use of synecological variables because: i) the geographic clustering of prediction errors derived from autoecology models is eliminated; ii) it allows using data on the distribution of rare species (which under normal conditions can not be modelled due to paucity of records and/or presence sites); iii) reliable estimates of synecological variables for a group or a few groups of species involve fewer operations than estimates for all species one-by-one, including rare ones, thus reducing the effort involved in putting together a valid conservation proposal. given the current state of ecological knowledge, there are also several theoretical advantages to the use of synecological variables: iv) they may reflect emergent assemblage properties not contained in the species level, but important to area selection, or even to the preservation of the characteristics of biodiversity phenomenon per se; v) whilst species distribution ranges are expected to shift in response to climate change, many biodiversity surrogates may be stable over time. of course details matter, but we should concentrate on why and where the hortal and lobo – synecological conservation framework 24 woods are, before worrying about individual trees (lawton 1999). practical advantages 1. errors in the predictions of species’ distributions as mentioned above, a few studies have explored the usefulness of predicted species distribution one-by-one for complementarity analysis. however, these interpolations present many problems (see gu and swihart 2004) that could accumulate in predictions for a large group of species. importantly, differing methodologies, extent, resolution, and/or kinds of distributions may produce varying estimates. the performance of distribution prediction methods varies according to the geographic distribution, the equilibrium or not-equilibrium with the environmental conditions, the environmental requirements and natural history of the studied species (guisan and theurillat 2000; guisan and zimmermann 2000; austin 2002; pearson and dawson 2003; segurado and araújo 2004; brotons et al. 2004; guisan and thuiller 2005; araújo et al. 2005a; seoane et al. 2005; soberón and peterson 2005). therefore, the choice of method should depend on the goals and kinds of distributions being modelled (segurado and araújo 2004; soberón and peterson 2005). whilst one group of predictive methods try to model present-day spatial probabilities of occurrence (environmental conditions occupied by the species), others estimate habitat suitability for each species (potentially suitable environmental conditions; see peterson et al. 1999). in single-species conservation planning, the latter are also useful for monitoring and re-introduction tasks (see hirzel 2001; hirzel et al. 2001, 2002, 2004; chefaoui et al. 2005; cassinello et al. in press). lehmann et al. (2002a) favored the use of niche-based models for conservation purposes, but theoretical models may not work in real scenarios. for example, the range size of one third of the species used by peterson et al. (2000) was overestimated by models based on realized niche. that is to say, models developed for single species failing to determine species absence from some territorial units with adequate environments. although theory and practice of predictive modelling is being continuously updated, drawbacks in data and in the theoretical assumptions under modelling techniques and predictors might cause the spatial accumulation of errors. the continuous development of tools and protocols is yet improving the power (and success) of modelling techniques (e.g., scott et al. 2002; anderson et al. 2003; elith et al. 2006), although no single modelling strategy outperforms the rest in all occasions. nevertheless, the ecological knowledge on the modelled species and the statistical skills of modellers might be more important than the method used for the correct prediction of species distributions (see austin et al. 2006). having said this, it is also known that errors in the predictor variables and/or, biases in observed species data and model inabilities to account for true distributions can produce a systematic bias (flather et al. 1997; fielding 2002). for example, most times data on the distribution of a species is biased to the center (either geographic or environmental) of its distribution (see martínez-meyer 2005). in addition, in some regions some parts of the geographic and environmental spaces have been repeatedly sampled through time, whereas others remained poorly known, being regional conditions underrepresented in the data at hand (cabrero-sañudo and lobo 2003; j. hortal, a. jiménez-valverde, j. f. gómez, j. m. lobo and a. baselga, unpublished results; see discussion in hortal and lobo 2005). recent work on the theoretical assumptions under environmental niche modelling (e.g., soberón and peterson 2005; araújo and guisan 2006; austin 2006) and the hortal and lobo – synecological conservation framework 25 implementation of such theoretical advances in modelling techniques will result in an improvement of the accuracy of predictions in the future. however, the systematic character of the biases in biological data can result in a spatial aggregation of model errors (see thuiller et al. 2004a,b and discussion in araújo et al. 2005a,b and soberón and peterson 2005). important evidence on such error accumulation was found by araújo et al. (2005a); species richness patterns obtained from the sum of all individual model predictions for british birds correlated only slightly with observed patterns. according to araújo and collaborators (2005a), aggregating predicted distributions for a number of species produces a propagation of false positives and negatives (model errors), which tend to concentrate at the edges of the observed distributions, as well as in some parts of the environmental spectrum. thuiller et al. (2004a) found a similar pattern when modelling three tree species from environmental data restricted to their distribution ranges. we hypothesize that the areas where errors could potentially accumulate would coincide with: i) range margins, where species usually suffer more stress (brown 1984; see loiselle et al. 2003; araújo et al. 2005a); or ii) areas where community-level and/or historic processes, hard to model even for a single species, have played an important role in modifying species distributions (woodward and beerling 1997; davis et al. 1998a,b; pearson and dawson 2003; see also hampe 2004; iverson et al. 2004; skov and svenning 2004; thomas et al. 2004; soberón and peterson 2005 and martínez-meyer 2005). therefore, most times species distributions are not in equilibrium with current climate conditions (see, e.g., araújo and pearson 2005), one of the basic assumptions of ecological niche modelling. prediction errors may come from either the model procedure or predictors used (gu and swihart 2004), or biased or fragmentary biological data (dennis et al. 1999; dennis and thomas 2000), apart from inadequacies in their theoretical assumptions. atlas data usually does not include real absences (see araújo and williams 2000; brotons et al. 2004; soberón and peterson 2005), where a given species has not been found despite exhaustive sampling. many prediction methods require absence information, so their use implies the assumed absence of a species from a given set of sites. although some promising alternatives exist which tries to identify absence points (engler et al. 2004; iverson et al. 2004; lobo et al. 2006), model parameter estimates in such cases are generally affected by the bias in such ‘added’ data (false positives), even if sites are selected at random to avoid spatial or environmental bias (see gu and swihart 2004). other procedures, such as enfa (ecological niche factor analysis, see hirzel et al. 2001), develop models of predicted species distribution using only presence data. although this approach gives a picture on the suitable area occupied by the species, it presents the drawback that the location of suitable conditions is not the only factor influencing species distribution. information on ecological and contingent constraints contained in, at least, a part of the absences, could be even more important for predicting species distribution range (see lobo et al. 2006). ignoring such information while modelling habitat suitability should probably produce less accurate predictions. model prediction errors are thus greater at environment and spatial range margins, and in the areas where historic processes have modified purely ecological distribution patterns, while reliable biological data is commonly biased towards well-sampled areas and common species. due to such biases, it is likely that model errors should not be located randomly in space, but should form a pattern, as they are the result of phenomena with a defined spatial distribution that effect the hortal and lobo – synecological conservation framework 26 distribution of each species differently. a recent work from fortin et al. (2005) needs a complex set of methods to identify reliable species range margins for a single species, an unaffordable task to be applied to a great number of species in conservation studies. thus, we hypothesize that complementarity analysis of a great number of predicted distributions together should involve the assumption that such errors will probably accumulate in space (see flather et al. 1997 and fielding 2002), having been summed repeatedly, and will give rise to major gaps in the biodiversity predicted for the selected areas. further work would be necessary to confirm such a hypothesis outside the bounds marked by thuiller et al. (2004a) and araújo et al. (2005a and b). 2. rare or insufficient-data species conservation planning based on species distribution predictions suffers from an inability to predict the distributions of many rare species, which may be underrepresented by complementarity selections when raw presence data from incomplete sampling effort is used (see gaston and rodrigues 2003). data on these species is usually insufficient for the development of a good model (peterson et al. 2000; stockwell and peterson 2002; lehmann et al. 2002b; gu and swihart 2004; hortal et al. 2005). when rare species are excluded, both areas selected by complementarity and rarity hotspots correlate poorly with those selected using information for all species (araújo et al. 2005a). thus, species-by-species predictions can exclude those rare species (lobo and hortal 2003), most of them critically endangered (peterson et al. 2000), which, in general, are responsible for a great proportion of total diversity (gaston 1994). for example, in the iberian peninsula, of the 21 insect species protected by the european habitats directive, eleven have been recorded in fewer than ten 10 x 10 km utm grid squares, and five species in 4 or fewer squares (galante and verdú 2000). to get around this problem, area selection should be based on both the predicted distribution of species that can be modelled and observed distributions for those species that can not (see, e.g., lombard et al. 2003; pressey 2004). however, this approach may underestimate both the distribution of geographically rare species, and of demographically rare (i. e., with small local abundance) but widely-distributed species (see loiselle et al. 2003), which may be more difficult to record throughout all their distribution range (but see gaston and rodrigues 2003). rare species, of great interest for conservation purposes, are more sensitive to human disturbance, store a multitude of potentially useful adaptations, and also are responsible for the replacement of inventories among assemblages (β-diversity). assemblage distribution limits within a given territory (i. e., boundaries between presence and absence of representative species) are not usually well defined, but rather diffuse, with the number of species belonging to an assemblage changing gradually. to protect all biodiversity in a given territory, account must be taken of areas where this community replacement occurs (ferrier 2002; spector 2002), where many rare species could persist while populations of their competitors shrink, and where species belonging to the various assemblages in the region can co-occur. the exclusion of rare species data from the analyses may reduce the chances of identifying such replacement areas, especially in cases where, due to other factors, species replacement has not led to high values of species richness, regardless of the rarity of species present. 3. estimates and predictions of synecological variables predictive maps of synecological variables can be obtained in three ways (ferrier and guisan 2006): (i) assemble first, predict later (i.e., aggregate species data into biodiversity hortal and lobo – synecological conservation framework 27 variables first, and then model these variables); (ii) predict first, assemble later (i.e., predict the distributions of species one-by-one and then aggregate these predictions into biodiversity variables); and (iii) assemble and predict together (i.e., use a joint procedure to establish relationships between species and predict their distribution using not only their environmental requirements, but also their patterns of co-occurrence). although the third strategy presents the attractive characteristic of potentially taking into account biological interactions, we believe that the systematic bias in model errors and the lack of representation of rare species discussed above (sections 3.1.1 and 3.1.2) would result in a misrepresentation of these interactions, and therefore in bad-quality predictions. in the same way, we also believe that the ‘first predicting and then assembling’ strategy would result in erroneous pictures of the distribution of biodiversity due to the effect of these two topics, underperforming the results obtained by assembling first, and then predicting. on the contrary, we hypothesize that both the aggregation of errors and the lack of representation of rare species can be overcome by using the first strategy (assemble first, predict later) if two steps are previously added to the modelling protocol: a sampling effort assessment to identify the well-sampledenough areas (and discard those with unreliable inventories), and the extrapolation (when possible) of the scores of the synecological variables to diminish the effects of incomplete inventories (i.e., checklists are unlikely to be complete even at well-sampled areas) in these areas (see hortal et al. 2004). modelling synecological variables, such as species richness, endemism, rarity, composition, and so on reduces the number of possible prediction errors to one per variable (instead of one per species). as these variables are based on data from many species in each territory unit, sampling effort assessment is easier. species accumulation curves (soberón and llorente 1993; colwell and coddington 1994) can help to identify well-sampled areas for a given group of species (lobo and martínpiera 2002; hortal et al. 2001, 2004; martínpiera and lobo 2003; jiménez-valverde and hortal 2003). they can also yield estimates of richness scores for each territory unit, from information restricted to that unit (in addition to the many other species richness estimators available; see, e.g., soberón and llorente 1993; colwell and coddington 1994; chazdon et al. 1998; gotelli and colwell 2001; chiarucci et al. 2003; brose et al. 2003). the results from several of these estimators could be comparable even when the biological information comes from heterogeneous sources (hortal et al. 2006). it should be pointed out that species richness alone is not the unique conservation target, but has to be integrated into a framework including other targets, such as rarity, endemism and, more importantly, complementarity. recently, the development of non-parametric estimators for jaccard and sørensen indices (chao et al. 2005), allows to avoid the erroneous estimates of compositional similarity between areas obtained from incomplete surveys, providing reliable complementarity measures from incomplete data. errors in scores from present inventories are thus reduced by using estimates for each synecological variable, usually allowing the inclusion of a larger number of territorial units from a given area while enlarging the spatial and environmental coverage of observations. reliability of models, built from more complete and robust data, is improved too. interestingly, there is a reduction, or even elimination, of sampling bias in scores from models developed from asymptoticallyestimated richness values (see hortal et al. 2004). moreover, the set of species considered is enlarged by the inclusion of rare species in the estimates of biodiversity surrogate scores, which become a major part of final values. rarity itself can be also estimated for each hortal and lobo – synecological conservation framework 28 territorial unit (see gaston 1994), as can other surrogates, such as β-diversity (faith and ferrier 2002; ferrier 2002; ferrier et al. 2002b), endemism (lumaret and lobo 1996), assemblage composition (ferrier 2002; hortal et al. 2003; araújo et al. 2004a; ferrier and guisan 2006). appropriate sampling effort assessment along with surveys designed to cover all spatial variations of biodiversity (see hortal and lobo 2005) could even make data from more taxonomic groups available, with a minor investment of finance and time. here, it is important to notice that several groups and higher taxa could serve as surrogates of other groups, for either local species richness (balmford et al. 2000; cardoso et al. 2004a,b), and geographic variations (macnally and fleishman 2004; fleishman et al. 2005; thomson et al. 2005; tognelli 2005; tognelli et al. 2005; larsen and rahbek 2005). therefore, a limited group of taxa could give a good picture of overall biodiversity variation, although their surrogacy presents several limitations (see, e.g., thomson et al. 2005; tognelli et al. 2005). prediction errors in the estimates may also be reduced through the use of available geostatistic methodologies (carroll and pearson 2000; ter steege et al. 2003), which model variables using spatial location as predictor (alone or together with environment variables, e.g. tognelli and kelt 2004). homogeneous estimates, such as the asymptotic species richness described above, can reduce the extreme sensitivity of these methodologies to input data error. also, the use of spatial location as predictor may allow the effect of unknown or unaccounted-for effects and/or historic processes (see, e.g., mac nally et al. 2003) to be included in environment-based models, and so improve the power and accuracy of estimates (see lobo and martín-piera 2002; hortal et al. 2001; lobo et al. 2001, 2002). theoretical advantages the three practical advantages discussed above indicate that biodiversity can readily be described by synecological variables. evidently, a theoretical advantage of using these variables to describe the geographic variations of a group of organisms that share a common evolutionary past is that the patterns in these compound variables arise from the response of their common adaptations and divergences to conditions in the territory studied. however, the most important benefit to conservation may come from the agreement of synecology with current developments in theoretical ecology. current evidence for chaotic structure in ecological systems suggests that the interrelationships among their components (species and individuals) may play a key role in ecosystem functioning. the complex interactions among these components would be the cause of the properties and dynamics of such systems (see, e.g., brown et al. 2002; bolliger et al. 2003; or gorshkov et al. 2004). therefore, these emergent properties pertain neither to individuals nor species, but to the whole of ecosystem biodiversity. we are going to focus on two different topics. recent theoretical advances indicate that biodiversity as a phenomenon may affect ecosystem functions such as productivity and resilience (e.g., species richness and composition; see tilman et al. 1996; yachi and loreau 1999; tilman 2000, 2001; kennedy et al. 2002; bond and chase 2002; loreau et al. 2003). moreover, in some places assemblages of several groups were stable through time (see brown et al. 2001; or fleishman and mac nally 2003); even through the dramatic holocene climate changes (see rodríguez 2004). in this context, ecosystem resilience has been related to diversity (tilman et al. 2006), providing a link between diversity, enhanced ecosystem functioning and community resistance (and thus stability) against environmental changes. other biodiversity aspects, such as assemblage hortal and lobo – synecological conservation framework 29 replacement, may remain invariant due to the particular geomorphology of a region (see spector 2002), or to the resilience of the system. 1. detection of evolutionarily important areas species populations go under ecophysiological stress (see hengeveld 1990) and are more threatened (araújo and williams 2000) in range margin areas. there, environmental conditions are more rigorous than in core areas, and populations become isolated easily at range margins. therefore, evolutionary processes could be enhanced due to ecophysiological stress (i.e., species adaptations are often pushed to the limit; brown 1984; brussard 1984; case et al. 2005) and the genetic differentiation caused by isolation. in addition, it is known that environmental stress increases retrotransposon activity (see sentís 2002), favoring symbiotic interactions (e.g., russell et al. 2003), and promoting lateral gene transmission (see nieto et al. 2004). in addition, many species may survive human disturbance or climate change in a given region (threatened or at range-margins) by adapting new life-history strategies, such as the modification of their micro-distributions and/or near environment (e.g. insects; danks 2002). therefore, it can be expected that some evolutionary processes may take place in these peripheral areas (see, e.g., thomas et al. 2001). given that evolutionary processes occur throughout space as well as over time, if conservation policies miss the areas where changes, re-sampling and additions to the evolutionary pool of the biota they intend to preserve are occurring, they might not conserve information and processes vital to nature’s resilience. policies based on single species distribution data may so fail to represent such phenomena, but those based on synecological variables may take them into account. their fewer dimensions summarize a greater proportion of the total variability attributable to all the species in a given territory, and should better represent the various facets of regional biodiversity. in this context, the environmental or geographic conditions limiting species distributions in their range margins often result in zones of ecological transition, where species replacement (β-diversity) and richness scores are high due to the joint appearance of different assemblages. therefore, we hypothesize that some areas of evolutionary importance could be identified by some combinations of these synecological variables. however, the importance of range margins for conservation is still unknown, as the probabilities of species persistence are lower in peripheral than in core areas (see araújo et al. 2002b). 2. community processes, emergent properties and stability along environmental changes species distribution ranges have shifted throughout the quaternary (see, e.g., hewitt 1996, 1999; bradshaw 1999) and the last century (lewis and bryant 2002; walther et al. 2002). such shifts are expected to continue due to human activity and climate change (see midgley et al. 2003, sparks et al. 2005). many of these shifts are being identified nowadays in a number of groups such as butterflies (parmesan et al. 1999; warren et al. 2001; hill et al. 2002), dung beetles (lobo 2001), birds (thomas and lennon 1999; walther et al. 2002) or dragonflies (hickling et al. 2005). such process gets even worse due the effects of land degradation by human impact (see pyke and fischer 2005). if these range shifts could be predicted with reliability, conservation policies could include both present and future distributions of biodiversity. the reliability of the extrapolation of future species distributions by applying the predictive models based on current environmental conditions to future climate scenarios is currently a matter of debate. although many recent works include the hortal and lobo – synecological conservation framework 30 implications of climate change on conservation planning (e.g. peterson et al. 2002; hannah et al. 2002; araújo et al. 2004b; pyke and fischer 2005), it has been repeatedly argued that the current state of knowledge of species range shifts can not yet be used to produce reliable extrapolations of future distributions (see pearson and dawson 2003, 2004; hampe 2004; thuiller et al. 2004a; but see araújo et al. 2004b, 2005b; pyke and fischer 2005). certainly, there is evidence on the coincidence of the environmental responses of some species in their recent and past distributions due to niche conservatism (martínez-meyer et al. 2004; a.t.peterson and e.martínez-meyer, unpublished results). such evidence indicates that it would be possible to extrapolate current environmental responses to future climate scenarios. however, we do not know how general is niche conservatism, and evidence is also available on increased evolutionary rates under changing conditions (see section 3.2.1). the accuracy of using current speciesenvironment relationships to estimate past climate conditions in palaeontology is doubtful, as the factors affecting past and present species distributions may have varied (rodríguez 1999; rodríguez and nieto 2003). by extension, present distribution response to present-day factors may differ from future distribution response to similar factors; thus, extrapolating future species distributions from their current relationship with climate might be problematic. in both cases, the role of historic (past and future) processes throughout time is unknown. therefore, modelling based on inherently co-linear environmental variables can not identify causal predictor relationships with species distribution, so extrapolation to other climate scenarios may not be meaningful (hampe 2004). taking this into account, conservation networks based on current species distributions may not be effective (margules et al. 1994; prendergast et al. 1999), and may need to be updated in a few decades. however, nature reserves are not yet selected and managed with an eye to the impact of climate change on species abundance and distribution (lawton 1997). short-term focused conservation will fail to protect on the long-term both biodiversity and the processes it generates and supports. however, community stability in structure and richness over time has been reported for various regions, extents, time periods, and taxa, despite changes in composition (brown et al. 2001; sax 2002), human induced impacts, or strong climate changes, even during glacial periods (rodríguez 2004). current evidence indicates that some of the many processes modifying biodiversity spatial distribution may remain relatively constant regardless of environmental variations. the relationship between productivity and richness has been extensively documented; the more productive the site, the larger the species richness (tilman and pacala 1993; tilman et al. 1996, 1997a,b; srivasta and lawton 1998; tilman 1999, 2000; lehman and tilman 2000). such high diversity acts as a reservoir of evolutionary solutions, improving the efficiency of ecosystem functioning (and, thus, productivity) in the presence of environmental changes (insurance hypothesis; see yachi and loreau 1999; loreau et al. 2003); systems with greater richness (or high species replacement, see below) may maintain productivity in spite of environmental change through the replacement of individuals of species in suboptimal conditions by those in optimum conditions (but see emmerson et al. 2005). high productivity (i.e., enhanced ecosystem functioning) provides resilience to environmental changes (tilman et al. 2006), and is also a key factor in maintaining the structure of ecological systems (brown et al. 2001). in addition, species-rich local communities are able to resist the establishment and effects of invasive species (levine and d’antonio 1999; burger et al. 2001; kennedy et al. 2002), favoring temporal hortal and lobo – synecological conservation framework 31 stability. nevertheless, dependence on initial conditions, characteristic of complex systems, may also lead to long-term structural stability in an ecosystem, in the absence of glaciations or other major disturbances to the processes occurring in a given area (e.g., quaternary mammal communities; rodríguez 2006). after each major perturbation, a new assembly, and development of assemblages, may lead to structurally different systems, which may, in turn, be stable despite minor perturbations. such inertia may be a synecological property of many systems or assemblages. shape, relief and location of a given site or region also affect the spatial distribution of biodiversity (burnett et al. 1998; nichols et al. 1998; jetz and rahbek 2001). environmental heterogeneity is in part the result of geomorphology and bedrock geology, which remain relatively stable during long periods of time. as commented above, the great variety of habitats present at environmentally heterogeneous areas promote higher species richness and replacement. for similar causes, some of these areas may become corridors for population range shifts of many different species under climate change conditions due to their particular geomorphology (e.g., latitudinal mountain chains, see lobo and halffter 2000; or spector 2002). these ‘biogeographic crossroads’ may also shelter high α and β-diversities, even while species composition change and populations fluctuate. therefore, species richness and replacement may remain constant throughout time in some of the areas important for the migration of most species during their range shifts. in the same way, some factors that favor endemism, such as isolation, may also stay constant (e.g., islands; see whittaker 1998; borges and brown 1999; emerson 2002; gillespie and roderick 2002). in sum, several biodiversity facets may remain constant throughout time, and could present resilience to human impact. identifying the synecological variables that are the output of these facets, and basing regional conservation assessments on them may increase the probability of conserving more biodiversity, and of protecting areas vital to evolutionary processes. synecological variables as a basis for regional systematic conservation planning as pimm and lawton (1998) stated, species may not be the correct units to use complementarity as conservation criterion. thus, other descriptors of geographic biodiversity variations must be found. as discussed above, the geographic range of species is neither invariant, nor easy to describe accurately. in most cases, observed single-species distribution is not reliable enough to be used in predictive modelling, and is even less so for complementarity-based selections. as we have shown, a pragmatic alternative could be the use of the synecological variables to describe the spatial distribution of biodiversity. while it might be argued that these variables fail to contain useful information contained in single-species distributions, it is widely known that an appropriate spatial scale reduces the ‘noise’ present in these variables due to differences in the distributions of a large number of species (see, e.g., brown et al. 2002; or willis and whittaker 2002). community ecology is a mess, with so much contingency that useful generalizations are hard to find. there are not as many kinds of population dynamics as species on earth, but a multitude of essentially trivial variations on a few common themes (lawton 1999). finding laws or patterns in nature’s chaotic organization, where species and environment dynamics and interactions combine to produce fractal complexity over time (e.g., brown et al. 2002; or bolliger et al. 2003), involves a proper choice of scale, organization level and extent examined (levin 1992; lawton 1999; hortal and lobo – synecological conservation framework 32 brown et al. 2002; willis and whittaker 2002). the biodiversity patterns described by synecological variables can be useful at virtually any grain size, as they provide information about processes at each level of resolution (blackburn and gaston 2002; willis and whittaker 2002). due to the fractal nature of the processes under these patterns, several “windows of order” appear along scale change. there, auto-organization increases, and noise is reduced, producing sharper patterns. one or several of these windows are the subject of study of biogeography and macroecology (see brown 1995; and brown and lomolino 1998). grids of 10to 100-km width (0.1 to 1 geographic degrees approximately), such as those used in most regional systematic conservation planning assessments can reasonably be expected to resolve macroecology and biogeography patterns (depending on group studied and purpose of analysis). predictions can be obtained for such composite variables as species richness (α-diversity), species turnover (β-diversity), community composition, rarity, endemism, etc. (ferrier 2002; lobo and hortal 2003), whose utility as descriptors of biodiversity is not open to discussion. one of these variables alone is unlikely to adequately describe the distribution and function of the whole of biodiversity, since it is unlikely that all the facets of biodiversity could be summarized by a single index (see gaston 1996). for example, both species number and composition have been related to ecosystem properties (see debate in tilman et al. 1996, 1997a,b; and wardle et al. 1997), and endemism may be related to the danger of extinction in an area (pimm et al. 1995), or to the rate of taxonomic replacement between two assemblages. however, maps of predicted values of a small set of these variables may display the main spatial patterns of biodiversity in a given area, patterns that may differ substantially from the ones derived from the aggregation of single-species information. even more, if these variables are the expression of several of the facets of biodiversity, it could be possible to find a reduced number of composite factors, namely ‘biodiversity factors’, that summarize all the different aspects covered by the related synecological variables that are today in use. however, their use as predictors would require first: i) design of sampling strategies to obtain reliable information on spatial biodiversity patterns using a minimum set of sites representative of environmental and spatial variability (austin 1998; dennis and hardy 1999; ferrier 2002; hortal and lobo 2005; see also araújo and guisan 2006); ii) inclusion of a measure of sampling effort or, at least, a surrogate, in the databases used to gather information for distribution atlases (austin 1998; dennis et al. 1999; dennis and thomas 2000); and iii) production of maps of as many predicted biodiversity components as possible, such as species richness, rarity, endemism or composition differences (carroll and pearson 1998a; pearson and carroll 1999; hortal et al. 2001, 2003, 2004; lobo and martín-piera 2002; ferrier 2002; ferrier et al. 2002b). traditionally, maps of biodiversity measures alone have not been used to depict all spatial variations in the assemblages of a given group, but only to identify hotspots and impoverished zones (see balmford 1998 or kitching 2000; and several examples at prendergast et al. 1993; gaston and david 1994; heikkinen and neuvonen 1997; araújo 1999; hortal et al. 2001, 2004; martín-piera 2001). this restricted use was partially amplified with the advent of complementarity, easier to apply directly to single-species distributions. tools to apply the complementarity criterion to synecology descriptions of biodiversity now exist. ferrier and collaborators (ferrier 2002; ferrier et al. 2002b) suggest using information on wellsurveyed sites to identify groups of sites with hortal and lobo – synecological conservation framework 33 similar assemblages, or groups of species occurring at similar sites, to subsequently model and extrapolate these multinomial variables. they also propose a promising procedure, the use of generalized linear matrix regression, to model composition dissimilarities between all pairs of survey sites as a function of the distances in one or several explanatory matrices (generalized dissimilarity modelling; gdm). the square matrix of the distances between well-surveyed localities can be used to build a triangular distance matrix that can be used in ordination procedures, to derive one or more continuous factors representing composition differences between localities. these scores can be modelled, as can any other continuous variable, and their values extrapolated to the unsampled territory (hortal et al. 2003). selecting localities representative of most of the variability represented in the matrix by modelling a number of synecological variables is a procedure similar to the one proposed by faith and walker (1996a; see also araújo et al. 2004a). the performance of synecology procedures should be compared with that obtained currently by modelling individual species to select areas for conservation (e.g. cabeza et al. 2004; sánchez-cordero et al. 2005). however, we believe that the synecology approach promises a theoretically robust use of biodiversity measures to decide where and how to locate protected areas. of course, the goodness-of-fit and utility of predicted synecology maps should increase with increasing scale (carroll and pearson 1998b; pearson and carroll 1999; murguía and villaseñor 2000), while the main patterns of variation may be obscured by small-scale processes at smaller spatial resolutions (see prendergast et al. 1999). nevertheless, using just biological data alone to locate protected areas is not practical, as the selected sites may not be adequate for reserve implementation, or the spatial resolution may be too small to provide areas able to host viable populations (pimm and lawton 1998). thus, regional conservation goals must be large-scale; then site networks should be designed in each important area on a smaller scale. in this step, downscaling techniques (see, e.g. araújo et al. 2005a) can be used to obtain a sharper picture of biodiversity patterns and the distribution of target species (see, e.g., barbosa et al. 2003). much national and continental atlas data can be referred to homogeneous territorial units, such as utm grid cells (e.g., 50 km width), or even broad geographic grid units, such as the units of 0.5º, which have approximately the same surface area except at extreme north and south latitudes. for most regions, even smaller grid sizes, such as 10x10or 20x20-km grids may be used, but before doing so it is necessary to develop a general, wideresolution basis for conservation planning. since no group can be used alone as an indicator for all biodiversity, as many groups as possible are needed to cover all the kinds of spatial dynamics, of the few that may exist (lawton 1999). as stated above, good atlas data is available only for a few groups, mainly plants and vertebrates. such atlases are very difficult to obtain for the groups of species, such as invertebrates, that represent the greater portion of overall biodiversity (see, e.g., wilson 2002). however, it is relatively easy to obtain good inventories of a number of these less well-known taxonomic groups from a limited number of localities or grid cells. if these sites are environmentally and spatially well-distributed (hortal and lobo 2005), such partial information can be used to obtain acceptable predictions of synecological biodiversity attributes. synecological variables, much more rapid to use, can produce more reliable estimates, and may also help to detect emergent properties of assemblages not contained in single-species distributions, nor thus even in cumulative single-species predictions. so a practical approach to the use of synecological variables for conservation assessment is hortal and lobo – synecological conservation framework 34 needed. during the 80s and 90s, the sound work of australian csiro and conservation agencies led to the development of regional, analytical and systematic territorial planning of biodiversity preservation. according to margules and pressey (2000), the six consecutive stages, including feedback from the last stages to improve the effectiveness of the previous ones, of this approach are: i) obtain regional biodiversity data, a stage consisting of compilation and evaluation of previously existing knowledge, sampling, and extrapolation (austin 1998); ii) identify regional conservation goals; iii) assess the effectiveness of current reserve network; iv) select additional areas for protection (see pressey and nicholls 1989; pressey et al. 1993, 1996; or faith and walker 1996a,b); v) implement conservation actions; and vi) maintain at appropriate levels in the protected areas the indicators that have been used as conservation goals. in the context of biological data scarceness described herein, we propose the use of the first four stages set out in margules and pressey 2000 (see also austin 1998 and ferrier 2002) with which to select areas effectively, from a synecological perspective, supported by currently available single-species data and environmental surrogacy to: i) obtain reliable data for as many groups as possible (see hortal and lobo 2005): (a) compile and analyze existing information to identify areas with reliable inventories; (b) design and run a survey to optimize data on biodiversity patterns; and (c) reexamine that data. ii) select the biodiversity attributes to be used, and estimate their scores in those territorial units with well-established information. gaps in distributional data should be filled in by selected environmental surrogates. iii) interpolate the scores of these biodiversity attributes to all regional territorial units by means of spatialenvironmental modelling. iv) analyze the effectiveness of current reserve network for biodiversity protection, and develop a proposal for new areas and structural elements (corridors and/or micro-reserves) to maximize this effectiveness, using both the complementarity criterion with species composition variables, and by selecting the areas with higher species richness, endemism, etc., according to the results of the research agenda proposed above. additional data on species of interest should play a central role in once the bulk of priority areas had been established using synecological variables, additional information on the distribution (either observed or predicted) of species of special interest (i.e., protected species or wellknown indicators) and on land types (descriptive of landscape composition) can be then added to area selection processes to improve the coverage of all conservation goals at stake, and to reserve design processes (part of margules and pressey v and vi stages) to be accounted for in the final spatial configuration of the reserve network. these four steps are relatively easy to carry out for a single group with limited staff, time, and funds (for a complete example, see hortal 2004). thus, effective reserve networks could be designed in a short period, with a small investment, and be integrated into biodiversity monitoring such as that expounded on by green et al. (2005). some a priori prioritization of synecological variables can be established from scratch, based on current theoretical knowledge on community ecology and assembly, and biodiversity and ecosystem functioning. for example, first use composition patterns to improve complementarity coverage, then try to cover functional diversity, and then select richest hortal and lobo – synecological conservation framework 35 areas within the selected to include more resilient/better functioning assemblages, then rarity and/or endemism, and so on (see some clues below). however, such decisions might rely on current prejudices/concepts on what to protect and how ecosystems are functioning, fields where some debate does exist (especially in the latter). therefore, a research agenda is needed to determine how these variables are able to identify and/or cover i) areas of higher resilience, ii) some stability in the patterns of species composition or complementarity, or iii) species, communities or landscapes identified as conservation goals. such research might necessarily involve the use of good quality real data taken at different periods of time to ensure that its conclusions could be taken as a baseline for conservation schemes. conclusions we have summarized a number of practical and theoretical advantages of using a synecological approach in systematic conservation planning. the conservation value of synecological variables is transparent and easy for policy makers to understand, as only a few variables are needed to design the core of a reserve selection. therefore, the advantages of the synecology approach make its use advisable for current reserve selection procedures. further studies would evaluate how this approach may be implemented most effectively (i.e., which biodiversity aspects should be prioritized and/or preferently used in complementarity approaches), to what extent the results obtained with selections based on other biodiversity surrogacy strategies are improved upon by the use of synecology, and, more importantly, to what extent propositions based on synecological variables provide resilient area networks able to face environmental changes through time. acknowledgements the authors wish to thank town peterson, richard cowling and several anonymous referees for their comments and suggestions on former versions of this work, and to james cerne for his exhaustive english review. this work was supported by the spanish mec project cgl2004-0439/bos and the fundación bbva project “yamana diseño de una red de reservas para la protección de la biodiversidad en américa del sur austral utilizando modelos predictivos de distribución con taxones hiperdiversos”. jh was also supported by a portuguese fct postdoctoral grant (bpd/20809/2004). literature cited airame, s., j. e. dugan, k. d. lafferty, h. leslie, d. a. mcardle, and r. r. warner. 2003. applying ecological criteria to marine reserve design: a case study from the california channel islands. ecological applications 13:s170-s184. andelman, s. j., and w. f. fagan. 2000. umbrellas and flagships: efficient conservation surrogates or expensive mistakes? proceedings of the national academy of sciences of the usa 97:5954-5959. anderson, r. p., d. lew, and a. t. peterson. 2003. evaluating predictive models of species' distributions: criteria for selecting optimal models. ecological modelling 162:211-232. andriamampianina, l., c. kremen, d. vane-wright, d. lees, and v. razafimahatratra. 2000. taxic richness patterns and conservation evaluation of madagascan tiger beetles (coleoptera: cicindelidae). journal of insect conservation 4:109-128. araújo, m. b. 1999. distribution patterns of biodiversity and the design of a representative reserve network in portugal. diversity and distributions 5:151-163. araújo, m. b., and p. h. williams. 2000. selecting areas for species persistence using occurrence data. biological conservation 96:331-345. araújo, m. b., and r. g. pearson. 2005. equilibrium of species' distributions with climate. ecography 28:693-695. araújo, m. b., and a. guisan. 2006. five (or so) challenges for species distribution modelling. journal of biogeography 33:doi:10.1111/j.13652699.2006.01584.x. araújo, m. b., c. j. humphries, p. j. densham, r. lampinen, w. j. m. hagemeijer, a. j. mitchellhortal and lobo – synecological conservation framework 36 jones, and j. p. gasc. 2001. would environmental diversity be a good surrogate for species diversity? ecography 24:103-110. araújo, m. b., p. h. williams, and a. turner. 2002a. a sequential approach to minimise threats within selected conservation areas. biodiversity and conservation 11:1011-1024. araújo, m. b., p. h. williams, and r. fuller. 2002b. dynamics of extinction and the selection of nature reserves. proceedings of the royal society of london b 269:1971-1980. araújo, m. b., p. j. densham, and p. h. williams. 2004a. representing species in reserves from patterns of assemblage diversity. journal of biogeography 31:1037-1050. araújo, m. b., m. cabeza, w. thuiller, l. hannah, and p. h. williams. 2004b. would climate change drive species out of reserves? an assessment of existing reserve-selection methods. global change biology 10:1618-1626. araújo, m. b., w. thuiller, p. h. williams, and i. reginster. 2005a. downscaling european species atlas distributions to a finer resolution: implications for conservation planning. global ecology and biogeography 14:17-30. araújo, m. b., r. j. whittaker, r. j. ladle, and m. erhard. 2005b. reducing uncertainty in projections of extinction risk from climate change. global ecology and biogeography 14:529-538. austin, m. p. 1998. an ecological perspective on biodiversity investigations: examples from australian eucalypt forests. annals of the missouri botanical garden 85:2-17. austin, m. p. 2002. spatial prediction of species distribution: an interface between ecological theory and statistical modelling. ecological modelling 157:101-118. austin, m. 2006. species distribution models and ecological theory: a critical assessment and some possible new approaches. ecological modelling in press:doi:10.1016/j.ecolmodel.2006.1007.1005. austin, m. p., l. belbin, j. a. meyers, m. d. doherty, and m. luoto. 2006. evaluation of statistical models used for predicting plant species distributions: role of artificial data and theory. ecological modelling in press:doi:10.1016/j.ecolmodel.2006.1005.1023. balmford, a. 1998. on hotspots and the use of indicators for reserve selection. trends in ecology and evolution 13:409. balmford, a., a. j. e. lyon, and r. m. lang. 2000. testing the higher-taxon approach to conservation planning in a megadiverse group: the macrofungi. biological conservation 93:209-217. barbosa, a. m., r. real, j. olivero, and j. m. vargas. 2003. otter (lutra lutra) distribution modelling at two resolution scales suited to conservation planning in the iberian peninsula. biological conservation 114:377–387. belbin, l. 1993. environmental representativeness: regional partitioning and reserve selection. biological conservation 66:223-230. blackburn, t. m., and k. j. gaston. 2002. scale in macroecology. global ecology and biogeography 11:185-189. bolliger, j., j. c. sprott, and d. j. mladenoff. 2003. self-organization and complexity in historical landscape patterns. oikos 100:541-553. bond, e. m., and j. m. chase. 2002. biodiversity and ecosystem functioning at local and regional spatial scales. ecology letters 5:467-470. bojórquez-tapia, l. a., i. azuara, e. escurra, and o. flores-villela. 1996. identifying conservation priorities in méxico through geographic information systems and modelling. ecological applications 5:215-231. borges, p. a. v., and v. k. brown. 1999. effect of island geological age on the arthropod species richness of azorean pastures. biological journal of the linnean society 66:373–410. bradshaw, r. h. w. 1999. spatial responses of animals due to climate change during the quaternary. ecological bulletins 47:16-21. brooks, t., g. a. b. da fonseca, and a. s. l. rodrigues. 2004a. protected areas and species. conservation biology 18:616-618. brooks, t., g. a. b. da fonseca, and a. s. l. rodrigues. 2004b. species, data, and conservation planning. conservation biology 18:1682-1688. brose, u., n. d. martinez, and r. j. williams. 2003. estimating species richness: sensitivity to sample coverage and insensitivity to spatial patterns. ecology 84:2364-2377. brotons, l., w. thuiller, m. b. araújo, and a. h. hirzel. 2004. presence–absence versus presenceonly based habitat suitability models for bird atlas data: the role of species ecology and prevalence. ecography 27:285–298. brown, j. h. 1984. on the relationship between abundance and distribution of species. american naturalist 124:255-279. brown, j. h. 1995. macroecology. university of chicago press, chicago. brown, j. h., and m. v. lomolino 1998. biogeography. second edition. sinauer associates, inc., sunderland, massachussets. brown, j. h., g. c. stevens, and d. m. kaufman. 1996. the geographic range: size, shape, boundaries, and internal structure. annual review of ecology and systematics 27:597-623. brown, j. h., s. k. m. ernest, j. m. parody, and j. p. haskell. 2001. regulation of diversity: maintenance hortal and lobo – synecological conservation framework 37 of species richness in changing environments. oecologia 126:321-332. brown, j. h., v. k. gupta, b.-l. li, b. t. milne, c. restrepo, and g. b. west. 2002. the fractal nature of nature: power laws, ecological complexity and biodiversity. philosophical transactions of the royal society of london b 357:619-626. brussard, p. f. 1984. geographic patterns and environmental gradients: the central-marginal model in drosophila revisited. annual review of ecology and systematics 15:25-64. burger, j. c., m. a. patten, t. r. prentice, and r. a. redak. 2001. evidence for spider community resilience to invasion by non-native spiders. biological conservation 98:241-249. burnett, m. r., p. v. august, j. h. brown jr., and k. t. killingbeck. 1998. the influence of geomorphological heterogeneity on biodiversity. i. a patch-scale perspective. conservation biology 12:363-370. cabeza, m., and a. moilanen. 2001. design of reserve networks and the persistence of biodiversity. trends in ecology and evolution 16:242-248. cabeza, m., m. b. araújo, r. j. wilson, c. d. thomas, m. j. r. cowley, and a. moilanen. 2004. combining probabilities of occurrence with spatial reserve design. journal of applied ecology 41:252262. cabrero-sañudo, f. j., and j. m. lobo. 2003. estimating the number of species not yet described and their charactersitics: the case of western palaearctic dung beetle species (coleoptera, scarabaeoidea). biodiversity and conservation 12:147-166. cardoso, p., i. silva, n. g. de oliveira, and a. r. m. serrano. 2004a. higher taxa surrogates of spider (araneae) diversity and their efficiency in conservation. biological conservation 117:453459. cardoso, p., i. silva, n. g. de oliveira, and a. r. m. serrano. 2004b. indicator taxa of spider (araneae) diversity and their efficiency in conservation. biological conservation 120:517-524. carroll, s. s., and d. l. pearson. 1998a. spatial modeling of butterfly species diversity using tiger beetles as a bioindicator taxon. ecological applications 8:531-543. carroll, s. s., and d. l. pearson. 1998b. the effects of scale and sample size on the accuracy of spatial predictions of tiger beetle (cicindelidae) species richness. ecography 21:401-414. carroll, s. s., and d. l. pearson. 2000. detecting and modeling spatial and temporal dependence in conservation biology. conservation biology 14:1893-1897. case, t. j., r. d. holt, m. a. mcpeek, and t. h. keitt. 2005. the community context of species' borders: ecological and evolutionary perspectives. oikos 108:28-46. cassinello, j., p. acevedo, and j. hortal. in press. prospects for population expansion of the exotic aoudad (ammotragus lervia; bovidae) in the iberian peninsula: clues from habitat suitability modelling. diversity and distributions. chao, a., r. l. chazdon, r. k. colwell, and t.-j. shen. 2005. a new statistical approach for assessing similarity of species composition with incidence and abundance data. ecology letters 8:148-159. chazdon, r. l., r. k. colwell, j. s. denslow, and m. r. guariguata. 1998. statistical methods for estimating species richness of woody regeneration in primary and secondary rain forests of ne costa rica. pp. 285-309 in f. dallmeier, and j. a. comiskey, eds. forest biodiversity research, monitoring and modeling: conceptual background and old world case studies. parthenon publishing, paris. chefaoui, r. m., j. hortal, and j. m. lobo. 2005. potential distribution modelling, niche characterization and conservation status assessment using gis tools: a case study of iberian copris species. biological conservation 122:327-338. chiarucci, a., n. j. enright, g. l. w. perry, b. p. miller, and b. b. lamont. 2003. performance of nonparametric species richness estimators in a high diversity plant community. diversity and distributions 9:283-295. church, r. l., d. m. stoms, and f. w. davis. 1996. reserve selection as a maximal covering location problem. biological conservation 76:105-112. clark, f. s., and r. b. slusher. 2000. using spatial analysis to drive reserve design: a case study of a national wildlife refuge in indiana and illinois (usa). landscape ecology 15:75-84. colwell, r. k., and j. a. coddington. 1994. estimating terrestrial biodiversity through extrapolation. philosophical transactions of the royal society of london b 345:101-118. cowling, r. m., r. l. pressey, a. t. lombard, p. g. desmet, and a. g. ellis. 1999. from representation to persistence: requirements for a sustainable system of conservation areas in the species-rich mediterranean-climate desert of southern africa. diversity and distributions 5:51-71. cowling, r. m., r. l. pressey, r. sims-castley, a. le roux, e. baard, c. j. burgers, and g. palmer. 2003a. the expert or the algorithm? comparison of priority conservation areas in the cape floristic region identified by park managers and reserve selection software. biological conservation 112:147-167. cowling, r. m., r. l. pressey, m. rouget, and a. t. lombard. 2003b. a conservation plan for a global biodiversity hotspot the cape floristic region, hortal and lobo – synecological conservation framework 38 south africa. biological conservation 112:191216. cowling, r. m., a. t. knight, d. p. faith, s. ferrier, a. t. lombard, a. driver, m. rouget, k. maze, and p. g. desmet. 2004. nature conservation requires more than a passion for species. conservation biology 18:1674-1676. cushman, s. a., and k. mcgarigal. 2002. hierarchical, multi-scale decomposition of species-environment relationships. landscape ecology 17:637-646. danks, h. v. 2002. modification of adverse conditions by insects. oikos 99:10-24. dasmann, r. f. 1972. towards a system for classifying natural regions of the world and their representation by national parks and reserves. biological conservation 4:247-255. davis, a. j., l. s. jenkinson, j. h. lawton, b. shorrocks, and s. wood. 1998a. making mistakes when predicting shifts in species range in response to global warming. nature 391:783-786. davis, a. j., j. h. lawton, b. shorrocks, and l. s. jenkinson. 1998b. individualistic species responses invalidate simple physiological models of community dynamics under global environmental change. journal of animal ecology 67:600-612. dennis, r. l. h. 1997. an inflated conservation load for european butterflies: increases in rarity and endemism accompany increases in species richness. journal of insect conservation 1:43-62. dennis, r. l. h. 2001. progressive bias in species status is symptomatic of fine-grained mapping units subject to repeated sampling. biodiversity and conservation 10:483-494. dennis, r. l. h., and p. b. hardy. 1999. targeting squares for survey: predicting species richness and incidence of species for a butterfly atlas. global ecology and biogeography 8:443-454. dennis, r. l. h., and c. d. thomas. 2000. bias in butterfly distribution maps: the influence of hot spots and recorder's home range. journal of insect conservation 4:73-77. dennis, r. l. h., and t. g. shreeve. 2003. gains and losses of french butterflies: test of predictions, under-recording and regional extinction from data in a new atlas. biological conservation 110:131139. dennis, r. l. h., t. h. sparks, and p. b. hardy. 1999. bias in butterfly distribution maps: the effects of sampling effort. journal of insect conservation 3:33-42. dinerstein, e., d. m. olson, d. j. graham, a. l. webster, s. a. pimm, m. a. bookbinder, and g. ledec. 1995. a conservation assessment of the terrestrial ecoregions of latin america and the caribbean. the world bank, washington, d. c. dobson, a. p., j. p. rodríguez, w. m. roberts, and d. s. wilcove. 1997. geographic distribution of endangered species in the united states. science 275:550-553. dyson, r. g. 2004. strategic development and swot analysis at the university of warwick. european journal of operational research 152:631-640. dyson, r. g., and f. a. o'brien 1998. strategic development: methods and models. wiley, chichester. edwards, j. l., m. a. lane, and e. s. nielsen. 2000. interoperability of biodiversity databases: biodiversity information on every desktop. science 289:2312-2314. elith, j., c. h. graham, r. p. anderson, m. dudik, s. ferrier, a. guisan, r. j. hijmans, f. huettmann, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. m. m. overton, a. t. peterson, s. j. phillips, k. richardson, r. scachetti-pereira, r. e. schapire, j. soberon, s. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species' distributions from occurrence data. ecography 29:129-151. emerson, b. c. 2002. evolution on oceanic islands: molecular phylogenetic approaches to understanding pattern and process. molecular ecology 11:951-966. emmerson, m., m. bezemer, m. d. hunter, and t. h. jones. 2005. global change alters the stability of food webs. global change biology 11:490-501. engler, r., a. guisan, and l. rechsteiner. 2004. an improved approach for predicting the distribution of rare and endangered species from occurrence and pseudo-absence data. journal of applied ecology 41:263-274. faith, d. p. 1994. phylogenetic pattern and quantification of organismal biodiversity. philosophical transactions of the royal society of london b 345:45-48. faith, d. p., and p. a. walker. 1996a. environmental diversity: on the best-possible use of surrogate data for assessing the relative biodiversity set of areas. biodiversity and conservation 5:399-415. faith, d. p., and p. a. walker. 1996b. how do indicator groups provide information about the relative biodiversity of different sets of areas?: on hotspots, complementarity and pattern-based approaches. biodiversity letters 3:18-25. faith, d. p., and s. ferrier. 2002. linking beta diversity, environmental variation, and biodiversity assessment. debate response to j. f. duivenvoorden et al. 2002. science 295:636b. faith, d. p., s. ferrier, and p. a. walker. 2004. the ed strategy: how species-level surrogates indicate hortal and lobo – synecological conservation framework 39 general biodiversity patterns through an 'environmental diversity' perspective. journal of biogeography 31:1207-1217. ferrier, s. 2002. mapping spatial pattern in biodiversity for regional conservation planning: where to from here? systematic biology 51:331-363. ferrier, s., and a. guisan. 2006. spatial modelling of biodiversity at the community level. journal of applied ecology 43:393-404. ferrier, s., g. watson, j. pearce, and m. drielsma. 2002a. extended statistical approaches to modelling spatial pattern in biodiversity in northeast new south wales. i. species-level modelling. biodiversity and conservation 11:2275-2307. ferrier, s., m. drielsma, g. manion, and g. watson. 2002b. extended statistical approaches to modelling spatial pattern in biodiversity in northeast new south wales. ii. community-level modelling. biodiversity and conservation 11:23092338. fielding, a. h. 2002. what are the appropriate characteristics of an accuracy measure? pp. 271280 in j. m. scott, p. j. heglund, j. b. haufler, m. morrison, m. g. raphael, w. b. wall, and f. samson, eds. predicting species occurrences: issues of accuracy and scale. island press, covelo, california. flather, c. h., k. r. wilson, d. j. dean, and w. c. mccomb. 1997. identifying gaps in conservation networks: of indicators and uncertainty in geographic-based analysis. ecological applications 7:531–542. fleishman, e., and r. mac nally. 2003. distinguishing between signal and noise in faunal responses to environmental change. global ecology and biogeography 12:395-402. fleishman, e., j. r. thomson, r. mac nally, d. d. murphy, and j. p. fay. 2005. using indicator species to predict species richness of multiple taxonomic groups. conservation biology 19:11251137. fortin, m.-j., t. h. keitt, b. a. maurer, m. l. taper, d. m. kaufman, and t. m. blackburn. 2005. species' geographic ranges and distributional limits: pattern analysis and statistical issues. oikos 108:7-17. galante, e., and j. r. verdú 2000. los artrópodos de la “directiva hábitat” en españa. ministerio de medio ambiente, madrid. gaston, k. j. 1994. rarity. chapman & may, london. gaston, k. j. 1996. what is biodiversity? pp. 1-9 in k. j. gaston, ed. biodiversity. a biology of numbers and difference. blackwell science, oxford. gaston, k. j. 2003. the structure and dynamics of geographic ranges. oxford university press, oxford. gaston, k. j., and r. david. 1994. hotspots across europe. biodiversity letters 2:108-116. gaston, k. j., and a. s. l. rodrigues. 2003. reserve selection in regions with poor biological data. conservation biology 17:188-195. gewin, v. 2002. all living things, online. nature 418:362-363. gillespie, r. g., and g. k. roderick. 2002. arthropods on islands: colonization, speciation, and conservation. annual review of entomology 47:595-632. gladstone, w. 2002. the potential value of indicator groups in the selection of marine reserves. biological conservation 104:211-220. gorshkov, v. g., a. m. makarieva, and v. v. gorshkov. 2004. revising the fundamentals of ecological knowledge: the biota-environment interaction. ecological complexity 1:17-36. gotelli, n. j., and r. k. colwell. 2001. quantifying biodiversity: procedures and pitfalls in the measurement and comparison of species richness. ecology letters 4:379-391. graham, c. h., s. ferrier, f. huettman, c. moritz, and a. t. peterson. 2004. new developments in museum-based informatics and applications in biodiversity analysis. trends in ecology & evolution 19:497-503. green, r. e., a. balmford, p. r. crane, g. m. mace, j. d. reynolds, and r. k. turner. 2005. a framework for improved monitoring of biodiversity: responses to the world summit on sustainable development. conservation biology 19:56-65. gu, w., and r. k. swihart. 2004. absent or undetected? effects of non-detection of species occurrence on wildlife-habitat models. biological conservation 116:195-203. guisan, a., and n. e. zimmermann. 2000. predictive habitat distribution models in ecology. ecological modelling 135:147-186. guisan, a., and j.-p. theurillat. 2000. equilibrium modeling of alpine plant distribution: how far can we go? phytocoenologia 30:353-384. guisan, a., and w. thuiller. 2005. predicting species distribution: offering more than simple habitat models. ecology letters 8:993-1009. hampe, a. 2004. bioclimate envelopes: what they detect and what they hide. global ecology and biogeography 13:469-471. hannah, l., g. f. midgley, and d. millar. 2002. climate change-integrated conservation strategies. global ecology and biogeography 11:485-495. heikkinen, r. k., and s. neuvonen. 1997. species richness of vascular plants in the subarctic landscape of northern finland: modelling hortal and lobo – synecological conservation framework 40 relationships to the environment. biodiversity and conservation 6:1181-1201. hengeveld, r. 1990. dynamic biogeography. cambridge university press, cambridge. hewitt, g. m. 1996. some genetic consequences of ice ages, and their role in divergence and speciation. biological journal of the linnean society 58:247276. hewitt, g. m. 1999. post-glacial re-colonization of european biota. biological journal of the linnean society 68:87-112. hickling, r., d. b. roy, j. k. hill, and c. d. thomas. 2005. a northward shift of range margins in british odonata. global change biology 11:502-506. higgins, j. v., t. h. ricketts, j. d. parrish, e. dinerstein, g. powell, s. palminteri, j. m. hoekstra, j. morrison, a. tomasek, and j. adams. 2004. beyond noah: saving species is not enough. conservation biology 18:1672-1673. hill, j. k., c. d. thomas, r. fox, m. g. telfer, s. g. willis, j. asher, and b. huntley. 2002. responses of butterflies to twentieth century climate warming: implications for future ranges. proceedings of the royal society of london b 269:2163-2171. hingston, r. w. g. 1932. a naturalist in the guiana forest. edward arnold & co., london. hirzel, a. h. 2001. when gis come to life. linking landscapeand population ecology for large population management modelling: the case of ibex (capra ibex) in switzerland. page 106. institute of ecology, laboratory for conservation biology. university of lausanne, lausanne. hirzel, a., v. helfer, and f. métral. 2001. assessing habitat-suitability models with a virtual species. ecological modelling 145:111-121. hirzel, a. h., j. hausser, d. chessel, and n. perrin. 2002. ecological-niche factor analysis: how to compute habitatsuitability maps without absence data? ecology 83:2027-2036. hirzel, a. h., b. posse, p.-a. oggier, y. crettenand, c. glenz, and r. arlettaz. 2004. ecological requirements of a reintroduced species, with implications for release policy: the bearded vulture recolonizing the alps. journal of applied ecology 41:1103-1116. hortal, j. 2004. selección y diseño de áreas prioritarias de conservación de la biodiversidad mediante sinecología. inventario y modelización predictiva de la distribución de los escarabeidos coprófagos (coleoptera, scarabaeoidea) de madrid. phd thesis. departamento de biología, facultad de ciencias. universidad autónoma de madrid. hortal, j., and j. m. lobo. 2005. an ed-based protocol for the optimal sampling of biodiversity. biodiversity and conservation 14:2913-2947. hortal, j., j. m. lobo, and f. martín-piera. 2001. forecasting insect species richness scores in poorly surveyed territories: the case of the portuguese dung beetles (col. scarabaeinae). biodiversity and conservation 10:1343-1367. hortal, j., j. m. lobo, and f. martín-piera. 2003. una estrategia para obtener regionalizaciones bióticas fiables a partir de datos incompletos: el caso de los escarabeidos (coleoptera) ibérico-baleares. graellsia 59:331-344. hortal, j., p. garcia-pereira, and e. garcía-barros. 2004. butterfly species richness in mainland portugal: predictive models of geographic distribution patterns. ecography 27:68-82. hortal, j., p. a. v. borges, and c. gaspar. 2006. evaluating the performance of species richness estimators: sensitivity to sample grain size. journal of animal ecology 75:274-287. hortal, j., p. a. v. borges, f. dinis, a. jiménezvalverde, r. m. chefaoui, j. m. lobo, s. jarroca, e. brito de azevedo, c. rodrigues, j. madruga, j. pinheiro, r. gabriel, f. cota rodrigues, and a. r. pereira. 2005. using atlantis tierra 2.0 and gis environmental information to predict the spatial distribution and habitat suitability of endemic species. pp. 69-113 in p. a. v. borges, r. cunha, r. gabriel, a. f. martins, l. silva, and v. vieira, eds. a list of the terrestrial fauna (mollusca and arthropoda) and flora (bryophyta, pteridophyta and spermatophyta) from the azores. direcção regional de ambiente and universidade dos açores, horta, angra do heroísmo and ponta delgada. howard, p. c., p. viskanic, t. r. b. davenport, f. w. kigenyi, m. baltzer, c. j. dickinson, j. s. lwanga, r. a. matthews, and a. balmford. 1998. complementarity and the use of indicator groups for reserve selection inuganda. nature 394:472475. iverson, l. r., and a. m. prasad. 1998a. predicting potential future abundance of 80 tree species following climate change in the eastern united states. ecological monographs 68:465-485. iverson, l. r., and a. m. prasad. 1998b. estimating regional plant biodiversity with gis modelling. diversity and distributions 4:49-61. iverson, l. r., m. w. schwartz, and a. m. prasad. 2004. how fast and far might tree species migrate in the eastern united states due to climate change? global ecology and biogeography 13:209-219. jetz, w., and c. rahbek. 2001. geometric constraints explain much of the species richness pattern in african birds. pnas 98:5661-5666. jiménez-valverde, a., and j. hortal. 2003. las curvas de acumulación de especies y la necesidad de hortal and lobo – synecological conservation framework 41 evaluar la calidad de los inventarios biológicos. revista ibérica de aracnología 8:151-161. kareiva, p., and m. marvier. 2003. conserving biodiversity coldspots. american scientist 91:344351. kennedy, t. a., s. naeem, k. m. howe, j. m. h. knops, d. tilman, and p. reich. 2002. biodiversity as a barrier to ecological invasion. nature 417:636638. kitching, r. 2000. biodiversity, hotspots and defiance. trends in ecology and evolution 15:484-485. larsen, f. w., and c. rahbek. 2005. the influence of spatial grain size on the suitability of the highertaxon approach in continental priority-setting. animal conservation 8:389-396. lawton, j. h. 1997. the science and non-science of conservation biology. oikos 79:3-5. lawton, j. h. 1999. are there general laws in ecology? oikos 84:177-192. lehman, c. d., and d. tilman. 2000. biodiversity, stability, and productivity in competitive communities. the american naturalist 156:534552. lehmann, a., j. m. overton, and m. p. austin. 2002a. regression models for spatial prediction: their role for biodiversity and conservation. biodiversity and conservation 11:2085-2092. lehmann, a., j. r. leathwick, and j. m. overton. 2002b. assessing new zealand fern diversity from spatial predictions of species assemblages. biodiversity and conservation 11:2217-2238. levin, s. a. 1992. the problem of pattern and scale in ecology. ecology 73:1943-1967. levine, j. m., and c. m. d'antonio. 1999. elton revisited: a review of evidence linking diversity and invasibility. oikos 87:15-26. lewis, o. t., and s. r. bryant. 2002. butterflies on the move. trends in ecology and evolution 17:351352. lobo, j. m. 2001. decline of roller dung beetle (scarabaeinae) populations in the iberian peninsula during the 20th century. biological conservation 97:43-50. lobo, j. m., and g. halffter. 2000. biogeographical and ecologic factors affecting the altitudinal variation of mountainous communities of coprophagous beetles (coleoptera: scarabaeoidea): a comparative study. annals of the entomological society of america 93:115-126. lobo, j. m., and f. martín-piera. 2002. searching for a predictive model for species richness of iberian dung beetle based on spatial and environmental variables. conservation biology 16:158-173. lobo, j. m., and j. hortal. 2003. modelos predictivos: un atajo para describir la distribución de la diversidad biológica. ecosistemas 2003/1. lobo, j. m., i. castro, and j. c. moreno. 2001. spatial and environmental determinants of vascular plant species richness distribution in the iberian peninsula and balearic islands. biological journal of the linnean society 73:233-253. lobo, j. m., j. p. lumaret, and p. jay-robert. 2002. modelling the species richness of french dung beetles (coleoptera, scarabaeidae) and delimiting the predicitive capacity of different groups of explanatory variables. global ecology and biogeography 11:265-277. lobo, j. m., j. r. verdu, and c. numa. 2006. environmental and geographical factors affecting the iberian distribution of flightless jekelius species (coleoptera: geotrupidae). diversity and distributions 12:179-188. loiselle, b. a., c. a. howell, c. h. graham, c. h. goerck, t. brooks, k. g. smith, and p. h. williams. 2003. avoiding pitfalls of using species distribution models in conservation planning. conservation biology 17:1591-1600. lombard, a. t., r. m. cowling, r. l. pressey, and a. g. rebelo. 2003. effectiveness of land classes as surrogates for species in conservation planning for the cape floristic region. biological conservation 112:45-62. lomolino, m. v. 2004. conservation biogeography. pp. 293-296 in m. v. lomolino, and l. r. heaney, eds. frontiers of biogeography: new directions in the geography of nature. sinauer associates, inc., sunderland, massachussets. loreau, m., n. mouquet, and a. gonzalez. 2003. biodiversity as spatial insurance in heterogeneous landscapes. proceedings of the national academy of sciences of the united states of america 100:12765-12770. lumaret, j. p., and j. m. lobo. 1996. geographic distribution of endemic dung beetles (coleoptera, scarabaeoidea) in the western palaearctic region. biodiversity letters 3:192-199. mac nally, r., and e. fleishman. 2004. a successful predictive model of species richness based on indicator species. conservation biology 18:646654. mac nally, r., e. fleishman, j. p. fay, and d. d. murphy. 2003. modelling butterfly species richness using mesoscale environmental variables: model construction and validation for mountain ranges in the great basin of western north america. biological conservation 110:21-31. mac nally, r., a. f. bennett, g. w. brown, l. f. lumsden, a. yen, s. hinkley, p. lilywhite, and d. ward. 2002. how well do ecosystem-based planning units represent different components of biodiversity? ecological applications 12:900-912. hortal and lobo – synecological conservation framework 42 mackey, b. g., h. a. nix, m. f. hutchinson, j. p. macmahon, and p. e. fleming. 1988. assessing representativeness of places for conservation reservation and heritage listing. environmental management 12:501-514. mackinnon, j., and k. mackinnon. 1986a. review of the protected areas system in the afrotropical realm. iucn, unep, gland, switzerland and cambridge, uk. mackinnon, j., and k. mackinnon. 1986b. review of the protected areas system in the indo-malayan realm. iucn, unep, gland, switzerland and cambridge, uk. mackinnon, j., g. child, and j. thorsell. 1986. managing protected areas in the tropics. iucn, gland, switzerland. margules, c. r., and r. l. pressey. 2000. systematic conservation planning. nature 405:243-253. margules, c. r., a. o. nicholls, and m. p. austin. 1987. diversity of eucalyptus species predicted by a multi-variable environmental gradient. oecologia 71:229-232. margules, c. r., a. o. nicholls, and r. l. pressey. 1988. selecting networks of reserves to maximise biological diversity. biological conservation 43:63-76. margules, c. r., a. o. nicholls, and m. b. usher. 1994. apparent species turnover, probability of extinction and the selection of nature reserves: a case study of the ingleborough limestone pavements. conservation biology 8:398-409. martínez-meyer, e. 2005. climate change and biodiversity: some considerations in forecasting shifts in species potential distributions. biodiversity informatics 2:42-55. martín-piera, f. 2001. area networks for conserving iberian insects: a case study of dung beetles (col., scarabaeoidea). journal of insect conservation 5:233-252. martín-piera, f., and j. m. lobo. 2003. database records as a sampling effort surrogate to predict spatial distribution of insects in either poorly or unevenly surveyed areas. acta entomológica ibérica e macaronésica 1:23-35. midgley, g. f., l. hannah, d. millar, w. thuiller, and a. booth. 2003. developing regional and specieslevel assessments of climate change impacts on biodiversity in the cape floristic region. biological conservation 112:87-97. molnar, j., m. marvier, and p. kareiva. 2004. the sum is greater than the parts. conservation biology 18:1670-1671. murguía, m., and j. l. villaseñor. 2000. estimating the quality of the records used in quantitative biogeography with presence–absence matrices. annales botanici fennici 37:289–296. nicholls, a. o. 1989. how to make biological surveys go further with generalised linear models. biological conservation 50:51-75. nicholls, a. o., and c. margules. 1993. an upgraded reserve selection algorithm. biological conservation 64:165-169. nichols, w. f., k. t. killingbeck, and p. v. august. 1998. the influence of geomorphological heterogeneity on biodiversity. ii. a landscape perspective. conservation biology 12:371-379. nieto, m., m. bastir, f. j. cabrero-sañudo, j. hortal, c. martínez-maza, and j. rodríguez. 2004. does evolution evolve? pp. 232-252. homenaje a emiliano aguirre, volúmen de paleoantropología. museo arqueológico regional de madrid, alcalá de henares, madrid. noss, r. f. 1987. from plant communities to landscapes in conservation inventories: a look at the nature conservancy (usa). biological conservation 41:11-37. oliver, i., a. holmes, j. m. dangerfield, m. gillings, a. j. pik, d. r. britton, m. holley, m. e. montgomery, m. raison, v. logan, r. l. pressey, and a. j. beattie. 2004. land systems as surrogates for biodiversity in conservation planning. ecological applications 14:485-503. olson, d. m., and e. dinerstein. 1998. the global 2000: a representation approach to conserving the earth's most biologically valuable ecoregions. conservation biology 12:502-515. parmesan, c., n. ryrholm, c. stefanescu, j. k. hill, c. d. thomas, h. descimon, b. huntley, l. kaila, j. kullberg, t. tammaru, w. j. tennent, j. a. thomas, and m. warren. 1999. poleward shifts in geographical ranges of butterfly species associated with regional warming. nature 399:579–583. pearson, d. l., and s. s. carroll. 1999. the influence of spatial scale on cross-taxon congruence patterns and prediction accuracy of species richness. journal of biogeography 26:1079-1090. pearson, r. g., and t. p. dawson. 2003. predicting the impacts of climate change on the distribution of species: are bioclimatic envelope models useful? global ecology and biogeography 12:361-371. pearson, r. g., and t. p. dawson. 2004. bioclimate envelopes: what they detect and what they hide response to hampe (2004). global ecology and biogeography 13:471-473. peterson, a. t., j. soberón, and v. sánchez-cordero. 1999. conservatism of ecological niches in evolutionary time. science 285:1265-1267. peterson, a. t., s. l. egbert, v. sánchez-cordero, and k. p. price. 2000. geographic analysis of conservation priority: endemic birds and mammals in veracruz, mexico. biological conservation 93:85-94. hortal and lobo – synecological conservation framework 43 peterson, a. t., e. martínez-meyer, c. gonzálezsalazar, and p. w. hall. 2004. modeled climate change effects on distributions of canadian butterfly species. canadian journal of zoology 82:851-858. peterson, a. t., m. a. ortega-huerta, j. bartley, v. sánchez-cordero, j. soberón, r. h. buddmeier, and d. r. b. stockwell. 2002. future projections for mexican faunas under global climate change scenarios. nature 416:626-629. pharo, e. j., and a. j. beattie. 2001. management forest types as a surrogate for vascular plant, bryophyte and lichen diversity. australian journal of botany 49:23-30. pimm, s. l., and j. h. lawton. 1998. planning for biodiversity. science 279:2068-2069. pimm, s. l., g. j. russell, j. l. gittleman, and t. m. brooks. 1995. the future of biodiversity. science 269:347-350. polasky, s., j. d. camm, a. r. solow, b. csuti, d. white, and r. ding. 2000. choosing reserve networks with incomplete species information. biological conservation 94:1-10. possingham, h. p., s. j. andelman, m. a. burgman, r. medellín, l. l. master, and d. a. keith. 2002. limits to the use of threatened species lists. trends in ecology and evolution 17:503-507. powell, g. v. n., j. barborak, and s. m. rodriguez. 2000. assessing representativeness of protected natural areas in costa rica for conserving biodiversity: a preliminary gap analysis. biological conservation 93:35-41. prendergast, j. r., r. m. quinn, and j. h. lawton. 1999. the gaps between theory and practice in selecting nature reserves. conservation biology 13:484-492. prendergast, j. r., r. m. quinn, j. h. lawton, b. c. eversham, and d. w. gibbons. 1993. rare species, the coincidence of diversity hotspots and conservation strategies. nature 365:335-337. pressey, r. l. 2004. conservation planning and biodiversity: assembling the best data for the job. conservation biology 18:1677-1681. pressey, r. l., and a. o. nicholls. 1989. efficiency in conservation evaluation: scoring versus iterative approaches. biological conservation 50:199-218. pressey, r. l., and r. m. cowling. 2001. reserve selection algorithms and the real world. conservation biology 15:275-277. pressey, r. l., c. j. humphries, c. r. margules, r. i. vane-wright, and p. h. williams. 1993. beyond opportunism: key principles for systematic reserve selection. trends in ecology and evolution 8:124128. pressey, r. l., h. p. possingham, and c. r. margules. 1996. optimality in reserve selection algorithms: when does it matter and how much? biological conservation 76:259-267. pyke, c. r., and d. t. fischer. 2005. selection of bioclimatically representative biological reserve systems under climate change. biological conservation 121:429-441. raxworthy, c. j., e. martínez-meyer, n. horning, r. a. nussbaum, g. e. schneider, m. a. ortega-huerta, and a. t. peterson. 2003. predicting distributions of known and unknown reptile species in madagascar. nature 426:837-841. rodrigues, a. s. l., and k. j. gaston. 2002. optimisation in reserve selection procedures why not? biological conservation 107:123-129. rodrigues, a. s. l., s. j. andelman, m. i. bakarr, l. boitani, t. m. brooks, r. m. cowling, l. d. c. fishpool, g. a. b. da fonseca, k. j. gaston, m. hoffman, j. long, p. a. marquet, j. d. pilgrim, r. l. pressey, j. schipper, w. sechrest, s. n. stuart, l. g. underhill, r. w. waller, m. e. j. watts, and x. yan. 2003. global gap analysis: towards a representative network of protected areas. advances in applied biodiversity science 5:100. rodríguez, j. 1999. use of cenograms in mammalian palaeoecology. a critical review. lethaia 32:331347. rodríguez, j. 2004. stability in pleistocene mediterranean mammalian communities. palaeogeography, palaeoclimatology, palaeoecology 207:1-22. rodríguez, j. 2006. structural continuity and multiple alternative stable states in middle pleistocene european mammalian communities. palaeogeography, palaeoclimatology, palaeoecology 239:355-373. rodríguez, j., and m. nieto. 2003. influencia paleoecológica en mamíferos cenozoicos: limitaciones metodológicas. coloquios de paleontología vol. ext. 1:459-474. rojas-soto, o. r., o. alcantara-ayala, and a. g. navarro. 2003. regionalization of the avifauna of the baja california peninsula, mexico: a parsimony analysis of endemicity and distributional modelling approach. journal of biogeography 30:449-461. russell, a. f., l. l. sharpe, p. n. m. brotherton, and t. h. clutton-brock. 2003. cost minimization by helpers in cooperative vertebrates. proceedings of the national academy of sciences of the united states of america 100:3333-3338. saetersdal, m., i. gjerde, and h. h. blom. 2005. indicator species and the problem of spatial inconsistency in nestedness patterns. biological conservation 122:305-316. sánchez-cordero, v., v. cirelli, m. munguial, and s. sarkar. 2005. place prioritization for biodiversity hortal and lobo – synecological conservation framework 44 content using species ecological niche modeling. biodiversity informatics 2:11-23. sarkar, s., j. justus, t. fuller, c. kelley, j. garson, and m. mayfield. 2005. effectiveness of environmental surrogates for the selection of conservation area networks. conservation biology 19:815-825. sax, d. f. 2002. equal diversity in disparate species assemblages: a comparison of native and exotic woodlans in california. global ecology and biogeography 11:49-57. scholes, r. j., and r. biggs. 2005. a biodiversity intactness index. nature 434:45-49. scott, j. m., p. j. heglund, j. b. haufler, m. morrison, m. g. raphael, w. b. wall, and f. samson, eds. 2002. predicting species occurrences: issues of accuracy and scale. island press, covelo, california. segurado, p., and m. b. araújo. 2004. an evaluation of methods for modelling species distributions. journal of biogeography 31:1555-1568. sentís, c. 2002. retrovirus endógenos humanos: significado biológico e implicaciones evolutivas. arbor 677:135-166. seoane, j., l. m. carrascal, c. l. alonso, and d. palomino. 2005. species-specific traits associated to prediction errors in bird habitat suitability modelling. ecological modelling 185:299-308. simberloff, d. 1998. flagships, umbrellas and keystones: is single-species management passé in the landscape era? biological conservation 83:247257. skov, f., and j.-c. svenning. 2004. potential impact of climatic change on the distribution of forest herbs in europe. ecography 27:366-380. soberón, j., and j. llorente. 1993. the use of species accumulation functions for the prediction of species richness. conservation biology 7:480-488. soberón, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species' distribution areas. biodiversity informatics 2:1-10. sparks, t. h., d. b. roy, and r. l. h. dennis. 2005. the influence of temperature on migration of lepidoptera into britain. global change biology 11:507-514. spector, s. 2002. biogeographic crossroads as priority areas for conservation. conservation biology 16:1480-1487. srivasta, d. s., and j. h. lawton. 1998. why more productive sites have more species: an experimental test of theory using tree-hole communities. the american naturalist 152:510529. stockwell, d. r. b., and a. t. peterson. 2002. effects of sample size on accuracy of species distribution models. ecological modelling 148:1–13. su, j. c., d. b. debinski, m. e. jakubauskas, and k. kindscher. 2004. beyond species richness: community similarity as a measure of cross-taxon congruence for coarse-filter conservation. conservation biology 18:167-173. ter steege, h., n. pitman, d. sabatier, h. castellanos, p. van der hout, d. c. daly, m. silveira, o. phillips, r. vasquez, t. van andel, j. duivenvoorden, a. a. de oliveira, r. ek, r. lilwah, r. thomas, j. van essen, c. baider, p. maas, s. mori, j. terborgh, p. núñez vargas, h. mogollón, and w. morawetz. 2003. a spatial model of tree �-diversity and tree density for the amazon. biodiversity and conservation 12:22552277. thomas, c. d., and j. j. lennon. 1999. birds extend their ranges northwards. nature 399:213. thomas, c. d., e. j. bodsworth, r. j. wilson, a. d. simmons, z. g. davies, m. musche, and l. conradt. 2001. ecological and evolutionary processes at expanding range margins. nature 411:577-581. thomas, c. d., a. cameron, r. e. green, m. bakkenes, l. j. beaumont, y. c. collingham, b. f. n. erasmus, m. f. de siqueira, a. grainger, l. hannah, l. hughes, b. huntley, a. s. van jaarsveld, g. f. midgley, l. miles, m. a. ortegahuerta, a. t. peterson, o. l. phillips, and s. e. williams. 2004. extinction risk from climate change. nature 427:145-148. thomson, j. r., e. fleishman, r. m. nally, and d. s. dobkin. 2005. influence of the temporal resolution of data on the success of indicator species models of species richness across multiple taxonomic groups. biological conservation 124:503-518. thuiller, w., l. brotons, m. b. araújo, and s. lavorel. 2004a. effects of restricting environmental range data to project current and future species distributions. ecography 27:165-172. thuiller, w., m. b. araujo, r. g. pearson, r. j. whittaker, l. brotons, and s. lavorel. 2004b. biodiversity conservation uncertainty in predictions of extinction risk. nature 430: doi:10.1038/nature02716. tilman, d. 1999. the ecological consequences of changes in biodiversity: a search for general principles. ecology 80:1455-1474. tilman, d. 2000. causes, consequences and ethics of biodiversity. nature 405:208-211. tilman, d. 2001. an evolutionary approach to ecosystem functioning. proceedings of the national academy of sciences of the usa 98:10979-10980. tilman, d., and s. w. pacala. 1993. the maintenance of species richness in plant communities in r. e. ricklefs, and d. schluter, eds. species diversity in ecological communities. historical and hortal and lobo – synecological conservation framework 45 geographical perspectives. the university of chicago press, chicago. tilman, d., d. wedin, and j. knops. 1996. productivity and sustainability influenced by biodiversity in grassland ecosystems. nature 379:718-720. tilman, d., j. knops, d. wedin, p. reich, m. ritchie, and e. siemann. 1997a. the influence of functional diversity and composition on ecosystem processes. science 277:1300-1302. tilman, d., s. naeem, j. knops, p. reich, e. siemann, d. wedin, m. ritchie, and j. lawton. 1997b. biodiversity and ecosystem properties. science 278:1865c. tilman, d., p. b. reich, and j. m. h. knops. 2006. biodiversity and ecosystem stability in a decadelong grassland experiment. nature 441:629-632. tognelli, m. f. 2005. assessing the utility of indicator groups for the conservation of south american terrestrial mammals. biological conservation 121:409-417. tognelli, m. f., and d. a. kelt. 2004. analysis of determinants of mammalian species richness in south america using spatial autoregressive models. ecography 27:427-436. tognelli, m. f., c. silva-garcía, f. a. labra, and p. a. marquet. 2005. priority areas for the conservation of coastal marine vertebrates in chile. biological conservation 126:420-428. van jaarsveld, a. s., s. freitag, s. l. chown, c. muller, s. koch, h. hull, c. bellamy, m. krüger, s. endrödy-younga, m. w. mansell, and c. h. scholtz. 1998. biodiversity assessment and conservation strategies. science 279:2106-2108. vane-wright, d. 1996. identifying priorities for the conservation of biodiversity. pp. 309-344 in k. j. gaston, ed. biodiversity. a biology of numbers and difference. blackwell science, oxford. vane-wright, r. i., c. j. humphries, and p. h. williams. 1991. what to protect?-systematics and the agony of choice. biological conservation 55:235-254. venevsky, s., and i. venevskaia. 2005. hierarchical systematic conservation planning at the national level: identifying national biodiversity hotspots using abiotic factors in russia. biological conservation 124:235-251. walther, g. r., e. post, p. convey, a. menzel, c. parmesan, t. j. c. beebee, j. m. fromentin, o. hoegh-guldberg, and f. bairlein. 2002. ecological responses to recent climate change. nature 416:389-395. ward, t. j., m. a. vanderklift, a. o. nicholls, and r. a. kenchington. 1999. selecting marine reserves using habitats and species assemblages as surrogates for biological diversity. ecological applications 9:691-698. wardle, d. a., o. zackrisson, g. hörnberg, and c. gallet. 1997. biodiversity and ecosystem properties. science 278:1865c. warren, m. s., j. k. hill, j. a. thomas, j. asher, r. fox, b. huntley, d. b. roy, m. g. telfer, s. jeffcoate, p. harding, g. jeffcoate, s. g. willis, j. n. greatorex-davies, d. moss, and c. d. thomas. 2001. rapid responses of british butterflies to opposing forces of climate and habitat change. nature 411:65-69. wessels, k. j., s. freitag, and a. s. van jaarsveld. 1999. the use of land facets as biodiversity surrogates during reserve selection at a local scale. biological conservation 89:21-38. whittaker, r. j. 1998. island biogeography. ecology, evolution and conservation. oxford university press, oxford. whittaker, r. j., m. b. araújo, p. jepson, r. j. ladle, j. e. m. watson, and k. j. willis. 2005. conservation biogeography: assessment and prospect. diversity and distributions 11:3-23. williams, p. h. 1999. worldmap 4.1 windows. software and user document 4. privately distributed, london. williams, p. h., and c. j. humphries. 1994. biodiversity, taxonomic relatedness, and endemism in conservation. pp. 269-287 in p. l. forey, c. j. humphries, and r. i. vane-wright, eds. systematics and conservation evaluation. clarendon press, oxford. williams, p. h., n. d. burgess, and c. rahbek. 2000. flagship species, ecological complementarity and conserving the diversity of mammals and birds in sub-saharan africa. animal conservation 3:249260. willis, k. j., and r. j. whittaker. 2002. species diversity scale matters. science 295:1245-1248. wilson, e. o. 2002. the future of life. alfred a. knopf, new york. woodward, f. i., and d. j. beerling. 1997. the dynamics of vegetation change: health warnings for equilibrium 'dodo' models. global ecology and biogeography letters 6:413-418. yachi, s., and m. loreau. 1999. biodiversity and ecosystem productivity in a fluctuating environment: the insurance hypothesis. proceedings of the national academy of sciences of the united states of america 96:1463-1468. zimmermann, n. e., and f. kienast. 1999. predictive mapping of alpine grasslands in switzerland: species versus community approach. journal of vegetation science 10:469-482. bridging the biodiversity data gaps: recommendations of the biodiversity informatics, 8, 2013, pp. 41-58 bridging biodiversity data gaps: recommendations to meet users’ data needs daniel p. faith (1)*, ben collen (2), arturo h. ariño (3), patricia koleff (4), john guinotte (5), jeremy kerr (6) and vishwas chavan (7) (1) australian museum, 6 college street, sydney, nsw 2010, australia. (2) institute of zoology, zoological society of london, regent’s park, london, nw1 4ry, united kingdom (3) department of zoology and ecology, university of navarra, 31080 pamplona, spain. (4) national commission for the knowledge and use of biodiversity (conabio), mexico (5) marine conservation biology institute, 2122 112 th ave ne, suite b-300, bellevue wa 98004, usa. (6) university of ottawa, canada (7) global biodiversity information facility secretariat, universitetsparken 15, dk 2100,copenhagen, denmark *corresponding author abstract – freely available high quality, data on species occurrence and associated variables are needed in order to track changes in biodiversity. one of the main issues surrounding the provision of such data is that sources vary in quality, scope, and accuracy. publishers of such data must face the challenge of maximizing quality, utility and breadth of data coverage, in order to make such data useful to users. with the global biodiversity information facility (gbif), we recently conducted a content needs assessment survey to consolidate and synthesize major user needs regarding biodiversity data. we find a broad range of recommendations from the survey respondents, principally concerning issues such as data quality, bias, and coverage, and ease of access. we recommend a candidate set of actions for the gbif that fall into three classes: 1) addressing data gaps, data volume, and data quality, 2) aggregating data types that are relatively new to gbif, to support emerging new applications, and 3) promoting ease-of-use and providing incentives for wider use. addressing the challenge of providing high quality primary biodiversity data potentially can serve the needs of national and international biodiversity initiatives. these include the “flexible framework” for addressing the new 2020 biodiversity targets of the convention on biological diversity, the global biodiversity observation network (geo bon) and the new intergovernmental science-policy platform on biodiversity and ecosystem services (ipbes). each of these presents opportunities for countries to define appropriate actions and corresponding data needs, with links from local to global scales. introduction biodiversity is the variety of life, extending from the level of genes, to species, to ecosystems (norse et al. 1986; wilson 1988; wilson 1999; margalef 2000). the extent of the biodiversity crisis – the potential loss of much of this living variation – was highlighted again recently by the third global biodiversity outlook (secretariat of the convention on biological diversity 2010). it documented a decline in the measures of biodiversity, and an increase over the past 30 years in the pressures that cause biodiversity loss (butchart et al. 2010). in 2012, the united nations conference on sustainable development (uncsd or “rio+20”) marked 20 years since the original conference that gave birth to the convention on biological diversity. the major outcome document from the uncsd conference (uncsd 2012) reinforced the critical links between biodiversity conservation and sustainability. the recent call by scientists for new efforts to “map the biosphere” (wheeler et al. 2012) also highlighted the importance of biodiversity data for addressing many basic research question in biodiversity science. effective strategies for addressing biodiversity loss, and progressing biodiversity faith et al— bridging biodiversity data gaps 42 science, require that biodiversity-relevant information is made available in a useful form. the range of existing available data must be extended by strategically filling biodiversity knowledge gaps (collen et al. 2008). further, biodiversity data (data relevant to variation at the genes, species, and ecosystems levels) will be of limited use on their own. such data must be integrated at various spatial scales, not only with other environmental data, but also with a wide variety of types of socioeconomic data, in order to address the most pressing questions in biodiversity and sustainability science. all of these considerations highlight the fact that useful biodiversity databases must go beyond lists of species, and provide integrated data of a quality required for a range of research and applications. biodiversity data typically may have been generated for one specific intention, but then used subsequently for other, multiple, purposes (chapman 2005). biological specimen data originally collected by museums and herbaria have contributed to spatial models describing broader biodiversity patterns, to species richness estimation, and to ecological/ environmental studies on attributes of species, including population size, geographic distribution,, habitat and behaviour (pyke and ehrlich 2010; gropp 2012). increasingly, such data have helped to assess the impacts of threats to species, including pollution, disease, and climate change. if we consider one of the most prominent of those challenges, climate change, it is clear that assessment of climate change impacts is limited by lack of species-occurrence data, including limitations in mechanisms for data discovery and access (tews 2006; chavan and ingwersen 2009). the shortfall in available data highlights the ongoing need for a mechanism to facilitate sharing of existing and future biodiversity data, both within and among countries (gaikwad and chavan 2006; chavan and ingwersen 2009). the global biodiversity information facility (gbif), an inter-governmental organization, facilitates a discovery and publishing mechanism for free and open exchange/sharing of existing (and future) biodiversity data, and is the predominant international, publically funded resource for species occurrence data. through gbif, institutions and countries are discovering and publishing their data resources online with common exchange standards, and are part of a growing global network of shared biodiversity data. more than 388 million primary biodiversity data records are accessible (as of october 2012) through the gbif global networks data portal (http://data.gbif.org/). there are many studies highlighting the variety of uses for shared biodiversity data (e.g. chapman 2005; including examples of its potential for new scientific research and for decision making related to natural resources management). however, deficiencies in species occurrence data coverage also have been highlighted (e.g., yesson et al. 2007; collen et al. 2008; gbif 2010a; boakes et al. 2010; gaiji et al. 2013, this volume). deficiencies in coverage and data quality (bortolus 2008) reduce the magnitude and utility of available data (escobar et al. 2009). data concerns raised in the routine feedback from gbif users have included: lack of sufficient data (volume, depth, and density), lack of fitnessfor-use (precision, accuracy and authenticity), and lack of mechanisms for data discovery, publishing, and accessing data (chavan and ingwersen 2009; ariño et al. 2012; gbif 2010c). a common concern is that applications may not provide defensible conclusions because of the largely opportunistic way in which existing biodiversity data have been collated, and published (pyke and ehrlich 2010). content needs assessment to address these issues in a holistic manner, gbif established the content needs assessment task group (cna tg) composed of the authors of this paper (co-chairs: faith & collen; gbif senior programme officer: chavan) with a remit to identify major areas of opportunity to mobilize data in a way that better considers users’ needs (gbif 2009a). the objective of cna tg therefore was to investigate user needs regarding biodiversity data within the broad arena of biodiversity research. the cna tg was mandated also to provide recommendations to gbif that would (a) determine the priority questions that gbif mobilised data should be able to address for various areas of science and policy in the near, medium, and long-term, from local to global level including thematic areas, (b) for the identified faith et al— bridging biodiversity data gaps 43 priority questions, evaluate the content needs (volume, depth, and density) and data fitness-foruse (precision, accuracy, authenticity) for specific uses, (c) assess what unique scientific and policy contributions of gbif mobilised data that cannot be met easily through other mechanisms, (d) identify gaps in accessible data mapped against data needs, and (e) recommend strategies and priorities for data discovery and publishing through gbif network. in this paper, we draw on the work of cna tg to propose some guidelines to gbif for addressing user needs. however, given the broad importance of primary biodiversity data to biodiversity science, our findings may be more broadly applicable. the sections below present and discuss a candidate set of recommendations resulting from the cna tg analysis and discussion of the results of an extensive survey among biodiversity information stakeholders. the many opinions and comments provided through the survey have resulted in a variety of detailed suggestions, and we have refined these to produce recommendations falling under several different themes or contexts. as a complement to this list of recommendations, we also have drawn upon the current relevant scientific literature, including that relating to international and national biodiversity initiatives, in order to provide additional broad context for these recommendations (see also gbif 2011). methods the survey (see ariño et al. 2013, this volume for full details), was launched in 2009 in six languages (english, spanish, french, chinese, and russian) (gbif 2009b; gbif 2009c; and gbif 2009d). the survey consisted of 21 questions covering (a) respondent profile, (b) uses of primary biodiversity data, (c) access to primary biodiversity data, (d) data quality and quantity requirements, (e) species level data requirements, and (f) usefulness of gbif mobilised data. surveymonkey (http://www.surveymonkey.com) was used to administer the survey and retrieve responses. the survey was widely circulated using biodiversity-related lists and portals. resulting raw answers were retrieved from surveymonkey, demographic information was collected, and individual responses anonymised before proceeding to analysis. the raw files were converted into a purpose-made database (ariño et al., this volume) amenable for deeper analysis and cross-referencing. responses sent in languages other than english were translated, and all answers coded homogeneously while retaining the language demographics. the coded database was crosstabulated and analyzed variously for frequencies, correlations, and statistically summarized over several dimensions (ariño et al., this volume) to produce frequency tables, plots, maps and summaries that helped discover trends and steer our discussions. responses to the cna tg survey formed the basis for deliberations by the cna tg and consultation with colleagues from the gbif and other related communities. this process resulted in 14 main recommendations. results the survey resulted in responses from more than 700 individuals, providing more than 48,000 individual answers and nearly four-thousand individual verbatim comments. our analysis of the responses revealed a vast array of uses of primary biodiversity data, underlining the importance of its provision. importantly, the survey highlighted some new types of data in high demand, particularly data used to add value to the geographic distribution and taxonomic data already mobilized by gbif (ariño et al. 2013, this volume). at the same time, the survey highlighted a common lack of awareness about the extent of availability of accessible primary data. of great concern were the responses that identified unreliable or incomplete data as the reason for not using data portals such as gbif. recommendations we note that these recommendations do not nominate any specific actor, as many of these recommendations require collective, simultaneous, and/or coordinated actions by multiple actors. we discuss the recommendations under three sections below. first, we examine important issues regarding data gaps, data volume, and data quality. in the second section, we focus on the visions for new kinds of data, and new potential applications. we extend that section by considering some of the related national and global challenges raised by international biodiversity initiatives, including the new http://www.surveymonkey.com/ faith et al— bridging biodiversity data gaps 44 biodiversity targets (http://www.cbd.int/sp/targets) of the convention on biological diversity (unep 2010a), the global biodiversity observation network (geo bon) and the new intergovernmental science-policy platform on biodiversity and ecosystem services (ipbes). finally, in the third section we consider issues and recommendations that relate to strategies for improving ease of use of primary biodiversity data, over the wide range of potential applications, and with particular reference to the data mobilization by the gbif network. data gaps, data volume, and data quality the cna tg survey revealed that a major concern for biodiversity data users is the quality and coverage of the available data (ariño et al. 2013, this volume). highest among their concerns were geographic and taxonomic gaps in data coverage, and the need for data quality assurance. given the existing trends of publishing data predominantly by the data rich countries as opposed to biodiversity rich countries, some biases towards particular taxa, places or periods of time, are inevitable (yesson et al. 2007; robertson 2008). our recommendations focus on ways to address bias by identifying and filling data gaps. the problem of data quality raises a number of issues. it is critically important that sources of error can be accounted for. the recommendations made under this section (table 1) suggest strategies to address these issues, in order to ensure that gbif mediated data are of the highest possible quality. the recommendations relating to data quality issues in particular, would help deliver an improved process for user feedback and reporting of errors (including the ability to pin-point malicious records). a key outcome from these recommendations would be increased credibility and increased confidence in use of gbif mobilised data. further, data producers and primary publishers will be in better control of assessment and improving ‘fitness-for-use’ of gbif mobilised data prior to its discovery and publishing over the network. table 1. recommendations relating to data gaps, data volume, and data quality 1. recognise that geographic, temporal, taxonomic, and ecosystem gaps exist in currently mobilised data, develop tools to spot and identify biases at multiple geographic scales, and expedite efforts to bridge these gaps from local to global scale. 2. facilitate ways to overcome the inevitable biased nature of the data discovered and mobilised through the gbif network, promoting more uniform spread of primary biodiversity data (geographic, taxonomic, and temporal). a. develop best practice guidelines to help participants and potential data publishers to prioritise data discovery, capture, digitization, and publishing on a demand-driven basis, so that such proposals make business sense for donor agencies. b. draw ‘local-to-global’ scale strategy and action plans for discovery and publishing of primary biodiversity data. 3. encourage and increase investment in retrospective discovery, digitization and publishing of historical and time series datasets. a. expedite digitization of natural history collections data, especially for type specimens. 4. ensure increased access to authoritative taxonomic catalogues and expedite progress towards a ‘global names architecture’. 5. add and enhance features in the gbif data portal, promoting improved ‘fitness-for-use’ of data discovered and accessed through the portal. http://www.cbd.int/sp/targets faith et al— bridging biodiversity data gaps 45 6. initiate the following steps to enhance the trust-worthiness of gbif mobilised data: a. expedite the tagging of data to highlight possible errors, data quality, and uncertainty levels. b. develop standards for annotations and data user feedbacks. c. strengthen data quality assessment and quality enhancement mechanisms. d. develop distributed, decentralized, yet coordinated (inter-connected) annotation and feedback mechanisms. e. develop a user-friendly standard vocabulary for types of errors and types of quality issues. f. develop mechanisms for users to document the ‘use confidence rating’ at data set and data record level. g. improve pathways for data publishers to provide warnings about biases or errors in the data at an early stage of discovery and publishing process. h. make provision for inclusion of data problems (lack of fitness-for-use) relevant annotations or feedbacks together with data, not only at dataset level but at the data record level. i. develop contributor specific best practices and/or mechanisms for quality assurances and quality control. j. expedite efforts in improving taxonomic and geo-spatial quality of gbif mobilised data. this task includes attention to geo-referencing. k. improve fitness-for-use of data at the data producer and/or primary publisher stage. l. establish links to the original datasets to enhance the verifiability of the data. there are approximately 388 million primary biodiversity data records published through the gbif network (as on october 2012). among the key recommendations in this section are those calling for efforts to bridge gaps (taxonomic and geographic, from local to global scales) in the coverage of these records, and to overcome the inevitable biased nature of the data discovered and mobilised through the network. the type of information collated by gbif is being used in many applications, including models of species diversity status and change and tracking progress in conserving biodiversity (boakes et al. 2010). one of the main advantages of species occurrence data is that it has the potential to give more than a simple snap shot of biodiversity status and distribution (e.g. mace et al. 2010). a sampling that is representative and balanced will greatly promote such analyses. therefore strategies to create a balanced spread of data (geographically, taxonomically, and temporally) are essential to facilitate meaningful analysis and interpretation of biodiversity as a whole, as opposed to “sectorial” biodiversity (for example, biodiversity of birds or biodiversity in fennoscandia). gbif mobilised data are currently dominated by terrestrial ecosystems and higher animals (gbif 2010a; gaiji et al. 2013, this volume). a balanced spread of records across ecosystems and taxa provides a stronger basis for broad generalisations about biodiversity patterns that are not limited to narrow taxon groups, or highly-sampled geographical areas. the patterns of use revealed by the survey, spanning monocots to algae, and lower animal phyla, reflect the lack of data on such taxa. gbif should investigate taxonomic, geospatial, and temporal biases and errors in data at the level of individual records (rather than just at the higher data-set level). for many biodiversity studies, the principal data-coverage limitations include the ‘linnaean shortfall’ (we only know a fraction of the planet’s species) and the ‘wallacean shortfall’ (even for described species, we know little about geographic distributions for a review, see brito 2010). for faith et al— bridging biodiversity data gaps 46 example, in canada, perhaps 2/3 of all species are known to science, and in australia, perhaps only about 1/3 of all species are known. the fact that only a small fraction of known species have mobilised, high quality data, covering their full geographic distribution makes the task of biodiversity assessments even more daunting (collen et al. 2009). large geographic and taxonomic biases in biodiversity data from historical inventories have been widely documented (hortal et al. 2008). boakes et al. (2010) observed that the details of sampling biases (or validation) may be difficult to find in databases, and concluded that “compensating for these biases will be important for any study aiming to draw conclusions about real trends in biodiversity over time and space. accounting for biases in biodiversity samples depends on a clear knowledge of the source and nature of those biases.” boakes et al. (2010) investigated patterns of spatial and temporal bias using a database covering more than 200 years, for species in the avian order galliformes. among their various sources of species distribution data (museum collections, scientific literature, ringing records, ornithological atlases, and website reports from “citizen scientists”), museum data provided the best historical coverage of species' ranges. on the other hand, they concluded that such data were time-intensive to collect. one of the quickest ways to enhance data coverage would be to mobilise additional natural history collections. type specimens are valuable data resources that provide baselines for comparison, and historical context that is often missing from biodiversity trend data (lotze & worm 2009). however, historical context requires capacity to carefully re-examine “old” data. although there is no readily available mechanism to back track to original datasets which form the index of the gbif mediated data, and some of these resources have been either moved away from their original locations or are no longer online (and some may no longer exist), the gbif index itself could eventually be used as a historical repository of published dynamic data (i.e., data published and thus usable within the time frame when they were on line) through its periodical snapshots. a persistent id attributed to each specimen would greatly facilitate this task, and alleviate the limit to the ability of users to verify the authenticity or validity of the resource. one of the principal advantages of using occurrence data to evaluate changing biodiversity is that it can potentially provide time series information (e.g. mace et al. 2010). time series perhaps represents a more challenging gap than taxonomic or geographic imbalances in data. boakes et al. (2010) argued that an historical context to current changes can be set into context by long-term trends that reveal major shifts in abundance and composition of biological communities. unfortunately, many datasets do not seem to survive the typical 2-5 years worth of data support linked to funding cycles or graduate student tenure (likens 1989), and if they do, another limit seems to be reached when the research career of particularly dedicated individuals end (warner et al. 1995; ariño and pimm 1995). if data such as the type mediated by gbif are to be used to evaluate biodiversity change as a whole, as well as securing historical data for the future and making it available for users, sampling biases need to be understood and addressed (boakes et al. 2010). wheeler et al. (2004) argued that “some naively see the information technology challenge as liberating data from cabinets. the reality is that for all but a few taxa, much data is outdated or unreliable. many specimens represent undescribed or misidentified species. rapid access to bad data is unacceptable; the challenge is not merely to speed data access but to expedite taxonomic research.” these arguments justify the dual focus, in the recommendations listed above, on filling taxonomic (and other) gaps, and, at the same time, reducing data errors. users of gbif mobilised data can take advantage of several developments that help overcome limitations imposed by gaps and errors, including the development of robust biodiversity modelling approaches (e.g. providing surrogates for overall biodiversity patterns; e.g., ferrier 2002), and related ‘gap’ analyses (gbif 2010a; gaiji et al. 2013, this volume; otegui et al. 2013, this volume) which promote strategic growth in mobilised biodiversity data (both through new surveys and priorities for digitising collections data) (berendsohn et al. 2010; berents et al. 2010). faith et al— bridging biodiversity data gaps 47 there is increasing support (e.g. see juutinen et al. 2008) for a surrogates framework which uses models to combine environmental and primary biotic data, in order to provide best-possible use of available data to serve urgent needs of decisionmaking (e.g. faith and walker 1996). the approach uses primary species data as a starting point, but integrates this with environmental data to make better inferences about overall biodiversity (faith et al. 2004). an example of this perspective is found in linke et al. (2011): “with increasing availability of both geographic information system (gis) data and new, userfriendly modelling techniques, it is rapidly becoming easier to produce modelled species surrogates or highly informed physical surrogates. \moreover, with the availability of predictive modelling techniques that are robust to data-poor inputs, the use of environmental rather than species surrogates is only necessary where species data are extremely limited and the area very large, for example in the amazon.... a key advantage of systematic approaches is that they make the best use of existing data and can be applied where data are limited to generate reliable but coarse assessments.” the surrogates approach is able to make effective use of even the small amounts of data available for some taxonomic groups, because the model depends only on the combination of species (across all groups) to produce an estimated pattern for overall biodiversity. this strategy greatly increases the information content for a broad range of gbif data. further, because the robust models integrate widely available environmental variables, the models based on the gbif mobilised data can be applied in a region that has little gbif mobilised data of its own. biodiversity surrogates strategies might suggest that limited existing data are all that are needed. however, better quality data will generally produce more accurate surrogate models. importantly, initial surrogate models can be used to strategically fill in taxonomic and geographic gaps. one method, “survey gap analysis” (see ferrier 2002) indicates places most in need of new data collection or integration into gbif – so promoting the mobilisation of new data. similarly, brito (2010) suggests that directing surveys towards areas of known data deficiency will likely result in the discovery of species new to science, helping to address the linnean shortfall (see also, raxworthy et al. 2003). another strategy suggests that priorities for data mobilisation should focus in part of how best to improve current surrogates models (faith 2005). of course, the potential gains from new data and knowledge are further enhanced when the possibility exists that the new knowledge will provide many future applications, not just one. these considerations all highlight issues linked to our recommendations to fill data gaps and improve data quality, even while using robust methods to make best-possible use of the currently available data. these arguments again highlight the need to analyze errors and biases in existing data. the approaches suggested by boakes et al. (2010) for coping with various biases (including differences in taxonomic efforts, differences in sampling efforts in different localities) are concordant with our recommendations for the gbif network (see also hortal et al. 2006). the comments also strongly support a strategy in which gbif increases its’ capacity to discover and publish data from a wide range of sources, including extended data types (including historical or time series). we note that uncertainties are relevant to other data types parr et al. (2011) in reviewing “evolutionary informatics” highlight challenges for phylogenetic information including the adequate communication of uncertainties: “accommodating topological disagreement where necessary, would consolidate taxon names, phenotypic and geographical distributional data across clades.” extending data types and applications our cna tg survey raised issues relating to the needs for “new” kinds of data serving a broader range of applications (ariño et al. 2013, this volume). these needs may be addressed here by four recommendations (table 2). again, a number of the recommendations have subsidiary recommendations (lettered entries). faith et al— bridging biodiversity data gaps 48 table 2. recommendations relating to extended data types and applications 7. develop a biodiversity landscape map depicting gbif’s position (including its participants, and data publishers), role, unique advantages and collaborative strategies amidst the myriad biodiversity, biodiversity informatics initiatives of local to global scale. a. establish or initiate collaborations with initiatives, networks, and organizations outside of the sphere of biodiversity. 8. develop initiatives for dealing with new or under-appreciated kinds of data. a. expand discovery and access to new data sources, types and themes and types (e.g. observations including absence only records, population richness data, and other data associated with different types such as multimedia objects and environmental impact assessment etc). b. focus attention on mobilising data that will be useful for multilateral environmental agreements, mainly for those closely related to biodiversity issues such as cbd and cites, etc. c. develop mechanisms for linkages with other data types and data resources. d. provide functionalities for export, linkages and retrieval of data associated with gis shape files and/or polygons. 9. develop initiatives for enhancing applications, demonstrating that existence of gbif mobilised data indeed makes a difference in biodiversity conservation and sustainable use. a. demonstrate use of gbif mobilised data for applied environmental sciences as well as socially relevant issues, rather than merely for basic scientific purposes. b. promote publication of data addressing a wide spectrum of applications or usages of biodiversity data. 10. conduct a ‘content needs assessment’ (local, regional, thematic, and global scale) at frequent intervals. a. gbif participants could conduct multilingual content needs assessment exercise at regular intervals. b. ensure improved coordination while conducting content needs assessment exercise. c. ensure increased participation of stakeholders, policy makers, administrators, natural resources managers, representatives of civil society and non-governmental organizations into content need assessment activities. gbif is regarded as having an important role in enabling a ‘research infrastructure’ based on the discovery and publishing of the world’s primary biodiversity data (bridgewater et al. 2010; peterson et al. 2010). this role can create complementary linkages with initiatives within and outside of the sphere of biodiversity. these linkages will work particularly well when there is seamless access to both biodiversity and nonbiodiversity data resources. this calls for extended data types from disciplines such as genetics, ecology, fisheries, agro-biodiversity, and environmental impact assessment. these will enhance the potential use of gbif mobilized data and will demonstrate scientific, ecological, social and economic relevance of gbif network (sutherland et al. 2009). ariño et al. (2013, this volume) noted several key data needs suggested by the survey. for example, genetics/genomics data are needed for faith et al— bridging biodiversity data gaps 49 integrated biodiversity studies, suggesting benefits will be gained from gbif links to existing molecular sequence databases, including genbank and barcode of life. page (2008) notes that genomics and phylogenetics research call for identifiers, such as specimen codes and genbank accession numbers, as companions to conventional taxonomic names. these genomics and phylogenetics databases in turn call for integration with other auxiliary data that supports applications. the importance of such variables is highlighted in “genomic observatories” (gos), which focus sequencing efforts in locations that are already wellstudied and rich in associated data (e.g. data on environmental variables) (davies et al. 2012). a wide range of environmental and ecological data will complement standard biodiversity data in multi-disciplinary biodiversity science. mace and baillie (2010), for example, argued that “available data also tend to emphasize taxonomic diversity, rather than ecosystem functions and services”. these arguments are reinforced by survey respondents’ call for mobilised data relating to species’ trait data and corresponding functions for species. biodiversity studies that integrate species data, genomics, functional traits, and other data will be supported by emerging tools for research and policy that can incorporate all these kinds of data (e.g. faith et al. 2009). equally, greater integration of different data types can be accompanied by a shared analytical toolbox. for example, the barcode of life data portal (sarkar and trizna 2011) currently makes available a range of tools specifically for analysis of dna barcoding data, but also could take advantage of tools commonly used for other data types. arguments for expanded biodiversity-related data sometimes are reflected in an expanded definition of “biodiversity”, beyond its core interpretation as “living variation”. for example, diaz et al. (2009) consider biodiversity as “the number, abundance, composition, spatial distribution, and interactions of genotypes, populations, species, functional types and traits, and landscape units in a given system”. each of these aspects defines potential new data types that need to be considered in an integrated way by the gbif network and its partners (see also hardisty et al. 2013). to expedite such integration, gbif might link to progress already made in digitizing some of these types of biodiversity data. other useful progress is found in compiled databases such as that for phylogenetic relationships (see page 2008). important new data links of this kind include, for example, observations data from the global mountain biodiversity assessment (gmba), a cross-cutting network of diversitas. gmba has produced a metadata catalogue (http://gmba.unibas.ch/index/index.htm?output=pri ntable) linked to gbif, that describes ‘the "who, what, where, when, why and how" pertaining to the collection of a given ecological dataset on mountain biodiversity”. another example is found in the extensive information contained in existing atlas data bases, which provide collections of spatially explicit data on species occurrences. robertson et al. (2010) advocate expanded databases of this kind: “the most useful atlases have a good measure of sampling effort; include data collected at a fine enough resolution to link to habitat variables of potential interest; have a sufficiently large sample size to work with in a multivariate context; and offer clear, quantitative indications of the quality of each record to allow for the needs of users who have specific demands for high-quality data.” when these requirements are met, ecological biogeography and other inferences can be drawn, as demonstrated by exemplary cases such as escala et al. (1997). gbif users would benefit from links to the rapidly expanding information on ecosystem services (any processes or benefits derived from ecosystems; tallis and polasky 2009). while conservation of ecosystem services does not necessarily provide support for conservation of biodiversity, it does represent a key way to reduce the opportunity costs of retaining intact localities (faith 2011). the need to integrate ecosystem services into regional-scale biodiversity assessments is a prime example of the gains to be made by integrated data. indeed, this fundamental requirement for effective trade-offs and synergies among biodiversity conservation and other needs of society points to the need for integration of a wide variety of different data types. “unaccounted-for ecosystem services” have been recognised as a major information challenge http://gmba.unibas.ch/index/index.htm?output=printable http://gmba.unibas.ch/index/index.htm?output=printable faith et al— bridging biodiversity data gaps 50 (tallis and polasky 2009; eigenbrod et al. 2010). at the same time, biodiversity data gaps must be addressed in order to properly find synergies and trade-offs among these different aspects of human well-being. the biodiversity measures used in many recent trade-offs studies do not yet combine ecosystem services information with the available primary biotic data that would produce effective surrogates or indicators for overall (“wholesale”) biodiversity patterns (faith et al. 2010). the millennium ecosystem assessment (2005) called for “a ‘‘calculus’’ of global and regional biodiversity. this would allow global biodiversity gains and losses to be quantified in a unified way over various response strategies, thereby clearly identifying the trade-offs involved at a regional level. such a calculus of biodiversity depends on effective biodiversity surrogates. these will be based upon the best possible use of a combination of environmental and species (for example, museum collections) data.” a calculus of biodiversity therefore must use surrogates information to provide estimates of useful measures relating to biodiversity change, in order for biodiversity to effectively be “on the table” for decision-makers (for examples, see faith et al. 2008). international initiatives (discussed below) make the comprehensive use of mobilised primary biodiversity data even more urgent. national and international biodiversity initiatives our recommendations covering “extending data types, applications” can support links between gbif and important national and international biodiversity initiatives. because these initiatives require integration of biodiversity conservation goals and other factors (ecosystem services, socioeconomic costs, etc.), they not only reinforce our recommendations about “making a difference” in biodiversity conservation, but also link to the need to consider “new” types of data and applications. below, we particularly highlight some of the linkages that build on the capacity for gbif mobilised data to develop a calculus of biodiversity. 1. geo bon the group on earth observations (geo; www.earthobservations.org) and its global earth observing system of systems (geoss) seek to improve the coordination of new and existing “earth observation” data sets. one of the geoss systems (addressing to one of geo’s nine designated societal benefit areas) is the new biodiversity observation network (geo bon; andrefouet et al. 2008). geo bon is to provide capabilities for sharing observations regarding biodiversity, from regional to global scales. gbif is already established as a key partner in pursuing these goals (andrefouet et al. 2008). in accord with our recommendations, future efforts could pursue links between existing gbif mobilised data and the other data types that are required within geo bon (including remote sensing, ecosystem services, vegetation maps, and genomics data). case studies, and early geo bon products, could highlight the value added by such integration of gbif mobilised data with other information. for example, gbif mobilised data on species distributions can be used to add value to existing vegetation/ecosystem classifications (by estimating biotic overlaps among types), so providing improved surrogates (proxy) information for global biodiversity patterns (faith and walker 1996). one important new context is the red list for ecosystems (rodríguez et al. 2012), which may use gbif mobilised data to interpret these assessments at the species level (faith 2012). geo bon may present many opportunities for such value-adding using gbif mobilised data (see also couvet et al. 2012). biodiversity surrogates models based on combining available environmental data and gbif mobilised species data (faith and walker 1996; ferrier 2002; soberón and peterson 2009) may provide “essential biodiversity variables” (pereira et al. 2013) describing patterns of “overall” or wholesale biodiversity. these models can support a core geo bon strategy in which the interpretation of remotely-sensed changes in the extent and condition of land (or water) localities are interpreted through the “lens” of these inferred spatial patterns in biodiversity (andrefouet et al. 2008; faith et al. 2009). soberón and peterson (2009) propose a simple variant of the lens faith et al— bridging biodiversity data gaps 51 approach “based on the increased availability of raw data about occurrences of species, cuttingedge modelling techniques for estimating distributional areas, and land-use information based on remotely sensed data to allow estimation of rates of range loss for species affected by landuse conversion.” an important point is that the developing toolbox for biodiversity models and inferences that make use of gbif mobilised species data opens the door to integration with other kinds of data. for example, existing biodiversity surrogates modelling approaches may be applied also to genomics data (faith et al. 2009). such a common modelling framework highlights the fact that these previously separate data types need to be linked, supporting our recommendations related to new data types. 2. ipbes while geo bon focuses on observations systems, other relevant international initiatives focus on assessment – broadly, the task of bringing science to bear on policy. the new intergovernmental science-policy platform on biodiversity and ecosystem services, ipbes (http://www.ipbes.net) is to identify key scientific information needed for policymakers “for the conservation and sustainable use of biodiversity, long-term human well-being and sustainable development” (unep 2010b). ipbes also will catalyze new research to fill knowledge gaps. significantly, ipbes will respond to requests from individual nations and from other stakeholders. in considering the gaps in the science-policy interface that might be addressed by ipbes, unep (2009) pointed to the need to build on gbif’s increasing “collaboration with a wide range of organizations in order to explore the value of the data available, and to seek to combine it with other data meaningfully.” the challenges raised by ipbes provide important context for our recommendations. ipbes is to “perform regular and timely assessments of knowledge on biodiversity and ecosystem services and their interlinkages, which should include comprehensive global, regional and, as necessary, subregional assessments” (unep 2009). further, the assessments are to be based on “a clear and transparent process for sharing and incorporating relevant data.” (unep 2009). ipbes will identify “core variables” related to biodiversity to support ongoing assessments of biodiversity and ecosystem services at multiple scales (van jaarsveld et al. 2011). these are to provide “some degree of standardization in approach across regional and subregional assessments”. these core variables logically should include various indices of biodiversity based on the “core observations” provided through gbif. this process would benefit from an expanded gbif global infrastructure for mobilising and sharing biodiversity information (at various geographic scales and across national boundaries), together with gbif’s provision of information to common standards and open access principles. ipbes therefore provides an important opportunity to address our recommendations that call for biodiversity conservation applications and for greater integration of other types of data. 3. convention on biological diversity the convention on biological diversity has adopted 20 targets as part of its new strategic plan (unep 2010a). the strategic plan calls for gbif inputs towards implementation of the plan: “the following are key elements to ensure effective implementation of the strategic plan: ...global monitoring of biodiversity: work is needed to monitor the status and trends of biodiversity, maintain and share data, and develop and use indicators and agreed measures of biodiversity and ecosystem change; the geobiodiversity observation network, with further development and adequate resourcing, could facilitate this, together with global biodiversity information facility and the biodiversity indicators partnership.” one goal of the new strategic plan is that “the science base and technologies relating to biodiversity, its values, functioning, status and trends, and the consequences of its loss, are improved, widely shared and transferred, and applied”(unep 2010a). gbif clearly can help address this broad goal. gbif also can play a specific role in helping to faith et al— bridging biodiversity data gaps 52 provide a biodiversity calculus, based on gbif mobilised species data, as a basic building block for addressing the 20 new cbd biodiversity targets. this can help ensure that strategies for individual targets address overall biodiversity. as one example, target 11 calls for conservation of “at least 17 per cent of terrestrial and inland water, and 10 per cent of coastal and marine areas, especially areas of particular importance for biodiversity and ecosystem services” (unep 2010a). a given percentage area protected can have large or small representation of biodiversity; consequently assessment requires some explicit measurement of biodiversity coverage by protected areas. gbif data could help to achieve a balance among the sometimes competing goals of conservation of ecosystems, biodiversity, and ecosystem services. thus, gbif mobilised data should play a key role in providing the information base to assess whether conserved areas include those of particular importance for biodiversity. promoting ease-of-use and incentives for wide-use our cna tg survey raised concerns relating to ease of use that may be addressed by four recommendations (table 3; again, a number of the recommendations have subsidiary recommendations). these recommendations focus on enhancements of the gbif global discovery and access portals (including useful outputs), on creating incentives for users and contributors, and on design aspects that better link different data sources within a de-centralised system. table 3. recommendations relating to promotion of ease-of-use and incentives for wide-use 11. implement a ‘data publishing framework’, including mechanisms for incentivizing efforts for data volume mobilization, as well as data quality enhancement. a. develop a ‘data citation mechanism’ to adequately credit and acknowledge the contributions of all players in the data ‘life cycle’, from data creation to dissemination. 12. enhance the data portals (http://data.gbif.org, and participant portals) to facilitate ready to use processed outputs such as maps as images for immediate use in publications and reports. a. explore production of customized maps as output options. b. provide more easy to use output formats and data retrieval processes. 13. enhance the gbif network with infrastructure, services, standards, and tools that facilitate rapid and cost-efficient discovery and publishing of ‘fit-for-use’ primary biodiversity data. a. implement persistent identifiers at dataset and data record level. b. establish the data hosting center infrastructure across the gbif network. c. ensure that time required for indexing data sets is minimized. similarly, the gap between updates should be minimised. d. develop and publish the gbif internationalization strategy and action plans, enhancing the ability of the network to discover, publish and use multilingual data resources. e. enhance gbif’s ability to discover and publish datasets in their entirety. f. promote more national, regional, and thematic portals (and access points) in addition to global data portal (data.gbif.org), with features and services that would satisfy the needs of the cross-sectional stakeholders. 14. enhance gbif’s role as a discovery service within a decentralised system, where users can discover data using data descriptions contained in distributed metadata catalogues. http://data.gbif.org/ faith et al— bridging biodiversity data gaps 53 in our recommendations, one of the key themes is incentives that can expedite progress in discovery and publishing of primary biodiversity data. progress would be promoted if data publishing was regarded as having the same status as traditional scientific (peer-reviewed) scholarly publishing (chavan and ingwersen 2009). in accord with this, we recommend an early uptake of recommendations of its data publishing framework task group (gbif 2009e; moritz et al. 2011) and incentivisation mechanisms proposed (chavan and penev 2011; chavan and ingwesen 2009; ingwersen and chavan 2011; goddard et al. 2011). another key theme relates to the increasing decentralisation of data and the need for meta-data strategies to help users find the data they need. in recent years, the focus of the gbif network has shifted from data access to broader data discovery (gbif 2008). therefore, further enhancements are essential to broaden the coverage of data types discovered through the gbif data portal (gbif 2009f; gbif 2009g; gbif 2009h; kelling et al. 2008; morris et al. 2013, this volume), and to promote meta-data catalogues. as mentioned earlier, local-to-global scale strategies and action plans for discovery and publishing of biodiversity data are essential (berendsohn et al. 2010; gbif 2010a; gbif 2010b; gaiji et al. 2013, this volume; otegui et al. 2013, this volume). this requires that gbif establish a network of distributed ‘data hosting centers’ to provide user-friendly infrastructure for data publishing (goddard et al. 2011). in tune with this, gbif needs to better facilitate indexing data, and its discovery through data portal. these efforts might be enhanced by appropriate training and capacity development (e.g. coetzer 2012). several of the key recommendations would help increase efforts to make it easy for a wide variety of users to discover gbif. an example strategy that would address this is forging links to a new initiate through plos (public library of science). plos argue that: “scientists are amassing details about the scope and status of life’s variation at an accelerating rate. this aids our understanding of species distributions and their interactions over time. however, if we are to address the consequences of global environmental change for life’s future, biodiversity data must be integrated and synthesized to a much greater degree than they are at present, and this can be promoted by enhanced communication among the interested parties, and raising public awareness.” plos have launched a biodiversity hub (http://hubs.plos.org/web/biodiversity) to promote such communication of biodiversity studies. the biodiversity hub is to add value to such studies in the form of data/images, etc. thus, these added value contributions might serve to highlight uses of gbif mobilised data. conclusions the gbif content needs assessment task group investigated user needs relating to biodiversity data, and identified some major areas of opportunity to mobilize data serving these needs. the recommendations for gbif fell into three classes: 1) data gaps, data volume, and data quality, 2) new kinds of data and new applications, and 3) promoting ease-of-use and incentives for wide-use. addressing these challenges can serve the needs of international biodiversity initiatives, including the new biodiversity targets of the convention on biological diversity, the global biodiversity observation network, geo bon, and the new intergovernmental science-policy platform on biodiversity and ecosystem services (ipbes). however, as is evident from the content needs assessment survey (ariño et al. 2013, this volume), any user needs survey realistically provides just one glimpse at the evolving user landscape. one limitation of even the most up-todate survey is that it cannot anticipate the full range of future applications and data needs. further, any one survey will always have some bias in geographic coverage and/or scale. future content needs assessments no doubt will reflect the changing needs within what is a rapidly evolving multi-disciplinary biodiversity science. in that spirit, we finish by quoting from the recent call for major efforts to explore earth’s species and map their distribution (wheeler et al. 2012): “there is ample evidence that clever scientists and advances in technology will continue to find new uses for museum specimens”. this perspective no doubt applies across the whole spectrum of biodiversity data to be mobilized by gbif. faith et al— bridging biodiversity data gaps 54 literature cited andrefouet, s., m.j. costello, d.p. faith, s. ferrier, g.n. geller, r. höft, n. jürgens, m.a. lane, a. larigauderie, g. mace, s. miazza, d. muchoney, t. parr, h.m. pereira, r. sayre, r.j. scholes, m.l.j. stiassny, w. turner, b.a. walther, and t. yahara. 2008: the geo biodiversity observation network: concept document. geo group on earth observations, geneva, switzerland. ariño, a.h., v. chavan, and d.p. faith (2013). assessment of user needs of primary biodiversity data: analysis, concerns, and challenges biodiversity informatics, 8, pp-pp ariño, a.h., and s.l. pimm (1995). on the nature of population extremes. evolutionary ecology, 9 (4): 429-443. berendsohn, w.g., v. chavan, and j. macklin (2010). recommendations of the gbif task group on the global strategy and action plan for the mobilization of natural history collections data. biodiversity informatics 7: 67-71. berents, p., m. hamer, and v. chavan (2010). towards demand driven publishing: approaches to the prioritization of digitization of natural history collections data. biodiversity informatics, 7(2): 113-119. boakes, e., p.j.k. mcgowan, r.a. fuller, d. changqing, n.e. clark, k. o'connor, and g.m. mace (2010). distorted views of biodiversity: spatial and temporal bias in species occurrence data. plos biology, 8:1-11. bortolus, a. (2008). error cascades in the biological sciences: the unwanted consequences of using bad taxonomy in ecology. ambio, 37: 114-118. bridgewater, p., s. knapp, c. prip, and m. macdavette. (2010). gbif review 2009. from prototype to full operation: managing expectations. copenhagen: global biodiversity information facility, may 3 (2010). 39pp. brito, d. (2010). overcoming the linnean shortfall: data deficiency and biological survey priorities. basic and applied ecology 11: 709–713. butchart, s.h.m., m walpole, b. collen, a. van strien., j.p.w. scharleman, r.e.a. almond, j.e.m.baillie, b.bomhard, c. brown, j. bruno, k.e. carpenter, g.m. carr, j. chanson, a. chenery, j. csirke, n.c. davidson, f. dentener, m. foster, a. galli, j.n. galloway, p. genovesi, r. gregory, m. hockings, v. kapos, j.-f. lamarque, f. leverington, j. loh, m.a. mcgeoch, l. mcrae, a. minasyan, m. hernández morcillo, t. oldfield, d. pauly, s. quader, c. revenga, j. sauer, b. skolnik, d. spear, d. stanwell-smith, s.n. stuart, a. symes, m. tierney, t.r. tyrrell, j.-c. vié, and r. watson. (2010). global biodiversity decline continues. science 328: 1164-1168. chapman, a.d. (2005). uses of primary species occurrence data, version 1.0. copenhagen: global biodiversity information facility. 106 pp. isbn: 87-92020-01-1 (available as a standalone pdf from http://www.gbif.org). chavan, v. and p. ingwersen. (2009). towards a data publishing framework for primary biodiversity data: challenges and potentials for the biodiversity informatics community. bmc bioinformatics, 10(suppl 14): s2, doi: 10.1186/1471-2105-10-s14s2. chavan, v. and l. penev. (2011). the data paper: a mechanism to incentivise publishing in biodiversity science. bmc bioinformatics, 12(suppl 15):s2, doi:10.1186/1471-2105-12-s15-s2. coetzer, w., m. hamer, and f., parker-allie. (2012). a new era for specimen databases and biodiversity information management in south africa. biodiversity informatics 8: 1-11. collen, b., m. ram, t. zamin and l. mcrae. (2008). the tropical biodiversity data gap: addressing disparity in global monitoring. tropical conservation science 1(2): 97-110. collen, b., m. ram, n. dewhurst, v. clausnitzer, v.j. kalkman, n. cumberlidge, and j.e.m. baillie. (2009). broadening the coverage of biodiversity assessments. in: wildlife in a changing world an analysis of the 2008 iucn red list of threatened species (ed. by j.-c. vié, c. hilton-taylor & s.n. stuart), pp. 67-76. iucn, gland. couvet, d., v. devictor, f. jiguet, and r. julliard. (2011). scientific contributions of extensive biodiversity monitoring. c. r. biologie 334: 370– 377. davies, n., c. meyer, j.a. gilbert, l. amaral-zettler, j. deck, m. bicak, p. rocca-serra, s. assuntasansone, k. willis, and d. field. (2012). a call for an international network of genomic observatories (gos). gigascience 2012: 1:5. diaz, s., a. hector, and d.a. wardle. (2009). biodiversity in forest carbon sequestration initiatives: not just a side benefit. current opinon in environmental sustainability, 1: 55-60. faith et al— bridging biodiversity data gaps 55 escala, m.c., j.c. irurzun, a. rueda, and a.h. ariño (1997). atlas of the insectivora and rodentia of navarra. biogeographical analysis. publ. biol. univ. navarra, ser. zool. 25: 1-79. eigenbrod, f, b.j. anderson, p.r. armsworth, a heinemeyer, s gillings, d.b. rob, c.d. thomas, and k.j. gaston. (2010). representation of ecosystem services by tiered conservation strategies. conservation letters, 3(3): 184-191. escobar, f., p. koleff, and m. rös. (2009). assessment of the capacities for knowledge: the national biodiversity information system as a case study. in: conabio-undp (eds.) mexico: capacities for conservation and sustainable use of biodiversity. national commission for the knowledge and use of biodiversity and the united nations development programme, mexico, pp 24-29. accessible at http://www.biodiversidad.gob.mx/v_ingles/country/pd f/capacities.pdf faith, d. p. (2005). indicator taxa in support of the 2010 biodiversity target. gbif seed money prioritization ecat electronic conference. (may 25 – june 1, 2005). faith, d. p. (2011). ecosystem services and biodiversity option values. science online. 10 february 2011. http://www.sciencemag.org/content/330/6012/1745/re ply faith, d. p (2012). ecosystems on the edge: ecologically distinctive globally endangered. http://australianmuseum.net.au/ecosystems-on-theedge-ecologically-distinctive-globallyendangered faith, d. p., and p.a. walker. (1996). environmental diversity: on the best-possible use of surrogate data for assessing the relative biodiversity of sets of areas. biodivers. conserv. 5: 399–415. faith, d.p., s. ferrier, and p.a. walker. (2004). the ed strategy: how species-level surrogates indicate general biodiversity patterns through an ‘environmental diversity’ perspective. j. biogeogr. 31: 1207–1217. faith, d.p., s. ferrier, and k.j. williams. (2008). getting biodiversity intactness indices right: ensuring that “biodiversity” reflects “diversity. global change biology 14: 207–217. faith d.p., c.a. lozupone, d. nipperess, and r. knight. (2009). the cladistic basis for the phylogenetic diversity (pd) measure links evolutionary features to environmental gradients and supports broad applications of microbial ecology’s “phylogenetic beta diversity” framework. int. j. mol. sci. 10:4723-4741. faith, d.p., s.magallón, a.p. hendry, e. conti, t yahara, m.j. donoghue. (2010). evosystem services: an evolutionary perspective on the links between biodiversity and human-well-being. curr. opin. environ. sustain. 2: 66–74. ferrier, s. (2002). mapping spatial pattern in biodiversity for regional conservation planning: where to from here? systematic biology 51: 331– 363. gaiji, s, v. chavan, a.h. ariño, j. otegui, e. robles, and n. king. (2013). content assessment of the primary biodiversity data through gbif network: status, challenges and potentials. biodiversity informatics, this volume. gaikwad, j. and v. chavan. (2006). open access and biodiversity conservation: challenges and potentials for the developing world. data science journal, 5: 1-17. gbif. (2008). gbif work programme 2009-2010. copenhagen, 59 pp. accessible at http://www2.gbif.org/wp2009-10.pdf gbif. (2009a). content needs assessment task group. copenhagen: global biodiversity information facility. accessible at http://www.gbif.org/informatics/primary-data/taskgroups/cna-tg/. gbif. (2009b). gbif content needs assessment survey (2009). accessible at http://www.gbif.org/communications/news-andevents/showsingle/article/gbif-content-needsassessment-survey-2009/. last accessed 2010.12.26. gbif. (2009c). gbif content needs assessment survey 2009 available in chinese. http://www.gbif.org/communications/news-andevents/showsingle/article/gbif-content-needsassessment-survey-2009-available-in-chinese/. last accessed 2010.12.26. gbif. (2009d). gbif content needs assessment survey 2009 available in russian. http://www.gbif.org/communications/news-andevents/showsingle/article/russkaja-versijavoprosnika-po-vyjasneniju-potrebnostei-v/ http://www.sciencemag.org/content/330/6012/1745/reply http://www.sciencemag.org/content/330/6012/1745/reply http://www.gbif.org/informatics/primary-data/task-groups/cna-tg/ http://www.gbif.org/informatics/primary-data/task-groups/cna-tg/ faith et al— bridging biodiversity data gaps 56 gbif. (2009e). data publishing framework task group. accessible at http://www.gbif.org/informatics/primary-data/taskgroups/dpf-tg/ gbif. (2009f). adoption of persistent identifiers for biodiversity informatics: recommendations of the gbif lsid guid task group. 23p. accessible at http://www2.gbif.org/persistent-identifiers.pdf gbif. (2009g). report of the gbif gbrds stakeholders planning workshop. 24p. accessible at http://www2.gbif.org/gbrdsworkshopreportfinal.pdf gbif. (2009h). report of the gbif metadata implementation framework task group. 38pp. accessible at http://www2.gbif.org/gbif-miftgreport.pdf gbif. (2010a). state-of-the-network 2010: discovery and publishing of the primary biodiversity data through the gbif network. authored by chavan, v. s., gaiji, s., hahn, a., sood, r. k., raymond, m., and n. king. (2010). copenhagen: global biodiversity information facility, 36 pp. isbn: 8792020-13-5. accessible online at http://www.gbif.org. gbif. (2010b). best practice guide for ‘data discovery and publishing strategy and action plans’ version 1.0. authored by chavan, v. s., sood, r. k., and a. h. ariño. (2010). copenhagen: global biodiversity information facility, 29 pp. isbn: 8792020-12-7. accessible online at http://www.gbif.org. gbif. (2010c). gbif position paper on future directions and recommendations for enhancing fitness-for-use across the gbif network, version 1.0. authored by hill, a. w., otegui, j., ariño, a. h., and r. p. guralnick. (2010). copenhagen: global biodiversity information facility, 25 pp. isbn: 87-92020-11-9. accessible on-line at http://www2.gbif.org/gpp-final.pdf. gbif. (2011). gbif strategic plan 2012–2016 seizing the future. copenhagen, denmark: global biodiversity information facility. goddard, a., p. cryer, n. wilson and g. yamashita. (2011). data hosting infrastructure for primary biodiversity data. bmc bioinformatics 12(suppl 15): s5. gropp, r.e. (2012). increasing access to biological collections. bioscience 62: 703. hardisty, a., roberts, d. and the biodiversity informatics community. 2013.. bmc ecology 2013, 13:16. http://www.biomedcentral.com/14726785/13/16 hortal, j., a. jiménez-valverde, j.f. gómez, j.m. lobo, and a. baselga. (2008). historical bias in biodiversity inventories affects the observed environmental niche of the species. oikos 117(6):847-858. ingwersen, p. and v. chavan. (2011). indicators for data usage index: an incentive for publishing primary biodiversity data through global information infrastructure. bmc bioinformatics 12 (suppl 15): s3. juutinen, a., m. mönkkönen, and m. ollikainen. (2008). do environmental diversity approaches lead to improved site selection? a comparison with the multi-species approach. forest ecology and management 255: 3750-3757. kelling, s., b. ingole, b. daly, b. stein, d. lepage, e. o’tuama, j. cooper, m. jones, t. lahti, and v. chavan. (2008). recommendations of the gbif observational data task group. september 2008, pp. 21. likens, g. e. (ed.) (1989). long-term studies in ecology. springer-verlag, new york. linke, s., e. turak, and j. nel. (2011). freshwater conservation planning: the case for systematic approaches. freshwater biology 56: 6–20. lotze, h. k. and b. worm. (2009). historical baselines for large marine mammals. trends in ecology & evolution, 24: 254–262. mace, g. m. and j. e. m. baillie. (2010). the 2010 biodiversity indicators: challenges for science and policy. conservation biology, 21 (6): 1406–1413. mace, g. m., b. collen, r.a. fuller, and e.h. boakes. (2010). population and geographic range dynamics: implications for conservation planning. philosophical transactions of the royal society of london b, 365; 3743-3751. margalef, r. (2000). diversidad y biodiversidad. in: la diversidad biológica en españa, f. pineda, j.m. de miguel, m.a. casado, j. montalvo (eds). prentice hall, madrid. millennium ecosystem assessment. (2005). ecosystems and human well-being: responses. world resources institute: washington, dc, usa (2005). http://www.millenniumassessment.org/en/reports.a spx morris, r.a., a. olson, g. riccardi, g. whitebread, v. barve, g. hagedorn, p. leary, i. teage, d. mozzherin, c. freeland, m. carausu, j. cuarda, and v. chavan (2012). discovery and publishing of primary biodiversity data associated with http://www.gbif.org/informatics/primary-data/task-groups/dpf-tg/ http://www.gbif.org/informatics/primary-data/task-groups/dpf-tg/ http://www2.gbif.org/persistent-identifiers.pdf http://www2.gbif.org/gbrdsworkshopreport-final.pdf http://www2.gbif.org/gbrdsworkshopreport-final.pdf http://www2.gbif.org/gbif-miftg-report.pdf http://www2.gbif.org/gbif-miftg-report.pdf http://www.gbif.org/ http://www2.gbif.org/gpp-final.pdf http://www.millenniumassessment.org/en/reports.aspx http://www.millenniumassessment.org/en/reports.aspx faith et al— bridging biodiversity data gaps 57 multimedia resources: the audubon core strategies and approaches. biodiversity informatics, 8 moritz, t., s. krishnan, d. roberts, p. ingwersen, d. agosti, l. penev, m. cockerill, , and v. chavan. (2011). towards mainstreaming of biodiversity data publishing: recommendations of the gbif data publishing framework task group. bmc bioinformatics, 12(suppl.15):s1.doi:10.1186/14712105-12-s15-s1. norse, e.a., k.l. rosenbaum, d.s. wilcove, b. a. wilcox, w.h. romme, d.w. johnston, and m. l. stout. (1986). conserving biological diversity in our national forests. washington, d.c.: the wilderness society. otegui, j, a.h. ariño, v. chavan, and s. gaiji. (2011). on the dates of the gbif mobilized primary biodiversity data records. biodiversity informatics, 8 page, r.d.m. (2008). biodiversity informatics: the challenge of linking data and the role of shared identifiers. brief bioinform 9: 345-354. doi: 10.1093/bib/bbn022 parr, c.s., r. guralnick, n. cellinese, r.d.m. page. (2011). evolutionary informatics: unifying knowledge about the diversity of life. trends in ecology & evolution, 27: 94-103. peterson, a.t., d. canhos, u. gardenfors, r.j. scholes, y. shirayama, m.s. graham, and f. pando. (2010). global biodiversity information facility: 2010 forward look report. copenhagen: global biodiversity information facility, 17 august (2010). 45pp. pyke, g.h., and p.r. ehrlich. (2010). biological collections and ecological/environmental research: a review, some observations and a look to the future. biological reviews. 85: 247 – 266. raxworthy, c.j., e. martinez-meyer, n. horning, r.a. nussbaum, g.e. schneider, m.a. ortega-huerta and a.t. peterson. (2003). predicting distributions of known and unknown reptile species in madagascar. nature 426: 837-841. robertson, m.p., g.s. cumming, and b.f.n. erasmus. (2010). getting the most out of atlas data. diversity and distributions 16: 363–375. robertson, d.r. (2008). global biogeographical data bases on marine fishes: caveat emptor. diversity and distributions 14: 891-892. rodríguez, j.p.k., m. rodríguez-clark, d.a. keith, e.g. barrow, j. benson, e. nicholson, and p. wit. (2012). iucn red list of ecosystems, s.a.p.i.en.s [online], 5.2 | 2012, accessed 27 august (2012). http://sapiens.revues.org/1286 sarkar, i.n., and m. trizna. (2011). the barcode of life data portal: bridging the biodiversity informatics divide for dna barcoding. plos one 6(7): e14689. doi:10.1371/journal.pone.0014689 secretariat of the convention on biological diversity (2010). global biodiversity outlook -3. isbn: 929225-220-b, accessible at http://gbo3.cbd.int/, montreal, 94 pp. soberón, j. and a.t. peterson. (2009). monitoring biodiversity loss with primary species-occurrence data: towards national level indicators for the 2010 target of the convention on biological diversity. ambio, 38(1): 29-34. sutherland, w.j., w.m. adams, r.b. aronson, r. aveling, t.m. blackburn, s. broad, g. ceballos, i.m. cote, r.m. cowling, g.a.b. da fonseca, e. dinerstein, p.j. ferraro, e. fleishman, c. gascon, m. hunter jr, j. hutton, p. kareiva, a. kuria, d. w. macdonald, k. mackinnon, f.j. madgwick, m.b. mascia, j. mcneely, e.j. milner-gulland, s. moon, c.g. morley, s. nelson, d. osborn, m. pai, e.c.m. parsons, l.s. peck, h. possingham, s.v. prior, a.s. pullin, m.r.w. rands, j. ranganathan, k.h. redford, j.p. rodriguez, f. seymour, j. sobel, n.s. sodhi, a. stoti, k. vance-borland, and a.r. watkinson. (2009). one hundred questions of importance to the conservation of global biological diversity. conservation biology 23: 557-567. tews, j. (2006). biodiversity and climate change: a modelling perspective. in: focus on biodiversity research (editor: jan schwartz), isbn: 1-60021372-3, pp. 15. tallis, h. and s. polasky. (2009). mapping and valuing ecosystem services as an approach for conservation and natural-resource management. annals of the new york academy of sciences, 1162: 265-283. uncsd. (2012). the future we want. united nations conference on sustainable development rio de janeiro, brazil. http://www.uncsd2012.org/content/documents/727 the future we want 19 june 1230pm.pdf cited 1 august 2012. unep. (2009). gap analysis for the purpose of facilitating the discussions on how to improve and strengthen the science-policy interface on biodiversity and ecosystem services. unep/ipbes/2/inf/1. faith et al— bridging biodiversity data gaps 58 unep. (2010a). unep/cbd/bs/copmop/2/wg.1/crp.1 updating and revision of the strategic plan for the post-2010 period. convention on biological diversity. 2010. available online: http://www.cbd.int/cop/cop10/doc/advance-final-unedited-texts/advanceunedited-version-strategic-plan-footnote-en.doc. unep. (2010b). intergovernmental science-policy platform on biodiversity and ecosystem services. unep/gc.26/6. van jaarsveld, a.s. et al. (2011). international science workshop on assessments for ipbes. united nations university, tokyo, japan 2529 july 2011 workshop report. warner, s.c., k.e. limburg, a.h. ariño, m. dodd, j. dushoff, k.i. stergiou, and j. potts. (1995). time series compared across the land-sea gradient. in: ecological time series, t.m. powell & j.h. steele (eds.). chapman & hall, new york. pp. 242-273. wheeler, q.d., p.h. raven, and e.o. wilson. (2004). taxonomy: impediment or expedient? science 303(5656): 285. wheeler, q.d., s. knapp, d.w. stevenson, j. stevenson, s.d. blum, b.m. boom, g.g. borisy, j.l. buizer, m.r. de carvalho, a. cibrian, m.j. donoghue, v. doyle, e.m. gerson, c.h. graham, p. graves, s.j. graves, r.p. guralnick, a.l. hamilton, j. hanken, w. law, d.l. lipscomb, t.e. lovejoy, h. miller, j.s. miller, s. naeem, m.j. novacek, l.m. page, n.i. platnick, h. portermorgan, p.h. raven, m.a. solis, a.g. valdecasas, s. van der leeuw, a. vasco, n. vermeulen, j. vogel, r.l. walls, e.o. wilson, and j.b. woolley. 2012. mapping the biosphere: exploring species to understand the origin, organization and sustainability of biodiversity. systematics and biodiversity 10:1, 1-20 wilson, e.o. (1988). the current state of biological diversity. in: biodiversity, e. o. wilson editor, pp. 3-18. national academy press, washington, d. c. 521 pp. wilson, e.o. (1999). the diversity of life (reissue edn.) w.w. norton, new york,. yesson, c., p.w. brewer, t. sutton, n. caithness, j.s. pahwa, et al. (2007). how global is the global biodiversity information facility? plos one, 2(11): e1124 (doi:10.1371/journal.pone.0001124). http://www.cbd.int/cop/cop-10/doc/advance-final-unedited-texts/advance-unedited-version-strategic-plan-footnote-en.doc http://www.cbd.int/cop/cop-10/doc/advance-final-unedited-texts/advance-unedited-version-strategic-plan-footnote-en.doc http://www.cbd.int/cop/cop-10/doc/advance-final-unedited-texts/advance-unedited-version-strategic-plan-footnote-en.doc fundamental niches, distribution areas and presence-only data biodiversity informatics, 7, 2010, pp. 17-44 organization of occurrence-related biodiversity resources based on the process of their creation and the role of individual organisms as resource relationship nodes steven j. baskauf department of biological sciences vanderbilt university, nashville, tn, usa, steve.baskauf@vanderbilt.edu abstract. kinds of occurrences (evidence of particular living organisms) can be grouped by common data and metadata characteristics that are determined by the way that the occurrence represents the organism. the creation of occurrence resources follows a pattern which can be used as the basis for organizing both the metadata associated with those resources and the relationships among the resources. the central feature of this organizational system is a resource representing the individual organism. this resource serves as a node which connects the organism's occurrences and any determinations of the organism's taxonomic identity. i specify a relatively small number of predicates which can define the important relationships among these resources and suggest which metadata properties should logically be associated with each kind of resource. key words. guid, individual, metadata, occurrence, rdf traditionally, biodiversity informatics has involved the discovery and compilation of specimen data, and the application of those data to determine where organisms live. with the development of computers and the internet, the capability now exists to aggregate and disseminate biodiversity data globally as well as to broaden the scope beyond that of establishing species ranges (kelling, 2008). currently there is a coordinated effort to create the data infrastructure, standardize metadata terms (darwin core task group, 2009; jones et al. 2009), and develop persistent, actionable identifiers (globally unique identifiers or guids; cryer et al. 2009, richards 2009) in order to facilitate the aggregation of not only specimen metadata, but metadata derived from observations and digital representations of live organisms such as images. an important aspect of the development of this global network is creating the capacity for users (and the computer applications that assist them) to find other resources that are related to a resource which is at hand, i.e. "linked data" 1 . implicit in the concept of persistent identifiers is the understanding such identifiers are also resolvable, meaning that knowledge of the identifier also provides a means to access the information (data and metadata) for the resource to 1 http://www4.wiwiss.fu-berlin.de/bizer/pub/linkeddatatutorial/ which the identifier refers. an important part of this information is the relationship of the resource to related resources, but using these links to related resources is only possible if the relationships are specified in a consistent way that can be understood by all applications that wish to use them. in this paper, i suggest a system of organizing biodiversity resources around individual organism records. these records serve as nodes for grouping related occurrence resources and their taxonomic determinations. because this system is based on the way occurrence resources are created, it connects the resources in a logically consistent way and associates metadata properties with resources in an efficient manner. definitions in this paper, these terms are used according to their standard meaning as a defined in the citations or footnotes. the exceptions are abstract, conceptual, and defined abstract resources which for the purposes of this paper have been assigned specific meanings as defined below. resource a physical, digital or conceptual entity which can be identified by a uniform resource identifier (uri) (berners-lee et al. 2005). information resource a resource for which all essential characteristics can be transmitted in a mailto:steve.baskauf@vanderbilt.edu http://www4.wiwiss.fu-berlin.de/bizer/pub/linkeddatatutorial/ baskauf organization of occurrence-related biodiversity resources 18 message (jacobs and walsh 2004), i.e. a digital resource. examples: text and digital images. data the content of the message that provides the representation of the information resource. non-information resource a resource that cannot be transmitted electronically. technically, a resource is defined as a non-information resource when an http get request for the resource does not result in a 2xx "success" response, i.e. no data are returned (w3c technical architecture group 2005). examples: persons, specimens, observations, and taxonomic concepts. metadata data about data 2 . all resources can have metadata that describe their properties. physical resource a non-information resource that is a material thing 3 . examples: living organisms, specimens, and 35mm slide images. abstract resource a non-information resource that does not represent a particular material thing. examples: observations, protein structures, state boundaries, and concepts. conceptual resource an abstract resource that is subject to varying interpretation. examples: relationships, taxonomic concepts, and properties. defined abstract resource a resource which represents a defined circumstance or abstract object, and which therefore is not subject to interpretation. examples: observations, determinations, and mathematical concepts. http uri an identifier beginning with "http://" used to name resources on the world wide web ("the web"). an http uri may or may not refer to a retrievable resource on the web. http uris uniquely identify a single resource. node a resource in a data structure which is linked to other resources by references to them. in diagrams in this paper, nodes are represented by geometric shapes and the links between them are represented by lines or arrows. resources in a biodiversity context characteristics of occurrences resources there are a number of categories of resources that are of interest to the biodiversity community and all of the types of resources defined above are represented in these resource categories. the primary focus of this paper is on resources that are 2 http://dublincore.org/metadata-basics/ 3 http://purl.org/dc/terms/physicalresource instances of the darwin core class occurrence, which is defined as "the category of information pertaining to evidence of an occurrence in nature, in a collection, or in a dataset (specimen, observation, etc.)" (darwin core task group 2009). although the name "occurrence" suggests that such resources serve the purpose of documenting that an individual or population of individuals occurred at a particular time and place, categorizing a resource as an occurrence does not imply fitness for that use or any other particular use. an occurrence resource could also be used to document character states, to serve as a type specimen, to be used as a logo, or any combination of these or other uses. it should be clarified that the scope of this paper is restricted to occurrences as a subset of a broader category of biodiversity resources defined by kelling (2008) as "observational data". the scope of the paper includes observational data that document the presence of a single individual at a given point in time, but does not include observational data that are place-based (e.g. species checklists and measures of abundance; kelling 2008). for the purposes of this paper, an occurrence is assumed to document a single individual. however, most of what is discussed here could also apply to small populations of individuals that are of the same species. fig. 1. components of a generic occurrence resource what occurrence resources have in common is that they assert that an organism occurred somewhere at some time. in this paper, i will describe the process by which occurrence resources come into existence and how this process can help us organize our thinking about the relationship between resources and the properties that should be associated with particular types of resources. determining the associations and properties that are appropriate for particular http://dublincore.org/metadata-basics/ http://purl.org/dc/terms/physicalresource baskauf organization of occurrence-related biodiversity resources 19 resource types is critical for the successful implementation of an actionable guid system. in this paper i will distinguish between several aspects of occurrence resources. each occurrence resource is the result of an event 4 which results in the creation of the resource. during the event, a representation of the organism may be created. the resource and the possible representation with which the resource may be associated have properties that comprise the metadata for that resource. thus the occurrence resource may have three aspects: a resource creation event, a representation of the organism, and metadata (fig. 1). occurrence resource creation events in a resource creation event, the occurrence resource must be created from another resource. this follows from the fact that an occurrence resource is evidence of the existence of an organism. a resource such as a generic illustration of a bird which is not created from another particular resource is not an occurrence because it does not represent a particular organism. the resource creation event could itself be considered a separate resource with its own persistent identifier. however, because of the oneto-one relationship between the creation event and the created resource, it is more straightforward to simply consider the metadata for the event (i.e. time and location information) as a part of the metadata for the occurrence resource itself. when i say that a resource was "created" i mean that the resource and any associated representation was caused to become a separate entity by whatever means is associated with that particular type of resource. for example, preserved specimens are collected, living specimens are planted or moved from their natural environment, digital images are photographed, dna sequences are sequenced, etc. this is not saying that the creator necessarily caused the representation itself to come into existence. for example, a tree twig was actually "created" as a physical object by the tree, not the person who 4 http://purl.org/dc/dcmitype/event note that the use of "occurrence" in the definition of event is in the sense of "something happening" rather than the specific use of occurrence sensu darwin core. collects it. however, the specimen that has the twig as its representation of the tree was created by the collector. thus in this sense it is appropriate for all occurrence resources to use the dublin core (dublin core metadata initiative, dcmi 5 ) term dcterms:creator 6 ("an entity primarily responsible for making the resource") 7 to indicate the person (or institution) who caused the resource to come into existence. it is also important here to make the distinction between the event that results in the creation of the occurrence resource and the occurrence resource itself. for some resources, this distinction is obvious. for example, when a specimen is collected, there is an event (a segment of a collecting trip) that takes place at a particular time and place and the result of this event is a resource that we call a specimen. when an image of an organism is recorded, an event (a "photo shoot") that takes place at a particular time and place results in a resource we call a stillimage. unfortunately this distinction is blurred in the case of an observation because we use the word "observation" to mean at least two different kinds of things. webster's new world dictionary provides two definitions of "observation" as "4 a) the act or practice of noting and recording facts and events, as for some scientific study b) the data so noted and recorded" (neufeldt and guralnik 1994). the first definition (a) represents an observation as an event (as in "we conducted an observation") and the second (b) represents an observation as the created resource (as in "our observations were written in a notebook"). observation occurrences are resources of the second type. they have resource creation events, but they are not themselves events. (it is sometimes suggested that observation occurrences should be classed in the dublin core class event, but they are no more events than are specimens or images). in the remainder of this paper, unless otherwise noted use of the term "observation" is intended to represent sense (b). 5 http://dublincore.org/ 6 namespace http://purl.org/dc/terms/ 7 http://purl.org/dc/terms/creator http://purl.org/dc/dcmitype/event http://dublincore.org/ http://purl.org/dc/terms/creator baskauf organization of occurrence-related biodiversity resources 20 fig. 2. comparison of the components of the three general categories of occurrence resources. organism representations associated with occurrence resources occurrence resources fall into three of the general categories listed in the definitions section based on the nature of the way in which they represent the organism (fig. 2, table 1). these categories are natural divisions that are based on whether and how a consumer can retrieve a digital representation of the organism from which the resource is derived. there are several general points regarding these categories of occurrence resources. regardless of whether a digital representation is ultimately available, all three of these types of resources have associated metadata and they have many metadata terms in common. thus despite their differences in representation, metadata from these three types of resources can be combined in the same database. it should also be stated explicitly that in the case of physical resources, it is desirable to create information resources that provide internetdeliverable representations of the material artifacts. this is generally recognized and much effort has been expended toward digitization of physical resources. thus it follows that a conceptual system describing occurrence resources must be able to relate physical resources to information resources that are derived from them. the definition of an observation resource given here is a functional one, i.e. if an occurrence resource can never ultimately provide a digital representation of the organism, then it is an observation resource. this definition is consistent with other definitions of observations that are not explicitly based on the ultimate result of an http get request. the biodiversity information standards (tdwg) observations task group provides a working definition of observational data: "… the outcomes of acts of measurement using particular protocols within the context of any objective scientific measurement activity. examples include data from survey or monitoring efforts, controlled experiments, and sensor-derived measurements. in each case, the basic or atomic notion of an observation represents the outcome of some measurement taken of a defined attribute or characteristic of some 'entity' (e.g., an organism 'in the field', a specimen, a sample, an experimental treatment, etc.), within some context (possibly given by other observations). every observation entails the measurement of one or more properties of some real-world entity or phenomenon." 8 this definition makes two major points. one is that object of observations includes real world entities, including organisms in the field, specimens, and samples. the other is that observations measure defined attributes or characteristics. to place these statements in the formal language of uri-identified resources, observations establish values of properties that describe resources which can include occurrences (sensu darwin core). an important implication of this is that as defined here, observation resources record metadata (i.e. data about resources that can be described using language such as resource description framework, rdf) and not data in the sense of what is returned in the resolution of a 8 http://wiki.tdwg.org/observational/ http://wiki.tdwg.org/observational/ baskauf organization of occurrence-related biodiversity resources 21 table 1. categories of occurrence resource and their associated representations resource category how an organism is represented by that category specific representations how a consumer can acquire a representation of the organism information resource by a digital file digital image digital audio digital video dna sequence http get returns data which is the representation in digital form physical resource by a material artifact film photograph film movie audiotape sound recording preserved specimen living specimen seed tissue sample http get does not return data, but a redirect may ultimately result in access to an information resource that is a representation derived from the artifact observation resource (kind of defined abstract resource) no representation none http get returns nothing. no redirect can result in return of a representation. guid. this concept is explicitly stated in pereira (2009) which states that "core objects in biodiversity information systems such as … observations are not associated with immutable sequence of bytes that could be returned in the lsid getdata() call. it is reasonable not to return anything in the getdata() call." this concept of observation resources places them in the category of abstract resources (because they are not associated with a corresponding material object), although not in the category of conceptual resources (because the observation is a defined entity that is not subject to interpretation). this concept of an observation is also narrower than the general category of "observational data" as defined by kelling (2008). the tdwg observation task group definition of "observation" also includes metadata resulting from measurements of specimens and samples. however, in the system for categorizing resources presented here, such observations would not themselves be considered to be metadata for separate observation resources, but rather metadata associated with other resources that are physical. by the functional definitions in table 1, a numeric measurement of an organism in the field (e.g. a measured bill length of a bird) could be considered an information resource if the measurement is considered a very simplified representation of the organism and is an immutable series of bytes (i.e. data as opposed to metadata which are subject to change). the important point is that the functional categorization of the three general types of occurrence resources determines the kinds of metadata that one should expect for them and the way that those metadata should be organized. metadata associated with all occurrence resources despite the differences among the three general categories of occurrence resources, they all have several properties in common. the metadata records for all occurrence resources should include a globally unique identifier (dwc:occurrenceid), information about who created the resource (dcterms:creator), information about when the resource creation event occurred (dcterms:created), information about where the resource creation event occurred (location information), and information about the particular nature of the resource. this summarizes in the most basic way the "who, what, where and when" baskauf organization of occurrence-related biodiversity resources 22 desired by the global biodiversity information facility (gbif) 9 . recently, much effort has been exerted to establish precise location information for occurrence records, i.e. to geolocate them (chapman and wieczorek 2006). the ultimate goal of this effort is to create descriptors of location that are unambiguous and mappable by software. for newly created occurrence resources or those which have been precisely geolocated, the following set of four darwin core terms (dwc: 10 ) unambiguously define the location where the occurrence occurred: dwc:decimallatitude, dwc:decimallongitude, dwc:geodeticdatum, and dwc:coordinateuncertaintyinmeters 11 . the first two terms define the numeric coordinates of the location, the third term describes the reference system used (for gps measurements the default reference system is wgs84 which has the epsg code of epsg:4326 12 ; "unknown" is also a valid value), and the fourth term provides an estimate of the uncertainty of the measurement. a fifth term, dwc:locality, provides a verbal description of the location at the lowest geographic level. a locality description is important because it can make it easier for an observer on the ground to return to the location and because it can be used to validate the coordinates (chapman and wieczorek 2006). darwin core also provides many other terms to describe location using large-scale geographic and political subdivisions and these descriptors may be included in an occurrence record along with the five terms listed above. since the values of those additional terms may be subject to change over time as political boundaries and names change, and because they can theoretically be generated by software from the four basic mathematical terms, they should be considered of secondary importance. however, for older pre-existing records it may be difficult or impossible to establish a precise location for the occurrence. in such cases, the additional terms may be the only means available for specifying location. a description of the nature of the resource is important in order to help potential users of the metadata to assess the fitness of the resource for 9 http://www.gbif.org/ 10 namespace http://rs.tdwg.org/dwc/terms/ 11http://code.google.com/p/darwincore/wiki/location#verbatimlati tude,_verbatimlongitude 12 http://spatialreference.org/ their intended use. there are three terms that can be used to specify the nature of a particular resource. the property rdfs:type 13 has been designated as the means by which the nature of the biodiversity resource should be specified to consuming applications (example given in pereira et al. 2009). since resources should be typed according to the tdwg ontology "or other accepted vocabularies" (richards 2009), the tdwg ontology class taxonoccurrence 14 is the appropriate value of rdfs:type for occurrence resources. the term dcterms:type 15 is well known beyond the biodiversity community and has a defined type vocabulary 16 . the values physicalobject, stillimage, movingimage, and sound are appropriate for occurrence resources. however, there is no dcterms:type vocabulary term that adequately describes observation resources. the property dwc:basisofrecord 17 provides a more specific description of the resource than dcterms:type and uses a defined type vocabulary 18 that is designed particularly for the biodiversity community. those terms describe many of the kinds of resources that are occurrences and include preservedspecimen, fossilspecimen, livingspecimen, humanobservation, and machineobservation. metadata associated with specific categories of occurrence resources physical resources are distinguished from other occurrence resources by the physical artifact that represents the organism. therefore, the metadata that are most relevant to a physical artifact answer the question "how can i find/look at/get the artifact?" in a biodiversity context, physical artifacts are generally a part of museum/herbarium/botanical garden/zoo collections, so all of the terms related to where and how these artifacts are stored in a collection are relevant exclusively to physical resources. information resources are the only category of occurrence resources where the representation 13 http://www.w3.org/2000/01/rdf-schema#type 14http://rs.tdwg.org/ontology/voc/taxonoccurrence#taxonoccurre nce 15 http://dublincore.org/documents/dcmi-terms/#terms-type 16 dcmi type vocabulary section of http://dublincore.org/documents/dcmi-terms/ 17 http://rs.tdwg.org/dwc/terms/index.htm#basisofrecord 18 http://rs.tdwg.org/dwc/terms/type-vocabulary/index.htm http://www.gbif.org/ http://code.google.com/p/darwincore/wiki/location%23verbatimlatitude,_verbatimlongitude http://code.google.com/p/darwincore/wiki/location%23verbatimlatitude,_verbatimlongitude http://spatialreference.org/ http://www.w3.org/2000/01/rdf-schema#type http://rs.tdwg.org/ontology/voc/taxonoccurrence#taxonoccurrence http://rs.tdwg.org/ontology/voc/taxonoccurrence#taxonoccurrence http://dublincore.org/documents/dcmi-terms/%23terms-type http://dublincore.org/documents/dcmi-terms/ http://rs.tdwg.org/dwc/terms/index.htm%23basisofrecord http://rs.tdwg.org/dwc/terms/type-vocabulary/index.htm baskauf organization of occurrence-related biodiversity resources 23 itself can be stored and distributed electronically. thus, the foremost requirement for a digital resource is that the digital representation exists as data (i.e. a retrievable file of immutable bytes) from at least at one location in the internet. in addition to the url from which the data can be retrieved, metadata terms that specify the format of the data and the circumstances under which the information resource can be used (i.e. copyright, licensing, and attribution information) are important. the tdwg/gbif multimedia resources task group (mrtg) 19 draft standard metadata schema 20 is appropriate for occurrence information resources and its recommendations should be followed. since observation resources have no representation associated with them, there are no specific metadata terms required to describe the representation as is the case for physical and information resources. since the focus of most observations is measurement data, specifying the properties, format, and units of the measurements will be important aspects of observation resources metadata. occurrence resource creation networks general patterns in occurrence resource creation in the previous section it was established that various categories of occurrence resources are created from other resources during resource creation events and that it is sometimes desirable to create other occurrence resources from those occurrence resources. these relationships can be diagrammed generally by fig. 3. the following points can be made about the relationships among occurrence resources: 1. every occurrence resource must have a resource from which it is derived during a resource creation event. 2. every occurrence resource is the result of a single resource creation event. 3. an occurrence resource may serve as the source of one or more derivative occurrence resources, or it may have none. 4. an occurrence resource may be associated with a physical/digital representation of an organism. if it has no representation, it is considered in this context to be an observation resource. 19 http://www.keytonature.eu/wiki/mrtg 20 http://www.keytonature.eu/wiki/submission_v0.9 5. an occurrence resource which is an information resource may be derived from either a physical resource or another information resource. 6. an occurrence resource which is a physical resource will only be derived from another physical resource. (there may be trivial examples where a physical resource could be derived from an information resource, such as a paper print of a digital image. but since a general goal is to create distributable electronic representations of physical representations, information resources generally reside at the end of a chain of resource creation.) 7. an occurrence resource which is an observation resource cannot be derived from a representation associated with an information or physical occurrence resource. if one were to make "observations" of a physical or information occurrence resource (e.g. to take measurements from a specimen or image) those "observations" would be metadata that should be associated with the representation, and should not be considered a separate observation occurrence resource. thus observation resources can only be derived directly from an organism. 8. information or physical occurrence resources cannot be derived from observation resources because a digital or physical representation cannot be created from something that has no representation. individuals as the ultimate source of occurrence resources although an occurrence resource can be derived from another occurrence resource, ultimately all occurrence resources must originate through their chain of derivation back to an organism in its environment. i call this original organism an "individual" because it plays the role described in the definition of the darwin core term dwc:individualid 21 , "an identifier for an individual or named group of individual organisms represented in the occurrence. meant to accommodate resampling of the same individual or group for monitoring purposes. may be a global unique identifier… ". in clonal species or small species such as mosses, it may be difficult to determine whether samples collected at a 21 http://rs.tdwg.org/dwc/terms/individualid http://www.keytonature.eu/wiki/mrtg http://www.keytonature.eu/wiki/submission_v0.9 http://rs.tdwg.org/dwc/terms/individualid baskauf organization of occurrence-related biodiversity resources 24 fig. 3 generic relationships among a focal occurrence resource and other resources related to it particular location were collected from a single organism or a small population of organisms of the same species. thus a group of organisms of the same species in the same location may be considered an "individual" even if technically they are several individual organisms. if an individual were always represented by a single occurrence resource, then considering the individual to be a separate resource would be superfluous. properties related to taxonomic identity could as easily be associated with the occurrence resource as they could be associated with the individual. however, it is possible for a single individual to be observed, sampled, or imaged multiple times or to be represented by several different types of occurrence resources. it would be redundant to assign the same taxonomic identity metadata to multiple occurrence resources derived from the same individual rather than to assign those metadata to the individual and then link the occurrence resources to the individual. in addition, it would be inconsistent to make a determination of one taxonomic identity for one occurrence resource and make a different determination of identity based on and assigned to another occurrence resource derived from the same individual. that individual can only have a single identity, so any determinations assigned based on one resource should apply to all other resources derived from the same individual. the most straightforward way to ensure this is to assign the determinations to the individual rather than to the separate occurrence resources. the resource representing a wild individual can also connect determinations to resources derived from organisms other than that individual if those organisms arise from the individual through cloning (e.g. vegetative propagation) or even sexual reproduction. those additional derivative organisms (which generally would generally not be in their natural environment and would therefore be classified as living specimens) must have the same taxonomic identity as the wild individual from which they originate, so any determination based on a derivative organism must apply to the individual itself and all other organisms derived from it. some examples of resource creation networks the following four figures illustrate some examples of increasingly complex networks of resource creation. fig. 4 illustrates the types of resources involved in a classical museum collection. in this situation an herbaceous plant is collected in its entirety, so when the preserved specimen is created, the individual ceases to exist in the environment. it is easy to consider the plant and the specimen to be one and the same thing. because of the simplicity of this network, a simple and flat table in a database can be used to compile the data/metadata associated with specimens of this type. in a typical herbarium, the specimen is assigned a taxonomic identity (name) and the baskauf organization of occurrence-related biodiversity resources 25 fig. 4. diagram showing the relationships among an individual, resources derived from it (a specimen and a specimen image), and the resources' associated creation events in a classic museum collection scenario. plant's location is noted in the database for the specimen. later, the specimen is imaged and a single reference to the image is made in the database record for the specimen. in such a scenario, there is little benefit to considering the individual to be a separate resource because nearly all relevant metadata can be associated with the specimen. in fig. 5, an insect is located in a field survey. before the insect is collected and killed, it is photographed live in its environment with a digital camera. after the survey trip, the insect is pinned and added to the museum's collection. as a part of the museum's efforts to document its collection, the pinned specimen is imaged digitally. the museum is also supporting a molecular systematics effort and donates a small sample from the insect from which dna is extracted, sequenced, and submitted to genbank. later, as a part of an effort to develop a database of characters, the pinned specimen is dissected andseparate specimens are created of various body parts, which are then digitally photographed under a microscope. the original pinned specimen ceases to exist as an entity upon the creation of the separate part specimens. as they have now become a permanent part of the collection, the separate part specimens are assigned their own catalog numbers. fig. 6 represents a hypothetical collection scenario for a botanical garden. during a field expedition, a survey crew locates an unusual tree and takes a number of digital images of the tree. a botanist on the team reviews the images and tentatively identifies the tree as a rare species. the survey crew returns to collect cuttings for propagation and branch/leaf specimens from the tree for the botanical garden's herbarium. another team returns later in the season after fruit has set and collects seeds. part of the seeds are imaged and stored as a part of an effort to preserve germplasm of rare species. other seeds are germinated and the resulting plants are added to the garden's permanent collection. dna is collected from one of the living specimens in the garden and the dna sequence is submitted to genbank. additional herbarium specimens may be collected from the living specimens in addition to those collected from the original tree and other individuals may be vegetatively propagated and sent to other botanical gardens. although the various botanical garden specimens are separate organisms from individual in the wild, those living specimens are derived clonally or genetically from the individual in the environment and thus can provide information regarding the wild individual's taxonomic identity. fig.7 involves an effort to track an individual of a large, rare species of bird (e.g. whooping crane). the young bird is initially observed visually in the nesting ground and recorded in an ornithologist's field records. several days later the decision is made to capture and band it. the following year, the information from the band is recorded by a different ornithologist in the overwintering grounds. at a later time the bird is recaptured. a tissue sample is collected for dna extraction and analysis and the bird is fitted with a radio tracking device. data/metadata are collected throughout migration using telemetry from the tracking device. historically, biodiversity records have been organized around specimens. in a simple scenario such as fig. 4, this is practical. however, in figs. 5-7, occurrence information resources (i.e. images) are collected directly from the individual and are not derived from any specimen. in fig. 7, no specimen is collected at all. thus in any collection scenario more complicated than fig 5 it makes sense to consider the individual as an identified resource (i.e. having an assigned guid). baskauf organization of occurrence-related biodiversity resources 26 fig. 5. a more thorough museum collection scenario involving an insect specimen. metadata associated with individuals in previous sections it was noted that metadata associated with occurrence resources describes where, when, and by whom the resource was created as well as the nature of the resource itself and how any representations associated with the resource can be accessed. since an individual is by definition in its natural state (any individual taken out of its natural state would become a living specimen), no human entity can be assigned as its creator. although an individual could be considered to have had a time of creation (i.e. when it was born, hatched, or germinated), it is not possible to know when nor where this occurred without creating an occurrence resource. if the time of the organism's birth were noted by a human observer, then an observation resource would be created with the time of birth represented by the resource creation event metadata for the observation resource. if the birth were digitally imaged, then an information resource (stillimage) would be created. it is, in fact, not possible to know anything about any time of existence or location of an individual without an occurrence resource creation event occurring. for this reason, particular dates and locations of events should not be assigned to an individual. however, it would be possible, if desired, to assign to the individual the set of dates and locations for all resource creation events resulting in resources derived directly from the individual. such a set can be considered to define the spatial and temporal range of that individual. in figs. 4-7, such resource creation events originating directly from individuals are highlighted in gray. in addition to the conceptual reason why location should be associated with derived occurrence resources rather than the individual organism, there are also several practical reasons. foremost among these is that some organisms move, and since some resource creation events (those creating observations and images) do not baskauf organization of occurrence-related biodiversity resources 27 fig. 6. collection scheme for a botanical garden. necessarily result in the destruction of the organism and the fixing of its location at the point of death, it may not be possible to assign a single location to a mobile organism. therefore, the set containing all locations where occurrence resources were created directly from an individual represents the geographic range of that moving individual. it also makes sense to record multiple resource creation locations even for sessile organisms. for example, an individual plant may be imaged several times by a gps-enabled digital camera. because of inaccuracies in the determination of latitude and longitude recorded with the images, each location recorded by the camera represents an estimate of the actual location of the individual. in that situation, the set containing all locations where resource creation events occurred allows determination of a sort of "standard error" of location. it is also possible that locations associated with the creation of different resources from the same individual may have different precisions. for example, a specimen may first be collected from a tree and the tree's location determined by reference to a topographic map. at a later time, a digital image may be taken of the same tree with the tree's location determined by gps. the precision of the gps determination is likely to be greater than the determination from the map, so it would be important to maintain both sets of geographic coordinates along with their precisions so that a user could assess the relative utility of each for re-locating the tree. for the reasons given above, it is clear that location is logically an aspect of the creation event of each particular occurrence resource. however, it should be noted that since some occurrence resources may not be directly derived from an individual organism in its natural state, it cannot necessarily be inferred that an occurrence resource creation event contributes to documenting the distribution of a taxon. for example, a preserved specimen may not indicate the presence of a taxon representative at the place of its creation if it was collected from a living specimen in a botanical garden. a digital still image may not indicate the baskauf organization of occurrence-related biodiversity resources 28 fig. 7. bird survey based primarily on observations. presence of a taxon representative if it is an image of a specimen, while it may if it is a picture of a living individual in the wild. there is currently no term in the darwin core standard that is suitable for indicating unambiguously that the date and location of an event serves to document a taxon distribution. it can be established from the metadata that an event asserts that a representative of a taxon was present at the location of an occurrence resource creation event if the identifier for resource from which occurrence is derived is the same as the identifier of the individual that the occurrence represents. however, to serve the purpose of clarifying whether the resource creation event of an occurrence documents the distribution, the term sernec:documentsdistribution 22 has been defined. it should also be made clear that the typing of an occurrence resource using dcterms:type, rdfs:type, or dwc:basisofrecord does nothing to restrict fitness of use of that resource to any particular purpose. regardless of whether or not an occurrence resource is used to establish that a 22http://bioimages.vanderbilt.edu/rdf/terms#documentsdistribution ; appropriate text values are "true" or "false" taxon representative is present at a location, there is no restriction prohibiting that occurrence resource from simultaneously being used for any of a number of purposes, e.g. to define a character, as a learning tool, to illustrate a habitat, to serve as a type specimen, etc. fitness of use for such purposes may be noted through the inclusion of appropriate properties in the metadata for the occurrence resource there is no limit to the number of such properties that may be included. it is useful at this point to differentiate dcterms:creator from the darwin core term dwc:recordedby. as discussed in the section on resource creation events, in the context of occurrences, dcterms:creator represents the entity that caused any occurrence resource to come into existence. dwc:recordedby is a "list … of names of people, groups, or organizations responsible for recording the original occurrence". thus dwc:recordedby carries the additional connotation that the resource created by creator of an "original occurrence" (i.e. a resource created directly from an individual -a resource creation event represented by a gray arrow in figs. 4-7) can be used to establish the presence of a taxon at a http://bioimages.vanderbilt.edu/rdf/terms%23documentsdistribution baskauf organization of occurrence-related biodiversity resources 29 location. thus occurrence resources derived directly from individuals (sernec:documentsdistribution value="true") will usually have values for both dcterms:creator and dwc:recordedby, while resources not derived directly from individuals (sernec:documentsdistribution value="false") will have no value for dwc:recordedby. the same criteria hold for dcterms:created and dwc:eventdate. if a resource is derived directly from an individual, it should be given identical values for dcterms:created and dwc:eventdate, while a resource not derived directly from an individual should only be assigned a value for dcterms:created. the only general category of metadata that should logically be associated with individuals is information related to taxonomic determinations. as we shall see later, there is benefit to considering these determinations as separate resources rather than as metadata of the resource representing the individual. thus individuals as resources serve primarily as nodes that connect their derivative occurrence resources to taxonomic determinations. with few exceptions, the metadata for resources representing individuals will consist primarily of statements indicating the relationship of that individual to other resources. one such exception would be assignment of a value for dwc:establishmentmeans to indicate whether the individual was native, naturalized, adventive, or cultivated. relationships among data/metadata resources relationships networks the conceptual scheme of resource creation was focused on networks which traced the series of events by which occurrence resources of various types were created from individual organisms. in this section, i will focus on the relationships among resources of various kinds and the ways those relationships can be specified. in fig. 8, each shape represents a resource as defined in the definition section. each line represents an important relationship between two resources in the network. the heavy lines below the individual represent the connections established between pairs of resources by the occurrence resource creation events discussed previously. a given focal occurrence resource is connected by the line above to the resource from which it was derived. lines below the resource connect it to occurrence resources that were derived from it. since taxonomic identity is a property of the individual organism, the individual should have one or more determinations which link that individual with a general taxonomic concept (franz 2008), identified as "taxon" in fig. 8. each determination may be considered a unique, defined abstract resource linked to that individual and having its own assigned guid (dwc:identificationid). the taxonomic concept (dwc: taxonconceptid) to which the determination is linked is a conceptual resource that may be applied to many determinations and which has its own http uri (e.g. devries, 2009). the taxonomic concept itself may be linked to one or more scientific and vernacular names (not represented in fig. 8) which may be considered separate abstract resources having complex relationships among themselves (page 2006). in common usage, an initial determination may be called the "identification" and subsequent determinations may be called "annotations". however, regardless of order, each determination represents a functionally equivalent assessment of the individual's identity, having properties of dcterms:created, dwc:identifiedby, dwc:identificationremarks, etc. each determination has a relationship with one or more resources that served as the evidence by which the determination was made. this "determination based on" relationship is represented by the smalldashed lines. it should be noted that a determination has a particular resource (or resources) on which it is based, but particular resources do not have a particular determination that represents the identity of their source individual. the resource is a representation of the individual and the individual will have all of the determinations as possible assessments of its identity. since taxonomic identity is assigned to a resource through its relationship with the individual, this identity must be accessed via the individual at the root of the resource creation chain. the individual can be located by following the lines through resources up the derivation chain. however, it may be convenient to specify directly baskauf organization of occurrence-related biodiversity resources 30 fig. 8. network of important relationships among resources in a complex resource creation scheme. occurrence resources could represent some of those in fig. 6 (i=tree in environment, o1=seed, o2=botanical garden specimen, o3=genbank accession, o4=live plant image, o5=herbarium specimen, o6=herbarium specimen image). the source individual from which the resource was ultimately derived. the "source individual" relationship for each derived resource is indicated by the long-dash lines. specifying relationships between resources resource description framework (rdf) the links among resources discussed in the previous section can be defined formally using resource description framework (rdf) 23 , which has a well-defined syntax for representing resources and relationships. rdf can be represented several ways 24 . relationships can be shown as graphs, while rdf serialized as xml 25 is also the default means for describing metadata when a guid is resolved (richards 2009). figs. 9-11 show rdf graphs for examples of three of the major types of resources diagrammed in fig. 8 (the fourth, taxonomic concept, is beyond the scope of this paper). in these graphs, ovals represent resources named by http uris. 23 http://www.w3.org/rdf/ 24 http://www.w3.org/tr/rdf-primer/ 25 http://www.w3.org/xml/ rectangles represent literals (composed of unicode characters). the resource in the center of the graph is the subject of the relationships. each arrow in the graphs represents the nature of the relationship (referred to as a "predicate" in rdf) with a description of that relationship next to the arrow. the objects of the relationships (shapes pointed to by the arrows) are the metadata associated with the subject (the object of the relationship). in many cases, the predicate describing the relationship is a property that is a term from a well-known vocabulary. in particular, the following prefixes are used for qualified namespaces representing these vocabularies. dcterms="http://purl.org/dc/terms/" dwc="http://rs.tdwg.org/dwc/terms/" in addition, there are several terms for relationships described in this paper but which have not previously been defined. these terms have been formally defined in rdfs 26 and they have been assigned to the qualified namespace: 26 http://bioimages.vanderbilt.edu/rdf/terms http://www.w3.org/rdf/ http://www.w3.org/tr/rdf-primer/ http://www.w3.org/xml/ http://bioimages.vanderbilt.edu/rdf/terms baskauf organization of occurrence-related biodiversity resources 31 fig. 9. rdf graph for an occurrence resource (based on relationships outlined in fig. 8) fig. 10. rdf graph for an individual (based on relationships outlined in fig. 8) baskauf organization of occurrence-related biodiversity resources 32 fig. 11. rdf graph for a determination (based on relationships outlined in fig. 8) sernec="http://bioimages.vanderbilt.edu/rdf/ter ms#" these terms are: derivedfrom indicates the resource from which the subject (an occurrence resource) was derived. may be an individual or another occurrence resource derivativeoccurrence indicates an occurrence resource that was derived from the subject (the subject may be an individual or an occurrence resource) identifiesindividual indicates an individual that is identified by a determination basedonoccurrence indicates an occurrence resource used as the basis for a determination. usedindetermination indicates a determination for which the subject (an occurrence resource) was used as a basis for identification the properties of an occurrence (fig. 9; appendix a) fall into four general categories. the first includes general descriptive metadata that applies to all resources, such as the guid of the resource, and a dcterms:description of what it is. the second is general metadata that applies to all occurrence resources, as described in the section "metadata associated with all occurrence resources". other metadata terms related to occurrences may be included if necessary or desired. the third category of metadata is that which applies to the specific kind of occurrence (table 1). in general, information resources require metadata describing intellectual property rights, and use and attribution guidelines described in the mrtg schema. physical resources may be described using specific darwin core terms which apply to that type of occurrence. the fourth general category is properties that describe how the occurrence resource is related to other resources. these properties relate the resource to other resources from which it is derived or which are derived from it, determinations that are based on that occurrence, and the individual from which the resource is ultimately derived. in the example of fig. 9 (and its corresponding appendix a), literals were used for the values of properties referring to entities (persons and baskauf organization of occurrence-related biodiversity resources 33 institutions). however, http uri references could also be used and would be preferable if they represented a "semantic web-aware" resource (such as a foaf 27 file) that can be de-referenced. the metadata of an individual (fig. 10; appendix b) consists almost entirely of properties that define the relationships between the individual and resources derived from it, and the individual and determinations that assign the individual to a taxon concept. this highlights the role of an individual as a node that connects other resources. one property of the individual that is assigned a literal value is dwc:establishmentmeans, which can be used to indicate whether the individual is wild or cultivated. the metadata of a determination includes a description of who made the identification and when. but the most important aspect of the determination metadata is a description of the connections formed by the determination between an individual and a taxon concept that is consistent with the facts of the individual as illustrated by its occurrences. this connection is supported by including links to the occurrence resources on which the determination was based. if those occurrences are information resources, a user could examine them to verify the determination. one advantage of this system of organization is that it requires very few predicates to define the relationships among occurrence resources, individuals, and their determinations. by using generic descriptions of relationships followed by tagging 28 , the overall system is simpler than the alternative of creating different relationship types for different kinds of object resources. for example, sernec:derivativeoccurrence is used rather than creating separate predicates for each type of derivative occurrence object (e.g. rather than using dwc:associatedmedia and dwc:associatedsequences specifically to indicate images and sequences derived from specimens and tissue samples). similarly, by considering both identifications and annotations as one type of resource (determination), a single predicate can be used to indicate the relationship of another resource with the determination. if desired, an application can infer which determination is the 27 http://www.foaf-project.org/ 28 http://wiki.tdwg.org/twiki/bin/view/tag/subclassornot initial identification by finding the determination with the earliest dwc:dateidentified value. it should also be noted that relationships involving an occurrence resource do not have to be limited to those illustrated here. for example, nodes representing habitat sites or character definitions could group occurrence resources derived from different individuals. appropriately defined predicates could link the occurrence resources to those nodes through entries in the rdf xml file for the occurrence. the relationships outlined here represent the minimum set of relationships that are needed to establish the taxonomic identity that should be associated with the occurrence and to link the occurrence to others that are closely related to it. conclusions the role of individuals. resources representing individuals serve as nodes which connect multiple occurrence resources that are derived from that same individual, such as duplicate specimens, direct images of the live organism, and multiple observations. they also provide a single resource to which determinations can be assigned, avoiding the necessity of repeating identification metadata in all of the occurrence records for the same organism and insuring that annotations are applied to all of those occurrences. one deficiency of the conceptual scheme described in this paper is that it does not adequately handle circumstances where a single occurrence resource contains evidence of multiple individuals. examples would include the contents of a pitfall trap which included several insects, an image of a habitat showing several individuals of different species, and a specimen of an organism that showed evidence of parasitism by an individual of a different species. clarification of artifact types and their relationships to resource classes. by explicitly categorizing occurrence resources according to the nature of their representation of an organism (physical, digital, or none), it becomes clearer which metadata elements are most appropriately assigned to various particular kinds of occurrence resources. this categorization also makes it clear in the process of guid resolution what metadata should be http://www.foaf-project.org/ http://wiki.tdwg.org/twiki/bin/view/tag/subclassornot baskauf organization of occurrence-related biodiversity resources 34 expected by a consuming application for those kinds of resources. clarification of what determines that an occurrence provides information about the presence of a representative of a taxon at a location. it is clear from figs. 4-7 that occurrence resources assert that a taxon representative is present at a location if the occurrence resources were derived directly from the individual (i.e. gray arrows) and not because those resources are of any particular kind. this clarification can be made explicit through use of the term sernec:documentsdistribution. role of resource creation networks in defining possible relationships among resources. it is a basic premise of this paper that the assignment of metadata to resources, the structuring of that metadata, and the relationships among biodiversity resources expressed in their metadata should reflect process used to collect and create occurrence resources (i.e. the resource creation network). since this process is general and applies generically to many kinds of occurrence resources, a few general metadata structures and relationship types (illustrated in the rdf graphs and examples in appendices a-b) can be adopted rather than creating multiple special resource classes, structures, and relationships. applicability of this framework. significant challenges remain before the promise of a worldwide network of biodiversity information can come to fruition. in particular, implementation of the guids necessary for the functioning of this network is a difficult task. it is hoped that the framework outlined in this paper will facilitate the development and application of guids to occurrences by clarifying what types of resources should be identified by guids and which metadata terms can be appropriately used to specify the properties of those resources. acknowledgements i began to think about the issues raised in this paper during an email discussion of the sernec (southeast regional network of expertise and collections) live plant imaging subgroup during 2007-08 (http://www.sernec.org/files/summary-ofdiscussion.pdf). in particular, the comments of debbie paul, boyce tankersley, alexey zinovjev, and john wieczorek were helpful in developing my thinking as were earlier comments made by austin mast. irina kadis provided examples in the context of a botanical garden. alexey zinovjev, irina kadis, john wieczorek, and donald hobern made helpful comments on drafts of the manuscript. two anonymous reviewers provided useful comments and references. literature cited berners-lee, t., r. fielding, and l. masinter, 2005. uniform resource identifier (uri): generic syntax (request for comments: 3986). the internet society. (accessed 2009-12-15) 29 chapman, a.d. and j. wieczorek (eds), 2006. guide to best practices for georeferencing. copenhagen: global biodiversity information facility. (accessed 2010-01-15) 30 cryer, p., r. hyam, c. miller, n. nicolson, é.ó tuama, r. page, j. rees, g. riccardi, k. richards, r. white, 2009. adoption of persistent identifiers for biodiversity informatics: recommendations of the gbif lsid guid task group, 6 november 2009 (accessed 2009-12-15) 31 . darwin core task group, 2009. darwin core (accessed 2009-12-15) 32 . devries, p.j. 2009. geospecies knowledge base available from http://lod.geospecies.org. (accessed 2009-12-15) 33 . franz, n., r.k. peet, and a.s. weakley, 2008. on the use of taxonomic concepts in support of biodiversity research and taxonomy in q.d. wheeler (ed.) the new taxonomy. the systematics association special volume series 76, boca raton: crc press. 34 jacobs, i. and n. walsh (eds.), 2004. architecture of the world wide web, volume one (w3c recommendation 15 december 2004, accessed 2009-12-15). 35 jones, m.b., n. bertrand, j. holetschek, v. hutchison, b.c. ko, á.suárez-mayorga, m. meaux, w. ulate, d. watts, t. robertson, and é.ó tuama, 2009. report of the gbif metadata implementation framework task group (miftg). 16 december 29 http://tools.ietf.org/html/rfc3986 30 http://www2.gbif.org/biogeomancerguide.pdf 31 http://www2.gbif.org/persistent-identifiers.pdf 32 http://rs.tdwg.org/dwc/ 33 http://about.geospecies.org/ 34 http://www.sernec.org/files/usetaxconcepts_0.pdf 35 http://www.w3.org/tr/webarch/ http://tools.ietf.org/html/rfc3986 http://www2.gbif.org/biogeomancerguide.pdf http://www2.gbif.org/persistent-identifiers.pdf http://rs.tdwg.org/dwc/ http://about.geospecies.org/ http://www.sernec.org/files/usetaxconcepts_0.pdf http://www.w3.org/tr/webarch/ baskauf organization of occurrence-related biodiversity resources 35 2009. copenhagen: global biodiversity information facility. (accessed 2010-05-12). 36 kelling, s., 2008. significance of organism observations: data discovery and access in biodiversity research. report for the global biodiversity information facility, copenhagen. 37 neufeldt, v. and d. b. guralnik (eds.), 1994. webster's new world dictionary of american english, 3rd edition. new york: prentice hall. page, r.d.m., 2006. taxonomic names, metadata, and the semantic web. biodiversity informatics 3:115. 36 http://www2.gbif.org/gbif-miftg-report.pdf 37 http://www2.gbif.org/observational_data.pdf richards, k., 2009. tdwg guid applicability statement draft standard. biodiversity information standards (tdwg) 3 september 2009 (accessed 2009-12-15). 38 pereira, r., k. richards, d. hobern, r. hyam, l. belbin, and s. blum, 2009. tdwg life sciences identifiers (lsid) applicability statement (draft standard). biodiversity information standards (tdwg) 3 september 2009 (accessed 2009-1215). 39 w3c technical architecture group (tag) 2005. httprange-14: what is the range of the http dereference function? (accessed 2009-12-15) 40 38 http://www.tdwg.org/stdtrack/article/download/150/51 39 http://www.tdwg.org/stdtrack/article/download/150/51 40 http://www.w3.org/2001/tag/issues.html#httprange-14 http://www2.gbif.org/gbif-miftg-report.pdf http://www2.gbif.org/observational_data.pdf http://www.tdwg.org/stdtrack/article/download/150/51 http://www.tdwg.org/stdtrack/article/download/150/51 http://www.w3.org/2001/tag/issues.html#httprange-14 baskauf organization of occurrence-related biodiversity resources 36 appendix a. sample rdf file for the occurrence resource in fig. 9 using the example of the living specimen (o2) in the botanical garden scheme (figs. 6 and 8). other examples featuring an image of that specimen (o4) and one of the live plant images in fig. 6 are also provided. the following guids were used to represent resources in these examples as well as the other appendices: http://bigbotanicalgarden.org/individual/10224=i=tree in environment http://saveourspecies.org/germplasm/54899= o1=seed http://bigbotanicalgarden.org/collection/25993=o2=botanical garden specimen http://www.ncbi.nlm.nih.gov/nuccore/bq12345= o3=genbank nucleotide sequence, accession number bq12345 http://bigbotanicalgarden.org/image/52011= o4=live plant image of the living specimen. http://stateunivherbarium.org/vascular/67988=o5=herbarium specimen http://stateunivherbarium.org/vascular/67988#img=o6=image of herbarium specimen http://bigbotanicalgarden.org/image/44321=live plant image of the tree in environment. http://bigbotanicalgarden.org/ individual/10224#abcd= d1=first determination for individual http://bigbotanicalgarden.org/individual/10224#wxyz= d2=second determination for individual note that when possible, several alternative terms from well-known vocabularies (i.e. darwin core, dublin core, xmp, foaf) are used to describe the same property. all of these guids are fictitious and should not be assumed to be resolvable. however, the actual rdf example files can be downloaded by appending the file name in the example to http://bioimages.vanderbilt.edu/rdf/examples/, e.g. http://bioimages.vanderbilt.edu/rdf/examples/25993.rdf --------------------------------------- living botanical garden specimen (o2) contents of file http://bigbotanicalgarden.org/collection/25993.rdf containing the metadata for a living specimen identified as http://bigbotanicalgarden.org/collection/25993 en living specimen of arborus rarus http://bigbotanicalgarden.org/collection/25993 joe fields 2007-09-27 false http://bioimages.vanderbilt.edu/rdf/examples/25993.rdf baskauf organization of occurrence-related biodiversity resources 37 41.09834 -121.17611 epsg:4326 10 big botanical garden rdf formatted description of the living specimen http://bigbotanicalgarden.org/collection/25993 big botanical garden 2008-09-08t12:01:30-0800 en 2009-10-07t09:14:08-0800 2009-10-07t09:14:08-0800 --------------------------------------- still image (o4) of a living specimen contents of file http://bigbotanicalgarden.org/image/52011.rdf containing the metadata for an image identified as urn:lsid:bigbotanicalgarden.org:image:52011 of a living specimen identified as http://bigbotanicalgarden.org/collection/25993 en living specimen image with guid http://bigbotanicalgarden.org/image/52011 http://bigbotanicalgarden.org/image/52011 fred fotografer 2009-08-03t14:24:02-0800 digitalstillimage false (c) 2009 big botanical garden big botanical garden image by fred fotografer, big botanical garden living specimen of arborus rarus in big botanical garden available under creative commons attribution-noncommercial-share alike 3.0 license http://creativecommons.org/licenses/by-ncsa/3.0/us/ image of living arborus rarus specimen (whole tree) http://bigbotanicalgarden.org/about.htm http://bigbotanicalgarden.org/logo.jpg 41.09834 -121.17611 epsg:4326 10 big botanical garden best quality http://www.morphbank.net/?id=987654&imgtype=jpeg image/jpeg 3456 2304 thumbnail http://bigbotanicalgarden.org/thumbs/52011.jpg image/jpeg 100 67 rdf formatted description of the specimen image http://bigbotanicalgarden.org/image/52011 big botanical garden 2009-08-09t09:44:57-0800 en 2009-10-07t09:14:08-0800 2009-10-07t09:14:08-0800 --------------------------------------- still image of a living organism in the wild (e.g. live plant image in fig. 6) contents of file http://bigbotanicalgarden.org/image/44321.rdf containing the metadata for an image identified as http://bigbotanicalgarden.org/image/44321of an individual in its environment identified as http://bigbotanicalgarden.org/individual/10224 en image with guid http://bigbotanicalgarden.org/image/44321 of individual in the wild http://bigbotanicalgarden.org/image/44321 frederico fotografo 1998-11-16t09:14:36-0300 digitalstillimage true frederico fotografo 1998-11-16t09:14:36-0300 (c) 1998 instituto de plantas nativas instituto de plantas nativas image by frederico fotografo, instituto de plantas nativas arborus rarus in rio grande do sul, brasil available under creative commons attribution-noncommercial-share alike 3.0 license http://creativecommons.org/licenses/by-ncsa/3.0/br/ image of living arborus rarus specimen (whole tree) http://bigbotanicalgarden.org/about.htm http://bigbotanicalgarden.org/logo.jpg -29.633905 baskauf organization of occurrence-related biodiversity resources 41 -50.204086 epsg:4326 10 east side of rs-484 7.9 km north of junction with br 101 best quality http://bigbotanicalgarden.org/raw-imports/dsc00234.jpg image/jpeg 1440 2160 thumbnail http://bigbotanicalgarden.org/thumbs/44321.jpg image/jpeg 67 100 rdf formatted description of the live organism image http://bigbotanicalgarden.org/image/44321 big botanical garden 1999-03-15t14:54:51-0800 en 2009-10-07t09:14:08-0800 2009-10-07t09:14:08-0800 --------------------------------------- for a functioning example of an occurrence guid, enter http://bioimages.vanderbilt.edu/baskauf/66921 into a web or rdf browser. through content negotiation, it will resolve to either http://bioimages.vanderbilt.edu/baskauf/66921.htm if content-type text/html is requested by a web browser, or http://bioimages.vanderbilt.edu/baskauf/66921.rdf if content-type application/rdf+xml is requested by an rdf browser. http://bioimages.vanderbilt.edu/baskauf/66921 http://bioimages.vanderbilt.edu/baskauf/66921.htm http://bioimages.vanderbilt.edu/baskauf/66921.rdf baskauf organization of occurrence-related biodiversity resources 42 appendix b. sample rdf file for the resource representing an individual in fig. 10 and the determination resource in fig. 11 using the example of the tree (i) and determination 2 (d2) in the botanical garden scheme (figs. 6 and 8). the same guids were used as in appendix a. contents of file http://bigbotanicalgarden.org/individual/10224.rdf containing the metadata for an individual organism identified as http://bigbotanicalgarden.org/individual/10224 and its determinations en field individual of arborus rarus --> native determination of arborus rarus var. microcarpa for the individual http://bigbotanicalgarden.org/individual/10224 baskauf organization of occurrence-related biodiversity resources 43 juanita naturalista 1998-11-17 arborus rarus microcarpa variety j. rodriguez determination of arborus rarus for the individual http://bigbotanicalgarden.org/individual/10224 jane jardin 2008-06-24 arborus rarus species l. rdf formatted description of the field individual http://bigbotanicalgarden.org/individual/10224 big botanical garden 1999-03-15t14:54:54-0800 en baskauf organization of occurrence-related biodiversity resources 44 2009-10-07t09:14:08-0600 2009-10-07t09:14:08-0600 --------------------------------------- for a functioning example of a guid for an individual organism, enter http://bioimages.vanderbilt.edu/ind-baskauf/66920 into a web or rdf browser. through content negotiation, it will resolve to either http://bioimages.vanderbilt.edu/ind-baskauf/66920.htm if content-type text/html is requested by a web browser, or http://bioimages.vanderbilt.edu/baskauf/66921.rdf if content-type application/rdf+xml is requested by an rdf browser. http://bioimages.vanderbilt.edu/ind-baskauf/66920 http://bioimages.vanderbilt.edu/ind-baskauf/66920.htm http://bioimages.vanderbilt.edu/baskauf/66921.rdf microsoft word sautteretalv2.doc biodiversity informatics, 3, 2006, pp. 46-58 46 a combining approach to find all taxon names (fat) in legacy biosystematics literature guido sautter1,3, klemens böhm1, and donat agosti2 1 department of computer science, universität karlsruhe (th), 76128 karlsruhe, germany; 2 division of invertebrate zoology, american museum of natural history, new york ny 100245192, and naturmuseum der burgergemeinde bern, 3005 bern switzerland; 3 sautter@ipd.uka.de abstract.— most of the literature on natural history is hidden in millions of pages stacked up in our libraries. various initiatives aim now at making these publications digitally accessible and searchable, applying xmlmark up technologies. the unique biological names play a crucial role to link content related to a particular taxon. thus discovering and marking them up is extremely important. since their manual extraction and markup is cumbersome and time-intensive, it needs be automated. in this paper, we present computational linguistics techniques and evaluate how they can help to extract taxonomic names automatically. we build on an existing approach for extraction of such names (koning et al. 2005) and combine it with several other learning techniques. we apply them to the texts sequentially so that each technique can use the results from the preceding ones. in particular, we use structural rules, dynamic lexica with fuzzy lookups, and word-level language recognition. we use legacy documents from different sources and times as test bed for our evaluation. the experimental results for our combining approach (fat) show greater than 99% precision and recall. they reveal the potential of computational linguistics techniques towards an automated markup of biosystematics publications. key words.— digital library, systematics, named entity recognition, taxonomic name extraction. the mass digitization of biosystematics literature is becoming a major issue (e.g., biodiversity heritage library, www.bhl.si.edu; american museum of natural history digital library; antbase.org). this body of literature with well over 10 million pages contains all the descriptions of the world’s biological taxa, that is the names and formal descriptions of the estimated 1.7 million species known today and their higher categories (maze 2004). the scientific names, latinized binomen composed of a generic and a specific name (iczn 2000: article 5.1), are important. this is because they are unique within animals, plants, bacteria, virus and fungi, and their applications are ruled by respective codes (e.g., international code of zoological nomenclature for animals). each of these names belongs in a spe cific position within the taxonomic hierarchy. within the life sciences, these scientific names are used to report the identity of the organisms upon which a study has been conducted. this potentially allows finding and linking all information on a particular species. thus, recognizing taxonomic names is highly relevant for the digitization process, since no complete list of all the names of living organisms exists yet. manual extraction of these names is timeconsuming, e.g., 80 hours of manual extraction versus 330 seconds automatic extraction (koning et al. 2005), and thus expensive. automated name recognition and extraction is the ultimate solution. this article describes a combining approach for taxonomic name extraction, i.e., it combines several existing techniques from machine learning etc. we have dubbed our approach fat, which is short for ‘finds all taxonomic names’. by reducing the average error of the base techniques by over 90%, our technique comes close to meeting the claim behind its name. name extraction (techniques) taxonomic names have some basic structural commonalities. the combination of its elements (see table 1) is not very restrictive and includes many optional parts and combinations. some of these are no longer used, such as quadrinomen, a variety of a subspecies of a species of a genus. but sautter et al. – find all taxon names in literature 47 nevertheless it is part of the history of names (see iczn 2000 for legalistic aspects). table 1: the parts of taxonomic names part example 1 example 2 genus prenolepis dolichoderus (subgenus) (nylanderia) species vividula decollatus (author) nylander (subspecie s) subsp. guatemalensis (author) forel (variety) var. itinerans (author) forel for example, both “prenolepis (nylanderia) vividula nylander subsp. guatemalensis forel var. itinerans forel“ and “dolichoderus decollatus“ are taxonomic names. there are only two mandatory parts in such a name: the genus and the species name. table 1 shows the deconstruction of the two examples. the parts with their names in brackets are optional. formally, the rules of the linnaean (binominal) nomenclature define the structure of taxonomic names as follows, exemplified using animal names: 1. the genus is mandatory. it is a capitalized word, often abbreviated by its first one or two letters, followed by a dot. in enumerations of several species of the same genus, the genus tends to appear explicitly only with the first species in the sequence. 2. the subgenus is optional. it is a capitalized word. in most cases, it is enclosed in brackets, but not always. 3. the species is mandatory. it is a lower case word, often followed by the name of the scientist who first described the species. 4. the subspecies is optional. it is a lower case word as well, preceded by an indicator word like subsp. or subspecies. it is often followed by the name of the scientist who first described the subspecies. in newer publications, the species is often abbreviated if a subspecies is given. in this case, the author name of the species is omitted. in addition, the indicator word can be omitted as well. 5. the variety is optional. it is a lower case word, preceded by an indicator word like var. or variety. it is often followed by the name of the scientist who first described it. since 1960, however, the indicator word var. or variety is not permitted anymore (iczn 2000). the main problem for the automated recognition of these names is to distinguish them from the surrounding text, including other named entities (ne). named entity recognition (ner) techniques can be employed to automatically identify scientific names (chieu & ng 2002). ner uses a variety of methods. most common are gazetteers, grammars, rules, and statistical methods like support vector machine (bikel et al. 1997; cuerzan & yarowsky 1999; mikheev et. al., 1999; isozaki et al. 2002; koning et al. 2005). tjong et al. (2003) introduce two typical ner tasks: the names of locations, persons, and organizations are to be extracted. one may perceive taxonomic names as a special case of ne. but their structure is more complex and more variable than the one of ‘typical’ ne, e.g., location names, despite some basic shared elements, such as a latin binomen surrounded by text in another language, the latin binomen often not being part of existing dictionaries. hence, common ner techniques tend to be too general to recognize taxonomic names. newer tasks like the one presented by carreras et al. (2005) do not consider more complex entities, but start dealing with relationships and semantic roles. therefore, we do not have the hope that general ner research will turn to the extraction of complex entity names in the near future. another problem of existing ner techniques is that they usually require pre-annotated training data (several hundred thousand words) to achieve good results (about 97 % precision and recall). – besides ner, the following techniques are used to extract taxonomic names. list-based ner techniques. palmer & day (1997) perform a lookup to determine whether a word is a ne of the category sought. the sole use of a thesaurus as a positive list is not an option for taxonomic names. all existing thesauri are incomplete. nevertheless, such a list allows recognizing known parts of taxonomic names. the inverse approach would be a list-based exclusion technique, e.g., a common english sautter et al. – find all taxon names in literature 48 thesaurus like wordnet serves as a list of known negatives. this in isolation is not an option either. it would not exclude proper names reliably. next, it would exclude parts of taxonomic names that also happen to be used in common english. this was the reason for the majority of errors in the evaluation of taxongrab (koning et al. 2005), which combines list-based exclusion with some rules. however, exclusion of sure negatives, i.e., words that are never part of taxonomic names, simplifies the classification process. rule-based techniques do not require any training data. instead, they try to find words or word sequences with a certain structure, e.g., regarding punctuation. yoshida et al. (1999) presents a technique that extracts the names of proteins and their abbreviations based on regular expressions. it makes use of the very distinctive syntax of protein names, e.g., “ng-monomethyl-l-arginine”. the syntax of taxonomic names is subject to certain rules as well, but they are less restrictive. due to the wide range of optional parts (see tab. 1), it is impossible to find a regular expression that matches all taxonomic names and at the same time provides a satisfactory degree of precision. koning et al. (2005) present an approach based on regular expressions and lexica. this technique (called taxongrab) performs satisfactorily compared to common ner approaches. but the conception of what is a positive is restricted. for instance, it simply leaves aside taxonomic names that do not specify a genus. however, the general idea of using rules to filter the phrases of documents is helpful. bootstrapping. jones et al. (1999) describe an approach to training classifiers without large amounts of labeled training data. some labeled seed data and a large unlabeled training corpus is taken as input. learning from the seed data yields automatic labeling of the corpus. jones et al. (1999) have shown that the performance of this approach is equal to the one of other techniques that require large amounts of labeled training data. bootstrapping is not readily applicable to our particular problem, however. niu et al. (2003) use an unlabeled corpus of 88,000,000 words to bootstrap a named entity recognizer. for our purpose, even unlabeled training data is not available in this order of magnitude, at least right now. active learning. the intention behind active learning (day et al. 1997) is to speed up the creation of large labeled training corpora from unlabeled documents. in particular, the system uses all of its knowledge during all phases of the processing. in this way, it can label many data items automatically, and the user has to label only pathologic cases. to increase data quality, such a user-interactive approach should be part of a taxonomic-name extractor as well. we make use of this approach in two ways: first, the output of each step serves as base data for the subsequent ones. second, the user manually classifies the few remaining cases after these automated steps. following the general idea of active learning, we feed these manual classifications back into the base data. the algorithm can then use them later when processing other documents. this improves the performance of the algorithm at runtime. in our evaluation, we will use a measure that quantifies the number of user interactions. this is to enable comparison to other components. word language recognition. language recognition is intended to determine the language a given text is written in. (sautter & böhm 2006) have shown that these techniques can be used to extract parts of taxonomic names from english text. in particular, modifications have been made to the standard techniques so that little training data is required and becomes applicable on word level. the technique is based on two statistics containing the n-gram distribution of taxonomic names and of common english. both statistics are built from examples from the respective languages. it applies active learning to reduce the need for annotated training data. the classifier is tunable towards precision or recall, as needed. in optimal configurations, both reach a level of 96%. this is the typical level of common up-to-date ner components. the active learning requires the user to classify about 3% of the words manually. although this is relatively low, compared to manually annotating an entire text, the absolute number of user interactions is still high. in addition, some training data is needed. thus, other techniques are used to (a) gather the required training examples and to (b) reduce the input to sautter et al. – find all taxon names in literature 49 this classifier as far as possible. in particular, it should be used only to deal with word sequences that cannot be labeled safely with the other techniques. gene and protein name extraction. the major focus of ner in biomedicine is the extraction of gene and protein names. tanabe & wilbur (2002) give a wide overview of the techniques used for this purpose. the most frequently used approaches are hidden markov models, lexicon lookups and structural rules. many of the techniques also include a part-of-speech tagger and use its output as additional evidence. however, there are significant differences between gene and protein names on the one hand and taxonomic names on the other hand: first the nomenclature rules for the latter are by far less restrictive and include a wide range of optional parts. for instance, they may include the names of the discoverer/author of a given part. second, there are parts of gene and protein names which are easy to distinguish from the surrounding text because of their structure. for the extraction of taxonomic names, we cannot rely on this type of evidence. consequently, the techniques for gene or protein name recognition are not feasible for the extraction of taxonomic names. an individual technique in isolation thus might not be sufficient for taxonomic name extraction. mikheev et al. (1999) have shown that a combining approach, i.e., one that integrates the results of several different techniques, is superior to the individual techniques for common ner. for this reason, we combine approaches for taxonomic name extraction. due to the active learning, the word-level language recognizer needs little training data. in addition, the manual effort induced by user interactions is high. thus, other techniques need be applied beforehand, for the following two reasons: first, to find sufficient training examples for the word-level classifier. second, to reduce the input to the classifier to as few words as possible. this last aspect is based on the idea to prevent as many words as possible from being prompted to the user to further reduce the manual effort. usage of the typical structure of taxonomic names allows achieving both goals. syntax-based rules are used to extract training examples from the documents. this leads to a reduction of the number of words the classifier has to deal with. however, it is not possible to find rules that extract taxonomic names with both high precision and recall, as we will show later. but we have found rules that fulfill one of these requirements very well. in what follows, we refer to these as precision rules and recall rules, respectively. material and methods the classification process the general idea of our approach is first to extract or exclude those parts of the text for which we can be sure about that they are either taxonomic names or not (precision and recall rules in fig. 1). we then use the parts already classified to build lexica and statistics, which we use to classify the rest of the text (data rules and word classifier in fig. 1). if there are still uncertain parts left after this step, we present them to the user for manual classification (user feedback in fig. 1). in more detail, our approach works as follows: i) in a first pass through the document, we apply the precision rules. every word sequence from the document that matches such a rule is a sure positive. ii) in a second pass, we apply the recall rules to the phrases that are not sure positives. a phrase not matching one of these rules is a sure negative. iii) third, we build lexica from the sure positives and sure negatives, and apply them in several ways to the phrases that are still uncertain. for instance, we filter out word sequences that contain at least one known negative word. iv) we collect a set of names from the set of sure positives. we then use these names to both include and exclude further word sequences. v) we train the word-level language recognizer with the surely positive and surely negative words. we then use the language recognizer to classify word/phrases that are still uncertain. figure 1 visualizes the classification process. red areas mark the flow of words/phrases which are uncertain at several stages, blue areas mark sure positives, while yellow areas mark sure negatives. the boxes with round corners represent sets of words/phrases, colored according to their state. initially, all words in the text (documents) are uncertain. after fat has finished, all sautter et al. – find all taxon names in literature 50 words/phrases are classified as sure positives (tax. names) or sure negatives (not tax. names). the gray boxes represent the different steps of the fat algorithm; the arrows depict the data flow. more specifically, the meaning of the arrows also depends on where an incoming arrow meets a box: an arrow meeting the box at its top represents data the step has to process. arrows going to the side of a box stand for words/phrases already classified and now serving as additional input. the state of the words/phrases after a step is visualized at the bottom of a box: + indicates that a word/phrase has matched the rule, indicates the opposite, and ? indicates that the particular step could not classify the word or phrase with certainty. the background color of the area behind the outgoing arrows also emphasizes this. an arrow that splits indicates that data goes two ways. as data comes out of the user feedback step and is finally classified, for instance, it goes into the sets of sure positives and sure negatives, respectively. additionally, the word classifier receives it as additional training data. two joining arrows signal that data comes from two sources. the training data for the word classifier, for example, comes from the sure positives and negatives as well as from the words/ phrases classified in the user feedback step. figure 1. the classification process the order of application enables the different techniques to profit from each other: the precision and recall rules extract the base data for the subsequent steps, so there is no need for training data at all. using the data rules and the word classifier before them would require manual preparation of lexica and training data for the classifier. so inverting our proposed order of application is not feasible. when processing a document, the fat algorithm does one pass applying the precision and recall rules. at the same time it collects the sure positives, candidates, and sure negatives. all further steps base on this initial trisection of the text. only the recall rule using the set of scientist names (see below) requires one further pass over the document. this approach is somewhat similar to the bootstrapping algorithm proposed by jones et al. (1999). the difference is that this process works solely with the document it actually processes. in particular, it does not need any external data or a training phase. the 107 documents forming the a.m.n. (american museum novitates) part of our test bed count about 8,100 words on average, which is less than 0.02% of the data used by niu et al. (2003). the entire test bed has less than 2,500,000 words, still less than 2% of the corpus used by niu et al. (2003). on the other hand, with the classification process proposed here, the accuracy of the underlying classifier must be very high from the start. this is because we do not have a training phase, but start from scratch with the first document we process. rules for structure of taxonomic names in order to make use of the structure of taxonomic names, we use rules that refer to this structure, see tab. 2. the syntax used here is the one of the java programming language, documented in the java online documentation (java 1.4.2). we use regular expressions for the formal representation of the rules. in this section, we develop a regular expression matching any word sequence that conforms to the linnaean rules of nomenclature (see 3.3). sautter et al. – find all taxon names in literature 51 table 2: abbreviations. _ one white space character [a-z](3,) [a-z](1,2). [a-z][a-z](2,) [a-z]{[a-z]}?. {_}(0,2) the taxonomic names are modeled as follows: 6. the genus part of a taxonomic name is a capitalized word, often abbreviated by its first one or two letters, followed by a dot. we denote it as , which stands for {|}. 7. the subgenus part of a taxonomic name is a capitalized word, optionally surrounded by brackets. we denote it as , which stands for |(). 8. the species part of a taxonomic name is the name of the species, a lower case word, optionally followed by a name. in newer publications, on the other hand, the species is often abbreviated if a subspecies is given. in this case, the name is omitted. we denote this structure as , which stands for {{_}?|}. 9. the subspecies part of a taxonomic name is a lower case word, preceded by the indicator word subsp. or subspecies, and optionally followed by a name. in newer publications, however, the subspecies is often abbreviated if a variety is given. in this case, the name is omitted. in addition, the indicator word subsp. or subspecies can be omitted as well. we denote this structure as , standing for {{{{subsp.|subspecies}_}?{_ }?}| }. 10. the variety part of a taxonomic name is a lower case word, preceded by the indicator word var. or variety, and optionally followed by a name. in newer publications, however, the indicator word var. or variety can be omitted. we denote this structure it as , which stands for {{{var.|variety}_}? {_}?}. a taxonomic name is now modeled as follows. we refer to the pattern as : {_}? _{_}? {_}? precision rules because matches any sequence of words that conforms to the linnaean rules, it is not very precise. the simplest match is a capitalized word followed by one or more in lower case. any two words at the beginning of a sentence are a match! thus, to have less false positives, we need more precise regular expressions. to accomplish this, we rely on the optional parts of taxonomic names. in particular, we classify a sequence of words as a sure positive if it contains at least one of the optional parts , and . for the last two, we additionally demand the subspecies or variety to be explicitly labeled or the part before them to be abbreviated. the second restriction is as secure as the first one. the reason is that normal text rarely continues in lower case after a dot. the hope is to exclude almost all phrases that are not taxonomic names. even though our regular expressions may classify a sequence of words as a sure positive erroneously, our evaluation will show that this happens very rarely. our set of precise regular expressions has three elements: 11. with subgenus in brackets, and optional: _() _{_}? {_}? 12. with given, and optional: {_}? __ {_}? 13. with mandatory, and optional: {_}? _{_}? {_} to classify a word sequence as a sure positive if it matches at least one of these regular expressions, we combine them disjunctively and call the result . it matches any sequence of sautter et al. – find all taxon names in literature 52 words we can classify as a taxonomic name simply because of its structure. when applying the precision rules, we test phrases of up to 10 words, plus punctuation. in many taxonomic publications, new genera, species, etc. are explicitly labeled. if dolichoderus decollatus is described for the first time, for instance, it is likely to labeled as a new species somewhere. the title of the description would be dolichoderus decollatus, new species. we use special forms of the precision rules to make use of these labels. in particular, we consider a match of a sure positive if it directly precedes a label in the text. because they rely on explicit labels, we refer to these special precision rules as label rules. a notion related to that of a sure positive is the one of a surely positive word. a surely positive word is a part of a taxonomic name that is not part of a scientist’s name. for instance, the taxonomic name prenolepis (nylanderia) vividula erin subsp. guatemalensis forel var. itinerans forel contains the surely positive words prenolepis, nylanderia, vividula, guatemalensis, and itinerans. further steps of our process assume that surely positive words exclusively appear as parts of taxonomic names. recall rules the recall rules basically consist of , which matches any sequence of words that conforms to the linnaean rules. when applying it, we again test phrases of up to 10 words, plus punctuation. but there is a further issue: enumerations of several species of the same genus tend to contain the genus only once. for instance, in pseudomyrma (minimyrma) arboris-sanctae emery, latinoda mayr and tachigalide forel we want to extract latinoda mayr and tachigalide forel as well. to address this, we make use of the surely positive words: we use them to extract parts of taxonomic names that lack the genus. we also extract the names of the scientists from the sure positives and collect them in an extra list (name lexicon). we regard a capitalized word in a sure positive as a name if it comes after the second position. in the example, we would extract pseudomyrma, minimyrma and arboris-sanctae from the sure positive pseudomyrma (minimyrma) arboris-sanctae emery. we would also add emery to the set of names. we cannot be sure that the list of sure positive words suffices to find all species names in an enumeration. hence, we additionally collect all lower-case words followed by a capitalized word contained in the set of names. in the example, we need to have mayr and forel in the set of names to extract latinoda mayr and tachigalide forel. data rules because we want to achieve close to 100% in recall, the recall rules are minimally restrictive. consequently, many word sequences that are not taxonomic names are considered uncertain. before the word-level language recognizer deals with them, we explore some more ways to find negatives. because the precision rules are very restrictive, they match only a fraction of the taxonomic names in a text. making use of the sure positive words, we can also find additional sure positives. sure negatives. as previously mentioned, matches any capitalized word followed by a word in lower case. this includes the start of any sentence. but making use of the sure negatives, we can recognize these phrases. in particular, we classify any word sequence as negative that contains a word which is also in the set of sure negatives. for instance, in sentence “additional evidence results from …”, “additional evidence” matches . another sentence contains “… an additional advantage …”, which does not match . thus, the set of sure negatives contains “an”, “additional”, and “advantage”. knowing that “additional” is a sure negative, we exclude the phrase “additional evidence”. names of scientists. though the names of scientists are valid parts of taxonomic names, they may also cause false matches. a misclassification occurs when they are matched with the genus or subgenus part – cannot exclude this. in addition, they might appear elsewhere in the text without belonging to a taxonomic name. similarly to sure negatives, we exclude a word sequence if the first or second word is contained in the set of names. for instance, in “…, and forel further concludes …”, “forel further” matches . if the set of names contains “forel”, we can exclude “forel further”. this is because we sautter et al. – find all taxon names in literature 53 know that “forel” is not the name of a taxonomic genus. sure positives. making use of the sure positives we have extracted with the precision rules, we can find additional sure positives. in particular, we mark an uncertain word sequence as a sure positive if it consists of surely positive words or abbreviations. if the precision rules have extracted prenolepis (nylanderia) vividula erin subsp. guatemalensis forel var. itinerans forel, for instance, we conclude that prenolepis vividula is a sure positive as well. stemming lookup catch. koning et al. (2005) used a common english dictionary to exclude negatives. as mentioned in the introduction, this leads to the exclusion of taxonomic names containing a common english word. for instance, this would exclude the taxonomic name formica minor because of minor. on the other hand, such a dictionary-based exclusion can help to catch the (very few, but still existing) erroneous matches of the regular expressions in . our observation shows that if a common english word is part of a taxonomic name, it is always used in its base form. thus, if we find a stemmed form of a word in a dictionary, we conclude that it is not part of a taxonomic name. consider the following sentence from an essay on dangerous insects: … in chaguanas (trinidad) another subspecies poisoned forel. except for the word in, this sentence matches the regular expression from where is mandatory. but we can recognize it as a false match because of the conjugated verb poisoned in the subspecies position. similar pathologic cases can occur for the variety part. but all these cases share a useful feature: they comprise modified forms of common english words. thus, to exclude these errors, we combine the dictionary lookup with stemming: for all the words matched to a lower case part of a regular expression in (, or ), we check if it could be a conjugated verb given its ending (most common endings are s, -ed, and -ing). if so, we apply stemming in order to obtain the infinitive. if contained in a common english dictionary, we exclude the match. in the example, the ending rule applies to poisoned. porter’s (1980) stemming algorithm produces poison as the word stem. because this word is contained in the dictionary, we can exclude the erroneous match. in in chaguanas (trinidad) another subspecies poisoned forel the stemming lookup catch is the only data rule that we also apply to the matches of the regular expression from . name completion making use of the scientists’ names, we also extract taxonomic names that lack the genus, e.g., from enumerations, such as pheidole pallidula, orbula, xantra. in addition, the rules allow genus abbreviations like ph. for pheidole in ph. cornutula. in order to determine the meaning of a taxonomic name, we need to complete the names with their full parts. if the genus part is missing, we have two options: first, we check if the species part appears elsewhere in the document, together with the genus it belongs to. if this is not the case, we use the last genus that we have extracted before the position of the name to complete. this is useful especially in case of enumerations: if several species of the same genus are enumerated, the genus is often given only with the first one. we then transfer the genus part to the subsequent taxon names. if the genus is abbreviated, we also have two options: first, we again check if the species part appears elsewhere in the document, together with the full name of the genus it belongs to. if this fails, we check if we have recognized any genus name that starts with the given abbreviation. if there is exactly one such genus name, we insert it. if there is more than one, i.e., the abbreviation is ambiguous, we use the one which appears closest before the abbreviation. classification of remaining words after applying the various rules, uncertain word sequences still remain. to deal with them, we use word-level language recognition (sautter & böhm 2006), a technique to classify words as parts of taxonomic names or as common english, respectively. it is based on two statistics containing the n-gram distribution of taxonomic names and of common english, that is how often short sequences of letters occur in a group of words in a text (e.g., 4-gram for formica = {form, ormi, rmic, mica}). both statistics are built from sautter et al. – find all taxon names in literature 54 examples from the respective languages. this technique achieves about 96% in precision and recall. it involves the user in the classification process to back up narrow decisions with the human expert knowledge. to train the classifier, we use the surely positive and surely negative words as the training data. instead of classifying every word separately, we compute the word-level classification score of all words of a sequence and then classify the sequence as a whole. this has several advantages: first, if one word of a sequence is uncertain, this does not automatically incur a user-feedback request. second, if a sequence of words is uncertain as a whole, the user gives feedback for the entire sequence. this results in several surely classified uncertain words at the cost of only one feedback request. in addition, a user can easier determine the meaning of a sequence of words than the one of a single word. experimental setup and test bed we run two series of experiments: we first process each document individually. we then process the documents incrementally, i.e., we do not clear the sets of known positives and negatives after each document. the same is true for the statistics of the word-level language recognizer. this is to measure the benefit of reusing data obtained from one document in the processing of subsequent ones. finally, we take a closer look at the effects of the individual steps and heuristics. the platform is completely implemented in java 1.4.2, and we have used the java.util.regex package to represent the rules. all tests presented here are based on three groups of annotated documents. first, we use 107 issues of the american museum novitates (a.m.n.), a natural science periodical published by the american museum of natural history. the second group is a recent publication representing a widely used standard in ant systematics (f.2000, fisher 2000), and the third one is the birds of congo (c.1932: chapin 1932, 1939, 1953, 1954), partitioned into four parts of similar size, which was used by koning et al, 2005. koning et al.’s test bed has been extended to include additional groups of publications with different ways of combining and abbreviating names. table 3 contains the relevant numbers on our test bed (rounded), the numbers on the a.m.n. and f.2000 parts are the result of manual counting, those on c.1932 originate from (koning et al. 2005). table 3: the test bed a. m. n. f. 2000 c. 1932 words 857,000 58,000 1,100,000 taxonomic names 12,000 175 21,000 evaluation measure in nlp, the f-measure is popular to quantify the performance of a word classifier, but we also need to measure the advantage the system gains from asking the user for feedback on narrow classifications. in particular, we use three measures to quantify our test results. as mentioned, the first one is the f-measure: • p(p) := positives classified as positive n(p) := positives classified as negative p(n) := negatives classified as positive n(n) := negatives classified as negative rp rp2 :fmeasure n(p) p(p) p(p) :r callre p(n)p(p) p(p) :p ecisionpr + ×× = + = + = but our combined technique has three possible outputs. if the decision between positive or negative is narrow, a word is classified as uncertain, and the user is prompted. this prevents misclassifications and thus induces a considerable advantage over fully automated techniques. in order to enable comparison to fully automated techniques, we use two further measures: • u(p) := positives not classified (uncertain) u(n) := negatives not classified (uncertain) given this, coverage c is defined as the fraction of all classifications that are not uncertain: )n(u)n(n)n(p)p(u)p(n)p(p )n(n)n(p)p(n)p(p :c +++++ +++ = to combine these two measures to a single measure for overall classification quality, we multiply f-measure and coverage and define quality q as cfmeasure:q ×= sautter et al. – find all taxon names in literature 55 this measure treats all uncertain votes as misclassifications, thus punishing every user interaction as if it was an error. this enables comparison to techniques that do not involve the user. it is very restrictive because a random guess might result in at least half of the uncertain words classified correctly. on the other hand, a correct vote from the user avoids misclassifications. in learning components, this keeps statistics clean because no errors are fed back. evaluation and discussion a combining approach gives rise to many questions in the context of taxonomic name extraction, e.g.: how does a word-level classifier perform with training data automatically generated? how does rule-based filtering affect precision, recall, and coverage? what is the effect of extending the lexica dynamically? which kinds of errors remain? tests with individual documents first, we test the combined classifier with each document individually. we omit fisher (2000) here because it consists of only one document. the results for this part of the test bed will be presented in the next section. table 4 contains the average results for the a.m.n. and c.1932. the combination of rules, dynamic lexica, and wordlevel classification provides very high precision and recall. the former is 99.7% on average, the latter 98.2%. the need for manual intervention is very low: the average coverage is 99.7%. the average results with individual issues of the a.m.n. (column doc in table 4) are significantly worse than with the other two parts of the test bed. a more detailed look at the results reveals significant differences between the individual documents. for more than half of the documents, the results are equal to those of the other two parts of our test bed. the rest of the a.m.n. issues points out a weakness of our combined technique: if the precision rules do not extract a sufficient number of sure positives, we run into two problems. we neither have enough data to successfully apply the data rules, nor do we have enough positive examples to train the word level classifier. the sum column in table 4 contains the summed up results for the a.m.n. it turns out that precision, recall and coverage are far better for the total numbers than the average per document. this is because the documents with few taxonomic names in them produce the poor results, while our combined technique performs better with the bigger documents. for c.1932, this effect is almost non-existent. the individual parts are big enough and contain sufficiently many sure positives for our technique to succeed. table 4: test with individual documents a. m. n. c.1932 doc sum doc sum words 857,000 1,100,000 taxonomic names 12,600 22,500 sure pos. 24 2,528 2,833 11,331 uncertain 356 38,177 11,545 46,179 data rules sp 85 9148 4983 19933 data rules uc 42 4475 674 2,697 scorings 17 1,836 181 723 precision 93.1% 97.5% 92.7% 99.7% recall 55.2% 93.3% 97.8% 99.8% f-measure 56.8% 95.4% 95.0% 99.7% coverage 87.8% 99.0% 96.6% 99.8% quality 54.4% 94.4% 91.9% 99.6% tests with entire corpus in the first test the classifier did not transfer any experience from one document to later ones. we now process the documents one after another, de facto concatenating all the documents to one big super-document, which is then analyzed as a whole. table 5 shows the results. as expected, the classifier performs better than with individual documents. this is true for both the a.m.n. and c.1932 test documents. the average recall increases to 99.9%, coverage improves to 99.8% on average. precision increases to an average of 99.5%. the effect of the incremental learning is obvious, especially for the a.m.n. part of the test bed: the false positives are less than 2% of those in the first test shown by a comparison of the recall values in tables 4 and 5. the effect on precision is sautter et al. – find all taxon names in literature 56 significant as well: the number of false negatives is only a third of that in the first test. finally, the number of words for which the technique has to ask for feedback is halved (compare coverage values). table 5: test with corpora a. m. n. f. 2000 c. 1932 words 857,000 58,000 1,100,000 taxonomic names 12,600 175 22,500 sure pos. 3,059 172 13,827 uncertain 37,028 2,368 42,180 data rules sp 11,819 175 2,2084 data rules uc 1,132 2 583 scorings 618 1 295 precision 99.2% 100% 99.8% recall 99.9% 100% 99.9% f-measure 99.6% 100% 99.8% coverage 99.5% 100% 99.9% quality 99.1% 100% 99.7% the reason for the improvement is obvious from documents where the number of word sequences in is low: data from other documents compensates the lack of positive examples. this reduces the number of false positives and false negatives as well as the user interactions. the data rules the lines uncertain in tables 4 and 5 contain the number of uncertain phrases after the application of the regular expressions, the lines data rules uc display how many phrases remain uncertain after the data rules were applied. the exclusion of word sequences containing a sure negative turns out to be effective to filter the matches of . on average, this step reduces the number of uncertain word sequences by about 75%. the lines sure pos. and data rules sp., in turn, provide the number of sure positives after the regular expressions and after the data rules, respectively. the data rules based on the sure positives are very effective as well: they reduce the uncertain word sequences by another 15 %, and at the same time enlarge the set of sure positives by 50% on average. in particular, they do not only reduce the uncertain sequences, but also obtain additional training data for the word level classifier. the lines labeled scorings display the number of distinct phrases that were classified by the statistical component. our experiments show that the manual effort incurred by uncertain statistical classifications and subsequent user feedback decreases significantly. all four data rules decrease the number of words the language recognizer has to deal with. this is because they produce additional training data and reduce the number of words classified as uncertain. comparison to word-level classifier and taxongrab a word-level classifier (wlc) is the core component of fat. we compare it in standalone use to the combining technique (comb) and to the taxongrab (t-grab) approach (koning et al., 2005), which is based on a set of regular expressions and lists. the results for taxongrab were obtained from the c.1932 part of our test set used for this evaluation (see tab. 6). fat is superior to both taxongrab and standalone wordlevel classification. the improved precision and recall results are due to the usage of greater variety of evidence. the better coverage results from the lower number of words that the word-level classifier has to deal with. on average, it has to classify only 0.1% of the words in a document. this also significantly reduces the user feedback and the number of potential errors of the wordlevel classifier. table 6: comparison to related approaches precisi on recall fmeasur e covera ge t-grab 96% 94% 95% wlc 97% 95% 96% 95% comb 99.7% 99.9% 99.8% 99.8% all these positive effects result in about 99.8% fmeasure and 99.8% coverage. this means the error sautter et al. – find all taxon names in literature 57 is reduced by 90% compared to word-level classification, and by 93% compared to taxongrab. misclassifications although fat achieves very high performance, some errors remain. in this section, we take a closer look at the latter and discuss how we can prevent them in the future. false negatives. false negatives can occur in the language recognition step. most of them contain two words the first of which is a genus. xenomyrmex varies … , for instance, could induce such an error: the word level classifier (correctly) recognizes the first word as a part of a taxonomic name. the second word is not typical enough to change the overall classification of the sequence. to avoid this type of false negatives, one might use pos-tagging, which would label varies as a verb. we could exclude word sequences containing words with a meaning that cannot occur in taxonomic names. a related problem results from literature references in arbitrary languages. chapin (1932, 1939, 1953, 1954), for instance, cites systema naturae (linnaeus 1758), a book written in latin. the word level classifier correctly recognizes the language as latin. the problem is that our assumption that the taxonomic names are the only parts of the text in latin does not always hold. other publications cite complete paragraphs from documents written in italian. sulla posizione sistematica is an excerpt from such a paragraph that happens to match . because italian is closer to latin than to english, the word level classifier recognizes this phrase as a taxonomic name. german citations raise yet another problem: because the capitalization rules of this language differ from the english ones, the regular expressions happen to match such text parts. in particular, all nouns are capitalized in german. this lets them match the genus and subgenus part of our regular expressions. false positives. the regular expression matches any word sequence that is a taxonomic name. but the subsequent exclusion mechanisms may misclassify a sequence of words. in particular, the word-level classifier does not always recognize taxonomic names if they have been formed from proper names of persons. this is because these words consist of n-grams that are typical for common english. wheeleria rogersi smith, for instance, is a fictitious but valid taxonomic name. to overcome this problem, we can construct these genera and species from the names we have extracted from the sure positives. conclusions this paper has shown how combined computer linguistic techniques can be applied to automatically extract taxonomic names from english text documents. fat yields a precision of up to 99.7% as opposed to taxongrab (96%). a promising future avenue is to study those names which have not been detected, and to start to integrate other languages. this is in fact necessary, since a large part of the heritage literature is written in languages different from english. acknowledgments the authors thank the members of the project team (national science foundation iis-0241229, deutsche forschungsgemeinschaft bib47) for their comments, and colleagues from the marine biological laboratory at woods hole, massachusetts for their discussions regarding issues of name structure and extraction. references bikel, d. m., s. miller, r. schwartz, and r. weischedel. 1997. nymble: a high-performance learning name-finder. proceedings of anlp-97, washington, usa. carreras, x. and l. marquez. 2005. introduction to the conll-2005 shared task: semantic role labeling. chapin, j. 1932. the birds of the belgian congo: part 1. bulletin of the american museum of natural history: 65. american museum of natural history, new york. chapin, j. 1939. the birds of the belgian congo: part 2. bulletin of the american museum of natural history: 75. american museum of natural history, new york. chapin, j. 1953. the birds of the belgian congo: part 3. bulletin of the american museum of natural history: 75a. american museum of natural history, new york. chapin, j. 1954. the birds of the belgian congo: part 4. bulletin of the american museum of natural history: 75b. american museum of natural history, new york. chieu, h. l., and h. t. ng. 2002. named entity recognition: a maximum entropy approach using sautter et al. – find all taxon names in literature 58 global information. proceedings of coling-02, taipei, taiwan. cucerzan, s., and d. yarowsky. 1999. language independent named entity recognition combining morphological and contextual evidence. in proceedings of sigdat-99, college park, usa.. day, d., j. aberdeen, l. hirschman, r. kozierok, p. robinson and m. vilain. 1997. mixed-initiative development of language processing systems. proceedings of the 5th conference on applied natural language processing. fisher, b. 2000. the malagasy fauna of strumigenys. pp. 612–710. in bolton, b., the ant tribe dacetini. memoirs of the american entomological insititute, 65 (2): 1-1028. isozaki, h., and h. kazawa. 2002. efficient support vector classifiers for named entity recognition. proceedings of coling-02, taipei, taiwan jones, r., a. mccallum, k. nigam, and e. riloff. 1999. bootstrapping for text learning tasks. proceedings of ijcai-99 workshop on text mining. iczn. 2000. international code of zoological nomenclature, fourth edition. international commission of zoological nomenclature, london. java 1.4.2: java online documentation, http://java.sun.com/j2se/1.4.2/docs/api/java/util/regex/ pattern.html. koning, d., n. sarkar, and t. moritz. 2005. taxongrab: extracting taxonomic names from text. biodiversity informatics 2: 79-82. linnaeus, c. 1758. systema naturae. ipsiae, lund. maze, g. 2004. the role of taxonomy in species conservation. philosophical transaction of the royal society london b 359: 711–719. mikheev, a., m. moens, and c. grover. 1999. named entity recognition without gazetteers. proceedings of eacl-99, bergen, norway. niu, ch., w. li, j. ding, and r. k. srihari. 2003. a bootstrapping approach to named entity classification using successive learners. proceedings of 41st annual meeting of the association for computational linguistics. palmer, d. d., and d. s. day. 1997. a statistical profile of the named entity task. proceedings of anlp-97, washington, usa.. porter, m.f. 1980. an algorithm for suffix stripping. program, 14: 130-137 sautter, g., and k. böhm. 2006. how helpful is wordlevel language recognition to extract taxonomic names? technical report, http://www.ipd.unikarlsruhe.de/~sautter/taxonomicnameextraction.pdf. tanabe, l., and w. j. wilbur. 2002. tagging gene and protein names in biomedical text, bioinformatics 18:1124-1132. tjong, e. f., k. sang, and f. de meulder. 2003. introduction to the conll-2003 shared task: language-independent named entity recognition, edmonton, canada, 2003. wordnet. a lexical database for the english language, http://wordnet.princeton.edu/. yoshida, m., k.-i. fukada, and t. takagi. 1999. pdad-css: a workbench for constructing a protein name abbreviation dictionary. proceedings of the 32nd hawaii international conference on system science biodiversity informatics, 9, 2014, pp. 1–12 character selection during interactive taxonomic identification: “best characters” *nadia talent1, richard b. dickinson1, timothy a. dickinson1,2 1department of natural history, royal ontario museum, 100 queen’s park, toronto, m5s 2c6, canada 2department of ecology and evolutionary biology, university of toronto, 25 willcocks street, toronto, m5s 3b2, canada corresponding author: nadia talent, nadia.talent@utoronto.ca abstract— software interfaces for interactive multiple-entry taxonomic identification (polyclaves) sometimes provide a “best character” or “separation” coefficient, to guide the user to choose a character that could most effectively reduce the number of identification steps required. the coefficient could be particularly helpful when difficult or expensive tasks are needed for forensic identification, and in very large databases, uses that appear likely to increase in importance. several current systems also provide tools to develop taxonomies or single-entry identification keys, with a variety of coefficients that are appropriate to that purpose. for the identification task, however, information theory neatly applies, and provides the most appropriate coefficient. to our knowledge, delta-intkey is the only currently available system that uses a coefficient related to information theory, and it is currently being reimplemented, which may allow for improvement. we describe two improvements to the algorithm used by delta-intkey. the first improves transparency as the number of remaining taxa decreases, by normalizing the range of the coefficient to [0,1]. the second concerns numeric ranges, which require consistent treatment of sub-intervals and their endpoints. a stand-alone bestchar program for categorical data is provided, in the python, r, and java languages. the source code is freely available and dedicated to the public domain. key words— separation coefficient, polyclave, multi-access key, entropy, delta-intkey, information theory in biodiversity informatics, one type of automated tool for taxon identification is the polyclave, also called a multiple-entry key or multi-access key, an information retrieval system (duke 1969) that sequentially accepts specifications of character-states observed, in any order (morse 1975; duncan and meacham 1986; pankhurst 1991), and “eliminate[s] taxa which disagree with the specimen to be identified” (pankhurst and aitchison 1975). some confusion surrounds these and related terms (gower 1975; hagedorn et al. 2010). a different type of system that is frequently said to “identify” (pankhurst 1975; macleod 2007) establishes groupings of known and/or unknown individuals by using similarity or dissimilarity measures; thus, it establishes the taxa that form the basis of the dataset used by a polyclave. this is “clustering”, or in the terminology of sneath and sokal (1973, p. 3), “classification”, rather than “identification”. sneath and sokal, whose terminology we follow, define identification (p.3) as “the allocation or assignment of additional unidentified objects to the correct class once a classification has been established.” software tools for classification are sometimes packaged together with tools for building polyclaves (e.g., eti bioinformatics undated). the long-standing tradition for printed keys, since de lamarck’s development of these tools (de lamarck 1778), is to define a key as “an artificial analytical device or arrangement whereby a choice is provided between two contradictory propositions” (voss 1952), thus requiring binary characters. the character set and character states are usually carefully chosen and limited in number to save space. considerable effort is required to choose the characters and build such keys (hagedorn et al. 2010). the characters are often chosen for ease of use, or else to form “diagnostic descriptions” (“irredundant character sets”) that concisely differentiate every taxon (reviewed by 1 mailto:nadia.talent@utoronto.ca talent, dickinson, and dickinson — best characters payne and preece 1980). usually, the characters chosen do not vary within a taxon. it is common, but unnecessary, to be consistent with a natural taxonomy, and it may be undesirable to limit the character selection in that way (voss 1952). a polyclave, however, may hold quite different data, possibly with multi-state characters, using characters that vary within a taxon, and including states coded as “missing data” for all taxa. practice has shown that polytomies in printed identification keys can lead to considerable confusion (voss 1952), and they are almost universally avoided in that context (hagedorn et al. 2010). many of the identification keys that are being made available online also either use binary characters only (toda et al. 2004; guala 2004–12; christensen 1999; rosatti undated) and/or are automated versions of traditional keys (martellos 2010) called “pathway keys” by walter and winterton (2006). software is available, however, for implementing general polyclaves that allow multi-state characters and an unconstrained sequence of character selection (dallwitz 1993; brach and song 2005; alexander 2006; lucidecentral.org 2010; ung et al. 2010). a polyclave can be difficult to use correctly, such as by novices who make frequent errors, or for polythetic taxa which are distinguished by possessing a majority of a set of character states, rather than by required states (morse 1971, 1975). these problems can be treated with polythetic polyclaves that provide an “error-tolerance” or “mismatch threshold” setting, and by weighting character states and taxa (specifying which character states or taxa are rare). here we consider those features to be secondary additions to a simple interface that allows character and character-state selection in any order. we concentrate on an efficient error-free identification process in a monothetic key. in both printed keys (walters 1975; lobanov et al. 1981) and polyclaves, permitting multi-way choices (polytomies, multi-state characters) can gain considerable efficiency in the sense that the expected number of choices required to arrive at an identification is greatly reduced (cover and thomas 2006). ideal characters for speeding up the identification are those that separate the taxa evenly into small groups (osborne 1963; lobanov et al. 1981; cover and thomas 2006). quite complex characters with overlapping states (including numeric ranges; dallwitz et al. 2002) can have good separating power. some of the software that provides a polyclave interface calculates a coefficient to rank the characters by their separating power, to guide the user to an efficient identification of an unknown specimen. the coefficients have names such as “best character” (dallwitz et al. 1998/2012), “separation coefficient” (eti bioinformatics undated), “dichotomizing value” (morse, 1971), or “best” (brach and song 2005; pankhurst, 1978; lucidecentral.org 2010). other software provides a choice of coefficients (ung et al. 2010), such as variants of a jaccard coefficient, or simple matching coefficient (also called “sokal & michener coefficient”, sneath and sokal 1973). those ranking coefficients are different from, and complementary to, a “differentiating characters” list (e.g., as provided in meka; duncan and meacham 1986; duncan and meacham 1987; christensen 1999; rosatti undated), which lists those (binary) character states that uniquely identify a single taxon. some systems use the same coefficients both for assessing a proposed taxonomy and for designing identification keys (eti bioinformatics undated). it should be emphasized that the coefficients used in polyclaves and their terminology are related to, but are used in a very much simpler setting than, various techniques in multivariate data exploration. the many techniques grouped under headings such as “character ranking” (podani 2000) or “feature selection” or “weighting” (dale et al., 1986), evaluate the contribution of characters to the most significant patterns present, as a preparation for reducing the dimensionality of the data. with those methods, characters are considered all together or in pairs, whereas for a polyclave the characters are assessed individually, producing a single numeric value for each. pure polyclave software aims only to provide a hint for the user, a set of numeric values that indicate how well each character differentiates the remaining taxa. although different types of identification software and embellishments to 2 talent, dickinson, and dickinson — best characters polyclave software have been constructed that take taxon probabilities and character-state probabilities into account, expressed as weighting schemes and conditional characters (called “controlling” and “dependant” characters in delta; dallwitz et al. 2010), those probabilities often cannot be known (reviewed by pankhurst, 1978). the user working with a particular specimen is often the only meaningful source of an assessment of a character’s usefulness, because it depends on whether it is feasible to assess the character state at all. the usefulness of a character ranking probably amounts only to indicating which characters are very good and which very poor at differentiating; nonetheless, correct calculation is needed to align the characters appropriately. here we review the calculation of the character-ranking coefficient derived from information theory and recommend its use for polyclaves. it is similar to but not the same as delta-intkey’s “best character” function (spooner and chapman 2007; dallwitz 2010). statements made here about the best character implementation in delta-intkey are based on our reverse engineering. dallwitz (1974) gives a formula that is similar to delta-intkey’s behaviour if one interprets “n sub j, the number of taxa in the jth subgroup” as a proportion.) informationtheoretic character ranking can be counter-intuitive for certain types of characters, and has never, to our knowledge, been fully implemented for taxonomic applications. choice of coefficient formulae for character-ranking coefficients fall into two types: (i) a separation number or separation coefficient reflects the number of groups that can be distinguished by using the character, whereas (ii) the information statistic or entropy h reflects both the number of groups and the number of taxa in each, that is, the evenness of the division (cover and thomas 2006). the separation number is a count of the number of pairs of taxa that can be distinguished using the character (table 1). the separation coefficient is the separation number divided by the number of possible pairs, so it is normalized to the table 1: various separation measures for characters in a simple constructed key. ‘/’ in a list of character states indicates ‘or’; n is the total number of states; bold-faced values are less informative than the corresponding values of some other coefficients. anther color petal color stamen # style # taxon 1 red red 10 2 taxon 2 pink pink/white 10 5 taxon 3 white white 20 5 taxon 4 yellow yellow 20 5 maximum value separation number 6 5 4 3 n(n−1)/2 separation coefficient 1 0.83 0.67 0.50 1 simple matching (xper2, observed) 1 1 0.67 0.50 n/a jaccard (xper2, observed) 1 1 0.67 0.50 n/a pairwise average jaccard distance 1 0.92 0.67 0.50 1 information (in bits) 2 1.63 1 0.81 log2 n normalized to the range [0,1] for 4 taxa 1 0.81 0.50 0.41 1 3 talent, dickinson, and dickinson — best characters range [0,1]. if the separation coefficient takes the value 1, this signals that the answer to a single question about the state of that character would be enough to distinguish every taxon, a feature that we believe could be helpful to the polyclave user, particularly when a large number of possible identities remain. most authors who have used this approach have excluded any taxa that overlap, i.e., that share some but not all character states (e.g., table 1, petal color = white exhibits overlap). morse (1971) extended the separation number to count the extent of overlap, but only for binary characters. the simple matching (“sokal & michener”) coefficient, the ratio of character-state matches to character-state pairs, and the average pairwise jaccard distance (which excludes “negative” , also called “absence” matches, i.e., shared notapplicable states) are similar, relatively simple, calculations. variations on these coefficients have also been advocated (e.g., the jaccard coefficient of the xper2 system of ung et al. 2010). if overlapping states are fully incorporated, then the coefficient becomes the same as the information content. when designing a pathway key, stateoverlap is best avoided if other characters are available that cleanly distinguish the taxa, and a separation coefficient shows those characters to advantage while the information content does not. the information coefficient shannon’s information theory (shannon 1948) has long been used for problems similar to the taxon-identification problem (e.g., shwayder 1971, 1974), and is often cited as an appropriate foundation for ranking the characters in a polyclave (pankhurst 1991). although it has been stated that “the coefficient is not defined for continuously-varying numerical data” (lance and williams 1966), this is not relevant for polyclaves because the identification database will treat numeric ranges in most respects like other discrete categories (see below). if an event e occurs with probability p(e), and we are told that e has occurred, then we have received: 𝐼(𝑒) = log 1 𝑝(𝑒) = − log𝑝(𝑒) (equation 1) units of information (abramson 1963). for example, if one of two equally likely events is specified, then one bit of information is obtained. the information content is a lower bound on the number of yes/no questions that will lead to each of the possible identifications (cover and thomas 2006 chapter 5). for a single question, equivalent to each individual character in a polyclave, the mutual information i(x;y) is “the reduction in uncertainty of x due to the knowledge of y”, where x is the taxon identity and y is the taxonomic character: 𝐼(𝑋;𝑌) = �𝑝(𝑥,𝑦)log 𝑝(𝑥,𝑦) 𝑝(𝑥)𝑝(𝑦) 𝑥,𝑦 (equation 2) where x ∈ x and y ∈ y are the possible values of random variables x and y entropy (h) “is the minimal descriptive complexity of a random variable”, and “mutual information is the relative entropy between the joint distribution and the product distribution” (cover and thomas 2006). for the example in table 2 see figure 1. table 2: taxa and character states can be specified as a joint probability distribution, by assuming that taxa are equiprobable, and that alternative states are equiprobable for a taxon. ‘/’ in a list of character states indicates ‘or’. flower color white orange pink red taxon 1 white taxon 1 1/3 0 0 0 1/3 taxon 2 orange/pink taxon 2 0 1/6 1/6 0 1/3 taxon 3 pink/red taxon 3 0 0 1/6 1/6 1/3 1/3 1/6 1/3 1/6 4 talent, dickinson, and dickinson — best characters a succinct interpretation of how the formulae apply to a polyclave is given by pankhurst (1991). if p1, p2, p3, … , pm are the proportions of states 1, 2, 3, … , m for a group of taxa, then “the ‘information’ that we get from seeing state 1 on a specimen is the effect this fact has on our opinion of what taxon we think we have. if p1 was 1 (i.e. all taxa show character state 1 only), then on seeing state 1 we would gain no information at all.” normalizing the coefficient in equations 1 and 2, the choice of the base for the logarithm is arbitrary, and the units of information are called bits if logarithms to base 2 are used, hartleys (abramson 1963) with base 10, and nats (cover and thomas 2006) (or nits (macdonald 1952)) if natural logarithms are used. for the example in table 2: 𝐼(𝑋;𝑌) = 2 3 log(3) + 1 3 log �3 2 � ≈ 1.251 bits, 0.38 nats the entropy h of a character takes a maximum value when the probabilities are uniformly distributed (cover and thomas 2006, theorem 2.6.4), and that maximum is log n where n denotes “the number of elements in the range”. this result has been stated as “maximal h… depends only on the number of states” (abbott et al. 1985, p.101), as “the base of the logarithms is an arbitrary choice” (lance and williams 1966), and as “if one wanted to confine h always to the range 0 to 1, then logarithms to base m can be used for characters with m states ... this is in effect just the same as multiplying h by a normalizing constant.” (pankhurst 1991, p. 192). however, the base of the logarithm has practical implications. for the example in table 2, the number of taxa is 3, and there are 4 states. using log3 gives h = 0.79 to 2 decimal places and using log4 gives h = 0.63. base m, the number of states, is not ideal (pankhurst’s (1991) use of m rather than n might perhaps have originated as a typographical error). it is better to normalize using base n, the number of taxa, so that the coefficient remains within the range [0,1] as the number of taxa decreases in the course of an identification, which gives the user a consistent impression of whether a character has significant separating power. normalizing by the number of states achieves a [0,1] range but removes the benefit of multi-state characters and can give a low rank to completely separable taxa if each taxon has multiple (non-shared) states (table 6 includes an example of such an anomalous ranking). subsequent authors may have noticed the problem, but to our knowledge none has implemented the solution; for example, delta-intkey uses log2 throughout. numeric ranges numeric ranges are treated in most respects like other discrete categories, with the extremely useful exception (dallwitz et al. 2002) that the user specifies a simple value. for example, if taxon1 allows leaf length 1–3 cm, a calculation when the database is loaded indicates whether this is distinct from the leaf lengths of other taxa or not. if overlap between taxa occurs, such as if taxon2 allows leaf length 2–3 cm, then discrete component intervals with open ‘()’ or closed ‘[]’ limits can be calculated, which are henceforth treated separately. when the user enters leaf length 2.16 cm, this needs to match a character state allowed for both taxon1 and taxon2. in principle, it is easy to translate numeric ranges into sets, and thence into equivalent characters with non-numeric categorical states. there are some problems with doing this, however. the first problem is a conceptual one: numeric ranges as used in keys impose artificial limits; a character such as length is not naturally categorical, but taxonomists routinely cope with the necessary conversions as they design character 𝐼(𝑋;𝑌) = 1 3 log� 1 3 1 3 × 1 3 �+ 1 6 log� 1 6 1 3 × 1 6 �+ 1 6 log� 1 6 1 3 × 1 3 �+ 1 6 log� 1 6 1 3 × 1 3 �+ 1 6 log� 1 6 1 3 × 1 6 � figure 1: calculation of the information content for the example in table 2 5 talent, dickinson, and dickinson — best characters states, choosing a total range or a “normal range”, or choosing states that describe a statistical distribution (jardine and sibson 1970). here we follow the approach used by delta-intkey, assuming that although ranges might be entered in a more complex format, they come to this component of the polyclave software as simple ranges, e.g., leaf length 2–3 cm. a second problem is that the computer programming involved in interpreting the specified numeric ranges is non-trivial, but that is also routinely dealt with in polyclave software. a complication here is that overlapping ranges can be divided into subintervals in more than one way; in the example above either [1,2)+[2,3] or [1,2]+(2,3] is possible, as are non-minimal subdivisions such as [1,2)+[2]+(2,3]. the software needs to sort all range endpoints for the character states of a character, then work through the sorted list creating the minimum number of required subranges, and when a choice is possible, consistently assigning closed limits on either the left or right ends of intervals. if the program assigns closed limits at the right ends, then the above example with just two taxa produces limits [1,2]+(2,3], and the entered data 2.16 must match character state (2,3] which is allowed for both taxon1 and taxon2. a third problem is the one we are most interested in here, that, in keeping with the assumptions of information theory, it is not correct to calculate the information content of numericrange data using intersection of sets (tables 3 and 4). the assumptions of information theory are that the taxa are equally likely, and that within each taxon the various character states are equally likely (abramson 1963; cover and thomas 2006; shannon 1948). the size of a numeric range is irrelevant (table 4). however, unshared discrete states, if shared states are also present, become more important if they are subdivided (table 3). consequently, when ranges are compared, it is important to count the number of intervals of overlap and non-overlap (table 4, table 5), which is a different calculation from that used in set theory, potentially producing a different character ranking. examples in a polyclave, normalizing the best character coefficient to the range [0,1] clarifies whether a character warrants evaluation, which could be helpful if a relatively expensive technique such as microscopy or molecular testing is required. we recommend a coefficient that is the information content normalized by the number of taxa (table 6). a coefficient of 1.0 signals that specifying that one character will resolve all taxa. the unnormalized value lacks this clarity, and normalizing by the number of character states gives an incorrect ranking of the different characters (e.g., flower color in table 6). a stand-alone bestchar program for categorical data is provided that calculates a variety of coefficients for a single categorical character. the source code is dedicated to the public domain, as per http://creativecommons.org/publicdomain/zero/1.0/ three versions are given, one in the python language (http://www.python.org/) for clarity in handling lists, one in the r language (http://www.r-project.org/) which is heavily used in biodiversity informatics (kindt and coe. 2005; rossi 2011; chamberlain and barve 2012; kembel 2012; chamberlain et al. 2013; hijmans et al. 2013; vanderwal et al. 2013), and source code for a java applet and application (http://docs.oracle.com/javase/7/docs). the program source files bestchar.py, bestchar.r, and bestchar.java, as well as charinput.txt, a sample input data set corresponding to table 5 char 2 are accessible at https://github.com/nadiatalent/bestchar. 6 http://creativecommons.org/publicdomain/zero/1.0/ http://www.python.org/ http://www.r-project.org/ http://docs.oracle.com/javase/7/docs https://github.com/nadiatalent/bestchar talent, dickinson, and dickinson — best characters discussion as polyclave identification systems become larger, with more varied types of data such as mixed morphological and dna characters, character ranking might become important and widely used. although a naïve user might have some difficulty grasping the assumptions on which it is based, and might therefore prefer to ignore the statistic, the experienced botanists and taxonomists whom we have asked generally favor the inclusion of such an automatically calculated feature. the single calculated value indicates to the user who is identifying a specimen whether their effort to evaluate a particular character is likely to yield a significant reduction in the remaining search space of possible identifications. a traditional alternative approach, embodied particularly in single-entry keys, is to use hard-coded expert opinion about which characters are most useful or more reliable. however, whenever serious effort is needed to assess the character states of a specimen, as might occur in forensic and some other applications, precoded character rankings may not be the best guide. few current taxon-identification systems as yet use a fully multi-entry (polyclave) structure, so the possibilities of character ranking are hardly explored. we are convinced that its potential utility will not be appreciated until it is widely and correctly implemented, until polyclave users have seen it provide a useful hint in difficult situations. if the formula used is clearly linked to the table 3: if shared states are present, subdividing a character state increases weighting of the character in the information statistic. ‘/’ in a list of character states indicates ‘or’. char 1 char 2 char 3 char 4 taxon 1 a a/b a/b/c a/b/c/d taxon 2 b a a a discrete sets of states 2 2 2 2 information (in bits) 1.0 0.25 0.33 0.38 table 4: the extent of a numeric range is irrelevant to the information content (char 1 and char 2, taxon 5), except in so far as it affects the complexity of how numeric ranges overlap. overlapping ranges that differ at one extremity (char 3) are readily converted to equivalent categories using the minimum number of subintervals needed to distinguish the character states. char 3 requires five subintervals [1,2), [2,3), [3,4), [4,5), [5], equivalent to the five categories of char 4. ‘/’ in a list of character states indicates ‘or’. char 1 char 2 char 3 char 4 taxon 1 1–2 1–2 1–5 a/b/c/d/e taxon 2 3–4 3–4 2–5 b/c/d/e taxon 3 5–6 5–6 3–5 c/d/e taxon 4 7–8 7–8 4–5 d/e taxon 5 9–10 9–10000 5 e intkey’s best character (observed) 2.32 2.32 0.41 0.41 information (in bits) 2.32 2.32 0.41 0.41 normalized to the range [0,1] for 5 taxa 1 1 0.18 0.18 7 talent, dickinson, and dickinson — best characters established literature on information theory rather than left undocumented or hidden in a proprietary formula, students may find that literature helpful, which could promote the use of the statistic. we suspect that the overlap in terminology and the use of the information coefficient in multivariate data exploration may have caused some biologists and software developers to assume that character ranking in polyclaves is a complex topic with many possible solutions, but as we have reviewed above, the aim and the approach are actually quite simple. the characters with the highest ranking are those that divide the remaining taxa into evenly sized small groups. characters with numeric-range values are difficult to deal with taxonomically, but if coded in such a way that they accurately describe taxa, there is no reason to exclude them from the rank calculation. we emphasize that ranking characters by their information content in an online polyclave is a distinct problem that has its own special requirements. a related problem is to distinguish taxa and devise a taxonomy, which can involve cluster analysis using similarity or dissimilarity measures, and is a large research focus in many areas apart from biological systematics. another related problem occurs in software aids for developing single-entry (usually binary) identification keys, but the requirements differ and a separation coefficient is commonly used (hill table 5: numeric ranges that overlap are not treated as equivalent to sets; rather, the areas of overlap and non-overlap may form a greater number of categories that separately contribute to the information statistic. non-overlap at one extremity (table 4) is a lower-entropy situation than with additional nonoverlap (char 1, char 2). delta-intkey incorrectly analyzes situations of complex overlap in a way that appears to be consistent with combining non-sequential intervals of non-overlap into a set of intervals, rather than treating the intervals separately, using five sets rather than nine components of char 1, to produce 0.41 rather than 0.47. ‘/’ in a list of character states indicates ‘or’; anomalous coefficients appear in bold-face; () and [] indicate range limits, open (omitting the endpoint) and closed (including the endpoint) respectively; ∪ indicates a set union operation on intervals. char 1 subintervals of char 1 sets of char 1 subintervals char 2 taxon 1 1–10 [1,2),[2,3),[3,4),[4,5),[5,6 ],(6,7],(7,8],(8,9],(9,10] [1,2)∪(9,10], [2,3)∪(8,9],[3,4)∪(7,8 ],[4,5)∪(6,7],[5,6] a/b/c/d/e/f/g/h/i taxon 2 2–9 [2,3),[3,4),[4,5),[5,6],(6,7 ],(7,8],(8,9] [2,3)∪(8,9],[3,4)∪(7,8 ],[4,5)∪(6,7],[5,6] b/c/d/e/f/g/h taxon 3 3–8 [3,4),[4,5),[5,6],(6,7],(7,8 ] [3,4)∪(7,8],[4,5)∪(6,7 ],[5,6] c/d/e/f/g taxon 4 4–7 [4,5),[5,6],(6,7] [4,5)∪(6,7],[5,6] d/e/f taxon 5 5–6 [5,6] [5,6] e intkey’s best character (observed) 0.41 0.47 information (in bits) 0.47 0.47 0.41 0.47 normalized to the range [0,1] for 5 taxa 0.20 0.20 0.18 0.20 8 talent, dickinson, and dickinson — best characters 1974; pankhurst 1991; burguiere et al. 2013; eti bioinformatics undated); in that situation, lookahead may be important (quinlan 1986), and polytomies may be undesirable. to our knowledge, delta-intkey is the only currently available system that uses a coefficient related to information theory, and it is currently being reimplemented (atlas of living australia 2011 onwards), which may allow for improvement. we have suggested that the coefficient should be normalized to the range [0,1], in effect dividing by the number of remaining taxa, to make it more clearly interpretable to the user. we have also discussed the treatment of characters with numericrange data, which requires special care to remain consistent with information theory. computational expense is an issue. the description of actkey (brach and song 2005) suggested that a pre-calculated replication of intkey’s best character would be used to sort the characters, and it would not be recalculated in later stages of the identification and would therefore become unreliable. we would argue against taking that approach if at all possible because character ranking could be particularly useful after the most readily available data have been used, when it may be necessary to decide whether to use an expensive test. at late stages like this, the number of remaining taxa is probably reduced, and the calculation therefore becomes more feasible. acknowledgements the authors are grateful to graeme hirst, sara scharf, and don tarnawski for helpful discussion, to two anonymous reviews for insightful suggestions, and to the ievobio2012 organizing committee for agreeing to include nt among their number, an experience that produced a couple of serendipitous insights that encouraged submission of this manuscript. table 6: if the character-ranking coefficient, normalized by the number of taxa, has a value of 1.0, this means that specifying the character state will resolve all taxa, as with the calyx edge character both before and after the user has specified the state of the life cycle character. this helpful hint to the user is not available from the unnormalized information (initially 1.58 bits). using the number of states to normalize is incorrect, not reflecting the relative effectiveness of the characters to resolve the taxa. ‘/’ in a list of character states indicates ‘or’; anomalous coefficients appear in bold-face. initial conditions after choosing life cycle=perennial calyx edge flower color life cycle calyx calyx edge flower color calyx taxon 1 crenate white annual glabrous taxon 2 dentate orange/pink perennial glabrous/ pubescent dentate orange/pink glabrous/ pubescent taxon 3 cuspidate pink/red perennial pubescent cuspidate pink/red pubescent #states 3 4 2 2 2 3 2 #taxa 3 3 3 3 2 2 2 information (in bits) 1.58 1.25 0.92 0.58 1.0 0.50 0.25 normalized by #states 1.0 0.63 0.92 0.58 1.0 0.32 0.25 normalized by #taxa 1.0 0.79 0.58 0.37 1.0 0.50 0.25 9 talent, dickinson, and dickinson — best characters references abbott, l. a., f. a. bisby, and d. j. rogers. 1985. taxonomic analysis in biology: computers, models, and databases. columbia university press, new york. abramson, n. 1963. information theory and coding. mcgraw-hill, new york. alexander, g. 2006. sliks-alike interactive key software (saiks). accessible at http://www.galexander.org/saiks/readme. atlas of living australia. 2011 onwards. open-delta: a java port of the delta – description language for taxonomy suite. accessible at http://code.google.com/p/open-delta/. brach, a. r. and h. song. 2005. actkey: a web-based interactive identification key program. taxon. 54:1041–1046. burguiere, t., f. causse, v. ung, and r. vignes-lebbe. 2013. ikey+: a new single-access key generation web service. syst. biol. 62:157–161. chamberlain, s., boettiger, c., ram, k. and barve, v. 2013. package ‘rgbif’: interface to the global biodiversity information facility api methods. accessible at http://cran.rproject.org/web/packages/rgbif/rgbif.pdf. chamberlain, s. and barve, v. 2012. package ‘rvertnet’: search vertnet database from r. accessible at http://cran.rproject.org/web/packages/rvertnet/rvertnet.pdf. christensen, k. i. 1999. meka – an introduction to the use of meacham’s multiple-entry key algorithm. university of copenhagen. cover, t. m. and j. a. thomas. 2006. elements of information theory. john wiley & sons, inc., hoboken, new jersey. dale, m.b., m. beatrice, r. venanzoni, and c. ferrari. 1986. a comparison of some methods of selecting species in vegetation analysis. coenoses, 1, 35–52. dallwitz, m., t. paine, and e. zurcher. 1998/2012. principles of interactive keys. accessible at http://delta-intkey.com/www/interactivekeys.htm. dallwitz, m. j. 1974. a flexible computer program for generating identification keys. systematic zoology. 23:50–57. dallwitz, m. j. 1993. delta and intkey. pp. 287–296 in r. fortuner, ed. advances in computer methods for systematic biology: artificial intelligence, databases, computer vision. the johns hopkins university press, baltimore. dallwitz, m. j. 2010. overview of the delta system. accessible at http://delta-intkey.com/www/overview.htm. dallwitz, m. j., t. a. paine, and e. j. zurcher. 2002. interactive identification using the internet. pp. 23– 33 in h. saarenmaa, and e. s. nielsen, eds. towards a global biological information infrastructure — challenges, opportunities, synergies, and the role of entomology, european environment agency technical report 70. eea, copenhagen. dallwitz, m. j., t. a. paine, and e. j. zurcher. 2010. user’s guide to the delta system: a general system for processing taxonomic descriptions. accessible at http://delta-intkey.com/www/uguide.htm. de lamarck, j. b. p. a. d. m. 1778. flore françoise; ou, description succincte de toutes les plantes qui croissent naturallement en france. disposée selon une nouvelle méthode d’analyse, & à laquelle on a joint la citation de leurs vertus les moins équivoques en médecine, & de leur utilité dans les arts. l'imprimerie royale, paris. accessible at http://www.biodiversitylibrary.org/item/38206. duke, j. a. 1969. on tropical tree seedlings i. seeds, seedlings, systems, and systematics. annals of the missouri botanical garden. 56:125–161. duncan, t. and c. a. meacham. 1986. multiple-entry keys for the identification of angiosperm families using a microcomputer. taxon. 35:492–494. duncan, t. and c. a. meacham. 1987. meka manual. university herbarium, university of california, berkeley, california, usa. eti bioinformatics. undated. linnaeus ii. accessible at http://www.eti.uva.nl/products/linnaeus.php. gower, j. c. 1975. relating classification to identification. pp. 251–263 in r. j. pankhurst, ed. biological identification with computers. academic press, london and orlando. guala, g. f. 2004–12. sliks: stinger’s lightweight interactive key software. accessible at http://www.stingersplace.com/sliks/. hagedorn, g., g. rambold, and s. martellos. 2010. types of identification keys. pp. 59–64 in p. l. nimis, and r. v. lebbe, eds. tools for identifying biodiversity: progress and problems. eut edizioni università di trieste, trieste. 10 http://www.galexander.org/saiks/readme http://code.google.com/p/open-delta/ http://cran.r-project.org/web/packages/rgbif/rgbif.pdf http://cran.r-project.org/web/packages/rgbif/rgbif.pdf http://cran.r-project.org/web/packages/rvertnet/rvertnet.pdf http://cran.r-project.org/web/packages/rvertnet/rvertnet.pdf http://delta-intkey.com/www/interactivekeys.htm http://delta-intkey.com/www/overview.htm http://delta-intkey.com/www/uguide.htm http://www.biodiversitylibrary.org/item/38206 http://www.eti.uva.nl/products/linnaeus.php http://www.stingersplace.com/sliks/ talent, dickinson, and dickinson — best characters hijmans, r.j., phillips, s., leathwick, j. and elith, j. 2013. package 'dismo': species distribution modeling. accessible at http://cran.rproject.org/web/packages/dismo/dismo.pdf. hill, l. r. 1974. theoretical aspects of numerical identification. int. j. syst. bacteriol. 24:494–499. jardine, n. and r. sibson. 1970. quantitative attributes in taxonomic descriptions. taxon 19:862–870. kembel, s. 2012. biodiversity analysis in r: csee r workshop 2012. accessible at http://phylodiversity.net/skembel/rworkshop/biodivr/sk_biodiversity_r.html. kindt, r. and r. coe. 2005. tree diversity analysis: a manual and software for common statistical methods for ecological and biodiversity studies. world agroforestry centre, nairobi, kenya. accessible at http://www.worldagroforestry.org/downloads/publi cations/pdfs/b13695.pdf. lance, g. n. and w. t. williams. 1966. computer programs for hierarchical polythetic classification (“similarity analyses”). the computer journal. 9:60–64. lobanov, a. l., w. f. schilow, and l. m. nikritin. 1981. zur anwendung von computern für die determination in der entomologie. dtsch. entomol. z. 28:29–43. lucidecentral.org. 2010. about lucid. accessible at http://www.lucidcentral.org/home/aboutlucid/tabi d/203/language/en-us/default.aspx. macdonald, d. k. c. 1952. information theory and its application to taxonomy. journal of applied physics. 23:529–531. macleod, n., ed. 2007. automated taxon identification in systematics: theory, approaches and applications. crc press, taylor and francis group, boca raton. martellos, s. 2010. multi-authored interactive identification keys: the frida (friendly identification) package. taxon. 59:922–929. morse, l. e. 1971. specimen identification and key construction with time-sharing computers. taxon. 20:269–282. morse, l. e. 1975. recent advances in the theory and practice of biological specimen identification. pp. 11–52 in r. j. pankhurst, ed. biological identification with computers. academic press, london and orlando. osborne, d. v. 1963. some aspects of the theory of dichotomous keys. new phytol. 62:144–160. pankhurst, r. j. 1975. identification by matching. pp. 79–91 in r. j. pankhurst, ed. biological identification with computers. academic press, london and orlando. pankhurst, r. j. 1978. biological identification: the principles and practice of identification methods in biology. edward arnold, london. pankhurst, r. j. 1991. practical taxonomic computing. cambridge university press, cambridge. pankhurst, r. j. and r. r. aitchison. 1975. a computer program to construct polyclaves. pp. 73–78 in r. j. pankhurst, ed. biological identification with computers. academic press, london and orlando. payne, r. w. and d. a. preece. 1980. identification keys and diagnostic tables: a review. j. roy. stat. soc. ser. a. (stat. soc.). 143:253–292. podani, j. 2000. introduction to the exploration of multivariate biological data. backhuys publishers, leiden. quinlan, j. r. 1986. induction of decision trees. machine learning. 1:81-106. rosatti, t. j. undated. electronic, interactive identification keys for california plants using meka (multiple-entry key algorithm). accessible at http://ucjeps.berkeley.edu/keys/. rossi, j.-p. 2011. rich: an r package to analyse species richness. diversity. 3:112–120. shannon, c. e. 1948. a mathematical theory of communication. the bell system technical journal. 27:379–423, 623–656. shwayder, k. 1971. conversion of limited-entry decision tables to computer programs — a proposed modification of pollack’s algorithm. communications of the acm. 14:69–73. shwayder, k. 1974. extending the information theory approach to converting limited-entry decision tables to computer programs. communications of the acm. 17:532–537. sneath, p. h. a. and r. r. sokal. 1973. numerical taxonomy: the principles and practice of numerical classification. w. h. freeman and company, san francisco. spooner, a. and a. chapman. 2007. delta intkey tutorial. western australian herbarium. accessible at http://florabase.dec.wa.gov.au/help/keys/intkey_tut orial.pdf. 11 http://cran.r-project.org/web/packages/dismo/dismo.pdf http://cran.r-project.org/web/packages/dismo/dismo.pdf http://phylodiversity.net/skembel/r-workshop/biodivr/sk_biodiversity_r.html http://phylodiversity.net/skembel/r-workshop/biodivr/sk_biodiversity_r.html http://www.worldagroforestry.org/downloads/publications/pdfs/b13695.pdf http://www.worldagroforestry.org/downloads/publications/pdfs/b13695.pdf http://www.lucidcentral.org/home/aboutlucid/tabid/203/language/en-us/default.aspx http://www.lucidcentral.org/home/aboutlucid/tabid/203/language/en-us/default.aspx http://ucjeps.berkeley.edu/keys/ http://florabase.dec.wa.gov.au/help/keys/intkey_tutorial.pdf http://florabase.dec.wa.gov.au/help/keys/intkey_tutorial.pdf talent, dickinson, and dickinson — best characters toda, m. j., k. matsushita, and s. f. mawatari. 2004. biological classification and identification system (biocis). neo-science of natural history: proceedings of international symposium on “dawn of a new natural history — integration of geoscience and biodiversity studies”, sapporo. accessible at http://hdl.handle.net/2115/38489. ung, v., g. dubus, r. zaragüeta-bagils, and r. vigneslebbe. 2010. xper²: introducing e-taxonomy. bioinformatics 26:703–704. vanderwal, j., falconi, l., januchowski, s., shoo, l. and storlie, c. 2013. package ‘sdmtools’: species distribution modelling tools: tools for processing data associated with species distribution modelling exercises. accessible at http://cran.rproject.org/web/packages/sdmtools/sdmtools.p df. voss, e. g. 1952. the history of keys and phylogenetic trees in systematic biology. journal of the scientific laboratories, denison university 43:1–25. walter, d.e. and s. winterton, s. 2006. keys and the crisis in taxonomy: extinction or reinvention? ann. rev. entomol. 52:193-208. walters, s. m. 1975. traditional methods of biological identification. pp. 3–8 in r. j. pankhurst, ed. biological identification with computers. academic press, london and orlando. 12 http://hdl.handle.net/2115/38489 http://cran.r-project.org/web/packages/sdmtools/sdmtools.pdf http://cran.r-project.org/web/packages/sdmtools/sdmtools.pdf http://cran.r-project.org/web/packages/sdmtools/sdmtools.pdf character selection during interactive taxonomic identification: “best characters” choice of coefficient the information coefficient normalizing the coefficient numeric ranges examples discussion acknowledgements references biodiversity informatics, 15, 2020, pp. 92-102 92 the abundant niche-centroid hypothesis: key points about unfilled niches and the potential use of supraspecfic modeling units carlos yañez-arenas1*, gerardo martín2, luis osorio-olvera3,4, jazmín escobar-luján1, sandra castaño-quintero1, xavier chiappa-carrara1 and enrique martínez-meyer5 1laboratorio de ecología geográfica, unidad de biología de la conservación, parque científico y tecnológico de yucatán, universidad nacional autónoma de méxico, sierra papacal, yucatán 97302, méxico. 2 mrc centre for global infectious disease analysis, department of infectious disease epidemiology, imperial college london, uk. 3biodiversity institute, university of kansas, lawrence, 66045, usa. 4instituto de ecología, universidad nacional autónoma de méxico, ciudad universitaria, mexico city 04500, méxico 5departmento de zoología, instituto de biología, universidad nacional autónoma de méxico, ciudad universitaria, mexico city 04510, méxico. abstract. correlative estimates of fundamental niches are gaining momentum as an alternative to predict species’ abundances, particularly via the abundant niche-centroid hypothesis (an expected inverse relationship between species’ abundance variation across its range and the distance to the geometric centroid of its multidimensional ecological niche). the main goal of this review is to recapitulate what has been done, where we are now, and where should we move towards in regards to this hypothesis. despite evidence in support of the abundance-distance to niche centroid relationship, its usefulness has been highly debated, although with little consideration of the underlying theory regarding the circumstances that might break down the relationship. we address some key points about the conditions needed to test the hypothesis in correlative studies, specifically in relation to niche characterization and configurations of the biotic-abiotic-mobility (bam) framework to illustrate the problem of unfilled niches. using a created supraspecific modeling unit, we show that species for which only a portion of their fundamental niche is represented in their area of historical accessibility (m)—i.e., when the environmental equilibrium condition is violated—it is impossible to characterize their true niche centroid. therefore, we strongly recommend to analyze this assumption prior to assess the abundant niche-centroid hypothesis. finally, we discuss the potential of using modeling units above the species level for cases in which environmental conditions associated with species’ occurrences may not be sufficient to fully characterize their fundamental niches. key words: ecological niche modeling, niche centrality, abundant niche-centroid hypothesis, abundance, population density. introduction species distribution models (sdms) and ecological niche models (enms) represent a set of tools and techniques in which georeferenced records of presence (and sometimes absence) of species are statistically related with a set of environmental predictors (e.g., temperature, precipitation, elevation) to infer their ecological requirements (i.e., their ecological niche), and project them onto the geography to estimate their potential distribution (peterson et al. 2011). the estimated distributions can be used for different purposes: to discriminate areas with and without biological potential for the species of interest (guisan et al. 2006), evaluate potential shifts in the geographic ranges of species as a consequence of environmental changes (thomas et al. 2004; peterson 2006), identify regions where invasive species could * corresponding author: lichoso@gmail.com biodiversity informatics, 15, 2020, pp. 92-102 93 establish (peterson 2003; thuiller et al. 2005), describe biodiversity patterns and carry out macroecological studies (guisan and rahbek 2011; calabrese et al. 2014), among many others. however, for certain research goals and questions, knowing the extent of occurrence is not sufficiently informative (gaston and rodrigues 2003; hurlbert and jetz 2007; hooker et al. 2011). for instance, one of the criteria recommended in the design and establishment of natural protected areas is to maximize the area of high-quality habitat for a target species (pearce and ferrier 2001). in a distribution map it is not possible to identify areas of high-quality habitat since all portions of the species’ distribution have the same weight. furthermore, distributional maps alone are useless to identify whether the protection of a fraction of the distribution (without associated information on demographic aspects) is sufficient to guarantee the viability of the species (rodrigues et al. 2004). population abundance or density are frequently good indicators of habitat quality (johnson 2007), reflecting factors such as reproductive rate, longevity, carrying capacity, and susceptibility of populations to extinction. therefore, modeling and mapping species’ abundance may be more informative for many researchers and stakeholders than just estimating geographic ranges (hobbs and hanley 1990). when abundance/density is available from several locations, it is preferable to model it directly as a function of key environmental predictors (boyce et al. 2001). however, obtaining abundance data is complicated and demanding, especially for rare and cryptic species (johnston et al. 2015). therefore, since sdms/ enms became popular, researchers have been interested in evaluating their capacity to infer species’ abundance (e.g., pearce and ferrier 2001; pearce and boyce 2006). outputs of many algorithms used in sdms/ enms are raster layers with continuous values that are usually interpreted as environmental suitability, that is: higher values should represent better environmental conditions for species. yet, this interpretation assumes the existence of a positive relationship between the estimated output of the algorithm and independent measures directly related to a species’ biological fitness, like abundance or population density (gil and lobo 2017). first attempts to evaluate this relationship, via a sdm framework, failed to consistently provide strong abundance-suitability correlations (pearce and ferrier 2001; pearce and boyce 2006; jiménez-valverde et al. 2009). probably, these inconsistencies may be due to the incapacity of some modeling methods to account for environmental suitability (jiménez-valverde et al. 2009). later, some studies found more promising results when using correlative enms (vanderwal et al. 2009; yañez-arenas et al. 2012; martínez-meyer et al. 2013; weber et al. 2017). among these works, the abundant niche-centroid hypothesis stands out as a key element to understand the abundance-environmental suitability relationship. the abundant niche-centroid hypothesis niche theory suggests that reproduction and survival of individuals should be higher in localities placed at optimal conditions of their ecological fundamental niche (nf); conceptualized as an n-dimensional hypervolume where each dimension represents a relevant variable that acts on the organism’s fitness (hutchinson 1973; peterson et al. 2011). this idea was initially suggested by hutchinson (1957), but maguire (1973) was the first one to explicitly propose that different regions of the nf space should correspond to different values of the species intrinsic population growth rate (r). if this is true and population abundance is an expression of fitness, then, species’ abundance should be explained by their position with respect to the centroid of their nf’s (i.e., greater abundance should be found in populations closer to the centroid and decreases towards the edges showing a negative correlation; fig. 1; martínez-meyer et al. 2013). initially the hypothesis accumulated some empirical support: 1) yañez-arenas et al. (2012) and martínez-meyer et al. (2013) observed negative correlations between population abundance/density and the distance to the niche centroid (estimated with correlative methods) of some vertebrates; 2) manthey et al. (2015), ureña-aranda et al. (2015) and martínez-gutiérrez et al. (2018) noted that populations tend to have positive growth rates closer to the niche center and negative at the margins. however, more recent comprehensive analyses have obtained contradicting results: dallas et al. (2017) and santini et al. (2019) found weak support to the niche-centrality hypothesis when tested for different taxa; in contrast osorio-olvera et al. (2020) observed that correlations between abundances and the distance to the niche centroid of north american birds were mostly biodiversity informatics, 15, 2020, pp. 92-102 94 negative. differences between findings could be explained by methodological artifacts and the quality of the abundance data used (knouft 2018; soberón et al. 2018). in any case, these studies tested the hypothesis in all possible species for which abundance data were available. however, we consider that it is necessary, in the first place, to assess some theoretical assumptions and circumstances under which the abundant niche-centroid hypothesis may be able to explain the geographic patterns of species’ abundance. on the problem of unfilled fundamental niches testing the abundant niche-centroid hypothesis requires an unbiased characterization of a species’ nf, which would allow the estimated centroid to truly represent its environmental optimum (osorio-olvera et al. 2019). however, estimating the nf via correlative techniques is not an easy task (peterson et al. 2011). under many circumstances a species’ niche may be geographically unfilled (strubbe et al. 2013; ashby et al. 2017). the bam diagram is a simple heuristic tool specifically designed to summarize the potential effects of biotic interactions (b), abiotic suitable conditions (a) and mobility characteristics (m) on a species’ distribution, and can be used to assess ex-ante whether it is possible to characterize the nf correlatively or not (soberón and peterson 2005). biological features of species and scales of analysis define relative sizes and positions of the bam components that lead to different configurations, some of which allow a better characterization of the nf (saupe et al. 2012). we use the bam components and configurations to exemplify situations in which distances to the niche centroid estimated from distributional data may and may not explain species abundances. the geographic expression of a species’ nf is termed the existing fundamental niche (nf*), which is equivalent to a in the bam framework (soberón and peterson 2005; soberón 2010). relationships between these terms can vary among species and geographical scales. if a species’ nf = nf* = a ⸦ m, and b has no significant effects on constraining the occupied geography (go, or the intersection of b, a and m in the traditional configuration), then niche centrality determined correlatively is likely to explain abundance patterns. such bam configuration is commonly called ‘the hutchinson’s dream’ (saupe et al. 2012) and it is the ideal scenario to test the abundant niche-centroid hypothesis (fig. 2; upper right panel). however, if nf* ⸦ nf, then there will be portions of the nf that do not exist in geography (g) and therefore cannot be characterized from species’ occurrences. here, an adequate estimation of the niche centroid will depend on the magnitude of the difference between nf and nf*. for instance, thermal tolerance limits of some reptiles and amphibians is roughly 10 oc above the optimum, which could constitute a general guideline for assessing the difference between nf and nf* (soberón and arfigure 1. the abundant niche-centroid hypothesis. left panel: bi-dimensional environmental space in which a hypothetical niche, its centroid and internal structure are shown. black circles represent localities with different population abundance (defined by the size of the circle). right panel: expected relationship between abundance and the distance to the niche centroid. biodiversity informatics, 15, 2020, pp. 92-102 95 royo-peña 2017). alternatively, if only a portion of a exhibits biotic features suitable for the survival of a species, then go = a ∩ b, and a reduction of the nf is also expected. hutchinson (1957) was the first one to explicitly describe the above relationships in environmental space when the concept of realized niche (nr) was introduced. recent empirical data reveals that key interspecific interactions may have significant effects on species’ geographic ranges (e.g., gaston 2003; louthan et al. 2006; gotelli et al. 2010; ashby et al. 2017; anderson 2017). such cases would result in a major challenge to study abundance-centrality relationships, because it is almost axiomatic that estimated nf will be reduced or biased. therefore, research on niche centrality is highly dependent on the assumption of the so-called ‘eltonian noise’ that is, when inter-specific interactions do not limit species’ ranges at coarse geographical scales (soberón and nakamura 2009; soberón 2010). as described above, an important assumption to estimate a species’ nf via correlative modeling is that individuals can move freely across g and occur in all abiotic suitable locations, while being absent from all unsuitable ones, i.e., the environmental equilibrium assumption (araújo and pearson 2005). however, under many real-life situations, species are limited by geographical barriers and dispersal abilities (soberón and peterson 2005; soberón 2010); hence, unfilled nf are expected under some bam configurations in which m is included, such as ‘classic bam’ (go = b ∩ a ∩ m; go = a ∩ m) (fig. 2; upper left panel) and ‘wallace’s dream’ (go = m ⸦ b ∩ a; go = m ⸦ a) (fig. 2; lower left panel). also, considering that m determines the set of regions and environments occupied by the species, the known set of presence records (g+) occur in geographic space (go), so the associated environments η(g+) are already reduced by m. as a consequence, niche models that have been calibrated in an m region will frequently under-characterize the nf (peterson and soberón 2012), and the estimation of their centroid will be biased. an exception may occur under the ‘full overlap’ configuration since go = m ≈ a and a complete characterization of the species nf is possible (fig. 2; lower right panel). a hypothetical example of m effects using two virtual entities we show a simple example of how the correlation between abundance and distance to the niche centroid may be profoundly affected by an incomplete characterization of the nf under two different m configurations. generating virtual entities we created two virtual entities that are based on the biology and real data of two mussels; mytilus edulis and m. galloprovincialis. we chose these species because they are distributed in two geographical regions separated by mainland europe which have different environmental conditions. both species are closely related, therefore for some niche axes they may have similar environmental tolerances (lee et figure 2. bam configurations in which we show, in each letter (a – h), the potential scenarios regarding the distribution of the niche centroid conditions (red dots) in geographic space. in all cases we have omitted b to match the components accounted for in the analyses. abiotic suitable conditions = red circles (a component of bam); historically accessible area = blue circles (m component of bam); go = occupied geography. upper left panel (a c) = classic bam; lower left panel (e – g) = wallace’s dream; upper right panel (d) = hutchinson’s dream; lower right panel (h) = full overlap. biodiversity informatics, 15, 2020, pp. 92-102 96 al. 2019). the former is native to northern europe, and the latter occur in the mediterranean sea (wonham 2004). based on the assumption of phylogenetic niche conservatism (harvey and pagel 1991), we hypothesized that both mussel species could have largely similar nf. therefore the bam configuration of our example is go = a ∩ m (the classic bam without b). inputs for generating the niche models we obtained presence records of m. edulis (hereafter “northern entity”) and m. galloprovincialis (“southern entity”) from gbif (https://www.gbif. org). all data points which were evidently miss-georeferenced (outside its known range or placed on land) were eliminated. data were filtered for duplicate points with the “clean_dup” function of the “ntbox” r package (osorio-olvera et al. 2020), using a distance threshold equal to the resolution of the environmental data (below). finally, we randomly selected 150 occurrence records of each species to remove from the dataset used to characterize nf. hence, we obtained for each modeling entity two spatially-independent datasets of presence points called “d_pres” (full database without the 150 removed points) and “d_abun” (the 150 records removed from the original filtered database). we also obtained 12 marine variables from the bio-oracle database v2.0 (assis et al. 2018; http://www.bio-oracle.org), at a spatial resolution of 10’ (~20 km). these surfaces describe annual trends of sea surface temperature and salinity for the period 2000–2014: average, range, average maximum, average minimum, maximum, and minimum. we used a principal components analysis (pca) to reduce multicollinearity applying the ‘pcaraster’ function of the ‘enmgadgets’ package (barve and barve 2013) in r (r core team 2018), and retained the first three components that explained 93% of the overall variance for further analyses. niche models first, we merged the “d_pres” datasets of occurrence records of both virtual entities to create the “coupled unit”. then, using a minimum volume ellipsoid (mve, van aelst and rosseeuw 2009) we generated a hypothetical nf that included 97.5% of the occurrence records (leaving out the presence records from atypical environmental conditions that might represent sink populations or undetected errors in the original data cleaning process; fig. 3; green ellipsoid). the assumption is that nf are convex, as suggested by abundant observational and physiological experimental data (maguire 1973; hooper et al. 2008; angilletta 2009; soberón and nakamura 2009), and theoretical arguments (drake 2015). the mve was generated with the function “cov.rob” of the mass package in r (r core team 2019), and graphed with “rgl” (adler et al. 2019). using the “cov_center” function of package “ntbox” (osorio-olvera et al. 2020), we calculated the volume and its centroid. then we projected the mahalanobis distance to the mve centroid with the “mahalnobis” function of r to project the niche in geographic space (g) and obtained a continuous map of environmental suitability. the inputs to the “mahalanobis” function were the covariance matrix of the environmental variables and the vector of the centroid coordinates. finally, we built the mves using the same described methods with the sets of presences for the northern and southern entities separately (fig. 3; blue and red ellipsoids, respectively). abundance data as abundance data we used the maximum number of individuals reported for either species in an area equal to the size of the grid cells. then, we matched abundance data with the distance to the centroid of the mve of the coupled unit. the result was figure 3. minimum volume ellipsoids defined by three marine variables (sea surface mean temperature, sea surface temperature range, sea surface salinity). green ellipsoid: hypothetical fundamental niche built with presence records of both species (“northern” and “southern”). blue ellipsoid: niche characterized only from occurrences of the northern species. red ellipsoid: niche characterized only from occurrences of the southern species. the estimated centroid is also presented for each entity. biodiversity informatics, 15, 2020, pp. 92-102 97 that the maximum observed abundance (12,500 virtual individuals) was perfectly, negatively correlated with the distance to the niche centroid (fig. 4). using this approach of virtual entities we eliminated the effects of biotic interactions, dispersion, human impact, and variability of physiological tolerances among populations. statistical analyses we performed spearman correlation tests with the “cor_test” r function between abundance and distance to the centroid of the northern and southern entities. the two real species (and their virtual entities) have geographically disjoint distributions, separated by mainland europe, which represents the main barrier to dispersal (m) and impedes their sympatry. simultaneously, europe interrupts the spatial continuity of the environmental conditions to which these species have historically developed physiological tolerance. therefore, these tests are equivalent to an assessment of the effect of a niche characterization limited by m. results the resulting environmental suitability from each mve had a distinct geographic pattern. in the coupled unit model, the highest environmental suitability occurred throughout most of the cantabric sea, a portion of the northern sea and a small region in the southern coast of france, in the mediterranean sea. in the northern entity model, the highest environmental suitability almost exclusively occurred in the northern sea, while in the southern entity there are high suitability values (small distances to the niche centroid) across the three southern european seas: adriatic, aegean and tyrrhenian. given that hypothetical abundance was derived from environmental suitability of the coupled unit model, the correlation between this and abundance was ρ = 0.99. the correlation between abundance and the environmental suitability estimated for the other models were ρ = 0.27 and ρ = 0.16, for the northern and southern entities, respectively (fig. 5). conclusion as expected, the correlation between abundance and environmental suitability characterized from occupied niches reduced and biased by m was significantly lower. this example shows that accessibility has a crucial effect on the characterization of the nf and the resulting environmental centroids. under bam configurations in which there is a previous suspicion that there is a strong effect of m over go, the abundant niche-centroid hypothesis has low expectations to explain the geographic patterns of abundance. then, the following questions arise from these problems: 1) should species with unfilled fundamental niches be completely avoided when measuring the centrality-abundance relationship? 2) what options exist to reconstruct the nf of species in which this problem exists? the potential use of supraspecific modeling units mechanistic modeling is a good alternative to estimate and map species’ nf truncated by geography. these techniques are based on data from controlled experiments for measuring the physiological effects of different manipulated variables to estimate their tolerance limits (kearney and porter 2004; 2009). building these models, however, requires complex and long experimental processes and equipment (gallien et al. 2010). for these reasons, mechanistic models are usually built only with one or two environmental variables, and are unlikely to capture all the biologically relevant environmental factors for a species (aragón et al. 2010). in addition their application is limited to species in which physiological experiments are feasible (larson et al. 2014). figure 4. distribution of the 300 locations with a hypothetical abundance value. blue circles represent abundance data of the northern species within its m (pink polygon) and red circles represent abundance data of the southern species within its m (yellow polygon). biodiversity informatics, 15, 2020, pp. 92-102 98 on the other hand, characterizing the nf from presence data and using correlative methods is limited under many bam scenarios, as we have shown in the previous sections. nevertheless, modeling with supraspecific entities (i.e., modeling units above species level), may represent an alternative to overcome the problem of “unfilled niches” (qiao et al. 2017; smith et al. 2018; castaño et al. in press). this approach is valid under the assumption of phylogenetic niche conservatism, which states that evolutionary patterns of species with common ancestors share a substantial portion of their biological and physiological characteristics that determine their nf. in other words, ancestral adaptations to a set of environmental conditions tend to be preserved by descendant species (harvey and pagel 1991). niche conservatism tends to break down over time, although there is vast evidence showing that at timescales comparable with speciation events, the most common pattern is that of little ecological divergence of nf (climatic preferences of species tend to be phylogenetically preserved; peterson 2011; pavoine and bonsall 2011). therefore, it is possible to assume that sister species have similar nf despite having natural geographical distributions with different climatic characteristics. under this scenario, each species would inhabit different portions and combinations of the entire environmental space, but complementary regarding their nf. this strategy of combining presence records of sister species is known as “lumping” (smith et al. 2018). recent analyses of invasive species have demonstrated that lumping allows better characterization of nf when it has been reduced in go by m or b (castaño et al. in press). lumping, therefore, appears to be a good strategy to improve predictability of biological invasions or estimating nf for evaluating the abundant niche-centroid hypothesis. the assumption is that when bam configurations produce unfilled niches, lumping would allow a better approximation to esfigure 5. top panels: environmental suitability estimated from the distances to the niche characterized by different mves (the gradient goes from highest to lowest suitability = from light blue to dark blue). panels below: correlations between hypothetical abundance and environmental suitability estimated by each modeling entity. biodiversity informatics, 15, 2020, pp. 92-102 99 timate the nf and its centroid, hence improving the chances that abundance is explained by the abundant niche-centroid hypothesis. a third possibility is that of modeling subspecific units (e.g., subspecies, populations, environmental units) when local adaptation plays an important role in determining local abundance. under this scenario, the single centroid would not represent the species’ optimum, instead, several subspecific units would have individual optima. this may occur in species with very extensive geographic distributions and wide ecological niches (yañez-arenas et al. 2012). concluding remarks the abundant niche-centroid hypothesis is currently a hot topic in biogeography, macroecology and distributional ecology. to-date, at least three studies have evaluated the generality of the idea, one of them finding strong support for the hypothesis (osorio-olvera et al. 2020), and the other two showing contrasting results (dallas et al. 2017; santini et al. 2019). some of the factors that may affect the expected negative correlation between species’ abundance and the distance to their niche centroid are related to methodological issues, such as the effect of sample size and bias in occurrences used to build niche models (yañez-arenas et al. 2014), the quality of abundance/ density data (soberón et al. 2018, knouft 2018), and decisions regarding the metrics used to compute environmental distances (soberón et al. 2018). others are inherent to the system of study. for instance, biotic interactions (e.g., high levels of competition, predation, herbivory, or parasitism), metapopulation dynamics (e.g., stochastic processes of individuals dispersal among populations, allee effects) and human impact (e.g., direct exploitation of species, habitat destruction/modification) can decrease species’ population abundances in localities that are environmentally near the centroid of the niche (osorio-olvera et al. 2016; weber et al. 2017; osorio-olvera et al. 2019). also, adaptation of populations to local environments would result in higher abundance in those localities (leimu and fischer 2008), some of which may be distributed at the edges of the niche where selection pressures are stronger (aguirre-liguori et al. 2017). in such cases, splitting lineages into subespecific units (e.g., environmentally similar units) and estimate niche models for each unit may be a good strategy to describe environmental relationships (yañez-arenas et al. 2012; smith et al. 2018). some of the mentioned factors are impossible to identify prior to test the abundant niche-centroid hypothesis. however, here we demonstrate that thinking about assumptions related to the problem of unfilled niches and analyzing the potential bam configuration could be used to decide ex-ante in which cases it is worth testing the hypothesis. we showed that m is important in limiting the nf of species and the explanatory power of the niche structure towards abundance within its distribution. thus, species in which environmental conditions within their m do not represent a subset of their nf are the best candidates to study abundance-niche relationships. finally, we introduce the idea that supraspecific units can help to overcome some of the limitations inherent to testing the abundant niche-centroid hypothesis. acknowledgments we thank town peterson for his useful comments on the supra-specific modeling idea. we thank the editors of biodiversity informatics, ángela cuervo and luis escobar, for their effort and patience to organize this special supplement, and anonymous reviewers whose comments and suggestions improved the quality of this manuscript. references adler, d., d. murdoch, o. nenadic, s. urbanek, m. chen, a. gebhardt, b. bolker, g. csardi, a. strzelecki, and a. senger. 2019. rgl: 3d visualization using opengl. r package version 0.96.16. aguirre‐liguori, j. a., m. i. tenaillon, a. vázquez‐lobo, b. s. gaut, j. p. jaramillo‐correa, s. montes‐hernández, and l. e. eguiarte, l. e. 2017. connecting genomic patterns of local adaptation and niche suitability in teosintes. molecular ecology 26: 4226-4240. anderson, r. p. 2017. when and how should biotic interactions be considered in models of species niches and distributions? journal of biogeography 44: 8-17. angilletta jr, m. j., and m. j. angilletta, m. j. 2009. thermal adaptation: a theoretical and empirical synthesis. oxford university press. oxford. aragón, p., a. baselga, and j. m. lobo. 2010. global estimation of invasion risk zones for the western corn rootworm diabrotica virgifera virgifera: integrating distribution models and physiological thresholds to assess climatic favourability. journal of applied ecology 47:1026-1035. araújo, m. b., and r. g. pearson. 2005. equilibrium of species’ distributions with climate. ecography 28:693-695. biodiversity informatics, 15, 2020, pp. 92-102 100 ashby, b., e. watkins, j. lourenço, s. gupta, and k. r. foster. 2017. competing species leave many potential niches unfilled. nature: ecology and evolution 1:1495. assis, j., l. tyberghein, s. bosch, h. verbruggen, e. a. serrão, and & o. de clerck. 2018. bio‐oracle v2. 0: extending marine data layers for bioclimatic modelling. global ecology and biogeography 27: 277-284. boyce, m. s., d. i. mackenzie, b. f. manly, m. a. haroldson, and d. moody. 2001. negative binomial models for abundance estimation of multiple closed populations. journal of wildlife management 65:498-509. calabrese, j. m., g. certain, c. kraan, and c. f. dormann. 2014. stacking species distribution models and adjusting bias by linking them to macroecological models. global ecology and biogeography 23: 99-112. castaño-quintero, s., j. escobar-luján, l. osorio-olvera, a. t. peterson, x. chiappa-carrara, e. martínez-meyer, and c. yañez-arenas. in press. supraspecific units in correlative niche modelling improves the prediction of biological invasions. peerj. dallas, t., r. r. decker, and a. hastings. 2017. species are not most abundant in the centre of their geographic range or climatic niche. ecology letters, 20:15261533. drake, j. m. 2015. range bagging: a new method for ecological niche modelling from presence-only data. journal of the royal society interface 12(107): 20150086. gallien, l., t. münkemüller, c. h. albert, i. boulangeat, and w. thuiller. 2010. predicting potential distributions of invasive species: where to go from here? diversity and distributions 16:331-342. gaston, k. j. 2003. the structure and dynamics of geographic ranges. oxford university press. oxford. gaston, k. j., and a. s. rodrigues. 2003. reserve selection in regions with poor biological data. conservation biology 17:188-195. gil, g., and j. lobo. 2017. ¿son útiles los modelos de distribución realizados a gran escala para determinar la variación en la abundancia local de especies emblemáticas? el caso del parque nacional iguazú. revista biodiversidad neotropical 7:269-283. gotelli, n. j., g. r. graves, and c. rahbek. 2010. macroecological signals of species interactions in the danish avifauna. proceedings of the national academy of sciences usa 107: 5030-5035. guisan, a., o. broennimann, r. engler, m. vust, n. g. yoccoz, a. lehmann, and n. e. zimmermann. 2006. using niche‐based models to improve the sampling of rare species. conservation biology 20: 501-511. guisan, a., and c. rahbek. 2011. sesam–a new framework integrating macroecological and species distribution models for predicting spatio‐temporal patterns of species assemblages. journal of biogeography 38:1433-1444. harvey, p. h., and m. d. pagel. 1991. the comparative method in evolutionary biology (vol. 239). oxford: oxford university press. oxford. hobbs, n. t., and t. a. hanley. 1990. habitat evaluation: do use/availability data reflect carrying capacity? the journal of wildlife management 515-522. hooker, s. k., a. cañadas, k. d. hyrenbach, c. corrigan, j. j. polovina, and r. r. reeves. 2011. making protected area networks effective for marine top predators. endangered species research 13: 203-218. hooper, h. l., r. connon, a. callaghan, g. fryer, s. yarwood-buchanan, j. bigg, and r. m. sibly. 2008. the ecological niche of daphnia magna characterized using population growth rate. ecology 89:1015-1022. hurlbert, a. h., and w. jetz. 2007. species richness, hotspots, and the scale dependence of range maps in ecology and conservation. proceedings of the national academy of sciences usa 104:13384-13389. hutchinson, g. 1957. population studies-animal ecology and demography concluding remarks. in: cold spring harbor symposia on quantitative biology. cold spring harbor lab press. bungtown, plainview, ny 11724:415-427. jiménez-valverde, a., f. diniz, e. b. de azevedo, and p. a. borges. 2009. species distribution models do not account for abundance: the case of arthropods on terceira island. annales zoologici fennici 46: 451-465. johnson, m. d. 2007. measuring habitat quality: a review. condor 109: 489-504. johnston, a., d. fink, m. d. reynolds, w. m. hochachka, b. l. sullivan, n.e. bruns, e. hallstein, m. s. merrifield, s. matsumoto, and s. kelling. 2015. abundance models improve spatial and temporal prioritization of conservation resources. ecological applications 25:1749-1756. kearney, m., and w. porter. 2004. mapping the fundamental niche: physiology, climate, and the distribution of a nocturnal lizard. ecology 85: 3119-3131. kearney, m., and w. porter. 2009. mechanistic niche modelling: combining physiological and spatial data to predict species’ ranges. ecology letters 12:334-350. knouft, j. h. 2018. appropriate application of information from biodiversity databases is critical when investibiodiversity informatics, 15, 2020, pp. 92-102 101 gating species distributions and diversity: a comment on dallas et al. (2018). ecology letters 21:1119-1120. larson, e. r., r. v. gallagher, l. j. beaumont, and j. d. olden. 2014. generalized “avatar” niche shifts improve distribution models for invasive species. diversity and distributions 20:1296-1306. lee, y., h. kwak, j. shin, k. seung-chul, t. kim, and j. k. park. 2019. a mitochondrial genome phylogeny of mytilidae (bivalvia: mytilida). molecular phylogenetics and evolution 139: 106533. leimu, r., and m. fischer. 2008. a meta-analysis of local adaptation in plants. plos one 3, e4010. louthan, a.m., d. f. doak, and a. l. angert. 2006. where and when do species interactions set range limits? trends in ecology and evolution 30:780–792. maguire, jr, b. 1973. niche response structure and the analytical potentials of its relationship to the habitat. american naturalist 107:213-246. manthey, j. d., l. p. campbell, e. e. saupe, j. soberón, c. m. hensz, c. e. myers, h. l. owens, k. ingenloff, a. t. peterson, n. barve, a. lira-noriega, and v. barve. 2015. a test of niche centrality as a determinant of population trends and conservation status in threatened and endangered north american birds. endangered species research 26: 201-208. martínez‐gutiérrez, p. g., e. martínez‐meyer, f. palomares, and n. fernández. 2018. niche centrality and human influence predict rangewide variation in population abundance of a widespread mammal: the collared peccary (pecari tajacu). diversity and distributions 24:103-115. martínez-meyer, e., d. díaz-porras, a. t. peterson, and c. yáñez-arenas. 2013. ecological niche structure and rangewide abundance patterns of species. biology letters 9:20120637. osorio‐olvera, l., lira‐noriega, a., soberón, j., peterson, a. t., falconi, m., contreras‐díaz, r. g., … barve, n. (2020). ntbox: an r package with graphical user interface for modeling and evaluating multidimensional ecological niches. methods in ecology and evolution, 11:1199-1206. osorio-olvera, l. a., m. falconi, and j. soberón. 2016. sobre la relación entre idoneidad del hábitat y la abundancia poblacional bajo diferentes escenarios de dispersión. revista mexicana de biodiversidad 87:1080-1088. osorio‐olvera, l., j. soberón, and m. falconi. 2019. on population abundance and niche structure. ecography 42:1415-1425. osorio-olvera l., yañez-arenas c., martinez-meyer e, and a. t. peterson. 2020. relationships between population densities and niche-centroid distances in north american birds. ecology letters doi: 10.1111/ ele.13453. pavoine, s., and m. b. bonsall. 2011. measuring biodiversity to explain community assembly: a unified approach. biological reviews 86:792-812. pearce, j. l., and m. s. boyce. 2006. modelling distribution and abundance with presence‐only data. journal of applied ecology 43:405-412. pearce, j., and s. ferrier. 2001. the practical value of modelling relative abundance of species for regional conservation planning: a case study. biological conservation 98:33-43. peterson, a. t. 2003. predicting the geography of species’ invasions via ecological niche modeling. the quarterly review of biology 78:419-433. peterson, a. t. 2006. uses and requirements of ecological niche models and related distributional models. biodiversity informatics 3:59-72. peterson, a.t. 2011. ecological niche conservatism: a time-structured review of evidence. journal of biogeography 38: 817–827. peterson, a. t., and j. soberón. 2012. integrating fundamental concepts of ecology, biogeography, and sampling into effective ecological niche modeling and species distribution modeling. plant biosystems-an international journal dealing with all aspects of plant biology 146:789-796. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo .2011. ecological niches and geographic distributions. princeton university press. princeton. qiao, h., a. t. peterson, l. ji, and j. hu. 2017. using data from related species to overcome spatial sampling bias and associated limitations in ecological niche modeling. methods in ecology and evolution 8:18041812. r core team. 2019. r: a language and environment for statistical computing. r foundation for statistical computing. vienna, austria. rodrigues, a. s., s. j. andelman, m. i. bakarr, l. boitani, t. m. brooks, r. cowling, l. d. fishpool, g. a. da fonseca, k. j. gaston, and m. hoffmann. 2004. effectiveness of the global protected area network in representing species diversity. nature 428:640. santini, l., s. pironon, l. maiorano, and w. thuiller. 2019. addressing common pitfalls does not provide biodiversity informatics, 15, 2020, pp. 92-102 102 more support to geographical and ecological abundant‐centre hypotheses. ecography 42: 696-705. saupe, e. e., v. barve, c. e. myers, j. soberón, n. barve, c. m. hensz, a. t. peterson, h. l. owens, and a. lira-noriega. 2012. variation in niche and distribution model performance: the need for a priori assessment of key causal factors. ecological modelling 237:1122. smith, a. b., w. godsoe., f. rodríguez-sánchez., h. h. wang, and d. warren. 2018. niche estimation above and below the species level. trends in ecology and evolution 34: 260-273 soberón, j. m. 2010. niche and area of distribution modeling: a population ecology perspective. ecography 33:159-167. soberón, j., and b. arroyo-peña. 2017. are fundamental niches larger than the realized? testing a 50-year-old prediction by hutchinson. plos one 12, e0175138. soberón, j., and m. nakamura. 2009. niches and distributional areas: concepts, methods, and assumptions. proceedings of the national academy of sciences usa 106:19644-19650. soberón, j., and a. t. peterson .2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodiversity informatics 2:1-10. soberón, j., a. t. peterson, and l. osorio-olvera. 2018. a comment on “species are not most abundant in the center of their geographic range or climatic niche”. biorxiv 266510. strubbe, d., o. broennimann, f. chiron, and e. matthysen. 2013. niche conservatism in non‐native birds in europe: niche unfilling rather than niche expansion. global ecology and biogeography 22:962-970. thomas, c. d., a. cameron, r.e. green, m. bakkenes, l. j. beaumont, y.c. collingham, b. n. erasmus, m. ferreira de siqueira, a. grainger, l. hannah, l. hughes, b. huntley, a. van jaarsveld, g. f. midgley, l miles, m. a. ortega-huerta, t. peterson, o. l. phillips, and l. hughes. 2004. extinction risk from climate change. nature 427:145. thuiller, w.,d. m. richardson, p. pyšek, g. f. midgley, g. o. hughes, and m. rouget. 2005. niche‐based modelling as a tool for predicting the risk of alien plant invasions at a global scale. global change biology 11: 2234-2250. ureña-aranda, c. a., o. rojas-soto, e. martínez-meyer, c. yáñez-arenas, r. l. ramirez, and a. e. de los monteros. 2015. using range-wide abundance modeling to identify key conservation areas for the micro-endemic bolson tortoise (gopherus flavomarginatus). plos one 10:e0131452. van aelst s., and p. rousseeuw. 2009. minimum volume ellipsoid. wiley interdisciplinary reviews: computational statistics 1:71-82. vanderwal, j., l. p. shoo, c. n. johnson, and s. e. william. 2009. abundance and the environmental niche: environmental suitability estimated from niche models predicts the upper limit of local abundance. american naturalist 174: 282-291. weber, m. m., r. d. stevens, j. a. f. diniz‐filho, and c. e. v. grelle. 2017. is there a correlation between abundance and environmental suitability derived from ecological niche modelling? a meta‐analysis. ecography 40:817-828. wonham, m. j. 2004. mini-review: distribution of the mediterranean mussel, mytilus galloprovincialis (bivalvia: mytilidae), and hybrids in the northeast pacific. journal of shellfish research 23:535-544. yañez-arenas, c., r. guevara, e. martínez-meyer, s. mandujano, and j. m. lobo. 2014. predicting species’ abundances from occurrence data: effects of sample size and bias. ecological modelling 294:36-41. yañez‐arenas, c., e. martínez‐meyer, s. mandujano, and o. rojas‐soto. 2012. modelling geographic patterns of population density of the white‐tailed deer in central mexico by implementing ecological niche theory. oikos 121:2081-2089. microsoft word cmr_plants_20171115.docx biodiversity informatics, 12, 2017, pp. 76-83 botanical sampling gaps across the cameroon mountains moses nsanyi sainge1, 2, jean-michel onana3,4, felix nchu5, david kenfack6 and a. townsend peterson7 1tropical plant exploration group (tropeg), p.o. box 18, mundemba, ndian, south west region, cameroon; 2department of environmental and occupational studies, faculty of applied sciences, cape peninsula university of technology, cape town campus, keizersgracht, p.o. box 652, cape town 8000, south africa; 3national herbarium of cameroon. p.o. box 1601 yaoundé, centre region, cameroon; 4department of plant biology, faculty of science, university of yaoundé 1. p.o. box 812 yaoundé, cameroon; 5department of horticultural sciences, faculty of applied sciences, cape peninsula university of technology, p.o box 1906, bellville 7535, south africa; 6center for tropical forest science, smithsonian institution global earth observatory, washington, dc 20560-0166 usa; 7biodiversity institute, university of kansas, 1345 jayhawk blvd., lawrence, kansas 66045 usa abstract.—with the emergence of a new field, biodiversity informatics, an important task has been to evaluate completeness of biodiversity information that is existing and available for various countries and regions. this paper offers a first and very basic assessment of sampling gaps and inventory completeness across the cameroon mountains. because digital accessible knowledge is severely limited for the region, we relied on qualitative evaluations of inventory completeness, supplemented by large amounts of data from the national herbarium of cameroon (ya) database. detailed botanical inventories have been developed for mt cameroon, the kupe-mwanenguba mountains, mt oku, and the mambila plateau, leaving substantial geographic and environmental coverage gaps corresponding to rumpi hills, mt nlonako, kimbi fungom national park, bali and bafut ngemba, mt bamboutos, kagwene, and tchabal mbabo. this paper provides a roadmap for a comprehensive botanical survey for this region. completing this survey plan, the resulting data will allow researchers to track changes in biodiversity and identify priority areas for conservation on the various mountain ranges that make up this important biodiversity hotspot. keywords.—biodiversity informatics, primary biodiversity data, sampling gaps, inventory completeness, mountains. the cameroon mountains, also known as the cameroon line (nono et al. 2004) or cameroon volcanic line (ayonghe et al. 1999, marzoli et al. 2000), comprise a chain of isolated volcanic and plutonic mountain peaks that covers ~40,877 km2, stretching from pagalu island in the gulf of guinea to the mandara mountains in the interior. the oceanic portion of the mountain chain covers ~2998 km2 in the form of 4 major islands (pagalu, 17 km2; são tomé, 854 km2; principe, 110 km2; bioko, 2017 km2; ayonghe et al. 1999, frodin 2001). the continental portion of the mountain chain is broader, covering ~37,879 km2, extending from mt cameroon (4095 m) and mt etinde (commonly called small mt cameroon, 1474 m) on the coast, through a string of peaks, extending to the north and east into the interior (figure 1). the largest part of the cameroon mountain chain lies within cameroon, which is the focus of this study, but it extends into adjacent nigeria at sites including the mambila plateau, obudu plateau, mbola hills, mt shebsi, and the biu plateau (hepper 1965, 1966, hall and medler 1975, ayonghe et al. 1999; figure 1). numerous protected areas cover at least some portions of the cameroon mountain chain, including mt cameroon national park, bakossi national park, santchou forest reserve, kimbifungom national park, mbam et djerem national park, and vallée de la mbéré national park. however, although endemism is high in the region overall, particularly for plants, these protected areas were established based on minimal information, and no quantitative conservation 76 sainge et al. – cameroon mountains plants figure 1. diagrammatic representation of the cameroon mountain chain, extending from islands in the gulf of guinea to the coastal mt cameroon, and inland along the cameroon-nigeria border. 77 sainge et al. – cameroon mountains plants prioritization has been developed for the region. botanists have invested considerable energy in basic surveys of the plants of this region, although few primary data are available openly from this work. reasonably complete botanical surveys have been developed for only a few sites: this study set out to characterize those points of light and the intervening gaps in detail. more generally, however, the cameroon mountains hold tropical montane forests known internationally for their rich floras and high levels of endemism, combined with high degrees of threat (myers et al. 2000, onana and cheek 2011). this biome is part of a mega-hotspot of western africa cited as having impressive species diversity and hosting numerous endemic taxa (mittermeier et al. 1999, myers et al. 2000, marchese 2015). still, only scanty primary biodiversity data are available for the region, and—to our knowledge—no quantitative analysis of conservation priority has been developed regarding any taxon for the region. as such, a detailed evaluation of gaps in knowledge regarding the biota of this region is presented here to lay a foundation for more detailed inventories and analyses in terms of structure and composition, species diversity, and richness. because major sectors of the botanical knowledge of this region remain in non-digital or non-shared formats, our analyses are necessarily qualitative at this point, in hope of transition to fully digital and quantitative in coming years. methods we focused on montane areas with elevations >1600 m, but we considered information from those areas down to 600-1000 m. we produced a spatial dataset summarizing the distribution of montane areas within the cameroon mountains from the gmted2010 digital elevation model (dem1), at a spatial resolution of 7.5” (~230 m at the equator). we reclassified the raw dem into elevational intervals of 0-1000 m (lowlands), 1000-1600 m (foothills), and >1600 m (montane areas), and created vector-format shape-files of areas >1000 m and >1600 m to facilitate analyses. finally, for analysis, we further subdivided two large areas into smaller subunits based on relatively low valleys that were nonetheless above 1600 m (figure 2). 1 http://topotools.cr.usgs.gov/gmted_viewer/. primary data on previous botanical surveys across the region were derived from published documents, floras, monographs, reports, and the databases of the national herbarium of cameroon (ya) (holmgren et al. 1990). from the latter source, 23,667 records were assessed for accuracy, and records falling in highland areas extracted. this information was summarized in spreadsheets that included latitude, longitude, elevation, number of data records, and number of species documented, but numbers of records and species were unconvincingly low; the bulk of existing information for the plants of the region remains in non-digital formats. as a consequence, we resorted to a qualitative evaluation (our consensus opinion) of completeness of the inventory of each area, as 0 = never (to our knowledge) visited, 1 = incidental data, 2 = some records, and 3 = a published ‘flora’ or otherwise comprehensive summary that appears to be reasonably complete; we focused on areas with completeness at only the latter level, which we considered, at least in a preliminary sense, as well-sampled. once we had established which polygons could be considered as well-sampled, we converted the shapefile to raster (geotiff) format using custom scripts in r (r foundation for statistical computing 2004). this raster coverage was the basis for our identification of gaps, as follows. we used the proximity (raster distance) function in qgis (version 2.4) to summarize geographic distances from all montane sites to those that are well-sampled. to create a parallel view of environmental difference from well-sampled areas, we plotted 5000 random points across the cameroon mountains (i.e., in areas >1600 m), and used the point sampling tool in qgis to link each point to the geographic distance raster, and to raster data layers at 30” (~1 km) spatial resolution summarizing annual mean temperature and annual precipitation from the worldclim climate data archive (hijmans et al. 2005). we first rescaled values of each environmental variable to the overall range of the variable as (xi – xmin) / (xmax xmin), where xi is the particular observed value in question, such that each environmental variable varied only 0-1. we calculated the minimum euclidean distance in the two-dimensional climate space to any well-sampled area (sousa-baena et al. 2014). finally, these environmental distances were 78 sainge et al. – cameroon mountains plants figure 2. summary of highland areas in the continental portion of the cameroon mountains region. areas above 1000 m elevation are shown in darker gray shading. areas above 1600 m are shown colored red and outlined in black. the more important such highland areas are labeled for reference to the text. figure 3. map of the continental portion of the cameroon mountains region, showing points sampled in the ya collections in blue. the different highland areas are indicated as well-sampled (dark red), some sampling (medium red), some information (light red) or no information (gray). 79 sainge et al. – cameroon mountains plants imported into qgis, and linked back to the random point shape-file. results among the 23,667 digital herbarium records from ya, 13,609 records corresponded to highlands within the cameroon mountains region, the latter data representing a total of 3995 species (figure 3). data derived from published papers, reports, floras, and monographs correspond to 3314 records of 1175 species in the cameroon mountains region. our analysis shows significant sampling gaps along this mountain chain, including the rumpi hills, mt nlonako, the lebialem highlands west of mt bamboutos, and the bamenda highlands (extending to tchabal mbabo and the adamawa plateau). our explorations of the ya database (figure 3) indicated that data were sparse for individual montane areas, preventing quantitative analysis (colwell and coddington 1994). as a result, we were forced to rely on our qualitative analysis, detailed below. a broad view of the cameroon mountains region with our qualitative inventory summary is presented in figures 1 and 2. we identified 3 highland sites that could be considered as wellinventoried: mt cameroon, mt mwanenguba (including bakossi national park), and the mambila plateau. sites for which at least some information exists include santchou forest reserve, mt bamboutous, mt lefo, bali and bafut ngemba forest reserve, ijim and mt oku, tchabal mbabo, tchabal gandaba, mandara mountains, and benue national park. other areas, however, are characterized by only scanty botanical information (e.g., mt nlonako, kagwene, rhum rock, faro reserve). for 3 sites, we could find no indication of any previous botanical work: nkambe, rhum rock, ngel nyaki mountain forest. see table 1 for a summary of all of this information. geographically, the sites most distant from well-known sites were the adamawa plateau and tchabal mbabo (figure 2). interestingly, in terms of environmental distances, the sites most different from well-known sites were quite different from the list based on geographic distances. specifically, the montane areas most distinct in terms of environments were the rumpi hills, mt nlonako, as well as the lebialem highlands west of mt bamboutos. these sites are curiously located more or less centrally in the cameroon mountains chain, and are geographically relatively close to wellknown sites (figure 4, table 1). discussion the continental part of the cameroon mountains with focus on cameroon is one of the most diverse sites in africa, and has been classified as a biodiversity hotspot (cheek et al. 2000, cheek et al. 2004, barthlott et al. 2005). this region is the only part of central africa with an elevational range from sea level to >4000 m, and holds a high diversity of plants >6000 species out of the ~9000 species so far known to occur in cameroon (cable and cheek 1998, cheek et al. 2000, cheek et al. 2004, onana and cheek 2011, onana 2011). at the continental level, this region has great affinity in species composition with other montane sites such as the mountains of west africa (e.g., bersama abyssinica; cheek et al. 2004) and east africa (e.g., polyscias fulva, strombosia scheffleri, schlefflera abyssinica, alangium chinense, maesa lanceolata; dowsettlemaire 1989, thomas and thomas 1996; cheek et al. 2000, cheek et al. 2004, sainge 2016). this region holds >200 species of plants that are considered as threatened, which is the highest in cameroon and perhaps the highest in west and central africa (cheek et al. 2004, onana and cheek 2011), with >80 species endemic (cheek et al. 2004, franke 2004, sainge et al. 2005, sainge et al. 2010, sainge 2012, sainge 2016). examples of threatened and endemic plant species occurring in the region are afrothismia saingei, a. fungiformis, rhaptopetalium geophylax, schefflera manni, syzyzium staudtii, ixora foliosa, gambeya korupensis, deinbollia angustifolia, and begonia pseudoviola. the region hosts 3 of the 7 genera endemic to cameroon (hamilcoa, medusandra, platytinospora); finally, the only endemic family in cameroon (medusandraceae) is represented here, with its two species: medusandra mpomiana and m. richardsiana (cheek et al. 2004, onana 2013). however well, this landscape has been visited by and sampled by botanists, although that work was mostly in terms of surveys of species’ occurrences; most of the data remain in nondigitized formats. this information thus remains inaccessible to scientists with interest on african 80 sainge et al. – cameroon mountains plants figure 4. summary of gaps in coverage in terms of inventories of plants in the cameroon mountains region. white indicates well-known sites; darker shades of blue indicate greater distance (geographic space) or difference (environmental space) from well-surveyed sites. left-hand column is the southwestern part of the chain; right-hand column is the northeastern part of the chain. table 1. highland areas in the continental portion of the cameroon mountains region with focus in cameroon. areas are presented roughly in order from southwest (mt. cameroon) to northwest. montane unit number of species recorded latitude longitude elevation (m) degree of documentation mt cameroon 2435 4.2833 9.2167 2600 3 rumpi hills 326 4.8886 9.2418 1749 1 mt kupe 10 4.7921 9.6779 966 1 mwanengumba (incl. bakossi np) 2412 4.9969 9.8564 2040 3 mt nlonako 351 4.9031 9.9578 1606 1 santchou forest reserve 134 5.2602 10.0470 801 2 mt. bamboutous (incl. lebialem highlands) 222 5.6008 10.0400 2176 2 mt lefo 95 5.8641 10.3453 1239 2 kagwene 70 6.2717 9.4338 1025 1 bali and bafut ngemba forest reserve 415 5.8291 10.1026 1615 2 ijim and mt oku 920 6.1146 10.2856 2200 2 kumbo 1 6.2032 10.6848 1650 1 nkambe 0 6.6101 10.6707 1572 0 rhum rock 0 6.4671 11.0387 1603 0 ngel nyaki mountain forest 0 7.0903 11.0667 1167 0 tchabal mbabo 215 7.2406 12.1458 2166 2 mambila plateau (incl. gashaka gumti np) 213 7.3838 11.7303 1140 3 tchabal gandaba 155 7.0911 14.4376 1013 2 faro reserve 17 7.7151 12.4251 1021 1 mandara mountains 518 10.4942 13.6359 1079 2 benue national park 309 8.0405 14.0066 851 2 81 sainge et al. – cameroon mountains plants biodiversity, particularly those based in africa. we attempted to develop quantitative views of botanical inventory completeness across the cameroon mountains region (after sousa-baena et al. 2014, idohou et al. 2015, kouao et al. 2015), but we were stymied by the small numbers of primary occurrence data that are available for the region. as a consequence, in our second effort, we used a literature review to detect and identify landmark studies that have documented cameroon mountain sites in good detail, but found little or no access to the primary data that underlay those publications and that document the individual specimens collected. the occurrence data points from the cameroonian national herbarium corroborated this view: good numbers of occurrence points concentrated in sites identified as well-sampled in the qualitative analysis. in sum, then, the online digital accessible knowledge (dak; sousa-baena et al. 2014) for cameroon remains entirely too sparse, and should be the focus of intensive data development efforts. with a conservative estimate of ~155,000 specimens (and likely many more) collected to date in the country (onana 2011), only ~65,000 (41.9%) are represented in the database of ya (j.m. onana, unpubl. data.). this gap exists because large-scale collections made from the colonial era into the 1950s and subsequent decades are deposited in major herbaria in europe and north america; although not without exceptions, most of these data remain undigitized and largely inaccessible to the broader community interested in african biodiversity. although some of the big data-holders have begun steps to make information available (le bras et al. 2017), the dak impediment thus remains significant, with much of the primary data from extensive botanical surveys in the country remaining offline. this blockage of information flow must be resolved if any quantitative analyses of biodiversity pattern, subregional endemism, and conservation priority are to be developed for the region. acknowledgments we thank jrs biodiversity foundation for funding training activities that led to this collaboration. this paper is a contribution of tropical plant exploration group (tropeg), cameroon. literature cited ayonghe, s. n., g. t. mafany, e. ntasin, and p. samalang. 1999. seismically activated swarm of landslides, tension cracks, and a rockfall after heavy rainfall in bafaka, cameroon. natural hazards 19:13-27. barthlott, w., j. mutke, d. rafiqpoor, g. kier, and h. kreft. 2005. global centers of vascular plant diversity. nova acta leopoldina 92:61-83. cable, s., and m. cheek. 1998. the plants of mount cameroon: a conservation checklist. royal botanic gardens, kew, united kingdom. cheek, m., j. m. onana, and b. j. pollard. 2000. the plants of mount oku and the ijim ridge, cameroon. royal botanic garden, kew, united kingdom. cheek, m., b. j. pollard, l. darbyshire, j. m. onana, and c. wild. 2004. the plants of kupe, mwanenguba, and the bakossi mountains, cameroon: a conservation checklist. royal botanic garden, kew, united kingdom. colwell, r. k. and j. a. coddington. 1994. estimating terrestrial biodiversity through extrapolation. philosophical transactions of the royal society of london b 335:101-118. dowsett-lemaire f. 1989. the flora and phytogeography of the evergreen forests of malawi. 1: afromontane and mid-altitude forest. journal of the national botanical garden, belgium 59:3-131 franke t. 2004. afrothismia saingei (burmanniaceae), a new myco-heterotrophic plant from cameroon. systematics and geography of plants 74:27-33. frodin, d. g. 2001. guide to the standard floras of the world, 2nd ed. cambridge university press. hall, j. b., and j. a. medler. 1975. the botanical exploration of the obudu plateau area. nigeria field 40:101-117. hepper, f. n. 1965. the vegetation and floral of the vogel peak massif of the northern nigeria. bulletin of the institut fondamental d'afrique noire, series a. 27:413–513. hepper, f.n. 1966. outline to the vegetation and flora of the mambila plateau, northern nigeria. bulletin of the institut fondamental d'afrique noire, series a. 28:91-127. hijmans, r.j., s. e. cameron, j. l. parra, p. g. jones, and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25:1965-1978. holmgren, p.k., n. h. holmgren, and l. c. barnett. 1990. index herbariorum, 8th ed. new york botanical garden, new york. idohou, r., a. ariño, a. assogbadjo, r.g. kakai, and b. sinsin. 2015. diversity of wild palms (arecaceae) in the republic of benin: finding the gaps in the national inventory combining field and digital 82 sainge et al. – cameroon mountains plants accessible knowledge. biodiversity informatics 10:45-55. kouao, j.k., a.f. kouassi, c.y. adouyao, a. bakayoko, i. josephipou, j. bogaert. 2015. the present state of botanical knowledge in côte d’ivoire. biodiversity informatics 10:56-64. le bras, g. l., m. pignal, m. l. jeanson, s. muller, c. aupic, b. carré, g. flament, m. gaudeul, c. gonçalves, v. r. invernón, f. jabbour, e. lerat, p. p. lowry, b. offroy, e. p. pimparé, o. poncy, g. rouhan, and t. haevermans. 2017. the french muséum national d’histoire naturelle vascular plant herbarium collection dataset. scientific data 4:170016. marchese, c. 2015. biodiversity hotspots: a shortcut for a more complicated concept. global ecology and conservation 3:297-309. marzoli, a., e. m. piccirillo, p. r. renne, g. bellieni, m. iacumin, j. b. nyobe, and a. t. tongwa. 2000. the cameroon volcanic line revisited: petrogenesis of continental basaltic magmas from lithospheric and asthenospheric mantle sources. journal of petrology 41:87-109. mittermeier, r. a., n. myers, c. g. mittermeier, and p. robles-gil. 1999. hotspots: earth's biologically richest and most endangered terrestrial ecoregions. cemex, sa, agrupación sierra madre, mexico city, science. myers, n., r.a. mittermeier, c.g. mittermeier, g.a.b. da fonseca, and j. kent. 2000. biodiversity hotspots for conservation priorities. nature 403:853-858. nono, a., e. njonfang, a. k. dongmo, d. g. nkouathio, and f. m. tchoua. 2004. pyroclastic deposits of the bambouto volcano (cameroon line, central africa): evidence of an initial strombolian phase. journal of african earth sciences 39:409414. onana, j. m. 2011. the vascular plants of cameroon. a taxonomic check list with iucn assessements. flore du cameroun 39. irad-national herbarium of cameroon, yaoundé. onana, j. m., and m. cheek. 2011. the red data book of the flowering plants of cameroon. royal botanic garden, kew, united kingdom. onana j.m. 2013. synopsis des espèces végétales vasculaires endémiques et rares du cameroun. check-liste pour la conservation de la biodiversité. in j. m. onana (ed.). flore du cameroun 40, ministry of scientific research and innovation (minresi), yaoundé. r foundation for statistical computing. 2004. r, version 3.1.2. https://www.r-project.org/. sainge, m.n., t. franke, and r. agerer. 2005. afrothismia korupensis (burmanniaceae, tribe thismieae) from korup national park, cameroon. wildenowia 35:287-291. sainge, m.n., t. franke, v. merckx, and j.-m. onana. 2010. distribution of myco-heterotrophic (saprophytic) plants of cameroon. in: x. van der burgt, j. van der maesen & j.-m. onana (eds), systematics and conservation of african plants, royal botanic gardens, kew. pp. 279-286. sainge m.n. 2012. systematics and ecology of thismiaceae of cameroon, master of science thesis, university of buea, cameroon. 110 pp. sainge, n.m. 2016. patterns of distribution and endemism of plants in the cameroon mountains: a case study of protected areas in cameroon: rumpi hills forest reserve (rhfr) and the kimbi fungom national park (kfnp). tropical plant exploration group (tropeg), buea, cameroon. sousa-baena, m.s., l. c. garcia, a. t. peterson. 2014. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. diversity and distributions 20:369-381. thomas, d.w., and j. thomas. 1996, tchabal mbabo botanical survey. world wide fund for nature (wwf). yaounde, cameroon. 83 microsoft word pam_based_indicessept2.docx biodiversity informatics, 10, 2015, pp. 22-34 indices of biodiversity pattern based on presence-absence matrices: a gis implementation jorge soberón1, 2 and jeff cavner1 1biodiversity institute, university of kansas, lawrence, ks 66045, usa 2department of ecology & evolutionary biology, university of kansas, lawrence, ks 66045, usa abstract.—in this paper we present mathematical notation and formulae relating a number of indices of the biodiversity pattern of an aggregate of species, one index of phylogenetic similarity, and an implementation of them as linked maps or phylogenies in a plug-in for the increasingly popular open source geographic information system quantum gis. we provide detailed formulae relating three indices of beta diversity, two of pattern of nestedness, one of checkerboard pattern, and two of ratios of variances. the above synthesis is achieved by deriving six vectors from the full presence-absence matrix. our gis implementation is done via web services, tapping the lifemapper platform for estimating potential distributions of species. key words.—biodiversity patterns, beta diversity, presence-absence matrix, phylogeny, lifemapper, arbor. ecologists, macroecologists and biogeographers use a variety of indices and statistics to describe biodiversity patterns. data used to describe a pattern may be continuous or discrete, and among the simplest are presence-absence data. if members of a set of species are present or absent in a set of localities (islands, countries, reserves, or cells in a grid, for example), the presence-absence matrix (pam), is defined as a n sites by s species binary matrix ,i jδ⎡ ⎤= ⎣ ⎦x with elements equal to 1 if species j is present in cell i, and 0 otherwise. the pam contains the data from which one can estimate a variety of metrics of biodiversity pattern. there are metrics that describe the local or alpha species numbers, and their sum, average and spatial variance and covariance (borregaard and rahbek 2010; graves and rahbek 2005; lande 1996; legendre et al. 2005; routledge 1977; schluter and ricklefs 1993; whittaker 1972). there are other numbers that have been used to describe the degree of commonality in the composition of communities, or the overlap between the ranges of distribution of different species (gotelli 2000; schluter 1984; stone and roberts 1990). finally, there are indices related to what is called “nestedness” (almeida-neto et al. 2007; rodríguez-gironés and santamaría 2006; ulrich et al. 2009) which is the degree to which species compositions of smaller communities are proper subsets of larger ones, or equivalently, the degree to which larger distributional ranges contain smaller distributions in a nested way. all of the above indices can be derived from manipulations of the pam and most of the calculations are straightforward. however, these metrics were first presented in the literature for small-sized pams, of the kind often found in ecological problems, or in biogeography when considering just a handful of regions or islands (simberloff and connor 1984). for instance, a typical problem could use tens of species in a few dozen sites, for pams of in the order of 102 to 104 elements. the same metrics, however, can be used to describe patterns at biogeographic extents and much higher resolutions, for instance, in grids of thousands of cells (arita et al. 2008; borregaard and rahbek 2010; orme et al. 2006; rahbek and graves 2000; soberón and ceballos 2011; villalobos et al. 2013). unfortunately, the size of the pams that are encountered when describing whole faunas over gridded regions is often very significant, not unusually of ~106 to 108 elements. the amount of calculations required to estimate some of these indices is large, and testing hypotheses about difference of observed pams using null-models requires randomizations that are beyond the capacity of existing software. moreover, to assemble such large extent, highresolution pams for hundreds or thousands of species one has to resort to very cumbersome manipulations of one of two basic sets of data. in townpeterson typewritten text 22 biodiversity informatics, 10, 2015, pp. 22-34 the first case, it is possible to use “extent of occurrence” (hurlbert and jetz 2007) maps as presented in many biodiversity information initiatives; an excellent example is natureserve1. typically these data come in the form of vectorformatted files (“shapefiles”), one per species, that must be overlaid on a given grid to create a pam of n = number of cells in the grid, and s = number of different species. manipulating and operating large sets of such files is conceptually simple, but practically cumbersome. another possible source of distributional data is species distribution modeling (sdm, franklin 2010), where occurrence data and ecological features are modeled together to get predicted ranges. generally speaking, it is unadvisable to use occurrence data directly due to the biases inherent to this type of data (lira-noriega et al. 2007), and therefore sdm is needed. the typical outputs are raster files that after suitable post processing, including thresholding (liu et al. 2005), can be used to create the presence-absence maps. then individual species maps are overlaid, or “stacked,” to create a pam. stacked maps (whether from sdm or from extent of occurrence datasets) are interpreted as pams under the very obvious hypothesis that interactions between species do not matter. although this assumption is probably false in many cases (araujo and luoto 2007), there are several others in which it may be valid to use it, at least as a first order approximation (peterson et al. 2011) to more realistic procedures. in this work we describe the operations required to estimate the main indices, showing some of their relationships, and then present a practical implementation of these operations, developed as web services. we demonstrate this implementation using the open source gis platform quantum gis. mathematical relations we will show how a number of indices of geographic biodiversity pattern, already published, can be obtained from operations on six objects derived from a pam, denoted by x (see summary in table 1). as well known (gower 1966), the post-multiplication and pre-multiplication of x by its transpose xt yield, respectively, a matrix of                                                                                                                           1  http://www.natureserve.org/conservation-­‐tools/data-­‐maps-­‐ tools/digital-­‐distribution-­‐maps-­‐mammals-­‐western-­‐hemisphere.     shared species in sites, and a matrix of cooccurrence of usage of sites by species. the n x n matrix containing the number of shared species between sites i and h is t=a xx (in the diagonal it contains the number of species in site i= 1, 2 … n). the s x s matrix containing the number of sites shared by species j and k is t=ω x x (the diagonal contains the incidence of species j=1, 2 … s, which for large extents and high resolution of cells approaches the size of the area of distribution). the vectors 1 2( , ,... )t nα α α=α and 1 2( , ,... )t sω ω ω=ω contain the species richness of each site, and the range size of every species. by defining vectors of k-ones: 1(1,1,...1)t k k×=1 , which are used to obtain sums, and using the symbol diag to denote the diagonal of a matrix, we get: ( ) ( ) s t n diag diag = = = = α a x1 ω ω x 1 (1) since ( ) ( )t ttrace trace=x x xx , then the sum of the elements in a is equal to the sum of the elements in ω and equal to the total number of ones in the x matrix, also called the fill, f. two other vectors (of size s and n) contain the total sum of shared range sizes, including the species with itself, and the total number of shared community compositions, including the species number of each site. these are called the mean proportional range-size, and the mean proportional species-diversity (arita et al. 2008; christen and soberón 2009; graves and rahbek 2005) and borregaard and rahbek (2010) refer to the first as the dispersion field. the vectors are defined as: n t s = = = = φ xω a1 ψ x α ω1 (2) their averages are 1 t nn ϕ = φ 1 and 1 t ss ψ = ψ 1 . we use an asterisk to define proportionality with respect to s and n, for example: 1 2* ( / , / ,... / )t ns s sα α α=α and 1 2* ( / , / ,... / )t sn n nω ω ω=ω . townpeterson typewritten text 23 biodiversity informatics, 10, 2015, pp. 22-34 table 1. summary of relationships among different indices of biodiversity pattern. name algebraic definition linear algebra 1 whittaker’s multiplicative beta * 1 wβ ω = ( )w sn trace β = ω 2 lande’s additive beta (1 1/ )a wsβ β= − ( )[1 ]a traces sn β = − ω 3 legendre’s beta ( ) t l ntraceβ = −ω φ 1 4 richness-field of a species , 1 n j i j i i ψ δ α = =∑ t s= =ψ x α ω1 5 dispersion-field of a locality , 1 s i i j j j ϕ δ ω = =∑ n= =φ xω a1 6 matrix of covariance of composition of sites , , 2 1 1( , ) s j k sites j l k l l j k s s α α δ δ = = −∑σ * *1 ( )tsites s = −σ a α α 7 matrix of covariance of ranges of species , , 2 1( , ) n i h sps i j h j j i h i n n ωω δ δ = = −∑σ * *1 ( )tspecies n = −σ ω ω ω 8 mean composition covariance * * 1 j j j w τ α ϕ β − = − 1 *1 wns β −= −τ φ α 9 mean range covariance * * 1 i i i w ρ ω ψ β − = − 1 *1 wns β −= −ρ ψ ω 10 schluter sitescomposition covariance * 2 * / 1/ / w sites w sv n ϕ β β ψ − = − ( ) t sites sites sites v trace = 1 σ 1 σ 11 schluter speciesranges covariance * 2 * / 1/ / w sps w nv s ψ β β ϕ − = − ( ) sps sps vsps trace = 1σ 1 σ 12 wright & reeves’ nestedness 1 1 ( 1) 2 ( ) 2 s c j j j w n n s ω ω ϕ β = = − = − ∑ 1 ( ) 2 t c w nsn β = −φ 1 13 stone & roberts cscore , , 1 2 ( )( ) ( 1) n i i h h i h i h i c s s ω ω ω ω = < ⎡ ⎤ = − −⎢ ⎥− ⎣ ⎦ ∑∑ ( ) 2 t s s nϕ⊗ −1 ω ω 1 2 1 ( ) / / s l w j j ss sn nβ β ω = ⎛ ⎞ = = −⎜ ⎟ ⎝ ⎠ ∑x townpeterson typewritten text 24 biodiversity informatics, 10, 2015, pp. 22-34 finally, the covariance matrix between the species inhabiting sites j and k, and the covariance matrix between ranges of distribution of species h and i are, respectively: * * * * 1 1 1( ) 1 1 1( ) t t sites t t species s s s n n n ⎛ ⎞= − = −⎜ ⎟ ⎝ ⎠ ⎛ ⎞= − = −⎜ ⎟ ⎝ ⎠ σ a α α a αα σ ω ω ω ω ωω (3) notice the use of the symbol small bold sigma (σ ) to denote covariance matrices. there are two matrices in each of the formulae (3), and their traces are useful. the trace of a and of ω equal the fill, as we saw, and 2 2 1 1 ( ) , ( )n st t i ji j trace traceα ω = = = =∑ ∑αα ωω . these sums of squares are: 2 1 2 1 t n t t t t t i s s s s si s t t t t i n n n n nj s n α ψ ω ϕ = = = = = = = = = = = = ∑ ∑ α α 1 x x1 1 ω1 1 ψ ω ω 1 xx 1 1 a1 1 φ (4) it is a matter of a few lines of algebra and substitutions of the above to show that the traces of the two covariance matrices are: * * ( ) ( ) sites w species w ntrace strace ψ β ϕ β = − = − σ σ (5) and that the vectors containing the averages of all their entries are: 2 * 1 1 1 1 1 1 1 sites n t n n w w n n s s fsn s s n β β = ⎛ ⎞= −⎜ ⎟ ⎝ ⎠ = − = − = − τ σ 1 a1 αα 1 τ φ α τ φ α φ α (6) and 2 * 1 1 1 1 1 1 1 species s t s s w w s s n n fsn n n s β β = ⎛ ⎞= −⎜ ⎟ ⎝ ⎠ = − = − = − ρ σ 1 ω1 ωω 1 ρ ψ ω ρ ψ ω ψ ω (7) the above orgy of greek letters and symbols will be used now to find a common thread among a variety of indices of biodiversity pattern, based on the matrices a and ω , and the vectors α , ω , ψ , ϕ , ρ and τ . relations to biodiversity pattern indices whittaker’s beta diversity the two traces in equations (1) are equal, by definition of trace. therefore: ( ) ( )n trace trace s fα ω= = = =a ω (8), which immediately implies ns α ω = . whittaker, (1972) defined his first measure of beta diversity as / /w s ns fβ α= = , which means that whittaker’s beta is mathematically equivalent to the reciprocal of the average proportion of the total region (n) occupied by the species inhabiting the region (arita and rodríguez 2002; routledge 1977; schluter and ricklefs 1993; vellend 2001). both α and wβ are invariant to randomization of the matrix values subject to the condition that the total count of presences and the dimensions of the matrix remain constant, as is obvious from (1), because both traces are simply the fill of the x matrix (the total number of ones). for this reason wβ is not a measure of turnover (in which case it would be sensitive to the actual physical positions of presences and absences), but rather it is a simple measure of how much more diverse is an entire collection than the average of its subsets (vellend 2001). townpeterson typewritten text 25 biodiversity informatics, 10, 2015, pp. 22-34 lande’s beta diversity another measure of beta diversity, the socalled additive beta diversity, (lande 1996) is being used again (veech et al. 2002): a sβ α= − . for presence-absence matrices this measure is related very simply to whittaker’s multiplicative measure by (1 1/ )a wsβ β= − (6) this equation is obtained by substitution of the value of α in the additive beta formula (lande 1996; veech et al. 2002; ricotta 2005), as βw, and for the same reason, if the dimensions of the pam and its fill remain constant, βa are insensitive to permutations of ones and zeroes. legendre’s beta diversity legendre et al. (2005), proposed as a new measure of beta diversity: the total sum of squares, or ss(x) = βl, of the species-composition matrix. this definition is inspired from equating the concept of beta diversity to “variation in species composition among sites…” (legendre et al. 2005) and is made formal by use of equation (1) of legendre et al. (2005) which shows that ss(x) is mathematically equivalent to the sum of squared euclidean distances among sites, divided by n, the number of sites. the total sum of squares mentioned in legendre et al. (2005) is, by definition, the trace of the species variance-covariance matrix: *( )l species w straceβ ϕ β = = −σ (7) schluter’s variance ratios schluter (1984) introduced a ratio of variances test to assess simultaneously whether species in a group are associated. the method compares the observed variance in the total number of species in samples, with the variance expected if occurrence of each species is independent of the others. for presence-absence data there are two such ratios, one to test the association of species in communities, and another to test overlap of ranges of distribution. in the notation of schluter, the rowsums, which we denote as αj are called tj. the column sums (ωj), in schluter (1984) are denoted by ni. making these substitutions and using the identities (4) yields: 2 * 2 1 ** * 1 2 * 2 1 ** * 1 1( ) ( ) / 1/ /(1 ) 1( ) ( ) / 1/ /(1 ) n ii w com s wj jj s jj w range n wi ji snv n nsv s α α ϕ β β ψω ω ω ω ψ β β ϕα α = = = = − − = = −− − − = = −− ∑ ∑ ∑ ∑ nestedness one of the earliest indices of nestedness of a pam is that of wright and reeves (1992) which depends only on the marginal values of the pam, and is defined, with the notation of this paper, as ( ) ( ) 1 2 1 1 1 ( 1) 2 1 2 1 2 2 j s c j jj s s jj j w n n f n s ω ω ω ω ϕ ϕ β = = = = − = − = − ⎛ ⎞ = −⎜ ⎟ ⎝ ⎠ ∑ ∑ ∑ a more recent index proposed by almeidaneto et al. (2008) is based on the idea that both nestedness of community-composition and of distributional ranges should be taken into account. they propose ordering a matrix by column and row marginals, and then comparing every pair i, h of rows such that αi > αh, and every pair of columns such that ωj > ωk, adding all the corresponding values of a and ω and dividing by the number of compared pairs. this index is more an algorithm than a formula, but it can easily be expressed in terms of the mathematical objects defined above. c-scores in 1990, stone and roberts (1990) addressed the problem of whether the distributions of pairs of species are “random” in the sense of not being different of what they would be if the species did not interact. their c index is defined using the townpeterson typewritten text 26 biodiversity informatics, 10, 2015, pp. 22-34 number of co-occurrences (entries in the ω matrix above). first they looked for “checkerboard units,” which are patterns of the form {1, 0}, {0, 1}, i.e., for two species, pairs of sites where they do not cooccur. using the notation of this paper, they define the number of checkerboard units as (stone and roberts 1990): , , ,( )( )j k j j k k j kc ω ω ω ω= − − and their c index is then , 1 2 ( 1) s j k j k j c c s s = > = − ∑∑ the above expands to three sums that can be shown to be, using all the above definitions and relationships: 2 1 , 1 2 , 1 ( ) ; 2 ( ) ( )( ) 2 s j k j k j s t j k j k j k j ts s s j k j k j s n n n ω ϕ ω ω ω ω ω ϕ ϕ ω = > = > = > − = + = − ⊗ − = ∑∑ ∑∑ ∑∑ ω ψ 1 ω ω 1 where the symbol ⊗ represents the hadamard product (element by element) of two compatible matrices. when the only available information is presence-absence, probably most other indices of biodiversity pattern can be reduced to operations between the a and ω matrices, or the vectors describing marginals (α and ω), covariances (ρ and τ ) or the closely related fields (ψ and ϕ). mathematical constraints the relationships among the above different measures of pattern are constrained by their mathematical properties (soberón and ceballos 2011), and the constrained graphs are very useful to display in an aggregated way the properties of biodiversity patterns (arita et al. 2008). moreover, when these graphs are linked to the corresponding cartographic information in such a way that “brushing” (in gis terminology) different parts of the graph highlights the corresponding parts areas in a map and vice versa, very interesting relationships are revealed, with biogeographic or conservation implications (soberón and ceballos 2011; villalobos et al. 2013). in the remainder of the paper, we describe a software implementation that allows inspection of scatterplots among pairs of indices. our software allows highlighting jointly the locations in maps where given combinations occur, and positions in the scatterplot. moreover, although we do not yet provide a theoretical background, our software also include indices of phylogenetic proximity, enabling in this way a linked perspective of geographic and phylogenetic structure and covariance among the corresponding indices. statistical testing is enabled by providing fast bootstrapping for large matrices. our software is the first implementation that makes full use of the theory described above. software implementation to implement the above, we extended the platform lifemapper 2 to include multi-species range and diversity experiment construction, calculations and hypothesis testing for pams in a module called lmrad (lifemapper range and diversity). the platform's services are made available through a customized lmrad plug-in for qgis that includes point and click functionality for building and analyzing pams in user-defined regions. this plug-in is already implemented and available in qgis. linked custom data visualization spaces are also implemented inside qgis. we will describe the details of using the plugin in a forthcoming paper. to see how the linkages work consider that the statistics derived from the dispersal field (ϕ) and richness (α) vectors are measures attached to each geographic locality and thus can be linked to a map using the geographic coordinates of the centroids of the localities (corresponding to the pam rows). on the other hand, statistics based on the diversity field ψ and the range size vector ω result in measures attached to each species in a pam, so that the column space of the pam can be linked to a phylogenetic tree, using the names as common field. scatter plots can be species-based (ω , ψ) or site-based (α and ω), (arita et al. 2008). in both cases, the dispersion of points in the plots is                                                                                                                           2 http://www.lifemapper.org. townpeterson typewritten text 27 biodiversity informatics, 10, 2015, pp. 22-34 determined by patterns of species' co-occurrence, and the key idea is to link plots of site-based, or species-based statistics to maps or phylogenetic trees respect-tively. we plot such associations in qgis through the plug-in and link to a custom phylogenetic viewer built in qgis, thus allowing a display of site-based and species-based statistics in plots that are dynamically linked to the map or the tree. for instance, in figure 1, we plot the proportional number of species vs. the mean proportional range-size (both site-based numbers). brushing points in this plot highlights sites in the pam maps that show proportional range size similarity for the species in those sites and speciesnumbers similarity among the sites. the lifemapper qgis tool also provides tools for mapping the entire site-based statistics from the pam so that biogeographic patterns in the maps can guide brushing. layering the statistics with any number of gis layers for the areas of interest can add to the visual analysis. the phylogenetic statistics from the tree for the species are also aggregated spatially in the tool and can show interesting patterns. the tool also links the shared community composition of species from the pam to taxon distance statistics derived from the tree in a correlation coefficient of taxon distance to sites shared between species. the site brushing is multidirectional and can go from map-to-plot and plotto-map (see figure 1). the site-based indices that we described above have a natural correspondence to maps because essentially every cell in the grid has a value for all the indices. for species-based analysis the natural corresponding structure would be a tree. the phylogenetic data structure that drives the tree visualization is used to calculate dynamically the mean nearest taxon distance for selected species in the species association plot. in this way the spatially derived statistics for diversity can be compared to the degree of phylogenetic relatedness within species communities. individual species from those communities can also be sub-selected and their ranges from the pam shown in the pam based maps. selections by clade directly in the tree will vice-versa select those points in the plot (see figure 2). linkages and visual analysis of range-diversity relationships derived from very large pams are achieved through custom visualization spaces inside of qgis, but the construction and the outputs from calculations on the pams (exceeding 108 elements) are achieved by compute modules that interact with the visualization environment in a client-server relationship through web-services architecture. compute services for these very large jobs are exposed as open geospatial consortium (ogc)3 web processing services, and restful web-services. this scheme permits to spread the computational load for working with thousands of pam inputs and calculations using node parallelization and data parallelization across remote distributed computing environments so that dealing with large matrix operations is not dependent on desktop resources, thus freeing the qgis tool to do what it does best, preparing spatial inputs and visualizing the outputs spatially to discern biogeographic patterns. the advantages of using qgis was to leverage an entry level but powerful and extensible gis tool for users to be able to work with spatial data, prepare them, do spatial analyses on outputs, and the tools to define and upload lmrad experiments through the tool. the lifemapper infrastructure is composed of a central management component, lmdbserver, which manages data and analysis operations with a “data pipeline” written in python 4 and a postgresql/postgis database; multiple instances of lmcompute that can be co-located across institutions (currently this is deployed at compute clusters at university of kansas, the university of florida, and san diego supercomputer center); and lmwebserver which manages all communications between the components and client applications. (see figure 3.) lmrad is a job-based infrastructure that is environmentally agnostic and its algorithms are portable across compute environments through configurable instances of lmcompute. lmwebserver contains a job server tier that feeds compute jobs to any compute environment that can sponsor an instance of lmcompute. the compute plug-in for a specific resource receives compute jobs for pams through a job controller which determines which of several plugins are appropriate for the type of calculation. the pipeline and lmdbserver are responsible for presenting jobs to the job server and moving jobs through the                                                                                                                           3 http://www.opengeospatial.org. 4 http://www.python.org.     townpeterson typewritten text 28 biodiversity informatics, 10, 2015, pp. 22-34 figure 1. a screencapture of the qgis lifemapper plugin showing a sites-based output. site similarity plot for amphibians of the philippines, showing highlands in luzon (a, in yellow). the value of the mean proportional range size (b), the emerged relief during the glacial maximum (c), and the “brushed” luzon cells in the scatterplot of richness vs mean proportional rangesize (d). figure 2. a species-based screen-capture of the qgis plugin showing part of a mammal phylogeny (a) connected to a 800+ mammals of africa pam, with a map of species richness (b), mean proportional species-diversity (c), mean nearest taxon distance (d) and a scatterplot of range-size vs. mean phylogenetic distance (e) with the “brushed” species highlighted (yellow species in (a) and (e)). (d) (c) (b) (a) (a) (b) (c) (d) (e) townpeterson typewritten text 29 biodiversity informatics, 10, 2015, pp. 22-34 system. at different stages in a lmrad experiment, dependencies and statuses are updated by the compute environment which posts back to the job server. pam construction has been parallelized across processors on any compute environment that receives a pam job. data products for large pams (extents >106 km2, resolutions 10 km or less, >103 species) can be constructed and analyzed in this way with reasonable response times. results from the experiment are then posted back to the job server from the compute environment and are written to the database and file system shared by the lmdbserver and lmwebserver. data parallelization across multi-core architectures on each of the nodes in a compute cluster helps to speed large pam construction jobs. pam construction uses a combination of rtree 5 and matplotlib's nxutils and gdal 6 for vector and raster based intersections, respectively. calculations on the matrices use numpy7 built with the basic linear algebra subprograms (blas). permutations on the pam matrices for hypothesis testing against null models use methods that are specific to binary matrices, where row and column totals can both be kept intact while changing the mix of species in sites and the range size of each species. data parallelization is not suited to these computations since the entire matrix needs to be taken into account. however, since several hundred permutations may be required per experiment, the current job based parallelization across compute nodes works well for computing these models in toto. another method in lmrad for permuting the matrix is perfectly suited for both types of parallelization. it uses a dye dispersion algorithm which is a 2-dimensional geometric-constraints model that assumes range continuity (jetz and rahbek 2001). since range allocations are reassembled individually for each species, those data can be split across cores on a single machine or across nodes. the lifemapper plugin allows qgis to operate as a web service client to the architecture described above, edit and submit data, parameterize inputs and request computations. experiment results can be pulled down as statistical and geospatial outputs and linked to phylogenetic trees and range                                                                                                                           5 https://pypi.python.org/pypi/rtree/. 6 http://www.gdal.org/. 7 http://www.numpy.org/.   diversity plots that depict the 6 vectors described above, species richness and range size vectors, whitakker's beta, legendre's beta, and lande additive beta. range-diversity plots are produced in the plug-in that summarize these fields as indexes of site similarity and the degree of association of species, allowing the user to experiment across scale, geographic extent and pam grid resolution. the range-diversity plots and the tree viewer are custom visualization spaces built for the plug-in inside of qgis something made possible because the user community for qgis is free to customize qgis through plug-in development to do a wide variety of analysis with access to all of the qgis functionality through a python api. an example of a unique environment for qgis is the tree viewer for lmrad. the tree viewer presents the phylogenetic data as interactive svg built dynamically from incrementally loaded javascript object notation (json) data using web-based techniques in a document driven javascript framework. using advances in webbased javascript visualization libraries alleviates the need for the user to install external libraries when installing the plugin. there are several tree formats, e.g. phyloxml 8 , newick 9 , nexus (maddison et al. 1997), nexml10 and nexson, used by the open tree of life11 for connections to web-services. since these allow construction of tree databases that are either json or are easily translated into json, they can be directly mapped to python dictionaries, and are easily transported back and forth from lmcompute for analysis and they are ideal for a document driven visualization framework. visualization then is made possible with the javascript library for data driven documents (d3) d3.js12 . d3 allows the json document to be dynamically bound to the document object model so that data-driven transformations can be applied to the document with smooth transitions and fluid interaction. such smooth transitions are especially useful when navigating large trees with many nodes and edges. the d3 based interactive tree is rendered in the plug-in through a qt dialog using qtwebkit.                                                                                                                           8 www.phyloxml.org. 9 http://bit.ly/1n6elcz. 10 http://nexml.org. 11 http://blog.opentreeoflife.org/. 12 http://d3js.org/.   townpeterson typewritten text 30 biodiversity informatics, 10, 2015, pp. 22-34 communication between the tree and the rest of the plug-in is effected by qtwebkit bridge. the bridge allows the javascript and pyqt objects to communicate with one another. the tree viewer is linked to the interactive range-diversity plots in matplotlib (hunter 2007) by simple pyqt signals and slots. a similar method connects the rangediversity plots for site-based statistics to the maps in qgis based on the pam. using javascript in pyqt dialogs for qgis allowed us to achieve fluid visual representations of trees for large clades, e.g. one tree used in testing is the entire phylogeny for the phylum mollusca with over 85,000 nodes. discussion availability of biodiversity data is increasing very rapidly allowing researchers, in principle, to perform a large number of analyses. however, as long as the different perspectives (phylogenetic, ecologic, morphologic, functional and others, see maclaurin & sterelny (2008)) remain unlinked, the analyses are mostly disconnected from one another. in fact, it is possible to say that the fundamental task of biodiversity informatics should be to enable “integration” of different perspectives of biodiversity (harmon et al. 2013; miller and jolley-rogers 2014; peterson et al. 2010). integration is not a well-defined term, but one possible meaning may be the capacity to display simultaneously different perspectives of the data (laffan et al. 2010), preferably in a linked way. for instance, in figure 2, a simultaneous and linked display of phylogenetic, ecological and geographical data is presented. software has been developed that links geographic display of data with a variety of mathematical graphs, but very few integrated systems exist that address biogeography, community assembly, ecological niche and phylogeny. web-based solutions for viewing phylogenies are popular but are limiting in that geospatial tools for the web that allow ad hoc analysis of range data as character traits for species in trees struggle to keep pace with desktop gis implementations. most of these implementations choose to focus on simple clade-area relationships. a few of these contain similar analysis to lmrad and have spatial components that can be compared with our tools. spatial analysis in macroecology (rangel et al. 2010) is a software alternative for biodiversity experiments that directly influenced lmrad. sam offers a comprehensive set of tools for spatial statistics, some simple mapping tools and advanced spatial autoregression models. it uses extremely optimized linear algebra libraries for large matrix operations. the data table in sam accommodates pam data, and can be formatted as figure 3. general architecture of lifemapper and lmrad.   townpeterson typewritten text 31 biodiversity informatics, 10, 2015, pp. 22-34 esri shapefiles. just as in lmrad, data grids can be prepared in sam at any extent and resolution as shapefiles and pams can be generated directly from the shapefile. addition-ally, as in lmrad, sam allows a user to link scatterplots and maps geographically, where grid cells can be selected in a map and then correspondingly highlighted in a scatter plot or vice versa, allowing a user to detect outliers. sam also has sophisticated tools for evaluating the changes in spatial correlation as they are affected by changes in scale. on the other hand, sam does not incorporate phylogenetic analysis or visualization into its data linkages and it is a windows-dependent desktop software reliant on the processing capacity of the desktop, where the lmrad plugin is cross-platform and interfaces with remote wps services so that little processing occurs on the user's machine. ecosim (gotelli and entsminger 2011) is another windows-based macroecology software built specifically for dealing with null model hypotheses testing and pam data. it provides a variety of randomization routines for each of its modules. the co-occurence module randomization routine allows row and column constraints, including fixed-sum similar to lmrad. ecosim also allows equiprobable, proportional, and weighted constraints (gotelli and entsminger 2011). lmrad currently has two randomization algorithms. ecosim has four different ways of dealing with sparse or degenerate matrices (with empty rows and columns; (ellison 2000). lmrad currently has a compression algorithm for compressing and re-expanding such matrices which speeds processing time. the limiting factor for ecosim seems to be the size of the pam; 240,000 cells, or approximately 800 by 300 rows and columns is an absolute limit. one of the core requirements first addressed by lmrad was being able to work with much larger matrices. initial tests in lmrad for randomization algorithms were done for matrices on the order of 6.0 x 108 cells. data products for large pams at high resolutions (10 km) with upwards of 800 species can be constructed and analyzed with reasonable response times in lmrad. biodiverse (laffan et al. 2010), an open-source project similar to the lifemapper plug-in, provides linked visualization across different dataspaces. biodiverse links species distributions in geographic, phylogenetic, taxonomic and environmental spaces. one advantage of biodiverse is that scale comparisons are achieved through a window analysis for endemism, phylogenetic diversity, and beta diversity. by varying the size of the windows one can start to understand the effects of scale on those statistics. currently the lifemapper plug-in uses a multi-grid approach where several subsets at different cell resolution can be built out within the same experiment, allowing comparisons across scale for the range and diversity statistics including beta diversity. opengeoda is a free and open source crosssplatform package for exploratory spatial analysis. its strengths are techniques for dynamic linking and brushing data across multiple data spaces (anselin et al. 2005). it is like lmrad in this respect and strives for integration across different measures of spatial association rather than the specific biodiversity focus in lmrad. opengeoda then represents a very powerful set of tools found in mature desktop gis applications that are integrated beyond what most gis applications provide. unlike most of the mentioned software, lmrad is natively oriented to assembling and organizing very large datasets, and, finally, because we provide a simple but robust set of mathematical formulae that uncover the relationships and constraints among many of the main biodiversity indices lmrad is unique in its integration of species relatedness and spatial components for range, diversity fields and dispersion fields. the next step in developing the tool is to allow for linking other perspectives of biodiversity as high-end web services, a task already under way (harmon et al. 2013). nevertheless, the future development of “integrating” views of biodiversity goes well beyond linking displays. ideally, integration should also strive to express theoretical relationships among different views of biodiversity, preferably in a statistical or mathematical way. it is already possible to analyze statistically multiple views of biodiversity, as it has been demonstrated by a number of authors (doledec et al. 2000; doledec et al. 1996; leibold et al. 2010). in this way a fuller understanding of how different aspects of biodiversity are related, and how, can be achieved. townpeterson typewritten text 32 biodiversity informatics, 10, 2015, pp. 22-34 acknowledgments nsf/bio/avatol award 1208472 supported both authors. we are grateful to our colleagues h. arita, p. rodríguez, f. villalobos and a. lira for endless conversations on the issue of biodiversity patterns, and to town peterson for encouragement and advice, and for his time helping us with editorial issues. literature cited almeida-neto, m., p. guimarães, p. r. guimarães, r. d. loyola, and w. ulrich. 2008. a consistent metric for nestedness analysis in ecological systems: reconciling concept and measurement. oikos 117:1227-1239. almeida-neto, m., p. r. guimarães, and t. m. lewinsohn. 2007. on nestedness analyses: rethinking matrix temperature and anti-nestedness. oikos 116:716-722. anselin, l., i. syabri, and y. kho. 2005. geoda: an introduction to spatial data analysis. geogr. anal. 38:5-22. araujo, m., and m. luoto. 2007. the importance of biotic interactions for modelling species distributions under climate change. glob. ecol. biogeogr. 16:743-753. arita, h., and p. rodríguez. 2002. geographic range, turnover rate and scaling of species diversity. ecography 25:541-550. arita, h. t., j. a. christen, p. rodríguez, and j. soberón. 2008. species diversity and distribution in presence-absence matrices: mathematical relationships and biological implications. amer. nat. 172:519-532. borregaard, m. k., and c. rahbek. 2010. dispersion fields, diversity fields and null models: uniting range sizes and species richness. ecography 33:402-407. christen, a., and j. soberón. 2009. anidamiento y los análisis rq y qr en pams. misc. mat. 49:51-61. doledec, s., d. chessel, and c. gimaret-carpentier. 2000. niche separation in community analysis: a new method. ecology 81:2914-2927. doledec, s., d. chessel, c. j. f. ter braak, and s. champeley. 1996. matching species traits to environmental variables: a new three-table ordination method. environ. ecol. stat. 3:143-166. ellison, a. m. 2000. ecosim: null models software for ecology. bull. ecol. soc. amer. 81:125-127. gotelli, n. j. 2000. null models analysis of species cooccurrence patterns. ecology 81:2602-2621. gotelli, n. j., and g. l. entsminger. 2011. ecosim: null models software for ecology. version 7.0. acquired intelligence inc. & kessey-bear, http://homepages.together.net/~gentsmin/ecosim.ht m. gower, j. c. 1966. some distance properties of latent root and vector methods used in multivariate analysis. biometrika 53:325-338. graves, g., and c. rahbek. 2005. source pool geometry and the assembly of continental avifaunas. proc. nat. acad. sci. usa 102:7871-7876. harmon, l., j. baumes, c. hughes, j. soberón, c. d. specht, w. turner, c. lisle, and r. w. thacker. 2013. arbor: comparative analysis workflows for the tree of life. plos currents 5 hunter, j. d. 2007. matplotlib: a 2d graphics environment. compu. sci. eng. 9:90-95. hurlbert, a. h., and w. jetz. 2007. species richness, hotspots, and the scale dependence of range maps in ecology and conservation. proc. nat. acad. sci. usa 104:13384-13389. jetz, w., and c. rahbek. 2001. geometric constraints explain much of the species richness pattern in african birds. proc. nat. acad. sci. usa 98:56615666. laffan, s. w., e. lubarsky, and d. f. rosauer. 2010. biodiverse, a tool for the spatial analysis of biological and related diversity. ecography 33:643647. lande, r. 1996. statistics and partitioning of species diversity, and similarity among multiple communities. oikos 76 5-13. legendre, p., d. borcard, and p. peres-neto. 2005. analyzing beta diversity: partitioning the spatial variation of community composition data. ecol. monogr. 75:435-450. leibold, m. a., e. p. economo, and p. peres-neto. 2010. metacommunity phylogenetics: separating the roles of environmental filters and historical biogeography. ecol. lett. 13:1290-1299. lira-noriega, a., j. soberon, a. g. navarro-siguenza, y. nakazawa, and a. t. peterson. 2007. scale dependency of diversity components estimated from primary biodiversity data and distribution maps. diversity distrib. 13:185-195. liu, c., p. m. berry, t. p. dawson, and r. g. pearson. 2005. selecting thresholds of occurrence in the prediction of species distributions. ecography 28:385-393. maclaurin, j., and k. sterelny. 2008. what is biodiversity? university of chicago press, chicago. maddison, d. r., d. l. swofford, and w. p. maddison. 1997. nexus: an extensible file format for systematic information. syst. biol. 46:590-621. miller, j. t., and g. jolley-rogers. 2014. correcting the disconnect between phylogenetics and biodiversity informatics. zootaxa 3754:195-200. townpeterson typewritten text 33 biodiversity informatics, 10, 2015, pp. 22-34 orme, d. l., r. davies, v. a. olson, g. h. thomas, t.t. ding, p. c. rasmussen, r. s. ridgeley, a. stattersfield, p. m. bennett, i. p. f. owens, t. m. blackburn, and j. k. gaston. 2006. global patterns of geographic range size. plos biol. 4:1276-1283. peterson, a. t., s. knapp, r. guralnick, j. soberon, and m. t. holder. 2010. the big questions for biodiversity informatics. syst. biodivers. 8:159168. peterson, a. t., j. soberón, r. g. pearson, r. anderson, e. martínez-meyer, m. nakamura, and m. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. rahbek, c., and g. graves. 2000. detection of macroecological patterns in south american hummingbirds is affected by spatial scale. proc. roy. soc. lond. b 267:2259-2265. rangel, t. f., j. a. diniz-filho, and m. bini. 2010. sam: a comprehensive application for spatial analysis in macroecology. ecography 33:46-50. ricotta, c. 2005. on hierarchical diversity decomposition. j. veg. sci. 16:223-226. rodríguez-gironés, m. a., and l. santamaría. 2006. a new algorithm to calculate the nestedness temperature of presence-absence matrices. j. biogeogr. 33:924-935. routledge, r. d. 1977. on whittaker's components of diversity. ecology 58 1120-1127. schluter, d. 1984. a variance test for detecting species associations, with some example applications. ecology 65:998-1005. schluter, d., and r. e. ricklefs. 1993. species diversity: an introduction to the problem. pp. 1-12 in r. e. ricklefs and d. schluter, eds. species diversity in ecological communities: historical and geographical perspectives. university of chicago press, chicago. simberloff, d., and e. f. connor. 1984. inferring competition from biogeographic data: a reply to wright and biehl. amer. nat. 124:429-436. soberón, j., and g. ceballos. 2011. species richness and range size of the terrestrial mammals of the world: biological signal within mathematical constraints. plos one 6:e19359. stone, l., and a. roberts. 1990. the checkerboard score and species distributions. oecologia 85:7479. ulrich, w., m. almeida-neto, and n. j. gotelli. 2009. a consumer's guide to nestedness analysis. oikos 118:3-17. veech, j. a., k. summerville, t. crist, and j. gering. 2002. the additive partitioning of species diversity: recent revival of an old idea. oikos 99:3-9. vellend, m. 2001. do commonly used indices of betadiversity measure species turnover? j. veg. sci. 12 545-552. villalobos, f., a. lira-noriega, j. soberón, and h. t. arita. 2013. range–diversity plots for conservation assessments: using richness and rarity in priority setting. biol. cons. 158:313-320. whittaker, r. h. 1972. evolution and measurement of species diversity. taxon 21:213-251. wright, d. h., and j. reeves. 1992. on the meaning and measurment of nestedness of species aseemblages. oecologia 92:1432-1939. townpeterson typewritten text 34 microsoft word emily_20160707.docx biodiversity informatics, 11, 2016, pp. 12-22 12 digital knowledge of kenyan succulent flora and priorities for future inventory and documentation emily wabuyele1, 2, *, simon kang’ethe3, 4 and leonard e. newton2 1national museums of kenya, p.o box 40658, nairobi 00100, kenya; 2kenyatta university, p.o box 43844, nairobi 00100, kenya; 3east african herbarium, p.o box 46166, nairobi 00100, kenya; 4world agroforestry centre (icraf), united nations avenue, gigiri, p. o. box 30677 nairobi 00100 kenya abstract.—biodiversity inventory in kenya has been ongoing for about a century and a half, coinciding with the arrival of naturalists from europe, america, and elsewhere outside africa. since the first collections in the mid-to-late 1800s, there has been a steady increase of plant surveys, frequency of inventory, and discovery of new species that have considerably increased knowledge of faunal and floristic elements. however, as in all other countries, such historical biological collection activities are more often than not, ad hoc, resulting in gaps in knowledge of species and their habitats. while kenya is relatively rich botanically, with a succulent flora of about 428 taxa, it is apparent that the list is understated owing to, among other factors, difficulty of preparing herbarium material and restricted access to some sites. this study investigated completeness of geographic knowledge of succulent plants in kenya, with the aim of establishing species distribution patterns and identifying gaps that will guide and justify priority setting for future work on the group. species data were filtered from the general brahms database at the east african herbarium and cleaned via an iterative series of inspections and visualizations designed to detect and document inconsistencies in taxonomic concepts, geographic coordinates, and dates of collection. eight grid squares fulfilled criteria for completeness of inventory: one in the city of mombasa, one in the kulal–nyiro complex, one in garissa, one in baringo, and four grid squares in the nairobi–nakuru–laikipia area. poorly-known areas, mostly in the west, north, and north-eastern regions of the country, were extremely isolated from wellknown sites, both geographically and environmentally. these localities should be prioritised for future inventory as they are likely to yield species new to science, species new to the national flora, and/or contribute new knowledge on habitats. to avoid inconsistencies and data leakage, biodiversity inventory and documentation needs streamlining to generate standardised metadata that should be digitised to enhance access and synthesis. key words: succulent plants, well-known sites, data leakage, inventory, species diversity kenya lies astride the equator, and covers a total land area of 582,646 km2 at latitudes of 5°n– 5°s and longitudes 34–41°e along the northeastern seaboard of africa. the country’s topography varies from sea level in the coastal zone with coral reefs and mangrove swamps; the rift valley lake system; the low-lying plains and chalbi desert in the north and northeast; to scattered upland massifs such as mount elgon, the mau escarpment, cherangani hills, and the aberdare ranges, including the highest point at 5194 m on mount kenya (lucas, 1968; mewnr and rda, 2015). the climate across kenya is heavily influenced by elevation, its equatorial location, proximity to the indian ocean to the east, lake victoria to the west, and the central highlands. average annual rainfall and seasonality vary widely with elevation; a meagre 150 mm are received in the low-lying drylands in the north and northeast, while over 2500 mm are received at high elevations such as on the slopes of mount kenya. two rainy seasons are experienced each year: the short rains from october to december, and long rains from march to may. temperatures vary with relief, season, rainfall, and cloud cover. the northern and eastern lowlands reach maximum average temperatures of more than 35° c and the central highlands of less than 18° c; sub-zero temperatures are experienced during the nights in the alpine zones of the highest peaks (lucas, 1968). phytogeographically, the country sits at the confluence of five major regions of white (1983); the zone between the rift valley and the coastal belt belongs to the somalia masai regional centre biodiversity informatics, 11, 2016, pp. 12-22 of endemism (rce), an extensive region that is only occasionally broken by the afro-montane archipelago-like rce to which the highlands on both sides of the rift valley belong. the western parts of the country are split between the sudanian rce and the lake victoria regional mosaic (rm), whereas the coastal strip belongs to the zanzibar-inhambane rm. in tandem with climatic patterns of kenya, the vegetation ranges from almost bare rock and sand dunes in the arid zones, through acacia-commiphora bushland to grassland with scattered trees, dry highland forests, tropical rainforests, and alpine vegetation at high elevations (maundu et al., 1999; beentje and smith, 2001). although not as diverse biologically as countries in the american or asian tropics, kenya holds a large diversity of species, estimated at 35,000 taxa of plants, animals, and microorganisms, many of which are endemic. this diversity has been attributed to long periods of geological stability, as well as the variety of microhabitats that exist across the country (hepper, 1979; nordal et al., 2001). a recent analysis of biodiversity hotspots for kenya by the ministry of environment, water and natural resource, estimated plant diversity at just over 7000 species of vascular plants, belonging to 1720 genera and 240 families (mewnr and rda, 2015). according to that report, most species are concentrated in three key areas that harbour exceptionally high plant diversity: mount elgon in the west, nairobi and its environs in the central highlands, and the coastal strip bordering the indian ocean, with 650–950 species per square degree of area. isolated mountain peaks such as marsabit and kulal in the northeast are rich in endemic species. the coastal forests and some of the isolated mountains in southeastern kenya are part of the eastern arc and coastal forests hotspot and the afromontane hotspots, respectively (myers et al., 2000; mittermeier et al. 2005). to a large extent, knowledge of biodiversity in general and plants in particular is dictated by the intensity of field surveys and inventory. for instance kuper et al. (2006) noted that, despite a relatively good taxonomic understanding of vascular plants, knowledge of geographic distribution of species is generally poor, even in large herbaria where 75% of the species are represented by less than 10 collection records. as noted by colwell and coddington (1994), traditional collection methods used in biodiversity surveys for museums and herbaria may intend to collect all species, but such a goal is neither easy to attain nor to monitor and measure. as such, the geographic distribution of tropical plant and animal diversity is still poorly documented, especially at spatial resolutions of practical use for conservation. as with neighbouring countries in the region, biodiversity inventory in kenya has been ongoing for just about a century and a half, coinciding with the arrival of naturalists from europe, north america and elsewhere outside of africa. since acquisition of the earliest collections on expeditions in the mid-to-late 1800s (newton, 2004a), plant collections and surveys have shown a steady increase, with discoveries of new species and documentation of occurrence patterns that have increased knowledge of faunal and floristic elements. however, preliminary analyses of field sampling patterns suggest bias towards collecting in accessible localities that are not necessarily the most important biodiversity hotspots of the country (beentje and smith, 2001). furthermore, most of the information remains inaccessible beyond the immediate research and scientific communities, effectively inaccessible in analogue formats in museum and herbarium cabinets: only 17-20% of specimens worldwide are available in sharable formats (ariño, 2010), further aggravating the problem of inadequate species inventory. of the total floristic diversity in kenya, about 5% of kenyan species are succulent in nature. newton (2003) listed 428 taxa, representing 370 species, 51 genera, and 20 families, of succulent plants in the country. this concentration of species is high compared to global totals of ~10,000 species (el-ghanie et al., 2014). as reported by sajeva and costanzo (1994), succulence is found in over 30 plant families globally, the largest being aizoaceae, cactaceae, crassulaceae, euphorbiaceae, apocynaceae, agavaceae, asphodelaceae, (now xanthorrhoeaceae), chenopodiaceae, and portulacaceae. in africa, succulent plants are distributed in most vegetation types except the tall forests of west africa and the miombo woodlands (oldfield, 1997). like the rest of the national and regional flora, conservation of kenyan succulent plants is hampered by a suite of knowledge gaps, including townpeterson typewritten text townpeterson typewritten text 13 biodiversity informatics, 11, 2016, pp. 12-22 14 difficulties associated with description, identification and general taxonomic documentation of species. preparation of specimens of succulent plant species requires additional skill and patience (bridson and forman, 1989; newton, 1995), which more often than not are inadequate amongst general plant survey and inventory teams. in addition, probable key sites of diversity and endemism for succulent plant species (the arid and semi-arid regions) are hard to access, both logistically and politically. succulent species nonetheless deserve special attention, as they grow in habitats with a fragile ecological equilibrium where harsh environmental conditions lead to slow rates of growth and low reproductive rates, including seed viability as low as 0.1% for some species (sajeva and costanzo, 1994). in addition, owing to commercial interest, unregulated removal of mature plants may result in unsustainable reproductive rates and accelerate extinction of species. in this paper, we therefore hypothesize that the kenyan list of succulent taxa is understated, and far from complete, owing to, among other factors, difficulty of preparation of herbarium material and restricted access to key sites. we explore a relatively detailed data set that documents known occurrences of succulent plant species, using techniques designed to assess completeness of geographic knowledge of succulent plants in kenya. our aim is to establish species’ distribution patterns and identify knowledge gaps that can guide and justify priority setting for future work on this iconic functional group. methods the east african herbarium (ea), where large numbers of specimens and associated data records for kenya, uganda, and tanzania are housed, has over the last two decades made progress in digitisation of its vascular plant specimen data as part of routine information management activities. to date, information has been captured from >200,000 specimens using the botanical research and herbarium management software1. apart from records associated with specimens housed at the ea, the database includes >10,000 records from the kew herbarium that have been repatriated as part of collaborative data 1 http://herbaria.plants.ox.ac.uk/bol/brahms/software. mobilisation initiatives. data corresponding to succulent plant species were prioritised for digitisation as a thematic group of interest in conservation and trade across the region; as such, we estimate that 90% of succulent plant records at the ea have now been captured and are treated in this report. general information on species numbers for the country was collated from existing published sources. prior to analysis, data were cleaned via an iterative series of inspections and visualizations designed to detect and document inconsistencies. (1) we explored concentrations of sampling by calculating species densities based on raw specimen counts. (2) we created lists of unique names in excel, and inspected them for repeated versions of the same taxonomic concepts: misspellings, name variants, different versions of authority information, etc. such duplicate names were flagged, checked via independent sources, and corrected to produce single scientific names that correctly referred to single species taxa. (3) we checked for geographic coordinates that fell outside of the country, but which were referred to as falling within kenya. (4) within the country, we checked for consistency between textual descriptions of major area (divisions) and locations of geographic coordinates. in each case, where possible, we corrected the data record; where no clear correction was possible, we discarded data, recording losses at each step in the cleaning process. finally, (5) we discarded data records for which information on year, month, or day of collection was lacking and created a unique ‘stamp’ of time as year_month_day. we then aggregated point-based occurrence data to 0.5° spatial resolution across the country. this spatial resolution was the product of a detailed analysis of balancing benefits of aggregating data (i.e., larger sample sizes), versus the loss of spatial resolution that accompanies broader aggregation areas that can make important geographic features imperceptible (i.e., 1° resolution is a square ~110 km on a side). details of this procedure are provided by sousa-baena et al. (2014) and explored and analyzed in more detail for african examples in idohou et al. (2015) and koffi et al. (2015). we produced the aggregation grid shapefiles in the vector grid module of qgis, version 2.4; added the coarse-resolution grid identification townpeterson typewritten text townpeterson typewritten text biodiversity informatics, 11, 2016, pp. 12-22 15 figure 1. species diversity amongst succulent plant families in kenya, based on records at the east african herbarium and published sources. figure 2. collecting trends of succulent plant specimens between the years 1888 and 2012. ,-./0/1012% 3456789% :;0..-<0/1012% !!=% >0?@ab;;ab10/1 012%!'=% 5cb/d?0/1012% 3456789% 8-cab;ef0/1012% "(=% g@a1;.2%"+=% biodiversity informatics, 11, 2016, pp. 12-22 16 codes to each occurrence datum; and aggregated each datum into the coarse-resolution aggregation squares. in excel, we explored associations between data on species identity, time, and aggregation grid square. we calculated (1) the total number of records (n) available from each grid square; (2) the total number of species recorded from each grid square (sobs); and (3) the number of species detected on exactly one day (a), and (4) the number of species detected on exactly two days (b). using equations provided by chao (1987), we calculated the expected number of species (sexp), as 𝑆!"# = 𝑆!"# + 𝑎! 2𝑏, and inventory completeness (c) as c = sobs / sexp. we explored plots of c versus n to assess practical, appropriate, and adequate definitions of relatively completely versus incompletely inventoried grid squares. once we had established criteria for which grid squares could be considered as wellsampled, in qgis, we linked the table with the grid square statistics (i.e., n, sobs, sexp, c) to the aggregation grid, and saved this file as a shapefile. applying the criteria for ‘well-sampled’ (i.e., c > 0.5 and n > 10), we created a shapefile of wellsampled grid squares, which we in turn converted to raster (geotif) format using custom scripts in r (r core team, 2013). this raster coverage was the basis for our identification of gaps. we then used the proximity (raster distance) function in qgis to summarize geographic distance across the country to any well-sampled grid square. to create a parallel view of environmental difference from well-sampled areas (i.e., how different the climate is from that of the most similar well-surveyed grid square), we plotted 5000 random points across the country, and used the point sampling tool in qgis to link each point to the geographic distance raster, and to raster coverages (2.5’ spatial resolution) summarizing annual mean temperature and annual precipitation drawn from the worldclim climate data archive (hijmans et al., 2005). we exported the attributes table associated with the random points, and analysed further in excel. we first standardized the values of each environmental variable to the overall range of the variable as (xi – xmin) / (xmax xmin), where xi is the particular observed value in question. we then created a matrix of euclidean distances in the two-dimensional climate space, relating all of the points with a geographic distance >0 to all of the points with geographic distance of zero. the latter represent points falling in wellsampled grid squares, whereas the former are scattered across the entire country; points falling in well-sampled grid squares were assigned (by definition) environmental distances of zero. finally, the environmental distances were imported into qgis, and linked back to the random points shapefile to create a new shapefile with broad sampling across the country, with a z-value that is the environmental distance associated with that point. to convert this vector-format dataset into raster format, with values interpolated across the entire region, we used a second-degree inversedistance weighting approach, although many other interpolation approaches could be explored. results on the basis of available literature and records in the ea database, the succulent flora of kenya is dominated by species of five families: euphorbiaceae, apocynaceae (subfamily asclepiadoideae), xanthorrhoeaceae (genus aloe), crassulaceae and ruscaceae. these families together account for 80% of kenya’s succulent flora. several other families (malvaceae, cactaceae, pedaliacaeae and icacinaceae) are represented by single species each; the aizoaceae (mesembryanthemaceae), a predominantly south african family, is represented by two species of the genus delosperma. many species listed for kenya in the flora and other publications were not represented by any records in the dataset analyzed. this study indicated that collecting activity of succulent plants began in the latter part of the nineteenth century: the first record of a succulent species was made in 1888. this early period lasted until about 1910, within which time only 12 specimens of succulent plants were collected. subsequently, however, collecting activity accelerated, with peaks in the 1970s and 1990s. over these years, numbers of collections increased, with annual averages of 90–100 specimens. this activity, however, dwindled to a meagre average of about 25 specimens per year in more recent years. generally, the raw data showed close spatial correspondence to the existing road network and major settlements and urban areas. collection ‘hotspots’ of the country thus include areas neighbouring the capital city nairobi, the central, western and coastal regions, and mountain peaks in biodiversity informatics, 11, 2016, pp. 12-22 17 figure 3. spatial patterns of succulent plant collecting vis a vis existing protected areas (black outlines) in kenya. darker (brown) zones have highest record density; lighter (pink) zones have lower record density. white zones were not represented by collections in datasets available for this study. figure 4. well-known grid squares (purple colour outlines) in relation to collecting sites (black dots) revealing areas of relatively complete inventory of succulent plants in kenya. biodiversity informatics, 11, 2016, pp. 12-22 the southeastern part of the country. areas to the north, northeast and northwest are visibly undersampled except for scattered mountain peaks such as kulal, nyiru and the ndotos. of interest is the fact that most of the sampling ‘hotspots’ fall outside the country’s conservation areas, such that little current information is available for conservation areas. in the course of data cleaning, ~20% of the records were discarded in light of gaps in information content and inconsistencies; the final dataset contained 4304 records of succulent plant specimens collected from across the country. the total area of kenya was contained in 206 grid squares, 31.5% of which had no succulent plant records. eight grid squares fulfilled criteria for completeness of inventory: one centred on the coastal city of mombasa, one in the kulal–nyiru complex in the north, one at garissa in the northeast, one at baringo in the northwest and four grid squares in the nairobi–nakuru–laikipia complex. in between well-known sites, broad regions constituted gaps in kenya’s succulent plant inventory. measurement of geographic distance from well-known grid squares revealed the most farflung under-sampled regions; the turkana and the moyale regions bordering ethiopia were the most isolated. the next level of isolation included the northeastern zone in general: wajir, marsabit, mandera (moyale), and tana river counties bordering somalia in the east and the lake victoria basin in the west, the region between the nairobi area and other well-sampled grid squares showed some geographic isolation as well. roughly congruent patterns were observed in terms of environmental distance from wellsampled sites: maximum environmental difference was evident for the turkana region, wajir and mandera in the north to northeast, the tana river– malindi complex on the north coast and the lake victoria basin in the west. notable here is the fact that some geographically distant sites, such as those along the border with ethiopia (mandera), south sudan (illemi triangle), and tanzania (kilimanjaro and amboseli) were not markedly distinct in environmental terms. discussion while specimens are undoubtedly the most accurate primary research archives documenting biological diversity on earth, they are inevitably subject to spatial bias resulting from ad hoc spatial accumulation of samples, with most data coming from easy-access localities (ponder et al, 2001; reddy and davlos, 2003; grand et al., 2007). in addition, botanists tend to concentrate research efforts in botanically diverse (hence interesting) areas, generally avoiding species-poor areas (soria-auza and kessler, 2008). finally, and more specifically to this paper, succulent plants are notoriously difficult to prepare and preserve (bridson and forman, 1989; newton, 1995), such that numbers of specimens of these groups tend to be lower than in other taxa. in line with global trends of taxonomic diversity, the kenyan succulent flora is dominated by the large families euphorbiaceae, apocynaceae (subfamily asclepiadoideae), xanthorrhoeaceae, and crassulaceae (oldfield, 1997). similarly, aizoaceae and cactaceae, predominantly south african and south american in terms of biogeographic origins and centers of diversity, respectively, are poorly represented, with only two and one species, respectively. the country also holds rich diversity of the otherwise small family ruscaceae, represented by species of the genera sansevieria and dracaena. (figure 1). according to newton (2004a), the earliest specimens of succulent plants were collected by thomas wakefield, an english missionary who lived on the kenyan coast for over 20 years beginning around 1862. this period (figure 2) was one of little or no knowledge of the kenyan flora, and is part of what has been termed as the ‘heroic period,’ during which scientists had to brave dangerous terrain to access the largely unexplored african interior (gillett, 1962). in subsequent years, however, further opening up of the interior, and particularly the arrival of trains and motor vehicles facilitated plant survey and collecting expeditions. one notable development in the 1930s was establishment of the corydon museum herbarium in nairobi, and hiring of its first keeper, who actively carried out collecting missions that added at least 4000 specimens to the collection (newton, 2004b). in addition, interest in the regional flora accelerated, leading to commencement of preparation of the flora of tropical east africa in the early 1950s. activities of the flora project peaked in the 1990s, by which time almost 70% of the family accounts had been published townpeterson typewritten text townpeterson typewritten text townpeterson typewritten text 18 townpeterson typewritten text townpeterson typewritten text biodiversity informatics, 11, 2016, pp. 12-22 19 figure 5. geographic distances across kenya to well-inventoried grid squares. dark blue zones are within well-known areas, light blue to yellow shows middle to high distances; red zones are extremely isolated from well-inventoried grid squares. figure 6. environmental distance showing ecological isolation of little known, poorly inventoried sites. dark blue zones are within well-known areas, light blue to yellow shows middle to high distances; red zones are extremely distinct from well-inventoried grid squares. biodiversity informatics, 11, 2016, pp. 12-22 20 (beentje and smith, 2001; newton, 2004a). however, field collection decelerated in succeeding years, probably owing to difficulties in the research permitting process, which remain almost prohibitive to date, among other impediments. based on our analyses, three well-known sites—the kulal–nyiro complex in the north, garissa in the northeast, and baringo in the northwest—are largely semi-arid to arid in nature, and therefore are good confirmation of universal patterns of succulent plant habitat preferences. oldfield (1997) noted that the somalia-maasai rce, to which most of this region belongs, has been documented as being especially diverse for succulent taxa. however, expansive geographic and environmental gaps exist among these few well-known sites, raising the possibility that concentrated survey efforts sited strategically in these gaps would result in more species new to science, or new to the flora, or contribute hitherto undocumented information on succulent plant diversity patterns. the rest of the well-known sites, the nairobi– nakuru–laikipia complex and that centred on mombasa at the coast, certainly reflect both ease of access and the concentration of infrastructure, including research institutions, personnel, and botanical gardens and herbaria. two of the largest herbaria, ea and the herbarium of the university of nairobi, are located in nairobi; the nairobi botanic garden, among the oldest such facilities in the country, is also located in nairobi. the coastal region has benefited from research capacity generated through the coast forest survey programme of the national museums of kenya for over two decades. as such, these regions are wellknown owing simply to concentrated sampling and documentation of the flora. while the plant inventory completeness patterns documented in this study are obviously determined by a combination of the natural richness of sites, as well as ease of access and proximity to research infrastructure, it is important to recall that only a fraction of the existing succulent plant data was available for the present analyses. some of the existing data was either not digital, had no geographic coordinates, or had no dates of collection, causing extensive ‘leakage’ from an otherwise more sizeable dataset. this reduction of digitally accessible knowledge therefore calls for need to improve documentation protocols and standards, including the processes of specimen collection, preparation, and storage, and subsequent management, digitisation, improvement, and publication of associated data. furthermore, the data analysed here correspond to collections held in only two herbaria, neither of which has been digitised completely. as observed by morat and lowry (1997), gaps such as those exhibited in the succulent flora of kenya dataset emphasize the need to take advantage of available expertise to verify, compare, and standardise information to generate reliable accounts of floristic elements of various regions. conclusions and recommendations the history of succulent plant collecting is apparently closely associated with patterns of general floristic exploration across east africa, and in kenya in particular, which undoubtedly is influenced by ease of access to localities. as the debate on climate change continues, it will be critical to reflect on the wealth of biodiversity of the region, and identify strategies for mitigation that will enhance resilience and survival of ecosystems in the face of the anticipated increased temperature and reduced rainfall. importantly, it will be critical to re-evaluate the extent to which unique elements of biodiversity, such as its succulent flora, are protected in the present conservation area network. as demonstrated in this study, knowledge of succulent plants in kenya is far from complete, hence the need for focused survey, inventory and documentation especially in sites that have been shown to be distant geographically and ecologically. with the writing of the flora out of the way, the time is ripe for development of research programmes that will translate existing information into conservation policy and action, and mobilise resources to enable survey and inventory teams to ramp up collection activity in isolated and little-known sites across the country. increased digitisation and publication (i.e., data sharing via data portals such as the global biodiversity information facility2 of information should be priority for ea and herbaria in general to enhance utility of research material. digitisation and open sharing of data–in effect data ‘repatriation’–from more herbaria with significant 2 http://www.gbif.org. biodiversity informatics, 11, 2016, pp. 12-22 21 kenyan holdings, including missouri botanical garden, botanic garden and botanical museum of berlin-dahlem, the smithsonian institution, and others, is also crucial. acknowledgments we thank the biodiversity informatics team, national museums of kenya, led by dr. geoffrey mwachala, director of research and collection, for sustained efforts in digitising collections of the museums. we thank the east african herbarium for permission to use the plants database, and kew gardens for their efforts in digitising and sharing their data. capacity building for this assessment was provided by the university of kansas biodiversity informatics training curriculum, funded by the jrs biodiversity foundation. references el-ghani, a. m., a. soliman and r. a. el-fatta. 2014. spatial distribution and soil characteristics of the vegetation associated with common succulent plants in egypt. turkish journal of botany 38: 550– 565. ariño, a. h. 2010. approaches to estimating the universe of natural history collections data. biodiversity informatics 7: 81–92. beentje, h. and s. smith. 2001. ftea and after. in (eds.). e. robbretht, j. degreef, and i. friis. plant systematics and phytogeography for the understanding of african biodiversity. proceedings of the xvith aetfat congress, held at the national botanical garden, belgium, 2000. bridson d. and l. forman.1989. the herbarium handbook. revised edition, 303 pp. royal botanic gardens, kew, uk. chao, a. 1987. estimating the population size for capture-recapture data with unequal catchability. statistica sinica 10: 227–246. colwell, r. k. and j. a. coddington. 1994. estimating terrestrial biodiversity through extrapolation. philosophical transactions of the royal society of london 345:101–118. gillett, j. b. 1962. the history of the botanical exploration of the area of “the flora of tropical east africa.” comptes rendus iv réunion aetfat. pp 205–229. hepper, f. n. 1979. second edition of the map showing extent of floristic exploration in africa south of the sahara, published by aetfat. in g. kunkel (ed.), taxonomic aspects of african economic botany. proceedings of the ix plenary meeting of aetfat. las palmas de gran canaria. pp 157– 162. hijmans, r. j., s. e cameron, j. l. parr, p .g. jones and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25: 1965–1978. idohou, r., a. h. ariño, a. e. assogbadjo, r. g. kakai, g. k. romain, b. sinsin. 2015. knowledge of diversity of wild palms (arecaceae) in the republic of benin: finding gaps in the national inventory by combining field and digital accessible knowledge. biodiversity informatics 10: 45–55. koffi, k. j., a. f kouassi, c. y. a. yao, a. bakayoko, i. j. ipou and j. bogaert. 2015. the present state of botanical knowledge in côte d’ivoire. biodiversity informatics 10: 56-64. küper, w., w. sommer, j. h., lovett, j. c. and w. barthlott. 2006. deficiency in african distribution data—missing pieces of the puzzle. botanical journal of the linnaean society 150: 355–368. lucas, g. l. 1968. kenya. acta phytogeographica suecica 54: 152–163. maundu, p. m., g. w. ngugi and h. s. kabuye. 1999. traditional food plants of kenya. national museums of kenya, nairobi. kenya. mewnr and rda. 2015. kenya biodiversity atlas ministry of environment, natural resources and regional development authorities. nairobi, kenya. mittermeier, r. a., p. r. gil, m. hofman, j. pilgrim, t. brooks, c. g. mittermeier, j. lamoreux and g. a. b. da fonseca. 2005. hotspots revisited: earth’s biologically richest and threatened terrestrial ecoregions. conservation international, washington, d.c. morat, p. and p. lowry. 1997. floristic richness in the africa madagascar region: a brief history and perspective. adansonia 19: 101–115. myers, n., r. a mittermeier, c. g mittermeier, g. a. b. da fonseca and j. kent. 2000. biodiversity hotspots for conservation priorities. nature 403: 853–858. newton, l. e. 2003. a check-list of kenyan succulent plants. succulenta east africa, nairobi, kenya. newton, l. e. 2004a. the history of succulent plants in kenya. succulenta east africa, nairobi, kenya. newton, l. e. 2004b. the first herbarium botanist in nairobi. journal of east african natural history 93: 49–55. newton, l.e. 1995. making herbarium specimens. ballya 2: 1-3. r development core team 2008. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. isbn 3-900051-07-0.3 3 http://www.r-project.org. biodiversity informatics, 11, 2016, pp. 12-22 22 nordal, i., sebsebe demissew and e. s. odd. 2001. endemism in groups of geophytes (liliiflorae). biologiske skrifter 54: 247–258. oldfield, s. (ed.). 1997. cactus and succulent plants– status survey and conservation action plan. iucn/ssc cactus and succulent specialist group. iucn. switzerland and cambridge, uk. ponder, w. f., g. a. carter, p. flemons, and r. r. chapman. 2001. evaluation of museum collection data for use in biodiversity assessment. conservation biology 15: 648–658. reddy, s., and l. m. davalos. 2003. geographical sampling bias and its implications for conservation priorities in africa. journal of biogeography 30: 1719–1727. sajeva, m., and m. costanzo. 1994. succulents: the illustrated dictionary. cassell, london uk. soria-auza, r. w., and m. kessler. 2008. the influence of sampling intensity on the perception of the spatial distribution of the spatial distribution of tropical diversity and endemism: a case study of ferns from bolivia. diversity and distributions 14: 123–130. sousa-baena, m. s., l. c. garcia and a. t. peterson. 2014. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. diversity and distributions 20: 369-381. white, f. 1993. the vegetation of africa. a descriptive memoir to accompany the unesco/aetfat/ unso vegetation map of africa/ unesco, paris. 8 biodiversity informatics, 14, 2019, pp. 8-13 climatestability: an r package to estimate climate stability from time-slice climatologies hannah l. owens1,2,* and robert p. guralnick1 1florida museum of natural history; university of florida; dickinson hall, gainesville, fl, 32611, usa; 2center for macroecology, evolution, and climate; copenhagen university; universitetsparken 15, bld. 3, 2nd floor dk-2100 copenhagen, denmark; *corresponding author: hannah.owens@gmail.com abstract. as continental and global-scale paleoclimate model data become more readily available, biologists can now ask spatially explicit questions about the tempo and mode of past climate change and the impact of those changes on biodiversity patterns. in particular, researchers have focused on climate stability as a key variable that can drive expected patterns of richness, phylogenetic diversity and functional diversity. yet, climate stability measures are not formalized in the literature and tools for generating stability metrics from existing data are nascent. here we define “deviation” of a climate variable as the mean standard deviation between time slices over time elapsed; “stability” is defined as the inverse of this deviation. finally, climate stability is the product of individual climate variable stability estimates. we also present an r package, climatestability, which contains tools for researchers to generate climate stability estimates from their own data. introduction many hypotheses have been offered to explain geographic patterns of biodiversity from evolutionary and ecological perspectives (macarthur, 1984). it has often been suggested that decreased extinction and increased speciation rates have led to an accumulation of species in the tropics (rohde, 1992; rolland, et al. 2014). proposed drivers of this diversification rate pattern include reduced seasonality and long-term climate stability in the tropics, which, in addition to reducing extinction due to stochastic climate events, may have allowed species to specialize in narrow abiotic ecological niches and allowed a greater diversity of species to accumulate. outside of the tropics, pockets of relatively stable conditions, introduced by heusser in 1955 as refugia, may have provided species with the means to escape climate-driven extinctions in the past. conversely, in temperate climates, seasonality and cyclical warming and cooling may have both driven species to remain generalists to accommodate varying conditions (gouveia, et al. 2013; sandel, et al. 2011) and extirpated those species that could not accommodate such climatic stochasticity. yet, it has also been suggested that climate instability has driven the formation of biodiversity hotspots in some systems, particularly as it pertains to allopatric speciation during glacial cycles (mayr and o’hara, 1986; nakazawa and peterson, 2015). with advances in climate modeling, more and higher-resolution climate data are available for biologists to infer how past climate change may have influenced modern biodiversity patterns. recently, loarie and colleagues (2009) proposed “climate velocity” (the speed an individual would have to travel to track analogous climatic conditions) as a useful way to estimate potential climate change impacts on the position of organisms’ ranges. a simplified version of this metric has also become common: dividing the difference between past and modern climate variable values (e.g. average temperature, average yearly rainfall) by the time elapsed between the two values (e.g. ornelas, et al. 2015). an even simpler metric of climate change is “climate anomaly”, the difference between climate variables between time slices of interest (e.g. feng, et al. 2017). however, while these metrics are useful for estimating the degree of climate change between two time points, none adequately characterizes the stability of climates over time. with the advent of paleoview (fordham, et al. 2017) and other paleoclimate assembly tools and datasets (e.g. brown, et al. 2018; fick and hijmans, 2017), as well as expanded computing capacity to process large climatological datasets, time slice climatologies from climate model data are now more accessible than ever. there is also an increasing wealth of uneven time slice data from empirical inferences of climate based on geologic or archeological proxies.1 while the time-frame and exact measure of climate 1https://www.ncdc.noaa.gov/paleo-search/ 9 stability may vary according to the research question, we propose here a clear definition of climate stability, as well as provide an implementation of our metric to facilitate macroecological analyses predicated on climate stability hypotheses. for each climate variable of interest, deviation through time is calculated as the mean of standard deviations between time slices divided by the time elapsed between time slices; “stability” is defined as the inverse of deviation through time. finally, climate stability is the product of individual variable stability estimates that have been re-scaled between 0 and 1. herein, we present an r package, climatestability, with functions to perform these calculations and visualize the results, as well as providing a set of world-wide climate temperature and precipitation deviation-through-time layers over the last 21kybp, from which climate stability can be quickly estimated. these data may then facilitate empirical tests of two key hypotheses in bioeographic theory: 1) diversity results from a cradle-like environment where extinction resulting from climatic stocasticity is infrequent, and 2) diversity results from repeated isolation events, often driven by climate change. we also discuss the challenges with temporal grain and interpretation when developing climate stability datasets. general package structure and functionality. the climatestability r package allows users to: 1. calculate stability for individual climate variables over even or uneven time slices. 2. calculate overall climate stability from individual climate variable stability estimates. 3. visualize resulting estimates either in geographic space, or as a graph showing the relationship between latitude and stability. for the climatestability workflow, the user first produces rasters depicting estimated deviation through time for a single climate variable of interest by loading a series of climatology estimates (as raster files) and stipulating the time elapsed between each time slice (either as a single value if the time elapsed between climatologies is the same for all, or as a vector of values specifying the time elapsed between climatologies in order). this is repeated for each climate variable of interest. next, for each variable, the user then calculates the inverse of the deviation and rescales the result to a scale between 0 and 1, in order to render comparable the stability estimates for each variable. overall climate stability can then be estimated by taking the product of stability estimates for each individual variable and rescaling the result to a scale between 0 and 1. finally, climatestability provides two functions to graph the relationship between a given latitudinal band and stability (for single variables or overall climate)—one for raw latitude (-90° to 90°), and one for the absolute value of latitude (0° to 90°). we have also provided two raster layers of deviation estimates for global precipitation and temperature from 21 to 1 kypb, as described below, and a vignette describing the basic climatestability workflow (appendix 2). methods climate data. paleoclimate estimates were derived from the trace21ka experiment (liu, et al. 2009; liu, et al. 2014; otto-bliesner, et al. 2014) implemented using the community climate system model, version 3 (ccsm3; collins, et al. 2006; otto-bliesner, et al. 2006; yeager, et al. 2006). more details on these data are described by fordham and colleagues (2017). briefly, the ccsm 3 is a mathematical model that simulates coupled atmosphere and ocean climatic conditions given a particular set of parameters, such as those from the trace21ka experiment, which sought to simulate the rapid global climate changes of the last 21,000 years. the model produces monthly snapshot estimates of the simulated climate over the duration of a particular run (20,000 years, in the case of trace21ka), which can then be processed into climatic means at a given time point. using paleoview, version 1.1 (fordham, et al. 2017) we generated 20 100-year climatological means for 1,000-year time slices between 1 and 21 kypb for annual precipitation and mean temperature at a resolution of 2.5 degrees. we also selected eight climatologies (from 1, 2, 3, 4, 5, 10, 15, and 21 kybp) from the original set of 20 to simulate an unevenly-sampled time slice series. this was done to demonstrate the properties of our climate stability measurement under such conditions as might exist when a researcher wishes to incorporate deeper-time slices or time slices based on empirical climate inferences. we chose these specific time slices to simulate a “pull of the present” scenario where data becomes scarcer deeper in time. calculation. for each time interval, we calculated the standard deviation of a given variable at the beginning and end of the interval and divided the result by the length of the interval to quantify deviation 10 over time for each time slice, then averaged the result across all time slices. we then took the inverse of this metric and scaled it to between 0 and 1 to estimate the relative stability for each variable. finally we multiplied the stability estimate for mean annual precipitation by the stability estimate for mean annual temperature to estimate overall climate stability over the last 21,000 years. we performed these calculations both for our original climate dataset and the thinned dataset (hereafter referred to as the “even dataset” and “uneven dataset”, respectively), and plotted the results as a both as raster maps and linear graphs of the relationship between stability versus latitude, all in the r programming platform (r core team, 2013). code for these functions, as well as the temperature and precipitation deviation-through-time layers we produced can be downloaded as the r package climatestability2. code for the analyses as described herein is available as appendix 3. results over the last 21,000 years, temperature has been most stable in the tropics and decays as latitude increases, but the patterns can be complex (e.g. low temperature stability in the amazon basin over the time period examined; fig. 1). other notable areas of low temperature stability are found in northern north america, europe, and asia as well as central australia; high temperature stability areas include sub-saharan africa and the indian subcontinent. there is no consistent pattern between precipitation stability and latitude in this dataset (fig. 2); instead, terrestrial precipitation stability appears highest in central asia and antarctica, whereas marine precipitation stability appears highest in areas with cold ocean currents in the temperate southern pacific and atlantic (fig. 1). overall climate stability appears more strongly dictated by precipitation stability than by temperature (figs. 1 and 2). generally, the uneven time slice dataset estimated congruent patterns of temperature, precipitation, and overall stability compared to the even dataset (supplementary figure 1). the most dramatic disagreements between the even and uneven datasets are found in the temperature stability estimates (supplementary figure 2)—the even dataset estimates higher stability in the tropics, whereas the uneven dataset estimates higher stability in the temperate zones, especially in the southern hemisphere. these 2https://github.com/hannahlowens/climatestability trends are also reflected in the mean stability by latitude plots (fig. 2). discussion our estimated geographic patterns of temperature stability largely reflect glaciation patterns over the past 21,000 years, conforming to a latitudinal gradient predicted in macroecological and biogeographic theory (slobodkin and sanders, 1969). conversely, patterns of precipitation stability appear to follow atmospheric circulation patterns, with areas of lowest stability occurring in areas of atmospheric upwelling and highest stability occurring with downwelling areas. temperature and precipitation do not play the same role in defining species’ suitable abiotic niches across the tree of life (araujo and guisan, 2006); therefore, it may be informative to consider these variables separately when assessing the degree to which climate stability influences species diversity. still, our metric of climate stability provides an figure 1. relative temperature, precipitation, and climate stability based on even dataset. darker colors indicate higher stability. 11 intuitive means of quantifying a key macroecological driver of distributions of biodiversity. our larger aim with this contribution has been to provide a more formal basis for generating climate stability measures. stability has been addressed more and more as a key hypothesis in the biodiversity literature, but there have been multiple types of measurements and underlying data used (garcia et al., 2014). while these uses are often reasonable in the context of research questions, it can be hard to evaluate across studies, and different metrics can support very different conclusions regarding how climate change affects assemblages, species, and populations at local and regional scales (garcia et al., 2014). we limit our definition here to climate stability. climate stability or lack thereof can be a driver of habitat (e.g. vegetation) stability or instability, which is more likely to be the ultimate driver of species’ extinction or persistence in the face of climate change (ashcroft, 2010). clear labeling of the stability metric employed will help tremendously not only to understand broader past trends, but also to make predictions about the effects of climate change in the future. as well, we urge reporting of temporal grain when considering climate stability metrics – stability or lack thereof is determined in part by temporal smoothing imposed by sampling grain. our even dataset used 1,000 year time slices, but over longer timescales than the late pleistocene and holocene, the lack of fine-grain and even climatological time-slices makes the inference of stability more challenging. as our simulated uneven dataset illustrates, estimates of climate stability may be greatly influenced by the rate of variation through time. generally, we would expect a dataset with fewer time slices to estimate lower climate stability than a dataset with more time slices due to temporal smoothing. however, the uneven dataset generally estimated higher temperature stability than the even dataset at all but the lowest latitudes (fig. 2); when these differences mapped, it is clear that the uneven dataset estimated higher temperature stability particularly at middle latitudes (supplemental fig. 2). given that the uneven dataset was largely biased toward modern time slices, this suggests that the rate of change in temperature decelerates at middle and high latitudes as we approach the present. finally, while the example provided herein used data from a single climate model under a single set of conditions, we recommend estimating climate stability with data from multiple climate models, when available. different climate models have different sensitivities to climatic conditions, and may differ in their estimates of local and/or extreme phenomena—this is true of both past and future climate simulations. to generate particularly robust estimates of climate stability, researchers may want to either perform additional analyses to test how sensitive their results are to the climate model used, or use data figure 2. plots of temperature, precipitation, and climate stability versus latitude. solid line: even dataset; dashed line: uneven dataset. 12 from ensemble model outputs such as those from the paleoclimate modelling intercomparison project (pmip)3 and coupled model intercomparison project (cmip)4. while our climate stability metric is not completely robust to the effects of uneven time slices, it consistently identifies areas of high and low stability regardless of temporal resolution and distribution of climate data. we hope the contribution presented here precipitates a step forward from the qualitative assessment that the tropics are generally climatically stable to a more quantitative assessment of where the tropics are stable, over what time periods, and at what temporal and spatial resolution. this will facilitate more thorough examination of the role of climate stability in driving diversity patterns, which may, in turn, provide critical calibration needed for understanding how continuing and accelerating climate instability may ultimately structure future biodiversity. software availability software available on cran5. install the latest release in r as install.packages(“climatestability”), and latest development version as devtools:: install_github(“hannahlowens/climatestability”). supplementary materials supplementary materials cited in the text are available at https://doi.org/10.17161/1808.28080. license: gnu general public license v. 3.0 acknowledgments we thank b. f. oliviera, collaboration with whom was the inspiration for this project. we were supported by the u.s. national science foundation (nsf) division of environmental biology (deb) grant # 1541500. competing interests the authors have declared that no competing interests exist. references araujo, m.b., and a. guisan. 2006. five (or so) challenges for species distribution modelling. j. biogeogr. 33:1677-88. 3https://pmip.lsce.ipsl.fr/ 4https://esgf-node.llnl.gov/projects/esgf-llnl/ 5https://cran.r-project.org/web/packages/climatestability/index.html ashcroft, m.b. 2010. identifying refugia from climate change. j. biogeogr. 37: 1407–1413. brown, j., hill, d.j., dolan, a.m., carnaval, a.c., and a.m. haywood. 2018. paleoclim, high spatial resolution paleoclimate surfaces for global land areas. sci. data 5: 180254. collins, w.d., bitz, c.m., blackmon, m.l., bonan, g.b., bretherton, c.s., carton, j.a., chang, p., doney, s.c., hack, j.j., henderson, t.b. and j.t. kiehl. 2006. the community climate system model version 3 (ccsm3). j. clim. 19: 2122–2143. feng, g., ma, z., benito, b.m., normand, s., ordonez, a., jin, y., mao, l., and j.-c. svenning. 2017. phylogenetic age differences in tree assemblages across the northern hemisphere increase with long-term climate stability in unstable regions. glob. ecol. biogeogr. 26: 1035–1042. fick, s.e. and r.j. hijmans, 2017. worldclim 2: new 1-km spatial resolution climate surfaces for global land areas. int. j. climatol. 37: 4302–4315. fordham, d.a., saltré, f., haythorne, s. , wigley, t.m., otto‐bliesner, b.l., chan, k.c. and b.w. brook. 2017. paleoview: a tool for generating continuous climate projections spanning the last 21,000 years at regional and global scales. ecography 40: 1348–1358. garcia, r.a., cabeza, m., rahbek, c., and m.b. araújo. 2014. multiple dimensions of climate change and their implications for biodiversity. science 344: 1247579. gouveia, s.f., hortal, j., cassemiro, f.a.s., rangel, t. f. and j.a.f. diniz-filho. 2013. nonstationary effects of productivity, seasonality, and historical climate changes on global amphibian diversity. ecography 36: 104–113. heusser, c.j. 1955. pollen profiles from the queen charlotte islands, british columbia. can. j. bot. 33: 429– 449. liu, z., otto-bliesner, b.l., he, f., brady, e.c., tomas, r., clark, p.u., carlson, a.e., lynch-stieglitz, j., curry, w., brook, e. and d. erickson. 2009. transient simulation of last deglaciation with a new mechanism for bølling-allerød warming. science 325: 310–314. liu, z., lu, z., wen, x., otto-bliesner, b.l., timmermann, a. and k.m. cobb. 2014. evolution and forcing mechanisms of el niño over the past 21,000 years. nature 515: 550. loarie, s.r., duffy, p.b., hamilton, h., asner, g.p., field, c.b., and d.d. ackerly. 2009. the velocity of climate change. nature 462: 1052–1055. macarthur, r.h. 1984. geographical ecology: patterns in the distribution of species. princeton university press, princeton, new jersey. 269 pp. 13 otto-bliesner, b.l., tomas, r., brady, e.c., ammann, c., kothavala, z. and g. clauzet. 2006. climate sensitivity of moderate-and low-resolution versions of ccsm3 to preindustrial forcings. j. clim. 19: 2567– 2583. mayr, e. and r.j. o’hara. 1986. the biogeographic evidence supporting the pleistocene forest refuge hypothesis. evolution 40: 55–67. nakazawa, y. and a.t. peterson. 2015. effects of climate history and environmental grain on species’ distributions in africa and south america. biotropica 47: 292–299. neiva, j., paulino, c., nielsen, m.m., krause-jensen, d., saunders, g.w., assis, j., bárbara, i., tamigneaux, é., gouveia, l., aires, t., marbà, n., bruhn, a., pearson, g.a. and e.a. serrão. 2018. glacial vicariance drives phylogeographic diversification in the amphi-boreal kelp saccharina latissima. sci. rep. 8: 1112. nevado, b., contreras-ortiz, n., hughes, c. & filatov, d.a. (2018) pleistocene glacial cycles drive isolation, gene flow and speciation in the high-elevation andes. new phytol. 219: 779–793. otto-bliesner, b.l., russell, j.m., clark, p.u., liu, z., overpeck, j.t., konecky, b., nicholson, s.e., he, f. and z. lu. 2014. coherent changes of southeastern equatorial and northern african rainfall during the last deglaciation. science 346: 1223–1227. ornelas, j.f., s.g.d. león, c. gonzález, y. licona-vera, a.e. ortiz-rodriguez and f. rodríguez-gómez. 2015. comparative palaeodistribution of eight hummingbird species reveal a link between genetic diversity and quaternary habitat and climate stability in mexico. folia zool. 64: 245–258. owens, h.l., lewis, d.s., dupuis, j.r., clamens, a.l., sperling, f.a., kawahara, a.y., guralnick, r.p. and f.l. condamine. 2017. the latitudinal diversity gradient in new world swallowtail butterflies is caused by contrasting patterns of out‐of‐and into‐the‐tropics dispersal. glob. ecol. biogeogr. 26: 1447–1458. r core team. 2013. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria.6 rangel, t.f., edwards, n.r., holden, p.b., diniz-filho, j.a.f., gosling, w.d., coelho, m.t.p., cassemiro, f.a., rahbek, c. and r.k. colwell. 2018. modeling the ecology and evolution of biodiversity: biogeographical cradles, museums, and graves. science 361: eaar5452. rohde, k. 1992. latitudinal gradients in species diversity: the search for the primary cause. oikos 65: 514-527. rolland, j., condamine, f.l., jiguet, f. and h. morlon. 2014. faster speciation and reduced extinction in the tropics contribute to the mammalian latitudinal diversity gradient. plos biol. 12: e1001775. sandel, b., arge, l., dalsgaard, b., davies, r.g., gaston, k.j., sutherland, w.j. and j.-c.svenning. 2011. the influence of late quaternary climate-change velocity on species endemism. science 334: 660-664. slobodkin, l.b., and h.l. sanders. 1969. on the contribution of environmental predictability to species diversity. brookhaven symp. biol. 22:82-95. yeager, s.g., shields, c.a., large, w.g. and j.j. hack. 2006. the low-resolution ccsm3. j. clim. 19: 25452566. 6http://www.r-project.org/ assessment of user needs of primary biodiversity data: analysis, concerns, and challenges biodiversity informatics, 8, 2013, pp 59 93 assessment of user needs of primary biodiversity data: analysis, concerns, and challenges arturo h. ariño (1)*, vishwas chavan (2), daniel p. faith (3) (1) department of zoology and ecology, university of navarra, pamplona, spain. (2) global biodiversity information facility secretariat, universitetsparken 15, dk 2100, copenhagen, denmark (3) australian museum, 6 college street, sydney, nsw, australia. *corresponding author abstract — a content needs assessment (cna) survey has been conducted in order to determine what gbif-mediated data users may be using, what they would be using if available, and what they need in terms of primary biodiversity data records. the survey was launched in 2009 in six languages, and collected more than 700 individual responses. analysis of the responses showed some lack of awareness about the availability of accessible primary data, and pointed out some types of data in high demand for linking to distribution and taxonomic data now derived from the gbif cache. a notable example was linkages to molecular data. also, the cna survey uncovered some biases in the design of user needs surveys, by showing demographic and linguistic effects that may have influenced the distribution of responses received in analogous surveys conducted at the global scale. introduction biodiversity research is becoming a dataintensive science (kelling et al., 2009). the more than 267 million primary biodiversity records that hundreds of data publishers were making openly available through gbif by the end of 2010 international biodiversity year (gbif 2010) is a significant asset, and a valuable form of scientific capital (borgman, 2003, 2007). scientific data are expensive to produce but can be of tremendous future value (borgman, 2007), although quantifying such future value is difficult without some indication about future data uses. nevertheless, the value for natural history collections (nhcs) has been demonstrated, in accord with long-standing predictions (grinnell, 1910). such values are seen as extending to various kinds of data derived from biodiversity research which need to be held in perpetuity as well. of course, not all available data are fit for all uses (hill et al., 2010). the expense of producing data, and maintaining the cyberinfrastructure needed for their open access, delivery, and dataintensive collaborative research (borgman et al., 2006), justifies increased efforts to assess what types of biodiversity data are most needed by researchers. this may help optimize resource allocation and research output. in 2009, gbif set up a content needs assessment task group (cna tg) to address this assessment (gbif 2009a). the objective of cna is to get a first-hand idea about the user needs of biodiversity data (chavan et al., 2010). two main tools are available for cna: information mining (including literature review), and surveys. while the former may be thought of as retrospective research, collating documented uses in response to specific needs, the latter can proceed both ways: describing researchers’ past, present and possible future requirements. in 2009, cna tg conducted a survey with the purpose of collecting information on the demography of use of biodiversity data, and understanding the myriad of broad ‘primary biodiversity data’ needs across user communities (gbif, 2009). the survey also sought input to determine the unique scientific and policy contributions of uses made of data mobilized and accessible through the gbif community. assessment of user needs – ariño et al. 60 design the survey contained 21 questions spread over 6 sections (table 1): (a) respondent profile, (b) uses of primary biodiversity data, (c) access to primary biodiversity data, (d) data quality and quantity requirements, (e) species level data requirements, and (f) usefulness of gbif mobilised data. the survey included an introduction succinctly describing gbif and the objective of cna (table 1). most questions were multiplechoice, although estimates were required for some quantity data. also, most questions included an option for a free-text answer not covered by available choices. table 1: list of questions and options overhead: gbif content needs assessment (cna) survey: introduction [objective of cna survey, description of gbif, estimated time to completion (21 questions, 18 minutes), anonymity assurance]. section (a): gbif content needs assessment (cna) survey: user profile user profile question q1. details of the person undertaking this survey. options/suboptions: name; organisation/institution affiliated with; street/po box; city; state; country; zip code; phone/mobile; email; web/url. (free-text answers) user profile question q2. describe your organization (please tick one or several options) options/suboptions: academic / educational institution; research institution; national agency; non governmental organisation (ngo); intergovernmental organisation (igo) or multilateral convention; private company; individual researcher or naturalists (e.g. citizen scientists); others (please specify). (exclusive multiple choice) user profile question q3. main interest/business of your organization (please tick one or several options) options/suboptions: conservation science (including taxonomic research); bioproductivity / bioprospecting (agriculture; fisheries; forestry; etc.); biodiversity; biomedical and/or public health; biotechnology; biosecurity; natural resources management; industrial / commercial use of natural resources; exhibition / educational / academic; others (please specify). (nonexclusive multiple choice) section (b): gbif cna survey: uses of primary biodiversity data. this section of the survey is designed to understand the purpose for which ‘primary biodiversity data’ is used by various stakeholders. definition: primary biodiversity data is defined as the digital text or multimedia data record detailing the instance of an organism – or the what, where, when, how and by whom of the organisms occurrence and recording. the uses of primary biodiversity data are wide and varied, and encompass virtually every aspect of human endeavor – food, shelter, health, recreation, art and history, society, science & politics, etc. furthermore, such data is essential for predicting the sustainable future of our planet, and therefore of all living beings. question (q) 01. list the ways in which you use primary biodiversity data (please choose one or several options) options/suboptions: taxonomy; biogeographic studies; species diversity & populations; life histories & phonologies; endangered, migratory and invasive species; impact of climate change; ecology, evolution & genetics; environmental regionalization; conservation planning; sustainable use; natural resources management; agriculture, fisheries, forestry and mining; nursery & pet industry; health & public safety; bioprospecting; forensics; border control and wildlife trade; education & public outreach; ecotourism; art & history; society and politics; recreation; human infrastructure planning; industrial use; environmental impact management; others (please specify). (non-exclusive multiple choice) q02. provide example documentation (reports/papers/presentations) where primary biodiversity data has been used by you/your group? note: please provide literature references, urls of web sites, news items. please also email us a copy of the report/paper etc. at contentneeds@gbif.org. options/suboptions: separate fields for multiple examples. (free-text answers) section (c): gbif cna survey: access to primary biodiversity data. in this section, gbif seeks to learn how users access primary biodiversity data (please choose one or several options). the objective is to understand the mechanisms employed and the frequency for accessing primary biodiversity data. q03. how do you access primary biodiversity data? options/suboptions: through your own field works/surveys; through hardcopy, literature survey (nondigital form); through primary publications (e.g. taxonomic monographs, maps of species observations); through access to offline digital data sets (cdrom/dvd/tapes etc.); through the gbif data portal (http://data.gbif.org); through other web based data mailto:contentneeds@gbif.org http://data.gbif.org/ assessment of user needs – ariño et al. 61 portals (please specify); through ftp sites (please specify); through institutional agreements; through payment basis; through free and open datasests within and outside of your institution; through reciprocal agreements with other groups/individuals; through others (please specify). (non-exclusive multiple choice) q04. frequency of access options/suboptions: daily basis; once a month; once a quarter; bi-annual; can not determine (on need basis); others (please specify). (exclusive multiple choice) q05. some of the datasets mobilised through gbif have multiple access points (e.g. obis mobilised data set can typically have three access points – gbif data portal, obis portal, and data sets own portal). how do you access such data sets? options/suboptions: only through gbif portal; only through thematic/regional aggregator portal(s); directly through datasets own portal(s); all of the above. (exclusive multiple choice) q06. if you are accessing datasets through access points other than the gbif data portal, why? options/suboptions: lack of awareness about accessibility through the gbif portal; have been using these access points for a long time; ease of use; more specific search features; workflow integration; others (please specify). (non-exclusive multiple choice) q07. select the data formats which you often choose to access the primary biodiversity data. options/suboptions: mysql (dump); excel; tab delimited; comma separated values; xml; maps as images; kml; others (please specify). (non-exclusive multiple choice) q08. gbif serve data in all the formats listed in the previous question. if gbif were to serve data in other formats, which would be your preference(s)? (freetext answers) q09. list the other types of data you use together with primary biodiversity data? (e.g. satellite imagery, environmental data layers such as salinity, temperature etc., land use data, infrastructure development such as housing, roads, dams, etc.) (free-text answers) section (d): gbif cna survey: quality and quantity requirements q10. types or nature of primary biodiversity data required? options/suboptions: taxonomic names/checklists; occurrence records (presence only); occurrence records (including absence records); population density/dynamics; species interaction data; species information (descriptive data); others (please specify). (non-exclusive multiple choice) q11. quantity of data required for each data type? options/suboptions: taxonomic names/checklists; occurrence records; population density/dynamics; multimedia resources; others (please specify). (choice matrix. exclusive column options for each option row: 1100 records; 101-1000 records; 1001-10000 records; 10000+ records) q12. for which type of environments do you use/need more primary biodiversity data? options/suboptions: marine: coasts; marine: oceans; marine: deep seas; marine: islands; marine: estuarine; inland: wetlands; inland: river basin; terrestrial: tropical forests; terrestrial: temperate forests; terrestrial: deserts; terrestrial: grasslands; terrestrial: agroecosystem; terrestrial: mountains; others (please specify). (choice matrix. exclusive column options for each option row: frequent use; less frequent use; occasionally required; not required) q13. which data at the ecosystem level are the most required by you and at what scale? options/suboptions: ecoregions; vegetation coverage; protected areas; temperature; precipitation; soil; watersheds; basins; others (please specify). (choice matrix. non-exclusive column options for each option row: global; regional; national; provincial; local) section (e): gbif cna survey: species-level data requirements. the objective of this section is to understand data on which taxa’s is most often required. q14. which data at the plant species level are most required by you and at what scale? please specify child taxa or common names in the box below. options/suboptions: plants: monocots; plants: dicots; plants: bryophytes; plants: pteridophytes; plants: gymnosperms; plants: algae; plants: others (please specify). (choice matrix. non-exclusive column options for each option row: global; regional; national; provincial; local) q15. which data at the animal species level are the most required by you and at what scale? please specify child taxa or common names in the box below. options/suboptions: phylum: acanthocephala; phylum: annelida; phylum: arthropoda; phylum: brachiopoda; phylum: cephalorhyncha; phylum: chaetognatha; phylum: chordata; phylum: cnidaria; phylum: ctenophora; phylum: echinodermata; phylum: echiura; phylum: ectoprocta; phylum: entoprocta; phylum: gastrotricha; phylum: gnathostomulida; phylum: hemichordata; phylum: mesozoa; phylum: mollusca; phylum: myxozoa; phylum: nematoda; phylum: nemertea; phylum: onychopora; phylum: phoronida; phylum: placozoa; phylum: platyhelminthes; phylum: porifera; phylum: rotifera; phylum: sipuncula; phylum: tardigrada; others (please specify). (choice matrix. nonassessment of user needs – ariño et al. 62 exclusive column options for each option row: global; regional; national; provincial; local) q16. which data at the fungi, virus and microbial species level are most required by you and at what scale? please specify child taxa or common names in the box below. options/suboptions: microbes; fungi; virus; others (please specify). (choice matrix. non-exclusive column options for each option row: global; regional; national; provincial; local) q17. what are the most important characteristics that you generally want for species occurrence data? options/suboptions: precise/accurate geo-referenced data; metadata on uncertainty about geographical/georeferenced data; pre-1990 data; post1990 data; type specimens in scientific collections; source of information; images; synonyms of species name; common name of species; species habitat descriptions; others (please specify). (non-exclusive multiple choice) section (f): gbif cna survey: usefulness of gbif mobilised data q18. does gbif mobilised data satisfy your needs? options/suboptions: no, i have not at all used gbif mobilised data; no, not at all useful for my applications; maybe, partially useful for my applications; yes, completely useful for my applications; please specify for what applications you use gbif data: (exclusive multiple choice) q19. if gbif mobilised data is partially or absolutely not useful for your applications, we would like to know which needs are not satisfied by the gbif mobilised data? options/suboptions: type of data; data volume/quantity; spatial extent; taxonomic coverage; georeference quality; age of data; sequence based associated occurrence data; others (please specify): (nonexclusive multiple choice) q20. what type of data would you like to see becoming increasingly discoverable and accessible through gbif? options/suboptions: taxonomic names/checklist data; specimen based occurrence data; observation based occurrence data; multimedia resources based occurrence data; other types of observations/occurrences data (e.g. agro-forestry, fish landing, migration etc.); names and occurrences extracted from publications; sequence based associated occurrence data; any other (please specify): (non-exclusive multiple choice) q21. if you have any comments not covered by the survey, feel free to enter them here. [free-text answers] surveymonkey (http://www.surveymonkey.com) was used to design and host the survey. on may7 th , 2009, the survey was launched in english, french and spanish. chinese (traditional and simplified) and russian versions of the survey were launched a week later on may 14 th , 2009 (gbif 2009b, gbif 2009c, gbif 2009d). while english, french and spanish survey versions were closed on june 12 th , 2009, chinese and russian versions were drawn to a close on june 19 th , 2009. survey announcements were widely circulated using, (a) gbif communications portal, (b) gbif mailing lists, (c) taxacom, (d) international commission on zoological nomenclature, (e) taxonomic database working group, and (f) expert centre for taxonomic identification mailing list. the gbif secretariat made a request to the convention on biological diversity (cbd) secretariat to disseminate the launch of the survey, and this request was implemented by the cbd secretariat. task group members also forwarded requests to other national, professional or subjectspecific lists and networks. analytical methods surveymonkey output was supplied as a set of excel tables for each version, recording individual respondents in rows and each single possible option for each question as a column. cells were filled with the selected, verbatim options (see fig. 1). as the number of options exceeded excel’s maximum column capacity, additional excel books were produced by the site holding additional columns. in all, twelve sheets (two for each distinct survey) were downloaded. http://www.surveymonkey.com/ assessment of user needs – ariño et al. 63 figure 1: a small section of one of the raw files as produced by the survey software, arranged in an excel spreadsheet. each row corresponds to one respondent (personal data obscured). as this layout was not amenable to direct analysis (chavan et al., 2010), a 48,767-record database was constructed where each record was an individual option or response supplied by each respondent to each question (see fig. 2). in order to nullify language differences between surveys, free-text answers coming from fixed options were then recoded homogeneously across all six surveys, and merged together into a single file. the original language was however retained as a field, allowing for grouping when the language factor was needed later in the analysis. also, verbatim responses (in their original language, before recoding) were retained for reference as fields. figure 2: data arranged as a database. each row is an individually selected option in the survey, with fields for language (surveylang), question and option number (vname), verbatim answer (vcontent), and recoded (language-free) answer (vname-c). assessment of user needs – ariño et al. 64 this recoding also allowed for the original number of variables in the survey output (one for each possible answer in multiple-choice questions) to be greatly reduced to one variable for each question. in the case of multiple-choice range questions, variables were created where a weighted index substituted several individual options within a range by the centroid of the chosen options. thus, a final, unified excel datasheet 1 was built from the database for subsequent analyses containing numerical data. in addition, 3,883 verbatim, free-text answers and comments 2 were compiled together after translating into english some 1,873 from the original traditional chinese (cn-t), simplified chinese (cn-s), spanish (es), french (fr), and russian (ru) languages. where appropriate, some of these answers were in turn coded to gather frequency data, in order to address emergent questions not included within the surveys at the outset. the unified datasheet was checked for duplicates, errors and mismanagement, and summary statistics and frequency data were compiled (fig. 3). a number of additional data were collected from other sources for further analysis, e.g. the respondent’s city’s coordinates were taken from geo-location facilities. for the majority of the questions, we analysed responses by frequency analyses, either directly on the data variables, or on cross-tabulations among variables. frequencies were plotted or mapped as appropriate in order to address trends from questions either originally designed in the survey’s goals, or emerging from the analytical process. results and discussion we will present the main results here along with a short discussion relevant to each result. more detailed discussion of the survey results, in the context of biodiversity conservation challenges, can be found in faith et al., this volume. 1 the unified excel datasheet has been archived and is available for further analysis on request. 2 the full set of 3,883 verbatim, free-text answers and comments has been archived and is available for further analysis on request. survey characteristics the survey received 750 distinct responses from 77 countries (table 2). however, most respondents were from taiwan (157), spain (124), usa (85), mexico (64), and canada (50). thirtyone countries (40%) provided a single response each. two-thirds of responses came from developed countries (advanced economies as defined by the international monetary fund, 2009), the number of responses appearing to be dependent on the economic power of the country (figs. 4 and 5), although slightly more so on size-dependent wealth (fig. 5, right) than relative wealth (fig. 5, left.) responses according to imf/un category advanced economies emerging and developing economies least developed economies figure 4: number of responses received according to the development status of the country. classes based on the imf database, 2009, and united nation’s office of the high representative for the least developed countries, landlocked developing countries and the small island developing states (un-ohrlls, 2010.) http://www.imf.org/external/pubs/ft/weo/2009/01/weodata/groups.htm http://www.unohrlls.org/en/ldc/related/62/ assessment of user needs – ariño et al. 65 figure 3: flow chart of the analytical design for content needs assessment (cna) survey. cn-s: simplified chinese; cn-t: traditional chinese; db: database; en: english; es: spanish; fr: french; ru: russian; qc: quality control. survey design localization cn-s cn-t ru fr es en survey review delocalization working db database checking and qc frequency analysis report recoding data collection en cn -s cn -t ru fr es en cn -s cn -t ru fr es en cn -s cn -t ru fr es format: surveymonkey reformatting and databasing en cn-s cn-t ru fr es format: database translation compilation su r v ey f ie ld in g cna model exercise [start] [end] free-text extraction en cn -s cn -t ru fr es en assessment of user needs – ariño et al. 66 table 2: list of countries of origin of the received answers and their iso 3166-1 alpha-3 three-letter code (iso, 2007.) code country arg argentina aus australia aut austria bdi burundi bel belgium bgd bangladesh bol bolivia bra brazil can canada cmr cameroon col colombia com comoros cri costa rica cub cuba cze czech republic che switzerland chl chile chn china deu germany dnk denmark dom dominican republic ecu ecuador egy egypt esp spain est estonia fin finland code country fra france gbr united kingdom gnq equatorial guinea gtm guatemala idn indonesia ind india irl ireland isl iceland isr israel ita italy jpn japan lby libya lca st. lucia lso lesotho lva latvia mex mexico mli mali mlt malta mus mauritius mwi malawi nga nigeria nld netherlands nor norway npl nepal nzl new zealand pak pakistan code country per peru phl philippines pol poland prt portugal reu réunion rom romania rus russia scg serbia and montenegro sgp singapore slv el salvador sur suriname svn slovenia swe sweden syc seychelles tgo togo tjk tajikistan tur turkey twn taiwan tza tanzania ury uruguay usa united states ven venezuela vnm vietnam zaf south africa zar congo, drc figure 5: number of responses according to gross national income per capita and gross domestic product (world bank, 2010.) note log scales. answers vs. relative wealth 1 10 100 1000 1 10 100 1000 10000 100000 1000000 approximate per capita income ($) n u m b e r o f a n s w e rs p e r c o u n tr y answers vs. economic power 1 10 100 1000 1 100 10000 1000000 100000000 gross domestic product (m$) n u m b e r o f a n s w e rs p e r c o u n tr y assessment of user needs – ariño et al. 67 the geographical spread of the respondents is depicted in figure 6. there is a high concentration of respondents from the northern hemisphere (developed countries), and there are also apparent geographical gaps, such as russia and china. most respondents used the english version (43%), followed by spanish (32%), chinese (19%), french (5%), and russian (1%). among gbif participant countries, 38 responded and provided most responses (89%), representing 50% of all responding countries, although about half of the participants provided very few responses, less than five each: che, nld, aut, cze, idn, per, pol, prt, cmr, egy, est, isl, jpn, nor, svn, tza, cri, irl, pak, phl. furthermore, eleven gbif participant countries did not respond: ben, bgr, gha, gin, kor, mdg, mar, nic, png and svk (fig. 7). in general, respondents appear to have used their own language to respond the survey (table 3), although some did select the en version even though a localised version was available. in fact, the most common assumed (vernacular, official, or widely used in the country of origin) language among all respondents was spanish (242 respondents, vs. 197 english speakers). it seems therefore apparent that the translation effort resulted in a higher turnout for the survey than if it had been in en only. -90 -60 -30 0 30 60 90 -180 -150 -120 -90 -60 -30 0 30 60 90 120 150 180 figure 6: geographical location of respondents. each dot represents one or more respondents. assessment of user needs – ariño et al. 68 0 25 50 75 100 125 150 twn esp usa mex can fra gbr deu ind aus zaf bra fin dnk ita swe arg col nzl ven bel chl chn rus ury che nld aut cze idn per pol prt tur bgd cmr ecu egy est gtm isl jpn nor rom svn tza dom bdi bol com cri cub gnq irl isr lby lca lso lva mli mlt mus mwi nga npl pak phl reu scg sgp slv sur syc tgo tjk vnm zar c o u n tr y ( is o c o d e ) number of respondents cn en es fr ru figure 7: breakdown of respondents per language and country. a comparison with a similarly-circulated survey by the global strategy and action plan for the digitisation of natural history collections task group that was issued in en only (berendsohn et al., 2010; vollmar et.al., 2010) shows that nonenglish speakers were much less responsive when lacking the localised surveys (figure 8). countries rom gtmecu bgd tur ita bra sw e arg fin dnk zaf aus ind deu gbr fra can m ex usa esp tw n col nzl ven chl chn rus ury bel in gbif not in gbif languages cn 19% en 43% es 32% fr 5% ru 1% assessment of user needs – ariño et al. 69 table 3: percent of speakers of a main language (rows) using the language-specific survey (columns) in the cna survey. assumed language of respondent language of survey en es cn fr ru en 185 3 0 8 1 es 21 221 0 0 0 cn 4 0 141 0 0 fr 2 0 0 28 0 ru 0 0 0 0 5 other 97 7 0 0 1 assumed vernacular/official languages of respondents to cna (outer) and nhc (inner) surveys other 23% ru 1% other 15% es 11% en 66% cn 20% es 33% en 27% fr 4% en es cn fr ru other figure 8: comparison between the assumed languages (vernacular, official, or widely used in the country of residence) of more than 700 respondents to the cna survey (outer ring) and more than 200 respondents to the gsap-nhc survey (inner ring). cna respondents could choose among six different surveys (en, es, cn-s, cn-t, fr, ru; for simplicity, both chinese surveys, traditional and simplified, have been merged here). gsap-nhc respondents were issued only an en version. respondents were less responsive to the en-only survey. for example, no responses to the gsap-nhc survey came from ru, fr or cn-speaking countries, and the es response was much higher when an es survey was available. in the gsap survey, “other” includes the following languages in descending frequency order: nl, sv, pt, de, da, fi, it, ms, ar, he, ja, sq, ur. (iso 639-1 codes.) assessment of user needs – ariño et al. 70 although countries mobilizing more data were also providing more responses, some countries had a very low turnout, with three or less responses each: cri, isl, jpn, nor, svn, aut, per, pol, prt (figure 9). one of the gbif participant countries (south korea) did mobilize data but did not provide any responses, but 49 non-participant countries did provide responses (fig. 10): bdi, bgd, bol, bra, chl, chn, cmr, col, com, cub, cze, dom, ecu, egy, est, gnq, gtm, idn, ind, irl, isr, ita, lby, lca, lso, lva, mli, mlt, mus, mwi, nga, npl, phl, reu, rom, rus, sgp, slv, sur, syc, tgo, tjk, tur, tza, ury, ven, vnm, zaf, zar. responses vs. amount of mobilised records 1 10 100 1000 1 100 10000 1000000 100000000 mobilised records n u m b e r o f a n s w e rs p e r c o u n tr y figure 9: responses from gbif participant countries vs. volume of data mobilization (gbif, 2009.) the above results, especially the low turnout from a number of gbif participant countries, suggest a need for improved coordination by the gbif participant nodes in conducting similar surveys. this highlights the gains for gbif as a community to be made from improved outreach and public relations. most of the survey respondents were academic (45%) or research (26%) (figure 11), but surprisingly, ngos were poorly represented (5%). this suggests that either ngos were not sampled adequately, or the ngos do not actually use the type of data mobilised by gbif. 68 respondents (9%) specified other types or made clarifications, although most could actually be included within the predefined types. the most common “other” types listed were those related with the administration or national, state, or county government (27) and museums, herbaria or botanical gardens (15), although many also included this institution within academic or research institution. a number of respondents made clarifications because it was not possible to tick more than one predefined answer. the majority of the respondents were active within biodiversity research (69%) or conservation science (59%), including taxonomic research (because the survey here posed a multiple-choice question, respondents could select more than one area). a second group of interest included management and education, chosen each by onethird of the respondents (figure 12). . assessment of user needs – ariño et al. 71 figure 10: countries mobilising data through gbif network (gbif, 2009) vs. countries providing responses. green: mobilising and responding; yellow: mobilising but not responding (kr); saffron: responding but not mobilising; blank: neither mobilising nor responding. 0 150 300 450 600 750 academic / educational institution research institution national agency non governmental organisation (ngo) individual researcher or naturalist (e.g. citizen scientists) private company intergovernmental organisation (igo) or multilateral convention others (please specify) number of responses ru fr en es cn figure 11: organisations that responded to survey (table 1: user profile –q2.) assessment of user needs – ariño et al. 72 0 150 300 450 600 750 biodiversity conservation science (including taxonomic research) natural resources management exhibition / educational / academic bioproductivity / bioprospecting (agriculture, fisheries, forestry, etc.) biotechnology biosecurity industrial / commercial use of natural resources biomedical and/or public health others (please specify) number of responses ru fr en es cn figure 12: main interest/business of respondent organisations. more than one option was available to each respondent. b io d iv e rs it y c o n s e rv a ti o n s c ie n c e ( in c lu d in g t a x o n o m ic r e s e a rc h ) n a tu ra l r e s o u rc e s m a n a g e m e n t e x h ib it io n / e d u c a ti o n a l / a c a d e m ic b io p ro d u c ti v it y / b io p ro s p e c ti n g (a g ri c u lt u re , f is h e ri e s , f o re s tr y , e tc .) b io te c h n o lo g y o th e rs b io s e c u ri ty in d u s tr ia l / c o m m e rc ia l u s e o f n a tu ra l r e s o u rc e s b io m e d ic a l a n d /o r p u b lic h e a lt h academic / educational institution 240 223 84 136 57 34 18 8 15 17 research institution 150 125 67 57 32 23 14 12 13 9 national agency 37 24 35 9 16 5 10 14 5 1 others 24 18 16 12 3 12 4 3 non governmental organisation (ngo) 27 21 9 7 4 3 6 3 2 1 individual researcher or naturalist (e.g. citizen scientists) 17 16 3 5 3 2 1 2 private company 9 7 10 1 2 1 1 2 3 intergovernmental organisation (igo) or multilateral convention 6 4 5 1 6 1 1 1 1 main interest/business of your organization d e s c ri b e y o u r o rg a n iz a tio n respondents 34% 31% 27% 24% 20% 17% 14% 10% 7% 3% none figure 13: correspondences between type of institution and their main interests (table 1: user profile – q2 & q3.) assessment of user needs – ariño et al. 73 a few respondents (3%) chose not to select any predefined answer but supplied an alternate definition. however, most of these answers could fit within the predefined categories (see annex). some exclusive answers that appeared very focused and could not be readily fit in other categories were: “application of environmental regulations”; “environmental policy and legislation”; “software development”; “sustainable design and construction”; “to promote environmental care and sustainable development”. educational/academic institutions seem proportionally more related to biodiversity and conservation science than their administration counterparts. ngos, in turn, seem more committed to this research or activity. management also lies within the administration, but not so much bioproductivity. (figure 13). uses of primary biodiversity data: using primary biodiversity data (q01) results of the survey (figure 14) show that there are three broad categories of uses for biodiversity data: 1. basic science, as represented by taxonomy, diversity, population dynamics, biogeography, ecology, evolution. these represent the majority of the uses. 2. more applied science, such as genetics, endangered species, studies dealing with migrations and invasions, conservation planning, natural resources management, environmental impact management or climate change impact. list the ways in which you use primary biodiversity data 0 100 200 300 400 500 600 700 taxonomy species diversity and populations biogeographic studies endangered, migratory and invasive species ecology, evolution and genetics conservation planning natural resources management life histories and phenologies education and public outreach impact of climate change environmental impact management sustainable use agriculture, fisheries, forestry and mining environmental regionalisation ecotourism bioprospecting forensics recreation border control and wildlife trade health and public safety human infrastructure planning society and politics art and history nursery and pet industry industrial use others (please specify) ru fr en es cn figure 14: uses of primary biodiversity data. frequency of responses to q01: “list the ways in which you use primary biodiversity data” (table 1.) assessment of user needs – ariño et al. 74 unspecified museum collections databases cited 17-32 times cited 9-16 times cited 5-8 times cited 3-4 times cited twice rest (cited once) it is ; 9 g b if; 10 anthos; 10 c o n a b io ; 1 7 tr o p ic o s ; 1 7 ipni; 14 m a n is ; 7 n a tu re s e rv e ; 6 g e n b a n k; 5 c s ic in de xe s; 5k ew ; 5 f is h b a se ; 5 usda p la nts; 4herpnet; 4 obis; 4 google; 4 mobot; 4 nybg sciweb; 4 paleo portal; 4 sp2k; 3 catbdb; 3 remib; 3 aluka; 3 o r nis; 2 b h l; 2 b ioo n e ; 2 e u r is c o ; 2 e r m s ; 2 t e l a b o ta n ic a ; 2 f a u n a e u ro p a e a ; 2 s y s t a x ; 2 c o l ; 2 in d e x f u n g o ru m ; 2 w o r m s ; 2 how do you access primary biodiversity data? 0 100 200 300 400 500 600 700 through your own field works/surveys: through hardcopy, literature survey (non-digital form): through primary publications (e.g. taxonomic monographs, maps of species observations): through other web based data portals (please specify) through the gbif data portal (http://data.gbif.org) through access to offline digital data sets (cdrom/dvd/tapes etc.) through free and open datasets within and outside of your institution: through reciprocal agreements with other groups/individuals through institutional agreements through payment basis through ftp sites (please specify) through others (please specify) ru fr en es cn figure 15: modes of access to primary biodiversity data. frequency of responses to q03: “how do you access primary biodiversity data?” (table 1.) figure 16: breakdown of database-type access to primary biodiversity data that were specified by 173 respondents. assessment of user needs – ariño et al. 75 3. societal issues, such as ecotourism, recreation, public health, infrastructure planning, etc. these have a low representation overall. these results must be viewed in the light of the types of respondents, which were heavily biased towards research/academic institutions. this accounts for numerous respondents’ links to basic science. (a) accessing data (q03) two main categories can be distinguished here (figure 15). first, data that are deemed trustworthy: one’s own data collected from field work, or surveys, and peer-reviewed data collected from literature. second, data sources assumed to be “less reliable” (because of potential lack of quality checks such as in a peer review, or because of intrinsic lack of confidence in other’s data), such as web portals (including gbif data portal) and other digital data sources. one-third of the respondents used gbif data, either directly from the gbif data portal, or similar access points. therefore, the remaining two-thirds of the respondents who use other resources define a group of potential future contributors to gbif (although many of them might be actually using gbif data as many of these portals are indeed associated with gbif). the fact that the majority of respondents were using portals other than gbif data portal may suggest that national, regional or thematic data portals should be encouraged as part of the gbif community. respondents answering the previous question were asked to provide detailed data. more than two hundred (209) respondents provided sources 3 of which 173 supplied 316 databased/electronic sources. figure 16 summarises these sources. this breakdown allows us to see both the relative importance of online sources, and what sources could eventually be most ‘profitably’ targeted by gbif for integration different types of users tend to use different access mechanisms (figure 17). for example, systematists tend to use data originating through their own work program. access through gbif 3 at the time of publishing of this report, these will be archived and made available for further analysis on request. (green in fig. 17) follows, in general, the same pattern as for other on-line data sources. most access of gbif data appears to be related to “hard”-science, i.e. taxonomy, biogeography, biodiversity, etc. however, the percentage oriented in this way is not as great as that for traditional access means (own/field work, hardcopy literature, etc.) frequency of access (q04) the majority of the respondent users were not able to determine the frequency of access (figure 18). however, nearly two hundred respondents indicated that they access data on a daily basis. further, another one hundred did so on a monthly basis. the breakdown of the frequency of access according to different uses of data (figure 19) shows that basic science data require access more frequently, along with outreach and environmental impact management needs. together, these findings may indicate what fraction of users appear to be depending on data availability. multiple access points (q05) although about one fourth of users like to use more than one data portal (figure 20), the survey results indicate that most users have their own preferred data portal. among these, the majority of the preferred, sole-use, portals is the gbif data portal. using other data access points (q06) among the reasons that respondents put forward for accessing data portals other than gbif, “tradition” was the most frequently cited (figure 21). it should be noted in this context that many data portals existed even before gbif data portal was put in place. thus, “tradition” reflects the “head start” gained by some data portals (“why should i go somewhere else?”). it is noteworthy that a widespread lack of awareness of the gbif data portal is revealed by the survey results. further, the survey reveals that some users are choosing other data portals because of “ease of use”. assessment of user needs – ariño et al. 76 there were 62 respondents (9% of total) providing textual reasons, often under the “other” option in the survey. a noteworthy outcome was that new, unforeseen reasons were put forward (figure 22). the most frequent reason provided qualifications on the basic rationale that it is better to use “known systems” (tradition). however, a number of responses point to gbif portal performance/design issues (15 respondents), or the data quality, coverage, or adequacy (23 respondents). this suggests that the data quality for other access points, as well as breadth, depth, richness and granularity, may be higher than that of the gbif data portal. y o u r o w n f ie ld w o rk s /s u rv e y s h a rd c o p y , lit e ra tu re s u rv e y ( n o n -d ig it a l fo rm ) p ri m a ry p u b lic a ti o n s ( e .g . ta x o n o m ic m o n o g ra p h s , m a p s o f s p e c ie s o b s e rv a ti o n s ) o th e r w e b b a s e d d a ta p o rt a ls t h e g b if d a ta p o rt a l (h tt p // d a ta .g b if .o rg ) a c c e s s t o o ff lin e d ig it a l d a ta s e ts (c d r o m /d v d /t a p e s e tc .) f re e a n d o p e n d a ta s e ts w it h in a n d o u ts id e o f y o u r in s ti tu ti o n r e c ip ro c a l a g re e m e n ts w it h o th e r g ro u p s /i n d iv id u a ls in s ti tu ti o n a l a g re e m e n ts p a y m e n t b a s is o th e rs f t p s it e s species diversity and populations 370 335 320 172 157 126 120 104 70 20 22 8 taxonomy 369 335 328 172 158 122 120 100 65 18 23 8 life histories and phenologies 366 318 310 176 140 146 128 106 70 20 26 8 biogeographic studies 337 310 309 162 150 117 109 101 63 14 20 6 endangered, migratory and invasive species 270 250 243 136 124 107 108 84 60 19 21 9 ecology, evolution and genetics 269 235 229 124 108 87 83 72 49 16 16 5 conservation planning 194 183 173 98 87 86 86 66 49 14 9 5 natural resources management 171 154 149 77 68 84 68 60 50 14 9 5 education and public outreach 159 149 151 91 73 73 68 55 41 13 16 4 impact of climate change 160 150 150 75 77 73 65 65 45 11 7 3 environmental impact management 139 128 131 62 51 62 61 52 35 9 8 5 sustainable use 98 98 86 51 52 51 44 33 33 11 7 4 agriculture, fisheries, forestry and mining 98 89 89 42 42 42 41 33 24 10 1 2 environmental regionalisation 77 76 76 40 40 39 36 34 24 4 3 3 ecotourism 71 64 58 26 23 38 29 24 22 7 5 2 bioprospecting 45 41 42 18 23 21 19 15 12 4 3 2 forensics 40 32 30 16 15 11 17 11 8 4 1 recreation 30 30 23 16 12 18 16 7 10 3 1 border control and wildlife trade 22 22 24 12 9 14 12 8 10 3 1 1 health and public safety 24 24 19 11 9 11 10 10 7 5 3 1 society and politics 18 18 18 12 13 19 11 10 9 1 1 1 human infrastructure planning 19 17 18 13 8 14 13 10 9 3 industrial use 7 11 10 7 8 6 7 2 2 1 nursery and pet industry 12 11 11 7 1 4 5 2 3 1 2 1 others (please specify) 3 5 2 6 2 1 3 2 1 [you do] acces primary biodiversity data [through:] l is t th e w a y s in w h ic h y o u u s e p ri m a ry b io d iv e rs ity d a ta respondents 55% 49% 44% 38% 33% 27% 22% 16% 11% 5% none figure 17. correspondences between declared uses of primary biodiversity data and approaches to access them: cross-frequencies of q01 and q03 (table 1.) assessment of user needs – ariño et al. 77 frequency of access 0 100 200 300 400 500 600 700 daily basis once a month once a quarter bi-annual cannot determine (on need basis) others (please specify) ru fr en es cn figure 18: frequency of access to biodiversity data. frequencies of responses to q04: “frequency of access” (table 1.) c a n n o t d e te rm in e ( o n n e e d b a s is ) d a ily b a s is o n c e a m o n th o n c e a q u a rt e r b ia n n u a l o th e rs species diversity and populations 171 153 76 16 6 17 taxonomy 164 161 73 14 6 16 life histories and phenologies 156 156 80 6 4 20 biogeographic studies 146 140 74 14 6 16 endangered, migratory and invasive species 129 106 51 9 5 16 ecology, evolution and genetics 126 99 54 11 2 15 conservation planning 93 82 42 7 5 12 natural resources management 101 58 36 8 3 10 education and public outreach 82 69 30 6 3 8 impact of climate change 75 63 32 10 2 10 environmental impact management 68 57 25 10 3 9 sustainable use 75 33 17 5 2 6 agriculture, fisheries, forestry and mining 61 30 19 7 5 environmental regionalisation 46 32 16 4 1 4 ecotourism 45 15 13 6 3 3 bioprospecting 25 15 8 3 2 forensics 30 8 4 3 1 recreation 16 6 5 2 1 5 health and public safety 13 9 5 1 2 human infrastructure planning 10 11 3 1 4 society and politics 13 8 5 1 1 border control and wildlife trade 9 12 2 3 nursery and pet industry 7 7 1 industrial use 5 7 1 2 others (please specify) 6 2 1 1 1 frequency of access l is t th e w a y s in w h ic h y o u u s e p ri m a ry b io d iv e rs ity d a ta respondents 26% 23% 21% 18% 15% 13% 10% 8% 5% 3% none figure 19: types of uses of data and their frequency of access: cross-frequency of q01 and q04. assessment of user needs – ariño et al. 78 some of the datasets mobilised through gbif have multiple access points (e.g. obis mobilised data set can typically have three access points gbif data portal, obis portal, and data sets own portal). how do you access such data sets? 0 100 200 300 400 500 600 700 only through gbif portal only through thematic/regional aggregator portal(s) directly through datasets own portal(s) all of the above ru fr en es cn figure 20: multiple access points. frequencies of responses to q05: “some of the datasets mobilised through gbif have multiple access points. how do you access such data sets?” (table 1.) if you are accessing datasets through access points other than the gbif data portal, why? 0 100 200 300 400 500 600 700 have been using these access points for a long time lack of awareness about accessibility through the gbif portal ease of use more specific search features workflow integration others (please specify) ru fr en es cn figure 21: reasons for accessing data through access points other than gbif: frequencies of responses to q06: “if you are accessing datasets through access points other than the gbif data portal, why?” (table 1.) data formats (q07, q08) a large user community would continue to use the popular excel sheet that has become a de facto standard for small world data keeping (“simple is better”). on the opposite side of the spectrum, specialty database formats were not highlighted by respondents (figure 23). a surprisingly high percentage of users would be using maps as images, or would like to have pdf files available from the portals. as these formats lend themselves poorly to data analysis, this suggests that these users simply want alreadyprocessed output that will be used without any further processing. this highlights well the need for continuous improvement of the data portal. forty-five respondents specified formats other than predefined, with a majority favouring access (17) and gis shape files (10). also, a demand for shape files or gis layers could be identified (figure 24). assessment of user needs – ariño et al. 79 own/others are adequate, sufficient, or trusted 37% gbif inadequate (formats, performance, interface) 25% gbif incomplete (data gaps) 13% gbif unreliable (errors) 12% gbif inadequate (data) 10% legal/agreement issues 3% figure 22: breakdown of 62 free-text answers to the question related to reasons for accessing data through access points other than gbif, after recoding (q04, last option; table 1.) select the data formats which you often choose to access the primary biodiversity data. 0 100 200 300 400 500 600 700 excel maps as images tab delimited comma separated values kml mysql (dump) xml others (please specify) ru fr en es cn figure 23: data formats for accessing data. frequencies of responses to q07: “select the data formats which you often choose to access the primary biodiversity data” (table 1.) access 48% shape file, gis map 29% specialty databases (delta, fasta) 6% other spreadsheets; statistical packages 6% other generalpurpose dbms (oracle, dbase, filemaker) 11% assessment of user needs – ariño et al. 80 if gbif were to serve data in other formats, which would be your preference(s)? 0 5 10 15 20 25 30 35 40 shape file, gis mapdata excel access text, csv, prn, matrix file image maps (jpg, gif) dbase pdf other office/oo files sql dumps kml xml filemaker postgres matlab formats not available in gbif formats already available in gbif figure 24: other data formats potentially required from gbif. frequencies of responses to q08: “if gbif were to serve data in other formats, which would be your preference?” (table 1.) other types of data (q09) when asked to list what other types of data are used along with primary biodiversity data, three broad groups are evident (figure 25): (i) geographically-explicit data, including satellite and aerial imagery and related data, environmental data, and land use/infrastructures data (the most sought after); (ii) speciesor habitat-related ecological and taxonomical data; (iii) other specialty data (molecular, genetics, collection and methods, historical, etc.) quality and quantity requirements primary biodiversity data required (q10, q11) more than two thirds of respondents required taxon names to be included among the retrieved data (figure 26), either because of the respondents’ disciplinary biases (a majority of biodiversity-related scientists) or because data meaning or usefulness would require some type of taxonomic ascription. this result is not surprising given that it has been recognized worldwide that the reliance on a correct name is absolute. occurrence data, and descriptive data about the species, both naturally linked to the taxon identifier, are the second group of required data. together, these two types of data appear to form the core of “biodiversity data”. a more specific type of data appears as a third requirement: distribution data that may be used for modelling, such as occurrence data (including absence data), and population and population interaction data. respondents free-texting “others” offered a wide range of options, although most could actually be included within the pre-defined types (fig. 27). among the particular types (but always with low frequency) were some that might not be properly considered primary biodiversity data, such as “risk status”, “invasiveness” or “interactive keys” (see annex for a full list). it is illustrative to observe the importance given to assessment of user needs – ariño et al. 81 certain types by certain language-specific surveys, such as “conservation/risk” or “habitat data”. respondents seemed to agree that for their uses, hundreds to thousands of pbrs seemed adequate (fig. 28) although requirements varied with the type of data: the biggest requirement seemed to be occurrence records, presumably for species monitoring programmes (1,000 datapoints). multimedia resources are required in less quantities, about 500 on average. data-intensive environments (q12, q13) participants were asked to identify which environments consumed more pbr in their experience. terrestrial environments dominate among respondents (fig. 29), and more so for mountain environments and temperate forests. this plot, however, may also reflect the transect across the interest fields of respondents, or might eventually depict the composition of the scientific body related to primary biodiversity data. it is significant, though, that the lowest frequency lies in deserts and deep seas (harsh environments). given that the set of respondents was not randomly stratified over different environmenttypes or biomes, we cannot draw conclusions about the most important, highest priority, context for new data requirements. nevertheless, the results do show that, no matter what the environment/habitat of interest, there is a general call for more/better primary biodiversity data. list the other types of data you use together with primary biodiversity data 0 50 100 150 200 250 satellite, aerial imagery environmental data and layers (climate, cover, salinity, etc.) gis, geo data, terrain, map-related features land use, territory organization, infrastructures habitat data, ecosystem classifications, region categories legal, list status of species and habitats, management data taxon-related info: morphology, phenology, abundance, behavior, etc. taxon images, photographs molecular-related data (dna, genetics, etc.) literature, bibliography data collection/collecting data biosecurity risk ru fr en es cn figure 25: types of data that are used along with primary biodiversity data (table 1, q09: “list the other types of data you use together with primary biodiversity data?”). respondents could use free-text and specify several types each. responses4 have been recoded into frequently-mentioned categories, or ascribed to them. 4 these 284 verbatim answers have been archived and are available for further analysis on request. assessment of user needs – ariño et al. 82 types or nature of primary biodiversity data required? 0 100 200 300 400 500 600 700 taxonomic names / checklists occurrence records (presence only) species information (descriptive data) occurrence records (including absence records) population density / dynamics data species interaction data others (please specify) ru fr en es cn figure 26: types of primary biodiversity data required. frequencies of respondents to q10: “types or nature of primary biodiversity data required?” (table 1.) types or nature of primary biodiversity data required? other (please specify) 0 20 40 60 habitat data taxonomy, phylogeny, morphology range, population, diversity data conservation status, threats collection/collector metadata uses, medical, economical importance biology invasiveness molecular, dna, genetic data images literature data, references fr en es cn figure 27: composition of the “new data types”, listed under “others” in q10 (fig. 26), and described by 77 respondents (table 1 and annex). assessment of user needs – ariño et al. 83 figure 28: average quantity of data required for each type (table 1, q11: “quantity of data required for each type?”). each respondent was given a choice of order-of-magnitude levels, but could select multiple levels. to allow comparisons, for each respondent selecting more than one option in the range, the centroid of the selected options was calculated (see methods). the coloured dots are the averages of the selected single ranges (or centroids of multiple ranges) across all respondents within the language-specific survey. the totals (big circles) follow the same rule but are not restricted to language-specific surveys. thus, they represent the average across all respondents (not across surveys) and yield the best estimates of quantity of required data for each type, based on the largest number of responses and irrespective of language. when respondents were asked to focus at the ecosystem level, and identify the scales at which these primary biodiversity data were needed or useful, no clear pattern emerged. the requirements were fairly well spread over all ranges and ecosystem types. as shown in figure 30, the range scale includes global (broadest) down to local (narrowest). we highlight the fact that, in some language-specific versions (fr, es) of our survey, the term “regional” may have been misunderstood (in these languages it means something below national level, not above it). given the high number of spanish-speakers among the respondents, this effect (that cannot be tested from the dataset alone) may have had a large effect on this distribution of responses. 0 1 2 3 4 taxonomic names/checklists occurrence records population density/dynamics multimedia resources others (please specify) quantity of data 10 102 103 104 cn fr en ru es total assessment of user needs – ariño et al. 84 figure 29: frequency of use or need for primary biodiversity data according to environment type: table 1, q12: “for which type of environments do you use/need more primary biodiversity data?”. see figure 26 for an explanation of the metrics. local provincial national regional global ecoregions 97 98 156 206 132 protected areas 135 116 196 149 76 vegetation coverage 144 127 158 163 63 temperature 160 114 131 145 74 precipitation 160 110 126 134 57 soil 155 103 95 103 39 watersheds 108 98 107 105 33 basins 97 89 94 95 25 others 15 9 12 15 10 required scale e c o s y s te m l e v e l respondents 45% 40% 36% 31% 27% 22% 18% 13% 9% 4% none figure 30: ecosystem level data requirements. frequencies of responses to q13: “which data at the ecosystem level are the most required by you and at what scale?” (table 1.) 0 1 2 3 4 marine: coasts marine: oceans marine: deep seas marine: islands marine: estuarine inland: wetlands inland: river basin inland: lakes terrestrial: tropical forests terrestrial: temperate forests terrestrial: deserts terrestrial: grasslands terrestrial: agro-ecosystem terrestrial: mountains others (please specify) frequency no occas. less freq. freq. cn fr en ru es total assessment of user needs – ariño et al. 85 species-level data requirements plants the survey revealed that data at all scales (from local to global) were in demand, with no taxonomic pattern across scales (please note the caveat on “regional”, above). taxonomically, higher plants were slightly more in demand, although not significantly so (fig. 31). this may well reflect either the spread of active taxonomists across groups, or an imbalance in the use of taxa for ecological studies. local provincial national regional global dicotyledons 151 121 162 135 104 monocotyledons 140 113 149 119 99 gymnosperms 109 84 118 93 70 pteridophytes 89 67 92 81 53 bryophytes 78 53 74 58 38 algae 71 41 70 59 43 plants: others 15 13 17 14 17 required scale p la n t s p e c ie s l e v e l respondents 40% 36% 32% 28% 24% 20% 16% 12% 8% 4% none figure 31: data requirements at plant species level. frequencies of responses to q14: “which data at the plant species level are most required by you and what scale?” (table 1.) animals as depicted in figure 32, three taxon groups heavily dominate the data needs: arthropods, vertebrates, and (in a lower tier) molluscs. the remaining groups are mentioned by respondents much less often. users seemed to demand data more often for higher animals, as well as species occurrence data related to ecological and public health factors. however, this outcome may reflect existing biases in the actual body of zoological knowledge and the taxonomic coverage: vertebrates have been traditionally well studied (constituting the vast majority of gbif-mediated available data, with birds and fish dominating), and are the focus of many conservation programs. at the same time, entomological data needs include the largest groups of pests and other species of interest. other taxa interest in other organisms was halved among respondents as compared to plants and animals (fig. 33). despite the ecological importance of microflora, data on these were only halfway in demand relative to fungi and viruses. again, this result may reflect the traditional paucity of ecosystem-level studies on these groups, rather than a genuine lack of interest in these extremely important groups. as before, no particular pattern was detected in the geographical width (scale) of the requirements assessment of user needs – ariño et al. 86 local provincial national regional global arthropoda 111 92 125 111 118 chordata 91 81 110 91 81 mollusca 60 42 63 49 44 annelida 33 25 42 29 26 cnidaria 30 16 29 29 27 echinodermata 26 17 33 24 18 brachiopoda 24 15 30 22 19 porifera 22 15 29 22 22 platyhelminthes 24 14 29 18 21 ctenophora 22 12 27 18 18 acanthocephala 19 12 28 19 15 nemata 20 11 23 18 18 rotifera 19 12 25 18 16 chaetognatha 20 9 26 18 14 tardigrada 20 11 21 15 17 sipuncula 20 11 24 16 12 nemertea 18 11 24 14 15 echiura 20 10 23 14 14 entoprocta 19 11 22 15 14 hemichordata 18 9 22 15 13 ectoprocta 17 10 22 13 14 gastrotricha 17 8 22 13 13 cephalorhyncha 15 6 22 15 12 mesozoa 16 7 21 13 13 myxozoa 16 6 21 13 13 phoronida 17 6 21 14 11 placozoa 15 6 20 14 14 cycliophora 16 6 20 13 13 onychopora 15 6 21 16 10 gnathostomulida 16 6 20 13 12 animals: others 11 5 10 10 11 required scale a n im a l s p e c ie s l e v e l respondents 36% 32% 29% 25% 22% 18% 14% 11% 7% 4% none figure 32: animal species level data requirements. frequencies of responses to q15: “which data at the animal species level are the most required by you and at what scale?” (table 1.) local provincial national regional global fungi 83 69 85 68 58 microbes 35 29 53 38 41 virus 21 13 34 20 29 others 3 1 5 5 8 required scale respondents 41% 37% 33% 28% 24% 20% 16% 12% 8% 4% none figure 33: data requirements at microbes, fungi and virus species level. frequencies of responses to q16: “which data at the fungi, virus and microbial species level are most required by you and at what scale?” (table 1.) assessment of user needs – ariño et al. 87 types of species-level occurrence data as expected, precise and accurate georeferenced data is the most important, desirable, property for most users who are interested in working with species occurrence data. this goal is followed by the desire for species habitat descriptions, taxonomic accuracy, and also ancillary data associated with specimens. perhaps surprisingly, users also look for images describing species and its habitat. this latter need may be satisfied by referring to appropriate pages (where available) from encyclopedia of life (encyclopedia or life, 2011). usefulness of gbif mobilised data cases where gbif provides adequate data as depicted in figure 35, more than half (55%) of the respondents suggest that gbif mobilised data met their needs, either completely or partially. a small minority of the respondents felt that gbifmobilised data were not useful for them at all (see below). at the same time, a large cluster of respondents (41%) had never used gbif-mobilised data. this highlights the obvious, but critical, point that there always is a need to undertake efforts to encourage the use of gbif mobilised data, especially if the main reason for lack of use is simply lack of awareness of the resource. we attempted a cross-tabulation of two fundamental factors relating to utility: the perceived usefulness of gbif mobilised data, and the types of data required by the users. the fact that the majority of respondents make at least partial use of gbif mobilised data suggests that these data are widely used for basic tasks, including initial exploratory analyses. it is interesting that there was a 50-50 split amongst the respondents who used gbif-mobilised taxonomic and/or occurrence data, and those who had never used such data. respondents who indicated that gbif mobilised data were useful had also expressed a strong need for descriptive information (at the species level). types or nature of primary biodiversity data required? 0 100 200 300 400 500 600 700 precise/accurate geo-referenced data species habitat description synonyms of species name type specimens in scientific collections images source of information post-1990 data pre-1990 data metadata on uncertainty about geographical/georeferenced data common name of species others please specify ru fr en es cn figure 34: required characteristics for species occurrence data. frequencies of responses to q17: “what are the most important characteristics that you generally want for species occurrence data?” (table 1.) assessment of user needs – ariño et al. 88 figure 35: percentages of people finding gbif-mobilised data useful. frequencies of responses to q18: “does gbif mobilised data satisfy your needs?” and breakdown of frequencies according to types or nature of data required (q10). cases where gbif data do not cover user needs figure 36 collates a number of concerns about why gbif data are not found useful by some respondents. this information here is recoded for analysis into homogeneous groups. it also includes recoded verbatim answers. the term, “more detailed data” collates requests for better coverage, finer detail, finer geo-referencing, gap filling, and the like. the term, “already satisfied by gbif” includes various requests mentioning items or data that in fact are actually provided by gbif. we note that in most cases respondents are not aware of this existing provision. the term, “excessive data” includes cases where respondents referred to data that was deemed not pertinent/not suitable for them, and expressed a preference to remove these from the databases (see annex for a full list). particular answers for the remaining subquestions can be roughly categorized according to the kind of perceived data limitations. three questions focused on the general, quantitative, availability of data (its span across themes), and three focused on the confidence the researchers would put in the data according to perceived quality. some respondents appreciate a general need for more data or better data, while others qualify this lack according to particular fields of research, taxa, or geographical areas. see annex for a full list. answers can be grouped according to these categories as shown in figure 37, where the shading reflects the number of responses placed in each category. most respondents who perceived a lack of quality attributed this to the whole dataset. however, in considering the issue of data coverage, respondents focussed on particular areas, probably according to expertise. it may be assumed reasonably that a number of respondents may be attributing to the whole dataset problems that pertain to their particular field of competence or geographical interests. 41% 4% 43% 12% assessment of user needs – ariño et al. 89 if gbif mobilised data is partially or absolutely not useful for your applications, we would like to know which needs are not satisfied by the gbif mobilised data? (please specify details in the text box in front of each category): type of data/other: 0 20 40 60 80 more detailed data don't know particular taxon groups excessive data, unreliable sources (already satisfied by gbif) names, synonyms species pictures, images, sound, morphology habitat data commercial/anthropic data, i.e. catch, planted, cultivars metadata: completness of data, qualifications type series reference associations between species species biology, distribution tracks fossil data legal status, i.e. protection level fr en es cn figure 36: needs not satisfied by gbif mobilised data. frequencies of responses to q19: “which needs are not satisfied by the gbif mobilised data?”. classes have been recoded from the original options and verbatim answers (see text). note, however, that the number of respondents finding issues is relatively low: from 3.1% of all respondents (age of data) to 6.7% (taxonomic quality). wish list the gbif network is most sought after for the discovery and access it facilitates to species occurrence and names data. demand for occurrence data, having an established basis through specimens, publications, observations, sequences and multimedia etc., is increasing rapidly. this demand is closely matched by that for the names or checklist data (fig. 38). we note that most of 32 respondents who suggested “other” data type in fact were actually suggesting existing types, i.e. images, georeferenced data, etc. some interesting example suggestions, however, could be identified, such as, for example: ornamental, commercial species data from any provenance (ledger-based occurrence data) images and historical data about type series raw literature data, i.e. direct links to electronic publications quantitatively-oriented data, such as in field ecology species management data. assessment of user needs – ariño et al. 90 focus data volume/ quantity spatial extent taxonomic coverage georeference quality taxonomic quality age of data general: more data/better data/all data 22 7 12 28 26 14 qualified: more/better from selected fields/areas/taxa 9 12 23 11 16 8 filling data voids, gaps 3 17 8 2 8 1 ok as of now 0 0 3 1 2 0 n/a 5 0 2 0 6 1 span reliability 0% 25% 50% 75% 100%p e rc e n t o f re s p o n d e n ts n e e d in g d a ta n o t s u p p li e d b y g b if figure 37: group categories for the perceived quality issues in the datasets. some representative examples of needs (interpreted) 5 can be found in box 1. data volume/quantity spatial extent taxonomic coverage georeference quality taxonomic quality age of data download limit not enough results/data scant data for selected taxa or regions spatial bias, gaps africa, tropics, pacific patchy local coverture often missing incomplete coverage of groups some groups absent/poorly covered: acari, geometridae, basidiomycota, etc. poor in places (i.e. africa); imprecise data errors in coordinates, i.e. +/ metadata lacking: datum high concern about identification s: id of taxonomist lacking concern about curation synonymies missing; choice of taxonomies data from all ages needed historical gaps recent revisions missing box 1: representative examples of user data need. 5 these verbatim answers have been archived, and are available for further analysis on request. assessment of user needs – ariño et al. 91 figure 38: demand for types of data among respondents, re-categorised from free-text answers (available upon request). frequencies of responses to q20: “what type of data would you like to see becoming increasingly discoverable and accessible through gbif?” another big issue is the call for dna sequence data and other genetic diversity data (particularly anticipating large scale, next generation sequencing studies.) linkages between dna barcoding of collected specimens, identified to species and geo-referenced, and gbif georeferenced data for the same species, will perhaps be one of the most efficient development for many potential users. conclusions the cna survey exercise has provided a good picture of what users know about the availability of gbif-mediated data, what information they would like to have and use, and what they perceive as lacking (or requiring improvements in availability). importantly, this survey also reveals the wide breadth and scope of interests among the community of potential and actual users. we conclude that gbif could build on its current services, and at the same time steer future developments in order to cater for the full spectrum of users. this goal may be achieved, most likely, by enhancing the gbif linkages to other data types and sources, such as molecular and environmental data. while providing much insight on the expectations of users, the cna exercise also uncovered survey properties that may raise concerns about whether all these aspects of userneeds have been accurately conveyed. a comparison of cna survey with a similar survey conducted in english only highlights dangers in not fully taking into account the diversity of users’ circumstances. this can reduce the degree to which the full spectrum of user-needs are expressed through the survey. in particular, our study clearly shows that the level of responses elicited from target users is strongly tied to the ability to address the user in his/her own language. future surveys should certainly be designed to cover most of the target user population in either their first of second language in order to ensure a wider coverage. the chosen languages in this study (un official languages) might be adequate in this respect. a challenge for future cna work, in setting out to capture the full spectrum of data needs for 0 50 100 150 multimedia resources based occurrence data other types of observations/occurrences data sequences based associated occurrence data observation based occurrence data names and occurrences extracted from publications specimen based occurrence data taxonomic names/checklist data yes, definitively yes, for this particular field yes, but/if/when/provided that… not at all assessment of user needs – ariño et al. 92 biodiversity research, will be to ensure reaching that maximum breadth of researchers, in a representative, homogeneous, manner. new data/information challenges are emerging through national and international programs and activities, including those related to the new post-2010 targets of the convention on biological diversity, the global biodiversity observation network (geo bon) (andrefouet et al., 2008), and the new intergovernmental science-policy platform on biodiversity and ecosystem services (ipbes, 2011). these challenges particularly involve needs to integrate biodiversity with ecosystem services and other needs of society, for research, observations, assessments, and policy development. we touch on some of these issues in the companion paper on cna recommendations (faith et al., this volume). acknowledgements we are grateful to the global biodiversity information facility for supporting and funding this survey through the content needs assessment task group. aha is grateful to the university of navarra for support and infrastructure. references andrefouet s., costello m.j., faith d.p., ferrier s., geller g.n., höft r., jürgens n., lane m.a., larigauderie a., mace g., miazza s., muchoney d., parr t., pereira h.m., sayre r., scholes r.j., stiassny m.l.j., turner w., walther b.a., yahara t., 2008. the geo biodiversity observation network concept document. geo group on earth observations, geneva, switzerland. 45 pp. berendsohn w., chavan v., macklin j., 2010. summary of recommendations of the gbif task group on the global strategy and action plan for the digitisation of natural history collections. biodiversity informatics, 7(2): 67-71. borgman c.l., 2003. from gutenberg to the global information infrastructure : access to information in the networked world. cambridge, ma, usa: mit press. 345 pp. borgman c.l., 2007. scholarship in the digital age: information, infrastructure, and the internet. the mit press. 336 pp. borgman c.l., wallis j.c., enyedy n., 2006. little science confronts the data deluge: habitat ecology, embedded sensor networks, and digital libraries. international journal on digital libraries, 7 (1-2): 17-30. chavan v., sood r., ariño a.h., 2010. best practice guide for ‘data discovery and publishing strategy and action plans’. gbif, copenhagen. encyclopedia of life, 2011. available from http://www.eol.org. last accessed 24 jan 2011. faith d.p., collen b., ariño a.h., koleff p.o., kerr j., guinotte j., chavan v., 2013. bridging the data gaps: recommendations of the gbif content needs assessment task group. biodiversity informatics, 8. global biodiversity information facility (gbif), 2009a. content needs assessment task group. accessible at http://www.gbif.org/informatics/primarydata/task-groups/cna-tg/. last accessed 2010.12.26 global biodiversity information facility (gbif) 2009b. gbif content needs assessment survey 2009. accessible at http://www.gbif.org/communications/news-andevents/showsingle/article/gbif-content-needsassessment-survey-2009/. last accessed 2010.12.26. global biodiversity information facility (gbif) 2009c. gbif content needs assessment survey 2009 available in chinese. http://www.gbif.org/communications/news-andevents/showsingle/article/gbif-content-needsassessment-survey-2009-available-in-chinese/. last accessed 2010.12.26. global biodiversity information facility (gbif) 2009d. gbif content needs assessment survey 2009 available in russian. http://www.gbif.org/communications/news-andevents/showsingle/article/russkaja-versijavoprosnika-po-vyjasneniju-potrebnostei-v/. last accessed 2010.12.26. global biodiversity information facility (gbif), 2010. data portal. countries, territories and islands. http://data.gbif.org/countries/last accessed 2010.12.21 grinnell j., 1910. the methods and uses of a research museum. popular science monthly, 77: 163-169. hill a.w, otegui j., ariño a.h., guralnick r.p., 2010. gbif position paper on future directions and recommendations for enhancing fitness-for-use across the gbif network, version 1.0. copenhagen: global biodiversity information facility, 25 pp. http://mitpress.mit.edu/catalog/item/default.asp?ttype=2&tid=11333 http://mitpress.mit.edu/catalog/item/default.asp?ttype=2&tid=11333 http://www.springerlink.com/content/?author=christine+l.+borgman http://www.springerlink.com/content/?author=jillian+c.+wallis http://www.springerlink.com/content/?author=noel+enyedy http://www.springerlink.com/content/1432-5012/ http://www.eol.org/ http://www.gbif.org/informatics/primary-data/task-groups/cna-tg/ http://www.gbif.org/informatics/primary-data/task-groups/cna-tg/ http://www.gbif.org/communications/news-and-events/showsingle/article/gbif-content-needs-assessment-survey-2009/ http://www.gbif.org/communications/news-and-events/showsingle/article/gbif-content-needs-assessment-survey-2009/ http://www.gbif.org/communications/news-and-events/showsingle/article/gbif-content-needs-assessment-survey-2009/ http://www.gbif.org/communications/news-and-events/showsingle/article/gbif-content-needs-assessment-survey-2009-available-in-chinese/ http://www.gbif.org/communications/news-and-events/showsingle/article/gbif-content-needs-assessment-survey-2009-available-in-chinese/ http://www.gbif.org/communications/news-and-events/showsingle/article/gbif-content-needs-assessment-survey-2009-available-in-chinese/ http://www.gbif.org/communications/news-and-events/showsingle/article/russkaja-versija-voprosnika-po-vyjasneniju-potrebnostei-v/ http://www.gbif.org/communications/news-and-events/showsingle/article/russkaja-versija-voprosnika-po-vyjasneniju-potrebnostei-v/ http://www.gbif.org/communications/news-and-events/showsingle/article/russkaja-versija-voprosnika-po-vyjasneniju-potrebnostei-v/ http://data.gbif.org/countries/last%20accessed%202010.12.21 http://data.gbif.org/countries/last%20accessed%202010.12.21 assessment of user needs – ariño et al. 93 international monetary fund (imf), 2009. world economic outlook database—weo groups and aggregates information. http://www.imf.org/external/pubs/ft/weo/2009/01/w eodata/groups.htm. last accessed 2010.12.21 international standards organisation (iso), 2007. iso 3166 maintenance agency (iso 3166/ma) iso's focal point for country codes. http://www.iso.org/iso/country_codes.htm. last accessed 2010.12.21 ipbes, 2011. intergovernmental science-policy platform on biodiversity and ecosystem services. available online: http://www.ipbes.net/. last accessed on 13 january 2011). kelling, s., w.m. hochachka, d. fink, m. riedewald, r. caruana, g. ballard, g. hooker. 2009. data intensive science: a new paradigm for biodiversity studies. bioscience 59:613-620. united nation’s office of the high representative for the least developed countries, landlocked developing countries and the small island developing states (un-ohrlls), 2010. list of countries. http://www.unohrlls.org/en/ldc/related/62/.last accessed 2010.12.21 world bank, 2010. world development indicators. gni per capita, atlas method (current us$). world bank national accounts data and oecd national account data files. http://data.worldbank.org/indicator/ny.gnp.pcap .cd. last accessed 2010.12.21 vollmar a., macklin j., ford l., 2010. natural history specimen digitization: challenges and concerns. biodiversity informatics, 7(2): 93-112. http://www.imf.org/external/pubs/ft/weo/2009/01/weodata/groups.htm http://www.imf.org/external/pubs/ft/weo/2009/01/weodata/groups.htm http://www.iso.org/iso/country_codes.htm http://www.ipbes.net/ http://www.unohrlls.org/en/ldc/related/62/ http://data.worldbank.org/indicator/ny.gnp.pcap.cd http://data.worldbank.org/indicator/ny.gnp.pcap.cd biodiversity informatics, 15, 2020, pp. 67-68 67 general theory and good practices in ecological niche modeling: a basic guide marianna simões1,2, daniel romero-alvarez2, claudia nuñez-penichet2, laura jiménez2, marlon e. cobos2 1 entomology department, centrum für naturkunde, university of hamburg, d‐20146 hamburg, germany 2 department of ecology & evolutionary biology and biodiversity institute, university of kansas, lawrence, ks, usa abstract. ecological niche modeling (enm) and species distribution modeling (sdm) are sets of tools that allow the estimation of distributional areas on the basis of establishing relationships among known occurrences and environmental variables. these tools have a wide range of applications, particularly in biogeography, macroecology, and conservation biology, granting prediction of species potential distributional patterns in the present and dynamics of these areas in different periods or scenarios. due to their relevance and practical applications, the usage of these methodologies has significantly increased throughout the years. here, we provide a manual with the basic routines used in this field and a practical example of its implementation to promote good practices and guidance for new users. key words— calibration area, geographical information system, maxent, occurrence data, species distribution modeling, variable selection. species distributions are understood to be determined by three limiting factors: movement capacities, abiotic conditions, and biotic interactions. the joint effects of these three factors have been summarized in the so-called bam diagram (soberón and peterson 2005). the principle of ecological niche modeling (enm) is to relate locations where a species is observed with the environmental characteristics of those locations, to estimate conditions that are favorable for the species and consequently its potential geographical range (peterson et al. 2011). recent years have seen an explosion of interest in building models of environmental suitability for species based on presence, absence, and abundance data, to estimate conditions within which a species can maintain populations without immigration subsidies (i.e., the ecological niche; peterson et al. 2011, warren 2012). however, the employment, documentation, and understanding of the rationale for configurations used on such estimates have been poorly documented, detracting robustness from enm studies (cobos et al. 2019, warren 2012). seeking to reduce mistakes, improve standardization, ensure consistency, and provide training on the application of these tools, we provide basic guidelines on the building of ecological niche models. this document is the second in the series of manuals in the biodiversity informatics training curriculum to provide students and researchers interested in the field with a practical guide to construct enms. the first manual discussed some of the steps involved in cleaning biodiversity data, and included a practical exercise (cobos et al. 2018). both manuals are intended to emphasize key theoretical and methodological considerations. the present document and the associated manual are not intended as a detailed treatment of geographical information systems, maxent functions, or to present thorough guidelines (araújo et al. 2019, feng et al. 2019, merow et al. 2013, phillips et al. 2006, phillips and dudík 2008, phillips et al. 2017); in fact, multiple free resources are available, and we are including references to many of them at the end of the manual. rather, we focus on common practices and techniques that are used when estimating the ecological niche and the potential distribution of a species. our goal is to provide a manual for educational purposes, building capacity, and basic comprehension for the broader audience exploring the field, similar to what has been developed for other modeling approaches in the biological sciences (e.g., phylogenetics: hall 2013, o’halloran 2014). here, we are providing theoretical background and a practical exercise to estimate areas of environmental suitabilibiodiversity informatics, 15, 2020, pp. 67-68 68 ty for our chosen model species: the poison dart frog, dendrobates auratus. the full manual and datasets used for the exercise are available at http://hdl.handle.net/1808/30276. acknowledgments we thank the broader university of kansas ecological niche modeling group, which was the context within which we developed this manual. we thank a. townsend peterson for comments on an earlier version of the manuscript. competing interests the authors have declared that no competing interests exist. references araújo, m. b., r. p. anderson, a. m. barbosa, c. m. beale, c. f. dormann, r. early, r. a. garcia, a. guisan, l. maiorano, b. naimi, r. b. o’hara, n. e. zimmermann, and c. rahbek. 2019. standards for distribution models in biodiversity assessments. sci. adv. 5(1):eaat4858. cobos, m. e., l. jiménez, c. nuñez-penichet, d. romero-alvarez, and m. simões. 2018. sample data and training modules for cleaning biodiversity information. biodiv. informatics 13:49–50. cobos, m. e., a. t. peterson, n. barve, and l. osorio-olvera. 2019. kuenm: an r package for detailed development of ecological niche models using maxent. peerj 7:e6281. feng, x., d. s. park, c. walker, a. t. peterson, c. merow, and m. papeş. 2019. a checklist for maximizing reproducibility of ecological niche models. nat. ecol. evol. 3:1382–1395. hall, b. g. 2013. building phylogenetic trees from molecular data with mega. mol. biol. evol. 30:1229– 1235. merow, c., m. j. smith, and j. a. silander. 2013. a practical guide to maxent for modeling species’ distributions: what it does, and why inputs and settings matter. ecography 36:1058–1069. o’halloran, d. 2014. a practical guide to phylogenetics for nonexperts. jove 84: e50975. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo. 2011. ecological niches and geographic distributions princeton university press, princeton, new jersey. phillips, s. j., r. p. anderson, and r. e. schapire. 2006. maximum entropy modeling of species geographic distributions. ecol. model. 190:231–259. phillips, s. j., r. p. anderson, m. dudík, r. e. schapire, and m. e. blair. 2017. opening the black box: an open-source release of maxent. ecography 40:887– 893. phillips, s. j. and m. dudík. 2008. modeling of species distributions with maxent: new extensions and a comprehensive evaluation. ecography 31:161–175. soberon, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodivers. informatics 2:1–10. warren, d. l. 2012. in defense of “niche modeling.” trends ecol. evol. 27:497–500. biodiversity informatics, 14, 2019, pp. 1-7 1 curso modelado de nicho ecológico, versión 1.0 a. townsend peterson1, robert p. anderson2, marlon e. cobos1, martín cuahutle3, angela p. cuervo-robayo3,4, luis e. escobar5, marc fernandez6, daniel jiménez-garcía7, andrés lira-noriega8, jorge m. lobo9, fernando machado-stredel1, enrique martínezmeyer4,10, claudia nuñez-penichet1, javier nori11, luis osorio-olvera10, maría teresa rodríguez3, octavio rojas-soto8, daniel romero-álvarez1, jorge soberón1, sara varela12, y carlos yañez-arenas13 1biodiversity institute, university of kansas, lawrence, kansas 66045 usa; 2department of biology, city college of new york, city university of new york, 160 convent avenue, new york, ny 10031 usa; doctoral program in biology, graduate center, city university of new york, 365 5th avenue, new york, ny 10016 usa; american museum of natural history, central park west at 79th street, new york, ny 10024 usa; 3comisión nacional para el conocimiento y uso de la biodiversidad (conabio), ciudad de méxico, c.p., méxico; 4instituto de biología, universidad nacional autónoma de méxico, méxico city 04510, méxico; 5department of fish and wildlife conservation, virginia tech, blacksburg, virginia 24061 usa; 6centre for ecology, evolution and environmental changes/azorean biodiversity group, and faculdade de ciências e tecnologia, universidade dos açores, ponta delgada, 9501-801, portugal; 7centro de agroecología y ambiente, instituto de ciencias, benemérita universidad autónoma de puebla. edif. val1, ecocampus-buap. km 1.7 carretera san baltazar tetela, san pedro zacachimalpa. c.p. 72960, puebla, puebla, méxico.; 8instituto de ecología, a.c.; carretera antigua a coatepec no. 351, el haya, 91070 xalapa, veracruz, méxico.; 9dept. biogeography and global change, national museum of natural sciences, c/ josé gutiérrez abascal, 2, 288006, madrid, spain; 10centro del cambio global y la sustentabilidad en el sureste ac, cp 86080, villahermosa, tabasco, mexico.; 11instituto de diversidad y ecología animal (idea-conicet) and centro de zoología aplicada, universidad nacional de córdoba,córdoba, argentina; 12museum für naturkunde leibniz institut für evolutions und biodiversitätsforschung. invalidenstraße 43. d-10115 berlin, germany; 13laboratorio de ecología geográfica, unidad de biología de la conservación, parque científico y tecnológico de yucatán, facultad de ciencias, universidad nacional autónoma de méxico, 97302 sierra papacal, yucatán, méxico. resumen.—el conjunto de ideas, métodos y programas informáticos que se conoce como “modelado de nicho ecológico” (mne) el relacionado “modelado de distribución de especies” (mds)— han sido objeto de intensa exploración e investigación en las últimas décadas. a pesar de existir al menos cuatro síntesis publicadas, este campo ha crecido tanto en complejidad, que la formación de nuevos investigadores es difícil. hasta ahora, dicha formación se ha hecho de manera presencial en cursos organizados por universidades o centros de investigación, de los que hemos formado parte como instructores. sin embargo, el acceso a este tipo de cursos especializados es restringido, por un lado, porque los cursos no se ofrecen en todas las universidades, y por otro, porque normalmente se imparten en inglés. para facilitar el acceso a una mayor comunidad de científicos de habla hispana, presentamos un curso en español, completamente digital y de acceso gratuito, que se realizó vía internet durante 23 semanas consecutivas en 2018. aunque las barreras intrínsecas al uso de internet pueden dificultar la accesibilidad a los materiales del curso, hemos usado diversos formatos para la divulgación de los contenidos académicos (video, audio, pdf) con el objetivo de eliminar la mayor parte de estos problemas. abstract.—the suite of ideas, protocols, and software tools that has come to be known as “ecological niche modeling” (enm)—as well as those for the related “species distribution modeling” (sdm)—has seen intensive exploration and research attention in recent decades. in spite of at least four syntheses, the field has grown so much in complexity that it is rather difficult to access for newcomers. until now, accessibility to this field was achieved by in-person courses organized by universities or research centers, in some of which we have participated as instructors. however, the access to these biodiversity informatics, 14, 2019, pp. 1-7 2 specialized courses is limited, on one hand because they are not offered in all universities, and on the other because normally they are taught in english. to expand the access to a wider community of spanish-speaking researchers, here we offer an entirely digital and free-of-charge course in spanish, which was presented over 23 weeks via internet in 2018. although intrinsic internet-related barriers may limit access to course materials, we have made them available in diverse formats (video, audio, pdf) in order to eliminate most of these problems. palabras claves: modelado de nicho ecológico, modelado de distribucion de especies, ecología, biogeografía, curso en línea key words: ecological niche modeling, species distribution modeling, ecology, biogeography, online course especialidad, la ecología distribucional, se ha convertido en un campo difícil de abarcar, provocando que los estudiantes e investigadores que buscan comenzar su formación en la materia puedan llegar a sentirse abrumados e incluso frustrados. aunque la mayoría de la literatura científica sobre la ecología distribucional ha sido publicada en inglés, con ciertas excepciones (mateo et al. 2011; varela et al. 2014; anderson 2015; soberón et al. 2017), investigadores de españa y latinoamérica han sido líderes en el establecimiento y crecimiento del campo. además, la gran biodiversidad en latinoamérica también ha contribuido en la demanda de fuentes de entrenamiento en estas técnicas en la región. gran parte de nosotros, individual y colectivamente, hemos realizado distintos cursos en el campo, contribuyendo con el establecimiento de los principios básicos del aporte digital que aquí tratamos. a pesar de que dichos cursos han mostrado una alta participación y han sido realizados en diversos países (argentina, brasil, chile, colombia, ecuador, españa, méxico, perú y venezuela), estas iniciativas presentan limitaciones obvias si se busca alcanzar a la mayoría de los hispanohablantes interesados en el campo. por lo tanto, elaboramos un curso en línea sobre modelado de nicho, gratuito y en español, durante el período marzo a agosto de 2018. dicho curso fue dictado por un grupo de 21 instructores que trabajan actualmente en seis países, realizando investigación activa y con amplia experiencia en la ecología distribucional. ya que esta iniciativa se encuentra en línea en formato de acceso abierto, presentamos esta contribución como un recurso que esperamos sea de utilidad para una comunidad más amplia. sin dudas, anticipamos que el presente material requerirá ser revisado y actualizado en el futuro, de manera que se mantenga como una herramienta pertinente en el campo de la ecología distribucional. muchas de las preguntas en ecología y biogeografía giran en torno a la posibilidad de disponer de un conocimiento detallado sobre la distribución de las especies. esto incluye la ecología poblacional (e.g., qué condiciones necesita una especie para mantener a una población), biogeografía (e.g., qué o cuántas especies se encuentran en diferentes lugares del planeta), conservación (e.g., qué áreas necesitan ser protegidas para asegurar la supervivencia de un grupo de especies) y el estudio de las invasiones biológicas (e.g., qué áreas pueden ser colonizadas por una especie invasora en particular). el gran valor de este tipo de información relacionada con la distribución geográfica de las especies incrementó recientemente debido a la popularidad de un tipo de métodos generados a finales de los años 70 y principios de los 80 (soto et al. 1984; nix 1986) y que se han desarrollado intensamente en los últimos 20 años (peterson et al. 1999; guisan y zimmermann 2000; elith et al. 2006). la base conceptual de este campo, que podría denominarse “ecología distribucional”, está en los trabajos clásicos de ecología y biogeografía (andrewartha y birch 1964; udvardy 1969), pero su sentido moderno fue desarrollado en las últimas dos décadas. pulliam (2000) desarrolló el primer conjunto de ideas modernas, que luego fue interpretado y reelaborado por soberón y peterson (2005). más tarde se editaron una serie de síntesis extensas (franklin 2010; peterson et al. 2011; peterson 2014; guisan et al. 2017) al mismo tiempo que se generó una serie de trabajos desarrollando nuevos conceptos o discutiendo los fundamentos del campo (araújo y guisan 2006; jiménez-valverde et al. 2008; anderson 2012). además, numerosos avances metodológicos han sido publicados en forma de artículos científicos, opiniones e ideas sobre diversas elecciones metodológicas (e.g., radosavljevic y anderson 2014; zhu y peterson 2017; qiao et al. 2018). por lo tanto, esta biodiversity informatics, 14, 2019, pp. 1-7 3 el curso el curso fue impartido durante 23 semanas. fue organizado en los siguientes tópicos generales: introducción al concepto de nicho ecológico, variables ambientales, datos de presencia, visualización, el área de calibración (m), el diagrama bam (i.e., factores del ambiente abiótico, del ambiente biótico y de movimiento, que afectan distribuciones de especies; sensu soberón y peterson 2005), algoritmos, evaluación de modelos, transferencia de modelos, comparaciones de nichos, aplicaciones y conclusiones. el material de cada semana fue presentado en una serie de video conferencias publicadas cada lunes, junto a versiones descargables en distintos formatos, así como una sesión posterior de preguntas y respuestas cada viernes. la tabla 1 resume el curso y los diferentes elementos incluidos. las presentaciones consistieron en grabaciones de audio de los instructores con captura simultánea de diapositivas. esto aseguró una calidad alta de audio y video, en lugar de instructores filmados frente a una pantalla como en cursos anteriores (peterson y ingenloff 2015). todo el material se encuentra disponible a través de un centro de intercambio de información de enlaces en la página web de biodiversity informatics training curriculum. para el final del tabla 1. resumen del curso de modelado de nicho ecológico, versión 1.0, incluyendo los temas principales, título específico, enlaces para las versiones en formato .pdf, .mp3, y .mp4, así como los enlaces a los videos en youtube, materiales descargables e instructor para cada tema. unidad título pdf mp3 youtube mp4 material adicional instructor introducción plan del curso pdf mp3 yt mp4 a. townsend peterson introducción a la ecología de distribuciones de especies pdf mp3 yt mp4 a. townsend peterson preguntas y respuestas yt todos elementos de una teoría del nicho grinnelliano pdf mp3 yt mp4 enlace jorge soberón mirando un mapa pdf mp3 yt mp4 jorge m. lobo preguntas y respuestas yt todos datos ambientales relación a teoría de nicho pdf mp3 yt mp4 enlace angela cuervo robayo datos ambientales – práctica pdf mp3 yt mp4 daniel jiménez garcía preguntas y respuestas yt todos datos de climas, con práctica en r pdf mp3 yt mp4 enlace angela cuervo robayo aogcms y diferencias entre worldclim y aogcms pdf mp3 yt mp4 sara varela datos de ambientes marinos y su dinámica pdf mp3 yt mp4 marc fernandez preguntas y respuestas yt todos datos de sensores remotos pdf mp3 yt mp4 teresa rodríguez y martín cuahutle procesamiento de datos ambientales: preparación, reducción, selección pdf mp3 yt mp4 claudia nuñez-penichet preguntas y respuestas yt todos datos de presencia relación a teoría de nicho ¿qué son presencias y qué son ausencias? pdf mp3 yt mp4 simões paper, saupe paper jorge soberón unidades de modelado pdf mp3 yt mp4 octavio rojas soto preguntas y respuestas yt todos fuentes de datos pdf mp3 yt mp4 a. townsend peterson georeferenciación pdf mp3 yt mp4 enlace daniel jiménez garcía preguntas y respuestas yt enlace todos control de calidad y reducción pdf mp3 yt mp4 enlace fernando machado stredel subconjuntos para evaluación pdf mp3 yt mp4 enlace a. townsend peterson preguntas y respuestas yt todos http://biodiversity-informatics-training.org/ http://biodiversity-informatics-training.org/ https://www.dropbox.com/s/sveqnux7b8nxeif/semana1_platica1_intro_pdf.pdf?dl=0 https://www.dropbox.com/s/hl0cdsztyoynqed/semana1_platica1_intro_audio.mp3?dl=0 https://youtu.be/br-e9t4pmes https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/cy7ewtpvzscqklz/semana1_platica2_introecoldistr_pdf.pdf?dl=0 https://youtu.be/xfaeybl2hv8 https://www.dropbox.com/s/qo1bbh921ibbx5f/semana1_platica2_introecoldistr_video.mp4?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://youtu.be/0oa2vdhqfjm https://www.dropbox.com/s/s51uj7hahw62a9r/semana2_platica1_introconceptual_final.pdf?dl=0 https://www.dropbox.com/s/t09eq3fyr7pocp9/semana2_platica1_introconceptual_final.mp3?dl=0 https://youtu.be/wuopdqkbmry https://www.dropbox.com/s/busx8yzp77qyupe/semana2_platica1_introconceptual_final.mp4?dl=0 http://nicho.conabio.gob.mx/ https://biodiversity.ku.edu/biodiversity-modeling https://www.dropbox.com/s/9x7e20v28jcuhps/semana2_platica2_mirandounmapa_final.pdf?dl=0 https://www.dropbox.com/s/ws0kt3laqb4mart/semana2_platica2_mirandounmapa_final.mp3?dl=0 https://youtu.be/nhcjkw67uii http://www.biogeografia.org/es/index.html https://youtu.be/fcbgn42vyku https://www.dropbox.com/s/r0sax3koofikt8t/semana3_platica1_relacionateoria_final.pdf?dl=0 https://www.dropbox.com/s/txy6g5kptigrzr2/semana3_platica1_relacionateoria_final.mp3?dl=0 https://youtu.be/48kkd_dgozm https://www.dropbox.com/s/ha9qmuteu4rlt6t/semana3_platica1_relacionateoria_final.mp4?dl=0 https://www.dropbox.com/s/09a1xwndm1x8x3x/ligasdatosambientes.xlsx?dl=0 https://www.researchgate.net/profile/angela_cuervo-robayo https://www.dropbox.com/s/zn9v21os31u7d79/semana3_platica2_datosambientalespractica_final.pdf?dl=0 https://www.dropbox.com/s/m820o25oo3n8bbw/semana3_platica2_datosambientalespractica_final.mp3?dl=0 https://youtu.be/2n0awzga7tg https://www.dropbox.com/s/hxk6z1hl1713xq1/semana3_platica2_datosambientalespractica_final.mp4?dl=0 https://www.researchgate.net/profile/daniel_jimenez-garcia/reputation https://youtu.be/t67iu5qaafs https://www.dropbox.com/s/3zj3d67p6ufte9i/semana4_platica1_cr_datosambientales.pdf?dl=0 https://www.dropbox.com/s/ejloqlpg1cz1pq9/semana4_platica1_cr_datosambientales.mp3?dl=0 https://youtu.be/yzaq2ylzpf4 https://www.dropbox.com/s/1ww618qe0lzmrbi/semana4_platica1_cr_datosambientales_new2.mp4?dl=0 https://www.dropbox.com/s/hvrxagi2pdjghk4/variables_espaciales_new.xlsx?dl=0 https://www.researchgate.net/profile/angela_cuervo-robayo https://www.dropbox.com/s/n8telwtojg2gbxf/semana4_platica2_sv_aogcms.pdf?dl=0 https://www.dropbox.com/s/ruknc3h9scqwpt7/semana4_platica2_sv_aogcms.mp3?dl=0 https://youtu.be/fmfrvbb0f8y https://www.dropbox.com/s/yyoaoj3w6usff80/semana4_platica2_sv_aogcms.mp4?dl=0 https://scholar.google.com/citations?user=8dcki70aaaaj https://www.dropbox.com/s/jj918gearhdyjh7/semana4_platica3_mf_datosmarinos.pdf?dl=0 https://www.dropbox.com/s/wfrj4ara7o72fig/semana4_platica3_mf_datosmarinos.mp3?dl=0 https://youtu.be/vo1ezg7djo8 https://www.dropbox.com/s/fisljobrq9wjvd7/semana4_platica3_mf_datosmarinos.mp4?dl=0 https://scholar.google.com/citations?user=z0koeduaaaaj&hl=es https://youtu.be/37dakg_qkho https://www.dropbox.com/s/hz325br1cokpvs7/semana5_platica1_sensoresremotos_full.pdf?dl=0 https://www.dropbox.com/s/0319ngx72hurrq8/semana5_platica1_sensoresremotos.mp3?dl=0 https://youtu.be/dqlkvrcfw58 https://www.dropbox.com/s/hif3l66akzvm9qw/semana5_platica1_sensoresremotos.mp4?dl=0 http://www.conabio.gob.mx/web/conocenos/cgia_dgg_perrem.html http://www.conabio.gob.mx/web/conocenos/cgia_dgg_perrem.html https://www.dropbox.com/s/sv8mma9o5tmlrzz/presentaci%c3%b3n_claudia.pdf?dl=0 https://youtu.be/chczgyk8age https://www.dropbox.com/s/0zllq5qmfpte32y/semana5_platica2_seleccionambientes_cnp.mp4?dl=0 https://eeb.ku.edu/claudia-nu%c3%b1ez-penichet https://youtu.be/s3qw-3hf3-w https://www.dropbox.com/s/x3fdrh4fhski9rb/semana6_platica1_puntospresencia.pdf?dl=0 https://www.dropbox.com/s/q265mj5d1ztugaf/semana6_platica1_puntospresencia_audio.mp3?dl=0 https://youtu.be/lfi7z3ebdeo https://www.dropbox.com/s/uyzzlci5cqe9lea/semana6_platica1_puntospresencia_final.mp4?dl=0 https://doi.org/10.1111/icad.12288 https://doi.org/10.1016/j.ecolmodel.2012.04.001 https://biodiversity.ku.edu/biodiversity-modeling https://www.dropbox.com/s/iu3jjkucceqfcjc/semana6_platica2_unidadesdemodelado.pdf?dl=0 https://www.dropbox.com/s/1kxg3jz71b3w0c5/semana6_platica2_unidadesdemodelado_audio2.mp3?dl=0 https://youtu.be/jeswm0lvh-c https://www.dropbox.com/s/67i1qf6emur9gvr/semana6_platica2_unidadesdemodelado_final.mp4?dl=0 http://www.inecol.edu.mx/personal/index.php/biologia-evolutiva/4-octavio-rojas-soto https://youtu.be/1ko-o4iagli https://www.dropbox.com/s/pkn4o0bz0i3rwya/semana7_platica1_fuentesdedatos.pdf?dl=0 https://www.dropbox.com/s/vgkrmwk833lqfmo/semana7_platica1_fuentesdedatos_final.mp3?dl=0 https://youtu.be/00aogpot0ti https://www.dropbox.com/s/3vdbd2t7n5zoqi3/semana7_platica1_fuentesdedatos_final.mp4?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/slp2rwmic8ex5i8/semana7_platica2_georeferenciacion.pdf?dl=0 https://www.dropbox.com/s/fst8re4zu6ri5yu/semana7_platica2_georeferenciacion.mp3?dl=0 https://youtu.be/coedhhlnlcs https://www.dropbox.com/s/6fb0jnp1zy2wqv1/semana7_platica2_georeferenciacion_final_1080p.mp4?dl=0 https://www.dropbox.com/s/jt8wunta4n1bmvu/semana7_platica2_georeferenciacion.zip?dl=0 https://www.researchgate.net/profile/daniel_jimenez-garcia/reputation https://youtu.be/0wv-s3juscs https://www.dropbox.com/s/mlnl69z4hbwlgsf/letal_ijhg_2012.pdf?dl=0 https://www.dropbox.com/s/vgiou6rozg0la4x/semana8_platica1_reduccion_final.pdf?dl=0 https://www.dropbox.com/s/join9yoqmkgucri/semana8_platica1_reduccion_final.mp3?dl=0 https://youtu.be/1txrrjyzg3u https://www.dropbox.com/s/drfctngobuspyez/semana8_platica1_reduccion_final.mp4?dl=0 https://www.dropbox.com/s/rcd2x93cpk25x1f/pnp_bboc_2004.pdf?dl=0 https://eeb.ku.edu/fernando-machado-stredel https://eeb.ku.edu/fernando-machado-stredel https://www.dropbox.com/s/xxh9n0hqqzfgbq2/semana8_platica2_subconjuntos.pdf?dl=0 https://www.dropbox.com/s/hcdrfzhfpjo8rza/semana8_platica2_subconjuntos_final.mp3?dl=0 https://youtu.be/wnffnm8zhzu https://www.dropbox.com/s/ie5scbvr8dmapdb/semana8_platica2_subconjuntos_final.mp4?dl=0 https://www.dropbox.com/s/77dmcmzmnw2dnui/muscarella_etal_2014_enmeval.pdf?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://youtu.be/bp3qugvqcac biodiversity informatics, 14, 2019, pp. 1-7 4 visualización visualización de los datos ambientales pdf mp3 yt mp4 datos daniel jiménez garcía intro a nichea pdf mp3 yt mp4 nichea luis escobar nichea: práctica pdf mp3 yt mp4 datos luis escobar preguntas y respuestas yt geoda, enlace m y configuración bam relación a teoría de nicho pdf mp3 yt mp4 enlace carlos yañez arenas estimados de m pdf mp3 yt mp4 enlace a. townsend peterson limitantes en que se puede modelar y que no pdf mp3 yt mp4 enlace a. townsend peterson preguntas y respuestas yt enlace todos algorithmos relación a teoría de nicho pdf mp3 yt mp4 jorge soberón “the good, the bad, and the ugly,” “un solo dios” y balas de plata pdf mp3 yt mp4 a. townsend peterson preguntas y respuestas yt todos incertidumbre: introducción pdf mp3 yt mp4 a. townsend peterson incertidumbre: conceptos básicos pdf mp3 yt mp4 enlace enrique martínez-meyer evaluación de incertidumbre pdf mp3 yt mp4 marlon cobos selección de modelos, y control de sobreajuste en calibración pdf mp3 yt mp4 robert anderson preguntas y respuestas yt ebola, climate models todos evaluación relación a teoría de nicho pdf mp3 yt mp4 a. townsend peterson datos para evaluación pdf mp3 yt mp4 enlace a. townsend peterson probabilidad, favorabilidad e idoneidad pdf mp3 yt mp4 jorge m. lobo preguntas y respuestas yt enlace todos validación, discriminación y calibración: ¿cómo podemos validar un modelo? pdf mp3 yt mp4 jorge m. lobo evaluación no dependiente de un umbral (roc y roc parcial) pdf mp3 yt mp4 enlace a. townsend peterson roc parcial, teoría y práctica. pdf mp3 yt mp4 luis osorio olvera preguntas y respuestas yt enlace todos más sobre evaluación: umbrales, pruebas dependientes de umbral, rendimiento, etc. pdf mp3 yt mp4 enlace a. townsend peterson preguntas y respuestas yt enlace todos transferencia transferencias 1 pdf mp3 yt mp4 enlace a. townsend peterson transferencias 2 pdf mp3 yt mp4 robert anderson pasado, presente, futuro pdf mp3 yt mp4 enrique martínez-meyer preguntas y respuestas yt enlace todos owens et al. – mop y extrapolación pdf mp3 yt mp4 enlace a. townsend peterson implementación de mop pdf mp3 yt mp4 luis osorio olvera preguntas y respuestas yt todos comparación relación a teoría de nicho pdf mp3 yt mp4 enlace a. townsend peterson visualización de modelos pdf mp3 yt mp4 luis escobar comparación de modelos pdf mp3 yt mp4 luis escobar preguntas y respuestas yt enlace todos https://www.dropbox.com/s/qieryq2dwmt3v3q/visualizaci%c3%b3n%20de%20datos.pdf?dl=0 https://www.dropbox.com/s/x9500gbbm7fdl1h/semana9_platica1_visdatosambientales.mp3?dl=0 https://youtu.be/7hfcwxnfqrm https://www.dropbox.com/s/5bhjuie92f41mjy/semana9_platica1_visdatosambientales.mp4?dl=0 https://www.dropbox.com/s/uo60avmj8kbgydz/visdatos.zip?dl=0 https://www.researchgate.net/profile/daniel_jimenez-garcia/reputation https://www.dropbox.com/s/yg0wbgzzsp6ijob/semana9_platica2_introverdatosene.pdf?dl=0 https://www.dropbox.com/s/qzr43nlmnt59l4k/semana9_platica2_introverdatosene.mp3?dl=0 https://youtu.be/mwddljbdpf0 https://www.dropbox.com/s/5wzq0jp1qgmq2zb/semana9_platica2_introverdatosene.mp4?dl=0 https://www.dropbox.com/s/emgajay69xd0s7c/gui%c3%8fa%20de%20instalacio%c3%8fn%20de%20niche%20analyst.pdf?dl=0 http://ecoguate2003.wixsite.com/escobar https://www.dropbox.com/s/bz2680gt5l5rys9/practica_nicheaa.pdf?dl=0 https://www.dropbox.com/s/ma1e22chrpkffbu/semana9_platica3_nichea.mp3?dl=0 https://youtu.be/7ahbcucxig8 https://www.dropbox.com/s/eubyohczdp43zlh/semana9_platica3_nichea_final.mp4?dl=0 https://www.dropbox.com/s/f13oaql816dvgsc/datos_nichea.7z?dl=0 http://ecoguate2003.wixsite.com/escobar https://youtu.be/chmacqqoetc https://gisgeography.com/geoda-software/ https://goo.gl/a6wuau https://www.dropbox.com/s/e038twbqfb3rfpv/semana10_platica1_bam.pdf?dl=0 https://www.dropbox.com/s/426v6pjz0s1y6ka/semana10_platica1_bam.mp3?dl=0 https://youtu.be/yqe0oencduq https://www.dropbox.com/s/ibuwlnzgtwetav9/semana10_platica1_bam.mp4?dl=0 https://journals.ku.edu/index.php/jbi/article/view/4 https://sites.google.com/site/cyanezarenas/ https://www.dropbox.com/s/scqw452unbyojar/semana10_platica2_m.pdf?dl=0 https://www.dropbox.com/s/k3ywrwy4vzrit3v/semana10_platica2_m.mp4?dl=0 https://youtu.be/bnt1kneknzo https://www.dropbox.com/s/k3ywrwy4vzrit3v/semana10_platica2_m.mp4?dl=0 https://www.dropbox.com/s/9u1iax3bbmaibpx/betal_em_2011.pdf?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/ktjnhq1dhmclrp6/semana10_platica3_m_constraints.pdf?dl=0 https://www.dropbox.com/s/paubxnb2wpjajq6/semana10_platica3_m_constraints_final.mp3?dl=0 https://youtu.be/h5vltenfbda https://www.dropbox.com/s/fe64akblkgvjp6i/semana10_platica3_m_constraints_final.mp4?dl=0 https://www.dropbox.com/s/6q8dc9bi2w13n27/setal_em_2012.pdf?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://youtu.be/j_x5g1mmbx4 https://www.dropbox.com/s/6sbvm5mc8avsddu/questions_literature.zip?dl=0 https://www.dropbox.com/s/0ovim0h5k2r05i7/semana11_platica1_teoriainterpretacion.pdf?dl=0 https://www.dropbox.com/s/4dinx9jl2usmu5m/semana11_platica1_teoriainterpretacion.mp3?dl=0 https://youtu.be/xhygjjd_8pk https://www.dropbox.com/s/tr6aba1ty2wqqua/semana11_platica1_teoriainterpretacion.mp4?dl=0 https://biodiversity.ku.edu/biodiversity-modeling https://www.dropbox.com/s/svigadld5k14e7m/semana11_platica2_unsolodios.pdf?dl=0 https://www.dropbox.com/s/zhdpv1heh8fz0yt/semana11_platica2_unsolodios.mp3?dl=0 https://youtu.be/2a-swifgwgg https://www.dropbox.com/s/skcfk5e0srbvcqy/semana11_platica2_unsolodios.mp4?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://youtu.be/f8fos5z-qzy https://www.dropbox.com/s/sjakabjbvhylc99/semana12_platica1_incertidumbre.pdf?dl=0 https://www.dropbox.com/s/tjcnalhjj26efj4/semana12_platica1_incertidumbre.mp3?dl=0 https://youtu.be/z0zn0fzg3xs https://www.dropbox.com/s/i2pfktty5z4yuxa/semana12_platica1_incertidumbre.mp4?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/b1nzwnivc8hcomh/semana12_platica2_incertidumbre.pdf?dl=0 https://www.dropbox.com/s/gsb0k7jv5779eti/semana12_platica2_incertidumbre.mp3?dl=0 https://youtu.be/c8mdzmdxbpy https://www.dropbox.com/s/k7qzf12acg2gb5s/semana12_platica2_incertidumbre.mp4?dl=0 https://www.dropbox.com/s/nnwo8ogz1626flr/literatura%20incertidumbre.zip?dl=0 http://www.ib.unam.mx/directorio/114 https://www.dropbox.com/s/52oywa0atjtj4za/semana12_platica3_estimaci%c3%b3nincertidumbre.pdf?dl=0 https://www.dropbox.com/s/g1b0sbphinknzm9/semana12_platica3_estimaci%c3%b3nincertidumbre.mp3?dl=0 https://youtu.be/lovj_xdevfe https://www.dropbox.com/s/v28vop37v5kd27v/semana12_platica3_estimaci%c3%b3nincertidumbre.mp4?dl=0 https://eeb.ku.edu/marlon-e-cobos https://www.dropbox.com/s/edombmk7sw6slro/semana12_platica4_seleccionmodelos.pdf?dl=0 https://www.dropbox.com/s/52gh7oafmjs2xsb/semana12_platica4_seleccionmodelos.mp3?dl=0 https://youtu.be/515aqjgcto0 https://www.dropbox.com/s/7lzqswvd3ukgx9d/semana12_platica4_seleccionmodelos.mp4?dl=0 http://www.andersonlab.ccny.cuny.edu/ http://youtu.be/dmvjob9mfga https://www.dropbox.com/s/21pol9e3r4g01ad/sp_at_2016.pdf?dl=0 https://journals.ku.edu/jbi/article/view/4955/4476 https://journals.ku.edu/jbi/article/view/4955/4476 https://www.dropbox.com/s/auqbfk6hvi3wdij/semana13_pl%c3%a1tica1_relacionateoria.pdf?dl=0 https://www.dropbox.com/s/w7l0c38vs63x6ag/semana13_pl%c3%a1tica1_relacionateoria.mp3?dl=0 https://youtu.be/wwf9_qkzw2w https://www.dropbox.com/s/fus2jdzv47awlch/semana13_pl%c3%a1tica1_relacionateoria.mp4?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/ifcqzabx0rb1mc7/semana13_pl%c3%a1tica2_datosparaevaluacion.pdf?dl=0 https://www.dropbox.com/s/mcl9my3sfg78460/semana13_pl%c3%a1tica2_datosparaevaluacion.mp3?dl=0 https://youtu.be/osppnf-hiqy https://www.dropbox.com/s/nxidph9vaggbtds/semana13_pl%c3%a1tica2_datosparaevaluacion.mp4?dl=0 https://www.researchgate.net/profile/robert_muscarella/publication/265297099_enmeval_an_r_package_for_conducting_spatially_independent_evaluations_and_estimating_optimal_model_complexity_for_maxent_ecological_niche_models/links/5a128d370f7e9bd1b2c11323/enmeval-an-r-package-for-conducting-spatially-independent-evaluations-and-estimating-optimal-model-complexity-for-maxent-ecological-niche-models.pdf https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/jbw2gbmlsolon1a/semana13_platica3_probabilidasfavorabilidadidoneidad.pdf?dl=0 https://www.dropbox.com/s/t03t18s4z2jjtea/semana13_pl%c3%a1tica3_probabilidadfavorabilidadidoneidad.mp3?dl=0 https://youtu.be/2gsgpgspphi https://www.dropbox.com/s/fucsj0teuhj7lia/semana13_pl%c3%a1tica3_probabilidadfavorabilidadidoneidad.mp4?dl=0 http://www.biogeografia.org/es/index.html https://youtu.be/da7ydhkii8e https://www.dropbox.com/s/ntahkdgqsunhzxo/phillipsdudik2008.pdf?dl=0 https://www.dropbox.com/s/hgj6d3z8mgbngbl/semana14_platica1_validationdiscriminationcalibration.pdf?dl=0 https://www.dropbox.com/s/174j1zyyj3ukp01/semana14_platica1_validationdiscriminationcalibration.mp3?dl=0 https://youtu.be/84pek-rhc9y https://www.dropbox.com/s/agc54pd1xydj4xp/semana14_platica1_validationdiscriminationcalibration.mp4?dl=0 http://www.biogeografia.org/es/index.html https://www.dropbox.com/s/qb3213kqcqdk7sg/semana14_platica2_evaluacionsinumbral.pdf?dl=0 https://www.dropbox.com/s/nghylz87n8dd7b0/semana14_platica2_evaluacionsinumbral.mp3?dl=0 https://youtu.be/av1ryplmtzq https://www.dropbox.com/s/hipzsuxoe7qm1hr/semana14_platica2_evaluacionsinumbral.mp4?dl=0 https://www.dropbox.com/s/tdx5xonmjny6jx8/semana14_lectura.zip?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/br7jax2868mw4gy/semana14_platica3_rocparcial_practica.pdf?dl=0 https://www.dropbox.com/s/bx239pm46fnn3hz/semana14_platica3_rocparcial_practica.mp3?dl=0 https://youtu.be/neystbwujsu https://www.dropbox.com/s/jg8ngz1hgdyj3fw/semana14_platica3_rocparcial_practica.mp4?dl=0 https://luismurao.github.io/ https://youtu.be/uuutyvayxk4 https://www.dropbox.com/s/dwr8bxuzne6dxfm/lecturas.zip?dl=0 https://www.dropbox.com/s/1o8t9yvly1gui50/semana15_platica1_massobreevaluacion.pdf?dl=0 https://www.dropbox.com/s/gn42c1nojn0f6jr/semana15_platica1_massobreevaluacion.mp3?dl=0 https://youtu.be/cn52i75vioe https://www.dropbox.com/s/lqkz0oqzgdzcbcf/semana15_platica1_massobreevaluacion.mp4?dl=0 https://www.dropbox.com/s/7asruqidiabirxa/lecturas_semana15.zip?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://youtu.be/rlxrsuyb2sm https://www.dropbox.com/s/a6qfgyoncgu0n0m/aetal_mve_2009.pdf?dl=0 https://www.dropbox.com/s/iu1uqwiqu1vh285/semana16_pl%c3%a1tica1_transferenciasyteoria.pdf?dl=0 https://www.dropbox.com/s/3mdkt58saq6viq1/semana16_pl%c3%a1tica1_transferenciasyteoria.mp3?dl=0 https://youtu.be/5ck6gsydi9i https://www.dropbox.com/s/e0q0827a6pd1tk4/semana16_pl%c3%a1tica1_transferenciasyteoria.mp4?dl=0 https://www.dropbox.com/s/zl8s3105zh63btl/lecturas.zip?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/l8b3wm4rw2m67u6/semana16_pl%c3%a1tica2_transferencias.pdf?dl=0 https://www.dropbox.com/s/ifbdfe814boj960/semana16_pl%c3%a1tica2_transferencias.mp3?dl=0 https://youtu.be/tm-v-x5g2zw https://www.dropbox.com/s/k3ybdaflnem59tc/semana16_pl%c3%a1tica2_transferencias.mp4?dl=0 http://www.andersonlab.ccny.cuny.edu/ https://www.dropbox.com/s/aap69pybg9fpeqi/semana16_pl%c3%a1tica3_transferenciasdenuevo.pdf?dl=0 https://www.dropbox.com/s/q6mzi0dnekc3830/semana16_pl%c3%a1tica3_transferenciasdenuevo.mp3?dl=0 https://youtu.be/-n39wonomjk https://www.dropbox.com/s/wc708zpw3pmze47/semana16_pl%c3%a1tica3_transferenciasdenuevo.mp4?dl=0 http://www.ib.unam.mx/directorio/114 https://youtu.be/m0ny_6_32yq https://www.dropbox.com/s/r9kaeieoder4u2z/lecturas2.zip?dl=0 https://www.dropbox.com/s/099h4fb7cilxivg/semana17_pl%c3%a1tica1_mop_otravez.pdf?dl=0 https://www.dropbox.com/s/ma2cu8fiwslg92l/semana17_pl%c3%a1tica1_mop_otravez.mp3?dl=0 https://youtu.be/vspo0jcqllm https://www.dropbox.com/s/t3q197bm80cg3xx/semana17_pl%c3%a1tica1_mop_otravez.mp4?dl=0 https://www.dropbox.com/s/8zlmqi0jct7oa6t/lecturas.zip?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/h1t1wtbqrh1z39s/semana17_pl%c3%a1tica2_mop_pr%c3%a1ctica.pdf?dl=0 https://www.dropbox.com/s/pwo3w0eh497pe2v/semana17_pl%c3%a1tica2_mop_pr%c3%a1ctica.mp3?dl=0 https://youtu.be/dwojuvgoovs https://www.dropbox.com/s/r63wbtf3q3wpm52/semana17_pl%c3%a1tica2_mop_pr%c3%a1ctica.mp4?dl=0 https://luismurao.github.io/ https://youtu.be/d4ksq7h5uyc https://www.dropbox.com/s/lz6cvkmiru83f62/semana18_pl%c3%a1tica1_teoriacomparaciones.pdf?dl=0 https://www.dropbox.com/s/dx1xx5ul6wf263c/semana18_pl%c3%a1tica1_teoriacomparaciones.mp3?dl=0 https://youtu.be/ryhvkx0fvmw https://www.dropbox.com/s/phmh7a505cgawhf/semana18_pl%c3%a1tica1_teoriacomparaciones.mp4?dl=0 https://www.dropbox.com/s/24mvb17tw4mii8u/lecturas.zip?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/ltw993gffov2l3z/semana18_pl%c3%a1tica2_visualizaci%c3%b3n.pdf?dl=0 https://www.dropbox.com/s/nqzn415s495dzip/semana18_pl%c3%a1tica2_visualizaci%c3%b3n.mp3?dl=0 https://youtu.be/r6psuxvgvgw https://www.dropbox.com/s/q1donvx9uvlpql2/semana18_pl%c3%a1tica2_visualizaci%c3%b3n.mp4?dl=0 http://ecoguate2003.wixsite.com/escobar https://www.dropbox.com/s/txw5mljd0dk3gr6/semana18_pl%c3%a1tica3_comparaci%c3%b3n.pdf?dl=0 https://www.dropbox.com/s/e4prfsohd5scbu7/semana18_pl%c3%a1tica3_comparaci%c3%b3n.mp3?dl=0 https://youtu.be/o7mz83lohjq https://www.dropbox.com/s/3aiwwcfqkpm4t6p/semana18_pl%c3%a1tica3_comparaci%c3%b3n.mp4?dl=0 http://ecoguate2003.wixsite.com/escobar https://youtu.be/hocob72ikc0 https://www.dropbox.com/s/y1qe39y9ekmpcfm/lecturas2.zip?dl=0 biodiversity informatics, 14, 2019, pp. 1-7 5 curso, los videos fueron vistos en youtube un aproximado de 64,000 veces, con un tiempo total de más de 12,500 horas de reproducción. este formato también permitió que los materiales del curso estuviesen disponibles para descarga directa como archivos de video (.mp4) y audio (.mp3) acompañados de las presentaciones utilizadas (.pdf; curso modelado de nicho ecológico 2018). anticipamos que debido a la diversidad de espectadores del curso, diferentes tipos de acceso iban a ser considerados como óptimos, por lo cual usamos diversos recursos digitales. en una encuesta de recapitulación del curso, con 354 participantes, 50.0% reportaron usar videos de youtube como su modo de acceso, 19.2% descargaron archivos .mp4 y 2.2% usaron archivos .mp3. el 28.5% de los participantes descargó las diapositivas en formato .pdf como auxiliar a las presentaciones en archivo de audio o video. todos los estudiantes tuvieron la oportunidad de enviar preguntas a los instructores a través de un formulario de google, los instructores pudieron responder una porción de estas preguntas— las más comunes—al final de cada semana. a pesar de no ser respondidas en su totalidad, las preguntas fueron una manera útil de verificar la participación (i.e., participantes que enviaron preguntas más de 11 semanas) y también para averiguar lo entendido o parcialmente asimilado por la audiencia en las presentaciones (fig. 1). al final del curso contamos aproximadamente 6,200 preguntas; un patrón consistente en la inquietud de los participantes del curso fue la selección de las variables ambientales necesarias para calibrar los modelos de nicho. las preguntas frecuentes fueron: cuáles variables usar? y cómo seleccionarlas? ¿cuántas usar? además, un tema de interés común fue la evaluación de la transferencia de modelos hacia otras regiones o períodos, que son aplicaciones de estas herramientas para evaluar zonas de riesgo para especies invasoras y potenciales efectos del cambio climático, respectivamente. otro tema particularmente dominante en las preguntas de los participantes, fue sobre la selección del algoritmo ideal para utilizar en los distintos estudios (fig. 1). hasta donde sabemos, este material curricular representa el conjunto más exhaustivo de recursos de entrenamiento para el campo disponible en cualquier idioma. el grupo de instructores representó diversas aplicaciones introducción a aplicaciones -mp3 yt mp4 a. townsend peterson estructura interna de nichos pdf mp3 yt mp4 jorge soberón descubrir poblaciones pdf mp3 yt mp4 octavio rojas soto preguntas y respuestas yt enlace todos cambio de clima i pdf mp3 yt mp4 angela cuervo robayo cambio de clima ii pdf mp3 yt mp4 enrique martínez-meyer preguntas y respuestas yt todos modelado de abundancias pdf mp3 yt mp4 enlace carlos yañez arenas planeación para conservación pdf mp3 yt mp4 enlace javier nori preguntas y respuestas yt todos invasiones de especies pdf mp3 yt mp4 andres lira noriega mapeo de enfermedades pdf mp3 yt mp4 enlace daniel romero-alvarez preguntas y respuestas yt enlace todos conclusiones mesa redonda con (casi) todos los instructores yt enlace todos figura 1. palabras más frecuentes encontradas en las preguntas de los participantes del curso de modelado de nicho ecológico. http://biodiversity-informatics-training.org/bi-curriculum/curso-modelado-de-nicho-ecologico-2018/ http://biodiversity-informatics-training.org/bi-curriculum/curso-modelado-de-nicho-ecologico-2018/ https://www.google.com/intl/es_us/forms/about/ https://www.dropbox.com/s/ar7r7tvc3vvffl4/semana19_pl%c3%a1tica1_comentariogeneral.mp3?dl=0 https://youtu.be/f_nkqbqjmjc https://www.dropbox.com/s/ftqzwia94k16k01/semana19_pl%c3%a1tica1_comentariogeneral.mp4?dl=0 https://www.researchgate.net/profile/andrew_peterson10 https://www.dropbox.com/s/z2cykt4ecs08h32/semana19_pl%c3%a1tica2_caracterizarnichos.pdf?dl=0 https://www.dropbox.com/s/zr6o9dn13tb0a4w/semana19_pl%c3%a1tica2_caracterizarnichos.mp3?dl=0 https://youtu.be/iae33ivzrwk https://www.dropbox.com/s/bfpue168omtg3ld/semana19_pl%c3%a1tica2_caracterizarnichos.mp4?dl=0 https://biodiversity.ku.edu/biodiversity-modeling https://www.dropbox.com/s/ehe9y2iyv0zmr7s/semana19_pl%c3%a1tica3_descubrirpoblaciones.pdf?dl=0 https://www.dropbox.com/s/s4yjuh30yvctjcb/semana19_pl%c3%a1tica3_descubrirpoblaciones_final.mp3?dl=0 https://youtu.be/i77dg8-hli8 https://www.dropbox.com/s/ff5kli1fulnachf/semana19_pl%c3%a1tica3_descubrirpoblaciones%20copy.mp4?dl=0 http://www.inecol.edu.mx/personal/index.php/biologia-evolutiva/4-octavio-rojas-soto https://youtu.be/eedvs69xo3q https://www.dropbox.com/s/wt475exzin0006g/lecturas.zip?dl=0 https://www.dropbox.com/s/hehirt6wj708ytl/semana20_pl%c3%a1tica1_cambiodeclima.pdf?dl=0 https://www.dropbox.com/s/fcfinu81btoerjk/semana20_pl%c3%a1tica1_cambiodeclima.mp3?dl=0 https://youtu.be/s8c9qz5c53s https://www.dropbox.com/s/pb7kthtzq11mk4e/semana20_pl%c3%a1tica1_cambiodeclima.mp4?dl=0 https://www.researchgate.net/profile/angela_cuervo-robayo https://www.dropbox.com/s/pao9srpy3l9utva/semana20_pl%c3%a1tica2_cambiodeclima_casodeestudio.pdf?dl=0 https://www.dropbox.com/s/ikucyc345ijgb4s/semana20_pl%c3%a1tica2_cambiodeclima_casodeestudio.mp3?dl=0 https://youtu.be/bivp5p4s7bo https://www.dropbox.com/s/6hd3lpihj4q3adj/semana20_pl%c3%a1tica2_cambiodeclima_casodeestudio.mp4?dl=0 http://www.ib.unam.mx/directorio/114 https://youtu.be/gjzvhord0dq https://www.dropbox.com/s/u9thyuvjpfp1jfp/semana21_pl%c3%a1tica1_abundancias_nuevo.pdf?dl=0 https://www.dropbox.com/s/k12w1fsjfn5mmgu/semana21_pl%c3%a1tica1_abundancias_nuevo.mp3?dl=0 https://youtu.be/hoic0r-4r1i https://www.dropbox.com/s/tzlvz1vkxpmyffz/semana21_pl%c3%a1tica1_abundancias_nuevo_final.mp4?dl=0 https://www.dropbox.com/s/qugfzj42lnqf668/abundancias.zip?dl=0 https://sites.google.com/site/cyanezarenas/ https://www.dropbox.com/s/8udzs0ex9fbsbur/semana21_pl%c3%a1tica2_conservaci%c3%b3n.pdf?dl=0 https://www.dropbox.com/s/vcsxda3zncpmvwq/semana21_pl%c3%a1tica2_conservaci%c3%b3n.mp3?dl=0 https://youtu.be/mug_nl--5fg https://www.dropbox.com/s/zm7alpvm7sr5mmb/semana21_pl%c3%a1tica2_conservaci%c3%b3n.mp4?dl=0 https://www.dropbox.com/s/69v5sjfyie4p7ie/literatura_conservacion.zip?dl=0 https://www.researchgate.net/profile/javier_nori https://youtu.be/gy81zt0jflg https://www.dropbox.com/s/rpj7y3b22wtlu2a/semana22_pl%c3%a1tica1_invasionesdeespecies.pdf?dl=0 https://www.dropbox.com/s/dv6fus0w44z858m/semana22_pl%c3%a1tica1_invasionesdeespecies.mp3?dl=0 https://youtu.be/pldypbyobjm https://www.dropbox.com/s/kv2koy1wqkx6zv0/semana22_pl%c3%a1tica1_invasionesdeespecies_final.mp4?dl=0 http://www.inecol.edu.mx/personal/index.php/moleculares/157-andres-lira-noriega https://www.dropbox.com/s/09kjruhub90t70r/semana22_pl%c3%a1tica2_enfermedades.pdf?dl=0 https://www.dropbox.com/s/ep8cge6egw1n7p8/semana22_pl%c3%a1tica2_enfermedades_newfinal.mp3?dl=0 https://youtu.be/puttqhj3dvq https://www.dropbox.com/s/l7xwgla1ifdtwns/semana22_pl%c3%a1tica2_enfermedades_newfinal2.mp4?dl=0 https://www.dropbox.com/s/zuluml07cwosdso/waller2007.pdf?dl=0 https://www.romerostories.com/research https://youtu.be/ayh35odeg-4 https://www.dropbox.com/s/gh8qluj8er4naps/petal_eje_2016.pdf?dl=0 https://youtu.be/ayh35odeg-4 https://www.dropbox.com/s/gh8qluj8er4naps/petal_eje_2016.pdf?dl=0 biodiversity informatics, 14, 2019, pp. 1-7 6 áreas del conocimiento de la ecología distribucional, permitiendo abarcar una diversidad considerable de temas y herramientas. consideramos que el curso fue efectivo, con base en la diversidad de países representados por instructores y participantes, el alto número de participantes y su compromiso al atender a la mayoría de las sesiones semanales. así, proponemos que esta estructura de educación puede ser empleada por otras disciplinas para facilitar la educación de acceso libre y a distancia, facilitando la trasferencia de tecnología e información hacia (y entre) países en desarrollo. por su forma modular, este curso se puede actualizar y evolucionar, según los intereses y deseos de los miembros de la comunidad. los videos de las presentaciones son accesibles en internet de forma gratuita, y pueden ser actualizados y mejorados a través del tiempo; además, esperamos la participación de la comunidad global para la traducción de los materiales generados en este curso a otros idiomas. agradecimientos agradecemos al resto de los miembros del grupo de modelado de nicho ecológico de la universidad de kansas por su aporte y asistencia. las presentes contribuciones representan difusión científica por parte de varias instituciones y agencias de financiamiento (incluyendo u.s. nsf dbi-1661510 para r.p. anderson y virginia tech startup funds para l.e. escobar). literature cited anderson, r. p. 2012. harnessing the world’s biodiversity data: promise and peril in ecological niche modeling of species distributions. annals of the new york academy of sciences 1260:66-80. anderson, r. p. 2015. el modelado de nichos y distribuciones: no es simplemente “clic, clic, clic”. biogeografía 8:4-27. andrewartha, h. g., and l. c. birch. 1964. the distribution and abundance of animals. university of chicago press, chicago. araújo, m. b., and a. guisan. 2006. five (or so) challenges for species distribution modelling. journal of biogeography 33:1677-1688. elith, j., c. h. graham, r. p. anderson, m. dudik, s. ferrier, a. guisan, r. j. hijmans, f. huettmann, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. m. overton, a. t. peterson, s. j. phillips, k. richardson, r. scachetti-pereira, r. e. schapire, j. soberon, s. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species’ distributions from occurrence data. ecography 29:129-151. franklin, j. 2010. mapping species distributions: spatial inference and prediction. cambridge university press, cambridge. guisan, a., w. thuiller, and n. e. zimmermann. 2017. habitat suitability and distribution models: with applications in r. cambridge university press, cambridge. guisan, a., and n. zimmermann. 2000. predictive habitat distribution models in ecology. ecological modelling 135:147-186. jiménez-valverde, a., j. m. lobo, and j. hortal. 2008. not as good as they seem: the importance of concepts in species distribution modelling. diversity and distributions 14:885-890. mateo, r. g., á. m. felicísimo, and j. muñoz. 2011. modelos de distribución de especies: una revisión sintética. revista chilena de historia natural 84:217-240. nix, h. a. 1986. a biogeographic analysis of australian elapid snakes. pp. 4-15 in r. longmore, ed. atlas of elapid snakes of australia. australian government publishing service, canberra. peterson, a. t. 2014. mapping disease transmission risk. johns hopkins university press, baltimore. peterson, a. t., and k. ingenloff. 2015. biodiversity informatics training curriculum, version 1.2. biodiversity informatics 10:65-74. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. peterson, a. t., j. soberón, and v. sánchez-cordero. 1999. conservatism of ecological niches in evolutionary time. science 285:1265-1267. pulliam, h. r. 2000. on the relationship between niche and distribution. ecology letters 3:349-361. qiao, h., x. feng, l. e. escobar, a. t. peterson, j. soberón, g. zhu, and m. papeş. 2018. an evaluation of transferability of ecological niche models. ecography 42:521-534. radosavljevic, a., and r. p. anderson. 2014. making better maxent models of species distributions: complexity, overfitting and evaluation. journal of biogeography 41:629-643. soberón, j., l. osorio-olvera, and a. t. peterson. 2017. diferencias conceptuales entre modelación de nichos y modelación de áreas de distribución. revista mexicana de biodiversidad 88:437-441. biodiversity informatics, 14, 2019, pp. 1-7 7 soberón, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodiversity informatics 2:1-10. soto, m., m. j. angulo, o. l. garduño, and m. hernández. 1984. bioclimatología y computación interactiva. ciencia y desarrollo 59:153-161. udvardy, m. d. f. 1969. dynamic zoogeography. van nostrand reinhold company, new york. varela, s., r. g. mateo, r. garcía-valdés, and f. fernández-gonzález. 2014. macroecología y ecoinformática: sesgos, errores y predicciones en el modelado de distribuciones. revista ecosistemas 23:46-53. zhu, g.-p., and a. t. peterson. 2017. do consensus models outperform individual models? transferability evaluations of diverse modeling approaches for an invasive moth. biological invasions 19:2519-2532. biodiversity informatics, 15, 2020, pp. 1-10 1 co-occurrence networks do not support identification of biotic interactions a. townsend peterson1, jorge soberón1, janine m. ramsey2 and luis osorio-olvera1,3 1biodiverisity institute, university of kansas, lawrence, kansas 66045 usa; 2instituto nacional de salud pública, tapachula, chiapas, méxico; 3centro de cambio global y sustentabilidad, concat, villahermosa, tabasco, méxico abstract. we assess a body of work that has attempted to use co-occurrence networks to infer the existence and type of biotic interactions between species. although we see considerable promise in the approach as an exploratory tool for understanding patterns of co-occurrence of species, we note and describe numerous problems in the step of inferring biotic interactions from the co-occurrence patterns. these problems are both theoretical and empirical in nature, and limit confidence in inferences about interactions rather severely. we examine a series of examples that demonstrates striking discords between interactions inferred from co-occurrence patterns and previous experimental results and known life-history details. introduction a series of contributions over the past decade has explored the application of networks of cooccurrence to identifying interactions among species (sánchez‐cordero et al. 2008; stephens et al. 2009; gonzález-salazar and stephens 2012; ibarra-cerdeña et al. 2017; stephens et al. 2019). that is, the authors purport to be able to use networks of co-occurrence or non-co-occurrence, based on spatial information in primary biodiversity databases, to infer biotic interactions (sánchez‐cordero et al. 2008). were such an inferential ability to prove feasible, it would represent an exciting new dimension in biodiversity informatics—for instance, with the emergence and existence of large-scale biodiversity information resources (canhos et al. 2004; stein and wieczorek 2004), now beyond 109 records easily and readily available for analysis, species’ interactions could be assessed and inferred on global scales, adding an important new dimension to what have been termed “essential biodiversity variables” (pereira et al. 2013). this paper, however, examines critically the basic proposition that spatial co-occurrence can be translated into a hypothesis and prediction of biotic interactions. specifically, we assess (1) whether conceptually such a connection (co-occurrence biotic interactions) should be expected to exist, and (2) the degree to which practical considerations (e.g., sampling biases) may further cloud and confuse any such relationships. finally, (3) we examine a series of examples of such analyses, and show that in many cases, by applying the proposed methodology, conclusions are reached that are even opposite to those of both common sense and actual experimental data. conceptual bases should a co-occurrence biotic interaction connection exist? imagining a (wonderful) situation in which the occurrence data used for inferences are comprehensive, complete, and unbiased, one could then estimate patterns of co-occurrence with some confidence. however, whether spatial co-incidence can be a proxy for biotic interactions is a theme that has been debated for decades—the co-occurrence work makes brief reference to these debates (gonzález-salazar and stephens 2012; stephens et al. 2016; stephens et al. 2017; stephens et al. 2019), but generally fails to cite, mention, or assume their crucial elements (e.g., connor and simberloff 1979; gilpin and diamond 1982; hubbell 2001; peres‐neto et al. 2001; ulrich 2004; ulrich et al. 2017). quite generally, spatial co-incidence of species’ distributions may be a consequence of geographic constraint, history, shared climate or substrate preferences, migratory patterns, or many other factors (morueta‐holme et al. 2016; freilich et al. 2018). in tandem with co-occurrence patterns, since non-co-occurrence may derive similarly from causes not related to biotic interactions, patterns of co-incidence and non-co-incidence are no indication of the processes causing them (bell 2005). as such, the network patterns that are the focus of the inferences from the co-occurrence network analyses are simply numerical representations of patterns of co-occurrence in a particular database. in a. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 2 sampling, nc invariably underestimates the actual numbers of counts that would be obtained from a systematic documentation of the range of species c, so nc/n will be, in the vast majority of cases, an underestimate of the true probability. the same problem occurs when estimating co-incidences of pairs or sets of species, which has been termed the “unseen shared species effect”—this problem is known to be serious, and detailed methods have been proposed to fix it (colwell and coddington 1994; chao et al. 2000; chao et al. 2005, 2006). however, these more robust estimation methods, which are clearly more appropriate than the simple, uncorrected nc/n, are not used in the co-occurrence network methodology. details the co-occurrence network work to date has used a mathematical notation that is extremely cumbersome to describe the methodology (stephens et al. 2017; stephens et al. 2019). for instance, xa represents a cell, but since only the sub-index matters… why use xa at all? we have also noted changes of meaning—for instance, in stephens et al. (2017), xa is used to denote a cell in one place, but elsewhere to denote a variable. the co-occurrence network methodology also falls into the trap of unnecessary introduction of symbols—in stephens et al. (2009), the symbol for variable i in cell xi is bi(xi), but later these same variables are denoted c and xi, which is confusing. similarly, the definition of the epsilon index is presented in two different symbologies in the same paper (stephens et al. 2009). much more appropriate would be to stick to the simpler and transparent symbols that are used in the software species. that is, nj should be used to denote the number of incidences of species j. then, nj,k can be used to denote the number of co-occurrences of species j and k; n is used to denote the total number of cells in the grid. everything in the methodology (including the cardinality of complement sets) can be calculated from these three quantities. for instance, epsilon is: the above formula immediately raises questions. first, it is asymmetric: that is, , and particular, their species tool (stephens et al. 2019), which calculates indices of co-occurrence of species in the form of a statistic epsilon, is certainly useful as an exploratory tool focused on patterns of co-occurrence. however, going beyond the pattern to make inferences of process, is immediately suspect. the co-occurrence network group ostensibly tested the predictive capacity of their method empirically by correlating their interaction coefficient epsilon values with numbers of positive tests for the presence of a parasite in independent samples (gonzález-salazar and stephens 2012). however, given almost 30 years of empirical and theoretical literature casting doubts on the robustness of such results, this claim cannot be accepted at face value and based on a single test, and rather needs to be examined critically, using a properly constructed battery of tests (gotelli and graves 1996; gotelli 2000). practicalities what other factors become important? the co-occurrence network approach is based on databases of primary biodiversity occurrences, which is attractive in that it is primary (i.e., based on data deriving directly from individual occurrence records of each species), rather than secondary information (i.e., deriving from some interpretation or synthesis of primary data). however, such data are well-known to be massively influenced by biases related to sampling, in terms of the diverse logistic, practical, historical, and political factors that structure how biologists have been able to sample biodiversity on earth, and report those data to the broader scientific community. these biases have been documented thoroughly in general (yesson et al. 2007; beck et al. 2013; gaiji et al. 2013; otegui et al. 2013a; otegui et al. 2013b; beck et al. 2014; idohou et al. 2015; anderson et al. 2016; asase and peterson 2016; peterson and soberón 2018), and specifically for mexico (bojórquez-tapia et al. 1995; peterson et al. 1998; soberón et al. 2000; soberón et al. 2007). to claim to estimate co-occurrence rates (let alone biotic interactions) robustly on the basis of such incomplete, sparse, uneven, and biased sampling is rather doubtful. for instance, in various descriptions of the co-occurrence network method (e.g., stephens et al. 2019), the “probability” of species c being present in a random cell is stated as nc/n, where nc is the number of counts of presences of c, in a total of n cells. however, given incomplete, uneven, and biased ,( / / ) ( | ) ( , ) (1 / ) / k i k k i i k k i i n n n n n b i i k n n n n n ε ε − = = − a. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 3 exploration of an example with mexican felids (see below) suggests that the asymmetry can be by as much as 20%. what is the meaning of these two values for the co-occurrence of the two species? second, epsilon is dependent on the size of the grid, n, even though the magnitude and direction of the interaction of the two species should not depend on such contextual information. third, in this formula, it is possible to have divisions by zero, which would produce undefined epsilon values. this undesirable situation will happen when ni = n, which is relatively likely for coarse-grained grids. what is the meaning of these undefined values? fourth, the authors claim that associations are significant when associated with epsilon values >1.96 (i.e., equivalent to two standard deviations of a normal distribution). still, epsilon(bi|lik) is most certainly not normally distributed, as any exploration of the species software will show, and therefore the use of a symmetric, normal distribution is not appropriate. finally, the gaps and biases in the sistema nacional de información sobre biodiversidad (snib) make the entire probabilistic argument behind the epsilon index very doubtful. the numbers in that formula, in general, cannot be regarded as probabilities, but as proportions of observations, for a given species, or proportions of observed co-occurrences, for pairs of species, in a particular database. since the quantities nxi, p(c | xi), and p(c) would mostly be underestimated relative to some hypothetical “true” value, often extremely so, the index epsilon will present serious problems. for example, imagine a grid of 32 km grid cells, it will have 6736 cells covering mexico. consider some species j with a true incidence of 60%, and some species k with a true incidence of 40%. if the true co-incidence of the two species is 20%, epsilon (j, k) is 3.79. however, if species j is sampled in only 10% of its localities, the epsilon value shifts to -0.75! worked examples the co-occurrence network-based methodology has been implemented elegantly on a platform called species, which has been connected to the sistema nacional para información de la biodiversidad (snib), maintained by the comisión nacional para el uso y conocimiento de la biodiversidad (conabio), of the mexican government (see the “species” site1). this facility offers the opportu1http://species.conabio.gob.mx nity to explore the methodology in relation to real data, and in diverse contexts. the outcome, however, is rather damning for the conclusions that one might wish to make from the co-occurrence network analyses. for example, we explored the “interactions” between two taxa—the family trogonidae (aves) and the family scarabeidae (insecta) across mexico, and found a complex set of attractions and interdependencies (figure 1). that is, we noted some scarab species that were tightly and significantly associated with particular trogon or quetzal species, and others that were not closely associated at all. the interesting feature, however, is that trogons are arboreal and frugivorous, in largest part, and have tiny feet that would not permit any terrestrial activity (forshaw 2009), whereas scarabs are terrestrial. we see no direct or indirect scenario that would lead to what could be termed biotic interactions between these two taxa, yet the epsilon index misinterprets distributional coincidence or distributional non-coincidence as positive or negative interactions. to provide a more concrete calibration of the co-occurrence network interaction coefficients, we explored situations in which actual field experiments have been conducted. that is, we explored the co-occurrence of the rodents dipodomys merriami and perognathus longimembris (based on data derived from the global biodiversity information facility; https://www.gbif.org/occurrence/download/0000059-190415153152247), which yielded a substantial positive epsilon value of 15.37, which is highly statistically significant using the species methodology. a positive epsilon should reflect positive interactions (i.e., mutualism, symbiosis), yet detailed field experiments (lemen and freeman 1983; lemen and freeman 1986) indicate that these species rather have a strong negative interaction, in which dipodomys depresses populations of perognathus dramatically. similarly, we compared three dipodomys spp. against a suite of six other rodent species (based on data derived from the global biodiversity information facility2) all of the pairwise epsilon values were >4.39, and all were statistically significant, yet heske et al. (1994) documented diverse, strong negative interactions among these same species. in a further example, we took the six cat species (family felidae) occurring in mexico, and calculated 2https://www.gbif.org/occurrence/download/0000059-190415153152247) http://species.conabio.gob.mx a. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 4 epsilon values for each pair; however, we used two distinct databases to construct the presence-absence matrices (pams) from which epsilon is calculated: one was the 2018 snib database, and the other was for the same species, but with distributional data derived from the iucn extent-of-occurrence datasets (iucn 2016). both pams have errors, of course, but they also have contrasting properties: the iucn data generally overpredict alpha (i.e., single-site) diversity and under-predict beta (i.e., among-site) diversity, whereas the opposite is true for the snib database (lira-noriega et al. 2007). the outcomes were quite contrasting (figure 2): epsilon values from the iucn data were centered on zero but quite variable, whereas those from the snib data were generally positive but less variable. this result highlights that the epsilon index is database-dependent, such that inferences about true processes will remain doubtful, even if, as commented above, the database were the outcome of perfect and comprehensive sampling. discussion importance in ecology finding a way by which to infer process from pattern has always been a holy grail in ecology. the specific challenge of inferring biotic interactions figure 1. summary and visualization of a co-occurrence network of trogons and quetzals (trogonidae; blue circles) and scarab beetles (scarabeidae, orange circles). the trogons and quetzals are labeled as to species, and highland species are aligned along the left, whereas the lowland species are aligned along the right. scarab species are not labeled, as they are quite numerous. inset: distribution of epsilon values, with significant values shown in blue. a. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 5 from data documenting only occurrences is complicated by the facts that interactions are scale-dependent, and that co-occurrences are determined by complex interactions among dispersal, physiology, behavior, habitat preferences, and evolutionary history. the bottom line is that epsilon, particularly when used with real-world data with the ever-present biases and gaps, produces highly doubtful results that are certainly not interpretable as indices of real co-incidence, much less of ecological interactions. now, fixes exist for this set of problems and challenges. (1) one can perform a completeness analysis for every grid cell, and use in analyses only those cells that have good completeness indices (sousa‐ baena et al. 2014). this approach will produce a database with a species x site matrix that has fewer rows (i.e., fewer sites), but those sites will have inventories that are more directly comparable. (2) one can use a species x site matrix created not from observations, but from expert data, such as iucn’s extent of occurrence maps (hurlbert and jetz 2007). this approach will probably overestimate incidences and co-incidences (nj, nk, and nj,k), but the error would be much less marked than when using raw occurrence data, which will inevitably underestimate these three quantities. (3) one can develop detailed species distribution models, and use careful interpretations of the outputs of the models to create the species x site matrix (rojas-soto et al. 2003; cooper and soberón 2018). these three approaches allow researchers to deal with the serious problems involved in using a database of simple occurrence data; epsilon values deriving from such analyses will be a truer measure of co-incidence of species j and k. if one desires to interpret patterns of co-occurrence as reflecting biotic interactions, approaches have been explored that would clarify and refine that interpretation. for instance, morueta‐holme et al. (2016) refined simple interpretations of networks by correcting for indirect effects of other species, avoiding spurious associations driven by regional‐ scale distributions, and describing associations in multi‐species contexts—these refinements permitted some degree of direct interpretation of co-occurrence networks in the context of interactions. such refinements, however, appear not to have been contemplated—much less implemented—by the co-occurrence network group. in sum, we perceive a set of points that would improve the ideas and methods of those who would wish to interpret co-occurrence networks as interactions. they should (1) refer to their epsilon values as simple measures of association, rather than making conclusions about interactions and other biological processes that may or may not be producing those associations. (2) they should make clear that the data in the snib are simply examples, and make explicit that unless the database is a “true,” complete, and comprehensive presence-absence database, their epsilon calculations are not appropriate measures of co-incidence, much less indicators of ecological interactions. finally (3), those using co-occurrence network properties should describe for what type of database and in what sense epsilon can be calculated and how this index can and should be interpreted. this latter point clearly affects much of the substance of the work, such that no doubt exists that those using these methods are seriously overinterpreting the meaning of the epsilon statistic. importance in epidemiology and public health the arena in which co-occurrence networks have been used most intensively is that of detecting species interactions relevant to transmission of pathogens from reservoir or host species to humans via insect vectors (stephens et al. 2009). the interaction networks of interest in this area are interactions between pathogens and vectors, pathogens and hosts, and figure 2. boxplots of epsilon values obtained for pairwise comparisons among the six species of felidae in mexico, from the sistema nacional para información de la biodiversidad (snib) database, and from the international union for the conservation of nature (iucn) database. both datasets were assessed at a spatial resolution of ½º, or about 55 km. a. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 6 vectors and hosts. in each case, knowledge of these biotic associations is too-often incomplete, fragmentary, and/or incorrect (woolhouse and gowtage-sequeria 2005; peterson et al. 2007; estrada-peña et al. 2015). in this sense, the potential inferences that these researchers explore are exciting, but only if they hold robust and well-founded promise of anticipating real associations. the co-occurrence network group has focused most on transmission of leishmania spp. parasites from vertebrates via sandflies of the genus lutzomyia to humans (e.g., gonzález-salazar and stephens 2012; stephens et al. 2016). the key assertion is that geographic co-occurrence implies biotic interactions such as vectoring and hosting pathogens—this proposition can be examined empirically for this same system. for instance, pech‐may et al. (2010), in detailed sampling at two sites in the calakmul region of southeastern mexico, found 7 species of lutzomyia species present, with leishmania infections common; however, those infections were common in only three species (lu. shannoni, lu. cruciata, lu. ylephiletor), and rare or absent in the remaining species. similar results obtained in a later study in the same region, with several species co-occurring with leishmania infections, but apparently not interacting with the pathogen (pech‐may et al. 2016). similarly, sánchez-garcía et al. (2010) collected large samples of 10 sandfly species in the chetumal region of southeastern mexico, but found only three of them to carry leishmania infections. these carefully designed studies demonstrate the differential vector capacity of phlebotomine species, despite co-occurrence, and the rash nature of conclusions based simply on distributional coincidence. in terms of leishmania-mammal interactions, rodríguez-rojas et al. (2017) sampled rodents and sandflies in northeastern mexico, and tested them for infection with leishmania. four of 10 rodent species were infected with le. mexicana at an overall infection rate of ~10%, yet 9 sandfly species—including two species (lu. cruciata, lu. shannoni) for which sampling was numerous enough that leishmania should have been detected—revealed no positive samples. this comment is not to assert that such infections do not exist, but rather to emphasize the tenuous, assumption-laden, and uncertain nature of the inferential linkage between co-occurrence in space and participation in biotic interactions, and more importantly, linkages to human disease risk. to consider a very different class of diseases, arboviruses are of increasing concern in many regions in light of the important disease burden that they create, particularly where they are invasive (gubler 1998). vectorial capacity and competence are key factors driving pathogen transmission. for instance, dengue virus is now well-established in the americas via the invasive vector species aedes aegypti and a. albopictus, and overlaps broadly in terms of geographic distribution with large numbers of mammal species, which those using co-occurrence networks would interpret as positive interactions (e.g., host-parasite interactions). nonetheless, dengue has been identified only tentatively in several bat species in mexico, which likely reflects incidental infection, since an elegant recent study demonstrates that those species are not competent reservoirs, and neither co-incidence, viremia or duration of the latter are sufficient for vector infection (vicente-santos et al. 2017)—bats are therefore dead-end hosts for the virus, despite broad co-occurrence. a recent review of zika virus (zikv) etiology (gutiérrez-bugallo et al. 2019) sets out current information regarding infection and competency of vector and host associations, focusing on regions where the virus circulates. despite known zikv infections in numerous species of bats in africa and asia, they have not been found to be competent reservoirs, contra the predictions of gonzález-salazar et al. (2017). experimentally infected rodents, particularly mice, are not susceptible to zikv infection and development of sustained viremia, thanks at least in part to differences in viral detection by the rodents’ innate immune system (ding et al. 2018). indeed, even among invertebrates found infected with zikv, only a handful are competent vectors (garcia-rejon et al. 2010). these results contrast and highlight the incomplete and weak inferences that derive from assumptions that co-occurrence implies biotic interactions that may impact human disease risk. general conclusions a key point in this debate, quite clearly, is why species co-occur and why do other species not co-occur? the naive application of the competitive exclusion principle (gause 1934) would suggest that two species with identical ecological niches will not be able to coexist. this idea immediately causes concern for the co-occurrence network methodology, which uses co-occurrence to infer interactions: coma. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 7 petitors would never be found co-occurring. however, the complexities that enter into this debate are quite daunting. that is, many species pairs indeed interact strongly in a negative sense, but are quite able to exist at different times, or on spatial scales finer than those that are considered by the co-occurrence network methodology. that is, a substantial body of ecological theory treats how the competitive exclusion principle is necessarily modified by the existence of spatial heterogeneity in a landscape (amarasekare 2003), or by the possibility of temporal partitioning to avoid strong, negative interactions between species (chesson and warner 1981). indeed, the classic book, geographical ecology (macarthur 1972) has entire chapters treating mechanisms of co-existence of species pairs that might otherwise compete. a related question is that of why species do not co-occur. the co-occurrence network methodology interprets non-co-occurrences as evidence of negative interactions between species. however, many other factors may enter the picture. for instance, the species concerned may simply have distinct ecological niches—that is, they may belong to lineages that have adapted to different sets of conditions, and for that reason do not co-occur. they may also have different areas of origin, which is still reflected in distinct distributional areas. the important point is that such species may never have even been close to one another, much less interacted (d’amen et al. 2018). overall, indeed, the species software is quite attractive in that it is fast and user-friendly, constituting an elegant exploratory tool for primary biodiversity occurrence datasets. the epsilon calculations are valid tools, but should be interpreted as indicating co-occurrence in a particular database only, which may have many meanings and interpretations, depending on the particular situation and conditions. we note that more sophisticated approaches to these questions of links between co-occurrence and interactions have been published (e.g., morueta‐holme et al. 2016) that take into account ecological niche differences and other factors (see review in dormann et al. 2018), and that simple, siteor pixel-based spatial coincidence could be refined via consideration of spatial topology or via fuzzy spatial matching (visser and de nijs 2006). as we have discussed and demonstrated above, the further inferential step of interpreting epsilon as summarizing the magnitude and direction of biotic interactions is quite inappropriate. acknowledgments the presence-absence matrices were compiled by raul sierra, a collaborator of stephens, as part of a collaboration to analyze the performance of species when applied to different types of databases. loo thanks conacyt for postdoctoral fellowship support (number 740751, cvu: 368747). this research was supported in part by a grant from the national science foundation (iia-1920946). competing interests the authors have declared that no competing interests exist. literature cited amarasekare, p. 2003. competitive coexistence in spatially structured environments: a synthesis. ecology letters 6:1109-1122. anderson, r. p., m. araújo, a. guisan, j. m. lobo, e. martínez-meyer, a. t. peterson, and j. soberón. 2016. data fitness for use in distribution modelling: are species occurrence data in global online repositories fit for modeling species distributions? the case of the global biodiversity information facility (gbif). global biodiversity information facility, copenhagen. asase, a., and a. t. peterson. 2016. completeness of digital accessible knowledge of the plants of ghana. biodiversity informatics 11:1-11. beck, j., l. ballesteros‐mejia, p. nagel, and i. j. kitching. 2013. online solutions and the ‘wallacean shortfall’: what does gbif contribute to our knowledge of species’ ranges? diversity and distributions 19:10431050. beck, j., m. böller, a. erhardt, and w. schwanghart. 2014. spatial bias in the gbif database and its effect on modeling species’ geographic distributions. ecological informatics 19:10-15. bell, g. 2005. the co‐distribution of species in relation to the neutral theory of community ecology. ecology 86:1757-1770. bojórquez-tapia, l. a., i. azuara, e. ezcurra, and o. a. flores v. 1995. identifying conservation priorities in mexico through geographic information systems and modeling. ecological applications 5:215-231. canhos, v. p., s. de souza, r. de giovanni, and d. a. l. canhos. 2004. global biodiversity informatics: setting the scene for a “new world” of ecological forecasting. biodiversity informatics 1:1-13. a. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 8 chao, a., r. l. chazdon, r. k. colwell, and t. j. shen. 2005. a new statistical approach for assessing similarity of species composition with incidence and abundance data. ecology letters 8:148-159. chao, a., r. l. chazdon, r. k. colwell, and t. j. shen. 2006. abundance‐based similarity indices and their estimation when there are unseen species in samples. biometrics 62:361-371. chao, a., w.-h. hwang, y.-c. chen, and c.-y. kuo. 2000. estimating the number of shared species in two communities. statistica sinica 10:227-246. chesson, p. l., and r. r. warner. 1981. environmental variability promotes coexistence in lottery competitive systems. american naturalist 117:923-943. colwell, r. k., and j. a. coddington. 1994. estimating terrestrial biodiversity through extrapolation. philosophical transactions of the royal society of london b 335:101-118. connor, e. f., and d. simberloff. 1979. the assembly of species communities: chance or competition? ecology 60:1132-1140. cooper, j. c., and j. soberón. 2018. creating individual accessible area hypotheses improves stacked species distribution model performance. glob. ecol. biogeogr. 27:156-165. d’amen, m., h. k. mod, n. j. gotelli, and a. guisan. 2018. disentangling biotic interactions, environmental filters, and dispersal limitation as drivers of species co‐occurrence. ecography 41:1233-1244. ding, q., j. m. gaska, f. douam, l. wei, d. kim, m. balev, b. heller, and a. ploss. 2018. species-specific disruption of sting-dependent antiviral cellular defenses by the zika virus ns2b3 protease. proceedings of the national academy of sciences usa 115:e6310-e6318. dormann, c. f., m. bobrowski, d. m. dehling, d. j. harris, f. hartig, h. lischke, m. d. moretti, j. pagel, s. pinkert, and m. schleuning. 2018. biotic interactions in species distribution modelling: 10 questions to guide interpretation and avoid false conclusions. global ecology and biogeography 27:1004-1016. estrada-peña, a., j. de la fuente, r. s. ostfeld, and a. cabezas-cruz. 2015. interactions between tick and transmitted pathogens evolved to minimise competition through nested and coherent networks. scientific reports 5:10361. forshaw, j. m. 2009. trogons: a natural history of the trogonidae. princeton university press, princeton. freilich, m. a., e. wieters, b. r. broitman, p. a. marquet, and s. a. navarrete. 2018. species co‐occurrence networks: can they reveal trophic and non‐trophic interactions in ecological communities? ecology 99:690699. gaiji, s., v. chavan, a. h. ariño, j. otegui, d. hobern, r. sood, and e. robles. 2013. content assessment of the primary biodiversity data published through gbif network: status, challenges and potentials. biodiversity informatics 8:94-172. garcia-rejon, j. e., b. j. blitvich, j. a. farfan-ale, m. a. loroño-pino, w. a. c. chim, l. f. flores-flores, e. rosado-paredes, c. baak-baak, j. perez-mutul, and v. suarez-solis. 2010. host-feeding preference of the mosquito, culex quinquefasciatus, in yucatan state, mexico. journal of insect science 10:32. gause, g. f. 1934. experimental analysis of vito volterra’s mathematical theory of the struggle for existence. science 79:16-17. gilpin, m. e., and j. m. diamond. 1982. factors contributing to non-randomness in species co-occurrences on islands. oecologia 52:75-84. gonzález-salazar, c., and c. r. stephens. 2012. constructing ecological networks: a tool to infer risk of transmission and dispersal of leishmaniasis. zoonoses and public health 59:179-193. gonzález-salazar, c., c. r. stephens, and v. sánchez-cordero. 2017. predicting the potential role of non-human hosts in zika virus maintenance. ecohealth 14:171-177. gotelli, n. j. 2000. null model analysis of species co‐ occurrence patterns. ecology 81:2606-2621. gotelli, n. j., and g. r. graves. 1996. null models in ecology. smithsonian institution, washington, d.c. gubler, d. j. 1998. resurgent vector-borne diseases as a global health problem. emerging infectious diseases 4:442. gutiérrez-bugallo, g., l. a. piedra, m. rodriguez, j. a. bisset, r. lourenço-de-oliveira, s. c. weaver, n. vasilakis, and a. vega-rúa. 2019. vector-borne transmission and evolution of zika virus. nature ecology and evolution 3:561-569. heske, e. j., j. h. brown, and s. mistry. 1994. long‐ term experimental study of a chihuahuan desert rodent community: 13 years of competition. ecology 75:438-445. hubbell, s. t. 2001. the unified neutral theory of biodiversity and biogeography. princeton university press, princeton, n.j. hurlbert, a. h., and w. jetz. 2007. species richness, hotspots, and the scale dependence of range maps in a. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 9 ecology and conservation. proc. natl. acad. sci. usa 104:13384-13389. ibarra-cerdeña, c. n., l. valiente-banuet, v. sánchez-cordero, c. r. stephens, and j. m. ramsey. 2017. trypanosoma cruzi reservoir—triatomine vector co-occurrence networks reveal meta-community effects by synanthropic mammals on geographic dispersal. peerj 5:e3152. idohou, r., a. arino, a. assogbadjo, r. g. kakai, and b. sinsin. 2015. diversity of wild palms (arecaceae) in the republic of benin: finding the gaps in the national inventory combining field and digital accessible knowledge. biodiversity informatics 10:45-55. iucn. 2016. spatial data download.3 international union for the conservation of nature, gland. lemen, c., and p. w. freeman. 1983. quantification of competition among coexisting heteromyids in the southwest. southwestern naturalist 28:41-46. lemen, c. a., and p. w. freeman. 1986. interference competition in a heteromyid community in the great basin of nevada, usa. oikos 46:390-396. lira-noriega, a., j. soberón, a. g. navarro-sigüenza, y. nakazawa, and a. t. peterson. 2007. scale-dependency of diversity components estimated from primary biodiversity data and distribution maps. diversity and distributions 13:185-195. macarthur, r. 1972. geographical ecology. princeton university press, princeton, n.j. morueta‐holme, n., b. blonder, b. sandel, b. j. mcgill, r. k. peet, j. e. ott, c. violle, b. j. enquist, p. m. jørgensen, and j. c. svenning. 2016. a network approach for inferring species associations from co‐occurrence data. ecography 39:1139-1150. otegui, j., a. h. ariño, v. chavan, and s. gaiji. 2013a. on the dates of gbif mobilised primary biodiversity records. biodiversity informatics 8:173-184. otegui, j., a. h. ariño, m. a. encinas, and f. pando. 2013b. assessing the primary data hosted by the spanish node of the global biodiversity information facility (gbif). plos one 8:e55144. pech‐may, a., f. j. escobedo‐ortegón, m. berzunza‐cruz, and e. a. rebollar‐téllez. 2010. incrimination of four sandfly species previously unrecognized as vectors of leishmania parasites in mexico. medical and veterinary entomology 24:150-161. pech‐may, a., g. peraza‐herrera, d. a. moo‐llanes, j. escobedo-ortegón, m. berzunza-cruz, i. becker-fauser, a. c. montes de oca-aguilar, and e. rebollar-tellez. 2016. assessing the importance of 3http://www.iucnredlist.org/technical-documents/spatial-data four sandfly species (diptera: psychodidae) as vectors of leishmania mexicana in campeche, mexico. medical and veterinary entomology 30:310-320. pereira, h. m., s. ferrier, m. walters, g. n. geller, r. h. g. jongman, r. j. scholes, m. w. bruford, n. brummitt, s. h. m. butchart, and a. c. cardoso. 2013. essential biodiversity variables. science 339:277-278. peres‐neto, p. r., j. d. olden, and d. a. jackson. 2001. environmentally constrained null models: site suitability as occupancy criterion. oikos 93:110-120. peterson, a. t., a. g. navarro-sigüenza, and h. benítez-díaz. 1998. the need for continued scientific collecting: a geographic analysis of mexican bird specimens. ibis 140:288-294. peterson, a. t., m. papeş, d. s. carroll, h. leirs, and k. m. johnson. 2007. mammal taxa constituting potential coevolved reservoirs of filoviruses. journal of mammalogy 88:1544-1554. peterson, a. t., and j. soberón. 2018. essential biodiversity variables are not global. biodiversity and conservation 27:1277-1288. rodríguez-rojas, j. j., á. rodríguez-moreno, m. berzunza-cruz, g. gutiérrez-granados, i. becker, v. sánchez-cordero, c. r. stephens, i. fernández-salas, and e. a. rebollar-téllez. 2017. ecology of phlebotomine sandflies and putative reservoir hosts of leishmaniasis in a border area in northeastern mexico: implications for the risk of transmission of leishmania mexicana in mexico and the usa. parasite 24:33. rojas-soto, o. r., o. alcantara-ayala, and a. g. navarro. 2003. regionalization of the avifauna of the baja california peninsula, mexico: a parsimony analysis of endemicity and distributional modelling approach. journal of biogeography 30:449-461. sánchez-garcía, l., m. berzunza-cruz, i. becker-fauser, and e. a. rebollar-téllez. 2010. sand flies naturally infected by leishmania (l.) mexicana in the periurban area of chetumal city, quintana roo, méxico. transactions of the royal society of tropical medicine and hygiene 104:406-411. sánchez‐cordero, v., d. stockwell, s. sarkar, h. liu, c. r. stephens, and j. giménez. 2008. competitive interactions between felid species may limit the southern distribution of bobcats lynx rufus. ecography 31:757764. soberón, j., r. jiménez, j. golubov, and p. koleff. 2007. assessing completeness of biodiversity databases at different spatial scales. ecography 30:152-160. soberón, j. m., j. b. llorente, and l. oñate. 2000. the use of specimen-label databases for conservation http://www.iucnredlist.org/technical-documents/spatial-data a. townsend peterson et al. – co-occurrence networks do not support identification of biotic interactions 10 purposes: an example using mexican papilionid and pierid butterflies. biodiversity and conservation 9:1441-1466. sousa‐baena, m. s., l. c. garcia, and a. t. peterson. 2014. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. diversity distrib. 20:369-381. stein, b. r., and j. wieczorek. 2004. mammals of the world: manis as an example of data integration in a distributed network environment. biodiversity informatics 1:14-22. stephens, c., v. sánchez-cordero, and c. gonzález salazar. 2017. bayesian inference of ecological interactions from spatial data. entropy 19:547. stephens, c. r., j. giménez-heau, c. gonzález, c. n. ibarra-cerdeña, v. sánchez-cordero, and c. gonzálezsalazar. 2009. using biotic interaction networks for prediction in biodiversity and emerging diseases. plos one 4:e5725. stephens, c. r., c. gonzález-salazar, v. sánchez-cordero, i. becker, e. rebollar-tellez, á. rodríguez-moreno, m. berzunza-cruz, c. d. balcells, g. gutiérrez-granados, and m. hidalgo-mihart. 2016. can you judge a disease host by the company it keeps? predicting disease hosts and their relative importance: a case study for leishmaniasis. plos neglected tropical diseases 10:e0005004. stephens, c. r., r. sierra‐alcocer, c. gonzález‐salazar, j. m. barrios, j. c. salazar carrillo, e. robredo ezquivelzeta, and e. del callejo canal. 2019. species: a platform for the exploration of ecological data. ecology and evolution 9:1638-1653. ulrich, w. 2004. species co‐occurrences and neutral models: reassessing jm diamond’s assembly rules. oikos 107:603-609. ulrich, w., w. kryszewski, p. sewerniak, r. puchałka, g. strona, and n. j. gotelli. 2017. a comprehensive framework for the study of species co‐occurrences, nestedness and turnover. oikos 126:1607-1616. vicente-santos, a., a. moreira-soto, c. soto-garita, l. g. chaverri, a. chaves, j. f. drexler, j. a. morales, a. alfaro-alarcón, b. rodríguez-herrera, and e. corrales-aguilar. 2017. neotropical bats that co-habit with humans function as dead-end hosts for dengue virus. plos neglected tropical diseases 11:e0005537. visser, h., and t. de nijs. 2006. the map comparison kit. environmental modeling and software 21:346-358. woolhouse, m. e. j., and s. gowtage-sequeria. 2005. host range and emerging and reemerging pathogens. emerging infectious diseases 11:1842. yesson, c., p. w. brewer, t. sutton, n. caithness, j. s. pahwa, m. burgess, w. a. gray, r. j. white, a. c. jones, f. a. bisby, and a. culham. 2007. how global is the global biodiversity information facility? plos one 2:e1124. biodiversity informatics, 15, 2020, pp. 61-66 61 reflections on the challenge of inferring ecological interactions from spatial data robert d. holt department of biology university of florida, gainesville, fl 32611. preamble dr. luis escobar asked me to provide a joint review of the submissions by stephens et al. (2019, this issue) and peterson et al. (2019, this issue) to this debate. i pulled thoughts together, but by the time i sent them, he had received other reviews and made an editorial decision. however, he felt my perspective might nevertheless warrant publishing as a commentary alongside these two pieces. my review was of the original submissions, which are now appearing with minor, mainly cosmetic changes. i have edited the text of my review only lightly, and added a few additional thoughts and pertinent references. neither group of authors has seen my commentary, and so i am responsible for any omissions or lapses in interpretation. introduction the basic bone of contention between the two contributions, by stephens et al. and peterson et al., is whether or not one can make inferences about interactions among species based on spatial patterns of co-occurrence (or not). this is of course a long-standing issue in ecology, going back at least to the ‘assembly rule and null model debate’ that raged in the 1970s and 80s (e.g., diamond 1975, connor and simberloff 1983; see sanderson and pimm 2015, for a review of this debate). in this commentary, i first briefly summarize the central points of the two papers, and then note points on which i disagree with each of them. i note at the outset that i respect all the authors. stephens et al. stephens et al.’s paper has a grandeur about it, as it deals with issues such as the meaning of ‘interactions,’ not just in ecology and biogeography, but more generally across the sciences. they make particular reference to physics and the four fundamental forces of nature, and define an interaction to exist if the spatial positions of two objects of study differ from a null expectation, across an ensemble of observations. they reflect briefly on the relationship between niche concepts and interactions, and touch on the issue of direct vs. indirect interactions. the bulk of the paper is devoted to developing a metric of co-occurrence, called epsilon, with a focus on binary information (presence/absence), and a bayesian framework for making inferences about interactions. they champion the use of a software platform (species), and present intriguing examples. one of these is for the bobcat (lynx rufus) in mexico, in which they argue that including information about other mammals greatly increased the predictive power of a distributional model. they then summarize previous work they have done on zoonoses, where the problem is to identify vectors and hosts associated with a particular pathogen. i read the manuscript with interest, and found their approach and examples intriguing, if not entirely convincing, for reasons laid out below. peterson et al. peterson et al. argue that patterns of co-occurrence shed no light at all on underlying process. they point out that large and poorly understood sampling biases exist in the kind of biodiversity data that stephens et al. use. they provide some specific critiques of the metrics and the notation used by stephens et al. they then examine the example of co-occurrences of trogons and scarab beetles across mexico, using the methodology of stephens et al., and find examples of tight co-occurrence (or non-co-occurrence). still, nothing in the natural history of these taxa suggests either strong mutualisms or competition. they likewise examine two taxa for which independent evidence of interactions exists – desert rodents, and felid cats. in the former case, the epsilon of stephens et al. indicates positive interactions, yet experiments show negative interactions. in the latter case, different results emerge from different databases. finally, peterson et al. walk through some examples from epidemiology, suggesting caution in the inference of interaction from co-occurrence data, as assessed by the stephen et al. epsilon metric. robert d. holt – reflections on the challenge of inferring ecological interactions from spatial data 62 discussion stephens et al. define interactions in terms of co-occurrence, and then try to identify interactions from data on spatial distributions. they use an analogy with physics to understand ecology (in their box 1), and state that fundamental interactions in physics are ‘direct’, whereas other interactions are ‘indirect’. they note that the fundamental data of physics are not just positions, but rather changes in positions over time, e.g., orbits in astronomy. spatiotemporal data—i.e., trajectories—provide much stronger information about presumptive causes of physical phenomena, than just spatial pattern data. the classical celestial mechanics of isaac newton focused on planetary orbits, i.e., spatiotemporal data, not the static patterns of the stars. in newtonian mechanics, if the initial conditions of non-interacting particles happened to have them arranged in some spatially correlated pattern with equal velocities and trajectories pointing in the same direction, with equal forces acting on them, as time goes on, the spatial correlation structure will remain unchanged. in other words, static measures of co-occurrence in a snapshot could reflect the imprint of history and shared responses of the particles to their physical environments, not ongoing interactive processes. this general issue is recognized by the authors, i think, but ends up somewhat lost in the flow of the paper, which focuses on analyses of static spatial patterns, not dynamical spatiotemporal patterns. all of these remarks have analogues in ecology. when available, spatiotemporal data provide much more powerful insights into processes than do static spatial data. moreover, interactions in particle physics do not consist just of changes in spatial position. the weak force for instance can flip the ‘flavor’ of quarks—important in processes such as beta nuclear decay, in which a neutron is converted to a proton. this is not a trivial process—it powers the stars. in other words, interactions in physics change state variables, not just spatial positions. again, this point holds in ecology, as well. stephens et al. include a paragraph touching on population genetics. their use of terms like ‘epistasis’ and ‘linkage’ unfortunately deviates from accepted usage of population genetics. however, leaving this terminological issue aside, it is worth noting that there are protocols in molecular population genetics that do analyze static patterns to infer a process—in particular, natural selection (e.g., the mcdonald-kreitman test for detecting the presence of selection on amino acid sequences, see ch. 6 in charlesworth and charlesworth 2010) the authors also refer to text-mining protocols in linguistics, such as inferences about syntax or semantics from positions of words in sentences. deciphering ancient languages requires interpreting ‘interactions’ among script elements arranged in a linear spatial pattern. these analogies with other disciplines do suggest that useful information about causal processes may be buried in static ecological patterns. however, such inferences rely not just on pattern analysis, but also on prior knowledge about processes (e.g., how living languages work). there are challenges in executing the stephens et al. approach in ecology, beyond those mentioned by peterson et al. unlike physical forces, ecological interactions are often context dependent. for instance, the qualitative sign of an indirect interaction of two prey species via a shared predator depends on (among other factors) whether or not that predator is constrained in its numerical response by higher-order predators. such constraints could prevent the occurrence of apparent competition (holt and bonsall 2017), and turn the indirect interaction from (-,-) to (+,+). quantitatively, even without a change in interaction sign, the impact of species a on species b depends on the abundance of species a, so mere co-occurrence is at best a crude assay of the strength of their interaction. in addition, interactions depend upon abiotic conditions (dunson and travis 1991) and networks of interactions can vary along environmental gradients (pellissier et al. 2018). hence, considerable contingencies likely exist in the strength and even signs of interspecific interactions, implying spatial and temporal variability in interactions among many species. the protocol of stephens et al., however, seems to assume that interactions are fixed (if i understand it correctly). ecologists by and large recognize the inherent difficulty in inferring interactions from descriptive data, and many authors are actively engaged in developing methods to do so, while recognizing the difficulties (e.g., cazellas et al. 2015; sander et al. 2017). stephens et al., however, do not engage with that literature, nor do they attempt to relate their proposed measure or definition of interaction to other metrics of interaction strength that are used widely in the ecological literature (see, e.g., novak et al. 2016). this linkage might be a goal in future developments robert d. holt – reflections on the challenge of inferring ecological interactions from spatial data 63 of their methodology. metrics of interactions should be linked to basic population dynamics—what is the effect of individuals of species a on per capita birth, death, movement, and stage transition rates of individuals of species b? this information might lead to changes in spatial relationships, but not always. a well-mixed chemostat for instance by definition rapidly destroys any spatial structure of the populations it contains, but there can be (for example) strong resource competition among algal species because one competitor reduces the mean-field abundance of a shared resource needed by another. under a giant canopy tree, spatial locations of each individual in a cluster of understory herbs may be driven by microenvironmental germination requirements, but the rate of their individual growth and ultimate seed production may be governed by light competition with that tree. this point matters at larger spatial scales, because seed production and dispersal are required by the herb for colonization of empty sites. nonetheless, the interspecific interaction need not be reflected in local spatial patterns under the dominant light competitor at all; instead, the interaction alters the internal states of the herbs. one issue largely ignored by stephens et al. is the importance of background spatial structure in the environment, and spatial autocorrelation. imagine two species that do not interact but that have distinct responses across a gradient, with one species more prevalent at one end, and the other species at the other. using the metric of stephens et al. one would (i think) conclude there was a negative interaction at play—but this conclusion would be incorrect. there needs to be attention paid to spatial autocorrelation and related issues. bar-massada and belmaker (2017) showed for tree species in the united states that co-occurrence varied across gradients, and conclude that pairwise analyses of co-occurrence (and thus, interactions) are scale-dependent. peterson et al. crisply lay out several problems in the stephens et al. protocol. however, i do think some of their categorical claims need to be qualified. one point about ‘inference’ is that inference is not ‘either/or.’ one can have a tentative inference, a weak inference, a reasonably convincing inference, and even a strong and nearly irrefutable inference. clearly, the latter is best, but that does not mean the former are worthless. so when peterson et al. state “… patterns of co-incidence and non-co-incidence are no indication of the processes causing them” (citing bell 2005), i think that that is too strong a conclusion. in conjunction with other information about a system (e.g., natural history), such patterns can provide some indication that something is going on. scientific inferences should when possible draw on a wide array of evidence, not just single sources, which sometimes can be quite convincing. both sets of authors refer to the celebrated dispute back in the 1970s between jared diamond on the one hand, and individuals like dan simberloff and ed connor, on the other, about inferring competition between species based on distributional patterns. neither paper refers, however, to the recent, incisive book by sanderson and pimm, patterns in nature (2015), which reviews that entire debate, and lays out more sophisticated versions of null models than were used in the past. sanderson and pimm provide a reasonably convincing case that some classic examples of checkerboard distributions of related species, and distributions along gradients, indeed reflect competition. the precise specification and analysis of appropriate null models is crucial—and non-trivial. if one could not use patterns of co-occurrence over time to make tentative inferences about causal processes, most of paleoecology would become an intellectually derelict discipline. wisz et al. (2013) cite many examples in which biotic interactions have large-scale, biogeographic consequences, many of which they drew from the paleontological record. the spread of homo sapiens across the globe, along with our (alas) symbionts such as rats, have had huge consequences for the persistence and geographic ranges of (for instance) large-bodied and highly edible vertebrates, and particularly for flightless birds. cases like this one of course involve not just patterns in spatial data, but spatiotemporal data (e.g., piles of moa bones in prehistoric sites of human habitation in new zealand). one issue missing in this interchange is a thorough consideration of time scales. a tight mutualism or asymmetric facilitation implies that species a cannot be present over even a single generation without species b. this effect should be manifested in ongoing interactions matching current distributions. by contrast, competition and predation can lead to elimination of species co-occurrences—the interactions may all be in the past, not in the present (dubbed “the ghost of competition past” by connell 1980). the point is that current distributions reflect not just current interactions, but past interactions. inferring inrobert d. holt – reflections on the challenge of inferring ecological interactions from spatial data 64 teractions from current distributions requires paying attention to ecological memory, as well as drawing on other avenues of ecological understanding. there are known strong positive associations between species that surely have distributional consequences. for instance, many epiphytic orchids cannot germinate and survive as seedlings without specific mycorrhizal fungal symbionts. there are ecological niche models for orchids in the literature, which successfully predict the distribution of orchid species based on climatic variables (along with bark and other habitat factors). however, it is a leap of faith to conclude that those models portray merely direct ecophysiological responses of the orchid to those specific abiotic factors—they could equally well involve responses of the orchid’s required symbiont. one problem in conservation of endangered orchids is that getting them to germinate and grow can be quite tricky, which apparently reflects problems in getting conditions ‘just right’ for the symbionts, as well as ensuring they are present in the first place. in this sense, a beautifully verified ecological niche model that uses only abiotic variables, might well have an underlying causal dynamic involving strong interspecific interactions. the protocol developed by stephens et al. could potentially provide a valuable tool helping to sort among initial hypotheses about potential mutualist partners in a community, i think (given some prior natural history or trait data). the exercise by peterson et al. relating distributional data on trogons to scarabs is, at best, underwhelming. one always has background hypotheses at play (e.g., the debate between diamond and simberloff focused on biogeographically relevant and phylogenetically related bird taxa, which have similar diets and likely share parasites, and so arguably might compete), but nothing is presented here to warrant this exercise. looking at enough taxa, across enough situations, will surely reveal some ‘significant’ associations. we all know that correlation need not imply causation, but certainly correlations can help to generate hypotheses. with respect to the bobcat, stephens et al. provide a reasonably convincing exposition that their protocol improves understanding the determinants of the distribution of this generalist predator. it would have been instructive for peterson et al. to focus on that case study. peterson et al. then present as a case study two desert rodents, a dipodomys and a perognathus, and show they have positive epsilon values. however, prior experiments by lemen and freeman (1983, 1986) had demonstrated (according to peterson et al.) strong competition. this outcome would seem to be a clear indication that the stephens et al. protocol is grossly misleading. however, i looked at these papers and note that lemen and freeman (1986, p. 390) in fact stated that “interference competition was present but weak,” and cited studies by other authors, at other sites, not showing competition, or with variable results among sites. lemen and freeman also noted that there is a “great deal of overlap in food habitats and habitat preferences” (p. 395). so it is not surprising that at larger spatial scales, there might a positive association in the spatial distributions of these two species. the final set of case studies examined by peterson et al. involve epidemiology. the stephens et al. protocol seems to overpredict host-vector-pathogen interactions. this is an important criticism. however, the epsilon metric might still help refine the pool of possibilities for identifying likely suspects; whether or not such is the case is not clear from this critique. this is the sort of system in which the gold standard of demonstrating unequivocally interspecific interactions—manipulative experiments—are likely either unfeasible or unethical. still, it is important to reach sensible conclusions about what potential suite of vectors should be monitored by public health agencies, and any tool that can help refine our bayesian priors about this matter should be tried repeatedly until it is found to be wanting. a quote from a forthcoming book by ovaskainen and abrego (in press) is apt here, however, with respect to inferring interactions from such exercises: “the results from species association analyses should always be interpreted with caution, and in light of ecological knowledge on the study system.” i want to make a final conceptual point, relevant to both species distribution models (sdms) in general, and the challenge of inferring interactions from static spatial data specifically. as stephens et al. note, many authors use sdms (correlative statistical models) to make inferences about species abiotic niches with no mention of biotic interactions, which in effect sweeps causal reasoning under the rug. consider a thought experiment. a prey species has an intrinsic growth rate r that is a function of abiotic environmental conditions e, formally described by r(e), and these conditions vary across a biogeographic region. the fundamental niche is that set of e for which robert d. holt – reflections on the challenge of inferring ecological interactions from spatial data 65 r(e) > 0. a generalist predator has uniform density p across that region and invariant per capita attack rates a; thus, the predator is not affected by our focal prey (i.e., that prey species suffers “incidental predation,” schmidt 2004). when rare, the prey per capita growth rate is (1/n)dn/dt = r(e) ap. the range of the prey will include those locations for which r(e) > ap. however, locations outside that range include not just sites with negative r(e), but also sites for which 0 < r(e) < ap. a correlative sdm might be constructed that does a perfect job of predicting the range, based solely on abiotic variables. however, it would be incorrect to conclude that only physiological factors cause the range limits, as removing predators in some locations would permit persistence otherwise impossible. conversely, given that p is invariant, no correlation can exist with the presence or absence of the prey species. from a statistical point of view, the predator is uninformative in ‘explaining’ the distribution of this prey species. in other words, one could not reveal this important interaction based purely on co-occurrence data, along the lines of the stephens et al. protocol. in short, my point is that “the absence of a correlation need not imply an absence of causation.” conclusions it can be useful to have provocative papers published, even if flawed, and even if we disagree with them. both of these papers are provocative. all the points i made above about physics carry over to distributional ecology. (i) the distinction between ‘direct’ and ‘indirect’ interactions depends on the fineness of resolution of information about causality. interference competition between terrestrial plants due to allelopathy, for instance, can be represented as a direct interaction (e.g., in a lotka-volterra model), or as an indirect interaction mediated by the concentration of an allelopathic compound. (ii) for inferring interactions, spatiotemporal data are much more insightful than purely spatial data. (iii) patterns of co-occurrence reflect many factors such as initial conditions (viz., history) and responses to external (and possibly unmeasured) environmental factors, and interactions can occur that will not be reflected in co-occurrence data. (iv) interactions should be expressed not just in spatial position, but in terms of all the state variables needed to capture the ‘forces’ driving dynamics of a system. advances in remote sensing are leading to terabytes of data on the physiological states of plants across environments (cavender-bares et al. 2017), and such data should be mined by distributional ecologists addressing range limits, and community ecologists teasing out interactions. it also would be valuable for metrics such as the epsilon of stephens et al. to be evaluated with ‘virtual ecologist’ approaches (see, e.g., zurell et al. 2010, ovoskainen and abrego, in press), in which one creates virtual, spatially explicit communities and ecosystems, and then beats the hell out of them with reference to assessing the utility of proposed metrics or data analytic methods. these two papers contribute to the evolving dialogue about how best to link community ecology and distributional ecology. acknowledgments i thank luis escobar for inviting me to turn a review into a commentary, and the university of florida foundation for support. conflict of interest none references bar-massada, a. and j. belmaker. 2017. non-stationarity in the co-occurrence patterns of species across environmental gradients. j. ecol. 105:391-399. bell, g. 2005. the co-distribution of species in relation to the neutral theory of community ecology. ecology 86:1757-1770. cavender-bares, j., j.a. gannon, s.e. hobbie, m.d. madritch, j.e. meireles, a.k. schweiger, and p. a. townsend. 2017. harnessing plant spectra to integrate the biodiversity sciences across biological and spatial scales. amer. j. botany 104:966-969. cazelles, k., m.b. araujo, n. mouquet, d. gravel. 2015. a theory for species co-occurrence in interaction networks. theor. ecol. 9:39-48. charlesworth, b. and d. charlesworth.2010. elements of evolutionary genetics. roberts and company publishers, greenwood village, colorado. connell, j. h. 1980. diversity and the coevolution of competitors, or the ghost of competition past. oikos 35:131-138. connor, e.f. and d. simberloff. 1983. interspecific competition and species co-ocurrence patterns on islands: null models and the evaluation of evidence. oikos 41:455-465. diamond, j.m. 1975. assembly of species communities. pp. 342-444 in ecology and evolution of communirobert d. holt – reflections on the challenge of inferring ecological interactions from spatial data 66 ties, eds. m.l. cody and j.m. diamond. harvard university press, cambridge. dunson, w.a. and j. travis. 1991. the role of abiotic factors in community organization. amer. nat. 138:10671091. holt, r.d. and m.b. bonsall. 2017. apparent competition. ann. rev. ecol., evol. syst. 48: 447-471. lemen, c. and p.w. freeman. 1983. quantification of competition among coexisting heteromyids in the southwest. southw. nat. 28:41-46. lemen, c.a. and p.w. freeman. 1986. interference competition in a heteromyid community in the great basin of nevada, usa. oikos 46:399-396. novak, m., j.d. yeakel, a.e. noble, d.f. doak, m. emmerson, j.a. estes, u. jacov, m.t. tinker, and j.t. wootton. 2016. characterizing species interactions to understand press perturbations: what is the community matrix? ann. rev. ecol., evol. syst. 47:409-432. ovaskainen, o. and n. abrego. in press. joint species distribution modeling with applications in r. cambridge university press. pellissier, l., albouy, c., bascompte, j., farwig, n., graham, c., loreau, m., maglianesi, m. a., melián, c. j., pitteloud, c., roslin, t., rohr, r., saavedra, s., thuiller, w., woodward, g., zimmermann, n. e. & gravel, d. 2018. comparing species interaction networks along environmental gradients. biol. rev. 93: 785-800. peterson, a.t., j. soberon, j.m. ramsey, and l. osorio. 2019. co-occurrence networks do not support identification of biotic interactions. biodiv. inf. 17:1-10. sander, e.l., j.t. wootton and s. allesina. 2017. ecological network inference from long-term presence-absence data. sci. rep. 7:7154. sanderson, j.g. and s.l. pimm. 2015. patterns in nature: the analysis of species co-occurrences. university of chicago press, chicago. schmidt, k.e. 2004. incidental predation, enemy-free space, and the coexistence of incidental prey. oikos 106:353-343. stephens, c.r., c. gonzalez-salazar, m. vaillalobos, and p.a. marquet. 2019. can ecological interactions be inferred from spatial data? biodiv. inf. 17:11-54. wisz, m.s. et al. 2013. the role of biotic interactions in shaping distributions and realised assemblages of species: implications for species distribution modelling biol. rev. 99:15-33. zurell, d., berger, u., cabral, j.s., jeltsch, f., meynard, c.n., münkemüller, t., nana pagel, n.j., reineking, b., schröder, b., and grimm, v. (2010). the virtual ecologist approach: simulating data and observers. oikos, 119(4), 622-635. microsoft word engler_roedder_final.doc biodiversity informatics, 8, 2012, pp. 30-40   30 disentangling interpolation and extrapolation uncertainties in ecologial niche models: a novel visualization technique for the spatial variation of predictor variable colinearity jan o. engler and dennis rödder zoologisches forschungsmuseum alexander koenig, adenauerallee 160, d 53113 bonn, germany abstract. – environmental niche models (enms) are increasingly used in many scientific fields, with most studies requiring the application of the enm to predict the likelihood of occurrence and/or environmental suitability in locations and time periods outside the range of the data set used to fit the model. uncertainty in the quality of enm predictions caused by errors of interpolation and extrapolation has been acknowledged for a long time, but the explicit consideration of the magnitude of such errors is, as yet, uncommon. among other issues, the spatial variation in the colinearity of the environmental predictor variables used in the development of enms may cause misleading predictions when applying enms to novel locations and time periods. in this paper, we provide a framework for the spatially explicit identification of areas prone to errors caused by changes in the inter-correlation structure (i.e. their colinearity) of environmental predictors used for enm development. the proposed method is compatible with all enm algorithms currently employed, and expands the available toolbox for assessing the uncertainties rising from enm predictions. we provide an implementation of the analysis as a script for the r statistical platform in an online appendix. key words. – climate change, environmental niche, residual distribution, plethodon, purv plots, r statistical platform introduction the development in computational resources during the last few decades, combined with an increasing availability of environmental data and information on species occurrences, has led to a boost in the application of environmental niche models (enms) in many biological fields (kozak et al. 2008, elith and leathwick 2009). in correlative enms, an idealized environmental niche of the target species is derived from the environmental conditions found at the locations where the species is known to occur and/or absence or pseudo-absence data reflecting the general conditions within this area (guisan and zimmermann 2000, peterson et al. 2011). analyzing the environmental conditions at the species’ occurrence records, enms can be used to estimate the target species’ potential distribution based on its realized niche, which is commonly a subset of its fundamental niche (soberón and peterson 2005, peterson et al. 2011). based on the enm, the derived habitat preferences can be subsequently projected into unsampled locations and time periods using a geographic information system (gis). various methods have been proposed for this purpose (guisan and zimmermann 2000, elith and leathwick 2009, franklin 2009, peterson et al. 2011). early enm approaches define the species’ niche as a multidimensional boxcar or convex hull envelope enclosing the environmental characteristics of the known occurrences of the species in environmental space (e.g., bioclim, busby 1991, domain, carpenter et al. 1993). however, more recently developed methods, such as artificial neural networks (olden et al. 2008), multivariate adaptive regression splines (friedman 1991), random forest (breiman 2001), and maximum entropy approaches (phillips et al. 2006), allow for the incorporation of even more complex interactions between predictor variables. these techniques commonly require absence or pseudoabsence data for model training. this extra flexibility has been shown to increase the accuracy of predictions computed with these methods, frequently outperforming more conventional approaches (elith et al. 2006, hernandez et al. 2006, wisz et al. 2008). however, appropriate selection of absence or pseudo-absence data requires special attention here since the choice may strongly influence the reliability and interpretation of the results (saupe et al. 2012). various authors have suggested that the most appropriate background data should reflect the environmental space that is potentially engler and rödder – visualizing predictor colinearity   31 colonizable by the target species (e.g., anderson and raza 2010, barve et al. 2011, saupe et al. 2012). although this restriction of the training area of an enm may result in a more specific model, it is at the same time commonly less generalizable across space and time. applications of enms are manifold, ranging, for example, from simple visualizations of species’ potential distributions (e.g., brown and twomey 2009), assessments of the potential distribution of invasive species (e.g., peterson and vieglais 2001, phillips et al. 2008), quantification of possible impacts of climate change (e.g., araújo et al. 2004, thomas et al. 2004), to analyses of niche evolution and speciation (e.g., kozak and wiens 2006, peterson 2011). many of these applications require some degree of model prediction to novel locations and time periods. however, with increasing model complexity, assessing the quality of the prediction resulting from such a transfer is becoming more difficult (randin et al. 2006, peterson 2007, elith et al. 2010, rödder and lötters 2010, peterson et al. 2011), and various possible error sources related to issues of methodology and the biological characteristics of the species have been identified (heikkinen et al. 2006, franklin 2009, jiménezvalverde et al. 2009, dobrowski et al. 2011, mcinerny and purves 2011, rocchini et al. 2011). among these, changes in the predictor colinearity through space and time may cause errors, not only when extrapolating, but also when interpolating enms within the area of interest (elith and leathwick 2009, jiménez-valverde et al. 2009, elith et al. 2010, rocchini et al. 2011). moreover, it has been suggested that some of the more sophisticated methods that use both information found from areas where the species is present and from where it is absent (presence/absence models) may characterize a species’ realized niche, where others, which only use the information contained in the records where species are known to occur (presence/pseudo-absence or presence-only models) quantify its potential distribution (rödder and lötters 2010, jiménez-valverde et al. 2011, peterson et al. 2011). these conceptual differences require additional attention when interpreting enm predictions and different types of uncertainties may arise when applying one over the other method. as stated above, the selection of appropriate predictor variables is a critical task during enm development and is one that is likely to dramatically influence the modeling results (e.g., peterson and nakazawa 2008, rödder et al. 2009, synes and osborne 2011, peterson et al. 2011, varela et al. 2011). generally, environmental predictors used in enms can be classified in a continuum spanning between proximal or distal predictors, depending on how they affect the fitness of the target species (austin 2002). proximal predictors are those actually affecting the physiology of the species, whereas distal variables do not directly affect the target species. unfortunately, it is commonly not possible to use the most proximal variables due to limited availability. hence, for most enm approaches, predictors are used that are rather intermediate or even distal within the continuum. nonetheless, within a complete set of imaginable environmental predictors, they may still be correlated with the most proximal predictors and therefore be informative. when assessing the correlation structure among all possible combinations of available predictors, this information can be used to estimate the correlation of the available predictors with those (unavailable) proximal predictors. distal predictors should only be used for enm development if a high degree of colinearity with proximal predictors exists – an assumption that is frequently fulfilled. modern enm algorithms search for a set of predictors or derived features thereof that best explain the variation in the dependent variable: the presence and/or absence of the target species across the different environments. as in a standard multiple regression model, these methods can supplement the supplied environmental predictors with extra predictor variables representing interactions between the predictors. within the region an enm is trained for, effects of proximal predictors can be estimated by distal predictors as long as both are correlated. these relationships and the implications for the transfer of the model to novel geographic regions are exemplified in figure 1: consider the high degree of correlation between the minimum temperature of the coldest month (assumed proximal variable) and the maximum temperature of the warmest month (assumed distal variable) within the ‘training area’ and its deviance in other areas. when transferring an enm onto environments differing from the training conditions, predictions may only be useful in those cases where the predictor colinearity between interacting variables in the enm as well as between proximal and distal variables remain stable. in those areas where proximal figure 1: illustration of the potential effects of local changes in the inter-correlation between two predictors across a longitudinal gradient in north america introduced by topographic factors as well as the underlying principles of purv (prediction uncertainty assessments using residual variation). assume that an enm is developed within one part of the study region (training area, black bar) and subsequently projected onto the rest of the gradient. the variability of the residuals within the training region can be used as reference to quantify the degree of inter-correlation between both predictors as incorporated in the enm. when projecting the enm through space or time, any change in the amplitude of residuals exceeding the variation present in the training area may indicate a change in the local inter-correlation, e.g. as present in projection area 1. projections outside of the training area of an enm may only be reliable when the amplitude of the residuals is not larger than in the training region (e.g. in projection area 2). the magnitude of the deviance can be quantified using the proportion of the residuals at a given grid cell exceeding the confidence limits within the training region. physiological constraints affecting a species’ survival may not be well reflected in the distal variables, model projections may become unreliable. in figure 1, we see that this is the case in ‘projection area 1’ where any prediction of an enm developed in the ‘training area’ may be misleading. in contrast, the degree of intercorrelation between both predictors is well within the training range in projection area 2 and, as a result, predictions may be much more reliable. engler and rödder – visualizing predictor colinearity   33 aside from possible uncertainties arising from the inter-correlation structure between proximal and distal variables, the inclusion of highly correlated variables may violate model assumptions in many regression techniques. therefore, a priori selection of a set of less correlated variables (e.g., r2 < 0.75) is often necessary, wherein the putatively biologically most relevant predictors should be selected (e.g., saupe et al. 2012). as long as the intercorrelation structure among selected and omitted predictors is stable, this information reduction is appropriate – but not when it changes, since in this case enm projections may become unreliable. although several authors mention likely problems caused by varying inter-correlation structures of predictors between training and projection conditions (e.g., heikkinen et al. 2006, peterson 2007, jiménez-valverde et al. 2009, elith et al. 2010), no technique allowing for a spatially-explicit evaluation of this error source is available, although elith et al. (2010) presented some general ideas. here, we propose the use of prediction uncertainty assessments using residual variation (purv) plots based on comparisons between the residual ranges of each pair of predictors within the training area of an enm and the projection areas to produce maps showing spatial explicit variations in correlation matrices. these plots can be used to identify areas where enm predictions may be prone to errors, which may lead to erroneous conclusions when interpreting prediction maps. methods species distribution models in order to illustrate the applicability of purv plots to assess variations in correlation structures, we developed enms for two sister species of north american salamanders (i.e. plethodon cylindraceus (harlan, 1825) and p. teyahalee hairston, 1950) as described in kozak and wiens (2006). a total of 180 georeferenced species records of p. cylindraceus and 761 records of p. teyahalee were obtained through the global biodiversity information facility (accessed through gbif data portal1; nmnh vertebrate zoology herpetology collections2; mvz herp catalog3; herp specimens4) and checked for possible georeferencing errors using                                                                                                                 1  www.gbif.org   2  http://data.gbif.org/datasets/resource/1838   3  http://data.gbif.org/datasets/resource/8123   4  http://data.gbif.org/datasets/resource/8956   diva-gis 7.1.6 (hijmans et al. 1999, hijmans et al. 2005a). information on current climate was extracted from the worldclim database (hijmans et al. 2005b5). the set of predictor variables used for enm development consisted of those variables identified by kozak and wiens (2006) as being biologically relevant to the salamanders: ‘mean annual temperature range’, ‘maximum temperature of the warmest month’, ‘minimum temperature of the coldest month’, ‘precipitation seasonality’, and ‘precipitation of the driest quarter’. using maxent 3.3.3e (phillips et al. 2006, phillips and dudík 2008, elith et al. 2010) we developed enms for both species. maxent requires the designation of a set of pseudoabsences for the calculation of species habitat preferences and, to this end, random background points were automatically sampled by maxent in a circular buffer around each record of 0.16° (~ 14 km radius), ensuring that each grid cell was considered only once to avoid pseudo-replication. this area reflects the geographic space potentially accessible for the species (phillips et al. 2009, anderson and raza 2010, barve et al. 2011, saupe et al. 2012). to assess the performance of the model, 100 enms per species were developed using the default maxent settings, but splitting the species records into 70 % used for training the model and 30 % to test model performances by calculating the area under the receiver operating curve (auc) (hanley and mcneil 1982, swets 1988). choosing the logistic output format with values ranging linearly from 0 (unsuitable) to 1 (optimal) (phillips and dudík 2008) the averages and standard deviations of the potential distributions of the species suggested by the 100 enms were projected onto a larger geographic area. the resulting maps show each species’ probability of occurrence per grid cell and its variation across all replicates. the degree of environmental novelty when extrapolating the enms into conditions outside those found in the training region was assessed using the application of multidimensional environmental similarity surfaces (mess) as described by elith et al. (2010) (i.e., using the relevant tool implemented in maxent 3.3.3e). in mess maps, values range from +100 to –∞, with positive values indicating grid cells with environmental conditions within the range of the                                                                                                                 5  www.worldclim.org   figure 2: principles of multivariate environmental similarity surfaces (mess; elith et al. 2010): if the observed value pi at grid cell pi lies within the range of variable vi in the training region, the similarity of this grid cell with respect to vi is defined as percentage deviance from the median of the observed range (arrow 1). this creates a gradient across the environmental envelope within the training area ranging from 100 in the center to 0 at the margins indicated by the black box. in multivariate environmental space this approach is equal to the bioclim approach (busby 1991). in those cases where any p is outside of this environmental envelope spanned by v within the training region, the distance to its margin of the most divergent variable vi is measured (arrow 2), wherein the scores reflect the quotient of the absolute deviance of pi from the margin of the envelope’s range. a score of -100 would indicate a deviance of pi equaling the range of vi within the training area. in multivariate space, the score of the most divergent vi at pi is assigned. in purv plots, grids showing residuals are used instead of environmental variables themselves. variables in the training region (analogous to the cells predicted as suitable in the bioclim algorithm; busby 1991) and negative values denoting grid cells with environmental conditions that fall outside the range of environments present in the training area (as described by the boxcar environmental envelope – see figure 2). note that the training region of an enm ideally comprises those areas that are potentially colonizable for the target species, wherein a number of different configurations are possible (saupe et al. 2012). mess maps were transformed into binary shapefiles, indicating those areas where at least one variable falls outside the range present in the training area of the enm. prediction uncertainty assessments using residual variation (purv) for explicit spatial comparisons of correlation structures among predictors across geographic space, we developed a function for the r statistical platform for the construction of purv plots (supplementary material appendix 1). as reference for subsequent analyses, we assessed the inter-correlation structure of each pair of z-standardized predictor variables within the training area of the enm. z-standardization referred to the average and standard deviation of the respective variable within the training area. therefore, we used simple linear regressions and determined the intercept and slope of the corresponding regression. we used linear instead of other, more complex, relationships as basic assumption for this analysis, as the true intercorrelation structure for a given projection area is a priori not known and might change between different predictor variable comparisons. therefore, assuming a linear relationship represents the most conservative approach for the construction of purv plots. for each pair of predictor variables, we created a grid covering the previously-defined projection area showing the spatial distribution of residuals in r 2.12 (r development core team 2011), which is parameterized based on the correlation structure in the training area only. these grids allow for a spatially-explicit quantification of the variation in the correlation structure within both the training and projection areas, where the magnitude of the variation within the residuals in the training area of the enm can be used as reference. in those projection areas where the residuals exceed the range observed within the training area, it can be assumed that the two variables are less correlated locally. the number of possible pair-wise comparisons of predictors can quickly become large, making detailed analyses time consuming. therefore, mess analyses can be used to summarize the maximum possible effect of intercorrelation changes as described by elith et al. (2010). the novelty here is that the mess plots are derived from the residual grids instead of the original environmental predictors. wherein traditional mess analyses ask whether the environmental conditions at a given site are within those conditions available within the engler and rödder – visualizing predictor colinearity   35 training area, our purv approach asks whether the inter-correlation structure of the environmental variables at a given site are similar to those within the training area or not. purv plots allow the quantification of errors in two different areas (compare figure 2): (1) in regions characterized by the same inter-correlation structure present in the training region, the plots show the relative distance of each grid cell to the center of the multivariate boxcar residual envelope, and (2) in areas in which the intercorrelation structure deviates from the intercorrelation structure within the training area, the plots show the distance to the border of the boxcar envelope of the most distant residual. in the former case positive values ranging from 100 (center of the boxcar envelope) to 0 (its margin) are assigned, while in the latter case, negative scores are assigned, identifying the maximum deviant change in the predictor correlation (for a more detailed description of the mess procedure see elith et al. 2010). enm projections into areas where the range of residuals in at least one predictor-pair is greater than that found within the training area can be interpreted as being less reliable. this becomes reasonable when considering two different cases: (1) interactions between two variables fitted by an enm refer to the correlation structure of the predictors within the training area. unreasonable response curves might result when projecting these relationships onto deviating correlation structures in other areas or time slices. (2) in many cases it is necessary to select only the putatively biologically most important variable for enm development if a set of variables is highly correlated. as long as the correlation structure remains stable this subjective selection may cause no negative effects, but in those cases where the correlation structure between a selected variable and one omitted variable changes, the selected variable cannot be used to estimate the omitted ones. the purv plots indicate spatial changes in the correlation structure among predictor variables irrespective and unaffected from the enm algorithm. for ease of interpretation, it may be helpful to transform the continuous scale of the residuals deviation into a binary classifier where a ‘1’ value is given to any cell with a residual deviation found within the confidence interval of the residual deviation of cells within the training region. all other cells would be given a value of ‘0’. results a summary of the model performances in terms of training and test aucs, relative contributions of the environmental predictors to the final enm, and presence/absence thresholds is given in table 1. auc scores on average indicate good model performance and enms developed for both plethodon cylindraceus and p. teyahalee depict the known distributions of the species very well. however, they also project high probabilities of occurrence in some areas outside the known range of each salamander species (figure 3a, b), many of which are situated in the known range of the other species. this is not unexpected since both are sister taxa and likely occupy similar environmental niches (kozak and wiens 2006), thus a mixture of dispersal limitation and biotic interaction most likely explains the lack of this species in areas deemed suitable. as indicated by the sd per grid cell, some of the projected potentially suitable areas for each species show a comparatively high variability (sdmax = 0.29 in both) (figure 3, middle row) suggesting a higher degree of model uncertainty. only mess analyses based on the original predictor variables highlight parts of them as prone to potential errors (figure 3, bottom). moreover, the areas identified by purv plots only partly overlap with them, indicating that the spatial distribution of both error sources is not necessarily coincident. this appears to be reasonable since the extrapolation situation identified by mess refers to the range of the variables within the training region, but the purv plots identify changes in the correlation structure among predictors. even if the range of predictors exceeds the environmental conditions within the training area as indicated by mess, the correlation structure might still be within the training range. vice versa, this might also explain the spatial independence of mess and purv plots. discussion our results highlight the importance of considering both the relative ranges of environmental parameters as well as their intercorrelation structures in the training region when projecting enms into novel locations and time periods. engler and rödder – visualizing predictor colinearity   36 figure 3: occurrence probability maps of plethodon cylindraceus (a) and p. teyahalee (b), corresponding maps of standard deviations of 100 maxent models (c, d) and prediction uncertainty assessment using residual variation analysis (purv) plots (e, f). regions highlighted by purv confidence limits are indicated as downward hatched, wherein areas requiring model extrapolation identified by multidimensional environmental similarity analyses (mess) are upward hatched. models were trained with species records (white dots) and random background points drawn within circular buffers (white outlines). extrapolation errors in both species, multivariate environmental similarity surfaces (mess) based on the original environmental predictors identified varying proportions of the studied areas in which at least one variable exceeded its range present in the training region (figure 3, appendix a). here, assigning any occurrence probability requires engler and rödder – visualizing predictor colinearity   37 extrapolation by the model. in most cases, the enms did not predict the species in those novel environments, but where they did, sd of the corresponding occurrence probabilities became much higher compared to other regions, thus depicting extrapolation areas quite well (e.g. in central areas in p. cylindraceus and in the northeastern corner of p. teyahalee). predictions in these areas are much more variable then interpolated predictions due the requirement of model extrapolation, and hence less reliable. other portions of the study area where the species were not predicted to occur but requiring extrapolation such as the south-eastern areas were correctly identified as being too hot in the warmest month in p. cylindraceus or too cold in the coldest month in p. teyahalee. this was correctly incorporated in the enms by fading response curves at the corresponding tail of the temperature range (appendix a). note that both predictors, the maximum temperature of the warmest month and the minimum temperature of the coldest month, had the highest overall variable contribution in the respective enms (table 1). uncertainties due to extrapolation may therefore only arise when high response values for a given parameter are extrapolated but not when physiological limits are correctly captured by fading response curves. for further discussion how the shape of response curves may affect the reliability of enms, see santika and hutchinson (2009). spatial changes in the inter-correlation structure purv plots indicated that the correlation structure of the variables in the training area was only partly coincident with those in the projection area (figure 3). a number of spots in the study areas were identified as showing substantial differences in their correlation structures. however, the species were only occasionally predicted to occur in the areas highlighted by purv plots (e.g. northern and north-eastern parts of the study area in p. teyahalee). unlike in areas requiring extrapolation of the enm beyond the parameter range within the training region, no relationship between sd in the enms and intercorrelation change was evident, suggesting that this kind of error source may frequently remain undetected in traditional analyses. jiménez-valverde et al. (2009) identified fairly-to-highly similar correlation structures on a continental scale in the most commonly used climate data sets using mantel tests, where the correlation structures were most similar when comparing different time slices than when comparing different areas. whilst these comparisons on a continental scale may indicate detectable but rather low variations, this may not necessarily be true on smaller scales. we obtained quite different purv plots for both plethodon species, suggesting that spatial position of the areas exhibiting the largest table 1. performance and characteristics of maxent models for plethodon cylindraceus and plethodon teyahalee. auc = area under the receiver operating characteristic curve; min train = minimum training presence; 10 % train = lowest 10 percentile training omission. species p. cylindraceus p. teyahalee model performance training auc 0.836 0.802 test auc 0.807 0.790 variable contributions [%] mean annual temperature range 17.4 6.8 maximum temperature warmest month 57.9 22.4 minimum temperature coldest month 5.3 61.9 precipitation seasonality 2.5 3.5 precipitation driest quarter 16.9 5.3 presence/absence thresholds min training 0.065 0.034 10 % training 0.239 0.252 changes in the correlation structure strongly depend on the position of the training area. on this small scale, the spatial variability of the residuals within the training areas can be much smaller than can be expected on a continental scale due to fine-scale topographic features. this increases the probability of locally-occurring high inter-correlations among variables resulting in a higher chance of deviations across space even within a small area. as a consequence, although the negative impacts caused by changes in the inter-correlation structure of predictors may be rather small on large scales, they may bear heavily on enm predictions for species occupying rather small ranges (e.g. the two plethodontid salamanders). conclusions as illustrated by our results, purv plots are able to highlight formerly undetected areas of uncertainty in projection areas of enms, which are per se independent of the employed algorithm. however, the impact of the uncertainty caused by changes in the correlation structure of predictors may strongly depend on the algorithm used. it may be absent when using simple profile algorithms, which do not incorporate variable interactions at all such as bioclim (busby 1991) or domain (carpenter et al. 1993), where an appropriate choice of relevant proximal predictors may be more important. in contrast, the impact of changes in the correlation structure of predictors may become highest when interactions among predictors are incorporated in the model, as is commonly the case with state-of-the-art algorithms (elith and leathwick 2009, franklin 2009, thuiller et al. 2009), such as artificial neural networks, boosted regression trees, generalized additive models, classification and regression trees, random forests or maxent. furthermore, projection areas identified using purv plots might become unreliable when interpredictor relationships deviate from simple colinearity into more complex relationships, since residual deviation may be increased. there is a need for future extensions incorporating more complex relationships into this approach, and the authors welcome useful comments from the scientific community. summarizing, purv plots in its current stage may be very helpful to identify areas prone to misleading predictions, an area of research that requires more attention in the future use of enms. acknowledgments we are grateful to kai f. abt, joseph chipperfield and ursula bott for helpful comments on an early version of the manuscript. d.r. is grateful for financial support by the ministry of education, science, youth and culture of the rhineland-palatinate state of germany within the framework of the project ‘implications of global change for biological resources, law and standards.’ comments by a. townsend peterson and an anonymous reviewer greatly improved the manuscript. references anderson, r.p., and a. raza (2010). the effect of the extent of the study region on gis models of species geographic distribution and estimates of niche evolution: preliminary tests with montane rodents (genus nephelomys) in venezuela. journal of biogeography 37: 1378-1393. araújo, m.b., m. cabeza, w. thuiller, l. hannah, and p.h. williams (2004). would climate change drive species out of reserves? an assessment of existing reserve-selection methods. global change biology 10: 1618-1626. austin, m.p. (2002). spatial prediction of species distribution: an interface between ecological theory and statistical modelling. ecological modelling 157: 101-118. barve, n., v. barve, a. jiménez-valverde, a. liranoriega, s.p. maher, a. t. peterson, j. soberón, and f. villalobos (2011). the crucial role of the accessible area in ecological niche modeling and species distribution modeling. ecological modelling 222: 1810-1819. breiman, l. (2001). random forests. machine learning 45: 5-32. brown, j.l., and e. twomey (2009). complicated histories: three new species of poison frogs of the genus ameerega (anura: dendrobatidae) from north-central peru. zootaxa 2049: 1-38. busby, j.r. (1991). bioclim a bioclimatic analysis and prediction system. in: margules, c.r. and m.p. austin (eds.), nature conservation: cost effective biological surveys and data analysis. csiro, melbourne, pp. 64-68. carpenter, g., a.n. gillison, and j. winter (1993). domain: a flexible modeling procedure for mapping potential distributions of plants and animals. biodiversity and conservation 2: 667680. dobrowski, s.z., j.h. thorne, j.a. greenberg, h.d. safford, a.r. mynsberge, s.m. crimmins, and a.k. swanson (2011). modeling plant ranges over 75 years of climate change in california, usa: temporal transferability and species traits. ecological monographs 81: 241-257. elith, j., and j.r. leathwick (2009). species distribution models: ecological explanation and prediction across space and time. annual engler and rödder – visualizing predictor colinearity   39 reviews in ecology, evolution and systematics 40: 677-697. elith, j., c.h. graham, r.p. anderson, m. dudík, s. ferrier, a. guisan, r.j. hijmans, f. huettmann, j.r. leathwick, a. lehmann, j. li, l.g. lohmann, b.a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. mcc.m. overton, a. townsend peterson, s.j. phillips, k. richardson, r. scachetti-pereira, r.e. schapire, j. soberón, s. williams, m.s. wisz, and n.e. zimmermann (2006). novel methods improve prediction of species´ distributions from occurrence data. ecography 29: 129-151. elith, j., m. kearney, and s. phillips (2010). the art of modelling range-shifting species. methods in ecology and evolution 1: 330-342. franklin, j. (2009). mapping species distributions: spatial inference and prediction. cambridge university press. friedman, j. (1991). multivariate adaptive regression splines. annals of statistics 19: 1-141. guisan, a., and n. zimmermann (2000). predictive habitat distribution models in ecology. ecological modelling 135: 147-186. hanley, j., and b. mcneil (1982). the meaning of the use of the area under a receiver operating characteristic (roc) curve. radiology 143: 2936. heikkinen,, r.k., m. luoto, m.b. araújo, r. virkkala, w. thuiller, and m.t. sykes (2006). methods and uncertainties in bioclimatic envelope modeling under climate change. progress in physical geography 30: 751-777. hernandez, p.a., c.h. graham, l.l. master, and d.l. albert (2006). the effect of sample size and species characteristics on performance of different species distribution modeling methods. ecography 29: 773-785. hijmans, r.j., m. schreuder, j. de la cruz, and l. guarino (1999). using gis to check coordinates of genebank accessions. genetic resources and crop evolution 46: 291-296. hijmans, r.j., l. guarino, a. jarvis, r. o’brien, p. mathur, c. bussink, m. cruz, i. barrantes, and e. rojas (2005a). diva-gis version 5.2 manual. hijmans, r.j., s.e. cameron, j.l. parra, p.g. jones, and a. jarvis (2005b). very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25: 1965-1978. jiménez-valverde, a., y. nakazawa, a. lira-noriega, and a.t. peterson (2009). environmental correlation structure and ecological niche model projections. biodiversity informatics 6: 28-35. jiménez-valverde, a., a.t. peterson, j. soberón, j.m. overton, p. aragón, and j.m. lobo (2011). use of niche models in invasive species risk assessments. biological invasions 13: 2785-2797. kozak, k., and j.j. wiens (2006). does niche conservatism promote speciation? a case study in north american salamanders. evolution 60: 2604-2621. kozak, k.h., c.h. graham, and j.j. wiens. (2008). integrating gis-based envrionmental data into evolutionary biology. trends in ecology and evolution 23: 141-148. mcinerny, g.j., and d.w. purves (2011). fine-scale environmental variation in species distribution modeling: regression dilution, latent variables and neighborly advice. methods in ecology and evolution 2: 248-257. olden, j.d., j.j. lawler, and n.l. poff (2008). machine learning methods without tears: a primer for ecologists. quarterly reviews in biology 83: 171-183. peterson, a.t. (2007). why not why where: the need for more complex models of simpler environmental spaces. ecological modelling 203: 527-530. peterson, a.t. (2011). ecological niche conservatism: a time-structured review of evidence. journal of biogeography 38: 817-827. peterson, a.t., and y. nakazawa (2008). environmental data sets matter in ecological niche modeling: an example with solenopsis invicta and solenopsis richteri. global ecology and biogeography 17: 135-144. peterson, a.t., and d.a. vieglais (2001). predicting species invasions using ecological niche modeling: new approaches from bioinformatics attack a pressing problem. bioscience 51: 363371. peterson, a.t., j. soberón, r.g. pearson, r.p. anderson, e. martínez-meyer, m. nakamura and m.b. araújo (2011). ecological niches and geographic distributions. princeton university press, princeton, 328 pp. phillips, b.l., j.d. chipperfield, and m.r. kearney (2008). the toad ahead: challenges of modelling the range and spread of an invasive species. wildlife research 35: 222–234. phillips, s.j., and m. dudík (2008). modeling of species distributions with maxent: new extensions and comprehensive evaluation. ecography 31:161-175. phillips, s.j., r.p. anderson, and r.e. schapire (2006). maximum entropy modeling of species geographic distributions. ecological modelling 190: 231-259. phillips, s.j., m. dudík, j. elith, c.h. graham, a. lehmann, j. leathwick, and s. ferrier (2009). sample selection bias and presence-only distribution models: implications for background and pseudo-absence data. ecological applications 19: 181-197. r development core team (2011). r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. r foundation for statistical computing. randin, c.f., t. dirnböck, s. dullinger, n.e. zimmermann, m. zappa, and a. guisan (2006). are niche-based species distribution models transferable in space? journal of biogeography 33: 1689-1703. engler and rödder – visualizing predictor colinearity   40 rocchini, d., j. hortal, s. lengyel, j.m. lobo, a. jiménez-valverde, c. ricotta, g. bacaro, and a. chiarucci (2011). accounting for uncertainty when mapping species distributions: the need for maps of ignorance. progress in physical geography 35: 211-226. rödder, d., and s. lötters (2010). explanative power of variables used in species distribution modelling: an issue of general model transferability or niche shift in the invasive greenhouse frog (eleutherodactylus planirostris). naturwissenschaften 97: 781-796. rödder, d., s. schmidtlein, m. veith, and s. lötters (2009). alien invasive slider turtle in unpredicted habitat: a matter of niche shift or of predictors studied? plos one 4, e7843. santika, t., and hutchinson, m.f. (2009). the effect of species response form on species distribution model prediction and inference. ecological modelling 220: 2365-2379. saupe, e., v. barve, c. myers, j. soberón, n. barve, c. hensz, a.t. peterson, h.l. owens, and a. lira-noriega (2012). variation in niche and distribution model performance: the need for a priori assessment of key causal factors. ecological modelling 237–238: 11-22. soberón, j., and a.t. peterson (2005). interpretation of models of fundamental ecological niches and species' distributional areas. biodiversity informatics 2: 1–10. swets, k. (1988). measuring the accuracy of diagnostic systems. science 240: 1285-1293. synes, n.w., and p.e. osborne (2011). choice of predictor variables as a source of uncertainty in continental-scale species distribution modelling under climate change. global ecology and biogeography 20: 904-914. thomas, c.d., a. cameron, r.e. green, m. bakkenes, l.j. beaumont, y.c. collingham, b. f.n. erasmus, m. ferreira de siqueira, a. grainger, l. hannah, l. hughes, b. huntley, a.s. van jaarsveld, g.f. midgley, l. miles, m.a. ortegahuerta, a.t. peterson, o.l. phillips, and s.e. williams (2004). extinction risk from climate change. nature 427: 145-148. thuiller, w., b. lafourcade, r. engler, and m.b. araújo (2009). biomod a platform for ensemble forecasting of species distributions. ecography 32: 369-373. varela, s., j.m. lobo, and j. hortal (2011). using species distribution models in palaeobiogeography: a matter of data, predictors and concepts. palaeogeography, palaeoclimatology, palaeoecology 310: 451-463. wisz, m.s., r.j. hijmans, j. li, a.t. peterson, c.h. graham, a. guisan, and nceas predicting species distributions working group (2008). effects of sample size on the performance of species distribution models. diversity and distributions 14: 763-773.   biodiversity informatics, 16, 2021, pp. 20-27 20 visualizing species richness and site similarity from presence-absence matrices jorge soberón1*, marlon e. cobos1, and claudia nuñez-penichet1 1department of ecology and evolutionary biology & biodiversity institute, university of kansas, lawrence, kansas, united states. abstract. species richness and similarity of biotas among distinct sites are important quantities in biogeography. indices derived from presence-absence matrices are used to represent these quantities in so-called range-diversity plots. the most commonly used range-diversity plot, however, has multiple special cases and its interpretation is cumbersome. here we present an equivalent formulation that is geometrically simpler and has no special cases. in addition, we introduce a method to identify the statistical significance of the dispersion field, an index that represents how similar species composition is in a cell with respect to the whole area. the new range-diversity plot is a promising tool to explore biodiversity and endemism in a region as the values shown in this plot and their statistical significance can also be represented in geography. key words: biodiversity index, dispersion field, range-diversity plot, endemism, pam introduction representing biodiversity is a complex challenge, mainly because the concept itself is poorly defined (sarkar 2002). most often, “biodiversity” is used to mean the list of species in a region. one way of organizing biodiversity data is by using presence-absence matrices (pams), where a one represents presence of species j in cell i, and a zero, absence (fig. 1). pams could be regarded as mathematical representations of lists of species per spatial unit. originally, pams were developed for dozens of species in a handful of islands in archipelagos (connor and simberloff 1979), which means that, originally, pams had a few hundreds of cells. however, it is possible and desirable to expand the concept of a pam to arbitrary grid partitions of large regions, with extents of 106 km2 or larger, and grids of resolutions in the order of 1 to a few thousand km2. using these grid partitions allows obtaining more detailed information on the incidence of species in a region, but also helps in visualizing biodiversity patterns via distinct plots of the indices that can be calculated from a pam (soberón and cavner 2015). pams have been analyzed with a variety of techniques (arita et al. 2008; christen and soberón 2009; ulrich and gotelli 2012; soberón and cavner 2015). one of the methods is the range-diversity plot (arita et al. 2008; borregaard and rahbek 2010; soberón and ceballos 2011). this plot displays jointly two indices describing the community composition of every cell in the grid: (1) the number of species, in absolute or relative numbers, and (2) the mean dispersion field (graves and rahbek 2005). the dispersion field is simply total range size of all the species occupying a cell. this can be proven to be equivalent to the total amount of overlap of the community of species in a cell, with all the other cells (soberón & ceballos, 2011). instead of total values, in the range-diversity plot, the mean value of the dispersion field is used. the mean is taken with respect of the species existing in each cell. the mean can be taken also with respect to all species in the region, as we shall see later. the range-diversity plot displays a large amount of data in a compact way that allows interpretations. the mathematical limits of the range-diversity plot are well-defined (arita et al. 2008; soberón and ceballos 2011) and constrains the scatterplot in a very strict way. we refer the reader to these articles for all mathematical details. moreover, these limits have a straightforward geographic interpretation (soberón and ceballos 2011). however, the geometric properties of a range-diversity plot are somewhat cumbersome and subject to many particular cases (i.e., plots are not linear and their shape depends on the specifics * corresponding author: jorge soberón department of ecology & evolutionary biology and biodiversity institute, university of kansas, lawrence, kansas 66045, usa. jsoberon@ ku.edu orcid: js: https://orcid.org/0000-0003-2160-4148 mec: https://orcid.org/0000-0002-2611-1767 cnp: https://orcid.org/0000-0001-7442-8593 https://scholar.google.com/citations?view_op=view_org&hl=es&org=13272385838766914177 https://orcid.org/0000-0003-2160-4148 https://orcid.org/0000-0002-2611-1767 https://orcid.org/0000-0001-7442-8593 biodiversity informatics, 16, 2021, pp. 20-27 21 of the parameters), making them more difficult to interpret. for instance, the sign of the minimum covariance determines the slope of the lines defining the limits, leading to rather contrasting shapes (fig. 2); and the relationship between minimum relative richness and minimum covariance determines whether the limits will be truncated or not (not illustrated). in this contribution, we introduce a formulation of the range-diversity plot, which is equivalent to, but geometrically much simpler than the previous one, and does not have particular cases, facilitating interpretations and further explorations. we also present a way to detect whether observed values of the index related to species ranges (the dispersion field) are statistically significant, allowing for further interpretations in regards to how species are distributed in a region, based on a pam. methods data to illustrate the application of the new method to produce range-diversity plots, we use a pam originally composed of s = 1595 species of terrestrial vertebrates, on a grid of n = 711 cells (resolution of 0.5°, or ~55 km at the equator) subdividing mexico. the original data used to produce the pam comes from the international union for the conservation of nature (iucn), which maintains a database of downloadable, machine-readable maps (in shapefile format) for the non-avian terrestrial (increasingly also aquatic) vertebrates of the world (iucn 20201). in this pam, each cell is represented by the geographic coordinates of its centroid and can be downloaded using the code provided to replicate analysis and plots (data and code are available2). after excluding cells with no species, or species with values of 0 (absence) for all cells, the final number of species and cells was 1573 and 711, respectively (fig. 1). calculations for new plotting method the basic idea for the range-diversity plot follows from an identity proven in soberón and cavner (2015) and based on the variance-covariance matrix of species in cells. this is the covariance in species composition among all cells in the grid. by forming the vector τ of the average covariance of every cell to all other cells, the following identity is proven: 1 https://www.iucnredlist.org/resources/spatial-data-download. 2 https://github.com/jsoberon/pams-mexico. where s is the number of species, φ* is the vector of normalized (divided by n) dispersion field values, α is the vector of numbers of species, and β is whittaker’s (1960) beta “diversity”. a division by n creates a proportion with respect to the size of the region in question. whittaker’s first equation for beta diversity is simply /s α , which, by a simple replacement is / /s ns fβ α= = where f is the “fill” of the pam, or the total number of ones. from the above arita et al. (2008) obtained the following equation, which is the base of their range-diversity plot: the range-diversity plot represents every cell in the grid, using as y-axis the quantity αi* (the asterisk represents that α has been normalized by dividing it by s) and as x-axis the mean dispersion field per cell, or * iϕ (with the asterisk representing division by n and to get the mean dispersion field per cell iϕ is divided by αi. as we shall exemplify below, the above equation implies that the data points in the scatterplot are enclosed within a tent-like “envelope” determined by the minimum and maximum mean covariances, and by the beta diversity (fig. 2). this envelope is non-linear in shape and has many special cases and one potential discontinuity. however, a much simpler and direct graph comes directly from equation (1). indeed, element-by-element equation (1) means: figure 1. graphical representation of a presence-absence matrix (top panel, presence = 1 in brown, absence = 0 in yellow), for the pam (representing 711 cells and 1573 species) of the terrestrial vertebrates of mexico; and geographic representation of indices of richness and dispersion field derived from it. https://www.iucnredlist.org/resources/spatial-data-download https://github.com/jsoberon/pams-mexico biodiversity informatics, 16, 2021, pp. 20-27 22 or, by an immediate rearrangement: this suggests that * /i sϕ vs. * iα creates a simpler graph, as it is indeed the case. moreover, the theoretical limits of this new range-diversity plot are straightforward. the limits form a parallelepiped with vertical lines at αmin and αmax, and straight lines with a slope of 1/β and intercepts at the minimum and maximum values for τ (fig. 2). notice the main difference is that instead of plotting the mean normalized dispersion field, with the mean taken by dividing by the local number of species ( * iα ), the mean is taken over all the species in the region. in the discussion, we further expand this point. statistical significance of dispersion field indices derived from a pam can be tested for significance. for instance, schluter (1984) provides a test for the similarity of the distributions of a number of species, using a variance ratio. however, a pam permits the calculation of a large variety of indices. a more general approach to the problem of statistical significance relies on the randomization of pams. pams can be randomized under a number of assumptions (miklós and podani 2004) and this allows calculation of the distributions of values of the dispersion field, the mean dispersion field, the distribution of richness, overlaps, or many other quantities, from those that cannot be distinguished from random expectations and others that differ significantly from such a distribution. randomizing pams requires making important decisions (strona et al. 2018), specifically how the marginal sums of the pam will be treated. these marginals represent the total number of species per site (richness, the vector α), when summing over columns, or the total number of sites each species uses (incidence, the vector ω) when summing over rows. the question then is whether to leave richnesses, or incidences, or both, constant at the time of randomizing. this is a problem well beyond the purpose of this contribution (see discussion), but we used the approach of randomizing subject to fixed richness and incidence. we used this approach because we assume that both incidence and richness are observed data, deriving from reliable sources of information. to perform the randomization process, we used the function “randomizematrix” from the package picante (kembel et al. 2010) in r 4.0.2 (r core team 2020). for the process of randomization, we used the null model “independentswap” which randomizes the matrix with the independent swap algorithm from gotelli (2000). considering how the randomization is done, before starting the process, we broke any potential geographic structure remaining in the pam (where closer rows are also closer cells in the geography) by rearranging randomly the rows (which represent geographic cells) but maintaining the identity of these cells in the results. after each randomization, the dispersion field is calculated from the resulting pam and it is normalized. the result of repeating this process a number of iterations (500 in our case) generates a distribution of values of the dispersion field under random expectations. then, the actual values of any index, for each cell, are compared to the random distribution of values, and the observed values that are as extreme or more extreme than the confidence limits of the null distribution (by default 5%: the 2.5% lower and 97.5% upper limits) are considered statistically significant. significance is treated differently when the actual value is below the lower limit or when it is above the upper limit as this has implications for interpretation (see discussion). calculations of biodiversity indices derived from the pam, and its randomization, were done using the function “prepare_pam_cs” from the package biosurvey (nuñez-penichet et al. 2020; available3) in r. geographic representations of the results obtained 3 https://github.com/claununez/biosurvey. figure 2. two special cases of the range-diversity plot. notice that in the old diagram the change of sign in the minimum covariance produces a very different-looking envelope. in the new diagram, it is the size that changes. https://github.com/claununez/biosurvey biodiversity informatics, 16, 2021, pp. 20-27 23 with the new method were done using the package maps (brownrigg et al. 2018) and other base functions in r. exploring the new range-diversity plot to illustrate the differences between the previous range-diversity plot and the new diagram proposed here, we exemplify two theoretical special cases of these plots. the two cases are: 1) a positive minimum covariance when comparing communities per site (grid cells in our case); and 2) the case with a negative minimum covariance. using the pam based on the terrestrial vertebrate species of mexico, we produced the original and the new range-diversity plots. based on the same pam, we also detected statistically significant values of the dispersion field and identified them in the new range-diversity plot and in a geographic representation for mexico. as the relationship between points in the range-diversity plot and areas in the geographic region of interest is of high importance, we created visualizations of block-like chunks of points in the diagram and their geographic projections. blocks in the range-diversity plot were divided using a grid (equal-size cells) of 4 rows and 4 columns using biosurvey in r. four of the resultant blocks were selected randomly to be used for plots. results in figure 2 we display the old and the new range-diversity plots. the plots in the top row in figure 2 are examples of the original range-diversity diagram proposed by arita et al. (2008). in this version, the data points are constrained to occur under a tentlike curved region determined by quantities derived from the pam. the shape and position of the tentlike permitted zone are determined by the beta diversity (whittaker), the minimum and maximum mean covariance of the community of species in a cell, relative to every other; and indirectly, by the minimum and the maximum number of species in the cells of the grid (fig 2, upper row). there are several possible particular cases of this diagram, as illustrated in figure 2 for two possibilities only. in the bottom row of figure 2, we present the new range-diversity plot, where the order of the axes is reverted, and the normalizations are not exactly the same. there is still a permitted region, determined by the same values, but the permitted region is now a simple parallelogram, always of the same shape. the area of the parallelogram, though, may change with the values of the parameters (minimum and maximum alpha, beta, and minimum and maximum mean covariances), but the general aspect of the plot is always the same. the range-diversity plot of the iucn pam for 1573 species of terrestrial vertebrates of mexico is displayed in figure 3 using the old and the new diagrams. notice that the points form clusters. as shown below, these clusters have a geographic meaning. figure 3. comparison between previous range-diversity plot and the new plot proposed created with iucn species ranges for mexico (summarized in a pam). https://www.zotero.org/google-docs/?broken=dtrmcf biodiversity informatics, 16, 2021, pp. 20-27 24 the differences between the old and the new diagrams are: (i) the variables used in the old are the normalized mean (with respect to local richness) dispersion field and the normalized species richness (of every cell), or αi* vs. φi*/αi. in the new one, the variables are the normalized dispersion field divided by the total number of species, which can be interpreted as a mean with respect to the number of species in the entire region, and the normalized (with respect to s) species richness. in symbols, we plot φi*/s vs. αi*. (ii) the limits of the permitted regions are determined by the minimum and maximum mean covariance in both diagrams, in the old one these are hyperbolic lines, whereas in the new plot they are straight lines of slope 1/β. in figure 4 we show the new range-diversity plot for the iucn pam, as well as a geographic representation of the results obtained. the top-left panel is the new diagram. the top-right panel shows the values of the dispersion field resulted from the randomizations of the pam. we randomized the pam, subject to fixed marginals. this randomization changes the values of the dispersion field, but not those of the richness (by construction: richness per cell is kept constant during randomizations). values above and below the 97.5% and the 2.5% percentiles of the random distribution are illustrated in the bottom-left panel, with the corresponding geographic representation in the bottom-right part of figure 4. circles in black are above the 97.5% percentile of the randomizations, and those closed circles in gray are below the 2.5%. open circles in gray are inside the 2.5% to 97.5% range of randomizations. notice that the map shows three clearly differentiated regions: in black are cells with dispersion-fields (remember that the mean dispersion field is identical to mean similarity) significantly larger than random, dark gray cells are those with dispersal fields smaller than random, and light gray cells are those indistinguishable from random. in general, our results show regions of higher than randomly expected similarity in the nearctic part of mexico, and lower than randomly expected similarities in the neotropical part. a transition region where similarities are as expected under a random model is also observed in the transition between neotropics and nearctic. finally, the range-diversity plot displays clusters or groups of points, which correspond to cells in geography with distinct patterns (some of them clustered but others disjoint; fig. 5). exploring this duality helps to understand how richness and similarity of species’ composition are distributed in the geography, since recognizing where points in the range-diversity plot are located geographically is not necessarily intuitive. discussion since its introduction by arita et al. (2008) the range-diversity plot has been used in a number of analyses (for instance, borregaard and rahbek 2010; villalobos et al. 2013a and b; 2017; zwiener et al. 2018), but it is a bit cumbersome to read. it has many special cases, contains discontinuities around the value of 1/β, and its “envelope” is non-linear. in this work, we propose a much simpler variant that displays the same information. it can be used, for instance, to obtain richness-rarity quartiles (villalobos et al. 2013a and b), and has the same properties of relation with maps in geographic space, but it is simpler to interpret and has no differently-shaped special cases (figs. 2-3). for instance, in the new plot, having negative covariances among-communities composition does not create, as in the previous plot, positive and negative clouds of points; now the entire cloud has a positive slope. the mean dispersion field is simply the sum of jaccard similarities among regions (soberón and ceballos 2011). when taking the mean relative to the s species in the region, rather than to iα , the local species number, one gets a smaller measure of similarity, figure 4. plotting options for the new range-diversity plot and representation of results in the geography. the non-statistically significant cells map into a region of mexico that represents the transition between tropical and temperate conditions. biodiversity informatics, 16, 2021, pp. 20-27 25 since s > any local α. this idea needs to be kept in mind when interpreting the new plot. we also propose an application of randomization of a pam to obtain intervals of confidence for one statistic derived from the pam, the normalized dispersion field. because our randomization process starts by reordering arbitrarily the rows in the matrix (cells or sites), spatial autocorrelation is, in some sense, broken, preventing artifacts due to it. although we do not prove it here, several metrics of a pam are invariant to randomization with fixed marginals (obviously the mean richness and whittaker’s beta, but also the mean dispersion field, the mean diversity field, and others). fortunately, the dispersion field is sensitive to such type of randomization, and since, as it has been proven in soberón and ceballos (2011), the dispersion field of a cell is mathematically equivalent to the total number of shared species with all the others cells, identifying those cells that are statistically significant is a useful result (figs. 4-5). the randomization of the entire pam allows testing the significance of any of the parameters derived from it. in this sense, our approach is more general than those that rely on the significance of specific indices. as said before, schluter (1984) presented a variance test to compare the composition of two regions. this has been applied, for instance, by villalobos et al. (2014). our approach would permit to test directly whether composition, or nestedness, or any parameter sensitive to the randomization of a pam has an extreme value, with respect to the null model of fixed marginals. the null models obtained by fixing either of the marginals of a pam have biological interpretations. in an interesting paper, gibert and escarguel (2019) suggest that fixing the species numbers marginal is equivalent to hypothesizing that the dominant community assembly processes are “niche” based, and fixing the species ranges marginal is a hypothesis about the dominance of dispersal factors. although we do not explore these ideas here, we notice that the range-diversity plots are ideally suited to do this exploration. statistically significant values of the dispersion field may have a geographic meaning (table 1). indeed, cells with observed values below the lower confidence limits of random expectations are those with smaller similarity to others, or in other words, inhabited with species of more restricted ranges. this can be regarded as cells with a species’ composition “more endemic.” those above the upper limits of random expectations indicate cells that share more species with others than expected randomly. this is, cells with a species’ composition shared widely with others, or less “endemic.” that the region indistinguishable from random appears to separate geographical regions of higher than random similarity from those of lower than random similarity is quite interesting. in other explorations (not presented), we also found that geographic regions with dispersion field values above or below random expectations are bordered or separated by the areas with values indistinguishable from random. we believe that our method would be useful to suggest transitional zones in terms of levels of endemism, but this requires further empirical exploration. for a given cell, the value of its dispersion field is positively correlated with the number of species in the cell, since it is calculated as the sum of the ranges of all the species in the cell (thus more species, more terms in the sum). therefore for a given matrix, figure 5. regions in the new range-diversity plot and their representations in the geography. there is a consistent relationship between clusters of cells in the diagram and geographic regions. blocks of points highlighted in the diagram were selected randomly. biodiversity informatics, 16, 2021, pp. 20-27 26 it is possible to have a large value of the dispersion field (relative to all in the matrix) that nevertheless is significantly low, for cells (rows in the matrix) with many species; and the opposite, relatively low values of the dispersion field can be significantly high, in cells with a small number of species. in other words, high values of the dispersion field can be identified to be below random expectations (therefore highly endemic) because the high value of this index is directly related to a high species richness in that particular cell, and vice versa (table 1, fig. 4). this last point is of special interest as it suggests that interpretations of low or high values of the dispersion should be more appropriate if accompanied by considerations of species richness patterns. we anticipate that the plot we proposed here, and the geographic projection of its results can aid in promoting better interpretations of patterns of endemism in a region. we showed that, for our data, the partition of dispersion fields in the upper, middle, and lower quantile ranges has a clear geographic interpretation. this is a fact first noticed by soberón & ceballos (2011) and used by villalobos et al. (2013a). the geographic pattern of significant values is dependent on the region of interest, as species composition changes and such changes have different patterns in distinct areas of the planet. further explorations using the range-diversity plot and its projections will provide more information on how the patterns of endemism are rescued from pams in distinct regions of the world (soberón and ceballos 2011), or in different parts of a phylogeny (villalobos et al. 2017). for instance, the geographic projection of points with high values of both indices in the range-diversity plot (fig. 5) corresponds with areas that have been described to harbor a large number of species, which ranges are among the most restricted of mexican avifauna (villalobos et al. 2013b). with the formulation of the new range-diversity plots, we also introduce software to prepare such figures and to randomize pams. this implementation is parallelized, and depending on the features of the computer allows randomization of large matrices, with millions of elements (i.e., thousands of cells and species). we anticipate that the software provided will help in the exploration of the effects of randomizing pams and the implications of the results that can be obtained for different regions of the planet. as in any other analysis, the quality of input data is one of the most important factors to consider. researchers should acknowledge limitations deriving from data quality and determine appropriate features of the pams to be created. selecting appropriate cell sizes in the pam, for instance, can help to avoid over-interpretations of resulting range-diversity plots. acknowledgments the authors would like to thank the comisión nacional para el conocimiento y uso de la biodiversidad (conabio), especially raul sierra, for facilitating access to data. fabricio villalobos and an anonymous reviewer gave useful suggestions. the authors declare that no competing interests exist. data accessibility statement all data and code used in this research are openly available at https://github.com/jsoberon/pams-mexico references arita, h. t., j. a. christen, p. rodríguez, and j. soberón. 2008. species diversity and distribution in presence‐ absence matrices: mathematical relationships and biological implications. american naturalist 172:519– 532. borregaard, m. k., and c. rahbek. 2010. dispersion fields, diversity fields and null models: uniting range sizes and species richness. ecography 33:402–407. brownrigg, r., t. p. minka, a. deckmyn, r. a. becker, and a. r. wilks. 2018. maps: draw geographical place in the null distribution low intermediate high below 2.5% ci of random expectations high endemism low richness high endemism intermediate richness high endemism high richness within 95% of random expectations cannot say much about endemism cannot say much about endemism cannot say much about endemism above 97.5% ci of random expectations low endemism low richness low endemism intermediate richness low endemism high richness table 1. interpretation of dispersion field and significance values in the range-diversity plot. https://github.com/jsoberon/pams-mexico https://github.com/jsoberon/pams-mexico biodiversity informatics, 16, 2021, pp. 20-27 27 maps. r package. https://cran.r-project.org/web/packages/maps/index.html christen, j. a., and j. soberón. 2009. anidamiento y los ananálisis rq y qr en pam’s. micelánea matemática 49:51–61. connor, e. f., and d. simberloff. 1979. the assembly of species communities: chance or competition? ecology 60:1132–1140. gibert, c., and g. escarguel. 2019. per-simper—a new tool for inferring community assembly processes from taxon occurrences. global ecology and biogeography 28:374–385. gotelli, n. j. 2000. null model analysis of species co-occurrence patterns. ecology 81:2606–2621. graves, g. r., and c. rahbek. 2005. source pool geometry and the assembly of continental avifaunas. pnas 102:7871–7876. iucn. 2020. the iucn red list of threatened species. version 2020-6.2. kembel, s. w., p. d. cowan, m. r. helmus, w. k. cornwell, h. morlon, d. d. ackerly, s. p. blomberg, and c. o. webb. 2010. picante: r tools for integrating phylogenies and ecology. bioinformatics 26:1463– 1464. miklós, i., and j. podani. 2004. randomization of presence-absence matrices: comments and new algorithms. ecology 85:86–92. nuñez-penichet, c., m. e. cobos, a. t. peterson, j. soberón, n. barve, v. barve, and t. gueta. 2020. biosurvey: tools for biological survey planning. r package. https://github.com/claununez/biosurvey. r core team. 2020. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. sarkar, s. 2002. defining “biodiversity”; assessing biodiversity. monist 85:131–155. schluter, d. 1984. a variance test for detecting species associations, with some example applications. ecology 65:998–1005. soberón, j., and j. cavner. 2015. indices of biodiversity pattern based on presence-absence matrices: a gis implementation. biodiversity informatics 10:22–34. soberón, j., and g. ceballos. 2011. species richness and range size of the terrestrial mammals of the world: biological signal within mathematical constraints. plos one 6:e19359. strona, g., w. ulrich, and n. j. gotelli. 2018. bi-dimensional null model analysis of presence-absence binary matrices. ecology 99:103–115. ulrich, w., and n. j. gotelli. 2012. a null model algorithm for presence-absence matrices based on proportional resampling. ecological modelling 244:20–27. villalobos, f., r. dobrovolski, d. b. provete, and s. f. gouveia. 2013a. is rich and rare the common share? describing biodiversity patterns to inform conservation practices for south american anurans. plos one 8:e56073. villalobos, f., a. lira-noriega, j. soberón, and h. t. arita. 2013b. range-diversity plots for conservation assessments: using richness and rarity in priority setting. biological conservation 158:313–320. villalobos, f., a. lira-noriega, j. soberón, and h. t. arita. 2014. co-diversity and co-distribution in phyllostomid bats: evaluating the relative roles of climate and niche conservatism. basic and applied ecology 15:85–91. villalobos, f., m. á. olalla‐tárraga, m. v. cianciaruso, t. f. rangel, and j. a. f. diniz‐filho. 2017. global patterns of mammalian co-occurrence: phylogenetic and body size structure within species ranges. journal of biogeography 44:136–146. whittaker, r. h. 1960. vegetation of the siskiyou mountains, oregon and california. ecological monographs 30:279–338. zwiener, v. p., a. lira‐noriega, c. j. grady, a. a. padial, and j. r. s. vitule. 2018. climate change as a driver of biotic homogenization of woody plants in the atlantic forest. global ecology and biogeography 27:298–309. https://cran.r-project.org/web/packages/maps/index.html https://cran.r-project.org/web/packages/maps/index.html biodiversity informatics, 15, 2020, pp. 69-80 69 presence-only and presence-absence data for comparing species distribution modeling methods jane elith1*, catherine h. graham2, roozbeh valavi1, meinrad abegg2, caroline bruce3, simon ferrier4, andrew ford5, antoine guisan6, robert j. hijmans7, falk huettmann8, lucia lohmann9, bette loiselle10, craig moritz11, jake overton12, a. townsend peterson13, steven phillips14, karen richardson15, stephen e. williams16, susan k. wiser17, thomas wohlgemuth2, niklaus e. zimmermann2 1school of biosciences, university of melbourne, australia. 2swiss federal research institute wsl, ch-8903 birmensdorf, switzerland. 3csiro land and water, cairns, queensland, australia. 4csiro land and water, canberra, australian capital territory (act), australia. 5csiro land and water, tropical forest research centre, atherton, queensland, australia. 6university of lausanne, 1015 lausanne, switzerland. 7university of california, davis, usa. 8ewhale lab, institute of arctic biology, biology & wildlife department, university of alaska fairbanks, fairbanks alaska 99775 usa. 9universidade de são paulo, brazil. 10college of agricultural and life sciences, university of florida, usa. 11research school of biology & center for biodiversity analysis, australian national university, australia. 12manaaki whenua—landcare research, hamilton, new zealand (current address: panthera, floor 18, 8 west 40 st, new york, usa 10018. 13biodiversity institute, university of kansas, lawrence, kansas 66045, usa. 14center for biodiversity and conservation, american museum of natural history, new york, usa. 15department of geography, planning and environment, concordia university, montreal, canada. 16centre for tropical environmental and sustainability science, james cook university, townsville, australia. 17manaaki whenua—landcare research, lincoln, new zealand *corresponding author: j.elith@unimelb.edu.au abstract. species distribution models (sdms) are widely used to predict and study distributions of species. many different modeling methods and associated algorithms are used and continue to emerge. it is important to understand how different approaches perform, particularly when applied to species occurrence records that were not gathered in structured surveys (e.g. opportunistic records). this need motivated a large-scale, collaborative effort, published in 2006, that aimed to create objective comparisons of algorithm performance. as a benchmark, and to facilitate future comparisons of approaches, here we publish that dataset: point location records for 226 anonymised species from six regions of the world, with accompanying predictor variables in raster (grid) and point formats. a particularly interesting characteristic of this dataset is that independent presence-absence survey data are available for evaluation alongside the presence-only species occurrence data intended for modeling. the dataset is available on open science framework and as an r package and can be used as a benchmark for modeling approaches and for testing new ways to evaluate the accuracy of sdms. from 2002 to 2005 a working group funded by the united states’ national center for ecological analysis and synthesis (nceas), and led by atp and cm, compared methods for fitting species distribution models (sdms). these models combine observations of species occurrence or abundance with environmental data and can be used to predict distributions across space and time. the authors of this current paper are the subset of the nceas working group who gathered and processed the data described here, alongside suppliers of those data; referred to here as “the nceas data group.” the data come from six regions of the world (fig. 1). for each region we gathered two types of species occurrence data: presence-only (po, also known as “collection”) data, and presence-absence (pa, or “survey”) data. we generated random locations for each study region, referred to as “background” (or elsewhere “pseudo-absence”) samples. these are necessary for model fitting for some modeling methods. we compiled spatially continuous environmental predictor variables (“raster data”) deemed relevant to the species, and sampled these rasters at all po, pa and background (bg) locations. the nceas working group designed a “baseline” study to compare 16 modeling algorithms (elith mailto:j.elith@unimelb.edu.au jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 70 et al. 2006), and also several experimental treatments that manipulated the datasets to explore the effects of sample size (wisz et al. 2008), spatial resolution (grain) of environmental data (guisan et al. 2007), error in po location (graham et al. 2008), bias in records (dudik and phillips 2009; phillips et al. 2009) and treatment of bg data (phillips et al. 2009) on model performance. models were fitted (trained) on po and optionally bg data. the environmental data for pa sites were provided, so modelers could predict environmental suitability for all species at these sites. in all studies, models were fitted by working group members with expertise in the respective methods, and modeling was blind to the species presence or absence at evaluation sites. the modelers sent their predictions to one group member who evaluated them against the pa observations at those sites. in subsequent years thirteen additional papers were published (detailed in supplementary information 1,1 exploring aspects of the data or of model performance. the nceas data group obtained and compiled the data 1http://hdl.handle.net/1808/30579 in 2002. in support of recent trends to make science more transparent and repeatable (national academy of sciences et al. 2009; zuckerberg et al. 2010; garzon-lopez et al. 2016; munafò et al. 2017) we have now obtained permissions to publish the data. these data are valuable for their spread across regions of the world and across species, and particularly for the complementary sets of po and pa species data. here we expand on the latter point. whilst sdms can be fitted to a range of data types, the use of po and presence-background (p-bg) are common (kissling et al. 2018). fit and evaluation of sdms is often achieved through cross-validation; that is, using subsets of the data iteratively for either training and fitting the model or for testing (evaluation) (hastie et al. 2009; hijmans 2012; roberts et al. 2017). a problem with po data is that they can be spatially biased with some areas sampled intensively and others not at all (reddy and dávalos 2003; hortal et al. 2008; amano and sutherland 2013; isaac and pocock 2015) and thus, may not be representative of the species distribution in the study area. when evaluating with po or p-bg data, such biases remain, thus figure 1: overview of data supplied & model workflow. numbers on map refer to number of species, and icons represent taxa (birds, trees, other plants, reptiles, bats), from each region indicated. the workflow at the bottom illustrates supply of presence-only species data with accompanying environmental covariates for modeling, and presence-absence (1/0) data at different sites, for evaluation. http://hdl.handle.net/1808/30579 jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 71 wrongly emphasising the suitability of some environments and under-reporting the suitability of others. a model trained and evaluated with biased po or p-bg data may appear to perform well, although it does not produce meaningful predictions of the true species distribution (el-gabbas and dormann 2018). whilst methods are available for dealing with these sorts of bias in model training (phillips et al. 2009; syfert et al. 2013; warton et al. 2013; dorazio 2014; fithian et al. 2015; stolar and nielsen 2015; qiao et al. 2017), none are problem-free, and sdm evaluation remains a challenge. although the spatial distribution of pa locations can also be biased, the bias is less problematic because a higher or lower spatial density of sites in some areas simply leads to a more or less precise undestanding of the species distribution (assuming that at least some sites exist across the major environmental gradients) (phillips et al. 2009). pa evaluation data are therefore very useful. they can allow for an independent and less biased view of whether models correctly predict species occurrences. they also can be used to calculate a broader suite of evaluation statistics than those that can be estimated from po data (lawson et al. 2014). whilst pa data are desirable for evaluation (el-gabbas and dormann 2018) they are often not available, so the data supplied here—with both po and pa data from independent sources are a valuable resource for testing new modeling methods and evaluation approaches. gathering and processing the data here we focus on the data used by the studies produced by the nceas working group (supplementary information 1). we are unable to release the original records supplied to us, and instead are releasing the cleaned version used in our modeling. we have permission to release anonymized species labels rather than real species names; this will not detract from the dataset as a benchmark resource for reproducing previous results or for assessing other aspects of fitting and evaluating sdms. some of the data preparation methods were reported in the original baseline modeling paper (elith et al. 2006), but we describe them here in full detail, to gather all the information in one place, and to ensure the descriptions are adequate for data re-use. this manuscript and the accompanying metadata should be treated as the authoritative description of the data supplied here. data suppliers were initially asked to select species encompassing a range of life forms, responses to the environment, geographic distributions, and rarity, and to attempt to find species that had at least 20 records in both po and pa datasets. the limit of 20 was set so we had enough information for training and testing models (harrell 2001). this was generally adhered to, with some exceptions as evident in tables presented in supplementary information 2.2 (“summaries of species data for each region”). po and pa datasets were to be from different collection efforts, and not have sites in common. suppliers were asked to find a set of predictor variables in raster (grid) format that they considered relevant to the distribution of the species, and typical for what a skilled distribution modeler in their region would use. we asked for between ten and 15 variables to enable meaningful predictions and limit duplication across predictor variables, at the finest spatial resolution (smallest raster cell size) available, with a minimum acceptable resolution of 1 km2. the minimum grain reflects the finest grain of global climate data available at that time. all datasets were cleaned by je and cg to these common properties agreed to by the group: (a) all data projected to a common projection for that region; (b) all raster data for a region aligned to the same extent and resolution, and only rasters with close to complete coverage in the region of interest retained; (c) species records reduced to a maximum of one record per raster cell using the following protocol: for po data: if there is at least one presence record in a cell, retain one presence record for that cell; for pa data: reduce to one record per cell using the rule: if presence(s) and absence(s) both occur in the same cell, retain one presence; (d) records checked and rectified if necessary to ensure that po and pa locations do not co-occur in a grid cell; (e) species records from locations with no environmental data removed. many sdms contrast the environment at locations of known occurrence of a species to that at a set of random locations in the study region (background, quadrature, or pseudo-absence points: (phillips et al. 2009; warton and shepherd 2010; barbet-massin et al. 2012; renner et al. 2015). the nceas data group therefore supplied modelers with a sample of 10,000 background points for each region. regional extents were delineated by the boundaries of countries or bioregions within countries, as deemed appropriate by the data suppliers. background points were selected spatially at random across each region and sampled irrespective of the location of any presence records. that is, by design, a presence record and a 2http://hdl.handle.net/1808/30581. http://hdl.handle.net/1808/30581 jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 72 background point might occur in the same raster cell. this approach aligns with recent interpretations of background samples as an approach for fitting a point process (renner et al. 2015). characteristics of the gathered data we sourced datasets from six regions of the world (figure 1 and table 1); the regions are hereafter referred to by the initials provided in figure 1 and in column 1, table 1. the six regions vary in size from approximately 24 to 12,223 thousand km2 (table 1). we gathered 11 to 13 predictor variables per region, most of which were continuous variables, but with four out of the six regions providing one or two categorical variables in their predictor sets (summarised in table 1, and details of variables in supplementary information 3).3 records span species of birds (awt, can, nsw), bats (nsw), plants (awt, nsw, nz, sa, swi) and reptiles (nsw) (table 1), totaling 226 species. the regions show useful variation in the amount of species data per region, with tens to thousands of po records per species, and 102 to 19 120 pa evaluation sites (table 1 and detailed summaries per region in supplementary information 2). this provided a diverse and representative data set for the nceas studies (supplementary information 1), and a benchmark set that we anticipate being broadly useful into the future. data sources for the po and pa species data are detailed in table 2. the different data sources used different sampling designs and methods which can provide insights into how data quality influences model outcomes/accuracy. for instance, some po locations are recorded by gps and therefore likely accurate (see awt, table 2), whereas others are typical po data from a range of sources where the level of geographic accuracy is unknown (see nsw, table 2). all the pa data are from intentional surveys, but vary in their design, age and number of data collectors. for instance, the swi data are from plots on a regular lattice, the sa data are all collected by one person, and the can data are breeding bird data collected over years by multiple people. these variations are typical of what is seen in ecological datasets further making this dataset a useful benchmark for sdm modelers. details of data format and location the data are available on open science framework (osf) and accompanied by human metadata and/or readme files, and machine-accessible meta3http://hdl.handle.net/1808/30582 data; most data are also available in an r package, as described below. all data, organised as described below, are available openly.4 a. osf data in overview, we have uploaded the data in separate directories. the environmental raster (gridded) data are separate from all other records, since their zip file is 561 mb in total and many users will not want to predict to the rasters but rather to the tabled environmental data for the evaluation sites. the species data (both po and pa), and the background samples are in separate folders within a “records” folder and include the data extracted from these rasters (i.e. all environmental conditions) at every site, and thus are ready for modeling and evaluation. all site-based data (po, bg, pa) are available as comma-separated text files (.csv). polygon outlines of each region are also supplied to give context to locations of species records, if users want to map them without using rasters. next, we detail the data locations and formats for 5 subsets of data within the data folder. at any level of the folder organisation, data within a user-selected folder can be downloaded as a zip file. 1. environmental rasters within the /data/environment folder at the location above, rasters (~ 1gb total unzipped) are arranged in folders, one per region, and supplied as .tif files. in each region’s folder, a metadata file explains all known details for each variable. within the /data/environment folder, a readme. txt file adds authors responsible for data preparation, and details of coordinate reference systems, units and raster cell sizes. 2. presence-only data—locations and environmental samples within the /data/records/train_po folder at the location above, there is one .csv file per region containing records for all species, and each is accompanied by a metadata file providing details for each column. users will find that file formats are consistent across regions (table 1). 3. background data—locations and environmental samples ten thousand background (bg) samples are supplied for each region, as outlined in an earlier section. within the /data/records/train_bg folder at 4https://osf.io/kwc4v/ http://hdl.handle.net/1808/30582 https://osf.io/kwc4v/ jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 73 co de re gi on d et ai ls ar ea (‘0 00 k m 2 ) ar ea lo ca tio n – re d po ly go ns sh ow lo ca tio ns w ith in c ou nt rie s / c on tin en ts n o. e nv v ar s (n o. ca te go ric al ) ap pr ox . gr id c el l re so lu tio n (m ) bi ol og ic al g ro up s & n um be r sp ec ie s m ea n no . re co rd s p er sp ec ie s n o. sit es : pa po pa aw t au st ra lia n w et tr op ic s, q ue en sla nd , au st ra lia 23 .9 7 13 (0 ) 80 b : b ird s: 2 0 15 5 97 34 0 p: v as cu la r p la nt s: 2 0 35 30 10 2 ca n o nt ar io , c an ad a 97 9. 34 11 (1 ) 1 00 0 bi rd s: 3 0 25 3 1 28 2 14 5 71 n sw n or th -e as t n ew so ut h w al es , au st ra lia 76 .1 8 13 (1 ) 10 0 ba : b at s: 7 27 76 57 0 db : d iu rn al b ird s: 8 18 9 57 70 2 nb : n oc tu rn al b ird s: 2 13 4 14 2 1 13 7 ot : o pe nfo re st tr ee s: 8 42 16 4 2 07 5 ou : o pe nfo re st u nd er st or ey va sc ul ar p la nt s: 8 21 35 8 1 30 9 rt : r ai nf or es t t re es : 7 9 21 2 1 03 6 ru : r ai nf or es t u nd er st or ey va sc ul ar p la nt s: 6 18 93 90 9 sr : s m al l r ep til es : 8 84 63 1 00 8 ta bl e 1: s um m ar y of d at a av ai la bl e ac ro ss re gi on s. jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 74 ta bl e 1: s um m ar y of d at a av ai la bl e ac ro ss re gi on s ( co nt in ue d) . n z n ew z ea la nd 26 5. 41 13 (2 ) 10 0 va sc ul ar p la nt s: 5 2 59 1 80 1 19 1 20 sa co nt in en ta l br az il, e cu ad or , co lo m bi a, bo liv ia , a nd p er u, so ut h am er ic a 12 22 3. 17 11 (0 ) 1 00 0 va sc ul ar p la nt s: 3 0 74 12 15 2 sw i sw itz er la nd 39 .5 6 13 (1 ) 10 0 tr ee s: 3 0 1 17 0 81 0 10 0 13 jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 75 table 2: information on sources of po and pa data (initials refer to co-author names; all other acronyms defined elsewhere in text) po data pa data awt birds: supplied by sw. incidental surveys. gps locations therefore more accurate than many po data. plants: supplied by af & kr. herbarium data, cleaned and corrected by af. reliability codes for sites selected were 2 (15%), 3 (75%) and 4 (10%). this means that most records are accurate to within 3km and were collected within the 40 years prior to 2002. birds: supplied by sw. field surveys. gps locations. plants: af’s survey sites. these have accurate locations (within 0.1km) and were collected in the 20 years prior to 2002. can birds from the ontario nest records database, royal ontario museum (rom). supplied by m. peck to fh. temporal span 1870-2002 (usually 1960-2001). coordinates derived from map by rom; some locations ground-truthed with gps. from breeding bird atlas (bba) for ontario, provided by m. cadman to fh. nsw 8 biological groups supplied by sf. fauna data from the atlas of nsw wildlife (a database of incidental sighting records); flora data: specimen records from both the university of new england herbarium and the sydney herbarium (royal botanic gardens). no information on collection dates or accuracy. supplied by sf. from designed surveys described elsewhere (ferrier and watson 1996; pearce et al. 2001) nz plants, mostly trees and shrubs from indigenous forests. supplied by jo, sw. records from allan herbarium, managed by manaaki whenua - landcare research supplied by jo, sw. records from national vegetation survey databank (wiser et al. 2001), nvs.landcareresearch.co.nz. sa plant species from the family bignoniaceae. supplied by bl and ll. from the missouri botanical garden database management system tropicos (http://www.mobot.org) and lucia lohmann (lohmann@mobot.org). species localities were calculated by tropicos and by l. lohmann using the getty thesaurus of geographical names browser (http://shiva.pub.getty.edu). supplied by bl and ll. survey data collected by al gentry over 22 years (1971-1993). swi 30 tree species supplied by nez & tw from a forest vegetation data base containing 14 800 irregularly and non-systematically sampled forest vegetation relevés throughout switzerland. records start in 1904 and ends in 1995. the majority (95%) of the plots was collected after 1940, and ~60% of the data were sampled between 1960 and 1995. the individual authors had their own local sampling design or used preferential sampling techniques (details in wohlgemuth 2012). species cover estimation prevails as performance measure (98%) and follows the braun-blanquet approach (braun-blanquet 1964). the data is part of the european vegetation archive eva (chytrý et al. 2016). around 14,100 relevés were selected from the original data base, targeting minimal data standards such as coordinates, and species extracted from these. 30 tree species supplied by nez & ma data extracted from the swiss national forest inventory (nfi). pa data is collected on accessible sample plots at a regular 1 km point lattice across switzerland. the data originate from the first national inventory, collected 1983-1985. for details, see brassel and lischke (2001) and eafv (1988). jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 76 the location above, there is one .csv file per region, with columns exactly matching those for the po data. hence, po and bg files can easily be combined for modeling, as needed. one readme.txt file points to relevant files explaining the data for all regions and explains additional details specific to the background data setup. 4. presence-absence data pa data are intended for evaluation and supplied in two sets of files both at the location above: one (/data/records/test_env/) containing the sampled environments for predicting to, and one (/ data/records/test_pa/) with the species data, in identical site (row) order to the environmental data. this two-file format reflects our original use of the data, keeping the evaluation data “blind” to the modeling (see comments on usage in the final section). where data span more than one biological group (awt, nsw), there are multiple files, one for each group. in each of the two folders of .csv files, there is a readme.txt file that points to relevant files explaining the data for all regions and providing details particular to the evaluation data. 5. polygons of region extents polygons defining the extent (i.e. the borders) of each region are provided at the location above, in the /data/borders/ folder. b. r package datasets 2 to 5, above, are also available in an r package, “disdat.”5 in the future we intend to submit it to cran.6 the r package contains the data, func5https://github.com/rspatial/disdat/blob/master/readme.md. 6https://cran.r-project.org/ figure 2: examples of plots that can be achieved with functions supplied in the r package vignette. top left: maps of species data, top right: an interactive map with pa site locations; bottom left: density plot showing the distribution of po data along one environmental gradient, compared with that of random points from the region; bottom right: pairwise correlations between variable. https://github.com/rspatial/disdat/blob/master/readme.md https://cran.r-project.org/ jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 77 tions to call the data, help files describing the data and pointing to this paper and its metadata, and two vignettes to assist modelers in data use. discussion of the data and its usage before we used these data for modeling, extensive efforts were made to prepare the data in an appropriate format. in retrospect—and as a comment for future data preparation exercises—some steps might have been done differently. for example, instead of removing duplicate records we could have marked records as “selected” and others as, for example, “not selected (duplicate in raster cell)”. in addition, we reduced the species records to one per cell across regions with varying cell sizes. instead across all regions a minimum distance between sites could have been set and applied to all. however, our cleaning is a common approach and not essentially flawed, so the data are still highly valuable. we no longer hold the original data, and our intention is to publish the dataset as used by publications listed in supplementary information 1, allowing comparison with these previous influential modeling efforts. we note that the environmental data in the species files were extracted from rasters that we are also releasing. this extraction was done almost two decades ago. when we now, with current software, extract environmental data at those same site locations we find some very minor discrepancies in some datasets. these will likely have negligible effects on models. nevertheless, we supply the values in the species files ‘as is’, because these are the data used for modeling in the most well-cited output of the nceas working group (elith et al. 2006). much can be learned both about different methods and about properties of data by iterative modeling and evaluation. readers can find analyses conducted to date on these data, in the publications summarized in supplementary information 1. one example of what we learned addresses bias in the species data. in 2002 there were very few published explorations of datasets like these, and their biases. this contrasts with the many published explorations and methods available now for handling bias in presence-only data (kadmon et al. 2004; phillips et al. 2009; kramer-schadt et al. 2013; syfert et al. 2013; warton et al. 2013; bird et al. 2014; boria et al. 2014; dorazio 2014; fithian et al. 2015; stolar and nielsen 2015; qiao et al. 2017). the nceas group first explored whether the cleaned data could reasonably be used to predict species occurrence, without any bias treatment. the nceas modeling group demonstrated (elith et al. 2006) that in some cases predictions had reasonable to very good accuracy, but that some regions’ datasets were clearly hampered by bias. this led to subsequent work, particularly that of phillips and co-authors (phillips et al. 2009) who explored the extent and impact of bias in these data, and presented and tested the “target-group” approach for dealing with bias. in other words, our understanding of the problem of bias developed from our first iteration of modeling and our analyses of the outputs. we look forward to future insights gained from working with these data. in the r package and associated vignettes, we present methods for exploring the supplied data to give insight into their properties. in the r package we provide a function for mapping all species, producing a “map book” of all po and pa data for all regions. in the data visualisation vignette we provide code for mapping any given species (po, pa data on a static map, plus an interactive map linked to satellite image data). in that vignette we also provide functions for exploring the distribution of sites in geographic and environmental space, and for analysing pairwise correlations between variables. figure 2 illustrates some of the outputs that can be produced with these functions. our data are available on on osf and github as detailed in earlier, and are easy to download. we kindly request that each user (even students within teaching exercises) download the data or r package individually because some data providers would like to track data downloads, to enable reporting on data usage as required by their funding agencies. part of the value of this dataset is that independent pa data are available for evaluation of models fitted with po data. in the publications shown in the table in supplementary information 1, the pa evaluation (test) data were kept independent as a “blind evaluation” set, that is, they were not used to tune models. in line with this setup, we have supplied the data in two distinct sets: 1) the environmental conditions at each evaluation site, which enables predictions to be made to the sites; 2) the actual pa observations at the evaluation sites for evaluation after modeling is complete. to facilitate future comparative research, we encourage users to provide clear documentation if they choose to use a different setup—for instance, tuning their models on some or all of the evaluation data. jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 78 whilst we provide 10,000 background points for each region, recent research has shown that in some regions larger background samples may be required to sufficiently represent all environments (renner et al. 2015). we supply the background data used in the main output of the nceas working group (elith et al. 2006), though note that such datasets are easy to re-create by sampling the environmental raster data that we supply. they also may be sampled according to designs other than spatially at random. the use of random background data is justified if the po occurrence data are a representative sample of the species’ distribution (phillips et al. 2009; elith et al. 2011), a condition that is likely not satisfied for many species and regions. one alternative approach is to create and use a target group background sample (tgb; phillips et al. 2009) where the target group contains all species in the identified group (e.g. all birds), including the species to be modeled. depending on how the po data are treated, the tgb sample might also need to be reduced to one sample per unique location. the “modeling nceas data” vignette in the r package includes example code for making a tgb sample. to replicate models in the main output of the nceas working group (elith et al. 2006), users should read both the full paper and the associated appendix (freely available7). since we were at that stage aiming to reflect expert use of models, different authors implemented different methods, starting with the data supplied here. as can be seen in the details recorded in that manuscript and its appendix, some modelers chose to use all predictors (with or without automated variable selection) whereas others chose a subset of the predictors based on pairwise correlations. specific settings for each method are also documented in the appendix of elith et al. 2006. whilst code is not available to reproduce those models, in future work we plan to reproduce several of the models and provide fully documented r code for those algorithms, for reproducibility. in the meantime, and in order to help less experienced modelers, we also provide example code in the r package vignette “modeling nceas data” for using the data for modeling and evaluation, applying one method across all species. acknowledgments the working group “testing alternative methodologies for modeling species’ ecological niches 7http://www.ecography.org/sites/ecography.org/files/appendix/e4596.pdf and predicting geographic distributions” (project id: 4980) was central to this project, funded and hosted by the national centre for ecological analysis and synthesis, santa barbara, california. competing interests the authors have declared that no competing interests exist. references amano, t., and w. j. sutherland. 2013. four barriers to the global understanding of biodiversity conservation: wealth, language, geographical location and security. proc. r. soc. b biol. sci. 280. barbet-massin, m., f. jiguet, c. h. albert, and w. thuiller. 2012. selecting pseudo-absences for species distribution models: how, where and how many? methods ecol. evol. 3:327–338. bird, t. j., a. e. bates, j. s. lefcheck, n. a. hill, r. j. thomson, g. j. edgar, r. d. stuart-smith, s. wotherspoon, m. krkosek, j. f. stuart-smith, g. t. pecl, n. barrett, and s. frusher. 2014. statistical solutions for error and bias in global citizen science datasets. biol. conserv. 173:144–154. boria, r. a., l. e. olson, s. m. goodman, and r. p. anderson. 2014. spatial filtering to reduce sampling bias can improve the performance of ecological niche models. ecol. model. 275:73–77. brassel, p., and h. lischke. 2001. swiss national forest inventory: methods and models of the second assessment. swiss federal institute wsl, birmensdorf. braun-blanquet, j. 1964. pflanzensoziologie. grundzüge der vegetationskunde. springer-verlag wein, new york. chytrý, m., s. m. hennekens, b. jiménez‐alfaro, i. knollová, j. dengler, f. jansen, f. landucci, j. h. j. schaminée, s. aćić, e. agrillo, d. ambarlı, p. angelini, i. apostolova, f. attorre, c. berg, e. bergmeier, i. biurrun, z. botta‐dukát, h. brisse, j. a. campos, l. carlón, a. čarni, l. casella, j. csiky, r. ćušterevska, z. d. stevanović, j. danihelka, e. d. bie, p. de ruffray, m. d. sanctis, w. b. dickoré, p. dimopoulos, d. dubyna, t. dziuba, r. ejrnæs, n. ermakov, j. ewald, g. fanelli, f. fernández‐gonzález, ú. fitzpatrick, x. font, i. garcía‐mijangos, r. g. gavilán, v. golub, r. guarino, r. haveman, a. indreica, d. i. gürsoy, u. jandt, j. a. m. janssen, m. jiroušek, z. kącki, a. kavgacı, m. kleikamp, v. kolomiychuk, m. k. ćuk, d. krstonošić, a. kuzemko, j. lenoir, t. lysenko, c. marcenò, v. martynenko, d. michalcová, j. e. moeslund, v. onyshchenko, h. pedashenko, a. pérez‐ http://www.ecography.org/sites/ecography.org/files/appendix/e4596.pdf jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 79 haase, t. peterka, v. prokhorov, v. rašomavičius, m. p. rodríguez‐rojo, j. s. rodwell, t. rogova, e. ruprecht, s. rūsiņa, g. seidler, j. šibík, u. šilc, ž. škvorc, d. sopotlieva, z. stančić, j.-c. svenning, g. swacha, i. tsiripidis, p. d. turtureanu, e. uğurlu, d. uogintas, m. valachovič, y. vashenyak, k. vassilev, r. venanzoni, r. virtanen, l. weekes, w. willner, t. wohlgemuth, and s. yamalov. 2016. european vegetation archive (eva): an integrated database of european vegetation plots. appl. veg. sci. 19:173–180. dorazio, r. m. 2014. accounting for imperfect detection and survey bias in statistical analysis of presence-only data. glob. ecol. biogeogr. 12:1472–1484. dudik, m., and s. j. phillips. 2009. generative and discriminative learning with unknown labeling bias. pp. 401–408 in advances in neural information processing systems 21 (d. koller, d. schuurmans, y. bengio, and l. bottou, eds.). neural information processing systems 2008, neural information processing systems foundation, inc. eafv. 1988. schweizerisches landesforstinventar: ergebnisse der erstaufnahme 1982-1986. el-gabbas, a., and c. f. dormann. 2018. improved species-occurrence predictions in data-poor regions: using large-scale data and bias correction with down-weighted poisson regression and maxent. ecography 41: 1161–1172. elith, j., c. h. graham, r. p. anderson, m. dudík, s. ferrier, a. guisan, r. j. hijmans, f. huettmann, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. mcc. overton, a. t. peterson, s. j. phillips, k. s. richardson, r. scachetti-pereira, r. e. schapire, j. soberón, s. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species’ distributions from occurrence data. ecography 29:129–151. elith, j., s. j. phillips, t. hastie, m. dudík, y. e. chee, and c. j. yates. 2011. a statistical explanation of maxent for ecologists. divers. distrib. 17:43–57. ferrier, s., and g. watson. 1996. an evaluation of the effectiveness of environmental surrogates and modelling techniques in predicting the distribution of biological diversity. consultancy report prepared by the nsw national parks and wildlife service for department of environment, sport and territories, environment australia, canberra. fithian, w., j. elith, t. hastie, and d. keith. 2015. bias correction in species distribution models: pooling survey and collection data for multiple species. methods ecol. evol. 6:424–438. garzon-lopez, c. x., l. bastin, g. m. foody, and d. rocchini. 2016. a virtual species set for robust and reproducible species distribution modelling tests. data brief 7:476–479. graham, c. h., j. elith, r. j. hijmans, a. guisan, a. t. peterson, and b. a. loiselle. 2008. the influence of spatial errors in species occurrence data used in distribution models. j. appl. ecol. 45:239–247. guisan, a., c. h. graham, j. elith, f. huettmann, and . nceas species distribution modelling group. 2007. sensitivity of predictive species distribution models to change in grain size: insights from a multi-models experiment across five continents. divers. distrib. 13:332–340. harrell, f. e. 2001. regression modeling strategies with applications to linear models, logistic regression and survival analysis. springer verlag, new york. hastie, t., r. tibshirani, and j. h. friedman. 2009. the elements of statistical learning: data mining, inference, and prediction, second edition. 2nd ed. springer-verlag, new york. hijmans, r. j. 2012. cross-validation of species distribution models: removing spatial sorting bias and calibration with a null model. ecology 93:679–688. hortal, j., a. jiménez-valverde, j. f. gómez, j. m. lobo, and a. baselga. 2008. historical bias in biodiversity inventories affects the observed environmental niche of the species. oikos 117:847–858. isaac, n. j. b., and m. j. o. pocock. 2015. bias and information in biological records. biol. j. linn. soc. 115:522–531. kadmon, r., o. farber, and a. danin. 2004. effect of roadside bias on the accuracy of predictive maps produced by bioclimatic models. ecol. appl. 14:401–413. kissling, w. d., j. a. ahumada, a. bowser, m. fernandez, n. fernández, e. a. garcía, r. p. guralnick, n. j. b. isaac, s. kelling, w. los, l. mcrae, j.-b. mihoub, m. obst, m. santamaria, a. k. skidmore, k. j. williams, d. agosti, d. amariles, c. arvanitidis, l. bastin, f. de leo, w. egloff, j. elith, d. hobern, d. martin, h. m. pereira, g. pesole, j. peterseil, h. saarenmaa, d. schigel, d. s. schmeller, n. segata, e. turak, p. f. uhlir, b. wee, and a. r. hardisty. 2018. building essential biodiversity variables (ebvs) of species distribution and abundance at a global scale. biol. rev. 93:600–625. kramer-schadt, s., j. niedballa, j. d. pilgrim, b. schröder, j. lindenborn, v. reinfelder, m. stillfried, i. heckmann, a. k. scharf, d. m. augeri, s. m. cheyne, a. j. hearn, j. ross, d. w. macdonald, j. mathai, j. eaton, a. j. marshall, g. semiadi, r. rustam, h. bernard, jane elith et al. – presence-only and presence-absence data for comparing species distribution modeling methods 80 r. alfred, h. samejima, j. w. duckworth, c. breitenmoser-wuersten, j. l. belant, h. hofer, and a. wilting. 2013. the importance of correcting for sampling bias in maxent species distribution models. divers. distrib. 19:1366–1379. lawson, c. r., j. a. hodgson, r. j. wilson, and s. a. richards. 2014. prevalence, thresholds and the performance of presence–absence models. methods ecol. evol. 5:54–64. munafò, m. r., b. a. nosek, d. v. m. bishop, k. s. button, c. d. chambers, n. percie du sert, u. simonsohn, e.-j. wagenmakers, j. j. ware, and j. p. a. ioannidis. 2017. a manifesto for reproducible science. nat. hum. behav. 1:0021. national academy of sciences, national academy of engineering, and institute of medicine. 2009. ensuring the integrity, accessibility, and stewardship of research data in the digital age. the national academies press, washington, dc. pearce, j. l., k. cherry, m. drielsma, s. ferrier, and g. whish. 2001. incorporating expert knowledge and fine-scale vegetation mapping into statistical modelling of faunal distribution. j. appl. ecol. 38:412–424. phillips, s. j., m. dudík, j. elith, c. h. graham, a. lehmann, j. leathwick, and s. ferrier. 2009. sample selection bias and presence-only distribution models: implications for background and pseudo-absence data. ecol. appl. 19:181–197. qiao, h., a. t. peterson, l. ji, and j. hu. 2017. using data from related species to overcome spatial sampling bias and associated limitations in ecological niche modelling. methods ecol. evol. 8:1804–1812. reddy, s., and l. m. dávalos. 2003. geographical sampling bias and its implications for conservation priorities in africa. j. biogeogr. 30:1719–1727. renner, i. w., j. elith, a. baddeley, w. fithian, t. hastie, s. j. phillips, g. popovic, and d. i. warton. 2015. point process models for presence-only analysis. methods ecol. evol. 6:366–379. roberts, d. w., v. bahn, s. ciuti, m. s. boyce, j. elith, g. guillera-arroita, s. hauenstein, j. j. lahoz-monfort, b. schroder, w. thuiller, d. warton, b. a. wintle, f. hartig, and c. f. dormann. 2017. cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. ecography 40:913–929. stolar, j., and s. e. nielsen. 2015. accounting for spatially biased sampling effort in presence-only species distribution modelling. divers. distrib. 21:595–608. syfert, m. m., m. j. smith, and d. a. coomes. 2013. the effects of sampling bias and model complexity on the predictive performance of maxent species distribution models. plos one 8:e55158. warton, d. i., i. w. renner, and d. ramp. 2013. model-based control of observer bias for the analysis of presence-only data in ecology. plos one 8:e79168. warton, d. i., and l. c. shepherd. 2010. poisson point process models solve the “pseudo-absence problem” for presence-only data in ecology. ann. appl. stat. 1383–1402. wiser, s. k., p. j. bellingham, and l. e. burrows. 2001. managing biodiversity information: development of new zealand’s national vegetation survey databank. n. z. j. ecol. 25:1–14. wisz, m. s., r. j. hijmans, j. li, a. t. peterson, c. h. graham, a. guisan, and nceas predicting species distributions working group. 2008. effects of sample size on the performance of species distribution models. divers. distrib. 14:763–773. wohlgemuth, t. 2012. swiss forest vegetation database. pp. 340–340 in dengler j., j. oldeland, f. jansen, m. chytry, j. ewald, m. fickh, f. glöckler, g. glopez-gonzalez, r. k. peet, and j. h. j. schaminée. (editors) vegetation databases for the 21st century. zuckerberg, b., f. huettman, and j. friar. 2010. chapter 3: proper data management as a scientific foundation for reliable species distribution modeling. pp. 45–70 in predictive species and habitat modeling in landscape ecology. springer new york. biodiversity informatics, 10, 2015, 65-74 65 biodiversity informatics training curriculum, version 1.2 a. townsend peterson and kate ingenloff biodiversity institute, university of kansas, lawrence, kansas 66045 usa abstract.—biodiversity informatics as a field lacks a synthetic, comprehensive summary, guide, and training resource, such as a textbook, a gap that impedes advance in the field. the biodiversity informatics training curriculum (bitc) represents the compilation of digital videos and accessory materials from 3 years of courses held across africa, on topics covering the full thematic breadth of the field of biodiversity informatics. instructors include experts on each topic, such that the instruction is at a high level. the bitc is presented as an open-access, fully digital, video-based teaching resource—in essence a first textbook for the field. the curriculum is constantly being updated and improved, but this version 1.2 represents a first version that covers all major subfields in biodiversity informatics. links to videos hosted on youtube are provided herein, but versions that do not require internet access are available upon request. key words.—biodiversity, informatics, teaching, digital video biodiversity informatics is, simply put, a new field (schalk 1998; oecd megascience forum 1999; soberón 1999; bisby 2000). informatics more generally is the science of creating, organizing, and analyzing digital data. although biodiversity scientists (systematics, taxonomists, ecologists, biogeographers) have used informatics processes in studies of biodiversity for centuries, the digital dimensions of the challenge are relatively new. indeed, it is only in the last decade or so that large-scale biodiversity databases have become broadly available and accessible. the major steps of innovation and advancement of the field—in the view of many—will derive from novel but meaningful linkages between disparate realms of information about biodiversity (peterson et al. 2010). given the newness of the field, biodiversity informatics remains largely without any formal synthesis, guide, textbook, or summary. that is, no textbook entitled something like “biodiversity informatics” exists to guide users through the steps involved in biodiversity informatics applications. further, no graduate programs exist that provide a comprehensive curriculum for the field. clearly, books and programs exist treating portions of the field (e.g., ecological niche modeling, peterson et al. 2011), but no overarching summary has been developed. the biodiversity informatics training curriculum (bitc) represents an effort to remedy this problem and fill the gap in availability of knowledge and experience in the field. with the generous support of the jrs biodiversity foundation, a series of detailed training courses has been held at sites across africa, each treating a different topic within the broader area of biodiversity informatics, with experts brought in from around the world, and trainees from across africa. full summaries of the courses, and the experts and trainees who attended, are available on the bitc website1. this document is designed to provide a summary and index to the first nearly-complete version of the full curriculum. each topic in the long list in the appendix is a hyperlink to a youtube video or playlist. the topics are more or less in a logical order, although some topics really belong in multiple categories: for example, georeferencing could go under data cleaning or data capture, or as introductory material in ecological niche modeling. above all, we note that this is a first more or less complete version, so we invite suggestions, criticisms, ideas, and revisions. if youtube does not work for you… very early in the history of the bitc initiative, the decision was made to house the video component of each course on youtube. this decision was made largely in light of the nearglobal availability and massive usership of that site. however, we are quite cognizant that youtube is not a universal solution to providing access to video content, as access to the site is prohibited by some institutions—and even some 1 http://biodiversity-informatics-training.org/. http://biodiversity-informatics-training.org/ biodiversity informatics, 10, 2015, 65-74 66 countries. furthermore, additional limitations may be imposed by infrastructural limitations (e.g., insufficient bandwidth) that may limit or impede long views of youtube-based movies. as such, we are seeking actively to provide multiple access solutions to bitc materials. as a more failsafe solution to the access challenge, we have ‘published’ the entirety of the bitc video set and ancillary materials on 32gb usb keys. of course, the challenge now becomes one of availability of this physical copy. for the moment, however, we offer to send the usb key wherever feasible by mail or courier; requests can be addressed to biodivtraining@gmail.com. help us improve the bitc subtitling many users of bitc materials have expressed frustration in the fact that, apart from a few courses that have been replicated in spanish and portuguese, the entire body of materials and resources is in english. in africa, where the bitc in-person courses were held, this english focus was a significant complication for participation of trainees from francophone countries. this more general failing in the design and content of the curriculum simply reflects the failings of the two authors of this summary in not having significant language abilities in french, arabic, chinese, russian, etc. as a step toward a solution to this shortcoming, we are exploring an automated process of adding subtitles to curriculum materials. we have laid out a protocol that involves the use of voicerecognition software to provide an initial template of english subtitles, which (of course) is qualitycontrolled by native english-speaking users. the english subtitles will then be placed on a crowdsourcing website (e.g., dotsub2), where users worldwide can help facilitate the development and improvement of subtitling into other languages. this effort is as-yet in its infancy, but has proven effective in preliminary tests involving translation into arabic, chinese, spanish, french, and portuguese. opportunities to assist in the development of these online resources in multiple languages will be announced via the bitc facebook group3. 2 http://dotsub.com. 3 https://www.facebook.com/groups/biodiversityinformatics/. problems assembling the online curriculum has been a long and difficult process, involving experimentation with multiple video recording and processing technologies, and considerable evolution in the management of video data. as such, it is entirely possible—indeed probable—that problems and errors have crept into this outcome. we request kindly that any problems that users note be reported to us at biodivtraining@gmail.com. we will make every effort to attend to each and every problem that is reported, at least to the extent that is possible. future plans the bitc content is being updated and expanded by a variety of means, particularly via a monthly online seminar series in biodiversity informatics4 (e.g., a series on the history of biodiversity informatics currently in process). one important next step that we are exploring is that of certifying the use and mastery of bitc materials by students and young trainees around the world. we are exploring the idea of a masters degree program based on bitc online materials, which would allow trainees to study the material online, develop a project or prepare for an examination, and then be examined by a panel of experts. the degrees would be granted by host universities around the developing world. this model is as-yet under exploration and development, and we expect a fully functioning prototype to be in place by late 2016. summary the bitc represents years of collective work by dozens of persons, including the two of us as co-directors, but perhaps more significantly by our expert course instructors, and more than 134 student trainees representing 23 countries across africa. the digital teaching resources that have been assembled are unique, in that no other comprehensive teaching resource exists for the field. needless to say, we are eager to see bitc usership continue to grow and expand. acknowledgments the bitc exists thanks to generous contributions of many instructors, including tanya abrahamse, 4 http://biodiversity-informatics-training.org/webinar-series/. mailto:biodivtraining@gmail.com http://dotsub.com/ https://www.facebook.com/groups/biodiversityinformatics/ mailto:biodivtraining@gmail.com http://biodiversity-informatics-training.org/webinar-series/ biodiversity informatics, 10, 2015, 65-74 67 arturo ariño, david blackburn, kyle braak, rafe brown, bilal butt, lindsay campbell, jacob cooper, vanderlei perez canhos, firkirte gebresenbet erda, eric fokam, lee hannah, leonard krishtalka, rafael loyola, enrique martínez-meyer, adolfo navarro-sigüenza, monica papeş, richard pearson, thiago rangel, mark robbins, laura russell, moses nsanyi sainge, jorge soberón, javier otegui tellechea, melissa tulig, kumara wakjira, kimberly watson, christiane weirauch, john wieczorek, and selwyn willoughby. funding was generously provided by the jrs biodiversity foundation literature cited bisby, f. a. 2000. the quiet revolution: biodiversity informatics and the internet. science 289:23092312. oecd megascience forum. 1999. final report of the working group on biological informatics. organisation for economic co-operation and development, paris. peterson, a. t., s. knapp, r. guralnick, j. soberón, and m. t. holder. 2010. the big questions for biodiversity informatics. systematics and biodiversity 8:159-168. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. schalk, p. h. 1998. management of marine natural resources through by biodiversity informatics. marine policy 22:269-280. soberón, j. 1999. linking biodiversity information sources. trends in ecology and evolution 14:291. biodiversity informatics, 10, 2015, 65-74 68 appendix biodiversity informatics training curriculum, version 1.2 introductory material  biodiversity informatics  introduction to the bitc  introduction to biodiversity informatics  what is biodiversity data?  building the biodiversity knowledge graph  writing scientific papers  introduction  choosing a journal  technology assists  words to avoid  figures  using color in science  tables  proofing and editing  literature cited  cover letter and reviewers  authorship  response to reviews  correcting proofs  copyright and open access  writing proposals biodiversity data capture and initial enrichment  biodiversity data capture  welcome to ghana  introduction to bitc  introduction to data capture course  what is biodiversity data?  biological data & collections: entomology perspective  data sharing standards: darwin core  data acquisition: gbif data  exercise: darwin core mapping  data models  authority files  persistent identifiers  exercise: data models  introduction to database management software systems  software highlight: arthropod easy capture  software highlight: brahms  software highlight: symbiota  software highlight: specify  data capture: strategy  data capture from images  data capture: demonstration  university of ghana herbarium tour  data capture: workflow analysis https://youtu.be/xmko5q-foyo https://www.youtube.com/playlist?list=plsm4rjxsj4anv3suizb6adas3nscnsevc https://www.youtube.com/playlist?list=plsm4rjxsj4amzp_odiceyfj3d-i2klv0o https://youtu.be/bgzadpwq1dc https://youtu.be/cw6-yovmyfu https://youtu.be/uvwajpmuopw https://youtu.be/ak9rfzlglss https://youtu.be/o8xnjch6zje https://youtu.be/odbyqct1itm https://www.youtube.com/playlist?list=plsm4rjxsj4ao749bytdbbi-xj5byel7lq https://youtu.be/o582ejhbzsc https://youtu.be/joieg9w8n1c https://youtu.be/4itn98s90cq https://youtu.be/e3g5vigeo8k https://youtu.be/wzjvzyso-a4 https://youtu.be/iuluzd01ksw https://youtu.be/fcezul43gcc https://youtu.be/1o1p7tmsu5m file:///c:/users/k747i383/appdata/local/temp/•%09https:/www.youtube.com/playlist%3flist=plsm4rjxsj4anwi8nrhgqg_rfpdkjlbmtp https://youtu.be/c3bebjpdhpw https://youtu.be/co39scurmgk https://youtu.be/6pjvjshfwlu https://www.youtube.com/playlist?list=plsm4rjxsj4apwvwpcgrzfhjftbf1dmfh1 https://www.youtube.com/playlist?list=plsm4rjxsj4ao7iuciwiza3liuh8mrzu3v https://www.youtube.com/playlist?list=plsm4rjxsj4ame54arluhgytvdtvlewkwh https://youtu.be/0xhaazcqafe https://www.youtube.com/playlist?list=plsm4rjxsj4amghqhyfthqjlat5ecouiin https://www.youtube.com/playlist?list=plsm4rjxsj4apl7lpvjlpr7igpcsrkyyee https://www.youtube.com/playlist?list=plsm4rjxsj4aobhgc-lb_ok2nggcc6u7xx https://www.youtube.com/playlist?list=plsm4rjxsj4andj-izohnkt8kaakjrxbcm https://www.youtube.com/playlist?list=plsm4rjxsj4appzcyfbkanqvqdgzfpkxsi https://youtu.be/cc7tr_d-3ui https://www.youtube.com/playlist?list=plsm4rjxsj4ao_zn_shwrzzexagmtfgw0o https://www.youtube.com/playlist?list=plsm4rjxsj4anubuyipj53nbzs4y_h0v5n https://www.youtube.com/playlist?list=plsm4rjxsj4apzsmirfs1rl2xj582yj8cu https://www.youtube.com/playlist?list=plsm4rjxsj4am1-0t8s9pavpmrrlckyqnm https://www.youtube.com/playlist?list=plsm4rjxsj4ao5g4s8mkgqql13ol1-t9eq https://www.youtube.com/playlist?list=plsm4rjxsj4apgvhhc2jexopuzhwtlujpn https://www.youtube.com/playlist?list=plsm4rjxsj4apeuaf4q_g6cuzhcupsjzz9 https://youtu.be/s-ppgqnxs3u https://www.youtube.com/playlist?list=plsm4rjxsj4apeluvhlfn0vufqhciswhxp biodiversity informatics, 10, 2015, 65-74 69  workflows large 2d  workflows large 3d  workflows small 3d  q&a  specimen imaging: image capture  specimen imaging: insects  specimen imaging: image processing  georeferencing from bitc  georeferencing: introduction  georeferencing: geographical concepts  georeferencing: map concepts  writing locality descriptions  locality types  software example: georeferencing calculator  georeferencing exercise: results  georeferencing collaboration & automation case studies  georeferencing from idigbio  collaboration to automation  geographic concepts: coordinate systems  point-radius method & best practices  using paper maps: named places  geolocate: the basics  introduction to geolocate  using geolocate  geolocate: batch processing  geolocate: collaborative georeferencing  collaborative georeferencing demonstration i  collaborative georeferencing demonstration part ii  georeferencing using specify 6.0 biodiversity inventories  welcome to cameroon  introduction to bitc  introduction to biodiversity inventories  inventories vs. sampling  inventories example: mexico & landscape change  why conduct biodiversity inventories?  long-term biodiversity studies  assessing inventory completeness part i  species accumulation curves  richness estimators: parametric vs. non-parametric probabilities  assessing inventory completeness part ii  results-based sampling  conducting inventories example: avifaunal  specimen preparation: ornithological  conducting inventories example: herpetofaunal i  conducting inventories example: herpetofaunal ii  specimen preparation: herpetological  conducting inventories example: botanical  inventory field methods: botanical  inventory data: ornithological  observational data: ornithological  inventory data: herpetological https://youtu.be/lt6ccgsdsxg https://www.youtube.com/playlist?list=plsm4rjxsj4apstvgt49itqrlu9m_e__6l https://www.youtube.com/playlist?list=plsm4rjxsj4amrfys4vs6wx1cdtv-669eq https://youtu.be/grku4u1jptc https://www.youtube.com/playlist?list=plsm4rjxsj4aphmxtfwh74tj7u5zzc1mef https://www.youtube.com/playlist?list=plsm4rjxsj4amjytg6hclxjrdggny7uys9 https://www.youtube.com/playlist?list=plsm4rjxsj4aoe2250e5zoo1trxawhmfxt https://www.youtube.com/playlist?list=plsm4rjxsj4anlsldkj238beokp8y3g-gp https://www.youtube.com/playlist?list=plsm4rjxsj4am5bykpp0ydtxzqokd_0ypd https://youtu.be/amsfqohz1ek https://www.youtube.com/playlist?list=plsm4rjxsj4apokhedpkscwvfkdwpzbp78 https://www.youtube.com/playlist?list=plsm4rjxsj4anolh0fo66fawjhrcvy55ta https://www.youtube.com/playlist?list=plsm4rjxsj4apsxnfmyyili330kacnzgb2 https://youtu.be/vkojotuzv9e https://www.youtube.com/playlist?list=plsm4rjxsj4an3-b0zwyvu9ik5cld3lwu8 https://vimeo.com/album/2163673/video/53006304 https://vimeo.com/album/2163673/video/53008556 https://vimeo.com/album/2163673/video/53006303 https://vimeo.com/album/2163673/video/71755411 https://vimeo.com/album/2163673/video/65222791 https://vimeo.com/album/2163673/video/53620662 https://vimeo.com/album/2163673/video/52703079 https://vimeo.com/album/2163673/video/65222618 https://vimeo.com/album/2163673/video/52708180 https://vimeo.com/album/2163673/video/96842881 https://vimeo.com/album/2163673/video/96842882 https://vimeo.com/album/2163673/video/50943206 https://www.youtube.com/playlist?list=plsm4rjxsj4andamdrbooqabavfe7rfeev https://www.youtube.com/playlist?list=plsm4rjxsj4amikcp-bssebq1jlz3zieyi https://www.youtube.com/playlist?list=plsm4rjxsj4aoeineo4q1jjmdnkn5zw55k https://youtu.be/bwy2qbkp_ls https://www.youtube.com/playlist?list=plsm4rjxsj4anv54gtmz1feqwfonbshaya https://www.youtube.com/playlist?list=plsm4rjxsj4anw2ysyk3vod7tmihc0mfgg https://www.youtube.com/playlist?list=plsm4rjxsj4aonpxrweo3g7a46a9wop2cm https://www.youtube.com/playlist?list=plsm4rjxsj4aptyuuod68hasfr-6lsax8j https://www.youtube.com/playlist?list=plsm4rjxsj4apwa_9_pic3yykcyl5v47bd https://youtu.be/ykxpcm8olrm https://www.youtube.com/playlist?list=plsm4rjxsj4ang7iozhnavrkt86iltym7l https://www.youtube.com/playlist?list=plsm4rjxsj4ap4vq9ib6tswvhtsakorxn_ https://www.youtube.com/playlist?list=plsm4rjxsj4amfo6mpkk0fzyrl6rgvdik2 https://www.youtube.com/playlist?list=plsm4rjxsj4aoszj-wa3awvxqzpshcqzto https://www.youtube.com/playlist?list=plsm4rjxsj4aomscrcumgemo6t9vf88pve https://www.youtube.com/playlist?list=plsm4rjxsj4amqmdil5zil0djnltzynwns https://youtu.be/ouo-sh4zemk https://www.youtube.com/playlist?list=plsm4rjxsj4aovstakpafi8pbifxkja5ep https://youtu.be/8x8dbpa1e4q https://www.youtube.com/playlist?list=plsm4rjxsj4aovegurxrss1xbutqks-lds https://youtu.be/m2vgl7da8-e https://www.youtube.com/playlist?list=plsm4rjxsj4apgf9m8mv22quiltpzy0zhh biodiversity informatics, 10, 2015, 65-74 70  daily lists  additional data & field notes  gps data  data security  course wrap-up biodiversity data cleaning and data publishing  basics  introduction to biodiversity data standards  darwin core: introduction  darwin core: detail  darwin core: archives  darwin core: mapping demonstration  darwin core: archives demonstration  cleaning  introduction to biodiversity data cleaning  data content considerations  openrefine: exploring real world data  assessing biodiversity data quality  primary biodiversity data: data precision  biodiversity data cleaning tools: info xy  why 'clean' biodiversity data?  error flagging  data cleaning demonstration: notepad  taxonomic assessment  data cleaning demonstration: refine and taxonomic names recognition  software demonstration: gbif names parser  software demonstration: iplant  software demonstration: refine and biodiversity datasets  software demonstration: refine and taxonomic names  repatriating improvements  publishing  introduction to data publishing  landscape of biodiversity data publishing tools  biodiversity data networks  extracting source data  ipt: importing biodiversity data  ipt: map to darwin core  ipt: metadata  ipt: metadata comments  ipt: darwin core archives  registering data with gbif  publishing sensitive data  the realities of sharing biodiversity data  data licensing  general q&a  trainee presentations  african conservation centre  annet nambuusi  yvette umurungi  tamene yohannes https://youtu.be/mkvl2ewcsxc https://www.youtube.com/playlist?list=plsm4rjxsj4angpuulitfhfsr7u4k9hxa2 https://www.youtube.com/playlist?list=plsm4rjxsj4amsipulgtfzwpv4xhxijmkm https://www.youtube.com/playlist?list=plsm4rjxsj4appgi5c9yfhrzq8vv_wuzyc https://youtu.be/ejldfn67o-g https://www.youtube.com/playlist?list=plsm4rjxsj4ao9ypkpyfnhha-d51fftnql https://youtu.be/m4gmq_fs7lq https://www.youtube.com/playlist?list=plsm4rjxsj4anuaptdln6vawrms6z-9x3f https://www.youtube.com/playlist?list=plsm4rjxsj4amz0izsjdtgmsuuixkr8_en https://www.youtube.com/playlist?list=plsm4rjxsj4anvy-w9riybsonibgg-bzyc https://www.youtube.com/playlist?list=plsm4rjxsj4amd31uun7npzsb1btcranyo https://youtu.be/0zwyzhofoze https://youtu.be/yefsft75j9e https://youtu.be/79_ivkyeuiy https://youtu.be/pis6kyti_6o https://www.youtube.com/playlist?list=plsm4rjxsj4am3un1udifvnrr59wa0g_pd https://www.youtube.com/playlist?list=plsm4rjxsj4anml-1hizndvc7atu-yl-6b https://www.youtube.com/playlist?list=plsm4rjxsj4anphbtwtbn5dfqv90mponzh https://www.youtube.com/playlist?list=plsm4rjxsj4apeglv47xjvurin1kyymyan https://youtu.be/qyf62eio704 https://www.youtube.com/playlist?list=plsm4rjxsj4apeqvlvqpevltdr6sfhsa57 https://www.youtube.com/playlist?list=plsm4rjxsj4ao2iiiukjquz6i2c0fll5oz https://youtu.be/h-t2yjdbnlc https://www.youtube.com/playlist?list=plsm4rjxsj4aodzexb4_hyblbjomus0ieo https://youtu.be/_7bnetm7gxy https://youtu.be/vgyyjydxbqc https://www.youtube.com/playlist?list=plsm4rjxsj4anr5z-oq6ctdmsoupljwhiz https://youtu.be/5fhtbs6wtqa https://www.youtube.com/playlist?list=plsm4rjxsj4anm3dtxxh-saszryto4hlhl https://youtu.be/9m5wrjmmfa0 https://youtu.be/9ems5mbozt0 https://youtu.be/ttsqj0xjizm https://youtu.be/1ekmqaprmzq https://youtu.be/bntqjegv6n0 https://youtu.be/bi6cfoh0gys https://youtu.be/ajdbya7cdfo https://youtu.be/iy3gxmst_h4 https://youtu.be/rnfvf_vib3s https://www.youtube.com/playlist?list=plsm4rjxsj4anq-zh9hllzqvl4qfgsh0-s https://www.youtube.com/playlist?list=plsm4rjxsj4ap7pa1fvl9m0fstam7pqr0n https://youtu.be/wbd2gjckklg https://youtu.be/qluvn6yuvaq https://youtu.be/tamnpjqcg4u https://youtu.be/ggbu38_yfyo https://youtu.be/spyarwwagsm biodiversity informatics, 10, 2015, 65-74 71 ecological niche modeling  introduction to enm course  concepts: distributional ecology of species  concepts: ecological niche modeling  concepts: species occurrence data  primary biodiversity data: considerations  primary biodiversity data: quality checks  environmental data: considerations  bias in occurrence data and model calibration  enm: model selection considerations  model calibration: bioclim and distance algorithms  introduction to complex algorithms  garp  model thresholding  no silver bullets  model evaluation: concepts & threshold-dependent approaches  model validation with small samples of species observation records  model evaluation: roc  m hypotheses  maxent outputs details  general q&a  applications  niche conservatism  rare species  conservation planning  diseases  invasive species  climate change  historical biogeography biodiversity data analysis  introduction to data analysis course  assessing biodiversity survey completeness o introduction to sampling biodiversity o sampling theory o examples: survey completeness & species accumulation curves o comparative analyses o software highlight: estimates o exercise: abundance & inventory completeness o survey site prioritization: single species o inventory completeness: survey gap analysis o survey gap analysis example: kenya o biodiversity data initiatives & data quality: gbif o exercises: data analysis  biodiversity summaries  macroecological analyses  conservation prioritization o introduction and history o prioritization example analyses o software highlight: zonation o software exercise: zonation https://youtu.be/lp_2r8_ptik https://www.youtube.com/playlist?list=plsm4rjxsj4aozrswjf9y8sey9bwwg2zu_ https://www.youtube.com/playlist?list=plsm4rjxsj4anux3bymlvfubkqt56bvygx https://youtu.be/5kvitakoz5o https://www.youtube.com/playlist?list=plsm4rjxsj4amxv4pspqpgsn_7ibzmxcoc https://www.youtube.com/playlist?list=plsm4rjxsj4aml-j_zjd3d1udtilel0qby https://www.youtube.com/playlist?list=plsm4rjxsj4anm5ekwvmaw7rimiiwdogwi https://youtu.be/bqb-li2_1v0 https://www.youtube.com/playlist?list=plsm4rjxsj4aos7xe3b0z-uqwwethqme1h https://www.youtube.com/playlist?list=plsm4rjxsj4anb9vxivnygiygbhqz8hyln https://www.youtube.com/playlist?list=plsm4rjxsj4amau6zem5kdy56e9nalvaun https://www.youtube.com/playlist?list=plsm4rjxsj4amtb60rjo4k6fformmefhpu https://youtu.be/znmrzkd89ik https://youtu.be/npq7rszv8ia https://www.youtube.com/playlist?list=plsm4rjxsj4an2jdobkn9ff_ayuohyk47j https://www.youtube.com/playlist?list=plsm4rjxsj4aorvxbyw6s1vuqisj9icvyx https://www.youtube.com/playlist?list=plsm4rjxsj4aovnoxzbyxwdwq1cvzbouxx https://youtu.be/moxcs2clcwm https://www.youtube.com/playlist?list=plsm4rjxsj4aob6b7hqv6xmvjgk2uymmw1 https://youtu.be/8sq5q3jwjsk https://www.youtube.com/playlist?list=plsm4rjxsj4apul7kah8kqd4pkwx1albiw https://www.youtube.com/playlist?list=plsm4rjxsj4ann43dy4-ecz2wpdeta94xe https://www.youtube.com/playlist?list=plsm4rjxsj4ap8zpypbzfziihlcoknpmpv https://www.youtube.com/playlist?list=plsm4rjxsj4aoft5w51zv4fcwyv9dcspp4 https://www.youtube.com/playlist?list=plsm4rjxsj4apbejljtidyvh2xfrvz05tq https://www.youtube.com/playlist?list=plsm4rjxsj4amw9ctucbcqrs68rmn4bxcl https://www.youtube.com/playlist?list=plsm4rjxsj4apd9dfg6wtbaygqyfllr71_ https://youtu.be/wwttg0fru84 https://youtu.be/20lc_-hwdxo https://www.youtube.com/playlist?list=plsm4rjxsj4amdx_mu30jao2x1p4oovxwb https://www.youtube.com/playlist?list=plsm4rjxsj4apr70kntmjqzgktgtzage3o https://www.youtube.com/playlist?list=plsm4rjxsj4ao9kbkpzjfugp5j3x4ofbds https://youtu.be/fssl3thrup4 https://www.youtube.com/playlist?list=plsm4rjxsj4ao80jfiptnixi78ab7mlklo https://www.youtube.com/playlist?list=plsm4rjxsj4ap5mwna_uq-vi3hhfbr3cdl https://www.youtube.com/playlist?list=plsm4rjxsj4anwv0opav3bgftmizzjyxff https://www.youtube.com/playlist?list=plsm4rjxsj4amxvzlhekpe6zajwac3aajo https://youtu.be/xxz8lzntmhy https://www.youtube.com/playlist?list=plsm4rjxsj4aoljy1batd86nwakauhbyzq https://www.youtube.com/playlist?list=plsm4rjxsj4am3mkoditf0uo9oyhs5bq2f https://www.youtube.com/playlist?list=plsm4rjxsj4apikple0yyr3gqvpzkuxv3c https://www.youtube.com/playlist?list=plsm4rjxsj4aozmwvb59ah2t7jny3vwxst https://www.youtube.com/playlist?list=plsm4rjxsj4amyb2fecwa7_c881ah_uizw https://www.youtube.com/playlist?list=plsm4rjxsj4amyq37jpd5ksp4bxkyiwg-y https://www.youtube.com/playlist?list=plsm4rjxsj4apavfc_8_jtyaziqaam7cb1 biodiversity informatics, 10, 2015, 65-74 72 species descriptions  introduction to species descriptions course  new species discovery case study: cameroon  nomenclatural codes  species descriptions: nuts & bolts part i  systematics principles  naming species example: platymantis  naming species example: bufo  naming species example: caprimulgidae  naming conventions & name selection  species descriptions: nuts & bolts part ii  species description example: frog 1  species description example: bird  species description example: frog 2 public health applications of biodiversity data  introduction to public health applications course  mapping diseases transmission risk  introduction to disease ecology  current pha toolkit  q&a: part i  distributional ecology primer  how disease systems are different  ecology and biogeography of diseases  q&a: part ii  niche modeling overview  example: mpx in congo basin  disease particulars  beyond suitability  evaluating risk maps  example: buruli ulcer  example: white nose syndrome  example: mycetoma  example: h5n1 flu  q&a: part iii  future constraints  opportunities and ideas  public health and biodiversity building biodiversity informatics institutions  welcome to sanbi  introduction to building biodiversity informatics institutions course  introduction from sanbi ceo  information is power  why start a biodiversity institution?  need for biodiversity knowledge  biodiversity policy  biodiversity information to policy  bi institution example: conabio part i  the biodiversity informatics initiative game  the universe of biodiversity informatics players  transitioning the scope of bi products: local to global  discussion: part i https://youtu.be/byfhrotpfdi https://youtu.be/qtbjr4aekh4 https://www.youtube.com/playlist?list=plsm4rjxsj4apirwlew-cw_nqnzomjw3iq https://youtu.be/epenehrtpxc https://www.youtube.com/playlist?list=plsm4rjxsj4amruaou6buxk0ea7hs8aogd https://youtu.be/mt51tz0n37k https://youtu.be/v0kvwk9yz64 https://youtu.be/yqr2f1oiak4 https://youtu.be/sqjhw3qkaeo https://youtu.be/sjgoffj9n40 https://youtu.be/nkxuiqmf3fm https://youtu.be/lueqx3x_se0 https://youtu.be/0zuuzjdwsjc https://youtu.be/i1fmoiyz6c4 https://youtu.be/mrh_tcgnc3s https://youtu.be/gpuhrt70np0 https://youtu.be/rgkqay_dje8 https://youtu.be/rh1ygzqualc https://youtu.be/d5vsck8pkow https://youtu.be/qpgamg7ofo4 https://youtu.be/fmzpf2bnpvi https://youtu.be/gwlhctuyzey https://youtu.be/j0rkf1l4ic8 https://youtu.be/q6v06ka8s5u https://youtu.be/hnbtpggpzky https://youtu.be/eevcprymq1i https://youtu.be/1_a5xm4oevk https://youtu.be/0lk1qnygeaw https://youtu.be/dz-qxkfjobk https://youtu.be/vt2cedcp_pk https://youtu.be/_ud7jmbcj8y https://youtu.be/k37tj1o5aic https://youtu.be/ycltk0y_ef4 https://youtu.be/esyyzuthhxo https://youtu.be/zh-p3orxq_a https://youtu.be/ebb1oehhvo0 https://youtu.be/umvmc755ave https://www.youtube.com/playlist?list=plsm4rjxsj4amtfhdigotdsdf47e21m2ehttps://www.youtube.com/playlist?list=plsm4rjxsj4aortfwalfuvxfszjcjimi37 https://www.youtube.com/playlist?list=plsm4rjxsj4apbqlu6glcld2hnbum6oizg https://www.youtube.com/playlist?list=plsm4rjxsj4amsyz7ziubwb9lr5ijvcka6 https://www.youtube.com/playlist?list=plsm4rjxsj4ao_-n7l2uwzryv-5wo7yhyc https://www.youtube.com/playlist?list=plsm4rjxsj4apubgbnli8p-v_xcip570n5 https://www.youtube.com/playlist?list=plsm4rjxsj4am__lw1ssx02uuwfrlp9dym https://youtu.be/jm9lekqjvtk https://www.youtube.com/playlist?list=plsm4rjxsj4anvwk_d5ipe6sn7uj87roiz https://www.youtube.com/playlist?list=plsm4rjxsj4amffuncefsdscyz-nf9xpen https://www.youtube.com/playlist?list=plsm4rjxsj4aoypl7romhcg84j9jrgocaa biodiversity informatics, 10, 2015, 65-74 73  policy & biodiversity science example: u.s.  ways forward in building bi institutions  bi demonstration: data quality  discussion: part ii  gbif: part i  bi data quality example: survey gap analysis using kenya gbif data  gbif: part ii  biodiversity informatics partnerships  bi institution example: conabio part ii  bi partnerships: centro de referência em informação ambiental (cria)  bi institution example: centro de referência em informação ambiental (cria)  bi institution example: university of kansas biodiversity institute national biodiversity diagnoses  introduction to national biodiversity diagnoses course  publication plans  data cleaning  data quality checks: taxonomy  data quality checks: geography  data quality checks: temporal  basic patterns  assessing inventory completeness  gaps: knowledge vs. data  environmental gaps in biodiversity inventories  environmental variation across study regions  environmental variation demonstration  richness and endemism  beta diversity  protocol for national biodiversity diagnoses biodiversity conservation implementation  introductory material o course introduction o why areas are protected  overview presentations o digital accessible knowledge about biodiversity and choosing sites for conservation action o conservation biology and conservation science o people and protected areas: past present and future o climate change and wine o conservation in ethiopia  case studies o maasai mara, kenya o conservation lessons o madre de dios region, peru o tehuacan-cuicatlan biosphere reserve, mexico o lions in nech sar national park, ethiopia  thematic presentations o history of pa conservation strategies o design strategy o balancing needs of local communities and conservation priorities o ethiopian protected areas o dynamic conservation planning o climate change and south african reserve example o indicator species https://www.youtube.com/playlist?list=plsm4rjxsj4aoxshrm7qtpnpvt2fkbpxgz https://youtu.be/cricuizqlaq https://www.youtube.com/playlist?list=plsm4rjxsj4apwm6miwrqabpyafi3zcb4l https://www.youtube.com/playlist?list=plsm4rjxsj4anyrqb2o0nkwtscgmmcprou https://www.youtube.com/playlist?list=plsm4rjxsj4amypllevhpnlwfrkrhjrbuo https://www.youtube.com/playlist?list=plsm4rjxsj4amzzro4ejmpe7zd08tyh_er https://www.youtube.com/playlist?list=plsm4rjxsj4apcfraexbbznypxr52d3sfe https://www.youtube.com/playlist?list=plsm4rjxsj4aokmyuarzaptpmfbplhzv0e https://www.youtube.com/playlist?list=plsm4rjxsj4aphhqkf5wl4yudqgnt6gsbq https://youtu.be/f6hje60mze0 https://www.youtube.com/playlist?list=plsm4rjxsj4anrlpyyvhvv7im0vxnsmgod https://www.youtube.com/playlist?list=plsm4rjxsj4ammbkfncmqs521zb6jxh1mw https://youtu.be/ds821dzucak https://youtu.be/h-bm7vf0xj4 https://www.youtube.com/playlist?list=plsm4rjxsj4amf9s6qjxhnobtaaho6a-ua https://www.youtube.com/playlist?list=plsm4rjxsj4anwmqo_h5qwp_sehk-7jhjg https://youtu.be/x8fqsggegvq https://youtu.be/nht4nj_xq_e https://www.youtube.com/playlist?list=plsm4rjxsj4apoxum-gdcrhbzz547btu2_ https://www.youtube.com/playlist?list=plsm4rjxsj4am8xe8api7yerg8avyaq364 https://www.youtube.com/playlist?list=plsm4rjxsj4amzhztl9_95kisut0qehk_p https://youtu.be/unz28g1ebqw https://youtu.be/r1bsi73h3qi https://youtu.be/k1k6hjnkoi0 https://www.youtube.com/playlist?list=plsm4rjxsj4apovf7uihbg4kwninwohmft https://www.youtube.com/playlist?list=plsm4rjxsj4aphqgnrdwcggqenuptd0stb https://www.youtube.com/playlist?list=plsm4rjxsj4ank7eqo8rgunh5ajndacndn https://youtu.be/cvy3jeoenaq https://youtu.be/2lit3yazswu https://youtu.be/2dvbdvawb9w https://youtu.be/ntor_4psijy https://youtu.be/iyngl73ocx4 https://youtu.be/6na6040jz8u https://youtu.be/336luawoooi https://youtu.be/gctcvwo9fqi https://youtu.be/agyl8q-5mwg https://youtu.be/s9mnkzlhpyk https://youtu.be/e6qjrspljbw https://youtu.be/yqu_93nio3w https://www.youtube.com/playlist?list=plsm4rjxsj4aon77x4oxkbqp9mgcqy2znq https://www.youtube.com/playlist?list=plsm4rjxsj4anvfpfxefynk-5dynh8yjc2 https://www.youtube.com/playlist?list=plsm4rjxsj4apmecjw-nvo8uixyjuno755 https://www.youtube.com/playlist?list=plsm4rjxsj4aob2wd1_a25sviuwy8qaf2q https://www.youtube.com/playlist?list=plsm4rjxsj4apvcuj2srzrsy--j8yyizoo https://www.youtube.com/playlist?list=plsm4rjxsj4am1x1eapl-ck5dil9kuldzs https://www.youtube.com/playlist?list=plsm4rjxsj4an_6dqgfpcb7cplyu3rccz4 biodiversity informatics, 10, 2015, 65-74 74 o maintaining conservation as climate changes o climate change primer o integrity of protected areas  participant case studies: o egypt o cameroon o liberia https://www.youtube.com/playlist?list=plsm4rjxsj4ao_anjjxlwvptwhegzncngo https://www.youtube.com/playlist?list=plsm4rjxsj4aoasmcktismmg9xgjb0hho2 https://www.youtube.com/playlist?list=plsm4rjxsj4aphtw6lg0oobeub-rdsdrzs https://youtu.be/8shtbj8zaqg https://youtu.be/czdhdh1vxgk https://youtu.be/xz6kg1iy4z0%20%ef%bb%bf%20werwe biodiversity informatics, 15, 2020, pp. 55-56 54 response to stephens et al. (2020) a. townsend peterson1, jorge soberón1, janine m. ramsey2 and luis osorio-olvera1.3 1biodiversity institute, university of kansas, lawrence, kansas 66045 usa 2instituto nacional de salud pública, tapachula, chiapas, méxico 3centro de cambio global y sustentabilidad, conacyt, villahermosa, tabasco, méxico readers of the contributions to this debate will no doubt be daunted by the length and density of the presentation of the co-occurrences-imply-interactions methodology by stephens et al. (2020). that is, stephens et al.’s (2020) presentation of the inspiration, concepts, and justification for the methodology is presented over too many pages, including considerable amounts of text that is lateral, peripheral, and/or extraneous to the main challenge of presenting, justifying, and defending a novel methodology in a debate. the overwhelming length and detail are distracting, and we are concerned that it may obscure certain crucial details (and failings) of the authors’ arguments. the argument for the “co-occurrences-imply-interactions” methodology centers on a rather peculiar set of definitions. that is, “biotic interactions” are accorded a rather holy place in ecology, being the central and defining processes in the entire field of community ecology (e.g., mittelbach and mcgill 2019). these interactions are defined in terms of the precise roles (e.g., predator-prey, pathogen-reservoir), or of relative benefits to each of the interacting species (e.g., symbiosis, mutualism, parasitism, etc.), which are then grouped more broadly based on impact, into positive, neutral, and negative interactions. stephens et al. (2020), however, have opted to recycle and redefine this rather important term in ecology, so that it fits with what can be estimated with their methodology. that is, they stated: in our methodology, an interaction is defined by quantifying the degree of co-occurrence of variables—biotic or abiotic—relative to that expected in the absence of the interaction. they also stated: we have defined an interaction as a deviation from an appropriate null hypothesis of the spatial distribution of a taxon conditioned on one or more abiotic and/or biotic variables. clearly, stephens et al. (2020) are using a definition of “interaction” that is quite distinct from that which is in universal use in ecology. rather than a definition that responds directly to the biological processes in question, such as one animal eating another (= predation), or an animal pollinating a plant, they have redefined “interaction” to refer simply to spatial co-occurrence. this empirical and observable definition might be useful were it to be termed “spatial attraction,” or some similar term, but it is quite deceptive and confusing because of its re-definition of such an important term in ecology. indeed, stephens et al. (2020) are aware of the challenges involved in the inferences that they are attempting to make. they stated: in ecology, as elsewhere, co-occurrences are a necessary condition for an interaction. for a predation event to occur, the predator and the prey must be in the same place at the same time. similarly, for pollination, or any other type of ecological micro interaction. we concur. obviously, co-occurrence is necessary for an interaction to occur naturally. predation or pollination cannot occur if the two species are not ever together in the same place. still, more information is necessary if one is to be able to make conclusions about the type of interaction—notice that, in the quotation above, both predation and pollination are mentioned, and both require co-occurrence to be positive, and yet one is a negative interaction and the other is a positive interaction. quite simply, more information is needed before one can make any concrete conclusions about the type, or even the general direction of the interactions between species. two recent empirical papers exploring these issues (sander et al. 2017; freilich et al. 2018) paint a very different picture, with the authors being careful and specific when reporting and interpreting their results. for instance, freilich et al. (2018) concluded, a. townsend peterson et al. – response to stephens et al. (2019) 55 “thus, as observed in previous empirical and theoretical studies, patterns of interactions in co-occurrence networks must be interpreted with caution.” similarly, sander et al. (2017) stated: our findings suggest that although these methods hold some promise for ecological network inference, presence-absence data does not provide enough signal for models to consistently identify interactions, and networks inferred from these data should be interpreted with caution. as such, other research groups have arrived at ambiguous and non-conclusive results in their empirical studies as a result of the many factors hindering detection of a proper, actual, biologically-defined interaction. we believe that this non-conclusion is not a function of the simpler need for better software, but rather is illustrative of careful and appropriate caution in interpretation of results. to summarize, although stephens et al. (2020) contribute to assembling an interesting and useful analysis tool in the species site, we disagree strongly with them as regards the interpretation that co-occurrence signals can predict species interactions. although stephens et al. (2020) have reinterpreted co-occurrence signals as “interactions,” co-occurrence alone does not carry sufficient information to permit rigorous interpretation as predictions of actual individual contact or different types of interactions. rather, co-occurrence should be interpreted as exactly that: co-occurrence that signals geographic distributional coincidence. interpretation as actual interactions—and determining types of interactions—requires further information that is generally unavailable from simple occurrence data. literature cited freilich, m. a., e. wieters, b. r. broitman, p. a. marquet, and s. a. navarrete. 2018. species co‐occurrence networks: can they reveal trophic and non‐trophic interactions in ecological communities? ecology 99:690-699. mittelbach, g. g., and b. j. mcgill. 2019. community ecology. oxford university press. sander, e. l., j. t. wootton, and s. allesina. 2017. ecological network inference from long-term presence-absence data. scientific reports 7:7154. stephens, c. r., c. gonzález-salazar, m. villalobos, and p. a. marquet. 2020. can ecological interactions be inferred from spatial data? biodiversity informatics 17:11-54. microsoft word alex1.docx biodiversity informatics, 13, 2018, pp. 27-37 assessment of biodiversity data holdings and user data needs for ghana alex asase1* and gladys o. schwinger2 1department of plant and environmental biology, university of ghana, p. o. box lg 55, legon, ghana. 2institute for environment and sanitation studies, university of ghana, p. o. box lg 209, legon, ghana. *email: aasase@ug.edu.gh. abstract.—data on biodiversity are important to addressing the challenges of sustainable development, and for decision-making about natural resources and environments. biodiversity information, when mobilized and shared openly, has the potential to impact science and conservation positively. however, biodiversity data mobilization is expensive, such that data mobilization and sharing activities must be prioritized to meet the needs of the user community. in this study, we undertook a detailed assessment of biodiversity data holdings and user needs in ghana through semi-structured questionnaire interviews, and focus-group discussions in the form of a workshop. most biodiversity data-holding organizations were at preliminary stages of digital biodiversity data mobilization and sharing. taxonomic, checklist, and geographic data on plants and animals were identified as most needed. priority thematic needs were as regards protected areas, invasive alien species, threatened species, economic species (timber and non-timber forest products), and pathogens and diseases. human and infrastructural capacities, and sustainable coordination were identified as the major challenges to biodiversity data management. this study provides a detailed case study of how assessing biodiversity data holdings and user data needs can be used to strategize biodiversity data mobilization, data publication, and data use activities. key words.—biodiversity, primary biodiversity data, data mobilization, digitization, ghana biological diversity may be defined as the full variation of living organisms from all sources including, inter alia, terrestrial, marine, and other aquatic ecosystems and the ecological complexes of which they are a part. this term thus includes diversity within species, between species, and of ecosystems (cbd, 2001). it includes genetic diversity, species diversity, ecosystem diversity, and associated evolutionary and ecological processes. biodiversity is a compound word derived from biological diversity, and therefore is considered to have the same meaning. biodiversity is important for human wellbeing: it provides tangible benefits, such as food, clothing, and shelter, as well as intangible benefits such as climate amelioration and clean water (brauman et al., 2007). however, biodiversity is being lost at unprecedented rates owing to a plethora of factors: deforestation, agricultural expansion, habitat loss, timber extraction, firewood collection, and mineral extraction (norris et al., 2010). data on biodiversity are crucial to addressing the challenges of sustainable development and decision-making about natural resources and environments (chapman, 2005; sousa-baena et al., 2013). biodiversity data include data on species inventories, distributions, images, sounds, specimens, and ecological interactions, as well as descriptions of datasets (i.e., metadata) (costello et al., 2013). biodiversity data are basically of two kinds (i.e., primary and secondary), and can be numerical, categorical (e.g., species or place names), or mediabased (costello et al., 2013). primary biodiversity data are data records that document the occurrence of a particular species at a place at a point in time. in contrast, secondary biodiversity data represent summaries, interpretations, or syntheses of primary biodiversity data. secondary biodiversity data, such as species atlases and range maps from field guides, often include subjective elements that reduce their utility when compared to primary biodiversity data. as such, primary biodiversity data have many applications: documenting basic biodiversity patterns (guralnick and hill, 2009), identifying priority areas for conservation efforts (myers et al., 2000), providing baseline information for detection of biotic change (peterson et al., 2015), and supporting modeling efforts that anticipate biotic responses to local and global change (ehrlen and morris, 2015). sources of primary biodiversity data are many, including labels associated with specimens in research collections of natural history museums and herbaria, and data from field studies and asase and schwinger – ghana biodiversity informatics activity observations made by scientists and researchers (peterson et al., 2011), as well as data from field observations made by citizen scientists (asase and peterson, 2016). types of primary biodiversity data include primary occurrence data that document presence (and absence) of organisms such as those on labels on herbarium sheets (peterson et al., 2011); sample-based data that have information on species occurrence and their abundance, such as those from ecological plot inventories; and multimedia data such as sound, images and videos. biodiversity data offer greatest information when they are integrated and used with other data types, such as environmental data and socioeconomic data, to address the most pressing questions in biodiversity and sustainability science (faith et al., 2013). nonetheless, the importance of natural history collections has been demonstrated clearly (e.g., ariño et al., 2013). at this point in time, biodiversity data must be digitized and shared openly (costello et al., 2013). certainly, past decades have witnessed massive progress in digitization of biodiversity data thanks to advances in information technology, development of efficient data digitization workflows, and changes in policies of owners of primary biodiversity data (asase and peterson, 2016). biodiversity data digitization and sharing of natural history collections include several stages: pre-digitization preparation, advance curation, image capture, processing and storage, capture of data records from either images or specimens, georeferencing, data cleaning, and data publication (nelson et al., 2015). it is expensive to produce and share biodiversity data, and not all data are fit for all uses (hills et al., 2010). consequently, tasks of digitization and sharing of primary biodiversity data must be prioritized to meet the needs of the user community. surveys of biodiversity data user needs are important to understanding data needs across diverse user communities (gaiji et al., 2013), and to measuring scientific and policy contributions of data mobilized (ariño et al., 2013). here, we undertook a detailed assessment of biodiversity data holdings and data needs across the user communities of scientists, researchers, curators of natural history collections, non-governmental organizations (ngos), and policy-makers in ghana. we explored how detailed assessment of biodiversity holdings and user needs of a country can be used to guide biodiversity data mobilization, data sharing, and data use activities. we used ghana as a case study in view of its active involvement in biodiversity data mobilization activities and biodiversity informatics initiatives, such as the global biodiversity information facility (gbif). methods data were collected via a combination of two common survey methods: semi-structured interviews, and focus group discussions. interviews were achieved via distribution of a questionnaire to major biodiversity stakeholder organizations in ghana for completion. stakeholders included university departments, government agencies, research institutions, ngos, and other groups. questionnaires were administered either in person or were sent to organizations for them to complete. we did not attempt online questionnaires owing to unreliable internet connectivity. the questionnaire was designed to assemble information on three main areas: (1) the profile of organization, (2) its data holdings, and (3) its biodiversity data needs. for data holdings, our focus was on the status of the holdings, strategies towards data mobilization, the digitization landscape, and attitudes about data publication, as well as data preservation and archiving. we focused on data needs regarding five broad categories: checklist, taxonomic data, geographic data, ecological data, and educational data (i.e., biodiversity data for public education, and training materials). within each broad category, we assessed level of importance, availability, and sustainability of the data source for major taxonomic groups (plants, animals, fungi, microorganisms, algae). in total, we received 22 fully completed questionnaires from stakeholder organizations out of >50 sent out. the stakeholders were 50% from academia and research institutions, 40% from governance and policy institutions, and 10% from ngos and other groups, and thus were broadly representative of major biodiversity stakeholder groups in ghana. focus group discussions were in the form of a national biodiversity stakeholders’ workshop. the aim of this workshop was to arrive at national consensus on biodiversity information needs for ghana, including identifying challenges. the workshop included three breakout sessions among different stakeholder types: (a) academic and research, (b) governance and policy, and (b) ngo and others. about 61% of the stakeholders that completed the questionnaire attended the workshop, and 31 stakeholder groups in total attended. each of the three asase and schwinger – ghana biodiversity informatics activity groups was tasked to brainstorm on data needs, opportunities, threats, and challenges in biodiversity data mobilization for ghana. after initial group discussions, the groups reconvened and discussed major findings of each group. the results of the discussions were summarized in the form of a “strengths, weaknesses, opportunities and threats” (swot) analysis, as well as with summaries of priority data needs and data challenges. workshop participants were well-trained people with capacity to answer questions and verify answers regarding biodiversity information in ghana; each participant presented views of their respective institutions. data were entered into microsoft excel, checked for consistency, and cleaned for errors such as duplications and name variants. using the pivottable function in excel, we analyzed responses to survey questions in terms of frequencies. summary statistics and frequencies data were compiled, and presented using appropriate visual aids. results data holdings, digitization, and publication of the 22 organizations interviewed, 15 had biological collections, 4 had only field data, and 3 were data users only. about 89.0% of the 19 organizations with data had clearly defined purposes for mobilizing data, and 36.8% had already assessed the scope and extent of biodiversity data digitization (fig. 1). similar proportions (47.4%) of the organizations had already achieved pre-digitization activities or had not undertaken pre-digitization activities; the remaining 5.3% did not know about predigitization activities. institutional data policies (e.g., as regards data sharing) were present in 26.3% of the organizations; most (63.2%) of the organizations had no such policies. most (68%) of the organizations had no data management systems in place. about 26% of the organizations had a digitization workspace, whereas 68.4% had no digitization workspace, while 5.3% were not sure (table 1). ten of the organizations were knowledgeable about the proportions of their holdings that had been digitized, whereas 47.4% did not know. pre-curation digitization activities had been carried out by only 36.8% of organizations, whereas 57.9% of organizations had not pursued such activities. digitization processes and technologies existed for 21.1% of organizations, whereas a majority (73.7%) of organizations had no such processes and technologies in place. out of the 19 data holders, only 42.1% had well-defined workflows for digitization. most (73.7%) organizations did not have staff adequately trained and equipped for digitization. more than half (57.9%) of organizations had no specifications for data quality control and standards. however, most (73.7%) organizations had selected a data platform for digitization, and formats for digital data storage were available for 42.1% of organizations. regarding data publication, 37.1% of organizations had considered end-users and web publication needs, whereas the rest had not made such considerations. about 74.1% of organizations had not selected any data publication tool or data licensing option, and 68.3% have no online platform for sharing data. access to long-term archival repositories and safeguards against obsolesce of data formats and applications were not available in most of the organizations. data needs and challenges the questionnaire interviews were helpful in identifying biodiversity data needs of the organizations interviewed. data on plants were most frequently cited as needed, followed by those on animals; data on fungi were least mentioned as needed (fig. 2). data types most frequently mentioned as needed were taxonomic data (26%), followed by checklists (25%) and geographic data (25%), then ecological data (18%), and lastly educational data (6%). at the same time, the most available data type was taxonomic data, while checklists were the least available (fig. 3). details on data needs according to data type and the major taxonomic groups are presented in table 2. about 49% of data sources were regarded as sustainable; 44.4% were unsustainable; for 6.6%, it was unknown. the stakeholders’ workshop was important in highlighting biodiversity data needs at the national level. results of the swot analysis revealed strengths (8 points), opportunities (6 points), weaknesses (5 points), and threats (2 points) regarding biodiversity data in ghana (table 3). the data requirements of the different user-groups indicated above varied (table 4). in general, stakeholders identified by consensus 7 priority areas of data need areas: (a) up to date biodiversity inventory data on protected areas such as forest reserves and national parks; (b) geographic data on biodiversity under land-use and climate change scenarios; (c) national red lists of biodiversity for asase and schwinger – ghana biodiversity informatics activity major groups such as plants, mammals, birds, and insects; (d) pathogens and microbes, and their effects on economic crops and livestock; (e) harvested biodiversity particularly timber and non-timber forest products (ntfps); (f) economic and useful biodiversity such as medicinal plants and edible mushrooms; and (g) invasive alien species, and their presence and distribution. major challenges to biodiversity data included poor human and infrastructural capacity, lack of funds for data acquisition, and lack of sustainable coordination. discussion biodiversity data mobilization landscape the need for biodiversity data to be made accessible, discoverable, and integrated cannot be overstated, as biodiversity research is rapidly becoming a data-intensive science (kelling et al., 2009). biodiversity data mobilization is an expensive enterprise, but once mobilized, the data can be of great value (borgman, 2007). the results of this study show that the biodiversity data mobilization landscape in ghana is at preliminary stages of operation, as most organizations have not yet begun strategizing for data mobilization. it is always important that proper strategies and planning be in place before embarking on data mobilization (frazier et al., 2008) to assure effective implementation of project tasks, and ultimately the success of projects. such strategies should consider the various project stakeholders, and their roles at the different phases of the project implementation. long-term policies and a strong vision towards gathering and using biodiversity data are required for any country such as ghana. currently, no such policies and vision have been articulated for ghana, which is worrying. the present study has been helpful in characterizing the major biodiversity data holdings in ghana. biodiversity data-holders may be either cooperative or non-cooperative. cooperative biodiversity data-holders are those willing to collaborate and share their data resources openly, whereas non-cooperative biodiversity data-holders do not want to participate and share their data resources. it is unclear which of the many biodiversity dataholders identified in ghana belongs to each of the two types; as such all data-holders must be encouraged, especially those that prove to be non-cooperative. information that is shared and accessible has the potential to impact science and conservation, as well as the care and curation of specimens (asase and peterson, 2016). barriers to sharing biodiversity data could be psychological and behavioral (including legal barriers), or may relate to describing information and data, or may spring from inadequate strategies and resources. non-cooperative data-holders would most likely become at least intermediate biodiversitydata holders willing to collaborate and share their data resources after the barriers have been identified, and they have been highly motivated and assured of incentives. another area of importance in terms of biodiversity data mobilization is the large amounts of data associated with significant collections of ghanaian biological and paleontological specimens held elsewhere in the world. data repatriation from european and north american organizations with large collections from ghana is an important potential source data on ghanaian biodiversity. for example, european and north american institutions, such as naturalis biodiversity centre in the netherlands, missouri botanic gardens in the united states, and the royal botanic gardens in the united kingdom have digital images and data records of botanical collections from ghana that they could make openly to the appropriate institutions in ghana. indeed, naturalis biodiversity centre has already provided large series of digital images of botanic collections from ghana to the ghana herbarium at university of ghana. digitization of biodiversity data refers to capture of information in electronic form from checklists, field notebooks, or specimens, or may be extracted from publications, documents, or other media. it may refer to electronic capture of an image of an object, or it can also be refer to capture of textual information about an object or extracted from an object that contains text (frazier et al., 2008). the advantages of biodiversity digitization are many: broad dissemination of data via open and accessible platforms; enabling natural history collections to be studied in different ways, including from outside of the museum or herbarium; enhancement of curatorial activities. this step also reduces future time spent on transcription of data records from specimens, and enhances visibility of institutions sharing data. biodiversity digitization workflows and protocols have been developed to maximize rates of specimen digitization without sacrificing the most useful information on each specimen (tulig et al., 2012). according to nelson et al. (2015), efficient workflows provide the foundation for successful digitization of asase and schwinger – ghana biodiversity informatics activity biodiversity collections and foster mobilization of increased quantities of specimen data for scientific research, natural resource management, education, and policy-making. application of available workflows for digitization of natural history collections in ghana will lead to better refinement, and additions that will increase availability of mobilized biodiversity data, and enhance specimenbased research in the country. the ghana herbarium at the university of ghana has started using some of these protocols. it is highly encouraged that other biodiversity data holding institutions in ghana, particularly those currently involved in biodiversity informatics projects e.g., plant genetic resource research institute of the council for scientific and industrial research (csir-pgrri), a rocha ghana and conservation alliance will apply such protocols in their digitization programmes. data needs and data use surveys of biodiversity data needs of various organizations and user groups can be a useful means of identifying common data needs versus priority data needs. the present study provided insights into data needs of various organizations and user groups across ghana. most organizations interviewed needed data on the taxonomy and geography of taxa. high demand for taxonomic data, particularly as regards identifications, was not surprising because it is fundamental to communicating about biodiversity (judd, 2008). also, data on names (taxonomic data) and place (geographic data) are important to exploring joint efforts that relate directly to applications such as ecological niche modeling and species distribution modelling (peterson et al., 2018). data on plants were the most required probably because of high human dependence on plants and plant products. data needs identified as priority areas at the national level in this study concerns protected areas, invasive alien species, threatened species, economic species, and pathogens and diseases. these data needs priority areas fall within the areas of interest of international organizations and bodies such as gbif, international union of conservation of nature, united nation environment programme–world conservation monitoring centre, global earth observation biodiversity observation network, intergovernmental science-policy platform on biodiversity and ecosystem services (ipbes), and 1 http://gef-connect.web-staging.linode.unep-wcmc.org/. convention on biological diversity. developing programmes and collaborating with these organizations will be useful in mobilizing primary biodiversity data about ghana. other priority areas should include agricultural biodiversity, given the fact that ghana is an agrarian country, and aquatic biodiversity, because it is less studied than terrestrial biodiversity. it is also important that data needs be aligned to meet national and international obligations such as the clearing house mechanism and aichi targets 2020 of the convention of biological diversity, and the sustainable development goals, to which ghana is a signatory. another area worthy of consideration are the so-called “essential biodiversity variables” for monitoring biodiversity change (pereira et al., 2013). our swot analysis was useful in identifying strengths, weaknesses, opportunities and threats regarding biodiversity data in ghana. this information can be used in formulating strategic management decisions concerning biodiversity data and ecosystem services as they relate to the societal needs for ghana. for example, ipbes has underscored the importance of integrating biodiversity data with data on ecosystem services. biodiversity data should be used in making policy decisions to support sustainable development; for research at universities, colleges, and research institutions; and to enable training at different levels of the educational ladder in ghana. mainstreaming biodiversity data into national policy decisions could provide insights about how data are used, and could also demonstrate the value of digital mobilization of biodiversity data. unfortunately, as pointed out in the ghana shared growth and development agenda (gsgda) ii (2014-2017) policy document, ghana sees weak integration of biodiversity issues in decision-making, especially at the local level, in ghana (ndpc, 2014). the unep-wcmc connect project1 in ghana, mozambique and uganda is one such model project on mainstreaming biodiversity into national policy decisions. awareness about the importance of biodiversity at schools could promote biodiversity conservation into the future. training and capacity enhancement limited human and infrastructural capacities were identified as challenges to biodiversity data mobilization in ghana. for example, ~74% of stakeholder organizations in this study do not have asase and schwinger – ghana biodiversity informatics activity staff adequately trained to digitize and share biodiversity data. this scenario is worrying, as it may lead to data quality issues, an issue of grave concern in biodiversity data management (veiga et al., 2017). adequate human expertise and skills are required to produce research-grade data for use: data must be properly captured, georeferenced, and cleaned before publication, such that data shared will have immediate applications (peterson et al., 2018). the biodiversity informatics community has developed tools and standards to support these challenges, such as the botanical research and herbarium management systems (brahms2) and specify3 for capturing data records; geolocate has been developed for geo-referencing data records (guralnick et al., 2006); darwincore was developed as a standard for publishing and integrating biodiversity information (wieczorek et al., 2012); and the integrated publishing toolkit (ipt) is a tool for sharing biodiversity datasets via the internet (robertson et al., 2014). training in use of these tools and standards is required to deliver usable dak (peterson et al. 2018). another key area of human capacity is in data analysis—e.g., multivariate statistics, place-prioritization efforts, and ecological niche modelling—as such skills are necessary if data mobilized are to be used to inform national and regional decision-making. biodiversity informatics is a young science, with few or no approved textbooks or academic programs developed. however, many biodiversity informatics initiatives and resources exist that could be of help in addressing the biodiversity informatics capacity challenges for ghana: e.g., the biodiversity informatics training curriculum (bitc) (peterson and ingenloff, 2015), and various opportunities available through gbif, gbif-africa, and biodiversity information standards (tdwg). a longterm solution to the human capacity challenge is to develop sustainable biodiversity science programs in ghanaian universities and colleagues. in africa, a pilot biodiversity informatics programme has started at the university of abomey-calavi in benin, and another is planned at the university of western cape in south africa. mobilizing biodiversity data to ensure maximum access and use requires a robust and easily usable infrastructure (robertson et al., 2014). it is therefore important that the various ghanaian biodiversity 2http://herbaria.plants.ox.ac.uk/bol/brahms/software. organizations should consider investing in this area of science. basic equipment and software for data digitization, data storage, and data analysis are useful in achieving desired outcomes in biodiversity data management. perhaps a long-term solution to this problem is for the national government to commit to establishment of a sustainable central organization with the needed facilities, more or less following the example set by the south african biodiversity institute in south africa. biodiversity data and expertise are currently unevenly distributed in ghana such that most institutions with data and expertise are in the southern half of the country. for example, of the six welldeveloped herbaria in ghana, only the recently established savanna herbarium at the university for development studies (nyankpala) is situated in northern half of the country. this situation is not surprising, as ghanaian biodiversity is richest in the southern part of the country, especially in the southwest, where well-preserved forest remnants can be found, whereas much of northern ghana is covered with savanna (mes, 2002). the geographic distribution of biodiversity holding data institutions and expertise is important to how decisions are made about biodiversity information management in ghana. conclusions this study presents the first detailed assessment of biodiversity holdings and user needs for ghana. although a number of biodiversity stakeholders did not participate in the survey, this study has highlighted pertinent issues about the biodiversity data landscape in ghana. most biodiversity dataholding organizations in ghana are at preliminary stages of data mobilization, and human capacity and infrastructure, as well as sustainable coordination, are the major challenges to data mobilization, data publication, and data use. we did not solicit responses regarding biodiversity data for different ecosystems, or regarding ethnobiological and molecular data, which can be considered in future studies. a next logical step will be to undertake data gap analyses, to identify discrepancies between current ideas and the state of the entire biodiversity science enterprise in ghana. an analysis of completeness of digital accessible knowledge exists for the plants of ghana (asase and peterson, 2016), but completeness of 3http://www.sustain.specifysoftware.org. asase and schwinger – ghana biodiversity informatics activity digital accessible knowledge for other taxa has not been assessed. this study illustrates how assessing biodiversity data holdings and user data needs can be used to strategize for biodiversity data mobilization, data publication, and data use activities, for a country and/or for a taxon. our findings are relevant to diverse biodiversity stakeholders, such as researchers, museum curators, and policy makers, in formulating strategic ideas and policies concerning their biodiversity data resources. acknowledgments the authors are thankful to all the biodiversity stakeholders in ghana who participated in the study. we are most grateful to our funders: gbif, biodiversity information for development (contract #: bid-af2015-0032-nac; supported by the european commission), and the jrs biodiversity foundation. we are grateful to two anonymous reviewers of this manuscript for their insightful comments, and to town peterson for edits to improve this manuscript. references ariño, a.h., v. chavan, and d.p. faith. 2013. assessment of user needs of primary biodiversity data: analysis, concerns, and challenges. biodiv. inf. 8:59-93. asase, a., and a.t. peterson. 2016. completeness of digital accessible knowledge of the plants of ghana. biodiv. inf. 11:1-11. borgman c.l. 2007. scholarship in the digital age: information, infrastructure, and the internet. mit press, boston. brauman, k.a., g.c. daily, t.k.e. duarte, and h.a. mooney. 2007. the nature and value of ecosystem services: an overview highlighting hydrologic services. annu. rev. environ. resour. 32:67-98. cbd. 2001. handbook of the convention on biological diversity. earthscan publications ltd., london. chapman, a.d. 2005. principles and methods of data cleaning: primary species and species occurrence data, version 1.0. global biodiversity information facility, copenhagen.4 costello, m.j., w.k. michener, m. gahegan, z.-q. zhang, and p. bourne. 2013. biodiversity data should be published, cited and peer-reviewed. trends ecol. evol. 28: 454–461. ehrlén, j., and w.f. morris. 2015. predicting changes in the distribution and abundance of species under environmental change. ecol. lett. 18:303-314. faith d.p., b. collen, a.h ariño, p.o. koleff, j. kerr, j. guinotte, and v. chavan. 2013. bridging the data gaps: 4http://www.gbif.org/resource/80528. recommendations of the gbif content needs assessment task group. biodiv. inf. 8:41-58. gaiji, s., v. chavan, a.h. ariño, j. otegui, d. hobern, r. sood, and e. robles. 2013. content assessment of the primary biodiversity data published through gbif network: status, challenges and potentials. biodiv. inf. 8:94-172. guralnick, r., and a. hill. 2015. biodiversity informatics: automated approaches for documenting global biodiversity patterns and processes. bioinformatics 25:421-428. guralnick, r.p., j. wieczorek, r. beaman, and r.j. hijmans. 2006. biogeomancer: automated georeferencing to map the world's biodiversity data. plos biology 4:e381. hill, a.w., j. otegui, a.h. ariño, and r. guralnick. 2010. gbif position paper on future directions and recommendations for enhancing fitness-for-use across the gbif network, version 1.0. global biodiversity information facility, copenhagen. judd, w.s. 2008. plant systematics: a phylogenetic approach, 3rd edition. sinauer associates, sunderland, mass. kelling, s., w.m. hochachka, d. fink, m. riedewald, r. caruana, g. ballard, and g. hooker. 2009. data intensive science: a new paradigm for biodiversity studies. bioscience 59:613-620. mes. 2002. national biodiversity strategy for ghana. ministry of environment and science, ghana. myers, n., r.a. mittermeier, c.g. mittermeier, g.a.b. da fonseca, and j. kent. 2000. biodiversity hotspots for conservation priorities. nature 403:853-858. nelson, g., p. sweeney, l.e. wallace, r.k. rabeler, d. allar, h. brown, j.r. carter, m.w. denslow, e.r. ellwood, c.c. germain-aubrey, g. ed, e. gillespie, l.r. goertzen, b. legler, d.b. marchant, t.d. marsico, a.b. morris, z. murrell, m. nazaire, c. neefus, s. oberreiter, d. paul, b.r. ruhfel, t. sasek, j. shaw, p.s. soltis, k. watson, a. weeks, and a.r. mast. 2015. digitization workflows for flat sheets and packets of plants, algae, and fungi. appl. plant sci. 3:1500065. national development planning commission (ndpc). 2014. ghana shared growth and development agenda (gsgda) ii, 2014-2017. volume 1. policy framework. 242. norris, k., a. asase, b. collen, j. gockowksi, j. mason, b. phalan, and a. wade. 2010. biodiversity in a forestagriculture mosaic—the changing face of west african rainforests. biol. conserv. 143: 2341-2350. pereira, h.m., s. ferrier, m. walters, g.n. geller, r.h.g. jongman, r.j. scholes, m.w. bruford, n. brummitt, s.h.m. butchart, a.c. cardoso, n.c. coops, e. dulloo, d.p. faith, j. freyhof, r.d. gregory, c. heip, h.r. asase and schwinger – ghana biodiversity informatics activity hurtt, w. jetz, d.s. karp, m.a. mcgeoch, d. obura, y. onoda, n. pettorelli, b. reyers, r. sayre, j.r.w. scharlemann, s.n. stuart, e. turak, m. walpole, and m. wegmann. 2013. essential biodiversity variables. science 339: 277–278. peterson, a.t., a. asase, d.l. canhos, s. de souza, and j. wieczorek. 2018. data leakage and loss in biodiversity informatics. in review. peterson, a.t., and k. ingenloff. 2015. biodiversity informatics training curriculum, version 1.2. biodiv. inf. 10:65-74 peterson, a.t., j. soberón, r.g. pearson, r.p. anderson, e. martínez-meyer, m. nakamura, and m.b. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. robertson, t., m. döring, r. guralnick, d. bloom, j. wieczorek, k. braak, j. otegui, l. russell, and p. desmet. 2014. the gbif integrated publishing toolkit: facilitating the efficient publishing of biodiversity data on the internet. plos one 9(8): e102623. sousa-baena, m.s., l.c. garcia, and a.t. peterson. 2013. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. div. distrib. 20:369-381. tulig, m., n. tarnowsky, m. bevans, a. kirchgessner, and b.m. thiers. 2012. increasing the efficiency of digitization workflows for herbarium specimens. zookeys 209:103-113. veiga a.k., a.m. saraiva, a.d. chapman, p.j. morris, c. gendreau, d. schigel, and t.j. robertson. 2017. a conceptual framework for quality assessment and management of biodiversity data. plos one 12: e0178731. wieczorek, j., d. bloom, r. guralnick, s. blum, m. doring, r. giovanni, t. robertson, and d. vieglais. 2012. darwin core: an evolving communitydeveloped biodiversity data standard. plos one 7:e29715. biodiversity informatics, 13, 2018, pp. 27-37 figure 1: status of biodiversity data mobilization strategies of data holders in ghana. error bars ± standard error. figure 2: biodiversity data needs in ghana according to major taxonomic groups. asase and schwinger – ghana biodiversity informatics activity figure 3: survey responses regarding availability of biodiversity data in ghana. error bars ± standard error. table 1: summary of the biodiversity data digitization landscape in ghana. items responses yes no unknown digitization workspace available 5 13 1 proportion of collections digitized 10 9 0 pre-curation digitization carried out 7 11 1 digitization processes and technologies defined 5 14 0 digitization workflows defined 8 9 2 staff adequately trained and equipped 4 14 1 quality control and standards specified 7 11 1 data platform for digitization selected 14 2 3 format digital data stored selected 8 3 8 asase and schwinger – ghana biodiversity informatics activity table 2: biodiversity data needs according to data types and major biodiversity groups. biodiversity data biodiversity groups / frequencies plants animals fungi microorganisms algae total checklist 84 65 28 36 37 250 ecological data 72 59 36 36 31 234 educational data 39 27 14 5 9 94 geographic data 98 79 36 41 48 302 taxonomic data 120 87 50 56 56 369 total 413 317 164 174 181 1249 table 3: swot analysis of biodiversity data needs landscape for ghana. strengths (internal) 1. existence of a national legal framework pertaining to data use and management. 2. high technical expertise on biodiversity science. 3. rich natural history collections on plants, mammals and insects. 4. both public and private organizations have data on ghanaian biodiversity. 5. willingness of stakeholders to be trained in biodiversity informatics. 6. active participation of stakeholders in national biodiversity activities. 7. many biodiversity research programmes and projects across the country. 8. possibility to upgrading the national biodiversity committee into a biodiversity commission. weaknesses (internal) 1. lack of motivation for biodiversity data-holders to share data. 2. poor infrastructural and human capacities in biodiversity informatics. 3. lack / inadequate funds to support biodiversity data mobilization. 4. weak institutional collaborations about biodiversity. 5. lack of a national sustainable and coordinating biodiversity information hub. opportunities (external) 1. biodiversity data about ghana is available online through outlets such as gbif etc. 2. repatriation of biodiversity data associated with collections of natural history museums in europe and north america. 3. training opportunities in biodiversity informatics (e.g. bitc programme, gbif, tdwg) 4. ghana is a signatory to international biodiversity conventions such as convention on biological diversity, gbif and ipbes. 5. availability of external funds for biodiversity science (e.g. unewcmc connect project, jrs biodiversity foundation, biodiversity for development (gbif-bid) programme). 6. networking and collaborations with external partners. threats (external) 1. non-cooperative biodiversity data-holders. 2. lack of funds to capture data in european / north america institutions. table 4: biodiversity data needs or gaps, and challenges according to three broad user-groups in ghana. biodiversity data user group data needs / gaps challenges academia and research plant distribution and phenology; aquatic biodiversity; data on pathogens and disease-causing organisms; medicinal plants; and nomenclatural changes. lack of motivation to share data, lack of taxonomists, minimal capacity for collection and curation of biodiversity, low expertise in biodiversity informatics, and lack of a sustainable biodiversity datacoordinating unit. governance and policy savanna biodiversity, agrobiodiversity, marine biodiversity, lower-taxon biodiversity (fungi, algae etc.), and invasive species. lack of motivation to share data, lack of taxonomists, inadequate financial resources, absence of policy on biodiversity data, and lack of law enforcement on data. non-governmental organizations and others herpetofauna, red list of biodiversity in ghana, species distribution maps, harvested biodiversity, and effects of climate and land use change on biodiversity lack of sustainable central biodiversity coordinating unit, and little collaboration between biodiversity data-holders. lowenberg_paper biodiversity informatics, 13, 2018, pp. 11-26 a metric to quantify analogous conditions and rank environmental layers peter löwenberg-neto instituto de ciências da vida e da natureza, unila, av. tarquínio joslin dos santos, 1000 cep 85870-901, foz do iguacu, parana, brasil. peter.lowenberg@unila.edu.br abstract.—analogous conditions in environmental variables are expected because environments are spatially autocorrelated and often present similar combinations over geographic space. that similar environmental combinations may be found at different localities provides a crucial basis for correlative species distribution modeling. an absolutely analogous variable is constant, while a non-analogous variable has no-repeating values, yet no current method allows researchers to quantify intermediate degrees of analogous conditions and rank environmental layers. i approached this issue from the perspective of dual-space correspondence, in which (a) variable range and modal frequency have a theoretical inverse relationship (y ∝ x-1), and (b) modal values of frequency are limited by the number of pixels in a given raster layer. for two geographic extents and two resolutions (2.5’ and 10’), i obtained range and modal frequency of 19 bioclimatic variables and 5 reference variables. then, i measured euclidean distances from candidate variables to the non-analogous variable as a metric for degree of analogous conditions, which were used to rank variables. bioclimatic layers were plotted in log-log scatterplots of range vs. modal frequency; variables were located inside the upper-right triangle (except for one set), and no layer fit the inverse model. temperature variables presented higher degrees of analogous conditions than precipitation for south america and the araucaria moist forests ecoregion. geographic extent and pixel resolution changed the degree of analogous conditions of derived variables (quarterly and monthly); however, a pattern of change was not observed, which suggested ad hoc hypotheses on geographic and temporal idiosyncrasies. variables with high contribution in previous sdm/enm studies (e.g., temperature seasonality and annual precipitation) showed low degree of analogous conditions. it is expected that heterogeneous layers would generate better correlational geographic distributional predictions than analogous variables, even though this hypothesis remains untested. ranking layers can provide grounds for selecting variables in distribution and niche modeling, particularly as regards interpreting spatial projection and transferability. alternatively, ranking can be used to compare degrees of analogous conditions of the same layer in different time spans. key words.—analog conditions, bioclimatic variables, environmental space, geographic space an environmental digital layer is a common object of geographic information systems, in which values of a continuous variable are stored in a spatially referenced matrix (chang 2017). in biogeography and macroecology, environmental layers include temperature, precipitation, humidity, radiation, soil, and human occupation. they can be used as a background in illustrating maps and as predictors in statistical analysis (williams et al. 2012). environmental layers are normally raster format objects, which implies some level of discretization of continuous space (hijmans and elith 2017). environmental variables are not distributed heterogeneously across space. variables like temperature are spatially autocorrelated, and show repeated or similar values over space (legendre 1993). additionally, combinations of variables can replicate more complex circumstances and cause localities to represent analogous conditions, for example, a monsoonlike climate in south america caused by heating and circulation regimes associated with topography (zhou and lau 1998). in the literature, ‘analogous’ or ‘analog’ has frequently been employed to designate equivalent climatic conditions through time. several studies have inferred how populations and communities responded to past (overpeck et al. 1992, jackson and overpeck 2000) and current climate change (garcia et al. 2014) by comparing to modern climates against analogous past and/or future climate scenarios. they have also estimated new and disappearing climatic combinations in scenarios of change (ohlemüller et al. 2006, williams et al. 2007, ackerly et al. 2010). 11 löwenberg-neto – analogous conditions contemporary analogous conditions, such as pixels with equal or similar values, provide grounds for species distribution modeling (guisan and zimmermann 2000). in a correlative approach, occurrences of species and digital environmental layers are used to estimate existing or realized niches of species (peterson et al. 2011). then, a given algorithm may search (elith and graham 2009) across the geographic extent for equivalent conditions (elith and leathwick 2009). an interesting topic relevant to correlational modeling is projection over space and the problem of non-analog climates (fitzpatrick and hargrove 2009). it is possible that native-range geographic range climate conditions are distinct from conditions in the invaded region (fitzpatrick et al. 2006, soberón and peterson 2011). in these cases, challenging the equilibrium assumption, including invasion stages (gallien et al. 2012), and controlling for possible niche shifts (guisan et al. 2014) in spatial projections may overcome the problem. correlational modeling procedures are based on the dual-space correspondence (peterson et al. 2011, soberón et al. 2017). dual-space correspondence, or hutchinson’s duality (colwell and rangel 2009), occurs when points from an input space (e.g. biotope space, geographic space) are plotted in a feature space (e.g. climatic space), where variables are axes and measures are coordinates (hutchinson 1978, colwell and rangel 2009). each input point g from geographic space (g) has a vector eg composed of measures of v variables, such as annual mean temperature, annual precipitation, and so on. the vector eg with v elements represents the coordinates of the point in a vdimensional space, the environmental space (e) (peterson et al. 2011). when enough variables are used at sufficient precision, points from g generally correspond one to one to points in e (apinall and lees 1994). however, this situation is not necessarily the case, and the same or very similar (analogous) climatic combinations may occur in separated geographic localities. in any case, the cloud of points can be interpreted as a particular realization of environmental conditions that occur across a given geographic extent at a particular time (i.e., the realized environment) (jackson and overpeck 2000). two additional features from e are of particular interest: (i) empty environmental spaces denote combinations of conditions that are missing, such as warm (>30°c) and wet (>3000 mm) climates in california (ackerly et al. 2010), and (ii) closelylocated points that are environmentally analogous localities (de oliveira et al. 2014). for a climatic variable, when many localities have the same value, their corresponding points pile up in a kernel in e, and represent high frequency regions (figure 1). in this case, when more than one locality has the same environmental-variable value, if a single point was mapped back, it would represent all pixels in g in an asymmetric relationship (one-to-many, a partial reciprocity, colwell and rangel 2009). on the other hand, a non-analogous (nonrepeating) variable in g has corresponding points (1:1) spread over e, with a maximum frequency of one. let n = |g| the number of points in the rasterization of a variable; y is the range of values of the variables (y = ymax –ymin) and x is the modal frequency of variable y. if the extent and resolution of the discretization of g do not change, then an inverse relationship (y ∝ x-1) between the range of a variable and its modal frequency is to be expected. in other words, if the range of a variable is small, most cells in the raster have values in that small range. if the range is broad, the distribution of frequencies will tend to be flat (figure 1). this effect occurs because (a) each variable’s range determines the span of its axis in e, and (b) the number of pixels is constant for a given extent and resolution, which provides a zero-sum scenario. hence, when the span is low, density will be concentrated in a small region, and kernel modal frequency will be high; when the breadth is high, density is spread over the axis and maximum frequency is low. this point is important because of hutchinson’s duality: the same niche breadths in regions of contrasting spans of values of environmental variables may predict contrasting sizes of areas of distribution. statistical selection of variables for species distribution modeling frequently aims at controlling variable collinearity and ranking variables (negrão and löwenberg-neto in prep.). current metrics for analogy of conditions 12 löwenberg-neto – analogous conditions figure 1. inverse relationship between variable range in g (geographic space) and modal frequency in e (environmental space). top row represents an absolutely analogous variable, zero-ranged, with modal frequency equal to number of pixels, and bipartite network asymmetric. middle row represents a layer with intermediate degrees. bottom row is a non-analogous layer: range equals the number of pixels minus one, modal frequency equals one; bipartite network is symmetric. figure 2. log-transformed scatterplots of variable range versus maximum frequency for two extents (south america, sa; araucaria moist forests, ar), and two pixel resolutions (2.5’, 10’). bioclimatic variables are labeled following hijmans et al. (2006), and referential variables as co = constant, sc = semi-constant and wide-range, rn = random normal, ht = heterogeneous, and nan = non-analogous. 13 löwenberg-neto – analogous conditions are available only in temporal frameworks and only for comparing pixels within layers (garcia et al. 2014), which does not allow comparison among layers. in the present paper, i present a layer-scoped metric that quantifies overall degree of analogy of environmental layers under hutchinson’s duality. i have then used the measurements to rank variables, and discuss the importance of these tools and ideas in the broader field of distributional ecology. methods i obtained 19 bioclimatic variables (hijmans et al. 2005), and calculated their ranges and modal values. measurements were developed for variables at two geographic extents: all of south america (sa) and the araucaria moist forests ecoregion (ar) in southern brazil (olson et al. 2002). for both extents, i analyzed bioclimatic variables at two resolutions: 2.5’ and 10’ (guisan and thuiller 2005). for each combination of extent and resolution, i created 5 reference variables: (1) constant variable (co) is a homogeneous variable, with a modal frequency equal to the number of pixels and zero for range. for the logtransformed distance (see below), i assigned variable range to one. (2) semi-constant, wide range (sc), is the second most homogeneous variable has a single, high modal frequency and a wide range with low frequencies. this variable is important variable because it controls for variables that are very homogeneous but that may mislead interpretation or metric quantification owing to its wide range of values. (3) random normal (rn) is a heterogeneous variable drawn from a normal distribution; its range is similar to the bioclimatic variable with the broadest range in all combinations, annual precipitation. (4) heterogeneous (ht) is the most heterogeneous variable, with a range similar to that of annual precipitation; it has repeating values with the lowest maximum frequency. finally, (5) non-analogous (nan) is the absolutely heterogeneous variable, with no repeating values, range equal to the number of pixels, and a modal frequency of one. i compiled variable ranges and modal frequencies into data matrices. for each dataset, i measured the euclidean distance matrix between rows using dist command and method = “euclidean” sqrt(sum((xi yi)^2)) in r version 3.5.0. the distance between a given variable to the non-analogous variable (nan) was used to quantify variable’s degree of analogous conditions; therefore, longer distances to nan denoted more homogeneous (analogous) variables. the same measurement procedure was done for a log-transformed data matrix (log distance). pearson’s correlations were calculated among distance, log distance, and secondary metrics, which included range, maximum frequency, and the shannon-weaver diversity index (shannon 2001). the last metric was calculated using the command diversity in the ‘vegan’ r package (oksanen et al. 2007). results variable histograms used to measure range and modal frequency are presented in appendix a; measurements are in appendix b. for each variable, range and modal frequency were logtransformed and plotted (figure 2). correlation analyses showed that distance and log distance were strongly positively correlated; log distance was negatively correlated with variable range (appendix c). for each combination of geographic extent and resolution, distances to the nan variable were used to compare variables. ranking showed that reference variables co and sc had the highest degree of analogous conditions while rn, ht and nan the least (figure 3a). statistical variables showed disparate degrees of analogous conditions (figure 3b): mean diurnal range, temperature annual range, and precipitation seasonality had their degree of analogy of conditions affected by geographic extents, whereas isothermality and temperature seasonality, which are standardized variables, were less affected. temperature variables presented higher degrees of analogy of conditions than precipitation in both regions (figure 3c and 3d). discussion a metric that quantifies overall degree of analogy of conditions for individual environmental layers was presented. by creating a non-analogous variable in which the range equals the number of pixels and modal frequency equals one, it was possible to plot and measure 14 löwenberg-neto – analogous conditions figure 3. ranking environmental variables by their euclidean distance to the non-analogous (nan) variable in a line graph for four combinations of extent and resolution (south america at 2.5’, south america at 10’, araucaria moist forests at 2.5’, araucaria moist forests at 10’). variables were displayed in four subsets: (a) reference, (b) statistical, (c) temperature, and (d) precipitation. bioclimatic variables were labeled following hijmans et al. (2006), and reference variables are as follows: co = constant, sc = semi-constant and wide-range, rn = random normal, ht = heterogeneous, and nan = non-analogous. 15 löwenberg-neto – analogous conditions the euclidean distance to the candidate variable as an index of dissimilarity to the non-analogous variable. in this sense, higher distances to nan denote that a variable has a high degree of analogous conditions. log-transformed plots provided a better visualization of the variables and their spatial positions in the kernel, and allow visualization of expected upper limits, which are based on the number of pixels the expected inverse relationship between variable range and modal frequency (figure 2). bioclimatic layers were located inside the right-angled triangle, and no layer fit the inverse model, which was expected for climatic variables. reference variables, especially sc and ht, showed consistent positions in all plots, providing internal references for the bioclimatic variables; conversely, rn was very close to realistic variables. three variables were placed outside the triangle envelope, beyond the vertical axis. this effect occurred because variables had a wider range than the numbers of pixels available. the ar10 treatment comprised 684 pixels, and ranges were above 1000 for temperature seasonality, annual precipitation, and rn. secondary metrics were not tested formally because i intended to provide a metric based on the duality ontology. nevertheless, their correlations with euclidean distance provided some information. for example, variable range was strongly negatively correlated with distance, which supports an exploratory approximation to analogous degree with no need for developing scatterplot. modal frequencies were less correlated with distance in treatments with finer pixels; diversity index showed a weak relation to distance; the index did not discern the semiconstant wide-range variable well. ranking bioclimatic layers showed that geographic extent and pixel resolution both affect the degree of analogy of conditions. it was expected that a change in grain size would not severely affect the degree of analogy of conditions (guisan et al. 2007). for the same extent, ranking showed that pixel resolution modified a variable’s ranking position, even though it showed no pattern of increasing analogy of conditions with decreasing resolution or vice-versa. regarding geographic extent, it was expected that extents limited the ranges of environmental variables (thuiller et al. 2004, randin et al. 2006), and that the smaller geographic extent would present higher degrees of analogy of conditions (anderson and raza 2010). this effect was observed only for a few variables, including statistical ones constructed by consideration of ranges of values (temperature diurnal and annual ranges) and coefficients of variation (precipitation seasonality). a third expectation was that variables arranged in temporal slices (quarterly and monthly) would show increasing degrees of analogy of conditions when compared to annual variables. in fact, annual mean temperature showed consistent ranking positions (8, 8, 8, 6, fig. 3c), whereas time-sliced variables showed changeable positions (e.g., mean temperature of wettest quarter, ranks 5, 6, 14, 13). the same effect was observed for precipitation variables (e.g., annual precipitation, ranks 21, 20, 21, 21; precipitation of wettest quarter, ranks 18, 21, 12, 14); however, i did not observe any general trend of increasing degree by decreasing temporal slice size. in sum, geographic extent and pixel resolution changed the degree of analogous conditions of derived variables whereas annual variables tended to maintain their rankings. no consistent trend of change between extension/resolution and increasing degree of analogous conditions was recognized, which suggested ad hoc hypotheses for geographic and temporal idiosyncrasies. for the purpose of species distribution modeling (sdm), variables showing high degrees of analogy of conditions tend to estimate broad geographic ranges and few values are frequent across geographic space (peterson et al. 2011). for a given range or niche-breadth, sdm models using variables close to the upper-left side of the triangle would predict larger geographic expanses than those in the lowerright part of the triangle. this observation thus offers a cautionary note for studies relating niche breadth to distributional area without controlling for variable degrees of analogy of conditions (slatyer et al. 2013). in fact, this statement depends on each band of the variable histogram having a correlation with occurrences of species (guisan and thuiller 16 löwenberg-neto – analogous conditions 2005). in any case, a study that summarized environmental variables that most contributed to estimating species’ geographic distributions showed that, for the worldclim dataset, temperature seasonality, annual precipitation, and precipitation of the driest month had highest mean contributions (bradie and leung 2017). interestingly, in this study, the former two variables were consistently ranked as showing low degree analogy for both extents and resolution; precipitation of the driest month showed an atypical trend, with high analogy for south america and low analogy for araucaria moist forests (fig. 3d), perhaps owing to odd contrasts in this variable in homogeneity across the two regions (appendix a). further, it is common in processing raster layers to transform decimal values into integers by multiplying by 10, 100, or 1000 (hijmans et al. 2005), which produces increasing variable heterogeneity. it is also common to use raster layers arranged into categorical, nominal, or ordinal classes (peterson et al. 2011), which dramatically homogenizes variables. increasing or decreasing numbers of bins on the environmental axis affects modal frequencies in the e-space kernel and therefore the degree of analogy of conditions of environmental layers. in this paper, i have focused on quantification of analogy of conditions within the same variable layer as ‘contemporary’ analogous conditions. this quantification approach can be used to compare degrees of analogy of conditions in different time spans (garcia et al. 2014). for a given variable with constant extent and resolution, the variable can be plotted for different spans, and degrees of analogous conditions in temporal scenarios of change can be ranked. by analyzing how analogous conditions were coded in g and e, i used two parameters to characterize degree of analogy of conditions. the euclidean distance between the candidate layer and the non-analogous layer provided a metric of dissimilarity used to rank and compare variables by their degree of analogy. the resulting information may be used to select layers and interpret results in species distribution and ecological niche modeling. 1 doi: 10.1002/9781118786352.wbieg0152. acknowledgments i am grateful to l.r.r. faria jr., t. vasconcelos, j. soberón, and town peterson, for suggestions that improved the manuscript. this research was conducted in the biogeography and macroecology lab (ilacvn/unila/brazil). references ackerly, d.d., s.r. loarie, w.k. cornwell, s.b. weiss, h. hamilton, r. branciforte, and n.j.b. kraft. 2010. the geography of climate change: implications for conservation biogeography. diversity and distributions, 16:476-487. aspinall, r. and b.g. lees. 1994. sampling and analysis of spatial environmental data. proceedings of the sixth international symposium on spatial data handling, vol 2. edinburgh. anderson, r. p., and raza, a. 2010. the effect of the extent of the study region on gis models of species geographic distributions and estimates of niche evolution: preliminary tests with montane rodents (genus nephelomys) in venezuela. journal of biogeography, 37:1378-1393. bradie, j., and b. leung. 2017. a quantitative synthesis of the importance of variables used in maxent species distribution models. journal of biogeography, 44:1344-1361. chang, k.-t. 2017. geographic information system. the international encyclopedia of geography.1 colwell, r.k. and t.f. rangel. 2009. hutchinson's duality: the once and future niche. proceedings of the national academy of sciences usa 106:19651-19658. elith, j., and graham, c. h. 2009. do they? how do they? why do they differ? on finding reasons for differing performances of species distribution models. ecography 32:66-77. elith, j., and j.r. leathwick. 2009. species distribution models: ecological explanation and prediction across space and time. annual review of ecology, evolution, and systematics 40:677697. fitzpatrick, m.c., and w.w. hargrove. 2009. the projection of species distribution models and the problem of non-analog climate. biodiversity and conservation 18:2255-2261. fitzpatrick, m.c., j.f. weltzin, n.j. sanders, and r.r. dunn. 2007. the biogeography of prediction error: why does the introduced range of the fire ant over-predict its native range? global ecology and biogeography 16:24-33. gallien, l., r. douzet, s. pratte, n.e. zimmermann, and w. thuiller. 2012. invasive species 17 löwenberg-neto – analogous conditions distribution models—how violating the equilibrium assumption can create new insights. global ecology and biogeography 21:1126-1136. garcia, r. a., m. cabeza, c. rahbek, and m.b. araújo. 2014. multiple dimensions of climate change and their implications for biodiversity. science 344:1247579. guisan, a., c.h. graham, j. elith, and f. huettmann. 2007. sensitivity of predictive species distribution models to change in grain size. diversity and distributions 13:332-340. guisan, a., b. petitpierre, o. broennimann, c. daehler, and c. kueffer. 2014. unifying niche shift studies: insights from biological invasions. trends in ecology and evolution 29:260-269. guisan, a., and n.e. zimmermann. 2000. predictive habitat distribution models in ecology. ecological modelling 135:147-186. guisan, a., and w. thuiller. 2005. predicting species distribution: offering more than simple habitat models. ecology letters 8:993-1009. hijmans, r.j., s.e. cameron, j.l. parra, p.g. jones, and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25:19651978. hijmans, r. and j. elith. 2017. species distribution modeling with r.2 hutchinson, g.e. 1978. an introduction to population ecology. yale university press, new haven. jackson, s.t. and j.t. overpeck. 2000. responses of plant populations and communities to environmental changes of the late quaternary. paleobiology 26:194-220. legendre, p. 1993. spatial autocorrelation: trouble or new paradigm? ecology 74:1659-1673. de oliveira, g., t.f. rangel, m.s. lima-ribeiro, l.c. terribile, and j.a.f. diniz-filho. 2014. evaluating, partitioning, and mapping the spatial autocorrelation component in ecological niche modeling: a new approach based on environmentally equidistant records. ecography 37:637-647. oksanen, j., f.g. blanchet, m. friendly, p. kindt, p. legendre, d. mcglinn, p.r. minchin, r.b. o'hara, g.l. simpson, p. solymos, m.h.h. stevens, e. szoecs, and h. wagner. 2007. community ecology package: ordination, diversity and dissimilarities, the vegan package for r, version 2.4.3 olson, d.m., e. dinerstein, e.d. wikramanayake, n.d. burgess, g.v. powell, e.c. underwood, j.a. d’amico, i. itoua, h.e. strand, j.c. morrison, c.j. loucks, t.f. allnutt, t.h. ricketts, j.f. 2 https://cran.r-project.org/web/packages/dismo/vignettes/sdm.pdf. lamoreux, w.w. wettengel, p. hedao, and k.r. kassem. 2001. terrestrial ecoregions of the world: a new map of life on earth. bioscience 51:933938. overpeck, j.t., r.s. webb, and t. webb. 1992. mapping eastern north american vegetation change of the past 18 ka: no-analogs and the future. geology 20:1071-1074. peterson, a.t., j. soberón, r.g. pearson, r.p. anderson, e. martínez-meyer, m. nakamura, and m.b. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. randin, c.f., t. dirnböck, s. dullinger, n.e. zimmermann, m. zappa, and a. guisan. 2006. are niche-based species distribution models transferable in space? journal of biogeography 33:1689-1703. shannon, c. e. 2001. a mathematical theory of communication. acm sigmobile mobile computing and communications review 5:3-55. slatyer, r.a., m. hirst, and j.p. sexton. 2013. niche breadth predicts geographical range size: a general ecological pattern. ecology letters 16:1104-1114. soberón, j. and m. nakamura. 2009. niches and distributional areas: concepts, methods, and assumptions. proceedings of the national academy of sciences usa 106:19644-19650. soberón, j., l. osorio-olvera, and a.t. peterson. 2017. diferencias conceptuales entre modelación de nichos y modelación de áreas de distribución. revista mexicana de biodiversidad 88:437-441. soberón, j., and a.t. peterson. 2011. ecological niche shifts and environmental space anisotropy: a cautionary note. revista mexicana de biodiversidad 82:1348-1355. thuiller, w., l. brotons, m.b. araújo, and s. lavorel. 2004. effects of restricting environmental range of data to project current and future species distributions. ecography 27:165172. williams, k.j., l. belbin, l., m.p. austin, j.l. stein, and s. ferrier. 2012. which environmental variables should i use in my biodiversity model? international journal of geographical information science 26:2009-2047. zhou, j. and k.-m. lau. 1998. does a monsoon climate exist over south america? journal of climate 11:1021-1040. 3 https://cran.r-project.org/web/packages/vegan/vegan.pdf. 18 löwenberg-neto – analogous conditions appendix 1: histograms of bioclimatic variables. figure a.1 bioclimatic variables and reference variables presented in histograms for the extent of all of south america at 2.5’ resolution, with 883,760 pixels. 19 löwenberg-neto – analogous conditions figure a.2 bioclimatic variables and reference variables presented in histograms for the extent of all of south america at 10’ resolution, with 55,377 pixels. 20 löwenberg-neto – analogous conditions figure a.3 bioclimatic variables and reference variables presented in histograms for the extent of the araucaria moist forest ecoregion at 2.5’ resolution, with 11,033 pixels. 21 löwenberg-neto – analogous conditions figure a.4 bioclimatic variables and reference variables presented in histograms for the extent of the araucaria moist forest ecoregion at 10’ resolution, with 684 pixels. 22 löwenberg-neto – analogous conditions appendix 2: variable measurements. table b.1. measurements for the extent of all of south america at a 2.5’ spatial resolution. bioclimatic and reference variables: co = constant, sc = semi-constant and wide range, rn = random normal, ht = heterogeneous, nan = non-analogous. variable range modal frequency diversity distance to nan log distance bio1 399 87311 13.6680 1087163.5 7.3081 bio2 202 33077 13.6732 1082891.1 7.1080 bio3 80 37357 13.6732 1083249.1 7.4754 bio4 6283 51005 13.3247 1076499.4 6.3375 bio5 355 88750 13.6736 1087391.9 7.3503 bio6 458 47793 13.6567 1083401.9 7.0026 bio7 286 48816 13.6397 1083680.7 7.1585 bio8 434 118219 13.6706 1091494.6 7.4176 bio9 380 55925 13.6552 1084081.0 7.1281 bio10 371 105921 13.6729 1089675.5 7.4149 bio11 435 79219 13.6589 1086189.6 7.2395 bio12 10577 42407 13.5000 1070686.8 6.1368 bio13 1197 34742 13.5276 1081751.6 6.5779 bio14 697 182027 13.0964 1104264.7 7.4798 bio15 259 26691 13.5697 1082557.0 6.9363 bio16 3450 59923 13.5211 1080650.0 6.5528 bio17 2319 149545 13.1747 1094966.9 7.0824 bio18 2574 38118 13.4865 1080237.2 6.4129 bio19 2962 198539 13.1451 1105818.1 7.1615 co 1 883760 13.6919 1530715.5 10.2994 sc 10600 873260 13.6896 1512443.2 7.6473 rn 17787 39148 13.4992 1061679.1 5.9963 ht 10600 8400 13.4977 1069447.7 5.3512 nan 883760 1 13.4988 0.0 0.0000 table b.2. measurements for the extent of all of south america at a 10’ spatial resolution. bioclimatic and reference variables: co = constant, sc = semi-constant and wide range, rn = random normal, ht = heterogeneous, nan = non-analogous. variable range modal frequency diversity distance to nan log distance bio1 355 5460 10.86252 67718.76 5.3070 bio2 167 2083 10.90318 67666.22 5.1040 bio3 55 2352 10.90312 67816.49 5.5296 bio4 6214 3221 10.55476 60341.14 4.4511 bio5 303 5531 10.90393 67790.77 5.3560 23 löwenberg-neto – analogous conditions bio6 414 2999 10.75527 67415.71 4.9916 bio7 280 1243 10.86969 67496.91 4.7193 bio8 395 7300 10.85650 67929.69 5.4130 bio9 340 3527 10.84392 67544.47 5.1200 bio10 318 6600 10.88888 67915.83 5.4236 bio11 394 5013 10.80168 67619.34 5.2398 bio12 9916 2665 10.73032 55773.64 4.2942 bio13 980 2199 10.75779 66676.81 4.6217 bio14 652 11355 10.32861 68451.50 5.4999 bio15 232 1657 10.79936 67569.00 4.9018 bio16 2787 1592 10.75131 64438.80 4.2316 bio17 2159 9357 10.40633 66178.05 5.1607 bio18 2427 2418 10.71755 64917.76 4.4653 bio19 2787 5381 10.37613 64745.49 4.8381 co 1 55377 10.92192 95914.04 8.2157 sc 9915 36308 10.85571 71256.49 5.6593 rn 23713 2443 10.72992 38895.48 4.1738 ht 9915 600 10.71872 55684.18 3.5234 nan 55377 1 10.72878 0.00 0.0000 table b.3. measurements for the extent of the araucaria moist forest ecoregion at a 2.5’ spatial resolution. bioclimatic and reference variables: co = constant, sc = semi-constant and wide range, rn = random normal, ht = heterogeneous, nan = non-analogous. variable range modal frequency diversity distance to nan log distance bio1 93 426 9.3029 13372.10 4.1008 bio2 70 464 9.3002 13402.14 4.2311 bio3 19 1397 9.3044 13560.81 5.1271 bio4 1092 302 9.3039 12144.04 3.2765 bio5 102 312 9.3044 13356.38 3.9409 bio6 85 365 9.2870 13379.19 4.0669 bio7 83 351 9.3035 13381.08 4.0589 bio8 125 284 9.2944 13327.28 3.8341 bio9 133 237 9.2950 13316.11 3.7384 bio10 98 409 9.3039 13365.19 4.0666 bio11 86 387 9.3014 13378.89 4.0870 bio12 1176 313 9.2993 12041.63 3.2797 bio13 139 259 9.3007 13309.38 3.7606 bio14 124 261 9.2798 13327.80 3.8017 bio15 40 1023 9.2325 13485.09 4.7450 bio16 397 587 9.3020 13009.46 3.8236 bio17 406 385 9.2853 12987.14 3.6204 24 löwenberg-neto – analogous conditions bio18 347 590 9.3013 13070.80 3.8596 bio19 402 317 9.2873 12989.29 3.5329 co 1 11003 9.3059 19056.02 7.0001 sc 1200 9813 9.2849 16987.09 5.0290 rn 1200 453 9.1950 12018.93 3.4600 ht 1200 100 9.1034 12006.79 2.7183 nan 11003 1 9.1128 0.00 0.0000 table b.4. measurements for the extent of the araucaria moist forest ecoregion at a 10’ spatial resolution. bioclimatic and reference variables: co = constant, sc = semi-constant and wide range, rn = random normal, ht = heterogeneous, nan = non-analogous. variable range modal frequency diversity distance to nan log distance bio1 72 30 6.5251 750.38 2.1695 bio2 65 36 6.5223 759.33 2.2804 bio3 18 86 6.5265 822.30 3.0589 bio4 1002 22 6.5260 390.32 1.6566 bio5 76 27 6.5265 745.33 2.1069 bio6 73 28 6.5100 749.05 2.1349 bio7 73 35 6.5255 749.48 2.2344 bio8 108 22 6.5167 705.92 1.9150 bio9 92 21 6.5170 725.46 1.9393 bio10 75 30 6.5261 746.71 2.1576 bio11 67 30 6.5237 756.50 2.1909 bio12 1062 21 6.5215 463.74 1.6362 bio13 104 20 6.5229 710.73 1.8822 bio14 117 18 6.5021 694.74 1.8016 bio15 39 67 6.4550 794.09 2.7061 bio16 307 46 6.5242 465.01 2.0806 bio17 379 28 6.5075 375.01 1.8000 bio18 274 19 6.5234 502.63 1.6400 bio19 377 26 6.5096 377.24 1.7617 co 1 684 6.5280 1182.99 4.9105 sc 230 456 6.4635 787.22 3.3077 rn 1062 19 6.3412 463.48 1.5835 ht 230 6 6.3369 556.07 1.1155 nan 684 1 6.3355 0.00 0.0000 25 löwenberg-neto – analogous conditions appendix 3: pearson correlation coefficients. table c.1. correlation coefficients for the parameters of two extents (south america sa, and araucaria moist forests ecoregion ar) and two spatial resolutions (2.5’, 10’). nan = non-analogous variable. sa2.5 modal frequency diversity distance to nan log distance range -0.1226 -0.0497 -0.8786 -0.8591 modal frequency 1 0.1363 0.5759 0.4805 diversity 0.1363 1 0.1488 0.1713 distance to nan 0.5759 0.1488 1 0.9259 sa10 range -0.1218 -0.0818 -0.9207 -0.8089 modal frequency 1 0.1520 0.4938 0.6196 diversity 0.1520 1 0.1674 0.1884 distance to nan 0.4938 0.1674 1 0.9445 ar2.5 range -0.0782 -0.6920 -0.7546 -0.8806 modal frequency 1 0.1317 0.6270 0.5536 diversity 0.1317 1 0.6466 0.6317 distance to nan 0.6270 0.6466 1 0.9278 ar10 range -0.204 -0.1892 -0.6936 -0.4737 modal frequency 1 0.0608 0.5563 0.8129 diversity 0.0608 1 0.4542 0.4736 distance to nan 0.5563 0.4542 1 0.8559 26 microsoft word 4914-9215-2-sm_final.docx biodiversity informatics, 10, 2015, 45-55     knowledge of diversity of wild palms (arecaceae) in the republic of benin: finding gaps in the national inventory by combining field and digital accessible knowledge rodrigue idohou1,2*, arturo h. ariño3, achille ephrem assogbadjo2, romain glele kakai1, brice sinsin2 1laboratory of biomathematics and forest estimations, university of abomey-calavi, benin. 2laboratory of applied ecology, faculty of agronomic sciences, university of abomey-calavi, 01 bp 526, cotonou, benin. 3department of environmental biology, university of navarra, pamplona, spain. *corresponding author: rodrigidohou@gmail.com. abstract.—despite efforts by researchers worldwide to assess the biodiversity of plant groups, many locations on earth remain poorly surveyed, resulting in inadequate or biased knowledge. robust estimates of inventory completeness could help alleviate the problem. this study aimed to identify areas representing gaps in current knowledge of african palms, with a focus on benin (west africa). we assessed the completeness of knowledge of african palms, targeting geographic distance and climatic difference from well-known sites. data derived from intensive fieldwork were combined with independent data available online. inventory completeness indices were calculated and coupled with other criteria. results showed a high overall value for inventory completeness, as well as an even distribution of well-known areas across the country. however, poorly-known areas were identified, which were in remote locations with low accessibility. this study illustrates how biodiversity survey and inventory efforts can be guided by existing knowledge. we strongly recommend the combination of digital accessible knowledge and fieldwork, coupled with expert knowledge, to obtain a better picture of inventory completeness in tropical ecosystems. key words.—biological databases, gis, inventory, sampling efficiency, spatial resolution. one of the greatest challenges that tropical biologists are facing now is how to conserve biological diversity in the current context of demographic pressure, increase of needs, overexploitation, climate change, and economic crisis (fao 2010). under these threats, without effective protection, much of tropical biodiversity is unlikely to survive, so strategies to promote its conservation are needed (bruner et al. 2001). measurements of biological diversity can provide baseline information on distribution, richness, and relative abundance of taxa that is required for taking appropriate conservation decisions (humphries et al. 1995; may 1988; magurran 1988; raven and wilson 1992). the national flora of benin is estimated at 2807 species (akoègninou et al. 2006). some of those species are of high socioeconomic importance and have been studied in depth. however, others remain not well assessed, such as wild palm species. wild palms are amongst the most diverse plant groups in the world (tomlinson 1990) and are species with significant cultural, social, economic, and ecological uses (monteiro et al. 2006). they serve as bio-indicators in many latin-american countries (kjaeret al. 2004; vormisto et al. 2004), and their occurrences could be used as climate trend proxies. in sub-saharan africa, and especially in benin, wild palms are not well documented. the species diversity is not well known, and ecological studies are rare. these data, together with a complete richness inventory, are nonetheless critical to planning informed conservation actions. many studies now exist on the use of primary biodiversity data that are both digital and accessible in standard formats (graham et al. 2004; guralnick et al. 2007; sousa-baena et al. 2014) providing access to more than 6 x108 data records. the magnitude of digital accessible knowledge (dak) is large though perhaps not sufficient when measured against global biodiversity (sousa-baena et al. 2014). in contrast, cases of use of extensive fieldwork data not obtained from online data portals (i.e., requiring time-consuming, expensive field surveys) are less frequent. in addition, assessing sampling effort across geographic space requires an understanding of how species assemblages differ among different environments, across biogeographic barriers, and as a result of townpeterson typewritten text 45 biodiversity informatics, 10, 2015, 45-55     dispersal limitation. species accumulation curves, species richness estimates, and diversity accumulation curves have been used to determine the level of survey completeness (thompson et al. 2007; ariño et al. 2008; de thoisy et al. 2008; aranda et al. 2010; lovell et al. 2010). the measured level of completeness can be compared to the desired level of completeness for the same locality, and some authors have defined particular targets that may be broadly appropriate (cardoso et al. 2009). several statistics are available for calculating species richness estimates, including non-parametric methods and extrapolations of species accumulation curves, that vary in their accuracy under different conditions, often having drawbacks that may prevent their use in common circumstances (e.g. low species density). other less well-known methods have been proposed to try to overcome some of these challenges, such as the generalization by ariño (2010) of the probability theory developed by seber (1982).these novel methods may help determining the completeness of the inventory and bring out gaps in sampled areas for further documentation (chao and jost 2012). we carried out this study on both available dak and extensive fieldwork inventory of wild palms (i) to describe the national species richness of this group, and (ii) to estimate the completeness of the inventory within the group. we assessed knowledge gaps across benin through estimation of geographic and environmental distances to wellknown localities. methods study area benin is a west african country located between 6°20’ and 12°25’n and 1° and 3°40’e. biogeographically, benin is subdivided into three contrasting phytochorological zones: the guineocongolean zone, the sudano-guinean transition zone and the sudanian zone (akoègninou et al. 2006; white 1983). rainfall is bimodal in the guineo-congolean zone. north of this zone, rainfall distribution becomes unimodal. human activities have resulted in a high level of degradation of the vegetation (figure 1). data sources our analyses are based on data from both extensive fieldwork carried out from may 2013 to april 2014, during which a megatransect covering the whole country was executed, comprised of daily transects, and data downloaded from the global biodiversity information facility1,   comprising data on 11 wild palm species (8 observed during our fieldwork, and 3 additional species appearing in the gbif dataset). the gbif search was done in january 2015 through the use of key fields such as palms, arecaceae, african palms, african native palms, borassus, eremo-spatha, hyphaene, laccosperma, phoenix, raphia, rattan, raffia, wild palms, etc. the initial dataset contained 1847 records from the two sources. the dataset was then cleaned via a series of inspections and visualizations designed to detect and document inconsistency, as follows. (1) we created lists of unique names in each dataset in microsoft excel, and manually inspected them for repeated versions of the same taxonomic concepts: misspellings, name variants, different versions of authority information, etc. such repeated name variants were flagged, checked via independent sources, and corrected to produce unique scientific names that we believed correctly referred to single taxa. (2) we checked for geographic coordinates that fell outside of the country, but which were referred to benin. (3) within the country, we checked for consistency between descriptions of district and position of geographic coordinates. in each case, where possible, we created a corrected version of the data record; where no clear correction was possible, we discarded data, recording data losses at each step in the cleaning process. in all, 1375 records were finally considered (1154 fieldwork + 221 gbif records; figure 2) which were constrained also to include only those with consistent coordinates. data analysis and interpretation we aggregated point-based occurrence data to ½° spatial resolution across the country, which near the equator corresponds to a square ~56 km on a side (figure 2). this spatial resolution was the product of a detailed analysis of balancing the benefits of aggregating data (i.e., larger sample sizes), versus the loss of spatial resolution that accompanies broader aggregation areas that can make imperceptible important geographic features. the procedure consists on examining the relative change in area-adjusted variance of the data versus                                                                                                                 1  http://www.gbif.org.     townpeterson typewritten text 46 biodiversity informatics, 10, 2015, 45-55     figure 1. geographic pattern of benin’s biogeographic zones (sudanian, sudano-guinean, and guineo-congolean) and soil types. townpeterson typewritten text 47 biodiversity informatics, 10, 2015, 45-55     figure 2. elevation map of benin, with the geographic locations of records of palms collected in the field and downloaded through gbif. grid squares delimit the ½° cells used to calculate completeness. townpeterson typewritten text 48 biodiversity informatics, 10, 2015, 45-55     increasing plot size, and selecting the smallest plot size at which the trend of the slope of the overall variance vs. area curve changed most. the concept is similar to selecting the largest sample size beyond which no significant increase in diversity is expected (ariño et al. 2008), and the resulting quadrat size was consistent with the spatial resolution used by sousa-baena et al. (2014) in their analysis of brazilian plant diversity and presentation of the idea of dak. we produced the ½° grid shapefiles in the vector grid module of qgis, version 2.62. next, we attributed each data record to the corresponding grid cell, and used a set of criteria on the aggregated number of records per cell, per taxon, to consider whether each cell was well-sampled. we calculated (1) the total number of records available from each grid square (termed n); (2) the number of distinct species recorded from each grid square (sobs) for species appearing exclusively within field data, exclusively as gbif records, and species recorded both as field data and gbif data records; and (3) the number of species whose occurrence was recorded exactly once (a) and exactly twice (b) at each grid cell. with that information we were able to use chao’s (chao et al. 2000) formula to calculate the corresponding expected number of species (sexp in chao’s work, which we will denote sc here) for all three cases: 𝑆! = 𝑆!"# + 𝑎! 2𝑏 then defined inventory completeness (c) according to chao as cc = sobs / sc. in addition, the probability theory developed by seber (1982) originally applied to the problem of recognizing how many tagged animals had lost their marks in a recapture experiment, and later generalized by ariño (2010) for estimating the number of missing data records from any number of overlapping datasets, was applied here. as the fieldwork data had not been already shared with available data from gbif, the total number of species existing in the study area would be: 𝑆! = 𝑆!"#$% + 𝑆!"#$ + 𝑆!"#$%∗!"#$ + 𝑆!                                                                                                                 2  http://www.qgis.org.     where sfield*gbif is the number of recorded species shared in both datasets, sfield is the number of species recorded in field data but not in gbif, sgbif the number of species recorded in gbif data but not in field data, and s0 the unknown number of species that weren’t recorded in either collection. sp cannot thus be known but it can be estimated (seber 1982) by probability theory on the intersection of the corresponding independent datasets (ariño 2010). in our case, with two datasets, the estimate is 𝑆! = ! !!! (𝑆!"#$% + 𝑆!"#$ + 𝑆!"#$%∗!"#$), where 𝑘 = !!"#$∗!!"#$% !!"#$!!!"#$%∗!"#$ (!!"#$%!!!"#$%∗!"#$) . completeness could then be calculated as cp = sobs/integer(sp) for samples not so small as to introduce large bias in k due to the estimates of the multinomial function used to derive it (seber and felton, 1981; “integer” indicates the whole number part of a real number). we then explored plots of cx versus n to assess appropriate and adequate definitions of relatively completely versus incompletely inventoried grid squares. as many cells either were not amenable to estimating sp for want of at least one parameter (commonalities or exclusivities), or would not yield sc for want of singletons (a) or doubletons (b), we decided to combine both approaches to derive completeness rather than rely solely on either chao’s or ariño’s approaches whenever possible. we decided to use a highly conservative criterion by estimating completeness on the highest available value for expected species (either sc or sp), when both could be calculated. thus, we obtained a lower limit for our completeness estimate as: 𝐶! = 𝑆!"#/𝑀𝐴𝑋 𝑆! , 𝑆! . we then classified each of the squares according to the completeness criteria. we deemed a square to be well-sampled if any of these was true: (1) cc > 0.5 and singletons/doubletons available, (2) cp > 0.5 and k available, or (3) expert judgment, based on the known sampling density. townpeterson typewritten text 49 biodiversity informatics, 10, 2015, 45-55     next, in qgis, we linked the table with the grid square statistics (i.e., sobs, sc, sp, cm, cc, cp) to the aggregation grid, and saved this file as a shapefile. the shapefile of well-sampled grid squares was further converted to raster (geotiff) format using custom scripts in r (r development core team 2013). this raster coverage was the basis for our identification of gaps, as follows. we used the proximity (raster distance) function in qgis to summarize geographic distance to any well-sampled area. to create a parallel view of environmental difference from wellsampled areas, we plotted 5000 random points across the country, and used the point sampling tool in qgis to link each point to the geographic distance raster, and to raster coverages (2.5’ resolution) summarizing annual mean temperature and annual precipitation drawn from the worldclim climate data archive (hijmans et al. 2005). we exported the attributes table associated with the random points, and analyzed it further in excel. we standardized values of each environmenttal variable to the overall range of the variable as (xi – xmin) / (xmax xmin), where xi is the particular observed value in question. we then calculated the environmental distance matrix by obtaining the euclidean distances for the climate variables between points falling in well-sampled cells (by definition, points having a geo-graphic distance of zero) and the points falling in the remaining cells (those points with non-zero geographic distances). hence, each random point in incomplete cells was defined by its distance in environmental space to the points in well-sampled cells. finally, the environmental distances were imported into qgis, and linked to the random points. the shapefile containing the random points was thus given a zvalue that is the environmental distance associated with that point. this vector file was then rasterized to provide continuous coverage across the region. results preliminary the raw data show a greater concentration of wild palms records in the northwestern and southernmost part of benin. however, data covered the whole country and did not appear to be particularly concentrated along points of access such as roads or rivers. inspecting the relationship between sobs and various c values, we observed a variation of outputs. for cc, completeness greater than 0.8 was observed for more than 10 expected species and a number of individuals between 0-50; for other completeness indices, more variation was observed in c values (figure 3). by definition, cells for which cm could not be calculated were declared as under-sampled. well-sampled areas according to the criteria defined in the methods could in turn be segregated into complete, with all species observed (either valid cm =1 or by expert judgment), and incomplete (0.5 0.8 and 21% had 0.5 < cm < 0.8 (figure 3). one-third of all cells covering the country were finally declared as under-sampled. table 1. decision table for levels of knowledge of palm species across benin. code frequency percent decisions 0 18 37 under-sampled: cm either <0.5 or cannot be calculated, and expert judgment not available 1 18 37 complete (cm = 1) 2 2 4 complete (based on expert judgment) 3 11 22 incomplete (sp or sc >sobs and cm valid) townpeterson typewritten text 50 biodiversity informatics, 10, 2015, 45-55     figure 3. completeness of inventories of palm trees in benin at ½°spatial resolution and environmental and geographic distances from well-sampled cells. bar diagrams in each cell represent the number of observed species (sobs), and upper limit for expected species (sexp), classified by completeness criteria. the density of shades represent the combination of the environmental distance (reds) and geographic distance (blues) between the shaded area and the well-sampled cells. townpeterson typewritten text 51 biodiversity informatics, 10, 2015, 45-55     we analysed completeness values for the ½° resolution which showed completeness evenly distributed across biogeographic zones in benin based on the geographic distance map. the environmental distance was higher in lowlands beyond the atacora mountains (northwestern part of the country), in remote areas, and in border regions. comparison of soil types in the country revealed variability across biogeographic zones (figure 1). ferralitic soils are found mostly in the southern part (and in the northeast corner) whereas more ferruginous soils are found northwards. wellknown areas covered much of these soil types, suggesting completeness of inventory of palm communities on soil types. homogenous, relatively stable climatic conditions were encountered across most ecosystems in the country. annual mean temperature was between 25-29oc and annual precipitation between 700-1300 mm. regions with higher temperature generally had lower precipi-tation and vice versa (figure 4). however, some areas were environmentally different, especially above 10o n. we calculated distances to well-known cells in climate space, which turned out to be roughly comparable to geographic distances to well-known sites (figure 3). combining both distances, we produced a view of areas that seem both poorly known and are both geographically remote and environmentally different from well-known sites (figure 3). discussion this study represents a first attempt in characterizing completeness of knowledge of palm community composition across benin through the use of data from our fieldwork and data available to the broader scientific community. inventory completeness was high across the country and most of the country’s ecosystems hosting palm species are thus well sampled. as such, the current state of the inventory of wild palms across ecosystems in benin is considered reasonably complete. this result comes from the concordance of the findings from different estimates. although palm records were more concentrated in some areas, these higher concentrations were not linked to accessibility features (e.g., roads), as has often been described in whole-region biodiversity studies (e.g., escala et al. 1997). however, the opposite was not true: the few sections where low completeness existed did not lack records, but often coincided with remote areas having a low density of roads or other access points, or being otherwise harder to reach. some of these incomplete sectors had also high environmental distances to well-sampled areas, and constituted gaps in sampling and knowledge. figure 4. scatterplot of precipitation vs. temperature at 5000 random points across benin, classified according to the relative geographic distance to well-sampled cells. townpeterson typewritten text 52 biodiversity informatics, 10, 2015, 45-55     gap areas are places that have not been well sampled (kier et al. 2005; stehmann 2009). gap analyses mostly focus on a particular taxon and its distribution and diversity across regions, ecoregions, or biomes (mora et al. 2008). meanwhile, it is important to know how complete areas of inventories are, in order to apply appropriate levels of confidence (colwell and coddington 1994). for benin, gaps resulting from wild palm inventory assessment are located in remote areas, such as in mountainous regions. these areas have not been previously mentioned as hosting palm biodiversity (akoègninou et al. 2006). the dryness of the climate could also explain the rarity of palm species in these areas. contrary to the findings of soria-auza and kessler (2008), palm diversity assessment in benin was not influenced by uneven collecting effort. the current study was based on intensive fieldwork through different seasons with the help of knowledgeable local people in the field. in addition, wild palms are recognizably distinct species, with little room for identification error: the species have long been described and few taxonomic misidentifications have been reported. as such, taxonomic bias is not likely to have affected the inventory, contrary to situations for other taxa (soberón et al. 2000; pyke and ehrlich 2010). the value of sharing data has been recognized for some time (nelson 2009). earlier, data were often safely and jealously kept by their owner (be it an individual, laboratory, or museum) and could only be accessed through remuneration of some sort, e.g. authorship (scoble 2000; ponder et al. 2001; wang et al. 2007). however, recent advances in information technology and an increased willingness to share primary biodiversity data are enabling unprecedented access (soberón and peterson 2004), as in case of gbif. this seachange makes the research more interesting and easy; as more data are available, more predictions and analyses can be developed. as the bioinformatics community pointed out, only by looking at vast databases that describe the whole of the system will we be able to understand the big picture, and find correlations and patterns (hardisty et al. 2013). however, more efforts should be made by data providers to assure the quality of the data that they are sharing, as most of these data require thorough cleaning (otegui et al. 2013). this study revealed insightful information that will potentially impact scientific knowledge and conservation efforts. even if exhaustive inventories of african palms are somehow feasible objectives for short-term fieldwork, our results demonstrate that, with the addition of digital accessible knowledge on top of existing survey data, a relatively complete picture about the group of interest could be obtained. this observation has important implications for sampling, as combination of available data source reduces the time, effort, and money required for new field surveys, which are nevertheless necessary to gather new data. further, re-visitation of the already studied areas would provide information to understand and appreciate the level of changes in the landscape where these palms are found. acknowledgments data collection for this research was supported by the university of abomey-calavi (republic of benin) through wild-palm project. analyses were fully developed and implemented through active collaboration with town peterson and lindsay campbell. we thank salako valère, akpona jean didier, and donou marcel for their contributions, and two anonymous referees whose comments improved this paper. references akoègninou, a., w. j. van der burg and l. j. g. van der maesen. 2006. flore analytique du bénin. blackhuys publishers, cotonou aranda, s. c., r. gabriel, p. a. borges and j. m. lobo. 2010. assessing the completeness of bryophytes inventories: an oceanic island as a case study (terceira, azorean archipelago). biodiversity & conservation 19: 2469-2484. ariño, a. h. 2010. approaches to estimating the universe of natural history collections data. biodiversity informatics 7: 81-92. ariño, a. h., c. belascoáin and r. jordana. 2008. optimal sampling for complexity in soil ecosystems. in: minai a., bar-yam y., unifying themes in complex systems iv. springer. pp 220230. bruner, a. g., r. e. gullison, r. e. rice, and g. a. da fonseca. 2001. effectiveness of parks in protecting tropical biodiversity. science 291: 125-128. cardoso, p., s. s. henriques, c. gaspar, l. c. crespo, r. carvalho, j. b. schmidt and t. szűts. 2009. species richness and composition assessment of spiders in a mediterranean scrubland. journal of insect conservation 13: 45-55. townpeterson typewritten text 53 biodiversity informatics, 10, 2015, 45-55     chao, a. and l. jost. 2012. coverage-based rarefaction and extrapolation: standardizing samples by completeness rather than size. ecology 93: 25332547. chao, a., w-h. hwang, y-c. chen and c-y kuo. 2000. estimating the number of shared species in two communities. statistica sinica 10: 227-246. colwell, r. k. and j. a. coddington. 1994. estimating terrestrial biodiversity through extrapolation. philosophical transactions of the royal society b, 345: 101-118. dethoisy, b., s. brosse and m. a. dubois. 2008. assessment of large-vertebrate species richness and relative abundance in neotropical forest using linetransect censuses: what is the minimal effort required? biodiversity & conservation 17: 26272644. escala, m.c., j.c. irurzun, a. rueda and a.h. ariño. 1997. atlas of the insectivora and rodentia of navarra: biogeographical analysis. publicaciones de biología de la universidad de navarra, serie zoológica 25: 1-79. fao. 2010. evaluation des ressources forestières mondiales. rapport principal. organisation des nations unies pour l'alimentation et l'agriculture. rome, italie. graham, c. h., s. ferrier, f. huettman, c. moritz and a. t. peterson. 2004. new developments in museum-based informatics and applications in biodiversity analysis. trends in ecology and evolution 19: 497-503. guralnick, r. p., j. wieczorek, r. beaman, r. j. hijmans and biogeomancer working group. 2006. biogeomancer: automated georeferencing to map the world's biodiversity data. plos biology 4: e381. hardisty, a., d. roberts and the biodiversity informatics community. 2013. a decadal view of biodiversity informatics: challenges and priorities. bmc ecology 13:16. hijmans, r.j., s.e. cameron, j.l. parra, p.g. jones and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25: 1965-1978. humphries, c. j., p. h. willams and r. i. vane-wright. 1995. measuring biodiversity value for conservation. annual review of ecology and systematics 26: 93-111. kier, g., j. mutke, e. dinerstein, t.h. ricketts, w. küper, h. kreft and w. barthlott 2005. global patterns of plant diversity and floristic knowledge. journal of biogeography 32: 1107–1116. kjær, a., a. s. barfod, c. b. asmussen and o. seberg. 2004. investigation of genetic and morphological variation in the sago palm (metroxylonsagu; arecaceae) in papua new guinea. ann bot-london 94: 109-117. lovell, s. j., m. l. hamer, r. h. slotow and d. herber. 2010. assessment of sampling approaches for a multi-taxa invertebrate survey in a south african savanna-mosaic ecosystem. austral ecology 35: 357-370. magurran, a.e. 1988. ecological diversity and its measurement. london: north-holland. may, r. m. 1988. how many species are there on earth? science 241: 1441-1449. monteiro, j. m., u. p. de albuquerque, e. m. de freitas lins-neto, e. l. de araújo and e. l. c. de amorim. 2006. use patterns and knowledge of medicinal species among two rural communities in brazil's semi-arid northeastern region. journal of ethnopharmacology 105: 173-186. mora, c. 2008. a clear human footprint in the coral reefs of the caribbean. proceedings of the royal society of london b 275: 767-773. nelson, b. 2009. data sharing: empty archives. nature: 461.  160-163 otegui, j., a. h. ariño, v. chavan and s. gaiji. 2013. on the dates of gbif-mobilised primary biodiversity records. biodiversity informatics 8: 173-184. pino-del-carpio, a., a. h. ariño and r. miranda. 2014a. data exchange gaps in knowledge of biodiversity: implications for the management and conservation of biosphere reserves. biodiversity & conservation 23: 2239-2258. ponder, w. f., g. a. carter, p. flemons and r. r. chapman. 2001. evaluation of museum collection data for use in biodiversity assessment. conservation biology 15: 648-657. pyke, g. h. and p. r. ehrlich. 2010. biological collections and ecological/environmental research: a review, some observations and a look to the future. biological review 85: 247-266. r development core team. 2013. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. raven, p.h. and e.o. wilson. 1992. a fifty-year plan for biodiversity surveys. science 258: 1099-1100. scoble, m. j. 2000. costs and benefits of web access to museum data. trends in ecology and evolution 15: 374. seber, g. a. f. 1982. the estimation of animal abundance and related parameters. macmillan, new york. soberón, j. m., j. b. llorente and l. oñate. 2000. the use of specimen-label databases for conservation purposes: an example using mexican papilionid and pierid butterflies. biodiversity & conservation 9: 1441-1466. townpeterson typewritten text 54 townpeterson typewritten text biodiversity informatics, 10, 2015, 45-55     soberón, j., and t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. philosophical transactions of the royal society b 359: 689-698. soria-auza, r. w. and m. kessler, 2008. the influence of sampling intensity on the perception of the spatial distribution of tropical diversity and endemism: a case study of ferns from bolivia. diversity & distributions 14: 123-130. sousa-baena, m. s., l. c. garcia and a. t. peterson. 2014. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. diversity & distributions 20: 369-381. stehman, s.v. 2009. sampling designs for accuracy assessment of land cover. international journal of remote sensing 30: 5243–5272. thompson, g. g., s. a. thompson, p. c. withers and j. fraser. 2007. determining adequate trapping effort and species richness using species accumulation curves for environmental impact assessments. austral ecology 32: 570-580. tomlinson, p. b. 1990. the structural biology of palms. oxford university press. oxford. vormisto, j., j. c. svenning, p. hall and h. balslev. 2004. diversity and dominance in palm (arecaceae) communities in terra firme forests in the western amazon basin. journal of ecology 92: 577-588. wang, y., l. m. aroyo, n. stash and l. rutledge. 2007. interactive user modeling for personalized access to museum collections: the rijksmuseum case study. in: user modeling. 2007 springer berlin heidelberg. pp. 385-389. townpeterson typewritten text 55 biodiversity informatics, 17, 2022, pp. 1-9 1 enm2020: a free online course and set of resources on modeling species niches and distributions a. townsend peterson1, matthew e. aiello-lammens2, giuseppe amatulli3, robert p. anderson4,5, marlon e. cobos1, josé alexandre f. diniz-filho6, luis e. escobar7, xiao feng8, janet franklin9, luiz m. r. gadelha jr.10,11, d. georges12, m. guéguen12, tomer gueta13,14, kate ingenloff1,15, scott jarvie16,17,*, laura jiménez1,18, dirk n. karger19, jamie m. kass20, michael r. kearney21, rafael loyola6,22, fernando machado-stredel1, enrique martínez-meyer23, cory merow24, maria luiza mondelli11, sara ribeiro mortara25,26, robert muscarella27, corinne e. myers28, babak naimi29, daniel noesgaard30, ian ondo31, luis osorio-olvera1,32, hannah l. owens33,34, richard pearson35, gonzalo e. pinilla-buitrago4, andrea sánchez-tapia26, erin e. saupe36, wilfried thuiller12, sara varela37,38, dan l. warren20, john wieczorek39, katherine yates40, gengping zhu41, gabriela zuquim16,42, damaris zurell43 1biodiversity institute, university of kansas, lawrence, ks 66045, usa 2department of environmental studies and science, pace university, pleasantville, ny 10570, usa 3yale university, school of the environment, 195 prospect street, new haven, ct, 06511, usa 4department of biology, city college of new york, city university of new york, new york, ny 10031, usa; ph.d. program in biology, graduate center, city university of new york, new york, ny 10016, usa 5division of vertebrate zoology (mammalogy), american museum of natural history, new york, ny 10024, usa 6departamento de ecologia, icb, universidade federal de goiás, goiânia, go, brazil 7department of fish and wildlife conservation, virginia tech, blacksburg, va, 24061, usa 8department of geography, florida state university, tallahassee, fl 32306, usa 9department of botany and plant sciences, university of california, riverside, ca 92521, usa 10friedrich-schiller-university jena, germany 11national laboratory for scientific computing, petrópolis, brazil 12univ. grenoble alpes, univ. savoie mont blanc, cnrs, leca, f-38000 grenoble, france 13department of civil and environmental engineering, the technion – israel institute of technology, haifa 3200003, israel 14steinhardt museum of natural history, tel aviv university, tel aviv 6997801, israel 15center for biodiversity and global change, yale university, new haven, ct, usa 16section for ecoinformatics & biodiversity, department of biology, aarhus university, ny munkegade 114, 8000 aarhus, denmark 17centre for biodiversity dynamics in a changing world (biochange), department of biology, aarhus university, ny munkegade 114, 8000 aarhus, denmark *otago regional council, dunedin, aotearoa, new zealand 18school of life sciences, university of hawai’i at manoa, honolulu, hi 96822, usa 19swiss federal institute for forest, snow and landscape research wsl, zürcherstrasse 111, 8903 birmensdorf, switzerland 20biodiversity and biocomplexity unit, okinawa institute of science and technology graduate university, tancha, onna-son, kunigami-gun, okinawa, japan 21school of biosciences, university of melbourne, victoria 3010, australia 22fundação brasileira para o desenvolvimento sustentável, rio de janeiro, rj, brazil a. townsend peterson et al. – free online course and set of resources on modeling species niches 2 23departamento de zoología, instituto de biología, universidad nacional autónoma de méxico, mexico city 04510, mexico 24eversource energy center and department of ecology and evolutionary biology, university of connecticut, storrs, ct, usa 25international institute for sustainability iis-rio, rio de janeiro, brazil 26instituto de pesquisas jardim botânico do rio de janeiro, núcleo de computação científica e geoprocessamento, rio de janeiro, brazil 27plant ecology & evolution, evolutionary biology centre, uppsala university, uppsala, sweden 28department of earth and planetary sciences, university of new mexico, albuquerque, nm 87131, usa 29department of geosciences and geography, university of helsinki, 00014 po box 64, helsinki, finland 30global biodiversity information facility (gbif) secretariat, universitetsparken 15, 2100 copenhagen ø, denmark 31department of biodiversity informatics and spatial analysis (bisa), royal botanic gardens, kew, richmond, surrey tw9 3ab, uk 32departamento de ecología de la biodiversidad, instituto de ecología, universidad nacional autónoma de méxico, ciudad de méxico, 04510, mexico 33center for macroecology, evolution and climate, globe institute, university of copenhagen, copenhagen, denmark 34florida museum of natural history, university of florida, gainesville, fl, 32611, usa 35centre for biodiversity & environment research, department of genetics, evolution & environment, university college london, uk 36department of earth sciences, university of oxford, s parks road, ox1 3an, uk 37departamento de ecoloxía e bioloxía animal. universidade de vigo, spain 38museum für naturkunde, leibniz institute for evolution and biodiversity science, berlin, germany 39vertnet, rauthiflor llc 40school of science, engineering and environment, university of salford, manchester, uk 41department of entomology, washington state university, pullman, wa 99163 usa 42department of biology, university of turku, finland 43institute of biochemistry and biology, university of potsdam, potsdam, germany abstract. the field of distributional ecology has seen considerable recent attention, particularly surrounding the theory, protocols, and tools for ecological niche modeling (enm) or species distribution modeling (sdm). such analyses have grown steadily over the past two decades—including a maturation of relevant theory and key concepts—but methodological consensus has yet to be reached. in response, and following an online course taught in spanish in 2018, we designed a comprehensive english-language course covering much of the underlying theory and methods currently applied in this broad field. here, we summarize that course, enm2020, and provide links by which resources produced for it can be accessed into the future. enm2020 lasted 43 weeks, with presentations from 52 instructors, who engaged with >2500 participants globally through >14,000 hours of viewing and >90,000 views of instructional video and question-and-answer sessions. each major topic was introduced by an “overview” talk, followed by more detailed lectures on subtopics. the hierarchical and modular format of the course permits updates, corrections, or alternative viewpoints, and generally facilitates revision and reuse, including the use of only the overview lectures for introductory courses. all course materials are free and openly accessible (cc-by license) to ensure these resources remain available to all interested in distributional ecology. key words: ecological niche model, species distribution model, course, open access, methods a. townsend peterson et al. – free online course and set of resources on modeling species niches 3 distributional ecology is a branch of biogeography that focuses on the fundamental question of why species are found where they are and identifying where they could occur given changing environmental, geographic, or biotic conditions. the modern renaissance of the field began in the 1970s (macarthur 1972, austin 1987, ferrier 2002), and has since led to many exciting insights (e.g., novel reflections on the frequency of ecological speciation) and useful products (e.g., detailed range maps and potential distribution maps) (guisan and zimmermann 2000, araújo and pearson 2005, peterson and navarro-sigüenza 2017). work in the field of distributional ecology includes theory, protocols, and tools drawn from many areas of inquiry, including ecology, biogeography, evolutionary biology, geographic information science, meteorology, hydrology, remote sensing, statistics, and computer science. despite at least three book-length syntheses (franklin 2010, peterson et al. 2011, guisan et al. 2017) and numerous synthetic papers on the subject (e.g., guisan and thuiller 2005, elith and leathwick 2009, anderson 2013, araújo et al. 2019, feng et al. 2019, zurell et al. 2020), the field still lacks a clearly established set of methodologies by which to guide future advances. this novelty and speed of development of the field, combined with intense interest, have led to a series of in-person and online training programs and courses, ranging from broad surveys of all biodiversity informatics (e.g., peterson and ingenloff 2015) to courses specifically on ecological niche modeling (e.g., peterson et al. 2019; note, this particular course was in spanish). however, given the somewhat dated nature of many existing courses and the substantial fees often involved, a significant gap was noted: a free, online course in english spanning the entire suite of theory, protocols, and tools in the ecological niche modeling toolkit. a team of 52 instructors that represents a great breadth of expertise in this area worked to generate the instructional format and content of a new course. instructors came from many countries and represented a variety of career stages. the course the course was delivered during january– november 2020. it was divided into 18 overarching themes or topics: introduction, applications, key tools, environmental data, occurrence data, visualization, distributional equilibrium, algorithms, uncertainty, evaluation, model selection, model transfers, model comparisons, reproducibility, abundances, frontiers, practicalities, and conclusions (table 1). each major topic was initiated by a talk designated as an “overview.” persons desiring deep knowledge could view all the talks in each set following the corresponding overview lecture, whereas individuals aiming for a general summary of the field had the option to only watch overview talks. completion certificates for the full course required participation in question-and-answer sessions and adequate performance on a short examination at the end of the course. the modular format of the course and its materials also facilitates later addition of updates, corrections, or alternative viewpoints, leading to a set of resources that should be adaptable and easily updated into the future. the course followed a set weekly schedule. lectures were pre-recorded to avoid technological and internet-related complications of live presentations, for both instructors and students. presentations were made available on monday morning (in the western hemisphere, utc-6) in various formats: youtube videos, .mp4 video files, .mp3 audio files, and .pdf slide decks. youtube videos had the advantage of automatically including the option of closed-captioning, which (though not perfect) can assist both deaf and hard-of-hearing individuals and persons for whom english is not a native language. presentations were accompanied by ancillary materials such as readings from the primary literature, example datasets, and programming code. participants’ questions were due by wednesday each week, and a live question-and-answer session among instructors was held each friday, with an archived version made available online directly upon conclusion. the course reached a large audience. in total, 2541 formal participants joined the course facebook group1, but many more took advantage of the materials. the youtube videos were viewed 90,938 times by participants from at least 72 countries worldwide (figure 1), representing 14,172 hours of viewership (as of 15 november 2020). in total, 3159 questions were submitted by course participants (see figure 2 for a word cloud summarizing the terms most frequently used in these questions; r code to produce this figure is provided via ku scholarworks2. our desire in developing this course was to facilitate and motivate current and future scholars in the field worldwide to explore and innovate in distributional ecology. we also hope that the open format— 1 https://www.facebook.com/groups/enm2020/. 2 http://hdl.handle.net/1808/32540. https://www.facebook.com/groups/enm2020/ http://hdl.handle.net/1808/32540 a. townsend peterson et al. – free online course and set of resources on modeling species niches 4 made available globally via the internet without cost—will serve to increase the diversity of participation in the field and improve educational equity by reducing barriers for all interested in these theories, protocols, and tools. lessons learned this course differs from the usual model for broad, extra-institutional courses in recent years. for enm2020, we assembled a large proportion of the leading experts in the field and created an open, free-to-all learning platform that presents much of the current knowledge and practice for a complex area of inquiry. many other courses in this and related areas are presented by one or a few researchers, and are often accompanied by substantial fees that constitute significant barriers to participation for many potentially-interested individuals. the key features of enm2020 were (1) broad participation by many leaders in the area of distributional ecology, as reflected in the long author list for this contribution, and (2) open access to the content. the hefty community participation in the instructor list gave the course an air of plurality towards different, and at times even opposing, ideas regarding particular topics. these differences and debates, while conducted civilly, can be perceived in the ideas presented in various talks in enm2020, and particularly in the question-and-answer sessions, where differences were at times debated more directly. open access to the content, which was facilitated by posting course materials on youtube and making materials available to participants on multiple platforms (e.g., videos for download or streaming, as well as .mp3 audio files and .pdf slide decks for those with poorer internet access), is also a key divergence from previous courses. in the process of developing and presenting this course, however, we noted ways in which the process could have been improved. specifically, even a modicum of direct funding might have permitted more sophisticated editing and preparation of the course videos before they were posted. additionally, figure 1. summary of enm2020 course video views by country, presented as orders of magnitude of viewership. note that some countries do not allow access to youtube, such that viewership from those countries shows as “no data”. figure 2. visualization showing representation of different words among 3159 questions submitted to the instructors by course participants during enm2020 (developed in r with package wordcloud; code included in supplemental information). a. townsend peterson et al. – free online course and set of resources on modeling species niches 5 developing specific exercises could have supported the learning experience more directly, particularly if they had been presented on a single learning platform (e.g., moodle) for more direct and easy access. notably, enm2020 was developed and presented with no funding other than support from the authors’ institutions and a few software-development grants in the form of the salaries and facilities for each of the instructors. acknowledgments we thank the course participants for their enthusiastic and inquisitive approach to such a large-scale and (at times) seemingly unending course. although all instructors and their institutions and employers donated the time they spent preparing and delivering the materials, many of them were supported by particular grants and fellowships that achieved widespread outreach through their participation: grants from the u.s. national science foundation (dbi1661510 for mal, rpa, jmk, cm, and gepb; iia1920946 for mec and atp), national aeronautics and space administration (80nssc18k0406 for mal, rpa, jmk, and gepb), and a japan society for the promotion of science postdoctoral fellowship for foreign researchers (jmk). conflicts of interest the authors declare that no conflicts of interest exist. literature cited anderson, r. p. 2013. a framework for using niche models to estimate impacts of climate change on species distributions. ann. ny acad. sci. 1297:8–28. araújo, m. b., r. p. anderson, a. m. barbosa, c. m. beale, c. f. dormann, r. early, r. a. garcia, a. guisan, l. maiorano, b. naimi, r. b. o’hara, n. e. zimmermann, and c. rahbek. 2019. standards for distribution models in biodiversity assessments. science adv. 5:eaat4858. araújo, m. b., and r. g. pearson. 2005. equilibrium of species’ distributions with climate. ecography 28:693-695. austin, m. 1987. models for the analysis of species’ response to environmental gradients. vegetatio 69:3545. elith, j., and j. r. leathwick. 2009. species distribution models: ecological explanation and prediction across space and time. ann. rev. ecol. evol. syst. 40:677697. feng, x., d. park, c. walker, a. t. peterson, c. merow, and m. papes. 2019. a checklist for maximizing reproducibility of ecological niche models. nature ecol. evol. 3:1382-1395. ferrier, s. 2002. mapping spatial pattern in biodiversity for regional conservation planning: where to from here? syst. biol. 51:331-363. franklin, j. 2010. mapping species distributions: spatial inference and prediction. cambridge university press, cambridge. guisan, a., and w. thuiller. 2005. predicting species distribution: offering more than simple habitat models. ecol. lett. 8:993-1009. guisan, a., w. thuiller, and n. e. zimmermann. 2017. habitat suitability and distribution models: with applications in r. cambridge university press, cambridge. guisan, a., and n. zimmermann. 2000. predictive habitat distribution models in ecology. ecol. mod. 135:147186. macarthur, r. 1972. geographical ecology. princeton university press, princeton. peterson, a. t., r. p. anderson, m. e. cobos, m. cuahutle, a. p. cuervo-robayo, l. e. escobar, m. fernández, d. jiménez-garcía, a. lira-noriega, j. m. lobo, f. machado-stredel, e. martínez-meyer, c. nuñezpenichet, j. nori, l. osorio-olvera, m. t. rodríguez, o. rojas-soto, d. romero-álvarez, j. soberón, s. varela, and c. yañez-arenas. 2019. curso modelado de nicho ecológico, versión 1.0. biodiv. inf. 14:1-7. peterson, a. t., and k. ingenloff. 2015. biodiversity informatics training curriculum, version 1.2. biodiv. inf. 10:65-74. peterson, a. t., and a. g. navarro-sigüenza. 2017. what bird specimens can reveal about species-level distributional ecology. stud. av. biol. 50:111-125. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. zurell, d., j. franklin, c. könig, p. j. bouchet, j. m. serradiaz, c. f. dormann, j. elith, g. fandos guzman, x. feng, g. guillera-arroita, a. guisan, p. j. leitão, j. j. lahoz-monfort, d. s. park, a. t. peterson, g. rapacciuolo, d. r. schmatz, b. schröder, w. thuiller, k. l. yates, n. e. zimmermann, and c. merow. 2020. a standard protocol for reporting species distribution models. ecography 43:1261-1277. a. townsend peterson et al. – free online course and set of resources on modeling species niches 6 table 1. summary of enm2020 course curriculum. note that the order of the talks has been rearranged somewhat to reflect better the flow of topics. talks marked with a cross (†) are “overview” talks, which can be viewed in sequence (without the talks that are not so marked) to provide a much-briefer introductory “short course.” this table can also be found in the website of biodiversity informatics training curriculum1. 1 http://biodiversity-informatics-training.org/. title youtube link instructor(s)* pdf link additional materials introduction course introduction video atp --introduction to ecological niche theory video js pdf -introduction to distributional ecology† video atp pdf -question and answer video loo, js, atp, hlo -materials applications applications† video rp pdf -niche structure and limits video js pdf -discovery of species and populations video atp pdf materials question and answer video js, mc, atp -materials climate change video emm pdf materials special discussion: climate change video js, emm, mp, mc, atp -materials reconstructing past distributions video cm1 pdf -invasive species applications video gz1 pdf materials question and answer video emm, js, ast, sm, atp pdf materials systematic conservation planning video rl1 pdf materials large-scale conservation/recovery projects video sj pdf materials public health video lee pdf materials question and answer video mp, atp -materials key tools what is the basic toolkit? video atp pdf materials tools for biodiversity data cleaning video srm, ast pdf exercise, readings question and answer video mp, tg, atp --niche toolbox video loo pdf readings, tutorial. sdmtoolbox video jb pdf materials question and answer video mp, loo, atp -materials environmental data environmental data / relation to theory† video sv pdf materials climate data video dk pdf -question and answer video mp, sv, atp -materials remote-sensing data video mp pdf -soils databases video gz2 -materials question and answer video mp, atp -materials topographic data video ga pdf materials marine environments video hlo pdf -question and answer video ga, hlo, mp, atp -nakazawa, list. paleoclimate data video ees pdf -question and answer video mc, mp, atp, ees -ribeiro, virtual-world, stigall, saupe, myers http://biodiversity-informatics-training.org/ https://youtu.be/vj8qto56rpa https://youtu.be/ilbjjtwqnde https://www.dropbox.com/s/k9nuhh4i66kx8x2/enm2020_w1t2_ecologicalnichetheory.pdf?dl=0 https://youtu.be/gbclzkqiimi https://www.dropbox.com/s/0jws2w7hyekv54n/enm2020_w1t3_distributionalecology.pdf?dl=0 https://youtu.be/mhapzicokum https://www.dropbox.com/s/07tv1a1uhh72b3c/w1_qa_materials.zip?dl=0 https://youtu.be/n5_ucywlmng https://www.dropbox.com/s/kdvd7xidadz5zg1/enm2020_w2t1_applications_overview_final.pdf?dl=0 https://youtu.be/qermoeowsgg https://www.dropbox.com/s/d2kov9ocpkcfdpf/enm2020_w2t2_nichestructurelimits.pdf?dl=0 https://youtu.be/sgwxhwficec https://www.dropbox.com/s/1q672oibhu7wi3n/enm2020_w2t3_discoveryspeciespopulations.pdf?dl=0 https://www.dropbox.com/s/rnt8d2ljlirx6ou/enm2020_w2t3_discoveryspeciespopulations.zip?dl=0 https://youtu.be/rd7ghfaflbu https://www.dropbox.com/s/2yjv0e9vmaa2ag4/enm2020 w2 questionandanswer_materials.zip?dl=0 https://youtu.be/cwdxpzbq7e0 https://www.dropbox.com/s/bt88c6tkfw7q40w/enm2020_w3t1_climatechange.pdf?dl=0 https://www.dropbox.com/s/ds5e2o69j18kaiw/enm2020_w3t1_climatechange.zip?dl=0 https://youtu.be/ih7gvpromvu https://www.dropbox.com/s/xk1w59y4ura7fx8/petal_nyas_2018.pdf?dl=0 https://youtu.be/771nfhoc5rk https://www.dropbox.com/s/xdua8uanpxb8qdf/enm2020_w3t2_pastdistributions.pdf?dl=0 https://youtu.be/rznt44phmci https://www.dropbox.com/s/eezebpa69rdbfax/enm2020_w3t3_invasivespecies.pdf?dl=0 https://www.dropbox.com/s/em6m2m6jorwmcv4/enm2020 w3 invasivespecies_literature.zip?dl=0 https://youtu.be/pl6t5eppe0a https://www.dropbox.com/s/16m4d2ecj15eoow/enm2020 w3 qa_answers.pdf?dl=0 https://www.dropbox.com/s/24wwkbvl1mzh469/enm2020%20w3 qa_materials.zip?dl=0 https://youtu.be/zks9qs1i69g https://www.dropbox.com/s/65l8km4om8hunxs/enm2020_w4t1_systematicconservationplanning.pdf?dl=0 https://www.dropbox.com/s/nevlqccgqyhq17e/enm2020_w4t1_systematicconservationplanning.zip?dl=0 https://youtu.be/fb8mfq_cfui https://www.dropbox.com/s/2fkpqr7nm9y31nn/enm2020_w4t2_restoration.pdf?dl=0 https://www.dropbox.com/s/k1ye4v66zaontmw/enm2020_w4t2_restoration.zip?dl=0 https://youtu.be/dnbksl7jfwq https://www.dropbox.com/s/5ssg9fpb3tuwrlw/enm2020_w4t3_publichealth.pdf?dl=0 https://www.dropbox.com/s/ls2xu86paoi2hdq/enm2020_w4t3_publichealth.zip?dl=0 https://youtu.be/yibtgs4n6jg https://www.dropbox.com/s/bd5g68f1etbjpox/enm2020 w4 questionandanswer.zip?dl=0 https://youtu.be/6s4ulajxpa4 https://www.dropbox.com/s/rd4tgt2cim1iqy9/enm2020_w5t1_toolkit.pdf?dl=0 https://www.dropbox.com/s/avpknof1jih1cjs/enm2020_w5t1_toolkit.zip?dl=0 https://youtu.be/266q56m-48w https://www.dropbox.com/s/t9qu2jr9h9u7dok/enm2020_w5t2_datacleaningtools.pdf?dl=0 https://github.com/saramortara/data_cleaning https://www.gbif.org/document/80528/principles-and-methods-of-data-cleaning-primary-species-and-species-occurrence-data https://youtu.be/fc5d3dsgnhi https://youtu.be/42rsg60rk-k https://www.dropbox.com/s/kgte46jj74pmm7k/ntbox_presentation.html?dl=0 https://www.dropbox.com/s/nh0kfz5571x6gea/enm2020_w6t1_ntbox_readings.zip?dl=0 https://www.dropbox.com/s/gph6bje0cqdk4ax/tutorial videos for niche toolbox.pdf?dl=0 https://youtu.be/0eyrycawi_u https://www.dropbox.com/s/g3nnzyff0fivc0r/detailed_guide_associated_w_video_sdmtoolbox.pdf?dl=0 https://www.dropbox.com/s/cq9sjzin6kp0swx/enm2020_w6t2_sdmtoolbox.zip?dl=0 https://youtu.be/cf4i9nfrnt4 https://www.dropbox.com/s/cftgsyxmnbc9u2l/enm2020_w6t3_questionsanswers_reading.pdf?dl=0 https://youtu.be/ixt0gcum0zk https://www.dropbox.com/s/daurdjpdifbjmgw/enm2020_w7t1_envdataoverview_presentation.pdf?dl=0 https://www.dropbox.com/s/r3blo5zxxaitrjb/enm2020_w7t1_envdataoverview_papers.zip?dl=0 https://youtu.be/i8n9qhccpyk https://www.dropbox.com/s/7jequ9mka7boljg/enm2020_w7t2_climatedata.pdf?dl=0 https://youtu.be/rlvrbjsogdm https://www.dropbox.com/s/apxamt8zn1pw26q/enm2020_w7t3_questions.zip?dl=0 https://youtu.be/aijbe91aqly https://www.dropbox.com/s/jyn22qilddlibx8/enm2020_w8t1_remotesensingdata.pdf?dl=0 https://youtu.be/icuyr8z3ewg https://www.dropbox.com/s/an95pox22tlmdve/enm2020_w8t2_soilsdata.zip?dl=0 https://youtu.be/zsq4jfskdvu https://www.dropbox.com/s/r1q9lmer53nlwv0/zuquim_answers.pdf?dl=0 https://youtu.be/bqauisbsmsa https://www.dropbox.com/s/0sxocxna7oz0faj/enm2020_w9t1_topographicdata.pdf?dl=0 https://www.dropbox.com/s/vr3xa1yqtoou018/enm2020_w9t1_topographicdata.zip?dl=0 https://youtu.be/uxaowrzxmdo https://www.dropbox.com/s/ad082u4l7iuxh9b/enm2020_w9t2_marinedata.pdf?dl=0 https://youtu.be/96wflge2f1w https://www.dropbox.com/s/8lxoo57k00dv7h4/netal_a_2004.pdf?dl=0 https://www.dropbox.com/s/guloa387mlhjf30/owens bibliography.pdf?dl=0 https://youtu.be/xsvixogf2wy https://www.dropbox.com/s/r5hvlj81w3sg2r7/enm2020_w10t1_paleoclimatedata.pdf?dl=0 https://youtu.be/-7kce3ghg0m https://www.dropbox.com/s/64ryw5soiqem5ap/4955-article text-9503-1-10-20150823.pdf?dl=0 https://www.dropbox.com/s/fi48jylrfigwa5k/setal_nee_2019.pdf?dl=0 https://www.dropbox.com/s/fi48jylrfigwa5k/setal_nee_2019.pdf?dl=0 https://scholar.google.com/citations?hl=en&user=nr0mg_uaaaaj https://scholar.google.com/citations?hl=en&user=zg45llaaaaaj https://scholar.google.com/citations?hl=en&user=hamv7wyaaaaj a. townsend peterson et al. – free online course and set of resources on modeling species niches 7 occurrence data occurrence data† video atp pdf materials relation to theory video js pdf -sources video jw pdf -question and answer video mp, jw, mc, atp --georeferencing video mp pdf -occurrence data cleaning i (simple consistency checks) video atp pdf materials occurrence data cleaning ii (automating the process) video tg pdf -question and answer video tg, mp, mc, atp -materials filtering and autocorrelation video mal pdf r code subsetting for evaluation video jmk pdf r code data citation video dn pdf -question and answer video mal, jmk, dn, js, mc, mp, atp --visualization nichea video lee, hq pdf -demo with nichea further visualization 3d video lee pdf -question and answer video lee, mc, mp, atp -geoda, reprints distributional equilibrium distributional equilibrium / relation to theory† video js pdf -estimating m video fms pdf -question and answer video mc, fms, mp, js, atp -grinnell, literature. bam and m and model success what you can and cannot model video atp pdf bam, literature. bam scenario exercises video atp pdf -question and answer video mc, fms, mp, atp -materials special discussion: extent and resolution video mc, js, atp -ecoclimate algorithms algorithms / relation to theory† video jf pdf hastie, elith, elith06 un solo díos video atp pdf materials question and answer video mc, atp -paper on uses, zhu. maxent 1 video cm2 pdf materials maxent 2 video cm2 pdf -question and answer video cm2, mp, mc, atp --modler platform demo video ast, srm pdf -wallace video jmk, gepb pdf -question and answer video jmk, gepb, sm, ast, mc, mp, atp --sdm video bn pdf r code biomod – introduction video wt pdf -biomod single species video dg pdf r code biomod multiple species video dg pdf r code biomod – specifics video mg pdf r code biomod – shiny video io --question and answer video bn, mc, mp, atp --https://www.youtube.com/watch?v=dscgrynweb4&feature=youtu.be https://www.dropbox.com/s/ltdrz3yz60nsm04/enm2020_w11t1_occurrence dataoverview.pdf?dl=0 https://www.dropbox.com/s/dnkbptcyn54a0s1/petal_bdj_2018.pdf?dl=0 https://youtu.be/8x80f-e37ue https://www.dropbox.com/s/2vngyva20p18unj/enm2020_w11t2_occurrencedataconcepts.pdf?dl=0 https://youtu.be/rr9gqaxb06w https://www.dropbox.com/s/3mfszyjjfrcnbl7/enm2020_w11t3_occurrencedatasources.pdf?dl=0 https://youtu.be/lzihzof6xqe https://youtu.be/oohaekoy-sa https://www.dropbox.com/s/v6slfmkvdo7pzg4/enm2020_w12t1_georeferencing.pdf?dl=0 https://youtu.be/gy6uvmpfcfa https://www.dropbox.com/s/9j44nl51iwxoldr/enm2020_w12t2_datacleaning.pdf?dl=0 https://www.dropbox.com/s/zbcwxb9gu1g66op/enm2020_w12t2_datacleaning.zip?dl=0 https://youtu.be/eojdyav0qqq https://www.dropbox.com/s/vmf21g6l0vpwxbs/enm2020_w12t3_datacleaning2.pdf?dl=0 https://youtu.be/xek5inojnxk https://www.dropbox.com/s/fc4ipyiel191srx/enm2020_w12t4_questions.zip?dl=0 https://youtu.be/8vl46bwen8m https://www.dropbox.com/s/2qrjfnchcva8xps/enm2020_wk13t1_filtering_autocorrelation.pdf?dl=0 https://www.dropbox.com/s/pxmbq0eh4h56ree/enm2020_wk13t1_filtering_autocorrelation-r-example.r?dl=0 https://youtu.be/nyxygsrnzlw https://www.dropbox.com/s/gfnkcu9oqpks8sj/enm2020_w13t2_datasubsetting.pdf?dl=0 https://www.dropbox.com/s/o3rlivtkn8akqpq/data_subsetting.r?dl=0 https://youtu.be/co7b477fo6w https://www.dropbox.com/s/wasyjgy696m8arg/enm2020_w13t3_datacitation.pdf?dl=0 https://youtu.be/3uitrybff3u https://youtu.be/vkqoxowalwo https://www.dropbox.com/s/ezbly9eb657dkqc/enm2020_w14t1_visualization.pdf?dl=0 https://youtu.be/ytpu_igra5m https://www.dropbox.com/s/oq5olibd8lmm9jj/enm2020_w14t2_visualization2.pdf?dl=0 https://youtu.be/t-mtaftw5ce https://geodacenter.github.io/ https://www.dropbox.com/s/jl2gf8i8n7tkrz6/enm2020_w14t3_questions.zip?dl=0 https://youtu.be/89z2aknxzoo https://www.dropbox.com/s/v88zehl2plvhx3e/enm2020_w15t1_distributionalequilibriumoverview.pdf?dl=0 https://youtu.be/uzcj5jhgrqs https://www.dropbox.com/s/vyl0eumg627gmfc/enm2020_w15t2_estimatingm.pdf?dl=0 https://youtu.be/1hzh-upzcoc https://github.com/fmachados/grinnell https://www.dropbox.com/s/fqfvmmflgtk281b/enm2020_w15t3_questions.zip?dl=0 https://youtu.be/62tye4qejis https://www.dropbox.com/s/0j4sxq9ua54b2tc/enm2020_w16t1_bam.pdf?dl=0 https://journals.ku.edu/jbi/article/view/4 https://www.dropbox.com/s/o03h0otfq1cuifl/enm2020_w16t1_bam.zip?dl=0 https://youtu.be/bzxnv8x1xxs https://www.dropbox.com/s/dk4qbsq7y5w3ssi/enm2020_w16t2_bamexamples.pdf?dl=0 https://youtu.be/olwsadyd2ke https://www.dropbox.com/s/0ag2t43kp75433o/betal_geb_2013.pdf?dl=0 https://youtu.be/kc1n8dhnfsk https://journals.ku.edu/jbi/article/view/4955 https://youtu.be/iro9wa6xkke https://www.dropbox.com/s/gzzzgomrhszd9po/enm2020_w17t1_algorithmsoverview.pdf?dl=0 https://web.stanford.edu/~hastie/papers/eslii.pdf https://www.dropbox.com/s/4f7gneg9zf8svng/eetal_e_2006.pdf?dl=0 https://www.dropbox.com/s/isedqmbfazvtxxw/elith_franklin_2013.pdf?dl=0 https://youtu.be/byx9kzvzrbk https://www.dropbox.com/s/oiqgoszzduwdz51/enm2020_w17t2_unsolodios.pdf?dl=0 https://www.dropbox.com/s/rv17whjzj0unc6n/enm2020_w17t2_unsolodios.zip?dl=0 https://youtu.be/va-mudjuyv4 https://journals.ku.edu/jbi/article/view/29 https://www.dropbox.com/s/s8g7z1pcavpe0mr/zp_bi_2017.pdf?dl=0 https://youtu.be/q-vhisr9qfi https://www.dropbox.com/s/lu5j2ga86em5qrd/enm2020_w18t1_maxent.pdf?dl=0 https://www.dropbox.com/s/l3d0qgptmnlfx6o/enm2020_w1t3_distributionalecology.mp3?dl=0 https://youtu.be/daiqtksskz4 https://www.dropbox.com/s/lu5j2ga86em5qrd/enm2020_w18t1_maxent.pdf?dl=0 https://youtu.be/aw4dl4suqnk https://youtu.be/4xw33tdivxa https://www.dropbox.com/s/a31kvunp34ak9o7/enm2020_w19t1_modler.pdf?dl=0 https://youtu.be/kwnynd2x1uo https://www.dropbox.com/s/0otjhx4772bcn3m/enm2020_w19t2_wallace.pdf?dl=0 https://youtu.be/qlm_whhzrbk https://youtu.be/tabxweev56i https://www.dropbox.com/s/zdr6s6654hu0tma/enm2020_w20t1_sdm.pdf?dl=0 https://www.dropbox.com/s/542a7wkumtgs616/enm2020_w20t1_sdm.r?dl=0 https://youtu.be/-iadf8vh6uy https://www.dropbox.com/s/5x8h777wlduk93v/biomod2_1_introduction_wthuiller.pdf?dl=0 https://youtu.be/qrwqhjgrbny https://www.dropbox.com/s/no1pruo7kuyqrad/biomod2_2_single_species_modelling_dgeorges.pdf?dl=0 https://www.dropbox.com/s/dk0muucfal5ncse/biomod2_2_single_species_modelling_dgeorges.r?dl=0 https://youtu.be/chxidjblxe0 https://www.dropbox.com/s/i99im5kuhfiqihs/biomod2_3_multi_species_modelling_dgeorges.pdf?dl=0 https://www.dropbox.com/s/5s3i256lz4hetki/biomod2_3_multi_species_modelling_dgeorges.r?dl=0 https://youtu.be/hejhdurry3o https://www.dropbox.com/s/p1mv5ffcu21zb47/biomod2_4_specificities_mgueguen.pdf?dl=0 https://www.dropbox.com/s/730h5epfjacqm1f/biomod2_4_specificities_mgueguen.r?dl=0 https://youtu.be/nhwpyelhxoa https://youtu.be/xa4lx5xddzc a. townsend peterson et al. – free online course and set of resources on modeling species niches 8 special discussion: algorithm choice video js, mc, mp, atp, emm --bonus: model fit in e video dw --uncertainty uncertainty in enm video atp pdf materials question and answer video mc, mp, atp --evaluation model evaluation / relation to theory† video rpa pdf materials prediction-based evaluations video atp pdf materials question and answer video rpa, jmk, mc, mp, atp -materials model evaluation not prediction-based video atp, sm pdf -question and answer video mc, mp, atp --model selection relation to theory video dw pdf materials enmeval video bm pdf -kuenm video mc pdf materials question and answer video mc, mp, atp -materials model transfers model transfers† video ky --relation to theory video hlo pdf -question and answer video hlo, mp, mc --past, present, future video emm pdf materials extrapolation and measuring extrapolation video hlo pdf materials collinearity video xf pdf -question and answer video -answers, literature model comparisons relation to theory: g or e spaces, niche overlap video mc, js pdf materials niche comparisons in geographic space video dw pdf materials question and answer video mc, mp, atp --niche comparisons in environmental space 1 video mc pdf -niche comparisons in environmental space 2 video mc pdf materials question and answer video mc, mp --repeatability introduction to repeatability in enm† video atp pdf -metadata standards video xf pdf materials quality standards video dz pdf materials question and answer video dz, mlm, mp, mc, atp --abundances relation to theory video js pdf -controversy video lee, acr pdf -methods and test results video emm pdf materials question and answer video mp, emm, js, mc, atp -materials https://youtu.be/zly9etdhigc https://youtu.be/ativruhnhlq https://youtu.be/cr0u7ojcgwm https://www.dropbox.com/s/0wy34vk2hr7es9x/enm2020_w21t1_uncertainty.pdf?dl=0 https://www.dropbox.com/s/81h9sqpfgzzezjt/enm2020_w21t1_uncertainty.zip?dl=0 https://youtu.be/r-t0xel1-q0 https://youtu.be/jg5bcr3jzma https://www.dropbox.com/s/uen5d8yzwd7tdi7/enm2020_w22t1_evaluationoverview.pdf?dl=0 https://www.dropbox.com/s/d90cr7k4dyn27u8/enm2020_w22t1_evaluationoverview.zip?dl=0 https://youtu.be/bifxb6mcnty https://www.dropbox.com/s/2bhjk9jmdfgzf0r/enm2020_w22t2_predictionevaluation.pdf?dl=0 https://www.dropbox.com/s/ds3lgt06qc4n111/enm2020_w22t2_predictionevaluation.zip?dl=0 https://youtu.be/dwjboemwjki https://www.dropbox.com/s/lcbbg85gbu7j1he/enm2020_w22t3_questions.zip?dl=0 https://youtu.be/p12snd_s-lu https://www.dropbox.com/s/9g1n4eoul7r6pc4/enm2020_w23t1_alternativeevaluations.pdf?dl=0 https://youtu.be/fsgso-fe_aw https://youtu.be/ezhqszaaksa https://www.dropbox.com/s/0004oi4sjtypyfx/enm2020_w24t1_modelselectionoverview.pdf?dl=0 https://www.dropbox.com/s/vk8cg00881neach/enm2020_w24t1_modelselectionoverview.zip?dl=0 https://youtu.be/tnyqwmmfmsc https://www.dropbox.com/s/70s1y3m2yxczlk7/enm2020_w24t2_enmeval.pdf?dl=0 https://youtu.be/dn4-g_-mxu8 https://www.dropbox.com/s/jc3tdwosk3fbjtd/enm2020_w24t3_kuenm.pdf?dl=0 https://www.dropbox.com/s/gvndyp1v174ck33/enm2020_w24t3_kuenm.zip?dl=0 https://youtu.be/-cctgawqpr8 https://www.dropbox.com/s/oly2omnyx1p2zbn/enm2020_w24t4_questions_warrenanswers.pdf?dl=0 https://youtu.be/hqtwie_2ppo https://youtu.be/e9n0erb0ba0 https://www.dropbox.com/s/cipb65ij6jvqjib/enm2020_w26t2_modeltransfertheory.pdf?dl=0 https://youtu.be/munewndlqgi https://youtu.be/epxz4g7f4ak https://www.dropbox.com/s/2rzaq39gtsai5ra/enm2020_w27t1_pastpresentfuture.pdf?dl=0 https://www.dropbox.com/s/8ve77301bvg7imt/enm2020_w27t1_pastpresentfuture.zip?dl=0 https://youtu.be/vuvl_fxdh3c https://www.dropbox.com/s/kf53xce0fmqoton/enm2020_w27t2_measuringextrapolation.pdf?dl=0 https://www.dropbox.com/s/wwg73ss10h4j38k/oetal_em_2013.pdf?dl=0 https://youtu.be/ueujovk_b8y https://www.dropbox.com/s/gqmtcjv6el7gezc/maxent collinearity v3.pptx_v1.pdf?dl=0 https://youtu.be/wn_gmu7zooa https://www.dropbox.com/s/bmgkyhfe3cvh80b/enm2020_w27t4_questions.pdf?dl=0 https://www.dropbox.com/s/azph31rv7d1zou9/enm2020_w27t4_questions.zip?dl=0 https://youtu.be/dnexvcexmao https://www.dropbox.com/s/q7vhfnppbl3gugn/enm2020_w28t1_nichecomparisonsintro.pdf?dl=0 https://www.dropbox.com/s/78q58q79xioapze/enm2020_w28t1_nichecomparisonsintro.zip?dl=0 https://youtu.be/ruca2ygrpxq https://www.dropbox.com/s/5vh54hbwg48ibi9/enm2020_w28t2_nichecompsg.pdf?dl=0 https://www.dropbox.com/s/n09bneexcqyolpq/enm2020_w28t2_nichecompsg_readings.pdf?dl=0 https://youtu.be/n_sxpsfctns https://youtu.be/eigvveszt8e https://www.dropbox.com/s/7lphxj7v3ue1eco/enm2020_w29t1_comparingniches.pdf?dl=0 https://youtu.be/id-o5uciff4 https://www.dropbox.com/s/7lphxj7v3ue1eco/enm2020_w29t1_comparingniches.pdf?dl=0 https://www.dropbox.com/s/pzamq0737zmpv7c/enm2020_w29t1_comparingniches.zip?dl=0 https://youtu.be/pwspiz8hp3m https://youtu.be/v1-qfryypgw https://www.dropbox.com/s/lstyjjv3wz9mndx/enm2020_w30t1_reproducibilityoverview.pdf?dl=0 https://youtu.be/opy_u-cjkve https://www.dropbox.com/s/p148p9dcxii36h1/enm2020_w30t2_reproducibility.pdf?dl=0 https://www.dropbox.com/s/qdv1adzjxjp9bfr/fetal_nee_2019.pdf?dl=0 https://youtu.be/mulboxmnyj8 https://www.dropbox.com/s/jcgfpfhx65mrws3/enm2020_w30t3_reproducibility.pdf?dl=0 https://www.dropbox.com/s/9mj5zdz02nqp3tk/zetal_e_2020.pdf?dl=0 https://youtu.be/rfpssqjoiro https://youtu.be/chjclzh8iu8 https://www.dropbox.com/s/gh33bg6kpu1gsbb/talknumber2nichestructure.pdf?dl=0 https://youtu.be/kd6ifux-oqu https://www.dropbox.com/s/jveqy546hz95za8/enm2020_w31t2_nicheabundance.pdf?dl=0 https://youtu.be/va5qfsog8ro https://www.dropbox.com/s/oerzep2sc84sb5v/enm2020_w31t3_nicheabundance_more.pdf?dl=0 https://www.dropbox.com/s/e630ll6xfbd3701/enm2020_w31t3_nicheabundance_more.zip?dl=0 https://youtu.be/0d8cuqrompk https://www.dropbox.com/s/4huel2k85fb5xb5/enm2020_w31t5_questions.zip?dl=0 a. townsend peterson et al. – free online course and set of resources on modeling species niches 9 frontiers frontiers† video jadf pdf materials fitting biologically realistic responses video lj, js pdf materials question and answer video mp, mc, js, lj, atp --mechanistic models. a. mechanistic niche model concepts video mk pdf -mechanistic models. b. heat budgets and microclimates video mk pdf -mechanistic models. c. energy budgets and their integration video mk pdf -mechanistic models. d. integration with correlative models video mk pdf -virtual species and virtual worlds video atp pdf materials question and answer video mk, js, mp, mc, atp -vespa scientific workflows in enm video mlm, lg pdf materials time-specific ecological niche modeling video kri pdf -special discussion: eltonian noise video mc, js, emm, atp -co-occurrence debate question and answer video mlm, lg, mc, mp, js, atp -iucn paper genetic basis of niches video dd, atp pdf materials question and answer video dd, atp, mc, js -materials practicalities publishing results from enms video atp, aa pdf publication course: intro, journal choice, technology, words, figures, color, tables, proofing, literature cited, cover letter and reviewers, authorship, response to reviewers, proofs, presentation for course. open access to scientific literature video atp pdf example writing effective proposals for funding video atp pdf short course conclusions special feature: interview with sara varela video sv, atp --course wrap-up video rp, js, cm1, emm, gz1, hlo, mc, mp, dz, lee, atp, jw --https://youtu.be/kwp8zkjxuow https://www.dropbox.com/s/izin5u9fqz1164u/enm2020_w32t1_frontiersoverview.pdf?dl=0 https://www.dropbox.com/s/y2rhkd5u4xyorv7/enm2020_w32t1_frontiersoverview.zip?dl=0 https://youtu.be/tnxjf9a-2xg https://www.dropbox.com/s/u331bwqytib7xew/jimenezlaura_fitting_biologically_realistic_responses.pdf?dl=0 https://www.dropbox.com/s/ye7ufivod5tpz79/enm2020_w32t2_realisticresponsetypes_reading.pdf?dl=0 https://youtu.be/u10bduoqpf8 https://youtu.be/hwcpzz23p5s https://www.dropbox.com/s/jaorvbl3sth5lh0/enm2020_w33t1a_mechanisticmodels_1.pdf?dl=0 https://youtu.be/nw5rkqxbe5m https://www.dropbox.com/s/4mx2mcrjo6mmlps/enm2020_w33t1b_mechanisticmodels_2.pdf?dl=0 https://youtu.be/dh7b_nhlj1c https://www.dropbox.com/s/z4bjhurp04ff1r9/enm2020_w33t1c_mechanisticmodels_3.pdf?dl=0 https://youtu.be/z0sjkp-1fho https://www.dropbox.com/s/8yn4c9cxrou8li5/enm2020_w33t1d_mechanisticmodels_4.pdf?dl=0 https://youtu.be/frx8k6fno3i https://www.dropbox.com/s/u0ywgvyyn2fk3xy/enm2020_w33t2_virtualworlds.pdf?dl=0 https://www.dropbox.com/s/as6i0h3tl9nc74u/enm2020_w33t2_virtualworlds.zip?dl=0 https://youtu.be/izs8idklpjc https://www.biorxiv.org/content/biorxiv/early/2020/08/12/2020.08.11.246991.full.pdf https://youtu.be/u5l-u03kjbk https://www.dropbox.com/s/xmt7u1q5lssrq5z/enm2020_w34t1_fullmodelreproducibility.pdf?dl=0 https://www.dropbox.com/s/ofgdlbbz8lf5te7/metal_fair_2019.pdf?dl=0 https://youtu.be/ldorc-ihagw https://www.dropbox.com/s/vm1cdq2ep1yiejd/enm2020_w34t2_timespecificenm.pdf?dl=0 https://youtu.be/4u6opknmgxk https://journals.ku.edu/jbi/issue/view/1714 https://journals.ku.edu/jbi/issue/view/1714 https://youtu.be/ic362n3fgt0 https://www.dropbox.com/s/pl13vt8v5i9ik0g/petal_cb_2016.pdf?dl=0 https://youtu.be/urlzwemasp0 https://www.dropbox.com/s/kggyjtnnx4uqtk6/enm2020_w35t1_nichegenomics.pdf?dl=0 https://www.dropbox.com/s/gkdzt6vo29ini4p/enm2020_w35t1_nichegenomics.zip?dl=0 https://youtu.be/2f6qz1klzms https://www.dropbox.com/s/gkdzt6vo29ini4p/enm2020_w35t1_nichegenomics.zip?dl=0 https://youtu.be/pkudqjs6en8 https://www.dropbox.com/s/2lcbc7er26b06ki/enm2020_w40t1_publishing.pdf?dl=0 https://www.google.com/url?q=http://youtu.be/cw6-yovmyfu&sa=d&ust=1602775062766000&usg=afqjcnheduh3qfyidec0rfi2lsellg1ztw https://www.google.com/url?q=http://youtu.be/uvwajpmuopw&sa=d&ust=1602775062767000&usg=afqjcnfbm3mpizuwpwlotaf2me0mlemr_g https://www.google.com/url?q=http://youtu.be/uvwajpmuopw&sa=d&ust=1602775062767000&usg=afqjcnfbm3mpizuwpwlotaf2me0mlemr_g https://www.google.com/url?q=http://youtu.be/ak9rfzlglss&sa=d&ust=1602775062767000&usg=afqjcnfnvr4cr35agiefyl_jwbn9vcmjlq https://www.google.com/url?q=http://youtu.be/ak9rfzlglss&sa=d&ust=1602775062767000&usg=afqjcnfnvr4cr35agiefyl_jwbn9vcmjlq https://www.google.com/url?q=http://youtu.be/o8xnjch6zje&sa=d&ust=1602775062768000&usg=afqjcnh-_xvi3etl3idjbopprge6mxcaiw https://www.google.com/url?q=http://youtu.be/odbyqct1itm&sa=d&ust=1602775062768000&usg=afqjcnfuysb-f2m1pzgiis8rp-umspa4ea https://www.google.com/url?q=http://www.youtube.com/watch?v%3di1bzqgpfs54&sa=d&ust=1602775062774000&usg=afqjcnhxvujfrtinq2kw4l5hqeo8rbw9nw https://www.google.com/url?q=http://youtu.be/o582ejhbzsc&sa=d&ust=1602775062769000&usg=afqjcngzwbiv8nrpz_0eawlx95-tgkdtaq https://www.google.com/url?q=http://youtu.be/joieg9w8n1c&sa=d&ust=1602775062769000&usg=afqjcnehnvywnszw_iw-yllvvmtbao9zwg https://www.google.com/url?q=http://youtu.be/4itn98s90cq&sa=d&ust=1602775062770000&usg=afqjcnfufsos-kl9nwo_rldqwyiorvbsga https://www.google.com/url?q=http://youtu.be/4itn98s90cq&sa=d&ust=1602775062770000&usg=afqjcnfufsos-kl9nwo_rldqwyiorvbsga https://www.google.com/url?q=http://youtu.be/e3g5vigeo8k&sa=d&ust=1602775062771000&usg=afqjcnh2w10ztt33fslpmpae8jxx4z26eq https://www.google.com/url?q=http://youtu.be/e3g5vigeo8k&sa=d&ust=1602775062771000&usg=afqjcnh2w10ztt33fslpmpae8jxx4z26eq https://www.google.com/url?q=http://youtu.be/wzjvzyso-a4&sa=d&ust=1602775062771000&usg=afqjcnfghtddnkkkdwaspgiyjqxntozcqw https://www.google.com/url?q=http://youtu.be/wzjvzyso-a4&sa=d&ust=1602775062771000&usg=afqjcnfghtddnkkkdwaspgiyjqxntozcqw https://www.google.com/url?q=http://youtu.be/iuluzd01ksw&sa=d&ust=1602775062772000&usg=afqjcnhg_vfuamwqjrlorlmj-qo-l5vrkg https://www.google.com/url?q=http://youtu.be/iuluzd01ksw&sa=d&ust=1602775062772000&usg=afqjcnhg_vfuamwqjrlorlmj-qo-l5vrkg https://www.google.com/url?q=http://youtu.be/fcezul43gcc&sa=d&ust=1602775062772000&usg=afqjcnedtosn4a7oaom57k91_mfuqnp8ia https://www.google.com/url?q=http://www.slideshare.net/townpeterson/publishing-scientific-results&sa=d&ust=1602775062773000&usg=afqjcnfky3_tpntkjdhrnttmz2f9h1bxuq https://www.google.com/url?q=http://www.slideshare.net/townpeterson/publishing-scientific-results&sa=d&ust=1602775062773000&usg=afqjcnfky3_tpntkjdhrnttmz2f9h1bxuq https://youtu.be/8wx5bh6nfgw https://www.dropbox.com/s/njg1bs95jowt66d/enm2020_w41t1_openaccess.pdf?dl=0 https://www.dropbox.com/s/8nn34sr039p1v6i/science magazine october 23%2c 2020 page 391.pdf?dl=0 https://youtu.be/i1h2mfarlr8 https://www.dropbox.com/s/4dbcx6su26r1aro/enm2020_w42t1_funding.pdf?dl=0 https://youtu.be/bqsbpmfprzi https://youtu.be/t2y21hphyga https://youtu.be/9oc_sqgpjhw biodiversity informatics, 18, 2024, pp. 13-23 13 interpreting and georeferencing the concept of “near” in biodiversity records peter d. campbell1 biodiversity institute and department of ecology and evolutionary biology, university of kansas, lawrence, ks, 66049 usa correspondence: pete.campbell8@gmail.com abstract. georeferencing historical biodiversity specimens is a difficult but necessary task to bring data accumulated in the course of past scientific efforts into full currency for modern use. textual locality descriptions vary widely, and are prone to error involved in interpretation of brief descriptions and often-unclear terms. each type of locality description requires particular georeferencing methods to maximize precision and accuracy of resulting coordinates and uncertainty. current “best practice” methods concerning textual descriptions referring to proximity to a locality (i.e., “near” a locality) are arbitrary, restrictive, or undefined. in this paper, i explore these challenges, and provide new methods for assigning geographic coordinates and uncertainty (with appropriate metadata) to such locality descriptions using point, line, or polygon shapes as the basis for voronoi diagrams. voronoi diagrams define the geographic space nearer to a given point than to any other point in a collection, making them ideally suited for determining the shape of such locality descriptions. key words: voronoi diagram, best practices, geographic information system, darwin core, coordinate uncertainty. 1 orcid: https://orcid.org/0000-0003-3967-3109 introduction understanding the geographic context from which a specimen originates is vitally important to understanding its ecological, environmental, and evolutionary history. modern specimens are often collected with gps-based geographic coordinates, such that their metadata document precise coordinates and information on uncertainty. however, many historical records were collected without the aid of gps technology, so geographic coordinates need to be assigned to them retrospectively. when a biodiversity specimen is collected without gps-based coordinates, textual locality descriptions allow researchers to relate the specimen back to where it originated. this description can take a variety of forms: a city, an address, a landscape feature like a mountain or river, or any other geographic feature to which the specimen can be related. although most specimens in natural history collections include some sort of associated locality description, many remain as textual descriptions without associated geographic coordinates necessary for many quantitative analyses (murphy et al. 2004). of the 215 million records in the global biodiversity information facility (gbif) data portal that correspond to museum specimens, 41.7% (90 million) do not have geographic coordinates associated. coordinate uncertainty has significant effects in quantitative analyses (graham et al. 2007, bloom et al. 2017, marcer et al. 2022); availability of such information, however, is even more limited than geographic coordinates. more than 79.1%, or 170 million, of specimens in the gbif data portal lack information on coordinate uncertainty. in short, without coordinates, a locality description remains unwieldy, subjective, and relatively imprecise in describing the geographic context of a specimen. the process of assigning geographic coordinates to a locality description, termed georeferencing, is a nuanced task. the point-radius method described in wieczorek et al. (2004) outlines and justifies a methodology behind the three numbers that correspond to the latitude, longitude, and coordinate uncertainty. wieczorek et al. (2004) use these three numbers to describe a circular area that ostensibly includes all geographic points from where the specimen might mailto:pete.campbell8@gmail.com https://orcid.org/0000-0003-3967-3109 campbell – interpreting and georeferencing the concept of “near” 14 have actually originated. this interpretation has become the gold standard, with leading data standards preferring such a point-radius representation (darwin core maintenance group 2023), although many efforts have been made to expand on the method and create best practices for georeferencing (wieczorek 2001, frazire et al. 2004, murphy et al. 2004, chapman & wieczorek 2006, guo et al. 2008, van erp et al. 2015, bloom et al. 2018). the most recent of these efforts are those of zermoglio et al. (2020) and chapman & wieczorek (2020). georeferencing is not a purely objective task, as it often requires the interpretation of imperfect, inexact, or incomplete locality descriptions. take, for instance, a locality description that reads, “found near the city of springfield,” which might be assigned a coordinate pair corresponding to the centroid of the city limits. nonetheless, uncertainties arise immediately: how close to springfield is “near”? could this description refer to a point within the city? although zermoglio et al. (2020) and chapman & wieczorek (2020) make many aspects of georeferencing straightforward and reproducible, their proposed techniques and solutions regarding these ideas of nearness or proximity can certainly be improved. the field of quantitative geography has an abundance of tools for diverse spatial analyses and inferences. particularly relevant to the question of proximity are algorithms for creating voronoi diagrams (also called thessian polygons or dirichlet tessellations). a voronoi diagram is a collection of polygons and points (or cells), wherein the polygons describe the geographic space that is closer to a particular point than to any other point in the collection of points. the polygons are referred to as voronoi polygons, which form a tessellation of the geometric space in question (okabe 2000; see example in figure 1). in this paper, i explore the potential use of voronoi diagrams to improve some georeferencing practices related to concepts of proximity. i use these methods to overcome the arbitrary or overly restrictive nature of existing methods, and create a method that addresses the characteristics of a previously undefined locality type. methods construction of a voronoi diagram all processing of shapes was done using the opensource gis platform qgis [ver. 3.22] and its prepackaged optional libraries. building collections of various points, lines, and polygons that describe localities was done based on esri shapefiles. shapefiles for cities and states were acquired from the u.s. census bureau (2021); shapefiles summarizing physical geographic features (e.g., rivers) were acquired from the u.s. geological survey (2019). collections of localities.—to create a voronoi diagram for georeferencing, i first found or assembled a single shapefile that provided a collection of localities categorically relevant to the described locality (e.g., the rivers surrounding a described river). an individual voronoi polygon is created in the context of its neighboring polygons, just as a locality description in a biodiversity record is created in a geographic context that conveys not only where the record was recorded, but also where it was not. for instance, if an individual scrub-jay (genus aphelocoma) was found in lagos de moreno, mexico, then that individual was not found in any other city or country on the globe. as such, the first step to creating a voronoi diagram for the purpose of georeferencing is to represent the desired locality as a member of a larger collection of contextually relevant localities (fig. 2. a). in doing so, the specified locality describes where the locality is, and the rest of the collection describes where the locality is not. as in figure 1, if one desires to find the voronoi polygon for fire station f, it is only possible when station f is considered relative to stations b and c. this problem of relativity means that it is important that collections be populated both accurately and precisely. an underpopulated collection will result in larger-than-necessary uncertainty figure 1. a voronoi diagram of hypothetical fire stations represented as points. each unique fire station is labeled with a letter and has a corresponding voronoi polygon. campbell – interpreting and georeferencing the concept of “near” 15 a b c d create a collection label localities densify extract vertices dissolve georeference voronoi polygons workflow for creating voronoi diagrams for georeferencing figure 2. a general workflow for creating voronoi diagrams for georeferencing. (a) a collection of localities relevant to x, all uniquely labeled. (b) an example of the necessity for densifying. the diagram on the left shows the voronoi polygon for shape x encroaching on shape y. point x1 is closer to parts of line y4,1 than either point y4 or point y1. the diagram on the right shows how densifying resolves this issue, by giving line y4,1 more representative points. (c) an intermediate voronoi diagram with one voronoi polygon per point, but several voronoi polygons per locality. (d) a final, post-dissolve, voronoi diagram. the voronoi polygon surrounding locality x is fit to be georeferenced. campbell – interpreting and georeferencing the concept of “near” 16 estimates downstream in the georeferencing process, and is therefore prone to type ii error. this sort of situation typically stems from poor sources of gis reference material (i.e. city boundary maps, collections of addresses, hydrology data, etc.). this issue can be ameliorated using official and comprehensive sources of data, such as national census data. an overpopulated collection will result in excessively confident uncertainty predictions, leading to type i errors and exclusion of potentially relevant geographic spaces. this phenomenon is a consequence of indiscriminate reference locality selection (e.g., including police stations when discussing fire stations, including secondary roads when referencing highways, etc.). a safe assumption to avoid this issue is to include in a collection the smallest class of described locality, and any class larger. for example, if the described locality is the confluence of two permanent streams, it is safe to include all streams and rivers in the collection, but intermittent streams and arroyos should be excluded, as they are smaller in magnitude. voronoi polygons are tessellations of the collective landscape, and as such, are defined strictly relative to the overall collection. converting the collection to points.—a voronoi diagram is created from a finite set of two or more distinct points in the euclidean plane (okabe 2000). if a collection is composed of shapes with higher dimensional complexity than points, then for this methodology the shapes should be reduced to points. while methods exist that can create voronoi diagrams from higher-dimensional input shapes such as lines, curved lines, and polygons (held 2001, culver et al. 2004, karavelas 2004), they are either deprecated, unnecessarily complicated, or too computationally intensive for the purposes of georeferencing. although voronoi diagrams can be created from collections of lines (okabe 2000, p. 169) or polygons (okabe 2000, p. 186), for my purposes aimed at georeferencing, i reduced such instances to points for a more universal and streamlined workflow that is compatible with common gis software. before creating points from a collection, i ensured that each locality targeted for georeferencing had a unique identifier, so that later processing stayed organized and traceable (richards et al. 2011). if a collection has to be reduced from a 2-dimensional shape (i.e., lines or polygons) to 1 dimensional points, the products developed from those points will need to be reassembled to accurately represent the original 2-dimensional form. if a collection was composed of lines or polygons, i densified the point-based representation of shapes to ensure that spaces between localities were assigned accurately. the geometry of 2-dimensional shapes is typically coded as a series of longitude and latitude pairs that represent vertices of the shapes. densifying adds additional points to the sequential list between the vertices that comprise the initial geometry description of a shape. to ensure that the farthest points of a shape’s voronoi polygon are accurate, the distance between the additional points along the shape’s line segments needs to be less than half the distance to the nearest foreign vertex (fig. 2. b). so, if the distance from shape x’s vertex to a line segment of shape y is 500 m, then shape y needs to have additional points no more than every 250 m along its line segments. in qgis, this step is achieved via the function “densify by interval.” creating a georeferencable voronoi diagram.—i used the curated or manufactured collection of points to seed an intermediate voronoi diagram. the voronoi tessellation algorithm in qgis is titled voronoi polygons, although other algorithms may refer to these same procedures as thiessen, voronoi, dirichlet polygons, tessellations, or diagrams (burrough et al. 2015, p.160). the qgis algorithm has two input parameters: the collection of points, and a measure of how far to expand the borders of the output voronoi diagram. the intermediate voronoi diagram shapefile resulting from the voronoi polygons algorithm consists of many polygons, with each unique polygon corresponding to one of the input points (fig. 2. c). rather than using hundreds of individual voronoi polygons to represent one original locality, these polygons can be grouped by their unique identifiers that they inherited from their input points, and merged to form one representative voronoi polygon. dissolving all the constituent/intermediate voronoi polygons together creates unified voronoi polygons that describe the geospatial area nearest each of the original input localities (fig. 2. d). these resulting shapes can be georeferenced like any of the standard localities mentioned in zermoglio et al. (2020); they can be easily described through both the point-radius method and in well known text (wkt) format. case study to demonstrate real-world applications of voronoi diagram-based methods, i georeferenced localities associated with a historical collection of tick records from kansas, representing collections campbell – interpreting and georeferencing the concept of “near” 17 assembled by the late prof. donald e. mock, kansas state university. between 1987 and 2003, mock advertised across the state of kansas, asking citizen scientists and ksu extension agents to mail ticks to him, in an attempt to catalog and understand distributions of various tick species. mock’s complete collection comprises 1551 records of a total of more than 6600 ticks. given the nature of the citizen scientist contributions, about 30% of the localities described in the collection are written using vague, relative, or imprecise terms. rather than discard these specimen descriptions, i sought to georeference them. using voronoi diagrams, i report on localities in the following categories: “near a feature,” “offset heading,” “paths, roads, and rivers,” “intersections, confluences, junctions, and crossings” and “on/near the border of.” i have used different sorts of georeferencing challenges manifested in this dataset to exemplify the utility of voronoi diagrams, as described below. it should be emphasized that the goal of these methods is not to increase or decrease uncertainty, as neither outcome is indicative of a better or worse technique; rather, the goal is to make existing techniques non-arbitrary and less assumption-laden in their interpretation. near a feature.—voronoi diagrams can be used to define geospatial areas that are more “near” to a given locality than to any other locality in a collection. imprecise locality descriptions can come with a preface of proximity such as: near, around, off the coast of, etc. (zermoglio et al. 2020). as an example from mock’s collection, one locality reads “... acquired farming near [the city of] salina…” the current methodology for georeferencing locality descriptions including the word “near” is to arbitrarily inflate (i.e., buffer) the borders of a locality, and georeference the result (zermoglio et al. 2020). rather than try to guess arbitrarily at the original author’s precision, i propose systematically defining the area that is nearest to the locality (fig. 3). in this case, the voronoi polygon associated with the city of salina will describe the geographic area closer to salina than to any other similar-sized or larger city in the regions around it. interpreting someone else’s writing can never be 100% objective, and some assumptions will be necessary regarding the original author’s intent. still, georeferencing the literal interpretation of the locality description in this way accommodates conservative interpretations of uncertainty (like within the city boundaries or the majority of mailing addresses), while remaining as objective as possible about the text to avoid being arbitrary. using voronoi diagrams to interpret the concept of “near” thus makes the process of georeferencing a vague locality description more systematic and the interpretation more consistent. offset heading.—voronoi polygons can also provide bounds within which relative headings can be georeferenced, as in the case of locality descriptions combining a locality and a heading. an examfigure 3. current practices vs proposed voronoi method. (left) a simple buffer arbitrarily made at 1km, constructed following current standard practice vs (right) voronoi polygons constructed using the vertices of each city’s polygon. campbell – interpreting and georeferencing the concept of “near” 18 ple would be a specimen recorded as “found north of salina.” although this description is much more precise than the “near a feature” example, it still remains vague and imprecise. the locality and heading act as a pair that combine to form a cone. the center of the named locality represents the apex of the cone, whereas the heading description determines the cone’s direction and arc. the current standard for georeferencing localities with a heading is to create this cone and extend it in the direction of the heading “...until reaching a constraining boundary imposed by other information in the locality record, or until reaching the proximity of another similar feature” (zermoglio et al. 2020). the methodology behind the standard just described is similar to the voronoi method, in that both rely on relevant locality neighbors to define the outer limits of the new representative shape being georeferenced. however, the current method is not explicit about where these limits lie. rather than relying on the discretion of the interpreter, voronoi polygons can define the “proximity” of neighboring localities objectively, and provide a constraining boundary. a voronoi diagram will define all of the area bordering all relevant localities, not just the first in the path of the cone. the intersection of salina’s voronoi polygon and the northward cone then explicitly describes the area only found north of salina, as opposed to west of new cambria, southeast of culver, or northeast of bavaria (see fig. 4). this slice of the voronoi polygon can then be georeferenced using standard protocols and described precisely using wkt formats. paths, roads, and rivers.—paths, roads, and rivers can be difficult to georeference accurately owing to their long, thin geometry. a simple point-radius figure 4. a voronoi diagram based around the city of salina. the purple lines show the boundaries of each nearby city’s voronoi polygon. the black-dashed, purple shaded region shows the subsection of salina’s voronoi polygon north of the city. the cone used to create this region starts at salina’s centroid (38.8147, -97.6110) and extends northward at the arc suggested in wieczorek et al. (2004). the resulting polygon has a centroid at (-97.6201, 38.8993), and a coordinate uncertainty of 9420 m. campbell – interpreting and georeferencing the concept of “near” 19 measurement based on a centroid does not describe well the uncertainty surrounding the terminal ends of a linear path due to its strictly non-circular shape, such that uncertainty is likely over appreciated in the middle and underappreciated at the ends. rather than use an arbitrary buffer to partially compensate for this problem at the terminal ends, a voronoi diagram describes the area closer to one path than to any other in a collection, which includes appropriate treatment of the ends of the path (fig. 5). creating a voronoi diagram for a collection of lines or polylines is straightforward, and follows the same process as for a collection of polygons as described above. while a voronoi polygon can be a useful tool for georeferencing paths, it may still fail to overcome the inherent linearity of the locality at the coarsest scales. voronoi polygons for exceedingly long paths (the mississippi river, an interstate railroad, interstate highways, etc.) may still not be described well by the point-radius method, which may highlight the utility of the footprintwkt darwin core field. in the end, if an excessive path is not described as hemmed in by some other type of border (e.g., found along the kansas river in douglas county), then the resulting uncertainty will be high no matter what metric is used. intersections, confluences, junctions, and crossings.—describing the area corresponding to an intersection or confluence using voronoi polygons is a more accurate and realistic interpretation of a locality description for such crossings. consider the locality description, “143rd st. and evening star rd.” this locality description taken literally would imply that a specimen was found somewhere in the middle of the 16 m2 section of pavement that makes up that particular intersection. indeed, that is the interpretation that is the current best practice for georeferencing crossings: to measure the shape that the two paths create via their overlap. although road kills represent a common source of specimen collections, in many cases, it is more reasonable to assume that a collector is using an intersection as a reference point and is collecting specimens nearby. using a voronoi diagram to describe areas surrounding intersections as part of a collection allows the georeferencer to define the area that is closest to each individual intersection. in effect, voronoi diagrams systematically and quantitatively enlarge areas of uncertainty associated with crossings according to the proximity of other such crossings in the collection under analysis. this alternative aims to err figure 5. an example of a voronoi diagram based on rivers. the red lines indicate the borders of each river’s (blue) voronoi polygon. notice the rivers only touch voronoi polygon borders when converging with another river. campbell – interpreting and georeferencing the concept of “near” 20 towards overestimating uncertainty, rather than attribute too much confidence to an overly narrow locality description (fig. 6). an intersection’s voronoi polygon includes the original standard pavement section, but would also describe the greater area across which a collector might consider that particular intersection to be a reasonable reference point. on / near the border of.—one type of locality without a specific method for georeferencing in zermoglio et al. (2020) is a border, such as “...killed by a rancher on osage/franklin co line.” this locality description effectively describes a 35 km line that spans the length of the border between osage county and franklin county. treating this line as a member of a collection made up of the remaining border of the two counties gives a discrete voronoi area for referencing. this is possibly the most consistent use for voronoi diagrams in georeferencing, as the collection is typically pre-defined by two 2-dimensional shapes (e.g., political boundaries). the desired voronoi polygon is then a subset of the combined area of the two shapes, representing all geographic areas closer to the border of interest than to any other border (fig. 7). figure 6. a comparison of current practice vs voronoi diagram interpretation of intersections. the orange circles indicate the current interpretation of an intersection, being the centroid of the two roads and the distance to the farthest corner of the intersection. the right panel shows how voronoi diagrams include the original orange circle interpretation, but expand on it and accommodate the surrounding environment. figure 7. an example of how a voronoi polygon can be used to quantify the area around a border. the solid black lines show the exteriors of osage county and franklin county. the dashed black line indicates the border adjoining the two counties, which is to be georeferenced. the blue polygon shows the voronoi polygon of this border, with the exterior borders acting as relevant localities in the collection. campbell – interpreting and georeferencing the concept of “near” 21 discussion assumptions voronoi polygons are a systematic and more objective approach to georeferencing, but it is ultimately impossible to escape some level of interpretation. a description that reads, “near the city of lawrence,” implies that the described sample is not as near to any other city as it is to lawrence. note, however, that i assumed that when a locality is described as a city, it negates the description of other cities. but, what if the author would have described the locality as the nearby lake if the collector had known about the lake’s presence? in such a case, my assumption produces larger uncertainty areas than what the collector was describing, because the collection of cities remains incomplete without lakes as well. a locality description that mentions a single specific type of locality does not inherently ignore any other category of locality. there is also the issue of magnitude. in the described process for making collections, i assume that all localities in a collection are of equal likelihood to be referenced. it is reasonable to imagine a scenario in which a sample is collected in shawnee (a suburb of kansas city), yet the locality might be described as, “near kansas city, kansas.” localities of larger physical extent or higher population are typically more commonly known, and are thus more likely to be thought of or used in writing a locality description. for this reason, i suggest adding all localities of the referenced locality’s magnitude or larger when creating a collection. put more simply, if someone knows about the city of chicago, illinois, they may or may not know about the neighboring city of elmhurst, illinois but, if someone knows of elmhurst, they almost certainly know about chicago. one possible solution to this issue may be the use of weighted voronoi diagrams. a weighted voronoi diagram is created from points with an additional value that varies based on the property being weighted (e.g., population size, square mileage, # of references, strahler stream order, etc.; okabe p. 119). rather than every locality in a collection being territorially isolated, a weighted collection could depict more commonly recognized localities as having significantly larger polygons that engulf much less recognized localities. the new weighted voronoi polygons thereby simulate the “influence” that some localities might have over the average collector’s perception of the “nearest” locality. more simply, although voronoi diagrams can be helpful in systematically georeferencing past efforts, they can not fully remove the subjective nature of interpreting someone else’s work. shortcomings voronoi diagrams rely heavily on the quality of the collection being used, and a sparse collection will make for imprecise locality data, particularly in the sense of being overly general. creating a collection composed of scarce locality types can make it difficult to fill a collection objectively. for instance, lakes in the state of kansas are relatively few. although many locality descriptions are simply the name of a lake, like “tuttle creek lake,” a tick is certainly not collected from the middle of the water, so the description must be referring to the area surrounding and associated with tuttle creek lake. however, only 0.6% of the land area of kansas is covered in water, making it the 4th driest state in the us (u.s. census bureau 2010), so a collection of lakes in kansas will be quite sparse, and the corresponding voronoi polygons will be enormous. on the other hand, michigan, which is 41.5% water, might have the opposite issue of magnitude, where a superabundance of small lakes could restrict the voronoi polygons of larger lakes that would more immediately draw a reference. although the methods described herein could be considered conceptual improvements on existing methods, they are still interpretations of inexact locality descriptions. the voronoi diagram methods do not necessarily increase or decrease coordinate uncertainty compared to existing methods, as the existing methods produce arbitrary uncertainty measures depending on explicit assumptions. a voronoi polygon could produce a larger or smaller uncertainty than an arbitrarily defined distance from a centroid, but conceptually, it has a contextual basis defined by the collection of localities and their spatial distributions. however, the collection on which the voronoi methods are based is still an interpretation of the locality description’s intentions. voronoi methods do not attempt to interpret another’s writing perfectly, only to make said interpretation more methodical and inclusive. future developments a promising opportunity for reporting information associated with voronoi polygons lies in the footprintwkt darwin core field (darwin core maintenance group 2021). this field holds an exact textual description of a spatial polygon that can be campbell – interpreting and georeferencing the concept of “near” 22 read precisely in a gis. currently, most data records that do provide uncertainty information provide only the three-value combination associated with the point-radius method (i.e., latitude, longitude, and uncertainty in meters), which reduces the complexity of a locality’s polygon down to a circle. though highly reproducible and simple, such a description is often too cursory. the implication of coordinate uncertainty measures associated with the point-radius method is that the specimen has an equal chance of having come from any part of the described area. the point-radius method is particularly poor in describing non-circular localities, like concave shapes or lines. the footprintwkt field circumvents this issue by explicitly defining the shape in question with no reduction in precision. novel methods have been developed around specimen polygons (peterson et al. 2006, smith et al. 2023), but without special effort of the investigator to create these shapefiles or footprint fields, the resources necessary to implement these methods remain unavailable. acknowledgments i would like to thank dr. c. olds (kansas state university) for locating the mock tick dataset, and for providing high-quality scans. i thank dr. a. t. peterson for crucial insight, vital resources, and unwavering confidence. this work was supported by the national science foundation (oia-1920946). references bloom, t. d. s., a. flower, and e. g. dechaine. 2018. why georeferencing matters: introducing a practical protocol to prepare species occurrence records for spatial analysis. ecology and evolution 8:765–777. burrough, p. a., r. mcdonnell, and c. d. lloyd. 2015. principles of geographical information systems, 3rd ed. oxford university press, oxford. chapman, a.d. and j. wieczorek. 2006. guide to best practices for georeferencing. copenhagen: global biodiversity information facility. chapman, a.d. and j. wieczorek. 2020. guide to best practices for georeferencing. copenhagen: global biodiversity information facility. culver, t., j. keyser, and d. manocha. 2004. exact computation of the medial axis of a polyhedron. computer aided geometric design 21:65–98. darwin core maintenance group. 2023. list of darwin core terms. biodiversity information standards (tdwg). http:// rs.tdwg.org/dwc/doc/list/2023-07-07 frazier, c., t. neville, j. t. giermakowski, & g. racz (2004). the inram protocol for georeferencing biological museum specimen records (1.3). zenodo. https://doi. org/10.5281/zenodo.3235003 graham, c. h., j. elith, r. j. hijmans, a. guisan, a. t. peterson, b. a. loiselle, and the nceas predicting species distributions working group. 2007. the influence of spatial errors in species occurrence data used in distribution models: spatial error in occurrence data for predictive modelling. journal of applied ecology 45:239–247. guo, q., y. liu, and j. wieczorek. 2008. georeferencing locality descriptions and computing associated uncertainty using a probabilistic approach. international journal of geographical information science 22:1067–1090. held, m. 2001. vroni: an engineering approach to the reliable and efficient computation of voronoi diagrams of points and line segments. computational geometry 18:95–123. karavelas, m. 2004. a robust and efficient implementation for the segment voronoi diagram. international symposium on voronoi diagrams in science and engineering. 51-62. marcer, a., a. d. chapman, j. r. wieczorek, f. xavier picó, f. uribe, j. waller, and a. h. ariño. 2022. uncertainty matters: ascertaining where specimens in natural history collections come from and its implications for predicting species distributions. ecography 2022: e06025. https://doi. org/10.1111/ecog.06025 murphy, p. c., r. p. guralnick, r. glaubitz, d. neufeld, and j. a. ryan. 2004. georeferencing of museum collections: a review of problems and automated tools, and the methodology developed by the mountain and plains spatio-temporal database-informatics initiative (mapstedi). phyloinformatics 3: 1-29. https://doi.org/10.5281/zenodo.59792 okabe, a. 2000. spatial tessellations: concepts and applications of voronoi diagrams. 2nd ed. wiley, chichester; new york. peterson, a. t., r. r. lash, d. s. carroll, and k. m. johnson. 2006. geographic potential for outbreaks of marburg hemorrhagic fever. the american journal of tropical medicine and hygiene, 75(1), 9–15. https://doi.org/10.4269/ajtmh.2006.75.1.0750009 richards, k., r. white, n. nicolson, and r. pyle. 2011. a beginner’s guide to persistent identifiers. global biodiversity information facility, copenhagen. smith, a. b., s. j. murphy, d. henderson, and k. d. erickson. 2023. including imprecisely georeferenced specimens improves accuracy of species distribution models and estimates of niche breadth. global ecology and biogeography 32:342–355. u.s. census bureau. 2021. cartographic boundary files shapefile. http://www.census.gov/geographies/mapping-files/ time-series/geo/carto-boundary-file.html. u.s. census bureau. 2010. state area measurements and internal point coordinates. https://www.census.gov/geographies/reference-files/2010/geo/state-area.html. http://rs.tdwg.org/dwc/doc/list/2023-07-07 http://rs.tdwg.org/dwc/doc/list/2023-07-07 http://rs.tdwg.org/dwc/doc/list/2023-07-07 https://doi.org/10.5281/zenodo.3235003 https://doi.org/10.5281/zenodo.3235003 https://doi.org/10.1111/ecog.06025 https://doi.org/10.1111/ecog.06025 https://doi.org/10.5281/zenodo.59792 https://doi.org/10.4269/ajtmh.2006.75.1.0750009 https://doi.org/10.4269/ajtmh.2006.75.1.0750009 https://www.census.gov/geographies/mapping-files/time-series/geo/carto-boundary-file.html https://www.census.gov/geographies/mapping-files/time-series/geo/carto-boundary-file.html https://www.census.gov/geographies/reference-files/2010/geo/state-area.html https://www.census.gov/geographies/reference-files/2010/geo/state-area.html campbell – interpreting and georeferencing the concept of “near” 23 u.s. geological survey. 2019. usgs tnm hydrography (nhd). https://apps.nationalmap.gov/services/. van erp, m., r. hensel, d. ceolin, and m. van der meij. 2015. georeferencing animal specimen datasets: georeferencing animal specimen datasets. transactions in gis 19:563–581. wieczorek, j. 2001. manis/herpnet/ornis georeferencing guidelines. http://georeferencing.org/georefcalculator/ docs/georefguide.html. wieczorek, j., q. guo, and r. hijmans. 2004. the point-radius method for georeferencing locality descriptions and calculating associated uncertainty. international journal of geographical information science 18:745–767. zermoglio, p., a. chapman, j. wieczorek, m. c. luna, and d. bloom. 2020. georeferencing quick reference guide. copenhagen: gbif secretariat. https://doi.org/10.35035/ e09p-h128 https://hydro.nationalmap.gov/arcgis/rest/services/nhd/mapserver https://hydro.nationalmap.gov/arcgis/rest/services/nhd/mapserver http://georeferencing.org/georefcalculator/docs/georefguide.html http://georeferencing.org/georefcalculator/docs/georefguide.html https://doi.org/10.35035/e09p-h128 https://doi.org/10.35035/e09p-h128 biodiversity informatics, 17, 2022, pp. 50-58 50 detecting signals of species’ ecological niches in results of studies with defined sampling protocols: example application to pathogen niches marlon e. cobos*, a. townsend peterson department of ecology and evolutionary biology & biodiversity institute, university of kansas, lawrence, kansas 66045, usa *corresponding author: marlon e. cobos, email: manubio13@gmail.com abstract. ecological niches are increasingly appreciated as a long-term stable constraint on the geographic and temporal distributions of species, including species involved in disease transmission cycles (pathogens, vectors, hosts). although considerable research effort has used correlative methodologies for characterizing niches, sampling effort (and the biases that this effort may or may not carry with it) considerations have generally not been incorporated explicitly into ecological niche modeling. in some cases, however, the sampling effort can be characterized explicitly, such as when hosts are tested for pathogens, as well as comparable situations such as when traps are deployed to capture particular species, etc. here, we present simple methods for testing the hypothesis that non-randomness in occurrence or detection exists with respect to environmental dimensions (= a detectable signal of ecological niche); i.e., whether a pathogen occurs nonrandomly with respect to environment, given the occurrence and sampling of its host. we have implemented a set of r functions that presents an overall test for nonrandom occurrence with respect to a set of environmental dimensions, and, a posteriori, a set of exploratory tests that identify in which dimension(s) and in which direction or form the nonrandom occurrence is manifested. our tools correctly detected signals of niche in most of our example cases. although such a signal may not be detectable in cases in which the niche of interest is broader than the universe sampled, such a possibility was correctly discarded in our analyses, preventing further interpretations. this kind of testing can constitute an initial step in a process that would conclude with development of a more typical ecological niche model. the particular advantage of the analyses proposed is that they consider the biases involved in sampling, testing, and reporting, in the context of nonrandom occurrence with respect to environment before proceeding to inferential and predictive steps. key words: ecological niche, host, niche position, niche breadth, non-parametric test, permanova the ideas, tools, and methods used under the rubrics of “ecological niche modeling” and “species distribution modeling” (here referred to as enm/sdm), have seen extensive application to understanding the geographic and environmental distributions of species (franklin 2010; peterson et al. 2011). most popular have been correlative approaches, in which environmental characteristics of places of known occurrences of species are subjected to a variety of model-fitting approaches, to create a classification of different parts of environmental space into suitable and unsuitable sets of conditions (peterson et al. 2011; enriquez-urzelai et al. 2019). a major challenge for these methods, however, has been the pervasive biases and gaps that characterize the sampling that produced the primary occurrence data, and how to avoid propagation of those biases through the analytical sequence to the results (anderson and gonzalez 2011; acevedo et al. 2012; araújo et al. 2019). some primary biodiversity occurrence data, however, may be connected to information that can characterize the sampling universe integrally. such data may take the form of occurrences of pathogens detected by testing hosts (e.g., eisen and paddock 2021), disease case data that come marlon e. cobos et al. – detecting signals of species’ ecological niches 51 from active surveillance (e.g., m’ikanatha et al. 2008), biodiversity data that are accumulated by trapping where trap data are recorded (meek et al. 2015), and biodiversity data that are accumulated by standardized sampling protocols (manley et al. 2005). in each case, the geographic and temporal distribution of sampling can be characterized precisely, and all positive records of the species of interest must necessarily derive from one of the sampling events. this additional information offers considerable promise in informing the modeling process precisely about the sampling universe, rather than relying on assumptions of random sampling (phillips et al. 2009) or an interpolated sampling bias surface (warren et al. 2014). to our knowledge, exploring signals of ecological niche differentiation in data for which the sampling universe is known has not been done before. most studies use traditional approaches to characterize ecological niches of species and compare such niches without explicit consideration of the sampling universe. in this contribution, we present a logic for a suite of analyses designed to take advantage of this additional information (the sampling universe) available for occurrence data that come from such controlled sampling schemes. we provide a methodological protocol that first tests for any overall niche difference, and then characterizes these differences in terms of a spectrum of possible changes in niches in each environmental dimension. we present the protocol in the form of a set of r functions, to facilitate wide use and incorporation in many other analyses. protocol description we offer two complementary approaches to detect signals of niche: (1) a multivariate analysis based on a permutational multivariate analysis of variance (permanova; (anderson 2017), and (2) a univariate non-parametric method based on descriptive statistics. to illustrate the utility of this approach, we use a suite of virtual species. for each, we created a sample of records that represent the universe of sampling (e.g., a host species that is to be tested for a particular pathogen), and then identify a subgroup of those records that may or may not be positive for the pathogen (see example application). that is, the data required to perform these analyses consist of a set of records representing the sampling universe, to which a test is applied that determines presence or absence of the species of interest (0 = negative and 1 = positive). each of these records carries with it a vector of relevant environmental conditions (see example in table s1). a typical such situation would be sampling a host species and testing for presence of a pathogen in each host, but many parallel applications exist. each host record has a geographic reference and potentially also information about collection time—this place and time information can be used to extract environmental data that is place-specific or place-and-time-specific from diverse raster data layers (e.g., data on climate, remote-sensing information, etc.) that are relevant in niche characterization (ingenloff and peterson 2021). multivariate test as a multivariate test to detect overall signals of niche, we propose an approach using a permanova. permanova is a non-parametric multivariate test that allows comparison of samples by testing a null hypothesis (h0) that the position and dispersion of the sample are equivalent to those of the sampling universe. rejecting h0 indicates that either the centroid (position) or spread (dispersion) is different, which would be indicative of a niche in the pathogen distinct from that of the host. similarity among groups is tested based on distances (e.g., euclidean or mahalanobis distances). in this application, the groups to be compared are records of the host of which a few are infected (i.e., all host records vs records of infected hosts). we chose to base our permanova analyses on mahalanobis distances, but other methods to calculate dissimilarities can be used to perform these processes (e.g., euclidean distances, bray-curtis, and jaccard indices; see code documentation). the permanova yields a result of significant or not, indicating rejection or acceptance of the null hypothesis of equivalency, but does not characterize the form of those differences. for this reason, the univariate tests described below are used to provide additional information about the form of these differences. univariate test to characterize niches in individual environmental dimensions, we use a comparison of observed values of descriptive statistics summarizing characteristics of distributions of environmental conditions associated with known-infected hosts marlon e. cobos et al. – detecting signals of species’ ecological niches 52 against a null distribution of those statistics derived from many similar-sized random samples drawn from the set of all host records. the descriptive statistics explored and used in this approach are the mean, median, standard deviation (sd), and range. the mean and median of the environmental values are used as estimators of niche position; the sd and the range are descriptors of the spread of environmental conditions comprising the niche. a null distribution of statistics is derived from many random samples (of size matching the number of positive tests) drawn from the set of all host records; this distribution informs about how common certain values of the descriptive statistics would be if the pathogen had no particular preference or bias from among the set of environmental conditions used by the host. therefore, the null hypothesis (h0) for this test is that the descriptive statistic calculated for the samples in which the pathogen was detected cannot be distinguished from the comparable statistic for the host (or the sampling universe). this h0 is tested for the mean, median, sd, and range, as different measures of characteristics of the distribution. the following is a sequential description of steps involved in running this analysis: 1. the number of infected host records is calculated (ni). 2. the mean, median, sd, and range of environmental conditions for the infected hosts are calculated. 3. a random sample of ni records is drawn from the entire set of host records. 4. the statistics of interest of environmental conditions are calculated for the sample in step 3. 5. steps 3 and 4 are repeated nt times (iterations; generally nt = 1000). 6. the full distributions of the statistics of interest are compiled and characterized, particularly as regards the 2.5% and 97.5% levels of the distribution. 7. the observed value of the statistic of interest for infected hosts (step 2) is compared against the null distribution of values (step 6) to establish whether it falls in the central 95% of the null distribution. 8. depending on the results from step 7, the statistic of interest for the niche of the pathogen is categorized as different or not from null expectations. 9. the direction of the difference is characterized by direct inspection to establish whether the pathogen’s niche is shifted upward or downward in the values of the particular environmental dimension (mean or median), or whether it has broadened or narrowed (sd or range). the results obtained from these steps allow us to accept or reject h0 (fig. 1). to reject h0, the value of the statistic observed for the positive records must be as extreme or more extreme than the 2.5% or 97.5% of the null distribution. we conclude that the pathogen niche (in terms of the statistic under test) is not distinct if h0 cannot be rejected. when h0 is rejected, the statistic under test can be lower or higher depending on in which tail of the null distribution the observed value falls. software we created a set of r functions to run the analyses described above. to aid interpretation, we also created functions to plot results from analyses. these functions are open-source tools that can be accessed following indications in software availability. proper documentation describing the data required to run analyses and how parameter values can be established is provided with the r scripts. example application example data to explore and test the performance of the protocols described above, we generated virtual niches for a host and seven simulated pathogens (figs. 2, s1) representing a distinct scenario of similarity of host and pathogen niches (table 1). one case (scenario 1) was designed to have a pathogen figure 1. representation of outcome and suggested interpretation of results from the univariate non-parametric test to detect signals of niche dissimilarity. in this case, we present mean temperature responses for scenario 5; the null hypothesis would be rejected, in favor of an alternative hypothesis of higher-than-null mean temperature. marlon e. cobos et al. – detecting signals of species’ ecological niches 53 with exactly the same niche as the host, in the other scenarios, the host and pathogen niches overlap in different ways (see figs. 2, s1). virtual niches were generated in r 4.1.1 (r core team 2021) using the package “evniche”1, which uses ellipsoids to create niches based on user-defined limits (variable ranges) and covariance values. we considered annual mean temperature and annual precipitation as the dimensions of our virtual niches. to make our simulations more realistic, when generating data from virtual niches, we considered suitability values derived from mahalanobis distances (based on multivariate normal distributions) to the centroids of the ellipsoidshaped niches, measured from points present in available environmental conditions in a region (for details see etherington 2019; nuñez-penichet et al. 2021). using ellipsoids and the multivariate normal transformations generates responses that are simple, symmetrical, and convex, which we consider appropriate to represent virtual fundamental niches; however, we emphasize that our methods are general, and do not depend on assumptions of normality, regardless of whether our example application makes such assumptions. as a result, the density of records generated from virtual niches increases towards the centroid of the ellipsoids, but it will also depend on the density of points representing available conditions across the area of analysis. available environmental conditions were represented by values of annual mean temperature and annual precipitation present across south america. we used two of the socalled “bioclimatic” data layers from the worldclim 1https://github.com/marlonecobos/evniche. host / pathogen scenario description temperature range (°c) precipitation range (mm) covariance pathogen prevalence host 12–26 700–2800 t: 5.44 p: 122,500 pathogen 1 pathogen niche is equal to host niche 12–26 700–2800 t: 5.44 p: 122,500 0.52 pathogen 2 pathogen niche has the same size as host niche but with changed position 14–28 800–2900 t: 5.44 p: 122,500 0.38 pathogen 3 pathogen niche changed in position and size (smaller) compared to host niche 13–22 900–2600 t: 2.25 p: 8,277 0.61 pathogen 4 pathogen niche smaller than host niche 15–23 1000–2500 t: 1.78 p: 62,500 0.68 pathogen 5 pathogen niche smaller and changed in position 18–25 1000–4000 t: 1.36 p: 250,000 0.47 pathogen 6 pathogen niche larger than host niche, but overlaps most of it 14–29 600–4500 t: 6.25 p: 422,500 0.29 pathogen 7 pathogen niche larger than host niche but contains it completely 8–30 200–3200 t: 13.44 p: 250,000 0.30 table 1. general description of parameters that define the host and pathogen virtual niches according to distinct scenarios of niche similarity. values of resulting pathogen prevalence in the host are also shown in the table. t = temperature; p = precipitation. figure 2. virtual niches (ellipses) of a host and 7 pathogen scenarios used in the example application. points in black and red represent records of host and pathogen, respectively, as if they were obtained from geographic records. host and pathogen records were derived from ellipses, and are overlaid on environmental conditions across south america (gray points). https://github.com/marlonecobos/evniche marlon e. cobos et al. – detecting signals of species’ ecological niches 54 database v1.4 (hijmans et al. 2005) and masked them to south america to perform the analyses described above. raster processing was done using the package “raster” (hijmans 2019). we generated populations of points via sampling from the centrality-weighted ellipsoids for the host and the pathogen niches separately, and then used as “pathogen-positive” records those host-niche points that coincided exactly with pathogen-niche points; for this purpose, we generated 200 host-niche points and 400 pathogen-niche points. because in scenario 1, host and pathogen niches were exactly the same, we simply subsampled 200 points from among the pathogen 400 points to exclude some of the host records from being considered as infected. because records derived from ellipsoids with distinct sizes and positions in the cloud of available environmental conditions for host and pathogens, we were not able to control pathogen prevalences (see table 1). niche comparisons we compared host and pathogen ecological niches considering the 7 pathogen-niche scenarios using both the multivariate and univariate approaches. multivariate comparisons were made using permanova analyses with 1000 iterations for calculation of statistical significance. for univariate comparisons, the mean, standard deviation, and range of values corresponding to infected hosts were compared to the distribution of the same statistics for 1000 random samples from the host records. to aid with interpretation, we created ellipsoids for the environmental distributions of the host and all pathogens. for pathogens, we considered the data used in analyses (i.e, not all records generated using virtual niches of pathogens, but rather only those of pathogens that match the host). we plotted all ellipsoids derived from the data to explore and visualize the position and spread of host and pathogen niches. results final datasets prepared for analysis consisted of 200 records of the host, of which 57-136 matched virtual pathogen records, and thus were considered as infected hosts (see table 1 for pathogen prevalences; see example dataset in table s1). environmental representations of datasets showed distinct levels of overlap between host records and infected ones, which helped us to understand the actual configurations of host and pathogen records that can be observed in real applications (fig. 3). no signal of a distinct pathogen niche was detected in 3 of the 7 scenarios using the permanova (scenarios 1, 6, and 7; fig. 4). that is, based on the multivariate analysis, the centroid and dispersion of infected hosts can be considered as non-distinguishable from those of all hosts for scenarios 1, 6, and 7. univariate analyses further indicated that host and pathogen in scenarios 1 and 7 were not distinct in any individual dimension. for all other scenarios, some signal of dissimilarity was detected (table 2). for these cases, the observed mean, median, standard deviation, or range derived from individual variable values of infected hosts fell outside of the central 95% of the null distribution of values derived from 1000 random samples of all hosts for one or both of the environmental dimensions (figs. s2-s5). considering the qualities of the 7 scenarios, both methods could not reject h0 when niches were exactly equal (scenario 1), or when the pathogen niche contained completely that of the host (scenario 7). however, the permanova did not detect niche dissimilarity for scenario 6, even though the original host and pathogen niches were different. both methods identified a signal of dissimilarity for scenario 2, yet the niche of the pathogen had only a slight change in position from that of the host. discussion both multivariate and univariate approaches performed well in detecting signals of niche dissimilarity in cases in which the pathogen niche represented a subgroup of that of the host (or the sampling universe). the two tests are complementary in the sense that they are based on different procedures and ideas, but both help to detect signals and interpret the type of signal detected. the permanovabased test seeks an overall signal of niche difference, and also considers covariation among variables, although a direct understanding of the sort of difference manifested does not derive directly from this test. the univariate analyses, in contrast, allow one to understand changes in niche position and breadth when an overall signal is detected, although it does not consider covariation among multiple environmental variables. graphical representations of results help considerably with interpretation and complement further the understanding of the signals detected. apart from the obvious differences between the univariate and multivariate analyses, another marlon e. cobos et al. – detecting signals of species’ ecological niches 55 figure 4. results from niche comparisons using permanova analyses. ellipses were reconstructed from the data created from virtual niches and the available background. values of statistical significance are shown for each comparison. figure 3. visualization of data derived from virtual niches representing hosts infected under 7 pathogen-niche scenarios. this view shows how data would look if derived from sampling the host and testing for pathogens, with no previous knowledge of host or pathogen niches. niche differences between host and pathogen can be noted as conditions under which hosts are not marked as infected, such as above 22°c in scenario 3. marlon e. cobos et al. – detecting signals of species’ ecological niches 56 important difference should be noticed. in the univariate analyses, as summary statistics from pathogen testing data are compared against a null distribution derived from sampling host data, conditions analogous to accessible environments (soberón and peterson 2005) are considered in the univariate non-parametric approach. that is, the univariate approaches are considering the set of environmental conditions for all hosts as those to which the pathogen could have had access. if some conditions used by the host have no pathogen records, it may be because of niche-based limitations, and these methods assess how nonrandom environmentally those gaps in pathogen records are. these limitations may be related to pathogen tolerance of or preference for certain environments, or environmental conditions limiting pathogen transmission from one host to another. they could also relate to inappropriate delimitation of relevant hosts for analysis, such as if a pathogen were recently introduced in a region, and has not yet reached all areas inhabited by the host. this last detail is important because critical biases can be introduced in analyses if the data have not been filtered carefully based on ecological and biological considerations (barve et al. 2011; machado-stredel et al. 2021). our protocols did not detect clear signals of dissimilarity in scenarios 6 and 7, in which the pathogen niche was larger than and overlapped with or included the host niche. this outcome derives from the type of information available for analysis—a set of host records of which some are infected—which makes it difficult to detect signals of dissimilarity because pathogen niches will be characterized incompletely. a more comprehensive characterization of pathogen niches may require consideration of a larger group of host species, which could inform about the type of conditions that are suitable or unsuitable for a pathogen. however, the fact that this latter set of information will be scarce in real applications highlights the utility of our protocols in exploring signals of niche in pathogens. we note that a topic of current interest is that of coinfections of multiple pathogen species (collinge and ray 2006)—although the current implementation of our methods is in terms of single pathogen species, a clear potential extension is that of simultaneous evaluation of environmental bias in distributions of multiple pathogen species. although we have presented these protocols in the context of tests for pathogen infections in host organisms, as mentioned in the introduction, these methods can be useful in any situation in which (1) the entire universe of sampling can be characterized, and (2) the set of positive records will be a strict subset of that universe of sampling. this situation comparison variable pathogen niche mean vs null distribution pathogen niche median vs null distribution pathogen niche sd vs null distribution pathogen niche range vs null distribution scenario 1 temperature – – – – precipitation – – – – scenario 2 temperature – – lower lower precipitation higher higher – – scenario 3 temperature lower lower lower lower precipitation – – lower lower scenario 4 temperature lower lower lower lower precipitation – – lower lower scenario 5 temperature higher higher lower lower precipitation higher higher – – scenario 6 temperature – – lower – precipitation higher – – lower scenario 7 temperature – – – – precipitation – – – – table 2. summary of results derived from univariate niche comparisons. comparisons identified as higher or lower were statistically significant (i.e., observed values from the pathogen were as extreme or more extreme than the central 95% of the null distribution). marlon e. cobos et al. – detecting signals of species’ ecological niches 57 would be manifested in cases such as an analysis based on a single sampling protocol (e.g., data from the u.s. breeding bird survey; sauer et al. 2013), or deriving from a single-investigator sampling protocol (e.g., regional trapping of insects, such that trap positions are known completely; sciarretta and trematerra 2014). as such, this set of approaches can be considered as a precursor to formal ecological niche modeling, testing at the outset whether any nonrandom environmental use (= ecological niche) is manifested by that species, at that extent, at that resolution, and in those environmental dimensions. acknowledgments we thank our colleagues eric ngeno, ram raghavan, and abdelkafar alkishe for the case studies that led to development of this paper. this research was supported by a grant from the national science foundation (oia-1920946). competing interests the authors have declared that no competing interests exist. software availability r scripts containing the functions needed to run analyses and plot results are provided in a github repository2. data availability the data required to run the analyses presented in this work can be obtained using the script available at github3. supplementary materials all supplementary materials can be accessed openly via ku scholarworks4. literature cited acevedo, p., a. jiménez-valverde, j. m. lobo, and r. real. 2012. delimiting the geographical background in species distribution modelling. j. biogeogr. 39:1383–1390. anderson, m. j. 2017. permutational multivariate analysis of variance (permanova). wiley statsref stat. ref. online 1–15. anderson, r. p., and i. gonzalez. 2011. species-specific tuning increases robustness to sampling bias in models of species distributions: an implementation with maxent. ecol. model. 222:2796–2811. 2 https://github.com/marlonecobos/host-pathogen. 3 https://github.com/marlonecobos/host-pathogen. 4 https://doi.org/10.6084/m9.figshare.16870159.v1. araújo, m. b., r. p. anderson, a. m. barbosa, c. m. beale, c. f. dormann, r. early, r. a. garcia, a. guisan, l. maiorano, b. naimi, r. b. o’hara, n. e. zimmermann, and c. rahbek. 2019. standards for distribution models in biodiversity assessments. sci. adv. 5:eaat4858. barve, n., v. barve, a. jiménez-valverde, a. lira-noriega, s. p. maher, a. t. peterson, j. soberón, and f. villalobos. 2011. the crucial role of the accessible area in ecological niche modeling and species distribution modeling. ecol. model. 222:1810–1819. collinge, s. k., and c. ray. 2006. disease ecology: community structure and pathogen dynamics. oxford university press, new york. eisen, r. j., and c. d. paddock. 2021. tick and tickborne pathogen surveillance as a public health tool in the united states. j. med. entomol. 58:1490–1502. enriquez-urzelai, u., m. r. kearney, a. g. nicieza, and r. tingley. 2019. integrating mechanistic and correlative niche models to unravel range-limiting processes in a temperate amphibian. glob. change biol. 25:2633–2647. etherington, t. r. 2019. mahalanobis distances and ecological niche modelling: correcting a chi-squared probability error. peerj 7:e6678. franklin, j. 2010. mapping species distributions: spatial inference and prediction. cambridge university press. hijmans, r. j. 2019. raster: geographic data analysis and modeling. r package.5 hijmans, r. j., s. e. cameron, j. l. parra, p. g. jones, and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. int. j. climatol. 25:1965–1978. ingenloff, k., and a. t. peterson. 2021. incorporating time into the traditional correlational distributional modelling framework: a proof-of-concept using the wood thrush hylocichla mustelina. methods ecol. evol. 12:311–321. machado-stredel, f., m. e. cobos, and a. t. peterson. 2021. a simulation-based method for identifying accessible areas as calibration areas for ecological niche models and species distribution models. front. biogeogr. 13:e48814. manley, p. n., m. d. schlesinger, j. k. roth, and b. van horne. 2005. a field-based evaluation of a presence-absence protocol for monitoring ecoregional-scale biodiversity. j. wildl. manag. 69:950–966. meek, p. d., g.-a. ballard, k. vernes, p. j. s. fleming, p. d. meek, g.-a. ballard, k. vernes, and p. j. s. fleming. 2015. the history of wildlife camera trapping as a survey tool in australia. aust. mammal. 37:1–12. m’ikanatha, n. m., r. lynfield, c. a. v. beneden, and h. de valk. 2008. infectious disease surveillance. john wiley & sons. nuñez-penichet, c., m. e. cobos, and j. soberon. 2021. nonoverlapping climatic niches and biogeographic barriers explain disjunct distributions of continental urania moths. front. biogeogr. 13:e52142. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo. 2011. 5 https://cran.r-project.org/package=raster. https://github.com/marlonecobos/host-pathogen https://github.com/marlonecobos/host-pathogen https://doi.org/10.6084/m9.figshare.16870159.v1 https://cran.r-project.org/package=raster marlon e. cobos et al. – detecting signals of species’ ecological niches 58 ecological niches and geographic distributions. princeton university press, princeton. phillips, s. j., m. dudík, j. elith, c. h. graham, a. lehmann, j. leathwick, and s. ferrier. 2009. sample selection bias and presence-only distribution models: implications for background and pseudo-absence data. ecol. appl. 19:181–197. r core team. 2021. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. sauer, j. r., w. a. link, j. e. fallon, k. l. pardieck, and d. j. ziolkowski jr. 2013. the north american breeding bird survey 1966–2011: summary analysis and species accounts. north am. fauna 79:1–32. sciarretta, a., and p. trematerra. 2014. geostatistical tools for the study of insect spatial distribution: practical implications in the integrated management of orchard and vineyard pests. plant prot. sci. 50:97–110. soberón, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodivers. inform. 2:1–10. warren, d. l., a. n. wright, s. n. seifert, and h. b. shaffer. 2014. incorporating model complexity and spatial sampling bias into ecological niche models of climate change risks faced by 90 california vertebrate species of concern. divers. distrib. 20:334–343. asasepeterson2016 biodiversity informatics, 11, 2016, 1-11 completeness of digital accessible knowledge of the plants of ghana alex asase1 and a. townsend peterson2 1department of botany, university of ghana, p. o. box lg 55 legon, ghana 2biodiversity institute, university of kansas, lawrence, ks 66045, usa abstract.—providing comprehensive, informative, primary, research-grade biodiversity information represents an important focus of biodiversity informatics initiatives. recent efforts within ghana have digitized >90% of primary biodiversity data records associated with specimen sheets in ghanaian herbaria; additional herbarium data are available from other institutions via biodiversity informatics initiatives such as the global biodiversity information facility. however, data on the plants of ghana have not as yet been integrated and assessed to establish how complete site inventories are, so that appropriate levels of confidence can be applied. in this study, we assessed inventory completeness and identified gaps in current digital accessible knowledge (dak) of the plants of ghana, to prioritize areas for future surveys and inventories. we evaluated the completeness of inventories at ½° spatial resolution using statistics that summarize inventory completeness, and characterized gaps in coverage in terms of geographic distance and climatic difference from well-documented sites across the country. the southwestern and southeastern parts of the country held many well-known grid cells; the largest spatial gaps were found in central and northern parts of the country. climatic difference showed contrasting patterns, with a dramatic gap in coverage in central-northern ghana. this study provides a detailed case study of how to prioritize for new botanical surveys and inventories based on existing dak. key words.—biodiversity informatics, primary data, inventory completeness, data gaps, ghana, flora, botanical surveys biodiversity informatics may be defined as the application of information technologies to the management, algorithmic exploration, analysis, and interpretation of primary data regarding life, with a particular focus at the species level of organization (soberón and peterson, 2004). it is a rather new field, with the earliest citation of the term only 17 years ago (schalk, 1998). the most important biodiversity information centers on primary data, such as records of occurrences of species (particularly when vouchered by specimens), although many secondary sources (e.g., atlases, species accounts, distribution maps) exist as well (costello et al., 2013). such data have accumulated over centuries, but only relatively recently have they been converted into digital formats (guralnick et al., 2007) and shared openly via data portals (graham et al., 2004). major activities in biodiversity informatics are currently data-centered, focused in three areas: (i) data extraction and capture, (2) data compilation and serving, and (3) data display and visualization (peterson et al., 2010). the past two decades have seen advances and improvements in information technology (e.g. large-capacity electronic storage media, internet and data portals, distributed database technology); development of efficient data digitization workflows; changes in policies of owners of primary biodiversity data (e.g., see large-scale initiatives toward digitization of specimen data in naturalis biodiversity centre, leiden, and muséum national d'histoire naturelle (mnhn), paris), as well as establishment of global and regional biodiversity information initiatives (e.g., global biodiversity information facility, gbif; atlas of living australia, ala). these initiatives have contributed to massive accumulation and serving of primary biodiversity data records via the internet: e.g., gbif currently serves >648m primary data records on its data portal (accessed 13 may 2016). however, this forward progress could be threatened if the data do not prove sufficiently useful to biodiversity researchers, managers, and decision makers (peterson et al., 2010). primary biodiversity data have myriad applications, providing an information base that is crucial to addressing challenges of sustainable development and decision-making about natural resources and environments (chapman, 2005; sousa-baena et al., 2013). digital accessible knowledge (dak) regarding biodiversity comprises primary data records that are in digital format, townpeterson typewritten text 1 biodiversity informatics, 11, 2016, 1-11 figure 1. graphs showing accumulation of records of ghanaian plants through time (a) years, and (b) during the year. townpeterson typewritten text 2 biodiversity informatics, 11, 2016, 1-11 accessible globally without cost, and integrated with the broader university of such data (sousa-baena et al., 2013). some exciting examples of uses of dak exist, including for prioritizing areas for conservation, assessing geographic potential for species invasions, and understanding ecological and evolutionary processes (e.g., mora et al., 2008; nakamura & soberón, 2008). in ghana, significant efforts have been invested in digitization of and providing access to primary biodiversity data on the plants of the country. although the national biodiversity strategy for ghana estimated 3227 plant species (2974 indigenous and 252 introduced) in ghana (ministry of environment and science, 2002), no consensus exists on how many plant species occur in the country. still, dak on the plants (including fungi and algae that are traditionally studied in the field of botany) of ghana is relatively large, based on >90% of primary biodiversity data records derived from specimens in ghanaian herbaria, plus data from other institutions served through biodiversity informatics initiatives such as gbif. data on plants of ghana have not been integrated and assessed to establish how complete are site inventories across the country, so that appropriate levels of confidence can be applied; these gaps in knowledge affect directly the fitness-for-use of the data (otegui et al., 2013). as a consequence, this study undertook detailed assessment of dak on plants of ghana to identify and highlight gaps in knowledge. methods data were obtained from two major sources: (1) a brahms database on plants of ghana that includes data associated with plant specimens in the collections of the university of ghana, resource support management centre (rsmc) of ghana forestry commission, aburi botanic gardens, and centre for scientific research into plant medicine (csrpm), ghana 1 and (2) records of plants collected from ghana downloaded from the gbif data portal (accessed 11 january 2015). the database on plants of ghana included a total of 53,509 records captured from ghanaian herbarium specimen sheets, including 10,765 records of ghanaian plants from university of wageningen (wag). the gbif data contributed a further 9673 records after cleaning (see below). 1 http://herbaria.plants.ox.ac.uk/bol/. data were cleaned via an iterative series of inspections and visualizations designed to detect and document inconsistencies. first, we created lists of unique names in each dataset in microsoft excel, and inspected them for repeated versions of the same taxonomic concepts: misspellings, name variants, different versions of authority information, etc. such repeated name variants were flagged, checked via independent sources, and corrected to produce single scientific names that correctly referred to single taxa. second, we checked for geographic coordinates that fell outside of the country, but that were referred to it. next, within the country, we checked for consistency between textual description of regions (equivalent to provinces or states in other countries) and geographic coordinates. in each case, where possible, we corrected the data record; where no clear correction was possible, we discarded data, recording data losses at each step in the cleaning process. lastly, we discarded data records for which information on year, month, or day of collection was lacking; we created a unique ‘stamp’ of time as year_month_day. we aggregated point-based occurrence data to ½° spatial resolution grid across the country. this choice of spatial resolution was the product of an analysis balancing benefits of aggregating data (e.g., larger sample sizes) versus disadvantages (e.g., loss of spatial resolution across larger areas). the procedure consists of examining the relative change in area-adjusted variance of the data with increasing grid-cell size, and selecting the finest resolution at which the trend of the slope of the overall variance versus area curve changed most (ariño et al., in prep.); it is similar to the concept of selecting the largest sample size beyond which no significant increase in diversity is expected (ariño et al. 2008). in this study, a ½° spatial resolution offered the best balance between spatial resolution and inventory completeness, and was consistent with the spatial resolution used in a previous analysis of brazilian plants (sousa-baena et al. 2013) and wild palms of benin (idohou et al. 2015). we produced aggregation grid shapefiles in the vector grid module of qgis, version 2.4, added the coarseresolution grid identification codes to each occurrence datum, and aggregated occurrence data into coarse-resolution aggregation squares. in excel, we explored relations between species identity, time, and aggregation grid-square. we calculated (1) the total number of records available from each grid townpeterson typewritten text 3 biodiversity informatics, 11, 2016, 1-11 figure 2. chronhorogram showing temporal course of accumulation of records of ghanaian plants. radial dimension shows year of collection, and angle indicates day of the year. color of dots: black = no records; a ramp of colors indicates numbers of records, from blue = few to red = many. figure 3. summary of distances to nearest road for the available dak for ghanaian plants (gray bars) and for 5000 random points across the country. the concentration of the dak along roads (i.e., within 1.1 km of roads) compared to random patterns across the country is clear. townpeterson typewritten text 4 biodiversity informatics, 11, 2016, 1-11 square (termed n); (2) the total number of distinct species recorded from each grid square (sobs); and (3) the number of species detected on one date only (a), and (4) the number of species detected on two dates only (b). via equations provided by chao (1987), we calculated the expected number of species (sexp), as 𝑆!"# = 𝑆!"# + !! !! , and inventory completeness (c) as c = sobs / sexp. we then explored plots of c versus n (hortal et al. 2007) to assess appropriate and adequate definitions of relatively completely versus incompletely inventoried grid squares—we used criteria of c ≥ 0.4 and n ≥ 1000 as final definitions of wellinventoried areas. once we had established criteria for which grid squares could be considered as well-sampled, in qgis, we linked the table with the grid square statistics (i.e., n, sobs, sexp, c) to the aggregation grid, and saved this file as a shapefile. applying the criteria for ‘well-sampled,’ we created a shapefile of well-sampled grid squares, which we in turn converted to binary-valued (0 = not well-sampled, 1 = well-sampled) raster (geotiff) format using custom scripts in r (r core team 2014). this raster coverage was the basis for our identification of gaps, as follows. we used the proximity (raster distance) function in qgis to summarize geographic distance across the country to the nearest well-sampled area. to create a parallel view of environmental difference from well-sampled areas, we plotted 5000 random points across the country, and used the point sampling tool in qgis to link each point to the geographic distance raster, and to raster coverages (2.5’ spatial resolution) summarizing annual mean temperature and annual precipitation drawn from the worldclim climate data archive (hijmans et al. 2005). we exported the attributes table associated with the random points, and analyzed further in microsoft excel. we first standardized the values of each environmental variable to the overall range of the variable as (xi – xmin) / (xmax xmin), where xi is the particular observed value in question, thus rescaling the two variables on the same magnitude of overall variation. we then created a matrix of euclidean distances in the two-dimensional climate space, relating all of the points with a geographic distance >0 (see above) to all of the points with geographic distance of zero. the latter represent points falling in well-sampled regions, whereas the former are scattered across the entire region; the points in wellsampled regions were assigned (by definition) environmental distances of zero. finally, the environmental distances were imported into qgis, and linked back to the random points shapefile. results the dak of plants of ghana consisted of a total of 38,400 cleaned and geo-referenced data records covering the period 1830-2012. number of records per year ranged between 1 and 241, with an average of 46.2 data records per year; ~66% of the records were from the period 1950–1977 (figure 1). seasonal patterns in the records showed that most records were from the dry season (october to december), whereas the fewest records were from the rainy season (june to august; figure 1). the chronological and seasonal pattern in data records can be visualized via the chronhorogram in figure 2: we observed a rapid increase in numbers of data records between 1940 and 1980, and decreasing numbers of records thereafter. the data showed an overwhelming tendency towards concentration of records in southern ghana. points of access (roads and rivers) were clearly visible in the spatial distribution of records, and indeed the dak was significantly concentrated close to roads compared to random points (figure 3). total number of grid cells across ghana for ½° resolution was 92; all held data and about 13% were classified as well-known sites. plots of c against number of records per grid cell showed sample-size dependency in the range of 500-1000 records (figure 4); hence, we defined well-sampled sites as those having ≥1000 records and c ≥ 0.4, because stricter criteria would identify massive swaths of territory as not-well-sampled (see sousa-baena 2013). generally, well-known grid cells were concentrated in the southwestern and southeastern areas of the country, and the largest gaps were in the westcentral and northeastern parts of the country (figures 5 and 6). climatic conditions are diverse in ghana, with more homogenous climates in the south compared to those of the central and northern parts of the country; distinct climatic conditions exist in the northern and northwestern areas of the country. townpeterson typewritten text 5 biodiversity informatics, 11, 2016, 1-11 figure 4. plot of inventory completeness (c) against sample size (n) for grid cells across ghana. red lines indicate criteria for “well-known” with respect to each of the axes. townpeterson typewritten text townpeterson typewritten text 6 biodiversity informatics, 11, 2016, 1-11 climatic differences from well-known cells were most pronounced in a broad swath of the northern part of the country (figure 6). combining these two views identified sites that are both geographically remote and environmentally different from well-known sites (figure 6). four areas fit these criteria: (1) northeastern ghana, including the entire upper west region; (2) the north-central part of the northern region of the country around tamale; (3) the west-central part of the country, including parts of the brong-ahafo region around bui national park; and (4) the eastcentral part of ghana, including parts of the northern volta region and the brong-ahafo region, including digya national park. discussion access to primary biodiversity data is critical to addressing challenges of sustainable development and decision-making (sousa-baena et al., 2013). most primary biodiversity data, in excess of 6.5 x 108 data records, that have been shared openly are based on biological collections held in herbaria and museums, as well as observational data from citizen scientists. our focus on dak emphasizes data that are available to the broader scientific community for analysis and exploration (sousa-baena et al. 2013), such that biological collections for which associated data have not been digitized or that are digital but remain broadly unavailable are ignored (costello et al., 2013). in contrast, information that is open and accessible has potential to impact science and conservation, as well as the care and curation of specimens (sousa-baena et al. 2013), such that digitization and sharing of primary biodiversity data is much to be encouraged (see article 17, convention on biological diversity). knowledge of inventory completeness is important to determine appropriate levels of confidence that can be applied to data-derived patterns of biodiversity across a region or a taxon (soberón and llorente, 1993; colwell and coddington, 1994; gotelli and colwell, 2001). recent analyses have attempted to summarize the state of knowledge of plant diversity (kier et al., 2005; mutke and barthlott, 2005), but few of these studies are based on primary biodiversity data (soberón et al., 2000; ariño et al., 2012; sousabaena et al. 2013). such studies have generally indicated high species richness at small numbers of well-sampled areas, and few sites that are wellknown and comprehensively documented, but broad areas that remain poorly sampled. particularly perplexing is when high species richness sites correspond closely to sites of high sampling intensity, as such situations suggest that “hotspots” in fact represent artifacts of incomplete sampling (tobler et al., 2007; ahrends et al., 2011). in ghana, few studies have evaluated the state of knowledge of distributions of plants; the few studies existing focused in the forest vegetation zone in the southern parts of the country, based on both herbarium records and observations from sampled plots, and characterized species’ distribution patterns (hall and swaine, 1981; hawthorne and abu-juam, 1995). although the savanna vegetation zone covers about two-thirds of ghana, no assessment has addressed knowledge of distributions of plants there. this study is also the first to be based on digital records that are openly accessible to the broader scientific community. although dak for ghanaian plants should improve with time, both in quantity and quality, numbers of new collections have been decreasing steadily over recent years. dak completeness focuses on the consistency of information that is available, and offers a useful index about why a site has few or many species recorded (soberón et al., 2000; sousa-baena et al. 2013). here, we addressed questions about sites in ghana where biodiversity knowledge is relatively reliable versus where information is incomplete. we found well-known sites principally in southeastern ghana relatively close to the location of the ghana herbarium, where botanists and students have developed intensive collections. southwestern ghana is considered richest in terms of plant species in ghana (hawthorne and abu-juam, 1995); as a consequence, many botanists and indeed many large-scale projects have collected plants from the area. in northwestern ghana, extensive plant collections have been undertaken around mole national park, such that that area is botanically well known. in this study, we identified knowledge gaps, and characterized them in terms of geographic distance and environmental difference from wellknown sites, and see these sites as priority areas for botanical sampling. these gaps frequently result from no previous collecting visits to sites, but may in some cases reflect lack of digital access to collections that in truth exist (sousa-baena et al. 2013). most herbaria in ghana now have their data townpeterson typewritten text 7 biodiversity informatics, 11, 2016, 1-11 figure 5. geographic patterns of inventory completeness across ghana based on ½° grid squares. shading indicates inventory completeness (c), in a spectrum from violet (as low as 0) to red-brown (as high as 0.63). thick dashed black outlines indicate those grid squares that fit the “well-known” criterion of c ≥ 0.4 and ≥1000 records. townpeterson typewritten text 8 biodiversity informatics, 11, 2016, 1-11 figure 6. visualization of geographic distance of ½° pixels across ghana from well-known sites; climatic difference of ½° pixels across ghana from well-known sites based on nearest-neighbor distance; and a combination of the two distances based on equal minimum-maximum scaling. townpeterson typewritten text 9 biodiversity informatics, 11, 2016, 1-11 records in digital formats and openly available: ~80% of herbarium specimen sheets have been captured, although carpological and other collections remain to be digitized. in addition, many specimens of ghanaian provenance are deposited in collections in other countries and remain to be digitized and shared. as such, the fastest way to improve dak of ghanaian plants is to fix data “leaks” among existing digital data records (sousa-baena et al. 2013). in particular, of the 53,509 records analyzed herein, 60 (0.1%) records had indeterminate names, 24,710 (46.2%) records had incomplete dates (missing year, month, or day), and 538 (1.0%) records lacked geographic coordinates, such that these records could not be included in our analyses. another area of importance in terms of dak improvement are the large amounts of botanical data associated with significant collections of ghanaian plants held elsewhere in the world. data repatriation from european and north american herbaria with large collections from ghana is an important potential source of dak of ghanaian plants. collaborative efforts linking west african botanists, and north american and european herbaria are now underway2, and promise to develop and enable rich new dak resources. this study illustrates the importance of assessing completeness of dak for prioritizing botanical surveys based on existing knowledge and benefits to be reaped from biodiversity data sharing and integration. field sampling efforts should focus in areas identified herein as both environmentally different and geographically distant from wellknown sites. such information exists for only a few countries, such as brazil (sousa-baena et al. 2013, idohou et al. 2015, koffi et al. 2015). this kind of information is important for strategic national policy (soberón and peterson, 2009), and is an important step towards meeting the aichi targets3 and national commitments to the clearing house mechanism of the convention on biological diversity4. acknowledgments we are most grateful to jrs biodiversity foundation for funding the first author in development of the database of the plants of ghana. the authors are thankful to several student data 2 http://jrsbiodiversity.org/grant/university-of-ghana-herbaria/. 3 https://www.cbd.int/sp/targets. 4 https://www.cbd.int/chm. capturists, particularly cindy asare, gloria oppong, tony asafo-agyei, amos asase, and joana adofo. preparation of this paper began at a workshop on national biodiversity diagnoses at entebbe in uganda supported by a jrs biodiversity foundation grant to the second author. we are grateful to arturo ariño, for developing the chronohorogram. we thank arturo ariño, lindsay campbell, and kate ingenloff for guidance and support in developing these analyses. literature cited ahrends, a., n.d. burgess, r.e. gereau, r. marchant, m.t. bulling, j.c. lovett, p.j. platts, v.w. kindemba, n. owen, e. fanning, and c. rahbek. 2011. funding begets biodiversity. div. and distrib. 17:191-200. ariño, a.h., c. belascoáin, and r. jordana. 2008. optimal sampling for complexity in soil ecosystems. pp. 220-230 in minai a., and y. bar-yam, unifying themes in complex systems iv. springer-verlag, berlin. ariño, a.h., j. otegui, a. villarroya, and a. pérez de zabalza. 2012. primary biodiversity data records in the pyrenees. env. eng. & man. j. 11:1059-1075. chao, a.1987. estimating the population size for capturerecapture data with unequal catchability. biometrics 43:783-791. chapman, a.d. 2005. uses of primary speciesoccurrence data, version 1.0. global biodiversity information facility, copenhagen. colwell, r.k., and j. a. coddington. 1994. estimating terrestrial biodiversity through extrapolation. phil. trans. roy. soc. b 335:101-118. costello, m.j., w.k. michener, m. gahegan, z.q. zhang, and p.e. bourne. 2013. biodiversity data should be published, cited, and peer reviewed. trends ecol. evol. 28:454-461. gotelli, n.j., and r. k. colwell. 2001. quantifying biodiversity: procedures and pitfalls in the measurement and comparison of species richness. ecol. lett. 4:379-391. graham, c.h., s. ferrier, f. huettman, c. moritz, and a. t. peterson. 2004. new developments in museumbased informatics and applications in biodiversity analysis. trends ecol. evol. 19:497-503. guralnick, r.p., a.w. hill, and m. lane. 2007. towards a collaborative, global infrastructure for biodiversity assessment. ecol. lett. 10:663-672. hall, j.b., and m. swaine. 1981. distribution and ecology of vascular plants in tropical rainforest: forest vegetation in ghana. w. junk, the hague. hawthorne, w. d., and m. abu-juam. 1995. forest protection in ghana. international union for the conservation of nature, gland. townpeterson typewritten text 10 biodiversity informatics, 11, 2016, 1-11 hijmans, r.j., s.e. cameron, j.l. parra, p.g. jones, and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. int. j. climatol. 25:1965-1978. hortal, j., j.m. lobo, and a. jimenez-valverde. 2007. limitations of biodiversity databases: case study on seed-plant diversity in tenerife, canary islands. cons. biol. 21:853-863. idohou r., a. h. ariño, a. assogbadjo, r. glele kakai, and b. sinsin. 2015. knowledge of diversity of wild palms (arecaceae) in the republic of benin: finding gaps in the national inventory by combining field and digital accessible knowledge. biodiv. inf. 10:45-55. kier, g., j. mutke, e. dinerstein, t.h. ricketts, w. kuper, h. kreft, and w. barthlott. 2005. global patterns of plant diversity and floristic knowledge. j. biogeogr. 32:1107-1116. koffi, k.j., a.f. kouassi, c.y. adou yao, a. bakayoko, i.j. ipou, and j. bogaert. the present state of botanical investigations in côte d’ivoire. biodiv. inf. 10:56-64. ministry of environment of science. 2002. national biodiversity strategy of ghana. ministry of environment and science, accra. mora,c., d.p. tittensor, and r. a. myers. 2008. the completeness of taxonomic inventories for describing the global diversity and distribution of marine fishes. proc. roy. soc. b 275:149-155. mutke, j., and w. barthlott. 2005. patterns of vascular plant diversity at continental to global scales. biologiske skrifter 55:521-537. nakamura, m., and j. soberón. 2008. use of approximate inference in an index of completeness of biological inventories. cons. biol. 23:469-474. otegui, j., a.h. ariño, m.a. encinas, and f. pando. 2013. assessing the primary data hosted by the spanish node of the global biodiversity informatics facility (gbif). plos one 8:e55144. pardo, i., m.p. pata, d. gómez, and m. b. garcía. 2013. a novel method to handle the effects of uneven sampling effort in biodiversity databases. plos one 8:e52786. peterson, a.t., a.g. navarro-sigüenza, and h. benítezdíaz. 1998. the need for continued scientific collecting: a geographic analysis of mexican bird specimens. ibis 140:288-294. peterson, a.t., s. knapp, r. guralnick, j. soberón, and m. t. holder. 2010. the big questions for biodiversity informatics. syst. biodiv. 8:159-168. r core team. 2014. r: a language and environment for statistical computing; http://www.r-project.org/. r foundation for statistical computing, vienna. schalk, p.h. 1998. management of marine natural resources through biodiversity informatics. marine pol. 22:269-280. soberόn, j., and a.t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. phil. trans. roy. soc. b 359: 689698. soberόn, j. and a. t. peterson. 2009. monitoring biodiversity loss with primary species-occurrence data: toward national-level indicators for the 2010 target of convention on biological diversity. ambio 38:29-34. soberόn, j., and j.e. llorente. 1993. the use of species accumulation functions for the prediction of species richness. cons. biol. 7:480-488. soberόn, j.m., j. b. llorente and l. oñate. 2000. the use of specimen-label databases for conservation purposes: an example using mexican papilionid and pierid butterflies. biodiv. & conserv. 9:1441-1466. sousa-baena, m.s., l.c. garcia, and a.t. peterson. 2013. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. div. distrib. 20:369-381. sousa-baena, m.s., l. c. garcia, a. t. peterson. 2014. knowledge behind conservation status decisions: data basis for ‘data deficient’ brazilian plant species. biol. conserv. 173:80-89. tobler, m.w., e. honorio, j. janovec, and c. reynel. 2007. implications of collection patterns of botanical specimens on their usefulness for conservation planning: an example of two neotropical plant families (moraceae and myristicaceae) in peru. biodiv. & conserv. 16: 659-677. townpeterson typewritten text 11 biodiversity informatics, 17, 2022, pp. 10-26 10 biodiversity and distribution of isopoda and polychaeta along the northwestern pacific ocean and the arctic ocean hanieh saeedi1*, angelika brandt1,2, nils l. jacobsen2 1senckenberg research institute and natural history museum; department of marine zoology, 60325 frankfurt am main, germany 2goethe university frankfurt, institute for ecology, diversity and evolution, 60438 frankfurt am main, germany *corresponding author: hanieh saeedi, hanieh.saeedi@senckenberg.de abstract. the northwestern pacific ocean is one of the hotspots of species richness and one of the high en‑ demicity areas of the world ocean. however, large-scale biodiversity patterns of major deep -sea taxa such as isopoda and polychaeta are still poorly studied. the goal of this research is to study the distribution, biodiver‑ sity, and community composition of isopoda and polychaeta (including siboglinidae and echiura) across the northwestern pacific ocean and the adjacent arctic ocean. the study area was divided into equal-sized hexag‑ onal cells (c. 700,000 km²), ecoregions, 5° latitudinal bands, and 200 m depth intervals as unit of analysis. our results revealed that the area around the philippines and the laptev sea had the highest isopod and polychaete’s species richness compared to the other geographic regions of our study, with a latitudinal decline of species richness in shallow waters in both taxa. in the deep sea, maximum species richness increased towards the tem‑ perate latitudes. gamma species richness (number of species per 200 m depth interval) also declined with depth. rarefied species richness of isopods peaked around 5000 m depth. rarefaction curves demonstrated a great po‑ tential for undiscovered richness across 5° latitudinal bands and depth intervals. in shallow waters, polychaetes with a pelagic larval phase had a wider distribution range compared to brooding isopods, but, in the deep sea, isopods had slightly wider distribution ranges compared to polychaetes. these results thus demonstrated that shallow water taxa with pelagic larvae and polychaete species with a wide vertical distribution range could potentially invade higher latitudes, such as species from the northwest pacific invading the arctic ocean under the rapid climate change and catastrophic reduction of sea ice cover. these changes might dramatically change the benthic communities of the arctic ocean and management of such should take an adaptive approach and apply measures that take potential extension and invasion of species into account. key words: biodiversity, northwest pacific ocean, arctic ocean, isopoda, polychaeta the northwestern pacific ocean (nwp) extends from the equator in the south to the arctic ocean (ao) in the north, and from asia and australia in the west to 180° longitude in the east. the nwp is characterized by many heterogeneous habitats and numerous oceanic islands and deep-sea trenches (saeedi and brandt 2020a; saeedi et al. 2020). the tropical and subtropical areas of the west pacific in‑ clude the indo–australian archipelago and host the highest number of marine species in the world ocean (renema et al. 2008; saeedi et al. 2019b) the bordering ao has an average depth of 987 m (ostenso 1962). it is strongly influenced by the nu‑ trient-rich water that is flowing through the bering strait (woodgate and aagaard 2005). estimates of annual primary production in the northern bering and the southern chukchi seas are very high, with an average of 470 g c m‑2 y‑1 (springer and mcroy 1993; sakshaug 2004; grebmeier et al. 2006). the bering and chukchi sea shelves are among the larg‑ est shelf areas worldwide hosting a rich benthic food web, which in turn supports benthic-feeding preda‑ tors (grebmeier et al. 2006). isopoda are a crustacean order in the superfam‑ ily peracarida with >10,000 species (westheide and rieger 1996). they are among the most common taxonomic groups in the world ocean and are major contributors to the deep-sea diversity (golovan et al. 2013). the suborder asellota is predominant below 200 m depth. in some cases, isopod samples from mailto:hanieh.saeedi@senckenberg.de hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 11 deep-sea basins mostly consist of asellota specimens (brandt et al. 2004; kaiser et al. 2009). one synapo‑ morphic structure for all peracarida including iso‑ pods is the marsupium, a brood chamber in which the females carry their offspring (westheide and rieg‑ er 1996; saeedi and brandt 2020a). after hatching, the isopods are in an early juvenile stage (manca) when they stay inside the marsupium (westheide and rieger 1996; saeedi and brandt 2020a). usually, the isopods leave the brood chamber during the second manca stage (wägele 1989). thus, isopods are effec‑ tively missing a pelagic larval phase which might re‑ sult in lower dispersal capabilities compared to taxa with a pelagic larval phase (siegel et al. 2003). polychaeta is a class in the phylum annelida. in this study, we refer to “polychaeta” because the term has been widely used in the available literature. to date, >12,000 polychaete species have been re‑ corded (read and fauchald 2020). polychaetes are a very abundant taxon and occur in almost any marine benthic habitat feeding on deposited organic matter and microfauna enriching the food web (hessler and jumars 1974; fauchald 1983; saeedi and brandt 2020a). polychaetes have a great diversity of de‑ velopmental modes, often fertilized externally by releasing both eggs and sperms into the water col‑ umn; thus, most polychaeta larvae are planktotrophic (fauchald 1983; wilson 1991). pelagic larvae can be divided into planktotrophic larvae and lecithotrophic larvae, the former being the most common strategy for benthic invertebrates which implies feeding on phytoplankton and/or zoo‑ plankton (wilson 1991; levin 2006). planktonic larval durations (pld) range from just a few hours to multiple months (carson and hentschel 2006; shanks 2009). a positive correlation between larval duration and dispersal distance has been recorded (siegel et al. 2003; shanks 2009). larval behavior is important to consider regarding dispersal capabili‑ ties (krug and zimmer 2004). species with a pelagic larval phase have the advantage that they can reach areas of favorable conditions more easily than brood‑ ers (pechenik 1999), which might reduce fluctuations in adult populations (eckert 2003). brood protection, on the contrary, can offer security in terms of shel‑ ter and food availability (levin and bridges 1995). these attributes might be especially favorable in the food-limited deep-sea ecosystems (bush et al. 2012). however, the global and regional distribution and diversity patterns of isopods and polychaetes should be studied further to enable us to better compare the dispersal abilities of these two taxa. it is generally assumed that species richness for all marine species including isopods and poly‑ chaetes peaks in the tropics and decreases towards higher latitudes (rex et al. 2000; brown 2014; val‑ entine and jablonski 2015). however, recent studies have shown that latitudinal global species richness pattern is bimodal, not only decreasing with latitude and depth, but also decreasing at the equator (saee‑ di and costello 2012; costello and chaudhary 2017; saeedi et al. 2017b; saeedi and costello 2019). in the deep sea, maximum species richness is recorded at higher latitudes (30–50°n; (mcclain et al. 2012; woolley et al. 2016; saeedi et al. 2017b)). species richness in shallow waters correlates positively with temperature (tittensor et al. 2010; chaudhary et al. 2017; costello and chaudhary 2017), whereas spe‑ cies richness in the deep sea is more likely influenced by chemical energy and food availability (brandt et al. 2015; brandt and malyutina 2015; woolley et al. 2016; yasuhara and danovaro 2016; barroso et al. 2018; golovan 2018; saeedi et al. 2019a). consider‑ ing global climate change and its impacts on distri‑ butions of marine species, further investigating the main drivers of species distributions and richness should be prioritized. the nwp and ao are experiencing rapid climate change; consequently, local species might experience a northward and southward distribution range shift (simões et al. 2021). the latitudinal and bathymet‑ ric species-richness gradients, as well as their causes, need to be (re)examined in these areas. an under‑ standing of present distributional patterns and their potential drivers is imperative to predicting future changes and their consequences. an informational basis is necessary for any conservational work or the implementation of marine protected areas (mpa). saeedi et al. (2019b) already investigated potential drivers of marine species richness in the nwp and the ao, but did not include the important role of dis‑ persal abilities of different taxa as a driver of species richness patterns in those areas (saeedi et al. 2019b; saeedi et al. in press). in addition, studies of the re‑ lationship between developmental mode and distri‑ bution of marine species are rare (mileikovsky 1971; krug and zimmer 2004). the goal of this study is thus to (1) analyze the latitudinal and bathymetric distribution of polychaeta and isopoda and identify hotspots of species richness, (2) compare species hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 12 distributions of isopoda and polychaeta, and (3) dis‑ cover if polychaeta with a pelagic larval phase have wider distributions compared to polychaeta with no pelagic larval phase and brooding isopoda. methods study area and data preparation the study area included large portions of the nwp and the adjacent ao at latitudes of 0–90°n and longitudes of 100–180°e. most records used in this study were extracted from the ocean biodiversity in‑ formation system (obis) and the global biodiver‑ sity information facility (gbif). we also drew data from four russian -german and german-russian benthic deep-sea expeditions in the nwp (malyuti‑ na and brandt 2013; brandt and malyutina 2015; malyutina and brandt 2018; saeedi et al. 2019e; saeedi and brandt 2020b). records were collected by a great variety of sampling methods (saeedi et al. 2019e; saeedi and brandt 2020b). the gear deployed during four benthic deep-sea expeditions included ctd, muc (multicorer); gkg (giant box corer); ebs (epibenthic sledge); agt (agassiz trawl), bt (bottom trawl) (brandt and malyutina 2014), bc (box corer) (brandt et al. 2010), and pn (plankton net) (chernyshev and polyakova 2018). standard‑ ized methodology was used for their deployment during all expeditions for reasons of comparison (brandt and malyutina 2014). data were merged and cleaned following saee‑ di et al. (2019a). records were checked for reliabil‑ ity with the r package “robis” (provoost and bosch 2020); doubtful coordinates were either corrected (e.g., when longitude and latitude were switched), or were removed (e.g., in the case of fossil records; (saeedi et al. 2019b; saeedi et al. 2019d)). all re‑ cords were taxon-matched against the world reg‑ ister of marine species (worms). old taxonomy was updated, and unconfirmed matches were crosschecked and incorporated into the dataset. following the worms taxonomy backbone for polychaetes, siboglinidae and echiura were included, but sipun‑ culas were not. after these steps, the dataset consisted of 6716 records, of which 5328 records were annelida and 1388 records were arthropoda. table s1 summarizes percentages of the records and available depth infor‑ mation for shallow water and deep sea in the nwp and ao (table s1): 1066 species from 105 families were represented in the dataset. all records were clas‑ sified as either shallow-water (0–500 m) or deep-sea (>500 m) records, based on the world register of deep-sea species (wordss; (glover et al. 2021)). the reason for this classification is the small vari‑ ability in physical parameters over different seasons and the negligible effect of sunlight on organisms at depths below 500 m. information about presence of a pelagic larval phase was extracted from the literature and worms, and added to the dataset for polychaetes. in addition to actual information available, an estimation method was used (carson and hentschel 2006): when >80% of a genus or family shared the same developmen‑ tal mode, that mode was assumed for other species in that group. from a total of 736 polychaete spe‑ cies, 321 were classified as having a pelagic larval phase according to worms and other literature (fauchald 1983; wilson 1991; bhaud 1998; carson and hentschel 2006; shanks 2009; kędra et al. 2013; read and fauchald 2020) which corresponds to 2401 records, out of a total of 5328 records. data analysis for data analysis and visualization, we used r 3.6.2 (r core team, 2020). the packages readxl (wickham et al. 2019), tidyr (wickham and henry 2020), dplyr (wickham et al. 2021), and ggplot2 (wickham 2016) were employed for importing, cleaning, manipulating, and plotting the data. with the help of the sf package (pebesma 2018), we cre‑ ated a map of the study area with an overlaying hex‑ agonal grid. each hexagonal cell covered roughly 700,000 km²; the study area was also divided into ecoregions to provide better insights for conserva‑ tional planning. these ecoregions were taken from the marine ecoregions of the world (meow) poly‑ gon layer (spalding et al. 2007). to study patterns of latitudinal and bathymetric distribution, the study area was divided into 5° latitudinal bands and 200 m depth intervals. sampling effort, species richness, and rarefied species richness es15 (see below) were calculated for each hexagonal cell, ecoregion, 5° latitudinal band, and 200 m depth interval. using an occurrence table, sampling effort was examined. alpha species richness (species richness per hexagonal cell) and gamma species richness (species richness per 5° latitudinal band) were calculated with a presence/ absence matrix. the rarefaction method (es15) was employed to account for sampling bias (saeedi et al. 2019b). it repeatedly re-samples 15 randomly cho‑ sen records from all records available and calculates hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 13 the average species number per 15 records (oksanen et al. 2019). a sample size of 15 was chosen over a sample size of 50, because the number of distribution records at a site needs to be higher than the chosen sample size in order to get a rarefied result. at many stations, the number of distribution records were too low for a sample size of 50. the function “rarefy” from the “vegan” r package oksanen et al. 2019) was used for the rarefaction method. the rarefaction method (es15) was also em‑ ployed to create rarefaction curves. rarefied expect‑ ed number of species shows the number of species in relation to the number of samples, by a random se‑ lection of samples along the study area. when a rar‑ efaction curve reaches an asymptote, it implies that most or all of the species in a given area have been found. if the curve does not show a plateau, more sampling is needed (chiarucci et al. 2008; oksanen et al. 2019). all isopod and polychaete records were grouped into suborders. suborder was chosen as a unit of comparison because it enabled informative compar‑ isons. violin plots were created based on the subor‑ ders to study community composition and density distribution ranges of isopods and polychaetes in shallow waters and in the deep sea. a hierarchical cluster analysis was performed to study similarity and difference among ecoregions. ordinary bootstrap resampling (bp) and multiscale bootstrap resampling (au) were handled by the function “pvclust” of the package (suzuki and shimodaira 2019)(suzuki et al., 2019). results distribution and diversity overall, the highest sampling effort (i.e., number of records) was recorded around south korea based on both hexagons and ecoregions plots of isopods and polychaetes (fig. 1; fig. s1-s20). alpha species richness peaked around south korea, the philippines, the central kuroshio current region south of japan, and in the laptev sea (fig. 1; fig. s1-s20). rarefac‑ tion (es15) showed that highest species richness was around the philippines, sea of japan, yellow sea, and the ao (fig. 1). isopod species had a narrower range compared to polychaete species in both shallow water and deep sea (fig. 2; fig. 3; fig. s3-s6). the distribution per‑ centages for isopods were 55.7% overall, 42.6% in shallow waters, and 29.5% in the deep sea. deep-sea sampling effort for isopods was highest in the oyashio current region (fig. s4). species richness and es15 values were higher around the kuril-kam‑ chatka trench, in the oyashio current region, and in the ao, compared to other regions (fig. s4). polychaeta were present in 83.6% of all hexa‑ gons in the study area (fig. 3; fig. s5-s6). in shal‑ low waters, this number was reduced to 77.1%, and figure 1: biodiversity patterns across all taxa; (a) sampling effort (number of distribution records), (b) alpha spe‑ cies richness (number of species per hexagon), and (c) es15 (expected number of species) for both isopoda and polychaeta in the nw pacific and the arctic ocean. hexagon cell size is 700,000 km². the areas with no hexagons had zero values. hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 14 in the deep sea to 42.6%. however, in the deep sea, polychaetes with a pelagic larval phase were only present in 24.6% of all hexagons. polychaetes with a known pelagic larval phase also had wider distri‑ bution overall (73.8% of all hexagons) and in shal‑ low waters (65.6% of all hexagons) compared to isopods. in the deep sea, sampling effort and alpha species richness were highest in the laptev sea (fig. s6). es15 values were highest in the laptev sea and around the southern point of the philippines. because of the scarce amount of deep-sea records, rarefaction worked only for a small number of hexagons, so no clear pattern was observable. polychaetes with a known pelagic larval phase showed distribution and diversity patterns similar to those of polychaetes overall (fig. s7). however, figure 2: biodiversity patterns of all isopoda; (a) sampling effort (number of distribution records), (b) alpha species richness (number of species per hexagon), and (c) es15 (expected number of species) for all isopoda in the nw pacific and the arctic ocean. hexagon cell size is 700,000 km². the areas with no hexagons had zero values. figure 3: biodiversity patterns of all polychaeta; (a) sampling effort (number of distribution records), (b) alpha species richness (number of species per hexagon), and (c) es15 (expected number of species) for all polychaeta in the nw pa‑ cific and the arctic ocean. hexagon cell size is 700,000 km². the areas with no hexagons had zero values. hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 15 species richness and es15 patterns in polychaetes with an unknown pelagic larval phase showed slight differences compared to polychaetes with a known pelagic larvae and overall polychaetes (fig. s7-s10). polychaetes with unknown pelagic larvae showed higher species richness in the ao as well the sea of japan and the yellow sea (fig. s8). however, poly‑ chaetes with known and unknown pelagic larvae still had a broad distribution, and hotspots of sampling effort and species richness were identical. since less data were available from the deep sea, the rarefaction method failed in more hexagons and ecoregions than before (fig. s10). nevertheless, es15 hotspots were still present in the same locations. latitudinal gradients between the latitudes of 55° and 75° n, few or no records were available in light of less or minimum available ocean area (fig. 4; fig. s21). sampling ef‑ fort was highest at latitudes 40° and 80° n. gamma species richness showed two peaks: one at 15° and another one at 40° n. es15 values were highest at latitudes 5–10° n, and were lowest at latitudes 20° and 75° n (fig. 4; fig. s22). values for the ao were slightly lower than those for the nwp. most of the records at 40° n had missing depth information (table s1), which is why this peak disap‑ peared when categorizing records by depth. in shal‑ low waters, highest sampling effort was at 80° n. for depths above 200 m, gamma species richness was highest at 15° and 80° n. also, a decline of species richness is visible towards higher latitudes, and with a dip near the equator. in shallow waters, a dip at 20° n in rarefaction values was visible, with slightly lower than average levels at 50° n and at ao lati‑ tudes (fig. s23). in the deep sea, sampling effort was figure 4: latitudinal distribution and diversity of isopoda and polychaeta in the nw pacific and the arctic ocean; (a) sampling effort (number of distribution records) and (b) gamma species richness (number of species per 5° latitudinal band) for isopoda (in red) and polychaeta (in blue); es15 values (expected number of species) for (c) isopoda and (d) polychaeta. hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 16 highest at latitudes 45° and 80° n. deep-sea gamma species richness was highest at 45° and 80° n. high‑ est deep sea es15 values were recorded at latitudes 15°, 40°, and 80° n, whereas lowest values were at 20°, 30°, and 85° n (fig. s24). for isopods, in shallow waters, sampling effort peaked at 80° n (fig. s25). gamma species richness decreased towards higher latitudes, with a small dip near the equator and a peak at latitude 80° n. in shal‑ low waters, es15 values were higher in the tropical zone, with peaks 35–45° and 80° n, compared to temperate and cold zones. deep-sea sampling effort peaked at 45° and 80° n (fig. s26). gamma species richness was highest between 40°, 55°, and 80° n. in the rarefaction plot (fig. s26, c), while show‑ ing a peak at 40° n and a drop at latitudes of 85° n and greater, data were too few for a proper rar‑ efaction estimation at many latitudinal bands in the deep sea. the rarefaction curve for all isopods per 5° latitudinal band showed a high species count for the latitudes 15–40° n (fig. s33). a flattening of the curve was clearly visible at latitudes 45° and 80° n. for ao latitudes, including 75°, 85° and 90° n, spe‑ cies counts and sample sizes were minimal. for polychaetes, shallow-water sampling effort was highest at latitude 80° n, in addition to nwp latitudes of 15°, 25°, and 35° n (fig. s25). in shal‑ low waters, species richness decreased with higher latitudes (except at 80° n), and a dip was recorded near the equator. the peak of gamma species rich‑ ness was at 80° n latitude. rarefaction produced a similar pattern to the es15 values for all polychaete records. deep-sea records showed a significant peak of sampling effort at 80° n, and high values in the temperate and arctic zones were also observed. two peaks of gamma species richness were noticeable, at 45° n and 80° n. es15 values in the deep sea were lowest between latitudes 15° and 35° n, and at lati‑ tudes 55° and 85° n, compared to other latitudes. bathymetric gradients from all 6716 records, ~45 % were from upper shallow depths above 200 m, ~18 % in the deep sea; ~37% were missing information about depth alto‑ gether. around 71% of records were from the nwp and ~29 % from the ao. isopoda and polychaeta were both distributed over the whole depth range, from shallow waters to deep sea (fig. s27-s28). re‑ cords came from the surface all the way to a depth of 10,170 m (the deepest record was a polychaete, bathykermadeca hadalis (kirkegaard, 1956), from 10,170 m near the philippines. for isopods, the deep‑ est occurrence was at 9910 m from the same loca‑ tion, referred to macrostylis galatheae wolff, 1956. in the ao, the deepest recorded specimen was from 4170 m, of aricidea (aricidea) albatrossae petti‑ bone, 1957. most records were from upper shallow waters above 200 m. almost half of isopods and ~44% of polychaetes were found in this zone. below the depth 200 m, records decreased continuously, but peaked again around 600 m. many isopods were re‑ corded around 5000 m in the nwp. species richness values were correlated with number of records. es15 had the highest values from depth 0 m to 600 m, and then gradually declined below that depth, reaching its lowest value at 2800 m (fig. s29-s30). rarefaction curves for both taxa showed simi‑ lar patterns (fig. s31-s36). the rarefaction analysis documented highest sample size and species count at 40° n. the 40° n curve and the 80° n curve moved towards their maximum species count, while all oth‑ ers still showed potential for growth. latitude 15° n was distinguished with a high species count for its sample size (fig. s31). highest sample sizes and spe‑ cies counts were in the range 0–200 m (upper shal‑ low waters), followed by 400–600 m and 200–400 m (fig. s32). a notable difference between isopods and polychaetes (except sample size and species count) was the depth range 5200–5400 m, where sample size was roughly equal to the 200–400 m range, with a flattened shape for isopods. isopods made up ~21% of all records. only ~12 % of isopod records had missing depth infor‑ mation. shallow records were ~49% and deep-sea records were ~39% of the total. nwp records made up 67% and ao records constituted 33%. isopods were represented by 330 species. of those, 302 spe‑ cies (~92%) were from the nwp, and 27 (~8 %) from the ao; only one species was found in both areas. around 45% of isopod species were classified as shallow-water species; ~17% of isopods were pres‑ ent only in the deep sea, and ~23% of isopods were eurybathic organism, found in both shallow waters and the deep sea. isopods are grouped into five suborders (fig. 5); of 37 families, 34 could be assigned to a suborder. asellota, cymothoida, and valvifera were dominant in the ao community, whereas cymothoida was most prominent in the nwp, followed by sphaero‑ matidea and limnoriidea. analyzing the bathymet‑ rical distribution revealed that asellota was by far the most common representative in shallow waters hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 17 and in the deep sea. all other suborders were present in shallow waters, with few records from the deep sea. asellota also had the widest distribution range (depthwise), with roughly twice the range of cy‑ mothoida, which fell in second place. for polychaetes, ~80% of records were of poly‑ chaetes. of those, over two thirds (71.8%) were reported from the nwp, and the rest (28.2%) from the ao. polychaete records had more missing depth information (43.8%) than isopod records. around 43% of polychaete records were from shallow water, compared to nearly 13% from deep sea. the dataset included 736 polychaete species. of those species, 624 (~85 %) were found solely in the nwp; only 70 species (~10%) were from the ao. in total, 42 species (6%) were present in both the nwp and the ao. half of the records were from shallow waters, and only 15% were from the deep sea; the rest had missing depth information. of total records, 154 spe‑ cies (21%) of polychaeta occurred in both shallow water and deep sea. the violin plot for taxonomic composition of polychaete species by depth showed that bonelliida was not the most common taxon in shallow waters (fig. 6), but was the dominant taxon in the deep sea. while not having many records in the deep sea, aphroditiformia had the widest distribu‑ tion, followed by bonelliida, compared to other taxa. similarity cluster the cluster analysis of ecoregions revealed six distinct groups (fig. 7); all had an au value of >95. the first cluster included the eastern philippines and figure 5: violin plot representing the (a) latitudinal and (b) bathymetric distribution density range of isopoda records for each suborder in the nw pacific and the arctic ocean. hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 18 figure 6: violin plot representing the (a) latitudinal and (b) bathymetric distribution density range of polychaeta records for each suborder in the nw pacific and the arctic ocean. hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 19 palawan north borneo (au = 100). the southern china ecoregion stood alone (au = 95). the third cluster consisted of ecoregions from the ao and the cold temperate nwp (sea of okhotsk, chukchi sea, east siberian sea, laptev sea, sea of japan, and the yellow sea; au = 97). the first part of cluster number 4 included only ecoregions from the central indo-pacific (south kuroshio, halmahera, sulawesi sea/makassar strait, sunda shelf/java sea, gulf of thailand, and southern vietnam), whereas the sec‑ ond part was more spread out, including ecoregions ranging from central indo-pacific and eastern in‑ do-pacific to the temperate northern pacific (central kuroshio current, marshall islands, north-eastern honshu, mariana islands, and ogasawara islands). cluster 5 (au = 97) included mostly ecoregions from the temperate northern pacific (east china sea, oyashio current, aleutian islands and kamchatka shelf and coast), plus two from the central indo-pa‑ cific (west caroline islands and northeast sulawe‑ si). cluster 7 consisted only of ecoregions from the central indo-pacific (gulf of tonkin, malacca strait, east caroline islands, and south china sea oceanic islands; au = 96). discussion distribution and diversity in this study, we used open-access data in com‑ bination with data from four russian–german and german–russian benthic deep-sea expeditions to the nwp. the open-access data were only available in english. the resulting dataset therefore suffered from taxonomic, language (i.e., no local languages), and geographic bias. we documented high alpha species richness (number of species per hexagonal cell), as well as high es15 values (expected species richness) for the area around the philippines. this concentration was observed for both shallow waters and the deep sea. these findings are supported by earlier studies (renema et al. 2008; costello and chaudhary 2017; saeedi et al. 2019b). in fact, the indo-australian ar‑ chipelago (iaa) is considered as the most diverse region of the world ocean (renema et al. 2008; chaudhary et al. 2016; chaudhary et al. 2017; saee‑ di et al. 2017a; saeedi et al. 2017b). the past and present hotspots of marine species richness mark the locations of collisions between tectonic plates (ren‑ ema et al. 2008). during the early stage of plate col‑ lisions, new shallow-water habitats are created and new islands appear (provoost and bosch 2020). this figure 7: cluster analysis (pvclust) of isopoda and polychae‑ ta in ecoregions of the nw pacific and the arctic ocean. the numbers above each edge show the probability of nodes below that edge occurring as a cluster in resampled trees, via ordinary bootstrap resampling (bp, green) or multi‑ scale bootstrap resampling (au, red); distance: correlation; cluster method: average. hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 20 process results in altered ocean circulation, more heterogeneous habitats, and opportunities for isola‑ tion of populations, which all have been identified as driving factors of species richness (jokiel and marti‑ nelli 1992; bellwood and hughes 2001; bellwood et al. 2005; leprieur et al. 2016). temperature has also been found to be significantly correlated with species richness (tittensor et al. 2010; costello and chaud‑ hary 2017; saeedi et al. 2017b; saeedi and costello 2019; saeedi et al. 2019c): higher temperatures can speed metabolism and mutation rates, which in turn may result in higher speciation potential (costello and chaudhary 2017), as well as allowing for a wid‑ er range of energetic lifestyles (clarke and gaston 2006). our results also revealed hotspots of species richness in isopoda and polychaeta in convergence zones and areas of high temperature. in addition to the philippine hotspot of species richness in the nwp, we have also found species richness hotspots (alpha and es15) in the ao for both isopoda and polychaeta. high alpha species richness values of marine species in the ao have been already reported in our previous study (saeedi et al. 2019b). in particular, values of es15 showed that the ao could be very rich if we account for sam‑ pling bias. as such, species richness in the ao might be higher than has been thought before, given lower sampling effort compared to other areas (bodil et al. 2011; saeedi et al. 2019b). this sampling gap is most likely the case for many other marine taxa, especially for deep-sea fauna. for example, analysis of deep-sea isopod records here presented a high alpha species richness as well as and high es15 value in the deep kuril-kamchat‑ ka trench (kkt) area. our results match earlier findings (brandt and malyutina 2015; golovan et al. 2018). this high diversity in the deep sea of the kkt has been reported to be most likely correlated with food availability (golovan et al. 2018). golo‑ van et al. (2018) reported that peracarid abundance and species richness increased with a high sediment organic carbon content. compared to deep-sea iso‑ pods, deep-sea polychaetes did not show high alpha species richness in the kkt area in our study. alpha species richness was expected be high‑ est in the iaa region in polychaeta. nonetheless, we found that polychaetes had highest alpha species richness in the laptev sea, in both shallow water and deep sea. higher polychaete species richness in the ao compared to the nwp has been already reported by saeedi et al. (2019b). older studies reported gen‑ erally low ao species richness (bilyard and carey jr 1980; kupriyanova and badyaev 1998), whereas recent studies revised species richness estimates up‑ ward (sirenko 2001; kędra et al. 2013). we suggest that the species richness in the ao is likely higher than expected before owing to low sampling effort. dispersal and distribution in this study, we compared brooding isopods with polychaetes that have pelagic larval phases (includ‑ ing planktotrophic and lecithotrophic larvae), and polychaetes with unknown pelagic larvae. in shallow waters, polychaetes with pelagic larval phases had a wider distribution compared to brooding isopods. the distributions of polychaetes with pelagic larvae and those with unknown pelagic larvae did not differ markedly. deep sea isopods had wider distributions compared to polychaete species. previous work has shown that polychaetes with pelagic larval phases have significantly wider distributions compared to brooding isopods in shallow waters, which could be a function of the dispersal capabilities of the poly‑ chaete larvae (grantham et al. 2003; siegel et al. 2003). grantham et al. (2003) studied the dispersal potential of marine invertebrates. they found that planktotrophic polychaetes stayed an average of 57 days in the planktonic phase and could disperse tens or hundreds of kilometers. they also discovered that lecithotrophs stayed on average 7 days in the plank‑ tonic phase, and suggested a lower dispersal rate for this group (grantham et al. 2003). planktonic larval durations are reported to be correlated with disper‑ sal distance (shanks 2009), which typically ranges from 20–40 km for short plds and >200 km for long plds (siegel et al. 2003). species with greater dispersal potential suggested a higher probability to reach areas of favorable conditions, as well as lower risk of extinction (pechenik 1999; eckert 2003). one difference between the two studied taxa was that ~5.7% of polychaete species were present in both nwp and the ao, but only 0.3% of isopod spe‑ cies were reported from both oceans. one possible explanation for this contrast might be the dispersal capabilities of the polychaete larvae (grantham et al. 2003; siegel et al. 2003), although it could result oth‑ er influences, such as sampling bias and taxonomic expert bias. in the deep sea; however, isopoda had a wider distribution compared to polychaeta with a pelagic larval phase. the possibly higher protection through hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 21 brooding (pechenik 1999) and active selection for favorable conditions (hunt and seibel 2000) might be responsible for this. intuitively, one would assume that greater dispersal capabilities always result in a wider distribution range, but the following variables must be considered as well. ecological niche-related constraints like temperature can limit the areas where larvae can survive (verween et al. 2007; talmage and gobler 2011; chaudhary et al. 2016; saeedi et al. 2017b); ocean currents and turbidity influence dispersal of larvae significantly (bhaud 1998; bhaud 2000; carson and hentschel 2006)); and interand intraspecific competition and resistance to pollutants and microbes must also be considered (economou 1991; pechenik 1999; mcedward 2020). the isopod suborder asellota had by far the highest density in shallow waters, as well as in the deep sea. it has been recorded that the diversity of the suborder asellota increases in some cases with depth, while the diversity of other isopod suborders decreases with depth (saeedi and brandt 2020a). in this study, the polychaete suborder bonelliida had highest density in the deep sea, compared to other polychaete suborders. they have been described as a characteristic part of the fauna in the abyssal and hadal zones (zenkevitch 1966), but information on these taxa is scarce and no recent studies are avail‑ able. further research is necessary to understand the mechanisms of larval dispersal and their implica‑ tions for distributions of adult organisms. different methods of marking larvae could allow for direct measurement of larval dispersal, such as using artifi‑ cial markers, environmental tags, or genetic markers (thorrold et al. 2002; cowen and sponaugle 2009). through these methods, dispersal distances can be observed, and larval trajectories can be estimated (jones et al. 1999). in this way, it is possible to in‑ vestigate the relationship between dispersal and dis‑ tribution, as well as population connectivity. these factors must be considered in the design of effective marine reserves. latitudinal gradients species richness decreased towards higher lat‑ itudes in both isopods and polychaetes across the nwp and the ao. the this pattern, in combination with a dip near the equator, has been reported to cor‑ relate with sea temperature (chaudhary et al. 2016; chaudhary et al. 2017; saeedi et al. 2017b). bivalve larvae, for example, could not stand temperatures >28°c, which led to higher embryo and larval mor‑ tality (verween et al. 2007; talmage and gobler 2011). a clear northward shift of gamma species richness was observed in deep-sea isopods. latitude 80° n had an es15 value of ~10, and was therefore just as species-rich as the temperate latitudes. a peak in species richness in temperate deep-sea latitudes has been found in other taxa (rex et al. 2000). since productivity in the deep sea is general‑ ly low (except for hydrothermal vent areas), input of particulate organic matter (pom) plays an important role as a driver of species richness (woolley et al. 2016; golovan 2018). sediment heterogeneity has been shown to increase species richness, and is im‑ portant to consider in the context of deep-sea habi‑ tats owing to the great number of deposit feeders that rely on organic detritus as a food source (levin et al. 2001). another important factor shaping deep-sea species richness is oxygen concentration (levin and gage 1998; saeedi et al. 2020): levin et al. (2001) stated that oxygen-depleted bottom water typically resulted in strongly reduced macrofaunal diversity. the same drivers of species richness, in both shallow waters and in the deep sea, must be considered when investigating polychaete distribution patterns. as documented for isopods, polychaete species richness in shallow waters decreased at higher lat‑ itudes and increased towards temperate latitudes in the deep sea. nonetheless, highest species richness, for shallow-water as well as deep-sea polychaetes, was recorded at latitude 80° n. this result is consis‑ tent with our previous findings (saeedi et al. 2019b), who attributed this effect to higher sampling and tax‑ onomy efforts in that area. the lower species richness at latitudes 55–75° n and >85° n might be explained by the relatively little available ocean area. bathymetric gradients in the present research, a clear decline in species richness from shallow waters to the deep sea was re‑ corded in both isopods and polychaetes. most sam‑ ples were reported in depths above 200 m, and were correlated with highest species richness values. both isopod and polychaete species richness peaked again at a depth of ~600 m. isopod sampling effort and spe‑ cies richness increased greatly a depth of ~5000 m. most isopod records at this depth originate from the kurambio expedition in 2012 (brandt and malyutina 2015). the peak around south korea was not visi‑ ble for records above or below 200 m, meaning that these records had missing depth information. hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 22 es15 values for depths 5000–5600 m were com‑ parable to es15 value for depths ~1000 m. in general, species richness decreases with depth (costello and chaudhary 2017; saeedi et al. 2019b). this linear de‑ cline of species richness with depth, in combination with peaks at ~500 m and again at abyssal depths, has appeared in other studies of different taxa (rex et al. 2005; yasuhara et al. 2012; jöst et al. 2019; saeedi et al. 2020). variability in species richness in relation to depth can be explained by heterogeneity of pom (golovan et al. 2018), sediment and habitat (levin et al. 2001), and oxygen concentration (levin and gage 1998; levin et al. 2001; saeedi et al. 2020). cluster analysis cluster analysis of the ecoregions revealed six significantly different groups of isopoda and poly‑ cheta distribution. the first consisted of palawan north borneo and the eastern philippines, which en‑ compasses the ocean area directly around the philip‑ pines and indicates a distinct community in this area. these finding are in agreement with earlier studies (roberts et al. 2002; saeedi et al. 2019b). the sec‑ ond cluster consisted only of the southern china sea, which was surprising in view of the great distance to cluster 6, which includes the gulf of tonkin and the south china sea oceanic islands. given their imme‑ diate geographic connection, higher similarity was expected. this distance could be explained by a high endemicity rate in benthic species, which has been reported for the gulf of tonkin (saeedi et al. 2019b). cluster number 3 revealed a high similarity between the sea of okhotsk and the arctic ocean, probably due to the waters that are flowing through the ber‑ ing strait which heavily influence the ao (woodgate and aagaard 2005). the cluster analysis revealed the generally expected groups which were explained by rates of endemicity (roberts et al. 2002; saeedi et al. 2019b) and hydrological characteristics (woodgate and aagaard 2005). conclusions the present research documented high species richness around the philippines and in the laptev sea in isopoda and polychaeta. these values were cou‑ pled with high es15 values, suggesting no high cor‑ relation between sampling bias and species richness in regards to the philippines and to the ao. latitudi‑ nal species richness in both isopoda and polychae‑ ta decreased towards higher latitudes, in agreement with recent studies, such as declining species rich‑ ness with latitude in shallow waters, an increase in species richness towards temperate latitudes in the deep sea, and the majority of species occurring in coastal depths. however, a peak in the expected number of species was observed for polychates in the ao, highlighting the importance of sampling bias in geographic studies of specific taxa. to conserve global biodiversity, we need to un‑ derstand underlying mechanisms that operate at this scale. the present study demonstrated a fundamental need for higher sampling effort across all latitudes and depth intervals. ao species richness is suspected to be higher than appreciated previously. thus, we advocate for increased sampling effort, with special consideration focus on the ao. the impact of larval dispersal on distributions of organisms should be in‑ tegrated into the design of marine protected areas. the information presented in this paper adds to the global picture of species richness patterns and func‑ tional biodiversity, which is necessary to prioritize areas for employing effective conservation efforts. declarations ethics approval and consent to participate not applicable consent for publication not applicable availability of data and materials the data used in this study retrieved from open-access databases including obis and gbif (dataset citations are in doi: https://doi.org/10.1038/ s41598-019-45813-9) in combination with our own expedition data published in obis (doi: http://ipt. iobis.org/obis-deepsea/resource?r=beneficial_deep‑ sea&v=1 and doi: https://www.gbif.org/data‑ set/24c56165-469b-426a-98bb-2cbdef596680). the cleaned dataset analyzed in this study, as well as the r scripts, are shared via github1. supplementary ma‑ terials are available also2. competing interests the authors declare that they have no competing interests. 1 https://github.com/haniehsaeedi/biodiversity-and-dis‑ tribution-of-isopoda-and-polychaeta. 2 http://hdl.handle.net/1808/32555. https://doi.org/10.1038/s41598-019-45813-9 https://doi.org/10.1038/s41598-019-45813-9 http://ipt.iobis.org/obis-deepsea/resource?r=beneficial_deepsea&v=1 http://ipt.iobis.org/obis-deepsea/resource?r=beneficial_deepsea&v=1 http://ipt.iobis.org/obis-deepsea/resource?r=beneficial_deepsea&v=1 https://www.gbif.org/dataset/24c56165-469b-426a-98bb-2cbdef596680 https://www.gbif.org/dataset/24c56165-469b-426a-98bb-2cbdef596680 https://github.com/haniehsaeedi/biodiversity-and-distribution-of-isopoda-and-polychaeta https://github.com/haniehsaeedi/biodiversity-and-distribution-of-isopoda-and-polychaeta http://hdl.handle.net/1808/32555 hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 23 funding this paper was part of the “biogeography of the nw pacific deep-sea fauna and their possible future invasions into the arctic ocean project (ben‑ eficial project)”. beneficial project (grant number 03f0780a) was funded by federal ministry for edu‑ cation and research (bmbf: bundesministerium für bildung und forschung) in germany. authors’ contributions hs developed the idea of the paper, contributed in running the analyses, and was a major contributor in writing the manuscript. nlj developed the idea of the paper, ran the analyses, and was a major con‑ tributor in writing the manuscript. ab contributed in developing the idea and helped in writing the paper. all authors read and approved the final manuscript. acknowledgements we would like to thank laura breitkreuz for the english proofreading of the manuscript. we would like to also thank the anonymous reviewers for their constructive comments, and the editor, townsend pe‑ terson, who has majorly edited the last version of this manuscript. references barroso, r., j. d. kudenov, k. m. halanych, h. saeedi, p. y. g. sumida, and a. f. bernardino. 2018. a new species of xylophylic fireworm (annelida: amphin‑ omidae: cryptonome) from deep-sea wood falls in the sw atlantic. deep sea research part i: oceano‑ graphic research papers 137:66-75. bellwood, d. r., t. hughes, s. connolly, and j. tanner. 2005. environmental and geometric constraints on indo‐pacific coral reef biodiversity. ecology letters 8:643-651. bellwood, d. r. and t. p. hughes. 2001. regional-scale assembly rules and biodiversity of coral reefs. sci‑ ence (new york, n.y.) 292:1532-1535. bhaud, m. 1998. the spreading potential of polychaete larvae does not predict adult distributions; conse‑ quences for conditions of recruitment. hydrobiologia 375:35-47. bhaud, m. 2000. some examples of the contribution of planktonic larval stages to the biology and ecology of polychaetes. bulletin of marine science 67:345-358. bilyard, g. r. and a. g. carey jr. 1980. zoogeography of western beaufort sea polychaeta (annelida). sarsia 65:19-25. bodil, b. a., w. g. ambrose, m. bergmann, l. m. clough, a. v. gebruk, c. hasemann, k. iken, m. klages, i. r. macdonald, and p. e. renaud. 2011. diversity of the arctic deep-sea benthos. marine biodiversity 41:87107. brandt, a., w. brökeland, s. brix, and m. malyutina. 2004. diversity of southern ocean deep-sea isopoda (crustacea, malacostraca)—a comparison with shelf data. deep sea research part ii: topical studies in oceanography 51:1753-1768. brandt, a., n. o. elsner, m. v. malyutina, n. brenke, o. a. golovan, a. v. lavrenteva, and t. riehl. 2015. abyssal macrofauna of the kuril–kamchatka trench area (northwest pacific) collected by means of a cam‑ era–epibenthic sledge. deep sea research part ii: topical studies in oceanography 111:175-187. brandt, a. and m. malyutina. 2014. german-russian deep-sea expedition kurambio (kurile kamchatka biodiversity study)–results and perspectives. brandt, a., m. malyutina, n. majorova, a. bashmanov, n. brenke, t. chizhova, n. elsner, o. golovan, c. göcke, and d. kaplunenko. 2010. the russian-ger‑ man deep-sea expedition (sojabio). brandt, a. and m. v. malyutina. 2015. the german-rus‑ sian deep-sea expedition kurambio (kurile kamchat‑ ka biodiversity studies) on board of the rv sonne in 2012 following the footsteps of the legendary expe‑ ditions with rv vityaz. deep sea research part ii: topical studies in oceanography 111:1-9. brown, j. h. 2014. why are there so many species in the tropics? journal of biogeography 41:8-22. bush, s. l., h. j. hoving, c. l. huffard, b. h. robison, and l. d. zeidberg. 2012. brooding and sperm storage by the deep-sea squid bathyteuthis berryi (cephalopoda: decapodiformes). journal of the marine biological association of the united kingdom 92:1629-1636. carson, h. s. and b. t. hentschel. 2006. estimating the dispersal potential of polychaete species in the southern california bight: implications for design‑ ing marine reserves. marine ecology progress series 316:105-113. chaudhary, c., h. saeedi, and m. j. castello. 2016. bi‑ modality of latitudinal gradients in marine species richness. trends in ecology & evolution 31:670676. chaudhary, c., h. saeedi, and m. j. costello. 2017. ma‑ rine species richness is bimodal with latitude: a reply to fernandez and marques. trends in ecology & evolution 32:234-237. chernyshev, a. v. and n. e. polyakova. 2018. nemerteans from deep-sea expedition sokhobio with description hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 24 of uniporus alisae sp. nov.(hoplonemertea: reptantia sl) from the sea of okhotsk. deep sea research part ii: topical studies in oceanography 154:121-139. chiarucci, a., g. bacaro, d. rocchini, and l. fattori‑ ni. 2008. discovering and rediscovering the sam‑ ple-based rarefaction formula in the ecological liter‑ ature. community ecology 9:121-123. clarke, a. and k. j. gaston. 2006. climate, energy and diversity. proceedings. biological sciences 273:22572266. costello, m. j. and c. chaudhary. 2017. marine biodiver‑ sity, biogeography, deep-sea gradients, and conserva‑ tion. current biology 27:r511-r527. cowen, r. k. and s. sponaugle. 2009. larval dispersal and marine population connectivity. annual review of marine science 1:443-466. eckert, g. l. 2003. effects of the planktonic period on marine population fluctuations. ecology 84:372-383. economou, a. n. 1991. is dispersal of fish eggs, embryos and larvae an insurance against density dependence? environmental biology of fishes 31:313-321. fauchald, k. 1983. life diagram patterns in benthic poly‑ chaetes. proceedings of the biological society of washington. glover, a. g., n. higgs, and t. horton. 2021. world reg‑ ister of deep-sea species (wordss). http://www. marinespecies.org/deepsea golovan, o. a. 2018. desmosomatidae (isopoda: asello‑ ta) from the kuril basin of the sea of okhotsk: first data on diversity with the description of the dominant species mirabilicoxa biramosa sp. nov. deep sea research part ii: topical studies in oceanography 154:292-307. golovan, o. a., m. błażewicz-paszkowycz, a. brandt, l. l. budnikova, n. o. elsner, v. v. ivin, a. v. lavrente‑ va, m. v. malyutina, v. v. petryashov, and l. a. tza‑ reva. 2013. diversity and distribution of peracarid crustaceans (malacostraca) from the continental slope and the deep-sea basin of the sea of japan. deep sea research part ii: topical studies in oceanography 86:66-78. golovan, o. a., m. v. malyutina, and a. brandt. 2018. first record of the deep-sea isopod family dendro‑ tionidae (isopoda: asellota) from the northwest pa‑ cific with description of two new species of dendro‑ munna. marine biodiversity 48:531-544. grantham, b. a., g. l. eckert, and a. l. shanks. 2003. dispersal potential of marine invertebrates in diverse habitats: ecological archives a013-001-a1. ecologi‑ cal applications 13:108-116. grebmeier, j. m., l. w. cooper, h. m. feder, and b. i. sirenko. 2006. ecosystem dynamics of the pacific-in‑ fluenced northern bering and chukchi seas in the amerasian arctic. progress in oceanography 71:331361. hessler, r. r. and p. a. jumars. 1974. abyssal commu‑ nity analysis from replicate cores in the central north pacific. pp. 185-209. deep sea research and ocean‑ ographic abstracts. elsevier. hunt, j. and b. seibel. 2000. life history of gonatus onyx (cephalopoda: teuthoidea): ontogenetic changes in habitat, behavior and physiology. marine biology 136:543-552. jokiel, p. and f. martinelli. 1992. the vortex model of coral reef biogeography. journal of biogeography:449-458. jones, g. p., m. milicich, m. emslie, and c. lunow. 1999. self-recruitment in a coral reef fish population. na‑ ture 402:802-804. jöst, a. b., m. yasuhara, c. l. wei, h. okahashi, a. os‑ tmann, p. martínez arbizu, b. mamo, j. svavarsson, and s. brix. 2019. north atlantic gateway: test bed of deep‐sea macroecological patterns. journal of bio‑ geography 46:2056-2066. kaiser, s., d. k. barnes, c. j. sands, and a. brandt. 2009. biodiversity of an unknown antarctic sea: assessing isopod richness and abundance in the first benthic sur‑ vey of the amundsen continental shelf. marine biodi‑ versity 39:27-43. kędra, m., k. pabis, s. gromisz, and j. m. węsławski. 2013. distribution patterns of polychaete fauna in an arctic fjord (hornsund, spitsbergen). polar biology 36:1463-1472. krug, p. j. and r. k. zimmer. 2004. developmental di‑ morphism: consequences for larval behavior and dis‑ persal potential in a marine gastropod. the biological bulletin 207:233-246. kupriyanova, e. k. and a. v. badyaev. 1998. ecological correlates of arctic serpulidae (annelida, polychaeta) distributions. ophelia 49:181-193. leprieur, f., p. descombes, t. gaboriau, p. f. cowman, v. parravicini, m. kulbicki, c. j. melián, c. n. de san‑ tana, c. heine, and d. mouillot. 2016. plate tecton‑ ics drive tropical reef biodiversity dynamics. nature communications 7:1-8. levin, l. a. 2006. recent progress in understanding larval dispersal: new directions and digressions. integrative and comparative biology 46:282-297. levin, l. a. and t. s. bridges. 1995. pattern and diversity in reproduction and development. ecology of marine invertebrate larvae 1:48. http://www.marinespecies.org/deepsea hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 25 levin, l. a., r. j. etter, m. a. rex, a. j. gooday, c. r. smith, j. pineda, c. t. stuart, r. r. hessler, and d. pawson. 2001. environmental influences on regional deep-sea species diversity. annual review of ecology and systematics 32:51-93. levin, l. a. and j. d. gage. 1998. relationships between oxygen, organic matter and the diversity of bathyal macrofauna. deep sea research part ii: topical stud‑ ies in oceanography 45:129-163. malyutina, m. v. and a. brandt. 2013. introduction to so‑ jabio (sea of japan biodiversity studies). deep sea research part ii: topical studies in oceanography 86:1-9. malyutina, m. v. and a. brandt. 2018. first records of deep-sea munnopsidae (isopoda: asellota) from the kuril basin of the sea of okhotsk, with description of gurjanopsis kurilensis sp. nov. deep sea research part ii: topical studies in oceanography 154:275291. mcclain, c. r., a. p. allen, d. p. tittensor, and m. a. rex. 2012. energetics of life on the deep seafloor. proceed‑ ings of the national academy of sciences 109:1536615371. mcedward, l. 2020. ecology of marine invertebrate lar‑ vae. crc press. mileikovsky, s. 1971. types of larval development in ma‑ rine bottom invertebrates, their distribution and eco‑ logical significance: a re-evaluation. marine biology 10:193-213. oksanen, j., f. g. blanchet, m. friendly, r. kindt, p. legendre, d. mcglinn, p. minchin, r. o’hara, g. simpson, and p. solymos. 2019. vegan: community ecology package (r package version 2.5-5) https:// cran.r-project.org/package=vegan. community ecology package. ostenso, n. a. 1962. geophysical investigations of the arctic ocean basin. the university of wiscon‑ sin-madison. pebesma, e. 2018. simple features for r: standardized support for spatial vector data, r j., 10, 439–446. pechenik, j. a. 1999. on the advantages and disadvantag‑ es of larval stages in benthic marine invertebrate life cycles. marine ecology progress series 177:269-297. provoost, p. and s. bosch. 2020. robis: ocean biodiversity information system (obis) client. . read, g. and k. fauchald. 2020. world polychaeta data‑ base. pseudopotamilla laciniosa. renema, w., d. r. bellwood, j. c. braga, k. bromfield, r. hall, k. g. johnson, p. lunt, c. p. meyer, l. b. mcmonagle, r. j. morley, a. o’dea, j. a. todd, f. p. wesselingh, m. e. j. wilson, and j. m. pandolfi. 2008. hopping hotspots: global shifts in marine biodiversity. science (new york, n.y.) 321:654-657. rex, m. a., j. a. crame, c. t. stuart, and a. clarke. 2005. large-scale biogeographic patterns in marine mol‑ lusks: a confluence of history and productivity? ecol‑ ogy 86:2288-2297. rex, m. a., c. t. stuart, and g. coyne. 2000. latitudinal gradients of species richness in the deep-sea benthos of the north atlantic. proceedings of the national academy of sciences of the united states of america 97:4082-4085. roberts, c. m., c. j. mcclean, j. e. veron, j. p. hawkins, g. r. allen, d. e. mcallister, c. g. mittermeier, f. w. schueler, m. spalding, and f. wells. 2002. marine biodiversity hotspots and conservation priorities for tropical reefs. science (new york, n.y.) 295:12801284. saeedi, h., z. basher, and m. j. costello. 2017a. model‑ ling present and future global distributions of razor clams (bivalvia: solenidae). helgoland marine re‑ search 70:23. saeedi, h., a. f. bernardino, m. shimabukuro, g. fal‑ chetto, and p. y. g. sumida. 2019a. macrofaunal community structure and biodiversity patterns based on a wood-fall experiment in the deep south-west atlantic. deep sea research part i: oceanographic research papers. saeedi, h. and a. brandt. 2020a. biogeographic atlas of the deep nw pacific fauna. advanced books 1. saeedi, h. and a. brandt. 2020b. nw pacific deep-sea benthos biodiversity (beneficial project), v1.2. deep-sea obis node. dataset/samplingevent. saeedi, h. and m. costello. 2012. aspects of global distri‑ bution of six marine bivalve mollusc families. clam fisheries and aquaculture:27-44. saeedi, h. and m. j. costello. 2019. a world dataset on the geographic distributions of solenidae razor clams (mollusca: bivalvia). biodiversity data journal 7. saeedi, h., m. j. costello, d. warren, and a. brandt. 2019b. latitudinal and bathymetrical species richness patterns in the nw pacific and adjacent arctic ocean. scientific reports 9:9303. saeedi, h., t. e. dennis, and m. j. costello. 2017b. bi‑ modal latitudinal species richness and high endemic‑ ity of razor clams (mollusca). journal of biogeogra‑ phy 44:592-604. saeedi, h., j. d. reimer, m. i. brandt, p.-o. dumais, a. m. jażdżewska, n. w. jeffery, p. m. thielen, and m. j. costello. 2019c. global marine biodiversity in the context of achieving the aichi targets: ways forward and addressing data gaps. peerj 7:e7221. https://cran.r-project.org/package=vegan https://cran.r-project.org/package=vegan hanieh saeedi et al. – biodiversity and distribution of isopoda and polychaeta 26 saeedi, h., m. simões, and a. brandt. 2019d. endemicity and community composition of marine species along the nw pacific and the adjacent arctic ocean. prog‑ ress in oceanography 178:102199. saeedi, h., m. simões, and a. brandt. 2020. biodiversi‑ ty and distribution patterns of deep-sea fauna along the temperate nw pacific. progress in oceanography 183:102296. saeedi, h., d. warren, and a. brandt. in press. the envi‑ ronmental drivers of benthic fauna diversity and com‑ munity composition. frontiers in marine science. saeedi, h., winterberg h, alalykina li, bergmeier sf, downey r, golovan o, jażdżewska a, kamenev g, kameneva n, maiorova a, malyutina m, minin k, mordukhovich v, petrunina a, schwabe e, and b. a. 2019e. nw pacific deep-sea benthos distribution and abundance (beneficial project). deep-sea obis node. dataset/samplingevent 1. sakshaug, e. 2004. primary and secondary production in the arctic seas. pp. 57-81. the organic carbon cycle in the arctic ocean. springer. shanks, a. l. 2009. pelagic larval duration and dispersal distance revisited. the biological bulletin 216:373385. siegel, d., b. kinlan, b. gaylord, and s. gaines. 2003. lagrangian descriptions of marine larval dispersion. marine ecology progress series 260:83-96. simões, m. v., h. saeedi, m. e. cobos, and a. brandt. 2021. environmental matching reveals non-uniform range-shift patterns in benthic marine crustacea. cli‑ matic change 168:1-20. sirenko, b. i. 2001. list of species of free-living inver‑ tebrates of eurasian arctic seas and adjacent deep waters. russian academy of science, zoological in‑ stitute. spalding, m. d., h. e. fox, g. r. allen, n. davidson, z. a. ferdana, m. finlayson, b. s. halpern, m. a. jorge, a. lombana, and s. a. lourie. 2007. marine ecore‑ gions of the world: a bioregionalization of coastal and shelf areas. bioscience 57:573-583. springer, a. m. and c. p. mcroy. 1993. the paradox of pelagic food webs in the northern bering sea—iii. patterns of primary production. continental shelf re‑ search 13:575-599. suzuki, r. and h. shimodaira. 2019. pvclust: hierarchical clustering with p-values via multiscale bootstrap res‑ ampling. r package version 2.2-0. the r foundation. vienna. talmage, s. c. and c. j. gobler. 2011. effects of elevated temperature and carbon dioxide on the growth and survival of larvae and juveniles of three species of northwest atlantic bivalves. plos one 6. thorrold, s. r., g. p. jones, m. e. hellberg, r. s. burton, s. e. swearer, j. e. neigel, s. g. morgan, and r. r. warner. 2002. quantifying larval retention and con‑ nectivity in marine populations with artificial and nat‑ ural markers. bulletin of marine science 70:291-308. tittensor, d. p., c. mora, w. jetz, h. k. lotze, d. ricard, e. v. berghe, and b. worm. 2010. global patterns and predictors of marine biodiversity across taxa. nature 466:1098. valentine, j. w. and d. jablonski. 2015. a twofold role for global energy gradients in marine biodiversity trends. journal of biogeography 42:997-1005. verween, a., m. vincx, and s. degraer. 2007. the effect of temperature and salinity on the survival of mytilopsis leucophaeata larvae (mollusca, bivalvia): the search for environmental limits. journal of experi‑ mental marine biology and ecology 348:111-120. wägele, j.-w. 1989. evolution und phylogenetisches sys‑ tem der isopoda. westheide, w. and r. rieger. 1996. spezielle zoologie, teil 1: einzeller und wirbellose. gustav fischer ver‑ lag, jena, new york [in german]. wickham, h. 2016. ggplot2: elegant graphics for data analysis. springer. wickham, h., j. bryan, m. kalicinski, k. valery, c. leiti‑ enne, b. colbert, d. hoerl, e. miller, and m. j. bryan. 2019. package ‘readxl’. wickham, h., r. françois, and l. henry. 2021. müller k. dplyr: a grammar of data manipulation. 2020. r pack‑ age version 0.8 4. wickham, h. and l. henry. 2020. tidyr: tidy messy data. r package version 1:397. wilson, w. h. 1991. sexual reproductive modes in poly‑ chaetes: classification and diversity. bulletin of ma‑ rine science 48:500-516. woodgate, r. a. and k. aagaard. 2005. revising the ber‑ ing strait freshwater flux into the arctic ocean. geo‑ physical research letters 32. woolley, s. n., d. p. tittensor, p. k. dunstan, g. guil‑ lera-arroita, j. j. lahoz-monfort, b. a. wintle, b. worm, and t. d. o’hara. 2016. deep-sea diversity patterns are shaped by energy availability. nature 533:393. yasuhara, m. and r. danovaro. 2016. temperature im‑ pacts on deep‐sea biodiversity. biological reviews 91:275-287. yasuhara, m., g. hunt, h. j. dowsett, m. m. robinson, and d. k. stoll. 2012. latitudinal species diversity gradient of marine zooplankton for the last three mil‑ lion years. ecology letters 15:1174-1179. zenkevitch, l. 1966. the systematics and distribution of abyssal and hadal (ultraabyssal) echiuroidea. gala‑ thea report 8:175-184. microsoft word jeanganglo_20160707.docx biodiversity informatics, 11, 2016, pp. 23-39 23 completeness of digital accessible knowledge of plants of benin and priorities for future inventory and data discovery jean cossi ganglo* and sunday berlioz kakpo laboratory of forest sciences, faculty of agricultural sciences, university of abomeycalavi, 03 bp 393, cotonou, benin. *corresponding author: ganglocj@gmail.com. abstract.—discovery of and access to primary biodiversity data are critical components in informed decision-making regarding sustainable use of biological resources and conservation of biodiversity. primary biodiversity data are increasingly available from benin, but information about completeness of this information across the country is still lacking for most groups. this study analyzed the digital accessible knowledge regarding the plants of benin to identify gaps in both geographic and environmental dimensions. many gaps exist in plant data for benin, particularly in the northern most departments; central and southern benin are better known, but some gaps remain even there. the resulting view of beninese digital accessible knowledge can guide future inventory and data discovery efforts. key words.—benin, inventory completeness, data cleaning, digital accessible knowledge, gaps, gbif. primary biodiversity data are data that combine three essential attributes: geographic location, collection date, and taxonomic identification (johnson, 2007; soberón and peterson, 2009; sousa-baena et al. 2013). digital accessible knowledge (dak) is the part of the existing primary biodiversity data that is (1) in digital format, (2) published and accessible worldwide for free, and (3) integrated into the broader global storehouse of biodiversity information (in essence, a transformation of data into “knowledge”; sousa-baena et al. 2013). dak is essential in enabling decision-making on natural resource management, especially biodiversity conservation (peterson et al. 2000; peterson et al. 2001; peterson et al. 2004; kremen et al. 2008; morris et al. 2013; gaiji et al. 2013). the global biodiversity information facility (gbif1) is by far the largest initiative assembling and sharing dak on biodiversity, with the aim of sustaining scientific research, conservation, and sustainable development (gaiji et al. 2013, peterson et al. 2015). gbif was created in 2001, and—as of may 2016—provides access to >640m records from more than 1.6m species. although plant collections have accumulated from benin since the 1780s (akoegninou et al., 2006), digitization and publication of primary occurrence data within the country began only recently (2010), in the framework of the activities of the global biodiversity information facility (gbif) through gbif-benin2. some dak for 1 http://www.gbif.org. 2 http://gbif-benin.org. 2 http://gbif-benin.org. benin was available previously through other sources3: indeed, institutions in 22 countries including benin presently contribute data about beninese biodiversity, with 170 occurrence datasets and >190,000 records. dak on the biodiversity of benin is of great importance to guiding and shaping priorities of decision-makers and engaging them in sustainable management of the limited and increasingly stressed natural resources of the country. biodiversity assessments versus dak availability in benin many studies have addressed biodiversity, in terms of its collections, patterns, and conservation, at various levels across benin. adomou (2005) studied distributional patterns of vegetation types of benin, and divided the country into 10 phytogeographic districts, and described the major vegetation types and conservation priorities. various individual vegetation types have been studied in terms of plant communities and species composition (ganglo et al. 1999; ganglo and lejoly 1999; sokpon et al. 2001; ganglo 2004; ganglo 2005; ganglo and de foucault 2006; tohngodo et al. 2006; awokou et al. 2009; noumon and ganglo 2005; noumon et al. 2009; aoudji et al. 2011; yêvidé et al.2011), resulting in classifications of forest sites, and useful recommendations for sustainable management. the diversity of neglected and underutilized species was assessed by dansi et al. (2012), who found 41 species of high importance in terms of nutrient content, medicinal value, 3 http://www.gbif.org/country/bj/about/datasets. biodiversity informatics, 11, 2016, pp. 23-39 24 figure 1. loss of dak records during the data-cleaning process for data regarding occurrences of the plants of benin. figure 2. spatial distribution of georeferenced plant records of benin. biodiversity informatics, 11, 2016, pp. 23-39 25 contribution to household income, poverty reduction, extent of consumption, degree of consumption, extent of production, availability during the year, contribution to empowerment of women, market value, and market use. several stands and populations of forest species, medicinal plants, and multipurpose species have been studied in benin in terms of structural characteristics, ecology, and usefulness (sokpon and biaou 2002; sokpon et al. 2006; gouwakinou et al. 2009; yêhouénou tessi et al. 2012; koura et al. 2011; koura et al. 2013a; koura et al. 2013b).the spatial genetic structure of baobab (adansonia digitata) populations from west african agroforestry systems was evaluated by kyndt et al. (2009): 11 populations from benin, ghana, burkina faso, and senegal had comparable levels of genetic diversity, and the organization of genetic diversity appeared to result essentially from spatially restricted gene flow, with some influences of human seed exchange. although of importance in advancing knowledge of vegetation types and structure, and their determinants and conservation priorities, these studies did not address representativeness of sampling of plant diversity across the country, which could help identify knowledge gaps, data resource limitations, and possible biases (sousabaena et al. 2013). the single exception is the palms (arecaceae), which were analyzed in detail in terms of completeness across benin by idohou et al. (2015). moreover, apart from scientific collections (ariño, 2010), which covered the whole country during the elaboration of the flora of benin (akoegninou et al. 2006), most of the studies stated above only addressed limited sectors of the country. worse yet, few, if any, of the occurrence data from those surveys are openly available to the scientific community. a survey of biodiversity data holders and users within benin (gbif benin, 2011) showed that the most important collection in terms of number of specimens is the national herbarium (50,000 specimens), but <30% of its specimens have been digitized and published by beninese institutions via gbif-benin and gbif. therefore, the representativeness of the dak with regard to the geography and species diversity of the country remains unclear. this study aims to develop a detailed assessment of dak for benin’s plants to assess its fitness for use in detailed analyses. methods data cleaning data used in this paper were downloaded from the gbif site in january 2015 using the filter “plantae” in the scientific name field on the benin page4. to ‘clean’ the data, to make optimal use of information available, we used an iterative series of cleaning steps, as follows. (1) we created lists of unique taxon names in each dataset in microsoft excel, and inspected them for multiple versions of the same taxonomic concepts: misspellings, name variants, different versions of authority information, synonyms, etc. such name variants were flagged, and checked with independent sources; a field was created to hold standardized preferred scientific names that correctly corresponded to single taxa. we used the list-matching service of the catalogue of life5 and prota6, both of which were accessed in 2015. (2) we checked for and removed incomplete or suspicious geographic coordinates (e.g., one or both of the coordinates either missing or falling outside of benin in spite of being referred to that country. (3) within the country, we checked for consistency between textual descriptions of departments (administrative divisions) and the position of geographic coordinates. in each case, where problems were detected, we created a corrected version of the data records; where no clear correction was possible, we discarded records, recording data losses at each step in the process. finally, (4) we discarded data records for which full information on year, month, and day of collection was lacking; we created a unique ‘stamp’ of time as year_month_day. inventory completeness next, we aggregated point-based occurrence data to 0.5° spatial resolution across the country. this spatial resolution was the product of detailed analyses of balancing the benefits of aggregating data (i.e., larger sample sizes) versus the loss of spatial resolution that accompanies broader aggregation areas and can make imperceptible important geographic features (i.e., 0.5° resolution is a square ~55 km on a side). those steps focused on choice of an optimal spatial resolution for analysis are detailed in another paper (ariño et al., in preparation). we produced the aggregation grid shapefiles in the vector grid module of qgis, version 2.6, 4 http://www.gbif.org/occurrence/search?country=bj#. 5 http://www.catalogueoflife.org/listmatching. 6 http://www.prota4u.info/protaindex.asp. biodiversity informatics, 11, 2016, pp. 23-39 35 table 1. inventory completeness in the departments of benin. departments percentage of grid cells or objects with no data (a) percentage of grid cells or objects not well sampled with n < 200 and c < 0.5 (b) percentage of grid cells or objects not well sampled with n ≥ 200 and c < 0.5 (c) total (a) + (b) + (c) percentage of grid cells or objects well sampled with 0.5 ≤ c < 0.6 (d) percentage of grid cells or objects well sampled with 0.6 ≤ c < 0.7 (e) percentage of grid cells or objects well sampled with 0.7 ≤ c < 0.8 (f) percentage of grid cells or objects well sampled with c ≥ 0.8 (g) alibori 1.27 9.28 0.00 10.55 0.00 0.00 1.27 2.95 borgou 1.27 5.91 0.00 7.18 0.00 0.00 6.33 2.53 atakora 0.42 5.06 0.00 5.48 0.00 0.00 5.91 2.11 kouffo 0.00 2.11 0.00 2.11 0.00 0.00 1.27 2.53 plateau 2.11 0.00 0.00 2.11 0.00 0.00 1.27 1.69 collines 0.00 1.69 0.00 1.69 0.00 1.69 2.11 3.80 donga 0.00 1.27 0.00 1.27 0.00 0.00 2.11 3.38 zou 0.00 0.84 0.00 0.84 0.00 0.00 2.11 5.06 mono 0.00 0.42 0.00 0.42 0.00 1.27 1.69 2.53 atlantique 0.00 0.00 0.00 0.00 0.00 2.11 1.69 5.06 ouémé 0.00 0.00 0.00 0.00 0.00 1.27 0.42 3.38 littoral 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.00 townpeterson typewritten text 26 biodiversity informatics, 11, 2016, pp. 23-39 27 table 2. inventory completeness in the municipalities of benin. department municipality percentage of grid cells or objects with no data (a) percentage of grid cells or objects not well sampled with n < 200 and c < 0.5 (b) percentage of grid cells or objects not well sampled with n ≥ 200 and c < 0.5 (c) total (a) + (b) + (c) percentage of grid cells or objects well sampled with 0.5 ≤ c < 0.6 (d) percentage of grid cells or objects well sampled with 0.6 ≤ c < 0.7 (e) percentage of grid cells or objects well sampled with 0.7 ≤ c < 0.8 (f) percentage of grid cells or objects well sampled with c ≥ 0.8 (g) alibori banikoara 0.00 2.11 0.00 2.11 0.00 0.00 0.00 0.42 gogounou 0.00 2.11 0.00 2.11 0.00 0.00 0.00 0.42 kandi 0.00 1.69 0.00 1.69 0.00 0.00 0.42 0.42 karimama 0.42 1.69 0.00 2.11 0.00 0.00 0.00 0.42 malanville 0.42 0.84 0.00 1.26 0.00 0.00 0.42 0.84 segbana 0.42 0.84 0.00 1.26 0.00 0.00 0.42 0.42 atakora boukoumbé 0.00 0.00 0.00 0.00 0.00 0.00 0.84 0.42 cobly 0.00 0.00 0.00 0.00 0.00 0.00 0.84 0.00 kérou 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.00 kouandé 0.00 0.00 0.00 0.00 0.00 0.00 0.84 0.42 matéri 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.00 natitingou 0.00 0.00 0.00 0.00 0.00 0.00 1.27 0.42 péhunco 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.42 tanguiéta 0.42 0.00 0.00 0.42 0.00 0.00 0.84 0.00 toucountouna 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 atlantique abomey-calavi 0.00 0.00 0.00 0.00 0.00 0.84 0.00 0.84 allada 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 kpomassè 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.42 ouidah 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.42 sô-ava 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.42 toffo 0.00 0.00 0.00 0.00 0.00 0.00 0.84 0.42 townpeterson typewritten text townpeterson typewritten text biodiversity informatics, 11, 2016, pp. 23-39 28 tori-bossito 0.00 0.00 0.00 0.00 0.00 0.00 0.42 1.27 zè 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.84 borgou bembéréké 0.00 1.27 0.00 1.27 0.00 0.00 0.00 0.42 kalalé 0.84 1.27 0.00 2.11 0.00 0.00 0.00 0.42 n'dali 0.00 0.84 0.00 0.84 0.00 0.00 1.69 0.42 nikki 0.42 0.42 0.00 0.84 0.00 0.00 0.42 0.42 parakou 0.00 0.00 0.00 0.00 0.00 0.00 0.84 0.00 pèrèrè 0.00 0.42 0.00 0.42 0.00 0.00 0.84 0.00 sinendé 0.00 0.84 0.00 0.84 0.00 0.00 0.42 0.42 tchaourou 0.00 0.84 0.00 0.84 0.00 0.00 2.11 0.42 collines bantè 0.00 0.84 0.00 0.84 0.00 0.42 0.42 0.00 dassa-zoumè 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.84 glazoué 0.00 0.00 0.00 0.00 0.00 0.42 0.84 1.27 ouèssè 0.00 0.00 0.00 0.00 0.00 0.00 0.84 0.42 savalou 0.00 0.84 0.00 0.84 0.00 0.42 0.00 0.42 savè 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.84 donga bassila 0.00 0.84 0.00 0.84 0.00 0.00 0.84 0.84 copargo 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.84 djougou 0.00 0.42 0.00 0.42 0.00 0.00 0.84 1.27 ouaké 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.42 kouffo aplahoué 0.00 0.84 0.00 0.84 0.00 0.00 0.42 0.42 djakotomey 0.00 0.42 0.00 0.42 0.00 0.00 0.42 0.00 klouékanmè 0.00 0.42 0.00 0.42 0.00 0.00 0.00 0.42 lalo 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.84 toviklin 0.00 0.42 0.00 0.42 0.00 0.00 0.42 0.84 littoral cotonou 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.00 mono athiémé 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.42 bopa 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 biodiversity informatics, 11, 2016, pp. 23-39 29 comè 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.42 dogbo 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.42 grand-popo 0.00 0.42 0.00 0.42 0.00 0.42 0.42 0.42 houéyogbé 0.00 0.00 0.00 0.00 0.00 0.42 0.42 0.42 ouémé adjarra 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.42 adjohoun 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 aguégués 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.42 akpromissérété 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 avrankou 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 bonou 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.42 dangbo 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 porto-novo 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 sèmè-kpodji 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.42 plateau adja-ouèrè 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.42 ifangni 0.42 0.00 0.00 0.42 0.00 0.00 0.00 0.42 kétou 0.84 0.00 0.00 0.84 0.00 0.00 0.42 0.42 pobè 0.42 0.00 0.00 0.42 0.00 0.00 0.42 0.00 sakété 0.42 0.00 0.00 0.42 0.00 0.00 0.00 0.42 zou abomey 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 agbangnizoun 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 bohicon 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 covè 0.00 0.00 0.00 0.00 0.00 0.00 0.42 1.27 djidja 0.00 0.84 0.00 0.84 0.00 0.00 0.00 0.84 ouinhi 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.00 za-kpota 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.42 zagnanado 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.42 zogbodomey 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.84 biodiversity informatics, 11, 2016, pp. 23-39 30 added the coarse-resolution grid identification codes to each occurrence datum, and aggregated each datum into the coarse-resolution grid squares. in excel, we explored relations between data on species identity, time (i.e, taking days as a unit of sampling effort), and aggregation grid square. we calculated (1) total number of records available from each grid square (termed n), (2) total number of distinct species recorded from each grid square (sobs), (3) number of species detected on exactly one day (a), and (4) number of species detected on exactly two days (b). via equations provided by chao (1987), we calculated the expected number of species (sexp), as 𝑆!"# = 𝑆!"# + !! !! , and inventory completeness (c) as c = sobs / sexp. we explored plots of c versus n to establish appropriate and adequate definitions of relatively completely versus incompletely inventoried grid squares. once we had established criteria under which grid squares would be considered as wellsampled, in qgis, we linked the table with the grid square statistics (i.e., n, sobs, sexp, c) to the aggregation grid, and saved this file as a shapefile. we created a shapefile of well-sampled grid squares, which we in turn converted to raster (geotiff) format using custom scripts in r. this raster coverage was the basis for our identification of gaps in coverage of geographic and environmental spaces, as follows. we used the proximity (raster distance) function in qgis to summarize geographic distance across the country to any well-sampled area, at 0.05° spatial resolution. to create a parallel view of environmental (climatic) difference from well-sampled areas, we used a detailed protocol, as follows. we plotted 5000 random points across the country, and linked each point to the geographic distance raster and to raster coverages (2.5’ spatial resolution) summarizing annual mean temperature and annual precipitation drawn from the worldclim climate data archive (hijmans et al., 2005) via the point sampling tool in qgis. we exported the attributes table associated with the random points, and imported it into excel. we standardized values of each environmental variable to the overall range of the variable among the random points as (xi – xmin) / (xmax xmin), where xi is the particular observed value in question. we then created a matrix of euclidean distances in the two-dimensional climate space, relating all of the points with a geographic distance >0 to all of the points with geographic distance of zero; the latter represent points falling in well-sampled regions, whereas the former are scattered across the entire country. points in well-sampled regions were assigned (by definition) environmental distances of zero. finally, environmental distances were imported into qgis, and linked back to the random point’s shapefile. this shapefile provided a broad sampling across the country, with a zvalue that is the environmental distance associated with that point. to convert this vectorformat dataset to raster format, with values across the entire region, we used inverse distance weighting, with a distance coefficient of 2.0. we explored these results further via relating them to benin’s municipalities to provide local contexts for future inventories. we further explored impacts of roads, waterways, and protected areas on data completeness of the country. this exploration was achieved by intersecting completeness (c) at 0.5° spatial resolution with data layers summarizing distributions of roads, waterways, and protected areas across benin. these data layers were downloaded from open sources for roads, waterways7, and protected areas8. the attributes tables of intersection layers were used to calculate respectively the length of roads and waterways as well as the surface of protected areas associated to each grid cell. we could then calculate the coefficients of correlation between c values and each of the parameters calculated. finally, we explored, in a preliminary manner, the impacts of absence of roads, waterways, and protected areas on data completeness. results initial numbers of plant records downloaded from gbif comprised 148,944 primary occurrence records. after cleaning and removing duplicates and records of exotic species, we were left with 84,350 records (56.6% of the original data set) corresponding to 3188 species (figure 1). of the original total of records, 87.6% were identified to the species level, and 80.4% had adequate geographic coordinates. after checking coordinates and displaying the records against the administrative limits of benin (figure 2), only 74.3% of records were located within benin’s 7 http://www.diva-gis.org/gdata. 8 http://www.protectedplanet.net/country/bj#. biodiversity informatics, 11, 2016, pp. 23-39 31 figure 3. floristic composition of dak for plants of benin in terms of representation of plant classes. figure 4. inventory completeness (c values) as a function of numbers of records available within 0.5° grid squares across benin. 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 0 1000 2000 3000 4000 5000 6000 7000 c va lu es number of records biodiversity informatics, 11, 2016, pp. 23-39 32 figure 5. geographic patterns of inventory completeness across benin (a) 0.5° spatial resolution; (b) geographic distance to the most well-known cells (n ≥ 200 and c ≥ 0.80); (c) environmental distance to the most well-known cells. (b) (a) (b) (c) (a) biodiversity informatics, 11, 2016, pp. 23-39 33 borders and 62.1% had full date information. most species belonged to the classes magnolipsida (85.2% of records) and liliopsida (14.2%; figure 3). the occurrence data corresponded to species in 188 families, principally fabaceae (25.1% of records), combretaceae (9.8%), rubiaceae (6.6%), sapotaceae (6.3%), and poaceae (5.1%). we inspected the relationship between c and numbers of records (n) at 0.5° spatial resolution (figure 4). the correlation was weak and not significant (r = 0.203, p = 0.157). to account for artifactual c values, (grid cells with high c values (i.e., >0.7) for few record numbers), and taking into account the relationship between c values and n (figure 4), we considered grid cells as well-sampled when n ≥ 200 records and c ≥ 0.5. the analysis of inventory completeness across benin using the well-sampled criteria set above (n ≥ 200 records and c ≥ 0.5) showed that only two grid cells in northern west benin are not well known (figure 5a). to be more consistent with the status and reality of data completeness of the country, we considered grid cells as bestknown when n ≥ 200 records and c ≥ 0.8. using the latter criteria we achieved the distribution patterns of inventory completeness at 0.5° spatial resolution with the corresponding distribution of geographic and environmental distances (figure 5). the analysis of inventory completeness at the level of departments shows that most departments of northern benin (alibori, borgou, atakora) had the largest percentages of grid cells with no data or not well-sampled (5-11%), whereas the best-known departments (c ≥ 0.8) were in southern and central benin (table 1, figure 5). analysis of inventory completeness at the level of municipalities (a lower-level administrative subdivision nested within departments) showed that the municipalities of banikoara, gogounou, karimama, kandi, mallanville, and segbana (department of alibori), and kalalé and bembérékè (department of borgou) had the most grid cells with no data or not well-sampled (table 2, figure 5). those municipalities thus represent priorities for future inventory efforts to fill the picture of inventory completeness across benin. the best-known municipalities (c ≥ 0.8) were djougou (department of donga, northern benin), tori-bossito (atlantique, southern benin), glazoué (collines, central benin), and covè (zou, central benin; table 2, figure 5). the intersection of inventory completeness with roads, waterways, and protected areas suggested positive impacts of all three on plant data completeness of benin (table 3, figure 6). among grid cells holding roads, only 28.5% had no data or were not well-sampled, compared with 50% across all grid cells; presence of roads in grid cells also increased the percentage of wellknown grids cells (34.8%) against only 22.4% across the country (table 3, figure 6a). however, the correlation between road length within grid cells and c values was weak and not significant (r = -0.027; p = 0.501). the intersection of the waterways data layer with grid cells revealed that 35.8% of the latter had no data or are not well-sampled against 50% for the whole grid cells across the country; the presence of waterways in grid cells also increased the percentage of best-known grids cells (36.8%) against only 22.4% across the country (table 3, figure 6b). absence of waterways also coincided with grid cells lacking data, although the correlation between the length of waterways and c values is weak and not significant (r = -0.033; p = 0.342). however, grid cells containing waterways are also those containing roads, so the relative impacts of the two factors remain unclear. the intersection of protected areas layer with the grid cells revealed that 35.0% of the latter had no data or were not well-sampled, compared with 50% for the whole grid cells across the country. presence of protected areas in grid cells also increased the percentage of well-known grids cells (39.4%) against only 22.4% across the country (table 3, figure 6c). however, the correlation between protected area and c was significant (r = -0.253; p = 0.035), but was negative in sign. again, grid cells containing protected areas also held roads and waterways, so the relative impacts of the three factors are difficult to evaluate. discussion completeness at different spatial levels we considered a grid cell as well-sampled if it was represented by ≥200 records and if completeness was ≥0.5. although this threshold is low compared to the 1000 records per grid cell, at the same spatial resolution, used by sousabaena et al. (2013) for analyses of plant communities in brazil, our threshold is higher than that of funk et al. (2005), who used 40 records per 50 x 50 km grid cell in their analyses across guyana. the additional criterion of c ≥ biodiversity informatics, 11, 2016, pp. 23-39 34 figure 6. geographic patterns of inventory completeness across benin at 0.5° spatial resolution overlaid on spatial patterns of (a) roads, (b) waterways, and (c) protected areas. (a) (b) (c) c values biodiversity informatics, 11, 2016, pp. 23-39 35 table 3: impacts of protected areas, waterways, and roads on data completeness in benin type of layer information percentage of grid cells or objects with no data (a) percentage of grid cells or objects not well sampled with n < 200 and c < 0.5 (b) percentage of grid cells or objects not well sampled with n ≥ 200 and c < 0.5 (c) total (a) + (b) + (c) percentage of grid cells or objects well sampled with 0.5 ≤ c < 0.6 (d) percentage of grid cells or objects well sampled with 0.6 ≤ c < 0.7 (e) percentage of grid cells or objects well sampled with 0.7 ≤ c < 0.8 (f) percentage of grid cells or objects well sampled with c ≥ 0.8 (g) half degree grid cells 12.07 37.93 0.00 50.00 1.72 5.17 20.69 22.41 half degree grid cells inter protected areas 1.41 33.80 0.00 35.21 0.00 1.41 23.94 39.44 half degree grid cells inter waterways 2.30 33.49 0.00 35.79 0.00 3.99 23.46 36.76 half degree grid cells inter roads 1.77 26.77 0.00 28.54 0.00 5.65 30.97 34.84 biodiversity informatics, 11, 2016, pp. 23-39 36 0.5 for a grid cell to be considered as well inventoried coincides with the threshold used by sousa-baena et al. (2013) in their study of inventory completeness of plants of brazil. the analysis of inventory completeness for different departments and municipalities revealed that most departments of northern benin, and— more precisely—the municipalities of banikoara, gogounou, karimama, kandi, mallanville, and segbana in the department of alibori, and the municipalities of kalalé and bembérékè in the department of borgou, were poorly known. these gaps stand in contrast to southern and central benin, where municipalities were betterknown (n ≥ 200, c ≥ 0.8). the municipalities of northern benin listed above are therefore of highest priority for future inventory work in the country. lack of infrastructure and institutions in the past in that region could explain our results in northern benin, as well as insufficiency of publishing of occurrence data from the studies of vegetation in that part of the country. indeed, in the region, the university of parakou was created only in 2001, and its faculty of agronomy has initiated biodiversity studies only since that time. exploring inventory completeness at 0.5° spatial resolution, we observed a positive impact of roads on data completeness. these results coincide with those of kadmon et al. (2004), who affirmed that bias in distributional data is common in the form of high concentrations of collection sites along roads. souza-baena et al. (2013) also found road bias effects on plant data completeness across brazil, and hijmans et al. (2000) reported that most gene bank accessions for wild potatoes in bolivia were collected within 2 km of roads, 3-fold greater than random expectations. ballesteros-mejia et al. (2013) reported positive effects of infrastructure (traffic access, road density, tourism) on inventory completeness for tropical insects of sub-saharan africa. analyses of distribution of dak with respect to waterways also showed a positive impact of the latter on data completeness. our further exploration revealed that the grid cells concerned here are served both by waterways and roads so that the positive impact of waterways on data completeness could be induced by the presence of roads instead. considering geographic patterns of inventory completeness in relation to protected areas, we found a significant impact of the latter on data completeness. the correlation of area protected and c was significant and negative, suggesting that data completeness was more effective in protected areas with smaller surfaces. we inferred that efforts of data inventories were more and more diluted when the surface of protected areas increases. our results support those of ballesteros-mejia et al. (2013), who reported positive effects of protected areas status on inventory completeness of sphinx moths (sphingidae) of sub-saharan africa. dak fitness for use after cleaning, our data set corresponded to 3188 species. the total species richness of plant species of benin has been estimated at 3200 species (adjanohoun et al., 1989), but the most complete and recent flora of benin included only 2807 species (akoegninou et al., 2006). our result can therefore complement the list of the flora of benin of akoegninou et al. (2006), via identification of several hundred species that have been recorded within the country. data fitness for use revolves around data precision, accuracy, and authenticity for specific uses (faith et al. 2013). as such, an important initial observation is that >40% of initial records downloaded from gbif proved unavailable or incomplete, and thus were not included in our analyses. this amount is considerable in terms of loss of data—we emphasize that the data records lost were already dak, and yet were not available for analysis owing to gaps or inconsistencies in content in the form of incomplete taxonomic determination, inadequate geographic coordinates, and missing time information. this result confirmed data concerns raised in feedback received by gbif from its community of data users: the recent (2010) survey by gbif’s content needs assessment task group (cnatg) revealed major concerns in terms of geographic and taxonomic gaps in data coverage, as well as the need for data quality assurance (faith et al. 2013). for benin, 12.4% of records lacked usable species names, higher than the 9.5% found across the broader gbif network (gaiji et al., 2013), although we perhaps used a stricter set of criteria for our filtering. records lacking coordinates in benin (19.6%) were comparable to the 18.5% found by cnatg as of december 2010, but higher than the 14.1% reported in february 2012 (gaiji et al. 2013). the percentage of records falling outside of benin’s borders (7.5% of georeferenced records) was high compared to the 3.6% reported by gaiji et al. (2013). numbers of records lacking full temporal information for biodiversity informatics, 11, 2016, pp. 23-39 37 benin (37.9%) was high compared to the 23.1% found by sousa-baena et al. (2013) for plant data in brazil, and the 30.8% found by gaiji et al. (2013) for gbif-mediated data globally. to improve dak quality for the biodiversity of benin, rigorous data capture protocols and detailed error detection and data cleaning workflows need to be implemented. these steps may depend in large part on sound capacitybuilding, such that data managers and publishers of data relevant to benin can work more effectively. for instance, error rates in taxonomic names can be reduced massively by using authority data held in authority lists like the catalogue of life and prota as controlled vocabularies. geographic coordinates can be resolved using tools like geolocate and biogeomancer to add coordinates to welldescribed localities where data have been collected (gaiji et al., 2013). acknowledgments the jrs biodiversity foundation has generously supported biodiversity informatics activities in benin. we address our sincere gratitude to town peterson of the university of kansas, who facilitated writing this paper through training in a course on national biodiversity diagnoses, and revised the manuscript profoundly. we also thank kate ingenloff and lindsay campbell for their assistance in analyses. references adjanohoun, e. j., v. adjakidjè, m. r. a. ahyi, l. ake assi, a akoegninou, j. d’almeida, f. apovo, k. boukef, m. chadare, g. cusset, k. dramane, j. eyme, j.-n. gassita, n. gbaguidi, e. goudoté, s. guinko, p. houngnon, l. issa, a. keita, h. v. kiniffo, d. koné-bamba, a. musampa nseyya, m. saadou, t. sodogandji, s. de souza, a. tchabi, c. zinsou dossa and t. zohoun. 1989. contribution aux études éthnobotaniques et floristiques en république populaire du bénin. agence de coopération culturelle et technique (acct), paris, france. adomou, a. c., 2005. vegetation patterns and environmental gradients in benin: implications for biogeography and conservation. ph.d. thesis, wageningen university, wageningen. aoudji, a. k. n., i. s. a. yêvidé, j. c. ganglo, g. atindogbé, s. m. toyi, c. de canniere, h. a. azontondé, v. adjakidjè, b. de foucault and a. b. sinsin. 2011. structural characteristics and forest sites identification in pahou forest reserve, south-benin. bois et forêts des tropiques 308:47-58. akoegninou, a., w. j. van der burg, l., j., g., van der maesen, v. adjakidjè, j. p. essou, b. sinsin and h. yédomonhan. 2006. flore analyti que du bénin. backhuys publishers. wageningen. ariño, a. h. 2010. approaches to estimating the universe of natural history collections data. biodiversity informatics 7: 81-92. awokou, k. s., j. c. ganglo, h. a. azontondé, v. adjakidjè and b. de foucault. 2009. caractéristiques structurales et écologiques des phytocénoses forestières de la forêt classée d’itchèdè (département du plateau, sud-est bénin). science & nature 6: 125-138. ballesteros-mejia, l., i. j. kitching, w. jetz, p. nagel and j. beck. 2013. mapping the biodiversity of tropical insects: species richness and inventory completeness of african sphingid moths. global ecology and biogeography 22: 586-595. chao, a. 1987. estimating the population size for capture recapture data with unequal catchability. statistica sinica 10: 227-246. chapman, a. d., m. e. s. muñoz and i. koch. 2005. environmental information: placing biodiversity phenomena in an ecological and environmental context. biodiversity informatics 2: 24-41. collen, b., m. ram, t. zamin and l. mcrae. 2008. the tropical biodiversity data gap: addressing disparity in global monitoring. tropical conservation science 1: 75-88. dansi, a., r. vodouhè, p. azokpota, h. yedomonhan, p. assogba, a. adjatin, y. l. loko, i. dossouaminon and k. akpagana. 2012. diversity of the neglected and underutilized crop species of importance in benin. scientific world journal 2012: 932947. faith, d. p., b. collen, a. h. ariño, p. koleff, j. guinotte, j. kerr and v. chavan. 2013. bridging biodiversity data gaps: recommendations to meet users’ data needs. biodiversity informatics 8: 4158. funk, v. a., k. s. richardon and s. ferrier. 2005. survey-gap analysis in expeditionary research: where do we go from here? biological journal of the linnaean society 85: 549-567. gaiji, s., v. chavan, d. hobern, a. h. ariño, r. sood, j. otegui and e. robles. 2013. content assessment of the primary biodiversity data published through gbif network: status, challenges and potentials. biodiversity informatics 8: 94-172. ganglo, c. j., j. lejoly and t. pipar. 1999. le teck (tectona grandis l. f.) au bénin, gestion et perspectives. bois et forêts des tropiques 261: 17-27. ganglo, c. j. and j. lejoly.1999. biotope et valeur indicatrice écologique de l’association à lecaniodiscus cupanioides et landolphia calabarica dans le sous-bois naturel des teckeraies du sud-bénin. acta botanica gallica 146: 227-245. biodiversity informatics, 11, 2016, pp. 23-39 38 ganglo, c. j. 2004. phytosociologie appliquée à l’aménagement des forêts: cas du périmètre forestier de toffo (sud-bénin, département de l’atlantique). rapport scientifique, université d’abomey-calavi, cotonou. ganglo j. c. 2005. groupements de sous-bois, identification et caractérisation des stations forestières: cas d’un bois au bénin. bois et forêts des tropiques 285: 35-46. ganglo, j. c. and b. de foucault. 2006. plant communities, forest site identification and classification in toffo reserve, south-benin. bois et forêts des tropiques 288: 25-38. gbif-benin. 2011. enquête sur les détenteurs et les utilisateurs de données et d’informations sur la biodiversité au bénin. gbif bénin, cotonou. gouwakinnou, n. g., kindomihou, v., assogbadjo, e. a. and b. sinsin. 2009. population structure and abundance of sclerocarya birrea (a. rich) hochst subsp. birrea in two contrasting land-use systems in benin. international journal of biodiversity and conservation 1: 194-201. johnson, n. f. 2007. biodiversity informatics. annual review of entomology 52: 421-38 hijmans, r. j., s. e. cameron, j. l. parra, p. g. jones and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. international journal of climatology 25: 1965-1978. hijmans, r. j., k. a. garrett, z. huaman, d. p. zhang, m. schreuder and m. bonierbale. 2000. assessing the geographic representativeness of gene bank collections: the case of bolivian wild potatoes. conservation biology 14: 1755-1765. idohou, r., a. h. ariño, a. e. assogbadjo, r. l. glèlè kakai and b. sinsin. 2015. diversity of wild palms (arecaceae) in the republic of benin: finding gaps in the national inventory combining field and digital accessible knowledge. biodiversity informatics 10: 45-55. kadmon, r. and a. danin. 2003. a systematic analysis of factors affecting the performance of climatic envelope models. ecological applications 13: 853-867. kadmon, r., o. farber and a. danin. 2004. effect of roadside bias on the accuracy of predictive maps produced by bioclimatic models. ecological applications 14: 401-413. koura, k., j. c. ganglo, a. e. assogbadjo and c. agbangla. 2011. ethnic differences in use values and use patterns of parkia biglobosa in northern benin. journal of ethnobiology and ethnomedicine 7: 42. koura, k., e. f. dissou and j. c. ganglo. 2013a. caractérisation écologique et structurale des parcs à néré [parkia biglobosa r. br.] du département de la donga au nord-ouest du bénin. international journal of biological and chemical sciences 7: 726-738. koura, k., y. mbaidé and j. c. ganglo. 2013b. caractéristiques phénotypique et structurale de la population de parkia biglobosa r. br. du nordbénin. international journal of biological and chemical sciences.7: 2317-2327 kremen, c., j. dransfield, jr. e. louis, m. vences, a. cameron, b. l. fisher, r. a. nussbaum, d. r.vieites, a. moilanen, f. glaw, t. c. good, c. j. raxworthy, p.c. wright, s. j. phillips, g. j. harper, c. d. thomas, r. j. hijmans, a. razafimpahanana, m. l. zjhra, h. beentje, d. c. lees and g. e schatz. 2008. aligning conservation priorities across taxa in madagascar with high-resolution planning tools. science 320: 222-226. kyndt, t., e. a. assogbadjo, j. olivier, j. o. hardy, r. glèlè kakaï, b. sinsin, p. van damme and g. gheysen. 2009. spatial genetic structuring of baobab (adansonia digitata, malvaceae) in the traditional agroforestry systems of west africa. american journal of botany 96: 950-957. martínez-meyer, e. 2005. climate change and biodiversity: some considerations in forecasting shifts in species’ potential distributions. biodiversity informatics 2: 42-55 morris, r. a., v. barve, m. carausu, v. chavan, j. cuadra, c. freeland, g. hagedorn, p. leary, d. mozzherin, a. olson, g. riccardi, i. teage and g. whitbread. 2013. discovery and publishing of primary biodiversity data associated with multimedia resources: the audubon core strategies and approaches. biodiversity informatics 8: 185-197. noumon, j. c. and j. c. ganglo. 2005. phytosociologie appliquée à l'aménagement des forêts: cas du périmètre forestier de koto (département du zou, centre-bénin). acta botanica gallica 152: 421-426. noumon, j. c., j. c. ganglo, h. a. azontondé, b. de foucault and v. adjakidjè. 2009. ecological and silvicultural indicatory value of plantcommunities of koto forest reserve (centrebenin). international journal of biological and chemical sciences 3: 367-377. peterson, a.t., s. l. egbert, v. sánchez-cordero and k. p. price. 2000. geographic analysis of conservation priorities using distributional modeling and complementarity: endemic birds and mammals in veracruz, mexico. biological conservation 93:85-94. peterson, a. t. 2001. predicting species’ geographic distributions based on ecological niche modeling. condor 103: 599-605. peterson, a. t., m. a. ortega-huerta, j. bartley, v. sanchez-cordero, j. soberón, r. h. buddemeier and d. r. b. stockwell. 2002. future projections for mexican faunas under global climate change scenarios. nature 416: 626-629. peterson, a.t., e. martínez-meyer, c. gonzálezsalazar and p. w. hall. 2004. modeled climate biodiversity informatics, 11, 2016, pp. 23-39 39 change effects on distributions of canadian butterfly species. canadian journal of zoology 82: 851-858. peterson, a. t., j. soberón and l. krishtalka. 2015. a global perspective on decadal challenges and priorities in biodiversity informatics. bmc ecology 15: 15. soberón, j., r. jiménez, j. golubov and p. koleff. 2007. assessing completeness of biodiversity databases at different spatial scales. ecography30: 152-160. soberón, j. and a. t. peterson. 2009. monitoring biodiversity loss with primary species-occurrence data: toward national-level indicators for the 2010 target of the convention on biological diversity. ambio 38: 29-34. sokpon, n., t. sinadouwirou, f. gbaguidi and h. s. biaou. 2001. les forêts édaphiques hygrophiles du bénin. belgian journal of botany 134: 79-93. sokpon, n. and s. h. biaou. 2002. the use of diameter distributions in sustained-use management of remnant forests in benin: case of bassila forest reserve in north benin. forest ecology and management 161: 13-25. sokpon, n., s. h. biaou, c. ouinsavi and o. hunhyet. 2006. bases techniques pour une gestion durable des forêts claires du nord-bénin: rotation, diamètre minimal d'exploitabilité et régénération. bois et forêts des tropiques 287: 45-57. sousa-baena, m. s., l. c. garcia and a. t. peterson. 2013. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. diversity and distributions 20: 369-381. tohngodo, c., j. c. ganglo, k. e. agbossou, b. de foucault and v. adjakidjè. 2006. caractéristiques structurales, identification et caractérisation des stations forestières de la forêt classée de bonou (sud-est bénin). sciences & nature 3: 39-47. yêhouénou tessi, r. d., s. g. akouèhou and j. c. ganglo. 2012. caractéristiques structurales et écologiques des populations de antiaris toxicaria (pers.) lesch et de ceiba pentandra (l.) gaertn dans les forêts reliques du sud-benin. international journal of biological and chemical sciences 6: 5056-5067. yêvidé, i. s. a., j. c. ganglo, k.a. aoudji, s. m. toyi, c. de cannière, b. de foucault, j.-l. devineau and b. sinsin. 2011. caractéristiques structurelles et écologiques des phytocénoses de sous-bois des plantations privées de teck du département de l’atlantique (sud-bénin, afrique de l’ouest). acta botanica gallica 158: 263-283. biodiversity  informatics,  10,  2015,  pp.  1-­‐21   ecoclimate: a database of climate data from multiple models for past, present, and future for macroecologists and biogeographers matheus s. lima-ribeiro1, sara varela2,3, javier gonzález-hernández4, guilherme de oliveira5, josé alexandre f. diniz-filho6, levi carina terribile1 1laboratório de macroecologia, universidade federal de goiás, regional jataí, cx. postal 03, 75804-020, jataí, go, brazil 2departamento de ciencias de la vida, edificio de ciencias, campus externo, universidad de alcalá, 28805 alcalá de henares, madrid, spain 3museum für naturkunde. leibniz institute for evolution and biodiversity science. invalidenstr. 43, 10115 berlin, germany 4carmelitas descalzos, 5. 45002, toledo, spain 5laboratório de biogeografia da conservação e comportamento animal, universidade federal do recôncavo da bahia (ufrb), centro de ciências agrárias, ambientais e biológicas (ccaab), setor de biologia, rua rui barbosa 710, centro 44380-000, cruz das almas, ba, brazil 6departamento de ecologia, icb, universidade federal de goiás, cx. postal 131, 74001-970, goiânia, go, brazil abstract.—studies in biogeography and macroecology have been increasing massively since climate and biodiversity databases became easily accessible. climate simulations for past, present, and future have enabled macroecologists and biogeographers to combine data on species’ occurrences with detailed information on climatic conditions through time to predict biological responses across large spatial and temporal scales. here we present and describe ecoclimate, a free and open data repository developed to serve useful climate data to macroecologists and biogeographers. ecoclimate arose from the need for climate layers with which to build ecological niche models and test macroecological and biogeographic hypotheses in the past, present, and future. ecoclimate offers a suite of processed, multi-temporal climate data sets from the most recent multi-model ensembles developed by the coupled modeling intercomparison projects (cmip5) and paleoclimate modeling intercomparison projects (pmip3) across past, present, and future time frames, at global extents and 0.5° spatial resolution, in convenient formats for analysis and manipulation. a priority of ecoclimate is consistency across these diverse data, but retaining information on uncertainties among model predictions. the ecoclimate research group intends to maintain the web repository updated continuously as new model outputs become available, as well as software that makes our workflows broadly accessible. key words.—climate data, raster, bioclimatic variables, general circulation models, kriging, downscaling, pleistocene, pliocene, holocene. the availability of spatially explicit data layers summarizing past, present, and future climate conditions has stimulated the fields of biogeography and macroecology greatly in the last two decades. for instance, the pioneering worldclim repository1 (hijmans et al., 2005) enabled researchers to integrate data on species’ geographic occurrences with detailed information on climate conditions through time to predict biological responses across large spatial and temporal scales. in tandem with the climate data, access to vast data                                                                                                                 1 http://www.worldclim.org. resources about biodiversity (e.g., gbif2, paleobiology database3, specieslink4), and exciting new computational tools (e.g., the new r packages; rgbif: chamberlain et al., 2013; ravis: varela et al., 2014a; paleobiodb: varela et al., 2014b), have facilitated fundamental analyses by macroecologists and biogeographers on broad scales. the multi-temporal climate data have been used to explore effects of past (varela et al., 2015a) and future (thomas et al., 2004)                                                                                                                 2 http://www.gbif.org. 3 http://www.paleobiodb.org. 4 http://splink.cria.org.br. townpeterson typewritten text townpeterson typewritten text 1 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   climate change; to understand past extinction events (lima-ribeiro et al., 2013a); and to test hypotheses concerning population dynamics (barrientos et al., 2014), evolutionary processes (araújo et al., 2013; saupe et al., 2014), and ecological dynamics (martínez-meyer et al., 2004; martínez-meyer and peterson, 2006). data layers summarizing climatic information are created by interpolating continuous surfaces from real (generally point-based) observations, or by modeling conditions based on complex simulations describing key processes of atmospheric and ocean circulation, the so-called general circulation models (gcms; braconnot et al., 2007). for instance, new et al. (2002), hijmans et al. (2005), and kriticos et al. (2012) provide interpolated data layers for modern climates based on observations from almost 50,000 weather stations worldwide; ~100 fossil pollen records from the last glacial maximum (lgm 21 ka, bartlein et al., 2011; harrison et al., 2014) and midpliocene (dowsett et al., 2012) made possible building layers for past continental climates. however, a dearth of detailed fossil data linking time and space has hindered building reliable data layers from paleobiological observations. obviously, future climatic conditions are not accessible via observation. consequently, climatologists have invested in gcms to simulate global climate over long time periods, and connect climates of the deep geological past through the present to future conditions. climatologists tune gcms based on boundary conditions such as orbital parameters, solar forcing, greenhouse gas concentrations, co2 emissions, land-use, and ice coverage, coupled with atmospheric, vegetation, and ocean dynamics (braconnot et al., 2007). via these intensive computer simulations, climatologists have predicted global climates for past (miocene, pliocene, late quaternary), present (preand postindustrial), and future conditions (end of 21st century; taylor et al., 2012). still, climate model outputs are not at all user-friendly for most macroecologists and biogeographers. outputs are normally formatted as complex text files (e.g., netcdf format), which are not trivial to process; even to understand the acronyms used for identifying the dozens of climatic variables and model parameters can be challenging. macroecologists and biogeographers are used to working with data layers in the form of raster images or simple ascii text files. spatial resolution also differs among gcms, precluding direct incorporation into spatial analyses without complex downscaling procedures. given, then, the often complex and inaccessible nature of gcm outputs, and yet the great interest in multi-temporal climate data and their enormous applicability to important questions in macroecology and biogeography, we decided to process climate layers from multiple gcms, and make them available on a free and open web repository: ecoclimate5. ecoclimate offers a wide suite of climate data layers from the most recent multi-model ensembles published as part of the coupled model inter-comparison project (cmip5; taylor et al., 2012) and the paleoclimate modeling intercomparison project (pmip3) across past, present and future time frames, at global extents and 0.5° spatial resolution. data are provided in formats easily incorporated in analyses via common platforms and popular gis software. methods raw climatic variables we accessed climatic simulations from most recent generations of coupled atmosphere-ocean general circulation models (aogcms) available in the cmip5 and pmip3 databases (table 1). the aogcms comprised multi-model ensembles for longterm experiments: past ("plioexp2a", "lgm", and "midholocene" experiments), present ("picontrol" and "historical" experiments) and future scenarios ("rcps" experiments for different co2 emission scenarios), although some specific outputs are not yet available from some modeling groups (table 2). aogcms run long-term simulations for preindustrial scenario (~1760), a control experiment (picontrol) for stabilizing climate predictions and model evaluation, and then simulate climates for different time slices, according to specific boundary conditions. besides pre-industrial, current conditions are also simulated for historical periods, providing climatic conditions for the 20th century (indeed, the industrial period from 1850 to 2005). future conditions are                                                                                                                 5 http://www.ecoclimate.org/. townpeterson typewritten text 2   ta bl e 1. d et ai ls o f th e co up le d a tm os ph er eo ce an g en er al c irc ul at io n m od el s (a o g c m s) a va ila bl e in t he e co c lim at e da ta ba se . a ll si m ul at io ns w er e ob ta in ed f ro m t he r 1i 1p 1 en se m bl e, e xc ep t g is s (r 1i 1p 15 1) . t he n at iv e a o g c m -s pe ci fic r es ol ut io ns , i n de ci m al d eg re es (lo ng itu de a nd la tit ud e) , a re c oa rs e (1 -4 º). m os t m od el s w er e ru n ov er ti m e se rie s of 1 00 y r a fte r i ni tia l c al ib ra tio n pe rio ds . s ou rc e: c m ip 5, th e fif th p ha se o f c ou pl ed m od el in te rc om pa ris on p ro je ct 12 a nd p m ip 3, th e th ird p ha se o f p al eo cl im at e m od el in g in te rc om pa ris on p ro je ct 13 .                                                                                                                 12 h ttp :// cm ip -p cm di .ll nl .g ov /c m ip 5/ . 13 h ttp :// pm ip 3. ls ce .ip sl .fr /. m od el id m od el in g c en te r r es ol ut io n n um be r o f ye ar s so ur ce r el ea se c c sm 4 n at io na l c en te r f or a tm os ph er ic r es ea rc h, u sa 0. 9° × 1 .2 5° 10 0 c m ip 5/ pm ip 3 20 12 c n r m -c m 5 c en tre n at io na l d e r ec he rc he s m et eo ro lo gi qu es / c en tre e ur op ee n de r ec he rc he e t f or m at io n a va nc ee s e n c al cu l s ci en tif iq ue , f ra nc e 1. 4° x 1 .4 ° 20 0 c m ip 5/ pm ip 3 20 12 c o sm o sa so (f u b ) fr ei e u ni ve rs itä t b er lin , g er m an y 3. 75 ° x 3. 7° 60 0 pm ip 3 20 12 g is se2 -r n a sa g od da rd in st itu te fo r s pa ce s tu di es , u sa 2. 5° x 2 .0 ° 10 0 c m ip 5/ pm ip 3 20 12 fg o a ls -g 2 n at io na l k ey l ab or at or y of n um er ic al m od el in g fo r a tm os ph er ic sc ie nc es a nd g eo ph ys ic al f lu id d yn am ic s ( la sg ) / in st itu te o f a tm os ph er ic p hy si cs (i a p) , c hi na 2. 8° × 2 .8 ° 10 0 c m ip 5/ pm ip 3 20 13 ip sl -c m 5a lr in st itu t p ie rr e si m on l ap la ce , f ra nc e 3. 75 ° x 1. 9° 20 0 c m ip 5/ pm ip 3 20 12 m ir o c -e sm a tm os ph er e an d o ce an r es ea rc h in st itu te (u ni ve rs ity o f t ok yo ), n at io na l i ns tit ut e fo r e nv iro nm en ta l s tu di es , a nd ja pa n a ge nc y fo r m ar in eea rth s ci en ce a nd t ec hn ol og y, ja pa n 2. 8° × 2 .8 ° 10 0 c m ip 5/ pm ip 3 20 12 m pi -e sm -p m ax p la nc k in st itu te fo r m et eo ro lo gy , g er m an y 1. 9° x 1 .9 ° 10 0 c m ip 5/ pm ip 3 20 11 m r ic g c m 3 m et eo ro lo gi ca l r es ea rc h in st itu te , j ap an 1. 1° x 1 .1 ° 10 0 c m ip 5/ pm ip 3 20 12 townpeterson typewritten text 3   table 2. ecoclimate layers processed as regards availability of outputs (! indicates available, ✗ indicates not available) from the cmip5 and pmip3 working groups. experiment acronyms: plio: mid-pliocene warm period (mpwp, ~3.3 to 3.0 million years ago); lgm: last glacial maximum (21,000 years ago); hol: mid-holocene (6000 years ago); picontrol: pre-industrial (~1760); historical (1900-1949); modern (1950-1999); rcps: representative carbon pathways, with emission scenarios for the end of the 21st century (2080-2100). aogcm past present future plio lgm hol picon. histor. modern rcp2.6 rcp4.5 rcp6.0 rcp8.5 ccsm ! ! ! ! ! ! ! ! ! ! cnrm ✗ ! ! ! ! ! ✗ ! ✗ ! cosmos ✗ ! ✗ ! ✗ ✗ ✗ ✗ ✗ ✗ fgoals ✗ ! ! ! ! ! ! ✗ ✗ ! giss ✗ ! ✗ ! ! ! ! ! ! ! ipsl ✗ ! ! ! ! ! ! ! ! ! miroc ✗ ! ! ! ! ! ! ! ! ! mpi ✗ ! ! ! ! ! ✗ ✗ ✗ ✗ mri ✗ ! ! ! ! ! ! ! ! ! table 3. spatial correlation (pearson's coefficient, r) among downscaled temperature (above diagonal) and precipitation (below diagonal) layers from distinct techniques. krige: ordinary kriging; idw: inverse distance weighting; splines: thin-plate spline; trend: trend surface with 12th-order polynomial regression. krige idw splines trend natural neighbor krige 0.99 0.99 0.99 0.99 idw 0.98 0.99 0.99 0.99 splines 0.99 0.97 0.99 1.00 trend 0.84 0.90 0.83 0.99 natural neighbor 0.99 0.97 0.99 0.83 figure 1. mean square errors (mse) among downscaling techniques for (a) temperature and (b) precipitation layers. note that lowest mses come from ordinary kriging method. krige: ordinary kriging; idw: inverse distance weighting; splines: thin-plate spline; trend: trend surface with 12th-order polynomial regression. townpeterson typewritten text 4 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   simulated in sequence, out to 2100, following different representative concentration pathways (rcps) across 21st century (rcp2.6, rcp4.5, rcp6.0, and rcp8.5 experiments). finally, experiments for past scenarios cover the lgm (bartlein et al., 2011) and midholocene (6 ka), key periods representing glacial and interglacial phases related to the last ice age, as well as the mid-pliocene warm period (mpwp, ~3.3-3.0 ma). from such aogcms (table 2), we downloaded 4 simulated monthly atmospheric variables: precipitation flux (pr), mean surface temperature (tas), maximum surface temperature (tasmax), and minimum surface temperature (tasmin). all outputs match ensemble member "r1i1p1" (except giss, which was r1i1p151), assuring compatible outputs among aogcms. because aogcms run long-term simulations, we averaged simulated monthly values from the entire native time-series for pre-industrial and past experiments (table 1) to guarantee reliable long-term means. for the historical experiments, we averaged climate predictions for 1900-1949 (hereafter the “historical” period) and 1950-1999 (hereafter the “modern” period). future predictions were averaged between 2080 and 2100, thus representing conditions for the end of the 21st century. temperature variables were transformed from kelvin to celsius, and precipitation flux (in mm m-2 s-1) was converted to total monthly precipitation (mm month-1), taking into account a month with 30 days according the specific calendar of 360 days year-1. the original netcdf files with raw aogcm outputs were processed using the ncdf package in r (pierce, 2014). scripts are available openly6. statistical downscaling: regridding aogcmspecific native outputs the long-term means for temperature and precipitation variables were downscaled to 0.5° resolution. our downscaling was actually a regridding procedure instead of a standard interpolation (implications discussed below). standard interpolations are commonly used for observed climatology to increase resolution, but mainly to create spatially continuous values across a regular grid of                                                                                                                 6 https://github.com/ecoclimate. cells (see details in harris et al., 2014). because weather stations are not regularly spatially, this continuity is crucial (see examples in new et al., 2002; and hijmans et al., 2005). in our case, aogcms outputs are already originally gridded and continuous, albeit at coarse resolutions (table 1), so we regridded raw variables from model-specific native resolutions to a global 0.5° grid. we thus produced climatic layers at a resolution relevant to the spatial scales at which macroecologists and biogeographers are interested, and on a comparable grid system among all aogcms. change-factor approach we followed the change-factor approach to downscaling (wilby et al., 2004). this approach comprises three steps: (i) compute the change-factor (also called climate change trends or anomalies) between past/future and baseline climate for each raw variable at model-specific native spatial resolution, (ii) downscale change-factor (“smoothing”) and the corresponding baseline climate from each aogcm to the standard 0.5° resolution, and (iii) apply downscaled change-factor to the downscaled baseline climate to reconstitute values and obtain downscaled layers for past and future climates. in the change-factor approach, current climate layers from weather station interpolations are often used to represent baseline climates from which change-factors are computed (hijmans and graham, 2006; kriticos et al., 2012). we considered three scenarios from aogcms as baseline climates (pre-industrial, historical, and modern), taking into account macroecological and biogeographic interests, as has been used in applications such as ecological niche modeling (terribile et al., 2012; collevatti et al., 2013a; 2013b; lima et al., 2014); they also cover time periods for ample biodiversity data exist, and so are potentially useful for models relating organisms to environments. in the first step, change-factors for temperature variables (t change-factors) were computed as the simple difference between past/future and baseline conditions (a standard anomaly in climatology) for each grid cell, for a given aogcm. for precipitation, changefactors (p change-factors) were computed as ratios of anomalies to corresponding baseline conditions. ratios are more robust in townpeterson typewritten text 5   figure 2. maps illustrating the 19 bioclimatic variables available in ecoclimate; red and blue titles indicate maps for temperature and precipitation variables, respectively. for simplicity, maps were built only for the lgm in south america, as predicted by aogcm ccsm4. however, ecoclimate offers climate data at global extents for multiple periods (figure 3) and 9 aogcms (table 2), including raw monthly variables. the name and unit of bioclimatic variables are provided in table 4. townpeterson typewritten text 6   figure 3. time series illustrating the global extent and multi-temporal characteristic of climate data available in ecoclimate. for illustration, maps show only annual mean temperature (bio1) across past, present and future time periods from aogcm ccsm4. however, ecoclimate offers similar time series for 9 aogcms (table 2) and 19 bioclimatic variables (table 4 and figure 2). the period "present" represents the baseline used to downscale time series (pre-industrial ~1760, historical 19001949, or modern 1950-1999; not shown here, but see details in figure 5). color scale is standardized across all maps; white, red and blue tones indicate near zero, positive and negative temperatures, respectively. townpeterson typewritten text 7 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   maintaining original patterns in downscaling when managing large values, as in precipitation (wilby et al., 2004). in the second step, we used ordinary kriging to downscale both raw t and p change-factors and their respective baseline climates statistically to the standard global 0.5° grid (see details below on kriging methods). in the third step, we applied the downscaled t and p change-factors to the corresponding downscaled baseline layers to obtain downscaled past and future scenarios. this step follows the inverse of the computation in the first step. for temperature, this involves adding the t change-factors to the correspondent baseline temperature value in each grid cell; for precipitation, the values are multiplied by baseline values to unpack ratios on the original precipitation scale. ordinary kriging method we automated downscaling by coupling functions from the gstat package (pebesma, 2004) in a script in r (r development core team, 2014). downscaling was performed using the “krige” kriging function, based on the 12 nearest observations to a given focal point (rather than fitting an inverse distance weighted power from the global neighborhood) and a variogram model. to model the spatial structure in data, we fit a variogram using the “fit.variogram” function, which fits ranges and sills from a variogram model (here a spherical variogram) to a sample variogram. the spherical variogram model was used because it shows a progressive decrease of spatial autocorrelation out to some distance, beyond which autocorrelation is zero, a common spatial structure in climate data (in an exponential model, for example, autocorrelation would disappear completely only at an infinite distance). the sample variogram was obtained using the “variogram” function, following the direction with the largest range (i.e., the omnidirectional model type) in each variable, and assuming a constant trend for variables (i.e., we did not specify predictors to fit sample variograms). integrating these functions from gstat package in r (pebesma, 2004) made automating the interpolation procedure possible. scripts are available openly at the internet link given in the footnote7. sensitivity analyses: comparing downscaling methods diverse statistical methods have been used for downscaling data and generate standardized, finer-resolution climate surfaces. we used ordinary kriging because it is known to produce reliable regridded surfaces by considering spatial structure in raw gridded variables to minimize variance of errors (hartkamp et al., 1999). considering spatial structure in data is an important advantage relative to other simple linear interpolation techniques (e.g., regression methods, see hartkamp et al., 1999, for an overview); ordinary kriging is desirable here for regridding climatic simulations that reflect the spatial structure of original gridded boundary conditions (e.g., ice sheet, topography, vegetation, insolation). moreover, because our dataset was based on gridded climatic simulations, it makes no conceptual sense to account for effects on observed climate patterns (e.g., coastal influence, terrain barriers, temperature inversions, which are explicitly accounted by the prism method, for example; see daly et al., 2002), nor linking weather stations along isoclines from irregularly spaced data points (which, for example, would be obtained by thin-plate spline-fitting techniques like anusplin; see details in hutchinson and xu, 2013; see applications in new et al., 2002; hijmans et al., 2005). however, to avoid doubt about our choices, we spatially downscaled raw temperature mean (tas) and precipitation (pr) variables from aogcm ccsm4, preindustrial experiment, using other 4 common methods: thin-plate splines, inverse distance weighting, trend surface (best results achieved with 12th-order polynomial regression), and natural neighbor. to compare methods, we correlated all downscaled layers each other, and found them highly spatially concordant with corresponding originally interpolated layers by ordinary kriging (all r > 0.96 for precipitation, except for trend surface method; all r > 0.98 for temperature; table 3). moreover, we evaluated the efficiency of each downscaling method by comparing                                                                                                                 7  https://github.com/ecoclimate.   townpeterson typewritten text 8   table 4. description of the 19 bioclimatic variables and contribution of raw monthly temperature and precipitation variables used in their calculations, all available in ecoclimate. tasmin: minimum surface temperature; tasmax: maximum surface temperature; tas: mean surface temperature; pr: precipitation flux. bioclimatic variables raw variables variable description (name/unit) tasmin (°c) tasmax (°c) tas (°c) pr (mm m-2) bio1 annual mean temperature (°c) ! bio2 mean diurnal range (°c) (mean of monthly (max temp min temp)) ! ! bio3 isothermality (%) (100*bio2/bio7) ! ! bio4 temperature seasonality (%) (standard deviation *100) ! bio5 max temperature of warmest month (°c) ! bio6 min temperature of coldest month (°c) ! bio7 temperature annual range (°c) (bio5-bio6) ! ! bio8 mean temperature of wettest quarter (°c) ! ! bio9 mean temperature of driest quarter (°c) ! ! bio10 mean temperature of warmest quarter (°c) ! bio11 mean temperature of coldest quarter (°c) ! bio12 annual precipitation (mm/m2) ! bio13 precipitation of wettest month (mm/m2) ! bio14 precipitation of driest month (mm/m2) ! bio15 precipitation seasonality % (coefficient of variation) ! bio16 precipitation of wettest quarter (mm/m2) ! bio17 precipitation of driest quarter (mm/m2) ! bio18 precipitation of warmest quarter (mm/m2) ! ! bio19 precipitation of coldest quarter (mm/m2) ! ! townpeterson typewritten text 9 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   values of 5000 random points (n, ~10% of original ccsm4 grid cells) from downscaled (x) and native gridded (z) layers using mean square errors, as mse = 1/n*σ(xi – zi)2. from mse, lower errors indicate more precise methods: downscaled estimates are more similar to corresponding original values. this procedure was repeated 1000 times. multiple regridded values (finer resolution) matching every aogcm-native grid cell (coarser resolution) were averaged to allow direct comparison. kriging showed lowest mses for both temperature and precipitation variables (figure 1). our sensitivity analyses showed that, although all methods produced downscaled climatic layers with similar spatial patterns (high correlations), ordinary kriging was the most precise for downscaling both temperature and precipitation variables (lowest mse). because the climate science community has not established best practices by which to develop higher-resolution climate layers (hall, 2014), our sensitivity analyses at least guarantee that ordinary kriging represents a good practical choice. building bioclimatic layers we used the downscaled layers of the 4 raw variables (precipitation, mean temperature, maximum temperature, minimum temperature) for the 12 months of the year to calculate the 19 core bioclimatic variables (table 4). we followed the standard equations used by worldclim8, except that bio1 (annual mean temperature) was obtained directly from aogcm simulations (variable tas), instead of as an average maximum and minimum temperatures, as implemented in the "biovars" function in the dismo package in r (hijmans et al., 2013). the ecoclimate database the web-repository we created a web-repository, ecoclimate9, to share downscaled bioclimatic layers (figure 2), as well as the long-term means for raw monthly temperature and precipitation variables. bioclimatic variables are commonly used in macroecological and biogeographic analyses, like ecological niche modeling. monthly variables are needed to compute                                                                                                                 8 http://www.worldclim.org/bioclim. 9 http:///www.ecoclimate.org. other desirable climate predictors, such as actual and potential evapotranspiration (aet/pet, see aet calculator10). the dataset includes standard 0.5° gridded climate layers for mid-pliocene (~3.3 to 3.0 ma), lgm (21 ka), mid-holocene (6 ka), pre-industrial (~1760), historical (1900-1949), modern (1950-1999) and future conditions (20802100, end of the 21st century), all from updated aogcms available in the most recent cmip5 and pmip3 climate modeling projects (figure 3). future simulations include four representative concentration pathways (rcps): rcp2.6 (low emissions scenario), rcp4.5 and rcp6.0 (intermediate emissions scenarios), and rcp8.5 (high emissions scenario; see details about climate scenarios in taylor et al., 2012). a distinctive attribute of ecoclimate is its multi-model and multi-temporal coverage (figure 4). this distinction makes ecoclimate potentially applicable to a multitude of questions commonly asked in macroecology and biogeography. however, in analyses using multi-temporal climatic scenarios, downscaled layers should be matched. for example, past and future layers are genuinely comparable only if downscaled from the same baseline condition (pre-industrial, historical or modern climate; figure 5). this detail is needed to ensure that climate layers are comparable, and not reflecting differences among baselines. by presenting and serving uniform data from different climatic simulations that are compatible through time, ecoclimate allows users to consider and evaluate apparent differences among aogcms (see details on climate modeling uncertainties in taylor et al., 2012, and harris et al., 2014). the multiple current climate data available in ecoclimate that were used across the downscaling procedure as baseline scenarios are specific to each aogcm, instead of a unique standard observed climate (e.g., from interpolations among weather stations), favors keeping modeling uncertainties intact. therefore, considering the spread of results as available in ecoclimate is crucial to perspectives on the range of potential signals of interest in macroecological and biogeographic studies (see discussion in varela et al., 2015b).                                                                                                                 10 http://geog.uoregon.edu/envchange/software.html. townpeterson typewritten text 10   fi gu re 4 . o ve rv ie w o f th e ec oc lim at e da ta ba se , da ta a va ila bi lit y, a ttr ib ut es , an d pe rs pe ct iv es . ec oc lim at e of fe rs c om pa tib le 0 .5 ° re gr id de d cl im at e la ye rs f ro m t he m os t re ce nt p ha se o f th e m od el in g pr oj ec ts c m ip 5 an d pm ip 3, e nc om pa ss in g m ul tip le a o g c m s an d ex pe rim en ts ac ro ss t he p as t, pr es en t, an d fu tu re . th e pe rs pe ct iv e fo r ke ep in g ec oc lim at e co m pl et e an d up -to -d at e w ill r eq ui re i nc or po ra tio n of f ur th er a o g c m s fr om n ex t-g en er at io n ex pe rim en ts ; s ee m ee hl e t a l. 20 14 . e xp er im en t a cr on ym s: p lio : m id -p lio ce ne w ar m p er io d (m pw p, ~ 3. 3 to 3. 0 m ill io n ye ar s ag o) ; lg m : la st g la ci al m ax im um ( 21 ,0 00 y ea rs a go ); m h : m id -h ol oc en e (6 00 0 ye ar s ag o) ; pi c on tro l: pr ein du st ria l (~ 17 60 ), h is t: h is to ric al ( or ig in al ly 1 85 020 05 , bu t in e co c lim at e av er ag ed i n h is to ric al 1 90 019 49 a nd m od er n 19 50 -1 99 9) ; r c ps : re pr es en ta tiv e ca rb on p at hw ay s w ith e m is si on s ce na rio s fo r th e en d of th e 21 st c en tu ry ( 20 80 -2 10 0) ; l ig : l as t i nt er g la ci al ( ~1 25 ,0 00 y ea rs ag o) ; m io : m id -m io ce ne c lim at ic o pt im um (m m c o , 1 7 to 1 4. 5 m ill io n ye ar s a go ). se e de ta il ab ou t a o g c m s i n ta bl e 1. townpeterson typewritten text 11   figure 5. schematic representation of the compatibility pattern among downscaled layers from ecoclimate database. for each aogcm, three groups of downscaled layers have been derived based on distinct baseline scenarios (pre-industrial, historical, modern). past and future climate layers are necessarily compatible for a same baseline, but incompatible among baselines. see experiment acronyms in figure 4. townpeterson typewritten text 12 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   using a scenario based on observed climate (see example in hijmans and graham, 2006) would reduce variability among aogcm outputs by relating downscaled past and future layers to the same unique baseline, and could hide potential macroecological and biogeographic responses or underestimate their variation or uncertainty. not considering the full diversity of potential responses to climate change may compromise many research questions, or even lead to invalid results in many cases (e.g., see the importance of using intact climate uncertainty for phylogeographic inference as discussed in collevatti et al., 2013b; 2015). potentiality, applicability, and relevance the data served via ecoclimate are potentially applicable to a multitude of research interests, including numerous questions in macroecology and biogeography, but also in diverse environmental, agricultural and paleobiological sciences (figure 6). ecological niche modeling (enm) and its numerous research purposes, for instance, offer an excellent example. ecological niche models estimate associations between environmental aspects (most often climate) and known occurrences of species to characterize the range of conditions under which the species’ populations are viable. this suite of methods and ideas has been applied to diverse research purposes: guiding discovery of populations of known (bourg et al., 2005; guisan et al., 2006) and unknown (raxworthy et al., 2003) species; understanding distributional dynamics under past (banks et al., 2008a; 2008b) and future (dormann, 2007; anderson, 2013) climates; anticipating climate change impacts on agricultural (fraga et al., 2013) and natural (nabout et al., 2011) extraction; mapping invasion risk (peterson, 2003; jiménez-valverde et al., 2011), pest distributions (venette et al., 2010; estay et al., 2014), and disease trasmission (peterson, 2014); estimating population parameters (tôrres et al., 2012; lima-ribeiro and dinizfilho, 2013; thuiller et al., 2014), species richness (wisz and rahbeck, 2007; limaribeiro et al., 2013b), and community composition (pellissier et al., 2012); analyzing biotic interactions (anderson et al., 2002; wheeler et al., 2015); illuminating patterns and processes of diversification and speciation (silva et al., 2014); characterizing dispersal (génard and lescourret, 2013; saltré et al., 2015); highlighting extinction (nogués-bravo et al., 2008; lima-ribeiro et al., 2013a); testing niche conservatism (martínez-meyer et al., 2004; 2006; peterson and nyári, 2007; jakob et al., 2010) and phylogeographic hypotheses (collevatti et al., 2013b; 2015; alvarado-serrano and knowles, 2014); establishing historical refugia (waltari et al., 2007; terribile et al., 2012); identifying biodiversity hotspots (carnaval and moritz, 2008; carnaval et al., 2009); choosing appropriate areas for biodiversity conservation (nobrega and de marco, 2011; williams et al., 2013) and translocation (martínez-meyer et al., 2006; richmond et al., 2010); etc. all of these applications depended on data such as those now served via ecoclimate. besides uses in enm, ecoclimate can be applicable to diverse other research areas in the natural sciences. quantifying and mapping historical climate signatures in relation to spatial (araújo et al., 2008) and temporal (lyons and wagner, 2009) biodiversity patterns, for example, is a general issue in macroecology to which ecoclimate data have much to offer. in the “new” paleobiology, a research field integrating paleontologists and evolutionary theorists, paleoclimatic simulations have provided opportunities for testing climatic controls on macroevolutionary patterns (eronen et al., 2009; myers and saupe, 2013). similarly, community phylogenetics has recently seen important advances by drawing information from climate models to understand community assembly on geographic scales (hawkins et al., 2014). also of current interest are climatechange-induced transformations on agricultural systems (image team, 2001; ramirez-villegas et al., 2013), not restricted to food supply (parry et al., 2004; elliott et al., 2014), but including conservation issues (hannah et al., 2013; zarco-gonzález et al., 2013). although not exhaustive, this list of research interests clearly exemplifies the relevance of ecoclimate to multiple studies in the natural sciences. challenges: resolution, scale and uncertainty building databases is challenging in multiple dimensions, including operational and intrinsic, data-related features. first, ecoclimate presents processed climate layers townpeterson typewritten text 13   figure 6. flowchart showing potential applications of ecoclimate in diverse areas of the natural sciences. townpeterson typewritten text 14 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   based on data derived from dynamic modeling groups who advance in model predictions by improving their earth system equations, model assumptions, input data, and parameters (haywood et al., 2011), so processed layers must be updated continuously. second, as should be apparent from the discussions above, several features of simulated climate data (e.g., spatial resolution, temporal scale, and modeling uncertainty) challenge researchers in every analysis. spatial resolution of native-gridded simulations is irregular among aogcms and often coarse, ranging from 1-4° of latitude and longitude. the regridding performed in developing ecoclimate data sets had two finalities: producing climate layers that are directly comparable across geography from a standard grid, and that are at resolutions relevant to macroecologists and biogeographers. the standard 0.5° grid fulfills these goals: although 0.5° resolution might be coarse for specific scenarios in more regional and local analyses, we decided not to produce high-resolution layers from the aogcms because simple downscaling do not produce any new information in the climate change signal and often creates artificially highresolution surfaces in which climate information can be no more reliable than the native, coarse-resolution simulations underlying them (harris et al., 2014). rather, climate signals in global circulation models are often spatially biased at regional scales (e.g., lacking certain features of atmospheric circulation like the jet stream) which may reduce credibility of downscaled data (hall, 2014). besides the peculiarities of climate noise per se, artificially highresolution climate surfaces may provide unreliable signals in macroecological and biogeographic models in the form of classical consequences of pseudoreplication for statistical results. that is, greater detail in artificially pseudoreplicated information across space does not imply more accurate information (taylor et al., 2012). this point holds in particular in regions with complex topography, such as where elevational gradients determine significant climate differentiation across local to regional scales (e.g., the andes in south america, the rocky mountains in north america, the alps in europe, and the himalayas in asia). such limitations, however, do not invalidate high-resolution downscaling, as long as their limitations are understood. actually, producing reliable climate layers at resolutions relevant at local to regional scales is possible via dynamic downscaling procedures (or regional modeling; pal et al., 2007). dynamic downscaling is based on regional climate models simulated from finerresolution surface features such as terrain, whereas simple downscaling uses transfer functions representing climate relationships at global scales (pielke and wilby, 2012). however, regional climate downscaling is challenging at broad extents, like the global climate layers in ecoclimate, because mesoscale models simulating dynamical regional climate features are not yet available for most regions worldwide (kerr, 2011). meanwhile, the climate modeling community suggests some caution with downscaled information: in general, careful researchers may wish to avoid consideration of downscaled information from the cmip5 models unless they have become sufficiently aware of the limitations of both the global models and the downscaling methods. (taylor et al. 2012: 496) such spatially finer information is needed for specific analyses at regional scales. indeed, macroecologists and biogeographers also need climate data at finer temporal resolutions, especially for past conditions to test hypotheses in evolutionary macroecology (diniz-filho et al., 2013) and paleobiogeography (varela et al., 2011). the climate modeling community simulates climates for key past periods during which boundary conditions are extreme (e.g., last glacial maximum, mid-pliocene warm period), however, for long intermediate periods, direct estimates are generally lacking. to solve this problem, researchers have interpolated conditions to finer time-series using climate proxies as covariables, instead of interpolating values linearly (see lawing and polly, 2011; and rödder et al., 2013). although temporal interpolations are relatively straightforward for some variables like temperature, they are conceptually challenging. specific proxy estimates (e.g., 18o-oxygen radioisotopes, a direct proxy for townpeterson typewritten text 15 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   temperature change) are scarce, spatially or temporally aggregated, and represent continental dynamics unusually (see lisiecki and raymo, 2005 for details on the most accurate globally distributed benthic δ18o records). more challenging is the case of climatic variables (e.g., precipitation) for which reliable or direct proxies describing their temporal variability are lacking. finally, the time scale of climatic simulations is temporally limited. for instance, lgm (21 ka) is the oldest period in the pmip3 simulations (taylor et al., 2012). parallel simulations have been developed for the mid-pliocene (~3.3-3.0 ma; haywood et al., 2011), and plans exist for development of mid-miocene (17-14.5 ma; goldner et al., 2014) simulations as well, maybe as a specific experiment for the next generation (cmip6; see meehl et al., 2014). regardless, the temporal limitations of aogcms challenge macroecological and biogeographic studies of evolutionary dynamics that are older than existing climatic simulations. lastly, predictive uncertainty in aogcm outputs represents another aspect that challenges macroecological and biogeographic analyses. gcms simulate climate conditions based on predetermined external forcings (e.g., earth's orbital parameters, anthropogenic activities, greenhouse gas concentrations), but exhibit variations owing to internal interactions of stochastic and nonlinear climate features manifested on a variety of time scales (e.g., el niño events; taylor et al., 2012). these unforced variations produce noise in climate at different spatial scales. hence, it is important that macroecologists and biogeographers consider all uncertainty from aogcms to derive realistic measures of confidence around predictions of biological signals. however, aogcms and particular variables can be selected a priori based on similarity patterns across specific regions and periods (see the recent proposal of varela et al., 2015b). perspectives many practical aspects that impeded comprehensive macroecological and biogeographic studies two or three decades ago have received special attention, and have been overcome at least partially over the past 15 years by making available detailed biological and environmental databases that are easily accessed, managed, and updated (wilson, 2000; hijmans et al., 2005). ecoclimate adds to this effort to build permanent research infrastructure in this field. ecoclimate arose from the need for detailed and rigorously-documented climate layers to build ecological niche models and test macroecological and biogeographic hypotheses through the past, present, and future (see fundamental questions in sutherland et al., 2013; seddon et al., 2014). ecoclimate offers a range of processed multitemporal climate layers from the most recent multi-model ensembles by cmip5 and pmip3 over diverse time frames, at global extents and 0.5° spatial resolution. our plan is to maintain the ecoclimate database updated continuously as new model outputs become available. when new experiments (e.g., miomip) and next generations of climate models from cmip and pmip (e.g., cmip6; see meehl et al., 2014) and their descendents are developed, they will be processed and incorporated into the ecoclimate database. what is more, as a key next step, we will also advance the processing of climate layers by developing bias-corrected layers, and build temporally interpolated sequences at 1 ka resolution. we are developing an r package providing functions to deal with all these processing phases along ecoclimate purpose. because many aogcmbased outputs were not processed for ecoclimate, including hundreds of variables (e.g., evaporation, relative humidity, heat flux, etc.; see the complete list11), simulated for distinct time frequencies (daily, every 6h, 3h) and realms (e.g., land, land-ice, ocean), an r package will surely help ecoclimate users to amplify their research horizons. acknowledgments we first acknowledge the world climate research programme's working group on coupled modeling by the cmip5 and pmip3, and we thank the climate modeling groups (table 1) for producing and making available model outputs. the u.s. department of energy's program for climate model diagnosis and intercomparison provides coordinating support and led development of software infrastructure in partnership with the global organization for earth system science                                                                                                                 11 http://cmip-pcmdi.llnl.gov/cmip5/docs/standard_output.pdf. townpeterson typewritten text townpeterson typewritten text 16 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   portals. we also thank thiago f. rangel (universidade federal de goiás, brazil) and a. townsend peterson (university of kansas) for help and invaluable suggestions. the feedback from users who used our previous climatic data sets was also valuable. finally, we dedicate ecoclimate to the memory of mariana rocha, who was enthusiastically interested in this project when developing the early ecoclimate team. we acknowledge financial support from cnpq, capes, and fapeg for supporting our work via multiple grants and fellowships, especially the research network genpac (geographical genetics and regional planning for natural resources in brazilian cerrado). references alvarado-serrano, d.f. and l.l. knowles. 2014. ecological niche models in phylogeographic studies: applications, advances and precautions. mol ecol resour 14: 233-248. anderson, r.p., a.t. peterson and m. gómezlaverde. 2002. using niche-based gis modeling to test geographic predictions of competitiveexclusion and competitive release in south american pocket mice. oikos 98: 316. anderson, r.p. 2013. a framework for using niche models to estimate impacts of climate change on species distributions. ann n y acad sci 1297: 8-28. araújo, m.b., f. ferri-yáñez, f. bozinovic, p.a. marquet, f. valladares and s.l. chown. 2013. heat freezes niche evolution. ecol lett 16: 1206-1219. araújo, m.b., d. nogués-bravo, j.a.f. dinizfilho, a.m. haywood, p. valdes and c. rahbeck. 2008. quaternary climate changes explain diversity among reptiles and amphibians. ecography 31: 8-15. banks, w.e., f. d'errico, a.t. peterson, m. kageyama and g. colombeau. 2008a. reconstructing ecological niches and geographic distributions of caribou (rangifer tarandus) and red deer (cervus elaphus) during the last glacial maximum. quat sci rev 27: 2568-2575. banks, w.e., f. d'errico, a.t. peterson, m. vanhaeren, m. kageyama, p. sepulchre, g. ramstein, a. jost and d. lunt. 2008b. human ecological niches and ranges during the lgm in europe derived from an application of ecocultural niche modeling. j archaeol sci 35: 481-491. barrientos, r., l. kvist, a. barbosa, f. valera, f. khoury, s. varela and e. moreno. 2014. refugia, colonization and diversification of an arid-adapted bird: coincident patterns between genetic data and ecological niche modelling. mol ecol 23: 390-407. bartlein, p.j., s.p. harrison, s. brewer, s. connor, b.a.s. davis, k. gajewski, j. guiot, t. harrison-prentice, a. henderson, o. peyron, i.c. prentice, m. scholze, h. seppa, b. shuman, s. sugita, r.s. thompson, a.e. viau, j. williams and h. wu. 2011. pollenbased continental climate reconstructions at 6 and 21 ka: a global synthesis. clim dyn 37: 775-802. bourg, n.a., w.j. mcshea and d.e. gill. 2005. putting a cart before the search: successful habitat prediction for a rare forest herb. ecology 86: 2793-2804. braconnot, p., b. otto-bliesner, s. harrison, s. joussaume, j.-y. peterchmitt, a. abe-ouchi, m. crucifix, e. driesschaert, th. fichefet, c.d. hewitt, m. kageyama, a. kitoh, a. laîné, m.-f. loutre, o. marti, u. merkel, g. ramstein, p. valdes, l. weber, y. yu and y. zhao. 2007. results of pmip2 coupled simulations of the mid-holocene and last glacial maximum – part 1: experiments and large-scale features. clim past 3: 261-277. carnaval, a. and c. moritz. 2008. historical climate modelling predicts patterns of current biodiversity in the brazilian atlantic forest. j biogeogr 35: 1201. carnaval, a.c., m.j. hickerson, c.f.b. haddad, m.t. rodrigues and c. moritz. 2009. stability predicts genetic diversity in the brazilian atlantic forest hotspot. science 323: 785789. chamberlain, s., c. boettiger, k. r. ram and v. barve. 2013. rgbif: interface to the global biodiversity information facility api methods., r package version 0.3.0. http://cran.r-project.org/package=rgbif. collevatti, r., m.s. lima-ribeiro, j.a.f. dinizfilho, g. oliveira, r. dobrovolski and l.c. terribile. 2013a. stability of brazilian seasonally dry forests under climate change: inferences for long-term conservation. am j plant sci 4: 792-805. collevatti, r.g., l.c. terribile, g. oliveira, m.s. lima-ribeiro, j.c. nabout, t.f. rangel and j.a.f. diniz-filho. 2013b. drawbacks to palaeodistribution modelling: the case of south american seasonally dry forests. j biogeogr 40: 345-358. collevatti, r., l.c. terribile, j.a.f. diniz-filho and m.s. lima-ribeiro. 2015. multi-model inference in comparative phylogeography: an integrative approach based on multiple lines of evidence. front genet 6: 1-8. daly, c., w.p. gibson, g.h. taylor, g.l. johnson and p. pasteris. 2002. a knowledge-based approach to the statistical mapping of climate. clim res 22: 99-113. townpeterson typewritten text 17 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   diniz-filho, j.a.f., s. gouveia and m.s. limaribeiro. 2013. evolutionary macroecology. front biogeogr 5: 195-203. dormann, c.f. 2007. promising the future? global change projections of species distributions. basic appl ecol 8: 387-397. dowsett, h.j., m.m. robinson, a.m. haywood, d.j. hill, a.m. dolan, d.k. stol, w.-l. chan, a. abe-ouchi, m.a. chandler, n.a. rosenbloom, b.l. otto-bliesner, f.j. bragg, d.j. lunt, k.m. foley and c.r. riesselman. 2012. assessing confidence in pliocene sea surface temperatures to evaluate predictive models. nat clim change 2: 365-371. elliott, j., d. deryng, c. muller, k. frieler, m. konzmann, d. gerten, m. glotter, m. flörke, y. wada, n. best, s. eisner, b.m. fekete, c. folberth, i. foster, s.n. gosling, i. haddeland, n. khabarov, f. ludwig, y. masaki, s. olin, c. rosenzweig, a.c. ruane, y. satoh, e. schmid, t. stacke, q. tang and d. wisser. 2014. constraints and potentials of future irrigation water availability on agricultural production under climate change. proceedings of the national academy of sciences of the usa 111: 3239-3244. eronen, j.t., m.m. ataabadi, a. micheels, a. karme, r.l. bernor and m. fortelius. 2009. distribution history and climatic controls of the late miocene pikermian chronofauna. proc natl acad sci of the usa 106: 1186711871. estay, s., f. labra, r. sepulveda and l. bacigalupe. 2014. evaluating habitat suitability for the establishment of monochamus spp. through climate-based niche modeling. plos one 9: e102592. fraga, h., a.c. malheiro and j. moutinho-pereira. 2013. future scenarios for viticultural zoning in europe: ensemble projections and uncertainties. int j biometeorol 57: 909-925. génard, m. and f. lescourret. 2013. combining niche and dispersal in a simple model (ndm) of species distribution. plos one 8: e79948. goldner, a., n. herold and m. huber. 2014. the challenge of simulating the warmth of the mid-miocene climatic optimum in cesm1. clim past 10: 523-536. guisan, a., o. broennimann, r. engler, m. vust, n.g. yoccoz, a. lehman and n.e. zimmermann. 2006. using niche-based models to improve the sampling of rare species. conserv biol 20: 501-511. hall, a. 2014. projecting regional change. science 346: 1461-1462. hannah, l., p.r. roehrdanz, m. ikegami, a.v. shepard, m.r. shawc, g. tabor, l. zhi, p. marquet and r.j. hijmans. 2013. climate change, wine, and conservation. proc natl acad sci usa 110: 6907-6912. harris, r.m.b., m.r. grose, g. lee, n.l. bindoff, l.l. porfirio and p. fox-hughes. 2014. climate projections for ecologists. wires clim change 5: 621-637. harrison, s.p., p.j. bartlein, s. brewer, i.c. prentice, m. boyd, i. hessler, k. holmgren, k. izumi and k. willis. 2014. climate model benchmarking with glacial and mid-holocene climates. clim dyn 43: 671-688. hartkamp, a. d., k. de beurs, a. stein and j. w. white. 1999. interpolation techniques for climate variables. nrg-gis series 99-01: cimmyt, mexico, d.f. hawkins, b.a., m. rueda, t.f. rangel, r. field and j.a.f. diniz-filho. 2014. community phylogenetics at the biogeographical scale: cold tolerance, niche conservatism and the structure of north american forests. j biogeogr 41: 23-38. haywood, a.m., h.j. dowsett, m.m. robinson, d.k. stoll, a.m. dolan, d.j. lunt, b. ottobliesner and m.a. chandler. 2011. pliocene model intercomparison project (pliomip): experimental design and boundary conditions (experiment 2). geosci model develop 4: 571-577. hijmans, r.j., s.e. cameron, j.l. parra, p.g. jones and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. int j clim 25: 1965-1978. hijmans, r.j., s. phillips, j.r. leathwick and j. elith. 2013. dismo: species distribution modeling. r package version 0. 9-3. http://cran. r-project. org/package=dismo. hijmans, r.j. and c.h. graham. 2006. the ability of climate envelope models to predict the effect of climate change on species distributions. glob change biol 12: 22722281. hutchinson, m. f. and t. xu. 2013. anusplin version 4.4: user guide. fenner school of environment and society (australian national university), canberra. image team. 2001. the image 2.2 implementation of the sres scenarios: a comprehensive analysis of emissions, climate change and impacts in the 21st century. netherlands environmental assessment agency (mnp) cd-rom publication 500110001, bilthoven, netherlands. jakob, s.s., c. heibl, d.r. dder and f.r. blattner. 2010. population demography influences climatic niche evolution: evidence from diploid american hordeum species (poaceae). mol ecol 19: 1423-1438. jiménez-valverde, a., a.t. peterson, j. soberón, j.m. overton, p. aragón and j.m. lobo. 2011. use of niche models in invasive species risk assessments. biol invasions 13: 2785-2797. townpeterson typewritten text 18 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   kerr, r.a. 2011. vital details of global warming are eluding forecasters. science 334: 173-174. kriticos, d.j., b.l. webber, a. leriche, n. ota, i. macadam, j. bathols and j.k. scott. 2012. climond: global high-resolution historical and future scenario climate surfaces for bioclimatic modelling. methods ecol evol 3: 53-64. lawing, a.m. and p.d. polly. 2011. pleistocene climate, phylogeny, and climate envelope models: an integrative approach to better understand species' response to climate change. plos one 6: e28554. lima, n.e., m.s. lima-ribeiro, c.f. tinoco, l.c. terribile and r.g. collevatti. 2014. phylogeography and ecological niche modelling, coupled with the fossil pollen record, unravel the demographic history of a neotropical swamp palm through the quaternary. j biogeogr 41: 673-686. lima-ribeiro, m. s. and j. a. f. diniz-filho. 2013. modelos ecológicos e a extinção da megafauna: clima e homem na américa do sul. cubo, são carlos. lima-ribeiro, m.s., d. nogués-bravo, l.c. terribile, b. persaram and j.a.f. diniz-filho. 2013a. climate and humans set the place and time of proboscidean extinction in late quaternary of south america. palaeogeogr palaeoclimatol palaeoecol 392: 546-556. lima-ribeiro, m.s., f.a.v. faleiro and d.p. silva. 2013b. current and historical climate signatures to deconstructed tree species richness pattern in south america. acta sci biol sci 35: 219-231. lisiecki, l.e. and m.e. raymo. 2005. a pliocenepleistocene stack of 57 globally distributed benthic δ18o records. paleoceanogr 20: pa1003. lyons, s.k., wagner, p., 2009. using a macroecological approach to the fossil record to help inform conservation biology. in: dietl, g.p., flessa, k.w. (eds.), conservation paleobiology: using the past to manage for the future. the paleontological society, pp. 141166. martínez-meyer, e., a.t. peterson, j.a. servin and l.f. kiff. 2006. ecological niche modelling and prioritizing areas for species reintroductions. oryx 40: 411-418. martínez-meyer, e. and a.t. peterson. 2006. conservatism of ecological niche characteristics in north american plant species over the pleistocene-to-recent transition. j biogeogr 33: 1779-1789. martínez-meyer, e., a.t. peterson and w.w. hargrove. 2004. ecological niches as stable distributional constraints on mammal species, with implications for pleistocene extinctions and climate change projections for biodiversity. glob ecol biogeogr 13: 305-314. meehl, g.a., r. moss, k.e. taylor, v. eyring, r.j. stouffer, s. bony and b. stevens. 2014. climate model intercomparison: preparing for the next phase. eos, trans am geophys union 95: 77-78. myers, c.e. and e.e. saupe. 2013. a macroevolutionary expansion of the modern synthesis and the importance of extrinsic abiotic factors. palaeontology 56: 1179-1198. nabout, j.c., g. oliveiro, m.r. magalhães, l.c. terribile and f.a.s. almeida. 2011. global climate change and the production of "pequi" fruits (caryocar brasiliense) in the brazilian cerrado. braz j nat conserv 9: 55-60. new, m., d. lister, m. hulme and i. makin. 2002. a high-resolution data set of surface climate over global land areas. clim res 21: 1-25. nobrega, c. and p. de marco. 2011. unprotecting the rare species: a niche-based gap analysis for odonates in a core cerrado area. div distrib 17: 491-505. nogués-bravo, d., j. rodríguez, j. hortal, p. batra and m.b. araújo. 2008. climate change, humans, and the extinction of the woolly mammoth. plos biol 6: 685-692. pal, j.s., f. giorgi, x. bi, n. elguindi, f. solmon, s.a. rauscher, x. gao, r. francisco, a. zakey, j. winter, m. ashfaq, f.s. syed, l.c. sloan, j.l. bell, n.s. diffenbaugh, j. karmacharya, a. konaré, d. martinez, r.p. rocha and a.l. steiner. 2007. regional climate modeling for the developing world: the ictp regcm3 and regcnet. bull am meteorol soc 88: 1395-1409. parry, m.l., c. rosenzweig, a. iglesias, m. livermore and g. fischer. 2004. effects of climate change on global food production under sres emissions and socio-economic scenarios. glob environ change 14: 53-67. pebesma, e.j. 2004. multivariable geostatistics in s: the gstat package. comp geosci 30: 683691. pellissier, l., j.-n. pradervand, j. pottier, a. dubuis, l. maiorano and a. guisan. 2012. climate-based empirical models show biased predictions of butterfly communities along environmental gradients. ecography 35: 684692. peterson, a. t. 2014. mapping disease transmission risk: enriching models using biogeography and ecology. johns hopkins university press, baltimore. peterson, a.t. and á. nyári. 2007. ecological niche conservatism and pleistocene refugia in the thrush-like mourner, schiffornis sp., in the neotropics. evolution 62: 173-183. townpeterson typewritten text 19 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   peterson, a.t. 2003. predicting the geography of species' invasions via ecological niche modeling. q rev biol 78: 419-433. pielke, r.a. and r.l. wilby. 2012. regional climate downscaling: what's the point? eos, trans am geophys union 93: 52-53. pierce, d. 2014. ncdf: interface to unidata netcdf data files. r package version 1.6.8. http://cran. r-project. org/package=ncdf. r development core team, 2014. r: a language and environment for statistical computing. in: r foundation for statistical computing, vienna, austria. isbn 3-900051-07-0, available at: http://www.r-project.org. ramirez-villegas, j., a.j. challinor, p.k. thornton and a. jarvis. 2013. implications of regional improvement in global climate models for agricultural impact research. environ res lett 8: 024018. raxworthy, c.j., e. martinez-meyer, n. horning, r.a. nussbaum, g.e. schneider, m.a. ortega-huerta and a.t. peterson. 2003. predicting distributions of known and unknown reptile species in madagascar. nature 426: 837-841. richmond, o., j. mcentee, r. hijmans and j. brashares. 2010. is the climate right for pleistocene rewilding? using species distribution models to extrapolate climatic suitability for mammals across continents. plos one 5: e12899. rödder, d., a.m. lawing, m. flecks, f. ahmadzadeh, j. dambach, j.o. engler, j.c. habel, t. hartmann, d. hornes, f. ihlow, k. schidelko, d. stiels and p.d. polly. 2013. evaluating the significance of paleophylogeographic species distribution models in reconstructing quaternary rangeshifts of nearctic chelonians. plos one 8: e72855. saltré, f., a. duputié, c. gaucherel and i. chuine. 2015. how climate, migration ability and habitat fragmentation affect the projected future distribution of european beech. glob change biol 21: 897-910. saupe, e.e., j.r. hendricks, r.w. portell, h.j. dowsett, a. haywood, s. hunter and b.s. lieberman. 2014. macroevolutionary consequences of profound climate change on niche evolution in marine mollusks over the past three million years. proc r soc biol sci ser b 281: 20141995. seddon, a.w.r., a.w. mackay, a.g. baker and et al. 2014. looking forward through the past: identification of 50 priority research questions in palaeoecology. j ecol 102: 256-267. silva, d.p., b. vilela, p. de marco and a. nemésio. 2014. using ecological niche models and niche analyses to understand speciation patterns: the case of sister neotropical orchid bees. plos one 9: e113246. sutherland, w.j., r.p. freckleton, h.c. godfray and et al. 2013. identification of 100 fundamental ecological questions. j ecol 101: 58-67. taylor, k.e., r.j. stouffer and g.a. meehl. 2012. an overview of cmip5 and the experiment design. bull am meteorol soc 93: 485-498. terribile, l.c., m.s. lima-ribeiro, m.b. araújo, n. bizão and j.a.f. diniz-filho. 2012. areas of climate stability in the brazilian cerrado: disentangling uncertainties through time. braz j nat conserv 10: 152-159. thomas, c.d., a. cameron, r.e. green, m. bakkenes, l.j. beaumont, y.c. collingham, b.f.n. erasmus, m.f. siqueira, a. grainger, l. hannah, l. hughes, b. huntley, a.s. van jaarsveld, g.f. midgley, l. miles, m.a. ortega-huerta, a.t. peterson, o.l. phillips and s.e. williams. 2004. extinction risk from climate change. nature 427: 145-148. thuiller, w., t. münkemüller, k.h. schiffers, d. georges, s. dullinger, v.m. eckhart, t.c. edwards jr., d. gravel, g. kunstler, c. merow, k. moore, c. piedallu, s. vissault, n.e. zimmermann, d. zurell and f.m. schurr. 2014. does probability of occurrence relate to population dynamics? ecography 1155-1166. tôrres, n.m., p. de marco, t. santos, l. silveira, a.t.d.a. jácomo and j.a.f. diniz-filho. 2012. can species distribution modelling provide estimates of population densities? a case study with jaguars in the neotropics. div distrib 18: 615-627. varela, s., j.m. lobo and j. hortal. 2011. using species distribution models in paleobiogeography: a matter of data, predictors and concepts. palaeogeogr palaeoclimatol palaeoecol 310: 451-463. varela, s., j. gonzález-hernández, e. casabella and r. barrientos. 2014a. ravis: an r-package for downloading stored in proyecto avis, a citizen science bird project. plos one 9: e91650. varela, s., j. gonzález-hernández and l. f. sgarbi. 2014b. paleobiodb: a package for downloading, visualizing and processing data from the paleobiology database. http://cran.rproject.org/web/packages/paleobiodb/. varela, s., m.s. lima-ribeiro, j.a.f. diniz-filho and d. storch. 2015a. differential effects of temperature change and human impact on european late quaternary mammalian extinctions. glob change biol 21: 1475-1481. varela, s., m.s. lima-ribeiro and l.c. terribile. 2015b. a short guide to the climatic variables of the last glacial maximum for biogeographers. plos one 10: e0129037. townpeterson typewritten text townpeterson typewritten text 20 biodiversity  informatics,  10,  2015,  pp.  1-­‐12   venette, r.c., d.j. kriticos, r.d. magarey, f.h. koch, r.h.a. baker, s.p. worner, n.n.g. raboteaux, d.w. mckenney, e.j. dobesberger, d. yemshanov, p.j. barro, w.d. hutchison, g. fowler, t.m. kalaris and j. pedlar. 2010. pest risk maps for invasive alien species: a roadmap for improvement. bioscience 60: 349-362. waltari, e., r.j. hijmans, a.t. peterson, á. nyári, s.l. perkins and r.p. guralnick. 2007. locating pleistocene refugia: comparing phylogeographic and ecological niche model predictions. plos one 2: e563. wheeler, h.c., j.d. chipperfield, c. roland and j.-c. svenning. 2015. how will the greening of the arctic affect an important prey species and disturbance agent? vegetation effects on arctic ground squirrels. oecologia 178: 915929. wilby, r.l., charles, s.p., zorita, e., timbal, b., whetton, p., mearns, l.o., 2004. guidelines for use of climate scenarios developed from statistical downscaling methods. in: ipcc task group on data and scenario support for impact and climate analysis (tgica)., available at: http://www.ipccdata.org/guidelines/dgm_no2_v1_09_2004.pd f (accessed 25 january 2014). williams, j.w., h.m. kharouba, s. veloz, m. vellend, j. mclachlan, z. liu, b. ottobliesner and f. he. 2013. the ice age ecologist: testing methods for reserve prioritization during the last global warming. glob ecol biogeogr 22: 289-301. wilson, e.o. 2000. a global biodiversity map. science 289: 2279. wisz, m.s. and c. rahbeck. 2007. using potential distributions to explore determinants of western palaearctic migratory songbird species richness in sub-saharan africa. j biogeogr 34: 828-841. zarco-gonzález, m.m., o. monroy-vilchis and j. alaníz. 2013. spatial model of livestock predation by jaguar and puma in mexico: conservation planning. biol conserv 159: 8087. townpeterson typewritten text 21 lima-ribeiro et al. ecoclimate 08-11-15 refoff lima-ribeiro et al. ecoclimate 08-11-15 refoff.2 lima-ribeiro et al. ecoclimate 08-11-15 refoff.3 lima-ribeiro et al. ecoclimate 08-11-15 refoff.4 lima-ribeiro et al. ecoclimate 08-11-15 refoff.5 biodiversity informatics, 17, 2022, pp. 59-66 59 rangemap: an r package to explore species geographic ranges marlon e. cobos1*, vijay barve2, narayani barve2, alberto jiménez-valverde3, claudia nuñez-penichet1 1department of ecology and evolutionary biology & biodiversity institute, university of kansas, lawrence, kansas 66045, usa 2florida museum of natural history, university of florida, gainesville, florida 32611, usa 3universidad de alcalá, departamento de ciencias de la vida, 28805, alcalá de henares, madrid, spain *corresponding author: marlon e. cobos, email: manubio13@gmail.com abstract. data exploration is a critical step in understanding patterns and biases in information about species’ geographic distributions. we present rangemap, an r package that implements tools to explore species’ ranges based on simple analyses and visualizations. the rangemap package uses species occurrence coordinates, spatial polygons, and raster layers as input data. its analysis tools help to generate simple spatial polygons summarizing ranges based on distinct approaches, including spatial buffers, convex and concave (alpha) hulls, trend-surface analysis, and raster reclassification. visualization tools included in the package help to produce simple, high-quality representations of occurrence data and figures summarizing resulting ranges in geographic and environmental spaces. functions that create ranges also allow generating extents of occurrence (using convex hulls) and areas of occupancy according to iucn criteria. a broad community of researchers and students could find in rangemap an interesting means by which to explore species’ geographic distributions. key words: area of occupancy, buffer, concave hull, convex hull, extent of occurrence, suitable areas, trend-surface analysis biodiversity conservation and research in biogeography, macroecology, disease ecology, and other fields, rely heavily on information about species’ geographic distributions. understanding species’ ranges is, however, a challenging task, because where a species is located depends on the existence of suitable environmental conditions, adequate dispersal ability, and appropriate biotic interactions (soberón and peterson 2005). multiple analyses can be done to characterize species’ distributions (e.g., correlative models, dispersal simulations, mechanistic models, etc.), and the quality of the results from such processes depend on many factors, including the algorithm used, knowledge of the species’ ecology and natural history, and quality and quantity of information available (clobert et al. 2012; radosavljevic and anderson 2014; qiao et al. 2015). species’ occurrence data are among the most numerous types of biological data available online thanks to recent technological developments and initiatives for data archiving and sharing, like the global biodiversity information facility, inaturalist, idigbio, and others (hardisty and roberts 2013; peterson et al. 2015). the existence of these data has facilitated use of diverse tools to model and simulate distributions and distributional dynamics, which has improved understanding of species’ ranges (franklin 2010; peterson et al. 2011). using these tools is data demanding, and requires considerable effort (anderson 2015), including intensive processes of data exploration. however, simple tools to facilitate such explorations of species distributional data are still scattered (i.e., multiple tools from distinct gis and statistical software may be required to perform such analyses). initial explorations are critical steps when working with distributional data (cobos et al. 2018). these processes consist of a series of analyses and visualizations that allow researchers to identify patterns and recognize potential errors (sillero et al. 2021). among the most common steps for data exploration, plotting records on a map offers a simple marlon e. cobos et al. – rangemap: an r-package to explore species geographic ranges 60 visualization that helps to recognize potential biases. creating simple range areas, derived from buffers or other types of polygons, also helps researchers to appreciate certain biases in sampling, geographic outliers, and areas related to species’ distributions. visualizations that integrate distributional data (occurrence data and/or simple range areas) and environmental information (in geographic or environmental spaces) give a general idea of conditions under which species have been detected and could exist. once a better understanding of the patterns of data has been gained by performing these explorations, more robust analyses to characterize species’ distributions can be performed. this is particularly relevant in conservation projects in which the quality of the data used and the modeling exercises determine de effectiveness of decisions that derive from the results of such applications. thanks to the relatively recent development of specialized packages, r (r core team 2021) is rapidly becoming an excellent alternative for analyzing spatial patterns in biodiversity data. taking advantage of these specialized packages and the versatility of r, we developed “rangemap,” a new r package to explore species’ distributions using simple algorithms. this set of tools offers handy and robust open-source options to generate and visualize ranges, which otherwise will require users to perform numerous analyses. the functions in rangemap facilitate automation of exploratory analyses, which can be helpful when working with large numbers of species. software description the rangemap r package contains tools that help researchers to explore species’ ranges based on occurrence data and simple algorithms. the main tools in this package help to perform analyses to generate hypotheses of species ranges (table 1).1 the functions in rangemap can be separated into two groups: (1) analysis functions, and (2) visualization functions. data required the main input for analysis functions is a set of occurrence data (geographic coordinates) in the form of a “data.frame”. other inputs vary depending on the function used, but they are common types of data used in geographic analyses (i.e., spatial polygons 1 https://cloud.r-project.org/web/packages/rangemap/rangemap.pdf. function group description rangemap_buff analysis generates a range by buffering species’ occurrences using a user-defined distance. rangemap_boundaries analysis creates a range by selecting all features of a spatial polygon layer in which the species is known to occur. individual polygons are selected considering geographic occurrences and/or by manually defining their names. rangemap_hull analysis generates ranges by creating convex hull (eddy 1977) or concave hull (park and oh 2012) polygons based on occurrence data. polygons can be split based on geographic clustering using hierarchical (everitt 1974) or k-means (hartigan and wong 1979) algorithms. final polygons can be buffered if needed. rangemap_enm analysis creates ranges by thresholding a continuous raster layer resulting from ecological niche modeling or species distribution modeling exercises. the threshold value is a user-specified level of omission or a specific value present in the continuous raster layer (nenzén and araújo 2011). rangemap_tsa analysis generates range polygons using species’ occurrences and trend surface analyses. trend surface analysis is a method based on low-order polynomials of spatial coordinates for estimating a regular grid of points from scattered observations (legendre and legendre 1998). this method assumes that all cells not occupied by occurrences are absences; hence its use depends on the quality of data and the completeness of sampling in the region of interest. rangemap_explore visualization creates simple figures to visualize occurrence data on top of a country map. all countries with at least one record are shown in the plot. rangemap_plot visualization produces plots of ranges resulting from analysis functions. species’ ranges, extents of occurrence, and occurrences can be plotted on the same map if needed. other aspects of a map can be added: legend, north arrow, scale bar, and axis values. ranges_emaps visualization plots one or more ranges of a species on one or more raster layers of environmental variables. ranges_espace visualization generates a three-dimensional plot, in environmental space, of ranges created using distinct algorithms. table 1. description of the main functions included in rangemap. more detailed documentation of these functions and other helper functions used to run analyses can be seen at1. https://cloud.r-project.org/web/packages/rangemap/rangemap.pdf marlon e. cobos et al. – rangemap: an r-package to explore species geographic ranges 61 and raster layers). examples of all types of data required to use the analysis functions are included in the package to help users to understand the structure and classes of such data. analysis toolset analysis functions are the ones used to generate estimates of species’ ranges. the five analysis functions use different approaches: buffers, feature selection in polygon layers, convex hulls, concave (alpha) hulls, trend surface analyses, and raster classification (see details in table 1). before creating spatial polygons of species ranges, these functions perform simple steps of data cleaning: (1) erasing duplicates; (2) deleting occurrences lacking coordinates; and (3) excluding records falling outside of a region of interest. by default, the regions of interest are spatial polygons representing country borders (using data included in the package maptools; bivand and lewin-koh 2021), but other spatial polygons can be provided by users if smaller geographic regions need to be considered or if working in marine areas. diverse arguments in the functions allow users to explore different parameterizations in the analyses (e.g., buffer distance, geographic projection, algorithm, occurrence clustering, etc.). other arguments allow users to produce spatial polygons to represent the extent of occurrence (using convex hulls) and the area of occupancy of species according to iucn criteria (iucn standards and petitions committee 2019). these functions also permit saving results directly in a directory, which could help to prevent memory issues if multiple analyses are performed, and the user does not want results to be stored in memory (see ways to use such arguments in the code provided to reproduce examples). all results obtained with analysis functions are returned in objects of class s4, together with information describing how the range polygons were created. users can save these results in multiple formats (according to the class of each of the results) if needed, or results can be used in further analyses. visualization tools four functions of rangemap help to visualize occurrence data and ranges generated with analysis functions (table 1). to start, users can explore the occurrences to be used and roughly identify potential problems with the data (“rangemap_explore”). after obtaining spatial polygons of species’ ranges, another function (rangemap_plot) can be used to produce maps of ranges, occurrences, and, if present in results, extents of occurrence. the generic maps that are obtained can also be modified using arguments from the plotting function that control some graphic and map attributes (e.g., color, legend, north arrow, scale bar, etc.). two special visualization tools in rangemap (“ranges_emaps” and “ranges_espace”; table 1) allow users to represent species’ ranges considering environmental information. these functions help to visualize the environmental implications of using distinct approaches to produce range estimates. environmental conditions are represented in geography using raster layers, and in environmental space using three-dimensional views of variable values present in records and ranges. for representations in environmental space, principal components derived from variables can also be used. to simplify comparisons of ranges in environmental space, ellipsoids (norris et al. 2006; nuñez-penichet et al. 2021), created based on environmental conditions within each range, are used instead of all points. example application the example consists of a step-by-step guide to explore, create, and plot species’ ranges using rangemap. all code to reproduce this example is presented in the supplementary material (file s1). the occurrence data, spatial polygons, and raster layers used in these examples can be obtained using the code provided. most of the data used are included in the rangemap package; this information can be loaded using the code provided and following other available guides (see software availability). data for examples occurrence data for four species were used in our example application: (1) dasypus kappleri (greater long-nosed armadillo), (2) peltophryne fustiger (western giant toad), (3) peltophryne empusa (cuban small-eared toad), and (4) amblyomma americanum (lone star tick). other data used in our examples are in raster format; one of these layers is the result of an ecological niche modeling exercise with the species a. americanum (raghavan et al. 2019); the other raster layers represent bioclimatic variables (hijmans et al. 2005): bio_5 = maximum temperature of marlon e. cobos et al. – rangemap: an r-package to explore species geographic ranges 62 warmest month; bio_6 = minimum temperature of coldest month; bio_13 = precipitation of wettest month; and bio_14 = precipitation of driest month. the spatial polygons used in analyses are prepared using data from the maptools package (bivand and lewin-koh 2021). layers to represent distinct administrative boundaries were downloaded using the function that performs this analysis (information available online2). 2 https://gadm.org/data.html. exploring data and generating range estimates we developed plots to explore the occurrence data using the function “rangemap_explore” (fig. 1). we generated polygons to represent ranges using approaches based on buffers, polygon selection, convex hulls, concave hulls, and trend-surface analysis. to demonstrate the use of ecological niche modeling outputs, we used the data for a. americanum. when ranges were created based on buffers or convex or concave hulls, we used the figure 1. examples of figures that can be produced to perform initial explorations of species occurrence data using the function “rangemap_explore.” https://gadm.org/data.html marlon e. cobos et al. – rangemap: an r-package to explore species geographic ranges 63 following buffer distances: 300 km for d. kappleri; 50 km for p. fustiger; 30 km for p. empusa; and 350 km for a. americanum. ranges generated by feature selection used the following administrative boundaries: countries for d. kappleri; municipalities for p. fustiger; provinces for p. empusa; and states for a. americanum. for the example case using the species d. kappleri and convex hulls, we used a distance of 1500 km to separate data in hierarchical clusters. visualization of ranges after the creation of range estimates, we produced figures to represent results in simple maps using the function “rangemap_plot” (fig. 2; s1–s5). to exemplify how ranges can be explored considering environmental conditions, we used range estimates created for d. kappleri and a. americanum and the functions “ranges_emaps” and “ranges_espace” (fig. 3–4). figures that represent ranges in environmental space were produced using three environmental variables and three principal components derived from four environmental variables (see section data for examples). discussion the tools included in the rangemap package allowed exploration of occurrence data (fig. 1) and distinct options to generate spatial polygons (fig. 2) that help to understand species’ ranges. the plots created with these tools showed the results from analyses and helped visualize environmental characteristics corresponding to species’ ranges (fig. 3–4). geographic patterns like clustering and density of records were easily identified using such explorations (see, e.g., species a. americanum; fig. 1–2). lack of sampling in certain areas of the regions of interest was also clear for d. kappleri (fig. 1–2). the recognition of patterns like clustering, disjunction, and lack of sampling, can also help to identify potential errors in occurrence records, which are rather common in this type of data. using different algorithms to generate range estimates has marked effects on the results that can be obtained (e.g., d. kappleri ranges deriving from buffers and feature selection; fig. 2). changing some parameters in the tools used also affected the outcome of analyses; for instance, allowing the algorithm to recognize hierarchical clusters of occurrence records based on a distance, showed a discontinuity figure 2. examples of simple range estimates generated based on buffers, feature selection of country polygons, concave hulls, a trend surface analysis (tsa), convex hulls, and thresholding of a continuous layer resulting from an ecological niche modeling exercise. in the range of d. kappleri (fig. s3; see also meyer et al. 2017 for applications of disjunct hulls). we recommend using multiple algorithms for range creation and different parameterizations of these tools to perform more complete explorations of the distributional data available. one of the parameters that could be of special importance is the one that sets the distance for buffers in some of our tools. although, in our examples, distances were used for purposes of demonstration only, we encourage users to pick values of distance based on appropriate considerations (e.g., dispersal abilities, home range, etc.). consideration of environmental dimensions adds valuable information when exploring species’ distributions. visualizing environmental conditions in plots and ranges on maps (fig. 3) could help to marlon e. cobos et al. – rangemap: an r-package to explore species geographic ranges 64 figure 3. comparison of ranges of the species dasypus kappleri created with distinct algorithms. ranges are shown on top of four bioclimatic variables. bio_5 = max temperature of warmest month; bio_6 = min temperature of coldest month; bio_13 = precipitation of wettest month; bio_14 = precipitation of driest month. figure 4. visualization of distinct hypotheses of ranges for amblyomma americanum in environmental space. environmental space is represented by three raw environmental dimensions (left) and by the three first principal components (pc) derived from such dimensions (right): bio_5 = maximum temperature of warmest month, bio_6 = minimum temperature of coldest month, bio_13 = precipitation of wettest month. marlon e. cobos et al. – rangemap: an r-package to explore species geographic ranges 65 identify potential biogeographic barriers derived from changes in such conditions in the continuous geographic region of interest. plots produced to represent ranges in environmental space help to visualize effects of considering results from one type of analysis or another. for instance, using buffers of 350 km to create a range for a. americanum translates to wider limits in terms of temperature as compared to a range deriving from selecting the states where the species has been detected (fig. 4). this type of visualization also helps to start understanding conditions used by the species (the occurrences) compared with conditions available nearby. in fact, some of the methods used to generate range estimates with our tools have also been proposed to create areas for model calibration in enm or sdm exercises (e.g., acevedo et al. 2012; gonzalez et al. 2021). this is, generating calibration areas as spatial polygons obtained from buffering points, creating convex or concave hulls, or selecting polygons from layers representing ecoregions, is not uncommon in the enm/sdm literature. however, these areas should be delimited based on more biologicallyrelevant considerations as they represent the regions that have been accessible to the species of interest for relevant periods of time (the m from the bam diagram; barve et al. 2011). for that reason, we recommend caution if the intention is to use our tools to generate such areas, as other, more advanced tools have been developed for this purpose (machadostredel et al. 2021). as previously established, the tools presented here aim to help users to perform initial explorations of species’ ranges. results from tools that use ecological niche modeling outputs could be interpreted as areas with suitable conditions for species (franklin 2010; peterson et al. 2011). however, we recommend caution in interpreting results, as no other ecological processes are considered in creation of species ranges. another consideration when using our tools is that buffers for polygons are based on distances measured on spatial objects converted to the azimuthal equidistant projection, centered on the geographic centroid of the occurrence data. when working with very large areas, distances far from the center will lose precision (snyder 1987). in sum, the rangemap package offers handy options to explore species’ distributions with minimal data requirements. after appropriate considerations of the limitations of the analyses performed to generate range estimates, the information deriving from the toolset presented here could be useful in diverse applications. researchers working on projects dealing with questions related to biogeographic patterns and conservation planning, for instance, could find in rangemap a friendly tool to perform initial but necessary steps to explore and filter occurrence data. although the number of analyses allowed in this package is currently limited, future versions will provide options to explore occurrence data and distribution ranges considering environmental dimensions in more detail. the rangemap package and its dependencies (table s1) are on cran (see software availability). acknowledgments we thank m. weber, the members of the kuenm working group at the university of kansas biodiversity institute, and students from the google code-in program, for their help in betatesting rangemap. students from the google codein program also helped to develop the logo of the package. this r package was developed during and with support from the 2018 google summer of code program; r organization, project “species range maps in r.” competing interests the authors have declared that no competing interests exist. software availability rangemap is available on cran3. other detailed guides to using rangemap can be seen at4. data availability all data used to produce the examples presented here can be obtained with the code provided as part of the example application. supplementary material all supplementary materials can be accessed at5. literature cited acevedo, p., a. jiménez-valverde, j. m. lobo, and r. real. 2012. delimiting the geographical background in species distribution modelling. j. biogeogr. 39:1383–1390. 3 https://cran.r-project.org/package=rangemap. 4 https://marlonecobos.github.io/rangemap/. 5 https://doi.org/10.6084/m9.figshare.16624228.v1. https://cran.r-project.org/package=rangemap https://marlonecobos.github.io/rangemap/ marlon e. cobos et al. – rangemap: an r-package to explore species geographic ranges 66 anderson, r. p. 2015. el modelado de nichos y distribuciones: no es simplemente “clic, clic, clic.” biogeografía 8:4–27. barve, n., v. barve, a. jiménez-valverde, a. lira-noriega, s. p. maher, a. t. peterson, j. soberón, and f. villalobos. 2011. the crucial role of the accessible area in ecological niche modeling and species distribution modeling. ecol. model. 222:1810–1819. bivand, r., and n. lewin-koh. 2021. maptools: tools for handling spatial objects. r package6. clobert, j., m. baguette, t. g. benton, and j. bullock (eds). 2012. dispersal ecology and evolution. oxford university press, croydon, uk. cobos, m. e., l. jiménez, c. nuñez-penichet, d. romeroalvarez, and m. simões. 2018. sample data and training modules for cleaning biodiversity information. biodivers. inform. 13:49–50. eddy, w. f. 1977. a new convex hull algorithm for planar sets. acm trans. math. softw. 3:398–403. everitt, b. 1974. cluster analysis. heinemann educational for social science research council, london. franklin, j. 2010. mapping species distributions: spatial inference and prediction. cambridge university press. cambridge, uk. gonzalez, v. h., m. e. cobos, j. jaramillo, and r. ospina. 2021. climate change will reduce the potential distribution ranges of colombia’s most valuable pollinators. perspect. ecol. conserv. 19:195–206. hardisty, a., and d. roberts. 2013. a decadal view of biodiversity informatics: challenges and priorities. bmc ecol. 13:16. hartigan, j. a., and m. a. wong. 1979. a k-means clustering algorithm. j. royal stat. soc. c 28:100–108. hijmans, r. j., s. e. cameron, j. l. parra, p. g. jones, and a. jarvis. 2005. very high resolution interpolated climate surfaces for global land areas. int. j. climatol. 25:1965– 1978. iucn standards and petitions committee. 2019. guidelines for using the iucn red list categories and criteria. version 147. legendre, p., and l. f. j. legendre. 1998. numerical ecology. 2nd ed. elsevier. amsterdam. machado-stredel, f., m. e. cobos, and a. t. peterson. 2021. a simulation-based method for identifying accessible areas as calibration areas for ecological niche models and species distribution models. front. biogeogr. 13:e48814. meyer, l., j. a. f. diniz-filho, and l. g. lohmann. 2017. a comparison of hull methods for estimating species ranges and richness maps. plant ecol. divers. 10:389–401. taylor & francis. 6 https://cran.r-project.org/package=maptools. 7 http://www.iucnredlist.org/documents/redlistguidelines.pdf. nenzén, h. k., and m. b. araújo. 2011. choice of threshold alters projections of species range shifts under climate change. ecol. model. 222:3346–3354. norris, j. r., s. t. jackson, and j. l. betancourt. 2006. classification tree and minimum-volume ellipsoid analyses of the distribution of ponderosa pine in the western usa. j. biogeogr. 33:342–360. nuñez-penichet, c., m. e. cobos, and j. soberón. 2021. nonoverlapping climatic niches and biogeographic barriers explain disjunct distributions of continental urania moths. front. biogeogr. 13:e52142. park, j.-s., and s.-j. oh. 2012. a new concave hull algorithm and concaveness measure for n-dimensional datasets. j. inf. sci. eng. 28:587–600. peterson, a. t., j. soberón, and l. krishtalka. 2015. a global perspective on decadal challenges and priorities in biodiversity informatics. bmc ecol. 15:15. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. qiao, h., j. soberón, and a. t. peterson. 2015. no silver bullets in correlative ecological niche modelling: insights from testing among many potential algorithms for niche estimation. methods ecol. evol. 6:1126–1136. r core team. 2021. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. radosavljevic, a., and r. p. anderson. 2014. making better maxent models of species distributions: complexity, overfitting and evaluation. j. biogeogr. 41:629–643. raghavan, r. k., a. t. peterson, m. e. cobos, r. ganta, and d. foley. 2019. current and future distribution of the lone star tick, amblyomma americanum (l.) (acari: ixodidae) in north america. plos one 14:e0209082. sillero, n., s. arenas-castro, u. enriquez‐urzelai, c. g. vale, d. sousa-guedes, f. martínez-freiría, r. real, and a. m. barbosa. 2021. want to model a species niche? a step-bystep guideline on correlative ecological niche modelling. ecol. model. 456:109671. snyder, j. p. 1987. map projections: a working manual. u.s. geological survey, washington, d.c. soberón, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodivers. inform. 2:1–10. https://cran.r-project.org/package=maptools http://www.iucnredlist.org/documents/redlistguidelines.pdf biodiversity informatics, 15, 2020, pp. 81-91 81 weighing the evidence for the abundant-center hypothesis tad a. dallas1,*, luca santini2,3, robin decker4, and alan hastings5,6 1department of biological sciences, louisiana state university, baton rouge, usa 2department of environmental science, institute for wetland and water research, faculty of science, radboud university, po box 9010, nl-6500 gl nijmegen, the netherlands 3national research council, institute of research on terrestrial ecossystems (cnr-iret), via salaria km 29.300, 00015, monterontondo (rome), italy 4department of integrative biology, university of texas, austin, usa 5department of environmental science and policy, university of california, davis, usa 6santa fe institute, santa fe, nm, usa abstract. the abundant-center hypothesis posits that species density should be highest in the center of the geographic range or climatic niche of a species, based on the idea that the center of either will be the area with the highest demographic performance (e.g., greater fecundity, survival, or carrying capacity). while intuitive, current support for the hypothesis is quite mixed. here, we discuss the current state of the abundant-center hypothesis, highlighting the relatively low level of support for the relationship. we then discuss the potential reasons for this lack of empirical support, emphasizing the inherent ecological complexity which may prevent the observation of the abundant-center in natural systems. this includes the role of non-equilibrial population dynamics, species interactions, landscape structure, and dispersal processes, as well as variable data quality and inconsistent methodology. the incorporation of this complexity into studies of the distribution of species densities in geographic or niche space may underlie the limited empirical support for the abundant-center hypothesis. we end by discussing potentially fruitful research avenues. most notably, we highlight the need for theoretical development and controlled experimental testing of the abundant-center hypothesis. key words: abundant-center, distance-abundance, macroecology, niche, species distribution. introduction based on observational evidence, early ecologists formulated the hypothesis that species abundance (or perhaps more accurately, species density) should be highest in the center of the species geographic range (geographic abundant-center; brown 1984; sagarin and gaines 2002). this makes the assumption that the geographic range is centered on the niche optimum. one way this could occur is through spatially-autocorrelated environmental conditions, the absence of dispersal limitation, and equilibrial population density within each local population across the landscape. this allows for the translation of geographic performance into niche space, where species growth rates are highest where conditions are the most favorable, coinciding with the center of their geographic distributions (gaston 2009; sagarin and gaines 2002). however, this may be an oversimplification, as it assumes a coherent environmental gradient, a lack of strong density-dependence, equilibrium population dynamics, no dispersal limitation, and the absence (or unimportance) of competition, predation, mutualism, and parasitism in regulating population demography. it is perhaps due to these strong assumptions that the abundant-center hypothesis has received such mixed empirical support (pironon et al. 2017; sagarin et al. 2006; dallas et al. 2017; santini et al. 2019). recently, it was proposed that variation in species density should not be viewed in geographic space, but instead of species’ niche space (niche abundant-center; martı́nez-meyer et al. 2013). this helps to remove the confounding effect of spatial autocorrelation on population processes, and links species density more closely with species abiotic tolerances (i.e., the niche). however, * corresponding author: tad.a.dallas@gmail.com biodiversity informatics, 15, 2020, pp. 81-91 82 few studies of the niche abundant-center hypothesis have been performed to date (but see waldock et al. 2019; dallas et al. 2017; martı́nez-meyer et al. 2013; santini et al. 2019). here, we review evidence for abundant-center ideas, discuss the potential reasons for variability in support, outline existing conceptual issues, and highlight potentially fruitful open questions related to the abundant-center hypothesis. what is the evidence for the abundant-center hypothesis? questions surrounding the spatial distribution of species density are central to population ecology and biogeography, leading to a great interest in understanding the constraints on population size across spatial and environmental gradients (hengeveld and haeck 1982; bell 2000). this interest has generated numerous hypotheses concerning the distribution of abundance across a species geographic range (hengeveld and haeck 1982) e.g., the geographic abundant-center hypothesis. support for the geographic abundant-center hypothesis from large scale analyses can vary from 10% (dallas et al. 2017) to values around 50% (pironon et al. 2017) of examined species. a systematic meta-analysis is outside of the scope of the current review, and is perhaps redundant considering the two other comprehensive reviews on the topic (sagarin and gaines 2002; pironon et al. 2017). numerous studies exist examining geographic abundant-center relationships for single or small set of species (scrosati and freeman 2019; samis and eckert 2007; feldhamer et al. 2012; virgós et al. 2011). these studies tend to find more positive support for abundant-center relationships, potentially as a result of increased sampling effort in these more focused studies, or due to a publication bias toward species providing support for the hypothesis. while the abundant-center hypothesis was first conceptualized with respect to species’ geographic distributions, a recent effort has suggested that species should instead be more dense in the center of their environmental niche space (niche abundant-center; martı́nez-meyer et al. 2013). the underlying idea here is to link the species’ niche to species performance curves from physiological ecology, in order to link individual survival or fecundity to niche axes. this makes intuitive sense, as individuals located towards the middle of the species environmental tolerances may have higher survival or fecundity rates than individuals towards the niche margins (but see pironon et al. 2017) for a thorough discussion of this idea). defining the abundant-center hypothesis in terms of species’ niche distances is a useful step, as it removes the conflation of geographic space and climate, and more closely connects with niche theory (santini et al. 2019). to date, this revised niche abundant-center hypothesis has received approximately the same level of support as the original (~20% of studied species; table 1). for instance, dallas et al. (2017) and santini et al. (2019) found limited support for abundant-center relationships given an examined set of 1419 and 108 species respectively, and support differed depending on if the abundant-center is examined in geographic—15.7% of species in santini et al. (2019) and 13.3% of species in dallas et al. (2017)—or environmental niche space—17.6% of species in santini et al. (2019) and 10.3% of species in dallas et al. (2017). interestingly, some species may also have higher densities on the edge of their geographic distributions (22.2% of the species in santini et al. (2019) and 10% of species in dallas et al. (2017) or climatic niches—12% of species in santini et al. (2019) and 8% of species in dallas et al. (2017). the prevalence of species following the opposite pattern to the predicted abundant-center is difficult to quantify, as previous systematic reviews have not reported this information (sagarin and gaines 2002), or methodological choices prohibit this from even being possible (waldock et al. 2019). a recent study on north american bird species claims to provide strong support for niche abundant-center relationships (osorio-olvera et al. 2020), which varied between 16-45% depending on the approach (osorio-olvera et al. 2020). in table 1, we report the results where the niche was defined as the mve of the first three pca axes, which found support for abundant-center relationships in 200 out of 442 bird species (45% of species), but go into further detail on methodological decisions and abundant-center support elsewhere (dallas, et al. 2020). this level of variation in support may stem from the taxonomic diversity of species analyzed, methodological choices in estimating abundant-center relationships, or due to publication biases promoting evidence in favor of the hypothesis. despite the mixed support and low predictive power of abundant-center relationships in geographic or niche space, there remains an inherent interest in the hypothesis. the sustained interest in the abundant-center hypothesis may stem from the conceptubiodiversity informatics, 15, 2020, pp. 81-91 83 al simplicity and intuitive nature of the relationship. further, the implications of the abundant-center relationship provide a means to estimate areas of high conservation priority using information on species occurrence data (oliveira et al. 2009; manthey et al. 2015; boakes et al. 2018). the ability to estimate species density from species occurrence data is a long-standing goal in ecology, as occurrence data are far more plentiful than species abundance data (ashcroft et al. 2017). however, it remains unclear if models which use species occurrence data to estimate species density can be predictive, as there is mixed evidence for a link between habitat suitability estimated from species distribution models trained on species occurrence data and species densities (weber et al. 2017; ashcroft et al. 2017; dallas et al. 2018; santini et al. 2019). in addition to examining empirical support for abundant-center relationships, linking the degree of support more closely to species ecology is important. specifically, understanding the phylogenetic and trait correlates of the abundant-center relationship may provide insight into the drivers of the observed mixed support for the abundant-center hypothesis. aspects related to species population demography and dispersal processes may influence the extent to which species follow abundant-center relationships, and quantifying these effects may help identify species most (or least) likely to result in abundant-centers. a series of other potential drivers exist, which may contribute to the disconnect between species environmental suitability and species densities. incorporating these potential drivers will provide a degree of ecological realism and may promote further theoretical development and empirical understanding of abundant-center relationships. what are the potential ecological drivers of the variation in support? despite this strong interest in mechanism, studies of the abundant-center hypothesis have traditionally lacked a strong theoretical basis (as discussed in osorio-olvera, soberón, and falconi 2019; holt 2019; dallas and santini 2020), or empirical support (table 1). this suggests that—under the assumption that the relationship can exist in controlled systems—variation in support may be driven by unmeasured environmental variation or factors that influence species density apart from niche limits. it is important to note that even in cases where abundant-center relationships are found, these relationships tend to be extremely weak, suggesting that distance from the niche center cannot be used to predict species densities (dallas and hastings 2018). there are many reasons for this, many of which relate to how the niche and corresponding center are estimated (discussed further in in the next section). here, table 1: a review of existing support for the abundant-center hypothesis suggests that approximately 20% of species demonstrate a relationship between species density and distance to a species geographic range or climatic niche center. source system niche species (n) support (n+) sagarin and gaines (2002) mixed* – 121 56 freeman and beehler (2018) birds – 17 2 waldock et al. (2019) reef fish + 181 56 dallas et al. (2017) birds, mammals, fish, & trees + / – 1419 118 martı́nez-meyer et al. (2013) birds & mammals + / – 11 10 rivadeneira et al. (2010) crabs – 5 2 pironon et al. (2015) plants + / – 3 0 santini et al. (2019) birds and mammals + / – 108 5-20 ** osorio-olvera et al. (2020) birds 442 200*** grand total 2307 464 *: sagarin and gaines (2002) was a meta-analysis combining 22 articles + / – : study tested for abundant-center in both niche and geographic space **: santini et al. (2019) used 9 different distance measures, range of support provided, and the upper value was used in the table totals. ***: here, we report the relationship using the first three pca axes and estimating the niche using minimum volume ellipsoids (mves). see osorio-olvera et al. (2020) for more information, and dallas, pironon, and santini (2020) for a critique. biodiversity informatics, 15, 2020, pp. 81-91 84 we highlight the potential ecological factors which limit the detectability and applicability of the abundant-center relationship to natural populations. these factors, which may drive the low empirical support for abundant-center relationships currently, may be incorporated into future studies, leading to clarifications or refinements on the existing abundant-center hypothesis. stochasticity populations are typically not in equilibrium, but fluctuate over time (lande et al. 2003; black and mckane 2012) due to stochasticity and flow of individuals between populations through immigration and emigration to and from nearby habitat patches. within a single population, demographic and environmental stochasticity move populations away from equilibrium (lande et al. 2003). this is important for at least two reasons. first, demographic stochasticity—variation in species numbers due to probabilistic birth-death processes—influences small populations more strongly, creating inherent spatial variation in species densities over time across the species range. secondly, environmental stochasticity—in which environmental variation influences probabilistic birthdeath processes—may disproportionately influence some populations in a species range more than others (ellner et al. 2016). species’ niche axes may only capture static environmental tolerances, failing to incorporate the role of temporal autocorrelation in environmental conditions and the existence of rare climatic events which can have large impacts on species densities (ellner et al. 2016). this presents a clear area for theoretical development, as existing models used in support of abundant-center theory do not incorporate stochasticity (osorio-olvera et al. 2019), and the incorporation of stochasticity strongly influences resulting abundant-center relationships (dallas and santini 2020). dispersal limitation in geographic space, barriers to dispersal can limit where species occur. reflecting boundaries such as impassable streams serve may influence the spatial distribution of species density by increasing species densities at barriers, and by not allowing access to suitable geographic areas beyond the dispersal barriers (nislow et al. 2011). this can further influence the estimation of a species’ niche, as the niche is commonly defined using data on geographic occurrence points of a species. species dispersal can be informed by social cues, habitat patch quality, and climatic events (clobert et al. 2009; reed et al. 1999; jacob et al. 2015; travis and dytham 1999), suggesting that population density will non-randomly fluctuate, independent of any abundant-center constraints. further, dispersal processes may create spatially-aggregated clusters of individuals which may not truly represent the potential distribution of the species, but are a result of limited dispersal distance (condit et al. 2000). the spatial arrangement of habitat patches, dispersal distance and barriers, and the distribution of populations in geographic or environmental space is also important, as frequent dispersal events may result in spatial synchrony (paradis et al. 1999), potentially influencing support for abundant-center relationships. finally, species with limited dispersal ability may occupy only a small part of their fundamental niche space even in the absence of dispersal barriers. this may still lead to abundant-center relationships, but only if the species has a strong core number of populations aggregated in geographic or environmental space. that is, if a species with poor dispersal ability has disjoint populations, this could create a bimodality in the resulting abundant-center relationship, where species are most abundant in the small regions of geographic or environmental space where populations are established. species interactions species do not exist in isolation from one another, but interact through competition, predation, mutualism, and parasitism. these interactions are geographically variable (early and keith 2019), such that even directly including the density of an interacting species in an abundant-center analysis may not capture the reality of the system. presently, no abundant-center hypothesis study has explicitly accounted for the influence of these antagonistic and mutualistic forces. despite this, ecologists have long recognized that competition (greiner la peyre et al. 2001; lawlor 1979), predator-prey dynamics (blasius et al. 1999; stenseth et al. 1997), mutualistic interactions (holland et al. 2002), and parasitism (hochberg and holt 1990; hudson et al. 1998) can all strongly influence species population dynamics. related to the importance of species interactions to determining local population size, the order of species’ arrival to a site can determine competitive outcomes and strongly influence species density (strauss 1991; biodiversity informatics, 15, 2020, pp. 81-91 85 fukami 2015). however, these priority effects (or historical contingencies) are difficult to incorporate into abundant-center analyses. most often, this is due to the limited data available on the order of species arrival, though the conceptual framework used to test the abundant-center hypothesis also doesn’t provide a clear way to incorporate these dynamic processes into the expected relationship between species density and geographic or niche distance. species traits even if we assume species do exist in relative isolation, responding to environmental variation independently from other species, species traits may influence abundant-center relationships. understanding how species-level abundant-center relationships are influenced by phylogenetic relationships and traits is an important research frontier (dallas et al. 2017; santini et al. 2019). for example, examining north american birds, osorio-olvera et al. (2020) found evidence that body mass, migratory status, and habitat (aquatic or terrestrial) influenced abundant-center relationships, while dallas et al. (2017) failed to detect an effect of body size on abundant-center relationships in birds, trees (measured as height), fishes, or mammals. species traits are expected to influence abundant-center relationships through the lens of demographic performance and distributional effects. that is, traits like body mass or species leaf area may be strongly related to maximum demographic rates (duncan et al. 2007), while traits relating to species dispersal distance (e.g., seed mass) may collectively influence the range and distributions of observable species densities (soudzilovskaia et al. 2013; harpole and tilman 2006). anthropogenic effects human effects on landscapes may strongly influence species abundances through the effects of disturbance, exploitation, and land use changes (sykes et al. 2019; benı́tez-lópez et al. 2010; benı́tez-lópez et al. 2017). changing land use can strongly influence species densities by altering resource availability or shifting community composition, as species will vary in their ability to tolerate human disturbance. however, the response of species densities to human-induced environmental changes can be mixed, as humans can often provide food supplementation, can act as predator deterrent, leading to increased density (sorace 2002; jesse et al. 2018; wang et al. 2015; rouco et al. 2019; šálek et al. 2015), while direct disturbance, exploitation, shifts in competitive community structure, mating opportunities, and reduced access to resources can also reduce population density (lemoine et al. 2007; benı́tez-lópez et al. 2010; benı́tez-lópez et al. 2017). anthropogenic barriers can alter the distribution and abundance of migratory species (said et al. 2016). humans have also altered the size and shape of species geographic ranges (di marco and santini 2015; channell and lomolino 2000), an effect which would obviously influence the detection of the geographic or niche center and thus abundant center-relationships. further, intensity of human effects on a landscape are not likely to be uniformly distributed across a species geographic range or within climatic niche space (sanderson et al. 2002; benı́tez-lópez et al. 2019), leading to considerable variation and potential divergence from abundant-center relationships. incorporating human population density or land use change more explicitly into abundant-center ideas could help disentangle how disturbance can influence the spatial distribution of species densities. however, whether or not these variables represent true niche axes depends on how the niche is conceptually defined, as niche concepts differ in the incorporation of abiotic and biotic axes. potential methodological drivers of variation in support apart from the above mechanisms, there are also a number of existing issues with how support for the abundant-center hypothesis is quantified. for instance, measuring species densities instead of demographic rates assumes that the population density is a proportional representation of demographic processes, which is not necessary in non-equilibrial populations where demographic and environmental stochasticity play a strong role. further, the original formulation of the abundant-center hypothesis did not explicitly operationalize the estimation of “center”. that is, some researchers relate species density to a) the distance from each sampled population to the geographic (or niche) center, b) distance to the geographic (or niche) edge, or c) one of numerous other measures (e.g., santini et al. (2019) use nine different measures). many of these measures require estimation of the center of a species geographic (or niche), as well as the use of an appropriate distance measure to quantify distance of a population from geographic range (or niche) center (dallas et biodiversity informatics, 15, 2020, pp. 81-91 86 al. 2017; soberón et al. 2018). aggregated or incomplete sampling of a species geographic range or climatic niche area then represents a clear issue in the detection of abundant-center relationships (pironon et al. 2017), as this may bias the estimation of the geographic range (or climatic niche) center. distance measures and niche delineation with respect to distance measures, numerous methods have been used, with mahalanobis distance being a current favorite since it incorporates the covariance between spatial or climatic axes (soberón et al. 2018; osorio-olvera et al. 2019). however, it is important to note that some distance measures may be highly correlated (dallas et al. 2018)—while others are not (santini et al. 2019)—suggesting that criticisms based on distance measure used have the potential to be meaningful or not. this creates a clear issue, as different measures may produce drastically different results, and researchers wishing to support abundant-center ideas could simply select a combination of measures which provide the strongest degree of support. further, the methodological decision of how to delineate geographic range (or niche) boundaries is another consideration, as each method of delineating a species geographic range or climatic niche makes assumptions. there are at least two conceptual approaches to this. the first involves training a species distribution model on environmental covariates, and using this model to delineate the species geographic range or climatic niche space. this approach may get around spatial sampling biases and more accurately capture the species’ niche and geographic distribution, but suffers from existing issues in species distribution modeling, such as the identification of relevant niche axes used as covariates, the choice of modeling approach (e.g., presence-background versus presence only approaches, different modelling algorithms (norberg et al. 2019), and thresholding the continuous prediction of species distribution models to binary estimate for range delineation. additionally, niche models attempt to model the grinellian niche (e.g. abiotic component), while the eltonian niche (e.g. biotic component) is generally only partly considered or entirely disregarded (soberón and nakamura 2009). finally, these models rely on the assumption that species are in equilibrium with the environment, but this assumption is often violated biasing niche estimates (faurby and araújo 2018). a second approach simply uses the sampled species occurrence points to delineate the geographic range and corresponding environmental niche space. methods of delineating species range from occurrence data such as convex hulls may overestimate range size and bias range center estimation (soberón et al. 2018), but benefit from being simply defined and not requiring parameterization. for instance, alpha hulls offer another way to delineate range, but they are highly sensitive to parameterization (joppa et al. 2016). lastly, estimation methods tend to assume a particular shape of the niche (e.g., ellipsoidal), while identifying the true shape of the niche would require far more data and far fewer assumptions. defining the niche one of the greatest strengths of the abundant-center hypothesis is the relation of species performance curves along environmental gradients to the concept of the species’ niche and corresponding geographic projection. however, this strength is complicated by the multiple definitions of the species’ niche (leibold 1995; colwell and rangel 2009). some of these niche definitions are at odds with abundant-center ideas. for instance, hutchinson (1957) described the niche as a persistence boundary, not a continuous surface of suitability or demographic performance. even when the niche may be estimated perfectly in -dimensional space, a variety of niche shapes could be estimated, with some resulting in a center that is not actually contained within the niche, or if the niche is discontinuous. this is true for the realized niche, which is the form of the niche that data on the known geographic distributions of species allows us to estimate, and may not be the case for the fundamental niche. for instance, there may be regions of climatic niche space that are not part of the niche, but are contained within niche (blonder 2016). finally, the identification of appropriate niche axes can strongly influence resulting support for abundant-center relationships, as osorio-olvera et al. (2020) demonstrated by defining the niche using all possible combinations of 2-3 climatic niche axes for a set of species. assuming species density, carrying capacity, or growth rates do peak in the center of a species’ niche, there remains the issue of which climatic covariates best define the niche. most studies of niche abundant-center relationships use temperature and precipitation variables to define the niche space, which may only capture mean conditions, and does not include biodiversity informatics, 15, 2020, pp. 81-91 87 information on habitat quality, predator density, or other factors which could strongly influence population density and temporal dynamics. these climatic covariates are able to capture species distributions based on occurrence data well (norberg et al. 2019), but may not be able to accurately estimate species densities (dallas and hastings 2018). data considerations data quality is a common issue affecting the empirical examination of numerous macroecological laws, and the abundant-center hypothesis is no exception. especially in macroecology, data are often difficult to obtain, as spatial and temporal coverage vary considerably, different data sources may not be directly comparable, and the inherent presence of sampling and detection biases (knouft 2018). studies aimed at examining consistency of abundant-center support to methodological decisions are important (santini et al. 2019), as inconsistent methodology can also strongly influence the resulting level of support (or non-support). a critical first step in addressing the influence of methodological decisions is to release all code and data to reproduce the analyses, allowing others to modify the original approach and determine change in support (dallas et al. 2017; dallas and hastings 2018; waldock et al. 2019; osorio-olvera et al. 2020). however, we should also bear in mind that finding a stronger abundant-center relationship does not validate a methodological approach or the quality of a data source, as the degree of support (or non-support) for the abundant-center relationship is not inherently a reflection on data quality or the suitability of a given method. what counts as abundant-center evidence? perhaps a more important current limitation is what we currently consider as evidence supporting the abundant-center idea. some studies have used rank correlation coefficients between distance and density to address the significance and strength of the abundant-center relationship (santini et al. 2019), while others have used a classification approach based on threshold values (waldock et al. 2019). that is, if species density at the range margins is less than some percentage the density of a species in the geographic or niche center, a species is said to follow an abundant-center distribution. this makes a rigorous meta-analysis of support for the abundant-center hypothesis difficult, as thresholds for statistical support differ greatly, and authors quantify “support” in many different ways. what should we do next? in light of these existing issues and the need for further theoretical development, it is worth outlining important next steps in the study of the abundant-center hypothesis. this list is neither exhaustive nor prescriptive, as this area has been, and will likely continue to be, of great interest to ecologists and biogeographers with different research approaches. the continued development of theory (holt 2019; dallas and santini 2020), as well as the use of controlled experimental approaches (e.g., microcosms), represent two useful advances to the study of abundant-center ideas. the combination of observational, experimental, and theoretical approaches will aid in addressing the numerous questions which currently exist, including: 1. how do demographic and environmental stochasticity influence the distribution of species density in geographic or niche space? 2. is there a species trait basis for variation in support for the abundant-center hypothesis? 3. what is the influence of human impacts on the support for the abundant-center hypothesis? 4. what are the relative roles of niche requirements and species interactions (e.g., predation, competition, parasitism) on abundant-center relationships? 5. do mutualistic species tend to both follow abundant-centers with respect to their own niche requirements, or as a function of interactor density (e.g., are plant-pollinator interactions driven by climate or availability of partners)? 6. instead of simple measures of population density, how do demographic rates vary as a function of distance from the geographic range or climatic niche center? (see pironon et al. 2017 for examples of this approach) 7. what is the role of seasonal environmental fluctuations to abundant-center biodiversity informatics, 15, 2020, pp. 81-91 88 ideas? seasonally fluctuating environments will move populations in niche space, resulting in different estimates of distance and variability in the relationship between distance and density. 8. how do abundant-center relationships change temporally? species with strong seasonal dynamics may occupy a different range or attain variable densities as a function of phenology and environmental drivers which may lead to a seasonally variable relationship between species density and distance from the geographic range or climatic niche (dallas and hastings 2018). 9. how does global change alter our ability to detect an abundant-center pattern? changes to environmental conditions shift the niche center in space, yet demographic responses may be lagged behind. but what if the goal of the abundant-center hypothesis is not to predict species densities, but simply to document the decline of species density within a species geographic range or climatic niche (i.e., to document a pattern)? many similar macroecological and biogeographical laws exist, where the current goal is to gauge support for the relationship across different taxa and environments (pironon et al. 2017). this is certainly a worthwhile endeavor, provided the conclusions are tempered by considering the degree of statistical support for abundant-center relationships, which are typically quite low pironon et al. (2017). concluding thoughts it is important to not become dogmatic in the assessment of abundant-center ideas. while intuitive, the mixed support for the hypothesis suggests that other factors are important in structuring the spatial distribution of species densities. even in studies showing support for the abundant-center relationship, the amount of variance in species density explained by the selected distance measure is quite small, which limits the utility of the relationship for prediction, and suggests that other processes may contribute far more strongly to predicting species densities than distance measures. further conceptual and theoretical development is necessary to understand when we would expect to find abundant-center relationships, and what other factors contribute to controlling species density at a given site. to this end, applying spatial population dynamic models is a fruitful path forward to explore the conditions which promote abundant-center relationships in controlled simulations. finally, the use of laboratory systems to test abundant-center ideas is a clear research need, as they allow the ability to define the species’ niche independent of geographic space and to control the amount of variation present in the system. this not only provides a baseline for how strong abundant-center relationships can be when all other environmental variation is ignored, but would also allow for demonstrations of the relative effects of temporal variation in environmental conditions, species interactions, and dispersal dynamics in structuring species densities across geographic or niche gradients. acknowledgments the research was funded by the academy of finland and the jane and aatos erkko foundation. we thank town peterson and an anonymous reviewer for their constructive comments on earlier drafts. this work has been performed with funding to t. dallas from the national science foundation (nsf-deb-2017826). author contributions all authors contributed to manuscript writing. data accessibility there is no data or code associated with this manuscript. conflict of interest the authors have no conflicts of interest to declare. references ashcroft, m.b., d.h. king, b. raymond, j.d. turnbull, j. wasley, and s.a. robinson. 2017. moving beyond presence and absence when examining changes in species distributions. global change biology 23:2929–2940. bell, g. 2000. the distribution of abundance in neutral communities. american naturalist 155:606–617. benı́tez-lópez, a., r. alkemade, a.m. schipper, d.j. ingram, p.a. verweij, j.a.j. eikelboom, and m.a.j. huijbregts. 2017. the impact of hunting on tropical mammal and bird populations. science 356:180–183. benı́tez-lópez, a., r. alkemade, and p.a. verweij. 2010. the impacts of roads and other infrastructure on biodiversity informatics, 15, 2020, pp. 81-91 89 mammal and bird populations: a meta-analysis. biological conservation 143:1307–1316. benı́tez-lópez, a., l. santini, a.m. schipper, m. busana, and m.a.j. huijbregts. 2019. intact but empty forests? patterns of hunting-induced mammal defaunation in the tropics. plos biology 17:e3000247. black, a.j., and a.j. mckane. 2012. stochastic formulation of ecological models and their applications. trends in ecology & evolution 27:337–345. blasius, b., a. huppert, and l. stone. 1999. complex dynamics and phase synchronization in spatially extended ecological systems. nature 399: 354. blonder, b. 2016. do hypervolumes have holes? american naturalist 187: e93–e105. boakes, e.h., n.j.b. isaac, r.a. fuller, g.m. mace, and p.j.k. mcgowan. 2018. examining the relationship between local extinction risk and position in range. conservation biology 32:229–239. brown, j.h. 1984. on the relationship between abundance and distribution of species. american naturalist 124:255–279. channell, r., and m.v. lomolino. 2000. trajectories to extinction: spatial dynamics of the contraction of geographical ranges. journal of biogeography 27:169–179. clobert, j., j.-f. le galliard, j. cote, s. meylan, and m. massot. 2009. informed dispersal, heterogeneity in animal dispersal syndromes and the dynamics of spatially structured populations. ecology letters 12:197– 209. colwell, r.k., and t.f. rangel. 2009. hutchinson’s duality: the once and future niche. proceedings of the national academy of sciences usa 106:19651–19658. condit, r., p.s. ashton, p. baker, s. bunyavejchewin, s. gunatilleke, n. gunatilleke, s.p. hubbell, r.b. foster, a. itoh, j.v. lafrankie, and h.s. lee. 2000. spatial patterns in the distribution of tropical tree species. science 288:1414–1418. dallas, t.a., and a. hastings. 2018. habitat suitability estimated by niche models is largely unrelated to species abundance. global ecology and biogeography 27:1448–1456. dallas, t.a., r.r decker, and a. hastings. 2017. species are not most abundant in the center of their geographic range or climatic niche. ecology letters 20:1526– 1533. dallas, t.a., r.r decker, and a. hastings. 2018. multiple data sources and freely available code is critical when investigating species distributions and diversity: a response to knouft (2018). ecology letters 21:1423–1424. dallas, t.a., s. pironon, and l. santini. 2020. weak support for the abundant niche-center hypothesis in north american birds. biorxiv doi: https://doi. org/10.1101/2020.02.27.968586. dallas, t.a., and l. santini. 2020. the influence of stochasticity, landscape structure, and species traits on abundant-center relationships. ecography 43:1341– 1351. di marco, m., and l. santini. 2015. human pressures predict species’ geographic range size better than biological traits. global change biology 21:2169–2178. duncan, r.p., d.m. forsyth, and j. hone. 2007. testing the metabolic theory of ecology: allometric scaling exponents in mammals. ecology 88:324–333. early, r., and s.a. keith. 2019. geographically variable biotic interactions and implications for species ranges. global ecology and biogeography 28:42–53. ellner, s.p, d.z. childs, and m. rees. 2016. environmental stochasticity. pp. 187-227 in data-driven modelling of structured populations. springer. faurby, s., and m.b. araújo. 2018. anthropogenic range contractions bias species climate change forecasts. nature climate change 8:252. feldhamer, g.a, d.b. lesmeister, j.c. devine, and d.i. stetson. 2012. golden mice (ochrotomys nuttalli) co-occurrence with peromyscus and the abundant-center hypothesis. journal of mammalogy 93:1042–1050. freeman, b.g., and b.m. beehler. 2018. limited support for the ‘abundant center’ hypothesis in birds along a tropical elevational gradient: implications for the fate of lowland tropical species in a warmer future. journal of biogeography 45:1884–1895. fukami, t. 2015. historical contingency in community assembly: integrating niches, species pools, and priority effects. annual review of ecology, evolution, and systematics 46:1–23. gaston, kevin j. 2009. geographic range limits: achieving synthesis. proceedings of the royal society b 276:1395–1406. la peyre, m.k.g., j.b. grace, e. hahn, and i.a. mendelssohn. 2001. the importance of competition in regulating plant species abundance along a salinity gradient. ecology 82: 62–69. hengeveld, r., and j. haeck. 1982. the distribution of abundance. i. measurements. journal of biogeography 9:303–316. biodiversity informatics, 15, 2020, pp. 81-91 90 hochberg, m.e., and r.d. holt. 1990. the coexistence of competing parasites. i. the role of cross-species infection. american naturalist 136:517–541. holland, j.n., d.l. deangelis, and j.l. bronstein. 2002. population dynamics and mutualism: functional responses of benefits and costs. american naturalist 159:231–244. holt, robert d. 2019. reflections on niches and numbers. ecography 43:387–390. hudson, peter j, andy p dobson, and dave newborn. 1998. prevention of population cycles by parasite removal. science 282:2256–2258. hutchinson, g evelyn. 1957. concluding remarks. cold spring harbor symposium on quantitative biology 22:415–427. jacob, s., a.s. chaine, n. schtickzelle, m. huet, and j. clobert. 2015. social information from immigrants: multiple immigrant-based sources of information for dispersal decisions in a ciliate. journal of animal ecology 84:1373–1383. jesse, w.a.m., j.e. behm, m.r. helmus, and j. ellers. 2018. human land use promotes the abundance and diversity of exotic species on caribbean islands. global change biology 24:4784–4796. joppa, l.n., s.h.m. butchart, m. hoffmann, s.p. bachman, h.r. akçakaya, j.f. moat, m. böhm, r.a. holland, a. newton, b. polidoro, and a. hughes. 2016. impact of alternative metrics on estimates of extent of occurrence for extinction risk assessment. conservation biology 30: 362–370. knouft, j.h. 2018. appropriate application of information from biodiversity databases is critical when investigating species distributions and diversity: a comment on dallas et al. ecology letters 21:1119–1120. lande, r., s. engen, and b.-e. saether. 2003. stochastic population dynamics in ecology and conservation. oxford university press. lawlor, l.r. 1979. direct and indirect effects of n-species competition. oecologia 43: 355–364. leibold, matthew a. 1995. the niche concept revisited: mechanistic models and community context. ecology 76:1371–1382. lemoine, n., h.-g. bauer, m. peintinger, and k. böhning-gaese. 2007. effects of climate and land-use change on species abundance in a central european bird community. conservation biology 21:495–503. manthey, j.d., l.p. campbell, e.e. saupe, j. soberón, c.m. hensz, c.e. myers, h.l. owens, k. ingenloff, a.t. peterson, n. barve, and a. lira-noriega. 2015. a test of niche centrality as a determinant of population trends and conservation status in threatened and endangered north american birds. endangered species research 26 (3): 201–8. martı́nez-meyer, e., d. dı́az-porras, a.t. peterson, and c. yáñez-arenas. 2013. ecological niche structure and rangewide abundance patterns of species. biology letters 9:20120637. nislow, k.h., m. hudy, b.h. letcher, and e.p smith. 2011. variation in local abundance and species richness of stream fishes in relation to dispersal barriers: implications for management and conservation. freshwater biology 56:2135–2144. norberg, a., n. abrego, f.g. blanchet, f.r. adler, b.j. anderson, j. anttila, m.b araújo, t.a. dallas, d. dunson, j. elith, and s.d. foster. 2019. a comprehensive evaluation of predictive performance of 33 species distribution models at species and community levels. ecological monographs 89:e01370. oliveira, g. de, j.a.f. diniz-filho, l.m. bini, and t.f.l.v.b. rangel. 2009. conservation biogeography of mammals in the cerrado biome under the unified theory of macroecology. acta oecologica 35:630– 638. osorio-olvera, l., j. soberón, and m. falconi. 2019. on population abundance and niche structure. ecography 42:1415–1425. osorio-olvera, l., c. yañez-arenas, e. martı́nez-meyer, and a.t. peterson. 2020. relationships between population densities and niche-centroid distances in north american birds. ecology letters 23:555-564. paradis, e., s.r. baillie, w.j. sutherland, and r.d. gregory. 1999. dispersal and spatial scale affect synchrony in spatial population dynamics. ecology letters 2:114–120. pironon, s., g. papuga, j. villellas, a.l. angert, m.b. garcı́a, and j.d. thompson. 2017. geographic variation in genetic and demographic performance: new insights from an old biogeographical paradigm. biological reviews 92:1877–1909. pironon, s., j. villellas, w.f. morris, d.f. doak, and m.b. garcı́a. 2015. do geographic, climatic or historical ranges differentiate the performance of central versus peripheral populations? global ecology and biogeography 24:611–620. reed, j.m., t. boulinier, e. danchin, and l.w. oring. 1999. informed dispersal. pp. 189-259 in current ornithology. springer. rivadeneira, m.m., p. hernáez, j.a. baeza, s. boltana, m. cifuentes, c. correa, a. cuevas, e. del valle, i. hinojosa, n. ulrich, and n. valdivia. 2010. testing biodiversity informatics, 15, 2020, pp. 81-91 91 the abundant-center hypothesis using intertidal porcelain crabs along the chilean coast: linking abundance and life-history variation. journal of biogeography 37:486–498. rouco, c., i.c. barrio, f. cirilli, f.s. tortosa, and r. villafuerte. 2019. supplementary food reduces home ranges of european wild rabbits in an intensive agricultural landscape. mammalian biology 95:35–40. sagarin, r.d., and s.d gaines. 2002. the ‘abundant center’ distribution: to what extent is it a biogeographical rule? ecology letters 5 (1): 137–147. sagarin, r.d., s.d. gaines, and b. gaylord. 2006. moving beyond assumptions to understand abundance distributions across the ranges of species. trends in ecology & evolution 21 (9): 524–530. said, m.y., j.o. ogutu, s.c. kifugo, o. makui, r.s. reid, and j. de leeuw. 2016. effects of extreme land fragmentation on wildlife and livestock population abundance and distribution. journal for nature conservation 34:151–164. šálek, m., l. drahnı́ková, and e. tkadlec. 2015. changes in home range sizes and population densities of carnivore species along the natural to urban habitat gradient. mammal review 45:1–14. samis, k.e., and c.g. eckert. 2007. testing the abundant center model using range-wide demographic surveys of two coastal dune plants. ecology 88:1747–1758. sanderson, e.w., m. jaiteh, m.a. levy, k.h. redford, a.v. wannebo, and g. woolmer. 2002. the human footprint and the last of the wild: the human footprint is a global map of human influence on the land surface, which suggests that human beings are stewards of nature, whether we like it or not. bioscience 52:891–904. santini, l., s. pironon, l. maiorano, and w. thuiller. 2019. addressing common pitfalls does not provide more support to geographical and ecological abundant-center hypotheses. ecography 42:696–705. scrosati, r.a., and m.j. freeman. 2019. density of intertidal barnacles along their full elevational range of distribution conforms to the abundant-center hypothesis. peerj 7:e6719. soberón, j., and m. nakamura. 2009. niches and distributional areas: concepts, methods, and assumptions. proceedings of the national academy of sciences usa 106:19644–19650. soberón, j., a.t. peterson, and l. osorio-olvera. 2018. a comment on ‘species are not most abundant in the center of their geographic range or climatic niche.’ rethinking ecology 3:13. sorace, a. 2002. high density of bird and pest species in urban habitats and the role of predator abundance. ornis fennica 79:60–71. soudzilovskaia, n.a., t.g. elumeeva, v.g. onipchenko, i.i. shidakov, f.s. salpagarova, a.b. khubiev, d.k. tekeev, and j.h.c. cornelissen. 2013. functional traits predict relationship between plant abundance dynamic and long-term climate warming. proceedings of the national academy of sciences usa 110:18180–18184. harpole, w.s., and d. tilman. 2006. non-neutral patterns of species abundance in grassland communities. ecology letters 9:15–23. stenseth, n.c., w. falck, o.n. bjørnstad, and c.j. krebs. 1997. population regulation in snowshoe hare and canadian lynx: asymmetric food web configurations between hare and lynx. proceedings of the national academy of sciences usa 94:5147–5152. strauss, s.y. 1991. indirect effects in community ecology: their definition, study and importance. trends in ecology & evolution 6:206–210. sykes, l., l. santini, a. etard, and t. newbold. 2019. effects of rarity form on species’ responses to land use. conservation biology 34:688–696. travis, j.m.j., and c. dytham. 1999. habitat persistence, habitat availability and the evolution of dispersal. proceedings of the royal society of london b 266:723–728. virgós, e., r. kowalczyk, a. trua, a. de marinis, j.g. mangas, j.m. barea-azcón, and e. geffen. 2011. body size clines in the european badger and the abundant center hypothesis. journal of biogeography 38:1546–1556. waldock, c., r.d. stuart-smith, g.j. edgar, t.j. bird, and a.e. bates. 2019. the shape of abundance distributions across temperature gradients in reef fishes. ecology letters 22: 685–696. wang, y., m.l. allen, and c.c. wilmers. 2015. mesopredator spatial and temporal responses to large predators and human development in the santa cruz mountains of california. biological conservation 190:23–33. weber, m.m., r.d. stevens, j.a.f. diniz-filho, and c.e.v. grelle. 2017. is there a correlation between abundance and environmental suitability derived from ecological niche modelling? a meta-analysis. ecography 40:817–828. biodiversity informatics, 13, 2018, pp. 38-48 38 from theory to practice: a photographic inventory of museum collections to optimize collection management jonas merckx1, martijn van roie1,2,*, jesús gómez-zurita3, wouter dekoninck4 1biodiversity inventory for conservation (binco) vzw, walmersumstraat 44, 3380 glabbeek, belgium. 2department of biology, ecosystem management research group, university of antwerp, universiteitsplein 1, 2610 wilrijk, belgium. 3animal biodiversity and evolution, institut de biologia evolutiva (csic-universitat pompeu fabra), pg. marítim de la barceloneta 37, 08003 barcelona, spain. 4royal belgian institute of natural sciences, vautierstraat 29, 1000 brussels, belgium. *corresponding author, martijnvanroie@hotmail.com abstract.—the digitization of museum specimens is a key priority in the digital era. digital databases help to avoid unnecessary manipulation hazards to delicate collections, increase their accessibility to third party researchers, and contribute to the ongoing documentation of global biodiversity. time, workforce and the need of specialized infrastructures limit the processing of the vast number of specimens in natural history collections. cheaper, easy-to-use methods and volunteer programs are developing quickly to help bridge the gap. we present the results of combining citizen science for the digitization of an entomological collection in conjunction with the cooperation of a taxonomic expert for the remote identification of samples. in addition, we provide an assessment of the avoided monetary costs and the time needed for each step of the process. a photographic inventory of specimens belonging to the leaf beetle genus calligrapha was compiled by volunteers using a low-cost compact camera and the species were identified using these images. using digital photographs allowed for a rapid screening of specimens in the collection and resulted in an updated taxonomic identification of the calligrapha collection at the royal belgian institute of natural sciences. the pictures of the specimens and their original labels, as well as the new information from this endeavor were placed in an online public catalogue. this study demonstrates a worked example of how digitization has led to a practical, useful outcome through cooperation with an end user and highlights the value of museum collection digitization projects. key words: mass digitization, volunteers, collections, stakeholder engagement, calligrapha introduction in the digital era, one of the challenges that museums must face is being able to sort and digitize the vast number of specimens stored in their repositories. many different approaches are being implemented toward this end, including projects run by museum staff (mathys et al. 2015) and projects which rely on volunteers (holmes 2003; flemons and berents 2012). considering the dimension of the problem and the typically reduced number of staff in museums, the latter is a promising solution. however, the digitization approach is still dependent on the infrastructure and equipment available to the institution, which can be expensive (brecko et al. 2014) and generally not available to a sufficient number of volunteers. recently, the image quality and speed necessary for macro photography have improved significantly in compact camera technology (see e.g. pratt 2015), allowing for low-budget approaches to museum collection digitization (mertens et al. 2017). thus, the combination of cheaper and easy-to-use methods and volunteer programs holds the potential for fast and qualitative mass digitization. photographic inventories of specimen collections and curated collection of individual specimen photographs with any accompanying taxonomic or geographic data hold great potential for the documentation of biodiversity. such inventories aid in species recognition and identification and provide with permanent digital copies of the specimens for future generations (joger 2018). this could increase accessibility of museum collections to third party researchers via online platforms or shared digital storages, and augment the documentation of biodiversity in general (beaman and cellinese 2012). one traditional way to sort and identify collection material involves loaning of relevant specimens to an expert, usually upon the expert request and related to a systematic revision, and restoring the specimens to the collection once they have been studied. although jonas merckx et al. – from theory to practice 39 convenient for the expert, there are several disadvantages to consider: (i) it takes time, effort and costs of personnel to sort the material to loan, prepare the parcels and carry out the administrative procedures, including loan forms and registries documenting the loan and subsequent follow-up; (ii) shipping specimens incurs the risk of losing parcels, damaging or even destruction of often irreplaceable and valuable specimens; and (iii) the shipped material might not all be relevant for the expert and, in any case, any of the specimens on loan remain unavailable for other experts. alternatively, the experts can visit the institute, but it is generally a less cost-effective option since one needs to consider expenses related to accommodation and travel. although these traditional approaches have advantages for taxonomic research, including the possibility to study the specimens up close for good taxonomic practice, screening of specimens for collection sorting via a photographic inventory can avoid shipment of specimens and avoid unnecessary research visits. even though several articles have been published about the advantages and approaches to specimen digitization (blagoderov et al. 2012; smith and blagoderov 2012), few focus on examples where digitization has led to a practical outcome by cooperation with an end user (barber et al. 2013). in this article, we show the results of implementing crowd sourcing in an efficient manner to help the digitization of entomological collections in immediate cooperation with a taxonomic expert for collection organization. a coupled digitization-identification workflow can create a mutually beneficial platform whereby the institute gains immediate feedback on species taxonomy and the researchers (including third party researchers and the public) have easy access to specimen data, while avoiding unnecessary risks to specimens for collection organization (instead of taxonomic research). we describe the results of the digitization of calligrapha specimens in the form of a specimen list and dorsal photographs of the specimens, as well as an assessment of the monetary cost avoided with the use of specimen digitization and give an estimation of the time needed for each step of the process. the applicability of such a workflow is discussed for other insect collections including other taxonomic groups. methods taxonomic choice to assess the feasibility of a digitization-identification workflow, we specifically selected a species-rich taxon with species that can be identified reliably based on photographic material. we therefore opted to digitize the specimens of the leaf beetle genus calligrapha chevrolat (chrysomelidae) from the collections in the royal belgian institute of natural sciences (rbins), see figure 1. the genus calligrapha currently comprises some 130 species and it is distributed from alaska to argentina (montelongoand gómez-zurita 2014). most species in this genus display a striking character, a contrasting pattern of intricate dark and pale markings on the elytra. these markings have recently been shown to hold strong systematic value (montelongo and gómez-zurita 2014). for the purposes of this project, the elytral markings are extremely useful for the identification of the majority of species, thus eliminating the need for genitalia dissection, the typical ‘gold-standard’ for beetle species identification (richmond et al. 2016). every specimen in the collection was identified remotely using digital photographs by one of the authors (jgz), an expert in the taxonomy of the group. specimen digitization for an overview of the followed digitization flow, see figure 2. every specimen was photographed from dorsal, lateral and frontal views using an olympus tg4 digital camera with focus stacking functionality. this rugged point-and-shoot camera has an in-camera focus-stacking feature with two stacking methods: internal stacking (in which a stacked picture is constructed by an internal stacking algorithm in the camera, based on a total of 10 pictures) and focus bracketing (in which the camera can take up to 29 pictures, which have to be stacked manually afterwards by means of external software). this camera is on the market for approximately us$400, some eightfold cheaper than a professional setup for photo-stacking used in documentation of collection specimens (for a comprehensive review of this method and a comparison with a professional canon-cognisys setup, see mertens et al. 2017). since the focus-bracketing feature allows for sharper pictures (mertens et al., 2017), we chose to use this method over internal stacking. pictures of the specimens were manually stacked using helicon focus software (heliconsoft ltd., kharkiv, ukraine), which has a one-time license fee of approximately us$ 120. original specimen labels were photographed and digitized as well, although for the majority of the specimens, some important information (such as jonas merckx et al. – from theory to practice 40 locality, collecting date and collector) was missing, as is often the case for labels of historical insect voucher specimens. despite the lack of this information, the catalogue of these specimens and their identifications are included in the results for archiving purposes. photographs were made available to jgz via a google drive folder where all specimen photographs and label data were given specific ids by the taxonomist. after identification, the pictures and data were made publicly available online at the rbins online database1. workflow assessment to help assessing the cost-effectiveness of the digitization-identification workflow, we quan1http://collections.naturalsciences.be/ssh-entomo/collections/ be-rbins-calligrapha-chevrolat-1836. tified the amount of time needed for each of the steps described above (taking photographs, manual stacking, label digitization, species identification), as well as the overall cost of the project (e.g. equipment material, transportation cost of the volunteers). results workflow assessment in all, 529 specimens of calligrapha stored in the rbins were digitized (figure. 3 shows the dorsal views of one specimen per species). as there were four photographs taken per specimen, 2,116 pictures were taken in total. this number is a conservative estimation as some specimens had to be photographed a second time (e.g., because some parts of it were out of focus). it required two volunteers (working separately) an estimated 100 working hours figure 1. whole-drawer photograph of one of the trays with calligrapha specimens as found in the royal belgian institute of natural sciences. http://collections.naturalsciences.be/ssh-entomo/collections/be-rbins-calligrapha-chevrolat-1836 http://collections.naturalsciences.be/ssh-entomo/collections/be-rbins-calligrapha-chevrolat-1836 jonas merckx et al. – from theory to practice 41 or 12 working days, to take all the photographs. the manual stacking using the helicon software took eight hours, and label digitization another 14 hours. for most specimens in the collection (80%), identification only required looking at the dorsal pictures and the whole process of downloading the picture, identifying the species and writing down the identification in an excel file took approximately 15-20 seconds per specimen. the remaining specimens required checking lateral pictures as well, to confirm identifications based on elytral spots close to the border of the elytron; in these cases, the whole process took an additional 5 seconds to finish. we estimate that the net amount of time devoted to species identification and cataloguing of the entire collection took approximately 3 hours of work. the highest monetary costs for the project were related to the acquisition of the required equipment, namely the camera setup (us$400) and a helicon software license (us$120; see mertens et al. (2017) for a detailed cost description). however, this was a once-off investment rather than a recurring expense which should be amortized relatively easily, as the equipment may be used in subsequent projects. additional costs were related to the transportation of volunteers to the museum (in our case, us$170, but this is highly dependent on mode of transportation, distance traveled, age of volunteers, and other factors). for the steps of the procedure involving specimen identification, the cost for one expert to travel to brussels to do the identification in situ (avoided thanks to our digitization setup) would have not been trivial, and it is estimated as: (i) a plane return ticket barcelona-brussels (us$150); and (ii) two/three days accommodation for the visiting taxonomist (us$500-800). this cost may even have to consider the equivalent of three days salary for an expert assessment when taxonomic work has a service component. these avoided costs are very much circumstantial, but could be easily estimated in the range of several hundred dollars per day (ca. us$300-350). moreover, several other non-trivial costs involving curatorial work for the collection and institution that houses the specimens were avoided. these are estimated as: (i) a salary for one full day to search, select and prepare 529 specimens (ca. us$250-300); and (ii) the cost to send the parcel by registered airmail (in our case, to spain, this would be estimated at us$200, but it depends on the chosen package delivery service). in summary, we estimated the avoided costs for this project to be in the range of us$1400-1800. catalogue of calligrapha at rbins in the original classification of calligrapha at the rbins, 39 taxa were considered. however, after the remote taxonomic reassessment of the specimens by the specialist, the new total was of 47 species and figure 2. scheme of the digitization-identification workflow. after an expert offers help with collection management, museum staff screens the collections for relevant material. after recruitment of volunteers, the digitization step starts with brushing off the specimens and taking pictures of the labels and all relevant angles of the specimens (here dorsal-lateral-frontal). resulting stacked pictures get adjusted if needed and labels get digitized in a database. after the expert receives the pictures, she/he can remotely aid in identifications and collection sorting. any more unidentifiable specimens can then be shipped to the expert if needed. additionally, all of the data can be made available online instantly, where they can be double-checked by public review. jonas merckx et al. – from theory to practice 42 figure 3. dorsal photographs of the species of calligrapha in the rbins collections. (a) c. aeneopicta; (b) c. alni; (c) c. amelia; (d) c. ancoralis; (e) c. annulata; (f) c. apicalis; (g) c. argus; (h) c. bajula; (i) c. bidenticola; (j) c. californica coreopsivora; (k) c. confluens; (l) c. curvilinea; (m) c. dislocata; (n) c. diversa; (o) elegantula; (p) c. felina; (q) c. fulvipes; (r) c. geographica; (s) c. ignara; (t) c. knabi; (u) c. labyrinthica; (v) c. limbaticollis; (w) c. lunata; (x) c. lunata hybrida; (y) c. matronalis; (z) c. multiguttata; (a) c. multipunctata bigsbyana; (b) c. notatipennis; (c) c. nupta; (d) c. pantherina; (e) c. percheroni; (f) c. philadelphica; (g) c. polyspila; (h) c. praecelsis; (i) c. pruni; (j) c. ramulifera; (k) c. rhoda; (l) c. rowena; (m) c. scalaris; (n) c. serpentina; (o) c. serpentina temaxensis; (p) c. sigmoidea; (q) c. spiraea; (r) c. sponsa; (s) c. suffriani; (t) c. verrucosa; (u) c. vicina; (v) c. vigintimaculata. jonas merckx et al. – from theory to practice 43 two subspecies (excluding five specimens which cannot as yet be identified to species level). the discrepancies between the old and new classifications were due to misidentifications (33.46%), naming of unidentified specimens (0.01%) and correction of old, invalid names (0.04%) rather than altered taxonomic status. we provide a catalogue of the specimens in the collection to the lowest taxonomical level as possible. the original labels can be found online at the rbins online database. calligrapha aeneopicta stål, 1859 (fig. 3a) chile: 2 specimens, coll. chapuis. mexico: 1 specimen jalapa, höge. 1 specimen, genin. 1 specimen, ghiesbrecht. unknown: 1 specimen, 1885, restité. 3 specimens, coll. duvivier. calligrapha alni schaeffer, 1928 (fig. 3b) canada: 3 specimens, env. de quebec, provanchers. 2 specimens, saguenay, v. huart. usa: 3 specimens, bayfield wisconsin, wickham. 1 specimen, d. leconte, coll. chapuis. 1 specimen, coll. chapuis. calligrapha amelia knab, 1909 (fig. 3c) usa: 1 specimen, california, 1937, f. heylemans. 2 specimens, massachusetts. 3 specimens, springfield, massachusetts, dimmock. calligrapha ancoralis stål, 1860 (fig. 3d) mexico: 1 specimen, ventanas, durango, höge. calligrapha annulata jacoby, 1903 (fig. 3e) bolivia: 1 specimen, restité, 1885. calligrapha apicalis notman, 1919 (fig. 3f) unknown: 1 specimen, coll. duvivier. calligrapha argus stål, 1859 (fig. 3g) guatemala: 4 specimens, chacoj, vera paz, champion. mexico: 3 specimens, coll. chapuis. unknown: 1 specimen. calligrapha bajula stål, 1860 (fig. 3h) el salvador: 1 specimen, el boqueron (san salvador), vii/1959, j. béchyné. calligrapha bidenticola brown, 1945 (fig. 3i) canada : 1 specimen, ottawa, vi/1949, r. de ruette. usa: 2 specimens, illinois. 1 specimen, massachusetts. 1 specimen, new york, coll. chapuis. 1 specimen, philadelphia, coll. chapuis. unknown: 1 specimen, restité, 1885. 1 specimen, candèze, coll. chapuis. calligrapha californica coreopsivora (linell, 1896) (fig. 3j) usa: 1 specimen, illinois, coll. chapuis. calligrapha confluens schaeffer, 1928 (fig. 3k) canada: 2 specimens, quebec, provanchers. 2 specimens, saguenay, 1876, v. huart. 2 specimens, saguenay, 1877, v. huart. 13 specimens, saguenay, v. huart. usa: 1 specimen, massachusetts. unknown: 1 specimen, coll. duvivier. calligrapha curvilinea stål, 1859 (fig. 3l) peru: 3 specimens, coll. chapuis. unknown: 1 specimen, coll. duvivier. calligrapha dislocata (rogers, 1956) (fig. 3m) brasil: 1 specimen, a. bau, coll. duvivier. mexico: 2 specimens, ocoyoacac, coll. l. legiest. 1 specimen, pachuca, hidalgo, höge. usa: 2 specimens, texas, coll. chapuis. unknown: 1 specimen, coll. duvivier. calligrapha diversa stål, 1859 (fig. 3n) costa rica: 1 specimen, coll. chapuis. guatemala: 1 specimen, capetilla, g.c. champion. mexico: 8 specimens, guanajuato, e. dugès. 2 specimens, xocomanatlan guerrero 7000ft., vii, h.h. smith. 1 specimen, 1891, e. dugès. unknown: 1 specimen, van lansberg. 3 specimens, coll. duvivier. 1 specimen, coll. chapuis. 5 specimens, restité, 1885. calligrapha elegantula jacoby, 1877 (fig. 3o) america: 1 specimen, coll. duvivier. brasil: 1 specimen, coll. camille van volxem. central america: 2 specimens, weyers. costa rica: 1 specimen, irazu 6-7000ft.[sic], h. rogers. 1 specimen, van patten, coll. duvivier. 1 specimen, van patten. unknown: 1 specimen, coll. thirot. 1 specimen, restité, 1885. calligrapha felina stål, 1860 (fig. 3p) mexico: 4 specimens, central states, 1891, e. dugès. 4 specimens, guanajuato, e. dugès. 1 specimen, coll. duvivier. unknown: 3 specimens, coll. duvivier. 1 specimen, restité, 1885. calligrapha fulvipes stål, 1859 (fig. 3q) belize: 9 specimens, ventes, 1893, j.c. stevens. brasil: 1 specimen, coll. camille van volxem. central america: 17 specimens, weyers. el salvador: 40 specimens, el boqueron (san saljonas merckx et al. – from theory to practice 44 vador), vii/1959, j. béchyné. guatemala: 1 specimen, purula, vera paz, champion. 2 specimens, 1943, j. rodriguez. honduras: 1 specimen, r. sarstoon & b. blancaneau. mexico: 1 specimen, oaxaca, hoege. 1 specimen, 1954, coll. thirot. 1 specimen, ghiesbrecht. 5 specimens, coll. thirot. 2 specimens, coll. chapuis. unknown: 1 specimen, coll. thirot. 3 specimens, coll duvivier. 2 specimens, coll. chapuis. 3 specimens, restité, 1885. calligrapha geographica stål, 1860 (fig. 3r) unknown: 1 specimen, coll. chapuis. 1 specimen, coll. duvivier. calligrapha ignara stål, 1860 (fig. 3s) bolivia: 1 specimen, chuquisaca, j. bechyné, coll. chapuis. calligrapha knabi brown, 1940 (fig. 3t) canada: 1 specimen, un. laval, prov. quebec. calligrapha labyrinthica stål, 1859 (fig. 3u) mexico: 1 specimen, coll. chapuis. nicaragua: 1 specimen, chontales, t. belt. unknown: 1 specimen, 1885, restité. 2 specimens, coll. chapuis. 2 specimens, coll. duvivier. 2 specimens, coll. thirot. 2 specimens. calligrapha limbaticollis stål, 1859 (fig. 3v) mexico: 2 specimens, coll. chapuis. 1 specimen, 2012. unknown: 1 specimen, 1885, restité. 2 specimens, coll. duvivier. calligrapha lunata (fabricius, 1787) (fig. 3w) america: 1 specimen, restité, 1. usa: 1 specimen, massachusetts. calligrapha lunata hybrida (fabricius, 1787) (fig. 3x) unknown: 4 specimens, iv/1923, f.s. carr, coll. ant. ball. calligrapha matronalis erichson, 1847 (fig. 3y) peru: 1 specimen, chanchamayo, 1896, ventes j.c. stevens. 2 specimens, chanchamayo. 1 specimen, duvivier. south-america: 1 specimen, rio santiago ’29, marquis de wavria. unknown: 1 specimen, coll. duvivier. 1 specimen. calligrapha multiguttata stål, 1859 (fig. 3z) mexico: 1 specimen, j. bechyné, coll. chapuis. calligrapha multipunctata bigsbyana (kirby, 1837) (fig. 3a) canada: 17 specimens, ottawa, v/1949, r. de ruette. 3 specimens, ottawa, vi/1949, r. de ruette. 1 specimen, ottawa, vii/1949, r. de ruette. 1 specimen, env. de quebec, provanchers. 1 specimen, un. laval, prov. quebec. 1 specimen, saguenay, v. huart. 1 specimen, l. burgeon, coll. l. burgeon. usa: 1 specimen, illinois, 1954, coll. chapuis. 1 specimen, springfield, massachusetts, dimmock. 1 specimen, springfield, massachusetts. 1 specimen, massachusetts. 2 specimens, coll. schramm. 1 specimen, coll. thirot, 1. unknown: 1 specimen, mont. roch, coll. chapuis. 1 specimen, coll. chapuis. 1 specimen, 1954, coll. duvivier. 3 specimens, coll. duvivier. 1 specimen, ct. calligrapha notatipennis stål, 1859 (fig. 3b) mexico: 1 specimen, jalapa, 1954, hoege. 1 specimen, jalapa, hoege. 1 specimen, jalapa, dr. a. fenyes. 1 specimen, ghiesbrecht. 2 specimens, coll. duvivier. 1 specimen, coll. chapuis. chile: 1 specimen. unknown: 3 specimens, 1885, restité. 1 specimen, roelofs. 2 specimens, coll. chapuis. 1 specimen, coll. duvivier. 2 specimens, coll. chapuis. 1 specimen. calligrapha nupta stål, 1859 (fig. 3c) colombia: 7 specimens, le bas. el salvador : 1 specimen, el boqueron, vii/1959, j. bechyné. central america: 1 specimen, weyers. unknown: 1 specimen, 1885, restité. calligrapha pantherina stål, 1859 (fig. 3d) america: 1 specimen, chevrolat. colombia: 1 specimen, 1954, coll. chapuis. guatemala: 1 specimen, candèze, coll. chapuis. 1 specimen, rodriguez. 2 specimens, coll. chapuis. mexico: 1 specimen, presidio, forrer. 1 specimen, amula, guerrero 6000 ft., h.h. smith. 1 specimen, coll. duvivier. unknown: 2 specimens, 1885, restité. 3 specimens, coll. chapuis. 4 specimens, coll. duvivier. 1 specimen, coll. thomson. calligrapha percheroni (guérin-méneville, 1830) (fig. 3e) bolivia: 1 specimen, santos varros, 2000m. ecuador: 1 specimen, 1971, e. de ville, coll. chapuis. 1 specimen, 1971, e. de ville. 1 specimen, coll. chapuis. peru: 1 specimen, coll. chapuis. unknown : 1 specimen, 1954, coll. duvivier. 1 jonas merckx et al. – from theory to practice 45 specimen, 1885, restité. 2 specimens, coll. duvivier. 1 specimen, coll. chapuis. 3 specimens. calligrapha philadelphica (linnaeus, 1758) (fig. 3f) canada: 1 specimen, ottawa, vii/1948, r. de ruette. 5 specimens, ottawa, v/1949, r. de ruette. 1 specimen, ottawa, vi/1949, r. de ruette. 3 specimens, saguenay, v. huart. 1 specimen, coll. chapuis. 1 specimen, coll. l. burgeon. usa: 1 specimen, illinois, coll. chapuis. 1 specimen, springfield, massachusetts, dimmock. unkown: 1 specimen, 1885, restité. 2 specimens, coll. duvivier. 1 specimen. calligrapha polyspila (germar, 1821) (fig. 3g) argentina: 2 specimens, buenos aires, palerma, coll. j. muller. 1 specimen, buenos aires coll. thirot. 2 specimens, buenos aires, coll. camille van volxem. 6 specimens, buenos aires. 1 specimen, east-chavarria, corrientes, xii/1904. 2 specimens, san nicolas, 1928, dr. ch. michel. brasil: 3 specimens, rio janeiro, c. hygin furey. 1 specimen, de lacerda. 2 specimens, santa catarina, 1893, ach. geilenkeuser. 1 specimen, bahia, de lacerda. 1 specimen, bahia, coll. chapuis. 1 specimen, thereyspolis. 5 specimens, de segueira. 2 specimens, roelofs. 1 specimen, coll. j. muller. 7 specimens. guyana: 1 specimen, p. mabile. uruguay: 3 specimens, montevideo, coll. chapuis. usa: 1 specimen, pennsylvania, 1920. south america: 1 specimen. unknown: 1 specimen, dr. a. breyer. 2 specimens, de segueira. 3 specimens, coll. chapuis. 4 specimens, coll. de borre. 1 specimen, coll. duvivier. 5 specimens, coll. j. muller. 6 specimens, coll. h. d’udekem d’acoz. 6 specimens. calligrapha praecelsis (rogers, 1856) (fig. 3h) usa: 1 specimen, kansas, 1885, restité. 1 specimen, kansas, coll. duvivier. unknown: 2 specimens. calligrapha pruni brown, 1945 (fig. 3i) usa: 1 specimen, illinois, duvivier. calligrapha ramulifera stål, 1859 (fig. 3j) guatemala: 1 specimen, coll. chapuis. unknown: 1 specimen, coll. duvivier. 1 specimen. calligrapha rhoda knab, 1909 (fig. 3k) usa: 2 specimens, illinois, coll. chapuis. unknown: 2 specimens, coll. duvivier. calligrapha rowena knab, 1909 (fig. 3l) canada: 1 specimen, saguenay, v. huart. 1 specimen, un. laval, prov. quebec. usa: 3 specimens, ithaca, new york, viii/1928, a. ball. calligrapha scalaris (le conte, 1824) (fig. 3m) canada: 1 specimen, montreal, v/1936, coll. j. muller. 2 specimens, saguenay, 1877, v. huart. mexico: 1 specimen, 1994, coll. f. heylemans. usa: 2 specimens, cambridge, massachusetts. 1 specimen, massachusetts. 1 specimen, coll. schram. unknown: 1 specimen, mont. roch, coll. chapuis. 1 specimen, 1885, restité. 2 specimens, coll. duvivier. calligrapha serpentina (rogers, 1856) (fig. 3n) mexico: 14 specimens, guanajuato, e. dugès. 1 specimen, mexico city, höge. 1 specimen, north sonora, morrison. 1 specimen, queretaro, dr. palmer. 1 specimen, saltillo, coahuila, dr. palmer. 1 specimen, genin. 1 specimen, reitter. 3 specimens, coll. f. heylemans. 2 specimens. usa: 3 specimens, albuquerque, new mexico. unknown: 2 specimens, coll. chapuis. 1 specimen. calligrapha serpentina temaxensis bechyné, 1952 (fig. 3o) mexico: 1 specimen, temax, north yucatan, gaumer. calligrapha sigmoidea (leconte, 1859) (fig. 3p) usa: 2 specimens, california, coll. duvivier. unknown: 1 specimen, 1885, restité. 2 specimens, coll. chapuis. 1 specimen, coll. duvivier. calligrapha ?simillima stål, 1860 unknown: 1 specimen, 1885, restité. 1 specimen, coll. chapuis. calligrapha spiraea (say, 1826) (fig. 3q) usa: 1 specimen. unknown: 1 specimen, coll. duvivier. 1 specimen, coll. wellens. calligrapha sponsa stål, 1859 (fig. 3r) panama: 1 specimen, chiriqui 25-4000 ft.[sic], champion. calligrapha suffriani jacoby, 1882 (fig. 3s) mexico: 2 specimens, central states, e. dugès. 1 specimen, ventanas, durango, höge, coll. duvivier. jonas merckx et al. – from theory to practice 46 calligrapha verrucosa (suffrian, 1858) (fig. 3t) unknown: 1 specimen, coll. chapuis. calligrapha vicina schaeffer, 1933 (fig. 3u) canada: 2 specimens, ottawa, v/1949, r. de ruette. 1 specimen, env. de quebec, provanchers. calligrapha vigintimaculata (chevrolat, 1833) (fig. 3v) guatemala: 3 specimens, capetillo, g.c. champion. 1 specimen, coll. chapuis. mexico: 1 specimen, caracas, bihimit, coll. duvivier. 1 specimen, roelofs, ghiesbrecht. unknown: 1 specimen, restité, 1885. 1 specimen, coll. chapuis. 1 specimen. discussion this project allowed for the digitization and taxonomic reassessment of 529 specimens of calligrapha present in the rbins collection. this allowed all species names to be updated and established that the collection contains 47 species and two subspecies of calligrapha, although initially only 39 species were listed, with 33.5% of specimens bearing incorrect names. despite of the genus calligrapha not being taxonomically challenging in general, this high proportion of misidentifications follows a consistent trend which has emerged after working on the collections of 17 museums in europe and north america visited for the systematic revision of this genus (gomez-zurita 2015, 2016). the results of this work now give access to a complete and reassessed catalogue with pictures of all the specimens and their labels. the digital interface produced by this project will allow taxonomists to easily consult the database and study specimens for more than one third of the known species of calligrapha. based on the experience gained in this project, we note some necessary prerequisites for successful volunteer-based digitization projects. (i) the first one would be engaging committed volunteers to digitize the specimens in a standardized way as a key aspect to ensure that the delicate and repetitive work required can be done efficiently and effectively. (ii) also, the availability of an expert in the taxonomy of the taxon of interest is essential to the correct identification and classification of the digitized specimens. (iii) the taxon targeted by the digitization initiative must be amenable to identification based on visible external structures. finally, (iv) a cost-effective online interface is needed to archive digitized pictures and facilitate remote consultation of the specimens (hill et al. 2012). the most crucial point in our workflow was recruiting capable volunteers. this task can be fairly challenging, but using institutional websites as platforms to promote the importance of digitization and its associated positive implications for biodiversity documentation can encourage recruitment. the importance of direct promotion through professional and amateur entomological meetings and nature enthusiast forums cannot be overstated in their potential to recruit volunteers. there is also the potential for projects to be launched directly by entomological societies such as the royal belgian entomological society2. increased time spent on a project translates into the perception of increased effort and cost, which are factors to take into account when recruiting volunteers or choosing to volunteer for a task (e.g. andow et al. 20163). during our work, two working volunteers processed all the specimens. with more cameras available (e.g. by renting or sponsorship), this procedure could be accelerated to accommodate multiple volunteers working simultaneously. a whole collection could be digitized in a relatively short amount of time. for the vast majority of animal groups, the simple approach adopted in this project is not feasible, since species identification may require dissection of specimens when no obvious external diagnostic characters are visible. in our work, this affected, for instance, the group of cryptic species ranked under c. scalaris, which includes several north american taxa that cannot be confidently identified by external characteristics without knowledge about their host plant (brown 1945; gómez-zurita 2015). as information on the plants where these beetles were collected was not recorded, the identification of these rbins specimens is compromised. although these specimens could not be identified to species level, they could be reliably classified as members of a species complex, an improvement on the original collection catalogue. as such, taxa that are difficult to classify or require dissection are not excluded from approaches like the one developed here, they only require additional volunteer power to prepare and photograph serial mounted slides of the structures of interest such as genitalia of lepidoptera or other insect groups like 2http://www.srbe-kbve.be. 3http://www.serviceleader.org/virtual. http://www.srbe-kbve.be http://www.serviceleader.org/virtual jonas merckx et al. – from theory to practice 47 boopidae4. in these cases, however, it might be more practical to send the specimens to an expert taxonomist. by using this digitization-identification approach to museum specimen identification, the transportation and accommodation costs associated with enlisting a taxonomist can be avoided. the cost of the photographic setup is a once-off expense that is amortized easily with the use of the equipment in subsequent projects. in terms of time spent, finishing the project took approximately 125 hours or 18 working days. in reality, the project took several months (due to the time schedule of the volunteer, corrections of mistakes in photographs, etc.) during which time the authors were pursuing other research endeavors. if time is a constraint, the digitization time as described in the methods section can be considered the only drawback of the digitization approach. however, the digitization approach not only allows for the same kind of taxonomic assessment that could be achieved working directly with the collection, but also provides with an online photographic inventory of calligrapha species for researchers and the public not able to visit the museum. democratizing the collections through digital archives is one of the most important goals of initiatives like ours, offering a solution to one serious conundrum traditionally faced by museums, i.e. that their collections, patrimony of the society, have to be secluded in order to protect and conserve them (e.g. milroy and rozefelds (2015)). this democratization is intrinsically associated to an opportunity for discussion and also puts in place a control mechanism for the inflexibility of the principle of authority, so strongly linked to taxonomic practice. lastly, the workflow proposed here avoids the risks of specimen loss or damage associated with shipping them to a taxonomist. in addition, the risk of infection by fungi, anthrenus spp. and other pest species is avoided as the specimens do not leave the conservation room where they are stored at ideal climatic conditions and under ipm integrated pest management protocol. this study highlights the importance of museum specimen digitization and demonstrates that very basic tasks like organizing a collection can be done remotely. good communication between the institute, digitizing volunteers and taxonomists results in more profitable research visits and avoids the hazards associated with specimen loans thus making 4https://digivol.ala.org.au/project/index/21451805. specimen digitization a valuable tool in facilitating museum collection management. acknowledgments we greatly acknowledge pol limbourg for logistic support and koen martens for help with collection digitization. jan mertens is acknowledged for providing the figures in the manuscript. furthermore, we thank the scientific service heritage (patrick semal) of the royal belgian institute for natural sciences, brussels for the logistic support to put the pictures online through the virtual collections platform. alex laking and alexandra evans are acknowledged for grammatic and semantic editing of the manuscript. lastly, we want to thank jonathan brecko for a first stage review of the paper, and the reviewers are thanked for their valuable feedback. references andow, d. a., e. borgida, t. m. hurley, and a. l. williams. 2016. recruitment and retention of volunteers in a citizen science network to detect invasive species on private lands. environmental management 58:606-618. barber, a., d. lafferty, and l. r. landrum. 2013. the salix method: a semi-automated workflow for herbarium specimen digitization. taxon 62:581-590. beaman, r. s., and n. cellinese. 2012. mass digitization of scientific collections: new opportunities to transform the use of biological specimens and underwrite biodiversity science. zookeys:7-17. blagoderov, v., i. kitching, l. livermore, t. simonsen, and v. s. smith. 2012. no specimen left behind: industrial scale digitization of natural history collections. zookeys 209:133-146. brecko, j., a. mathys, w. dekoninck, m. leponce, d. vandenspiegel, and p. semal. 2014. focus stacking: comparing commercial top-end set-ups with a semi-automatic low budget approach: a possible solution for mass digitization of type specimens. zookeys 464:1-23. brown, w. j. 1945. food-plants and distribution of the species of calligrapha in canada, with descriptions of new species (coleoptera, chrysomelidae). canadian entomologist 77:117-133. flemons, p., and p. berents. 2012. image based digitisation of entomology collections: leveraging volunteers to increase digitization capacity. zookeys:203-217. gómez-zurita, j. 2015. systematic revision of the genus calligrapha chevrolat (coleoptera: chrysomelidae, https://digivol.ala.org.au/project/index/21451805 jonas merckx et al. – from theory to practice 48 chrysomelinae) in central america: the group of calligrapha argus stal. zootaxa 3922:1-71. gómez-zurita, j. 2016. systematic revision of calligrapha chevrolat (coleoptera: chrysomelidae) with pale spots on dark elytra and description of two new species. zootaxa 4072:61-89. gómez-zurita, j. 2015. what is the leaf beetle calligrapha scalaris (leconte)? breviora 541:1-19. hill, a., r. guralnick, a. smith, a. sallans, r. gillespie, m. denslow, j. gross, z. murrell, t. conyers, p. oboyski, j. ball, a. thomer, r. prys-jones, j. de la torre, p. kociolek, and l. fortson. 2012. the notes from nature tool for unlocking biodiversity records from museum records through citizen science. zookeys 209:219-233. holmes, k. 2003. volunteers in the heritage sector: a neglected audience? international journal of heritage studies 9:341-355. joger, u. 2018. research collections in germany: modern trends in methods of sorting, preserving, and research. pp. 17-28 in l. a. beck, ed. zoological collections of germany: the animal kingdom in its amazing plenty at museums and universities. springer international publishing, cham. mathys, a., j. brecko, d. vandenspiegel, l. cammaert, and p. semal. 2015. bringing collections to the digital era three examples of integrated high resolution digitisation projects. pp. 155-158. 2015 digital heritage. mertens, j., m. van roie, j. merckx, and w. dekoninck. 2017. the use of low cost compact cameras with focus stacking functionality in entomological digitization projects. zookeys 712:141-154. milroy, a., and a. rozefelds. 2015. democratizing the collection: paradigm shifts in and through museum culture. australasian journal of popular culture 4:115-130. montelongo, t., and j. gómez-zurita. 2014. multilocus molecular systematics and evolution in time and space of calligrapha (coleoptera: chrysomelidae, chrysomelinae). zoologica scripta 43:605-628. pratt, s. 2015. mirrorless, dslr or point and shoot: which camera is best for macro photography?, https://digital-photography-school.com/mirrorlessdslr-or-point-and-shoot-which-camera-is-best-formacro-photography. richmond, m. p., j. park, and c. s. henry. 2016. the function and evolution of male and female genitalia in phyllophaga harris scarab beetles (coleoptera: scarabaeidae). journal of evolutionary biology 29:2276-2288. smith, v. s., and v. blagoderov. 2012. bringing collections out of the dark. zookeys 209:1-6. biodiversity informatics, 16, 2021, pp. 1-19 1 predicting multi-species bark beetle (coleoptera: curculionidae: scolytinae) occurrence in alaska: openaccess big gis-data mining to provide robust inference khodabakhsh zabihia,b,*, falk huettmanna, and brian d. youngc a ewhale lab, biology and wildlife department, institute of arctic biology, university of alaska fairbanks (uaf), fairbanks, ak 99775-7000, usa b extemit-k, faculty of forest and wood sciences, czech university of life sciences (culs) prague, kamýcká 129, 165 00 praha 6-suchdol, czech republic c stem department, landmark college, putney, vt 05346, usa abstract. native bark beetles (coleoptera: curculionidae: scolytinae) are a multi-species complex that ranks among the key disturbances of coniferous forests of western north america. many landscape-level variables are known to influence beetle outbreaks, such as suitable climatic conditions, spatial arrangement of incipient populations, topography, abundance of mature host trees, and disturbance history that includes former outbreaks and fire. we assembled open-access data for understanding the ecology of bark beetles in alaska. we used boosted classification and regression trees as a machine-learning data-mining algorithm to predict relationships between 838 occurrence records of 68 bark beetle species and 14 environmental variables, compared to pseudo-absence locations across alaska. environmental variables included topographyand climate-related predictors as well as feature proximities and anthropogenic factors. we were able to model, predict, and map multi-species bark beetle occurrences across alaska at 1-km spatial resolution: about 16% of the mixed forest and 59% of evergreen forest are expected to be occupied by the bark beetles based on current climatic conditions and biophysical landscape attributes. the open-access dataset that we prepared, and the machine learning modeling approach that we used, can provide a foundation for future research not only on scolytines but for other multi-species questions of concern, such as forest defoliators, and wildlife species assemblages worldwide. key words: scolytinaes, pest insects, outbreaks, boosted classification and regression tree, forest ecology, spatial modeling, machine-learning algorithm. introduction historic background forests are a major terrestrial ecosystem of global relevance encompassing about 30% of the land area on the earth (schmitt et al. 2009; liu et al. 2018). forest ecosystems play a critical role in ecological services and reducing the threat of natural disasters, such as floods, droughts, and landslides (uy and shaw 2012). at the global scale, forests can mitigate climate change impacts via carbon sequestration and rainfall infiltration that safeguards drinking water supplies (uy and shaw 2012). leaves and needles play major roles in such equations, but they can be consumed or defoliated by insects, specifically bark beetle species. native bark beetles (coleoptera: curculionidae: scolytinae) consist of several species complexes acting together as a community, and constitute one of the key disturbances of coniferous forests of western north america (bentz et al. 2010; seidl et al. 2014; morris et al. 2016). in spite of traditional singlespecies views and subsequent efforts regarding the ‘bark beetle problem’ (e.g., bentz et al. 2009), bark beetles can be perceived as a community organism that together plays a wider role as ecosystem engineers (martikainen et al. 1999; jonasova and pracha 2004; müller et al. 2008). since 1990, bark beetles have killed billions of coniferous trees across millions of hectares in north american forest ecosystems from mexico to alaska (raffa et al. 2008; bentz et al. 2009; bentz et al. 2010). although insects such as bark beetles constitute an inherent element of ecosystems, initiating succession and providing food sources for predators such as woodpeckers, bark beetle outbreaks can cause severe direct and indirect impacts on forest ecosystems and species that are dependent on or that interact with forests. direct effects include tree mortality * corresponding author biodiversity informatics, 16, 2021, pp. 1-19 2 and the ensuing changes in forest composition and structure, increased chance of wildfire owing to creation of broad areas with large quantities of dead trees, and a greater chance of windthrow of living residual trees (schowalter 2012). indirect impacts can include, for example, reduction of timber value resulting from accelerated salvage harvest activities following outbreaks, decreased carbon sequestration, degradation of fish and wildlife habitats, and reduced recreational capacity (schowalter 2012). rapid widespread tree mortality leaves longer-term impacts on structural and functional aspects of ecosystems with ongoing influences on climate, habitats, species, and land use (kurz et al. 2008; mcdowell et al. 2008; bentz et al. 2010). from the ecological and biodiversity point of view, insects such as bark beetles represent an inherent part of the wider ecological and global web, but usually are seen as pests to be gotten rid of. an increasing body of evidence suggests that bark beetle outbreaks and post-outbreak conditions form part of the wider successional pattern in a landscape, and can have some positive impacts on ecosystem services. for example, tree loss can increase water yield (bearup et al. 2014; morris et al. 2016), improve foraging for livestock and wildlife species, and thus increase the population of some big game species and provide more wildlife viewing and hunting opportunities (saab et al. 2014; morris et al. 2016). bark beetles also serve as major prey species for many insectivores, particularly for woodpeckers (bonnot et al. 2009; saab et al. 2014). in addition, bark beetles have fascinating ecological relationships with various fungal species (munro et al. 2019). several studies found that bark beetle outbreaks appear to occur when factors such as drought, aging, and density attenuate trees’ resistance abilities against bark beetle attacks (christiansen and bakke 1988; raffa 1988; fettig et al. 2007; bentz et al. 2010). although treeand stand-level characteristics such as tree vigor and size and stand density are critical for local infestation under endemic conditions (raffa and berryman 1983; simard et al. 2011), landscapelevel factors influence the transition of infestation from endemic to epidemic conditions: a local eruption to regional outbreaks (wallin and raffa 2004; raffa et al. 2008; simard et al. 2011). strong correlations with environmental parameters provide local patches of suitable habitats that enhance potential for scolytine growth; this growth can lead to outbreaks under certain circumstances (aukema et al. 2006; raffa et al. 2008). the spread of outbreaks is spatially and temporally autocorrelated regardless of host tree vigor (aukema et al. 2006; aukema et al. 2008; simard et al. 2011). several landscape-level variables facilitate scolytine beetle outbreaks, such as suitable climatic conditions, spatial arrangement of incipient populations, topography, abundance of mature host trees, and disturbance history that includes outbreak and fire history (aukema et al. 2006). identifying these variables could aid in understanding the epidemiology and ecology of these species and improve management strategies for outbreak control (simard et al. 2011). rationale as bark beetle issues are ecologically complex and multidimensional, computational approaches such as data mining, machine learning, and geographic information system (gis) could be helpful to study and map the current and future spatial distributions of bark beetles. use of these methods is on the rise (bhattacharya 2013), which is particularly needed in forest ecology and management. they have been recently applied in quantitative modeling of macroscale ecological niches of tree species (prasad 2018), mapping aboveground biomass of trees within the alaskan boreal forest (young et al. 2018), inference questions regarding wildlife (huettmann et al. 2018b), and variance assessment of predicted climate-scapes based on topographic variation (huettmann 2018a). however, in a high-dimensional multivariate (i.e., 10-20 predictors) investigation, tools such as data-mining and machine-learning have not yet been applied to understand current and future presence of scolytines, considered as a community organism. this point is particularly true for broad-scale prediction and mapping for a large region such as alaska, which contains part of the largest forest area in the world (the boreal forest) and the world’s largest temperate rainforest (tongass). machine-learning, in contrast to parametric methods such as regression models, can use many algorithms in ensembles (humphries et al. 2018). algorithms are developed to comprehend complex, nonlinear relationships in the data without requiring prior model assumptions such biodiversity informatics, 16, 2021, pp. 1-19 3 as normally distributed model residuals and freedom from spatial autocorrelation (huettmann 2018c). natural phenomena do not necessarily follow known distributions and generally have no linear interactions (ver hoef et al. 1993). in addition, in regression models including logistic regression, use of residual analysis as a measure of model fitness is unreliable when modeling a high-dimensional data set, with nonlinear combinations of the variables (breiman 2001). as such, to fulfill model assumptions and have easily interpretable information in parametric models, a widely used but faulty approach is to weed out less-important predictor variables and consider effects of unmeasured variables as “noise” (breiman 2001). this approach can lead to wrong conclusions and consequently poor management strategies, more a reflection of the model’s mechanism rather than a true emulation of nature (breiman 2001). in contrast, algorithmic models such as boosted classification and regression tree are not affected by limited and biased sets of predictors or by nonlinear relationships among predictors. this characteristic results in more accurate and informative conclusions (breiman 2001; elith et al. 2008; tyralis 2019), very important for biodiversity conservation and forest management. the commonly-used, single-species approach of species distribution modeling assumes that species respond individually to environmental gradients. however, species distributions may be influenced by biotic interactions, such as competition, with other species within a community (araújo and luoto 2007; heikkinen et al. 2007; chapman and purse 2011). in addition, as different species within a community may have similar responses to environmental changes at regional scales (golicher et al. 2008; azeria et al. 2009; chapman and purse 2011), communitylevel analyses of spatial patterns of biodiversity may provide beneficial information (chapman and purse 2011) for biodiversity conservation and natural resource management. for example, several bark beetle species, e.g., polygraphus rufipennis and pityophthorus nitidulus, attack the same host trees (e.g., white spruce and black spruce). in this sense, the common, individual-based approach of modeling and mapping spatial distributions of bark beetle species may not provide comprehensive information required for efficient forest management. in addition, the individual-species-based approach may not help in understanding the broader, overall diversity of bark beetle species. in this study, we set out to map forest landscapes that could favor presence of multi-species bark beetles, with attention to other, more poorly studied landscape reservoirs, such as arctic tundra shrublands. the resulting map of bark beetle species occurrence will provide useful resources for forest managers and policy makers to prioritize their forest management measures spatially in a timeand costeffective manner. we assess many environmental variables and proxies, as well as tabulated metadata for 68 species of bark beetles that together comprise a large dataset that we make openly accessible to the scientific community and the broader public. we seek to develop an example of how large-scale environmental data-mining on a big set of species presence and pseudo-absence locations, machinelearning, and geographic information systems can help to predict and map presence of scolytines, as a community organism. we develop this work in alaskan landscapes, forested and non-forested, without prior model assumptions or consequent data perturbation resulting from violations of assumptions. material and methods study area the study was conducted across the state of alaska, united states, covering a total area of ~152 x 106 ha (figure 1) ranging from approximately 54º to 71º n and 130º w to 173º e (figure 1). within figure 1. hillshade map of the study area, the state of alaska, along with 838 locations of 68 bark beetle species, the presence points for developing and validating the three models, provided by the university of alaska museum (uam; white circles). a different record of 68 locations of 3 bark beetle species, the presence points for additionally assessing/testing the three models, surveyed by u.s. department of forest service (usfs; black circles). see appendix i for species details. biodiversity informatics, 16, 2021, pp. 1-19 4 that area, ~49 x 106 ha (~32%) is forested (defined as areas with >10% tree cover; hutchison 1968, adfg 2018). most of the forested area (~43 x 106 ha) is in interior alaska, and is classified as “boreal forest.” the remaining forestland, ~5 x 106 ha, occurs along the southern coast of alaska and is classified as “coastal temperate rainforest,” which covers the regions of kodiak, prince william sound, and the islands and mainland of south alaska, including the world’s largest temperate rainforest in southeast alaska (tongass) and the chugach national forest in south-central alaska (adfg 2018). coniferous communities of boreal forests are dominated by spruce trees including white spruce (picea glauca (monech) voss) and black spruce (p. mariana (miller) britton, sterns & poggenburg) (adfg 2018). the coastal temperate rainforest includes western hemlock (tsuga heterophylla (rafinesque) sargent), sitka spruce (p. sitchensis (bongard) carrière), mountain hemlock (t. mertensiana (bongard) carrière), alaska yellow cedar (chamaecyparis nootkatensis (d. don) farjon & d.k. harder), western red cedar (thuja plicata (donn. ex d. don)), and lodgepole pine (pinus contorta (douglas ex loudon)) (adfg 2018). alaska has a glacial history, including many refugia. the glaciated regions of alaska include 11 mountain ranges: coast mountains, saint elias mountains, chugach mountains, kenai mountains (including montague island), aleutian range, wrangell mountains, talkeetna mountains, alaska range, wood river moun tains, kigluaik mountains, and the brooks range (usgs 2017). the heavily glaciated alaska range spans an arc ~965 km long that extends from the alaska-canada border towards the alaska peninsula. mount mckinley (6195 m), known as denali, is the highest mountain in the alaska range and north america (usgs 2017), with extensive alpine and forest cover. within the study region, large-scale tree mortality is often caused by different insect species and diseases, with bark beetle species being among the important elements (usda 2008). for example, a massive mortality event that covered >1.3 x 106 ha during 1990–1999 was caused by spruce beetles (usda 2008). research design the year the species were identified varied from 1953 to 2018, with nearly half coming from after 2011 (figure 2), which justifies our use of 2011 land cover data. our inferences and conclusions regarding alaska’s potential habitats for the bark beetle organism therefore focus primarily on the recent period after 2010. this study follows a study design that has been applied previously for predictive modeling of the distribution of white spruce [picea glauca (monech) voss] and small mammals across alaska (ohse et al. 2009, baltensperger and huettmann 2015a). we compiled 838 records of 68 bark beetle species, as the occurrence points for the model, provided by the university of alaska museum (uam1; figure 1 and appendix i) in a separate record for each species. the occurrence data we use are not balanced by species nor do they come from a systematic sampling, but they still serve as a ‘presence record’ of the bark beetle community in alaska. in absence of detailed knowledge about many of the bark beetle species of north american landscapes, this compilation helps to shed light on habitat preferences and dynamics of bark beetles in general. among the pooled bark beetle species, dominant genera were dryocoetes (n = 133), trypodendron (n = 107), ips (n = 104), and dendroctonus (n = 83) (appendix i). the most common species were trypodendron lineatum (n = 74), dryocoetes affaber (n = 69), dendroctonus rufipennis (n = 66), and ips perturbatus (n = 52) (appendix i). the host evergreen trees from which the bark beetle specimens were collected included white spruce (picea glauca), black spruce (p. mariana), sitka spruce (p. sitchensis), western hemlock (tsuga heterophylla), lodgepole pine (pinus contorta), mountain hemlock (t. mertensiana), lutz spruce (p. x lutzii), tamarack (larix laricina), yellow cedar (cupressus nootkatensis), and western redcedar (thuja plicata) (appendix i). some specimens were collected within non-evergreen forests. we extracted land cover types underlying 838 beetle locations 1 http://arctos.database.museum/specimensearch.cfm. figure 2. frequency distributions of the year of observation for the 838 bark beetle presence records. http://arctos.database.museum/specimensearch.cfm biodiversity informatics, 16, 2021, pp. 1-19 5 using the raster map of the national land cover database (nlcd 2011; table 1).23 following ohse et al. (2009), baltensperger and huettmann (2015a), and young et al. (2018), we created a background dataset with which to compare bark beetle presences, from across alaska. we established 5000 random point locations using arcmap 10.4 (esri inc., redlands, ca), with a minimum euclidean distance from each other of 1 km across alaska. the 838 bark beetle presence locations and 5000 random points were then used to extract underlying pixel values of the environmental variables, as independent variables, for comparison with the binary response variable of presence/ absence of bark beetles. we created a lattice point grid with a 1-km euclidean distance, in qgis (version 3.4.0), for the entire study area, for a total number of 1,522,655 points. the lattice points were used to extract underlying environmental variable values for mapping and predicting bark beetle occurrences across the study area based on the model. for additional model assessment, we used 68 independently surveyed locations of 3 bark beetle species—spruce beetle (dendroctonus rufipennis), western balsam bark beetle (dryocoetes confusus swaine), and northern spruce engraver (ips perturbatus)—with sample sizes of 57, 8, and 3, respectively, from surveys conducted by the u.s. department of forest service in 2016–2017 (figure 1). the host trees of these observed bark beetle species included white spruce (picea glauca 2 https://mrlc.gov/data/legends/national-land-cover-database-2011-nlcd2011-legend. 3 open water is a pixel and includes coastal and island locations. (monech) voss), subalpine fir (abies lasiocarpa), and sitka spruce (p. sitchensis (bong.) carrière). environmental variables to assess environmental requirements of bark beetles quantitatively, we used 14 independent variables (table 2). we used aspects of mean monthly maximum and minimum temperature and total precipitation as our climate data. slope and aspect maps were derived from the available 2 arc-second digital elevation model (dem) at 60 m spatial resolution provided by united states geological survey (usgs). we used a soil map unit that aggregates different soil components, including multiple soil classes and miscellaneous areas delineated together in spatial polygons. generally, one to four soil series (or soil taxonomic classes), along with non-soil areas, named as miscellaneous areas, are attached to each map unit (nauman and thompson 2013). we also used the land cover raster map of nlcd (2011) that contains the following classes: water, including open water and perennial ice/snow; developed area, including open space, low-, medium-, and high-intensity developed regions; barren land; forest, including deciduous, evergreen, and mixed forest; shrubland, including dwarf shrub and shrub/scrub classes; grassland/herbaceous; emergent herbaceous; and woody and non-woody wetlands (table 1). land status, indicating ownership of the land, was drawn from a vector file provided by the bureau of land management (2013), and had the following categories: private or municipal, state, bureau of land management, native, national park service, nlcd 2011 designated no. land cover type number of bark beetles found percent of bark beetles found 11 open water3 47 5.6 12 perennial ice/snow 5 0.6 21 developed, open space 7 0.8 22 developed, low intensity 70 8.4 23 developed, medium intensity 39 4.7 24 developed, high intensity 8 1.0 31 barren land 28 3.3 41 deciduous forest 68 8.1 42 evergreen forest 251 30.0 43 mixed forest 81 9.7 51 dwarf shrub 62 7.4 52 shrub/scrub 74 8.8 71 grassland/herbaceous 16 1.9 90 woody wetlands 62 7.4 95 emergent herbaceous wetlands 19 2.3 table 1. number and percent of bark beetle specimens collected at different geographic locations calculated by different land cover types derived from nlcd 20112 https://mrlc.gov/data/legends/national-land-cover-database-2011-nlcd2011-legend https://mrlc.gov/data/legends/national-land-cover-database-2011-nlcd2011-legend biodiversity informatics, 16, 2021, pp. 1-19 6 variable type spatial resolution d esignated nam e o riginal source secondary source soil m ap unit vector and categorical n /a soilakab https://m rlc.gov/data http://hdl.handle.net/11122/10870 land status vector and categorical n /a landstatakab https://sdm s.ak.blm .gov/sdm s/ http://hdl.handle.net/11122/10871 n ational land cover data in 2011 r aster and categorical 30 m landc o11akab https://m rlc.gov/data http://hdl.handle.net/11122/10872 m ean annual precipitation in 2010 r aster and continuous 60 m tem pakab https://uaf-snap.org/ http://hdl.handle.net/11122/10873 m ean annual tem perature in 2010 r aster and continuous 60 m precipakab https://uaf-snap.org/ http://hdl.handle.net/11122/10874 d igital elevation m odel (d em ) r aster and continuous 60 m d em 60m akab https://ned.usgs.gov http://hdl.handle.net/11122/10875 slope r aster and continuous 60 m slope60m akab http://hdl.handle.net/11122/10876 a spect r aster and continuous 60 m a spect60m akab http://hdl.handle.net/11122/10877 euclidean distance to coastline r aster and continuous 60 m d istc oast https://gis.data.alaska.gov http://hdl.handle.net/11122/10878 euclidean distance to lakes and rivers r aster and continuous 60 m d istlaker iver https://usgs.gov/ http://hdl.handle.net/11122/10879 euclidean distance to drainage netw ork r aster and continuous 60 m d istd rin et https:// usgs.gov/ http://hdl.handle.net/11122/10880 euclidean distance to tow ns r aster and continuous 100 m d isttow ns https://gis.data.alaska.gov/ http://hdl.handle.net/11122/10881 euclidean distance to m ain roads r aster and continuous 100 m d istr oads https://gis.data.alaska.gov/ http://hdl.handle.net/11122/10882 euclidean distance to infrastructure r aster and continuous 100 m d istinfrastruct https://gis.data.alaska.gov/ http://hdl.handle.net/11122/10883 m odel results type spatial resolution d esignated nam e o riginal source secondary source m odel 1 (a ppendix iii) r aster and continuous 1 km m odel1 http://hdl.handle.net/11122/10921 m odel 2 (figure 3) r aster and continuous 1 km m odel2 http://hdl.handle.net/11122/10922 m odel 3 (a ppendix iv ) r aster and continuous 1 km m odel3 http://hdl.handle.net/11122/10923 b inary m ap (figure 6) r aster and categorical 1 km b inarym ap http://hdl.handle.net/11122/10924 final m ap (figure 7) r aster and categorical 1 km finalm ap http://hdl.handle.net/11122/10925 a dditional data type spatial resolution d esignated nam e o riginal source secondary source b ark b eetle presence points vector and non-categorical n /a b b eetleakab http://arctos.database.m useum / http://hdl.handle.net/11122/10927 b ark b eetle pseudo-absence points vector and non-categorical n /a r apointsakab http://hdl.handle.net/11122/10928 1-km space grid points vector and non-categorical n /a g ridpointsakab http://hdl.handle.net/11122/10929 b ark b eetle assessm ent/test points vector and non-categorical n /a b b eetlevalidakab https://fs.usda.gov/ http://hdl.handle.net/11122/10930 table 2. d escription of the environm ental variables and additional dataset used to develop and test the m odels, the three m odel results as raster m aps, and their open-access sources. http://hdl.handle.net/11122/10870 http://hdl.handle.net/11122/10871 http://hdl.handle.net/11122/10872 http://hdl.handle.net/11122/10873 http://hdl.handle.net/11122/10874 https://ned.usgs.gov http://hdl.handle.net/11122/10875 http://hdl.handle.net/11122/10876 http://hdl.handle.net/11122/10877 https://gis.data.alaska.gov http://hdl.handle.net/11122/10878 http://hdl.handle.net/11122/10879 http://hdl.handle.net/11122/10880 http://hdl.handle.net/11122/10881 http://hdl.handle.net/11122/10882 http://hdl.handle.net/11122/10883 http://hdl.handle.net/11122/10921 http://hdl.handle.net/11122/10922 http://hdl.handle.net/11122/10923 http://hdl.handle.net/11122/10924 http://hdl.handle.net/11122/10925 http://hdl.handle.net/11122/10927 http://hdl.handle.net/11122/10928 http://hdl.handle.net/11122/10929 https://www.google.com/search?q=https://fs.usda.gov/&spell=1&sa=x&ved=2ahukewjh0mbw_yzxahwpk4skhacvaywqbsgaegqiaraw http://hdl.handle.net/11122/10930 biodiversity informatics, 16, 2021, pp. 1-19 7 national wildlife refuge, military, national forest service, and state and native together. we also derived raster maps from vector files of coastline, lakes and rivers, drainage network, towns, main roads, and infrastructures, using euclidean distance toolbox in arcmap. the drainage network represents more ephemeral water channels, whereas lakes and rivers encompass more permanently standing or flowing water. the most dominant infrastructures in the state, based on the measured length, were pipelines, gas lines, winter trail, transmission line, highway, tractor trail, old railroad, and summer trail, respectively. we used nad 1983 alaska albers geographic projection for all vector and raster layers; raster maps were prepared at spatial resolutions of 60-m and 100-m; open-access sources are provided for all datasets (table 2). model development and assessment we used a treenet gradient boosting model (salfordsystems, san diego, ca; ohse et al. 2009) to infer relationships between environmental variables and presence of the multi-species bark beetle. bark-beetle presence and random points were set as 1 and 0, respectively, as the binary categorical response variable of the model. treenet uses a boosting classification and regression tree approach (ohse et al. 2009; baltensperger and huettmann 2015a; young et al. 2018; humphries et al. 2018) to model relationships between model predictors and a response variable. we developed three models based on different sets of predictors. treenet settings were the same for all three models: the well-tested ‘default setting’ in the salford predictive modeler (spm) software suite in which the number of trees and maximum nodes per tree was set 200 and 6, respectively, with10-fold cross-validation. this setting is known to perform very well as a standard on most data (salford systems, san diego, ca; ohse et al. 2009; baltensperger and huettmann 2015a; humphries et al. 2018). in model 1, we included all environmental variables as predictors (termed the ‘full model’). in model 2 (‘ecological model’), we excluded distance-based variables and included only predictors expressing local conditions: soil map unit, land status, land cover type, mean annual precipitation, mean annual temperature, elevation, slope, and aspect. finally, in model 3, we excluded roads (euclidean distance to roads) to minimize impact of any sampling bias (kadmon et al. 2004; ohse et al. 2009; baltensperger and huettmann 2015a). in our dataset, almost half of the recorded bark beetle presences were <20 km from roads (appendix ii). each model was applied to the 1-km spaced grid points across alaska. every point on the grid was assigned a relative index of occurrence (rio) ranging 0–1. for better visualization of the study area overall, we used inverse distance weighting (idw) algorithm in arcmap 10.4 to interpolate the value of rio across the state. performance of the three models were assessed by comparing the area under the receiver operating characteristic curve (auc), and percent of correctly predicted presences (% corr.) derived from the confusion matrix (fielding and bell 1997). variable importance (ranking) in each model was assessed using computed relative importances in spm; partial dependency plots were used to illustrate relationship between the response and predictor when all other variables were held to average values. however, the variable importance (ranking) score should not be used to conclude the absolute informative value of a variable; rather, the scores indicate the amount of contribution that each variable (either via a primary role in splitting tree branches or in a substitute role to any of the primary splitters) makes in classifying or predicting the target variable (d. steinberg; salford systems, san diego, ca). as such, the variable rankings are highly related to the performance metric and a chosen tree structure (d. steinberg; salford systems, san diego, ca). our focus, rather, is on deriving inference from the predictions (breiman 2001; salford systems, san diego, ca). conventionally, models have been tested using part of the primarily records of the species presence used to build the model; however, this approach may overestimate a model accuracy, especially if the collection of datasets was systematically biased (newbold et al. 2010). therefore, it is better to assess a model using an independent dataset (newbold et al. 2010). we used an independent dataset of 68 locations of 3 bark beetle species surveyed by the u.s. department of forest service during 2016–2017 (usfs 2019) to additionally assess the three models. the model-predicted occurrences of bark beetle species underlying these 68 widely-spaced locations in the study area were compared using descriptive statistics and box plots in r (r core team 2018; humphries et al. 2018). biodiversity informatics, 16, 2021, pp. 1-19 8 results predicted distribution map most occurrence data were in evergreen forest, mixed forest, shrub/scrub, low intensity developed area, and deciduous forest, in decreasing order (table 1). our results comprise three maps summarize the rio of bark beetle suitability patterns across alaska (figure 3, appendices iii and iv). after comparing the distribution maps resulting from the full model and the ecological model, the core predicted bark beetle occurrence falls in three hotspot regions: south-central alaska, southeastern alaska, and interior alaska. model 3 showed a pattern similar to models 1 and 2: hotspots of occurrence corresponded to the vicinity of human settlements and humanbuilt infrastructure such as towns, highways, and railroads (figure 3, appendices iii and iv). the core hotspots on all three maps (areas close to the urban centers of fairbanks, anchorage, and juneau) likely contributed to the emergence of multi-species bark beetles. in addition to urban areas, other hotspots were along linear features, such as rivers, of central and western alaska. the overall similarity in the prediction patterns using different combinations of predictors for the three models may signal the generic prediction strength gained from machinelearning algorithms. model performance and inferences the auc values for the three models were very close, all at ~0.99. the percent correctly predicted presences for the three models were 91.9%, 93.2%, and 94.4%, respectively. relative importances of predictor variables were computed for the three models (table 3). in model 1, distance to infrastructure, soil, land status, distance to roads, distance to towns, and land cover type were the most explanatory variables. model 3 (in which roads were excluded) showed a similar pattern to that observed in model 1. model 2 revealed variable importance patterns similar to those in the full and ecological models even though the spatially-dependent predictors were excluded from model development. in the ecological model, soil, land status, and land cover type became topcontributing variables. in all three models, land status of “state and native” and “private or municipal” explained the most presence locations whereas “national wildlife refuge” explained pseudo-absence locations. land cover categories of low and medium development intensities appeared to favor bark beetle presence most. developed areas with a low or medium intensity most commonly include singlefamily houses with a mixture of constructed materials and vegetation, and the impervious surfaces from the total cover ranged from 20-49% and 50-79%, respectively. in models 1 and 3, distances below 5-6 kilometers from infrastructures explained the most bark beetle presences. a similar pattern was revealed for distance-to-roads in model 1 and distance-to-towns in models 1 and 3, so that distances below 20-25 km from roads and towns favored bark beetle presences. soil types that with a texture of silt loam, schrock (usually found on stream terraces with a slope of 0-2%), and typic haplocryands (typically found on 1-8% slopes) (usda 1998, 2005), were the most important soil types in providing suitable habitats for bark beetle host tree species. in all three models, figure 3. predicted distribution map of bark beetles in alaska using model 2 (ecological model). predicted maps of model 1 (full model) and model 3 (model with excluded roads) are included as appendices iii and iv, respectively. variable model 1 model 2 model 3 distance to infrastructure 100.00 100.00 map unit soil 82.50 100.00 76.94 land status 44.19 63.52 47.73 distance to main roads 36.40 distance to towns 34.75 37.95 land cover 2011 30.89 54.17 31.56 mean annual temperature 16.03 51.25 14.60 distance to drainage network 14.22 14.39 elevation (dem) 12.70 19.72 11.92 distance to lakes and rivers 10.25 12.37 distance to coastline 9.30 13.99 mean annual precipitation 8.13 25.26 17.47 aspect 5.69 14.69 4.81 slope 4.49 15.42 4.07 table 3. relative importance of predictor variables included in model 1 (full model), model 2 (ecological model), and model 3 (model with roads excluded). top three predictors are shown in bold, and dash indicates nonincluded predictors biodiversity informatics, 16, 2021, pp. 1-19 9 regions with mean annual temperatures >-3.0ºc and mean annual precipitation <350 mm were those that favored the distribution of bark beetles. in addition, areas with slope range <40º and aspect of 100–300º correspond to most bark beetle presences. elevation showed a different pattern from other predictors: in all three models, areas at elevations <2000 m and >4000 m favored bark beetle presence, whereas elevations of 2000-4000 m was not occupied by bark beetle species. in models 1 and 3, proximity at 4 km from rivers and lakes with standing water, 500 m from the drainage network, and 10 km from the coastline, favored bark beetle presences. model assessment although the auc and percent correct prediction statistics revealed that all three models performed well in predicting bark beetle presence and pseudo-absence locations, testing with newly surveyed independent data (2016-2017 records of 68 presence points) revealed differences in their performance. model 2 was most successful in predicting the independent test points (figure 4). the median rios received by test points using models 1, 2, and 3 were 0.018, 0.021, 0.033, respectively. it should be stated that those rios are an index and not probabilities and thus, it includes a range of values that are neither symmetrical nor always reaching 1. that is due to the tree nature of the algorithm used. the model assessment with the alternative test data shows the validity of those concepts (figure 4; kandel et al. 2015). given that the ecological model predicted the bark beetle occurrences marginally better than two other models, we additionally included the frequency distribution of the predicted rio for the 68 surveyed points using the ecological model (figure 5). the frequency plot revealed that about 60% of surveyed bark beetles (assessment/test points) received a predicted index greater than 0.1. predicted values less than a 0.0049 threshold excluded 5% of the test points within the 95% confidence interval (95% of predicted presence were >0.0049; figure 5). we followed pearson et al. (2004) and newbold et al. (2010) to present the prediction map in a binary format but using the 95% confidence interval of the newly surveyed independent presence points, aiming to incorporate current variations in species distribution across the landscape due to temporal-scale changes in the environment (figure 6). the omission rate was figure 4. box plots of models 1-3 used to describe the statistics of the rio gained by 68 assessment/test points of 3 bark beetle species. the dots in model 1 shows the outliers. the dark thick line within the boxes represents median value within the range of predicted index. 25% of dataset are below the median (1st quantile between the two straight dark lines), and 75% of the dataset are above the median (3rd quantile ends by the upper edge of boxes). the whisker on top of the boxes and the lower triangular shapes below the second dark straight line represent the maximum and minimum values. note that machine learning models produce a rio which is not a probability nor symmetrical. figure 5. frequency distribution of gained predicted rios for 68 locations of bark beetle presence surveyed by the usfs (assessment/test points) using the ecological model (model 2). figure 6. classified prediction map of multi-species bark beetle occurrence using 95% confidence interval of assessment/ test points to differentiate predicted index of relative occurrence (rio) of the ecological model. value 1 (presence) represents the favorable habitats and value 0 (absence) represents regions that may not be occupied by scolytines community based on the current climatic conditions and biophysical attributes of the landscape. biodiversity informatics, 16, 2021, pp. 1-19 10 zero after overlaying the 838 presence points on the binary map. about 60.3% of the surface area of alaska received a value of ‘presence’ and is expected to provide favorable habitats for scolytines. those habitats are not solely forested landscapes but include shrublands as well (figure 6). we additionally overlaid the mixed and evergreen forest types, extracted from the nlcd 2011 map, on the binary map (figure 7): ~16% of the mixed forests and 59% of evergreen forests are expected to be suitable for bark beetles, based on current climatic conditions and biophysical attributes of the landscape (figure 7). discussion our approach of studying several species of bark beetles, considered as a community organism was new. it should be emphasized that bark beetles live not only on trees but also on shrubs and similar species that together may create poorly studied landscape reservoirs for bark beetles (mcdermott et al., 2021). for instance, the arctic tundra shrubs would allow bark beetle species to live beyond the tree line in northern areas of alaska. these landscape reservoirs are apparent in our final classified prediction maps (figures 6 and 7). this finding would be critical for a better understanding of the ecology of bark beetle communities, in addition to designing a better forest pest management strategy. the forest and non-forested landscapes (e.g., arctic shrublands) that are predicted to favor bark beetle communities represent potential habitats that may support intraand inter-specific competition within and among bark beetle species. from a biodiversity standpoint, these various favorable habitats for multi-species bark beetles may help to preserve or even promote biodiversity, as well as co-evolution within the bark beetle community. on the other hand, forested areas that are supposed not to be occupied by bark-beetle species (green shading, figure 7) would also be important from a forest management perspective to be protected against future anthropogenic disturbances that may promote the infestations. we were able to assemble open-access data for understanding the ecology of bark beetle communities on a broad scale within the immense geographic area of alaska. this dataset, together with the machine-learning modeling approach that we used, can provide a foundation for future research use. the methodology applies not only to scolytines, but also to other multi-species questions of concern, such as forest defoliators and small and big game wildlife species worldwide (see, e.g., huettmann and schmid 2014, humphries et al. 2018). the boosted classififigure 7. classified prediction map of multi-species bark beetle occurrences in different forest types: mixed and evergreen forests that predicted not to favor bark beetle occurrences (green color), mixed forests that expected to favor bark beetle occurrences (red color), and evergreen forests that predicted to be occupied by different bark beetle species (blue color). the 2011 nlcd was the reference map to extract forest type and area across the state of alaska. biodiversity informatics, 16, 2021, pp. 1-19 11 cation and regression tree approaches that we used are particularly useful in dealing with such ecological and environmental datasets, with the common characteristics of being big, complex, and spatially autocorrelated (breiman 2001; elith et al. 2006; humphries et al. 2018). furthermore, our results will allow us, in the future, to focus on predictions based on climate change scenarios for future time periods (see baltensperger and huettmann 2015b). the higher predictive power of model 3, relative to the full model (model 1), may highlight potential effects of sampling bias in collecting bark beetle occurrence data that results from opportunistic sampling along roads or inconsistent sampling effort over time and space (yost et al. 2008; zabihi et al. 2017). often, species occurrence data have such characteristics owing to lack of awareness or sampling bias in the geographic space (stockwell and peterson 2002; graham et al. 2004; elith et al. 2006; yost et al. 2008; ohse et al. 2009; baltensperger and huettmann 2015a; zabihi et al. 2017). however, the predictive performance of the three models, using the machine-learning algorithm, is sufficiently high that it likely reveals signals in the environmental landscape that can be captured even with sampling bias of species presence. this strength of algorithmic models in finding associations between environmental variables and species occurrences are evident in all three models. for example, even though we lowered the number of predictors from 14 to 13 to 8 in models 1, 2, and 3, respectively, the relative importance of predictor variables did not change (table 3). in addition, the predictive strength of the three models is evident in the three resulting species distribution maps, in which the hotspots of bark beetle suitability are in the southeast, south, and interior of alaska (figure 3, appendices iii and iv). in contrast to the algorithmic model that we used, more traditional parametric models can become unstable by removing less important predictor variables from the model and consequently lead to wrong conclusions (breiman 2001). although one advantage of algorithmic models is in including more predictors to make more information available for prediction (breiman 2001), we selected model 2 for producing a binary predictive map in view of its slightly higher predictive ability using additional independent test points. this model may have had higher predictive ability, thanks to removal of multicollinearity effects between distance-based predictors (e.g., roads, drainage networks, towns, infrastructures, and coastlines) and those predictors considered in the model. for example, drainage network is a function of elevation, slope, and aspect (ohse et al. 2009), which were included in all three models. also, soils and landcover can be a function of drainage networks; for example, rich soils are usually found in well-drained locations such as valleys, whereas mountaintops and alpine zones usually do not have fertile soils, and vegetation classes reflect those correlations and interactions indeed. as is evident from the top predictor variables and visually from the maps in all three models, human settlements and infrastructures are important factors in shaping the distribution of bark beetle species. the hotspot of bark beetle occurrences in the north corresponds well with the periphery of established pipelines; those in the south and southeast are around the cities of anchorage and juneau, respectively. for example, tongass national forest has one of the highest densities of road networks in southeastern alaska, where roads have been used for logging and deer hunting since the mid-1950s (brinkman 2009). the interior hotspots correspond to the vicinity of fairbanks and along highways. we further found that land ownership and management, such as state lands, native lands, and private lands, are closely associated with bark beetle occurrence, likely as a consequence of land use practices that may disturb forest landscapes. for example, state lands, managed primarily by various divisions of the alaska department of natural resources (akdnr), have been influenced by designated land use in the form of sale or lease to the public; lease for commercial, industrial, and recreational use; selling minerals; and temporary permits for use and access (alaska department of transportation and public facilities, northern region, 2018). native lands are aboriginal lands that are owned by individual village corporations having regional rights of exploiting minerals. private lands are owned by individual entities, municipalities, and boroughs, and are generally concentrated close to cities, villages, and populated regions along highways and roads (alaska department of transportation and public facilities, northern region, 2018). the emergence and spreads of bark beetle attacks in the vicinity of human settlements and recreation sites, with public use and infrastructure developments such as pipelines, roads, and hiking trails, could be related to the associated disbiodiversity informatics, 16, 2021, pp. 1-19 12 turbances and compaction damages in soil structure. these soil disturbances consequently compromise tree roots and may eventually lead to higher chance of infections (fs-r10-fhp 2019). urbanization in arctic and sub-arctic regions increases impervious surfaces, creating urban heat islands (chandler 1960; oke 1988) that may impact the climate and associated ecosystem components, such as the spread of bark beetle infestations. in sum, anthropogenic factors seem to be closely associated with, and perhaps accelerate, beetle outbreaks in different ways. for example, human-induced climate change results in warmer and dryer summers that reduce tree resiliency, and milder winters that decrease beetle mortality (müller et al. 2008; müller 2011). in addition, untreated spruce slashand-debris, due to, for instance, highway and powerline constructions, can elevate spruce beetle populations (schmid 1977; werner et al. 2006). logs, slash, or dead and dying trees favor several bark beetle species, such as ips spp., because of little or no host resistance against beetle attacks (fettig et al. 2007). the soil texture of silt loam was closely associated with bark beetle presences, likely due to providing suitable habitats for host trees. for example, schrock and haplocryands soils provide habitats for white spruce (usda 1998, 2005). across the study area, regions with a higher mean annual temperature and lower annual precipitation, relative to other areas in alaska, were more closely associated with bark beetle occurrence. this finding mirrors that of økland et al. (2019), in which high summer temperatures and low precipitation favored the flight period and reproduction rate of most bark beetle species, even those at high latitudes with cooler climates. the aspect range (100–300º), including southand west-facing slopes, is likely to provide a warmer and more favorable microclimate for bark beetle activities in addition to providing favorable habitats for host tree species. for example, white spruce occurrence was concentrated on south-facing slopes in previous studies (viereck and little 2007; ohse et al. 2009). our community-based modeling approach could be debated based on variations in species’ interactions with local environments at fine scales of individual host trees or stands, mostly considered as issues of spatial autocorrelation. however, our modeling approach of using a non-parametric model of boosted classification and regression tree that uses many algorithms, ensembles, and responses (humphries et al. 2018) aimed to learn and model these complex, nonlinear relationships in the data without prior assumptions such as being free of spatial autocorrelation (huettmann 2018c). in addition, different species within a community may have similar responses to changes in the environment at regional scales (golicher et al. 2008; azeria et al. 2009; chapman and purse 2011), so community-level analyses of spatial patterns of biodiversity may be beneficial (chapman and purse 2011) for biodiversity conservations and natural resource management purposes. we used a historical collection of bark beetle specimens from uam without consideration of sampling design strategies or assumptions such as balanced sample sizes for different species. although unbalanced samples may represent true populations of species across landscapes, future work might test these ideas by removing different species from the model. however, our approach treating the species presences across all bark beetles represents a way of dealing with numerous small sample-size species in our dataset. a model with high sensitivity, even if it results in some overpredictions, will minimize omission of sites that are actually suitable, which is particularly meaningful for rare species (engler et al. 2004; barbet-massin et al. 2012). the 2011 nlcd map that we used does not provide detailed information about different types of conifer species; preparing and using such a map in our models could provide additional detail as regards host tree communities and their effects on bark beetle assemblages. acknowledgments this study was supported by grant no. cz.02. 1.01/0.0/0.0/15_003/0000433, “extemit – k,” financed by the operational program research, development and education (op rde). we greatly appreciate the constructive comments and suggestions from the reviewers. we are grateful to all data contributors, namely the alaska division of forestry and the united states department of forest service, and derek sikes, curator of insects at the university of alaska museum, for helping with accessing the data used in this study, as well as for providing comments on an earlier version of the manuscript. we further thank john morton for discussions, and iab-uaf, namely brian barnes and his team, for support for khoda zabihi to carry out this research in the ewhale lab. huettmann thanks hazel berrios, thor chrome, maya hera, and ‘cub’ sparks poa. this is ewhale lab publication # 210. https://onlinelibrary.wiley.com/action/dosearch?contribauthorstored=%c3%98kland%2c+bj%c3%b8rn biodiversity informatics, 16, 2021, pp. 1-19 13 literature cited acia. 2004. impacts of a warming arctic climate impact assessment. cambridge university press, uk. accessed 5 april 2019.4 adfg (alaska department of fish and game). 2018. key habitats of featured species. accessed 5 april 2019.5 dot&pf (alaska department of transportation and public facilities), northern region. land ownership and management (appendix c). 2018. accessed 5 april 2019.6 araújo, m. b., and m. luoto. 2007. the importance of biotic interactions for modelling species distributions under climate change. glob. ecol. biogeogr. 16:743753. azeria, e. t., d. fortin, c. hébert, p. peres‐neto, d. pothier, and j. c. ruel. 2009. using null model analysis of species co‐occurrences to deconstruct biodiversity patterns and select indicator species. divers. distrib. 15:958-971. baltensperger, a., p., and f. huettmann. 2015a. predictive spatial niche and biodiversity hotspot models for small mammal communities in alaska: applying machine-learning to conservation planning. landsc. ecol. 30:681-679. baltensperger, a. p., and f. huettmann. 2015b. predicted shifts in small mammal distributions and biodiversity in the altered future environment of alaska: an open access data and machine learning. plos one 13:e0194377. barbet-massin, m., f. jiguet, c. h. albert, and w. thuiller. 2012. selecting pseudo-absences for species distribution models: how, where and how many? methods ecol. evol. 3:327-338. bearup, l. a., r. m. maxwell, d. w. clow, and j. e. mccray. 2014. hydrological effects of forest transpiration loss in bark beetle-impacted watersheds. nat. clim. change 4:481-486. bentz, b. j., j. logan, j. macmahon, c. allen, et al. 2009. bark beetle outbreaks in western north america: causes and consequences. proceedings: bark beetle symposium, november 15-17, 2005, snowbird, utah. university of utah press, chicago, il. p. 42. bentz, b. j., j. régnière, c. j. fettig, et al. 2010. climate change and bark beetles of the western united states and canada: direct and indirect effects. bioscience 60:602-13. 4 http://www.library.arcticportal.org/1299. 5 https://www.adfg.alaska.gov/static/species/wildlife_action_plan/ appendix5_forest_habitats.pdf. 6 http://www.dot.state.ak.us/nreg/westernaccess/documents/corridor_ planning_report_appx_c.pdf. bhattacharya, m. 2013. machine learning for bioclimatic modelling. international journal of advanced computer science and applications, 4:1-8. bonnot, t. w., j. j. millspaugh, and m. a. rumble. 2009. multi-scale nest-site selection by black-backed woodpeckers in outbreaks of mountain pine beetles. forest ecol. manag. 259:220-228. breiman, l. 2001. statistical modeling: the two cultures. stat. sci. 16:199-215. brinkman, t. j., f. s. iii. chapin, g. p. kofinas, and d. k. person. 2009. linking hunter knowledge with forest change to understand changing deer harvest opportunities in intensively logged landscapes. ecol. and soc. 14:36. chandler, t. j. 1960. wind as a factor of urban temperatures: a survey in north-east london, weather, 15:204-213. chapman, d. s. and b. v. purse. 2011. community versus single‐species distribution models for british plants. j. biogeogr. 38:1524-1535. christiansen, e., and a. bakke. 1988. the spruce bark beetle of eurasia. in: berryman a. a., (eds), dynamics of forest insect populations: patterns, causes, and implications. plenum, new york, pp 480-504. elith, j., c. h. graham, r. p. anderson, m. dudík, s. ferrier, a. guisan, r. j. hijmans, f. huettmann, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. mcc. overton, a. t. peterson, s. j. phillips, k. s. richardson, r. scachetti-pereira, r. e. schapire, j. soberón, s. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species’ distributions from occurrence data. ecography 29:129-151. elith, j., j. r. leathwick, and t. hastie. 2008. a working guide to boosted regression trees. j. of anim. ecol. 77:802-813. engler, r., a. guisan, and l. rechsteiner. 2004. an improved approach for predicting the distribution of rare and endangered species from occurrence and pseudo-absence data. j. of appl. ecol. 41:263-274. fettig, c. j., k. d. klepzig, r. f. billings, a. s. munson, t. e. nebeker, j. f. negrón, and j. t. nowak. 2007. the effectiveness of vegetation management practices for prevention and control of bark beetle infestations in coniferous forests of the western and southern united states. forest ecol. and manag. 238:24-53. fielding, a.h., and j.f. bell. 1997. a review methods for the assessment of prediction errors in conservation http://www.library.arcticportal.org/1299/ https://www.adfg.alaska.gov/static/species/wildlife_action_plan/appendix5_forest_habitats.pdf https://www.adfg.alaska.gov/static/species/wildlife_action_plan/appendix5_forest_habitats.pdf http://www.dot.state.ak.us/nreg/westernaccess/documents/corridor_planning_report_appx_c.pdf http://www.dot.state.ak.us/nreg/westernaccess/documents/corridor_planning_report_appx_c.pdf biodiversity informatics, 16, 2021, pp. 1-19 14 presence/absence models. env. cons. 24:38-49. flynn, m., j. d. ford, t. pearce, and s. harper. ihacc research team. 2018. participatory scenario planning and climate change impacts, adaptation and vulnerability research in the arctic. environ. sci. policy 79:45-53. fs-r10-fhp. 2019. forest health conditions in alaska 2019. anchorage, alaska. u.s. forest service, alaska region. publication r10--pr-45. 68 pp. accessed february 19, 2021.7 golicher, d. j., l. cayuela, j. r. m. alkemade, m. gonzález‐espinosa, and n. ramírez‐marcial. 2008. applying climatically associated species pools to the modelling of compositional change in tropical montane forests. glob. ecol. biogeogr. 17:262-273. graham, c. h., s. ferrier, f. huettmann, c. moritz, and a. t. peterson. 2004. new developments in museum-based informatics and applications in biodiversity analysis. trends ecol. evol. 19:497-503. heikkinen, r. k., m. luoto, r. virkkala, r. g. pearson, and j. körber. 2007. biotic interactions improve prediction of boreal bird distributions at macro‐ scales. glob. ecol. biogeogr. 16:754-763. huettmann, f. 2018a. advanced data mining (cloning) of predicted climate-scapes and their variances with machine learning: an example from southern alaska shows topographical biases and strong differences. in: g. r. w., humphries, d. r., magness, and f. huettmann (eds.), machine learning for ecology and sustainable natural resource management (1st edition, pp. 227-241). cham, switzerland: springer nature switzerland ag, 441 pp. huettmann, f., and m. schmid. 2014. publicly available open access data and machine learning model-predictions applied with open source gis for the entire antarctic ocean: a first meta-analysis and synthesis from 53 charismatic species. in: b. veress and j. szigethy (eds.), horizons in earth science research. volume 11 (chapter 3): pp. 24-34. huettmann, f., e. h. craig, k. a. herrick, a. p. baltensperger, g. r. w. humphries, d. j. lieske, k. miller, t. c. mullet, s. oppel, c. resendiz, i. rutzen, m. s. schmid, m. k. suwal, and b. d. young. 2018c. use of machine learning (ml) for predicting and analyzing ecological and ‘presence only’ data: an overview of applications and a good outlook. in: g. r. w., humphries, d. r., magness, and f. huettmann (eds.), machine learning for ecology and sustainable natural resource management (1st edition, pp. 315-332). cham, switzerland: springer nature switzerland ag, 7 https://www.fs.usda.gov/internet/fse_documents/fseprd712413. pdf 441 pp. huettmann, f., mi. chunrong, and y. guo. 2018b. ‘batteries’ in machine learning: a ffirst experimental assessment of inference for siberian crane breeding grounds in the russian high arctic based on ‘shaving’ 74 predictors. in: g.r.w., humphries, d.r., magness, & f. huettmann (eds.), machine learning for ecology and sustainable natural resource management (1st edition, pp. 315-332). cham, switzerland: springer nature switzerland ag, 441 pp. humphries, g. r. w., d. r. magness, and f. huettmann (eds.). 2018. machine learning for ecology and sustainable natural resource management (1st edition). cham, switzerland: springer nature switzerland ag, 441 pp. hutchison, o. k. 1968. alaska’s forest resource. institute of northern forestry, usfs, pacific northwest forest and range experiment station. 74 p. kadmon, r., o. farber, and a. danin. 2004. effect of road-side bias on the accuracy of predictive maps produced by bioclimatic models. ecol. appl. 14:401-413. kandel, k., f. huettmann, m. k. suwal, g. r. regmi, v. nijman, k. a. i. nekaris, s. t. lama, a. thapa, h. p. sharma, and t. r. subedi. 2015. rapid multi-nation distribution assessment of a charismatic conservation species using open access ensemble model gis predictions: red panda (ailurus fulgens) in the hindu-kush himalaya region. biodivers. conserv. 181:150-161. kurz, w. a., c. c. dymond, g. stinson, g. j. rampley, e. t. neilson, a. l. carroll, t. ebata, and l. safranyik. 2008. mountain pine beetle and forest carbon feedback to climate change. nature 452:987-990. liu, z. l., c. h. peng, t. work, j. n. candau, a. desrochers, and d. kneeshaw. 2018. application of machine-learning methods in forest ecology: recent progress and future challenges. environ. rev. 26:339350. martikainen, p., j. siitonen, l. kaila, p. punttila, and j. rauh. 1999. bark beetles (coleoptera, scolytidae) and associated beetle species in mature managed and old-growth boreal forests in southern finland. for. ecol. manage. 116:1-3. mcdermott, m.t., p. doak, c.m. handel, g.a. breed, c.p.h. mulder. 2021. willow drives changes in arthropod communities of northwestern alaska: ecological implications of shrub expansion. ecosphere 12:e03514. mcdowell, n., et al. 2008. mechanisms of plant survival and mortality during drought: why do some plants survive while others succumb to drought? new phyhttps://www.fs.usda.gov/internet/fse_documents/fseprd712413.pdf https://www.fs.usda.gov/internet/fse_documents/fseprd712413.pdf https://esajournals.onlinelibrary.wiley.com/doi/pdf/10.1002/ecs2.3514 https://esajournals.onlinelibrary.wiley.com/doi/pdf/10.1002/ecs2.3514 https://esajournals.onlinelibrary.wiley.com/doi/pdf/10.1002/ecs2.3514 biodiversity informatics, 16, 2021, pp. 1-19 15 tologist 178:719-739. morris, j. l., s. cottrell, c. j. fettig, w. d. hansen, r. l. sherriff, v. a. carter, j. l. clear, j. clement, r. j. derose, j. a. hicke, p. e. higuera, k. m. mattor, a. w. r. seddon, h. t. seppä, j. d. stednick, and s. j. seybold. 2016. managing bark beetle impacts on ecosystems and society: priority questions to motivate future research. j. appl. ecol. 54:750-760. müller, j., h. busler, m., t. rettelbach, and p. duelli. 2008. the european spruce bark beetle ips typographus in a national park: from pest to keystone species. biodivers. conserv. 17:2979-3001. müller, m. 2011. how natural disturbance triggers political conflict: bark beetles and the meaning of landscape in the bavarian forest. glob. environ. change 21:935946. munro, h. l., b. t. sullivan, c. villari, and k. j. k. gandhi. 2019. a review of the ecology and management of black turpentine beetle (coleoptera: curculionidae), environ. entomol. 48:765-783. nauman, t. w., and j. a. thompson. 2013. semi-automated disaggregation of conventional soil maps using knowledge driven data mining and classification trees. geoderma 2014:385-399. newbold, t., t. reader, a. el-gabbas, w. berg, w.m. shohdi, and s. zalat, 2010. testing the accuracy of species distribution models using species records from a new field survey. oikos, 119:1326-1334. ohse, b., f. huettmann, s. ickert-bond, and g. juday. 2009. modeling the distribution of white spruce (picea glauca) for alaska with high accuracy: an open access role-model for predicting tree species in last remaining wilderness areas. polar biol. 32:1717-29. oke, t. r. 1988. the urban energy balance, prog. phys. geogr. 12:471-508. pearson, r.g., t.p. dawson, and c. liu. 2004. modelling species distributions in britain: a hierarchical integration of climate and land-cover data. ecography 27:285-298. prasad, a. m. 2018. machine learning for macroscale ecological niche modeling a multi-model, multi-response ensemble technique for tree species management under climate change. in: g. r. w., humphries, d. r., magness, and f. huettmann (eds.), machine learning for ecology and sustainable natural resource management (1st edition, pp. 315-332). cham, switzerland: springer nature switzerland ag, 441 pp. r core team, 2018. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. 8 8 http://www.r-project.org/. raffa, k. f. 1988. the mountain pine beetle in western north america. in: berryman, a.a. (ed.), dynamics of forest insect populations: patterns, causes, and implications. plenum, new york, pp. 556-576. raffa, k. f., and a. a. berryman. 1983. the role of host plant resistance in the colonization behavior and ecology of bark beetles (coleoptera: scolytidae). ecol. monogr. 53:27-49. raffa, k. f., b. h. aukema, b. j. bentz, a. l. carroll, j. a. hicke, m. g. turner, and w. h. romme. 2008. crossscale drivers of natural disturbances prone to anthropogenic amplification: the dynamics of bark beetle eruptions. bioscience 58:501-517. jonasova, m., and k. pracha. 2004. central-european mountain spruce (picea abies (l.) karst.) forests: regeneration of tree species after a bark beetle outbreak. ecol. eng. 23:15-27. saab, v. a., q. s. latif, m. m. rowland, t. n. johnson, a. d. chalfoun, s. w. buskirk, et al. 2014. ecological consequences of mountain pine beetle outbreaks for wildlife in western north american forests. forest sci. 60:539-559. schmid, j. m. 1977. guidelines for minimizing spruce beetle populations in logging residuals. u.s. for. ser. res. pap. rm-185, fort collins, colorado. schmitt, c. b., n. d. burgess, l. coad, a. belokurov, c. besancon, l. boisrobert, et al. 2009. global analysis of the protection status of the world’s forests. biol. conserv. 142:2122-2130. schowalter, t. d. 2012. ecology and management of bark beetles (coleoptera: curculionidae: scolytinae) in southern pine forests. j. integr. pest manage. 3:a1a7. seidl, r., m. j. schelhaas, w. rammer, and p. j. verkerk. 2014. increasing forest disturbances in europe and their impact on carbon storage. nature climate change 4:806-810. steinberg, d. what is the variable importance measure? [blog post]. salford systems, san diego, ca. accessed august 2, 2019.9 stockwell, d. r. b., and a. t. peterson. 2002. controlling bias in biodiversity data. in: j. m., scott, p. j., heglund, m. l., morrison, j. b., haufler, m. g., raphael, w. a., wall, and f. b., samson (eds.), predicting species occurrences: issues of accuracy and scale. island press, washington, pp 537-546. stucky, b., j. balhoff, n. barve, v. barve, l. brenskelle, m. brush, g. dahlem, j. gilbert, a. kawahara, o. keller, a. lucky, p. mayhew, d. plotkin, k. selt9 https://www.salford-systems.com/blog/dan-steinberg/item/37regression-tree-ensembles. http://http://www.r-project.org/ https://www.salford-systems.com/blog/dan-steinberg/item/37-regression-tree-ensembles/ https://www.salford-systems.com/blog/dan-steinberg/item/37-regression-tree-ensembles/ biodiversity informatics, 16, 2021, pp. 1-19 16 mann, e. talamas, g. vaidya, r. walls, m. yoder, g. zhang, and r. guralnick. 2019. developing a vocabulary and ontology for modeling insect natural history data: example data, use cases, and competency questions. biodivers. data j. 7:e33303. tyralis, h., g. papacharalampous, and a. langousis. 2019. a brief review of random forests for water scientists and practitioners and their recent history in water resources. water 11:910. usda (united states department of agriculture), forest service, pacific northwest research station, 2009. the western bark beetle research group: a unique collaboration with forest health protection: proceedings of a symposium at the society of american foresters conference, october 23-28, portland, oregon. accessed february 17, 2021.10 usda (united states department of agriculture). 1998. soil survey of yentna area, alaska. accessed april 2, 2019.11 usda (united states department of agriculture). 2005. soil survey of yentna area, alaska. accessed april 2, 2019.12 usda (united states department of agriculture). 2008. insects and diseases of alaskan forests. accessed august 29, 2019.13 usfs (united states forest service). 2019. forest health conditions in alaska reports & ads damage maps (2002-2019). accessed september 12, 2019. usgs (united states geological survey). 2017. alaska range; wood river mountains; kigluaik mountains; brooks range. accessed august 4, 2019.14 10 https://www.fs.fed.us/pnw/pubs/pnw_gtr784.pdf. 11 https://www.nrcs.usda.gov/internet/fse_manuscripts/alaska/ ak631/0/yentna.pdf. 12 https://www.nrcs.usda.gov/internet/fse_manuscripts/alaska/ ak653/0/saintpaul.pdf 13 https://www.fs.usda.gov/internet/fse_documents/ stelprdb5315942.pdf. 14 https://pubs.usgs.gov/pp/p1386k/pdf/10_1386k_alaskarange.pdf. ver hoef, j. m., n. cressie, r. n. fisher, and t. j. case. 2001. uncertainty and spatial linear models for ecological data. spatial uncertainty in ecology: implications for remote sensing and gis applications. c. t., hunsaker, m. f., goodchild, m. a., friedl, and t. j., case (eds.), pp. 214-237. springer, new york. viereck, l. a., and e. l., little. 2007. alaska trees and shrubs. snowy owl books, fairbanks. wallin, k. f., and k. f., raffa. 2004. feedback between individual host selection behavior and population dynamics in an eruptive herbivore. ecol. monogr. 74:101-116. werner, r. a., e. h. holsten, s. m. matsuoka, and r. e. burnside. 2006. spruce beetles and forest ecosystems in south-central alaska: a review of 30 years of research. forest ecol. manage. 227:195-206. yost, a. c., et al., 2008. predictive modeling and mapping sage-grouse (centrocercus urophasianus) nesting habitat using maximum entropy and a long-term dataset from southern oregon. ecol. informatics 3:375386. young, b. d., j. yarie, d. verbyla, f. huettmann, and s. chapin iii. 2018. mapping aboveground biomass of trees using forest inventory data and public environmental variables within the alaskan boreal forest. in: g. r. w., humphries, d. r., magness, and f. huettmann (eds.), machine learning for ecology and sustainable natural resource management (1st edition, pp. 315-332). cham, switzerland: springer nature switzerland ag, 441 pp. zabihi, k., g. b. paige, a. l. hild, s. n. miller, a. wuenschel, and m. j. holloran. 2017. a fuzzy logic approach to analyse the suitability of nesting habitat for greater sage-grouse in western wyoming. j. spat. sci. 62:215-234. https://www.fs.fed.us/pnw/pubs/pnw_gtr784.pdf https://www.nrcs.usda.gov/internet/fse_manuscripts/alaska/ak631/0/yentna.pdf/ https://www.nrcs.usda.gov/internet/fse_manuscripts/alaska/ak631/0/yentna.pdf/ https://www.nrcs.usda.gov/internet/fse_manuscripts/alaska/ak653/0/saintpaul.pdf https://www.nrcs.usda.gov/internet/fse_manuscripts/alaska/ak653/0/saintpaul.pdf https://www.fs.usda.gov/internet/fse_documents/stelprdb5315942.pdf https://www.fs.usda.gov/internet/fse_documents/stelprdb5315942.pdf https://pubs.usgs.gov/pp/p1386k/pdf/10_1386k_alaskarange.pdf biodiversity informatics, 16, 2021, pp. 1-19 17 species host tree species sample size genus genus sample size scierus annectans white spruce (picea glauca) 1 scierus 1 alniphagus aspericollis n/a 1 alniphagus 1 carphoborus andersoni black spruce (picea mariana) and white spruce 5 carphoborus 12 carphoborus carri black spruce and white spruce 5 carphoborus intermedius white spruce 1 carphoborus sp. white spruce 1 cryphalus ruficollis white spruce 7 cryphalus 7 crypturgus borealis black spruce and white spruce 18 crypturgus 18 dendroctonus punctatus white spruce 11 dendroctonus 83 dendroctonus rufipennis white spruce 66 dendroctonus simplex tamarack (larix laricina) 4 dendroctonus sp. white spruce 2 dolurgus pumilus sitka spruce (picea sitchensis) 22 dolurgus 23 dolurgus sp. n/a 1 dryocoetes affaber black spruce, white spruce, sitka spruce, western hemlock (tsuga heterophylla), and lodgepole pine (pinus contorta) 69 dryocoetes 133 dryocoetes autographus black spruce, white spruce, sitka spruce, western hemlock, mountain hemlock (tsuga mertensiana) 51 dryocoetes caryi lutz spruce (picea x lutzii) 10 dryocoetes sp. n/a 3 gnathotrichus retusus n/a 2 gnathotrichus 3 gnathotrichus sp. n/a 1 hylastes nigrinus n/a 2 hylastes 2 hylurgops rugipennis white spruce, sitka spruce, lodgepole pine, and western hemlock 36 hylurgops 52 hylurgops sp. western hemlock 14 hylurgops subcostulatus n/a 2 ips borealis white spruce 9 ips 104 ips perroti black spruce and white spruce 2 ips perturbatus white spruce 52 ips pini lodgepole pine 7 ips sp. white spruce 3 ips perturbatus white spruce 3 ips tridens sitka spruce, white spruce, and lutz spruce 28 lymantor alaskanus n/a 1 lymantor 1 orthotomicus caelatus white spruce, sitka spruce, and lodgepole pine 13 orthotomicus 14 orthotomicus sp. white spruce 1 phloeosinus cupressi yellow cedar (cupressus nootkatensis) 3 phloeosinus 17 phloeosinus pini white spruce 7 phloeosinus punctatus western redcedar (thuja plicata) 3 phloeosinus sequoiae yellow cedar and western redcedar 2 phloeosinus sp. yellow cedar and western redcedar 2 phloeotribus lecontei black spruce and white spruce 3 phloeotribus 12 phloeotribus piceae black spruce and white spruce 9 pityophthorus bassetti white spruce 2 pityophthorus 74 pityophthorus murrayanae white spruce 2 pityophthorus nitidulus balck spruce, white spruce, lutz spruce, sitka spruce, and lodgepole pine 21 pityophthorus nitidus black spruce and white spruce 5 pityophthorus opaculus white spruce 4 pityophthorus pulchellus lodgepole pine 1 pityophthorus recens lutz spruce 1 pityophthorus sp. white spruce, black spruce, sitka spruce, and lodgepole pine 34 pityophthorus tuberculatus lodgepole pine 2 pityophthorus borealis white spruce 1 pityophthorus venustus white spruce 1 polygraphus convexifrons white spruce and lutz spruce 6 polygraphus 59 polygraphus rufipennis white spruce, black spruce, and sitka spruce 52 polygraphus sp. n/a 1 procryphalus mucronatus n/a 1 procryphalus 3 procryphalus utahensis n/a 2 pseudips concinnus sitka spruce and lutz spruce 16 pseudips 21 pseudips mexicanus lodgepole pine 3 pseudips sp. n/a 2 pseudohylesinus granulatus n/a 1 pseudohylesinus 52 pseudohylesinus sericeus lodgepole pine 5 pseudohylesinus sitchensis sitka spruce 2 pseudohylesinus sp. western hemlock 27 pseudohylesinus tsugae western hemlock and mountain hemlock 17 scierus annectans n/a 10 scierus 14 scierus pubescens white spruce 4 scolytinae sp. white spruce and black spruce 8 scolytinae 8 scolytus piceae white spruce, black spruce, and tamarack 12 scolytus 12 trypodendron betulae white spruce and black spruce 4 trypodendron 107 trypodendron lineatum white spruce, black spruce, sika spruce, and western hemlock 74 trypodendron retusum white spruce 9 trypodendron rufitarsis mountain hemlock and white spruce 5 trypodendron sp. white spruce and sitka spruce 15 trypophloeus populi white spruce 1 trypophloeus striatulus n/a 2 trypophloeus 3 xylechinus montanus white spruce 2 total sample size 838 xylechinus 2838 appendix i. bark beetle species, host trees species, and beetle sample size used as presence points to extract model inputs/predictors. n/a represents collected bark beetle specimens with no host tree species included. https://en.wikipedia.org/wiki/picea_mariana biodiversity informatics, 16, 2021, pp. 1-19 18 appendix ii. frequency distribution of euclidean distance to roads (km) for 838 bark beetle presence locations. the peak distance close to the roads could indicate an ecological corridor for the spread of bark beetles. appendix iii. predicted distribution map of bark beetle using model 1 (full model) biodiversity informatics, 16, 2021, pp. 1-19 19 appendix iv. predicted distribution map of bark beetle using model 3 (model with excluded roads) content assessment of the primary biodiversity data published through gbif network: status, challenges and potentials biodiversity informatics, 8, 2013, pp. 94-172 content assessment of the primary biodiversity data published through gbif network: status, challenges and potentials samy gaiji (1) *, vishwas chavan (1) , arturo h. ariño (2) , javier otegui (2) , donald hobern (1) , rajesh sood (1), estrella robles (2) (1) global biodiversity information facility secretariat, universitetsparken 15, dk-2100, copenhagen, denmark (2) university of navarra, pamplona, spain *corresponding author, email: sgaiji@gbif.org abstract —with the establishment of the global biodiversity information facility (gbif) in 2001 as an inter-governmental coordinating body, concerted efforts have been made during the past decade to establish a global research infrastructure to facilitate the publishing, discovery, and access to primary biodiversity data. the participants in gbif have enabled the access to over 377 million records of such data as of august 2012. this is a remarkable achievement involving efforts at national, regional and global levels in multiple areas such as data digitization, standardization and exchange protocols. however concerns about the quality and ‘fitness for use’ of the data mobilized in particular for the scientific communities have grown over the years and must now be carefully considered in future developments. this paper is the first comprehensive assessment of the content mobilised so far through gbif, as well as a reflexion on possible strategies to improve its ‘fitness for use’. the methodology builds on complementary approaches adopted by the gbif secretariat and the university of navarra for the development of comprehensive content assessment methodologies. the outcome of this collaborative research demonstrates the immense value of the gbif mobilized data and its potential for the scientific communities. recommendations are provided to the gbif community to improve the quality of the data published as well as priorities for future data mobilization. keywords— primary biodiversity data, content assessment, and gap analysis. introduction free and open access to primary biodiversity data is essential both to enable effective decisionmaking and to empower those concerned with the conservation of biodiversity and the natural world (bisby, 2000; gaikwad and chavan, 2005; gbif, 2008). however, the history of publishing of primary biodiversity data is very recent. with the establishment of the global biodiversity information facility (gbif) in 2001, concerted efforts to publish primary biodiversity data using community driven and agreed standards and tools gained momentum. gbif was created to facilitate free and open access to biodiversity data worldwide, via the internet, to underpin scientific research, conservation and sustainable development. the gbif network, through its data portal (http://data.gbif.org), already facilitates access to over 377 million records from more than 400 data publishers 1 . the progress achieved in gbif’s first decade indicates that the development of a global informatics infrastructure, facilitating free and open access to biodiversity data, is indeed a realistic aspiration. one of the key future challenges for gbif is now to ensure that such volume of knowledge about biodiversity on earth is indeed of high relevance for the scientific communities. 1 as of august 2012. gaiji et al. content assessment of the primary biodiversity data 95 why assess the content of gbif-mobilised data? despite gbif’s achievements, questions are frequently raised about whether it can yet be considered a global facility (yesson et al., 2007), and about the usefulness of the data mobilised. gbif has been criticised for the taxonomic, thematic, geospatial as well as temporal biases in the data mobilised by its network of data publishers (johnson, 2007). there have been isolated studies to assess gaps, quality and fitness for use of gbif-mobilised data (e.g. guralnick et al., 2007; collen et al., 2008; gbif, 2010a). in 2010, an initial overview of the data published through the gbif network (gbif, 2010b) provided a first set of indicators on the content mobilized so far as well as major bias such as in the taxonomy and temporal areas. recognising this, the gbif-constituted content needs assessment task group (cnatg) recommended that assessment of gbif-mobilised content at various levels (global, regional, national and thematic) is crucial for determining the demanddriven approach for data mobilisation (faith et. al., 2013, 2013). in 2011, in response to these recommendations, a series of improvements to the gbif infrastructure were made such as the rework of the gbif ‘backbone taxonomy’ with up-to-date checklists and taxonomic catalogues such as the catalogue of life 2011 2 . other improvements such as the automated interpretation of the coordinates, country location and scientific names used in published records have been improved to screen out inaccuracies – for example, ensuring that records identified as coming from a particular country are shown as occurring within the borders and territorial waters of that country. the current study attempts to assess the gaps and fitness for use of the gbif-mobilised data. it aims to provide a comprehensive overview of the ‘state of the 2 ruggiero m., gordon d., bailly n., kirk p., nicolson d. (2009). the catalogue of life taxonomic classification, edition 2, part a. in: species 2000 & itis catalogue of life, 3rd february 2012 (bisby f.a., roskov y.r., culham a., orrell t.m., nicolson d., paglinawan l.e., bailly n., appeltans w., kirk p.m., bourgoin t., baillargeon g., ouvrard d., eds). dvd; species 2000: reading, uk. network’ for data published through the gbif network in 2012. such assessment is aimed at demonstrating the value of the content mobilised and how it can contribute to our improved understanding of biodiversity in particular by the scientific community. to achieve this objective and taking into account the large volume of information to be analysed, the authors of this study have adopted two complementary methodologies. one approach led by the gbif secretariat (gbifs) focused on two temporal complete studies (december 2010 and february 2012) while the department of zoology and ecology at the university of navarra (unzyec) focused on processing random samples of the full content. the research outputs of these two studies were compared and complemented each other. the outcomes of these two complementary exercises are presented in three categories: (a) data quality assessment, (b) trends/patterns assessment, and (c) fitness-for-use assessment. data flow of the gbif network as of august 2012, the gbif network is comprised of 419 data publishers from 44 countries and 15 international organisations. together they publish through gbif 10,028 occurrence based data resources (or datasets). figure 1 depicts the typical flow of the data publishing processes through the gbif network. data publishers can use a variety of tools and protocols (e.g. digir 3 , biocase 4 , tapir 5 , gbif integrated publishing toolkit 6 ) and data standards (e.g. dwc 7 and abcd 8 ) in order to publish primary occurrence records to gbif. after successful registration of their resources through the central registry, gbif centrally indexes a limited but essential number of core data elements 3 http://www.digir.net/ 4 http://www.biocase.org/ 5 http://wiki.tdwg.org/tapir/ 6 http://www.gbif.org/orc/?doc_id=2935 7 http://rs.tdwg.org/dwc/index.htm 8 http://www.tdwg.org/standards/115/ gaiji et al. content assessment of the primary biodiversity data 96 detailing the ‘what’ (species), ‘when’ (date/time), ‘where’ (location), “with what evidence” (basis of record) and ‘by whom’ (collector/observer) of the primary biodiversity data published by the gbif network (also called gbif-mediated data). the list of core data elements (table 1) follows a common data standard: the darwin core standard 9 . this data standard has been used for the discovery of the vast majority of specimen occurrence and observational records published through the gbif network. the darwin core standard was originally conceived to facilitate the discovery, retrieval, and integration of information about modern biological specimens, their spatio-temporal occurrence, and their supporting evidence housed in collections (physical or digital). these elements are compiled into a central database (also called gbif index) and their discovery and access is enabled through the gbif data portal (http://data.gbif.org) as well as through web services (http://data.gbif.org/tutorial/services). such a global discovery system is aimed at promoting access to the original information sources owned by each single publisher participating in the gbif network, where more information can be found (e.g. media, richer data etc.). while all data publishers are expected to follow common standards (e.g. dwc), their data resources discoverable through the gbif infrastructure have varying precision and quality. this could be explained by incomplete information at the publisher level, errors during the publishing processes (e.g. formatting of date information) as well as errors during the central gbif harvesting and indexing procedures. in order to assess the content mobilised through the gbif network, this study will focus on using the content of the gbif index as a proxy to the information published by the contributing publishers. 9 http://rs.tdwg.org/dwc/index.htm content assessment of gbif-mobilised data methodology in the last two decades, the informatics field has evolved to a stage where the handling of very large volume of data is becoming the central component of data discovery 10 . the capacity to store, manage and analyse a large volume of data is becoming a fundamental requirements in the field of biodiversity informatics and in particular for infrastructures like gbif 11 . today, technologies like hadoop 12 and hive 13 offer the ability to process such huge volumes of information on certain kinds of distributable problems using a large number of computers. the assessment carried out by gbifs used this new technology to process and analyse the full gbif index is depicted in figure 2. the full gbif index was extracted in the form of hive tables in december 2010 and february 2012. all outputs of the data-mining processes were stored in mysql tables for easy processing and visualisation. the results of these analyses were kept so that in the future similar experiments could be repeated and compared temporally. the hadoop/hive technology allowed the processing and analysis of the full gbif index in a reasonable amount of time compared to conventional technologies like relational database using known database management systems like mysql. however such methodology requires a dedicated infrastructure with sufficient it expertise and understanding of the processes involved in manipulating such large volume of information at once. unzyec used two separate approaches in their assessment (figure 3). in one, a random sample of the gbif index was obtained by issuing 10 jiawei han and jing gao, “research challenges for data mining in science and engineering", in h. kargupta, et al., (eds.), next generation of data mining, chapman & hall/crc, 2009, pp. 3-28. 11 http://www.gbif.org/communications/news-andevents/showsingle/article/important-quality-boost-for-gbif-data-portal/ 12 http://hadoop.apache.org/ 13 http://hive.apache.org/ gaiji et al. content assessment of the primary biodiversity data 97 an automated set of queries through the portal’s web services 14 . this approach mimics an ecological sampling where a vast amount of data is represented by a subset, thus greatly reducing the data processing requirements. in another approach, mirrors of both the gbif index and the raw data harvested from the participants were queried using standard sql statements and scripts. although much more taxing in terms of resources, this approach enabled the authors to finely track the flow of information (not just data) from the publishers to the index. in this way, gaps caused by the data processing flow can be detected. the unzyec team made queries and samplings during a three-year period, over ten versions of the gbif index. however, for the purpose of this assessment, analyses were made mostly on the november, 2010-released mirror, in order to provide an independent comparison of gbifs-obtained results. limitations of the methodologies the methodology used in this article enables the fast data mining of the gbif data index but does not address issues such as: the level of accuracy of the data (e.g. precision in geospatial coordinates). the risk of misidentification of taxa. duplicate records that can arise from: i. datasets being unwittingly published repeatedly, ii. duplicate records within and between datasets, iii. multiple digital records derived from the same physical specimen, such as a specimen being physically split and stored in multiple museums. computing interpretation errors in the data harvesting and indexing routines. 14 http://data.gbif.org/tutorial/services for example, depending on the data schema used (darwin core or abcd) and their versions, an occurrence date may be represented as a datetime stamp, an iso-formatted date, a simple text string in varying formats, or composed of individual fields (day, month, year). the mapping of the data by the publisher may therefore introduce additional error or ambiguity, if for example month and day are swapped. in order to overcome this difficulty, we assumed the level of error of the year within a malformed date-time stamp as sufficiently low to be considered as a good proxy to assess the temporal dimension. with regards to the conversion and validation of taxonomical information (e.g. genus, species, scientific names) the challenges are more complex. during the harvesting and indexing procedures, the taxonomical information is checked against the most up-to-date gbif taxonomical backbone. until end 2011, gbif used the catalogue of life (col) 2007 as its core taxonomical backbone and when unmatched names were identified during the harvesting/indexing procedures they were simply added to the backbone. in november 2011, gbif has entirely refreshed its taxonomical backbone and uses now primarily the latest version of the catalogue of life in addition to other resources (table 2). today, unmatched names are not added to the core backbone and whenever possible, expert taxonomists are consulted. therefore the study undertaken in terms of taxonomical comparison (in 2010 and 2012) should be undertaken taking into account this particular bias due to the improvement of the gbif taxonomical backbone and resolution services. material for the purpose of this study, elements covering three dimensions (“what”, “where” and “when”) were extracted from the gbif index by gbifs and unzyec in december 2010, and also from raw data as supplied by the providers by unzyec for some specific analysis. further analyses using the february version of the gbif index were undertaken by gbifs. gaiji et al. content assessment of the primary biodiversity data 98 the elements covered in these analyses are: source of the data: the assessment has taken into account the identifiers of the data publisher and data resources. however, due to incompleteness and lack of accuracy of entries in the institution id, collection id and catalogue fields in the gbif index, we have decided to exclude these fields from the analysis. taxonomic data: taxonomic ranks such as kingdom, phylum, class, family, genus and species are included. the assessments have also taken into account the synonyms as recorded in the gbif index, in order to provide the most accurate estimate of the number of species. data from multiple synonyms get merged during the harvesting and indexing routines. geospatial data: latitude and longitude information was used when available. however, due to scarce information provided by data publishers, it was not possible to consider precision. this is a serious limitation that will need to be addressed in future analysis. temporal data: limited to the field year of observation/collection. the assessments ignored the day and month recorded in the date field, except for analysing possible causes of year misassignment. other data: the basis of records, a descriptive term indicating whether the record represents an object or observation, was included in the analysis. the basis of record actually contains useful information such as the level of evidence and other categories that may be considered enhanced subclasses of information. results of the content assessment of the gbifmobilised data we present the salient outcomes of these two independent exercises in four categories, namely: (a) data quality, (b) trends/patterns and (c) fitnessfor-use assessments. in most cases, both exercises reached similar conclusions and therefore validate each other. in some instances, significant differences arose and were assessed. a. data quality assessment: taxonomy: until november 2011, the processing of taxonomical references was made against some taxonomical references such as the checklist of catalogue of life 2007 (http://www.catalogueoflife.org/annualchecklist/2007/) or the international plant names index (http://www.ipni.org). during the discovery of unmatched taxonomical references against the accumulated gbif taxonomical backbone, these are automatically added. therefore, the 2010 gbif taxonomical backbone contained accepted names (e.g. from col 2007) and new names discovered during the indexing process. this also means that in our december 2010 assessment, we had limited capacity to distinguish between authoritative names (e.g. referring to catalogue of life 2007 version) and added names, which had no validation against any taxonomical reference. in november 2011, the gbif taxonomical backbone was rebuilt using primarily the latest version of the catalogue of life as well as many new taxonomical authoritative references (table 2). therefore the february 2012 assessment on taxonomical names can be considered as much more accurate. matching against the catalogue of life using a less advanced interpretation techniques developed in 2006 by the gbifs, the backbone taxonomy that covers the occurrence records has 1,946,429 concepts at species or lower ranks, of which 458,716 (24%) is provided by the catalogue of life 2007 annual checklist 15 . a more recent study made in december 2010 16 showed that 52 per cent of the distinct canonical names found in the gbif index matched to a name in the col 2010 using straight, case insensitive matches. this can be slightly increased to 54% if a ‘fuzzy’ matching with a maximum difference of 10% in characters is used. in february 2012, a 15 gbifs personal communication (march 2011) 16 http://code.google.com/p/gbifoccurrencestore/wiki/taxonomicintegration http://www.catalogueoflife.org/annual-checklist/2007/ http://www.catalogueoflife.org/annual-checklist/2007/ gaiji et al. content assessment of the primary biodiversity data 99 similar study (table 3) showed than 53.47% of names were straight, case insensitive matched of the canonical names in the catalogue of life 2011 annual checklist. completeness of the taxonomical classification in order to study the completeness of taxonomical classification in the gbif index, we assessed for each rank (kingdom, phylum, class, order, family, genus and species) the valid references generated after the harvesting and indexing routines. the level of completeness is therefore based on valid taxonomical references within the gbif taxonomical backbone. in cases where for example a family name wasn’t mapped correctly, a ‘null’ value is assigned to this field in the published occurrence record. for each rank, we evaluated the number of occurrences and species (or lower taxa) having incomplete or unknown taxonomical status – or ‘null’ values (e.g. counting all occurrences having an `unknown` status for the kingdom rank). table 4.a provides a summary of our findings in december 2010 and table 4.b the summary for february 2012. in 2010, a total of 114,721 species or lower taxa corresponding to 15 million occurrences representing 5.6% of the gbif index were not ‘mapped’ against the gbif taxonomical backbone at the kingdom level. similar trends are observed for other taxonomical ranks with somehow a variation in amplitude of incompleteness (e.g. 14.5% for species and lower taxa at the family level and 7.4% at the species level). this analysis confirmed similar results obtained in 2008 and 2010 (gbif, 2010b and ariño and otegui, 2008). however some of the correctly matched names against the gbif taxonomy backbone may not be valid names if referred to authoritative references such as catalogue of life. the reasons being that some of these names if not matched to the existing gbif taxonomy backbone during the harvesting and indexing processes were simply added as valid references. the mixing of valid taxonomical references with new unverified references with limited capacity to track such changes over time caused serious difficulties to our study. the assessment summarized in table 4.a provides therefore more a status of incompleteness of the taxonomical backbone rather than a real comparison to any authoritative taxonomical references. in december 2010, our preliminary findings suggested the need for an urgent review of the gbif taxonomical backbone in particular against the most critical taxonomical authorities such as the annual checklist catalogue of life 2010 (http://www.catalogueoflife.org/) and other sources such as the interim register of marine and nonmarine genera (irmng). the decision not to mix unverified names with existing authoritative names was critical. in november 2011, gbifs successfully upgraded its taxonomical backbone against the latest version of the catalogue of life (2011) and other authoritative references. this resulted in our february 2012 study in a more accurate assessment of the taxonomical gaps within the gbif index. the results of this analysis are presented in table 4.b. the percentages of incompleteness observed in 2012 were significantly lower (i.e. 0,35%, 1.81%, 2.82%, 2.17% respectively at the kingdom, class, family and genus levels) than the once observed in december 2010 (i.e. 7.0%, 14,5%, 14,5% and 4.7% respectively at the kingdom, class, family and genus levels) with the exception of the species rank. similar trends are observed taking into account occurrences. therefore a high number of unmapped taxonomical ranks from kingdom to genus levels were resolved using the upgraded gbif taxonomical backbone. the higher number of taxonomical references used to construct the gbif taxonomic backbone largely explains this. the observed percentages of unresolved names at the species level represents 9.15% in 2012 while in 2010 this percentage was of 7.4%. however these numbers can’t be compared because of the changes in the taxonomical backbone between these dates. taking into account these improvements in taxonomical name resolution, we have tried to assess the additional data quality improvements gaiji et al. content assessment of the primary biodiversity data 100 that could be undertaken. to achieve this, we have looked at the top 10 possible misidentifications (at the kingdom level) by number of occurrences (table 5). the three species within the genus zonotrichia listed as within the plantae kingdom are wrongly assigned. these species belong to the american sparrows group of the family emberizidae 17 . this misidentification is due to the generic homonym zonotrichia being both present in the plantae and animalia kingdom. this misidentification is being resolved in the gbif taxonomical backbone and these obvious misidentifications progressively corrected 18 . for the other cases listed in table 5, the discrepancy with col 2011 version is resolved in the latest version of the col (february 2012) or other taxonomical authorities (i.e. marine species identification portal). once these changes are implemented we estimate that 1,808,488 occurrences would be correctly mapped and the total of occurrences with ‘unknown’ status at the species level would decrease from 25,343,834 to 23,535,346. this shows that while the gbif index has grown from 267 to 324 million occurrences (+21.3%) from december 2010 to february 2012, corrections on the top 10 species misidentifications in february 2012 would have resolved a substantive volume of the gbif index: the growth in occurrences with ‘unknown’ status at the species rank would have grown of only 2.3% (from 23,015,905 to 23,535,346). it is therefore reasonable to extrapolate that: a large portion of the gaps identified in table 4.b will in the future be resolved with newest versions of the taxonomical authorities used to build the gbif taxonomic backbone. the rate of resolved names should in principle directly be correlated with the growth in volume of the taxonomic authoritative references used by gbif. table 6.a provides a summary of the taxonomical misidentification at the kingdom level and an indication of the total number of associated occurrences affected. for example, 17 http://en.wikipedia.org/wiki/zonotrichia 18 http://dev.gbif.org/issues/browse/clb-119 correcting the wrong assignment of 90 species from the kingdom plantae to animalia will impact more than 1.3 million occurrences within the gbif index as of february 2012. on the other hand correction of the wrong assignments to animalia of 26 species will only affect 1,536 occurrences. similar breakdowns are provided for phylum (table 6.b) and class (table 6.c). this table shows that the effort in correcting misidentifications at a high taxonomical rank (e.g. kingdom) will impact a limited number of occurrences (1.3 million representing less than 0.5% of the gbif index) only 9.15% of the discovered scientific names in the gbif network have not been mapped to a taxonomic reference at the species level. such volume of unknown references includes for example species not yet endorsed by existing authoritative references used to construct the gbif taxonomy backbone, as well as misidentified or wrongly spelled names. this represents 7.82% of the gbif index in terms of volume of occurrences (i.e. 25.3 million occurrences). we have also demonstrated that this volume of unmapped scientific names has grown less than the growth of the gbif index: +9.9% (25.3 million in 2012 against 23 million in 2010) while the gbif index has grown in the same period of +21% (267 million occurrences in 2010 and 323 million in 2012). the study also demonstrated that compared to the largest authoritative reference the catalogue of life (col) – only half (53.47%) of the species names known to gbif would have been recognized. the other half are mostly names known to other taxonomical references but unknown to col. geospatial: during the harvesting and indexing routines, these geo-referenced occurrences are checked in particular for wrong assignments (e.g. when the latitude and longitude information is not corresponding to the country where the occurrence was observed/collected). in the context of this study, we considered geo-referenced occurrences as a record in the gbif index with the latitude and gaiji et al. content assessment of the primary biodiversity data 101 longitude within the earth-bounding box (i.e. 90<=latitude<=90 and -180<=longitude<=180). this amounted to 99.67% of all occurrence records where geo-spatial information was provided in the gbif index in december 2010the remainder being ‘extra-terrestrial’ (otegui et al., 2009). this includes a substantial number of records being reported as 0.0n, 0.0e and therefore suspicious. this could happen for example when the publisher maps a zero value to the latitude or longitude fields instead of a ‘null’ value. in order to solve such problems, publishers should in addition to ensuring that the mapping of the fields is appropriate, provide for example the country in which the observation/collection has occurred. this would greatly facilitate the validation of geo-referenced occurrences during the harvesting and indexing routines. in end of 2010, 18.45% of the mobilised data were not geo-referenced. this percentage was lower (14.1%) in our assessment of february 2012. as shown in figure 4, the rate of geo-referenced records is increasing over time. such rate is higher for recent years of collection/observation (e.g. from 1973 the rate is constantly greater than 80% for the february 2012 assessment). for older occurrences the rate of geo-referencing is decreasing substantially. for example, before 1930, the rate of geo-referencing was largely lower than 50% and this can be explained by technology limitations (e.g. absence of gps), absence or limited data collection standards covering geolocation (e.g. latitude/longitude, or location fields in collection forms) or simply due to the absence of such information in the collection forms. during the harvesting and indexing procedures, a series of verifications on geospatial fields (e.g. latitude, longitude, country boundaries etc.) are performed enabling for example identification of potential latitude/longitude incorrect assignments. this can be the case for occurrences where longitude and latitude values were swapped; or simply when the longitude value was incorrectly assigned causing for example occurrences originally collected/observed in north america to appear on the asian continent. in february 2012, we estimated that less than 3.6% of the total georeferenced occurrences are falling in this category. in addition we estimated that occurrences without latitude and longitude information but with information for the locality represent 11.1% of the total gbif index. taking into consideration that gbifs is not mandated to apply corrections to the original published occurrence records, these records with possibly wrong coordinates are therefore only flagged during the harvesting and indexing routines. these occurrences aren’t displayed on maps through the gbif data portal but the original occurrences records are kept intact. users of the gbif index (e.g. scientists) should be aware of this limitation and ensure that they consult the ‘geospatial issue’ flag provided by the gbifs. while this addresses partly the problem, it is important to note that the verification and correction of the original occurrences records lies with the publishers. the availability of better guidelines 19 and practices in recording biodiversity observations/specimens should support publishers in this effort. in addition, the use of tools like biogeomancer and geolocate should be recommended and more widely promoted. such tools parse place name descriptions in multiple languages and provide in return a set of longitude/latitude coordinates associated with that description. the data curators can therefore enrich their database content, increase the quality and accuracy of the content mobilised through gbif and thus makes it suitable for wider uses. the high percentage of georeferenced records within the gbif index as well as the observed positive improvements in our two assessments is an important quality stamp of the gbif mobilised data. the study showed that the rate of georeferencing in the gbif-index is increasing over time due mostly to better data quality checking 19 principles of data quality arthur chapman http://www.gbif.org/orc/?doc_id=1229 gaiji et al. content assessment of the primary biodiversity data 102 activities both at the publisher and central levels (figure 4). the improvements are observed for all decades since 1900. however, the variability is very high and older occurrences are expected to have a lower probability of valid geospatial information (e.g. prior to 1930: <50%, 1930 – 1960: <70%, 1980-2010: >80%). more importantly, our study shows that the percentage of potential wrong geo-spatial records is very low (<3.6%). in most cases, such situation can be explained by wrong latitude or longitude sign assignments and these can be easily resolved by swapping coordinates. temporal: as detailed in table 7, 30.8% of the gbif index contains records with null or not valid year in the date time stamp field. the unzyec analysis (table 8) estimated a similar percentage (31%) although it distinguished invalid years (i.e. before 1750 or in the future). the breakdown provided in this analysis shows that 4.3% had not valid date stamp data and 26.7% had no data or null values. however, the comparison between raw data and processed data uncovered some issues on date processing, such as mismatches between the published and interpreted date stamp. for example, 8.6% of the records with a value in the date stamp field were nullified during the harvesting and indexing procedures. in addition, 5.0% of the null values in the publisher data were converted to valid date stamp values after harvesting and indexing. more details about this mismatch can be found in otegui & al., 2013 (this volume). thus, according to the unzyec study (table 8), 36.1% of the records would be either undated or doubtfully dated at the year level. these analyses show that a large volume of date stamp information does not convert to a valid date stamp, or at least lacks information about the year of collection/observation. in 2011, these preliminary findings were taken into account by the gbifs and existing processes to interpret date stamp at the publisher level were reviewed and improved. table 7 shows also the comparison between the assessment made in december 2010 and february 2012. while during this period of time, the gbif index has grown by 21% in total (respectively 267 in december 2010 to 324 million occurrences in february 2012), the total number of occurrences with no year provided in the final gbif index has decreased by 47.9%. this amounted to 13.3% of the total gbif index compared to 30.8% in december 2010. most of these improvements relate to improved interpretation of malformed data stamp information in the published resources during the harvesting and indexing routines. temporal information is useful for two classes of questions: 1) biogeographic changes over time and 2) phenological. the year information is the most important element within temporal date stamp information to note long-term changes. however, the month and day elements provide additional accuracy in particular when looking at migratory species moving for example from feeding to reproduction areas during the same year. partial date, as found on many older specimens may be useful for one or the other of these purposes even if they cannot serve all needs. such gaps in the temporal attributes are a limitation for certain types of analysis, such as population cycles or changes in migration patterns related to climate change. alone, the low percentage of occurrence records without temporal information (13.2%) is not considered as a major limitation. however, combined with other parameters like georeferencing, it could become a serious limitation for scientists in particular when dealing with analyses requiring the combination of these (e.g. ecological niche modelling). as shown in table 9, if we consider only presence of valid temporal and geospatial information as determinants of ‘fitness for use’ in the context of ecological niche modelling analysis, 78.8% of the gbif index is meeting these criteria. this total represents 484,963 (48.6%) species from the total identified in the full gbif index of 995.974 species as of february 2012. but this also indicates that 51.4% of the species recorded in the gbif index don’t gaiji et al. content assessment of the primary biodiversity data 103 have a single information on the temporal*geospatial dimensions. background ‘noise’: in december 2010, we estimated that 121.7 million records had missing, doubtful or wrong information in at least one of the three key attributes (i.e. taxonomy, georeferencing and temporal). this represented 45.6% of the gbifmobilised data records (267 million records). although this was an improved figure compared to the 50.1% calculated in may 2008, it calls for concerted efforts firstly to sensitise data publishers of the need to ensure that all available temporal, taxonomical and geospatial information are correctly mapped during the publishing process to gbif. in 2011, gbifs had greatly improved its harvesting and indexing processes in order to optimize its ability to interpret as accurately as possible the information of publishers. in february 2012, the taxonomical backbone was greatly improved and the indexing processes fine-tuned. this has led to a lower percentage (21.3%) of the gbif index with absence of information in at least one of the three variables: temporal, taxonomical and geospatial. while these data quality trends are promising (figure 5), they are mostly due to technical improvements in the gbif it infrastructure and much more efforts are required at the level of the data publishers within the gbif community. collection curators should be encouraged to explore ways to improve the quality of the published information in particular for three dimensions, namely: taxonomical, temporal and geospatial. many tools are aimed at helping curators to identify possible errors and to standardise data in accordance with authoritative references. some key examples are:  specieslink developed by cria (centro de referência em informação ambiental) available at: http://splink.cria.org.br/  biogeomancer coordinated by the university of california at berkeley (http://www.biogeomancer.org)  diva-gis developed by robert hijmans (http://www.diva-gis.org/)  biddsat developed at unzyec (otegui & ariño, 2012) available at: http://www.unav.es/unzyec/mzna/biddsat/ duplicates concerns about the amount of record duplicates in the gbif index were also raised over recent years (hobern, 2003; page, 2012). such situation could happen for example when the same dataset is published more than one time through gbif. comparing datasets on criteria like taxonomy, temporal and geospatial information can easily identify these cases. to assess these cases, we assumed that a duplicate record would be identified when the values respectively for taxonomical (species id), temporal (timestamp date e.g. yyyymmdd) and geospatial (latitude and longitude) are identical. based on this assumption, we calculated in february 2012 the total amount of potential duplicates between resources. the results are summarized in table 10. we have identified 42 combinations of datasets with at least 100,000 potential duplicate occurrences representing a total of more than 30 million occurrences. this represents more than 9.5% of the gbif index. the top 20 potential duplicate combinations are listed in table 11. in all cases (e.g. inbio, cnin/lepidoptera, pelagic fish observations 1968-1999, birds (kiee-bi)) it appeared that the resources were republished twice to gbif but with a different name (e.g. ‘pelagic fish observations 1968-1999’ and ‘pelagic fish observations 19681999 (australian antarctic data centre)’). what appears very surprising is that most of these potentially duplicated resources were registered with very similar names (e.g. ‘cnim/lepidoptera’ and ‘colección de referencia de lepidópteros diurnos mexicanos de la cnin’). when a new resource is registered, a simple text comparison between the title of the new resource with existing published ones would have enabled rapid http://www.cria.org.br/ http://www.cria.org.br/ http://splink.cria.org.br/ http://www.biogeomancer.org/ http://www.des.ucdavis.edu/facultyinfo.aspx?id_number=83 http://www.diva-gis.org/ gaiji et al. content assessment of the primary biodiversity data 104 identification of obvious possible duplication. this has never been implemented up to now in gbif but efforts are underway to automate this process as well as to resolve the already identified potential duplicates in close communication with the respective gbif publishers. an improved monitoring of the resource at the time of registration is indeed an immediate solution but ultimately the adoption of persistent identifiers for each resource published (e.g. doi), with proper metadata, would have been a much more robust solution. b. trends and patterns assessment taxonomy: in december 2010, of the 267 million occurrences records accessible through the gbif network 62% belonged to kingdom animalia, followed by kingdom plantae (23%), fungi (1.55%), protozoa (0.67%), and bacteria (0.59%) (figure 6.a). a similar assessment in february 2012 (figure 6.b) showed that the major variation was the increase for the plantae from 23% to 30%. between these two assessments the gbif taxonomical backbone was reviewed with the latest version of the catalogue of life. monitoring of the taxonomical name resolution during the gbif harvesting and indexing procedures has shown that a large proportion of names previously classified as ‘unranked/unknown’ were now reclassified in particular within the kingdom plantae (gbif, personal communication). in december 2010, as depicted in figure 6.a, 52% of the occurrences belonged to phylum chordata (kingdom animalia) followed by 17.7% belonging to phylum magnoliophyta (kingdom: plantae), and 9.8% to phylum arthropoda (kingdom animalia). a breakdown at the class rank (figure 8.a, figure 9) shows that the largest class in the gbif mobilised data is aves (43%). this is mostly due to field observation from the ornithological community as depicted in figure 10. in december 2010, the bird observation checklist database represented 42.21 million occurrences or 15.7% of the total gbif index at that time. within this top 5, four resources are related to bird watching activities (e.g. bird observation checklist database, project feederwatch, great backyard bird count, southern african bird atlas project). while these figures clearly indicate the dominance of bird observations among the data accessible through the gbif network, it also demonstrate the effectiveness of a given specialized network to leverage on the existence of gbif to enable the publishing, discovery and access to such type of biodiversity observations. these figures also show that the spread of occurrences across various taxonomical levels is also rather heterogeneous (figure 7.a). some phyla are extremely underrepresented, while specific classes such as aves dominate, or even orders within the class hexapoda (insects) (figure 8.a and b, figure 9). the hierarchy of the most represented groups (irrespective of taxonomic level) shows classes aves, actinopterygii (bony fishes), poales (grasses), mammals and asterales as the largest groups, followed by order lepidoptera within the class hexapoda. on the other extreme, for example phyla zygomycota, nemata or platyhelminthes, or kingdom bacteria, have marginal occurrence despite their natural abundance. however it is important to note that many of these apparently overrepresented taxa are species-rich, and have greater biomass and greater visibility, but also a higher number of competent specialists and observers. having such large amount of data for a relatively small number of taxa should also be considered as a positive asset in particular when looking at temporal species distribution, provided that these taxa are ecologically diverse as well as representative. the availability of such high-density information for fewer taxa should not be under-estimated. the analysis of the temporal spread of gbifmediated data for the two dominant kingdoms (animalia and plantae, figure 11) shows that the exponential increase observed from 1960 is mostly explained by the abundance of occurrences for the kingdom animalia. this increase of bird gaiji et al. content assessment of the primary biodiversity data 105 observation data records exceeds the mobilised data from all other classes from year 2000 onwards. in the same period (1960-2010), we also observed that beside a peak in 1999, the trend for plantae is stable varying from 1.5 to 2.1 million occurrences observed/collected per year. as shown in figure 13, the exponential increase of data records in the gbif index in recent years is largely explained by the growth of occurrences in the class aves. figure 14 provides a breakdown of occurrence records by basis of record within the class aves. since 1960, bird observation data have been growing almost exponentially while the trend remains stable for specimen and other types of data. if these trends are confirmed in upcoming years, it is expected that the growth of data records in aves will be the main driver behind the growth of gbif index in terms of volume. this phenomenon is even more revealing when listing the top 15 species by the number of data records. tables 12.a and 12.b show that all of the top 15 species are birds, mostly published through networks like ‘ebird bird observation checklist’, or other similar resources (e.g. project feederwatch, great backyard bird count). in order to demonstrate the difference between the kingdom animalia and plantae, we have generated two sub-indexes for each kingdom from the february, 2012 version of the gbif index. each sub-index was subdivided in new subsets based on the range of occurrence numbers for each species. table 13 provides the summary of the results. for example, from the total of 457,340 animalia species in the gbif index, 400,088 are species with less than 100 occurrences each, and represent 2.4% of the total number of occurrences in the kingdom animalia. the breakdown of species by occurrences did not show any major differences between the two kingdoms except for 20 species in the kingdom animalia holding more than 1 million occurrences each, while no species had as many occurrences within the kingdom plantae. however, occurrences themselves diverged between kingdoms. the set of 20 species having more than 1 million occurrences each identified in the animalia kingdom accounted for 15.9% of all animalia occurrences (zero for plantae), and for the species in the range 100,000-1 million occurrences a higher percentage was also observed for animalia (39.8%). plantae occurrences concentrated around species represented each by less than 100,000 occurrences. we conclude that the abundance of occurrences records in the gbif index for a few animalia species is representing a significant portion of the full gbif index. these records are mostly represented by bird observation data. however this trend shouldn’t under-estimate the amount of species from all kingdoms having less than 1 million and more than 1,000 occurrences, since these do represent a large portion of the gbif index (74.5% of animalia and 78.7% of plantae). we have compared the distributions of the year of collection/observation of occurrences for both plantae and animalia kingdoms (see figure 12.a. and 12.b) taking into consideration the december 2010 and february 2012 versions of the gbif index. both figures show that over time the rate of data mobilised per year tend to increase in both cases. however, for plantae (figure 12.a) we observed that, aside a few artefacts (e.g. year 1999) the rate of mobilisation is increasing at a slower rate to even stagnate from year 2000 compared to the animalia kingdom (figure 12.b). on the other hand, we observed that the evolution for the animalia kingdom was approximately exponential in both versions of the gbif index. as indicated previously, this is attributed to the increased proportion of bird data in the gbif index in particular in the last decade, as shown in figure 13. this confirms the fact that the rapid growth of the volume of occurrences in the gbif index is mostly driven by the bird observation data. the spread of other large publishers is perhaps wider, the main difference being the concentration of bird data towards recent years and few publishers (see otegui & ariño, 2013). the value of observational data in comparison to voucher specimen in museums or accessions stored in genebanks is a subject for another gaiji et al. content assessment of the primary biodiversity data 106 discussion. however this study (figure 14) demonstrates the over-representation of observational occurrences in the gbif index. the ratio between observation and specimen was very close to 1:1 until 1975. thenceforth, the amount of observation occurrences has grown exponentially while the trend for specimen data was very much stagnating until 2000, where we observed a decline. in the last decade, the proportion of observation occurrences represented more than 90% of the yearly collected/observed occurrences. the dominance of bird observational data in the last decades, as well as the drop for data mobilized in recent years for other classes during the last decades, is cause of concern. while on one hand the availability of such large volume of bird data will enable advanced research in temporal trends of bird populations, it also reveal the difficulty to undertake such valuable research in other classes. part of this can be explained by rapid data mobilisation of the “low-hanging fruits” (or relatively easy to digitise and publish) vouchered specimen data (berendsohn et. al., 2010). many of the large natural history museums have digitised their main historical records (ariño, 2010) and published them through gbif. it is therefore expected if this situation of dichotomy between bird observation data and the other classes will increase in the next years. taking into consideration the existing major threats to biodiversity, the gbif community needs to greatly strengthen its capacity to assess trends also for all non-bird biodiversity records. for example, gbif could evaluate the opportunity to develop a list of priority species based on known references, for example the iucn red list; gather information about their distribution, and evaluate for each the availability of rich yet still undigitized or electronically unavailable occurrence data in the gbif community. this approach would lead to a series of strategic data mobilisation strategies for each priority species. geospatial: in the december 2010 assessment we observed (table 14) that the majority of the occurrences present in the gbif index were located in northern america (28.19%), northern (30.06%) and western (11.48%) europe. this represents a total of 69.73% of the gbif index. in february 2012, we observed the same trend where these three regions represented 70.8% of the gbif index, with minor variations in the order (e.g. northern america was classified as the second region in 2010 while it became first in the 2012 assessment). there are multiple reasons that can explain this distribution. the comparison between existing financial contributions to the gbif secretariat (table 15, as of year 2010) on a regional basis shows that the sum of the contributions of these three regions equals 64.9% of the total gbif operational budget, which is very similar to the percentage of occurrences collected/observed in these regions (69.73%). the major discrepancy observed in this table is the financial contribution of eastern asia countries (22.1%) for only 2% of the occurrences in the gbif index. this can be explained by the contribution of japan within a region where the rate of data mobilisation is still low. in 1999, the oecd biodiversity informatics subgroup in its working group on biological informatics report made major recommendations for the establishment of gbif. it is therefore not surprising today to observe (table 16) that the majority of the occurrences in the gbif index are located in oecd countries (84.45%). taking into consideration megadiverse countries, large countries like brazil, china, democratic republic of congo, indonesia, malaysia, papua new guinea or venezuela are not yet members of gbif in 2012 making it difficult for information from these countries to be published through gbif. thus, gbif mobilised data are very much biased towards it original founders, mostly oecd countries. a clear example can be found in otegui et al. (2009), where the geospatially-explicit provenance of data gaiji et al. content assessment of the primary biodiversity data 107 contributed by european publishers in the 2008 sample nicely matches the publisher’ country (figure 19). as shown on table 17, not surprisingly the majority (85%) of the occurrences were located in high-income countries, 11.7% in upper-middle income countries and less than 4% in lower-middle and low income countries. the distribution of occurrences along latitude (figure 16) confirmed also the large proportion of occurrences located in northern hemisphere, where the three regions contributing most records are located (table 14). the peak observed in the southern hemisphere is mostly explained by the recent publication of a large volume of occurrences from south africa, australia, and in particular through the atlas of living australia. however, the species richness, as measured by density of species per half-degree of latitude (figure 17), showed a slightly different trend. we did not observe the large dichotomy between the two hemispheres that appeared in the density of occurrences, and the species richness ranged from 10,000 to 40,000 species per half-degree. figure 18 provides a justification for these trends. the average number of occurrences in the southern hemisphere did not exceed 35 occurrences per species at that latitude range resolution, while this rate exceeded 50 for much of the latitudes north of 50ºn, and even reached peaks higher than 160 occurrences per species per half-degree. we therefore conclude that despite a bias of occurrences towards northern hemisphere, the species richness observed is equally distributed between hemispheres. we also conclude that species in northern hemisphere had a higher rate of occurrences/species than in southern hemisphere. this can suggest a wider distribution of temporal data for these species in the northern hemisphere, and therefore the availability of information more suitable for studying the temporal trends of species distribution in these regions. for the southern hemisphere, we also conclude that many species may not have sufficient occurrences to perform such analysis. temporal: the temporal evolution in the gbif index is summarized in figure 20. with the exception of a few artefacts (1950, and 1987 for the december 2010 curve), we observed that the availability of occurrence data over time grew almost exponentially. a striking feature in this trend was the presence of large peaks in specific years. these peaks seemed to respond to a combination of a provider effect and a possible mismatch between published data and indexed data arising from the date processing algorithms, that is explained in detail in otegui et al., 2013 (this volume). the drop observed in the last period (between 2007 and 2010) for the december 2010 assessment can be attributed to the lag time required between the data collection/observation, digitization and publishing through gbif. the same lag time (3 years) was later confirmed for the february assessment. we conclude that the amount of biodiversity data collected or observed tends to be greater for more recent years than for any older period (e.g. prior to 1970-1980). we also analysed the evolution of such trend by comparing the december 2010 and february 2012 assessments (figure 21). the two horizontal lines represent the average growth in the gbif index for all occurrences and for occurrences having temporal information. the difference can be explained by two factors: (1) the improvement of the gbif indexing processes in 2011, which enabled greater recovery of malformed date-stamp fields; and (2) the greater percentage of well-formed temporal fields (e.g. date of collection/observation) in the recently published data. the graph shows that for more recent decades (e.g. 1971-1980 onward) the growth of data in the gbif index is faster than for older data. more remarkably, we observe that for the latest decade (2000-2010) the variation is of 89.6%, which is the highest growth rate ever observed. the exponential growth of recent data in the gbif network content is particularly driven by the availability of bird observational data during the last decade; this growth in recent content is sometimes termed a 'data deluge'. gaiji et al. content assessment of the primary biodiversity data 108 the trends for the number of species collected/observed every year since 1900 (figure 22) for both plantae and animalia were very similar. we observed an increase until the 1990’s, and then stagnation followed by a drop from year 200 (with the exception of few artefacts). the drops observed in 1914 to 1917 as well as from 1939 to 1942 can be easily explained by the effect of the two world wars. what is troublesome, however, is the drop in both curves from 2000 onwards. the drop for the animalia is even more severe than for plantae. while the volume of occurrences mobilized every year is increasing until 2009, we note that at the same time these occurrences belonged to fewer species across both kingdoms. one possible explanation could have been related to a lower number of data resources publishing since year 2000 (figure 23), but it should be noted that the decline in species richness started more than one decade earlier. we have also calculated for each year the rate of geospatial occupancy in a grid with a resolution of halfdegree (figure 24). we observed that in all cases the grid occupancy for animalia species was higher than for plants. from 1963 to 1993, grid occupancy for animalia was stable followed by a peak in 2000. for plantae, grid occupancy was stable from 1970 until 2000. we also noted that in both cases grid occupancy started to decrease in 2000. a detailed analysis of these trends will be further presented in a separate study. c‘fitness-for-use’ assessment: assessing the value of the gbif mobilised data for a variety of usages is challenging. in this study, we decided to focus on the most common uses for gbif-mobilised data reported in the scientific literature: ecological niche modelling (enm) (grinnell, 1917; fernández et al., 2009; peterson and vieglais, 2001) and related analyses. the compilation of scientific literature using or citing gbif is available since 2011 on-line at: http://www.mendeley.com/groups/1068301/gbifpublic-library/. such modelling techniques (e.g. using maxent) required occurrence records with proper temporal attributes, correct geo-referencing attributes as well as sufficient volume of welldistributed data-points. the minimum number of distinct data-points for a niche modelling analysis is in the range of 10 to 20 (pearson et al., 2007; grantham et al. 2008). recent studies on gbifmediated data using more than 19,000 plant species showed that a preferred threshold of 20 to 40 points is recommended (jarvis, personal communication). maxent models generated for species meeting these criteria have an area under the curve (auc) greater than 0.75 in more than 95% of the cases. if time series are part of the models, then the requirements on number of data points can be an order of magnitude higher. for example, ariño and pimm (1995) showed that successful modelling of the evolution of population extremes require a minimum of 15 distinct time-dependent population estimates. in terms of enm, it could be argued that if using cell frequencies in the enm as an indicator of potential population estimates, at least 15 independent models, each time-constrained, should be needed to adequately characterize any time-dependent changes in the model. this may hold for both terrestrial and marine models, despite their intrinsic differences (warner et al., 1995). in this study we have decided to use the threshold of “presence in at least 20 distinct cells in a 1/10 degree grid” to define whether a species has sufficient occurrences in the gbif index to be used for ecological niche modelling. we used this threshold for temporal/spatial requirements to assess the number of species suitable for such ecological niche modelling analysis, but make no attempt yet to assess whether each selected species can be adequately modelled over time. tables 18.a, b and c provide a distribution of the number of species falling in various categories of grid occupancy. our analysis was based on the february 2012 version of the gbif index due to the improved accuracy of the taxonomical matching, the greater resolution of date stamp as gaiji et al. content assessment of the primary biodiversity data 109 well as for geo-referencing attributes. for the full gbif index, more than 995,975 species (table 18.a) were recorded in the gbif index (with at least one occurrence record). however only 603,532 species had at least one occurrence present in the gbif index with at least one presence in a distinct 1/10-degree grid. this means that 39.4% of the species recorded had no geo-referenced attributes. 747,988 species had at least one occurrence with a valid temporal attribute. this number dropped to 485,105 species if we added the condition of at least one geo-referenced attribute within a 1/10-degree grid. if we consider the enm threshold of 20 presences in 1/10-degree grid with valid temporal attributes, the total number of species that were suitable for enm analysis fell to 81,057. this represents 8.1% of the species recorded in the gbif index. while this percentage could be interpreted as a low percentage, the number of species falling in this category is already very high for many scientists and researchers interested in estimating the actual species distribution as well a projections in the future taking into account future climatic scenario. the use of such information is extremely valuable already for advanced scientific research and in particular in support of global biodiversity assessments such as the strategic plan on biodiversity of the meas (also called aichi targets). taking into account the constant growth of the gbif index with the addition of new datasets for example, it is logical to expect that this amount of ‘eligible’ species will increase over time. what was also remarkable was the number of species with a presence in at least 100 1/10-degree cells with valid temporal attribute: 14,041. this rich reservoir of species with high quality occurrences is already an important message to the research community seeking to assess the species distribution evolution over time as well as future predictions. as shown in table 18.b and c, this number is somehow equally distributed between species within the two dominant kingdoms: plantae (6,100) and animalia (6,756). figure 24 also shows the temporal trends between these two kingdoms. even if the grid occupancy is constantly higher on a yearly basis for animalia, the trends between these two kingdoms are very similar. the same trends are also observed for a low threshold of 20 1/10-degree grid presences. we also noted no major differences between kingdoms in the breakdown assessment of the animalia (table 18.b) and plantae (table 18.c): 36,462 animalia species were suitable for enm (7.9% of the total number of animalia species recorded in the gbif index) against 37,730 plantae species (8.7%). therefore the concerns about the over-representation of bird observation data within the kingdom animalia are contradicted here in terms of ‘fitness for use’ since a large number of plant species were already meeting the enm suitability criteria. figure 25 proves how much the two kingdoms can’t be distinguished when looking at a presence higher than 20 in 1/10degree grids. in this study, we have also tried to assess the grid occupancy at class (table 19), and family (table 20) levels. our objective was to assess the percentage of species within each rank suitable for enm. for example, within the class aves 45.2% of the recorded species in the gbif index were already suitable for enm studies. in this particular case, we can conclude that the gbif index as of today can be used to estimate the biodiversity of most species within the class aves. taking into account the actual trend in terms of bird observation data, it is therefore expected that this percentage will grow in the future. such volume of information can now open new opportunities such as studies on the over-sampled areas (e.g. north america or western europe) and recommendations for new areas where collection of new specimen/observation is required. studies on the estimated number of species (chao2) at a regional and global level could now be performed. table 19 also shows that other classes are eligible for such analysis: cartilaginous fish (elasmobranchii, holocephali), bryophyte plants (marchantiophyta). for many, the limited number gaiji et al. content assessment of the primary biodiversity data 110 of species within that class can explain this. however, for some larger classes like actinopterygii (ray-finned fishes), elasmobranchii (cartilaginous fish) or pinophyta (conifers) the gbif index holds sufficient information for more advanced enm or other biodiversity assessment analysis. figure 26 provides a visual representation of these trends. with the exception of the class aves, the other classes shared a similar trend. such analysis may thus include a bias: the number of species representing each taxon group at a given level. large families or classes, e.g. hexapoda (insects), may not be listed in the top 10 or 20 lists. figure 27 shows the distribution of families taking into account the number of species within each family. while we observed that classes aves and actinopterygii were listed as the ones with the highest suitability for large-scale enm, families within other classes, such as some insects (e.g. cryptophagidae beetles), or mammals (phyllostomidae new world leafnosed bats) are also to be considered. for even larger families (e.g. with a number of species greater than 1,000) it was not surprising to observe that the percentage of species suitable for enm was lower. however for such large families, the suitability for enm of 5-10% of their known species is probably a good proxy to initiate an assessment of the full class. in this category, in addition to families of insects we observe some large families of reptiles (e.g. scincidae – lizards, colubridae – snakes). for plants (figure 28), the distribution of families is somehow distorted due to the high to very high number of species found within each class. discussion the idea that birthed gbif ten years ago remains as simple and powerful now as it was then: to make the world’s biodiversity information freely and universally available for science, society and a sustainable future (oecd, 1999). after 10 years of existence, the gbif network represents the largest resource of primary biodiversity data that is freely accessible to all. with over 377 million occurrence records about nearly one million species (as of august 2012), the gbif mobilised data provides a data-driven window to the state of the world’s biodiversity. access to such large volume of data opens for example new research avenues from assessing the state of biodiversity, identifying the potential threats up to monitoring trends and predicting future evolution and composition of biodiversity and ecosystems (rödder & lötters, 2010; ramírez-villegas et al., 2010; ready et al.,2010). since 2008 till june 2012, over 600 scientific peer reviewed papers have been published which are based on analysis and interpretations of gbif mediated data (gbif, 2012a). to become such a truly ‘global biodiversity information facility’, gbif needs now to take into consideration the primary applications it originally intended to offer to the public such as in policy formulation, economic development, environmental protection, education, and scientific research. in order to ensure its relevance for such applications, the gbif community needs to warrant that the information it delivers is of relevance to address the major science, societal and policy challenges. while these needs are very diverse and difficult to categorize they do have in common essential pre-requisites that can be summarized as follows:  “can i trust the information provided?”  “is the information representative of biodiversity on earth?”  “can i use the data to model biodiversity over time?” the present study was therefore aimed at assessing the data quality, bias and ‘fitness-foruse’ of the gbif mobilized content. these challenging questions were addressed by tasking two separate teams to evaluate the content using different methodologies. the results from the gbifs and unzyec teams were similar and they both demonstrated the validity of the conclusions presented hereby. gaiji et al. content assessment of the primary biodiversity data 111 are the gbif-mediated data scientifically credible/reliable? gbif mobilised data is often criticized for errors (yesson et al., 2007; otegui et al., 2009). however, these errors are reflections of the data as collected, collated, and published by the heterogeneous data publishers across the globe. the role of gbif is to provide a discovery window on the published data. such a role requires reconciling, interpreting and publishing the essential key attributes: taxonomic, temporal and geospatial. in assessing the state of data quality in the gbif index over time, inevitably such study will combine data quality improvements at the level of the data publishers as well as at the central discovery point. the recent improvements made by gbif in the re-building of its taxonomical backbone and data quality checking routines have positively impacted on the level of data quality in the gbif index. however these improvements are explained by the improvements of the informatics infrastructure and processing algorithms, but these are not addressing the most critical underlying causes of poor data quality: accuracy and gaps. informatics routines alone can eventually spot but cannot recover missing attributes,(if anything, perhaps hint or guess), in particular when these attributes were not mapped correctly at the publisher level or if they weren’t even digitized from the original voucher specimen. taxonomy: a majority of the scientific names published by the gbif network are now recognized as valid references against a collection of authoritative taxonomic catalogues. the current gbif taxonomy backbone provides an appropriate resolution service to the large majority of the scientific names discovered by gbif. given the fact that gbif taxonomic backbone is a combination of multiple authoritative taxonomic catalogues (e.g. col, worms, ipni, ncbi, and itis etc.), it has potential to serve larger systematicians communities than any specific taxonomic group alone. while doing so, informatics approaches are proven to be effective; however, questions about future improvements can be raised. linkages with more authoritative taxonomic catalogues (recommendation 4 in faith et al., 2013) and involvement of taxonomic expertise will soon be required to resolve taxonomic discrepancies. in order to continuously assess the effectiveness of its taxonomic backbone, gbif secretariat should perform regular estimation of completeness at all taxonomic ranks as described in table 4.b. such an analysis should in particular assess the amount of mis-identifications (e.g. species within genus zonotrichia). gbif should also improve its reporting services to the original publishers so that potential taxonomic mis-identifications are reported (recommendation 6 in faith et al., 2013). gbif should also monitor over time the taxonomic data quality improvements made in the gbif index (e.g. indicators of taxonomic completeness at the class, order or family levels). in addition, gbif should provide means to assess the effectiveness of its taxonomic names resolution services used during the harvesting and indexing processes. all taxa mis-identifications should be documented and calls to expert groups (e.g. marine biologists, crop wild relatives experts) should be considered in order to tap into taxonomist expertise and increase their engagements (chavan et al., 2005) in improving the quality of such valuable global resource. temporal and geospatial: setting an ideal target for the rate of georeferenced occurrences within the gbif index is a difficult task. while the ideal scenario would be that all occurrences are georeferenced, the reality is that in many cases the original records for example specimens in zoological or botanical collections itself won’t have such information. some voucher specimens (especially older ones) have in general a lower percentage of geo-referenced records compared to recent field observation records. there is a high variability between data resources within the gbif index. however and as shown in gaiji et al. content assessment of the primary biodiversity data 112 figure 4, the average percentage of georeferenced records has increased between 1-5% in average during the period 1990-2010 and is consistently higher in february 2012 than in december 2010. the data publisher community is therefore addressing this challenge in particular for recent records. while gbif’s role is to enable the discovery of primary biodiversity data from a network of publishers (recommendation 14, faith et al., 2013), it is not mandated to undertake or correct the content published. however, this can be questioned in particular when a simple correction, such as a sign correction on a longitude or latitude field, could be undertaken within the gbif index and therefore immediately improve the quality of data published. taking into consideration the growing difficulties in communicating with a large network of publishers, such option may be considered for the most obvious data corrections. while this can be seen as a limitation, one way forward would be to set targets by periods where we observe low variation of the geo-referencing average (e.g. 1900-1930, 1930-1960, 1960-1990 and 1990-today). within each period, a georeferencing target could be set based on a subset of data resources (e.g. comparing all datasets publishing insect occurrences against the top 10% best georeferenced datasets). however, any decision on such baselines would need to be discussed and agreed with the community of publishers. experts could investigate datasets falling well under these baselines and reports with recommendations on possible corrections should be sent to the original publishers. however, this approach would require engagement from expert groups as well as willingness and availability of data owners to undertake more accurate verifications such as getting back to the original voucher specimen (recommendation 3, faith et al., 2013). as shown in figure 5, we demonstrated that in december 2010 approximately 50% of the occurrences in gbif index had at least one of the taxonomic, temporal and/or geospatial attributes missing; this percentage dropped to less than 22% by february 2012. this means that occurrence records with essential attributes represent now more than three quarters of the gbif index. taking into consideration that more recent occurrences tend to be of such quality, it is expected that over time this percentage will continue to increase. the most critical priority for the gbif network in this field is now to engage the data publisher community at large (including data curators and original collectors) to be (1) aware of the importance of data quality and accuracy; (2) alerted of the possible data gaps and/or quality issues identified centrally; and (3) investigate and fix these whenever possible (e.g. by checking the data publishing process up to involving the original curators and specimen) (recommendation 6, faith et al., 2013). to achieve this, a distributed annotation service will be required whereby reports on possible data quality issues are communicated to the original publishers. however such service would in turn require the promotion of effective identification of data objects such as persistent identifiers and sustainable resolution services (recommendation 13, faith et al., 2013). gbif should therefore place the use and re-use of persistent identifiers as a high priority activity and possibly as mandatory for all datasets (gbif, 2009). are the gbif-mediated data increasingly representative of “some” biodiversity on earth? the gbif index has recorded information about 995,974 species, which is a remarkable amount compared for example with the catalogue of life, which contains, as of june 2012, more than 1.3 million species. however, the gbif index is facing a bias toward the kingdom animalia and to a less extent towards plants. other kingdoms like fungi, protozoa and even bacteria are underrepresented within the data mobilized so far. therefore, the gbif index is not representative of all kingdoms and can’t be used yet as a proxy to all biodiversity on earth (recommendations 1 & 2, faith et al., 2013). gaiji et al. content assessment of the primary biodiversity data 113 the over-representation of bird observation data should not be considered as problematic, as the “over-“ bit means just by comparison to other groups. our study shows that the bird observation community has managed in particular over the last two decades to mobilize a vast amount of information on a number of species. such volume is remarkable and of great value to understand not only the distribution of species at a given time but also on a temporal basis. this is, for example, of immense value when dealing with monitoring potentially invasive alien species or reaction to climate change over time (sullivan et al., 2009). citizen scientists are using such bird observation network (e.g. ebird) data to monitor the biological patterns and the environmental and anthropogenic factors that influence them. these networks are providing today a near real-time observational network, a model to be followed by many other networks. however, what is of greater concern to the authors is the flat data mobilization rate since the 1990’s for classes other than aves (figure 13), a phenomenon masked by the approximately exponential growth of bird observation data, or even more critically (figure 14) by the exponential growth of the observation/specimen ratio. taking into consideration the greater intrinsic value of specimen versus observation data (e.g. accuracy, taxonomic validation, validation by experts, and availability of voucher specimen for verification), the stagnation of such valuable resources over the last 2 decades is a priority that needs to be addressed by gbif (recommendation 3, faith et al., 2013). instead of focusing on a volume target, such as the one and two-billionrecords goals as set by the gbif governing board in 2007 and 2009 respectively gbif (gbif, 2008, 2012b), gbif should instead focus on an optimal distribution of such volume across kingdoms, classes, orders, families and genera (recommendations 1, 2 & 8, faith et al., 2013). it is unrealistic to hope that the gbif network will manage within the next decade to mobilize as many data for all classes as what has been mobilized so far for birds. the volume of records has always been an easy and tempting target. we argue here that this is not an appropriate indicator of the success of gbif as being a window on earth biodiversity. we demonstrated in our study that even though the volume of occurrences has a clear bias towards the northern hemisphere (figure 18: 70% for northern america, western and northern europe table 15), in terms of species richness (figures 17) we do not observe such bias. therefore, even with fewer occurrences per species, the southern hemisphere shows to have a similar amount of species richness in the gbif index than the northern hemisphere. this can be explained by the recent addition of species-rich datasets from south africa (sanbi), australia (atlas of living australia), and costa rica (inbio). therefore if gbif needs to become a window on earth biodiversity, and taking into account the stagnation of specimen records and the exponential growth of observational data, one way forward would be to engage non-bird observation networks (e.g. flowering plants, snails, fish, butterflies etc.) to participate as actively as the bird observation community. because, the additions of biodiversityrich datasets with fewer records per species can make a large difference in the representativeness of gbif index of the earth biodiversity distribution (recommendations 1, 2, 3 & 8, faith et al., 2013]. another complementary strategy would be to focus on priority species derived from key scientific-policy priorities (e.g. reducing threats caused by invasive alien species, reducing the loss of threatened species etc.) and priority regions/areas (e.g. biodiversity hotspots, protected areas, high biodiversity regions) (recommendation 8, faith et al., 2013). therefore we conclude that if gbif wants to become the main window on earth biodiversity it needs to articulate data mobilization strategies engaging the full gbif network on a list of priority species and regions (recommendation 2, faith et al., 2013). to decide on these two major priorities, gbif needs to undertake a more advanced data gaiji et al. content assessment of the primary biodiversity data 114 gap analysis looking at what biodiversity needs to be monitored and in which locations (recommendations 1 & 10, faith et al., 2013). this would be somehow a radical shift from the former ‘opportunistic’ to a more pragmatic demand-driven approach (berents et al., 2010). this shift in strategy may however have some serious financial implications since most of the high biodiversity regions in the world (e.g. amazonia, tropical africa, etc.) are difficult of access or even dangerous (e.g. war zones or unstable regions). are the gbif-mediated data opening new opportunities for the scientific communities to assess the state of biodiversity as well as the pressure it faces and its response? our study on the ‘fitness-for-use’ of gbifmediated data can be summarized in figure 29. from a large volume of 323 million occurrences and 995 thousand species, the gbif index can be synthetized to a smaller volume of information (71.4 million occurrences) covering less than 50% of the known species in gbif. this first filter of the gbif index is based primarily on the availability of valid taxonomical references, temporal and geospatial elements. less than 200,000 species have sufficient occurrences (requiring a minimum presence in at least 10 distinct 1/10 degree grids) to be used to assess their distribution through ecological niche modelling (enm) analysis. still, this represents a large volume of valuable information that can be immediately used to assess for example the status of biodiversity for a group of species within a given ecosystem. our study also showed that many classes and families already have many species meeting these enm requirements. the growth of scientific literature using gbif mediated data in recent years is also an additional indicator that the gbif mediated data is a valuable resource. therefore we concluded that the gbif index is a valuable resource that can already be used by the scientific community to assess the status of biodiversity at least for major groups such as for birds, fish, plants and insects (rödder & lötters, 2010; ramírez-villegas et al., 2010; ready et al., 2010). enhancing the fitness-for-use and trustworthiness of gbif mobilised data is a natural course of action in this direction, that needs be attained urgently at all levels of data management chain (recommendation 5 & 6, faith et al., 2013). we further opine that gbif as a community needs to proactively advocate the use of gbif mediated data in scientific analysis, which may result into sound decision making and effective conservation and sustainable uses of biological resources (recommendation 9, faith et al., 2013). in order to encourage the cross-sectional scientific and naturalist communities in publishing primary biodiversity data through the gbif network, a comprehensive ‘data publishing framework’ (chavan & ingwersen, 2009; moritz et al., 2011) needs to be promoted and implemented (recommendation 11, faith et al., 2013). conclusion the third global biodiversity outlook (gbo3) published by the cbd in 2010 concluded that the target agreed by the world’s governments in 2002, “to achieve by 2010 a significant reduction of the current rate of biodiversity loss at the global, regional and national level as a contribution to poverty alleviation and to the benefit of all life on earth”, was not met (scbd, 2010). the loss of biodiversity is an issue of profound concern for its own sake, but biodiversity also underpins the functioning of ecosystems, which provide a wide range of services to human societies (scbd, 2010). its continued loss, therefore, has major implications for current and future human wellbeing. the lack of a consistent baseline data and ongoing monitoring of biodiversity has been often cited as a major obstacle towards improving the scientific evidence of the consequences of biodiversity loss. gbif through its mission is providing a mean to achieve greater improvements in this evidence base in particular through its global network and the discovery, access and use gaiji et al. content assessment of the primary biodiversity data 115 of the largest resources of primary biodiversity data. this study has clearly demonstrated the great value of the gbif mediated data in various aspects from improved data quality and accuracy, progressive reduction of gaps in the content and increased fitness-for-use. more importantly, the gbif mediated data is offering today opportunities to undertake scientific research that has never been possible before. access to three fundamental and essential biodiversity variables (i.e. taxonomy, geospatial and temporal) opens opportunities for example in assessing today’s distribution of species as well as predicting their future distribution. taking ecological niche modelling as the model application of gbif-mediated data already opens a myriad of research opportunities, such as assessing the state of biodiversity (e.g. threatened species assessments, genetic diversity, etc.) or the pressures (e.g. invasive alien species, effect of climate change or land cover change) as well as the responses. gbif is therefore uniquely positioned today to become the ‘data to science’ interface in support of major scientific research trends such as in support of the strategic plan on biodiversity as agreed in nagoya in 2010 (scbd, 2012). gbif must therefore take concrete steps toward effective monitoring of current trends in science and policy, such that it is maximally responsive and effective as a mega-science data infrastructure. this is how gbif’s focus can remain on being the single most important infrastructure for primary biodiversity data at the organism level, accompanied by strong, effective, and targeted links to data at the genetic, genomic, and ecosystem levels. acknowledgements authors are grateful to their colleagues in the global biodiversity information facility secretariat, and university of navarra for their support in this study and providing access to the gbif index. references ariño a.h. (2010). approaches to estimating the universe of natural history collections data. biodiversity informatics, 7: 81-92. ariño a.h., otegui j. (2008). sampling biodiversity sampling. proceedings of tdwg, 2008: 77-78. ariño a.h., pimm s. (1995). on the nature of population extremes. evolutionary ecology, 9: 429443. berendsohn w., chavan v., macklin j.a. (2010). recommendations of the gbif task group on the global strategy and action plan for the mobilization of natural history collections data. biodiversity informatics, 7: 67-71. berents p., hamer m., chavan v. (2010). towards demand-driven publishing: approaches to the prioritization of digitisation of natural history collections data. biodiversity informatics 7, 113119. chavan v., ingwersen p. (2009). towards a data publishing framework for primary biodiversity data: challenges and potentials for the biodiversity informatics community. bmc bioinformatics, 10 (suppl. 14): s2. chavan v., rane n., watve a., ruggiero m. (2005). resolving taxonomic discrepancies: role of electronic catalogues of known organisms. biodiversity informatics 2: 70-78. collen b., ram m., zamin t., louise m. (2008). the tropical biodiversity data gap: addressing disparity in global monitoring. tropical conservation science, 1(2): 75-88. faith d., collen b., ariño a.h., koleff p., guinotte j., kerr j., chavan v. (2012). bridging the biodiversity data gaps: recommendations of the gbif content needs assessment task group. biodiversity informatics, 2013. fernández m.a., blum s.d., reichle s., guo q., holzman b., hamilton h. (2009). locality uncertainty and the differential performance of four gaiji et al. content assessment of the primary biodiversity data 116 common niche-based modeling techniques. biodiversity informatics 6: 36-52. bisby. f.a. (2000). the quiet revolution: biodiversity informatics and the internet. science 289: 23092312. gaikwad j. chavan v. (2005). open access and biodiversity conservation: challenges and potentials for the developing world. data science journal, 5: 1-17. gbif (2008). gbif work programme 2009-2010. copenhagen: global biodiversity information facility, 59pp. accessible at http://www2.gbif.org/wp2009-10.pdf. gbif (2009) adoption of persistent identifiers for biodiversity informatics recommendations of the gbif lsid guid task group, 6 november 2009 (http://imsgbif.gbif.org/cms_orc/?doc_id=2956) gbif (2010a). gbif position paper on future directions and recommendations for enhancing fitness-for-use across the gbif network, version 1.0. authored by hill, aw, otegui j, ariño ah, and rp guralnick. 2010, copenhagen: global biodiversity information facility, 25 pp. isbn: 8792020-11-9. accessible online at http://www2.gbif.org/gpp-final.pdf. gbif. (2010b). state-of-the-network 2010: discovery and publishing of the primary biodiversity data through the gbif network. authored by chavan, v. s., gaiji, s., hahn, a., sood, r. k., raymond, m., and n. king. 2010. copenhagen: global biodiversity information facility, xx pp. isbn: 8792020-13-5. accessible online at http://www.gbif.org. gbif (2010b). best practice guide for ‘data discovery and publishing strategy and action plans’ version 1.0. authored by chavan, vs, sood rk, and ah arino. 2010. copenhagen: global biodiversity information facility, 29 pp. isbn: 87-92020-12-7. accessible online at http://www.gbif.org/bestpracticeguide-final.pdf. gbif (2012.a). monthly statistics accessible online at: http://www.gbif.org/communications/resources/mo nthly-statistics/ as of 21 june 2012 gbif (2012.b). history oif gbif –m accessible online at: http://www.gbif.org/communications/press/historyof-gbif/ grantham h.s., moilanen a., wilson k.a., pressey r.l., rebelo t.g., possingham h.p. (2008). diminishing returns on investment for biodiversity data conservation planning. conservation letters, 1(4): 190-198. grinnell j. (1917). field tests of theories concerning distributional control. american naturalist 51:115. guralnick r.p., hill a.h., lane m. (2007). towards a collaborative, global infrastructure for biodiversity assessment. ecology letters, 10(8): 663-672. hanson t., brooks t.m, da fonseca g.a.b, hoffmann m., lamoreux j.f, machlis g, mittermeier c.g, mittermeier r.a, pilgrim j.d. (2009). warfare in biodiversity hotspots. conservation biology, volume 23, no. 3, 578–587 hobern d. (2003). an introduction to gbif biodiversity informatics (version 1.0 – final). gbif, 25 pp, accessible at: http:// http://www.gbif.es/ficheros/gbifbiodiversityinfor maticsintroduction-v1.0-final.pdf iucn and unep (2009). the world database of protected areas (wdpa). 2010 bip indicator: coverage of protected areas. unep-wcmc. cambridge, uk. http://www.wdpa.org/statistics.aspx johnson n.f. (2007). biodiversity informatics. annual review of entomology, 52: 421-438. oecd (1999). final report of the oecd megascience forum working group on biological informatics, january 1999, pp. 74, accessible at http://www.oecd.org/dataoecd/24/32/2105199.pdf. otegui j., ariño a.h. (2013). on the time series of records mobilised through gbif otegui j, ariño a.h. (2012b). biddsat: visualizing the content of biodiversity data publishers in the gbif network. bioinformatics, 2012: doi: 10.1093/bioinformatics/bts359. otegui j., robles e., ariño a.h. (2009). noise in biodiversity data. poster presented at e-biosphere 09. e-biosphere conference 2009, conference abstracts, london; mcleod n. & j. edwards, eds. p. 190. http://www.e-biosphere09.org/assets/files/ebiosphere%20abstracts%20volume%20%20final.pdf mittermeier r.a., robles-gil p., hoffmann m., pilgrim j., brooks t., mittermeier c.g., lamoreux j., da http://www2.gbif.org/wp2009-10.pdf http://www.gbif.org/ http://www.gbif.org/bestpracticeguide-final.pdf http://www.gbif.org/communications/resources/monthly-statistics/ http://www.gbif.org/communications/resources/monthly-statistics/ http://www.gbif.org/communications/press/history-of-gbif/ http://www.gbif.org/communications/press/history-of-gbif/ http://www.oecd.org/dataoecd/24/32/2105199.pdf gaiji et al. content assessment of the primary biodiversity data 117 fonseca g.a.b. (2004). hotspots revisited. cemex, mexico. page r. (2012). how many specimens does gbif really have? retrieved august 2, 2012, from http://iphylo.blogspot.com.es/2012/02/how-manyspecimens-does-gbif-really.html pearson r.g., raxworthy c.j., nakamura m., townsend peterson a. (2007). predicting species distributions from small numbers of occurrence records: a test case using cryptic geckos in madagascar. journal of biogeography, 34, 102-117. peterson a., vieglais d. (2001). predicting species invasions using ecological niche modelling: new approaches from bioinformatics attsack a pressing problem. bioscience 51: 363-371. pino-del-carpio a., villarroya a., ariño a.h., puig j., miranda r. (2011). communication gaps of knowledge of freshwater fish biodiversity: implications for the management and conservation of mexican biosphere reserves. journal of fish biology, 79(6): 1563-1591. ramírez-villegas j., khoury c., jarvis a., debouck d. g., guarino l. (2010). a gap analysis methodology for collecting crop genepools: a case study with phaseolus beans. plos one, 5(10), e13497. public library of science. retrieved from http://dx.doi.org/10.1371/journal.pone.0013497 ready j., kaschner k., south a.b., eastwood p.d., rees t., rius j., agbayani e., et al. (2010). predicting the distributions of marine organisms at the global scale. ecological modelling, 221(3), 467-478. doi:doi: 10.1016/j.ecolmodel.2009.10.025 rödder d., lötters s. (2010). potential distribution of the alien invasive brown tree snake, boiga irregularis (reptilia: colubridae). pacific science, 64(1), 11-22. university of hawai’i press. doi:10.2984/64.1.011 secretariat of the convention on biological diversity (2010). global biodiversity outlook 3. montreal, 94 pages. secretariat of the convention on biological diversity (2012). strategic plan for biodiversity 2011-2020, including aichi biodiversity targets accessible at: http://www.cbd.int/sp/ sullivan b.l., wood c.l., iliff m.j., bonney r.e., fink d., kelling s. (2009). ebird: a citizen-based bird observation network in the biological sciences. biological conservation, 142(10), 2282-2292. doi:doi: 10.1016/j.biocon.2009.05.006 warner s., linburg k., ariño a.h., dushoff j., dodd m., stergiou k., potts j. (1995). time series compared across the land-sea gradient. pp. 242273 en: powell t.m. & steele j.h. (eds.:) ecological time series. chapman & hall, new york etc. yesson c., brewer p.w., sutton t., caithness n., pahwa j.s., et al. (2007) how global is the global biodiversity information facility? plos one 2(11): e1124. doi:10.1371/journal.pone.0001124 list of tables table 1. essential core data elements (in the gbif-index occurrence table). table 2. resources currently available through gbif checklistbank used to build the gbif taxonomical backbone. table 3. taxonomical rank matching with catalogue of life 2011 (february 2012) table 4.a. scientific names and occurrences summary for each ‘unknown’ taxonomic rank (as of december 2010). table 4.b. scientific names and occurrences summary for each ‘unknown’ taxonomic rank (as of february 2012). table 5: potential misidentification at the kingdom rank and tentative resolution through col 2011 and more recent version (february 2012) http://iphylo.blogspot.com.es/2012/02/how-many-specimens-does-gbif-really.html http://iphylo.blogspot.com.es/2012/02/how-many-specimens-does-gbif-really.html http://dx.doi.org/10.1371/journal.pone.0013497 gaiji et al. content assessment of the primary biodiversity data 118 table 6.a: estimation of the taxonomical misidentification at the kingdom level (february 2012) table 6.b: estimation of the taxonomical misidentification at the phylum level (february 2012) table 6.c: estimation of the taxonomical misidentification at the class level (february 2012) table 7. percentage (%) temporal quality of the gbif mobilised data records according to gbifs methodology. table 8. breakdown of the year provided in the gbif mobilised data records. table 9. breakdown of the temporal and geospatial data availability. table 10. summary of potential resource duplicates. table 11. top 20 potential resources duplicates. table 12.a: top 15 species with highest number of data records (december 2010). table 12.b: top 15 species with highest number of data records (february 2012). table 13: breakdown of the species and occurrences richness for kingdom animalia and plantae (february 2012). table 14. breakdown of occurrences by continents. table 15. proportion of financial contribution to gbif (2010) by continents. table 16. distribution of occurrences per members to the organisation for economic co-operation and development (oecd). table 17. distribution of occurrences per country income status. table 18.a grid (1/10 degree) occupancy (february 2012). table 18.b grid (1/10 degree) occupancy for animalia. table 18.c grid (1/10 degree) occupancy for plantae. table 19. grid (1/10 degree) occupancy by classes (animalia and plantae). table 20. grid (1/10 degree) occupancy by families (animalia and plantae). list of figures figure 1. typical flow of data discovered and published through the gbif network. figure 2. data mining methodology employed during content assessment exercise carried out by the gbif secretariat. figure 3. data mining methodologies employed during content assessment exercise carried out by the university of navarra. figure 4. evolution of the percentage of geo-referenced records (1800-2010). figure 5. evolution of incomplete records (taxonomical*temporal*geospatial) in the gbif-index figure 6.a. data records by kingdom (december 2010). figure 6.b. data records by kingdom (february 2012). figure 7.a. data records by phylum (december 2010). figure 7.b. data records by phylum (february 2012). figure 8.a. data records by class (december 2010). figure 8.b. data records by class (february 2012). figure 9. treemap of occurrences according to taxon group, down to class. cell surface proportional to number of occurrences. blue: invertebrates; purple: vertebrates; green: higher plants; yellow: algae and ferns; brown: fungi; red: unicellular organisms. figure 10. top 10 data resources publishing maximum number of data records (december 2010). gaiji et al. content assessment of the primary biodiversity data 119 figure 11. kingdom wise distribution of data records by year. figure 12.a. plantae kingdom wise distribution of occurrences by year (february 2012). figure 12.b. animalia kingdom wise distribution of occurrences by year (february 2012). figure 13. aves wise distribution of occurrences by year (february 2012) figure 14. basis of records distribution of occurrences by year (february 2012) figure 15. distribution of voting and associate country participants (february 2012) figure 16. distribution of occurrences by latitude (february 2012) figure 17. distribution of species richness by latitude (february 2012) figure 18. distribution of the average number of occurrences per species by latitude (february 2012) figure 19. georeferenced records (dots) published by the largest data providers in europe. colors represent distinct providers. (from otegui et al., 2009). figure 20. records available in the gbif-index by date figure 21. variation of the records available in the gbif-index by date (december 2010 – february 2012) figure 22. distribution of distinct species collected/observed over time (february 2012) figure 23. distribution of distinct resources contributing to the gbif-index over time (february 2012) figure 24. grid occupancy (1/2 degree grid) of the gbif-index over time (february 2012) figure 25. grid occupancy (1/10 degree grid) of the gbif-index by kingdom (february 2012) figure 26. grid occupancy (1/10 degree grid) of the gbif-index by classes (february 2012) figure 27. distribution of families for ffu-enm – kingdom animalia (february 2012) figure 28. distribution of families for ffu-enm – kingdom plantae (february 2012) figure 29. selection of fit-for-use records in the gbif index for ecological niche modelling (enm). gaiji et al. content assessment of the primary biodiversity data 120 table 1. essential core data elements (in the gbif-index occurrence table). title description publisher publisher of the resource/dataset dataset resource/dataset institution the name (or acronym) in use by the institution having custody of the object(s) or information referred to in the record. collection the name, acronym, code, or initials identifying the collection or data set from which the record was derived. catalogue number an identifier (preferably unique) for the record within the data set or collection. scientific name the full scientific name, with authorship and date information if known. when forming part of identification, this should be the name in lowest level taxonomic rank that can be determined. taxon author the authorship information for the scientific name. taxon rank the taxonomic rank of the most specific name in the scientific name. recommended best practice is to use a controlled vocabulary. kingdom the full scientific name of the kingdom in which the taxon is classified. phylum the full scientific name of the phylum or division in which the taxon is classified. class the full scientific name of the class in which the taxon is classified. order the full scientific name of the order in which the taxon is classified. family the full scientific name of the family in which the taxon is classified. genus the full scientific name of the genus in which the taxon is classified. species epithet the name of the first or species epithet of the scientific name. gaiji et al. content assessment of the primary biodiversity data 121 infraspecific epithet the name of the lowest or terminal infraspecific epithet of the scientific name, excluding any rank designation. latitude the geographic latitude (in decimal degrees) of the geographic center of a location. positive values are north of the equator; negative values are south of it. legal values lie between -90 and 90, inclusive. longitude the geographic longitude (in decimal degrees) of the geographic center of a location. positive values are east of the greenwich meridian, negative values are west of it. legal values lie between -180 and 180, inclusive. coordinate precision a decimal representation of the precision of the coordinates given in the latitude and longitude. maximum altitude the upper limit of the range of elevation (altitude, usually above sea level), in meters. minimum altitude the lower limit of the range of elevation (altitude, usually above sea level), in meters. altitude precision a decimal representation of the precision of the altitude. minimum depth the lesser depth of a range of depth below the local surface, in meters. maximum depth the lesser depth of a range of depth below the local surface, in meters. depth precision a decimal representation of the precision of the depth. continent or ocean the name of the continent in which the location occurs. recommended best practice is to use a controlled vocabulary such as the getty thesaurus of geographic names or the iso 3166 continent code. recommended best practice is to use a controlled vocabulary such as the getty thesaurus of geographic names. country the name of the country or major administrative unit in which the location occurs. recommended best practice is to use a controlled vocabulary such as the getty thesaurus of geographic names. state or province the name of the next smaller administrative region than country (state, province, canton, department, region, etc.) in which the location occurs. county the full, unabbreviated name of the next smaller administrative region than state or province (county, shire, department, etc.) in which the location occurs. gaiji et al. content assessment of the primary biodiversity data 122 name of collector/observer a list (concatenated and separated) of names of people, groups, or organizations responsible for recording the original occurrence. locality the specific description of the place. less specific geographic information can be provided in other geographic terms. this term may contain information modified from the original to correct perceived errors or standardize the description. year of collection the four-digit year in which the collection or observation event occurred, according to the common era calendar. month of collection the ordinal month in which the collection or observation event occurred. day of collection the integer day of the month on which the collection or observation event occurred. basis of record the specific nature of the data record. recommended best practice is to use a controlled vocabulary such as the darwin core type vocabulary (http://rs.tdwg.org/dwc/terms/type-vocabulary/index.htm). name of identifier a list (concatenated and separated) of names of people, groups, or organizations that assigned the taxon to the subject. identification date the date on which the subject was identified as representing the taxon. recommended best practice is to use an encoding scheme, such as iso 8601:2004(e). date of creation timestamp of creation of this raw occurrence record in the index. date of modification timestamp of last update of this raw occurrence record in the index. date of deletion timestamp of deletion of this raw occurrence record in the index (obsolete). http://rs.tdwg.org/dwc/terms/type-vocabulary/index.htm gaiji et al. content assessment of the primary biodiversity data 123 table 2. top 10 resources currently available through gbif ‘checklistbank’ used to build the gbif taxonomical backbone. title version families genera species the catalogue of life 2012-01-14 8,149 129,461 1,379,178 register of marine and nonmarine genera (irmng) 2012-01-13 34,119 790,025 1,017,851 international plant names index 2011-07-13 791 59,766 1,317,317 ncbi taxonomy 2012-01-13 7,223 59,404 668,915 the integrated taxonomic information system (itis) 2012-01-14 6,972 45,531 306,358 world register of marine species 2012-05-02 6,370 41,293 233,811 index fungorum 2011-07-13 2,926 10,569 267,553 fauna europaea 2011-07-13 37,214 131,671 wikipedia species pages english 2011-09-04 grin taxonomy for plants 2012-01-14 492 12,909 58,773 a full up-to-date list can be accessed at: http://ecat-dev.gbif.org/ gaiji et al. content assessment of the primary biodiversity data 124 table 3. taxonomical rank matching with catalogue of life 2011 (february 2012) taxonomical rank matching with catalogue of life 2011 percentage of the gbifindex percentage of the total number of species k in g d o m p h y lu m c la ss o rd er f am il y g en u s s p ec ie s (324,247,283 occurrences) (995,974 species in total) 0.05% 0.27% ✔ 1.38% 0.29% ✔ ✔ 0.53% 0.89% ✔ ✔ ✔ 0.77% 1.35% ✔ ✔ ✔ ✔ 0.76% 2.40% ✔ ✔ ✔ ✔ ✔ 4.54% 13.36% ✔ ✔ ✔ ✔ ✔ ✔ 9.13% 27.98% ✔ ✔ ✔ ✔ ✔ ✔ ✔ 82.83% 53.47% 100.00% 100.00% table 4.a. scientific names and occurrences summary for each ‘unknown’ taxonomic rank (as of december 2010). taxonomy scientific name with ‘unknown’ status % of total species recorded in gbif occurrences with ‘unknown’ status % of total occurrences recorded in gbif index kingdom 114,721 7.0% 15,030,014 5.6% phylum 223,433 13.8% 22,180,639 8.3% class 235,857 14.5% 23,071,180 8.6% order 261,706 16.1% 24,605,925 9.2% family 235,089 14.5% 21,508,688 8.1% genus 76,416 4.7% 8,665,178 3.2% species 120,362 7.4% 23,015,905 8.6% gaiji et al. content assessment of the primary biodiversity data 125 table 4.b. scientific names and occurrences summary for each ‘unknown’ taxonomic rank (as of february 2012). taxonomy scientific name with ‘unknown’ status % of total species recorded in gbif occurrences with ‘unknown’ status % of total occurrences recorded in gbif index kingdom 5,153 0.35% 167,208 0.05% phylum 11,305 0.77% 4,640,252 1.43% class 26,266 1.81% 3,963,750 1.22% order 52,007 3.58% 6,304,444 1.94% family 41,932 2.82% 6,015,636 1.86% genus 31,565 2.17% 8,959,016 2.76% species 133,086 9.15% 25,343,834 7.82% table 5: potential misidentification at the kingdom rank and tentative resolution through col 2011 and more recent version (february 2012) kingdom species occurrences col 2011 plantae zonotrichia albicollis 775,671 accepted name in col 2012 in animalia kingdom plantae zonotrichia leucophrys 362,767 accepted name in col 2012 in animalia kingdom protozoa neogloboquadrina pachyderma 141,720 not in col 2011 accepted in col 2012 plantae zonotrichia atricapilla 106,804 accepted name in col 2012 in animalia kingdom protozoa globigerinoides ruber 86,563 not in col 2011 accepted name in col 2012 protozoa globigerina bulloides 82,643 not in col 2011 accepted name in col 2012 protozoa globigerinita glutinata 74,617 not in col 2011 accepted in col 2012 gaiji et al. content assessment of the primary biodiversity data 126 protozoa globorotalia truncatulinoides 64,707 not in col 2011 not col 2012 identified in marine species identification portal (as of feb 2012) 20 protozoa globorotalia inflata 57,706 not in col 2011 not in col 2012 identified in marine species identification portal (as of feb 2012) 21 protozoa orbulina universa 55,290 not in col 2011 not in col 2012 identification portal (as of feb 2012) 22 table 6.a: estimation of the taxonomical misidentification at the kingdom level (february 2012) incorrect kingdom assignment correct kingdom in col 2011 occurrences species plantae animalia 1,308,111 90 animalia plantae 1,536 26 chromista animalia 1,504 1 chromista plantae 310 3 animalia fungi 190 23 fungi animalia 186 10 protozoa chromista 100 8 plantae fungi 98 11 plantae protozoa 61 2 plantae chromista 43 6 20 http://species-identification.org/species.php?species_group=zsao&id=1387 21 http://species-identification.org/species.php?species_group=zsao&id=1384 22 http://species-identification.org/species.php?species_group=zsao&id=1397 gaiji et al. content assessment of the primary biodiversity data 127 fungi plantae 41 2 animalia chromista 26 3 bacteria protozoa 22 1 plantae bacteria 13 5 protozoa plantae 9 5 protozoa fungi 6 1 fungi protozoa 2 1 total 1,312,258 198 table 6.b: estimation of the taxonomical misidentification at the phylum level (february 2012) incorrect phylum assignment correct phylum in col 2011 occurrences species bryophyta magnoliophyta 17,488 24 magnoliophyta arthropoda 2,788 36 cnidaria chordata 2,213 12 ochrophyta arthropoda 1,504 1 chordata magnoliophyta 833 5 cyanobacteria proteobacteria 312 5 ochrophyta rhodophyta 309 2 arthropoda magnoliophyta 297 10 arthropoda chlorophyta 244 1 magnoliophyta chordata 201 5 magnoliophyta cnidaria 176 2 labyrinthista sarcomastigophora 116 4 ascomycota chordata 115 1 marchantiophyta bryozoa 114 1 annelida tardigrada 111 1 arthropoda ascomycota 93 6 arthropoda rhodophyta 82 2 magnoliophyta ascomycota 80 2 gaiji et al. content assessment of the primary biodiversity data 128 mollusca ascomycota 48 10 bryozoa magnoliophyta 47 4 sarcomastigophora ochrophyta 46 4 ascomycota magnoliophyta 40 1 ascomycota arthropoda 39 3 chlorophyta magnoliophyta 34 2 brachiopoda ascomycota 32 2 mollusca arthropoda 30 2 arthropoda nematoda 26 1 arthropoda ochrophyta 25 2 annelida magnoliophyta 24 1 pinophyta arthropoda 23 1 platyhelminthes arthropoda 22 9 rhodophyta arthropoda 21 1 basidiomycota arthropoda 19 2 arthropoda bacillariophyta 17 5 chlorophyta rhodophyta 16 2 chlorophyta cyanobacteria 13 5 echinodermata arthropoda 12 1 bacteroidetes proteobacteria 11 3 magnoliophyta ochrophyta 8 5 bryophyta rhodophyta 8 1 ascomycota bryozoa 7 2 echinodermata cnidaria 6 1 magnoliophyta rotifera 6 1 ciliophora chlorophyta 4 1 arthropoda mollusca 4 1 arthropoda pinophyta 4 1 ascomycota bacillariophyta 3 2 euglenozoa rhodophyta 3 2 gaiji et al. content assessment of the primary biodiversity data 129 annelida bacillariophyta 3 1 ascomycota cnidaria 3 1 ascomycota echinodermata 3 1 chlorophyta arthropoda 2 1 platyhelminthes bacillariophyta 2 1 magnoliophyta bacillariophyta 2 1 platyhelminthes acanthocephala 1 1 pteridophyta arthropoda 1 1 ascomycota chlorophyta 1 1 cnidaria ochrophyta 1 1 arthropoda platyhelminthes 1 1 dinophyta rhodophyta 1 1 total 27,695 210 table 6.c: estimation of the taxonomical misidentification at the class level (february 2012) incorrect class assignment correct class in col 2011 occurrences species bryopsida andreaeopsida 21,966 82 bryopsida liliopsida 17,488 24 magnoliopsida insecta 2,429 12 hydrozoa arachnida 2,213 12 phaeophyceae insecta 1,504 1 actinopterygii magnoliopsida 808 2 insecta malacostraca 664 11 malacostraca insecta 359 1 phaeophyceae florideophyceae 309 2 insecta trebouxiophyceae 244 1 liliopsida insecta 239 15 lecanoromycetes dothideomycetes 220 1 gaiji et al. content assessment of the primary biodiversity data 130 magnoliopsida hydrozoa 176 2 liliopsida andreaeopsida 175 1 insecta magnoliopsida 167 3 insecta liliopsida 116 6 labyrinthulea polycystina 116 4 jungermanniopsida gymnolaemata 114 1 polychaeta eutardigrada 111 1 insecta florideophyceae 82 2 magnoliopsida lecanoromycetes 78 1 liliopsida maxillopoda 59 1 insecta lecanoromycetes 58 1 magnoliopsida arachnida 57 7 lobosa coscinodiscophyceae 54 4 stenolaemata magnoliopsida 47 4 zoomastigophora craspedophyceae 45 3 lecanoromycetes magnoliopsida 40 1 lecanoromycetes insecta 35 1 insecta leotiomycetes 35 5 chlorophyceae liliopsida 34 2 rhynchonellata lecanoromycetes 32 2 ostracoda secernentea 26 1 actinopterygii liliopsida 24 2 polychaeta liliopsida 24 1 pinopsida insecta 23 1 turbellaria arachnida 22 9 florideophyceae insecta 21 1 agaricomycetes insecta 19 2 magnoliopsida actinopterygii 18 2 insecta agaricomycetes 17 5 chlorophyceae florideophyceae 16 2 gaiji et al. content assessment of the primary biodiversity data 131 maxillopoda liliopsida 14 1 asteroidea entognatha 12 1 sphingobacteria alphaproteobacteria 8 2 bryopsida florideophyceae 8 1 insecta phaeophyceae 8 1 lecanoromycetes gymnolaemata 7 2 magnoliopsida eurotatoria 6 1 granuloreticulosea lecanoromycetes 6 1 magnoliopsida reptilia 6 1 magnoliopsida coscinodiscophyceae 5 3 ciliatea chlorophyceae 4 1 magnoliopsida entognatha 4 1 ostracoda gastropoda 4 1 dothideomycetes insecta 4 2 magnoliopsida liliopsida 4 2 insecta pinopsida 4 1 rhabditophora turbellaria 4 1 dothideomycetes asteroidea 3 1 polychaeta bacillariophyceae 3 1 dothideomycetes eurotiomycetes 3 1 euglenida florideophyceae 3 2 flavobacteria gammaproteobacteria 3 1 leotiomycetes hydrozoa 3 1 liliopsida actinopterygii 2 1 magnoliopsida agaricomycetes 2 1 turbellaria bacillariophyceae 2 1 lecanoromycetes granuloreticulosea 2 1 ulvophyceae insecta 2 1 pezizomycetes leotiomycetes 2 2 magnoliopsida phaeophyceae 2 1 gaiji et al. content assessment of the primary biodiversity data 132 liliopsida coscinodiscophyceae 1 1 eurotiomycetes dothideomycetes 1 1 dinophyceae florideophyceae 1 1 filicopsida insecta 1 1 appendicularia liliopsida 1 1 gastropoda orbiliomycetes 1 1 neoophora palaeacanthocephala 1 1 anthozoa phaeophyceae 1 1 zoomastigophora synurophyceae 1 1 arachnida turbellaria 1 1 50,434 290 table 7. percentage (%) temporal quality of the gbif mobilised data records according to gbifs methodology. december 2010 february 2012 difference occurrences with no year provided 82,300,746 42,890,654 -47.9% percentage of the gbif index 30.8 % 13,2% gaiji et al. content assessment of the primary biodiversity data 133 table 8. breakdown of the year provided in the gbif mobilised data records. raw refers to records as supplied by the publisher, whereas occ indicates records available through the portal after processing. according to unzyec methodology. “not valid” year includes years supplied as <1750 (including explicit zero) or in the future. “null” includes records with year provided as null value but do not include years explicitly stated as a numerical zero value. “matching/not matching” indicates whether the value for year in a record matches between the raw data collected from providers (raw) and the processed data made available through the portal (occ). valid occ not valid occ null value % raw valid matching 63,9% --63,9% not matching -0,1% 8,6% 8,7% not valid matching -0,2% -0,2% not matching 0,1% -0,4% 0,5% null matching --17,7% 17,7% not matching 5,0% 3,9% -8,9% total 69,1% 4,3% 26,7% 100,0% 31% table 9. breakdown of the temporal and geospatial data availability with year without year total georeferenced 255,4 (78.8%) 24.4 (7.5%) 279.8 (86.3%) not georeferenced 25.9 (8%) 18.5 (5.7%) 44.4 (13.7%) total 281.3 (86.8%) 42.9 (13.2%) 324.4 (100%) (in million occurrences) table 10. summary of potential resource duplicates estimated potential duplicates number of resources combinations total number of ‘potential’ duplicates’ >100.000 42 30.905.772 10.000 – 100.000 215 5.460.179 gaiji et al. content assessment of the primary biodiversity data 134 1.000 – 10.000 560 1.989.482 100 – 1.000 820 309.663 1.637 38.665.096 table 11. top 20 potential resources duplicates resource name (1) resource name (2) 'potential' duplicates biodiversidad de costa rica especímenes inbio 12,993,467 cnin/lepidoptera colección de referencia de lepidópteros diurnos mexicanos de la cnin (lepidoptera: papilionoidea) (ibunam) 1,779,872 (appendix 1) planktonic foraminifera abundances in odp site 181-1123 (appendix 1) census data of planktic foraminiferal faunas together with estimates of mean annual sst for odp site 181-1123 1,620,028 planktic foraminifera counts of sediment core md95-2040 planktonic foraminifera, stable isotope record and temperature reconstruction of sediment core md95-2040 1,514,214 pelagic fish observations 1968-1999 pelagic fish observations 1968-1999 (australian antarctic data centre) 1,288,625 (fig. 2) abundance of neogloboquadrina pachyderma sinistral in sediment core md95-2040 planktic foraminifera counts of sediment core md95-2040 1,132,425 (appendix b5) distribution of planktic foraminifera in dsdp site 90-594 east of new zealand (appendix 1) census data of planktic foraminiferal faunas together with estimates of mean annual sst for dsdp site 90-594 770,149 planktic foraminifera counts of sediment core md95-2040 (appendix 4) stable oxygen isotope record of globigerina bulloides and abundances of neogloboquadrina pachyderma and ice-rafted debris in sediment core md95-2040 761,421 birds (kiee-bi) birds (mnhm-bi) 593,674 gaiji et al. content assessment of the primary biodiversity data 135 benthic foraminifera abundance in counts of hole prad1-2 benthic foraminifera abundance in per cent of hole prad1-2 585,766 (table 3) distribution and abundance of selected planktonic foraminifera of the pliocene dsdp hole 41-366a (table 4) distribution and abundance of selected planktonic foraminifera of the pliocene dsdp hole 41-366a 515,705 (appendix 3) assemblage of benthic foraminifera in sediment core m5/2_kl15 relative abundance of benthic foraminifera in sediment core m5/2_kl15 480,479 birds (uwep-bi) birds (kiee-bi) 479,080 planktic foraminifera counts of sediment core md95-2040 (fig. 8g-h, 11) abundance of planktonic foraminifera and estimation of sea surface temperature and export production of sediment core md952040 420,615 planktic foraminifera abundance in counts of hole prad1-2 planktic foraminifera abundance in per cent of hole prad1-2 395,733 (fig. 2) abundance of neogloboquadrina pachyderma sinistral in sediment core md95-2040 planktonic foraminifera, stable isotope record and temperature reconstruction of sediment core md95-2040 368,550 (table 3) occurrences of planktonic foraminifers in samples from odp hole 105-647a (table 2) occurrences of planktonic foraminifers in samples from odp hole 105-647a 321,750 planktic foraminifera abundance of hole 41369a (table 2) distribution and abundance of selected planktonic foraminifera of the pliocene dsdp hole 41-369a 286,390 hatikka observation data gateway tiira information service 281,119 (appendix a) stable carbon and oxygen isotopes and paleoproductivity reconstructions for the last 550 kyr of odp hole 130-807a from the ontong java plateau, pacific ocean (appendix a) benthic foraminiferal assemblages in sediments of the last 550 kyr of odp hole 130-807a from the ontong java plateau, pacific ocean 249,300 gaiji et al. content assessment of the primary biodiversity data 136 table 12.a: top 15 species with highest number of data records (december 2010). occurrences (% of total) georeferenced (%) zenaida macroura 2,163,341 0.8% 99.8% cardinalis cardinalis 1,953,522 0.7% 99.8% passer domesticus 1,892,301 0.7% 94.9% sturnus vulgaris 1,852,357 0.7% 98.0% junco hyemalis 1,735,767 0.6% 99.4% cyanocitta cristata 1,697,922 0.6% 99.8% picoides pubescens 1,673,374 0.6% 99.8% carduelis tristis 1,671,365 0.6% 99.8% carpodacus mexicanus 1,611,847 0.6% 99.7% poecile atricapillus 1,518,874 0.6% 99.9% corvus brachyrhynchos 1,369,418 0.5% 99.8% baeolophus bicolor 1,259,308 0.5% 99.9% sitta carolinensis 1,244,000 0.5% 99.8% turdus migratorius 1,220,210 0.5% 99.5% anas platyrhynchos 1,202,864 0.4% 99.0% 24,066,470 9.0% gaiji et al. content assessment of the primary biodiversity data 137 table 12.b: top 15 species with highest number of data records (february 2012). occurrences (difference with 2010) (% of total) georeferenced (%) zenaida macroura 2,270,891 (+4,9%) 0.7% 99.9% sturnus vulgaris 2,171,136 (+17,2%) 0.7% 99.7% passer domesticus 2,029,427 (+7.2%) 0.6% 99.0% cardinalis cardinalis 1,779,316 (-8.9%) 0.5% 99.9% picoides pubescens 1,776,269 (+6.2%) 0.5% 99.9% junco hyemalis 1,731,413 (-0.2%) 0.5% 99.3% cyanocitta cristata 1,695,019 (-0.1%) 0.5% 99.8% carduelis tristis 1,666,477 (-0.3%) 0.5% 99.8% poecile atricapillus 1,612,214 (+6.2%) 0.5% 99.9% carpodacus mexicanus 1,609,953 (-0.1%) 0.5% 99.8% gaiji et al. content assessment of the primary biodiversity data 138 table 13: breakdown of the species and occurrences richness for kingdom animalia and plantae (february 2012) species occurrences number of occurrences/species animalia plantae animalia plantae <100 400.088 (87.5%) 374.524 (86.5%) 4.608.908 (2.4%) 5.655.807 (6%) 100-1.000 46.264 (10.1%) 49.643 (11.5%) 14.226.733 (7.4%) 14.455.647 (15.3%) 1.000-10.000 9.178 (2%) 7.428 (1%) 25.228.569 (13.1%) 18.437.768 (19.5%) 10.000-100.000 1.483 (<0.1%) 1.126 (<1%) 41.730.120 (21.6%) 33.901.539 (35.9%) 100.000-1.000.000 306 (<0.1%) 129 (<0.1%) 76.851.930 (39.8%) 22.025.055 (23.3%) >1.000.000 20 (<0.01%) 30.641.798 (15.9%) total 457.340 (100%) 432.851 (100%) 193.288.058 (100%) 99.475.816 (100%) gaiji et al. content assessment of the primary biodiversity data 139 table 14. breakdown of occurrences by continents % of gbif index region (december 2010) (february 2012) northern america 32.26% 28.19% northern europe 30.24% 30.06% western europe 8.34% 11.48% southern africa 4.12% 3.75% central america 3.61% 4.43% south america 2.72% 2.63% southern europe 2.35% 2.21% australia and new zealand 2.09% 7.03% eastern asia 1.72% 2.00% eastern africa 1.00% 0.66% eastern europe 0.87% 0.90% south-eastern asia 0.70% 0.64% caribbean 0.47% 0.44% melanesia 0.47% 0.42% antartica 0.37% 0.30% western africa 0.31% 0.25% western asia 0.31% 0.30% middle africa 0.30% 0.28% southern asia 0.30% 0.33% northern africa 0.14% 0.15% micronesia 0.06% 0.06% polynesia 0.05% 0.05% central asia 0.04% 0.04% unknown 7.16% 1.08% total 100% 100% gaiji et al. content assessment of the primary biodiversity data 140 table 15. proportion of financial contribution to gbif (2010) by continents region % of contribution (2011) oecd countries % of gbif index (february 2012) northern america 23.0% 2 28.19% northern europe 17.8% 8 30.06% western europe 24.1% 7 11.48% southern africa 1.0% 3.75% central america 1.4% 1 4.43% south america 0.5% 1 2.63% southern europe 6.3% 5 2.21% australia and new zealand 3.6% 2 7.03% eastern asia 22.1% 2 2.00% eastern africa <0.1% 0.66% eastern europe <0.1% 3 0.90% south-eastern asia <0.1% 0.64% caribbean <0.1% 0.44% melanesia <0.1% 0.42% antartica <0.1% 0.30% western africa <0.1% 0.25% western asia <0.1% 2 0.30% middle africa <0.1% 0.28% southern asia <0.1% 0.33% northern africa <0.1% 0.15% micronesia <0.1% 0.06% polynesia <0.1% 0.05% central asia <0.1% 0.04% unknown 1.08% total 100% 100% gaiji et al. content assessment of the primary biodiversity data 141 table 16. distribution of occurrences per members to the organisation for economic co-operation and development (oecd). occurrences % oecd countries 261,377,957 84.45% non-oecd countries 48,126,509 15.55% (based on occurrences where geospatial information is provided) table 17. distribution of occurrences per country income status. country income status occurrences % high 263,073,917 85.00% upper middle 36,197,975 11.70% lower middle 7,292,802 2.36% low 2,939,563 0.94% table 18.a grid (1/10 degree) occupancy (february 2012) presence in grid all with valid date % in at least 1 grid 603,532 (60.6%) 485,105 (48.7%) in at least 10 grids 150,771 (15.1%) 127,408 (17.0%) in at least 20 grids 95,783 (9.6%) 81,057 (10.8%) in at least 50 grids 38,520 (3.9%) 31,596 (3.2%) in at least 100 grids 18,072 (1.8%) 14,041 (4.2%) total number of species 995,975 (100%) 747,988 (100%) [75.1%] gaiji et al. content assessment of the primary biodiversity data 142 table 18.b grid (1/10 degree) occupancy for animalia. presence in grid all with valid date % in at least 1 grid 307,871 (67.3%) 213,498 (72.3%) in at least 10 grids 68,475 (15.0%) 55,561 (18.8%) in at least 20 grids 43,761 (9.6%) 36,462 (12.3%) in at least 50 grids 17,747 (3.9%) 15,071 (5.1%) in at least 100 grids 7,893 (1.7%) 6,756 (2.3%) total number of species 457,600 (100%) 295,380 (100%) [64.6%] table 18.c grid (1/10 degree) occupancy for plantae. presence in grid all with valid date % in at least 1 grid 253,477 (58.5%) 233,720 (62.5%) in at least 10 grids 71,332 (16.5%) 61,771 (16.5%) in at least 20 grids 44,572 (10.3%) 37,730 (10.1%) in at least 50 grids 17,325 (4.0%) 13,454 (3.6%) in at least 100 grids 8,690 (2.0%) 6,100 (1.6%) total number of species 433,174 (100%) 373,885 (100%) [86%] table 19. grid (1/10 degree) occupancy by classes (animalia and plantae). class total number of species in gbif-index number of species with sufficient grid presence (>=20) % aves 12,065 5,452 45.2% holocephali 54 19 35.2% gaiji et al. content assessment of the primary biodiversity data 143 marchantiopsida 156 54 34.6% cephalaspidomorphi 47 14 29.8% elasmobranchii 1,275 286 22.4% sphagnopsida 293 64 21.8% jungermanniopsida 1,750 381 21.8% actinopterygii 26,417 5,047 19.1% pinopsida 1,064 199 18.7% phascolosomatidea 43 8 18.6% sipunculidea 95 16 16.8% nuda 7 1 14.3% bryopsidophyceae 431 52 12.1% asteroidea 1,664 199 12.0% ulvophyceae 716 69 9.6% liliopsida 65,196 6,237 9.6% thaliacea 75 7 9.3% bangiophyceae 108 9 8.3% holothuroidea 1,178 98 8.3% echinoidea 1,290 106 8.2% appendicularia 51 4 7.8% hydrozoa 2,704 210 7.8% amphibia 5,665 409 7.2% haplomitriopsida 14 1 7.1% magnoliopsida 324,251 22,677 7.0% filicopsida 21,260 1,477 6.9% anthocerotopsida 58 4 6.9% bivalvia 11,956 801 6.7% cephalopoda 2,871 174 6.1% insecta 223,933 13,146 5.9% malacostraca 23,167 1,203 5.2% gastropoda 33,760 1,701 5.0% myxini 64 3 4.7% gaiji et al. content assessment of the primary biodiversity data 144 crinoidea 552 25 4.5% maxillopoda 7,864 345 4.4% cubozoa 25 1 4.0% anthozoa 7,747 278 3.6% null 5,548 159 2.9% aplacophora 223 4 1.8% stenolaemata 859 14 1.6% demospongiae 4,589 63 1.4% neoophora 201 2 1.0% rhynchonellata 2,033 12 0.6% cycadopsida 352 2 0.6% hexactinellida 504 1 0.2% leiosporocerotopsida 1 sarcopterygii 36 remipedia 17 cephalocarida 9 somasteroidea 2 eucycliophora 2 lingulata 163 pleurastrophyceae 7 pedinophyceae 2 pararotatoria 2 cestoda 150 pauropoda 185 eoacanthocephala 25 phylactolaemata 32 myxosporea 72 eutardigrada 107 collembola 1 archoophora 5 gaiji et al. content assessment of the primary biodiversity data 145 table 20. grid (1/10 degree) occupancy by families top 10 (animalia and plantae). taxonomical ranks (kingdom –phylum – class – order) families with sufficient grid presence (>=20) for a number of species greater than 100 number of species % of total total number of species in the family animalia chordata aves anseriformes anatidae 133 (73.1%) 182 animalia chordata actinopterygii myctophiformes myctophidae 173 (68.4%) 253 animalia chordata actinopterygii perciformes carangidae 103 (66.9%) 154 animalia chordata aves ciconiiformes laridae 102 (66.2%) 154 animalia chordata aves passeriformes thraupidae 139 (48.9%) 284 animalia chordata aves passeriformes tyrannidae 212 (48.0%) 442 animalia chordata aves passeriformes emberizidae 159 (46.8%) 340 animalia chordata aves piciformes picidae 115 (46.4%) 248 animalia chordata aves ciconiiformes accipitridae 132 (43.7%) 302 animalia chordata aves passeriformes furnariidae 128 (40.3%) 318 gaiji et al. content assessment of the primary biodiversity data 146 figure 1. typical flow of data discovered and published through the gbif network. gaiji et al. content assessment of the primary biodiversity data 147 figure 2. data mining methodology employed during content assessment exercise carried out by the gbif secretariat. gaiji et al. content assessment of the primary biodiversity data 148 figure 3. data mining methodologies employed during content assessment exercise carried out by the university of navarra. gaiji et al. content assessment of the primary biodiversity data 149 figure 4. evolution of the percentage of geo-referenced records (1800-2010). gaiji et al. content assessment of the primary biodiversity data 150 figure 5. evolution of incomplete records (taxonomical*temporal*geospatial) in the gbif-index gaiji et al. content assessment of the primary biodiversity data 151 figure 6.a. data records by kingdom (dec 2010). figure 6.b. data records by kingdom (february 2012). gaiji et al. content assessment of the primary biodiversity data 152 figure 7.a. data records by phylum (december 2010). figure 7.a. data records by phylum (february 2012). gaiji et al. content assessment of the primary biodiversity data 153 figure 8.a. data records by class (december 2010). figure 8.b. data records by class (february 2012). gaiji et al. content assessment of the primary biodiversity data 154 figure 9. treemap of occurrences according to taxon group, down to class. cell surface proportional to number of occurrences. blue: invertebrates; purple: vertebrates; green: higher plants; yellow: algae and ferns; brown: fungi; red: unicellular organisms. gaiji et al. content assessment of the primary biodiversity data 155 figure 10. top 10 data resources publishing maximum number of data records (december 2010). gaiji et al. content assessment of the primary biodiversity data 156 figure 11. kingdom wise distribution of data records by year (feb 2012). gaiji et al. content assessment of the primary biodiversity data 157 figure 12.a. plantae kingdom wise distribution of occurrences by year. gaiji et al. content assessment of the primary biodiversity data 158 figure 12.b. animalia kingdom wise distribution of occurrences by year (february 2012). gaiji et al. content assessment of the primary biodiversity data 159 figure 13. aves wise distribution of occurrences by year (february 2012) (red: aves, blue: other classes, black: all classes) gaiji et al. content assessment of the primary biodiversity data 160 figure 14. basis of records distribution of occurrences by year (february 2012) (red: observation, blue: specimen, black: all types) gaiji et al. content assessment of the primary biodiversity data 161 figure 15. distribution of voting and associate country participants (february 2012) gaiji et al. content assessment of the primary biodiversity data 162 figure 16. distribution of occurrences by latitude (february 2012) figure 17. distribution of species richness by latitude (february 2012) gaiji et al. content assessment of the primary biodiversity data 163 figure 18. distribution of the average number of occurrences per species by latitude (february 2012) gaiji et al. content assessment of the primary biodiversity data 164 figure 19. georeferenced records (dots) published by the largest data providers in europe. colors represent distinct providers. (from otegui et al., 2009). gaiji et al. content assessment of the primary biodiversity data 165 figure 20. records available in the gbif-index by date figure 21. variation of the records available in the gbif-index by date (december 2010 – february gaiji et al. content assessment of the primary biodiversity data 166 2012) figure 22. distribution of distinct species collected/observed over time (february 2012) (red: plantae, green: animalia, black: both) figure 23. distribution of distinct resources contributing to the gbif-index over time (february 2012) gaiji et al. content assessment of the primary biodiversity data 167 (red: plantae, green: animalia) figure 24. grid occupancy (1/2 degree grid) of the gbif-index over time (february 2012) gaiji et al. content assessment of the primary biodiversity data 168 figure 25. grid occupancy (1/10 degree grid) of the gbif-index by kingdom (february 2012) gaiji et al. content assessment of the primary biodiversity data 169 figure 26. grid occupancy (1/10 degree grid) of the gbif-index by classes (february 2012) gaiji et al. content assessment of the primary biodiversity data 170 figure 27. distribution of families for ffu-enm – kingdom animalia (february 2012) gaiji et al. content assessment of the primary biodiversity data 171 figure 28. distribution of families for ffu-enm – kingdom plantae (february 2012) gaiji et al. content assessment of the primary biodiversity data 172 figure 29. selection of fit-for-use records in the gbif index for ecological niche modelling (enm). biodiversity informatics, 19, 2025, pp. 109-119 109 field+genomics workshop: an initiative to build nanopore sequencing capacity in field-based host-pathogen research marlon e. cobos1, carlos carrion bonilla2, santiago f. burneo3, joseph a. cook2, jonathan l. dunnum2, lexi e. frank4, mackenzie grover1, alexander d. hey1, peter a. larsen4, ben j. wiens1, m. alejandra camacho3*, jocelyn p. colella1* 1biodiversity institute & department of ecology and evolutionary biology, university of kansas, lawrence, kansas, united states. lawrence, kansas 66045, usa 2biology department & museum of southwestern biology, university of new mexico, albuquerque, new mexico 87131, usa 3sección de mastozoología, museo de zoología, pontificia universidad católica del ecuador, quito, pichincha, ecuador 4department of veterinary and biomedical sciences, university of minnesota, st. paul, mn 55108, usa abstract. wildlife disease surveillance has received considerable attention following recent emergence of highconsequence zoonotic pathogens in humans. increased portability and affordability of sequencing technologies over the last decade have made real-time sequencing of wild animals and their pathogens a reality. wildlife samples screened for pathogens, however, are rarely permanently archived in museum biorepositories, which limits potential for scientific validation and prevents extension by related disciplines (e.g., ecology, evolution, conservation). to better connect biodiversity and biomedical sciences, the museums and emerging pathogens in the americas (mepa) network developed the field+genomics workshop to build capacity for surveillance of wildlife and their pathogens in biodiverse countries. here, we share workshop resources, in english and spanish, to facilitate reproducibility and expansion of the workshop into the future. the workshop lasted 10 days, 6 days of fieldwork and 4 days of molecular lab and bioinformatic techniques. the field component emphasized the importance of holistic collecting—that is, permanently preserving many parts and symbionts from each sampled organism—as a critical step in wildlife and pathogen surveillance and to build foundational scientific infrastructure. the molecular component of the workshop used samples collected during the field portion to identify hosts and pathogens in real-time. for this component, we trained participants in methods of dna extraction, library preparation, and nanopore adaptive sampling (a software feature for real-time selective enrichment or depletion of target sequences). bioinformatic training consisted of a basic introduction to computational genomics, a worked example to analyze a small sequence dataset, and an exercise using data generated from samples collected during the workshop. in total, the workshop cost ~$37k (~$3k per participant), however, ~25% of those funds are invested in basic equipment and infrastructure that is reusable in future workshops (e.g., sequencer, computer, etc.). this workshop highlights the effort and expertise required to conduct voucher-backed surveillance of wildlife and their pathogens and the many benefits of uniting biodiversity and biomedical sciences to build local capacity. keywords: bioinformatics, capacity building, latin america, metagenomics, nanopore adaptive sampling, open science * co-corresponding authors: jocelyn p. colella, colella@ku.edu, and m. alejandra camacho, macamachom@puce.edu.ec. mailto:colella@ku.edu mailto:macamachom@puce.edu.ec marlon e. cobos et al. – field+genomics workshop 110 background wildlife disease surveillance has received considerable attention and funding following the recent emergence of high-consequence zoonotic pathogens in humans (e.g., ebola virus, sars-cov-2, h5n1) (watsa 2020; lawson et al. 2021; mazzamuto et al. 2022). host and pathogen biodiversity are greatest in the tropics, potentially leading to increased risk of spillover and a greater need for regular, specimen-backed wildlife surveillance (cook 2018). the microscopic nature of most pathogens (e.g., viruses, bacteria, fungi, protozoa) requires molecular diagnostic tools for detection and characterization (aiewsakun and simmonds 2018). access to these tools, however, is uneven, posing significant challenges to scientists in regions with higher zoonotic risk who also face difficulties in securing and affording such technologies (rodriguez-morales et al. 2021). the increased portability and affordability of genomic sequencing technologies over the last decade— notably, oxford nanopore technology’s (ont) minion sequencer—has made real-time sequencing of wildlife hosts and their pathogens possible (lu et al. 2016; frank et al. 2023; kipp et al. 2023; de muelenaere et al. 2024). pairing holistic sampling of hosts–that is, permanently preserving many parts from each sampled organism–with molecular screens for pathogens allows for the integration of ecological and genomic information to provide a more comprehensive one health perspective on the ecology and evolution of emerging pathogens (colella et al. 2023). in addition to health applications, real-time sequencing can further guide field-based sampling efforts through real-time identification of hybrids or detection of undescribed diversity. historically, wildlife samples screened for pathogens were rarely permanently archived in museum biorepositories (colella et al. 2020; thompson et al. 2021). that lack of sample preservation prevents validation of the original science and further hampers scientific extension by related disciplines (e.g., biodiversity, conservation). to bridge that gap between two otherwise complementary sciences—biodiversity and biomedicine—the museums and emerging pathogens in the americas network (mepa1), a virtual community of practice and a branch of project echo2 (extensions for community healthcare outcomes), aims to better connect biorepositories with biomedical initiatives across the americas and build capacity in surveillance of wildlife and their pathogens in biodiverse countries (colella et al. 2021). to that end, we developed a field+genomics workshop, designed to train the next generation of molecularly-enabled field 1 https://mepa-network.weebly.com/. 2 https://projectecho.unm.edu/. researchers and promote greater health equity through voucher-backed wildlife surveillance. the workshop values traditional specimen-based wildlife sampling as a means of building foundational in-country biodiversity infrastructure and pairs holistic sampling practices with cutting-edge molecular methods of species identification and pathogen detection. here, we summarize the 2024 mepa field+genomics workshop and provide links by which teaching resources can be accessed and adapted for future workshops. the 2024 workshop took place at the reserva otongachi field station, located close to union del toachi, on the border of the provinces santo domingo and pichincha, in ecuador and lasted 10 days, with 6 days of fieldwork and 4 days of molecular laboratory and bioinformatic techniques (fig. 1). twenty participants and instructors attended, representing 8 different countries across the americas. the field component emphasized the importance of holistic collecting as an important first step in wildlife and pathogen surveillance that builds foundational scientific infrastructure for one health applications. the genomic component of the workshop used samples collected during the first week of fieldwork to identify hosts and detect pathogens in real-time and train participants in methods of dna extraction, library preparation, nanopore adaptive sampling (nas), and basic bioinformatics. all course materials are openly available online in english and spanish to enable replication, adaptation, or extension of this workshop (see data availability). figure 1. schematic of the workshop workflow. field collection of small mammals occurred during the first 6 days and utilized a variety of collection methods, followed by 4 days of wet lab and bioinformatic techniques. https://mepa-network.weebly.com/ https://projectecho.unm.edu/. marlon e. cobos et al. – field+genomics workshop 111 the workshop field component on day 1, group introductions were followed by a tour of a local museum biorepository (museo de zoología de la pontificia universidad católica del ecuador [qcaz]) and discussion regarding the role of museums in biomedical and biodiversity research. a brief welcome lecture (introduction to mepa & holistic collecting) introduced participants to the mepa network and established expectations for the workshop. then, the group traveled to the field station. that evening, 500 small mammal traps (live [sherman box traps] and lethal [museum special, rat traps]; fig. 2) and 6 mist nets (length = 12 m (2), 6 m (4); height = 3 m, gauge = 38 mm) were deployed, with nets monitored regularly following guidelines by sikes et al. (2016) and then closed at the end of the night. day 2 began early by checking traplines. a morning lecture on specimen preparation introduced students to five assembly line-style “preparation stations”: (1) measurements and data collection, (2) ectoparasites and fungal swabs, (3) tissue collection, (4) endoparasite necropsy, and (5) voucher specimen preparation. captured small mammals were then processed holistically (fig. 3, galbreath et al. 2019), with tissues preserved in liquid nitrogen or dna/rna shield (zymo research, irvine, california, usa) within 15 minutes of euthanasia to maximize molecular quality. euthanasia was performed by trained personnel (instructors) and followed approved iacuc procedures per sikes et al. (2016) and the american veterinary medical association (2020). the day concluded with each participant preparing a voucher specimen (e.g., study skin and skeleton), followed by a wrap-up lecture on the value of building biorepository infrastructure. days 3 through 5 involved checking traps at least twice daily, morning and evening, and holistically preparing collected specimens. participants rotated among preparation stations each day, such that they were able to learn each skill. as time permitted, lectures focused on local species ecology, mammals and disease ecology, field parasitology, natural history collection databasing, and holistic specimen research applications. the time and order of lectures varied based on the number of animals collected each day to accommodate required processing time. during downtime, participants gave a 3–5 minute informal, oral presentation about their background, research interests, and future plans for the application of information learned in the workshop. the number of field days included in the workshop will depend on the scientific goals of the expedition and capture success. for training purposes, we recommend no less than 3 to 4 days of fieldwork but, minimally, specimen yields must be sufficient for strategic genomic sequencing during the second half of the workshop. all lecture slides are available in english3 and in spanish4. genomic component wet-lab.—day 6 began the genomic component of the workshop with dna extractions. our dna extraction procedure takes approximately 4 h (3 h active, 1 h wait time), followed by a concentration step of variable duration depending on extraction yields5. we started extractions in the morning to ensure sufficient time for concentrating in the afternoon. lectures occurred during breaks in the extraction procedure and included an introduction to genomic sequencing methods and applications of nanopore sequencing. on day 7, hands-on nanopore library preparation began at 08:00 h. a maximum of 24 samples can be multiplexed with the ont native barcoding kit (sqk3 https://doi.org/10.6084/m9.figshare.28736396.v2. 4 https://doi.org/10.6084/m9.figshare.28736417.v2. 5 https://doi.org/10.6084/m9.figshare.28744844.v1. figure 2. baiting and setting a sherman live trap in the field. workshop participant dr. camila acosta-lópez, assistant professor at universidad central del ecuador, sets a trap in the cloud forest transect. https://doi.org/10.6084/m9.figshare.28736396.v2 https://doi.org/10.6084/m9.figshare.28736417.v2 https://doi.org/10.6084/m9.figshare.28744844.v1 marlon e. cobos et al. – field+genomics workshop 112 al. 2023) for two reasons: (1) the mitochondrion is smaller (~16 kbp) and (2) there are many more copies of mitochondria per cell compared to the nuclear genome (naue et al. 2024). to adaptively sequence mammalian mitochondrial genomes, we budgeted 24 h of run time. during mitogenome sequencing, we iteratively checked output data volume and flow cell health (e.g., number of available pores), and performed test assemblies to estimate sequencing depth. we then performed a depletion experiment, using the genome assembly of a common vampire bat (desmodus rotundus, ncbi refseq assembly: gcf_022682495.1) to deplete host dna, thereby enriching the metagenomic community for downstream pathogen identification. we ran the depletion experiment until <10% of pores on the flow cell were producing sequence data (~32 h). bioinformatics.—day 8 opened with an introduction to bioinformatics lecture. while the sequencer ran, we used a “worked example” to practice a basic bioinformatic workflow for processing nas data. kipp et al. (2023) published a straightforward example of how nas can be used to selectively sequence dna-based bacterial pathogens in black-legged ticks (ixodes scapularis). for this exercise, we used kipp et al.’s depletion dataset, which was adaptively sequenced such that tick sequences were selectively removed (depleted), increasing the probability of nbd114.24) and sequenced simultaneously. the number of samples selected for sequencing will vary depending on the scientific goals of the expedition, trap success, desired sequencing depth, and diversity of species sampled, among other variables (e.g., frank et al. 2023, kipp et al. 2023). limited by the size of our heat block (n = 12) and number of pipettes available to the group (n = 8), we divided participants into two groups and processed 24 samples in two batches of 12. that structure allowed group two to watch each step of the procedure performed by group one, before trying it themselves. day 7 culminated in loading the flow cell (> r10.4.1) and starting sequencing, ideally around 18:00 h (fig. 4). duration of sequencing depends on the scientific goals, with longer run times yielding more data. adaptive sampling is a software-based enrichment or depletion method unique to ont that selectively sequences a targeted subset of genetic material from a genomic sample (martin et al. 2022). here, we performed two sequencing runs: (1) an enrichment run for host and cestode mitochondrial genomes to confirm host and cestode species identity, and (2) a depletion run to recover potential pathogens and microbial symbionts. adaptive sampling of mitochondrial genomes, which are commonly used for species identification, requires less time than adaptive sampling of a larger part of the nuclear genome (wanner et al. 2021; kipp et figure 3. demonstration of a holistic specimen preparation assembly line that included five stations: (1) species identification and measurement, (2) ectoparasite examination and fungal swab collection, (3) tissue necropsy, (4) endoparasite examination, and (5) voucher specimen preparation. on day 2, workshop participants observed the process. on day 3, participants started rotating through each station. marlon e. cobos et al. – field+genomics workshop 113 detecting pathogen sequence information. we then guided participants in how to use bwa (li 2013) and minimap2 (li 2018) to map reads to a reference sequence, samtools (danecek et al. 2021) to filter out residual host sequences, and ncbi’s basic local alignment search tool (blast6) to perform basic homology searches against known pathogen sequences. on day 9, we terminated the mitogenome sequencing run and initiated the pathogen depletion run. while the sequencer ran, participants and instructors explored the sequencing output of the mammal mitogenome run. 6 https://blast.ncbi.nlm.nih.gov/blast.cgi. we aligned raw reads to a local bat (sturnira bakeri, qcaz-m18241) reference mitogenome (genbank: on357724.1) with bwa mem (-x ont2d), and called consensus in samtools. consensus sequences were imported into mega v. 4.0 (tamura et al. 2007) for alignment and visualization as a neighbor-joining phylogeny. while for publication purposes we recommend using alternative alignment and phylogenetic estimation software, this workflow is simple, even for those without bioinformatic experience, does not require access to high-powered computers, and provides an initial assessment of evolutionary relationships that can be used to direct field sampling in real-time or guide future analyses. figure 4. demonstration of nanopore adaptive sampling, molecular lab, and bioinformatic protocols. alexander hey, university of kansas graduate student, shows workshop participants details of nanopore sequencing. table 1. sequencing results for two nanopore adaptive sampling (nas) experiments conducted during the 2024 museums and emerging pathogens in the americas (mepa) field+genomics workshop. the pathogen depletion run, rejected reads that matched the host (e.g., common vampire bat, desmodus rotundus; ncbi refseq assembly: gcf_022682495.1). the mitogenome enrichment run sequenced reads that matched mammalian and cestode mitogenomes with at least 70% sequence identity. nas summary statistics pathogen “depletion” run mitogenome “enrichment” run total data produced (pass + fail) 58.4 gb 57.57 gb estimated bases 4.85 gb 4.83 gb reads generated 1.5 m 1.7 m estimated n50 5.69 kb 5.34 kb passed bases called 4.5 gb 4.51 gb failed bases called (q < 9) 418.46 mb 372.54 mb run duration 18 hrs, 24 mins 18 hrs, 24 mins https://blast.ncbi.nlm.nih.gov/blast.cgi marlon e. cobos et al. – field+genomics workshop 114 overall, the pathogen (depletion) and host mitogenome (enrichment) experiments generated 58.40 gb and 57.57 gb of sequence data, respectively. table 1 provides summary metrics output by the sequencer at the end of each run. after filtering for reads that mapped to the mitochondria, phylogenies were generated from 5.20 gb of data. at the end of the workshop, participants were asked to complete an optional evaluation survey to gauge learning outcomes and identify opportunities for improvement. survey results and ideas on how to overcome challenges identified during the workshop are detailed in the “lessons learned” section below. how to replicate this workshop putting this workshop together and making the materials available online required significant funding and collaborative participation. we have compiled a package of materials for researchers and instructors aiming to extend this workshop or create similar training opportunities. these materials include a public website7 with detailed directions for software installation, downloading data for the worked example, and step-by-step bioinformatics. additional materials are available on figshare8, and include: (1) workshop resources (field waiver, code of conduct, pre-workshop emails, equipment list and budget, workshop application and evaluation forms9), (2) eight lectures in english10 and spanish11, (3) template datasheets used by the university of kansas biodiversity institute’s division of mammals (specimen datasheets, gazetteer, ancillary datasheets, trapline datasheets, and bat net datasheets12), and (4) protocols used during the workshop (dna extraction, flow cell wash, library preparation, and concentration and dilution procedures13). the materials are openly available and free to use with appropriate acknowledgement and attribution. logistical considerations a field site where electricity and a classroom are available nearby is ideal. room and board accommodations for participants and instructors can also make this workshop more accessible and inclusive. accommodations can be made individually or for the group by the lead organization. it is possible, however, to lead this workshop 7 https://alexanderhey.github.io/mepa_field-genomics/. 8 figshare.com/projects/mepa_field_genomics_workshop/243968. 9 doi: 10.6084/m9.figshare.28744208.v1. 10 doi: 10.6084/m9.figshare.28736396.v2. 11 doi: 10.6084/m9.figshare.28736417.v2. 12 doi: 10.6084/m9.figshare.28744412.v1. 13 doi: 10.6084/m9.figshare.28744844.v1. in more remote settings, by using a generator to support a sequencing and informatics-capable computer and camping outdoors. the list of resources made available in the next section facilitates the replication of this workshop, regardless of local amenities, but a generator would be required if electricity is not available. the area (size) of the field site should be adjusted based on the planned sampling methods, target taxa, and number of participants. the lead organization must obtain the appropriate national, international, and institutional permissions and approvals to legally and ethically work with wild animals. that may include but is not limited to: an approved institutional animal care and use committee (iacuc) protocol, in-country collecting permits (state, federal), and, if specimens are to be moved across international borders, a material transfer agreement (mta) outlining mutually agreed upon terms (mat) and prior informed consent (pic) negotiated under the nagoya protocol for access and benefit-sharing. if the lead organization intends to sample wildlife in a foreign country, we strongly encourage them to partner with an in-country biorepository (i.e., a natural history museum or similar entity) early in the planning process to facilitate in-country permitting, logistics, and equitable specimen sharing (ramírez-castañeda et al. 2022). permit documentation differs for each country, government, and institution, and can take more than 6 months to process. workshop participants should be selected at least 3 months in advance, and ideally 6 months ahead of the workshop. depending on the country of origin, participants may require a visa to enter some countries. a sponsored invitation letter may speed up that process and should be considered when possible. safety considerations safety is critical to the success of this workshop. although fieldwork has inherent risks, organizers can take steps in advance to set their team up for success (campbell et al. 2025; ramírez-castañeda et al. 2022). field safety includes a priori consideration of potential hazards, including environmental and interpersonal risks, as well as identifying the actions, resources, and responses to be undertaken in case of an adverse event (kuebbing et al. 2021; rudzki et al. 2022). biosafety is another dimension of field safety that not only protects participants but also helps prevent the unintentional introduction of new pathogens into fragile ecosystems (islam et al. 2025). biosafety requirements differ across institutions, countries, and for different taxa (e.g., sikes et al. 2016; shapiro et al. 2024). we recommend advanced development of a field safety document tailored to the context of each workshop. each https://alexanderhey.github.io/mepa_field-genomics/ http://figshare.com/projects/mepa_field_genomics_workshop/243968 marlon e. cobos et al. – field+genomics workshop 115 participant must read and sign the field safety document to confirm understanding of the risks associated with fieldwork, acknowledge available resources in the event of an emergency, and limit liability of the lead organization. additional waivers or health documentation (e.g., proof of vaccination) may be required depending on the location of the workshop, lead or hosting institutions, and taxa to be sampled. as an example, proof of vaccination against yellow fever virus may be required in yellow fever endemic areas and rabies vaccines are generally required prior to handling carnivores or bats (aguilar-setién et al. 2022). the lead organization is responsible for collecting and maintaining digital and print copies of emergency contact information for instructors and participants in advance of the workshop. at a minimum, emergency contact information should include a contact name, phone number with country code, email, and mailing address, as well as pertinent medical information (e.g., allergies, medications, conditions, etc.). additional emergency phone numbers and maps to health services (e.g., hospital, urgent care center, etc.) nearest the field site should be included in the emergency response packet compiled and maintained by the lead organization. a code of conduct should be developed, shared with, and agreed upon by all participants in advance that outlines standards for professional behavior, steps to prevent harassment, and appropriate contacts and links for reporting. workshop resources and materials a comprehensive list of physical equipment and materials required for this workshop is included as supplementary material and online via figshare (project #243968). required field gear will vary depending on the geographic location, taxa being sampled, and the number of participants. in general, data sheets, tags, tubes, and specimen preparation supplies were provided by the museum biorepository that has agreed to accession and preserve the resulting specimens in perpetuity. in-country museum biorepositories may also have traps and specimen preparation gear that can support the field component of the workshop. participants were expected to bring their own laptop computers for the bioinformatic portion of the workshop. prior to the start of the workshop, participants were instructed on how to independently download and install required software and data files to their local machine to allow the workshop to operate off-line, as needed. detailed instructions for software installation on linux, macos, and windows operating systems are included on the workshop website and were emailed to participants two weeks in advance. we recommend that instructors also carry a backup digital copy of all required software and data files. to ensure that participants were able to perform the bioinformatic workflows during the workshop, the “worked example” exercise was designed to be reproducible on computers with mediumto high-computing capabilities (e.g., ram => 8 gb, i5 or similar processor, no dedicated graphics card required). given the nature and volume of nas data, however, higher computing power may be needed to work with data generated during the workshop. in such cases, participants can work in small groups and take turns watching and completing steps on more capable computers. participants were responsible for bringing their own field gear. a list of required gear was provided in advance and included snake gaiters, tall rubber boots, long pants, long-sleeve shirts, leather gloves, headlamp, rain gear, a hat, and personal toiletries, including sunscreen, bug spray, and personal medications. the list of required gear will vary significantly depending on the season, duration, weather, lodging, and amenities at the field site. budget major workshop expenses included field gear, curatorial supplies, molecular equipment and reagents, airfare, transportation, and room and board for participants and instructors. in 2024, purchasing all workshop equipment and reagents from scratch required approximately $4500 for field gear and $13,500 for genomic lab equipment and reagents, totaling $18,000. while molecular equipment and reagents were the most expensive category, a large proportion of the required equipment and infrastructure only needs to be purchased once, and some could be borrowed from in-country collaborators. for example: a minion sequencer ($1000), pipettes (8 x $355 each = $2840), 4 tb hard drive ($280), portalyzer ($300, peck et al. 2022), and m3 macbook computer ($4000) with sufficient gpu to perform adaptive sampling, are one-time purchases, whereas, mammal traps may be provided by in-country collaborators. in 2024, round-trip airfare per person (home country to/from ecuador) was approximately $1000 (12 international plane tickets); room and board at the reserva otongachi field table 2. high-level cost breakdown for the 2024 museums and emerging pathogens in the americas (mepa) field+genomics workshop (20 people). see supplementary materials for a detailed breakdown of expenses. item cost field equipment $4500 molecular equipment and reagents $13,500 airfare (~$1000 per participant) $12,000 lodging ($29/person/day) $5800 board ($6/person/day) $1200 transportation ($20/person) $400 total $37,400 marlon e. cobos et al. – field+genomics workshop 116 station were $35 per person per day ($7000 for a group of 20 people); and group transportation by bus was $20 per person ($400). in total, individual attendance for an international participant could be sponsored for about $1360, however, total costs will vary depending on local conditions and accommodations. lessons learned need for partnerships that build in-country infrastructure partnering with an in-country museum biorepository is essential to the success of this type of workshop and directly supports the growth and long-term maintenance of in-country research infrastructure, as a form of benefit sharing. by permanently archiving materials in a local biorepository, the samples collected during the workshop can be made available to the local scientific community to address additional questions (colella et al. 2020). further, involvement of local participants, institutions, and community stakeholders in advance can help meaningfully shape the research focus and experiments to be conducted during the workshop (hetu et al. 2019, vilaça et al. 2024). for example, although this workshop focused on dna-based pathogens and host mitochondria, the methods could be extended to screen for rna-based pathogens. including local personnel in the workshop further fosters local network development and scientific capacity, ensuring the long-term transfer of knowledge and skills (martin et al. 2022). by strengthening both infrastructure and workforce, this type of workshop can create a more equitable and sustainable framework for future research and empower local communities to address their own biodiversity and health challenges (ramírez-castañeda et al. 2022). such partnerships also facilitate in-country logistics, international communication, and inform research priorities. workshop participants identified local access to equipment, resources, and infrastructure as the major barrier to scaling such workflows and technologies in lower-resourced countries. for example, field workflows require access to traps (mist nets, traps, etc.; >$2500), sampling media (250 ml dna/rna shield; $250), and consumables (data pages, tags, cryotubes, etc., ca $500), as well as space and storage infrastructure to maintain samples long-term. similarly, adaptive sequencing requires access to both a minion sequencer ($1000), flow cells ($700/ ea.), laboratory consumables (e.g., pipette tips, tubes, etc., $1500), and a gpu-enabled computer (>$3000). additionally, sequencing data and intermediary analysis files consume significant digital storage. not all countries and institutions have access to liquid nitrogen for cryogenic tissue preservation in the field. dry ice or shelf-stable buffers, like dna/rna shield or rna later, may be used as alternatives depending on in-country regulations. during the workshop, we performed adaptive sampling and live base calling/demultiplexing simultaneously on a gpu-enabled apple macbook pro m3 (2023) in minknow v. 23.11.2. the latest software update (minknow v. 24.02), however, disabled simultaneous base calling and adaptive sampling for macos due to strain on the apple silicon processor (pers. comm. ont). as a result, the sequence output generated during this workshop may not perfectly reflect results obtained in future workshop iterations. future iterations should leverage recent software updates and perform adaptive sampling in real-time, followed by later base calling and demultiplexing if using minknow v. 24.02 on macos. teaching new teachers this workshop was designed to teach field and molecular techniques to the next generation of scientists in latin america, however, an ideal extension of this effort would be to teach workshop instructors, such that in-country instructors are able to replicate the workshop. as with any skill, regular application and use are key to retention and adoption. therefore, we recommend recruiting participants with a high probability of using these tools and skills in the near future. plan for post-workshop communication establish a plan for group communication after the completion of the workshop. choose a platform that is accessible, informal, and multilingual, and, ideally, one that allows sharing of images and files. for the 2024 workshop, we established a whatsapp group to continue discussion and problem-solving after the workshop. for collaborative writing to be successful, we identified co-leads and co-senior authors, representing pairs of north-south colleagues, to lead discrete products stemming from the data generated during the workshop. for manuscript development, we used google docs as it is freely accessible to all participants and allows multi-user editing, text change tracking, and commenting. co-leads were tasked with data analysis, visualization, and writing of respective projects, and encouraged to reach out to other participants and instructors for help, as necessary. conclusion workshops that integrate complementary field and molecular skill development build new capacity in biodiverse countries to address critical societal questions related to emerging pathogens, biodiversity loss, climate marlon e. cobos et al. – field+genomics workshop 117 change, and food security, among others. by (1) developing and enhancing the skills, knowledge, and abilities of local scientists, (2) building international, multi-disciplinary collaborative networks, and (3) growing in-country biodiversity infrastructure by contributing samples to local biorepositories, such workshops can help stimulate solutions to the public health and biodiversity challenges of each region. by linking field-based collection and permanent specimen archival to real-time genomic surveillance of hosts and pathogens, workshops like this can improve detection of zoonotic pathogens while also aligning with global one health initiatives. this workshop aimed to demonstrate the level of effort, expense, and expertise involved in conducting voucher-backed surveillance of wildlife and their pathogens, coupling fieldwork for specimen collection with downstream molecular applications. such training workshops are one mode of benefit sharing (colella et al. 2023) under the nagoya protocol for access and benefit sharing and will serve local communities beyond the duration of a single workshop, grant, or project. such collaborations further contribute to in-country infrastructure in the form of natural history museum biorepositories (colella et al. 2020), while exposing non-museum-based researchers to the benefits of long-term data and sample archival. together, these outcomes highlight how collaborations that unite fieldwork, genomics, and museum science can advance both local capacity and global health preparedness. acknowledgments we thank the course participants for their enthusiasm, patience, and work ethic. we especially thank the sección de mastozoología del museo zoología (qcaz; quito-católica-zoología) at pontificia universidad católica del ecuador (puce) for providing in-country logistical support. this workshop was funded by the u.s. national science foundation (nsf) pursuit award to jocelyn p. colella (#2302677) and the yates field fund (university of new mexico). we thank the international community of practice, museums and emerging pathogens in the americas, for organizational support. mec was partially supported by picante (pathogen informatics center for analysis, networking, translation, and education award to joseph a. cook; nsf #2155222). data availability all supplementary materials are available on figshare14. declaration of conflicts of interest the authors declare that no competing interests exist. 14 https://figshare.com/projects/mepa_field_genomics_ workshop/243968. references aguilar-setién, a., aréchiga-ceballos, n., balsamo, g. a., behrman, a. j., frank, h. k., fujimoto, g. r., gilman duane, e., hudson, t. w., jones, s. m., ochoa carrera, l. a., powell, g. l., smith, c. a., triantis van sickle, j., and vleck, s. e. (2022). biosafety practices when working with bats: a guide to field research considerations. applied biosafety: journal of the american biological safety association, 27(3), 169–190. aiewsakun, p., and simmonds, p. (2018). the genomic underpinnings of eukaryotic virus taxonomy: creating a sequence-based framework for family-level virus classification. microbiome, 6(1), 38. american veterinary medical association (2020). avma guidelines for the euthanasia of animals: 2020 edition15. campbell, l. g., montgomery, n. r., king, m. r., hall, j., and struwe, l. (2025). students in the wild: safety instruction practices in distance-taught biological laboratory and field classes. american biology teacher 87(1): 13–19. colella, j. p., stephens, r. b., campbell, m. l., kholi, b. a., parsons, d. j., and mclean, b. s. (2020). the open-specimen movement. bioscience, 71(4), 405–414. colella, j. p., bates, j., burneo, s. f., camacho, m. a., carrion bonilla, c. a., constable, i., d’elía, g., dunnum, j. l., greiman, s., hoberg, e. p., lessa, e. p., liphardt, s. w., londoño-gaviria, m., losos, e., lutz, h., ordóñez garza, n., peterson, a. t., martin, m. l., ribas, c. c., struminger, b., torres-pérez,f., thompson, c.w., weksler, m., cook, j. a. (2021). leveraging natural history collections as a global, decentralized pathogen surveillance network. plos pathogens, 17(6), e1009583. colella, j. p., cobos, m. e., salinas, i., cook, j. a., and picante consortium. (2023). advancing the central role of non-model biorepositories in predictive modeling of emerging pathogens. plos pathogens, 19(6), e1011410. cook, j. (2018). primary infrastructure for mammalogy in the anthropocene. mastozoologia neotropical, 25(2), 267–268. danecek, p., bonfield, j. k., liddle, j., marshall, j., ohan, v., pollard, m. o., whitwham, a., keane, t., mccarthy, s. a., davies, r. m., and li, h. (2021). twelve years of samtools and bcftools. gigascience, 10(2), giab008. 15 https://www.avma.org/resources-tools/avma-policies/avmaguidelines-euthanasia-animals. https://figshare.com/projects/mepa_field_genomics_workshop/243968 https://figshare.com/projects/mepa_field_genomics_workshop/243968 https://www.avma.org/resources-tools/avma-policies/avma-guidelines-euthanasia-animals https://www.avma.org/resources-tools/avma-policies/avma-guidelines-euthanasia-animals marlon e. cobos et al. – field+genomics workshop 118 de meulenaere, k., cuypers, w. l., gauglitz, j. m., guetens, p., rosanas-urgell, a., laukens, k., and cuypers, b. (2024). selective whole-genome sequencing of plasmodium parasites directly from blood samples by nanopore adaptive sampling. mbio, 15(1), e0196723. frank, l. e., lindsey, l. l., kipp, e. j., faulk, c., stone, s., roerick, t. m., moore, s. a., wolf, t. m., and larsen, p. a. (2023). rapid molecular species identification of mammalian scat samples using nanopore adaptive sampling. journal of mammalogy, 105(5), 965-975. galbreath, k. e., hoberg, e. p., cook, j. a., armién b., bell k. c., campbell m. l., dunnum j. l., dursahinhan a. t., eckerlin r. p., gardner s. l., greiman s. e., henttonen h., jiménez f. a., koehler a. v. a., nyamsuren b., tkach v. v., torres-pérez f., tsvetkova a., and hope a. g. 2019. building an integrated infrastructure for exploring biodiversity: field collections and archives of mammals and parasites. journal of mammalogy 100(2): 382–93. hetu m., koutouki k., and joly y. 2019. genomics for all: international open science genomics projects and capacity building in the developing world. frontiers in genomics 10:95. islam, s., kangoyé, m., diallo, a. h., katani, r., and escobar, 2025. a field guide for sampling bats (chiroptera) for eco-epidemiological studies. frontiers in veterinary science 12:1605150. kipp, e. j., lindsey, l. l., khoo, b., faulk, c., oliver, j. d., and larsen, p. a. (2023). metagenomic surveillance for bacterial tick-borne pathogens using nanopore adaptive sampling. scientific reports, 13(1), 10991. kuebbing, s., clark, d., gharaibeh, b., janecka, m., kohl, k., kramp, r., ohmer, m., olmsted, c., martinez, k. p., rudzki, e., turcotte, m. m., and richards-zawacki, c. (2021). field safety manual.16 lawson, b., neimanis, a., lavazza, a., lópez-olvera, j., tavernier, p., billinis, c., duff, j. p., mladenov, d. t., rijks, j., savić, s., wibbelt, g., ryser-degiorgis, m., and kuiken, t. (2021). how to start up a national wildlife health surveillance programme. animals: an open access journal from mdpi, 11. li, h. (2013). aligning sequence reads, clone sequences and assembly contigs with bwa-mem. arxiv. doi: 10.48550/arxiv.1303.3997 li, h. (2018). minimap2: pairwise alignment for nucleotide sequences. bioinformatics (oxford, england), 34(18), 3094–3100. 16 https://www.ple.pitt.edu/sites/default/files/documents/pitt_ biological_sciences_field_safety_manual_-_6-10-2025.pdf. lu, h., giordano, f., and ning, z. (2016). oxford nanopore minion sequencing and genome assembly. genomics, proteomics and bioinformatics, 14(5), 265–279. mazzamuto, m. v., schilling, a.-k., and romeo, c. 2022. wildlife disease monitoring: methods and perspectives. animals: an open access journal from mdpi, 12. martin, a. r., stroud, r. e., abebe, t., akena, d., alemayehu, m., atwoli, l., chapman, s. b., flowers, k., gelaye, b., gichuru, s., kariuki, s. m., kinyanjui, s., korte, k. j., koen, n., koenen, k. c., newton, c. r. j. c., olivares, a. m., pollock, s., post, k., singh, i., stein, d. j., teferra, s., zingela, z., and chibnik, l. b. 2022. increasing diversity in genomics requires investment in equitable partnerships and capacity building. nature genetics 54:740–745. martin, s., heavens, d., lan, y., horsfield, s., clark, m. d., and leggett, r. m. (2022). nanopore adaptive sampling: a tool for enrichment of low abundance species in metagenomic samples. genome biology 23: 11. naue j., xavier c., hörer s., parson w., and lutz-bonengel s. 2024 assessment of mitochondrial dna copy number variation relative to nuclear dna quantity between different tissues. mitochondrion 74:101823. peck, c., jackobs, f., and smith, e. (2022). the portalyzer, a diy tool that allows environmental dna extraction in the field. hardwarex, 12(e00373), e00373. ramírez-castañeda, v., westeen, e. p., frederick, j., amini, s., wait, d. r., achmadi, a. s., andayani, n., arida, e., arifin, u., bernal, m. a., bonaccorso, e., bonachita sanguila, m., brown, r. m., che, j., condori, f. p., hartiningtias, d., hiller, a. e., iskandar, d. t., jiménez, r. a., khelifa, r., márquez, r., martínez-fonseca, j. g., parra, j. l., peñalba, j. v., pinto-garcía, l., razafindratsima, o. h., ron, s. r., souza, s., supriatna, j., bowie, r. c. k., cicero, c., mcguire, j. a., tarvin, r. d. (2022). a set of principles and practical suggestions for equitable fieldwork in biology. proceedings of the national academy of sciences of the united states of america, 119(34), e2122667119. rodriguez-morales, a. j., paniz-mondolfi, a. e., faccinimartínez, á. a., henao-martínez, a. f., ruiz-saenz, j., martinez-gutierrez, m., alvarado-arnez, l. e., gomez-marin, j. e., bueno-marí, r., carrero, y., villamil-gomez, w. e., bonilla-aldana, d. k., haque, u., ramirez, j. d., navarro, j.-c., lloveras, s., arteaga-livias, k., casalone, c., maguiña, j. l., escobedo, a. a., hidalgo, m., bandeira, a. c., mattar, https://www.ple.pitt.edu/sites/default/files/documents/pitt_biological_sciences_field_safety_manual_-_6-10-2025.pdf https://www.ple.pitt.edu/sites/default/files/documents/pitt_biological_sciences_field_safety_manual_-_6-10-2025.pdf marlon e. cobos et al. – field+genomics workshop 119 s., cardona-ospina, j. a., suárez, j. a. (2021). the constant threat of zoonotic and vector-borne emerging tropical diseases: living on the edge. frontiers in tropical diseases, 2, 676905. rudzki, e. n., kuebbing, s. e., clark, d. r., gharaibeh, b., janecka, m. j., kramp, r., kohl, k. d., mastalski, t., ohmer, m. e. b., turcotte, m. m., and richards-zawacki, c. l. (2022). a guide for developing a field research safety manual that explicitly considers risks for marginalized identities in the sciences. methods in ecology and evolution, 13(11), 2318–2330. shapiro, j. t., phelps, k., racey, p., vicente, a., viquez-r, l., walsh, a., weinberg, m., and kingston, t. (2024). iucn ssc bat specialist group guidelines for field hygiene. zenodo.17 sikes, r. s., and the animal care and use committee of the american society of mammalogists. (2016). 2016 guidelines of the american society of mammalogists for the use of wild mammals in research and education. journal of mammalogy, 97(3), 663–688. tamura, k., dudley, j., nei, m., and kumar, s. (2007). mega4: molecular evolutionary genetics analysis (mega) software version 4.0. molecular biology and evolution, 24(8), 1596–1599. 17 http://doi.org/10.5281/zenodo.12169384. thompson, c., phelps, k., allard, m., cook, j. a., dunnum, j. l., ferguson, a., gelang, m., khan, f. a. a., paul, d., reeder, d., simmons, n., vanhove, m., webala, p., weksler, m., and kilpatrick, c. w. (2021). preserve a voucher specimen! the critical need for integrating natural history collections in infectious disease studies. mbio, 12(1), e02698-20. vilaça, s. t., vidal, a. f., pavan, a. c. d., silva, b. m., carvalho, c. s., povill, c., luna-lucena, d., nunes, g. l., figueiró, h. v., mendes, i. s., bittencourt, j. a. p., côrtes, l. g., canesin, l. e. c., oliveira, r. r. m., damasceno, r. p., vasconcelos, s., barreto, s. b., tavares, v., oliveira, g., martins, a. b., and aleixo, a. 2024 leveraging genomes to support conservation and bioeconomy policies in a megadiverse country. cell genomics 4:100678. wanner, n., larsen, p. a., mclain, a., and faulk, c. (2021). the mitochondrial genome and epigenome of the golden lion tamarin from fecal dna using nanopore adaptive sequencing. bmc genomics, 22(1), 726. watsa, m., and wildlife disease surveillance group. (2020). rigorous wildlife disease surveillance. science, 369, 145–147. http://doi.org/10.5281/zenodo.12169384 biodiversity informatics, 19, 2025, pp. 120-143 120 letsrept: an r package to access the global reptile database and facilitate taxonomic harmonization joão paulo dos santos vieira-alencar1,*, h. christoph liedtke2, shai meiri3, 4, uri roll5, peter uetz6, javier nori7,8 1 centro de ciências naturais e humanas, universidade federal do abc, são bernardo do campo, sp, brazil 2 department of ecology and evolution. biological station of doñana, csic, calle américo vespucio 26, 41092 seville, spain 3 school of biosciences, university of melbourne, parkville, victoria, australia 4 school of zoology, tel aviv university, tel aviv, israel 5 mitrani department of desert ecology, the jacob blaustein institutes for desert research, ben-gurion university of the negev, midreshet ben-gurion, israel 6 center for biological data science, school of life sciences / reptile database, virginia commonwealth university, richmond, va, united states of america 7 geobio, instituto de diversidad y ecología animal (idea, conicet), córdoba, argentina 8 facultad de ciencias exactas físicas y naturales, universidad nacional de córdoba, córdoba, argentina abstract. taxonomy is a highly dynamic science upon which most biodiversity studies rely. constant revisions of species delimitation hypotheses, using ever-growing amounts of data and tools cause species numbers and identities to continuously and rapidly change. reptiles are the most species rich terrestrial vertebrate group and are amongst the most threatened and least known vertebrate taxa, representing nearly half of all datadeficient terrestrial vertebrate species. every year hundreds of new species are described and dozens are revised, resulting in synonymizations, splittings, generic reassignments, or elevation from synonymy or from subspecies into species status. the nomenclature of this group is therefore highly dynamic and consequently, to integrate available reptile datasets generally requires extensive nomenclature review, especially for broad scale analyses. letsrept is a new r package that integrates the reptile database – the best curated and reliable global taxonomic reference for reptiles – into the r programming environment. its main functions allow users to retrieve the most up-to-date taxonomic information in real time, to compare lists of species names to current nomenclature, and to detect names that have been changed by either lumping or splitting, all through web scraping techniques. additional functions allow to produce quick taxonomic summaries, access species accounts, retrieve full reference lists and more. by permitting to embed the reptile database directly into r workflows, the letsrept package improves the integration of datasets from different sources, with authoritative taxonomy, reducing data loss due to nomenclature mismatch and improving the consistency in biodiversity analyses. keywords: data acquisition, data management, taxonomy, the reptile database, web scraping * corresponding author: joaopaulo.valencar@gmail.com. mailto:joaopaulo.valencar@gmail.com joão paulo dos santos vieira-alencar et al. – letsrept 121 introduction cataloguing the diversity of life on earth is fundamental to the scientific endeavour, yet, we have only formally described a fraction of all living organisms (mora et al., 2011). in some taxa, such as nematodes, our knowledge shortfalls remain vast, with most species still undiscovered or undescribed (larsen et al., 2017). this is termed a ‘linnean shortfall’, which describes the discrepancy between the number of species that exist and the number of species that were formally described and are known to science to date (brown and lomolino, 1998; hortal et al., 2015). vertebrates are amongst the best-studied animals, and such knowledge deficit is probably relatively small in this group. yet, even within vertebrates there is active discussion on the delimitation and nomenclature of species (wüster et al., 2024). furthermore, in several vertebrate taxa, such as reptiles, species continue to be discovered and described at accelerating rates (meiri, 2016; uetz et al., 2021). these taxonomic efforts result in ever increasing numbers of recognized species, and frequent changes of perceived species identities and taxon nomenclature. this creates taxonomic mismatches across major large databases (e.g., genbank, gbif, and the iucn red list), which pose significant challenges for researchers, the public at large, and conservation practitioners (cordier et al., 2024; nori et al., 2022a). moreover, advances in our knowledge of species’ natural histories, evolutionary relationships, geographical distribution, and conservation status and needs, among others, are contingent on establishing current and unified taxonomic reference across different data sources (baranzelli et al., 2023). to maintain consistent and up-to-date taxonomic references, several databases are being actively curated, such as the catalogue of life (col; bánki et al., 2025) and the integrative taxonomic information system (itis1). additionally, quick nomenclature verification by querying a list of species names is available in several web servers, such as the global names verifier (gnv2) and the gbif name parser3. they are not, however, easily integrated into workflows in statistical programming environments, a problem which been addressed by, for example, the r packages taxize (chamberlain and szöcs, 2013; chamberlain et al., 2020), and taxadb (boettiger et al., 2023). these packages integrate taxonomic information from general databases, such as the previously mentioned col and itis, into the r environment. although the taxize and taxadb packages employ sophisticated matching algorithms, and draw from multiple sources, they are limited in their ability to address 1 https://www.itis.gov/. 2 http://resolver.globalnames.org/. 3 https://www.gbif.org/tools/name-parser. taxonomy ambiguity (see the taxize outputs in the example application section below). taxonomic ambiguity may arise from species synonymization, when accumulated evidence suggests that a given taxonomic entity no longer bears enough distinctive characters to be considered a separate species. conversely, new evidence may support the division of a previously unified taxonomic entity into multiple newly recognized species, leading to taxonomic splitting. complex taxonomic synonymization or splitting make database nomenclature matching difficult. for example, with regards to database management, cases of taxonomic synonymization would involve deciding on how to merge information that was previously regarded as multiple taxonomic entities. on the other hand, cases of taxonomic splitting would require detailed revision to ensure which portion of the information previously assigned to a single species should now be attributed to distinct taxonomic entities. moreover, these challenges will vary depending on the type of data meant to be merged or separated. distribution data from a junior synonym will in most cases just require merging the records to those of the senior synonym, conversely, cases of taxonomic splitting would require a careful revision, not only of the records related to the vouchers used in the new species description, but to some extent, most nearby records. meanwhile, databases on species traits might involve distinct decisions with respect to, for example, categorical or continuous data. in both cases such decisions are dependent on the taxonomic authority followed. given the dynamic nature of taxonomy, especially for historically overlooked groups such as amphibians and reptiles, taxon-specific databases, that are regularly updated with high standards of data curation, such as “amphibian species of the world” (frost, 2025), and “the reptile database” (uetz et al., 2025) are indispensable to ensure nomenclature consistency. liedtke (2018) created amphinom, an r package dedicated to retrieving and synchronizing taxonomic information, and species synonyms, from the amphibian species of the world dataset (frost, 2025). conversely, reptiles still lack any comparable tool to facilitate access to their primary taxonomic database. properly managing reptile taxonomic information is essential, not only because they represent the richest terrestrial vertebrate group with nearly 12,500 recognized species (uetz et al., 2025), but also because their conservation is constantly affected by taxonomic changes. reptiles are among the most threatened vertebrate groups, with 1846 species (~18%) currently classified as threatened (iucn, 2025) and high numbers of datadeficient, recently described, and still undescribed species https://www.itis.gov/ http://resolver.globalnames.org/ https://www.gbif.org/tools/name-parser joão paulo dos santos vieira-alencar et al. – letsrept 122 suspected to be threatened (meiri, 2016; caetano et al., 2022; meiri et al., 2023). many reptile lineages exhibit high taxonomic instability, with alpha diversity estimates continuously changing due to new species descriptions and reclassifications (e.g., uetz et al., 2020; nori et al., 2022b), with a clear impact on their management and conservation (cordier et al., 2021). the number of known reptiles is currently increasing, with one in six (17.0%) of all recognized species being described since 2014 and an nearly 100 additional species described between january and july of 2025 (fig. 1; uetz et al., 2025). the reptile database (rdb), founded nearly 30 years ago (uetz et al., 2021), has become the central resource for reptile taxonomy. it serves as a key reference for taxonomic information, on which other databases (such as col) are reliant. besides highly curated taxonomy, the rdb also provides species synonyms, and their respective historical use in the literature, type series information, countries of species occurrences, a reference list for each species, and more (uetz et al., 2025). with over 50,000 monthly users, and a rapidly growing number of citations, the rdb continues to increase in importance within the herpetological community (uetz et al., 2021). given the close ties between taxonomy and fields such as systematics and conservation – and considering the fast pace of taxonomic revisions and persistent knowledge gaps, particularly in species-rich regions (e.g., melville et al., 2021) – the ability to track changes and provide expert, up-to-date taxonomic decisions is essential. the rdb has established itself as the most reliable and comprehensive source of such high-quality information for reptiles worldwide. a growing number of databases now provide valuable information on reptile traits and distributions, such as squambase (meiri, 2024) and repttraits (oskyrko et al., 2024) for ecological traits, iucn red list (iucn, 2025) for threats and the global assessment of reptile distributions (roll et al., 2017; caetano et al., 2022) for species level distribution data. genetic information is also accessible in genbank (benson et al., 2015). together, all these provide valuable information used to access reptile diversity patterns and to track threats and the evolutionary history of reptile species. however, the dynamic nature of reptile taxonomy, and frequent nomenclature changes, hamper their straightforward use, leading to potential data loss due to nomenclature mismatch and to the time-consuming (but indispensable) step of careful nomenclature unification. therefore, the development of a dedicated tool to address the challenges related to reptile taxonomy can greatly facilitate access and comparison of all known valid species names and their respective synonyms, as well as use of the most reliable sources of reptile information, making it both timely and valuable. consequently, we developed ‘letsrept’, an r package to serve as a tool to fig 1. annual number of species descriptions for squamata. points represent species counts per year, and the smoothed lines show trends over time using loess regression. joão paulo dos santos vieira-alencar et al. – letsrept 123 letsrept is primarily designed to facilitate the comparison and matching between datasets where the species is the primary unit of classification, and the current valid names in the reptile database. the package queries a vector of species names and resolves nomenclatural mismatches, including identification of cases of taxonomic splitting or synonymization. to avoid oversimplification the package highlights cases of ambiguity and enables users to focus on and investigate complex taxonomic changes. it also extracts the full taxonomic information available to allow exploratory analyses (e.g., on higher taxonomic levels). letsrept is also equipped with built-in datasets that comprise the full reptile taxonomic information available in the reptile database, including a full synonym list, that can be accessed without an active internet connection. additionally, the package provides versions of squambase (meiri, 2024) and repttraits (oskyrko et al., 2024) comprising all information originally available within these datasets, with former and current suggested nomenclature (september 2025 version, uetz et al., 2025) along with the equivalent nomenclatural status as obtained from letsrept. we describe below a regional case study with detailed examples of main functions of letsrept (table 1), a flowchart summarizing step-by-step the nomenclature update is illustrated in figure 2. the script along with all necessary input data are available in the supporting information. integrate the reptile database with the r environment to provide a fast and efficient way to retrieve and synchronize reptile species taxonomic information. software tool description letsrept is a package written in the r programming language (r core team, 2025) specifically designed to facilitate access to most reptile information available in the reptile database. by extracting and structuring information directly from the reptile database website programmatically sending queries to the web server, it enables users to retrieve summarized taxonomic information from species lists obtained from advanced searches (e.g. a country or high taxonomic level species lists); explore synonyms and other content available in the species account and track nomenclatural changes (table 1). the package leverages existing http and web scraping tools: httr (wickham, 2023a), rvest (wickham, 2024), and xml2 (wickham et al., 2025). it uses dplyr (wickham et al., 2023b), stringr (wickham, 2023c), and tidyr (wickham et al., 2024) to manage and manipulate character strings and data frames. it uses parallel (r core team, 2025) to implement parallel processing and enhances user experience with progress bars incorporated from pbapply (solymos and zawadzki, 2023) and pbmcapply (kuang et al., 2022). table 1. functions available within the letsrept package. function description reptsearch queries the reptile database (rdb) for information about a single reptile species using its binomial name. reptadvancedsearch creates a search url for retrieving species lists from rdb based on multiple filters. reptspecies retrieves a list of reptile species from the reptile database (rdb) based on a search url. it optionally returns detailed taxonomic information and urls for each species for further use. allows parallel processing. reptstats summarizes higher taxonomic information from a list of species reptsynonyms retrieves a data frame containing the current valid names of reptile species along with all their recognized synonyms and chresonyms, as listed in the reptile database (rdb). optionally, it returns the references citing each entry. allows parallel processing. reptcompare compares a list species with the current nomenclature and highlights mismatched names that require nomenclature review. reptsync queries a user-provided list of reptile species binomials known to require review against the current nomenclature as defined in the reptile database and returns a list of valid names, highlighting the status of the queried nomenclature in comparison to the current. allows parallel processing. reptsplitcheck queries a user-provided list of reptile species binomials and check them as synonyms of species described after a user-defined date. allows parallel processing. repttidysyn prints the outputs of herpsync and herpsplitcheck in a tidy way, with an optional filter to the status column. reptrefs retrieves the list of references from a species account, optionally including their respective access links (when available) and summarize in a data frame. joão paulo dos santos vieira-alencar et al. – letsrept 124 fig 2. summary of steps performed to update the atlas of brazilian snakes (nogueira et al., 2019) nomenclature. joão paulo dos santos vieira-alencar et al. – letsrept 125 note that future uses of these scripts may produce different outcomes due to future taxonomic changes. a guide to install the package and reproduce all the examples along with all data used is accessible4. detailed documentation and vignettes are also available through standard r help systems. example application we used the atlas of brazilian snakes (nogueira et al., 2019) as a case study to illustrate the utility of letsrept to reconcile outdated nomenclature with current taxonomic standards within an extensive and taxonomically challenging dataset. nogueira et al. (2019) mapped the distribution of 411 snake species that occur in brazil including georeferenced information for all available type localities (see supplementary material, table s3 therein). however, brazilian snake nomenclature has changed since 2019, with synonymizations, taxonomic splitting, and genus and species revalidations. to use this important resource, it now requires a detailed nomenclatural review. to sample a subset of species from the reptile database we retrieved snake species known to occur in the brazilian territory using the function reptadvancedsearch. this function replicates the filtering logic of the reptile database’s web interface that allow users to construct advanced queries directly from r: snakes_br_link < reptadvancedsearch(location = “brazil”, higher = “snakes”) by specifying the location and taxonomic group of interest, we generated a search url and obtained a count of 450 recognized brazilian snake species. next, we used reptspecies to retrieve detailed taxonomic information for each species returned by the search. this function accepts the search url and returns a vector of species names or, optionally, a data frame including species higher taxonomic information and the urls to access species account: snakes_br 1.96. here, peterson et al. seem to be under the misapprehension that it is the normality of the distribution of values of ε itself that are being discussed. the distribution of a statistic and the probability distribution of the data from which the statistic is derived, however, are not the same thing. the next line of criticism is that “the entire probabilistic argument behind the ε index very doubtful” as “the numbers in that formula, in general, cannot be regarded as probabilities, but as proportions of observations, for a given species, or proportions of observed co-occurrences, for pairs of species, in a particular database.” of course, the numbers in that formula refer to proportions that is how a probability is defined in the frequentist perspective of observations in a particular database, but not, contrary to peterson et al.’s comments, to the “true” probability representing the “true” species distribution in their terms. obviously, the probability that is calculated from a database represents an under sampling relative to the true distribution. all peterson et al.’s worked examples are based on the incorrect assumption that our work states that statistically significant co-occurrence is a sufficient condition for a biotic interaction. as the “labels” associated with the possibly interacting species are important to understand the nature of any interaction, it is a rather fruitless enterprise to actively seek examples where the associated labels do not naturally lead towards a hypothesis of a likely biotic interaction, as in the case of peterson et al.’s example using the families trogonidae (aves) and scarabeidae (insecta), where they claim that there can be no biotic component to the interaction as scarabs are terrestrial and trogons are frugivorous. however, contrary to their statements, not all scarabs are terrestrial (vulinec et al., 2007) and trogons are not exclusively frugivorous (remsen et al.,1993). thus, some possible biotic interactions are: some trogon species consume some scarab species, or some fruit consuming scarabs (reyes and morón, 2005) consume the same food source as some trogons. this does not mean, christopher r. stephens et al. – a reply to peterson et al. (2020) 59 of course, that co-occurrence is due exclusively to these potential biotic interactions. the second example considers the explanation of the significant co-occurrence of two rodents dipodomys merriami and perognathus longimembris. peterson et al. wish to infer, wrongly, that a large positive value of ε signifies within our methodology “mutualism, symbiosis” while, in contrast, there is evidence of competition between these two species which comes from controlled experiments where one set of species were excluded from a given area and the relative numbers of any remnant species were recorded. the observation that the removal of some species led to increases in the abundance of some others was interpreted as evidence of competition, although this might not be the case, as shown by davidson et al. (1984). this, micro-level competition however, will not necessarily lead to manifestations at a macro level, where, on the contrary, one might expect to see one species being a niche variable for the other if, for example, they share food resources. one can investigate this phenomenon by creating niche and community models using both climate and rodent species as covariates and noting that the interaction is not explainable by both species being adapted to the same climate. it is then an interesting exercise to construct the niche models and networks for both species by including in potential food sources. one finds for example, that larrea tridenta (creosotebush), prosopis glandulosa (honey mesquite), gutierrezia californica and gutierrezia ramulosa all have significant ε values with both species. thus, one can verify that one reason, among various, for a positive interaction may be due to shared food resources. the final worked example of peterson et al. concerns calculating ε values for six cat species occurring in mexico but from two distinct data sources. the criticism there is that ε is “database dependent,” as the results they derive from the snib and the iucn extent-of occurrence datasets (iucn, 2016) are different.” we completely agree, and so it should be. iucn data are not primary data: it is the result of a subjective model based on expert opinion, so to call them two different data bases as if they were two different representations of the same thing, both using primary data, is extremely misleading. in their discussion of our work in epidemiology and public health, peterson et al. state that our “key assertion is that geographic co-occurrence implies biotic interactions such as vectoring and hosting pathogens.” no! our key assertion, again, is that geographic co-occurrence is a necessary but not sufficient condition for such biotic interactions to be present. thus, it is not a “rash” conclusion to state that a leishmania vector must co-occur with one or more leishmania hosts and vice versa and, contrary to their claim, we have never asserted that all lutzomyia species are vectors, or that all are of the same competence. in the case of flavivirus, peterson et al. state: “despite known zikv infections in numerous species of bats in africa and asia, they have not been found to be competent reservoirs, contra the predictions of gonzález-salazar et al. (2017).” this is a false statement. in that paper, we state clearly: “finally, we emphasise again that the scope of the model of the present paper is to serve as a focus for future studies and show that potentially useful information can be gleaned from the method, which, at this level, is not capable of predicting detailed elements such as potential host competency.” in particular, our methodology is clearly predictive of contact between pathogen and potential mammalian hosts. the question is more, what is the result of that contact? given the enormous importance of emerging and re-emerging diseases, anything that can focus attention on potential risk factors is extremely valuable. in closing, we note that network reconstruction based on co-occurrence of living entities or their activities in space and time is becoming a major research direction in the life sciences, from biochemistry, to genomics, to ecology. the reason for this is because it is a fundamental topic, which will provide important clues on phenomena such as co-existence and ecosystem functioning. several techniques have been developed, such as maxent methods, boolean networks, and bayesian approaches. as most research directions and theories in ecology, they proceed by continuous refinement and approximation of complex natural systems. we do not claim that ours is the best approach; we limit ourselves to present it and show that it makes sensible predictions that can be tested, and we acknowledge that, as with most approaches, it can be refined. isn’t that what science is about? … increasing our knowledge of nature by developing theories and models and testing the hypothesis derived from their application in order to refine them? christopher r. stephens et al. – a reply to peterson et al. (2020) 60 references atauchi, p. j., peterson, a. t. and j. flanagan. 2018. species distribution models for peruvian plantcutter improve with consideration of biotic interactions. journal of avian biology, 49:e01617. bell, g. 2005. the co‐distribution of species in relation to the neutral theory of community ecology. ecology 86:1757-1770 berzunza-cruz m, rodríguez-moreno a, gutiérrez-granados g, gonzález-salazar c, stephens c.r., hidalgo-mihart m., et al. 2015. leishmania (l.) mexicana infected bats in mexico: novel potential reservoirs. plos neglected tropical diseases 9:1-15. davidson, d. w., r. s. inouye, and j. h. brown. 1984. granivory in a desert ecosystem: experimental evidence for indirect facilitation of ants by rodents. ecology 65,:1780-86. gotelli, n.j. 2000. null model analysis of species co-occurrence patterns. ecology 81:2606-2621 remsen jr, j.v., hyde, m.a., and a. chapman. 1993. the diets of neotropical trogons, motmots, barbets and toucans. condor 95: 178-192. peterson, a. t. soberón, j., ramsey, j. r., and l. osorio-olvera. 2020. co-occurrence networks do not support identification of biotic interactions. biodiversity informatics 17:1-10. reyes, n.e. and m.a. morón. 2005. fauna de coleoptera melolonthidae y passalidae de tzucacab y conkal, yucatán, méxico. acta zooligica mexicana 21:15-49. vulinec, k., mellow, d. j. and c.r.v da fonseca. 2007. arboreal foraging height in a common neotropical dung beetle, canthon subhyalinus harold (coleoptera: scarabaeidae). coleopterists’ bulletin 61:75-81. biodiversity informatics, 18, 2024, pp. 43-55 43 species distribution model accuracy is strongly influenced by the choice of calibration area sergio luna1, alexander peña-peniche2,3 & roberto mendoza-alfaro1* 1universidad autónoma de nuevo león, facultad de ciencias biológicas, laboratorio de ecofisiología, nuevo león, méxico 2universidad autónoma de nuevo león, facultad de ciencias biológicas, laboratorio de biología de la conservación y desarrollo sustentable, nuevo león, méxico 3centro de investigación científica de yucatán, a. c., unidad de recursos naturales, mérida, yucatán, méxico *corresponding author: roberto mendoza-alfaro, email: roberto.mendoza@yahoo.com abstract. species distribution models (sdm) are widely used tools in ecology and conservation aimed at predicting the potential distribution of a species based on its environmental requirements and occurrence data. sdm face many challenges and uncertainties that influence their accuracy. selecting the ideal calibration area is one of these difficulties. this study analyzes the influence of the extent of the calibration area on the accuracy of sdm through simulations with virtual species. using bioclimatic variables, 100 virtual species were gener-ated. occurrence probabilities were determined based on environmental suitability, spatial sampling bias, and accessible areas. sdm were built using maxent, varying size of calibration area, spatial filtering of occurrence records, predictor collinearity treatment, and regularization parameter. model performance was assessed in terms of functional accuracy (true model accuracy) and discrimination accuracy (model ability to separate oc-currence from random sites). results show that the extent of the calibration area was the most influential factor (explaining 50% of the variance in functional accuracy), while regularization multiplier, predictor collinearity, and spatial thinning had minimal impact (about 4% of explained variance combined). overall, larger calibration areas generally led to higher functional accuracy, although it varies across species. the correlation between functional and discrimination accuracy was relatively low, indicating that models performing well in one metric may not excel in the other. in conclusion, this research advances the discussion on calibration area selection, providing insights on its substantial effects on model accuracy. our findings demonstrate that the size of the calibration area is one of the most critical factors affecting the accuracy of models, surpassing the influence of other factors. these insights highlight the importance of select appropriate calibration areas to improve model predictions and ensure more reliable applications of the models. key words: species distribution models, ecological niche, calibration area, model accuracy, virtual species, simulations, maxent introduction species distribution models (sdm) are widely used tools in ecology and conservation aimed at predicting the potential distribution of a species based on its environmental requirements and occurrence data (peterson and soberón 2012). sdm provide valuable information for assessing biodiversity patterns (weaver et al. 2006), identifying conservation priorities (srinivasulu et al. 2021), evaluating climate change impacts (brodie et al. 2022), and forecasting species invasions (duque-lazo et al. 2016). however, sdm also face many challenges and uncertainties that may impair their reliability (see sillero et al. (2021) and sillero and barbosa (2021) for details). one significant challenge is the occurrence thinning process, which aims to reduce artificial clustering in species occurrence records. while thinning is necessary to mitigate biases introduced by uneven sampling efforts, it can also reduce the number of data points available for modeling, potentialluna et al. – species distribution model accuracy influenced by calibration area 44 ly affecting the model’s performance and accuracy (veloz, 2009; boria et al. 2014; varela et al. 2014). another critical challenge that significantly impacts model performance is managing multicollinearity among predictor variables. high levels of correlation between variables can lead to overfitting, where the model becomes excessively tailored to the training data, reducing its ability to generalize to new data. additionally, multicollinearity can result in unreliable response curves, where the effect of one variable is confounded by its correlation with others, making it difficult to discern the true influence of each predictor (dormann et al. 2013; de marco and nóbrega, 2018; feng et al. 2019). furthermore, the challenge of selecting the ideal calibration area adds another layer of complexity, as it directly influences the geographic region where the model is calibrated using observed data (machado-stredel et al. 2021). this is particularly important for some presence-background algorithms, such as ecological niche factor analysis (enfa; hirzel et al. 2002), genetic algorithm for rule-set production (garp; stockwell and noble, 1992), and maximum entropy (maxent; phillips et al. 2006; merow et al. 2013), which are influenced by the selection of the calibration area. it is important to emphasize that while the selection of the calibration area is a critical challenge for some algorithms, not all methods face this problem (sillero et al. 2021). for instance, mahalanobis distance (clark et al. 1993) and bioclim (booth et al. 2014) do not require a calibration area due to their focus on occurrence data. presence-background methods evaluate the environmental conditions across the study area (background) and compare them with the conditions where the species is found, as indicated by its occurrence records (phillips et al. 2006; merow et al. 2013). the choice of calibration area can significantly influence model accuracy, as different areas may have different environmental conditions and sampling biases (acevedo et al. 2012; owens et al. 2013). the ideal calibration area is where the species is in equilibrium with its environment (guisan and zimmermann, 2000). this means that the species has adapted to the current environmental conditions and as a result its distribution reflects these conditions (araújo and pearson 2005; barve et al. 2011). under the biotic-abiotic-mobility framework (bam), the m fulfills this assumption (soberon and peterson 2005; barve et al. 2011); however, for most species, its estimation is complicated (acevedo et al. 2012). despite its significance, there has yet to be a consensus on how to select an appropriate calibration area, and different criteria and methods have been used in the literature (machado-stredel et al. 2021; rotllan-puig and traveset 2021). some approaches include selecting political regions, polygons, or rectangles around occurrence records (shcheglovitova and anderson 2013), selecting bioclimatic or hydrogeographical regions occupied by the species (espindola et al. 2019; sillero et al. 2021), or generating buffers around occurrence records (zhu et al. 2014). in many studies, the calibration area used is not explicitly stated (e.g., masin et al. 2014). these different methods can result in regions of varying sizes, which can significantly impact subsequent model characteristics. according to several studies, either too constrained or overly expansive calibration areas may compromise the accuracy of model predictions (vanderwal et al. 2009, acevedo et al. 2012). other studies have found that smaller calibration areas may yield superior model accuracy, as they mitigate the risk of overfitting to conditions near occupied localities or exclude regions with suitable conditions that remain unoccupied due to dispersal limitations and biotic interactions (anderson and raza 2010). furthermore, selecting a background from larger areas leads to changes in variable importance, resulting in models becoming increasingly simplified and dominated primarily by just a few variables (vanderwal et al. 2009). considering this context, there is a strong need for a more systematic evaluation of the effect of calibration area on sdm performance. throughout this study, we refer to “model performance” in terms of functional accuracy (the model’s ability to infer the true occurrence probability) and discrimination accuracy (the model’s ability to correctly differentiate between presences and background in geographic space) (warren et al. 2020). in this research, we address this through a simulation approach. virtual species simulation allows us to create artificial species with known occurrence probabilities and environmental responses, allowing the accurate isolation of the effects of targeted factors (meynard et al. 2019). our objectives were 1) to compare models calibrated over areas of various sizes in terms of functional and discrimination accuracy (warren et al. 2020); 2) to test whether knowing the accessible area to a species leads to improved models; and 3) to compare the relative contribution of calibration area size to model luna et al. – species distribution model accuracy influenced by calibration area 45 performance against other factors such as occurrence thinning and collinearity. by using virtual species with known environmental responses and probabilities of occurrence, this study provides insights into how several factors influence sdm accuracy and highlights the importance of careful calibration area selection. methods virtual species simulations we simulated 100 virtual species using the virtualspecies r package (leroy et al. 2016) and the 19 bioclimatic variables from worldclim 2.0 (fick and hijmans 2017). these variables encompass critical climatic factors influencing species distributions and are widely utilized in sdm studies. we used a resolution of 2.5 arc-minutes to simulate the species’ niches, as this resolution provides a balance between spatial detail and computational efficiency. the region considered was restricted to the american continent due to its broad range of habitats and climatic conditions. we chose to include variables that combine precipitation and temperature, despite a recent trend of excluding them from sdm studies (booth 2022). these variables have been demonstrated to significantly influence species distributions. for instance, precipitation during the wettest and driest months are key factors shaping global plant distributions (huang et al. 2021), while the mean temperature of the driest quarter has been identified as the most influential variable in explaining continental fish distributions (castillo-torres et al. 2017). the ecological niche concept is based on the principle that each species has an optimal set of environmental conditions where it can survive, grow, and reproduce (peterson and soberon, 2012). to define each species’ environmental preferences, two to five of the 19 bioclimatic variables were randomly selected and then subjected to a principal component analysis (pca) using singular value decomposition, and the two principal components derived were used to define the species niche. different species are influenced by different environmental factors (huang et al. 2021), and the primary drivers of their distributions can differ significantly. randomly selecting 2 to 5 variables helps simulate this variability. for the first two principal components, a gaussian response curve was defined, with randomly selected values for the mean (representing the most suitable values for the species) and the standard deviation (indicating the niche breadth, or physiological tolerance of the species). the initial suitability function, which describes the species–environment relationship and can have any range of values, was converted to an occurrence probability with a logistic function. after this transformation, the probability of occurrence is bounded between 0 and 1 (i.e., the likelihood that the species is present at a site given the specific set of environmental conditions on that site). specifically, we used a logistic transformation with β (inflexion point) fixed at 0.5 and α (steepness of the slope) randomly selected. this method serves as an equivalent to applying thresholds but offers a more ecologically realistic representation by smoothing the transition between suitable and unsuitable conditions and simulating stochastic processes acting on species occurrences (leroy et al. 2016; meynard et al. 2019). the magnitude of the spatial sampling bias was simulated using the human influence index dataset (wcs and ciesin, 2005), representing the likelihood of selection per cell. these values were divided by the maximum value, resulting in a sampling bias range between 0 and 1. an essential aspect, according to niche theory, is the distinction between the fundamental, potential, and realized niches. the fundamental niche represents the full range of environmental conditions under which a species can survive and reproduce, while the potential niche corresponds to the subset of those conditions that actually exists at a given time (jackson and overpeck, 2000). the potential niche includes areas that are environmentally suitable but may not be occupied by the species due to additional factors such as dispersal constraints and/or the influence of biological interactions (barve et al. 2011; soberón and nakamura, 2009). the realized niche is a subset of the potential niche after considering these factors (jiménez and soberón 2022; soberón and arroyo-peña 2017). dispersal capability, in particular, can limit a species’ ability to explore favorable regions. to account for this, we defined the accessible area of each species by programming a cellular automaton (ca) to find potential areas where the species would be at environmental equilibrium. to achieve this, each species was assigned a random dispersal capability within the range of 50 to 250 km. this dispersal capability is fixed for each species and represents the maximum distance it can disperse in each iteration. using this information, the accessible luna et al. – species distribution model accuracy influenced by calibration area 46 area was simulated by selecting cells through a bernoulli trial based on their occurrence probability. the point with the highest occurrence probability across the landscape served as the initial selection point, and a buffer was generated around it with a radius equivalent to the assigned dispersal capability. subsequent cells were selected within this buffer through additional bernoulli trials, generating a new buffer. this process was repeated until no additional cells could be added (environmental equilibrium). the resulting buffer was considered the accessible area for the species, representing the sites it has potentially explored. the ca function is based on the presence of cells with sufficiently high probabilities of occurrence. in regions with high occurrence probabilities, the buffer will continue to expand, allowing the species to explore more adjacent areas. conversely, in regions where occurrence probabilities are lower, fewer cells are added, and the buffer’s growth tends to diminish. this process is analogous to real ecological dynamics, where species are more likely to colonize and establish in regions that provide optimal conditions for survival and reproduction. for each species, 100 occurrence records were obtained by randomly selecting cells (with replacement) with a bernoulli trial with a probability of success equal to: p(x) = s(x)b(x)r(x) where p(x) is the sampling probability of the cell, s(x) is the occurrence probability of the species in that cell, b(x) is the relative strength of spatial sampling bias, and r(x) is a binary variable with values 1 inside the accessible area of the species and 0 outside (warren et al. 2020). this process was repeated five times for each species to account for stochasticity in the sampling process. sdm the simulation of each virtual species aimed to represent realistic aspects affecting real species occurrences. when testing the models, our goal was to reflect real-world modeling scenarios, acknowledging that the exact parameters influencing species distributions are often unknown (meynard et al. 2019). for example, including spatial bias in the simulation of each virtual species occurrence was intentional to mimic real-world uneven sampling conditions due to factors like accessibility and observer effort. the thinning process applied during modeling (see below) is one of the existing methods intended to reduce this bias (aiello‐lammens et al. 2015). similarly, variable selection is a critical step in species distribution modeling (sdm), yet the true explanatory variables driving species distributions are rarely known (inman et al. 2021). by excluding variables based on collinearity rather than relying on using the true influencing factors, we aimed to reflect the inherent uncertainty that exists in real-world variable selection scenarios. while simulating all aspects that influence species distributions is conceptually and computationally challenging, we aimed to include some essential aspects to provide a more realistic evaluation framework. sdm were built for each of the occurrence datasets using maxent v3.4.4 (phillips et al. 2006) in the dismo r package (hijmans et al. 2017). the different levels of the four factors set out below involved the evaluation of 2000 models per species (400 models x 5 occurrence replicates). the remaining configurations were set to their default settings. calibration area.—given that maxent works with a maximum entropy principle, it is necessary to provide random data from the environment for its characterization (merow et al. 2013). as the extent of the calibration area directly influences the range of environmental conditions available for model training, it is of great importance to carefully consider the geographic space in model calibration. given that the accessible area for a species is usually unknown, the effect of the calibration area extent was evaluated by randomly sampling 10,000 background cells (default value) within a radius of 25, 50, 100, 200, 300, 500, 700, and 1000 km around the occurrence records. all available cells were used for calibration areas with less than 10,000 background points. spatial filtering of occurrence records.—spatial filtering of occurrence records is one of the main methods used to reduce the sampling bias in datasets derived from biological collections (taylor et al. 2020). occurrence records were filtered using the spthin package (aiello‐lammens et al. 2015) with the following filtering distances: 0 (no filtering), 5, 10, 15, and 20 km. predictor collinearity.—there is still a lack of consensus regarding how predictor collinearity should be treated in sdm, given that using highly correlated variables may influence model performance (feng et al. 2019). we compared the effect luna et al. – species distribution model accuracy influenced by calibration area 47 of dealing with predictor collinearity by calibrating models using all available variables or selecting variables with pearson correlation coefficients below 0.7 using the vifcor function of the usdm package (naimi et al. 2014). regularization parameter (rm).—maxent uses lasso regularization to constrain the modeled distributions to lie within a specific interval around the empirical mean instead of matching it exactly. this overfitting can be reduced by specifying a rm value that penalizes the use of additional parameters (phillips et al. 2006; warren and seifert 2011). the effect of this factor was evaluated by using rm values of 0.5, 1 (default), 2, 3, and 5. in addition to assessing the relative contributions of the previously mentioned factors, we examined the accuracy of a model for each species constructed under a scenario where the ecological characteristics of the species are well understood (hereafter termed unbiased model). the considerations for building this model included utilizing 1000 occurrence data points (instead of 100) sampled without spatial sampling bias, using the accessible area of the species as the calibration area, employing only the environmental variables that precisely define a species’ niche as predictors, and evaluating the regularization parameter with the same 5 previously defined values. model performance model performance was evaluated using both functional and discrimination accuracy (warren et al. 2020). functional accuracy (true model accuracy) was calculated as the spearman rank correlation between the true occurrence probability and the occurrence probability inferred from maxent (with the complementary log-log (cloglog) transformation) across the accessible area. for large areas, 25,000 cells were selected at random due to computational constraints. preliminary trials showed that using 25,000 randomly selected cells produced results very close to those obtained using the entire area, with a pearson correlation of 0.997 between values from the whole area and those from the 25,000 randomly selected cells. discrimination accuracy was calculated using cross-validation, where the data were divided into four groups according to two criteria: randomly and by geographic blocks, using the enmeval r package (kass et al. 2021). the boyce index was used as the evaluation metric (hirzel et al. 2006), calculated for each of the four groups. the average boyce index across these groups was used as the overall measure of discrimination accuracy. the boyce index was selected as an additional metric to assess whether it is possible to identify the best models based on discrimination accuracy, particularly when considering variations in the datasets that might not be perfectly aligned with the “known” truth. while we have a ‘true’ model, the use of the boyce index allows us to evaluate how models perform in a comparative context, offering an additional perspective on model performance in scenarios that simulate real-world conditions. data analysis we chose not to rely on p-values to evaluate the significance of our findings due to well-documented criticisms of their use (hurlbert et al. 2019) and their inappropriateness in the context of simulation studies (white et al. 2014). instead, we adopted an approach that focuses on the relative contributions and relationships among the examined factors through linear models, utilizing the lmg method implemented in the relaimpo r package. this method is based on sequential r2 and addresses the dependency of regressor orderings by averaging over these orderings using simple unweighted averages (grömping 2006). the response variable was the functional accuracy of sdm models, and the explanatory variables were the extension of the calibration area, the spatial filtering of occurrence records, the predictor collinearity, and the regularization parameter β treated as categorical variables. this approach allowed us to analyze the variability and contributions of different factors influencing model performance more robustly, rather than relying on traditional hypothesis testing with p-values. additionally, we evaluated the correlation between functional and discrimination accuracy using pearson correlation to assess the capacity of discrimination metrics to select the best-performing models using withheld data. these analyses were performed individually for each species, given that we do not expect all species to be affected by the evaluated factors in the same way; some may experience more pronounced sampling bias, others may have smaller accessible areas, etc. results overall, the unbiased models exhibited high accuracy within the species accessible area, with 80% of these models showing functional accuracy values luna et al. – species distribution model accuracy influenced by calibration area 48 exceeding 0.9. the median functional accuracy of these models was high, with a spearman correlation of 0.968, and the range of accuracy scores varied from 0.560 to 0.999 (fig. 1). in contrast, the rest of the models demonstrated more variable performance within and across species. the maximum functional accuracy across species ranged from 0.178 to 0.996, with a median of 0.902. the minimum functional accuracy across species ranged from -0.959 to 0.699, with a median of -0.256. this considerable variability in accuracy was also reflected in the range of functional accuracy values within species (i.e. the difference between the maximum and minimum value for each species), showing values from 0.291 to 1.87, with a median of 1.045. overall, these findings highlight the considerable range in model accuracy between species. among the analyzed species, 30 consistently exhibited models with positive functional accuracy, characterized by spearman correlation coefficients greater than 0. the other 70 species showed at least one model with negative functional accuracy. furthermore, 13 species were particularly notable, as most of their models yielded negative functional accuracy values. this disparity in model accuracy underscores the diverse responses of species to the calibration area and other factors, leading to variations in the accuracy of the generated models. explained variance the extent of the calibration area turned out to be the most important factor in terms of true model accuracy, with a substantial median of 50.46% explained variance (range: 2.49% to 92.99%). following this, the regularization parameter (rm) played a less prominent role, with a median explained variance of 3.65% (range: 0.01% to 48.05%). predictor collinearity and spatial thinning exhibited a negligible impact on true model accuracy, each contributing with a median explained variance of 0.41% and 0.04%, respectively (fig. 2). functional and discrimination accuracy we explored the correlation between functional and discrimination accuracy for the 100 species under two different data partitioning scenarios (fig. 3). when considering random data partitioning, a wide range of correlations was observed. the minimum and maximum values were -0.58 and 0.82. the median correlation between these two metrics was 0.46. remarkably, 19 species exhibited negative correlations. when data partitioning was based on figure 1. functional accuracy of the models per species (400 models × 5 replicates). the y-axis represents functional accuracy (measured as the spearman rank correlation between true and predicted suitability) and the x-axis represents different species (without specific names as they are virtual). data points are represented with five different colors (one for each replicate). species are ordered from highest to lowest based on their best-performing model. the accuracy of the unbiased model is denoted for each species with a "+" symbol. the red dashed line represents the expected value by chance. luna et al. – species distribution model accuracy influenced by calibration area 49 figure 2. percentage of variance explained by the four factors analyzed and the residual (unexplained) variance. the x-axis represents the percentage of variance explained, and the y-axis shows the factors analyzed. the boxplots summarize the distribution of these percentages for all species. area: calibration area; rm: regularization multiplier; cor: predictor collinearity; thin: occurrence thinning. figure 3. relationship between model functional and discrimination accuracy for the 100 species under varying calibration area extents and partitioning methods. the x-axis represents the functional accuracy, divided by the extent of the calibration area, the y-axis represents discrimination accuracy, divided by the partitioning method (random and geographical blocks). the blue lines, obtained through ordinary least squares (ols) regression, provide a visual representation of the correlation between the two accuracy metrics for each subset of points. luna et al. – species distribution model accuracy influenced by calibration area 50 geographical blocks, the correlations also displayed variability, ranging from -0.73 to 0.69, with a median correlation of 0.32. here, 25 species showed negative correlations between functional and discrimination accuracy. however, a noteworthy finding emerged despite the relatively low correlation between functional and discrimination accuracy. for 76 species, the best models based on the random data partitioning evaluation showed functional accuracies exceeding 0.5. in parallel, for 71 species, the best models selected with the geographical block data partitioning strategy achieved functional accuracies exceeding 0.5. discussion our research advances a more systematic evaluation of how varying extents of calibration area affect sdm accuracy. indeed, the observed variability in model performance within and across species underscores the need for tailored approaches, considering species-specific characteristics. in this context, the most critical factor influencing the accuracy of species distribution models turned out to be the size of the calibration area. overall, models calibrated with larger areas tend to show higher functional accuracy than models calibrated with smaller areas. although it is not straightforward to compare different algorithms, due to their reliance on different types of data (e.g., presence-only vs. presence-absence), statistical methodologies (e.g., classification vs. regression), or evaluation strategies (e.g., rocauc vs boyce index), results are often compared in the literature (bucklin et al. 2015; valavi et al. 2022) and our major findings are in line with other previous research. vanderwal et al. (2009), for instance, used buffers with increasing distances ranging from 10 to 500 km around the species’ occurrences to study the impact of various calibration areas working with rainforest vertebrate from the australian wet tropics (awt). they found a rapid increase in accuracy as the background size expanded from 10 to 100 km (roc-auc > 0.93), with subsequent expansions beyond this threshold showing only marginal improvements (roc-auc > 0.99). however, an important drawback, acknowledged by the authors, was the potential overestimation of model accuracy when assessed over a large geographical extent (lobo et al. 2008). vanderwal et al. (2009) recognized this phenomenon and used a fixed evaluation area to calculate auc values across all species. their findings showed that the “fixed-area” accuracy was maximum at a background size around 200 km, and it gradually decreased as points were generated from larger regions. it is important to note that the use of simulated data with known occurrence probabilities allows us to directly assess the accuracy of the models without relying solely on the discrimination metrics, thereby ensuring that our findings are not artifacts of these metrics. in a similar way, acevedo et al. (2012) conducted a study using data from four ungulate species in spain to evaluate the predictive accuracy of sdm calibrated over varying extents of calibration areas. their results showed that while calibration accuracy (miller’s statistic) declined with the expansion of the calibration region, discrimination accuracy (rocauc) increased. this approach allowed for the generation of purely environmental models that, when projected onto a new scenario, depicted the potential distribution of the species. more recently feng (2023), working with 87 hummingbird species, evaluated the effect of a series of buffers created around occurrences (from 5 to 5000 km) as calibration areas. the models calibrated with spatial buffers were compared with models calibrated with regions considered areas accessible to species. as a result, discrimination accuracy increased when the size of the calibration area was larger, but it reached a species-specific saturation threshold. although the evaluation method affected this criterion, it was typically estimated to be less than 200 km. surprisingly, model accuracy based on areas accessible to species was comparable to the saturation accuracy of models when spatial buffers were used. in the present study, the comparison between unbiased models within each species’ accessible areas and the rest of the models highlights a significant challenge: determining the accessible area of a species. unbiased models consistently demonstrated excellent performance, achieving high functional accuracy across all species. in contrast, the rest of the models exhibited more variable accuracy, with some species displaying even negative functional accuracies. this discrepancy underscores the importance of considering species-specific characteristics and selecting appropriate calibration areas to ensure accurate model predictions. the findings suggest that modeling within species’ accessible areas can mitigate biases and improve model reliability, highlighting the potential benefits of adopting unbiased approaches in sdm studies. luna et al. – species distribution model accuracy influenced by calibration area 51 in contrast to our findings, lobo and tognelli (2011) reported different results. they investigated the impacts of spatial sampling bias, and the number and location of pseudo-absences on model accuracy using virtual species. their results indicated that the number of pseudo-absences and the presence of spatial bias in sampling localities, along with their interaction, exerted a substantial influence on model accuracy (interpreted here with roc-auc, but they also evaluated sensitivity and specificity). as expected, higher number of pseudo-absences coupled with an absence of spatial bias yielded superior models. the location of pseudo-absences, whether distributed across the entire study area or restricted to regions outside the environmental envelope of the species, had a relatively smaller effect on model accuracy. they acknowledged that this might be attributed to the low relative occurrence area (only 3.5% of the total study area inhabited by the species). when contrasting our work with theirs, some differences stand out. they did not account explicitly for the dispersal capabilities of the species, potentially resulting in an overestimation of the realized distribution (araújo and pearson 2005). in addition, they employed thresholded maps instead of considering the more accurate occurrence probability (leroy et al. 2016; meynard et al. 2019). also, the sdm was calibrated using the same bioclimatic variables that were used to create the virtual species niche. they used a threshold once more for model evaluation, restricting the use of data pertaining to the actual probability of occurrence. finally, they simulated a single virtual species, while our study encompassed the results of 100 virtual species (in our work the lower relative contribution for the calibration area was 2.5% and the highest accounted for 93%). these distinctions highlight the complexities involved in modeling species distributions and the importance of considering multiple factors to enhance the robustness and ecological relevance of such models. increasing the extent of the calibration area involves incorporating data that are environmentally more distant (on average) from the occurrences. consequently, the discrimination accuracy of the model may increase due to the ease to parameterize models with good discrimination capacity but that are low in useful information (barve et al. 2011; acevedo et al. 2012). this could be the result of larger calibration areas covering places with appropriate environmental conditions that are unoccupied because of biotic interactions and/or dispersal constrains, which could induce overfitting to conditions close to the occupied localities (anderson and raza 2010). on the other hand, the importance of coarse-scale factors such as climate may be underestimated at small calibration areas (barve et al. 2011; acevedo et al. 2012). our results confirm previous research highlighting the impact of the calibration area on model accuracy and show the complexity involved in its selection (vanderwal et al. 2009). the environmental equilibrium assumption, wherein the species is adapted to its current environmental conditions, emphasizes the importance of choosing a calibration area that accurately reflects these conditions (araújo and pearson 2005). nonetheless, our findings, along with earlier research, indicate that there is no agreement on the ideal calibration area (rotllan-puig and traveset 2021; machado-stredel et al. 2021). as some correlative methods for estimating ecological niches rely on contrasting the environmental characteristics of known occurrence sites with those from the available conditions across the study area, it becomes imperative to delineate and comprehend the potential range the species might have explored. this is crucial because the absence of a species outside its accessible area is not necessarily due to abiotic or biotic factors. instead, a species may be absent from suitable regions simply due to its inability to disperse and reach those areas (anderson and raza 2010; barve et al. 2011). the way in which sdm handle collinearity between the predictor variables (feng et al. 2019), sample bias (ranc et al. 2017; inman et al. 2021), and model complexity are additional aspects that could impact the model accuracy (merow et al. 2014). however, when viewed in a multifactorial way, our results demonstrate that these factors have a significantly smaller impact than the selection of the calibration area. concerning this, barbet‐massin et al. (2012) observed that the impact of different methodological decisions in model quality varied depending on the specific sdm employed. for machine learning techniques (boosted regression trees and random forest), the number of pseudoabsences explained a greater amount of deviance (between 20 and 85%) than the weighting scheme and the method for selecting pseudo-absences (less than 15%). thus, exploring a range of modeling techniques beyond maxent is needed to further understand their differential responses and implications for sdm techniques. regarding the correlation between functional and discrimination accuracy, our results show that luna et al. – species distribution model accuracy influenced by calibration area 52 this correlation is relatively low. discrimination accuracy based on both data partition schemes were a misleading measure of functional accuracy. however, contrary to our expectations, models selected with random partitioning demonstrated a better correlation between functional and discrimination accuracy. this is surprising because we expected lower correlation in this scenario due to the lower degree of independence between data used for evaluation and calibration. these findings suggest that random partitioning might be more effective in selecting models that accurately predict the true suitability values across the landscape, despite the theoretical advantages of block partitioning in creating geographically independent evaluation sets. one possible reason is that certain ranges of environmental values may be geographically clustered and so are not utilized during calibration in block partitioning, unlike random partitioning, which avoids this stratification (kass et al. 2021). this contrasts with the findings of warren et al. (2020), who observed a better correlation using geographical block partitioning. despite the relatively low correlation between functional and discrimination accuracy. the fact that the best models selected based on random and block data partitioning exhibited functional accuracies exceeding 0.5 highlights that, in certain cases, the choice of data partitioning strategy can lead to the selection of models with reasonably high functional accuracy, even when their overall discrimination accuracy showed limited alignment with the functional accuracy metric. the inappropriate selection of the calibration area has significant implications for modeling applications (barve et al. 2011). the consequences extend to critical aspects such as the inaccurate estimation of the extent of occurrence and area of occupancy, which may lead to misguided conservation priorities (vanderwal et al. 2009), failure to generate appropriate mechanistic hypotheses about the parameters governing species distributions (vanderwal et al. 2009), and distorted response curves (thuiller et al. 2004). therefore, it is essential to give careful thought and choose the calibration area to guarantee the validity and robustness of sdm. constructing spatial buffers around known occurrences, reflecting the potential spatial range a species could explore, offers a straightforward method for delineating a calibration area (feng 2023). this approach aligns more closely with the theoretical considerations of species’ mobility (holloway and miller 2017), providing a more realistic foundation for sdm exercises. in line with other authors (vanderwal et al. 2009; barve et al. 2011; feng 2023), we recommend that species distribution modeling exercises should initiate with exploratory analyses of the calibration area, assessing the extent that can yield both the most accurate results and a biologically meaningful fit between species occurrence and predictor variables. to sum up, this research advances the discussion of calibration areas, shedding light on their nuanced impacts on sdm accuracy. acknowledging the complexity of these considerations, our findings contribute to the ongoing refinement of sdm practices, emphasizing the need for tailored approaches in different ecological contexts. the present study has demonstrated that the area of calibration is one of the most important factors affecting the functional accuracy of species distribution models using maxent. other factors, such as the value of the regularization parameter and the presence of collinearity between the predictor variables, have a much smaller impact. acknowledgments sl thanks the national council of research and technology (conahcyt) for the doctoral scholarship provided. we also thank alejandra arreola-triana and two anonymous reviewers for their helpful suggestions to earlier versions of this manuscript. competing interests the authors have declared that no competing interests exist. references acevedo, p., a. jiménez-valverde, j. m. lobo, and r. real. 2012. delimiting the geographical background in species distribution modelling. j. biogeogr. 39:1383–1390. aiello‐lammens, m. e., r. a. boria, a. radosavljevic, b. vilela, and r. p. anderson. 2015. spthin: an r package for spatial thinning of species occurrence records for use in ecological niche models. ecography 38:541–545. anderson, r. p., and a. raza. 2010. the effect of the extent of the study region on gis models of species geographic distributions and estimates of niche evolution: preliminary tests with montane rodents (genus nephelomys) in venezuela: effect of study region on models of distributions. j. biogeogr. 37:1378–1393. araújo, m. b., and r. g. pearson. 2005. equilibrium of species’ distributions with climate. ecography 28:693–695. barbet‐massin, m., f. jiguet, c. h. albert, and w. thuiller. 2012. selecting pseudo‐absences for species distribution models: how, where and how many? methods ecol. evol. 3:327– 338. luna et al. – species distribution model accuracy influenced by calibration area 53 barve, n., v. barve, a. jiménez-valverde, a. lira-noriega, s. p. maher, a. t. peterson, j. soberón, and f. villalobos. 2011. the crucial role of the accessible area in ecological niche modeling and species distribution modeling. ecol. model. 222:1810–1819. booth, t. h. 2022. checking bioclimatic variables that combine temperature and precipitation data before their use in species distribution models. austral ecol. 47:1506–1514. booth, t. h., h. a. nix, j. r. busby, and m. f. hutchinson. 2014. bioclim: the first species distribution modelling package, its early applications and relevance to most current maxent studies. divers. distrib. 20:1–9. boria, r. a., l. e. olson, s. m. goodman, and r. p. anderson. 2014. spatial filtering to reduce sampling bias can improve the performance of ecological niche models. ecol. model. 275:73–77. brodie, s., j. a. smith, b. a. muhling, l. a. k. barnett, g. carroll, p. fiedler, s. j. bograd, e. l. hazen, m. g. jacox, k. s. andrews, c. l. barnes, l. g. crozier, j. fiechter, a. fredston, m. a. haltuch, c. j. harvey, e. holmes, m. a. karp, o. r. liu, m. j. malick, m. pozo buil, k. richerson, c. n. rooper, j. samhouri, r. seary, r. l. selden, a. r. thompson, d. tommasi, e. j. ward, and i. c. kaplan. 2022. recommendations for quantifying and reducing uncertainty in climate projections of species distributions. glob. change biol. 28:6586–6601. bucklin, d. n., m. basille, a. m. benscoter, l. a. brandt, f. j. mazzotti, s. s. romañach, c. speroterra, and j. i. watling. 2015. comparing species distribution models constructed with different subsets of environmental predictors. divers. distrib. 21(1):23–35. castillo-torres, p. a., e. martínez-meyer, f. córdova-tapia, and l. zambrano. 2017. potential distribution of native freshwater fish in tabasco, mexico. rev. mex. biodivers. 88:415–424. clark, j. d., j. e. dunn, and k. g. smith. 1993. a multivariate model of female black bear habitat use for a geographic information system. j. wildl. manage. 57(3):519–526. de marco, p. jr, and c. c. nóbrega. 2018. evaluating collinearity effects on species distribution models: an approach based on virtual species simulation. plos one 13(9):e0202403. dormann, c. f., j. elith, s. bacher, c. buchmann, g. carl, g. carré, j. r. garcía marquéz, b. gruber, b. lafourcade, p. j. leitão, t. münkemüller, c. mcclean, p. e. osborne, b. reineking, b. schröder, a. k. skidmore, d. zurell, and s. lautenbach. 2013. collinearity: a review of methods to deal with it and a simulation study evaluating their performance. ecography 36:27–46. duque-lazo, j., h. van gils, t. a. groen, and r. m. navarro-cerrillo. 2016. transferability of species distribution models: the case of phytophthora cinnamomi in southwest spain and southwest australia. ecol. model. 320:62–70. espindola, s., j. l. parra, and e. vázquez-domínguez. 2019. fundamental niche unfilling and potential invasion risk of the slider turtle trachemys scripta. peerj 7:e7923. feng, x. 2023. a test of species’ mobility hypothesis in ecological niche modelling. j. biogeogr. 50:1955–1966. feng, x., d. s. park, y. liang, r. pandey, and m. papeş. 2019. collinearity in ecological niche modeling: confusions and challenges. ecol. evol. 9:10365–10376. fick, s. e., and r. j. hijmans. 2017. worldclim 2: new 1‐km spatial resolution climate surfaces for global land areas. int. j. climatol. 37:4302–4315. grömping, u. 2006. relative importance for linear regression in r: the package relaimpo. j. stat. softw. 17(1):1–27. guisan, a., and n. e. zimmermann. 2000. predictive habitat distribution models in ecology. ecol. model. 135:147–86. hijmans, r. j., s. phillips, j. leathwick, and j. elith. 2017. package ‘dismo’. circles 9(1):1–68. hirzel, a. h., j. hausser, d. chessel, and n. perrin. 2002. ecological-niche factor analysis: how to compute habitat-suitability maps without absence data? ecology 83(7):2027– 2036. hirzel, a. h., g. le lay, v. helfer, c. randin, and a. guisan. 2006. evaluating the ability of habitat suitability models to predict species presences. ecol. model. 199:142–152. holloway, p., and j. a. miller. 2017. a quantitative synthesis of the movement concepts used within species distribution modelling. ecol. model. 356:91–103. huang, e., y. chen, m. fang, y. zheng, and s. yu. 2021. environmental drivers of plant distributions at global and regional scales. glob. ecol. biogeogr. 30:697–709. hurlbert, s. h., r. a. levine, and j. utts. 2019. coup de grâce for a tough old bull: “statistically significant” expires. am. stat. 73:352–357. inman, r., j. franklin, t. esque, and k. nussear. 2021. comparing sample bias correction methods for species distribution modeling using virtual species. ecosphere 12(3): e03422. jiménez, l., and j. soberón. 2022. estimating the fundamental niche: accounting for the uneven availability of existing climates in the calibration area. ecol. model. 464:109823. jackson, s. t., and j. t. overpeck. 2000. responses of plant populations and communities to environmental changes of the late quaternary. paleobiology 26(suppl): 194–220. kass, j. m., r. muscarella, p. j. galante, c. l. bohl, g. e. pinilla‐buitrago, r. a. boria, m. soley‐guardia, and r. p. anderson. 2021. enmeval 2.0: redesigned for customizable and reproducible modeling of species’ niches and distributions. methods ecol. evol. 12:1602–1608. leroy, b., c. n. meynard, c. bellard, and f. courchamp. 2016. virtualspecies, an r package to generate virtual species distributions. ecography 39:599–607. lobo, j. m., a. jiménez‐valverde, and r. real. 2008. auc: a misleading measure of the performance of predictive distribution models. glob. ecol. biogeogr. 17:145–151. lobo, j. m., and m. f. tognelli. 2011. exploring the effects of quantity and location of pseudo-absences and sampling biases on the performance of distribution models with limited point occurrence data. j. nat. conserv. 19:1–7. luna et al. – species distribution model accuracy influenced by calibration area 54 machado-stredel, f., m. e. cobos, and a. t. peterson. 2021. a simulation-based method for selecting calibration areas for ecological niche models and species distribution models. front. biogeogr. 13(4):e48814. masin, s., a. bonardi, e. padoa-schioppa, l. bottoni, and g. f. ficetola. 2014. risk of invasion by frequently traded freshwater turtles. biol. invasions 16:217–231. merow, c., m. j. smith, and j. a. silander. 2013. a practical guide to maxent for modeling species’ distributions: what it does, and why inputs and settings matter. ecography 36:1058–1069. merow, c., m. j. smith, t. c. edwards, a. guisan, s. m. mcmahon, s. normand, w. thuiller, r. o. wüest, n. e. zimmermann, and j. elith. 2014. what do we gain from simplicity versus complexity in species distribution models? ecography 37(12):1267–81. meynard, c. n., b. leroy, and d. m. kaplan. 2019. testing methods in species distribution modelling using virtual species: what have we learnt and what are we missing? ecography 42:2021–2036. naimi, b., n. a. s. hamm, t. a. groen, a. k. skidmore, and a. g. toxopeus. 2014. where is positional uncertainty a problem for species distribution modelling? ecography 37:191–203. owens, h. l., l. p. campbell, l. l. dornak, e. e. saupe, n. barve, j. soberón, k. ingenloff, a. lira-noriega, c. m. hensz, c. e. myers, and a. t. peterson. 2013. constraints on interpretation of ecological niche models by limited environmental ranges on calibration areas. ecol. model. 263:10–18. peterson, a. t., and j. soberón. 2012. species distribution modeling and ecological niche modeling: getting the concepts right. nat. conserv. 10:102–107. phillips, s. j., r. p. anderson, and r. e. schapire. 2006. maximum entropy modeling of species geographic distributions. ecol. model. 190:231–259. ranc, n., l. santini, c. rondinini, l. boitani, f. poitevin, a. angerbjörn, and l. maiorano. 2017. performance tradeoffs in target-group bias correction for species distribution models. ecography 40(9):1076–87. rotllan-puig, x., and a. traveset. 2021. determining the minimal background area for species distribution models: minbar package. ecol. model. 439:109353. sillero, n., and a. m. barbosa. 2021. common mistakes in ecological niche models. int. j. geogr. inf. sci. 35(2):213–226 sillero, n., s. arenas-castro, u. enriquez‐urzelai, c. gomes vale, d. sousa-guedes, f. martínez-freiría, r. real, and a. márcia barbosa. 2021. want to model a species niche? a step-by-step guideline on correlative ecological niche modelling. ecol. model. 456:109671. soberón, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodivers. inform. 2:1–10. soberón, j., and b. arroyo-peña. 2017. are fundamental niches larger than the realized? testing a 50-year-old prediction by hutchinson. plos one 12(4):e0175138. shcheglovitova, m., and r. p. anderson. 2013. estimating optimal complexity for ecological niche models: a jackknife approach for species with small sample sizes. ecol. model. 269:9–17. soberon, j., and a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodivers. inform. 2:1–10. soberón, j., and m. nakamura. 2009. niches and distributional areas: concepts, methods, and assumptions. proc. natl. acad. sci. usa 106(2):19644–19650. srinivasulu, a., b. srinivasulu, and c. srinivasulu. 2021. ecological niche modelling for the conservation of endemic threatened squamates (lizards and snakes) in the western ghats. glob. ecol. conserv. 28:e01700. stockwell, d. r. b., and i. r. noble. 1992. induction of sets of rules from animal distribution data: a robust and informative method of data analysis. math. comput. simul. 33:385–390. taylor, a. t., t. hafen, c. t. holley, a. gonzález, and j. m. long. 2020. spatial sampling bias and model complexity in stream‐based species distribution models: a case study of paddlefish (polyodon spathula) in the arkansas river basin, usa. ecol. evol. 10:705–717. thuiller, w., l. brotons, m. b. araújo, and s. lavorel. 2004. effects of restricting environmental range of data to project current and future species distributions. ecography 27:165– 172. valavi, r., g. guillera‐arroita, j. j. lahoz‐monfort, and j. elith. 2022. predictive performance of presence‐only species distribution models: a benchmark study with reproducible code. ecol. monogr. 92(1):e01486. varela, s., r. p. anderson, r. garcía-valdés, and f. fernández-gonzález. 2014. environmental filters reduce the effects of sampling bias and improve predictions of ecological niche models. ecography 37:1084–1091. veloz, s. d. 2009. spatially autocorrelated sampling falsely inflates measures of accuracy for presence-only niche models. j. biogeogr. 36:2290–2299. vanderwal, j., l. p. shoo, c. graham, and s. e. williams. 2009. selecting pseudo-absence data for presence-only distribution modeling: how far should you stray from what you know? ecol. model. 220:589–594. warren, d. l., n. j. matzke, and t. l. iglesias. 2020. evaluating presence‐only species distribution models with discrimination accuracy is uninformative for many applications. j. biogeogr. 47:167–180. warren, d. l., and s. n. seifert. 2011. ecological niche modeling in maxent: the importance of model complexity and the performance of model selection criteria. ecol. appl. 21:335–342. weaver, k. f., t. anderson, and r. guralnick. 2006. combining phylogenetic and ecological niche modeling approaches to determine distribution and historical biogeography of black hills mountain snails (oreohelicidae). divers. distrib. 12:756–766. luna et al. – species distribution model accuracy influenced by calibration area 55 white, j. w., a. rassweiler, j. f. samhouri, a. c. stier, and c. white. 2014. ecologists should not use statistical significance tests to interpret simulation model results. oikos 123:385–388. wildlife conservation society, and center for international earth science information network. 2005. last of the wild project, version 2, 2005 (lwp-2): global human influence index (hii) dataset (geographic). palisades, new york: nasa socioeconomic data and applications center (sedac). zhu, g.-p., d. rédei, p. kment, and w.-j. bu. 2014. effect of geographic background and equilibrium state on niche model transferability: predicting areas of invasion of leptoglossus occidentalis. biol. invasions 16:1069–1081. microsoft word rees-et-al-irmng-revised-2.docx irmng 2006–2016: 10 years of a global taxonomic database tony rees11, leen vandepitte2, wim decock2 and bart vanhoorne2 1private address, new south wales, australia. 2flanders marine institute/vlaams instituut voor de zee, wandelaarkaai 7, 8400 oostende, belgium abstract.—irmng, the interim register of marine and nonmarine genera, was commenced in 2006 as an initiative of the australian obis node (obis australia) following an analysis of the taxonomic names management needs of the ocean biogeographic information system (obis). the main objectives were to produce a hierarchical classification of all life, both extant and fossil, to at least generic level (and to species as data were readily available) and to provide a tool to distinguish marine from nonmarine, and extant from fossil taxa. over its first 10 years of operation irmng has acquired almost 487,000 of an estimated 510,000 published genus names (including both valid names and synonyms) in addition to almost 1.8 million species names, of which 1.3 million are considered valid. throughout this time irmng data have been available for public query via a dedicated web interface based at csiro in australia, as well as being supplied as bulk downloads for use by a range of global biodiversity projects. over the period 2014-2016 responsibility for the system has been passed to the data centre division of the flanders marine institute (vliz) in belgium, which is continuing the maintenance and development of irmng at its new web location, www.irmng.org. with its present estimated holdings of >95% of all published genus names (plus associated authorities and years of publication) across all taxonomic domains, including fossil as well as extant taxa, within an internally consistent taxonomic hierarchy, irmng is at present uniquely placed to provide an overview of “all life” to at least generic level, to permit the discovery of trends in publication of genera through time, to provide preliminary information on the marine vs. nonmarine and extant vs. fossil status of the taxa concerned, and to generate lists of both unique and non-unique names (homonyms sensu lato) for the benefit of users of biodiversity data. key words.—taxonomic databases; biodiversity; marine taxa; terrestrial taxa; extant taxa; fossil taxa; biological classification. 1 email address for correspondence: tony.rees@marinespecies.org. biodiversity informatics, 12, 2017, pp. 1-44 1 introduction the desire to obtain an overview of taxon names for “all life”, whether extant-only or also including named fossil taxa, is one which has repeatedly emerged since the time of linnaeus in the eighteenth century, and has been revisited at intervals for various groups e.g., for animals (agassiz 1848, sherborn 1902-1933, neave 1939-1996), plants (hooker and jackson 1895, andrews 1970 plus supplements, willis 1973, farr et al. 1979), prokaryotes (euzéby 1997current), and viruses (international committee on taxonomy of viruses 1971-current). since the advent of the internet age a number of these works have been migrated to the web and/or new initiatives started, in particular the integrated taxonomic information system itis)2, the “species 2000” collective3, and their combined initiative the catalogue of life4, currently in its 15th annual edition (2016). for extinct (fossil) taxa, the paleobiology database5 is also becoming progressively more complete. nevertheless, both of these latter two compilations still have some gaps: the 2016 catalogue of life currently covers only 84% of world diversity6, while coverage of the paleobiology database is limited by its user contributions, and currently missing at least 100,000 valid fossil species names (data presented later in this paper). on account of such gaps in coverage it is still possible to encounter both species names and those of genera, either within the scientific literature or included among names submitted to biological information systems as identifiers for accompanying data, not presently held in the above compilations and which cannot therefore be placed taxonomically, and/or ancillary information discovered, without additional manual effort. genesis of this work biodiversity data aggregation projects such as the ocean biogeographic information system (obis7) and the global biodiversity information facility (gbif8) commenced in the early 2000s with an aim to bring together available occurrence data on extant species in either just 2 http://www.itis.gov/ 3 http://www.sp2000.org/ 4 http://www.catalogueoflife.org/ 5 https://paleobiodb.org/ 6 http://www.catalogueoflife.org/content/frequently-asked-questions#8 7 http://www.iobis.org/ 8 http://www.gbif.org/ the marine domain (obis) or all habitats (gbif). incoming data to such projects typically comprise location data (latitude-longitude) together with an associated species name and potentially other information. obis and gbif therefore require a taxonomic hierarchy in which to place the incoming data by means of the supplied taxonomic names, as well as (in obis’ case) the capacity to discriminate marine from nonmarine taxa (and by extension, extant taxa from fossil) to support the intended purpose of displaying content for marine, extant species only. in 2003-4 a process was commenced to extract all species names in the then-latest version of the catalogue of life (2003 edition), manually assign them a marine/nonmarine flag, and use this process to place names held in obis in the catalogue of life taxonomic hierarchy along with an indication of their marine/nonmarine status (see rees and zhang 2007). however, at that time a non-trivial proportion (around 30%) of obis names held could not be taxonomically resolved using the catalogue of life, prompting a search for an alternative method of name resolution together with a means to query their associated marine/nonmarine, and extant/fossil status. the solution adopted was to attempt to create a more comprehensive index of genera only. this was considered more tractable for two reasons: first, the number of names to be compiled (valid names plus synonyms) is potentially an order of magnitude smaller for genera than for species, and second, a number of genus-level compilations were already in existence, which in combination could offer greater taxonomic completeness than available equivalents at specific level. for taxa covered by the zoological code (international commission on zoological nomenclature 1999) there exists a digitised online version of neave’s nomenclator zoologicus, updated to the end of 20049, for taxa covered by the botanical code (mcneill et al. 2012) the online version of index nominum genericorum10, while prokaryotes are covered by the list of prokaryotic names with standing in nomenclature11 and viruses by the taxonomy releases of the international committee on taxonomy of viruses12. nomenclator zoologicus 9 http://ubio.org/nomenclatorzoologicus 10 http://botany.si.edu/ing/ 11 http://www.bacterio.net/ 12 http://www.ictvonline.org/taxonomyreleases.asp biodiversity informatics, 12, 2017, pp. 1-44 2 and index nominum genericorum incorporate indications of extant/fossil (but not marine/nonmarine) status in most cases, plus family assignment is provided in the case of the botanical, prokaryote and virus name compilations. in practice, the project was started without relying on these particular resources, but they were made available within the first 2 years of the project (see appendix 1). the value of an all-genera list for indexing species data lies in the binomial format of species names in all domains except viruses, whereby the including genus is represented as the first element of the binomial name. therefore, even if information on a particular species is not held explicitly, by inference it can frequently be presumed to inherit particular traits including taxonomic placement and marine vs. nonmarine, and extant vs. fossil information, to the extent that the latter two aspects are unambiguous (e.g., marine only, extant only). further, the process of creation of relevant trait information within the database can be simplified via hierarchical propagation from higher to lower levels, (for example all echinoderms are marine, all trilobites are extinct), so that detailed manual effort is reserved only for “mixed cases”, namely genera or higher taxa containing both extant and fossil, and/or both marine and nonmarine components. a supplementary benefit of creating and populating a taxonomic framework to genuslevel is that the process of later adding species (when required) is simplified, since in the majority of cases their appropriate parent at generic level will already be present and can be used as an attachment point for the species concerned, with no additional effort needed to decide on their higher taxonomic placement. the name applied to the project, the interim register of marine and nonmarine genera (irmng), was intended to reflect that of erms, the european register of marine species (costello et al. 2001), a precursor to the world register of marine species or worms (worms editorial board 2007-current), while indicating that the geographic coverage was now global and the scope extended to include nonmarine as well as marine taxa, however with the focus switched to genera in the first instance. the inclusion of the term “interim” in the project name was intended to convey the fact that in this instance, the focus would be on relatively rapid assembly of at least a “first pass” product for use by clients in the short term, which could then be improved as an iterative process over time, rather than waiting for every included data item to be as exhaustively checked as might be the case when the ultimate in data quality is required. irmng goals the goals of irmng as envisaged at the project outset were as follows: 1. assemble as completely as possible a list of published genus names covering all domains of life, within a taxonomic hierarchy constructed as a set of logically consistent relationships (allkin et al. 1992): e.g., a given genus cannot be in multiple families simultaneously; all taxa (apart from the root, i.e. “all life”) must have a parent record which must be at a higher rank; all records are discoverable by traversing the taxonomic tree from top downwards; a parent cannot be deleted while it possesses “live” child records; etc. non-current names (synonyms), when known to the compilers, should be pointed to their relevant valid name as information is at hand. 2. as many genus names as possible are then to be flagged marine or nonmarine (or both), extant or fossil (or both), for the use of clients such as obis and others. 3. as species data are readily available (for example as already compiled in the catalogue of life), add these to the system connected to relevant genera, and continue the flagging to specific level by either automatic propagation (where possible) or manual flagging as needed. 4. genus names would be subject to reasonable scrutiny upon acquisition to remove duplicates, provide author citations in a consistent form, and insert names into the most appropriate position in the irmng taxonomic framework. for reasons of limited resourcing and to facilitate the assembly of an “interim” product within a realistic time frame, species names from generally reliable sources including the catalogue of life, regional compilations, and some museum databases would generally be accepted without additional checks. 5. make irmng content available in bulk form as download files, either of the entire database content or of selected data items as appropriate, for upload into relevant client systems. as the project progressed, additional aims were incorporated, including: 6. make irmng content available for external query—and subsequently, web-based edit for designated content editors—via a publicly accessible website. biodiversity informatics, 12, 2017, pp. 1-44 3 7. include original publication information for names (at generic level in the first instance) as held in external resources: initially as “microcitations” (typically a journal or publication name, volume, and page in work, imported as a text string), later expanded with the capability to hold article titles and/or full (atomised) bibliographic citations in a separate “literature” module. 8. include numeric identifiers for names as held in selected other systems (initially aphia/ worms, subsequently ion: see below) to enable cross linkages to those systems via the web as desired. 2006 catalogue of life identifiers for species, originally uploaded to irmng, are not persistent and have subsequently been deprecated. 9. include the ability to generate lists of homonymous names (i.e., where the same name has been used to denote different taxa on different occasions) via the web interface, in response to several user requests. 10. include a “near” or “fuzzy” name matching process so that incoming misspelled names could be matched against correctly spelled target names, when held. this capability was found to be very valuable to users wishing to detect misspelled names on their own lists as well as providing the opportunity for the system administrator to detect and rationalise variant spellings of the same name already held within the system. 11. include a range of online search methods to suit user requirements, such as search by full or partial scientific name, or by authority or year published, filter by higher taxonomic group, etc. also, provide an option for remote users to input lists of names for bulk real-time matching to irmng holdings. 12. provide the capability to report statistics on irmng data holdings at any time, including number of names held at different ranks, number of valid names vs. known synonyms vs. unresolved names in particular groups, numbers of marine/nonmarine and extant/fossil names per group, and more. estimations of the ultimate data volume are given later in this work, but at the outset it was considered sensible to plan for up to at least millions of names (i.e., comparable with the catalogue of life and other current systems), with the expectation that several hundred thousand of these would be genera. selected irmng design principles what is a “taxonomic name” in irmng? since irmng aspires to hold only one record per “taxonomic name” (or more specifically, taxonomic name instance), it is necessary to articulate what is meant by this term in the irmng context. by taxonomic name instance we mean the scientific name used by a particular author at a point in a specific publication to formally describe a new taxon, thus representing a unique combination of taxonomic name, author and position of the relevant entry (typically page, or sometimes even line) in the particular cited work. it follows, therefore, that variations in the representation of an author name (such as presence or absence of initials, abbreviated form versus full form, etc.) and/or subgenus inserted into a species name do not comprise multiple taxonomic name instances under this definition, neither do variations of how the work itself is cited. as an example, in patterson et al. (2010; their figure 1) numerous possible variant representations of the name cyclotrachelus sodalis are given including “cyclotrachelus sodalis (le conte)”; “cyclotrachelus (e.) sodalis (lec.)”; “c. (evarthrus) sodalis (lec. 1848)” and many more; for irmng purposes these would all be rationalized to a single name instance, which in this case would be represented as cyclotrachelus sodalis (le conte, 1848). similarly, the names “acanthoperla cavalier-smith” and “acanthoperla cavalier-smith in cavalier-smith & chao, 2012” would be considered the same name instance and represented by only one irmng record, as would ficus röding, 1798 versus ficus bolten, 1798, the two variants referring to the same work which has been ascribed variously to either bolten or röding in the past (the latter is now the accepted author of this work). on the other hand, if the same (or different) author has published the same name as new on multiple occasions (perhaps first as a nomen nudum, followed by a subsequent valid publication) then this represents multiple instances of “name+author+position in cited work” and would be indexed accordingly as multiple records in irmng. this approach of creating only one record per “taxonomic name instance” contrasts with the concept of “name strings” as discussed, e.g., in patterson et al. (2010), and as collected by indexers of such name strings including ubio13, the index to organism names (ion)14, and the global names index (gni)15, with the result that indexing operations such as 13 http://www.ubio.org/ 14 http://www.organismnames.com/ 15 http://gni.globalnames.org/ biodiversity informatics, 12, 2017, pp. 1-44 4 those repositories potentially contain many more names-as-strings, each typically with its own designated identifier (such as a life science identifier or lsid) than is the case for irmng. the corollary of this is that incoming names to irmng may frequently need deduplication (multiple representations—of authorities in particular—being reconciled to a single preferred form), but that on occasion, care must also be taken not to amalgamate name instances which look similar (or even identical) but in fact refer to different taxa, i.e., homonyms in the broad sense. this includes both within-code homonyms (homonyms sensu stricto) and duplicate names across different nomenclatural codes (“transregnal homonyms” in patterson et al., 2016), for example the same name used for both an animal and a plant, or a plant and a bacterium. a practical example is provided by the genus name ceratium, which has been used multiple times for different taxa, with each instance also potentially cited slightly differently in the different compilations that may be used as input to irmng. rationalising and deduplicating these citations (name strings) is key to deciding how many irmng records should be created for this name. combining data from five different sources (nomenclator zoologicus, index nominum genericorum, index fungorum, index to organism names and global names index) we can arrive at the following rationalised irmng list (table 1). while rules mandate that homonyms within the same nomenclatural code are not permitted (junior homonyms requiring a replacement name or nomen novum), no such restrictions apply to the same name being currently valid across different codes and so (a single instance of) the same name can be concurrently valid in “plants” (historic usage, i.e. including algae and fungi), animals, and prokaryotes, (although probably not in viruses, whose genus names all end in “virus”). thus in the ceratium example given above, the oldest genus name (ceratium schrank, 1793, the dinoflagellate), if considered a zoological name, preoccupies subsequent identically named genera in zoology but not in botany, for which ceratium j.b. albertini & l.d. schweinitz, 1805 remains a validly published name, with ceratium blume, 1825 an invalid (illegitimate) junior homonym. ranks in the irmng data structure at the project outset and over the period represented by this report, for simplicity in data handling and also following the then-current edition of the catalogue of life, the irmng data structure above genus included only the “linnaean” ranks, i.e, kingdom, phylum (= division in botany), class, order and family, other intermediate ranks being dropped when supplied (but in some cases captured as an accompanying text remark). with the move of irmng content into the vliz data structure, the capability to easily add intermediate ranks exists and this has been taken up to a limited degree from 2017 on, commencing with the addition of subphyla such as insecta, crustacea and myriapoda in arthropoda, and vertebrata, urochordata and cephalochordata in chordata at this time. protozoa and chromista have also been revised, with intermediate ranks added thus far between the ranks of kingdom and class, following the scheme of ruggiero et al. (2015). over time the entry of further names at intermediate ranks can be expected; however, at time of writing the “linnaean” ranks mentioned above remain the only ones consistently populated for all names. also, for reasons of efficiency in data collection, ranks below the level of species have been ignored at the present time, although this may be revisited at a future date. treatment of subgenus names at present, irmng does not include the rank of subgenus in its design concept, although this may be revisited in future versions. however, in zoology, names published as subgenera are deemed to be simultaneously available (published) at generic level—also at any other level in the “genus group”—via the principle of coordination (international commission on zoological nomenclature 1999, article 43). as an example, the subgenus abyssopinna created by schultz & huber (2013) within the genus pinna linnaeus, 1758 is also available (without change in cited authorship) if subsequent authors wish to use it at generic rank on account of a change in taxonomic opinion. hence, within irmng, abyssopinna is included as a published name at generic rank, although at that level it is listed as a synonym of its containing genus pinna. a second issue regarding subgenera is that, in zoology, at specific level it is legitimate to include a subgenus in parenthesis in between the genus name and the specific epithet, thus (taking the case of abyssopinna as introduced above) the same species name can be represented as both pinna (abyssopinna) epica and pinna epica. in biodiversity informatics, 12, 2017, pp. 1-44 5 table 1. irmng records for genus “ceratium”, with equivalents in selected other available data sources, namely nomenclator zoologicus (“a”), index nominum genericorum (“b”), index fungorum (“c”), ion (“d”) and global names index (“e”). irmng numeric identifiers (“irmng id”) are allocated to the names on addition to the irmng system and are persistent (also are independent of identifiers for the same names in other systems). note: where the irmng “preferred form” of the cited author and year differs from some or all of the sources used, the format has been selected and/or adjusted according to the principles set out in appendix 2. taxonomic assignment verbatim records in sources a-e, arranged in one group per “taxonomic name instance” (refer caption for abbreviations used) equivalent name + authority as stored (in irmngpreferred form) irmng id remarks class dinophyceae (dinoflagellates) ceratium schrank 1793 (a, d) ceratium schrank (b, d) ceratium f. schrank, 1793 (e) ceratium schrank, 1793 (e) ceratium schrank, 1793 1274897 class myxogastrea (myxomycetes) ceratium albertini & schweinitz 1805 (a, d, e) ceratium albertini et schweinitz (b) ceratium alb. & schwein. (c, e) ceratium albertini & s. 1805 (d) ceratium j.b. albertini & l.d. schweinitz, 1805 1273955 currently a synonym of ceratiomyxa j. schröter in engler & prantl, 1889 family orchidaceae (angiosperms) ceratium blume (b) ceratium blume, 1825 (e) ceratium blume, 1825 1274128 currently a synonym of eria j. lindley, 1825 family pyralidae (moths) ceratium thienemann 1828 (a, d, e) ceratium thienemann, 1828 1274013 currently a synonym of phycita curtis, 1828 family pyralidae (moths) ceratium gistl 1848 (a, e) ceratium gistl, 1848 1274070 a later usage of ceratium thienemann, 1828 and thus also a synonym of phycita curtis, 1828 phylum rotifera (rotifers) ceratium agassiz 1846 (a, d, e) ceratium agassiz, 1846 1274194 currently a synonym of keratella bory st. vincent, 1827 biodiversity informatics, 12, 2017, pp. 1-44 6 accordance with present irmng data conventions (appendix 2), and again following the earlier practice of the catalogue of life, only the binomial version is used for species, and where incoming species names include the subgenus this will be removed before uploading the name to the system (it may however be checked to see whether or not the subgenus name is held, a least for zoological names, and where missing this should also be uploaded as a generic name in its own right). by contrast, subgenera in botany and bacteriology are not automatically available for use as genera from their original publication; they are therefore not indexed as genera in irmng unless subsequently formally raised to that rank via a separate nomenclatural action. botanical and bacteriological subgenus names are rarely encountered as a portion of a species name but where they are, the relevant element would be removed: for example “moraxella (subg. branhamella) caviae” would be entered simply as moraxella caviae. unavailable names in addition to available names (the zoological term, equivalent to validly published names in botany), irmng includes a component of unavailable names including some nomina nuda, original or subsequent literature misspellings, unjustified emendations, suppressed names, as well as bacterial names without standing in prokaryotic nomenclature (for definitions refer the relevant nomenclatural codes), chiefly when these have been indexed in nomenclators such as nomenclator zoologicus or included in data compilations such as the catalogue of life. such names are retained in irmng since they may appear in published literature and/or other taxonomic data compilations and may be required for correct assignment of accompanying information; whenever possible they are pointed to the current accepted name for the same taxon, when known. a few names in irmng are also classified as later usages: this applies when a name is indexed as new (for example in nomenclator zoologicus) but additional investigation indicates that the original valid publication of the name in fact occurs in an earlier work by the same, or a different author. as an example, in nomenclator zoologicus the name onychites is credited to zakrzewski, 1886 but, according to other published sources, in fact dates from a publication by quenstedt, 1856. in this case, “onychites zakrzewski, 1886” is retained in irmng but cited as a later usage of onychites quenstedt, 1856, for which a separate entry is created. taxonomic uncertainty and divergent views taxonomy is not an exact science and the views of multiple authors regarding the taxonomic placement, rank, or status of a particular name or taxon do not always necessarily coincide, and may also change through time with advancing knowledge. certain taxonomic information systems attempt to accommodate this by supporting multiple taxonomic views—an example being the present encyclopedia of life16, which is capable of displaying classifications for any included taxon from multiple sources that may not always agree with each other. by contrast, as is the case with other compendia such as catalogue of life and itis, irmng supports a single taxonomic view at this time which, if not always completely up-to-date, can be (and in some cases has already been) upgraded to follow arrangements generally recognised as “authoritative” for the group in question. thus, angiosperm taxonomy, previously following the apg iii treatment (angiosperm phylogeny group 2009) has recently been upgraded to follow apg iv (angiosperm phylogeny group 2016); the higher classification of fungi was updated in 2009 to follow the then-latest version of index fungorum (kirk 2001-current), extant fishes follow the 2009 version of eschmeyer’s catalog of fishes (eschmeyer 2000-current), and families of extant crustaceans, gastropod molluscs, and coleoptera have been adjusted to follow recent treatments (de grave et al. 2009; bouchet and rocroi 2005; bouchard et al. 2011). when alternative views exist and are known to the irmng compilers, these are typically captured in a taxonomic note appended to the record, for example the taxonomic remark for the linnaean class “aves” (irmng id: 1142) presently reads: “treated by some recent authorities, and ruggiero et al., 2015, as a subclass of reptilia; maintained as class at the present time in common with catalogue of life (2016), worms (2017) and elsewhere”. changes introduced in the most recent treatments (for example combining or splitting families, genera or species) will not always be 16 http://eol.org/ biodiversity informatics, 12, 2017, pp. 1-44 7 immediately reflected in irmng but, as resources permit, are hoped to be incorporated either directly from the literature or as they make their way into sources used as input to irmng, such as catalogue of life, the world register of marine species, and others. irmng attributes as introduced above, a key initial driver for the compilation of irmng was the requirement for obis to distinguish between marine vs. nonmarine, and extant vs. fossil taxa. accordingly, a system of “flags” was incorporated into the initial (2006-2014) version of irmng whereby a single “habitat flag” could be set to values corresponding to “marine,” “nonmarine,” “marine plus nonmarine,” or null, i.e. not yet entered, together with a source for the relevant setting (“habitat flag source”). for extant/fossil status, a flag could be set to values corresponding to extant, fossil, extant plus fossil, or null. the concept of “marine” is intended to conform to the definition used in erms, namely “…broadly defined to include intertidal (littoral) and brackish water habitats, defined as up to the strandline or splash zone above the high tide mark and down to 0.5 ppt [parts per thousand] salinity in estuaries” (costello 2000). this definition excludes salt marshes, which were considered in erms to be out-of-scope by virtue of being elsewhere included in terrestrial ecosystems. in irmng, “marine” is considered to exclude primarily terrestrial or freshwater species that may be found at sea in a nonobligate manner (for example certain ducks) but to include species which spend a regular portion of their life in the marine environment such as shorebirds and waders. species which spend portions of their life cycle both at sea and on land (e.g. marine mammals, birds and reptiles that return to the land to breed) are assigned both marine and nonmarine status, as are species which alternate between marine and freshwater habitats at different stages of their life cycles. from 2016 onwards, irmng supports additional options for habitat flagging in that “marine” is further divided into marine and brackish which can be separately assigned the states yes, no, or unknown, while “nonmarine” has been replaced by the categories freshwater and terrestrial which can be assigned equivalent states as required. to transfer legacy data most effectively from the initial version of irmng to the current one, names previously flagged “marine only” (which includes a component of brackish water organisms) are represented as marine = yes, brackish = unknown, while former “nonmarine only” has been represented as marine = no, freshwater = unknown, terrestrial = unknown. while this representation is suboptimal for some purposes, it still permits discrimination of marine vs. nonmarine taxa by interrogating the flag for “marine” = yes or no, and is capable of being upgraded through time as additional resources are available. the irmng extant/fossil flag (“extant” being renamed “recent” from 2016 onwards) supports the options of recent only, recent plus fossil, fossil only, and unknown. “recent” is defined as including taxa alive at any time since the beginning of modern scientific investigation, broadly interpreted as post 1500 a.d., thus including some species that have become extinct since that date (such as the dodo, raphus cucullatus, last seen alive in 1662) but excluding species such as the extinct moas of new zealand (of dinornis and related genera), never seen alive by naturalists and known only from sub-fossil deposits older than c.1500 a.d. irmng implementation environment, programming language, virtual (web) and physical location to address irmng goals, for the initial version a custom oracle® database was constructed in august-september 2006 at csiro marine laboratories in hobart, australia, together with an application developed in the oracle pl/sql programming language to address the requirements for data query and web display, web editing, and other administrative functions. the database was made available for live query via the web in october 200617 and was subsequently upgraded to also incorporate input name parsing and fuzzy matching using the algorithm “taxamatch” as described in rees (2008a,b). throughout 2006-2014 new content was added to the system (more detail given below) until 2014 when a process was commenced to transfer the database to the flanders marine institute (vliz) in belgium, as part of an agreed migration between the two institutions. over the next two years the content was progressively transferred to new data tables in belgium and a new domain name was 17 http://www.cmar.csiro.au/datacentre/irmng/ biodiversity informatics, 12, 2017, pp. 1-44 8 established for the project18, with release of the new system at vliz announced in july 2016 after completion of suitable testing. meanwhile, content acquired for the system post-july 2014 has been added to the copy at vliz but not to the csiro copy, which is thus effectively “frozen” as at that date and is planned to be deprecated once its functionality and use as a target for links from third party compilations is no longer required. in its new location at vliz, irmng shares the same data table structure and code base as that developed for aphia, the data store at vliz which supports worms and other taxonomic databases (vandepitte et al., 2015); however, the data content of irmng and aphia currently remains separate. this means that (for example) irmng ids are not interchangeable with worms (=aphia) ids, irmng and aphia displayed content for the same taxon may not be identical, and references are not shared between the two systems, although these aspects could change in the future. the underlying database system at vliz is microsoft sql server with a web interface implemented in the php programming language. the complete database design includes over 400 data fields spread over 81 tables (for a more detailed description see vandepitte et al. 2015, also a diagrammatic summary of aphia table relationships as shown via the worms website19); only a subset of these are presently used for irmng, of which the principal ones are described in the next section. the actual migration process for the irmng data was nontrivial, and involved comparison of the existing irmng and aphia data structures, mapping of the relevant data fields, import of the data, and then an appraisal and iterative adjustment process to ensure that fields had been mapped in the most meaningful way and no content deemed important was lost during the migration. separately from the actual data migration, new web interfaces for general information, data search, and online editing of both taxonomic names and relevant literature were constructed (based on pre-existing templates but customized to suit irmng requirements), and tested / adjusted as needed prior to public release. with the move of the database to vliz, a number of additional search options (some via 18 http://www.irmng.org 19 http://www.marinespecies.org/structure/relationships.html “advanced search”) developed for aphia/ worms but not previously available in irmng have also been implemented. new features include, among others, a browsable taxon tree; search by irmng id as well as taxonomic name; search by taxon status (i.e., accepted, unaccepted, nomen nudum, etc.); search by habitat and/or extant/fossil flag status; limiting a search to above or below a designated taxon rank; search by note type and any particular included text desired; and literature (source) search, in addition to filter by higher taxon and search by scientific name and/or name author, which were previously also offered in the original search interface at csiro. by default, searches via the new interface execute across, and return results from, multiple ranks simultaneously, an improvement on the original capability at csiro which only returned results from the same taxonomic rank as specified by the user. a previous capability for input of multiple names to the standard “search” box has been replaced by a more flexible taxon match tool, based on that previously developed for worms, via which a user can submit up to 1500 rows of data (taxon names) for matching and be returned a file in csv format containing a flexible, user-selected range of output variables including any or all of irmng_id, scientific name, authority, accepted name, higher classification, quality status, taxon status, environment, and citation (= publication details for the name). additional description of the detailed operation of this tool is available via the taxon match tool user manual20. examples of the current (may 2017) irmng “basic search” interface, with a sample search result and example taxon page, are given in figures 1-3, with the additional options offered via “advanced search” in figure 4. irmng data fields irmng holds a set of data fields associated with every taxonomic name instance, plus (in additional tables) details for references/sources, values displayed in picklists used at data entry time, etc.; in addition, it holds internal administrative data on persons who have made or can make changes. available data fields associated with each taxonomic name instance are shown in table 2. 20 http://www.marinespecies.org/tutorial/taxonmatch.php biodiversity informatics, 12, 2017, pp. 1-44 9 figure 1. irmng “basic search” interface at vliz, current (may 2017) version (http://www.irmng.org/aphia.php?p=search). figure 2. initial portion of the search results page produced by the search as specified in figure 1. the search operates simultaneously across all ranks (unless otherwise requested), includes names from botany, zoology bacteriology and viruses as appropriate, and also both extant and fossil taxa, the latter indicated with the dagger symbol (†) as a suffix. biodiversity informatics, 12, 2017, pp. 1-44 10 figure 3. example generic level taxon page produced by clicking on the link to “ammopemphix loeblich, 1952” from the search results shown in figure 2 (a small number of cited sources are omitted for clarity). note that, as is the case with the majority of genus-level pages in irmng, the information displayed is aggregated from multiple sources, in this case systema naturae 2000 and for the initial genus name and authority, nomenclator zoologicus for the original publication details, genus spelling and authority verification, and a taxonomic remark, worms for the family allocation, status as an accepted (current) name, and worms identifier (included in the preformatted link as displayed), ruggiero et al. (2015) for the classification between kingdom and order, and the ion database for the ion identifier which is used to create relevant deep links to both ion and the bionames database. the two displayed child species are sourced from aphia2006/erms and worms (2013 version), respectively. additional information on sources used is given in appendix 1. biodiversity informatics, 12, 2017, pp. 1-44 11 figure 4. irmng “advanced search” interface at vliz, current (2017) version. name status “accepted” corresponds to available names (in botany: validly published names) not presently known to be synonyms, in other words both valid (current) names plus a subsidiary component of names not yet scrutinized for taxonomic status. figure 5. growth in irmng content, 2006-2016, as numbers of records held at family, genus and species levels. biodiversity informatics, 12, 2017, pp. 1-44 12 table 2. irmng data fields for taxonomic names. “required” fields are those that must be populated before a new irmng record cannot be saved; “desirable” are fields that should be completed when information is available, either at initial upload or as a subsequent activity. priority field name/s remarks auto-created record attributes irmng id unique, numeric identifier (auto-allocated, as next available number). record creator and creation date record status by default, “active”; can subsequently be altered to either “in quarantine” or “deleted” if required. required fields scientific name if genus rank and above, the uninomial name; if species rank, the specific epithet only (to be combined with its associated genus for record display, searching, and data export). rank taxonomic rank of the name, selected from a set of pre-defined options (kingdom through species) parent parent taxon, as taxonomic name selected from a picklist (then held internally as the relevant irmng id), permitting placement in the taxonomic hierarchy. top-level taxa, presently kingdoms, have the single parent “biota”, which has irmng id =1. status name status i.e. one of the following options as per the present worms standard: • accepted • unaccepted • nomen nudum • alternate representation • nomen dubium • temporary name • taxon inquirendum • interim unpublished. habitat flag/s, i.e., marine, brackish, freshwater, terrestrial for names entered 2016 onwards, at least one of these four flags must be set to either “yes” or “no” (permitted values for each flag are “yes”, “no” or unknown”). accepted name if the taxon status is entered as “accepted”, this is by default the taxon name entered; if the name status is “unaccepted”, the currently accepted name is required, selected via a picklist drawn from names already in the system. desirable fields authority the authorship associated with the scientific name, in irmng-preferred form (refer appendix 2) extant/fossil status permitted options are recent only; fossil only; recent + fossil; unknown. linked source a repeatable link to an irmng reference (entry in “sources” table via relevant source id, see below), biodiversity informatics, 12, 2017, pp. 1-44 13 categorized under one of the following headings: • original description • basis of record • additional source • source of synonymy • redescription • new combination reference • status source • toxicology source • taxonomy source • ecology source • identification resource • subsequent type designation • misapplication • original description (unavailable nomenclaturally) • emendation (re-diagnosis of genus) • verified source for family • verified source for genus • current name source • extant flag source • habitat flag source • context source a number of these fields had no equivalent in the original (csiro) version of irmng and are not populated at the present time, but are available for use from 2016 onwards. “original description” corresponds to the “publication” field in the initial version of irmng, and may be populated with a full bibliographic citation or simply with an abbreviated “microcitation” (e.g., journal, volume, page) as shown in the example given in figure 3 above. “basis of record” is used as required to identify the source from which the name was acquired, as per the examples given in appendix 1. unaccept reason a free text field to supply more information to the user on the reason for a given name having the status “unaccepted”; examples: junior synonym, junior homonym, name published in suppressed work, misspelling, etc. supporting information as available notes additional notes (e.g., nomenclatural or taxonomic comments, or information on habitat or geologic range), either as present in the source used, or added by an irmng editor. notes take the form of a repeatable free text field (optionally with an associated source), assigned to one of the following categories: • classification • descriptive information biodiversity informatics, 12, 2017, pp. 1-44 14 • habitat • nomenclatural status • nomenclature • subfamily • taxonomic remark • type species as cited • validity nomenclatural status is completed where applicable, for example: orthographia (misspelling), nomen novum, nomen nudum, etc. “nomenclature” defines the nomenclatural code governing the use of the name, chosen from iczn, icbn (historic name for the current botanical code or icnafp), bacteriological code (bc) and virus code (vir). “type species as cited” is provided to hold any text providing this information (e.g. as available in an irmng data source) without the requirement for the name to presently exist within the irmng system, and could later be converted to a relevant irmng id as available. “taxonomic remarks” may include any comments regarding alternative taxonomic views of the taxon in question, remarks on the availability of a replacement name if the name is a homonym, presentation including cited authorship, etc. id in other system/s system-specific identifiers for the same name in other systems, currently comprising a dedicated field for itis taxonomic serial number (tsn), plus a repeatable “links” field presently holding web links (including relevant identifiers) to the world register of marine species (via aphia id), and the index to organism names, and bionames, both of which can be addressed using ion ids. type species for genera, indication of the type species (selected from a picklist of relevant names already entered), plus optional type designation method selected from a picklist. original name for species (e.g., new combinations), indication of the original name on which the new combination is based. gender for genera and species, indication of the grammatical gender of the relevant genus name or epithet, selected from the options masculine/feminine/neuter/ unknown. updater name, date, aspect changed record update history, when changed since original entry (field repeatable as needed). details on specific aspects changed are held in the database and can be reviewed by administrators as required. biodiversity informatics, 12, 2017, pp. 1-44 15 it should also be noted that the order of record creation proceeds downwards through the taxonomic hierarchy; for example, if a new family, genus and species are required, the family must be created first in order that it can be cited as the parent of the genus, and the same for the genus and species records. further, since every record “knows” the identity of its parent in the irmng system, lists of child taxa at any level are generated automatically for display as part of a taxon page (e.g., species of a particular genus) without any requirement for data entry at the level of the parent. “active” records are those displayed in response to user searches. the status of a record can be changed (by a relevant irmng editor) to either “in quarantine” or “deleted” as required. records in quarantine are basically in-progress, “work” records requiring more effort or scrutiny before a decision is made to make them publicly viewable, while records may be deleted once a decision has been made that they are no longer required, for example being duplicate entries, “temporary” parent records that no longer have any children, or records based on erroneous external data that are not required to be kept. deleted records remain on the system so that first, their numbers are never re-used, which could lead to confusion; second, their current status and reason for deletion can be tracked for administrative purposes, and incoming queries against them can be redirected to a designated active record; and third, they can be reactivated (undeleted) if required, for example if the record was deleted in error, or if the reason for its deletion ceases to apply. irmng sources are held in a dedicated module labelled “literature” which is available for user search via its own dedicated interface linked from the irmng home page. an individual source can be either a free text statement such as “cavalier-smith, 1992” (typically a legacy entry from the previous version of irmng), a microcitation (as per the example shown in figure 3), or a fully atomised bibliographic reference, and can optionally include an online link to a more complete version of the cited work on an external accessible site. irmng sources have a numeric id which is allocated on source creation and may be linked to multiple statements about multiple taxa: for example the source “nomenclator zoologicus” is presently linked to >300,000 such statements. with the move of irmng to vliz and the data structure developed for aphia, certain other fields have become available for use within the irmng data structure, namely vernacular names (and relevant language), distribution, specimen details for a species, feeding type, images, and “contexts”, the last being a set of “tags” controlling within which of the multiple systems supported by aphia the relevant record is to be displayed. currently, none of these fields is populated for irmng but they may be utilized in the future. since names may be uploaded with only the “minimal” information in the first instance, an important aspect of ongoing irmng population (in addition to adding new names not previously held) is to add or upgrade missing information items for names already within the system. accordingly, a web-based data edit interface is provided (to suitably authorized editors) which supports both the creation of new records, and upgrade or alteration of any aspect of existing records. data can also be added or amended as bulk operations on request to the database administrators, which can be preferable where particular tasks would be inefficient for editors to perform via the online edit interface on a recordby-record basis; examples of the latter might be supplying a list of names/irmng ids with new information to be entered against each (such as a habitat or extant/fossil flag where these are currently missing), or requiring the move of a substantial number of child taxa from one parent to another. population of the irmng system the initial task in developing irmng in 2006 was to construct a higher taxonomic framework (kingdom through family) into which incoming lists of genera and/or species could be placed. concentrating on the “linnaean ranks”, the family-level treatment of parker (1982) was used for extant taxa, as the most comprehensive then available despite being some decades outof-date (a few minor groups accidentally omitted from the latter were added as required from other sources). subsequently, fossil-only families were also added using a digital summary of data from benton (1993). input of the family treatment of parker was facilitated by digitization of the relevant portion of the printed work in collaboration with the team at the sealifebase project21, while for the fossil families, a spreadsheet version of the holdings from benton 21 http://www.sealifebase.org/ biodiversity informatics, 12, 2017, pp. 1-44 16 (1993) was used as cited in appendix 1. habitat flags were created for all families based on indications in both works, those in the benton compilation already being present within the electronic file used. with the initial family and higher taxonomic framework in place (subject to later revision as required), available sets of genus and/or species names were formatted for addition to irmng and progressively uploaded to the system (refer appendix 1 for full details). in addition to the major sources listed in appendix 1, a large number of smaller compilations and individual papers have been consulted, either as sources of additional names or to obtain supplementary information about specific name instances. as sources are consulted, their details are added to the literature module as described above, from which their details can then be displayed on relevant taxon pages. over the ten-year period covered by this report, irmng holdings have grown to 486,652 genus names, of which 362,597 are presently designated “accepted”, within a higher taxonomic structure consisting of 23,243 families, 3,107 orders, 587 classes, 162 phyla and 7 kingdoms, including a small number of higher taxon names currently regarded as synonyms. the period of most rapid growth for families was 2006-2007 including uploads from parker (1982) and benton (1993), for genera 2006-2009 (including initial uploads from systema naturae 2000, catalogue of life—the latter without authorities, index nominum genericorum and nomenclator zoologicus), and for species 2006-2007 (corresponding to the addition of over 1.2 million names from the catalogue of life), supplemented by a boost between 2011 and 2013 corresponding to the addition of more species names from the hallan biology catalog and the 2013 edition of worms. overall growth in irmng holdings over time at family, genus and specific levels is illustrated in figure 5. figure 6 shows a breakdown of the genuslevel content as at december 2016 by major taxonomic group, also indicating the proportions of genus names within each group flagged extant (= recent) only, extant + fossil, fossil only, and unknown extant/fossil status as held at that time. irmng data loading, standardization, and quality checks for the initial components of irmng content to be uploaded, the data were formatted to suit the irmng data structure and then added to the system, with a unique irmng identifier created for each name. as information was available, synonyms and some unavailable names were pointed to their equivalent accepted name (valid taxon) as appropriate. habitat and extant/fossil flags were added whenever possible by flagging entire families using available information from the compilations by parker, benton, or elsewhere, and otherwise as discrete exercises based on a range of published literature and/or web sources. subsequent batches of candidate data for addition to the system are first checked to see if names present are already held on the system before the remainder are prepared for upload by adding the minimum required fields, plus as many of the desirable/optional fields listed in table 2 as can easily be populated. some names in incoming lists might not be loaded if, for example, they are misspellings or simply variants of names already held, while on occasion, names apparently already held will still be loaded if they represent a new instance, for example a previously unknown homonym of a name already in the system, possibly even at a different rank (such as the protist phylum sagenista cavalier-smith, 1995, homonym of the genus sagenista bohart, 1967 in hymenoptera). the matching of candidate names for upload with names already held on the system may be complicated by varying degrees of authority match, for which a decision must be made on a case by case basis aided on occasion by prior experience and/or supplementary research. for instance, on some occasions apparently dissimilar cited authorities can in fact represent the same publication instance, for example ficus röding, 1798 vs. ficus bolten, 1798 (in this case, both röding and bolten have been cited as the author of the work in question, so only one taxonomic name instance is involved). conversely, apparently similar authorities may represent different name publication instances, for example two instances of sosxetra walker, 1862, but with slightly different publication details (in lepidoptera and hymenoptera, respectively), or enhydrus macleay, 1825 (a beetle) vs. enhydrus macleay, 1925 (a mammal). to assist in authority comparisons, where names are identical, the “authority match” component of taxamatch (rees 2014) can be employed to quantify and pre-sort names according to the degree of authority similarity, as biodiversity informatics, 12, 2017, pp. 1-44 17 figure 6. irmng genus holdings at december 2016 by major taxonomic category, with breakdown by extant/fossil status as held in the system. extended scale at bottom of chart (100,000-180,000) applies to the second bar of hexapoda only. categories displayed represent a mix of formal plus informal groups as devised for the csiro implementation; to replicate these via the new search interface may require some component groups to be summed (for example to produce totals for non-protist algae, “pisces”, and plantae as displayed). figure 7. breakdown of irmng genus holdings at december 2016 by key attributes entered. for additional detail refer text. biodiversity informatics, 12, 2017, pp. 1-44 18 a precursor to manual scrutiny of the differences thus detected. in practice, higher calculated degrees of author similarity (for example >0.7 on a 0-1 scale) suggest that two name instances may be the same (enabling generally more rapid scrutiny), with the converse applying for lower levels of similarity (e.g., <0.3), suggesting that the name instances may well be different, again assisting the speed of a potential accept/reject decision. names with an intermediate degree of author match are set aside for more detailed review as required. incoming data to irmng are typically reviewed for errors and/or inconsistencies, and may be edited as needed to suit in-house conventions or correct obvious data errors as required. present conventions used for data standardization in irmng are listed in appendix 2. where possible, the bulk of the desired editing, such as restoring missing diacritical marks on author names, expansion of abbreviated botanical authors and addition of missing dates, is carried out prior to data upload. another area of quality assurance is provided by the irmng taxon match tool (which again uses taxamatch) which, when presented with an incoming name not found in other reputable sources, can return suggestions as to names already in the system of which it may be a misspelling. once loaded, the availability of more extensive sets of comparative irmng data provides opportunity for additional scrutiny via review of names sorted binned in different ways (such as alphabetic sort by scientific name or author, review of all genera in a family or species in a genus), which can further assist in revealing inconsistencies or errors (such as exact or near duplicate entries originating from multiple sources, the same author name cited in multiple ways, or an inconsistency in cited dates) which can then be addressed as resources are available. irmng completeness and limitations numbers of records held, versus anticipated “complete” data holdings in an ideal situation it would be helpful to know in advance how many published family, genus, and species names exist so that first, the eventual size of the system when fully populated could be gauged, and second, the degree of irmng completeness could be estimated at any given time. parker (1982) gives a list of approximately 6,900 extant families, with a further 3,600 fossil-only families listed in benton (1993), both listings considered essentially complete by the relevant authors at those times, and all uploaded to irmng as described above. a further 12,700 family names (including 1,400 presently regarded as synonyms, plus an unknown number of variant spellings or synonyms not yet detected) have subsequently been added from a range of sources, suggesting that irmng is likely to be quite complete for accepted names at family level. a comparison of irmng extant, accepted family names with the 9,650 extant families listed in ruggiero (2014) would also be instructive but has yet to be undertaken. it should, of course, be noted that a large number of family names proposed in the past are presently regarded as synonyms, but since irmng does not aspire to include these exhaustively, the extent of this situation is not relevant here. for genera, the most complete vetted (i.e., deduplicated) recent listings are the approximately 357,000 zoological names indexed by nomenclator zoologicus up to 2004, together with the 69,000 botanical genera in index nominum genericorum 2012 version (minus overlap with nomenclator zoologicus in 2,500 cases), plus prokaryote genus names from the catalogue of life and the list of prokaryotic names with standing in nomenclature (2,200 names in 2008) and virus names from the catalogue of life and the ictv virus database (420 names in 2011). these totals include both accepted and unaccepted names (valid names plus synonyms) and also, in the case of nomenclator zoologicus, a component of unavailable names including published misspellings and nomina nuda. together, these total approximately 426,100 names, to which should be added first, totals for names within scope for, but missed by, the nomenclators as given above, and second, the number of names published since the cut-off dates for those compilations and their respective versions. estimates for the first of these are not available; however, irmng thus far contains approximately 22,000 dated names in zoology and botany published before 2004 plus a further 1,000 undated names, not in relevant nomenclators. if it is reasonable to suggest that over the past ten years, irmng has uncovered perhaps 50-70% of such missing names, then a further 10,000-20,000 may exist to be found and indexed through time. over the period since 2004, an estimate of irmng completeness is somewhat easier to biodiversity informatics, 12, 2017, pp. 1-44 19 determine since annual publication rates of new genus names have been relatively stable at around 2,500 per year for all groups (irmng data, 2000-2009 average) so, for the period 2010-2016 inclusive a total of around 17,500 published names would be expected. irmng presently holds 9,831 published names for the period 2010-2016 leaving an estimated shortfall of around 7,700 published names to be acquired in this respect. overall, the shortfalls estimated comprise around 18,000-28,000 generic names suggesting that a total of around 510,000 published genus names may exist, with present irmng holdings (487,000 genera) therefore comprising an estimated 95% of all such names published. at specific level, estimates of accepted (valid) names do exist, namely those of chapman (2009) for extant taxa (1.9 million) and raup (1986), alroy (2002) and prothero (2013) for fossils (250,000-300,000 depending on source consulted), giving a present published total of up to 2.2 million accepted species as at 2009, also increasing by up to around 20,000 new names per year22. in the context of “all names” this number must be expanded to include both synonyms and previous or alternative binomials (i.e., genus+species combinations) now outdated by subsequent genus transfers. the precise extent of the “synonym problem” is unknown. in perhaps the most extensive, expert-vetted compilation to date for a major portion of the living world (vascular plants and bryophytes), version 1.1 of “the plant list” 23 contains 350,699 species names listed as “accepted”, 470,624 as known synonyms, and a further 242,712 as “unresolved”, the majority of which are probably also synonyms. such values may be typical across other portions of the taxonomic realm, for example benton (2008) reports a synonymy/error rate within named dinosaur species of 48.2% (not including alternative combinations) while in sources considered by patterson et al. (2016) a synonymy rate of around 3 synonyms per accepted name is regarded as the most typical value. using this estimate, in addition to the approximate figure of 2.2 million species a further 6.6 million synonyms may exist, giving a total of approaching 9 million species names published to date. on this basis at 1.78 million species names irmng presently 22 http://www.esf.edu/species/ 23 http://www.theplantlist.org/ holds around 20% of estimated published species names at this time, leaving a further approximately 7 million to be collected should this be an eventual design goal. overall record completeness (at generic level) completeness of irmng data holdings should not just be assessed by simple numbers of records held, but also by the relative completeness of key attributes of those records. for this purpose, relevant data are presented in figure 7, with remarks against the individual measures used given below. regarding the attributes shown in figure 7, the following notes are applicable: • name verified from trusted source: this applies to incoming names which can also be found in major nomenclators for the group in question, or have been verified from their citation in the primary literature, either from the original description or as part of a formal taxonomic treatment. a small number of genus names (approx. 10,000) has entered irmng from sources not considered “trusted” in this sense, such as museum databases or third-party compilations, and do not match entries in major nomenclators on a first pass; such names are identified as candidates for further investigation and will either be accepted or rejected for continued irmng use in due course. • taxonomic position fully resolved: this means that a genus has been allocated to a known family; names not yet allocated to a family are associated in irmng with the next higher category for which a placement is available, such as “mammalia (awaiting allocation)”, “arthropoda (awaiting allocation)”, etc. this attribute is less than 100% complete because one significant source in particular (nomenclator zoologicus) does not contain family allocation for its included names, meaning that that the latter must be backfilled from other sources. • taxonomic status known: this covers names that are determined to be either accepted names for current taxa (valid name in zoology, current name in botanical usage), or unaccepted names (taxonomic synonyms, or unavailable names for other reasons as determined in relevant nomenclatural codes). names acquired solely from nomenclator zoologicus lack this information, which again must be backfilled from other sources. • marine/nonmarine status entered: once again, this information is not available via standard nomenclators; however, where the taxonomic biodiversity informatics, 12, 2017, pp. 1-44 20 position has been resolved to family (or higher taxon in some cases) this can frequently be allocated by inheritance of the relevant status from containing higher taxa. • extant/fossil status entered: major nomenclators such as index nominum genericorum and nomenclator zoologicus do contain indicators of extant/fossil status; however, those in the latter compilation are not exhaustive so that while a fossil indication can be relied on, names without a fossil flag have been found to be not exclusively extant. therefore, this attribute is not entered in irmng as “extant” until names sourced from nomenclator zoologicus lacking a fossil flag therein have been independently checked in other sources, leading to a lesser degree of completeness for this attribute than would otherwise be the case. • publication details entered: animal and protistan names from nomenclator zoologicus do have associated publication details (as microcitations), which is the chief function of such a nomenclator and these have been carried through to the relevant irmng records. equivalent publication details for botanical names were not included in the original download file from index nominum genericorum provided for irmng use in 2007 and so were not uploaded at that time; however, they are available via the latter’s website and may be added to irmng as an enhancement at a future time. • ion id held: ion covers the zoological subset of taxonomic names and creates ion identifiers for any names encountered during the creation of the zoological record compilation which indexes both the primary literature (in this regard, descriptions of new taxa from 1864 onwards) and the secondary literature. consequently it has created a large number of identifiers (over 3.5 million at the present time, with perhaps one tenth of these being genera); however, many of these are essentially duplicates in the irmng sense (different ids for versions of the same name differing only by slight variants in their cited authority), or in some cases represent literature misspellings. irmng has harvested only the ion ids which have been created from the original published descriptions of the taxa involved since these are of the highest quality and also of the most benefit to users, in that following the ion id as a deep link to either the ion or bionames database24 will lead to a more complete citation of the original description of the taxon involved. at 24 http://www/bionames/org/ present, ion ids are held for 43% of irmng genera, but this increases to 66% if those not in scope for ion (non-animal names, and animal names published prior to 1864) are excluded. • worms (aphia) id held: aphia ids were uploaded to irmng for linking purposes as part of a cross mapping exercise in 2013 and this process is intended to be repeated at intervals as content is added to both systems. since the worms system is currently limited to marine taxa in the main, also with relatively few fossils, the proportion of irmng genera with this attribute populated will not be likely to exceed about 12% (the proportion of accepted, extant eukaryotic species currently estimated to be marine using data from appeltans et al., 2012 and chapman, 2009), presuming that genus representation is similar to that for species and that trends for synonyms follow those for accepted names. in practice the proportion will be less again to the degree that fossil names are currently under-represented in worms. at present, worms/aphia ids are held for 12% of irmng genera, but this increases to 48% if those not in scope for worms (irmng genera not flagged marine = yes) are excluded. data are not presented above for species because this rank is not a primary goal for irmng at the present time, but since the majority of species records currently held are associated with genera already verified, taxonomically resolved, and possessing relevant habitat and extant/fossil flags, it can be expected that overall, species records are more “complete” than the values shown for genera for the displayed attributes with the exception of publication details and ion ids. publication details for species were generally not available in the sources used to compile irmng to date, though some will be available via deep links to ion and/or worms, while ion ids have only been entered where these link to original publication details for the name in question, thus excluding new combinations which (in zoology) are generally not tracked via changes in authorship and associated bibliographic citations. other known limitations and data gaps apart from the indications of present database completeness presented above, some other present limitations of irmng content should be noted, in particular: biodiversity informatics, 12, 2017, pp. 1-44 21 • despite reasonable efforts to avoid this, some genus records may contain inaccuracies in assigned habitat and extant/fossil flags, and in the currency or correctness of stated synonymy assertions and family assignments. • records between phylum and generic level (i.e., class, order, family) have not been subject to the same level of scrutiny as genera and may contain errors (e.g., misspellings and out-of-date names) or other inconsistencies. • the higher taxonomy of many groups may be outdated to varying degrees, and has only been revised to follow the latest published sources for selected groups (such as extant angiosperms) at this time. • updates to the status of irmng species can lag behind those for genera, so some “unaccepted” genera may still contain species flagged as “accepted” at this time. as an example, over 100 “accepted” species of the genus michelia linnaeus, 1753 were uploaded from the 2006 version of catalogue of life; however, according to more recent sources, that genus is now treated as a synonym of magnolia. transfer of the affected species has not yet been made in irmng, and in any case would not be done until a relevant source is located in which the relevant revised combinations are supplied. • to the extent that a particular group contains presently unallocated genera (for example, phylum mollusca currently contains 8,600 such names, order coleoptera 3,800, and class reptilia 1,900), listings of genera for its families may be incomplete, with some otherwise valid family names having few or even no listed genera in the irmng data structure. (over time, this situation should improve as “unallocated” irmng genera are scrutinized and further taxonomically resolved). • species records have received less scrutiny than genera at this time and are known to include some duplicates and misspellings as uploaded from the various data sources utilized. an effort will be made to further deduplicate and rationalize this element of irmng data holdings in the future. data gaps in irmng at generic level principally comprise a proportion of genera published since 2010 (2009 for prokaryotes and fungi), with coverage for animal genera continuing through to 2014 at a decreasing level. at levels above genus, coverage should be essentially complete with the exception of a small number of families or other higher taxa recently established or resurrected. as previously noted, at specific level irmng does not presently aspire to completeness and in general, gaps exist in the terrestrial area for some groups not covered in the irmng sources used to date (marine species are generally covered via recent updates from worms) and, more particularly, for fossils in many groups, although the latter could be addressed in part via future imports of relevant data, for example, as contained in the paleobiology database and elsewhere. homonyms in irmng homonyms in the strict sense are recognised only within a particular nomenclatural code, for example “botanical” groups (mcneill et al. 2012), zoological names (international commission on zoological nomenclature 1999), and prokaryotes (lapage et al. 1992; parker et al. 2015) with the same name being legitimately available (and therefore not technically a homonym) between codes. for example, the genus ficus is the current name for both a flowering plant and a gastropod, while peranema is both a protist and a fern. homonymy in the strict sense also excludes unavailable names such as nomina nuda and subsequent misspellings, which may nevertheless be found in the literature used as identifiers for taxa. in irmng, therefore, we use the term homonym in an expanded sense to cover any multiple instances of the use of the same name for different taxa whether within or between codes (also whether or not available), so that users can be alerted to sources of potential confusion and given a pointer to the fact that the taxonomic placement of a particular named taxon in external data may require to be checked further. a capability has accordingly been created within irmng to generate lists of homonymous (i.e., duplicate) names at any rank which is constructed on-the-fly from the database, in other words, as a duplicate name is entered to the system (or removed) the list is automatically updated without any requirement for manual curation. present statistics indicate that there are around 77,000 homonymous genus name instances (including small numbers of misspellings and nomina nuda in addition to validly published names), plus around 190 homonymous family names as encountered with incoming data during irmng construction (the list at family level is by no means exhaustive since many older / non-current family names are not presently included in irmng). a set of homonymous species names can also be biodiversity informatics, 12, 2017, pp. 1-44 22 generated on demand; this list is quite extensive (around 100,000 records / 30,000+ names, which may however also include some entries requiring deduplication) but the number reduces to just 160 presently known to irmng if it is restricted to species-level homonyms associated with different genus instances (e.g., abronia aurita the reptile, versus abronia aurita the angiosperm). an example from the present list of known family-level homonyms is shown in figure 8; lists for family, genus and specific level can be generated on demand via the “homonyms” link indicated on all present irmng pages, or directly via this url25. the incidence of genus-level homonymy can be illustrated by inspection of a list of irmng genera names simply sorted alphabetically (example shown in figure 9), in which the occurrence of names sharing the same spelling is readily apparent. furthermore, such sets are not limited to name pairs: the case of ceratium, with 6 instances, has been discussed earlier, while the dubious honour of the name with the highest level of homonymy presently held in irmng goes to “wagneria” with 14 separate instances (figure 10). it should also be noted that at specific level, certain slightly different epithet spellings are “deemed to be identical” under the zoological code (article 58) such as caeruleus/coeruleus/ ceruleus, or litoralis/littoralis. these are, however, not reported as homonyms in the irmng species lists at this time, which is presently restricted to exactly matching epithets and associated generic names. homonyms at generic level remain the biggest source of potential confusion for both acquisition of irmng content (e.g., determining to what genus name instance to attach incoming species data) and for users wishing to resolve taxonomic names to a position in the taxonomic hierarchy, especially at specific level where the author of the genus (as opposed to the species name) will normally not be included. the ability to generate lists of homonymous names, or simply to check an existing name to determine whether or not any homonyms may exist, is also a useful feature for taxonomists and has been employed in a number of cases to date in order to discover previously unsuspected cases of homonymy at generic level in particular, 25 http://www.irmng.org/homonyms.php resulting in the proposal of replacement names as required (e.g., ng and low 2010; zeidler 2017). irmng current clients and use cases web clients of irmng data over the past ten years have included a wide range of representatives of museums and herbaria, individual researchers, and members of the public from most countries in the world, as evidenced by accumulated logs of user searches by ip address and search term entered, plus email communications from specific users, typical enquiry rates via the web being in the order of tens to hundreds per day. a second, significant class of users comprises administrators of taxonomic information systems who prefer to receive a bulk download of irmng content (currently as data files, in future potentially also via a web service) for ingestion and re-use within their own systems. once loaded there, re-usage via such “bulk clients” is of course supplementary to that recorded in user logs recording traffic via the irmng web interface. the number of such bulk clients has grown over the past ten years and currently includes obis, gbif (döring 2017), the atlas of living australia (ala)26, worms, the open tree of life project (otol27; see also hinchcliffe et al. 2015), the global names index (gni), and the encyclopedia of life (eol)28. from 2017, the use of irmng data is also being investigated to potentially fill gaps in the present algal coverage of the catalogue of life29. a further capability, recently released at vliz, involves making irmng data accessible to machine-based query via dedicated web services, which opens up the potential to embed irmng queries as one element in a chain of machine-based reasoning: for example, a given query could seek habitat or extant/fossil status from irmng on a particular taxon before deciding whether to proceed further and extract additional information from a separate data source. specific uses of irmng include: • parse incoming or stored names of initially unknown taxonomic affinity, allocate to position in a taxonomic hierarchy based on their full species name when held, or the genus portion of the species name as applicable (also alerting to possible 26 http://www.ala.org.au/faq/species-data/ 27 https://tree.opentreeoflife.org/about/open-tree-of-life 28 http://eol.org/content_partners/676 29 http://www.catalogueoflife.org/col/details/database/id/501 biodiversity informatics, 12, 2017, pp. 1-44 23 figure 8. initial portion of the irmng family-level homonyms list as generated in may 2017; unaccepted names are indicated by the red circle enclosing an ‘x’, entirely extinct taxa by the dagger suffix (†). in the case of the two instances of family amphiporidae within the same nomenclatural code (one in nemertea based on amphiporus ehrenberg, 1831, one in porifera based on amphipora schulz, 1883), an application has been made (özdikmen and demir 2011) to remove the homonymy by emending the junior name (in porifera) to amphiporaidae; at the time of writing this case still awaits a decision from the relevant commission (international commission on zoological nomenclature 2017). biodiversity informatics, 12, 2017, pp. 1-44 24 figure 9. initial portion of the “complete” irmng genera list, sorted alphabetically, showing the presence of 4 homonym pairs (8 names), as indicated by red outlines, within the first 33 names listed as at may 2017. figure 10. irmng holdings for genus name = “wagneria” as at may 2017. note: synonymy has been researched for many, but not all irmng name instances at the present time, therefore a subset of names is not yet flagged as synonym in the above list. biodiversity informatics, 12, 2017, pp. 1-44 25 homonyms, i.e. a non-unique generic name within the irmng system): also referred to as a taxonomic name resolution service (see discussion). • obtain near (“fuzzy”) matches to an input name or names to cope with candidate misspelled names in incoming or presently stored data (a second component of a taxonomic name resolution service). • allocate incoming or stored names to categories i.e. extant, fossil or both; marine, nonmarine or both, for filtering and/or sorting purposes as desired. • generate top-down views of “all life” to at least generic level by traversing a taxonomic tree of kingdom through genus (and species where available), for comparison/integration with equivalent data from other sources. • generate lists of taxa and names on demand based on any characteristic held in the database (for example name begins with…, authority contains…), also generate lists of duplicate names (homonyms) as required. • provide cross walks / deep links to data held remotely in other systems (at this time worms, ion, and bionames) by holding those systems’ taxon or name identifiers within relevant irmng records. a schematic view of irmng data flows including sources, editor actions, and the range and nature of current clients as indicated above is represented in figure 11. irmng editing since its inception in 2006, irmng data compilation and editing has been the responsibility of obis australia, whose personnel have been assigned privileges at the time to edit all aspects of content in their designated taxonomic group/s, either as direct operations on the database or on offline copies of relevant data files that have then been uploaded to the live system. from 2016 onwards, a new web-based edit system has been constructed which permits authorized editors to similarly alter any aspects of taxa under their control including the creation and quarantining of records, correcting spelling errors, adding, changing, or deleting attributes, sources and links, moving child records to a different parent, and so on. responsibility for allocation of edit privileges to relevant persons for the future now resides with the vliz data management team who will be managing this aspect in tandem with equivalent procedures for worms (for more details see discussion). irmng data availability data dumps of irmng content in dwc-a and/or native database format have been made available to clients on request since 2007, with a new dedicated download location created at the vliz instance via the url30, from which (at time of writing) the last three “snapshot” versions of irmng are available dating from january 2013, january 2014, and april 2017. meanwhile, the master version of irmng content is always accessible via the web and may contain updates that post-date any particular snapshot, plus in addition some supplementary information (principally sources for assertions used, together with the searchable literature module) not included in the darwin core data file(s). historic data dumps from irmng are also available via other locations; for example, at time of writing the 2013 version is searchable via the global names resolver31 and the 2014 version is searchable via a copy hosted at gbif32, plus an archived version is accessible via the holdings of the open tree of life33. irmng data are released without any irmng-issued copyright assertion, although unfortunately the same is not true for some of its constituent sources. worms data is currently cc-by (data are freely available for re-use but attribution is required), the catalogue of life and index fungorum declare their content to be available for re-use by non-commercial users only34,35, while (for example) ion, zoological record and algaebase assert that their content “may not be downloaded or replicated by any means” without appropriate permission36,37. nevertheless, as argued by patterson et al. (2014), a particular taxonomic name, its cited authority, taxonomic position, status and so on are simply facts or opinions sourced from the primary scientific literature and should not therefore be copyrightable by any downstream compilations. an exception may exist for comments added by record editors but arguably these should also be reproducible under “fair use”, with appropriate attribution, as is the case with such comments from other sources reproduced in irmng. for so long as the ipr 30 http://www.irmng.org/download.php 31 http://resolver.globalnames.org/ 32 http://www.gbif.org/dataset/0938172b-2086-439c-a1dd-c21cb0109ed5 33 http://purl.org/opentree/ott/ott2.8/inputs/irmng_dwc-2014-01-30.zip 34 http://www.catalogueoflife.org/content/terms-use 35 http://www.indexfungorum.org/names/indexfungorumpartnership.htm 36 http://organismnames.com/terms.htm 37 http://www.algaebase.org/copyright/ biodiversity informatics, 12, 2017, pp. 1-44 26 figure 11. schematic overview of irmng data flows. a wide variety of specialized and more general sources are used to populate the irmng database. all name-related sources go through a number of pre-processing steps prior to upload, including comparison with already available irmng data and assigning of the relevant flags. irmng content can also be added or upgraded remotely by authorized editors using an online edit interface. exports of irmng content to external systems are arranged either through darwin core archive (dwc-a) files or through web services, while human users can enter search queries and be returned relevant results via the irmng web portal. biodiversity informatics, 12, 2017, pp. 1-44 27 situation remains untested, irmng data downloads are presently accompanied by a statement “irmng data to specific level incorporates some content from the catalogue of life, the world register of marine species (worms) and other providers and may be subject to their respective terms of use” and are not accompanied by a free-reuse license such as creative commons cc038; however, it is to be hoped that this may change in the future. irmng does not require attribution for its data (which in any case originate almost exclusively from other sources) although an acknowledgement is requested if users find its services of value in their work. discussion other taxonomic compilations irmng is presently unique in that it aspires to completeness, at generic level at least, for “all life,” within a coherent hierarchical data structure which attempts to conform to current taxonomic opinion, to the degree that present resources permit. in this respect it differs from strict nomenclatural compilations such as nomenclator zoologicus, index nominum genericorum, the international plant name index (ipni)39, ion and others which, in addition to their self-imposed taxonomic coverage limits, are concerned in the first instance with the date, authorship, and place of publication of scientific names (facts) and less on their valid name/synonym status and current taxonomic placement (both opinions, also subject to change through time). compilations that do share these goals include, e.g., itis, the catalogue of life (chiefly for extant taxa) and the paleobiology database (for fossils), as well as numerous more selective compendia such as worms. conceptually, irmng presently fills a gap that eventually should be occupied by content from a combination of the catalogue of life and the paleobiology database, if these compilations were complete and if the catalogue of life were extended to hold information on genera, and thus it is pertinent to assess their present levels of completeness with respect to irmng, in particular at generic level. some relevant comparisons with these and other available sources are given in table 3. 38 https://wiki.creativecommons.org/wiki/cc0/ 39 http://www.ipni.org/ from the data presented in table 3 it can be seen that irmng is currently the most extensive, vetted compendium of genus names across all groups and in addition, contains the extant/fossil and marine/nonmarine status flags of interest to a range of users for over 80% of its entries at generic level. only gni (not utilized as an irmng source to date) is likely to contain a potentially useful component of genus names not yet in irmng, but the effort to extract these would most likely be considerable, bearing in mind that gni contains a range of misspelled, malformed, or even non-names (e.g., vernacular names harvested by mistake) as well as multiple variants of existing names. in addition, since the gni data structure does not distinguish genus names from uninomials at other ranks, each candidate “new” name would have to be researched manually to see if it were even a genus before considering further for aspects of interest to irmng such as correct cited authority, taxonomic status, current placement and more. on the other hand, potential sources of additional species names include such compilations as ion, catalogue of life, the paleobiology database, ipni and more, from which the expert vetted component(s) would certainly be useful to add to irmng at some point in the future. large scale aggregators such as global names, gbif, encyclopedia of life, and open tree of life are a special case in that they are consumers of irmng content (among others) which they then integrate to produce “supersets” including irmng data in combination with other information (e.g., see döring 2017; rees and cranston 2017), in the same way that irmng is itself a data aggregator. this then generates a question as to whether such aggregators could, or should, then be used as a preferred alternative to irmng for general taxonomic name resolution. positives of this approach include the fact that a user will encounter non-irmng as well as irmngsourced data which can be useful to fill data gaps, especially at specific level. on the other hand, the irmng content in such systems will generally be a time-stamped snapshot, and “live” irmng data may have been adjusted/improved in the intervening time; irmng may have richer content on a given taxonomic name than is preserved in the version made available elsewhere; and irmng may also be searchable via ways not offered in the data aggregator in question. by the same logic, biodiversity informatics, 12, 2017, pp. 1-44 28 ta bl e 3. ir m n g h ol di ng s a nd a ttr ib ut es c ov er ed c om pa re d w ith m aj or o th er re le va nt c om pi la tio ns a s a t e nd 2 01 6. d at a so ur ce sc op e n o. o f g en us na m es h el d n o. o f s pe ci es na m es h el d g en us au th or iti es he ld ? fa m ily al lo ca tio n he ld ? (f or su bs id ia ry ra nk s) a cc ep te d na m e/ sy no ny m st at us he ld ? e xt an t/ fo ss il st at us he ld ? m ar in e/ no nm ar in e st at us he ld ? r em ar ks ir m n g a ll gr ou ps , e xt an t + fo ss il, ki ng do m s t o sp ec ie s 48 7, 00 0 1. 78 m ill io n ye s ye s ( ge ne ra 78 % , s pe ci es >9 5% ) ye s ( ge ne ra 67 % , sp ec ie s >9 5% ) ye s ( ge ne ra 85 % , sp ec ie s >9 5% ) ye s (g en er a 88 % , sp ec ie s >9 0% ) pu bl ic at io n de ta ils (a s m ic ro ci ta tio ns ) h el d fo r m an y ge ne ra , m os tly in z oo lo gy . c at al og ue o f li fe , 2 01 6 ve rs io n 85 % o f a ll ex ta nt g ro up s, sm al l n um be r o f f os si ls , ki ng do m s t o su bs pe ci es 19 3, 00 0 3. 1 m ill io n no ye s sp ec ie s: ye s; ge ne ra : n o ye s no ta xo no m ic c ov er ag e of e xt an t g ro up s c ur re nt ly in co m pl et e (e .g ., al l a lg ae , s om e ar th ro po ds m is si ng ). n o pu bl ic at io n de ta ils h el d fo r a ny n am es . g lo ba l n am es in de x a ll gr ou ps , e xt an t + fo ss il, al l r an ks 40 0, 00 060 0, 00 0? c. 1 6 m ill io n ye s no no no no n o ex ac t g en us c ou nt a va ila bl e. in cl ud es m an y du pl ic at ed n am es w ith m in or a ut ho rit y va ria nt s, an d m is sp el lin gs . u ni qu e bi no m ia l n am e co un t i s 6 .8 m ill io n if au th or iti es a re ig no re d (d . m oz zh er in pe rs . c om m . 2 01 6) . io n zo ol og ic al n am es , e xt an t + fo ss il, a ll ra nk s 30 0, 00 040 0, 00 0? c. 3 m ill io n ye s fo r s ub se t on ly (< 30 % ) no no no n o ex ac t g en us c ou nt a va ila bl e. m an y na m es h el d m ul tip le ti m es (a ut ho rit y va ria nt s n ot d ed up lic at ed ), m is sp el lin gs n ot c or re ct ed . pa le ob io lo gy d at ab as e, 20 16 v er si on zo ol og ic al p lu s b ot an ic al na m es , a lm os t a ll fo ss il, a ll ra nk s 83 ,0 00 15 9, 00 0 ye s ye s ye s ye s ye s n o de ta ile d in fo rm at io n on w eb si te , v al ue s a re fr om 20 16 d at a du m p (a us tra lia n r es ea rc h c ou nc il 20 16 ). d et ai le d ge ol og ic a nd e nv iro nm en ta l in fo rm at io n is h el d as a va ila bl e; m an y na m es h av e fu ll bi bl io gr ap hi c ci ta tio ns . n om en cl at or zo ol og ic us zo ol og ic al n am es , e xt an t + fo ss il, g en er a an d su bg en er a on ly , c ut -o ff d at e 20 04 35 7, 00 0 0 ye s no no ye s no h ig he r t ax on om ic a llo ca tio n ty pi ca lly g iv en to ph yl um (o r c la ss / or de r) o nl y, a cc or di ng to g ro up . fo ss il in di ca tio n in cl ud ed fo r m an y bu t n ot a ll ap pl ic ab le n am es . in de x n om in um g en er ic or um b ot an ic al n am es , e xt an t + fo ss il, g en er a on ly , 1 75 320 12 (+ ?) 69 ,0 00 0 ye s ye s no ye s no h ig he r t ax on om ic a llo ca tio n ty pi ca lly g iv en to fa m ily , b ut o ut da te d in so m e ca se s. biodiversity informatics, 12, 2017, pp. 1-44 29 ip n i b ot an ic al n am es e xc lu di ng al ga e, fu ng i a nd b ry op hy te s, ex ta nt o nl y, a ll ra nk s 30 ,0 00 +? 0. 81 m ill io n? ye s ye s no ye s ( al l pr es um ed ex ta nt ) no o ve r 1 .3 m ill io n na m es a t a ll ra nk s, in cl ud in g in fr as pe ci es . g en er ic n am es w ill la rg el y ov er la p co ve ra ge o f e qu iv al en t g ro up s i n in de x n om in um g en er ic or um . it is m os t g ro up s, al l e xt an t, al l ra nk s 50 ,0 00 15 0, 00 0? 52 6, 00 0 ye s ye s ye s ye s ( al l pr es um ed ex ta nt ) no 70 8, 00 0 na m es a t a ll ra nk s a s a t d ec em be r 2 01 5, in cl ud in g 52 6, 00 0 sp ec ie s b in om ia ls (g ua la 2 01 6) . w or m s m ar in e ta xa (p lu s s om e no nm ar in e) , e xt an t + so m e fo ss il, a ll ra nk s 47 ,0 00 47 0, 00 0 ye s ye s ye s ye s ye s fo ss il ta xa a re in cl ud ed fo r s el ec te d gr ou ps o nl y. biodiversity informatics, 12, 2017, pp. 1-44 30 irmng data may not always represent the most current version of its own externally sourced content, and it is also recommended to visit those sites too for their most current content. that aspect aside, irmng aspires to provide a single point of entry to a comprehensive overview of “all life”, particularly to generic level, with added, machine-readable basic extant/fossil and habitat indicators, as well as associated original publication information, taxonomic comments and more, that may not be readily accessible elsewhere in such an integrated form. irmng as a research entity a compilation of genus names through time such as irmng offers a rich source of content for further study, such as the numbers of genera per taxonomic group published through time, the lexical character and variation of scientific names, overall contributions by specific workers, reporting the size of different taxonomic groups, ratios of valid names to synonyms once the latter are more completely annotated, compilation of lists of within-code and between-code homonyms, and much more. such topics are largely outside the scope of the present paper but may be discussed elsewhere in due course. in this respect irmng can function as a literary corpus, which is a standardized collection of text upon which specific linguistic investigations can be performed and results reported in a repeatable fashion (biber et al. 1998). a further valuable use of the collection of names assembled within irmng has been in providing a reference suite of data to support the development and extensive testing of algorithms for name matching, including both taxamatch and some subsidiary methods, as reported in more detail in rees (2014). representing as it does the current largest, deduplicated and preliminarily vetted set of genus names so far compiled across all taxonomic domains, virtually all with an associated authority and year, irmng is presently unique in offering potential insights into the nature of such names as well as the history of contributions by different scientists through time. irmng as a source of “marine” content irmng taxon names currently flagged “marine = yes” that are not presently held in worms represent a potential source of additional content to the latter system. a workflow has been commenced whereby such names within groups of interest to worms can be passed to relevant worms experts for scrutiny as needed and then added to worms holdings once approved to do so. a separate, but related issue, is to compare environmental flags as presently held in the two systems with an emphasis on upgrading these, in particular within worms. this work (recommended by the worms steering committee, dec. 2015) has commenced, and in this context the flags already present in irmng for many taxa have already proved a valuable resource to assist both the vliz data management team and worms taxonomic editors. irmng as a taxonomic name resolution service with its associated tools (name parsing and fuzzy matching) accessing a reference database of hierarchically arranged scientific names, the web “search” function provided to irmng users since 2006 can be viewed as one of the earliest implementations of a taxonomic name resolution service, which has been characterized as “a service that corrects variant and erroneous spellings, disambiguates homonyms by means of higher taxonomic filtering, and updates synonyms with reference to authoritative taxonomic sources” (boyle et al. 2013). such a service (currently in version 4.040) is also offered by the iplant taxonomic name resolution service (tnrs) described by boyle et al., which has been operational since 2011. this service offers many of the same features as irmng but differs in scope, since it is currently set up to access five databases of land plants only (i.e., extant bryophytes through angiosperms) and thus currently excludes algal, fungal, protozoan, and animal names, as well as all fossils. there are some minor conceptual differences between the iplant tnrs and irmng in that in the iplant tnrs, fuzzy matching is extended to family level and name parsing is performed by a dedicated module developed elsewhere, the global names index scientific names parser41. most significantly, in the case of the iplant tnrs, all of the reference datasets searched exist as separate external resources whose curation and maintenance are not the responsibility of the iplant development team. irmng sustainability and future activities as with any biodiversity project intended to last beyond just a few years, succession planning 40 http://tnrs.iplantcollaborative.org/ 41 http://gni.globalnames.org/parsers/new biodiversity informatics, 12, 2017, pp. 1-44 31 and the sustainability of the project are important considerations, especially as the project lifetime extends (costello et al. 2014). initially the project was developed to address the taxonomic and associated trait requirements of obis and progressed in accordance with that project’s direction (e.g., unesco 2012), as a contribution from the obis australia regional node which was at that time (and is still) located at csiro in australia. when it was realised that this contribution was potentially coming to an end, a succession plan was introduced which involved transferring the data and internal relationships to the servers at vliz where they could be sustained by that institution into the future in tandem with that organisation’s pre-existing commitment to maintain worms and associated taxonomic databases42. from 2016, the vliz data management team (dmt) have assumed responsibility for ongoing maintenance and development of both irmng and the it platform that supports it, addressing most of the concerns expressed in costello et al. (2014) with the strictly it-related costs essentially already covered under the general operation of existing vliz data systems and therefore not requiring separate resourcing. with this move, irmng has also evolved from a system largely developed and maintained principally by a single author to one supported by a larger team and with the potential for distributed content maintenance and enhancement in the future, again a key recommendation of the combined expertise represented by the authors contributing to costello et al. (2014). in its new location as part of the vliz “family” of taxonomic data systems, a number of new as well as continuing activities are envisaged. first, the public web presence and search-based functions are now available at the new location, modified to reflect and include functionality already developed for worms and associated data systems that also run on the aphia platform. second, export routines for irmng data are being continued and new versions will be made available to present and also potential new irmng clients. third, irmng data will be made available to external automated clients via soap/wsdl-based web services as already offered for other vliz-hosted databases (soap: simple object access 42 http://www.marinespecies.org/about.php protocol; wsdl: web services description language; for additional explanation see e.g., weerawarana et al. 2005). fourth, co-location of the irmng and worms/obis data systems is facilitating closer linkages between these projects, as well as with the taxonomic backbone for the european lifewatch project which is also being developed at vliz (dekeyzer et al. 2014, see also the lifewatch page at vliz43 and the lifewatch “taxonomic information” page44). fifth, the remote editing interfaces already developed for use by worms editors have been adapted for use with irmng and are already operational and in use by relevant accredited users. lastly, vliz has also expressed a desire to develop a network of remote editors for irmng along the lines of that already in place for worms, a process which is anticipated to be commenced over the period 2017-2018. it is also intended to avoid duplication of effort so far as is possible, in other words if a worms editor already exists for a predominantly marine group, worms would be the natural continued vehicle to hold this information from which it can be ported to irmng in due course without the requirement for additional data entry. on top of additional functionality and linkages for the new version of irmng, the requirement for ongoing maintenance and content enhancement will continue. this can conceptually be divided into three areas: (1) ongoing addition of new names as these are published—the previously noted approx. 2,500 names annually at generic rank (plus new families as well), in addition to between 15,000 and 20,000 new species names per year, to the extent that the latter are considered desirable to be held; (2) “catch-up” acquisition of legacy names not yet held—particularly additional species if these are desired—using available sources such as more recent editions of the catalogue of life, paleobiology database, and others; and (3) upgrade of content already held, with respect to the present data limitations as discussed earlier. one aspect which should receive particular attention is a review of the present irmng classification above family level, which (with the exception of specific groups as detailed in appendix 1) retains some inconsistencies on 43 http://www.vliz.be/en/lifewatch 44 http://www.lifewatch.be/en/taxonomic-information biodiversity informatics, 12, 2017, pp. 1-44 32 account of being imported from a range of different sources over time, in many cases because no “standard” higher taxonomy existed at that time. with the emergence of some recent standardized treatments such as zhang (2011, 2013), ruggiero (2014), and ruggiero et al. (2015), as well as the ongoing contributions by the angiosperm phylogeny group already incorporated, the opportunity exists to review and standardize the irmng classification to a higher degree than previously implemented, and thereby facilitate better interoperability and data exchange with other systems that utilize the same classifications. to date, prioritization and resourcing of the various areas of activity indicated above have been the responsibility of the first author in the main, aided by in-kind contributions from csiro plus grants from key irmng clients including obis, gbif and the atlas of living australia. from 2016 onwards, future directions and resourcing for continued irmng development will be determined by the vliz data management team (which includes the remaining authors of this paper) in liaison with their network of other interested parties including obis, worms and its editors, catalogue of life and more. concluding remarks when irmng was commenced in 2006, its main role was to provide a means for obis and related data aggregation projects to manage incoming taxonomic data in a consistent manner, and to distinguish marine from nonmarine, and extant from fossil taxa. as an accessory function, the system has also served as an effective taxonomic name resolution service for these projects and a range of other users over the succeeding ten years. during that time, there has been a growth in the availability and/or completeness of other systems dealing with species data in particular, such as the catalogue of life annual editions, the paleobiology database, the world register of marine species, and the plant list, which are currently more complete than irmng for species although less so for genera. the existence of “super aggregators” such as the global biodiversity information facility, the global names index and global names resolver, the encyclopedia of life, and the open tree of life also permits such projects to claim more comprehensive coverage than irmng in the simple metric of “numbers of names held”, although a number of these ingest irmng data and therefore rely to some degree on irmng to continue to acquire and supply them with new content. in its location at vliz from 2016 onwards, and with new arrangements for governance, data sharing and distributed editing, the method of operation for irmng is certain to evolve and change in some respects as well as responding to changes in the overall biodiversity informatics landscape. this review of the first ten years of irmng’s operation serves as a document of processes and activity in content building to date and provides a baseline against which future progress can be measured, and in addition may potentially provide some ideas and discussion points of value to other biodiversity projects that operate in a similar area, both now or in the future. author contributions tony rees established the original irmng data compilation, designed, constructed, and populated the initial irmng data tables at csiro and created the associated web application for data entry and retrieval. leen vandepitte collaborated with tony to manage the transfer of irmng data to vliz, while wim decock and bart vanhoorne carried out the technical operations of data mapping between the csiro and aphia systems, migration to the vliz servers and construction of the web interfaces for the new version of irmng hosted at vliz, and serve as administrators for the system from 2016 onwards. all authors contributed to pre-release testing, release, and ongoing improvement of the new version of irmng at vliz during 2016. tony rees wrote the initial version of this manuscript which was then developed further with the benefit of contributions from the other authors. acknowledgements a project of the scale of irmng could not be attempted within any reasonable time frame without leveraging the activities of previous workers who have already compiled and checked vast quantities of nomenclatural and taxonomic data used as input to irmng, much of it (particularly at specific level) without further checking at this time. we thank, among others, nicolas bailly and deng palomares (obis, incofish and sealifebase, phillippines), sheila brands (systema naturae 2000 project, the netherlands), ely wallis (museum victoria, australia), the late frank bisby (catalogue of life, united kingdom), ellen farr (index biodiversity informatics, 12, 2017, pp. 1-44 33 nominum genericorum, u.s.a.), edward vanden berghe (formerly worms/vliz), david remsen (formerly gbif), dennis gordon (niwa, new zealand), john wiersema (grin, u.s.a.), joel hallan (hallan biology catalog, u.s.a.), paul kirk (index fungorum, u.k.) and rod page (bionames, u.k.) for generous provision of content from data systems under their control, and helen morgan, anna povey, gary poore, robin wilson, and stevie davenport (australia) for additional contributions to irmng data holdings. dmitry mozzherin (u.s.a.) provided totals of species names (binomials) in the global names index, and wouter addink (sp2000 secretariat, netherlands) provided counts for genera in the 2016 catalogue of life. thanks are also due to obis for providing the scientific context within which the system was initially developed and for financial support to attend relevant meetings. miroslaw ryba and steven edgar provided excellent technical support and advice as required at csiro, australia. financial and inkind support for portions of this work were provided at different times by csiro marine and atmospheric research, obis australia, the atlas of living australia, the global biodiversity information facility and the flanders marine institute; the transfer and hosting of irmng to vliz was supported by lifewatch belgium, part of the e-science european lifewatch infrastructure for biodiversity and ecosystem research. at vliz, we thank francisco hernandez for his willingness to accept irmng as a component of the vliz database portfolio and for committing to ensure its persistence and ongoing support within that institute’s data management activities. lastly, we thank two anonymous referees for their useful comments which have helped to improve this manuscript. literature cited agassiz, l. 1848. nomenclatoris zoologici index universalis: continens nomina systematica classium, ordinum, familiarum et generum animalium omnium, tam viventium quam fossilium, secundum ordinem alphabeticum unicum disposita, adjectis homonymiis plantarum, nec non variis adnotationibus et emendationibus. soloduri: sumptibus jent et gassman. allkin, r., r.j. white, and p.j. winfield. 1992. handling the taxonomic structure of biological data. mathl. comput. modelling 16(6/7):1-9. alroy, j. 2002. how many named species are valid? proc. natl. acad. sci. usa 99:3706-3711. andrews, h.n. 1970. index of generic names of fossil plants, 1820-1965. geological survey bulletin 1300. us government printing office, washington; plus supplements (geological survey bulletin 1396: index of generic names of fossil plants 1966-1973 by a. m. blazer, 1975, and geological survey bulletin 1517: index of generic names of fossil plants 1974-1978 by a. d. watt, 1982). angiosperm phylogeny group. 2009. an update of the angiosperm phylogeny group classification for the orders and families of flowering plants: apg iii. bot. j. linn. soc. 161:105-121. angiosperm phylogeny group. 2016. an update of the angiosperm phylogeny group classification for the orders and families of flowering plants: apg iv. bot. j. linn. soc. 181:1-20. appeltans, w., s.t. ahyong, g. anderson, m.v. angel, t. artois, et al. 2012. the magnitude of global marine species diversity. curr. biol. 22:2189-2202. australian research council. 2016. the paleobiology database (july 5 2016 version), hosted by the gbif secretariat. online web resource, available at http://www.gbif.org/dataset/c33ce2f2-c3cc43a5-a380-fe4526d63650. benton, m.j., ed. 1993. the fossil record 2. chapman & hall. london. benton, m.j. 2008. how to find a dinosaur, and the role of synonymy in biodiversity studies. paleobiology 34:516-533. biber, d., s. conrad, and r. reppen. 1998. corpus linguistics—investigating language structure and use. cambridge university press, cambridge. bisby, f.a., m.a. ruggiero, y.r. roskov, m. cachuela-palacio, s.w. kimani, p.m. kirk, a. soulier-perkins, and j. van hertum, eds. 2006. species 2000 & itis catalogue of life: 2006 annual checklist. cd-rom; species 2000, reading. bouchard, p., y. bousquet, a. davies, m. alonsozarazaga, j. lawrence, c. lyal, a. newton, c. reid, m. schmitt, a. slipinski, and a. smith. 2011. family-group names in coleoptera (insecta). zookeys. 88:1-972 bouchet, p., and j.-p. rocroi. 2005. classification and nomenclator of gastropod families. malacologia 47:1-397. boyle, b., n. hopkins, z. lu, j.a. raygoza garay, d. mozzherin, t. rees, n. matasci, m.l. narro, w.h. piel, s.j. mckay, s. lowry, c. freeland, r.k. peet, and b.j. enquist. 2013. the taxonomic name resolution service: an online tool for automated standardization of plant names. bmc bioinformatics 14:16. chapman, a.d. 2009. numbers of living species in australia and the world. australian government, biodiversity informatics, 12, 2017, pp. 1-44 34 department of the environment, canberra, 2nd edition. costello, m.j. 2000. developing species information systems: the european register of marine species (erms). oceanography 13(3):48-55. costello, m.j., w. appeltans, n. bailly, w.g. berendsohn, y. de jong, m. edwards, r. froese, f. huettmann, w. los, j. mees, h. segers, and f.a. bisby. 2014. strategies for the sustainability of online open-access biodiversity databases. biol. conserv. 173:155-165. costello, m.j., c. emblow, and r.j. white, eds. 2001. european register of marine species: a check-list of the marine species in europe and a bibliography of guides to their identification. collection patrimoines naturels, 50. muséum national d'histoire naturelle, paris. de grave, s., n.d. pentcheff, s.t. ahyong, t.-y. chan, k.a. crandall, p.c. dworschak, d.l. felder, r.m. feldmann, c.h.j.m. fransen, l.y.d. goulding, r. lemaitre, m.e.y. low, j.w. martin, p.k.l. ng, c.e. schweitzer, s.h. tan, d. tshudy, and r. wetzer. 2009. a classification of living and fossil genera of decapod crustaceans. raffles bull. zool., suppl. 21:1-109. dekeyzer, s., l. vandepitte, s. claus, k. deneudt, and f. hernandez. 2014. the lifewatch taxonomic backbone: supporting the marine biodiversity and ecosystem functioning research community [abstract]. p. 106 in a variety of interactions in the marine environment. abstracts volume from 49th european marine biology symposium, september 8-12, 2014. st. petersburg, russia. döring, m. 2017. gbif backbone—february 2017 update. gbif developer blog, available at http://gbif.blogspot.com.au/2017/02/gbifbackbone-february-2017-update.html. eschmeyer, w.n. 2000-current. catalog of fishes online database. online web resource, available at http://researcharchive.calacademy.org/research/ic hthyology/catalog/fishcatmain.asp. euzéby, j.p., ed. plus successors. 1997-current. list of prokaryotic names with standing in nomenclature (originally: list of bacterial names with standing in nomenclature). online web resource, available at http://www.bacterio.net. edited by j. p. euzéby, 1997-2013, a. c. parte from 2013. farr, e.r., j.a. leussink, and f.a. stafleu, eds. 1979. index nominum genericorum (plantarum). scheltema & holkema, utrecht, and w. junk, the hague. 3 volumes. gordon, d.p., ed. 2009-2011. new zealand inventory of biodiversity. volume 1: kingdom animalia: radiata, lopotrochozoa, deuterostomia; volume 2: kingdom animalia: chaetognatha, ecdysozoa, ichnofossils; volume 3: kingdoms bacteria, protozoa, chromista, plantae, fungi. canterbury university press, canterbury, new zealand. guala, g.f. 2016. the importance of species name synonyms in literature searches. plos one 11(9):e0162648. hinchliff, c.e., s.a. smith, j.f. allman, j.g. burleigh, r. chaudhary, l.m. coghill, k.a. crandall, j. deng, b.t. drew, r. gazis, k. gude, d.s. hibbett, l.a. katz, h.d. laughinghouse, e.j. mctavish, p.e. midford, c.l. owen, r.h. ree, j.a. rees, d.e. soltis, t. williams, and k.a. cranston. 2015. synthesis of phylogeny and taxonomy into a comprehensive tree of life. proc. nat. acad. sci. usa 112:12764-12769. hooker, j.d., and b.d. jackson. 1895. index kewensis plantarum phanerogamarum nomina et synonyma omnium generum et specierum a linnaeo usque ad annum mdccclxxxv… oxford university press, oxford. international commission on zoological nomenclature. 1999. international code of zoological nomenclature, 4th edition. international trust for zoological nomenclature, london. international commission on zoological nomenclature. 2017. open cases. available online at http://iczn.org/content/list-open-cases, date of query 25 may 2017. international committee on taxonomy of viruses (ictv). 1971-current. report(s) on virus classification: includes 7th report (1999), 8th report (2005), 9th report (2009). kirk, p., compiler. 2001-current. index fungorum. online web resource, available at http://www.indexfungorum.org/. lapage, s.p., p.h.a. sneath, e.f. lessel, v.b.d. skerman, h.p.r. seeliger, and w.a. clark, eds. 1992. international code of nomenclature of bacteria: bacteriological code, 1990 revision. asm press, washington (dc), u.s.a. martin, j.w., and g.e. davis. 2001. an updated classification of the recent crustacea. science series 39, natural history museum of los angeles county. mcneill, j., f.r. barrie, w.r. buck, v. demoulin, w. greuter, d.l. hawksworth, p.s. herendeen, s. knapp, k. marhold, j. prado, w.f. prud’homme van reine, g.f. smith, j.h. wiersema, and n.j. turland. 2012. international code of nomenclature for algae, fungi, and plants (melbourne code). regnum vegetabile 154. koeltz scientific books. mullins, g.l., ed. 2007. the phytopal taxonomic database (taxon list), version 02 february 2007. 1132 pp. originally downloaded from http://www.le.ac.uk/geology/glm2/phytopal/taxa. pdf (no longer available from that source). neave, s.a., ed. 1939-1996. nomenclator zoologicus. a list of the names of genera and subgenera in zoology from the tenth edition of linnaeus, 1758, biodiversity informatics, 12, 2017, pp. 1-44 35 to the end of 1935 (plus succeeding volumes to end 1994). zoological society of london, london. 9 volumes. ng, p.k.l., and m.e.y. low. 2010. on the generic nomenclature of nine brachyuran names, with four replacement names and two nomina protecta (crustacea: decapoda). zootaxa 2489:34-46. özdikmen, h., and demir, h. 2011. case 3540 amphiporidae rukhin, 1938 (porifera, stromatoporata, amphiporida): proposed emendation to amphiporaidae to remove homonymy with amphiporidae mcintosh, 1873 (nemertea, hoplonemertea). bull. zool. nomenclature 68:167-169. parker, c.t., g.m. garrity, and b.j. tindall. 2015. international code of nomenclature of prokaryotes. int. j. syst. evol. microbiol. november 2015. parker, s.p., ed. 1982. synopsis and classification of living organisms. mcgraw-hill book company, new york. 2 volumes. patterson, d.j., j. cooper, p.m. kirk, r.l. pyle, and d.p. remsen. 2010. names are key to the big new biology. trends ecol. evol. 25:686-691. patterson, d.j., w. egloff, d. agosti, d. eades, n. franz, g. hagedorn, j.a. rees, and d.p. remsen. 2014. scientific names of organisms: attribution, rights, and licensing. bmc res. notes 7:79. patterson, d., d. mozzherin, d.p. shorthouse, and a. thessen. 2016. challenges with using names to link digital biodiversity information. biodiv. data j. 4:e8080. prothero, d.r. 2013. bringing fossils to life: an introduction to paleobiology. 3rd edition. columbia university press, new york. raup, d.m. 1986. biological extinction in earth history. science 231:1528-1533. rees, j.a., and cranston, k. 2017. automated assembly of a reference taxonomy for phylogenetic data synthesis. biodiversity data journal 5:e12581. rees, t., 2008a. applications of fuzzy (approximate string) matching in taxonomic database searches, with an example multi-tiered approach. p. 11-13 in worcester, t., l. bajona, and b. branton, eds. proceedings of a conference on ocean biodiversity informatics (2-4 october 2007). bedford institute of oceanography, dartmouth, nova scotia. rees, t. 2008b. irmng—the interim register of marine and nonmarine genera [abstract]. p. 72 in weitzman, a. and l. belbin, eds. proceedings of tdwg (2008), fremantle, australia. biodiversity information standards (tdwg) and missouri botanical garden. rees, t. 2014. taxamatch, an algorithm for near (‘fuzzy’) matching of scientific names in taxonomic databases. plos one 9(9):e107510. rees, t., and y. zhang. 2007. evolving concepts in the architecture and functionality of obis, the ocean biogeographic information system. pp. 167-176 in vanden berghe, e. et al., eds. proceedings ocean biodiversity informatics: international conference on marine biodiversity data management, hamburg, germany 29 november to 1 december, 2004. vliz special publication 37. ruggiero, m.a. 2014. families of all living organisms, version 2.0.a.15, (4/26/14). expert solutions international, llc, reston, va, usa. tabulated data file (as “ggi family data”) available via the global genome initiative (ggi) knowledge portal, http://ggi.eol.org/downloads. ruggiero, m.a., d.p. gordon, t.m. orrell, n. bailly, t. bourgoin, r.c. brusca, t. cavalier-smith, m.d. guiry, and p.m. kirk, 2015. a higher level classification of all living organisms. plos one 10(4):e0119248. schultz, p.w., and m. huber. 2013. revision of the worldwide recent pinnidae and some remarks of fossil european pinnidae. acta conch. 13:1-164. sepkoski. j.j., jr. 2002. a compendium of fossil marine animal genera. bull. amer. paleontol. 363:1-560. sherborn, c.d. 1902-1933. index animalium; sive, index nominum quae ab a. d. mdcclviii generibus et speciebus animalium imposita sunt, societatibus eruditorum adiuvantibus. cambridge, u.k. unesco. 2012. iode steering group for obis (sg-obis), second session, 19-21 november 2012. reports of meetings of experts and equivalent bodies, (english), unesco. vandepitte, l., b. vanhoorne, w. decock, s. dekeyzer, a.t. verbeeck, l. bovit, f. hernandez, and j. mees. 2015. how aphia—the platform behind several online and taxonomically oriented databases—can serve both the taxonomic community and the field of biodiversity informatics. j. mar. sci. eng. 3:1448-1473. weerawarana, s., f. curbera, f. leymann, t. storey, and d.f. ferguson. 2005. web services platform architecture: soap, wsdl, ws-policy, wsaddressing, ws-bpel, ws-reliable messaging and more. prentice hall ptr, upper saddle river, nj, usa. willis, j.c. 1973. a dictionary of the flowering plants and ferns. cambridge university press, cambridge, u.k., 8th edition. worms editorial board. 2007-current. world register of marine species. available from http://www.marinespecies.org at vliz. zhang, z.-q., ed. 2011. animal biodiversity: an outline of higher-level classification and survey of taxonomic richness. zootaxa 3148:1-237. zhang, z.-q., ed. 2013. animal biodiversity: an outline of higher-level classification and survey of taxonomic richness (addenda 2013). zootaxa 3703:1-82. biodiversity informatics, 12, 2017, pp. 1-44 36 zeidler, w. 2017. validation of the replacement name eusceliotes stebbing, 1888 for the pelagic hyperiidean amphipod genus euscelus claus, 1879 (crustacea: amphipoda: hyperiidea: parascelidae), preoccupied by euscelus schoenherr, 1833 (insecta: coleoptera: attelabidae). zootaxa 4254:377-378. biodiversity informatics, 12, 2017, pp. 1-44 37 a pp en d ix 1 : s ig n if ic a n t so u r c es u se d in ir m n g d a ta c o m pi la ti o n , 2 00 620 16 d at e ad de d so ur ce a nd w eb lo ca tio n (a s a pp lic ab le ) c on te nt a dd ed r em ar ks 1/ 08 /2 00 6 pr in t w or k: “ sy no ps is a nd c la ss ifi ca tio n of l iv in g o rg an is m s” (p ar ke r 1 98 2) , ta xo no m ic in de x se ct io n di gi tiz ed b y d . pa lo m ar es a nd te am (s ea li fe b as e) a s a co nt rib ut io n to ir m n g 6, 90 0 ex ta nt fa m ili es c on ta in s v al id fa m ily n am es o nl y, w ith n o as so ci at ed au th or iti es . a sm al l n um be r o f g ro up s n ot in cl ud ed . 20 /0 9/ 20 06 sy st em a n at ur ae 2 00 0 (s n 20 00 ) co ur te sy d r. s. b ra nd s, co m pi le r ( 20 06 ve rs io n) 45 23 4 ad di tio na l f am ili es , 1 11 ,0 00 ge nu s n am es m ix o f d at a fr om a ra ng e of o rig in al so ur ce s, in cl ud in g 62 ,0 00 “ ve rif ie d” a nd 4 9, 00 0 “u nv er ifi ed ” ge nu s n am es , m os t w ith p re se nt fa m ily a llo ca tio n. s pe ci es n am es n ot re qu es te d. 20 /0 9/ 20 06 c at al og ue o f l ife , 2 00 6 ed iti on (b is by e t al . 2 00 6) 2, 10 0 ad di tio na l f am ili es , 3 6, 00 0 ad di tio na l e xt an t g en us n am es p lu s 1, 28 2, 00 0 ex ta nt sp ec ie s n am es in cl ud in g sy no ny m s ( in cl ud in g 12 0, 00 0 la te r d ep re ca te d as du pl ic at es ) in cl ud es n um er ou s g ro up s, vi ru se s t o ve rte br at es a lth ou gh ho ld in gs in co m pl et e es pe ci al ly fo r i ns ec ts . f am ily al lo ca tio ns in cl ud ed ; n o au th or iti es g iv en fo r g en us n am es , la tte r n ee d to b e ad de d fr om o th er so ur ce s a s a va ila bl e. 20 /0 9/ 20 06 p. b oc k b ry oz oa h om e pa ge (2 00 6 ve rs io n) 46 13 0 ad di tio na l f am ili es 17 /1 0/ 20 06 m us eu m v ic to ria k em u da ta ba se co ur te sy d r. e. w al lis , m us eu m v ic to ria 47 9, 30 0 ad di tio na l e xt an t+ fo ss il ge nu s na m es , p lu s 5 6, 30 0 sp ec ie s n am es in cl ud es b ot h ex ta nt a nd fo ss il ge nu s p lu s s pe ci es n am es . 6/ 11 /2 00 6 m ar tin & d av is 2 00 1 c ru st ac ea cl as si fic at io n (m ar tin a nd d av is 2 00 1) 13 0 ad di tio na l f am ili es 21 /1 1/ 20 06 pr in t w or k: “ th e fo ss il r ec or d 2” (b en to n 19 93 ) a s r ep re se nt ed b y de riv ed sp re ad sh ee t a va ila bl e vi a th e fo ss il r ec or d 2 w eb si te 48 3, 20 0 (o f 5 ,6 00 ) a dd iti on al , l ar ge ly fo ss il fa m ili es ex ta nt /fo ss il fla gs a ls o up gr ad ed fo r f am ili es a lre ad y he ld , as n ee de d 45 h ttp :// sn 20 00 .ta xo no m y. nl /, se e al so h ttp :// ta xo no m ic on .ta xo no m y. nl / 46 h ttp :// w w w .b ry oz oa .n et / 47 o nl in e ve rs io n su bs eq ue nt ly a va ila bl e at h ttp :// co lle ct io ns .m us eu m vi ct or ia .c om .a u/ 48 h ttp :// pa la eo .g ly .b ris .a c. uk /fo ss ilr ec or d2 /fo ss ilr ec or d/ biodiversity informatics, 12, 2017, pp. 1-44 38 7/ 12 /2 00 6 a ph ia d at ab as e (2 00 6 ve rs io n) , c ou rte sy d r. e. v an de n b er gh , v li z 11 9 ad di tio na l f am ili es , 3 ,3 00 ad di tio na l g en us n am es a nd 4 5, 00 0 ad di tio na l s pe ci es n am es in cl ud es th e eu ro pe an r eg is te r o f m ar in e sp ec ie s ( er m s) pl us a n um be r o f o th er re gi on al d at a co m pi la tio ns . 6/ 02 /2 00 7 se pk os ki (2 00 2) c om pi la tio n of m ar in e fo ss il an im al g en er a, v ia a p ub lic ly av ai la bl e on lin e da ta su m m ar y49 27 ,0 00 a dd iti on al m ar in e fo ss il an im al g en us n am es n o au th or iti es g iv en ; g en er a cl as si fie d to o rd er o nl y. se pk os ki fo ss il fla g an d ge ol og ic ra ng e ad de d fo r a ll in cl ud ed g en us n am es in cl ud in g th e su bs et a lre ad y he ld . 29 /0 5/ 20 07 in de x n om in um g en er ic or um (2 00 7 ve rs io n) , c ou rte sy d r. e. f ar r, sm ith so ni an in st itu tio n (o nl in e ve rs io n at 50 ) 40 1 ad di tio na l f am ili es , 3 6, 00 0 ad di tio na l e xt an t+ fo ss il bo ta ni ca l ge nu s n am es g en us n am es a re a llo ca te d to fa m ily in m os t c as es , t he la tte r su bs eq ue nt ly a dj us te d us in g ot he r s ou rc es e .g ., a pg ii i a nd g r in . 11 /0 6/ 20 07 th e ph yt op al t ax on om ic d at ab as e, f eb 20 07 v er si on (m ul lin s 2 00 7) 40 0 ad di tio na l g en er a, 5 ,5 00 sp ec ie s na m es u se fu l s ou rc e fo r a cr ita rc h ge ne ra a nd sp ec ie s ( al l f os si l). 23 /0 8/ 20 07 a us tra lia n fa un al d ire ct or y (2 00 7 ve rs io n) 51 19 2 ad di tio na l f am ili es , 9 ,8 00 ad di tio na l g en us n am es , 5 5, 00 0 ad di tio na l s pe ci es n am es 20 /0 9/ 20 07 sp ec ie s 2 00 0 n ew z ea la nd te xt fi le (p re pu bl ic at io n ve rs io n) , c ou rte sy d r. d . g or do n, n iw a , n ew z ea la nd (s ub se qu en tly p ub lis he d as g or do n 20 09 20 11 ) 54 a dd iti on al fa m ili es , 1 ,8 00 ad di tio na l g en us n am es , 1 0, 60 0 ad di tio na l s pe ci es n am es 24 /0 6/ 20 08 eu zé by , l is t o f p ro ka ry ot ic n am es w ith st an di ng in n om en cl at ur e (l ps n )52 , 20 08 v er si on 77 a dd iti on al fa m ili es , 4 50 a dd iti on al ge nu s n am es sp ec ie s n am es n ot u pl oa de d (m an y/ m os t w ill b e in c ol 20 06 ). 14 /0 4/ 20 09 n om en cl at or z oo lo gi cu s, di gi ta l v er si on as c om pi le d by th e ub io p ro je ct 53 39 4 ad di tio na l f am ili es , 2 05 ,0 00 ad di tio na l a ni m al g en us /s ub ge nu s na m es n am es h av e au th or iti es a nd m ic ro ci ta tio ns , n o fa m ili es . c ite d au th or s, pu bl ic at io n in fo rm at io n al so u se d to u pg ra de c. 1 20 ,0 00 g en us n am es a lre ad y he ld . d at a w er e pr epr oc es se d pr io r t o up lo ad to in tro du ce c or re ct io ns a s n ot ed , se pa ra te n om in a nu da fr om c ite d fir st in st an ce o f v al id pu bl ic at io n, a nd re m ov e so m e o c r e rr or s. 49 h ttp :// st ra ta .g eo lo gy .w is c. ed u/ ja ck / 50 h ttp :// bo ta ny .si .e du /in g/ 51 h ttp :// w w w .e nv iro nm en t.g ov .a u/ bi od iv er si ty /a br s/ on lin ere so ur ce s/ fa un a/ af d/ 52 h ttp :// w w w .b ac te rio .n et / 53 h ttp :// ui o. m bl .e du /n om en cl at or zo ol og ic us / biodiversity informatics, 12, 2017, pp. 1-44 39 12 /0 5/ 20 09 es ch m ey er c at al og o f f is he s54 , o nl in e 20 09 v er si on 80 0 ad di tio na l g en us n am es fa m ily a llo ca tio ns fo r n am es a lre ad y he ld u pd at ed a s ne ed ed , g en us sy no ny m st at us a ls o ch ec ke d. s pe ci es n am es no t u pl oa de d (m an y/ m os t w ill b e in c ol 2 00 6) . 13 /0 7/ 20 09 in de x fu ng or um 55 , o nl in e 20 09 v er si on 17 1 ad di tio na l f am ili es , 1 ,8 00 ad di tio na l g en us n am es fa m ily a llo ca tio ns fo r n am es a lre ad y he ld u pd at ed a s ne ed ed , g en us sy no ny m st at us a ls o ch ec ke d. s pe ci es n am es no t u pl oa de d (m an y/ m os t w ill b e in c ol 2 00 6) . 17 /0 2/ 20 10 c la ss ifi ca tio n of li vi ng a nd fo ss il ge ne ra of d ec ap od c ru st ac ea ns (d e g ra ve e t a l. 20 09 ) 75 a dd iti on al fa m ili es , 3 30 a dd iti on al ge nu s n am es fa m ily a llo ca tio ns fo r n am es a lre ad y he ld u pd at ed a s ne ed ed . 14 /0 4/ 20 10 g b if d at a— va rio us in cl ud in g pa le od b , y al e u ni ve rs ity p ea bo dy m us eu m , ot he rs ; d at a co ur te sy d . r em se n, g b if . 1, 10 0 ad di tio na l f am ili es 22 /0 4/ 20 10 c la ss ifi ca tio n an d no m en cl at or o f ga st ro po d fa m ili es (b ou ch et a nd r oc ro i 20 05 ) 92 a dd iti on al fa m ili es 6/ 12 /2 01 0 a pg ii i c la ss ifi ca tio n (a ng io sp er m ph yl og en y g ro up 2 00 9) 39 a dd iti on al fa m ili es se le ct ed a ng io sp er m fa m ili es m er ge d as p er th is tr ea tm en t. 19 /0 1/ 20 11 th e pl an t l is t, ve rs io n 1. 056 1, 40 0 ad di tio na l g en us n am es n ew g en us n am es c he ck ed a ga in st o th er so ur ce s ( ip n i, tr op ic os ) f or a ut ho rit ie s ( no t i nc lu de d in t he p la nt l is t) an d as so ci at ed m ic ro ci ta tio ns . s pe ci es n am es n ot u pl oa de d. 29 /0 4/ 20 11 a ni m al g en us n am es a nd o rig in al w or ks as re po rte d in se ar ch o f “ zo ol og ic al r ec or d” 57 fo r n ew ly p ub lis he d na m es , 20 05 -2 01 1, p lu s s up pl em en ta ry w eb ch ec ks fo r m is si ng d ia cr iti cs a nd c or re ct au th or ity c ita tio n 4, 00 0 ad di tio na l g en us n am es b ib lio gr ap hi c de ta ils fo r a rti cl es a ls o up lo ad ed , t o ac co m pa ny n ew g en us n am es . 9/ 05 /2 01 1 n ew fo ss il pl an t g en er a as p ub lis he d in w or ks li st ed in th e k an sa s u ni ve rs ity on lin e b ib lio gr ap hy o f p al eo bo ta ny 58 46 3 ad di tio na l g en us n am es b ib lio gr ap hi c de ta ils fo r a rti cl es a ls o up lo ad ed , t o ac co m pa ny n ew g en us n am es . 54 h ttp :// w w w .c al ac ad em y. or g/ sc ie nt is ts /p ro je ct s/ ca ta lo gof -f is he s 55 h ttp :// w w w .in de xf un go ru m .o rg / 56 o rig in al ly a t h ttp :// w w w .th ep la nt lis t.o rg /, ar ch iv ed c op y at h ttp :// w w w .th ep la nt lis t.o rg /1 /. 57 h ttp :// w ok in fo .c om /p ro du ct s_ to ol s/ sp ec ia liz ed /z r/ 58 h ttp :// pa le ob ot an y. bi o. ku .e du /b ib lio o fp al eo .h tm biodiversity informatics, 12, 2017, pp. 1-44 40 3/ 8/ 20 11 ic tv v iru s d at ab as e (2 01 1 ed iti on ), do w nl oa de d fr om 59 90 a dd iti on al g en us n am es 7/ 10 /2 01 1 g er m pl as m r es ou rc es in fo rm at io n n et w or k (g r in ) t ax on om y fo r p la nt s (s ep 2 01 1 ve rs io n) 60 1, 60 0 ad di tio na l g en us n am es a ll ex ta nt h ig he r p la nt g en er a up da te d to fo llo w g r in fa m ily a llo ca tio ns a s w el l a s c ur re nt n am e/ sy no ny m st at us . 4/ 05 /2 01 2 h al la n b io lo gy c at al og (o nl in e ve rs io n as a t m ar 2 01 2) 61 72 8 ad di tio na l f am ili es , 9 ,4 00 ad di tio na l g en us n am es , 2 08 ,0 00 ad di tio na l s pe ci es n am es 26 /0 6/ 20 12 es ch m ey er o nl in e (j un 2 01 2 ve rs io n) 65 0 ad di tio na l g en us n am es u pd at e on ly (f or n am es n ot in 2 00 9 ve rs io n) , g en er a on ly . 29 /8 /2 01 2 in de x n om in um g en er ic or um (j ul 2 01 2 ve rs io n) 70 0 ad di tio na l g en us n am es u pd at e on ly (f or n am es n ot in 2 00 7 ve rs io n) . 9/ 04 /2 01 3 in de x fu ng or um (u pd at e su pp lie d co ur te sy p . k irk ) 60 0 ad di tio na l g en us n am es u pd at e on ly (f or n am es n ot in 2 00 9 ve rs io n) , g en er a on ly . 3/ 06 /2 01 3 w or m s62 (m ar 2 01 3 ve rs io n) 1, 15 0 ad di tio na l f am ili es , 3 ,4 00 ad di tio na l g en er a, a nd 2 26 ,0 00 ad di tio na l s pe ci es n am es 6/ 05 /2 01 4 th e pl an t l is t v .1 .1 63 10 a dd iti on al fa m ili es , 1 00 a dd iti on al ge ne ra sp ec ie s n am es n ot u pl oa de d. 8/ 06 /2 01 5 io n d at a du m p vi a b io n am es (d at a to 20 12 )64 , p lu s s up pl em en ta ry w eb c he ck s as re qu ire d 9, 60 0 ad di tio na l g en us n am es pa rs ed fo r g en us a nd /o r s ub ge nu s n am es n ot p re vi ou sl y he ld , d at e ra ng e 19 90 -2 01 2. t itl es o f o rig in al p ub lic at io ns al so u pl oa de d, to a cc om pa ny re le va nt n ew g en us n am e re co rd s. 6/ 01 /2 01 6 io n d at a du m p vi a b io n am es (d at a to 20 14 )65 , p lu s s up pl em en ta ry w eb c he ck s as re qu ire d 5, 70 0 ad di tio na l g en us n am es u pd at e on ly (f or n ew n am es si nc e 20 12 v er si on ), ge ne ra on ly . 13 /0 9/ 20 16 r ev ie w o f a su bs et o f g en er a pr es en tly la ck in g an e xt an t/f os si l f la g ex ta nt /fo ss il st at us a dd ed fo r 1 6, 00 0 ge nu s n am es p re vi ou sl y la ck in g th is in fo rm at io n a ro un d 1, 00 0 na m es re se ar ch ed in di vi du al ly , o th er s i nf er re d fr om a ut ho r/p ub lic at io n de ta ils . 59 h ttp :// w w w .ic tv on lin e. or g/ (a s a t 2 01 1) 60 d at a do w nl oa de d fr om ft p: //f tp .a rs -g rin .g ov /p ub /m is c/ ta x, a pp ar en tly n o lo ng er a va ila bl e in th at fo rm . c ur re nt v er si on se ar ch ab le o nl in e vi a ht tp s: //n pg sw eb .a rs -g rin .g ov /g rin gl ob al /ta xo n/ ta xo no m ys ea rc h. as px . 61 h ttp :// bu g. ta m u. ed u/ re se ar ch /c ol le ct io n/ ha lla n/ 62 h ttp :// w w w .m ar in es pe ci es .o rg / 63 h ttp :// w w w .th ep la nt lis t.o rg /, ne w v er si on s ep te m be r 2 01 3 64 d ow nl oa de d fr om h ttp :// w w w .b io na m es .o rg /d at a in a ug us t 2 01 4— no lo ng er a va ila bl e as a d is cr et e do w nl oa d, b ut c on te nt is in cl ud ed in 2 01 5 ve rs io n be lo w , w ith a dd iti on al d at a 65 h ttp :// w w w .b io na m es .o rg /d at a (a pr il 20 15 v er si on ), al so a va ila bl e at h ttp s: //z en od o. or g/ re co rd /1 68 63 . o rig in al d at a so ur ce d fr om io n , h ttp :// w w w .o rg an is m na m es .c om . biodiversity informatics, 12, 2017, pp. 1-44 41 4/ 10 /2 01 6 w or m s (s ep te m be r 2 01 6 ve rs io n) 2, 30 0 ad di tio na l g en us n am es , p lu s at tri bu te s / ta xo no m ic p la ce m en t up gr ad ed fo r a fu rth er 4 3, 00 0 ta xa . u pd at e on ly (f or n ew n am es si nc e 20 13 v er si on ), ge ne ra on ly . 16 /1 2/ 20 16 a pg iv (a ng io sp er m p hy lo ge ny g ro up 20 16 ) 5 ne w a ng io sp er m o rd er s, c. 1 0 ne w fa m ili es ; r el ev an t m ov es , s pl its a nd m er ge rs c ar rie d ou t a t f am ily le ve l, so m e ge ne ra re lo ca te d. su pe rs ed es p re vi ou s t re at m en t f ro m g r in a nd a pg ii i, w he re d iff er en t. biodiversity informatics, 12, 2017, pp. 1-44 42 appendix 2: irmng data conventions (2006-current) scientific names • scientific names incorporating diacritical marks (other than the optional diaeresis in botany) are not permitted under the relevant nomenclatural codes, refer international commission on zoological nomenclature (1999) article 27 (zoology); mcneill et al. (2012) article 60.6 (botany). where supplied (sometimes via older works) they are standardized according to rules specified in the botanical code (icnafp), such that mülleri becomes muelleri, and so on. o exception: for botanical names incorporating the diaeresis (example: isoëtes, the quillwort), a “plain form” of the name (isoetes in this case) is created for use within irmng and the original version maintained as an alternative spelling with its own entry, and listed as a synonym of the plain form / present accepted name, as applicable. • specific epithets with an initial capital letter (archaic / incorrect usage) are normalized to lowercase. o example: hippopotamus madagascariensis is normalized to hippopotamus madagascariensis. • subgenera are removed where supplied in parentheses between a genus name and specific epithet (refer example cited earlier, under "treatment of subgenus names")— species are presently attached to their parent name at generic, not subgeneric rank. authorities • author surnames are preferably spelled out in full—including authors of botanical genera (not yet implemented for botanical species, but may be in the future)—and include publication year when known. o examples (generic level): § homo linnaeus, 1758 (zoology) § prunus linnaeus, 1753, not “prunus l.” (botany). • zoological authors are given without initials (except as noted below), for botanical and bacteriological authors these are included when known. o example: § uharella taylor, casadío & gordon, 2008 (zoology); § scutifolium d.w. taylor, g.j. berner & s.h. basha, 2008 (botany). in zoology, an exception is made for a small number of cases where confusion is possible due to authors sharing the same surname working on the same group at a similar time, such as h. adams vs. a. adams (molluscs), m. sars vs. g.o. sars (marine crustaceans). • author combinations are conjoined using the ampersand character consistently. o example: “cavalier-smith & chao”, not “cavalier-smith et chao” (convention elsewhere in botanical works), “cavalier-smith and chao”, etc. • “long form” citations (i.e., author a in [work authorship] a & b, etc.) are preferred to “short form” citations, when known to be applicable. o example: “cavalier-smith in cavalier-smith & chao, 2010”, not “cavaliersmith, 2010”. • diacritics in author names are added back where missing in supplied data. o examples: chujo, 1969 becomes chûjô, 1969, etc. biodiversity informatics, 12, 2017, pp. 1-44 43 • representation of non-english author names is checked against additional sources and adjusted/corrected as required. o example: “lu junchang, pu hanyong, xu li, wu yanhua & wei xuefang, 2012” (as given in ion) is corrected to “lu, pu, xu, wu & wei, 2012”; “vilela cruz, falcao salles & hamada, 2013” (in ion) is corrected to “cruz, salles & hamada, 2013”. • multiple authorships are spelled out in full where these do not exceed five authors. o example: “cruz, salles & hamada, 2013”, not “cruz et al., 2013” (“et al.” is used from the sixth author onwards). • a comma is inserted before all dates in authorship citations (this is optional according the zoological code but standardized for irmng use), e.g., “linnaeus, 1758”, not “linnaeus 1758”. • “curly” apostrophes are replaced with “plain” forms within author names for consistency of data entry and text searching (e.g., “o'donohue”, not “o’donohue”). literature citations from 2016, literature citations are being entered according to worms/aphia conventions, with separate, searchable fields for author name(s), year, article title, journal title, volume, article pagination, and page in work, also doi (digital object identifier) and online link as available. these are then reassembled and formatted as required (e.g., with the journal title italicized, and doi / online link presented as hyperlinks) for presentation on relevant source and taxon pages. citations entered prior to 2016 are slowly being converted to the new standard, commencing with those used for large numbers of taxon names or as the source for the most recent treatments for certain groups. biodiversity informatics, 12, 2017, pp. 1-44 44 biodiversity informatics, 18, 2024, pp. 24-27 24 guideline materials and documentation for the genetic diversity indicators of the monitoring framework for the kunming-montreal global biodiversity framework alicia mastretta-yanes1,2, sofía suárez3, rebecca jordan4, sean hoban5,6, jessica m. da silva7,8, luis castillo-reina9, myriam heuertz10, fumiko ishihama11, viktoria köppä12, linda laikre12, anna j. macdonald13, joachim mergeay14,15, ivan paz-vinas16, gernot segelbacher17, alicia knapps17, henry rakoczy17, amelie weiler17, angelica atsaves17, kira cullmann17, simone bagnato17, brenna r. forester18 1 consejo nacional de humanidades ciencias y tecnología (conahcyt), avenida insurgentes sur 1582, crédito constructor, benito juárez, ciudad de méxico. c.p. 03940. mexico.* 2 departamento de ecología de la biodiversidad, instituto de ecología, universidad nacional autónoma de méxico. av. ciudad universitaria 3000, 04510, coyoacán, ciudad de méxico, mexico 3 laboratorio de genética de la conservación, jardín botánico, instituto de biología, universidad nacional autónoma de méxico. ciudad de méxico, mexico 4 csiro environment, 15 college rd, sandy bay 7005, tasmania, australia 5 center for tree science, the morton arboretum, lisle, usa 6 committee on evolutionary biology, the university of chicago, chicago, usa 7 south african national biodiversity institute, kirstenbosch research centre, private bag x7 claremont 7735, cape town, south africa 8 centre for ecological genomics and wildlife conservation, department of zoology, university of johannesburg, auckland park 2006, johannesburg, south africa 9 department of biology, faculty of science,ku leuven, leuven, belgium 10 univ. bordeaux, inrae, biogeco, f-33610 cestas, france 11 national institute for environmental studies, onogawa16-2, tsukuba, ibaraki, japan 12 department of zoology, stockholm university, se10691 stockholm, sweden 13 australian antarctic division, department of climate change, energy, the environment and water, kingston, tasmania 7050, australia 14 research institute for nature and forest, gaverstraat 4, 9500 geraardsbergen, belgium 15 ecology, evolution and biodiversity conservation, ku leuven, charles deberiotstraat 32, box 2439, leuven, belgium 16 universite claude bernard lyon 1, lehna umr 5023, cnrs, entpe, f-69622, villeurbanne, france 17 university of freiburg, wildlife ecology and management, 79106 freiburg im breisgau, germany 18 u.s. fish and wildlife service, fort collins, co, usa *corresponding author: alicia mastretta-yanes, email: amastretta@iecologia.unam.mx abstract. genetic diversity is fundamental to biological diversity, vital for species’ health and adaptation to environmental change. under the recently adopted kunming-montreal global biodiversity framework (gbf), 196 parties committed to report the status of genetic diversity for both wild and domesticated species. for this, three genetic diversity indicators were developed, two of which focus on processes contributing to genetic diversity conservation: ensuring that populations are large enough to maintain genetic diversity (effective population size ne 500 indicator) and maintaining genetically distinct populations (populations maintained, pm indicator). a third indicator focuses on the number of species being monitored using dna-based methods. adopted by 196 cbd parties in december 2022, gbf integrated ne 500 and pm as headline and complementary indicators, respectively. to aid nations in quantifying these indicators, a detailed set of guideline materials was developed, encompassing species selection, data compilation, and indicator computation. these guidelines draw from the collaborative efforts of the first multinational assessment of genetic divermastretta-yanes et al. – guidelines for genetic diversity indicators for the kunming-montreal framework 25 sity indicators that was recently completed and that will be refined continually through a versioning system, as more experience is gained and shared. the materials aim to support the global monitoring framework established by the cbd and are accessible online for utilization and updates. the guidelines are available at https://ccgenetics.github.io/guidelines-genetic-diversity-indicators/ key words: biodiversity indicators, kunming montreal global biodiversity framework, biodiversity monitoring, cop15, effective population size, population maintained, populations. genetic diversity is the foundation of all biological diversity. it is necessary for populations of both wild and domesticated species to remain healthy and be able to adapt to environmental change, and for conserving nature’s contributions to people (des roches et al., 2021). starting in 2020, during preparation of what would become the kunming montreal global biodiversity framework (gbf), three genetic diversity indicators were developed (hoban et al., 2020, 2021; laikre et al., 2020): (1) effective population size (ne) 500 indicator, which measures the proportion of populations within a species that are of sufficient size (ne > 500) to maintain genetic diversity and adaptive potential within that species; (2) populations maintained (pm) indicator, which measures the proportion of populations that still exists compared to the total number of populations that used to occur; and (3) a dna-based monitoring indicator, which is a count of the number of species in which genetic diversity has been or is being monitored using dna-based methods. the first two indicators focus on processes contributing to genetic diversity conservation: ensuring that populations are large enough to maintain genetic diversity (ne 500 indicator) and maintaining genetically distinct populations (pm indicator). these two indicators were adopted in 2022 by gbf as headline a4 and complementary indicators, respectively, which means that gbf parties will use these indicators to report on their progress over the next decade (cbd, 2022b, 2022a). they also cover two key aspects of the gbf: conserving genetic diversity both within populations and between populations. the gbf commitment to monitoring genetic diversity for all species, instead of only socioeconomically and culturally valuable taxa (as was required in 2010-2020), represents a significant milestone for conservation genetics, but comes with new challenges. by focusing on processes underlying the generation and maintenance of genetic diversity, the pm and ne 500 indicators alleviate some of these challenges because they can be estimated using both genetic and non-genetic data (hoban et al., 2020; laikre et al., 2020). genetic data include dna-based molecular markers to estimate ne or delimit population boundaries, whereas non-genetic data include census population size (nc), which could be transformed to ne using a ne/nc ratio, as well as occurrence data and knowledge of the species’ biology, history and dispersal to define populations geographically (hoban et al., 2023, 2024). integrating data from these diverse sources, formats and disciplines, from global databases to local knowledge, would make it possible to monitor genetic diversity across the world, much faster than is possible with genetic studies alone (mastretta-yanes et al., 2024). gathering and sharing biodiversity data has its own difficulties (blair et al., 2020; enke et al., 2012). however, integrating the diverse data needed to estimate genetic diversity indicators also requires capacity building to integrate genetic principles to new fields, designing new data collection protocols, and mobilizing data for indicator calculation in a reliable, transparent, and inclusive way, without stretching the personnel, time, and financial resources of the agencies in charge of reporting them. to contribute to meeting these challenges, with colleagues, we co-developed the following guideline materials, as part of the first multinational assessment of the genetic diversity indicators (mastretta-yanes et al., 2024). part of the guidelines were described previously in hoban et al. (2023), but after implementing them across nine countries (australia, belgium, colombia, france, japan, mexico, south africa, sweden, and the united states of america), several improvements were made. these improvements incorporate feedback from around 80 participants (students, practitioners and researchers) who gathered data for the indicators, as well as feedback from 13 international webinars & seminars with hundreds of participants. this publication leverages our shared experience, with more detailed guidelines in an online documentation format, that will be kept updated through a versioning system as more teams share insights. the materials are intended to assist nations in quantifying genetic indicator values https://ccgenetics.github.io/guidelines-genetic-diversity-indicators/ https://www.zotero.org/google-docs/?kkvp8a https://www.zotero.org/google-docs/?kkvp8a https://www.zotero.org/google-docs/?xco5ok https://www.zotero.org/google-docs/?xco5ok https://www.zotero.org/google-docs/?jf50bo https://www.zotero.org/google-docs/?jf50bo https://www.zotero.org/google-docs/?kvvcc2 https://www.zotero.org/google-docs/?kvvcc2 https://www.zotero.org/google-docs/?dchu3i https://www.zotero.org/google-docs/?dchu3i https://www.zotero.org/google-docs/?qq8jdq https://www.zotero.org/google-docs/?novdxv https://www.zotero.org/google-docs/?novdxv https://www.zotero.org/google-docs/?tosgnj https://www.zotero.org/google-docs/?tosgnj https://www.zotero.org/google-docs/?07bfcz mastretta-yanes et al. – guidelines for genetic diversity indicators for the kunming-montreal framework 26 at every stage of the process: from species selection to data compilation to indicator calculation. we hope they become useful as a reference point, from which countries can adjust their protocols to their own needs and preferences. the guidelines are available as an online documentation at https://ccgenetics.github.io/guidelines-geneticdiversity-indicators/ the online documentation consists of eight content sections, as follows: (1) background on the population genetics rationale behind the genetic diversity indicators, (2) quickstart guide summarizing steps needed to estimate the indicators, (3) discussion on how many and which species to include in the species list to evaluate the indicators, (4) how-to guides with practical examples showing how to perform the most common tasks involved in assessing the genetic diversity indicators, (5) example assessments of real-life species, (6) data collection advice with a ready-to-use template for a web tool for data collection using kobotoolbox, (7) equations, scripts and examples for calculations and reporting of the indicators, and (8) glossary. acknowledgments these materials are based on the co-creation experience of the first pilot multinational assessment of the genetic diversity indicators, and on interactions with practitioners, researchers and students of several institutions across the world. we are particularly grateful to the swedish environmental protection agency, the ad hoc technical expert group on indicators for the kunming-montreal global biodiversity framework, and to the following people for providing feedback and ideas: akio takenaka, alejandra domínguez álvarez, alexander llanes-quevedo, alice hughes, ana wegier, ashley hamilton, atsaves angelica, austin koontz, bastian silva, belma kalamujić stroil, caitlin miller, catherine e grueber, w chris funk, emma suzuki spence, erica robertson, eugenia zarza, fleur visser, gaëlle brahy, georgina wood, glenn m shea, henrik thurfjell, hesiquio benitez, irene ramos, iris lang, isa-rita russo, juan francisco ornelas, katie millette, keiichi fukaya, kira cullmann, libertad arredondo-amezcua, lily durkee, lucía ruíz, luke dedecke, malte julius benedikt lehmann, malte lehmann, margaret e. hunter, maria alejandra rodriguez-morales, maría camila latorre, marlien van der merwe, matt desaix, meg mahoney, metztli arcila santiago, mónica alegre, paulette bloomer, per sjögren-gulve, philipp ungar, robyn e. shaw, santiago ramírez-barahona, sheela turbek, simone bagnato, sofía treviño, tanya latty, taylor stack, and victor julio rincon-parra. author contributions am-y developed the github repository and coordinated the overall project with brf. ss and am-y loaded the content into the github repository and programmed the kobo-form. am-y, ss, rj, sh, jmds and brf wrote most of the updated version of the guidelines and made figures. am-y, rj, sh, jmds, brf, lc-r, mh, fi, vk, ll, ajm, jm, ip-v wrote the initial version of the guidelines and made figures. gs coordinated the testing of the guidelines and contributed to the test. ak, hr, aw, aa, kc and sg tested the guidelines and provided step-bystep examples following them. all authors contributed to the design of the guidelines, the kobo-form and proof-read the guidelines and the manuscript. competing interests the authors have declared that no competing interests exist. references blair, j., gwiazdowski, r., borrelli, a., hotchkiss, m., park, c., perrett, g., & hanner, r. (2020). towards a catalogue of biodiversity databases: an ontological case study. biodiversity data journal, 8. https://doi.org/10.3897/bdj.8.e32765 cbd. (2022a). decision adopted by the conference of the parties to the convention on biological diversity cbd/cop/ dec/15/4 kunming-montreal global biodiversity framework. cbd/cop/dec/15/4. https://www.cbd.int/doc/decisions/cop-15/cop-15-dec-04-en.pdf cbd. (2022b). decision adopted by the conference of the parties to the convention on biological diversity cbd/cop/ dec/15/5 monitoring framework for the kunming-montreal global biodiversity framework. cbd/cop/dec/15/5. https://www.cbd.int/doc/decisions/cop-15/cop-15-dec-05en.pdf des roches, s., pendleton, l. h., shapiro, b., & palkovacs, e. p. (2021). conserving intraspecific variation for nature’s contributions to people. nature ecology & evolution, 1–9. https://doi.org/10.1038/s41559-021-01403-5 enke, n., thessen, a., bach, k., bendix, j., seeger, b., & gemeinholzer, b. (2012). the user’s view on biodiversity data sharing—investigating facts of acceptance and requirements to realize a sustainable use of research data. ecological informatics, 11, 25–33. https://doi.org/10.1016/j. ecoinf.2012.03.004 https://ccgenetics.github.io/guidelines-genetic-diversity-indicators/ https://ccgenetics.github.io/guidelines-genetic-diversity-indicators/ https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu mastretta-yanes et al. – guidelines for genetic diversity indicators for the kunming-montreal framework 27 hoban, s., bruford, m., d’urban jackson, j., lopes-fernandes, m., heuertz, m., hohenlohe, p.a., paz-vinas, i., sjögrengulve, p., segelbacher, g., vernesi, c., aitken, s., bertola, l.d., bloomer, p., breed, m., rodríguez-correa, h., funk, w.c., grueber, c.e., hunter, m.e., jaffe, r., liggins, l., mergeay, j., moharrek, f., o’brien, d., ogden, r., palma-silva, c., pierson, j., ramakrishnan, u., simo-droissart, m., tani, n., waits, l., laikre, l., (2020). genetic diversity targets and indicators in the cbd post-2020 global biodiversity framework must be improved. biological conservation, 248, 108654. https://doi.org/10.1016/j.biocon.2020.108654 hoban s, da silva jm, hughes a, hunter me, kalamujić stroil b, laikre l, mastretta-yanes a., millette k, paz-vinas i, ruiz bustos l, shaw re, vernesi c, the coalition for conservation genetics (2024). too simple, too complex, or just right? advantages, challenges, and guidance for indicators of genetic diversity. bioscience, biae006. hoban, s., da silva, j. m., mastretta-yanes, a., grueber, c. e., heuertz, m., hunter, m. e., mergeay, j., paz-vinas, i., fukaya, k., ishihama, f., jordan, r., köppä, v., latorre-cárdenas, m. c., macdonald, a. j., rincon-parra, v., sjögren-gulve, p., tani, n., thurfjell, h., & laikre, l. (2023). monitoring status and trends in genetic diversity for the convention on biological diversity: an ongoing assessment of genetic indicators in nine countries. conservation letters, 16(3), e12953. https://doi.org/10.1111/conl.12953 hoban, s., paz-vinas, i., aitken, s., bertola, l. d., breed, m. f., bruford, m. w., funk, w. c., grueber, c. e., heuertz, m., hohenlohe, p., hunter, m. e., jaffé, r., fernandes, m. l., mergeay, j., moharrek, f., o’brien, d., segelbacher, g., vernesi, c., waits, l., & laikre, l. (2021). effective population size remains a suitable, pragmatic indicator of genetic diversity for all species, including forest trees. biological conservation, 253, 108906. https://doi.org/10.1016/j.biocon.2020.108906 laikre, l., hoban, s., bruford, m. w., segelbacher, g., allendorf, f. w., gajardo, g., rodríguez, a. g., hedrick, p. w., heuertz, m., hohenlohe, p. a., jaffé, r., johannesson, k., liggins, l., macdonald, a. j., orozco-terwengel, p., reusch, t. b. h., rodríguez-correa, h., russo, i.-r. m., ryman, n., & vernesi, c. (2020). post-2020 goals overlook genetic diversity. science, 367(6482), 1083–1085. https:// doi.org/10.1126/science.abb2748 mastretta-yanes, a., da silva, j.m., grueber, c.e., castillo-reina, l., köppä, v., forester, b.r., funk, w.c., heuertz, m., ishihama, f., jordan, r., mergeay, j., paz-vinas, i., rincon-parra, v.j., rodriguez-morales, m.a., arredondo-amezcua, l., brahy, g., desaix, m., durkee, l., hamilton, a., hunter, m.e., koontz, a., lang, i., latorre-cárdenas, m.c., latty, t., llanes-quevedo, a., macdonald, a.j., mahoney, m., miller, c., ornelas, j.f., ramírez-barahona, s., robertson, e., russo, i.-r.m., santiago, m.a., shaw, r.e., shea, g.m., sjögren-gulve, p., spence, e.s., stack, t., suárez, s., takenaka, a., thurfjell, h., turbek, s., van der merwe, m., visser, f., wegier, a., wood, g., zarza, e., laikre, l., hoban, s. (2024). multinational evaluation of genetic diversity indicators for the kunming-montreal global biodiversity framework. ecology letters, 27(7), e14461. https://doi.org/10.1111/ele.14461 https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://www.zotero.org/google-docs/?zbcizu https://doi.org/10.1111/ele.14461 biodiversity informatics, 19, 2025, pp. 85 85 book review the game of species. an introduction to biodiversity. j. s. lópez. pelagic publishing, 2022 sara gamboa mapaslab, ecology and animal biology department, university of vigo, spain the game of species: an introduction to biodiversity offers a distinctive synthesis of evolutionary and ecological thought, presented through the powerful metaphor of a game. the “pieces” are species, the “boards” are ecological and evolutionary arenas, and the “rules” are the processes that generate and sustain diversity. this framing, at once playful and rigorous, provides an accessible yet sophisticated introduction to biodiversity science. the book is an english-language adaptation and expansion of lópez’s earlier el juego de las especies (2019), originally published in spanish. in revising the text for an international audience, the author has updated examples, integrated recent literature, and refined the pedagogical flow. the result is not merely a translation, but a thoughtful reworking that situates the book ideally within broader debates in biodiversity and conservation. each chapter builds progressively: from the origin of species (on the origin of pieces), to their ecological functions (on the roles of pieces), the environments that structure them (on the boards), and the dynamic processes that animate them (on the dynamics of the game). the treatment of coevolution and the red queen hypothesis (van valen), the presentation of the classic framework of island biogeography (macarthur & wilson), and the discussion of adaptive radiations place the book firmly in the lineage of modern evolutionary ecology. darwin’s legacy is present throughout, not in a purely historical sense but as an ongoing influence on how we ask and answer questions about the diversity of life. one of the book’s strengths is its ability to convey complexity without resorting to unnecessary jargon. lópez explains why biodiversity is unevenly distributed across the planet, why extinction is as important as speciation, and why ecological interactions create functional rather than merely numerical diversity. the discussion is enriched by careful illustrations that serve not only as visual aids, but as integral components of the narrative. inclusion of a glossary further enhances the accessibility of the text, making it particularly valuable for students and non-specialists who may be encountering these concepts for the first time. at the same time, the book does not shy away from presenting the field’s major theoretical frameworks. readers are introduced to the red queen hypothesis, island biogeography, ecological niches, keystone species, and functional diversity. these explanations are faithful to the original sources while presented in a way that highlights their continuing relevance. for evolutionary biologists, this balance of synthesis and pedagogy is particularly appealing: the text revisits familiar debates while offering a fresh metaphorical lens. the closing chapters (redux and the future of life) provide both a synthesis and a call to action. lópez emphasizes that understanding biodiversity as a game entails recognizing our own disruptive role as players who alter the rules. the urgency of conservation emerges as a natural consequence of the scientific narrative, rather than as a moral afterthought. the intended audience is broad: undergraduate and graduate students in biology will find it an accessible entry point; educators will appreciate its clarity and illustrative richness; and professionals will enjoy the reflective value of revisiting foundational questions. it also works as a piece of high-level science communication for readers outside academia, who will find in it both intellectual stimulation and aesthetic pleasure. if there is a limitation, it is the book’s relatively cursory treatment of microbial diversity. this gap is acknowledged by the author from the outset. still, the book’s candor about its scope, its elegant prose, and its integration of illustration and glossary compensate amply for this omission. in conclusion, the game of species is a lucid, engaging, and beautifully produced introduction to biodiversity science. by combining metaphor, illustration, classical theory, and contemporary concerns, lópez has produced a work that informs, inspires, and challenges. for those of us trained in evolutionary biology, it is a welcome reminder of why biodiversity matters: not only as an object of study, but as the living game in which we, too, are players. biodiversity informatics, 18, 2024, pp. 56-77 56 addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases elisa marchetto1,*, martina livornese1,*, francesco maria sabatini1,2, enrico tordoni3, daniele da re4,5 jonathan lenoir6, riccardo testolin1, giovanni bacaro7, roberto cazzolla gatti1, alessandro chiarucci1, giles m. foody8, lukáš gábor9,10, quentin groom11, jacopo iaria1, marco malavasi12, vítězslav moudrý9, diletta santovito1, petra šímová9, piero zannini1, and duccio rocchini1,9 1alma mater studiorum university of bologna, department of biological, geological and environmental sciences, via irnerio 42, 40126 bologna, italy 2czech university of life sciences prague, department of forest ecology, faculty of forestry and wood sciences, kamýcka 129, 165 21 prague, czech republic 3university of tartu, institute of ecology and earth science, j. liivi 2, 50409 tartu, estonia 4university of trento, center agriculture food environment, via edmund mach, 1, 38098 san michele all’adige, italy 5research and innovation centre, fondazione edmund mach, san michele all’adige, italy 6umr cnrs 7058 ecologie et dynamique des systèmes anthropisés (edysan), université de picardie jules verne, 1 rue des louvels, 80037 amiens, france 7university of trieste, department of life sciences, via l. giorgieri 10, 34127 trieste, italy 8university of nottingham, school of geography, university park, nottingham ng7 2rd, uk 9czech university of life sciences prague, faculty of environmental sciences, department of spatial sciences, kamýcka 129, 16500 praha suchdol, czech republic 10yale university, new haven, ct 06520, united states 11meise botanic garden, nieuwelaan 38, 1860 meise, belgium 12university of sassari, department of chemistry, physics, mathematics and natural sciences, via vienna 2, 07100 sassari, italy *authors equally contributed to the manuscript corresponding author: duccio rocchini, duccio.rocchini@unibo.it orcid id 0000-0003-0087-0594 biome lab, department of biological, geological and environmental sciences (bigea), alma mater studiorum university of bologna, piazza di porta s. donato, 1, 40126 bologna, italy marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 57 abstract. the availability of biodiversity databases is expanding at unprecedented rates. nevertheless, species occurrence data can be intrinsically biased and contain uncertainties that impact the accuracy and reliability of biodiversity estimates. in this study, we developed a reproducible framework to assess three dimensions of bias—taxonomic, spatial, and temporal—as well as temporal uncertainty associated with data collections. we utilized the vegetation plot data located in europe, from splotopen, an open-access database, as a case study. the metrics proposed for estimating bias include completeness of the species richness for taxonomic bias, nearest neighbor index for spatial bias, and pielou’s index for temporal bias. additionally, we introduced a new method based on a negative exponential curve to model the temporal decay in biodiversity data, aiming to quantify temporal uncertainty. finally, we assessed the sampling bias considering the influence of various spatial variables (i.e, road density, human population count, natura 2000 network and topographic roughness). we discovered that the facets of bias and the temporal uncertainty varied throughout europe, as did the different roles played by spatial variables in determining biases. splotopen showed a clustered distribution of the vegetation plots, and an uneven distribution in sampling completeness, year of sampling and temporal uncertainty. the facets of bias were significantly explained mainly by the presence of natura 2000 network and marginally by the human population count. these results suggest that employing an efficient procedure to examine biases and uncertainties in data collections can enhance data quality and provide more reliable biodiversity estimates. key words: biodiversity; community composition; data quality; spatial bias; taxonomic bias; temporal bias; temporal uncertainty introduction biodiversity and ecosystem functioning are experiencing a widespread degradation globally. the main drivers of biodiversity decline are represented by an increase in the intensity of human activities such as land and sea-use, the exploitation of organisms and natural resources, atmospheric and water pollution as well as the introduction of alien species (ipbes 2019). together with climate change, whose impact on biodiversity is expected to increase in the coming years (di marco et al. 2019), these factors pose a significant threat to the integrity of ecosystems and biodiversity. to monitor biodiversity change, we need records that capture the occurrence and/or co-occurrence (i.e. community composition) of species within specific time frames and geographical locations. these raw records, now increasingly available through global biodiversity collections such as the bien and splot database (enquist et al. 2016; bruelheide et al. 2019), play a crucial role in ecological research and represent essential sources of information for guiding and monitoring actions aimed at meeting global biodiversity targets (boakes et al. 2010; meyer et al. 2015). their utility spans over a wide range of applications, including investigations into species redistribution (jandt et al. 2022b), community reassembly (bertrand et al. 2011), threat assessment and conservation planning (ricci et al. 2024), as well as the study of invasive species propagation (turbelin et al. 2017). since the 2000s, the number of publicly available biodiversity databases has risen, alongside their use (ball-damerow et al. 2019). data availability alone, however, is not sufficient to ensure reliable ecological inferences. as a matter of fact, data quality should be considered and checked, both in terms of spatial and temporal representativeness (wüest et al. 2020). one common issue with biodiversity databases relates to the way in which data are collected. frequently, these databases contain opportunistic collections of data, which are characterized by uneven sampling effort and might hide subtle sources of bias and uncertainties (daru and rodriguez 2023; garcía-roselló et al. 2023; rocchini et al. 2023). when these limitations are not accounted for, our ability to describe and analyse biodiversity might be compromised (hortal et al. 2015). bias and uncertainty are terms developed in the statistical literature, and refer to the theory of sampling (walther and moore 2005). bias occurs when the sampling is unrepresentative of the target statistical population. it might depend on uneven sampling across geographic areas, taxonomic groups or time periods (walther and moore 2005). uncertainty, on the other hand, refers to the lack of precision in measurements, which also affects the degree to which data can represent reality (hortal et al. 2015). biodiversity data are particularly prone to these problems, and considerations on the bias and uncertainty of the data marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 58 acquire particular relevance across three specific dimensions: taxonomic, spatial and temporal (meyer et al. 2016). while assessments of the limitations posed by the use of biodiversity databases do exist (monsarrat et al. 2019; colli‐silva et al. 2020; ronquillo et al. 2020), most studies focus on one dimension at the time, commonly spatial or taxonomic (but see (meyer et al. 2016) for a multidimensional approach), and often consider only bias but not their related uncertainty. taxonomic bias is a well-known issue in biodiversity research, where the study of specific taxa is favoured over others (troudet et al. 2017) (e.g. vertebrates over invertebrates and vascular plants over bryophytes and lichens). as a result, biodiversity databases may overand under-represent different taxonomic groups (garcía-roselló et al. 2023). in the geographical space, taxonomic bias can be analysed using measures of inventory or sampling completeness, which estimate taxonomic coverage of the collected data within a given surface area (chao and jost 2012). traditionally, sampling completeness is calculated using parametric or non-parametric estimators of the expected species richness within a given spatial unit and then computing the ratio of observed versus expected species richness (chesshire et al. 2023). alternatively, a metric of completeness is given by the final slope of species accumulation curves for the investigated geographic unit (yang et al. 2013; girardello et al. 2019). reliable methods for species richness estimation based on a combination of probabilistic and opportunistic data are now available (chiarucci et al. 2018) but can hardly be applied only using opportunistically collected data. spatial bias arises when data distribution and density are uneven in space, as a result of an unbalanced sampling design (tessarolo et al. 2014; rocchini et al. 2023). the spatial distribution of collected data is often the result of socio-economic factors such as accessibility and the presence of road networks (oliveira et al. 2016), uneven financial investments in research across regions (meyer et al. 2015), but also the preference for sampling in nature protected areas hosting rare or charismatic species (yang et al. 2014). the spatial distortion of the data resulting from these factors might yield inaccurate modelling outputs, especially when modelling species distribution (bazzichetto et al. 2023; rocchini et al. 2023). for being aggregated over long time periods, considerations on biodiversity data should take into account the temporal dimension. this aspect is gaining attention as reliably estimating biodiversity loss and change in time stand as a paramount challenge in ecological research (jandt et al. 2022b). however, surveys are often not conducted systematically over time, leading to collections characterized by uneven data coverage and large temporal gaps where no record is present. like bias, uncertainty is present in all the components of biodiversity data and can stem from various sources. for instance, in the taxonomic dimension uncertainty may arise from imprecise or equivocal species names (stropp et al. 2022), whereas in the geographic space, positional inaccuracy of survey locations is recognized as a contributor to the overall uncertainty in the data (gábor et al. 2020). while these aspects of taxonomic and spatial uncertainty are routinely considered in macroecological research, the uncertainty derived from the temporal dimension of the data is often neglected. natural communities are not constant over time and exhibit spatial and/ or compositional shifts in response to natural variability and/or human-induced alteration in land use, climate and introduction of alien species (newbold et al. 2015). because of the dynamism of ecological systems, the information associated with any data on the occurrence of a certain species or species assemblage in a specific area inevitably decays with time (tessarolo et al. 2017). understanding this process of information decay becomes particularly relevant when biodiversity records are used in conservation planning, where accurate and up-to-date knowledge is essential (boitani et al. 2011). given all the above factors, it is important to recognize the different limitations of biodiversity databases and identify new approaches to tackle them. here, we showcase how different aspects of bias and uncertainty can be quantified. as an example, we used vegetation plot data in europe from the openaccess database splotopen (sabatini et al. 2021b). we assessed four specific aspects of error through the use of different metrics: taxonomic bias, spatial bias, temporal bias and temporal uncertainty, so to explore the geographical pattern of these sources of error. finally, we explored how these sources of error relate to a set of geographic variables, namely human population count and road density, the occurrence of protected areas and topographic roughness. the ultimate goal is to provide a workflow (fig. 1) that can be generalized and applied to other biodiversity databases, regardless of the spatial scale of the analysis. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 59 material and methods data preparation splotopen is an open-access, stratified subset of the splot database. it includes only vascular plant species and was built based on climatic and soil variables as resampling strata (sabatini et al. 2021b). the stratified resampling used to build splotopen specifically focuses on maximizing the representativeness of the vegetation plot data in the environmental space, at the expense of the geographical space. after accessing splotopen (march 2023 version 2.0, (sabatini et al. 2021a)), we exclusively extracted data 1) located in europe and within the boundaries of laea europe coordinates system (wgs84 bounds: -16.1, 32.88, 40.18, 84.73), 2) having coordinates uncertainty lower than 250 m, and 3) with a year of recording equal to or greater than 1992. we did this to minimize errors coming from the inaccurate location of the plots, mainly deriving from possible errors of data georeferencing, and to be consistent with the year of establishment of the natura 2000 network. this filtering phase reduced the data from 94,951 to 9,481 vegetation plots. we superimposed a grid of 0.5 degree resolution (epsg:4326) over the european extent and projected it to laea europe coordinates system (etrs89-extended, epsg:3035). accordingly, the resolution of the grid cells was transformed from 0.5 degrees to 39.5 km. finally, we assigned each vegetation plot to its corresponding grid cell. bias we measured and represented three facets of bias (taxonomic, spatial and temporal) and we plotted them in a trivariate map (appendix). taxonomic bias.—we represented the spatial distribution of the taxonomic bias, according to the taxonomic coverage of the vascular plants in splotopen, in terms of completeness in species richness. using chao’s formula to estimate the total number of species in a grid cell, we calculated the sample completeness as the ratio of the observed species in a sample to the true species richness (observed plus undetected) in the entire assemblage (chao et al. 2020). we used the r package inext (version 3.0.1) (hsieh et al. 2016), to determine the species richness for each grid cell of 39.5 km. for each grid cell, the input data comprise the number of sampling units (t) (i.e., vegetation plots), the observed incidence frequencies and the hill number q = 0. we set k, the equally spaced knots (samples sizes), to 5 and we removed to the input data all those grid cells containing two vegetation plots or less. the values of the completeness of species richness were calculated without considering that plots varied in size, within and across grid cells. spatial bias.—we estimated the degree of spatial bias (or geographical sampling bias) by estimatfigure 1: methodological workflow to assess the presence of taxonomic, spatial and temporal shortfalls in biodiversity databases. the assessment of bias in raw data involves the following measurements: sampling completeness for taxonomic bias, nearest neighbour index for spatial bias, pielou’s index for temporal bias. the temporal uncertainty is calculated using a negative exponential curve. the different facets of bias and the temporal uncertainty are computed for grid cells of 39.5 km. the response of biases to spatial variables is estimated by fitting generalized additive models. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 60 ing the spatial pattern of the plots locations within each grid cell through the nearest neighbor index (nni) (clark and evans 1954). we used the r package spatstat (version 3.0.8) and the package spatstat. explore (revision: 1.21, date: 2023/10/17). the nni was computed using the function clarkevanscalc (baddeley et al. 2016) and it evaluates whether the plots exhibit a clustered or random distribution. the nni is expressed as the ratio of the observed average distance between each plot and its nearest neighbor and the expected average distance in a random distribution with the same number of plots. values of the index less than one indicate clustering i.e., higher spatial bias, values around one a random distribution i.e., lower spatial bias, whereas values greater than one imply overdispersion (e.g., systematic distribution). we also modified the original clarkevans.test function in clarkevans.test2 to calculate the gridbased nni with standardized effect size (nni ses) as the difference between the observed nni and the mean of nni simulations divided by the standard deviation of the simulations. we used monte carlo approach to generate 999 populations of plots location under the condition of a complete spatial randomness (csr) of the observed number of plots. then, for each valid simulation we calculated nni within the extent of the grid cell. temporal bias.—pielou’s index (j) is a metric commonly used in ecology to assess how equitable or even the abundance of species is within a specific community or ecosystem (pielou 1966). in this work, we used pielou’s evenness to estimate the temporal bias of plot data based on the years of different plots were recorded for each grid cell. we computed the metric using the functions provided by the r package vegan (version 2.6.6). pielou’s index is calculated as (1), where h is the shannon-wiener index and it is calculated as (2). traditionally, n represents the total number of species and pi is their relative abundances for each species i ∈ {1, . . . ,n}. the maximum value of shannon’s index is expressed as: hmax = lnn. it is the value that indicates an even distribution, which is attained when all species have equal relative abundances. in our study, n refers to the total number of years of recording, where i is the ith year of recording, and pi is the proportion of plots in a grid cell being sampled in year i. this means that the pielou’s evenness was calculated by taking into account the number of plots per grid cell, instead of the number of individuals, that share the same year of recording. higher is the value of pielou’s index lower is the temporal bias. temporal uncertainty the information associated with any biodiversity data decays with time. we modelled the temporal decay of the information by applying a negative exponential transformation to our data. the function is defined as (3), where y is the temporal precision, i.e., the remaining information associated with a vegetation plot, and t is the difference between the year of the most recent surveyed plot (i.e., 2014) and the date of recording of the data point. since there is no way of knowing the actual rate of information decay for a vegetation plot, we calculated our results using three different exponents (i.e., z1= -1, z2= -1/5, z3= -1/25) so that the curves decrease with different rates (according to the slope, appendix: supplementary fig. s1). therefore, for each plot we calculated three values of temporal precision. finally, we quantified the temporal uncertainty of the vegetation plot data in a given grid cell as the median value of 1 temporal precision of each plot. we chose negative exponential functions, as they have four desirable properties, when compared to other linear transformations. first, negative exponentials are consistent with the assumption that the information associated to a vegetation plot can only decrease (or be stable) with time (i.e., is monotonically decreasing), and that this information will never reach zero. this corresponds to the reasonable assumption that having vegetation plot data for an area, no matter how old the data is, will always provide more information than having no data at all. second, negative exponentials can be used to constrain the amount of remaining information to a 0-1 interval, which is intuitive and easy to communicate. third, negative exponentials are simple and versatile functions that can assume a range of shapes, including a linear shape for short time intervals. finally, negative exponentials have often been used to model the decrease of a quantity against time or space. radioactive decay is the most typical example, but see xu et al. (2019) for an application on population decrease over time, or newling (1969) for the decrease in population as a function of the distance from the city center. however, we also tested the temporal decay of the plot information (i.e., temporal uncertainty) as a linear function of the median value of the differences between the year of the most recent surveyed plot (i.e., 2014) and the year of recording of the ith plot (see appendix for further details). marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 61 spatial variables of bias we selected a number of variables (number of plots, human population count, road density, natura 2000 network, and topographic roughness), which are likely to be related to the facets of bias (taxonomic, spatial, temporal) in splotopen data. we chose these variables because they have already been tested as sources of bias in several studies (ballesteros‐ mejia et al. 2013; geldmann et al. 2016; girardello et al. 2019). human population count: the human population count per pixel at 0.0083 degrees of spatial resolution for the year 2014 (year of the most recent plots in the database) was obtained from world pop1 (stevens et al. 2015). we calculated the human population count for each grid cell of 39.5 km as the mean value of the human population counts at the plot locations to be consistent with the method applied to calculate the facets of bias. accordingly, we extracted the values of the variable at 0.0083 degrees for each plot within the grid cell then, we calculated the mean value. road density: road density was employed as a metric to quantify the level of accessibility at the collection sites; road data shapefile for the european network were obtained from the global roads inventory project (grip)2 (meijer et al. 2018) and filtered by retaining only highways, primary and secondary roads. the road density was then calculated with a kernel density estimation (kde) at 1 km of spatial resolution through the spatstat package (baddeley et al. 2016). kernel density function is frequently employed to produce a continuous, smooth surface that depicts the spatial density of data points. we obtained the road density at 39.5 km by extracting the values from the original raster layer for each plot location then, we calculated the mean of the values included in each grid cell. natura 2000 network: we measured the relative number of plots inside the natura 2000 network to detect if the locations of the records were biased toward natura 2000 areas. the polygon layer of the natura 2000 network was obtained from the european environment agency.3 for each grid cell, we calculated the ratio between the number of plots located inside the natura 2000 area and the total number of 1 https://hub.worldpop.org/geodata/listing?id=64. 2 https://www.globio.info/download-grip-dataset. 3 https://www.eea.europa.eu/data-and-maps/data/natura-13, published: 6 oct 2022, temporal coverage: 2021. plots present in that grid cell, so as to obtain a gridbased measure of the number of plots inside the protected area which accounts also for the records size. topographic roughness: it refers to the variation in elevation and the spatial distribution of landform elements. this variable, which measures the topographic heterogeneity, was taken from (amatulli et al. 2018). we selected the topographic heterogeneity cause it determines the establishment of different habitats and diverse microenvironments that support different species (stein et al. 2014; barajas‐barbosa et al. 2020). therefore, if the sampling is not appropriately distributed across these different habitats, it can underestimate or lose certain species. the variable at 39.5 km of spatial resolution was obtained by extracting the values from the original rater layer with a spatial resolution of 0.4 degrees for each plot location then, calculating the mean of the values included in each grid cell. finally, we used these variables as predictors in three generalized additive models (gams), one for each measure of taxonomic, spatial and temporal bias (i.e., completeness of species richness, nni, and pielou’s evenness). we used the thin plate splines as spline-based technique for each smooth term of gam. the variables of gams were standardized to zero mean and one standard deviation before rescaling to a 0-1 range. we also considered the spatial autocorrelation including the term s(x,y) to the gam, where s is a smoothing spline and x and y are the longitude and latitude coordinates of the centroid of the grid cell. to control for the varying number of vegetation plots across grid cells, we added sampling effort as an additional explanatory variable to the models. sampling effort was calculated as the number of plots within each grid cell. results bias the taxonomic bias, described by the completeness of the species richness, was not evenly distributed over europe (fig. 2, appendix: supplementary fig. s3), following a similar pattern as the number of vegetation plots recorded per grid cell (appendix: supplementary figs. s4, s9). besides, the spatial distribution of the plots, measured through the nearest neighbor index (nni), was clustered almost everywhere in europe (fig. 2, appendix: supplementary fig. s6). most grid cells (97.4%) exhibited a clustered spatial pattern. the values of the nni https://hub.worldpop.org/geodata/listing?id=64 https://www.globio.info/download-grip-dataset https://www.eea.europa.eu/data-and-maps/data/natura-13 marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 62 confirmed that the effect size was large, pointing out that the magnitude of the deviation from the random expectation was substantial (appendix: supplementary fig. s7). furthermore, we observed that the temporal bias, calculated using pielou’s index to estimate the distribution of data across years, followed a different and independent pattern from the taxonomic and spatial bias (fig. 2, appendix: supplementary fig. s8). however, it highlighted a heterogeneous evenness of plots inventory over time. indeed, surveys turned out to be evenly distributed (i.e., lower bias) in several countries such as slovakia, netherlands and czech republic. overall, the european data in splotopen had high spatial clustering and heterogeneous temporal evenness and completeness of the species richness. additionally, the prevalence of one type of bias over another varied across geographic areas in europe, with some countries being characterized by the prevalence of one facet of bias over another (appendix: supplementary fig. s2). the completeness of the species richness (i.e., low taxonomic bias) showed to be preponderant in norway. an high temporal evenness ( i.e., low temporal bias) was observed in some plots in lithuania and in the netherlands, while low values of spatial bias were detected in czech republic. temporal uncertainty the different negative exponential functions being used for calculating the temporal uncertainty revealed that different exponents (i.e., z1= -1, z2= -1/5, z3= -1/25) allow for discriminating in different ways the pattern and intensity of the hotspots of temporal uncertainty (fig. 3). the temporal uncertainty measured using the exponent z1 = -1 was high across the entire european extent, except for some grid cells primarily distributed in estonia. the temporal uncertainty calculated with z2 = -1/5 highlighted new areas with lower uncertainty values, namely the danish peninsula and bulgaria. finally, the temporal uncertainty calculated using z3 = -1/25 smoothed out the values of temporal uncertainty, making uncertainty hotspots less visible compared to the uncertainties based on the other exponents. this exponent most closely approximated the negative exponential curve to a linear trend. indeed, its pattern of values was comparable with that obtained by calculating the uncertainty as median difference of the year of recording of the plot with the most recent one (appendix: supplementary fig. s11). spatial variables of bias the generalized additive models showed that most facets of bias are related to the presence of natura 2000 areas. the regression models of taxonomic, spatial and temporal bias had respectively a deviance explained of 49.5%, 14.8% and 22.2% (table 1). only natura 2000 network and human population count contributed to influencing the three facets figure 2: grid-based map of three facets of bias. the map shows a the uneven distribution of the taxonomic bias, b the distribution of the vegetation plots through the nearest neighbour index (spatial bias) and, c the heterogeneous distribution of the temporal bias. nni values greater than 1 indicate a random distribution of plots within a grid cell, while values less than 1 indicate a clustered distribution; high completeness of the species richness implies low taxonomic bias; high values of pielou’s index reveals low temporal bias. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 63 figure 3: the map shows the median temporal uncertainty of the vegetation plots per grid cell of 39.5 km; the intensity of the temporal uncertainty changes according to the exponents being used setting the exponential negative function (i.e., exponents: z1= -1, z2= -1/5 and z3= -1/25). table 1: terms of quality and fitting process of generalized additive models, as well as, overall significance of explanatory variables. n2k refers to relative number of plots in natura 2000 network, pop to human population count, road to road density, rough to topographic roughness. significance codes: ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1. reml refers to restricted maximum likelihood, r-squared to coefficient of determination, deviance expl. to deviance explained. taxonomic bias spatial bias temporal bias r-squared 0.461 0.117 0.187 deviance expl. 49.5% 14.8% 22.2% reml -115.35 -89.511 150.18 f p value f p value f p value n2k 2.927 < 0.05 * 3.807 < 0.05 * 4.117 < 0.01 ** pop 0.809 0.358 3.547 < 0.01 ** 1.903 0.118 road 0.081 0.918 1.392 0.239 2.377 0.124 rough 2.088 0.102 0.425 0.687 1.700 0.171 marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 64 of bias (fig. 4). specifically, the relative number of plots inside the natura 2000 network significantly explained the variability in all response variables while human population count was a significant predictor only for spatial bias. concerning the relative number of plots in natura 2000 network, lower values were associated with higher completeness of the species richness (lower bias), nevertheless the relationship was not linear (effective degree of freedom (edf) = 3.083); the completeness slightly decreased when the share of plots in natura 2000 areas increased from about 0.30 to 0.65 and, then increased again. also the nni did not follow a complete linear relationship with natura 2000 protected area (edf = 2.208) showing higher bias (low nni value) where the share of natura 2000 areas was higher. instead, the temporal bias reached its lowest value (highest pielou’s index) when the plots were almost evenly distributed both inside and outside the natura 2000 network; the degree of non-linearity was low with an edf value of 2.606. finally, the spatial bias decreased to about 0.30 of the human population count and then increased until it reached almost stability as the covariate increased (edf = 3.830). the control variable sampling effort had a significant effect on the variability of the three biases and the same applied to the term s(x,y) except for the spatial bias. overall, about 47% of the vegetation plots were inside natura 2000 protected areas, although this network only accounts for 18% of eu’s land area. this showed how vegetation plots were not uniformly distributed inside and outside natura 2000 areas (appendix: supplementary fig. s10). discussion biodiversity big data are being increasingly used to understand ecological patterns and monitor biodiversity trends (garcía-roselló et al. 2023). yet, these figure 4: trends of significative predictors with respect to the response variables of gams. n2k refers to the relative number of plots inside natura 2000 network, pop refers to the human population count. the plot a represents the estimated values of taxonomic bias (i.e., completeness of species richness) at each value of n2k, b the estimated values of spatial bias (i.e., nni) at each value of n2k, c the estimated values of temporal bias (i.e., pielou’s index) at each value of n2k, d the estimated values of spatial bias at each value of pop. the estimated values of the response variable are represented in the y-axis while the observed values of the spatial variable in the x-axis. the ”ticks” in the x-axis indicate the distribution of the values. finally, the line shows the estimated smooth and the point the partial residuals. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 65 large collections of opportunistic data come with intrinsic sources of bias, that require careful considerations (caldwell et al. 2024). here, we proposed a methodological framework and a set of useful metrics to quantify three different dimensions of bias (taxonomic, spatial, and temporal), as well as the underappreciated dimension of temporal uncertainty in biodiversity data, using vegetation plot data from the open-access database splotopen as an example. we found that the completeness of the species richness estimates varied across grid cells in europe, and vegetation plot data varied both in terms of their level of spatial clustering, and their level of temporal unevenenness at the european extent. in addition, the prevalence of one dimension of bias over the others also exhibited a non-uniform distribution, highlighting the presence of several hotspots of bias. in splotopen, the taxonomic bias varied unevenly across europe and accordingly to the plot size (appendix: supplementary fig. s5). as expected, we found that the observed species richness is significantly influenced by the sample size (chao and jost 2012), with high completeness occurring in grid cells with a high number of plots. however, the sampling completeness still presents some limitations. in particular, the species accumulation curve assumes that there is no spatial and temporal autocorrelation between the species occurrences (gotelli and colwell 2001; yang et al. 2013), and the values of the completeness of the species richness do not represent the degree of sampling of different habitat types (lobo et al. 2018). regardless, to address this constraint, a measure of the dark diversity, i.e., the species that are potentially present in a given community but have not yet been detected, can provide a more complete representation of the taxonomic sampling bias by associating it with the value of sampling completeness (carmona and pärtel 2021). despite these limitations, the use of sampling completeness is particularly common. its use appears in several applications such as for calculating the taxonomic gaps of species records at both multi (la sorte and somveille 2020) and single-taxa level (chesshire et al. 2023), and in assessing the efficacy of a sampling method (pelayo-villamil et al. 2018). here, we provided an example of how sampling completeness can be employed to depict the distribution of the taxonomic information gaps, based on the taxonomic coverage of the vascular plants, at the continental scale. as far as we know, the use of the nearest neighbor index to assess the spatial bias of raw data is not widespread (e.g., (geldmann et al. 2016; oliveira et al. 2016; hughes et al. 2021; rocchini et al. 2023)). in splotopen, we observed a high spatial bias, where most of the grid cells had a clustered distribution of plots. consequently, a high spatial bias in data collection can alter the current representation of community composition and environmental conditions, as well as the potential distribution of a species (michalcová et al. 2011; bazzichetto et al. 2023). however, the high clustering we found may depend on the environmental-based resampling of splotopen and possibly on the further filtering we applied to the database which, may have promoted the process of concentration of the plots in a restricted area. furthermore, the nni displayed values different from random expectations, suggesting a clustered pattern which, can have been determined by multiple factors, such as the sampling within the network of protected areas. although the sampling effort is the most commonly used method to represent the spatial bias of raw data, recent studies (sumner et al. 2019; boyd et al. 2021) have proposed the nni as a suitable index to measure and represent it. in this regard, combining the nni with the sampling effort can complement our understanding of spatial bias in its possible facets. here, we also represented the temporal bias, calculated using pielou’s index. in splotopen, the temporal bias follows a heterogeneous distribution across europe; high values (i.e., low bias) indicate a more uniform distribution of data across years. however, most of the studies tested the effect of irregular collection over time of raw data in ecological modelling or indices. examples are the temporal variation of the inventory completeness (stropp et al. 2016; ronquillo et al. 2020), the temporal change in species occupancy (powney et al. 2019; outhwaite et al. 2020), the temporal coverage of the species records (meyer et al. 2016; daru and rodriguez 2023), or the temporal variation of species distribution models due to biased sampling of species records under land-use change (bowler et al. 2022). here, we propose a new method to quantify temporal bias using a common metric employed in ecology, i.e., the pielou’s index, focusing on the distribution of the year of recording of the plots data rather than determining the impact of an uneven sampling over time of the species records. in our study, we tested three metrics commonly used in ecology to measure the bias of raw data at different dimensions. nevertheless, many other apmarchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 66 proaches exist to assess gaps and biases in biodiversity data and one does not exclude the others. some methods use directly raw data to evaluate the errors, others use predictions or estimations. for instance, ruete (2015) proposed an ignorance score representing the sampling effort of raw data; oliver et al. (2021) developed indicators of biodiversity data coverage and sampling effectiveness; moura and jetz (2021) analyzed one aspect of taxonomic and geographic knowledge gaps by modelling species discovery probability. eventually, it is even possible to face biases in raw data by a pre-processing procedure through their standardization and filtering to improve the accuracy of the inferences (ronquillo et al. 2023). in this study, we also provide a measure of the temporal uncertainty. to account for the wide uncertainties in the process of temporal decay, we quantified temporal precision using different negative exponential curves. with the method proposed, it is possible to appreciate different patterns of temporal uncertainty based on the exponents used. as lower z-values are used, the rate of decay of information increases. this allows us to identify areas where temporal uncertainty is always low and the information contained is consistently more precise. on the other side, it is possible to notice how areas that appeared to be more precise with higher z-values (e.g., -1/25) become highly uncertain with lower exponents. however, temporal precision is likely to decrease with different rates across different regions and vegetation types, due to many possible drivers of changes, such as anthropogenic pressures, climatic changes, or successional trajectories. this means that using the same function to model information decay across large areas is just an approximation since different contexts might be subjected to different drivers and intensities of change. in future research, it would be interesting to relate the rate of biodiversity information decay to rates of habitat loss and species assemblage turnover (jandt et al. 2022a,b). only a few studies paid attention to the temporal uncertainty of raw data (meyer et al. 2016; tessarolo et al. 2021; d’antraccoli et al. 2022). for instance, when creating a map of ignorance (rocchini et al. 2011) for species distribution models, tessarolo et al. (2021) calculated the temporal decay of the information provided by each occurrence record through a kernel gaussian function that increases the uncertainty for the increment in years since the last recording date. to our knowledge, no study has modelled temporal uncertainty using negative exponential functions. however, future research should investigate how to calibrate the most appropriate set of decay functions to model information loss across regions and vegetation types rather than arbitrarily choosing the exponent. it is most likely that the biases and uncertainties of the vegetation plots we found in splotopen reflect those of european vegetation archive (eva) (chytrý et al. 2016); in fact, the integration into eva database is necessary before european data can be contributed to splot. eva is an archive of multiple databases, and has continued accumulating, compared to the version splotopen was built upon. although many of the gaps in geographic coverage and representation of specific vegetation types might have been filled in the meantime (chytrý et al. 2014; sporbert et al. 2019), it is likely that some aspects of spatial, taxonomic or temporal bias remain. the resulting biases inevitably stem from errors embedded in individual contributing databases as well as challenges related to integrating data from databases with different objectives and adhering to diverse national and regional rules for structuring them. the relative number of plots inside the natura 2000 network and the human population count play a role in determining some facets of bias. ballesteros‐ mejia et al. (2013); girardello et al. (2019) showed how the sampling collection in protected areas increases the completeness of the species richness, as well as, ricci et al. (2024) demonstrated the effectiveness of natura 2000 protected area in increasing the species diversity. furthermore, we found that, as the number of plots inside the natura 2000 network increases, the distribution of the plots is more clustered (i.e., higher spatial bias). regarding the temporal evenness of the record collection, we found a non-linear relationship with the number of plots inside the natura 2000 network, with data collection being more even in time where plots are located both inside and outside the network of protected areas (fig. 4). in any case, the initial removal of vegetation plots in splotopen to maximize the representation of the environmental space may have altered the current representation of bias dimensions from that of the original splot database and their subsequent relationship with the spatial variables we considered. nevertheless, our outcomes show the strength of the presence of protected areas in shaping the three facets of bias and in influencing the sampling location of the vegetation plots (boakes et al. 2010). howevmarchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 67 er, the role played by each spatial variable is limited by its release year, which does not reflect the entire temporal period covered by the plots considered in the analysis. it is crucial to note that in many studies the taxonomic and spatial bias of biodiversity databases correlates with human population density and road density (ballesteros‐mejia et al. 2013; geldmann et al. 2016; mair and ruete 2016). this was partially observed in our models. in fact, only the spatial bias was significantly influenced by the human population count. this can probably depend on the initial environmental-based resampling of splotopen or by the possible masking effect that the sampling effort had on the other spatial variables in explaining the variability of the models. eventually, it is likely that human population count and presence of roads are better predictors of the spatial bias in sampling effort across grid cells, rather than predicting the level of clustering within cells (geldmann et al. 2016; mair and ruete 2016; oliveira et al. 2016). different facets of bias and uncertainty can be present in biodiversity databases because of many natural and anthropogenic factors that influence the choice of collecting data in a specific place and at a specific time. not accounting for these sources of errors in biodiversity data could create knowledge shortfalls and hinder our capacity to monitor real trends in biodiversity and consequently develop effective conservation strategies. it is, therefore, necessary to take into consideration the different facets of bias and uncertainty in biodiversity data by incorporating a routine to check for their presence. here, we proposed and tested a methodological framework that can be reproduced and applied at different spatial scales (local, ecoregions, biomes, global) and for other databases such as vegetation plots, or simple occurrence data, as those contained in gbif (gbif, 2020). we argue that our framework can be useful for quantifying, making visible, and possibly addressing different sources of bias and uncertainty transparently both when creating a new biodiversity database, and when highlighting priorities for gap-filling in existing ones. for instance, it can be helpful to point out where more actions to fix gaps and sources of errors could be allocated and to provide guidance to data users on how to avoid falling into potential pitfalls and drawing biased inferences. declaration of conflict of interest the authors have declared no competing interests exist. data availability statement the data that support the findings of this study are openly available in zenodo.4 acknowledgments f.m.s. gratefully acknowledges financial support from the italian ministry of university and research, within the rita-levi montalcini 2019 program. r.t., r.c.g., a.c., j.i., d.s. and d.r. were supported by the european union—nextgenerationeu, under the national recovery and resilience plan (nrrp), project title “national biodiversity future center -nbfc” (project code cn_00000033) cup j33c22001190001. d.r. was also partially funded by the horizon europe project b3-biodiversity building blocks for policy (grant agreement 101059592). d.r., v.m., and p.s. were partially funded by the horizon europe project earthbridge (grant agreement 101079310). views and opinions expressed are, however, those of the authors only, and do not necessarily reflect those of the european union or the european research council executive agency. neither the european union nor the granting authority can be held responsible for them. author contributions e.m. and m.l. equally contributed at developing the idea and performed the formal analyses. they also wrote the first version of the manuscript and headed the review editing. f.m.s., e.t., d.r. provided substantial input on the conceptualization and the analytical and methodological framework. j.l., d.d.r., r.t. provided methodological revision and gave considerable suggestions on writing – review and editing. g.b., r.c.g., a.c., g.m.f., l.g., q.g., j.i., m.m., v.m., d.s., p.s., p.z. contributed to writing – review and editing. literature cited amatulli, g., s. domisch, m.-n. tuanmu, b. parmentier, a. ranipeta, j. malczyk, and w. jetz. 2018. a suite of global, cross-scale topographic variables for environmental and biodiversity modeling. sci data 5:180040. baddeley, a., e. rubak, and r. turner. 2016. spatial point patterns: methodology and applications with r. crc press, boca raton london new york. ball-damerow, j. e., l. brenskelle, n. barve, p. s. soltis, p. sierwald, r. bieler, r. lafrance, a. h. ariño, and r. p. guralnick. 2019. research applications of primary biodiversity 4 https://zenodo.org/doi/10.5281/zenodo.12179384. https://zenodo.org/doi/10.5281/zenodo.12179384 marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 68 databases in the digital age. plos one 14:e0215794. ballesteros‐mejia, l., i. j. kitching, w. jetz, p. nagel, and j. beck. 2013. mapping the biodiversity of tropical insects: species richness and inventory completeness of a frican sphingid moths. global ecology and biogeography 22:586–595. barajas‐barbosa, m. p., p. weigelt, m. k. borregaard, g. keppel, and h. kreft. 2020. environmental heterogeneity dynamics drive plant diversity on oceanic islands. journal of biogeography 47:2248–2260. bazzichetto, m., j. lenoir, d. da re, e. tordoni, d. rocchini, m. malavasi, v. barták, and m. g. sperandii. 2023. sampling strategy matters to accurately estimate response curves’ parameters in species distribution models. global ecol biogeogr 32:1717–1729. bertrand, r., j. lenoir, c. piedallu, g. riofrío-dillon, p. de ruffray, c. vidal, j.-c. pierrat, and j.-c. gégout. 2011. changes in plant community composition lag behind climate warming in lowland forests. nature 479:517–520. boakes, e. h., p. j. k. mcgowan, r. a. fuller, d. chang-qing, n. e. clark, k. o’connor, and g. m. mace. 2010. distorted views of biodiversity: spatial and temporal bias in species occurrence data. plos biol 8:e1000385. boitani, l., l. maiorano, d. baisero, a. falcucci, p. visconti, and c. rondinini. 2011. what spatial data do we need to develop global mammal conservation strategies? phil. trans. r. soc. b 366:2623–2632. bowler, d. e., c. t. callaghan, n. bhandari, k. henle, m. benjamin barth, c. koppitz, r. klenke, m. winter, f. jansen, h. bruelheide, and a. bonn. 2022. temporal trends in the spatial bias of species occurrence records. ecography 2022:e06219. boyd, r. j., g. d. powney, c. carvell, and o. l. pescott. 2021. occassess: an r package for assessing potential biases in species occurrence data. ecology and evolution 11:16177– 16187. bruelheide, h., j. dengler, b. jiménez‐alfaro, o. purschke, s. m. hennekens, m. chytrý, v. d. pillar, f. jansen, j. kattge, b. sandel, i. aubin, i. biurrun, r. field, s. haider, u. jandt, j. lenoir, r. k. peet, g. peyre, f. m. sabatini, m. schmidt, f. schrodt, m. winter, s. aćić, e. agrillo, m. alvarez, d. ambarlı, p. angelini, i. apostolova, m. a. s. arfin khan, e. arnst, f. attorre, c. baraloto, m. beckmann, c. berg, y. bergeron, e. bergmeier, a. d. bjorkman, v. bondareva, p. borchardt, z. botta‐dukát, b. boyle, a. breen, h. brisse, c. byun, m. r. cabido, l. casella, l. cayuela, t. černý, v. chepinoga, j. csiky, m. curran, r. ćušterevska, z. dajić stevanović, e. de bie, p. de ruffray, m. de sanctis, p. dimopoulos, s. dressler, r. ejrnæs, m. a. e. m. el‐sheikh, b. enquist, j. ewald, j. fagúndez, m. finckh, x. font, e. forey, g. fotiadis, i. garcía‐mijangos, a. l. de gasper, v. golub, a. g. gutierrez, m. z. hatim, t. he, p. higuchi, d. holubová, n. hölzel, j. homeier, a. indreica, d. işık gürsoy, s. jansen, j. janssen, b. jedrzejek, m. jiroušek, n. jürgens, z. kącki, a. kavgacı, e. kearsley, m. kessler, i. knollová, v. kolomiychuk, a. korolyuk, m. kozhevnikova, ł. kozub, d. krstonošić, h. kühl, i. kühn, a. kuzemko, f. küzmič, f. landucci, m. t. lee, a. levesley, c. li, h. liu, g. lopez‐gonzalez, t. lysenko, a. macanović, p. mahdavi, p. manning, c. marcenò, v. martynenko, m. mencuccini, v. minden, j. e. moeslund, m. moretti, j. v. müller, j. munzinger, ü. niinemets, m. nobis, j. noroozi, a. nowak, v. onyshchenko, g. e. overbeck, w. a. ozinga, a. pauchard, h. pedashenko, j. peñuelas, a. pérez‐ haase, t. peterka, p. petřík, o. l. phillips, v. prokhorov, v. rašomavičius, r. revermann, j. rodwell, e. ruprecht, s. rūsiņa, c. samimi, j. h. j. schaminée, u. schmiedel, j. šibík, u. šilc, ž. škvorc, a. smyth, t. sop, d. sopotlieva, b. sparrow, z. stančić, j. svenning, g. swacha, z. tang, i. tsiripidis, p. d. turtureanu, e. uğurlu, d. uogintas, m. valachovič, k. a. vanselow, y. vashenyak, k. vassilev, e. vélez‐martin, r. venanzoni, a. c. vibrans, c. violle, r. virtanen, h. von wehrden, v. wagner, d. a. walker, d. wana, e. weiher, k. wesche, t. whitfeld, w. willner, s. wiser, t. wohlgemuth, s. yamalov, g. zizka, and a. zverev. 2019. splot – a new tool for global vegetation analyses. j vegetation science 30:161–186. caldwell, i. r., j.-p. a. hobbs, b. w. bowen, p. f. cowman, j. d. dibattista, j. l. whitney, p. a. ahti, r. belderok, s. canfield, r. r. coleman, m. iacchei, e. c. johnston, i. knapp, e. m. nalley, t. m. staeudle, and á. j. láruson. 2024. global trends and biases in biodiversity conservation research. cell reports sustainability 1:100082. carmona, c. p., and m. pärtel. 2021. estimating probabilistic site‐specific species pools and dark diversity from co‐occurrence data. global ecol. biogeogr. 30:316–326. chao, a., and l. jost. 2012. coverage‐based rarefaction and extrapolation: standardizing samples by completeness rather than size. ecology 93:2533–2547. chao, a., y. kubota, d. zelený, c. chiu, c. li, b. kusumoto, m. yasuhara, s. thorn, c. wei, m. j. costello, and r. k. colwell. 2020. quantifying sample completeness and comparing diversities among assemblages. ecological research 35:292–314. chesshire, p. r., e. e. fischer, n. j. dowdy, t. l. griswold, a. c. hughes, m. c. orr, j. s. ascher, l. m. guzman, k. j. hung, n. s. cobb, and l. m. mccabe. 2023. completeness analysis for over 3000 united states bee species identifies persistent data gap. ecography 2023:e06584. chiarucci, a., r. m. di biase, l. fattorini, m. marcheselli, and c. pisani. 2018. joining the incompatible: exploiting purposive lists for the sample-based estimation of species richness. ann. appl. stat. 12. chytrý, m., s. m. hennekens, b. jiménez‐alfaro, i. knollová, j. dengler, f. jansen, f. landucci, j. h. j. schaminée, s. aćić, e. agrillo, d. ambarlı, p. angelini, i. apostolova, f. attorre, c. berg, e. bergmeier, i. biurrun, z. botta‐dukát, h. brisse, j. a. campos, l. carlón, a. čarni, l. casella, j. csiky, r. ćušterevska, z. dajić stevanović, j. danihelka, e. de bie, p. de ruffray, m. de sanctis, w. b. dickoré, p. dimopoulos, d. dubyna, t. dziuba, r. ejrnæs, n. ermakov, j. ewald, g. fanelli, f. fernández‐gonzález, ú. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 69 fitzpatrick, x. font, i. garcía‐mijangos, r. g. gavilán, v. golub, r. guarino, r. haveman, a. indreica, d. işık gürsoy, u. jandt, j. a. m. janssen, m. jiroušek, z. kącki, a. kavgacı, m. kleikamp, v. kolomiychuk, m. krstivojević ćuk, d. krstonošić, a. kuzemko, j. lenoir, t. lysenko, c. marcenò, v. martynenko, d. michalcová, j. e. moeslund, v. onyshchenko, h. pedashenko, a. pérez‐haase, t. peterka, v. prokhorov, v. rašomavičius, m. p. rodríguez‐rojo, j. s. rodwell, t. rogova, e. ruprecht, s. rūsiņa, g. seidler, j. šibík, u. šilc, ž. škvorc, d. sopotlieva, z. stančić, j. svenning, g. swacha, i. tsiripidis, p. d. turtureanu, e. uğurlu, d. uogintas, m. valachovič, y. vashenyak, k. vassilev, r. venanzoni, r. virtanen, l. weekes, w. willner, t. wohlgemuth, and s. yamalov. 2016. european vegetation archive (eva): an integrated database of european vegetation plots. applied vegetation science 19:173–180. chytrý, m., l. tichý, s. m. hennekens, and j. h. j. schaminée. 2014. assessing vegetation change using vegetation‐plot databases: a risky business. applied vegetation science 17:32–41. clark, p. j., and f. c. evans. 1954. distance to nearest neighbor as a measure of spatial relationships in populations. ecology 35:445–453. colli‐silva, m., m. reginato, a. cabral, r. c. forzza, j. r. pirani, and t. n. d. c. vasconcelos. 2020. evaluating shortfalls and spatial accuracy of biodiversity documentation in the atlantic forest, the most diverse and threatened brazilian phytogeographic domain. taxon 69:567–577. d’antraccoli, m., g. bedini, and l. peruzzi. 2022. maps of relative floristic ignorance and virtual floristic lists: an r package to incorporate uncertainty in mapping and analysing biodiversity data. ecological informatics 67:101512. daru, b. h., and j. rodriguez. 2023. mass production of unvouchered records fails to represent global biodiversity patterns. nat ecol evol 7:816–831. di marco, m., t. d. harwood, a. j. hoskins, c. ware, s. l. l. hill, and s. ferrier. 2019. projecting impacts of global climate and land‐use scenarios on plant biodiversity using compositional‐turnover modelling. global change biology 25:2763–2778. enquist, b. j., r. condit, r. k. peet, m. schildhauer, and b. m. thiers. 2016. cyberinfrastructure for an integrated botanical information network to investigate the ecological impacts of global climate change on plant biodiversity. peerj preprints. gábor, l., v. moudrý, v. lecours, m. malavasi, v. barták, m. fogl, p. šímová, d. rocchini, and t. václavík. 2020. the effect of positional error on fine scale species distribution models increases for specialist species. ecography 43:256– 269. garcía-roselló, e., j. gonzález-dacosta, and j. m. lobo. 2023. the biased distribution of existing information on biodiversity hinders its use in conservation, and we need an integrative approach to act urgently. biological conservation 283:110118. gbif: the global biodiversity information facility (year) what is gbif?. available from https://www.gbif.org/what-is-gbif [13 january 2020] geldmann, j., j. heilmann‐clausen, t. e. holm, i. levinsky, b. markussen, k. olsen, c. rahbek, and a. p. tøttrup. 2016. what determines spatial bias in citizen science? exploring four recording schemes with different proficiency requirements. diversity and distributions 22:1139–1149. girardello, m., a. chapman, r. dennis, l. kaila, p. a. v. borges, and a. santangeli. 2019. gaps in butterfly inventory data: a global analysis. biological conservation 236:289–295. gotelli, n. j., and r. k. colwell. 2001. quantifying biodiversity: procedures and pitfalls in the measurement and comparison of species richness. ecology letters 4:379–391. hortal, j., f. de bello, j. a. f. diniz-filho, t. m. lewinsohn, j. m. lobo, and r. j. ladle. 2015. seven shortfalls that beset large-scale knowledge of biodiversity. annu. rev. ecol. evol. syst. 46:523–549. hsieh, t. c., k. h. ma, and a. chao. 2016. inext: an r package for rarefaction and extrapolation of species diversity ( h ill numbers). methods ecol evol 7:1451–1456. hughes, a. c., m. c. orr, k. ma, m. j. costello, j. waller, p. provoost, q. yang, c. zhu, and h. qiao. 2021. sampling biases shape our view of the natural world. ecography 44:1259–1269. ipbes. 2019. global assessment report on biodiversity and ecosystem services of the intergovernmental science-policy platform on biodiversity and ecosystem services. [object object]. jandt, u., h. bruelheide, c. berg, m. bernhardt-römermann, v. blüml, f. bode, j. dengler, m. diekmann, h. dierschke, i. doerfler, u. döring, s. dullinger, w. härdtle, s. haider, t. heinken, p. horchler, f. jansen, t. kudernatsch, g. kuhn, m. lindner, s. matesanz, k. metze, s. meyer, f. müller, n. müller, t. naaf, c. peppler-lisbach, p. poschlod, c. roscher, g. rosenthal, s. b. rumpf, w. schmidt, j. schrautzer, a. schwabe, p. schwartze, t. sperle, n. stanik, h.-g. stroh, c. storm, w. voigt, a. von heßberg, g. von oheimb, e.-r. wagner, u. wegener, k. wesche, b. wittig, and m. wulf. 2022a. resurveygermany: vegetation-plot time-series over the past hundred years in germany. sci data 9:631. jandt, u., h. bruelheide, f. jansen, a. bonn, v. grescho, r. a. klenke, f. m. sabatini, m. bernhardt-römermann, v. blüml, j. dengler, m. diekmann, i. doerfler, u. döring, s. dullinger, s. haider, t. heinken, p. horchler, g. kuhn, m. lindner, k. metze, n. müller, t. naaf, c. peppler-lisbach, p. poschlod, c. roscher, g. rosenthal, s. b. rumpf, w. schmidt, j. schrautzer, a. schwabe, p. schwartze, t. sperle, n. stanik, c. storm, w. voigt, u. wegener, k. wesche, b. wittig, and m. wulf. 2022b. more losses than gains during one century of plant biodiversity change in germany. nature 611:512–518. la sorte, f. a., and m. somveille. 2020. survey completeness of a global citizen‐science database of bird occurrence. ecogmarchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 70 raphy 43:34–43. lobo, j. m., j. hortal, j. l. yela, a. millán, d. sánchez-fernández, e. garcía-roselló, j. gonzález-dacosta, j. heine, l. gonzález-vilas, and c. guisande. 2018. knowbr: an application to map the geographical variation of survey effort and identify well-surveyed areas from biodiversity databases. ecological indicators 91:241–248. mair, l., and a. ruete. 2016. explaining spatial variation in the recording effort of citizen science data across multiple taxa. plos one 11:e0147796. meijer, j. r., m. a. j. huijbregts, k. c. g. j. schotten, and a. m. schipper. 2018. global patterns of current and future road infrastructure. environ. res. lett. 13:064006. meyer, c., h. kreft, r. guralnick, and w. jetz. 2015. global priorities for an effective information basis of biodiversity distributions. nat commun 6:8221. meyer, c., p. weigelt, and h. kreft. 2016. multidimensional biases, gaps and uncertainties in global plant occurrence information. ecology letters 19:992–1006. michalcová, d., s. lvončík, m. chytrý, and o. hájek. 2011. bias in vegetation databases? a comparison of stratified-random and preferential sampling: stratified-random and preferential sampling. journal of vegetation science 22:281–291. monsarrat, s., a. f. boshoff, and g. i. h. kerley. 2019. accessibility maps as a tool to predict sampling bias in historical biodiversity occurrence records. ecography 42:125–136. moura, m. r., and w. jetz. 2021. shortfalls and opportunities in terrestrial vertebrate species discovery. nat ecol evol 5:631–639. newbold, t., l. n. hudson, s. l. l. hill, s. contu, i. lysenko, r. a. senior, l. börger, d. j. bennett, a. choimes, b. collen, j. day, a. de palma, s. díaz, s. echeverria-londoño, m. j. edgar, a. feldman, m. garon, m. l. k. harrison, t. alhusseini, d. j. ingram, y. itescu, j. kattge, v. kemp, l. kirkpatrick, m. kleyer, d. l. p. correia, c. d. martin, s. meiri, m. novosolov, y. pan, h. r. p. phillips, d. w. purves, a. robinson, j. simpson, s. l. tuck, e. weiher, h. j. white, r. m. ewers, g. m. mace, j. p. w. scharlemann, and a. purvis. 2015. global effects of land use on local terrestrial biodiversity. nature 520:45–50. newling, b. e. 1969. the spatial variation of urban population densities. geographical review 59:242. oliveira, u., a. p. paglia, a. d. brescovit, c. j. b. de carvalho, d. p. silva, d. t. rezende, f. s. f. leite, j. a. n. batista, j. p. p. p. barbosa, j. r. stehmann, j. s. ascher, m. f. de vasconcelos, p. de marco, p. löwenberg‐neto, p. g. dias, v. g. ferro, and a. j. santos. 2016. the strong influence of collection bias on biodiversity knowledge shortfalls of b razilian terrestrial biodiversity. diversity and distributions 22:1232–1244. oliver, r. y., c. meyer, a. ranipeta, k. winner, and w. jetz. 2021. global and national trends, gaps, and opportunities in documenting and monitoring species distributions. plos biol 19:e3001336. outhwaite, c. l., r. d. gregory, r. e. chandler, b. collen, and n. j. b. isaac. 2020. complex long-term biodiversity change among invertebrates, bryophytes and lichens. nat ecol evol 4:384–392. pelayo-villamil, p., c. guisande, a. manjarrés-hernández, l. f. jiménez, c. granado-lorencio, e. garcía-roselló, j. gonzález-dacosta, j. heine, l. gonzález-vilas, and j. m. lobo. 2018. completeness of national freshwater fish species inventories around the world. biodivers conserv 27:3807–3817. pielou, e. c. 1966. the measurement of diversity in different types of biological collections. journal of theoretical biology 13:131–144. powney, g. d., c. carvell, m. edwards, r. k. a. morris, h. e. roy, b. a. woodcock, and n. j. b. isaac. 2019. widespread losses of pollinating insects in britain. nat commun 10:1018. ricci, l., m. di musciano, f. m. sabatini, a. chiarucci, p. zannini, r. c. gatti, c. beierkuhnlein, a. walentowitz, a. lawrence, a. r. frattaroli, and s. hoffmann. 2024. a multitaxonomic assessment of natura 2000 effectiveness across european biogeographic regions. conservation biology 38:e14212. rocchini, d., j. hortal, s. lengyel, j. m. lobo, a. jiménez-valverde, c. ricotta, g. bacaro, and a. chiarucci. 2011. accounting for uncertainty when mapping species distributions: the need for maps of ignorance. progress in physical geography: earth and environment 35:211–226. rocchini, d., e. tordoni, e. marchetto, m. marcantonio, a. m. barbosa, m. bazzichetto, c. beierkuhnlein, e. castelnuovo, r. c. gatti, a. chiarucci, l. chieffallo, d. da re, m. di musciano, g. m. foody, l. gabor, c. x. garzon-lopez, a. guisan, t. hattab, j. hortal, w. e. kunin, f. jordán, j. lenoir, s. mirri, v. moudrý, b. naimi, j. nowosad, f. m. sabatini, a. h. schweiger, p. šímová, g. tessarolo, p. zannini, and m. malavasi. 2023. a quixotic view of spatial bias in modelling the distribution of species and their diversity. npj biodivers 2:10. ronquillo, c., f. alves-martins, v. mazimpaka, t. sobral-souza, b. vilela-silva, n. g. medina, and j. hortal. 2020. assessing spatial and temporal biases and gaps in the publicly available distributional information of iberian mosses. bdj 8:e53474. ronquillo, c., j. stropp, n. g. medina, and j. hortal. 2023. exploring the impact of data curation criteria on the observed geographical distribution of mosses. ecology and evolution 13:e10786. ruete, a. 2015. displaying bias in sampling effort of data accessed from biodiversity databases using ignorance maps. bdj 3:e5361. sabatini, f. m., j. lenoir, h. bruelheide, and the splot consortium. 2021a. splotopen – an environmentally-balanced, open-access, global dataset of vegetation plots. german centre for integrative biodiversity research (idiv) halle-jena-leipzig. sabatini, f. m., j. lenoir, t. hattab, e. a. arnst, m. chytrý, j. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 71 dengler, p. de ruffray, s. m. hennekens, u. jandt, f. jansen, b. jiménez‐alfaro, j. kattge, a. levesley, v. d. pillar, o. purschke, b. sandel, f. sultana, t. aavik, s. aćić, a. t. r. acosta, e. agrillo, m. alvarez, i. apostolova, m. a. s. arfin khan, l. arroyo, f. attorre, i. aubin, a. banerjee, m. bauters, y. bergeron, e. bergmeier, i. biurrun, a. d. bjorkman, g. bonari, v. bondareva, j. brunet, a. čarni, l. casella, l. cayuela, t. černý, v. chepinoga, j. csiky, r. ćušterevska, e. de bie, a. l. de gasper, m. de sanctis, p. dimopoulos, j. dolezal, t. dziuba, m. a. e. m. el‐sheikh, b. enquist, j. ewald, f. fazayeli, r. field, m. finckh, s. gachet, a. galán‐de‐mera, e. garbolino, h. gholizadeh, m. giorgis, v. golub, i. g. alsos, j. grytnes, g. r. guerin, a. g. gutiérrez, s. haider, m. z. hatim, b. hérault, g. hinojos mendoza, n. hölzel, j. homeier, w. hubau, a. indreica, j. a. m. janssen, b. jedrzejek, a. jentsch, n. jürgens, z. kącki, j. kapfer, d. n. karger, a. kavgacı, e. kearsley, m. kessler, l. khanina, t. killeen, a. korolyuk, h. kreft, h. s. kühl, a. kuzemko, f. landucci, a. lengyel, f. lens, d. v. lingner, h. liu, t. lysenko, m. d. mahecha, c. marcenò, v. martynenko, j. e. moeslund, a. monteagudo mendoza, l. mucina, j. v. müller, j. munzinger, a. naqinezhad, j. noroozi, a. nowak, v. onyshchenko, g. e. overbeck, m. pärtel, a. pauchard, r. k. peet, j. peñuelas, a. pérez‐haase, t. peterka, p. petřík, g. peyre, o. l. phillips, v. prokhorov, v. rašomavičius, r. revermann, g. rivas‐ torres, j. s. rodwell, e. ruprecht, s. rūsiņa, c. samimi, m. schmidt, f. schrodt, h. shan, p. shirokikh, j. šibík, u. šilc, p. sklenář, ž. škvorc, b. sparrow, m. g. sperandii, z. stančić, j. svenning, z. tang, c. q. tang, i. tsiripidis, k. a. vanselow, r. vásquez martínez, k. vassilev, e. vélez‐ martin, r. venanzoni, a. c. vibrans, c. violle, r. virtanen, h. von wehrden, v. wagner, d. a. walker, d. m. waller, h. wang, k. wesche, t. j. s. whitfeld, w. willner, s. k. wiser, t. wohlgemuth, s. yamalov, m. zobel, and h. bruelheide. 2021b. splotopen – an environmentally balanced, open‐access, global dataset of vegetation plots. global ecol. biogeogr. 30:1740–1764. sporbert, m., h. bruelheide, g. seidler, p. keil, u. jandt, g. austrheim, i. biurrun, j. a. campos, a. čarni, m. chytrý, j. csiky, e. de bie, j. dengler, v. golub, j. grytnes, a. indreica, f. jansen, m. jiroušek, j. lenoir, m. luoto, c. marcenò, j. e. moeslund, a. pérez‐haase, s. rūsiņa, v. vandvik, k. vassilev, and e. welk. 2019. assessing sampling coverage of species distribution in biodiversity databases. j vegetation science 30:620–632. stein, a., k. gerstner, and h. kreft. 2014. environmental heterogeneity as a universal driver of species richness across taxa, biomes and spatial scales. ecology letters 17:866–880. stevens, f. r., a. e. gaughan, c. linard, and a. j. tatem. 2015. disaggregating census data for population mapping using random forests with remotely-sensed and ancillary data. plos one 10:e0107042. stropp, j., r. j. ladle, t. emilio, t. lessa, and j. hortal. 2022. taxonomic uncertainty and the challenge of estimating global species richness. journal of biogeography 49:1654– 1656. stropp, j., r. j. ladle, a. c. m. malhado, j. hortal, j. gaffuri, w. h. temperley, j. olav skøien, and p. mayaux. 2016. mapping ignorance: 300 years of collecting flowering plants in africa. global ecol. biogeogr. 25:1085–1096. sumner, s., p. bevan, a. g. hart, and n. j. b. isaac. 2019. mapping species distributions in 2 weeks using citizen science. insect conserv diversity 12:382–388. tessarolo, g., r. j. ladle, j. m. lobo, t. f. rangel, and j. hortal. 2021. using maps of biogeographical ignorance to reveal the uncertainty in distributional data hidden in species distribution models. ecography 44:1743–1755. tessarolo, g., r. ladle, t. rangel, and j. hortal. 2017. temporal degradation of data limits biodiversity research. ecology and evolution 7:6863–6870. tessarolo, g., t. f. rangel, m. b. araújo, and j. hortal. 2014. uncertainty associated with survey design in species distribution models. diversity and distributions 20:1258–1269. troudet, j., p. grandcolas, a. blin, r. vignes-lebbe, and f. legendre. 2017. taxonomic bias in biodiversity data and societal preferences. sci rep 7:9132. turbelin, a. j., b. d. malamud, and r. a. francis. 2017. mapping the global state of invasive alien species: patterns of invasion and policy responses. global ecol. biogeogr. 26:78–92. walther, b. a., and j. l. moore. 2005. the concepts of bias, precision and accuracy, and their use in testing the performance of species richness estimators, with a literature review of estimator performance. ecography 28:815–829. wüest, r. o., n. e. zimmermann, d. zurell, j. m. alexander, s. a. fritz, c. hof, h. kreft, s. normand, j. s. cabral, e. szekely, w. thuiller, m. wikelski, and d. n. karger. 2020. macroecology in the age of big data – where to go from here? journal of biogeography 47:1–12. xu, g., l. jiao, m. yuan, t. dong, b. zhang, and c. du. 2019. how does urban population density decline over time? an exponential model for chinese cities with international comparisons. landscape and urban planning 183:59–67. yang, w., k. ma, and h. kreft. 2014. environmental and socio‐economic factors shaping the geography of floristic collections in c hina. global ecology and biogeography 23:1284–1292. yang, w., k. ma, and h. kreft. 2013. geographical sampling bias in a large distributional database and its effects on species richness–environment models. journal of biogeography 40:1415–1426. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 72 appendix supplementary methods trivariate map we represented in a trivariate map, which is a graphic representation that shows the relationship between three variables at once, the three dimensions of bias. we used the functions provided by “tricolore” r package for creating the map. the variables selected for the trivariate map were the nearest neighbour index, the completeness of the species richness and the pielou's evenness. the variables were first standardized to have a mean of zero and a standard deviation of one, then rescaled to a 0-1 range and subsequently mapped over our study area. we also removed all grid cells that had missing values for at least one facet of bias. figure s1: negative exponential curves fitting using three different exponents, i.e., z1=-1, z2=-1/5 and z3=-1/25. the trivariate map highlighted those area where the prevalence of one type of bias prevail to the others. the grid cells with different colours from those of vertices (e.g., brown) tend to be more and more influenced uniformly by the three dimensions of bias as the colour approaches the center of the triangle. facets of bias single grid-based map of each metric of bias with unstandardized grids number. temporal uncertainty here we represented the temporal uncertainty by assessing the difference between the most recent year in the database (2014) and the year of each record. the higher the difference, the more uncertainty we have. the uncertainty per grid cell was calculated as the median value of the differences between the year of the most recent surveyed plot (i.e., 2014) and the date of recording of the ith plot within the grid cell. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 73 figure s2: grid-based trivariate map of taxonomic bias (i.e., completeness of species richness, abbrv. legend comp), spatial bias (i.e., nni, abbrv. legend nni) and temporal bias (i.e., pielou's evenness, abbrv. legend j). each grid cell has a spatial resolution of 39.5 km. the highest sampling completeness is represented by light green color (low taxonomic bias), the highest temporal evenness by light blue color (low temporal bias), the highest uniform distribution of the plots (low spatial bias) by pink color. figure s3: completeness of the species richness per grid cell of 39.5 km. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 74 figure s4: a) completeness of the species richness and b) logarithm to base 10 of the number of plots per grid cell of 39.5 km with standardized number of grid cells. figure s5: a) completeness of the species richness per grid cell of 39.5km including only the vegetation plots with area less than or equal to 150 m2. b) completeness of the species richness for plots with an area greater than 150 m2. the area size was determined by relying on sabatini et al. 20221 and to have a comparable number of plots belonging to the two categories. 1sabatini, f. m., jiménez-alfaro, b., jandt, u., chytrý, m., field, r., kessler, m., ... & bruelheide, h. (2022). global patterns of vascular plant alpha diversity. nature communications, 13(1), 4683. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 75 figure s6: nni per grid cell of 39.5 km. nni values greater than 1 indicate a random distribution of plots within a grid cell, while values less than 1 indicate a clustered distribution. figure s7: spatial distribution of the vegetation plots per grid cell. a represents the map of nni and b represents the map of nni with a standardized effect size. nni values greater than 1 indicate a random distribution of plots within a grid cell, while values less than 1 indicate a clustered distribution. there is no value of nni with a standardized effect size between – 0.8 and 0.8, meaning that the effect size is large. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 76 figure s8: pielou’s index per grid cell of 39.5 km. figure s9: logarithm to base 10 of the number of plots per grid cell of 39.5 km. marchetto et al. – addressing multiple facets of bias and uncertainty in continental-scale biodiversity databases 77 figure s10: map of the relative number of plots in the natura 2000 network per grid cell. figure s11: median of the temporal distance between the most recent year (i.e., 2014) and the year of each record per grid cell of 39.5 km. biodiversity informatics, 19, 2025, pp. 61-84 61 ecological niche modeling applications to infectious diseases shariful islam1,2, mariana castaneda-guzman1, diego soler-tovar3, and luis e. escobar1,2,4* 1department of fish and wildlife conservation, virginia tech, blacksburg, va, usa 2global change center, virginia tech, blacksburg, va, usa 3faculty of agricultural sciences, universidad de la salle, bogotá, colombia 4center for emerging zoonotic and arthropod-borne pathogens, virginia tech, blacksburg, va, usa abstract. ecological niche modeling (enm) is a widely used analytical approach for predicting species distributions and has been applied to study the spatial epidemiology of infectious diseases. nevertheless, research evaluating the key components and assumptions of enm in disease systems remains limited, raising concerns about its robustness, reproducibility, and transparency. to address this limitation, we conducted a systematic review and evaluated articles on enm applications to infectious diseases between 2020 and 2022. we reviewed 78 articles to extract information following a standard protocol for reporting enm analysis and summarized the information for each component (e.g., study subject, location, duration). the spatial extent of study areas varied from village to global scales, temporal duration ranged from 1 to 101 years, and the organismal levels ranged from individuals (57.7%) to populations (33.3%). less frequently reported components included temporal autocorrelation tests (2.66%), algorithmic uncertainty (28.21%), temporal resolution (35.90%), background data selection (44.87%), coordinate reference system (41.02%), model performance from validation data (46.15%), and model averaging (20.51%). our findings highlight a lack of consistency and transparency in disease ecology and disease biogeography studies, which may lead to misleading enm applications in spatial epidemiology. researchers and reviewers applying enm to disease systems should clearly report key modeling components to ensure biologically sound outputs. this article identified trends and gaps in reporting enm protocols for mapping disease transmission risk. keywords: biogeography, disease, ecological niche, epidemiology, reproducibility, spatial * corresponding author: address: 1015 life science circle, blacksburg, virginia, 24061, united states. escobar1@vt.edu. orcid: https://orcid.org/0000-0001-5735-2750 introduction emerging infectious diseases are increasing in frequency and represent a major threat to global public health and global economies (dobson et al. 2020). human-animal interaction is considered an important driver of disease emergence along with climate change, biodiversity loss, and socio-economic conditions (keesing and ostfeld 2024). nevertheless, our ability to understand and accurately forecast disease emergence and spread remains limited (escobar and craft 2016). one primary reason for inaccurate disease forecasting is the limited standards in data availability, model parameterization, and the inherent stochasticity (i.e., randomness or unpredictability) of disease transmission (escobar 2020; peterson 2014). disease ecology and biogeography can be used to understand why diseases emerge in some areas but not others (peterson 2008). the “ecological niche”, defined as the set of favorable conditions (biotic and abiotic) allowing an organism to persist and disperse in the long term, is a key concept in disease ecology and disease biogeography (escobar 2020; peterson 2014). the ecological niche of pathogens, vectors, hosts, or disease transmission events is complex considering host-pathogen relationships. in some cases, the distribution of a pathogen is shaped by the distribution of its host. nevertheless, there is also evidence that some pathogens do not follow the niches of their hosts and instead exhibit their own distinct ecological niche (maher et al. 2010; astorga et al. 2018). to estimate or charactermailto:escobar1@vt.edu https://orcid.org/0000-0001-5735-2750 shariful islam et al. – ecological niche modeling applications to infectious diseases 62 ize the ecological niche of an organism causing an infectious disease, researchers commonly use ecological niche modeling (enm). for example, romero-alvarez et al. (2023) assessed the potential spread of oropouche virus in humans, and escobar et al. (2017) identified areas of transmission risk for heterosporis, a fish disease caused by the microsporidian parasite heterosporis sutherlandae. a complete enm framework consists of five major sections (overview/conceptualization, data, model fitting, assessment, and prediction), each containing components essential for building robust, transparent, and reproducible models (zurell et al. 2020). with increasing access to enm training materials, tools, and data, the implementation of enm in the study of diseases is becoming more common (frans and liu 2024). for instance, a comprehensive open-access online course facilitates the learning and development of enm for biodiversity and disease (peterson et al. 2019; peterson and ingenloff 2016). as the use of enm in disease ecology and biogeography continues to expand (escobar and morand 2021), it is increasingly important to assess whether studies are adhering to best practices and reporting important modeling components transparently. nevertheless, researchers may overlook key modeling components (e.g., study objects, location, and duration) and assumptions during enm implementations (escobar 2020), which can lead to an untrustworthy model and the misinterpretation of infectious diseases systems (araújo et al. 2019; peterson 2014; soley-guardia et al. 2024). for instance, a study conducted by brito-hoyos et al. (2013) aiming to map the risk of bat-borne rabies outbreaks in colombia employed a climate-based enm. the authors, however, skipped critical steps of enm (escobar and peterson 2013). one of them was the use of all available occurrence points for model calibration without considering spatial clustering and sampling bias. additionally, brito-hoyos et al. (2013) used 15 climatic variables for their model building without considering the multicollinearity among the variables and spatial lags. later, escobar and peterson (2013) revised the brito-hoyos et al. (2013) predictions and found that key biogeographic principles were missing, resulting in a misinformation of rabies transmission risk in colombia. similarly, studies have identified artifacts in four bioclimatic variables (bio8, bio9, bio18, and bio19), evidenced by unusual spatial anomalies and inconsistencies among adjacent pixels (booth 2022; escobar et al. 2014). therefore, it has been recommended to assess the discontinuities among the four problematic variables for the specific study areas before final inclusion in an enm. nevertheless, olivera et al. (2021) still included all the available bioclimatic variables in their enm without quality check, which may fail to identify actual suitable areas. in a systematic review, feng et al. (2019a) found a lack of critical modeling information (e.g., source of occurrence data, spatial extent, source of environmental data) in enm applications to biodiversity. enms of infectious diseases is inherently more complex than traditional applications to biodiversity (escobar and craft 2016). understanding disease ecology, particularly transmission dynamics involving hosts, vectors, and pathogens, is often complicated by small sample sizes, sampling biases, and limited geographic and environmental data (escobar 2020). moreover, failing to incorporate key biological components in enm for disease systems may hinder the prediction of disease distribution and transmission risk (escobar and craft 2016; peterson, 2006). given the growing role of enm in the understanding of the ecology of disease systems and informing spatial epidemiology, it is essential to evaluate how researchers incorporate fundamental components and assumptions to ensure robust and biologically meaningful model results. a standardized protocol may promote consistency and best practices in model development (zurell et al. 2020). previous studies have shown that the transparency and methodological rigor of individual-based or agent-based models have improved following the introduction of the overview, design concepts, and details (odd) protocol (grimm et al. 2010). similarly, the adoption of the darwin core standard has had a positive influence on data sharing practices (wieczorek et al. 2012). considering the importance of having a standard checklist for ecological niche studies, and to support and guide key steps in enm, zurell et al. (2020) adapted the odd protocol into the overview/ conceptualization, data, model fitting, assessment, and prediction (odmap) framework. the odmap protocol complements and integrates the range model metadata standard (rmms) by providing a step-by-step guide along with a comprehensive checklist of components required to build an ecological niche model (merow et al. 2019, zurell et al. 2020; fitzpatrick et al. 2021). the odmap protocol also includes a shiny web application that offers guidance for developing standardized enm documentation using the rmms dictionary and allows users to download the odmap table (zurell et al. 2020; fitzpatrick et al. 2021). as an updated and adaptable metadata framework, odmap is well-suited for documenting enm components and is designed to automatically accommodate future updates to the dictionary (merow et al. 2019, zurell et al. 2020; fitzpatrick et al. 2021). zurell et al. (2020) demonstrated the utility of their odmap proshariful islam et al. – ecological niche modeling applications to infectious diseases 63 tocol through nine case studies. it has been expected that odmap will have a similar impact like odd protocol in biodiversity informatics (zurell et al. 2020; fitzpatrick et al. 2021). the odmap protocol is designed to assist beginners in enm, provides a structured workflow for advanced modelers, and serves as an efficient tool for editors and reviewers to evaluate model-based studies (zurell et a. 2020). the aim of this study was to assess recent enm applications to disease systems using the odmap checklist. to our knowledge, no studies have been conducted to evaluate the quality standards of the implementation of enm in disease ecology and disease biogeography. based on key enm requirements, data limitations, and the complexities of disease transmission systems, we hypothesize that most articles fail to account for critical components during the modeling of infectious disease. materials and methods search strategy we conducted a systematic literature review on web of sciences following the prisma framework (moher et al. 2009) (figure 1, date 2/15/2023). we recovered articles published between 2020 and 2022 using the terms “ecological niche modeling” or “species distribution modeling” and “disease”. we selected this period since it corresponds to the release of a disease-specific article published with recommendations on enm applications to describe and forecast disease (escobar 2020). the study period also provides an overview of the early 2020s as a proxy of contemporary methods, including the period of explosion of covid-19 articles among disciplines (riccaboni and verginer 2022). to avoid missing any articles, we repeated the search following the same terms on february 29, 2024. figure 1: flow diagram of systematic article search and selection for the data synthesis. from top to bottom, different background color indicates the article search on web of science, articles identification, article screening following inclusion and exclusion criteria, and final list of articles selected to include and data synthesis. shariful islam et al. – ecological niche modeling applications to infectious diseases 64 after extracting information from all selected articles, we saved the dataset as a .csv file and imported it into r statistic software version 4.4.1 for the descriptive analysis (r core team 2024). we used the ggplot2 package for visualization (wickham 2016). for components with discrete or categorical data (e.g. yes/no), we used bar plot, while for continuing data, we used box plots. for categorical components with only one observation, we excluded that category from the bar plot. we used vosviewer (www.vosviewer.com) version 1.6.20, to conduct keyword network analysis and visualization. we considered the minimum of 2 occurrences of keywords as a threshold for the keyword network in vosviewer. results we collected a total of 491 published articles between 2020 and 2022. after screening articles following inclusion and exclusion criteria, 78 articles were selected for review and evaluated for data collection. results were categorized into five sections, including overview, data, modelling, assessment, and prediction (odmap) following protocol by (zurell et al. 2020). the journals in which the selected articles were published had impact factors ranging from 0.6 to 14.8 (n=75; mean=3.2) (figure s1). based on keywords network analysis, the application of enm in the field of infectious disease research appears to be increasing. the larger node “ecological niche model” and “species distribution model” highlight its recurrent and central themes in the analysis of the reviewed articles (figure 2). overview for the overview section there were twenty-two components, only eight components were explained by all 78 articles (36.4%; n=22) (figure 3). for instance, only a small subset of articles provided information about code availability (9.0%; n=7), spatial extent (i.e., longitude, latitude) (26.9%; n=21), temporal resolution (35.9%; n= 8), data availability (36.4%; n=44), and model assembling (64.1%; n=50). study selection all articles retrieved were downloaded and revised using detailed inclusion criteria (table 1). we screened each article to assess its eligibility for this study, including language, publication period, research article, and application of enm to understand the disease systems. articles retained were read in full and data were extracted and standardized in terms of units to allow comparisons. data collection and preparation following zurell et al. (2020), we assessed and extracted data from each article following five standard sections to gather data or model parameterization, including an overview, data, modeling, assessment, and prediction (odmap) (supplementary material). for each section model components or parameters are required for robust enm calibration and validation. besides, we collected the most recent impact factor available for each journal in which the selected articles were published, using information from the respective journal websites. all data were entered and processed using ms excel (microsoft office professionalplus 2021, microsoft corporation, washington) and saved as .csv file. pieces of information were continuous or binary (yes or no) for some parameters. if the answer was “yes,” we recorded the corresponding details. for example, if the question was whether the author mentioned the statistical algorithm used for enm building, and the answer was “yes,” we noted the algorithm employed in the study. similarly, if the question was whether the author mentioned the temporal duration of the study, and the answer was “yes,” we recorded the duration as a quantitative value. to standardize the geographical extent of the studies, we categorized them as regional if the study area included more than one country, or state-level if articles specified a sub-area within a country. some articles reported spatial resolution in arcminutes. we converted the resolution from arcminutes to meters using conversion factors provided by the united states geological survey (usgs 2011). table 1: inclusion criteria considered selecting and removing the articles during article selection step. inclusion criteria exclusion criteria articles written in english articles written in a language other than english articles published between 2020 to 2022 removed all the articles outside the study period research articles removed review articles, commentaries, rebuttals, opinion letters or editorial notes study used enm to understand the ecology and distribution of a pathogen, or vector, or reservoir or definite host of an infectious disease study fails to define a component of the disease system, non-infectious disease (e.g., cancer, mental health) full text article is available full text article is not available shariful islam et al. – ecological niche modeling applications to infectious diseases 65 figure 2: keyword network analysis for the selected articles (n=78). we used default settings in vosviewer and considered minimum 2 occurrences of keywords as a threshold. the node size indicates its frequency in the analysis of the reviewed articles. color gradient indicates the temporal trends of use of different keywords in reviewed articles along with the publication years. the spatial resolution of data ranged from 6 to 17,000 meters, with a median pixel size of 1000 meters (figure s2a). articles varied in spatial extent, from village to district, state, country, multi-country (regional), and global extents. most of the studies were conducted in the united states (figure s2b). the temporal extent of the articles varied between one and 101 years, with a median of 11 years (figure s2c). articles used different temporal resolutions (time durations between two sequential data collections), from hourly to annual datasets. most articles used data that were collected monthly, seasonally or annually (figure s2d). researchers applied various statistical methods to implement enm theory and examine the ecological niches of their disease study subjects. sixty articles (76.9%) used only maxent as a modeling technique. other articles also used maxent in combination with other modeling techniques, such as generalized linear models, random forest, boosted regression trees, and generalized additive models (table s1). considering subject organisms, most articles aimed to reconstruct the ecological niche of mosquitoes (20.5%; n=16) and ticks (14.1%; n=11). only a limited number of articles used enm to understand the ecological niches of pathogens in tandem with their hosts (figure s2e). out of 78 articles, six used biotic (e.g., human footprint) in tandem with abiotic (e.g., temperature) variables to study infectious disease systems. data four of the 32 components used for modeling were present, including ecological level, taxon names, predictor variables and spatial extent, representing 12.5% (n=32) of figure 3: articles containing or missing information about different components of the overview section. the overview section contained key information about analysis. cyan-blue: number of articles with overview information. green-cyan: number of articles missing overview information. the most complete components in the overview section were types of extent boundary, target output, response data type, predictor types, observation type, model objective, specific study location, focal taxon, and affiliation of first author. while the least reported components were biotic variables, code availability, spatial extent (lon, lat), temporal resolution, and data availability. shariful islam et al. – ecological niche modeling applications to infectious diseases 66 the articles reviewed. a limited number of articles provided information on background data (44.9%; n=35), sample design (48.7%; n=38), addressing errors and biases (61.5%; n=48), temporal resolution (24.4%; n=19), and extent of transfer data (30.8%; n=24). articles examined pathogens and their hosts at different levels of biological organization. the majority of the research was done at the individual level (57.7%; n=45), followed by population (33.3%; n=26) and community (9.0%; n=7) levels (figure s3.1a). articles used a variety of sampling designs, such as field surveys (n=8), random sampling (n=11), and data from online databases (n=7), but most (n=41) of the 78 articles reviewed did not disclose details regarding their sampling methodology (figure s3.1c). the sample sizes used in enm research varied from 7 (bacillus cereus) to 1,394,279 (tuberculosis) occurrence records, with a median of 233 occurrences (figure s3.1b). most articles mentioned data partitioning procedures during model evaluation but did not mention percentage or number of observations or occurrences records used for model training, and testing. data partitioning information for model validation was explained in 54 articles (69.2%), and training data in 59 (75.6%) articles (figure 4). articles employed either geographic coordinate reference systems (n=12) or projected coordinate systems (n=16) to spatially identify occurrence records. only 32 (42.2%) articles provided details regarding the coordinate reference system (figure s3.2a). of the 78 articles, 74 used bioclimatic variables, either alone or in tandem with other variables. in addition to bioclimatic variables, articles used sociodemographic, topography, elevation, normalized difference vegetation index, and population density data as enm predictor variables (figure s3.2b). variable data were mostly derived (n=57) from the worldclim database, either by itself or in conjunction with other data sources (figure s3.2c; table s2). articles also downloaded bioclimatic data from merraclim (n=3), chelsa (n=4), and other sources (figure s3.2c; table s2). almost all articles (96.2%; n=75) mentioned model transference. a limited number of articles provided information about the spatial resolution and extent of variables for model transference. similarly, a small number of articles mentioned the temporal resolution (24.4%; n=19) and temporal extent (30.8%; n=24) of their study. only 38.5% of articles (n=30) provided information about environmental scenarios used in model transferring (figure 4). modelling there were twelve distinct components in the modeling section, but only one of them (8.3%, n=12) was explained in 100% of articles (figure 4). information about nested data was provided in 2.7% (n=2), temporal autocorrelation in 2.7% (n=2), model averaging in 20.5% (n=16), model ensemble in 39.7% (n=31), and model setup for extrapolation 46.2% (n=36). most articles (n=53) did not provide information about the pre-selection of the predictor variables. there were 15 articles merely referring to “correlation analysis” without explaining the precise methodology used to assess variable redundancy. studies used seven different methods or techniques for the pre-selection of the variables (figure s4a). for variables, the most commonly used methods were principal component analysis (n=13), pearson correlation analysis (n=12), spearman correlation (n=5), and variance inflation factor (n=4) (figure s4a). articles used eight different techniques to identify the variable of importance (figure s4b). for identifying variable importance, studies implemented the jackknife test (n=20), correlation score (n=6), principal component analysis (n=6), figure 4: data section is divided into four different subsections (biodiversity data, predictor variables, data partitioning and transfer data). this sections summaries the information about species and predictor variables, and data processing. cyan-blue: number of articles with data information. green-cyan: number of articles missing data information. the most complete components in the data section were taxon names, spatial extent (predictor), ecological level, predictor variables, and data sources. while the least reported components were temporal resolution (transfer), temporal extent (transfer), absence data, quantification of novelty, and models and scenarios. * is for predictor, # is for transfer. shariful islam et al. – ecological niche modeling applications to infectious diseases 67 permutation importance (n=6), and cross-validation (n=3) most commonly (figure s4b). articles mentioned that they checked for multicollinearity among the variables, though most of them did not mention specific methods they had used (n=26). we found that the articles used eight different techniques for multicollinearity check. the most common techniques for multicollinearity check in enm were principal component analysis (n=7), pearson correlation (n=6), and variance inflation factor (n=5) (figure s4c). assessment in the assessment section, five components were included, which are useful to explain the estimated relationship between biodiversity and environmental data (figure 5). none of the five components of assessment section were explained in all articles (figure 6). a limited number of articles provided information about expert judgement about enm assessment (11.5%; n=9), performance on validation data (46.2%; n=36), and response shape of predictor variables (62.8%; n=49). prediction in enm articles, seven components were considered for prediction. only the information related to prediction unit was explained in all articles (figure 7). only a limited number of articles addressed algorithmic uncertainty (28.2%; n=22), novel environment (38.2%; n=29), and scenario uncertainty (38.5%; n=30). figure 5: articles containing or missing information about different components of the modelling section. the modelling section contained information about repeatability of the model building. cyan-blue: number of articles with modelling information. green-cyan: number of articles missing modelling information. the most completely reported components in the modelling section were model setting fitting, threshold selection, model selection, variable importance, and parameter uncertainty. while the least reported components were temporal autocorrelation, nested data, model averaging, model ensembles, and coefficients. shariful islam et al. – ecological niche modeling applications to infectious diseases 68 figure 6: articles containing or missing information about different components of the assessment section. the assessment section contained key information about the performance of the built enm. cyan-blue: number of articles with assessment information. green-cyan: number of articles missing assessment information. the most complete components of assessment section were performance of test data and performance on training data. while the least reported components were expert judgement and performance on validation data. figure 7: articles containing or missing information about different components of the prediction section. the prediction section contained information about inter or extrapolating of the developed enm. cyan-blue: number of articles with prediction information. green-cyan: number of articles missing prediction information. the most complete components in the prediction section were prediction unit, input data uncertainty, and parameter uncertainty. while the least reported components were algorithmic uncertainty, novel environment, and scenario uncertainty. shariful islam et al. – ecological niche modeling applications to infectious diseases 69 discussion we evaluated published articles for consistency and transparency in reporting key modeling and data components in ecological niche modeling (enm) applications in disease ecology and disease biogeography. the use of enm in infectious disease research is increasing with the journals where articles are being published varying in impact factor. a previous study reported that species distribution modeling research is published across a wide range of journals, with 12% annual growth (vasconcelos et al. 2024). the increasing trends of articles publication in diverse journals reflect the interdisciplinary nature and broad applicability of enm. we found that most articles missed important components during enm building and were biased towards a certain modeling technique (i.e., maxent), disease system (i.e., vector-borne), and study region (i.e., united states). our study identified methodological gaps of published articles, which could limit the replicability of studies. our findings highlight an urgent need to improve enm protocols and workflows among researchers, reviewers, and editors to as a means to improve rigor and reproducibility. our analysis revealed that many infectious disease articles lacked comprehensive reporting of key components essential for the proper implementation of enm across the odmap sections (feng et al. 2019a; zurell et al. 2020). for instance, temporal (n=19; 24.4%) and spatial resolution (n=41, 52.6%), which directly influence modeling outputs, were inadequately addressed in most articles (figure 4). ecological niches are highly context dependent and vary across different spatial and temporal scales. to understand pathogen transmission dynamics, comprehensive and scale-aware approaches are important for understanding ecology and physiology of the organisms involved in disease transmission (escobar 2020; zarzo-arias et al. 2023). the absence of a clear description of model parameterization limits the accurate validation and reproducibility of enm research, posing challenges considering the fields rapid growth and the lack of reproducibility standards (araújo et al. 2019; feng et al. 2019a; hallgren et al. 2019; kass et al. 2025; sandel et al 2011; schmolke et al. 2010; sillero et al. 2021; soley-guardia et al. 2024; valavi et al. 2022; veronica and liu 2024). previous case studies have similarly highlighted the unstructured and incomplete explanation of components (e.g., braga et al. 2014; escobar and peterson 2013; escobar et al. 2014; escobar et al. 2018; kass et al. 2025; merow et al. 2013; peterson and nakazawa 2008; xu et al. 2024; zurell et al. 2020). to address these issues, infectious disease modelers should adhere to a standardized framework that ensures detailed reporting of fundamental modeling components, facilitating both effective communication among modelers and better understanding by readers. numerous mathematical and statistical algorithms are available to estimate the ecological niches, with some being more frequently used on the assumption that they are the most effective approaches (araújo et al. 2019; beery et al. 2021; blonder et al. 2018; drake 2015; elith et al. 2008; escobar 2020; kass et al. 2025; phillips et al. 2006; qiao et al. 2015; soley-guardia et al. 2024; valavi et al. 2022; wiens et al. 2009). we found that most articles utilized maxent as the predominant enm algorithm (table: s1). nevertheless, evidences demonstrate that no single modeling algorithm is universally superior (qiao et al., 2015; valavi et al., 2022). maxent is the most common algorithm used by researchers modeling ecological niches and spatial distributions (barker and macisaac 2022; campos et al. 2023; feng et al. 2019b; lippi et al. 2023a; valavi et al. 2022). previous studies have shown that the performance of enm varies with changes in species or calibration areas (elith et al., 2006; qiao et al., 2019; valavi et al., 2022) and algorithm employed (escobar et al. 2018). strikingly, the selection of modeling algorithms depends on the preferences of the modelers, instead of more scientific approaches, with some modelers preferring default settings and others careful model tuning (valavi et al. 2022). the preference maybe responds to the modelers’ limited understanding of how algorithms work (joppa et al. 2013). to construct an enm, the choice of modeling techniques and predictor variables should be informed by the physiology, ecology, and biogeographic history of the organism (escobar 2020; peterson et al. 2011). joppa et al. (2013) recommended that the community should prioritize the understanding of the algorithms over the development of newer, fancier modeling workflows. we found that most studies used predictor variables from a particular data source (i.e., worldclim) (hijmans et al. 2005). reliance on a specific source suggests that researchers are potentially biased on the use of similar types of readily available predictors, despite differences in study questions, model requirements, and mismatching study periods (araújo et al. 2019; booth 2022; morales‐barbero et al. 2019; oliver and morecroft 2014; regos et al. 2019). previous studies also found worldclim data to be the most frequently used (e.g., barker and macisaac 2022; bobrowski et al. 2021; booth 2022; datta et al. 2020; escobar et al. 20014; lippi et al. 2023a; maria et al. 2017; merkenschlager et al. 2023; morales‐barbero et al. 2019; poggio et al. 2018). worldclim temperature and precipishariful islam et al. – ecological niche modeling applications to infectious diseases 70 tation data are derived from global weather station averages and interpolated with covariables such as elevation, distance to the coast, and satellite derived data to generate high-resolution bioclimatic variables (e.g., 30 arc seconds, approximately 1 kilometer) (fick and hijmans 2017; hijmans et al. 2005; waltari et al. 2014; xu and hutchinson 2011). because weather station coverage varies geographically (hijmans et al. 2005; waltari et al. 2014), worldclim interpolations at ~1 km resolution generate 99.99% model-generated, interpolated climatic data (peterson 2014). in contrast, remote sensing offers consistent, largescale monitoring of surface temperature and precipitation, including day and night measurements (adler et al. 2000; jones et al. 2010; waltari et al. 2014). a study reported that enm built with predictor variables from merra performed as well as or better than models with worldclim data (waltari et al. 2014). as such, the selection of predictor variables for enm should align with the study question and spatiotemporal scale of the calibration data (pearson and dawson 2003; peterson et al. 2011; regos et al. 2019). using multiple predictors, including non-climatic predictors, improves ecological understanding of organisms (regos et al. 2019). employing diverse data sources, while considering species behavior and calibration area, helps reduce uncertainty and avoid overreliance in enm construction (regos et al. 2019). a next frontier in enm applications to disease should be the combination of abiotic with biotic predictors, but generation of biotic variables is warranted (peterson et al. 2019). we identified significant gaps in published articles regarding clear explanations of occurrence record data collection, cleaning processes, removal of spatial biases, and data partitioning into training/testing datasets. data quality is a critical factor in building accurate and reliable enm (feng et al. 2019a; escobar 2020; soley-guardia et al. 2024; zurell et al. 2020). ecological processes vary across spatial scales, and errors in selecting appropriate resolution and extent for occurrence and predictor variables can violate biological and statistical assumptions (feng et al. 2019a; mcgill 2010; soberón and nakamura 2009; soley-guardia et al. 2024). biotic and abiotic predictors influence species distributions, but the level of influence vary with spatial resolution of variables, often requiring modification of original resolution of variables for consistency (soberón and nakamura 2009; sunday et al. 2012). as such, researchers should follow standard procedures to describe how spatial and environmental biases were addressed (boria et al. 2014; park and davis 2017), how background data were selected (barbet-massin et al. 2012; phillips et al. 2009), how calibration data were collected and prepared (anderson and raza 2010; escobar 2020; saupe et al. 2012), and how multicollinearity among predictors were managed (bucklin et al. 2015; pliscoff et al. 2014; synes and osborne 2011; zeng et al. 2016). a higher number of articles utilized enm to study the ecology and distribution of vector-borne disease caused by mosquitoes and ticks (figure s2e). our results align with lippi et al. (2023b) and van de vuurst et al. (2023), who reported a growing focus on modeling vector distribution over the past decade. the higher number of enm articles may be driven by factors such as the high disease burden posed by vectors, the critical need for risk area identification, policy mandates, and biased allocation of resources, or higher availability of vector datasets. enm studies of infectious diseases are also geographically biased. we found a strong focus of enm applications to disease in the americas (figure s2b). lippi et al. (2023a) and van de vuurst et al. (2023) observed a higher number enm modeling efforts in north america and europe. paradoxically, vector-borne disease are predominantly identified in tropical, low-income countries, highlighting the need for integrated, geographically focused studies that account for environmental justice (rosenberg et al. 2013). the limited number of enm in regions like southeast asia may stem from scarce publicly available data, limited research resources, and fewer research effort compared to north and south america, and africa (rosenberg et al. 2013). in contrast, the americas benefit from robust systematic surveillance systems like the national ecological observatory network, which provide long-term, freely available pathogen, vector, and host data for the united states (springer et al. 2016). global databases like the global biodiversity information facility also support parasite, vector, and host research, though their data quality and geographic coverage are influenced by research and funding biases (gbif secretariat and iaia, 2020). more efforts may be needed to standardize and release pathogen occurrence data. overall, the public health and economic significance of infectious diseases, availability of long-term environmental data, and extensive occurrence records are likely to promote more enm studies in years to come. the use of enm methods, such as maxent, have been commonly used to study the ecology of vector-borne and directly transmitted zoonotic diseases. for instance, maxent has been effectively applied to map the potential distribution of schistosomiasis hosts and to guide eradication efforts for tsetse flies and trypanosomiasis (dicko et al. 2014, singleton et al. 2024). nevertheless, enm based solely on correlative relationships with climatic variables, shariful islam et al. – ecological niche modeling applications to infectious diseases 71 like those from worldclim, may fall short in capturing the full complexity of disease ecology (figure 8) (cuervo et al. 2023). for example, a parsimonious study of avian influenza could include only one viral lineage (e.g., h5n1), one primary host (e.g., one bird), and one secondary host (e.g., pig) in a low-dimensional environmental space (e.g., one environmental variable). this simple model to forecast transmission from birds to other species including humans, could identify just a portion of the actual transmission risk (figure 8a). to develop more comprehensive and biologically meaningful models, it is essential to understand the fundamental ecology of the disease systems and incorporate data on original hosts, secondary hosts, pathogen diversity, and relevant human features in enm applications (figure 8b). maxent limitations derive from its vulnerability to extrapolate, which may produce overly simplistic or misleading results (escobar et al. 2018; qiao et al. 2019). for example, maxent models can unrealistically predict the survival of malaria vectors beyond 100°c (boiling water) temperature (owens et al 2013). therefore, future enm to understand the ecology of infectious disease, including pathogen, vector, and host, should consider the physiology of the organisms to generate sound and realistic forecast. caveats the odmap primarily serves as a framework to enhance the reproducibility and transparency of spatial distribution models by standardizing the reporting of modeling procedures and data (zurell et al. 2020). its core purpose is to allow other researchers to evaluate, reproduce, and build upon published work rigorously, rather than inherently guaranteeing the ecological validity or comprehensive capture of complex disease ecological processes. since pioneering applications to infectious diseases (i.e., peterson et al. 2002), use of enm in infectious disease research has emerged as a new theoretical and analytical framework (srivastava et al. 2019; escobar 2020; vasconcelos et al. 2024). nevertheless, key enm components (e.g., study area, occurrence points, environmental variables, analytical framework) could mislead model results (araújo et al. 2019; barker and macisaac 2022). our analfigure 8: schematic representation of simple and complex ecological niche model development. a. building and geographical projection of ecological niche model using a subtype of avian influenza virus, a type of predictor, and a bird species may result in an incomplete understanding of the ecology of zoonotic avian influenza. b. building and geographical projection of ecological niche using multiple subtypes of avian influenza virus, diverse predictors and multiple bird species for a more comprehensive understanding of the ecology of zoonotic avian influenza. shariful islam et al. – ecological niche modeling applications to infectious diseases 72 ysis focused solely on description of the analytical framework in the literature based on the odmap protocol. we did not account for the theoretical framework employed for the biological interpretation of the models. odmap protocol suggests describing enm parameterization in the main article, so that we focused on the details in the main articles. it is plausible that enm tuning was rigorous for some studies but described only in the supplementary material, not in the main article. to address this caveat, we thoroughly examined the supplementary materials available in the selected articles and found no methodological descriptions as supplementary materials. overall, we found a considerable lack of details to foster reproducibility and also minimum efforts in model parameterization to find the best model that reconstructs nature and allow the prediction of disease events (levin 1992). this review covered articles published between 2000 and 2022 as a sample of current trends. although, previous seminal work in the field and recent applications were not included, we argue that they likely suffer similar flaws in the analytical framework. future enm reviews should consider broader publication timeframes and future enm applications should adopt enm protocols to enhance reproducibility and transparency (feng et al. 2019a; zurell et al. 2020). conclusions this study highlights flaws in enm applications in disease ecology and disease biogeography. specifically, many published articles provide insufficient explanations of key model components, such as spatial and temporal extent, biodiversity data sources, and mitigation of collinearity, which reduce study reproducibility and transparency. while enm applications have advanced our understanding of disease ecology and disease biogeography, our findings underscore the importance of adopting standard protocols to ensure consistency and reproducibility. our study aimed to lay the groundwork for subsequent discussions and research on effectively incorporating important parameters in enm applications to disease ecology and disease biogeography. acknowledgements this study was supported by the national science foundation human-environment and geographical sciences program (2116748) and career (2235295) awards, and by the national institute of allergy and infectious diseases of the national institutes of health (k01ai168452). this project was also supported by seed grants from the virginia tech institute for critical technology and applied science, pandemic prediction and prevention destination area, and the center for emerging, zoonotic, and arthropod-borne pathogens. the content is solely the authors’ responsibility and does not necessarily represent the official views of the national institutes of health. the authors would like to thank dana hawley, william mark ford, sarah karpanty, paanwaris paansri, and paige van de vuurst for their support on the development and review of this manuscript. data availability the data table on which the analyses in this study are based is available via ku scholarworks.1 competing interests the authors have declared that no competing interests exist. author contributions s.i. provided substantial input on the conceptualization, performed the formal analyses and wrote the first version of the manuscript and headed the review editing. m.c.g., d.s.t., l.e.e. provided methodological revision and gave considerable suggestions on writing – review and editing. l.e.e. developed the idea. all authors approved the last version of this article. references adler, robert f., george j. huffman, david t. bolvin, scott curtis, and eric j. nelkin. 2000. “tropical rainfall distributions determined using trmm combined with other satellite and rain gauge information.” journal of applied meteorology 39(12): 2007-2023. https://doi.org/10.1175/1520-0450(2001) 040%3c2007:trddut%3e2.0.co;2 anderson, robert p., and ali raza. 2010. “the effect of the extent of the study region on gis models of species geographic distributions and estimates of niche evolution: preliminary tests with montane rodents (genus nephelomys) in venezuela.” journal of biogeography 37(7): 1378–1393. https://doi. org/10.1111/j.1365-2699.2010.02290.x araújo, miguel b., robert p. anderson, a. márcia barbosa, colin m. beale, carsten f. dormann, regan early, raquel a. garcia, et al. 2019. “standards for distribution models in biodiversity assessments.” science advances 5(1): eaat4858. https://doi.org/10.1126/ sciadv.aat4858 astorga, francisca, luis e. escobar, daniela poo-muñoz, joaquin escobar-dodero, sylvia rojas-hucks, mario alvarado-rybak, melanie duclos et al. 2018. “distri1 https://hdl.handle.net/1808/36112. https://doi.org/10.1175/1520-0450(2001)040%3c2007:trddut%3e2.0.co;2 https://doi.org/10.1175/1520-0450(2001)040%3c2007:trddut%3e2.0.co;2 https://doi.org/10.1111/j.1365-2699.2010.02290.x https://doi.org/10.1111/j.1365-2699.2010.02290.x https://doi.org/10.1126/sciadv.aat4858 https://doi.org/10.1126/sciadv.aat4858 https://hdl.handle.net/1808/36112 shariful islam et al. – ecological niche modeling applications to infectious diseases 73 butional ecology of andes hantavirus: a macroecological approach.” international journal of health geographics 17: 1-12. https://doi.org/10.1186/ s12942-018-0142-z barbet-massin, morgane, frédéric jiguet, cécile hélène albert, and wilfried thuiller. 2012. “selecting pseudo-absences for species distribution models: how, where and how many?” methods in ecology & evolution 3(2): 327–338. https://doi.org/10.1111/j.2041210x.2011.00172.x barker, justin r., and hugh j. macisaac. 2022. “species distribution models applied to mosquitoes: use, quality assessment, and recommendations for best practice.” ecological modelling 472(10):110073. https://doi.org/10.1016/j.ecolmodel.2022.110073 beery, sara, elijah cole, joseph parker, pietro perona, and kevin winner. 2021. “species distribution modeling for machine learning practitioners: a review.” in acm sigcas conference on computing and sustainable societies (compass) (compass ’21), june 28-july 2, 2021, virtual event, australia. acm, new york, ny, usa, 20 pages. https://doi. org/10.1145/3460112.3471966 blonder, benjamin, cecina babich morrow, brian maitner, david j. harris, christine lamanna, cyrille violle, brian j. enquist, and andrew j. kerkhoff. 2018. “new approaches for delineating n‐dimensional hypervolumes.” methods in ecology and evolution 9(2): 305-319. https://doi.org/10.1111/2041-210x.12865 bobrowski, maria, johannes weidinger, and udo schickhoff. 2021. “is new always better? frontiers in global climate datasets for modeling treeline species in the himalayas.” atmosphere 12(5): 543. https://doi. org/10.3390/atmos12050543 booth, trevor h. 2022. “checking bioclimatic variables that combine temperature and precipitation data before their use in species distribution models.” austral ecology 47(7): 1506-1514. https://doi-org. ezproxy.lib.vt.edu/10.1111/aec.13234 boria, robert a., link e. olson, steven m. goodman, and robert p. anderson. 2014. “spatial filtering to reduce sampling bias can improve the performance of ecological niche models.” ecological modelling 275(3): 73–77. https://doi.org/10.1016/j.ecolmodel.2013.12.012 braga, guilherme basseto, josé henrique hildebrand grisi-filho, bruno meireles leite, elaine fátima de sena, and ricardo augusto dias. 2014. “predictive qualitative risk model of bovine rabies occurrence in brazil.” preventive veterinary medicine 113(4): 536-546. https://doi.org/10.1016/j.prevetmed.2013.12.011 brito-hoyos, diana marcela, edilberto brito sierra, and rafael villalobos alvarez. 2013. “geographic distribution of wild rabies risk and evaluation of the factors associated with its incidence in colombia, 1982-2010.” pan american journal of public health 33(1): 8–14. https://doi.org/10.1590/s102049892013000100002 bucklin, david n., mathieu basille, allison m. benscoter, laura a. brandt, frank j. mazzotti, stephanie s. romañach, carolina speroterra, james i. watling, and wilfried thuiller. 2015. “comparing species distribution models constructed with different subsets of environmental predictors.” diversity & distributions 21(1): 23–35. https://doi.org/10.1111/ddi.12247 campos, joão c., nuno garcia, joão alírio, salvador arenas-castro, ana c. teodoro, and neftalí sillero. 2023. “ecological niche models using maxent in google earth engine: evaluation, guidelines and recommendations.” ecological informatics 76(9):102147. https://doi.org/10.1016/j.ecoinf.2023.102147 cuervo, pablo fernando, patricio artigas, jacob lorenzo-morales, maría dolores bargues, and santiago mas-coma. 2023. “ecological niche modelling approaches: challenges and applications in vector-borne diseases.” tropical medicine and infectious disease 8(4): 187. https://doi.org/10.3390/tropicalmed8040187 datta, arunava, oliver schweiger, and ingolf kuhn. 2020. “origin of climatic data can determine the transferability of species distribution models.” neobiota 59: 61-76. https://doi.org/10.3897/neobiota.59.36299 dicko, ahmadou h., renaud lancelot, momar t. seck, laure guerrini, baba sall, mbargou lo, marc jb vreysen et al. 2014. “using species distribution models to optimize vector control in the framework of the tsetse eradication campaign in senegal.” proceedings of the national academy of sciences usa 111(28): 10149-10154. https://doi.org/10.1073/ pnas.1407773111 dobson, andrew p., stuart l. pimm, lee hannah, les kaufman, jorge a. ahumada, amy w. ando, aaron bernstein, et al. 2020. “ecology and economics for pandemic prevention.” science 369(6502): 379–381. https://doi.org/10.1126/science.abc3189 drake, john m. 2015. “range bagging: a new method for ecological niche modelling from presence-only data.” journal of the royal society interface 12(107): 20150086. http://dx.doi.org/10.1098/rsif.2015.0086 elith, jane, catherine h. graham*, robert p. anderson, miroslav dudík, simon ferrier, antoine guisan, robert j. hijmans, et al. 2006. “novel methods improve https://doi.org/10.1186/s12942-018-0142-z https://doi.org/10.1186/s12942-018-0142-z https://doi.org/10.1111/j.2041-210x.2011.00172.x https://doi.org/10.1111/j.2041-210x.2011.00172.x https://doi.org/10.1016/j.ecolmodel.2022.110073 https://doi.org/10.1145/3460112.3471966 https://doi.org/10.1145/3460112.3471966 https://doi.org/10.1111/2041-210x.12865 https://doi.org/10.3390/atmos12050543 https://doi.org/10.3390/atmos12050543 https://doi-org.ezproxy.lib.vt.edu/10.1111/aec.13234 https://doi-org.ezproxy.lib.vt.edu/10.1111/aec.13234 https://doi.org/10.1016/j.ecolmodel.2013.12.012 https://doi.org/10.1016/j.ecolmodel.2013.12.012 https://doi.org/10.1016/j.prevetmed.2013.12.011 https://doi.org/10.1590/s1020-49892013000100002 https://doi.org/10.1590/s1020-49892013000100002 https://doi.org/10.1111/ddi.12247 https://doi.org/10.1016/j.ecoinf.2023.102147 https://doi.org/10.3390/tropicalmed8040187 https://doi.org/10.3390/tropicalmed8040187 https://doi.org/10.3897/neobiota.59.36299 https://doi.org/10.1073/pnas.1407773111 https://doi.org/10.1073/pnas.1407773111 https://doi.org/10.1126/science.abc3189 http://dx.doi.org/10.1098/rsif.2015.0086 shariful islam et al. – ecological niche modeling applications to infectious diseases 74 prediction of species’ distributions from occurrence data.” ecography 29(2): 129–151. https://doi. org/10.1111/j.2006.0906-7590.04596.x elith, jane, john r. leathwick, and trevor hastie. 2008. “a working guide to boosted regression trees.” journal of animal ecology 77(4): 802-813. https:// doi.org/10.1111/j.1365-2656.2008.01390.x escobar, luis e. 2020. “ecological niche modeling: an introduction for veterinarians and epidemiologists.” frontiers in veterinary science 7: 519059. https://www.frontiersin.org/articles/10.3389/ fvets.2020.519059 escobar, luis e., and serge morand. 2021. “disease ecology and biogeography.” frontiers in veterinary science 8: 765825. https://doi.org/10.3389/ fvets.2021.765825 escobar, luis e., andrés lira-noriega, gonzalo medina-vogel, and a. townsend peterson. 2014. “potential for spread of the white-nose fungus (pseudogymnoascus destructans) in the americas: use of maxent and nichea to assure strict model transference.” geospatial health 9(1): 221. https://doi.org/10.4081/ gh.2014.19 escobar, luis e., and meggan e. craft. 2016. “advances and limitations of disease biogeography using ecological niche modeling.” frontiers in microbiology 07(8):1174. https://doi.org/10.3389/ fmicb.2016.01174 escobar, luis e., and a. townsend peterson. 2013. “spatial epidemiology of bat-borne rabies in colombia.” pan american journal of public health 34(2): 135–136. escobar, luis e., huijie qiao, javier cabello, and a. townsend peterson. 2018. “ecological niche modeling re-examined: a case study with the darwin’s fox.” ecology and evolution 8(10): 4757–4770. https://doi.org/10.1002/ece3.4014 escobar, luis e., huijie qiao, christine lee, and nicholas bd phelps. 2017. “novel methods in disease biogeography: a case study with heterosporosis.” frontiers in veterinary science 4: 105. https:// doi.org/10.3389/fvets.2017.00105 feng, xiao, daniel s. park, cassondra walker, a. townsend peterson, cory merow, and monica papeş. 2019a. “a checklist for maximizing reproducibility of ecological niche models.” nature ecology & evolution 3(10): 1382–1395. https://doi.org/10.1038/ s41559-019-0972-5 feng, xiao, daniel s. park, ye liang, ranjit pandey, and monica papeş. 2019b. “collinearity in ecological niche modeling: confusions and challenges.” ecology and evolution 9(18): 10365–10376. https://doi. org/10.1002/ece3.5555 fick, stephen e., and robert j. hijmans. 2017. “worldclim 2: new 1-km spatial resolution climate surfaces for global land areas.” international journal of climatology 37(12): 4302–4315. https://doi. org/10.1002/joc.5086 fitzpatrick, matthew c., susanne lachmuth, and natalie t. haydt. 2021. “the odmap protocol: a new tool for standardized reporting that could revolutionize species distribution modeling.” ecography 44(7): 1067-1070. https://doi.org/10.1111/ecog.05700 frans, veronica f., and jianguo liu. 2024. “gaps and opportunities in modelling human influence on species distributions in the anthropocene.” nature ecology & evolution 8(7): 1365-1377. https://doi-org.ezproxy. lib.vt.edu/10.1038/s41559-024-02435-3 gbif secretariat and international association for impact assessment. 2020. “best practices for publishing biodiversity data from environmental impact assessments.” https://doi.org/10.35035/doc-5xdm-8762 grimm, volker, uta berger, donald l. deangelis, j. gary polhill, jarl giske, and steven f. railsback. 2010. “the odd protocol: a review and first update.” ecological modelling 221(23): 2760-2768. https:// doi.org/10.1016/j.ecolmodel.2010.08.019 hallgren, w., f. santana, s. low-choy, y. zhao, and b. mackey. 2019. “species distribution models can be highly sensitive to algorithm configuration.” ecological modelling 408(9): 108719. https://doi. org/10.1016/j.ecolmodel.2019.108719 hijmans, robert j., susan e. cameron, juan l. parra, peter g. jones, and andy jarvis. 2005. “very high resolution interpolated climate surfaces for global land areas.” international journal of climatology 25(15): 1965–1978. https://doi.org/10.1002/joc.1276 jones, lucas a., craig r. ferguson, john s. kimball, ke zhang, steven tsz k. chan, kyle c. mcdonald, eni g. njoku, and eric f. wood. 2010. “satellite microwave remote sensing of daily land surface air temperature minima and maxima from amsr-e.” ieee journal of selected topics in applied earth observations and remote sensing 3(1): 111–23. https:// doi.org/10.1109/jstars.2010.2041530 joppa, lucas n., greg mcinerny, richard harper, lara salido, kenji takeda, kenton o’hara, david gavaghan, and stephen emmott. 2013. “troubling trends in scientific software use.” science 340(6134): 814-815. https://doi.org/10.1126/science.1231535 kass, jamie m., adam b. smith, dan l. warren, sergio vignali, sylvain schmitt, matthew e. aiello‐lamhttps://doi.org/10.1111/j.2006.0906-7590.04596.x https://doi.org/10.1111/j.2006.0906-7590.04596.x https://doi.org/10.1111/j.1365-2656.2008.01390.x https://doi.org/10.1111/j.1365-2656.2008.01390.x https://www.frontiersin.org/articles/10.3389/fvets.2020.519059 https://www.frontiersin.org/articles/10.3389/fvets.2020.519059 https://doi.org/10.3389/fvets.2021.765825 https://doi.org/10.3389/fvets.2021.765825 https://doi.org/10.4081/gh.2014.19 https://doi.org/10.4081/gh.2014.19 https://doi.org/10.3389/fmicb.2016.01174 https://doi.org/10.3389/fmicb.2016.01174 https://doi.org/10.1002/ece3.4014 https://doi.org/10.3389/fvets.2017.00105 https://doi.org/10.3389/fvets.2017.00105 https://doi.org/10.1038/s41559-019-0972-5 https://doi.org/10.1038/s41559-019-0972-5 https://doi.org/10.1002/ece3.5555 https://doi.org/10.1002/ece3.5555 https://doi.org/10.1002/joc.5086 https://doi.org/10.1002/joc.5086 https://doi.org/10.1111/ecog.05700 https://doi-org.ezproxy.lib.vt.edu/10.1038/s41559-024-02435-3 https://doi-org.ezproxy.lib.vt.edu/10.1038/s41559-024-02435-3 https://doi.org/10.35035/doc-5xdm-8762 https://doi.org/10.1016/j.ecolmodel.2010.08.019 https://doi.org/10.1016/j.ecolmodel.2010.08.019 https://doi.org/10.1016/j.ecolmodel.2019.108719 https://doi.org/10.1016/j.ecolmodel.2019.108719 https://doi.org/10.1002/joc.1276 https://doi.org/10.1109/jstars.2010.2041530 https://doi.org/10.1109/jstars.2010.2041530 https://doi.org/10.1126/science.1231535 shariful islam et al. – ecological niche modeling applications to infectious diseases 75 mens, eduardo arlé et al. 2025. “achieving higher standards in species distribution modeling by leveraging the diversity of available software.” ecography 2(2025): e07346. https://doi.org/10.1111/ ecog.07346 keesing, felicia, and richard s. ostfeld. 2024. “emerging patterns in rodent-borne zoonotic diseases.” science 385(6715): 1305–1310. https://doi.org/10.1126/ science.adq7993 levin, simon a. 1992. “the problem of pattern and scale in ecology: the robert h. macarthur award lecture.” ecology 73(6): 1943-1967. https://doi. org/10.2307/1941447 lippi, catherine a., stephanie j. mundis, rachel sippy, j. matthew flenniken, anusha chaudhary, gavriella hecht, colin j. carlson, and sadie j. ryan. 2023a. “trends in mosquito species distribution modeling: insights for vector surveillance and disease control.” parasites & vectors 16(1): 302. https://doi. org/10.1186/s13071-023-05912-z lippi, catherine a, samuel s c rund, and sadie j ryan. 2023b. “characterizing the vector data ecosystem.” journal of medical entomology 60(2): 247–254. https://doi.org/10.1093/jme/tjad009 maher, sean p., christine ellis, kenneth l. gage, russell e. enscore, and a. townsend peterson. 2010. “rangewide determinants of plague distribution in north america.” the american journal of tropical medicine and hygiene 83(4): 736. https://doi.org/10.4269/ ajtmh.2010.10-0042 maria, bobrowski, and schickhoff udo. 2017. “why input matters: selection of climate data sets for modelling the potential distribution of a treeline species in the himalayan region.” ecological modelling 359(9): 92-102. https://doi.org/10.1016/j.ecolmodel.2017.05.021 mcgill, brian j. 2010. “matters of scale.” science 328(5978): 575–576. https://doi.org/10.1126/science.1188528 merkenschlager, christian, freddy bangelesa, heiko paeth, and elke hertig. 2023. “blessing and curse of bioclimatic variables: a comparison of different calculation schemes and datasets for species distribution modeling within the extended mediterranean area.” ecology and evolution 13(10): e10553. https://doi.org/10.1002/ece3.10553 merow, cory, matthew j. smith, and john a. silander jr. 2013. “a practical guide to maxent for modeling species’ distributions: what it does, and why inputs and settings matter.” ecography 36(10): 1058-1069. https://doi.org/10.1111/j.1600-0587.2013.07872.x merow, cory, brian s. maitner, hannah l. owens, jamie m. kass, brian j. enquist, walter jetz, and rob guralnick. 2019. “species’ range model metadata standards: rmms.” global ecology and biogeography 28(12): 1912-1924. https://doi-org.ezproxy.lib. vt.edu/10.1111/geb.12993 moher, david, alessandro liberati, jennifer tetzlaff, douglas g. altman, and the prisma group. 2009. “preferred reporting items for systematic reviews and meta-analyses: the prisma statement.” plos medicine 6(7): e1000097. https://doi.org/10.1371/ journal.pmed.1000097 morales‐barbero, jennifer, and julia vega‐álvarez. 2019. “input matters matter: bioclimatic consistency to map more reliable species distribution models.” methods in ecology and evolution 10(2): 212-224. https://doi.org/10.1111/2041-210x.13124 oliver, tom h., and mike d. morecroft. 2014. “interactions between climate change and land use change on biodiversity: attribution problems, risks, and opportunities.” wires climate change 5(3): 317–335. https://doi.org/10.1002/wcc.271 olivera, leonela, eugenia minghetti, and sara i. montemayor. 2021. “ecological niche modeling (enm) of leptoglossus clypealis a new potential global invader: following in the footsteps of leptoglossus occidentalis?” bulletin of entomological research 111(3): 289-300. https://doi.org/10.1017/ s0007485320000656 owens, hannah l., lindsay p. campbell, l. lynnette dornak, erin e. saupe, narayani barve, jorge soberón, kate ingenloff et al. 2013. “constraints on interpretation of ecological niche models by limited environmental ranges on calibration areas.” ecological modelling 263(8): 10-18. https://doi.org/10.1016/j. ecolmodel.2013.04.011 park, daniel s., and charles c. davis. 2017. “implications and alternatives of assigning climate data to geographical centroids.” journal of biogeography 44(10): 2188–2198. https://www.jstor.org/stable/26626941 pearson, richard g., and terence p. dawson. 2003. “predicting the impacts of climate change on the distribution of species: are bioclimate envelope models useful?” global ecology and biogeography 12(5): 361–371. https://doi.org/10.1046/j.1466822x.2003.00042.x peterson, a. townsend. 2008. “biogeography of diseases: a framework for analysis.” naturwissenschaften 95(6): 483–491. https://doi.org/10.1007/s00114-0080352-5 https://doi.org/10.1111/ecog.07346 https://doi.org/10.1111/ecog.07346 https://doi.org/10.1126/science.adq7993 https://doi.org/10.1126/science.adq7993 https://doi.org/10.2307/1941447 https://doi.org/10.2307/1941447 https://doi.org/10.1186/s13071-023-05912-z https://doi.org/10.1186/s13071-023-05912-z https://doi.org/10.1093/jme/tjad009 https://doi.org/10.4269/ajtmh.2010.10-0042 https://doi.org/10.4269/ajtmh.2010.10-0042 https://doi.org/10.1016/j.ecolmodel.2017.05.021 https://doi.org/10.1016/j.ecolmodel.2017.05.021 https://doi.org/10.1126/science.1188528 https://doi.org/10.1126/science.1188528 https://doi.org/10.1002/ece3.10553 https://doi.org/10.1111/j.1600-0587.2013.07872.x https://doi-org.ezproxy.lib.vt.edu/10.1111/geb.12993 https://doi-org.ezproxy.lib.vt.edu/10.1111/geb.12993 https://doi.org/10.1371/journal.pmed.1000097 https://doi.org/10.1371/journal.pmed.1000097 https://doi.org/10.1111/2041-210x.13124 https://doi.org/10.1002/wcc.271 https://doi.org/10.1017/s0007485320000656 https://doi.org/10.1017/s0007485320000656 https://doi.org/10.1016/j.ecolmodel.2013.04.011 https://doi.org/10.1016/j.ecolmodel.2013.04.011 https://www.jstor.org/stable/26626941 https://www.jstor.org/stable/26626941 https://doi.org/10.1046/j.1466-822x.2003.00042.x https://doi.org/10.1046/j.1466-822x.2003.00042.x https://doi.org/10.1007/s00114-008-0352-5 https://doi.org/10.1007/s00114-008-0352-5 shariful islam et al. – ecological niche modeling applications to infectious diseases 76 peterson, a. townsend, robert p. anderson, marlon e. cobos, martín cuahutle, angela p. cuervo-robayo, luis e. escobar, marc fernández et al. 2019. “curso modelado de nicho ecológico, versión 1.0.” biodiversity informatics 14: 1-7. https://doi.org/10.17161/ bi.v14i0.8189 peterson, a. townsend, and kate ingenloff. 2016. “biodiversity informatics training curriculum, version 1.2.” biodiversity informatics 11:1. https://doi. org/10.17161/bi.v11i1.5008 peterson, a. t., and y. nakazawa. 2008. “environmental data sets matter in ecological niche modelling: an example with solenopsis invicta and solenopsis richteri.” global ecology and biogeography 17(1): 135-144. https://doi-org.ezproxy.lib.vt.edu/10.1111/ j.1466-8238.2007.00347.x peterson, a. townsend. 2014. mapping disease transmission risk. baltimore: johns hopkins university press. https://doi.org/10.1353/book.36167 peterson, a. t., victor sánchez-cordero, c. ben beard, and janine m. ramsey. 2002. “ecologic niche modeling and potential reservoirs for chagas disease, mexico.” emerging infectious diseases 8(7): 662– 667. https://doi.org/10.3201/eid0807.010454 peterson, a.t., j. soberón, r.g. pearson, r.p. anderson, e. martínez-meyer, m nakamura, and m.b. araújo. 2011. “ecological niches and geographic distributions.” princeton university press, new jersey, usa. https://press.princeton.edu/books/paperback/9780691136882/ecological-niches-and-geographic-distributions-mpb-49 phillips, steven j., miroslav dudík, jane elith, catherine h. graham, anthony lehmann, john leathwick, and simon ferrier. 2009. “sample selection bias and presence-only distribution models: implications for background and pseudo-absence data.” ecological applications 19(1): 181–197. https://doi. org/10.1890/07-2153.1 phillips, steven j., robert p. anderson, and robert e. schapire. 2006. “maximum entropy modeling of species geographic distributions.” ecological modelling 190(3-4): 231-259. https://doi.org/10.1016/j. ecolmodel.2005.03.026 pliscoff, patricio, federico luebert, hartmut h. hilger, and antoine guisan. 2014. “effects of alternative sets of climatic predictors on species distribution models and associated estimates of extinction risk: a test with plants in an arid environment.” ecological modelling 288(9):166–177. https://doi.org/10.1016/j. ecolmodel.2014.06.003 poggio, laura, enrico simonetti, and alessandro gimona. 2018. “enhancing the worldclim data set for national and regional applications.” science of the total environment 625: 1628-1643. https://doi. org/10.1016/j.scitotenv.2017.12.258 qiao, huijie, xiao feng, luis e. escobar, a. townsend peterson, jorge soberón, gengping zhu, and monica papeş. 2019. “an evaluation of transferability of ecological niche models.” ecography 42(3): 521– 534. https://doi.org/10.1111/ecog.03986 qiao, huijie, jorge soberón, and a. townsend peterson. 2015. “no silver bullets in correlative ecological niche modelling: insights from testing among many potential algorithms for niche estimation.” methods in ecology and evolution 6(10): 1126–1136. https:// doi.org/10.1111/2041-210x.12397 r core team. 2024. “r: a language and environment for statistical computing.” r foundation for statistical computing. 2024. https://www.r-project.org/ regos, adrián, laura gagne, domingo alcaraz-segura, joão p. honrado, and jesús domínguez. 2019. “effects of species traits and environmental predictors on performance and transferability of ecological niche models.” scientific reports 9(1): 4221. https:// doi.org/10.1038/s41598-019-40766-5 riccaboni, massimo, and luca verginer. 2022. “the impact of the covid-19 pandemic on scientific research in the life sciences.” plos one 17(2): e0263001. https://doi.org/10.1371/journal.pone.0263001 romero-alvarez, daniel, luis e. escobar, albert j. auguste, sara y. del valle, and carrie a. manore. 2023. “transmission risk of oropouche fever across the americas.” infectious diseases of poverty 12(1): 47. https://doi.org/10.1186/s40249-023-01091-2 rosenberg, ronald, michael a. johansson, ann m. powers, and barry r. miller. 2013. “search strategy has influenced the discovery rate of human viruses.” proceedings of the national academy of sciences usa 110(34): 13961–13964. https://doi.org/10.1073/ pnas.1307243110 sandel, brody, l. arge, bo dalsgaard, r. g. davies, k. j. gaston, w. j. sutherland, and j-c. svenning. 2011. “the influence of late quaternary climate-change velocity on species endemism.” science 334(6056): 660-664. https://doi.org/10.1126/science.1210173 saupe, e. e., v. barve, c. e. myers, j. soberón, n. barve, c. m. hensz, a. t. peterson, h. l. owens, and a. lira-noriega. 2012. “variation in niche and distribution model performance: the need for a priori assessment of key causal factors.” ecological modelling 237–238(7):11–22. https://doi.org/10.1016/j. ecolmodel.2012.04.001 schmolke, amelie, pernille thorbek, donald l. deangelis, and volker grimm. 2010. “ecological models suphttps://doi.org/10.17161/bi.v14i0.8189 https://doi.org/10.17161/bi.v14i0.8189 https://doi.org/10.17161/bi.v11i1.5008 https://doi.org/10.17161/bi.v11i1.5008 https://doi-org.ezproxy.lib.vt.edu/10.1111/j.1466-8238.2007.00347.x https://doi-org.ezproxy.lib.vt.edu/10.1111/j.1466-8238.2007.00347.x https://doi.org/10.1353/book.36167 https://doi.org/10.3201/eid0807.010454 https://press.princeton.edu/books/paperback/9780691136882/ecological-niches-and-geographic-distributions-mpb-49 https://press.princeton.edu/books/paperback/9780691136882/ecological-niches-and-geographic-distributions-mpb-49 https://press.princeton.edu/books/paperback/9780691136882/ecological-niches-and-geographic-distributions-mpb-49 https://doi.org/10.1890/07-2153.1 https://doi.org/10.1890/07-2153.1 https://doi.org/10.1016/j.ecolmodel.2005.03.026 https://doi.org/10.1016/j.ecolmodel.2005.03.026 https://doi.org/10.1016/j.ecolmodel.2014.06.003 https://doi.org/10.1016/j.ecolmodel.2014.06.003 https://doi.org/10.1016/j.scitotenv.2017.12.258 https://doi.org/10.1016/j.scitotenv.2017.12.258 https://doi.org/10.1111/ecog.03986 https://doi.org/10.1111/2041-210x.12397 https://doi.org/10.1111/2041-210x.12397 https://www.r-project.org/ https://doi.org/10.1038/s41598-019-40766-5 https://doi.org/10.1038/s41598-019-40766-5 https://doi.org/10.1371/journal.pone.0263001 https://doi.org/10.1186/s40249-023-01091-2 https://doi.org/10.1073/pnas.1307243110 https://doi.org/10.1073/pnas.1307243110 https://doi.org/10.1126/science.1210173 https://doi.org/10.1016/j.ecolmodel.2012.04.001 https://doi.org/10.1016/j.ecolmodel.2012.04.001 shariful islam et al. – ecological niche modeling applications to infectious diseases 77 porting environmental decision making: a strategy for the future.” trends in ecology & evolution 25(8): 479–486. https://doi.org/10.1016/j.tree.2010.05.001 sillero, neftalí, salvador arenas-castro, urtzi enriquez‐ urzelai, cândida gomes vale, diana sousa-guedes, fernando martínez-freiría, raimundo real, and a. márcia barbosa. 2021. “want to model a species niche? a step-by-step guideline on correlative ecological niche modelling.” ecological modelling 456(9): 109671. https://doi.org/10.1016/j.ecolmodel.2021.109671 singleton, alyson l., caroline k. glidden, andrew j. chamberlin, roseli tuan, raquel gs palasio, adriano pinter, roberta l. caldeira et al. 2024. “species distribution modeling for disease ecology: a multiscale case study for schistosomiasis host snails in brazil.” plos global public health 4(8): e0002224. https://doi.org/10.1371/journal.pgph.0002224 soberón, jorge, and miguel nakamura. 2009. “niches and distributional areas: concepts, methods, and assumptions.” proceedings of the national academy of sciences usa 106(supplement_2): 19644–19650. https://doi.org/10.1073/pnas.0901637106 soley-guardia, mariano, diego f. alvarado-serrano, and robert p. anderson. 2024. “top ten hazards to avoid when modeling species distributions: a didactic guide of assumptions, problems, and recommendations.” ecography 2024(4): e06852. https://doi. org/10.1111/ecog.06852 springer, yuri p., david hoekman, pieter t. j. johnson, paul a. duffy, rebecca a. hufft, david t. barnett, brian f. allan, et al. 2016. “tick-, mosquito-, and rodent-borne parasite sampling designs for the national ecological observatory network.” ecosphere 7(5): e01271. https://doi.org/10.1002/ecs2.1271 srivastava, vivek, valentine lafond, and verena c. griess. 2019. “species distribution models (sdm): applications, benefits and challenges in invasive species management.” cabi reviews 2019: 1-13. https://doi. org/10.1079/pavsnnr201914020 sunday, jennifer m., amanda e. bates, and nicholas k. dulvy. 2012. “thermal tolerance and the global redistribution of animals.” nature climate change 2(9): 686–690. https://doi.org/10.1038/nclimate1539 synes, nicholas w., and patrick e. osborne. 2011. “choice of predictor variables as a source of uncertainty in continental-scale species distribution modelling under climate change.” global ecology and biogeography 20(6): 904–914. https://doi.org/10.1111/j.14668238.2010.00635.x united states geological survey. 2011. “usgs ofr 2011-1127: construction of a 3-arcsecond digital elevation model for the gulf of maine, conversion factors.” conversion factors. 2011. https://pubs. usgs.gov/of/2011/1127/convert.html valavi, roozbeh, gurutzeta guillera-arroita, josé j. lahoz-monfort, and jane elith. 2022. “predictive performance of presence-only species distribution models: a benchmark study with reproducible code.” ecological monographs 92(1): e01486. https://doi.org/10.1002/ecm.1486 van de vuurst, paige, and luis e. escobar. 2023. “climate change and infectious disease: a review of evidence and research trends.” infectious diseases of poverty 12(1): 51. https://doi-org.ezproxy.lib.vt.edu/10.1186/ s40249-023-01102-2 vasconcelos, rodrigo n., taimy cantillo-pérez, washington js franca rocha, william moura aguiar, deorgia tayane mendes, taíse bomfim de jesus, carolina oliveira de santana, mariana mm de santana, and reyjane patrícia oliveira. 2024. “advances and challenges in species ecological niche modeling: a mixed review.” earth 5(4): 963-989. https://doi. org/10.3390/earth5040050 waltari, eric, ronny schroeder, kyle mcdonald, robert p. anderson, and ana carnaval. 2014. “bioclimatic variables derived from remote sensing: assessment and application for species distribution modelling.” methods in ecology and evolution 5(10): 1033–1042. https://doi.org/10.1111/2041-210x.12264 wickham, hadley. 2016. ggplot2: elegant graphics for data analysis. second edition. use r! cham: springer international publishing. wieczorek, john, david bloom, robert guralnick, stan blum, markus döring, renato giovanni, tim robertson, and david vieglais. 2012. “darwin core: an evolving community-developed biodiversity data standard.” plos one 7(1): e29715. https://doi. org/10.1371/journal.pone.0029715 wiens, john a., diana stralberg, dennis jongsomjit, christine a. howell, and mark a. snyder. 2009. “niches, models, and climate change: assessing the assumptions and uncertainties.” proceedings of the national academy of sciences usa 106(2): 1972919736. https://doi.org/10.1073/pnas.0901639106 xu, tingbao, and michael hutchinson. 2011. “anuclim vversion 6.1 user gguide.” fenner school of environment and society, australian national university. xu, quanli, xiao wang, junhua yi, and yu wang. 2024. “bias correction in species distribution models https://doi.org/10.1016/j.tree.2010.05.001 https://doi.org/10.1016/j.ecolmodel.2021.109671 https://doi.org/10.1016/j.ecolmodel.2021.109671 https://doi.org/10.1371/journal.pgph.0002224 https://doi.org/10.1073/pnas.0901637106 https://doi.org/10.1111/ecog.06852 https://doi.org/10.1111/ecog.06852 https://doi.org/10.1002/ecs2.1271 https://doi.org/10.1079/pavsnnr201914020 https://doi.org/10.1079/pavsnnr201914020 https://doi.org/10.1038/nclimate1539 https://doi.org/10.1111/j.1466-8238.2010.00635.x https://doi.org/10.1111/j.1466-8238.2010.00635.x https://pubs.usgs.gov/of/2011/1127/convert.html https://pubs.usgs.gov/of/2011/1127/convert.html https://doi.org/10.1002/ecm.1486 https://doi-org.ezproxy.lib.vt.edu/10.1186/s40249-023-01102-2 https://doi-org.ezproxy.lib.vt.edu/10.1186/s40249-023-01102-2 https://doi.org/10.3390/earth5040050 https://doi.org/10.3390/earth5040050 https://doi.org/10.1111/2041-210x.12264 https://doi.org/10.1371/journal.pone.0029715 https://doi.org/10.1371/journal.pone.0029715 https://doi.org/10.1073/pnas.0901639106 shariful islam et al. – ecological niche modeling applications to infectious diseases 78 based on geographic and environmental characteristics.” ecological informatics 81(7): 102604. https:// doi.org/10.1016/j.ecoinf.2024.102604 zarzo-arias, alejandra, britta uhl, daniel s. maynard, and manuel b. morales. 2023. “the ecological niche at different spatial scales.” frontiers in ecology and evolution 11: 1296340. https://doi.org/10.3389/ fevo.2023.1296340 zeng, yiwen, bi wei low, and darren c. j. yeo. 2016. “novel methods to select environmental variables in maxent: a case study using invasive crayfish.” ecological modelling 34(12):5–13. https://doi. org/10.1016/j.ecolmodel.2016.09.019 zurell, damaris, janet franklin, christian könig, phil j. bouchet, carsten f. dormann, jane elith, guillermo fandos, et al. 2020. “a standard protocol for reporting species distribution models.” ecography 43(9): 1261–1277. https://doi.org/10.1111/ecog.04960 https://doi.org/10.1016/j.ecoinf.2024.102604 https://doi.org/10.1016/j.ecoinf.2024.102604 https://doi.org/10.3389/fevo.2023.1296340 https://doi.org/10.3389/fevo.2023.1296340 https://doi.org/10.1016/j.ecolmodel.2016.09.019 https://doi.org/10.1016/j.ecolmodel.2016.09.019 https://doi.org/10.1111/ecog.04960 shariful islam et al. – ecological niche modeling applications to infectious diseases 79 figure s1: distribution of impact factors of the journals where selected articles for this study were published. supplemental material shariful islam et al. – ecological niche modeling applications to infectious diseases 80 figure s2: some of the selected components from the overview section for a comprehensive understanding. a. boxplot shows the ranges of spatial resolutions (meters) used in the 78 articles; b. bar plot shows the spatial extent of the study area mentioned in 78 articles; c. boxplot shows the temporal extent considered in different articles; d. bar plot shows the temporal resolutions used in different articles; and e. bar plot shows the study subjects considered in different articles to study the ecological niche. we plotted the bars having value more than one only. shariful islam et al. – ecological niche modeling applications to infectious diseases 81 figure s3.1: some of the selected components related to biodiversity data from the data section for a comprehensive understanding. a. bar plot shows the articles using ecological niche models at different biological organization levels; b. box plot shows the ranges of sample sizes used in different articles; and c. bar plot shows the different sampling techniques used in different articles. we plotted the bars having value more than one only. igure s3.2: some of the selected components related to predictor data from the data section for a comprehensive understanding. a. bar plot shows the articles using different coordinate reference system; b. bar plot shows the different predictor variables used in different articles for ecological niche model development; and c. bar plot shows the different sources for the predictor data used in different articles. we plotted the bars having value more than one only. shariful islam et al. – ecological niche modeling applications to infectious diseases 82 figure s4: some of the selected components from the modeling section for a comprehensive understanding. a. bar plot shows the articles using different techniques used to preselect the predictor variables; b. bar plot shows the different techniques used to identify the variable importance in ecological niche modeling in different articles; and c. bar plot shows the different techniques used to check for multicollinearity among the predictor variable during ecological niche modeling in different articles. we plotted the bars having value more than one only. shariful islam et al. – ecological niche modeling applications to infectious diseases 83 serial no modeling techniques number of articles 1 maximum entropy modeling (maxent) 60 2 artificial neural networks (ann), surface range envelope (sre), flexible discriminant analysis (fda), general linear models (glm), general additive models (gam), general boosted models (gbm), classification tree analysis (cta), multiple adaptive regression splines (mars), random forests (rf), and maxent 1 3 ann, sre, fda, glm, gam, gbm, cta, mars, rf, and maxent 1 4 bayesian additive regression trees (barts) 1 5 bioclimatic envelope, maxent, logistic regression, and domain model 1 6 bioclimatic envelope (bioclim), maxen, glm, mars, cart, mixture discriminant analysis (mda), rf, brt, and gam 1 7 brt, gam, mars, and maxent. 1 8 bioclim, and bioclim true/false 1 9 sre, maxent, glm, gam, mars, cta, rf, and gbm 1 10 gbm, glm, rf, and mars 1 11 glm, mars, rf, and maxent 1 12 glm, gbm, rf, and maxent 1 13 logistic regression model 1 14 maxent, svm, environmental distance (ed) and climate space model (csm) 1 15 maxent, glm, cta, mars, svm, gam, gbm, and ann 1 16 maxent, rf, and brt 1 17 glm, gam, maxent, r), and brt 1 18 simple logistic regression model; rf 1 19 none 1 table s1: different ecological niche modeling techniques used to study the ecological niches of diseases shariful islam et al. – ecological niche modeling applications to infectious diseases 84 serial no sources of the predictor variables number of articles worldclim 36 chelsa 2 chelsa; isro 1 chelsa; isric; modis; nasa; soilgrid; worldgrids 1 merraclim 1 merraclim; soilgrids 1 merraclim; nasa; soilgrids 1 aafc; cwi; nrcan 1 agri-geomatics service of agriculture and agri-food canada; agriculture and agri-food canada; canada centre for mapping and earth observation 1 argentinian space agency 1 aws open data terrain tiles; iucn; prism; sric soilgrids 1 brazilian geomorphometric database; worldclim 1 calenviroscreen; gap; usgs 1 cgiar-csi; cgls; dhs; ethiopia’s survey data in 2016; isric; sdsm; worldclim 1 cgiar-csi; gridded population of the world, modis; worldclim; worldgrids 1 cgiar-csi; fao; hydrosheds; worldclim 1 cgiar-csi; worldclim 1 china resource and environment science and data center; worldclim 1 climatena 1 copernicus global land service archive; worldclim 1 dem; diva-gis; isric; modis; nasa; noaa; worldclim 1 dem; diva-gis; iucn; worldclim 1 digital soil mapof the world; sgs; sgs fews net; worldclim 1 dtm; worldclim 1 geographical survey institute 1 global land cover database; worldclim 1 government databases; worldclim 1 isric; soilgrids; worldclim 1 modis; nlcd; prism; tiger 1 nasa; spot4; worldclim 1 national land numerical information 1 national mapping organization; worldclim 1 pakistan bureau of statistics; worldclim 1 paleoclim 1 sedac; srt; worldclim 1 shuttle radar topography mission; worldclim 1 soilgrids; worldclim 1 srtm 1 srtm; worldlcim 1 usgs; worldclim 1 none 2 biodiversity informatics, 17, 2022, pp. 27-49 27 best practices for data management in citizen science: an indian outlook thomas vattakaven1*, vijay barve2, geetha ramaswami3, priya singh4, suneha jagannathan5, balasubramanian dhandapani6 1strand life sciences, ground floor, uas alumni association building, veterinary college campus, bellary road, bangalore, 560 024 2nature mates nature club, 6/7, bijoygarh, kolkata 700032 west bengal, india 3nature conservation foundation, 1311,”amritha”, 12th main vijayanagar 1st stage, mysore, 570 017. 4researchers for wildlife conservation (rwc), national centre for biological sciences, gkvk campus, bangalore, 560065. 5dakshin foundation, #2203, d block, 8th main, 16th d cross, sahakar nagar, bengaluru, 560092. 6french institute of pondicherry, 11, st. louis street, pondicherry, 605001. *corresponding author: thomas vattakaven, email: thomas.vee@gmail.com abstract. citizen science has been in practice since the 1800s and is an important source of data for scientists and other applied users. it plays a vital role in democratizing science, providing equitable access to scientific participation and data, helps build the capacity of its participants, inculcates the spirit of scientific endeavor and discovery and sensitizes participants towards species and habitat conservation, creating a sense of stewardship towards nature. in recent years, citizen science, especially in biodiversity, has rapidly developed with the rising popularity of smartphones, and widespread access to the internet, leading to wider adoption globally. india has also witnessed a surge in the number of new citizen science projects being initiated and increased participation in these projects. with more proponents looking at initiating such projects, there is little documentation from an indian perspective on setting up, collecting, managing, and maintaining biodiversity-focused citizen science projects, especially in a data-management context. we have attempted to fill this void by examining the best practices across the data life cycle of citizen science projects while keeping in mind sensitivities and scenarios in india. we hope this will prove to be an important reference for citizen science practitioners looking to better manage their data in their projects. key words: biodiversity, citizen science, data, data management, india, licensing, metadata, protocol, quality assurance, standards, data lifecycle citizen science has evolved as a significant field of practice, and its role in contributing to new knowledge on biodiversity is well established (e.g., kobori et al. 2016; schuttler et al. 2019). while originating in the west, it has spread globally, particularly in india in the last decade (sekhsaria and thayyil 2019). the evolution of citizen science as a field of practice and research has necessitated an inquiry towards overarching insights, standards, vocabulary, and guidelines (vohland et al. 2021). in 2020, india hosted its first citsci india conference for biodiversity1, bringing citizen science 1 https://citsci-india.org/. practitioners, researchers, educators, students, and policymakers interested in biodiversity and citizen science. this virtual meeting was hosted as part of the preparatory phase of the national mission on biodiversity and human well-being proposed by the biodiversity collaborative, with the national biodiversity authority as a nodal agency. citsci india 2020 was a starting point to bring together the citizen science community in india under one platform to share experiences, inspire each other, and engage in discussions related to citizen science. two prominent topics that surfaced during these discussions were the importance of ethics, diversity and inclusion, and mailto:thomas.vee@gmail.com https://citsci-india.org/ thomas vattakaven et al. – best practices for data management in citizen science 28 data in citizen science. in this context, we focus on best practices in data management for indian practitioners. following global trends, citizen science efforts involving biodiversity in india have rapidly gained pace over the past few years. with larger and smaller-scale citizen science projects increasingly launched in india each year, voluminous data is generated on various aspects of biodiversity (see list of indian citizen science projects2). however, this also raises several issues regarding data, such as ownership, accessibility, attribution, storage, interoperability, and quality. the working group on citizen science data was tasked with identifying significant aspects related to data on which project proponents should have clear procedures and policies. to put together this document, we surveyed existing global practices and standards and described various options that projects could adopt, with some guidance about benefits and costs associated with each option. the document is intended to form a toolkit for citizen science practitioners in india and elsewhere who seek to make informed decisions on various aspects of data. for this document, we use a definition of ‘citizen science’ provided by guerrini et al. (2019), as per which, citizen science “... generally refers to an approach to scientific inquiry in which members of the public participate in one or more steps of the research process other than, or in addition to, allowing personal data or biospecimens to be collected from them for analysis by others’’. at this juncture, it is worthwhile to mention that there is an effort to replace the usage of the word “citizen” with “community” science to be more inclusive (cooper et al. 2019), although some prefer to retain the distinction between these two terms (dosemagen and parker 2019). for the context of this paper, we primarily limit our reference to citizen science projects within biodiversity that at least partially utilize online participatory mediums with databases and servers that make data and its products accessible online. types of citizen science projects citizen science initiatives vary extensively in their aims and objectives across disciplines and citizen engagement. an attempt to classify common citizen science projects can be undertaken based on parameters such as the research question, modes of participation, medium of participation, and mode of the survey, as discussed in detail below. 2 https://citsci-india.org/citizen-science-projects/. research question/focus. although citizen science has traditionally been used to address targeted research questions and hence involves specified protocols, the advent of online mediums and the ability to crowdsource content has paved the way for more open-ended platforms. such platforms may engage citizen scientists in tasks such as gathering sightings of species or transcribing or classifying data for which the uses may be unknown or changing (lukyanenko et al. 2016). based on the above criteria, projects may be classified as generalist or specialist projects. an alternate way of looking at this type of focus may be to classify projects based on the taxa of focus. there are larger generalist initiatives that have little or no restriction based on the taxonomic focus (e.g., india biodiversity portal3, ibp), while targeted projects often tend to focus on selected or a single taxonomic group or species (e.g., biodiversity atlas india4, bird count india5, wild canids-india project6, marine life of mumbai7). modes of participation. citizen science initiatives can vary in terms of who initiates a project or the level and stage of involvement of volunteers or the general public in an initiative depending on the project’s objectives. project initiators play an essential role in defining the nuances of a project and, hence, determining the end goals that influence ‘the political authority of science’ (kimura and kinchy 2016). similarly, the composition and training of citizen science initiators vary across projects. for this document, we highlight different types of public-scientist collaborations that qualify as citizen science engagements based on veeckman et al. 2019. citizen science projects can be categorized based on multiple criteria: the extent of citizen participation, taxonomic focus of the project, or medium of participation. table 1 describes different types of citizen science programs based on the extent of involvement of citizens, as described in veeckman et al. 2019. medium of participation. currently, two distinct channels allow establishing a citizen science initiative. these include: 3 https://indiabiodiversity.org/. 4 https://www.bioatlasindia.org/. 5 https://birdcount.in/. 6 https://www.wildcanids.net/. 7 https://www.marinelifeofmumbai.in/. https://citsci-india.org/citizen-science-projects/ https://indiabiodiversity.org/ https://www.bioatlasindia.org/ https://birdcount.in/ https://www.wildcanids.net/ https://www.marinelifeofmumbai.in/ thomas vattakaven et al. – best practices for data management in citizen science 29 a. independent platforms via web or smartphone-based methods using protocols built explicitly for the project context. such platforms allow for flexibility in developing independent protocols tailor-made to suit the requirements of the study. b. larger aggregator platforms with the ability to host independent projects within them, e.g., india biodiversity portal, biodiversity atlas india, inaturalist8, citsci.org9. these platforms usually host a range of projects that collectively benefit from an existing user-base of citizen science contributors, are easy to use with access to pre-vetted guidelines, instructions of usage, terms, and conditions, and other legal and technical formalities addressed. they are also equipped with measures to ensure data security and data-quality regulations. all these features allow them to be used with ease across a diversity of projects and overcome the lack of technical know-how amongst project managers. a few initiatives are platform-independent and use social media and mobile messaging applications such as facebook, whatsapp, or email to gather biodiversity data. data collected through such mediums are primarily not structured by default, nor are they controlled environments with binding data policies or licenses. most of these are still emergent, and al8 https://www.inaturalist.org/. 9 https://citsci.org/. though there is potential to crowdsource content using these increasingly popular mediums, due to their free-form nature of the interaction, much effort will be needed to extract, curate, and cleanse the content before use as meaningful, structured data. mode of survey. citizen science may be carried out in a conventional scientific framework with a standardized field protocol. however, the most popular citizen science initiatives, especially those that allow for data entry through online interfaces and recruit online participation, are increasingly being done without standardized field protocols, giving rise to the term “opportunistic sampling.” we discuss below some pros and cons of opportunistic versus structured data sampling from the perspective of participant motivation and data quality. citizen science and data-life cycle like most scientific data, citizen science data follow a general data life cycle. for the purpose of this paper, we have chosen to adopt the high-level science data lifecycle model (sdlm) (faundeen et al. 2014), to illustrate how data management activities relate to citizen science data workflows and recommend actions and activities at each stage of the model. sdlm comprises primary elements that proceed sequentially and cross-cutting elements performed across stages of the data life cycle. the sdlm and its elements are summarized in figure 1. keeping in mind the above data model, we have attempted to structure our paper into broad sections table 1. types of citizen science programs based on the extent of participation by citizens, as described in veeckman et al. (2019). type of project extent of citizen participation crowdsourcing volunteers remain passive while contributing time and equipment only distributed intelligence volunteers are involved with simple interpretations or categorizing material from gathered data participatory science volunteers define a problem, collect data and assist scientists in analyzing the data. interpretation and analysis handled by scientists extreme citizen science volunteers and scientists collectively determine stages of the project, with the former handling all tasks related to the study and executing them. scientists only act as facilitators on these projects. contributory project volunteers are invited to contribute data, while scientists decide the research focus of the study, and analyze and interpret data. collaborative project flexible projects where the scientist involved may identify the research focus of a project, while volunteers participate at different stages of the study based on their interest co-created project aimed at influencing public policy or educational agenda. volunteers identify a set of questions, answers to which are thereafter pursued in consultation with scientists on the project. https://www.inaturalist.org/ https://citsci.org/ thomas vattakaven et al. – best practices for data management in citizen science 30 that cater to aspects of data within citizen science: before starting a project (planning), during the implementation of the project (acquiring), and after gathering data (processing, analyzing, preserving, and publishing data). certain aspects covered below have cross-cutting implications and may be relevant at multiple stages of the data life cycle but may be covered in more detail in one section to avoid repetition. before starting in conformity with scientific practice, before initiating a citizen science project, it is essential to premeditate on a project’s objectives, develop hypotheses, identify methods of acquiring, analyzing, and interpreting data, and ideal ways of disseminating results. potential challenges need to be identified to ensure that project end goals are achieved effectively. this section summarises points to keep in mind at different stages of a citizen science project. identify project goals and means of implementation citizen science programs vary in their primary goals and the extent and role of public participation, with some projects exclusively aiming to achieve public engagement (refer to types of citizen science projects). projects can have a broad focus, such as creating generic biodiversity repositories or targeting focused research questions. preliminary intensive review of the topic being pursued allows determination of the suitability of citizen science as a study technique. identify target participants/stakeholders delineating target participants helps one design appropriate strategies to recruit, train and engage volunteers for a program. these can be selected based on need (not all citizen science projects require a targeted volunteer base), required skill sets (such as swimming, diving, climbing, identifying species), access to technology (such as smartphones), or age (adults or children). at times, engagement with intermediaries (such as schools, colleges, tourism ventures) may be required to enlist participation. when soliciting involvement from local communities, planning for localizing content in regional languages can help enhance participation and outreach. building online and offline infrastructure the backbone of a citizen science project is the infrastructure that it requires to function: to maintain registers of participants; collect, manage, curate, figure 1: the science data life cycle model has primary and cross-cutting elements that help determine action at different stages of data collection, preservation and analyses. primary elements include planning, acquiring, processing, analyzing, preserving, and publishing data, while the cross-cutting elements run parallelly throughout the data life cycle and involve describing metadata, managing data quality, and data security/backup (faundeen et al 2014). the primary elements of the data life cycle can be further mapped to ‘before’ data acquisition (pink arrow), ‘during’ a project (orange arrow), and ‘after’ data are acquired (green arrows). thomas vattakaven et al. – best practices for data management in citizen science 31 and store data; and disseminate results and maintain regular communication with participants. online infrastructure such as websites, smartphone applications, or even simple forms of communication like whatsapp groups allows contributing data. suitable back-end databases that store data in appropriate formats and enable interfaces to query and retrieve data efficiently need to be chosen at this stage. establishing offline infrastructure requires high human effort, such as for collaborations, acquiring permits for access to protected areas (if needed), and initial outreach to gauge interest within target participants. it is also crucial to design data collection protocols (explained in greater detail in data collection methodologies) for citizen science projects and pilot them within small focus groups to devise appropriate data collection methods. volunteer recruitment, engagement, and outreach citizen science initiatives depend extensively on volunteer engagement to be successful. this engagement can be divided into three general phases, as follows: 1) volunteer recruitment involves outreach to the target participants, testing out protocols with focus groups, and seeking volunteer feedback on initial processes. typical methods employed are social media outreach, publicity articles in print media, tapping email list services, and presentations at target institutions such as nature clubs, schools, or colleges. in the case of projects targeting niche species that are uncommon or restricted in their distribution, establishing partnerships with local communities or tour operators may also be considered. 2) volunteer education and capacity building: in addition to gathering data, citizen science initiatives also endeavor to increase scientific and ecological literacy among the public. in some cases, volunteers might require specific knowledge (species identification skills, basic survey skills) to participate in a program. the extent of skill training and knowledge exchange often depends on the data collection methodology. volunteer education is a long-term exercise that needs to be carried out regularly. practitioners need to ensure that volunteers and contributors understand the scientific problem being addressed, are trained well in collecting information, can use technology (if any) required for data contribution, and collect data in a standardized manner. errors can be minimized by training and reiterating the collection protocol. the contribution process needs to be tested periodically to recognize new error sources and iteratively update training and contribution processes. 3) volunteer retention: citizen science efforts benefit from retaining volunteers over the long term, as their expertise and skill are likely to increase with time. however, this exercise requires innovative methods to sustain the interest of long-term volunteers. leaderboards (to track the highest participation) or games and contests may encourage participation. not all projects start with a captive volunteer base, and citizen participants may see turnover throughout the project duration. for long-term projects renewing interest in the project to recruit newer participants is crucial. data management, analysis, and dissemination maintaining, curating, and analyzing data are critical aspects of a citizen science program. data management involves data storage, curation, and backup techniques, ensuring that data are not lost once a project is deemed to be complete. it is important to note that citizen science is a long and evolving effort the goals of a project might change over its lifetime. considering this, one must follow data standards to maintain the usefulness of the data collected. furthermore, it is essential to ensure that data are analyzed and visualized by participants, including non-experts, in an engaging manner. mechanisms for user interaction with the data, roles, and permissions for data validators, strategies to flag erroneous data, etc., need to be thought of at this point and must be a continuous endeavor throughout the project, evolving with time. data collection methodologies citizen science projects vary in the rigor of sampling protocol, from simple occurrence reporting to more structured data collection techniques. this presents practitioners with a trade-off between volunteer participation and the quality of data collected. often, programs with rigorous volunteer training and sampling protocols obtain better quality data but with reduced levels of participation. in contrast, those with simple data collection methods report higher participation, often with biased and noisy data. techniques such as data collection forms or semi-structured surveys can reduce the need for rigorous training while still ensuring that data are collected in a prescribed format (bonney et al. 2009; kelling et al. 2019). incentivizing quality over the number of observations in leaderboards and gamification techniques can be helpful. gamification is used to motivate parthomas vattakaven et al. – best practices for data management in citizen science 32 ticipants to contribute data to maximize contributions and enhance volunteer retention. it can range from adding a point system to ranking, creating leader-boards, giving badges or rewards, to creating an actual game that requires enhanced engagement from the participant. prioritization of spatial and temporal scales where there are data gaps, rather than species or numbers of records, could result in more even distribution of biodiversity records, thus reducing spatial and temporal biases (callaghan et al. 2019). planning for data quality assurance the credibility and quality of citizen science data is often questioned even though recent studies have challenged this view (e.g., barve 2014). hence, data quality assurance is an important topic to consider at each data cycle stage, i.e., data collection, upload and ingestion, storage, management, and analysis. data should be accurate, precise, and representative to enhance credibility10, but when the desired accuracy and precision are not achieved, it is important to document its known quality as data quality and usability depends on the users’ questions (chapman 2005b,a). data needs to be thoroughly vetted for accuracy by curators or professionals pre-identified by the project and reflect reality. consistency and replicability can ensure data precision, while spatial and temporal representativeness is necessary for any scientific exercise. other ancillary information such as date, location, time of observation, weather conditions, etc. further aid in improving data quality and reliability. maintaining records with data provenance allows preserving information about the evolution of data and methodologies used to acquire it, which is vital for debugging, tracking changes, auditing, and evaluating quality. wiggins et al. (2011) categorize potential biases in ecological data obtained using citizen science. some of these include (a) positive spatial bias induced by areas with a higher human population, better or easier accessibility, and better accessibility to the internet in urban spaces (geldmann et al. 2016). this may also be true for frequently visited hotspots (boakes et al. 2016; tiago et al. 2017); (b) positive temporal bias induced by a more significant number of records during weekends, holidays, and contests, which is particularly of concern in phenology studies, where data seasonality is an integral part of the research (courter et al. 2013). (c) taxonomic bias is noticeable, with rare, cryptic, or difficult-to-identify 10 https://citizenscienceguide.com/design-sample-collection. species often being underrepresented or remaining unidentified leading to a paucity of data for such species (falk et al. 2019). in addition, very commonly observed species tend to be overlooked and underreported, while highly sought after species may also be over-reported, leading to non-representative sampling (troudet et al. 2017; callaghan et al. 2021). finally, (d) observer bias induced by individual perceptions and levels of experience (gonsamo and d’odorico 2014; callaghan et al. 2021). identifying sources of bias is necessary to help managers design appropriate strategies at different phases of projects to mitigate or manage them effectively. a variety of end-users can utilize data from citizen science projects and it is important to identify the likely end-users of the data (e.g., scientists, policymakers, amateur naturalists) at the initial stages of a project. it should be noted that the onus of maintaining data quality is the shared responsibility of the project as well as data-users who must examine and put thought into the use of data and correct for biases, error rates, or quirks. data quality and minimization of biases can be accounted for before data collection and during the data contribution stage. baker et al. 2021, summarize the types of data and levels of evidence at the data contribution stage which would require verification a) levels of evidence: reporting of sightings without evidence, and reporting sightings with evidence such as photo/video/audio/specimen, b) types of observations: direct observation, where the taxon is observed directly, and indirect observation, wherein taxon signs (such as tracks, dung, etc.) are recorded. the mechanisms and criteria required for data validation need to be considered at this stage and suitably incorporated into the collection procedure to provide for the availability of fields or the target precision levels to be achieved. the ability to validate or curate records may be contingent on the presence of such information fields and without which data may be unverifiable. data analyses should also be planned and anticipated before data collection and should be appropriate for the kind of data collected and driven by the project’s goals (wiggins et al. 2011; balázs et al. 2021). factors affecting data quality need to be identified. this may include improper data collection, incorrect implementation of data collection protocols, a mismatch between project goals and data collection protocols, incomprehensive protocols that do not match end-user expectations, and use of data thomas vattakaven et al. – best practices for data management in citizen science 33 in wrong contexts. balázs et al. (2021) suggest the following at the planning stage of a citizen science project to ensure data quality and to make data conducive for further analyses: (a) simple and intuitive data collection protocol supplemented by a simple user interface design that is engaging and can be applied across a diverse group of users with varied skills, (b) calibrating and standardizing devices and recognizing limitations of technology, (c) appropriate documentation, and (d) metadata to prevent misuse of data in incorrect contexts. conferring with experts could enhance the quality of analyses. inferences should be cautious and consider all the caveats of data accuracy and analysis. it is also beneficial to get the analyses reviewed by experts and peer groups. quality assurance through following standards. it is essential to incorporate data collection methods and protocols, fitness for use, and data quality assessment as part of the metadata/documentation (assumpção et al. 2018). adapting and adhering to standards helps in improving data quality and usability due to breaking up data attributes into appropriate terms and following controlled vocabularies to ensure each term conveys the correct meaning. in the biodiversity realm, standards developed by the biodiversity information standards, originally called the taxonomic databases working group (tdwg11), like darwin core and audubon core (table 1) and tools built around them, are readily available for citizen science projects to use and adapt. planning for data infrastructure contemporary citizen science projects are mostly born-digital, conceived and implemented predominantly in the digital ecosystem of information technology platforms, software applications, and toolchains. there are multiple online and offline infrastructure concerns. proponents may wish to choose between larger aggregator platforms that allow projects within them vis-a-vis building independent applications. fast-growing mobile application technologies have made it possible to quickly deploy data collection and integration tools with little effort (lemmens et al. 2021). on the other hand, several large biodiversity data aggregating platforms are well established and have gained a reputation across the globe (ebird12, inaturalist), or country-level data (india biodiversity portal and biodiversity atlas for 11 https://www.tdwg.org/. 12 https://ebird.org/home. india, atlas of living australia13 for australia, national gbif nodes, etc.). others look at specific taxa or simply to address a particular question. there are apparent advantages of using an existing platform for biodiversity data collection and aggregation, as they readily provide technological infrastructure, communities, and tested infrastructures across data lifecycle (de sherbinin et al. 2021). depending on the project’s larger goals, there could be challenges in fitting the needs of a citizen science project to pre-existing templates and applications provided by such platforms. such larger-scale citizen science initiatives should allow for flexibility in engaging at different ecological levels, different aspects of ecosystem changes, and conservation issues (devictor et al. 2010). many large aggregator platforms already support infrastructure that allows such flexibility. for example, the ibp allows creating groups within its infrastructure for any theme of interest, such as a taxonomic group. forms for gathering data can be extended with the capability to include custom queries and fields, as required. data infrastructures for citizen science projects need to be adaptable to address the unique nature of each citizen science project. the data infrastructure should generally allow for data collection, aggregation, analysis, and dissemination, thus covering the whole data life-cycle management or digital information supply chain (brenton et al. 2018). for example, citizen science projects could use a phone or web application to collect data, a cloud server to store the data, automated code to verify data, a web portal to promote interaction among contributors, and a backend database structure such that it can be aggregated with other types of data. this would also mean that the infrastructure enables the participation of citizen scientists in the full range of scientific methods from problem definition, research design, analysis, and action (mcquillan 2014). citizen science projects in biodiversity tend to collect data across a wide assortment of attributes. these include taxonomic, evolutionary, biogeographic, functional, and interspecific interaction attributes of a taxon (könig et al. 2019). data infrastructure should be flexible enough to accommodate the diversity of data types in a variety of formats such as text, tabular, geo-spatial, and varying media types, including images, audio, and video. such capabilities will influence the scalability of storage required, particularly for long-term projects. cloud-based storage and content delivery networks in the mainstream it 13 https://www.ala.org.au/. https://www.tdwg.org/ https://ebird.org/home https://www.ala.org.au/ thomas vattakaven et al. – best practices for data management in citizen science 34 ecosystem have matured enough to ensure scalability and high availability across geographies. apart from these fundamental concerns on data models and data storage, one must ensure that the platform is stable and provides continuous access to participants with minimum downtime. platforms need to cater to data security with well-defined data access policies, user authentication systems with defined roles, transparent workflows, and user-centered design (bowser et al. 2020). regular backup of data with multiple copies in multiple locations and a consistent preservation policy across sites is essential for the security and integrity of citizen science data. in keeping with the spirit of open science, one can also insist on using, developing, and deploying free/ open-source technology stacks to help collaboratively build, share and replicate developed technologies for wider and unrestricted use. data ownership data ownership is crucial for a project and needs deliberation at the planning stage. participant perception of data ownership can influence their motivation in future participation. yet, studies have indicated ambivalence in how participants feel about data ownership. on the one hand, most participants appear far removed from thoughts of data and its ownership. each record is more of a personal nature experience and less so as data with legal ownership. ganzevoort et al. 2017, best summarize this as constituting “an “imagined contract” between volunteer naturalists and nature, based on respect and wonderment…”. on the other hand, participants believed that “...data extracted from nature should properly be used towards its preservation” and hence “wrong” use of data can result in citizens being resentful and withholding contribution. in some instances of data sharing, moral rights may get infringed, especially if the user of such data distorts or mutilates the data contributed by volunteers through re-use or if some private/sensitive information gets accidentally disclosed. although most participants surveyed were against the unconditional use of data generated using citizen science, most participants were undecided on ownership, with some feeling data is nobody’s property and others that it could be owned by the organization conducting the study (ganzevoort et al. 2017). it may also be said that participants may feel strongly about data in ways that are not covered under legal ownership and may not qualify for legal protection. however, it may be possible to validate such feelings outside of traditional law through policies that put in practice exclusive or non-exclusive access to or control over data (guerrini et al. 2019). sometimes traditional knowledge belonging to communities related to bio-resources or conservation practices might be part of such data. establishing who owns this knowledge can be very challenging. in the case of community-held knowledge, it is easy to attribute ownership to a particular community. still, there is ambiguity when the traditional knowledge is from an unidentifiable source or shared between communities spread across large territories. the knowledge may also be based on specific practices, beliefs, and linguistic representations, which may get lost in translation. one must be mindful of specific communities’ cultural sensitivities and secretiveness to divulge knowledge. the communities must have the freedom to say no to sharing their knowledge if they wish, and if they agree, they should be allowed to choose how their knowledge is used. data accessibility it is essential to have prior clarity on data accessibility regarding who can access the data, at what stages, and for what purposes. accessibility to data generated through citizen science projects is a core aspect. open access to data allows for democratizing science and upholding the values of universal and equitable access to scientific data, especially when gathered through public participation. just as we strive to make citizen science accessible to a diversity of participants and make ‘doing science’ inclusive, the resulting data must be accessible in ways that support reproducible science and can influence policy through bridging gaps between knowledge and action. there are outlying concerns that, more often than not, a citizen scientist’s contribution disappears into the closed databases within institutions, and particular emphasis needs to be paid to alleviate these concerns. what is open data? there are variable interpretations of the term ‘open data’. as stated by the open knowledge foundation, “data is open if it can be freely accessed, used, modified and shared by anyone for any purpose14 subject only, at most, to requirements to provide attribution and/or share-alike”. specifically, open data is defined by the open definition and requires that the data be, both, a) legally open: where it is made available under an open (data) license that allows anyone to freely access, reuse and 14 https://opendefinition.org/. https://opendefinition.org/ thomas vattakaven et al. – best practices for data management in citizen science 35 redistribute the data; b) technically open: where the data is made available freely or at a cost, no more than what is required for its reproduction and in formats that are in bulk and machine-readable. therefore, open data means that it is complete, preferably downloadable over the internet in a convenient and modifiable format without requiring proprietary software to process. it should also be “provided under terms that permit reuse, redistribution, allow intermixing with other datasets and must not discriminate against fields of endeavour or against persons or groups such as against commercial use.” in this context, the data should conform to the fair open data principles to be findable, accessible, interoperable, and reusable. why is citizen science data not always open? due to the varied nature of citizen science projects, proponents, and contexts of funding, not all projects may likely be in a position to adhere entirely to the tenets of open data. some disagree with open data with justifications that vary in their context. with data serving as the currency for competition between scientists for limited funding and prestige through publications, the conventions of traditional academic publishing have resulted in a tendency to hoard data in closed silos (hampton et al. 2013). others cite the burden and expense of running massive data projects, curating data and processing, and managing people involved as a justification for exclusive access and reaping the resulting benefits (walker et al. 2016). the data can be used as leverage to fund further project activities or, more importantly, to obtain acknowledgment, particularly as authors on publications. some might be willing to share data but on request to keep track of how it is being used, hence not publishing them under an open-access license. other important reasons for data not being open are projects anchored at institutions having restrictive blanket data policies, especially concerning intellectual property rights for work generated as a part of the institution. similarly, funding bodies sometimes impose conditions on data release as a part of their terms, which may be restrictive. finally, privacy concerns regarding those of the participants generating the data and when the data is about a species of concern may be a key consideration in limiting open access (groom et al. 2017). why it is recommended that citizen science data be open-access. there are many reasons for recommending citizen science data be open-access. groom et al. 2017, state that “the voluntary aspect of the time invested by citizen scientists is generally interpreted as being motivated primarily by its contribution to society and that society should profit from this effort through openly accessible data.” open-access allows participants to track their participation alongside aggregated data from other participants, learn from it, and incorporate the learning into improving their knowledge. opening the data has also been shown to motivate participants with greater frequency and depth (bonney et al. 2009). the availability of open data allows easy and quick access for citizens and decision-makers to use as evidence towards influencing policy without waiting for formal assessments to emerge and closing the gap between knowledge and action. open citizen science data thus enable participants to be at the “forefront of socially relevant science” (hampton et al. 2013). open data also supports reproducible science. ethical considerations to be made at this stage some of the key ethical considerations in the planning stage of a citizen science initiative cover the realms of recognizing contributor rights in citizen science and designing socially inclusive projects. it also includes information on making data public or open-access versus limiting access (addressed in data accessibility). with the growing popularity of citizen science across geographies and academic disciplines, it is being rapidly incorporated as a methodology to identify scientific queries and find means of answering socially relevant questions with the aid of citizen contributors or collaborators. this has led to a growing recognition of contributor rights and the need to address power imbalances between project handlers and contributors. hence, project managers must recognize the participatory nature of citizen science, where volunteers contribute data and time obligingly and ensure that projects are socially inclusive. projects must include participants irrespective of gender, geographical location, socio-cultural, religious, linguistic, and academic backgrounds (paleco et al. 2021). simultaneously, participation at all stages of a project must be permitted. to maximize participation, project designers should identify means of reaching out to all potential stakeholders. these could include interested citizen participants, such as members of the public, established citizen scientists or those from the scientific fraternity, academic institutions/ organizations, policy experts, etc. (veeckman et al. 2019). collaborating with schools, thomas vattakaven et al. – best practices for data management in citizen science 36 communities directly associated with the study subject, or government bodies also helps in increasing participation (veeckman et al. 2019). project implementation data and metadata standards and data infrastructure once the study design, data quality, adherence to standards, data infrastructure, and accessibility are planned for in a citizen science project, implementation is next. at this stage, data infrastructure should be appropriate, scalable, and highly available for organizing the acquired data. it is also essential to incorporate data standards at this stage. biodiversity data standards are shared rules and conventions to describe, record, and structure biodiversity data to enable data aggregation and exchange across different organizations generating and managing different data sets. data standards enforce unambiguous definitions of what kind of data are being collected, follow well-defined ontologies and vocabularies, and standardize the usage of established protocols. it is recommended that citizen science projects follow prevalent international standards and adopt recommended storage formats and protocols. the use of biodiversity data standards has two key objectives, as follows. (1) data standards provide a comprehensive set of relevant attributes for most projects and meeting individual project needs for the collection and management of data. (2) data standards aid in identifying a subset of core biodiversity data attributes that can be used to aggregate data. the different international standards available for citizen science data are described in table 2. user-contributed data are typically restructured slightly to adhere to project-specific standards before storing in databases. however, this standardization and large-scale aggregation may lead to a loss of contextual richness (turnhout and boonman-berson 2011; ganzevoort et al. 2017). apart from adhering to standards, it is good practice to ensure that data are findable, accessible, interoperable and reusable (fair, wilkinson et al. 2016) and consider other principles like collective benefit, authority to control, responsibility and ethics (care, carroll et al. 2021) in the context of data originating from indigenous communities. the fair data principles facilitate the discovery of knowledge, integration, and use by the larger scientific community. fair principles ensure that data are discoverable to computational agents (e.g., computers and apps) and humans through standardized protocols – otherwise referred to as ‘machine actionable’ data (wilkinson et al. 2016). in highly linguistically diverse contexts, such as in india, data infrastructure should be developed with support for multiple languages to enable data input and support accesses to facilitate large-scale participation. quality assurance and quality control citizen science projects with high data quality adhere to good practices related to standards, metadata, and documentation. such projects also ensure that data errors, biases, uncertainty, and ethical concerns are addressed through volunteer training and validation, calibration of data collection tools, iterative evaluation, and enhancements, and flagging of erroneous records via appropriate validation methods (kosmala et al. 2016; ratnieks et al. 2016; downs et al. 2021). while integrating data from different citizen science programs, incompatible design and inconsistencies in nomenclature can also affect data quality (campbell et al. 2020). the following fail-safes can ensure that data collection is accurate before and during data collection: profiling contributors and assessing their skill levels, piloting a citizen science project to get a sample of data and potential sources of errors and biases, following standardized methods of data collection and adopting established standards for terminology, participant training, auto-correction (e.g., erroneous geocoding), data verification, and facilitating access to data use (balázs et al. 2021). projects using devices should calibrate sensors, perform initial checks on devices and ascertain the ability of observers to use these devices to make accurate observations (de sherbinin et al. 2021). depending on the types of observation, post-collection data verification is a crucial step to ensure data accuracy. community/peer consensus, expert verification, automated verification, model-based verification (statistical models address random/individual variation, residual errors, and uncertainty of devices to flag erroneous observations), and linked data analysis (combine existing datasets to serve as reference data and use data mining tools to flag erroneous observations; kelling et al. 2015; balázs et al. 2021) are some of the existing methods that can be used for assessing data quality. in biodiversity data, species (identity, geographic co-occurrence with other species, rarity), environmental (time, date, location) and expertise (experience of the recorder) contexts should be verified through one or more methods thomas vattakaven et al. – best practices for data management in citizen science 37 standard description url darwincore (dwc) “a glossary of identifiers, labels, and definitions that facilitate the sharing of biodiversity information. dwc is based on taxa and their distribution documented through observations, specimens, samples, and related information. it is being regularly improved with the addition of terms as well as the development of extensions to map various sources of data accurately.” https://www.tdwg.org/standards/dwc/ audubon core multimedia resources metadata schema (ac): “a set of vocabularies designed to represent metadata for biodiversity multimedia resources and collections, with the aim of determining the suitability of the media for specific biodiversity science applications. among others, the vocabularies address such concerns as the management of the media and collections, descriptions of their content, their taxonomic, geographic, and temporal coverage, and the appropriate ways to retrieve, attribute and reproduce them.” https://www.tdwg.org/standards/ac/ the access to biological collections data (abcd) “an evolving comprehensive standard for the access to and exchange of primary biodiversity data (i.e. specimens and observations)” https://www.tdwg.org/standards/abcd/ ecological metadata language (eml) “defines a comprehensive vocabulary and a readable xml markup syntax for documenting research data. eml includes modules for identifying and citing data packages, for describing the spatial, temporal, taxonomic, and thematic extent of data, for describing research methods and protocols, for describing the structure and content of data within sometimes complex packages of data, and for precisely annotating data with semantic vocabularies.” https://eml.ecoinformatics.org/ taxonomic concept transfer schema (tcs) “a schema to allow the representation of taxonomic concepts as defined in published taxonomic classifications, revisions and databases. it specifies the structure for xml documents to be used for the transfer of defined concepts. currently, this standard is not followed widely.” https://www.tdwg.org/standards/tcs/ table 2. international standards for citizen science data. https://www.tdwg.org/standards/ac/ https://www.tdwg.org/standards/abcd/ https://www.tdwg.org/standards/tcs/ thomas vattakaven et al. – best practices for data management in citizen science 38 (baker et al. 2021). observer biases can also be accounted for in the post-collection stage through data filters and models that account for levels of contributor expertise and using ai-based techniques to reduce biases in model training data (johnston et al. 2018; chen and gomes 2019; steen et al. 2019). data that pass through the validation stages need to be curated. this involves processing raw data in terms of end-user requirements, ensuring that data meet reproducibility standards (for analyses), and lending themselves well to being combined with other standardized datasets. if the end-use of the data includes re-use or integration, data credibility can be increased by doing analyses on sampling approaches and quality and triangulating against other data sources. it is ideal to store citizen science data in the most disaggregated form with minimal privacy concerns and documented data quality assurance protocols (de sherbinin et al. 2021; downs et al. 2021). licensing licensing is necessary to ensure that media contributed by users to a citizen science project is used appropriately and data is appropriately cited. copyright is a state-guaranteed right covering ‘work’ that includes intellectual creations, such as text, photographs, diagrams, maps, movies, etc. ideas, knowledge, information, or data are traditionally not copyright-protected, and scientists have traditionally been content with being cited for their original work (hagedorn et al. 2011), to facilitate public access and dissemination of knowledge. although it is commonly assumed that data with no license applied is free for unrestricted use, this is not the case. the lack of a license poses ambiguity in its reuse, especially where the data usage terms have to be made explicit, especially for commercial usage (groom et al. 2017), and may lead to unwitting copyright violations. data must be made available under carefully crafted licenses where the terms and conditions for its reuse are made clear. adopting open, machine-readable licenses are recommended to meet the fair data principles discussed earlier (de sherbinin et al. 2021). the most common license employed in citizen science data is the creative commons license (cc15). this license seeks to find a ‘balance between public and private interests, and between the free flow of expressions of ideas and knowledge and state-guaranteed control and monopolies’ (hagedorn et al. 2011). creative commons licenses are not an alternative to copyright 15 https://creativecommons.org/. and work alongside copyright, enabling one to modify copyright terms to best suit their needs. a violation of a cc license is a copyright violation. the cc licenses provide standardized terms-of-use definitions that have been adapted for various jurisdictions and upheld in court in several countries (hagedorn et al. 2011). the license has been adapted for india under the aegis of wikimedia india, centre for internet and society, and acharya narendra dev college16. cc licenses by default allow people to reuse, remix and adapt original works while still providing attribution to the original author. however, it understands that no single license can cover all use cases and instead offers a set of licenses to cover a wide range of use cases (table 3). the cc licenses are accordingly adopted in whole or part by large data repositories such as the global biodiversity information facility (gbif17 ), wikipedia, and wikimedia commons, among others. cc0, cc‐by, and cc‐by‐ nc are the only cc license options recommended by gbif. as scientific data is primarily facts and is not copyrightable, cc0 is the recommended license for data18. if any media are contributed as part of the data, the terms of use of the platform gathering the data should be clear on the applicability of the cc license to such media as well. another such relevant license is the open data commons (odc) maintained by the open knowledge foundation19. although open data commons licenses are more suitable for data licensing, they are more specific to databases and apply only to database frameworks and structures, not to the particular content within a database. it allows for the “distinction between the data (base) and material (content) generated from it (“produced works”)”. odc provides three types of licenses (table 4). india’s open government data initiative started with the notification of the national data sharing and accessibility policy (ndsap), by the department of science and technology to the union cabinet in 2012 and the subsequent launch of the open government data platform india. the recommended licenses to be used for datasets published under ndsap through the ogd platform remained unspecified until the release of the government open data 16 https://wiki.creativecommons.org/wiki/india. 17 https://www.gbif.org/. 18 https://wiki.creativecommons.org/wiki/cc0_use_for_data. 19 https://opendatacommons.org/. https://creativecommons.org/ https://wiki.creativecommons.org/wiki/india https://www.gbif.org/ https://wiki.creativecommons.org/wiki/cc0_use_for_data https://opendatacommons.org/ thomas vattakaven et al. – best practices for data management in citizen science 39 icon right description attribution (“by”) is a part of all cc licenses and requires users to give appropriate attribution to the creators of a work share-alike (“sa”) allows the distribution of derivative works, but requires that all such works must also be shared under the same conditions ensuring that more restrictive licenses are not applied to derivatives. no derivative works (“nd”) states that the user “may not alter, transform, or build upon this work” non-commercial (“nc”) states that one “may not use this work for commercial purposes” types of cc licenses icon name applicable rights cc by22 attribution cc by-sa23 attribution-sharealike cc by-nd24 attribution-noderivatives cc by-nc25 attribution-noncommercial cc by-nc-sa26 attribution-noncommercial-sharealike cc by-nc-nd27 attribution-noncommercial-noderivatives table 3: creative commons rights and licenses for data. the creative commons logo and icons used are from wikipedia21 and are under the public domain. name applicable rights open data commons open database license28 (odbl) attribution share-alike open data commons attribution license29 (odc-by) attribution open data commons public domain dedication and license30 (pddl) public domain (all rights waived) 2021222324252627 20https://en.wikipedia.org/wiki/creative_commons_license. 21https://creativecommons.org/licenses/by-sa/4.0. 22https://creativecommons.org/licenses/by-nd/4.0. 23https://creativecommons.org/licenses/by-nc/4.0. 24https://creativecommons.org/licenses/by-nc-nd/4.0. 25https://opendatacommons.org/licenses/odbl/. 26 https://opendatacommons.org/licenses/by/. 27 https://opendatacommons.org/licenses/pddl/. table 4: open data commons licenses. https://en.wikipedia.org/wiki/creative_commons_license https://creativecommons.org/licenses/by-sa/4.0 https://creativecommons.org/licenses/by-nd/4.0 https://creativecommons.org/licenses/by-nc-nd/4.0 https://opendatacommons.org/licenses/odbl/ https://opendatacommons.org/licenses/by/ https://opendatacommons.org/licenses/pddl/ thomas vattakaven et al. – best practices for data management in citizen science 40 license india28, which is governed by indian law. it allows end-users to “use, adapt, publish (either in original, or in adapted and/or derivative forms), translate, display, add value, and create derivative works (including products and services), for all lawful commercial and non-commercial purposes.” however, the terms of the license remain ambiguous and have been criticized for being incomplete in many aspects, such as privacy and accountability of data providers (kodali 2017). in addition, it is possible to set up custom bespoke licenses for a citizen science project. however, this is not a trivial endeavor and will almost certainly have to include legal offices and organizational research departments (ball 2011). such cases are usually unnecessary considering the availability of standard licenses as documented above, except when exceptional circumstances require the same. creating additional bespoke licenses adds to the burden on end-users of the data in ensuring compliance and adhering to multiple license requirements. once a suitable license has been decided upon, one must attach that license to the data. this mainly involves putting out a statement that the data is released under the chosen license or public domain and a mechanism for retrieving the full text of the license itself. the rights statement must be displayed prominently to avoid ambiguity and confusion. adding the rights statement within downloaded zip files in an rdf/xml format for machine recognition is highly recommended (ball 2011). ethical considerations at this stage some of the key ethical considerations in the stage of data acquisition include clear, prior communication with potential participants before collecting data, information on data licenses, encouraging participants to contribute data collected following fair practices, and legal and social conformity to data being incorporated from indigenous communities. prior communication regarding objectives of a project, terms of data usage, methods for data storage, recognition of participant roles, etc., should be communicated to participants at one or more stages during project implementation. before participation, participant consent should be sought to ensure that contributed data are not misused, and participants are aware of data licenses. while participant data-usage trends can be used to communicate about project developments, unauthorized usage of participant personal in28 https://data.gov.in/government-open-data-license-india. formation should be prevented (sullivan et al. 2014). project managers need to be aware of regional laws and other legal components that govern the usage of information pertaining to indigenous communities. traditional knowledge must be handled sensitively and collected only after receiving consent from indigenous groups. national laws related to copyright or protection of imagery and text narratives must be well understood before accessing such data and complied with. efforts must be made to ensure project participants abide by government and community laws and regulations while accessing contributed data. in the case of biodiversity projects, the safety of biodiversity and participants should be a priority over data collection. gaming elements are often used in citizen science projects to positively influence participant engagement by creating an environment of fun, competition or both (bowser et al. 2013; iacovides et al. 2013). gamification may reward participants for attaining high scores and can enhance user participation. however, it may also demotivate participants who do not achieve competitive targets, distract participants from scientific data collection, trigger unfair practices to inflate competitive scores, and withhold ‘winning’ information, conflicting with citizen science principles of open data. skills and resources required to participate in games may create inequity by putting some participants at an advantage over others. ponti et al. (2018), note that game design and its influence on participant strategies, contributor values and motivations, and acknowledgment of participation outside of the gaming context are essential aspects of citizen science gamification. at the data-validation stage, artificial intelligence techniques such as deep learning and convolutional neural networks are now used in citizen science projects to classify images, especially for species identification. ai applications may be used to automatically classify visual, acoustic, and spatial information through learning algorithms that utilize vast datasets and extract and organize images from social media (lamba et al. 2019; august et al. 2020). here, ethical challenges arise when black-boxed artificial intelligence systems are trained with citizen science contributed open data but exclude citizens from understanding how their data contributions are used. transparency in such ai systems is essential and can help detect biases in training datasets, thus improving the efficacy of these systems. for example, ebird’s human/computer learning network https://data.gov.in/government-open-data-license-india thomas vattakaven et al. – best practices for data management in citizen science 41 (kelling et al. 2012) is cited as an example of such a transparent system (mcclure et al. 2020). what to do with the data once data begins accumulating, considerations related to data storage, processing, analysis, and dissemination in standardized and verified forms and ensuring its longevity have to be implemented. data infrastructure managing physical risks associated with data storage, ensuring security, and ensuring access to data across its lifecycle is critical. the usgs data lifecycle model recommends that such security measures cover “raw and processed research data, original science plan, data management plan, data acquisition strategy, processing procedures, versioning, analysis methods, published products, and associated metadata” (faundeen et al. 2014). the bouchout declaration for open biodiversity knowledge management29, which aims “to promote free and open access to data and information about biodiversity by people and computers and to bring about an inclusive and shared knowledge management infrastructure,” lists some core principles that are vital towards data perseverance: ● an agreed infrastructure, standards, and protocols to improve access to and use of open data; ● persistent identifiers for data objects and physical objects such as specimens, images, and taxonomic treatments with standard mechanisms to take users directly to content and data; ● tracking the use of identifiers in links and citations to ensure that sources and suppliers of data are assigned credit for their contributions; ● registers for content and services to allow discovery, access, and use of open data; ● linking data using agreed vocabularies, both within and beyond biodiversity, that enable participation in the linked open data cloud. these approaches to data lifecycle management point to implementing various strategies pertaining to aggregation and processing of data for analysis, transformative action, and tracking usage. multiple techniques in deploying persistent identifiers and urls, digital object identifiers, lifescienceid, and personally identifiable information are adopted across platforms to ensure that data in its various 29 http://www.bouchoutdeclaration.org/declaration/. types and stages are traceable. implementation of such persistent identifiers will become a norm soon and will help ensure data quality, access and accreditation. trustworthy data repositories like zenodo/dryad, mendeley data, among others, could be considered for storing citizen science data that has been curated for research quality. developed as part of efforts from research data alliance, a set of harmonized common requirements for certification of research data repositories certifies that these remain trustworthy (coretrustseal standards and certification board 2019). for occurrence data, global repositories like gbif, ebird, and ibp in india could act as apt data repositories to ensure the perpetuity of data. while many such data repositories are evolving with long term ecological observatories30 and other state-sponsored initiatives, it is pertinent to note the significance of archiving citizen science initiatives with their raw data and the context within which they are conducted (williams et al. 2018). this will ensure the dual goals of securing the perpetuity of citizen science data and maximizing re-use. such public data archiving for citizen science initiatives are required but a challenge to build (pearce‐higgins et al. 2018). data standards data standards play an important role in biodiversity data publishing. following data standards makes publishing either through aggregators like gbif or in the form of data papers simple. it saves effort in describing metadata and makes published data readily usable for the intended user base. data papers and data repositories often require the metadata to be marked up in standardized formats such as eml. independent projects may use software such as r or morpho31 to markup the metadata from their datasets. many larger platforms like inaturalist or ibp serve as an archive and a publishing platform. they already have some standardization inbuilt within their structure, allowing data downloads to be served under such standards. such platforms also have arrangements on publishing the data to global biodiversity repositories such as gbif through common standards. data accessibility as stated earlier, open access data means that “data must be freely available for download online.” 30 https://lteo.iisc.ac.in/. 31 https://old.dataone.org/software-tools/morpho. http://www.bouchoutdeclaration.org/declaration/ https://lteo.iisc.ac.in/ thomas vattakaven et al. – best practices for data management in citizen science 42 this also implies that the data is accessible in formats that do not need proprietary software to open and must have an open license for reuse. the csv format is generally used for tabular data download and ensures compatibility for machine-reading of the data in a machine-readable format. many sites require prior registration or serve download requests via a user’s registered email. imposing registration for data downloads is an accepted means of tracking data usage, ensuring compliance with the project’s policies and the site’s data licensing. from the accessibility perspective, there is a need to involve citizens beyond the mere act of data collection and provide them with opportunities and incentives to interact with the data they have generated. participants are rarely given opportunities beyond data collection, such as data analysis or interpretation (kennett et al. 2015; lukyanenko et al. 2016). activities such as data consumption influence learning and conservation outcomes and may lead to better user retention in the project (cooper et al. 2017). such interaction can be achieved through participatory data analysis and visualization that can be user-generated as per their needs and variables of interest. many projects are increasingly gravitating towards developing such interactive visualizations for participant engagement. however, since data analysis is usually an end-user’s specific perspective, generic visualizations and analyses inbuilt into portals may be limited as they are usually set up to predefined criteria. such limitations can be overcome through developing and offering apis and client packages for popular data analysis software such as r or python. some examples of such packages are the ‘rgbif’32 and ‘pygbif’33 clients for interfacing with gbif and the ‘galah’ r package34 for acquiring data from the atlas of living australia. this capability would allow users to fetch data flexibly, do further analysis and generate custom visualizations as per their needs. dissemination of knowledge gathered through citizen science citizens should not be viewed only as data contributors in the scientific endeavor; they are also the end-users in many situations. while the purpose of a citizen science project may vary (publishing a scientific paper, data repositories, outreach to the public, etc.), knowledge generated through citizen science 32 https://cran.r-project.org/web/packages/rgbif/index.html. 33 https://github.com/gbif/pygbif. 34 https://atlasoflivingaustralia.github.io/galah/index.html. must find its way back to its contributors. citizen science participation can be enhanced by incorporating clear channels of communication and data dissemination (vohland et al. 2021). this allows access to a wide audience, makes people aware of the project, and keeps them in continued engagement with the project. the traditional means of disseminating scientific knowledge through publication in peer-reviewed journals can often be too technical for the lay public to understand. involving the public in science is one of the core principles in citizen science, and hence knowledge should also reach the public in a digestible manner. this can be done through activities such as creating data visualizations to communicate results attractively e.g., ebird status and trends abundance animations that reveal migratory pathways of birds35.; writing articles in popular media sources like newspapers, magazines, and online magazines; visual communication of knowledge through art, videos, and graphic design and using social media to disseminate results. it is worth noting that disseminating knowledge to the public is crucial to ensure long-term participation and collaboration in any citizen science program through various means, targeting multiple stakeholder communities. however, excessive emails or other means of contacting participants can adversely affect and discourage participation. data attribution attribution is the act of giving credit to data providers during publication. author attribution has historically been a tricky issue across disciplines and this has only been accentuated with the advent of big data and data papers with proper guidelines on giving authorship not being stabilized even today (venkatraman 2010; escribano et al. 2018). while protocols such as the science commons advocate publishing data openly, there is no mention of providing attribution. authors typically negotiate their order within the author list, assuming that the first author is the most coveted and has led the publication idea. the last typically is the head of the lab and the point of contact (venkatraman 2010). as the contributor list grows, especially in large collaborative projects, the contribution order becomes less understandable and meaningless. some journals provide a separate text or list stating individuals’ roles and contributions instead of authorship. there is much ambiguity in citizen science as to who should get attributed and how and whether indi35 https://ebird.org/science/status-and-trends/abundance-animations. https://cran.r-project.org/web/packages/rgbif/index.html https://github.com/gbif/pygbif https://atlasoflivingaustralia.github.io/galah/index.html https://ebird.org/science/status-and-trends/abundance-animations thomas vattakaven et al. – best practices for data management in citizen science 43 vidual citizens will be acknowledged in publications. the joint declaration of data citation principles (crosas 2013), states that when cited, there should be ‘legal attribution to all contributors to the data, but recognizes that a single style or mechanism of attribution may not be applicable to all data’. however, large datasets or data involving many contributors, such as citizen science data, are prone to the issue of ‘attribution stacking’ where citing every person involved in the generation of the dataset may become unwieldy and difficult to manage. this issue is further magnified when citizen science and other data from multiple projects are combined for further use. ensuring the correct citation formats are maintained manually or by machines itself becomes challenging. to tackle this, it becomes necessary to allow for ‘lightweight attribution mechanisms’ (ball 2011). in this context, it is also worth considering that citizens may be less likely to be motivated by citation in academic journals as against acknowledgment of their contribution that is visible to their local peers and that projects should support attribution in a way that matters to the citizen scientists. some sites, such as ebird, provide the option to hide user names and anonymize them. however, in such cases, attribution for the data is not provided to the contributor for apparent reasons. attribution and user privacy are interlinked, and setting conditions on one of these usually has inverse effects on the other. data policy having clear and robust data policies is a means of ensuring that the data collected through citizen science projects are stored, shared, attributed, and utilized ethically. citizen science project proponents should be mindful of different stakeholders, from contributors to end-users of data and data policies, of the differing rights and responsibilities that each party may possess. while the definition of citizen science is still evolving, it generally encompasses participation from individuals without specific scientific training who participate as volunteers in activities. such activities may cover the breadth of the data life, including study design, data collection and analysis, and dissemination of results (guerrini et al. 2018). this information is then used in ways that may or may not be fully understood by volunteers, and so informed consent must be obtained from volunteers on how the data will be used and what credit they will receive for it. informed consent and refusal are some of the essential components of research ethics that the volunteer willingly gives themselves up for use as a resource (reiheld and gay 2019). informed consent can be ensured by using easy-to-understand documents with minimal text and ensuring participants have agreed to the project terms. it is advisable to place the documents in a conspicuous place on the portal. these documents are a collection of guidelines that constitute the project’s policies that determine how a citizen science project and the users, a website, or a citizen science volunteer may interact or transact. such documents are usually presented as different types of formalized policy documents (bowser et al. 2013). these include: terms of use these form the conditions that a user is expected to know and accept before they begin using the portal. it also encompasses guidelines for acceptable behavior between the user and the portal. terms and conditions may be explicit, requiring the user to accept and consent to the site’s terms before proceeding with registration and usage (clickwrap), or it may be implicit, assuming that the user agrees to the terms simply by continued use of the portal (browsewrap). the terms and conditions set out the conditions of usage of the portal, covering aspects along the lifecycle of the data. it indicates the portal’s stand on data ownership, data access, reuse, and providing attribution to users or recommended citation policies. clarity on aspects of data ownership, including any media uploaded by the user, is imperative. further, the terms need to specify how owners of the site will use the data. it would also need to indicate terms of being contacted for communication regarding outreach or marketing purposes, acceptance of terms and conditions of any third party website linked to the portal (such as youtube or google maps), liability clauses that protect the owner of the portal from any inappropriate content posted on the website by a third party and indemnity clauses against harm caused to any third party from the content of the portal. it would be beneficial to list all the activities that are prohibited on the portal, similar to what is observed in the european citizen science portal36. additional terms of use may allow the portal to block a user in case they violate the terms of use. in some countries, the project may need to clarify if it is merely an intermediary where the adminis36 https://eu-citizen.science/terms/. https://eu-citizen.science/terms/ thomas vattakaven et al. – best practices for data management in citizen science 44 trator does not initiate the transmission by posting the information or select who will be able to view the information or make changes to the information, thereby claiming exemption from liability arising out of the conduct of its users. otherwise, the portals should safeguard themselves from potential legal liability through clear terms of use for all classes of users with clear contracts. legal policies this would cover information on how the site deals with the legal aspects such as its obligations to national or local laws, liabilities of the project, disclaimers, and waivers. it is best practice to include or link to texts containing specific legal or non-legal documentation. privacy policies this covers information on how and what kind of information the project gathers from participants, including information collected during registration, data upload, and how such information is saved, used, and kept confidential. it would also need to disclose the usage of cookies, whether for functionality within the portal such as for login and role-based permissions or through the usage of features provided by third-party sites such as social media networks or advertising providers. privacy concerns in citizen science much attention has been paid to privacy concerns about citizen science data involving medical and genetic information participants. however, data obtained as part of biodiversity inventories or ecological phenomena may also require close perusal for violations of the privacy rights of participants and federal laws that prevent sharing of sensitive information that could jeopardize the safety of endangered species. when collecting biodiversity-related information, privacy breaches can occur at two levels: ● personal information of the observer ● georeferenced data associated with a species record being contributed most projects collect basic personal information of participants, such as names, email ids, and addresses to keep them informed of the progress of the project. through these mediums citizen science projects wittingly or unwittingly end up with personally identifiable information (pii) of participants in their projects. additionally, smartphones equipped with tools that utilize cameras, audio-recorders, and location-capturing applications to capture biodiversity-related information often end up revealing pii (cartwright 2016), that may reveal near real-time information about their locations, patterns of daily or weekend travel, types of phones used, etc. geo-locations of species, commonly required by biodiversity inventories, may reveal sensitive information related to endangered species. information on the location of species could lead to poaching, unethical collection, or disturbance through excessive attention from nature enthusiasts and photographers. this is particularly important when dealing with range-restricted, endangered, frequently traded, or breeding populations of uncommon species. although participants are generally aware of these issues while contributing data (bowser et al. 2013), it is still imperative to get informed consent and brief them on the terms of service employed by the project. a recent study showed that 51% of projects that did not focus exclusively on people data often overlooked the fact that they were still collecting pii (cooper et al. 2019). the personal genome project37 (pgp) has been globally acclaimed for its approach to informed consent that transcends traditional boundaries. the project proponents ensure that all participants pass an examination that tests their knowledge of genomic science and privacy issues. after that, they sign access to their personal and genomic data for the project (angrist 2009). the us and the eu have implemented legal provisions to safeguard the privacy of citizen science contributors. under the us privacy laws, citizen science project managers are mandated to make users aware of their rights and are provided with the privacy act statement. under the children’s online privacy protection rule, collection of personal information of children below the age of 13 is illegal; and the freedom of information and the privacy acts require cleansing all personal information of participants from data collected by projects supported by the federal government before such databases are made public. in the eu, the general data protection regulation (gdpr) seeks the right to be informed, the right of access, the right to rectification, the right to erasure, the right to restrict processing, the right to data portability, the right to object and rights around automated decision making and profiling. under gdpr, project managers are mandated to get fully informed consent from contributors and inform them of how data contributed by them would be used. 37 https://www.personalgenomes.org/. https://www.personalgenomes.org/ thomas vattakaven et al. – best practices for data management in citizen science 45 such existing and upcoming legal provisions have potential implications for the privacy of participants in citizen science portals (ganzevoort et al. 2017). conclusions including the above considerations, every project has to consider its unique situation in terms of biodiversity such as between the explored and unexplored, the documented and undocumented, conservation threats, along with specific challenges to discover, document, and disseminate. each design is significant on its own, reflecting the needs of the socio-ecological system that information technology has to integrate into and co-evolve. although citizen science is rapidly gaining popularity, data generated through it still deals with a perceived “image problem” regarding data quality. while the debate around this issue rages, several studies have indicated that with the appropriate data quality checks in place, citizen science data is no less reliable than data gathered by experts (jordan et al. 2012; ganzevoort et al. 2017). the sole objective of a citizen science project is not necessarily data. through the duration of the project, it builds the capacity of its participants and inculcates the spirit of scientific endeavor and discovery while also sensitizing participants towards species and habitat conservation, creating a sense of stewardship towards nature. another challenge with citizen science is ensuring sustained participation both from citizens and scientists to help validate the data (irwin 2018). from this perspective, imposing too much rigor in data collection and quality can reduce inclusivity and lead to reduced participation. as one of citizen science’s objectives involves broader participation, holding participants to unrealistic scientific standards could mean missing out on opportunities to “fully engage with people in the core objective of discovery” (lukyanenko et al. 2016). multiple competing citizen science initiatives operating within the same region and data sharing between various sources often result in duplication of data contributed in multiple places. this issue will need attention and effort to identify and de-duplicate. global aggregators such as gbif are already investing effort in algorithms to identify potentially related records and cluster them. identifying individual contributors across portals such as through an orcid id can also help in these efforts, although this is still not widely used beyond the academic community yet. to conform to the expectations of its varied user bases, citizen science has to meet the dual objectives of providing high-quality summarized data to the general public as well as spatially, temporally, and taxonomically explicit data to the research community. these have to be achieved while protecting sensitive information and providing privacy protection. achieving these objectives requires significant investment in technology solutions, clear data policies, and transparency. anhalt-depies et al. (2019), give a set of recommendations that may be apt to cater to data quality, privacy, transparency, and trust in citizen science. these include constant communication and consultation with stakeholders, addressing volunteer needs on aspects such as data sharing and user privacy through clear policy documents that evolve through iterative evaluation based on user feedback. among other resources, we refer readers to the 10 principles of citizen science developed by the european citizen science association, which set out the key principles that underlie good practice in citizen science38. in the indian context, it would be ideal for envisaging a directory of citizen science projects and a repository for citizen science projects, which could allow design, host, store, and archive initiatives. this is necessitated by the nature of present-day data infrastructures, which are stretched to provide the full set of features for citizen science practitioners to engage through all the stages of the data lifecycle. many act as platforms for data collection, organization, and aggregation but for various reasons focus less on providing tools to analyze collected data by citizen science practitioners. given the immense potential to contribute to biodiversity monitoring at different scales, a culture of integration covering various tenets of biodiversity information, technical design, and stakeholder networks needs to be promoted (kühl et al. 2020). this is truer for small, focused, and independent citizen science projects for which there is a dire need in a mega-diverse country like india. technology and data infrastructures need to evolve in a direction where modular, decentralized, and federated architectures are imagined and attempted. such architectures will help address the spatial, temporal, and taxon bias and empower communities in sensitive socio-ecological systems to participate in biodiversity conservation effectively. such infrastructure could help transform data infrastructure into knowledge infrastructures, helping enhance 38 https://eu-citizen.science/about/. https://eu-citizen.science/about/ thomas vattakaven et al. – best practices for data management in citizen science 46 the biodiversity knowledge commons and shape policy and practice. acknowledgments we are grateful to suhel quader, pankaj sekhsaria, farida tampal, shannon olsson and prabhakar rajagopal, all members of the organizing committee of citsci india for conceptualizing the idea of this working group and providing guidance and feedback at various stages of developing this toolkit. we thank akshata pradhan for facilitating the functioning of the working group. mridula vijairaghavan, sushmitha viswanathan (wildlife conservation societyindia), and shyama kuriakose (wildlife conservation society-india) vetted the legal components of this document for accuracy and provided additional inputs. we are most grateful to townsend peterson, naveen thayyil, and shannon olsson for reviewing an early version of this document and providing helpful feedback. we acknowledge the role of the citsci india conference participants for sharing their thoughts, and thank them all for their time and for enhancing the quality of this document. competing interests the authors have declared that no competing interests exist. literature cited angrist, m. 2009. eyes wide open: the personal genome project, citizen science and veracity in informed consent. pers. med. 6:691–699. anhalt-depies, c., j. l. stenglein, b. zuckerberg, p. a. townsend, and a. r. rissman. 2019. tradeoffs and tools for data quality, privacy, transparency, and trust in citizen science. biol. conserv. 238:108195. assumpção, t. h., i. popescu, a. jonoski, and d. p. solomatine. 2018. citizen observations contributing to flood modelling: opportunities and challenges. hydrol. earth syst. sci. 22:1473–1489. august, t. a., o. l. pescott, a. joly, and p. bonnet. 2020. ai naturalists might hold the key to unlocking biodiversity data in social media imagery. patterns 1:100116. baker, e., j. p. drury, j. judge, d. b. roy, g. c. smith, and p. a. stephens. 2021. the verification of ecological citizen science data: current approaches and future possibilities. citiz. sci. theory pract. 6:12. ubiquity press. balázs, b., p. mooney, e. nováková, l. bastin, and j. jokar arsanjani. 2021. data quality in citizen science. pp. 139–157 in k. vohland, a. land-zandstra, l. ceccaroni, r. lemmens, j. perelló, m. ponti, r. samson, and k. wagenknecht, eds. the science of citizen science. springer international publishing, cham. ball, a. 2011. how to license research data. digital curation centre, edinburgh. barve, v. 2014. discovering and developing primary biodiversity data from social networking sites: a novel approach. ecol. inform. 24:194–199. boakes, e. h., g. gliozzo, v. seymour, m. harvey, c. smith, d. b. roy, and m. haklay. 2016. patterns of contribution to citizen science biodiversity projects increase understanding of volunteers’ recording behaviour. sci. rep. 6:33051. bonney, r., c. b. cooper, j. dickinson, s. kelling, t. phillips, k. v. rosenberg, and j. shirk. 2009. citizen science: a developing tool for expanding science knowledge and scientific literacy. bioscience 59:977–984. bowser, a., c. cooper, a. de sherbinin, a. wiggins, p. brenton, t.-r. chuang, e. faustman, m. (muki) haklay, and m. meloche. 2020. still in need of norms: the state of the data in citizen science. citiz. sci. theory pract. 5:18. bowser, a., a. wiggins, and r. d. stevenson. 2013. data policies for public participation in scientific research: a primer. dataone public participation in scientific research working group. brenton, p., s. von gavel, e. vogel, and m.-e. lecoq. 2018. technology infrastructure for citizen science. pp. 63–80 in citizen science: innovation in open science, society and policy. ucl press. callaghan, c. t., a. g. b. poore, m. hofmann, c. j. roberts, and h. m. pereira. 2021. large-bodied birds are over-represented in unstructured citizen science data. sci. rep. 11:19073. callaghan, c. t., j. j. l. rowley, w. k. cornwell, a. g. b. poore, and r. e. major. 2019. improving big citizen science data: moving beyond haphazard sampling. plos biol. 17:e3000357. public library of science. campbell, d. l., a. e. thessen, and l. ries. 2020. a novel curation system to facilitate data integration across regional citizen science survey programs. peerj 8:e9219. peerj inc. carroll, s. r., e. herczog, m. hudson, k. russell, and s. stall. 2021. operationalizing the care and fair principles for indigenous data futures. sci. data 8:108. cartwright, j. 2016. technology: smartphone science. nature 531:669–671. chapman, a. 2005a. principles and methods of data cleaning – primary species and species-occurrence data, version 1.0. report for the global biodiversity information facility. chapman, a. 2005b. principles of data quality. global biodiversity information facility. chen, d., and c. p. gomes. 2019. bias reduction via end-to-end shift learning: application to citizen science. proc. aaai conf. artif. intell. 33:493–500. cooper, c., l. larson, k. k. holland, r. gibson, d. farnham, d. hsueh, p. culligan, and w. mcgillis. 2017. contrasting the views and actions of data collectors and data consumers in a volunteer water quality monitoring project: implications for project design and management. citiz. sci. theory pract. 2:8. ubiquity press. thomas vattakaven et al. – best practices for data management in citizen science 47 cooper, c., l. shanley, t. scassa, and e. vayena. 2019. project categories to guide institutional oversight of responsible conduct of scientists leading citizen science in the united states. citiz. sci. theory pract. 4:7. coretrustseal standards and certification board. 2019. coretrustseal trustworthy data repositories requirements 2020–2022. doi: 10.5281/zenodo.3638211. zenodo. courter, j. r., r. j. johnson, c. m. stuyck, b. a. lang, and e. w. kaiser. 2013. weekend bias in citizen science data reporting: implications for phenology studies. int. j. biometeorol. 57:715–720. crosas, m. 2013. joint declaration of data citation principles final. force11. de sherbinin, a., a. bowser, t.-r. chuang, c. cooper, f. danielsen, r. edmunds, p. elias, e. faustman, c. hultquist, r. mondardini, i. popescu, a. shonowo, and k. sivakumar. 2021. the critical importance of citizen science data. front. clim. 3. frontiers. devictor, v., r. j. whittaker, and c. beltrame. 2010. beyond scarcity: citizen science programmes as useful tools for conservation biogeography: citizen science and conservation biogeography. divers. distrib. 16:354–362. dosemagen, s., and a. j. parker. 2019. citizen science across a spectrum: broadening the impact of citizen science and community science. sci. technol. stud. 32. downs, r. r., h. k. ramapriyan, g. peng, and y. wei. 2021. perspectives on citizen science data quality. front. clim. 3. frontiers. escribano, n., d. galicia, and a. h. ariño. 2018. the tragedy of the biodiversity data commons: a data impediment creeping nigher? database j. biol. databases curation 2018:bay033. falk, s., g. foster, r. comont, j. conroy, h. bostock, a. salisbury, d. kilbey, j. bennett, and b. smith. 2019. evaluating the ability of citizen scientists to identify bumblebee bombus species. plos one 14:e0218614. public library of science. faundeen, j., t. e. burley, j. a. carlino, d. l. govoni, h. s. henkel, s. l. holl, v. b. hutchison, e. martín, e. t. montgomery, c. ladino, s. tessler, and l. s. zolly. 2014. the united states geological survey science data lifecycle model. u.s. geological survey, reston, va. ganzevoort, w., r. j. g. van den born, w. halffman, and s. turnhout. 2017. sharing biodiversity data: citizen scientists’ concerns and motivations. biodivers. conserv. 26:2821–2837. geldmann, j., j. heilmann-clausen, t. e. holm, i. levinsky, b. markussen, k. olsen, c. rahbek, and a. p. tøttrup. 2016. what determines spatial bias in citizen science? exploring four recording schemes with different proficiency requirements. divers. distrib. 22:1139–1149. gonsamo, a., and p. d’odorico. 2014. citizen science: best practices to remove observer bias in trend analysis. int. j. biometeorol. 58:2159–2163. groom, q., l. weatherdon, and i. r. geijzendorffer. 2017. is citizen science an open science in the case of biodiversity observations? j. appl. ecol. 54:612–617. guerrini, c. j., m. lewellyn, m. a. majumder, m. trejo, i. canfield, and a. l. mcguire. 2019. donors, authors, and owners: how is genomic citizen science addressing interests in research outputs? bmc med. ethics 20:84. guerrini, c. j., m. a. majumder, m. j. lewellyn, and a. l. mcguire. 2018. policy for citizen science. science 361:134– 136. hagedorn, g., d. mietchen, r. morris, d. agosti, l. penev, w. berendsohn, and d. hobern. 2011. creative commons licenses and the non-commercial condition: implications for the re-use of biodiversity information. zookeys 150:127– 149. pensoft publishers. hampton, s. e., c. a. strasser, j. j. tewksbury, w. k. gram, a. e. budden, a. l. batcheller, c. s. duke, and j. h. porter. 2013. big data and the future of ecology. front. ecol. environ. 11:156–162. iacovides, i., c. jennett, c. cornish-trestrail, and a. l. cox. 2013. do games attract or sustain engagement in citizen science? a study of volunteer motivations. pp. 1101–1106 in chi ’13 extended abstracts on human factors in computing systems. association for computing machinery, paris, france. irwin, a. 2018. no phds needed: how citizen science is transforming research. nature 562:480–482. johnston, a., d. fink, w. m. hochachka, and s. kelling. 2018. estimates of observer expertise improve species distributions from citizen science data. methods ecol. evol. 9:88– 97. jordan, r. c., w. r. brooks, d. v. howe, and j. g. ehrenfeld. 2012. evaluating the performance of volunteers in mapping invasive plants in public conservation lands. environ. manage. 49:425–434. kelling, s., d. fink, f. a. la sorte, a. johnston, n. e. bruns, and w. m. hochachka. 2015. taking a ‘big data’ approach to data quality in a citizen science project. ambio 44:601–611. kelling, s., j. gerbracht, d. fink, c. lagoze, w.-k. wong, j. yu, t. damoulas, and c. gomes. 2012. ebird: a human/computer learning network for biodiversity conservation and research. p. in twenty-fourth iaai conference. kelling, s., a. johnston, a. bonn, d. fink, v. ruiz-gutierrez, r. bonney, m. fernandez, w. m. hochachka, r. julliard, r. kraemer, and r. guralnick. 2019. using semistructured surveys to improve citizen science data for monitoring biodiversity. bioscience 69:170–179. kennett, r., f. danielsen, and k. m. silvius. 2015. citizen science is not enough on its own. nature 521:161–161. kimura, a. h., and a. kinchy. 2016. citizen science: probing the virtues and contexts of participatory research. engag. sci. technol. soc. 2:331–361. kobori, h., j. l. dickinson, i. washitani, r. sakurai, t. amano, n. komatsu, w. kitamura, s. takagawa, k. koyama, t. thomas vattakaven et al. – best practices for data management in citizen science 48 ogawara, and a. j. miller-rushing. 2016. citizen science: a new approach to advance ecology, education, and conservation. ecol. res. 31:1–19. kodali, s. 2017. not open or accountable: the government open data use license is flawed.39 könig, c., p. weigelt, j. schrader, a. taylor, j. kattge, and h. kreft. 2019. biodiversity data integration—the significance of data resolution and domain. plos biol. 17:e3000183. kosmala, m., a. wiggins, a. swanson, and b. simmons. 2016. assessing data quality in citizen science. front. ecol. environ. 14:551–560. kühl, h. s., d. e. bowler, l. bösch, h. bruelheide, j. dauber, david. eichenberg, n. eisenhauer, n. fernández, c. a. guerra, k. henle, i. herbinger, n. j. b. isaac, f. jansen, b. könig-ries, i. kühn, e. b. nilsen, g. pe’er, a. richter, r. schulte, j. settele, n. m. van dam, m. voigt, w. j. wägele, c. wirth, and a. bonn. 2020. effective biodiversity monitoring needs a culture of integration. one earth 3:462–474. lamba, a., p. cassey, r. r. segaran, and l. p. koh. 2019. deep learning for environmental conservation. curr. biol. 29:r977–r982. lemmens, r., v. antoniou, p. hummer, and c. potsiou. 2021. citizen science in the digital world of apps. pp. 461–474 in k. vohland, a. land-zandstra, l. ceccaroni, r. lemmens, j. perelló, m. ponti, r. samson, and k. wagenknecht, eds. the science of citizen science. springer international publishing, cham. lukyanenko, r., j. parsons, and y. f. wiersma. 2016. emerging problems of data quality in citizen science. conserv. biol. 30:447–449. mcclure, e. c., m. sievers, c. j. brown, c. a. buelow, e. m. ditria, m. a. hayes, r. m. pearson, v. j. d. tulloch, r. k. f. unsworth, and r. m. connolly. 2020. artificial intelligence meets citizen science to supercharge ecological monitoring. patterns 1:100109. mcquillan, d. 2014. the countercultural potential of citizen science. mc j. 17. paleco, c., s. g. peter, n. s. seoane, j. kaufmann, and p. argyri. 2021. inclusiveness and diversity in citizen science. p. 529 in k. vohland, a. land-zandstra, l. ceccaroni, r. lemmens, j. perelló, m. ponti, r. samson, and k. wagenknecht, eds. the science of citizen science. springer. pearce‐higgins, j. w., s. r. baillie, k. boughey, n. a. d. bourn, r. p. b. foppen, s. gillings, r. d. gregory, t. hunt, f. jiguet, a. lehikoinen, a. j. musgrove, r. a. robinson, d. b. roy, g. m. siriwardena, k. j. walker, and j. d. wilson. 2018. overcoming the challenges of public data archiving for citizen science biodiversity recording and monitoring schemes. j. appl. ecol. 55:2544–2551. ponti, m., t. hillman, c. kullenberg, and d. kasperowski. 2018. getting it right or being top rank: games in citizen science. citiz. sci. theory pract. 3:1. ratnieks, f. l. w., f. schrell, r. c. sheppard, e. brown, o. e. bristow, and m. garbuzov. 2016. data reliability in citizen 39 https://thewire.in/102905/open-data-licensegovernment/. science: learning curve and the effects of training method, volunteer background and experience on identification accuracy of insects visiting ivy flowers. methods ecol. evol. 7:1226–1235. reiheld, a., and p. l. gay. 2019. coercion, consent, and participation in citizen science. arxiv190713061 phys. schuttler, s. g., r. s. sears, i. orendain, r. khot, d. rubenstein, n. rubenstein, r. r. dunn, e. baird, k. kandros, t. o’brien, and r. kays. 2019. citizen science in schools: students collect valuable mammal data for science, conservation, and community engagement. bioscience 69:69–79. sekhsaria, p., and n. thayyil. 2019. citizen science in ecology in india an initial mapping and analysis. dst centre for policy research, indian institute of technology delhi. steen, v. a., c. s. elphick, and m. w. tingley. 2019. an evaluation of stringent filtering to improve species distribution models from citizen science data. divers. distrib. 25:1857– 1869. sullivan, b. l., j. l. aycrigg, j. h. barry, r. e. bonney, n. bruns, c. b. cooper, t. damoulas, a. a. dhondt, t. dietterich, a. farnsworth, d. fink, j. w. fitzpatrick, t. fredericks, j. gerbracht, c. gomes, w. m. hochachka, m. j. iliff, c. lagoze, f. a. la sorte, m. merrifield, w. morris, t. b. phillips, m. reynolds, a. d. rodewald, k. v. rosenberg, n. m. trautmann, a. wiggins, d. w. winkler, w.-k. wong, c. l. wood, j. yu, and s. kelling. 2014. the ebird enterprise: an integrated approach to development and application of citizen science. biol. conserv. 169:31–40. tiago, p., a. ceia-hasse, t. a. marques, c. capinha, and h. m. pereira. 2017. spatial distribution of citizen science casuistic observations for different taxonomic groups. sci. rep. 7:12832. troudet, j., p. grandcolas, a. blin, r. vignes-lebbe, and f. legendre. 2017. taxonomic bias in biodiversity data and societal preferences. sci. rep. 7. turnhout, e., and s. boonman-berson. 2011. databases, scaling practices, and the globalization of biodiversity. ecol. soc. 16. the resilience alliance. veeckman, c., talboom, s., gijsel, l., devoghel, h., and duerinckx, a. 2019. communication in citizen science. a practical guide to communication and engagement in citizen science. scivil; leuven, belgium. venkatraman, v. 2010. conventions of scientific authorship. sci. aaas, doi: https://www.science.org/careers/2010/04/ conventions-scientific-authorship. vohland, k., a. land-zandstra, l. ceccaroni, r. lemmens, j. perelló, m. ponti, r. samson, and k. wagenknecht (eds). 2021. the science of citizen science. springer international publishing, cham. walker, d., c. mccord, n. stradiotto, m. zhou, and d. singh. 2016. citizen’s guide to open data. wiggins, a., g. newman, r. d. stevenson, and k. crowston. 2011. mechanisms for data quality and validation in citizen science. pp. 14–19 in 2011 ieee seventh international conference on e-science workshops. thomas vattakaven et al. – best practices for data management in citizen science 49 wilkinson, m. d., m. dumontier, ij. j. aalbersberg, g. appleton, m. axton, a. baak, n. blomberg, j.-w. boiten, l. b. da silva santos, p. e. bourne, j. bouwman, a. j. brookes, t. clark, m. crosas, i. dillo, o. dumon, s. edmunds, c. t. evelo, r. finkers, a. gonzalez-beltran, a. j. g. gray, p. groth, c. goble, j. s. grethe, j. heringa, p. a. c. ’t hoen, r. hooft, t. kuhn, r. kok, j. kok, s. j. lusher, m. e. martone, a. mons, a. l. packer, b. persson, p. rocca-serra, m. roos, r. van schaik, s.-a. sansone, e. schultes, t. sengstag, t. slater, g. strawn, m. a. swertz, m. thompson, j. van der lei, e. van mulligen, j. velterop, a. waagmeester, p. wittenburg, k. wolstencroft, j. zhao, and b. mons. 2016. the fair guiding principles for scientific data management and stewardship. sci. data 3:160018. williams, j., c. chapman, d. g. leibovici, g. loïs, a. matheus, a. oggioni, s. schade, l. see, and p. p. l. van genuchten. 2018. maximising the impact and reuse of citizen science data. pp. 321–336 in citizen science: innovation in open science, society and policy. ucl press. biodiversity informatics, 16, 2021, pp. 28-38 28 global land-use and land-cover data: historical, current and future scenarios mariana m. vale1,2,3, mattheus s. lima-ribeiro3,4, and tainá c. rocha3,5* 1ecology department, federal university of rio de janeiro, brazil 2brazilian research network on global climate change (rede clima), são josé dos campos, brazil 3national institute of science and technology on ecology, evolution and biodiversity conservation (inct eecbio), federal university of goiás, goiânia, brazil 4biodiversity department, federal university of jataí (ufj), goiás, brazil 5botanical garden research institute of rio de janeiro, rio de janeiro, brazil abstract. land-use land-cover (lulc) data are important predictors of species occurrence and biodiversity threat. although there are lulc datasets available under current conditions, there is a lack of such data under historical and future climatic conditions. this hinders, for example, projecting niche and distribution models under global change scenarios at different time scenarios. the land use harmonization project (luh2) is a global terrestrial dataset at 0.25o spatial resolution that provides lulc data from 850 to 2300 for 12 lulc state classes. the dataset, however, is compressed in a file format (netcdf) that is incompatible for many analyses and intractable for most researchers, requiring layer extractions and transformations of this format. here we selected and transformed the luh2 in a standard gis format data to make it more user-friendly. we provide lulc for every year from 850 to 2100, and from 2015 on, the lulc dataset is provided under two shared socioeconomic pathways (ssp2-4.5 and ssp5-8.5). we provide two types of files for each year: separate files with continuous values for each of the 12 lulc state classes, and a single categorical file with all state classes combined. to create the categorical layer, we assigned the state with the highest value in a given pixel among the 12 continuous data. luh2 predicts a pronounced decrease in primary forest, particularly noticeable in the amazon, the brazilian atlantic forest, the congo basin and the boreal forests, an equally pronounced increase in secondary forest and non-forest lands, and in croplands in the brazilian atlantic forest and sub-saharan africa. the final dataset provides lulc data for 1251 years that will be of interest for macroecology, ecological niche modeling, global change analysis, and other applications in ecology and conservation. key words: conservation biogeography, ecological niche modelling, macroecology, cmip6, climate change, deforestation introduction land-use and land-cover (lulc) change has been one of the main drivers of environmental change at multiple scales and is currently recognized as an important predictor of anthropogenic impacts and biodiversity threats (maxwell et al. 2016; prestele et al. 2016; gomes et al. 2020, 2021; rosa et al. 2021). mapping land-use land-cover (lulc) changes through time is, therefore, important and desirable to predict these threats and propose effective conservation policies (jetz et al. 2007). lulc is also an important predictor of species’ occurrence and, thus extensively used in ecological and conservation studies (eyringet al. 2016; ruiz-benito et al. 2020; sobral-souza et al. 2021). there are several lulc datasets available at a global scale under current conditions, such as the copernicus (buchhorn et al. 2020), global land survey, the 30 meter global land cover, and the globeland30 (gutman et al. 2013; pengra et al. 2015; brovelli et al. 2015), as well as the near historical period, such as the esa climate change initiative (1992 to 2015), the finer resolution observation, monitoring of global land cover (1984 to 2011) (hollmann et al. 2013; gong et al. 2013) and gcam (20152100) (chen et al. 2020). these datasets are usually available in standard geographic information system (gis) formats (e.g. tif or kmz), routinely used by landscape ecologists, macroecologists, biogeographers, and others (eyringet al. 2016; ruiz-benito et al. 2020; sobral-souza et al. 2021). however, there is an important gap of historical lulc data covering pre-industrial periods (i.e. older than 1700) and, more importantly, projecting lulc changes into the future. current-* corresponding author: taina013@gmail.com. mailto:taina013@gmail.com biodiversity informatics, 16, 2021, pp. 28-38 29 ly, only two initiatives provide future projections: global change analysis model (chen et al. 2020) and land-use harmonization project1 (hurtt et al. 2006; 2011; 2020), and only the last one provides a long historical time-series. the absence of compatible dataset across past, present and future scenarios, for example, hinders the use of lulc predictors in projections of ecological niche and species distribution models throughout the time and hamper global change analyses (escobar et al. 2018). the recent and robust lulc dataset called land-use harmonization project is part of the coupled model intercomparison project (cmip) (hurtt et al. 2006, 2011, 2020), which coordinates modeling experiments worldwide used by the intergovernmental panel on climate change (ipcc) (eyring et al. 2016). the data is an input to earth system models (esms) to estimate the combined effects of human activities on the carbon-climate system. currently, cmip datasets are available in netcdf format, a quite complex file format for most researchers. a few studies used or analyzed the cmip lulc (xia & niu 2020 and references therein), as opposed to cmip’s climate data already simplified on standard gis formats available in worldclim2 (fick and hijmans 2017) and ecoclimate3 (lima-ribeiro et al. 2015). the land-use harmonization project (luh2) provides the most complete data in term of time-series and scenarios of climate change. the data covers a period from 850 to 2300 at 0.25o spatial resolution (ca. 30 km). the first generation of models (luh1, hurtt et al. 2006; 2011) made future land-use land-cover projections under cmip5’s representative concentration pathways greenhouse gas scenarios (rcps, see vuuren et al. 2011), and the current generation of models (luh2, hurtt et al. 2020) makes projection under cmip6’s shared socioeconomic pathways greenhouse gas scenarios (ssp, see popp et al. 2017). both provide data on 12 land-use land-cover state classes, including different categories of natural vegetation, agriculture and urban areas. in order to make the global land-use harmonization data more accessible and readily usable, here we filtered, combined and transformed it in standard gis formats, making the dataset accessible for users with standard gis skills. besides providing the land-use harmonization data in regular gis format 1 https://luh.umd.edu/data.shtml. 2 https://www.worldclim.org/. 3 https://www.ecoclimate.org/. at yearly temporal resolution covering 1251 years of past, present and future (from 850 to 2100), we also derived new data based on the existing dataset. methods we downloaded the 12 land-use land-cover state layers (state.nc) provided in network common data form (netcdf) from the land-use harmonization project (luh2): forested primary land (primf), non-forested primary land (primn), potentially forested secondary land (secdf), potentially non-forested secondary land (secdn), managed pasture (pastr), rangeland (range), urban land (urban), c3 annual crops (c3ann), c3 perennial crops (c3per), c4 annual crops (c4ann), c4 perennial crops (c4per), c3 nitrogen-fixing crops (c3nfx). the “forested” and “non forested” land-use states are defined on the basis of the aboveground standing stock of natural cover; where “primary” are lands previously undisturbed by human activities, and “secondary” are lands previously disturbed by human activities and currently recovered or in process of recovering of their native aspects (see hurtt et al. 2006; 2011; 2020 for more details). they were computed using an accounting-based method that tracks the fractional state of the land surface in each grid cell as a function of the land surface at the previous time step through historical data. because it deals with a large and undetermined system, the approach was to solve the system for every grid cell at each time step, constraining with several inputs including land-use maps, crop type and rotation rates, shifting cultivation rates, agriculture management, wood harvest, forest transitions and potential biomass and biomass recovery rates (see fig. s1 in the supplementary material for details). to manipulate the netcdf files, we used the ncdf4 and rgdal packages in r environment (r core team 2020, pierce 2019; hijmans et al. 2020; bivand et al. 2021). we also used the panoply software version 4.84 for quick visualization of the original data (states.nc) (schmunk, 2017). we created two sets of files for each year, the continuous “state-files” and the categorical “lulcfiles” (fig.1, fig.2 and fig. s2 of supplemental material). the statefiles are the same data provided in the original luh2 dataset (states.nc), transformed into tag image file format (tiff) and standardized for ranging from 0 to 1. we built the new lulcfiles, also in tiff format, assigning the highest val4 https://www.giss.nasa.gov/tools/panoply/. https://luh.umd.edu/data.shtml https://www.worldclim.org/ https://www.ecoclimate.org/ http://hdl.handle.net/1808/31846 http://hdl.handle.net/1808/31846 http://hdl.handle.net/1808/31846 https://www.giss.nasa.gov/tools/panoply/ biodiversity informatics, 16, 2021, pp. 28-38 30 ue among the 12 available states to each pixel. for instance, if the highest value in a given pixel is the forest state value, it was categorically set as a forest pixel. thus, the lulc-files present categories ranging from 1 to 12, which represents each one of the 12 existing states in the dataset (table s1 in supplementary material). we generated states-files and lulcfiles for every year from 850 to 2100 for two greenhouse gas scenarios: an intermediate (ssp2-4.5) and a pessimistic (ssp5-8.5) (see fig. s2 in supplementary material for the workflow to create state files and lulc-files). the ssp2-4.5 scenario, a.k.a “middle of the road”, represents a 4.5 w/m2 radiative forcing by 2100, where historical development patterns continue throughout the 21st century, susceptibility to societal and environmental changes remains, and greenhouse gas emissions are at intermediate levels. the ssp5-8.5, a.k.a. “fossil-fueled development”, on the other hand, represents the upper limit of the ssp scenarios spectrum economic, where social development is coupled with the exploitation of abundant fossil fuel resources, an energy-intensive lifestyles, and high levels of greenhouse gas emissions (popp et al. 2017; meinshausen et al. 2020; gatti et al. 2021). we performed an accuracy assessment of our classification for the lulc-files following olofsson et al.’s (2014) good practices, for the all continents together and for newton and dale’s (2001) zoogeographic regions separately. we compared our classified lulc-file for the year 2000 with that of the global land cover share (glc-share) data, used as the ground truth reference data in the accuracy assessment. the glc-share was built from a combination of “best available” high resolution national, regional and/or sub-national land cover databases (latham et al. 2014), and has a finer spatial resolution (1 km) than the luh2 (30 km). glc-share has 11 classes that are very similar with those from the luh2 database: artificial surfaces (01), cropland (02), grassland (03), tree covered areas (04), shrubs covered areas (05), herbaceous vegetation, aquatic or regularly flooded (06), mangroves (07), sparse vegetation (08), bare soil (09), snow and glaciers (10), and water bodies (11). to make the two datasets comparable, we reclassified luh2 and glc-share to the following classes: forest, crops, open areas and urban (fig. 3, table s1 in supplementary material). we also masked-out ice and water areas from glc-share, as they do not have an equivalent in the luh2 datafigure 1. example of state-files data. continuous forested primary land state for 2020 (top) and 2100 (bottom) under ssp5-8.5 greenhouse gas scenario, as originally provided by the land-use harmonization (luh2) project. state values range from 0 to 1, roughly representing the likelihood a pixel is occupied by the land-use land-cover class depicted in the map. all other state-files have the same structure. http://hdl.handle.net/1808/31846 http://hdl.handle.net/1808/31846 http://hdl.handle.net/1808/31846 http://hdl.handle.net/1808/31846 http://hdl.handle.net/1808/31846 biodiversity informatics, 16, 2021, pp. 28-38 31 set. thus, greenland was removed from analysis and is absent in the lulc-files. we performed the accuracy assessment in qgis 3.20 through a confusion matrix error, quantifying the commission and omission errors for each class, and then computing three primary metrics: overall accuracy (oa), producer accuracy (pa) and user accuracy (ua). we also provide other supplemental metrics, such as kappa, allocation disagreement and quantity disagreement using map accuracy tools (salk et al. 2018) so that users can choose the best metric given their purpose (see supplementary material, accuracies). all codes to perform the analysis are available on the github platform (https://github.com/tai-rocha/ luh2_data). the entire resulting dataset is freely available for download at the ecoclimate repository, an open database of processed environmental data in a suitable resolution and user-friendly format (lima-ribeiro et al. 2015). results we generated 17.394 files, 16.056 of which are the luh2 original (continuous data) states files transformed into tiff (fig. 1), and the other 1.338 are new (categorical data) lulc-files created by combining the 12 states files (). the lulc-files had good results for most zoogeographic regions and land-use land-cover classes, but not for all (fig. 3, table 1). the overall accuracy (oa) was over than 70% for global scale and for most regions, except figure 2. example of lulc-files data. categorical lulc for 2020 (top) and 2100 (bottom) under ssp5-8.5 greenhouse gas scenarios, as a result of the combination of the 12 luh2 original state classes (state-files) into a single map. crops forest open areas urban oa pa ua pa ua pa ua pa ua global 71.7 79.7 47.3 70.5 66.8 71.2 82.7 55.5 13.2 afrotropical 70,9 72.2 15.1 72.4 42.2 70.6 93.9 50.0 2.0 australasian 82.0 80.5 54.9 91.2 47.0 80.0 98.0 83.3 20.0 indomalayan 77.7 90.0 77.0 83.2 83.0 58.2 71.3 35.7 9.8 neartic 71.7 83.1 59.4 61.1 84.3 81.2 66.9 80.9 27.9 neotropical 65.4 89.5 14.8 87.3 66.9 47.7 88.3 39.2 40.7 afrotropical 71.4 71.3 53.1 67.7 64.7 73.5 81.2 30.3 4.0 table 1. classification accuracy (expressed as percentages) for lulc classes at global scale and biogeographical regions. oa: overall accuracy; pa: producer accuracy; ua: user accuracy. see the confusion matrix and accuracy metrics in accuracies.xlsx supplemental file. http://hdl.handle.net/1808/31846 https://github.com/tai-rocha/luh2_data https://github.com/tai-rocha/luh2_data http://hdl.handle.net/1808/31846 biodiversity informatics, 16, 2021, pp. 28-38 32 for the neotropics, with 65 % overall accuracy. australasia had the highest oa, with 82% accuracy (see table 1 and supplemental material s3 for all metrics of accuracy). the producer accuracy (pa) and user accuracy (ua) for land-use land-cover classes in zoogeographic regions showed some interesting patterns (table 1 and supplemental tables s3). for crops, there was good pa (71% to 90%) and poor or moderated ua (14% to 59%), except for the indomalayan region (ua = 77%). forest had moderate to good pa (61% to 91%) and poor to good ua (42% to 84%). open area had poor to good pa (47% to 81%), moderate to good ua (71% to 93%). urban areas had poor to good pa (30% to 83%) and very poor or poor ua (2% to 40%). the land-use harmonized project shows important changes in lulc through time (fig. 1 and 2), although with no noticeable difference between greenhouse gas scenarios within the same year (fig. 4). it predicts a pronounced decrease in primary forest, and an equally pronounced increase in secondary forest and non-forest lands (fig. 4). the decrease in primary forest is particularly noticeable in the amazon, the brazilian atlantic forest, the congo basin and the boreal forests (fig. 1), coupled with an increase in secondary forest in these regions (fig. 2). a predicted increase in c4 annual, c3 nitrogen-fixing and c3 perennial crops is especially pronounced in the brazilian atlantic forest and sub-saharan africa (fig. 2). these crops will apparently replace managed pastures in africa’s great lakes region. finally, there is also a specially pronounced predicted decrease in non-forested primary land (fig. 4), especially in northern africa and in the horn of africa (fig. 2). discussion this data paper is an important contribution, making the land-use harmonization project data more accessible. here, we provide a global scale lulc dataset with yearly time resolution over a period of 1251 years (from 850 to 2100), and considerable spatial resolution (0.25o long/lat). we contributed not only by transforming the data into standard gis file format, but also by providing new categorical data on land-use land-cover through a long time period. this lulc database provides support for several research fields in ecology and biodiversity, by disseminating open datasets/open-source tools for a quality, transparent and inclusive science. our open, ready-to-use and user-friendly database will enable a more robust integration between climate and land-use change figure 3. data used in the accuracy assessment of lulc-files. the accuracy of the classification of the lulc-file (bottom) assessed using the glc-share as reference data (top). to make the two datasets comparable, both were reclassified to four land-use land-cover states for the year 2000 (see table 1 for reclassification scheme). http://hdl.handle.net/1808/31846 http://hdl.handle.net/1808/31846 biodiversity informatics, 16, 2021, pp. 28-38 33 within biodiversity science (titeux et al. 2017; albert et al. 2020; hanna et al. 2020). given that overall accuracy is still a widely used metric (e.g., curtis et al. 2018; gong et al. 2019; kafy et al. 2021; liu et al. 2021), our lulc-files provide good quality data (70% to 82% oa), especially for large and coarse scale studies. besides, we follow the best practices suggested by olofsson et al. (2014) for validation, considering a reference map with higher quality than the map classification. validation requires the matching of both maps in terms of classes. thus, we carefully choose a reference map (glcshare) that shared similarities with luh2 in terms of number of classes, which we believe reduced the biases in the reclassification process. in any case, we suggest that users consult table 1 and supplemental file accuracies.xlsx for classes’ accuracy at different zoogeographic regions when performing regional analysis. the most pronounced changes predicted by the land-use harmonized project between years 2020 and 2100 are the decrease in primary forest and the increase in secondary forest and non-forested lands (fig.4, ssp2-4.5 and ssp5-8.5). it is important to note that “primary” refers to intact land, undisturbed by human activities since 850, while “secondary” refers to land undergoing a transition or recovering from previous human activities (hurrt et al. 2006; 2011; 2020). a major concern regarding the reduction of primary forest is, obviously, habitat loss and associated biodiversity decline, specially of rarer species (chase et al. 2020; horta and santos 2020; lima et al. 2020), in addition to increased greenhouse gas emissions (mackey et al. 2020) and likelihood of pandemics associated with viral spillover from wildlife to humans (dobson et al. 2020). predicted forest loss is noticeable in the amazon, brazilian atlantic forest, congo basin and boreal forests, especially under the ssp5-8.5 (fig .1 and fig. 2), which is in agreement with recent findings. svensson et al. (2019) found, for example, a decrease from 75% to 38% in boreal forests between years 1973 and 2013, and shapiro et al. (2021) showed that over 24 million hectares of forest were degraded in the congo basin between years 2000 and 2016. similar or worse scenarios are happening in the amazon and atlantic forest (junior et al. 2021; rosa et al. 2021). this is happening particularly inside brazil, where recent governmental actions have promoted deforestation and forest fires (escobar 2019; 2020; amigo 2020; silva et al. 2021; frança et al. 2021; qin et al. 2021; vale et al. 2021), with record deforestation rates in the amazon (junior et al. 2021). although not captured quantitatively at the global analysis (fig.4), another relevant regional level prediction is the increase in c4 annual, c3 nitrogen-fixing, and c3 perennial crops in the brazilian atlantic forest and sub-saharan africa (fig. 2),. other studies have similar predictions (zabel et al. 2019), and the trend is already observed in the atlantic forest (rosa et al. 2021). the data provided here provides support for several analysis in ecology and biodiversity. the continuous data in the state-files may be particularly useful as predictors in ecological niche modeling (peterson et al. 2011) or can be combined to species distribution models to reconstruct changes in species distributions (sofaer et al. 2019; cazaca et al. 2020). the forested primary land state, for example, can be used to model the distribution of forest-defigure 4. land-use land cover comparison among years and scenarios. data for the lulc-files for year 2020 and 2100 for the optimistic (ssp2-4.5, top) and pessimistic (ssp5-8.5, bottom) greenhouse gas scenarios, arranged in decreasing order of class area in 2020. http://hdl.handle.net/1808/31846 http://hdl.handle.net/1808/31846 biodiversity informatics, 16, 2021, pp. 28-38 34 pendent species, as in birds from the atlantic forest biodiversity hotspot (vale et al. 2018). this data has the advantage of being represented in continuous values, as opposed to most discrete land cover data (e.g. all datasets cited in this paper), overcoming the shortcoming of using categorical data as layers in ecological niche modeling (peterson 2001). more importantly, it allows for the use of land cover data in projections of species distribution under future climate change scenarios. additionally, the categorical data in the lulc-files can be useful in ecosystem services mapping, especially when working with the widely-used invest modeling tool5, which is highly dependent on land-use land-cover data (sharp et al. 2020). the lulc-files can also be used in studies of global change impacts from other perspectives (mantyka-pringle et al. 2015; titeux et al. 2017; newbold 2018; clerici et al. 2019; hong et al. 2019; jetz et al 2007; powers and jetz 2019). least, but not least, the data can help decision-makers in the construction of evidence based mitigation and conservation policies (martinez-fernández et al. 2015; dong et al. 2018). we hope that the dataset provided here, which is freely available for download at the ecoclimate repository, can foster the use of land-use land-cover data in many and different fields of study. author contributions mv and tr designed the study, tr carried out the analyses, mv, tr and ml-r wrote the paper. acknowledgments this initiative was possible due to the high-quality data that are maintained and made publicly available by land-use harmonization. we are grateful to ritvik sahajpa for helping in the early stages of data acquisition and transformation. this study was developed in the context of the national institute for science and technology in ecology, evolution and conservation of biodiversity (inct eecbio, cnpq grant n 465610|2014-5, fapeg 201810267000023) and the brazilian network on global climate change research (rede clima). mmv and ml-r received support from the national council for scientific and technological development (cnpq grants no. 304309/2018-4 and 301514/2019-4, respectively) and tr from capes (coordination for the improvement of higher education personnel grant no. 88887.373031/2019-00) and open life science program (ols-2). 5 https://naturalcapitalproject.stanford.edu/software/invest. data accessibility all files will be made freely available online in ecoclimate database6. conflict of interest the authors have declared that no competing interests exist. literature cited albert, c. h., m. hervé, m. fader, a. bondeau, a., leriche, a. c. monnet, and w. cramer. 2020. what ecologists should know before using land use/cover change projections for biodiversity and ecosystem service assessments. reg. environ. change 20:1-12. doi:10.1007/s10113-020-01675-w amigo, i. 2020. when will the amazon hit a tipping point? nature 578(7796): 505-508. doi:10.1038/d41586020-00508-4 ay, j. s., j. guillemot, j. martin‐stpaul, n. doyen, l. and p. leadley. 2017. the economics of land use reveals a selection bias in tree species distribution models. glob. ecol. biogeogr. 26(1): 65-77. doi: 10.1111/ geb.12514 bivand, r., t. keitt, b. rowlingson, e. pebesma, m. sumner, r. hijmans, e. rouault, and m.r bivand. 2015. rgdal: bindings for the geospatial data abstraction library. r package version 1.5-12. brovelli, m. a., m. e. molinari, e. hussein, j. chen, and r. li. 2015. the first comprehensive accuracy assessment of global and 30 at a national level: methodology and results”. remote sens. 7:4191-4212. doi:10.3390/rs70404191 buchhorn, m., m. lesiv, n. e. tsendbazar, m. herold, l. bertels, and b. smets. 2020. copernicus global land cover layers—collection 2. remote sens. 12(6): 1044. doi: 10.3390/rs12061044 casazza, g., f. malfatti, m. brunetti, v. simonetti, and a. s. mathews. a. s. 2021. interactions between land use, pathogens, and climate change in the monte pisano, italy 1850–2000. landsc. ecol. 36(2): 601-616. doi:10.1007/s10980-020-01152-z chase, j. m., s.a. blowes, t. m knight, k. gerstner, and f. may, f. 2020. ecosystem decay exacerbates biodiversity loss with habitat loss. nature 584(7820): 238243. doi: doi.org/10.1038/s41586-020-2531-2 chen, m., c. r. vernon, n. t. graham, m. hejazi, m. huang, y. cheng, and k. calvin. 2020. global land use for 2015–2100 at 0.05° resolution under diverse socioeconomic and climate scenarios. sci. data 7:111. doi:10.1038/s41597-020-00669-x. 6 https://www.ecoclimate.org/. https://naturalcapitalproject.stanford.edu/software/invest http://doi.org/10.1007/s10113-020-01675-w http://doi.org/10.1038/d41586-020-00508-4 http://doi.org/10.1038/d41586-020-00508-4 https://doi.org/10.1111/geb.12514 https://doi.org/10.1111/geb.12514 http://doi.org/10.3390/rs70404191 https://doi.org/10.3390/rs12061044 https://doi.org/10.1007/s10980-020-01152-z https://doi.org/10.1038/s41586-020-2531-2 http://doi.org/10.1038/s41597-020-00669-x. https://www.ecoclimate.org/ biodiversity informatics, 16, 2021, pp. 28-38 35 clerici, n., f. cote-navarro, f. j. escobedo, k. rubiano, and j. c. villegas. 2019. spatio-temporal and cumulative effects of land use-land cover and climate change on two ecosystem services in the colombian andes. sci. total enviro. 685:1181-1192. doi: 10.1016/j.scitotenv.2019.06.275 curtis, p. g., c. m. slay, n. l. harris, a. tyukavina, and m. c. hansen. 2018. classifying drivers of global forest loss. science, 361(6407), 1108-1111. doi:10.1126/ science.aau3445 dobson, a.p., s.l. pimm, l. hannah, l. kaufman, j. a. ahumada, a. w. ando, a. bernstein, j. busch, p. daszak, j. engelmann, m. f. kinnaird, b. v. li, t. loch-temzelides,t. lovejoy, k.nowak, p. r. roehrdanz, and m. m. vale. 2020. ecology and economics for pandemic prevention. science, 369(6502): 379381. dong, n., l. you, w. cai, g. li, and h. lin. 2018. land use projections in china under global socioeconomic and emission scenarios: utilizing a scenario-based land-use change assessment framework. glob. environ. change 50, 164-177. doi:10.1016/j.gloenvcha.2018.04.001. escobar, l. e., h. qiao, j. cabello, and a. t. peterson. 2018. ecological niche modeling reexamined: a case study with the darwin’s fox. ecol. evol. 8:4757-4770. doi:10.1002/ece3.4014 escobar, h. 2019. amazon fires clearly linked to deforestation, scientists say. science, 853-853. doi:10.1126/ science.365.6456.853 escobar, h. 2020. deforestation in the brazilian amazon is still rising sharply. science, 613. doi: 10.1126/science.369.6504.613 eyring, v., s. bony, g. a. meehl, c. a. senior, b. stevens, r. j. stouffer, and k. e. taylor. 2016. overview of the coupled model intercomparison project phase 6 (cmip6) experimental design and organization. geosci. model dev. 9:1937-1958. doi:10.5194/gmd-91937-2016. frança, f., r. solar, a. c. lees, l. p. martins, e. berenguer, and j. barlow. 2021. reassessing the role of cattle and pasture in brazil’s deforestation: a response to “fire, deforestation, and livestock: when the smoke clears”. land use policy 108:105195. doi:10.1016/j. landusepol.2020.105195 gatti l.v., l. s. basso , j. b. miller , m. gloor , l.g. domingues , h. l. g. cassol , g. tejada , l. e. o. c. aragão , c. nobre, w. peters, l. marani e. arai, a. h. sanches , s. m. corrêa , l. anderson, c.v. randow, c. s. c. correia , s. p. crispim, and r. a. l. neves. 2021. amazonia as a carbon source linked to deforestation and climate change. nature 595: 388-393. doi:10.1038/s41467-019-10775-z gong, p., j. wang, l. yu, y. zhao, y. zhao, l. liang, z. niu, x. huang, h. fu, s. liu, c. li, x. li, w. fu, c. liu, y. xu, x. wang, q. cheng, l. hu, w. yao, h. zhang, p. zhu, z. zhao, h. zhang, y. zheng, l. ji, y. zhang, h. chen, a. yan, j. guo, l. yu, l. wang, x. liu, t. shi, m. zhu, y. chen, g. yang, p. tang, b. xu, c. giri, n. clinton, z. zhu, j. chen, and j. chen. 2013. finer resolution observation and monitoring of global land cover: first mapping results with landsat tm and etm+ data. int. j. remote sens. 34:26072654. doi:10.1080/01431161.2012.748992. gutman, g., c. huang, g. chander, p. noojipady, and j. g. masek. 2013. assessment of the nasa-usgs global land survey (gls) datasets. remote sens. environ. 134:249-265. doi:10.1016/j.rse.2013.02.026. hanna, d. e. l., c. raudsepp-hearne, and e. m. bennett. 2020. effects of land use, cover, and protection on stream and riparian ecosystem services and biodiversity. conserv. biol. 34:244-255. doi:10.1111/ cobi.13348. hong, j., g. s. lee, , j. j. park, , h. h. mo, and k. cho. 2019. risk map for the range expansion of thrips palmi in korea under climate change: combining species distribution models with land-use change. j. asia pac. entomol. 22(3): 666-674. doi:10.1016/j.aspen.2019.04.013 hortal, j., and a. m. santos. 2020. rethinking extinctions that arise from habitat loss. nature 584:194-195. doi:10.1038/d41586-020-02210-x fick, s. e., and r. j. hijmans. 2017. worldclim 2: new 1‐ km spatial resolution climate surfaces for global land areas. int. j. climatol. 37: 4302-4315. doi:10.1002/ joc.5086. hijmans, r. j. 2020. raster: geographic data analysis and modeling. r package version 3.3.13. hollmann, r., c. j. merchant, r. saunders, c. downy, m. buchwitz, a. cazenave, e. chuvieco, p. defourny, g. de leeuw, r. forsberg, t. holzer-popp, f. paul, s. sandven, s. sathyendranath, m. van roozendael, and w. wagner. 2013. the esa climate change initiative: satellite data records for essential climate variables. bull. am. meteorol. soc. 94:1541-1552. doi:10.1175/ bams-d-11-00254.1 hurtt, g. c., l. p. chini, s. frolking, r. a. betts, j. feddema, g. fischer, j. p. fisk, k. hibbard, r. a. houghton, a. janetos, c. d. jones, g. kindermann, t. kinoshita, kees klein goldewijk, k. riahi, e. shevliakova, s. smith, e. stehfest, a. thomson, p. thornton, d. p. van vuuren, and y. p. wang. 2011. harmonization https://doi.org/10.1016/j.scitotenv.2019.06.275 https://doi.org/10.1016/j.scitotenv.2019.06.275 http://doi.org/10.1126/science.aau3445 http://doi.org/10.1126/science.aau3445 http://doi.org/10.1016/j.gloenvcha.2018.04.001. http://doi.org/10.1016/j.gloenvcha.2018.04.001. http://doi.org/10.1002/ece3.4014 http://doi.org/10.1126/science.365.6456.853 http://doi.org/10.1126/science.365.6456.853 http://doi.org/10.1126/science.369.6504.613 http://doi.org/10.1126/science.369.6504.613 http://doi.org/10.1080/01431161.2012.748992. http://doi.org/doi:10.1016/j.rse.2013.02.026 https://doi.org/10.1016/j.aspen.2019.04.013 https://doi.org/10.1016/j.aspen.2019.04.013 https://doi.org/10.1038/d41586-020-02210-x http://doi.org/doi:10.1002/joc.5086 http://doi.org/doi:10.1002/joc.5086 http://doi.org/10.1175/bams-d-11-00254.1 http://doi.org/10.1175/bams-d-11-00254.1 biodiversity informatics, 16, 2021, pp. 28-38 36 of land-use scenarios for the period 1500-2100: 600 years of global gridded annual land-use transitions, wood harvest, and resulting secondary lands. clim. change 109:117-161. doi: 10.1007/s10584-0110153-2.7 hurtt, g. c., l. chini, r. sahajpal, s. frolking, b. l. bodirsky, k. calvin, j. c. doelman, j. fisk, s. fujimori, k. k. goldewijk, t. hasegawa, p. havlik, a. heinimann, f. humpenöder, j. jungclaus, j. kaplan, j. kennedy, t. krisztin, d. lawrence, p. lawrence, l. ma, o. mertz, j. pongratz, a. popp, b. poulter, k. riahi, e. shevliakova, e. stehfest, p. thornton, f. n. tubiello, d. p. van vuuren, and x. zhang. 2020. harmonization of global land use change and management for the period 850-2100 (luh2) for cmip6. geosci. model dev. 13:5425-5464. doi: 10.5194/ gmd-13-5425-2020 hurtt, g. c., s. frolking, m. g. fearon, b. moore, e. shevliakova, s. malyshev, s. w. pacala, and r. a. houghton. 2006. the underpinnings of land-use history: three centuries of global gridded land-use transitions, wood-harvest activity, and resulting secondary lands. glob. change biol. 12:1208-1229. doi:10.1111/j.1365-2486.2006.01150.x jetz, w., d. s. wilcove, and a. p. dobson. 2007. projected impacts of climate and land-use change on the global diversity of birds. plos biol. 5:1211-1219. doi:10.1371/journal.pbio.0050157 junior, c. h. s., a. c. pessôa, n. s. carvalho, j. b. reis, l. o. anderson, and l. e. aragão. 2021. the brazilian amazon deforestation rate in 2020 is the greatest of the decade. nat. ecol. evol. 5(2): 144-145. doi:10.1038/s41559-020-01368-x kafy, a.a., a. al rakib, k. s. akter, z. a. rahaman, a. a. faisal, s. mallik, n. r. nasher, m. i. hossain, and m. y. ali. 2021. monitoring the effects of vegetation cover losses on land surface temperature dynamics using geospatial approach in rajshahi city, bangladesh. environ. challenges 100187. doi:10.1016/j. envc.2021.100187 latham, j., r. cumani, i. rosati, and m. bloise. 2014. global land cover share (glc-share). fao: rome, italy. version 1.0-20147. lima-ribeiro, m. s. 2015. ecoclimate: a database of climate data from multiple models for past, present, and future for macroecologists and biogeographers. biodiv. inf. 10:1-21. doi:10.17161/bi.v10i0.4955. lima, r. a., a. a. oliveira, g. r. pitta, a. l. de gasper, a. c. vibrans, j. chave, h. ter steege, and prado, p.i., 2020. the erosion of biodiversity and biomass in 7 http://www.fao.org/uploads/media/glc-share-doc.pdf. the atlantic forest biodiversity hotspot. nat. comm. 11(1): 1-16. doi:10.1038/s41467-020-20217-w liu, l., x. zhang, y. gao, x. chen, x. shuai, and j. mi. 2021. finer-resolution mapping of global land cover: recent developments, consistency analysis, and prospects. j. remote sens. 2021(5289697):1-38. doi:10.34133/2021/5289697 mackey, b., c. f. kormos, h. keith, w. r moomaw, r. a. houghton, r. a. mittermeier, d. hole, and s. hugh. 2020. understanding the importance of primary tropical forest protection as a mitigation strategy. mitig. adapt. strateg. glob. change 25: 763-787. doi:10.1007/s11027-019-09891-4 maxwell, s. l., r. a. fuller, t. m. brooks, and j. e.m. watson. 2016. biodiversity: the ravages of guns, nets and bulldozers. nature 536:143-145. doi: 10.1038/536143a mantyka-pringle, c. s., p. visconti, m. di marco, t. g. martin, c. rondinini, and j. r. rhodes. 2015. climate change modifies risk of global biodiversity loss due to land-cover change. biol. conserv. 187: 103111. doi:10.1016/j.biocon.2015.04.016 martinez-fernandez, j., p. ruiz-benito, and m. a. zavala. 2015. recent land cover changes in spain across biogeographical regions and protection levels: implications for conservation policies. land use policy 44: 62-75. doi:doi.org/10.1016/j.landusepol.2014.11.021 meinshausen, m., z. nicholls, j. lewis, m. j. gidden, e. vogel, m. freund, u. beyerle, c. gessner, a. nauels, n. bauer, j. g. canadell, j. s. daniel, a. john, p. krummel, g. luderer, n. meinshausen, s. a montzka, p. rayner, s. reimann, s. j smith, m. van den berg, g. j. m. velders, m. vollmer, and h. j. wang. 2020. the shared socio-economic pathway (ssp) greenhouse gas concentrations and their extensions to 2500. geosci. model dev.13(8), 3571-3605. doi:10.5194/gmd-13-3571-202 newton i, and l. dale. 2001. a comparative analysis of the avifaunas of different zoogeographical regions. j. zool. 254:207-218. doi:10.1017/s0952836901000723 newbold, t. 2018. future effects of climate and land-use change on terrestrial vertebrate community diversity under different scenarios. proc. r. soc. b, 285(1881): 20180792. doi:10.1098/rspb.2018.0792 qin y., x. xiao , j. p. wigneron, p. ciais , m. brandt , l. fan , x. li , s. crowell , x. wu , r. doughty , y. zhang , f. liu , s. sitch , and b. moore. 2021. carbon loss from forest degradation exceeds that from deforestation in the brazilian amazon. nat. clim. change, 11(5): 442-448. doi:10.1038/s41558-021-01026-5 http://doi.org/10.1007/s10584-011-0153-2.7 http://doi.org/10.1007/s10584-011-0153-2.7 http://doi.org/10.5194/gmd-13-5425-2020 http://doi.org/10.5194/gmd-13-5425-2020 http://doi.org/10.1111/j.1365-2486.2006.01150.x http://doi.org/10.1371/journal.pbio.0050157 http://doi.org/10.1038/s41559-020-01368-x https://doi.org/10.1016/j.envc.2021.100187 https://doi.org/10.1016/j.envc.2021.100187 http://doi.org/10.17161/bi.v10i0.4955. http://www.fao.org/uploads/media/glc-share-doc.pdf http://doi.org/10.34133/2021/5289697 https://doi.org/10.1007/s11027-019-09891-4 https://doi.org/10.1038/536143a https://doi.org/10.1016/j.biocon.2015.04.016 https://doi.org/10.1016/j.landusepol.2014.11.021 http://doi.org/10.5194/gmd-13-3571-202 https://doi.org/10.1098/rspb.2018.0792 https://doi.org/10.1038/s41558-021-01026-5 biodiversity informatics, 16, 2021, pp. 28-38 37 pengra, b., j. long, d. dahal, s. v. stehman, and t. r. loveland. 2015. a global reference database from very high resolution commercial satellite data and methodology for application to landsat derived 30m continuous field tree cover data. remote sens. environ. 165:234-248. doi:10.1016/j.rse.2015.01.018. peterson, a. t. 2001. predicting species’ geographic distributions based on ecological niche modeling. condor 103:599-605. doi:10.1093/condor/103.3.599 peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. araújo. b. 2011. ecological niches and geographic distributions. princeton university press. princeton. pierce, d., and m. d. pierce. 2019. ncdf4: interface to unidata netcdf. r package version 1.17. available at: https://www.vps.fmvz.usp.br/cran/web/packages/ncdf4/ncdf4.pdf popp, a., k. calvin, s. fujimori, p. havlik, f. humpenöder, e. stehfest, b. l. bodirskyah, j. p. dietrich, j. c. doelmann, m. gusti, t. hasegaw, p. kyl, m. obersteiner, a. tabeau, k. takahashi, h. valin, s. waldhoff, i. weindl, m. wise, e. kriegler , h. lotze-campen , o. fricko , k. riahi , d. van vuuren. 2017. land-use futures in the shared socio-economic pathways. glob. environ. change, 42:331-345. doi:10.1016/j.gloenvcha.2016.10.002. powers, r. p., and w. jetz. 2019. global habitat loss and extinction risk of terrestrial vertebrates under future land-use-change scenarios. nat. clim. change 9(4): 323-329. radinger, j., f. essl, f. hölker, p. horký, o. slavík, and c. wolter. 2017. the future distribution of river fish: the complex interplay of climate and land use changes, species dispersal and movement barriers. glob. change biol. 23(11):4970-4986. doi:10.1111/ gcb.13760 r core team. 2020. r: a language and environment for statistical computing. r foundation for statistical computing. vienna, austria. rosa, m.r., p. h. brancalion, r. crouzeilles, l. r tambosi, p. r. piffer, f. e. lenti, m. hirota, e. santiami,and j. p. metzger, j.p. 2021. hidden destruction of older forests threatens brazil’s atlantic forest and challenges restoration programs. sci. adv. 7(4): eabc4547. doi:10.1126/sciadv.abc4547 ruiz-benito, p., g. vacchiano, e. r. lines, c. p. o. reyer, s. ratcliffe, x. morin, f. hartig, a. mäkelä, r. yousefpour, j. e. chaves, a. palacios-orueta, m. benito-garzón, c. morales-molino, j. j. camarero, a. s. jump, j. kattge, a. lehtonen, a. ibrom, and m. a. zavala. 2020. available and missing data to model impact of climate change on european forests. ecol. model. 416:108870. doi:10.1016/j.ecolmodel.2019.108870 salk, c., s. fritz, l. see, c. dresel, i. and mccallum. 2018. an exploration of some pitfalls of thematic map assessment using the new map tools resource. remote sens. 10(3):376. doi:10.3390/rs10030376 schmunk, r. b. 2017. panoply: netcdf, hdf and grib data viewer. nasa goddard institute for space studies. version 4.7. shapiro, a. c., h. s. grantham, n. aguilar-amuchastegui, n. j murray, v. gond, d. bonfils,and o. rickenbach. 2021. forest condition in the congo basin for the assessment of ecosystem conservation status. ecol. indic.122:107268. doi:10.1016/j.ecolind.2020.107268 sharp, r., j. douglass, s. wolny, k. arkema, j. bernhardt, w. bierbower, n. chaumont, d. denu, d. fisher, k. glowinski, r. griffin, r. g. guannel, a. guerry, j. johnson, p. hamel, c. kennedy, c. k. kim, m. lacayo, e. lonsdorf, l. mandle, l. rogers, j. silver, j. toft, g. verutes, a. l. vogl, s. wood, and k. wyatt. 2020. invest 3.9.0.post177+ug.gb77feed user’s guide. the natural capital project, stanford university, university of minnesota, the nature conservancy, and world wildlife fund8. silva, r. d. o., l. g. barioni,and d. moran. 2021. fire, deforestation, and livestock: when the smoke clears. land use policy 100:104949. sobral-souza, t., j. p. santos, m.e. maldaner,m. s. lima-ribeiro, and m. c. ribeiro. 2021. ecoland: a multiscale niche modelling framework to improve predictions on biodiversity and conservation. perspect. ecol. conserv. (in press). doi:10.1016/j.pecon.2021.03.008 sofaer, h.r., c. s. jarnevich, i. s. pearse, r. l. smyth, s. auer, g. l. cook, t. c. edwards jr, g. f. guala, t. g. howard, j. t. morisette, and h. hamilton. 2019. development and delivery of species distribution models to inform decision-making. bioscience 69(7):544557. doi:10.1093/biosci/biz045 svensson, j., j. andersson, p. sandström, g. mikusiński, and b. g. jonsson. 2019. landscape trajectory of natural boreal forest loss as an impediment to green infrastructure. conserv. biol. 3(1):152-163. taylor, k. e., r. j. stouffer, and g. a. meehl. 2012. an overview of cmip5 and the experiment design. bull. am. meteorol. soc. 93:485-498. doi:10.1175/ bams-d-11-00094.1. titeux, n., k. henle, j. b. mihoub, a. regos, i. r. geijzen-dorffer, w. cramer, p. h. verburg, and l. bro8 https://storage.googleapis.com/releases.naturalcapitalproject.org/invest-userguide/latest/index.html. http://doi.org/10.1016/j.rse.2015.01.018 https://doi.org/10.1093/condor/103.3.599 http://doi.org/10.1016/j.gloenvcha.2016.10.002. http://doi.org/10.1016/j.gloenvcha.2016.10.002. https://doi.org/10.1111/gcb.13760 https://doi.org/10.1111/gcb.13760 http://doi.org/10.1126/sciadv.abc4547 http://doi.org/10.1016/j.ecolmodel.2019.108870 http://doi.org/10.1016/j.ecolmodel.2019.108870 https://doi.org/10.3390/rs10030376 http://doi.org/10.1016/j.ecolind.2020.107268 http://doi.org/10.1016/j.pecon.2021.03.008 https://doi.org/10.1093/biosci/biz045 http://doi.org/10.1175/bams-d-11-00094.1. http://doi.org/10.1175/bams-d-11-00094.1. https://storage.googleapis.com/releases.naturalcapitalproject.org/invest-userguide/latest/index.html biodiversity informatics, 16, 2021, pp. 28-38 38 tons. 2017. global scenarios for biodiversity need to better integrate climate and land use change. divers. distrib. 23:1231-1234. doi:10.1111/ddi.12624. vale, m. m., l. tourinho, m. l. lorini, h. rajão, and m. s. l. figueiredo. 2018. endemic birds of the atlantic forest: traits, conservation status, and patterns of biodiversity. j. field ornithol. 89: 193-206. doi:10.1111/ jofo.12256. vale, m. m., berenguer, e., de menezes, m. a., de castro, e. b. v., de siqueira, l. p., and, portela, r.c.q. 2021. the covid-19 pandemic as an opportunity to weaken environmental protection in brazil. biol. conserv. 255:108994. doi:10.1016/j.biocon.2021.108994 vuuren, d. p. van, j. a. edmonds, m. kainuma, k. riahi, and j. weyant. 2011. a special issue on the rcps. clim. change, 109: 1-4. doi:10.1007/s10584-0110157-y. zabel, f., delzeit, r., schneider, j. m., seppelt, r., mauser, w., and václavík, t. 2019. global impacts of future cropland expansion and intensification on agricultural markets and biodiversity. nat. comm.10(1): 1-10. doi:10.1038/s41467-019-10775-z http://doi.org/10.1111/ddi.12624. http://doi.org/doi:10.1111/jofo.12256 http://doi.org/doi:10.1111/jofo.12256 https://doi.org/10.1016/j.biocon.2021.108994 http://doi.org/10.1007/s10584-011-0157-y. http://doi.org/10.1007/s10584-011-0157-y. https://doi.org/10.1038/s41467-019-10775-z biodiversity informatics, 18, 2024, pp. 1-12 1 the ibdata web system for biological collections: design focused on usability miguel murguía-romero1*, bernardo serrano-estrada2, gerardo a. salazar1, gerardo e. sánchez-gonzález3, ubaldo melo-samper-palacios1, david s. gernandt1, susana magallón1, víctor sánchez-cordero1 1instituto de biología, universidad nacional autónoma de méxico, circuito zona deportiva s/n, ciudad universitaria, coyoacán, 04510 mexico city, mexico. 2seres sistemas especializados, membrillos mz. 120 lt. 38. col. ojo de agua, tecámac, 55770 estado de méxico, mexico. 3facultad de ciencias, universidad nacional autónoma de méxico. *corresponding author: miguel murguía-romero, email: miguel.murguia@ib.unam.mx abstract. the software design process must put users at the core of the process to enable them to meet their specific objectives effectively, efficiently, and successfully. thus, a software design for a computing system to consult biological collections guided by the concept of usability will result in an effective and efficient biodiversity informatics tool. here, we introduce ibdata, a web system to consult biological collections, developed using a design approach based on the architecture of three layers: database, business rules, and user interface. the user interface design was guided by the concept of usability focused on four core concepts: simplicity, adaptability, guide the user through the journey, and feedback. the ibdata web system that we developed is composed of three modules (query, capture and editing, and administration), permitting it to query a database with about 1.7 million specimen records. biodiversity data query systems must be effective and efficient and should meet the user’s expectations. software design methodologies play a central role in achieving these goals, and, in this context, interface design techniques that put the user at the core of development are valuable, as in the development of the ibdata web system. key words: biodiversity informatics, biological databases, geotax search, software design, user experience, ux design, ui design. biological collections document the biodiversity of our planet, and in many cases, they are the result of the efforts of many people over decades or centuries (penn et al., 2018). in the face of the biodiversity crisis (wilson, 1985; sandor et al., 2022), biological collections may represent the only places where recently extinct species are found. biological collections are the primary source of information for many types of research such as taxonomy, wildlife, floristics, and conservation biology studies, biodiversity analyses, phylogenetic and even phylogeographic analyses, as well as a variety of studies to investigate the evolutionary processes associated with the origin and maintenance of biological diversity (castillo-figueroa, 2018). for this reason, computer systems for consulting information in biological collections are central tools for conducting research and generating knowledge. these tools should be built by putting their users at the center of their design. in recent years, user interface design has benefited from the concept of usability (bevan et al., 2016; iso, 2018), understood as the degree to which specific users can use a system, product, or service to achieve specific objectives with effectiveness, efficiency, and satisfaction. thus, the process of building web systems for biological collections may be guided by the concept of usability if the goal is to provide efficient systems for the user. many biological collections are framed in educational contexts, being central in the training of undergraduate and graduate students. in addition, they play an important role in raising awareness in society about conservation and biodiversity issues (wen et al., 2015). for these reasons, it is important to make the associated data of the specimens available through the internet for universal access. in addition to the availability of mailto:miguel.murguia@ib.unam.mx miguel murguía-romero et al. – the ibdata web system 2 biodiversity data online, it is important that the computer systems through which this is achieved ensure aspects of correctness and speed. with computer and communication technologies such as databases, the world wide web and the internet, biological collections are accessible to more users if they can be consulted virtually. a clear example is the global biodiversity information facility gbif1 (gaiji et al., 2013), an intergovernmental effort to establish a global infrastructure for access to primary biodiversity data. initiatives aimed at digitizing biological collections around the world are very diverse. although some of the largest collections are already available in public databases, for many collections especially from developing countries, whose biological diversity may be very high, these efforts are just beginning. for example, the william and lynda steere herbarium at the new york botanical garden holds 7.8 million specimens, of which 4 million have been digitized2, while in the herbarium of escuela nacional de ciencias biológicas, instituto politécnico nacional (encb: mexico’s national school of biological sciences of the national polytechnic institute of mexico), with about one million specimens (thiers, 2018), only three families of flowering plants have been digitized3. the biological collections of the instituto de biología, universidad nacional autónoma de méxico (ibunam: institute of biology, national autonomous university of mexico), which includes zoological, mycological, botanical, and ethnobiological specimens, represent one of the most important resources for the study of mexico’s biodiversity. the vast majority are national collections, received by the unam at the time of its foundation as part of the national heritage of mexico. two examples are the national herbarium (mexu), which, with over 1.5 million specimens, houses the most representative collection of plants collected in mexico, and among the zoological collections, the national mammal collection, the most important of its kind in the country. during the last 60 years, efforts have been made to digitize the biological collections of the ibunam (e.g., scheinvar et al., 1967, 1968; gómez-pompa et al., 1975). in the 2000s, various databases were created, mainly within the framework of projects supported by the comisión nacional para el conoci1 https://www.gbif.org/ 2 https://sweetgum.nybg.org/science/ 3 https://www.encb.ipn.mx/herbario/ miento y uso de la biodiversidad de méxico4 (conabio: national commission for the knowledge and use of biodiversity). these efforts resulted in the contribution of about 200,000 records to the sistema nacional de información sobre biodiversidad database (snib: national biodiversity information system5). however, these efforts represented isolated initiatives and many specimens remained to be digitized. it was not until 2012 that the ibunam successfully undertook a comprehensive project aimed at digitizing most of its collections with the financial support from conabio and the coordinación de la investigación científica, unam (cic: coordination for scientific research, unam; gernandt et al., 2014; sánchez-cordero et al., 2021). to date, the total number of digitized records is approximately 1.7 million, which represents about 85% of the herbarium and most of the zoological collections, except for insects. the efforts of working groups in the construction of these systems have been great, since they have covered needs correctly (both in functionality and validity of the information) and with reasonably fast responses to online queries. conabio has developed various computer platforms for promoting the development of knowledge on the biodiversity of mexico, and the experience accumulated is valuable. among the most important examples of their systems are the snib, the red mundial de información sobre biodiversidad (remib: world network of information on biodiversity6); and the biótica system (conabio, 2012) for the capture and editing of biodiversity data, for example, from the specimen labels of biological collections (jiménez et al., 2016). the remib, along with north america biodiversity information network (nabin) from the biodiversity research center at the university of kansas, were important as examples for the subsequent development of gbif (soberón, 1999). to our knowledge, there is no web system for accessing biodiversity records that explicitly incorporates usability concepts into its design process. unfortunately, there is no documentation on the design process of any biodiversity web system. the only reference found is that of a system to build taxonomic identification keys on the web (murguía-romero et al., 2021), which reports the use of usability in its design process incorporating responsive web design technology. therefore, the documentation of the de4 https://www.biodiversidad.gob.mx/conabio/ 5 https://www.snib.mx/ 6 http://www.conabio.gob.mx/remib_ingles/doctos/remib_ing.html https://www.gbif.org/ https://sweetgum.nybg.org/science/ https://www.encb.ipn.mx/herbario/ https://www.biodiversidad.gob.mx/conabio/ http://www.snib.mx http://www.conabio.gob.mx/remib_ingles/doctos/remib_ing.html miguel murguía-romero et al. – the ibdata web system 3 sign process of a biodiversity system that incorporates usability concepts may benefit the community involved in the development of this type of tools. here we present the development process of a web system to consult biological collections guided by the concept of usability focused on four core concepts: simplicity, adaptability, guide the user through the journey, and feedback. the development of the web system for consulting the biological collections data presented here owes much to previous systems, among them, and very importantly, those developed by conabio (remib, snib, and biótica). unam’s portal de datos abiertos7 (open data portal, unam) was also a key platform for the development of this system, mainly in the design of how to summarize the list of records that results from a search. software engineering the objective of software engineering is to propose and systematize techniques and tools for the development of computer systems that seek to optimally use human, economic, and time resources to achieve the construction of computer programs that meet the objectives for which they are created. a group of very important tools in this discipline are the “development methodologies,” which emerged in the 1970s as a response to the “software crisis,” a term that was used to designate the problem of the high proportion of failures in the generation of computer programs that did not meet their objectives, or that were never completed. building software is a complex process, and therefore, if it does not follow a formal development methodology, it is likely to fail. although there are many differences in the estimates of the proportion of projects that are successful, approximately 50% have difficulty reaching their objectives and only 16% to 18% of software projects are successful, that is, they come to fruition (glass, 2005). usability in the last decade, ergonomics has gained importance in the development of computer systems. specifically, it has been shown that evaluating the efficiency of user interfaces contributes to the success of the system, in that the system meets the tasks and objectives for which it was created. usability is the degree to which specific users can use a system, product, or service to achieve specific objectives with effectiveness, efficiency, and 7 https://datosabiertos.unam.mx/biodiversidad/ satisfaction in a specific context of use (bevan et al., 2016; iso, 2018). it is a relative measure because it provides an evaluation by comparing different designs; it is also subjective, since it depends on the user and on the information that she or he is interested in manipulating (allanwood & beare, 2014). various dimensions of usability have been defined; among the most important are: • learning ability: how easy it is for users to perform basic tasks the first time they interact with the system. • efficiency: how fast users can perform tasks. • memorability: when users return to the system after a period of inactivity, how easily they can remember how to use it. • errors: how many errors users make, how serious they are, and how easy it is to recover from them. • satisfaction: how satisfying the users find the system. • attractiveness: how much the interface invites the user to interact and how pleasant it is to use. usability tests determine if interactions through the interface meet the needs and objectives of the user and are an important element of the design methodology and interfaces (allanwood & beare, 2014; jackson & ciolek, 2017). system development process an “agile development methodology” was followed to build the ibdata web system (beck et al., 2001; abrahamsson et al., 2002). the development team was formed in march 2018, when work began formally. the documentation of the process was intended to be as agile and summarized as possible, basically consisting of minutes from meetings, drafts or layout of the interface design, executive presentations to the working group, user guides, analysis of the information to explore the best ways to process and present it in the interface, and suggestions for modifications to the system. users of various profiles of the system were represented on this team. the functionalities that the system should have were decided based on the experience of information needs of the participants in the project and by making comparisons with other systems that had already been built. the main functionalities included in the system design specification were classified into three types: query, update, and administration: https://datosabiertos.unam.mx/biodiversidad miguel murguía-romero et al. – the ibdata web system 4 query: • obtaining lists of records of specimens filtered by the most important fields of the database, such as taxonomic (family, genus, species), the geographic location of the collection (country, state, municipality, coordinates), the date of collection (year, month, day), or by administrative data (collection, catalog number). • viewing a summary at the level of the specimen record with the most relevant fields. • visualization of the image associated with the specimen when it exists. • summaries of the content of the database in three formats: graphs, lists (of taxa or geographic entities), total number of records by different types of grouping (e.g., by taxa or by geographic entities). update: • editing of existing records. • capture of new records. • deletion of records. administration: • user registration in the system and password recovery. • updating by the user of their contact information and password. • specification of differential permissions per user. • reports on system usage statistics. the limits of the desired functionalities were defined by specifying what is not expected of the system; for example, common or frequent functionalities that are desired from a system that stores information on biological collections but that can be covered by existing systems; or, one that is highly complex and its implementation would increase the required resources, putting the system at risk to come to fruition or simply because it departs from the central objectives for which the system is intended to be used. some of the limits that were established were as follows. (1) the system will not provide analysis tools; there are multiple packages that can be used for this with data exported from the system. (2) the system will not attempt to solve particular requirements of the academic staff (i.e. requirements that only solve the needs of one researcher and are not useful to others), as it is intended to be the institutional computer repository of the biological collections of the ibunam. (3) the system will be released incrementally, so that it was expected that the first versions would only provide some basic functionalities, and more functionalities would be added in later versions. integrating usability into the development process usability is directly related to the user experience, and therefore the user interface is one of the main components of the system where developers can participate, seeking the objectives of usability: that users can use the system to achieve their specific objectives with effectiveness, efficiency, and satisfaction. the way we achieved this was to start with the selection of four key usability concepts (simplicity, feedback, adaptability, and journey guidance). the second step was to define the way in which these four concepts could be represented in the user interface, taking into account the technological elements available in the development of interfaces (such as windows, types of controls, ways to navigate, colors, among others). the third step was to integrate the different components into a model (information architecture), arranging them with a congruent signature both visually and functionally. the fourth step was to implement that mockup into a functional system. the fifth and final step was to develop usability tests that guided adjustments to the user interface. the process of building the user interface is explained below, which together with the database and business rules make up a three-layer system. user interface design (ui design) one of the central aspects for the success of a system is that it is used by the people for whom it is designed; that is, that the users use the system. there are other central aspects that cannot be ignored for a successful computer system, such as meeting the needs for which the system is built, and doing it correctly and reasonably quickly, i.e., correctness and speed. when the development team works with appropriate methodologies and those two aspects are taken care of, the question of the user interface is central to achieve the success of the development. usability is precisely the concept that guides the application of design techniques to achieve a good user experience. the following four axes were defined as usability criteria to achieve a good user experience in the system (figure 1): miguel murguía-romero et al. – the ibdata web system 5 • axis 1: simple exploration. commands available to the user with minimal hierarchy. • axis 2: guide the user in their journey. searches focused on geography and taxonomy as the main search criteria. • axis 3: feedback in the queries (query feedback). dialectic between the queries and their results. • axis 4: adaptability to the user. responsive and multilingual. axis 1 of usability: simple exploration (commands available to the user with minimal hierarchy). every interface designer wants their product to be simple to use by all the targeted users. but the criteria for simplicity of use are relative, both to the user and to the state of the technology available at the time it is developed. simplicity can be seen as the intelligent management of complexity (allanwood & beare, 2014). this conception is important especially if it is considered within the development team, because if the project management has that in mind, it will be able to recognize that the solutions that provide simplicity to the system are the result of a non-random and conscious analysis that shows the user in a simple way the intrinsic complexity of the system and the information it contains. the specification of the system design was defined so that the simplicity focused on ease of access to available commands. this accessibility should be visual and as far as possible, immediate; that is, a panel should be provided that presents a complete, explicit, and organized picture as possible of the operations that could be requested from the system in the current state of the session. thus, two of the criteria for the design of the interface were to a) avoid implicit commands, that is, ensure that they had a visual reference in the interface, and b) arrange the commands with minimal hierarchy. the proposed solution was to present the application-specific commands to the user on an always-visible bar containing them. the commands for general use, which can be found in other applications, such as login, logout, registration, and “i forgot my password,” among others, were placed in the usual way as a menu on the top bar and as a hamburger menu. keeping the largest number of commands in a place of common use is something that provides simplicity in the interface, since the users find what they need where they expect it. but placing the specific commands of the application in such spaces would require a user-oriented search effort. for this reason, it was decided that they should be as explicit and visually obvious in the interface (figure 1 “simplicity” section). therefore, its expression in the interface would be buttons with associated icons and tooltips. if the number of buttons required was excessive and difficult to accommodate in the interface, they could be ranked in no more than one level. axis 2 of usability: guide users in their journey; searches focused on geography and taxonomy. the interface should help the user to decide how to conduct a search, suggesting the types of filters that can be the simplest and most effective. there are different types of information that are customarily integrated in biodiversity databases, for example, administrative information (such as the name of the collection or the catalog number of the specimens), temporal information (e.g., the collection date), or description of the biology of the specimen (such as the phenological state or its size as measured by the collector), among other types. in this context, the interface is designed to interact with the user bearing in mind that errare humanum est, and that it is by means of a user-interface dialectic that the system can become, through successful and unsuccessful cycles of use, an efficient and effective tool. axis 3 of usability: feedback in the queries (dialectic between the queries and their results). by feedback in the queries, we mean the return of the system’s output information to the query that generated figure 1. the four axes considered in the design of the system to achieve a high level of usability. simplicity: the available commands to consult the database are visible. feedback: visual and operative dialectic between the queries and their results. adaptability: responsive interface changing its aspect according to the dimensions of the device. journey guidance: the system informs the user where and how to perform a query of the database. miguel murguía-romero et al. – the ibdata web system 6 it. the user-system feedback focused on the user’s queries and the corresponding results reported by the system. there are other types of feedback in a user interface, such as the current state of the system (e.g., waiting, working, on, selected, among many others), the indication of consistency (or inconsistency) of the combination or sequence of commands selected by the user, the drill-down, warnings or additional notes that support or guide the user in performing tasks, among others. of all the types of feedback that are usually given in a user interface, it was decided to focus on the one that occurs between the user’s requests (queries) and the corresponding results that the system provides. it was decided to implement the other types of feedback following the standards suggested in the design area, without analyzing or specifying in greater depth how to improve it. this feedback idea takes its heuristics from another system whose theoretical foundation is a “dynamic identification model” (murguía-romero, 1992; murguía-romero et al., 2021). the results can be queried; that is, it is possible to filter them so that the user can refine the query in a recursive, or spiral process. in scientific databases, which are usually extensive and rich (that is, with many records and with many tables, fields, and relationships), the user’s queries must be interpreted for their intention rather than taking them as a fixed structure, an immovable contract between the user and the system, from which a complete and correct answer is expected. the results of the queries should help the user to understand the content of the database and provide guidance to continue exploring the data (kersten et al., 2011). axis 4 of usability: adaptability to the user (responsive and multilingual). by adaptability it is understood that the system recognizes the characteristics of the type of use that the user gives to the system, or the preferences that the user establishes on it. this adaptability can occur automatically or be directed by the user. in the design specification, two specific criteria of adaptability were established: that the system was responsive; that is, that the interface was adapted according to the dimensions of the screen of the device in which it is used, and that the user could choose the interface language with at least two options. the first specification is automatic, while the second, the language, is chosen by the user. another specification of the design on adaptability that has not yet been implemented is the user profile; that is, the system stores the preferences of the environment indicated by the user. for example, when the user logs in, the system restores the environment according to the user’s preferences, including the interface language, the collection to consult by default, the display of the most used search type, or the default taxonomic family searched, among others. usability testing the design was incrementally implemented in functional prototypes. in each, usability tests were carried out, which guided the adaptation of the interface, its functionality, the graphic elements, and the associated texts, among many other elements. some examples of the modifications are the adjustment of the location, shape, size, and color of each of the interface elements, as well as its structure and arrangement, among many others, seeking in each cycle to improve the user experience. figure 1 shows the general layout that guided the design and from which the interface was implemented in programming languages. the usability tests consisted of inviting 15 faculty and students from ibunam to use the ibdata web system. after an average of ten days, users were interviewed by the designers to record their observations, which were incorporated by making pertinent changes in the interface. this process was repeated three times incrementally, that is, inviting new users to join those who participated in the previous tests. the usability measurement consisted of a measure relative to each user, verifying that they reported an increasingly better user experience. the context of each user was also taken into account, depending on the type of information or reports that interested them. in this sense, they were classified into three types of users: curator, taxonomist, and ecologist. prior to the usability tests, three respective fictitious characters were created that guided the design (allanwood & beare, 2014). finally, these characters were reflected in the data view of the interface in the form of a table, which shows a different set of fields depending on the chosen profile. information architecture the design must define and plan the way in which the interface will present the flow of information or its subsets to the user, either through requests that they make or at the suggestion of the system. this specification, known as “information architecture”, must differentiate between information that is relevant to the user and that which is secondary. it is in this design space that usability-oriented design specifications were concentrated. the resulting information architecture considered the integration of each of the four chosen usability criteria axes (figure 2), miguel murguía-romero et al. – the ibdata web system 7 and its objective is to show, grosso modo, the desired appearance of the interface in its functional version. institutional standards (unam, 2016) imposed demands on design, such as the location and size of the logos, copyright notice, and a visits counter. these were attended by locating the elements involved in an upper and a lower band (header and footer bands in figure 2). implementation of designed components architecture of three layers the web system, called ibdata8, allows the consultation of a database with so far, 1.7 million speci8 http://ibdata.ib.unam.mx/ men records of the biological collections of the ibunam (murguía-romero et al., 2023). it is made up of three modules: queries, capture and editing, and administration. the three layer architecture is implemented through a) a database with 70 tables, b) the program code in various programming languages, and c) the interface that the user accesses through the web. user interface the web interface (figure 3) takes into account the specifications of the information architecture (figure 2). based on this architecture, the various previously defined elements were adapted, mutatis figure 2. the design of the information architecture for ibdata. the main structure of the interface presented to the user (information architecture) is composed of 5 areas: three bands (header, tools, and footer) and two windows (query window and results window). figure 3. components of the ibdata bifenestra user interface. the result of the query “specimens of the orchidaceae of mexico in the vascular plants collection” is shown through “advanced search” (left window) showing the results in the “web view” (right window). http://ibdata.ib.unam.mx/ miguel murguía-romero et al. – the ibdata web system 8 mutandis, in a more specific way. it is made up of three bars, two above and one below; and two windows, one to the left and one to the right, which we called a “bifenestra” user interface (figure 3). the first upper bar fulfils the institutional requirement that web pages must contain official names and logos. the second upper bar contains the menus and commands that web systems usually have, such as the user menu (with the login and logout commands, access to edit the account data, among others), an icon for language selection, and a hamburger menu with commands to view general system information and credits. in this second bar an icon of a drum was also included, which allows the user to select the collection to consult. this considers a display on a large screen, but because the system is responsive, these elements change in their arrangement and presentation on smaller screens. between the top and bottom bars, and occupying most of the general window, are the query window (left) and the results window (right). each of these windows contains in its upper part a list of buttons with which commands are accessed, either to perform queries or to manipulate the results, respectively. the result of a search is presented in the right window, typically a list of specimen records in a summarized format, which has been called “web view”. the format can be changed to “table view” by choosing the corresponding command in the results window. clicking on any of the records displays the “specimen data summary sheet.” if the specimen has an associated image, it is shown as a thumbnail, which upon clicking on it, displays the high resolution image in a dialog where it can be explored by zooming in on the portions of interest to the user. table 1 summarizes the most important characteristics of the web system that were imposed as system requirements in the design stage, always guided by the purpose of achieving a good user experience. query module the interface reserves the left window as the space for the specification of user queries. queries obtain lists of specimens filtered by different methods: simple search (filter by genus or species name), geotax search (family-genus-species and country-state-county hierarchy), and advanced search, in which the user can choose any set of fields to filter the specimens (figure 4). it is in this space that it was decided to focus the guide of the user journey, because through text boxes the user is invited to make a successful query depending on the level of knowledge they have about his or her question or the available data. table 1. characteristics of the ibdata web system. feature description 1 responsive system the interface adapts to the size of the device screen 2 compatibility with the most used operating systems windows, mac os, linux kernel, ios, and android 3 compatibility with the most used internet browsers edge, firefox, google chrome, opera, and safari 4 level 1 navigation most of the functionalities are on the same page; it is not necessary to navigate outside of the current page 5 100% web user does not need to install components 6 user login grants different access permissions and customizes some aspects of the interface 7 multilingual spanish / english / italian with the ability to add more languages easily 8 multi-collection stores data from more than one collection 9 darwin core standard field names according to the darwin core standard 10 query module advanced simple queries, lists, totals; export to pdf, excel, and csv 11 administrative reports reports for the system administrator on usage and capture statistics 12 managing users accounts assigns user permissions 13 capture and edit module add, edit, and delete records miguel murguía-romero et al. – the ibdata web system 9 the simple search allows searching for specimens by species name. if users are taxonomists (or know the taxonomic names) and want information on the specimens of a particular species, the simple search indicates in the text box to type the name of the species (figure 4, “simple search” panel). the geotax search is a recurring requirement of users when they want to display the records of specimens filtered by some category of political geography (country, state, or county) or by some taxonomic category (family, genus, or species). it has two variants or options, free and structured. the free search is oriented to situations in which the user does not require guidance in specifying the taxon name, or when she or he precisely know the name of the political geography entity for which they wish to apply a filter (figure 4 “free geotax search” panel). if the user is uncertain about the names of the taxa or geographic entities to consult, the structured geotax search provides drop-down lists with which the user can hierarchically choose the name of the geographic entity (country, state, or county) or taxon (family, genus, and species) to be included in the query (figure 4 “structured geotax search” panel). the advanced search allows users to choose the fields by which to filter the specimens (figure 4 “advanced search” panel). if the field chosen is of numeric type, for example the year of collection, the interface shows two boxes in which the lower and upper limits can be specified. if the field is of text type, for example for collection location, the interface interprets the text specified in the corresponding box as a substring of the location, that is, it selects those records that in the location field contain the text written by the user. the user can choose and add any number of fields to the query. when more than one field is included in the query, it is interpreted as a conjunction (as a logical “and”), requiring that all the conditions represented by the values indicated for each field must be met. edit and capture module the ibdata edit and capture module can be used from any computer connected to the internet. it was designed considering various features aimed at the best efficiency for users. the capture interface is not available to all users, only to those who have been granted the corresponding permissions. it contains basic validations to ensure its quality. in some cases, the values must be chosen from a list whose source is a catalog of the database, for example, the country or the state; for fields which the user can write in a free format, validations are made according to rules that ensure integrity and congruence. administrative module the administrative functions are basically of two types: assigning permissions to users and their administration (delete, block, or add users) and display of usage reports and record capture. the administrator has access to the list of registered users in that interface, and by sorting by names, surnames, or some other field, can access the permissions assigned to edit them. lessons learned systems for consulting biodiversity data, such as specimen records of biological collections, are complex, and subject to failure risk. therefore, the use of development methodologies arising from the field of software engineering are essential tools. in its planning, the architecture must be defined, among other aspects, to maintain the highest possible data independence. likewise, the design of user experiences should benefit the greatest number of users, supporting them in efficiently obtaining the required information. the concept of usability is the central element to achieve better experiences. in this work it has been shown how the choice of four usability axes can guide the construction of a system that meets these goals. figure 4. query options in the ibdata web system. simple search: the user enters the name of a genus or species. geotax free search: the user enters search terms in one or more text boxes out of the seven displayed (collection, family, genus, species, country, state, or county). geotax structured search: the user must specify hierarchically, choosing from lists. advanced search: the user can select the fields with which to perform the filter. miguel murguía-romero et al. – the ibdata web system 10 conclusions in the development of systems for querying biodiversity data, such as those that allow the query and administration of biological collections, it is a conditio sine qua non to involve the user in all stages of development if such systems are to meet user expectations effectively and efficiently. software design methodologies play a central role in achieving these goals, and in this context, interface design techniques that put the user at the center of development are valuable tools. ibdata is an institutional system developed by the instituto de biología, unam, for consulting, capturing, and editing specimen information of the collections that it houses. its creation followed a software development methodology focused on data independence, and the concept of usability was considered at all stages to meet the expectations of its users as much as possible. acknowledgments many faculty members and students participated in the usability tests of the system, which contributed to the success of its development. the comisión nacional para el conocimiento y uso de la biodiversidad (conabio project “ke002: digitization and systematization of the national biological collections of the institute of biology, unam”) and the coordinación de la investigación científica, unam financially supported the digitization of the collections of the institute of biology, unam. javier nori and three anonymous reviews made suggestions to the manuscript that improved it. author contributions m.m.r. wrote the first draft with contributions from b.s.e. who developed the program code of the system; m.m.r. and b.s.e. designed the system database; all authors contributed to the design of the system and proof-read the manuscript. competing interests the authors have declared that no competing interests exist. literature cited glass, r.l. it failure rates—70 percent or 10–15 percent? ieee software 22, 3 (may–june 2005). abrahamsson, p., salo, o., ronkainen, j. & warsta, j. (2002) agile software development methods: review and analysis, vtt publication 478, espoo, finland, 107p. allanwood, g., and p. beare. 2014. basics interactive design: user experience design: creating designs users really love. a&c black. beck, k., m. beedle, a. bennekum van, a. cockburn, w. cunningham, m. fowler, j. grenning, j. highsmith, a. hunt, r. jeffries, j. kern, b. marick, r. martin, s. mellor, k. schwaber, j. sutherland and d. thomas (2001). “manifesto for agile software development.” 2002(22.3.2002) http:// agilemanifesto.org. bevan, n., j. carter, j. earthy, t. geis, and s. harker. 2016. new iso standards for usability, usability reports and usability measures. lecture notes in computer science (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics), 9731, 268–278. https://doi. org/10.1007/978-3-319-39510-4_25 castillo-figueroa, d. 2018. beyond specimens: linking biological collections, functional ecology and biodiversity conservation. revista peruana de biología, 25(3), 343-348. coordinación de colecciones universitarias digitales (ccud). 2017. manual de datos abiertos de colecciones universitarias digitales. secretaría de desarrollo institucional, universidad nacional autónoma de méxico (sdi-unam). méxico. 117pp. conabio. 2012. sistema de información biótica. manual de usuario. versión 5.0.3 méxico, d.f. 321pp. http:// www.conabio.gob.mx/biotica/cms/descargas/biotica50/ manualbiotica50/biotica50.pdf gaiji, s., v. chavan, a. h. ariño, j. otegui, d. hobern, r. sood, and e. robles. 2013. content assessment of the primary biodiversity data published through gbif network: status, challenges and potentials. biodiversity informatics, 8(2). gernandt, d. s., v. sánchez-cordero, u. melo samper palacios, o. j. giménez-héau, and g. a. salazar. 2014. digitalización del herbario nacional de méxico: avances y retos del futuro. revista digital universitaria, vol. 15, no. 4, pp. 1-13. grattarola, f., g. botto, i. da rosa, n. gobel, e. m. gonzález, j. gonzález, j. gonzállez, d. hernández, g. laufer, r. maneyro, j. a. martínez-lanfranco, d. e. naya, a. l. rodales, l. ziegler, and d. pincheira-donoso. 2019. biodiversidata: an open-access biodiversity database for uruguay. biodiversity data journal, 7. glass, r. l. 2005. it failure rate70% or 10-15%? ieee software 22, 3 (may-june 2005) gómez-pompa, a., j. a. toledo, and m. soto. 1975. electronic data processing of herbarium specimens data for the flora of veracruz program. in computers in botanical collections (pp. 35-51). springer, boston, ma. gries, c., e. e. gilbert, and n. m. franz. 2014. symbiota a virtual platform for creating voucher-based biodiversity information communities. biodiversity data journal, (2), e1114. doi:10.3897/bdj.2.e1114 iso. 2018. 9241-11: 2018 ergonomics of human-system interaction—part 11: usability: definitions and concepts. international organization for standardization, https://kebs. http://agilemanifesto.org http://agilemanifesto.org https://kebs.isolutions.iso.org/obp/ui#iso:std:iso:9241:-11:ed-2:v1:en miguel murguía-romero et al. – the ibdata web system 11 isolutions.iso.org/obp/ui#iso:std:iso:9241:-11:ed-2:v1:en, 9241(11). jackson, c., and n. ciolek. 2017. digital design in action: creative solutions for designers. ak peters/crc press. james, s. a., p. s. soltis, l. belbin, a. d. chapman, g. nelson, d. l. paul, and m. collins. 2018. herbarium data: global biodiversity and societal botanical needs for novel research. applications in plant sciences, 6(2), e1024 jiménez, r., p. koleff et al. 2016. la informática de la biodiversidad: una herramienta para la toma de decisiones, en capital natural de méxico, vol. iv: capacidades humanas e institucionales. conabio, méxico, pp. 143-195. kersten, m. l., s. idreos, s. manegold, and e. liarou. 2011. the researcher’s guide to the data deluge: querying a scientific database in just a few seconds. proceedings of the vldb endowment, 4(12), 1474-1477. morville, p., and l. rosenfeld. 2006. information architecture for the world wide web: designing large-scale web sites. o’reilly media, inc. murguía-romero, m. 1992. métodos en la identificación biológica automatizada (master’s thesis). facultad de ciencias, universidad nacional autónoma de méxico. méxico d.f. http://132.248.9.195/pmig2017/0187269/index.html murguía-romero, m., b. serrano-estrada, e. ortiz, and j. l. villaseñor. 2021. taxonomic identification keys on the web: tools for better knowledge of biodiversity. revista mexicana de biodiversidad, 92: e923592. murguía-romero, m., serrano-estrada, b., salazar, g. a., gernandt, d. s., melo-samper-palacios, u., sánchezgonzález, g. e., sánchez-cordero, v., and magallón, s. 2023. ibdata versión 3 “helia bravo hollis”: manual de uso. instituto de biología, universidad nacional autónoma de méxico. https://www.ib.unam.mx/ib/programa-editorial/publicaciones-especiales/; doi.org/10.22201/ ib.9786073073189e.2023 nielsen, j. 2012. usability 101: introduction to usability. nielsen norman group. available from: https://www.nngroup. com/articles/usability-101-introduction-to-usability/. accessed on sep 10, 2021. penn, m. g., cafferty, s., and carine, m. 2018. mapping the history of botanical collectors: spatial patterns, diversity, and uniqueness through time. systematics and biodiversity, 16, 1–13. https://doi.org/10.1080/14772000.2017.1355854 sánchez-cordero, v., s. magallón, a. contreras-ramos, g. salazar, d. s. gernandt, e. gonzález, u. melo samper, j. giménez, d. pérez, c. reséndiz, and m. murguía. 2021. digitalización y sistematización de las colecciones biológicas nacionales del instituto de biología, unam. universidad nacional autónoma de méxico. instituto de biología. informe final snib-conabio proyecto ke002. ciudad de méxico, méxico. http://www.conabio.gob.mx/ institucion/proyectos/resultados/infke002.pdf sandor, m. e., elphick, c. s., and tingley, m. w. 2022. extinction of biotic interactions due to habitat loss could accelerate the current biodiversity crisis. ecological applications, 32(6), e2608 scheinvar, l., a. gómez-pompa, and l. alonso. 1967. sistema automático de recuperación de información para el herbario nacional del instituto de biología de la u.n.a.m. anales del instituto de biología de la universidad nacional autónoma de méxico, 38, ser. bot. (1): 203-250. scheinvar, l., l. alonso, and a. gómez-pompa. 1968. proyecto piloto de recuperación automática de información del herbario nacional de la unam. anales del instituto de biología de la universidad nacional autónoma de méxico 38 ser. bot. soberón, j. 1999. linking biodiversity information sources. trends in ecology and evolution, 14(7), 291. the new york botanical garden. 2018. c. v. starr virtual herbarium. http://sweetgum.nybg.org/science/vh/. accessed on agu/9/2018. thiers, b. m. 2018. the world’s herbaria 2017: a summary report based on data from index herbariorum. william and lynda steere herbarium, the new york botanical garden. disponible en: http://sweetgum.nybg.org/science/docs/ the_worlds_herbaria_2017_5_jan_2018.pdf unam. 2016. lineamientos para sitios web institucionales de la unam. 22p. available from: https://www.visibilidadweb. unam.mx/normateca/normaunam/%20lineamientos_sitioswebinstitucionales_catic_octubre2016.pdf accessed on aug/9/2018. wen, j., s. m. ickert‐bond, m. s. appelhans, l. j. dorr, and v. a. funk. 2015. collections‐based systematics: opportunities and outlook for 2050. journal of systematics and evolution, 53(6), 477-488. wieczorek, j., d. bloom, r. guralnick, s. blum, m. doöring et al. 2012. darwin core: an evolving community-developed biodiversity data standard. plos one 7(1): e29715. doi:10.1371/journal.pone.0029715 wilson, e. o. 1985. the biodiversity crisis: a challenge to science. issues in science and technology, 2(1), 20-29. https://kebs.isolutions.iso.org/obp/ui#iso:std:iso:9241:-11:ed-2:v1:en https://www.ib.unam.mx/ib/programa-editorial/publicaciones-especiales/ https://www.ib.unam.mx/ib/programa-editorial/publicaciones-especiales/ http://doi.org/10.22201/ib.9786073073189e.2023 http://doi.org/10.22201/ib.9786073073189e.2023 https://doi.org/10.1080/14772000.2017.1355854 http://www.conabio.gob.mx/institucion/proyectos/resultados/infke002.pdf http://www.conabio.gob.mx/institucion/proyectos/resultados/infke002.pdf http://www.conabio.gob.mx/institucion/proyectos/resultados/infke002.pdf http://www.conabio.gob.mx/institucion/proyectos/resultados/infke002.pdf http://www.conabio.gob.mx/institucion/proyectos/resultados/infke002.pdf http://www.conabio.gob.mx/institucion/proyectos/resultados/infke002.pdf http://sweetgum.nybg.org/science/vh/ https://www.visibilidadweb.unam.mx/normateca/normaunam/ lineamientos_sitioswebinstitucionales_catic_octubre2016.pdf https://www.visibilidadweb.unam.mx/normateca/normaunam/ lineamientos_sitioswebinstitucionales_catic_octubre2016.pdf https://www.visibilidadweb.unam.mx/normateca/normaunam/ lineamientos_sitioswebinstitucionales_catic_octubre2016.pdf miguel murguía-romero et al. – the ibdata web system 12 appendix 1: glossary agile development methodology.—computer systems development methodology that contrasts with classical or previous methodologies in which less importance is given to documentation and more to the development of functional prototypes before meeting all the requirements. architecture of three layers.—the way in which the elements of a computer system are classified and arranged into three large groups: user interface, business rules, and database. bifenestra user interface.—two-window interface. technical solution to the “level 1 navigation” feature, which seeks to avoid the user having to navigate through multiple submenus to reach the desired functionality and from having to leave the current page. business rules.—specification of the restrictions and processes of the objective of a computer system. darwin core.—standard that specifies the nomenclature and meaning of the fields in biodiversity databases proposed by the taxonomic databases working group, now biodiversity information standards (tdwg) and that is widely used in public biodiversity databases. development methodology.—sequence of steps to build a computer system. drill-down.—analytical capacity that allows users to change immediately from a general view of the data to a more detailed one by clicking on a metric on the dashboard or report in which said view shows the data. geotax search.—type of search in ibdata that considers two sets of fields, geographic and taxonomic, in a three-level hierarchy: country, state, county, and family, genus and species. because it is a set of fields that is very common for the user to search, the interface makes it more efficient by displaying these fields so that the user simply enters the values by which they want to filter. information architecture.—the way in which the interface presents the flow of information to users, either by requests that they make or by suggestion of the system. navigation level 1.—characteristic of user interface consisting of presenting most of the options to the user in a single panel, without the need to navigate through submenus. in a technical way, in ibdata it is implemented through the bifenestra user interface, that is, a two-window interface. responsive system.—characteristic of the interface design that consists of adapting to the size of the screen to show the amount and shape of elements that allow more efficient and pleasant communication with the user. software engineering.—computer science branch that studies and proposes methods for the complex process of computer systems development. thumbnail.—small version of an image that aims to show a preview to make computing more efficient. usability test.—interface design technique to verify that the interface elements are adequately understood by users; this way, possible improvements, changes, or unnecessary elements are identified. usability.—degree to which specific users can use a system, product, or service to achieve specific objectives with effectiveness, efficiency, and satisfaction in a specific context of use. user experience design (ux design).—user interface design area that applies rules, techniques, and methods to create attractive and functional interfaces for users. user interface.—place where the user and the system interact and which considers the necessary elements depending on the type of user, as well as the social and technological moment. biodiversity informatics, 17, 2022, pp. 67-95 67 guide francophone pour la modelisation de niches ecologiques anaïs vignoles1* 1umr-5199 pacea, université de bordeaux, bâtiment b2, allée geoffroy saint-hilaire, cs 50023, 33615 pessac cedex (france) résumé. la modélisation corrélationnelle de niches écologiques (en anglais « ecological niche modeling »; enm) est un ensemble de méthodes populaire dans le champ de l’écologie de la distribution d’espèces et est employée pour une multitude d’applications. si le cadre conceptuel et méthodologique de l’enm a été largement décrit dans la littérature, il n’existe pas de synthèse exhaustive en langue française. dans cet article, j’expose les bases théoriques de l’enm à travers un historique du concept de niche écologique et ses implications pour l’étude de la distribution macro-géographique des espèces. je décris ensuite les différentes étapes d’une étude enm, en insistant tout d’abord sur l’importance de contrôler la qualité des données d’entrées. différentes préconisations concernant le choix des algorithmes, la calibration et l’évaluation des modèles ainsi que les analyses postérieures, telles que les comparaisons de niches ou le transfert à d’autres périodes/régions, sont présentées. j’insiste en particulier sur 1/ le fonctionnement de l’algorithme maxent – l’algorithme le plus usité dans la littérature actuellement – et la nécessité d’un processus de réglage de ses paramètres, 2/ l’importance du choix de l’aire de calibration m, 3/ la nécessité de prendre en compte les environnements accessibles (associés à l’aire de calibration m) dans le transfert et la comparaison des modèles, et 4/ l’importance d’évaluer et de présenter la variabilité des résultats en fonction de choix méthodologiques à différentes étapes (partitionnement des données d’occurrences, choix d’un modèle climatique, choix de l’algorithme, choix de l’aire de calibration, etc.). en conclusion, je rappelle l’importance d’ancrer toute étude employant l’enm dans un cadre théorique et méthodologique clair et explicite afin de garantir la pertinence des interprétations ultérieures. mots-clés: modélisation de niches écologiques; bonnes pratiques; cadre conceptuel; calibration et évaluation des modèles; transfert de modèles; comparaison de modèles abstract: correlational ecological niche modeling (enm) is a popular group of methods in the field of distributional ecology and is employed for a variety of applications. although the conceptual and methodological framework of enm has been widely described in the literature, there is still no exhaustive synthesis of it in the french language. in this article, theoretical bases of enm are exposed through a history of the concept of ecological niche as well as its implications for the study of species macroscale distributions. then, the different steps of enm are described, emphasizing on the importance of controlling the quality of input data. various recommendations concerning algorithm choice, model calibration and evaluation as well post-modeling analyses, such as niche comparison and transfer to other periods/regions, are presented. particular emphasis is placed on 1/ the operation of maxent – the most used algorithm in the literature today – and the need for parameter tuning prior modeling, 2/ the importance of the choice of the m calibration area, 3/ the need to take into account accessible environments (associated with the m calibration area) for model transfer and comparison, and 4/ the importance of evaluating and presenting the variability of models resulting from methodological choices at different stages (occurrence data partitioning, choice of a climate model, choice of algorithm, choice of the calibration area, etc.). to conclude, contextualizing any enm study in a clear and explicit theoretical and methodological framework is paramount to ensure the pertinence of subsequent interpretations. key words: ecological niche modeling; good practices; conceptual framework; model calibration and evaluation; model transfer; model comparison *corresponding author: anais.l.vignoles@gmail.com anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 68 la modélisation de niches écologiques (enm) désigne un ensemble de méthodes corrélatives visant à reconstituer la niche écologique d’un taxon ou d’une population. les modèles obtenus s’appuient sur des corrélations entre des occurrences connues pour ledit taxon et des variables environnementales présentées sous forme de cartes à plus ou moins haute résolution. ils sont utilisés pour une grande diversité d’applications et de problématiques, telles que la découverte de nouvelles espèces ou populations (e.g., raxworthy et al., 2003; de siqueira et al., 2009; peterson et navarro-sigüenza, 2009), la planification des politiques de conservation de la biodiversité (e.g., sohn et al., 2013; sobral-souza et al., 2021), les dynamiques spatiales d’espèces invasives (e.g., zhu et al., 2012; escobar et al., 2014; alkishe et al., 2020; nuñez-penichet et al., 2021a), les mécanismes de macroévolution des espèces (e.g., saupe et al., 2019), la paléoécologie et paléo-distribution d’espèces fossiles (e.g., myers et al., 2015; gibert, vignoles et al., 2022), la diffusion de maladies via des vecteurs animaux (e.g., sweeney et al., 2006; escobar et al., 2017; marques et al., 2020, 2021), l’impact du changement climatique sur les espèces animales ou végétales (e.g., warren et al., 2014; ashraf et al., 2017), ou encore les relations cultureenvironnement chez les chasseurs-cueilleurs du paléolithique (e.g., banks et al., 2006, 2009, 2011, 2021; vignoles et al., 2021). dans les années 2000, le nombre d’études employant l’enm a explosé (lobo et al., 2010). la démocratisation progressive de cette méthodologie a malheureusement donné suite à de nombreuses applications erronées d’un point de vue conceptuel ou insuffisamment robustes d’un point de vue méthodologique. si aujourd’hui de nombreux articles publiés en langue anglaise proposent des clarifications terminologiques (e.g., soberón et peterson, 2005; soberón et nakamura, 2009; peterson et soberón, 2012; warren, 2012; araújo et peterson, 2012) et des guides résumant les principales étapes et les bonnes pratiques associées en enm (e.g., peterson et al., 2011; feng et al., 2019a; araújo et al., 2019; sillero et barbosa, 2020; sillero et al., 2021), il n’existe aucune référence récapitulant les concepts et les précautions nécessaires à l’emploi de cette approche en langue française (voir cependant pierrat, 2011; antunes, 2015; vignoles, 2021). l’application de l’enm se divise en quatre principales étapes. la première (1) est de se positionner clairement vis-à-vis du cadre théorique de cette approche. les concepts et la terminologie employés posent en effet le cadre de l’étude; ils délimitent le type de problématique qui peut être résolu par l’enm, et donnent aux utilisateurs une référence pour le protocole de modélisation et pour correctement interpréter les modèles. la seconde étape (2) vise à recueillir les données à l’origine des modèles de niches. celles-ci sont de deux natures: tout d’abord, les données d’occurrences permettent de décrire la répartition géographique de l’espèce ou de la population considérée, par l’identification de localités dans lesquelles des spécimens ont été observés à un instant t. ensuite, les données environnementales représentent l’environnement dans lequel les spécimens ont évolué, par le biais de cartes de répartition de différentes variables (e.g., climat, topographie, etc.). ces données doivent être sélectionnées avec soin, pour éviter le phénomène de « garbage-in, garbage-out »: l’emploi de données trop éloignées de la réalité à ce stade se répercutera nécessairement sur les modèles de niches, qui seront alors moins pertinents pour répondre aux problématiques de départ. la troisième étape (3) est la création de modèles par la mise en relation de ces deux types de données via des méthodes statistiques ou la modélisation de surfaces de réponse. les algorithmes pouvant être utilisés sont nombreux et la pertinence de leur utilisation dépend du type de données employées (e.g., données d’occurrences de type présence-seule, présence-absence, etc.) et des objectifs de l’étude (e.g., modéliser une distribution potentielle, visualiser des volumes dans un espace environnemental, etc.). enfin, la dernière étape (4) consiste à projeter le modèle dans un nouveau contexte environnemental – comme une autre région ou une autre période. à ce stade, des modèles peuvent être comparés dans l’objectif d’évaluer leur degré de similarité ou de recouvrement. cette dernière étape est particulièrement intéressante pour prédire l’impact de changements environnementaux sur une population (e.g., ashraf et al., 2017) ou pour estimer la capacité dispersive d’une espèce potentiellement invasive dans une région donnée (e.g., nuñezpenichet et al., 2021a). dans cet article, je propose une synthèse des concepts et de la terminologie adéquate à l’enm, ainsi que les principales précautions à prendre en compte lors de la mise en œuvre de cette approche à un contexte d’application, quel qu’il soit – depuis le choix des données jusqu’à la création, le transfert et la comparaison des modèles. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 69 concepts et terminologie niche grinellienne et niche eltonienne la première définition du concept de niche est communément attribuée au zoologue américain j. grinnell, dans un article portant sur « les relations de niche du moqueur de californie » (grinnell, 1917; chase et leibold, 2003). dans cette publication, l’auteur met en relation la distribution géographique particulière du moqueur de californie (toxostoma redivivum), avec sa « niche », qu’il définit comme l’expression géographique de ses prérequis environnementaux (en anglais, « requirements »), telles que ses tolérances physiologiques ou ses habitudes alimentaires. dans les années 1930, une vision de la « niche » assez différente est proposée par l’écologue anglais c. elton, considéré comme l’un des pionniers de l’écologie des populations et communautés. il utilise ce terme pour désigner le rôle fonctionnel d’un animal au sein de la chaîne trophique (elton, 1927). cette définition se concentre alors sur l’impact d’un organisme sur son environnement, par exemple par le biais de la consommation de ressources ou de la compétition avec d’autres organismes (« food and enemies »). cette double signification du mot « niche » – « niche as requirements versus niche as impacts » – est probablement à l’origine d’une confusion autour de ce concept, qui provoque d’importantes controverses épistémologiques dans la seconde moitié du xxème siècle, allant jusqu’à son quasi-abandon dans la recherche en écologie (e.g., whittaker et al., 1973; chase et leibold, 2003; pocheville, 2015). il faut attendre les années 2000 pour qu’un cadre théorique réconciliant ces deux aspects du concept soit proposé (chase et leibold, 2003; peterson et al., 2011). la définition d’une niche proposée par chase et leibold (2003) s’appuie à la fois sur l’intégration des prérequis environnementaux de l’espèce, représentés par les variables non interactives (e.g., le climat, la topographie...), et sur les variables interactives, c’est-à-dire les interactions de cette espèce avec l’environnement (e.g., la consommation de ressources). cette dernière composante est en effet fondamentale puisqu’à une échelle locale, les espèces consomment des ressources et interagissent entre elles. un tel modèle doit donc représenter un espace environnemental comportant une composante statique (prérequis) et une composante dynamique (interactions). cela signifie que l’environnement change constamment suivant l’impact de l’espèce sur lui. or, l’échelle à laquelle se développe l’étude des distributions biogéographiques par l’enm ne permet pas de s’ancrer dans un tel cadre théorique (araújo et guisan, 2006; peterson et al., 2011). d’une part, ces modèles de niches sont essentiellement statiques et explorent des corrélations à un instant t, ce qui rend les mécanismes de rétroaction difficiles à évaluer (araújo et guisan, 2006). d’autre part, ces interactions sont difficiles à mesurer empiriquement ou à détecter pour de grandes échelles de temps et d’espace, puisqu’elles se manifestent de façon très locale avec parfois d’importants changements sur de faibles distances (araújo et guisan, 2006; soberón, 2007; peterson et al., 2011; araújo et rozenfeld, 2013). il est possible de contourner ce problème en définissant deux grandes classes de niches en fonction du type de variables utilisées pour les modéliser: la niche grinnellienne, définie par des variables noninteractives, et la niche eltonienne, définie par les variables interactives (soberón, 2007). ces deux classes de niches ne se manifestent pas à la même échelle géographique, étant donné la résolution à laquelle les deux types de variables sont mesurées. l’estimation des variables non interactives (ou scénopoétiques) concerne des échelles régionales à macro-régionales, voire mondiales, tandis que les variables interactives sont mesurées à une échelle locale (araújo et guisan, 2006; soberón, 2007; peterson et al., 2011). dans ce contexte, j. soberón (2007) propose l’hypothèse du bruit eltonien (en anglais: « eltonian noise hypothesis »). celle-ci s’appuie sur l’observation de l’effet généralement limité ou peu significatif des interactions biotiques sur la distribution d’une espèce à de grandes échelles spatiales, et propose de ne pas les prendre en compte dans la reconstitution de niches dans le cadre de questionnements sur la distribution macrogéographique d’un taxon ou population. c’est ce point de vue qui est adopté dans le cadre théorique que je présente ici, puisque celui-ci s’intéresse à des phénomènes d’ordre macro-géographique (en référence à peterson et al., 2011). le concept hutchinsonien de niche écologique l’une des principales avancées dans la définition de niche écologique est le fruit du travail de l’écologue anglais g. evelyn hutchinson qui a proposé dans les années 1950 une approche plus mathématique du concept (hutchinson, 1957). il définit une niche anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 70 comme « un hypervolume de n-dimensions […], au sein duquel chaque point correspond à un état de l’environnement qui permettrait à l’espèce [...] d’exister indéfiniment. » (ibid. p. 416). elle s’exprime dans l’espace n-dimensionnel des niches n, dont chaque dimension correspond à une variable environnementale. cette définition correspond à la niche fondamentale (nf), qui, en théorie, définit les propriétés écologiques intrinsèques de l’espèce considérée en l’absence d’interactions, à partir de variables scénopoétiques (également appelées abiotiques; figure 1).2 ces dernières sont typiquement les variables climatiques ou géographiques (peterson et al., 2011). g. e. hutchinson précise également que cet hypervolume ne serait pas une façon de « binariser » l’espace environnemental entre les points où l’espèce a la même probabilité d’exister, qui matérialiseraient nf (valeur sélective ou fitness = 1), et les points où l’espèce n’a aucune probabilité d’exister, et qui sont donc en dehors de nf (fitness = 0). il faudrait plutôt penser la niche comme un gradient de la valeur sélective d’une espèce au sein de n. autrement dit, la population considérée a une plus grande probabilité de persistance dans certaines zones de sa niche que dans d’autres (hutchinson, 1957). l’hypothèse la plus communément admise de nos jours est que la niche fondamentale prend la 2 source photo: www.butterfliesandmoths.org. forme d’un objet convexe, comme un polyèdre ou un ellipsoïde (e.g., van aelst et rousseeuw, 2009; escobar et al., 2014, 2017; qiao et al., 2016; jiménez et al., 2019; soberón et peterson, 2020; nuñezpenichet et al., 2021b; figure 1a). cette supposition repose sur l’idée que le bord de nf doit être de forme convexe, car la tolérance physiologique d’un individu à une variable environnementale est considérée comme étant unimodale par nature (angilletta, 2009; drake, 2015; jiménez et al., 2019; soberón et peterson, 2020; figure 1b). cependant, certains travaux remettent en partie en cause cette hypothèse, notamment dans le cas d’espèces annuelles vivant dans un environnement fortement marqué par la saisonnalité ou dans le cas d’espèces migratrices (nakazawa et al., 2004; soberón et peterson, 2020; ingenloff 2020). en effet, un organisme peut « changer » de niche, par exemple entre sa forme juvénile et sa forme adulte (grubb, 1977; soberón et arroyo-peña, 2017) ou dans le cadre d’un cycle saisonnier/migratoire (e.g., soberón et peterson, 2020; ingenloff 2020). il faut donc considérer un modèle de niche comme un modèle statique valable à un instant t (hutchinson, 1957, p. 417). cela implique un certain nombre de précautions pour le choix des données à l’origine de la modélisation. en s’appuyant principalement sur les travaux de v. volterra (1926) et g.f. gause (1934), g. e. figure 1. a. modèle de la niche du phalène urania fulgens (occurrence 530489 observée le 23 mars 20111) sous la forme d’un ellipsoïde dans un espace environnemental tri-dimensionnel composé de variables bioclimatiques significativement corrélées à sa distribution géographique (nuñez-penichet et al., 2021b). b. représentation schématique de la tolérance physiologique unimodale d’un individu à une variable environnementale. http://www.butterfliesandmoths.org anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 71 hutchinson inclut dans son modèle une composante biotique qui rend compte de la possibilité de compétition entre deux espèces pour les mêmes portions de l’espace environnemental à une localité donnée. il formule donc le principe de volterragause, selon lequel l’occurrence de deux espèces au sein d’une même localité signifie qu’elles doivent forcément occuper des niches écologiques différentes, sans quoi il y a compétition et, in fine, disparition de l’une des deux espèces à la localité considérée. ce principe implique la définition de la niche réalisée (nr), qui, hypothétiquement, est une réduction de nf par les interactions de compétition avec d’autres espèces (hutchinson, 1957). si l’on considère l’espace environnemental (e), il est possible d’identifier des régions correspondant aux conditions environnementales présente dans une ou plusieurs localités de l’espace géographique (g) 3, notées η(gi). (figure 2; hutchinson, 1957; peterson et soberón, 2012). à l’inverse, chaque localité de l’espace géographique (g) ne correspond qu’à un point de l’espace environnemental (e). ces environnements présents dans g sont notés η-1(e). ce principe est appelé la dualité de hutchinson, et 3 bien que g. e. hutchinson utilise la notation b et n pour désigner le biotope et l’espace des niches, j’emploirai plutôt les abréviations g pour espace géographique et e pour espace environnemental dans la suite de ce texte, en référence à un article récent de clarification théorique (peterson et soberón, 2012). a des conséquences théoriques et pratiques cruciales (colwell et rangel, 2009). en effet, il implique que la niche s’exprime à la fois dans l’espace environnemental et dans l’espace géographique, et donc que ces deux espaces sont intimement liés. cette relation n’est toutefois pas réciproque. de façon logique, à une localité de (g) ne peut correspondre qu’une seule combinaison de variables environnementales dans (e), puisqu‘elle est unique dans g. en revanche, une combinaison de variables environnementales dans (e) peut correspondre à plusieurs localités de (g); autrement dit, les mêmes conditions environnementales peuvent être présentes à différentes localités géographiques. la conséquence de cette relation non-réciproque est qu’un hypervolume continu dans n peut correspondre à plusieurs zones discontinues dans g, et viceversa (figure 2). c’est pourquoi il est primordial de prendre en compte ces deux espaces à la fois, g et e, lorsque l’on cherche à aborder la question des distributions biogéographiques en relation avec les niches écologiques (peterson et al., 2011). la relation entre les deux types de niches de g. e. hutchinson peut s’écrire en termes mathématiques selon l’inégalité suivante (soberón et arroyo-peña, 2017): figure 2. illustration de la dualité de hutchinson dans le cas d’une espèce virtuelle et d’un espace environnemental simplifié, dont les deux variables sont la température moyenne annuelle et la précipitation moyenne annuelle. dans l’espace environnemental, le nuage de points représente les combinaisons de températures et précipitations disponibles en france de nos jours (19792013), d’après les simulations issues du modèle chelsa (résolution spatiale: 10 arc min.; karger et al., 2017). le rectangle vert représente la niche de l’espèce virtuelle. les localités correspondant à ces conditions favorables sont représentées dans l’espace géographiques sous la forme de pixels verts. il est intéressant de noter ici que certaines conditions environnementales incluses dans la niche n’existent pas en france actuellement (zones grises contenue dans le rectangle vert). anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 72 cette inégalité prédit que les combinaisons de variables environnementales correspondant aux localités où l’espèce est présente (nr) sont contenues dans l’hypervolume de nf, ce qui se vérifie, ou pas, empiriquement (voir discussion soberón et arroyopeña, 2017). or, il est possible que des combinaisons de variables présentes dans nf n’existent pas dans une zone et à un instant donné (hutchinson, 1957; pulliam, 2000; jackson et overpeck, 2000; soberón et peterson, 2011; figure 2). c’est pourquoi jackson et overpeck (2000) ont formalisé l’idée de « niche potentielle », aussi nommée « niche disponible » (green, 1971) ou « niche fondamentale existante », annotée n* f (peterson et al., 2011, peterson et soberón, 2012). elle se définit par l’intersection entre nf et l’espace environnemental disponible à l’espèce considérée, soit en termes mathématiques: en d’autres termes, n* f correspond à l’ensemble des combinaisons de variables qui existent dans la région et la période considérée et qui coïncident avec les exigences écologiques de l’espèce. ce troisième type de niche peut être ajouté à l’inégalité de hutchinson de la façon suivante: cette inégalité établit un lien explicite entre niche écologique et répartition géographique puisque, par définition, nf est avant tout définie dans l’espace environnemental, alors que n* f et nr sont définies à partir des occurrences géographiques de l’espèce considérée. il convient maintenant de décrire plus précisément cette relation entre niches écologiques et répartitions géographiques. relation entre niche et distribution considérons le diagramme bam (biotique, abiotique, mobilité) proposé par soberón et peterson (2005, 2012), qui délimite un cadre théorique pour réfléchir aux facteurs influençant la distribution géographique d’un taxon ou une population (figure 3). la projection de nf dans l’espace géographique g, c’est-à-dire les localités géographiques présentant des conditions environnementales incluses dans nf, identifie les endroits où le taxon pourrait rencontrer des conditions scénopoétiques favorables (a). la distribution géographique du taxon peut également être réduite par les interactions biotiques (b). cependant, si nous nous plaçons dans l’hypothèse du bruit eltonien, ce facteur est peu limitant. enfin, le dernier facteur est l’accessibilité potentielle du taxon à ces conditions favorables sur une période pertinente (m; pulliam, 2000). ces dernières peuvent en effet être trop éloignées de l’aire de répartition de la population source ou séparées de celle-ci par une barrière géographique infranchissable, comme un océan ou une chaîne de montagne. l’intersection de a et b est l’aire de distribution potentielle (gp), qui est l’expression géographique de n* f. l’intersection de gp avec m définit l’aire de distribution occupée (go), qui est l’expression géographique de nr. enfin, la zone de gp, qui n’est pas occupée, est appelée l’aire de distribution envahissable (gi), c’est-à-dire que les conditions sont favorables et que le taxon n’y pas encore accès, mais qu’il pourrait théoriquement y perdurer. dans ce cadre, il est possible d’établir au moins quatre configurations du diagramme bam (peterson et al., 2011; saupe et al., 2012; figure 4): i) la configuration classique, dans laquelle a et m se recoupent partiellement; ii) la configuration de hutchinson, dans laquelle la présence de conditions figure 3. diagramme bam illustrant les facteurs influençant la distribution macro-géographique d’une espèce dans le cas de l’hypothèse du bruit eltonien (d’après soberón et peterson, 2005, 2012, modifié). les cercles représentent les différents facteurs et les points noirs représentent la distribution géographique de l’espèce. g: espace géographique; a: variables scénopoétiques; b: interactions biotiques; m: zones géographiques accessibles à l’espèce; gp: aire de distribution potentielle; go: aire de distribution occupée; gi: aire de distribution potentiellement envahissable. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 73 favorables constitue le principal facteur limitant à l’établissement de l’espèce dans un endroit et à un moment donnés, soit m englobe a; iii) la configuration de wallace, dans laquelle le principal facteur limitant est l’accès aux conditions favorables, soit a englobe m; et enfin, iv) la configuration de superposition, dans laquelle a et m se superposent parfaitement. le choix d’une configuration parmi ces quatre-là peut avoir un impact significatif sur le résultat de la modélisation de la niche (saupe et al., 2012). en effet, une étude a montré que les modèles sont plus performants dans la configuration classique et dans celle de hutchinson (ibid.); il conviendra donc de se placer dans l’une ou l’autre de ces hypothèses. il faut noter ici l’importance que prend m comme facteur contraignant de la distribution géographique d’un taxon ou d’une population, en plus de la présence de conditions favorables. il sera donc capital d’en donner une estimation fiable pour la bonne conduite d’une modélisation de niche. elle doit correspondre à une réalité biogéographique plutôt qu’à des frontières administratives; ses limites dépendent des capacités dispersives du taxon et sont par exemple matérialisées par des barrières géographiques naturelles, telles que la présence d’une chaîne de montagne, d’un océan ou d’une rivière (pulliam, 2000; barve et al., 2011; machado-stredel et al., 2021). que modélise-t-on en enm ? la modélisation de niches écologiques grinnelliennes s’appuie sur deux grands types de données: d’une part, les localités géographiques où le taxon considéré a été observé, et d’autre part, les conditions environnementales présentes dans l’aire géographique accessible au taxon. à partir de ces données, l’objectif est de déterminer, dans chaque pixel de la région étudiée, la probabilité d’occurrence en fonction des conditions environnementales présentes dans ce pixel. en d’autres termes, nous créons un modèle pour la fonction qui décrit la relation entre les variables environnementales et les occurrences du taxon (i.e., la niche; peterson et al., 2011). pour cela, nous employons un ou des algorithmes prédictifs qui établissent des règles décrivant la relation entre différentes variables ou paramètres, puis qui utilisent ces règles pour reconstruire la niche. afin d’interpréter correctement les modèles résultants, il est primordial de déterminer le ou les types de niches – nf, n* f ou nr – qui peuvent potentiellement correspondre à ces données de sortie. la nature des données utilisées (données d’occurrences et variables environnementales) réduit d’ores et déjà la possibilité de modéliser tous les types de niches présentés précédemment. en effet, ces données sont avant tout géographiques. or, nous avons vu que la niche fondamentale est définie en premier lieu dans l’espace environnemental, et que dans l’espace géographique, elle peut être réduite par au moins deux types de facteurs: les interactions avec d’autres espèces et l’aire géographique accessible au taxon. cela signifie que la niche fondamentale ne peut être estimée à partir de données géographiques seulement, étant donné que la distribution géographique du taxon ne correspond pas nécessairement à a (jiménez et al., 2019). elle doit être avant tout définie à partir de données expérimentales visant à déterminer les conditions limites tolérées par les individus (soberón et arroyopeña, 2017). la niche modélisée dans le cadre de l’enm se situe plus probablement quelque part entre la niche fondamentale existante et la niche réalisée en fonction de l’algorithme et des données utilisés. la nature de la démarche conditionne également l’interprétation des cartes de répartition issues de ces modélisations. la modélisation de distribution d’espèces (species distribution modeling, abrév. sdm) peut parfois employer les mêmes jeux de données et algorithmes qu’en enm (peterson et figure 4. quatre scénarios théoriques du diagramme bam (d’après saupe et al., 2012, modifié). le cercle vert représente les conditions scénopoétiques favorables (a); le cercle aux hachures ondulées représente l’aire accessible à l’espèce (m). a. configuration classique. b. configuration de hutchinson. c. configuration de wallace. d. configuration de superposition. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 74 soberón, 2012). toutefois, le cadre théorique et les concepts sous-jacent à la sdm sont différents et n’impliquent pas les mêmes axiomes – en particulier dans l’interprétation des cartes de répartition. en sdm, les cartes de sortie sont interprétées comme la distribution potentielle d’une espèce, tandis qu’en enm, elles correspondent à des projections géographiques d’une niche définie dans l’espace environnemental à partir de la répartition géographique connue d’une espèce (ibid.; peterson et al., 2011; sillero, 2011; araújo et peterson, 2012; warren, 2012; sillero et al., 2021). cette différence interprétative a des conséquences sur le choix des données d’occurrences et des algorithmes; ainsi, les études de sdm utilisent généralement des données de présence-absence, tandis qu’en enm, les données de présence-seule et de présence-background sont préférées (ibid.; cf. 3.2.). la qualité d’une étude employant l’enm repose donc avant tout sur l’exposé clair et explicite de ses fondements conceptuels (peterson et soberón, 2012). ceux-ci permettent de guider le choix des données à l’origine des modèles ainsi que la conception des analyses, pour que les modèles finaux se rapprochent au maximum des phénomènes que l’on souhaite étudier, qu’ils soient écologiques ou culturels. dans la suite de cet article, je propose un certain nombre de recommandations quant au choix des données et des algorithmes en enm. je me focaliserai donc principalement sur les méthodes de présence-seule ou présence-background, et plus particulièrement sur maxent, l’algorithme le plus populaire en enm. choisir les donnees a l’origine des modeles de niches ecologiques: quelques precautions methodologiques données environnementales les données environnementales utilisées en enm sont des cartes raster (i.e., une carte de données spatiales organisées sous la forme de pixels. à chaque pixel est attribué un set de coordonnées spatiales et une valeur attributaire) de différentes variables pertinentes pour décrire l’environnement du taxon ou de la population considéré. le choix des variables environnementales comporte trois aspects essentiels. en premier lieu, il est important de sélectionner des variables abiotiques susceptibles d’influencer la répartition géographique de l’espèce ou population considérée (peterson et al., 2011), une nécessité qui découle du diagramme bam lorsque l’on se place dans l’hypothèse du bruit eltonien (figure 3). en général, il s’agit de variables climatiques et géographiques (pearson et dawson, 2003), mais d’autres types peuvent être employés en fonction de l’espèce étudiée. par exemple, il est particulièrement pertinent d’inclure des variables décrivant les propriétés du sol pour modéliser la niche écologique de végétaux (e.g., hengl et al., 2017; zuquim et al., 2020), ou encore les propriétés de l’eau dans le cas d’espèces aquatiques (e.g., domisch et al., 2015; sbrocco et barber, 2013). en revanche, l’élévation n’est généralement pas une variable pertinente en enm en raison de son importante corrélation avec la température. l’utilisation de variables corrélées n’a pas nécessairement de conséquences majeures sur les modèles (en tout cas, avec l’algorithme maxent; feng et al., 2019b). cependant, la corrélation entre température et élévation n’est pas constante et diffère dans l’espace (notamment par rapport à la latitude) et dans le temps (peterson et al., 2011, p. 85). cela est problématique lorsque l’on cherche à transférer le modèle à une autre région ou une autre période. il est plutôt recommandé d’utiliser des indicateurs topographiques calculés à partir de l’élévation, mais qui sont indépendants des variables climatiques (e.g., amatulli et al., 2018, 2019). dans tous les cas, la plupart de ces variables correspondent elles-mêmes à des modèles, étant donné qu’il n’est pas possible de les mesurer en continu sur d’aussi grandes surfaces (figure 5), et ce particulièrement concernant des variables paléoclimatiques qui reconstituent le climat et l’environnement pour le passé (e.g., valdes et al., 2017; boucher et al., 2020), pour lequel la densité spatio-temporelle des proxys climatiques est très faible. en conséquence, les variables basées sur des modèles peuvent comporter un certain nombre d’incertitudes dont il faut être conscient lors de leur emploi en enm (varela et al., 2015). il est alors important de consulter leurs métadonnées afin de vérifier la fiabilité globale de ces modèles, mais également leur fiabilité à l’échelle de la région d’intérêt. par exemple, pour des jeux de variables climatiques mondiales telles que worldclim (fick et hijmans, 2017), certaines zones sont plus fiables que d’autres en raison de l’inégale densité de répartition des stations météorologiques (figure 5). dans le cas de variables paléoclimatiques, certaines périodes sont également mieux documentées que d’autres, ce qui permet une meilleure reconstitution du climat (e.g., anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 75 le dernier maximum glaciaire en europe; varela et al., 2015). ensuite, il est important de prendre en compte la dimensionnalité de l’environnement modélisé, c’est-à-dire le nombre de variables employées. il est préconisé d’utiliser un nombre restreint de variables non ou peu corrélées. l’utilisation d’une grande quantité de variables rend le modèle plus complexe et peut entraîner une trop grande adéquation entre le modèle de niche et les données d’occurrences – overfitting en anglais (peterson et al., 2011, p. 87) – ce qui peut masquer certains aspects écologiques de la niche modélisée. la trop grande complexité d’un modèle de niche diminue la qualité des modèles (e.g., warren et seifert, 2011), et donc les possibilités d’effectuer des analyses postérieures, tel que le transfert à une autre région ou une autre période (peterson et nakazawa, 2007). il est en ce sens primordial que le nombre de variables environnementales soit inférieur au nombre d’occurrences (sillero et al., 2021). en outre, l’utilisation de variables hautement corrélées entraîne l’impossibilité d’utiliser certains types d’algorithmes (guisan et zimmerman, 2000). enfin techniquement, utiliser un trop grand nombre de variables augmente automatiquement le temps de calcul et de calibration des modèles, une contrainte purement matérielle mais non négligeable dans la conduite d’un projet inscrit dans un temps limité. il existe de nombreuses propositions afin de sélectionner les meilleures variables pour un modèle de niche donné, en fonction des objectifs de l’étude (voir cobos et al., 2019a pour une revue de la littérature). figure 5. localisation des stations météorologiques utilisées pour produire les variables climatiques par interpolation, exemple de la température moyenne et de la radiation solaire du jeu de données worldclim2 (fick et hijmans, 2017). l’inégale densité de stations météorologiques implique que la reconstitution des variables sera plus précise dans certaines régions que d’autres. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 76 enfin, le dernier élément à prendre en compte est la résolution spatiale et l’étendue chronologique des variables employées. cellesci doivent impérativement correspondre à la résolution des données d’occurrences, au risque de produire un modèle de niche incohérent d’un point de vue spatio-temporel (peterson et al., 2011, p. 91; sillero et barbosa, 2020). si les cartes de données environnementales présentent une résolution spatiale trop grossière (i.e., des pixels trop larges) par rapport à l’échelle à laquelle se manifeste le phénomène étudié, il sera nécessaire de réaliser une descente d’échelle statistique afin d’augmenter la résolution. cette étape est assez délicate, car simplement subdiviser les pixels afin d’obtenir un maillage plus fin n’augmente pas la résolution intrinsèque de l’information donnée par la variable en question (sillero et barbosa, 2020). il est nécessaire pour cela d’avoir « accès à la relation d’interpolation utilisée pour […] estimer [ladite variable] et un modèle d’élévation à la résolution adéquate » (ibid., p. 5). plusieurs méthodes existent, des plus simples – comme l’interpolation avec le plus proche voisin – aux plus complexes – comme la descente d’échelle par modélisation gam (generalized additive model; e.g., vrac et al., 2007; antunes, 2015; latombe et al., 2018) ou la méthode de correction delta (e.g., beyer et al., 2020). d’autre part, la correspondance chronologique entre les données d’occurrences et les données environnementales est essentielle, au risque d’associer les mauvaises combinaisons de valeurs aux occurrences recensées. par exemple, dans le cas d’espèces migratrices, l’emploi d’un jeu de données trop imprécis temporellement peut conduire à des modèles peu informatifs sur leur comportement écologique (ingenloff, 2020; ingenloff et peterson, 2021). plusieurs auteurs ont souligné l’impact du choix des variables environnementales sur les modèles de niches sous-jacents (e.g., peterson et nakazawa, 2007; diniz-filho et al., 2009; varela et al., 2015; cobos et al., 2019a, b; vignoles, 2021). en effet, en ce qui concerne les variables climatiques, de nombreux modèles du climat régional ou global existent, plus ou moins complexes, ne se basant pas sur les mêmes types de calculs ni les mêmes processus influençant le climat (e.g., modèles de circulation globale; kageyama et al., 2005; singarayer et valdes, 2010; valdes et al., 2017; boucher et al., 2020; modèles de système terrestre de complexité intermédiaire; claussen et al., 2002; goosse et al., 2010; roche et al., 2014). chaque modèle fournit donc des jeux de simulations plus ou moins différents, influençant de fait les modèles de niches sur lesquels ils seront basés (varela et al., 2015). ce facteur de variabilité est donc important à prendre en compte, en particulier lorsque l’étude se propose d’évaluer l’impact du changement climatique sur la répartition géographique des aires environnementales en adéquation avec les nécessités écologiques d’un taxon ou d’une population (e.g., diniz-filho et al., 2009; peterson et al., 2018). il est donc préconisé de toujours proposer une évaluation de la variance des modèles en fonction des différentes sources d’incertitudes (e.g., diniz-filho et al., 2009; cobos et al., 2019b); il est également intéressant de présenter la répartition spatiale de la variance ou de l’écart-type, afin d’identifier les régions dans lesquelles la variabilité est plus forte ou, au contraire, plus faible (e.g., peterson et al., 2018; cobos et al., 2019b; warren et al., 2021a). données d’occurrences les données d’occurrences permettent de rendre compte de la répartition géographique de l’espèce ou de la population étudiée. elles se caractérisent par l’observation de la présence ou de l’absence d’un taxon dans une localité géographique donnée et à un moment donné. cette observation est soumise à plusieurs types de facteurs dont il faut être conscient lorsque l’on tente d’estimer la répartition géographique d’une espèce (peterson et al., 2011, p. 63). les facteurs qui intéressent les études d’enm sont les facteurs biologiques (figure 6), qui concernent donc la valeur sélective de l’espèce ou de la population étudiée – en d’autres termes, il s’agit de la nf, dont la projection géographique est a dans le diagramme bam modifié (figure 3). cette répartition géographique est également conditionnée par l’accessibilité de l’espèce ou population à a (m). enfin, le dernier type de facteurs est plus problématique, car il introduit des biais qui ne dépendent pas de l’écologie de l’espèce considérée, mais plutôt du fonctionnement de la recherche (figure 6). il s’agit par exemple de la différence de détectabilité entre espèces (e.g., royle et dorazio, 2006) – certaines espèces sont tout simplement plus faciles à détecter que d’autres – ou encore la disparité dans les efforts de collecte de anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 77 certains taxons ou dans certaines régions (peterson et al., 2011, p. 70-71). un biais spatial courant est par exemple la concentration d’occurrences le long de routes, car il est statistiquement plus fréquent d’y recenser une occurrence que dans des zones moins fréquentées (kadmond et al., 2004). un autre biais possible est la qualité du recensement de l’occurrence: l’identification d’un taxon peut comporter des erreurs qui dépendent de l’expérience du collecteur ou de l’état d’avancement de la recherche (localité recensée avant la découverte dudit taxon). dans le cas de données paléontologiques ou archéologiques, un biais important est également la conservation différentielle des vestiges liée aux conditions d’enfouissement et de fossilisation. celles-ci conduisent à un biais d’échantillonnage drastique s’ajoutant aux autres biais susmentionnés, car seule une très faible minorité de vestiges rencontrent des conditions suffisantes pour se conserver dans le sol. ce biais est d’autant plus fort que l’étendue spatiotemporelle du taxon ou population étudié est faible: plus un taxon a une répartition géographique et une durée de vie limitées, moins il a de probabilité d’être échantillonné (signor et lipps, 1982). ces quelques exemples montrent que les données d’occurrences ne peuvent être interprétées simplement comme la documentation de l’absence et présence d’un taxon. elles résultent en réalité d’une complexe intrication de facteurs écologiques (qui nous intéressent) et de facteurs extérieurs (dont il faut s’affranchir au maximum; figure 6). ces derniers auront pour conséquence d’introduire des biais spatiaux dans l’estimation de la répartition géographique de l’espèce, qui seront plus ou moins limitant pour la suite de l’étude. ceux-ci ne se répercutent pas nécessairement sur l’échantillonnage des environnements occupés par le taxon étudié (peterson et al., 2011). la présence plus ou moins figure 6. événements probabilistes menant à l’observation d’une présence ou d’une absence, en incluant des résultats erronés. chaque barre représente un choix, les cercles blancs représentent « oui » et les cercles gris représentent « non ». l’observation d’une présence ne requiert pas seulement que l’espèce soit présente dans une localité en fonction des trois processus biologiques représentés dans les trois premières colonnes, mais également que la localité ait été visitée par des observateurs et que cette observation ait été réalisée correctement (ex. pas d’erreur d’identification ou erreur typographique). plusieurs opportunités existent menant à des erreurs. dans cet exemple sont soulignées une fausse absence due à une mauvaise observation, deux absences liées à deux causes biologiques radicalement différentes, et une fausse présence résultant d’une mauvaise observation ou enregistrement (figure et légende issues de peterson et al., 2011, fig. 5.1, modifiées). l’analyse probabiliste de cet arbre par les auteurs de la figure conclut que la part de présence dans l’estimation de la répartition géographique d’une espèce est plus fiable que la part d’absence, puisque la probabilité d’identifier une fausse présence est moins importante que celle d’identifier une fausse absence. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 78 importante de ces biais conditionne également le type d’algorithme qui sera employé par la suite. en particulier, l’addition de biais d’échantillonnage successifs aux facteurs écologiques conduit généralement à une estimation plus fiable de la part de présence (avec moins de probabilités d’aboutir à une fausse présence) que de la part d’absence (plus de probabilité d’aboutir à une fausse absence; ibid.; figure 6). les jeux d’occurrences marqués par de tels biais devront donc être employés avec des algorithmes adaptés, dits de présence-seule ou de présence-background. en enm, les données d’absence sont en effet rarement employées, car les méthodes de présence-absence ont plutôt pour objectif de « distinguer les conditions environnementales entre habitats occupés et non occupés, résultant en la probabilité de rencontrer l’espèce à chaque localité » (sillero et al., 2021, p. 4; sillero, 2011). cela découle de la moins bonne fiabilité des données d’absence par rapport aux données de présence: une donnée d’absence peut en effet correspondre à une localité aux environnements favorables, mais inaccessibles à l’espèce (peterson et al., 2011, p. 76). les modèles résultants correspondent de ce fait à des cartes de probabilité de présence – un objectif de la sdm plutôt que de l’enm. (ibid.; sillero et al., 2021). en enm, il est préférable de n’employer que des modèles utilisant des données de présence-seule ou de présence-background. en outre, les biais d’échantillonnage peuvent se répercuter sur l’évaluation du modèle. en effet, la surreprésentation de l’espèce dans une zone va conduire à ce que les occurrences servant à tester le modèle soient géographiquement et/ou environnementalement trop proches de celles servant à le calibrer. ces effets spatiaux et/ou environnementaux vont donc conduire à de moins bonnes prédictions. pour limiter ces biais, le procédé de raréfaction spatiale ou filtration spatiale (en anglais, spatial thinning ou spatial filtering) est le plus communément utilisé et recommandé (e.g., anderson et gonzalez, 2011; boria et al., 2014). il s’agit d’éliminer les occurrences trop proches en fonction d’une distance choisie par l’utilisateur. plusieurs méthodes de raréfaction spatiale existent (e.g., aiello-lammens et al., 2015). toutefois, ce procédé n’est pas toujours suffisant pour s’affranchir de biais d’échantillonnage, en particulier si la région d’étude est très hétérogène d’un point de vue environnemental (varela et al., 2014). il est en effet important que cette raréfaction des jeux de données d’occurrences ne conduise pas à l’élimination de conditions environnementales pertinentes pour la création du modèle, en raison de leur proximité géographique avec d’autres points. il est alors conseillé de recourir à un filtre environnemental en éliminant les points présentant des combinaisons de valeurs trop proches (ibid.). enfin, la combinaison de données issues de différentes sources (e.g., programmes citoyens de science, inventaires de spécimens conservés dans les musées, collecte systématique dans le cadre d’une étude de terrain) peut se révéler problématique, car chaque source est affectée de biais différents (e.g., fletcher et al., 2019). il peut alors s’avérer difficile de limiter ces derniers dans un jeu de données unifié, puisque chaque source n’aura pas les mêmes motifs ni les mêmes sources de biais. plusieurs protocoles existent toutefois pour tenter d’intégrer différentes sources au sein d’une même étude de façon plus pertinente, par exemple en accordant plus de poids à certaines sources par rapport à d’autres ou encore, en comparant des modèles basés sur différentes sources (ibid.). une donnée d’occurrence est caractérisée par trois informations: sa localisation géographique, son attribution taxonomique et sa temporalité. selon les sources employées et la qualité intrinsèque des données d’occurrences, ces informations peuvent comporter plus ou moins d’erreurs ou être plus ou moins précises. il est donc particulièrement important d’évaluer la fiabilité de ces trois informations pour chaque occurrence, afin d’uniformiser le jeu de données, de s’assurer que chaque type d’information possède le même niveau de précision et, le cas échéant, de supprimer les occurrences dont la qualité est trop faible pour l’étude concernée (e.g., sillero et al., 2021). par exemple, la localisation de l’occurrence peut être plus ou moins précise selon la qualité et la source de l’information: elle peut se résumer à une indication vague, comme le nom de la commune ou du pays de collecte, ou à l’inverse être associée à une position géoréférencée par gps. dans ce cas, il est par exemple conseillé d’associer à chaque occurrence une notion d’incertitude (wieczorek et al., 2004; chapman et al., 2020), puis de ne garder que les occurrences se situant en dessous d’un seuil prédéfini (que l’on peut fixer par exemple à la résolution des données environnementales). il est également fréquent que les coordonnées de localisation comportent des erreurs, comme une inversion de la latitude/longitude, ou anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 79 encore une incohérence écologique de la localisation (cas d’occurrences situées dans un océan pour une espèce terrestre ou occurrence isolée très éloignée du reste de la répartition géographique; sillero et al., 2021; figure 7).4 ces occurrences doivent donc systématiquement être vérifiées, et écartées si les informations erronées ne peuvent être corrigées. le même type de problème peut apparaître pour les informations de temporalité et d’attribution taxonomique, et nécessiteront également un protocole permettant d’estimer l’incertitude des informations pour sélectionner les données les plus adaptées aux questionnements et à la résolution de l’étude. à l’issue de cette étape, il est possible de se rendre compte que le taxon étudié ne se prête pas à la problématique de départ. par exemple, étudier l’évolution d’un groupe au niveau spécifique n’est possible que dans la mesure où suffisamment d’occurrences sont identifiées au niveau de l’espèce. cet exemple peut sembler caricatural, mais il existe une infinité de cas pour lesquels mesurer l’adéquation des données à la problématique de départ est plus délicat. la construction de ce jeu de données d’occurrences doit donc s’accompagner d’une réflexion sur cette 4 https://doi.org/10.15468/dl.rz7r5t. possibilité d’inadéquation, et potentiellement à une adaptation des questionnements de départ. une autre question pouvant se poser à ce stade est le nombre d’occurrences qui composent le jeu de données final. il n’existe pas de règle fixe quant au nombre minimal d’occurrence nécessaires, si ce n’est qu’il doit être le plus élevé possible sans pour autant sacrifier la qualité des occurrences. une étude a par exemple montré que 3 occurrences peuvent suffire à produire un modèle de niche robuste pour des espèces à faible étendue spatiale, versus 13 occurrences pour des espèces plus largement répandues (proosdij et al., 2016). toutefois, si trop peu d’occurrences sont utilisées (< 25 localités), la variabilité des modèles et des statistiques d’évaluation augmente (e.g., hernandez et al., 2006; wisz et al., 2008). plusieurs protocoles existent pour limiter les problèmes induits par un trop faible nombre d’occurrences dans la création et l’évaluation des modèles de niches: par exemple, l’évaluation par jackknife ou « leave-one-out » (pearson et al., 2006; shcheglovitova et anderson, 2013; gibert, vignoles et al., 2022) permet d’évaluer la variabilité du modèle en fonction de l’occurrence utilisée pour l’évaluation du modèle. pour prendre en compte plus de variables environnementales que le nombre d’occurrences, il est aussi possible d’utiliser les « ensembles de petits modèles » (breiner et al., 2015), qui consistent à combiner plusieurs modèles n’utilisant qu’une petite partie des variables environnementales à chaque fois (e.g., modèles bivariés). enfin, l’algorithme maxent semble bien adapté aux petits jeux de données d’occurrences, comme l’ont démontré plusieurs études comparatives (e.g., hernandez et al., 2006; papeş et gaubert, 2007; wisz et al., 2008) quelques elements a prendre en compte dans le processus de modelisation quels algorithmes pour l’enm ? il existe de nombreux algorithmes permettant de proposer un modèle de niche écologique, chacun caractérisé par des propriétés intrinsèques qu’il faut connaître afin de le choisir. le choix d’un algorithme dépend du type de données d’occurrences: de présence-absence (e.g., les forêts d’arbres décisionnels [random forest]; breiman, 2001; les modèles hiérarchiques [hierarchical models]; royle et dorazio, 2006; les modèles linéaires généralisés [glm] ou les modèles additifs généralisés [gam]; guisan et al., 2002), de présence-seule (e.g., bioclim; figure 7. résultats d’une recherche d’occurrences dans l’agrégateur de données d’occurrences gbif pour le chimpanzé commun (pan troglodytes) entre 1970 et 2022. la recherche fournit 3 001 occurrences de présence (points jaunes à orange), mais une partie d’entre elles comportent probablement des erreurs ou des incertitudes trop importantes pour être prises en compte dans une étude employant l’enm. les erreurs les plus évidentes que nous pouvons relever visuellement sont indiquées par des flèches. la flèche fushia indique une occurrence aux coordonnées (0°, 0°), tandis que les flèches bleu clair indiquent des occurrences trop éloignées du reste de l’aire de répartition géographique. ces quatre localités sont clairement situées dans des environnements défavorables à la survie de pan troglodytes – à savoir l’océan pour la flèche fushia et des régions tempérées pour les flèches bleu clair. le reste du jeu de données nécessitera une vérification plus poussée de la fiabilité des informations taxonomiques, géographiques et temporelles. source: gbif.org (24 march 2022) gbif occurrence download3 https://doi.org/10.15468/dl.rz7r5t anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 80 booth et al., 2014; domain; carpenter et al., 1993; habitat; walker et cocks, 1991; ou les modèles basés sur la distance de mahalanobis; farber et kadmon, 2003) ou de présence-background (e.g., l’algorithme généralisé pour la production de groupes de règles [garp]; stockwell et noble, 1992; ou maxent; philips et al., 2006, 2017) et des objectifs de l’étude. comme souligné par plusieurs auteurs (e.g., antunes, 2015; qiao et al., 2015), il n’existe pas d’algorithme objectivement et universellement meilleur qu’un autre en enm. leur performance dépend principalement de leur adéquation avec les données et les objectifs de l’étude. dans cet article, je propose de nous focaliser sur le fonctionnement de l’algorithme maxent (philips et al., 2006, 2017; philips et dudik, 2008), en raison de sa popularité dans les études enm. celle-ci résulte de sa facilité d’utilisation, rendue possible grâce à un logiciel stand-alone5, à de nombreux packages r spécialisés dans son application et son optimisation (e.g., kuenm, maxnet, enmeval, enmtools, etc.; muscarella et al., 2014; cobos et al., 2019c; kass et al., 2021; warren et al., 2021b; phillips, 2021) et à ses bonnes performances par rapport à d’autres algorithmes (e.g., phillips et al., 2006; elith et al., 2006; hernandez et al., 2006). cette popularité masque les limites de cet algorithme ainsi que ses subtilités d’application – en particulier la nécessité de dépasser les paramètres par défauts de l’algorithme (morales et al., 2017). dans les paragraphes suivants, je présenterai donc plusieurs considérations permettant de correctement l’employer; je propose également une courte section sur une alternative à ce type de modèle, en particulier lorsque l’on veut se focaliser les dynamiques et comportements des niches écologiques dans l’espace environnemental. fonctionnement de l’algorithme maxent. « maxent » est le nom d’un algorithme utilisant le principe d’entropie maximale (phillips et al., 2006, 2017; phillips et dudik, 2008), c’est-à-dire que « la distribution de probabilité estimée doit être en accord avec ce qui est connu […], mais doit éviter les suppositions qui ne sont pas compatibles avec les données. » (peterson et al., 2011, p. 109). il est adapté aux données d’occurrences de type présenceseule, mais emploie des informations sur la variation de l’environnement (dénommé « background ») autour des points d’occurrences dans la construction du modèle: c’est donc un modèle de type « présencebackground ». concrètement, maxent va établir 5 https://github.com/mrmaxent/maxent. une relation entre l’occurrence d’une présence et la densité des variables environnementales qui lui sont associées, tout en prenant en compte la densité des variables environnementales associées à des pixels sélectionnés aléatoirement dans la zone étudiée (points de background). il en résulte une distribution de probabilité, qui fait ensuite l’objet d’une transformation logarithmique pour représenter la probabilité de présence de conditions adéquates (en anglais, « suitability », que je traduirai dans cet article par « favorabilité »; ibid., p. 109), soit des conditions proches des valeurs moyennes observées pour les points d’occurrences (elith et al., 2011). en termes de statistiques, maxent va estimer le ratio entre la densité conditionnelle des variables environnementales au niveau des points de présence f1(z), et la densité marginale (i.e., non conditionnelle) des variables environnementales au niveau des points de background f(z). parmi toutes les formules de f1(z) possibles, maxent choisit celle qui est la plus proche de f(z), selon le principe d’entropie maximale (ibid., p. 47). in fine, la fonction doit coller au maximum avec les données de présence, tout en apportant le moins d’informations possible sur le reste de l’aire prise en compte. en réalité, maxent calibre le modèle final non pas à partir des variables environnementales brutes, mais à partir de transformations de ces variables, appelées fonctions (features). cette fonctionnalité permet à maxent de modéliser des relations complexes entre les données de présence et la densité des variables environnementales. il existe cinq classes de features: linear, quadratic, product, threshold et hinge (phillips et al., 2006, 2017). une linear feature correspond à la variable continue en elle-même. le carré de cette variable constitue la quadratic feature. product est le produit de deux variables. threshold consiste à attribuer la valeur de 1 si f est au-dessus d’une certaine valeur, ou 0. hinge est similaire à threshold, mais utilise une fonction linéaire au lieu d’une fonction en escalier. le choix des types et du nombre de features utilisés nécessite un travail de réglage par l’utilisateur car les paramètres par défaut sont souvent inadaptés (e.g., anderson et gonzalez, 2011; elith et al., 2011; warren et seifert, 2011; merow et al., 2013; shcheglovitova et anderson, 2013; muscarella et al., 2014; cobos et al., 2019c; vignali et al., 2020; kass et al., 2021). de plus, une trop grande complexité du modèle entraîne parfois une trop grande adéquation entre le modèle et les https://github.com/mrmaxent/maxent anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 81 données (overfitting).ce comportement du modèle peut en effet limiter sa capacité à généraliser (e.g., peterson et al., 2007; elith et al., 2010), car celuici interprétera du bruit ou des biais comme faisant partie du motif recherché au lieu de ne pas les prendre en compte dans la formule (peterson et al., 2011; merow et al., 2014). les causes de l’overfitting sont par exemple une trop grande aire de calibration par rapport à la répartition géographique des occurrences, mais également un paramétrage excessif (trop grande complexité du modèle) ou un trop faible nombre de d’occurrences. outre l’utilisation des features, maxent va imposer des contraintes au modèle lui permettant de tenir compte des données d’occurrences tout en évitant l’overfitting. en effet, le modèle doit correspondre aux données mais également conserver sa capacité de généraliser, notamment dans l’objectif de transférer le modèle à une autre période ou une autre zone géographique. dans le cas d’un overfit du modèle aux données, les courbes de réponse auront tendance à trop s’ajuster aux données de calibration, diminuant alors leur pouvoir prédictif. il est donc nécessaire, d’une part, de lisser la distribution afin de limiter l’overfitting et, d’autre part, d’éliminer un certain nombre de features qui complexifieraient le modèle à outrance. cette étape est contrôlée par un multiplicateur de régularisation l1 (parfois appelé « multiplicateur β »), dont le choix nécessite lui aussi une étape d’optimisation, la valeur par défaut n’étant pas toujours adéquate (e.g., elith et al., 2011; warren et seifert, 2011; warren et al., 2014; merow et al., 2013; shcheglovitova et anderson, 2013; muscarella et al., 2014; radosavljevic et anderson, 2014; cobos et al., 2019c; vignali et al., 2020; kass et al., 2021). pour terminer, la niche modélisée par l’algorithme maxent et sa projection géographique se situera quelque part entre nr et n* f en fonction de la configuration du diagramme bam dans laquelle on se situe (figure 4). le premier type de niche sera approché dans la configuration de superposition: le modèle de sortie pourra être interprété comme une carte de probabilité d’occurrence (peterson et al., 2011, p. 109), puisque dans ce cas gp équivaut à go. dans les configurations classiques et de hutchinson, il est possible, mais plus délicat d’estimer nr, et donc go, car cela nécessite d’autres données (e.g., données de dispersion, vraies données d’absence; peterson et soberón, 2012). il est donc plus probable que le modèle de sortie soit proche de n* f, étant donné qu’il représenterait plutôt la probabilité de présence de conditions adéquates, alias favorabilité. évaluation. l’évaluation est une étape-clé dans la modélisation de niches par le biais d’algorithmes prédictifs tels que maxent. c’est au cours de cette phase que l’on va déterminer si un modèle est robuste. pour cela, il est nécessaire de quantifier deux aspects principaux: la performance et la signification. la performance d’un modèle désigne sa capacité à atteindre un objectif précis dans l’absolu. dans le cas de l’enm, il s’agit de vérifier si le modèle prédit correctement des données d’occurrences indépendantes du jeu de données utilisé pour la calibration. ce type de données est rarement disponible, et a fortiori impossible à obtenir pour des taxons paléontologiques ou des unités archéologiques. plus généralement, on procède à un partitionnement du jeu de données d’occurrence initial afin de créer deux jeux de données supposés indépendants (fielding et bell, 1997). le but de cette manœuvre est tout d’abord de calibrer le modèle à partir de points de calibration (model training en anglais), puis de vérifier que les points de test sont correctement prédits par le modèle (model testing). cette évaluation s’effectue par le biais de statistiques permettant de rendre compte de la capacité prédictive du modèle (i.e., sa capacité à prédire les points de test à partir des points de calibration), comme le taux d’erreur d’omission (anderson et al., 2003; peterson et al., 2008). il est primordial que les points de test et de calibration soient suffisamment éloignés spatialement (donc potentiellement écologiquement) afin de ne pas artificiellement gonfler les mesures de performance. en d’autres termes, il ne faut pas que les points de test se situent dans les mêmes pixels que les points de calibration, afin de garantir l’indépendance des points de test et de calibration. plusieurs méthodes de partitionnement existent et sont plus ou moins adaptées à des situations précises (e.g., peterson et al., 2011; shcheglovitova et anderson, 2013; muscarella et al., 2014; radosavljevic et anderson, 2014; roberts et al., 2017; valavi et al., 2018). par exemple, la méthode du jackknife ou leave-one-out (pearson et al., 2006; shcheglovitova et anderson, 2013), qui consiste à évaluer la performance de chaque modèle calibré à partir des (n – 1) points d’occurrence avec le nième point, est adaptée pour les très petits jeux de données. une autre méthode actuellement préconisée est le partitionnement par blocs spatiaux de validation croisée (roberts et al., anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 82 2017; valavi et al., 2018), et permet notamment de prévenir l’autocorrélation spatiale entre les données de calibration et de test, garantissant ainsi mieux leur indépendance (sillero et barbosa 2020). déterminer la signification statistique d’un modèle consiste à évaluer si sa performance est meilleure que ce à quoi l’on pourrait s’attendre dans le cadre de l’hypothèse nulle. en d’autres termes, il s’agit de tester statistiquement si les prédictions des points d’évaluation par le modèle ne sont pas aléatoires compte tenu des données de calibration et d’évaluation (peterson et al., 2011). de la même façon que pour la mesure de la performance, il existe plusieurs méthodes pour tester la signification statistique d’un modèle (e.g., ibid., fielding et bell, 1997; peterson et al., 2008), la plus couramment utilisée de nos jours étant le calcul de l’aire sous la courbe (area under curve, abrév. auc) du receiver operating characteristic (abrév. roc; e.g., fielding et bell, 1997). le roc est une courbe illustrant la variation de la sensibilité du modèle (la proportion de présences connues correctement prédites, i.e., 1 – taux de faux négatifs) en fonction de 1 – sa spécificité (la proportion d’absences connues prédites comme présentes, i.e., taux de faux positifs). l’auc du roc est ensuite comparée à la courbe appartenant à un modèle noninformatif – dont l’auc est égal à 0.5 – c’est-àdire qu’il ne peut discriminer les vraies des fausses présences (elith et al., 2006; peterson et al., 2008; figure 8). or, lorsque le jeu de données n’est constitué que de présences, il n’est pas possible de correctement mesurer 1 – la spécificité du modèle. ainsi, il est préférable d’utiliser l’approche du roc modifiée par peterson et al. (2008), qui 1) s’adapte à l’approche présence/background, 2) se restreint au domaine prédit par l’algorithme sans se soucier des zones sujettes à l’extrapolation (en dehors de m), et 3) se restreint au domaine dans lequel le taux d’erreurs d’omissions (e) est suffisamment faible en fonction d’une valeur prédéfinie par l’usager. enfin, il faut également prêter une attention figure 8. exemple de receiver operating characteristic (roc) d’une prédiction maxent pour évaluer sa signification statistique. la droite en pointillés rouges représente le roc d’un modèle-nul non informatif et a donc une aire sous la courbe (auc) de 0.5. la courbe noire représente le roc du modèle empirique. les zones colorées représentent son auc. les points blancs constituent les données servant à créer le modèle empirique. dans l’approche du roc partiel proposé par peterson et al. (2008), les aires de non-prédiction (c’est-à-dire en dehors des données de calibration) en rouge ne sont pas prises en compte dans le calcul de l’auc. d’autre part, l’aire sous la courbe correspondant à une proportion d’erreurs d’identification dans les données de calibration (e) en jaune est également écartée du calcul (figure modifiée d’après peterson et al., 2011, fig. 9.4, p.175). anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 83 particulière à la possibilité d’overfitting du modèle, c’est-à-dire une trop grande adéquation entre le modèle et les données. celui-ci peut être en partie dû à une trop grande complexité du modèle liée au paramétrage. dans le cadre de l’enm, il est généralement préférable d’opter pour des modèles plus simples, car susceptibles de mieux refléter la réponse d’une espèce aux variables environnementales (merow et al., 2014, jiménez et al., 2019). de même que pour la performance et la signification, plusieurs méthodes existent pour sélectionner les modèles en fonction de leur complexité (e.g., ibid., warren et seifert, 2011; warren et al., 2014; muscarella et al., 2014; cobos et al., 2019c). la métrique la plus courante est le critère d’information d’akaike (akaike information criteria) corrigé en fonction de la taille de l’échantillon (abrév. aicc; akaike, 1974; warren et seifert, 2011; warren et al., 2014). cette métrique consiste en la standardisation des scores d’adéquation de telle sorte que la somme de ces scores dans l’espace géographique soit égale à 1. ensuite, la probabilité des données par rapport au modèle est calculée par le produit des scores d’adéquation des pixels contenant une présence (ibid.). choix du modèle final. une fois les meilleurs modèles déterminés par le protocole d’évaluation, deux approches principales permettent à l’utilisateur d’aboutir à un modèle final qui sera employé dans la suite des analyses (transferts et comparaisons). certains défendent le choix du modèle obtenant les meilleures performances (e.g., qiao et al., 2015). cependant, cette approche comporte des biais, notamment liés à la difficulté de choisir un modèle parmi ceux ayant des performances égales ou similaires (antunes, 2015). il est donc souvent préconisé de créer un modèle de consensus de tous les modèles sélectionnés (e.g., ibid.; cobos et al., 2019b, c). cette méthode permet d’accéder à une estimation de l’incertitude associée à ces différents paramétrages d’un algorithme ou de différents algorithmes. identifier les sources d’incertitudes et leur localisation dans l’espace géographique est une précaution importante pour l’interprétation des modèles de niches et pour les analyses ultérieures (transferts et comparaisons; peterson et al., 2018). plusieurs méthodes de consensus existent (antunes, 2015). enfin, le modèle final nécessite un dernier traitement, appelé seuillage, avant d’être interprété et analysé. comme nous l’avons vu, les données d’occurrences peuvent comporter des erreurs d’identification qu’il est nécessaire de prendre en compte dans la calibration du modèle. cette prise en compte se fait notamment au moment d’évaluer la performance du modèle via le taux d’erreur d’omission. par exemple, supposons que le taux d’erreurs d’identifications au sein d’un corpus soit d’environ 5 %. il est alors nécessaire de permettre au modèle de se « tromper » (c’est-à-dire, d’omettre une occurrence dans la prédiction) dans maximum 5 % des cas (peterson et al., 2008). la mise en place de ce seuil doit nécessairement se retrouver dans le modèle final, par le biais du seuillage de la prédiction (antunes, 2015; sillero et al., 2021). dans mon exemple, nous pouvons considérer que la prédiction se trompe dans au plus 5 % des cas, et donc que 5 % des occurrences doivent se trouver en dehors des aires de favorabilité. de façon pratique, il s’agit de classer les occurrences en fonction de leur score de favorabilité, puis de retenir la valeur de favorabilité de la nième occurrence correspondant à 5 % du corpus. tous les pixels dont le score de favorabilité est égal ou inférieur à cette valeur seront classés comme nuls (non adéquats; figure 9). cette méthode est dénommée le seuil de sensibilité fixe (fixed sensitivity; peterson et al., 2011, p. 119), mais il existe d’autres façon de définir un seuil (ibid.). modéliser et visualiser la niche fondamentale existante dans l’espace environnemental. malgré la popularité de maxent, la forme des niches qu’il modélise reste généralement trop complexe pour correspondre à une approximation de n* f ou de nf (merow et al., 2014). en effet, si la tolérance d’une espèce à une variable est assez vraisemblablement unimodale (drake, 2015), alors nf devrait théoriquement prendre une forme convexe, comme un ellipsoïde ou un polyèdre (e.g., van aelst et rousseeuw, 2009; escobar et al., 2014, 2017; qiao et al., 2016; jiménez et al., 2019; soberón et peterson, 2020; nuñez-penichet et al., 2021b; jiménez et soberón, 2022). ces dernières années, plusieurs études se sont concentrées sur l’emploi de l’ellipsoïde en enm (qiao et al., 2016; jiménez et al., 2019; nuñez-penichet et al., 2021b; banks et al., 2021; vignoles, 2021; jiménez et soberón, 2022). il s’agit en effet d’un objet simple à modéliser et à quantifier, car il comporte deux paramètres: le centroïde et la matrice de covariance (jiménez et al., 2019). la simplicité du modèle permet à la fois de réduire les hypothèses quant au paramétrage anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 84 (contrairement à maxent; merow et al., 2013) et d’effectuer facilement des analyses comparatives du volume ou de la position des niches dans l’espace environnemental (nuñez-penichet et al., 2021b, p. 4). en cela, ce type de modèle est plus adapté à des problématiques centrées sur les dynamiques de niches écologiques dans l’espace environnemental que des modèles plus complexes, comme maxent – ces derniers seraient théoriquement plus pertinents pour modéliser nr par exemple (merow et al., 2014). importance d’estimer l’aire accessible m dans le processus d’enm la définition de m – la portion de paysage accessible à l’espèce étudiée en référence au diagramme bam (figure 3) – représente une étape majeure en enm. c’est généralement l’aire qui est utilisée pour la calibration du modèle (anderson et raza, 2010; barve et al., 2011; jiménez et soberón, 2022). elle va en effet représenter une zone dans laquelle l’espèce a pu se déplacer, donc dans laquelle son absence est a priori significative pour discriminer figure 9. illustration du seuillage de prédiction avec la méthode du seuil de sensibilité fixe. le modèle de niche maxent présenté correspond à celui d’une espèce virtuelle pour laquelle j’ai créé des données d’occurrences aléatoires (au nombre de 25). les variables environnementales utilisées pour représenter l’espace environnemental sont la température moyenne annuelle, la précipitation moyenne annuelle et la saisonnalité de la précipitation modélisées pour la france de nos jours (1979-2013; modèle chelsa, résolution spatiale de 10 arc min.; karger et al., 2017). j’ai employé le package r enmtools (warren et al., 2021b) pour créer le modèle maxent en utilisant les paramètres par défaut et en employant 20 % des données d’occurrences pour l’évaluation. dans cet exemple, le seuil de sensibilité est fixé à e = 5 %. a. prédiction « brute » à l’issue de la calibration: la favorabilité est présentée de manière continue. b. quatre premières occurrences du jeu de données classées par favorabilité croissante. les 5 % des occurrences associées aux plus faibles valeurs de favorabilité sont au nombre d’une seule; sa valeur de favorabilité est de 0.29 et sera retenue comme seuil pour le reclassement de la prédiction. c. prédiction seuillée: les pixels aux valeurs inférieures ou égales à 0.29 sont considérés comme non-adéquats (blancs) tandis que les valeurs supérieures sont divisées en faible/moyenne/haute favorabilité. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 85 les environnements favorables des environnements non-favorables (barve et al., 2011; peterson et al., 2011, p. 126). toutefois, sa taille impacte grandement les résultats de la calibration et de l’évaluation (lobo et al., 2008; vanderwal et al., 2009; anderson et raza, 2010). en effet, si l’aire de calibration est trop petite et inférieure à m, le modèle peut sous-estimer les facteurs macro-géographiques qui influencent la répartition géographique de l’espèce – notamment a (barve et al., 2011). lorsqu’elle est trop étendue, les statistiques d’évaluation sont exagérément bonnes, parce que l’étendue de la prédiction sera beaucoup plus importante. pour le formuler autrement, il y a plus de chances que les points servant à tester la capacité prédictive du modèle soient prédits correctement, sans pour autant que la réponse des variables soit biologiquement cohérente (vanderwal et al., 2009). la définition de m influence aussi les étapes visant à transférer et comparer les modèles de niches (e.g., warren et al., 2008; owens et al., 2013; nuñez-penichet et al., 2021a, b), car ces deux opérations nécessitent d’être relativisées par rapport aux conditions environnementales présentes dans l’aire de calibration (donc m). pour guider le choix d’une aire de calibration qui représente au mieux les zones accessibles au taxon étudié (i.e., m), il est nécessaire de se reposer sur des hypothèses biogéographiques claires plutôt que sur des limites administratives ou une quelconque autre aire arbitraire (peterson et al., 2011; barve et al., 2011). ces hypothèses doivent prendre en compte la capacité dispersive de l’espèce considérée et notamment l’effet limitant de barrières naturelles (océan, montagne, fleuve…) empêchant l’accès des individus à des conditions favorables (e.g., barve et al., 2011). cette aire est généralement estimée manuellement, par exemple en traçant une zone de dispersion maximale autour des occurrences en fonction de ce qui est connu de l’espèce ou de la population étudiée et en écartant les régions séparées des occurrences par une barrière géographique (e.g., escobar et al., 2018; nuñez-penichet et al., 2021b; vignoles, 2021). une autre proposition est d’utiliser des régions biotiques ou biogéographiques – c’està-dire des régions présentant un cortège spécifique particulier et différent de régions voisines (barve et al., 2011), ou encore la distribution du ou des écotopes dans lesquels l’espèce a été observée (soberón, 2010). ces estimations restent subjectives, car elles ne permettent pas de précisément prendre en compte l’histoire et les capacités dispersives du taxon. en effet, la période pendant laquelle l’espèce a été présente dans la zone d’étude impacte fortement ses possibilités d’accéder à des conditions environnementales favorables plus ou moins éloignées du cœur de sa répartition géographique (barve et al., 2011; machado-stredel et al., 2021). plus récemment, une méthode permettant d’estimer une aire m plus réaliste a été développée, en se basant sur des simulations de dispersion (ibid.). cette approche permet de mieux approximer l’aire occupée (go) pendant la durée d’existence de l’espèce dans une région, en fonction des fluctuations climatiques qu’elle a traversées et qui ont certainement influencé son accessibilité à certaines zones. dans le cas de l’application de l’enm au registre fossile ou archéologique, la définition de l’aire de calibration m peut être un défi d’autant plus grand. en effet, les connaissances limitées des comportements de dispersion d’organismes ou de populations aujourd’hui disparues obligent à se reposer sur des hypothèses d’autant plus spéculatives. certaines propositions existent toutefois dans la littérature. par exemple, dans le cas d’espèces fossiles, il est suggéré d’employer comme aire de calibration l’étendue spatiale des couches affleurantes dans laquelle l’organisme a été échantillonné (myers et al., 2015). en dehors de ces zones, il n’est en effet pas possible de savoir si les conditions environnementales étaient ou non favorables à l’implantation de l’organisme. concernant l’application de l’enm aux données archéologiques, l’estimation de m reste un point méthodologique encore peu exploré. les études récentes se basent sur des hypothèses biogéographiques – par exemple, en excluant les zones recouvertes de glaciers ou en incluant des portions de continent alors émergés – mais aussi culturelles, en se référant notamment au principe d’actualisme (i.e., les distances maximales parcourues par des groupes de chasseurs-cueilleurs subactuels; vignoles et al., 2021) ou à la provenance de matériaux présents au sein des sites (silex, parures en coquillages; vignoles, 2021). ces estimations de m restent toutefois approximatives. il conviendrait à l’avenir d’améliorer cette étape par le développement de méthodologies conduisant à des hypothèses de dispersion plus réalistes, en s’inspirant par exemple de travaux qui cherchent à modéliser les territoires parcourus par des groupes de chasseurs-cueilleurs par le biais de la modélisation d’agents (agentanaïs vignoles – guide francophone pour la modelisation de niches ecologiques 86 based modeling) et/ou du chemin le moins couteux (least cost path; e.g., gravel-miguel et wren, 2018; vaissié, 2021). transfert du modèle à d’autres régions ou périodes: précautions une fois le modèle calibré dans m, il est souvent nécessaire de projeter ce modèle dans d’autres contrées ou d’autres périodes. cette opération peut se révéler assez délicate en raison de la présence d’environnements inconnus à l’aire de calibration dans la nouvelle zone, ou dans la même zone mais sous un climat différent (e.g., randin et al., 2006; williams et jackson, 2007; williams et al., 2007; zurell et al., 2012; owens et al., 2013; peterson et al., 2018). ces conditions non-analogues peuvent être dues à la présence de nouvelles valeurs – par exemple, les valeurs de température moyenne annuelle varient entre 5°c et 10°c dans l’aire de calibration, tandis qu’elles varient entre 10°c et 20°c dans la zone où le modèle sera projeté (williams et al., 2007) – ou à la présence d’une nouvelle combinaison entre variables – par exemple, la combinaison de 5°c de température moyenne annuelle avec 350 mm.an-1 de précipitation moyenne annuelle existe dans la zone de projection mais pas dans la zone de calibration (zurell et al., 2012). or, ces conditions n’ayant pas été prises en compte dans la calibration du modèle, ce dernier sera obligé d’extrapoler lorsqu’il y sera confronté (owens et al., 2013; peterson et al., 2011). les prédictions du modèle dans ces conditions nouvelles auront de grandes chances d’être hautement incertaines, car le comportement du modèle ne peut être vérifié ou anticipé au-delà des conditions de calibration (figure 10). la réponse d’un individu à la variation d’une variable environnementale peut être représentée comme une courbe unimodale. or, lorsque cette courbe est tronquée en raison de l’absence d’une partie des valeurs de la variable environnementale dans l’aire de calibration, il est impossible de déterminer si la valeur sélective (fitness) va augmenter, rester constante ou diminuer (figure 10). la troncature des courbes de réponse est assez fréquente (peterson et al., 2011, 2018) puisque la modélisation corrélative de niches ne permet d’accéder qu’à des portions de nf (n* f ou nr en fonction des approches et données). malheureusement, les approches corrélatives ne sont pas forcément idéales pour déterminer des relations mécanistiques entre occurrences et environnement, ce qui rend l’extrapolation en dehors de m incertaine (peterson et al., 2011; owens et al., 2013). afin d’opérer un transfert limitant l’extrapolation du modèle, il est important d’évaluer la transférabilité du modèle à une autre zone ou une autre période (elith et al., 2010; zurell et al., 2012; owens et al., 2013). il s’agit de cartographier les conditions analogues (i.e., similaires) à ce qui existe dans m dans ce nouvel environnement. la métrique la plus récemment développée s’appelle mobilityoriented parity (abrév. mop) et « identifie les aires de stricte extrapolation et calcule la similarité environnementale entre les régions de calibration et de projection » (owens et al., 2013, p. 13). cette méthode se base sur le calcul de distances multivariées (e.g., distance de mahalanobis; mahalanobis, 1936; etherington, 2019) entre les environnements associés aux points de la région de projection et une proportion (prédéfinie par l’utilisateur) des environnements associés aux points de la région de calibration. cette proportion est généralement restreinte à une portion du nuage de points de m proche du nuage de points correspondant à la région de projection (ibid.). la carte de répartition des environnements similaires figure 10. illustration du risque d’extrapolation de prédictions de niches hors de la gamme environnementale des données de calibration (domaine vert). en dehors de la gamme environnementale de calibration (i.e., non représentée dans la région m, domaine rouge), il n’est pas possible de vérifier si la courbe de réponse (en pointillés) va continuer à monter, rester constante ou redescendre. il s’agit donc d’une courbe de réponse tronquée (figure et légende modifiées d’après peterson et al., 2011, fig.7.6, p. 127). anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 87 et des zones d’extrapolation stricte issue du mop pourra ensuite être superposée aux prédictions de niches, afin d’identifier les valeurs de favorabilité relevant de l’extrapolation pure et simple. comment comparer deux modèles de niches ? la comparaison de modèles de niches est une analyse très couramment mise en œuvre en enm, car elle permet d’aborder de nombreuses problématiques ayant trait à l’évolution des espèces ou à leurs similarités écologiques. il s’agit toutefois d’une étape délicate, car le choix d’une méthode de comparaison peut largement influencer les résultats (e.g., warren et al., 2008; rödder et engler, 2011). de façon générale, il est important de prendre en compte la dualité de hutchinson lors de cette étape, c’est-à-dire de proposer des comparaisons à la fois dans l’espace géographique et environnemental, étant donné que les propriétés des niches dans ces deux espaces peuvent être très différentes (figure 2). dans l’espace environnemental, il existe plusieurs approches en fonction de si l’on cherche à comparer leur recouvrement (i.e., les niches partagent-elles une partie de leurs environnements ?), leur proximité (i.e., la distance entre les niches dans l’espace multivarié des niches est-elle petite ?; mammola, 2019), ou encore leur similarité (i.e., les niches présentent-elles une forme similaire ?), leur équivalence (i.e., les niches sont-elles parfaitement identiques d’un point de vue environnemental?; warren et al., 2008). les deux premiers types de comparaisons se focalisent sur la localisation relative des niches au sein de l’espace environnemental et donnent généralement des résultats similaires (sauf dans le cas de volumes totalement disjoints: les mesures de recouvrement sont alors inefficaces; mammola, 2019). en revanche, les deux autres types de comparaison se focalisent sur la forme des enveloppes, sans référence à leur localisation dans l’espace environnemental, ce qui permet alors de répondre à d’autres types de questions. dans l’espace géographique, plusieurs types de comparaisons sont également proposés pour analyser la similarité de la distribution géographique de modèles de niches (e.g., warren et al., 2008; broennimann et al., 2011; nuñez-penichet et al., 2021a). ces méthodes vont se concentrer sur la comparaison de distributions de probabilités ou de favorabilité données par un ou des algorithmes; il ne s’agit donc pas de comparer des niches d’un point de vue environnemental, mais plutôt d’évaluer la similarité de la répartition d’environnements favorables ou occupés par les deux espèces considérées. la méthodologie la plus simple consiste à binariser les prédictions entre les pixels favorables et non-favorables, puis à la soustraire afin d’identifier les zones présentant un gain, une perte ou une absence de changement de la favorabilité (e.g., nuñez-penichet et al., 2021a; banks et al., 2021). elle ne permet toutefois pas de comparer le degré de favorabilité de différentes zones. pour ce faire, des statistiques plus complexes doivent être utilisées, qui prennent en compte les scores de favorabilité dans la comparaison (e.g., warren et al., 2008; broennimann et al., 2011). quelque soit le type de comparaison effectué, la signification des valeurs obtenues doit ensuite être vérifiée statistiquement (figure 11). en particulier, la proportion des environnements partagés par l’aire de calibration des deux modèles joue un rôle crucial dans la pertinence de la comparaison (e.g., warren et al., 2008; banks et al., 2021). si les environnements accessibles à deux espèces comparées sont trop différents, l’absence ou quasi-absence de recouvrement/similarité entre leurs niches sera peu, voire pas significative, car la comparaison de niches modélisées à partir des données d’occurrences prélevées aléatoirement dans les aires de calibration respectives aurait donné le même résultat (par exemple, figure 11b). en d’autres termes, l’absence de recouvrement résulterait de la chance plutôt que d’une véritable différence entre les niches par rapport aux environnements disponibles aux deux espèces dans leurs m respectifs. à l’inverse, si les environnements accessibles présentent un recouvrement important, mais que les modèles de niches ne se recouvrent que peu, alors celle-ci sera probablement significative étant donné que des niches modélisées à partir de données prélevées aléatoirement dans les aires de calibration respectives ont plus de chances de se recouvrir. la signification d’une comparaison est généralement testée par le biais de tests de randomisation (figure 11; e.g., warren et al., 2008, 2021; nuñez-penichet et al., 2021b; banks et al., 2021). ceux-ci consistent à comparer la valeur empirique de la métrique employée avec une distribution nulle de valeurs-nulles mesurées pour la comparaison de paires de modèles aléatoires générés à partir du background (i.e., environnements accessibles à chaque espèce comparée). afin de anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 88 figure 11. exemples de comparaisons du recouvrement d’ellipsoïdes en employant l’indice de jaccard (cf. mammola, 2019), dans le cadre de l’évaluation de l’impact du choix d’un type de simulation paléoclimatique sur les modèles de niches associés à une culture archéologique du paléolithique supérieur appelée rayssien (31 900 – 26 900 ans calibrés avec le présent; vignoles, 2021). cette culture est caractérisée par une méthode de fabrication d’armes de chasse particulière, qui n’était employée que dans une aire géographique restreinte et des environnements relativement spécifiques. ces analyses ont été réalisées avec le package r ellipsenm51 (e.g., nuñez-penichet et al., 2021b; banks et al., 2021). a. comparaison d’ellipsoïdes modélisés à partir de deux jeux de simulations paléoclimatiques transitoires (armstrong et al., 2019) moyennées sur 30 ans (vert clair) et 100 ans (vert foncé). le recouvrement des ellipsoïdes est important mais n’est pas significatif au regard de la proportion d’environnements partagés (fort recouvrement entre les background). b. comparaison d’ellipsoïdes modélisés à partir d’un jeu de simulations paléoclimatiques transitoires (vert clair; armstrong et al., 2019) et d’un jeu de simulations paléoclimatiques calculées à l’équilibre (vert foncé; beyer et al., 2020). le recouvrement des ellipsoïdes est nul, mais n’est pas significatif au regard de la proportion d’environnements partagés (quasi-absence de recouvrement entre les background). c. aire de calibration définie par l’intersection entre le trait de paléo-côte (-90 m il y a ca. 30 000 ans; siddall et al., 2003), l’extension maximale des calottes glaciaires (ehlers et gibbard, 2004) et la distance maximale d’approvisionnement en silex identifiée pour le rayssien (220 km; vignoles, 2021); points d’occurrences du rayssien. 5 https://github.com/marlonecobos/ellipsenm/. https://github.com/marlonecobos/ellipsenm/ anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 89 rejeter l’hypothèse de similarité (ou de différence), il faut alors que la valeur empirique se situe en dessous du seuil de signification défini par l’utilisateur (généralement 5%, i.e., p < 0.05) conclusions dans cet article, j’ai présenté un cadre théorique dans lequel inscrire toute étude en enm. j’ai pour cela rappelé les principaux concepts et terminologie de la théorie des niches, ainsi que sa relation avec la distribution géographique d’un taxon ou une population. les différentes étapes pratiques de l’enm ont ensuite été décrites: la construction de jeux de données d’occurrences et le choix des données environnementales; la création et l’évaluation des modèles de niches; leur transfert à d’autres régions ou périodes et leur comparaison. bien que loin d’être exhaustive, cette synthèse récapitule les principaux points de vigilance et éléments à prendre en compte lors de la mise en œuvre de l’enm, quel que soit le contexte d’application. j’ai particulièrement insisté sur (1) l’importance d’un regard critique sur les données de départ. les données d’occurrences peuvent être affectées par de nombreux biais qu’il convient d’identifier et de prendre en compte. de même, les données environnementales sont majoritairement issues de modèles et comportent un certain degré d’incertitude. ma synthèse a également mis en évidence (2) la diversité de pratiques qui existent à toutes les étapes de la chaîne. que ce soit dans le choix des prédicteurs environnementaux, de l’algorithme, de l’aire de calibration, de l’évaluation ou du paramétrage du modèle, de nombreuses approches peuvent être adoptées. c’est pourquoi je préconise, dans la mesure du possible, d’évaluer et de présenter la variabilité découlant de tout ou partie de ces choix méthodologiques sur les modèles finaux et leurs interprétations (diniz-filho et al., 2009; warren et al., 2021a). enfin, une fois les modèles calibrés, (3) leur transfert et leurs comparaisons nécessitent des précautions afin d’éviter que ces analyses et leurs interprétations soient trop fortement biaisées. ces opérations reposent en effet sur une référence à l’aire de calibration du modèle, qui doit servir à relativiser les comparaisons et tester leur significativité. ces préconisations se fondent sur un principe à garder à l’esprit lors de l’emploi de modèles mathématique: leurs résultats ne sont qu’une approximation du réel. ils dépendent majoritairement des données et des méthodes mathématiques de création et d’analyse des modèles. c’est pourquoi une vigilance constante quant à l’adéquation des données et des choix méthodologiques avec le cadre théorique de référence et les problématiques de l’étude est primordiale pour modéliser au mieux le phénomène étudié. remerciements je remercie a. townsend peterson de m’avoir proposé d’écrire cet article pour biodiversity informatics. merci également à amélie challier pour sa relecture de la première version du manuscrit et ses suggestions d’amélioration. merci enfin à lucas buffan et un relecteur anonyme pour leurs commentaires et suggestions qui m’ont permis de clarifier et de compléter cette synthèse. l’article est une adaptation de plusieurs chapitres de ma thèse de doctorat, dont l’élaboration a été financée par le projet région nouvelle-aquitaine « gravettoniches » (dir. w. e. banks) data availability les scripts r et les données ayant permis la réalisation de figures-exemples 2 et 9 sont accessibles à / the r scripts and data used for example-figures 2 and 9 are accessible at https://osf.io/3nwqk/?view_ only=7e51054abaf74582a7701253ec37704f declaration of competing interests the author has declared that no competing interests exist. literature cited aiello-lammens, m. e., r. a. boria, a. radosavljevic, b. vilela, et r. p. anderson. 2015. spthin: an r package for spatial thinning of species occurrence records for use in ecological niche models. ecography 38 (5): 541-45. akaike, h. 1974. a new look at the statistical model identification. ieee trans. automat. contr. ac-19 (6): 716-23. alkishe, a., m. e. cobos, a. t. peterson, et a. m. samy. 2020. recognizing sources of uncertainty in disease vector ecological niche models: an example with the tick rhipicephalus sanguineus sensu lato. perspect. ecol. conserv. 18 (2): 91-102. amatulli, g., s. domisch, m. tuanmu, b. parmentier, a. ranipeta, j. malczyk, et w. jetz. 2018. a suite of global, cross-scale topographic variables for environmental and biodiversity modeling. sci. data 5 (1): 180040. amatulli, g., d. mcinerney, t. sethi, p. strobl, et s. domisch. 2020. geomorpho90m, empirical evaluation and accuracy https://osf.io/3nwqk/?view_only=7e51054abaf74582a7701253ec37704f https://osf.io/3nwqk/?view_only=7e51054abaf74582a7701253ec37704f anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 90 assessment of global high-resolution geomorphometric layers. sci. data 7 (1): 162. anderson, r. p., et i. jr. gonzalez. 2011. species-specific tuning increases robustness to sampling bias in models of species distributions: an implementation with maxent. ecol. modell., 222: 279-2811. anderson, r. p., et a. raza. 2010. the effect of the extent of the study region on gis models of species geographic distributions and estimates of niche evolution: preliminary tests with montane rodents (genus nephelomys) in venezuela. j. of biogeogr. 37 (7): 1378-93. anderson, r. p, d. lew, et a. t. peterson. 2003. evaluating predictive models of species’ distributions: criteria for selecting optimal models. ecol. modell. 162 (3): 211-32. angilletta, m. j. 2009. thermal adaptation: a theoretical and empirical synthesis. oxford university press, oxford. antunes, n. 2015. application d’algorithmes prédictifs à l’identification de niche écoculturelles des populations du passé : approche ethnoarchéologique. ph.d thesis, université de bordeaux, pessac (france). araújo, m. b., et a. guisan. 2006. five (or so) challenges for species distribution modelling. j. of biogeogr. 33 (10): 1677-88. araújo, m. b., et a. t. peterson. 2012. uses and misuses of bioclimatic envelope modeling. ecology 93 (7): 1527-39. araújo, m. b., et a. rozenfeld. 2013. the geographic scaling of biotic interactions. ecography 37 (5): 406-15. araújo, m. b., r. p. anderson, a. m. barbosa, c. m. beale, c. f. dormann, r. early, r. a. garcia, et al. 2019. standards for distribution models in biodiversity assessments. sci. adv. 5 (1): eaat4858. ashraf, u., a. t. peterson, m. n. chaudhry, i. ashraf, z. saqib, s. r. ahmad, et h. ali. 2017. ecological niche model comparison under different climate scenarios: a case study of olea spp. in asia. ecosphere 8 (5): e01825. armstrong, e., p. o. hopcroft, et p. j. valdes. 2019. a simulated northern hemisphere terrestrial climate dataset for the past 60,000 years. sci. data 6 (1): 265. banks, w. e., f. d’errico, h. i. dibble, l. krishtalka, d. west, d. i. olszewski, a. t. peterson, et al. 2006. eco-cultural niche modeling: new tools for reconstructing the geography and ecology of past human populations. paleoanthropology, 68−83. banks, w. e., j. zilhão, f. d’errico, m. kageyama, a. sima, et a. ronchitelli. 2009. investigating links between ecology and bifacial tool types in western europe during the last glacial maximum. j. archaeol. sci. 36 (12): 2853-67. banks, w. e., t. aubry, f. d’errico, j. zilhão, a. liranoriega, et a. t. peterson. 2011. eco-cultural niches of the badegoulian: unraveling links between cultural adaptation and ecology during the last glacial maximum in france. j. anthropol. archaeol. 30 (3): 359-74. banks, w. e., m.-h. moncel, j.-p. raynal, m. e. cobos, d. romero-alvarez, m.-n. woillez, j.-p. faivre, et al. 2021. an ecological niche shift for neanderthal populations in western europe 70,000 years ago. sci. rep. 11 (1): 5346. barve, n., v. barve, a. jiménez-valverde, a. lira-noriega, s. p. maher, a. t. peterson, j. soberón, et f. villalobos. 2011. the crucial role of the accessible area in ecological niche modeling and species distribution modeling. ecol. modell. 222 (11): 1810-19. beyer, r. m., m. krapp, et a. manica. 2020. an empirical evaluation of bias correction methods for palaeoclimate simulations. clim. past 16 (4): 1493-1508. booth, t. h., h. a. nix, j. r. busby, et m. f. hutchinson. 2014. bioclim : the first species distribution modelling package, its early applications and relevance to most current maxent studies. diversity distrib. 20 (1): 1-9. boria, r. a., l. e. olson, s. m. goodman, et r. p. anderson. 2014. spatial filtering to reduce sampling bias can improve the performance of ecological niche models. ecol. modell. 275: 73-77. boucher, o., j. servonnat, a. l. albright, o. aumont, y. balkanski, v. bastrikov, s. bekki, et al. 2020. presentation and evaluation of the ipsl‐cm6a‐lr climate model. j. adv. model. earth syst. 12 (7): e2019ms002010. breiman, l. 2001. random forests. mach. learn. 45: 5-32. breiner, f. t., a. guisan, a. bergamini, et m. p. nobis. 2015. overcoming limitations of modelling rare species by using ensembles of small models. methods ecol. evol. 6 (10): 1210-18. broennimann, o., m. c. fitzpatrick, p. b. pearman, b. petitp., l. pellissier, n. g. yoccoz, w. thuiller, et al. 2011. measuring ecological niche overlap from occurrence and spatial environment data. glob. ecol. biogeogr. 21 (4): 481-97. carpenter, g., a. n. gillison, et j. winter. 1993. domain: a flexible modelling procedure for mapping potential distributions of plants and animals. biodivers. conserv. 2 (6): 667-80. chapman, a., j. wieczorek, p. zermoglio, m. luna, et d. bloom. 2020. improved georeferencing: three essential guiding documents. biodiv. information sci. stand. 4: e58983. chase, j. m., et m. a. leibold. 2003. ecological niches: linking classical and contemporary approaches. interspecific interactions. university of chicago press, chicago. claussen, m., l. mysak, a. weaver, m. crucifix, t. fichefet, m.f. loutre, s. weber, et al. 2002. earth system models of intermediate complexity: closing the gap in the spectrum of climate system models. clim. dyn. 18 (7): 579-86. cobos, m. e., a. t. peterson, l. osorio-olvera, et d. jiménezgarcía. 2019a. an exhaustive analysis of heuristic methods for variable selection in ecological niche modeling and species distribution modeling. ecol. inform. 53: 100983. cobos, m. e., l. osorio-olvera, et a. t. peterson. 2019b. assessment and representation of variability in ecological niche model predictions. biorxiv, 603100. cobos, m. e., a. t. peterson, n. barve, et l. osorio-olvera. 2019c. kuenm: an r package for detailed development of ecological niche models using maxent. peerj 7: e6281. colwell, r. k., et t. f. rangel. 2009. hutchinson’s duality: the once and future niche. proc. natl. acad. sci. u.s.a. 106 (2): 19651-58. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 91 diniz-filho, j. a. f., l. m. bini, t. f. rangel, r. d. loyola, c. hof, d. nogués-bravo, et m. b. araújo. 2009. partitioning and mapping uncertainties in ensembles of forecasts of species turnover under climate change. ecography 32 (6): 897-906. domisch, s., g. amatulli, et w. jetz. 2015. near-global freshwater-specific environmental variables for biodiversity analyses in 1 km resolution. sci. data 2 (1): 150073. drake, j. m. 2015. range bagging: a new method for ecological niche modelling from presence-only data. j. r. soc. interface 12 (107): 20150086. ehlers, j., et p. gibbard. 2004. quaternary glaciations extent and chronology, part i: europe. elsevier, amsterdam. elith, j., c. h. graham, r. p. anderson, m. dudík, simon ferrier, a. guisan, r. j. hijmans, et al. 2006. novel methods improve prediction of species’ distributions from occurrence data. ecography 29 (2): 129-51. elith, j., m. kearney, et s. phillips. 2010. the art of modelling range-shifting species. methods ecol. evol. 1 (4): 330-42. elith, j., s. j. phillips, trevor hastie, m. dudík, yung en chee, et c. j. yates. 2011. a statistical explanation of maxent for ecologists. divers. distrib. 17 (1): 43-57. elton, c. 1927. animal ecology. the macmillan company, new york. escobar, l. e., a. lira-noriega, g. medina-vogel, et a. t. peterson. 2014. potential for spread of the white-nose fungus (pseudogymnoascus destructans) in the americas: use of maxent and nichea to assure strict model transference. geospat. health 9 (1): 221-29. escobar, l. e., h. qiao, c. lee, et n. b. d. phelps. 2017. novel methods in disease biogeography: a case study with heterosporosis. front. vet. sci. 4: 105. escobar, l. e., h. qiao, j. cabello, et a. t. peterson. 2018. ecological niche modeling re-examined: a case study with the darwin’s fox. ecol. evol. 8 (10): 4757-70. etherington, t. r. 2019. mahalanobis distances and ecological niche modelling: correcting a chi-squared probability error. peerj 7: e6678. farber, o., et r. kadmon. 2003. assessment of alternative approaches for bioclimatic modeling with special emphasis on the mahalanobis distance. ecol. modell. 160 (1-2): 115-30. feng, xiao, d. s. park, cassondra walker, a. t. peterson, cory merow, et m. papeş. 2019a. a checklist for maximizing reproducibility of ecological niche models. nat. ecol. evol. 3 (10): 1382-95. feng, x., d. s. park, y. liang, r. pandey, et m. papeş. 2019b. collinearity in ecological niche modeling: confusions and challenges. ecol. evol. 9 (18): 10365-76. fick, s. e., et r. j. hijmans. 2017. worldclim 2: new 1‐km spatial resolution climate surfaces for global land areas. int. j. climatol. 37 (12): 4302-15. fielding, a. h., et j. f. bell. 1997. a review of methods for the assessment of prediction errors in conservation presence/ absence models. environ. conserv. 24 (1): 38-49. gause, g. f. 1934. the struggle for existence. the williams & wilkins company, baltimore. gibert, c., a. vignoles, c. contoux, w. e. banks, d. barboni, j.r. boisserie, o. chavasseau, et al. 2022. climate-inferred distribution estimates of mid-to-late pliocene hominins. glob. planet. change 210: 103756. goosse, h., v. brovkin, t. fichefet, r. haarsma, p. huybrechts, j. jongma, a. mouchet, et al. 2010. description of the earth system model of intermediate complexity loveclim version 1.2. geosci. model dev. 3 (2): 603-33. gravel-miguel, c., et c. d. wren. 2018. agent-based leastcost path analysis and the diffusion of cantabrian lower magdalenian engraved scapulae. j. archaeol. sci. 99: 1-9. green, r. h. 1971. a multivariate statistical approach to the hutchinsonian niche: bivalve molluscs of central canada. ecology 52 (4): 543-56. grinnell, j. 1917. the niche-relationships of the california thrasher. auk 34 (4): 427-33. grubb, p. j. 1977. the maintenance of species-richness in plant communities: the importance of the regeneration niche. biol. rev. 52 (1): 107-45. guisan, a., t. c. edwards, et t. hastie. 2002. generalized linear and generalized additive models in studies of species distributions: setting the scene. ecol. modell. 157 (2-3): 89-100. guisan, a., et n. e. zimmermann. 2000. predictive habitat distribution models in ecology. ecol. modell. 135 (2-3): 147-86. hengl, t., j. mendes de jesus, g. b. m. heuvelink, m. ruiperez gonzalez, m. kilibarda, a. blagotić, w. shangguan, et al. 2017. soilgrids250m: global gridded soil information based on machine learning. plos one 12 (2): e0169748. hernandez, p. a., c. h. graham, l. l. master, et d. l. albert. 2006. the effect of sample size and species characteristics on performance of different species distribution modeling methods. ecography 29 (5): 773-85. hutchinson, g. e. 1957. population studies: animal ecology and demography. bull.math. biol. 53 (1-2): 193-213. ingenloff, k. 2020. enhancing the correlative ecological niche modeling framework to incorporate the temporal dimension of species’ distributions. ph.d thesis, university of kansas, lawrence. ingenloff, k., et a. t. peterson. 2021. incorporating time into the traditional correlational distributional modeling framework: a proof-of-concept using the wood thrush (hylocichla mustelina). methods ecol. evol.12: 311-21. jackson, s. t, et j. t overpeck. 2000. responses of plant populations and communities to environmental changes of the late quaternary. paleobiology 26 (4): 194-220. jiménez, l., et j. soberón. 2022. estimating the fundamental niche: accounting for the uneven availability of existing climates in the calibration area. ecol. modell. 464: 109823. jiménez, l., j. soberón, j. a. christen, et d. soto. 2019. on the problem of modeling a fundamental niche from occurrence data. ecol. modell. 397: 74-83. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 92 kadmon, r., o. farber, et a. danin. 2004. effect of roadside bias on the accuracy of predictive maps produced by bioclimatic models. ecol. appl. 14 (2): 401-13. kageyama, m., n. combourieu nebout, p. sepulchre, o. peyron, g. krinner, g. ramstein, et j.-p. cazet. 2005. the last glacial maximum and heinrich event 1 in terms of climate and vegetation around the alboran sea: a preliminary model-data comparison. cr geosci. 337 (10-11): 983-92. karger, d. n., o. conrad, j. bönher, t. kawohl, h. kreft, r. w. soria-auza, n. e. zimmermann, h. p. linder, et m. kessler. 2017. climatologies at high resolution for the earth’s land surface areas. sci. data 4: 170122. kass, j. m., r. muscarella, p. j. galante, c. l. bohl, g. e. pinilla‐ buitrago, r. a. boria, m. soley‐guardia, et r. p. anderson. 2021. enmeval 2.0: redesigned for customizable and reproducible modeling of species’ niches and distributions. methods ecol. evol. 12: 1602-1608. latombe, g., a. burke, m. vrac, g. levavasseur, c. dumas, m. kageyama, et g. ramstein. 2018. comparison of spatial downscaling methods of general circulation model results to study climate variability during the last glacial maximum. geosci. model dev. 11 (7): 2563-79. lobo, j. m., a. jiménez-valverde, et r. real. 2008. auc: a misleading measure of the performance of predictive distribution models. glo. ecol. biogeogr. 17 (2): 145-51. lobo, j. m, a. jiménez-valverde, et j. hortal. 2010. the uncertain nature of absences and their importance in species distribution modelling. ecography 33: 103-14. machado-stredel, f., m. e. cobos, et a. t. peterson. 2021. a simulation-based method for selecting calibration areas for ecological niche models and species distribution models. front. biogeogr. 13 (4): e48814. mahalanobis, p.c. 1936. on the generalized distance in statistics. j. asiatic soc. bengal 26: 541-88. mammola, stefano. 2019. assessing similarity of n‐ dimensional hypervolumes: which metric to use? j. of biogeogr. 46 (9): 2012-23. marques, r., r. f. krüger, a. t. peterson, l. f. de melo, n. vicenzi, et d. jiménez-garcía. 2020. climate change implications for the distribution of the babesiosis and anaplasmosis tick vector, rhipicephalus (boophilus) microplus. vet. res. 51 (1): 81. marques, r., r. f. krüger, s. k. cunha, a. s. silveira, d. m.c.c. alves, g. d. rodrigues, a. t. peterson, et d. jiménezgarcía. 2021. climate change impacts on anopheles (k.) cruzii in urban areas of atlantic forest of brazil: challenges for malaria diseases. acta trop. 224: 106123. merow, c., m. j. smith, t. c. edwards jr, a. guisan, s. m. mcmahon, s. normand, w. thuiller, r. o. wüest, n. e. zimmermann, et j. elith. 2014. what do we gain from simplicity versus complexity in species distribution models? ecography 37 (12): 1267-81. merow, c., m. j. smith, et j. a. silander. 2013. a practical guide to maxent for modeling species’ distributions: what it does, and why inputs and settings matter. ecography 36 (10): 1058-69. morales, n. s., i. c. fernández, et v. baca-gonzález. 2017. maxent’s parameter configuration and small samples: are we paying attention to recommendations? a systematic review. peerj 5: e3093. muscarella, r., p. j. galante, m. soley-guardia, r. a. boria, j. m. kass, m. uriarte, et r. p. anderson. 2014. enmeval: an r package for conducting spatially independent evaluations and estimating optimal model complexity for maxent ecological niche models. methods ecol. evol. 5 (11): 1198-1205. myers, c. e., a. l. stigall, et b. s. lieberman. 2015. paleoenm: applying ecological niche modeling to the fossil record. paleobiology 41 (2): 226-44. nakazawa, y., a. t. peterson, e. martínez-meyer, et a. g. navarro-sigüenza. 2004. seasonal niches of nearticneotropical migratory birds: implications for the evolution of migration. auk 121 (2): 610. nuñez-penichet, c., l. osorio-olvera, v. h. gonzalez, m. e. cobos, l. jiménez, d. a. deraad, a. alkishe, et al. 2021a. geographic potential of the world’s largest hornet, vespa mandarinia smith (hymenoptera: vespidae), worldwide and particularly in north america. peerj 9: e10690. nuñez-penichet, c., m. e. cobos, et j. soberón. 2021b. nonoverlapping climatic niches and biogeographic barriers explain disjunct distributions of continental urania moths. front. biogeogr. 13 (2): e5214. owens, h. l., l. p. campbell, l. l. dornak, e. e. saupe, n. barve, j. soberón, k. ingenloff, et al. 2013. constraints on interpretation of ecological niche models by limited environmental ranges on calibration areas. ecol. modell. 263: 10-18. papeş, m., et p. gaubert. 2007. modelling ecological niches from low numbers of occurrences: assessment of the conservation status of poorly known viverrids (mammalia, carnivora) across two continents: ecological niche modelling of poorly known viverrids. divers. distrib. 13 (6): 890-902. pearson, r. g., et t. p. dawson. 2003. predicting the impacts of climate change on the distribution of species: are bioclimate envelope models useful? glob. ecol. biogeogr. 12 (5): 361-71. pearson, r. g., c. j. raxworthy, m. nakamura, et a. t. peterson. 2006. predicting species distributions from small numbers of occurrence records: a test case using cryptic geckos in madagascar. j. biogeogr. 34 (1): 102-17. peterson, a. t., et y. nakazawa. 2007. environmental data sets matter in ecological niche modelling: an example with solenopsis invicta and solenopsis richteri. glo. ecol. biogeogr. 17: 135-144. peterson, a. t., et a. g. navarro-sigüenza. 2009. making biodiversity discovery more efficient: an exploratory test using mexican birds. zootaxa 2246 (1): 58-66. peterson, a. t., et j. soberón. 2012. species distribution modeling and ecological niche modeling: getting the concepts right. nat. conserv. 10 (2): 102-7. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 93 peterson, a. t., m. papeş, et m. eaton. 2007. transferability and model evaluation in ecological niche modeling: a comparison of garp and maxent. ecography 30 (4): 550-60. peterson, a. t., m. papeş, et j. soberón. 2008. rethinking receiver operating characteristic analysis applications in ecological niche modeling. ecol. modell. 213 (1): 63-72. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martinez-meyer, m. nakamura, et m. b. araújo. 2011. ecological niches and geographic distributions. princeton university press, princeton. peterson, a. t., m. e. cobos, et d. jiménez-garcía. 2018. major challenges for correlational ecological niche model projections to future climate conditions: climate change, ecological niche models, and uncertainty. ann. n. y. acad. sci. 1429: 66-77. phillips, s. j. 2021. maxnet: fitting “maxent” species distribution models with “glmnet”. r package version 0.1.4, available at https://cran.r-project.org/web/packages/ maxnet/maxnet.pdf phillips, s. j., et m. dudík. 2008. modeling of species distributions with maxent: new extensions and a comprehensive evaluation. ecography 31 (2): 161-75. phillips, s. j., r. p. anderson, et r. e. schapire. 2006. maximum entropy modeling of species geographic distributions. ecol. modell. 190 (3-4): 231-59. phillips, s. j., r. p. anderson, m. dudík, r. e. schapire, et m. e. blair. 2017. opening the black box: an open-source release of maxent. ecography 40 (7): 887-93. pierrat, b. 2011. macroécologie des échinides de l’océan austral: distribution, biogéographie et modélisation. ph.d thesis, université de bourgogne, dijon (france). pocheville, a. 2015. the ecological niche: history and recent controversies. in handbook of evolutionary thinking in the sciences, édité par t. heams, philippe huneman, g. lecointre, et marc silberstein, 547-86. dordrecht: springer netherlands. proosdij, a. s. j., m. s. m. sosef, j. j. wieringa, et n. raes. 2016. minimum required number of specimen records to develop accurate species distribution models. ecography 39 (6): 542-52. pulliam, h.r. 2000. on the relationship between niche and distribution. ecol. lett. 3 (4): 349-61. qiao, h., j. soberón, et a. t. peterson. 2015. no silver bullets in correlative ecological niche modelling: insights from testing among many potential algorithms for niche estimation. methods ecol. evol. 6 (10): 1126-36. qiao, h., a. t. peterson, l. p. campbell, j. soberón, l. ji, et l. e. escobar. 2016. nichea: creating virtual species and ecological niches in multivariate environmental scenarios. ecography 39 (8): 805-13. radosavljevic, a., et r. p. anderson. 2014. making better maxent models of species distributions: complexity, overfitting and evaluation. j. of biogeogr. 41 (4): 629-43. randin, c. f., t. dirnböck, s. dullinger, n. e. zimmermann, m. zappa, et a. guisan. 2006. are niche-based species distribution models transferable in space? j. of biogeogr. 33 (10): 1689-1703. raxworthy, c. j., e. martinez-meyer, n. horning, r. a. nussbaum, g. e. schneider, m. a. ortega-huerta, et a. t. peterson. 2003. predicting distributions of known and unknown reptile species in madagascar. nat. 426 (6968): 837-41. rödder, d., et j. o. engler. 2011. quantitative metrics of overlaps in grinnellian niches: advances and possible drawbacks: quantitative metrics of niche overlap. glob. ecol. biogeogr. 20 (6): 915-27. roberts, d. r., v. bahn, s. ciuti, m. s. boyce, j. elith, g. guillera-arroita, s. hauenstein, et al. 2017. crossvalidation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. ecography 40 (8): 913-29. roche, d. m., c. dumas, m. bügelmayer, s. charbit, et c. ritz. 2014. adding a dynamical cryosphere to iloveclim (version 1.0): coupling with the grisli ice-sheet model. geosci. model dev. 7 (4): 1377-94. royle, j. a., et r. m. dorazio. 2006. hierarchical models of animal abundance and occurrence. j. agric. biol. environ. stat. 11 (3): 249-63. saupe, e.e., v. barve, c.e. myers, j. soberón, n. barve, c.m. hensz, a.t. peterson, h.l. owens, et a. lira-noriega. 2012. variation in niche and distribution model performance: the need for a priori assessment of key causal factors. ecol. modell. 237-238: 11-22. saupe, e. e., c. e. myers, a. t. peterson, j. soberón, j. singarayer, p. valdes, et h. qiao. 2019. spatio-temporal climate change contributes to latitudinal diversity gradients. nat. ecol. evol. 3 (10): 1419-29. sbrocco, e. j., et p. h. barber. 2013. marspec: ocean climate layers for marine spatial ecology: ecological archives e094-086. ecology 94 (4): 979-979. shcheglovitova, m., et r. p. anderson. 2013. estimating optimal complexity for ecological niche models: a jackknife approach for species with small sample sizes. ecol. modell. 269: 9-17. siddall, m., e.j. rohling, a. almogi-labin, ch. hemleben, d. meischner, i. schmelzer, et d.a. smeed. 2003. sea-level fluctuations during the last glacial cycle. nature 423: 853-58. signor, p. w., et j. h. lipps. 1982. gradual extinction patterns and catastrophes in the fossil record. geol. soc. am. bull. 190: 291-96. sillero, n. 2011. what does ecological modelling model? a proposed classification of ecological niche models based on their underlying methods. ecol. modell. 222 (8): 1343-46. sillero, n., et a. m. barbosa. 2020. common mistakes in ecological niche models. int j geogr. inf. sci. 35 (2): 213226. sillero, neftalí, salvador arenas-castro, urtzi enriquez‐urzelai, cândida gomes vale, diana sousa-guedes, fernando martínez-freiría, raimundo real, et a.m. barbosa. 2021. https://cran.r-project.org/web/packages/maxnet/maxnet.pdf https://cran.r-project.org/web/packages/maxnet/maxnet.pdf anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 94 want to model a species niche? a step-by-step guideline on correlative ecological niche modelling. ecol. modell. 456: 109671. singarayer, j. s., et p. j. valdes. 2010. high-latitude climate sensitivity to ice-sheet forcing over the last 120 kyr. quat. sci. rev. 29: 43-55. de siqueira, m. f., g. durigan, p. de marco jr, et a. t. peterson. 2009. something from nothing: using landscape similarity and ecological niche modeling to find rare plant species. j. nat. conserv. 17 (1): 25-32. soberón, j. 2007. grinnellian and eltonian niches and geographic distributions of species. ecol. lett. 10 (12): 1115-23. soberón, j. m. 2010. niche and area of distribution modeling: a population ecology perspective. ecography 33 (1): 159-67. soberón, j., et b. arroyo-peña. 2017. are fundamental niches larger than the realized? testing a 50-year-old prediction by hutchinson. plos one 12 (4): e0175138. soberón, j., et m. nakamura. 2009. niches and distributional areas: concepts, methods, and assumptions. proc. natl. acad. sci u. s. a. 106 (2): 19644-50. soberón, j., et a. t. peterson. 2005. interpretation of models of fundamental ecological niches and species’ distributional areas. biodiv. inform. 2: 1-10. ———. 2011. ecological niche shifts and environmental space anisotropy: a cautionary note. rev. mex. biodivers., no 82: 1348-55. ———. 2020. what is the shape of the fundamental grinnellian niche? theor. ecol. 13 (1): 105-15. sobral-souza, t., j. pereira santos, m. e. maldaner, m. s. limaribeiro, et m. c. ribeiro. 2021. ecoland: a multiscale niche modelling framework to improve predictions on biodiversity and conservation. perspect. ecol. conserv. 19 (3): 362-68. sohn, n., m. h. fernandez, m. papes, et m. anciães. 2013. ecological niche modeling in practice: flagship species and regional conservation planning. oecologia aust. 17 (3): 429-40. stockwell, d. r.b., et i. r. noble. 1992. induction of sets of rules from animal distribution data: a robust and informative method of data analysis. math. comput. simul. 33 (5-6): 385-90. sweeney, a. w., n. w. beebe, r. d. cooper, j. t. bauer, et a. t. peterson. 2006. environmental factors associated with distribution and range limits of malaria vector anopheles farauti in australia. j. med. entomol. 43 (5): 1068-75. vaissié, e. 2021. mobility of paleolithic populations: biomechanical considerations and spatiotemporal modelling. paleoanthropology 1: 120−144. valavi, r., j. elith, j. j lahoz-monfort, et g. guillera-arroita. 2018. blockcv: an r package for generating spatially or environmentally separated folds for k‐fold cross‐ validation of species distribution models. methods in ecol. evol.10 (2): 225-32. valdes, p. j., e. armstrong, m. p. s. badger, c. d. bradshaw, f. bragg, m. crucifix, t. davies-barnard, et al. 2017. the bridge hadcm3 family of climate models: hadcm3@ bristol v1.0. geosci. model dev. 10 (10): 3715-43. van aelst, stefan, et peter rousseeuw. 2009. minimum volume ellipsoid. wiley interdiscip. rev. comput. stat. 1 (1): 71-82. vanderwal, j., l. p. shoo, c. graham, et s. e. williams. 2009. selecting pseudo-absence data for presence-only distribution modeling: how far should you stray from what you know? ecol. modell. 220 (4): 589-94. varela, s., r. p. anderson, r. garcía-valdés, et f. fernándezgonzález. 2014. environmental filters reduce the effects of sampling bias and improve predictions of ecological niche models. ecography 37: 1084-91. varela, s., m. s. lima-ribeiro, et l. c. terribile. 2015. a short guide to the climatic variables of the last glacial maximum for biogeographers. plos one 10 (6): e0129037. vignali, s., a. g. barras, r. arlettaz, et v. braunisch. 2020. sdmtune: an r package to tune and evaluate species distribution models. ecol. evol. 10 (20): 11488-506. vignoles, a. 2021. trajectoires technologiques et dynamiques de niches éco-culturelles du gravettien moyen au gravettien récent en france. ph.d thesis, université de bordeaux, pessac (france). vignoles, a., w. e. banks, l. klaric, m. kageyama, m. e cobos, et d. romero-alvarez. 2021. investigating relationships between technological variability and ecology in the middle gravettian (ca. 32-28 ky cal. bp) in france. quat. sci. rev. 253: 106766. volterra, v. 1926. variazioni e fluttuazioni del numero d’individui in specie animali conviventi. academia nationale dei lincei, roma (italia). vrac, m., p. marbaix, d. paillard, et p. naveau. 2007. non-linear statistical downscaling of present and lgm precipitation and temperatures over europe. clim. past 3: 669-82. walker, p. a., et k. d. cocks. 1991. habitat: a procedure for modelling a disjoint environmental envelope for a plant or animal species. glob. ecol. biogeogr. 1 (4): 108-118. warren, d. l. 2012. in defense of ‘niche modeling’. trends ecol. evol. 27 (9): 497-500. warren, d. l., et s. n. seifert. 2011. ecological niche modeling in maxent: the importance of model complexity and the performance of model selection criteria. ecol. appl. 21 (2): 335-42. warren, d. l., r. e. glor, et m. turelli. 2008. environmental niche equivalency versus conservatism: quantitative approaches to niche evolution. evol. 62 (11): 2868-83. warren, d. l., a. n. wright, s. n. seifert, et h. b. shaffer. 2014. incorporating model complexity and spatial sampling bias into ecological niche models of climate change risks faced by 90 california vertebrate species of concern. divers. distrib. 20 (3): 334-43. warren, d. l., a. dornburg, k. zapfe, et t. l. iglesias. 2021a. the effects of climate change on australia’s only endemic pokémon: measuring bias in species distribution models. methods ecol. evol. 12 (6): 985-995. anaïs vignoles – guide francophone pour la modelisation de niches ecologiques 95 warren, d. l., n. j. matzke, m. cardillo, j. b. baumgartner, l. j. beaumont, m. turelli, r. e. glor, et al. 2021b. enmtools 1.0: an r package for comparative ecological biogeography. ecography 44: ecog.05485. whittaker, r. h., s. a. levin, et r. b. root. 1973. niche, habitat, and ecotope. am. nat. 107 (955): 321-38. wieczorek, j., q. guo, et r. hijmans. 2004. the point-radius method for georeferencing locality descriptions and calculating associated uncertainty. int. j. geogr. inf. sc. 18 (8): 745-67. williams, j. w., et s. t. jackson. 2007. novel climates, noanalog communities, and ecological surprises. front. ecol. environ. 5 (9): 475-82. williams, j. w., s. t. jackson, et j. e. kutzbach. 2007. projected distributions of novel and disappearing climates by 2100 ad. proc. natl. acad. sc. u. s. a. 104 (14): 5738-42. wisz, m. s., r. j. hijmans, j. li, a. t. peterson, c. h. graham, a. guisan, et nceas predicting species distributions working group†. 2008. effects of sample size on the performance of species distribution models. divers. distrib. 14 (5): 763-73. zhu, g., w. bu, y. gao, et g. liu. 2012. potential geographic distribution of brown marmorated stink bug invasion (halyomorpha halys). plos one 7 (2): e31246. zuquim, g., f. r. c. costa, h. tuomisto, g. m. moulatlet, et f. o. g. figueiredo. 2020. the importance of soils in predicting the future of plant habitat suitability in a tropical forest. plant soil 450 (1-2): 151-70. zurell, d., j. elith, et b. schröder. 2012. predicting to new environments: tools for visualizing model behaviour and impacts on mapped distributions. divers. distrib. 18 (6): 628-34. biodiversity informatics, 19, 2025, pp. 86-108 86 remote sensing enables accurate assessment of functional diversity rather than species diversity in sandy grasslands wen li,, yu peng*, and xiaoyue zhang college of life & environmental sciences, minzu university of china, beijing, china 100081 abstract. the prediction of grassland plant diversity using satellite imagery has been the subject of intensive research. however, the accuracy of functional diversity (fd) predictions remains unclear. to address this, high-spatial-resolution worldview-3 (wv-3) multispectral data were used to predict species diversity and fd at the pixel scale (1.2 × 1.2 m) in the central hunshandak sandland, inner mongolia, northern china. data collected from 120 field plots (6 × 6 m) were employed to train and validate several statistical learning methods, with the primary objective of establishing links between 156 satellite-derived spectral and texture indices and 6 plant diversity indices. among the various diversity indices tested, functional trait diversity—specifically functional attribute diversity (fad1) and modified functional attribute diversity (mfad)—were predicted most effectively (with coefficients of determination of approximately 0.29 and 0.14, respectively; n=48) using texture indices. in contrast, species diversity (richness, h, e, or d) and other fd metrics were not well predicted by wv-3 data. overall, wv data did not significantly improve the accuracy of plant diversity predictions in sandy grasslands. additionally, high plot-level vegetation coverage was found to enhance the performance of spectral indices in predicting h, e, d, and fd. these results underscore the importance of accounting for variability across field conditions and demonstrate the potential of high-spatial-and-spectral-resolution satellite imagery for monitoring plant functional diversity in sandy grasslands. key words: plant diversity, species diversity, functional diversity, texture, remote sensing, sandy grassland * corresponding author: yu peng, yuupeng@163.com. introduction the survival of both humans and animals depends on plant diversity (radhamoni et al., 2023). plant diversity can be measured at three different spatial scales: within-habitat diversity (α-diversity), between-habitat diversity (β-diversity), and regional diversity (γ-diversity) (socolar et al., 2016). species richness (r), shannon-wiener index (h), and simpson index (d) are used to measure α diversity. recently, leveraging the spectral characteristics and variations of different plant species, plant diversity has been assessed on a large scale using remote sensing techniques—exhibiting distinct advantages over traditional field measurement methods. different plant species demonstrate different spectral traits. the spectral species approach assumes that there are unique, definable spectral types (“spectral species”) that can be distinguished in image processing, thus, facilitating biodiversity estimation (rocchini & 2004; féret & asner, 2014). the linkage between species and “spectral species” can be validated using relatively simple to sophisticated spectral heterogeneity measures, which include measures of spectral entropy (gillespie et al., 2008), statistical dispersion (gould, 2000; palmer et al., 2002), mean euclidean distances between spectral clusters derived from principal components analysis (pca) (oldeland et al. 2010; rocchini, 2007), and the use of firstand second-order image texture analysis (culbert et al., 2012; viedma, et al., 2012; wood et al., 2013). the red (red; 630-690nm) and near infrared (nir; 760-900 nm) bands in the multi-spectral images are usually selected to assess species diversity (schowengerdt, 2007; white et al., 2010; peng et al 2019). various spectral vegetation indices generated from multiple bands, such as variation in normalized difference vegetation index (ndvi) (gould, 2000; bawa et al., 2002; xu, 2004; chawla et al., 2010, kiran and mudaliar, 2012), enhanced vegetation index (evi) (cabacinha, 2009; gao x, 2000; gallardo-cruz et al., 2012), infrared index (iri), middle infrared index (miri), atmospheric resistance vegetation index (arvi), and soil adjusted vegetation index (savi) can predict plant diversity with considerable high accuracy (nagendra, 2001; bawa, et al., 2002; schowengerdt, 2007; cabacinha, 2009). mailto:yuupeng@163.com wen li et al. – remote sensing to assess functional diversity in sandy grasslands 87 multispectral data from satellite platforms such as landsat satellites (sun et al., 2023; mapfumo et al., 2016), the sentinel series (mpakairi et al., 2022; xin et al 2024), quickbird (rocchini, 2007), worldview series (cho et al., 2012; rocchini, 2007), spot series (fauvel et al., 2020), and china’s gaofen-3 and haisi-1 satellites (gu et al., 2024) have been widely used for vegetation monitoring and plant diversity estimation. the wv-3 satellite has eight visible near infrared (vnir) bands (400–1040 nm) at 1.24 m resolution, has advantage in assessing plant diversity (ferreira et al., 2016). the emergence of nearearth hyperspectral spectroscopy has addressed the scale mismatch between ground-based species information and earlier satellite-based multispectral monitoring (schneider et al., 2017). hyperspectral sensors offer the technical advantage of integrating fine spectral resolution with hundreds of spectral bands, improves the detection and identification of subtle differences at the plant species, functional, and genetic levels (asner, 1998; zhang et al., 2023). although species diversity has been extensively explored through remote sensing data, functional diversity (fd), i.e., diversity in plant species adaptation to survive in different environments and various strategies in reproduction, pollination processes, seed-dispersal methods and life forms (ewers & didham, 2006; lindborg et al., 2012), has been rarely assessed using satellite-based high spatial and spectral resolution remote sensing imageries. fd is closely related to ecosystem service and ecological process, can also be regarded as an important indicator for biodiversity conservation (diaz and cabido, 2001). a plant’s life form is one predictor for indicating the survival ability of plants in severely environment (evju et al., 2015). plants with suitable life forms will gradually replace unsuitable plants in a poor environment. for example, in plant communities in the arid environments such as deserts, shrubs and semi-shrubs increasingly substitute annual plants. species with low offspring and colonization rates are easily influenced (higgins et al., 2003; henle et al., 2004) and therefore, dominant species in a community are usually highly adapted to suit their environments. therefore, this study aimed to assess the ability of wv-3 data in assessing fd in the sandy grasslands of hunshandak sandland, by comparing the performances with those fig. 1 location of the study area, land use types, distribution of sampled plots and subplots in the field, and corresponding pixels in the wv-3 image of the study area. each plot consists of 4 subplots (black circles) and corresponds to a 5×5 grid of wv-3 pixels. wen li et al. – remote sensing to assess functional diversity in sandy grasslands 88 of species diversity across various plot-level vegetation coverages. methods study area this study was conducted in the temperate sandy grasslands of the hunshandak sandland (41°46′–43°69′ n, 114°55′–116°38′ e), located in inner mongolia, northern china (fig. 1). the region has a temperate semi-arid climate, with an annual mean temperature of 1.7°c. the monthly diurnal minimum and maximum temperatures are -18.3°c and 18.7°c, respectively, and the annual precipitation ranges from 250 to 350 mm, 80–90% of which occurs between may and september (wang, 2016a). the hunshandak sandland features a unique landscape comprising fixed sandy dunes, semi-fixed sandy dunes, mobile dunes, and lowlands, all of which support relatively rich plant diversity. the study area also includes other land cover types such as water ponds and constructed land. these diverse landscape elements, combined with the region’s relatively uniform elevation, make it an ideal site for testing the capacity of wv-3 data to assess both plant species diversity and functional diversity (fd) in sandy grasslands under complex background conditions. field sampling field surveys were conducted across 120 plots in the study area during july–august 2016, which corresponds to the peak growing season. each plot (6 × 6 m, equivalent to a 5 × 5 grid of wv-3 pixels) was divided into four subplots, each with a diameter of 0.8 m (fig. 1). in total, 480 subplots (corresponding to 480 wv-3 pixels), distributed across the 120 plots, were surveyed in this study. global positioning system (gps) data were differentially corrected to achieve high-precision positioning within a geographic information system (gis) environment. vascular plant species composition was recorded at the plot scale, with macroplot-level species lists compiled as the aggregated set of species identified across the four subplots within each plot. this nested sampling design for plot configuration has also been adopted in studies by duccio rocchini (2007) and fauvel et al. (2020). all plants were identified to the species level, and the abundance of each species was recorded for each subplot. details of the plot characteristics are provided in table s1 in the appendix. plant species diversity in july and august 2016, the abundance, cover, and height of each plant species, as well as habitat categories (fixed sandy dunes, semi-fixed sandy dunes, mobile sandy dunes, lowlands, water bodies, and constructed land), were recorded. for each subplot, the number of individuals was counted for species whose stems were either fully or partially within the subplot. for clonal species, individuals were considered separate if their stems or culms were more than 20 cm apart from others of the same species. canopy cover of all species within the subplot was visually estimated, and consistency in these visual estimates was ensured by having the same observer (y. peng) conduct all assessments across plots. based on the collected data on plant species abundance, four biodiversity indices were calculated: richness (the number of plant species in a subplot), the shannon– wiener index (h), simpson’s species evenness index (d), and the pielou index (e) (magurran, 2004), using formulae (1)–(3). h = – n n n n is i i ln 1 ∑ = ln n n n n is i i ln 1 ∑ = (1), where ni is the number of individuals of the ith species, n is the total number of individuals of all the species, and ln is the natural logarithm. the value of h ranges from 0, meaning only one species is present, to 4.6, signifying high species richness and also signifying that different species in the quadrat or a community are equally abundant (magurran, 2004). d = 1 – ∑ = s i i n n 1 )( 2 (2) the values of the simpson species evenness index range from 0 (completely uneven) to 1 (different species occur in equal numbers). e = h/lns (3), where s is the total number of species recorded (γ diversity) and h is the shannon–wiener index. the plant diversity and dominant species of study plots are listed in table s1 in the appendix. functional trait diversity functional traits recorded for each species included life form (annual, biennial, perennial grass, shrub, or woody plant), seed dispersal method (gravitational, regular, wind, or animal-mediated), pollination mode (self-pollination, wind-pollination, or insect-pollination), flowering period (in months), flower longevity (in days), photosynthetic pathway (c3, c4, or cam), and nitrogen-fixing ability (n-fixing or non-n-fixing). based on these multiple functional traits, seven plant functional diversity (fd) indices were calculated using the fdiversity package (casanoves et al., 2011; spasojevic et al., 2014): functional attribute diversity (fad1), modified functional attribute diversiwen li et al. – remote sensing to assess functional diversity in sandy grasslands 89 image window or kernel) (anderson et al., 2009; duro et al., 2014; levin et al., 2007; lucas & carter, 2008). here, we used the cv to link spectral heterogeneity within 5×5 image pixels to field-measured species diversity and functional diversity (fd) at each plot. before conducting the calculations, an ndvi threshold of ≥0.2 was applied to distinguish vegetation from desert areas, in accordance with the method proposed by xin et al. (2024). first, 35 spectral indices (table s2 in the appendix; details are provided in the envi manual) were calculated at the pixel level using envi software. second, for each spectral index, the cv (cv), mean value (mean), majority (maj), and variance (var) were computed using all pixels within a plot, yielding 140 spectral diversity values at the plot level. this approach allowed us to quantify the spectral diversity of each plot. finally, spectral diversity values derived from all pixels within the plots were compared with field-measured species diversity and fd data. image texture metrics, derived from multi-scale spectral values, are effective predictors of plant species richness. in this study, we also computed spectral texture values for each plot. in image texture analysis, the value of a central pixel within a moving window is determined by the spectral variability of its neighboring pixels (hall-beyer, 2017). we calculated one first-order texture metric and six second-order (_sec) texture metrics: angular second moment (mom), contrast (cont), dissimilarity (dis), homogeneity (hom), entropy (ent), and correlation (corr, _cor). these metrics were derived from panchromatic images and principal components (pcs) of multispectral images, where the pcs were obtained via principal component analysis (pca) of the eight spectral bands. second-order textures are based on the gray-level co-occurrence matrix, thus accounting for the spatial arrangement and relationships among neighboring pixels (farwell et al., 2020). detailed descriptions of these image texture metrics are available in the envi software manual. a single moving window size (5×5 pixels, corresponding to 6×6 m²) was selected for the analysis, as texture metrics across different window sizes are highly correlated and exhibit similar relationships with plant species richness (st-louis et al., 2006; culbert et al., 2012). additionally, the cv (cv), mean (mean), majority (maj), and variance (var) of the pcs were calculated for each plot. in total, 140 spectral indices, 12 texture indices, and 4 pc indices were computed for each plot. statistical analysis of the 120 plots, five were excluded as their pixels primarily covered shifting sandy land. we employed ty (mfad), rrao, functional evenness (feve), functional divergence (fdiv), functional dispersion (fdis), and functional specialization (fspe). fad1 represents the number of distinct attribute combinations present in the community, with values always less than or equal to species richness. mfad is a modified index calculated as the sum of standardized distances between all pairs of species in trait space. rrao is derived from ultrametric trait distances and the abundance distribution of species within the community. feve measures the regularity of spacing between species in trait space. fdiv quantifies the spread of trait values across the range of the trait space. fdis is a multidimensional index based on multi-trait dispersion. fspe, associated with threat categories, quantifies the average distinctiveness of all threatened species. the calculation formulas for each fd index are detailed in casanoves et al. (2011). satellite image acquisition a cloud-free worldview-3 (wv-3, digitalglobe, inc.) image covering the study area was acquired on 12 september 2015 (fig. 1). at this time, most grasses remained leafy, making them easily distinguishable in the imagery due to their strong contrast with the surrounding sandy terrain. the wv-3 dataset included a panchromatic image with a spatial resolution of 0.30 m, accompanied by a multispectral image with a 1.20 m spatial resolution, spanning 8 spectral bands: coastal blue (427 nm), blue (482 nm), green (547 nm), yellow (604 nm), red (660 nm), red-edge (723 nm), near-infrared 1 (824 nm), and near-infrared 2 (914 nm). the wv-3 data were delivered at level 2a. wv-3 data processing followed the methods described by lelong et al. (2020) and cerrejón et al. (2023). first, radiometric calibration was performed to convert digital numbers to absolute radiance using the gain and offset values for each spectral band. next, absolute radiance was converted to top-of-atmosphere (toa) reflectance, and the data were orthorectified to correct geometric distortions while minimizing topographical effects. finally, the multispectral image was fused with the panchromatic image to generate a pansharpened multispectral image with a resolution of 0.30 m. spectral and textural analysis this study assessed plant species diversity and functional trait diversity based on the spectral diversity theory, which posits that higher spectral diversity corresponds to greater plant diversity. numerous studies have shown that dispersion metrics, such as the coefficient of variation (cv), serve as simple yet effective indicators of spectral heterogeneity within a sampling unit (e.g., an wen li et al. – remote sensing to assess functional diversity in sandy grasslands 90 a two-step approach to explore the potential of spectral and texture indices for assessing plant diversity. the first step involved selecting indices with significant pearson’s correlation coefficients using 67 plots. the second step validated these selected indices across different vegetation coverages using the remaining 48 plots. pearson’s correlation coefficients were used to evaluate relationships between plant species diversity, functional trait diversity indices, and spectral/texture indices derived from satellite imagery. this provided an assessment of the potential of these spectral indices and texture metrics for estimating plant diversity or functional diversity. spectral indices or texture metrics were considered optimal if they showed a statistically significant association (p < 0.05) with plant diversity indices based on pearson’s correlation analysis. potential indices that exhibited significant relationships with plant diversity were designated as final optimal indices only if they passed the validation test. for this test, the remaining 48 randomly selected plots were used for model validation. the potential spectral indices were required to show high consistency in predicting plant diversity across different plant communities. we used vegetation coverage to examine its influence on the consistency of the selected indices in estimating plant diversity. plot-level vegetation coverage was categorized into three classes based on values: 0–15%, 16–26%, and 27–100%, representing low, moderate, and high community coverage, respectively. the performance of the selected indices in the validation test was evaluated using two metrics derived from the validation dataset: the correlation coefficient (r) and the significance level (p) between predicted and observed richness values. indices with the highest r and p < 0.01 were deemed the best predictors. additionally, cluster analysis was used to identify groups of indices with similar performance, aiming to explore the underlying mechanisms of the best-performing indices in assessing plant diversity. these groups were not predefined prior to the analysis, and no a priori assumptions were made regarding the distribution of variables (indices). the indices were z-transformed for the analysis, and results are presented as a dendrogram, with groupings based on a squared euclidean distance matrix. results cluster tree the top 51 spectral and texture indices, which exhibited the highest number of significant correlations (p < 0.05) with plant diversity were clustered into five distinct groups (fig. 2). indices showing significant positive correlations with the species diversity indices (d, h) and the functional diversity index (fad1) included correlation (corr), principal component correlation (pca-cor), variance (var), principal component variance (pca-var), entropy (ent), principal component entropy (pca-ent), dissimilarity (dis), contrast (cont), principal component dissimilarity (pca-dis), and principal component contrast (pca-con). these indices all belong to image texture measures. notably, spectral indices based on spatial variability (e.g., coefficient of variation [cv], such as evi-cv) showed no significant relationships with species or functional diversity indices. indices with the most significant relationships were identified as potential optimal indices and selected for model construction and validation. based on this criterion, corr, pca-cor, var, and pca-var were chosen as potential indices for further analysis. model development six texture and spectral indices that passed the aforementioned tests were retained for model development. the resulting models exhibited high coefficients of determination (r2) and significant relationships between the indices and plant diversity indices (p < 0.05), indicating their potential as optimal indices (table 1). among these models, simpson’s index (d) was significantly predicted by variance (var), correlation (corr), principal component variance (pca-var), principal component correlation (pca-cor), and var-mean. all functional diversity (fd) indices—functional attribute diversity (fad1), modified functional attribute diversity (mfad), and functional specialization (fspe)—were significantly predicted by variance (var), correlation (corr), principal component variance (pca-var), and principal component correlation (pca-cor). model validation using the six identified optimal spectral and texture indices listed in table 1, plant diversity indices were calculated for the 48 model validation plots using the spectral and texture dataset. at the plot level, linear correlations between diversity estimates derived from spectral and texture indices and field-surveyed diversity were analyzed (fig. 3). the six selected spectral and texture indices were further compared in terms of the consistency of their relationships with plant species or functional diversity across different vegetation coverage classes (0–15%, 16–26%, and 27–60%). indices that exhibited significant correlations across all vegetation coverage classes were designated as the best-performing indices. all models showed non-significant correlations (p > 0.05) between recorded and predicted plant diversiwen li et al. – remote sensing to assess functional diversity in sandy grasslands 91 fig. 2 cluster dendrogram showing pearson’s correlation coefficients between spectral/texture indices and plant diversity indices. deep red and blue indicate significant positive and negative correlations, respectively (p < 0.05); light colors indicate non-significant correlations (p > 0.05). wen li et al. – remote sensing to assess functional diversity in sandy grasslands 92 fig. 3 linear regressions between the natural log-transformed field-measured values (x-axis) and natural log-transformed predicted values (y-axis) for plant species diversity and functional diversity, using the validation dataset from the central hunshandak sandland, northern china. predicted values were derived from the best-performing models listed in table 1. “low,” “mid,” and “high” indicate plot-level vegetation coverage classes of 0–15%, 16–26%, and 27–60%, respectively. wen li et al. – remote sensing to assess functional diversity in sandy grasslands 93 ty indices, except for the correlations between modified functional attribute diversity (mfad) and four texture indices, and between variance (var) and simpson’s index (d) (fig. 3). when comparing correlation coefficient (r) values across vegetation coverage classes, all spectral and texture models exhibited similarly high r values (>0.2) for high-coverage plots (>27%), followed by low-coverage plots, while models for medium-coverage plots had the lowest r values. among functional diversity indices, mfad was most effectively predicted by the selected texture indices. discussion among the 156 spectral indices and texture metrics, only four performed relatively well in estimating plant species diversity and functional diversity (fd). these four indices were primarily texture measures: variance (var), principal component variance (pca-var), var-mean, and correlation (corr). these indices, which reflect canopy structure and functional traits, can effectively indicate fd in sandy grasslands. notably, leaf traits related to light capture and growth—such as photosynthetic pigments, nutrients, and leaf mass—primarily absorb and scatter light in the 350–700 nm spectral range (ollinger, 2011). in contrast, secondary metabolites (e.g., lignin, cellulose, phenols, and tannins), which contribute to foliar defense and longevity, actively absorb and scatter near-infrared (nir) and shortwave infrared (swir) radiation (kokaly et al., 2009). this may partially explain why spectral indices show stronger associations with plant fd than with species diversity. previous studies have observed that visible spectrum (vis) variability (e.g., captured by var, var-mean, and pca-var) is significantly higher in sandy landscapes compared to moist or wetland environments (somers et al., 2015). the visible spectrum, dominated by pigment absorption and canopy structure, exhibits less variability among plant species but greater variability among functional traits (kattenborn et al., 2017; pacheco-labrador et al 2022). additionally, biomass variation—closely linked to life form, a key functional trait—is primarily explained by vegetation structure (hernández-stefanoni et al., 2014) and can be well-captured by texture metrics. in sandy grasslands, environmental conditions strongly drive functional differentiation and adaptation in plants. the nir region, reflecting leaf cellular structure (which scatters most incident energy) and canopy biochemical traits (e.g., leaf nutrient content), can distinguish plant strategies such as c3, c4, or cam photosynthesis, or n-fixing capacity (castillo-riffart et al., 2017; pachecolabrador et al 2022). collectively, these factors explain why fd, rather than species diversity, is better estimated by spectral and texture indices. a pixel represents a discrete spatial unit containing multiple objects, and the proportion of mixed objects increases with coarser resolution. a study on speciesspectral diversity relationships across spatial grains, using north american floristic data, showed that high spatial and spectral resolution imagery improves plant species diversity estimation accuracy (rocchini et al., 2014). coarser resolution data, however, suffer from mixed-pixel issues and are less sensitive to spatial complexity (rocchini, 2007). previous research has found that estimation accuracy improves when spectral pixel resolution is finer than the size of the target object (gholizadeh et al., 2019; lopatin et al., 2017; wang et al., 2018). finer spatial resolution, such as that of wv-3 imagery, enhances the representation of “pure” objects within sampling areas (rocchini, 2007; pacheco-labrador et al 2022) and captures more detailed spectral information. consequently, wv-3 images exhibited higher spectral variability in 6×6 m plots compared to aster or landsat etm+ imagery, likely due to reduced pixel mixing and larger effective sampling size. wv-3’s high spectral resolution further strengthens its capacity to estimate plant diversity. table 1 selected spectral indices, texture indices, models, and parameters for estimating plant species diversity and functional diversity using the training dataset richness h d fda1 mfad fspe var y = 4.761-0.001x y = 1.249-0.000317x y = 0.650+0.00014x y = 4.814-0.001x y = 1.007+0.000307x y = 2.366+0.000346x r2=0.048, p = 0.041 r2=0.081, p = 0.011 r2=0.094, p = 0.007 r2=0.061, p = 0.025 r2=0.053, p = 0.034 r2=0.183, p = 0.000 corr y = 3.947+0.852x y = 1.046+0.206x y = 0.569+0.08x y = 3.945+0.911x y = 0.797+0.219x y = 2.534-0.154x r2=0.074, p = 0.026 r2=0.072, p = 0.016 r2=0.061, p = 0.025 r2=0.076, p = 0.014 r2=0.059, p = 0.027 r2=0.068, p = 0.021 pca-var y = 4.755-0.001x y = 1.247-0.000318x y = 0.649+0.000141x y = 4.807-0.001x y = 1.005+0.000305x y = 2.369+0.00034x r2=0.047, p = 0.044 r2=0.079, p = 0.012 r2=0.093, p = 0.007 r2=0.059, p = 0.028 r2=0.051, p = 0.038 r2=0.172, p = 0.000 pca-cor y = 1.045+0.207x y = 0.569+0.080x y = 3.941+0.915x y = 0.796+0.219x y = 2.534-0.155x r2=0.071, p = 0.016 r2=0.06, p = 0.025 r2=0.076, p = 0.015 r2=0.059, p = 0.028 r2=0.068, p = 0.021 ari-maj var-mean y = 3.944+0.855x y = 0.534+0.346x r2=0.06, p = 0.026 r2=0.065, p = 0.021 wen li et al. – remote sensing to assess functional diversity in sandy grasslands 94 in this study, texture metrics were derived from principal components (pcs) generated via principal component analysis (pca) of 12 spectral bands. pcs maximize variance and are non-spatial, and it can be assumed that the first two pcs from 12 bands likely correlate with the most variable top-of-canopy structural and chemical traits (roth et al., 2016; dahlin, 2016). our results align with previous studies using image texture to distinguish vegetation structural patterns between habitats (wood et al., 2012), which also reported weak relationships (0.01 < r² < 0.3). other studies have confirmed that image texture metrics effectively characterize vegetation heterogeneity, including foliage height diversity (wood et al., 2012), successional stage (jakubauskas, 1997), and structural complexity (guo et al., 2004)—consistent with the use of vegetation indices (e.g., ndvi, evi) to quantify biophysical aspects of functional traits. texture metrics from wv-3 can capture spatial variations in vertical structures of mixed grasslands under different grazing regimes (guo et al., 2004), supporting the utility of medium-resolution image textures for detecting vegetation heterogeneity. we used 5×5 pixel windows to calculate texture metrics, a size well-suited for grassland ecosystems. a study on north american plant species richness found that spectral diversity explained little variance, whereas the spatial extent of sampling units (floras) explained much of the variance in plant diversity (rocchini et al., 2014). in the present study, correlations between observed and predicted values were weak. for plant species richness, the strongest predictor was pca-var (r² = 0.19), which is lower than the spectral diversity indices (r² = 0.4–0.6) derived from airborne hyperspectral data (wang et al., 2018). two key limitations likely contribute to this weakness. first, strong noise from the desert background in the study area weakens the correlation between wv-3derived spectral indices and plant diversity. rocchini et al. (2010, 2017) noted that pixels should be at least as large as the sampling unit, particularly when using spectral heterogeneity to estimate local species diversity. additionally, the trade-off between noise from high resolution and information loss from low resolution must be considered. we speculate that wv-3’s fine resolution may exacerbate this issue: in vegetation-rich areas, numerous pixels may capture small shade patches cast by canopies (nagendra & rocchini, 2008), while in sparse vegetation, pixels are dominated by strong desert signals—both reducing the perceived strength of vegetation signals. second, vegetation indices like ndvi struggle to detect plant signals in desert-vegetation mosaics, making it difficult to extract vegetation texture from mixed spectra, especially with fixed thresholds used in preprocessing (hernández-stefanoni et al., 2014). in this study, an ndvi threshold of 0.2 was used to distinguish vegetation from desert, but this may have excluded small-stemmed plants. a study conducted in a subalpine grassland of the italian alps showed that spectral diversity calculated via the spectral angle mapper (sam) proved to be a more effective proxy for biodiversity within the same ecosystem—whereas the broader spectral diversity approach failed to estimate α-diversity (imran et al., 2024). this finding suggests that introducing additional novel spectral indices may enhance the accuracy of plant diversity estimation (cherif et al., 2023). in addition, remote sensing for assessing forest functional diversity has typically demonstrated high estimation accuracy (cimoli et al., 2024; zeng et al., 2023). for example, in a subtropical evergreen and deciduous broad-leaved mixed forest, airborne lidar-derived parameters were found to correlate well with in situ plot-level morphological data (r² ≥ 0.67) (zeng et al., 2023). currently, improving the estimation accuracy of both species and functional diversity in grasslands remains a key challenge for future research. conclusion we demonstrate that texture metrics derived from high-spatial-resolution satellite imagery effectively predict plot-level patterns of plant functional traits in sandy grasslands, highlighting their potential to capture environmental heterogeneity not detected by more conventional heterogeneity metrics. while positive correlations were observed between texture metrics and both plant species richness and functional diversity, these relationships varied across diversity indices and texture metrics. notably, for the functional index mfad, the relationship remained consistently significant and positive. our results underscore the complexity of the heterogeneity-diversity relationship, emphasizing the need for further investigation into these relationships using different satellite data sources and across diverse ecosystems. we conclude that texture metrics show promise as a tool for modeling functional diversity—rather than species diversity—in areas with sparse vegetation. acknowledgements we thank wn zhang, fc li, tt yang, j li, and c ma for their field survey. this research was funded by the general program of the national natural science foundation of china (32271555). data and code availability data will be made available on request. wen li et al. – remote sensing to assess functional diversity in sandy grasslands 95 competing interests the authors declare no conflict of interest. references asner, g. p., 1998. biophysical and biochemical sources of variability in canopy reflectance. remote sens. environ. 64(3): 23–253. https://doi.org/10.1016/s00344257(98)00014-5 asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., and liddell, m. j., 2009a. leaf chemical and spectral diversity in australian tropical forests. ecol. appl. 19: 236–253. https://doi.org/10.1890/08-0541.1 bawa, k., rose, j., ganeshaiah, k. n., barve, n., kiran, m. c., and umashanker, r., 2002. assessing biodiversity from space: example from western ghats, india. conserv. ecol. 6(1): 7. https://www.ecologyandsociety.org/vol6/iss1/art7/ cabacinha, c. d., and de castro, s. s., 2009. relationships between floristic diversity and vegetation indices, forest structure and landscape metrics of fragments in brazilian cerrado. for. ecol. manag. 257: 2157– 2165. https://doi.org/10.1016/j.foreco.2009.02.024 carlson, k. m., asner, g. p., hughes, r. f., ostertag, r., and martin, r. e., 2007. hyperspectral remote sensing of canopy biodiversity in hawaiian lowland rainforests. ecosystems 10: 536–549. https://doi. org/10.1007/s10021-006-9013-5 casanoves, f., pla, l., di rienzo, j. a., and díaz, s., 2011. fdiversity: a software package for the integrated analysis of functional diversity. methods ecol. evol. 2: 233–237. https://doi.org/10.1111/j.2041210x.2010.00088.x castillo-riffart, i., galleguillos, m., lopatin, j., and perez-quezada, a. j. f., 2017. predicting vascular plant diversity in anthropogenic peatlands: comparison of modeling methods with free satellite data. remote sens. 9(7): 681–692. https://doi.org/10.3390/ rs9070681 castro-esau, k. l., sánchez-azofeifa, g. a., and rivard, b., 2006b. comparison of spectral indices obtained using multiple spectroradiometers. remote sens. environ. 103: 276–288. https://doi.org/10.1016/j. rse.2006.05.005 ceccato, p., gobron, n., flasse, s., and asner, g. p., 2002. designing a spectral index to estimate vegetation water content from remote sensing data: part i: theoretical approach. remote sens. environ. 82(2–3): 188– 197. https://doi.org/10.1016/s0034-4257(02)00074-7 cerrejón, c., valeria, o., and fenton, n. j., 2023. estimating lichen αand β-diversity using satellite data at different spatial resolutions. ecol. indic. 149: 110173. https://doi.org/10.1016/j.ecolind.2023.110173 chawla, a., kumar, a., rajkumar, j., singh, r. d. r., thukral, a. k., and ahuja, p. s., 2010. correlation of multispectral satellite data with plant species diversity vis-à-vis soil characteristics in a landscape of western himalayan region, india. appl. remote sens. 1: 013503. https://doi.org/10.1117/1.3486043 cherif, e., feilhauer, h., berger, k., asner, g. p., martin, r. e., ollinger, s. v., wessman, c. a., and schmidtlein, s., 2023. from spectra to plant functional traits: transferable multi-trait models from heterogeneous and sparse data. remote sens. environ. 292: 113580. https://doi.org/10.1016/j.rse.2023.113580 cho, m. a., mathieu, r., asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., liddell, m. j., and hughes, r. f., 2012. mapping tree species composition in south african savannas using an integrated airborne spectral and lidar system. remote sens. environ. 125: 214–226. https://doi.org/10.1016/j.rse.2012.05.015 cimoli, e., lucieer, a., malenovsk, z., woodgate, w., janoutová, r., turner, d., asner, g. p., and martin, r. e., 2024. mapping functional diversity of canopy physiological traits using uas imaging spectroscopy. remote sens. environ. 302: 113958. https://doi. org/10.1016/j.rse.2024.113958 culbert, p. d., radeloff, v. c., st-louis, v., flather, c. h., rittenhouse, c. d., albright, t. p., asner, g. p., and martin, r. e., 2012. modeling broad-scale patterns of avian species richness across the midwestern united states with measures of satellite image texture. remote sens. environ. 118: 140–150. https://doi. org/10.1016/j.rse.2011.11.013 diaz, s., and cabido, m., 2001. vive la difference: plant functional diversity matters to ecosystem processes. trends ecol. evol. 16: 646–655. https://doi. org/10.1016/s0169-5347(01)02224-5 doğan, h. m., and doğan, m., 2020. understanding and modeling plant biodiversity of nallıhan (a3-ankara) forest ecosystem by means of geographic information systems and remote sensing. doctoral dissertation, ankara university, ankara, turkey. available from: https://openaccess.ankara.edu.tr/handle/20.500.12573/26455 duro, d. c., girard, j., king, d. j., fahrig, l., mitchell, s., lindsay, k., and tischendorf, l., 2014. predicting species diversity in agricultural environments using landsat tm imagery. remote sens. environ. 144: 214–225. https://doi.org/10.1016/j.rse.2013.12.013 wen li et al. – remote sensing to assess functional diversity in sandy grasslands 96 evju, m., blumentrath, s., skarpaas, o., stabbetorp, o. e., and sverdrup-thygeson, a., 2015. plant species occurrence in a fragmented grassland landscape, the importance of species traits. biodivers. conserv. 24(3): 547–561. https://doi.org/10.1007/s10531-014-0847-2 ewers, r. m., and didham, r. k., 2006. confounding factors in the detection of species responses to habitat fragmentation. biol. rev. 81: 117–142. https://doi. org/10.1017/s1464793105006950 farwell, l. s., elsen, p. r., razenkova, e., pidgeon, a. m., and radeloff, v. c., 2020. habitat heterogeneity captured by 30‐m resolution satellite image texture predicts bird richness across the united states. ecol. appl. 30(8): e02157. https://doi.org/10.1002/ eap.2157 fauvel, m., lopes, m., dubo, t., rivers-moore, j., frison, p.-l., gross, n., and ouin, a., 2020. prediction of plant diversity in grasslands using sentinel-1 and -2 satellite image time series. remote sens. environ. 237: 111536. https://doi.org/10.1016/j.rse.2019.111536 feilhauer, h., faude, u., and schmidtlein, s., 2011. combining isomap ordination and imaging spectroscopy to map continuous floristic gradients in a heterogeneous landscape. remote sens. environ. 115: 2513– 2524. https://doi.org/10.1016/j.rse.2011.05.020 feilhauer, h., thonfeld, f., faude, u., he, k., rocchini, d., and schmidtlein, s., 2013. assessing floristic composition with multispectral sensors: a comparison based on monotemporal and multiseasonal field spectra. int. j. appl. earth obs. geoinf. 21: 218–229. https://doi.org/10.1016/j.jag.2012.11.005 féret, j.-b., and asner, g. p., 2014. mapping tropical forest canopy diversity using high-fidelity imaging spectroscopy. ecol. appl. 24: 1289–1296. https://doi. org/10.1890/13-0897.1 ferreira, m. p., zortea, m., zanotta, d. c., shimabukuro, y. e., and filho, c. r. d. s., 2016. mapping tree species in tropical seasonal semi-deciduous forests with hyperspectral and multispectral data. remote sens. environ. 179: 66–78. https://doi.org/10.1016/j. rse.2016.04.025 gallardo-cruz, j. a., meave, j. a., gonzález, e. j., lebrija-trejos, e. e., romero-romero, m. a., asner, g. p., martin, r. e., and hughes, r. f., 2012. predicting tropical dry forest successional attributes from space: is the key hidden in image texture? plos one 7(2): e30506. https://doi.org/10.1371/journal. pone.0030506 gao, x., huete, a. r., ni, w. g., and miura, t., 2000. optical-biophysical relationships of vegetation spectra without background contamination. remote sens. environ. 74: 609–620. https://doi.org/10.1016/s00344257(00)00103-3 gholizadeh, h., amani, m., townsend, p. a., zygielbaum, a. i., helzer, c. j., asner, g. p., martin, r. e., and ollinger, s. v., 2019. detecting prairie biodiversity with airborne remote sensing. remote sens. environ. 221: 38–49. https://doi.org/10.1016/j.rse.2018.10.030 gillespie, t. w., foody, g. m., rocchini, d., giorgi, a. p., and saatchi, s., 2008. measuring and modeling biodiversity from space. prog. phys. geogr. 32: 203–221. https://doi.org/10.1177/0309133308094513 gould, w., 2000. remote sensing of vegetation, plant species richness, and regional biodiversity hotspots. ecol. appl. 10(6): 1861–1870. https://doi.org/10.189 0/1051-0761(2000)010[1861:rsovps]2.0.co;2 gu, c., liang, j., liu, x. y., asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., and liddell, m. j., 2024. application and prospects of hyperspectral remote sensing in monitoring plant diversity in grassland. chin. j. appl. ecol. 35(5): 1397–1407. https://doi.org /10.13287/j.1001-9332.202405.024 guisan, a., and thuiller, w., 2005. predicting species distribution: offering more than simple habitat models. ecol. lett. 8: 993–1009. https://doi.org/10.1111/ j.1461-0248.2005.00787.x hall-beyer, m., 2017. glcm texture: a tutorial v. 3.0, march 2017. university of calgary, calgary, alberta, canada. available from: https://www.ucalgary.ca/ hallbey/tutorial he, k. s., rocchini, d., neteler, m., and nagendra, h., 2011. benefits of hyperspectral remote sensing for tracking plant invasions. divers. distrib. 17: 381–392. https://doi.org/10.1111/j.1472-4642.2011.00794.x hector, a., and bagchi, r., 2007. biodiversity and ecosystem multifunctionality. nature 448: 188–190. https:// doi.org/10.1038/nature05987 henle, k., davies, k. f., kleyer, m., margules, c., and settele, j., 2004. predictors of species sensitivity to fragmentation. biodivers. conserv. 13: 207–251. https:// doi.org/10.1023/b:bioc.0000025823.16415.69 hernández-stefanoni, j. l., dupuy, j. m., johnson, d. k., birdsey, r., tun-dzul, f., peduzzi, a., asner, g. p., and martin, r. e., 2014. improving species diversity and biomass estimates of tropical dry forests using airborne lidar. remote sens. 6: 4741–4763. https:// doi.org/10.3390/rs6064741 higgins, s. i., lavorel, s., and revilla, e., 2003. estimating plant migration rates under habitat loss and fragmentation. oikos 101: 354–366. https://doi. org/10.1034/j.1600-0706.2003.12158.x wen li et al. – remote sensing to assess functional diversity in sandy grasslands 97 imran, h. a., sakowska, k., gianelle, d., rocchini, d., dalponte, m., scotton, m., and vescovo, l., 2024. assessing plant trait diversity as an indicator of species αand β-diversity in a subalpine grassland of the italian alps. remote sens. ecol. conserv. 10: 328– 342. https://doi.org/10.1002/rse2.370 john, r., chen, j., lu, n., guo, k., liang, c., wei, y., asner, g. p., and martin, r. e., 2008. predicting plant diversity based on remote sensing products in the semi-arid region of inner mongolia. remote sens. environ. 112(5): 2018–2032. https://doi.org/10.1016/j. rse.2007.12.014 kattenborn, t., fassnacht, f., pierce, s., lopatin, j., grime, j. p., and schmidtlein, s., 2017. linking plant strategies and plant traits derived by radiative transfer modelling. j. veg. sci. 28: 42–49. https://doi. org/10.1111/jvs.12533 kiran, g. s., and mudaliar, a., 2012. remote sensing & geo-informatics technology in evaluation of forest tree diversity. asian j. plant sci. res. 2(3): 237–242. https://doi.org/10.5958/2249-7429.2012.00041.1 kokaly, r. f., asner, g. p., ollinger, s. v., martin, m. e., and wessman, c. a., 2009. characterizing canopy biochemistry from imaging spectroscopy and its application to ecosystem studies. remote sens. environ. 113(s1): s78–s91. https://doi.org/10.1016/j. rse.2008.10.017 lelong, c. c. d., tshingomba, u. k., and soti, v., 2020. assessing worldview-3 multispectral imaging abilities to map the tree diversity in semi-arid parklands. int. j. appl. earth obs. geoinf. 93: 102211. https:// doi.org/10.1016/j.jag.2020.102211 levin, n., shmida, a., levanoni, o., tamari, h., and kark, s., 2007. predicting mountain plant richness and rarity from space using satellite-derived vegetation indices. divers. distrib. 13: 692–703. https://doi. org/10.1111/j.1472-4642.2007.00366.x lindborg, r., helm, a., bommarco, r., heikkinen, r. k., kuhn, i., asner, g. p., martin, r. e., and ford, a. j., 2012. effect of habitat area and isolation on plant trait distribution in european forests and grasslands. ecography 35: 356–363. https://doi.org/10.1111/j.16000587.2011.06962.x lopatin, j., fassnacht, f. e., kattenborn, t., and schmidtlein, s., 2017. mapping plant species in mixed grassland communities using close range imaging spectroscopy. remote sens. environ. 201: 12–23. https:// doi.org/10.1016/j.rse.2017.05.034 lucas, k., and carter, g., 2008. the use of hyperspectral remote sensing to assess vascular plant species richness on horn island, mississippi. remote sens. environ. 112: 3908–3915. https://doi.org/10.1016/j. rse.2008.06.014 magurran, a. e., 2004. measuring biological diversity. oxford: blackwell publishing. https://doi. org/10.1002/9780470996267 mapfumo, r. b., asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., liddell, m. j., hughes, r. f., and ostertag, r., 2016. the relationship between satellite-derived indices and species diversity across african savanna ecosystems. int. j. appl. earth obs. geoinf. 52: 306–317. https://doi.org/10.1016/j.jag.2016.07.013 mpakairi, k. s., dube, t., dondofema, f., asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., and liddell, m. j., 2022. spatial characterization of vegetation diversity in groundwater dependent ecosystems using in-situ and sentinel-2 msi satellite data. remote sens. 14: 2995. https://doi.org/10.3390/rs14132995 nagendra, h., 2001. using remote sensing to assess biodiversity. int. j. remote sens. 22: 2377–2400. https:// doi.org/10.1080/01431160110043459 nagendra, h., and rocchini, d., 2008. satellite imagery applied to biodiversity study in the tropics: the devil is in the detail. biodivers. conserv. 17: 3431–3442. https://doi.org/10.1007/s10531-008-9457-0 oldeland, j., wesuls, d., rocchini, d., schmidt, m., and jürgens, n., 2010. does using species abundance data improve estimates of species diversity from remotely sensed spectral heterogeneity? ecol. indic. 10: 390– 396. https://doi.org/10.1016/j.ecolind.2009.11.006 oldeland, j., dorigo, w., wesuls, d., and jürgens, n., 2010. mapping bush encroaching species by seasonal differences in hyperspectral imagery. remote sens. 2: 1416–1438. https://doi.org/10.3390/rs2071416 ollinger, s. v., 2011. sources of variability in canopy reflectance and the convergent properties of plants. new phytol. 189: 375–394. https://doi.org/10.1111/j.14698137.2010.03527.x palmer, m. w., earls, p. g., hoagland, b. w., white, p. s., and wohlgemuth, t., 2002. quantitative tools for perfecting species lists. environmetrics 13: 121–137. https://doi.org/10.1002/env.507 pacheco-labrador, j., migliavacca, m., ma, x., mahecha, m., carvalhais, n., weber, u., asner, g. p., and martin, r. e., 2022. challenging the link between functional and spectral diversity with radiative transfer modeling and data. remote sens. environ. 280: 113170. https://doi.org/10.1016/j.rse.2022.113170 peng, y., fan, m., bai, l., sang, w., feng, j., zhao, z., and tao, z., 2019. identification of the best hyperwen li et al. – remote sensing to assess functional diversity in sandy grasslands 98 spectral indices in estimating plant species richness in sandy grasslands. remote sens. 11: 588. https://doi. org/10.3390/rs11050588 radhamoni, h. v. n., queenborough, s. a., arietta, a. z. a., asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., and liddell, m. j., 2023. localand landscape-scale drivers of terrestrial herbaceous plant diversity along a tropical rainfall gradient in western ghats, india. j. ecol. 111(5): 1021–1036. https://doi. org/10.1111/1365-2745.14010 rocchini, d., dadalt, l., delucchi, l., neteler, m., and palmer, m. w., 2014. disentangling the role of remotely sensed spectral heterogeneity as a proxy for north american plant species richness. community ecol. 15(1): 37–43. https://doi.org/10.1080/1745097 9.2014.881461 rocchini, d., 2007. effects of spatial and spectral resolution in estimating ecosystem α-diversity by satellite imagery. remote sens. environ. 111: 423–434. https://doi.org/10.1016/j.rse.2007.05.014 rocchini, d., balkenhol, n., carter, g. a., foody, g. m., gillespie, t. w., he, k., asner, g. p., and martin, r. e., 2010. remotely sensed spectral heterogeneity as a proxy for species diversity: recent advances and open challenges. ecol. inform. 5: 318–329. https:// doi.org/10.1016/j.ecoinf.2010.03.004 rocchini, d., chiarucci, a., and loiselle, s. a., 2004. testing the spectral variation hypothesis by using satellite multispectral images. acta oecol. 26: 117–120. https://doi.org/10.1016/j.actao.2004.03.005 rocchini, d., marcantonio, m., and riccota, c., 2017. measuring rao’s q diversity index from remote sensing: an open source solution. ecol. indic. 72: 234– 238. https://doi.org/10.1016/j.ecolind.2016.10.021 rocchini, d., ricotta, c., and chiarucci, a., 2007. using satellite imagery to assess plant species richness: the role of multispectral systems. appl. veg. sci. 10: 325–331. https://doi.org/10.1111/j.1654109x.2007.00316.x roth, k. l., casas, a., huesca, m., ustin, s. l., alsina, m. m., mathews, s. a., and whiting, m. l., 2016. leaf spectral clusters as potential optical leaf functional types within california. remote sens. environ. 184: 229–246. https://doi.org/10.1016/j.rse.2016.08.015 schmidtlein, s., feilhauer, h., and bruelheide, h., 2012. mapping plant strategy types using remote sensing. j. veg. sci. 23: 395–405. https://doi.org/10.1111/j.1654109x.2011.01230.x schneider, f. d., morsdorf, f., schmid, b., asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., and liddell, m. j., 2017. mapping functional diversity from remotely sensed morphological and physiological forest traits. nat. commun. 8: 1441. https://doi.org/10.1038/ s41467-017-01532-8 schowengerdt, r., 2007. remote sensing: models and methods for image processing. oxford: elsevier. 515 p. https://doi.org/10.1016/b978-012374372-1.00001-1 socolar, j. b., gilroy, j. j., kunin, w. e., asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., and liddell, m. j., 2016. how should beta-diversity inform biodiversity conservation? trends ecol. evol. 31(1): 67–80. https://doi.org/10.1016/j.tree.2015.10.004 somers, b., asner, g. p., martin, r. e., anderson, c. b., knapp, d. e., wright, s. j., ostertag, r., and hughes, r. f., 2015. mesoscale assessment of changes in tropical tree species richness across a bioclimatic gradient in panama using airborne imaging spectroscopy. remote sens. environ. 167: 111–120. https://doi. org/10.1016/j.rse.2015.06.032 somers, b., verbesselt, j., ampe, e. m., sims, n., verstraeten, w. w., and coppin, p., 2010b. spectral mixture analysis to monitor defoliation in mixed aged eucalyptus globulus labill plantations in southern australia using landsat 5tm and eo-1 hyperion data. int. j. appl. earth obs. geoinf. 12: 270–277. https://doi.org/10.1016/j.jag.2010.01.006 spasojevic, m. j., grace, j. b., harrison, s., and damschen, e. i., 2014. functional diversity supports the physiological tolerance hypothesis for plant species richness along climatic gradients. j. ecol. 102: 447– 455. https://doi.org/10.1111/1365-2745.12214 sun, w., liu, w., wang, y., asner, g. p., martin, r. e., ford, a. j., metcalfe, d. j., and liddell, m. j., 2023. progress and prospects of global hyperspectral remote sensing research on wetlands from 2010 to 2022. j. remote sens. 27(6): 1281–1299. https://doi. org/10.11834/jrs.20230374 tuanmu, m. n., and jetz, m., 2015. a global, remote sensing-based characterization of terrestrial habitat heterogeneity for biodiversity and ecosystem modeling. global ecol. biogeogr. 24: 1329–1339. https://doi. org/10.1111/geb.12316 vaglio laurin, g., chan, j. c.-w., chen, q., lindsell, j. a., coomes, d. a., asner, g. p., martin, r. e., and ford, a. j., 2014. biodiversity mapping in a tropical west african forest with airborne hyperspectral data. plos one 9(6): e97910. https://doi.org/10.1371/ journal.pone.0097910 viedma, o., torres, i., pérez, b., and moreno, j. m., 2012. modeling plant species richness using reflectance and texture data derived from quickbird in a recently burned area of central spain. remote sens. wen li et al. – remote sensing to assess functional diversity in sandy grasslands 99 environ. 119: 208–221. https://doi.org/10.1016/j. rse.2011.12.015 wang, r., gamon, j. a., montgomery, r. a., and townsend, p. a., 2016a. seasonal variation in the ndvi–species richness relationship in a prairie grassland experiment (cedar creek). remote sens. 8: 128– 131. https://doi.org/10.3390/rs8020128 wang, r., gamon, j. a., cavender-bares, j., townsend, p. a., and zygielbaum, a. i., 2018. the spatial sensitivity of the spectral diversity-biodiversity relationship: an experimental test in a prairie grassland. ecol. appl. 28: 541–556. https://doi.org/10.1002/eap.1724 white, j. c., gómez, c., wulder, m. a., and coops, n. c., 2010. characterizing temperate forest structural and spectral diversity with hyperion eo-1 data. remote sens. environ. 114: 1576–1589. https://doi. org/10.1016/j.rse.2010.02.018 wood, e. m., pidgeon, a. m., radeloff, v. c., and keuler, n. s., 2012. image texture as a remotely sensed measure of vegetation structure. remote sens. environ. 121: 516–526. https://doi.org/10.1016/j. rse.2012.02.020 wood, e. m., pidgeon, a. m., radeloff, v. c., and keuler, n. s., 2013. image texture predicts avian density and species richness. plos one 8(5): e63211. https://doi. org/10.1371/journal.pone.0063211 xin, j. x., li, j. n., zeng, q. q., peng, y., asner, g. p., martin, r. e., ford, a. j., and liddell, m. j., 2024. high-precision estimation of plant alpha diversity in different ecosystems based on sentinel-2 data. ecol. indic. 166: 112527. https://doi.org/10.1016/j. ecolind.2024.112527 zhang, y. w., guo, y. p., tang, r., tang, z. y., asner, g. p., martin, r. e., ford, a. j., and liddell, m. j., 2023. progress and trends of application of hyperspectral remote sensing in plant diversity research. natl. remote sens. bull. 27(11): 2467–2483. https://doi. org/10.11834/jrs.20231048 zheng, z., schmid, b., zeng, y., schuman, m. c., zhao, d., schaepman, m. e., asner, g. p., and martin, r. e., 2023. remotely sensed functional diversity and its association with productivity in a subtropical forest. remote sens. environ. 290: 113530. https://doi. org/10.1016/j.rse.2023.113530 wen li et al. – remote sensing to assess functional diversity in sandy grasslands 100 appendix table s1 coverage, habitat type, plant diversity, and dominant species in study plots plotid coverage altitude habitat type richness h d e dominant species 2 4 1331.0 semi-fixed 3 0.80 0.45 0.72 salsola collina 3 14 1330.0 semi-fixed 3 0.76 0.43 0.70 ulmus pumila 5 53 1334.0 semi-fixed 4 1.13 0.63 0.81 polygonum divaricatum 6 11 1329.0 semi-fixed 5 1.09 0.53 0.67 salsola collina 7 26 1328.0 semi-fixed 3 0.78 0.45 0.71 polygonum divaricatum 10 6 1329.0 semi-fixed 4 0.96 0.51 0.69 salsola collina 12 11 1330.0 semi-fixed 4 1.28 0.69 0.92 salsola collina 15 17 1323.0 semi-fixed 5 1.46 0.74 0.91 polygonum divaricatum 16 26 1323.4 semi-fixed 8 1.71 0.78 0.82 potentilla supina 17 9 1326.0 semi-fixed 5 1.08 0.56 0.67 salsola collina 18 20 1329.1 semi-fixed 5 1.46 0.74 0.91 lappula myosotis 19 25 1328.0 semi-fixed 4 1.24 0.69 0.89 polygonum divaricatum 20 8 1321.0 semi-fixed 5 1.07 0.53 0.66 echinops davuricus 21 8 1324.0 semi-fixed 4 0.77 0.39 0.56 salsola collina 22 5 1322.0 semi-fixed 4 1.05 0.58 0.76 agropyron mongolicum 23 4 1323.7 semi-fixed 3 1.05 0.64 0.96 potentilla supina 24 22 1324.0 semi-fixed 5 1.43 0.74 0.89 polygonum divaricatum 25 11 1322.0 semi-fixed 6 1.59 0.77 0.89 phragmites australis 26 40 1321.0 semi-fixed 5 1.45 0.74 0.90 bromus ircutensis 27 37 1321.0 semi-fixed 4 1.08 0.62 0.78 potentilla supina 28 12 1326.4 semi-fixed 5 0.94 0.46 0.59 leymus chinensis 29 27 1325.0 semi-fixed 4 1.27 0.69 0.92 polygonum divaricatum 30 13 1326.0 semi-shift 3 0.96 0.59 0.87 artemisia ordosica 31 15 1327.4 semi-shift 4 1.05 0.60 0.76 artemisia ordosica 32 22 1330.0 semi-shift 2 0.68 0.49 0.98 artemisia ordosica 33 10 1326.8 semi-shift 3 0.63 0.34 0.57 leymus secalinus 34 24 1329.0 semi-shift 3 0.93 0.57 0.84 artemisia ordosica 35 14 1325.0 semi-shift 4 1.26 0.69 0.91 artemisia ordosica 36 13 1323.3 semi-shift 5 1.26 0.66 0.79 artemisia ordosica 37 10 1324.0 semi-shift 4 1.21 0.67 0.87 agropyron mongolicum 38 11 1323.6 semi-shift 4 1.24 0.68 0.89 artemisia ordosica 39 15 1321.3 semi-shift 3 0.95 0.56 0.86 polygonum divaricatum 40 32 1324.0 semi-shift 2 0.64 0.44 0.92 artemisia ordosica 42 29 1323.0 semi-shift 4 1.27 0.69 0.92 polygonum divaricatum 43 9 1322.0 semi-shift 6 1.37 0.66 0.76 agropyron mongolicum 44 8 1322.0 semi-shift 3 0.74 0.45 0.67 salsola collina 45 14 1319.0 lowland 5 1.49 0.75 0.93 agropyron mongolicum 46 25 1322.5 lowland 4 1.00 0.53 0.72 leymus secalinus 47 25 1319.2 lowland 6 1.38 0.69 0.77 leymus secalinus 48 26 1319.0 lowland 6 1.31 0.61 0.73 agropyron mongolicum wen li et al. – remote sensing to assess functional diversity in sandy grasslands 101 49 27 1317.0 lowland 3 0.82 0.53 0.74 carex duriuscula 50 16 1320.0 lowland 6 1.68 0.80 0.94 salsola collina 51 45 1319.7 lowland 5 1.27 0.67 0.79 carex duriuscula 52 19 1319.0 lowland 6 1.46 0.71 0.81 agropyron mongolicum 53 34 1320.6 lowland 7 1.17 0.54 0.60 agropyron mongolicum 54 39 1322.0 lowland 5 1.44 0.73 0.89 carex duriuscula 55 10 1320.0 lowland 5 1.38 0.70 0.85 artemisia ordosica 56 85 1326.0 lowland 7 1.55 0.73 0.80 agropyron mongolicum 57 41 1318.7 lowland 6 1.27 0.65 0.71 carex duriuscula 58 20 1320.0 lowland 4 1.06 0.59 0.77 agropyron mongolicum 59 25 1319.4 lowland 4 0.86 0.50 0.62 agropyron mongolicum 60 38 1334.0 fixed-land 4 1.03 0.61 0.74 leymus chinensis 61 78 1328.0 fixed-land 2 0.28 0.15 0.40 leymus secalinus 62 23 1329.0 fixed-land 2 0.53 0.35 0.76 leymus chinensis 63 24 1329.3 fixed-land 5 1.33 0.70 0.83 agropyron mongolicum 64 16 1324.5 fixed-land 4 1.21 0.67 0.87 corispermum stauntonii 65 18 1326.4 fixed-land 4 1.15 0.64 0.83 corispermum stauntonii 66 14 1326.7 fixed-land 3 1.06 0.64 0.96 artemisia ordosica 67 31 1327.1 fixed-land 5 1.32 0.68 0.82 carex duriuscula 68 23 1326.1 semi-fixed 4 1.18 0.65 0.85 artemisia ordosica 69 13 1325.0 fixed-land 4 1.12 0.61 0.81 corispermum stauntonii 70 19 1324.4 fixed-land 4 1.08 0.61 0.78 artemisia ordosica 71 43 1326.3 fixed-land 5 1.04 0.50 0.65 carex duriuscula 72 42 1330.2 semi-fixed 4 1.14 0.65 0.82 artemisia ordosica 73 17 1327.6 semi-fixed 3 0.92 0.57 0.84 artemisia ordosica 74 32 1327.4 semi-fixed 7 1.68 0.79 0.86 agropyron mongolicum 75 28 1328.1 semi-fixed 6 1.61 0.77 0.90 artemisia ordosica 76 30 1325.5 semi-fixed 7 1.90 0.84 0.98 artemisia ordosica 77 22 1325.9 semi-fixed 3 0.51 0.26 0.46 artemisia ordosica 78 23 1324.4 semi-fixed 4 1.22 0.68 0.88 artemisia ordosica 79 15 1325.0 semi-fixed 5 1.20 0.60 0.75 chenopodium glaucum 80 22 1324.4 semi-fixed 5 1.33 0.68 0.83 agropyron mongolicum 81 8 1325.0 semi-fixed 4 0.99 0.54 0.71 agropyron mongolicum 82 23 1324.1 semi-fixed 2 0.48 0.30 0.70 artemisia ordosica 83 18 1324.0 semi-fixed 5 1.49 0.75 0.93 artemisia ordosica 84 35 1319.0 fixed-land 8 1.48 0.68 0.71 carex duriuscula 85 22 1322.3 fixed-land 4 1.18 0.64 0.85 agropyron mongolicum 86 27 1322.0 fixed-land 5 1.39 0.72 0.87 artemisia ordosica 87 32 1324.0 fixed-land 4 1.07 0.59 0.77 artemisia ordosica 88 16 1324.0 fixed-land 5 1.47 0.74 0.91 artemisia ordosica 89 17 1325.0 semi-fixed 5 1.39 0.72 0.87 artemisia ordosica 90 15 1324.7 semi-fixed 3 1.04 0.63 0.94 artemisia ordosica 91 6 1326.0 fixed-land 4 1.39 0.75 1.00 ulmus pumila 92 33 1322.3 fixed-land 4 1.23 0.68 0.89 agropyron mongolicum 93 19 1337.0 lowland 5 1.30 0.67 0.81 leymus chinensis 94 45 1335.0 lowland 8 1.63 0.74 0.79 cleistogenes caespitosa 95 20 1330.0 lowland 5 1.49 0.75 0.92 agropyron cristatum wen li et al. – remote sensing to assess functional diversity in sandy grasslands 102 96 58 1327.0 lowland 4 1.07 0.61 0.77 cleistogenes caespitosa 97 19 1327.7 lowland 5 1.27 0.65 0.79 asparagus schoberioides 98 36 1328.0 lowland 8 1.65 0.75 0.79 asparagus schoberioides 99 29 1327.0 lowland 8 1.80 0.80 0.87 cleistogenes caespitosa 100 37 1327.0 lowland 6 1.18 0.63 0.66 cleistogenes caespitosa 101 28 1329.0 lowland 5 1.28 0.66 0.80 thymus serpyllum 102 36 1327.0 lowland 5 1.42 0.72 0.88 artemisia frigida 103 19 1326.0 lowland 5 1.32 0.69 0.82 agropyron mongolicum 104 28 1329.0 lowland 6 1.64 0.79 0.92 artemisia ordosica 105 17 1326.0 lowland 5 1.36 0.70 0.84 cleistogenes caespitosa 106 34 1326.0 lowland 4 0.89 0.53 0.64 artemisia frigida 107 37 1326.0 lowland 5 1.52 0.76 0.94 agropyron mongolicum 108 13 1336.0 fixed-land 5 1.54 0.78 0.96 leymus chinensis 109 8 1332.4 fixed-land 4 1.04 0.60 0.75 leymus chinensis 110 5 1326.2 fixed-land 4 1.19 0.66 0.86 leymus chinensis 111 7 1327.0 fixed-land 5 1.44 0.73 0.89 leymus chinensis 112 17 1325.6 fixed-land 5 0.90 0.43 0.56 carex duriuscula 113 7 1327.4 fixed-land 4 1.21 0.67 0.87 leymus chinensis 114 20 1326.7 fixed-land 4 0.94 0.48 0.68 lappula myosotis 115 6 1326.7 fixed-land 4 1.29 0.70 0.93 leymus chinensis 116 5 1324.0 fixed-land 4 1.21 0.66 0.88 leymus chinensis 117 17 1326.0 fixed-land 4 0.45 0.19 0.32 carex duriuscula 118 15 1326.0 fixed-land 3 0.91 0.54 0.83 cleistogenes caespitosa 119 20 1325.0 fixed-land 4 1.30 0.71 0.94 thymus serpyllum 120 11 1325.0 fixed-land 4 1.03 0.55 0.75 carex duriuscula 121 12 1324.6 fixed-land 3 0.90 0.53 0.82 lappula myosotis 122 29 1326.0 fixed-land 6 0.92 0.45 0.51 carex duriuscula wen li et al. – remote sensing to assess functional diversity in sandy grasslands 103 table s2 t he nam e, form ula, and ecological m eanings of spectral indices used in the present study o rder n am e form ula ecological m eanings r eference 1 a nthocyanin r eflectance index 1 (a r i1) one m easure of stressed vegetation g itelson et al., 2001 2 a nthocyanin r eflectance index 2 (a r i2) is a m odification to the a r i1 that detects higher concentrations of anthocyanins in vegetation. liu et al., 2022 3 a tm ospherically r esistant vegetation index (a rv i) an enhancem ent to the n d v i that is relatively resistant to atm ospheric factors k aufm an et al., 1992 4 enhanced vegetation index (ev i) an im provem ent over n d v i by optim izing the vegetation signal in areas of high leaf area index h uete et al., 2002 5 g reen a tm ospherically r esistant index (g a r i) is m ore sensitive to a w ide range of chlorophyll concentrations and less sensitive to atm ospheric effects than n d v i. g itelson et al., 1996 6 g reen c hlorophyll index (g c i) to estim ate leaf chlorophyll content across a w ide range of plant species. g itelson et al., 2003 7 g reen d ifference vegetation index (g d v i) this index w as originally designed w ith color-infrared photography to predict nitrogren requirem ents for corn. sripada et al., 2006 8 g reen leaf index (g li) to m easure w heat cover, w here the red, green, and blue digital num bers (d n s) range from 0 to 255. louhaichi et al., 2001 9 g reen n orm alized d ifference vegetation index (g n d v i) this index w as originally designed w ith color-infrared photography to predict nitrogren requirem ents for corn. this index is m ore sensitive to chlorophyll concentration than n d v i. g itelson et al., 1998 wen li et al. – remote sensing to assess functional diversity in sandy grasslands 104 10 g reen o ptim ized soil a djusted vegetation index (g o sav i) this index is sim ilar to n d v i except that it m easures the green spectrum from 540 to 570 nm instead of the red spectrum . this index is m ore sensitive to chlorophyll concentration than n d v i. sripada et al., 2006 11 g reen r atio vegetation index (g rv i) sensitive to photosynthetic rates in forest canopies sripada et al., 2006 12 g reen soil a djusted vegetation index (g sa v i) this index w as originally designed w ith color-infrared photography to predict nitrogren requirem ents for corn. it is sim ilar to sav i, but it substitutes the green band for red. sripada et al., 2006 13 g reen vegetation index (g v i) this index m inim izes the effects of background soil w hile em phasizing green 14vegetation k authet al., 1976 14 g lobal environm ental m onitoring index (g em i) is sim ilar to n d v i but is less sensitive to atm ospheric effects. pinty et al., 1992 15 infrared percentage vegetation index (ipv i) like n d v i, is com putationally faster. values range from 0 to 1. c rippen et al., 1990 16 m odified c hlorophyll a bsorption r atio index (m c a r i) this index indicates the relative abundance of chlorophyll.it is designed prim arily to am plify the leaf-level chlorophyll signal, thereby indicating vegetation “physiological health and nutritional status,” rather than canopy structure or biom ass. zarco-tejada et al., 2005 https://envi.geoscene.cn/help/subsystems/envi/content/vegetation analysis/broadbandgreenness.htm#soil wen li et al. – remote sensing to assess functional diversity in sandy grasslands 105 17 m odified c hlorophyll a bsorption r atio index im proved (m c a r i2) this index can m ore accurately and consistently characterize the canopy’s “green biom ass” and its “gross prim ary productivity potential.” it rem ains sensitive to variations in leaf area index even in dense, high-biom ass stands w here n d v i tends to saturate, and it m inim izes the influence of changing chlorophyll concentrations h aboudane et al., 2004 18 m odified n on-linear index (m n li) this index is an enhancem ent to the n on-linear index (n li) that incorporates the soil a djusted vegetation index (sav i) yang et al., 2021 19 m odified soil a djusted vegetation index 2 (m sav i2) it reduces soil noise and increases the dynam ic range of the vegetation signal. q i et al., 1994 20 m odified triangular vegetation index (m tv i) suitable for la i estim ations h aboudane et al., 2004 21 m odified triangular vegetation index im proved (m tv i2) a better predictor of green la i eitelet al., 2007 22 n orm alized d ifference m ud index (n d m i) this index highlights m uddy or shallow w ater pixels liu et al., 2022 23 o ptim ized soil a djusted vegetation index (o sav i) is best used in areas w ith relatively sparse vegetation peng et al., 2018 24 pc a -cv c oeffi cients of variation in principle com ponents 25 r enorm alized d ifference vegetation index (r d v i) to highlight healthy vegetation r oujean et al., 1995 26 r ed g reen r atio index (r g r i) indicates the relative expression of leaf redness caused by anthocyanin to that of chlorophyll. g am on et al., 1999 27 soil a djusted vegetation index (sav i) this index is best used in areas w ith relatively sparse vegetation h uete et al., 1988 wen li et al. – remote sensing to assess functional diversity in sandy grasslands 106 28 sum g reen index (sg i) sg i is the m ean of reflectance across the 500 nm to 600 nm portion of the spectrum . this index is used for detecting changes in vegetation greenness. lobell et al., 2003 29 transform ed c hlorophyll a bsorption r eflectance index (tc a r i) this index indicates the relative abundance of chlorophyll.it is highly sensitive to chlorophyll concentration; by am plifying the chlorophyll absorption trough, it rapidly and quantitatively reflects vegetation nitrogen status and early physiological stress. h aboudane et al., 2004 30 transform ed d ifference vegetation index (td v i) for m onitoring vegetation cover in urban environm ents. b annari et al., 2002 31 triangular g reenness index (tg i) it is suitable for r g b cam eras. r aym ond et al., 2011 32 v isible a tm ospherically r esistant index (va r i) to estim ate the fraction of vegetation in a scene g itelson et al., 2002 33 w ide d ynam ic r ange vegetation index (w d r v i) is m ore sensitive to a w ider range of vegetation fractions and to changes in la i. g itelso n et al., 2004 34 w orldview im proved vegetative index(w iv i) the com m on range for green vegetation. w olf et al., 2010 35 w orldview n on-h om ogeneous feature d ifference (w n h fd ) to identify features that contrast highly against the background. w olf et al., 2010 wen li et al. – remote sensing to assess functional diversity in sandy grasslands 107 reference bannari, a., asalhi, h., and teillet, p. m., 2002. transformed difference vegetation index (tdvi) for vegetation cover mapping. proc. ieee int. geosci. remote sens. symp. (igarss 2002), toronto, on, canada, 24–28 june 2002, vol. 5: 3053–3055. crippen, r. e., 1990. calculating the vegetation index faster. remote sens. environ. 34(1): 71–73. https://doi.org/10.1016/0034-4257(90)90009-3 eitel, j. u. h., long, d. s., gessler, p. e., and smith, a. c., 2007. using in-situ measurements to evaluate the new rapideye™ satellite series for prediction of wheat nitrogen status. remote sens. environ. 108(2–3): 4183–4190. https://doi. org/10.1016/j.rse.2007.08.013 gamon, j. a., surfus, j. s., 1999. assessing leaf pigment content and activity with a reflectometer. new phytol. 143: 105–117. https://doi. org/10.1046/j.1469-8137.1999.00480.x gitelson, a. a., merzlyak, m., 1998. remote sensing of chlorophyll concentration in higher plant leaves. adv. space res. 22(5): 689–692. https:// doi.org/10.1016/s0273-1177(98)00144-8 gitelson, a. a., gritz, y., merzlyak, m. n., 2003. relationships between leaf chlorophyll content and spectral reflectance and algorithms for non-destructive chlorophyll assessment in higher plant leaves. plant physiol. 160(3): 271–282. https:// doi.org/10.1071/pp02082 gitelson, a. a., kaufman, y. j., merzlyak, m. n., 1996. use of a green channel in remote sensing of global vegetation from eos-modis. remote sens. environ. 58(3): 289–298. https://doi. org/10.1016/0034-4257(96)00042-7 gitelson, a. a., merzlyak, m. n., chivkunova, o. b., 2001. optical properties and nondestructive estimation of anthocyanin content in plant leaves. photochem. photobiol. 74(1): 38–45. https://doi. org/10.1562/0031-8655(2001)074[0038:opaneo]2.0.co;2 gitelson, a. a., 2004. wide dynamic range vegetation index for remote quantification of biophysical characteristics of vegetation. j. plant physiol. 161(2): 165–173. https://doi.org/10.1078/01761617-01084 haboudane, d., miller, j. r., pattey, e., and smith, a. c., 2004. hyperspectral vegetation indices and novel algorithms for predicting green lai of crop canopies: modeling and validation in the context of precision agriculture. remote sens. environ. 90(3): 337–352. https://doi. org/10.1016/j.rse.2004.01.011 huete, a. r., didan, k., miura, t., and smith, a. c., 2002. overview of the radiometric and biophysical performance of the modis vegetation indices. remote sens. environ. 83: 195–213. https:// doi.org/10.1016/s0034-4257(02)00096-2 huete, a. r., 1988. a soil-adjusted vegetation index (savi). remote sens. environ. 25(3): 295–309. https://doi.org/10.1016/0034-4257(88)90106-x kaufman, y. j., tanré, d., 1992. atmospherically resistant vegetation index (arvi) for eos-modis. ieee trans. geosci. remote sens. 30(2): 261–270. https://doi.org/10.1109/36.129241 kauth, r. j., 1976. tasselled cap—a graphic description of the spectral-temporal development of agricultural crops as seen by landsat. lars symp.: 41–51. https://ageconsearch.umn.edu/record/222424 liu, l. x., pang, y., sang, g. q., and smith, a. c., 2022. applying high resolution remote sensing to access tree species diversity of monsoonal broad-leaved evergreen forest in puer city. acta ecol. sin. 42(20): 8398–8410. https://doi. org/10.5846/stxb202106231710 lobell, d., asner, g., 2003. hyperion studies of crop stress in mexico. proc. 12th annu. jpl airborne earth sci. workshop, pasadena, ca. https:// earth.jpl.nasa.gov/files/airs2003_abstracts.pdf louhaichi, m., borman, m. m., johnson, d. e., 2001. spatially located platform and aerial photography for documentation of grazing impacts on wheat. geocarto int. 16: 65–70. https:// doi.org/10.1080/10106040108544267 peng, y., fan, m., song, j. y., and smith, a. c., 2018. assessment of plant species diversity based on hyperspectral indices at a fine scale. sci. rep. 8: 4776. https://doi.org/10.1038/s41598-01823183-5 pinty, b., verstraete, m., 1992. gemi: a non-linear index to monitor global vegetation from satellites. vegetatio 101(1): 15–20. https://doi. org/10.1007/bf00044317 qi, j., chehbouni, a., huete, a. r., kerr, y. h., and sorooshian, s., 1994. a modified soil adjusted vegetation index. remote sens. environ. 48(2): 119–126. https://doi.org/10.1016/00344257(94)90134-1 raymond hunt, e., daughtry, c. s. t., eitel, j. u. h., and smith, a. c., 2011. remote sensing leaf chlorophyll content using a visible band index. agron. j. 103(4): 1090–1099. https://doi. org/10.2134/agronj2010.0477 roujean, j.-l., breon, f.-m., 1995. estimating par absorbed by vegetation from bidirectional reflectance measurements. remote sens. environ. wen li et al. – remote sensing to assess functional diversity in sandy grasslands 108 51(3): 375–384. https://doi.org/10.1016/00344257(95)00048-4 sripada, r. p., heiniger, r. w., white, j. g., and smith, a. c., 2006. aerial color infrared photography for determining early in-season nitrogen requirements in corn. agron. j. 98(4): 968–977. https://doi.org/10.2134/agronj2005.0146 wolf, a., 2010. using worldview 2 vis-nir msi imagery to support land mapping and feature extraction using normalized difference index ratios. unpublished report, longmont, co: digitalglobe. available from: https://www. maxar.com/sites/default/files/documents/worldview-2-imagery-land-mapping.pdf yang, z., willis, p., and mueller, r., 2008. impact of band-ratio enhanced awifs image on crop classification accuracy. proc. 2008 asprs annu. conf., portland, or, usa, 28 april–2 may 2008. https://www.asprs.org/wp-content/uploads/pers/ 2008conference/proceedings/0295.pdf zarco-tejada, p. j., ustin, s., whiting, m., 2005. temporal and spatial relationships between within-field yield variability in cotton and high-spatial hyperspectral remote sensing imagery. agron. j. 97: 641–653. https://doi.org/10.2134/ agronj2004.0148 microsoft word kelly_cal_5.docx biodiversity informatics, 11, 2016, pp. 40-62 40 rescuing and sharing historical vegetation data for ecological analysis: the california vegetation type mapping project maggi kelly1,2,3,4, kelly easterday1, giovanni rapacciuolo5, michelle s. koo4,6, patrick mcintyre7, james thorne8 1134 mulford hall #3114, department of environmental sciences, policy and management, university of california, berkeley. berkeley, ca 94720-3114. 2geospatial innovation facility, university of california. berkeley, ca 94720. 3university of california division of agriculture and natural resources. 4berkeley initiative in global change biology, university of california. berkeley, ca 94720. 5department of ecology and evolution, 650 life sciences building, stony brook university, stony brook, ny 11789. 6museum of vertebrate zoology, 3101 valley life sciences, university of california, berkeley, ca 94720-3160. 7biogeographic data branch, california department of fish and wildlife, 1416 9th street, suite 1266, sacramento, ca 95814. 8department of environmental science and policy, university of california, davis, ca 95616. maggi kelly corresponding author: maggi@berkeley.edu. abstract.—research efforts that synthesize historical and contemporary ecological data with modeling approaches improve our understanding of the complex response of species, communities, and landscapes to changing biophysical conditions through time and in space. historical ecological data are particularly important in this respect. there are remaining barriers that limit such data synthesis, and technological improvements that make multiple diverse datasets more readily available for integration and synthesis are needed. this paper presents one case study of the wieslander vegetation type mapping project in california and highlights the importance of rescuing, digitizing and sharing historical datasets. we review the varied ecological uses of the historical collection: the vegetation maps have been used to understand legacies of land use change and plan for the future; the plot data have been used to examine changes to chaparral and forest communities around the state and to predict community structure and shifts under a changing climate; the photographs have been used to understand changing vegetation structure; and the voucher specimens in combination with other specimen collections have been used for large scale distribution modeling efforts. the digitization and sharing of the data via the web has broadened the scope and scale of the types of analysis performed. yet, additional research avenues can be pursued using multiple types of vtm data, and by linking vtm data with contemporary data. the digital vtm collection is an example of a data infrastructure that expands the potential of large scale research through the integration and synthesis of data drawn from numerous data sources; its journey from analog to digital is a cautionary tale of the importance of finding historical data, digitizing it with best practices, linking it with other datasets, and sharing it with the research community. keywords.—vtm; vegetation mapping; historical data; california; climate change; application programming interface; api the grand challenges that human societies face globally—chiefly biodiversity loss, climate change and conflicting land use—are at the intersection of natural and social systems, and will require a synthesis of past, present and projected future data to understand their complex interactions and mitigate the effects of rapid environmental change. the synthesis of historical and contemporary ecological data with predictive modeling approaches is a powerful approach to understand the complex response of species, communities, and landscapes to changing biophysical conditions through time and in space. synthesis of historical and current data can create new knowledge through novel combinations of datasets and modeling approaches (krebs et al. 2001; schulte and mladenoff 2001; mladenoff et al. 2002; peters 2010; fox and hendler 2011); can help guide management and conservation by improving the temporal transferability of predictive models biodiversity informatics, 11, 2016, pp. 40-62 41 (kittinger et al. 2013; rapacciuolo et al. 2014a); and can provide crucial information for adaptation planning, policy, and management of natural systems (stein et al. 2010; whipple et al. 2011; higgs et al. 2014). while calls for digital infrastructures that allow for interoperability between multiple ecological data have been heard since at least the mid 1990s (e.g., michener et al. 1997), there are remaining barriers that limit such data synthesis. for most of the last century, ecological and geographical databases were mostly disparate: developed and maintained by individuals or small academic groups, on focused areas, and concentrated in time (michener et al. 1997; frehner and braendli 2006; michener 2006). this paradigm, while still prevalent, is poorly suited to answer large multi-scale interdisciplinary questions that require data synthesis, large data files, user interactivity, dynamic and shared updates, and collaboration between scientists and stakeholders across social and ecological domains (jones et al. 2006; tenopir et al. 2011; borgman 2012; hampton et al. 2013). what is needed are data frameworks that facilitate these aspects of science through novel developments in data curation and integration that rely on web applications, open standards, and application programming interfaces (apis) that make the interaction between scientist and multiple diverse data archives as well as the building of analytical tools via the web easier. these technological improvements and standardized protocols make multiple diverse and large datasets more readily available for integration, searching, visualization, and analysis by researchers (frehner and braendli 2006; peters 2010; reichman et al. 2011; borgman 2012). by focusing on one valuable and, at times in its history, vulnerable dataset, we highlight the importance of data rescue, digitization and sharing of historical datasets for environmental science. we also make the case that data, open data frameworks, modeling, and scientific collaboration are critical for modern applied ecological exploration in order to understand past flora and land use, to search for mechanisms for decadalscale vegetation changes, and to predict and plan for the future. we present a brief description of the wieslander vegetation type mapping (vtm) project and the digitization of its components, and review the ways in which the vtm collection has been used in ecological and geographical analysis. we close with a discussion of its journey from paper record to digital database and api as an example of many of the critical themes facing ecologists today: data, data sharing, data synthesis, and modeling. vtm: a case study from california california provides a unique case study for studying the interactions between climate, land use, and the distribution and abundance of populations and communities due to its diversity in climate, biota, topography and land use and its long tradition of natural history recording and collection (rapacciuolo et al. 2014b; chornesky et al. 2015). moreover, the state is at the forefront in a range of climate adaptation measures (chornesky et al. 2015). these complex and interacting socioecological phenomena are evidenced in the current landscape but result from historical processes (solé and bascompte 2006). future planning increasingly requires syntheses of historical and contemporary data as well as modeling approaches to understand the complex nature of environmental processes and the response of species, communities, and landscapes to changing biophysical conditions through time and in space (beck et al. 2012). although many important historical datasets describing california vegetation exist, most exist in analog form and in isolation. small-scale county maps of forest cover of the more heavily forested portions of the counties in northern california began to be made in the late 19th century (colwell 1977; keeler-wolf 2007) and much of the sierra nevada was mapped for forest cover in the first decade of the 20th century by (leiberg 1902). however, for the most part, california lacked the coverage provided in much of the midwestern and eastern portions of the us by the general land office surveys (galatowitsch 1990; schulte and mladenoff 2001). one exception to this is the mapping collection we consider in this paper: the wieslander vegetation type mapping (vtm) collection, named after its director albert wieslander. the vtm collection covered california in the 1920s–1940s and has been described as “the most important and comprehensive botanical map of a large area ever undertaken anywhere on the earth’s surface” (kuchler 1967; jepson et al. 2000) and the “most ambitious attempt ever made biodiversity informatics, 11, 2016, pp. 40-62 42 figure 1. examples of the components of the original vtm collection: (a) vegetation maps (placer co.); (b) a plot card (imperial co.), and (c) a landscape photograph (san mateo co.), (d) an herbarium specimen (arctostaphylos morroensis, san luis obispo co.), and (e) the digital representation of maps, plots and photographs in the vtm website showing an area covering part of lake tahoe and the tahoe national forest. background color = vegetation polygons, red dots = plot locations, black icons = locations of georeferenced photographs. figure 2. schematic of the berkeley ecoengine api, holos and vtm websites. elements related to vtm are in bold and italic font. biodiversity informatics, 11, 2016, pp. 40-62 43 to describe the complex vegetation of california” (critchfield 1971). the collection “remains to this date the most exhaustive and detailed effort of mapping vegetation in the state” (keeler-wolf 2007). the collection provides a detailed picture of much of california land cover and vegetation in the early 20th century, and it continues to contribute to the study, characterization, and understanding of historical, contemporary, and future california landscapes. beginning in the 1920s, albert everett wieslander, an employee of the u.s. forest service (usfs) california forest and range experiment station located in berkeley, began an effort to map california’s wildlands (colwell 1977). by the time world war ii interrupted the project, 40 million hectares had been visited, covering most of the state’s natural areas exclusive of the deserts and the larger agricultural areas (wieslander 1961; colwell 1977). at the end of the war, the project resumed soil-vegetation survey maps (griffin and critchfield 1972), but relied on aerial photography techniques. the priority of the original vtm project was to draw detailed vegetation type maps, but the mapping protocol evolved over time with need. it soon added the collection of detailed floristic (trees and shrubs) and environmental data from plots with locations recorded on usgs topographic quadrangles (wieslander 1986); black and white landscape photographs and maps showing the vantage point of the photographs; and herbarium specimens for species recorded on the vegetation maps or in the sample plots (ertter 2000; kelly et al. 2005). the vegetation maps, plots, plot maps, photographs and herbarium specimens are now curated at university of california libraries and herbaria system. however, over the 20th century, parts of the collection have been at risk of loss. in his oral history, wieslander related how many of the original vegetation maps were thrown away by the university press, and a few were rescued from destruction by a concerned professor at uc berkeley (wieslander 1986). in a second example, in the 1990s the plot data collection was almost discarded to create collection space until a university librarian contacted a professor at berkeley and together they moved the paper collection into the professor’s office (norma kobzina, pers. comm.). examples of the collection are found in figure 1; and more details of the collection can be found several publications (kelly et al. 2005; thorne et al. 2008; thorne and le in press). vegetation maps based on direct field observations from vantage points, dominant vegetation types were mapped directly onto u.s. geological survey (usgs) topographic quadrangles with a minimum mapping unit (mmu) of 16 ha, and supplemented by sample plots (described below). the vegetation mapping scheme was driven by dominant overstory vegetation and included vegetation mosaics: complex vegetation conditions that resulted from fire or other disturbances, and pure and mixed stand conditions which were associated with “natural plant associations” (wieslander 1935b, a, 1961; thorne and le in press). the mapped products include 215 maps with the major vegetation types shown in different colors and separated by ink lines printed on 7.5’, 15’, and 30’ usgs topographic quadrangles (colwell 1977; thorne and le in press) (figure 1). plot data and plot maps the vtm crews visited over 18,000 vtm plots statewide; these are concentrated along the central and southern coastal ranges, and in the sierra nevada mountains. these plots were surveyed as a check on the vegetation polygons, and also to provide details on species composition, size and stand density of trees and shrubs and depth of leaf litter. plots were rectangular (800 m2 in forests, 400 m2 in shrub and chaparral communities), ran upslope and were divided into milacre sampling units in which dominant species and height characteristics were recorded. trees greater than 10 cm in diameter breast height within 10 m of either side of the centerline were tallied by species and diameter class. in both forest and understory plots, slope, soil characteristics, and year of last burn were recorded (wieslander 1935b, a). the plots cover a gradient of vegetation types and include data regarding tree stand structure (number per diameter class), percent cover of dominant overstory and understory vegetation by species, soil type, parent material, leaf litter, elevation, slope, aspect, parent material, and other environmental variables. all plot data were stored on paper data sheets and individual plots were numbered according to map name, quad section biodiversity informatics, 11, 2016, pp. 40-62 44 table 1. scientific focus of reviewed vtm studies. asterisks (*) indicate ph.d. dissertations. focus n items references conifer forest / mixed conifer forest / conifer trees 15 plots: maps: photos: minnich 1978*; minnich 1995; bouldin 1999*; goforth and minnich 2008; fellows and goulden 2008; lutz et al. 2009; lutz et al. 2010; swanson et al. 2013; dolanc et al. 2013a; dolanc et al. 2013b; dolanc et al. 2014; maxwell et al. 2014 walker 2000* dodge 1975*; taylor 2000 conifer and hardwood forest 1 plots: maps: photos: mcintyre et al. 2015 none none shrubland / chaparral 9 plots: maps: plots & ma ps: photos: minnich and dezzani 1998; franklin 2002; franklin et al. 2004; keeley 2004; taluto and suding 2008; syphard and franklin 2009 lippit et al. 2012; cox et al. 2014 bradbury 1974* none oaks / oak woodland / rangelands 4 plots: maps: photos: allen et al. 1991; allen-diaz and holzman 1991; vayssieres et al. 2000; conlisk et al. 2012 none none vascular plants (including both tree and shrub) 5 plots: maps: photos: syphard and franklin 2010; crimmins et al. 2011; dobrowski et al. 2011; crimmins et al. 2013; crimmins et al. 2014 none none land cover 6 plots: maps: photos: none thorne et al. 2004; thorne et al. 2008; rubidge et al. 2011; thorne et al. 2013; santos et al. 2014a; santos et al. 2014b; none grasslands 1 plots: maps: photos: none freunenberger et al. 1987 none various 3 plots: maps: kelly et al. 2008 preston et al. 2012; davis and sims 2013 biodiversity informatics, 11, 2016, pp. 40-62 45 number and plot number. individual plots locations were denoted by 3.5 mm hollow circles stamped in red ink on usgs topographic maps (editions of 1893–1920, reprinted in the 1930s) that had been cut into sections and mounted on canvas for ease of use in the field. there are about 150 15’ (1:62,500 scale) and 30’ (1:125,000 scale) plot maps (figure 1). photographs and metadata the vtm crews also captured about 3100 black and white landscape and stand scale photographs (9.2 x 13.6 cm), taken during 1920– 1941. the location of origin of many of these photographs are marked on usgs topographic maps in red pen, with an arrow marking the vantage point and view of the photo. the photograph captions typically includes a description of the location and subject of the photograph including relevant genus and species, timber stand conditions, and examples of cultivation, grazing, logging, mining and fire, and quad name. the photographer, date of the photograph, and occasionally township and range are included (figure 1). herbarium specimens the vtm crews collected over 23,000 herbarium specimens as part of their efforts. a primary motivation was to voucher specimens for regional identification of taxa observed in plots (wieslander 1935b, a). the vtm specimens span 3157 taxa, representing 40% of the ~7600 plant taxa recorded from california. primary holdings are at the university and jepson herbaria (uc berkeley), with duplicates housed at various institutions, most notably at the university of california santa cruz where vtm specimens helped form the initial holdings of the herbarium. the specimens are databased and georeferenced but awaiting integration with other vtm data on the vtm.berkeley.edu website. vtm digitization process the digitization of the vegetation maps, plots, plot maps, photographs, and locations of herbarium specimens was led by people in several groups in the university of california. the digitizing of the vegetation maps was led by james thorne at uc davis over a 10-year period. this involved finding many of them at various repositories and obtaining the same edition topographic maps as the vegetation maps were drawn on from university libraries. these were digitized and converted to gis data. the photograph collection was almost completely intact in the 2000s in the marian koshland biosciences museum at uc berkeley, and were scanned and uploaded online. in 2014, the efforts to georeference the photographs were led by michelle koo of the museum of vertebrate zoology at uc berkeley. the plot data cards and the majority of the vegetation maps were physically located in the 1990s in the laboratory of barbara allen-diaz at uc berkeley, and her lab led the digitization of the plot data. the plot maps were physically also located in dr. allen-diaz’ lab and maggi kelly led the scanning and georeferencing, and the linking of the map data with the plot data. the herbaria specimens are physically located in the jepson herbarium at uc berkeley, and georeferencing of the specimen locations was led by koo following darwincore standards and best practices for georeferencing. the collection parts were digitized and georeferenced with similar protocols resulting in digital shapefiles of polygons (vegetation maps) or points (plots, photographs and specimens) linked to respective databases. more details on the digitization and georeferencing protocols can be found in (kelly et al. 2005; thorne and le in press). map georeferencing was done by a suite of analysts using a collection of tie-points gathered from the vtm maps and current usgs topographic quadrangles. the relevant features for each part of the collection (e.g. polygons for the vegetation maps or points for the plots) were transcribed manually from the digital version of the respective map. vegetation species codes from vegetation maps and vegetation plot cards were transcribed using the manual of california vegetation types (sawyer and keeler-wolf 1995), and the california wildlife habitat relationships models (whr) (mayer and laudenslayer 1988) for land cover classifications. the locations of each photograph depicted on an accompanying map usgs topographic map were georeferenced by measuring the distance and bearing of each marked point from the known southwest corner and calculating its location. specimens were digitized primarily to township, range and section centroids. biodiversity informatics, 11, 2016, pp. 40-62 46 figure 3. locations in california of scientific projects that have used vtm data. biodiversity informatics, 11, 2016, pp. 40-62 47 web stack and api development all digital spatial vtm data are made available in an open source web-mapping application1 developed with an entirely open source software stack by uc berkeley’s geospatial innovation facility. all vtm data are stored using postgresql, a relational database that supports the storage of and analysis of geospatial vector data through the postgis extension. the map interface was built using leaflet, a lightweight javascript mapping library with open street map providing base layers (figure 2). the website presents a user interface for exploring, searching, aggregating, and downloading the vtm data collection. the vtm website was built using the berkeley ecoinformatics engine application programming interface (ecoengine api2). a web api is an application that serves machine-readable data and functionality to applications that represent the data to users. the ecoengine api is a directory and gateway to many diverse biological and environmental collections at uc berkeley, part of a recent trend that makes use of digitized museum collections for the study of anthropogenic loss of biodiversity, climate change, and changes to ecosystems (pyke & ehrlich 2010). the ecoengine api serves about 5 million records from museum specimens, soil and pollen data, field station records, sensor readings, as well as biophysical base layers such as climate and land use and the vtm data. the vtm website accesses the vtm data collection through three main ecoengine api resources: vtm photos, vtm plots and vtm vegetation (figure 2). for each resource, the user can explore a map that shows geographic locations of the resource and when the user clicks on a feature, a pop up window displays greater detail about the feature. querying the ecoengine api retrieves this detail. the original vegetation maps themselves display vibrant color schemes (figure 1), but rather than try to replicate this variety (although that might be possible in the future) the polygons are rendered using a modification from the usgs nlcd land cover color scheme (fry et al. 2008). all original vtm vegetation types were cross-walked with california wildlife habitat relationships (cwhr) information. 1 http://vtm.berkeley.edu. 2 https://ecoengine.berkeley.edu/. 2 https://ecoengine.berkeley.edu/. uses of the vtm collection we found 39 peer-reviewed journal articles and 5 ph.d. dissertations that discussed the use of vtm vegetation maps, plots or photographs before 2016 using google scholar. herbarium specimens have been used in combination with other numerous georeferenced specimens available through the consortia of california herbaria to assess the responses of different groups of plant species across the entire flora of california to temperature change (loarie et al. 2008; wolf et al. 2016). we did not evaluate reports or grey literature. of the documents reviewed, the majority (n = 31) of studies focused on the vtm plot data, 12 used vegetation maps, and two used photographs. the majority of the studies focused on forested landscapes (conifer or hardwood), but many focused on shrublands or rangelands, or on land use. research using plot data completed after data digitization in the mid-2000s (kelly et al. 2005) generally focused on regional or statewide analyses. a summary of the uses of vtm data is provided in table 1, with their study areas shown in figure 3. use of vtm data to develop vegetation reference systems the floristic detail found in the plot database has been used since the 1940s to develop californian vegetation community classification schemes in forests (maxwell et al. 2014), chaparral (franklin 2002), oak and rangeland communities (allen et al. 1991; allen-diaz and holzman 1991), as well as general land cover classification schemes (thorne et al. 2004). more recently, the plot data has been used to reconstruct past reference conditions or to establish a historical range of variability (hrv) for a particular time period to target ecosystem restoration or landscape management. for example, maxwell et al. (2014) used 399 vtm plots in conjunction with dendroecological reconstructions in the lake tahoe basin area to reconstruct forest structure, forest fuels, and fire regimes representing pre-settlement conditions as an aid to current forest management. herbarium specimens helped establish baseline distributions reference material for the distribution and identification of california trees and shrubs (mcminn 1951; griffin and critchfield 1972). biodiversity informatics, 11, 2016, pp. 40-62 48 20th century changes to vegetation communities the vtm collection has been used recently to understand the patterns and causes of changes in vegetation structure over decadal scales. authors have examined shifts in abundance and composition, as well as decadal scale changes in the elevational distribution of vegetation communities (i.e. community shifts upslope or downslope). this type of work is critical for predicting how vegetation communities will respond to a changing climate, as well as planning for community resilience in the face of disturbances. changes to structure and composition of california forests.—comparison of the vtm plots with current vegetation (using both relocated plots and contemporary vegetation maps) revealed consistent evidence of an increase in young-growth and small-diameter trees across the state and decreases in large trees (minnich et al. 1995; fellows and goulden 2008; goforth and minnich 2008; lutz et al. 2009; dolanc et al. 2013a, 2014; mcintyre et al. 2015), as well as changes in forest composition (minnich et al. 1995; dolanc et al. 2013a; mcintyre et al. 2015). much of this evidence of change in forest age structures derives from resurveyed vtm plots (table 1). goforth and minnich (2008) resurveyed three vtm plots of mixed conifer forest on cuyamaca mountain in the peninsular range of southern california four years after a stand replacing fire. they replicated measurements at multiple sites around the expected locations of the three original plots and covered similar environ-mental conditions as described in the original vtm plots. their analysis showed significant changes in forest composition, increases in tree density, and decreases in stem diameter over a 75-yr period. they augmented these results with an analysis of repeat aerial photographs from 1928 and 1995 that also show significant increase in canopy cover. dolanc et al. (2013b) resurveyed 139 vtm plots in the subalpine zone of the sierra nevada, and compared historical and modern climatic conditions using two high-elevation climate stations nearby. they found fewer larger trees and increases in smaller size classes. more evidence of increased tree density and change to age structure has been found comparing vtm data with modern plot data, typically provided by the forest inventory and analysis (fia) program (u.s. forest service 2016b) (table 2). the primary objective of the fia program is to determine the extent, condition, volume, growth, and use of trees on us forest land through repeated and widespread sampling. congress mandated the fia program in 1928, but it was not until 1999 that collections on a network of plots began to be made annually. fellows and goulden (2008) compared 269 vtm forest plots with 260 fia plots from the 1990s to quantify changes in aboveground biomass for california forests. they found that the size structure of the forests changed dramatically in 70 years, with large increases in stem density driven by increase in number of smaller trees and a net loss of large trees, with a concordant decrease in aboveground carbon stocks, estimated allometriccally. they attributed this change to fire suppression. lutz et al. (2009) compared 655 vtm and 210 modern vegetation plots surveyed by national park service field crews from 1988-1999 in yosemite national park. they found that largediameter tree density in yosemite declined by 24%, and declines were greatest in subalpine and upper montane forests, and least in lower montane forests. dolanc et al. (2013a) examined changes in abundance and composition across nine vegetation types in northern sierra nevada using 4321 vtm plots and 1000 modern fia plots. tree density was significantly higher in the modern plots in eight of nine vegetation types. they also found a shift in dominance toward shade-tolerant conifers and evergreen oaks. they attributed some of these changes to fire suppression, although not all, as alpine forests, which do not experience fire, also showed increases in density. dolanc et al. (2014) conducted a similar study of the entire elevational range of west-facing slopes in the central and south sierra nevada, further confirming significant patterns of change in forest structure and composition. mcintyre et al. (2015) used both vtm plot data (n = 6,572) from the entire state, and compared baselines in forest structure and composition in oaks (quercus) and pines (pinus) with those derived from fia data (n = 1909). they found declines in large trees in every ecoregion of the state, but also found forest composition has shifted toward increased dominance by oaks relative to pines. they showed that declines in large trees were more severe in areas experiencing greater increases in climatic water deficit since the 1930s. loss of shrub and chaparral communities.— shrublands are vegetation communities dominated ta bl e 2. d at a co m m on ly u se d to c om pa re w ith th e v tm c ol le ct io n. d at as et d at e a ge nc y sc al e ex te nt v tm 19 28 -1 94 0 u .s . f or es t s er vi ce 1: 12 5, 00 0, 1: 62 ,5 00 c al ifo rn ia v tm p lo t c om pa ris on fi a * 19 30 -p re se nt u .s . f or es t s er vi ce pl ot d at a u ni te d st at es v tm v eg et at io n m ap c om pa ris on c al v eg * 19 79 -1 98 1 u .s f or es t s er vi ce p sw f or es t a nd r an ge ex pe rim en t s ta tio n, u .s f or es t s er vi ce r em ot e se ns in g la bo ra to ry m m u 6 -8 00 a cr es c al ifo rn ia (e xc ep t so ut he rn in te rio r) h ar dw oo d r an ge la nd m ap pi ng p ro gr am 19 81 -1 99 1 c al ifo rn ia d ep t. of f or es try a nd f ire p ro te ct io n 1: 24 ,0 00 – 1: 58 ,0 00 c al ifo rn ia h ar dw oo d an d sa va nn ah (< 50 00 ft ) c a lg a p* 19 90 -2 00 8 u sg s 1: 10 0, 00 0 (1 k m m m u ) c al ifo rn ia n lc d * 19 92 -2 00 120 06 -2 01 120 13 u sg s 30 m u ni te d st at es * fi a = f or es t i nv en to ry a nd a ss es sm en t, c al v eg = c la ss ifi ca tio n an d a ss es sm en t w ith l an ds at o f v is ib le e co lo gi ca l g ro up in gs ; c a lg a p = c al ifo rn ia g ap a na ly si s p ro gr am ; n lc d = n at io na l l an d c ov er d at ab as e. biodiversity informatics, 11, 2016, pp. 40-62 townpeterson typewritten text townpeterson typewritten text townpeterson typewritten text 49 biodiversity informatics, 11, 2016, pp. 40-62 50 by short evergreen woody shrubs with sclerophyll leaves. these chaparral and sage scrub plant communities, once dominant in southern california, are now threatened by changing land use, fire regimes, nitrogen deposition, and invasive species. the vtm collection has played a large role in understanding 20th century changes to these communities. the original priority of the vtm project was the vegetation type maps and forest cover, but the mapping protocol developed over time to suit emerging needs. around 1927, the forest plot protocol was extended to include information on shrub communities in response to requests from southern california-based forest service (wieslander 1986). the earliest examination of changes to chaparral communities in san diego county that made use of the vtm plots is bradbury (1974), who resurveyed vtm plots 40 years after the original vtm crews and found slight changes to chaparral communities, most of which he related to disturbance (bradbury 1974; franklin 2002). freudenberger et al. (1987) used vtm vegetation maps as a guide to help interpret historical aerial photography in a project examining the shifting mosaic of grassland and shrubland in the los angeles basin over 50 years. franklin et al. (2004) resurveyed 649 plots to investigate the role of fire in changes in species composition and cover in chaparral communities in san diego county. they report a complex interplay between oak woodlands, chaparral communities and fire response. more recent uses of the vtm collection explore a range of disturbances driving the loss of southern california shrublands including fire, invasive species and nitrogen deposition. talluto and suding (2008) resampled 54 vtm shrub plots in southern california and found strong evidence of non-native grassland invasion at the expense of coastal sage scrub. grassland encroachment was positively correlated with increased fire frequency and, in areas with low fire frequencies, air pollution (likely nitrogen deposition). lippitt et al. (2013) investigated the role of fire in chaparral community resilience in southern california. they used the vtm maps and calveg (united states forest service 2016a) to identify areas of chaparral that experienced high frequency fires, finding that the number of burns and decreases in mean fire interval over time increased the chance for alteration and type conversion of chemise chaparral communities. cox et al. (2014) used a vtm vegetation map to examined the factors leading to dynamics between coastal sage scrub and grassland between 1930 and 2002. nitrogen deposition and fire were important to understanding and predicting the recovery of the coastal sage scrub system. with the exception of cox et al. (2014), these studies relied on the relocation of vtm plots. plot relocation is known to be difficult (e.g. minnich et al. 1995), especially in shrublands (keeley 2004). however, there seems to be ample evidence of successful relocation, particularly in efforts focused over larger scales that incorporate larger numbers of plots which reduce overall spatial uncertainty (kelly et al. 2008). additionally, there is evidence that relocation of vtm plots and fia data (e.g. mcintyre et al. 2015) provide very similar results (dolanc et al. 2013b). range shifts in vegetation communities.— pioneering work to digitize, georeference and attribute the countless vegetation type polygons from the original vtm collection has enabled analyses of shifts in vegetation classes over time (thorne and le in press). for example, thorne et al. (thorne et al. 2008) compared the digital vtm vegetation maps to contemporary remote-sensing products such as calveg; (schwind and gordon 2001) and found significant changes in vegetation types in the placerville quadrangle on the west slope of the sierra nevada. they used the california wildlife habitat relationships (cwhr) classes (mayer and laudenslayer 1988) as a crosswalk between the two data types, and found significant shifts in whr classes. at lower elevations below 700 m, annual grassland expanded, and low elevation hardwoods and conifers, particularly blue oak woodland and blue oak-foothill pine, contracted. at higher elevations above 700 m, ponderosa pine contracted while montane hardwood, montane hardwood-conifer and douglas fir expanded. a different approach was taken by crimmins et al. (2011), who examined the altitudinal distributions of 64 californian vascular plant species using 13,746 vtm plots and modern fia plots. they used logistic regression to estimate species’ optimum elevations in each period and altitudinal shifts were measured as the difference in optimum elevation between periods. contrary to the expectation that species from lower elevations will townpeterson typewritten text townpeterson typewritten text townpeterson typewritten text biodiversity informatics, 11, 2016, pp. 40-62 51 shift upslope in response to warming (e.g. kelly and goulden 2008; lenoir et al. 2008), they found that climate changes in california have resulted in a significant downward shift in these species’ optimum elevations. they explain this downhill shift by regional changes in climatic water balance. similarly, dolanc et al. (2013b) who resampled vtm plots in the sierra nevada in a study of forest structure changes also found no evidence of upslope shifts in vegetation communities: no species move into or out of the study area when comparing vtm plots with modern resurveys. use of vtm photography.—there are only two published examples that highlight the use of the vtm photographs to understand vegetation change, and no recent examples. wieslander himself discussed the possibilities of using repeat photography (wieslander 1986). dodge (1975) used vtm photographs to study historical changes in san diego county vegetation, with mixed results. he reported that most photographs in the area were stand level rather than landscape view inhibiting his ability to perform exact reshoots. he compared vegetation changes evident from rephotographing the general area of each photograph and reported that vegetation increased in density and large accumulation of dead material on the ground in areas with no evidence of fires, and in areas that experienced wildfires in the intervening years there was near total destruction of coniferous forests and extensive damage to oak woodlands. taylor (2000) had modest success in relocating four 1925 photographs on prospect peak in lassen national park; in three of them he found evidence of increased forest density, as a result of fire suppression. forecasting future change the use of past and contemporary data to understand biogeographic responses to recent climate change is fundamental for improving our predictions of likely future responses (rapacciuolo et al. 2014a). several recent papers have used the vtm plot data to gather representative samples of species to parameterize current distributions as a precursor to modeling future distributions. lutz et al. (2010) used data from 655 vtm plots to gather representative tree species data for yosemite national park. they correlated climatic water deficit with past and current tree distribution and projected tree distribution into the future using standard climate scenarios. they suggest that ongoing changes in forest structure and composition can be related to changes in climate water balance. while past increases in temperature since the little ice age have been offset by increasing precipitation, projected future temperature increases coupled with likely decreases in precipitation will increase water deficit, with detrimental impacts on many species. they showed projected future increases in water deficit of 23% across all plots in yosemite, which may disproportionately affect western white pine (pinus monticola) and mountain hemlock (tsuga mertensiana). dolanc et al. (2013b) used vtm data for a retrospective examination of climate related changes to subalpine forests, and they commented on the possible future changes to these communities. their work examining changes in tree abundance and composition in subalpine forests of the sierra nevada did not find evidence of change in the direction predicted by vegetation models linked to future climate scenarios. they explain this discrepancy by suggesting a possible lag effect, whereby lower-elevation species may eventually replace higher-elevation species, but only after decades or even centuries. conlisk et al. (2012) used vtm plot data and herbarium records with the species distribution modeling algorithm maxent (phillips et al. 2006) to find locations and carrying capacities of metapopulation patches of quercus engelmannii in eastern san diego county. they combined maxent models with a demographic model to determine the population dynamics within and between these patches in the future. they predicted the dramatic reduction of suitable habitat for q. engelmannii in 2100 under two climate scenarios, with suitable habitat patches predicted to shrink in extent and move to higher elevations. model assessment and validation species distribution models (sdm) correlate environmental predictors such as climate and topography to the known distribution of a species, and can be used to generate spatial predictions of the suitability or probability of presence for a species given the predictors. the past decade has seen an expansion in the types and sophistication of species distribution modeling approaches that can be used to forecast changes to species and biodiversity informatics, 11, 2016, pp. 40-62 52 communities based on locality data and baseline conditions (graham et al. 2004; guisan and thuiller 2005; guo et al. 2005; elith et al. 2006; elith and leathwick 2009). sdms have many challenges when forecasting species responses under changing environments, including the potential absence of a species-environment equilibrium, the difficulty to account for dispersal limitations, biotic interactions, phenotypic plasticity and evolutionary changes, and the incidence of novel environments outside the range of conditions used to calibrate the models (elith and leathwick 2009). despite these obstacles, sdms have become an integral part of conservation planning, resource management, and land decision-making processes. in recent years, the vtm plot data have been increasingly used in sdm studies to provide temporally independent data to validate model predictions and explore some of the methodological nuances of these models. specifically, vtm data have been used to evaluate model selection and methodology (single vs. ensemble models, or model vs. model) (vayssières et al. 2000; crimmins et al. 2013), to assess model uncertainty over time and incorporate spatial autocorrelation (e.g. swanson et al. 2013), to assess model transferability through time (syphard and franklin 2009, 2010; dobrowski et al. 2011; crimmins et al. 2014), to test how sdm performance varies with species’ traits and environmental predictors (syphard & franklin 2009, 2010), and to explore how sdms can be coupled with demographic models to forecast population responses (conlisk et al. 2012). linkage between vegetation and animal distributions through time studies using the vtm data predominantly examine change in vascular plant species. however, a small but growing number of studies have utilized vtm data to understand the impact of long-term vegetation changes on animal distributions. rubidge et al. (2011) highlighted the promise of using vtm maps with historical and resurveyed grinnell resurvey project transects3 to investigate drivers of change in alpine chipmunk (tamias alpinus) distributions. the detailed descriptions found on the vtm maps allowed vegetation polygons to be matched to california 3 http://mvz.berkeley.edu/grinnell/index.html. wildlife habitat relationships (cwhr) classes, generating comparable historical and current habitat maps for use in models of small mammal distribution changes. though rubidge et al. (2011) found only a limited effect of vegetation change on small mammal range shifts compared to climate, that study paved the way for the use of vtm maps in this context. as a follow-up to that study, santos et al. (2014b) investigated the relationship between species’ traits and habitat types using vtm maps as historical habitat. the authors found that omnivorous species responded in greater synchrony with shifts in their habitat types than any other diet guild. preston et al. (2012) used vtm maps to investigate the drivers of distribution patterns of an endangered butterfly (euphydryas editha quino). the study used vtm maps to aid in the correlation of local scale quino checkerspot population extinctions with human population growth and land use change, while climate variables determined population distributions at the broader regional scale. the vtm maps have also been used to understand proximate land use hazards associated with commercially important species. davis and sims (2013) analyzed the effect of land use change and precipitation events over time on the prevalence of landslides and gullies, which contribute sediment to the san pedro creek watershed in san mateo county, a key habitat for salmonids (davis and sims 2013). land use and land cover history and urban planning many currently urban areas in california have faced dramatic histories of land use and land cover change as agricultural and natural areas adjacent to cities have been encroached upon. the san francisco bay area exemplifies this complex pattern of land cover change over the previous one hundred and fifty years and is a mosaic of urban, natural and agricultural land. thorne et al. (2013) examined the dynamics of urban growth and the protection of open spaces—two often contentious interests—using vtm maps and urban growth models to predict how their dynamics under three contrasting urban growth policy scenarios. the authors determine that a historical trend of urban expansion will likely continue into the future, with new development converting an additional 43-53% biodiversity informatics, 11, 2016, pp. 40-62 53 figure 4. increase in number of plots used in research over time. townpeterson typewritten text 53 biodiversity informatics, 11, 2016, pp. 40-62 54 of existing agricultural and grassland areas. however, the growth policy enforced will have a large effect on the overall impact of urban growth on open space. santos et al. (2014a) used vtm vegetation maps to analyze the timeline of conservation land acquisition in the san francisco bay area from 1850-2010. describing the evolution of these land acquisitions and the cover types they protected, the authors found that the network of bay area conservation lands was highly representative of the diversity of land cover types, encompassing 20% or more of each land cover type in the region. the acquisition of land in the 19th century followed what the authors termed the “fill-in effect”: the purchase of fewer larger properties while the contemporary network was filled in through time by several smaller properties consisting of underrepresented land cover types. scale of vtm studies the geographic scope of study using vtm data has grown over time, particularly after 2005 when much of the data became available in digital form (kelly et al. 2005) (figure 4). studies prior to 2000 were geographically focused and used small numbers of plots (e.g. minnich and dezzani 1998; bouldin 1999). kelly et al. (2008) were the first to use the complete plot database (n = 18,000) in their study focusing on accuracy assessment. since the release of the digitized dataset, a multitude of regional scale projects focusing on the central sierra nevada forests (dobrowski et al. 2011; dolanc et al. 2013a; dolanc et al. 2013b) and the sierra nevada and the coastal ranges (crimmins et al. 2011; crimmins et al. 2013; crimmins et al. 2014) have been published, as well as studies focusing on the state and using the complete plot database (fellows and goulden 2008; mcintyre et al. 2015). another perhaps unforeseen result of the digitization process has been an alteration in the way change is assessed using historical and current data. many authors have commented on difficulties in re-locating individual plots (e.g. (e.g. allen-diaz and holzman 1991; minnich et al. 1995; walker 2000; franklin 2002; keeley 2004); and recent efforts have focused less on relocation of individual plots, and more on comparisons with contemporary data (mcintyre et al. 2015). with some exceptions, a similar trend is seen in the use of vegetation maps: earlier work focused on a single quadrangle (bradbury 1974), while more recent work is regional in scope (thorne et al. 2013; santos et al. 2014a). this kind of scalingup of focus is only realistically possible with digital infrastructures to make data discoverable, searchable, and downloadable. discussion historical ecological data is often required in contemporary ecological analyses for understandding patterns, establishing baselines and past conditions, and detecting changes. these analyses are crucial when we consider planning for possible futures. as a case study, the vtm collection reveals the promise of historical data for ecological analysis. here, we discuss prospects for increased and integrated usage of the dataset; for novel synergies with other data; and highlight through a discussion of the vtm digitization process the importance of open data frameworks for ecological analysis. challenges with the use of historical data the use of historical ecological data present at least three types of challenges for the modern researcher that result from field methods used at the time of data collection, scientific terminology and taxonomy, as well as errors introduced through the digitization process. first, historical field protocols are often different from contemporary ones. the vtm crew used four size classes to bin the diameter at breast height of trees recorded in the plot data; modern forest inventory and analysis (fia) data used many more. comparisons must deal with these differences, and manipulate modern data to best compare with historical data (sensu dolanc et al. 2013a; mcintyre et al. 2015). in contrast, the vtm vegetation maps have a much higher species detail than is provided on modern land cover maps, which requires simplification of the historical vegetation classes and cross-walking of historical and modern classification schemes prior to contemporary analyses (sensu thorne et al. 2008; thorne et al. 2013). second, ecological taxonomy has changed since the early 20th century (barbour et al. 2007), and confusion resulting from species name changes can occur, and careful crosswalking between historical and modern species names is required. third, as the historical data is scanned, digitized, and georeferenced, error is introduced. best practices must be used to minimize, estimate, and report the total spatial biodiversity informatics, 11, 2016, pp. 40-62 55 error in the final digital data product (kelly et al. 2008). such reported measures are critical to guide researchers on the use of the vtm data (thorne et al. 2008). interactions between the academy and the public sector in addition to its use in academic research over multiple decades, the vtm collection has had a parallel track of use by government resource agencies. the protocols developed by wieslander and his crew became the foundation for additional surveys of california land including 4.6 million ha of land covered by the state cooperative soilvegetation surveys from 1947-1977 (critchfield 1971; thorne et al. 2007). these early surveys paved the way for many of the contemporary vegetation classification schemes used today in california, including the manual for california vegetation, the national vegetation classification system, and the california gap analysis program. the calveg product, discussed in this paper, had its origins with vtm protocols (barbour et al. 2007). wieslander’s belief in constructing “an allpurpose map” that “would then be valuable for any kind of management of wildland, for whatever purpose” set the foundation for large scale mapping and biological surveys even serving as the impetus for the forest service’s “program of mapping all the national forests” (wieslander 1986). the value of comprehensive and detailed vegetation surveys made apparent through the work of wieslander and his crew remain a key component of agency mission statements and objectives (e.g. forest service, national park service, and california native plant society). we hope that the digitization and sharing of the vtm data will continue to expand the reach of this dataset in order to bridge traditional academic and agency silos for advances in research and conservation management practices. increased usage of the integrated dataset the plot data in digital form haves been available since 2008, but for the first time are complete (or as complete as possible) vegetation maps and georeferenced photographs are available, and a revival in their integrated use should be encouraged. it is our hope that the combined collection made available via an api and website will facilitate their use in a number of novel ways. first, the georeferenced photographs are an underused resource. only two studies focus on the photography collection (dodge 1975; taylor 2000), yet repeat photography can provide a clear record of vegetation and landscape change. only a handful of the georeferenced photographs have been retaken in modern times, yet these photographic pairs provide compelling stories about vegetation change (figure 5a-f). additionally, these images provide a window into our past and a sense of place: many of the vtm photographs are not of grand landscapes or vistas, but show the field data collection procedures used by the crews, or forest management practices they encountered in the early 20th century (figure 5g-h). thus these images provide unexplored cultural connections to our work and lives today (higgs et al. 2014). second, the plots provide data on structure and composition of vegetation, but they also comment on the terrain, soil types, fire and disturbances, and other environmental variables found at each site. these data might be mined for more nuanced understanding of historical fire behavior (e.g. mcintyre et al. 2015) that complements other early data sources of disturbances. third, there is great potential in the use of more than one vtm data type at a time. there is only one published example of research using more than one component of the vtm dataset (bradbury 1974), yet researchers might consider the combined use of plots and maps, or plots and photographs, or all of the components of the collection. there are numerous areas in the state where the three data types coexist including the sierra nevada mountains (e.g. figure 1), the central coastal ranges, the transverse ranges, and the san bernardino mountains (figure 3). the central coast ranges posses all components of vtm collection, and remain under-examined. new developments in data synthesis we should also expect and encourage new developments in data synthesis. data synthesis necessitates an extension of the data beyond its original purpose, increasing the potential to broaden the horizons of scientific inquiry and expand the potential for discovery (jones et al. 2006; tenopir et al. 2011). the recent proliferation in publicly available scientific data underscores the need to better integrate and synthesize these biodiversity informatics, 11, 2016, pp. 40-62 556 56 figure 5. example of vtm photography. examples a-f are pairs of vtm and contemporary images: french lake in the tahoe national forest in nevada co. taken by (a) n.h. french in june 1934, and (b) joyce gross in 2014; looking southwest from highway between drytown and amador in amador co. taken by (c) a.e. wieslander in may 1940, and (d) joyce gross in september 2013; and the site of what is now lake isabella in kern co., taken by (e) l.m. correll in july 1927, and (f) joyce gross in june 2013. examples g, h are examples of general interest photographs: (g) an example of a local mill and redwood lumber products cut from redwood second growth in santa cruz co.; and (h) topographic mapping crew on orleans mountain, in siskiyou co. biodiversity informatics, 11, 2016, pp. 40-62 57 datasets if they are to be effectively used in addressing issues such as climate and land cover and land use change which both span spatial and temporal dimensions (tingley and beissinger 2009). many of the studies reviewed here employ data synthesis between historical vtm and modern vegetation data (e.g., fia, calveg) to discuss change both qualitatively and quantitatively in vegetation structure (e.g. goforth and minnich 2008; dolanc et al. 2013b), composition (e.g. minnich et al. 1995), and abundance (e.g. fellows and goulden 2008; table 2). there are additional datasets and approaches that might be used. the vtm database can be linked to much of the ecological data collected at uc berkeley, among them about 5m records from museum specimen, soil and pollen data, field station records, sensor readings, as well as biophysical base layers such as climate and land use (figure 2). there is great potential in these data synergies to further expand the scope and scale of vtm analysis. for example, first, synergies between genome sequencing of biological specimens with historical and future vegetation allow us to move beyond single species vegetation type analysis and understand the dynamic nature of ecological communities (rubidge et al. 2011; preston et al. 2012; bi et al. 2013). second, more can be done using historical vegetation data with current and future climates using new species distribution models. for example, the floristic detail found in both the maps and plot data might be used to look at historical ranges of important taxa. third, there has been a significant focus on interactions between fire and vegetation in the vtm scholarship for example in conifer forest (dodge 1975; minnich 1978; taylor 2000; fellows and goulden 2008; goforth and minnich 2008) and shrubland communities (franklin et al. 2004; talluto and suding 2008; lippitt et al. 2013). mcintyre et al. (2015) used vtm field notes to identify plots with recent fires and were able to detect coarse differences in tree density in multiple size classes between burned and unburned plots within the vtm data and across time in comparisons of burned and unburned fia plots. more might be done to examine historical vegetation structure and pattern in modern mega fires such as the rim fire. conditions influencing fire overlap, fire severity and fire regimes might be explored through synthesis with contemporary fire data (collins et al. 2007, 2009). open data frameworks the multiple recent discussions in the scientific community around open, transparent, and reproducible science (jones et al. 2006; wolkovich et al. 2012) have fostered technological advances in more flexible and user friendly data storage and documentation systems that incentivize researchers to enter, store, and make datasets available to the broader community. such participation has engendered a need for a more transparent system of data collection and distribution. contemporary large-scale citizen science efforts promote and alleviate some issues of data transparency and sharing by promoting open data collection, distribution, and assessment (kearns et al. 2003; kelly et al. 2012). technological advances in computer science have increased our ability to query, visualize and use large heterogeneous collections in meaningful ways (baird 2010; fox and hendler 2011; reichman et al. 2011; hampton et al. 2013). new web applications such as apis linked to structured ecological databases allow the rapid generation of maps, charts, timelines, graphs, word clouds, search interfaces, rss feeds, and many others capabilities (fox and hendler 2011). the voyage from paper to api although focused on a single dataset, the vtm collection’s voyage from paper archive to api can be seen as a cautionary tale about the importance of finding, archiving, and sharing historical ecological datasets that should resonate more broadly. despite being well-known and welldocumented, in the decades since its creation, the vtm collection has faced the possibility on several occasions of partial destruction (wieslander 1986). even today, locations of portions of the collection remain unknown. thus, its journey from analog to digital is exemplary of a number of important themes facing the scientific community today: the importance of finding and rescuing historical or “dark” data; the need for best practices and standards for data digitization including uncertainty estimation and error control; the value of spatial data visualization and webbased portals for data sharing; the priority in modeling of data fusion and analytical integration; and the critical role of data infrastructures such as apis for sharing scientific data. other important detailed records of past biological, ecological, agricultural, and management conditions may exist biodiversity informatics, 11, 2016, pp. 40-62 58 across the state of california in paper archives, historical imagery, and/or physical biological specimens. these hidden “dark archives” are currently invisible to researchers, but with the kind of focused work described here, become invaluable. conclusions there is ample evidence that the rescuing, digitizing, and sharing of historical ecological data is an important scientific endeavor. these data provide benchmarks from which to compare change, they can be linked to modern ecological data to create new knowledge, and they can be modeled to help predict future changes. a. everett wieslander anticipated many modern uses of the vtm data in 1935. he wrote that it provided: 1) a partial explanation of the current (as of 1930s) distribution of vegetation types and dominant species; 2) a better understanding of vegetation changes that have occurred in the past, those now in progress, or those to be expected to occur in the future; and 3) further contributions to the knowledge of the values of certain plants and vegetation types as indicators of particular soil and climatic conditions (wieslander 1935b). he was perceptive in this analysis, but did not anticipate all of the uses of the data. the maps, plot data, and photographs have been used in isolation or paired with contemporary data to great effect to study california’s historical flora and land use, for documenting and finding mechanisms for decadalscale vegetation changes, and for predicting and planning for california’s future. the digitization and sharing of the vtm collection has expanded the scope and scale of possible analyses. any analyses larger than a single or few quads were impossible when the analog data were scattered around the state in libraries and research collections; or when plot or map data required laborious digitization. currently, most papers using vtm plot data explore the full complement of scale and detail. yet there is more that can be done. researchers might also use more of the vegetation map data, as well as the georeferenced photographs; they might embark on modeling that fuses data from more than one part of the collection, and synthesizes data from other sources. additionally, we hope more researchers will explore the connections between vegetation change and other biological signals of change such as isotopic signatures derived from spatially coincident animal specimens (e.g. rubidge et al. 2011; bi et al. 2013). finally, we want to highlight the increasingly critical role of data structures that foster scientific sharing and collaboration, such as apis, especially in their capacity to link at-risk historical data with contemporary ecological data. the digital vtm collection is an example of a web-based data framework that expands the potential of large-scale research through the integration and synthesis of data drawn from numerous data sources. we suggest here that the potential linkages and multidisciplinary connections waiting to be made with the use of the collection are numerous and important. understanding past, present, and future interrelationships between flora, fauna, land use, society, and climate is of paramount importance in ecology. the vtm dataset serves as a valuable and underutilized resource in this regard. the digital and shared data are an expansive historical reference that continues to provide exciting avenues for the modern geographer and ecologist to create connections between diverse datasets to tell the story of a region’s ecological past and better inform the future. acknowledgments the authors would like to thank several funders, including: the william m. keck foundation, berkeley institute for global change biology, uc division of agriculture and natural resources informatics and geographic information systems program, and the us forest service. much of the work was conducted in uc berkeley’s geospatial innovation facility. several individuals have worked on the digitization process early on; we are indebted to b. allen-diaz, k. ueda, b. thomas, t. de chant, e. zeledon, l. schile, k. tuxenbettman, and a. huber, among others. we thank joyce gross for use of her photo retakes. we also thank sarah hinman and the many undergraduate students who georeferenced the photographs and specimens in the museum of vertebrate zoology and the university and jepson herbaria. references allen, b. h., c. a. holzman, and r. r. evett. 1991. a classification system for california's hardwood rangelands. hilgardia 59:2. biodiversity informatics, 11, 2016, pp. 40-62 59 allen-diaz, b. h., and b. a. holzman. 1991. blue oak communities in california. madrono 38:80-95. baird, r. c. 2010. leveraging the fullest potential of scientific collections through digitization. biodivers. infor. 7:130-136. barbour, m. g., t. keeler-wolf, and a. a. schoenherr. 2007. terrestrial vegetation of california. univ of california press. beck, j., l. ballesteros‐mejia, c. m. buchmann, j. dengler, s. a. fritz, b. gruber, c. hof, f. jansen, s. knapp, and h. kreft. 2012. what's on the horizon for macroecology? ecography 35:673-683. bi, k., t. linderoth, d. vanderpool, j. m. good, r. nielsen, and c. moritz. 2013. unlocking the vault: next‐generation museum population genomics. mol. ecol. 22:6018-6032. borgman, c. l. 2012. the conundrum of sharing research data. j. am. soc. info. sci. tech. 63:10591078. bouldin, j. 1999. twentieth-century changes in forests of the sierra nevada, california. plant biol. university of california, davis. bradbury, d. 1974. vegetation history of the ramona quadrangle, san diego county, california (19311972). university of california, los angeles, los angeles. chornesky, e. a., d. d. ackerly, p. beier, f. w. davis, l. e. flint, j. j. lawler, p. b. moyle, m. a. moritz, m. scoonover, and k. byrd. 2015. adapting california's ecosystems to a changing climate. bioscience 65:247-262. collins, b., j. miller, m. kelly, j. w. van wagtendonk, and s. l. stephens. 2009. interactions among wildland fires in a long-established sierra nevada natural fire area. ecosystems 12:114-128. collins, b. m., m. kelly, j. w. van wagtendonk, and s. l. stephens. 2007. spatial patterns of large natural fires in sierra nevada wilderness areas. landscape ecol. 22:545-557. colwell, w. l. 1977. the status of vegetation mapping in california today. john wiley & sons, sacramento. conlisk, e., d. lawson, a. d. syphard, j. franklin, l. flint, a. flint, and h. m. regan. 2012. the roles of dispersal, fecundity, and predation in the population persistence of an oak (quercus engelmannii) under global change. plos one 7:e36. cox, r. d., k. l. preston, r. f. johnson, r. a. minnich, and e. b. allen. 2014. influence of landscape-scale variables on vegetation conversion to exotic annual grassland in southern california, usa. global ecol. conserv. 2:190-203. crimmins, s. m., s. z. dobrowski, j. a. greenberg, j. t. abatzoglou, and a. r. mynsberge. 2011. changes in climatic water balance drive downhill shifts in plant species’ optimum elevations. science 331:324-327. crimmins, s. m., s. z. dobrowski, and a. r. mynsberge. 2013. evaluating ensemble forecasts of plant species distributions under climate change. ecol. model. 266:126-130. crimmins, s. m., s. z. dobrowski, a. r. mynsberge, and h. d. safford. 2014. can fire atlas data improve species distribution model projections? ecol. appl. 24:1057-1069. critchfield, w. b. 1971. profiles of california vegetation. pp. 54. pacific southwest forest & range experiment station, forest service, u.s. department of agriculture, berkeley, ca. davis, j. d., and s. m. sims. 2013. physical and maximum entropy models applied to inventories of hillslope sediment sources. j. soils sed. 13:17841801. dobrowski, s. z., j. h. thorne, j. a. greenberg, h. d. safford, a. r. mynsberge, s. m. crimmins, and a. k. swanson. 2011. modeling plant ranges over 75 years of climate change in california, usa: temporal transferability and species traits. ecol. monogr. 81:241-257. dodge, j. m. 1975. vegetational changes associated with land use and fire history in san diego county. university of california, riverside. dolanc, c. r., h. d. safford, s. z. dobrowski, and j. h. thorne. 2013a. twentieth century shifts in abundance and composition of vegetation types of the sierra nevada, ca, us. appl. veg. sci. 17:442-455. dolanc, c. r., h. d. safford, j. h. thorne, and s. z. dobrowski. 2014. changing forest structure across the landscape of the sierra nevada, ca, usa, since the 1930s. ecosphere 5:1-26. dolanc, c. r., j. h. thorne, and h. d. safford. 2013b. widespread shifts in the demographic structure of subalpine forests in the sierra nevada, california, 1934 to 2007. global ecol. biogeogr. 22:264-276. elith, j., c. h. graham, r. p. anderson, m. dudik, s. ferrier, a. guisan, r. j. hijmans, f. huettmann, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. m. overton, a. t. peterson, s. j. phillips, k. richardson, r. scachetti-pereira, r. e. schapire, j. soberon, s. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species' distributions from occurrence data. ecography 29:129-151. elith, j., and j. r. leathwick. 2009. species distribution models: ecological explanation and prediction across space and time. annu. rev. ecol., evol. syst. 40:677. biodiversity informatics, 11, 2016, pp. 40-62 60 ertter, b. 2000. our undiscovered heritage: past and future prospects for species-level botanical inventory. madrono 47:237-252. fellows, a. w., and m. l. goulden. 2008. has fire suppression increased the amount of carbon stored in western u.s. forests? geophys. res. lett. 35:l12404. fox, p., and j. hendler. 2011. changing the equation on scientific data visualization. science 331:705-708. franklin, j. 2002. enhancing a regional vegetation map with predictive models of dominant plant species in chaparral. appl. veg. sci. 5:135-146. franklin, j., c. l. coulter, and s. j. rey. 2004. change over 70 years in a southern california chaparral community related to fire history. j. veg. sci. 15:701-710. frehner, m., and m. braendli. 2006. virtual database: spatial analysis in a web-based data management system for distributed ecological data. environ. model. software 21:1544-1554. freudenberger, d. o., b. e. fish, and j. e. keeley. 1987. distribution and stability of grasslands in the los angeles basin. bull. south. calif. acad. sci. 86:13-26. galatowitsch, s. m. 1990. using the original land survey notes to reconstruct presettlement landscapes in the american west. great basin nat. 50:181-191. goforth, b. r., and r. a. minnich. 2008. densification, stand-replacement wildfire, and extirpation of mixed conifer forest in cuyamaca rancho state park, southern california. for. ecol. manage. 256:36-45. graham, c. h., s. ferrier, f. huettman, c. moritz, and a. t. peterson. 2004. new developments in museum-based informatics and applications in biodiversity analysis. trends in ecology and evolution 19:497-503. griffin, j. r., and w. b. critchfield. 1972. the distribution of forest trees in california. u.s.d.a. forest service, pacific southwest forest and range experiment station, berkeley, ca. guisan, a., and w. thuiller. 2005. predicting species distribution: offering more than simple habitat models. ecol. lett. 8:993-1009. guo, q., m. kelly, and c. h. graham. 2005. support vector machines for predicting distribution of sudden oak death in california. ecol. model. 182:75-90. hampton, s. e., c. a. strasser, j. j. tewksbury, w. k. gram, a. e. budden, a. l. batcheller, c. s. duke, and j. h. porter. 2013. big data and the future of ecology. front. ecol. environ. 11:156-162. higgs, e., d. a. falk, a. guerrini, m. hall, j. harris, r. j. hobbs, s. t. jackson, j. m. rhemtulla, and w. throop. 2014. the changing role of history in restoration ecology. front. ecol. environ. 12:499506. jepson, w. l., r. beidleman, and b. ertter. 2000. willis linn jepson's "mapping in forest botany". madrono 47:269-272. jones, m. b., m. p. schildhauer, o. reichman, and s. bowers. 2006. the new bioinformatics: integrating ecological data from the gene to the biosphere. annu. rev. ecol., evol. syst. 37:519-544. kearns, f. r., m. kelly, and k. a. tuxen. 2003. everything happens somewhere: using webgis as a tool for sustainable natural resource management. front. ecol. environ. 1:541-548. keeler-wolf, t. 2007. the history of vegetation classification and mapping in california. terrestrial vegetation of california. university of california press, berkeley:1-42. keeley, j. e. 2004. vtm plots as evidence of historical change: goldmine or landmine? madrono 51:372378. kelly, a. e., and m. l. goulden. 2008. rapid shifts in plant distribution with recent climate change. proc. natl. acad. sci. 105:11823-11826. kelly, m., b. allen-diaz, and n. kobzina. 2005. digitization of a historic dataset: the wieslander california vegetation type mapping project. madrono 52:191-201. kelly, m., s. ferranto, s. lei, k. ueda, and l. huntsinger. 2012. expanding the table: the web as a tool for participatory adaptive management in california forests. j. environ. manage. 109:1-11. kelly, m., k. ueda, and b. allen-diaz. 2008. considerations for ecological reconstruction of historic vegetation: analysis of the spatial uncertainties in the california vegetation type mapping dataset. plant ecol. 194:37-49. kittinger, j. n., k. s. v. houtan, l. e. mcclenachan, and a. l. lawrence. 2013. using historical data to assess the biogeography of population recovery. ecography 36:868-872. krebs, c. j., r. boonstra, s. boutin, and a. r. e. sinclair. 2001. what drives the 10-year cycle of snowshoe hares? bioscience 51:25-35. kuchler, a. w. 1967. vegetation mapping. ronald press, new york. leiberg, j. b. 1902. forest conditions in the northern sierra nevada, california. department of the interior, us geological survey. lenoir, j., j.-c. gégout, p. marquet, p. de ruffray, and h. brisse. 2008. a significant upward shift in plant species optimum elevation during the 20th century. science 320:1768-1771. lippitt, c. l., d. a. stow, j. f. o’leary, and j. franklin. 2013. influence of short-interval fire occurrence on post-fire recovery of fire-prone biodiversity informatics, 11, 2016, pp. 40-62 61 shrublands in california, usa. int. j. wildland fire 22:184-193. loarie, s. r., b. e. carter, k. hayhoe, s. mcmahon, r. moe, c. a. knight, and d. d. ackerly. 2008. climate change and the future of california’s endemic flora plos one 3:e2502. lutz, j. a., j. w. van wagtendonk, and j. f. franklin. 2009. twentieth-century decline of large-diameter trees in yosemite national park, california, usa. for. ecol. manage. 257:2296-2307. lutz, j. a., j. w. van wagtendonk, and j. f. franklin. 2010. climatic water deficit, tree species ranges, and climate change in yosemite national park. j. biogeogr. 37:936-950. maxwell, r. s., a. h. taylor, c. n. skinner, h. d. safford, r. e. isaacs, c. airey, and a. b. young. 2014. landscape-scale modeling of reference period forest conditions and fire behavior on heavily logged lands. ecosphere 5:1-28. mayer, k. e., and w. f. laudenslayer. 1988. a guide to wildlife habitats of california. pp. 166. state of california, resources agency, department of fish and game, sacramento, ca. mcintyre, p. j., j. h. thorne, c. r. dolanc, a. l. flint, l. e. flint, m. kelly, and d. d. ackerly. 2015. twentieth-century shifts in forest structure in california: denser forests, smaller trees, and increased dominance of oaks. proc. natl. acad. sci. 112:1458-1463. mcminn, h. 1951. an illustrated manual of california shrubs. univ of california press. michener, w. k. 2006. meta-information concepts for ecological data management. ecol. infor. 1:3-7. michener, w. k., j. w. brunt, j. j. helly, t. b. kirchner, and s. g. stafford. 1997. nongeospatial metadata for the ecological sciences. ecol. appl. 7:330-342. minnich, r., and r. dezzani. 1998. historical decline of coastal sage scrub in the riverside-perris plain, california. west. birds 29:366-391. minnich, r. a. 1978. the geography of fire and conifer forests in the eastern transverse ranges, california. university of california, los angeles. minnich, r. a., m. g. barbour, j. h. burk, and r. f. fernau. 1995. sixty years of change in californian conifer forests of the san bernardino mountains. conserv. biol. 9:902-914. mladenoff, d. j., s. e. dahir, e. v. nordheim, l. a. schulte, and g. g. guntenspergen. 2002. narrowing historical uncertainty: probabilistic classification of ambiguously identified tree species in historical forest survey data. ecosystems 5:539533. peters, d. p. 2010. accessible ecology: synthesis of the long, deep, and broad. trends in ecology and evolution 25:592-601. phillips, s. j., r. p. anderson, and r. e. schapire. 2006. maximum entropy modeling of species geographic distributions. ecol. model. 190:231-259. preston, k. l., r. a. redak, m. f. allen, and j. t. rotenberry. 2012. changing distribution patterns of an endangered butterfly: linking local extinction patterns and variable habitat relationships. biol. conserv. 152:280-290. rapacciuolo, g., d. b. roy, s. gillings, and a. purvis. 2014a. temporal validation plots: quantifying how well correlative species distribution models predict species' range changes over time. methods ecol. evol. 5:407-420. rapacciuolo g., s.p. maher, a.c. schneider, t.t. hammond, m.d. jabis, r.e. walsh, k.j. iknayan, g.k. walden, m.f. oldfather, d.d. ackerly, and s.r. beissinger. 2014b. beyond a warming fingerprint: individualistic biogeographic responses to heterogeneous climate change in california. global change biology 20: 2841–2855. reichman, o., m. b. jones, and m. p. schildhauer. 2011. challenges and opportunities of open data in ecology. science 331:703-705. rubidge, e. m., w. b. monahan, j. l. parra, s. e. cameron, and j. s. brashares. 2011. the role of climate, habitat, and species co-occurrence as drivers of change in small mammal distributions over the past century. global change biol. 17:696708. santos, m. j., j. h. thorne, j. christensen, and z. frank. 2014a. an historical land conservation analysis in the san francisco bay area, usa: 1850–2010. landscape urban plann. 127:114-123. santos, m. j., j. h. thorne, and c. moritz. 2014b. synchronicity in elevation range shifts among small mammal and vegetation over the last century is stronger for omnivores. ecography 37:001-013. sawyer, j., and t. keeler-wolf. 1995. a manual of california vegetation. california native plant society press, sacramento, california. schulte, l. a., and d. j. mladenoff. 2001. the original us public land survey records: their use and limitations in reconstructing presettlement vegetation. j. for. october:5-10. schwind, b., and h. gordon. 2001. calveg geobook: a comprehensive information package describing california’s wildland vegetation, version 2. usda forest service, pacific southwest region, remote sensing lab, sacramento, ca. solé, r. v., and j. bascompte. 2006. self-organization in complex ecosystems. princeton university press. stein, e. d., s. dark, t. longcore, r. grossinger, n. hall, and m. beland. 2010. historical ecology as a tool for assessing landscape change and informing biodiversity informatics, 11, 2016, pp. 40-62 62 wetland restoration priorities. wetlands 30:589601. swanson, a. k., s. z. dobrowski, a. o. finley, j. h. thorne, and m. k. schwartz. 2013. spatial regression methods capture prediction uncertainty in species distribution model projections through time. global ecol. biogeogr. 22:242-251. syphard, a. d., and j. franklin. 2009. differences in spatial predictions among species distribution modeling methods vary with species traits and environmental predictors. ecography 32:907-918. syphard, a. d., and j. franklin. 2010. species traits affect the performance of species distribution models for plants in southern california. j. veg. sci. 21:177-189. talluto, m. v., and k. n. suding. 2008. historical change in coastal sage scrub in southern california, usa in relation to fire frequency and air pollution landscape ecol. 23:803-815. taylor, a. h. 2000. fire regimes and forest changes in mid and upper montane forests of the southern cascades, lassen volcanic national park, california, usa. j. biogeogr. 27:87-104. tenopir, c., s. allard, k. douglass, a. u. aydinoglu, l. wu, e. read, m. manoff, and m. frame. 2011. data sharing by scientists: practices and perceptions. plos one 6:e21101. thorne, j., t. r. kelsey, j. honig, and b. morgan. 2007. the development of 70-year-old wieslander vegetation type maps and an assessment of landscape change in the central sierra nevada. california energy commission, sacramento, ca. thorne, j. h., j. a. kennedy, j. f. quinn, m. mccoy, t. keeler-wolf, and j. menke. 2004. a vegetation map of napa county using the manual of california vegetation classifications and its comparison to other digital vegetation maps. madrono 51:343363. thorne, j. h., and t. n. g. le. in press. the wieslander vegetation type maps, an historic legacy for california landscape change analyses. madrono. thorne, j. h., b. j. morgan, and j. a. kennedy. 2008. vegetation change over sixty years in the central sierra nevada, california, usa. madrono 55:223237. thorne, j. h., m. j. santos, and j. h. bjorkman. 2013. regional assessment of urban impacts on landcover and open space finds a smart urban growth policy performs little better than business as usual. plos one 8:e65258. tingley, m. w., and s. r. beissinger. 2009. detecting range shifts from historical species occurrences: new perspectives on old data. trends in ecology and evolution 24:625-633. united states forest service. 2016a. calveg (classification and assessment with landsat of visible ecological groupings). united states forest service. 2016b. pacific northwest forest inventory and analysis. vayssières, m. p., r. e. plant, and b. h. allen-diaz. 2000. classification trees: an alternative nonparametric approach for predicting species distributions. j. veg. sci. 11:679-694. walker, r. e. 2000. investigations in vegetation map rectification, and the remotely sensed detection and measurement of natural vegetation changes. pp. 249. geography. university of california, santa barbara, santa barbara, ca. whipple, a. a., r. m. grossinger, and f. w. davis. 2011. shifting baselines in a california oak savanna: nineteenth century data to inform restoration scenarios. restor. ecol. 19:88-101. wieslander, a. e. 1935a. the first steps of the forest survey in california. j. for. 33:877-884. wieslander, a. e. 1935b. a vegetation type map of california. madrono 3:140-144. wieslander, a. e. 1961. california's vegetation maps: recent advances in botany. pp. 4. university of toronto press, toronto, canada. wieslander, a. e. 1986. a.e. wieslander, california forester: mapper of wildland vegetation and soils (an oral history conducted in 1985 by ann lange). regional oral history office, bancroft library, university of california, berkeley, berkeley, ca. wolf, a., n. b. zimmerman, w. r. anderegg, p. e. busby, and j. christensen. 2016. altitudinal shifts of the native and introduced flora of california in the context of 20th‐century warming. global ecol. biogeogr. wolkovich, e. m., j. regetz, and m. i. o'connor. 2012. advances in global change research require open science by individual researchers. global change biol. 18:2102-2110. biodiversity informatics, 18, 2024, pp. 28-42 28 enmpa: an r package for ecological niche modeling using presence-absence data and generalized linear models luis f. arias-giraldo1,2,*, marlon e. cobos3 ¹institute for sustainable agriculture, spanish national research council (csic), córdoba, spain. 2 programa de doctorado ingeniería agraria, alimentaria, forestal y de desarrollo rural sostenible. university of córdoba, córdoba, spain. 3department of ecology and evolutionary biology & biodiversity institute, university of kansas, lawrence, ks, united states of america. the authors contributed equally to this work *corresponding author: luis f. arias-giraldo e-mail: lfarias.giraldo@gmail.com abstract. here, we present the new r package “enmpa,” which includes a range of tools for modeling ecological niches using presence-absence data via logistic generalized linear models. the package allows users to calibrate, select, project, and evaluate models using independent data. we have emphasized a comprehensive search for ideal predictor combinations, including linear, quadratic, and two-way interaction responses, to provide more detailed and robust model calibration processes. we demonstrate the use of the package with an example of a simulated pathogen and its niche. since enmpa is designed specifically to work with presence-absence data, our tools are particularly useful for studies with data derived from a detection or non-detection sampling universe, such as pathogen testing results. enmpa can be downloaded from cran, and the source code is freely available on github. key words: ecological niche modeling, general linear models, open-source software, model calibration, model selection, model projections. introduction ecological niche modeling (enm), often also referred to as species distribution modeling (sdm), constitutes a range of analytical methods employed extensively in ecological research (guisan and zimmermann 2000; franklin 2010; peterson et al. 2011). these methods have been proven particularly useful in characterizing and predicting grinnellian (abiotic or non-interactive) ecological niches of species. applications of these methods span various fields, including conservation planning (franklin 2013; hannah et al. 2020), climate change impact assessment (searcy and shaffer 2016; blowes et al. 2019), potential biological invasions (jiménezvalverde et al. 2011; park and potter 2015; cordier et al. 2020), and disease risk mapping (peterson 2014). several modeling methods are available within the enm framework, which can be classified by the types of data that they use: presence-only, presence and background (or “pseudoabsence”), and presenceabsence data (elith et al. 2006; peterson et al. 2011). more generally, methods for enm fall into three broad categories: ‘profile,’ ‘regression,’ or ‘machinelearning.’ profile methods consider presence data only. regression and machine-learning methods use both presence-absence or presence-background data. an essential consideration in ecological studies is understanding the meaning of the outputs generated by algorithms. a common objective of these studies is to model the probability of presence of a particular species of interest. however, it is crucial to note that estimating probabilities of occurrence requires rigorous comparisons of presence and absence data (ward et al. 2009). modeling applications that utilize presence-only data can, at best, estimate relative suitability (ferrier et al. 2002). although availability of occurrence data poses a significant challenge in ecological niche modeling, biodiversity data-portals mailto:lfarias.giraldo@gmail.com arias-giraldo et al. – enmpa 29 like the global biodiversity information facility (gbif) can serve as valuable sources of occurrence data records, at least potentially including presences and absences. among the different modeling methods that are used in enm, generalized linear models (glms), an extension of classical multiple regression, have yielded reliable results in ecological research in estimating probability of occurrence of species (guisan et al. 2002; bolker et al. 2009; rupprecht et al. 2011; ghanbarian et al. 2019). determination of species’ responses to environmental gradients is of particular interest for most biological questions in enm, which is given by the shapes of response curves (guisan et al. 2002; oksanen and minchin 2002; austin 2007; santika and hutchinson 2009). fitting glms via a logistic link function for presence-absence binomial data offers results approximating gaussian curves according to the principles of ecological niche theory (austin 2007; santika and hutchinson 2009). however, response curves in glms may be inappropriate or unrealistic when models are not tuned adequately (austin et al. 1990). it is crucial to explore and determine appropriate parameter settings for models, instead of simply using default settings (warren and seifert 2011). this task can be accomplished through model calibration and selection processes (radosavljevic and anderson 2014; hao et al. 2020). models resulting from calibration exercises help to describe the phenomenon of interest better while achieving a robust fit to the data, high predictive performance, and generalizable model terms (cobos et al. 2019a). to our knowledge, well-defined tuning routines have yet to be developed for glms in enm, unlike methods like maxent (phillips et al. 2006; muscarella et al. 2014; phillips et al. 2017; cobos et al. 2019a). to bridge this methodological gap (glm tuning routines for enm), we introduce the enmpa r package (r core team 2022). this package is designed to refine the calibration and parameter tuning processes, offering a solution to the challenges in modeling robust ecological niches with presenceabsence data. via enmpa, we explore the entire range of possible model configurations to identify the most suitable and practical parameterizations for enms. package descripton the enmpa r package provides a set of tools to automate various enm steps using logistic regressions via glms, focusing on fitting linear and quadratic relationships and multiplicative interactions between predictors. the response (dependent) variable is a set of presence and absence records, and the predictor (independent) variables can be defined according to the question (e.g., bioclimatic variables). major steps enabled via enmpa include model calibration (candidate model fitting and evaluation), model selection, model transfers, and model evaluation with independent data. data required the input data required consist of presenceabsence records associated with values of the independent variables. the dependent variable is the set of presence-absence observations. the data must be structured as a data.frame in r to ensure the proper functioning of the package. raster layers are required if model predictions need to be done for geographic areas of interest. exploration of variables for models we adapted methods developed by cobos and peterson (2022) to identify relevant variables for characterizing species’ ecological niches. these methods include two complementary statistical analyses: (1) a multivariate approach based on a permutational multivariate analysis of variance (permanova) (anderson 2017) and (2) a univariate non-parametric method based on descriptive statistics and randomizations. these methods are designed to allow characterizations of signals of ecological niches while considering the sampling universe explicitly. as enmpa uses presence-absence data through logistic regression, we consider it appropriate to implement this method as a potential variable selection step prior to modeling. together with proper considerations of variable biological relevance, this step can help to reduce initial numbers of predictors considered for enm analysis (cobos et al. 2019b). model calibration the model calibration step aims to determine which combination of parameter settings best represents the phenomenon of interest via exploration of performance metrics that characterize how well models fit the data (steele and werndl 2013). the tools in enmpa automate a process that includes three main steps: (1) fitting candidate models with distinct parameter settings, (2) evaluating their performance, arias-giraldo et al. – enmpa 30 and (3) selecting the most robust candidates based on predefined criteria. candidate model fitting in this package, we propose exploring distinct parameter settings by producing multiple model formulas to fit models to the data. these formulas derive from combinations of predictors that can be obtained using the original independent variables and response types: linear (l), quadratic (q), and product (p). this approach allows exhaustive exploration of all possible combinations of predictors, enabling a detailed analysis of the entire predictor setting space (cobos et al. 2019b). users can produce all these formulas manually or use functions in enmpa designed to automate the process considering two main inputs: variable names and response types. users can also define the permutation strategy to create formulas according to the desired level of intensiveness in exploring setting options (e.g., only increasing complexity of variable combinations, or all independent and combinatorial options). candidate model evaluation to evaluate candidate models, we use three complementary approaches. the first tests predictive power using a k-fold cross-validation approach (hastie et al. 2009). the original dataset is partitioned into k subsets (folds) aiming for equal size and maintaining the original prevalence (ratio of presences and absences). the algorithm performs k iterations of training, in which each iteration uses k 1 folds for training and keeps one fold for testing. this process evaluates model discrimination and classification capacities. discrimination is measured using the area under the receiver operating curve (roc-auc), a nonthreshold dependent metric used, in our case, solely to detect models that perform better than random expectations (lobo et al. 2008). classification ability is measured via several metrics deriving from an estimated confusion matrix, including sensitivity, specificity, accuracy, false positive rate, and true skill statistic (tss), all of them according to three thresholds (i.e., equal sensitivity and specificity, sensitivity of 90%, and maximum tss) (fielding and bell 1997; manel et al. 2001; allouche et al. 2006; liu et al. 2011). means and standard deviations of these metrics are calculated to summarize the model’s predictive performance and help to select the best models. the second approach uses the akaike information criterion (aic) (akaike 1998; warren and seifert 2011; warren et al. 2014) to assess model goodness-of-fit, accounting for model complexity. this metric offers a relative quality measure to other candidate models based on the same dataset with different parameters. aic increases with information loss, so the best model for a set of occurrence data is the one with the lowest aic. to compare aic values of multiple models directly, we also calculate δaic (wagenmakers and farrell 2004) by subtracting the aic of the best model (the one with the lowest aic) from the aic of each model being compared, as follows: δaici = aici − aicmin. a δaic of 0 indicates that the model in question is the best model; models with δaic values ≤2 have substantial support, and can be considered almost as good as those with the lowest aic. subsequently, akaike weights are derived from δaic for model averaging, representing a model’s relative likelihood: akaike weight (wi) is calculated as an exponent of the negative half of its δaic value, and the relative likelihoods are normalized so that their sum across all models compared equals 1. this later step is achieved by dividing the relative likelihood of each model by the sum of the relative likelihoods for all models, producing the akaike weight for each model. we note that the implementation of aic calculations for glm consider models goodness-of-fit, whereas that for maxent, which has seen considerable use, aic is based on the model predictions (warren and seifert 2011). finally, enmpa incorporates an extra evaluation step involving analyzing the response curves of quadratic terms. quadratic features are optimal in studies of responses of species to variable gradients, and fit well with ecological niche theory (austin 2007; santika and hutchinson 2009). however, a limitation of quadratic features is that they can be concave upward, yielding a bimodal response, which does not fit well with theory (austin 2007). by investigating the coefficients of a second-degree equation (y = β0 + β1x + β2x 2), we can infer the shape of the curve. a positive β2 suggests a u-shaped bimodal curve, while a negative β2 suggests a gaussian-shaped unimodal curve. as such, we implemented a filter to retain only those candidate models in which any quadratic responses are β2-negative. model selection to choose the best candidate models, we follow a set of criteria, which are prioritized as follows: (1) arias-giraldo et al. – enmpa 31 we only consider models with roc-auc > 0.5; (2) from among such models, we only keep those that have an acceptable predictive ability (tss ≥ 0.4); and (3) from among all the models passing the first two filters, we chose those with good fitting and appropriate complexity, those with δaic ≤ 2 (burnham and anderson 1998). in addition, enmpa includes an optional filter that takes into account the shape of quadratic response curves. users have the option to consider only models with predictors that have unimodal responses. if this filter is used, it is applied before the first three filters. even if the bimodal response of species could have interesting and meaningful interpretations, it is tricky to determine optima and understand species’ tolerances, which need to be evaluated individually for particular species. therefore, we recommend considering only models with unimodal or monotonic responses. variable contribution and response curves variable contribution and analysis response curves are model outputs with practical relevance to researchers interested in interpreting model outputs. two well-established methods for determining the importance of predictor variables in glms are used to evaluate the individual contributions of variables (murray and conner 2009) and visualize predicted responses of species to specific predictor variables (elith et al. 2005). a response curve represents the relationship between the probability of occurrence of a species (dependent variable) and environmental variables (independent variables) included in the model. this curve describes the predicted probabilities across a range of values for a given environmental variable. the enmpa package calculates the probabilities along a single environmental gradient to estimate response curves while holding all other gradients constant at their mean values (elith et al. 2005). to identify the most relevant predictors in our models, we use a variable contribution analysis based on the deviance explained by predictors relative to the complete model deviance (guisan and zimmermann 2000; clouvel et al. 2023). deviance is a measure of the model’s lack of fit, so a decrease in deviance indicates an improvement in model fit when a predictor variable is included. to assess predictor importance, then, we implemented the following procedure: (1) a glm including all predictors is fitted; (2) the initial deviance is calculated; (3) each predictor is removed iteratively from the model, and a new deviance is calculated; (4) the decrease in deviance after adding each predictor is calculated; (5) the decrease in deviance is normalized and expressed as a ratio; and (6) predictors are ranked based on how they help to decrease model deviance: variables with higher ratios are considered more important for the model. model projections the enmpa package facilitates transferring selected models to different areas or scenarios, with three options: free extrapolation, extrapolation with clamping, and no extrapolation. these options are available for all variables or for just selected variables. when using free extrapolation, predictions will follow the response patterns when variable values are outside the ranges of the environmental data under which models were calibrated. in contrast, extrapolation with clamping limits the response to the level manifested at the boundaries of calibration values. model consensus one way of selecting a model is to choose the “best” one for the data based on one or a set of predictive performance metrics (elith et al. 2006). however, an alternative method is to use a consensus of models (thuiller 2003; qiao et al. 2015). consensus approaches may provide more robust predictions by leveraging the general agreement among models with similar performance. a consensus result can be calculated as the mean, median, or weighted average of any set of models. mean and median results are calculated using the predictions of all selected models, whereas a weighted average is calculated using these predictions and the aic weights for each model. models with higher aic weights contribute more to the consensus when the weighted average option is used. model evaluation with independent data ideally, the final model should be evaluated using independent test data, although it is often challenging to find data independent from those used to create the models (araújo and guisan 2006; peterson et al. 2011). in cases in which independent data are available, enmpa facilitates evaluation of final model predictions. users can provide presence-absence or arias-giraldo et al. – enmpa 32 functions group description niche_signal & plot_ niche_signal analysis/ visualization this implementation determines the detectability of a niche signal by analyzing one or multiple variables. it is based on the methodologies introduced by cobos & peterson (2022), specifically tailored for discerning niche signals in presence-absence data. a niche signal is defined as the non-random pattern in species detection concerning environmental dimensions, taking sampling into account. the accompanying function generates plots to facilitate interpreting results obtained from niche_signal tests. get_formulas analysis provides a simple and efficient way to create a comprehensive set of standardized glm formulas for statistical models. with this function, users can explore the entire range of formula combinations. it considers the following feature classes: linear (l), quadratic (q), and product (p) responses. model_validation analysis evaluates a glm model using an entire data set or with a k-fold cross-validation procedure. models are assessed based on discrimination ability (roc-auc), classification capability (fpr, accuracy, sensitivity, specificity, and tss), and the balance between goodness-of-fit and complexity (aicc). optimize_metrics analysis finds threshold values to produce three optimal metrics. the metrics true skill statistic (tss), sensitivity, and specificity are explored by comparing actual vs predicted values to find threshold values that produce sensitivity = specificity, maximum tss, and a sensitivity value of 0.9. calibration_glm analysis wrapped function that automates and simplifies the exploration, fitting, evaluation, and selection of robust models in the entire parameter space. fit_selected & predict_ selected analysis the initial function streamlines the process of fitting numerous generalized linear models (glms), while the subsequent function aids in forecasting the chosen models across diverse time frames or geographic areas. within the predict_selected function,two options, clamping and no extrapolation, are integrated to mitigate potential undesired extrapolation effects when fresh data extends beyond the calibrated range of the model. furthermore, this tool can construct consensus models derived from the predicted models. this consensus is formed by employing statistical metrics such as mean, median, or weighted mean. var_importance & plot_ importance analysis/ visualization the initial function computes the relevance of predictor variables based on explained deviance. the second function generates straightforward graphics demonstrating the outcomes for single or multiple models. these visuals represent the relative importance of each predictor variable within a model, facilitating a comparison of their contributions across various models. response_curve visualization visualization of the response of a variable based on a single model or multiple models. it illustrates the probabilities of species presence across a diverse range of values for a specific environmental variable. this visualization enhances understanding of how species occur as a function of the environmental conditions captured by the variable. independent_eval1 & independent_eval01 analysis these functions evaluate models using independent datasets, generating tailored metrics for the specific occurrence data available. for presence-only data, the evaluation includes partial roc and omission error metrics. for presence-absence data, it focuses on the auc-roc curve and classification capability, assessed through a confusion matrix and three threshold criteria: maxtss, ess, and sen90. table 1. description of the main functions included in the r package enmpa. arias-giraldo et al. – enmpa 33 presence-only records for this evaluation. evaluation in cases involving presence-only data includes partial roc and omission error (e) metrics (cobos et al. 2019a). for presence-absence data, the metrics returned are the roc auc, sensitivity, specificity, accuracy, false positive rate, and true skill statistic (tss), and for three thresholds (maximum tss, equal sensitivity and selectivity, and a value that ensures a sensitivity of 90%). example application here, we provide a guide on how to use enmpa, in the form of a worked example. this example includes the processes of variable exploration, model calibration and selection, transfers to specific areas and scenarios of interest, and post-modeling analyses. the code required to reproduce this example can be found in a github repository (available at https://github.com/ luisagi/enmpa_test). the presence-absence data, raster layers, and an independent test dataset used in these examples are included in the package (available at https://cran.r-project.org/package=enmpa). data the example data include 500 presence records of a virtual pathogen detected in a total of 5627 virtual host samples (i.e., 5127 absences), for a prevalence of 8.9%. the example data was generated based on ellipsoidal virtual niches for the host and the pathogen using the package evniche (https://github.com/ marlonecobos/evniche) in r v4.2.2 (r core team 2022). the pathogen was designed to have a higher prevalence towards warm and dry environmental conditions. example data to illustrate independent tests were generated the same way but during a posterior set of analyses. the environmental data related to occurrence records were extracted from raster layers (resolution = 10 arc-minutes) that represent two bioclimatic variables (annual mean temperature and annual precipitation, bio-1 and bio-12), from the worldclim database v2.0 (fick and hijmans 2017). the two datasets associated with environmental values are included as part of example data in enmpa. analysis variable exploration for niche signal detection.-we started with exploratory analyses to detect whether the environmental variables can describe the niche of the virtual pathogen (cobos and peterson 2022). we conducted the two tests (multivariate and univariate approaches) to assess whether the position and spread of the pathogen niche differed from those of the host, considering the two environmental variables in the example. calibration and model selection.--in all 31 candidate models were created using combinations of the two environmental variables, with linear, quadratic, and product responses. each model was tested for model complexity (aic) using the whole dataset and with a k-fold cross-validation (k = 5) to evaluate its performance in terms of discrimination (roc-auc) and classification capacity (false positive rate, accuracy, sensitivity, specificity, and tss). the classification metrics were based on three relevant thresholds: equal sensitivity and specificity (ess), the sensitivity of 90% (sen90), and maximum tss (maxtss). the best models were selected according to enmpa criteria: only models with only convex quadratic responses were assessed, then, models with roc-auc > 0.5 were retained, and after that, only models with a good classification capacity (tss > 0.4) were selected. finally, only the models with δaicc values ≤ 2 were chosen as the final selected models. post-modeling analyses.--we analyzed the contributions of predictors to the model and explored variable response curves for each model and the consensus results. projections to the 48 contiguous united states were produced for all models selected using the aforementioned criteria. we produced model projections in conditions outside calibration ranges allowing for free extrapolation, extrapolation with clamping, and no extrapolation. we used three approaches to generate consensus results: mean, median, and weighted average. to represent model variability, we also calculated the variance among selected models. finally, the independent data were used to evaluate the models selected and the consensus results. example results both tests for environmental sensitivity consistently performed well for the two bioclimatic variables for the virtual pathogen case. multivariate analysis based on permanova effectively detected niche dissimilarities between the pathogen and host, with the pathogen niche forming a subgroup nested within the host niche (fig. a1 of appendix). using the mean as the comparison metric for the univariate non-parametric test, we found that the pathogen niche was shifted towards high temperature and https://github.com/luisagi/enmpa_test https://github.com/luisagi/enmpa_test https://cran.r-project.org/package=enmpa https://github.com/marlonecobos/evniche https://github.com/marlonecobos/evniche arias-giraldo et al. – enmpa 34 lower precipitation values. comparisons using the standard deviation showed that the pathogen niche was narrower than the host niche in both variables (fig. a2 of appendix). after model calibration, eight models were excluded because they presented concave quadratic responses. the remaining 23 models met the next criterion with a roc-auc larger than 0.5 (table 2). of these models, 19 performed well, with a tss larger than 0.4. finally, only two models met all of the evaluation criteria regarding discrimination, prediction ability, and fitting considering complexity (δaic ≤ 2). response curves for the two climatic variables presented a well-defined gaussian shape (fig. 1). the peak probability for bio-1 occurs at id models threshold roc-auc sensitivity specificity tss aicc δ aic wi bimodality 1 b1 0.07 ± 0.01 0.87 ± 0.02 0.86 ± 0.02 0.73 ± 0.04 0.59 ± 0.03 2509.56 323.88 0.00 2 b12 0.09 ± 0.02 0.69 ± 0.03 0.76 ± 0.16 0.54 ± 0.18 0.30 ± 0.05 3164.57 978.89 0.00 3 i(b1^2) 0.06 ± 0.01 0.86 ± 0.02 0.86 ± 0.03 0.73 ± 0.03 0.58 ± 0.03 2599.64 413.96 0.00 i(b1^2) 4 i(b12^2) 0.10 ± 0.02 0.69 ± 0.03 0.76 ± 0.16 0.54 ± 0.18 0.30 ± 0.04 3191.21 1005.53 0.00 5 b1:b12 0.09 ± 0.00 0.65 ± 0.03 0.71 ± 0.09 0.58 ± 0.11 0.29 ± 0.04 3373.98 1188.30 0.00 6 b1 + b12 0.08 ± 0.01 0.90 ± 0.02 0.85 ± 0.03 0.82 ± 0.03 0.67 ± 0.04 2334.00 148.32 0.00 7 b1 + i(b1^2) 0.07 ± 0.02 0.87 ± 0.02 0.87 ± 0.03 0.73 ± 0.05 0.60 ± 0.03 2464.79 279.11 0.00 8 b1 + i(b12^2) 0.09 ± 0.01 0.90 ± 0.02 0.84 ± 0.03 0.83 ± 0.03 0.68 ± 0.04 2308.20 122.52 0.00 9 b1 + b1:b12 0.08 ± 0.01 0.90 ± 0.02 0.85 ± 0.03 0.82 ± 0.04 0.67 ± 0.04 2370.91 185.23 0.00 10 b12 + i(b1^2) 0.08 ± 0.00 0.90 ± 0.02 0.85 ± 0.04 0.82 ± 0.01 0.66 ± 0.05 2431.46 245.78 0.00 i(b1^2) 11 b12 + i(b12^2) 0.09 ± 0.03 0.69 ± 0.03 0.76 ± 0.16 0.54 ± 0.18 0.30 ± 0.05 3163.68 978.00 0.00 i(b12^2) 12 b12 + b1:b12 0.10 ± 0.02 0.89 ± 0.02 0.82 ± 0.03 0.82 ± 0.03 0.65 ± 0.06 2289.81 104.13 0.00 13 i(b1^2) + i(b12^2) 0.08 ± 0.01 0.90 ± 0.02 0.84 ± 0.04 0.83 ± 0.03 0.67 ± 0.05 2411.86 226.18 0.00 i(b1^2) 14 i(b1^2) + b1:b12 0.07 ± 0.01 0.90 ± 0.02 0.86 ± 0.04 0.80 ± 0.04 0.67 ± 0.04 2500.62 314.94 0.00 i(b1^2) 15 i(b12^2) + b1:b12 0.11 ± 0.01 0.90 ± 0.02 0.84 ± 0.03 0.84 ± 0.03 0.68 ± 0.04 2282.01 96.33 0.00 16 b1 + b12 + i(b1^2) 0.09 ± 0.02 0.90 ± 0.02 0.86 ± 0.04 0.81 ± 0.04 0.68 ± 0.04 2257.71 72.03 0.00 17 b1 + b12 + i(b12^2) 0.09 ± 0.01 0.90 ± 0.02 0.85 ± 0.04 0.83 ± 0.03 0.68 ± 0.04 2295.67 109.99 0.00 18 b1 + b12 + b1:b12 0.10 ± 0.01 0.89 ± 0.02 0.82 ± 0.03 0.83 ± 0.03 0.65 ± 0.05 2281.90 96.22 0.00 19 b1 + i(b1^2) + i(b12^2) 0.10 ± 0.02 0.90 ± 0.02 0.85 ± 0.05 0.83 ± 0.03 0.68 ± 0.04 2227.10 41.42 0.00 20 b1 + i(b1^2) + b1:b12 0.09 ± 0.02 0.89 ± 0.01 0.86 ± 0.04 0.81 ± 0.03 0.67 ± 0.03 2287.32 101.64 0.00 21 b1 + i(b12^2) + b1:b12 0.10 ± 0.01 0.90 ± 0.02 0.83 ± 0.04 0.85 ± 0.03 0.68 ± 0.04 2243.06 57.38 0.00 22 b12 + i(b1^2) + i(b12^2) 0.08 ± 0.00 0.90 ± 0.02 0.84 ± 0.04 0.84 ± 0.03 0.68 ± 0.04 2407.17 221.49 0.00 i(b1^2) 23 b12 + i(b1^2) + b1:b12 0.10 ± 0.02 0.89 ± 0.02 0.83 ± 0.03 0.82 ± 0.03 0.65 ± 0.06 2291.80 106.12 0.00 i(b1^2) 24 b12 + i(b12^2) + b1:b12 0.10 ± 0.00 0.89 ± 0.02 0.84 ± 0.04 0.82 ± 0.01 0.66 ± 0.05 2241.64 55.96 0.00 25 i(b1^2) + i(b12^2) + b1:b12 0.11 ± 0.01 0.90 ± 0.02 0.83 ± 0.03 0.84 ± 0.03 0.68 ± 0.04 2268.56 82.88 0.00 i(b1^2) 26 b1 + b12 + i(b1^2) + i(b12^2) 0.11 ± 0.02 0.90 ± 0.02 0.85 ± 0.04 0.83 ± 0.03 0.68 ± 0.04 2212.31 26.63 0.00 27 b1 + b12 + i(b1^2) + b1:b12 0.09 ± 0.01 0.90 ± 0.02 0.86 ± 0.02 0.80 ± 0.02 0.67 ± 0.04 2226.49 40.81 0.00 28 b1 + b12 + i(b12^2) + b1:b12 0.10 ± 0.01 0.90 ± 0.02 0.84 ± 0.04 0.82 ± 0.02 0.66 ± 0.05 2237.05 51.37 0.00 29* b1 + i(b1^2) + i(b12^2) + b1:b12 0.10 ± 0.02 0.90 ± 0.02 0.86 ± 0.04 0.82 ± 0.03 0.68 ± 0.04 2186.70 01.02 0.38 30 b12 + i(b1^2) + i(b12^2) + b1:b12 0.10 ± 0.01 0.89 ± 0.02 0.84 ± 0.05 0.82 ± 0.02 0.66 ± 0.05 2243.56 57.88 0.00 31* b1 + b12 + i(b1^2) + i(b12^2) + b1:b12 0.10 ± 0.02 0.90 ± 0.02 0.86 ± 0.04 0.83 ± 0.02 0.68 ± 0.04 2185.68 0.00 0.62 table 2. model calibration summary. the table displays the evaluation of the main metrics for the 31 candidates evaluated using a cross-validated k-fold (k = 5) analysis. the most robust models selected using three criteria implemented in enmpa are indicated in bold. roc-auc, sensitivity, specificity, and tss values are presented as mean ± sd. threshold values were estimated based on the maximum tss criteria. the bimodality column displays the predictors that demonstrate a concave response curve. the name of variables was shortened as follows b = bio. * best models based on the selection criteria. arias-giraldo et al. – enmpa 35 figure 1. bioclimatic variable response and importance across models. the upper half illustrates the response curves of the species to variables bio-1 (annual mean temperature ºc) and bio-12 (annual precipitation mm) across the two best models (id 29 and id 31), and the vertical dashed lines mark environmental limits of the calibration data. the lower half of the figure shows a boxplot detailing the distribution of variable importance values among selected models and highlighting the variability and median importance of predictors. the importance of variables, measured by explained deviance, incorporates linear, quadratic terms (denoted as i(var^2)) and two-way interaction. the y-axis represents the importance values, with boxplots delineating the interquartile range and median value. the frequency of predictor inclusion across models is also noted, providing insight into their relevance for models. arias-giraldo et al. – enmpa 36 around 20ºc, whereas for bio-12, the maximum probability was observed for values below 500 mm. extrapolation of these curves to values outside calibration ranges showed decreasing probabilities, indicating safe model extrapolations rather than perpetual increments of probability towards extreme environmental conditions. variable importance analysis highlighted the linear term of bio-1 and the quadratic terms of both variables as the most relevant predictors (fig. 1). the interaction between variables, along with the linear component of bio-12, were found to be less relevant. despite their minor contribution, models incorporating these predictors were selected based on their superior goodness of fit compared to models excluding them. the geographic predictions of the two selected models showed consistent patterns, with higher probability values in the southwestern parts of the united states (fig. 2). however, some discrepancies were noted in florida, where one model estimates higher probability values than the other (fig. 3). the effects of distinct types of extrapolation were minor and mainly noticeable in florida (figs. 2, 3, fig. a3 of appendix). independent data validation confirmed the comparable performance of all projected models, although with slightly different threshold values estimated for each model (tables 3 and table a1 of appendix). discussion understanding how species are distributed in different environments and predicting how species will respond to changes is crucial in research in ecology. ecological niche modeling helps in this task, and this contribution introduces enmpa, an r package that facilitates calibration of ecological niche models via glm. the package integrates a set of complex methodological developments in the enm field, using logistic glms to estimate the probability of a species figure 2. geographic depiction of estimated probability of occurrence for the pathogen virtual species deriving from the two final selected models allowing free extrapolation. maps are shown at a spatial resolution of 10’ (~20 km at the equator). arias-giraldo et al. – enmpa 37 figure 3. consensus geographic projection of probability of occurrence for the pathogen virtual species allowing free extrapolation. the figure displays the probability of occurrence derived from the two final selected models using averaging metrics: mean, median, and weighted average based on akaike weights. in this example, since the averaging is calculated from two models, the median coincides with the mean. the figure displays the variance among the three consensus projections. the maps are shown at a spatial resolution of 10‘ (~20 km at the equator). model threshold criteria threshold roc-auc false positive rate accuracy sensitivity specificity tss id 29 ess 0.140 0.957 0.090 0.910 0.909 0.910 0.819 maxtss 0.117 0.957 0.101 0.910 1.000 0.899 0.899 sen90 0.131 0.957 0.101 0.900 0.909 0.899 0.808 id 31 ess 0.158 0.957 0.090 0.910 0.909 0.910 0.819 maxtss 0.135 0.957 0.101 0.910 1.000 0.899 0.899 sen90 0.144 0.957 0.101 0.900 0.909 0.899 0.808 consensus (weighted average) ess 0.150 0.957 0.090 0.910 0.909 0.910 0.819 maxtss 0.127 0.957 0.101 0.910 1.000 0.899 0.899 sen90 0.139 0.957 0.101 0.900 0.909 0.899 0.810 table 3. evaluation of the two selected models and the consensus using an independent data set with presence and absence records. the classification capacities of the final prediction were calculated using the confusion matrix based on three threshold criteria. occurring in a particular environment (austin 2002, 2007; ward et al. 2009). the tools developed in this package are particularly interesting for studies that involve data derived from detection/non-detection sampling protocols, such as pathogen test results, detections of species on controlled-protocol surveys, etc. absence of a species is difficult to demonstrate because there are multiple reasons and circumstances that can lead not to detect a species (mackenzie 2005; feng and papeş 2017). however, when the sampling protocol is controlled (e.g., sampling effort and methods are comparable) non-detections are a valuable source of information in characterizing environments that are favorable or not for the occurrence of a species (cobos and peterson 2022). these non-detections are precisely what enmpa can use as absence records, which together with presences can help to develop more robust models. the package enmpa runs most analyses using data prepared by the users as a “data.frame” that contains presence-absence records associated with values of environmental variables. the steps of model fitting, training, and testing do not require variables as raster layers, which are only needed if mode geoarias-giraldo et al. – enmpa 38 graphic predictions are to be produced. therefore, raster layer resolution has few implications in the functionality of enmpa, which makes it computationally efficient. one of the most notable features of enmpa is that it allows users to explore a wide range of predictor combinations in a glm framework to find the set of combinations that better fit the data and explain the phenomenon of study. focusing on exploring different predictor features, including linear, quadratic, and two-way interaction responses, enabling a detailed analysis of the entire parameter space (cobos et al. 2019b). apart from the functionalities corresponding to the main steps in enm, enmpa implements two novel methods that allow users to select variables based on niche signal detection (cobos and peterson 2022) and filter those models with response shapes that do not align with ecological theory (austin 2002, 2007; peterson et al. 2011; merow et al. 2014). the first method helps users to discard potentially non-relevant variables before modeling, which prevents over-parameterization of models and makes the calibration step easier by avoiding exploring irrelevant predictor combinations that do not form part of the species’ niche (cobos and peterson 2022). although forward, backward, or stepwise selection processes have been commonly used methods for selecting predictors in modeling (efroymson 1960), they have been criticized for misapplication of a single-step statistical test in a multi-step procedure (harrell 2001; flom and cassell 2007; smith 2018). this problem may lead to the selection of nuisance variables or models that perform worse with independent data than in calibration (smith 2018). the second method seeks to stay in line with the standard of niche theory, in which species’ fitness responds to environmental conditions with a unimodal response. extreme environmental values lead to low fitness, while intermediate environmental values are optimal for the species (jiménez-valverde et al. 2011; escobar 2020). however, the fitting of quadratic terms can be limited if insufficient sampling information is available and one extreme of the curve is not captured, which can lead to estimation of odd response shapes. although bimodal responses of species to variables may have interesting and meaningful interpretations (e.g., indicating that a niche is incompletely represented by the data), they need to be evaluated in detail for particular species. therefore, excluding quadratic predictors with bimodal behavior is crucial in most cases to avoid misleading conclusions. the example of a virtual pathogen species demonstrates the usefulness and effectiveness of enmpa. the meticulous steps involved in model calibration, selection, and evaluation, in combination with the consideration of response curves and variable contributions, collectively contributed to a refined understanding of the ecological niche of the virtual species. for example, the results suggest that the virtual pathogen would thrive in warmer climates with lower rainfall. the example also highlights the practicality and accuracy of enmpa in modeling species’ niches, when records derived from sampling protocols of detection and non-detection are available. acknowledgments we thank the kuenm working group for valuable discussions. lfag thanks the biodiversity institute, university of kansas, for hosting him during the development of this project. lfag thanks juan a. navas-cortés and blanca b. landa for their support during the development of this work. competing interests the authors have declared that no competing interests exist. funding this research was funded partially by an international mobility grant for phd candidates awarded to lfag by the university of córdoba, spain. lfag was also supported by the consejo superior de investigaciones científicas intramural (project 202340e021). the u.s. national science foundation also supported the work via grant oia1920946. literature cited akaike, h. 1998. information theory and an extension of the maximum likelihood principle. pp. 199–213 in e. parzen, k. tanabe, and g. kitagawa, eds. selected papers of hirotugu akaike. springer, new york, ny. allouche, o., a. tsoar, and r. kadmon. 2006. assessing the accuracy of species distribution models: prevalence, kappa and the true skill statistic (tss). j. appl. ecol. 43:1223– 1232. anderson, m. j. 2017. permutational multivariate analysis of variance (permanova). pp. 1–15 in wiley statsref: statistics reference online. john wiley & sons, ltd. araújo, m. b., and a. guisan. 2006. five (or so) challenges for species distribution modelling. j. biogeogr. 33:1677–1688. arias-giraldo et al. – enmpa 39 austin, m. 2007. species distribution models and ecological theory: a critical assessment and some possible new approaches. ecol. model. 200:1–19. austin, m. p. 2002. spatial prediction of species distribution: an interface between ecological theory and statistical modelling. ecol. model. 157:101–118. austin, m. p., a. o. nicholls, and c. r. margules. 1990. measurement of the realized qualitative niche: environmental niches of five eucalyptus species. ecol. monogr. 60:161–177. blowes, s. a., s. r. supp, l. h. antão, a. bates, h. bruelheide, j. m. chase, f. moyes, a. magurran, b. mcgill, i. h. myerssmith, m. winter, a. d. bjorkman, d. e. bowler, j. e. k. byrnes, a. gonzalez, j. hines, f. isbell, h. p. jones, l. m. navarro, p. l. thompson, m. vellend, c. waldock, and m. dornelas. 2019. the geography of biodiversity change in marine and terrestrial assemblages. science 366:339–345. american association for the advancement of science. bolker, b. m., m. e. brooks, c. j. clark, s. w. geange, j. r. poulsen, m. h. h. stevens, and j.-s. s. white. 2009. generalized linear mixed models: a practical guide for ecology and evolution. trends ecol. evol. 24:127–135. burnham, k. p., and d. r. anderson. 1998. practical use of the information-theoretic approach. pp. 75–117 in k. p. burnham and d. r. anderson, eds. model selection and inference: a practical information-theoretic approach. springer, new york, ny. clouvel, l., b. iooss, v. chabridon, m. il idrissi, and f. robin. 2023. a review on variance-based importance measures in the linear regression context. hal-04102053v2 cobos, m. e., and a. t. peterson. 2022. detecting signals of species’ ecological niches in results of studies with defined sampling protocols: example application to pathogen niches. biodivers. inform. 17:50–58. cobos, m. e., a. t. peterson, n. barve, and l. osorio-olvera. 2019a. kuenm: an r package for detailed development of ecological niche models using maxent. peerj 7:e6281. peerj inc. cobos, m. e., a. t. peterson, l. osorio-olvera, and d. jiménezgarcía. 2019b. an exhaustive analysis of heuristic methods for variable selection in ecological niche modeling and species distribution modeling. ecol. inform. 53:100983. cordier, j. m., r. loyola, o. rojas-soto, and j. nori. 2020. modeling invasive species risk from established populations: insights for management and conservation. perspect. ecol. conserv. 18:132–138. efroymson, m. a. 1960. multiple regression analysis. math. methods digit. comput. 191–203. john wiley & sons. elith, j., s. ferrier, f. huettmann, and j. leathwick. 2005. the evaluation strip: a new and robust method for plotting predicted responses from species distribution models. ecol. model. 186:280–289. elith, j., c. h. graham*, r. p. anderson, m. dudík, s. ferrier, a. guisan, r. j. hijmans, f. huettmann, j. r. leathwick, a. lehmann, j. li, l. g. lohmann, b. a. loiselle, g. manion, c. moritz, m. nakamura, y. nakazawa, j. mcc. m. overton, a. townsend peterson, s. j. phillips, k. richardson, r. scachetti-pereira, r. e. schapire, j. soberón, s. williams, m. s. wisz, and n. e. zimmermann. 2006. novel methods improve prediction of species’ distributions from occurrence data. ecography 29:129–151. escobar, l. e. 2020. ecological niche modeling: an introduction for veterinarians and epidemiologists. front. vet. sci. 7. ferrier, s., m. drielsma, g. manion, and g. watson. 2002. extended statistical approaches to modelling spatial pattern in biodiversity in northeast new south wales. ii. community-level modelling. biodivers. conserv. 11:2309– 2338. feng, x., and papeş, m. 2017. can incomplete knowledge of species’ physiology facilitate ecological niche modelling? a case study with virtual species. diversity and distributions, 23: 1157-1168. fick, s. e., and r. j. hijmans. 2017. worldclim 2: new 1-km spatial resolution climate surfaces for global land areas. int. j. climatol. 37:4302–4315. fielding, a. h., and j. f. bell. 1997. a review of methods for the assessment of prediction errors in conservation presence/ absence models. environ. conserv. 24:38–49. flom, p. l., and d. l. cassell. 2007. stopping stepwise: why stepwise and similar selection methods are bad, and what you should use. p. in northeast sas users group inc 20th annual conference. franklin, j. 2010. moving beyond static species distribution models in support of conservation biogeography. divers. distrib. 16:321–330. franklin, j. 2013. species distribution models in conservation biogeography: developments and challenges. divers. distrib. 19:1217–1223. ghanbarian, g., m. r. raoufat, h. r. pourghasemi, and r. safaeian. 2019. 9 habitat suitability mapping of artemisia aucheri boiss based on the glm model in r. pp. 213–227 in h. r. pourghasemi and c. gokceoglu, eds. spatial modeling in gis and r for earth and environmental sciences. elsevier. guisan, a., t. c. edwards, and t. hastie. 2002. generalized linear and generalized additive models in studies of species distributions: setting the scene. ecol. model. 157:89–100. guisan, a., and n. e. zimmermann. 2000. predictive habitat distribution models in ecology. ecol. model. 135:147–186. hannah, l., p. r. roehrdanz, p. a. marquet, b. j. enquist, g. midgley, w. foden, j. c. lovett, r. t. corlett, d. corcoran, s. h. m. butchart, b. boyle, x. feng, b. maitner, j. fajardo, b. j. mcgill, c. merow, n. morueta-holme, e. a. newman, d. s. park, n. raes, and j.-c. svenning. 2020. 30% land conservation and climate action reduces tropical extinction risk by more than 50%. ecography 43:943–953. hao, t., j. elith, j. j. lahoz-monfort, and g. guillera-arroita. 2020. testing whether ensemble modelling is advantageous for maximising predictive performance of species distribution models. ecography 43:549–558. arias-giraldo et al. – enmpa 40 harrell, f. e. 2001. regression modeling strategies: with applications to linear models, logistic regression, and survival analysis. springer. hastie, t., r. tibshirani, and j. friedman. 2009. model assessment and selection. pp. 219–259 in the elements of statistical learning: data mining, inference, and prediction. springer new york, new york, ny. jiménez-valverde, a., a. t. peterson, j. soberón, j. m. overton, p. aragón, and j. m. lobo. 2011. use of niche models in invasive species risk assessments. biol. invasions 13:2785– 2797. liu, c., m. white, and g. newell. 2011. measuring and comparing the accuracy of species distribution models with presence-absence data. ecography 34:232–243. lobo, j. m., a. jiménez-valverde, and r. real. 2008. auc: a misleading measure of the performance of predictive distribution models. glob. ecol. biogeogr. 17:145–151. mackenzie, d. i. 2005. was it there? dealing with imperfect detection for species presence/absence data. australian & new zealand journal of statistics. 47(1), 65-74. manel, s., h. c. williams, and s. j. ormerod. 2001. evaluating presence-absence models in ecology: the need to account for prevalence. j. appl. ecol. 38:921–931. merow, c., m. j. smith, t. c. edwards jr, a. guisan, s. m. mcmahon, s. normand, w. thuiller, r. o. wüest, n. e. zimmermann, and j. elith. 2014. what do we gain from simplicity versus complexity in species distribution models? ecography 37:1267–1281. murray, k., and m. m. conner. 2009. methods to quantify variable importance: implications for the analysis of noisy ecological data. ecology 90:348–355. muscarella, r., p. j. galante, m. soley-guardia, r. a. boria, j. m. kass, m. uriarte, and r. p. anderson. 2014. enmeval: an r package for conducting spatially independent evaluations and estimating optimal model complexity for maxent ecological niche models. methods ecol. evol. 5:1198–1205. oksanen, j., and p. r. minchin. 2002. continuum theory revisited: what shape are species responses along ecological gradients? ecol. model. 157:119–129. park, d. s., and d. potter. 2015. why close relatives make bad neighbours: phylogenetic conservatism in niche preferences and dispersal disproves darwin’s naturalization hypothesis in the thistle tribe. mol. ecol. 24:3181–3193. peterson, a. t. 2014. mapping disease transmission risk. johns hopkins university press. peterson, a. t., j. soberón, r. g. pearson, r. p. anderson, e. martínez-meyer, m. nakamura, and m. b. araújo. 2011. ecological niches and geographic distributions (mpb-49). princeton university press. phillips, s. j., r. p. anderson, m. dudík, r. e. schapire, and m. e. blair. 2017. opening the black box: an open-source release of maxent. ecography 40:887–893. phillips, s. j., r. p. anderson, and r. e. schapire. 2006. maximum entropy modeling of species geographic distributions. ecol. model. 190:231–259. qiao, h., j. soberón, and a. t. peterson. 2015. no silver bullets in correlative ecological niche modelling: insights from testing among many potential algorithms for niche estimation. methods ecol. evol. 6:1126–1136. r core team. 2022. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. radosavljevic, a., and r. p. anderson. 2014. making better maxent models of species distributions: complexity, overfitting and evaluation. j. biogeogr. 41:629–643. rupprecht, f., j. oldeland, and m. finckh. 2011. modelling potential distribution of the threatened tree species juniperus oxycedrus: how to evaluate the predictions of different modelling approaches? j. veg. sci. 22:647–659. santika, t., and m. f. hutchinson. 2009. the effect of species response form on species distribution model prediction and inference. ecol. model. 220:2365–2379. searcy, c. a., and h. b. shaffer. 2016. do ecological niche models accurately identify climatic determinants of species ranges? am. nat. 187:423–435. the university of chicago press. smith, g. 2018. step away from stepwise. j. big data 5:32. steele, k., and c. werndl. 2013. climate models, calibration, and confirmation. br. j. philos. sci. 64:609–635. the university of chicago press. thuiller, w. 2003. biomod – optimizing predictions of species distributions and projecting potential future shifts under global change. glob. change biol. 9:1353–1362. wagenmakers, e.-j., and s. farrell. 2004. aic model selection using akaike weights. psychon. bull. rev. 11:192–196. ward, g., t. hastie, s. barry, j. elith, and j. r. leathwick. 2009. presence-only data and the em algorithm. biometrics 65:554–563. warren, d. l., and s. n. seifert. 2011. ecological niche modeling in maxent: the importance of model complexity and the performance of model selection criteria. ecol. appl. 21:335–342. warren, d. l., a. n. wright, s. n. seifert, and h. b. shaffer. 2014. incorporating model complexity and spatial sampling bias into ecological niche models of climate change risks faced by 90 california vertebrate species of concern. divers. distrib. 20:334–343. arias-giraldo et al. – enmpa 41 appendix: supplementary figures and tables figure a1. results from niche comparison using permanova analysis. the figure on the left represents positive and negative records in the environmental space for bio 1 and bio 2. on the right, ellipsoids are derived from the data to explore and visualize the position and spread of host and pathogen niches. the p-value represents the statistical significance of the permanova test. figure a2: visualization of the results obtained from the univariate non-parametric test for detecting signals of the virtual pathogen’s niche. the top-left and top-right panels depict the mean distribution of the niche position of the species about the null distribution for the bio 1 and bio 12 variables. bottom-left and bottom-right panels depict the range of environmental conditions of the niche about the null distribution, as represented by the standard deviation (sd). the vertical dotted blue lines signify the observed value associated with presences of the pathogen. the vertical dotted gray lines represent the lower and upper 95% confidence limits of the null distribution. the barplot histogram represents the null distribution. arias-giraldo et al. – enmpa 42 figure a3: geographic projections of the probability of occurrence for the virtual pathogen species, based on two final selected models and their weighted average consensus. the projections are displayed under two conditions: no extrapolation (ne) and extrapolation with clamping (ec). each row represents different models (model id 29 and model id 31) and their weighted averages, with the left column showing ne conditions and the right column showing ec conditions. maps are presented at a spatial resolution of 10' (~20 km at the equator). models threshold criteria threshold omission error mean auc ratio at 5% p-value proc id 29 ess 0.137 0.091 1.648 <0.0001 maxtss 0.117 0.000 1.648 <0.0001 sen90 0.131 0.091 1.648 <0.0001 id 31 ess 0.158 0.091 1.631 <0.0001 maxtss 0.135 0.000 1.630 <0.0001 sen90 0.144 0.091 1.636 <0.0001 consensus (weighted average) ess 0.150 0.091 1.635 <0.0001 maxtss 0.127 0.000 1.640 <0.0001 sen90 0.139 0.091 1.640 <0.0001 table a1. evaluation of the two selected models and the consensus using an independent data set using presences-only data. biodiversity informatics, 17, 2022, pp. 96-107 96 biodiversity informatics for public policy: the case of conabio in mexico jorge soberón1 1biodiversity institute and department of ecology & evolutionary biology, university of kansas, 1345 jayhawk blvd., lawrence, kansas 66045, usa (orcid https://orcid.org/0000-0003-2160-4148) abstract. in this paper, i present and review the development of the biodiversity information system that was developed in mexico. i describe briefly the organization that made the system possible and some of its history. then, i focus on the principles of design of the information system, and a few of its major uses. i provide data on costs and usage, and end with some reflections on the fragility of such institutional systems. key words: biodiversity information systems; mexico biodiversity is the aggregate of ways in which life manifests itself in the planet (brooks et al. 2006). biodiversity is a complex concept that can be defined from multiple perspectives (maclaurin and sterelny 2008; sarkar 2002). a comprehensive perspective is to regard biodiversity as an aggregate of elements, how are they structured, and how they function, at scales from the sub-individual to the planetary (noss 1990). for instance, at certain scales, the elements of biodiversity are individuals of species, the structure is their spatiotemporal locations, and the functioning is their interactions. at a different scale, elements of biodiversity may be biomes, structure would be their spatial extents, and functioning would be the biogeochemical processes taking place in them. from this comprehensive perspective, conservation of biodiversity requires actions and policies at multiple scales. historically, however, biodiversity has been managed mostly at relatively local scales (i.e., at the scale of activities of human groups of small size), by indigenous peoples, farmers, fishermen and such local actors (gadgil et al. 1993). this “management” has taken place for thousands of years, such that, overall, indigenous and traditional cultures generally have deep knowledge of their environments and respectful attitudes towards nature (toledo 2001). this proximity leads to a mostly sustainable management of components of biodiversity (gadgil et al. 1993), since many of the impacts were spatially concentrated, and were reversible in nature. moreover, traditionally, natural resources were often the subject of strict governance (ostrom et al. 1999), as opposed to the naïve view of traditionally managed resources as open-access “commons” (hardin 1968). traditional governance is, in the end, highly conducive to sustainable use (gadgil et al. 1993). in modern times (i.e., over the last ~400 years), however, the rate at which human activities have impacted biodiversity has accelerated (butchart et al. 2010; ehrlich 1995; mcneely et al. 1990; steffen 2015). actors beyond the local now exert substantial impacts on different components of biodiversity, sometimes in ways that are spatially very extended or have long-term effects, and that sometimes are irreversible. governance of common-pool resources of global extent is challenging (ostrom et al. 1999). managing and conserving biodiversity, therefore, requires participation of stakeholders at many different levels, which creates problems of obtaining and assembling the required information. at first, emphasis was placed on spatially structured information, in effect “putting biodiversity on the map” (bibby 1992; edwards et al. 2002; reid 1998; scott 1993). in practice, however, this emphasis meant putting the biodiversity of developed countries on the map, and biodiversity loss was not abated elsewhere (peterson and soberón 2018). still, some voices have insisted that, without biodiversity data, management would be difficult or impossible (balmford et al. 2005). indeed, when viewed from a multilevel perspective, management of the multitude of entities and processes comprising biodiversity is impossible without an overarching perspective. in the context of widespread loss of the https://orcid.org/0000-0003-2160-4148 jorge soberón – biodiversity informatics for public policy 97 components, structure, and functioning of biodiversity, the countries of the world negotiated a “convention on biological diversity” (koester 2002), which stressed a dire need for globally relevant biodiversity data to be made available openly to the broadest community (laihonen 2004). what are “biodiversity data” then, how can biodiversity data be managed in accessible ways, and what can they be used for? the core of this paper is an attempt to answer these questions, from a mainly historical perspective, using the case of the mexican national biodiversity agency (the comisión nacional para el uso y conocimiento de la biodiversity, or conabio) as an example. from its creation in 1992 until 2005, i served as the executive secretary of conabio. it is from this perspective that i write this paper. biodiversity data as stated above, “biodiversity” is a complex concept, being both multi-scale and multi-perspective. numerous perspectives on biodiversity have been documented in countless books, papers, images, recordings, and databases regarding protein structure, genetic sequences, species diversity, community ecology, etc. however, in practice, the key, focal concept has been that of records of occurrence of a species, otherwise known as primary biodiversity data (peterson et al. 2010; soberón and peterson 2004; sousa‐baena et al. 2014). the key idea of primary biodiversity data is that each record comprises a date, a description of a locality, and a taxonomic identity (johnson 2007; soberón and peterson 2004). the locality data allow linking to geographic information, and the taxonomic identity provides an index to genetic, demographic, systematic, or cultural data. the importance of the taxonomic identity in linking databases cannot be overemphasized. solving all the “knowledge shortfalls” described for biodiversity (hortal et al. 2015) is predicated on having a consistent and stable system of names, which in biology is based on linnean taxonomic schemes. the names constitute a “hinge feature” of primary biodiversity data, in fact linking geography with a multiplicity of perspectives, via the name, which is of fundamental importance (chapman 1991; peterson et al. 2010). in what follows, i will be focusing on this core of primary biodiversity data, mainly because in practice it has been the focus of most large-scale biodiversity informatics initiatives (coetzer 2012; conabio 2012; sandlund 1991). the beginnings of conabio in june of 1992, the united nations organized the conference on environment and development (also known as the “earth summit”). this took place in rio de janeiro, brazil. in preparation for this, the then-president of mexico asked the chancellor of the national university, josé sarukhán, the foremost ecologist of mexico, to provide some possible initiatives to present in rio. in february 1992, a meeting was organized in mexico (sarukhán and dirzo 1992) to begin designing a national initiative on biodiversity for the country. as a consequence of this meeting of international experts (mostly in biodiversity conservation), two of the most prominent ecologists of mexico (daniel piñero and rodolfo dirzo, both researchers in the institute of ecology of the national university) worked with sarukhán to propose to the president of mexico to create a highlevel government agency in charge of biodiversity. in 1992, an inter-ministerial commission was created (conabio), composed of ten cabinet-level ministers, and presided ex-officio by the president of mexico. i was appointed executive secretary of conabio, a role in which i served for 13 years. although conabio is formally a multi-ministry federal government agency, it operates via an executive secretariat that was allowed to establish a private trust fund via which to operate. this hybrid structure, combining private and public aspects, gave conabio not only the capacity to address challenging technical tasks, but also to act as a trusted and necessary government interlocutor. conabio was given a number of tasks. the most important was: “to synthesize information relative to the biological resources of the country, in a database that should be kept permanently updated.” this activity was the initial and major focus of conabio: to this end, the first step was to take stock of similar initiatives elsewhere in the world. the conabio team obtained information by visiting four existing organizations around the world. first, we consulted one in india, now extinct, that had worked entirely based on secondary information (bibliography). although the system was open to the public, it was entirely based on secondary data, making that consultation a dead end. a map was a page in a publication (as opposed to a machine-readable jorge soberón – biodiversity informatics for public policy 98 geospatial dataset), and a list of occurrence localities was an image of some text. this system was essentially a bibliographic consult system, and was not a useful lesson for mexico. the second system that we studied was that of the heritage methodology of the nature conservancy (groves 1995). this system was based on primary data, obtained from public museums in the united states and canada, among other sources. the fact that it used primary data meant that a variety of operations could be performed (e.g., performing statistical analyses or visualization of patterns in maps and graphs) on the data (stein et al. 2000), but the data were not available to the public. however, it was regularly used in for-profit consultations, leading to widespread resentment among museum officials, who had provided the data for free, without imagining a for-profit use. therefore, eventually, many sources of data closed to this system, and it clearly was not a model that we wanted to follow in mexico. this unfortunate situation has seldom been discussed in the literature, but anecdotally it is well known in the community. the experience led to another principle in conabio: if the data were to be used for a for-profit purpose, the user would need to consult with the original sources. a third system was that of costa rica’s instituto nacional de biodiversidad (inbio). this database, which was still in a design phase when we visited, was based on primary biodiversity data, mostly obtained from de novo collections performed and maintained by inbio (tangley 1990). at the time of our visit, the system was still in incipient stages. also, although the system was based on primary data, it had a rather narrow focus on bioprospecting for pharmaceutical products (sittenfeld and r.villers 1993). finally, in 1992, personnel of conabio, as well as an international group including kenyan, indonesian, costa rican, and u.s. american scientists (chapman 2001), visited the environmental resources information network (erin), in australia (kaye et al. 1997). erin has since disappeared, although many of its capabilities were replaced by the atlas of living australia (belbin 2021). in the 1990s, australia was without a doubt the most advanced country in the world in biodiversity informatics. their system was based on primary biodiversity data, provided in largest part by the network of australian herbaria and museums. they had sophisticated bioinformatics capacities for taxonomic descriptions (dallwitz 1993), species distribution modeling (booth 2018; busby et al. 1991; nix 1986), prioritizing sites for conservation (pressey et al. 1993), and more generally for organization, visualization and analysis of large databases of primary biodiversity data. the erin system was open to the public (even at a time when html was not yet operational), which in practice was principally academic users, though the users were many, and the types of applications were varied (e.g., designing conservation plans, and surveying poorly explored localities). the australian experience, compared with the others, suggested great potential for a biodiversity information system based on two key principles: primary biodiversity data. that is, the data should be as little interpreted as possible. essentially a name, a date, and a locality associated with a physical specimen. combining the data, interpreting, visualizing, and analyzing the data is the responsibility of the users (soberón and peterson 2004). data publicly available. data should be completely and openly accessible to everyone. when conabio was launched, in 1992, the world wide web was just being developed (berners-lee 1992), but computer scientists at conabio were already aware of it, and appreciated its potential to allow efficient public access to what was going to be a large amount of data. neither of these two fundamental points had been obvious at that time. regarding the utility of primary data, there were many expressions of doubt. most advisors to conabio were used to reading books and papers, not to performing their own analysis using large databases (recall that large-scale, publicly available databases of primary data were basically non-existent at this point in time). nevertheless, conabio opted for primary data, following the experience of australia, and what would eventually become the case in costa rica. on public access, at the time at which conabio was starting, attention to the problem of so-called biopiracy (reid 1996; ten kate 1999) was at its most intense. the authorities of conabio were under considerable pressure not to allow public access to the information, in case commercial agents might misuse it. moreover, some museum curators opposed releasing collections-associated data (graves 2000) for other reasons. for instance, it was argued that scientists working with vertebrates might be targeted for animal-rights concerns, or that the data were of monetary value. several prominent mexican biolojorge soberón – biodiversity informatics for public policy 99 gists were similarly adamant in their refusal to share data. after almost a year of such intense discussions, conabio convened a meeting of mexican museum curators and directors, to discuss the issue of public access to data via the internet. in november 1993, in oaxaca, mexico, a declaration was issued1 stating that the mexican museums and herbaria were committed to computerizing and distributing biodiversity data. as such, an important political battle had been won. however, at that time, in mexico (and indeed worldwide), very few biological collections had been digitized. what is more, no effective implementations existed for sharing data on the internet. finally, despite having signed the oaxaca declaration, many curators still had serious misgivings (expressed in private) about public access to biodiversity data! nevertheless, the signed commitment by mexican scientists gave conabio the legitimacy to start computerizing collections and developing technologies by which to share the data. the sistema nacional de información de la biodiversidad (snib) building a robust and stable computer system capable of dealing with the millions of data elements about biodiversity took conabio more than 10 years (sarukhán et al. 2014; soberón and koleff 2000). the cost was substantial, since most of the data were not yet digitized, and that process required resources to pay experts to travel to collections, acquire computers, and curate data. the cost of digitizing specimens (figure 1) was on the order of millions of dollars, paid by the mexican federal government. the figure shows the cost and yield (i.e., number of biodiversity records) for each of 221 projects supported by conabio between 1993 and 2000 that digitized or produced records for the main database. the other major element making up the snib was remote sensing, mostly oriented toward monitoring at the ecosystem level. to this end, the mexican government purchased a satellite dish and associated hardware and software, capable of downloading images from the moderate resolution imaging spectroradiometer (modis) in the terra and aqua satellites in real time. this purchase was an investment on the order of many hundreds of thousands of dollars (paid by the mexican government), and required hiring foreign experts familiar with remote-sense technology. the foreign experts were paid mostly by the german 1 http://www.conabio.gob.mx/remib/doctos/declaracion.html. gtz cooperation agency, with a symbolic contribution from conabio. the german experts came to work at conabio under the “shared experts” scheme of the gtz, which guaranteed several years of work in the host country. this long-term participation was key to the success of the project. although acquiring the remote-sensing infrastructure was costly, delays inherent in acquiring the same images from commercial or noncommercial foreign sources made the purchase necessary, mainly for initiatives to monitor disasters such as wildfires. by 2005, conabio had spent about us$10m of taxpayer’s money in acquiring data and remote-sensing hardware, and the computers and system engineers required to run the system. the cost of acquiring primary biodiversity data remained constant per project, on average, at us$5,500 per project. but the cost per specimen is inversely related to the size of the collection (figure 1), which means that is more efficient to computerize large collections. on the other hand, the experience in mexico was often that the larger institutional collections tended to be less willing to participate in these initiatives. figure 1. cost (in contemporary us dollars) of digitizing biodiversity collections, as a function of the number of specimens digitized (note that data are on a log-log scale). each point is a digitization project. the data for this figure were sourced from internal conabio reports. the “rugs” along each axis show the distribution of points. http://www.conabio.gob.mx/remib/doctos/declaracion.html jorge soberón – biodiversity informatics for public policy 100 by 2005, conabio had accumulated a substantial storehouse of data comprising primary biodiversity records, satellite images, photographs, maps, and textual data (table 1). the snib is the computer system that organizes all of these information resources, to assure both efficient access and open sharing.2 an outline of the technical details of the system has been published elsewhere (sarukhán and jiménez 2016), but stressing that the system is based on the two principles stated above: the backbone is primary biodiversity data, and all data are openly available. the sheer amount of data means that the expenditures involved are substantial, in terms of hardware and human resources. more precisely, the mexican taxpayer, and some foreign agencies (the german gtz, specifically) invested more than us$10m in the system. for comparison, the convention on biological diversity spent $12,300 per country on its “biodiversity clearing house mechanism” (reed 2017). the resources spent by conabio included not only expenses involved in capturing, organizing, and analyzing data, but also in design and implementation of the computer system to manage it (soberón et al. 2010). the need to keep the data updated means that hundreds of mexican (and some foreign) scientists’ participation was crucial to the success of the system. maintaining such participation requires money, time, and effort. despite the fact that much was developed inhouse, snib is compliant with important international efforts. specifically, the data architecture follows the “darwin core” (wieczorek et al. 2012). data quality control was influenced by the work of chapman (2005) and wieczorek et al. (2004); and 2 https://www.gob.mx/cms/uploads/attachment/file/548546/informeconabio-2017-2019.pdf. the primary data can be accessed via the global biodiversity information facility (lane and edwards 2007). digitizing data on the labels of millions of specimens was accomplished mostly by hand, often (mostly in herbaria) by taking photographs of the specimen sheets and capturing the data in mexico. digitizing specimens is now a major activity all over the world (asase et al. 2020; canhos 2017; nelson and ellis 2019; siebert and smith 2004), one that is increasingly technological (beaman and cellinese 2012; tegelberg et al. 2014). the snib is more than just a data repository, complex as this task is. there are serious analytical capacities developed in the area of biodiversity informatics. among the principal skills are those related to visualizing data (stephens et al. 2017), predicting species’ geographic distributions (conabio 2012), assembling complex remote-sensing products (gonzalez et al. 2014; hruby et al. 2016), monitoring wildfires (ressl et al. 2009) and others. biodiversity informatics, in a wide sense, is now a major activity in conabio, with engineers, mathematicians, taxonomists, and remote-sensing experts collaborating in the activities. usage of snib the primary data that conabio has assembled have been used regularly for many government purposes. this is also the case in other parts of the world (guisan et al. 2013), but the mexican examples are very illustrative. before discussing some examples of use of data for policy, it is interesting to mention that much of the data are used without conabio knowing the purpose. that is, the primary data of conabio are accessed very frequently. indeed, conabio’s website is accessed many thousands of times per week data type number link primary data records 14,000,000 https://www.snib.mx/ejemplares/descarga/ images 155,000 http://www.conabio.gob.mx/otros/cgi-bin/herbario.cgi taxonomy controlled vocabularies 103,000 https://www.snib.mx/taxonomia/descarga/ remote sensing images 582,000 http://www.conabio.gob.mx/informacion/gis/ digital maps 14,000 http://www.conabio.gob.mx/informacion/gis/ technical data about species 4000 https://www.gob.mx/conafor/documentos/fichas-tecnicas-especies-exoticas-invasoras; https://enciclovida.mx/ table 1. main informational elements in the sistema nacional de información de la biodiversidad of mexico (snib, based on the 2017-2019 conabio activities reports2) https://www.gob.mx/cms/uploads/attachment/file/548546/informe-conabio-2017-2019.pdf https://www.gob.mx/cms/uploads/attachment/file/548546/informe-conabio-2017-2019.pdf https://www.snib.mx/ejemplares/descarga/ http://www.conabio.gob.mx/otros/cgi-bin/herbario.cgi https://www.snib.mx/taxonomia/descarga/ http://www.conabio.gob.mx/informacion/gis/ http://www.conabio.gob.mx/informacion/gis/ https://www.gob.mx/conafor/documentos/fichas-tecnicas-especies-exoticas-invasoras https://www.gob.mx/conafor/documentos/fichas-tecnicas-especies-exoticas-invasoras jorge soberón – biodiversity informatics for public policy 101 (figure 2), with data being downloaded at the level of gigabytes (internal communication), although the organization is not aware of the purpose of the use of data downloads. one concern at the beginning of conabio was that most users of open biodiversity data would be foreign “biopirates” (ten kate 1999). in table 2, i show the data on access, over the last four years, by country domain. it shows that (by a factor of ~100fold), most users are mexicans, not foreigners. anecdotally, it is known that most users of conabio data are researchers, ngos, or mexican government agencies. planting permits for gmos in mexican legislation, planting genetically modified organisms (gmos) is forbidden if there is a risk of introgressions of modified sequences into wild relatives. conabio implemented a system of predicting the risk which is based on ecological niche modeling (a computational method used to predict areas of distribution) applied to wild relatives country users sessions average time (s) mexico 173,613 409,283 104 united states 2,501 3,908 67 colombia 1,099 1,421 53 peru 900 1,171 60 spain 644 918 70 ecuador 588 724 45 argentina 424 586 68 canada 298 555 141 guatemala 294 417 81 total (4 years) 184,148 424,825 103 figure 2. number of unique users of conabio website who had at least one session within 7-day time periods between april 2018 and august 2022. table 2. statistics on visits to conabio’s website over the last four years, with data sourced from google analytics in august 2022. note that most users of conabio databases are in mexico. jorge soberón – biodiversity informatics for public policy 102 of candidate species (soberón et al. 2002). this system has proved to have predictive ability (wegier et al. 2011), it is transparent and empirical (i.e., based on data), and was adopted by the ministries of the environment and of agriculture of mexico. the system is complicated, in the sense that it uses a variety of databases, predictive algorithms and software tools (acevedo et al. 2016). however, it is practical, and it has been accepted by major stakeholders. by 2005, more than 1000 permit applications had been assessed with the corresponding recommendations issued to the authority in the ministry of agriculture. invasive species a major use of conabio’s databases and capabilities in biodiversity informatics has been in assessing the risk of invasive species, mostly plants of economic importance (goettsch et al. 2021). the first example originated with an information request from the u.s. department of agriculture, about any known occurrences of the moth cactoblastis cactorum, a well-known pest of cacti (zimmermann et al. 2000) in mexico. this request (via mexico’s ministry of agriculture) lead to one of the first niche modeling exercises (simonson et al. 2005; soberón et al. 2001) performed by conabio. after several attempts at convincing the mexican government about the importance of the problem, the ministry of agriculture of mexico finally organized a campaign of monitoring and control for this pest species (hernández et al. 2007). wildfire monitoring mexico is a large country, with complex topography and large forested and inaccessible regions. monitoring of wildfires is done by conabio via its remote sensing capabilities (conabio 2011). the system, entirely developed at conabio (ressl et al. 2009), uses daily data from the modis sensor, and state of the art algorithms, to produce maps (published daily online) of “hot points” across mexico, central america and the southern united states. the software automatically issues emails to relevant local authorities in areas of mexico where wildfires are spotted. it may be interesting to note that the capacities of conabio for remote sensing, as applied to wildfires, were the first test of the power and promise of a biodiversity informatics-focused organization. the daily data about the occurrence of wildfires over the entirety of mexico was a test not only of the technical capacities of the organization, but also of its political clout, since data about wildfires involved major budget investments, issues of federalism, and even issues of national security. conabio was, on a daily basis, monitoring the entire country, and issuing daily reports of direct relevance. one of the first tests of conabio’s commitment to open data was the wildfires system, since many powerful agents in the federal government were staunchly opposed to what eventually happened: the wildfires reports were made public, daily, over the internet. wildfires monitoring was also one of the first occasions for using biodiversity informatics in a diplomatic context, since conabio was monitoring wildfires also in central america. whether or not to share such information required diplomatic negotiations. ecosystem monitoring the capacity to monitor wildfires lead quickly to other monitoring initiatives. specifically, conabio initiated efforts to monitor mangrove cover (valderrama et al. 2014), marine photosynthetic activity (cerdeira-estrada and lópez-saldaña 2008), and ecosystem health (garcía-alaniz et al. 2017; gebhardt et al. 2014). the capacity to use remote sensing to monitor functioning of ecosystems is of great utility to government agencies. however, since biodiversity is a multi-scale phenomenon, the components and processes at the local scales should not be forgotten. monitoring at the scale of populations and their interactions is a significant challenge, as i outline in the next section. wildlife monitoring recently, conabio has started attempts to monitor wildlife. in 2010, working as partners of the national commission of forestry (conafor, comisión nacional forestal) and of the national commission of protected areas (conanp, comisión nacional de areas naturales protegidas). conafor runs a forestry monitoring scheme, and conabio began adding recorders and infrared cameras to >3000 of the 25,000 monitoring sites that conafor maintains (medellín and corrales 2019)3. although 3000 monitoring sites appears to be a large number, mexico is a large country, with nearly 2m km2, so the density is only 0.0015 sites/km2. despite this low density, hundreds of thousands of sound or 3 https://sipecamdata.conabio.gob.mx/mapa. https://sipecamdata.conabio.gob.mx/mapa jorge soberón – biodiversity informatics for public policy 103 image files have been processed (dirzo et al. 2021)4. processing the deluge of data produced by cameras and recorders has required that conabio recruit experts in artificial intelligence and pattern recognition. moreover, the system requires active participation of local stakeholders, of ngos, and of government agencies at federal and state levels. this effort is at the level of pioneer, and its applications to policy are still in the future. biodiversity exploration where to conduct biodiversity explorations, which are expensive in funds, time, and personnel, was one of the first questions that conabio had to answer, to use public resources in an efficient way. this work was accomplished using the primary data repositories, combined with remote-sensing information about land use (soberón et al. 2004). essentially, conabio worked to identify areas that were simultaneously poorly sampled and with low human impact, to prioritize for exploration. for instance, there were large regions in the western sierra madre that were both unexplored (i.e., no specimens reported in any of the databases) and relatively well preserved, being very mountainous areas with few human settlements and no roads. this region was highlighted as a priority for exploration, and a call for relevant projects was issued in 2000. figure 3 illustrates the case for the state of durango (much of it covered in montane sierra madre ecosystems), which was identified as of high priority for retrospective data capture and digitization and de novo biodiversity explorations (soberón et al. 2004). the red line shows the point at which conabio began assigning priorities for funding based on existing databases. with a delay, key data started pouring into the system. scientific articles one last use of conabio’s data that should not be forgotten is to enhance capacity for research by the mexican biodiversity science community. this research community has taken good advantage of the massive, new, and unprecedented availability of data (peterson et al. 2016; rodríguez et al. 2017). this effect is illustrated in the graphs in figure 4. conclusions the national biodiversity agency of mexico performs a large variety of functions, including diplomatic, legislative and educational (sarukhan 2018). 4 https://sipecamdata.conabio.gob.mx/manual. however, the core of its capacities, what truly distinguishes it from other government agencies in mexico, is its solid empirical grounding in primary data. the time, money, and human effort (the result of literally hundreds of years of biological research about mexico, nationally and internationally) spent building a powerful, comprehensive data system provide the agency with its credibility. this credibility is one of the keystones of the process of translating from science to policy-making (cash et al. 2003; soberón 2004). when the scientists and negotiators of conabio argue in mexico’s congress, or negotiate in an international forum, they have the credibility that comes from positions solidly grounded on primary, verifiable, open data. moreover, the amount of research that the data made available by conabio has enabled is difficult to quantify. one can count number of papers published, but the number of internal reports in government agencies, dissertations, and other “gray” uses of data is impossible to quantify. anecdotally, however, it is known that the system of conabio is widely used. conabio was made possible by the vision of pioneers, and a very singular political environment that allowed mexico to create a politically and economically independent organization, capable of issuing science-based opinions at a high governmental level. political circumstances have changed, however, such that now conabio has been deprived of its figure 3. number of specimens in conabio’s databases for the state of durango, identified as a high priority in the year 2000 (red dashed line) https://sipecamdata.conabio.gob.mx/manual jorge soberón – biodiversity informatics for public policy 104 economic independence. it may be in the process of losing its political independence as well. the costa rican inbio has also disappeared, or collapsed (fonseca 2015), and the indian initiative on bioinformatics is also non-existent. of the original biodiversity institutions that visited erin in 1992, only the australian initiative survives, in the form of the atlas of living australia project. the long-term survival of any institution depends on a combination of political, economic, and social factors. conabio was created by the fortunate combination of a diplomatic need for mexico to have something to present at the earth summit conference, and the fact that the most prominent ecologist of mexico was also the chancellor of the national university at the time. given its hybrid private-public design, conabio was able to build an impressive capacity to assemble, organize, and analyze biodiversity data. moreover, the organization was acting as a bridge (cash et al. 2003; soberón 2004) between academia and decision-making in the federal government. this combination, however, has not survived changes in the political world of mexico. it is difficult to speculate what combination of factors could have maintained conabio as an independent, fully funded government agency. conabio’s hybrid design allowed it to maintain some of its assets (i.e., computing cluster, remote-sensing capacities, databases…) as private, thus providing some degree of permanence, but the cross-cutting multiple ministries character and conabio’s budgetary and political independence are probably gone for good. it is to be hoped that the huge data resources of conabio, still openly available on-line, will remain so, via mirrors like gbif and others, although even a multinational initiative like gbif is vulnerable to budgetary constraints. it is now clear that if scientists want to keep primary data openly available, databases probably will need to be spread over many independent organizations, to minimize the risk of collapse due to failure of one main participant. this perhaps will protect, at least, the purpose of the data, if not the organizations as such. literature cited acevedo, f., e. huerta, and c. burgeff. 2016. biosafety and environmental releases of gm crops in mesoamerica: context does matter in r. lira, a. casas and j. blancas, eds. ethnobotany of mexico. interactions of people and plants in mesoamerica. springer, new york. asase, a., m. sainge, r. radji, u. omokafe, and t. peterson. 2020. a new model for efficient, need‐driven progress in generating primary biodiversity information resources. applications in plant sciences 8:e11318. balmford, a., p. crane, a. dobson, r. green, e., and g. mace. 2005. the 2010 challenge: data availability, information needs and extraterrestrial insights. philosophical transactions of the royal society b 360:221-228. figure 4. impact of conabio on biodiversity research outputs. numbers of publications and citations were drawn from web of science, using the key word conabio in title or in funding source. left graph, number of papers published. right graph, citations to papers published, since 2000. jorge soberón – biodiversity informatics for public policy 105 beaman, r., and n. cellinese. 2012. mass digitization of scientific collections: new opportunities to transform the use of biological specimens and underwrite biodiversity science. zookeys 209:7-17. belbin, l., wallis, e. hobern, d., zerger, a. 2021. the atlas of living australia: history, current state and future directions. biodiversity data journal 9:e65023. berners-lee, t. 1992. the world-wide web. computer networks and isdn systems 25:454-459. bibby, c. j. e. a. 1992. putting biodiversity on the map: priority areas for global conservation. international council for bird preservation, cambridge, uk. booth, t. 2018. why understanding the pioneering and continuing contributions of bioclim to species distribution modelling is important. austral ecology 43:852-860. brooks, t. m., r. mittermeier, g. da fonseca, j. gerlach, m. hoffmann, j. lamoreux, c. mittermeier, j. d. pilgrim, and a. s. rodrigues. 2006. global biodiversity conservation priorities. science 313:58-61. busby, j. r., c. r. margules, and m. p. austin. 1991. bioclim a bioclimate analysis and prediction system. pp. 64 in c. r. margules and m. p. austin, eds. nature conservation. cost-effective biological surveys and data analysis, canberra, australia. butchart, s. h. m., m. walpole, b. collen, a. van strien, j. p. w. scharlemann, r. e. a. almond, j. e. m. baillie, b. bomhard, c. brown, j. bruno, k. e. carpenter, g. m. carr, j. chanson, a. m. chenery, j. csirke, n. c. davidson, f. dentener, m. foster, a. galli, j. n. galloway, p. genovesi, r. d. gregory, m. hockings, v. kapos, j.-f. lamarque, f. leverington, j. loh, m. a. mcgeoch, l. mcrae, a. minasyan, m. h. morcillo, t. e. e. oldfield, d. pauly, s. quader, c. revenga, j. r. sauer, b. skolnik, d. spear, d. stanwell-smith, s. n. stuart, a. symes, m. tierney, t. d. tyrrell, j.-c. vié, and r. watson. 2010. global biodiversity: indicators of recent declines. science 328:1164-1168. canhos, d. 2017. data management plan: brazil’s virtual herbarium. research ideas and outcomes 3:e14675. cash, d. w., w. c. clark, f. alcock, n. m. dickson, n. eckley, d. h. guston, j. jager, and r. b. mitchell. 2003. science and technology for sustainable development special feature: knowledge systems for sustainable development. proceedings of the national academy of sciences usa 100:80868091. cerdeira-estrada, s., and g. lópez-saldaña. 2008. automatic processing of near-real time operational modis ocean products applied to mexico seas monitoring. 2008 5th international conference on electrical engineering, computing science and automatic control, 545-549. chapman, a. 1991. the role of specimen-backed information in environmental decision making -the australian experience. pp. 1. symposium at the australian national botanical gardens, canberra, australia. chapman, a. 2001. biodiversity informatics, biota/fapesp and the future: a personal view. biota neotropica 1:1-9. chapman, a. 2005. principles and methods of data cleaning-primary species and species-occurrence data. global biodiversity information facility, copenhagen. coetzer, w. 2012. a new era for specimen databases and biodiversity information management in south africa. biodiversity informatics 8:1-11. conabio. 2011. sistema de alerta temprana de incendios forestales en méxico y centroamérica. national commission on biodiversity, mexico conabio. 2012. conabio: two decades of history, 19922012. pp. 1-36 in. ed. comision nacional para el conocimiento y uso de la biodiversidad, mexico d. f., mexico. dallwitz, m. j. 1993. delta and intkey. pp. 287-296 in r. fortuner, ed. advances in computer methods for systematic biology: artificial intelligence, databases, computer vision. johns hopkins university press, baltimore, usa. dirzo, r., o. lópez, p. maeda, r. mejía, m. munguía-carrara, e. robredo, and m. schmidt. 2021. manual de monitoreo: sitios permanentes de calibración y monitoreo de la biodiversidad. pp. 85. conabio, mexico city. edwards, j. l., m. lane, and e. nielsen. 2002. interoperability of biodiversity databases: biodiversity information on every desktop. science 289 2312. ehrlich, p. 1995. the scale of the human enterprise and biodiversity loss. pp. 233 in j. h. lawton and r. m. may, eds. extinction rates. oxford university press, oxford. fonseca, p. 2015. a major center of biodiversity research crumbles. scientific american e the sciences section. http://www. scientificamerican. com/article/a-major-center-of-biodiversityresearch-crumbles gadgil, m., f. berkes, and c. folke. 1993. indigenous knowledge for biodiversity conservation. ambio 22:151-156. garcía-alaniz, n., m. equihua, o. pérez-maqueo, j. e. benítez, p. maeda, f. p. urrutia, j. j. f. martínez, s. a. v. gaytán, and m. schmidt. 2017. the mexican national biodiversity and ecosystem degradation monitoring system. current opinion in environmental sustainability 26:62-68. gebhardt, s., t. wehrmann, m. a. m. ruiz, p. maeda, j. bishop, m. schramm, r. kopeinig, o. cartus, j. kellndorfer, and r. ressl. 2014. mad-mex: automatic wall-to-wall land cover monitoring for the mexican redd-mrv program using all landsat data. remote sensing 6:3923-3943. goettsch, b., t. urquiza-haas, p. koleff, f. acevedo, a. aguilar-melendez, and e. al. 2021. extinction risk of mesoamerican crop wild relatives. plants people planet 3:775795. gonzalez, c., e. mora, and m. munguia. 2014. modeling ecological integrity with bayesian belief networks. pp. 1-3 in conabio, ed. conabio, mexico. graves, g. 2000. costs and benefits of web access to museum data. trends in ecology and evolution 15:374. groves, c., klein, m. breden, t. 1995. natural heritage programs: public-private partnerships for biodiversity conservation. wildlife society bulletin 23:784-790. guisan, a., r. tingley, j. b. baumgartner, i. naujokaitis-lewis, p. r. sutcliffe, a. i. t. tulloch, t. j. regan, l. brotons, e. http://www jorge soberón – biodiversity informatics for public policy 106 mcdonald-madden, c. mantyka-pringle, t. g. martin, j. r. rhodes, r. maggini, s. a. setterfield, j. elith, m. w. schwartz, b. a. wintle, o. broennimann, m. austin, s. ferrier, m. r. kearney, h. p. possingham, and y. m. buckley. 2013. predicting species distributions for conservation decisions. ecology letters 16:1424-1435. hardin, g. 1968. the tragedy of the commons. science 162:1243-1248. hernández, j., h. sánchez, a. bello, and g. gonzález. 2007. preventive programme against the cactus moth cactoblastis cactorum in mexico. pp. 345-350 in m. j. b. vreysen, a. s. robinson and j. hendricks, eds. area-wide control of insect pests. iaea, vienna, austria. hortal, j., f. de bello, j. a. f. diniz-filho, t. m. lewinsohn, j. m. lobo, and r. j. ladle. 2015. seven shortfalls that beset large-scale knowledge of biodiversity. annual review of ecology, evolution, and systematics 46:523-549. hruby, f., s. melamed, r. ressl, and d. stanley. 2016. mosaicking mexico. the big-picture of big-data. pp. 407-411. international archives of the photogrammetry, remote sensing & spatial information sciences, prague, czech republic. johnson, n. f. 2007. biodiversity informatics. annual review of entomology 52:421-438. kaye, p., s. noble, and w. slater. 1997. environmental information for intelligent decisions. pp. 245-258. intelligent environments. elsevier. koester, v. 2002. the five global biodiversity-related conventions: a stocktaking. review of european community & international environment law 11:96-103. laihonen, p. k., r. salo, j. 2004. the biodiversity information clearing-house mechanism as a global effort. environmental science and policy 7:99-108. lane, m., and j. edwards. 2007. the global biodiversity information facility. pp. 1-4 in c. humphries, ed. systematics association special volume. crc press, boca raton, ca. maclaurin, j., and k. sterelny. 2008. what is biodiversity? the university of chicago press, chicago. mcneely, j., k. r. miller, w. v. reid, r. mittermeier, and t. werner. 1990. conserving the world’s biological diversity. the world bank, washington, d. c. medellín, c., and l. corrales. 2019. sistemas de monitoreo forestal en méxico. pp. 84. serie técnica. boletín técnico. catie, turrialba, costa rica. nelson, g., and s. ellis. 2019. the history and impact of digitization and digital data mobilization on biodiversity research. philosophical transactions of the royal society b 374:20170391. nix, h. a. 1986. a biogeographic analysis of australian elapid snakes in r. longmore, ed. atlas of elapid snakes of australia. australian government publishing service, canberra. noss, r. f. 1990. indicators for monitoring biodiversity: a hierarchical approach. conservation biology 4 355-364. ostrom, e., j. burger, c. b. field, r. b. norgaard, and d. policansky. 1999. revisiting the commons: local lessons, global challenges. science 284:278-282. peterson, a. t., s. knapp, r. guralnick, j. soberón, and m. t. holder. 2010. the big questions for biodiversity informatics. systematics and biodiversity 8:159-168. peterson, a. t., a. g. navarro-sigüenza, and a. gordillo-martínez. 2016. the development of ornithology in mexico and the importance of access to scientific information. archives of natural history 43:294-304. peterson, a. t., and j. soberón. 2018. essential biodiversity variables are not global. biodiversity and conservation 27:1277-1288. pressey, r. l., c. j. humphries, c. r. margules, r. i. vanewright, and p. h. williams. 1993. beyond opportunism: key principles for systematic reserve selection. trends in ecology & evolution 8 124-128. reed, g. 2017. the clearing-house mechanism: an effective tool for implementing the convention on biological diversity? pp. 115-126. governing global biodiversity. routledge. reid, w. l., s., meyer, c. gamez, r. sittenfeld, a., janzen, d. gollin, d. juma, c. 1996. biodiversity prospecting. pp. 142-173 in w. reid, ed. biodiversity prospecting. using genetic resources for sustainable development. world resources institute, washington, dc. reid, w. v. 1998. biodiversity hotspots. trends in ecology & evolution 13:275-280. ressl, r., g. lopez, i. cruz, r. colditz, m. schmidt, s. ressl, and r. jimenez. 2009. operational active fire mapping and burnt area identification applicable to mexican nature protection areas using modis and noaa-avhrr direct readout data. remote sensing of environment 113:11131126. rodríguez, p., f. villalobos, a. sánchez-barradas, and m. correa-cano. 2017. la macroecología en méxico: historia, avances y perspectivas. revista mexicana de biodiversidad 88:52-64. sandlund, o. t. 1991. costa rica’s inbio: towards sustainable use of natural biodiversity. pp. 1-25 in n. i. f. naturforskning, ed, trondheim. sarkar, s. 2002. defining biodiversity; assessing biodiversity. the monist 85:131-155. sarukhan, j. 2018. conabio, 25 years of evolution. pp. 160. comisión nacional para el conocimiento y uso de la biodiversidad, mexico. sarukhán, j., and r. dirzo. 1992. mexico ante los retos de la biodiversidad. pp. 343. universidad autonoma de chapingo, mexico df. sarukhán, j., and r. jiménez. 2016. generating intelligence for decision making and sustainable use of natural capital in mexico. current opinion in environmental sustainability 19:153-159. sarukhán, j., t. urquiza-haas, p. koleff, j. carabias, r. dirzo, e. ezcurra, s. cerdeira-estrada, and s. jorge. 2014. strategic actions to value, conserve, and restore the natural capital of megadiversity countries: the case of mexico. bioscience 65:164-173. jorge soberón – biodiversity informatics for public policy 107 scott, j. m. 1993. gap analysis: a geographic approach to protection of biological diversity. wildlife monographs 123:141. siebert, s., and g. smith. 2004. lessons learned from the sabonet project while building capacity to document the botanical diversity of southern africa. taxon 53:119-126. simonson, s. e., t. stolhgren, l. tyler, w. p. gregg, m. rachel, and j. garrett. lynn. 2005. preliminary assessment of the potential impacts and risks of the invasive cactus moth, cactoblastis cactorum berg, in the u.s. and mexico international atomic energy agency, vienna, austria. sittenfeld, a., and r.villers. 1993. exploring and preserving biodiversity in the tropics: the costa rican case. biotechnology 4 280-285. soberón, j. 2004. translating life’s diversity: can scientists and policymakers learn to communicate better? environment 46:10-20. soberón, j., p. dávila, and j. golubov. 2004. targeting sites for biological collections in r. r. smith, j. b. dickie, s. linington, h. pritchard and r. probert, eds. seed storage: turning science into practice. royal botanic gardens, kew, uk. soberón, j., j. golubov, and j. sarukhan. 2001. the importance of opuntia in mexico and routes of invasion and impact of cactoblastis cactorum (lepidoptera: pyralidae). florida entomologist:486-492. soberón, j., e. huerta, and l. arriaga. 2002. the use of databases to assess the risk of gene flow: the case of mexico. pp. 61-67 in c. r. roseland, ed. lmos and the environment. organization for economic cooperation and development, paris. soberón, j., r. jiménez, p. koleff, and j. golubov. 2010. la informática sobre la biodiversidad: datos redes y conocimiento. pp. 354 in v. m. toledo, ed. la biodiversidad de méxico. el fondo de cultura económica, méxico d. f. soberón, j., and p. koleff. 2000. the national biodiversity information system of mexico. pp. 625 in p. raven, ed. nature and human society: proceedings of the 1997 forum on biodiversity. national academies press, washington, d. c. soberón, j., and a. t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. philosophical transactions of the royal society b 35:689-698. sousa‐baena, m. s., r. garcía, l. couto, and a. t. peterson. 2014. completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. diversity and distributions 20:369-381. steffen, w. w. b., l. deutsch, o. gaffney, c. ludwig. 2015. the trajectory of the anthropocene: the great acceleration. the anthropocene review 2:81-98. stein, b., l. kutner, and j. s. adams. 2000. precious heritage. the status of biodiversity in the united states. the nature conservancy and oxford university press, new york. stephens, c., r. sierra-alcocer, c. gonzález-salazar, j. barrios, j. c. salazar, e. robredo, and e. del callejo. 2017. species: a platform for the exploration of ecological data. ecology and evolution 9:1638-1653. tangley, l. 1990. cataloging costa rica’s diversity. bioscience 40:633-636. tegelberg, r., t. mononen, and h. saarenmaa. 2014. high‐performance digitization of natural history collections: automated imaging lines for herbarium and insect specimens. taxon 63:1307-1313. ten kate, k. 1999. legal aspects of regulating access to genetic resources and benefit-sharing: the convention on biological diversity, national and regional laws and material transfer agreements. earthscan, london, uk. toledo, v. m. 2001. indigenous peoples, and biodiversity. pp. 1181-1203 in s. levin, ed. encyclopedia of biodiversity. academic press. valderrama, l., c. troche, m. rodriguez, d. marquez, b. vázquez, s. velázquez, a. vázquez, m. cruz, and r. ressl. 2014. evaluation of mangrove cover changes in mexico during the 1970–2005 period. wetlands 34:747-758. wegier, a., a. piñeyro-nelson, j. alarcón, a. gálvez-mariscal, e. r. álvarez-buylla, and d. piñero. 2011. recent long-distance transgene flow into wild populations conforms to historical patterns of gene flow in cotton (gossypium hirsutum) at its centre of origin. molecular ecology 20:4182-4194. wieczorek, j., d. bloom, r. guralnick, s. blum, m. doring, r. giovanni, t. robertson, and d. vieglais. 2012. darwin core: an evolving community-developed biodiversity data standard. plos one 7:e29715. wieczorek, j., q. guo, and r. j. hijmans. 2004. the point-radius method for georeferencing locality descriptions and calculating associated uncertainty. international journal of geographic information science 18:745-767. zimmermann, h., v. moran, and j. hoffmann. 2000. the renowned cactus moth, cactoblastis cactorum: its natural history and threat to native opuntia floras in mexico and the united states of america. diversity and distributions 6:259-269. biodiversity informatics, 12, pp. 58-75 crowdsourcing natural history archives: tools for extracting transcriptions and data katherine mika*, joseph deveer, and constance rinaldo ernst mayr library, museum of comparative zoology, harvard university, 26 oxford st, cambridge, ma 02142 usa. *corresponding author: kmika@fas.harvard.edu abstract.—this paper surveys the landscape of current, successful, and innovative crowdsourcing platforms for obtaining full text transcriptions and structured datasets hidden in manuscript items in the biodiversity heritage library. transcribing manuscripts are optimal tasks for crowdsourcing programs because they require intellectual engagement and thoughtful decision making to produce meaningful content. by offering full text transcriptions, digital collections are opened up to new types of searching, sorting, categorizing, and pattern finding. research derived from these new datasets can illustrate changes over time across much larger magnitudes of collections and types of information resources. a targeted analysis of methods, tools, and programs for crowdsourcing manuscript transcriptions describes the challenges and opportunities in developing a project that produces machine readable facsimiles and can support structured data extraction from natural history libraries and special collections content. key words.—field notes, transcription, data capture, primary biodiversity data, crowd-sourcing the biodiversity heritage library (bhl) is a global collaborative digital library established in 2006, with a mission to improve research methodology by making biodiversity literature openly available and inspiring discovery through free access to biodiversity knowledge (gwinn and rinaldo 2009). the library is adding manuscript collections to its collection in order to expose hidden and often encrypted observation and occurrence data. these data are collected from machine readable text transcriptions and document the environment, climate, and biodiversity over wide temporal and geographic domains. bhl in partnership with the library of congress’ national digital stewardship residency program has made it a priority to add text from manuscript transcriptions to enhance the quality and completeness of machine readable biodiversity data in the digital library portal and developer interface. manuscript items contain a wealth of occurrence data which cannot be investigated without reliable transcriptions. a general survey of transcription utilities revealed that there are many types of projects with dramatically different goals and user needs. therefore, a targeted survey will provide better guidance on methods, tools, and programs that can support structured natural history data extraction as well as free text transcription. current crowdsourced transcription platforms build on the successes of earlier projects and learn from their mistakes and challenges. because a consortium digital library depends on contributions from its members, a transcription platform must be easy to implement, generate output that is compatible with various types of transcription files, and function within several layers and types of encoding. the selected program will likely be used by libraries to transcribe local items in addition to content for bhl. several successful platforms have been identified as potential tools with which to develop a cohesive transcription program for bhl. the bhl digital library portal is built on a relational database that defines and indexes relationships between titles, items, pages, and segments and their administrative, structural, and descriptive metadata (figure 1). as scientists increasingly turn to computational methods to answer large scale questions about 58 mika et al. – crowdsourcing natural history archives the natural world, bhl plans to update its data model to provide better access to collection data and taxonomic, bibliographic, and descriptive metadata. natural history libraries are mandated to preserve and maintain the published literature that is critical to the discovery, revision, and naming of life. over time, nomenclature specialists generated species citations according to discipline and sub-discipline specific abbreviations independent of related species projects or the expertise of those collecting and providing accessto taxonomic literature. concurrently, librarians developed and implemented metadata policies without consulting taxonomists or the scientists using the information (pilsk et al. 2010). this has resulted in a corpus of heritage literature and special collections that preserves taxonomic information and citations but is relatively inaccessible to the modern scientist. in order to mitigate this, bhl relies on the global names recognition and discovery service1 to mine and index the literature’s text for stringsof potential linnaean binomials, resolve them to existing taxonomic databases, and link the strings to pages in the bhl portal. field notes the biodiversity heritage library is adding unpublished manuscript materials from its contributors in the form of collector’s field notes, diaries, and correspondence. because optical character recognition (ocr) in its current state of development does not accurately render handwritten materials, manual transcription of digitized manuscripts is essential for the creationof text files for each page (kalfatovic and rinaldo 2016). natural history field notebooks and diaries are iconic symbols of scientific work and are important as a sequential record along with correspondence. notebooks and correspondence showcase individual work but also provide insight into the personality of the writer and events of the time. natural history, 1http://gnrd.globalnames.org/. defined as biodiversity and environment of a region, is part of the cultural heritage that characterizes a nation or region and informs policy making (europe 2005; rinaldo and smith 2014). the recorded information in field notes and correspondence also reflects changes in the philosophy of the study of natural history through time. figure 1: adapted from bhl data model to illustrate how titles, items, pages and segments are related in the bhl portal2. field notes and correspondence contain a wealth of raw scientific data including unpublished observations, species occurrence records, habitat descriptions, climatological data, phenological recordings, sketches, weather reports, and travel narratives: these records are primary source data at its most raw and unevaluated. historical collections of field notes may be the only documentation of a scientist’s thought processes, ideas, and observations, particularly if only some of the material was ever published (deveer, rinaldo, and ford 2013). when researched over time, biologists’ field notes often document local environmental conditions and may help to identify gradual changes in number and presence of species over 2 https://github.com/gbhl/bhl-us/blob/master/documentation/datamodel. 59 mika et al. – crowdsourcing natural history archives many years. comparisons of species lists and descriptions of environmental conditions with current conditions provide valuable information about landscape and environmental changes and may pinpoint historical changes such as for species migrations. digitized field notes provide an opportunity to review historical, cultural, and weather events of the times described in these documents in comparison with current conditions. figure 2: excerpt from the waldo l. schmitt papers that contains valuable observation and meteorological data and illustrates the creator’s dynamic handwriting (schmitt 1962). extracting specific information from handwritten documents using ocr is difficult, if not impossible. some notebooks are crowded with writing in many different directions in an attempt to use as muchof the blank page as possible (figure 2). to index this information and make it discoverable, machine-readable text files for manuscripts are needed in the bhl database so that text files can be produced for indexing as they are for published items such as books and journal volumes. the ultimate goal for bhl is to mine the vast amount of biodiversity data locked away in field notes, correspondence, images, and published texts and make the data accessible from the bhl corpus. available information in bhl for a given species would then potentially include the original description of that species in the published literature with a link to an account of the field collection of the type specimen used as a basis for the description and to images in texts. eventually, links to museum specimens and other non-literaturebased data and objects will collocate all the elements of biodiversity information. crowdsourcing digitized manuscript items are often hidden and inaccessible in digital libraries because their descriptive metadata is frequently minimal and their unique content is not discoverable without a machine-readable facsimile. indexing transcriptions facilitates discovery of historical records and improves catalog search results. by offering full text transcriptions, digital collections are opened up to new types of searching, sorting, categorizing, and pattern finding. research derived from these new datasets can illustrate changes over time across much larger magnitudes of collections and types of information resources. this is particularly important when considering biodiversity heritage literature and archives collections due to their significant value in documenting species occurrences, botanical observations, climate patterns, and meteorological events. transcriptions facilitate the manipulation of this data and support research that extracts knowledge from formal and informal collecting and observation events. transcription projects for collections are time consuming, intellectually intensive, and expensive for an organization to 60 mika et al. – crowdsourcing natural history archives facilitate. crowdsourcing has been identified as a sustainable model for generating transcriptions for large collections and institutions with diverse holdings, and may improve data collection from a diverse range of users to enhance descriptive metadata. after reviewing various types of collaborations and crowdsourcing in libraries, sally ellis argues that librarianship is strengthened by “cooperation with, and contributions from, users; relationships among libraries, archives, museums (lam) and other information institutions; and the use of emerging technologies to facilitate these associations and interactions” (ellis 2014). crowdsourcing ensures a future for libraries, museums, and archives because it solves problems, strengthens collections and communities, and engages users. glam institutions (galleries, libraries, archives, and museums) are well prepared and appropriately situated to implement successful crowdsourcing projects due to their commitment to and history of facilitating public engagement. while discovery systems, online catalogs, and web 2.0 tactics have been widely adopted to enhance remote access to collections, cultural heritage institutions are often reluctant to cede control of content selection, description, discovery, and use to digital volunteers (holley 2010). collaboration across institutions and users enhances access, facilitates use, and strengthens collections. daren brabham defines the term as an “online, distributed problemsolving and production model that leverages the collective intelligence of online communities to serve specific organizational goals” (brabham 2013). tech companies, businesses, and academic research institutions have embraced crowdsourcing as a legitimate tool for generating new products, content, and operations, as well as for public relations and marketing purposes (brabham 2013). the internet’s speed, reach, temporal flexibility, anonymity, interactivity, and convergence brings people into conversation with each other, lowers barriers to information by creating easier access to professional bodies of knowledge, increases access to useful tools, and enables an online participatory culture (brabham 2013). by externalizing transcriptions of manuscript items, we can leverage the collective intelligence and wisdom of crowds and exploit a large and diverse set of skills, tools, and ideas to bear on archival materials and special collections. the internet encourages ongoing co-creation of new ideas in which content is generated through a mix of bottom-up (from the people) and top-down (policy-makers, businesses, and media organizations) processes (brabham 2013). glams are ideal institutions to encourage and utilize crowdsourcing initiatives due to their unique placement at the intersection of these processes. libraries and cultural heritage institutions have the advantages of mission statements and codified ideologies dedicated to enriching knowledge as well as the organizational structures to mobilize, energize, and capitalize reciprocally on the capabilities of its users. this symbiotic relationship is not only mutually beneficial, but is likely one of the spaces in which glams can thrive in the digital age. biodiversity research has a strong background in relying on non-scientist community members to collect data. these citizen scientist programs and the resulting data are understood as a “public good that is generated through increasingly collaborative tools and resources while supporting public participation in science and earth stewardship” (dickinson et al. 2012). tracking and understanding biodiversity at varying scales requires fine-grain data to be collected over regions and continents, years and decades. professional scientists alone are not generally capable of delivering the volume of data, analysis, and interpretation needed to support large-scale biodiversity research questions (theobald et al. 2015). the study of sweeping patterns in nature requires vast amounts of data to be collected across an array of locations and habitats over span of years and often decades (bonney et al. 2009) (figure 3). 61 figure 3: excerpt from the robert e. silberglied papers, 1960-1982 illustrates the complexity of data capture in field notes and manuscript content. this page includes varying date and location information, latin and common scientific names, occurrence and collection data, a map, and a photo foldout among other potential data points (silberglied 1965). figure 4: list of observed species from ornithologist william brewster’s field notes that illustrate the transcription difficulties in deciphering shorthand and abbreviations that are especially prevalent in scientist’s notes (brewster 1892). 62 mika et al. – crowdsourcing natural history archives transcribing field notebooks and manuscripts, correcting ocr output, and curating collections occurrence data are optimal tasks for crowdsourcing and citizen science programs because they require intellectual engagement and thoughtful decision making to produce meaningful content. collecting every potential data point from all possible documents generates datasets that are difficult to understand and obscured with irrelevant information. crowdsourcing applies standardized processes for transforming content germane to scientific inquiry and collections’ scopes into machine readable data points formatted for computational research processes. crowdsourcing transcriptions can be understood as a method of gathering data over wider geographical and temporal spaces. field notes become powerful and rich sources of biodiversity information when the existing data is transformed into machine readable format. by transcribing and generating structured datasets from field notes, scientists of yore can be recruited for current research projects. bhl’s content spans hundreds of years and the entire globe, creating a potentially vast pool of observation data that can inform current research (figure 3). in the same way that science departments have turned to public participation to enlist the community in creating scientific knowledge, crowdsourcing transcriptions creates global networks that can generate data to be analyzed for population trends, range changes, shifts in phenologies, and more (dickinson and bonney 2012). the crowdsourcing platforms discussed in this paper are the tools the biodiversity heritage library is considering in order to develop a consortium-wide standard for extracting data from digitized items to improve the discoverability of hidden collections. the ndsr bhl transcription work addresses similar goals to the art of life (rose-sandler 2012) and purposeful gaming (rose-sandler 2015) projects that sought to enrich the metadata of items to better facilitate access to collections. the work of art of life sought to “liberate natural history illustrations from the digitized books and journals in the online biodiversity heritage library through the development of software tools for automated identification and description of visual resources” (rose-sandler 2012). images in bhl include page level structural metadata, facilitating navigation by human users and citation resolvers, but they lack sufficient descriptive metadata to enable dynamic filtering and inquiry. the art of life project built new software tools and algorithms to automatically identify illustrations found within the text pages of the bhl corpus and push those illustrations to crowdsourcing environments like flickr and wikimedia commons for their description. similarly, full text searching of texts is significantly hampered by poor output from ocr software, and historic literature has proven to be particularly problematic because of its tendency to have varying fonts, typesetting, and layouts that make it difficult to accurately render. purposeful gaming was developed to identify a method for quick and efficient harnessing of large numbers of users to review and correct particularly problematic works by presenting the task as a game. both projects improve the discoverability of and access to digital texts by enriching descriptive metadata for items at the page level to support full-text searching, data mining, and markup of content in bhl collections. the ndsr transcription project complements art of life and purposeful gaming by developing a similar method for generating machine readable content that will enhance access to handwritten text, a final category of “hidden content” in bhl. transcription in the realm of manuscript transcription crowdsourcing platforms there are essentially two types of projects: record-based and document-based (brumfield 2012). recordbased projects are more closely aligned with citizen science crowdsourcing because they seek to extract tabular data from handwritten materials (figure 4). the output from these 63 mika et al. – crowdsourcing natural history archives projects can easily be stored in databases and searched, sorted, and categorized for findability and disseminated via apis (application programming interfaces). content in these items is usually transcribed in online forms that structure the data in appropriate schemas depending on the project goals and the types of relevant information. users and creators of these projects generally understand in advance what type of data is going to be produced from their collections and the research and analysis the data needs to support. record-based projects are also not usually interested in representing the content of an entire document.(brumfield 2012) structured datasets generated from projects like old weather (blaser 2014), familysearch indexing (holley 2010; hanse et al. 2012), and the north american bird phenology program and other specimen transcription citizen science projects (miller-rushing,primack,andbonney 2012; dickinson and bonney 2012) follow this record-based strategy in order to support research over very broad geographic and temporal scales. document-based projects, in contrast, produce transcriptions that try to replicate the original digitized item—all aspects of it—as closely as possible, and in a machine-readable format (brumfield 2012). managers of these types of projects are usually interested in a full text output (encoded or unencoded) that records all of the content on a page. platforms generally invite volunteers to type their transcriptions into a free-form text box that can support some markup conventions and may include toolbars to encourage consistency (brumfield 2012). there are some encoding schemas and conventions that make it easier to communicate non-textual information digitally, but the kinds of markup for these types of projects are idiosyncratic at best, and do not adhere to a semantic standard or professional best practices. many digital humanities and archives institutions are rapidly adopting tei (text encoding initiative) 3http://eol.org/. 4https://www.gbif.org/. metadata schema for transcription markup, and while it is a good option that structures datafor display effectively (line breaks, annotations, insertions, deletions, etc.), it is fairly limited in its ability to digitally represent informational content outside traditional archives scopes, like natural history. the biodiversity heritage library is considering combining record-based and document-based transcription in order to represent all of the relevant informational content in a document in addition to extracting specific structured datasets that are of particular use to its researchers and users (figure 5). bhl users include staff members at consortium institutions, other digital systems that harvest bhl content and data, and individuals that access items at all levels. the encyclopedia of life3, the global biodiversity information facility4, biostor5, and the global names architecture6 are systems users that link to bhl’s content, often via taxon names, for bibliographies and foundational literature (mcclanahan 2017). georeferencing and including transcribed field notes and archival collections will further enhance interoperability and connections between types of biodiversity data. in addition to traditional access, individual users pull datasets from bhls publicly available apis in order to mine contents (primarily) for species occurrences. currently, the bhl data model (figure 6) describes titles, pages, and segments by defining relationships with keywords, authors, and scientific names. authors and keywords can describe segments or titles, but not pages or items. scientific names are attached to pages, but not titles, items, or segments. for field notes and other manuscript collections additional metadata tables such as locations, dates, added personal names and corporate bodies, and perhaps expeditions could be added. currently dates are defined at the title level via an imported marc record, and expeditions are 5http://biostor.org/. 6http://globalnames.org/. 64 figure 5. image features manuscript page from ornithologist william brewster’s diary (1865), the automatically generated ocr, and its transcription with potential species occurrence and description data points that bhl would like to capture highlighted (brewster 1892). figure 6. illustrates the relationships between titles, items, pages, and segments, and associated keywords, authors, and scientific names. https://github.com/gbhl/bhl-us/blob/master/documentation/datamodel. 65 mika et al. – crowdsourcing natural history archives linked through curated bibliographies. bhl will need to consider where to add new tables, what description level they should be at, and whether to add or modify description levels of the existing metadata tables. early tools two early platforms, scripto7, developed by the roy rosenzweig center for history and new media at george mason university (rrchnm) and transcribe bentham8 shaped the landscape for successful transcription projects. archivists at rrchnm synthesized web 2.0 tactics and crowdsourcing successes from outside of the cultural heritage and glam arena to build an application to transcribe the papers of the war department (pwd)9. in addition to leveraging max evans’ concept of commons-based peer-production (evans 2007) the pwd project hoped to use transcriptions to enhance the collection’s findability. the rrchnm tool allows users to “easily submit those transcriptions and their knowledge back to the archive... and...draw upon the wisdom of the thousands of interested researchers, scholars, and students who work with these materials” (leon 2014). this combination of metadata enrichment and user engagement is the defining feature of crowdsourced humanities transcription platforms. natural history transcriptions can benefit from an additional layer that draws from the record-based implementation of transcription tools to generate structured datasets, which is most effectively accomplished via markup. scripto is available as a plugin for common content management systems (cms) including omeka, wordpress, and drupal, and best serves projects that use a cms as a repository for their project content. this model, adopted by the university of iowa and wellesley college creates a platform outside of the digital library interface to which images of the digitized items 7http://scripto.org/. 8http://blogs.ucl.ac.uk/transcribe-bentham/. 9http://wardepartmentpapers.org/. are uploaded and attached to a digital item record that includes a transcription file. transcriptions are typed by users into a text box with some style convention and slated for verification upon completion. in addition to transcribing activities, volunteers can communicate about items via disqus10 annotations or a discussion form that facilitates communication between transcribers and administrators. the university of iowa’s first transcription project, civil war diaries and letters11, was launched in 2011 and successfully enhanced collection access and usability by “enabling fulltext search of the content, and engaging the general public by allowing them to interact with the materials in new ways” (saylor and wolfe 2011). after the popularity of the program crashed their servers, ui sought a more user friendly and efficient solution in rrchnm’s customizable scripto. the transcriptions of anne whitney’s correspondence collection12 was developed out of a wellesley college undergraduate seminar to provide the wellesley community witha lifetime learning resource and enable access to the documents for a wider audience (bartle 2014). transcribed text allowed prof. jacqueline musacchio to use a range of dh tools to visualize whitney’s travel experience as clearly as possible by capturing evidence of movement through space and time and applying that evidence to historical maps (musacchio 2014). this is a critical similarity between humanities research and natural history research that transcriptions support. by providing access to transcribed text that connects to specimen occurrences with specific dates and geographic locations, natural historians can identify a broader and more complete picture of biodiversity. as part of the development process for scripto, rrchnm drew from other successful crowdsourcing programs including university college london’s (ucl) transcribe bentham. 10https://disqus.com/. 11https://diyhistory.lib.uiowa.edu/collections/show/8. 12http://omeka.wellesley.edu/whitneytranscribe/home. 66 mika et al. – crowdsourcing natural history archives archivists and researchers at ucl identified the need for a fully transcribed facsimile of jeremy bentham’s papers in order to better facilitate research and eventually publish a complete scholarly edition of bentham’s collected works for wider dissemination. the transcribe bentham digital platform, known as the “transcription desk,” is a mediawiki customization that supports tei encoding via a toolbar for ease of use. while mark-up was not required from volunteers, its adoption and use was significant and demonstrated that “volunteer labor can be used to undertake the type of detailed...tasks generally perceived to be the preserve of those trained in xml (extensible markup language) and tei” (causer and terras 2014). the validated transcriptions produced by transcribe bentham were uploaded to ucl’s digital repository and linked to the relevant manuscript item, enhancing access to the collection and improving primary source research. the ucl project was among the first large scale crowdsourcing projects for transcribing special collections items, and reflected a new focus in digital humanities scholarship: increasing and encouraging user engagement and providing open source tools that can be repurposed for other projects (causer and terras 2014). transcribe bentham was instrumental in demonstrating the sustainability of crowdsourcing approaches to both humanities scholarship and enhancing access to content in libraries in special collections. smithsonian transcription center the smithsonian’s transcription center13, built in 2012, is a popular system that is designed to extract full text transcriptions from archival collections. smithsonian libraries created a flexible program for 19 museums and archives that includes transcription, translation, and discussion features. to accommodate different formats, the team developed several different data structures for field notebooks, diaries, 13https://transcription.si.edu/. botanical specimen records, and numismatic proofs. the transcription center generates json (javascript object notation) files from text entered into a single data field. volunteers can utilize a wysiwyg-like toolbar that applies some tei-compliant markup but minimizes ui interference with the actual process of transcribing. the json-stored data allows any type of data to be stored in one database field instead of across several specific tables and can fairly easily interact with xml systems (gunther, schall, and wang 2016). one of the most significant impacts from the transcription center has been its contribution tounderstanding and leveraging motivations of theirvolunteers. dr. meghan ferriter, a former platform coordinator, has written and been interviewed extensively about the value of understanding volunteer motivations and customizing crowdsourcing activities to best address them (decker 2016; parilla and ferriter 2016; floyd 2017; ashenfelder 2016; ferriter 2016, 2014). the smithsonian outreach and engagement strategies are essential to a project’s success and must include communicating in a sincere and authentic way, volunteering information and content, and asking for help. the biodiversity heritage library has strong foundation of cooperative dialogue with its users and would likely be successful in managing a transcription project in a similar way. learning from the contributions and experience of the smithsonian transcription center offers valuable insight into data management and volunteer engagement practices. the transcription center, however, exclusively serves smithsonian institutions, which is incompatible with the larger organizational structure of bhl. the zooniverse the largest crowdsourcing program, the zooniverse14, redesigned the traditional model for document-based transcription projects. the zooniverse draws from a long legacy of citizen 14https://www.zooniverse.org. 67 mika et al. – crowdsourcing natural history archives science and co-creation projects from across scientific disciplines. what began as a web application to invite members of the public to classify and describe the universe’s photographed galaxies15 turned into the world’s largest platform for “people powered research”. the zooniverse has largely been focused on digital datasets that require hundreds of thousands if not millions of hours to investigate and classify by leveraging human intuitions and pattern finding abilities. with transcriptions, the image of a handwritten document becomes digital data and the zooniverse can use a similar model to apply and extract classifications, metadata structures, and bodies of text from manuscript collections. the zooniverse’s project builder and scribe (beaudoin 2015) utility that was developed in partnership with the new york public library both rely on the concept of microtasking to break up labor intensive transcriptions that require high levels of intelligence and concentration. by splitting tasks into more manageable chunks with varying degrees of difficulty, citizen scientists can engage with the project at whatever level they desire. breaking up the tasks also improves the data quality by mitigating against user fatigue and boredom. in order to achieve varying project goals successfully, zooniverse has developed three similar systems that each combine some degree of a mark, transcribe, and verify workflow (zooniverse 2015; snakeweight 2016; simpson, page, and roure 2014). instead of inviting volunteers to type complete page transcriptions into a text box, they break up the process into three separate tasks. page words and lines must first be marked by users to identify text locations in the image to maintain the author’s explicit layout and formatting choices; marked sections are then transcribed, which transforms the information into machine readable content and preserves the relationship between pixels and text; and transcribed text must finally be verified for 15https://www.galaxyzoo.org/. quality control. output data can be harvested raw (from each task) or algorithmically aggregated (from the whole set of mark and transcribe tasks for a given image) along with the level of the zooniverse’s confidence in the accuracy of the transcription. the output data is structured similarly to the transcription center’s (json), but does not depend on a specific database management system to manage data between the transcription platform and the collection repository. fromthepage fromthepage16 is a lightweight, open source, collaborative transcription platform. it’s defining feature is its use of wiki style markup to link references and subjects within texts to dynamically index terms. the design is optimized for archives projects, is a simple tool and can be deployed quickly. it has a very clean interface for viewing, transcribing, and coding people, places, and subjects across a collection of documents (figure 7). since 2005, fromthepage has hosted projects across the humanities and natural sciences and its creators ben and sara brumfield have become important community members that write, blog, and present extensively on their work and the larger themes or trends within the field of manuscript transcription and collaborative digitization. indeed, fromthepage has built a very good reputation among librarians, archivists, and other project managers over the last decade for its creative implementation of the wiki-like annotation index (lawson 2012). several natural history and botanic institutions have selected fromthepage to transcribe archival manuscript content including the san diego museum of natural history, university of california, berkeley’s museum of vertebrate zoology, and recently, the new york botanical garden’s luesther t. mertz library. each have identified the necessity of transcribing field notes to support research in the natural sciences by locating information about the 16https://fromthepage.com/. 68 figure 7. image of the university of california, berkeley’s museum of vertebrate zoology collection of joseph grinnell’s field notes that displays the item’s image and transcription with mediawiki encoded metadata tags for people, locations, dates, organizations, and could potentially be used to tag common and latin scientific names. https://www.fromthepage.com/display/display_page?page_id=4254. figure 8. image of the museum of comparative zoology’s collection of ornithologist william brewster’s digivol transcription project that displays the platform’s ability to add scientific names identified in the image as structured data. http://volunteer.ala.org.au/task/show/17794890. 69 mika et al. – crowdsourcing natural history archives quantity of species in an area, how often they were observed, and physical attributes of specimen. the san diego museum of natural history has digitized the field notes of noted herpetologist laurence m. klauber and commenced their transcription project in 2010. by march 2014 10,000 subjects had been identified for classification and volunteers made 24,000 page edits and 42,000 links between individual observations, species names, and personal names (brumfield 2014). similarly, the museum of vertebrate zoology at berkeley transcribed field notes of joseph grinnell and preserve his unique method of recording field observations. grinnell and other mvz scientists recorded observation notes with a particular philosophy (grinnell 1910) and using a precise structure that does not vary across creators, and, therefore, fromthepage’s tagging and indexing feature can be leveraged with tremendous success to create structured datasets for the collections. the new york botanical garden is also in the first steps of developing a transcription project for the john torrey papers that includes correspondence, manuscripts, notes, and botanical illustrations. digivol digivol17 was built by the australian museum as an atlas of living australia project and is designed primarily for transcription of biodiversity-related materials (kearney and wallis 2015). it combines a simple and attractive viewing and transcription interface with tools for extracting specimen data from items. there is no easy process for marking up text, but the platform features a form that invites volunteers to enter scientific names of specimens with dates and locations of their collection or observation. this generates a csv document that retains valuable information in a structured format. digivol is available for use by any institution, and simply requires establishment of a free account. project managers are given administrative privileges that enable uploading 17http://volunteer.ala.org.au/. of content, creation of custom guidelines and tutorials for each project, review and validation of the work of transcribers, and export of transcription files in csv format. digivol is nicely designed to facilitate communication between volunteers and project administrators. each transcription page has a comment box so that volunteers can make notes or ask questions about that page. there is a forum on which transcribers can raise topics and post comments or questions about any field notebook, and these are visible to administrators as well as other volunteers. managers can proactively use the forum to call attention to challenging pages, or address commonly made errors. volunteers can also email project administrators directly with queries or suggestions. the ernst mayr library of the museum of comparative zoology has been using digivol since 2014 for transcription of the field notes of william brewster, an amateur ornithologist of the late 19th and early 20th centuries. both digivol and fromthepage were initially used by the library in 2014-2015 to transcribe brewster materials for the purposeful gaming project. to date, over 6300 pages have been transcribed via digivol, and the resulting exports will form part of the initial load of transcription files to bhl. species names and other structured occurrence data, captured apart from the transcriptions and exported as a separate csv file, will be retained for possible future use in bhl pending development of a new data model (figure 8). early transcription platforms were generally tied to specific manuscript collections of significant value or at high risk for long term preservation. these projects depended on curation and were optimized for smaller programs for which significant manual data and community management practices were plausible. workflows for these projects also often require significant manual processes including the selection and curation of collections, potentially digitizing content if 70 mika et al. – crowdsourcing natural history archives necessary, uploading image files to external transcription platform, conducting outreach activities to recruit volunteers, writing collections specific tutorials, answering questions regarding transcriptions and data entry processes, validating completed works and performing qa, exporting text and associated metadata, transforming text and data into schema and formats optimal for a specific digital repository, uploading transcriptions to repository, ensuring that access is appropriate, and engaging in general troubleshooting. these models do not typically scale well, even with the expertise of a dedicated community or program manager. outputs for document-based projects are largely simple plain text documents that improve readability (by humans and machines) and do not structure data outputs. conversely, record-based programs enrich collections with structured data and metadata, but are limited in their capacity to accurately represent works in their entirety. the time investment required for collections-based transcription programs and their technological limitations prohibit or limit their use for large digital collections like the biodiversity heritage library. as the field of manuscript transcriptions developed and crowdsourcing proved to be a reliable method for generating machine readable text, platforms evolved to meet the demands of larger digital collections. the zooniverse and digivol incorporated citizen science practices and ideas and fromthepage and the smithsonian transcription center included markup and tagging features to structure text and data to improve access and enrich collections metadata. discussion the web is a co-creative digital experience, and glam organizations need to be prepared to engage with users’ knowledge and experience to build and augment online content (terras 2016). successful crowdsourcing platforms do more than invite users to donate time and expertise to specific projects; they cultivate digital environments that encourage open access and the clear and open transfer of ideas. in digital libraries, collections are not hidden behind glass exhibition cases but are living texts and documents that operate as part of a wider ecosystem of knowledge sharing and cocreation. bhl plans to design a workflow for its partners to generate high quality transcriptions which can be deployed with relative ease. field notebook and manuscript transcriptions need to fit into the larger bhl objective of producing and making available large scale datasets for users and researchers to study and manipulate. while generating only full-text transcriptions will improve discoverability, manuscript items will remain isolated from the other literature collections. one way to link knowledge produced in field notes and information derived from books and journals in bhl is to extend the database by adding tables and relationships for common access points including common species names, locations, dates, identified people and organizations, and events (figure 6) (studer and rinaldo 2014). item level catalog records for monographs and serials contain publication information that enhances access and provides context. special collections content, however, often does not include basic item level contextual information like creator (author), creation (publication) date, subject headings, location/geographic information, or language. as it designs a transcription and data collection program, bhl can leverage lessons learned from established crowdsourcing projects in the cultural heritage sector as well as citizen science initiatives that are common in biodiversity and natural history domains (table 1). as a digital library focused on scientific inquiry, bhl is well situated to join the “collections as data” movement by developing programs, including transcription and data collection, to support computational analyses and distant reading of texts (zwaard 2017). in order to truly capitalize on the power of digitization, future versions of bhl will aim to support text and data mining of its content and collections metadata. 71 table 1: comparison of transcription tools treated in the text. platform projects advantages disadvantages scripto1 papers of the war department2; anne whitney papers3 can be integrated with content management systems, including omeka, wordpress, and drupal plugins collection-based; does not scale up particularly well transcribe bentham’s transcription desk4 transcribe bentham5 designed as a research project; reveals much about project design collection based; requires significant amount of customization and development smithsonian transcription center6 william m. mann field notes fiji and british solomon islands, 191519167; albert spear hitchcock field notes8 extremely successful institutional program with significant number of volunteers; potential for connecting data from special collections and archives with smithsonian specimen collections only available for use by smithsonian institution organizations zooniverse9 anno.tate10 shakespeare’s world11; beyond words12 largest volunteer base; deeply connected to the citizen science universe; scales up well mark and transcribe workflow interrupts user interface connection to document; aggregation method requires multiple volunteers to complete tasks; does not provide immediate user access to output fromthepage13 joseph grinnell’s field notes14; c.s. pierce manuscripts15 mediawiki tags are flexible and simple to use; cleanest user interface; flexible; open source development; supports ocr corrections in addition to manuscript transcriptions external platform requires some development to improve interoperability digivol16 william brewster ornithological journals17 includes optional forms for entering structured occurrence data optimized for natural history collections occurrence data must be entered in addition to transcription 1http://scripto.org/. 2http://wardepartmentpapers.org/. 3http://omeka.wellesley.edu/whitneytranscribe/home. 4http://www.transcribe-bentham.da.ulcc.ac.uk/td/transcribe_bentham. 5http://www.ucl.ac.uk/transcribe-bentham. 6https://transcription.si.edu/. 7https://transcription.si.edu/project/9708. 8https://transcription.si.edu/project/10894. 9https://www.zooniverse.org/. 10https://anno.tate.org.uk. 11https://www.shakespearesworld.org. 12http://beyondwords.labs.loc.gov. 13https://fromthepage.com/. 14https://fromthepage.com/cfidler/transcribing-the-field-notes-of-the-museum-of-vertebrate-zoology. 15https://fromthepage.com/jeffdown1/c-s-peirce-manuscripts. 16https://digivol.ala.org.au/. 17https://digivol.ala.org.au/institution/index/11740375. 72 mika et al. – crowdsourcing natural history archives digitizing items in archives and special collections and adding them to online repositories promised dramatically enhanced access but did not deliver. images of content are still described at collection levels and are perhaps not as useable as once imagined. transcribing textual information contained in these images facilitates indexing and searching. texts can be mined to enrich metadata attributes, and context can be applied to records to better connect items according to content. transcriptions, however, are also in danger of disappearing into digital repositories. without some kind of imposed intellectual framework, digitized items are lost to the “dank cellar of electronic texts” (shillingsburg 2006), which bhl is working to avoid by developing crowdsourcing initiatives in concert with redesigning the portal’s metadata framework, image delivery system, and taxonomic backbone. literature cited ashenfelder, m. 2016. ’volun-peers’ help liberate smithsonian digital collections. the signal blog. december.18 bartle, j. 2014. the letters of anne whitney: using archives in digital scholarship. feminist collections 35 (3/4). beaudoin, p. 2015. scribe: toward a general framework for community transcription. nypl labs. november.19 blaser, l. 2014. old weather: approaching collections from a different angle. in crowdsourcing our cultural heritage, edited by mia ridge, 45–55. ashgate publishing ltd., surrey. bonney, r, c.b. cooper, j. dickinson, s. kelling, t. phillips, k.v. rosenberg, and j. shirk. 2009. citizen science: a developing tool for expanding science knowledge and scientific literacy. bioscience 11:977–984. doi:10.1525/bio.2009.59.11.9. 18http://blogs.loc.gov/thesignal/2016/12/volun-peers-help-liberate-smithsoniandigital-collections/. 19https://www.nypl.org/blog/2015/11/23/scribe-framework-communitytranscription. 20https://www.biodiversitylibrary.org/bibliography/77525. brabham, d.c. 2013. crowdsourcing. mit press, cambridge. brewster, w. 1892. journals of william brewster, 1871-1919. museum of comparative zoology.20 brumfield, b. 2012. what does it mean to ’support tei’ for manuscript transcription? collaborative manuscript transcription. november 10.21 brumfield, b. 2014. wikilinks in fromthepage. collaborative manuscript transcription. march 14.22 causer, t., and m. terras. 2014. crowdsourcing bentham: beyond the traditional boundaries of academic history. international journal of humanities and arts computing 8:46–64. doi:10.3366/ijhac.2014.0119. decker, j. 2016. exploring the smithsonian institution transcription center [special issue]. collections: a journal for museum and archives professionals 12(2). rowman/littlefield. deveer, j.m., c.a. rinaldo, and l. ford. 2013. primary source material in science: the importance of archival field notes.23 dickinson, j.l., and r. bonney, eds. 2012. citizen science: public participation in environmental research. cornell university press, ithaca. dickinson, j.l., j. shirk, d. bonter, r. bonney, r.l. crain, j. martin, t. phillips, and k. purcell. 2012. the current state of citizen science as a tool for ecological research and public engagement. frontiers in ecology and the environment 10(6). doi:10.1890/110236. ellis, s. 2014. a history of collaboration, a future in crowdsourcing: positive impact of cooperation on british librarianship. libri 64(1):1–10. doi:10.1515/libri-2014-0001. europe, council of. 2005. council of europe framework convention on the value of cultural heritage for society. council of europe treaty series no. 199. evans, m. 2007. archives of the people, by the people, for the people. american archivist 70(2):387–400. doi:10.17723/aarc.70.2.d157t6667g54536g. ferriter, m. 2014. growing to a community of volunpeers: communication and discovery. smithsonian institution archives. july.24 21http://manuscripttranscription.blogspot.com/2012/11/what-does-it-mean-tosupport-tei-for.html. 22https://manuscripttranscription.blogspot.com/2014/03/. 23http://escholarship.umassmed.edu/esciencesymposium/2013/posters/14. 24https://siarchives.si.edu/blog/growing-community-volunpeerscommunication-discovery. 73 mika et al. – crowdsourcing natural history archives ferriter, m. 2016. volunpeers: hashtag, identity, and collaborative engagement. meghan in motion blog. april.25 floyd, s. 2017. engaging with volunteers: smithsonian transcription center. fromthepage blog, april 2017.26 grinnell, j. 1910. the methods and uses of a research museum. popular science monthly 77:163–169. gunther, a, m. schall, and c.-h. wang. 2016. the creation and evolution of the transcription center: smithsonian institution’s digital volunteer platform. collections 12(2): 87–96. gwinn, n.e., and c.a. rinaldo. 2009. the biodiversity heritage library: sharing biodiversity literature with the world. ifla journal 35(1):25–34. hanse, d., j. gehring, p. schone, and m. reid. 2012. improving indexing efficiency and quality: comparing a-b-arbitrate and peer review. family history technology workshop, brigham young university, provo. holley, r. 2010. crowdsourcing: how and why libraries should do it. d-lib magazine 16 (3/4). doi: 10.1045/march2010-holley. kalfatovic, m., and c. rinaldo. 2016. enabling progress in global biodiversity research: the biodiversity heritage library. in libraries: enabling progress, proceedings of the eighth shanghai international library forum. shanghai scientific / technological literature press. pp. 406–418. kearney, n., and e. wallis. 2015. transcribing between the lines: crowd-sourcing historic data collection. mwa2015: museums and the web asia 2015.27 lawson, k.l. 2012. crowdsourcing transcription: fromthepage and scripto. chronicle of higher education.28 leon, s.m. 2014. build, analyze, and generalize: community transcription of the papers of the war department and the development of scripto. pp. 89-111 in crowdsourcing our cultural heritage (m. ridge, ed.). ashgate publishing ltd., surrey. mcclanahan, p. 2017. getting to know the bhl users. may.29 25http://meghaninmotion.com/2016/04/05/volunpeers-hashtag-identityengagement/. 26http://content.fromthepage.com/smithsonian_volunpeers/. 27http://mwa2015.museumsandtheweb.com/paper/transcribing-between-the-linescrowd-sourcing-historic-data-collection/. 28http://www.chronicle.com/blogs/profhacker/crowdsourcing-transcriptionfromthepage-and-scripto/38028. miller-rushing, a., r. primack, and r. bonney. 2012. the history of public participation in ecological research. pp. 285-290 in citizen science: public participation in environmental research (j.l. dickinson and r. bonney, eds.). cornell university press, ithaca. doi:10.1890/110278. musacchio, j.m. 2014. project narrative: anne whitney abroad, 1867-1868.30 parilla, l., and m. ferriter. 2016. social media and crowdsourced transcription of historical materials at the smithsonian institution: methods for strengthening community engagement and its tie to transcription output. american archivist 79(2). doi:10.17723/0360-9081-79.2.438. pilsk, s.c., m.a. person, j.m. deveer, j.f. furfey, and m.r. kalfatovic. 2010. the biodiversity heritage library: advancing metadata practices in a collaborative digital library. journal of library metadata10: 136–155. doi:10.1080/19386389.2010.506400. rinaldo, c.a., and j. smith. 2014. moving through time and culture with the biodiversity heritage library. pp. 95-108 in migrating heritage: experiences of cultural networks and cultural dialogue in europe (p. innocenti, ed.). ashgate publishing ltd., surrey. rose-sandler, t. 2012. the art of life: data mining and crowdsourcing the identification and description of natural history illustrations from the biodiversity heritage library. biodiversity heritage library.31 rose-sandler, t. 2015. purposeful gaming and bhl: engaging the public in improving and enhancing access to digital texts. biodiversity heritage library.32 saylor, n., and j. wolfe. 2011. experimenting with strategies for crowdsourcing manuscript transcription. research library issues: a quarterly report from arl, cni, and sparc.33 schmitt, w.l. 1962. palmer peninsula (antarctica) survey, 1962-1963: miscellaneous notes (2 of 4). series: sia ru007231.34 shillingsburg, p. 2006. from gutenberg to google: electronic representations of literary texts. cambridge university press, cambridge. 29https://ndsrbhl.wordpress.com/2017/05/17/getting-to-know-the-bhl-users/. 30http://www.w2ww.19thc-artworldwide.org/index.php/autumn14/musacchioproject-narrative. 31http://biodivlib.wikispaces.com/art+of+life. 32http://biodivlib.wikispaces.com/purposeful+gaming. 33https://eric.ed.gov/?q=ed527715&id=ed527715. 34https://www.biodiversitylibrary.org/bibliography/130464#/summary. 74 mika et al. – crowdsourcing natural history archives silberglied, r.e. 1965. field notes, mexico, julyaugust 1965. sia ru007316.35 simpson, r., k.r. page, and d. de roure. 2014. zooniverse: observing the world’s largest citizen science platform. pp. 1049-1054 in proceedings of the 23rd international conference on world wide web. association for computing machinery. doi:10.1145/2567948.2579215. snakeweight. 2016. ’what’s up with those grey dots?’ you ask. shakespeare’s world blog. february.36 studer, m., and c.a. rinaldo. 2014. from historical field notes to mobile field guides: the encyclopedia of life and the biodiversity heritage library team up to connect biodiversity-related content across the centuries to support today’s ecological research and education needs. ecological society of america.37 terras, m. 2016. crowdsourcing in the digital humanities. pp. 420-439 in a new companion to digital humanities. wiley-blackwell, chichester. theobald, e.j., a.k. ettinger, h.k. burgess, l.b. debey, n.r. schmidt, h.e. froehlich, c. wagner, j. hillerislambers, j. tewksbury, m.a. harsch, and j.k. parrish. 2015. global change and local solutions: tapping the unrealized potential of citizen science for biodiversity research. biological conservation 181:236–244. doi: 10.1016/j.biocon.2014.10.021. zooniverse. 2015. one line at a time: a new approach to transcription and art history. zooniverse. september.38 zwaard, k. 2017. collections as data and national digital initiatives. library of congress. august. 39 35https://www.biodiversitylibrary.org/bibliography/95422#/summary. 36https://blog.shakespearesworld.org/2016/02/24/whats-up-with-those-grey-dotsyou-ask/. 37https://eco.confex.com/eco/2014/webprogram/paper47651.html. 38https://blog.zooniverse.org/2015/09/01/one-line-at-a-time-a-new-approach-totranscription-and-art-history/. 39https://blogs.loc.gov/thesignal/2017/08/collections-as-data-and-national-digitalinitiatives/. 75 bhl_crowdsourcingtranscription_final2 bhl_crowdsourcingtranscription_figs3and4 bhl_crowdsourcingtranscription_figs5and6 bhl_crowdsourcingtranscription_figs7and8 bhl_crowdsourcingtranscription_table biodiversity informatics, 18, 2024, pp. 78-100 78 enhancing ecological education: utilizing agent-based modeling to simplify the impacts of deforestation on amphibians martín otálora-löw1*, j. nicolás urbina-cardona1, mauricio gonzález-méndez1 1 pontificia universidad javeriana, facultad de estudios ambientales y rurales, departamento de ecología y territorio. carrera 7 n 40 – 62, bogotá, colombia. *correspondence: m_otalora@javeriana.edu.co abstract. understanding how deforestation and changes in habitat boundaries affect biodiversity is essential for developing conservation solutions. these topics are central to biology and ecology programs, where students learn to apply their knowledge in real-world conservation efforts. higher education plays a crucial role in strengthening this understanding, particularly in life sciences programs. given the complexity of ecological processes in altered landscapes, agent-based modeling provides an interactive and engaging way to simplify and visualize the effects of land use changes. in this study, we integrate amazonian anurans, highly sensitive to temperature and humidity fluctuations, with an agent-based model to simulate the impacts of deforestation, habitat restoration, and land abandonment on species survival and movement. their ectothermic nature and dependence on pulmocutaneous respiration make them especially vulnerable to the drier and more variable conditions caused by deforestation. integrating this model into conservation biology courses has enhanced learning by encouraging independent exploration, both in and out of the classroom. this tool, an agent-based model, is particularly suited for university-level ecology and conservation courses, and can also serve as an effective awareness tool in environmental education and decision-making workshops, highlighting the negative effects of human-made habitat changes on biodiversity. key words: agent-based modeling, landscape transformation, education classroom, edge effect. introduction habitat loss is one of the leading causes of species extinctions worldwide (arroyo-rodríguez et al., 2020; didham et al., 2012; tscharntke et al., 2012). when native forest is cleared for agriculture or other human activities, the newly formed forest edges expose species to novel environmental conditions for which they may not be adapted (ries et al., 2004). these edge effects result in an ecotone between agricultural land and the remaining native forest patches, characterized by abrupt changes in temperature, humidity, and light that can significantly alter species distributions and abundances (harper et al. 2005). the severity of edge effects depends on several factors, including the age and contrast of the forest edge—determined by the degree of structural difference between the native forest and adjacent human-modified areas—as well as the shape and size of the remaining forest patch (ries et al., 2004). these complex interactions require careful study due to the high number of spatiotemporal variables involved. amphibians, particularly small-bodied species, are especially vulnerable to habitat loss and the creation of forest edges, largely due to their high dehydration rates and limited dispersal ability (pfeifer et al., 2017; schneider-maunoury et al., 2016). in fragmented landscapes, larger-bodied amphibians may be able to disperse between forest patches through anthropogenic matrices such as pastures, while smaller forest specialists often become isolated in remnant patches (pfeifer et al., 2021; zabala-forero and urbina-cardona 2021). over time, natural regeneration of abandoned lands or active ecological restoration can help restore amphibian populations, but movement between patches remains limited for species with low dispersal capacity (díaz-garcía et al., 2017; hernández-ordóñez et al., 2015). the impact of anthropogenic landscape transformation on biodiversity are inherently complex, particularly because they involve numerous environmental, structural, and biotic interactions that are difficult to fully convey in classroom settings without field-based experiential learning (tscharntke et al., 2012). furthermore, scientific research on this topic is often published in specialized journals, which may be inaccessible or not easily interpreted by non-specialists, such as students, policymakers, or rural people (ferraz et al., 2021; redford et al., 2012). this underscores the need for simplified, interactive tools that synthesize these complex processes and foster otálora-löw et al. – enhancing ecological education 79 collaborative learning environments. agent-based models (abms) provide a valuable approach to illustrating environmental interactions that are difficult to observe directly in nature, making it an effective tool for communicating complex ecological concepts in educational settings (carey and gougis 2017; farrell and cayelan, 2018). abms are computational simulations that represent individual agents (animals, plants, or humans) within an environment, allowing them to interact according to defined behavioral rules. abms help simulate complex systems by modeling the interactions between agents and their environments. for example, abms have been used to study the impact of agricultural practices on biodiversity by simulating the behavior of species in response to changes in land use, such as deforestation or reforestation efforts (chopin et al., 2019), and are increasingly integrated into classroom e-learning to reinforce concepts that traditionally require fieldwork (brewer, 2006; murphy et al., 2020). here we developed an interactive, user-friendly abm to illustrate the influence of edge effects on the distribution and abundance of native amphibian species in the amazon ecosystem of colombia. we delineated specific behavioral patterns and energy consumption of species in pasture, forest edge, and interior habitats, organizing this information within a user interface designed for educational purposes. this model simplifies the complexities of species movement and reproduction in response to deforestation, restoration, and habitat fragmentation, providing a visual and engaging way to communicate ecological concepts to students. methods agent-based model description the model was made in the programming language and integrated development environment netlogo (tisue and wilensky, 2004). the model was described in depth using the overview, design concepts, and detail (odd) protocol proposed by grimm (2010), as it has been used to organize and explain models so that they are understandable to the public and replicable by other researchers. the full odd can be found in the appendix 1. the user configures the transition parameters in land use type of pasture, rainforest edge, or interior from a panel with icons that can be slid to set the scenario (figure 1). in each tick, as the simulation runs, results can be observed on a board with patches representing different habitats (represented by different colored pixels) and on which the different types of anurans are scattered and reproduced. in which a tick is a count of times the code has been read, the figure 1. netlogo web model with description of the tabs and buttons. otálora-löw et al. – enhancing ecological education 80 processes have been accomplished and don’t represent real-time. however, processes like reproduction are linked to the ticks to show that reproduction is a discrete process. a patch is a stationary agent, the space in which the processes happen, and it can have its variables. this area does not have real measurements. classification of anuran species by forest edge response the mobile agents are the anura (frogs) and were configured based on the field abundances of the amphibian species registered by palomino-cuellar (2019) in 90 linear transects of 15m in length along a distance gradient from 120 m into the pastures and 400 m into the amazon rainforest interior. the field phase for amphibian sampling was carried out between november 2018 and february 2019 for a total effort of 240 person-hours (for more detail of the study area look for appendix 2). within the amphibian assemblage, species were identified that showed changes in their abundance between the pasture, forest edge, and the forest interior habitats and were classified into three types of response to the edge (schneider-maunoury et al., 2016), as follows: 1) mathiasson’s treefrog (dendropsophus mathiassoni) affine for the pasture habitat; 2) tiny tree toad (amazophrynella minuta), and variable robber frog (pristimantis variabilis) affine for the forest edge habitat; and 3) peters’ dwarf frog (engystomops petersi), canelos robber frog (p. acuminatus), and chirping robber frog (p. conspicillatus) affine for the forest interior habitat. two and three species represented the border edge (from 0 to 50m) and forest interior habitat (beyond 100m from the forest edge), respectively, as each species has a small number of individuals, so we grouped them to maximize the information on environmental variables and functional (morphological and natural history) traits per habitat. there were two criteria to group species: 1) similar leg/weight ratio, 2) similar reproductive behavior. agents, patches, and variables the board was divided into three habitats for the initial model: pastures, rainforest edge, and rainforest interior. each of these habitats has initial values from the average values of the environmental variables (temperature, relative humidity, and canopy cover) measured at the site where each of the individuals of the species was found by palomino-cuellar (2019) (table 1). these three habitats change according to a rate set with a slider by the user, allowing to see faster or slower processes of spatio-temporal change. the user can determine in the model the changes in land use and land cover adjacent to the remaining native forest with transitions such as ecological restoration, deforestation, or natural regeneration after land abandonment. every time a patch changes, it acquires a random new number for environmental temperature, canopy cover, and relative humidity within the range of the neighbors and the values observed in the field for each variable in each habitat (table 1) each amphibian species was assigned a categorical value of dispersal capacity, reproduction, and a limited supply of energy to accomplish survival. in this sense, each agent in the model can die from a differential lack of energy (exhaustion), and so, depending on natural history traits, like parental care, seasonality in reproduction, number of eggs, type of nest, between other factors (vitt and caldwell 2013), a species´ group could consume more or less energy when reproducing (table 2). specifically, the relative energy consumption in reproduction is used in table 1. environmental variables that characterize the habitat of the anuran species group from temperature, relative humidity, and canopy cover. these data were taken from the field measurements made by palomino-cuellar (2019). species environmental dimension minimum maximum average dendropsophus mathiassoni(15) temperature 22.70 28.50 26.54 relative humidity 84.50 94.80 90.36 canopy 0.00 7.54 0.90 pristimantis variabilis (1), amazophrynella minuta (3) temperature 26.70 33.20 30.22 relative humidity 67.20 88.20 75.70 canopy 81.90 88.14 85.97 engystomops petersi (2), p. conspicillatus (1), p. acuminatus (1) temperature 24.40 28.20 26.26 relative humidity 91.20 100.00 94.00 canopy 82.42 88.66 85.23 otálora-löw et al. – enhancing ecological education 81 the model as the rate at which a new amphibian is hatched per species group. the relative energy consumption values assigned to each species group during reproduction are based on a combination of natural history traits, such as reproductive seasonality, number of eggs, type of nest, and parental care, which are known to influence energy expenditure in anurans (vitt and caldwell, 2013). for example, species with high reproductive seasonality and larger clutch sizes (e.g., e. petersi and a. minuta) are assigned higher energy consumption values (categories 4-5), reflecting their significant investment in reproduction within a constrained time frame. in contrast, species with lower reproductive frequency or smaller clutch sizes (e.g., p. conspicillatus) have moderate energy consumption values (categories 2-3). these categorizations allow the model to simulate differential energy expenditure based on species-specific reproductive strategies, which are well-documented in herpetological literature (toledo et al., 2009). the leg-weight ratio was calculated as and indicates how capable is a frog to move their body, meaning that a lower number shows that the legs are longer or the frog´s body is smaller. then, all the results were standardized and multiplied by a value of 5 would reflect the maximum ability of an individual to explore his environment at a maximum of 5 pixels around the game board. for the model, the average between the individuals of the species in the same group was used (table 2). processes within the model firstly, this configuration assumes that changes in the environment´s structure will affect reproduction chances (cayuela et al., 2014). likewise, we assumed that the edge effect acted in a set boundary, when in reality it´s impact varies within variables and distance. at the beginning of the simulation, the three groups of species (pasture, forest edge, or forest interior affine frogs) are distributed within their habitat at random (pasture, forest edge, and interior) (figure 2). still, they may become locally extinct depending on changes in the parameters made by the user. the user can modify the change rate with the provided sliders, adjusting levels of deforestation and restoration (figure 3). all agents are assumed to be adults with reproductive capacity. when the user sets up the simulation, the patches acquire established values for the variables (table 1). all anuran individuals have a base movement, meaning that they are never still and always consume at least 1 unit of energy per tick. subsequently, when the simulation starts, the agents check their surroundings in a radius ranging from 1 to 5 pixels depending on their leg-weight ratio (table 2), the individuals inspect if the variables are within their suitable range to a greater or lesser number of pixels in the surrounding area. the frog will “run” facing the nearest suitable patch if they’re not within their range. the leg-weight ratio and available energy determine running, and thus, if energy is 0, the frog dies and disappears from the board (figure 4). table 2. morphological and reproductive values per anuran species inhabiting the rainforest in the caquetá department. the categorical value of relative energy consumption in reproduction ranges from 1 to 5, in which one would represent low energy consumption a species number of individuals relative energy consumption in reproduction (1-5) reproduction snoutvent length legweight ratio movement in model response group to forest edge dendropsophus mathiassoni 15 3 21.00 1.5 3.4 pasture amazophrynella minuta 3 4 seasonal 18.04 1.3 3.9 edge pristimantis variabilis 1 2 16.47 0.9 5.5 edge engystomops petersi 2 5 seasonal 27.37 4.6 1.1 forest interior pristimantis acuminatus 1 3 23.15 3.2 1.5 forest interior p. conspicillatus 1 2 34.62 5.6 0.9 forest interior otálora-löw et al. – enhancing ecological education 82 figure 2. at setup, the code first creates the environment with each variable on each habitat. then, it creates the agents, locates them in the previously created environment, and gives values to their variables. figure 3. illustration of the initial default settings of the agent-based model. each habitat is proportional at set up, having 30 x 60 pixels. in brown, we have the pasture; in yellow, we have the forest edge; in green, we have the forest interior. anuran individuals have the shape of frogs, representing one of the three species groups mentioned: yellow species are affine for pasture; blue is affine for forest edge; and green is affine for forest interior. at go, patches would turn orange to represent natural regeneration, blue, to represent restoration or brown, to represent deforestation. otálora-löw et al. – enhancing ecological education 83 in this sense, the values used in the simulation for energy thresholds and their corresponding behaviors (sleep, dispersal, reproduction, or death) are based on physiological principles observed in anurans. as ectotherms, anurans’ energy consumption is highly influenced by environmental conditions such as temperature and humidity (feder and burggren, 1992; hillman et al., 2009;). during periods of low energy, anurans reduce their activity to conserve resources, which is reflected in the model when energy levels are between 1 and 50. this is consistent with studies that show anurans become inactive under adverse conditions to preserve energy (wells, 2007). reproduction is highly energy-intensive, requiring substantial investment in behaviors such as courtship and gamete production (toledo et al., 2009), justifying the need for energy levels above 80 to reproduce in the simulation. additionally, dispersal capacity is linked to body size and the leg-weight ratio, which influences the distance anurans can move in search of optimal conditions (marsh and trenham, 2001). these values ensure that the simulation accurately reflects the physiological and ecological realities of anuran species. the described model can be found on the web.1 on the contrary, if the area surrounding an anuran individual is optimal, it will wait on the set reproduction time and acquire energy in the meantime (figure 4). all anuran individuals have a 50% chance to reproduce, which a new individual on the board evidences. it is also assumed that half of the population is female and must have at least 80 units of energy to generate a new individual; the new anuran individual hatches with a third of the parent’s energy 1 https://modelingcommons.org/browse/one_model/7476#model_tabs_ browse_nlw. figure 4. flowchart of the decision-making of anuran species in the model. the cloud represents a connection with figure 5. otálora-löw et al. – enhancing ecological education 84 and a random direction. once the anuran individuals have responded, the environment is altered, provoking a change in the proportion of habitat in the board. then, the model randomizes the environmental conditions at the forest edge according to the range of values in the variables observed in the field (figure 5). depending on the user, three land use/land cover transition scenarios can be set to dominate the landscape transformation processes in each scenario, all of which have been reported to affect the amphibian populations: deforestation dominance (agudelo-hz et al., 2019), natural regeneration after land abandonment dominance (herrera-montes and brokaw 2010; hernández-ordóñez et al., 2015), or ecological restoration dominance (brodman et al., 2006; díaz-garcía et al., 2017). application of the abm in a classroom a pilot study was conducted to probe changes in student responses during the three stages of evaluation: (a) before the module classes, (b) after teaching students the concepts and case studies in the 4-hour class on edge effects, and (c) after using and interacting with the agent-based model during a workshop guided by the model developer. the same person figure 5. flowchart of the patches change by neighbors. the change generates new variables in the environment linked with the anura decision-making, represented through the cloud in the top right. taught all the classes. during each stage, students were asked to answer the same four questions (two qualitative and two quantitative, using the likert scale (matas, 2018) about the intensity of the edge effect on different native forest patch sizes and shapes, and under different vegetation cover bordering the remnant native forest (appendix 3). the abm tool was applied during the first semester of 2022 to undergraduate students of ecology at pontificia universidad javeriana in bogotá (colombia), in the class of conservation ecology (7th semester). it was also taught to first-semester students of the master’s degree in conservation and use of biodiversity of the same university, in the class of conservation biology, for a total of 24 respondents that participated voluntarily of the evaluation during each of the class’s activities. before the abm tool was shown and before the class on edge effects, all students received 4 hours of class on the effects of habitat loss and fragmentation on biodiversity. to familiarize the students with the tool beforehand, an explanatory video in spanish was given2. this video is in spanish and it seeks to aid in the use of the abm tool. 2 https://youtu.be/kosyfkt5pzq. otálora-löw et al. – enhancing ecological education 85 students receive a 30-minute class on the use of the abm tool to explain edge effects. it was expected that as students attended the class and interacted with the model, their responses (the value of the impact of edge effects) changed. once the students had seen the agents’ behavior under different configurations in the netllogo platform1, they were asked for feedback on their use experience and suggestions for improvement of the graphical user interaction interface. from the students’ answers to questions 3 and 4 (appendix 3), two multivariate response variables were obtained to evaluate changes between the three stages of evaluation in (a) the value (likert scale from 1 to 5) of the impact of the edge effect on native forest fragments of different shape and size (in question 3), and (b) the value (likert scale from 1 to 5) of the impact of the edge effect on native forest fragments of different shape and adjacent land use (in question 4). given the multidimensional nature of the student responses, which involved both qualitative and quantitative measures, we required a statistical method capable of assessing multivariate differences across multiple evaluation stages. the use of euclidean distance matrices allowed us to compute dissimilarities between student responses across these different stages. permanova was chosen as it is a non-parametric method suitable for analyzing multivariate data without the stringent assumptions of normality or homoscedasticity. this is particularly crucial in educational studies, where data often violate these assumptions. furthermore, permanova, when combined with type iii sum of squares, allows for the partitioning of variation in the responses to specific factors—in this case, the evaluation stages—while accommodating unbalanced designs and complex, hierarchical data structures, as was present in our study (anderson et al., 2008). using 9,999 permutations enhances the robustness of the results, ensuring that statistical inferences are based on data-derived distributions rather than theoretical ones, thus aligning with the exploratory nature of our study. the ability of permanova to handle both qualitative and quantitative shifts in response patterns made it an ideal tool for evaluating changes in students’ perceptions of the edge effect after each interaction phase with the agent-based model. the experimental design consisted of the evaluation stage as a fixed factor (with three levels: before and after the edge effects module and after the interactive experience with the model). given that individuals of different academic levels and genders may have varying levels of prior ecological knowledge, shaping their perceptions and attitudes towards biodiversity (moreno-rubiano et al., 2023; vergara-ríos et al., 2021), we conducted separate analyses for men and women, as well as for undergraduate and graduate students. additionally, gender-based differences in educational environments have been shown to influence learning styles and concept appropriation (matas, 2018). however, due to the small sample size, the degrees of freedom were insufficient to evaluate interactions between gender and academic level in a two-way permanova. consequently, we opted to analyze men and women, as well as undergraduate and graduate students, separately. this approach allowed us to consider potential differences in learning outcomes across these demographic factors, despite the limitations imposed by the sample size. we compared the levels with statistical differences within the factors with a posterior pairwise comparison with the t-statistic based on 9999 permutations (vergara-ríos et al. 2021). results qualitative description of user experiences the 14 students who had seen the conservation ecology class in the undergraduate program in ecology were more reserved when giving feedback during the workshop compared to the ten students from the master’s degree conservation biology class. after the workshop, some students wanted more details about the model, so an additional demonstration was given. on the contrary, the interaction with the students of the master’s program was more dynamic, with more active participation from questions, comments, and suggestions during the presentation of the model in the classroom. graduate students were most interested in why things were happening in the model instead of waiting for the teacher to ask them; they also asked to generate more scenarios in which they could see land cover and use transitions. both approaches of the students exemplify the usefulness of the model, as different attitudes were evident in the questioning of the three scenarios presented (pastures created after the deforestation process, land abandonment, and ecological restoration) and the whole usage of the model as a teaching tool. otálora-löw et al. – enhancing ecological education 86 model behavior as a function of landscape transformation scenarios deforestation scenario.—when the students favored deforestation as a dominant process in the landscape (figure 6), they were able to see how fast the pasture overtakes the other landcover type (rainforest edge and interior), forcing the frogs that inhabited the forest edge and interior to disperse into the remaining forest, and finally visualizing the extinction of rainforest associated species. at the same time, while the population size increases for the pasture affine species. the user can visualize that the yellow species that inhabit the pasture reproduce very quickly, but no large amounts are scattered across the pixels; this happens as they are programmed to move to the best patch possible, but as they consider that the ones near them are optimal, and there is no anuran density limit for the patches, they agglomerate at the left of the screen leaving most of the pasture habitat without frogs (figure 7). figure 6. model setup for deforestation scenario. the slider is at 100 for deforestation, while ecological restoration is at 0. figure 7. end of the simulation of deforestation. pastures created after deforestation overtook the other habitats and extirpated the other two groups of native forest affine anuran species. otálora-löw et al. – enhancing ecological education 87 land abandonment scenario.—as the class progressed, students asked what happened to the anurans when there was no deforestation and no restoration, just the abandonment of the pasture to make way for natural regeneration (figure 8). land abandonment as a dominant process in the landscape (figure 6) shows how the natural regeneration (orange pixels) grows, overtaking the pasture and extirpating the group 1 of anuran species affine for pasture. in contrast, while both forest edge and interior affine species overgrow in abundance (figure 9). they asked why the interior was growing, and the explanation was that this happens because natural regeneration serves as a buffer, giving extra protection to the interior. figure 8. model setup for land abandonment and natural regeneration. all sliders are at 0. figure 9. end of the land abandonment simulation. the pasture was overtaken by natural regeneration, in color orange, eliminating group 1 of pasture affine frogs while native forest affine anurans groups 2 and 3 grew over time. land abandonment = orange pixels. otálora-löw et al. – enhancing ecological education 88 restoration scenario.—when looking at ecological restoration as a dominant process in the landscape (figure 10), students saw restoration (dark green pixels) emerge and forest interior borders move quickly to the left. in contrast native forest edge affine frogs and forest interior frogs usually reproduce. when restoration overtakes the whole pasture, the pasture affine frog begins to die; eventually, there are only blue and green pixels with green and blue anuran species (figure 11). students’ answers between stages of evaluation for question 3 about the degree of influence of the edge effect on the abundance of a specialist native forest species on the interaction between the size of the native forest and its shape (question 3 appendix 3), there were no statistical differences in the responses between stages of evaluation for undergraduate students (pseudo-f = 0.462; p-perm = 0.782). for graduate students, there were statistical differences between evaluation stages (pseudo-f = figure 10. model setup for restoration. the restoration slider is set to 100, and deforestation is set to 0. figure 11. end of the ecological restoration simulation. the pasture was overtaken by restoration, dark blue pixels, eliminating group 1 of frogs, and the forest interior grew. groups 2 and 3 of native forest affine anuran species grew over time. ecological restoration = dark green pixels. otálora-löw et al. – enhancing ecological education 89 2.4625; p-perm = 0.0242). differences were present when comparing responses after the class and after interacting with the netlogo model (t-statistic = 1.823; p-perm = 0.0121), but there were no differences between before and after the class (t-statistic = 1.428; p-perm = 0.139). when considering all undergraduate and graduate students as a whole, it was found that the responses of women did not vary between stages of evaluation (pseudo-f = 0.88; p-perm = 0.473), but men did vary in their responses (pseudo-f = 3.32; p-perm = 0.0041), and differences were present when comparing between men responses before and after the class (t-statistic = 1.606; p-perm = 0.036) and after the class and after interacting with the netlogo model (t-statistic = 2.153; p-perm = 0.0007). when the interaction between these factors was tested, differences between stages of evaluation were found only for the male respondents in the graduate program (pseudo-f = 3.248; p-perm = 0.023); and differences were present when comparing between responses after the class and after interacting with the netlogo model (t-statistic=2.281; p-perm = 0.0003). for question 4 about the degree of influence of the edge effect on the abundance of a specialist native forest species on the interaction between the type of land use adjacent to the native forest patch and the shape of the native forest patch (question 4 appendix 3), there were statistical differences in graduate students between stages of evaluation (pseudo-f = 2.256; p-perm = 0.045) and differences were present when comparing between responses before the class and after interacting with the netlogo model (t-statistic = 1.938; p-perm = 0.013). however, no differences were found in the responses between evaluation stages for undergraduate students (pseudo-f = 1.101; p-perm = 0.344). when considering all undergraduate and graduate students, it was found that the responses of women did not vary between stages of evaluation (pseudo-f = 1.798; p-perm = 0.089), as was found in the men’s responses (pseudo-f = 1.695; p-perm = 0.11). when the interaction between these factors was tested, differences between stages of evaluation were found only for the responses of males in the graduate program (pseudo-f = 2.748; p-perm = 0.0234); and differences were present when comparing between responses before the class and after interacting with the netlogo model (t-statistic = 2.27; p-perm = 0.0028). discussion the model developed in this study works realistically as it considers the habitat’s environmental variables and the frogs’ morphological traits measured on the field (palomino-cuellar 2019). because of it, when used in class, it shows how some anuran species react differently to the native forest edges after landcover transitions. it also helped the students better understand the concept of edge effect, as corroborated by the answers of a simple poll answered by 24 of his students, where 95.8% of the students stated that using the model was beneficial. overall, using the model in class allowed the students to approach a complex subject with relative ease, enabling a more dynamic class, as students asked several questions about the model as they interacted with it. the following describes how the model was improved after feedback from the students as users. agent-based models to explain landscape transformation and effects on biodiversity considering the synthesized understanding of edge effects and their implications for biodiversity education, this study’s discussions probe the intersection of ecological concepts and pedagogical objectives. our discourse embarks on a nuanced exploration of how innovative educational technology intersects with ecological discourse. anchored by the objectives and outcomes of the study, the ensuing discussions dissect the transformative potential of agent-based modeling in conveying complex ecological phenomena to diverse educational contexts while allowing for the integration of technology and meaningful educational applications. through a pedagogical lens, we delve into the implications of these findings for fostering a deeper understanding of edge effects and promoting eco-literacy among students and educators alike. depending on the user´s needs, agent-based modeling can have varying degrees of complexity. using a “simple” agent model, with an assumption of homogeneity, is enough when trying to show a problem in such a way that it can be used in a pedagogical environment (brown et al., 2004); this, of course, implies that the models need to leave out information, that for the scientific community is needed to improve the rigor and robustness of information regarding the research topic (brown and robinson 2006). as such, agent modeling has been used to test diverse scientific hypotheses through otálora-löw et al. – enhancing ecological education 90 simulations, especially regarding biodiversity. a study from chetcuti et al., (2021) used agent-based modeling to show the impact of fragmentation regardless of habitat loss, concluding that different matrices can interact differently with the fragments, having beneficial or harmful consequences for biodiversity. however, the information provided by this article falls on the complexity side that cannot be used in class without prior conceptual training (walls and gabor 2019), even when the code is free and available, allowing trials and corrections from others wanting to understand the model. in contrast, horiuchi and takasaki’s (2012) agent modeling sought to understand how species take advantage of their group space. they did it in such a way that helped decision-makers acquire the information to understand the behavior and environment used by the japanese macaque, beginning a dialogue to promote its conservation. with this, even though there is still room for complex models in the academy, easily understandable explanations are still needed and have been proven helpful to ensure that processes of landscape transformation and their impact on biodiversity could be considered in decision-making (carey and gougis 2017; farrell and cayelan, 2018; read et al. 2016). in our case, the model aims to simplify a complex topic so that the students could approach it. information regarding some environmental variables was left behind, prioritizing the ease of use. species functional traits and responses to edge grouping species according to their responses to the edge allows for more effective targeting of conservation and restoration actions within a biotic community (pfeifer et al. 2017; schneidermaunoury et al., 2016). forest-exclusive anurans are exceptionally sensitive to the edge, as richness increases with distance from it. the environmental variables that explain this effect were changes in basal area and canopy coverage (cortés-gómez et al. 2013). however, daily temperature and understory density were also found to have explanatory power for the changes in richness (pearman, 1997). alternatively, studies show that the edge effect impacts the species level, as some are unaffected while others exhibit a decrease in their abundance due to factors such as seasonality (dry or wet) and availability of sites for reproduction and predation (gascon 1993; tsuji-nishikido and menin, 2011). likewise, palomino-cuellar (2019) noticed that environmental factors were the main drivers for the diversity and composition of anuran assemblages in our study region in the caquetá department (colombia), especially temperature, light penetration, distance to the border, relative humidity, and distance to the nearest body of water. additionally, it has been shown that energy consumption varies according to body size and temperature, even when the anuran is at rest (wells, 2007), meaning that environmental changes affect the rate at which amphibians consume energy (wells, 2007). this was not contemplated in the present model, as representing energy differently through the groups was thought to hinder the ability of the model to explain the edge effect, as it would have increased complexity. however, the effectiveness of conservation actions depends on knowledge of the functional traits that allow species to respond to changes in their habitat in transformed landscapes (zabala-forero and urbina-cardona, 2021). functional ecology is an approach to understanding the functions and responses of species within their environment; this acknowledges the roles within a habitat and how they react to changes in their environment regarding their functional traits (salgado-negret and paz, 2016). functional traits of anurans, such as snout-vent length, leg length, and body weight, are crucial for understanding the response of anurans to environmental filters (álvarez-grzybowska et al., 2020). traits used for the model (table 2) align with previous literature, as size and energy are crucial for describing reactions to novel environmental conditions. the selected variables match the knowledge of how environmental variables impact amphibian ensembles, where temperature (álvarezgrzybowska et al. 2020; harper et al. 2005), canopy coverage (cortés-gómez et al., 2013) and relative humidity (santos-barrera and urbina-cardona, 2011; lehtinen et al., 2003), are amongst the more revised. other variables like understory density (pearman, 1997) and wind (lehtinen et al., 2003) have affected the anura but were not included in the model to remain simple. using the ambs to understand edge effects better our results show that this model can be helpful, describing the complexity of the edge effect simply and intuitively for the students. of all the possible variables that are affected by the edge effect on the habitat of the species (broadbent et al., 2008) in the present research, we chose to simplify the model based on three variables: temperature, canopy cover, otálora-löw et al. – enhancing ecological education 91 and relative humidity. moreover, these variables have also been used to describe the impacts of the edge effect on other taxa, as they (amongst other variables) change in a gradient that impacts the microclimate (harper et al., 2005; ries et al., 2004). using three key environmental variables and applying this model in teaching activities was helpful, as it provided a tool for students to understand a complex topic from the interaction with the model and the discussions that arose in the class. however, our results suggest that the degree of interaction with the model and its usefulness in reinforcing knowledge about edge effects vary between undergraduate and graduate students and women and men. some of the testimonials from the students participating in the workshop show the usefulness of this tool in teaching. “being clearer and graphical helped me to understand” (a male in the undergraduate class). this kind of model has been used very little (murphy et al., 2020) to mediate meaningful learning in students, as a teaching method and a way to understand learning. still, it has been used when describing complex biological and social processes in scientific publications (koster et al., 2016), translating complex scientific literature to keep students motivated through simple models and case studies. a male student in the master’s class mentioned, “it helped me to visualize the number of variables and the level of interactions between them.” utilizing agent-based models also enables the dialogue to “initiate a discussion between experts and stakeholders, bringing together different expertise” (van berkel and verburg, 2012), allowing for better, more enriched decisions. empirically, the statistical results have been corroborated, as some male students indicated a personal gain from using the model. at the same time, the female group didn’t report acquiring new knowledge, instead, they expressed feeling more confident about what they already knew. for example, a female student in the master’s class mentioned that “i understood the concept in class, but seeing it didactically supports what was taught.” a male student in the undergraduate class mentioned that “it improved, the concept became clearer and how they interact with multiple factors.” recently, there have been some approaches to improve and expand the usage of agent-based modeling as a teaching and training tool (romanowska et al., 2021) as the extent of field investigations and classes were reduced due to the covid-19 pandemic (murphy et al., 2020). murphy (2020) breaks down the main components of an agent model and explains its benefits to students and teachers so that one can see different approaches to it. a female student in the undergraduate class mentioned, “the froggy that’s on the forest can move if there is restoration. it can displace,” suggesting that the model allows to better understand the concepts of edge effects, appropriate them, and generate more profound questions. conclusion this study was based on the amphibian data previously collected in the field by palomino-cuellar (2019), which was the support that allowed the creation of an abm that was able to represent the responses of three groups of amphibians to the edge effect. this model not only considered the functional traits that determine the movement and reproduction of the study species but also constrained the dispersal of individuals based on the values of three environmental habitat variables (temperature, canopy cover, and relative humidity) that are important not only for amphibians but in general for biodiversity. additionally, the model had several stages, in which it was intended to maintain the core functionality of the model but to make it as easy to use as possible so that it could be used in class by students. however, the need to establish clear rules of action must first be considered, as they are the steppingstones to maintain the model’s realism (to a certain extent). the workshops within the undergraduate and master’s classes served as pilot studies that allowed not only to improve the model´s graphical interface, but also to explore whether there were additional changes in the understanding of the edge effects seen in class after interacting with the model. the edge effect is a highly complex, and important topic to understand as it is intertwined with fragmentation. so, the main objective of this work was to generate a visual representation of it through modelling. the present model is versatile and flexible. it can incorporate new parameters and change scenarios as the dynamics of edge-effect classes generate new interactions for the teacher and his students. one such incorporation could be to provide a manual on how to use the model to promote independent learning. with enough usage and feedback, this model can be included in introductory university courses in ecology to aid in explaining this topic, giving more tools to both the teachers and the students, bettering the retention of the concept. a final recommendation for the readers and users otálora-löw et al. – enhancing ecological education 92 is to interact as a community with the model and the code, as it is free; they can learn more from the model by using it, adding and subtracting lines from the code, and asking the code the questions they want to answer. acknowledgments we want to thank all the students that participated in the survey, all the comments and feedback that brought the model to the state presented here. to all who use the model seeking to improve it or to understand the concept better. conflict of interest the authors have declared that no competing interests exist. references agudelo-hz, w., j. n. urbina-cardona, and d. armenteras-pascual. 2019. critical shifts on spatial traits and the risk of extinction of andean anurans: an assessment of the combined effects of climate and land-use change in colombia. perspectives in ecology and conservation, 17(4), 206–19. https://doi.org/10.1016/j. pecon.2019.11.002 álvarez-grzybowska, e., j. n. urbina-cardona, f. córdova-tapia, and a. garcía. 2020. amphibian communities in two contrasting ecosystems: functional diversity and environmental filters. biodiversity and conservation, 29(8), 2457–85. https://doi.org/10.1007/s10531-020-01984-w arroyo-rodríguez, v., l. fahrig, m. tabarelli, j. watling, l. tischendorf, m. benchimol, e. cazetta, et al. 2020. designing optimal human-modified landscapes for forest biodiversity conservation. ecology letters, 23(9), 1404–20. https://doi. org/10.1111/ele.13535 brewer, c. 2006. translating data into meaning: education in conservation biology. conservation biology, 20(3), 689–91. http://www.jstor.org/stable/3879232 broadbent, e. n., g. p. asner, m. keller, d. e. knapp, p. j. c. oliveira, and j. n. silva. 2008. forest fragmentation and edge effects from deforestation and selective logging in the brazilian amazon. biological conservation, 141(7), 1745–57. https://doi.org/10.1016/j.biocon.2008.04.024 brodman, r., m. parrish, h. kraus, and s. cortwright. 2006. amphibian biodiversity recovery in a large-scale ecosystem restoration. herpetological conservation and biology, 1, 101–8. brown, d. g., and d. t. robinson. 2006. effects of heterogeneity in residential preferences on an agent-based model of urban sprawl. ecology and society, 11, 1. https://doi.org/10.5751/es01749-110146 carey, c. c., and r. d. gougis. 2017. simulation modeling of lakes in undergraduate and graduate classrooms increases comprehension of climate change concepts and experience with computational tools. journal of science education and technology, 26, 1–11. https://doi.org/10.1007/ s10956-016-9644-2 cayuela h., a. besnard, e. bonnaire, h. perret, j. rivoalen, c. miaud, and p. joly. 2014. to breed or not to breed: past reproductive status and environmental cues drive current breeding decisions in a long-lived amphibian. oecologia, 176, 107– 16. https://doi.org/10.1007/s00442-014-3003-x chetcuti, j., w. e. kunin, and j. m. bullock. 2021. matrix composition mediates effects of habitat fragmentation: a modelling study. landscape ecology, 36(6), 1631–46. https://doi. org/10.1007/s10980-021-01243-5 chopin p., g. bergkvist, and l. hossard. 2019. modelling biodiversity change in agricultural landscape scenarios: a review and prospects for future research. biological conservation, 235, 1–17. https://doi.org/10.1016/j.biocon.2019.03.046 cortés-gómez, a. m., f. castro-herrera, and j. n. urbina-cardona. 2013. small changes in vegetation structure create great changes in amphibian ensembles in the colombian pacific rainforest. tropical conservation science, 6(6), 749–69. díaz-garcía, j. m, e. pineda, f. lópez-barrera, and c. e. moreno. 2017. amphibian species and functional diversity as indicators of restoration success in tropical montane forest. biodiversity and conservation, 26(11), 2569–89. https://doi. org/10.1007/s10531-017-1372-2 didham, r. k., v. kapos, and r. m. ewers. 2012. rethinking the conceptual foundations of habitat fragmentation research. oikos, 121(2), 161–70. https://doi.org/10.1111/j.16000706.2011.20273.x farrell, k. j., and c. c. carey. 2018. power, pitfalls, and potential for integrating computational literacy into undergraduate ecology courses. ecolootálora-löw et al. – enhancing ecological education 93 gy and evolution, 8(16), 7744-7751. https://doi. org/10.1002/ece3.4363 feder, m. e., and w. w. burggren. 1992. environmental physiology of the amphibians. university of chicago press. ferraz, k. m. p. m. b. f., r. g. morato, a. a. a. bovo, c. o. r. da costa, y. g. g. ribeiro, r. cunha de paula, a. l. j. desbiez, c. s. c. angelieri, and k. traylor‐holzer. 2021. bridging the gap between researchers, conservation planners, and decision makers to improve species conservation decision‐making. conservation science and practice, 3(2), e330. https://doi.org/10.1111/ csp2.330 gascon, c. 1993. breeding-habitat use by five amazonian frogs at forest edge. biodiversity and conservation 2(4), 438–44. https://doi.org/10.1007/ bf00114045 grimm, v., u. berger, d. l. deangelis, j. g. polhill, j. giske, and s. f. railsback. 2010. the odd protocol: a review and first update. ecological modelling, 221(23), 2760–68. https://doi. org/10.1016/j.ecolmodel.2010.08.019 harper, k. a., s. e. macdonald, p. j. burton, j. chen, k. d. brosofske, s. c. saunders, e. s. euskirchen, d. roberts, m. s. jaiteh, and p. a. esseen. 2005. edge influence on forest structure and composition in fragmented landscapes. conservation biology, 19(3), 768–82. https://doi. org/10.1111/j.1523-1739.2005.00045.x hernández-ordóñez, o., j. n. urbina-cardona, and m. martínez-ramos. 2015. recovery of amphibian and reptile assemblages during old-field succession of tropical rain forests. biotropica, 47(3), 377–88. https://doi.org/10.1111/btp.12207 herrera-montes, a., and n. brokaw. 2010. conservation value of tropical secondary forest: a herpetofaunal perspective. biological conservation, 143(6), 1414–22. https://doi.org/10.1016/j.biocon.2010.03.016 hillman, s. s., p. c. withers, r. c. drewes, and s. d. hillyard. 2009. ecological and environmental physiology of amphibians. oxford university press. horiuchi, s., and h. takasaki. 2012. boundary nature induces greater group size and group density in habitat edges: an agent-based model revealed. population ecology, 54(1), 197–203. https://doi. org/10.1007/s10144-011-0279-0 lehtinen, r. m., j. b. ramanamanjato, and j. g. raveloarison. 2003. edge effects and extinction proneness in a herpetofauna from madagascar. biodiversity and conservation 12 (7), 1357–70. https://doi.org/10.1023/a:1023673301850 marsh, d. m., and trenham, p. c. (2001). metapopulation dynamics and amphibian conservation. conservation biology, 15(1), 40-49. matas, antonio. 2018. diseño del formato de escalas tipo likert: un estado de la cuestión. revista electronica de investigacion educativa 20 (1), 38–47. https://doi.org/10.24320/redie.2018.20.1.1347 moreno-rubiano, m. c., moreno-rubiano, j. d., robledo-buitrago, d., de luque-villa, m. a., urbina-cardona, j. n., and granda-rodriguez, h. d. (2023). perception and attitudes of local communities towards vertebrate fauna in the andes of colombia: effects of gender and the urban/rural setting. ethnobiology and conservation, 12. https://doi.org/10.15451/ec2023-06-12.09-1-20 murphy, kilian j., s. ciuti, and a. kane. 2020. an introduction to agent-based models as an accessible surrogate to field-based research and teaching. ecology and evolution 10 (22), 12482– 12498. https://doi.org/10.1002/ece3.6848 palomino-cuellar, j. v. 2019. diversidad funcional y taxonomica de anfibios en diferentes coberturas de la tierra en una porcion de la llanura amazonica caqueteña. trabajo de grado de la maestría en conservación y uso de biodiversidad, pontificia universidad javeriana, bogotá, colombia. pearman, p. b. 1997. correlates of amphibian diversity in an altered landscape of amazonian ecuador. conservation biology 11 (5), 1211–1225. https:// doi.org/10.1046/j.1523-1739.1997.96202.x pfeifer, m., v. lefebvre, c. a. peres, c. banks-leite, o. r. wearn, c. j. marsh, s. h.m. butchart, et al. 2017. creation of forest edges has a global impact on forest vertebrates. nature 551 (7679), 187–191. https://doi.org/10.1038/nature24457 read, e. k., m. o’rourke, g. s. hong, p. c. hanson, l. a. winslow, s. crowley, c. a. brewer, and k. c. weathers. 2016. building the team for team science. edited by d. p. c. peters. ecosphere 7 (3). https://doi.org/10.1002/ecs2.1291 redford, k. h., c. groves, r. a. medellin, and j. g. robinson. 2012. conservation stories, conservation science, and the role of the intergovernmental platform on biodiversity and ecosystem services. conservation biology 26 (5), 757–759. https://doi.org/10.1111/j.15231739.2012.01925.x ries, l., r. j. fletcher, j. battin, and t. d. sisk. 2004. ecological responses to habitat edges: mechaotálora-löw et al. – enhancing ecological education 94 nisms, models, and variability explained. annual review of ecology, evolution, and systematics 35, 491–522. https://doi.org/10.1146/annurev. ecolsys.35.112202.130148 romanowska, i., c. wren, and s. a. crabtree. 2021. agent-based modeling for archaeology. sfi press. https://doi.org/10.37911/9781947864382 salgado-negret, b., and h. paz. 2016. escalando de los rasgos funcionales a procesos poblacionales, comunitarios y ecosistémicos. in la ecología funcional como aproximación al estudio, manejo y conservación de la biodiversidad: protocolos y aplicaciones. instituto de investigación de recursos biológicos alexander von humboldt. bogotá, colombia. santos-barrera, g., and j. n. urbina-cardona. 2011. the role of the matrix-edge dynamics of amphibian conservation in tropical montane fragmented landscapes. revista mexicana de biodiversidad 82 (2), 679–87. https://doi.org/10.22201/ ib.20078706e.2011.2.463 schneider-maunoury, l., v. lefebvre, r. m. ewers, g. f. medina-rangel, c. a. peres, e. somarriba, n. urbina-cardona, and m. pfeifer. 2016. abundance signals of amphibians and reptiles indicate strong edge effects in neotropical fragmented forest landscapes. biological conservation 200 (august), 207–215. https://doi.org/10.1016/j. biocon.2016.06.011 tisue, s, and u. wilensky. 2004. netlogo: design and implementation of a multi-agent modeling environment. proceedings of the agent 2004 conference on social dynamics: interaction, reflexivity and emergence, chicago. toledo, l. f., sazima, i., and haddad, c. f. b. 2009. behavioral defenses of anurans: an overview. ethology ecology and evolution, 21(1), 1-12. tscharntke, t., j. m. tylianakis, t. a. rand, r. k. didham, l. fahrig, p. batáry, j. bengtsson, y. clough, t. o. crist, c. f. dormann, r. m. ewers, j. fründ, r. d. holt, a. holzschuh, a. m. klein, d. kleijn, c. kremen, d. a. landis, w. laurance, d. lindenmayer, c. scherber, n. sodhi, i. steffan-dewenter, c. thies, w. h. van der putten, c. westphal. 2012. landscape moderation of biodiversity patterns and processes eight hypotheses. biological reviews 87 (3), 661–685. https://doi. org/10.1111/j.1469-185x.2011.00216.x tsuji-nishikido, b. minoru, and m. menin. 2011. distribution of frogs in riparian areas of an urban forest fragment in central amazonia. biota neotropica 11 (2), 63–70. https://doi.org/10.1590/ s1676-06032011000200007 van berkel, derek b., and peter h. verburg. 2012. combining exploratory scenarios and participatory backcasting: using an agent-based model in participatory policy design for a multi-functional landscape. landscape ecology 27 (5), 641–658. http://dx.doi.org/10.1007/s10980-012-9730-7 vergara-ríos, d., montes-correa, a. c., urbina-cardona, j. n., de luque-villa, m., cattan, p. e., and granda, h. d. (2021). local community knowledge and perceptions in the colombian caribbean towards amphibians in urban and rural settings: tools for biological conservation. ethnobiology and conservation, 10. https://doi. org/10.15451/ec2021-05-10.24-1-22 vitt, l. j., and j. p. caldwell. 2013. herpetology: an introductory biology of amphibians and reptiles. academic press. walls, susan c., and caitlin r. gabor. 2019. integrating behavior and physiology into strategies for amphibian conservation. frontiers in ecology and evolution 7, 234. http://dx.doi.org/10.3389/ fevo.2019.00234 wells, k. d. (2007). the ecology and behavior of amphibians. university of chicago press. wells, kentwood d. 2007. metabolism and energetics. in the ecology and behavior of amphibians, 184–189. chicago. zabala-forero, f., and n. urbina-cardona. (2021). taxonomic and functional diversity responses to landscape transformation: relationship of amphibian assemblages with land use and land cover changes. revista mexicana de biodiversidad, 92: e923443. https://doi.org/10.22201/ ib.20078706e.2021.92.3443 otálora-löw et al. – enhancing ecological education 95 appendix 1: overview, design, details 1.1 overview 1.1.1 purpose (1) what is the purpose of the study? the purpose of this model is to show, to the general public, the effect of natural forest edges on amphibian species, in such a way that they can easily grasp the core concepts through visual representation. the model must be user-friendly, meaning that buttons, sliders, and other interactive interfaces have to be easy to use and labeled. even though this model uses specific species the model can, to a certain extent, be used to explain other species. this kind of model is necessary because the basic ecological information incorporated into the model presents a high degree of complexity and its design in such a way that someone not akin to science will not be able to integrate the different components (from species traits to habitat variables) to correctly interpret the pattern from the original data. the information mustn’t be inaccessible, on the contrary, information needs to go to where it’s needed the most, the next generation of students and researchers, the governors and agencies, decision-makers, farmers, between other stakeholders. (2) for whom is the model designed? this model is designed for those who are affected and/or interested in understanding the edge effect caused by deforestation but don’t have a strong ecological background. 1.1.2 entities, state variables, and scales the agents are based on real amphibian species: 1) dendropsophus mathiassoni for the pasture habitat, 2) amazophrynella minuta, leptodactylus petersii and pristimantis variabilis for the border habitat and 3) engystomops petersi, p. acuminatus and p. conspicillatus for the forest interior habitat. the first species is alone as it has a lot of data, giving a very good representation of the pasture habitat. the border and forest interior habitat have 3 species each, this is because alone they have a small number of reports but thanks to their similarity in reaction to the edge, we can group them in order to maximize the information per habitat. thus, we have 3 representative groups. each habitat has an independent number of the same environmental variables. the variables that change each tick: (1) heading of the agent, and (2) position (x and y coordinates). the variables that change between groups: idle movement – all frogs have an idle movement, this means that they always move, different from running. this is placed to represent the type of foraging of the frogs, some are active forages meaning that they consume energy looking for food, while others are passive (sit and wait forage strategy; represented by low movement) and waiting for food (cortés-gómez et al. 2015). running capacity – this number is set between 1 to 10. it is based on the physiological attributes of the group, such as leg length and body weight – length ratio (cortés-gómez et al. 2015). it determines how many pixels can the agent move in each tick. reproduction – reproduction hatches a new frog, with a random heading and a third of the energy of the parent. the parent loses a substantial amount of energy, this depends on the type of reproduction (if the type of reproduction is different between species of the same group the method that implies more energy consumption is used to overestimate the consumption instead of understanding it). reproduction has a set time for each group, and it is determined by the mechanisms employed at amplexus. the frogs can only reproduce when a set time has passed. energy – energy goes from 0 to 100 and represents the percentage of available energy that the agent must perform movement o reproduction. the usage of energy is determined by 2 things, the environment in which the amphibian stands (is it optimal?) and the quantity of available energy. if there is now enough energy for reproduction, a new turtle is hatched with a third of the energy from the parent. the frogs inhabit one of 3 habitats: pasture, forest edge, and forest interior. at setup, the turtles spawn in their optimal habitat distributed randomly. the habitat patches have three main environmental variables: temperature, canopy cover, and relative humidity. these factors change as they interacted with other habitats, constituting the edge effect. in go, the habitats change randomly meaning that spawning patches (probably) won’t have the same conditions when compared to the setup. all variables (except idle movement) can be changed by the user of the model, including the rate at which the habitats change. the model has no explicit time or spatial scale. however, to make a more otálora-löw et al. – enhancing ecological education 96 realistic approach a limit to population growth can be established this can be done by giving the pixels a spatial value. for example, if each pixel is 1 x 1 meter then we can conduct a comparison to the real field effort of palomino. the area covered in the study from which the data was taken was 450m2 per habitat (palomino 2019), so our 30 pixels by 60 pixels grid equates to an area of 1600 m2. with this, if the study found 15 dendropsophus mathiassoni in 450 m2 then, proportionally, we can estimate that in 1600 m2 there are ~54; the model then can be stopped at 55 individuals of that species. some exogenous factors that exist but are not represented in the model included social and cultural processes, economics, and water flow dynamics. process overview and scheduling time is measured in discrete ticks. before starting: set the initial variables of the patches, the number of frogs per habitat, and their energy. all amphibians do this for each tick: the frogs look at their surrounding in a radius of 5 and check if the environment satisfies their established needs. if the environment isn’t optimal, face the nearest patch with the conditions and runs towards it (run capacity). if no energy, die, disappearing the frog form the screen. if the environment is optimal wait until the reproduction time. acquire energy while waiting. when the time comes, reproduce. the new frog gets the same reproduction time as the parent and a third of its energy. the patches then change according to the change fact of a slider. these changes are done through neighbors and generate different temperatures, relative humidity, and canopy cover and the border moves. end of the tick 1.2 design concepts 1.2.1 emergence there are two main outputs from the model: • count patches habitat coverage graph is a graph that periodically changes according to the count of patches, which changes by the value set at the beginning of the simulation • count frog – with this graph we can see how the populations are changing over time. even though the parameters on which the agents operate are fixed, both outcomes are, to a certain extent, unpredictable as there is a random generation of values that can determine the survival of an agent. reproduction may be affected by random changes in the environment. 1.2.2 adaptation as the environment changes, the agents will look for the nearest patch that has the appropriate conditions and run towards it. the idea behind it is that the goal of the agents is just to survive, and so when they see that the environment is out of its optimal range (causing a loss in energy and ultimately death) they’ll seek the best conditions as fast as possible. 1.2.3 objective the agents don’t have an inherent objective. the model just wants to show the interaction of the amphibian with the with the edge that is created between a native forest and an anthropogenic production system, and they are programmed to look for the best place to survive. they can only perceive the neighbors (radius of 5 pixels) and don’t recall where they have been before. 1.2.4 learning agents do not learn. 1.2.5 prediction agents are unable to predict future conditions 1.2.6 sensing agents can check the environmental characteristics (temperature, humidity, and canopy cover) of the patch they’re in and the neighbor patches (5-pixel radius). 1.2.7 interactions agents do not interact with each other nor affect their actions. patches interact with each other as there is a chance that a patch changes from one habitat to another, this is chance is determined through the interface. patches interact with the agents. 1.2.8 stochasticity movement direction is set by a random heading. hatched offspring has a random float. the border takes the minimum value found on field studies for each variable and adds a random number the maximum of which would leave the variable at the average. agents spawn at random within the boundaries of their corresponding habitat. 1.2.9 collectives three groups were created, composed of various species that have similar responses to the natural forest edge. in the model, the pasture group of species otálora-löw et al. – enhancing ecological education 97 is represented by color yellow, forest edge species by the blue color, and forest interior species by the green color. groups don’t interact with each other, and the quantity of individuals only affects the growth rate. 1.2.10 observation in information used in the model was taken from a study of caquetá done by palomino (2019). now, the interface shows at the top the setup and go button. then, there are some sliders where the user can input the initial number of amphibians per group, their initial energy, and their movement. they can also select the change rate of the patches, selecting restoration, abandonment, and deforestation. there is also a button at the end to establish neutral values, these are the ones found in the study of palomino and the rates of habitat change are balanced. in the world, the user can see the movement of the amphibians, their reproduction, and the changes of the patches by their color. on the right, two graphs show the changes in population and the changes in patches. 1.3 details 1.3.1 initialization the model has set values for the environmental conditions when in setup and there is a button to establish the neutral values for the rate of change of the patches and the exact number of individuals found on field research. however, everything is changeable through the sliders (aside from the initial values of temperature, canopy, and humidity) 1.3.2 input data the model doesn’t have external data files. 1.3.3 submodels no other models were explicitly integrated. otálora-löw et al. – enhancing ecological education 98 appendix 2: study area caquetá is a department located in the amazon region in southern colombia (figure 1). it has approximately 89.000.000 km2, and the predominant ecosystem is the tropical rainforest. with a deforestation rate of 0.77% (murad and pearse, 2018), the caquetá department has seen a very aggressive ramp-up in deforestation due to a disorganized colonization process (etter et al. 2006). primarily, this colonization has introduced cattle, turning a substantial amount of area into pastures so that the cattle can be productive (etter et al., 2008). this process has caused 12647 ha of deforestation in caquetá between 2022 and 2023, ranking it as the most deforested department in colombia (ideam, 2023), causing population declines as some species of amphibians are unable to escape or adapt to the sudden degradation of their environment (palominocuellar, 2019). appendix: literature cited etter, a., c. mcalpine, s. phinn, d. pullar, and h. possingham. 2006. unplanned land clearing of colombian rainforests: spreading like disease? landscape and urban planning, 77, 240–54. https://doi.org/10.1016/j.landurbplan.2005.03.002 etter, a., c. mcalpine, and h. possingham. 2008. historical patterns and drivers of landscape change in colombia since 1500: a regionalized spatial approach. annals of the association of american geographers, 98, 2–23. https://doi.or g/10.1080/00045600701733911 ideam. 2023. monitoreo de la superficie de bosque y la deforestación en colombia 2023 (resumen de resultados). murad, c. a., and j. pearse. 2018. landsat study of deforestation in the amazon region of colombia: departments of caquetá and putumayo. remote sensing applications: society and environment, 11, 161–71. https://doi.org/10.1016/j. rsase.2018.07.003 palomino-cuellar, j. v. 2019. diversidad funcional y taxonómica de anfibios en diferentes coberturas de la tierra en una porción de la llanura amazónica caqueteña. trabajo de grado de la maestría en conservación y uso sustentable de la biodiversidad. facultad de estudios ambientales y rurales. pontificia universidad javeriana. 78 p. figure 1. map of the study area where palomino-cuellar (2019) collected the data. solano, caquetá, colombia. otálora-löw et al. – enhancing ecological education 99 appendix 3: format of the survey the first two questions ask about the definition of the edge effect and the aspects that can influence the impact of this effect for biodiversity. the third question consisted of a matrix of three native forest fragment shapes (round, rectangular and irregular) vs. different native fragment sizes (big, medium or small) in which the students had to give a number between 1 and 5, were 1 is low edge effect impact and 5 high edge effect impact in a specialist species of the native forest; the students also had the chance of not putting anything if they were not confident with their answer. this matrix shows the knowledge regarding both shape and size of the patch which are key concepts when understanding the edge. the fourth question asked students to fill out a matrix similar to the previous one but contrasting the same three native forest fragment shapes but in this case limiting with five different land uses (pasture, shade coffee plantation, cocoa plantations and secondary native forest). students had to judge the impact on the edge effect in a specialist species of the native forest of different vegetation cover contexts bordering the remnant native forest in a scale of 1 to 5, were 1 is low edge effect impact and 5 high edge effect impact, they also had the chance of not answering if they were not confident. edge effect: survey phase 1 name: _______________________________ question 1: how did the concept of edge effect changed after interacting with the model? question 2: list some aspects that can be affected by edge effect question 3: thinking of a specialist specie of native forest: fill the blanks on the following table, giving values between 1, when you consider that the edge effect would be weak, and 5, when you consider that the edge effect is strong. the numerical value evidences the grade of influence of the edge effect when it interacts with (a) size of the native forest and its shape; and (b) interaction between the land use adjacent with the native forest and the shape of the native forest. 1 __________________5 weak strong otálora-löw et al. – enhancing ecological education 100 additionally, after phase 3, once the students interact with the netlogo model, they were asked to answer two additional questions: question 1: how did the concept of edge effect changed after interacting with the model? question 2: what new aspects did you identify, that are affected by the edge effect, after using the model? biodiversity informatics, 19, 2025, pp. 7-32 7 stopover hotspots for migratory birds in north and central america shi feng1, qinmin yang1*, huijie qiao2*, luis e. escobar3, xuan yan4 1state key laboratory of industrial control technology, college of control science and engineering, zhejiang university, hangzhou, 310007, pr china. 2state key laboratory of animal biodiversity conservation and integrated pest management, institute of zoology, chinese academy of sciences, beijing, 100101, pr china. 3department of fish and wildlife conservation, 1015 life science cir, virginia tech, blacksburg, va 24061, usa. 4independent researcher, 7912 heritage palms trl, mckinney, tx 75070, usa. abstract. despite the large body of literature on avian migratory behavior, there is little information about stopover sites during bird movement, including the population-level drivers of breeding grounds and wintering grounds. stopovers play an essential role in bird migratory site chains for energy supply and rest. there is an urgent need to identify and protect stopover sites to secure the long-term sustainability of migratory network connectivity and stability. to address this challenge, we reconstructed a migration network and identified geographic hotspots denoted as stopover sites. and we analyzed the high-density population movements of 52 focal migratory bird species using comprehensive observation data from ebird through pagerank algorithm. furthermore, potential alternative stopover sites were explored using a word embedding technique based on geo-functional similarity. our study was conducted in north and central america during a three-year period and revealed three key stopover areas, including florida peninsula and its inland, the region of central america, and the region near puget sound. results from this study can be used for conservation prioritization guidance, active surveillance of bird pathogens, and bird management. keywords: central america, bird conservation, ebird, migration network, stopover introduction understanding migration pathways is critical for bird management and conservation. migratory birds travel thousands of miles between breeding and wintering grounds, facing numerous threats along their routes, including habitat loss, climate change, and hunting (nemes et al., 2023). by analyzing migration pathways, conservationists can identify key stopover sites and habitats that are essential for bird survival. information about key stopover sites enables the creation of targeted protection strategies, such as establishing protected areas, restoring habitats, and implementing international agreements to safeguard migratory routes (higuchi et al., 2012). additionally, migration pathway studies can help predict the impacts of environmental changes on bird populations for proactive measures to mitigate potential threats (la sorte et al., 2016) and to implement precision epidemiology of avian influenza. usually, migration depends on a suit of interconnected sites (runge et al., 2015), and their conservation requires a deep understanding of the connectivity in migration networks. that is, untangling migration patterns requires to answer how migratory species connect their breeding and non-breeding grounds through their trajectories. migratory connectivity can describe the spatiotemporal link of individuals and populations between sites caused by migratory movements, which can influence both long-term evolutionary responses and short-term population dynamics (webster et al., 2002). research on bird migratory connectivity generally focuses on geographical patterns that describe linkages between breeding and wintering grounds and the influencing drivers, such as resource requirements and ecological relationships (kramer et al., 2018). stopover sites, such as wetlands, forest fragments or grasslands, are essential for birds to rest and *corresponding authors: qinmin yang, email: qmyang@zju.edu.cn, huijie qiao, email: qiaohj@ioz.ac.cn mailto:qmyang@zju.edu.cn shi feng et al. – stopover hotspots in north and central america 8 refuel during their migration journey. these sites are considered to be key nodes in migration network connectivity and require targeted attention (guo et al., 2024). traditionally, migration connectivity research relies on tracking techniques. for example, knight et al. (2018) constructed a migration network for tree swallow (tachycineta bicolor) with 133 geolocators to assess the spatial connection between wintering and breeding areas to help develop optimal conservation strategies in north america. similarly, xu et al. (2020) applied high-resolution gps tracking data of 10 whooper swans (cygnus cygnus), 81 swan geese (anser cygnoides), 93 bar-headed geese (anser indicus), and 54 greater white-fronted geese (anser albifrons) with gps loggers during 2005-2018 to study the migration network connectivity in the central and east asian-australasian flyways. lastly, catry et al. (2024) followed the migratory trajectories of 20 grey plovers (pluvialis squatarola) with tracking devices to highlight important stopover sites and potential bottlenecks. these studies focused on the individual-level flyway and network, limiting the insight into population-level movements and being restricted to tracking costs and bird body size. tracking a few individuals is not suitable for reconstructing the movement of a species. similarly, individual tracking covers limited species due to the difficulty of capturing and carrying the track devices for small-sized ones. based on this data limitation of animal movement, broad-scale monitoring, such as radar or citizen science data, can provide comprehensive data for wider views of signals about the biogeography of bird movement. for example, bonter et al. (2009) identified critical stopover sites with weather-surveillance radar images from 2000 to 2001 in the great lakes basin. guo et al. (2024) applied five years of weather surveillance radar data to map the stopover densities of land birds during spring and autumn migration across the eastern united states. radar data, however, fail to identify specific bird species and only provide information on total bird biomass. this coarse-scale estimation limits conservation and management strategies for specific bird species, and is subjected to the radar coverage region. ai et al. (2024) developed a portable stereo vision observer for bird flocks in field scenarios based on feature and sensor methods. in their study, ai et al. captured birds natural flocking behaviors such as foraging and convergent flying in mudflats and seashores, which is helpful for stopovers detection. nevertheless, observer activities are assumed to be scenario-oriented and can be disturbed by temperature and wind, and this observer focuses on shortterm movement of specific individuals within the observation site, failing to capture dynamics of the long-distance migration process. ebird (https://ebird.org/) is the world’s largest birdwatching data repository and has comprehensive information on the distribution and movement patterns of avian species at the population level. lin et al. (2020) used ebird data to discover priority stopover sites (psss) and quantified the potential benefit for resident species in terms of species abundance focusing on three north american countries, including canada, the united states (us), and mexico. nicol et al. (2023) proposed a hidden semi-markov model to infer crucial stopover nodes based on count data from ebird to estimate the most likely migration links among regions for an endangered shorebird named far eastern curlew (numenius madagascariensis) in the east asian-australasian flyway. the current literature on stopover sites, however, evaluates the importance of sites based on population abundance or potential distributions combining the site protection status in isolation, but not from a systematic perspective. the spatial and temporal information of linkages among stopover sites for multiple species within a migratory network is usually overlooked. for example, zhang et al. (2023) explored the backbone nodes of migration network among different regions using data from the global biodiversity information facility (gbif, https://www.gbif.org/) for 1862 species in 26 bird orders. in their study, zhang et al identified the relative importance of nodes by betweenness (wang et al., 2008) which is the number of shortest paths passing through the individual nodes in a network. the zhang et al study, however, focused on the role that nodes play in the shortest paths of the network rather than paths across the network. therefore, methods to comprehensively reconstruct overall stopover site connection are needed for biologically realistic migration network analysis (wyborn and evans, 2021). pagerank algorithm and word embedding technique can help address the need of more accurate network modeling methods for stopover detection. pagerank was originally used for webpage ranking in google (page et al., 1999), where pagerank considers the number of connections to a webpage (represents a node in the users’ clicking network). pagerank emphasizes the position and link relationship of nodes in the overall structure of the network. it measures node importance by analyzing the relationship between nodes, accounting for multiple indicators such as linkage quantity and quality, link distribution, damping factor, and initial pagerank, among others. through the comprehensive calculation of these indicators, pagerank can determine the relative importance of nodes in the network. usually, nodes linked to a high-pagerank node will have an increasing pagerank value and are considered more important in the network. in human society networks, pagerank values have been shi feng et al. – stopover hotspots in north and central america 9 used to identify important user social nodes (hong et al., 2023), and analyze urban mobility network (wang et al., 2017). word embedding is a natural language processing (nlp) technique used to represent words as dense vectors (mikolov et al., 2013). these vectors capture semantic relationships and contextual information about words, allowing computers to process and understand human language more effectively. word embedding methods have been widely used for vector representations of words in documents, social analyses, and location analyses in human society (jin et al., 2014). mikolov et al. (2013) introduced word2vec algorithm to represent each word as a vector using large-scale document datasets. words with similar contexts will have similar representations in the vector space. it is popular in the nlp applications, including machine translation (qian et al., 2019), citation visualization (berger et al., 2016), and sentiment analysis (liu et al., 2020; xiong et al., 2018). meanwhile, the method has also been widely applied in human mobility data for geo-functional similarity exploration based on semantic similarity (zhu et al., 2019), where a location can be defined as a word, and a set of successive and previous locations in a trajectory can be defined as its context. considering that trajectory data is a type of sequential data, exploring the function of a location through its context is similar to understanding the meaning of a word in a sentence. for example, zhu et al. (2019) proposed a location representation location2vec method based on word2vec to capture the similarity relationships among locations. based on these, the aim of this work was to investigate the stopover hotspots during the bird migration process and identify alternative stopover sites to increase the site connectivity along bird migratory networks. we assessed the importance of stopover sites with the network-level linkage quantity and quality through the pagerank algorithm. this study provides a comprehensive population-level understanding for migration trajectories for diverse taxa using ebird data. we introduced the word embedding technique to analyze trajectory sequence contexts and detected potential alternative sites based on geo-functional similarity. discovery of specific stopover hotspots are expected to inform migratory bird conservation, surveillance, and management for north and central america birds. materials and methods data acquisition and preparation bird occurrence data for 52 focal migratory bird species (table s1) were collected from the ebird (https:// ebird.org/) citizen-science database for the period 20172022. the 52 migratory species are selected from 82 focal species list (schrimpf et al., 2021) where observation data are abundant for analysis. data from 2017-2019 worked as the study dataset, and the data from 2020-2022 were used for network validation. observation data covered continental north and central america with metadata including date and location (longitude and latitude) for each record. the study was carried out in matlab r2021b for pagerank value calculation, and python 3.10 for trajectory estimation and word2vec analysis. the specific packages are shown below for the full process: pandas (version 1.4.3; reback et al., 2022), numpy (version 1.23.5; harris et al., 2020), datetime (version 4.4), matplotlib (version 3.5.3; hunter, 2007), xlwt (version 1.3.0; machin, 2017), pyproj (version 3.3.1; whitaker, 2022), pygam (version 0.8.0; marín, 2018), scikit-learn (version 1.1.2; grisel et al., 2022), basemap (version 1.3.4; whitaker, 2022), tqdm (version 4.65.0; yorav-raphael & da costa-luis, 2024), xlrd (version 2.0.1; withers, 2020), openpyxl (version 3.1.2; gazoni & clark, 2023), genism(version 4.3.1; rehurek, 2023), and xlwings(version 0.30.6; zumstein,2023). we randomly subsampled the original dataset to retain 100 records per day to keep the balance between computation cost and information available. migration trajectories of high-density populations for each species during each migration cycle were estimated using a minimum cost analysis (feng et al., 2021; somveille et al., 2021). data were first preprocessed through mean location interpolation for observations of missing dates, using rolling-window smoothing (zivot and wang, 2007), and outliers were detected and removed using space local deviation factor algorithm (zhang and wang, 2011). observations were then discretized by mean-shift clustering algorithm (derpanis, 2005) for dense population centroids. centroids were grouped according to the minimum cost principle and trajectories were fitted using a generalized additive model (hastie and tibshirani, 1990). migration trajectory sequence conversion to convert the migration trajectories into geographic cell-id sequences, we set an origin point in position (0.1° n, 134.2° w) which was (10799.54, -14796209.33) in the spherical mercator map (fig.1), and built a grid coordinate system for 35 x 36 cells divided by: (-14796209.33 + 200000m, 10799.54 + 200000n) where m ϵ [1,35], n ϵ [1,36] to cover the continental north and central america (fig.1). then the number for each cell was calculated by: m + 35 * (n – 1). shi feng et al. – stopover hotspots in north and central america 10 migration network construction and stopover hotspots mining in this study, each geographic cell was defined as a node in the migration network with links resulting from cells with occurrences in adjacent positions in trajectory sequences. node importance in the bird migration network was assessed through birds’ movement between cells with the pagerank algorithm. we set the initial pagerank values for all nodes as 1/n, where n was the total number of nodes (rogers, 2002). in order to avoid infinite iteration, it is usually necessary to set a convergence threshold. that is, if the pagerank value of each web page changes less than a preset threshold (0.0001 in our case), its value is considered to have converged, and the iteration can be stopped. alternatively, by setting the maximum number of iterations as an iteration termination condition (100 in our case). the damping factor is constant with an empirical value of 0.85, which is used to simulate the movement probability between different nodes and avoid the problem of infinite loops. after setting the movement relationship between nodes, the pagerank value calculation result of each node can be obtained according to the set threshold and the maximum number of iterations. pagerank values could represent the relative importance of each node in the migration network, which was used for node ranking. the mathematical expression for pagerank algorithm is: ( ) ( ) (1 ) ( ) ( ) i i b a i b pr node pr node d d l node = − + ∑ where, pr(nodea) is the pagerank of node a; d is the damping factor; pr ( ) ibl node are the pagerank values of nodes that link to node a; ( ) ibl node is the number of outbound links on node bi. the cells in the grid coordinate system after geographic partition in this study were considered as nodes and the sequence of trajectories formed links between different nodes to construct a migration network. we defined nodes with the sojourn time between 5-20 days in the trajectory sequences as stopover sites based on the empirical distribution of bird stopover days during migration (kaiser, 1999; fig.s1). this stopover period allowed us to differentiate stopover sites from breeding and wintering sites. potential alternative stopover sites exploration cell-id sequences were analyzed using a word embedding technique to explore the potential alternative sites for stopover hotspots identified through pagerank algorithm. a cell-id in the trajectory sequence was defined as a word, while a trajectory sequence was defined as a sentence, with multiple trajectories defined as a document corpus. this ensemble of cells allowed us to encode a migratory trajectory as: tr = {cell – id1, cell – id2, ..., cell – idk, k = 1,2,...} figure 1. schematic diagram for geographic partition by building 35×36 cells to cover continental north and central america. the number for each cell is 1-35 in the bottom line, then 36-70 for the second line, and so on. through the grid coordinate system after geographic partition, trajectories are presented as geographic cell-id sequences according to daily locations. for example, the ellipse trajectory in the diagram above can be represented by a sequence from point a to point b and back to point a as follows (sojourn time ignored): 543-544-580-615-650-684-719-754-788-822856-891-854-818-783-748-714-679-644-610-576-542-543. shi feng et al. – stopover hotspots in north and central america 11 all the trajectory sequences were encoded as: tr = {tr1, tr2,...trn, n = 1,2,...} then we confined the context of word cell – idi as: ( 1) 1 1 ( 1)( ) { , , , ,..., , }i i m i m i i i m i mcontext cell id cell id cell id cell id cell id cell id cell id− − − − + + − +− = − − − − − − ( 1) 1 1 ( 1)( ) { , , , ,..., , }i i m i m i i i m i mcontext cell id cell id cell id cell id cell id cell id cell id− − − − + + − +− = − − − − − − where m is the window size. figure 2 shows a case when the window size is set as five. we applied the skip-gram model (mccormick, 2016), which is part of word2vec (mikolov et al., 2013), to build word representations (embeddings) that can predict the surrounding context words for a given target word. skipgram model can help capture the semantic relationships between words, discovering the geo-functional similarity relationship between sites in the migration network. the model takes a target word as input and predicts the previous and following words expected to appear in the context. word prediction is conducted using a user-specified window size. model training involves adjusting the word embeddings to maximize the probability of correctly predicting the context words surrounding the target word. here, we defined the object of the skip-gram model to maximize the average log probability as: log ( ( ) | ) cell id a l p context cell id cell id − ∈ = − −∑ where a is all the words in the document corpus, context( cell id− ) is the context of the word cell id− with its size 2 1m + .we then moved a contextual window of length 2 1m + across the documents to maximize the co-occurrence probability of the words that appeared within a window. we assumed that the sequence of words was identically distributed and independent, where the probability of its contextual words was: ( ) context( ) (context( ) ) ccell id c cell id cell id cell idp cell id cell id p ∈− − = −− −− ∏∣ ∣ where ccell id− is a contextual word of cell id− . the probability of ( )ccell id cell idp − −∣ was calculated with softmax function: ( )| cell id cell id t c c t u u w p cell i d v vd cell i v v − ∈ − =− − ∑ where , ,c ucell id cell id cell id− − − are the vectors of word , ,c ucell id cell id cell id− − − , w denote the set of all words. ucell id− is one of the total words, and “t” is the transposition for the vector of ucell id− , then the dot products of the vector cell id− and cell id− can be calculated. in the skip-gram model, the computational complexity of the output layer was high because it required calculating the probability distribution over the entire vocabulary. here, the vocabulary represented the entire datasets of cell-ids, which was large and in turn, time-consuming. we applied a negative sampling to selectively update model parameters (mikolov et al., 2013), allowing the model to update only a small subset of parameters in each training step. the negative samples accelerated the training process. for the target word represented by cell id− in the formula, we selected its contextual words context(p) as positive samples and p words that did not belong to context( cell id− ) as negative samples (zhu et al., 2019). a logistic regression was applied to train the model using positive and negative samples, aiming to maximize the prediction probability for positive samples while minimizing it for negative samples. we defined the object function as: ( ) ( ) l = log ( ) log(1 ( )) p t t x cell id p cell id cell id a x context cell id cell id neg cell id v v v vσ σ− − − ∈ ∈ − − ∈ −   + −     ∑ ∑ ∑ figure 2. a diagram for the context of a word with the window size is five. the context of a word icell id− with its window size is five can be defined as: 5 1 1 5( ) { ,..., , ,..., }i i i i icontext cell id cell id cell id cell id cell id− − + +− = − − − − , where a contextual word ( ccell id ) can be one of 5icell id −− , 4icell id −− , 3icell id −− , 2icell id −− , 1icell id −− , 1icell id +− , 2icell id +− , 3icell id +− , 4icell id +− , 5icell id +− . shi feng et al. – stopover hotspots in north and central america 12 where: 1(x)= (1 exp( ))x σ + − , ,x p cell idv v v − is the vectors of word , ,x pcell id cell id cell id− − − . we use xcell id− to present a word included in the context of cell id− and use pcell id− to present a word not included in the context of cell id− . “t” also presents the transposition of the vector for dot products between vectors. finally, a high-dimensional vector for each word was got after optimizing the object function on the trajectory documents. the similarity of two vectors ,a bv v was then calculated with the cosine distance in the vector space as: ( ) 2 2 similarity , a b a b a b v vv v v v ⋅ = ⋅   the cosine distance calculation results for similarity representation between each two high-dimensional vectors in the trajectory corpus were shown in the heat map for geographic interpretation. geo-functional similarity mining was done by finding the two cell-ids closest to the important stopover sites (hotspots got through pagerank algorithm) in the vector space with the minimum cosine distances as the potential alternative stopover sites. when applying word2vec for vector presentation, the following two parameter settings need to be noted: (1) vector dimensionality: low dimensionality usually means information loss. we set 100 as the embedding dimension to facilitate information preserving and compactness according to empirical knowledge and experiments. (2) window size: a larger window tends to capture more overall information, and a smaller window gets local syntactic contexts (levy and goldberg, 2014). for mobility data, a larger window reveals mobility behaviors, and a smaller window captures geographical similarity (zhu et al., 2019). according to the length of trajectories, we set the window size as 5 to focus on the local geographical similarity of neighbor cells. analytical framework the overall analytical framework is described in figure.3. migratory trajectories were estimated from field occurrence data in the form of geographic coordinates from the ebird repository (https://ebird.org/; feng et al., 2021). geographic coordinates of migration trajectories were analyzed as spatio-temporal words using a situation-aware figure 3. overall structure for the method. (a) trajectory estimation based on observations from ebird. the modelling process begins with estimating bird migration trajectories based on citizen-science observation data stored in ebird platform. the data collection forms the foundation for understanding the migration paths of birds. (b) geographic partitioning for migration network construction. after estimating the trajectories, the migration routes are geographically partitioned, and a migration network can be constructed between different cells which form nodes, and edges represent the migration paths between these cells. (c) stopover hotspots mining based on pagerank algorithm. the pagerank algorithm is then applied to the migration network to identify important stopover hotspots. pagerank ranks nodes based on their linkage quantity and quality, thereby identifying critical nodes, especially key stopover sites in the migration network. (d) potential alternative sites exploration based on word embedding technique. a word embedding technique (word2vec) is used to explore potential alternative sites that have similar geo-functional properties to the identified stopover hotspots. this technique identifies locations that could serve as alternative stopover sites based on spatial and functional similarities. (e) stopover protection prioritization guidance. stopover protection prioritization can be inferred based on the stopover hotspots and alternative stopover sites. right: (a)-(b) preparation. the first two steps are part of the data preparation phase, which focuses on trajectory estimation and preparation for network construction. (c) stopover hotspots based on pagerank. the third step utilizes the pagerank algorithm to rank and identify stopover hotspots. (d) geo-functional similarity sites based on word2vec. the fourth step explores alternative sites with similar geo-functional characteristics using the word2vec technique. (e) protection guidance. the final step is for stopover protection prioritization guidance based on sites predicted. shi feng et al. – stopover hotspots in north and central america 13 analysis (zhu et al., 2019). the situation-aware analysis assessed dynamic bird mobility across a grid coordinate system of continental north and central america after geographic partition. trajectory coordinates were converted to cell-id (the number for each cell) sequences. cell-id sequences were used to build a migratory network with each cell as a node. stopover hotspots were identified using the pagerank algorithm (page et al., 1999) and potential stopover alternative hotspots were determined using a word embedding technique (mikolov et al., 2013). stopover hotspots and alternative stopover sites were projected on maps to identify areas for protection prioritization and analysis for migratory birds. stopover hotspots characteristics the effect of landscape configuration on the network was explored based on landscape variables and stopover hotspots. we overlaid the artificial light (https://www. nasa.gov/image-article/earth-night/), the topographic map (https://apps.nationalmap.gov/viewer/), and the world database on protected areas (https://www.protectedplanet.net/en/thematic-areas/wdpa?tab=wdpa) in north and central america for protection prioritization reference. and we used the coverage of the total species and trajectory numbers to evaluate the stopover hotspots importance, and the overlapped species to assess the results of geographic functional similarity. then we used the datasets of the global land cover estimation (glance; stanimirova et al., 2023) product and human footprint (hfp) datasets (mu et al., 2022) to analyze the land cover classes and the interference degree of human activities on the ecosystem about the stopover hotspot conditions quantitatively. finally, we combined the stopover hotspot areas with the important bird and biodiversity areas (ibas) in north and central america (donald et al., 2019; https://datazone.birdlife.org/) for protection strategy comparison. results migration network structure and pagerank value results we modeled 52 migratory species during the period from 2017 to 2019, and estimated 540 migration trajectories in 2017, 435 in 2018, and 435 in 2019. the migration network structure resulted in a total of 1410 sequences during the study period (fig.4), including their respective pagerank values at the cell level (fig.5). the top 10 important stopovers nodes with the highest pagerank values were: cells 237, 589, 554, 520, 519, 952, 271, 624, 236,and 918. potential alternative stopover sites a heatmap of cosine distance revealed calculation results between each two high-dimensional vectors of cells by word embedding technique (fig.6). we performed geo-functional similarity mining for the top 10 important stopover cell-ids (i.e., 237, 589, 554, 520, 519, 952, 271, 624, 236, 918) and the results were shown in table 1. figure 4. migration network structure for 2017-2019. each node can be located and show the in-degree and out-degree. take node 1299 in the black square for example, it represents cell with number 1299 in the geographic partition map (fig.1) with its in-degree and out-degree are 14. arrows show movement direction between nodes. inset: the in (out) degree can reflect the number of arrows move to (from) each node. the figure above can help us obtain the degree and structure information for each node in the migration network which works as the basis of pagerank algorithm. https://www.nasa.gov/image-article/earth-night/ https://www.nasa.gov/image-article/earth-night/ https://datazone.birdlife.org/ https://datazone.birdlife.org/ shi feng et al. – stopover hotspots in north and central america 14 figure 5. the pagerank value results. (a) the pagerank value curve of 1091 cell-ids included in the avian trajectory sequences, which is ordered from the largest to smallest. the y-axis is the pagerank values, and the x-axis is cell-ids; (b) the pagerank value curve for the top 10 important stopovers cells, where 10 cells: 237, 589, 554, 520, 519, 952, 271, 624, 236, 918 are included. similarly, the y-axis is the pagerank values, and the x-axis is the cell-ids of top 10 important stopovers. (a) curve for pagerank values of each cell (b) curve for the top 10 pagerank values of stopover cells figure 6. heatmap of cosine distance between vectors of cells based on word2vec. the image on the left is a partial eagle eye diagram of the right one for a clearer presentation. the color of the grid can show the similarity between two cells. the smaller the cosine distance is, the higher the similarity is. table 1. top 10 important stopovers and their potential alternative sites with cosine distances analysis for 2017-2019 top 10 important stopovers potential alternative stopover sites (cosine distance based) 237 236(0.0367) 165(0.0406) 589 659(0.1367) 624(0.1454) 554 519(0.1330) 624(0.1437) 520 519(0.0661) 485(0.0749) 519 520(0.0661) 485(0.1238) 952 917(0.0498) 987(0.0606) 271 237(0.0610) 236(0.0618) 624 659(0.1240) 694(0.1291) 236 235(0.0261) 237(0.0367) 918 883(0.0493) 952(0.0611) shi feng et al. – stopover hotspots in north and central america 15 post modeling interpretation the map for the stopover hotspots and their potential alternative sites in north and central america (fig.7) revealed three areas highlighted for higher protection prioritization: florida peninsula and its inland (fl), the region of central america (ca), and the region near puget sound (ps) rich in national parks and forests. the overlapped map with the artificial light map and topographic map can be found in fig.s2. in the stopover hotspots result assessment for 20172019, the cells with top 10 (i.e., 10/1091; 0.9%) pagerank values were extracted, covering 777 trajectories of 46 species (table s2). trajectories accounted for 88.5% (46/52) species and 55.1% (777/1410) trajectory sequences. independent data from 2020-2022 were used for model evaluation focusing on key stopover sites detection. the top 10 important cells covered 729 (729/1353; 53.9%) trajectories of 41 (41/52; 78.9% ) species (table s3). these important cells revealed that the sites during our threeyear study period also stand out as important stopovers in the test years. also, we assessed the results of geographic functional similarity through the overlapped species for 20172019. the number of overlapped species revealed that the more species overlap, the higher the similarity is (tables 2 and s2). data from 2020-2022 revealed the effectiveness of geo-functional similarity. that is, the similarity results revealed that the study period and the testing period share almost the same outcomes for potential alternative sites exploration (tables 3 and s3). furthermore, fig.8 showed different landcover classed in the three hotspots areas, including water, developed, barren/sparsely vegetated, tress, shrub and herbaceous with all the three areas having the most percentages of trees, where “ca” represented the central america area, “fl” represented the florida peninsula and its inland, and “ps” represented the region near puget sound. fig 9 showed the hfp levels of the stopover hotspots: (1) ps: the distribution was concentrated at a low value with its median about 10, which indicated that human interference was relatively low and stable; (2) the florida peninsula and its inland (fl): its median was about 17 with its maximum hfp value could be greater than 35, indicating that there was a rather strong human interference to parts of this area; (3) central america (ca): its median was about 15 locating between fl and ps indicating moderate interference levels. finally, the overlapping percentage between the stopover hotspot cells and ibas was 13/18 (72.2%, grids in red and yellow) with 5/18 (27.8%) not identified as ibas (fig.10). figure 7. stopover hotspots and potential alternative sites for them overlaid on the protected area map. the red solid circles represent the top 10 important stopovers acquiring the top 10 pagerank values. the orange circles represent potential alternative stopover sites. green polygons denote terrestrial and inland water protected areas. blue polygons represent marine protected areas. three areas are highlighted and recommended as higher protection prioritization sites for targeted attention, including florida peninsula and its inland, the region of central america, and the region near puget sound. table 2. top 10 important stopovers and their potential alternative sites with overlapped species analysis for 2017-2019 top 10 important stopovers (included species number) potential alternative stopover sites (overlapped species number) 237(10) 589(25) 554(19) 520(9) 519(15) 952(13) 271(12) 624(19) 236(10) 918(20) 236(8) 659(17) 519(13) 519(9) 520(9) 917(11) 237(9) 659(18) 235(3) 883(14) 165(2) 624(17) 624(14) 485(6) 485(7) 987(10) 236(9) 694(17) 237(8) 952(13) table 3. top 10 important stopovers and their potential alternative sites with overlapped species analysis for 2020-2022 top 10 important stopovers (included species number) potential alternative stopover sites (overlapped species number) 237(9) 589(20) 554(20) 520(14) 519(16) 952(11) 271(11) 624(19) 236(7) 918(18) 236(7) 659(17) 519(15) 519(11) 520(13) 917(10) 237(9) 659(18) 235(3) 883(17) 165(1) 624(18) 624(15) 485(7) 485(7) 987(10) 236(7) 694(15) 237(7) 952(10) shi feng et al. – stopover hotspots in north and central america 16 figure 8. landcover classes analysis with the glance product. “ca” represents the central america area, “fl” represents the florida peninsula and its inland, and “ps” represents the region near puget sound. the figure above can show different landcover classed in the three hotspots areas, including water, developed, barren/sparsely vegetated, tress, shrub and herbaceous. and all the three areas have the most percentages of trees. figure 9. human disturbance analysis with the human footprint dataset. the gray dotted lines represent the 0%, 1%, 5%, 25%, 50%, 75%, 95%, 99%, and 100% quantiles of the dataset which are 0.006, 0.316, 5.440, 10.511, 15.911, 24.520, 30.927, and 41.259 respectively. for the three stopover hotspots: (1) the region near puget sound (ps): the distribution is concentrated at a low value with its median about 10, which indicates that human interference is relatively low and stable; (2) the florida peninsula and its inland (fl): its median is about 17 with its maximum hfp value can be greater than 35, indicating that there is a rather strong human interference to parts of this area; (3) central america (ca): its median is about 15 locating between fl and ps indicating moderate interference levels. shi feng et al. – stopover hotspots in north and central america 17 discussion this study combined the booming of public citizen science observation data with the urgent need to understand bird migration (rosenberg et al., 2019). in this article, we presented a systematic method considering the network structure during migration for multi-species at the population level to help develop precision conservation strategies. effective stopover hotspots mining and potential alternative sites exploration can help manage migration network connectivity and decrease migration risks, especially for long-distance migratory birds (zurell et al., 2018). we applied the pagerank algorithm for stopover sites importance assessment considering both the linkage quantity and quality across the migration network for stopover hotspots discovery. we also explored the alternative sites for each hotspot based on geo-functional similarity using a word embedding technique. our study effectively mined both the spatiotemporal and semantic information of the trajectory sequences and offered effective and prospective guidance for protection prioritization decisions. our method employed 52 focal migratory species in north and central america during 2017-2019, and generated rigorous estimations for stopover hotspots and potential alternative sites during migration. species analysis for stopover hotspots and their potential alternative sites the larger number of species included in a stopover site can be used as a proxy of the high importance of the site. the more species overlapped between sites, the higher the similarity between the sites. that is to say, species sharing a stopover site could be a proxy of biogeographic analogy even in spatially disparate zones. and the species overlapped can indicate the similarity of geographic functions, which means areas attract more same species, the more similar their geographic functions is. we note that cell 237 acquires a higher pagerank value for its crucial location as the “bottleneck” of the american flyway. this area is around the central american isthmus, includes plains and central volcanic ridge covered with rich forests and some farmlands as well as both water and well-protected land resources, which can provide support and evidence for the inference of central america’s significant value for migratory bird network connectivity (bayly et al., 2018). detection of relevant cells can give refined and scientific targets for central america protection against potential threats from forest loss and expansion of land conversion (cohen et al., 2007). figure 10. the stopover hotspots overlapped with the ibas map. the overlapping percentage between the stopover hotspot cells and ibas is 13/18 (72.2%, grids in red and yellow) with 5/18 (27.8%) not identified as ibas. shi feng et al. – stopover hotspots in north and central america 18 stopover hotspots characteristics analysis with the growing attention to migratory bird protection, there is a great need for informed protection of stopover hotspots across migratory paths. we found fundamental stopover sites across a migration network and potential alternative sites for further management accounting for human activity. our finding mirrored other efforts for data-driven biodiversity conservation studies have explored relevant decision-making plans. for example, thomson et al. (2020) developed a landscape-scale spatial conservation action planning tool (scap) to offer advances for conservation actions in heterogenous landscapes. guo et al. (2024) discussed the relationship between protection status and light pollution with the stopover hotspots distribution in the eastern us. migratory birds need water to drink and feed. thus, areas close to water sources, such as coastlines, freshwater lakes, rivers, and wetlands, are often crucial stopovers (boere et al., 2006). different species of migratory birds have different habitat needs, thus, heterogenous habitats can attract more migratory species (tu et al., 2020). the three stopover hotspot areas (fl, ca, and the region near ps) were found close to water sources and food sources like forests. the stopover sites ca and ps are currently surrounded by terrestrial and marine protected areas. the regions around fl, however, are not under sufficient protection. considering the high-level population density and economic development (mu et al., 2022), it may be necessary to involve proper urban development planning and private lands involvement for improved stopover protection in the eastern us. results highlighted the importance of forest conservation for bird migration (guo et al., 2023; mehlman et al., 2005). also, the land cover of ca, fl and the region near ps is dominated by trees, calling for more attention while making bird conservation decisions (fig.8). hfp index levels of the region near ps suggest a lower level of human disturbance. ca shows moderate but widespread human presence. in contrast, fl shows a wide range of hfp values, with many cells experiencing high levels of anthropogenic impact, which needs more protection activities (fig.9). lastly, the three hotspot areas overlap with the protected areas in north and central america (fig.10, 72.2% overlap). mcclure et al. (2018) suggested enhancing cooperation through policy mechanisms that account for protected areas for raptor species along the central america area. kirby et al. (2008) called for attention to agricultural intensification, human infrastructure development, and forest protection based on protected areas for migratory landbird and waterbird species. it is unclear whether birds select migration stopovers in protected areas or if protected areas are established in sites recognized as important for seasonal bird congregation. future directions one important limitation of our work is the limited capacity to detect areas with small pagerank values that are still relevant to migration connectivity. for example, some endangered species could have low ebird records resulting in weak signal of stopover hotspots, but could be the species in higher need of management and protection, such as kirtland’s warbler (setophaga kirtlandii) breeding in the great lakes. areas not covered by ebird observations are not assessed and, instead, oversampled areas provide more information at the cost of bias. assisting with individual tracking data, human movement data or other observation datasets (ai et al., 2024) could complement our stopover hotspots map. additionally, pagerank algorithm and word embedding technique only consider the link relationship between nodes, and ignored other information, such as human and bird behavioral data. finally, ebird data has the biases and limitations that citizen science data have, such as the bias from regional differences in reporting rates that depend heavily on species’ overlap with ebird users’ activities (robinson et al., 2022). to address this data limitation, information from different databases and additional taxa during longer periods would help detect and mitigate systematic bias from a single data platform. conclusion our novel analytical approach to access migratory patterns revealed important stopover nodes in the migratory network for a comprehensive range of bird taxa in north and central america. we extracted important stopover signals based on the network-level evaluation method by pagerank algorithm, and explored alternative stopover sites that accounted for spatial and temporal context information of trajectory sequences between sites with word embeddings, avoiding studying the sites as isolated ones. targeted protection and recharging of areas near water and forests were recommended in the sites identified as hot stopover sites of bird migration, especially in developed land in the eastern united states. discoveries and analytical approaches presented here could be useful for migratory connectivity protection in the full migration cycle and inform multinational conservation prioritization. data and code availability data and code used to perform this study are available at github (https://github.com/ash0920-git/stopover-hotspots). shi feng et al. – stopover hotspots in north and central america 19 acknowledgements this work is supported by the national key r&d program of china (2022yff0802300), national natural science foundation of china (u21a20478), zhejiang high-level talents special support program (2021r52002). lee was supported by national science foundation career (2235295) and hegs (2116748) awards, nih k01ai168452 award, virginia tech da ppp, cezap, and ictas grants, and the chinese academy of sciences pifi project 2024pvc0085. the content is solely the responsibility of the authors and does not necessarily represent the official views of the national institutes of health. competing interests the authors have declared that no competing interests exist. references ai, y., zhai, h., sun, z., yan, w., and hu, t., 2024. flockseer: a portable stereo vision observer for bird flocking. iet cyber syst. robot. 6(1): e12118. https://doi. org/10.1049/csy2.12118 artificial light map, 2024. https://www.nasachina.cn/ apod/19232.html (accessed june 14, 2024) bayly, n. j., rosenberg, k. v., easton, w. e., gomez, c., carlisle, j. a. y., ewert, d. n., and goodrich, l., 2018. major stopover regions and migratory bottlenecks for nearctic-neotropical landbirds within the neotropics: a review. bird conserv. int. 28(1): 1-26. https://doi.org/10.1017/s0959270917000296 berger, m., mcdonough, k., and seversky, l. m., 2016. cite2vec: citation-driven document exploration via word embeddings. ieee trans. vis. comput. graph. 23(1): 691-700. https://doi.org/10.1109/ tvcg.2016.2598667 boere, g. c., galbraith, c. a., and stroud, d. a. (eds.)., 2006. waterbirds around the world: a global overview of the conservation, management and research of the world’s waterbird flyways. bonter, d. n., gauthreaux jr, s. a., and donovan, t. m., 2009. characteristics of important stopover locations for migrating birds: remote sensing with radar in the great lakes basin. conserv. biol. 23(2): 440-448. https://doi.org/10.1111/j.1523-1739.2008.01085.x catry, t., correia, e., gutiérrez, j. s., bocher, p., robin, f., rousseau, p., and granadeiro, j. p., 2024. low migratory connectivity and similar migratory strategies in a shorebird with contrasting wintering population trends in europe and west africa. sci. rep. 14(1): 4884. https://doi.org/10.1038/s41598-024-55501-y cohen, e. b., barrow jr, w. c., buler, j. j., deppe, j. l., farnsworth, a., marra, p. p., and moore, f. r., 2017. how do en route events around the gulf of mexico influence migratory landbird populations? condor ornithol. appl. 119(2): 327-343. https://doi.org/10.1650/ condor-17-20.1 derpanis, k. g., 2005. mean-shift clustering. lecture notes. [online]. available:http://www.cse.yorku. ca/%7ekosta/compvis_notes/mean_shift.pdf (accessed 10 march 2023) donald, p. f., fishpool, l. d., ajagbe, a., bennun, l. a., bunting, g., burfield, i. j., ... and wege, d. c., 2019. important bird and biodiversity areas (ibas): the development and characteristics of a global inventory of key sites for biodiversity. bird conserv. int. 29(2), 177-198. https://doi.org/10.1017/ s0959270918000102 ebird, 2023. https://ebird.org/ (accessed april 22, 2023) feng, s., yang, q., hughes, a. c., chen, j., and qiao, h., 2021. a novel method for multi-trajectory reconstruction based on lomct for avian migration in popu-lation level. eco. inform. 63: 101319. https:// doi.org/10.1016/j.ecoinf.2021.101319 gazoni, e. and clark, c, 2023. openpyxl 3.1.2. available from: https://pypi.org/project/openpyxl/ (accessed 10 feb 2024) gbif. available from https://www.gbif.org/ (accessed 20 august 2023) grisel, o., mueller,a., … and eren.k., 2022. scikit-learn/ scikit-learn: scikit-learn 1.1. 2. available from: https:// zenodo.org/records/6968622 (accessed 10 feb 2023) guo, f., buler, j. j., smolinsky, j. a., and wilcove, d. s., 2023. autumn stopover hotspots and multiscale habitat associations of migratory landbirds in the eastern united states. proc. natl. acad. sci. 120(13): e2203511120. https://doi.org/10.1073/ pnas.2203511120 guo, f., buler, j. j., smolinsky, j. a., and wilcove, d. s., 2024. seasonal patterns and protection status of stopover hotspots for migratory landbirds in the eastern united states. curr. biol. 34(1): 235-244. https://doi. org/10.1016/j.cub.2023.11.033 harris, c.r., millman, k.j., van der walt, s.j. et al., 2020. array programming with numpy. nature 585, 357– 362. https://doi.org/10.1038/s41586-020-2649-2 hastie, t. j., and tibshirani, r. j., 1990. generalized additive models (vol. 43). crc press, london. higuchi, h., 2012. bird migration and the conservation of the global environment. j. ornithol. 153(suppl 1), 3-14. https://doi.org/10.1007/s10336-011-0768-0 shi feng et al. – stopover hotspots in north and central america 20 hong, l., qian, y., gong, c., zhang, y., and zhou, x., 2023. improved key node recognition method of social network based on pagerank algorithm. comput. mater. continua. 74(1). https://doi.org/10.32604/ cmc.2023.029180 hunter, j.d. 2007. matplotlib: a 2d graphics environment. comput. sci. eng. 9 (3), 90-95. https://doi.org/ 10.1109/mcse.2007.55 jin, y. t., you, j., wakamiya, s., and kwon, h. y., 2024. analyzing user reactions using relevance between location information of tweets and news articles. epj data sci. 13(1): 44. https://doi.org/10.1140/epjds/ s13688-024-00465-2 kaiser, a., 1999. stopover strategies in birds: a review of methods for estimating stopover length. bird study. 46 (supplement): s299-s308. https://doi. org/10.1080/00063659909477257 kirby, j. s., stattersfield, a. j., butchart, s. h., evans, m. i., grimmett, r. f., jones, v. r., ... and newton, i., 2008. key conservation issues for migratory landand waterbird species on the world’s major flyways. bird conserv. int., 18(s1), s49-s73. https://doi. org/10.1017/s0959270908000439 knight, s. m., bradley, d. w., clark, r. g., gow, e. a., bélisle, m., berzins, l. l., ...and norris, d. r., 2018. constructing and evaluating a continent-wide migratory songbird network across the annual cycle. ecol. monogr. 88(3): 445-460. https://doi.org/10.1002/ ecm.1298 kramer, g. r., andersen, d. e., buehler, d. a., wood, p. b., peterson, s. m., lehman, j. a., ... and streby, h. m., 2018. population trends in vermivora warblers are linked to strong migratory connectivity. proc. natl. acad. sci. 115(14): e3192-e3200. https://doi. org/10.1073/pnas.1718985115 la sorte, f. a., fink, d., hochachka, w. m., and kelling, s., 2016. convergence of broad-scale migration strategies in terrestrial birds. proc. r. soc. b. 283(1823): 20152588. https://doi.org/10.1098/rspb.2015.2588 levy, o., and goldberg, y., 2014. dependency-based word embeddings. proc. 52nd annu. meet. assoc. comput. linguist. https://doi.org/10.3115/v1/p14-2050 lin, h. y., schuster, r., wilson, s., cooke, s. j., rodewald, a. d., and bennett, j. r., 2020. integrating season-specific needs of migratory and resident birds in conservation planning. biol. conserv. 252: 108826. https://doi.org/10.1016/j.biocon.2020.108826 liu, f., zheng, l., and zheng, j., 2020. hienn-dwe: a hierarchical neural network with dynamic word embeddings for document-level sentiment classification. neurocomputing. 403: 21-32. https://doi. org/10.1016/j.neucom.2020.04.084 machin, j., 2017. xlwt 1.3.0. available from: https://pypi. org/project/xlwt/ (accessed 12 feb 2022) marín, s. d., 2018. pygam 0.8.0. available from: https:// pypi.org/project/pygam/ (accessed 11 feb 2022) matlab, r2021b. version 9.11.0.1873467, natick, massachusetts: the mathworks inc. mcclure, c. j., westrip, j. r., johnson, j. a., schulwitz, s. e., virani, m. z., davies, r., ... and butchart, s. h., 2018. state of the world’s raptors: distributions, threats, and conservation recommendations. biol. conserv., 227, 390-402. https://doi.org/10.1016/j.biocon.2018.08.012 mccormick, c., 2016. word2vec tutorial-the skipgram model. [online]. available: http://mccormickml.com/2016/04/19/word2vec-tutorial-the-skip-gram-model mehlman, d. w., mabey, s. e., ewert, d. n., duncan, c., abel, b., cimprich, d., ... and woodrey, m., 2005. conserving stopover sites for forest-dwelling migratory landbirds. the auk, 122(4), 1281-1290. https:// doi.org/10.1093/auk/122.4.1281 mikolov, t., chen, k., corrado, g., and dean, j., 2013. efficient estimation of word representations in vector space. arxiv preprint arxiv:1301.3781. https://doi. org/10.48550/arxiv.1301.3781 mikolov, t., sutskever, i., chen, k., corrado, g. s., and dean, j., 2013. distributed representations of words and phrases and their compositionality. neurips. 26. https://doi.org/10.48550/arxiv.1310.4546 mu, h., li, x., wen, y., huang, j., du, p., su, w., ... and geng, m., 2022. a global record of annual terrestrial human footprint dataset from 2000 to 2018. sci. data, 9(1), 176. https://doi.org/10.1038/s41597-02201284-8 nemes, c. e., cabrera-cruz, s. a., anderson, m. j., degroote, l. w., desimone, j. g., massa, m. l., and cohen, e. b., 2023. more than mortality: consequences of human activity on migrating birds extend beyond direct mortality. ornithol. appl. 125(3), duad020. https://doi.org/10.1093/ornithapp/duad020 nicol, s., cros, m. j., peyrard, n., sabbadin, r., trépos, r., fuller, r. a., and woodworth, b. k., 2023. flywaynet: a hidden semi-markov model for inferring the structure of migratory bird networks from count data. methods ecol. evol. 14(2): 265-279. https://doi. org/10.1111/2041-210x.14011 page, l., brin, s., motwani, r., and winograd, t., 1999. the pagerank citation ranking: bringing order to the web. stanford infolab. [online]. available: http://ilpubs.stanford.edu:8090/422/1/1999-66.pdf python 3.10. python software foundation, 2021. shi feng et al. – stopover hotspots in north and central america 21 qian, m., liu, j., li, c., and pals, l., 2019. a comparative study of english-chinese translations of court texts by machine and human translators and the word2vec based similarity measure’s ability to gauge human evaluation biases. proc. mt summit xvii: translator, project and user tracks. pp: 95-100. https://aclanthology.org/w19-6714.pdf reback, j., jbrockmendel, mckinney, w., van den bossche, j., roeschke, m., … and augspurger, t. pandas dev/pandas: pandas1.4.3. available from: https://zenodo. org/record/3509134 (accessed 2 feb 2023) rehurek, r., 2023. gensim 4.3.1. available from: https:// pypi.org/project/gensim/ (accessed 2 feb 2024) robinson, o. j., socolar, j. b., stuber, e. f., auer, t., berryman, a. j., boersch-supan, p. h., ... and johnston, a., 2022. extreme uncertainty and unquantifiable bias do not inform population sizes. proc. natl. acad. sci. 119(10): e2113862119. https://doi.org/10.1073/ pnas.2113862119 rogers, i., 2002. the google pagerank algorithm and how it works. [online]. available: https://cs.wmich.edu/ gupta/teaching/cs3310/lecturenotes_cs3310/pageran%20explained%20correctly%20with%20examples_www.cs.princeton.ed_~chazelle_courses_bib_ pagerank.pdf rosenberg, k. v., dokter, a. m., blancher, p. j., sauer, j. r., smith, a. c., smith, p.a., and marra, p. p., 2019. decline of the north american avifauna. science, 366(6461):120-124. https://doi.org/10.1126/science. aaw1313 runge, c. a., watson, j. e., butchart, s. h., hanson, j. o., possingham, h. p., and fuller, r. a., 2015. protected areas and global conservation of migratory birds. science, 350(6265): 1255-1258. https://doi.org/10.1126/ science.aac9180 schrimpf, m. b., des brisay, p. g., johnston, a., smith, a. c., sánchez-jasso, j., robinson, b. g., and koper, n., 2021. reduced human activity during covid-19 alters avian land use across north america. sci. adv. 7(2): eabf5073. https://doi.org/10.1126/sciadv. abf5073 somveille, m., bay, r. a., smith, t. b., marra, p. p., and ruegg, k. c., 2021. a general theory of avian migratory connectivity. ecol. lett. 24(9): 1848-1858. https://doi.org/10.1111/ele.13817 stanimirova, r., tarrio, k., turlej, k., mcavoy, k., stonebrook, s., hu, k. t., ... and friedl, m. a., 2023. a global land cover training dataset from 1984 to 2020. sci. data, 10(1), 879. https://doi.org/10.1038/s41597023-02798-5 thomson, j., regan, t. j., hollings, t., amos, n., geary, w. l., parkes, d., ... and white, m., 2020. spatial conservation action planning in heterogeneous landscapes. biol. conserv. 250, 108735. https://doi. org/10.1016/j.biocon.2020.108735 tu, h. m., fan, m. w., and ko, j. c. j., 2020. different habitat types affect bird richness and evenness. sci. rep., 10(1), 1221. https://doi.org/10.1038/s41598020-58202-4 usgs geographic map, 2024. https://apps.nationalmap. gov/viewer/ (accessed june 17, 2024) wang, h., hernandez, j. m., and van mieghem, p., 2008. betweenness centrality in a weighted network. phys. rev. e stat. nonlin. soft matter phys., 77(4), 046105. https://doi.org/10.1103/physreve.77.046105 wang, m., yang, s., sun, y., and gao, j., 2017. discovering urban mobility patterns with pagerank based traffic modeling and prediction. physica a: stat. mech. appl. 485: 23-34. https://doi.org/10.1016/j. physa.2017.04.155 webster, m. s., marra, p. p., haig, s. m., bensch, s., and holmes, r. t., 2002. links between worlds: unraveling migratory connectivity. trends ecol. evol. 17(2), 76-83. https://doi.org/10.1016/s01695347(01)02380-1 whitaker, j., 2022. basemap 1.3.4. available from: https:// pypi.org/project/basemap/ (accessed 3 feb 2023) whitaker, j., 2022. pyproj 3.3.1. available from: https:// zenodo.org/record/3509134 (accessed 7 feb 2023) withers, c., 2020. xlrd 2.0.1. available from: https://pypi. org/project/basemap/ (accessed 13 feb 2022) world database on protected areas, 2024. https:// www.protectedplanet.net/en/thematic-areas/wdpa?tab=wdpa (accessed may 15, 2024) wyborn, c., and evans, m. c., 2021. conservation needs to break free from global priority mapping. nat. ecol. evol. 5(10): 1322-1324. https://doi.org/10.1038/ s41559-021-01540-x xiong, s., lv, h., zhao, w., and ji, d., 2018. towards twitter sentiment classification by multi-level sentiment-enriched word embeddings. neurocomputing, 275:2459-2466. https://doi.org/10.1016/j.neucom.2017.11.023 xu, y., si, y., takekawa, j., liu, q., prins, h. h., yin, s., ... and de boer, w. f., 2020. a network approach to prioritize conservation efforts for migratory birds. conserv. biol. 34(2): 416-426. https://doi.org/10.1111/ cobi.13383 yorav-raphael, n. and and da costa-luis, c., 2024. tgdm: a fast, extensible progress bar for python and cli, https://github.com/tqdm/tgdm, version 4.65.0. https://doi.org/10.1111/ele.13817 https://pypi.org/project/basemap/ https://pypi.org/project/basemap/ https://zenodo.org/record/3509134 https://zenodo.org/record/3509134 https://pypi.org/project/basemap/ https://pypi.org/project/basemap/ https://doi.org/10.1038/s41559-021-01540-x https://doi.org/10.1038/s41559-021-01540-x https://doi.org/10.1016/j.neucom.2017.11.023 https://doi.org/10.1016/j.neucom.2017.11.023 https://doi.org/10.1111/cobi.13383 https://doi.org/10.1111/cobi.13383 shi feng et al. – stopover hotspots in north and central america 22 zhang, t. y., and wang, x. l., 2011. outlier detection algorithm based on space local deviation factor. comput. eng. 14. https://doi.org/10.3969/j.issn.10003428.2011.14.096 (in chinese with english abstract) zhang, w., wei, j., and xu, y., 2023. prioritizing global conservation of migratory birds over their migration network. one earth, 6(11): 1340-1349. https://doi. org/10.1016/j.oneear.2023.08.017 zhu, m., chen, w., xia, j., ma, y., zhang, y., luo, y., and liu, l., 2019. location2vec: a situation-aware representation for visual exploration of urban locations. ieee trans. intell. transp. syst. 20(10): 3981-3990. https://doi.org/10.1109/tits.2019.2901117 zivot, e., and wang, j., 2007. modeling financial time series with s-plus® (vol. 191). springer science & business media. zumstein,f.,2023. xlwings 0.30.6. available from: https://pypi.org/project/xlwings/ (accessed 13 feb 2024) zurell, d., graham, c. h., gallien, l., thuiller, w., and zimmermann, n. e., 2018. long-distance migratory birds threatened by multiple independent risks from global change. nat. clim. chang. 8(11): 992-996. https://doi.org/10.1038/s41558-018-0312-9 https://doi.org/10.1016/j.oneear.2023.08.017 https://doi.org/10.1016/j.oneear.2023.08.017 https://pypi.org/project/xlwings/ shi feng et al. – stopover hotspots in north and central america 23 figure s1. frequency distribution of minimum stopover length during autumn migration periods figure s2. stopover hotspots overlapped with the artificial light map and topographic map (a) stopover hotspots overlapped with the artificial light map (b) stopover hotspots overlapped with the topographic map shi feng et al. – stopover hotspots in north and central america 24 table s1. species list and their habitats information* number species name 1 accipiter cooperii breeds in forested areas; more common in suburban areas. 2 aix sponsa found in wetlands and flooded woods. 3 anas platyrhynchos found anywhere with water, including city parks, backyard creeks, and various wetland habitats. 4 archilochus colubris readily comes to sugar water feeders and flower gardens. 5 ardea alba ponds, marshes, and tidal mudflats. 6 ardea herodias occurs in almost any wetland habitat, from small ponds to marshes to saltwater bays. 7 branta canadensis occurs in any open or wetland habitat. 8 bucephala albeola found in bays, estuaries, reservoirs, and lakes in winter. travels to boreal forest and nests in cavities in summer. 9 buteo jamaicensise dges of trees. 10 buteo lineatus often in forested areas. 11 catharus ustulatus breeds in the boreal forest. 12 charadrius vociferus often in fields with short grass or barren dirt. 13 colaptes auratus often seen feeding on the ground in open areas, foraging for ants and worms. 14 cyanocitta cristata pairs or small groups travel through mature deciduous or coniferous woodlands. 15 dumetella carolinensis especially thickets or second-growth at the edge of forests, often near water. 16 egretta thula found in a variety of wetland habitats, especially shallow marshy pools and mudflats. 17 fulica americana ponds, city parks, marshes, reservoirs, lakes, ditches, and saltmarshes. 18 geothlypis trichas found in shrubby wet areas, including marshes, forest edges, and fields. 19 haliaeetus leucocephalusnear near bodies of water. 20 hirundo rustica especially large fields and wetlands. 21 icterus galbula breeds in deciduous trees in open woodlands, forest edges, orchards, riversides, parks, and backyards. 22 larus delawarensis found along lakes, rivers, ponds, and beaches. 23 leiothlypis celata found in scrubby areas, woodland edge, and thickets. shi feng et al. – stopover hotspots in north and central america 25 number species name 24 leiothlypis ruficapilla breeds in coniferous or mixed forests. 25 mareca strepera typically found in pairs or small flocks in shallow wetlands, ponds, or bays. 26 megaceryle alcyon edges of streams, lakes, and estuaries. 27 melospiza melodia especially edges of fields, often near water. 28 molothrus aterwoods farmland, and stockyards. 29 myiarchus crinitus deciduous forests. 30 pandion haliaetus on top of channel markers, utility poles and high platforms near water. 31 passerina cyaneathe edge of forests and fields. 32 pheucticus ludovicianus especially in deciduous forests. 33 pipilo erythrophthalmus inhabits scrubby areas and forest edges. 34 podilymbus podiceps occurs on ponds and marshes. 35 polioptila caerulea breed in deciduous woodlands, often near water. 36 quiscalus quiscula forages in fields, scrubby areas, and open woods. 37 sayornis phoebe woodland edge, brushy fields, or edges of ponds. 38 setophaga americana breeds in mature coniferous or deciduous forests, especially near water. 39 setophaga coronata mixed forests, often near clearings or edges. in migration and winter, found in any woodland or open shrubby area, including coastal dunes, fields, parks, and residential areas. 40 setophaga palmarum breeds in bogs and clearings in the boreal forest. 41 setophaga petechia near water, often foraging in shrubs fairly low to the ground. 42 sialia sialis favors fields and open woods. 43 spatula discors usually found in shallow wetlands or marshes. 44 spinus tristis found in weedy fields, cultivated areas, roadsides, orchards, and backyards. 45 spizella passerina usually found in open woodlands, scrubby areas, or even in suburban settings. 46 stelgidopteryx serripennis often seen near water, sometimes in mixed flocks with other swallows. shi feng et al. – stopover hotspots in north and central america 26 number species name 47 sturnus vulgaris often abundant, gathering in large flocks in open agricultural areas and towns and cities. 48 toxostoma rufum shrubby habitats, especially second-growth woodland, thickets, and forest edge. 49 troglodytes aedon open or semiopen habitats, including suburbs, parks, rural farmland, and woodland edge with thick tangles. 50 turdus migratorius in gardens, parks, yards, golf courses, fields, pastures, and many other wooded habitats. 51 zenaida macroura found in a variety of habitats from agricultural fields to lightly wooded areas. 52 zonotrichia leucophrys breeds in brushy areas or thickets in open forest, often with conifers. *the habitat information is collected from ebird (https://ebird.org/). shi feng et al. – stopover hotspots in north and central america 27 table s2. cell-ids with the included species for 2017-2019 top 10 important stopovers potential alternative stopover sites cell-id species cell-id species cell-id species 237 catharus ustulatus egretta thula hirundo rustica myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia sialia sialis spatula discors ardea alba 236 catharus ustulatus hirundo rustica myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia sialia sialis spatula discors troglodytes aedon icterus galbula 165 pheucticus ludovicianus setophaga petechia 589 archilochus colubris ardea alba buteo jamaicensis dumetella carolinensis egretta thula geothlypis trichas leiothlypis ruficapilla myiarchus crinitus pandion haliaetus passerina cyanea pheucticus ludovicianus podilymbus podiceps polioptila caerulea quiscalus quiscula sayornis phoebe setophaga americana setophaga palmarum spatula discors stelgidopteryx serripennis bucephala albeola catharus ustulatus charadrius vociferus icterus galbula hirundo rustica setophaga petechia 659 buteo lineatus geothlypis trichas icterus galbula myiarchus crinitus pandion haliaetus polioptila caerulea sayornis phoebe setophaga americana setophaga coronata setophaga palmarum dumetella carolinensis podilymbus podiceps quiscalus quiscula stelgidopteryx serripennis buteo jamaicensis catharus ustulatus hirundo rustica leiothlypis ruficapilla megaceryle alcyon pheucticus ludovicianus spizella passerina 624 buteo jamaicensis buteo lineatus geothlypis trichas leiothlypis ruficapilla myiarchus crinitus pandion haliaetus pheucticus ludovicianus podilymbus podiceps polioptila caerulea sayornis phoebe setophaga americana setophaga palmarum stelgidopteryx serripennis dumetella carolinensis icterus galbula spatula discors catharus ustulatus hirundo rustica quiscalus quiscula 554 archilochus colubris buteo jamaicensis dumetella carolinensis egretta thula geothlypis trichas myiarchus crinitus pandion haliaetus passerina cyanea pheucticus ludovicianus podilymbus podiceps polioptila caerulea setophaga americana setophaga palmarum anas platyrhynchos catharus ustulatus leiothlypis ruficapilla hirundo rustica setophaga petechia spatula discors 519 buteo jamaicensis fulica americana myiarchus crinitus pandion haliaetus passerina cyanea podilymbus podiceps setophaga americana setophaga palmarum catharus ustulatus egretta thula leiothlypis ruficapilla pheucticus ludovicianus stelgidopteryx serripennis hirundo rustica setophaga petechia 624 buteo jamaicensis buteo lineatus geothlypis trichas leiothlypis ruficapilla myiarchus crinitus pandion haliaetus pheucticus ludovicianus podilymbus podiceps polioptila caerulea sayornis phoebe setophaga americana setophaga palmarum stelgidopteryx serripennis dumetella carolinensis icterus galbula spatula discors catharus ustulatus hirundo rustica quiscalus quiscula shi feng et al. – stopover hotspots in north and central america 28 top 10 important stopovers potential alternative stopover sites cell-id species cell-id species cell-id species 520 myiarchus crinitus pandion haliaetus passerina cyanea podilymbus podiceps setophaga americana setophaga palmarum leiothlypis ruficapilla pheucticus ludovicianus setophaga petechia 519 buteo jamaicensis fulica americana myiarchus crinitus pandion haliaetus passerina cyanea podilymbus podiceps setophaga americana setophaga palmarum catharus ustulatus egretta thula leiothlypis ruficapilla pheucticus ludovicianus stelgidopteryx serripennis hirundo rustica setophaga petechia 485 buteo jamaicensis passerina cyanea podilymbus podiceps setophaga americana myiarchus crinitus pheucticus ludovicianus setophaga palmarum 519 buteo jamaicensis fulica americana myiarchus crinitus pandion haliaetus passerina cyanea podilymbus podiceps setophaga americana setophaga palmarum catharus ustulatus egretta thula leiothlypis ruficapilla pheucticus ludovicianus stelgidopteryx serripennis hirundo rustica setophaga petechia 520 myiarchus crinitus pandion haliaetus passerina cyanea podilymbus podiceps setophaga americana setophaga palmarum leiothlypis ruficapilla pheucticus ludovicianus setophaga petechia 485 buteo jamaicensis passerina cyanea podilymbus podiceps setophaga americana myiarchus crinitus pheucticus ludovicianus setophaga palmarum 952 aix sponsa anas platyrhynchos bucephala albeola geothlypis trichas larus delawarensis leiothlypis celata megaceryle alcyon melospiza melodia setophaga coronata spinus tristis turdus migratorius zonotrichia leucophrys branta canadensis 917 anas platyrhynchos bucephala albeola buteo lineatus geothlypis trichas larus delawarensis leiothlypis celata melospiza melodia setophaga coronata setophaga palmarum turdus migratorius zonotrichia leucophrys aix sponsa branta canadensis leiothlypis ruficapilla molothrus ater 987 aix sponsa bucephala albeola geothlypis trichas haliaeetus leucocephalus leiothlypis celata megaceryle alcyon melospiza melodia setophaga coronata turdus migratorius zonotrichia leucophrys branta canadensis 271 ardea alba catharus ustulatus hirundo rustica icterus galbula myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia sialia sialis spatula discors archilochus colubris spizella passerina 237 catharus ustulatus egretta thula hirundo rustica myiarchus crinitus pandion haliaetus pheucticusludovicianus setophaga petechia sialia sialis spatula discors ardea alba 236 catharus ustulatus hirundo rustica myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia sialiasialis spatula discors troglodytes aedon icterus galbula shi feng et al. – stopover hotspots in north and central america 29 top 10 important stopovers potential alternative stopover sites cell-id species cell-id species cell-id species 624 buteo jamaicensis buteo lineatus geothlypis trichas leiothlypis ruficapilla myiarchus crinitus pandion haliaetus pheucticus ludovicianus podilymbus podiceps polioptila caerulea sayornis phoebe setophaga americana setophaga palmarum stelgidopteryx serripennis dumetella carolinensis icterus galbula spatula discors catharus ustulatus hirundo rustica quiscalus quiscula 659 buteo lineatus geothlypis trichas icterus galbula myiarchus crinitus pandion haliaetus polioptila caerulea sayornis phoebe setophaga americana setophaga coronata setophaga palmarum dumetella carolinensis podilymbus podiceps quiscalus quiscula stelgidopteryx serripennis buteo jamaicensis catharus ustulatus hirundo rustica leiothlypis ruficapilla megaceryle alcyon pheucticus ludovicianus spizella passerina 694 buteo lineatus egretta thula geothlypis trichas icterus galbula megaceryle alcyon myiarchus crinitus pandion haliaetus pheucticus ludovicianus podilymbus podiceps polioptila caerulea sayornis phoebe setophaga americana setophaga coronata setophaga palmarum spizella passerina stelgidopteryx serripennis dumetella carolinensis pipilo erythrophthalmus quiscalus quiscula toxostoma rufum buteo jamaicensis catharus ustulatus hirundo rustica passerina cyanea sialia sialis 236 catharus ustulatus hirundo rustica myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia sialia sialis spatula discors troglodytes aedon icterus galbula 235 hirundo rustica catharus ustulatus setophaga petechia 237 catharus ustulatus egretta thula hirundo rustica myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia sialia sialis spatula discors ardea alba 918 aix sponsa anas platyrhynchos bucephala albeola geothlypis trichas larus delawarensis leiothlypis celata leiothlypis ruficapilla megaceryle alcyon melospiza melodia pandion haliaetus setophaga coronata sturnus vulgaris turdus migratorius zonotrichia leucophrys branta canadensis colaptes auratus spinus tristis ardea herodias mareca strepera molothrus ater 883 aix sponsa anas platyrhynchos ardea herodias bucephala albeola geothlypis trichas leiothlypis celata leiothlypis ruficapilla pandion haliaetus setophaga coronata setophaga palmarum sturnus vulgaris zonotrichia leucophrys branta canadensis melospiza melodia turdus migratorius molothrus ater 952 aix sponsa anas platyrhynchos bucephala albeola geothlypis trichas larus delawarensis leiothlypis celata megaceryle alcyon melospiza melodia setophaga coronata spinus tristis turdus migratorius zonotrichia leucophrys branta canadensis shi feng et al. – stopover hotspots in north and central america 30 table s3. cell-ids with the included species for 2020-2022 top 10 important stopovers potential alternative stopover sites cell-id species cell-id species cell-id species 237 catharus ustulatus myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia spatula discors hirundo rustica icterus galbula ardea alba 236 catharus ustulatus hirundo rustica icterus galbula pheucticus ludovicianus setophaga petechia myiarchus crinitus spatula discors 165 setophaga petechia 589 archilochus colubris buteo jamaicensis dumetella carolinensis geothlypis trichas myiarchus crinitus pandion haliaetus pheucticus ludovicianus polioptila caerulea quiscalus quiscula sayornis phoebe setophaga americana setophaga palmarum setophaga petechia spatula discors ardea herodias catharus ustulatus icterus galbula leiothlypis ruficapilla podilymbus podiceps stelgidopteryx serripennis 659 ardea herodias buteo lineatus dumetella carolinensis geothlypis trichas icterus galbula myiarchus crinitus pandion haliaetus pheucticus ludovicianus podilymbus podiceps polioptila caerulea quiscalus quiscula sayornis phoebe setophaga americana setophaga coronata setophaga palmarum setophaga petechia spizella passerina leiothlypis ruficapilla sialia sialis spatula discors stelgidopteryx serripennis toxostoma rufum egretta thula 624 buteo jamaicensis buteo lineatus dumetella carolinensis geothlypis trichas icterus galbula leiothlypis ruficapilla myiarchus crinitus pandion haliaetus pheucticus ludovicianus polioptila caerulea quiscalus quiscula sayornis phoebe setophaga americana setophaga palmarum setophaga petechia spatula discors ardea herodias podilymbus podiceps stelgidopteryx serripennis 554 archilochus colubris buteo jamaicensis dumetella carolinensis egretta thula myiarchus crinitus pandion haliaetus passerina cyanea pheucticus ludovicianus setophaga americana setophaga palmarum setophaga petechia spatula discors stelgidopteryx serripennis ardea herodias catharus ustulatus fulica americana icterus galbula leiothlypis ruficapilla polioptila caerulea quiscalus quiscula 519 archilochus colubris buteo jamaicensis leiothlypis ruficapilla myiarchus crinitus passerina cyanea podilymbus podiceps setophaga americana setophaga petechia stelgidopteryx serripennis ardea herodias catharus ustulatus pheucticus ludovicianus setophaga palmarum dumetella carolinensis fulica americana icterus galbula 624 buteo jamaicensis buteo lineatus dumetella carolinensis geothlypis trichas icterus galbula leiothlypis ruficapilla myiarchus crinitus pandion haliaetus pheucticus ludovicianus polioptila caerulea quiscalus quiscula sayornis phoebe setophaga americana setophaga palmarum setophaga petechia spatula discors ardea herodias podilymbus podiceps stelgidopteryx serripennis shi feng et al. – stopover hotspots in north and central america 31 top 10 important stopovers potential alternative stopover sites cell-id species cell-id species cell-id species 520 buteo jamaicensis egretta thula leiothlypis ruficapilla passerina cyanea setophaga americana setophaga petechia catharus ustulatus myiarchus crinitus podilymbus podiceps fulica americana icterus galbula megaceryle alcyon pandion haliaetus setophaga palmarum 519 archilochus colubris buteo jamaicensis leiothlypis ruficapilla myiarchus crinitus passerina cyanea podilymbus podiceps setophaga americana setophaga petechia stelgidopteryx serripennis ardea herodias catharus ustulatus pheucticus ludovicianus setophaga palmarum dumetella carolinensis fulica americana icterus galbula 485 buteo jamaicensis passerina cyanea setophaga americana setophaga petechia catharus ustulatus fulica americana pandion haliaetus 519 archilochus colubris buteo jamaicensis leiothlypis ruficapilla myiarchus crinitus passerina cyanea podilymbus podiceps setophaga americana setophaga petechia stelgidopteryx serripennis ardea herodias catharus ustulatus pheucticus ludovicianus setophaga palmarum dumetella carolinensis fulica americana icterus galbula 520 buteo jamaicensis egretta thula leiothlypis ruficapilla passerina cyanea setophaga americana setophaga petechia catharus ustulatus myiarchus crinitus podilymbus podiceps fulica americana icterus galbula megaceryle alcyon pandion haliaetus setophaga palmarum 485 buteo jamaicensis passerina cyanea setophaga americana setophaga petechia catharus ustulatus fulica americana pandion haliaetus 952 aix sponsa bucephala albeola buteo lineatus geothlypis trichas leiothlypis celata megaceryle alcyon melospiza melodia spinus tristis turdus migratorius zonotrichia leucophrys anas platyrhynchos 917 anas platyrhynchos branta canadensis bucephala albeola buteo lineatus leiothlypis celata melospiza melodia spinus tristis turdus migratorius zonotrichia leucophrys aix sponsa megaceryle alcyon 987 aix sponsa branta canadensis bucephala albeola colaptes auratus geothlypis trichas icterus galbula leiothlypis celata megaceryle alcyon melospiza melodia spinus tristis turdus migratorius zonotrichia leucophrys anas platyrhynchos haliaeetu leucocephalus larus delawarensis 271 catharus ustulatus hirundo rustica icterus galbula myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia spatula discors stelgidopteryx serripennis ardea alba spizella passerina 237 catharus ustulatus myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia spatula discors hirundo rustica icterus galbula ardea alba 236 catharus ustulatus hirundo rustica icterus galbula pheucticus ludovicianus setophaga petechia myiarchus crinitus spatula discors shi feng et al. – stopover hotspots in north and central america 32 top 10 important stopovers potential alternative stopover sites cell-id species cell-id species cell-id species 624 buteo jamaicensis buteo lineatus dumetella carolinensis geothlypis trichas icterus galbula leiothlypis ruficapilla myiarchus crinitus pandion haliaetus pheucticus ludovicianus polioptila caerulea quiscalus quiscula sayornis phoebe setophaga americana setophaga palmarum setophaga petechia spatula discors ardea herodias podilymbus podiceps stelgidopteryx serripennis 659 ardea herodias buteo lineatus dumetella carolinensis geothlypis trichas icterus galbula myiarchus crinitus pandion haliaetus pheucticus ludovicianus podilymbus podiceps polioptila caerulea quiscalus quiscula sayornis phoebe setophaga americana setophaga coronata setophaga palmarum setophaga petechia spizella passerina leiothlypis ruficapilla sialia sialis spatula discors stelgidopteryx serripennis toxostoma rufum egretta thula 694 ardea herodias buteo lineatus dumetella carolinensis egretta thula geothlypis trichas icterus galbula megaceryle alcyon myiarchus crinitus pandion haliaetus podilymbus podiceps polioptila caerulea quiscalus quiscula sayornis phoebe setophaga americana setophaga petechia spizella passerina leiothlypis ruficapilla setophaga coronata sialia sialis toxostoma rufum setophaga palmarum turdus migratorius 236 catharus ustulatus hirundo rustica icterus galbula pheucticus ludovicianus setophaga petechia myiarchus crinitus spatula discors 235 catharus ustulatus hirundo rustica icterus galbula 237 catharus ustulatus myiarchus crinitus pandion haliaetus pheucticus ludovicianus setophaga petechia spatula discors hirundo rustica icterus galbula ardea alba 918 aix sponsa ardea herodias branta canadensis bucephala albeola colaptes auratus geothlypis trichas megaceryle alcyon melospiza melodia pandion haliaetus spinus tristis sturnus vulgaris turdus migratorius zonotrichia leucophrys anas platyrhynchos leiothlypis celata leiothlypis ruficapilla molothrus ater setophaga coronata 883 aix sponsa anas platyrhynchos ardea herodias branta canadensis bucephala albeola colaptes auratus geothlypis trichas megaceryle alcyon molothrus ater spinus tristis sturnus vulgaris turdus migratorius zonotrichia leucophrys leiothlypis celata leiothlypis ruficapilla melospiza melodia setophaga coronata 952 aix sponsa bucephala albeola buteo lineatus geothlypis trichas leiothlypis celata megaceryle alcyon melospiza melodia spinus tristis turdus migratorius zonotrichia leucophrys anas platyrhynchos biodiversity informatics, 19, 2025, pp. 33-45 33 welcome to the machine: a pan-continental overview of machine learning applications in ecology and conservation juliano a. bogoni1*, derick victor de souza campos2*, claumir c. muniz2, manoel dos santos-filho1, jéssica eloá poletto3 1 universidade do estado de mato grosso, centro de pesquisa de limnologia, biodiversidade e etnobiologia do pantanal, programa de pós-graduação em ciências ambientais, laboratório de mastozoologia, cáceres, mt, brazil. 2 universidade do estado de mato grosso, centro de pesquisa de limnologia, biodiversidade e etnobiologia do pantanal, programa de pós-graduação em ciências ambientais, laboratório de investigação ambiental do pantanal norte – lipan, cáceres, mt, brazil. 3 faculdade de ciências médicas da universidade estadual de campinas, piracicaba, sp, brasil. abstract. machine-learning emerged as an excellent alternative to understanding ecological patterns and processes at different spatiotemporal scales. the study aimed to offer a global overview of the status quo on the use of machine-learning in ecology and conservation globally. using keywords in the scopus engine, we indexed all publications in ecology and conservation using machine-learning. we employed descriptive statistics and regressions models to provide an overview and predict geopolitical patterns. the majority of manuscripts were condensed in economically affluent countries, such as the united states (usa) and china (chn) which together amount to 91 (36.8%) studies. there is a spatial aggregation in the authors’ affiliations, once 182 (73.7%) studies derived from both nearctic and palearctic teams, whereas tropical teams published 65 (26.3%) manuscripts and the most-cited papers also are concentrated in northern regions. in ecology and conservation, machine-learning first appear in the literature in 2003. since then, the number of publications has increased exponentially, from 09 manuscripts in 2010, to 120 manuscripts 10 years later. most studies (n = 173; 70.1%) focused on landscape and vertebrate ecology. the primary aims of the publications were widely variable but strongly adherent to providing the best-information on both landscape-scale classifications and species distribution modelling. the manuscripts encompass different methods, from maximum entropy to boosted regression trees and random forest, sometimes using a range of deep-learning architectures. finally, the predictive variables (i.e., mammal diversity and per capita gdp) do not exert significant influences on the number of studies published. finally, we recommend a well-structured and collaborative agenda aiming to integrate less-resourced countries into scientific advancements, fostering more equitable and effective responses to global environmental challenges. keywords: data analysis, global-scale, informatics, numerical ecology, tropical forest. *corresponding authors: bogoni.ja@gmail.com | derick@unemat.br mailto:bogoni.ja@gmail.com mailto:derick@unemat.br juliano a. bogoni et al. – machine-learning applications in ecology and conservation 34 introduction ecology is a relatively young science that fundamentally seeks to understand the causes and consequences (i.e., processes) of diversity patterns and species distributions across global environments (brown, 1995; haeckel, 1866). universal features of ecological processes exhibit mathematical properties that are inherently non-linear and complex, historically addressed only through mathematical approximations (bogoni et al., 2019; conway, 1977; may, 1976). since the 1920s, explicit models as in volterra (1926), lotka (1925), elton (1924) have been employed in ecology to predict and describe synchronization mechanisms in animal behaviour as araujo et al. (2013), predator-prey relationships and dynamics in sherratt et al. (1997), kar et al. (2010), host-parasitoid interactions in hassell (2000), and various other ecological dynamics, such as seed predation and dispersal (bogoni et al., 2019) and species distribution (bogoni and tagliari, 2021; tagliari et al., 2023). as global biodiversity faces unprecedented and widespread declines in modern history (bogoni et al., 2020; ceballos et al., 2020), addressing contemporary ecological challenges—such as biodiversity loss, climate change, and the growing demand for ecosystem services—has become a pressing priority for ecologists (rammer and seidl, 2019). this issue is particularly critical in the global tropics, where habitat loss is most severe. for example, within just half a decade in the early 21st century, from 2000 to 2005, tropical forests in both wet and dry regions experienced a staggering loss of approximately 475,000 km² of forest cover (around 50%) (hansen et al., 2010). thus, understanding how to mitigate natural habitat degradation processes, particularly across tropical landscapes, is imperative for safeguarding biodiversity on a global scale. machine-learning methods have emerged as an excellent alternative for predicting and understanding ecological patterns and processes across various spatiotemporal scales. these methods consist of computational algorithms designed to analyse the often non-linear structures of complex data, thereby generating predictive models based on estimated patterns (rammer and seidl, 2019; olden et al., 2008). in comparison to classical statistical approaches (e.g., regression models), machine learning relies on computational power to identify and model complex relationships, prioritizing predictive accuracy over the estimation of parameters and confidence intervals (breiman, 2001). positioned at the intersection of statistics and computer science—and driven by advances in artificial intelligence and data science—the application of machine learning in ecology and conservation has emerged as a rapidly expanding field (rammer and seidl, 2019; jordan and mitchell, 2015). however, the conducted search reveals a lack of studies examining how machine-learning methods are being applied in ecology and conservation or providing a pan-continental overview of how these methods can assist ecologists in ensuring biodiversity persistence across global landscapes. additionally, disparities between local and global scales have historically biased scientific advancements (bogoni et al., 2021; hughes et al., 2021). modern ecological methods, such as machine learning, tend to be more widely utilized in regions with more affluent economies, where access to advanced computational resources is readily available (e.g., bogoni et al., 2021). in contrast, countries where scientific development is still in its early stages—and which often host the highest levels of biodiversity—have likely made less use of these computational advancements. this limitation is particularly concerning when it comes to understanding ecological patterns and processes and addressing the severe biodiversity crises in these regions. our primary aim was to provide a global pictorial overview of the use of machine learning methods in ecology and conservation by assembling a comprehensive database through an extensive literature review. specifically, we sought to: (1) quantitatively and qualitatively evaluate how machine learning methods have been applied in peer-reviewed ecology and conservation publications across a pancontinental scale; (2) assess the research questions, focal areas, and taxonomic groups addressed in studies involving machine learning within the disciplines of ecology and conservation; and (3) discuss how countries with lower investment in science or less frequent use of modern analytical methods can leverage these theoretical and analytical advancements to address their environmental and ecological challenges. we tested the following hypotheses: (1) the vast majority of publications originate from authors affiliated with economically affluent countries; (2) the application of machine learning in ecology and conservation is predominantly focused on species distribution models and landscape ecology; and (3) at a predictive level, a country’s per capita income positively correlates with the number of its publications, while mammal diversity (used as a proxy for country-scale biodiversity) is negatively correlated with publication output, as poorer countries often harbor greater biodiversity than wealthier nations. our goal is to provide a global synthesis of how machine learning has been integrated into ecologijuliano a. bogoni et al. – machine-learning applications in ecology and conservation 35 cal and conservation research. by identifying geographic, taxonomic, and methodological patterns in the literature, it discusses existing disparities in scientific output and highlights opportunities for greater inclusion of biodiversity-rich yet underrepresented countries. material and methods our study is characterized as data gathering research, with a posteriori application of plans with descriptive statistical analysis method and mapping. thus, in early may (04th-may-2023), we used the scopus search engine1 to index and aggregate all publications (published or in press) in ecology and conservation that used machine-learning methods. to do so, we used a conjunct of keywords allied to and/or operators, as follows: “machine learning” and/or “deep learning” and “ecology” and/ or “conservation (similar to magioli et al., 2015). this allowed us to compile the number of scientific publications that included these terms in their manuscript titles, abstracts, and keywords. after capturing the studies, we extracted the following information: (1) the total number of studies; (2) the publication year of each manuscript; (3) the taxonomic group(s) studied or the primary approach (e.g., landscape ecology); (4) the country of primary affiliation of the first author; (5) the city-based geolocation of the study in latitudinal and longitudinal decimal degrees; (6) the set of study keywords; (7) the study title; and (8) the total number of citations for each manuscript. documents with incomplete information, missing location, author, or keywords, were not consider in our analysis. once the database was consolidated, we used descriptive statistics and mapping to provide a pancontinental overview of the state of machine-learning applications in ecology and conservation. we analysed the number of published articles, the annual publication trends, and the distribution of articles by country. additionally, we generated a word-cloud based on the top-100 keywords provided by the authors. we also explored the proportion of taxonomic groups or approaches studied in a pictorial synthesis of studies. then, we identified the top-cited manuscripts (i.e., all those with more than 100 citations) and provided detailed insights into their primary aims, methods, and the specific sub-disciplines of ecology and conservation they addressed. the data analyses were performed in r v.4.0.5 (r core team, 2023) using stats (r core team, 2023), sp (bivand et al., 2013; pebesma and bivand, 2005), maptools (bivand and lewin-koh, 2021) and the dependent 1 https://www.scopus.com/home.uri. r-package. additionally, we utilized data from the international union for conservation of nature (iucn) red list of threatened species2 to obtain country-level mammal richness as a proxy for national biodiversity. mammals were chosen for this analysis as they represent one of the most extensively studied and well-documented taxonomic groups worldwide (bogoni et al., 2021). we also sourced data from the world bank3 to obtain country-level per capita gross domestic product, hereafter, gdp. both variables mentioned above were subsequently used to predict the number of studies published per country. to do so, we used a generalized linear regression model (glm), using poisson distribution under the data variance given that the response variable is a count (non-negative integers) (dobson, 1990), correcting the predictive data asymmetry by log x + 1. thus, the glm was performed as follows: studies(ith country) ~ log(mammalrichness(ith country )) + log(gdp(ith country)) where studies(ith country) is the number of studies published in the country, log(mammalrichness(ith country )) is the natural log of mammalian species richness in the country, and log(gdp(ith country)) is the natural log of the gross domestic product for country. then, deriving a bivariate effect plot for both predictive variables vs. number of studies. the statistics of glm model were presented for zand p-values, estimates and standard errors. the regression analysis and bivariate plots were performed in r v.4.0.5 (r core team, 2023) based on the stats r-package (r core team, 2023). results our literature search identified 251 documents related to machine learning and/or deep learning with a focus on ecology and conservation. after applying an intuitive filter to exclude documents with incomplete information, our study was able to analyse pancontinental trends based on 247 manuscripts. this overview revealed that the vast majority of published manuscripts were concentrated in economically affluent countries, such as the united states (usa) and china (chn), which together accounted for 91 studies (36.8%) (fig. 1a). furthermore, there was a notable spatial concentration in the primary affiliations of corresponding authors: 182 studies (73.7%) originated from teams in the nearctic and palearctic regions, particularly europe, while neotropical and paleotropical teams 2 https://www.iucnredlist.org/. 3 https://www.worldbank.org/en/home. https://www.scopus.com/home.uri https://www.iucnredlist.org/ https://www.worldbank.org/en/home juliano a. bogoni et al. – machine-learning applications in ecology and conservation 36 contributed 65 studies (26.3%). of these, only 11 studies (4.5%) were published in the biodiversity-rich neotropical realm (fig. 1b). similarly, the most-cited papers also predominantly originated from northern countries, mirroring the geographic bias in publication origins (fig. 1b). in terms of publication accumulation, the machine-learning publications in ecology and conservation first appear on literature in 2003. yet, this topic increased exponentially since the 2010s (fig 1b). in 2010 this overview indicated nine published manuscripts, whereas 10yrs later this number reached 120 publications (fig. 1c), an increment of 111 publications. moreover, in 2021 and 2022, these techniques seem to have become a hot-topic and consolidating more than 100 (>40.5%) publications in ecology and conservation (≥50 manuscripts per year; fig 1c). across the countries, most studies (n = 173; 70.1%) are fundamentally focused on landscape ecology (which here includes species distribution), bird ecology and conservation, wildlife ecology with more than one taxonomic group, mammal ecology and conservation, plant ecology and conservation, fish ecology and conservation, and invertebrate ecology and conservation (fig 2a). proportionally, these areas respond to 28, 10, 10, 7, 6, 5, and 5%, respectively, of all approaches (fig. 2b). log(n of studies) c ou nt ry c od e 0 50 100 150 200 250 20 03 20 07 20 09 20 11 20 13 20 15 20 17 20 19 20 21 20 23 year c um m ul at iv e pu bl ic at io ns a b c usa chn aus gbr ind can fra esp jpn deu ita tur zaf bra mex nld fin prt lka rus col irn idn swe svn mar vnm hkg pol che kor lva ury cze svk bel ben sgp nzl eth npl egy srb rou mys 0 1 2 3 4 figure 1. number of studies (log-scaled) involving machinelearning in ecology and conservation among countries (a), its global geo-distribution and citation number (log-scaled; b), and the cumulative publications over the years (c). source: original search results. landscape.ecology bird.ecology wildlife.ecology mammal.ecology plant.ecology fish.ecology invertebrate.ecolgoy climate.change data.obtaining freshwater.ecology marine.ecology urban.ecology acoustic.ecology soil.conservation citizen.science disease.ecology green.energy reptile.ecology amphibian.ecology behaviour.ecology conservation.ecology genetic multidisciplinary transdisciplinary agroecology education polution usa chn aus gbr ind can fra esp jpn bra deu ita tur zaf fin lka mex nld prt col idn irn rus svn swe mar bel ben che cze egy eth hkg kor lva mys npl nzl pol rou sgp srb svk ury vnm a b landscape ecology: 28% bird ecology: 10% wildlife ecology: 10% mammal ecology: 7% plant ecology: 6% fish ecology: 5% invertebrate ecology: 5% climate change: 3% data obtaining: 3% freshwater ecology: 3% marine ecology: 3% urban ecology: 3% acoustic ecology: 2% other approachers: <2% figure 2. (a) network between publication-based country vs. major approach of machine-learning studies in ecology and conservation; and (b) percentage (and main authors keywords) of approaches in ecology and conservation using machine-learning methods. source: original search results juliano a. bogoni et al. – machine-learning applications in ecology and conservation 37 analysis of the top-cited manuscripts (n = 14; 6%), totaling 9887 citations, reveals that the primary objectives of these publications varied significantly, yet consistently focused on providing high-quality information on both landscape-scale classifications and species distribution modeling (table 1). whereas, some top-cited manuscripts provide comparisons between machine-learning algorithms in solving an ecological or conservation issue (table 1). these 100% northern-based researches, published in high-impact journals, encompasses different methods, from maximum entropy models to boosted regression trees and random forest, and sometimes uses a vast gamma of deep learning architectures (e.g., alexnet, nin, vgg, googlenet, and resnet; see norouzzadeh et al. (2018); table 1). finally, the mammal diversity (species richness) exerts significant influences on the number of country-scale studies, whereas per capita gdp does not. mammal diversity had a positive and significant tendency [z-score = 2.362; estimates = 0.41; se = 0.17; p = 0.02] while per capita gdp had a positive but non-significant tendency [z = 1.357; estimate = 0.14; se = 0.10; p = 0.175] in influencing the patterns of global distribution of studies (fig. 3). log(mammal diversity) n um be r o f s tu di es log(per capita gdp) nearctic neotropic palearctic paleotropic 1.0 1.5 2.0 2.5 3.0 3.5 4.0 biogeographic realm lo g (n um be r o f s tu di es ) nearctic neotropic palearctic paleotropic 3.5 4.0 4.5 5.0 5.5 6.0 6.5 biogeographic realm lo g (m am m al d iv er si ty ) nearctic neotropic palearctic paleotropic 1 2 3 4 5 6 7 biogeographic realm lo g (p er c ap ita g pd ) c d e 0 5 10 15 20 4 5 6 5 10 15 2 4 6 ba figure 3. bivariate plot between the number of studies per country vs. mammal diversity (a) and per capita gdp (b), and distribution of per country studies (c), mammal diversity (d), and per capita gdp (e) across the major biogeographic realms. source: original search results juliano a. bogoni et al. – machine-learning applications in ecology and conservation 38 this model presented deviance and goodness-of-fit (g-statistic) with p < 0.001, low overdispersion (od = 0.346) and diagnostic residuals plots suggest that the distribution is satisfactory in terms of residuals. discussion the complexities of ecosystems and the challenges of managing and protecting biodiversity require constant computational advances (chaves, 2013), in which, machine-learning approaches, poses as strong candidate to help ecologists perform sophisticated data analysis. machine-learning algorithms have a high capacity of processing large datasets, identify patterns, and make predictions that can support in decision-making and resource allocation (rammer and seidl, 2019). our main results showed that machine-learning has been widely used in habitat mapping particularly within landscape ecology, predictive modelling, and other ecological data analysis devoted to understanding patterns and drivers of diversity and species distribution, especially in vertebrate and plant ecology. the integration of satellite imagery and machine-learning provides a powerful tool for mapping and monitoring habitats, tracking deforestation, and assessing land-use changes (mclaren et al., 2018; kampichler et al., 2010). moreover, machine-learning can predict species distribution, movement patterns, and migration routes, supporting in conservation planning (pittman and brown, 2011). in the context of ecological data analysis, machine-learning approaches were used to process large datasets derived from field observations and sensors (e.g., christin et al., 2019; stowell and plumbley, 2014) as to identify complex relationships and patterns in ecological systems. these main approaches address important ecological and conservation challenges. both landscape ecology and predictive modelling are essential tools for understanding and responding to the accelerating changes in habitats and climate. ecologists are increasingly required to improve their analytical skills to better predict biodiversity and species distribution responses to highly modified landscapes and pervasive climatic changes. understanding the process that influences the species distributions is a critical issue in ecology and conservation (franklin, 2009; hutchinson, 1957), especially for species under threats (phillips et al., 2004). in this context, machine-learning algorithms have become widely used in biogeography, ecology, and conservation biology to estimate the relationship between species occurrences and environmental variables (elith et al., 2011; elith and leathwick, 2009; franklin, 2009). for instance, species distribution models (sdms) enable researchers to explore key questions in conservation, ecology, and evolution, such as: (1) determining priority areas for conservation and contributing to the protection of species (faleiro et al., 2013; garcía, 2006; chen and peterson, 2002); (2) understanding invasive process of species (giovanelli et al., 2008; ficetola et al., 2007); (3) generating ecocultural niche modelling that reflects ecological influences on past human culture distributions (banks et al., 2008); (4) proposing past distributions of species (carnaval and moritz, 2008; hugall et al. 2002); and (5) predicting future distributions of species under changing of climatic and environmental characteristics (tagliari et al., 2023; bogoni and tagliari, 2021; siqueira and peterson, 2003). our results on the use of machine learning methods in ecology and conservation have shown an exponential increase since 2017. in 2010, our search identified nine published manuscripts, whereas ten years later, this number had increased by 111 publications. furthermore, our findings reveal a strong geographic concentration of ecological studies using machine learning approaches in northern countries. historically, the vast majority of scientific publications are located in these regions. for instance, only in 2020, the us authors signed about ~755,000 scientific publications, while brazilian authors reached ~100,000 publications on scopus (see search logs4), representing 86.6% less. these values are reflected in the context of machine learning approaches in ecology, as our results indicate that brazilian authors published 92.0% fewer manuscripts than their counterparts in the united states. brazil is undoubtedly earth’s most biodiverse country, harbouring the largest set of known and unknown species (moura and jetz, 2021), but still presenting glaring knowledge gaps such as the multispecies wallacean shortfalls (bogoni et al., 2021). addressing these discrepancies requires substantial investment in scientific research. in contrast, with the exception of 2022 – when investments in science were temporarily resumed –, the brazilian government has systematically deepened budget cuts for science since 2017 (see angelo, 2017; escobar, 2017). in terms of machine-learning, we advocated that the investments should be addressed to form and solidify human resources and increase the computational power of institutions, especially all those located across the countryside, given that both these factors are critical to performing high-complex analysis. moreover, our findings based on proxies of biodiversity and financial resources in 4 https://www.scopus.com/results/results.uri?st1=united+states&st2= &s=affilcountry%28united+states%29&limit=10&origin=resul tslist&sort=plf-f&src=s&sot=b&sdt=cl&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 | https:// www.scopus.com/results/results.uri?st1=united+states&st2=&s=af filcountry%28brazil%29&limit=10&origin=searchbasic&sort =plf-f&src=s&sot=b&sdt=b&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28united+states%29&limit=10&origin=resultslist&sort=plf-f&src=s&sot=b&sdt=cl&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28united+states%29&limit=10&origin=resultslist&sort=plf-f&src=s&sot=b&sdt=cl&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28united+states%29&limit=10&origin=resultslist&sort=plf-f&src=s&sot=b&sdt=cl&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28united+states%29&limit=10&origin=resultslist&sort=plf-f&src=s&sot=b&sdt=cl&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28brazil%29&limit=10&origin=searchbasic&sort=plf-f&src=s&sot=b&sdt=b&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28brazil%29&limit=10&origin=searchbasic&sort=plf-f&src=s&sot=b&sdt=b&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28brazil%29&limit=10&origin=searchbasic&sort=plf-f&src=s&sot=b&sdt=b&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28brazil%29&limit=10&origin=searchbasic&sort=plf-f&src=s&sot=b&sdt=b&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 https://www.scopus.com/results/results.uri?st1=united+states&st2=&s=affilcountry%28brazil%29&limit=10&origin=searchbasic&sort=plf-f&src=s&sot=b&sdt=b&sessionsearchid=1f0679027590dc7c447bf72ed3d85980&yearfrom=2020&yearto=2020 juliano a. bogoni et al. – machine-learning applications in ecology and conservation 39 predicting the pancontinental patterns in the application of machine-learning in ecology showed no significant influence on the number of studies published globally. however, mammal diversity showed a positive tendency whereas per capita gdp had a negative tendency in potentially influencing the distribution patterns of published studies. economically affluent regions — despite having low biodiversity when compared to tropical countries — typically account for the vast majority of scientific projects research fundings, while biodiversity-rich countries tend to have lower per-capita incomes. furthermore, despite the highly-recognised quality of researches led by global south scientists—often conducted under highly adverse field and financial conditions and whose legacy chips away at much of our knowledge gaps (bogoni et al., 2021)—rarely is a tropical country a protagonist in terms of publications number. in this context, for example, canada and china together harbour a total of 753 mammalian species (195 and 558, respectively; smith and xie, 2013; banfield, 1974), whereas brazil has 775 mammal species described (abreu et al., 2022). therefore, a kind of per capita publication in relation to mammalian species reaches 0.07 and 0.006, respectively, representing 13-fold more publications lead by these northern countries weighted by its biodiversity when compared to brazil. yet, our results indicate that this is not an exclusivity for neotropical countries, such as brazil. excepting for australia, paleotropical countries amount to 25 studies, representing only ~10% of the total. moreover, the top-10 richest countries in terms of gdp per capita (70% of them located in the nearctic of palearctic) amount by 38.1% of publications (n = 94), while the top-10 poorer countries (60% located in the paleotropical region) signed only 10.1% of scientific publications (n = 27) focused in machine-learning in ecology. our initial hypotheses were partially corroborated. our results indicated that the vast majority of publications derive from authors affiliated in economically affluent countries, and the major focus of machine-learning in ecology and conservation was widely applied to species distribution models and landscape ecology. in predictive terms, although there is a clear tendency for richer countries to publish more than poorer ones, and a negative relationship between biodiversity and publication output, mammal diversity showed a positive and significant trend, while per capita gdp showed a positive but non-significant trend. based on our global overview of the use of machine learning methods in ecology and conservation, we conclude that a wide range of ecological and conservation issues has already been addressed using these techniques. machine learning thus represents a strong promise for the coming years, contingent only on the availability of human and financial resources. this promise can therefore be solidified by a fine-tuned agenda aiming to initiate a discussion concerning countries with lower investments in scientific research and those that employ fewer modern analytical methods. such an agenda would also allow for an exploration of how these nations can benefit from theoretical and analytical advancements to address their persistent environmental and ecological challenges. finally, collaboration between scientific communities across different economic levels could enable more equitable and effective responses to shared global challenges. declarations funding not applicable conflicts of interest/competing interests the authors have no conflicts of interest to declare. availability of data and material data used here are widely available at scopus search engine (https://www.scopus.com/home.uri). code availability see supporting information code 1 author contributions statement jab and dvsc: conceptualization, data acquisition, data analysis and figures, writing & revising the original draft; dvsc, ccm, msf & jep: conceptualization, review and editing. references abreu, e.f., casali, d., costa-araújo, r., garbino, g.s.t., libardi, g.s., loretto, d., loss, a.c., marmontel, m., moras, l.m., nascimento, m.c., oliveira, m.l., pavan, s.e., tirelli, f.p., 2022. lista de mamíferos do brasil (2022-1) [dataset]. zenodo. https://doi. org/10.5281/zenodo.7469767 angelo, c., 2017. scientists plead with brazilian government to restore funding. nature 550: 166–167. https:// doi.org/10.1038/nature.2017.22757 araujo, s.b.l., rorato, a.c., perez, d.m., pie, m.r., 2013. a spatially explicit model of synchronization in fiddler crab waving displays. plos one 8(3): e57362. https://doi.org/10.1371/journal.pone.0057362 banfield, a.w.f., 1974. the mammals of canada. university of toronto press. banks, w.e., d’errico, f., peterson, a.t., vanhaeren, m., kageyama, m., sepulchre, p., ramstein, g., jost, a., https://www.scopus.com/home.uri https://doi.org/10.5281/zenodo.7469767 https://doi.org/10.5281/zenodo.7469767 https://doi.org/10.1038/nature.2017.22757 https://doi.org/10.1038/nature.2017.22757 https://doi.org/10.1371/journal.pone.0057362 juliano a. bogoni et al. – machine-learning applications in ecology and conservation 40 lunt, d., 2008. human ecological niches and ranges during the lgm in europe derived from an application eco-cultural niche modeling. j. archaeol. sci. 35: 481–491. https://doi.org/10.1016/j.jas.2007.05.011 bogoni, j.a., navarro, a.b., graipel, m.e., peroni, n., 2019. modeling the frugivory of a plant with inconstant productivity and solid interaction with relictual vertebrate biota. ecol. model. 408: 108728. https:// doi.org/10.1016/j.ecolmodel.2019.108728 bogoni, j.a., tagliari, m.m., 2021. potential distribution of piscivores across the atlantic forest: from bats and marsupials to large-bodied mammals under a trophic-guild viewpoint. ecol. inform. 64: 101357. https://doi.org/10.1016/j.ecoinf.2021.101357 bogoni, j.a., peres, c.a., ferraz, k.m.p.m.b., 2021. mediumto large-bodied mammal surveys across the neotropics are heavily biased against the most faunally intact assemblages. mammal rev. 52(2): 221–235. https://doi.org/10.1111/mam.12274 bivand, r.s., pebesma, e., gomez-rubio, v., 2013. applied spatial data analysis with r. 2nd ed. springer, ny, usa. bivand, r.s., lewin-koh, n., 2021. maptools: tools for handling spatial objects. r package version 1.1-1. https://cran.r-project.org/package=maptools (accessed 2 mar 2023). breiman, l., 2001. statistical modeling: the two cultures. stat. sci. 16: 199–231. brown, j.h., 1995. macroecology. 1st ed. university of chicago press, chicago, il, usa. carnaval, a.c., moritz, c., 2008. historical climate modelling predicts patterns of current biodiversity in the brazilian atlantic forest. j. biogeogr. 35: 1187–1201. https://doi.org/10.1111/j.1365-2699.2007.01870.x ceballos, g., ehrlich, p.r., raven, p.h., 2020. vertebrates on the brink as indicators of biological annihilation and the sixth mass extinction. proc. natl. acad. sci. usa 117: 13596–13602. https://doi.org/10.1111/ j.1365-2699.2007.01870.x chaves, j., 2013. the problem of pattern and scale in ecology: what have we learned in 20 years? ecol. lett. 16(1): 4–16. https://doi.org/10.1111/ele.12048 chen, g., peterson, a.t., 2002. prioritization of areas in china for the conservation of endangered birds using modeled geographical distributions. bird conserv. int. 12: 197–209. https://doi.org/10.1017/ s0959270902002125 christin, s., hervet, é., lecomte, n., 2019. applications for deep learning in ecology. methods ecol. evol. 10(10): 1632–1644. https://doi.org/10.1111/2041210x.13256 conway, g.r., 1977. mathematical models in applied ecology. nature 269: 291–297. dobson, a.j., 1990. an introduction to generalized linear models. chapman and hall. elith, j., leathwick, j.r., 2009. species distribution models: ecological explanation and prediction across space and time. annu. rev. ecol. evol. syst. 40: 677–697. https://doi.org/10.1146/annurev.ecolsys.110308.120159 elith, j., phillips, s.j., hastie, t., dudík, m., chee, y.e., yates, c.j., 2011. a statistical explanation of maxent for ecologists. divers. distrib. 17: 43–57. elton, c.s., 1924. periodic fluctuations in the numbers of animals: their causes and effects. br. j. exp. biol. 2: 119–163. https://doi.org/10.1111/j.14724642.2010.00725.x escobar, h. 2017. in brazil, researchers struggle to fend off deepening budget cuts. scienceinsider. available at: https://www.science.org/content/article/brazil-researchers-struggle-fend-deepening-budget-cuts faleiro, f.v., machado, r.b., loyola, r.d., 2013. defining spatial conservation priorities in the face of landuse and climate change. biol. conserv. 158: 248–257. https://doi.org/10.1016/j.biocon.2012.09.020 ficetola, g.f., thuiller, w., miaud, c., 2007. prediction and validation of the potential global distribution of a problematic alien invasive species – the american bullfrog. divers. distrib. 13: 476–485. https://doi. org/10.1111/j.1472-4642.2007.00377.x franklin, j., 2009. mapping species distributions: spatial inference and prediction. cambridge university press. garcía, a., 2006. using ecological niche modeling to identify diversity hotspots for the herpetofauna of pacific lowlands and adjacent interior valleys of mexico. biol. conserv. 130: 25–46. https://doi.org/10.1016/j. biocon.2005.11.030 giovanelli, j.g.r., haddad, c.f.b., alexandrino, j., 2008. predicting the potential distribution of the alien invasive american bullfrog (lithobates catesbeianus) in brazil. biol. invasions 10: 585–590. https://doi. org/10.1007/s10530-007-9154-5 haeckel, e., 1866. generelle morphologie der organismen. g. reimer, berlin, germany. hansen, m.c., stehman, s.v., potapov, p.v., 2010. quantification of global gross forest cover loss. proc. natl. acad. sci. usa 107(19): 8650–8655. https://doi. org/10.1073/pnas.0912668107 hassell, m.p., 2000. host-parasitoid population dynamics. j. anim. ecol. 69: 543–566. https://doi.org/10.1046/ j.1365-2656.2000.00445.x https://doi.org/10.1016/j.jas.2007.05.011 https://doi.org/10.1016/j.ecolmodel.2019.108728 https://doi.org/10.1016/j.ecolmodel.2019.108728 https://doi.org/10.1016/j.ecoinf.2021.101357 https://doi.org/10.1111/mam.12274 https://cran.r-project.org/package=maptools https://doi.org/10.1111/j.1365-2699.2007.01870.x https://doi.org/10.1111/j.1365-2699.2007.01870.x https://doi.org/10.1111/j.1365-2699.2007.01870.x https://doi.org/10.1111/ele.12048 https://doi.org/10.1017/s0959270902002125 https://doi.org/10.1017/s0959270902002125 https://doi.org/10.1111/2041-210x.13256 https://doi.org/10.1111/2041-210x.13256 https://doi.org/10.1146/annurev.ecolsys.110308.120159 https://doi.org/10.1146/annurev.ecolsys.110308.120159 https://doi.org/10.1111/j.1472-4642.2010.00725.x https://doi.org/10.1111/j.1472-4642.2010.00725.x https://www.science.org/content/article/brazil-researchers-struggle-fend-deepening-budget-cuts https://www.science.org/content/article/brazil-researchers-struggle-fend-deepening-budget-cuts https://doi.org/10.1016/j.biocon.2012.09.020 https://doi.org/10.1111/j.1472-4642.2007.00377.x https://doi.org/10.1111/j.1472-4642.2007.00377.x https://doi.org/10.1016/j.biocon.2005.11.030 https://doi.org/10.1016/j.biocon.2005.11.030 https://doi.org/10.1007/s10530-007-9154-5 https://doi.org/10.1007/s10530-007-9154-5 https://doi.org/10.1073/pnas.0912668107 https://doi.org/10.1073/pnas.0912668107 https://doi.org/10.1046/j.1365-2656.2000.00445.x https://doi.org/10.1046/j.1365-2656.2000.00445.x juliano a. bogoni et al. – machine-learning applications in ecology and conservation 41 hugall, a., moritz, c., moussalli, a., stanisic, j., 2002. reconciling paleo distribution models and comparative phylogeography in the wet tropics rainforest land snail gnarosophia bellendenkerensis (brazier 1875). proc. natl. acad. sci. usa 99: 6112–6117. https://doi.org/10.1073/pnas.092538699 hughes, a.c., orr, m.c., ma, k., cotello, m.j., waller, j., provoost, p., yang, q., zhu, c., qiao, h., 2021. sampling biases shape our view of the natural world. ecography 44(9): 1259–1269. https://doi.org/10.1111/ ecog.05926 hutchinson, g.e., 1957. concluding remarks. cold spring harb. symp. quant. biol. 22: 415–427. jordan, m.i., mitchell, t.m., 2015. machine learning: trends, perspectives, and prospects. science 349(6245): 255–260. https://www.science.org/ doi/10.1126/science.aaa8415 kampichler, c., wieland, r., calmé, s., weissenberger, h., et al., 2010. classification in conservation biology: a comparison of five machine-learning methods. ecol. inform. 5(6): 441–450. https://doi.org/10.1016/j. ecoinf.2010.06.003 kar, t.k., chakraborty, k., pahari, u.k., 2010. a prey– predator model with alternative prey: mathematical model and analysis. can. appl. math. q. 18(2): 137–168. lotka, a.j., 1925. elements of physical biology. williams and wilkins co., baltimore, md, usa. mclaren, j.d., buler, j.j., schreckengost, t., smolinsky, j.a., et al., 2018. artificial light at night confounds broad-scale habitat use by migrating birds. ecol. lett. 21(3): 356–364. https://doi.org/10.1111/ele.12902 magioli, m., ribeiro, m.c., ferraz, k.m.p.m.b., rodrigues, m.g., 2015. thresholds in the relationship between functional diversity and patch size for mammals in the brazilian atlantic forest. anim. conserv. 18(6): 499–511. https://doi.org/10.1111/acv.12201 magurran, a.e., 2004. measuring biological diversity. blackwell publishing, malden, ma, usa. may, r.m., 1976. thresholds and breakpoints in ecosystems with a multiplicity of stable states. nature 269: 471–477. https://doi.org/10.1038/269471a0 moura, m.r., jetz, w., 2021. shortfalls and opportunities in terrestrial vertebrate species discovery. nat. ecol. evol. 5: 631–639. https://doi.org/10.1038/s41559021-01411-5 olden, j.d., lawler, j.j., poff, n.l., 2008. machine learning methods without tears: a primer for ecologists. q. rev. biol. 83: 171–193. https://www.journals.uchicago.edu/doi/10.1086/587826 pebesma, e.j., bivand, r.s., 2005. classes and methods for spatial data in r. r news 5(2). https://cran.r-project.org/doc/rnews/ (accessed 2 mar 2023). pittman, s.j., brown, k.a., 2011. multi-scale approach for predicting fish species distributions across coral reef seascapes. plos one 6(5): e20583. https://doi. org/10.1371/journal.pone.0020583 r core team, 2023. r: a language and environment for statistical computing. r foundation for statistical computing. http://www.r-project.org/ (accessed 2 mar 2023). rammer, w., seidl, r., 2019. harnessing deep learning in ecology: an example predicting bark beetle outbreaks. front. plant sci. 10: 1327. https://doi. org/10.3389/fpls.2019.01327 sherratt, j.a., eagan, b.t., lewis, m.a., 1997. oscillations and chaos behind predator-prey invasion: mathematical artifact or ecological reality? philos. trans. r. soc. lond. b biol. sci. 352: 21–38. https://doi. org/10.1098/rstb.1997.0003 siqueira, m.f., peterson, a.t., 2003. consequences of global climate change for geographic distributions of cerrado tree species. biota neotrop. 3(2): 1–14. https://doi.org/10.1590/s1676-06032003000200005 smith, a.t., xie, y., hoffmann, r.s., lunde, d., mackinnon, j., wilson, d.e., wozencraft, w.c., gemma, f. (illustrator), wang, s. (honorary editor), 2013. mammals of china. princeton university press. (princeton pocket guides). stowell, d., plumbley, m.d., 2014. automatic large-scale classification of bird sounds is strongly improved by unsupervised feature learning. peerj 2: e488. https:// doi.org/10.7717/peerj.488 tagliari, m.m., bogoni, j.a., blanco, g.d., cruz, a.p., peroni, n., 2021. disrupting a socio-ecological system: could traditional ecological knowledge be the key to preserving the araucaria forest in brazil under climate change? clim. change 176: 1–20. https://doi. org/10.1007/s10584-022-03477-x volterra, v., 1926. variazioni e fluttuazioni del numero d’individui in specie animali conviventi. mem. acad. lincei roma 2: 31–113. https://doi.org/10.1073/pnas.092538699 https://doi.org/10.1111/ecog.05926 https://doi.org/10.1111/ecog.05926 https://www.science.org/doi/10.1126/science.aaa8415 https://www.science.org/doi/10.1126/science.aaa8415 https://doi.org/10.1016/j.ecoinf.2010.06.003 https://doi.org/10.1016/j.ecoinf.2010.06.003 https://doi.org/10.1111/ele.12902 https://doi.org/10.1111/acv.12201 https://doi.org/10.1038/269471a0 https://doi.org/10.1038/s41559-021-01411-5 https://doi.org/10.1038/s41559-021-01411-5 https://www.journals.uchicago.edu/doi/10.1086/587826 https://www.journals.uchicago.edu/doi/10.1086/587826 https://cran.r-project.org/doc/rnews/ https://cran.r-project.org/doc/rnews/ https://doi.org/10.1371/journal.pone.0020583 https://doi.org/10.1371/journal.pone.0020583 http://www.r-project.org/ https://doi.org/10.3389/fpls.2019.01327 https://doi.org/10.3389/fpls.2019.01327 https://doi.org/10.1098/rstb.1997.0003 https://doi.org/10.1098/rstb.1997.0003 https://doi.org/10.1590/s1676-06032003000200005 https://doi.org/10.7717/peerj.488 https://doi.org/10.7717/peerj.488 https://doi.org/10.1007/s10584-022-03477-x https://doi.org/10.1007/s10584-022-03477-x juliano a. bogoni et al. – machine-learning applications in ecology and conservation 42 journal im pact factor year a uthors m ain country n citation prim ary aim m achine learning m ethod(s) or fram ew ork m ain approach ecography 6.802 2006 elith et al. a u s 6282 to provide effective guidance on how best to use this inform ation in the context of num erous approaches for m odelling distributions m axim um entropy m odels (m a x en t and m a x en t-t) and boosted regression trees (b rt, also called stochastic gradient boosting) landscape ecology & species distribution ecological m odelling 3.512 2009 k ocev et al. sv n 139 to applied various m achine learning m ethods to the problem of predicting the condition or quality of the rem nant indigenous vegetation across an extensive area of south-eastern a ustralia r egression trees, m ulti-target regression trees, ensem bles, b agging, r andom forests landscape ecology & species distribution ecological inform atics 4.498 2010 k am pichler et al. m ex 120 to com pare five m achine-learning based classification techniques (classification trees, random forests, artificial neural netw orks, support vector m achines, and autom atically induced rulebased fuzzy m odels) in a biological conservation context techniques of classification trees, random forests, artificial neural netw orks, support vector m achines, and autom atically generated fuzzy classifiers b ird ecology & conservation plos o n e 3.752 2011 pittm an & b row n u sa 181 to: (1) d eterm ine w hether the influence of environm ental predictors on species’ distribution w as scale-dependent; (2) evaluate the utility of environm ental data from a single rem ote sensing device com bined w ith m etrics for surface m orphology to predict and m ap fish species distributions across a com plex coral reef ecosystem ; (3) determ ine w hich om ponents of rem otely sensed seafloor structure contribute m ost to the species distribution m odels; (4) identify threshold effects w here changes in environm ental variables abruptly influence species occurrence; and (5) evaluate the perform ance of tw o different m achine-learning m odelling algorithm s for spatial predictions of m arine fish distributions b oosted regression trees (b rt) and m axim um entropy m odelling (m axent) fish ecology & distribution table 1. journal, the im pact factor (if), year of publication (yr), authors, m ain country, num ber of citations, prim ary aim , m achine-learning m ethods or fram ew ork, and m ain approach of the top-cited m anuscripts (> 100 citations; n = 14) using m achine-learning approach in ecology and conservation across the globe. juliano a. bogoni et al. – machine-learning applications in ecology and conservation 43 r em ote sensing of environm ent 13.85 2012 d ronova et al. u sa 139 c lassify poyang lakew etland pfts using o b ia and to determ ine w hich im age segm entation scales and m achinelearning algorithm s optim ize class discrim ination six m achine-learning algorith representing m achine-learning principles: a probabilistic b ayes m ethod (n aïveb ayessim ple in w eka notation), a logistic regression m ethod (sim plelogistic in w eka), an artificial neural netw ork algorithm (m ultilayerperceptron (m lp)), a support vector m achine tool w ith polynom ial kernel and com plexity param eter value of 10 (sm o inw eka), a k -n earest n eighbors (ib k inw eka) and a tree-based classifier (r andom forest). landscape ecology & species distribution journal of b iogeography 4.81 2012 w erneck et al. u sa 174 to investigate the historical distribution of the c errado across q uaternary clim atic fluctuations and to generate historical stability m aps to test: (1) w hether the ‘historical clim ate’ stability hypothesis explains squam ate reptile richness in the c errado; and (2) the hypothesis of pleistocene connections betw een savannas located north and south of a m azonia m axim um -entropy c lim ate change m ethods in ecology and evolution 8.335 2012 b arbetm assin et al. fr a 1401 to conduct a com prehensive com parative analysis based on sim ple sim ulated species distributions to propose guidelines on how, w here and how m any pseudo-absences should be generated to build reliable species distribution m odels. b oosted regression trees (b rt) and random forest (r f) landscape ecology & species distribution integrative zoology 2.083 2013 li & w ang c h n 120 to applying various algorithm s for species distribution m odelling m ultivariate adaptive regression splines, m ixture discrim inant analysis, a rtificial neural netw orks, g eneralized boosting m odels, c lassification and regression tree, r andom forest, h ierarchical m odelling, g enetic algorithm for rule set production, and m axim um entropy landscape ecology & species distribution juliano a. bogoni et al. – machine-learning applications in ecology and conservation 44 peerj 3.061 2014 stow ell & plum bley g b t 185 to introduce a technique for feature learning from large volum es of bird sound recordings, inspired by techniques that have proven useful in other dom ains spectral features and feature learning b ird ecology & conservation plos o n e 3.752 2016 de souza et al. c a n 152 to assess the accuracy of satellite-based a is (a utom atic identification system ) and v m s (vessel m onitoring system ) in correctly identifying individual fishing events or ‘sets’ by com paring against expert-labeled data h idden m arkov m odel (h m m ) and d ata m ining (d m ) fish ecology & distribution energy and b uildings 7.201 2018 touzani et al. u sa 212 to propose an energy consum ption baseline m odeling m ethod based on a gradient boosting m achine w asproposed. d ecision trees and g radient boosting m achine g reen energy & energy econom y ecology letters 11.274 2018 m claren et al. u sa 102 to disentangle anthropogenic and landscape-related factors affecting stopover density, and thereby assess w hether artificial light at night m ight be affecting selection of stopover habitat, w e estim ated responses in seasonalm ean reflectivity to geographic, land cover and anthropogenic predictors using additive regression m odels fit by gradient boosting, a m achine-learning technique a dditive regression m odels fit by gradient boosting b ird ecology & conservation juliano a. bogoni et al. – machine-learning applications in ecology and conservation 45 proceedings of the n ational a cadem y of sciences 12.799 2018 n orouzzadeh et al. u sa 493 test how w ell deep learning can autom ate inform ation extraction from cam era-trap im ages d eep learning architectures: a lexn et, n in , v g g , g ooglen et, r esn et w ildlife ecology m ethods in ecology and evolution 8.335 2019 c hristin et al. c a n 187 to review existing im plem entations and show that deep learning has been used successfully to identify species, classify anim al behaviour and estim ate biodiversity in large datasets like cam era‐trap im ages, audio recordings and videos tensorflow, pytorch, k eras, m icrosoft c ognitive toolkit (c n tk ), d eeplearning4j, m atla b + d eep learning toolbox, a pache m x n et, plaidm l w ildlife ecology biodiversity informatics, 19, 2025, pp. 1-6 1 deriving best use data from neon for mosquito research applications: a practical guide with code amely m. bauer1*, sara paull2, robert p. guralnick3, lindsay p. campbell1 1florida medical entomology laboratory, department of entomology and nematology, ifas, university of florida, vero beach, fl, usa 2battelle, national ecological observatory network, boulder, co, usa 3florida museum of natural history, department of natural history, university of florida, gainesville, fl, usa abstract.—the national ecological observatory network (neon) is a long-term monitoring program at the continental scale designed to understand and forecast ecological responses to environmental change at local to broad scales. however, despite robust and nearly continuous collections, there are several challenges to deriving analysis-ready data sets from the available raw data that must be overcome to maximize the use of neon mosquito collections. here, we provide species-level estimated abundances for nighttime collected female mosquitoes derived from the mosquitoes sampled from co2 traps. by including zero counts, our derived data complement existing data sets and provide an analysis-ready time series useful for investigating mosquito phenology, abundances, and diversity at the species or community level. we also outline a set of considerations specific to filtering neon mosquito data by sex and for day or nighttime collections, highlighting factors that could introduce uncertainty to abundance estimates. along with the data set, we provide an r markdown file that includes annotated code and documents our data filtering and quality assurance and quality control (qa/ qc) steps, as well as data files used to filter the mosquito data based on qa/qc criteria. all files are freely available for download through the environmental data initiative data portal. our reproducible and fully documented workflow can be easily adapted for specific needs or other neon surveillance data. our work aims to enhance the accessibility and use of neon’s rich, long-term monitoring data. key words.—culicidae, vector abundance, mosquito diversity, population, spatiotemporal data, ecology of vector-borne disease, modeling mosquitoes are a highly diverse group, with >3700 species globally (harbach, 2018). multiple species are a substantial public and veterinary health concern because they are capable of transmitting disease-causing pathogens to humans and animals (mullen & durden, 2019). as ectotherms, mosquito abundances and phenology link closely to abiotic environmental conditions and can serve as a sentinel to global environmental change. in addition, mosquito species distributions are changing, with movements facilitated by greater global connectivity. mosquito monitoring is often conducted at local or regional levels to support understanding mosquito populations and potential disease risk. the national ecological observatory network (neon) is a long-term monitoring program at the continental scale and is designed to understand and forecast ecological responses to environmental change at local to broad scales (kao et al., 2012). neon began routine mosquito collections at a subset of sites in 2014, with all 47 core and terrestrial gradient sites collecting samples by 2019 (hoekman et al., 2016). sampling follows a standardized protocol to facilitate broad-scale analyses of e.g. mosquito abundances, diversity, and phenology (levan et al., 2022; paull et al., 2024). neon mosquito collections provide a unique opportunity to understand and predict drivers of mosquito population dynamics under a range of environmental conditions. despite the raw data reflecting robust and nearly continuous collections, producing analysis-ready data sets is a multi-step process that requires careful review of the neon documentation and considerable effort (atkins et al., 2025). providing a derived mosquito abundance data set will eliminate these steps, while substantially broadening the scope of analyses that can be performed with the data (atkins et al., 2025; nagy et al., 2021). * corresponding author: amelybauer92@gmail.com. mailto:amelybauer92%40gmail.com?subject= amely m. bauer et al. – deriving best use data from neon for mosquito research applications 2 one challenge to maximizing the use of neon mosquito collections is that the data are provided in a raw format after initial quality assurance and quality control (qa/qc) by neon staff, so appropriate filtering and processing steps must be taken prior to conducting analyses. neon and the broader community provide tutorials and two r packages to facilitate assembly of raw data files (lunch, laney, et al., 2024; lunch, sokol, et al., 2024; neon 2025). however, implementing assembly steps still requires data wrangling skills. a second challenge is that the neon raw data includes a free text field that contains comments about factors that can introduce uncertainty into abundance estimates during trapping, sorting, and taxonomic identification. these comments provide additional insights beyond the initial qa/qc check in the raw data tables, but require comprehensive assessment. the third challenge is that zero counts and species-level absences are not explicitly entered into the identification data tables and must be inferred from trap collection records. an absence of mosquito counts can reflect a lack of trapping effort or indicate that no mosquitoes were captured for an individual species during a trapping event. inferring these values is essential for many downstream analyses, including occupancy models or other approaches that require either presence/absence or abundance-based counts. producing a derived data set from the robust, continental-scale raw time-series data that neon publishes facilitates increased use of the neon mosquito collections data. here, we provide species-level estimated abundances and estimated mean number of mosquitoes per trap hour for nighttime collected female mosquitoes derived from the mosquitoes sampled from co2 traps (dp1.10043.001) 2024 data release (neon, 2024). we focus on nighttime host-seeking female mosquitoes because they are the primary focus of mosquito control programs and offer the opportunity for broader data integration with monitoring and surveillance efforts outside of the neon (nagy et al., 2021), including with multiple digitized control program collections accessible through the vectorbase repository (vectorbase 2025). this data set complements an existing data set available through the `neondivdata` r package (li et al., 2022), which provides standardized species counts adjusted for trapping effort for multiple neon biological collections but does not include species-level zero counts for trap events (li et al., 2022; o’brien et al., 2021). we also outline a set of considerations or “best practices” specific to filtering neon mosquito data by sex and for day or nighttime collections and highlight additional factors that could introduce uncertainty to abundance estimates during trapping, sorting, and taxonomic identification. along with the data set, we provide a corresponding r markdown file with annotated code and two .csv files that we used to exclude records that did not meet our qa/qc criteria for full reproducibility. the aim of this work is to provide an analysis-ready data set that can be used for a variety of different studies of nighttime collected female mosquitoes, as well as an adaptable workflow to process neon biodiversity surveillance data. neon mosquito collections neon mosquito collections are distributed across 20 terrestrial core and 27 terrestrial gradient sites and trapped using co2-baited centers for disease control (cdc) light traps with no light attractant. typically, mosquitoes are collected at ten plots within each neon site, and traps are set at a minimum distance of 310 m from one another. to account for seasonal activity patterns at neon sites where mosquitoes are absent during winter months when temperatures are low, trap collections follow a “field season” and “off season” sampling protocol. during the field season, terrestrial core sites collect mosquitoes every two weeks and terrestrial gradient sites collect mosquitoes every four weeks. following three consecutive trapping events with zero mosquitoes collected at terrestrial core sites, regular field season sampling within the neon domain ends and the sampling protocol shifts to the off-season schedule. during the off-season, three traps instead of ten traps are sampled weekly at terrestrial core sites and no collections are made at the terrestrial gradient sites within the domain. the increase in trapping frequency during the off-season at fewer trap locations is designed to more precisely capture the return of mosquito flight activity following winter dormancy. once mosquito activity returns at a terrestrial core site, the regular field season protocol resumes for both terrestrial core and gradient sites within the domain (levan et al., 2022). in line with neon’s mission as long-term monitoring program, sampling protocols for sample collection, sorting, and taxonomic identification steps have largely remained constant over time. minor adjustments to protocols that can lead to small changes in the records between version releases are outlined in detail in the product release details1 and documentation (paull et al., 2024). the mosquitoes sampled from co2-baited cdc light traps (dp1.10043.001) data product is available for download from the neon data portal2 (neon, 2024). functions in the `neonutilties` and `neonos` r packages 1 https://www.neonscience.org/data-samples/data-management/datarevisions-releases/release-2024. 2 https://data.neonscience.org/data-products/dp1.10043.001/ release-2024. https://www.neonscience.org/data-samples/data-management/data-revisions-releases/release-2024 https://www.neonscience.org/data-samples/data-management/data-revisions-releases/release-2024 https://data.neonscience.org/data-products/dp1.10043.001/release-2024 https://data.neonscience.org/data-products/dp1.10043.001/release-2024 amely m. bauer et al. – deriving best use data from neon for mosquito research applications 3 can then be used to stack and join the downloaded files in preparation for data assembly (lunch, laney, et al., 2024; lunch, sokol, et al., 2024). the mosquito collection data are contained in comma-delimited files that must be joined, then filtered, and processed to estimate abundances of nighttime collected female mosquitoes (paull & levan 2024). the files are organized in long format, with each row in the trapping and sorting files corresponding to a single event, and each row in the taxonomic file corresponding to species, genus, or family-level counts by sex. the files contain information about the location, sampling dates, trapping hours, day or nighttime collection, collection time, sex, species, counts, and proportion of the trap sample identified needed for data assembly. in addition, each file contains a ̀ samplecondition` field and a `remarks` field where neon staff report whether the collections were compromised, damaged, or lost (paull et al., 2024). the `remarks` field is a free text field that contains important information that can be used to evaluate uncertainty in the mosquito collection data, and we examined these remarks carefully to filter records to assemble the nighttime collected female mosquito data set. estimating species-level female abundances neon has a specific approach for assessing mosquito abundances. for each trap event, expert taxonomists identify a representative subsample of up to ~200 mosquitoes. the identified subsample can then be used to estimate species-level mosquito abundances of the trap collection. the `proportionidentified` field indicates the proportion of the total trap collection (based on sample weights) identified by the taxonomic laboratory (levan et al., 2022; paull & levan 2024). we use this information to estimate the total number of mosquitoes for each species and sex, calculated as `individualcount` divided by `proportionidentified`. this approach assumes that the counts of individual species by sex in the subsample scale proportionally to the full trap sample. when assembling the data set for species-level nighttime collected female mosquitoes, an important consideration is the proportion of mosquitoes for which sex was not determined to ensure accurate representation of the female mosquito population in the subsample. as part of the taxonomic identification process, neon taxonomists indicate whether a mosquito is a female, male, or is unidentified to sex (levan et al.; 2022). here, we calculated the proportion of mosquitoes with sex identification for each trapping event, and kept only trap events where the sex of at least 90% of the mosquito subsample was determined. similarly, the proportion of mosquitoes identified to the species-level is important for estimating individual species abundances in trap events. we decided to include only trapping events where at least 90% of the mosquito subsample was identified to the species level. we then aggregated subspecies records at the species level and removed all records that identified mosquitoes to the family or genus level. to filter the data, we implemented two specific variables in the provided r code. changing the values of these variables allows users to easily adjust the resulting data set to their specific needs. filtering to nighttime collections the neon sampling protocol includes daytime and nighttime collections, generally with each trap event occurring at an individual plot during a 24-hour sampling bout (or 40 hours prior to 2018: paull et al., 2024). within the data set, the `nightorday` field indicates whether the trap was set for daytime or nighttime collection. nighttime traps are set approximately at dusk and collected the following morning, near dawn. a four-hour time window is allowed surrounding trap setting and collection times to accommodate unexpected logistical constraints. nighttime collections accounted for approximately half of all trap events included in the neon 2024 data release. however, we found that filtering trap events to nighttime collections identified by the `nightorday` field alone can exclude some trap events that are important for including zero counts. in order to retain the maximum number of trap events, we first filtered records that indicated nighttime collection events. then, we filtered records with no indicator in this field but with zero trap hours, which accounted for 15% of all trap events in the 2024 data release. these records occur when sampling is considered “impractical” because the weather was too cold for mosquitoes to be active (i.e., a zero for mosquito abundance is appropriate), the site could not be reached, or some other factor precluded sampling at the scheduled time. we evaluated the records in these events in a later step to determine whether to include a zero count because sampling was impractical due to cold weather. these steps can be easily updated using the code provided in the corresponding r markdown file if users are interested in filtering for daytime collections. considerations for including or excluding a trap event the neon mosquito collection data contain multiple qa/qc fields that provide useful information about the trap event, sample sorting, and taxonomic identification. multiple factors can introduce uncertainty into abundance estimates across these categories. we found that considering information within designated qa/qc fields, as well amely m. bauer et al. – deriving best use data from neon for mosquito research applications 4 as free text comments from neon staff members included in the `remarks` columns, was critical for determining whether to include or exclude a trap event. for example, in addition to mosquito sampling, a subset of neon mosquito collections are tested for pathogens, which requires that a strict cold chain is maintained during sample transport, sorting, and taxonomic identification. here, we found that filtering records based on `samplecondition` alone can eliminate trap events that were compromised for pathogen testing but not necessarily compromised for abundance estimation. in addition, in some cases, the `samplecondition` field does not indicate a compromised trap, but the free-text field providing comments from neon staff in the ̀ remarks` columns contained information showing that uncertainty could be introduced to abundance estimates. when considering factors that could introduce uncertainty to abundance estimates, we focused on general conditions. first, we checked qa/qc fields containing information about the status of the trap fan, co2 sublimation, and catch cup condition, and we eliminated trap events where these factors were compromised. next, we examined the `remarks` fields for trapping, sorting, and taxonomic identification. general criteria for excluding a trap event included traps that were tipped over and on the ground, the presence of ants or other predators in the sample, mosquitoes frozen to the sides of catch cups, sample spills where it was unclear if all mosquitoes were recovered, or issues affecting taxonomic identification. because these remarks are free text comments entered by neon staff across three categories of the sample collection process, we exported unique remarks to a separate .csv file and then examined them using the open-source software program openrefine. we used faceting, clustering, and text search functions within openrefine v3.7.9 (delpeuch et al., 2024) to facilitate examination of unique remarks, and we indicated whether the trap event should be kept or excluded in a `keepevent` field. we then joined the `keepevent` field back to the neon data to narrow collection records based on our qa/qc criteria for estimating abundances. using this approach, we removed 7% of trap events from the active nighttime mosquito collections data set. for ~40% of the excluded events, `samplecondition` indicated no known compromise during sample collection, sorting, and identification, but our comprehensive evaluation of the free text and trap status fields provided more detailed information that we determined could affect confidence in estimated abundances. we also identified that ~40% of the active trap events that we retained during this step did not include a compromise status, meaning that these records would be eliminated during a filtering step based on the `samplecondition` fields alone. the result is a qa/qc data set that is comprehensive to estimating species level nighttime collected female mosquito abundances. including zero counts one contribution of the neon mosquito collection data is relatively consistent and standardized sampling, which allows for the inclusion of zero counts when no mosquitoes are collected or when an individual species is not captured in a trap. the raw data downloaded from neon does not explicitly incorporate zero counts in the taxonomic data tables, but these values can be inferred from the trap collections data (e.g., by checking the trap hours and values of the `targettaxapresent` columns). here, we added zero counts for individual trap events and for individual species within each trap using a set of criteria. we considered two factors to determine whether to include a zero for a trap event. first, we identified whether a trap was active, indicated by the `traphours` field containing a value greater than zero. if the trap was active but no female mosquitoes were collected, we included a zero for the count for the trap event. then, we examined trap events where the trap was inactive, indicated by the `traphours` field equaling zero. for these events, we observed the `samplingimpractical` and `remarks` columns and added zero counts for traps that were inactive because of low temperature conditions or snow cover preventing access to trap locations. the `samplingimpractical` field was introduced to the neon data in september of 2019 to better indicate when planned sampling did not take place (paull et al., 2024). for records prior to 2020, we relied on information provided in the remarks field alone. combined, these considerations allowed us to include almost 18,000 zero count events that make up ~40% of all trap events in the final data set. we also included a zero count for individual species when a trap was set but the species was not collected. the neon mosquito data includes a field `nativestatuscode` indicating native and non-native species, which can be used as an indicator of whether a species may have moved into a region after a specific date. rather than trying to determine the timing of the distribution of non-native species across neon sites, we opted to include a zero count for a species in a trap event if the species was collected at the trap site at least once within the given sampling year. one limitation of this approach is that interannual variation in a species’ presence may be incomplete because no record will be included if the species was not collected at least once during a trapping year. this approach could be updated in the future, for example using more specific information amely m. bauer et al. – deriving best use data from neon for mosquito research applications 5 about arrival dates of non-native species to a geographic area. r markdown file and “remarks_or.csv” files along with the final filtered and qa/qc-processed data set of species-level, nighttime collected female mosquitoes with corresponding zero counts, we provide an r markdown file containing annotated code used to filter and conduct qa/qc steps specific to the inclusion/exclusion criteria outlined above. this file provides information for data access and download, as well as code required to reproduce the final data set made available here. the file also documents the general steps used to process the remarks field in openrefine, with screenshots illustrating where the facet, clustering, and text search functions are located within the software. in addition to this information, further minor qa/qc steps are outlined in the r markdown documentation, including checks for duplicate trap events and collection records, including events where counts of a mosquito species occur more than once for a given event. the documentation includes a list of all used r packages, with package versions reported in the markdown file reference section. we also provide the two .csv files exported from openrefine that indicate whether a trap event should be included or excluded in the `keepevent` field. the purpose of the markdown and .csv files is to provide a fully reproducible workflow for the final data set. further, the code and steps used may be altered in future iterations to filter neon data under different user defined criteria and these steps are easily adapted to derive abundance data sets from similar raw neon data. the derived data set, the two .csv files containing processed remarks, as well as the r markdown file in form of its executable .rmd script and rendered .hmtl version are available for download through the environmental data initiative data portal3 (bauer et al.,2025). conclusions the derived species-level nighttime female abundance data set described here provides a comprehensive resource for examining broadscale mosquito population dynamics at neon sites across the u.s. the inclusion of zero counts for trap events for which no mosquitoes are collected provides an analysis-ready time series useful for examining mosquito phenology, abundances, and diversity at the species or community level, and the reproducible workflow, including annotated coding steps, provides a blueprint for users to alter filtering steps for specific needs. we hope the derived data set and corresponding code will help to maximize accessibility and use of this rich, long3 https://doi.org/10.6073/pasta/5bd343a4d15fe7ac70107b0c4a0171b7 term monitoring data set. acknowledgments the national ecological observatory network is a program sponsored by the national science foundation and operated under cooperative agreement by battelle. this material is based in part upon work supported by the national science foundation through the neon program. competing interests the authors have declared that no competing interests exist. literature cited atkins, j. w., aho, k. s., chen, x., elmore, a. j., fiorella, r., luo, w., lombardozzi, d., lunch, c., manak, l., pablo, l. x. de, myers‐pigg, a. n., record, s., qiu, t., reed, s., ruddell, b., strange, b., torrens, c. l., yule, k., & richardson, a. d. (2025). recommendations for developing, documenting, and distributing data products derived from neon data. ecosphere, 16(1), article e70159. https://doi.org/10.1002/ ecs2.70159 bauer, a. m., paull, s., guralnick, r. p., & campbell, l. p. (2025). species-level estimated abundances and zero counts of nighttime collected female mosquitoes 2014 2022 (derived from neon mosquitoes sampled from co2 traps (dp1.10043.001, release-2024)) (version 2). environmental data initiative. delpeuch, a., morris, t., huynh, d., weblate, mazzocchi, s., jacky, guidry, t., elebitzero, stephens, o., matsunami, i., sproat, i., santos, s., larsson, a., allanaaa, kushthedude, fauconnier, s., mishra, e., beaubien, a., magdinier, m., . . . chandra, l. (2024). openrefine (version 3.7.9). zenodo. https://doi. org/10.5281/zenodo.10644943 harbach, r. e. (2018). culicipedia: species-group, genus-group and family-group names in culicidae (diptera). cabi publishing. hoekman, d., springer, y. p., gibson, c., barker, c. m., barrera, r., blackmore, m. s., bradshaw, w. e., foley, d. h., ginsberg, h. s., hayden, m. h., holzapfel, c. m., juliano, s. a., kramer, l. d., ladeau, s. l., livdahl, t. p., moore, c. g., nasci, r. s., reisen, w. k., & savage, h. m. (2016). design for mosquito abundance, diversity, and phenology sampling within the national ecological observatory network. ecosphere, 7(5), e01320. https:// doi.org/10.1002/ecs2.1320 kao, r. h., gibson, c. m., gallery, r. e., meier, c. l., barnett, d. t., docherty, k. m., blevins, k. k., travers, p. d., azuaje, e., springer, y. p., thibault, k. m., https://doi.org/10.6073/pasta/5bd343a4d15fe7ac70107b0c4a0171b7 https://doi.org/10.5281/zenodo.10644943 https://doi.org/10.5281/zenodo.10644943 https://doi.org/10.1002/ecs2.1320 https://doi.org/10.1002/ecs2.1320 amely m. bauer et al. – deriving best use data from neon for mosquito research applications 6 mckenzie, v. j., keller, m., alves, l. f., hinckley, e.-l. s., parnell, j., & schimel, d. (2012). neon terrestrial field observations: designing continental‐ scale, standardized sampling. ecosphere, 3(12), 1–17. https://doi.org/10.1890/es12-00196.1 li, d., record, s., sokol, e. r., bitters, m. e., chen, m. y., chung, y. a., helmus, m. r., jaimes, r., jansen, l., jarzyna, m. a., just, m. g., lamontagne, j. m., melbourne, b. a., moss, w., norman, k. e. a., parker, s. m., robinson, n., seyednasrollah, b., smith, c., . . . zarnetske, p. l. (2022). standardized neon organismal data for biodiversity research. ecosphere, 13(7), e4141. https://doi.org/10.1002/ ecs2.4141 levan, k., hoekman, d., & gibson, c. m. (2022). tos science design for mosquito abundance, diversity, and phenology: neon.doc.000910vc. national ecological observatory network (neon). lunch, c., laney, c., mietkiewicz, n., sokol, e., cawley, k., & national ecological observatory network. (2024). neonutilities: utilities for working with neon data (version 2.4.2). https://cran.r-project. org/package=neonutilities lunch, c., sokol, e., robinson, n., & national ecological observatory network. (2024). neonos: basic data wrangling for neon observational data (version 1.1.0). https://cran.r-project.org/package=neonos mullen, g. r., & durden, l. a. (2019). medical and veterinary entomology (third edition). academic press. nagy, r. c., balch, j. k., bissell, e. k., cattau, m. e., glenn, n. f., halpern, b. s., ilangakoon, n., johnson, b., joseph, m. b., marconi, s., o’riordan, c., sanovia, j., swetnam, t. l., travis, w. r., wasser, l. a., woolner, e., zarnetske, p., abdulrahim, m., adler, j., . . . zhu, k. (2021). harnessing the neon data revolution to advance open environmental science with a diverse and data‐capable community. ecosphere, 12(12), article e03833. https://doi.org/10.1002/ ecs2.3833 neon (national ecological observatory network). (2024). mosquitoes sampled from co2 traps (dp1.10043.001). https://data.neonscience.org/data-products/dp1.10043.001/release-2024 neon (national ecological observatory network). (2025). resources: educational resources. https:// www.neonscience.org/resources o’brien, m., smith, c. a., sokol, e. r., gries, c., lany, n., record, s., & castorani, m. c. (2021). ecocomdp: a flexible data design pattern for ecological community survey data. ecological informatics, 64, 101374. https://doi.org/10.1016/j.ecoinf.2021.101374 paull, s., & levan, k. (2024). neon user guide to mosquitoes sampled from co2 traps (dp1.10043.001) and mosquito‐borne pathogen status (dp1.10041.001), version f. national ecological observatory network (neon). paull, s., levan, k., tsao, k., hoekman, d., & springer, y. p. (2024). tos protocol and procedure: mos – mosquito sampling: neon.doc.014049vn. neon (national ecological observatory network). https://doi.org/10.1890/es12-00196.1 https://doi.org/10.1002/ecs2.4141 https://doi.org/10.1002/ecs2.4141 https://cran.r-project.org/package=neonutilities https://cran.r-project.org/package=neonutilities https://cran.r-project.org/package=neonos https://data.neonscience.org/data-products/dp1.10043.001/release-2024 https://data.neonscience.org/data-products/dp1.10043.001/release-2024 https://www.neonscience.org/resources https://www.neonscience.org/resources biodiversity informatics, 19, 2025, pp. 46-60 46 evaluating knowledge gaps in reptile records in nayarit, northwestern mexico maría daniela arvizu1, arturo ruiz-luna2*, beatriz yáñez-rivera3, césar alejandro berlanga-robles2, ángela p. cuervo-robayo4 1programa de doctorado en ciencias, centro de investigación en alimentación y desarrollo (ciad), subsede mazatlán, mazatlán, mexico. 2 laboratorio de manejo ambiental, centro de investigación en alimentación y desarrollo. subsede mazatlán, mazatlán, mexico. 3instituto de ciencias del mar y limnología, universidad nacional autónoma de méxico, mazatlán, mexico. 4departamento de zoología, instituto de biología, universidad nacional autónoma de méxico, mexico city, mexico. abstract. mexico hosts a great diversity of reptile species; yet, many reptiles are either threatened or endangered. complete and updated information is required to implement appropriate management and conservation actions; however, species inventories can include taxonomic, geographic, and temporal gaps. therefore, this study aimed to evaluate the magnitude of these gaps in digitally accessible information on reptiles from the state of nayarit, located in northwestern mexico. a database was generated using information from the national biodiversity information system (snib) of the national commission for the knowledge and use of biodiversity (conabio). the growth rate of new species descriptions was calculated, and the completeness of the inventory was evaluated in 10-km grid cells across various time periods, considering biogeographic and physiographic regions. the species description growth rate was low. in addition, approximately 40% of the surface of nayarit exhibited information gaps among reptile records, particularly in mountainous and hard-toreach areas. notably, the least amount of information was recorded between 1981 and 2000. our results lay the groundwork for future research and the development of effective strategies to conserve and manage the natural resources of nayarit. key words: taxonomic gap; spatial gap; temporal gap; geographic information systems; nayarit state. *corresponding author: arturo ruiz-luna, email: arluna@ciad.mx introduction the class reptilia comprises a diverse group of vertebrates, including the orders testudines (turtles), rhynchocephalia (tuataras), squamata (lizards and snakes), and crocodylia (crocodiles and caimans) that collectively encompass 12,386 living species. of these, approximately 97% belong to the order squamata (uetz et al., 2025). in addition to their intrinsic value as biotic components of the ecosystems they inhabit, reptiles provide multiple ecosystem services (millennium ecosystem assessment, mae, 2005). among other supporting and regulating services, reptiles participate in nutrient cycling, bioturbation, biological control, and seed dispersal (valencia-aguilar et al., 2013; cortés-gómez et al., 2015; de miranda, 2017). reptiles also contribute provisioning services by acting as sources of food, raw materials, and medicine. finally, reptiles hold notable importance for various cultures and are reflected in numerous cultural and religious expressions (valencia-aguilar et al., 2013; valdez-rentería et al., 2023). despite their importance, reptile species worldwide face major threats, with habitat destruction and fragmentation among the most important (cox et al., 2022; farooq et al., 2024). other major threats include defaunation (nijman et al., 2012; young et al., 2016; marshall et al., 2020; finn et al., 2023), the introduction of exotic or invasive species (pyšek et al., 2020), and the spread of diseases, which can lead to local population declines (okoh et al., 2021; schilliger et al., 2023). climate change is also a key threat that can directly or indirectly affect reptile development by altering temperature and precipitation patterns and water availability (pieau, 1996; booth, 2006; newbold, 2018). mailto:arluna@ciad.mx maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 47 from 1971 to 2000, the extinction rate of reptiles was ~0.03% per year; however, this rate is expected to increase to 2.93% over the next century (alroy, 2015). consequently, the international union for conservation of nature (iucn) has estimated that ~21.1% of reptile species are at risk of extinction (cox et al., 2022). notably, mesoamerica has the highest number of threatened species identified to date (alroy, 2015). given the multiple factors that threaten reptile diversity and conservation, it is crucial to understand how species are distributed across time and space. indeed, accurate species identification and monitoring are essential for future decision-making, although obtaining this information can be costly in terms of time and resources. understanding species distributions provides key insights for detecting, monitoring, measuring, and predicting variations in biological diversity and their effects on ecosystems (wheeler et al., 2012). several authors have warned of an impending mass extinction or biodiversity crisis related to human activities that negatively impact habitats and promote the overexploitation of natural resources (rull, 2022; sandor et al., 2022). notably, this biodiversity crisis is also partly due to a lack of public awareness of its existence (dirzo et al., 2022). thus, mapping the biosphere provides valuable information that can be used to generate potential solutions to the biodiversity crisis (zhang et al., 2017). in this context, biodiversity information repositories, such as the global biodiversity information facility (gbif), play a crucial role in compiling, storing, and safeguarding vast amounts of biological data. these resources, combined with various analyses, such as modeling species distributions and the effects of climate change, offer a means to better understand the ecology of individual species and large taxonomic groups (luo et al., 2021; lajeunesse & fourcade, 2023). in particular, such tools are essential for addressing information gaps and improving sampling and monitoring strategies (neves et al., 2019; grattarola et al., 2020). despite the large volume of recorded data, the information stored in these repositories contains multiple gaps, with taxonomic, spatial, and temporal gaps being the most common (hortal et al., 2015; isaac & pocock, 2015; marshall et al., 2024). taxonomic gaps refer to potential omissions or absences in the species composition of a given region (hortal et al., 2015). these are typically evaluated through the annual species description rate, which serves as an indicator of new species discoveries (marshall et al., 2024). spatial gaps occur when distribution data is incomplete at the level of individual species, groups, or taxocenoses (lobo et al., 2018; nori et al., 2023). these gaps can be quantified by assessing inventory completeness, which measures the proportion of recorded species in a given area relative to the estimated total number of species present (sousa-baena et al., 2014; escribano et al., 2019; huang et al., 2020; chesshire et al., 2023). temporal gaps refer to inconsistencies in species inventories over time (tessarolo et al., 2017; escribano et al., 2019), which can be measured by analyzing the frequency of species records and inventory completeness (shirey et al., 2021). geographic bias may be present in biodiversity data, meaning that species records have been disproportionately collected from certain regions, environments, or ecosystems (loiselle et al., 2008; hortal et al., 2015). identifying and evaluating this bias can provide a more balanced understanding of the natural capital of a region, allowing researchers and conservationists to pinpoint areas with high potential for implementing future research efforts and management actions (soberón et al., 2007; asase & peterson, 2016; ganglo & kakpo, 2016). mexico is one of the most biodiverse countries in the world. in addition, mexico hosts over 8% of all reptile species (~1,023 recorded species), ranking second globally (johnson et al., 2017; uetz et al., 2024). most of these species (~60.1%) are endemic and adapted to various habitats throughout the country (johnson et al., 2017; smith & lemos-espinal, 2022; lemos-espinal & smith, 2023; ramírez-bautista et al., 2023; suazo-ortuño et al., 2023). currently, ~15% of reptile species in mexico are threatened or endangered (semarnat, 2010; smith & lemos-espinal, 2022; lemos-espinal & smith, 2023; suazo-ortuño et al., 2023; iucn, 2024), while the conservation status of ~13% remains uncertain (iucn, 2024). previous studies have provided insights into species diversity across various regions and states in mexico (flores-villela & goyenechea, 2003; flores-villela & garcía-vázquez, 2014; aguilar-lópez et al., 2016; lemos-espinal & smith, 2020; 2023), including nayarit state, which is the focus of this study. according to ramírez-bautista et al. (2023), nayarit exhibits an intermediate level of reptile richness (135 recorded species). located in northwestern mexico, nayarit spans the neotropical and mexican transition zone biogeographic regions, where species from the nearctic and neotropical regions converge (morrone, 2019). in addition diverse geoforms, land cover types, and geological processes, converge in nayarit, creating a unique geographic identity although previous studies have examined the reptile composition of nayarit (luja et al., 2014; woolrich-piña et al., 2016; ramírez-bautista et al., 2023; loc-barragán et al., 2024), no assessment has been conducted to validate the sufficiency and quality of the available information. maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 48 therefore, the objective of this study was to determine the magnitude of taxonomic, spatial, and temporal gaps in digitally accessible reptile data from nayarit. this assessment aimed to validate existing inventories, and its results provide a foundation for designing future sampling and conservation plans. materials and methods study area the state of nayarit (23°05’04” to 20°36’12” n and 103°43’15” to 105°45’37” w) is located on the northwestern coast of mexico (fig. 1). the state borders the pacific ocean to the west and the states of sinaloa, durango, zacatecas, and jalisco to the north and east. nayarit covers an approximate area of 28,000 km², representing 1.4% of the national territory, and has a coastline of 296 km. its elevation ranges from 0 to 2,400 meters above sea level. physiographic provinces provide a means to divide the surface of the earth into regions that share common geological and geomorphological characteristics. nayarit hosts four physiographic provinces: eje neovolcánico, llanura costera del pacífico (including the tres marías islands subprovince), sierra madre del sur, and sierra madre occidental. the predominant climate in 60.6% of the territory is tropical savanna (aw), while 22% has a humid subtropical climate (cwa). the remaining areas exhibit subtropical highland climate (cwb) and hot semi-arid climate (bsh) (beck et al., 2018). the annual mean temperature ranges between 14 °c and 28 °c, while the average annual precipitation varies from 600 mm to 2,500 mm (inegi, 2022). the main land cover types of nayarit are tropical or sub-tropical broadleaf deciduous forest (27%), croplands (22%), and temperate or sub-polar broadleaf deciduous forest (18%) (cec, 2023). presence data a comprehensive review of reptile records in nayarit was conducted using the biodiversity information system of mexico (snib1), managed by the national commission for the knowledge and use of biodiversity (conabio). snib-conabio serves as the official global biodiversity information facility (gbif2) node for mexico, providing national biodiversity data on a global scale. the snib-conabio database contains biological records from mexico dating back from the mid-19th century up to 2023, data from scientific collections, information from citizen science initiatives (conabio, 2024), and data stored in vertnet3. the information in snib-conabio is curated, validated, updated periodically, and available for 1 https://www.snib.mx/ejemplares/descarga/ 2 https://www.gbif.org/ 3 http://www.vertnet.org/ download in all versions. therefore, this article focused on analyzing only this biodiversity information system. the present review was conducted using the advanced search system of snib-conabio with filters for taxonomic group (reptiles), country (mexico), and state (nayarit). the basic download included 49 selected fields. data processing once the snib database was generated and downloaded, a thorough cleaning process was conducted to remove duplicate records, entries lacking a scientific name at the species level, and records without a registration date. records containing errors or taxonomic synonyms (nomenclatural updates) were corrected whenever possible using validated names from the reptile database (uetz et al., 2024). additionally, records without geographic coordinates or those located outside the boundaries of nayarit were excluded. the snib database (v. 2024-02-26) contained 8,864 reptile records for nayarit, with the oldest dating back to 1861 and the most recent pertaining to 2023. of these, 34.4% lacked a registration date, and 3.0% were missing a species-level name. after the data-cleaning process, a total of 5,761 valid records were obtained for analysis. additionally, 279 records from tres marías islands, sourced from the snib database (v. 2021-05-28), were incorporated, resulting in a final dataset of 6,040 records. the processed records were integrated into comma-separated value (*.csv) spreadsheets for further analysis and processing in geographic information system (gis) applications, specifically qgis v. 3.34.6 (qgis, 2025). taxonomic gap analysis the taxonomic gap analysis was conducted by constructing a species accumulation curve based on the year figure 1. study area: state of nayarit, mexico. land use classification by the commission for environmental cooperation (modified from cec [2023]). natveg is natural vegetation. https://www.snib.mx/ejemplares/descarga/ https://www.gbif.org/ http://www.vertnet.org/ maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 49 number of species recorded only twice in the same quadrat during the study period. the inventory completeness (c) index was calculated as follows: . eq. (4) the c index ranges from 0 (lacking) to 1.0 (when observed and expected richness are equal). following the methods of peterson et al. (2016) and troia & mcmanamay (2016), cells with c values ≥0.8 were considered well-surveyed. using similar classification criteria as arvizu & ruiz-luna (2024) for amphibians in nayarit, the c values were categorized into three levels: complete (c ≥ 0.8), partially complete (0.5 ≤ c < 0.8), and incomplete (0 < c < 0.5). these classifications were visualized in maps generated using qgis. in addition to evaluating the proportion of well-inventoried areas for reptiles at the state level, the c index was analyzed across biogeographic regions (a geographic unit defined by the biotic and ecological characteristics shared by the species that inhabit it, shaped by evolutionary history) and physiographic provinces (a geographic area defined by its geological and geomorphological c haracteristics). biogeographic regions are the largest ecozones that divide the earth and are based on the distribution of terrestrial organisms, biogeographic subdivisions are called provinces, which are areas defined by endemism and by their ecological and physiographic identities (morrone, 2019). in nayarit, the two biogeographic regions were the neotropical region, with one province (pacific lowlands), and the mexican transition zone, with two provinces (sierra madre occidental and the transmexican volcanic belt). on the other hand, considering the physiographic provinces, the nayarit limits include four of them, named eje neovolcánico, llanura costera del pacífico (including the tres marías islands subprovince), sierra madre del sur, and sierra madre occidental. given that each of these regions and provinces contains distinct vegetation types, soil conditions, climates, and elevations, each has the potential to harbor unique reptile communities. both regionalization were used to generate comprehensive knowledge of the state of nayarit, which can be used by decision-makers to develop management and conservation plans, and support policy for the study and conservation of biological diversity. to assess geographic bias, the relationship between reptile records and the major roads and urban areas in nayarit was analyzed. for this, we calculated the proportion of records located near roads, including roads surof each description. the analysis was performed using microsoft excel (microsoft corporation, 2013), and the best-fit model was determined using solver, a microsoft excel tool. the gompertz model (tjørve, 2003) provided the best fit and is described as follows: , eq. (1) where n(t) represents the recorded number of species over time t, a is the final population size, b is a constant related to the initial points, c is the growth constant, and e is the exponential function. the goodness-of-fit for the gompertz model was evaluated using deviance (fox, 2015; dobson & barnett, 2018) calculated as: , eq. (2) where yi represents the observed cumulative number of species, and ŷi is the estimated cumulative number of species. to prevent convergence issues in logarithmic calculations, cases of ŷi< 1 were excluded. deviance (d) was used to evaluate the model fit (fox, 2015). under the null hypothesis (h₀), it was assumed that the model adequately described the observed data. the asymptotic distribution of d followed a chi-square (x2) distribution with n–r degrees of freedom, where n is the number of data points and r is the number of estimated parameters (in this case, r = 3). the annual rate of species accumulation was determined following the methodology proposed by marshall et al. (2024). the average species accumulation rate from 1990 to the present was calculated as an approximation of the probability of discovering new species. spatial gap analysis the spatial analysis of the data was conducted using qgis v. 3.34.6 (qgis, 2025). a 10-km resolution grid (100 km² per cell) was generated for the study area, following the methodologies previously applied in nayarit and similarly sized regions (baselga & novoa, 2006; gonzález et al., 2007; arvizu & ruiz-luna, 2024). from the generated grid, expected species richness (sexp) per cell was estimated using the equation proposed by chao (1984) and colwell & coddington (1994): , eq. (3) where sobs is observed species richness, a is the number of species recorded only once in a quadrat, and b is the maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 50 rounding and within urban centers, using buffer zones of 100 m and 250 m perpendicular to both sides of the road network. these buffer zones were created using vector geometry tools in qgis. in this study, the exclusive focus on distance to roads carries a sampling bias, which was due to a prior data analysis. other potential geographic biases include proximity to urban centers, especially those with educational institutions, and protected areas. temporal gap analysis to establish a timeline of species records, the c index was evaluated over time to identify trends in species documentation and the periods during which the most important contributions to reptile inventory completeness in nayarit occurred. the dataset was divided into five specific periods (i.e., 1861–1960, 1961–1980, 1981–2000, 2001–2020, and 2021–2023), and the c index was calculated for each. although the first period spans 100 years, the largest proportion of data (85%) came from the 20year period of 1941–1960 (933 records). the remaining periods spanned 20 years, with the exception of the final period, which spanned 3 years. this final 3-year period (2021–2023), which ensured our analysis was as current as possible, included 26% of all records. results based on the snib records for nayarit, a total of 136 reptile species were taxonomically validated for this study. these species were described in nayarit between 1861 and 2023. the species accumulation curve followed a sigmoidal pattern, indicating a decline in the rate of new species records reported in recent years. a notable inflection point was observed in the third decade of the last century, with the curve approaching an asymptote of nearly 140 species (fig. 2). the best fit for this curve was obtained using the gompertz model with a growth constant estimated at 0.025: . in the first three years of the analyzed period, the estimated accumulated species count (ŷi) was less than one, thus these observations were excluded from the goodnessof-fit test using deviance. of the 71 years analyzed, the test was conducted with 68 years, resulting in 65 degrees of freedom. the fit of the data to the gompertz model was significant (d = 47.8, p = 0.945), supporting the ability of the model to describe the accumulation of species records over the study period (fig. 2). for the period of 1990– 2022, the average species description growth rate, which was used as a proxy for the probability of the occurrence of a new species, was estimated to be 0.26 ± 0.46%. figure 2. species accumulation curve for reptiles in nayarit by year of description. the black diamonds represent accumulated species, and the red line represents the gompertz model. maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 51 approximately 60% of the 356 cells (10-km resolution) in the study area contained at least one reptile record within the analyzed period. however, due to methodological constraints, it was not possible to estimate the c index for all cells with species records. among the cells where the c index could be estimated, only 3.4% exhibited c values ≥0.8, indicating that these cells were well-inventoried. notably, these cells were heterogeneously distributed across the state. additionally, 17.4% of the cells had c index values of 0.5–0.8, indicating that these cells contained partially complete inventories. these cells were more evenly distributed throughout the state (fig. 3a). overall, only a little more than 20% of nayarit can be considered completely or partially inventoried in terms of reptile species. differences were found in inventory completeness between the two biogeographic regions of nayarit, despite their similarities in territorial extent. the neotropical region included 183 grid cells, with 70% of them containing at least one record of a reptile species. in contrast, the mexican transition zone included 173 grid cells, with reptile species records in only 50%. of these, only 27% and 14% of the grid cells corresponded to the neotropical region (50) and mexican transition zone (24), respectively, exhibiting c index values ≥0.5, with less than 5% of cases exhibiting c index values ≥0.8 (table 1a). when each region and its respective provinces were evaluated globally, the c index values ranged between 0.77 to 0.83 (fig. 3b), indicating the regions were partially to fully inventoried, with the highest species richness (110) observed in the neotropical region. figure 3. inventory completeness in the biogeographic regions and physiographic provinces of nayarit. a) cells representing the completeness index categories in the biogeographic regions: incomplete (0 < c < 0.5; red), partially complete (0.5 ≤ c < 0.8; yellow), and complete (c ≥ 0.8; green). b) inventory completeness values by biogeographic region: neotropical (pacific lowlands province; orange) and mexican transition zone (sierra madre occidental and trans-mexican volcanic belt provinces; green). c) cells representing the completeness index categories in the physiographic provinces: incomplete (0 < c < 0.5; red), partially complete (0.5 ≤ c < 0.8; yellow), and complete (c ≥ 0.8; green). d) inventory completeness values by physiographic province: llanura costera del pacífico (orange), sierra madre occidental (yellow), eje neovolcánico (purple), and sierra madre del sur (gray). maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 52 the analysis of the physiographic provinces indicated that more than 50% of grid cells (190) were located within the sierra madre occidental province, which also covered the largest area of nayarit (fig. 3c). however, only 45% of these cells contained records of at least one reptile species, with 11.6% of the cells showing c index values ≥0.5. the highest level of inventory completeness was observed in the eje neovolcánico province, which spatially represented 20% of the total grid cells and had c index values ≥0.5 in 35.7% (table 1b). for the sierra madre del sur and llanura costera del pacífico provinces, which together accounted for 27% of the total area, 28.1% of grid cells exhibited c index values ≥0.5. in general, the proportion of grid cells with c index values ≥0.8 ranged from 2.6% to 9.7% across these provinces. when the territory of the physiographic provinces was considered globally, the c index values exceeded 0.72 in all cases. the eje neovolcánico and sierra madre del sur provinces could be considered fully inventoried (c ≥ 0.8), while the two remaining provinces exhibited c index values indicating that they were partially inventoried. in terms of completeness, these followed the order of sierra madre occidental > llanura costera del pacífico (fig. 3d). it is worth noting that while the neotropical region exhibited the highest number of reptile records and the greatest species richness, this pattern was not mirrored among the physiographic provinces. the llanura costera del pacífico province had the highest number of reptile records, but the eje neovolcánico province exhibited the highest species richness (100 species; table 2). in relation to the above, a close association was observed between the grid cells with c index values and the locations of urban and rural settlements and road infrastructure. the proportion of records located near the main roads of nayarit was estimated at perpendicular distances of 100 and 250 m, with 34% and 51% of the records, respectively, within these limits (fig. 4). according to this distribution, approximately 72% of grid cells with c index values ≥0.5 (61% of the cells had c index values of 0.5–0.8, and 11% of those had c index values ≥0.8) were located near towns and roads. table 1. number of reptile species records (rsr), number of cells with records (cwr), and number of cells by inventory completeness index (c) category by (a) biogeographic region (br) and (b) physiographic province (php). biogeographic regions: neotropical (ntp) and mexican transition zone (mtz). physiographic provinces: sierra madre del sur (sms), llanura costera del pacífico (lcp), eje neovolcánico (env), and sierra madre occidental (smo). (a) br cells rsr cwr c index category 0 < c < 0.5 0.5 ≤ c < 0.8 c ≥ 0 .8 ntp 183 4616 128 27 42 8 mtz 173 1424 87 10 20 4 (b) php sms 31 1340 18 4 7 3 lcp 65 2287 55 12 15 2 env 70 1741 57 12 23 2 smo 190 672 85 9 17 5 table 2. number of reptile species records (rsr), observed species richness (sobs), number of species found once (a), number of species found twice (b), expected species richness (sexp), and inventory completeness index (c) by biogeographic region, biogeographic province, and physiographic province of nayarit. type name rsr sobs a b sexp c bdc adc biogeographic region and province* mexican transition zone (1) 210 63 55 22 15 71 0.773 mexican transition zone (2) 1214 102 85 21 13 101 0.833 neotropical (3) 4616 151 110 24 11 136 0.808 physiographic provinces sierra madre occidental 672 105 88 27 16 111 0.794 llanura costera del pacífico 2287 106 80 25 10 111 0.719 eje neovolcanico 1741 125 100 20 17 112 0.895 sierra madre del sur 1340 65 58 10 4 71 0.823 * (1) sierra madre occidental, (2) trans-mexican volcanic belt, and (3) pacific lowlands biogeographic provinces (morrone 2019). bdc: species richness before data cleaning; adc: species richness after data cleaning. andrew peterson line maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 53 the analysis of the species records per grid cell across different time periods indicated that only 18% of the 356 grid cells contained species records during the earliest period (1861–1960), with c index values ≥0.5 in ~3.3% of the territory (table 3). during this period, most records originated from research that had been primarily conducted by international institutions, such as the museum of comparative zoology at harvard university, the natural history museum in london, the san diego natural history museum, and the biodiversity institute and natural history museum at the university of kansas, with smaller contributions from national institutions. for the following period (1961–1980), the number of reptile records increased to include 22% of the grid cells, although those with c index values ≥0.5 represented only 3.9% of nayarit, reflecting only a marginal increase from the prefigure 4. records within the 100-m and 250-m buffers on both sides of the main roads in nayarit. table 3. number of reptile species records (rsr), number of cells with records (cwr), and number of cells by inventory completeness index (c) category by time period. period rsr cwr c index category 0 < c < 0.5 0.5 ≤ c < 0.8 c ≥ 0.8 1861–1960 969 64 11 8 4 1961–1980 1155 81 8 13 1 1981–2000 310 61 3 5 8 2001–2020 2012 167 20 27 14 2021–2023 1594 116 9 19 13 vious period. during 1961–1980, most contributions came primarily from foreign institutions, such as the natural history museum of los angeles county, the museum of natural history at the university of illinois, and the museum of vertebrate zoology at the university of california, berkeley, with limited participation from national institutions. in the next period (1981–2000), the number of grid cells with records declined to the lowest proportion in the entire study (17%). this period was characterized by a greater presence of national institutions, such as museo de zoología alfonso l. herrera of the facultad de ciencias and instituto de biología of the universidad nacional autónoma de méxico (unam), with a smaller contribution from international institutions. in the two most recent periods (2001–2020 and 2021– 2023), the number of records increased substantially andrew peterson line maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 54 (46.9% and 32.6%, respectively), despite the most recent period covering only three years. this increase was mainly attributed to the participation of citizen science, particularly through the inaturalistmx4 platform, along with contributions from national and international academic institutions. an increase in inventory completeness was also evident, with c index values ≥0.5 in 11.5% (2001–2020) and 9.0% (2021–2023) of the grid cells, while the proportion of cells with c index values ≥0.8 decreased to just under 4% for both periods. when integrating the entire study period, 20.8% of grid cells exhibited c index values ≥0.5, with only 3.4% having c index values ≥0.8 (fig. 5). discussion gaps in the taxonomic, spatial, and temporal knowledge of the reptile inventory in nayarit were identified based on information stored in the snib database of conabio, which includes records from the mid-19th century to the present. the information generated in the present study is representative of the actual conditions in nayarit and serves as a baseline for future research aimed at assessing reptile diversity in the state, which is crucial for developing adequate conservation measures and managing natural resources. the gap in taxonomic knowledge represents a potential lack of information for all species; in this study, specifically the reptiles found in nayarit. the data adjustment for newly incorporated reptile species showed a sigmoid trend with a reliable confidence level, indicating that the theoretical maximum number of species inhabiting the state has nearly been reached. however, the estimated growth rate of new species records for the past thirty years suggests that approximately one new species has been added to the list every four years. despite these results, since 2000, eight new species (order squamata) have been described that belong to the families phrynosomatidae (sceloporus huichol flores-villela, smith, campillo-garcía, martínez-méndez & campbell, 2022), phyllodactylidae (phyllodactylus cleofasensis ramírezreyes, barraza-soltero, nolasco-luna, flores-villela & escobedo-galván, 2021), scincidae (marisora aquilonaria mccranie, matthews & hedges, 2020), colubridae (tantilla ceboruca canseco-márquez, smith, ponce-campos, flores-villela & campbell, 2007), natricidae (thamnophis rossmani conant, 2000), viperidae (crotalus campbelli bryson jr, linkem, dorcas, lathrop, jones, alvarado-díaz, grünwald & murphy, 2014), and kinosternidae (kinosternon vogti lópez-luna, cupul-magaña, escobedo-galván, gonzález-hernández, centenero-alcalá, rangel-mendoza, ramírez-ramírez & cazares-hernán4 https://mexico.inaturalist.org/ dez, 2018; kinosternon cora loc-barragán, reyes-velasco, woolrich-piña, grünwald, venegas de anaya, rangel-mendoza & lópez-luna, 2020). as of 2020, four of these new species were included in the herpetofauna of nayarit due to morphological analyses and molecular techniques employed to distinguish cryptic species (loc-barragán et al., 2024). it was assumed that these species were previously captured and recorded for the state inventory but misidentified. additionally, loc-barragán et al. (2024) compiled a list of 19 reptile species distributed in areas adjacent to nayarit, namely in the states of jalisco, zacatecas, durango, and sinaloa, which may eventually be incorporated into the records of nayarit. these cases were not considered in the methods used to estimate the taxonomic gap, as they only included formally accredited records; thus, future adjustments may be needed. nonetheless, the taxonomic gap for the reptiles of nayarit is close to being resolved. the analysis of spatial gaps indicated that ~40% of the grid cells in nayarit lacked reptile records, particularfigure 5. inventory completeness of the reptiles in nayarit for each time period. the gray area represents cells where the inventory completeness (c) index could not be measured or where no reptile records were present. https://mexico.inaturalist.org/ andrew peterson line maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 55 ly those in mountainous and difficult-to-access regions. these areas may serve as refuges that host high endemism due to their isolation and environmental conditions, even in tropical zones, which makes them more sensitive to environmental changes (spehn, 2011; silveira et al., 2019). additionally, a correlation was observed between the number of records and the proximity to roads and human settlements (roads within urban centers). this sampling bias has been documented by kadmon et al. (2004), who noted that random surveys over large areas are rare, leading to species distribution models often being based on biased and incomplete data. similarly, the areas with the highest c index values (≥0.8) were mainly located near urban areas (tepic, santiago ixcuintla, ruíz, nuevo vallarta, sayulita, santa maría del oro, and guadalupe victoria), a trend that has been observed in other regions in which well-surveyed grid cells have been geographically associated with cities, rivers, and major roads (ballesteros-mejía et al., 2013; stropp et al., 2016; arvizu & ruiz-luna, 2024). it is likely that the greater inventory completeness associated with certain biogeographic regions or physiographic provinces, such as the neotropical region or eje neovolcánico, results more from their accessibility due to the presence of roads and urban areas than from the use of effective sampling strategies. this scenario has also been suggested by ficetola et al. (2014) when evaluating reptile records on mediterranean islands. regardless of this bias, the neotropical region, which includes humid and semi-humid tropical areas, was identified as having a higher proportion of cells with species records and higher inventory completeness (c ≥ 0.8) than the mexican transition zone, where nearctic and neotropical biota overlap. given its characteristics, the mexican transition zone would be expected to have a higher level of endemism (morrone, 2019). the physiographic province of sierra madre occidental, the largest and most rugged province located in eastern nayarit, exhibited the greatest geographic gap. in contrast, sierra madre del sur, the smallest province, exhibited the highest proportion of fully inventoried grid cells and, simultaneously, a higher c index value at the regional scale. the eje neovolcánico province exhibited a lower proportion of fully inventoried grid cells and, in contrast, exhibited the highest c index value at the regional level. additionally, eje neovolcánico was identified as the province with the greatest reptile diversity in nayarit (loc-barragán et al. 2024). at the regional scale, sierra madre occidental also exhibited a high c index value despite having the largest geographic gap, suggesting that using this value at the regional scale to estimate completeness is not advisable, as conclusions may be highly biased. finally, the temporal gap analysis provided insights into how the contributions to the reptile inventory of nayarit have evolved over time, as well as the main sources of data. of the 6,040 verified reptile records for the continental and insular zones of nayarit, ~60% were contributed during recent periods (2001–2020 and 2021–2023), which coincides with the increase in citizen science initiatives. with citizen science contributions, which are typically hosted through open-access online platforms, studies like the present one are able to compile a larger volume of data for analysis. furthermore, given that citizen science initiatives are based on public participation, they inherently increase the number of opportunities to improve environmental education efforts focused on conserving biodiversity (peter et al., 2019). however, there are potential risks or challenges associated with citizen science in biodiversity documentation that must be addressed, such as observer differences, reporting preferences, false positive errors, data validation, and detectability (johnston et al., 2023). during the period of 1981–2000, the lowest number of records were contributed to the reptile inventory of nayarit; therefore, the fewest number of grid cells with c index values ≥0.8 were present. this result may be explained by the data sources of the period, which were mainly national institutions. during the first two periods (1861–1960 and 1961–1980), data were primarily sourced from a diverse range of international and national institutions. thus, the trend observed during 1981–2000 suggests that a lower sampling effort was present during the period, which was possibly associated with the reduced participation of international institutions (meyer et al., 2015) and, consequently, reduced financial support for biodiversity conservation. this trend has been observed in various countries across africa, the americas, asia, europe, and oceania (waldron et al., 2013). however, the efforts of researchers, educational institutions, and research centers in mexico should be recognized for their contributions to biological inventories despite the challenges associated with securing funding for such studies when they are not prioritized in national policies or budgets. studies like the present one are important for identifying knowledge gaps in biological inventories, establishing a foundation for future research, and optimizing resources to fill notable gaps. importantly, the results of the present study indicate that few reptile species remain to be described in the state of nayarit. in addition, records of this terrestrial vertebrate group are influenced by accessibility maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 56 to sampling areas and proximity to urbanization. thus, future sampling efforts should focus on remote, mountainous, and difficult-to-access areas. lastly, the use of citizen science applications to increase the number of biodiversity records over the past two decades has notably contributed to the completeness of the reptile inventory in nayarit. acknowledgments this work was supported by the board of trustees of the universidad autónoma de nayarit and conahcyt [cvu 102166], which provided doctoral scholarships to the first author. we thank the centro en investigación en alimentación y desarrollo, a. c. and the laboratorio de manejo ambiental for aiding in the development of this study. we thank a.l. mactavish for english language editing. competing interests the authors have declared that no competing interests exist. references aguilar-lópez, j. l., pineda, e., luría-manzano, r., & canseco-márquez, l. (2016). species diversity, distribution, and conservation status in a mesoamerican region: amphibians of the uxpanapa-chimalapas region, mexico. tropical conservation science, 9(4), 1940082916670003. https:// doi.org/10.1177/1940082916670003 alroy, j. (2015). current extinction rates of reptiles and amphibians. proceedings of the national academy of sciences, 112(42), 13003-13008. https://doi. org/10.1073/pnas.1508681112 arvizu, m. d., & ruiz-luna, a. (2024). completeness of the amphibian inventory of nayarit, mexico, assessed from biodiversity information system records. ecoscience, 31(3), 165-178. https://doi.org/10.1080/1195 6860.2024.2426335 asase a., & peterson a. t. (2016). completeness of digital accessible knowledge of the plants of ghana. biodiversity informatics. 11(1), 1-11. https://doi. org/10.17161/bi.v11i1.5860 ballesteros‐mejia, l., kitching, i. j., jetz, w., nagel, p., & beck, j. (2013). mapping the biodiversity of tropical insects: species richness and inventory completeness of african sphingid moths. global ecology and biogeography, 22(5), 586-595. https://doi.org/10.1111/ geb.12039 baselga, a., & novoa, f. (2006). diversity of chrysomelidae (coleoptera) in galicia, northwest spain: estimating the completeness of the regional inventory. biodiversity & conservation, 15, 205-230. https:// doi.org/10.1007/s10531-004-6904-x beck, h. e., zimmermann, n. e., mcvicar, t. r., vergopolan, n., berg, a., & wood, e. f. (2018). present and future köppen-geiger climate classification maps at 1-km resolution. scientific data, 5(1), 1-12. https:// doi.org/10.1038/sdata.2018.214 booth, d. t. (2006). influence of incubation temperature on hatchling phenotype in reptiles. physiological and biochemical zoology, 79(2), 274-281. https://doi. org/10.1086/499988 cec. (2023). “2020 land cover of north america at 30 meters”. commission for environmental cooperation. north american land change monitoring system. canada centre for remote sensing (ccrs), u.s. geological survey (usgs), comisión nacional para el conocimiento y uso de la biodiversidad (conabio), comisión nacional forestal (conafor), instituto nacional de estadística y geografía (inegi). ed. 1.0, raster digital data [30-m]. [accessed 2024 jul 10]. http://www.cec.org/north-american-environmental-atlas/land-cover-30m-2020/ chao, a. (1984). nonparametric estimation of the number of classes in a population. scandinavian journal of statistics, 265-270. chesshire, p. r., fischer, e. e., dowdy, n. j., griswold, t. l., hughes, a. c., orr, m. c., ascher, j. s., guzman, l. m., hung, k. l. j., cobb & mccabe, l. m. (2023). completeness analysis for over 3000 united states bee species identifies persistent data gap. ecography, 2023(5), e06584. https://doi.org/10.1111/ecog.06584 colwell, r. k., & coddington, j. a. (1994). estimating terrestrial biodiversity through extrapolation. philosophical transactions of the royal society b: biological sciences, 345, 101-118. https://doi.org/10.1098/ rstb.1994.0091 conabio. (2024). data from: sistema nacional de información sobre biodiversidad. registros de ejemplares [dataset version 2024 feb 26]. comisión nacional para el conocimiento y uso de la biodiversidad. [accessed 2024 jul 08]. https://www.snib.mx/ejemplares/ descarga/ cortés-gomez, a. m., ruiz-agudelo, c. a., valencia-aguilar, a., & ladle, r. j. (2015). ecological functions of neotropical amphibians and reptiles: a review. universitas scientiarum, 20(2), 229-245. https://doi.org/10.11144/javeriana.sc20-2.efna cox, n., young, b. e., bowles, p., fernandez, m., marin, j., rapacciuolo, g., böhm, m., brooks, t. m., hedges, s. b., hilton-taylor, c., hoffmann, m., jenkins, r. k. b., tognelli, m. f., alexander, g. j., allison, a., ananjeva, n. b., auliya, m., avila, l. j., chapple, d. g., cisneros-heredia, d. f., cogger, h. g., colli, g. r., de silva, a., eisemberg, c. c., els, j., fong g., a., https://doi.org/10.1177/1940082916670003 https://doi.org/10.1177/1940082916670003 https://doi.org/10.1073/pnas.1508681112 https://doi.org/10.1073/pnas.1508681112 https://doi.org/10.1080/11956860.2024.2426335 https://doi.org/10.1080/11956860.2024.2426335 https://doi.org/10.17161/bi.v11i1.5860 https://doi.org/10.17161/bi.v11i1.5860 https://doi.org/10.1111/geb.12039 https://doi.org/10.1111/geb.12039 https://doi.org/10.1007/s10531-004-6904-x https://doi.org/10.1007/s10531-004-6904-x https://doi.org/10.1086/499988 https://doi.org/10.1086/499988 http://www.cec.org/north-american-environmental-atlas/land-cover-30m-2020/ http://www.cec.org/north-american-environmental-atlas/land-cover-30m-2020/ https://doi.org/10.1111/ecog.06584 https://doi.org/10.1098/rstb.1994.0091 https://doi.org/10.1098/rstb.1994.0091 https://www.snib.mx/ejemplares/descarga/ https://www.snib.mx/ejemplares/descarga/ https://doi.org/10.11144/javeriana.sc20-2.efna maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 57 grant, t. d., hitchmough, r. a., iskandar, d. t., kidera, n., martins, m., meiri, s., mitchell, n. j., molur, s., nogueira, c. de c., ortiz, j. c., penner, j., rhodin, a. g. j., rivas, g. a., rödel, m. o., roll, u., sanders, k. l., santos-barrera, g., shea, g. m. spawls, s., stuart, b. l., tolley, k. a., trape, j. f., vidal, m. a., wagner, p., wallace, b. p., & xie, y. (2022). a global reptile assessment highlights shared conservation needs of tetrapods. nature, 605(7909), 285-290. https://doi.org/10.1038/s41586-022-04664-7 de miranda, e. b. (2017). the plight of reptiles as ecological actors in the tropics. frontiers in ecology and evolution, 5, 309533. https://doi.org/10.3389/ fevo.2017.00159 dirzo, r., ceballos, g., & ehrlich, p. r. (2022). circling the drain: the extinction crisis and the future of humanity. philosophical transactions of the royal society b, 377(1857), 20210378. https://doi.org/10.1098/ rstb.2021.0378 dobson, a., & barnett, a.g. (2018) an introduction to general linear models. chapman & hall/crc press. escribano n., galicia d., & ariño a. h. (2019). completeness of digital accessible knowledge (dak) about terrestrial mammals in the iberian peninsula. plos one, 14(3): e0213542. https://doi.org/10.1371/ journal.pone.0213542 farooq, h., harfoot, m., rahbek, c., & geldmann, j. (2024). threats to reptiles at global and regional scales. current biology, 34(10), 2231-2237.e2. https://doi.org/10.1016/j.cub.2024.04.007 ficetola, g. f., cagnetta, m., padoa‐schioppa, e., quas, a., razzetti, e., sindaco, r., & bonardi, a. (2014). sampling bias inverts ecogeographical relationships in island reptiles. global ecology and biogeography, 23(11), 1303-1313. https://doi.org/10.1111/geb.12201 finn, c., grattarola, f., & pincheira‐donoso, d. (2023). more losers than winners: investigating anthropocene defaunation through the diversity of population trends. biological reviews, 98(5), 1732-1748. https:// doi.org/10.1111/brv.12974 flores-villela, o., & garcía-vázquez, u. o. (2014). biodiversity of reptiles in mexico. revista mexicana de biodiversidad, 85, s467-s475. https://doi. org/10.7550/rmb.43236 flores-villela, o., & goyenechea, i. (2003). patrones de distribución de anfibios y reptiles en méxico. una perspectiva latinoamericana de la biogeografía, 289296. fox, j. (2015) applied regression analysis and generalized linear models. sage publications. ganglo, j. c., & kakpo, s. b. (2016). completeness of digital accessible knowledge of plants of benin and priorities for future inventory and data discovery. biodiversity informatics, 11(1), 23-39. https://doi. org/10.17161/bi.v11i1.5053 gonzález, j., baselga, a., & novoa, f. (2007). diversity of water beetles (coleoptera: gyrinidae, haliplidae, noteridae, hygrobiidae, dytiscidae, and hydrophilidae) in galicia, northwest spain: estimating the completeness of the regional inventory. the coleopterists bulletin, 61(1), 95-110. https://doi.org/10.1649/919.1 grattarola, f., martínez-lanfranco, j. a., botto, g., naya, d. e., maneyro, r., mai, p., hernández, d., laufer, g., ziegler, l., gonzález, e. m., da-rosa, i., gobel, n., gonzález, a., gonzález, j., rodales, a. l., & pincheira-donoso, d. (2020). multiple forms of hotspots of tetrapod biodiversity and the challenges of open-access data scarcity. scientific reports, 10(1), 1-15. https://doi.org/10.1038/s41598-020-79074-8 hortal, j., de bello, f., diniz-filho, j. a. f., lewinsohn, t. m., lobo, j. m., & ladle, r. j. (2015). seven shortfalls that beset large-scale knowledge of biodiversity. annual review of ecology, evolution, and systematics, 46, 523-549. https://doi.org/10.1146/annurev-ecolsys-112414-054400 huang, x., lin, c., & ji, l. (2020). the persistent multi-dimensional biases of biodiversity digital accessible knowledge of birds in china. biodiversity and conservation, 29, 3287-3311. https://doi.org/10.1007/ s10531-020-02024-3 inegi. (2022). anuario estadístico y geográfico por entidad federativa 2021. instituto nacional de estadística y geografía. inegi. méxico. 644 pp. isaac, n. j., & pocock, m. j. (2015). bias and information in biological records. biological journal of the linnean society, 115(3), 522-531. https://doi.org/10.1111/ bij.12532 iucn. (2024). the iucn red list of threatened species. international union for conservation of nature. version 2024-1. [accessed 2024 jul 5]. https://www. iucnredlist.org johnston, a., matechou, e., & dennis, e. b. (2023). outstanding challenges and future directions for biodiversity monitoring using citizen science data. methods in ecology and evolution, 14(1), 103-116. https:// doi.org/10.1111/2041-210x.13834 johnson, j. d., wilson, l. d., mata-silva, v., garcía-padilla, e., & desantis, d. l. (2017). the endemic herpetofauna of mexico: organisms of global significance in severe peril. mesoamerican herpetology, 4(3), 544-620. https://doi.org/10.1038/s41586-022-04664-7 https://doi.org/10.3389/fevo.2017.00159 https://doi.org/10.3389/fevo.2017.00159 https://doi.org/10.1098/rstb.2021.0378 https://doi.org/10.1098/rstb.2021.0378 https://doi.org/10.1371/journal.pone.0213542 https://doi.org/10.1371/journal.pone.0213542 https://doi.org/10.1016/j.cub.2024.04.007 https://doi.org/10.1111/geb.12201 https://doi.org/10.1111/brv.12974 https://doi.org/10.1111/brv.12974 https://doi.org/10.7550/rmb.43236 https://doi.org/10.7550/rmb.43236 https://doi.org/10.17161/bi.v11i1.5053 https://doi.org/10.17161/bi.v11i1.5053 https://doi.org/10.1649/919.1 https://doi.org/10.1038/s41598-020-79074-8 https://doi.org/10.1146/annurev-ecolsys-112414-054400 https://doi.org/10.1146/annurev-ecolsys-112414-054400 https://doi.org/10.1007/s10531-020-02024-3 https://doi.org/10.1007/s10531-020-02024-3 https://doi.org/10.1111/bij.12532 https://doi.org/10.1111/bij.12532 https://www.iucnredlist.org https://www.iucnredlist.org https://doi.org/10.1111/2041-210x.13834 https://doi.org/10.1111/2041-210x.13834 maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 58 kadmon, r., farber, o., & danin, a. (2004). effect of roadside bias on the accuracy of predictive maps produced by bioclimatic models. ecological applications, 14(2), 401-413. https://doi.org/10.1890/025364 lajeunesse, a., & fourcade, y. (2023). temporal analysis of gbif data reveals the restructuring of communities following climate change. journal of animal ecology, 92(2), 391-402. https://doi.org/10.1111/13652656.13854 lemos-espinal, j. a., & smith, g. r. (2020). a checklist of the amphibians and reptiles of sinaloa, mexico with a conservation status summary and comparisons with neighboring states. zookeys, 931, 85. https:// doi.org/10.3897/zookeys.931.50922 lemos-espinal, j. a., & smith, g. r. (2023). an analysis of the inter-state similarity of the herpetofaunas of mexican states. nature conservation, 53, 223-256. https:// doi.org/10.3897/natureconservation.53.106732 lobo, j. m., hortal, j., yela, j. l., millán, a., sánchez-fernández, d., garcía-roselló, e., gonzález-dacosta, j., heine, j., gonzález-vilas, l., & guisande, c. (2018). knowbr: an application to map the geographical variation of survey effort and identify well-surveyed areas from biodiversity databases. ecological indicators, 91, 241-248. https://doi. org/10.1016/j.ecolind.2018.03.077 loc-barragán, j. a., smith, g. r., woolrich-piña, g. a., & lemos-espinal, j. a. (2024). an updated checklist of the amphibians and reptiles of nayarit, mexico with conservation status and comparison with adjoining states. herpetozoa, 37, 25-42. https://doi. org/10.3897/herpetozoa.37.e112093 loiselle, b. a., jørgensen, p. m., consiglio, t., jiménez, i., blake, j. g., lohmann, l. g., & montiel, o. m. (2008). predicting species distributions from herbarium collections: does climate bias in collection sampling influence model outcomes?. journal of biogeography, 35(1), 105-116. https://doi.org/10.1111/ j.1365-2699.2007.01779.x luja, v. h., ahumada-carrillo, i., ponce-campos, p., & figueroa-esquivel, e. (2014). checklist of amphibians of nayarit, western mexico. check list, 10(6), 1336-1341. https://doi.org/10.15560/10.6.1336 luo, m., xu, z., hirsch, t., aung, t. s., xu, w., ji, l., quin, h., & ma, k. (2021). the use of global biodiversity information facility (gbif)-mediated data in publications written in chinese. global ecology and conservation, 25, e01406. https://doi.org/10.1016/j. gecco.2020.e01406 marshall, b. m., strine, c., & hughes, a. c. (2020). thousands of reptile species threatened by under-regulated global trade. nature communications, 11(1), 1-12. https://doi.org/10.1038/s41467-020-18523-4 marshall, l., leclercq, n., carvalheiro, l. g., dathe, h. h., jacobi, b., kuhlmann, m., potts, s. g., rasmont, p., roberts, s. p. m., & vereecken, n. j. (2024). understanding and addressing shortfalls in european wild bee data. biological conservation, 290, 110455. https://doi.org/10.1016/j.biocon.2024.110455 microsoft corporation. 2013. microsoft excel. [online version] meyer, c., kreft, h., guralnick, r., & jetz, w. (2015). global priorities for an effective information basis of biodiversity distributions. nature communications, 6(1), 1-8. https://doi.org/10.1038/ncomms9221 millennium ecosystem assessment (2005). ecosystems and human well-being: wetlands and water. world resources institute, washington, dc. morrone, j. (2019). regionalización biogeográfica y evolución biótica de méxico: encrucijada de la biodiversidad del nuevo mundo. revista mexicana de biodiversidad, 90:e902980. https://doi.org/10.22201/ ib.20078706e.2019.90.2980 newbold, t. (2018). future effects of climate and landuse change on terrestrial vertebrate community diversity under different scenarios. proceedings of the royal society b, 285(1881), 20180792. https://doi. org/10.1098/rspb.2018.0792 neves, i. q., mathias, m. d. l., & bastos-silveira, c. (2019). mapping knowledge gaps of mozambique’s terrestrial mammals. scientific reports, 9(1), 1-14. https://doi.org/10.1038/s41598-019-54590-4 nijman, v., shepherd, c. r., & sanders, k. l. (2012). over-exploitation and illegal trade of reptiles in indonesia. the herpetological journal, 22(2), 83-89. nori, j., cordier, j. m., osorio‐olvera, l., & hortal, j. (2023). global knowledge gaps of herptile responses to land transformation. frontiers in ecology and the environment, 21(9), 411-417. https://doi.org/10.1002/ fee.2625 okoh, g. s. r., horwood, p. f., whitmore, d., & ariel, e. (2021). herpesviruses in reptiles. frontiers in veterinary science, 8, 642894. https://doi.org/10.3389/ fvets.2021.642894 peter, m., diekötter, t., & kremer, k. (2019). participant outcomes of biodiversity citizen science projects: a systematic literature review. sustainability, 11(10), 2780. https://doi.org/10.3390/su11102780 https://doi.org/10.1890/02-5364 https://doi.org/10.1890/02-5364 https://doi.org/10.1111/1365-2656.13854 https://doi.org/10.1111/1365-2656.13854 https://doi.org/10.3897/zookeys.931.50922 https://doi.org/10.3897/zookeys.931.50922 https://doi.org/10.3897/natureconservation.53.106732 https://doi.org/10.3897/natureconservation.53.106732 https://doi.org/10.1016/j.ecolind.2018.03.077 https://doi.org/10.1016/j.ecolind.2018.03.077 https://doi.org/10.3897/herpetozoa.37.e112093 https://doi.org/10.3897/herpetozoa.37.e112093 https://doi.org/10.1111/j.1365-2699.2007.01779.x https://doi.org/10.1111/j.1365-2699.2007.01779.x https://doi.org/10.15560/10.6.1336 https://doi.org/10.1016/j.gecco.2020.e01406 https://doi.org/10.1016/j.gecco.2020.e01406 https://doi.org/10.1038/s41467-020-18523-4 https://doi.org/10.1016/j.biocon.2024.110455 https://doi.org/10.1038/ncomms9221 https://doi.org/10.22201/ib.20078706e.2019.90.2980 https://doi.org/10.22201/ib.20078706e.2019.90.2980 https://doi.org/10.1098/rspb.2018.0792 https://doi.org/10.1098/rspb.2018.0792 https://doi.org/10.1038/s41598-019-54590-4 https://doi.org/10.1002/fee.2625 https://doi.org/10.1002/fee.2625 https://doi.org/10.3389/fvets.2021.642894 https://doi.org/10.3389/fvets.2021.642894 https://doi.org/10.3390/su11102780 maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 59 peterson a. t., navarro-sigüenza a. g., & martínez-meyer, e. (2016). digital accessible knowledge and wee-inventories sites for birds in mexico: baseline sites for measuring faunistic change. peerj. 4:e2362. https://doi.org/10.7717/peerj.2362 pieau, c. (1996). temperature variation and sex determination in reptiles. bioessays, 18(1), 19-26. https:// doi.org/10.1002/bies.950180107 pyšek, p., hulme, p. e., simberloff, d., bacher, s., blackburn, t. m., carlton, j. t., dawson, w., essl, f., foxcroft, l. c., genovesi, p., jeschke, j. m., kühn, i., liebhold, a. m., mandrak, n. e., meyerson, l. a., pauchard, a., pergl, j., roy, h. e., seebens, h., van kleunen, m., vilà, m., wingfield, m. j., & richardson, d. m. (2020). scientists’ warning on invasive alien species. biological reviews, 95(6), 1511-1534. https://doi.org/10.1111/brv.12627 qgis.org. 2025. qgis geographic information system. qgis association. http://www.qgis.org ramírez-bautista, a., torres-hernández, l. a., cruz-elizalde, r., berriozabal-islas, c., hernández-salinas, u., wilson, l. d., johnson, j. d., porras, l. w., balderas-valdivia, c. j., gonzález-hernández, a. j. x., & mata-silva, v. (2023). an updated list of the mexican herpetofauna: with a summary of historical and contemporary studies. zookeys, 1166, 287. https://doi. org/10.3897/zookeys.1166.86986 rull, v. (2022). biodiversity crisis or sixth mass extinction? does the current anthropogenic biodiversity crisis really qualify as a mass extinction?. embo reports, 23(1), e54193. https://doi.org/10.15252/ embr.202154193 sandor, m. e., elphick, c. s., & tingley, m. w. (2022). extinction of biotic interactions due to habitat loss could accelerate the current biodiversity crisis. ecological applications, 32(6), e2608. https://doi.org/10.1002/ eap.2608 schilliger, l., paillusseau, c., françois, c., & bonwitt, j. (2023). major emerging fungal diseases of reptiles and amphibians. pathogens, 12(3), 429. https://doi. org/10.3390/pathogens12030429 semarnat. (2010). norma oficial mexicana nom059-semarnat-2010: protección ambiental-especies nativas de méxico de flora y fauna silvestres-categorías de riesgo y especificaciones para su inclusión, exclusión o cambio-lista de especies en riesgo secretaría del medio ambiente y recursos naturales. méxico: diario oficial de la federación, 2010 dec 30. shirey, v., belitz, m. w., barve, v., & guralnick, r. (2021). a complete inventory of north american butterfly occurrence data: narrowing data gaps, but increasing bias. ecography, 44(4), 537-547. https://doi. org/10.1111/ecog.05396 silveira, f. a. o., barbosa, m., beiroz, w., callisto, m., macedo, d. r., morellato, l. p. c., neves, f. s., nunes, y, r. f., solar, r. r., & fernandes, g. w. (2019). tropical mountains as natural laboratories to study global changes: a long-term ecological research project in a megadiverse biodiversity hotspot. perspectives in plant ecology, evolution and systematics, 38, 64-73. https://doi.org/10.1016/j.ppees.2019.04.001 smith, g. r., & lemos-espinal, j. a. (2022). factors related to species richness, endemism, and conservation status of the herpetofauna (amphibia and reptilia) of mexican states. zookeys 1097: 85–101. https://doi. org/10.3897/zookeys.1097.80424 spehn, e. m., rudmann-maurer, k., & körner, c. (2011). mountain biodiversity. plant ecology & diversity, 4(4), 301-302. https://doi.org/10.1080/17550874.201 2.698660 soberón, j., jiménez, r., golubov, j., & koleff, p. (2007). assessing completeness of biodiversity databases at different spatial scales. ecography, 30(1), 152-160. https://doi.org/10.1111/j.0906-7590.2007.04627.x sousa-baena, a. s., garcia, l. c., & peterson, a. t. (2014). completeness of digital accessible knowledge of the plants of brazil and priorities for survey and inventory. diversity and distributions. 20, 369-381. https://doi.org/10.1111/ddi.12136 stropp, j., ladle, r. j., malhado, a. c. m., hortal, j., gaffuri, j., temperley, w. h., skøien, j. o., & mayaux, p. (2016). mapping ignorance: 300 years of collecting flowering plants in africa. global ecology and biogeography, 25(9), 1085-1096. https://doi.org/10.1111/ geb.12468 suazo-ortuño, i., ramírez-bautista, a., & alvarado-díaz, j. (2023). amphibians and reptiles of mexico: diversity and conservation. in mexican fauna in the anthropocene (pp. 105-127). cham: springer international publishing. tessarolo, g., ladle, r., rangel, t., & hortal, j. (2017). temporal degradation of data limits biodiversity research. ecology and evolution, 7(17), 6863-6870. https://doi.org/10.1002/ece3.3259 tjørve, e. (2003). shapes and functions of species–area curves: a review of possible models. journal of biogeography, 30(6), 827-835. https://doi.org/10.1046/ j.1365-2699.2003.00877.x troia, m. j., & mcmanamay, r. a. (2016). filling in the gaps: evaluating completeness and coverage of open-access biodiversity databases in the united states. ecology and evolution. 6(14), 4654-4669. https://doi.org/10.1002/ece3.2225 https://doi.org/10.1002/bies.950180107 https://doi.org/10.1002/bies.950180107 https://doi.org/10.1111/brv.12627 http://www.qgis.org https://doi.org/10.3897/zookeys.1166.86986 https://doi.org/10.3897/zookeys.1166.86986 https://doi.org/10.15252/embr.202154193 https://doi.org/10.15252/embr.202154193 https://doi.org/10.1002/eap.2608 https://doi.org/10.1002/eap.2608 https://doi.org/10.3390/pathogens12030429 https://doi.org/10.3390/pathogens12030429 https://doi.org/10.1111/ecog.05396 https://doi.org/10.1111/ecog.05396 https://doi.org/10.1016/j.ppees.2019.04.001 https://doi.org/10.3897/zookeys.1097.80424 https://doi.org/10.3897/zookeys.1097.80424 https://doi.org/10.1080/17550874.2012.698660 https://doi.org/10.1080/17550874.2012.698660 https://doi.org/10.1111/j.0906-7590.2007.04627.x https://doi.org/10.1111/ddi.12136 https://doi.org/10.1111/geb.12468 https://doi.org/10.1111/geb.12468 https://doi.org/10.1002/ece3.3259 https://doi.org/10.1046/j.1365-2699.2003.00877.x https://doi.org/10.1046/j.1365-2699.2003.00877.x https://doi.org/10.1002/ece3.2225 maría daniela arvizu et al. – evaluating knowledge gaps in reptile records in nayarit 60 uetz, p., freed, p, aguilar, r., reyes, f., kudera, j. & hošek, j. (eds.) (2025) the reptile database. [accessed 2025 feb 24]. http://www.reptile-database.org valdez-rentería, s. y., domínguez-vega, h., trujillo-mendoza, v., morón-garcía, c. e., gómez-ortiz, y., fernández-badillo, l., & gómez-sánchez, d. (2023). importancia cultural de las serpientes: más allá de la visión ecológica. herpetología mexicana, 5: 1-16. https://doi.org/10.69905/xbqkar19 valencia-aguilar, a., cortés-gómez, a. m., & ruiz-agudelo, c. a. (2013). ecosystem services provided by amphibians and reptiles in neotropical ecosystems. international journal of biodiversity science, ecosystem services & management, 9(3), 257272. https://doi.org/10.1080/21513732.2013.821168 waldron, a., mooers, a. o., miller, d. c., nibbelink, n., redding, d., kuhn, t. s., roberts, j. t., & gittleman, j. l. (2013). targeting global conservation funding to limit immediate biodiversity declines. proceedings of the national academy of sciences, 110(29), 1214412148. https://doi.org/10.1073/pnas.1221370110 wheeler, q. d., knapp, s., stevenson, d. w., stevenson, j., blum, s. d., boom, b. m., borisy, g. g., buizer, j. l., de carvalho, m. r., cibrian, a., donoghue, m. j., doyle, v., gerson, e. m., graham, c. h., graves, p., graves, s. j., guralnick, r. p., hamilton, a. l., hanken, j., law, w., lipscomb, d. l., lovejoy, t. e., miller, h., miller, j.s., naeem, s., novacek, m. j., page, l. m., platnick, n. i., porter-morgan, h., raven, p. h., solis, m. a., valdecasas, a. g., van der leeuw, s., vasco, a., vermeulen, n., vogel, j., walls, r. l., wilson, e. o., & woolley, j. b. (2012). mapping the biosphere: exploring species to understand the origin, organization and sustainability of biodiversity. systematics and biodiversity, 10:1, 1-20. https:// doi.org/10.1080/14772000.2012.665095 woolrich-piña, g.a., ponce-campos, p., loc-barragán, j., ramírez-silva, j.p., mata-silva, v., johnson, j.d., garcía-padilla e, y wilson, l.d. (2016). the herpetofauna of nayarit, mexico: composition, distribution, and conservation. mesoamerica herpetology, 3, 376448. young, h. s., mccauley, d. j., galetti, m., & dirzo, r. (2016). patterns, causes, and consequences of anthropocene defaunation. annual review of ecology, evolution, and systematics, 47(1), 333-358. https://doi. org/10.1146/annurev-ecolsys-112414-054142 zhang, y. b., wang, y. z., phillips, n., ma, k. p., li, j. s., & wang, w. (2017). integrated maps of biodiversity in the qinling mountains of china for expanding protected areas. biological conservation, 210, 64-71. https://doi.org/10.1016/j.biocon.2016.04.022 http://www.reptile-database.org https://doi.org/10.69905/xbqkar19 https://doi.org/10.1080/21513732.2013.821168 https://doi.org/10.1073/pnas.1221370110 https://doi.org/10.1080/14772000.2012.665095 https://doi.org/10.1080/14772000.2012.665095 https://doi.org/10.1146/annurev-ecolsys-112414-054142 https://doi.org/10.1146/annurev-ecolsys-112414-054142 https://doi.org/10.1016/j.biocon.2016.04.022 biodiversity informatics, 15, 2020, pp. 11-54 11 can ecological interactions be inferred from spatial data? christopher r. stephens1,2,*, constantino gonzález-salazar1,3, maría del carmen villalobos-segura4 and pablo a. marquet1,5 1c3—centro de ciencias de la complejidad, universidad nacional autónoma de méxico, mexico city, mexico; 2instituto de ciencias nucleares, universidad nacional autónoma de méxico, mexico city, mexico;* 3departamento de ciencias ambientales, cbs universidad autónoma metropolitana, unidad lerma; estado de méxico, mexico; 4laboratorio ecología de enfermedades y una salud, facultad de medicina veterinaria y zootecnia, universidad nacional autónoma de méxico, mexico city, mexico; 5departamento de ecología, facultad de ciencias biológicas, pontificia universidad católica de chile, santiago, chile; instituto de ecología y biodiversidad (ieb), santiago, chile; and the santa fe institute, 1399 hyde park road, santa fe, nm 8731, usa abstract. the characterisation and quantification of ecological interactions, and the construction of species distributions and their associated ecological niches, is of fundamental theoretical and practical importance. in this paper we give an overview of a bayesian inference framework, developed over the last 10 years, which, using spatial data, offers a general formalism within which ecological interactions may be characterised and quantified. interactions are identified through deviations of the spatial distribution of co-occurrences of spatial variables relative to a benchmark for the non-interacting system and based on a statistical ensemble of spatial cells. the formalism allows for the integration of both biotic and abiotic factors of arbitrary resolution. we concentrate on the conceptual and mathematical underpinnings of the formalism, showing how, using the naive bayes approximation, it can be used to not only compare and contrast the relative contribution from each variable, but also to construct species distributions and niches based on arbitrary variable type. we show how the formalism can be used to quantify confounding and therefore help disentangle the complex causal chains that are present in ecosystems. we also show species distributions and their associated niches can be used to infer standard “micro” ecological interactions, such as predation and parasitism. we present several representative use cases that validate our framework, both in terms of being consistent with present knowledge of a set of known interactions, as well as making and validating predictions about new, previously unknown interactions in the case of zoonoses. keywords— ecology, naive bayes, spatial data mining, inference, interaction, biotic interactions, distribution modelling. introduction darwin’s entangled bank analogy is an adequate pictorial representation of the complexity of interactions that occur in ecological systems. their inference and characterization have been a recurrent theme and a vexing problem in ecology, and one where theory has usually been ahead of empiricism. the seminal work of alfred lotka and vito volterra provided a theory for competition and predation that has been experimentally tested (gause, 1934), refined (arditi and ginzburg, 1989; gilpin and ayala, 1973; holling, 1959) and expanded in many directions, including multi-species communities (case, 1990; gilpin, 1975; wilson et al., 2003). empirical analyses of species interactions have progressed at a slower pace. earlier work tried to estimate interaction strengths using simple measures of resource overlap as a proxy for competition coefficients (macarthur and levins, 1967), or estimate them using simple regression of the abundance of pairs of species across space or time (crowell and pimm, 1976; schoener, 1974). these methods, however, have been largely abandoned as they make strong assumptions (abramsky et al., 1986; dayton, 1973; rosenzweig et al., 1985). *stephens@nucleares.unam.mx mailto:stephens@nucleares.unam.mx christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 12 another method, inferring interactions from species co-occurrence data, which is the focus of this contribution, has survived the test of history and provides, as we aim to show, a valid alternative. the question of whether or not, or to what degree, ecological interactions can be identified and characterised using spatial data and, in particular, co-occurrence data, has had a controversial history (connor and simberloff, 1979; connor et al., 2013; freilich et al. 2018; morales-castilla, 2015; pollock et al., 2014; royan et al., 2016). there have been multiple perspectives and opinions with respect to the hypothesis and multiple methodologies used to investigate it (araújo et al., 2011; araújo and rozenfeld, 2014; borthagaray et al., 2014; cazelles et al., 2016; clark et al. 2017; gonzález-salazar et al., 2013; mohd et al. 2017; pollock et al., 2014; stephens et al., 2009)1. however, after more than 40 years there is still no consensus. over the last ten years a methodology has been developed, both related to and distinct from others, that has been used to identify, characterise and quantify ecological interactions (gonzález-salazar et al., 2013; stephens et al., 2009; stephens et al., 2019), both in terms of being consistent with known interactions as well as predicting previously unknown ones (berzunza-cruz et al., 2015; rengifo-correa et al., 2017; stephens et al., 2016). in this paper we will present an overview of the conceptual and theoretical underpinnings of this methodology and illustrate its utility using several representative use cases. that the framework has not been more widely seen or adopted in the ecology community is perhaps linked to the belief that point collection data, in particular, is not capable of identifying, characterising and quantifying ecological interactions (morales-castilla, 2015). however, it has been widely accepted that such data are sufficient to characterise the spatial distributions and corresponding niches of taxa when the niche variables are restricted to abiotic variables, linked to the fundamental niche (peterson et al., 2011). it has not been generally accepted, however, that such data can be used to model the effects of biotic factors as niche variables (peterson et al., 2018). from a modelling point of view this is, of course, somewhat jarring—that it is fine to represent a class variable (the species you want to model) using a certain data type, but not to represent the predictors 1some, such as pollock et al., (2014) and clark et al., (2017), use an approach whereby biotic factors are modeled jointly using abiotic factors as niche variables rather than as predictors themselves as in our approach. (the species that are potential niche variables) with that data type. specially so when we know that no species exists in isolation of others, and that the occurrence of a species in a given place is the result of the interaction between physiological tolerances, interactions with other species, historical effects and dispersal limitation (soberón and peterson, 2005). thus, it is questionable that niche models are really able to obtain a representation of the fundamental niche of a species without having a way of assessing the relative importance of biotic and abiotic factors in accounting for the presence of species across sites (soberón and nakamura, 2009). historically, species co-occurrence analysis has been central to community ecology theory (diamond, 1975; ovaskainen et al., 2010), where it was used to test whether a set of species co-occur more or less than would be expected at “random”, where the question of what is random has also had a controversial history (colwell and winkler 1984; gotelli, 2000). thus, if co-occurrence patterns over the whole set deviate from the random benchmark, it has been interpreted as evidence that a structural aspect of the community is driven by biotic interactions (brown et al., 2002; diamond, 1975). although this conclusion has been challenged (connor and simberloff, 1979; gotelli and mccabe, 2002), a generally accepted idea is that signals of species interactions can be inferred from survey data at local scales (gotelli et al., 2010). however, currently, an emerging issue in ecology and biogeography is to understand the interplay between the geographic distribution of species and their interactions at macro-scale levels (aragón and sánchez-fernández., 2013; gotelli et al., 2010). a commonly accepted idea is that climatic variables (grinellian niche) are the main determinant of the geographical distribution of species, whereas biotic variables (eltonian niches) operate at local scales, and their influence at large scales can be disregarded (eltonian noise hypothesis) (soberón and nakamura, 2009). consequently, biotic interactions are often neglected in spatial modelling. despite growing evidence that biotic interactions may determine the distribution of species (alvarez-martínez et al., 2015; godsoe and harmon, 2012; gonzález-salazar et al., 2013; heikkinen et al., 2007), the debate remains open as to whether they should be considered in ecological niche modelling, and, if so, how should they be quantified? of course, including biotic variables in spatial modelling opens up several important theoretical and methodological christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 13 issues. for instance, which biotic variables should be included? in general, the number of potential interactions for a single species will be much more than the known ones. therefore, we need a framework that allows one to include different types of data (e.g., collection points, environmental layers) in order to infer, compare and contrast potential interactions. we believe that an important barrier to making further progress is that of developing and agreeing on a deeper and more quantitative understanding of what an “interaction” is, at least in the context of patterns of co-occurrence, and how interactions can be manifested at different spatio-temporal scales. in our methodology, an interaction is defined by quantifying the degree of co-occurrence of variables—biotic or abiotic—relative to that expected in the absence of the interaction. in these simple terms the underlying modus operandi is no different than the original motivations of diamond (1975), where it was hoped that co-occurrence data could be used to reflect inter-specific competition. however, the logic is quite universal across all areas of science—physics, chemistry, linguistics, genetics, epidemiology—interactions always affect the positions of the objects that interact. the chief difference between the different disciplines is what objects are interacting and how should co-occurrences be defined so as to characterise the interactions? in the case of ecology, the use of point collection data for determining co-occurrence, and the interpretation of the associated analysis to infer biotic interactions, has been controversial and, especially recently, has generated many papers—see, for instance, (wisz et al., 2013) for a recent discussion. although our characterisation of interaction is definitional, it is important to determine to what degree such a characterisation captures the intuition associated with the standard classification of ecological interactions—such as predation, mutualism, commensalism etc. the latter are associated with the relative impact of the interaction on each participant—positive for the predator, negative for the prey for example—and are linked to a specific set of “labels” that mark each participant, such as predator = yes/no, prey = yes/no, with each interaction being associated with a particular label. these ecological interactions are “micro” interactions, in that they are most manifest at the level of individuals, such as predation events, where a bobcat kills and eats a rabbit for example. as all of these interactions are local and direct, they should all be amenable to an analysis in terms of a statistical ensemble of suitably defined co-occurrences. however, they offer a rather poor representation in terms of predicting the “macro” distributions that are an emergent property of the micro interactions, where by macro we mean how the relation between the spatio-temporal distributions of the bobcat and the rabbit as species are affected by these micro interactions. one reason why they offer a poor representation is because the macro distribution of a species depends on many labels, the majority of which are unknown. for instance, just how many species are hosts of a given zoonosis but are unknown as such? is the label of “prey” sufficient to characterise a predator-prey interaction at the macro level? what about the potential relevance of other labels, such as adult/young, male/female, large/small, strong/weak, fast/slow? all of which may be relevant to quantifying the relative success of the predator and therefore its spatio-temporal distribution. effectively, as in many other areas of science, we are pointed in the direction of attempting to deduce and relate interactions at a micro scale to those at a macro scale, in the knowledge that the macro scale emerges at the collective level from the combined effect of very many micro events. a vital link between the two scales is the concept of a niche. if a species is an important niche variable for another, then, by definition, it favours the presence of the species and thereby affects its spatial distribution (giannini et al., 2013). however, a niche dimension also captures an intuition as to why at the micro level it is a niche dimension. thus, we can accept that a predator-prey interaction is the underlying cause of the fact that presence of the prey species is a niche dimension for the predator and affects its distribution. however, it does not have to be so. individual predation events may, in fact, be due to completely random encounters between predator and prey. this, in turn, would then leave no imprint at the macro level. we could not then speak of the prey as being an important niche variable of the predator, in spite of the fact that there existed micro-level interactions between the two. we have natural selection to thank for the fact that this type of situation would be the exception rather than the rule. a predator that captures prey randomly would soon be out-competed by a predator that is better adapted. although micro-scale interactions are potentially easier to measure than macro-scale ones, there are just too many to measure. for n species we may christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 14 imagine that there are n(n−1) inter-specific interactions. however, it is much worse than that, as for any pair of taxa there are potentially as many interactions as they have relevant labels. it is because of this that macro-level interactions are potentially more amenable to analysis as, using a proxy such as point collection data, we can compute the deviations from a given null hypothesis for any pair of taxa from among a vast number. the relevant question then goes in reverse: instead of trying to characterise macro-level interactions as emerging from underlying micro interactions we can try to deduce properties of the micro interactions from the macro level data. once again, it is the concept of a niche that gives hope to this endeavour as it links the two levels. of course, a macro-level interaction may be a result of many different types of micro interaction, thus leading to confounding. this is no different than in many other areas of science. the spatial distribution of disease is, epidemiologically speaking, a result of many underlying micro interactions. however, many of these micro interactions may not be manifest and our understanding is first pointed to relate the spatial distribution of disease with the spatial distribution of risk factors (niche variables). subsequently, one may then attempt to give a micro explanation to the macro relations or vice versa. confounding potentially plays an important role in ecology, where it has been suggested that apparent biotic interactions may be confounded by abiotic factors (purse and golding, 2015), which can lead to the right prediction for the wrong reason (dayton, 1973). the bayesian formalism we have developed allows for a detailed investigation of this phenomenon. of course, there is a fundamental question as to whether a set of spatial data gives explicit information about an interaction versus information from which an interaction is to be inferred and potentially characterised. the vast majority of spatial data that is available has not been generated with the specific intention of analysing a particular interaction. rather, the data is used to infer the existence and nature of an interaction using a suitable mathematical framework for making statistical inferences. bayesian inference (berger, 1985) provides an appropriate framework for this task, where bayes’ theorem is used to update the probability of a hypothesis, such as the existence and nature of an interaction, as more evidence or information becomes available. this is particularly appropriate in our cases of interest where we can deduce more information about the interaction by including in more spatial information. although we will try to couch much of our discussion in general terms, our main concern is to apply these ideas to ecological interactions and, particularly, in the context of niche descriptions. of course, the use of co-occurrence data in ecology has a long history, with the particulars depending on whether we are talking about abiotic or biotic interactions. in the case of climatic data, species distribution modelling has used the co-occurrence between a point collection of a target species and the specific environmental conditions at that point, the latter being modelled as environmental layers at the pixel level (elith and leathwick, 2019; peterson et al., 2011). many different algorithms have been used to model the relation (qiao et al., 2015). this paper is based on a methodological framework (gonzález-salazar and stephens, 2012; gonzález-salazar et al., 2013; sánchez-cordero et al., 2008, sierra and stephens, 2012; stephens et al., 2009) that has been developed to determine, characterise and quantify ecological interactions of any type, abiotic or biotic, using data of arbitrary spatial resolution. it allows one to characterise the full ecological niche of a taxon, data permitting, while comparing and contrasting the contribution of each niche variable. the formalism has been applied successfully, chiefly in the area of zoonoses, where it has led to the prediction and confirmation of many previously unknown vector-host interactions in several emerging or re-emerging diseases (gonzález-salazar et al., 2017; rengifo-correa et al., 2017; stephens et al., 2016). in spite of its success in this important application area, it is not well known in a wider context and, importantly, its conceptual underpinnings and its general applicability to the general area of identifying and characterising ecological interactions have not been exploited. as a complement to these papers, we will here concentrate on its conceptual and mathematical basis and, in particular, how it can be used to infer causal chains and identify confounding factors in any given ecological setting and to predict micro ecological interactions. the format of the paper is as follows: in the second section, we discuss interactions in a wider context, as it is important to see that relating micro interactions to macro interactions permeates all of science. fundamentally, there is nothing different between doing so in physics versus in ecology, other christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 15 than ecology is much more complex in terms of the number of different types of interaction and the large array of factors that characterise them. what links them all is that interactions lead to a different spatial distribution of the objects that are interacting than would be the case in the absence of the interaction. then, we present the empirical definition of an interaction that has been at the root of our efforts—that an interaction can be identified using statistical ensembles of spatio-temporal data to show that in the presence of the interaction the spatial distribution of the members of the ensemble is different to that in the absence of the interaction. such a notion is universal in science, the most fundamental formulation being associated with newton’s laws of motion, where forces (interactions) can be identified from the spatio-temporal trajectory of an object relative to the null hypothesis that all objects not subject to an interaction are at rest or in inertial motion. in order to not distract the reader, we include discussion of co-occurrence and interaction in several text boxes that can be read separately and independently of the main text. next, we discuss how the notion of a co-occurrence can be used as a fundamental variable for distinguishing between interacting and non-interacting systems. we also show how co-occurrences can be compared and contrasted between variables that have radically different spatial resolutions and also different data types by making all variables, abiotic and biotic, binomial and bringing them all to the same spatial resolution. next, we discuss the bayesian modelling framework that is the heart of our methodology. we show how, based on the fact that we can bring all variables to the same type and resolution, we may compute the relative weight of any variable to the probability to find a given taxon, thereby determining its importance as a niche variable and simultaneously as a determinant of the spatial distribution of the species. moreover, we show how the formalism can be used to compare and contrast the relative degree of confounding of one variable and another, showing how this permits us to begin to disentangle the complex causal chains that exist in ecological systems. in particular, we will be able to show that, generally, biotic factors are confounders for abiotic factors, not vice versa. further, we discuss the relation between micro and macro interactions, showing how, and under what circumstances, micro interactions may be inferred from macro data. in addition, we present several use cases to give ample support to all the assertions previously made. finally, in the last section we draw some conclusions. what are interactions and should we be able to characterise them through spatio-temporal data? the most general notion of interaction across the sciences is simply that one thing affects another— mutually—with the main differences being what we mean by “thing” and what we mean by “affect”. in physics, for example, the presence of one electrically charged particle affects the presence of another, and vice versa. in ecology, the presence of a predator affects the presence of a prey, and vice versa. at a fundamental level all interactions are local, i.e., the interacting entities are located at the “same” place at the “same” time2 and therefore co-occur. interactions, if they can be characterised, are given names: electromagnetism, gravity, predation, parasitism etc. what they have in common is that the presence of one element in the system affects the state of the other and, again, vice versa. however, in each case the state variables that are affected by the interaction may be quite different. for example, in the case of predation the most important state change of the prey is living to dead, while a state change of the predator is, for example, a transition from hunger to satiation. we believe there is value in understanding in more detail how co-occurrence and interaction are related in sciences other than ecology. however, so as to not distract from the main text we have separated this discussion into boxes. an individual interaction may, in principle, be directly observable, as is often the case in ecology. this requires that the interaction is directly characterizable in terms of measurable variables. for instance, predation is an interaction that may be observed directly given the abrupt change in state variable of the prey: live → dead. often, however, the interaction must be inferred using data and reasoning, as the change in state variable that best characterises the interaction may not be directly observable or difficult to observe. for instance, although one may actually be present at the act of predation, in other circumstances it may well have to be inferred indirectly, by, say, an examination of the faeces of the predator. similarly, the feeding of an hematophagic insect on 2there are, of course, subtleties involved in defining what we mean by “same”, such as the question of simultaneity in relativity or what looks like action at a distance. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 16 a mammal may not be directly observed, but a blood analysis of the insect may reveal which species it was feeding from and allow for an inference of the corresponding interaction. in these cases, the interaction per se is known, but due to data considerations its characterisation must be inferred. however, there are many cases where the interaction is not previously known, where it is inferred first and then characterised and understood later. this is what happened with the fundamental interactions of physics. they were inferred from data first. in the absence of data on the state variable changes that characterise an interaction, a state variable that is almost inevitably affected as a consequence of an interaction, especially in mobile organisms, is the position in space and time of the interacting entities relative to what they would have been in the absence of the interaction. indeed, it may well be that, to a very good approximation, position is the only observable state variable that changes, or is readily observable due to the interaction. this absence of interaction then represents a “null hypothesis,” with respect to which the interaction may be benchmarked. as discussed in box 1, in many sciences other than ecology, co-occurrence, though oftentimes not defined explicitly as such, has been used successfully to characterise interactions. so why is it that in ecology, as discussed above, the use of position data and, in particular, point collection data to deduce the nature of ecological interactions has been so controversial and, apparently, not sufficiently successful to be accepted by the wider community? we will provide an answer to this question in the following sections. we believe that one problem with the notion of interaction in ecology is that there has been no adequate characterisation of how standardly accepted ecological interactions, such as predation, mutualism and parasitism, can be applied at the collective level. in other words, how do we characterise the “interaction” between two species, predator-prey, from knowing that there is an interaction at the level of each individual—a specific predation event? first, we need a notion of interaction that is applicable at all observation scales. given the incontrovertible evidence from myriad disciplines that interactions lead to changes of state in the objects that are interacting and, in particular, the state variables associated with the objects’ positions, which are distinct relative to the case—null hypothesis—where the interaction is absent, we will define an interaction to be present if the spatial distribution of the objects of study is different to this null hypothesis. this represents a purely empirical characterisation, dependent on the null hypothesis chosen, but which makes no a priobox 1: interactions and co-occurrence outside ecology seen from the perspective of other sciences, the answer to the question as to whether interactions are characterizable through spatio-temporal data, is an unequivocal yes; the reason being that we have been doing it successfully in many scientific disciplines for centuries. in physics, where the notion of interaction has been most precisely quantified, all the principal fundamental interactions—gravity, electromagnetism, strong and weak nuclear forces— have been identified and characterised by observations of the relative positions, as a function of space and time, of objects, such as planets, electric charges, nucleons etc. the enormous success of this endeavour has been due, in large part, to the fact that each interaction is characterizable in terms of a very small number of parameters. among these are “labels” for the objects, such as mass and electric charge, as well as universal constants which are measures of the strength of the interactions. the characterisation of interactions through the observation of the positions in space and time of the interacting objects is not restricted to the fundamental interactions. in atomic and molecular physics and chemistry, for example, effective interactions that emerge from the underlying fundamental interactions can also be characterised by the positions in space and time of different types of object—atom, molecule, macromolecule, planets etc. in comparison with the fundamental interactions, which are all direct and local, all interactions in physics and chemistry are, by definition, indirect. thus, chemical interactions, such as hydrogen bonding, covalent bonding, van der waals forces etc. are all emergent, indirect interactions. however, they are generally considered to be “direct”. firstly, because the first principles derivation of this indirect interaction from the underlying fundamental direct interactions is too complicated to carry out and, secondly and importantly, a better understanding of how the interaction was mediated would not necessarily help us to better understand atomic physics, molecular physics or chemistry. thus, the question of to what degree it is convenient to characterise an interaction as indirect versus direct is to a large extent christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 17 a question of convenience. for instance, is it important to understand the nature of any intermediation, or are the intermediating states readily observable? an important requisite for identifying the presence of an interaction using observations of the positions of the involved objects is that one must have, a priori, a notion of what those positions should be in the case of a non-interacting system. in the case of physics this is enshrined in newton’s first law—that an object will remain at rest or in uniform motion in a straight line unless acted upon by an external force. thus, deviations from this null hypothesis serve as a definition of the presence of a force, i.e., an interaction. with this null hypothesis goes the idea that in the non-interacting state the degree of co-occurrence of two objects should be different to the case when they interact. thus, the fact that electrons in a given atom co-occur in space and time with the nucleus of that atom, relative to the null hypothesis that they are independent of the nucleus, is an indication of the existence of an interaction—the electromagnetic attraction between negatively charged electrons and positively charged nucleus. the fact that the planets in the solar system co-occur in space and time with the sun, relative to the null hypothesis that they are independent of the sun, is also an indication of the existence of an interaction—the gravitational attraction between sun and planets. note that any comparison between the observations of a system and the null hypothesis must be done at the statistical level, where an appropriate statistical ensemble of observations must be formed. one single observation is not sufficient. for example, tycho brahe’s observations of the regularities in the dynamics of the planets led to kepler’s phenomenological laws. kepler could not have deduced those laws from just one entry in brahe’s notebooks. an ensemble of entries as a function of time was required. later, newton, using the null hypothesis of galileo that bodies not subject to an interaction (force) stay at rest or in uniform motion, deduced that planets are subject to an interaction as they do not follow that null hypothesis. newton’s laws of motion, along with kepler’s observations, allowed that interaction to be characterised, deducing that it depends on the masses of the interacting bodies and is weaker (1/r2) as a function of their separation. this is probably the clearest example of the logic of identifying and deducing the nature of interactions through observations of object positions. the statistical ensemble of observations in this case was the set of positions on the celestial sphere across time of the interacting objects. of course, understanding interactions through examination of the relative positions of objects is not restricted to physics and chemistry. as discussed in stephens et al., (2017a), in standard population genetics, where genes are viewed as beads on a string, the concept of interaction is associated with the notion of epistasis (phillips, 2008), where, in this setting, the degree to which two genes are linked, i.e., they co-occur, can be used as a measure of such epistasis. in other words, we measure interaction by to what degree two genes actually co-occur relative to their expected distribution if they were independent—the no-interaction null hypothesis. these genetic interactions may also have differing degrees of directedness. for example, it may occur that two genes are linked, where the linkage is not direct but through the intermediation of a third gene with which the two are directly linked. similarly, in text mining, syntactic and semantic interactions can be deduced from the co-occurrence of textual elements, such as words or phrases or other linguistic objects. it has also long been used in epidemiology. indeed, perhaps the founding event of modern epidemiology, the analysis of john snow of the broad street cholera outbreak of 1854 (snow, 1855) was based on a co-occurrence analysis, where the positions of disease cases and potential disease sources were mapped and an interaction—that the events were clustered around a certain water pump as opposed to being randomly distributed—was identified. although the concept of interaction, especially in physics, is naturally tied to an ensemble of observations in space and time this is not a prerequisite. in the absence of temporal data, we can and must infer interactions using only a spatial ensemble. in the case of genetics or language, for example, the ensemble is a specification of the positions of genetic objects—genes, exons, nucleotides etc.—or syntactic objects—nouns, articles, verbs etc.—in an ensemble of such objects, such as genomes or texts. the logic, however, is identical to that of brahe-kepler-newton: that relative to a suitable null hypothesis, the spatial distribution of these objects is significantly different. thus, articles precede nouns in english and not vice versa and an analysis of texts will identify this grammatical “interaction” relative to the null hypothesis that articles and nouns are randomly distributed. thus, in all cases, the question of whether a given spatio-temporal distribution of objects is distinct to our null hypothesis is a question of statistical inference. what we may infer about an interaction is then very much related to the precise nature of the ensemble we use. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 18 ri reference as to the nature of the interaction or its properties. we will usually think of the interaction as binary, in that it relates to two types of object. however, as we will see, we may readily extend the notion to multiple types of object and thus capture the idea of the interaction between an object and its “niche” or environment. we might believe that we can observe an interaction using only a single observation. as discussed in box 1, however, this is not true. the comparison between the observations of a system and the null hypothesis must be done at the statistical level, where an appropriate statistical ensemble of observations must be formed. for example, in the case of, say, predation, the actual act, observationally, requires an ensemble in time. before the actual predation, both the predator and prey would be described by position coordinates changing as a function of time. however, after the act of predation, the position coordinates of only the predator would change. thus, a characterisation of the event requires a history (an ensemble in time)—a before and after. as emphasised, this definition of interaction is scale independent, in that we can apply it to objects at multiple resolutions. in the absence of temporal data however, we must infer interactions using only a spatial ensemble. this situation also frequently occurs in other sciences, as discussed in box 1. in all cases, however, the question of whether a given spatio-temporal distribution of objects is distinct to our null hypothesis is a question of statistical inference. what we may infer about an interaction is then very much related to the precise nature of the ensemble we use. classifying interactions given our empirical definition of interaction, all we may deduce is that a set of spatiotemporal data is consistent with the presence or absence of an interaction; but this tells us nothing about the nature or properties of that interaction and therefore does not necessarily help us to understand it. to characterise the interaction we must seek parameters, or state variables, on which the interaction depends and determine if there are changes in the interaction as those state variables or labels change. unlike physics, as discussed in box 2, in ecology, there are many labels that are relevant for potentially characterising a given empirical interaction, box 2: classifying interactions an important element in understanding interactions is to characterise the properties of an object that give rise to an interaction in the first place. each such property can be associated with a “label”. for instance, in physics each fundamental interaction—gravity, electromagnetism etc.—is characterised by one and only one label—mass, electric charge etc. however, the nature of those labels goes a long way towards allowing us to characterise and understand the phenomenology of the interaction. thus, although gravity is much weaker than electromagnetism, it can manifest itself at large scales because mass—the gravitational “charge”—is always positive, whereas electric charge is positive or negative and macroscopic matter is neutral. a consequence of this is that the fundamental interactions manifest themselves at very different spatial scales, a fact which lends itself to an enormous simplification when trying to disentangle their relative effects. however, the effect of a fundamental interaction, such as electromagnetism, can become much more subtle and complicated at the collective level. for example, two atoms can repel at a small scale while attracting at a larger, molecular, scale. as the complexity of the interacting objects increases so does the potential number of labels that characterise the objects involved in the interaction. so, in molecular physics we must not only specify the atomic components but must understand the three-dimensional structure of the atoms and molecules in order to understand their interactions. in the end though, physics and chemistry are relatively simple, in that, at any given scale, there is usually one dominant interaction and correspondingly, one, or a few, most relevant labels. however, which labels are relevant can change radically from one observational scale to another. an important property of an interaction with respect to these labels is its degree of universality. thus, the fundamental forces are fundamental because they are completely universal, depending only on one label, in all places and at all times. they are not contextual. the gravitational force between the earth and the sun depends only on their masses and is independent of any other parameter. we can deduce this fact by observing properties that are consequences of the interaction—position in space and time for instance—and noting that predictions based on our characterisation of this interaction are independent of other labels. thus, the force of gravity is independent of the state of the gravitationally interacting body—so that labels, such as gas giant, rock, hot, cold etc. are unimportant christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 19 such as prey/predator, male/female, parent/offspring, carnivore/herbivore, mammal/reptile, old/young, fast/slow, etc. as well as the taxonomic names of the involved species, all of which may affect the spatio-temporal distribution of the organisms. indeed, the deviation of the spatio-temporal distribution of objects from a null hypothesis, that is at the heart of our empirical characterisation of an interaction, could be the result of potentially many distinct interaction types. furthermore, in ecology, we do not even have a complete, accepted set of labels to use. the existence of a large number of relevant labels for a given organism, that encompass the set of possible interactions with other organisms, makes the full empirical characterisation of all its interactions by direct observation completely impossible. an important property of an interaction with respect to these labels is its degree of universality. thus, the fundamental forces in physics are fundamental because they are completely universal, depending only on one label, in all places and at all times. they are not contextual. in ecology, however, the degree of universality is much less. for example, for a given predator there will be a label “prey = yes/no” associated with prey species of the predator. this label characterises the interaction and would be consistent with the classification of predation as having a positive effect on the predator and a negative one on the prey. however, the label “prey = yes/no” is only one of many that may be relevant for characterising the interaction. for example, for one prey it may be that 70% of attempted predation events are successful, while for another it is only 10%. also, at the collective level, such as at the species level, two different prey species may form substantially different parts of the diet of the predator. thus, at the scale of an individual predator and an individual prey, the interaction may well be “repulsive”, when viewed in terms of the trajectories of the individuals, reflecting the fact that a prey may try to avoid the predator, while at the species level, the interaction can be attractive, meaning that the predators at the collective level are attracted to where the prey species are located. although the taxonomy of the fundamental interactions in physics is clear and well established, where at the most gross, phenomenological level, we may speak of an interaction in terms of whether it is positive or negative (attraction/repulsion), the “charges” (labels) on which it depends, its strength and its relative importance, in ecology interactions have been classified in a somewhat different way (lidicker, 1979; wisz et al., 2013), based on earlier work in the social sciences (haskell, 1949), where the characterisation of the interaction is principally based on the impact it has on the interacting taxa, where the impact is considered in terms of whether it has a “positive” (benefit) or “negative” (cost) effect (araújo and rozenfeld, 2014; belmaker et al., 2015). this cost/ benefit, in turn, must be evaluated in terms of some measurable function, such as reproductive success. however, this classification in no way exhausts the set of labels or other parameters that are potentially relevant for classifying the interaction. there has also been work on trying to quantify the notion of interaction strength (see for example paine, 1992; wootton and emmerson, 2005), where the measures are mainly linked to experimental procedures, such as removing a species from an environment, or on mathematical models, but not as measured directly from spatial distributions. niches and interactions: from micro to macro we have defined interactions as being identifiable from deviations in the spatiotemporal distributions of objects from a suitable null hypothesis. we have also emphasised that a statistical ensemble of observations is necessary. clearly, what we may deduce about an interaction depends on the nature of those observations and what data represents them. in for this interaction. hence, the gravitational attraction of one single 10 kg mass is the same as that of ten 1 kg masses combined, while the electrostatic force exerted by one charge of 10 coulomb is the same as that of ten charges together of 1 coulomb. the taxonomy of the fundamental interactions in physics is clear and well established. at higher levels of organisation there are also established taxonomic classifications of interactions, e.g., covalent versus ionic bonding, and labels, such as from the periodic table, that allow us to characterise and quantify the interactions. at the most gross, phenomenological level, we may speak of an interaction in terms of whether it is positive or negative (attraction/repulsion), the “charges” (labels) on which it depends, its strength and its relative importance. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 20 particular, it depends on the scale or spatio-temporal resolution of those observations. we may consider two distinct scales—the “micro” and the “macro”, though we use these as relative not absolute terms. we know, particularly in physics, as discussed in box 3, that the nature of interactions at one scale can radically change when passing to a different scale. in ecology, in the case of a predator-prey interaction, for example, the “micro” level would naturally correspond to observations of two individuals—one predator and one prey. following their trajectories in space and time would allow us to determine that there is an interaction present. moreover, labels such as “predator” and “prey” would allow us to determine its biological plausibility and characterise it. however, this ensemble of observations tells us nothing about the nature of the interaction at the collective level, at another spatial resolution—the “macro” level. in general, there is no fundamental reason why an interaction at one scale should manifest itself in an obvious and analogous way at another scale. in this sense the interaction is context (environment) dependent. similarly, as mentioned, predator and prey may “repel” at the micro-level, in that the prey tries to avoid the predator, but may be attracted at the macro level, in that the spatio-temporal distribution of the predator species is attracted to that of the prey species. in this case, the attraction between predator and box 3: from the micro to the macro physics is the discipline where the quantitative and qualitative relations between micro and macro variables is best understood, with micro and macro being relative terms. for instance, we may consider the intra-atomic scale as being micro and the inter-atomic scale as macro. thus, as an example, at the intra-atomic scale, there is the strong interaction between electrons and nucleus in a helium atom, while at the inter-atomic scale there are no significant interactions between the helium atoms themselves. a relevant label for the interaction is that of electric charge. at the micro level the electrons have a label corresponding to one unit of negative charge while the nucleus has a label corresponding to two units of positive charge. on the other hand, the helium atoms have a label corresponding to zero charge. of course, we fully understand how the zero electromagnetic interaction at the atomic “macro” level emerges from the underlying strong interaction at the “micro” level. similarly, electrons repel as free particles, but they can “attract” in the case of a covalent bond. thus, the nature of interactions at one scale can radically change when passing to a different scale, with the interactions at a more macro scale being an emergent phenomenon relative to the interactions present at the micro level. additionally, in physics instead of talking about the interaction between two objects we may speak of the interaction between an object and its environment when, for instance, the environment consists of an ensemble of objects, such as atoms, where in a solid say we may consider the interaction between an atom and its environment as represented by the ensemble of other atoms in the system. a particular atom or other structure in a solid has its “niche” in the same way as a species has its niche. the environment in both cases is the net, emergent effect of a large set of individual niche variables. thus, in principle, we may consider interactions along a spectrum, from between two individual objects of definite types, to between an object and any conglomeration of objects that represent its niche/environment. in physics, the fundamental interactions are direct. for example, the interaction between an electron and a proton in a hydrogen atom can be thought of as being direct, being describable directly in terms of the fundamental electrostatic attraction between the positively charged proton and the negatively charged electron. two hydrogen atoms, though, may form a hydrogen molecule where, unlike the direct electrostatic interaction, the interaction between these atoms is “indirect” and is a consequence of the presence of other intermediating elements—the electrons. in a potential abusive use of ecological terminology, we could say that the repulsive “competitive” interaction between individuals of the proton species were turned into an attractive interaction by the “facilitation” of individuals of the electron species. an effective, phenomenological model of a chemical bond between atoms as a “direct” interaction may well be much more useful than trying to simultaneously model the multiple, underlying, more fundamental, direct interactions between the atomic constituents. however, if we examine the molecule exhibiting the chemical bond at higher energies then these more fundamental interactions will become apparent. the important point to make is that how an interaction in a system is characterised in physics and chemistry, and the degree to which we describe it as direct or indirect, is very much a function of the scale at which we make observations of that system. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 21 prey at the species level is an emergent phenomenon relative to the discrete predation events observed at the individual level. in ecology, it is the concept of a niche that provides a powerful and intuitive framework for understanding emergence. that a biotic variable, xα, is considered and observed to be an important niche component of a species, c, will have the consequence that the presence of xα favours the presence of c. the conditional probability p(c|xα) is a suitable mathematical representation of this relationship, where we will specify below what ensemble of observations is suitable for calculating p(c|xα) and, in particular, the dependence on the spatio-temporal scale of the observations. however, the intuition is that micro-level interactions between c and xα can manifest themselves at a larger scale by xα being a niche variable for c. this concept of niche is relevant for understanding the relations between objects in any area of study. built into the concept of both the eltonian (elton, 1927) and hutchinsonian (hutchinson, 1957) niches is the notion that biotic interactions between a target species and other species in its niche affect the spatial distribution of the target species and vice versa. thus, to what degree a species persists in a given geographical area is affected by its interactions with other species. here, at the niche level, in distinction to the notion of interaction between two discrete objects, we consider the interaction between an object and its environment. thus, the concept of a niche links in a profound way the notion of micro interactions to macro interactions, as defined by our concept of deviations of the spatial distribution of a taxon from a null hypothesis, with the micro interactions being the supposed drivers of the spatial distribution as explained by its niche. direct versus indirect interactions the concept of direct versus indirect interaction occurs in all the examples we have mentioned. as stated, labelling an interaction as direct or indirect is to some degree a question of convenience and a question of observational scale. in ecology, determining whether an interaction is direct or indirect is difficult. this can be perfectly illustrated in the context of a simple food chain, such as: carnivore ← herbivore ← plant ← sun. one would be tempted to argue that the interaction between carnivore and herbivore was more direct than between carnivore and plant. however, a perfectly acceptable predictive model for the carnivore distribution might be built using plants as niche variables. moreover, the principle effect of climate on the carnivore distribution will be an indirect interaction, intermediated by the plant and herbivore distributions, rather than a direct one (see for example (rebolledo et al., 2019). we must then determine how from data we may disentangle these interactions and characterise their degree of indirectedness. as a prelude to a later discussion (“a bayesian framework for causal inference”), we may intuit the degree of indirectedness of an interaction between two taxa, c and xα, by trying to determine if there exists one of more other variables, xβ, that are more directly linked to c or xα and therefore act as confounders. thus, if c is a carnivore, with an important prey species, xα, and xβ is the principal plant food source of xα then the interaction between c and xβ will be intermediated by xα. co-occurrence as a measure of interaction having defined an interaction as being associated with a spatial distribution of two or more sets of objects that differs from a null hypothesis where the interaction is absent, we must define observable, measurable parameters that allow for a comparison between an observed distribution and the corresponding null hypothesis. an extremely useful measure is that of a co-occurrence. in ecology, as elsewhere, co-occurrences are a necessary condition for an interaction. for a predation event to occur, the predator and the prey must be in the same place at the same time. similarly, for pollination, or any other type of ecological micro interaction. of course, this does not deny the possibility of non-local, action-at-a-distance type interactions that are intermediated by other variables such as climate teleconnections or nutrient transports across continents or global scale phenomena such as climate change and commercial trade. (bradley et al., 2012; bristow et al., 2010; wang et al. 2000). for instance, two species may interact through an abiotic intermediary, where the interaction between the species and the intermediary is direct and local but the effective interaction between the species is non-local and indirect. however, the fact remains that any direct interaction is local. defining co-occurrence co-occurrence is a notion about discrete events happening in space and time, either at the same place, christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 22 or same time, or both. here we will consider only a two-dimensional space. to define same place in space and time we first specify some partition of a space, a, and time interval, t, into cells. a natural, though not obligatory, partition of a is that of an array of squares of a fixed linear dimension, while that of t would be into fixed intervals of time. a cell i thus defines an area, ∆a and a time interval ∆t. to the question of what is co-occurring we may consider variables (x1(x1,t1),x2(x2,t2),...,xm(xm,tm)), which, in principle, may be discrete or continuous and where xi = (xi,yi) are the two-dimensional coordinates for geo-referencing. a co-occurrence of any subset, v = {x•(x•,t•),xβ (xβ,tβ),...}, of these variables can then be defined by an indicator function as i = 1 ⇐⇒ all (xj,tj) ∈ v are observed in the cell i and is zero otherwise. thus, i is a boolean function. for example, for two variables, c(x,t) and xα(x´,t´), a co-occurrence in a cell i is such that i = 1 ⇐⇒ c and xα are observed in the cell i, with x,x ́∈ i, and is zero otherwise. for a single variable, xα(x,t), we can also use the indicator function, where now i = 1 ⇐⇒ all (x,t) ∈ v are observed in a cell i. the set, s, of n cells represents a statistical ensemble. we may then count events on this ensemble. for any single observable xα, such as the presence of a species, we may count the occurrences of xα on a × t as 𝑁𝑁𝑥𝑥∝ = ∑ ∑i(𝑋𝑋∝(𝑥𝑥, 𝑡𝑡)) 𝑡𝑡∈t𝒙𝒙∈a (1) similarly, the number of co-occurrences of two variables xα and xβ is given by 𝑁𝑁𝑥𝑥∝𝑥𝑥𝛽𝛽 = 1 2 ∑ ∑ ∑ ∑ i(𝑋𝑋∝(𝑥𝑥, 𝑡𝑡), 𝑋𝑋𝛽𝛽(𝑥𝑥′, 𝑡𝑡′) 𝑡𝑡′∈t𝒙𝒙′∈a𝑡𝑡∈t𝒙𝒙∈a (2) in this way, we could count multiple variable pairs within the same spatio-temporal cell. we can also count by simply counting in each cell occurrences of a given type, independently of how many examples ____________ 1 this null hypothesis is one that leads to lower rates of type i errors. of the type there are in the cell. in this case, the number of co-occurrences is given by 𝑁𝑁𝑥𝑥∝𝑥𝑥𝛽𝛽 = ∑i(𝑋𝑋∝(𝑖𝑖), 𝑋𝑋𝛽𝛽(𝑖𝑖)) 𝑖𝑖 (3) and the number of occurrences nxα by 𝑁𝑁𝑥𝑥∝ = ∑i(𝑋𝑋∝(𝑖𝑖), ) 𝑖𝑖 (4) in the case of a purely spatial ensemble, s represents a set of n spatial cells and (3) and (4) are calculated over this set. note that i, as a boolean function, may represent any composition of variables. for instance, we may consider i(xα(x,t),xβ(x´,t´),xγ(x´´,t´´)) = 1 ⇐⇒ ((x,t) and (x´,t´) ∈ v) or ((x,t) and (x´´,t´´) ∈ v). for example, α could represent a species while β and γ represented two species in the same genus and we were counting co-occurrences of α with that genus. in fact, we may consider the co-occurrence nxαx, where x = (xβ1, xβ1,...,xβm) can represent, for example, any set of potential niche variables. x in this sense can be seen as one single, composite variable. with nxα and nxαxβ in hand we may also calculate any associated probability distribution over our ensemble, such as p(xα) = nxα/n, p(xαxβ) = nxαxβ/n or p(xα|xβ) = nxαxβ/nxβ. the null hypothesis so, how do we infer that a given observed value of nxαxβ or p(xα|xβ) indicates the presence of an interaction? to do so we need a no-interaction null hypothesis. the relative merits of different null hypotheses have been the subject of much study (gotelli, 2000). in our methodology, we take it to be that, in the absence of any interaction, we expect the distribution of xα to be governed by the probability distribution p(xα). this is equivalent to the null hypothesis of type sim2 in the classification of gotelli (2000)1 and corresponds in the framework of presence-absence matrices to keeping the number of observations fixed but randomising their location. thus, in an ensemble of size nxβ we would expect to see nxβp(xα) events of type xα. put another way, our christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 23 null hypothesis is that p(xαxβ) = p(xα)p(xβ), i.e., the events xα and xβ are independent. in the bayesian sense, as we will see below, the null hypothesis is associated with p(xα) being identified as a prior distribution, while p(xα|xβ) represents a posterior distribution after the addition of the information xβ. making the world binary in the above we have implicitly had in mind that a set of variables of special interest in ecology are the binomial variables that represent presence/absence or presence/no presence of a taxon. thus, we will count co-occurrences of presence or absence/no presence of different taxa. however, there are many variables that could be relevant labels for characterising and quantifying an interaction that are not intrinsically binomial. for instance, a phenotypic variable, such as size, may be continuous, or abundances, or abiotic variables such as temperature, where we could potentially speak of interactions, as defined by deviations in spatial distributions relative to a null hypothesis. thus, we may speak of the interaction between the presence of a species and a given climatic condition—temperature or precipitation—or, indeed, any abiotic variable, noting that by so doing we are emphasising the pragmatic, empirical nature of our characterisation of interaction rather than forcing it to accord with the standard taxonomy of ecological interactions. the problem with a continuous variable, xβ, relative to that of a binomial variable, xα, such as presence, is that the number of co-occurrences nxαxβ will be very small, its magnitude depending on the resolution of the measurement of the continuous variable. indeed, with arbitrary resolution, there will be no repetitions of the same value and nxαxβ = 0, 1. in this case, direct statistical inference by counting is impossible, although by assuming a particular functional form of the distribution of the continuous variable as a function of position, it may be possible to proceed. an alternative approach is to coarse grain any continuous variable into a set of discrete values and consider each discrete range as a new binomial dummy variable. thus, for example, we can divide temperature into 10 bins with intervals tmin + n(tmax − tmin)/10, n = [1,10]. temperature in a given range is representable by a variable xβn(x) = 0, 1, as in cell we may ask if there is a “presence” of a temperature ____________ 2 in this case we can truly say absence as we may identify those cells where the temperature definitely isn’t in a chosen range. in any given range. thus, the abiotic variable can now be treated as 10 presence/absence variables2. by doing this we avoid any model bias associated with an assumption about the relationship between independent and dependent variable, as would be the case in a regression analysis. there is, of course, a question of how many bins to choose and what should be their ranges? too few bins risks missing potentially relevant information about variation of the variable within the bin, while too many bins risks losing statistical significance by having too few data points in a bin. the number of bins should also be motivated by underlying biological or ecological factors. for instance, we may ask if the spatial distribution of a biota is sensitive to variations in average annual temperature of 0.1°c? if not, then there is no need to use such resolution. by making every variable binary we may represent any state of a given cell by a vector of binary variables x = (x1,x2,...,xm), where xi = 0, 1. an equivalent coding is to consider a single, composite variable of cardinality 2m. either can be used to represent any set of niche variables—abiotic and/or biotic. cell size and the problem of variables of different resolutions we have defined co-occurrences with respect to a spatial cell of a given size. besides the problem of the dependence of our results on this cell size3 we must also ask how variables of quite different resolutions can be compared given a fixed cell size? the natural resolution of an environmental raster is at the pixel level, where there may be hundreds of thousands or even millions of pixels associated with a geographic area of interest. however, if we are to proxy biotic data by, say, point collection data, then for a given species we may have only tens or hundreds or data points. if we choose a pixel level resolution, then the interactions associated with this variable will be calculated from a sample size which is enormously greater than that for the biotic variable. more importantly, in any predictive model the contribution of the low-resolution variable will be very small relative to that of the high resolution variable in terms of its coverage, i.e., how many pixels are affected by the biotic variable. we may then be led to conclude that the abiotic interaction was much more important than the biotic one. this is a pure 3 a problem (gehlke and biehl, 1934) in spatial data mining known as the “modifiable areal unit problem” (maup) (openshaw, 1983). christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 24 effect of sample size associated with the resolution of the data however, without reflecting any underlying real difference. it is not the same as comparing the effect of a geographically restricted prey species versus a geographically disperse one. in this case, the disperse prey species would be associated with more cells than the restricted one. however, this reflects an underlying reality. the fact that reflects a difference in the importance of their interactions for a predator. thus, an environmental raster is equivalent to having corresponding presence/absence variables at every pixel. there are two alternatives to bringing variables to the same scale: i) extrapolate variables at a lower resolution to a higher one, or; ii) coarse grain variables of a higher resolution to a lower one. in the former, we must make assumptions associated with how this interpolation is done. for instance, one may use the proposed distributions for mammals of hall (1981). a more convoluted way is to convert a discrete distribution into a raster by using a species distribution model that was built using abiotic rasters (araújo et al., 2014; atauchi et al., 2018). in option ii), however, there is no interpolation or assumption. the coarse graining is done using the dummy variables defined elsewhere (“making the world binary”). thus, at a given cell resolution, for any pixel-level raster we ask in that cell how many presence-absence dummy variables are represented. for instance, if in the cell there are pixels that lie in two of the ten temperature ranges then those two ranges are present in the cell and the other eight are absent. the more general effect of cell size can be appreciated if we consider the limits of very large or very small cells. there are two general considerations: one is the effect of cell size on the effective size of the statistical sample to be analysed, and the other relates to its effect on the number of co-occurrences. for the former, as we are carrying out a statistical analysis and a corresponding hypothesis testing, it is natural to take advantage of the samples at our disposal as much as possible. if we have n events, then the maximum size of the sample of cells is also n. however, if the cell size is such that multiple, assuming they are independent, events occur in a given cell then the effective sample size is reduced. for instance, for a random distribution of events, if the cell size is such that, on average, the number of events per cell is 4, then a reduction in the cell size by a factor of 2 will probably lead to cells where the expected number of events per cell is closer to one. in other words, all else being equal, the number of cells should naturally scale as the number of events. for the case of co-occurrences, for a finite set of events, if we go to the limit of very small cells, it is clear that eventually we will end up with zero co-occurrences. on the other hand, in the limit of very large cells we will end up with only one co-occurrence, as all the events will be in one cell. the choice of cell size has been investigated empirically (sierra and stephens, 2012), where it has been determined that although an optimal resolution exists that maximises the number of co-occurrences, our results are robust to the precise cell size. testing the null hypothesis we now have a means to quantify the difference between the spatial distribution of two taxa, c and xα, and the distribution of either one of them in the absence of the interaction using either i1(cxα) = (p(c|xα) − p(c)) or, equivalently, i2(cxα) = (p(cxα) − p(c)p(xα)), with i2 = p(xα)i1. as i1 and i2 represent deviations from the null hypothesis of no interaction, any non-zero value might be interpreted as evidence of an interaction. however, as this question is being answered with respect to a statistical ensemble, we must determine its degree of statistical validity. various diagnostics may in principle we used. however, here, given our emphasis on converting everything to binomial variables, we will use a simple binomial test based on our null hypothesis. specifically, we will use (5) in the case where the binomial distribution may be approximated by a normal distribution, then |ε(c|xα)| > 1.96 corresponds to the 95% confidence interval for consistency with the null hypothesis. when a normal approximation is inadequate then a more sophisticated approximation may be used, such as the wilson intervals (wilson, 1927). note that ε(c|xα) ≠ ε(xα|c), i.e., it is asymmetric in its arguments, thus modelling the fact that the effect of xα on c may not be the same as that of c on xα. if we had used i2 instead of i1 to quantify deviations from the null hypothesis then the corresponding diagnostic is christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 25 (6) which is a measure of the deviations from the null hypothesis that p(cxα) = p(c)p(xα). of course, when there are sufficiently many xα, then some ε(c|xα) will be statistically significant for any chosen p-value. this is nothing new and there are many methods that may be used to ameliorate this effect, such as applying a bonferroni correction, or an anova analysis. importantly, one should also have some intuitive notion of why it should be significant in the first place. note that the principal role of ε(c|xα) is to inform us that the spatial distributions of c and xα are such that they are not consistent with the null hypothesis that the distribution of c is independent of the distribution of xα and that, therefore, by our definition, there is an interaction between them. what else can it tell us? note that it has both an intensive and an extensive character. i1, or equivalently i2, can be used to begin to characterise the intrinsic strength of the interaction in that the larger is i1 the more the presence of species α is correlated with the presence of species c. on the other hand, by virtue of taking into account the sample size, nxα, it allows us to infer a greater importance for the interaction as the sample size increases, if the sample size is representative of the underlying relative geographic distribution of the taxon. as an example, consider the relation between a vector, c, and a potential host species, α. the greater is i1, the more likely it is that the vector occurs when the host species is present relative to a random benchmark. on the other hand, for a fixed i1, the greater the value of nxα, the higher the expected number of observed instances of c relative to the null hypothesis. so, ε(c|xα) captures two different facets of the interaction, the first associated with the strength of the interaction via i1 and the second with the coverage of the interaction via nxα. it can be used to compare and contrast multiple interactions. for example, we may have three distributions c, xα(x) and xβ(x), such that but that ε(c|xα) ≈ ε(c|xβ), due to the fact that . in this case, we would conclude that the interaction between c and α is stronger than that between c and β. note that (5) and (6) generalise to the case where xα represents a composite variable, equivalent to a set of variables. thus, 𝜀𝜀(𝐶𝐶|𝑿𝑿) = 𝑁𝑁𝒙𝒙(𝑃𝑃(𝐶𝐶|𝑿𝑿) − 𝑃𝑃(𝐶𝐶)) √𝑁𝑁𝒙𝒙𝑃𝑃(𝐶𝐶)(1 − 𝑃𝑃(𝐶𝐶))) (7) represents a measure of the interaction between a taxon c and its niche variables x. indeed, we may use equation (7) to quantify the niche. the greater is the difference i1 = (p(c|x) − p(c)), for a given configuration of the niche variables x, the more that configuration represents more favourable niche conditions for the presence of c. similarly, if i1 < 0, the more negative it is, the more the corresponding configuration of variables represents “anti-niche” conditions, i.e., conditions that are unfavourable for the presence of c. however, the problem with considering multiple niche variables within this simple diagnostic directly is that nx = 0, 1 when many niche variables are included. note that as a measure of non-random co-occurrence relative to the null hypothesis, ε(c|xα) is quite different to the checkerboard matrix elements (roberts, 1990) ccxα = (nc − ncxα)(nxα − ncxα), as the latter is independent of the size of the statistical ensemble chosen. additionally, in the case of the checkerboard, the null hypothesis is applied to the entire matrix and a single index—the checkerboard score—is evaluated relative to that null hypothesis. what can we infer from co-occurrences? the interpretation of co-occurrences is very much related to the characterisation of our cells and our objects. to fix intuition: in the case of a spacetime ensemble, a predation event could be characterised by the dynamics of an individual of a predator species, represented by a presence variable, xα(x,t) = 0, 1, and an individual of a prey species represented by a presence variable, xβ(x´,t´) = 0, 1, which represent the trajectories of predator and prey in space and time. in this case, predation would be associated with the fact that nxαxβ = 1, where xα(x,t) = xβ(x,t) = 1 and after the predation event xβ(x´,t´) = 0 for any t´ > t. in other words, the predator and prey have trajectories that intersect, such that after the co-occurrence the prey species individual disappears. the statistical ensemble of events here is aschristopher r. stephens et al. – can ecological interactions be inferred from spatial data? 26 sociated with the positions, x(t), of the two individuals. this ensemble, however, would tell us nothing about how typical this particular interaction was nor, indeed, whether or not it was just a chance encounter. in other words, it would tell us nothing beyond the interaction of those two individuals. at the population level, we must account for the fact that a prey species may be easy to catch but be relatively rare, or difficult to catch and widely distributed, or any combination of these and other characteristics. thus, an ensemble of co-occurrences of predatorprey would potentially yield much more information about their interaction beyond the individual level. from such ensembles we may deduce the likelihood of success of an attempted predation for instance, or the importance in the overall diet of the predator of a given prey species, all of which are relevant for characterising the interaction and for considering the prey species as a niche variable of the predator. there is also the possibility that an interaction is more, or less, manifest at the level of one spatial resolution/ensemble versus another. for instance, it may be that the predator-prey interaction is manifest at the level of the trajectories of the individuals, but that there is no correlation between these events. in other words, that the distributions of predator and prey are random, and that any predation event is the result of a chance encounter. this is a common occurrence in physics. for instance, a gas of helium atoms will exhibit a strong interaction between the positively charged nucleus and the two electrons that make up the atom but there will be no resultant interaction between the neutral atoms themselves. as well as depending on the ensemble chosen, the inference that may be drawn from any event, or set thereof, depends on the spatio-temporal resolution of the cells. if ∆a is 1 m2 and ∆t = 1 second then we might infer, knowing that one individual was a predator species and one was a prey species, that disappeared after the co-occurrence, that a predation event had occurred. the trajectories before the co-occurrence might also give extra information that supported the hypothesis. was the prey species trying to avoid the predator species for example? there are multiple null hypotheses that could be used for comparison associated with the expectation of the trajectories of the individuals in the absence of the interaction. for instance, that the degree of correlation between the trajectories is zero. we could also construct an ensemble of co-occurrences of the two species with the same spatial resolution and just count the number of cells in which predator and prey co-occurred where subsequently the prey disappeared. however, if the spatial resolution were 1 km2 and the temporal resolution were 1 year then the co-occurrence only tells us that the predator and prey were in the same 1 km2 area at some point in the last year. this, of course, is not sufficient to infer a particular predation event. how might the relation between predator and prey be inferred now? in this case, we must infer any relation from a different statistical ensemble. if we consider a spatial ensemble of n cells, we have stated that ε can be used to determine the existence of an empirically defined interaction between two taxa. what else can it tell us? as noted, it has both an intensive and an extensive character. first, i1 can be used to begin to characterise both the sign and the intrinsic strength of the interaction, in that the larger is i1 the more the presence of species c is predictive of the presence of species α. if i1 > 0 we will say that the interaction is attractive, in that the probability to find c and α co-occurring is higher than our null hypothesis, and repulsive in the contrary case. furthermore, the larger the magnitude of i1, the stronger the interaction. finally, considering nxα allows us to infer a greater importance for the interaction as nxα increases. these interpretations naturally depend on the fact that the underlying data of the ensemble is representative of the underlying relative geographic distribution of the species. as an example, consider the relation between a predator, c, and two potential prey species, α and β. the greater is i1, the more likely it is that the predator occurs when the prey species is present relative to the null hypothesis. thus, if i1(c|xα) = i1(c|xβ) we say that the two interactions have the same strength. this is a measure of the interaction between any given individual predator and any individual prey taken from the ensemble. however, if nxα > nxβ we will say that the interaction between c and α is more important than that between c and β. this is a measure of the interaction at the population level. so, the process by which we may infer an interaction is the following: i) construct a statistical ensemble of cells associated with a given spatial and/or temporal resolution; ii) compute co-occurrences of presence variables on that ensemble, where the presence variables may have an associchristopher r. stephens et al. – can ecological interactions be inferred from spatial data? 27 ated set of labels—taxonomic, phenotypic, behavioural etc.; iii) compute one or more statistics based on the distribution of co-occurrences; iv) compare that statistic with a suitable null hypothesis; v) if the statistic and the null hypothesis are statistically significantly different conclude that there is a macro interaction; vi) use the labels and other data to deduce the nature of the interaction and relate it to micro interactions. what spatial data? essentially, we have highlighted the predation example to emphasise that, as in physics, spatiotemporal data can be used to identify and characterise ecological interactions. the real question is: what spatio-temporal data is necessary or sufficient? must we have micro-data associated with each, individual event across space and time? how are interactions manifest at different resolutions? clearly, we are not in the position of being able to track in space and time the positions of representative sets of all taxa. there are just too many potential interactions to be identified and quantified individually. thus, in the absence of true, detailed observational data on biotic interactions at a macro-scale, as derived aggregated micro data, one must resort to proxy data. this is similar to the dilemma faced when contrasting epidemiological considerations against clinical or physiological variables, where the causal distance between a disease, say, and its symptoms, may be much less than that between the same disease and its risk factors. an important consideration then is what data will represent the spatio-temporal distributions of biota? in standard niche modelling, where the output variable is the spatial distribution of a taxon, the proxy of choice has been a database of point collection data. such data may be bespoke, in that it represents a dedicated, controlled study that tries to be as unbiased as possible, versus museum collection data, where the data, although ample and widespread, is potentially biased and unrepresentative. however, in spite of all its shortcomings, which have been amply discussed (hortal et al., 2008; soberón and peterson, 2004), such data has been used ubiquitously to calculate species distributions (peterson et al., 2011). moreover, as mentioned, other sets of potentially biased biotic data have then been used as dependent variables to produce models using abiotic independent variables that are then input as independent variables as potential proxies of biotic interactions in the calculation of the distribution of a species of interest (araújo et al., 2014; atauchi et al., 2018). the ultimate test has to be to validate point-collection data by the predictions of models that use it as input. there are two possibilities: i) use it to calculate species distributions, using abiotic and/or biotic data, and infer interactions as defined herein, and then determine to what extent those distributions and interactions are consistent with known ecological interactions and, furthermore, if they lead to more precise species distribution models and a better ecological understanding of the niche in such models; ii) use it to predict unknown ecological interactions and use experimental protocols incorporating field work and potentially laboratory work to validate those predictions. a bayesian framework for analysing ecological interactions our definition of an interaction is probabilistic, based on analysing the difference between p(c|x(t)) and a null hypothesis p(c) for a taxon c. this is equally applicable in the case that x represents just one niche variable, xα, versus many. in either case, with a statistically significant deviation between them we infer the existence of an interaction between c and x that we must then further understand. we take it as an axiom that the spatial distribution of a taxon is a result of all the interactions, both abiotic and biotic, that impact on that distribution and that there exists an underlying, potentially dynamic, probability distribution that predicts and explains the distribution of the species. of course, there are many reasons why a predicted distribution may not be a good representation of the real distribution. first, it may be that the underlying distribution is dynamic and approximating it by a static (“equilibrium”) distribution is inadequate. secondly, it may be that important variables have been omitted from the model, such as biotic variables. thirdly, it may be that the data representation of the variables is inadequate, because of data bias, and, finally, it may be that the mathematical relation between those variables is not being modelled correctly. thus, any model and hypothesis about interactions must be validated both through its predictions and how it increases our understanding. in this christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 28 context we would like to know how different variables xα ∈ x contribute to p(c|x). in other words, from overall information about the interaction between a taxon and its niche, how may we deduce and characterise the relative effects of interactions with individual niche factors? in other words, how does an overall interaction taxon-niche, as proxied by p(c|x), emerge from its component interactions p(c|xα)? the theoretical framework we use to answer these questions is probabilistic and bayesian based. the bayesian formulation of probability theory and statistical inference has had an enormous impact (berger, 1985). one of its most important advantages is the natural way in which qualitative in formation about beliefs can be incorporated as bayesian priors, as well as the way in which quantitative information may then be naturally incorporated into a posterior probability using bayes’ theorem. a second advantage is how it can naturally incorporate new information and beliefs and update posterior probabilities in the light of this new information. a third important advantage is that it gives a very natural setting for considering issues of causality (pearl, 2000). the fundamental basis for the bayesian formulation of probability is bayes theorem (8) where p(c) is the prior probability for the event c in the absence of the information x, which can be represented by a vector of variables x = (x1,x2,...,xm). p(x|c) represents the likelihood of observing the information x given the event c, and p(c|x) is the posterior probability that takes into account how the data x allows one to adjust the expectation of c relative to its prior. the evidence, p(x), is a normalisation factor independent of c. from a frequentist viewpoint all these probabilities may, in principal, be calculated directly from data. for example, p(c) = nc/n, where nc is the number of events of type c and n is the total number of events, i.e., the size of the statistical ensemble. however, p(c) may also represent our belief about the probability of the event c. this is relevant when the concept of an ensemble, measurable in frequentist terms, is not readily available. our diagnostic, equation (5), for characterising the presence of an interaction has a very natural bayesian interpretation, as i1 measures the deviation of the posterior probability p(c|xα) in the presence of the information xα from the prior distribution p(c). this is equally true if xα represents a single or a composite variable. from a hypothesis testing perspective, equation (5) provides us with an estimate of the degree of confidence we may have as to whether the information xα leads to an improvement in our estimation of the probability of c. often, to get rid of the c-independent evidence function, p(x), the following “score” function is used 𝑆𝑆(𝐶𝐶|𝑿𝑿) = 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝐶𝐶|𝑿𝑿)𝑃𝑃(𝐶𝐶̅|𝑿𝑿)) = 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝑿𝑿|𝐶𝐶)𝑃𝑃(𝑿𝑿|𝐶𝐶̅)) + 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝐶𝐶)𝑃𝑃(𝐶𝐶̅)) (9) where 𝐶𝐶̅ is the set complement of c (absences or no presences of c) and hence p(𝐶𝐶̅ ) = 1 − p(c). if s(c|x) > 0 it is more likely that the information indicates the presence of the event c, and vice versa for s(c|x) < 0. the x independent constant term on the right-hand side just accounts for the different class weightings. so, if the class is a small percentage then s naturally leans towards classifying instances into 𝐶𝐶̅ rather than c. seen as a rigid classifier, s(c|x) > 0 indicates that the instance x should be assigned to the class c. in the context of ecological interactions, s(c|x) > 0 indicates that the conditions x are favourable for the presence of c, with the higher the value of s the more favourable the conditions and vice versa for s(c|x) < 0. in the context of species distribution modelling, the estimation of p(c|x) or s(c|x), or an equivalent, such as p(c,x), can be done using different algorithms. the well-known maxent procedure (phillips et al., 2004, phillips et al., 2006) is one. even a more “black-box” procedure, such as garp (stockwell, 1999), is effectively doing the same thing. of course, s(c|x), as representing the niche in the hutchinsonian sense, can be used to map back into geographic space and thus provides a species distribution model whose performance can then be measured using any one of many metrics, such as the area under the receiver operating curve christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 29 (auc), or confusion matrix statistics, such as sensitivity and specificity etc. each has its pros and cons. unfortunately, when x is of high dimension then neither p(c|x) nor p(x|c) may be estimated reliably from data as ncx is either 0 or 1. in other words, if we define an environment with sufficient precision then there is either no co-occurrence with c or only 1 due to the fact that we will not find exactly the same environment in two different places. so, how to proceed? a general approximation, that maintains the bayesian philosophy is to approximate the likelihood function by assuming that the variables (x1,x2,...,xm) are independent. hence, ). in this case equation (9) becomes 𝑆𝑆(𝐶𝐶|𝑿𝑿) = ∑ 𝑠𝑠(𝑋𝑋∝) 𝑚𝑚 ∝=1 + 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝐶𝐶)𝑃𝑃(𝐶𝐶̅)) (10) where is ) the contribution (“score”) to the overall s(c|x) from the variable xα. if s(xα) > 0, < 0 then the factor xα contributes positively/negatively to the occurrence of the event c. this is just the wellknown naive bayes approximation (nba) that is still extensively used in many, many applications as it is easy to implement, computationally very efficient and very transparent (broos et al., 2011; burak and ayse, 2009; wang et al., 2007; wei et al., 2011). the maximum entropy algorithm and the nba are closely related (alfonso and vilar., 2007), as the latter can be written in maximum entropy form, but with a gibbs distribution that is factorizable. within this approximation we determine precisely how a measure of the overall interaction between taxon and niche variables, s(c|x), is composed of measures, s(xα), of the individual interactions between c and any given variable xα. thus, the higher/lower the value s(xα) the more it determines favourable/unfavourable conditions for the presence of c. if s(c|x) = 0 then the taxon c is effectively distributed randomly, the variables x having no overall effect. however, it does not follow that each s(xα) = 0. the composition of a set of potential niche variables can be such that their net effect is neutral by cancellation between positive and negative contributions. note that both s(c|x), are s(xα) are measures of interaction strength, in that they make no reference to how widely distributed are x or xα. in machine learning terms the latter is more related to the coverage of the features x or xα, meaning how many data instances, in this case cells, are represented by them. thus, s(xα) may be very positive for a given xα but, if this represents a rare species, nxα may be very small relative to n. in other words, although it is an important variable, it is not widespread. similarly, one may have a widespread variable that has only a weak interaction. our interaction diagnostic (5) considers both aspects of the interaction: strength and coverage. the chief criticism of the nba is its strong assumption of feature independence, where in the case of ecology one certainly knows that many niche variables will be highly correlated. in spite of this it works, almost unreasonably, well. as has been pointed out and quantified (stephens et al., 2017b), one reason the nba works better than expected is that the correlations between features can be both positive or negative across different feature combinations. in fact, it can be generalised, so that different features may be combined. this leads to an improved approximation, both in terms of predictive accuracy as well as enhanced understanding of which variables are correlated (stephens et al., 2017b). we will present some basic elements of this below (“beyond the naive bayes approximation”), as it is very relevant for the important question of confounding and causality. in practice, we must decide what data to use to estimate ε(c|xα), s(xα) or their niche counterparts with xα → x. the fundamental components are counts: ncxα, nxα and nc, that are taken from a statistical ensemble. the ensemble of most interest is that of n cells within which we count presence or no presence/absence. in the case of small samples, we may have that ncxα = nxα or that ncxα = 0. in these cases, s(xα) = ∞ or −∞ respectively. to avoid this the probabilities may be smoothed using a correction factor, such as the laplace correction (chen, 1996), whereupon ncxα → ncxα + a and nc → nc + b, where a and b are constants. a common choice is a = 1 and b = 2. model selection prior probabilities and model selection.—an important question is: which variables should be included in the model? a niche model that only accounts for climatic data, using a feature vector xa to christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 30 describe the distribution of a species c, would, relative to a uniform prior p(c), calculate the posterior probability to be 𝑃𝑃(𝐶𝐶|𝑿𝑿𝑎𝑎) = 𝑃𝑃(𝑿𝑿𝒂𝒂|𝐶𝐶)𝑃𝑃(𝐶𝐶) 𝑃𝑃(𝑿𝑿𝑎𝑎) (11) we may then ask how biotic factors, xb, may be added. in the bayesian formulation, we may take the posterior probability, p(c|xa), after including in abiotic variables and use it as a new prior probability for which we then compute the likelihood associated with the new biotic information, xb, and subsequently calculate the new posterior distribution, p(c|xbxa), that includes both biotic and abiotic factors. thus, 𝑃𝑃(𝐶𝐶|𝑿𝑿𝑎𝑎𝑿𝑿𝑏𝑏) = 𝑃𝑃(𝑿𝑿𝒃𝒃|𝐶𝐶, 𝑿𝑿𝒂𝒂)𝑃𝑃(𝐶𝐶|𝑿𝑿𝒂𝒂) 𝑃𝑃(𝑿𝑿𝑎𝑎|𝑿𝑿𝑏𝑏) (12) similarly, for the score function we have 𝑆𝑆(𝐶𝐶|𝑿𝑿𝑎𝑎𝑿𝑿𝑏𝑏) = 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝐶𝐶|𝑿𝑿𝑏𝑏𝑿𝑿𝑎𝑎) 𝑃𝑃(𝐶𝐶̅|𝑿𝑿𝑏𝑏𝑿𝑿𝑎𝑎)) = 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝑿𝑿𝑏𝑏|𝐶𝐶, 𝑿𝑿𝑎𝑎) 𝑃𝑃(𝑿𝑿𝑏𝑏|𝐶𝐶̅, 𝑿𝑿𝑎𝑎)) + 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝐶𝐶|𝑿𝑿𝑎𝑎) 𝑃𝑃(𝐶𝐶̅|𝑿𝑿𝑎𝑎)) (13) in the nba, ), where the first equality uses the fact that in this approximation the abiotic and biotic factors act independently. hence, 𝑆𝑆(𝐶𝐶|𝑿𝑿𝑎𝑎𝑿𝑿𝑏𝑏) = ∑ 𝑠𝑠(𝑋𝑋∝ 𝑏𝑏) + 𝑁𝑁𝑏𝑏 ∝=1 ∑ 𝑠𝑠(𝑋𝑋∝ 𝑎𝑎) + 𝑁𝑁𝑎𝑎 ∝=1 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝐶𝐶) 𝑃𝑃(𝐶𝐶̅)) (14) where na and nb are the numbers of abiotic and biotic variables respectively. we may now better understand what an approximation where only abiotic variables are used implies. essentially, it is equivalent to p(c|xaxb) = p(c|xa) and hence p(c|xb) = p(c). in the nba, we hence have ∑ 𝑠𝑠(𝑋𝑋∝ 𝑏𝑏) = 0𝑁𝑁𝑏𝑏 ∝=1 , which is most naturally interpreted as 𝑠𝑠(𝑋𝑋∝ 𝑏𝑏) = 0 for each 𝑋𝑋∝ 𝑏𝑏. if this is not true, then the approximation of omitting biotic variables will not be a good one. the same would be true if we omitted abiotic variables and considered only biotic ones. our methodology of bringing everything to the same spatial resolution and making every variable binomial allows us to make a direct comparison of the contributions from every single variable. furthermore, as we will see, we may use the predictions of the model to validate the inclusion or exclusion of certain variables using the pragmatic criterion of whether or not they improve model performance. a further subtlety is that although we leave out a certain set of variables that does not mean their influence is absent, due to the fact that there may be correlations between omitted and included variables such that the latter include effects from the former. this is just the effect of confounding. evaluating predictability = ranking interactions.—independently of the approximation used to calculate the likelihood, an important task is to compare and contrast the contributions of the different features, and/or feature combinations, in order to obtain a better understanding of their relative importance. different criteria may be used, both data based and “belief” based. the belief-based component is as discussed above—selection is based on supposed knowledge of what are important variables to include. thus, a modeller of the niche and species distribution of a predator may decide to include in as biotic factors only its two “known” principle preys. of course, this biased model should be compared with others to determine if it leads to better model performance and better understanding. for instance, if, in fact, the predator has several other preys that are unknown then their inclusion would be expected to improve model performance. to choose variables for inclusion in an overall model, an appropriate measure of the relative importance and associated statistical significance of a niche variable xα is just the binomial test (5). similarly, for a combination of features x we may use (7). we may use the same statistical criteria on ε, such as |ε| > 1.96, to decide whether or not to include the variable in our prediction model. thus, we link a machine learning-based concept—feature selection—with our definition of interaction, so that christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 31 features (niche variables) are included only if they represent significant interactions. beyond the naive bayes approximation.—in spite of the robust nature of the nba, it is important to try and determine which niche variables are correlated and whether a prediction model may be improved by considering them together. even more importantly, such correlations will help us study aspects of causality in order to help distinguish between causation and correlation. we consider first the impact of a pair of features, xα and xβ, following the same procedure as for one variable. in this case, for the joint probability distribution, p(c, xα,xβ) = p(xαxβ|c)p(c), we define (stephens et al., 2017b) ∆(xαxβ|c) = (p(xαxβ|c) − p(xα|c)p(xβ|c)) (15) as a measure of correlation between niche variables. significant deviations from the null hypothesis ∆(xαxβ|c) = 0 indicate the presence of significant correlations which we interpret as the fact that the niche variables xα and xβ are not independent with respect to the taxon c. in this case the interaction is ternary from the point of view of xα, xβ and c, though one can also consider it to be binary if we consider xαxβ as a combination variable. in analogy with ε(c|xα), we can use a binomial test to determine the degree of consistency with the null hypothesis, considering (stephens et al., 2017b) (16) equation (16) determines when there is a statistically significant correlation between the features xα and xβ. instead of considering the likelihoods, we may also consider the posterior probability p(c|xαxβ) directly. to determine the impact of a pair of features, xα and xβ, we may follow the same procedure as for one variable, considering (17) once again, in the case where the binomial distribution may be approximated by a normal distribution, then |ε(c|xαxβ)| > 1.96, which corresponds to the 95% confidence interval for testing consistency with the null hypothesis. in this case however, (17) cannot distinguish between the relative contributions of xα versus xβ. this can be done, though, by considering alternative null hypotheses. considering as null hypothesis p(c|xα) and p(c|xβ) in turn, we have (18) and, similarly, (19) equations (18) and (19) can be used as measures of the relative impact of one variable versus another. for instance, (18) expresses the contribution of the variable xβ to the posterior probability p(c|xαxβ), i.e., in the presence of the feature xα, but relative to the contribution marginalised over xα. these equations, in principle, allow us to determine which, of a combination of two variables, is the most important in terms of prediction of the class variable. moreover, they facilitate the determination and analysis of confounding variables. as an example, imagine we determine that ε(c|xα) is significant, but hypothesise that there exists a confounding variable xβ. in this case, we may consider ε(c|xαxβ;xα) and ε(c|xαxβ;xβ). if xβ is, indeed, a confounding variable, then we should find that ε(c|xαxβ;xα) > ε(c|xαxβ;xβ) as the real influence on c is from xβ and so, if we take p(c|xβ) as null hypothesis, we should find little residual predictability. note that (16) can be used to identify those feature combinations that should be considered together, and this information used to generate a generalized bayes approximation (stephens et al., 2017b), where the factorization of the likelihoods is not maximal, and which leads to an improved predictive model. inferring causality a criticism that has been levelled against our methodology is that it does not allow one to infer causality, in that an apparently important contribution from a biotic niche variable may, for example, just be reflecting the existence of an underlying abiotic factor that acts as a confounder (purse and golding, 2015). of course, the question of distinguishing correlation from causation is a topic of christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 32 great import across many branches of science and has a large literature (pearl, 2000), especially in the social and medical sciences, where there is a “classical” approach (hill, 1965) and a “modern” approach (rosenbaum and rubin, 1983; rubin, 1974; rubin, 1978). in the classical approach (hill, 1965), a set of criteria were introduced for judging epidemiological evidence of a causal relationship between a presumed cause and an observed effect. they are: 1. strength (effect size): a small association does not mean that there is not a causal effect, though the larger the association, the more likely that it is causal. we will use this in the context of comparing bayesian score contributions from different variable types. 2. consistency (reproducibility): consistent findings observed by different persons in different places with different samples strengthens the likelihood of an effect. this could be checked by considering different sample populations for a given hypothesis. 3. specificity: causation is likely if there is a very specific population at a specific site and disease with no other likely explanation. the more specific an association between a factor and an effect is, the bigger the probability of a causal relationship. we will consider this in the context of the “modern” approach by inserting sets of potential confounders. 4. temporality: the effect has to occur after the cause (and if there is an expected delay between the cause and expected effect, then the effect must occur after that delay). this requires time ordered data and is, in principle, possible with ecological data if it has such time ordering. 5. biological gradient: greater exposure should generally lead to greater incidence of the effect. however, in some cases, the mere presence of the factor can trigger the effect. in other cases, an inverse proportion is observed: greater exposure leads to lower incidence. 6. plausibility: a plausible mechanism between cause and effect is helpful. that there is some sound ecological/biological underpinning. 7. coherence: coherence between epidemiological and laboratory findings increases the likelihood of an effect. in other words, that a laboratory experiment, such as a determination of the positivity of a species with respect to infection by a pathogen is consistent with an inference of a relation between host and vector from point collection data. 8. experiment: “occasionally it is possible to appeal to experimental evidence”. 9. analogy: the effect of similar factors may be considered. for the purposes of illustration, an example where we will apply the above framework is that of the relations between climate, vegetation, herbivore and carnivore. in particular, below (“some representative results”), we will consider as a specific example of a causal chain: lynx rufus as a carnivore; sylvilagus floridanus as a known, important prey of the l. rufus; microchloa kunthii as a known, important food source of sylvilagus floridanus (hudson, 2005); and, finally, climate as a known, important factor in the presence and abundance of the plant species. by the nature of this causal chain, we are led to hypothesise that the presence/no presence of sylvilagus floridanus is more important to the presence of lynx rufus than the presence/no presence of microchloa kunthii, which in turn is more important than the presence/no presence of a particular climatic configuration. the question will be if the nature of this causal chain may be deduced from spatial data and to what degree we can characterise confounding? a bayesian framework for causal inference with the bradford-hill criteria in mind, we may use the formalism presented elsewhere (“beyond the naive bayes approximation”), to see how to infer causality. consider two factors, biotic or abiotic, xα and xβ, and their impact on the distribution of taxon c, as modelled by the conditional probability p(c|xαxβ). we would like to determine the combined impact of xα and xβ relative to some null hypothesis and use this analysis to determine the relative importance of xα and xβ in predicting the presence of c. the most natural starting point is to consider xα and xβ together as one composite variable and calculate (20) this can be compared with ε(c|xα) or ε(c|xβ) separately. what conclusions could we draw? if i1(c|xαxβ) > i1(c|xα) or i1(c|xβ) we would infer that the interaction of species c with α and β together is christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 33 stronger than that with either α or β separately. moreover, we may consider different combinations of xα = 0, 1 and xβ = 0, 1. thus, i1(c|xα = 1 xβ = 1) is the strength of the interaction in the presence of both factors, while i1(c|xα = 1 xβ = 0) and i1(c|xα = 0 xβ = 1) are the strengths in the presence/absence of α and absence/presence of β. finally, i2(c|xα = 0 xβ = 0) is the strength of the interaction in the absence of either factor. if i1(c|xα = 1 xβ = 0) > i1(c|xα = 0 xβ = 1) this tells us that the interaction between c and α is stronger than that between c and β. using bradford-hill criterion number 1 (hill, 1965] in the section entitled “inferring causality and identifying cofounders,” we will take that as an indication that α is causally closer to c than β is. we will provide some specific examples of this in the section entitled “some representative results.” note that it is not appropriate to compare the values of ε(c|xαxβ) itself directly with those of ε(c|xα) or ε(c|xβ), as, all else being equal, ε(c|xαxβ) will be less than ε(c|xα) or ε(c|xβ) simply because nxαxβ < nxα or nxβ. thus, ε(c|xαxβ) is naturally commensurate with ε(c|xαxγ) but not with ε(c|xα) or ε(c|xβ). predicting interactions we have defined an interaction as a deviation from an appropriate null hypothesis of the spatial distribution of a taxon conditioned on one or more abiotic and/or biotic variables. we have argued that the statistical diagnostic ε allows us to intuit a measure of the strength of this interaction as well as its importance (coverage), while s(xα) is more a direct measure of its strength. we have also argued that its interpretation depends on the statistical ensemble of data used to compute it, emphasising that point collection data is a set of special interest given its wide availability. although this definition of interaction is not scale dependent when we speak of using point collection data and corresponding species data to compute ε then we are naturally speaking of “macro” level interactions. as noted, we may view interactions between individual taxa or between a taxon and its niche. the score functions s(xα) provide a means for determining the relative contribution of xα to s(c|x) and we may consider ranking all included potential niche variables xα, α = [1,n], from highest, smax, to lowest, smin, as a means of comparing their relative strengths, with smax being the strongest positive interaction—most favourable niche variable—and smin the strongest negative interaction—most unfavourable niche variable. we may do the same using ε(c|xα), ranking them in descending order, from εmax to εmin, with the most positive values of ε(c|xα) representing the strongest and most important attractive interactions and the most negative values representing the strongest and most important repulsive interactions. the extra component in ε(c|xα) relative to s(xα), as has been amply discussed above, is its dependence on nxα which is a measure of the geographic coverage of the niche variable. our characterisation of interaction is such that if there is a deviation in the spatial distribution of a taxon relative to one or more niche variables at a given spatial resolution then we define that as representing an interaction. we must now ask—do these interactions represent ecological interactions in the standard sense? by ecological interaction here we refer to the standard micro interactions, such as predation, mutualism, commensalism etc. the first task is to determine if this empirical interaction has a natural biological interpretation. this is related to criterion 6 of bradford-hill. the naturalness of the interpretation depends on the labels associated with c and the xα. with a hypothesis in hand as to the potential ecological interaction we may then determine the degree to which the ranked list based on macro data represents known results about micro interactions, and also to what degree it offers new predictions that may subsequently be checked with a suitable experimental protocol. an illustrative example, that has been used frequently in applications of the methodology (stephens et al., 2009, stephens et al., 2016), is that of predicting hosts of a zoonosis, host range being an important parameter in understanding its biology, its transmission cycle and what are appropriate potential public health interventions. although the interaction of interest here is that of host-vector, there are multiple facets to this complex interaction which can be validated in different ways, depending on which part of the transmission cycle is used. thus, the micro interaction where the potential host is a blood meal for the vector is a necessary cooccurrence condition that there is an interaction pathogen-host/pathogen-vector. similarly, if a potential host is positive for a pathogen then it must have co-occurred at some point with the vector of that pathogen. co-occurrence is then a necessary condition for the pathogen to be passed from vector to host and vice versa. infection events are the analogues here of predation events— christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 34 a “micro” interaction. the question is then: what imprint, if any, do these micro interactions leave on the macro distributions of the involved species? we hypothesise that the host-vector micro interaction will result in any host being a relevant, favourable niche dimension for the vector. this implies that, as a niche dimension, there will be more “attraction” between hosts and the vector versus nonhosts and the vector, with the intuition that more co-occurrence will, all else being equal, lead to more potential micro interactions and, as a consequence, a systematic sampling of different potential hosts will yield the result that the probability of finding an infected individual of one potential host versus another is higher for a species with a higher value of ε(c|xα) or of s(xα). this presumes that sampling can be sufficiently intensive so that a sufficient number of individuals of each potential host species are collected. however, to increase our statistical power, we may go further and consider collections of species. thus, with the same logic, if we divide the ranked list into groups, we would expect more species to be found to be positive in the group with the highest values of ε versus the lowest. this just accounts for the fact that a species that co-occurs significantly with the vector and is widely distributed should lead to more positives than a species that does not co-occur significantly and is restricted in its range. note that this analysis is not concerned with whether or not a potential host could be physiologically able to be a host. potential hosts may be infected by a vector in the laboratory even though they never encounter one another in nature. note also that this does not simply imply that the most widespread species will be most likely to be the most highly ranked. this has been checked explicitly in several examples where the ranked list by ε and the ranked list in terms of pure geographic coverage are radically different, with the former leading to a much more precise prediction model. note, at this level, we are considering only two labels: vector and potential host = yes/no. however, there are many other potentially relevant labels, such as the competence of the potential hosts, or if certain potential hosts are blood meals only at a certain time of the year etc. that would be relevant for determining their role in the transmission cycle. it is for these reasons that we added the caveat “all else being equal” above. we can think analogously about the question of predation. if we take a predator, c, we can rank all potential prey species, xα, by ε(c|xα) or s(xα). cooccurrence is a necessary condition for predation, and we hypothesise that prey species will be an important niche dimension for the predator in that there will be more “attraction” between prey species and the predator versus non-prey species and the predator. again, we apply the intuition that more cooccurrence will, all else being equal, lead to more potential micro interactions and, as a consequence, a systematic sampling of different potential preys will yield the result that the probability of finding a confirmed prey from one potential prey versus another is higher for a species with a higher value of ε(c|xα) or of s(xα). in the case of both vector-host and predatorprey, the hypothesis is that species, xα, that represent important niche dimensions for the target species c should be correlated with more micro interaction events. in other words, that the micro interactions leave an imprint at the macro level due to the fact that they are associated with important niche dimensions. of course, it is not required that in all circumstances micro interactions should leave an imprint at the macro level. an important element to emphasise here is that the macro level characterisation of micro level interactions is a statistical inference. this occurs ubiquitously in epidemiology, where we may infer a link, for instance, between diabetes and obesity, or, with an even more indirect link, between diabetes and socio-economic status. the macro link between these is an imprint of an underlying set of micro physiological events. however, being a statistical inference, the correlation is not 100%. not all diabetics are obese and not all obese are diabetics. not all poor people are diabetics and not all diabetics are poor. similarly, not all species, xα, that have a large value of ε(c|xα) or s(xα) with a predator are by definition prey species. a principle reason for this is the existence in ecology of multiple micro interactions, and hence multiple labels, that influence their spatial distributions and therefore contribute to the overall macro interaction. thus, a predator distribution may also be affected by potential competitors, or climate, or many other factors. both considerations of the ecological significance of the labels for the species involved, as well as the inclusion of other potentially correlated niche variables, can help us further refine a list of micro interaction candidates, as we will see below. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 35 one last point concerns how we would identify micro interactions in the first place. from what data? are the micro interactions to be exhaustively tested? if so, how are we sure we have identified all interactions both direct and indirect? for any given micro interaction between a taxon c and a set of other taxa xα, α ∈ [1, m], all m potential interactions must be checked observationally. if we consider n taxa, c, then we might think that the number of potential micro interactions is nm. however, this too is potentially a grave underestimate, as any given taxon has multiple labels and each label can represent a different micro interaction. if we cannot exhaustively check all interactions how may we predict them using a model? independently of the type of model used its performance must be evaluated. to do so we must always start with some benchmark. what should that benchmark be and how do we make a precise specification of it? a natural benchmark is to make a hypothesis that the micro interaction only manifests itself uniformly within a particular set of candidate taxa, as identified using one or more labels that we believe to be important for the presence of the interaction. such a restriction is equivalent to a bayesian prior. for instance, for prey species of the bobcat, one may restrict to mammals, or lagomorphs. in the first case we would miss prey species that are not mammals, such as some birds. in the second, we would also miss mammal preys other than lagomorphs, such as rodents. in modelling terms, such a strategy introduces, by definition, false negatives. the safest bet of course is to consider all species, as then all possible candidates are included and there is no underlying bias from a chosen prior. by uniformly, we mean that the prior probability that the micro interaction exists is equal for every candidate taxon. for this probability we may have pre-existing information about the species that interact, such as a set of known preys of the bobcat or known hosts of a zoonosis. however, the quality of this benchmark depends on how well the set of observed interactions represents the full set of underlying interactions. supervised versus unsupervised learning models in statistical modelling terms, the above methodology using ε to predict interactions is a form of unsupervised learning, as no information whatsoever about the class we wish to predict—those with a particular micro interaction, such as parasitism or predation—enters into the model creation. the model is based only on the logic: micro interactions between taxa lead to a niche association between them which in its turn leaves an imprint in the spatial distributions of the taxa. if we are to use a supervised learning model however, then we must have some data that the model can learn from. in this case any list of potential candidates for a given micro interaction that enter into the model have to be labelled as such if they are already known cases. thus, for a predator-prey interaction, we use the label prey = yes for the known prey species on our list of candidates. we then need other data to be used as predictors. naturally, we can use the already defined parameters, ε(c|xα) and s(xα), but we may also appeal to other labels than prey if they are available. these labels could indeed already have been used in terms of a model selection, where sets of labels were omitted, as discussed in the section entitled “prior probabilities and model selection.” which labels to use depends on several factors—first and foremost, are they available? potentially relevant labels, such as those associated with phenotypic characteristics, or trophic guild, are not widely available, at least not at the level of covering large numbers of species. however, a set of labels that is always available is that corresponding to linneaean taxonomy. a species name, or indeed any higher order taxon, is a shorthand notation for a host of characteristics that make that species, or higher taxon, distinguishable from other species, or higher taxa. we would argue that implicit in these taxonomic names are associated labels that are relevant for many micro interactions, as well as identifiers as niche dimensions. with a set of labels in hand we may take a list of n species, already labelled with respect to the micro interaction of interest, i, and create a supervised learning model. the predictors are labels, xi, other than the label, ci, associated with the micro interaction. we may then use the bayesian framework of the section entitled “a bayesian framework for analysing ecological interactions,” and, in particular, the nba. now, however, we are using no spatial information whatsoever, just the labels of the species. for a set of predictor labels x, we may determine the overall score where 𝑆𝑆(𝐶𝐶i|𝑿𝑿) =∑𝑠𝑠(𝑋𝑋𝑖𝑖) + 𝑚𝑚 𝑖𝑖=1 𝑙𝑙𝑙𝑙 (𝑃𝑃(𝐶𝐶i)𝑃𝑃(𝐶𝐶i̅̅ ̅) ) (21) christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 36 ) is the contribution (“score”) to the overall s(ci|x) from the label xi. if s(xi) > 0, < 0 then the label xi contributes positively/negatively to the class ci. p(ci) = nci/n, with nci being the number of species in the list labelled with the micro interaction label. p(𝐶𝐶i̅̅ ̅ )= 1−p(ci) is the fraction of species on the list that do not have the micro interaction label. p(xi|ci) = ncixi/nci, where ncixi is the number on the list that have the micro interaction label and the label xi. the model may be trained on a subset of data and then performance measured on a test set. however, the true out-of-sample test would be to use the model to predict from among non-identified species which are most likely to have micro interactions with the target and then make a systematic and representative geographic sampling to determine which species are correctly identified. an example, that we will consider in more detail in the section entitled “identifying specific ecological micro interactions,” is where the label ci represents the class of known prey species of a given predator, such as the bobcat, the specific interaction being predation of the bobcat on other species. what about the labels xi for the candidate preys? as mentioned, there are myriad labels of possible relevance. the set we will use in our examples are linneaean classifications at the level of kingdom, phylum, class, order, family and genus. thus, sylvilagus, lutzomyia and macadamia are potential labels at the genus level, while mammalia and aves would be two labels at the class level. thus, the total score for a given candidate prey represented by a classification x = (xkingdom, xphylum, xclass, xorder, xfamily, xgenus), where xclass denotes in which taxonomic group is the candidate prey species, would be (22) the n species may be ranked with respect to this score function and a suitable performance measure, such as a confusion matrix or a roc curve calculated. in using supervised learning, we are subject to biases that are not necessarily present in the unsupervised models. for instance, the representativity of the set of species labelled with ci could be a cause for concern. if observations of the micro interaction i have been biased towards a certain subset of the n species, say mammalian preys of the bobcat have been much more studied than non-mammalian preys, then the model results will be subject to this bias. however, as this bias would potentially be present in both training and test sets model performance might not be affected. for this reason, we emphasised using the model in a completely out-ofsample set wherein the most likely preys that have not been observed as such are studied. naturally, the performance of the supervised model may be compared and contrasted with the unsupervised model that uses only spatial information. however, we may also adopt a meta-model viewpoint, constructing a model that is a mix of supervised and unsupervised bayesian models, using labels, such as the taxonomic labels mentioned above, as well as spatial information through s(xα) or ε(c|xα) and compare its performance to either the supervised or unsupervised model. we will do this explicitly in the section entitled “identifying specific ecological micro interactions.” the advantage of this is that the unsupervised model, considering only single species as niche variables and without further information on their labels, identifies only an overall interaction that may be due to the superposition of several micro interactions. on the other hand, the supervised model identifies relevant labels but does not account for the fact that a micro interaction can only take place if there is a co-occurrence. a mixed model can compensate for the individual defects of each separate model type. interpretation issues our proposal is that biotic interactions between organisms may be statistically inferred from their relative positions in space and time, where by statistical inference we mean to reach a conclusion based on evidence (a statistical ensemble of observations) and reasoning. a possible objection to this thesis is that: there are many other factors that affect the relative positions of biota in space and time, other than biotic interactions, which can act as confounders. a second objection, in the case where we use point collection data to model the positions and distributions of species, is that such data is biased and therefore not reliable. our answer to this objection is two-fold: first, and most importantly, does the model make predictions that can be tested and what is the model performance? secondly, the same christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 37 of the distributions show no interaction across time, the snapshot shows a state that gives the appearance of an interaction. without time dependent data this cannot be analysed. in that sense, by using point collection data without a time dependence, we are assuming that species are in “equilibrium.” of course, in the case of multiple niche variables it is very, very unlikely that all snapshots of all species pairs exhibit interactions by accident. a niche versus a network perspective up to now, we have emphasised macro interactions as being fundamentally related to the concept of a niche. in other words, if two taxa c and xα exhibit an interaction in terms of their co-occurrences it is because xα is an important niche dimension for c. we have thus framed everything in terms of “binary” interactions—between one taxon and one or more taxa, i.e., between c and xα, or c and x. however, our methodology has also been extensively applied at the community level, where we consider multiple target taxa, c, such that we have a set of target taxa and a set of niche taxa x. each c and xα can be represented as nodes of a network and ε(c|xα) used to weight links between these nodes. the resultant network we term a complex inference network (cin) (gonzález-salazar and stephens, 2012; stephens et al., 2009). the term “inference” is used as it does not represent specific, previously identified micro interactions between taxa, such as in a food web, but, rather, represents the set of inferred interactions between the taxa as defined by our ε diagnostic in equation (5). to what degree any given interaction represents an underlying micro interaction is discussed amply in previous sections. the advantage of cins is that they allow for an analysis of interactions at the community level. for instance, in figure 1 we see the network that results from using the ecological community functionality of the species platform3 to analyse the relation between two predators—the bobcat and the coyote—as target species, and species of the order lagomorpha, artiodactyla and rodentia as potential prey species. this example will be of relevance for the section “some representative results.” for the network links, only links corresponding to values of ε > 8 are considered thus representing the most important possible positive interactions. 3http://species.conabio.gob.mx biased data is already being used ubiquitously in fundamental niche modelling. in other words, if the data bias is sufficient so as to invalidate its use for modelling a biotic niche variable xα then it is also inadequate to model a target taxon c. in reference to the first objection, we believe our methodology in the sections entitled “a bayesian framework for analysing ecological interactions” and “a bayesian framework for casual inference” provides a framework for iteratively adding the presence/no presence or absence of different variables in order to test which is playing the dominant role, the intuition being that the more causal is a factor, the more important it is likely to be as a predictor. of course, confounding is an element that may be mentioned in the context of the analysis of any complex adaptive system, be it ecological, social, physiological etc. it is impossible to even list the full set of factors that may influence the presence of a given species. it is also impossible to determine unambiguously, from a statistical inference, that there does not exist a factor that has not been accounted for that is relevant or a confounder. to do so we would first have to have a consensus on the full set of potential confounders. we would then have to have a data representation of them in order to include them in the models. we would then have to check them one by one using our bayesian inference framework, or other, to determine their relative importance and to see if one variable con-founds another. an alternative and more sensible approach is to make a hypothesis about a specific, potential confounder for the appear-ance of a biotic interaction in macro distribu-tions, such as “shared history, geography, mi-gratory patterns, climate preferences...,” pro-vide the data that allows us to include that fac-tor in our model, and then test the hypothesis. to show the feasibility of this approach, we explicitly show in the section entitled “inferring causality and identifying confounders” an example where we can prove that rather than climatic factors being a confounder for apparent biotic interactions, on the contrary, it is biotic factors that are confounders for abiotic effects. one factor that can cause problems however, is the restriction to time independent data. not only because micro interactions themselves may be time dependent, but because we are taking the distributions of biota as “snapshots.” thus, two taxa might interact according to our definition, but this may be a result of the fact that, although the time trajectories http://species.conabio.gob.mx christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 38 some representative results the success of any scientific theory must be judged on its ability to explain and predict phenomena. theory can be used to explain and further understanding of already existing data or may be used to propose hypotheses and associated predictions that require new data to be collected. here, we will present a small, representative sample of the results that have been obtained using the present formalism that comply with all these requirements. we will show that: i) biotic factors can be generally and successfully introduced into species distribution and niche modelling; ii) confounding can be analysed by comparing the simultaneous presence/absence of multiple niche factors; iii) specific ecological micro interactions can be inferred from macro (point collection) data. we could exhibit the advantages of our approach in myriad examples, several of which are already in the literature, trying in each example type to show what links it to other examples and what makes it different. here, we can only try to show a pair of representative examples that show the power of our methodology hoping that they show the applicability to any other possible use case. we also hope that the reader will test these ideas in their own use case of interest using the powerful species platform (stephens et al., 2019) that implements our methodology in an open, easy-to-use, on-line environment that uses data from the sistema nacional de información de la biodiversidad (snib4) and, more recently, north american data from gbif.5 including biotic factors into niche and species distribution predictions perhaps the least controversial application of our methodology is to include biotic factors into the characterisation of species’ niches and their associated geographic distributions, without discussing any subsequent interpretation in terms of micro ecological interactions. in this case we use the bayesian models of 4 to determine p(c|x), where x can be any combination of abiotic or biotic variables. to create models, we use the species platform,6 which incorporates north american point collection data from both the snib and gbif databases. the snib itself contains data on over 57,000 species. as a first example, we will consider the construction of the niche of the bobcat lynx rufus and use it to illustrate many of the points we have discussed 4http://www.snib.mx/ 5https://www.gbif.org/ 6http://species.conabio.gob.mx figure 1. complex inference network between the bobcat (blue circle to the left) and the coyote (blue circle to the right) and the set of potential prey species (orange circles) from the orders lagomorpha, artiodactyla and rodentia. only the most important interactions corresponding to ε> 8 are shown. http://www.snib.mx/ https://www.gbif.org/ http://species.conabio.gob.mx/ christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 39 above. the data used for this example are from the snib and consider data for mexico only. the first question that arises is: which niche dimensions should be included in x? in principal, in species we may include in as covariates/niche variables for the bobcat all species other than the bobcat, as well as worldclim as a set of abiotic variables, and other environmental layers. as discussed above, by bringing all variables to the same spatial resolution and making all variables binomial, all potential niche dimensions can be fairly compared and contrasted. in figure 2 we see the relative performance of three models: one including only mammal species as biotic niche variables, one including only abiotic variables (worldclim) and another including both. what is shown there, for the two groups of variables chosen, is the total score from equation (10) in subsets (deciles) of spatial cells. all cells are ranked by their total score and then divided into deciles. thus, decile 10 represents the 10% of spatial cells, wherein the total score is highest by category, corresponding to those conditions that are most favourable to the presence of the bobcat. decile 9 is the next 10% of highest score cells—more favourable than decile 8 but less favourable than decile 10. decile 1 is the 10% of cells with lowest scores, the least favourable cells, or most “anti”-niche. the relative performance of the models is evaluated using: recall = true positive/(true positives + false negatives). in species, model performance is gleaned from a 70%-30% training-test split which is repeated five times. thus, any performance metric is an average over five such iterations. clearly the performance of the biotic model is far superior to that of the abiotic model, with the combination of the two being even better. simply put, mammals as niche variables are much more predictive than climatic data for presence of the bobcat. we should emphasise here that this exercise, as implemented in the species platform can, quickly and efficiently, be used to validate or invalidate the eltonian noise hypothesis (soberón and nakamura, 2009) for any of the many thousands of species present as target taxa. indeed, one will quickly convince oneself that the eltonian noise hypothesis is rarely valid when all potential biotic interactions are considered and that biotic factors in general are more important niche dimensions than climatic factors, with the latter characterising the anti-niche more than the niche, i.e., they affect much more the unfavourability of a place than its favourability. of course, worldclim data is only a representative subset of all abiotic factors, not an exhaustive set. similarly, mammals only represent a subset of biotic variables. however, using our methodology, in the species platform we may consider any arbitrary subset of potential niche variables and in any combination. thus, we may compare a particular abiotic factor with a particular biotic one, or any group thereof, or compare all 57,000 biotic variables with all abiotic variables. even the mere form of the distribution of scores as a function of decile is informative. note that the relative contribution of climate is much more significant in the first deciles than in the tenth. this confirms that climate affects only weakly the niche of the bobcat but is much more influential in determining the anti-niche. a subsequent analysis of the highest ranked versus lowest ranked mammals shows that known prey species of the bobcat appear with high values of ε (gonzález-salazar et al., 2013). besides comparing biotic and abiotic contributions, we can also compare different types of biotic variable. for instance, in figure 3 we see the performance of a model that includes mammals as biotic factors with another that includes the order magnoliales—a class of flowering plant—as biotic factors. these are chosen as an example of a group of biotic figure 2. performance of predicted distribution models for the bobcat based on abiotic variables only (worldclim), biotic variables only (mammals) and a combination. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 40 factors where there is no a priori reason to expect any relevant direct micro interaction with the bobcat. as we can see, the performance of the model using mammals is much superior to that using magnoliales, as might be expected for flowering plants as predictors of the bobcat distribution. from the form of the curves we may see that the plants play no role as niche variables for the bobcat, although their effect as anti-niche variables is apparent. of the 120 species of magnoliales as potential niche factors, only two have a positive score and have ε > 1.96. how should we interpret the fact that the presence of magnoliales is negatively correlated with the presence of the bobcat? that there is a repulsive interaction between the bobcat and these flowering plants? that they are, for some reason, in competition? of course not. first, we must interpret the apparent macro interaction from the point of view of biological plausibility. second, we must check for confounding and causality to determine the extent to which climate confounds the negative interaction of these flowering plants with the bobcat. this can be done using the methodology presented in the section “a bayesian framework for causal inference.” inferring causality and identifying confounders we will show that our methodology is capable of untangling the underlying causal relationships between niche variables in the context of a specific example: the characterisation of the niche of the bobcat; considering particularly the relations between a known prey species of the bobcat—sylvilagus floridanus, microchloa kunthii a known food source of s. floridanus and, finally, climate, as proxied by worldclim. an important reason for doing this arises from the criticism that a particular biotic niche factor may be confounded by climatic or other variables. in other words, that the reason why two species co-occur is because they share the same habitat preferences rather than that there is a particular biotic interaction between them (ovaskainen et al., 2010; royal et al., 2016). in this case we consider p(c|xαxβxγ) and analogously for ε(c|xαxβxγ), where xα = 0, 1 represents presence/no presence of sylvilagus floridanus, xβ = 0, 1 represents presence/no presence of microchloa kunthii and xγ = 0, 1 will range over the decile ranges for two significant climatic variables from worldclim: mean annual temperature (tmp r) and mean annual precipitation (prec r), with decile 10 representing the highest temperature/precipitation and decile 1 the lowest. thus, for a given cell, xγ = 1, if the temperature range denoted by xγ is present in the cell and zero otherwise. for each of the ten temperature and precipitation variables there is one presence/ absence variable. in table 1 we analyse the combined effects of the presence/no presence/absence of the two biotic factors and two abiotic factors in terms of heatmaps. the colour with the corresponding scales corresponds to p(c|xαxβxγ), the probability of a cell having a presence of the bobcat given the corresponding configuration of niche variables. remember that this is not an absolute probability to detect a bobcat in a given cell, but a relative measure based on point collection data. the numbers in each cell of the graph are the corresponding ε(c|xαxβxγ) values, which allow us to determine if a given value of p(c|xαxβxγ) is statistically significant or not. they also allow us to infer the coverage of the corresponding niche variable combination. the sign of ε also allows us to infer if the interaction is positive or negative with a positive/negative value indicating that p(c|xαxβxγ) >, < p(c)—the null hypothesis. by reading vertically from top to bottom we can see the effect of decreasing temperature or precipitation, concentrating on the presence variables figure 3. performance of predicted distribution models for the bobcat based on two classes of biotic variables: group bio 1 = mammalia and group bio 2 = magnoliales, and their combination. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 41 sociated with favourable conditions. however, when the prey is present and the plant not, we see that the bobcat is present at a frequency statistically significantly greater than the null hypothesis for almost all temperature ranges. only the highest temperature ranges, r10=1 and r9=1, correspond to conditions such that, despite the presence of the prey, the bobcat is present less than would be expected by the null hypothesis. the same is true when both biotic factors are present. we can also note that the presence of the plant in the absence of the prey is less significant in indicating the presence of the bobcat than when the prey is present and the plant not. so, we may note that lower/higher temperatures are associated with more niche/anti-niche conditions, while presence of a food source of the prey of the bobcat is a positive niche factor but less significant than the presence of the prey itself. however, the presence of both together leads to even more favourable niche conditions. the same considerations apply for precipitation. the presence of the prey is more important than the presence of the plant, which in turn is more important than precipitation. we can see that both high precipitation (r10=1, r9=1) and very low precipitation (r1=1) are anti-niche conditions. that we can isolate the effects of confounding in this case can be even more plainly seen in table 2, where we compare p(c|xαxβxγ) for xα = 0 (rabbit not present) and xβ = 0 (plant not present) against p(c|xαxβxγ) for xα = 0 or 1 (rabbit present or not present) and xβ = 0 or 1 (plant present or not present). in this latter case p(c|xαxβxγ) represents the marginalised probability p(c|xγ), where xγ is a purely climatic factor. similarly, for ε(c|xαxβxγ) and ε(c|xγ). by for temperature or precipitation as denoted by ri = 1. the values ri = 0 correspond to the absence of the corresponding climatic range. similarly, by reading horizontally from left to right, we may see the effect of increasing the overall presence of the biotic factors from both not present to both present. what is immediately clear is that the biotic factors play a much more important role than climate in determining what are favourable conditions, i.e., they play a preponderant role in determining the niche of the bobcat. this is fully consistent with our findings in the previous section, where we found that biotic factors were more generally associated with determining the niche, while abiotic factors were more relevant for the anti-niche. this is equally true here. the gradient in the probabilities left to right is much greater than the gradient top to bottom. in the absence of both biotic factors, corresponding to cells without a presence of either species, there is no temperature range that corresponds to a statistically significant positive niche factor, where the probability to find the bobcat is greater than the null hypothesis. on the other hand, the temperature ranges r10-r6 are associated with probabilities to find the bobcat that are less than the null hypothesis and correspond to a statistically significant negative interaction between the bobcat and these higher temperatures. in fact, in ranges r10 and r9 there is very little probability of finding the bobcat, independent of whether the biotic factors are present or not. turning to conditions where the plant is present, but the prey is not, we see that once again the probabilities to find the bobcat are consistent with the null hypothesis, as no temperature range is astable 1. probability p(c|xαxβxγ) and ε(c|xαxβxγ) for the bobcat with respect to a prey species, a food source of that prey species and climate. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 42 marginalising we are omitting the direct influence of the biotic factors. this is equivalent, in the bayesian sense, of a model selection where only abiotic variables are chosen, as is done in standard niche and species distribution modelling. far from it being the case that abiotic factors are confounders for biotic ones, here we see that, on the contrary, the abiotic factors are confounded by the biotic factors in terms of determining the niche. as can be seen comparing the two cases, the same tendency is observed in the case of conditioning on the no presence of rabbit and plant versus with no conditioning, in that colder temperatures and moderate precipitation are more favourable for the presence of the bobcat. however, the strength of the interaction is very different. if we consider only abiotic variables, the apparent probabilities for presence of the bobcat and their corresponding statistical significance are greatly enhanced, due to the fact that climate is correlated with the biotic niche factors. in summary, we can characterise the niche of the bobcat, quantify the contribution of each niche factor and deduce their correlations. presence of the prey species and/or the presence of one of the prey species food sources are positive niche factors, as well as non-extreme temperatures and precipitation. we can deduce the causal chain of factors, noting that the factor closest causally to the bobcat—its prey—plays a much more important role than the factor which is indirectly linked—the prey’s food source, which, in turn, is more important in determining the niche than the climate. this ranking of the relative importance of the niche factors: prey > prey food source > climate, is as you would expect for a vagile mammal such as the bobcat. thus, the direct interaction between bobcat and prey, is stronger than the indirect interaction between bobcat and food source of the prey which, in its turn, is stronger than the even more indirect interaction between bobcat and climate. moreover, we can see that climate is confounded by the underlying presence of relevant biotic factors. this is prima facie evidence that standard fundamental niche modelling, based only on abiotic variables, does not represent the true statistical relationship between climate and species distributions. rather, the relationship between climate and bobcat reflects the presence of important biotic confounders such as the bobcat’s prey species. thus, fundamental niche models need to include biotic factors with subsequent analysis of the potential confounding between one type of factor and another. table 2: probability p(c|xαxβxγ) and ε(c|xαxβxγ) for the bobcat in the absence of biotic factors, and p(c|xγ) and ε(c|xγ), where xγ represents just climate. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 43 of labels that is universally available are the standard linneaean taxonomic labels. indeed, these were used implicitly in the choice of the three proposed models with class = mammalia being used in one case and family = lagomorpha in the other. we will use those taxonomic labels here. our examples used georeferenced point collection data from the snib with corresponding taxonomic labels. a bibliographic search was conducted to determine the set of known (identified) prey species of the bobcat in mexico. sixty-seven prey species were identified. the data was divided into 70%/30% training/test and for each training set scores for each taxonomic label calculated using equation (22). these scores were applied to the test set and model performance was then calculated using the area under the roc curve (auc). one hundred iterations of this process were carried out and an average performance calculated along with its standard error. finally, combined models were created—one that used the score from the taxonomic labels as well as ε(c|xα) and s(xα) and another that used the score from the taxonomic labels and ε(c|xα) only. the results can be seen in table 3, where the correct class is prey = 1. what these results clearly show is that a statistical inference of this particular micro interaction, using only ε as a measure of interaction, is very accurate, with aucs of 0.98, 0.91 and 0.95 for the all, mammal and lagomorph groups. as can be seen, it is actually much better than the supervised model in the case of mammals and lagomorphs. of the 67 identified prey species, which corresponds to 0.12% of the total number of all species, 22.4% can be found in the top 0.1% (54 species) of the ranked list by ε, while 70% are found in the top 1% and 95.5% in the top 10%. similarly, for the mammals only model that incorporates 496 mammals: 40% of known prey species can be found in the top 10% by ε. in table 4 we see the list of the top 0.1% of species from the all list. the statistical ensemble here is composed of 26,944 cells of dimension 16 × 16 km. ni = 238 is the total number of cells with presence of the bobcat, nj is the total number of cells with a presence of the potential prey species, j, and nij is the number of cells with a co-occurrence. this list gives good insight into why it is perhaps difficult to accept that ecological interactions can be identified using point collection data. in terms of statistical inference, the model performance is outidentifying specific ecological micro interactions as we have used throughout an empirical characterisation of an interaction, defined via a deviation in the spatial distribution of a taxon from some “non-interaction” null hypothesis, it is not clear to what degree this definition of an interaction accords with the ecological definition in those cases where an ecological micro interaction has been identified and characterised. we will consider a test case where a verified ecological interaction is known7 and show that our empirical characterisation accords with the ecological one, using the unsupervised learning technique presented in the section “predicting interactions.” specifically, we consider the prediction of prey species of the bobcat8 we have emphasised the importance, from a bayesian perspective, of model selection: which variables are to be included in as possible prey species in the first place? we will consider three sets, in order of increasing bias: all species (53,722 species), mammals (496 species) and lagomorphs (14 species). by bias here, we mean that in the second and third groups we include the assumption (bayesian prior) that preys of the bobcat are only to be found among mammals, or among lagomorphs, respectively. thus, for each group we rank the included species by ε, with the hypothesis that those species most likely to have been identified as preys of the bobcat will have higher values of ε, as they have a higher rate of co-occurrence and are more disperse. there is no information that enters in this unsupervised model other than the distribution of biota, as proxied by point collection data. in particular, we use no information about any labels that might be at hand. as discussed in the “clarifying interactions,” potentially relevant labels could be big/small, slow/ fast, nocturnal/diurnal, terrestrial/aerial etc. such labels are not widely available in point collection databases for large numbers of species. however, one set 7by “known” here we mean that we know it exists with respect to a given label and have a set of examples. this does not imply, however, that those examples necessarily form a complete set, or we have a complete set of relevant labels. 8we have also carried out a similar analysis for other examples: i) pollination—leptonycteris curasoae, a bat species that pollinates agaves; ii) pollination—dalechampia scandens, a twining vine that rewards insect pollinators; iii) mutualism—aechmea bracteacta, a tank bromeliad that provides an ideal habitat for the development and refuge of aquatic and terrestrial organisms; iv) facilitation—neobuxbaumia mezcalaensis, a plant that depends on nurse plants to have a favourable microhabitat and avoid humidity loss due to direct contact with the sun. detailed results will be presented in another publication. note that the species platform can be used to consider and validate any other example where appropriate date exists. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 44 table 3: model performance for 4 model types: i) ε (unsupervised); ii) taxonomic labels (supervised); iii) ε, s and taxonomic labels (meta-model); and iv) ε and taxonomic labels (meta-model). species nij nj ni n epsilon score class order prey canis latrans 106 400 238 26944 54.75 3.7 mammalia carnivora 0 urocyon cinereoargenteus 85 535 238 26944 37.09 3.05 mammalia carnivora 0 taxidea taxus 32 87 238 26944 35.79 4.18 mammalia carnivora 0 lepus californicus 64 383 238 26944 33.1 3.11 mammalia lagomorpha 1 peromyscus maniculatus 99 871 238 26944 33.06 2.67 mammalia rodentia 1 otospermophilus variegatus 54 339 238 26944 29.61 3.06 mammalia rodentia 1 procyon lotor 56 371 238 26944 29.25 2.99 mammalia carnivora 1 tadarida brasiliensis 66 520 238 26944 28.78 2.79 mammalia chiroptera 0 sylvilagus audubonii 58 417 238 26944 28.43 2.9 mammalia lagomorpha 1 puma concolor 33 143 238 26944 28.36 3.52 mammalia carnivora 0 mephitis macroura 46 279 238 26944 27.86 3.1 mammalia carnivora 1 odocoileus virginianus 71 633 238 26944 27.78 2.65 mammalia artiodactyla 0 bassariscus astutus 45 270 238 26944 27.72 3.11 mammalia carnivora 0 sayornis saya 92 1045 238 26944 27.36 2.38 aves passeriformes 0 thomomys bottae 51 351 238 26944 27.32 2.95 mammalia rodentia 0 haemorhous mexicanus 118 1648 238 26944 27.23 2.16 aves passeriformes 0 conepatus leuconotus 41 236 238 26944 27.07 3.16 mammalia carnivora 1 bubo virginianus 62 519 238 26944 26.93 2.72 aves strigiformes 0 dipodomys merriami 79 814 238 26944 26.9 2.49 mammalia rodentia 1 corvus corax 110 1504 238 26944 26.65 2.18 aves passeriformes 0 spizella passerina 96 1197 238 26944 26.39 2.28 aves passeriformes 0 regulus calendula 89 1044 238 26944 26.39 2.35 aves passeriformes 0 icterus parisorum 65 590 238 26944 26.31 2.63 aves passeriformes 0 reithrodontomys megalotis 58 488 238 26944 25.97 2.72 mammalia rodentia 1 sylvilagus floridanus 62 564 238 26944 25.66 2.63 mammalia lagomorpha 1 ursus americanus 16 43 238 26944 25.46 4.2 mammalia carnivora 0 lanius ludovicianus 108 1573 238 26944 25.36 2.11 aves passeriformes 0 accipiter cooperii 81 939 238 26944 25.36 2.36 aves accipitriformes 0 table 4: the top 57 highest ranked species by ε corresponding to those species with the most important interaction with the bobcat. the true positive rate in this group is 22.4% compared to the null (random) benchmark of 0.1%. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 45 standing. however, let us examine what, apparently, are false positives on the list. first and foremost, we must be careful about judging that a false positive is just that—that the corresponding species is not a prey species, versus it has not yet been identified as such. for instance, taxidea taxus may well be a prey species (skinner, 1990) that has not been so identified in mexico. there are also several bird species that are very highly ranked and therefore posited to be potential prey species but that have not been identified as such. in this case, our list represents a set of predictions for new potential preys that have not been previously discovered. secondly, our methodology is based on the fact that there is an interaction by virtue of the “attraction” between these species as niche dimensions and the bobcat. this does not mean, however, that the interaction has perforce to represent a predator-prey interaction with the bobcat as predator. it does not even mean that the interaction has to be direct. for instance, the strong interaction between the coyote (canis latrans) and the bobcat may principally be through shared preys. this hypothesis is fully consistent with the cin shown in figure 1, where we see that there are many more strong interactions (potential preys) of both the coyote and the bobcat (orange nodes connecting the blue nodes of the bobcat (left) and coyote (right)) than are available to only one or the other. so, the bobcat and coyote have a substantial niche overlap in terms of their prey species. however, there are studies that show that, despite this overlap, there is little competition between them (major, 1987), except in conditions where food resources are scarce. in principle, the nature of the interaction between the bobcat and the coyote can be analysed further using the formalism of the “a bayesian framework for causal inference.” in other words, we may consider p(bobcat = present|coyote = present,prey = present) versus p(bobcat = present|coyote = present,prey = absent), just as was done for the case of the bobcat in “a bayesian framework for causal inference.” the above considered only spatial information. we can also adjoin labels and potentially improve the prediction model based only on ε. in table 5 we see the top five species of the all group as ranked by ε. the ranking, using the score function as deduced by a supervised learning model, is quite different, euphagus cyanocephalus 59 532 238 26944 25.16 2.64 aves passeriformes 0 peromyscus eremicus 63 604 238 26944 25.08 2.57 mammalia rodentia 0 buteo jamaicensis 119 1922 238 26944 24.87 2 aves accipitriformes 0 colaptes auratus 81 971 238 26944 24.84 2.32 aves piciformes 1 pipilo maculatus 61 594 238 26944 24.45 2.55 aves passeriformes 0 phainopepla nitens 67 709 238 26944 24.38 2.46 aves passeriformes 0 tyrannus vociferans 91 1237 238 26944 24.33 2.19 aves passeriformes 0 sayornis nigricans 92 1279 238 26944 24.12 2.16 aves passeriformes 0 passer domesticus 107 1684 238 26944 23.99 2.03 aves passeriformes 0 tyto alba 54 490 238 26944 23.98 2.63 aves strigiformes 0 myotis californicus 37 242 238 26944 23.95 3.01 mammalia chiroptera 0 zonotrichia leucophrys 65 692 238 26944 23.92 2.45 aves passeriformes 0 neotoma mexicana 49 411 238 26944 23.92 2.72 mammalia rodentia 1 turdus migratorius 70 799 238 26944 23.8 2.38 aves passeriformes 0 geococcyx californianus 73 865 238 26944 23.75 2.34 aves cuculiformes 0 perognathus flavus 41 299 238 26944 23.71 2.88 mammalia rodentia 0 setophaga coronata 108 1751 238 26944 23.63 2 aves passeriformes 1 auriparus flaviceps 70 809 238 26944 23.62 2.36 aves passeriformes 0 aegolius acadicus 17 56 238 26944 23.57 3.89 aves strigiformes 0 table 4: (continued from previous page) christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 46 with the coyote and the grey fox in particular being substantially downgraded in the list. although these species have an important macro interaction with the bobcat, principally indirect through shared prey species, as is consistent with table 5, they do not have taxonomic labels that are as highly correlated to known prey species, such as lagomorpha. however, from table 3 we see that the supervised model with taxonomic labels is much inferior to the unsupervised model for the mammal and lagomorph groups. so why does this supervised learning model performance decay so significantly? when all species are considered, there are taxonomic labels that are very predictive. for example, as the bobcat is a carnivore, any plant species will clearly receive a very negative score. however, when we get to the mammal group, the taxonomic labels within this set lose predictive power, as the bobcat has preys in multiple mammal genera, families and orders and, obviously, when we get to the lagomorph group the taxonomic labels lose relevance even further. interestingly, the spatial models alone using ε contain implicit information about the diet of the bobcat’s diet! the highest ranked plants in the list of all species are muhlenbergia wrightii (rank 61) and bromus carinatus (rank 99). as both are grass species that are potential food sources for many of the main prey species of the bobcat, we see that the interaction in terms of these plants as niche dimensions of the bobcat is indirect, with the bobcat’s interaction being intermediated by its prey species as confounders. once again, using the methods of the section entitled “a bayesian framework for causal inference,” this confounding can be analysed to better understand the true causal nature of the interactions. note also the extremely non-random nature of the ranking of different taxonomic groups in the all list. as plants represent about 50% of the overall set of species the odds of not hitting a plant until the 61st place in the list if the species were distributed randomly would be astronomically small. so, an unsupervised model based only on macro interactions, as measured by ε, and without reference to any characteristics of the potential prey species, yields an extremely predictive model for identifying the micro interaction between the bobcat and its preys. this is because an important prey species will be an important niche dimension, and this will manifest itself in the species’ distributions. however, ε alone is not equipped to distinguish between the different potential micro interactions that give rise to the observed macro interactions. on the other hand, suitable labels, such as the taxonomic labels used here, can identify characteristics of the known prey species, but then cannot distinguish between those that are niche dimensions that affect the distribution of the bobcat and those that do not. in other words, a lagomorph, such as sylvilagus brasiliensis, has the right characteristics to be a prey but does not share niche with the bobcat, as can be seen in table 6, where there are zero co-occurrences with the bobcat. a combination of unsupervised and supervised models is a way of including both macro level interactions and useful labels for distinguishing between different micro level interactions as contributors to the macro distributions. so, in table 5 we see that a model that uses both ε and the scores from the taxonomic labels enhances the rank of those species—lepus californicus and peromyscus maniculatus—that have both an important micro interaction with the bobcat and taxonomic characteristics that are identified with known prey species, while, at the same time, suppressing those species—canis latrans and urocyon cinereoargenteus—that have an important macro interaction with the bobcat but do not have taxonomic labels that are consistent with that interaction having the micro interaction predator-prey as an important contributing source. finally, considering only lagomorphs as potential prey, we see the list of all mexican lagomorphs ranked by ε in table 6. note that once again the model is extremely good, with a sensitivity of 83.3% and a specificity of 87.5%. identifying disease hosts we presented the case of predation above as a test case as it is distinct to what has been, up to now, the main area of application of the methodology— zoonoses. as mentioned, the transmission cycle of lynx/prey rank by (ε) rank by taxonomic labels rank by epsilon, score and taxonomic labels rank by epsilon and taxonomic labels canis latrans 1 174 170 172 urocyon cinereoargenteus 2 145 80 81 taxidea taxus 3 76 102 102 lepus californicus (prey) 4 2 2 2 peromyscus maniculatus (prey) 5 34 39 3 table 5: impact of taxonomic labels on ranking by ε for most important macro interactions with the bobcat. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 47 a zoonosis involves several local and direct micro interactions: host-vector, pathogen-host and pathogen-vector, that are related in a complex way. both the host range (the number and type of hosts) and the vector range (the number and type of vectors) are important factors in the transmission cycle and highly relevant for indicating how to combat a zoonosis. it is prohibitively costly to try and test every possible host and every possible vector to determine if and how it enters in the transmission cycle of a pathogen and even more costly to determine its relative importance. we hypothesise that the micro interactions between pathogen-vector-host are such that each biotic element enters as a potential niche dimension for the others and that will leave an imprint of the interaction at the macro level when considering the relative spatial distributions of vector and host. although we have considered multiple zoonoses—leishmaniasis (berzunza-cruz et al., 2015; stephens et al., 2009; stephens et al., 2016), chagas disease (ibarra-cerdeña et al., 2017; rengifo-correa et al., 2017), zika virus (gonzález-salazar et al., 2017), yellow fever, st. louis encephalitis, dengue and west nile virus, we will consider here, as a representative example, only leishmaniasis. in the case of leishmaniasis, the vector-host relation is accepted to be between a hematophagous insect vector—of the genus lutzomyia—and a mammalian host. in mexico, until recently, the number of confirmed hosts was only nine, of more than 430 possible candidate mammal species (stephens et al., 2009). following the logic that a necessary condition for the presence of the pathogen is the presence of the host and that the disease hosts will be favourable niche dimensions for the vectors we created a list, ranked by ε, of all mammals in mexico (stephens et al., 2009). the list was then checked against current knowledge in terms of the nine confirmed hosts, all of which corresponded to high values of ε, indicating a strong positive interaction in terms of our empirical definition. as with the example bobcat-prey, the predictive value of this unsupervised model is very high. however, as with the predation example, we must ask whether the set of confirmed hosts is representative and if it is complete, and if so, what is the nature of the false positives or false negatives in the model? in table 7 we see the 150 most highly ranked (most important) interactions by ε between the genus lutzomyia and potential mammal hosts. the previously confirmed hosts are denoted by “yes” in the column “conf.” taking as an example classification criterion that any mammal in the top 5% (corresponding to rank 21 in the list) is predicted to be a host then, if we accept that the only mammal hosts are the nine already confirmed hosts, the sensitivity (recall) of the model is 3/21 = 14.3%, which may be compared with the null hypothesis that there is no interaction between the distributions of vectors and hosts, wherein the probability to find confirmed hosts in any group would be 9/419 = 2.1%. thus, this simple model as a classification model for identifying known disease hosts of leishmaniasis is almost 681% better than that of a random model benchmark. if we take a larger group, the top 20%, then the sensitivity drops off, as species nij nj ni n epsilon genus prey lepus californicus 64 383 238 26944 33.1 lepus 1 sylvilagus audubonii 58 417 238 26944 28.43 sylvilagus 1 sylvilagus floridanus 62 564 238 26944 25.66 sylvilagus 1 romerolagus diazi 9 19 238 26944 21.66 romerolagus 1 sylvilagus cunicularius 17 147 238 26944 13.84 sylvilagus 1 sylvilagus bachmani 9 53 238 26944 12.52 sylvilagus 0 lepus callotis 11 100 238 26944 10.81 lepus 1 lepus alleni 6 67 238 26944 7.06 lepus 0 sylvilagus graysoni 1 5 238 26944 4.57 sylvilagus 0 lepus flavigularis 1 13 238 26944 2.62 lepus 0 sylvilagus insonus 0 3 238 26944 -0.16 sylvilagus 0 sylvilagus mansuetus 0 2 238 26944 -0.13 sylvilagus 0 sylvilagus gabbi 0 1 238 26944 -0.09 sylvilagus 0 sylvilagus brasiliensis 0 61 238 26944 -0.74 sylvilagus 0 table 6: ranking by ε of all lagomorphs. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 48 it should, with a sensitivity of 7/84 = 8.3%. thus, this model does a good job at predicting previously identified hosts. this analysis presupposes however, that no other mammal is a host beyond those already identified. if we take the list as a prediction model and a geographically systematic and random sampling is made, then from a given number of samples we would expect the highest ranked species to lead to more positives than the lower ranked species. this sampling was, indeed, carried out (stephens et al., 2016), with the result that 922 individuals from 70 species were collected and tested for the presence of the leishmania pathogen. of the 70 species tested, 22 that tested positive were previously unknown hosts of leishmania in mexico (shown in blue in the table 7). our unsupervised model now yields a sensitivity of 12/21 = 57.1%, compared to 2.1% for the random benchmark, and 14.3% if we assume that only previously confirmed species were positive. in the top 20% of the list, the corresponding sensitivity is 26.5%. one may argue that the percentage, 42.9%, of “false” positives associated with the 5% of highest ranked candidate hosts is very high, but this would be very misleading. take, for example, the bat species molossus rufus that was collected without presenting any positive individuals. this species is very highly ranked and therefore considered to be an important niche dimension for the lutzomyia genus. is this to be considered a false positive? to do so we must have a hypothesis about the expected infection rate. the infection rate over the whole sample of 922 individuals was 6.7%. taking this as the null hypothesis, given that only one individual of this species was collected, the probability that it represented a true negative was only 6.7%, far from a 95% confidence interval. indeed, of the 70 collected species, none could be discarded as a potential host at this level of confidence. in other words, no species of the top 5% or top 20% can be considered as a false positive, either because it has not been collected and tested or because it has not been collected in sufficient quantitable 7: list of top 150 most highly ranked potential mammal hosts of leishmania. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 49 ty. independently of this, from a statistical inference viewpoint, the model works extremely well. conclusion we have here tried to summarise both the essential conceptual and theoretical elements that enter in our formalism for defining, characterising, quantifying and predicting ecological interactions, in the hope that it will stimulate researchers to test the methodology on their own problems of interest and judge for themselves its merits. we believe that the representative use cases we have discussed, along with others in the literature, prove its worth. as many basic elements of the methodology are now available in an online platform—species—literally thousands of test cases, each with hundreds or thousands of covariates, can be produced, each in a matter of minutes. as a concrete example, the eltonian noise hypothesis may be validated or rejected for any species that exists in the snib (mexico) or gbif (north america) databases using any combination of biotic and abiotic variables. one can further compare and contrast the accuracy of any corresponding species distribution model associated with these combinations. this is equivalent to determining which variables are more niche-like—positively correlated with the species of interest—and which are anti-niche-like—negatively correlated with the species of interest. we can compare and contrast the role of every niche variable from a statistically level playing field by bringing every variable to the same spatial resolution and making every variable binomial. we can compare and contrast abiotic with biotic variables, or use different groups of biotic variables, grouped by any criterion we choose (given the data is to hand), such as by taxonomic label, or phenotype or genotype or, by ecological interaction labels or, indeed, whatever we choose. we can go further and use the methodology to analyse correlations between niche variables and thereby begin to study confounding between different variable types and answer questions about who confounds who? all of this can be done, which is a benefit to species distribution and niche modelling, without ever mentioning the word “interaction”. however, without understanding the role of interaction in ecology, all of this is just building a better mousetrap. the concept of interaction is fundamental to understanding how ecology at the micro level emerges into the macro level and manifests itself in the relative distributions of species or other taxa as a function of position and time. neutral theory (hubell, 2001) has provided us with a null hypothesis, based on the first principles of stochasticity, that allows us to have access to how the world would be if all species were similar in their per capita rates of birth and death, and hence provided a benchmark against which to compare empirical patterns in the abundance and diversity of species (e.g. marquet et al., 2014) and infer the relative importance of niche-related processes. our approach makes use of a similar philosophy by comparing observed patterns to the appropriate benchmark and then making inferences about the significance of their deviation in order to understand and assess to what extent micro interactions may leave an imprint upon macro distributions of taxa. the tension between neutral and niche also exists in physics. some systems, such as helium atoms, have strong (intra-atomic) micro interactions and very weak (interatomic) macro interactions. on the other hand, sodium and chlorine atoms have strong (intra-atomic) micro interactions and also have strong (but less strong) inter-atomic interactions. that’s how we get salt. so, is ecology more like helium or more like salt? clearly, the answer is that some ecological systems are more like helium and some are more like salt. how can we distinguish one from the other? first, by defining, as in many other areas of science, interactions with respect to the effect they have on the constituents of the system that is interacting. in particular, on their positions in time and/ or space. this is our path to defining interactions at the macro level when we cannot directly derive the macro interaction from the micro. thus, we defined interactions as being present when the spatio-temporal distributions of the things that are interacting is different to that in the absence of the interaction—the null hypothesis. without data on the macro distributions however, we can go no further. luckily, large point collection databases, notwithstanding questions about data biases, are a wonderful potential source of information about what is where and when. with this data we can determine the effect of one or any number of (niche) variables on a taxon of interest, be they abiotic or biotic. the problem is that the spatio-temporal distribution of one species is an emergent result of the micro level interactions with all potential niche variables. thus, there is no chance of isolating the effect of abiotic christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 50 variables (the fundamental niche) on a species distribution, as the latter is a result of those variables and all the biotic ones too. this is a problem of all complex adaptive systems, not just ecology. there are just too many variables involved to be able to isolate the effects of one variable by “controlling” for the rest. however, point collection data does allow us to set up a potentially infinite set of hypotheses by considering combinations of niche variables. for instance, we can determine the effect of the presence of a species of prey on the bobcat distribution by “controlling” for the effect of temperature, or precipitation, as was done in the section entitled “inferring causality and identifying confounders.” to link this to micro interactions however, we need labels for the niche variables that are ecologically relevant, and which can be used to interpret the macro interactions in micro terms. thus, sylvilagus floridanus has an important macro interaction with lynx rufus and therefore is a relevant niche variable in the pragmatic sense that where the rabbit is present so is the bobcat. however, it is our understanding of its role as a prey species of the bobcat that provides the micro property that allows us to understand why it is an important niche dimension. so, a macro interaction may be absent even though there is an underlying micro interaction—the helium scenario. however, there can be no macro interaction if there is no micro interaction. macro interactions can only emerge from the collective effects of multiple micro interactions. it is this fact that allows us to attempt to infer the existence and nature of micro interactions from the macro data. we gave several concrete examples of this. the base model there (our unsupervised learning ε model) is just based on the logic that things can’t interact if they don’t co-occur, neither at the micro nor the macro level. this led to very predictive statistical inference models for the considered interaction. by considering the various ecological labels of the involved species we could improve those models by determining which labels are associated with false positives versus false negatives. this can be done by hand—deciding for instance that the coyote cannot be a prey of the bobcat and removing it from the model or, in a more principled way, by using a supervised learning model trained on these labels. we believe that our methodology, and its implementation, available to all in the species platform, open up new horizons for a large set of analyses that simply were not possible before. what is more, the methodology is equally applicable to any spatio-temporal data of any resolution and of any data type, including public health data, census data, commercial data etc. the data just needs to be incorporated into the species platform. thus, we may ask not just what the niche of a vector of a disease is, but also what is the niche of the disease itself by using geo-referenced cases and their associated labels. given that it works in predicting new, unknown interactions it can be used to rank candidates for furthermore detailed analysis from large numbers of such candidates. acknowledgments this work has benefited from many fruitful collaborations from the theoretical perspective as well as the experimental one. in particular, we thank raúl sierra, victor sánchez-cordero, angel rodríguez, ingeborg becker, carlos ibarra, laura rengifo, marie josé tolsá, fabiola nieto, jesús sotomayor, gabriel garcía, gerardo suzán, benjamin roche. we are grateful for financial support from dgapa-papiit grant ig200217 and afb17008 (conicyt-chile). conflicts of interest the authors declare no conflict of interest. references abramsky, z., bowers, m.a. and m.l. rosenzweig. 1986. detecting interspecific competition in the field: testing the regression method. oikos 47: 199-204. alfonso, j. and d. vilar. 2007. bridging the gap between naive bayes and maximum entropy. in proceedings of the pris 2007, funchal (portugal), pp 59-65 álvarez-martínez, j. m., suárez-seoane, s., palacín, c., sanz, j. and j.c. alonso. 2015. can eltonian processes explain species distributions at large scale? a case study with great bustard (otis tarda). div. dis. 21: 123-138. aragón, p. and d. sánchez-fernández. 2013. can we disentangle predator-prey interactions from species distributions at a macro-scale? a case study with a raptor species. oikos 122: 64-72. aranda, m., rosas, o., ríos, j.d.j. and n. garcía, n. 2002. análisis comparativo de la alimentación del gato montés (lynx rufus) en dos diferentes ambientes de méxico. acta zool. mex. 87: 99-109. araújo, m.b., rozenfeld, a., rahbek, c., and p.a. marquet. 2011. using species co-occurrence networks to christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 51 assess the impacts of climate change. ecography 34: 897-908. araújo, m.b. and a. rozenfeld. 2014. the geographic scaling of biotic interactions. ecography 37: 406-415. araújo, c.b., marcondes-machado, l.o. and g.c. costa. 2014. the importance of biotic interactions in species distribution models: a test of the eltonian noise hypothesis using parrots. j. biogeogr. 41: 513-523. arditi, r., and l.r. ginzburg. 1989. coupling in predator-prey dynamics: ratio-dependence. j. theor. biol. 139: 311-326. atauchi, p. j., peterson, a. t. and j. flanagan. 2018. species distribution models for peruvian plantcutter improve with consideration of biotic interactions. j. avian biol. 49: 01617. belmaker, j., zarnetske, p., tuanmu, m.n., zonneveld, s., record, s., strecker, a. and l. beaudrot. 2015. empirical evidence for the scale dependence of biotic interactions. global ecol. biogeogr. 24: 750761. bethan v.p. and n. golding. 2015. tracking the distribution and impacts of diseases with biological records and distribution modelling. biol. j. linn. soc. 115: 664-677. berzunza-cruz m, rodríguez-moreno a, gutiérrez-granados g, gonzález-salazar c, stephens c.r., hidalgo-mihart m., et al. 2015. leishmania (l.) mexicana infected bats in mexico: novel potential reservoirs. plos neglect. trop. d. 9: 1-15. berger, j.o. 1985. statistical decision theory and bayesian analysis. berlin: springer-verlag. bradley, b.a., blumenthal, d.m., early, r., grosholz, e.d., lawler, j.j., miller, l.p., et al. 2012. global change, global trade, and the next wave of plant invasions. front. ecol. environ 10: 20-28. bristow, c.s., hudson-edwards, k.a. and a. chappell. 2010. fertilizing the amazon and equatorial atlantic with west african dust. geophys. res. lett. 37(14). broos, p.s., getman, k.v., povich, m.s., townsley, l.k., feigelson, e.d. and g.p. garmire. 2011. a naive bayes source classifier for x-ray sources. astrophys. j. suppl. s. 194: 4. borthagaray, a. i., arim, m. and p.a. marquet. 2014. inferring species roles in metacommunity structure from species cooccurrence networks. p. roy. soc. b-biol. sci. 281: 20141425. brown, j.h., kelt, d.a. and b.j. fox. 2002. assembly rules and competition in desert rodents. am. nat. 160: 815-818. burak, t. and b. ayse. 2009. analysis of naive bayes assumptions on software fault data: an empirical study. data knowl. eng. 68: 278-290. case, t.j. 1990. invasion resistance arises in strongly interacting species-rich model competition communities. p. natl. acad. sci. usa 87: 9610-9614. cazelles, k., araújo, m.b., mouquet, n. and d. gravel. 2016. a theory for species co-occurrence in interaction networks. theor. ecol. 9: 39-48. chen, s.f. and j. goodman. 1996. an empirical study of smoothing techniques for language modeling. proceedings of the 34th annual meeting on association for computational linguistics. clark, j.s., nemergut, d., seyednasrollah, b., turner, p.j., and s. zhang. 2017. generalized joint attribute modeling for biodiversity analysis: median-zero, multivariate, multifarious data. ecol. monogr. 87: 3456 colwell, r. k. and d.w. winkler. 1984. a null model for null models in biogeography. pp. 344-359. in: strong, d.r., simberloff d, abele, l.g., thistle (eds.) ecological communities: conceptual issues and the evidence, princeton university press, princeton, nj. connor, e.f. and d. simberloff. 1979. the assembly of species communities: chance or competition. ecology 60: 1132-1140 connor, e.f., collins, m.d. and d. simberloff. 2013. the checkered history of checkerboard distributions. ecology, 94: 2403-2414. crowell, k.l., and s.l. pimm. 1976. competition and niche shifts of mice introduced onto islands. oikos 27: 251-258. dayton, p.k. 1973. two cases of resource partitioning in an intertidal community: making the right prediction for the wrong reason. am. nat. 107: 662-670. delibes, m., zapata, s.c., blázquez, m.c. and r. rodríguez-estrella. 1997. seasonal food habits of bobcats (lynx rufus) in subtropical baja california sur, mexico. can. j. zool. 75: 478-483. diamond, j.m. 1975. assembly of species communities. p. 342-444 in: ecology and evolution of communities. m.l. cody and j.m. diamond (eds.). harvard university press, cambridge elith j. and j.r. leathwick. 2009. species distribution models: ecological explanation and prediction across space and time. annu. rev. ecol. evol. s. 40: 677-697. elton, c. 1927. animal ecology. sidgwick and jackson, ltd, london, 56. freilich, m.a., wieters, e., broitman, b.r., marquet, p.a. and s.a. navarrete. 2018. species co-occurrence networks: can they reveal trophic and non-trophic interactions in ecological communities? ecology 99: 690-699. garcía, j.a.m., martínez, g.d.m., plata, p.f.x., rosas, o.c.r., arámbula, l.a.t. and l.c. bender. 2014. christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 52 use of prey by sympatric bobcat (lynx rufus) and coyote (canis latrans) in the izta-popo national park, mexico. southwest. nat. 59: 167-172. gause, g.f. 1934. experimental analysis of vito volterra’s mathematical theory of the struggle for existence. science 79: 16-17. gehlke, c.e. and k. biehl. 1934. certain effects of grouping upon the size of the correlation coefficient in census tract material. j. am. stat. assoc. 29: 169-170. giannini, t.c., chapman, d.s., saraiva, a.m., alvesdos-santos, i. and j.c. biesmeijer. 2013. improving species distribution models using biotic interactions: a case study of parasites, pollinators and plants. ecography 36: 649-656. gilpin, m.e. 1975. limit cycles in competition communities. am. nat. 109: 51-60. gilpin, m.e., and f.j. ayala. 1973. global models of growth and competition. p. natl. acad. sci. usa 70: 3590-3593. godsoe, w. and l.j. harmon. 2012. how do species interactions affect species distribution models? ecography 35: 811-820. gonzález-salazar, c., stephens, c.r. and p.a. marquet. 2013. comparing the relative contributions of biotic and abiotic factors as mediators of species-distributions, ecol. model. 248: 57-70. gonzález-salazar, c. and c.r. stephens. 2012. constructing ecological networks: a tool to infer risk of transmission and dispersal of leishmaniasis. zoonoses public. hlth. 59, s2, 179-193. gonzález-salazar, c., stephens, c.r. and v. sánchezcordero. 2017. predicting the potential role of non-human hosts in zika virus maintenance. ecohealth 14: 171-177. gotelli, n.j. 2000. null model analysis of species co-occurrence patterns. ecology 81: 2606-2621 gotelli, n.j. and d.j mccabe,. 2002. species cooccurrence a meta analysis of j m diamonds assembly rules model. ecology 83: 2091-2096. gotelli, n.j., graves, g.r. and c. rahbek. 2010. macroecological signals of species interactions in the danish avifauna. p. natl. acad. sci. usa 107: 5030. hall, e.r. 1981. the mammals of north america, wiley, new york. haskell, e.f. 1949. a clarification of social science. main currents in modern thought 7, 45-51. heikkinen, r.k., luoto, m., virkkala, r., pearson, r.g. and j.h. korber. 2007. biotic interactions improve prediction of boreal bird distributions at macro-scales. global ecol. biogeogr. 16: 754-763. hill, a. b. 1965. the environment and disease: association or causation? j. roy. soc. med. 108: 32-37 holling, c.s. 1959. some characteristics of simple types of predation and parasitism. can. entomol. 91: 385398. hortal, j., jiménez-valverde, a., gómez, j.f., lobo, j.m. and a. baselga. 2008. historical bias in biodiversity inventories affects the observed environmental niche of the species. oikos 117: 847-858 hudson, r., rodríguez-martínez, l., distel, h., cordero, c., altbacker, v. and martínez-gómez. 2005. a comparison between vegetation and diet records from the wet and dry season in the cottontail rabbit sylvilagus floridanus at ixtacuixtla, central mexico. acta theriol. 50: 377-389 hubbell, s.p. 2001. the unified neutral theory of biodiversity and biogeography. princeton university press hutchinson, g.e. 1957. concluding remarks. cold spring harbor symposia on quantitative biology. 22: 415427. ibarra-cerdeña, c.n., valiente-banuet, l., sánchez-cordero, v., stephens, c.r. and j.m. ramsey. 2017. trypanosoma cruzi reservoir-triatomine vector co-occurrence networks reveal meta-community effects by synanthropic mammals on geographic dispersal. peerj 5, e3152. lidicker, w.z. 1979. a clarification of interactions in ecological systems. bioscience 29: 475-477. macarthur, r., and r. levins. 1967. the limiting similarity, convergence, and divergence of coexisting species. am. nat. 101: 377-385. major, j.t. and j.a. sherburne. 1987. interspecific relationships of coyotes, bobcats, and red foxes in western maine. j. wildlife manage. 51: 606-616. marquet, p.a., allen, a.p., brown, j.h., dunne, j.a., enquist, b.j., gillooly, j.f., gowaty, p.a. green, j.l., harte, j., hubbell, s.p., et al. 2014. on theory in ecology. bioscience 64: 701-710. mohd, m.h., murray, r., plank, m.j., and w. godsoe. 2017. effects of biotic interactions and dispersal on the presence-absence of multiple species. chaos soliton fract. 99: 185-194. morales-castilla, i., matias, m.g., gravel, d. and m.b. araújo. 2015. inferring biotic interactions from proxies. trends ecol. evol. 30: 347-356. openshaw, s., 1983. the modifiable areal unit problem. concepts and techniques in modern geography. norfolk, uk: geo books. ovaskainen, o., hottola, j. and j. shtonen. 2010. modeling species co-occurrence by multivariate logistic christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 53 regression generates new hypotheses on fungal interactions. ecology 91: 2514-2521. paine, r.t. (1992). food-web analysis through field measurement of per capita interaction strength. nature 355: 73. peterson, a.t. soberón, j., pearson, r.g. robert p.a., martínez-meyer, e., nakamura, m. and m.b. araújo. 2011. ecological niches and geographic distributions, princeton university press. (mpb-49) (monographs in population biology) princeton university press peterson, a. t., cobos, m. e. and d. jiménez-garcía. 2018. major challenges for correlational ecological niche model projections to future climate conditions. ann. ny. acad. sci. 1429: 66-77. pearl, j. 2000. causality, cambridge university press. phillips, p.c. 2008. epistasis: the essential role of gene interactions in the structure and evolution of genetic systems. nat. rev. genet. 9: 855-867. phillips, s.j., dudík, m. and r.e. schapire. 2004. a maximum entropy approach to species distribution modeling. in proceedings of the twentyfirst international conference on machine learning, pages 655-662. phillips, s.j., anderson., r.p. and r.e. schapire. 2006. maximum entropy modeling of species geographic distributions. ecol. model. 190: 231-259. pollock, l.j., tingley, r., morris, w.k., golding, n., o´hara, r.b., parris k.m. et al. 2014. understanding co-occurrence by modelling species simultaneously with a joint species distribution model (jsdm). methods ecol. evol. 5: 397-406. qiao, h., soberón, j. and a.t. peterson. 2015. no silver bullets in correlative ecological niche modelling: insights from testing among many potential algorithms for niche estimation. methods ecol. evol. 6: 11261136. rebolledo, r., navarrete, s.a., kofi, s., rojas, s., and p.a. marquet. 2019. an open-system approach to complex biological networks. siam j. app. math. 79: 619-640. rengifo-correa, l., stephens, c.r; morrone, j.j., téllez-rendon, j.l. and c. gonzález-salazar. 2017. understanding transmissibility patterns of chagas disease through complex vector-host networks. parasitology 144: 760-772. roberts, a. and l. stone. 1990. island-sharing by archipelago species. oecologia 83: 560-567. rosenzweig, m. l., abramsky, z., and b. kotler. 1985. can interaction coefficients be determined from census data? oecologia, 66: 194-198. rosenbaum p.r. and d.b. rubin. 1983. the central role of the propensity score in observational studies for causal effects. biometrika 70: 41-55. royan, a., reynolds, s.j., hannah, d.m., prudhomme, c., noble, d.g. and j.p. sadler. 2016. shared environmental responses drive cooccurrence patterns in river bird communities. ecography 39: 733-742. rubin, d.b. 1974. estimating causal effects of treatments in randomized and nonrandomized studies. j. educ. psychol. 66: 688-701. rubin, d.b. 1978. bayesian inference for causal effects: the role of randomization. ann. stat. 6: 34-58. sánchez-cordero, v., stockwell, d., sarkar, s., liu, h., stephens, c.r. and j. giménez. 2008. competitive interactions between felid species may limit the southern distribution of bobcats lynx rufus. ecography 31: 757-764. schoener, t.w. 1974. competition and the form of habitat shift. theor. popul. biol. 6: 265-307. sierra, r. and c.r. stephens. 2012. exploratory analysis of the interrelations between co-located boolean spatial features using network graphs. int. j. geogr. inf. sci. 26: 444-468. skinner, s. 1990. earthmover. wyoming wildlife. 54: 4-9. snow, j. 1855. on the mode of communication of cholera. london: john churchill. soberón j. and a.t. peterson. 2005. interpretation of models of fundamental ecological niches and species distributional areas. biodiversity informatics 2: 1-10. soberón, j. and m. nakamura. 2009. niches and distributional areas: concepts, methods, and assumptions. p. natl. acad. sci. usa 106: 19644-19650. soberón j. and a.t. peterson. 2004. biodiversity informatics: managing and applying primary biodiversity data. philos. t. roy. soc. b. 359: 689-698. stephens c.r., heau j.g., gonzález, c., ibarra-cerdeña c.n., sánchez-cordero v. and c. gonzález-salazar. 2009. using biotic interaction networks for prediction in biodiversity and emerging diseases. plos one 4, e5725. stephens, c. r., sierra-alcocer, r., gonzález-salazar, c., barrios, j. m., salazar carrillo, j.c., robredo e.e., and e. del callejo. 2019. species: a platform for the exploration of ecological data. ecol. evol. 9: 16381653. stephens, c.r., gonzález-salazar, c., sánchez-cordero, v., becker, i., rebollar-tellez, e., rodríguez-moreno, a., berzunza-cruz, m., balcells, c. d., gutiérrez-granados, g. and m. hidalgo-mihart. 2016. can you judge a disease host by the company it keeps? christopher r. stephens et al. – can ecological interactions be inferred from spatial data? 54 predicting disease hosts and their relative importance: a case study for leishmaniasis. plos neglect. trop. d. 10: e0005004. stephens, c.r., sánchez-cordero, v. and c. gonzález-salazar. 2017a. bayesian inference of ecological interactions from spatial data. entropy 19: 547. stephens, c.r, flores, h.h, and a. ruiz linares. 2017b. when is the naive bayes approximation not so naive? mach. learn. 107: 397-441 stockwell, d. 1999. the garp modelling system: problems and solutions to automated spatial prediction. int. j. geogr. inf. sci. 13: 143-158. wang, q., garrity, g.m., tiedje, j.m. and j.r. cole. 2007. naive bayesian classifier for rapid assignment of rrna sequences into the new bacterial taxonomy. appl. environ. microb. 73: 5261-5267. wang, b., wu, r., and x. fu. 2000. pacific east asian teleconnection: how does enso affect east asian climate? j. climate 13: 1517-1536. wei, w., visweswaran, s. and g.f. cooper. 2011. the application of naive bayes model averaging to predict alzheimer’s disease from genome-wide data. j. am. med. inform. assn. 18: 370-375. wilson, e.b. 1927. probable inference, the law of succession, and statistical inference. j. am. stat. assoc. 22: 209-212. wilson, w.g., lundberg, p., vazquez, d.p., shurin, j.b., smith, m.d., langford, w., gross, k.l. and g.g. mittelbach. 2003. biodiversity and species interactions: extending lotka-volterra community theory. ecol. lett. 6: 944-952. wisz, m.s., pottier, j., kissling, w.d., pellissier, l., lenoir, j., damgaard, c.f., dormann, c.f., forchhammer, m.c., gryntes, j.a., guisan,a., heikkinen, r.k., hoye, t.t., kühn, i., luoto, m., maiorano, l., nilsson, m.c., et al. 2013. the role of biotic interactions in shaping distributions and realised assemblages of species: implications for species distribution modelling. biol. rev. 88: 15-30. wootton, j.t. and m. emmerson. 2005. measurement of interaction strength in nature. annu. rev. ecol. evol. s. 36: 419-444.