v2i1 Volume 2, Issue 1 How to Explain Behavior: Author Precis Sam S. Rakover DOI: 10.33014/issn.2640-5652.2.1.rakover.1 Abstract This is an author precis of the book How to Explain Be- havior: A Critical Review and a New Approach by Sam S. Rakover (2018, Lexington). The precis has two sections. The first, basic methodological origin-points, treats the fun- damental ideas and premises concerning explanation and understanding. The second section outlines the book’s ar- rangement, and summarizes the content of each chapter. 1 Introduction This section describes the infrastructure on which I built the entire book. It covers matters from the philosophy of explanation in general to explanations in psychology more particularly. I start the discussion by emphasizing the following two gen- eral cornerstone ideas in scientific research methodology (ci- tations are taken from Einstein & Infeld, 1938/1950. See also Popper, 1959/1995): (A) The scientific attempt is to explain and understand re- ality: "The purpose of any physical theory is to explain as wide a range of phenomena as possible. It is justi- fied in so far as it does make events understandable." (P. 43) "All the essential ideas in science were born in dra- matic conflict between reality and our attempts at un- derstanding." (P. 280) (B) Scientific knowledge is provisional: "There are no eter- nal theories in science. It always happens that some of the facts predicted by a theory are disproved by experiment. Every theory has its period of gradual de- velopment and triumph, after which it may experience a rapid decline." (P. 77) Now I move on to discuss several methodological ideas that are more specific to the subject of the book: (1) To o!er an explanation for the studied phenomenon, a suitable scientific procedure has to be used. To propose explanations one has to know how, that is, sci- entists need a justified procedure for giving an explana- tion that is appropriate for the given phenomenon under scrutiny, just as they need a procedure suitable for making an observation or performing an experiment. In a way similar to the justifications for the use of the pro- cedure of an experiment that are based on several logical, epistemological, ontological considerations, so the proce- dure for giving scientific explanations is justified by various considerations—logical, epistemological, ontological, and so on up to common-sense considerations of daily life. A model of scientific explanation is not like a scientific theory, but is like a set of rules, norms, on how to behave. From this per- spective it follows that an explanation model, procedure, may not be tested empirically like a theory (which is sub- ject to an empirical test) but is evaluated by theoretical considerations and the degree of match of the procedure to the field under study. (2) A theory (or a hypothesis, model, law, mechanism) in itself is not able to propose an explanation. To explain a phenomenon, suitable explanatory informa- tion is needed. In science this information is provided by a theory, hypothesis, model and so on. However, the theory in itself is not able to explain the occurrence of the studied phenomenon. For the theory to do this, it has to function in the framework of a procedure created to provide scientific explanations. (3) The functioning of an explanation-procedure is based on rational relations. Rational relations are of di!erent kinds: from logical, math- ematical, statistical and causal to practical. To propose a scientific explanation based on non-rational relations is im- possible because what one will get is fortuity and chaos. (4) A procedure of explanation is an integral part of the 18 How to Explain Behavior: Author Precis methodology of science and is closely connected espe- cially to a procedure of empirical testing. The methodology of science is a collection of methods/pro- cedures for research whose main purpose is to bring about understanding of the world. Explanation-procedures are in- dissolubly bound to a procedure for empirical testing, such as the Hypothetico-Deductive (H-D) method (see Hempel, 1965, 1966). Without an empirical test we will never be able to know if the given theory is false or true. However, it should be stressed (as mentioned above) that scientific re- search never ends, and anyone who believes that a certain theory is wholly correct in fact drops out of the empirical science game (see Popper, 1972, 1959/1995). (5) The scientific explanation is an attempt to approximate the hidden process involved in the occurrence of the studied phenomenon. This can be expressed by the fol- lowing three basic assumptions: (a) There exists in nature some kind of Unknown Real-Process (URP) that is responsible for the occurrence of the studied phenomenon under certain conditions. (Note that except for the assumption regarding reality, the rest is unknown.); (b) Science proposes a theory (hypothesis, model, mechanism) which tries to approximate the URP and its function, thereby giving a reasonable explanation for the studied phe- nomenon; (c) The proposed theory is based on the scientific knowledge collected in the area under discus- sion. The degree of proximity of this theory to the URP is estimated by the degree of suc- cess/failure of this theory. The approach here (and of course in the entire book) is realistic in the broad sense. That is, I perceive science as a system that attempts to answer the question how nature (inanimate, plant and animal) functions by suggesting a theory that attempts to decipher nature’s code: how it is built, how it works, and according to what. (6) To date no one has been able to develop a theory that will describe and explain the connection between neu- rophysiological activity in the brain and consciousness, i.e., no one has succeeded in developing a mind/body theory. This carries the following implications: (a) Behavior cannot be explained solely by mech- anistic explanations (which are accepted in the sciences, in behaviorism, in cognitive psy- chology, and in physiological psychology), be- cause it neglects conscious behavior; (b) To propose a better account of behavior one must consider the conscious aspect explana- torily. This may be done as follows: (1) A procedure must be developed for proposing mentalistic explanations (e.g., explanations based on one’s will and belief); (2) It has to be shown that a procedure for mentalistic explanation meets the accepted methodological require- ments of science; (3) A theoretical framework has to be developed that allows coherent use of these two kinds of explanation (mechanistic, mentalistic). This account will be better than a purely mechanistic explanation. Idea (6) is one of the most important ideas sustaining the present book. Chapter 5 summarizes the empirical- theoretical arguments showing that the mind/body theory has not yet been developed, and chapter 6 describes in de- tail (a) Methodological Dualism (MD), which shows that like the mechanistic explanation, the mentalistic explana- tion meets the methodological requirements of science, and (b) the Multi-Explanation Framework (MEF), which allows one to construct a coherent theory in any psychological do- main based on these two kinds of explanation: the mecha- nistic and the mentalistic. The book does not aim to o!er another solution for the mind/body or consciousness/brain problem. Quite the re- verse: the goal is to circumvent this problem, thereby to achieve another goal, but no less important, namely to im- prove the explanation of behavior. This purposeful e!ort takes place within the framework of science by constructing an argument showing that a mentalistic explanation can be encompassed by the accepted methodological framework of science. 2 Chapter Summaries The book’s seven chapters plot a course of study from chap- ter 1, on the relation between explanation and understand- Volume 2, Issue 1 2. CHAPTER SUMMARIES 19 ing, to chapter 7, on a general procedure for giving expla- nations for the behavior of the individual. Chapter 1 discusses the relation between explanation and understanding. The prevalent view from 1948 (the year Hempel and Oppenheim published their influential paper) was that scientific understanding was based on scientific ex- planation. About fifty years later the literature began to suggest that scientific understanding was not built entirely on explanation, as it had additional sources such as the ability of the scientist herself. The article by Lipton (2009) even proposed the idea of understanding without explana- tion. In the spirit of that article I proposed a relatively new approach, namely the Dual Theory of Understanding, which posits a certain non-dependence between explanation and understanding. First, scientific understanding is based on consciousness of the scientific content associated with the phenomenon un- der study. Hence a sophisticated robot may be imagined that will supply an appropriate explanation for a given phenomenon without understanding the significance of the explanation. Accordingly, while giving an explanation is a mechanical rational process, understanding the explanation requires a human’s consciousness. Secondly, even before 1948 (e.g., in antiquity), the human being tried to understand the world and himself by means of certain procedures that di!ered from the explanation- procedures (models) known today. The chapter sets forth several illustrations of procedures for an understanding of that kind: a match between the phenomenon under study and some scheme that exists in the individual’s cognitive system; understanding person A when person B gets into person A’s shoes; the passage of a square through a round hole when the diagonal of the square is smaller than the diameter of the circle; etc. These procedures are based on the underlying assumption that there is some Unknown Real-Process (URP) that is responsible for the appearance of the phenomenon in question (in ancient times and even today God represents URP). The chapter also suggests that compared with understand- ing, explanation is interwoven with scientific knowledge based on the development of a rational scientific method- ology, which emphasizes for instance the importance of the development of a scientific model. As an example, the chap- ter considers the model developed by Rakover & Cahlon (1989) called the "Catch model." It is intended to reproduce from a witness’s memory the picture of the accused’s face. The phenomenon whose understanding is being sought is presented as the "reality-model," and what explains it (the reality-model) is a "theoretical-model" (hence the Catch model is based on two sub-models: the reality-model and the theoretical-model). Chapter 2 reviews critically nine models of explanation discussed in the philosophical literature, which (a) exerted great influence on the professional literature; and (b) are likely to have a great e!ect on providing explanations in psychology. The chapter opens with a review and discussion of the three models proposed by Hempel: the Deductive-Nomological (D-N) model, the Deductive-Statistical (D-S) model and the Inductive-Statistical (I-S) model. Together these are known as the "covering-law theory" because the common idea is that a natural law, a statistical law or a statistical generalization cover explanatorily the studied phenomenon. The three models gave rise to strong criticism, and as a result other researchers proposed alternative models of ex- planation. The first alternative model is the Statistical-Relevance (S- R) model proposed by Salmon (1971, 1990). The important ideas behind this model are that an explanation need not necessarily rest first on a logical argument, and secondly on high probabilities close to 1, as Hempel’s I-S model re- quires. This model experienced sharp criticism, and even- tually Salmon himself abandoned it for a new explanatory model based on the concept of causality. The Causal-Mechanical (C-M) model proposed by Salmon (1984) is based on the idea that the world is a net- work of interactive causal processes that spread through space and time. The explanation places the studied phe- nomenon within this network. Another causal model is the Manipulationist-Interventionist (M-I) model suggested by Woodward (2003). Its basic idea is the causal connec- tion between X and Y through a certain manipulation that changes X, hence Y changes accordingly and systematically; thus the causal explanation allows seeing what is likely to happen as a result of the counterfactual operation of a cer- tain manipulation. An additional explanatory model is the Kairetic model proposed by Strevens (2008). Its basic idea is the application of a method of elimination to explanation: di!erent factors are removed from the list of causes poten- tially responsible for the studied phenomenon; the goal is to see the elimination of which factor causes the phenomenon to disappear (or be likely to disappear). The elimination is repeated until only the factors (or factor) are (is) left with- out which the given phenomenon would not have occurred. The chapter concludes with a discussion of two models. One is the Unificationist model developed by Kitcher (1989), the other is the Pragmatic model proposed by van Fraassen (1980). The former posits that understanding the world depends on the ability to derive from the same patterns 20 How to Explain Behavior: Author Precis of inference many phenomena that are observed in scien- tific research (e.g., the Newtonian theory). van Fraassen’s Pragmatic model draws attention to the practical aspects, the background factors, of explanation as giving an answer to the questions: the subject of the question, the contrast- class (the possible alternatives for giving an answer), and the relevance-relation, a kind of relation between giving the answer and the question (e.g., “Compared with the alterna- tives of the answer, the phenomenon occurred because...”). An important conclusion that arises from this review is that there is no single concept for scientific explanation, which is enormously complex and multi-dimensional. Chapter 3 examines three additional models of explana- tion: rule-explanation; functional, teleological models of explanation; and the new-mechanistic model of explana- tion. Although these models are intended to handle hu- mans’ and animals’ behavior, they are unable to handle the most important component of this behavior, namely conscious-experience. My proposal how to deal with this appears in the last three chapters of the book. It is not possible to apply any of the models reviewed in the foregoing chapter to human behavior, for example, to the following cases: Dan stopped his car when the tra"c light changed to red; or Dan stood on a chair in order to take a book down from the shelf. This is because for these kinds of behavior no law, theory or appropriate empirical generalization exists. For example, stopping at a red light is not explained by an appeal to a law or empirical general- ization that covers behavior of this kind, but by looking up a tra"c rule created by a human being and which a driver is obliged to obey. Seven important di!erences between be- havioral rules and laws are discussed: for example, laws are universal and rules are local; laws are subject to an empir- ical test and rules are not (rules are evaluated according to the purpose for which they were devised). Although not all rules are present in consciousness (e.g., basic rules for generating language are innate), they di!er from laws of nature in that they concern information. Explanation by rules is also di!erent from teleological ex- planation. For example, people are not always aware of the aim of the rules and they obey rules automatically. Two approaches attempt to characterize functional, teleological, explanations as causal. The approach developed by Nagel (1961) tries to answer the question why a certain component appears in a given system; Cummins’ (1975, 1983) approach tries to answer the question how a certain component suc- ceeds in contributing causally to the e"cient function of a given system. The chapter ends with a discussion of the new-mechanistic explanatory model. According to this model, an expla- nation for the given phenomenon is proposed by an ap- peal to a mechanism. It can be decomposed into di!erent components with di!erent functions, unique in their array and interaction, which produce the studied phenomenon. Bechtel (2008a, b) draws a distinction between a new- mechanistic explanatory model, which suits the sciences, and one that suits psychology, that is, a mentalistic mecha- nism, which processes information—the Information Pro- cessing (IP) mechanism. There is an important di!er- ence between Bechtel’s (2008b, c) approach to a mentalis- tic mechanism and Rakover’s approach. While according to Bechtel processing information occurs on the non-conscious cognitive level, according to Rakover mentalistic activation takes place on the conscious level. Chapter 4 discusses whether an explanation model in it- self can be subjected to empirical test. The chapter raises arguments against the idea that the correctness of the ex- planation model can be tested empirically. First, the prediction that emerges from a certain theory under certain conditions can be confirmed or refuted. In both cases the empirical result carries no implications for the explanatory model itself, but only for the theory in- serted in the explanatory model. The reason is that the explanation model is just a procedure that guides the re- searcher on how to propose scientific explanations. An ex- planation model is not like a hypothesis, theory or scientific model, which may pass/fail an empirical test. A procedure is assessed by its rational course of action (e.g., it does not create situations that contradict each other) and its appro- priateness for a given field. In this sense, a procedure for giving explanations is akin to an administrative procedure, such as tra"c rules that advise people how to behave in certain situations. A rule of the road is assessed as suit- able or unsuitable for a certain field of behavior but it is not judged right or wrong. The chapter goes on to discuss this distinction as against the accepted distinction between descriptive, normative, theories and prescriptive ones. Secondly, we may look at the argument that slight use of an explanation model Mx attests to its incorrectness (i.e., Mx is refuted). Here the question that arises is how we may ex- plain this finding of infrequency of Mx use. Clearly, it can- not be explained by Mx itself (because it has already been refuted), so we must seek another model—My. However, it is reasonable to suppose (according to those who assume that an explanation model can be tested empirically) that this model too fails empirically. Why? Because there is very high probability that very often theories that used My proposed predictions that were disconfirmed. Hence, not only is My not suitable for use as a procedure for giving explanations, no explanation model exists that can explain Volume 2, Issue 1 2. CHAPTER SUMMARIES 21 the finding of slight use of Mx —because all explanation models at some time have been associated with refutation of empirical results of di!erent theories. Clearly then, em- pirical success/failure has no implications for the degree of correctness of a procedure for explanation. Based on a review of the explanation models found in the literature, chapter 4 o!ers a list of several constraints, char- acteristics and properties for constructing an explanation model, which have been described above. These make up the infrastructure for building explanation models for psy- chology (see especially chapters 6 and 7). Chapter 5 discusses the following question: is it possi- ble to comprehend mentalistic concepts and explanations (based on the individual’s inner world) with the aid of mechanistic concepts and explanations (based on and ac- cepted in the sciences)? The answer is negative, and it forges the way to the development of a new methodological approach to proposing explanations, namely the Method- ological Dualism (MD) and Multi-Explanation Framework (MEF), whose principal purpose is to propose a procedure to explain behavior that is saturated with consciousness. This methodology is based on the fact that to date no one has succeeded in developing a mind/body theory. Were it possible to propose such a theory, it would be possible to forgo the DM & MEF approach, because everything would obtain a mechanistic explanation. But as chapter 6 shows, this is not the situation so far. Chapter 5 breaks down the above fundamental question into the following four questions: (1) Can human behavior be explained mechanistically alone, without recourse to the conscious-experience concept? (2) Has a theory been successfully developed that explains the connection between neurophysiological processes in the brain and conscious-experience? (3) Can mentalistic explanations anchored to the concept of conscious-experience be reduced to mechanistic ex- planations (e.g., to a neurophysiological explanation)? (4) Has a very complex and sophisticated computer proven able, like us, to behave with conscious-experience? An a"rmative answer to any of these four questions means that there is no need to develop MD and MEF because the standard methodology developed in the sciences can han- dle research in psychology. Only a negative answer to all these questions will oblige one to think about the develop- ment of a new explanatory approach suited to the field of psychological research. The answers to these questions are negative. Here are three popular arguments that support the negative answers. First, it seems that mentalistic theories cannot be reduced to mechanistic theories. One of the important factors that prevents this kind of reduction is multiple re- alizations. If a mental state or process is defined by its function, it may be realized in various ways and with var- ious materials. Hence, if reduction of a mentalistic theory to a mechanistic depends on the possibility of identifying a concept in one theory with another, this possibility is de- stroyed because according to multiple realizations there are many ways to realize and identify mentalistic concepts. Secondly, a computer, as sophisticated as it may be, is inca- pable of reaching a condition similar to humans’ conscious- experience. One of the most powerful and famous argu- ments in support of the negative answer is Searle’s (1980) "Chinese Room" thought experiment. According to this, activation of physical signs by means of syntactic rules (which are operated by the computer) cannot impart se- mantic meaning to these signs. Finally, a very important obstacle blocking the development of a mind/body theory is "consciousness epiphenomenal- ism," which states that mental concepts and explanations are inessential. The reason is that the entire explanation of behavior falls on the shoulders of the neurophysiological process. Chapter 6 describes and justifies the approach for Methodological Dualism (MD) and Multi-Explanation Framework (MEF), which is based on chapter 5. Chap- ter 6 is divided into several sections. The first proposes an argument justifying why MD & MEF should be developed. Namely, if we attempt to o!er an improved explanation for behavior saturated with consciousness there is no way other than adding to a mechanistic explanation a mentalistic ex- planation. The second part of the chapter develops Methodological Dualism (MD). The chapter focuses particularly on the tele- ological explanation that combines the individual’s will, be- lief and behavior. It presents a new proposition to conceive the following as a mentalistic explanation-model – a proce- dure for creating specific will/belief explanations: [Will/Belief]: If X wants G and believes that behavior B will realize his/her will, then X will perform B. This proposition immediately raises the following ques- tion: does this mentalisitic explanation-procedure satisfy the methodological requirements for explanation accepted in science? The answer is yes. 22 How to Explain Behavior: Author Precis The third section develops a theoretical framework, the Multi-Explanation Framework (MEF), by means of which it will be possible to construct a specific theory (a specific "MEF-theory") for a specific research project based on two kinds of explanation: mechanistic and mentalistic. Because the theory rests on multiple explanation models, a major di!erence arises between (a) the way a theory supplies an explanation in science and (b) the way an MEF theory sup- plies an explanation in psychology. An MEF-theory is based on a match between explanation- procedures (mechanistic or mentalistic) and behavior (it is called the "explanation/behavior match" for short). While there are behaviors that can be accounted for satisfacto- rily by an appeal to mechanistic explanations only, most behaviors and their decomposed components must be ap- proached by both mechanistic and mentalistic explanations. The coherence of a specific MEF-theory is achieved by matching appropriately the explanation-procedure to the behavior and its components by means of several guidelines based on theoretical-empirical knowledge and the princi- ple of explanation-matching. According to the latter, the components of a mechanistic behavior (a whole behavior that has been accounted for mechanistically) cannot be ex- plained by mentalist explanations, while the components of a mentalistic behavior (a whole behavior that has been accounted for mentalistically) can be explained mentalis- tically as well as mechanistically. The appropriate expla- nation/behavior match solves several methodological prob- lems raised by an MEF-theory. Finally, the last section summarizes the main ideas of the present approach and compares it to other relevant ap- proaches. Chapter 7 presents A General Explanation Procedure (GEP) for understanding an individual’s behavior. As can be seen from the previous chapters, psychology has used a fairly large number of explanation models or pro- cedures: di!erent variations of Hempel’s models, di!erent kinds of causal processes, mechanisms based on the com- puter metaphor, and di!erent kinds of procedures based on neurophysiological processes. This employment of multi- plicity of explanation models raises the following question: is it possible to propose for psychology a general procedure for giving an explanation that is likely to spread its wings over other models too? The answer is a"rmative: the GEP. It is based on three major theoretical ideas: (a) It is assumed that in the individual a certain real-process exists that is involved in the generation of the behavior un- der study; although the nature of this process is unknown, it is further assumed that it presents the real explanation for the behavior in question [this process is called the "Un- known Real-Process" (URP)]; (b) A researcher may approximate the URP by construct- ing a theory (or a hypothesis, model, mechanism, etc.) that portrays the properties and operation of the URP in gen- erating the studied behavior. If the proximity between the theory and URP is high, it may be proposed that the theory provides us with a good account of the behavior in question; (c) The degree of proximity of the theory to the URP is estimated by the development of a "proximity index.” It is based on the theory’s successful predictions. Clearly, a confirmed theory is closer than a refuted theory to the URP and to the real explanation. The generality of GEP is expressed in two respects. First, the URP can be approximated by diverse theories or hy- potheses, models and mechanisms; Secondly, the GEP may handle various explanation- procedures (models) that researchers have been using in their studies. These models may be viewed as various routes for approaching the real explanation. To illustrate the GEP, the chapter analyzes several explana- tions given for learning in animals, perception and recogni- tion of faces, and everyday purposive behavior (will/belief accounts). Finally, the chapter discusses the basic philo- sophical concepts and assumptions that underlie the GEP. References Bechtel, W. (2008a). Explanation: Mechanism, modular- ity, and situated cognition. In P. Robbins & M. Ayd- ede (Eds.), Cambridge handbook of situated cognition. Cambridge, UK: Cambridge University Press. Bechtel, W. (2008b). Mechanisms in cognitive psychology: What are the options? Philosophy of Science, 75, 9830- 9994. Bechtel, W. (2008c). Mental mechanisms: Philosophical perspectives on cognitive neuroscience. London: Rout- ledge. Cummins, R. (1975). Functional analysis. The Journal of Philosophy, 72, 741-765. Cummins, R. (1983). The nature of psychological explana- tion. Cambridge, MA: The MIT Press. Einstein, A. & Infeld, L. (1938/1950). The evolution of physics. London: The Scientific Book Club. Hempel, C. G. (1965). Aspects of scientific explanation and other essays in the philosophy of science. New York: The Free Press. Hempel, C. G. (1966). Philosophy of natural science. En- Volume 2, Issue 1 References 23 glewood Cli!s, NJ: Prentice-Hall. Hempel, C. G. & Oppenheim, P. (1948). Studies in the logic of explanation. Philosophy of Science, 15, 135-175. Kitcher, P. (1989). Explanatory unification and the causal structure of the world. In P. Kitcher & W. Salmon (Eds.), Scientific explanation (pp. 410-505). Min- neapolis: University of Minnesota Press. Lipton, P. (2009). Understanding without explanation. In H. W. De Regt, S. Leonelli, K. Eigner (Eds.), Scientific understanding: Philosophical perspectives (pp. 43-63). Pittsburgh: University of Pittsburgh Press. Nagel, E. (1961). The structure of science: Problems in the logic of explanation. London: Routledge & Kegan Paul. Popper, K. R. (1959/1995). The logic of scientific discov- ery. London and New York: Routledge. Popper, K. R. (1972). Objective knowledge: An evolution- ary approach. Oxford: Oxford University Press. Rakover, S. S. (2018). How to explain behavior: A critical review and new approach. New York: Lexington Books. Rakover, S. S. & Cahlon, B. (1989). To catch a thief with a recognition model: The model and some empirical results. Cognitive Psychology, 21, 423-468. Salmon, W.C. (1971). Statistical explanation. In W. Salmon et al. (Eds.), Statistical explanation and sta- tistical relevance (pp. 29-87). Pittsburgh: University of Pittsburgh Press. Salmon, W. C. (1984). Scientific explanation and the causal structure of the world. Princeton: Princeton University Press. Salmon, W. C. (1990). Four decades of scientific explana- tion. Minneapolis: University of Minnesota Press. Searle, J. R. (1980). Minds, brains and programs. The Behavioral and Brain Sciences, 3, 417-457. Strevens, M. (2008). Depth: An account of scientific expla- nation. Cambridge, MA: Harvard University Press. van Fraassen, B.C. (1980). The scientific image. Oxford: Clarendon Press. Woodward, J. (2003). Making things happen: A theory of causal explanation. Oxford: Oxford University Press. Volume 2, Issue 1 Two Methods of Calculating Axiom Size Jonathan Bartlett DOI: 10.33014/issn.2640-5652.2.1.bartlett.1 Abstract 1 Introduction Modeling insight has been a long-standing problem for cog- nitive psychology. Bartlett (2012) noted that humans tend to be able to infer axioms about a system in order to solve problems. However, the framework proposed there su!ers from a few theoretical problems. The first is that axioms can overlap. Therefore, identifying distinct axioms is di"- cult. Second, the framework relies on an ordering of axioms from “small” to “large” (i.e., easy to di"cult), but fails to provide a mechanism to measure the size of axioms. Here we identify two possible mechanisms for calculating the size of an axiom, and how it may help future progress of cognitive research. 2 Dependent Axioms Bartlett (2012) di!erentiated between two kinds of axioms—independent axioms and dependent axioms. An independent axiom is one which defines the parameters of a formal axiomatic system. However, as Gödel points out, any formal axiomatic system has truths which are not provable within the formal axiomatic system (Raatikainen, 2018). These truths cannot be theorems within the ax- iomatic system, because then they would be provable within the system. However, these truths are dependent on the structure of the axiomatic system. In other words, given a particular axiomatic system A, some truth B is set for certain, but cannot be proved within A. Therefore, B is not a theorem within A, but its truth value is set by the parameters of A. Because of this, we identify such a truth as a dependent axiom. If B is added to A as an axiom, A becomes A!. This new system can prove B (because B is an axiom), and may also be able to prove other new truths as well. However, there are an infinite number of dependent axioms, so no finite set of dependent axioms will cover all of the truths of A. 3 The Halting Problem The most practical outgrowth of this line of reasoning is in Turing’s Halting Problem. Turing showed that, given a program p in a general-purpose programming language (where p includes both the program and all its input), there is no general, finite program h such that providing h with p as a parameter will tell if p will ever complete (Turing, 1936). More specifically, the problem comes in the fact that, while we can tell if p does complete, we can never tell if it won’t complete. The reason is simple—if we run p a step at a time, at any given moment, p may complete. At that point, p will be known to complete. However, we won’t know if p doesn’t complete until we have waited an infinitely long time, and it hasn’t happened. Therefore, while knowledge of completion takes a finite amount of time, knowledge of non-completion takes an infinite amount of time. If p will not ever complete, knowledge of that fact can be considered a dependent axiom. Knowledge of whether a particular p or class of ps will not ever complete can be added to h as dependent axioms. However, because there are an infinite number of dependent axioms, we can never add all of the axioms to h and have h remain a finite pro- gram. 4 Example Axioms Let us presume that there is a machine language such that each instruction is one byte long, and the hexadecimal code CC means “jump relative,” which looks at the next byte and modifies the program counter based on that signed value. Therefore, if the byte sequence CC-00 occurs, the machine will go into an infinite loop, because it will jump relative to the current instruction, but the amount of the jump will be 0, because it doesn’t go anywhere. 26 Two Methods of Calculating Axiom Size There are several ways of constructing an axiom which cov- ers this fact. Non-exhaustively, these include: 1. If the entirety of the program contains exactly the code CC-00 it will not complete. 2. If the program contains the code CC-00 within a line of code that will provably execute (i.e., contains no prior jumps or some similar criterion), then the program will not complete. 3. If the program contains the code CC-XX, where X X is a value less than or equal to zero which moves the program to a place that provably returns back to the same line of code, then the program will not complete. Let us note several things about these di!erent dependent axioms. First of all, note that all of the proposed axioms include the extraordinarily simple program CC-00. Thus, there is definitely overlap among the axioms. Also note that while Axiom 1 is specifically for a single program, Ax- iom 2 and Axiom 3 both match multiple programs. In fact, Axiom 3 matches a much larger number of programs than Axiom 2. Also of interest, in these examples, the longer the descrip- tion of the item, the more potential programs are matched. This is not necessarily true (i.e., there could be a shorter program that captures more potential programs), but it is true that more length allows for more coverage. That is, if we had a shorter program with more coverage, we could get yet even more coverage by adding bits to our program. 5 Measuring the Axioms I have identified two potential ways of measuring the “size” of each of these axioms, though there are probably yet oth- ers. 5.1 Measuring Occurrence Probabilities The first way is to measure the probability of occurrence of the axiom. That is, what is the likelihood of a given arbitrary program that contains the axiom occurring? This probability could be converted into bits in order to give a size of an axiom. To see how this would work, look at Axiom 1. Imagine that programs were encoded as prefix-free codes.1 As a simplifi- cation, the prefix-free code will be achieved by saying first N bits are a tally of the number of bytes in the code, ter- minated by a single 0 bit. Therefore, the prefix-free ver- sion of the program identified by Axiom 1 would be 0b110 CC-00.2 The probability of this program would be 1 524288 , or 19 bits. Thus, the size of Axiom 1 using this metric is 19 bits. The size of the other two axioms would require significantly more investigation (and specification as to the programming system) to determine. However, for an axiom that matches a finite number of programs, the size of the axiom using this method is relatively easy to determine. 5.2 Measuring the Size of the Axiom Identifier Another method for size measurement is to identify the shortest possible program which is able to identify the given axiom in code. Using this method, Axiom 1 could be iden- tified by the pseudo-code, “if the size of the program is two bytes, and the first byte is CC and the second byte is 00, then return true, otherwise return false.“ Obviously, we would need to spell out the rest of the language in order to determine the size of this pseudo-code, but the procedure is straightforward. 6 Implications The most important result here is that dependent axioms are in fact measurable. Whether these are the most impor- tant measures of axioms is only partially relevant—the idea that they can be measured has been validated. Nonetheless, it seems that these measurements do have some importance. Measuring the occurrence probability measures the likelihood that the axiom will come into ef- fect in any given program, thus whether or not someone will likely be thinking about the axiom. Measuring the size of the detector has even more interesting applications. It was noted in Bartlett (2012) that some axioms require other axioms to be known before discovery. Measuring detector size fits in well with this idea because, given an existing detector A, and a goal detector B, there may exist some mutual information between A and B which 1This is a requirement so that the individual programs can be represented as probabilities (Kraft, 1949). 2Note that the first part of the program is not necessarily fall on an even byte mark, so it is just given as the list of bits, while the rest of the code is given in hexadecimal. Volume 2, Issue 1 REFERENCES 27 makes the size of B smaller given a detector for A. The proposed equation for generating axioms using insight in Bartlett (2012) was A = I (Q, p, i, B) (1) where I is the “insight” function of cognition, Q is the deci- sion problem you are trying to solve (i.e., the halting prob- lem), p is the program with i as the input, and B is a set of axioms. The point of the function is to say that, given a su"ciently inclusive B, the remaining axiom(s) A needed to decide Q for p and i can be found. Using the detector size metric, we might posit that the critical feature is the size of A given B. That is, insight allows an amount of information of a detector to be built. Given a specifically inclusive B, the amount of mutual information between B and A can reduce the number of bits required to build the detector. This provides some amount of directionality for insight problems, while still allowing for them to be found in various ways. References Bartlett, J (2012). “Using Turing Oracles in Cognitive Mod- els of Problem-Solving”. In: Engineering and the Ulti- mate: An Interdisciplinary Investigation of Order and Design in Nature and Craft. Ed. by J Bartlett, D Halsmer, and M R Hall. Broken Arrow, OK: Blyth In- stitute Press, pp. 99–122. Kraft, L G (1949). “A Device for Quantizing, Grouping, and Coding Amplitude-modulated Pulses”. PhD thesis. Massachusetts Institute of Technology. url: http:// hdl.handle.net/1721.1/12390. Raatikainen, P (2018). “Gödel’s Incompleteness Theo- rems”. In: The Stanford Encyclopedia of Philoso- phy. Ed. by E N Zalta. url: https : / / plato . stanford . edu / archives / fall2018 / entries / goedel-incompleteness/. Turing, A M (1936). “On Computable Numbers, with Ap- plication to the Entscheidungsproblem”. In: Proceedings of the London Mathematical Society 42, pp. 230–265. Volume 2, Issue 1 Letters and Notes Divergent series and its assigned value in a hyperreal context Bas van der Reijden DOI: 10.33014/issn.2640-5652.2.1.reijden.1 Abstract This letter discusses the deep connection between the infi- nite sum of natural numbers and the value " 1 12 . Aside of more widely known facts, we consider a nontrivial way in which we show the veracity of this connection; more pre- cisely this concerns the BGN method (Bartlett, Gaastra, and Nemati, 2020) applied on the so-called damped oscil- lated Abel summed variant of the series. Moreover, we have found a generalization of this method which ‘correctly’ as- signs finite values to other divergent series. We conclude with some questions concerning whether and how we can analytically relate our hyperreal terms to frame the method in a more justifiable and applicable context. It is obvious that the sum of natural numbers 1+2+3+ · · · tends to infinity and can thus not be equal to " 1 12 . There does however exist some connection between this series and value and it is highly probable that this connection is implicitely used (i.e. ‘under the hood’) in e.g. physics (which often turns out to be perfectly justifiable, as can be shown by various experiments). The first evidence of this connection is retrieved when one considers the Riemann zeta function ! (s). It is known that ! (s) is equal to " 1 12 when s = "1 and it is interesting that one retrieves the sum of natural numbers when one ‘plugs in’ s = "1 at the defining series of the Riemann zeta function !# k=1 1 k s . Plugging s = "1 in the above series is unfortunately not justifiable (given that ! (s) is only equal to this series when $(s) > 1) but it remains an interesting thing to mention. Another evidence of the connection can be revealed when one considers the ‘smoothed version’ of the partial sums !N k=1 k, see also Tao, 2010. It turns out that these smoothed partial sums have the same behaviour as the Figure 1: Smoothed partial sums !N k=1 k with a y- intercept of " 1 12 . regular partial sums (i.e. they have the same asymptotic expansion) and thus tend to infinity when N % #. However, one probably recognizes the constant value in its asymptotic expansion, which is (according to Tao (2010)) given by CN2 " 1 12 +O( 1 N ) (with C some coe"cient of little importance in this case). Moreover, it is given that " 1 12 is attained when one looks at the intersection with the y-axis in Figure 1. The first two evidences we mentioned are quite widely known but a more unknown fact can be observed when one considers a damped oscillating variant of the series 1 + 2 + 3 + · · · , namely #" k=1 ke"k! cos(k" ). (1) This variant was also discussed in a previous letter (Bartlett and Khurshudyan, 2019). In this letter it was also mentioned that, in the context of hyperreal numbers by introducing # := # (i.e. by appying the BGN method on it), (1) can be written in a closed-form expression (or at 30 Letters and Notes least as an approximation of it). It remained however still unclear which value/magnitude the infinitesemal quantity " must be1 in order that the BGN method applied on (1) equals the ‘appropriate’ value " 1 12 ; only numerical evidence was given. In particular, it was shown that if " = 1 " , the computing software “Wolfram Mathematica” will include the constant " 1 12 in its BGN expression (which is similar to the observation of the previous paragraph). At the time that Bartlett and Khurshudyan (2019) was written, it only seemed clear that there is a numer- ical evidence that (1) equals " 1 12 when we take " in a su"ciently small interval. Recently, we have found that Sugiyama (2014) (Section 2.3) provides a more theoretical derivation of this matter. Although the website and its choice of words are somewhat vague and confusing, the derivation seems correct. In this derivation there is being made use of a so-called ‘damped oscillated Abel summation method’, which is a kind of generalization of the more common Abel summation method used to assign finite values to divergent series. In this article, this method of ‘damped oscillated Abel summation’ is consequently used on a larger class of divergent series as well; furthermore it turns out that the ‘damping’ and ’vibrating’ constant should not be necessarily equal to each other. We thus in fact have that (see also Section 5.2 and Section 6.1 of Sugiyama (2014), we here write " instead of x) !# k=1 k i can be transformed to (letting i & 1 be an integer) "" k=1 k ie"k! cot ! 2i+2 cos(k" ) and "" k=1 k ie"k i+1 2 ! cos(k i+1 2 " ) and consequently taking the limit " % 0 yields the ‘ap- propriate’ assigned value; we also numerically verified this2. It remains of course interesting how this damped os- cillated Abel summation method can be stated in our more ‘detailed’ hyperreal context; i.e. in which we know the exact values of " (possibly in terms of #) in order that the BGN method assigns the ‘appropriate’ value to a divegent series. Unless it is still untrivial which values " must have in order that the mentioned method assigns this value, we can however say from Equation (5.57) in Sugiyama (2014) that in general 1 ! must be a lot smaller than # (this was also shown by numerical experiments: if we set " = 0.01, # must be a lot larger than 100). In conclusion, we can thus say that the connection 1in relation to " 2by again letting ! be in a su!ciently small interval between the often assigned value of a divergent series is hidden in its asymptotic expansion. Furthermore, some slight variations (performed in the context of hyperreals) of the terms in the divergent series will alterate its asymptotic expansion in such a manner that that the BGN method assigns the ‘desired value’ to it. As it is at this point still untrivial when equality holds, and how in this case " and # thus must be related, remains an interesting topic for further research. To state this in a more general and mathematically way: Consider a divergent series with BGN expansion A(" (#))#2 +C +O(1/#) (here A is a value dependent of " which is in turn dependent of # and C is the ‘appropriate’ value we want to have), the question is now which variations (in terms of " (#)) we have to make in order to make A(" (#))#2 equal to zero. Bartlett, J, L Gaastra, and D Nemati (2020). “Hyperreal Numbers for Infinite Divergent Series”. In: Communi- cations of the Blyth Institute 2.1, pp. 7–16. Bartlett, J and A Khurshudyan (2019). “Numberphile’s Proof for the Sum 1+2+3+...” In: Communications of the Blyth Institute 1.1, pp. 54–55. Sugiyama, K (2014). New proof that the sum of natural numbers is -1/12 of the zeta function. url: https : / / xseek - qm . net / Regularization _ e . htm # _Toc524947400 (visited on 11/16/2019). Tao, T (2010). The Euler-Maclaurin formula, Bernoulli numbers, the zeta function, and real-variable analytic continuation. url: https://terrytao.wordpress. com/2010/04/10/the-euler-maclaurin-formula- bernoulli - numbers - the - zeta - function - and - real-variable-analytic-continuation/ (visited on 01/07/2019). Do Mutation Rates Match the Kelly Criterion? Jonathan Bartlett DOI: 10.33014/issn.2640-5652.2.1.bartlett.2 The Kelly Criterion defines an optimal betting strategy for games that have a defined risk and payo!. It was developed by John Kelly, Jr. at Bell Labs (Kelly, 1956). Given a bet with a probability of success P and a payout of B, the Kelly Volume 2, Issue 1 On Logic of Being and Wigner’s astonishment regarding the applicability of Mathematics 31 criterion tells you the size of your bet compared to your overall bankroll. The Kelly criterion is given as PB + P " 1 B (1) This formula can be derived from a formula for an expected total payo! of the bet given by the equation T = A(1 + B f )NP (1 " f )N (1"P) (2) where T is the total winnings, A is your starting amount, N is the number of trials, and f is the bet size. Optimizing for T yields Equation 1. Many researchers have discussed the concept of mutations in populations as “bet hedging.” (Philippi and Seger, 1989; Bartlett, 2008; Simons, 2011; Grimbergen et al., 2015) Since the Kelly criterion allows one to at least theoreti- cally calculate the optimum bet size for each configuration, it might be possible to calculate various optimum mutation rates at di!erent sites and compare them to their optimal size according to the Kelly criterion, or an adjusted version of it. Most analysis of bet hedging has merely checked to see if the hedging strategy is empirically beneficial (Childs, Met- calf, and Rees, 2010; Simons, 2011) or potentially evolvable (King and Masel, 2007). Applying the Kelly criterion may be able to help determine how optimal organisms’ various bet hedging strategies are. One possible experimental approach would be to provide organisms with a long-term, continually-varying environ- ment. After many generations, it would be interesting to check if the mutation rates for adaptive switching between environments had any relation to the theoretical considera- tions of the Kelly criterion, or any other theoretical hedging system. Bartlett, J L (2008). “Statistical and Philosophical Notions of Randomness in Creation Biology”. In: Creation Re- search Society Quarterly 45, pp. 91–99. Childs, D Z, C J E Metcalf, and M Rees (2010). “Evolution- ary bet-hedging in the real world: empirical evidence and challenges revealed by plants”. In: Proceedings fo the Royal Society B: Biological Sciences 277, pp. 3055– 3064. doi: 10.1098/rspb.2010.0707. Grimbergen, A J et al. (2015). “Microbial bet-hedging: the power of being di!erent”. In: Current Opinion in Mi- crobiology 25, pp. 67–72. Kelly, J L (1956). “A New Interpretation of Informa- tion Rate”. In: The Bell System Technical Journal 35, pp. 917–926. url: http://www.herrold.com/ brokerage/kelly.pdf. King, O D and J Masel (2007). “The evolution of bet- hedging adaptations to rare scenarios”. In: Theoretical Population Biology 72, pp. 560–575. doi: 10.1016/j. tpb.2007.08.006. Philippi, T and J Seger (1989). “Hedging One’s Evolution- ary Bets, Revisited”. In: Trends in Ecology and Evolu- tion 4.2, pp. 41–44. Simons, A M (2011). “Modes of response to environmental change and the elusive empirical evidence for bet hedg- ing”. In: Proceedings of the Royal Society B: Biological Sciences 278.1712, pp. 1601–1609. doi: 10.1098/rspb. 2011.0176. url: https://royalsocietypublishing. org/doi/10.1098/rspb.2011.0176. On Logic of Being and Wigner’s astonishment regarding the applicability of Mathematics Gordon E Mullings DOI: 10.33014/issn.2640-5652.2.1.mullings.1 The Nobel Prize winning Physicist, Eugene Wigner, fa- mously posed a powerful challenge (1960) by asking why is mathematics so e!ective, especially in the physical sci- ences (Wigner, 1960). It is possible that the reason for the e!ectiveness of mathematics is not because mathematics is in any way causative, but instead because mathematics studies the structure of logical possibility and constraint. When plugged into a possible world, mathematics gives us the tools to analyze the logically possible outcomes. There- fore, when a possible world that is expressed mathemati- cally su"ciently aligns with reality, mathematics becomes e!ective at expressing relationships and outcomes. For example, beings (as well as possible beings and things impossible of being) can be understood in the context of possible worlds. A “possible world” is a su"ciently complete description of possible states of a!airs described through chains of propositions. We may observe that things impos- sible of being, such as a square circle, have in them mutu- ally inconsistent required core attributes; they cannot be realised in any possible world. Possible beings would exist in at least one possible world were it actualised. For instance, a contingent being B that depends on C might exist in a world W and not in a closely neighbouring one W ! 32 Letters and Notes if C is present in W but not W !; C thus being an enabling, necessary causal factor for B. By contrast, a necessary be- ing F will exist in all possible worlds, showing itself to be a framework element for such a world. A key insight is that for any world W to be distinct from W ! it requires some factor A in W that is absent in W !. We may then partition the factors of W as W = {A|¬A}. After parti- tioning, we will have two distinct groups—the factor A and all of the factors which are not A. The null set corresponds to zero. Each particular set in the partition can be counted as the number one, and the combination of both partitions (even in a single world where A is an empty set) is two. Thus, for any particular possible world W , the quantities 0, 1, 2 are necessarily present. Taking the von Neumann construction, immediately we find N, thence (using addi- tive inverses) Z, so also (taking ratios) Q and (summing convergent power series) R; where Z provides unit-stepped mileposts in R. That is, a structured core of quantities will be present in any W , and we may regard mathematics as the study of the logic of structure and quantity. Extensions to the hyperreals R' follow by construction of some H that has as reciprocal h = 1 H closer to 0 than 1 n for any n in N. Therefore, relationships and linked operations across such quantities will also be present, or may be constructed as needed. Illustrating, after Abraham Robinson (Robinson, 1966), hyperreals allow calculus to be treated as extensions of algebra in R'. Thus, while bare distinct identity and coherence focused on quantities will not cause things by the inherent potential or action of such entities, they instead are logical constraints on being and are tied to what can or must be or cannot be or happens not to be. So, too, we may see that the abstract logic model worlds that we may construct then lead to key entities that if necessary are framework to any possible world; thus applicable to our common world. By contrast, if certain quantities and relationships are merely part of the contingencies of some W !! that is close enough to our own, they may provide adequate analogies for mod- elling. As a result, we have good reason to expect that mathe- matical reasoning and core entities will in many cases be highly relevant to and have powerful predictive power for our common world. Robinson, Abraham (1966). Non-Standard Analysis. Ams- terdam: North-Holland Publishing Company. Wigner, Eugene (1960). “The Unreasonable E!ectiveness of Mathematics in the Natural Sciences”. In: Communica- tions in Pure and Applied Mathematics 13.1, pp. 1–14. doi: 10.1002/cpa.3160130102. Independence Conservation and Evolutionary Algorithms Eric Holloway DOI: 10.33014/issn.2640-5652.2.1.holloway.2 Levin’s Law of Independence Conservation Leonid Levin’s 1984 article (Levin, 1984) is the first to this author’s knowledge to prove a fully stochastic conservation of information law. Levin titled his law ’independence con- servation’ which he considered fairly obvious, describing it as “Torturing an uninformed witness cannot give informa- tion about the crime!” Levin’s law is not well known, which is unfortunate since the more commonly known conservation laws are focused either only on the random or deterministic case. Levin’s law is remarkable because it unifies both the random and deterministic cases, showing that the combination also can- not result in information increase. The second remarkable thing about his law is how easy it is to prove, given some preliminaries about algorithmic information. Algorithmic Information Theory Background First is required the notion of algorithmic information, which is defined on bitstrings. Algorithmic information is the length of the shortest program that generates a partic- ular bitstring. K (x) := min y |U (y)=x |y |. (1) The shortest program is itself known as the elegant program for that particular bitstring. Each bitstring has a unique elegant program. Volume 2, Issue 1 Independence Conservation and Evolutionary Algorithms 33 y' := arg min y |U (y)=x |y |. (2) All programs either terminate after a fixed amount of time, or never terminate. All the programs in question are known as prefix free, which means that no terminating program begins another terminating program. Algorithmic mutual information is the length of this pro- gram if we are also provided another bitstring as input, subtracted from the length if we are not provided the extra input bitstring. I (x : y) := K (y) " K (y |x). (3) Unfortunately, this basic definition of algorithmic mutual information is only symmetric under a logarithmic error, because we have to mark where one bitstring starts and the other ends. This requires a number of bits logarithmic on the size of the shortest bitstring, which is x in this case. I (x : y) " I (y : x) = O(log(x)). (4) We can improve the definition of algorithmic mutual infor- mation to be completely symmetric under a constant that is independent of the bitstrings we are looking at, which in other words means we don’t have to worry about the con- stant and the algorithmic mutual information is symmetric as far as we are concerned. This improvement is to use the elegant program of the input bitstring instead of the bitstring itself. I'(x : y) := K (y) " K (y |x'). (5) Since the elegant program halts once it has generated the input bistring, we know we can start on the next bitstring, so we avoid having to encode the bitstring length. This saves us from having to use the logarithimic term. I'(x : y) " I'(y : x) = O(1). (6) One final fascinating point on algorithmic information is that we can also use it to create a universal distribution. m(x) := 2"K (x) . (7) “Universal” means is that we have a distribution that pro- vides the highest probability for every bitstring possible, within a multiplicative constant, under the assumption that we are only dealing with computable generating sources for the bitstrings. m(x) & p(x) 'O(1). (8) The computable distribution assumption is a reasonable as- sumption for dealing with physical phenomena, since (as far as we know) everything physical can be modeled to theo- retically perfect accuracy with enough computational re- sources. Proving Levin’s Deterministic Law Alright, so now onto proving Levin’s law. We first start with a simple lemma, that providing more information can only decrease conditional algorithmic in- formation. In other words, the more we know about y, the less information we need to describe y. K (y |x) & K (y |x, z) +O(1). (9) We now introduce another simple lemma that with a pro- gram to generate x, namely z which we execute with Turing machine U to generate x, x = U (z), (10) we can generate both x and z. Thus, the joint information is the same between z and x, z. K (z) = K (x, z) +O(1). (11) This also means to generate the triple {y, x, z} we only need y and z. K (y, x, z) = K (y, z) +O(1). (12) 34 Letters and Notes Since it is the case that K (y |x, z) = K (y, x, z) " K (x, z). (13) Then performing replacements to Equation 13 with Equa- tions 11 and 12, we get K (y |x, z) = K (y, z) " K (z) (14) = K (y |z). (15) Combining Equations 9 and 14 shows us that x can never tell us more about y than z. K (y |x) & K (y |x, z) (16) = K (y |z). (17) We can then use Equation 16 to show running a program f on i does not increase mutual information with y. The notation U ( f .i) to mean we’ve run program f with input i using a universal Turing machine U. First we decompose the mutual information. I'( f (i) : y) = I'(U ( f .i) : y) (18) = K (y) " K (y |U ( f .i)). (19) (20) Now, we set z = f .i and x = U (z) = U ( f .i), and then apply Equation 16 to Equation 18. I'( f (i) : y) = I'(x : y) (21) = K (y) " K (y |x) (22) ( K (y) " K (y |x, z) (23) = K (y) " K (y |z) (24) = I'(z : y) (25) = I'( f , i : y). (26) Giving the final concise result, I'( f (i) : y) ( I'( f , i : y). (27) This Equation 27 states that executing function f on in- put i does not produce any more information about y than the function and input before they are executed. In other words, running a program doesn’t produce any information. Proving Levin’s Random Law Now with the deterministic version out of the way, we can move onto the random version. The random version asks, what if we generate f randomly, could that result in an information gain? This question is based on the fact that generating f randomly will result in an f with a lot of algorithmic information, since it will be incompressible. So, even though running U ( f .i) doesn’t give us anything new, the initial selection of f may start us o! with a good amount of information about y. Levin’s second step in proving the random law shows this intuition is false. To prove the second law, we will rely on the dominance property of the universal distribution in Equation 8. " f p( f )I'( f , i : y) = " f p( f ) log2 m( f , i |y) m( f , i) (28) ( " f p( f ) log2 m( f , i |y) m( f )m(i) (29) ( log2 " f p( f )m( f , i |y) m( f )m(i) (30) ( log2 ! f m( f , i |y) m(i) (31) = log2 m(i |y) m(i) (32) = K (i) " K (i |y) (33) = I'(i : y) (34) Which gives us the result that randomly generating an f is not expected to provide any information about y. " f p( f )I'( f , i : y) ( I'(i : y). (35) To wrap up the independency conservation law, we apply Equation 27 to Equation 35. " f p( f )I'( f (i) : y) ( " f p( f )I'( f , i : y) ( I'(i : y). (36) Volume 2, Issue 1 CrowdRank: A Simple Ranking Algorithm for Crowdsourced Rating Systems with Uneven Participation 35 Evolutionary Algorithms What sort of impact, if any, does the law of independency conservation in Equation 36 have on evolutionary algo- rithms? One of the simplest implications is that if there is a target area independently designated by y, then it is not possi- ble to randomly evolve population i towards y without any fitness information. However, there are also implications even if there is fitness information provided for the evolution of i. Let’s say that f represents one round of an evolutionary algorithm applied to i, which consists of the following steps: 1. crossover population 2. vary population 3. select population i is defined as a population of bitstrings selected at random. We then define y as the set of bitstrings rated at a certain level of fitness. Part of f stays constant, and another part is varied ran- domly, such as mutation and which bitstring sections are crossed. We denote the random part as r. This gives us a surprising result. " r p(r)I'( f (i, r) : y) ( I'( f , i : y). (37) Equation 37 states that each round of randomized evolution provides no further information about the fitness region y than already existed in the initial conditions of the original population i and the evolutionary algorithm f . Thus, this equation proves that evolutionary algorithms cannot gen- erate algorithmic mutual information, even in regards to regions specified according to fitness. Levin, Leonid A (1984). “Randomness conservation inequal- ities; information and independence in mathematical theories”. In: Information and Control 61.1, pp. 15–37. CrowdRank: A Simple Ranking Algorithm for Crowdsourced Rating Systems with Uneven Participation Jonathan Bartlett DOI: 10.33014/issn.2640-5652.2.1.bartlett.3 Introduction Public rating systems are di"cult to score well. Voting systems tend to simply favor what is already popular. Av- eraging systems tend to have significant variance if there are not enough people scoring. For instance, let’s say that I run a songwriting contest and have 100 entries. I then put it out to a public vote on the Internet to see who wins. Most people are not going to lis- ten to all 100 songs. If I do a simple “thumbs up” approach and count how many votes a song has, then whichever song- writer has the best existing following will simply tell their fans to vote for them, and it will simply devolve into a popularity contest. Let’s say instead I do a rating system where you can rate a song between 0 and 100. Now, songs by popular artists will actually be negatively weighted because they will have more visibility for negative ratings. It is not hard for a few votes to be all 100s, but it is hard for a thousand votes to be that way. Thus, those who have fewer ratings have an advantage. The goal, then, is to come up with a fair way of handling public ratings which takes into account both the average score that people assign and the relative certainty that we have that the score is representative of the “true” score. The Model This problem actually becomes rather easy once an appro- priate mental model is devised. Assuming a normal distri- bution of actual scores that come in around a “true” value for a particular score for an entry, what is the range of possible score values based on the scores that have been submitted so far? 36 Letters and Notes Take a concrete example. Let’s say that Song A has 12 votes with an average score of 60. What is the range that the “real” score should lie in? The main open question when dealing with statistics is what confidence level we want to deal with. For this example, let’s say that we want to main- tain a 95% confidence interval. That means that we want to know what the range is of two standard deviations from the mean. With only 12 samples, this leads to a fairly wide interval, with the real score being between 32 and 88. However, as we add more samples, this range narrows in to the aver- age. If we have 24 samples and maintain the same average, then our range is restricted to between 40 and 80. At 144 samples, the range narrows to 52–68. So, with a few scores, the possible “real” score has a very wide range. However, as more and more scores come in, the range narrows further and further. Now, even though these rankings get tighter variances with more scores, the average value for the scores remain what they were. So how do we convert this into a more legitimate ranking system than we had before? What we can do is simply rank the songs using their lowest possible scores according to the chosen confidence interval. That is, we have established statistically what the lower bound for their score is. Therefore, we can definitively give them that score because we know they have earned at least that score. This minimal defensible score will be called the CrowdRank score. Let’s say that Song A has 144 rankings that average to 60, and Song B has 25 rankings that average to 70. Which song should be ranked higher? As we have already noted, Song A’s “real” score has a potential range of 52–68. Song B, because it has fewer score submissions, has a wider potential range of 50–90. Since the lowest defensible score of Song A is 52, and the lowest defensible score of Song B is 50, that means that Song A will be ranked higher than Song B. The actual ranking will be dependent on the confidence level that is chosen for the rankings. The higher confidence levels will take many more rankings for the scores to ap- proach their averages. The Calculation The calculation of each entry’s score is fairly straightfor- ward. It is basically the inverse of standard statistical scores. p The population size n The number of samples (i.e., number of rankings on a particular entry) z The confidence level desired, expressed as a z-value (the number of standard deviations that a given confidence level uses—2.58 for 99% confidence, 1.96 for 95% con- fidence, etc.) e The margin of error for the confidence interval, expressed as a decimal (i.e., 0.25 for ±25%) s The average score of the samples expressed as a real num- ber between 0 and 1. In the present example we would divide all scores by 100. m The expected value. Choosing 0.5 is a “most-safe” value. Typically, the number of needed samples is determined from the desired margin of error, using n = z2m(1 " m) e2 . (1) Rearranging to find the margin of error from the sample size, we find e = # z2m(1 " m) n . (2) Since our results are distributed as a percentage anyway (a score of zero to one), the crowdrank is just the score s " e. Simplified using m = 0.5, the CrowdRank calculation for a particular entry is CrowdRank = s " # 0.25 z2 n . (3) If the samples are taken from a restricted population of size P (say, all the members of a club), you can get an even better measurement from the following: CrowdRank = s " # 0.25 z2 n P " n P " 1 (4) Di!culties There are two primary di"culties with this system. The first is that, if there are too few rankings for each entry, the confidence level will fall o! to zero. This can be mitigated by varying the desired confidence level based on the average rankings per entry. Volume 2, Issue 1 CrowdRank: A Simple Ranking Algorithm for Crowdsourced Rating Systems with Uneven Participation 37 The other di"culty is in communicating the results to end- users. It is di"cult for them to understand why having 144 people all giving a ranking of 60 might translate to a Crow- dRank of 52. Having scores whose origin is not transparent can lead to a lack of confidence in the system. However, be- cause the discount to the scores is fixed for the number of entries, you can communicate this as the number of points that are discounted for a given number of entries. For in- stance, if you are using the 95% confidence interval, then you can post that receiving 23–25 entries will result in a 20 percentage point discount. Conclusion This note introduced a system of averaging crowdsourced rankings that appropriately discounts ranking averages based on the number of submissions. This can be used in any place where a variable number of crowdsourced rank- ings might be received. It removes the “popularity contest” problem of simple voting, as well as the problem of having too few rankings available in a generic averaging system. Volume 2, Issue 1 News Eric Holloway Publishes Paper on Specified Complexity David Nemati and Blyth Institute fellow Eric Holloway recently published a new paper in the journal BIO- Complexity (Nemati and Holloway, 2019). The paper, titled “Expected Algorithmic Specified Complexity,” explores the expected Algorithmic Specified Complexity (ASC) of a ran- dom variable, concluding that the expected ASC is always less than zero. This is true both of random variables and of “processed” random variables—variables which have under- gone some sort of transformation. This indicates that the existence of positive ASC always counts as “surprise,” and therefore always requires explanation. New Thinking on Human Origins This past year witnessed a plethora of new thinking on hu- man origins. First up, new research by Ola Hössjer and Ann Gauger recently showed a unique model for looking at biodiversity (Hössjer and Gauger, 2019). Essentially, they looked at the ways in which minor allele frequency distri- butions can arise, and how long di!erent distributions take to arise. New alleles essentially start out as one-o! events, arising as a mutation in a single individual and later spread- ing through the population or dying out (or somewhere in- between). The frequency of these minor alleles can pro- vide some amount of information about the history of the species in question. It takes time for an allele to spread through a population. Therefore, initially, from an initially homozygous pair, the low-frequency side of the minor allele frequency spectrum will contain all of the new mutations. The mutations will take time to spread throughout the pop- ulation. What Hössjer and Gauger discovered, though, is that a heterozygous initial pair creates an allele frequency spectrum that looks much older than it is. This is because an allele can actually start as either 25%, 50%, 75%, or 100% of the population before any mutations even occur. Thus, the frequency spectrum will fill very quickly from the initial pair, and the initial heterozygosity will look equiva- lent to ancient mutations. Using standard assumptions, Hössjer and Gauger calcu- lated that the current allele frequency spectrum could be attained from an initial starting pair in 100,000–500,000 years. Using other alternative assumptions about the na- ture of the starting pair could result in attaining the current allele frequency spectrum in an even shorter period of time. Another interesting paper was done by Nathaniel Jeanson and Ashley Holland, which analyzed the human Y chromo- some (Jeanson and Holland, 2019). Analyzing the human genome as a whole leads to a number of model-specific is- sues. If we imagine a starting pair for humanity, were the original chromosomes identical or did they house diversity? Additionally, might the first human female have eggs that had additional diversity in their chromosomes? However, nearly all considerations of single-couple human origins have exactly one starting Y chromosome, leading to fewer model-specific considerations. Jeanson and Holland aimed to improve the data available about Y chromosome mutation rates by examining pedigree-based studies which used high-coverage sequencing. According to Jeanson and Holland, (a) a Y chromosome molecular clock exists, and (b) it suggests a paternal history of the human race of about 4,500 years. Finally, the year ended with the publication of Joshua Swamidass’s new book titled The Genealogical Adam and Eve (Swamidass, 2019). This book aims to show that the practical di!erence between popular and scientific concep- tions of “Adam and Eve” are not too far o!. His viewpoint is that, although, according to the consensus view, there could be no single-couple origin of humans, there could be a single couple to whom all modern humans could trace their genealogies. Communicating Science Through New Venues New media becomes old media very quickly in the modern age. Email was once the best technological way to com- municate interpersonally. This is now often replaced with various social media platforms such as Facebook, Twitter, and others. Likewise, YouTube was once at the center of 40 News video-based media. While YouTube is still the de facto des- tination for video content, it is useful to look at some of the newer players. First of all, some video is being distributed through apps on standard platforms. Developing an app for Roku, Amazon Fire, or Apple TV is a new way to get content to users. Additionally, streaming services such as Twitch, and now Mixer, are becoming the dominant social video platforms for the youth. Science communicators should explore ways of reaching the public through these outlets. One possibil- ity would be to play video games with scientists, discussing what they do while shooting zombies, or discussing the lat- est theories of the universe while racing Nascar. Growth opportunities are always with the latest technology. Asserting a strong position while they are still unproven is the best way to establish leadership for the future. Austrian Society Zentrum für BioKomplexität & NaturTeleologie Opens with Special Symposium A new Austria-focused scientific society, Zentrum für BioKomplexität & NaturTeleologie, recently formed and held their first symposium. The symposium featured many European scientists and mathematicians, as well as a few from the United States as well. The organizers of the con- ference were Günter Bechly, a prominent German paleon- tologist who presently works as a senior scientist for the Biologic Institute, and Siegfried Scherer, a professor of mi- crobiology and chair of microbial ecology at the Technical University of Munich. While the society is based in Austria, the sympo- sium itself was largely in English. Many of the talks from the three day symposium are available on YouTube, at https://www.youtube.com/playlist?list= PLkaKqUjdyg2JHNqeWQHnVTXZ-37h0tJMb. Breaking the Weismann Barrier and Closing the Loop for Lamarckian Evolution in Multicellular Organisms A recent review paper in the Royal Society shows that the Weismann barrier is crumbling. In “The active role of sper- matozoa in transgenerational inheritance,” Sciamanna et al review the mounting evidence that there is a Lamarckian feedback loop in DNA inheritance (Sciamanna et al., 2019). The paper reviews evidence that mammalian somatic tis- sues release RNA-containing vesicles, and that these vesi- cles are then passed to epididymal spermatozoa. Addition- ally, epididymal spermatozoa are known to be able to inter- nalize foreign nucleic acids into their nuclei. This completes the communication channel between somatic cells and germ cells required for Lamarckian inheritance. This communication channel was original proposed in the 1990’s. Steele et al’s Lamarck’s Signature suggested that so- matic mutations might be passed to germ line cells through an RNA channel (Steele, Lindley, and Blanden, 1999). Af- ter decades of work across multiple groups, evidence is mounting that Steele’s hypothesis was largely correct. Overcoming Entrenched Dogma About Pseudogenes Biologists have long known that not all non-coding DNA is junk DNA. However, pseudogenes have long been consid- ered the standard bearer for the junk DNA concept. Since pseudogenes look like defective, non-coding copies of ordi- nary genes, it has been often assumed that these are evolu- tionary leftovers—genes which once coded for something, or a copy of a gene that once coded for something, but which some accident of mutation incapacitated its activity. While the evidence for the activity of pseudogenes has long been known, the idea that pseudogenes indicate junk DNA has been ingrained in biologists. A group of researchers re- cently published a paper pointing out that the prejudices that biologists have about the status of pseudogenes as junk DNA is impeding the progress of understanding the way that the genome functions (Cheetham, Faulkner, and Dinger, 2019). References Cheetham, S W, G J Faulkner, and M E Dinger (2019). “Overcoming challenges and dogmas to understand the functions of pseudogenes”. In: Nature Reviews Genetics. doi: 10.1038/s41576-019-0196-1. Volume 2, Issue 1 REFERENCES 41 Hössjer, O and A Gauger (2019). “A Single-Couple Human Origin is Possible”. In: 2019 (1), pp. 1–20. Jeanson, N T and A D Holland (2019). “Evidence for a Human Y Chromosome Molecular Clock”. In: Answers Research Journal 12, pp. 393–404. Nemati, D and E Holloway (2019). “Expected Algorithmic Specified Complexity”. In: BIO-Complexity 2019 (2), pp. 1–10. Sciamanna, I et al. (2019). “The active role of spermatozoa in transgenerational inheritance”. In: Proceedings of the Royal Society B: Biological Sciences 286.20191263. Steele, E J, R A Lindley, and R V Blanden (1999). Lamarck’s Signature: How Retrogenes are Changing Darwin’s Natural Selection Paradigm. Basic Books. Swamidass, S J (2019). The Genealogical Adam and Eve: The Surprising Science of Universal Ancestry. IVP Academic.