Visual Recognition with a Large Scale Network of Dynamical Systems by Socrates Dimitriadis B.Sc. in Computer Science, University of Ioannina, Greece, 1999 M.Sc. in Computer Science, University of Crete, Greece, 2002 M.Sc. in Cognitive Science, Brown University, Providence RI, 2006 Thesis Submitted in partial fulfillment of the requirements for the Degree of Doctor of Philosophy in the Department of Cognitive, Linguistic and Psychological Sciences at Brown University PROVIDENCE, RHODE ISLAND MAY 2011 This thesis by Socrates Dimitriadis is accepted in its present form by the Department of Cognitive, Linguistic and Psychological Sciences as satisfying the thesis requirements for the degree of Doctor of Philosophy Date___________ __________________________________ James Anderson, PhD, Advisor Date___________ __________________________________ Fulvio Domini, PhD, Reader Date___________ __________________________________ Thomas Serre, PhD, Reader Approved by the Graduate Council Date___________ __________________________________ Peter M. Weber, PhD, Dean of the Graduate School ii AUTHORIZATION TO LEND AND REPRODUCE THE THESIS As the sole author of this thesis, I authorize Brown University to lend it to other institutions or individuals for the purpose of scholarly research. Date May 1, 2011 __________________________________ Socrates Dimitriadis, Author I further authorize Brown University to reproduce this thesis by photocopying or other means, in total or in part, at the request of other institutions or individuals for the purpose of scholarly research. Date May 1, 2011 __________________________________ Socrates Dimitriadis, Author iii Curriculum Vitae Personal Data Last name Dimitriadis First name Socrates Father name Pantelis Mother name Ekaterini Date of birth August 11, 1976 Nationality Greek Website www.socrates.name e-mail email@socrates.name Education 2003-2006 M.Sc. in Cognitive Science – Brown University, USA 2000-2002 M.Sc. in Computer Science – University of Crete, Greece GPA 9.28/10.0 Major in Computer Vision and Robotics Minor in Software Engineering and Information Systems 1995-1999 B.Sc. in Computer Science – University of Ioannina, Greece 1st in class '99, GPA 8.13/10.0 Foreign Languages English 1990, First Certificate in English 2002, T.O.E.F.L. (CBT 260) French 1992, Certificat de langue francaise German 1993, Zertifikat, Deutsch als fremdsprache Italian 1995, Diploma di lingua Italiana Athletics Long Distance and Marathon Runner (participation in more than 100 races) Bronze Medal in 1500m, National Student Championship 1998 Entry to six International Marathons (PR: 3.05.57) iv Awards and Distinctions 2007-2008 Brown University Corinna Borden Keen Research Fellowship 11/2006 OTEnet Prize in competition Innovation 2006 2004-2009 Brown University Teaching and Research Assistantships 2003-2004 Brown University Fellowship 2001-2003 ICS-FORTH Research Assistantship 2001-2003 University of Crete Teaching Assistantship 2000-2001 University of Crete Fellowship 1999 Hellenic Telecommunications Organization Scholarship (student with the highest GPA in the 8th semester of studies) 1998 Hellenic Telecommunications Organization Scholarship (student with the highest GPA in the 7th semester of studies) 1998 Greek State Scholarships Foundation (IKY) Prize 1998 Scholar of the Greek State Scholarships Foundation 1995-1999 Voyatzis Foundation Scholarship Research Projects 2008-2009 Brown University Laboratory for Engineering Man/Machine Systems Visual Change Detection in Satellite Images 2004-2010 Brown University Ersatz Brain Project Computational Cognitive Neuroscience Modeling 2001-2003 University of Crete & Institute of Computer Science – FORTH Content Based Image Retrieval 2001-2002 Institute of Computer Science – FORTH Computational Vision and Robotics Laboratory Development of a Robotic Soccer Platform for Robocup 2000 Institute of Cultural and Educational Technology, Xanthi Test and Evaluation of Teleconference-Teleducation Platforms 1998-1999 University of Ioannina EU Program EPEAEK Distance Learning Course for Neural Networks v Refereed Publications and Presentations Socrates Dimitriadis, “Visual processing as a large scale integration of associative neural networks”, 4th Computational Cognitive Neuroscience Conference, Boston MA, USA, November 2009 Socrates Dimitriadis, “Cortically-inspired parallel processing”, 13th SIAM conference on Parrallel Processing for Scientific Computing, Atlanta GA, USA, March 2008 Socrates Dimitriadis and Fulvio Domini, “Bayesian inference of visual depth based solely on primitive visual variables”, 30th European Conference on Visual Perception, Arezzo, Italy, August 2007 & Perception Vol.36 James A. Anderson, Paul Allopenna, Gerald S. Guralnik, David Scheinberg, John A. Santini, Socrates Dimitriadis, Benjamin B. Machta, and Brian T. Merritt, “Programming a Parallel Computer: The Ersatz Brain Project”, In Wlodzislaw Duch and Jacek Mandziuk (Eds.), Challenges for Computational Intelligence, Vol. 63, Springer: Berlin, 2007 Socrates Dimitriadis, Kostas Marias and Stelios Orphanoudakis, “A Multiagent Platform for Content Based Image Retrieval”, Journal of Multimedia Tools and Applications, Special Edition: Distributed Adaptation, Representation and Processing of Multimedia Information, Vol. 33, No. 1, 57-72, April 2007 Socrates Dimitriadis and James Anderson, “A network of networks model of pop-out phenomena in visual attention”, 2nd Computational Cognitive Neuroscience Conference, Houston TX, USA, November 2006 John Moustakas, Kostas Marias, Socrates Dimitriadis and Stelios Orphanoudakis, “A Two-level CBIR Platform with Application to Brain MRI Retrieval”, IEEE International Conference on Multimedia & Expo (ICME), Amsterdam, The Netherlands, July 2005 Socrates Dimitriadis, Kostas Marias and Stelios Orphanoudakis, “A Versatile Image Retrieval Platform Based on a Multiagent Architecture”, 6th International Conference on Visual Information Systems, Miami FL, USA, September 2003 Workshops Parallel Programming, Brown University, Providence, RI, October 18-19, 2004, Organized by Brown University's Technology Center for Advanced Scientific Computing and Visualization (TCASCV) Nonlinear Methods in Psychology, George Mason University, Fairfax, VA, October 24-25, 2003, Organized by National Science Foundation (NSF) vi Other Publications Socrates Dimitriadis et al., “2006 Review of Greek Information and Communication Technologies”, Hellenic Informatics Union, February 2007 Socrates Dimitriadis, “A Network of Networks Approach to Computational Modeling of Pop-out Phenomena in Visual Attention”, MSc. Thesis, Department of Cognitive and Linguistic Sciences, Brown University, May 2006 John Moustakas, Socrates Dimitriadis and Kostas Marias, “A Cognitive Architecture for Semantically Based Medical Image Retrieval”, ERCIM News No. 62, July 2005 Socrates Dimitriadis, Kostas Marias and Stelios Orphanoudakis, “Retrieval of Images based on Visual Content: A Biologically Inspired Multi-Agent Architecture”, ERCIM News No. 53, July 2003 Socrates Dimitriadis, “An Image Retrieval Platform Based on a Biologically Inspired Architecture”, MSc. Thesis, Department of Computer Science, University of Crete, November 2002 Socrates Dimitriadis, “Human-Computer Interaction with Gestures”, BSc. Thesis, Department of Computer Science, University of Ioannina, June 1999 Teaching Experience 2004-2009 Teaching Assistant Brown University, Department of Cognitive and Linguistic Sciences Courses: Quantitative Methods in Psychology (fall 2004), Neural Modeling Laboratory (spring 2005/2006/2007/2009), Visualizing Vision (fall 2005), Cognition (fall 2006) 2002-2003 Teacher of Computer Science 6th School of Technical Education of Heraklion, Greece Courses: Multimedia, Operating Systems-UNIX, Computer Applications 2000-2002 Teaching Assistant University of Crete, Department of Computer Science Courses: Computer Networks, Computational Vision, Intelligent Systems 2001 Teacher of Computer Science Vocational Training Institute of Kilkis, Greece Courses: Programming in C, Sound Processing and Synthesis vii Other Working Experience since 2006 Founder and owner of Greek-Movies.com, a search engine and directory for Greek videos 1999-2001 Hellenic Army, Reserve Officer of Communications 1994 Water Supply and Sewerage Company of Kilkis 1993-1994 Radio Producer in Akroama Radio Station and Euro Channel Memberships 2007-2009 Graduate student representative at the Computing Advisory Board of Brown University 2004-2005 Graduate student representative at the Board of the Department of Cognitive and Linguistic Sciences of Brown University since 2002 Member of the Hellenic Informatics Union 1998-1999 Undergraduate student representative at the Senate of the University of Ioannina 1997-1999 Undergraduate student representative at the Board of the Department of Computer Science of University of Ioannina viii Inspired by the words of an unknown poet When things go wrong, as they sometimes will When the road you're trudging seems all uphill When the funds are low and the debts are high And you want to smile, but you have to sigh When care is pressing you down a bit Rest, if you must, but don't you quit Life is queer with its twists and turns As every one of us sometimes learns And many a failure turns about When he might have won had he stuck it out Don't give up though the pace seems slow You may succeed with another blow Often the goal is nearer than It seems to a faint and faltering man Often the struggler has given up When he might have captured the victor's cup And he learned too late when the night slipped down How close he was to the golden crown Success is failure turned inside out The silver tint of the clouds of doubt And you never can tell how close you are It may be near when it seems so far So stick to the fight when you're hardest hit It's when things seem worst that you must not quit ix Acknowledgments First of all, I would like to thank my advisor, Jim Anderson, for believing in me and for giving me the opportunity to work my own way in this research. He has always been very supporting and infused me with confidence. A great thanks goes to the rest of the members of the Ersatz Brain Group and in particular to Paul Allopenna, John Santini, and Gerry Guralnik. The endless yet fruitful discussions we had in our weekly meeting all these years helped me shape many aspects of this thesis. I'm really grateful to the readers of my dissertation, Fulvio Domini and Thomas Serre, for the excellent comments and the valuable feedback they provided. I couldn't hope for a more expert reviewer than Thomas and I wish he had come earlier to our department. Fulvio Domini was a lot more than a reader, a true friend, and I will miss the running we did together. My research wouldn't be possible without the equipment and the technical support of the Center for Visualization and Computation and I'm grateful they let me use their systems free of charge for many years. I would also like to thank Joe Mundy for the research assistantship he offered to me in his lab for one semester. Last but not least, I would like to thank all my friends that helped me keep my sanity and my happy mood all these times – and there were plenty of them – that research didn't work the way I was expecting it to. My life in Providence wouldn't be the same without Dimitris' jocularity, Yannis' stresslessness, Babis' clumsiness, and Foteini's love. I will also never forget the fun we had with Christos, Panayotis, Vasilis, Olga, Nikos, Dimitra, Aris, Jeff, Arnold, Aggeliki, Adrian, Amit, Justin, Jae-Yung, Raj, Elena, Naomi, Giorgos, Maria, Danny, Vita, Katerina, Sophocles, Menia, Carmen, Qun, Paris, Stavroula, Kostas, Andrea, Sakis and Basilis, to name just a few! x Contents 1.Introduction..........................................................................................................1 2.Review of visual recognition models....................................................................5 2.1.Representation.............................................................................................6 2.2.Encoding.......................................................................................................8 2.3.Architecture.................................................................................................12 2.4.Processing flow...........................................................................................14 2.5.Learning......................................................................................................19 3.The proposed approach.....................................................................................23 3.1.Representation...........................................................................................24 3.2.Architecture.................................................................................................25 3.2.1.The basic processing unit...................................................................27 3.2.2.The organization of a layer..................................................................29 3.2.3.Multi-layer organization.......................................................................31 3.2.4.Connectivity map.................................................................................34 3.3.Stimulus encoding......................................................................................37 3.3.1.Geometry of the encoding...................................................................39 3.3.2.Content of the encoding......................................................................41 3.4.Learning......................................................................................................42 3.5.Dynamics....................................................................................................46 3.5.1.Alternative forms of dynamics.............................................................53 3.6.Discussion..................................................................................................55 3.6.1.Class of models...................................................................................55 3.6.2.A network of networks is not simply a large network..........................56 3.6.3.Why do we need a dynamical systems approach...............................57 3.6.4.Perception without action....................................................................59 4.Qualitative comparison with other approaches..................................................60 4.1.Specificity and variance..............................................................................60 4.2.Time............................................................................................................63 4.3.Learning......................................................................................................66 4.4.Parallel processing.....................................................................................68 4.5.Contribution to other approaches...............................................................69 xi 5.Experimental results..........................................................................................72 5.1.Methodology...............................................................................................72 5.2.Stimuli.........................................................................................................74 5.3.Visualization................................................................................................76 5.4.Study I – Sensitivity and specificity............................................................79 5.4.1.Experiment 1.......................................................................................80 5.4.2.Experiment 2.......................................................................................85 5.4.3.Experiment 3.......................................................................................92 5.4.4.Discussion...........................................................................................95 5.5.Study II – Affine transformations................................................................98 5.5.1.Experiment 4.....................................................................................100 5.5.2.Experiment 5.....................................................................................106 5.5.3.Discussion.........................................................................................111 5.6.Study III – Reaction time..........................................................................114 5.6.1.Experiment 6.....................................................................................114 5.6.2.Experiment 7.....................................................................................123 5.6.3.Experiment 8.....................................................................................129 5.6.4.Discussion.........................................................................................132 6.Conclusions.....................................................................................................135 7.Bibliography.....................................................................................................139 xii List of Figures Figure 1: Abstract view of an attractor network....................................................28 Figure 2: Energy landscape of an attractor network with four fixed points...........29 Figure 3: A network of attractor networks.............................................................30 Figure 4: Hierarchical organization of the hetero-associative network.................32 Figure 5: A processing unit over the respective receptive field............................38 Figure 6: Geometric characteristics of the stimulus encoding..............................39 Figure 7: Magnitude of 100 eigenvalues..............................................................51 Figure 8: Input stimuli and pre-processing...........................................................75 Figure 9: Experiment 1. Testing stimuli with 70% random ablations....................81 Figure 10: Experiment 1: Testing stimuli with 90% random ablations..................84 Figure 11: Experiment 2: Testing stimuli with ablated blocks...............................87 Figure 12: Experiment 2: Metastability in visual recognition................................90 Figure 13: Experiment 3: Stimulus specificity.......................................................93 Figure 14: Experiment 4: A random subset of the 1600 training stimuli.............100 Figure 15: Experiment 4: Position invariance with translated stimuli.................102 Figure 16: Experiment 4: Shift/size invariance with translated stimuli...............104 Figure 17: Experiment 4: False recognition of non-learned stimuli....................105 Figure 18: Experiment 5: A subset of the 45 training stimuli..............................107 Figure 19: Experiment 5: Position invariance with scaled stimuli.......................108 Figure 20: Experiment 5: Shift/size invariance with scaled stimuli.....................110 Figure 21: Experiment 5: Performance on non-learned stimuli..........................111 Figure 22: Experiment 6: Testing stimuli with 6 degrees of ablation..................116 Figure 23: Experiment 6: Quasi-stable recognition (8 associations)..................118 Figure 24: Experiment 6: Quasi-stable recognition (16 associations)................121 Figure 25: Experiment 7: Results for 8 hetero-associations...............................125 Figure 26: Experiment 7: Results for 16 hetero-associations.............................127 Figure 27: Experiment 8: Testing stimuli versus the resulted recognition..........131 xiii 1. Introduction Visual perception is one of the most fundamental abilities of our cognitive system. For the visual world to be perceived the way it is tasks like visual attention, pattern recognition, and visual cognition have to be seamlessly integrated and coordinated all the time. Visual processing, although seemingly effortless, is the result of massive and complex distributed computations that involve millions of neurons and a disproportionate amount of our brain circuitry. Joint research in many fields over the past few decades has generated a lot of discoveries for this marvelous cognitive modality that is probably the most thoroughly studied today. Numerous secrets, from visual neurobiology to high level vision, have been revealed experimentally so far. However, despite the huge number of details we already know about vision, we still lack a larger theoretical picture that will put all the disparate pieces together and connect the dots (Olshausen 2005). Most findings are dissociated from each other and most theories and models have a very restricted scope. The visual system is still a complex puzzle. The primary goal of this thesis is to begin to articulate a computational theory that will provide a neuronal-level explanation of the macro-level function of the visual system. In order to bridge the gap between neurons and behavior we have to shed the light not only on the pieces themselves but on how the puzzle assembles these pieces as well. Therefore we choose to experiment with a large scale integration of basic visual constituents instead of modeling certain details of the components of the visual system. Our belief is that understanding the fusion 1 of the information is sometimes equally, or even more, important than the information itself. And in order to understand the multifarious encapsulation of information that takes place in our brain we have to decipher the way that neuronal assemblies and network structures are formed. Studying the details without having a theory of how things may be put together is hopeless. Besides, from what we know about brain's robustness and adjust-ability, there is a strong indication that the integration process should be equally crucial and to some extent independent of the details of the individual components themselves. If this is indeed the case then our lack of knowledge for certain details at the neuronal level of processing, and the subsequent simplifications and assumptions we would have to inevitably make during this integration, should not be critical for the primary goal of this study. The second goal of this thesis is to propose and defend a computational framework that is not only biologically plausible and computationally feasible, but it also provides explanations at multiple levels. The proposed approach is able to provide the mechanistic details of a large scale neuronal integration while also accounting for a variety of behavioral phenomena at the level of visual cognition. It reconciles the notion of statistics on large image sets with the evidence from large numbers of distributed action potentials. We call it a framework rather than a model because different instantiations of the same architecture can exhibit diverse behaviors. Formally, it is a network of nonlinear dynamical systems with a well described mathematical formulation that is simple and easy to understand and analyze. Visually, the framework consists of a large network of orderly interconnected neural networks that are arranged in a grid-like topography. This 2 topological arrangement is inspired by the cortical columnar organization (Hubel and Wiesel 1974, Mountcastle 2003). Each constituent network has a specific function that abstracts the functionality of a cortical column, or minicolumn, depending on the level of abstraction we choose. The human cortex has powerful associative properties (Hebb 1949), hence we use associative neural networks as building blocks in our system. Each of these attractor networks learns to auto- associate visual stimuli in a local receptive field and forms stable point attractors that correspond to learned patterns of activity. In addition to internal module activity, networks are also interconnected and form hetero-associations of their internal states. This hybrid association drives the formation of network-wide attractors that correspond to pattern assemblies. Therefore, attractors that tend to co-occur cause a reciprocal excitation – conversely an inhibition – and facilitate the recognition of incomplete or noisy visual patterns. In the presented version of this framework there is only one layer of the proposed architecture, but in the general case there should be multiple layers organized hierarchically similarly to the visual association areas (Felleman and Van Essen 1991). The choice of the domain of visual perception for the experimentation with the proposed framework is not imperative. The overall approach could be equally suitable for other modalities like auditory perception or for integrating multimodal sensory information as well. Based on the principles of associativity and interconnectivity that are ubiquitous in the brain, an appropriate encoding of the stimuli along with a customization of the distribution of the dynamical systems network, should in principle be able to provide a diverse behavior to the proposed framework. 3 In the following chapters we first going to give a concise review of the trends that exist in the literature regarding the modeling of visual pattern recognition. We'll put an emphasis on a comparative analysis rather than on the individual details. This will give us an overview of the current situation as well as an understanding of how we got here and where we are heading to. Next, we're going to give a comprehensive description of the proposed framework, provide the reasoning behind its architecture, and discuss the implications of such an approach. Following that, we will put our framework side by side to the most popular models and examine qualitatively the points of agreement, but mostly, the ones of disagreement. There are important differences, conceptual and pragmatic, that have significant implications not only in how we perceive a model of vision but also in the scope and the limitations of the process of modeling itself. Last, we will present our experimental methodology and the results we got. We will discuss the findings and see how close they are to the ones we were expecting, as well as the predictions they make. We will conclude with some general thoughts, ideas for potential improvements, and future directions. 4 2. Review of visual recognition models The literature in the domain of visual pattern and object recognition is very large and rich. Since we argue for a meso-level approach that is able to account for high level behavioral phenomena while using low level mechanisms, it is necessary to examine a wide range of models that have a focus spanning from the neurophysiology of vision to visual cognition. Although this review will be far from comprehensive the main goal is to give a brief overview of the most influential models in the literature and examine the details of their operation, the constraints they work under, and the assumptions they make. This will provide an insight to their explanatory and predictive power that will help us make a comparison with our approach later. A review of the literature reveals a classification of the visual recognition models mainly according to the following basic dimensions: the representation of the object, the encoding of the stimulus, the architecture of the model, and the flow of processing. Based on these principal criteria we see distinctions between an object-centered and a view-based representation, a generic versus a class- specific encoding, a monolithic versus a modular architecture, and a feed-forward versus a feedback processing (Logothetis and Sheinberg 1996, Tarr and Bulthoff 1998, Wallis and Bulthoff 1999, Riesenhuber and Poggio 2000b). These distinctions may not always have an obvious implication on the performance of the various models but they have a profound significance on their plausibility and feasibility as models of human behavior. Next, we will review several models based on their stance towards a variety of these principles. 5 2.1. Representation One of the fundamental questions in visual recognition concerns the representation of the object or pattern that is recognized. On this aspect there are two main schools of thought. The first one argues for an object-centered representation that holds information about the full object in 3D space while the second one for a view-based representation that is limited by the number of views it has collected. In the first case we talk about a structural description that is formed by extracting visual information and building a view invariant representation. Once stored, the representation is uniquely defined, independent of the viewing conditions (viewpoint, lighting, etc.), and can be tested against previously stored or new objects. This idea was first introduced by Marr and Nishihara 1978 who argued that a shape representation for recognition should use an object-centered coordinate system and should include volumetric primitives of varied sizes. In this direction they suggested that the process of deriving a shape description such that must involve a means for identifying the natural axes of a shape and a mechanism for transforming viewer-centered axis specifications to specifications in an object-centered coordinate system. Moreover, regarding the process of recognition itself, they assumed a number of indexes that allow a new description to be associated to the stored ones. Clearly, the whole approach was influenced by the computer metaphor of the time. Few years later Biederman proposed an object-centered theory of visual recognition named Recognition by Components (Biederman 1987, Hummel and Biederman 1992). The components in this proposal are geometrical volumes, called geons, that correspond to regions of deep concavity and constitute the 6 building blocks for representing objects in 3D space. Their basic assumption is that these generalized cones can be derived from only five properties of a 2D image (curvature, collinearity, symmetry, parallelism and co-termination) which are invariant over viewing position and image quality. Therefore, since geons are reusable and can be freely assembled to form a variety of visual configurations, a modest set of them can theoretically be adequate for the representation of a far larger set of objects. Depending on the richness of this geon bank and the limitations of the structural description the recognition by components can thus be very powerful. This theory, however, lies in a very high level of description without providing a formal account for the recognition of the geons themselves. Moreover, the extraction of a structural description is dependent on the number of disparate views the observer would be able to collect from an object. Therefore, the so claimed view-invariant recognition of an object-centered approach cannot be completely different from a view-based approach. On the other side, the view-based recognition argues for a representation that consists of a collection of image views. A view doesn't necessarily correspond to a two-dimensional image but it can be any kind of viewpoint-dependent extraction of visual features, including depth and stereo information. Apparently, this type of representation is subject to noise caused by illumination and the viewing conditions in general. Moreover, the efficacy of this representation greatly depends on the number of disparate views the observer had the chance to collect in the course of time. If sampling of the scene is restricted to certain viewpoints only, extrapolation to new ones will probably be difficult. Although, this approach seems less powerful than the object-centered one and, in a sense, 7 counterintuitive to our own visual recognition abilities, experimental evidence favors the hypothesis that visual recognition in both humans and animals is viewpoint dependent (Logothetis et al. 1994, Tarr 1995, Tarr 1999, Tarr et al. 1998; but see Bar 2001 for a different interpretation). Psychophysical experiments demonstrate that humans are able to recognize unseen object views by interpolating views they previously learned (Bulthoff and Edelman 1992). This is also supported by model predictions (Poggio and Edelman 1990) as well as neurophysiological data showing that the majority of inferior temporal neurons are tuned to single views of the training objects with only a very small number of neurons being view invariant (Logothetis et al. 1995). Most models in the literature adhere to a view-based approach and regardless of the many other differences they have they all get trained on a set of image views coming either from within the same class of objects, if the purpose is the identification, or from different classes, if the purpose is the categorization. Some of the most known approaches in this direction are the models by Poggio and Edelman 1990, Perrett and Oram 1993, Neocognitron (Fukushima 1988), SEEMORE (Mel 1997), VisNet (Wallis and Rolls 1997), and HMAX (Riesenhuber and Poggio 1999). 2.2. Encoding The encoding of the input stimulus is of fundamental importance and probably one of the most critical issues, not only for visual recognition but for visual processing in general. There exists a great variety of proposed methods for encoding the input such as, intensity values of the raw image, color distribution, 8 shape information (e.g. different moments), texture, spatial frequencies, templates, etc. More important, though, the encoding of these features can either have a local scope (e.g. the color value of a particular pixel or receptive field) or a global one (e.g. the color distribution of the whole image given by the histogram). In either case, the goal of the encoding is to reduce the dimensionality of the highly variable input signal in order to facilitate the categorization or identification during visual recognition (Field 1994). By using a rich set of generic features with a local scope it is easier to flexibly encode a larger input space for a variety of objects from different classes. This, of course, comes at the expense of a higher complexity during the process of integrating these features across the image. On the other hand, the use of wide scope features, that result either from the encoding of the whole image or from a segmentation of it into regions of specific interest, has several limitations but with fewer problems during the information fusion. This latter approach reduces the high complexity arising from the full joint distribution over the features but is restricted to certain configurations or classes of objects. So, there is a clear trade-off and most models of visual recognition choose one of these two approaches. They either start with a local scope encoding of the two-dimensional input signal and ascend the path of integration by assembling these features, or they assume some kind of preprocessing that provides a wide scope class-specific image encoding that eliminates most of the input variability and focuses on higher-level relations within the image. An example of the latter is the eigenfaces recognition (Turk and Pentland 1991) that takes a holistic approach and describes each face as a combination of eigenvectors produced by a high-dimensional space of face images. The role of 9 the encoding here is simply to extract the principal components (eigenfaces) of a covariance matrix formed by concatenating the training images. The approach is global in nature and doesn't take into account any of the features that make a face what it is. The most well known model of visual recognition that works with high-level information is probably the fragment-based recognition (Ullman and Sali 2000, Ullman et al. 2002). The model doesn't use absolutely global image information, though, but it partitions the images into fragments that are associated with a certain class of objects. The fragments have a varying size that depends on the class of objects they target and their selection is made with the goal of optimizing their mutual information with a particular class. Recently, the model has been extended to incorporate a hierarchy of fragments that argues for a possible framework for categorization, recognition and segmentation in human vision (Ullman 2006). All these approaches emerged primarily from a computer vision background and the need for automated systems that can borrow some of the human abilities. Apparently, by working with more abstract visual information one can avoid the issues that arise from processing the low-level highly-variable input signal, and possibly achieve a better recognition performance. However, the assumptions that these models make have no biological ground and so far seem pretty implausible. Although, it is true that in the higher association areas of the visual cortex (e.g. V4) the concept of features becomes very complex; and it is also true that there exist selective cells in the inferotemporal cortex that respond to complete object views (Logothetis et al. 1995, Tanaka 2003), there is no evidence so far that these complex features or object views correspond to image 10 fragments of arbitrary sizes. Even if in the future such an encoding proves to be correct, and therefore justifies the approach that these models take, the fundamental question (in computational terms) of how do we arrive at this level of encoding remains unanswered by these approaches. Contrary to these approaches, most of the models in the literature adopt some variation of a local-feature encoding. Some of them use receptive fields with predetermined position and size (Fukushima 1988, Wallis and Rolls 1997, Riesenhuber and Poggio 1999). Some others use receptive fields with random positions and varying size (Edelman 1993, Edelman 1996), while other models use local patches without committing to the concept of receptive fields (Nelson and Selinger 1998, Amit and Geman 1999). As for the feature extraction itself, most models apply some form of filtering on the raw signal that detects edges or orientation bars, while others use Gabor wavelets which according to neurophysiological evidence model the filter response profiles of the simple cells in V1 (Daugman 1980, Marcelja 1980, Lee 1996). But regardless of the numerous details, the basic idea in most models is the same. The processing of the stimulus starts independently on an array of local fields and the information is incrementally fusing during the flow of the visual processing. Similar to what happens in the mammalian visual system, from the retina, to LGN, to V1 and the visual association areas, these models start with small pieces of encoded information and gradually fuse the information. 11 2.3. Architecture Although we have a pretty good idea of the overall architecture of the visual cortex, the details within each cortical area and especially the details of their interconnection are not well known (Felleman and Van Essen 1991). This is the main reason why, contrary to the representation and the encoding of the stimuli, the architectures of the various visual recognition models demonstrate a great diversity. For some models, this differentiation is justified by the level of abstraction they are targeting at. For others, though, it's only an assumption that simplifies modeling. Although there is no agreed-upon classification of the various architectures we will attempt to give one that ties better with the rest of the model features we are reviewing in this section. Given that the input to the visual system is a two-dimensional signal that is processed in subsequent stages, we can distinguish architectures in the horizontal and the vertical dimension. In the horizontal dimension we see differences among the architectures of the models in the way they compartmentalize the two-dimensional signal and the approach they later use to bind the processed pieces together. In the vertical dimension the most common variation regards whether a model is employing a single- or a multi-layer architecture. The two dimensions of the architecture that most models assume are usually independent. So for instance, we can see both monolithic and modular architectures with both single- and multi-layer models. An example of a monolithic horizontal architecture is the model by Hinton et al. 2006. This model assumes a single sensory input for the whole image. Despite 12 the fact that it uses multiple layers to process the information in both a bottom-up and a top-down fashion, it never breaks the unity of the visual stimulus. Apart from being biological implausible, this type of horizontal architecture has difficulties in scaling since the dimensionality of the system increases with the square of the input (Geman and Bienenstock 1992). By contrast, most of the models of visual recognition assume some form of a modular horizontal architecture. In this case the compartmentalization can have several forms like, receptive fields (Fukushima 1988, Edelman 1996, Mel 1997, Wallis and Rolls 1997, Riesenhuber and Poggio 1999, Serre et al. 2007b), patches (Nelson and Selinger 1998, Amit and Geman 1999), attentional windows (Anderson and Van Essen 1987, Olshausen et al. 1993), or even fragments of the image (Ullman and Sali 2000, Ullman et al. 2002, Ullman 2006). However, more critical in a modular architecture is not as much the name or the type of the constituents but the approach the model is using to bind them together. This is where the so called vertical architecture kicks in. In the vertical dimension most models have either some form of a layered structure or a single layer that incorporates an extra functionality. Characteristic examples of a vertically flat architecture are the models by Nelson and Selinger 1998, Edelman 1993, and Ullman and Sali 2000. These models do not focus on the biological plausibility of the architecture so they use a single layer on which they apply a series of computer vision processes without attributing them to a particular stage of visual processing. So typically, they use a monolithic vertical architecture but it is so complex that if we wanted to interpret it in biological terms we would probably couldn't do it without resorting to some form of a multi-layer 13 architecture. On the other side of the spectrum, the models that exhibit some form of modularity in the vertical dimension have either a stack of layers or a hierarchical architecture. For instance, the model by Anderson and Van Essen 1987 belongs to the first category. Its layers serve the purpose of routing the information from the bottom to the top instead of performing a hierarchical binding. By contrast, most of the models with a modular vertical architecture, like Neocognitron (Fukushima 1988, Fukushima 2005), VisNet (Wallis and Rolls 1997), and HMAX (Riesenhuber and Poggio 1999, Jiang et al. 2006), use the sequence of layers to gradually integrate the information and arrive to a final percept at the top layer. A modular architecture, though, does not only mean an information binding that occurs between adjacent layers. The brain architecture implies a rich interconnection among layers as well as among neurons within a layer (e.g interneurons) that play a significant role in the associative nature of this complex system (Thomson and Bannister 2003). Although there are some models that make a first attempt to account for this kind of connectivity (e.g. Wang 2001) most models choose not to touch this aspect of the cortical architecture. 2.4. Processing flow The brain architecture implies a system with multiple components that are interconnected and interdependent. In the visual system the graph of the various visual areas is densely connected with almost 40% of the cortices having links with all other areas and the majority of them being reciprocal (Felleman and Van Essen 1991, Braitenberg and Schuz 1998). This connectivity wouldn't be a big 14 problem if we were dealing with a directed acyclic graph. In that case the flow of processing would only depend on the point of entrance and the routing at each node. However, the brain circuitry involves many loops, both thalamo-cortical (Mumford 1991) and cortico-cortical ones (Mumford 1992). Hence, the flow of processing is not merely a function of the input and the routing, but a function of the time too. The more time we allow for the activity in the system to propagate across the network the more interactions among the areas will occur. Thus, the behavior that emerges is not the result of a simple input-output system but the outcome of a dynamical process that may or may not converge to a final state. And this, of course, is a matter of dynamics – the time evolution of physical processes. Therefore, time is a critical parameter that constraints the flow of processing and has significant subsequent implications on the types of behavior a model can exhibit. The flow of processing as a direct consequence of the temporal dynamics is probably one of the major differences among the various models of visual recognition. Behavioral models work in a more abstract level and do not usually provide a detailed account for the flow of processing. Computational models, on the other hand, provide a more formal description of the processing flow but they take a dichotomous position towards time. On one side we have models that decide to ignore it at all and adopt some kind of feed-forward processing, while on the other side, we have models that incorporate time in various forms usually as feedback or recurrent activity. The reasoning behind fully dropping the temporal component of a dynamical system like brain is based on the argument that visual recognition is too fast for feedback loops to have a significant effect in the 15 processing. The feed-forward models do not claim, of course, that feedback connectivity does not play some role in visual recognition, just that the speed of processing imposes a constraint to the modeling and makes feedback an unnecessary complexity. The fact that experimental evidence from EEG studies (Thorpe et al. 1996) have shown that humans can perform object recognition withing 150 ms, which is comparable to the latency (150-270 ms) of responses of IT units (Fuster 1990), is interpreted by some researchers (e.g. Riesenhuber and Poggio 2000b) as a strong indication for an immediate object recognition. However, they agree that this doesn't rule out the use of feedback processing. In a series of experiments that Serre et al. 2007a did, they found out that a feed- forward mechanism could only model the performance of humans when the possibility of feedback and top-down effects was eliminated with the use a stimulus onset asynchrony that was less than 50 ms. Apparently, the role of feedback connections is not to be neglected if a more complete account of processing is what we seek. Wallis and Rolls 1997 argue for a feed-forward visual recognition on the basis of a speed of processing account too. Their argument is based on neurophysiological studies by Tovee et al. 1993 that examined the responses of face selective neurons in the temporal lobe of rhesus macaques. Information theoretical analysis on neuronal spike trains in this study showed that a firing period of as little as 50 or even 20 ms is sufficient to give a reasonable estimate of the firing rate of the neuron. By doing a principal component analysis they showed that the information available in a short window of only 50 ms of firing rate can be as high as 84% of the information available in a firing period of over 16 400 ms. Even a window as small as 20 ms could account for 44% of the total information. In general, this analysis provided evidence that the start of the neuronal response provides a reasonable proportion of the total information that would be available if a long period of neuronal firing were utilized to extract it. There is also some evidence from visual backward masking experiments that a cortical area can perform the necessary computation for visual recognition in as little as 20-30 ms (Rolls and Tovee 1994). All this evidence combined led Wallis and Rolls to conclude that a feed-forward processing should be sufficient for visual recognition. However, a more recent model (Deco and Rolls 2004) makes a heavy use of both feed-forward and feedback connections, so it seems that they value more the role of feedback connections after all. In either way, none of the above evidence regarding the speed of processing is strong enough to support the idea that feedback processing is not participating in visual recognition, most probably the opposite. An information content of 40 to 80 percent in the first milliseconds is far from a complete 100%. It may be true that the initial response is more critical but that doesn't mean that the rest of the processing is useless or purposeless. For instance, the more time we allow for processing the more complex tasks we are able to solve. As for the latencies in the visual areas, given that feedback activity cannot be easily isolated when running experiments, it's quite tricky to assume that the measurements are free of the effects of the cortico-cortical loops. To the contrary, it's very likely that the latencies reported include the contribution of the feedback connections as well. Besides, there exist recurrent connections within visual areas that are involved in feedback processing without having to travel back to distant visual areas. And 17 last, there are analyses that suggest that even the local recurrent neocortical circuits can produce very rapid dynamics with small latencies (Treves 1993). Hence, feed-forward processing is not necessarily the only way to achieve the observed speed of recognition. Moreover, most models of visual recognition separate the process of recognition from learning and attention. It's hard to imagine how a model without a feedback process or a top-down influence could engage in visual learning, attention, storage and recall of stimuli. This may simplify things and help analyze independently the various visual processes but is biologically implausible and unrealistic. Even if it works for simple cases its limitations will uncover in the course of time. In this direction there is evidence from a complexity level analysis of visual search that speaks against a bottom-up processing (Tsotsos 1990). Architectural constraints suggest that attentional mechanisms that perform top- down tuning and selection of the stimulus are necessary during visual processing. Overall, the flow of processing causes a major split in the models of visual recognition because it differentiates them substantially from each other. The majority of models probably assumes a feed-forward processing with the most representative examples being Poggio and Edelman 1990, Perrett and Oram 1993, Neocognitron (Fukushima 1980, Fukushima 1988, Fukushima 2005), SEEMORE (Mel 1997), VisNet (Wallis and Rolls 1997), HMAX (Riesenhuber and Poggio 1999, Jiang et al. 2006), and Ranzato et al. 2007. On the other end, some of the most representative models that apply a feedback processing are 18 Mumford 1992, Olshausen et al. 1993, Rao and Ballard 1997, Deco and Rolls 2004, and Hinton et al. 2006. The feed-forward models have a sequential processing which is simpler and more intuitive so one can find many similarities among them. By contrast, in the case of feedback models, there is no general agreement on how the recurrence should work (within cortical areas, between cortical areas, both, etc.). Each one implements its own approach and one can see many differences among them. 2.5. Learning Learning in visual recognition is not a domain of debate or dichotomy among the various models. Most of them choose an approach that better suits their own architecture and not some general criteria. This is mostly due to the fact that the macroscopic properties of learning are not fully understood. We do know many microscopic details regarding synaptic plasticity and long term potentiation but when it comes to visual recognition which involves large groups of neurons and complicated circuitry it is hard to make any strong claims. Moreover, since all models of visual recognition work with neuronal spike rates that are easier to simulate and avoid the intricacies of spike-timing dependent plasticity (Markram et al. 1997, Zhang et al. 1998), any debate regarding the actual neural basis of learning would be highly speculative. In a more abstract, computational, level of description where most of the learning of the visual recognition models takes place, the main classification of learning mechanisms is between supervised and unsupervised methods. Supervised methods are easier to control but are not biologically plausible since they require 19 training examples that are labeled. Unsupervised or adaptive learning (Kohonen 1982), on the other hand, seems much more plausible but is of limited use to the task of visual recognition since it works as a method of clustering or dimensionality reduction that requires extra assumptions in order to account for recognition. An alternative approach is a hybrid method called semi-supervised learning (Blum and Mitchell 1998) that combines these two approaches. It first uses labeled examples to generate the best possible classifier and then gradually incorporates into the model its best predictions on the unlabeled data. A different type of learning that has both a biological underpinning and a computational realization is Hebbian learning (Hebb 1949). It is a form of associative learning that works in a distributed self-supervised way and, although it's not as powerful as the other methods, it has been used successfully for constructing content- addressable associative memories (Hopfield 1982, Anderson 1993). All these types of learning have been used in diverse ways in the various models of visual recognition. For example, Neocognitron (Fukushima 1988) is composed of layers of simple (S) and complex (C) cells that are arranged alternately in a hierarchical network. In this model only the layers with the S cells have their input connections modified through learning and the training can be either supervised or unsupervised (both methods have been proposed). A similar approach is followed by HMAX (Riesenhuber and Poggio 1999) which also has alternate layers of simple and complex cells. In this case the S cells show a Gaussian-like tuning and can be trained in an unsupervised fashion with a variation of the Hebbian rule while the C cells perform a nonlinear max operation and therefore their synaptic weights are all uniform and fixed to 1. A more recent modified 20 version of the same model (Jiang et al. 2006) is using an extra layer on top of the hierarchy in order to perform a supervised categorization of the features extracted from the view-tuned units. In a rather different model Hinton et al. 2006 use multiple layers of Boltzmann machines (Hinton and Sejnowski 1983) to create a deep belief network. The top two layers form an associative memory while the remaining hidden layers are trained with an unsupervised method called wake-sleep algorithm (Hinton et al. 1995). Although the learning algorithm is unsupervised it can also be applied to labeled data by learning a model that generates both the label and the data. Another example of a layered architecture (Ranzato et al. 2007) is having two stages, one for the encoding of the stimuli and one for the decoding, where learning proceeds in an EM-like fashion by using a gradient descent algorithm in a supervised mode. A more complex model with a multilayer architecture (Deco and Rolls 2004) is using both feed-forward and feedback connections, as well as intra-layer lateral connectivity. The feed-forward connections in this case are trained with Hebbian learning while the back-projections, which are symmetric and reciprocal in their connectivity with the forward connections, are set to a fixed (single parameter) fraction of the strength of the forward connections. As for the intra-modular local competition it is implemented as a lateral inhibition in a local neighborhood of neurons with a Gaussian-like weighting that is a function of distance. Apparently, there is a rich variety of learning methods and, unlike the other principles we reviewed so far, there seems to be no consensus on the approach 21 to be taken. Furthermore, it is obvious that learning is not a standalone issue but is highly related to decisions regarding the model's architecture and the flow of processing. For example, Hebbian learning is more suitable for bidirectional connectivity and recurrent processing while a supervised learning is easier to apply to omni-directional links and feed-forward processing. 22 3. The proposed approach In this thesis we present a computational approach to visual recognition that is based on a large scale dynamical system with a hierarchical architecture that has modular layers. The system employs a recurrent processing scheme applied directly to low-level visual information and generates a view-based distributed representation of visual patterns. The proposed framework is in agreement with the concepts of topographical representation and hierarchical processing that most models advocate but introduces a new paradigm regarding the flow of processing and the intra-layer communication and learning. Overall, it puts an emphasis on the information fusion that occurs within a particular stage of processing, an aspect that is less explored and usually overlooked despite being ubiquitous in the visual system and the brain in general. The key idea behind the proposed approach is coming from the work on Network of Networks (Anderson et al. 2007, Anderson and Sutton 1995). Conventional neural networks do not scale easily to larger dimensions, so the workaround is to create large scale networks by integrating small ones. At an abstract level, this idea is quite simple and there is plenty of evidence that the brain circuitry is following it too. However, there are many details (e.g. coupling of the networks, architecture, etc.) that can take many different forms which makes the proposal quite challenging. In the Ersatz Brain Group at Brown University we have worked on this approach for many years, mainly in the domains of language and cognition. This thesis was an attempt to apply the concept of Network of Networks to the domain of visual perception. It started a few years ago with the 23 modeling of pop-out phenomena in visual attention, and through various modifications over the time it evolved to the current form. One thing it showed is that the overall approach is quite generic and can be used equally well to model various modalities. But the major contribution of this thesis to the work on the general concept of Network of Networks, is not as much the domain of application. What makes it more valuable is the experimentation with a fully dynamical system. Due to the high computational demands of the proposed idea, it was quite difficult in the past to test very large structures without resorting to some sort of simplification (e.g. discretization of attractors, scalar couplings, substitution of automata for dynamics, and others). In the current work, with the use of High Performance Computing, we were able to fully implement all aspects of dynamics and of parallel distributed processing for very large scales, even up to millions of networks. But most important, we had to resolve numerous issues that arise during large integrations of dynamical systems that may seem trivial when a simplification is applied but are quite crucial when the complex dynamics comes into play. Next, we present the approach we take in the proposed framework regarding all major dimensions of the visual recognition models as we discussed them in the previous chapter. 3.1. Representation The proposed framework relies on a view-based representation and makes no assumptions about the structure and the configuration of the visual patterns. The visual stimuli correspond to two-dimensional images taken as snapshots from a 24 particular viewpoint and they are generated from a high resolution sampling and subsequent digitization of the raw image signal. There is no preprocessing in the form of clustering or segmentation that could extract more abstract visual information like object fragments or image components that facilitate a higher level non-pictorial representation. This absence of any internal representation of the visual patterns to be learned entails that the performance of the system is highly dependent on the richness and the complexity of the visual stimuli. This is also true for the capacity of the visual recognition as well as the invariance of the system towards viewing conditions. Everything will depend on the number and the content of the views that the system has the chance to be trained on. On the other hand, this type of representation has no bias towards a particular visual configuration or class of patterns in either 2D or 3D space. In other words, the system is not coming with any pre-built knowledge about the visual representation of objects, and whatever representations it will form they will be the mere result of the training and the statistics of the visual world. 3.2. Architecture Choosing an architecture is probably the most difficult decision to be made since it will set the scaffold for everything else to be based upon. If we go after an extremely detailed approach that is modeling properties at a level as low as the molecular one (e.g. the Blue Brain Project, Markram 2006) then we run the risk of never being able to interpret the findings in a more abstract computational level of description. If, on the other hand, we take a serial discrete view of behavior biased by the computer metaphor, we run the risk of not being able to tell 25 whether the findings have anything to say about the actual brain. Keeping the balance between a highly distributed system and a system that is devoid of the unnecessary complexities of the neuronal level is the primary goal of the proposed architecture. In order to bridge the gap between neural activity and behavior we need a multifaceted architecture that provides explanations at multiple levels. The architecture must retain its low-level biological plausibility while at the same time being able to account for high-level cognitive phenomena. In such a meso-scale approach we cannot neglect the principles dictated by any level of abstraction nor can we dissociate these levels like Marr did in his three- level analysis by claiming an independence between computation and implementation (Marr 1982). To face all these challenges the proposed framework is based on a multi-layer modular architecture with a hierarchical organization. In the horizontal dimension each layer has a spatial modularity with the constituent units exhibiting a massive communication that produces a highly hetero-associative network. In the vertical dimension the constituent modules of neighboring layers have the same functional interaction that they have within their layers but they are organized in a hierarchical fashion that produces different scales of abstraction. The most critical issue that arises from this approach is not as much the hierarchical organization as the specifics of the within-layer modularity and organization. This is the point where we'll put our focus both in theory and in the experiments. First, we will describe the structure of a single processing unit, then the organization of a whole layer, and finally, how multiple layers can stack on top of each other to form the full system. 26 3.2.1. The basic processing unit Recent evidence from neuroscience favors the argument that neural computation involves a lot more than just an aggregate of individual neurons. The neuron doctrine is under continuous scrutiny and the experiments lately put a great effort on exploring the coordinated behavior of groups of neurons (by using several techniques like multiple-cell recordings, fMRI, etc.). Moreover, it has been known for a long time that there exist in the cortex columns with a discrete anatomical organization (Mountcastle et al. 1955). These columns seem to have a nested structure that contains about 50-100 minicolumns with each minicolumn having roughly 80 neurons. More important, though, is that these cortical columns exhibit a functional specialization and operate as discrete processing units. This is also suggested by their connectivity which is dense internally (i.e. along the vertical dimension of the cortical sheet) and more sparse externally (i.e. along the horizontal dimension in their communication with other columns). In the proposed framework, the basic processing unit is a single cortical column and not a neuron. Our goal is to depart from the organizational level of networks of neurons (i.e. neural networks) and reach the level of network of networks. Although the full details of the organization of the cortical columns are not yet known, their functional role in some cases has been clearly revealed. One such example is the orientation selectivity columns discovered by Hubel and Wiesel 1962. In this case it is easy to abstract the functionality of an otherwise complex column with a simple neural network that performs a specific function. In our case this neural network will take the form of an auto-associative attractor network. The reason we take this approach is that we want a unit that is able to 27 learn its function directly from the stimuli, and we want this to happen in an unsupervised fashion too. Brain is a powerful associative system that learns with several variations of the Hebbian rule, so the choice of an auto-associative attractor network should be sufficient to model an abstracted version of a simplified visual cortical column. The idea that neural activity behaves like dynamical attractors has been suggested in the literature (Amit 1989) while there exists increasing evidence that many areas in the brain, like hippocampus (Wills et al. 2005) and olfactory bulb (Niessing and Friedrich 2010), form attractor representations. An attractor network (Figure 1) is a fully connected neural network with no input or output layers and no need for mapping patterns to labeled responses. A set of neurons is interconnected with Hebbian synapses and learns to associate a set of stimuli with themselves. Intuitively, it's a content addressable storage that is able to retrieve a partially completed or corrupted memory. Figure 1: Abstract view of an attractor network Formally, it's a dynamical system that is able to form stable attractors in a hyper- dimensional space that correspond to learned stimuli of the respective 28 dimensionality. For example, Figure 2 depicts the energy landscape that an attractor network has formed after learning four different patterns. The fixed points (in red) will attract any network activity that lies within their basin of attraction (depicted as dotted circle for one of them). If a corrupted or noisy variation of a learned stimulus is given as input to the system the attractor network will be able to retrieve the original stimulus by driving the network activity to the closest attractor. Figure 2: Energy landscape of an attractor network with four fixed points 3.2.2. The organization of a layer Now that we have our basic building block, the attractor network, we have to see how we put many of these processing units together to form a single layer of the proposed framework. Figure 3 depicts an overview of the organization of this network of attractor networks (Dimitriadis 2008). Within a layer, the attractors 29 form a rectangular grid and connect to other attractors in their vicinity. For illustration purposes, in this visualization we only see a partial connectivity for only two of the attractor networks (depicted with dark grey color). The omni- directional arrows represent the axonal projections, from one cortical column to another, that are bundled in single links. The synaptic strengths of these grouped connections are assembled into matrices and depicted in the figure with small squares. Figure 3: A network of attractor networks 30 If a stimulus is repeatedly presented to the system, the synaptic matrices will essentially get tuned to the neural activity of the corresponding cortical columns and will learn to hetero-associate the states of the respective attractor networks. So, the attractor networks have both internal and external dynamics. Internally, they work as local auto-associative systems that recognize small visual patches extracted from their receptive fields. Externally, they interact through hetero- association with a number of columns in order to fine-tune both their own and the distal network's dynamic states. This dual dynamics occurs simultaneously and continuously for all the networks in a given layer. Therefore, we have a hetero- associative network built from auto-associative attractor networks. 3.2.3. Multi-layer organization This dynamical system can be extended in the vertical dimension by creating a stack of layers on top of each other. Each of the layers in this case will have the exact same internal organization that we've seen before. Regarding the inter- layer connectivity, the associations between columns in different layers will follow the same principle of hetero-associativity that we've seen for the intra-layer organization. Hence, the attractor networks will not only form horizontal hetero- associations within their layer but vertical hetero-associations that span across layers as well. All the hetero-associations in this approach, both horizontal and vertical, operate with the same principles. The only thing that changes is the content of these associations and their spatial distribution, two issues that we will discuss further later. As for the number of layers, it is a variable that depends on the scales of abstraction that we want to implement and therefore it's a function of the specifics of the encoding that we apply to the stimulus. However, the 31 power of the proposed framework can be shown with as little as one layer only. Figure 4 gives an illustration of the full-fledged implementation of the framework. In the first layer each of the attractor networks (green columns) processes its own local receptive field (yellow cone) and constructs through auto-association a set of attractors that correspond to the set of the recognized stimuli the respective network can handle. At the same time, the whole bottom layer, in a collective way that we described in Figure 3, is forming a hetero-associative network that binds together the internal states of the attractor networks forming by that an assembly of networks instead of simple neurons. An example of such Figure 4: Hierarchical organization of the hetero-associative network 32 an assembly might be the five green columns depicted in the bottom layer of the scheme in Figure 4. These attractor networks might have responded to some coherent configuration that existed in the stimulus (e.g. an object part) and formed an assembly of their internal states. In the vertical dimension, if we assume that the layers are organized in a hierarchical fashion as depicted in the figure with the many-to-one connections between layers, as we ascend the levels of the hierarchy we would expect assemblies of a larger scale to gradually be formed. So, for instance, the red columns in the second layer pool several hetero-associations of the layer beneath (green columns) together in order to form an assembly of a higher level of abstraction. In a similar way the blue columns in the third layer might pool different sets of hetero-associations from the second layer (red columns) and form an even higher assembly of attractor states, and apparently, this can go on and on to higher levels of abstraction and scale. As we reach higher levels of this hierarchy the attractor networks presumably encode larger parts of a visual stimulus. This also means that in the higher levels we would expect to see assemblies that consist of a small number of hetero-associative networks or even of a single column that encodes a whole object in a very localized way. Evidence for this kind of representation is coming from physiological data in the inferotemporal cortex, one of the later stages of processing in the ventral visual stream (Tanaka 2003). In the current version of the framework we put the focus of our experimental work on a single layer of the proposed architecture in order to acquire a comprehensive analysis of all the details pertaining to such a complex dynamical system. Although we won't have the resources to demonstrate the behavior of an 33 immense network of networks that is the result of a vertical integration of many such layers, we believe that the essence and the power of the proposed organization would nevertheless become apparent. Moreover, even with a single layer is still possible to demonstrate some of the effects of the multitude of abstractions by experimenting with a variety of stimulus resolutions and scales. 3.2.4. Connectivity map In the previous sections we described the general architecture of the system and specifically the structure of the basic processing unit, the organization of the layer, as well as the integration of a set of layers. What remains to complete a full description of our architecture is the interconnectivity of the units. A major scientific question towards this direction is the genesis of the synaptic connections in the brain. The resonance hypothesis, probably the first theory on this issue, suggested that neuronal connections during stages of early development were nonspecific and the cortical circuits were formed by the activity-dependent rewiring of the initially random connections. However, a series of experiments on the frog's visual system held in the 1940's gave rise to a contrasting hypothesis called chemoaffinity (Sperry 1963) which is the most widely accepted theory of neuronal wiring today. According to this theory, growing axons recognize chemical markers produced by target cells and by that they connect to specific neurons or groups of neurons. Thus, connections are not initiated randomly but are predominantly predetermined into the identity of the differentiated cells. This means that in the early stage of development circuits of the brain are largely hardwired. 34 However, this only answers the question of how the genesis of the connections is realized and not what kind of connectivity pattern is produced. Despite the large body of statistical data regarding neuronal connectivity (Braitenberg and Schuz 1998), our knowledge in this matter is mostly quantitative. We know for instance that cortex has a dense connectivity with roughly 40% of the areas having reciprocal connections with other areas but we don't really know any topographic details regarding these connections other than which areas connect with each other. Therefore, there is a need to devise and test a set of plausible and feasible connectivity patterns. As the chemoaffinity hypothesis suggests these patterns will be fixed for any given test. However, we will be experimenting with several patterns in order to study a variety of properties that will possibly help us discern the suitability of the various connectivity maps for certain situations. The simplest connectivity map is probably a fully connected graph of all the attractor networks. Although this kind of map has been already proposed for classic neural networks and has been shown to work quite well for small-scale practical applications (Kohonen 1982), it is biologically implausible and computationally extremely expensive for a large scale implementation of the size of the proposed framework. A more plausible map would suggest a full connectivity only for the units within a certain radius R1. If we assume a rectangular grid (as the one we described earlier) with each node having 8 neighbors within its layer and 26 neighbors in the multi-layer structure, the degree of connectivity would scale according to 4R2+4R for a single layer, and 1 The unit of radius R in this case is the distance between two adjacent nodes in the hetero- associative network and not the real distance of the units in space. 35 according to 8R3+12R2+6R for a non-hierarchical multi-layer organization. Given the rate of growth of these numbers and the fact that cortical neurons connect to an average of 103-104 other neurons, it is obvious that this type of connectivity can only be feasible for a hetero-associative connectivity of a rather small scale. Of course, with a more plausible hierarchical organization connectivity constraints would reduce the cubic growth but even in this case the degree of connectivity would retain a quadratic form which is highly unrealistic for a full connectivity among the attractors in a neighborhood that exceeds a small range. A more efficient connectivity pattern can have a Gaussian, or a similar distribution, where the density of the connections tapers as a function of the distance R from the column. One advantage of this pattern is that the number of hetero-associations can grow in a fully controlled way by manipulating the mass and the width of the distribution of the connections. More important, though, is that it allows the attractor networks to form associations at much longer distances than the previous connectivity patterns would permit with the given limitations. Of course, this ability comes at the expense of a connectivity that becomes sparser as we depart from the origin. However, since the target columns are not isolated processing units but they form hetero-associations of their own neighborhood as well, it should theoretically be possible for a column to have an adequate interaction with a distant area without preserving a full connectivity with all possible target units in it. Due to the dynamics of processing an attractor network has the ability to receive the distant activities in a progressive way and this allows the target column to incorporate into their communication its gradually formed neighboring activity. 36 A connectivity pattern is not necessarily exclusively spatial but can combine a functional role too. Psychophysical evidence regarding lateral interactions between receptive fields indicates that the interaction of receptive fields with the same functionality is a lot stronger when they are situated along vertical and horizontal meridians as contrasted to random locations relative to each other (Polat and Sagi 1994). If we wanted to convert this finding to a proposal for a possible connectivity map, that would mean that the distribution of the hetero- associations cannot only depend on the spatial properties of the attractor networks but it should take into account their function too. But according to our initial hypothesis, the function of each attractor network is to recognize a set of forms in its receptive field. So, the connectivity of the networks cannot be static. It should either have a dynamic switch depending on the attractor that a network is into at every time, or we should assume a connectivity where the hetero- associative matrices contain functionally-inactive hetero-associations. From a biological point of view the first approach seems highly implausible whereas the second approach is much more realistic. Besides, as we will see later in the section of learning, there is no process of associative learning that could store in a single hetero-associative matrix all possible hetero-associations. So, the idea of having some of the hetero-associations set a-priori to be inactive serves both a statistical and a biological purpose. 3.3. Stimulus encoding Regarding the stimulus encoding, we follow a local scope approach that is based on receptive fields with a uniform distribution over the input signal. Given that the 37 proposed framework is focused on the processes underlying the integration and the fusion of the visual information, it is necessary to make as few initial assumptions as possible. We choose to start with the most biologically plausible encoding, the one that resembles the receptive fields of the visual system, and do not rely on any kind of informative image segments or components that may facilitate processing but will defer the crucial issues and will leave a lot of important questions unanswered. Figure 5: A processing unit over the respective receptive field Figure 5 illustrates a single processing unit (i.e. the attractor network) and the respective receptive field that receives input from the underlying image patch. The signal in this case is a binary image that may result from an edge detection filtering process. This is a simplified depiction of a generic receptive field (RF) 38 based on the discoveries by Hubel and Wiesel 1962 and we will be using it throughout this document. 3.3.1. Geometry of the encoding Figure 6 depicts the geometric and topographic characteristics of the receptive fields and the encoding of the visual stimulus. For illustration purposes the shape of the receptive fields in this figure is round. The proposed framework, however, allows for elliptical RFs of any variable size. For computational efficiency, and in order to simplify the vector calculations, we are going to approximate the ellipsoid of an RF having axes sx and sy with a rectangular patch of sx × sy (depicted with red frames in the figure). Figure 6: Geometric characteristics of the stimulus encoding 39 This may modify a bit the properties of the local processing but, overall, it shouldn't have any effect in the macro-scale processing. The receptive fields within a certain stage of processing in the visual system (e.g. V1) have a variable size that increases with the distance from the center of the projected visual field. However, for small visual angles (less than 5 degrees) the size of the receptive fields can be considered quite uniform across a visual map 2. This is especially true for the part of a cortical area that undertakes the foveal visual field where most of the processing occurs. So it wouldn't be too risky to assume that for visual recognition – unlike visual attention and other mechanisms that recruit the peripheral parts of the visual field – a visual patch roughly the size of the fovea with receptive fields of equal size is sufficient for our modeling purposes. Regarding the macro-scale properties of the topography of the encoding, we will be using a uniform grid of receptive fields with a variable but equal degree of overlap across the visual maps. The proposed framework allows for the quantities dx and dy (Figure 6) to vary depending on the purpose of each experiment but their value has to remain fixed for all RFs for any given test. This means that the receptive fields will always be forming a rectangular lattice. Although the topographical arrangement of the receptive fields in the cortex demonstrates a hexagonal pattern of local alignment this is almost certainly a result of a developmental biological constraint and there shouldn't be any reason for not adopting a technically simplified rectangular alignment. Last, we will 2 In the case of V1, in particular, the increasing size of the receptive fields as we depart from the point of fixation serves mostly as a compensation for the lower density of the ganglion cells in the periphery of the retina. So the actual number of cells per receptive field, even for bigger visual angles, is probably less variable in this preliminary stage of processing. 40 always assume some degree of overlap between adjacent receptive fields that will provide the necessary redundancy and robustness in the system. So the distance of the receptive fields will be, 0