Dr. Fırat Soylu

Artificial Neural Network Demos and Examples

Across the different classes I teach, the question of "how artificial neural networks (ANN) work?" often comes up. I usually draw a very simple ANN and try to explain how learning takes place by adjusting weights among the neurons. This usually results in further confusion and with me thinking "I should have some very basic demos showing how ANNs work." Well, I finally started working on this and developed some basic examples.

I am not an expert on ANNs. More than two decades ago, when I was a first year grad student at METU, Ankara, I took a neurocomputing course from Dr. Erol Şahin, in which we covered fundamentals of ANNs. I probably could not understand half of the content we covered, even though Dr. Şahin was a great instructor. The content was difficult and it was nothing like what I thought AI was about. I thought AI was about writing conditional statements explaining what to do under different conditions. Here we were learning about these things called neural networks, somehow mimicking biological neurons and networks. Coincidentally, I also knew nothing about how biological neurons worked. Overall, as an aspiring cognitive science student, I knew nothing about some of the most important things I should have known about.

The following year I moved to the US and started the PhD program at IU Bloomington. The cogsci courses I took at IU helped me further understand the significance of the connectionist paradigm, not because I was particularly interested in AI or cognitive modeling, but knowing a thing or two about ANNs helped me make sense of, first, how studies in neurophysiology impacted cognitive science/modeling and AI (e.g., McCulloch & Hebb ) and second, how biological networks function. ANNs and biological networks are quite different, but understanding the way ANNs work provides a nice scaffold for later learning about the biological ones.

Long story short, I will present some basic neural network models both in Matlab and Javascript. The JS ones work on the browser. You need to download the Matlab files and run them locally. They come with no guarantees. Again, I am not an ANN researcher. I am putting these here, so that I can use them next time I want to show someone a simple example for what an ANN is and how it works.

Perceptrons: The Little Neuron That Could

A perceptron is a very basic neural net that has only one neuron (like politicians). Read the Wikipedia article for more info. McCulloch and Pitts first proposed the idea of artificial neurons in 1943. McCulloch was a neurophysiologist (the word neuroscience was not used back then). When they first met, Pitts was a young homeless logician in Chicago. McCulloch was a much older, established professor. McCulloch let Pitts stay in his home, which allowed them to collaborate. Back then it was customary for people to come up with world-changing scientific inventions with their homeless guests. The artificial neuron they proposed was very simple. It received inputs in the forms of 1s or 0s and the neuron added up the inputs and compared the sum to the threshold. If this sum was greater than or equal to the threshold, the neuron fired, meaning it produced 1 as its output. Such a neuron can do simple operations, for example the logical operation AND. The AND operation requires two inputs and produces one output. If both inputs are 1 then the output is 1. For the other three input configurations (1-0, 0-1, 0-0), it outputs 0. It is similar to how we use the conjunction word "and" in daily language. You can only leave home if you find your wallet and your keys. You only found your wallet and not your keys? Too bad, you are staying home. For the AND operation the McCulloch and Pitts (MCP) neuron added up the inputs and compared to the threshold value of 2. If the sum was smaller than 2 it didn't fire (output is 0), if it was equal or bigger than 2, it fired (output is 1). In the case of AND the max sum was 2 of course. But you can imagine the general logic of the MCP neuron. It was simple. Add all inputs and compare to the threshold. If the sum equals to or bigger than the threshold then fire. If not, don't fire. The MCP neuron was capable of doing calculations, but it was not capable of learning. Regardless, it was a revolutionary idea. McCulloch has a 1961 paper, fantastically titled "What is a number, that a man may know it, and a man, that he may know a number," where he argued that "An interconnected set of logical units can realize different numerical operations." The core idea that is important here is that when you connect multiple very simple computational units that receive simple inputs and produces a simple output (1s or 0s), they can do amazing things together. This idea originates from early neurophysiologists' studies, like McCulloch and Hebb, characterizing biological neurons as simple computational units, receiving inputs and producing an output. Of course, the electrophysiological mechanisms for how biological neurons work are much more complex. Characterizing a biological neuron as a simple computational unit is a reduction. Which means we sometimes ignore the complexity of the thing that we are studying and just focus on the most important aspects of it.

Around 1957 Frank Rosenblatt physically built a neuron that was capable of learning from trial and error. The machinery he built had inputs connected to a neuron, just like the MCP neuron, but this time the strength of the connections were adjustable (with varying electrical resistance). The amazing thing with this setup was that when the neuron produced the wrong output (e.g., when it suggested that you can leave home with just you wallet, but without the keys), it adjusted the weights between the inputs and the neuron, so that it became more likely that the neuron produced the correct output next time. The threshold logic was similar to the MCP neuron, but this time the inputs were multiplied with the weights and then were added up. If the sum was higher than the threshold it fired, otherwise it stayed silent. This was pretty amazing, because he invented an artificial neuron that could learn. Also, his neuron received many more inputs. Imagine images of shapes (e.g., circles vs. squares) on a 20x20 grid. Each cell value could be black or white (0 or 1). So, the neuron received 400 inputs, each input representing a grid cell of the pictures with values of 0 or 1 (black or white). In a way, the neuron had an "eye" with 400 photocells. The neuron was able to learn how to distinguish between two different shapes, which we now call a binary classifier. He coined the term Perceptron for this machine, which means perceiving machine. If I were him, I would have named it LNTC (The Little Neuron That Could).

There was a lot of enthusiasm about the promise of ANNs back then. There are some wild quotes about what ANNs were and would be capable of doing in the future. A 1958 article in The New Yorker reads, "...we’d like to tell you about an even more remarkable machine, the perceptron, which, as its name implies, is capable of what amounts to original thought." At the time they implemented the Perceptron on an IBM computer, which meant that it didn't need the photocells or resistance adjusted weights. The perceptron was simulated on a computer: it became software. Back then, people thought these little artificial brains had endless promise. They believed that soon enough the artificial brains would surpass human intelligence and that we would reach some kind of tech utopia. Sounds familiar?

The enthusiasm died off when Papert and Minsky published the book Perceptrons in 1969, where they showed the limitations of the perceptron. It couldn't even learn the logical operation XOR! I actually bought and tried to read that book after the neurocomputing class. It was too heavy for me and I couldn't understand it. But, I was able to understand that a neural net with no hidden layers is limited in what it can do. First, it only can solve linearly separable problems (the demos below will further explain this) and second, it has limitations when dealing with global properties. For example, it cannot determine whether a shape is a closed or open curve simply by combining local properties, because it cannot "see" the big picture.

The book named Perceptrons ironically quenched the enthusiasm about perceptrons and neural nets in general, because it showed that nets with no hidden layers were limited in their capability and they had neither the hardware nor the knowledge of algorithms to train complex nets with hidden layers. This period, which lasted until the 80s is called the first AI Winter: a period when funding for research on ANNs disappeared. The Lighthill Report, published in the UK, made the point that the lavish promises of AI had failed and was very influential (the report should have been named "Winter is coming for AI").

Many years later ANNs made a big comeback in the 80s, with the realization that backpropagation (another concept I will detail soon) makes it possible to train networks with hidden layers. They also had much better computers by that time, so both issues, algorithms and computers, were solved, opening the floodgates for the next AI craze. Unfortunately, Rosenblatt did not witness this comeback. He passed in a boating accident two years after the publication of the damning book: Perceptrons. By now, you might be thinking, "I thought this was about how ANNs can help with understanding brains and cognition: why am I reading about the history of AI? One reason is that the author (me) is experiencing what he calls “keyboard diarrhea,” the uncontrollable urge to keep writing in freeform (because this is my website and not a journal publication). The other, and more compelling, reason is that I believe knowing about the history of the concepts covered in the demos below are helpful. History also gives us some context and we learn better when things are contextualized. Also, comeback stories are always interesting and dramatic. The current craze we are going through with the large language models (LLMs) is an extension of this drama. LLMs are ANNs as well. They are the descendants of The Little Neuron That Could.

A brief caution about what ANNs can and cannot tell us about brains

Before I go on, I am inclined to include some warning statements here. I am an embodied cognition researcher. I don't think any sufficiently complex neural net can lead to human-like intelligence or consciousness. AI can obviously be very smart (and smarter than us; depending on how you define smart), but that doesn't mean that it is human-like (see the Chinese room argument). Also, we can never be sure that even other humans have consciousnesses similar to ours (see the Problem of other minds). I just conflicted myself. Even though I cannot be sure that my cat experiences the world somehow similar to me; I highly doubt that a disembodied AI (or with a very different body[ies]) will have a similar life experience to mine. My point is: We are organisms with bodies and an evolutionary histories. Biological neurons are extremely complex across multiple levels (quantum, molecular, cellular and systems levels). To make this complexity more manageable when we are doing science, we have to focus on a level or two, and think about the neurons as more simple things than they "really" are. Also, the nervous system does not only include neurons. There are as many glia as neurons. Non-nervous system cells, like microglia (immune cells) also carry out important functions in the nervous system (e.g., pruning, plasticity, learning), in addition to doing regular immune system tasks like cleaning up damaged cells. The nervous system also heavily interacts with the rest of the body, particularly, digestive, immune, and vascular systems. I believe the whole body and its situatedness in environments is what leads to consciousness. Reducing the human brain into a large network of small computational units does not make sense. I don't think it is even theoretically possible to create a model of one's brain, by mapping each neuron and all of their connections, uploading it to some super powerful computer, and expecting that person to wake up with the same consciousness or even a Temu version that thinks that it is that original person. Here is the video of an interesting debate (2000), where Doug Hofstadter is arguing against some of these fantastical ideas. More recently, Hofstadter (who is considered one of the grandads of AI) shared some very pessimistic views about how the AI explosion we are experiencing now will impact humanity. I share his pessimism. I am going on a tangent here. All I am saying here is that the reason why I find ANNs helpful in understanding how the brains work is not because I believe they are equivalent: I do not believe that. For me, ANNs are simple models that are good pedagogical tools. Just like a 3D plastic brain model. I use a plastic 3D model brain in my teaching all the time, but I've never thought that it if someone made a plastic model of my brain, I can wake up in a plastic brain. Though that would be an interesting science fiction scenario.

Basic ANN Model

What does this model do?

The model shows how different basic binary classification problems can be solved (or not solved) with a perceptron (no hidden layers) or with an ANN that has a hidden layer and uses backpropagation and gradient descent. The problems presented are logical operations: OR, AND, XOR. If you don't remember what these are, the table shows the two inputs and the one output. Think about OR as "If the lion cage is open OR the lion is on a leash, I am safe." For AND: "If I have my wallet AND my keys, I can leave the house." XOR is a bit tricky: think about a light controller that has two switches and the light turns on if only one of the switches is on, otherwise it is off. I can explain these things to you with examples, because you have previous experiences that you can relate to. How are you going to teach these operations to a brain with one neuron? This is what Rosenblatt accomplished and that is what we will do.

Logic Gates

Input A Input B OR AND XOR
0 0 0 0 0
0 1 1 0 1
1 0 1 0 1
1 1 1 1 0

You can interact with the model below. You can also open it on a separate window if you are so inclined: Perceptron + ANN Model. It shows the perceptron on the right. The X1 and X2 are inputs (1 or 0). For example, if you select AND for the problem, then for 1 and 0 as the inputs, you should get a 0 from the output, because you can't leave home with just your keys. You need the wallet as well. Y is the output neuron that tells you if you can leave home. 1 is a go, 0 you are staying (go find your wallet). Now, at the beginning, the model doesn't know what AND means. So it gives a random response. The random responses will be correct roughly half of the time (like a coin flip). You need to teach it what AND means. This will happen with trial and error, and learning. Click the "Next Step" button to run the model. In every step, the model is given the inputs and if it gives the wrong answer the perceptron learning algorithm changes the weights (and the bias, which is another value that helps with the learning) a bit to get closer to the right answer. Keep clicking "Next Step" until the target and output values are always the same, which means the error is zero. That means that the model learned what AND means. Then try doing the same with OR ("Problem" drop-down menu). Keep clicking Next Step until error is always 0. Then try it with XOR. Just like the confused electrician who wired your switches, the model will face some difficulties. This is one of the things Papert and Minsky figured out.

Now that you tried all three problems, you might have realized that Perceptron is good at figuring out AND and OR, but not XOR. Why? Well, because XOR is strange (like the light switches in your room). I have another demo for you to figure out why it is strange. But before that, let's reflect on what this means for nervous systems. For example, imagine an organism that responds to two environmental cues for the objects around: 1) whether the object is large and 2) whether the object is moving.

This is essentially the XOR problem and the organism cannot solve it with a single layer network, like the perceptron. There are indeed organisms, like the C. elegans (little worms with around 350 neurons in their nervous systems), that have sensorimotor neurons. These neurons receive sensory signals from the environment and can directly trigger action, without any further intermediation through hidden neurons. Not every neuron in C. elegans is like that, but some of them are. C. elegans can encounter both terrestrial and aquatic environments, and its nervous system can trigger crawling or swimming responses depending on the environmental signals it receives. But, not every problem organisms face are as simple as that. Most require a more complex neural architecture. XOR is one such problem. It requires combining the sensory signals in a nonlinear way, meaning that there is need for a hidden layer of neurons that can help distinguish the more nuanced situation the organism is facing. Let's look at what a network with a single layer can do, compared to a network with a hidden layer that intermediates between the inputs and outputs.

C elegans

C. elegans: the little worm with ~350 neurons

Check out the model below. Or open it on a separate window: Perceptron linear decision boundary model---> Now your task is to move the sliders for w1 (weight for input1), w2 (weight for input2), and b (bias; a constant added to the sum) to move and rotate the line around, so that the filled dots and the unfilled dots are on different sides of the line. Remember, the output neuron multiplies the inputs (x1 and x2) with the respective weights (w1 and w2), and sums them together with the bias constant, before comparing the total sum to the threshold value, which determines if the output neuron will fire (output=1) or not (output=0). So, the total Sum (z) is calculated as:

\[ z = w_1x_1 + w_2x_2 + b \]

That sum determines what the output will be and since the organism (I meant the network) cannot change the inputs, it needs to adjust the weights and the bias value to find the boundary line that correctly separates the space. Again, remember, if the sum is bigger than or equal to 0, then the neuron fires (output = 1), if smaller than zero it doesn't (output = 0).

\[ y = \begin{cases} 1 & \text{if } z \geq 0 \\ 0 & \text{if } z < 0 \end{cases} \]

The line separates the space into two and the two different types of dots should be in two different areas separated by the line. What does that have to do with what we were talking about before? Well, the locations of the dots represent the inputs. X1 (input 1) is on the x-axis, X2 (input 2) is on the y-axis. Now, let's talk about the boundary line. At the boundary the Sum above should add up to zero. Why? Because, because 0 is the decision threshold. If we leave x2 alone (because it is on the y-axis, and we conventionally write line formulas as y = ax +b), we get the boundary line formula below.

\[ w_1x_1+w_2x_2+b=0 \] \[ x_2=-\frac{w_1x_1+b}{w_2} \]

Check the model below: The color of the dot (filled vs. empty) represents the output. Filled is 0, empty is 1. You are the perceptron. You are learning when the output will be 1 or 0, depending on the inputs. For the AND operation all conditions, except for X1=1 and X2=1, leads to Y=0 (you can't leave, except when you have both your keys and your wallet). Now, for the AND operation move the sliders so that the line separates the space so that filled dots are on one side and the empty one on the other. Then switch to the OR operation and do the same. Then do it for XOR.

What happened? You cannot find a line that separates the space into two so that the filled and empty dots are on separate areas divided by the line, right? This is the issue of linear separability. Perceptron can only separate that space with a line and because XOR is not linearly separable, it cannot learn the XOR operation.

Now go back to the previous model (Perceptron + ANN) and from the model menu select "ANN + hidden layer." When you do that you will see that your ANN now has two hidden neurons (between the inputs and the output neuron). Now select XOR for the problem and keep clicking on the "Next Step" button. As you click, you will see the two dotted lines moving. These are the lines that are suggested by the hidden neurons to the output neuron. The hidden neurons transform the inputs, and the output neuron combines these transformations to create a curved/nonlinear decision boundary. The decision boundary is what separates the space into two. For the perceptron that is a line, which is why it fails at XOR. When you have hidden layers with backpropagation, the decision boundary can become nonlinear. Backpropagation here refers to an algorithm that can go back to the weights between the neurons and decide how much the weight needs to be changed based on the error. Every time an error happens it goes back and changes the weights a little bit so that the curve gets closer to separating the space into two correctly (filled dots on one side, empty ones on the other), which means it makes the loss smaller. The loss is something similar to error, but it is a continuous measure of how far the model's output is from the correct output. For example, the error is the difference between what the model predicts and what the correct output is. For the XOR, if the inputs are 1 and 1, and the output is 1, then then error is 0 - 1 = -1. Why? Because the output was supposed to be 0 (both light switches are on, the light should be off). The loss is different. The output neuron calculates a continuous activation value that can be negative (e.g., -1.89) or positive (0.56). If it is negative, the output is 0; if it is positive, the output is 1. This is how it calculates the output. The loss is the difference between the activation value calculated and the output. As long as the model calculates a probability that equates to the right output after the rounding, we call it a sucess, because we don't have to be a perfectionist. However, we still calculate the loss value. If you run the model for many steps the loss value will get close to zero. Though you don't need to do that. We consider the model having learned when it always gives the right answer and that is when the nonlinear boundary curve separates the space correctly into two (filled and empty dots are on different sides of the curve.)

How does this relate to the biological brains?

The human brain has many many layers of interconnected neurons. We have sensory organs and different motor capacities. I am very careful about not calling these as inputs and outpus, because all the research and philosophizing in the phenomenology, ecological psychology and 4E (embodied, embedded, enactive, and extended) cognition literature, on top of all the neuroscience work have shown us that we are not input->processing->output machines like the simple perceptron or ANN you played with. Our perception is guided by our embeddedness in the environment (motor movement empowers 3D vision) as well as our internal states (you will recognize food more readily when you are starving). We are also constantly predicting what is going to happen next (you are usually not even aware of that) based on environmental cues, previous experiences, and internal states. There are many feedback loops in the nervous system and in the entire organism. All of these make it nonsensical to characterize humans and other biological organisms as input-output devices, simply responding to environmental stimuli. We seek and selectively interact with the environmental stimuli. However, these simple ANNs are a good way to start understanding where we might be coming from. The first principle we learn from these models is that neural architecture requires a certain level of complexity given the problems that it faces. Over the last 4 billion years of life on earth, we faced many challenges. Our brain has about 85 billion neurons, not to mention as many glia. So, yes, this simple models are indeed simplistic, but they help acquire some fundamental insights.

Next, I will introduce a more complex problem: discriminating across shapes. We will build an ANN that has eyes and that can recognize shapes! It will learn to look at the shapes like the ones below and decide which ones are circles, triangles, and squares: a form of discrimination learning. This gets us a bit closer to some living things, but still far far away.

shapes