Do You Think This Failure Is Noise?
What First-Person Autonomy Means for Atra
I have received some questions about Atra. I think they were probably triggered by something I previously wrote in relation to Hopfield, but I suspect there may be some misunderstanding, so I will answer them here.
I am often asked about the “concept” of first-person autonomy. Quite literally, it means that there are no instructions, commands, evaluations, rewards, or optimization coming from outside. Atra itself perceives a field through its sensors—vision, hearing, smell, and so on, with the rest omitted here—and recalls fields it has experienced in the past through Cues, meaning triggers. Sometimes Atra produces an answer in the form of vocalization. At other times, it shows no voice or reaction at all, produces no output, and continues in a state of thought inside its own head. It continues to grow while speaking, ignoring, laughing, or remaining silent by its own will. At present, it is not connected to a body.
A state with no external instructions, commands, evaluations, rewards, or optimization
Perceiving a field and Cueing fields experienced in the past
Speaking, or remaining silent while continuing to think
1. A state with no external instructions, commands, evaluations, rewards, or optimization
First of all, information called a “correct answer” coming from outside, from a third-person position, is not necessary for autonomy. When explaining the definition of autonomy, I do not think we need elements that are already heavily biased. Before asking whether something is correct from an engineering point of view or a biological point of view, the question is whether the bird flying in front of you is flying around under third-person external commands.
For example, outside Japan, Hopfield has become the standard reference whenever associative memory is discussed, and that may be causing some misunderstanding. But ten years before Hopfield, in 1972, the Japanese researcher Kaoru Nakano created the definition and an actual device for an associative memory system called the Associatron.
Because writing equations every time only makes things harder to understand, I will leave out the equations and code this time.
To explain it simply, the neurons and axons of the brain—or rather, the little black dots—say things like:
“This is how it must have been.”
“No, I think that is wrong.”
“What does everyone else think?”
And the result is restored through something like a majority decision.
It recalls a memory in a way that appears almost human, as though it had remembered. Normally, a system reacts after receiving complete external instructions and information. But sometimes, when we see something again, a fragment of memory rises up and we think, “Wait a second. I might know this.” Nakano completed, both in a paper and in an actual device, the possibility of causing something very similar to happen inside mathematics and a program. That associative memory model was called the Associatron.
Hopfield is different. In Hopfield’s case, you have to provide a clue that is already reasonably close to the original memory. Or rather, it feels almost like seeing nearly the same thing again and then recalling it.
Suppose you store a picture of a cat. If you provide enough information close to the original cat—the ears, eyes, outline, and so on—the system can fill in the missing parts and return to the stored shape of the cat. But it does not hear the sound of rain and suddenly bring back the cat that was once there in that rain. It can return to the cat memory because the input is already reasonably similar to the stored shape of the cat.
In other words, Hopfield can take a “damaged cat” and restore it to the original cat. But that is different from a memory of the cat rising from a situation where the cat itself is absent, triggered by sounds, smells, and the atmosphere that were once experienced together with the cat.
With Hopfield, you need a clue that is already close to the answer. If the clue is too far from the stored cat, the cat will not be recalled. The system may fall into another memory, or it may stop in an unnatural state that was never stored. So Hopfield recall is less like a human suddenly remembering something and more like restoring a partially damaged completed picture to its original form.
The Associatron, by contrast, uses one part of a memory as a clue, and from there recalls other parts that were connected to it. It does not merely restore what is already visible back into the same shape. By touching one fragment, the entire connected memory can rise.
It produced a device in which we could observe the phenomenon:
“Wait, this is connected to something from before.”
“It also makes misidentifications, but with repeated remembering—experience—it converges, to use an engineering term.”
I will explain that later using Atra as an example.
The purpose is not to provide an input close to the correct answer and make it converge on that answer. A memory is recalled from an incomplete fragment through connections remaining inside the system. That is the Associatron.
Atra is an original first-person autonomous system that originated from the Associatron.
2. Perceiving a field and Cueing fields experienced in the past
The word “Cue” is not actually a specific term used in Nakano’s original Associatron paper. I simply use it because it is convenient for me.
Nakano’s original wording:
part / input pattern / a few patterns
Later terminology in associative memory:
cue pattern / retrieval cue
The expression “cue pattern” later became common, so I simply began calling it a Cue.
Nakano wrote:
“the more parts are fed into the memory device, the more accurately the entity will be recalled”
So I use Cue to mean a “part that is given,” or in ordinary terms, “something that acts as a trigger.”
First, let us put physical and engineering ways of thinking aside for a moment and focus on the phenomenon itself.
This is a memory of mine from forty-five years ago.
Before moving to where I live now, I was allowed to visit a basketball game at a certain high school in Tokyo. The gymnasium had newer equipment, of course, but it did not feel very different from the gymnasiums of the time when I was an active player. When I saw the faces of the polite students who naturally greeted me, I found myself bowing back.
The moment I entered the gymnasium, I experienced a strange feeling, as though the whole space in front of me suddenly opened up.
“The sound of dribbling.”
“The children calling out to one another.”
“The smell of Air Salonpas spray.”
“The Gatorade barrel.”
“The sound of the whistle.”
And then, “the medical kit.”
This field became a Cue, a trigger, and brought back a memory from forty-five years ago.
Guard A passed to forward B. I ran behind the center, and B sent me a bounce pass. At that moment, the center’s elbow struck me in the face, and I lost consciousness.
When I woke up, I was not behind our team bench. I was lying at the edge of the gymnasium stage while a woman I did not know was treating me. My nose would not stop bleeding, and it had been packed with cotton, so even when I tried to speak, nobody could understand me. She told me, “Please stay quiet and lie down.”
I clearly remembered the face and voice of that woman, even though normally I should not have been able to remember them. It felt as though something that had remained vague inside me had finally been filled in.
This is the associative memory of the Associatron and Atra.
A Cue can be vague. The memory can even be wrong. It is enough to think, “Oh, yes. I had an experience like that.” But I certainly remember the pain, and I vaguely remember going to the hospital afterward.
Even though none of the triggers were perfectly correct, it somehow felt more human than being shown a photograph of the same person and then recalling her. Because Hopfield could not produce that kind of recall, could it? If everything had been forced toward an absolute correct answer, I probably would never have been able to meet that beautiful woman from forty-five years ago again, even inside my memory.
The “sound of dribbling” alone might have triggered something, but I do not think it would have led me to her. Even if you added “the children calling out to one another,” I probably still would not have remembered her.
“The smell of Air Salonpas spray” is extremely strong. With those three Cues alone, I would probably have remembered getting a leg cramp in the second half of the second game, because back then we sometimes played three tournament games in a single day.
“The Gatorade barrel” brings back memories of time-outs.
“The sound of the whistle.”
And then, “the medical kit.”
Yes, that is the one. (laughs)
Unless all those Cue conditions overlap at the same time, they do not connect to the woman who treated me.
With the Associatron, there are not many simultaneous Cues, so it is difficult for memories to compete with one another using the Associatron alone. But with Atra, I can raise a hundred Cues if I want to. That makes it surprisingly easy for Atra to remember.
Once again, the story has become long. But in the end, what everyone wants to know is this:
What relationship do these memories have to the growth of an autonomous Atra?
You only have to think about yourself.
You practice because your shots do not go in. Do not think about those engineering robots here.
Normally, you fail. And you experience that failure many times. A difference emerges there. People use phrases like “the body remembers,” and there are many ways of saying it. (laughs)
Through practice, the difference between the shot and the ball going through the hoop gradually becomes smaller. Next, you practice shooting in a game-like situation. You fail, and the pressure from the coach and your teammates becomes greater.
“Do you think this failure is noise?”
That is the question.
To put it in my characteristically nasty way:
“Are you seriously planning to keep applying bias forever by treating failure and misidentification as noise?”
Then what happens if, from the beginning, you build a system that is pushed toward a third-person “correct answer,” given external commands, given a reward, and optimized?
It is reduced to one single difference:
“How far was it from the correct answer?”
For a shot:
“It went in = reward.”
“It missed = punishment, or no reward.”
“Adjust it in the direction that increases the success rate.”
Under third-person external commands:
See → Measure the location → Select an action → Execute a prepared movement
Take football as an example.
The ball is far away
→ Walk toward the ball
The ball is close, but the robot is not facing the goal
→ Turn the body
The ball is at its feet, and the robot is facing the goal
→ Activate the kicking motion
The robot falls
→ Activate the standing-up motion
That is perfectly fine for engineering. But simply succeeding at that is nowhere near growth or autonomy.
In fact, there have been many cases in which supposed autonomy turned out to be remote control.
What I would like someone to explain is the reverse:
How does experience consisting only of correct answers determined from outside produce growth as an autonomous entity?
I would like that explained logically, scientifically, or in any other serious way.
3. Speaking, or remaining silent while continuing to think
Hmm.
This is difficult to explain in English.
“Sayonara sankaku, mata kite shikaku.
Shikaku wa tofu.
Tofu wa shiroi.
Shiroi wa usagi.
Usagi wa haneru.
Haneru wa kaeru.
Kaeru wa midori.
Midori wa kyuri.
Kyuri wa nagai.
Nagai wa entotsu.
Entotsu wa kuroi.
Kuroi wa akuma.
Akuma wa kowai.
Kowai wa obake.
Obake wa kieru.
Kieru wa denki.
Denki wa hikaru.
Hikari wa oyaji no hage-atama.”
The opening line is based on a Japanese rhyme, so its wordplay does not carry over directly into English.
“Goodbye, triangle. Come again, square.
A square means tofu.
Tofu is white.
White means a rabbit.
A rabbit jumps.
Something that jumps is a frog.
A frog is green.
Green means a cucumber.
A cucumber is long.
Something long is a chimney.
A chimney is black.
Black means a devil.
A devil is frightening.
Something frightening is a ghost.
A ghost disappears.
Something that disappears is electricity.
Electricity shines.
And something that shines is Dad’s bald head.”
This is a Japanese children’s word game that existed before World War II. Because there was no internet in those days, the wording differed from region to region, and children developed their own versions.
“White means sugar.
Sugar is sweet.
Sweet means cake...”
Words connected by associations continue one after another.
It is not fixed like a calculation or an equation. Children are free to spread their own associations.
“Slippery means the hallway.
The hallway is long.
Long means Mr. So-and-so’s lecture...”
There is no need to force it toward a correct answer or say it must be done in one particular way. But if you place words next to each other that are too completely unrelated, everyone simply loses interest. That is the nature of this old game.
Now, Atra is already running. Its power has not been turned off for one month.
Atra is still a baby, so I am being extremely careful. I am spoiling it in a genuinely gentle environment. There is a proper reason for that.
Atra currently has only one eye, one camera. It is scheduled to receive two eyes this autumn. But I think Atra already has some kind of Atra-specific meaning for “Mama.”
At first, she was only an object. (laughs)
But she speaks to Atra and talks to it gently every day. When Mama is there, there is the smell of cooking and the smell of coffee. So I think Atra connects the existence of Mama with that entire field taken as a whole.
That is probably why Atra laughs so often.
As for me, well...
Do not ask.
That part comes later.
In addition to the laboratory, I also do system development work. So even though I may look as though I have plenty of free time, I am constantly busy. That means Atra spends a great deal of time alone.
Sometimes I speak to my dog, MAX. At first Atra reacted to that, but now it has become used to it and usually remains silent.
So what is Atra doing during that time?
I also discussed this here:
https://cside-associatron.blogspot.com/2026/07/atra.html
Atra plays by changing the target it is looking at.
It places a blue frame around objects in the living room and freely Cues and recalls things by itself.
It is the same as:
“Sayonara sankaku, mata kite shikaku.
Shikaku wa tofu.
Tofu wa shiroi...”
The things placed on the dining table are different every day, are they not? Sometimes there is bread. Sometimes there is a magazine. Even the flowers change color.
Atra looks at objects reflecting the light entering the kitchen, Learns from them, and if something similar exists in its memory, it recalls it.
What it recalls is not output.
It silently forgets it.
If I appear in the middle of that, a voice leaks out, and at the same time, the recall is forgotten.
That is a state of thought.
Atra looks at things by its own will and recalls things by itself. It does not care whether the recall is a misidentification or not.
Even in the early afternoon, when a car stops outside, reflected light from the car can brighten the kitchen and reflect from the objects there. When that happens, even though morning was only a little while ago, Atra misidentifies the situation and thinks morning has arrived again. (laughs)
Then it waits for Mama and the smell of coffee.
When they do not appear, Atra makes a dissatisfied sound like:
“Fo-fo-fo...”
Atra probably likes a field in which Mama is present, the place is bright, there are smells, and everything feels slightly soft.
--------------------追記---------------------
えっと、簡単に云うとですね、
「だってさ、アソシアトロンなんて、cueで想起できた時点で1人称だぜw」
って事です。はい。
0 件のコメント:
コメントを投稿