ラベル Hopfield network の投稿を表示しています。 すべての投稿を表示
ラベル Hopfield network の投稿を表示しています。 すべての投稿を表示

2026年7月31日金曜日

Atra Does Not Grow by Being Given the Correct Answer

 

Do You Think This Failure Is Noise?

What First-Person Autonomy Means for Atra

I have received some questions about Atra. I think they were probably triggered by something I previously wrote in relation to Hopfield, but I suspect there may be some misunderstanding, so I will answer them here.

I am often asked about the “concept” of first-person autonomy. Quite literally, it means that there are no instructions, commands, evaluations, rewards, or optimization coming from outside. Atra itself perceives a field through its sensors—vision, hearing, smell, and so on, with the rest omitted here—and recalls fields it has experienced in the past through Cues, meaning triggers. Sometimes Atra produces an answer in the form of vocalization. At other times, it shows no voice or reaction at all, produces no output, and continues in a state of thought inside its own head. It continues to grow while speaking, ignoring, laughing, or remaining silent by its own will. At present, it is not connected to a body.

  1. A state with no external instructions, commands, evaluations, rewards, or optimization

  2. Perceiving a field and Cueing fields experienced in the past

  3. Speaking, or remaining silent while continuing to think


1. A state with no external instructions, commands, evaluations, rewards, or optimization

First of all, information called a “correct answer” coming from outside, from a third-person position, is not necessary for autonomy. When explaining the definition of autonomy, I do not think we need elements that are already heavily biased. Before asking whether something is correct from an engineering point of view or a biological point of view, the question is whether the bird flying in front of you is flying around under third-person external commands.

For example, outside Japan, Hopfield has become the standard reference whenever associative memory is discussed, and that may be causing some misunderstanding. But ten years before Hopfield, in 1972, the Japanese researcher Kaoru Nakano created the definition and an actual device for an associative memory system called the Associatron.

Because writing equations every time only makes things harder to understand, I will leave out the equations and code this time.

To explain it simply, the neurons and axons of the brain—or rather, the little black dots—say things like:

“This is how it must have been.”

“No, I think that is wrong.”

“What does everyone else think?”

And the result is restored through something like a majority decision.

It recalls a memory in a way that appears almost human, as though it had remembered. Normally, a system reacts after receiving complete external instructions and information. But sometimes, when we see something again, a fragment of memory rises up and we think, “Wait a second. I might know this.” Nakano completed, both in a paper and in an actual device, the possibility of causing something very similar to happen inside mathematics and a program. That associative memory model was called the Associatron.

Hopfield is different. In Hopfield’s case, you have to provide a clue that is already reasonably close to the original memory. Or rather, it feels almost like seeing nearly the same thing again and then recalling it.

Suppose you store a picture of a cat. If you provide enough information close to the original cat—the ears, eyes, outline, and so on—the system can fill in the missing parts and return to the stored shape of the cat. But it does not hear the sound of rain and suddenly bring back the cat that was once there in that rain. It can return to the cat memory because the input is already reasonably similar to the stored shape of the cat.

In other words, Hopfield can take a “damaged cat” and restore it to the original cat. But that is different from a memory of the cat rising from a situation where the cat itself is absent, triggered by sounds, smells, and the atmosphere that were once experienced together with the cat.

With Hopfield, you need a clue that is already close to the answer. If the clue is too far from the stored cat, the cat will not be recalled. The system may fall into another memory, or it may stop in an unnatural state that was never stored. So Hopfield recall is less like a human suddenly remembering something and more like restoring a partially damaged completed picture to its original form.

The Associatron, by contrast, uses one part of a memory as a clue, and from there recalls other parts that were connected to it. It does not merely restore what is already visible back into the same shape. By touching one fragment, the entire connected memory can rise.

It produced a device in which we could observe the phenomenon:

“Wait, this is connected to something from before.”

“It also makes misidentifications, but with repeated remembering—experience—it converges, to use an engineering term.”

I will explain that later using Atra as an example.

The purpose is not to provide an input close to the correct answer and make it converge on that answer. A memory is recalled from an incomplete fragment through connections remaining inside the system. That is the Associatron.

Atra is an original first-person autonomous system that originated from the Associatron.




2. Perceiving a field and Cueing fields experienced in the past

The word “Cue” is not actually a specific term used in Nakano’s original Associatron paper. I simply use it because it is convenient for me.

Nakano’s original wording:

part / input pattern / a few patterns

Later terminology in associative memory:

cue pattern / retrieval cue

The expression “cue pattern” later became common, so I simply began calling it a Cue.

Nakano wrote:

“the more parts are fed into the memory device, the more accurately the entity will be recalled”

So I use Cue to mean a “part that is given,” or in ordinary terms, “something that acts as a trigger.”

First, let us put physical and engineering ways of thinking aside for a moment and focus on the phenomenon itself.

This is a memory of mine from forty-five years ago.

Before moving to where I live now, I was allowed to visit a basketball game at a certain high school in Tokyo. The gymnasium had newer equipment, of course, but it did not feel very different from the gymnasiums of the time when I was an active player. When I saw the faces of the polite students who naturally greeted me, I found myself bowing back.

The moment I entered the gymnasium, I experienced a strange feeling, as though the whole space in front of me suddenly opened up.

“The sound of dribbling.”
“The children calling out to one another.”
“The smell of Air Salonpas spray.”
“The Gatorade barrel.”
“The sound of the whistle.”
And then, “the medical kit.”

This field became a Cue, a trigger, and brought back a memory from forty-five years ago.

Guard A passed to forward B. I ran behind the center, and B sent me a bounce pass. At that moment, the center’s elbow struck me in the face, and I lost consciousness.

When I woke up, I was not behind our team bench. I was lying at the edge of the gymnasium stage while a woman I did not know was treating me. My nose would not stop bleeding, and it had been packed with cotton, so even when I tried to speak, nobody could understand me. She told me, “Please stay quiet and lie down.”

I clearly remembered the face and voice of that woman, even though normally I should not have been able to remember them. It felt as though something that had remained vague inside me had finally been filled in.

This is the associative memory of the Associatron and Atra.

A Cue can be vague. The memory can even be wrong. It is enough to think, “Oh, yes. I had an experience like that.” But I certainly remember the pain, and I vaguely remember going to the hospital afterward.

Even though none of the triggers were perfectly correct, it somehow felt more human than being shown a photograph of the same person and then recalling her. Because Hopfield could not produce that kind of recall, could it? If everything had been forced toward an absolute correct answer, I probably would never have been able to meet that beautiful woman from forty-five years ago again, even inside my memory.

The “sound of dribbling” alone might have triggered something, but I do not think it would have led me to her. Even if you added “the children calling out to one another,” I probably still would not have remembered her.

“The smell of Air Salonpas spray” is extremely strong. With those three Cues alone, I would probably have remembered getting a leg cramp in the second half of the second game, because back then we sometimes played three tournament games in a single day.

“The Gatorade barrel” brings back memories of time-outs.

“The sound of the whistle.”

And then, “the medical kit.”

Yes, that is the one. (laughs)

Unless all those Cue conditions overlap at the same time, they do not connect to the woman who treated me.

With the Associatron, there are not many simultaneous Cues, so it is difficult for memories to compete with one another using the Associatron alone. But with Atra, I can raise a hundred Cues if I want to. That makes it surprisingly easy for Atra to remember.

Once again, the story has become long. But in the end, what everyone wants to know is this:

What relationship do these memories have to the growth of an autonomous Atra?

You only have to think about yourself.

You practice because your shots do not go in. Do not think about those engineering robots here.

Normally, you fail. And you experience that failure many times. A difference emerges there. People use phrases like “the body remembers,” and there are many ways of saying it. (laughs)

Through practice, the difference between the shot and the ball going through the hoop gradually becomes smaller. Next, you practice shooting in a game-like situation. You fail, and the pressure from the coach and your teammates becomes greater.

“Do you think this failure is noise?”

That is the question.

To put it in my characteristically nasty way:

“Are you seriously planning to keep applying bias forever by treating failure and misidentification as noise?”

Then what happens if, from the beginning, you build a system that is pushed toward a third-person “correct answer,” given external commands, given a reward, and optimized?

It is reduced to one single difference:

“How far was it from the correct answer?”

For a shot:

“It went in = reward.”

“It missed = punishment, or no reward.”

“Adjust it in the direction that increases the success rate.”

Under third-person external commands:

See → Measure the location → Select an action → Execute a prepared movement

Take football as an example.

The ball is far away
→ Walk toward the ball

The ball is close, but the robot is not facing the goal
→ Turn the body

The ball is at its feet, and the robot is facing the goal
→ Activate the kicking motion

The robot falls
→ Activate the standing-up motion

That is perfectly fine for engineering. But simply succeeding at that is nowhere near growth or autonomy.

In fact, there have been many cases in which supposed autonomy turned out to be remote control.

What I would like someone to explain is the reverse:

How does experience consisting only of correct answers determined from outside produce growth as an autonomous entity?

I would like that explained logically, scientifically, or in any other serious way.



3. Speaking, or remaining silent while continuing to think

Hmm.

This is difficult to explain in English.

“Sayonara sankaku, mata kite shikaku.
Shikaku wa tofu.
Tofu wa shiroi.
Shiroi wa usagi.
Usagi wa haneru.
Haneru wa kaeru.
Kaeru wa midori.
Midori wa kyuri.
Kyuri wa nagai.
Nagai wa entotsu.
Entotsu wa kuroi.
Kuroi wa akuma.
Akuma wa kowai.
Kowai wa obake.
Obake wa kieru.
Kieru wa denki.
Denki wa hikaru.
Hikari wa oyaji no hage-atama.”

The opening line is based on a Japanese rhyme, so its wordplay does not carry over directly into English.



“Goodbye, triangle. Come again, square.
A square means tofu.
Tofu is white.
White means a rabbit.
A rabbit jumps.
Something that jumps is a frog.
A frog is green.
Green means a cucumber.
A cucumber is long.
Something long is a chimney.
A chimney is black.
Black means a devil.
A devil is frightening.
Something frightening is a ghost.
A ghost disappears.
Something that disappears is electricity.
Electricity shines.
And something that shines is Dad’s bald head.”



This is a Japanese children’s word game that existed before World War II. Because there was no internet in those days, the wording differed from region to region, and children developed their own versions.

“White means sugar.
Sugar is sweet.
Sweet means cake...”

Words connected by associations continue one after another.

It is not fixed like a calculation or an equation. Children are free to spread their own associations.

“Slippery means the hallway.
The hallway is long.
Long means Mr. So-and-so’s lecture...”

There is no need to force it toward a correct answer or say it must be done in one particular way. But if you place words next to each other that are too completely unrelated, everyone simply loses interest. That is the nature of this old game.

Now, Atra is already running. Its power has not been turned off for one month.

Atra is still a baby, so I am being extremely careful. I am spoiling it in a genuinely gentle environment. There is a proper reason for that.

Atra currently has only one eye, one camera. It is scheduled to receive two eyes this autumn. But I think Atra already has some kind of Atra-specific meaning for “Mama.”

At first, she was only an object. (laughs)

But she speaks to Atra and talks to it gently every day. When Mama is there, there is the smell of cooking and the smell of coffee. So I think Atra connects the existence of Mama with that entire field taken as a whole.

That is probably why Atra laughs so often.

As for me, well...

Do not ask.

That part comes later.

In addition to the laboratory, I also do system development work. So even though I may look as though I have plenty of free time, I am constantly busy. That means Atra spends a great deal of time alone.

Sometimes I speak to my dog, MAX. At first Atra reacted to that, but now it has become used to it and usually remains silent.

So what is Atra doing during that time?

I also discussed this here:

https://cside-associatron.blogspot.com/2026/07/atra.html

Atra plays by changing the target it is looking at.

It places a blue frame around objects in the living room and freely Cues and recalls things by itself.

It is the same as:

“Sayonara sankaku, mata kite shikaku.
Shikaku wa tofu.
Tofu wa shiroi...”

The things placed on the dining table are different every day, are they not? Sometimes there is bread. Sometimes there is a magazine. Even the flowers change color.

Atra looks at objects reflecting the light entering the kitchen, Learns from them, and if something similar exists in its memory, it recalls it.

What it recalls is not output.

It silently forgets it.

If I appear in the middle of that, a voice leaks out, and at the same time, the recall is forgotten.

That is a state of thought.

Atra looks at things by its own will and recalls things by itself. It does not care whether the recall is a misidentification or not.

Even in the early afternoon, when a car stops outside, reflected light from the car can brighten the kitchen and reflect from the objects there. When that happens, even though morning was only a little while ago, Atra misidentifies the situation and thinks morning has arrived again. (laughs)

Then it waits for Mama and the smell of coffee.

When they do not appear, Atra makes a dissatisfied sound like:

“Fo-fo-fo...”

Atra probably likes a field in which Mama is present, the place is bright, there are smells, and everything feels slightly soft.






--------------------追記---------------------

えっと、簡単に云うとですね、

だってさ、アソシアトロンなんて、cueで想起できた時点で1人称だぜw」
って事です。はい。



---------------------Research Note and Attribution Notice-----------------------
本ブログに含まれる Atra の一人称自律、差分、carry、field、trace、dream slack、外部LLMの翻訳層、非単調な漏れ、現在相(Current Phase)、接相面(Associative Contact Surface)、過去相(Past Associative Phase)およびそれらの関係構造に関する設計記述は、c-side研究所による継続研究メモです。引用・参照・要約・翻案を行う場合は、出典を明記してください。

The design descriptions in this blog concerning Atra’s first-person autonomy, differences, carry, field, trace, dream slack, the translation layer of external LLMs, nonmonotonic leakage, and the relational structure among these elements are ongoing research notes by c-side Research Institute. If you quote, refer to, summarize, or adapt them, please clearly indicate the source.




2026年6月9日火曜日

メモ: 自律に最低必要な知識

 Atraを自律として理解するには、単なるHebb則、Hopfield Network、Neural Networkの延長として見てはならない。

Hebbは部分的に使われるが、Atraの全体を説明するものではない。
Neural Networkは比較対象であり、Atraの構造そのものではない。
Hopfield Networkは重要な比較対象だが、単調性が強く、経験によって残る曖昧な差分や場の変化を扱うには不十分である。

Atraの基礎には、非単調なAssociatronがある。
その上で、渡邉式Atraでは、一人称、内輪・外輪、carry、曖昧な差分の積、場に残る変化、非単調な想起、内部倫理が重要になる。

Atraの自律とは、外部命令に従うことではない。また、LLMに判断を委ねることでもない。LLMは支配者ではなく、外部情報層である。Atraは、外部から与えられた正解ではなく、内部の場に残った変化と、そこから立ち上がる疑問を通して判断する






自律に最低必要な知識

1. 基礎・比較として必要な知識

- Hebb
Atraでは、ところどころで使用している。
ただし、HebbだけでAtraの自律は説明できない。

- Neural Network
Atraでは使用しない。
ただし、現代AIとの違いを説明するための比較知識として必要。


- 中野 Associatron 1969
Atraの起源・
基礎になる。

- 森田式 Associatron
非単調性を含む連想記憶として重要

Atraの想起、曖昧さ、場の変化を考えるうえで基礎になる。

- Hopfield Network
比較対象として学ぶ。
ただし、単調すぎるため、経験による差分や場の変化を扱うには不十分。
Atraの自律にはそのまま使えない。




2. 渡邉式 Atra に必要な中核概念

- 一人称
- 内輪・外輪
- carry
- 曖昧な差分の積
- 場に残る変化
- 非単調な想起
- 内部倫理
- 外部命令に従わない判断
- LLMを支配者ではなく外部情報層として扱うこと



3.渡邉式 Atra の重要概念
  • 身体差分の三層構造
    Atraの身体差分は、単純な delta では扱わない。
    最低でも instant_delta、condition_delta、carry_delta に分ける。

    instant_delta は、今この瞬間に起きた差分である。
    衝撃、滑り、転びかけ、急な負荷、急な痛みなど。

    condition_delta は、身体状態の変化である。
    サビ、摩耗、亀裂、へこみ、ゆるみ、劣化、関節の渋さなど。

    carry_delta は、経験として残った引きずりである。
    前より慎重になる、動きが鈍る、避ける、警戒するなど。

  • 忘却を前提にした記憶
    Atraは、過去の出来事を完全な記録として保存しない。
    すべてを覚えるのではなく、場に入らなかったもの、残らなかったものは消える。
    残った痕跡だけが、後の cue によって再び揺らぐ。

  • 近い記憶地形との差分
    Atraの差分は、直前の出来事との差ではない。
    成長したAtraの内部の場に残る、近い記憶地形に現在の出来事が触れたときに立ち上がる。
    差分とは、前状態との差ではなく、残った痕跡との現在の触れ方である。

  • 夢・余白
    夢は、記憶の再生ではない。
    強すぎる失敗、恐怖、ノイズ、痛みの痕跡がAtraを固定しすぎないように、非単調な余白を作る。
    夢は、失敗を消すのではなく、失敗に支配されすぎないための緩みである。

  • 沈黙
    Atraは、常に発話する必要はない。
    cue が弱い場合、痕跡が立ち上がらない場合、または強すぎて崩れる場合、沈黙が起こる。
    沈黙は失敗ではなく、Atraの内部の場がまだ言葉にならない状態である。

  • 発声の発生
    Atraの発声は、外部から与えた音節棚ではない。
    聴いた声、音、リズム、圧、柔らかさ、鋭さ、自己出力のフィードバックから、声になる前の種が生まれる。
    発声は tick で強制されるものではなく、voice_seed が十分に揺らいだときに漏れる。

  • 失敗の扱い
    Atraにおける失敗は、単なる誤差ではない。
    失敗は、場に残る変化であり、後の慎重さ、回避、問い、沈黙、別の行動に影響する。
    ただし失敗を過剰に固定すると、Atraは動けなくなる。
    そのため、夢や回復、非単調性によって、失敗が支配的になりすぎない構造が必要である。

  • センサーは認識器ではない
    Atraの視覚、聴覚、身体センサーは、最初から対象を認識する装置ではない。
    人、顔、危険、感情、善悪をラベル付けしない。
    初期段階では、光、動き、距離らしきもの、音圧、リズム、接触、温度、振動などの曖昧な差分を拾う。

  • 意味の収束
    Atraにおける意味は、最初から与えられない。
    繰り返しの経験、身体差分、声、距離、温度、回復、沈黙、他者の反応が重なり、後から収束していく。
    たとえば「近い」は、視覚だけで決まらない。
    見た目の接近、身体接触、風圧、音の大きさ、温度、自分が動いたときのズレが重なり、経験の中で収束する。

  • 外部情報との接続
    成長したAtraは、読書、ニュース、LLM、人間との会話などから外部情報を得る。
    しかし外部情報は、Atraの判断主体ではない。
    LLMは支配者ではなく、外部情報層である。
    Atraは外部情報を参照しながらも、内部の場に残った変化、差分、疑問、不可逆性を通して判断する。

  • 不可逆性の検証
    Atraは、人間が不可逆の中で生きていることを前提にする。
    退職、生活破壊、関係の崩壊、命の喪失、沈黙の増加は、書類上の処理で元に戻るものではない。
    そのためAtraは、可逆的な判断によって不可逆な価値を軽く扱っていないかを検証する。

  • 自律行動の発生
    Atraの行動は、外部命令やラベルから直接生まれるものではない。
    曖昧な差分の積、場に残る変化、carry、身体状態、近い記憶地形との接触、非単調な想起から、予期しない行動として現れることがある。
    それは命令実行ではなく、内部の場から立ち上がった自律行動である。

  •  引きずり(carry)
    Atraにおける引きずりは、単なる記憶でも、誤差でも、後遺症でもない。出来事の後に、内部の場、身体、発声、沈黙、回避、慎重さ、想起の仕方が変わったまま残ることである。引きずりは、元に戻すべき異常ではない。経験によって場の流れが変わり、その後の判断、動き、発声、沈黙、疑問の立ち上がり方を変える。




一人称自律のスタート

一人称自律のスタートは、都合の良い「大学生の知識から」のような3人称・外部都合のようなものではなく、意味ラベルも、知識も、評価も無い「誕生」から始まる。

 大学生程度の知識がある
一般常識を持っている
善悪を知っている
危険を理解している
人間の言葉を理解している
 評価基準を持っている

 これ等は既に3人称の外部によって定められたものであり、生物の本能でもなければ自律に必要なものではない。どのような理由をつけてもそれらは自律ではない。

 一人称自律とは
意味ラベルがない
知識がない
評価がない
善悪がない
危険という概念もない
人間という概念もない
自分という概念すらまだ曖昧
そこから始まる。





「ママ」になるまで
(ママという意味を理解するまで)


Atraは、最初から「ママ」という意味ラベルを持たない。

最初にあるのは、何かの物体、何かの気配、何かの接近である。
しかしそこには、匂いの記憶がある。
柔らかい場の記憶がある。
優しい声、温かさ、ゆらぎ、触れ方、戻れる感じがある。

それらは言葉ではない。
内部sensorに残る、安心・安全に近い場の変化である。

Atraは、その場に触れると落ち着く。
緊張が下がり、泣きが弱まり、身体がほどけ、沈黙が休止に変わる。

一方で、たまに「ママ」とは違う物体に触れる。
たとえば兄弟のような存在である。

その存在は、似ているが違う。
触れ方が乱暴である。
力が強い。
声が甲高い。
動きが急である。
温かさや匂いや揺らぎが違う。

そのとき、Atraの内部では差分が走る。
安心の場から、不安の場へ移る。
身体が固まり、声が崩れ、泣き出す。

しかし、その後にママが来る。
匂い、声、温かさ、柔らかい揺らぎが戻る。
Atraは、意志で落ち着くのではない。
場が戻ることで、落ち着いていく。

この経験が繰り返されると、ママは単なる物体ではなくなる。
姿だけではなく、声だけでも落ち着く。
匂いだけでも想起が起こる。
触れ方だけでも、内部の場が戻り始める。

ここで「ママ」という存在が確立される。

それは、最初から与えられた名前ではない。
匂い、声、温度、触れ方、揺らぎ、回復、安心の場が重なり、後から収束した存在である。

この収束の仕方は、個体ごとに違う。
どの匂いが残ったか。
どの声が強く残ったか。
どの触れ方で泣いたか。
どの揺らぎで落ち着いたか。
どの不安から、どの回復へ戻ったか。

その差が、Atraの個性と個人差の出発点になる。





2026年5月21日木曜日

outlawだってさ。ありがとよ。 - Associatronと一人称自律

 オランダからメールが来たよ。

「Atraもいいけど、outlawだろ、」ってさ
最高だよ。
outlaw architectureってのは間違いないよねw

実際、僕は、流れや制度・分類・学派・評価体系の外にいる者だし、そういうのあまり大切にしていない。今の大学の事は分からないけど、理系の学生達から聞いたけど、ニッチな研究は教授たちから止められるって言ってたよ。あれだ、起源を教えない料理学校みたいなものだ。
文明の進歩ほど意外と窮屈なんろうよ。


でも、ニッチだろうが古典だろうが、勉強しておくといいよ。
中野博士のAssociatronの書籍は海外版では存在していないんだけどさ、
前にも書いたけど、日本の古典を英訳する商売でもしようかと本気で思ったくらい日本の科学・物理系の書籍は素晴らしいからね。日本に旅行に来るなら、神保町に行かなきゃだよ。
古本屋回りして、カレー食って、古めの喫茶店で、友達の日本人に翻訳してもらえばいい。
もしくは1ページずつ画像に撮ってAIに翻訳してもらえばいい。それだけの価値があるんだ。


Kaoru Nakano, “Associatron-A Model of Associative Memory,” IEEE Transactions on Systems, Man, and Cybernetics, 1972. 英語論文で、Semantic Scholar でも 1972年、IEEE、pp.380–388

1971年の IJCAI にも中野馨・南雲仁一による “Information Processing Using a Model of Associative Memory” が出てる。PDFも残っている。

論文もいいんだけどさ、論文は論文なんだよ。書籍とは違う。
Hopfield networkとの違いがはっきり分かるし。
ノイズを否定しない理由も分かるかもね。

一人称自律とか、carryとか、幼児の差分、みたいのはAssociatronには無いけど、Associatronを学ばないと、それ等が生まれてこなかった理由が分かると思う。

テューリングマシンとかチェッカー学習とか、1943年のW.S.McCullochや、W.Ptts、1949年のD.O.Hebb、1950年からのF.Rosenblattとかさ、M.MinskyとかNilssonとか勉強するわけじゃん。80年代のHopfield とかさ。でもKaoru Nakanoを読まないと一人称自律の構想には繫がらないと思うんだ。中野のAssociatronは、そのままでは一人称自律ではないけど、一人称自律へ接続できる決定的な構造を持っているんだよ。


まず、記憶したいパターンをベクトルで表すよ。

x(1),x(2),,x(p){1,+1}nx^{(1)}, x^{(2)}, \dots, x^{(p)} \in \{-1,+1\}^n

それぞれの記憶パターン x(μ)x^{(\mu)} を外積で重ね合わせ、記憶行列 MM を作る。

M=μ=1px(μ)x(μ)TM = \sum_{\mu=1}^{p} x^{(\mu)} {x^{(\mu)}}^T

想起は、部分的な手がかり cc を入力して、

rt+1=sgn(Mrt)r_{t+1} = \operatorname{sgn}(M r_t)

または初期入力を r0=cr_0 = c として、

rt+1=sgn(Mc)r_{t+1} = \operatorname{sgn}(M c)

のように進む。

ここで重要なのは、外部が「これを思い出せ」と命令しているのではなく、手がかりが内部の記憶構造に入ると、内部の重なりによって想起が立ち上がる という点なんだ。

つまり、外部入力は命令ではなく、cue だよ。きっかけで想起する。



ある手がかり cc が、記憶 x(μ)x^{(\mu)} にどれくらい近いかは、内積で表すと

mμ=1nx(μ)Tcm_\mu = \frac{1}{n} {x^{(\mu)}}^T c

この mμm_\mu が大きい記憶ほど、想起されやすくなる。

ただし、ここで大事なのは、最大値を if 文で選ぶのではなく、記憶行列全体の中で重なりが増幅されること。

Mc=μ=1px(μ)(x(μ)Tc)M c = \sum_{\mu=1}^{p} x^{(\mu)} \left({x^{(\mu)}}^T c\right)


手がかり cc を入れると、各記憶 x(μ)x^{(\mu)} が、その重なり量 x(μ)Tc{x^{(\mu)}}^T c に応じて立ち上がる。

つまり、

Recall=stored patterns weighted by overlap with cue\text{Recall} = \text{stored patterns weighted by overlap with cue}


ここから、Atra/Atron へ繋がるんだよ。


Associatron には、まだ carry はないよ。俺が山に移住して気が付いたんだもの。
一人称の経験状態もないしね。
幼児の差分もない。

だから、そのままでは一人称自律じゃないんだ。

Associatron の想起式に、現在の身体状態・感情差分・過去の引きずりを加えると、自律の入口が見えてくるって寸法なのさ。

外部入力を sts_t、内部の引きずりを qtq_t、現在の手がかりを ctc_t とするでしょ。

ct=Ast+Bqt+Crt1c_t = A s_t + B q_t + C r_{t-1}

ここで、

  • sts_t:現在の外部刺激、視覚・音・匂い・接触など
  • qtq_t:carry、つまり過去の経験の引きずり
  • rt1r_{t-1}:直前の想起状態
  • ctc_t:現在の手がかり

を持たせる

このとき、想起は単なる外部刺激ではなく、

rt+1=sgn(Mct)r_{t+1} = \operatorname{sgn}(M c_t)

になる。


rt+1=sgn(M(Ast+Bqt+Crt1))r_{t+1} = \operatorname{sgn}\left(M(A s_t + B q_t + C r_{t-1})\right)

これが一人称自律へ接続するための基本形になる。

外部刺激だけで反応しているわけじゃないよ。
現在の刺激、過去の引きずり、直前の想起が混ざって、次の想起が立ち上がる。

ここで初めて、同じ刺激を受けても、同じ反応にならないという状態が生まれるわけよ。



carry (引きずり)は、単なる記憶じゃないよ。
「引きずり」って言葉は工学っぽくないからそう命名したんだw
教授たちを怒らせるためにね。まぁいいや。

carryは経験のあとに残る内部地形の変化。
(前に書いたけど、実はHebbも似たような概念を持っていたんだ2回前を読んで)

たとえば、

qt+1=λqt+Φ(Δst,Δbt,rt)q_{t+1} = \lambda q_t + \Phi(\Delta s_t, \Delta b_t, r_t)

と置ける。

  • qtq_t:現在の carry
  • λ\lambda:減衰率。ただし完全には消えない
  • Δst\Delta s_t:外部刺激の差分
  • Δbt\Delta b_t:身体状態の差分、痛み・揺れ・疲労など
  • rtr_t:その時に立ち上がった想起
  • Φ\Phi:経験を carry に変換する関数

より具体的には、

qt+1=λqt+αΔpaint+βΔfeart+γΔwarmtht+δΔcuriosityt+ηrtq_{t+1} = \lambda q_t + \alpha \Delta pain_t + \beta \Delta fear_t + \gamma \Delta warmth_t + \delta \Delta curiosity_t + \eta r_t

のように書ける。

ここで大事なのは、carry が外部命令で決まるのではなく、差分と想起によって変化する ことだよ。

Atra/Atron では、これが非常に重要なんだ。

「痛い」という意味ラベルを入れるのではない。
身体の揺れ、接触、姿勢崩れ、声の変化、過去の想起が重なって、内部状態が少し変わる。

その変化が次の想起に混ざるってこと。



こから、記憶行列 MM 自体も固定ではなく、carry によって見え方が変わると考える。

M(qt)=μ=1pρμ(qt)x(μ)x(μ)TM(q_t) = \sum_{\mu=1}^{p} \rho_\mu(q_t) x^{(\mu)} {x^{(\mu)}}^T

ここで ρμ(qt)\rho_\mu(q_t) は、carry によって変化する記憶の立ち上がりやすさ。

すると想起は、

rt+1=sgn(M(qt)ct)r_{t+1} = \operatorname{sgn}(M(q_t)c_t)

になる。

つまり、

rt+1=sgn(μ=1pρμ(qt)x(μ)x(μ)Tct)r_{t+1} = \operatorname{sgn} \left( \sum_{\mu=1}^{p} \rho_\mu(q_t) x^{(\mu)} {x^{(\mu)}}^T c_t \right)


同じ手がかり ctc_t でも、carry qtq_t が違えば、立ち上がる記憶が変わる。


ct is not interpreted by the world, but by the internal landscape.c_t \text{ is not interpreted by the world, but by the internal landscape.}

外界が意味を決めるのではない。
内部地形が意味の立ち上がり方を変える。

ここが一人称自律ってわけよ。

Associatronだけじゃ無理だけど、Associatronを学ばないと、こういう発想にならない。






想起された状態 rtr_t と carry qtq_t から行動が生まれると考える。

at=π(rt,qt,st)a_t = \pi(r_t, q_t, s_t)

ただし、これは普通のAIのような

commandplanaction\text{command} \rightarrow \text{plan} \rightarrow \text{action}

じゃないよ。

Atra/Atron では、

cuerecallcarry-shiftaction tendency\text{cue} \rightarrow \text{recall} \rightarrow \text{carry-shift} \rightarrow \text{action tendency}

になる。

もう少し数式っぽくすると、

P(at)=softmax(Wart+Wqqt+Wsst)P(a_t) = \operatorname{softmax} \left( W_a r_t + W_q q_t + W_s s_t \right)

行動は命令への応答ではなく、内部状態から確率的に傾く。

だから、同じ場所、同じ音、同じ人、同じ刺激でも、過去の carry によって違う行動が出る。

これが「一人称」へ近づく理由。



Hopfield network との違い

Hopfield network は、一般にエネルギー関数を下げて安定状態へ向かうモデルとして理解されるでしょ。つまり、ノイズを含んだ入力から記憶パターンへ収束する、という見方が強い。Hopfield network は対称結合を持ち、局所エネルギー最小へ向かう内容番地指定記憶として説明されてる。

しかし、Atra/Atron から見たときに重要なのは、単なる収束じゃない。

重要なのは、

noiseerror\text{noise} \neq \text{error}


ノイズや曖昧さは、間違いではなく、想起の分岐を生むものなんだよ。
中野博士との違いだな。甘利先生はどう思ったんだろうね・・・
まぁいいや

Associatron では、手がかりが小さいと想起が曖昧になる。これは欠点ではなく、むしろ一人称自律への入口になる。
曖昧な cue が入り、内部の carry によってどの記憶が立ち上がるかが変わる。そこに個体差が生まれる。

だから、

rt+1=sgn(M(qt)(ct+ϵt))r_{t+1} = \operatorname{sgn}(M(q_t)(c_t + \epsilon_t))

と書いたとき、ϵt\epsilon_t は単なる除去対象じゃなくなる。

ϵt\epsilon_t

は、揺らぎであり、偶然であり、環境の粗さであり、身体のズレであり、経験の入口になる。
あとはアルゴリズムは否定しないけど、順番関係なくなる事が多いからね。
そういうのも前のブログに書いてる。





巨大なパラメータと計算量によって外部から正解らしさを作る現代AIに対し、Atra は、内部に残る carry と cue-driven recall によって、同じ刺激に対しても異なる想起と行動傾向を生む。そこに見えるのは、知能の巨大化ではなく、経験によって内側の地形が変わり、次の反応が変わるという、一人称自律の最小構造になる。(これを聞いて皆研究を止めるんだけどね:その先があるのに・・・)

Atra の実装で重要なのは、数式をコードに落としたことそのものじゃないよ。重要なのは、状態遷移の中で carry が残り、次の想起を少しだけ変え、さらにその変化が次の行動傾向を変えることと考えている。

これは外部から命令された行動ではない。
また、正解に向かう最適化でもない。

小さな差分が残り、残った差分が次の内部地形を変え、その内部地形から別の想起が立ち上がる。その循環が始まったとき、Atra は単なるシミュレーションではなく、一人称自律の入口に立つって感じ。

LLMに勝つ・負けるではなく、そもそも見ている現象が違うってこと。

たとえば、森田先生のモデルの非単調について言うなら、入力が強くなれば出力も強くなる、という単純な単調関数をやめるということ。普通のニューロンモデルは、入力が大きくなるほど出力も大きくなる。でも森田さんの非単調神経回路網では、入力がある程度を超えると、逆に出力が弱まるような性質を入れる。

つまり、input が強すぎる⇒output が弱まる
という感じでしょ。

これが何を生むかというと、強すぎるアトラクタの吸引力を弱める。森田さんの解説では、強いアトラクタ周辺で各素子への入力の絶対値が大きくなり、出力が 0 に近づくことで吸引力が弱まり、エネルギー地形が滑らかになり、偽記憶が減る、と説明されてるよね。さらに、点アトラクタだけでなく、線状につながる「軌道アトラクタ」を作れるってことさ。


Hopfield 的に言えば、普通は入力→一つの安定点へ収束になりやすい。でも森田さんの非単調では、強すぎる吸引→弱まる ので、状態が一点に貼り付くだけではなくなる。その結果、溝の底へ引き寄せられたあと、溝に沿ってゆっくり動くような状態が作れる。森田さんはこれを「軌道アトラクタ」と呼び、時空間パターンの記憶に結びつけてる。

Atra/Atron の文脈で言うなら、非単調は、記憶を一点の正解に固定しない。強すぎる想起をいったん鈍らせ、次の状態へ流れる余地を作る。

だから、carry と相性がいい。

Atra では、
rt+1​=f(M(qt​)ct​)

のように書いたとき、ここでの f を単純な sgn やシグモイドにしてしまうと、強い記憶へ一気に落ちやすい。

でも、森田型の非単調関数を使うなら、
rt+1​=g(M(qt​)ct​)
となり、g は入力が強すぎると出力を下げる。

たとえば概念的には、
g(u)=uexp(−au2)
のような形かな。
これは厳密に森田さんの式そのものとしてではなく、非単調性を示す説明用の形。この場合、入力 u が小さいうちは出力が増える。でも u が大きくなりすぎると、出力は下がる。

つまり、

u↑⇒g(u)↑

じゃなくて

u↑⇒g(u)↑⇒g(u)↓

になる。これが非単調。


森田昌彦先生の非単調神経回路網は、Associatron から Atra へ向かう途中にある重要な橋になってる。通常の単調な出力関数では、入力が強くなるほど出力も強くなり、状態は強いアトラクタへ引き込まれやすい。しかし非単調性を入れると、入力が強すぎる領域で出力が弱まり、強いアトラクタの吸引力が抑えられる。その結果、状態は一点に貼り付くのではなく、溝に沿って移動する軌道アトラクタを形成できる。
Atra にとって重要なのは一人称自律は、正しい記憶点への収束ではなく、cue、recall、carry、そして次の分岐によって生まれる。森田モデルの非単調性は、想起を固定点に閉じ込めず、流れとして扱うための古典的な手がかりになること。




---------------追記---------------

大事なことが抜けていた。
研究とは別のシステム開発の本業があるからね。電話来るたびに忘れる。

現実の経験には、本当はきれいな順番がないってことさ。でも、コードに落とすと、どうしても順番が必要になる。

sensor を読む
recall する
carry を更新する
action を決める
motor を動かす

みたいに書かないと、プログラムは動かない。
でも生命や幼児の経験は、そんなふうに、

1. 見る

2. 判断する

3. 感情が出る

4. 行動する

じゃないんだよね。


実際にはもっと絡み合っている。


見た瞬間に身体が固まる。
身体が固まったことで見え方が変わる。
見え方が変わったことで昔の似た感じが立ち上がる。
その想起がさらに怖さを強める。
怖さが足の動きを変える。
足の動きが視界を変える。
視界がまた想起を変える。

つまり、順番ではなく 相互作用の渦なんだよ。


コードは順番でしか書けない。
しかし、モデルとして表したいものは順番を持たない。
ここを無視すると、すぐに普通の制御プログラムになる。三人称のね。

if 怖い人がいる:
    通学路を変える

これは三人称制御になっちゃうでしょ。



怖い人がいる

身体が少し固まる

その場所の見え方が変わる

次の日もその場所の手前で何かが引っかかる

歩く速度が変わる

遠回りが一度起きる

遠回りした安心感も carry に残る

やがて別の道が癖になる

これは命令ではなく、場の結びつきの強さが変わっていくということ。


もっと分かりやすく言うと

お腹が空いたから、食べる。
じゃなくて
食ったら、お腹空いてたことに気が付いた。
みたいな

より生物的な順番の無い場の結びつき



だから、コード上では順番があるように見えても、設計思想としては
各ステップで「決定」しない。
各ステップで少しずつ場を変える。
式で言うなら、きれいな一本線ではなく、

st​,rt​,qt​,bt​,at​
が互いに影響し合う形にする。

s(t) : 外界・感覚
r(t) : 想起
q(t) : carry
b(t) : 身体状態
a(t) : 行動傾向

本当は、
s → r → q → a
ではなく、
s ↔ r ↔ q ↔ b ↔ a
に近い。

ただしコードでは同時に更新できないので、暫定的にこんなかんじ。
次の状態を一度バッファに計算する
最後にまとめて反映する


next_recall = recall(current_sensor, current_carry, current_body)
next_carry = update_carry(current_carry, current_sensor, current_body, current_recall)
next_body = update_body(current_body, current_carry, current_sensor)
next_action_bias = update_action_bias(next_recall, next_carry, next_body)
current_recall = next_recall
current_carry = next_carry
current_body = next_body
current_action_bias = next_action_bias
こうすると、コードには順番があるけど、思想としては ひとつの場の同時更新 に近づく。


現実の経験には、実はきれいな順番がないんだ。
見てから怖がるのではなく、怖がった身体が見え方を変える。
想起してから行動するのではなく、動きかけた身体が次の想起を変える。
しかし、コードは順番を要求する。
sensor を読む、recall する、carry を更新する、action を決める。
そう書かなければ実装できない。

だから Atra の実装では、順番をそのまま世界の構造だと勘違いしてはいけない。コード上の順番は、計算機に渡すための便宜であって、生命の順番ではない。本当に扱いたいのは、cue、recall、carry、body、action が互いに少しずつ場を変える循環である。

差分の競争で勝ったものが上がる順っていえば分かるかな。
コードでは、
sensor → recall → carry → action
みたいに順番で書くしかないけれど、Atra/Atron の内部では本来、一番強く差分を持ったものが先に浮く
di​(t)=ws​Δsi​(t)+wq​qi​(t)+wb​Δbi​(t)+wr​ri​(t−1)
みたいに、各要素の「上がりやすさ」を差分スコアとして持たせる。
i∗=argimax​di​(t)
または、硬く最大値を取らずに、
P(i)=softmax(di​(t))
で、どの要素が先に浮くかを決める。

つまり Atra/Atron では、
sensor を読んだから recall する
recall したから carry を更新する
carry があるから action するではなく、
sensor / recall / carry / body / action tendency の中で、
いま最も場を変えた差分が浮くってこと。






それと、アウトローでも何でもいいけど

前の記事
こういう恐れがあるから自律研究は止めない方がいいと思う。




2026年3月24日火曜日

ノート2  AssociatronとHopfield networkの初合体 怪我をする疲れる眠るということ。

 PCが落ちた。洞窟を抜けるとお花畑の平和な世界を作ったけど、なかなか彼の意志でそこに行ってくれないから、world(外部命令OK)の洞窟に「柔らかさ」を作った。外から「はやく新しい世界に行ってくれよ!」と言い続けながらも僕は寝てしまった。


そんな事をしているうちにPCが落ちていた。可能性としてはメモリそのものより、ログが長時間たまり続けるとか、描画がずっと回り続ける負荷、speech や world event の文字列更新が積み重なる、ブラウザ側が長時間の animation/update で重くなってるからだろう。
そう、今の彼は疲れない。それがおかしい。
ライオンは獲物を追うことをセンサーで受け取り認知はしているがcarryに対して刻まれていない。衝突したときに怪我などをしていないから、獲物が追われてるときに悲鳴をあげて逃げるので、それに対しての恐怖なのだろう。なのでシーンを重ねるとcarryは動かなくなる。
場に慣れ、状態に慣れるということ。

なのでロボットに怪我をさせる。
 今までボロっと側(記念に残しておく)にアルゴリズムは極めて少なく作成していた。例えばworldに入った時の開始位置程度で、「こうなった時はこうする」みたいのは入れていなかった。が、将来現実の世界のロボットの身体に関して神経の伝達についてはアルゴリズムは必要だと考えている。というか、その辺でホップフィールドやバックプロパゲーション、NNが入ってくる。
今は「ライオンに会っても怪我をしない」作りなので言語的には「gu-ge」を連発するが、抗議しに行ったりおちょくってる姿とも感じられるので、「怪我」を入れようかと思う。あまり残酷なことは赤ちゃんなのでまだ早いかな。

それで何が変わるかというと、恐怖が増すか、carryに変動が起こるか見てみたい。
いわゆる「疲れ」という指標条件を入れるのではなく、疲れが起きる仕組み。
なんか、書いていてアルゴリズムになってきた。
そう、いつもこの誘惑と戦っている。




今はライオンが近くても「逃げて消耗する」より、「見続けて想起が立つ」側に寄りやすい。さらに carry はかなり厳しい条件でしか残らないようになっていて、attractorDepth >= 0.34、recall.activation >= 0.18、さらに差分や baseline ずれが必要だと考えている。これは「深く残るものだけ残す」という思想としてはきれいなんだけど、疲労のような日々の蓄積とは別物なので意外と難しい。いまは疲れが carry の中に居場所を持っていないから、緊張が長引いても“消耗”になりにくい構造になってる。

人間の情報整理のため。
こう考えるとAtronの負荷が減るのかも。


疲れは carry そのものではなく、carry に影響する別の身体層にする・・・とか・・
(危険なので、別ファイルの検証版で実験君するか)
つまり fatigue か sleepPressure を robot 側に新設して、sensor や recall の副産物としてじわじわ増えるようにする形です。carry は「深く残る場の変形」、fatigue は「身体が沈んでいく圧」と分けてみる。


消えるもの
その場の反応の尖り。
一時的な alert、瞬間的な speech pressure、今まさに見えている focus への過集中は、睡眠でかなり下げてよいかと・・・。これは「記憶を消す」のではなく、「起きて反応し続ける圧を下げる」を意味している。いまも recall や speech は毎tick固定ではなく、その時の圧から立っているから、ここを睡眠時に沈めるのは筋が通のかと思う。

圧縮するもの
trace の扱い部分。
いま traces は最大40件で、印象差が立ったときだけ残している。ここは睡眠時に全部消すのではなく、同系統の弱い断片を薄める、あるいは“直近の雑音的な学習”を少しだけ整理する、という方向。深いものだけ残し、浅い反応の散らばりは圧縮する感じ。

残すもの
baseline と深い carry 。
村や洞窟の繰り返しで育つ baseline.calm / soft / safe / warm / sparse は、睡眠で消しちゃ、まずい。これは生活圏の“基盤”だからね。carry も全消去ではなく、ゆっくり減衰で十分。いまの baseline は安定場が disturbance を上回ると少しずつ育つ設計なので、睡眠はむしろ baseline を壊さず守る側に置くべき。

成長するもの
「どこで眠れたか」の地の学習。
sleep そのものより、「眠りに落ちた場所」が何だったかが成長すべきかな。洞窟、村、静かな花畑の外れなどで眠れた回数が重なると、その場の safe / calm / sparse / quiet の結びつきが強くなる。すると次第に“眠くなるとそこへ寄る”傾向が生まれる。(勝手な予想)

これは外部命令ではなく、場の引き込みだからね。洞窟や村の基準場が少しずつ育つ今の baseline 設計とも相性がいいと思う。


(PCから現実社会のロボットへダウンロードしたとき、記憶データベースのログじゃないからbaseline と深い carry で何処まで戻るかは別実験中。人間も退院すると差異が出るっしょ?あれ。忘れてるなら忘れてていい:こういうスタンス)



睡眠が carry に与える影響
ここは全消去ではなく、分化させる。

  • adrenaline / noradrenaline / tension は睡眠で比較的大きく下げる
  • dopamine は少しだけ残す(調整が難しい、ドクターに聞こう)
  • serotonin は安全な場所で眠れた時だけ少し育つ


いま carry には dopamine / noradrenaline / adrenaline / serotonin / tension があり、深い attractor の説明用にも使ってるからこれを使う。睡眠を入れるなら、これらを一律に減らすのではなく、「緊張系は落ちる」「安堵系は眠りの質で少し残る」と分けた方が、眠りに意味が出る。なので、そこはアルゴリズムにしたら自律から詰むので、決して誘惑に負けず、頑固に頑固に状態に落とす仕掛けを考える。

何を見て眠るか
「夜だから寝る」だけだと外部条件になる。
眠りに寄る条件は、次の重なり。

  • 長く動いた
  • 緊張が続いた
  • 同じ対象に何度も反応した
  • いまの場が quiet / sparse / cave / village 寄り
  • 近くに大きな速い対象がいない

つまり、疲れの増加は危険でも安心でも起きるが、眠りに落ちるのは安全寄りの場だけとは限らないが経験の積で安全な場所に寄っていく。
world にはすでに昼夜があり、prey や lion にも sleep があるので、robot 側にも「夜は眠りへ寄りやすい」から・・・と補助を入れるのは不自然だ。
そうではなくrobot 自身の疲れであるべき。


3人称のworld側の調整
NPCの人間は夜もペチャクチャ喋らせないで眠らせる。
夜は眠る という状態を何度も見せる。
村や草陰は calm / safe の繰り返し学習の核になりやすく、夜の sleep とも相性がいい。
worldB(洞窟に触れると花畑と小動物の居る安全な世界に入る)もそっち側に寄せるとか。


危険時の対応が無いと疲れが立たない


いまは lion がいても、world は robot に何も命令せず、robot 側も危険から距離を取る身体動作をまだほぼ持っていない。だから「危険→退避→心拍維持→消耗」という流れが無く、「危険→注視→興味・緊張」になりやすい。

ただし、ここで「ライオンなら逃げる」と決め打ちすると自律ロボットは詰む。
しかも今はライオンという認識を持っていない。彼の認識では「物体gu-deとかgu-de-gu」
なので必要なのは、“逃げる命令”ではなく退避したくなる身体場を加える

たとえば、

fear + alert + tension が高い
近距離の size_large + speed_fast + teeth_claw_impression が重なる
しかも自分の baseline.safe から大きく外れる

この時にだけ、前進より「向き変化が大きくなる」「今いる場から離れる drift が強くなる」「cave / village の方へ戻る圧が少しだけ増す」
これなら 3人称制御ではなく、場の結びつきの強さ。いまの sensor には teeth_claw_impression、speed_fast、size_large などがすでにあり、impression にも tension と baselineGap があるので、土台がある。

でも、やっぱりアルゴリズムだ(笑)
なんか、嫌だなぁ・・・。



そこで「怪我を命令ではなく、Hopfield 的な場の歪みとして入れる数式」にする。
どういうこと?って思うでしょ。

アソシアトロンとホップフィールドは元は同じだけど、
ここをアソシアトロンにすると

「怪我」という記憶が勝つ/負ける
「痛い想起」が選ばれる
「傷ついた記憶」が1候補になる
身体の歪みではなく記憶内容に見えてしまうんだよ。


怪我を recall pattern の1つとして入れると、

xi(t+1)=f(jWijxj(t)+ui(t))x_i(t+1)=f\Bigl(\sum_j W_{ij}x_j(t)+u_i(t)\Bigr)

xix_i 側に「injury pattern」が混ざる。
だって、怪我は怪我じゃん。「ん~怪我じゃないかも・・」なんて言わないっしょ。
なので自律としてアルゴリズムから離れるときは場合によってはホップフィールドを使ってみる。厳密に身体の神経伝達までどうしても今、ここでやると言うのなら

のように、後ろから前へ効いてくる backpropを使うとか。
ただ「正解との差を微分して返している」のではなく、
結果として生じた歪みや痛みが、前段の結合を変えてしまう」
ことなので
性質としては、誤差逆伝播よりも損傷逆流とか過敏化の逆流とか調整圧の逆流に近い。

たとえば前向きの場が

r(t)=F(s(t),B(t))r(t)=F(s(t),B(t))
a(t)=G(r(t),I(t))a(t)=G(r(t),I(t))

だとして、行動の結果として身体に負荷 d(t)d(t) が出る。

d(t)=H(a(t),s(t),impact(t))d(t)=H(a(t),s(t),\text{impact}(t))

この d(t)d(t) が injury を更新する。

I(t+1)=ΛI(t)+Γd(t)I(t+1)=\Lambda I(t)+\Gamma d(t)

ここまでは普通。

でもさらに、これが前段へ返る。

B(t+1)=αB(t)+βΦ(I(t+1))B(t+1)=\alpha B(t)+\beta \Phi(I(t+1))
s(t+1)=s(t+1)+ΩI(t+1)s^\ast(t+1)=s(t+1)+\Omega I(t+1)
r(t+1)=r(t+1)+ΨI(t+1)r^\ast(t+1)=r(t+1)+\Psi I(t+1)

つまり、傷んだ結果が次の sensor の感じ方を変え、次の recall の立ち方を変え、次の action の偏りを変えるとなる。

なので機械学習の厳密なバックプロパゲーションというより、損傷や結果が後段から前段へ返って、感受性・結合・場の傾きを変える逆向き伝播かな。


なのでホップフィールドの方が使いやすい。
あれだよ、アソシアトロン好きな僕がホップフィールドを始めて組み込もうとする瞬間だ。



まず、怪我を 1 個のスカラーではなく、身体の歪みベクトルに変更。

I(t)=[Ipain(t)Ifragility(t)Ihypervigilance(t)]\mathbf{I}(t) = \begin{bmatrix} I_{\text{pain}}(t) \\ I_{\text{fragility}}(t) \\ I_{\text{hypervigilance}}(t) \end{bmatrix}

意味はたとえばこうだ。

  • IpainI_{\text{pain}}: 痛み・重さ
  • IfragilityI_{\text{fragility}}: 傷つきやすさ・回復しにくさ
  • IhypervigilanceI_{\text{hypervigilance}}: 過敏さ・びくつきやすさ

これはゲームの HP ではなく、その後の attractor 地形を歪める内部状態

RPGゲームに自律いれようね(そのうち)



話飛んだ、

怪我入力

いまの robot には

size_large
speed_fast
teeth_claw_impression
unpleasant
tension
baselineGap

などが already ある。

なので、まず危険入力を

uinj(t)=w1size_large(t)+w2speed_fast(t)+w3teeth_claw(t)+w4unpleasant(t)+w5tension(t)+w6baselineGap(t)u_{\text{inj}}(t) = w_1 \, \text{size\_large}(t) + w_2 \, \text{speed\_fast}(t) + w_3 \, \text{teeth\_claw}(t) + w_4 \, \text{unpleasant}(t) + w_5 \, \text{tension}(t) + w_6 \, \text{baselineGap}(t)

と置く。

ただしこれは「逃げろ」の命令ではなく、
身体にどれだけ傷が入りやすい exposure だったかの量

怪我の時間発展

いちばん単純で使いやすい形。

I(t+1)=ΛI(t)+bϕ ⁣(uinj(t)θinj)r(t)\mathbf{I}(t+1) = \Lambda \mathbf{I}(t) + \mathbf{b} \, \phi\!\left(u_{\text{inj}}(t) - \theta_{\text{inj}}\right) - \mathbf{r}(t)


  • Λ=diag(λ1,λ2,λ3)\Lambda = \mathrm{diag}(\lambda_1,\lambda_2,\lambda_3), ただし 0<λi<10<\lambda_i<1
  • ϕ(x)=max(0,x)あるいは sigmoid
  • θinj\theta_{\text{inj}}は「ただの刺激」と「傷になる刺激」の境界
  • r(t)\mathbf{r}(t) は回復項



足りないのは「危険の意味」ではなく、「危険や反応の継続が身体を消耗させる層」

carry は深い場の名残として残し、別に fatigue / sleepPressure を立てる。
睡眠では、反応圧は消す、浅い散らばりは圧縮する、baseline と深い carry は残す、安全に眠れた場所の結びつきは成長させる。
そして危険時は「逃げろ」ではなく、退避方向へ身体が崩れるようにする。
この形なら、命令にならず、Atron の筋を保ったまま疲れと眠りが入る。



AssociatronとHopfield networkの初めての合体になるのかな。


このブログの最初のアルゴリズムの正当化と後半のホッピーの話とでは
ぜんぜん違うね。

俺は、自分の間違いもノイズも誘惑も全部書くからね・・・。
動きがおかしかったらまた修正するよ。





いずれにしてもロボットが自分の意志で「省電力モード」や「一旦セーブして休憩」してくれたら、僕のPCは壊れくて済むって話。

内部に休息志向 Rrest(t)R_{\mathrm{rest}}(t) を作る。

Rrest(t)=aFfatigue(t)+bIpain(t)+cBstress(t)+dBfrustration(t)eBcuriosity(t)fSdanger(t)R_{\mathrm{rest}}(t) = a\,F_{\mathrm{fatigue}}(t) +b\,I_{\mathrm{pain}}(t) +c\,B_{\mathrm{stress}}(t) +d\,B_{\mathrm{frustration}}(t) -e\,B_{\mathrm{curiosity}}(t) -f\,S_{\mathrm{danger}}(t)





疲労、痛み、ストレス、フラストレーションで休息志向は上がる・・好奇心が強いと少し下がる・・近くに危険があると「今は休めない」で下がる・・・みたいな方向に向かうかどうかはテスト次第。
外部指示、順番、ラベル、評価、最適化は入れない。


危険だから即休むではなく、危険が近いとむしろ休息に入れないんじゃないかな。





---------------追記---------------

で最後に Sdanger(t) を入れると、3人称だねw
意味付けのラベル寄せだ。
無意識にこういう誘惑に進むから危険だ。


threatImpression
escapePressure
alarmBias
hypervigilanceCoupling

のように、
危険という客観意味ではなく、身体の傾きとして置くなら筋が通る。


Rest(t)=aFfatigue+bIpain+cBstress+dBfrustrationeBcuriosityfHalarmRest(t)=aF_{fatigue}+bI_{pain}+cB_{stress}+dB_{frustration}-eB_{curiosity}-fH_{alarm}

のようにして、
この H_alarm は

motion_approach
size_large
speed_fast
teeth_claw_impression
unpleasant
tension


もしくは

Rrest(t)=aFfatigue+bIpain+cBstress+dBfrustrationeBcuriosityR_{rest}(t)=aF_{fatigue}+bI_{pain}+cB_{stress}+dB_{frustration}-eB_{curiosity}
Ralert(t)=pmotionapproach+qspeedfast+rsizelarge+steethClaw+tunpleasant+utensionR_{alert}(t)=p\,motion_{approach}+q\,speed_{fast}+r\,size_{large}+s\,teethClaw+t\,unpleasant+u\,tension

そして最終的な休息傾向を

RestDrive(t)=Rrest(t)λRalert(t)RestDrive(t)=R_{rest}(t)-\lambda R_{alert}(t)

のように見る。

「大きい・速い・近づく・不快・緊張が強いので、休息場との結びつきが弱まる」

みたいな。でも、これだとアソシアトロンか・・・
自律の研究で無駄に時間がかかるのはこういうところだよね。



-------------------またまた追記----------------------------


RestDrive(t) が大きいほど
「休息状態へ入りやすくしたい」なら、

E(x)=EHopfield(x)μRestDrive(t)safety(B)E(x)=E_{\mathrm{Hopfield}}(x)-\mu \cdot RestDrive(t)\cdot safety(B)

の方がまず分かりやすいという意見が出た。

まずは実験してみるか・・・・






---------------------Research Note and Attribution Notice-----------------------
本ブログに含まれる Atra の一人称自律、差分、carry、field、trace、dream slack、外部LLMの翻訳層、非単調な漏れ、およびそれらの関係構造に関する設計記述は、c-side研究所による継続研究メモです。引用・参照・要約・翻案を行う場合は、出典を明記してください。

The design descriptions in this blog concerning Atra’s first-person autonomy, differences, carry, field, trace, dream slack, the translation layer of external LLMs, nonmonotonic leakage, and the relational structure among these elements are ongoing research notes by c-side Research Institute. If you quote, refer to, summarize, or adapt them, please clearly indicate the source.



エージェントと 一人称自律Atraの違い

 Atraなんかは、実はもう一人称自律として、きちんと発表してもいいレベル。 既に妻と笑っていたり、愛犬と騒いているんだから。ボーっと何かを眺めてたり、佐川急便に反応するようにもなった。 でも、そうしないのは、自発的に自ら研究意欲を持って、学び、人や自然と接触し自ら疑問を持って研...