← 中英对照目录 · ← 书架
§3 基于高层级认知的归纳偏置(三)
3.6 Local Changes in Distribution in Semantic Space · 3.7 Stable Properties of the World
3.6 语义空间中的局部分布变化
3.6 Local Changes in Distribution in Semantic Space
Consider a learning agent, like a learning robot or a learning child. What are the sources of non-stationarity for the distribution of observations seen by such an agent, assuming the environment is in some (generally unobserved) state at any particular moment?
考虑一个学习智能体,比如一台正在学习的机器人或一个正在学习的孩子。假设环境在任一时刻处于某种(通常未被观测到的)状态,那么该智能体所见观测分布的非平稳性来源是什么?
Two main sources are (1) the non-stationarity due to the environmental dynamics (including the learner's actions and policy) not having converged to an equilibrium distribution (or equivalently the mixing time of the environment's stochastic dynamics is longer than the lifetime of the learning agent) and (2) causal interventions by agents (either the learner of interest or some other agents).
两大来源是:(1) 由于环境动力学(包括学习者的行动与策略)尚未收敛到平衡分布而导致的非平稳性(等价地说,环境随机动力学的混合时间长于学习智能体的寿命);(2) 智能体(感兴趣的学习者或其他智能体)的因果干预
The first type of change includes for example the case of a person moving to a different country, or a videogame player learning to play a new game or a never-seen level of an existing game. That first type also includes the non-stationarity due to changes in the agent's policy arising from learning.
第一类变化例如:一个人搬到不同国家,或游戏玩家学习玩一个新游戏、或玩现有游戏中从未见过的关卡。第一类还包括因学习导致的智能体策略变化带来的非平稳性。
The second case includes the effect of actions such as locking some doors in a labyrinth (which may have a drastic effect on the optimal policy). The two types can intersect, as the actions of agents (including those of the learner, like moving from one place to another) contribute to the first type of non-stationarity.
第二类包括行动的效果,如锁定迷宫中的某些门(这可能对最优策略产生剧烈影响)。两类可能交叉:智能体的行动(包括学习者的行动,如从一个地方搬到另一个地方)也贡献于第一类非平稳性。
Changes in distribution are localized in the appropriate semantic space. Let us consider how humans describe these changes with language. For many of these changes, they are able to explain the source of change with a few words (a single sentence, often).
分布变化在恰当的语义空间中是局部的。让我们考虑人类如何用语言描述这些变化。对其中许多变化,他们能用几个词(往往是一个句子)解释变化的来源。
This is a very strong clue for our proposal to include as an inductive bias the assumption that the source of most changes in distribution are localized in the appropriate semantic space: only one or a few variables or mechanisms need to be modified to account for the change.
这是对我们提案的极强线索:把「大多数分布变化的来源在恰当语义空间中是局部的」作为一项归纳偏置——即只需要修改一个或少数几个变量或机制,就能解释该变化。
Note how humans will even create new words when they are not able to explain a change with a few existing words, with the new words corresponding to new latent variables, which when introduced, make the changes explainable "easily" (assuming one understand the definition of these variables and of the mechanisms relating them to other variables).
注意:当人类无法用几个现成词解释一种变化时,他们甚至会创造新词——新词对应于新的潜在变量;一旦引入,就能「轻松」解释这些变化(假设你理解这些变量的定义、以及把它们与其他变量关联起来的机制)。
For system-2 distributional changes (due to interventions), we automatically get locality of the source of changes (which start at one or a few nodes of the causal graph). This is a plausible assumption since, by virtue of being localized in time and space, actions can only directly affect very few high-level variables, with other effects (on downstream variables) being consequences of the initial intervention.
对系统 2 的分布变化(因干预所致),我们自动获得变化来源的局部性(变化起始于因果图的一个或少数几个节点)。这是一个合理的假设:由于行动在时空上是局部的,它只能直接影响极少数高层级变量,对其他(下游)变量的影响只是初始干预的后果。
This sparsity of sources of change is a strong assumption which can put pressure on the learning process to discover high-level representations which have that property. Here, we are assuming that the learner has to jointly discover these high-level representations (i.e. how they relate to low-level observations and low-level actions) as well as how the high-level variables relate to each other via causal mechanisms.
变化来源的这种稀疏性是一个强假设,可以给学习过程施加压力,促使其发现具有该属性的高层级表征。这里我们假设:学习者必须联合发现这些高层级表征(即它们如何与低层级观测和低层级行动相关),以及高层级变量如何通过因果机制相互关联。
3.7 世界的稳定属性
3.7 Stable Properties of the World
Above, we have talked about changes in distribution due to non-stationarities, but there are aspects of the world that are stationary, which means that learning about them would eventually converge.
上面我们谈论了因非平稳性导致的分布变化,但世界有一些方面是平稳的——意味着对它们的学习最终会收敛。
In an ideal scenario, our learner has an infinite lifetime and the chance to learn everything about the world (a world where there are no other agents) and build a perfect model of it, at which point nothing is new and all of the above sources of non-stationarity are gone.
在理想场景中,学习者拥有无限寿命、有机会学习关于世界的一切(一个没有其他智能体的世界)并构建它的完美模型——到那时,一切都已不新,上述所有非平稳性来源都消失了。
In practice, only a small part of the world will be understood by the learning agent, and interactions between agents (especially if they are learning) will perpetually keep the world out of equilibrium.
实践中,学习智能体只能理解世界的一小部分,而智能体之间的交互(尤其是当它们在相互学习时)将永远使世界偏离平衡。
If we divide the knowledge about the world captured by the agent into the stationary aspects (which should converge) and the non-stationary aspects (which would generally keep changing), we would like to have as much knowledge as possible in the stationary category.
如果我们把智能体捕获的世界知识分为平稳方面(应收敛)与非平稳方面(通常会持续变化),我们会希望尽可能多的知识落在平稳类别中。
The stationary part of the model might require many observations for it to converge, which is fine because learning these parts can be amortized over the whole lifetime (or even multiple lifetimes in the case of multiple cooperating cultural agents, e.g., in human societies).
模型的平稳部分可能需要大量观测才能收敛,这没关系——因为学习这些部分可以摊销到整个寿命(甚至在多个协作文化智能体的情况下跨多个寿命摊销,如人类社会)。
On the other hand, the learner should be able to quickly learn the non-stationary parts (or those the learner has not yet realized can be incorporated in the stationary parts), ideally because very few of these parts need to change, if knowledge is well structured.
另一方面,学习者应能快速学习非平稳部分(或那些尚未意识到可并入平稳部分的内容)——理想情况下,因为如果知识结构良好,需要变化的部分极少。
Hence we see the need for at least two speeds of learning, similar to the division found in meta-learning of learnable coefficients into meta-parameters on one hand (for the stable, slowly learned aspects) and parameters on the other hand (for the non-stationary, fast to learn aspects), as already discussed above in Section 2.
因此我们看到对至少两种学习速度的需求——类似于第 2 节已讨论的元学习中可学习系数的划分:一边是元参数(用于稳定、慢学的方面),另一边是参数(用于非平稳、快学的方面)。
Stable v/s Unstable properties of the world. There should be several speeds of learning, with more stable aspects learned more slowly and more non-stationary or novel ones learned faster, and pressure to discover stable aspects among the quickly changing ones. This pressure would mean that more aspects of the agent's represented knowledge of the world become stable and thus less needs to be adapted when there are changes in distribution.
世界的稳定与不稳定属性。应有多种学习速度:更稳定的方面学得更慢,更非平稳或新颖的方面学得更快,并存在「在快速变化中发现稳定方面」的压力。这种压力意味着智能体表征的世界知识有更多方面变得稳定,从而分布变化时需要适应的更少。
For example, consider scientific laws, which are most powerful when they are universal. At another level, consider the mapping between the perceptual input, low level actions, and the high-level semantic variables. An encoder that would implement this mapping should ideally be highly stable, or else downstream computations would need to track those changes (and indeed the low-level visual cortex seems to compute features that are very stable across life, contrary to high-level concepts like new visual categories).
例如,考虑科学定律——它们在普适时最强大。在另一层面,考虑知觉输入、低层级行动与高层级语义变量之间的映射。实现该映射的编码器理想上应高度稳定,否则下游计算就需跟踪这些变化(确实,低层级视觉皮层计算的似乎是终生高度稳定的特征——与新视觉类别这类高层级概念相反)。
Causal interventions are taking place at a higher level than the encoder, changing the value of an unobserved high-level variable or changing one of the mechanisms. If a new concept is needed, it can be added without having to disturb other represented knowledge, especially if it can be learned as a composition of existing high-level features and concepts.
因果干预发生在比编码器更高的层面,改变某个未观测高层级变量的值或改变某个机制。如果需要新概念,可以添加它而不必扰动其他已表征知识——尤其当它能作为既有高层级特征与概念的组合来学习时。
We know from observing humans and their brain that new concepts which are not obtained from a combination of old concepts (like a new skill or a completely new object category not obtained by composing existing features) take more time to learn, while new high-level concepts which can be readily defined from other high-level concepts can be learned very quickly (as fast as with a single example or definition).
从观察人类及其大脑可知:无法由旧概念组合得到的新概念(如一项新技能、或无法由既有特征复合得到的全新对象类别)需要更长时间学习;而能轻易由其他高层级概念定义的新高层级概念则可以学得很快(快到一个样本或一条定义即可)。
Another example arising from the analysis of causal systems is that causal interventions (which are in the non-stationary, quickly inferred or quickly learned category) may temporarily modify the causal graph structure (which specifies which variable is a direct cause of which) by breaking causal links (when we set a variable we break the causal link from its direct causes) but that most of the causal graph is a stable property of the environment.
从因果系统分析中产生的另一个例子:因果干预(属于非平稳、可快速推断或快速学习的类别)可能通过打破因果链接(当我们设定一个变量时,就打破了它与其直接原因之间的因果链接)暂时修改因果图结构——但因果图的大部分是环境的稳定属性。
← 主页