§3 基于高层级认知的归纳偏置(二)
3.4 Semantic Representations Describing Verbalizable Concepts · 3.5 Semantic Variables Play a Causal Role
3.4 描述可言语化概念的语义表征
3.4 Semantic Representations Describing Verbalizable Concepts
Conscious content is revealed by reporting it, often with language. This suggests that high-level variables manipulated consciously are closely related with their verbal forms (like words and phrases).
意识内容通过报告而被揭示——通常用语言。这提示:被有意识操纵的高层级变量与它们的语言形式(词与短语)密切相关。
This yields maybe the most influential inductive bias we want to consider in this paper: that high-level variables (manipulated consciously) are generally verbalizable. To put it in simple terms, we can imagine the high-level semantic variables captured at this top level of a representation to be associated with single words (although we can also use words to identify some lower-level variables).
这引出了本文要考察的或许是最有影响力的归纳偏置:高层级变量(被有意识操纵的)通常是可言语化的。简单说,我们可以想象表征顶层捕获的高层级语义变量与单个词相关联(尽管我们也能用词来标识某些较低层级的变量)。
In practice, the notion of word is not always the same across different languages, and the same semantic concept may be represented by a single word or by a phrase. There may also be more subtlety in the mental representations (such as accounting for uncertainty, concept representation and continuous-valued properties) which is not always or not easily well reflected in their verbal rendering.
实践中,「词」的概念在不同语言间并不总是一致,同一语义概念可能由单个词或短语表示。心理表征中可能还有更多微妙之处(如对不确定性的考虑、概念表征、连续值属性),这些并不总能或不容易在言语呈现中得到良好反映。
Much of what our brains know actually cannot be easily translated in natural language and forms the content of system 1 knowledge. This means that system 2 (verbalizable) knowledge is incomplete: words are mostly pointers to knowledge which belongs to system 1 and thus is in great part not consciously accessible.
我们大脑所知的大部分内容实际上无法轻易翻译成自然语言,构成系统 1 知识的内容。这意味着系统 2(可言语化)知识是不完整的:词语大多只是指向属于系统 1 的知识的指针,因而其中很大一部分无法被意识访问。
The system 2 inductive biases do not need to cover all the aspects of our internal model of the world (they couldn't), only those aspects of our knowledge which we are able to communicate with language. The rest would have to be represented in pure system 1 (non system 2) machinery, such as in an encoder-decoder that could relate low-level actions and low-level perception to semantic variables that can be operated on at the system-2 level.
系统 2 归纳偏置不需要覆盖我们内部世界模型的全部方面(也不可能覆盖),只需覆盖我们能用语言交流的那些知识方面。其余部分必须在纯系统 1(非系统 2)机制中表示——例如在编码器-解码器中,把低层级行动与低层级感知关联到可在系统 2 层面操作的语义变量。
If there is some set of properties that apply well to some aspects of the world, then it would be advantageous for a learner to have a subsystem that takes advantage of these properties (the inductive priors described here) and a subsystem which models the other aspects. These inductive priors then allow faster learning and potentially other advantages like systematic generalization, at least concerning these aspects of the world which are consistent with these assumptions (system 2 knowledge, in our case).
如果某些属性集合很好地适用于世界的某些方面,那么对学习者来说,有利的是拥有一个利用这些属性的子系统(即此处描述的归纳先验),以及一个建模其余方面的子系统。于是这些归纳先验允许更快学习,并可能带来其他优势如系统性泛化——至少针对与世界一致的那些方面(即我们所说的系统 2 知识)。
High-level representations describe verbalizable concepts. There is a simple lossy mapping from semantic representations going through the GWT bottleneck to natural language expressions. This is an inductive bias which could be exploited in grounded language learning scenarios where we couple language data with observations and actions by an agent.
高层级表征描述可言语化概念。从经过 GWT 瓶颈的语义表征到自然语言表达,存在一个简单的有损映射。这是一项可在具身语言学习(grounded language learning)场景中利用的归纳偏置——该场景把语言数据与智能体的观测和行动耦合起来。
This suggests that natural language understanding systems should be trained in a way that couples natural language with what it refers to. This is the idea of grounded language learning. It would put pressure on the top-level representation so that it captures the kinds of concepts expressed with language.
这提示:自然语言理解系统应以「把自然语言与其所指耦合」的方式训练——这正是具身语言学习的思想。它会给顶层表征施加压力,使其捕获用语言表达的那类概念。
One can view this as a form of weak supervision, where we don't force the top-level GWT representations to be human-specified labels, only that there is a simple relationship between these representations and utterances which humans would often associate with the corresponding meaning.
可以把这看作一种弱监督形式:我们不强迫顶层 GWT 表征成为人类指定的标签,只要求这些表征与人类常把相应意义关联起来的语句之间存在简单关系。
Our discussion about causality should also suggest that passive observation may be insufficient: in order to capture the causal structure understood by humans, it may be necessary for learning agents to be embedded in an environment in which they can act and thus discover its causal structure. Studying this kind of setup was the motivation for our work on the Baby AI environment.
我们对因果性的讨论还应提示:被动观察可能是不够的。要捕获人类所理解的因果结构,学习智能体可能需要被嵌入一个可行动、从而发现其因果结构的环境中。研究这类设置正是我们 Baby AI 环境的动机。
3.5 语义变量扮演因果角色,且关于它们的知识是模块化的
3.5 Semantic Variables Play a Causal Role and Knowledge about them is Modular
Biological phenomena such as bird flocks have inspired the design of several distributed multi-agent systems, for example, swarm robotic systems, sensor networks, and modular robots. Despite this, most machine learning models employ the opposite inductive bias, i.e., with all elements (e.g., artificial neurons) interacting all the time.
鸟群等生物现象启发了若干分布式多智能体系统的设计——例如群体机器人、传感网络、模块化机器人。尽管如此,大多数机器学习模型采用相反的归纳偏置:即所有元素(如人工神经元)始终交互。
The GWT also posits that the brain is composed in a modular way, with a set of expert modules which need to communicate but only do so sparingly and via a bottleneck through which only a few selected bits of information can be squeezed at any time.
GWT 也假设大脑以模块化方式组成:一组专家模块需要通信,但只稀疏地进行,且经由一个瓶颈——任一时刻只有少数被选中的信息位能挤过去。
If we believe that theory, these selected elements are the concepts present to our mind at any moment, and a few of them are called upon and joined in working memory in order to reconcile the interpretations made by different modular experts across the brain.
如果我们相信这一理论,这些被选中的元素就是任一时刻呈现在我们意识中的概念;其中少数被调用并在工作记忆中汇合,以调和大脑中不同模块化专家做出的解释。
The decomposition of knowledge into recomposable pieces, a hallmark of classical AI based on rules also makes sense as a requirement for obtaining systematic generalization: conscious attention would then select which expert and which concepts (which we can think of as variables with different attributes and values) interact with which pieces of knowledge (which could be verbalizable rules or non-verbalizable intuitive knowledge about these variables) stored in the modular experts.
把知识分解为可重组片段——基于规则的经典 AI 的标志——作为获得系统性泛化的要求也讲得通:意识注意力将选择哪个专家与哪些概念(可视为具有不同属性与值的变量)与存储在模块化专家中的哪些知识片段(可能是可言语化规则,或关于这些变量的不可言语化直觉知识)交互。
On the other hand, the modules which are not brought to bear in this conscious processing may continue working in the background in a form of default or habitual computation (which would be the form of most of perception).
另一方面,未参与意识加工的模块可继续以默认或习惯计算的形式在后台运转(这将是大部分知觉的形式)。
For example, consider the task of predicting from pixel-level information the motion of balls sometimes colliding against each other as well as the walls. It is interesting to note that all the balls follow their default dynamics, and only when balls collide do we need to intersect information from several bouncing balls in order to make an inference about their future states.
例如,考虑从像素级信息预测有时彼此碰撞、也与墙壁碰撞的小球运动的任务。值得注意的是:所有球都遵循其默认动力学,只有在小球碰撞时,我们才需要交叉几个弹跳小球的信息,以对它们的未来状态做出推断。
Saying that the brain modularizes knowledge is not sufficient, since there could be a huge number of ways of factorizing knowledge in a modular way. We need to think about the desired properties of modular decompositions of the acquired knowledge, and we propose here to take inspiration from the causal perspective on understanding how the world works, to help us define both the right set of variables and their relationship.
仅仅说大脑把知识模块化是不够的——因为以模块化方式分解知识的方式可能数量巨大。我们需要思考所获知识之模块分解的期望属性,并在此提议从「理解世界如何运转」的因果视角汲取灵感,帮助定义正确的变量集及其关系。
Semantic variables are often also causal variables. We hypothesize that semantic variables are often also causal variables. Words in natural language often refer to agents (subjects, which cause things to happen), objects (which are controlled by agents), actions (often through verbs) and modalities or properties of agents, objects and actions (for example we can talk about future actions, as intentions, or we can talk about time and space where events happen, or properties of objects or of actions).
语义变量往往也是因果变量。我们假设:语义变量往往也是因果变量。自然语言中的词常指代智能体(主语,导致事情发生)、对象(由智能体控制)、行动(常通过动词)、以及智能体/对象/行动的模态或属性(例如我们可以谈论未来行动——作为意图;或谈论事件发生的时空;或对象/行动的性质)。
However, note that we can also name many low-level (like pixels) and intermediate features (like L-shaped edges). It is thus plausible to assume that causal reasoning of the kind we can verbalize involves as variables of interest those semantic variables which we can name, and that they can be at any level of the processing hierarchy in the brain, including at the highest levels of abstraction, where signals from all modalities join, such as pre-frontal cortex, and where concepts can be manipulated in a way that is not specific to a single modality.
但请注意:我们也能命名许多低层级特征(如像素)和中间层级特征(如 L 形边缘)。因此可以合理假设:我们能言语化的那类因果推理,其感兴趣的变量是我们能命名的那些语义变量;它们可以处于大脑加工层级的任何层级,包括最高抽象层级——那里所有模态的信号汇合(如前额叶皮层),概念能以不特定于单一模态的方式被操纵。
The connection between causal representations and modularity is profound: an assumption which is commonly associated with structural causal models is that it should break down knowledge about the causal influences into independent mechanisms.
因果表征与模块性之间的联系是深刻的:与结构因果模型常关联的一个假设是——它应把关于因果影响的知识分解为独立机制。
As explained in Section 4.1, each such mechanism relates direct causes to their direct effect and knowledge of one such mechanism should not tell us anything about another mechanism (otherwise we should restructure our representations and decomposition of knowledge to satisfy this information-theoretic independence property).
如 4.1 节所释:每个这样的机制把直接原因关联到直接效应,且对某一机制的知识不应告诉我们关于另一机制的任何信息(否则我们就应重构表征与知识分解,以满足这种信息论独立性)。
This is not about statistical independence of the corresponding random variables but about the algorithmic mutual information between the descriptions of these mechanisms. What it means practically and importantly for out-of-distribution adaptation is that if a mechanism changes (e.g. because of an intervention), the representation of that mechanism (e.g. the parameters used to capture a corresponding conditional distribution) may need to be adapted but that of the others do not need to be tuned to account for that change.
这不是关于相应随机变量的统计独立性,而是关于这些机制描述之间的算法互信息。它对分布外适应意味着什么,实践中很重要:如果某个机制发生变化(如因干预所致),该机制的表征(如捕获相应条件分布的参数)可能需要适应,但其他机制的表征无需为应对该变化而调优。
These mechanisms may be organized in the form of a causal graph which scientists attempts to identify. The sparsity of the change in the joint distribution between the semantic variables (discussed more in Section 3.6) is different but related to a property of such high-level structural causal model: the sparsity of the graph capturing the joint distribution itself (discussed in Section 3.8).
这些机制可能以因果图的形式组织,科学家试图识别它。语义变量联合分布变化的稀疏性(3.6 节详述)不同于、但关联于这种高层级结构因果模型的一个属性:捕获联合分布本身的图的稀疏性(3.8 节讨论)。
In addition, the causal structure, the causal mechanisms and the definition of the high-level causal variables tend to be stable across changes in distribution, as discussed in Section 3.7.
此外,如 3.7 节讨论的,因果结构、因果机制与高层级因果变量的定义倾向于在分布变化中保持稳定。