← 中英对照目录 · ← 书架
第十三章 · Section 13
Discussion · 讨论
This framework provides a structured, quantifiable methodology for evaluating Artificial General Intelligence (AGI), moving beyond narrow, specialized benchmarks to assess the breadth (versatility) and depth (proficiency) of cognitive capabilities. By operationalizing AGI through ten core cognitive domains inspired by the CHC theory, we can systematically diagnose the strengths and profound weaknesses of current AI systems. The estimated AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) illustrate both the rapid progress in the field and the substantial gap remaining before achieving human-level general intelligence.
本框架为评估通用人工智能(AGI)提供了一种结构化、可量化的方法论,它超越了狭窄的专用基准,去评估认知能力的广度(通用性)与深度(熟练度)。通过借鉴 CHC 理论、用十个核心认知域来操作化 AGI,我们能系统地诊断当前 AI 系统的优势与深层弱点。估算出的 AGI 分数(如 GPT-4 为 27%、GPT-5 为 57%)既展示了该领域的快速进步,也展示了在达成人类级通用智能之前仍然存在的巨大差距。
"锯齿状"的 AI 能力与关键瓶颈
"Jagged" AI Capabilities and Crucial Bottlenecks
The application of this framework reveals that contemporary AI systems exhibit a highly uneven or "jagged" cognitive profile. While models demonstrate high proficiency in areas that leverage vast training data—such as General Knowledge (K), Reading and Writing (RW), and Mathematical Ability (M)—they simultaneously possess critical deficits in foundational cognitive machinery. This uneven development highlights specific bottlenecks impeding the path to AGI. Long-term memory storage is perhaps the most significant bottleneck, scoring near 0% for current models. Without the ability to continually learn, AI systems suffer from "amnesia" which limits their utility, forcing the AI to re-learn context in every interaction. Similarly, deficits in visual reasoning limit the ability of AI agents to interact with complex digital environments.
应用这一框架揭示出,当代 AI 系统展现出一种高度不均衡或"锯齿状"的认知画像。模型在利用海量训练数据的领域——如常识(K)、读写(RW)与数学(M)——表现出高熟练度,但同时,在基础认知机制上存在关键缺陷。 这种不均衡的发展凸显出阻碍通向 AGI 之路的具体瓶颈。长时记忆存储或许是最大的瓶颈,当前模型得分接近 0%。没有持续学习的能力,AI 系统就受困于"失忆",这限制了它们的效用,迫使 AI 在每次交互中重新学习上下文。同样,视觉推理的缺陷限制了 AI 智能体与复杂数字环境交互的能力。
能力扭曲与"通用性幻觉"
Capability Contortions and the Illusion of Generality
The jagged profile of current AI capabilities often leads to "capability contortions," where strengths in certain areas are leveraged to compensate for profound weaknesses in others. These workarounds mask underlying limitations and can create a brittle illusion of general capability. • Working Memory vs. Long-Term Storage: A prominent contortion is the reliance on massive context windows (Working Memory, WM) to compensate for the lack of Long-Term Memory Storage (MS). Practitioners use these long contexts to manage state and absorb information (e.g., entire codebases). However, this approach is inefficient, computationally expensive, and can overload the system's attentional mechanisms. It ultimately fails to scale for tasks requiring days or weeks of accumulated context. A long-term memory system might take the form of a module (e.g., a LoRA adapter) that continually adjusts model weights to incorporate experiences. • External Search vs. Internal Retrieval: Imprecision in Long-Term Memory Retrieval (MR)—manifesting as hallucinations or confabulation—is often mitigated by integrating external search tools, a process known as Retrieval-Augmented Generation (RAG). However, this reliance on RAG is a capability contortion that obscures two distinct underlying weaknesses in an AI's memory. First, it compensates for the inability to reliably access the AI's vast but static parametric knowledge. Second, and more critically, it masks the absence of a dynamic, experiential memory—a persistent, updatable store for private interactions and evolving contexts in a long time scale. While RAG can be adapted for private documents, its core function remains retrieving facts from a database. This dependency can potentially become a fundamental liability for AGI, as it is not a substitute for the holistic, integrated memory required for genuine learning, personalization, and long-term contextual understanding.
当前 AI 能力的锯齿状画像常常导致"能力扭曲"——利用某些领域的优势去弥补其他领域的深层弱点。这些变通手段掩盖了底层局限,可能制造出一种脆弱的"通用能力幻觉"。 • 工作记忆 vs. 长时存储:一个突出的扭曲是依赖巨大的上下文窗口(工作记忆 WM)来弥补长时记忆存储(MS)的缺失。从业者用这些长上下文来管理状态、吸收信息(如整个代码库)。然而,这种方法低效、计算昂贵,还可能压垮系统的注意力机制。对于需要数天或数周累积上下文的任务,它最终无法扩展。一个长时记忆系统可能采取"模块"的形式(如 LoRA 适配器),持续调整模型权重以吸收经验。 • 外部搜索 vs. 内部检索:长时记忆检索(MR)的不精确——表现为幻觉或虚构——常通过集成外部搜索工具来缓解,即"检索增强生成"(RAG)。然而,对 RAG 的依赖是一种能力扭曲,掩盖了 AI 记忆的两处底层弱点。第一,它补偿了无法可靠存取 AI 庞大但静态的参数化知识。第二,也是更关键的,它掩盖了"动态经验记忆"的缺失——一个持久、可更新、用于存放长期尺度上的私有交互与演化语境的存储。虽然 RAG 可以适配到私有文档,但其核心功能仍是"从数据库检索事实"。这种依赖可能成为 AGI 的根本性负债,因为它无法替代真正学习、个性化与长期语境理解所需的整体性、一体化记忆。
Mistaking these contortions for genuine cognitive breadth can lead to inaccurate assessments of when AGI will arrive. These contortions can also mislead people to assume that intelligence is too jagged to be understood systematically.
把这些扭曲误认为真正的认知广度,会导致对"AGI 何时到来"的错误判断。这些扭曲还可能误导人们以为:智能太"锯齿状"了,无法被系统性地理解。
引擎类比
The Engine Analogy
Our multifaceted view of intelligence suggests an analogy to a high-performance engine, where overall intelligence is the "horsepower" (Jensen, 2000). An artificial mind, much like an engine, is ultimately constrained by its weakest components. Currently, several critical parts of the AI "engine" are highly defective. This severely limits the overall "horsepower" of the system, regardless of how optimized other components might be. This framework identifies these defects to guide our assessment and how far we are from AGI.
我们对智能的多面性观点提示了一个类比:一台高性能发动机——整体智能就是"马力"。一个人工心智,就像发动机一样,最终被它最弱的部件所约束。目前,AI"发动机"的几个关键部件严重缺陷。无论其他部件优化得多好,这都严重限制了系统的整体"马力"。本框架识别出这些缺陷,用来指导我们的评估,以及我们离 AGI 有多远。
社会智能 / 能力的相互依赖
Social Intelligence / Interdependence of Cognitive Abilities
Social Intelligence. Interpersonal skills are represented across these broad abilities. For example, cognitive empathy is captured in K's "commonsense" narrow ability. Facial emotion recognition is necessary for proficiency in V's "image captioning." And theory of mind is tested in on-the-spot reasoning (R). Interdependence of Cognitive Abilities. While this framework dissects intelligence into ten distinct axes for measurement, it is crucial to recognize that these abilities are deeply interdependent. Complex cognitive tasks rarely utilize a single domain in isolation. For example, solving advanced mathematical problems requires both Mathematical Ability (M) and On-the-Spot Reasoning (R). Theory of Mind questions require On-the-Spot Reasoning (R) as well as General Knowledge (K). Image recognition involves Visual Processing (V) and General Knowledge (K). Understanding a movie requires the integration of Auditory Processing (A), Visual Processing (V), and Working Memory (WM). Consequently, various batteries of narrow abilities test cognitive abilities in combination, reflecting the integrated nature of general intelligence.
社会智能:人际技能在这些宽泛能力中都有体现。例如,认知共情被 K 的"常识性知识"窄能力所涵盖;面部情绪识别是 V 的"图像描述"达到熟练所必需的;心智理论则在即时推理(R)中被测试。 能力的相互依赖:虽然本框架把智能拆解为十个独立测量轴,但必须认识到这些能力深度相互依赖。复杂的认知任务很少孤立地使用单一领域。例如,解决高等数学问题需要数学能力(M)与即时推理(R)两者;心智理论问题需要即时推理(R)以及常识(K);图像识别涉及视觉处理(V)与常识(K);理解一部电影需要听觉处理(A)、视觉处理(V)与工作记忆(WM)的整合。因此,各种窄能力测验是以组合方式测试认知能力,这正反映了通用智能的一体化本质。
污染防护 / 解数据集 vs 解任务 / 歧义消解
Contamination / Solving the Dataset vs. Solving the Task / Ambiguity Resolution
Contamination. Sometimes AI corporations "juice" their numbers by training on data highly similar to or identical to target tests. To defend against this, evaluators should assess model performance under minor distribution shifts (e.g., rephrasing the question) or testing on similar but distinct questions. Solving the Dataset vs. Solving the Task. Our operationalization relies on task specifications. We occasionally elaborate on these task specifications with specific datasets, and we usually treat them as necessary but not sufficient for solving the task. Moreover, solving our illustrative examples do not imply the task is solved, as our collection of examples are not exhaustive. It is the default for automatic evaluations to inadequately cover their target phenomena, so our operationalization is far more likely to be robust and stand the test of time compared to existing automated evaluations. Since we couch our definition in a collection of tasks rather than heavily depend on specific existing datasets, we can test AI systems using the best available tests at the time. Ambiguity Resolution. The batteries in the operationalization have varying levels of precision. However, the descriptions and examples should be clear enough that people can grade the AI systems themselves. Consequently, different people could issue their own estimates of the AGI score, and people can decide whether they find the grader's judgment reasonable.
污染防护:有时 AI 公司用与目标测验高度相似甚至相同的数据训练,"注水"自己的分数。为抵御这一点,评估者应在轻微分布偏移下评估模型表现(如改写问题),或用相似但不相同的问题测试。 解数据集 vs 解任务:我们的操作化依赖"任务规格"。我们偶尔用特定数据集来细化这些任务规格,但通常把它们视为"解任务的必要条件而非充分条件"。此外,解出我们的示例问题并不等于任务已解——因为我们的示例集并不穷尽。自动评估的常态恰恰是"无法充分覆盖其目标现象",因此我们的操作化相比现有自动评测,更可能稳健且经得起时间检验。由于我们把定义寄托于"一组任务"而非重度依赖特定现有数据集,我们就能用当时最好的测试来评估 AI 系统。 歧义消解:操作化中的测验套件精度不一。但描述与示例应当足够清晰,让任何人可以自行给 AI 系统打分。因此,不同的人可以给出自己对 AGI 分数的估计,而人们可以自行判断某一评分者的判断是否合理。
相关工作 / 局限
Related Work / Limitations
Related Work. Ilić and Gignac (2024) and Ren et al. (2024) find that a variety of AI systems' capabilities are highly correlated with pre-training compute. Gignac and Szodorai (2024) discuss human psychometrics and testing the intelligence of AI systems. Turing (1950) argues that the Turing Test can indicate general ability. Gubrud (1997) proposed an early definition of AGI in 1997. Marcus et al. (2016) discuss the need to move beyond the Turing Test to capture the multidimensional nature of intelligence. Jin et al. (2025) connect factor analysis of human cognition to AI capabilities. Morris et al. (2023) articulate levels of AGI based on performance percentiles. Legg and Hutter (2007) discuss various tests for general machine intelligence. Limitations. First, our conceptualization of intelligence is not exhaustive. It deliberately excludes certain faculties, such as the kinesthetic abilities proposed in alternative frameworks like Gardner's theory of multiple intelligences. Second, our illustrative examples are specific to the English language and are not culturally agnostic. Future research could involve adapting these tests across diverse linguistic and cultural contexts. Furthermore, our operationalization has inherent constraints. The General Knowledge (K) tests are necessarily selective and do not assess the full breadth of possible subject areas. A 100% AGI score represents a "highly proficient" well-educated individual who has achieved mastery across these tested dimensions, rather than well-educated in the sense of having a college degree. Moreover, while the scoring weights we employ are necessary for quantitative measurement, they represent one of many possible configurations. We give equal weight to each broad ability (10%) to prioritize breadth, but more discretionary weighting schemes could be reasonable. The results are contingent on these methodological choices, and future work could explore alternative collections of tasks and weighting schemes. Finally, while the aggregate AGI Score is provided for convenience, it could be misleading. A simple summation can obscure critical failures in bottleneck capabilities. For example, an AI system with a 90% AGI Score but 0% on Long-Term Memory Storage (MS) would be functionally impaired by a form of "amnesia," severely limiting its capabilities despite a high overall score. Therefore, we recommend reporting the AI system's cognitive profile and not just its AGI Score.
相关工作:Ilić 与 Gignac(2024)、Ren 等(2024)发现多种 AI 系统的能力与预训练算力高度相关。Gignac 与 Szodorai(2024)讨论人类心理测量学与 AI 系统的智能测试。Turing(1950)主张图灵测试可指示通用能力。Gubrud(1997)在 1997 年提出过一个早期的 AGI 定义。Marcus 等(2016)讨论超越图灵测试、捕捉智能多维度本质的必要性。Jin 等(2025)把人类认知的因子分析连接到 AI 能力。Morris 等(2023)基于表现百分位阐述 AGI 的层级。Legg 与 Hutter(2007)讨论通用机器智能的各种测试。 局限:第一,我们对智能的概念化并不穷尽。它有意排除了某些官能,如加德纳多元智能理论等替代框架提出的"身体运动智能"。第二,我们的示例针对英语、并非文化中立。未来研究可把这些测试适配到多样化的语言与文化语境。此外,我们的操作化有内在约束:常识(K)测试必然是选择性的,不评估所有可能学科领域的完整广度;100% 的 AGI 分数代表一个在这些被测维度上达到精通的"高熟练"受良好教育者,而非"有大学学历"意义上的受良好教育者。而且,我们采用的评分权重是量化测量所必需的,但只是众多可能配置之一——我们给每个宽泛能力等权(10%)以优先广度,但更灵活加权的方案也可能合理。结果取决于这些方法学选择,未来工作可探索替代的任务集合与权重方案。最后,聚合的 AGI 分数虽为方便而设,但可能误导:简单求和会掩盖瓶颈能力的致命失败。例如,一个 AGI 分数 90% 但长时记忆存储(MS)为 0% 的 AI 系统,会被一种"失忆"在功能上削弱——尽管总分很高,其能力仍严重受限。因此,我们建议报告 AI 系统的认知画像,而不只是它的 AGI 分数。
相关概念的定义:AI 能力阶梯
Definitions of Related Concepts
Some types of strategically relevant AI can arrive before or after AGI. As follows are some particularly noteworthy types of AI: 1. Pandemic AI is an AI that can engineer and produce new, infectious, and virulent pathogens that could cause a pandemic. 2. Cyberwarfare AI is an AI that can design and execute sophisticated, multi-stage cyber campaigns against critical infrastructure (e.g., energy grids, financial systems, defense networks). 3. Self-Sustaining AI is an AI that can autonomously operate indefinitely, acquire resources, and defend its existence. 4. AGI is an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult. 5. Recursive AI is an AI that can independently conduct the entire AI R&D lifecycle, leading to the creation of markedly more advanced AI systems without human input. 6. Superintelligence is an AI that greatly exceeds the cognitive performance of humans in virtually all domains of interest (Bostrom, 2014). 7. Replacement AI is an AI that performs almost all tasks more effectively and affordably, rendering human labor economically obsolete.
某些战略相关的 AI 类型可能比 AGI 更早或更晚到来。以下是一些特别值得注意的 AI 类型: 1. 大流行病 AI:能设计并制造可引发大流行病的新型、传染性、高致病病原体的 AI。 2. 网络战 AI:能针对关键基础设施(如电网、金融系统、国防网络)设计并执行复杂多阶段网络攻击的 AI。 3. 自维持 AI:能无限期自主运行、获取资源并捍卫自身存在的 AI。 4. AGI:能匹配或超越"受过良好教育的成年人"的认知通用性与熟练度的 AI。 5. 递归 AI:能在无人类输入下独立完成整个 AI 研发生命周期、从而创造出显著更先进 AI 系统的 AI。 6. 超级智能:在几乎所有重要领域都远超人类认知表现的 AI。 7. 替代性 AI:能以更高效率与更低成本完成几乎全部任务、使人类劳动在经济上过时的 AI。
Our AGI definition is about human-level AI, not economically-valuable AI, nor economy-level AI. OpenAI and Microsoft have reportedly considered AGI to be an AI that can generate $100 billion in profit (TechCrunch, 2024). We do not conflate AGI with economically valuable AI because narrow technologies, such as the iPhone, can generate billions in economic value, despite not being generally intelligent. Meanwhile, Replacement AI is about economy-level AI, and it includes physical tasks, unlike AGI. Recursive AI removes the need for human researchers and "closes the loop" on AI R&D, enabling rapid, recursive capability gains (an "intelligence recursion") without human scientific input and could potentially lead to Superintelligence.
我们的 AGI 定义关乎"人类级 AI",而非"有经济价值的 AI",也非"经济级 AI"。据报道,OpenAI 与微软曾把 AGI 视为能产生 1000 亿美元利润的 AI。我们不把 AGI 与有经济价值的 AI 混为一谈——因为像 iPhone 这样的窄技术,尽管不是通用智能,也能产生数十亿美元的经济价值。同时,"替代性 AI"关乎"经济级 AI",且包含体力任务——这不同于 AGI。 递归 AI 消除了对人类研究者的需求,"闭合"了 AI 研发的回路,在无人科学输入下实现快速、递归的能力增益(一种"智能递归"),并可能通向超级智能。
通往 AGI 的障碍
Barriers to AGI
Achieving AGI requires solving a variety of grand challenges. For example, the machine learning community's ARC-AGI Challenge aiming to measure abstract reasoning is represented in On-the-Spot Reasoning (R) tasks. Meta's attempts to create world models that include intuitive physics understanding is represented in the video anomaly detection task (V). The challenge of spatial navigation memory (WM) reflects a core goal of Fei-Fei Li's startup, World-Labs. Moreover, the challenges of hallucinations (MR) and continual learning (MS) will also need to be resolved. These significant barriers make an AGI Score of 100% unlikely in the next year.
达成 AGI 需要解决一系列宏大挑战。例如,机器学习社区旨在衡量抽象推理的 ARC-AGI 挑战,体现在即时推理(R)任务中;Meta 创建包含直觉物理理解的世界模型的尝试,体现在视频异常检测任务(V)中;空间导航记忆(WM)的挑战,反映了李飞飞创办的 World-Labs 的核心目标。此外,幻觉(MR)与持续学习(MS)的挑战也需解决。这些重大障碍使"明年内 AGI 分数达到 100%"不太可能。
← 主页