第六章 · Section 6
On-the-Spot Reasoning (R) · 即时推理
On-the-Spot Reasoning (R): The deliberate but flexible control of attention to solve novel "on the spot" problems that cannot be performed by relying exclusively on previously learned habits, schemas, and scripts.
即时推理(R):刻意但灵活地控制注意力,去解决无法仅靠先前习得的习惯、图式与脚本完成的"现场"新问题。
Reasoning from general statements or premises to reach a logically guaranteed conclusion.
Sample task: David knows Mr. Zhang's friend Jack, and Jack knows David's friend Ms. Lin. Everyone of them who knows Jack has a master's degree, and everyone of them who knows Ms. Lin is from Shanghai. Who is from Shanghai and has a master's degree?
从一般陈述或前提出发,推出逻辑上必然成立的结论。
样例问题:David 认识张先生的朋友 Jack,而 Jack 认识 David 的朋友林女士。他们中认识 Jack 的人都有硕士学位,认识林女士的人都来自上海。谁既来自上海又有硕士学位?
Discovering the underlying principles or rules that determine a phenomenon's behavior.
Sample task: The can of Pringles has moldy chips in it. Mary picks up the can in the supermarket and walks to the cashier. Is Mary likely to be aware that "The can of Pringles has moldy chips in it"?
发现决定某一现象行为的底层原理或规则。
样例问题:一罐品客薯片里有发霉的薯片。Mary 在超市拿起这罐薯片走向收银台。Mary 很可能知道"这罐品客薯片里有发霉的薯片"吗?
Attributing mental states to others and understanding how those states may differ from one's own.
把心理状态归因于他人,并理解这些状态如何与自己不同。
Devising a sequence of actions to achieve a specific goal.
Sample task: You plan a 14-day trip to 3 European cities, taking only direct flights between. You'll stay 4 days in London, 5 days in Bucharest, and 7 days in Reykjavik. You need to meet a friend in Bucharest between days 10 and 14. Direct flights are available between London and Bucharest, and between London and Reykjavik. Find a 14-day travel plan that satisfies these conditions.
设计一连串动作以达到特定目标。
样例问题:你计划一趟 3 个欧洲城市的 14 天旅行,城市间只坐直飞航班。你将在伦敦待 4 天、布加勒斯特 5 天、雷克雅未克 7 天。你需要在第 10 到 14 天之间在布加勒斯特见一位朋友。伦敦-布加勒斯特、伦敦-雷克雅未克之间都有直飞航班。找出满足这些条件的 14 天旅行方案。
The ability to infer unstated classification rules from a sequence of simple performance feedback.
Sample task: Wisconsin Card Sorting Test—classify cards according to an inferred rule that changes over time, based only on "right/wrong" feedback.
从一系列简单的表现反馈中推断未明说的分类规则。
样例问题:威斯康星卡片分类测验——仅凭"对/错"反馈,按一条随时间变化的推断规则对卡片分类。
评估细节与 AI 表现
Assessment Details & AI Performance
Assessment Details. See Appendix D for further details on how to assess on-the-spot reasoning capabilities concretely.
AI System Performance. The table summarizes current AI system performance on On-the-Spot Reasoning (R) tasks. GPT-4 has negligible on-the-spot reasoning capabilities, while GPT-5 only has some remaining gaps.
Model | Deduction (2%) | Induction (4%) | Theory of Mind (2%) | Planning (1%) | Adaptation (1%) | Total
GPT-4 | 0% | 0% | 0% | 0% | 0% | 0%
GPT-5 | 2% | 2% | 2% | 1% | 0% | 7%
评估细节:如何在具体层面评估即时推理,参见附录 D。
AI 系统表现:下表汇总当前 AI 系统在即时推理(R)任务上的表现。GPT-4 的即时推理能力几乎为零,GPT-5 仍有部分缺口。
模型 | 演绎(2%) | 归纳(4%) | 心智理论(2%) | 规划(1%) | 适应(1%) | 合计
GPT-4 | 0% | 0% | 0% | 0% | 0% | 0%
GPT-5 | 2% | 2% | 2% | 1% | 0% | 7%