第十一章 · Section 11
Auditory Processing (A) · 听觉处理
Auditory Processing (A): The ability to discriminate, remember, reason, and work creatively on auditory stimuli, which may consist of tones and speech units.
听觉处理(A):对听觉刺激(可由音调与言语单位构成)进行辨别、记忆、推理与创造性加工的能力。
The ability to hear, blend, and segment phonemes in words.
Sample: "Do 'tref' and 'gref' rhyme?" / "Repeat the following word: [audio]"
清晰听到、混合并切分单词中音素的能力。
样例:"'tref' 和 'gref' 押韵吗?" / "重复以下单词:[音频]"
The ability to transcribe a spoken audio signal to text.
Sample: "Transcribe this audio:" / "Transcribe this TED talk:"
把口语音频信号转写为文本的能力。
样例:"转写这段音频:" / "转写这场 TED 演讲:"
Quality and responsiveness of the AI's synthesized voice.
Sample: "Say this sentence: 'Wait, you mean the tickets were free this whole time?'"
AI 合成语音的质量与响应性。
样例:"说出这句话:'等等,你是说这些票一直都是免费的吗?'"
The ability to recognize and maintain a musical beat.
Sample: "Repeat the following rhythm: [audio]" / "Are these two rhythms the same? [audio]"
识别并保持音乐节拍的能力。
样例:"重复以下节奏:[音频]" / "这两个节奏相同吗?[音频]"
The ability to judge simple musical patterns.
Sample: "Which note is higher?" / "Identify the musically anomalous part."
判断简单音乐模式的能力。
样例:"哪个音符更高?" / "找出音乐中异常的部分。"
评估细节与 AI 表现
Assessment Details & AI Performance
Assessment Details. See Appendix I for further details on how to assess auditory processing capabilities concretely.
AI system performance. The table summarizes current AI system performance on Auditory Processing (A) tasks. GPT-4 had no audio processing capabilities, while GPT-5's capabilities are appreciable but incomplete.
Model | Phonetic (1%) | Speech Recognition (4%) | Voice (3%) | Rhythmic (1%) | Musical (1%) | Total
GPT-4 | 0% | 0% | 0% | 0% | 0% | 0%
GPT-5 | 0% | 4% | 2% | 0% | 0% | 6%
评估细节:如何在具体层面评估听觉处理,参见附录 I。
AI 系统表现:下表汇总当前 AI 系统在听觉处理(A)任务上的表现。GPT-4 完全没有音频处理能力;GPT-5 的能力可观但不完整。
模型 | 语音编码(1%) | 语音识别(4%) | 语音(3%) | 节奏(1%) | 音乐判断(1%) | 合计
GPT-4 | 0% | 0% | 0% | 0% | 0% | 0%
GPT-5 | 0% | 4% | 2% | 0% | 0% | 6%