Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

· Source: Computation and Language · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Natural Language Processing · Depth: Advanced, quick

Summary

This study presents a preliminary comparison of advanced automatic speech recognition (ASR) systems and Dutch native listeners on "diverse" Dutch speech, including child and older adults' speech, and Flemish accents. Contrary to the assumption that humans represent the upper-bound, ASR systems demonstrated similar performance to listeners, with Google Telephony specifically outperforming other ASR systems and, in certain instances, even surpassing human recognition. Slight performance variations between ASR and human listeners were observed concerning speaker's age, regional accents, and utterance length. The research also highlighted the significant impact of specific test sets on benchmarking conclusions. Future work should enhance ASR robustness to acoustic variability from aging and regional accents.

Key takeaway

For machine learning engineers developing ASR systems for diverse populations, this research indicates that current ASR can achieve human-level performance, especially with systems like Google Telephony. You should rigorously benchmark your models against human listeners using varied speech datasets, particularly focusing on child, older adult, and regional accent speech. Prioritize improving model robustness to acoustic variability related to speaker age and regional accents to ensure broader applicability and superior real-world performance.

Key insights

ASR can match or exceed human speech recognition on diverse speech, challenging traditional performance benchmarks.

Principles

Method

The study compared ASR systems and Dutch native listeners on Dutch child, older adult, and Flemish speech samples.

In practice

Topics

Best for: AI Engineer, Research Scientist, AI Scientist, Machine Learning Engineer, NLP Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.