System and method for multi-user GPU-accelerated speech recognition engine for client-server architectures

Invention Grant

US10453445B2 System and method for multi-user GPU-accelerated speech recognition engine for client-server architectures 有权

Please log in to see more content

Patent Title: System and method for multi-user GPU-accelerated speech recognition engine for client-server architectures
Application No.: US15435251

Application Date: 2017-02-16
Publication No.: US10453445B2

Publication Date: 2019-10-22
Inventor: Ian Richard Lane , Jungsuk Kim
Applicant: CARNEGIE MELLON UNIVERSITY
Applicant Address: US PA Pittsburgh
Assignee: CARNEGIE MELLON UNIVERSITY
Current Assignee: CARNEGIE MELLON UNIVERSITY
Current Assignee Address: US PA Pittsburgh
Agent Michael G. Monyok
Main IPC: G10L15/00
IPC: G10L15/00 ; G10L15/14 ; G10L15/28

System and method for multi-user GPU-accelerated speech recognition engine for client-server architectures

Abstract:

Disclosed herein is a GPU-accelerated speech recognition engine optimized for faster than real time speech recognition on a scalable server-client heterogeneous CPU-GPU architecture, which is specifically optimized to simultaneously decode multiple users in real-time. In order to efficiently support real-time speech recognition for multiple users, a “producer/consumer” design pattern is applied to decouple speech processes that run at different rates in order to handle multiple processes at the same time. Furthermore, the speech recognition process is divided into multiple consumers in order to maximize hardware utilization. As a result, the platform architecture is able to process more than 45 real-time audio streams with an average latency of less than 0.3 seconds using one-million-word vocabulary language models.

Public/Granted literature

US20170236518A1 System and Method for Multi-User GPU-Accelerated Speech Recognition Engine for Client-Server Architectures Public/Granted day:2017-08-17

Information query

Espacenet

IPC分类:

G	物理
G10	乐器；声学
G10L	语音分析或合成；语音识别；语音或声音处理；语音或音频编码或解码
G10L15/00	语音识别（G10L17/00优先）