G10L17/20

Speech recognition method, electronic device, and computer storage medium

A speech recognition method includes segmenting captured voice information to obtain a plurality of voice segments, and extracting voiceprint information of the voice segments; matching the voiceprint information of the voice segments with a first stored voiceprint information to determine a set of filtered voice segments having voiceprint information that successfully matches the first stored voiceprint information; combining the set of filtered voice segments to obtain combined voice information, and determining combined semantic information of the combined voice information; and using the combined semantic information as a speech recognition result when the combined semantic information satisfies a preset rule.

ELECTRONIC DEVICE AND SPEAKER VERIFICATION METHOD OF ELECTRONIC DEVICE
20230016465 · 2023-01-19 ·

An electronic device is provided. The electronic device includes a microphone configured to receive an audio signal including a voice of a user, a sensor configured to detect a vibration signal generated by the user, at least one processor, and a memory configured to store an instruction executable by the processor. The at least one processor may be configured to determine a noise level included in the audio signal, calculate a verification score based on the noise level, the audio signal, and the vibration signal, and perform speaker verification for the user based on the verification score.

ELECTRONIC DEVICE AND SPEAKER VERIFICATION METHOD OF ELECTRONIC DEVICE
20230016465 · 2023-01-19 ·

An electronic device is provided. The electronic device includes a microphone configured to receive an audio signal including a voice of a user, a sensor configured to detect a vibration signal generated by the user, at least one processor, and a memory configured to store an instruction executable by the processor. The at least one processor may be configured to determine a noise level included in the audio signal, calculate a verification score based on the noise level, the audio signal, and the vibration signal, and perform speaker verification for the user based on the verification score.

Method and apparatus for implementing speaker identification neural network

A method and apparatus for generating a speaker identification neural network include generating a first neural network that is trained to identify a first speaker with respect to a first voice signal in a first environment, generating a second neural network for identifying a second speaker with respect to a second voice signal in a second environment, and generating the speaker identification neural network by training the second neural network based on a teacher-student training model in which the first neural network is set to a teacher neural network and the second neural network is set to a student neural network.

METHODS FOR IMPROVING THE PERFORMANCE OF NEURAL NETWORKS USED FOR BIOMETRIC AUTHENTICATIO
20220405363 · 2022-12-22 ·

A method of generating a biometric signature of a user for use in authentication using a neural network, the method comprising: receiving (110) a plurality of biometric samples from a user;

extracting at least one feature vector using the plurality of biometric samples; using the elements of the at least one feature vector as inputs for a neural network; extracting the corresponding activations from an output layer of the neural network; and generating a biometric signature of the user using the extracted activations, such that a single biometric signature represents multiple biometric samples from the user.

SYSTEMS AND METHODS FOR ENABLING VOICE-BASED TRANSACTIONS AND VOICE-BASED COMMANDS
20220406313 · 2022-12-22 ·

Aspects of the present disclosure involve processing audio signals to determine the presence and proximity of a user to a computing device, such as a voice-controlled computing device located within an environment. When the proximity of the user in comparison to the computing device is within an acceptable threshold, a voice command is detected that is associated with the user of a plurality of users located in the environment. In some instances, a device command is generated based on the voice command. The device command is executed, for example, at the computing device.

SYSTEMS AND METHODS FOR ENABLING VOICE-BASED TRANSACTIONS AND VOICE-BASED COMMANDS
20220406313 · 2022-12-22 ·

Aspects of the present disclosure involve processing audio signals to determine the presence and proximity of a user to a computing device, such as a voice-controlled computing device located within an environment. When the proximity of the user in comparison to the computing device is within an acceptable threshold, a voice command is detected that is associated with the user of a plurality of users located in the environment. In some instances, a device command is generated based on the voice command. The device command is executed, for example, at the computing device.

MACHINE LEARNING FOR IMPROVING QUALITY OF VOICE BIOMETRICS

Methods and systems are disclosed herein for improving the quality of audio for use in a biometric. A biometric system may use machine learning to determine whether audio or a portion of the audio should be used as a biometric for a user. A sample of the user's voice may be used to generate a voice signature of the user. Portions of the audio that do not meet a similarity threshold when compared with the voice signature may be removed from the audio. Additionally or alternatively, interfering noises may be detected and removed from the audio to improve the quality of a voice biometric generated from the audio.

MACHINE LEARNING FOR IMPROVING QUALITY OF VOICE BIOMETRICS

Methods and systems are disclosed herein for improving the quality of audio for use in a biometric. A biometric system may use machine learning to determine whether audio or a portion of the audio should be used as a biometric for a user. A sample of the user's voice may be used to generate a voice signature of the user. Portions of the audio that do not meet a similarity threshold when compared with the voice signature may be removed from the audio. Additionally or alternatively, interfering noises may be detected and removed from the audio to improve the quality of a voice biometric generated from the audio.

Speaker recognition method and system

A speaker recognition system for assessing the identity of a speaker through a speech signal based on speech uttered by said speaker is provided. The system includes a framing module that subdivides the speech signal over time into a set of frames, and a filtering module that analyzes the frames of the set to discard frames affected by noise and frames which do not comprise a speech, based on a spectral analysis of the frames. A feature extraction module extracts audio features from frames which have not been discarded, and a classification module processes the audio features extracted from the frames which have not been discarded for assessing the identity of the speaker.