G10L25/87

Speech recognition method of and system for determining the status of an answered telephone during the course of an outbound telephone call

A system for determining the status of an answered telephone during the course of an outbound telephone call includes an automated telephone calling device for placing a telephone call to a location having a telephone number at which a target person is listed, upon the telephone call being answered, initiating a prerecorded greeting which asks for the target person and receiving a spoken response from an answering person and a speech recognition device for performing a speech recognition analysis on the spoken response to determine a status of the spoken response. If the speech recognition device determines that the answering person is the target person, the speech recognition device initiates a speech recognition application with the target person.

Speech recognition method of and system for determining the status of an answered telephone during the course of an outbound telephone call

A system for determining the status of an answered telephone during the course of an outbound telephone call includes an automated telephone calling device for placing a telephone call to a location having a telephone number at which a target person is listed, upon the telephone call being answered, initiating a prerecorded greeting which asks for the target person and receiving a spoken response from an answering person and a speech recognition device for performing a speech recognition analysis on the spoken response to determine a status of the spoken response. If the speech recognition device determines that the answering person is the target person, the speech recognition device initiates a speech recognition application with the target person.

ACCELEROMETER-BASED ENDPOINTING MEASURE(S) AND /OR GAZE-BASED ENDPOINTING MEASURE(S) FOR SPEECH PROCESSING
20230197071 · 2023-06-22 ·

An overall endpointing measure can be generated based on an audio-based endpointing measure and (1) an accelerometer-based endpointing measure and/or (2) a gaze-based endpointing measure. The overall endpointing measure can be used in determining whether a candidate endpoint is an actual endpoint. Various implementations include generating the audio-based endpointing measure by processing an audio data stream, capturing a spoken utterance of a user, using an audio model. Various implementations additionally or alternatively include generating the accelerometer-based endpointing measure by processing a stream of accelerometer data using an accelerometer model. Various implementations additionally or alternatively include processing an image data stream using a gaze model to generate the gaze-based endpointing measure.

ADDING BACKGROUND SOUND TO SPEECH-CONTAINING AUDIO DATA
20170352361 · 2017-12-07 ·

An editing method facilitates the task of adding background sound to speech-containing audio data so as to augment the listening experience. The editing method is executed by a processor in a computing device and comprises obtaining characterization data that characterizes time segments in the audio data by at least one of topic and sentiment; deriving, for a respective time segment in the audio data and based on the characterization data, a desired property of a background sound to be added to the audio data in the respective time segment, and providing the desired property for the respective time segment so as to enable the audio data to be combined, within the respective time segment, with background sound having the desired property. The background sound may be selected and added automatically or by manual user intervention.

Method for user voice input processing and electronic device supporting same

According to an embodiment, disclosed is an electronic device including a speaker, a microphone, a communication interface, a processor operatively connected to the speaker, the microphone, and the communication interface, and a memory operatively connected to the processor. The memory stores instructions that, when executed, cause the processor to receive a first utterance through the microphone, to determine a speaker model by performing speaker recognition on the first utterance, to receive a second utterance through the microphone after the first utterance is received, to detect an end-point of the second utterance, at least partially using the determined speaker model. Besides, various embodiments as understood from the specification are also possible.

Method for user voice input processing and electronic device supporting same

According to an embodiment, disclosed is an electronic device including a speaker, a microphone, a communication interface, a processor operatively connected to the speaker, the microphone, and the communication interface, and a memory operatively connected to the processor. The memory stores instructions that, when executed, cause the processor to receive a first utterance through the microphone, to determine a speaker model by performing speaker recognition on the first utterance, to receive a second utterance through the microphone after the first utterance is received, to detect an end-point of the second utterance, at least partially using the determined speaker model. Besides, various embodiments as understood from the specification are also possible.

Method and Device for Voice Activity Detection
20220375493 · 2022-11-24 ·

In accordance with an example embodiment of the present invention, disclosed is a method and an apparatus for voice activity detection (VAD). The VAD comprises creating a signal indicative of a primary VAD decision and determining hangover addition. The determination on hangover addition is made in dependence of a short term activity measure and/or a long term activity measure. A signal indicative of a final VAD decision is then created.

Method and Device for Voice Activity Detection
20220375493 · 2022-11-24 ·

In accordance with an example embodiment of the present invention, disclosed is a method and an apparatus for voice activity detection (VAD). The VAD comprises creating a signal indicative of a primary VAD decision and determining hangover addition. The determination on hangover addition is made in dependence of a short term activity measure and/or a long term activity measure. A signal indicative of a final VAD decision is then created.

METHOD AND APPARATUS FOR PROVIDING INTERPRETATION SITUATION INFORMATION

A method is provided. The method includes receiving a speech input in a first language from a first device; obtaining, by using an artificial intelligence (AI) model, an estimated interpretation time that indicates a time expected to be required to interpret the speech input in the first language into a second language; transmitting, based on the estimated interpretation time, interpretation situation information to at least one of the first device or a second device; interpreting the speech input in the first language into the second language; and transmitting, to the second device a result of the interpreting of the speech input into the second language.

METHOD AND APPARATUS FOR PROVIDING INTERPRETATION SITUATION INFORMATION

A method is provided. The method includes receiving a speech input in a first language from a first device; obtaining, by using an artificial intelligence (AI) model, an estimated interpretation time that indicates a time expected to be required to interpret the speech input in the first language into a second language; transmitting, based on the estimated interpretation time, interpretation situation information to at least one of the first device or a second device; interpreting the speech input in the first language into the second language; and transmitting, to the second device a result of the interpreting of the speech input into the second language.