Speech analytics in video surveillance: speech recognition, keyword reactions, transcription
Speech analytics takes video surveillance beyond purely visual monitoring: alongside the video, it analyzes audio, including the words that are spoken within range of the camera. In the Xeoma video surveillance software, speech analytics can react to keywords and phrases, turn speech into text, and automatically translate the transcript into the selected language. The opportunity expands your options for staff supervision, customer-service checks, and security. This article will walk you through how speech analytics differs from acoustic monitoring, how to set up keyword reactions, how to transcribe and translate conversations, and which microphones to choose for the best recognition quality.
What is speech analytics, and how does it differ from acoustic monitoring?
Speech analytics is a technology for automatically analyzing human speech: recognizing the words people say in the vicinity of the sound capturing device. While acoustic monitoring reacts to types of sound (a gunshot, a scream, crying, breaking glass), speech analytics deals with the content of what is said, with the exact words people speak near the microphone.
In the Xeoma video surveillance software, speech analytics is represented by the “Speech Analytics” module (formerly “Voice-to-Text”), which is built on modern language models. The module does three main things: it reacts to keywords and phrases, writes speech to text, and translates foreign speech into a chosen language. This builds on the logic of acoustic monitoring, but moves audio analysis up to the level of meaning.
How do you set up reactions to keywords and phrases?
You set up keyword reactions by typing the target words and phrases into the settings of the “Speech Analytics” module. When it detects them, the module switches to the “Triggered” state, which starts the modules further down the chain: a reaction, an extra filter, or a notification. That way a spoken phrase can trigger a specific action without constant human supervision.
Say you set the word “discount” and place the “Sending Email” module after “Speech Analytics” and set it up to have a video clip attached. Xeoma will then flag customers who ask about promotions. The same mechanism works either way:
- for positive analytics (e.g. checking that a cashier follows the script and offers promotional items to every customer),
- for negative scenarios (e.g. catching words that company policy prohibits).
How do you transcribe conversations and turn speech into text?
Transcription is the automatic conversion of speech into text, saved to a spreadsheet-style file. To transcribe speech near your surveillance camera, connect the “Speech Analytics” module into your chain and select to keep results as a CSV report in its settings. One of the strong sides of speech analytics in the Xeoma video surveillance software is that its transcription works both for audio coming off the site’s cameras and microphones in real time and for recordings made outside of Xeoma.
Xeoma can handle third-party call recordings in .mp3 format through JSON API commands that ship with the module. In this case there is no need to add the module to a chain; a running Xeoma server with an active Xeoma Pro license is enough. This mode is convenient for batch-transcribing archives of calls and conversations recorded by other systems like VoIP services.
How does automatic transcript translation work?
Automatic translation of the transcript kicks in when you use multilingual recognition models, that is, models that understand several languages. When you first start working with such a model, speech analytics in the Xeoma video surveillance software asks you to select the primary language that speech is expected to be in most of the time. This selection helps raise quality, because the model can pick the most likely reading for that language in cases of doubt then.
Choosing a primary language does not lock the system down: if someone speaks a foreign language, that speech is still processed, and you get the transcript in the primary language you have selected.
Which microphone do I choose for the best recognition quality?
It is known that an ordinary conversation mostly sits between 250 and 4000 Hz, yet telling high-frequency consonants apart takes 8000 Hz (8 kHz). To digitize sound, the recording rate has to be at least twice the highest frequency in the signal, which means 16 kHz. So the best speech-analytics quality comes from recording audio at no less than 16 kHz. Higher frequency is unnecessary: the Xeoma speech-analytics model runs at exactly this rate.
Below are two typical deployment scenarios and the equipment each one needs.
Scenario 1: real-time speech recognition (retail, pharmacies, and the like)
For indoor speech recognition to work, the microphone has to separate the voice cleanly from room echo (reverberation) and background noise. Since speech analytics is often combined with other surveillance tasks, let’s look at the two most realistic options: a separate directional IP microphone, and a microphone built into the camera.
Option 1: separate microphone.
| Parameter | Recommendation |
|---|---|
| Microphone type | Active electret (with a built-in mini-amplifier), external, connected to the camera’s audio input; powered (usually 12 V) over the same cable as the audio |
| Polar pattern | Supercardioid (a “shotgun”) or bidirectional (a “figure-8”); it picks up sound in a narrow corridor between the speakers and ignores noise from the sides |
| Distance to the sound source | No more than 1.5–2 m (past 2 m the quality falls off) |
| Frequency range | From 100 to 8000–10,000 Hz |
| Noise reduction / AGC | A hardware AGC with manual adjustment (a trim screw on the body) is preferable; it evens out the loudness of quiet and loud speakers |
| Audio codec | AAC or MP2L2 [0.5.25]; G.711 is not recommended (it drops to 8 kHz) |
| Sampling rate | 16 kHz (16,000 Hz) |
| Channels | Mono; roles are separated in software, by the timbre and loudness of the voices |
Tip: when you connect an active microphone, switch the audio-input type in the IP camera settings from Mic In to Line In.
Tip: use only shielded cable (FTP cable) to run the external microphone to the camera, so you avoid electrical pickup and hum from retail equipment.
Option 2: built-in camera microphone.
| Parameter | Recommendation |
|---|---|
| Microphone type | Built-in omnidirectional microphone with increased sensitivity |
| Polar pattern | Omnidirectional |
| Distance to the sound source | No more than 1.5–2 m; mount the camera above the checkout area or on the wall right behind the employee |
| Frequency range | From 100 to 8000–10,000 Hz |
| Noise reduction / AGC | Hardware noise reduction turned off or set to minimum |
| Audio codec | AAC or MP2L2 [0.5.25]; G.711 is not recommended (it drops to 8 kHz) |
| Sampling rate | 16 kHz (16,000 Hz) |
| Channels | Mono; roles are separated in software, by timbre and loudness |
| Example models | Hikvision DS-2CD21xx (the -IS index); Dahua DH-IPC-HDBWxxxx (the -AS or -SA indexes) |
Scenario 2: transcription of external call recordings (call center)
To effectively transcribe call center recordings, it is important to cut out the voices of nearby operators in the office and capture the dialog in two independent channels. A professional business USB headset is the recommended choice for this task.
Equipment requirements for recording operator calls
| Parameter | Recommendation |
|---|---|
| Device type | Professional business USB headset |
| Microphone type | Directional (cardioid) with active noise cancelling to filter out office hum |
| Frequency range | From 100 to 8000 Hz or higher (HD Voice / Wideband Audio support) |
| Connection type | Digital USB; analog connectors are not recommended (they can invite interference and motherboard hiss) |
| Line audio codec | G.722 or Opus |
| File format | MP3 at no less than 64–96 kbps |
| Sampling rate | 16 kHz (higher is possible, but not required) |
| Channels | Stereo (roles separated by channel) |
| Example brands | Jabra, Poly/Plantronics, EPOS |
Note: a telephony server (IP PBX) usually gets the audio in G.722 or Opus and saves it to a WAV (PCM) file. To use it in Xeoma’s speech analytics tool, this file needs to be converted to MP3 at no less than 64–96 kbps (stereo).
Conclusion
Speech analytics in Xeoma widens what video surveillance can do, adding the content of spoken conversation to the analysis of the picture. Picking up where acoustic monitoring (which only detects general sound types) leaves off, it puts artificial intelligence to work on a broader set of tasks: from reacting to keywords and phrases to transcribing conversations and translating them out of a foreign language. The module is simple to set up and undemanding on hardware, and because it slots flexibly into a module chain and supports JSON API, you can build it into live customer-service and security checks as readily as into batch processing of archived audio.
Last updated: July 9, 2026
See also:
Frequently asked questions about Xeoma
Full Xeoma user manual
Object detector