Voice-to-Text
Xeoma’s intellectual module for speech recognition
The module ‘listens’ to the audio stream from a camera or a separate microphone, recognizes speech, reacts to
keywords and saves the transcript of conversations in a CSV report or overlays it as text on the live
video/archive depending on your preferences. It can also transcribe .mp3 files: recordings of conversations, training videos, etc.
Important: beta version
The module is available starting from Xeoma 24.8.12 and is in the beta state, so in some cases it can skip words or contain loops. No specialized equipment is required: the sound stream from any camera or microphone and a regular off-the-shelf computer with a GPU are suitable.
Application scenarios
Voice-to-Text helps improve security in private life, the life of the city and citizens, as well as in the commercial sphere, and contributes to the optimization of business operations.
Taking care of the elderly
The ability to instantly react to a cry for help or protect them from scammers.
Any business
Automated control of the customer service quality (for example, detection of swear words).
Call center
Transcription of call recordings to monitor compliance with the company policy and conversation scripts.
City surveillance
Counter-terrorism security and recognition of words that promise danger.
Parental control
Assistance in protecting a child against bullying or communicating with scammers.
Police
In addition to body-worn cameras — transcription of conversations with a suspect and detection of threats.
Research and analytics
Background collection of statistics for speech-related studies and word-use frequency.
Marketing
Find out whether customers are discussing a promotional campaign, their reaction to a banner, etc.
Filtering and automation
Detection of unwanted or key phrases for targeted control of conversations — without having to listen to all of them.
Advantages of the Voice-to-Text module
No special equipment required
Works on regular computers with almost any camera and any graphics card.
Simply flexible
Various reactions (including your own programmable ones) and integration with third-party systems.
Real-time work
On-the-fly work with streams, without any latency. Works on your own hardware.
Affordable solution
The module can be purchased in addition to the Standard or Pro license.
How it works
Build a chain
Since the module works with an audio stream, you need a sound source in the chain.
- Sound source: a microphone built into the camera, or a separate USB or IP microphone
- Sample chain: “Universal Camera” – “Voice-to-Text” – “Preview and Archive”
- The module is licensed additionally to a Xeoma Pro or Standard license
Download AI resources
Click on the Voice-to-Text icon in the chain to open the module settings. The download of additional resources will start automatically when you first open the settings. When the “Downloading in progress” message disappears, everything is ready.
- Additional resources contain data arrays for artificial intelligence and are downloaded on request from FelenaSoft servers
- Resources are downloaded only when needed — this keeps the program size small
Choose a recognition model and keywords
Several AI models are available for different languages, differing in size, recognition quality and load on the hardware. In the “Speech recognition model” field, select the language of the transcript (the language of the speech itself does not need to be specified) and, if required, specify keywords for detection in the field below. The “Display of recognized speech during real-time viewing” and “Save recognized speech in subtitles when recording/exporting the archive” switches let you overlay the transcript on the live video and in the archive.
- “Keywords for detection” — the module reacts only to the words or phrases you need
- Transcript overlay in live preview and in the archive
- “Save data in CSV report” — the transcript is saved to a spreadsheet file on the disk
- Connect a reaction after the module: recording, notification, sending a command
- The
%VOICE%macro is available in the “HTTP Request Sender”, “Email Sending” and “Application Runner” modules
Integration with external programs
Voice-to-Text can be used from external programs — for example, to transcribe VoIP conversations. Give an .mp3 file to the module and get the result as text — even on operator workstations where there is no Xeoma or cameras. Important: only .mp3 files are supported.
Example of a request via the Xeoma API (JSON):
curl -F "audio_file=@speech.mp3" "http://192.168.0.67:10090/api?login=Administrator&password=123&speech_recognition=recognition&model=qwen3-asr-1.7b&ttl=10&language=en"
Where:
audio_file=@speech.mp3— the path to the audio file on your computer192.168.0.67— the IP address of the Xeoma server10090— the Xeoma server portAdministrator/123— the Administrator login (always Administrator) and passwordmodel=…— the recognition model (see the list below)language=en— the language of the transcript. If it differs from the actual speech language, the transcript will be automatically translated to the language you specified
To save the result as a text file, add >savetext.txt after the command, where savetext.txt is the name of the text file.
Recognition models for the API:
| Name | Description | API code |
|---|---|---|
| Russian basic | 140 MB | tone |
| Russian enhanced | punctuation, 215 MB | giga |
| English basic | 70 MB | zipformer |
| Multilingual basic | punctuation, 1 GB, 0.6B parameters | qwen3-asr-0.6b |
| Multilingual enhanced | punctuation, 2.2 GB, 1.7B parameters | qwen3-asr-1.7b |
| Gemma4 multilingual basic | punctuation, 2.2 GB, 1.7B parameters | gemma4-e2b |
| Gemma4 multilingual enhanced | punctuation, 7 GB, 12B parameters | gemma4-12b |
List of language codes
en — English, zh — Chinese, de — German, es — Spanish, ru — Russian, ko — Korean, fr — French, ja — Japanese, pt — Portuguese, tr — Turkish, pl — Polish, ca — Catalan, nl — Dutch, ar — Arabic, sv — Swedish, it — Italian, id — Indonesian, hi — Hindi, fi — Finnish, vi — Vietnamese, he — Hebrew, uk — Ukrainian, el — Greek, ms — Malay, cs — Czech, ro — Romanian, da — Danish, hu — Hungarian, ta — Tamil, no — Norwegian, th — Thai, ur — Urdu, hr — Croatian, bg — Bulgarian, lt — Lithuanian, la — Latin, mi — Maori, ml — Malayalam, cy — Welsh, sk — Slovak, te — Telugu, fa — Persian, lv — Latvian, bn — Bengali, sr — Serbian, az — Azerbaijani, sl — Slovenian, kn — Kannada, et — Estonian, mk — Macedonian, br — Breton, eu — Basque, is — Icelandic, hy — Armenian, ne — Nepali, mn — Mongolian, bs — Bosnian, kk — Kazakh, sq — Albanian, sw — Swahili, gl — Galician, mr — Marathi, pa — Punjabi, si — Sinhala, km — Khmer, sn — Shona, yo — Yoruba, so — Somali, af — Afrikaans, oc — Occitan, ka — Georgian, be — Belarusian, tg — Tajik, sd — Sindhi, gu — Gujarati, am — Amharic, yi — Yiddish, lo — Lao, uz — Uzbek, fo — Faroese, ht — Haitian Creole, ps — Pashto, tk — Turkmen, nn — Nynorsk, mt — Maltese, sa — Sanskrit, lb — Luxembourgish, my — Myanmar, bo — Tibetan, tl — Tagalog, mg — Malagasy, as — Assamese, tt — Tatar, haw — Hawaiian, ln — Lingala, ha — Hausa, ba — Bashkir, jw — Javanese, su — Sundanese, yue — Cantonese.
How to test
Download and launch Xeoma. Use the Trial edition or activate Xeoma Standard/Xeoma Pro license + “Voice-to-text” additional module license.
Add a camera manually or wait while Xeoma finds cameras in your network automatically. For a separate microphone, connect the “Microphone” module and select the appropriate sound source.
Add the “Voice-to-Text” module to the chain, set up a reaction to keywords or continuous speech transcription and, if needed, enable the transcript overlay on the camera image and recordings.
Set a reaction to the event — saving to CSV, mobile notifications or any other.
demo license
Other modules that process audio streams
These modules work on their own or together with Voice-to-Text.
Select a USB microphone or a separate IP microphone as the sound source.
Analyzes audio streams and triggers when the sound level exceeds a specified limit.
Recognizes car alarms, a child crying, gunshots, screams, breaking glass.
Voice-to-Text: Speech Recognition in Xeoma
Try Xeoma for free
Enter your name and email to get a free demo license. No credit card required.
Get Xeoma free demo licenses to email
Full functionality. Sent instantly.
We urge you to refrain from using emails that contain personal data, and from sending us personal data in any other way. If you still do, by submitting this form, you confirm your consent to the processing of your personal data
Do you need something else?
The module doesn’t quite fit your conditions? We can develop the functions you need and add them into Xeoma as paid development. See details.
Have questions? Need help? Please contact us! We’ll be happy to help!
Ready to start?
Test Voice-to-Text with a free demo license — no credit card and no personal data required.
