AI-powered audio analytics

Voice-to-Text

Xeoma’s intellectual module for speech recognition

Voice-to-Text module icon The module ‘listens’ to the audio stream from a camera or a separate microphone, recognizes speech, reacts to keywords and saves the transcript of conversations in a CSV report or overlays it as text on the live video/archive depending on your preferences. It can also transcribe .mp3 files: recordings of conversations, training videos, etc.

Voice-to-Text: Xeoma's Intellectual Module for Speech Recognition
90+ transcription languages
CSV report saved to disk
Live streams and .mp3 files
AI additional module
Important note: the module is in beta

Important: beta version

The module is available starting from Xeoma 24.8.12 and is in the beta state, so in some cases it can skip words or contain loops. No specialized equipment is required: the sound stream from any camera or microphone and a regular off-the-shelf computer with a GPU are suitable.

Use cases

Application scenarios

Voice-to-Text helps improve security in private life, the life of the city and citizens, as well as in the commercial sphere, and contributes to the optimization of business operations.

Taking care of the elderly

The ability to instantly react to a cry for help or protect them from scammers.

Any business

Automated control of the customer service quality (for example, detection of swear words).

Call center

Transcription of call recordings to monitor compliance with the company policy and conversation scripts.

City surveillance

Counter-terrorism security and recognition of words that promise danger.

Parental control

Assistance in protecting a child against bullying or communicating with scammers.

Police

In addition to body-worn cameras — transcription of conversations with a suspect and detection of threats.

Research and analytics

Background collection of statistics for speech-related studies and word-use frequency.

Marketing

Find out whether customers are discussing a promotional campaign, their reaction to a banner, etc.

Filtering and automation

Detection of unwanted or key phrases for targeted control of conversations — without having to listen to all of them.

Advantages

Advantages of the Voice-to-Text module

No special equipment required

Works on regular computers with almost any camera and any graphics card.

Simply flexible

Various reactions (including your own programmable ones) and integration with third-party systems.

Real-time work

On-the-fly work with streams, without any latency. Works on your own hardware.

Affordable solution

The module can be purchased in addition to the Standard or Pro license.

Workflow

How it works

A sample of a chain with the Voice-to-Text intellectual module
Step 1

Build a chain

Since the module works with an audio stream, you need a sound source in the chain.

  • Sound source: a microphone built into the camera, or a separate USB or IP microphone
  • Sample chain: “Universal Camera” – “Voice-to-Text” – “Preview and Archive”
  • The module is licensed additionally to a Xeoma Pro or Standard license
Downloading additional AI resources
Step 2

Download AI resources

Click on the Voice-to-Text icon in the chain to open the module settings. The download of additional resources will start automatically when you first open the settings. When the “Downloading in progress” message disappears, everything is ready.

  • Additional resources contain data arrays for artificial intelligence and are downloaded on request from FelenaSoft servers
  • Resources are downloaded only when needed — this keeps the program size small
Choosing the AI model and transcription language
Step 3

Choose a recognition model and keywords

Several AI models are available for different languages, differing in size, recognition quality and load on the hardware. In the “Speech recognition model” field, select the language of the transcript (the language of the speech itself does not need to be specified) and, if required, specify keywords for detection in the field below. The “Display of recognized speech during real-time viewing” and “Save recognized speech in subtitles when recording/exporting the archive” switches let you overlay the transcript on the live video and in the archive.

  • “Keywords for detection” — the module reacts only to the words or phrases you need
  • Transcript overlay in live preview and in the archive
  • “Save data in CSV report” — the transcript is saved to a spreadsheet file on the disk
  • Connect a reaction after the module: recording, notification, sending a command
  • The %VOICE% macro is available in the “HTTP Request Sender”, “Email Sending” and “Application Runner” modules
API

Integration with external programs

Voice-to-Text can be used from external programs — for example, to transcribe VoIP conversations. Give an .mp3 file to the module and get the result as text — even on operator workstations where there is no Xeoma or cameras. Important: only .mp3 files are supported.

Example of a request via the Xeoma API (JSON):

curl -F "audio_file=@speech.mp3" "http://192.168.0.67:10090/api?login=Administrator&password=123&speech_recognition=recognition&model=qwen3-asr-1.7b&ttl=10&language=en"

Where:

  • audio_file=@speech.mp3 — the path to the audio file on your computer
  • 192.168.0.67 — the IP address of the Xeoma server
  • 10090 — the Xeoma server port
  • Administrator / 123 — the Administrator login (always Administrator) and password
  • model=… — the recognition model (see the list below)
  • language=en — the language of the transcript. If it differs from the actual speech language, the transcript will be automatically translated to the language you specified

To save the result as a text file, add >savetext.txt after the command, where savetext.txt is the name of the text file.

Recognition models for the API:

Name Description API code
Russian basic140 MBtone
Russian enhancedpunctuation, 215 MBgiga
English basic70 MBzipformer
Multilingual basicpunctuation, 1 GB, 0.6B parametersqwen3-asr-0.6b
Multilingual enhancedpunctuation, 2.2 GB, 1.7B parametersqwen3-asr-1.7b
Gemma4 multilingual basicpunctuation, 2.2 GB, 1.7B parametersgemma4-e2b
Gemma4 multilingual enhancedpunctuation, 7 GB, 12B parametersgemma4-12b

List of language codes

en — English, zh — Chinese, de — German, es — Spanish, ru — Russian, ko — Korean, fr — French, ja — Japanese, pt — Portuguese, tr — Turkish, pl — Polish, ca — Catalan, nl — Dutch, ar — Arabic, sv — Swedish, it — Italian, id — Indonesian, hi — Hindi, fi — Finnish, vi — Vietnamese, he — Hebrew, uk — Ukrainian, el — Greek, ms — Malay, cs — Czech, ro — Romanian, da — Danish, hu — Hungarian, ta — Tamil, no — Norwegian, th — Thai, ur — Urdu, hr — Croatian, bg — Bulgarian, lt — Lithuanian, la — Latin, mi — Maori, ml — Malayalam, cy — Welsh, sk — Slovak, te — Telugu, fa — Persian, lv — Latvian, bn — Bengali, sr — Serbian, az — Azerbaijani, sl — Slovenian, kn — Kannada, et — Estonian, mk — Macedonian, br — Breton, eu — Basque, is — Icelandic, hy — Armenian, ne — Nepali, mn — Mongolian, bs — Bosnian, kk — Kazakh, sq — Albanian, sw — Swahili, gl — Galician, mr — Marathi, pa — Punjabi, si — Sinhala, km — Khmer, sn — Shona, yo — Yoruba, so — Somali, af — Afrikaans, oc — Occitan, ka — Georgian, be — Belarusian, tg — Tajik, sd — Sindhi, gu — Gujarati, am — Amharic, yi — Yiddish, lo — Lao, uz — Uzbek, fo — Faroese, ht — Haitian Creole, ps — Pashto, tk — Turkmen, nn — Nynorsk, mt — Maltese, sa — Sanskrit, lb — Luxembourgish, my — Myanmar, bo — Tibetan, tl — Tagalog, mg — Malagasy, as — Assamese, tt — Tatar, haw — Hawaiian, ln — Lingala, ha — Hausa, ba — Bashkir, jw — Javanese, su — Sundanese, yue — Cantonese.

Quick start

How to test

1
Launch Xeoma

Download and launch Xeoma. Use the Trial edition or activate Xeoma Standard/Xeoma Pro license + “Voice-to-text” additional module license.

2
Add a camera

Add a camera manually or wait while Xeoma finds cameras in your network automatically. For a separate microphone, connect the “Microphone” module and select the appropriate sound source.

3
Add the module

Add the “Voice-to-Text” module to the chain, set up a reaction to keywords or continuous speech transcription and, if needed, enable the transcript overlay on the camera image and recordings.

4
Set up reactions

Set a reaction to the event — saving to CSV, mobile notifications or any other.

Voice-to-Text module settings in Xeoma
The Trial edition has a number of limitations. To get to know Voice-to-Text better, we recommend requesting a free demo license. It gives access to all features without limitations and without any obligations. Only an email address is required. Click the button below to get started:
Get a free
demo license
Video

Voice-to-Text: Speech Recognition in Xeoma

Free trial

Try Xeoma for free

Enter your name and email to get a free demo license. No credit card required.

Get Xeoma free demo licenses to email

Full functionality. Sent instantly.

We urge you to refrain from using emails that contain personal data, and from sending us personal data in any other way. If you still do, by submitting this form, you confirm your consent to the processing of your personal data

Need the module adapted to your needs?

Do you need something else?

The module doesn’t quite fit your conditions? We can develop the functions you need and add them into Xeoma as paid development. See details.

Have questions? Need help? Please contact us! We’ll be happy to help!

Ready to start?

Test Voice-to-Text with a free demo license — no credit card and no personal data required.