Intelligence and Voice: the three components of an AI4CALL assistant
The Intelligence and Voice section of an AI4CALL assistant configures three components. The model processes the request and decides the answer. Speech recognition turns the caller's voice into text. Text-to-speech delivers the answer as audio.
Each component has its own tab, but what the caller actually experiences depends on how the three work together. In this video tutorial we look at where the settings live, what each parameter does and how to check the result before saving.
LLM model: how the assistant decides the answer
Open Configure in the LLM Model tab. The model follows the prompt and decides when to call the available tools: it therefore determines both the content of the answer and the moment the assistant uses a tool.
AI4CALL provider or your own API keys
With a provider supplied by ai for call you use the services available on the platform: you select the model from the catalogue and check the price list for the related consumption.
The your own API keys mode connects your own account with the provider instead. You choose the provider and the wallet configuration and, if it does not exist yet, you create it from this screen. In this mode you are the one who keeps the credentials valid and monitors the provider's credit and limits. Consider response times too: on a phone call, every wait is audible.
Advanced settings: Temperature, Max Tokens and generation
In the advanced settings you can leave the fields empty and use the model's default values. Temperature controls how much the answers vary, Max Tokens limits the amount of text generated. The other parameters cover token selection, repetition and stopping the generation. What is actually available depends on the model and provider you choose.
- To assess a change, adjust one parameter at a time and compare similar conversations.
- A limit set too low can cut off a useful answer.
- Advanced parameters are no substitute for a clear prompt.
- For short answers, state the style and length you want in the instructions as well.
- Leaving the fields empty means using the model's default values.
Speech recognition: transcribing what the caller says
Speech recognition transcribes what is said during the call. Here you choose the engine and the caller's language. Check the result using the names, numbers and addresses that are typical of your service.
The recognition language is a separate choice from the voice the assistant will speak with: two distinct settings, each to be checked on its own.
Text-to-speech: the voice the caller hears
In the TTS tab you choose the provider, the voice and the gender: these are the settings for the voice the assistant speaks with during calls. The name shown in this example belongs to the assistant's configuration. Choose the voice based on language, clarity and naturalness.
The voice menu includes audio previews. Compare the candidates using sentences that represent your service, and pay particular attention to proper names, English terms and numbers. For a full check, then try the voice in the context of a call.
Background audio: when to use it and how to check it
The Background tab lets you choose ambient or musical audio: you can browse the available backgrounds and their previews. No selection means no background. If you do use one, make sure the voice always stays understandable.
The final check before saving the assistant
Before you finish, review the model, recognition and text-to-speech together: the model has to answer correctly, recognition has to understand the caller, the voice has to sound clear. Confirm the windows you changed and save the assistant.
A test call shows how the three components work together: it is the most direct way to see whether the configuration holds up in real conditions.
The three components at a glance
<b>Model:</b> processes the request, decides the answer and when to call the available tools. <b>Speech recognition:</b> transcribes what the caller says into text, according to the engine and language you selected. <b>Text-to-speech:</b> delivers the answer as audio, with the provider, voice and gender you selected. Confirm the windows you changed, save the assistant and check the result with a test call.
Frequently asked questions
What is the difference between speech recognition and text-to-speech?
Speech recognition turns the caller's voice into text, text-to-speech delivers the answer as audio. They are two distinct settings: the recognition language is a separate choice from the voice the assistant will speak with.
Should I use the AI4CALL provider or my own API keys?
With a provider supplied by ai for call you use the services available on the platform: you select the model from the catalogue and check the price list for consumption. With your own API keys you connect your own account with the provider, and you are the one who keeps the credentials valid, monitors credit and limits, and evaluates response times.
What happens if I leave the advanced settings empty?
Leaving the fields empty means using the model's default values. That is the recommended choice until you have a specific reason to intervene: advanced parameters are no substitute for a clear prompt.
Let's configure your voice assistant together
Model, recognition and voice: we show you how to set them up for your use case and how to check them with a test call.