Skip to content

Agents overview

An agent combines a conversation engine, voice, instructions, knowledge, call controls, and post-call analysis. Provider-specific settings remain isolated, so switching an engine does not silently copy incompatible values.

EnginePipelineBest forKey controls
Classic PipelineSpeech-to-text → LLM → text-to-speechBroad provider and voice choiceTTS provider, voice, language, speed, transcription languages
Vani UltraWhisper → Gemma → OmniVoiceVaniAgent-managed multilingual voice callsVani Ultra voice, vocabulary, language, generation steps
Ultra RealtimeGemini Live speech-to-speechLow-latency natural conversationGemini voice, realtime language, provider-specific prompt
  1. Identity — name, active state, assigned phone number, and optional avatar.
  2. Conversation — fixed or variable greeting, fallback greeting, system prompt, and engine-specific prompt.
  3. Knowledge — attach one or more published knowledge bases.
  4. Speech — voice, language, speed, transcription languages, and pronunciation vocabulary.
  5. Call behaviour — interruptions, background music, and hard call-duration limit.
  6. Analysis — default, structured, or advanced post-call extraction.
  7. Advanced flow — optional flow PDL with off, shadow, or active runtime mode.

Create the agent, configure its engine, test it in the browser or with a controlled call, attach a phone number, then activate it for inbound calls or campaigns. Deactivate an agent before major production changes if it is already receiving traffic.

  • Default produces the standard call summary and outcome.
  • Structured extracts explicitly configured fields for predictable workflows.
  • Advanced uses a custom analysis prompt and configuration.

Analysis is separate from the live conversation prompt. Changing analysis instructions does not alter what the caller hears.