Sarvam Conv AI SDK

The Sarvam Conversational AI SDK is a Python package that helps developers build and extend conversational agents. It provides core components to manage conversation flow, language preferences, and messaging, making it easier to develop interactive and context-aware AI experiences.

Overview

The Sarvam Conv AI SDK enables developers to create tools that can:
  • Facilitate agentic capabilities like API calling in the middle of a conversation.
  • Manage agent-specific variables
  • Control and modify the language used during conversations
  • Send dynamic messages to both the user and the underlying language model (LLM)

Installation

Basic Installation

Install the SDK via pip:

Audio Support (Optional)

If you want to use audio streaming features (microphone input and speaker output), you need to install PyAudio. This requires system-level dependencies:

Option 1: Install with audio support

Note: You’ll need to install PortAudio first:

Option 2: Use without PyAudio

The SDK works without PyAudio for non-playback environments; audio capture/playback features will not be available. You can still:
  • Use the WebSocket client for real-time voice conversations (provide your own audio I/O)
  • Build backend proxies for frontend applications

AsyncSamvaadAgent

Build real-time voice with a small set of inputs.
  • You provide InteractionConfig: who the user is, which app to talk to, interaction type, and audio sample rate; optionally include overrides like agent_variables and initial language/state.
  • You create AsyncSamvaadAgent with your API key, config, and optional audio interface plus callbacks for text/audio/events.
  • Start the agent: it fetches a signed WebSocket URL, sends interaction_start, and streams audio/text both ways.

Key features

  • Real-time voice interaction — natural speak and listen
  • Automatic audio management — built-in microphone input and speaker output
  • Async/await support — non-blocking operations
  • Callback handling — process text/audio/events asynchronously
  • Connection management — robust WebSocket handling
Minimal example:

AsyncSamvaadAgent parameters

Methods:
  • await agent.start() — start and connect
  • await agent.stop() — stop and cleanup
  • await agent.wait_for_connect(timeout: float | None = 5.0) — wait until connected
  • await agent.wait_for_disconnect() — wait until disconnected or stopped
  • agent.is_connected() — connection status
  • await agent.send_audio(audio_bytes: bytes) — send raw 16‑bit PCM audio
  • agent.get_interaction_id() — current interaction id or None
Audio interface (optional): AsyncDefaultAudioInterface(input_sample_rate: int = 16000)
  • Methods: start(input_callback), output(audio: bytes, sample_rate?: int), interrupt(), stop()
  • Audio: LINEAR16 (16‑bit PCM mono). Supported sample rates: 16000.

What you must provide: InteractionConfig

Required fields:
  • user_identifier_type: One of CUSTOM, EMAIL, PHONE_NUMBER, UNKNOWN
  • user_identifier: The identifier value (string; phone/email/custom id) # This id can be used to see logs in the log analyser
  • org_id: Your organization, e.g., “sarvamai”
  • workspace_id: Your workspace, e.g., “default”
  • app_id: The target application id
  • interaction_type: InteractionType.CALL (voice)
  • sample_rate: 16000 (16-bit PCM mono)
  • version: int (Optional)
Important
If version is not provided, the SDK uses the latest committed version of the app.
The connection will fail if the provided app_id has no committed version.
Optional overrides (applied server-side at start):
  • agent_variables: dict of key/value to seed the agent context
  • initial_language_name: e.g., “English”, “Hindi” (must be allowed by app)
  • initial_state_name: starting state name, if your app uses states
  • initial_bot_message: first message from the agent
Example config:

Quick start: local voice test

  1. Install dependencies
  1. Set credentials (or pass directly in code)
  1. Run the example
The example uses AsyncDefaultAudioInterface to capture mic at 16kHz and play responses. You can override base_url in AsyncSamvaadAgent if you use a different environment.

Headless mode (no PyAudio)

Use your own audio I/O. Create the agent without audio_interface and push raw 16‑bit PCM mono chunks that match config.sample_rate.

Connect your frontend (backend proxy pattern)

See the section above for AsyncSamvaadAgent usage. For a full backend bridge, follow the same pattern in your server. Message shapes:
  • Frontend → backend (init):
  • Frontend → backend (text):
  • Frontend → backend (audio):
Bridge essentials on the backend:
  • Build InteractionConfig from init context; create AsyncSamvaadAgent with callbacks.
  • Decode base64 and forward audio via await agent.send_audio(audio_bytes).
  • In text/audio/event callbacks, websocket.send_json back to the frontend.
Minimal sketch:

Requirements for Async Audio

  1. PyAudio installation:
  2. System dependencies:
    • macOS: brew install portaudio
    • Ubuntu/Debian: sudo apt-get install portaudio19-dev
    • Windows: download from http://www.portaudio.com/download.html
  3. Environment variables (optional convenience):

Complete Example

See sarvam_conv_ai_sdk/examples/async_audio_example.py for a full, runnable script with mic capture, callbacks, and clean shutdown.

Text-Based Conversations

In addition to voice interactions, the SDK supports text-based conversations for chat applications, messaging platforms, and other text-only use cases.

Key Features

  • Real-time text conversation — send and receive text messages asynchronously
  • Voice note support — send audio recordings with automatic transcription
  • Same callback pattern — consistent API with audio mode
  • Event handling — track conversation state and transitions
  • Async/await support — non-blocking text I/O

Basic Text Example

Text vs Audio Configuration

The main differences between text and audio modes:

Text-Specific Methods

  • await agent.send_text(text: str) — Send a text message to the agent
    • Accepts plain string messages
    • Non-blocking, returns immediately
    • Messages are queued and sent over WebSocket
  • await agent.send_voice_note(audio_data: bytes, transcribe: bool = False) — Send a voice note in text conversations
    • Accepts raw PCM audio bytes (16-bit PCM mono at the sample_rate of 48000)
    • When transcribe=True, the server transcribes the audio to text before processing
    • The transcribed text is returned via the text_callback
    • Non-blocking, returns immediately
    • The default sample rate for text interaction type is 48000
    • Useful for adding voice input to text-based conversations

Interactive Text Loop

For continuous chat experiences, use an input loop:

Voice Notes in Text Conversations

Text mode supports voice notes — users can send audio recordings that are transcribed and processed by the agent.

Key Features

  • Audio recording — Capture voice input using a microphone
  • Automatic transcription — Server-side speech-to-text conversion
  • Seamless integration — Voice notes work alongside regular text messages
  • Same conversation flow — Transcribed text is processed just like typed text

Voice Note Flow

When a voice note is sent with transcribe=True:
  1. User records audio at 48000 Hz
  2. Audio is sent to the server via send_voice_note()
  3. Server transcribes the audio to text
  4. Transcribed text is returned via the text_callback
  5. Server processes the transcription and generates agent response
  6. Agent response is delivered via text_callback

Example: Voice Note in Text Conversation

Recording Audio for Voice Notes

To capture audio from the microphone, you can use PyAudio:

Audio Format Requirements

  • Format: LINEAR16 (16-bit PCM mono)
  • Sample Rate: 48000 Hz (must match for voice notes)
  • Channels: Mono (single channel)
  • Encoding: Raw PCM bytes (no headers)

Interactive Voice Note Example

Combine text and voice input in an interactive loop:

Voice Note Dependencies

Voice note recording requires PyAudio:
System dependencies: Note: PyAudio is only needed for recording audio. If you have audio from another source (e.g., web upload, mobile app), you can send it directly via send_voice_note() without PyAudio.

Text Message Types

The text_callback receives ServerTextMsgType which can be:
  • ServerTextChunkMsg — Streaming text chunks (status: pending/completed/failed)
  • ServerTextMsg — Complete text messages
Both contain:
  • text: str — The text content
  • type: ServerMsgType — Message type identifier

Quick Start: Text Chat Test

  1. Install SDK
For text-only (no voice notes):
For text + voice notes:
  1. Set credentials
  1. Run the text example

Use Cases for Text Mode

  • Chat applications — Web chat widgets, mobile messaging
  • Messaging platforms — WhatsApp, Telegram, Slack bots
  • Backend proxies — Bridge between your frontend and Sarvam AI
  • Headless environments — Servers without audio hardware
  • Testing & development — Faster iteration without audio setup
  • Multi-modal apps — Support both voice and text channels
  • Voice messaging — Text conversations with voice note transcription
  • Accessibility — Enable users to choose between typing and speaking

Custom Tools

Example Usage


Base Classes

The SDK exposes three base classes for tool development:

1. SarvamTool

Primary base class for all operational tools invoked during conversation flow. Example:

2. SarvamOnStartTool

Executed at the beginning of a conversation, typically for initialization. The class must be named OnStart.

3. SarvamOnEndTool

Executed at the end of a conversation, typically for cleanup or post-processing. The class must be named OnEnd.

Context Classes and Methods

SarvamToolContext

The context object passed to SarvamTool.run() methods.

Variable Management

  • get_agent_variable(variable_name: str) -> Any Retrieve the value of a variable.
  • set_agent_variable(variable_name: str, value: Any) -> None Update a variable’s value.

Language Control

  • get_current_language() -> SarvamToolLanguageName Returns the current language of the agent.
  • change_language(language: SarvamToolLanguageName) -> None Update the language preference.

Conversation Flow

  • set_end_conversation() -> None Explicitly end the conversation.

State Management

  • get_current_state() -> str Returns the current state of the conversation.
  • change_state(state: str) -> None Transition to a new state. Note: The new state must be one of the next valid states defined in the agent configuration.

Engagement Metadata

  • get_engagement_metadata() -> EngagementMetadata Retrieve the engagement metadata containing information about the current interaction.

SarvamOnStartToolContext

The context object passed to SarvamOnStartTool.run() methods.

Variable Management

  • get_agent_variable(variable_name: str) -> Any Retrieve the value of a variable.
  • set_agent_variable(variable_name: str, value: Any) -> None Update a variable’s value.

User Information

  • get_user_identifier() -> str Get the user identifier.

Telephony Information

  • provider_ref_id: Optional[str] The reference ID from the channel provider. For telephony providers, this would contain the Call SID (Session ID) which uniquely identifies a specific phone call. For other channel providers, this would contain their respective reference IDs. Defaults to None for channels that don’t provide reference IDs.

Initialization Methods

  • set_initial_bot_message(message: str) -> None Set the first message sent by the agent when the conversation starts.
  • set_initial_state_name(state_name: str) -> None Set the initial state from which the agent should start.
  • set_initial_language_name(language: SarvamToolLanguageName) -> None Define the initial language preference for the user.

Engagement Metadata

  • get_engagement_metadata() -> EngagementMetadata Retrieve the engagement metadata containing information about the current interaction.

SarvamOnEndToolContext

The context object passed to SarvamOnEndTool.run() methods.

Variable Management

  • get_agent_variable(variable_name: str) -> Any Retrieve the value of a variable.
  • set_agent_variable(variable_name: str, value: Any) -> None Update a variable’s value.

User Information

  • get_user_identifier() -> str Get the user identifier.

Telephony Information

  • provider_ref_id: Optional[str] The reference ID from the channel provider. For telephony providers, this would contain the Call SID (Session ID) which uniquely identifies a specific phone call. For other channel providers, this would contain their respective reference IDs. Defaults to None for channels that don’t provide reference IDs.

Engagement Metadata

  • get_engagement_metadata() -> EngagementMetadata Retrieve the engagement metadata containing information about the current interaction.

Interaction Reattempt

  • set_retry_interaction The user will be reattempted with the same agent. Useful when any business goal has not been met.

Interaction Transcript

  • get_interaction_transcript() -> SarvamInteractionTranscript Retrieve the conversation history containing user and agent messages in English and the timestamp when the conversation began and ended. Format: yyyy-mm-dd hh:mm:ss
Example transcript:

Return Types

SarvamToolOutput

The return type for SarvamTool.run() methods. Contains:
  • message_to_user: Optional[str] - Message that is sent directly to the user
  • message_to_llm: Optional[str] - Message that is sent to the LLM, which then responds
  • context: SarvamToolContext - The updated context object
Note: At least one of message_to_llm or message_to_user must be set. Important: When both message_to_user and message_to_llm are set, only the message_to_user is actually sent to the user, but the message_to_llm overrides the message_to_user when adding to the chat thread for the LLM’s context.

EngagementMetadata

The engagement metadata object that can be retrieved from context objects using get_engagement_metadata(). Contains:
  • interaction_id: str - Unique identifier for each conversation between user & agent.
  • attempt_id: Optional[str] - Unique identifier for each attempt created on the platform
  • campaign_id: Optional[str] - Campaign ID for the interaction
  • interaction_language: SarvamToolLanguageName - The language used for the interaction (defaults to English)
  • app_id: str - Application identifier of the agent for the interaction
  • app_version: int - Version number of the agent
  • agent_phone_number: Optional[str] - Phone number associated with the conversational agent application

Supported Languages

The SDK supports multilingual conversations using the SarvamToolLanguageName enum. Available languages include:
  • Bengali
  • Gujarati
  • Kannada
  • Malayalam
  • Tamil
  • Telugu
  • Punjabi
  • Odia
  • Marathi
  • Hindi
  • English
Note: The allowed languages are actually a subset that is preselected while defining the agent configurations.

Best Practices

  1. Always implement run(): The run() method is the entry point for tool execution logic.
  2. Use Field() for parameters: Ensures type safety and adds descriptive metadata necessary for LLM to use in the prompt.
  3. Gracefully handle errors: Avoid accessing unset variables or using invalid types.
  4. Return the appropriate type: SarvamTool.run() must return SarvamToolOutput, while SarvamOnStartTool.run() and SarvamOnEndTool.run() return their respective context objects.
  5. Write meaningful docstrings: Clearly describe what each tool is intended to do as this directly impacts the performance of tool calling capabilities of the agent.
  6. Use async operations for I/O: For the best performance, use async/await for external API calls to avoid blocking.
  7. Use context methods: Use the provided context methods for variable management, language control, and messaging instead of directly accessing context attributes.

Testing Your Tools

After creating a tool, you can test it locally to ensure it works as expected. Here’s how to test your tools:

Testing Steps

  1. Create the ToolContext: Initialize the appropriate context object with test data
  2. Instantiate the tool class: Use tool.model_validate(tool_args) to create a tool instance
  3. Run the tool: Call the tool’s run() method with the context
  4. Observe the returned object: Check if the necessary changes have been made to the context

Example Test: SarvamTool

Example Test: OnStart Tool

For SarvamOnStartTool, the testing approach is similar but it returns the context object directly:

Example Test: OnEnd Tool

Requirements for Async Audio

  1. PyAudio Installation:
  2. System Dependencies:
  3. Environment Variables:

Best Practices for Async Audio

  1. Use proper event loop setup for PyAudio compatibility:
  2. Handle connection states gracefully:
  3. Implement proper cleanup in finally blocks:
  4. Use appropriate sample rates (typically 16000 Hz for input)
  5. Handle interruptions with KeyboardInterrupt:

Complete Example

See sarvam_conv_ai_sdk/examples/async_audio_example.py for a complete working script.