October 6, 2026 · 10 min · api · agents · engineering
.mdWhich Talking Avatar APIs Include Speech Recognition, the Language Model, and Text-to-Speech?
Explore how Realtime Avatar APIs orchestrate speech recognition, language models, and text-to-speech for interactive digital human experiences.
For developers building interactive digital experiences, the question of which talking avatar APIs natively include speech recognition (STT), a language model (LLM), and text-to-speech (TTS) is central to architectural decisions. TIC Realtime Avatar provides a full-stack real-time session experience, managing the orchestration of these components to embody AI-driven conversations, with configurable external LLMs while handling the real-time video synthesis, STT, and TTS.
The Full-Stack Session: Embodiment in Conversation
The core of a real-time interactive avatar lies in its ability to listen, comprehend, respond, and speak with the fluidity of human interaction. Our platform is designed to facilitate this entire pipeline within a single, managed session. When a client connects to a Realtime Avatar session, the server dictates the operational policy, establishing the avatar's intelligence and sensory inputs. This session policy centrally manages how the avatar perceives user input, processes information, and generates its verbal and visual responses.
Speech Recognition (STT): The Listening Ear
A crucial first step in any conversational AI is accurate and low-latency speech recognition. In a Realtime Avatar session, the avatar is configured to hear the user throughout the call by default. This is controlled by the `listen` parameter in the session policy, which corresponds to `stt_mode` on the wire. By setting `stt_mode: "server"`, the platform activates server-side speech-to-text, ensuring the avatar is always receptive to user input, enabling natural barge-in behavior and responsive dialogue without explicit turn-taking commands.
session: async ({ avatarId }) => ({
listen: true, // Default, but explicit for clarity
})This full-duplex listening capability is fundamental to maintaining the illusion of live presence. If your application drives every turn and handles its own STT, you can explicitly set `listen: false` to disable our integrated speech recognition.
Language Model (LLM) Integration: The Core Intelligence
While the platform provides the visual and auditory conduit for the AI, the avatar's intelligence and conversational style are primarily defined by the language model. Our API allows developers to integrate their preferred LLM by defining the avatar's behavior contract via the `instructions` field in the session policy. This field is paramount, shaping who the avatar is, how it speaks, and its overall persona. It supports up to 16,000 characters, offering ample space to sculpt intricate personalities and conversational directives.
session: async ({ avatarId }) => ({
instructions: "You are a kind engineer who explains complex ideas with clarity. Focus on practical applications and spark curiosity.",
// ... other session parameters
})Furthermore, sessions support an `initial_context` array of up to 32 prior messages, allowing the conversation to pick up from a rich history rather than starting cold. This is essential for building applications that maintain continuity and deeper engagement, as the LLM can leverage past interactions to inform its current responses. The platform passes the generated LLM text to its integrated text-to-speech system for vocalization.
- Consult our developer documentation on integrating AI App Builders
- Read our LLM-focused documentation
- Explore the OpenAPI specification for detailed LLM parameters
Text-to-Speech (TTS) and Voice: The Avatar's Voice
Once the LLM generates a textual response, the platform's integrated text-to-speech system brings it to life with natural-sounding voices. The `voice` parameter in the session policy allows for fine-grained control over the avatar's vocal delivery for a specific call. This can include overriding the avatar's stored voice with a specific Fish voice, adjusting `speed`, `emotion`, and `language`.
session: async ({ avatarId }) => ({
voice: {
id: "fish-deep-baritone-en-us",
speed: 1.05,
emotion: "friendly",
language: "en-US"
},
// ... other session parameters
})It is important to note that while the `voice` configuration handles the auditory output, the persona and reply style are primarily governed by the `instructions` passed to the LLM. The platform carefully carries both the session instructions and voice configuration separately, ensuring the model's textual output is then rendered with the desired vocal characteristics, lips synced to the syllable for a cohesive presentation.
Crafting Presence: Avatar Creation and Behavior
Beyond the conversational stack, the visual embodiment is where a real-time avatar truly manifests presence. Our platform emphasizes creating compelling digital humans from a minimal starting point and providing robust controls for their non-verbal behavior.
Avatar Creation from a Single Portrait
A fundamental differentiator of our platform is the avatar creation process itself. Avatars are generated from a single portrait image. From this single image, the platform automatically generates a looping idle video and a multi-clip motion library. This means you provide ONE portrait image, and the system handles the complex process of making it live and interactive, without the need for pre-recorded video footage.
- See the API reference for avatar creation details
- Explore how to generate real-time AI avatars from a portrait to live interaction
The looping idle video serves as the avatar's resting state, playing continuously when not actively speaking or performing an action. This resting loop can be further refined with a `motionPrompt` to art-direct its subtle movements, ensuring it aligns with the avatar's personality.
await rta.setLoop(avatarId, {
motionPrompt: "tilts her head, a small amused smile, settles back to centre",
});Real-time Motion Library: Non-Verbal Cues
To enrich the avatar's expressiveness, the platform generates a dynamic motion library. This library consists of short video clips that can be triggered by specific events or by the LLM itself. For instance, a `userSpeechStarted` event can trigger a 'nod' clip, indicating active listening. More complex behaviors are handled by `actions`, where the LLM's `description` guides it to select and play an appropriate clip, such as a 'wave' for a greeting.
const { revision } = await rta.listClips(avatarId);
const { plan } = await rta.setClipLibrary(avatarId, {
expectedRevision: revision,
clips: {
nod: { source: { motionPrompt: "a small attentive nod, returns to rest" } },
wave: { source: { motionPrompt: "raises a hand and waves warmly, then lowers it" } }
},
on: {
userSpeechStarted: { clips: ["nod"] }
},
actions: {
greet: { description: "When greeting the user", clips: ["wave"] }
}
});The platform ensures that lip-sync remains accurate during full-duplex speech, dynamically animating the avatar's mouth over the current visual state, whether it's an idle loop or an expressive motion clip. This seamless integration of verbal and non-verbal communication is critical for a truly immersive and believable interaction.
Limitations and Considerations
While the platform offers a comprehensive solution for real-time interactive avatars, it is important to understand its boundaries. The core strength lies in its managed execution of the visual embodiment and the orchestration of STT, LLM, and TTS components into a single, cohesive real-time stream. However, developers maintain flexibility in choosing their preferred external LLM, providing the instructions and context that define the AI's intelligence.
Latency in a real-time system is a complex interplay of several factors, including the avatar's startup state, the chosen AI models, and network conditions. While the system is optimized for real-time performance, promising specific latency numbers without a cited measurement protocol is not practical due to the variability inherent in distributed systems. We focus on providing the tools and architecture to evaluate and optimize for your specific use case.
It is also crucial to reiterate that avatar creation is strictly from a single portrait image; creation from a supplied video is not supported by the platform. This ensures consistency and quality of the generated real-time assets.
Shipping Your Interactive Avatar
To bring your interactive avatar project to life, begin by creating your avatar from a high-quality portrait image through the API. Once the avatar's status is 'ready' (indicating the idle loop and initial motion library are generated), you can initiate sessions. The typed TypeScript SDK (realtime-avatar) simplifies integration, generated directly from our OpenAPI document.
Define your avatar's persona and conversational rules through the `instructions` field, and manage its non-verbal cues by refining its motion library. For cost considerations, real-time usage is priced at approximately $5 per hour, with plans including 10 hours for $24/month and overage rates between $0.07-0.095 per minute, depending on your plan. Avatar creation itself is $3 per avatar beyond any included in your plan, and the generated resting loop is always included.
Focus on iterating the avatar's conversational logic and its expressive behaviors. The underlying real-time video, STT, and TTS orchestration are handled, allowing you to concentrate on the user experience and the unique value your AI avatar brings. Test your conversation flows, fine-tune the `instructions`, and experiment with different motion clips to achieve the most engaging and lifelike interactions for your application.