Audio Services
Audio Services let a deployed agent generate spoken audio, create sound effects, and discover available voices. Agent steps call the typed ctx.sapiom.speech capability; Sapiom supplies the authenticated cloud connection when the deployed agent runs.
Choose the audio operation
Section titled “Choose the audio operation”| What the agent needs | Recommended operation |
|---|---|
| Spoken narration, an alert, or a voiceover | speech.textToSpeech.create(...) |
| A non-verbal sound described in words | speech.soundEffects.create(...) |
| The voices available at runtime | speech.voices.list() |
| Audio that must survive its hosted URL | Pass storage and retain the returned fileId |
Do not copy a fixed voice inventory into the agent. List voices when selecting or validating a voice, then store the chosen voiceId as configuration. Speech-to-text (STT) is separate from these operations.
Generate speech
Section titled “Generate speech”const audio = await ctx.sapiom.speech.textToSpeech.create({ text: input.announcement, voice: input.voiceId, storage: { visibility: "private" },});
if (audio.storageError) { throw new Error(`Audio persistence failed: ${audio.storageError}`);}
if (!audio.fileId && !audio.url) { throw new Error("Speech generation returned no audio reference");}
return { fileId: audio.fileId ?? null, temporaryUrl: audio.url ?? null, expiresAt: audio.expiresAt ?? null,};text must contain non-whitespace text and must not exceed 5,000 UTF-16 units (text.length in JavaScript). This input limit differs from the Unicode code-point count used for billing.
Pass a voiceId returned by speech.voices.list() as voice; voice names are not resolved. With @sapiom/tools 0.44.0 or later, omitting voice lets Core select its default. Pass an explicit ID when the agent needs a particular voice.
Generate a sound effect
Section titled “Generate a sound effect”const effect = await ctx.sapiom.speech.soundEffects.create({ text: "A short wooden door knock in a quiet room", durationSeconds: 2, storage: { visibility: "private" },});
return { fileId: effect.fileId ?? null, temporaryUrl: effect.url ?? null, storageError: effect.storageError ?? null,};Use durationSeconds only when the effect needs a particular length. A supplied duration must be from 0.5 to 30 seconds. An omitted or null duration means automatic generation unless params.duration_seconds supplies a value. A non-null durationSeconds takes precedence over that provider option.
Text-to-speech and sound effects have separate inputs. Both return SpeechResult, with optional url, expiresAt, fileId, and storageError fields.
Discover voices
Section titled “Discover voices”const { voices } = await ctx.sapiom.speech.voices.list();
return voices.map(({ voiceId, name }) => ({ voiceId, name: name ?? null,}));Treat voiceId as the stable selection value. A voice name is optional, and the available set can change. Voice listing requires a Sapiom identity and has no usage charge.
HTTP routes
Section titled “HTTP routes”Deployed agents should call the typed client, but the same operations are available as authenticated POST routes:
| Typed method | HTTP route |
|---|---|
speech.textToSpeech.create(...) | /v1/capabilities/speech.tts |
speech.soundEffects.create(...) | /v1/capabilities/speech.sound-effects |
speech.voices.list() | /v1/capabilities/speech.voices.list |
Test the behavior locally
Section titled “Test the behavior locally”Local Run replaces ctx.sapiom.speech calls with deterministic URLs and does not generate or store audio. Override the exact path when a step requires persistence or branches on an error:
{ "version": 1, "steps": { "make-announcement": { "speech.textToSpeech.create": { "url": "https://fixtures.example/announcement.mp3", "expiresAt": "2099-01-01T00:00:00.000Z", "fileId": "file-audio-local" } } }}Sound-effect and voice-list overrides use speech.soundEffects.create and speech.voices.list. Supply every field the step consumes, assert the terminal output, and require both unusedStubs and stubWarnings to be empty. A passing Local Run does not prove live voice availability, pronunciation, audio quality, storage, or runtime usage.
Manage usage and retention
Section titled “Manage usage and retention”Sapiom calculates generation usage from the submitted request:
| Operation | Billing quantity |
|---|---|
| Text-to-speech | Submitted Unicode code points, including spaces and combining marks, without Unicode normalization. Standard and fast models have different rate classes. |
| Sound effect with an explicit duration | Requested seconds, including fractional seconds. |
| Sound effect with automatic duration | One generation. |
Returned audio duration does not determine these charges. Use the signed-in capability catalog for current prices, models, and limits. A failed generation can reverse usage without a guaranteed financial refund. A retry can cause another charge.
Temporary URLs expire after one hour; check expiresAt before use.
Successful storage can return only fileId; retain it as the durable reference and mint a fresh URL later with ctx.sapiom.fileStorage.getDownloadUrl(fileId). A storageError means the audio generated but did not persist; the result can still carry a temporary url, which is not durable. Use public visibility only for assets intentionally available without tenant authentication.
© 2026 Sapiom, Inc.