Skip to content
Go To Dashboard

Audio Services

Audio Services let a deployed agent generate spoken audio, create sound effects, and discover available voices. Agent steps call the typed ctx.sapiom.speech capability; Sapiom supplies the authenticated cloud connection when the deployed agent runs.

What the agent needsRecommended operation
Spoken narration, an alert, or a voiceoverspeech.textToSpeech.create(...)
A non-verbal sound described in wordsspeech.soundEffects.create(...)
The voices available at runtimespeech.voices.list()
Audio that must survive its hosted URLPass storage and retain the returned fileId

Do not copy a fixed voice inventory into the agent. List voices when selecting or validating a voice, then store the chosen voiceId as configuration. Speech-to-text (STT) is separate from these operations.

const audio = await ctx.sapiom.speech.textToSpeech.create({
text: input.announcement,
voice: input.voiceId,
storage: { visibility: "private" },
});
if (audio.storageError) {
throw new Error(`Audio persistence failed: ${audio.storageError}`);
}
if (!audio.fileId && !audio.url) {
throw new Error("Speech generation returned no audio reference");
}
return {
fileId: audio.fileId ?? null,
temporaryUrl: audio.url ?? null,
expiresAt: audio.expiresAt ?? null,
};

text must contain non-whitespace text and must not exceed 5,000 UTF-16 units (text.length in JavaScript). This input limit differs from the Unicode code-point count used for billing.

Pass a voiceId returned by speech.voices.list() as voice; voice names are not resolved. With @sapiom/tools 0.44.0 or later, omitting voice lets Core select its default. Pass an explicit ID when the agent needs a particular voice.

const effect = await ctx.sapiom.speech.soundEffects.create({
text: "A short wooden door knock in a quiet room",
durationSeconds: 2,
storage: { visibility: "private" },
});
return {
fileId: effect.fileId ?? null,
temporaryUrl: effect.url ?? null,
storageError: effect.storageError ?? null,
};

Use durationSeconds only when the effect needs a particular length. A supplied duration must be from 0.5 to 30 seconds. An omitted or null duration means automatic generation unless params.duration_seconds supplies a value. A non-null durationSeconds takes precedence over that provider option.

Text-to-speech and sound effects have separate inputs. Both return SpeechResult, with optional url, expiresAt, fileId, and storageError fields.

const { voices } = await ctx.sapiom.speech.voices.list();
return voices.map(({ voiceId, name }) => ({
voiceId,
name: name ?? null,
}));

Treat voiceId as the stable selection value. A voice name is optional, and the available set can change. Voice listing requires a Sapiom identity and has no usage charge.

Deployed agents should call the typed client, but the same operations are available as authenticated POST routes:

Typed methodHTTP route
speech.textToSpeech.create(...)/v1/capabilities/speech.tts
speech.soundEffects.create(...)/v1/capabilities/speech.sound-effects
speech.voices.list()/v1/capabilities/speech.voices.list

Local Run replaces ctx.sapiom.speech calls with deterministic URLs and does not generate or store audio. Override the exact path when a step requires persistence or branches on an error:

{
"version": 1,
"steps": {
"make-announcement": {
"speech.textToSpeech.create": {
"url": "https://fixtures.example/announcement.mp3",
"expiresAt": "2099-01-01T00:00:00.000Z",
"fileId": "file-audio-local"
}
}
}
}

Sound-effect and voice-list overrides use speech.soundEffects.create and speech.voices.list. Supply every field the step consumes, assert the terminal output, and require both unusedStubs and stubWarnings to be empty. A passing Local Run does not prove live voice availability, pronunciation, audio quality, storage, or runtime usage.

Sapiom calculates generation usage from the submitted request:

OperationBilling quantity
Text-to-speechSubmitted Unicode code points, including spaces and combining marks, without Unicode normalization. Standard and fast models have different rate classes.
Sound effect with an explicit durationRequested seconds, including fractional seconds.
Sound effect with automatic durationOne generation.

Returned audio duration does not determine these charges. Use the signed-in capability catalog for current prices, models, and limits. A failed generation can reverse usage without a guaranteed financial refund. A retry can cause another charge.

Temporary URLs expire after one hour; check expiresAt before use.

Successful storage can return only fileId; retain it as the durable reference and mint a fresh URL later with ctx.sapiom.fileStorage.getDownloadUrl(fileId). A storageError means the audio generated but did not persist; the result can still carry a temporary url, which is not durable. Use public visibility only for assets intentionally available without tenant authentication.