STT
Speech to text. Picks a model from any connected STT provider (Fish Audio, ElevenLabs, xAI) the same way the LLM picks its model. lang carries the language the provider detected.
Fired. It runs when an event or a value lands on trigger, and pulls its other inputs at that moment.
Inputs
Section titled “Inputs”| Port | Type | Notes |
|---|---|---|
trigger |
event | Fires the node |
audio |
audio |
Outputs
Section titled “Outputs”| Port | Type | Notes |
|---|---|---|
text |
text | |
lang |
lang | Optional |
trigger |
event |
| Knob | Kind | Default | Notes |
|---|---|---|---|
model |
model | none | stays a knob (cannot become an input) |
Models that ship with Boltjar
Section titled “Models that ship with Boltjar”Boltjar ships a manifest for each of these models, and the picker lists them. Picking a model adds its own settings to the node as knobs, and each of those can become an input. Models explains manifests and how to add one.
| Model | Id | Needs | Can |
|---|---|---|---|
| ElevenLabs Scribe v2 | elevenlabs/scribe_v2 |
ELEVENLABS_API_KEY |
|
| Fish Audio ASR | fish/asr |
FISH_API_KEY |
|
| xAI STT (grok-voice-transcribe 2.0) | xai/stt |
XAI_API_KEY |
Example
Section titled “Example”Transcribes a clip. Wire a data: audio value into audio, from an Audio node or an Audio Input, and fire trigger. text carries the words and lang the language the provider heard; wire lang into a TTS lang to answer in the same language with an xAI voice. Every STT model needs its provider’s key.
More in AI
Section titled “More in AI”- Embed: Turn text into an embedding vector with the embed model you pick, the way the LLM picks its model.
- LLM: A chat / multimodal model.
- Rerank: Reorder candidates by how well each one answers the query, and keep the best.
- Tool: A tool the LLM can call.
- Tool Args: Where a tool’s body starts.
- TTS: Text to speech.