Skip to content
Sponsor

Voice

The graph answers your messages out loud: the model writes a short reply, Strip takes out what should not be read aloud, and TTS turns it into a clip that plays in a Preview.

Download the graph and save it in user/graphs/, or build it below. The graph is MIT-0: copy it and license what you build from it however you like.

A voice costs money, so the downloaded graph picks no model and ships its TTS and Speaker disabled: it turns On as a text chat, and step 4 gives it its voice.

In Settings › AI Providers, under AI Providers, paste your xAI key. A TTS node without its key fails with an error that names the missing key.

Node Settings
Chat Input, named Message none
Template, named Prompt the text below
LLM model: pick Gemma 4 e4b (Ollama), or another model you pulled
Strip, named Speakable every toggle on
TTS model: pick xAI TTS (Grok voices), voice eve
Preview, named Speaker autoplay on
Chat none
You are a friendly voice assistant. Answer in one or two short sentences, with no lists and no markdown.
User: {message}
From To
Message trigger Prompt trigger
Message text Prompt message
Prompt trigger LLM trigger
Prompt out LLM prompt
LLM response Speakable in
LLM trigger TTS trigger
Speakable out TTS text
TTS trigger Speaker trigger
TTS audio Speaker in

Plus the four Chat wires, so the text shows too: Message into user and user_trigger, the LLM into reply and reply_trigger.

The {message} tag is named after the node its wire comes from. One message fires the Prompt, then the LLM, then the TTS, each once, so every reply is spoken once.

  1. Pick a model on the LLM node.
  2. Pick xAI TTS (Grok voices) on the TTS node, then right-click it and pick Enable.
  3. Right-click the Speaker and pick Enable.

Enable the TTS before the Speaker: a Speaker whose TTS is disabled gets nothing to play, and the graph stays Off until the TTS is back.

Press On and send a message. The reply shows in the Chat node, and a moment later the Speaker plays it. Try other voices from the TTS node’s voice knob; with the language on auto, xAI picks up the language of the text.

STT turns a clip into text: pick a speech-to-text model on it, wire a data: audio value into its audio, fire its trigger, and its text goes where the Message node’s text went. Its lang output can go into the TTS lang input, so an xAI voice answers in the language you spoke.

The editor has no microphone button yet, so a clip comes from one of two places:

  • A file. An Audio node reads a clip you put in user/data/files, for example question.wav. A Manual trigger can fire the STT.
  • Another program. An Audio Input node takes clips posted to it while the graph is On, as JSON with the install token:
Terminal window
curl -X POST http://127.0.0.1:8770/audio/<graph>/<Audio Input node> \
-H "Authorization: Bearer $(cat user/data/token)" \
-H "Content-Type: application/json" \
-d '{"audio": "data:audio/wav;base64,UklGR...", "lang": ""}'
  • Fish Audio or ElevenLabs. Pick their TTS model on the node and add the key. Fish Audio needs a voice reference id, ElevenLabs a voice id; both detect the language themselves.
  • Faster first words. Split a long reply with Sentences and speak it through a For-each, one clip per sentence.