Voice
LUNA’s voice is sovereign-first. Speech is transcribed and synthesised on Runink’s own servers. The phone’s or browser’s own speech engines are only a fallback, used when the server path declines or the platform can’t do it.
Talking to Luna
Tap the mic button and speak. The app records raw 16 kHz mono audio. It
detects the start of speech, ends the turn after about 1.8 seconds of silence,
and stops after 30 seconds at most. The audio goes to
VoiceService.Transcribe, which runs Voxtral speech-to-text on the
platform’s own inference plane. The transcript is sent to Luna as an ordinary
chat turn marked source: "voice", and it gets the same server-side
prompt-injection hygiene as typed text.
If nothing was said, Luna replies “I didn’t catch any words in that — try again a little closer.”
Hearing Luna
The speaker toggle next to the composer controls whether replies are
spoken (“Luna speaks aloud — tap to silence” / “Replies are silent — tap to
give Luna a voice”). The setting is remembered. When it is on, each reply is
turned into plain text (markdown is not read aloud) and sent to
VoiceService.SynthesizeSpeech. That forwards to the platform’s shared
ttsd service (Piper, running in-cluster) and returns WAV audio.
Where each path runs
| Speech-to-text | Text-to-speech | |
|---|---|---|
| Web build (debug-only) | Server (Voxtral) | Server (ttsd), played in the browser |
| Android / iOS builds | Server (Voxtral) | On-device (flutter_tts). Native builds have no WAV player yet, so server speech can’t be played there. |
When the app falls back
| Situation | What happens |
|---|---|
| The server can’t transcribe (voice plane down or not configured) | Luna says “My own ears are unreachable — using the browser’s for this one. Say it again?” and uses the device or browser recogniser for that turn. |
| The platform can’t capture raw audio (microphone refused, encoder unsupported) | The device recogniser is used. |
| Your session has expired | The app ends the session and asks you to sign in again. It does not quietly switch your voice to the device recogniser. |
| Server speech fails or is disabled | The reply is spoken by the on-device engine. |
Limits
- Voice calls share an hourly budget,
LUNA_VOICE_CALLS_PER_HOUR, which defaults to 120 and covers both transcription and speech. A reply that was already synthesised is served from a cache and doesn’t count. Over the budget, the server answers “voice transcription is rate-limited on the shared plane — try again in ~N min” (or “voice is rate-limited…” for speech). Text chat keeps working. - One transcription accepts up to 16 MiB of audio.
Transcribestreams with a heartbeat because transcription on the shared CPU plane can take minutes.SynthesizeSpeechis a single request, because it takes seconds.- Piper has one voice. The
voicefield on the request is only a cache key and doesn’t select a voice.
For operator settings (VOXTRAL_REMOTE_URL, TTSD_URL), see
Configuration.