Skip to content

Voice

LUNA’s voice is sovereign-first. Speech is transcribed and synthesised on Runink’s own servers. The phone’s or browser’s own speech engines are only a fallback, used when the server path declines or the platform can’t do it.

Talking to Luna

Tap the mic button and speak. The app records raw 16 kHz mono audio. It detects the start of speech, ends the turn after about 1.8 seconds of silence, and stops after 30 seconds at most. The audio goes to VoiceService.Transcribe, which runs Voxtral speech-to-text on the platform’s own inference plane. The transcript is sent to Luna as an ordinary chat turn marked source: "voice", and it gets the same server-side prompt-injection hygiene as typed text.

If nothing was said, Luna replies “I didn’t catch any words in that — try again a little closer.”

Hearing Luna

The speaker toggle next to the composer controls whether replies are spoken (“Luna speaks aloud — tap to silence” / “Replies are silent — tap to give Luna a voice”). The setting is remembered. When it is on, each reply is turned into plain text (markdown is not read aloud) and sent to VoiceService.SynthesizeSpeech. That forwards to the platform’s shared ttsd service (Piper, running in-cluster) and returns WAV audio.

Where each path runs

Speech-to-textText-to-speech
Web build (debug-only)Server (Voxtral)Server (ttsd), played in the browser
Android / iOS buildsServer (Voxtral)On-device (flutter_tts). Native builds have no WAV player yet, so server speech can’t be played there.

When the app falls back

SituationWhat happens
The server can’t transcribe (voice plane down or not configured)Luna says “My own ears are unreachable — using the browser’s for this one. Say it again?” and uses the device or browser recogniser for that turn.
The platform can’t capture raw audio (microphone refused, encoder unsupported)The device recogniser is used.
Your session has expiredThe app ends the session and asks you to sign in again. It does not quietly switch your voice to the device recogniser.
Server speech fails or is disabledThe reply is spoken by the on-device engine.
The fallbacks are third parties. On most Android phones, the device speech recogniser is Google’s, and the device text-to-speech engine receives the text of Luna’s replies. No model provider receives your voice through Luna’s backend, but the fallback paths do hand voice or text to the platform’s engines. Because native builds have no WAV player, every spoken reply on a phone currently goes through the device TTS engine.

Limits

  • Voice calls share an hourly budget, LUNA_VOICE_CALLS_PER_HOUR, which defaults to 120 and covers both transcription and speech. A reply that was already synthesised is served from a cache and doesn’t count. Over the budget, the server answers “voice transcription is rate-limited on the shared plane — try again in ~N min” (or “voice is rate-limited…” for speech). Text chat keeps working.
  • One transcription accepts up to 16 MiB of audio.
  • Transcribe streams with a heartbeat because transcription on the shared CPU plane can take minutes. SynthesizeSpeech is a single request, because it takes seconds.
  • Piper has one voice. The voice field on the request is only a cache key and doesn’t select a voice.

For operator settings (VOXTRAL_REMOTE_URL, TTSD_URL), see Configuration.