Live speech-to-speech translation for calls

Hear every call in your language.

They speak English on Google Meet. You hear Hindi — in a natural voice, under a second behind, while they are still talking. Nothing to install on their side.

EN
HI
fyv phrase end → voice 0.74 s
0.7 smedian delay from the end of a phrase to translated speech, measured in the browser against real vendors
250 msbetween word-level recognition updates — translation starts mid-sentence
0installs, accounts or consent dialogs for the person speaking

How it works

It never waits for the sentence to end.

Human interpreters start speaking before the speaker finishes. So does fyv: every stage streams into the next, and each of them is on while the one before it is still working.

phrase end+0.25 s+0.5 s+0.75 s
  1. speech“…does that work for you?”
  2. recognisewords committed as they stabilise
  3. translatefirst token 140 ms after the final
  4. voicefirst byte ~220 ms
  5. you hearक्या यह आपके लिए ठीक है?
One real segment from the benchmark. Time runs left to right; the original speaker is quietened, not muted, while the Hindi plays.
  1. 01

    Listen

    Call audio is captured in the browser. On-device voice-activity detection marks phrase ends in milliseconds, before anything leaves your machine.

  2. 02

    Recognise

    Turn-aware streaming recognition commits words the moment they stop changing — not when the speaker pauses.

  3. 03

    Translate

    A low-latency model translates with the conversation as context, so names, numbers and questions survive the trip.

  4. 04

    Speak

    Streaming speech starts voicing the first words of the translation before the last words exist.

  5. 05

    Stay in step

    If translation falls behind, the voice speaks a touch faster and playback stretches without changing pitch. You stay within a sentence of the speaker.

Languages

Any pair. Spoken, not subtitled.

Engineering

Latency is measured, not promised.

Every translated phrase carries timestamps from the end of speech to your speaker. A replay benchmark runs on every change and fails the build if the median moves. The whole pipeline is open source.

Read the source on GitHub
32-second continuous English monologue, real vendors, phrase end → translated speech
Pipelinep50p95
fyv today0.86 s1.6 s
fyv, first release1.10 s2.2 s
fyv, v0 (gpt-4o-mini)1.6 s—

Roadmap

Where this goes.

  1. Now

    Chrome extension for Google Meet

    You hear them in your language. Works on any tab with audio.

  2. Next

    Speak back

    They hear you in theirs — your voice, their words.

  3. Then

    Desktop app

    A virtual microphone and speaker, so fyv works in Zoom, Teams, Discord — everywhere.

Stop losing people to language.

Get the Chrome extension

Free during the beta with an access code. Open source under MIT — or run your own relay.