Skip to content

Voice & Wakeword

Talk to your companion out loud and hear it talk back. Saying a wake word starts the conversation, you speak your message, and the AI replies in the companion’s own voice. Everything happens on your own machine: your voice is never streamed to a cloud service.

How It Works

Say your wake word (the one your admin set up), then speak normally. The app listens, turns your speech into text, sends it to your companion, and reads the reply aloud.

Wake Word

The wake word is a short phrase that activates the microphone. Out of the box it is “Hey MaiPai,” and you can see which wake word is active in your profile settings.

Each companion can have its own wake word, so saying a different name can bring up a different companion. Admins can pick from a set of ready-made wake words or train a brand-new one (even a made-up name) right inside the app.

Talking Back and Forth

  • Keep going without repeating the wake word. After your companion finishes a reply, it keeps listening for a few seconds. If you just keep talking, the conversation continues. If you go quiet, it returns to waiting for the wake word.
  • Interrupt anytime. Start talking while your companion is speaking and it stops to listen to you (this is called barge-in).
  • Say “stop” (or “cancel,” “quiet,” “never mind”) to end the turn right away.

Voice Settings

In your profile you can:

  • Choose the companion voice
  • Adjust how fast the voice speaks
  • Turn wake word detection on or off

On-Device and Private

Wake word detection runs right in your browser. Speech-to-text and the spoken reply run in a small local voice service on the server. Nothing about your voice leaves your home network.

Tips

  • Speak clearly at a normal pace. The speech-to-text handles accents and everyday speech well.
  • You don’t need to raise your voice; normal conversational volume works.
  • If the wake word isn’t responding, check that your browser has microphone permission.
  • The voice files download once on first setup (about 300 MB), so the first run takes a moment before voice is ready.