OpenHuman Guide
← Back to Guides

OpenHuman Voice & Google Meet Agent — STT, TTS, and Live Meeting Integration

OpenHuman is voice-first when you want it to be. Speech-to-text, text-to-speech, and the live Google Meet agent are part of the core — not third-party plugins.

🎙️ Speech-to-Text (STT)

  • Push-to-talk hotkey: Hold a key and speak. Release to send.
  • Cross-platform mic capture: Voice-activity detection built in.
  • Streaming transcription: Words appear as you speak.
  • Hallucination filter: Strips "Thanks for watching" and other silence-induced artifacts.
  • Post-processing: Punctuation, capitalization, dictation cleanup.

Dictation can replace the active text input on your desktop or be sent straight into a chat with the agent.

🔊 Text-to-Speech (TTS)

Reply speech routes through a hosted TTS model — ElevenLabs by default. The agent's responses are spoken back in a voice you choose, with natural timing and prosody. The mascot avatar lip-syncs to the audio stream via a viseme map.

  • Voice selection is configurable per user
  • Natural speed and prosody
  • Mascot lip-syncs in real time

🎥 Live Google Meet Agent

This is OpenHuman's flagship voice integration. The mascot joins your Google Meet as a real participant — not just a listener.

Capabilities

  • Joins via embedded webview: Drop a Meet link and the agent enters the call.
  • Real-time transcription: Streams audio to STT, transcribes everyone on the call, and writes structured notes into the Memory Tree as the meeting progresses.
  • Speaks back: When asked (or when it decides it has value to add), it generates speech through TTS and plays it into the meeting as an audio stream — other participants hear it live.
  • Uses tools during calls: Queries data, looks up documents, sends messages while on the call.

How to use it

  1. Ensure the mascot is enabled (Settings → Mascot)
  2. Grant microphone + camera permissions (system prompt on first use)
  3. Share a Google Meet link in conversation: "Join this meeting"
  4. The mascot opens the call, streams audio, transcribes, and participates

💡 Example prompt

"Join my 2 PM sprint retro and take notes" → Mascot finds the calendar event, joins the Meet, transcribes the entire session, and stores key decisions in Memory Tree.

🔒 Privacy

  • Audio capture is local. Streaming STT goes through the OpenHuman backend; no recording is retained beyond the live transcript.
  • TTS audio is streamed and discarded — nothing is stored.
  • Meeting transcripts land in your local Memory Tree, just like any other source.

⚡ Model routing for voice

During meetings, the agent uses hint:fast routing for low-latency conversational turns. Complex analysis still routes to reasoning models automatically.

🛠 Setup steps

  1. Open Settings → Voice tab
  2. Enable STT (choose mic input) and TTS (choose voice)
  3. For Meet Agent: ensure camera/mic permissions are granted to OpenHuman
  4. Test: "What time is it?" spoken aloud → hear the mascot respond