← All projectsLet’s talk about it ↗
MOMO_AI
AI & systems / Sep 2025
MOMO_AI
A collaborative multimodal assistant connecting wake-word detection, speech recognition, language models, and voice output.
My contribution
Contributed AI/ML components, data-processing experiments, debugging, and architectural refinement in a collaborative fork.
The approach
- Detect a wake word, capture speech, and transcribe it with faster-whisper.
- Route requests through Gemini and developer-command modules.
- Return spoken responses using ElevenLabs, with desktop UI and session handling.
Scope & perspective
Voice interaction ties together audio timing, transcription, command handling, and feedback. The expressive 3D avatar remains an extension rather than a completed integrated feature.
Collaborative project; this fork documents Arnav's contributions and experiments.
Explore the repository ↗