./run geena-assistente-vocale-raspberry-wake-word-voce-reale
Geena hears my Mac better than me: 7 “Gina”s out of 40
Technical diary of a voice assistant on a Raspberry Pi 5: barge-in at 0.72 s, speaker recognition, a local model that got rejected, and a wake word that can't hear me.
- status
- wip
- project
- geena
- updated
- 2026-09-28
- tags
For days I piled up success clips: “Gina” spoken, beep, turn started. Three hundred recordings, statistics going up. Then I listened to the clips. Almost all of them were the synthetic voice of the Mac I use for automated tests, not mine. Every number built on that pile was false, and the truth measured on my forty real recordings is this: 7 out of 40. The same word, spoken by the Mac, reaches 0.86.
The idea
Geena lives on the desk; it’s not a smart speaker scattered around the house. A single device: a Raspberry Pi 5 with 8 GB, a two-microphone board and a WM8960 codec. I corrected the goal along the way: I started from “better than Alexa” and ended up at “better than J.A.R.V.I.S.” — someone to talk to, not a voice remote control.
The structural choice that holds everything up: the local program is not a brain. It’s ears, mouth and LEDs. The brain is an LLM agent with real tools (terminal, reminders, memory, network), on a dedicated profile. Every command goes through it. Before this there was a prototype sitting on top of a home-automation hub: thrown away, because an assistant that can only turn on the lights is not someone to talk to.
How it actually works
microfono ──► dsnoop (ALSA) ──┬─► cancellazione eco (AEC3) ─► wake word
│ │
│ └─► rilevatore "qualcuno le parla sopra"
│ ─► pausa ─► ASR conferma «Gina»
└─► registratore comando ─► ASR ─► testo
│
impronta vocale (chi parla)
▼
agente LLM + strumenti ──► esecuzione reale
│
altoparlante ◄── TTS ◄───────────────┘ + LED + pulsante
Recognizer: Parakeet TDT 0.6B v3 via whisper.cpp. Voice: Piper, medium Italian voice. Wake word: openWakeWord. Every feature has a switch in features.yaml and a rollback command tested in both directions, stacked in order: you roll back one piece at a time.
The part that works: interrupting by voice
The piece I’m proudest of is barge-in — you talk over her while she’s answering and she goes quiet. Measured: 28 successful interruptions out of 30, 0.72 s median from the end of the wake word to the beep, 0 false triggers in 19 attempts with similar words (“Tina”, “Gino”, “Regina”), 0 self-interruptions in 13 minutes of continuous speech.
The counterintuitive discovery is here: echo cancellation works too well. While she speaks, the AEC cancels the echo from her speaker — and along with it, it also cancels the voice of whoever is talking over her. The fix wasn’t a better filter, but making her breathe: about 0.45 s of pause between sentences. It’s in her silences that the recognizer hears the human.
Who is speaking
I needed her to tell her brain who is talking to her. A voice embedding model in ONNX (~26 MB), on the same onnxruntime already there for the wake word; a Kaldi-style mel front-end rewritten in numpy so as not to install anything in the agent’s environment; voiceprint = 256-dimensional centroid from 8 sentences, compared by cosine.
def session():
global _sess
if _sess is None:
opts = ort.SessionOptions()
opts.intra_op_num_threads = 2 # lascia core alla catena audio
_sess = ort.InferenceSession(str(MODEL), sess_options=opts,
providers=["CPUExecutionProvider"])
return _sess
def cmd_verify(path, name=None):
e = embed(path)
scores = {k: float(e @ np.array(d["centroid"])) for k, d in load_voices().items()}
Serious validation, with sentences held out of the voiceprint: +0.857 and +0.801 for my voice, +0.10 for Geena’s own voice, from −0.21 to +0.16 for synthetic voices. Operating threshold 0.55. Live, the synthetic voice played through speaker, room and microphone scored +0.526 in one test and +0.085 in the final one: the channel matters; the voiceprint also captures microphone and room, not just timbre.
The first validation, the one done overnight, gave 0.92 versus 0.10. It was skewed by the same wrong clips. The same mistake twice, on the same pile of data.
The product choice: aware, not a gatekeeper. Geena answers everyone, but tells her brain who is talking to her. No command is refused.
The wake word: the problem still open
| Condition | Recognized (threshold 0.4) |
|---|---|
| Normal voice | 2/10 |
| Quiet voice | 3/10 |
| Turned sideways | 0/10 |
| From a distance | 2/10 |
| Total | 7/40 (17%) |
Many scores are exactly 0.00: it’s not a threshold issue, the model simply doesn’t react to my voice. It’s trained on synthetic English speech.
What I tried, in order of decreasing optimism:
- Personal classifier on my 40 recordings: 11–45% holding out one condition at a time, with one false trigger every 1000 windows. Better than 17%, not enough.
- Italian synthetic dataset: 288 positives (system voices at three speeds plus commercial voices) and 822 confusable negatives — Dina, Gianna, cucina, pagina, macchina. Trained on that alone and measured on my voice: AUC 0.42, below chance. Synthetic voices don’t transfer to real voices, and adding more of them does nothing.
- Direct similarity comparison, no training: 0% at the same false-alarm rate.
Three paths remain: many more samples of my own, a longer wake phrase, or using transcription as the recognizer — that one does catch the word, and it’s the same pipeline already built for barge-in.
The local model: tried and rejected
I tried serving a 2B model on the Pi itself to remove network latency. Contention on the four cores, measured: command transcription went from 1.5–1.6 s to 2.4–3.3 s (+68%) while it generates; first spoken word from 0.8–0.9 s to 1.8–2.0 s (+125%). Four threads are counterproductive — 2 threads 7.89 tokens/s, 3 threads 7.09, 4 threads 6.31: generation is bound by memory bandwidth, not compute.
But the rejection came from somewhere else: zero tool calls in 6 real turns, 4 of which needed them. Instead of executing, it made things up: “it’s [current time]”, “volume at seven” without having touched anything. With reasoning turned back on it does use the tools, but drops to 0.88 tokens/s. For an assistant that has to do things, a model that hallucinates execution is the worst possible flaw. I went back to a remote model.
Honest gotchas (almost all hardware)
- The “mysterious” failures were power.
throttledregister at 0x50000,Undervoltage detected!right before every crash, 25 drops in a single boot. Lesson learned the hard way: total watts don’t matter, what matters is how much it delivers at 5 V — lots of 65 W power supplies rate that power at 9/12/20 V and stop at 3 A at 5 V. - The hardware watchdog was useless. During an overnight hang systemd kept giving it the tap while the system could no longer start a login process. Kernel alive, userspace dead: regular heartbeat, dead patient.
- The microSD was wrongly accused. Read-only filesystem three times, always under heavy writes.
f3wrote and read back 42 GB without losing a byte; the power supply held 5.10–5.16 V at full load. The test that settled it was removing one variable: without the audio board, 56 GB written at full speed, zero errors; with the board fitted, read-only after 80 seconds. With the board refitted on the updated kernel, seven 30-minute runs all passed. The culprit was almost certainly a kernel I had pinned by hand — the only thing that changed between the crashes and stability. The precise cause inside that kernel is not proven, and I have no intention of pretending otherwise. - I had pinned the kernel to fix a different failure (audio dead after an update: the board’s overlay declared the master clock on the wrong node, the frequency stayed at zero and every audio open failed with
-22). When the real cause emerged, the reason for the pin no longer existed. But the pin had stayed there doing damage. - Early mistakes, for completeness: recording always lasting 20 s because the median RMS of room noise (278) was above the silence threshold (200); commands executed in UTC, two hours off; 44.1/48 kHz audio brought down to 16 kHz by averaging 2–3 samples, which distorts exactly the voice frequencies.
How it’s going
It works, and it’s measured by voice: every feature is tested by having another computer in the room speak and recording the response with a second microphone, then comparing the real timings. Wake word, barge-in, continuous conversation (after answering she keeps listening, the next sentence doesn’t need “Gina”), reminders, timers, alarms, announcements on her own initiative, speaker recognition. Never silent and never cut off: holding phrases at 60 and 120 s, cancellation only after 300 s of silence. The only piece that doesn’t hold up is the front door itself — the wake word on my voice.
What I learned
Look at the data before the metrics. Three hundred “successful” clips that belonged to another speaker gave me two false validations in a row. Now, before trusting a number, I listen to or transcribe the sample.
Synthetic data doesn’t teach a real voice. 288 synthetic positives produced AUC 0.42 — worse than a coin toss. Forty recordings of mine are worth more than a thousand from the Mac, and the fix isn’t generating another thousand.
Remove one variable at a time. The last failure looked every bit like a dying microSD, and it wasn’t. No hypothesis solved it: physically removing one piece and trying again did.
An assistant that fakes execution is worse than a slow one. I’d rather pay for a network round-trip than a model that tells me it turned up the volume without touching it.
Failing silently is forbidden. Every new feature has a switch and a rollback tested in both directions; if it doesn’t work, we go back to the previous behavior and write it in the log. Better Geena as she was than Geena gone silent.
Updates
2026-09The second ear: from 7 to 32 “Gina”s out of 40
I closed the open wake word problem with the transcriber I was already using for commands: Parakeet had always understood my “Gina”s. On the same 40 real recordings from the post I go from 7 to 32, with zero false alarms over 95 s of real conversation.
| Condition | Wake model | Parakeet |
|---|---|---|
| Normal voice | 2/10 | 10/10 |
| Low voice | 3/10 | 9/10 |
| From afar | 2/10 | 9/10 |
| Turned sideways | 0/10 | 4/10 |
| Total | 7/40 | 32/40 (80%) |
parakeet-serve: the CLI reloaded the model on every call and took 5.3 s. I wrote a dedicated resident process that loads the model only once (1 s, 653 MB) and answers in 0.57 s on a 4 s clip and in 0.2 s on a short call. It takes paths on stdin and returns one line per request.earshot.py: a basic voice detector (energy against a continuously estimated background noise) cuts out the speech. It only sends segments under 2 s for transcription, and of a long sentence only the beginning, because “Gina, che ore sono” has the name up front. If the transcript contains a word from the family (gina, dina, gino, china…), it’s a wake-up. I added the low-voice truncations (“Gin.”, “Ging.”, “Gene.”) after checking that they never appear in 353 words of real conversation in the room. Everything sits behindasr_wakeinfeatures.yaml, with a dedicated rollback.- Two traps: the microphone has a constant offset of −968, so silence weighed 964 RMS and exceeded every threshold; I removed it with a 12 Hz high-pass filter. Then the double wake-up: both detectors heard the same word and the second one cancelled the turn the first had just opened. Now the second ear checks the daemon’s last wake-up.
- The test that can’t be done: I played the 40 recordings through the speakers of another computer in the room and got 3/40, against 32/40 on the same files read from disk. Going through speaker, room and microphone adds a second reverberation on top of the one already recorded, the consonants disappear and the transcripts become “Yeah.” or stay empty. A recording made in the room can’t be re-measured by playing it back in the room. With a synthetic voice, on the other hand, it can, because it’s clean at the source.
- Speaker threshold from 0.55 to 0.40: with a 7-sentence voiceprint, measured on the eighth one held out, I get 8/8 with 1.5 s of speech and 7/8 below about 1 s. On “Gina” alone (half a second) no score out of 40 exceeds 0.55, so recognition runs on the command that follows. At 0.40 it also catches short commands and stays at double the best impostor observed (0.19). However, the measured impostors are synthetic voices and room conversation, not a second real person.
“Turned sideways” at 4/10 remains open: the HAT has two microphones and I only use one, so merging them is the natural path. The definitive check will be saying “Gina” live, not a measurement on files. The voice is the next chapter.


