There are a great many ESP32 voice assistant projects. Most of them are a wiring diagram and a hope. This one, by jayesh_nawani and covered by oeadmin at Open Electronics, is documented at the level where you can learn something from it — specifically about where the hard parts are, which is not the AI.
The pipeline
- An I2S microphone picks up speech
- The clip is packaged as a WAV file and sent to the OpenAI Whisper API for transcription
- GPT writes a reply
- TTS audio streams back out through a MAX98357A amplifier into a 3W speaker
- A 0.96”, 128×64 OLED shows two animated eyes reacting to the state of the conversation
Recording starts when sustained volume passes a threshold and stops after a short silence, with a two-second cap.
Capture runs at 16,000 Hz; the text-to-speech audio arrives at 24,000 Hz, streamed in 512-byte chunks.
The detail that makes it work
The eye animation lives in its own FreeRTOS task on core 0 while the audio and network pipeline runs on core 1, so blocking HTTP calls do not stall it.
This is the whole lesson of the project and it generalises to almost every interactive object anyone builds.
The problem: an HTTPS request to a cloud API takes hundreds of milliseconds to several seconds, and the network stack’s calls are blocking. In a single-threaded loop(), everything stops while that request completes. Your animation freezes, your buttons stop responding, your LEDs hold their last state.
And a frozen interface does not read as waiting. It reads as broken. A user who sees nothing happening assumes the device is dead and starts pressing things — which is exactly the worst moment for them to do that.
The fix: the ESP32 has two cores and FreeRTOS underneath Arduino. Pin the animation to one core, the network to the other, and the eyes keep moving while the request is in flight. The eyes then are the progress indicator — they can look up while thinking, blink while listening, and that state is legible from across a room without text.
The general rule worth extracting: anything a human looks at or touches belongs on a different task from anything that waits on a network. This applies to installations, instruments, kiosks, wearables — anything with both an interface and a remote dependency. It is the single most common structural mistake in hobby-grade interactive hardware, and the fix is twenty lines.
The second detail: the sample rates disagree, deliberately
16kHz in, 24kHz out. Not an oversight.
16kHz is the right capture rate for speech recognition. Whisper and essentially every ASR model are trained on 16kHz audio; sending 44.1kHz means the API downsamples it anyway, and you paid in RAM and upload time for nothing. Speech intelligibility lives below 8kHz, so 16kHz sampling covers it per Nyquist.
24kHz is a reasonable playback rate for synthesised speech, where you want a bit more top end so the result does not sound like a telephone.
So the device is running two different audio rates in two different directions, which the I2S peripheral has to be reconfigured for between phases. Getting that wrong gives you speech that plays back at the wrong pitch and speed — a classic and very confusing bug.
512-byte chunks for the streaming playback is the other practical number: small enough that playback starts almost immediately rather than after a full download, large enough that you are not doing per-sample bookkeeping.
The voice activity detection is deliberately dumb, and that is correct
Start on sustained volume above a threshold, stop after a short silence, cap at two seconds.
No neural VAD, no semantic endpointing, no wake word. An amplitude threshold with hysteresis and a hard timeout.
For a device on a desk, this is the right engineering. It costs almost nothing, has no model to load, and fails in a way the user immediately understands — if it does not trigger, you speak louder; if it cuts you off, you were too slow. The two-second cap is the unglamorous hero: it guarantees the device can never get stuck recording, which is the failure mode that otherwise requires a power cycle.
It is worth contrasting this with the more sophisticated approaches we have covered — semantic VAD that distinguishes a thinking pause from an actual end of turn. That is better, and it needs a model and the compute to run it. On a device where every byte of RAM is accounted for, a threshold is the answer.
What this is and isn’t
It is not a local AI device. The intelligence is entirely in the cloud: Whisper transcribes, GPT answers, TTS speaks. The ESP32 is a well-engineered microphone, speaker, screen and state machine with a network connection. Which means it needs Wi-Fi, costs money per interaction, and sends everything said near it to a third party.
That is a real limitation and worth stating plainly, particularly alongside the direction we wrote about yesterday — generative models running on bare-metal microcontrollers with no network at all. The two projects sit at opposite ends of the same spectrum.
But as a template for the physical layer of a voice-driven object, this is genuinely instructive. The two eyes on an OLED are doing more interaction design work than most people would credit — they give the device a state, a direction of attention, and a personality, from a $3 display and a few hundred bytes of animation data.