Live music-genre classification on a SensorTile (STEVAL-STLCS01V1, STM32L476JG Cortex-M4F at 80 MHz). Sound comes in through the board's MEMS microphone; features and a FastGRNN recurrent network run on the chip; a genre comes out every 992 ms. Nothing leaves the chip.
The model is trained on a laptop (FMA-small, 8,000 tracks, 8 genres). The board only runs inference.
EdgeEar is the clean rebuild of the Spektrer project: same measured pipeline, rewritten to be easy to read, held to MISRA C:2012, and documented for learning.
flowchart LR
MIC["MP34DT01-M mic<br/>1-bit PDM, 2.048 MHz"] --> DFSDM["DFSDM1<br/>SINC4 ÷128 → 16 kHz"]
DFSDM --> DMA["DMA ping-pong<br/>2 × 256 samples"]
DMA -->|"every 16 ms"| FE["frontend.c<br/>256 samples → 32 log-mel"]
FE --> RNN["fastgrnn_q15.c<br/>one timestep"]
RNN -->|"every 61 steps"| GATE{"gate.c<br/>real sound?"}
GATE -->|yes| OUT["LED blinks + g_status"]
Speed of one FastGRNN step. Each build changes one thing from the one
before it (make IMPL=n, measured by firmware/tools/bench_all.py):
| build | cycles per step | vs baseline |
|---|---|---|
0 · baseline: arm_mat_mult_f32 ×2 + libm |
235,862 | — |
| 1 · + one fused matrix-vector product | 194,346 | −17.6 % |
| 2 · + lookup-table sigmoid/tanh | 173,974 | −26.2 % |
| 3 · + q15 fixed point with SMLALD SIMD (shipping build) | 96,304 | −59.2 % |
| 4 · + weights copied to SRAM (memory experiment) | 65,501 | −72.2 % |
Real-time deadline. One DMA half is 256 samples = 16 ms = 1,280,000 cycles. Live, with the microphone running:
| cycles | share of 16 ms | |
|---|---|---|
| frontend (window, FFT, mel, log, normalise) | 62,073 | 4.85 % |
| FastGRNN step | 96,800 | 7.56 % |
| everything for one hop | 161,305 | 12.60 % |
mic_overruns = 0, task_overruns = 0.
Memory. Flash 152,680 B of 1 MiB. SRAM1 36,688 B of 96 KiB (no heap: all memory is static). SRAM2 31 KiB, used only by the host test-vector buffer.
Board matches Python. python/parity.py pushes a clip into the board over
SWD and compares against numpy: features agree to 1.65e-6 (relative), logits
to 4.7e-4, and the predicted class matches.
Accuracy (FMA-small test split, 8 classes, chance 12.5 %): 35.65 % for one 992 ms window, 47.38 % voting over a 30 s clip. The q15 build agrees with the float model on 99.55 % of decisions.
# build and flash the shipping firmware
cd firmware
make # needs arm-none-eabi-gcc
make flash # needs st-flash (stlink); board on the ST-LINK
python3 tools/ee_host.py status
# checks
make -C ../tests # unit tests on the laptop
make misra # MISRA C:2012 check (cppcheck)
python3 ../python/parity.py --impl 3 # board vs numpy, on hardware
python3 tools/bench_all.py # the speed table above (flashes 5 builds)Play music near the board. The LED blinks (class + 1) times per decision
(Electronic = 1 … Rock = 8); in a quiet room it gives one short blink a second.
ee_host.py history prints the last 128 decisions.
Start at docs/00-start-here.md. The docs go from hardware to audio to features to the network to timing, in plain language with diagrams:
| 01 Hardware | the board, pins, memory map, which datasheet says what |
| 02 Build and flash | toolchain, make targets, safe flashing, restoring |
| 03 Audio path | PDM, DFSDM, DMA ping-pong, the 16 ms deadline |
| 04 Features | the streaming log-mel frontend |
| 05 FastGRNN | RNN → GRU → FastGRNN, q15 fixed point, SIMD |
| 06 Optimisation ladder | what each build changed and what it bought |
| 07 Activity gate | "is anything playing?", and why spectral flux failed |
| 08 RTOS and timing | tasks, interrupts, static memory, cycle counting |
| 09 Host tools | talking to the board over SWD, parity, benchmarks |
| 10 Training | dataset, training, exporting the model to C |
| MISRA | the rules, how they are checked, the 10 documented deviations |
firmware/
inc/ headers: edgeear.h is the pipeline contract
src/ EdgeEar's C code, one job per file (see docs/00)
libc/ the few C-library headers this bare toolchain lacks
third_party/ ST HAL, CMSIS, CMSIS-DSP, FreeRTOS, unmodified
tools/ ee_host.py, bench_all.py, air_vs_file.py, live_demo.py
misra/ accepted MISRA findings, each tied to docs/MISRA.md
Makefile make / make flash / make misra / make all-impls
python/ frontend (single source of truth), model, training, export
models/ the trained model (fastgrnn.npz) and its metrics
tests/ laptop unit tests: maths, lookup tables, activity gate
docs/ the learning docs and the raw benchmark JSON
data/ cache/ the 16 GB dataset and 1 GB feature cache (links, not in git)
backups/ flash images read off the board (not in git)
The board's flash was read out before EdgeEar first wrote to it:
backups/board-before-edgeear-20261005.bin (full 1 MiB). The factory image is
kept as backups/factory-allmems1.bin, and make restore writes it back.
Nothing here ever touches the option bytes, the one realistic way to make an
STM32L4 unrecoverable.
MIT, see LICENSE. The vendor code in firmware/third_party/ keeps
its own licences (BSD-3-Clause, Apache-2.0, MIT).
- Spektrer: the original prototype EdgeEar was rebuilt from. Not public.
- TinyMLDelta: an over-the-air model-patching engine. Porting it so EdgeEar's weights can be updated in flash is the planned next step and will live in its own fork; it is not part of this repository yet.