Skip to content
hdparmarPublic

About

Live music-genre classification on an STM32L476 SensorTile

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

EdgeEar

Live music-genre classification on a SensorTile (STEVAL-STLCS01V1, STM32L476JG Cortex-M4F at 80 MHz). Sound comes in through the board's MEMS microphone; features and a FastGRNN recurrent network run on the chip; a genre comes out every 992 ms. Nothing leaves the chip.

The model is trained on a laptop (FMA-small, 8,000 tracks, 8 genres). The board only runs inference.

EdgeEar is the clean rebuild of the Spektrer project: same measured pipeline, rewritten to be easy to read, held to MISRA C:2012, and documented for learning.

flowchart LR
  MIC["MP34DT01-M mic<br/>1-bit PDM, 2.048 MHz"] --> DFSDM["DFSDM1<br/>SINC4 ÷128 → 16 kHz"]
  DFSDM --> DMA["DMA ping-pong<br/>2 × 256 samples"]
  DMA -->|"every 16 ms"| FE["frontend.c<br/>256 samples → 32 log-mel"]
  FE --> RNN["fastgrnn_q15.c<br/>one timestep"]
  RNN -->|"every 61 steps"| GATE{"gate.c<br/>real sound?"}
  GATE -->|yes| OUT["LED blinks + g_status"]
Loading

Results, measured on the board (2026-10-05)

Speed of one FastGRNN step. Each build changes one thing from the one before it (make IMPL=n, measured by firmware/tools/bench_all.py):

build cycles per step vs baseline
0 · baseline: arm_mat_mult_f32 ×2 + libm 235,862 —
1 · + one fused matrix-vector product 194,346 −17.6 %
2 · + lookup-table sigmoid/tanh 173,974 −26.2 %
3 · + q15 fixed point with SMLALD SIMD (shipping build) 96,304 −59.2 %
4 · + weights copied to SRAM (memory experiment) 65,501 −72.2 %

Real-time deadline. One DMA half is 256 samples = 16 ms = 1,280,000 cycles. Live, with the microphone running:

cycles share of 16 ms
frontend (window, FFT, mel, log, normalise) 62,073 4.85 %
FastGRNN step 96,800 7.56 %
everything for one hop 161,305 12.60 %

mic_overruns = 0, task_overruns = 0.

Memory. Flash 152,680 B of 1 MiB. SRAM1 36,688 B of 96 KiB (no heap: all memory is static). SRAM2 31 KiB, used only by the host test-vector buffer.

Board matches Python. python/parity.py pushes a clip into the board over SWD and compares against numpy: features agree to 1.65e-6 (relative), logits to 4.7e-4, and the predicted class matches.

Accuracy (FMA-small test split, 8 classes, chance 12.5 %): 35.65 % for one 992 ms window, 47.38 % voting over a 30 s clip. The q15 build agrees with the float model on 99.55 % of decisions.

Quick start

# build and flash the shipping firmware
cd firmware
make                       # needs arm-none-eabi-gcc
make flash                 # needs st-flash (stlink); board on the ST-LINK
python3 tools/ee_host.py status

# checks
make -C ../tests           # unit tests on the laptop
make misra                 # MISRA C:2012 check (cppcheck)
python3 ../python/parity.py --impl 3   # board vs numpy, on hardware
python3 tools/bench_all.py             # the speed table above (flashes 5 builds)

Play music near the board. The LED blinks (class + 1) times per decision (Electronic = 1 … Rock = 8); in a quiet room it gives one short blink a second. ee_host.py history prints the last 128 decisions.

Learn it

Start at docs/00-start-here.md. The docs go from hardware to audio to features to the network to timing, in plain language with diagrams:

01 Hardware the board, pins, memory map, which datasheet says what
02 Build and flash toolchain, make targets, safe flashing, restoring
03 Audio path PDM, DFSDM, DMA ping-pong, the 16 ms deadline
04 Features the streaming log-mel frontend
05 FastGRNN RNN → GRU → FastGRNN, q15 fixed point, SIMD
06 Optimisation ladder what each build changed and what it bought
07 Activity gate "is anything playing?", and why spectral flux failed
08 RTOS and timing tasks, interrupts, static memory, cycle counting
09 Host tools talking to the board over SWD, parity, benchmarks
10 Training dataset, training, exporting the model to C
MISRA the rules, how they are checked, the 10 documented deviations

Layout

firmware/
  inc/            headers: edgeear.h is the pipeline contract
  src/            EdgeEar's C code, one job per file (see docs/00)
  libc/           the few C-library headers this bare toolchain lacks
  third_party/    ST HAL, CMSIS, CMSIS-DSP, FreeRTOS, unmodified
  tools/          ee_host.py, bench_all.py, air_vs_file.py, live_demo.py
  misra/          accepted MISRA findings, each tied to docs/MISRA.md
  Makefile        make / make flash / make misra / make all-impls
python/           frontend (single source of truth), model, training, export
models/           the trained model (fastgrnn.npz) and its metrics
tests/            laptop unit tests: maths, lookup tables, activity gate
docs/             the learning docs and the raw benchmark JSON
data/ cache/      the 16 GB dataset and 1 GB feature cache (links, not in git)
backups/          flash images read off the board (not in git)

Safety

The board's flash was read out before EdgeEar first wrote to it: backups/board-before-edgeear-20261005.bin (full 1 MiB). The factory image is kept as backups/factory-allmems1.bin, and make restore writes it back. Nothing here ever touches the option bytes, the one realistic way to make an STM32L4 unrecoverable.

Licence

MIT, see LICENSE. The vendor code in firmware/third_party/ keeps its own licences (BSD-3-Clause, Apache-2.0, MIT).

Related

  • Spektrer: the original prototype EdgeEar was rebuilt from. Not public.
  • TinyMLDelta: an over-the-air model-patching engine. Porting it so EdgeEar's weights can be updated in flash is the planned next step and will live in its own fork; it is not part of this repository yet.

About

Live music-genre classification on an STM32L476 SensorTile

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages