A compact, multiplier-free, event-list-driven accelerator for spiking neural
network inference on RISC-V. The accelerator runs a complete SNN inference
autonomously: a CPU loads input currents, writes GO, and reads the classified
digit after the accelerator raises DONE.
This repository is the continuation and full superset of Custom RISC-V Instructions for Spiking Neural Networks. It retains every Phase-1 artifact—training code, quantization studies, custom RISC-V instruction coprocessor, tests, simulations, and synthesis flow—then extends that foundation with the Phase-2 autonomous accelerator, RISC-V SoC, Arty Z7-20 board top, and generated bitstream.
| Phase 1: instruction coprocessor | Phase 2: event-driven accelerator | |
|---|---|---|
| Repository role | Foundation, also published separately | This repository: continuation and complete project |
| CPU interface | CORE-V-XIF custom instructions | Standard memory-mapped OBI peripheral |
| Granularity | One neuron / four synapses per instruction | One complete inference per GO command |
| CPU role | Executes loops and calls instructions | Loads 512 currents, starts the accelerator, polls completion |
| Sparsity | Add work skipped, but every group is visited | Silent hidden neurons are never delivered or visited |
The two implementations run the same 784→512→10, 25-step LIF SNN on the same CV32E40X. That makes the tight-versus-loose coupling comparison a measured result rather than an architectural claim.
The Phase-2 accelerator contains:
- a time-multiplexed hidden-layer LIF neuron engine with 512 membrane states;
- an event FIFO holding indices of fired hidden neurons;
- BRAM-backed input-current and packed layer-2 weight memories;
- ten parallel output accumulators and output LIF neurons;
- a sequential argmax/vote stage;
- a memory-mapped OBI wrapper with control, status, result, vote, and input current windows;
- a CV32E40X-based SoC with program BRAM and address routing;
- an Arty Z7-20 top level with 125 MHz→50 MHz MMCM clocking, reset sequencing, and LED result display.
It uses int8/Q8.8 arithmetic throughout the SNN datapath:
U_next = U - (U >>> 4) + I
spike = U_next >= threshold
U_next = 0 when spike is asserted
The leak is implemented using a shift and subtraction, while binary spikes turn synaptic multiplication into add-or-skip logic. Both Phase 1 and Phase 2 therefore use zero DSP multipliers.
| Metric | Phase 1 coprocessor | Phase 2 accelerator |
|---|---|---|
| LUTs | 151 | 1,466 |
| BRAM | 0 | 2.5 tiles (1 × RAMB36 + 3 × RAMB18) |
| DSP multipliers | 0 | 0 |
| Dynamic power estimate | 2 mW | 54 mW |
| Block timing | 100 MHz, met | 100 MHz post-route, +0.624 ns |
| End-to-end cycles/image | 4,205,230 | 19,313–20,711 |
| End-to-end speedup | — | ≈213× |
Area, power, and timing are measured on the synthesized netlist with the trained layer-2 weights present in Block RAM. Cycle counts, speedup, and accuracy are functional results and are unchanged.
Accuracy over the full MNIST test set:
| Model | Accuracy |
|---|---|
| Float snnTorch baseline | 95.04% |
| Integer, instruction-exact model | 95.03% |
The event-list architecture measured 3.99× fewer layer-2 synaptic operations on a 2,000-image study. The full 10,000-image firing rate was 26.0%, implying approximately 3.8× less synaptic work than a dense clock-driven traversal.
The complete CV32E40X + accelerator system was placed and routed for the Zynq-7020:
| Full-system metric | Result |
|---|---|
| LUTs | 5,619 (10.6% of the target FPGA) |
| BRAM | 34.5 tiles (33 × RAMB36 + 3 × RAMB18) |
| Timing | 50 MHz, met with +0.369 ns slack |
| Bitstream | results/vivado_board/arty_z7_top.bit |
The Arty Z7-20 board top was additionally verified in XSim: the MMCM locks,
reset releases only after clock lock, the SoC runs at approximately 50 MHz,
and the baked-in digit-7 demo drives LEDs to 0111.
The project uses a layered verification flow:
- Python LIF and instruction golden models
- Bit-exact quantized full-network evaluation
- Verilator unit, accelerator, and CPU-driven system simulations
- Vivado XSim four-state checks for X/reset bugs
- Vivado synthesis, implementation, timing, power, and bitstream generation
Important engineering issues found and resolved include BRAM inference blocked by resettable memory logic, a timing-breaking combinational argmax chain, and a read-to-clear OBI status race.
From WSL/Linux:
# Phase 1 coprocessor tests and benchmarks
bash core/build.sh test
bash core/build.sh bench
bash core/build.sh mnist
# Phase 2 accelerator and CPU-integrated system
bash tb/run_core_tb.sh
bash core/build_system.shFrom a Vivado shell or batch invocation:
vivado -mode batch -source vivado/build_phase2.tcl
vivado -mode batch -source vivado/build_soc.tcl
vivado -mode batch -source vivado/build_bitstream.tclrtl/snn_fu/ Phase-1 CORE-V-XIF coprocessor
rtl/accel/ Phase-2 neuron engine, event FIFO, core, and OBI wrapper
rtl/soc/ Synthesizable CV32E40X + RAM + accelerator SoC
rtl/board/ Arty Z7-20 board top and LED output path
sw/ Training, golden models, firmware, and data exporters
tb/ Accelerator and memory-mapped-interface testbenches
core/ RISC-V simulation harnesses and build scripts
vivado/ Synthesis, implementation, XSim, constraints, bitstream flow
docs/ Architecture, ISA, verification, and full results
results/ Generated evaluation, waveform, and Vivado artifacts
The RTL, SoC, board top, and bitstream are simulation-, synthesis-, and post-route-verified. Physical FPGA implementation remains outstanding: the bitstream has not yet been flashed or validated on an Arty Z7-20 board.
Other deliberate limitations:
- layer 1 is dense CPU preprocessing rather than spike-encoded input;
- CPU-to-accelerator input transfer is 18% of runtime; DMA is future work;
- the benchmark is a compact MNIST SNN, not a large event-camera workload;
- the CPU core is third-party CV32E40X RTL; this project's contribution is the coprocessor, accelerator, integration, verification, and board flow.
See docs/RESULTS.md for the detailed measured comparison and the verification record.