Skip to content

Latest commit

 

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Event-Driven SNN Accelerator for RISC-V

A compact, multiplier-free, event-list-driven accelerator for spiking neural network inference on RISC-V. The accelerator runs a complete SNN inference autonomously: a CPU loads input currents, writes GO, and reads the classified digit after the accelerator raises DONE.

This repository is the continuation and full superset of Custom RISC-V Instructions for Spiking Neural Networks. It retains every Phase-1 artifact—training code, quantization studies, custom RISC-V instruction coprocessor, tests, simulations, and synthesis flow—then extends that foundation with the Phase-2 autonomous accelerator, RISC-V SoC, Arty Z7-20 board top, and generated bitstream.

Project lineage: tight coupling to loose coupling

Phase 1: instruction coprocessor Phase 2: event-driven accelerator
Repository role Foundation, also published separately This repository: continuation and complete project
CPU interface CORE-V-XIF custom instructions Standard memory-mapped OBI peripheral
Granularity One neuron / four synapses per instruction One complete inference per GO command
CPU role Executes loops and calls instructions Loads 512 currents, starts the accelerator, polls completion
Sparsity Add work skipped, but every group is visited Silent hidden neurons are never delivered or visited

The two implementations run the same 784→512→10, 25-step LIF SNN on the same CV32E40X. That makes the tight-versus-loose coupling comparison a measured result rather than an architectural claim.

What is built

The Phase-2 accelerator contains:

  • a time-multiplexed hidden-layer LIF neuron engine with 512 membrane states;
  • an event FIFO holding indices of fired hidden neurons;
  • BRAM-backed input-current and packed layer-2 weight memories;
  • ten parallel output accumulators and output LIF neurons;
  • a sequential argmax/vote stage;
  • a memory-mapped OBI wrapper with control, status, result, vote, and input current windows;
  • a CV32E40X-based SoC with program BRAM and address routing;
  • an Arty Z7-20 top level with 125 MHz→50 MHz MMCM clocking, reset sequencing, and LED result display.

It uses int8/Q8.8 arithmetic throughout the SNN datapath:

U_next = U - (U >>> 4) + I
spike  = U_next >= threshold
U_next = 0 when spike is asserted

The leak is implemented using a shift and subtraction, while binary spikes turn synaptic multiplication into add-or-skip logic. Both Phase 1 and Phase 2 therefore use zero DSP multipliers.

Results

Metric Phase 1 coprocessor Phase 2 accelerator
LUTs 151 1,466
BRAM 0 2.5 tiles (1 × RAMB36 + 3 × RAMB18)
DSP multipliers 0 0
Dynamic power estimate 2 mW 54 mW
Block timing 100 MHz, met 100 MHz post-route, +0.624 ns
End-to-end cycles/image 4,205,230 19,313–20,711
End-to-end speedup — ≈213×

Area, power, and timing are measured on the synthesized netlist with the trained layer-2 weights present in Block RAM. Cycle counts, speedup, and accuracy are functional results and are unchanged.

Accuracy over the full MNIST test set:

Model Accuracy
Float snnTorch baseline 95.04%
Integer, instruction-exact model 95.03%

The event-list architecture measured 3.99× fewer layer-2 synaptic operations on a 2,000-image study. The full 10,000-image firing rate was 26.0%, implying approximately 3.8× less synaptic work than a dense clock-driven traversal.

Full SoC and board build

The complete CV32E40X + accelerator system was placed and routed for the Zynq-7020:

Full-system metric Result
LUTs 5,619 (10.6% of the target FPGA)
BRAM 34.5 tiles (33 × RAMB36 + 3 × RAMB18)
Timing 50 MHz, met with +0.369 ns slack
Bitstream results/vivado_board/arty_z7_top.bit

The Arty Z7-20 board top was additionally verified in XSim: the MMCM locks, reset releases only after clock lock, the SoC runs at approximately 50 MHz, and the baked-in digit-7 demo drives LEDs to 0111.

Verification

The project uses a layered verification flow:

  1. Python LIF and instruction golden models
  2. Bit-exact quantized full-network evaluation
  3. Verilator unit, accelerator, and CPU-driven system simulations
  4. Vivado XSim four-state checks for X/reset bugs
  5. Vivado synthesis, implementation, timing, power, and bitstream generation

Important engineering issues found and resolved include BRAM inference blocked by resettable memory logic, a timing-breaking combinational argmax chain, and a read-to-clear OBI status race.

Reproducing key runs

From WSL/Linux:

# Phase 1 coprocessor tests and benchmarks
bash core/build.sh test
bash core/build.sh bench
bash core/build.sh mnist

# Phase 2 accelerator and CPU-integrated system
bash tb/run_core_tb.sh
bash core/build_system.sh

From a Vivado shell or batch invocation:

vivado -mode batch -source vivado/build_phase2.tcl
vivado -mode batch -source vivado/build_soc.tcl
vivado -mode batch -source vivado/build_bitstream.tcl

Repository layout

rtl/snn_fu/        Phase-1 CORE-V-XIF coprocessor
rtl/accel/         Phase-2 neuron engine, event FIFO, core, and OBI wrapper
rtl/soc/           Synthesizable CV32E40X + RAM + accelerator SoC
rtl/board/         Arty Z7-20 board top and LED output path
sw/                Training, golden models, firmware, and data exporters
tb/                Accelerator and memory-mapped-interface testbenches
core/              RISC-V simulation harnesses and build scripts
vivado/            Synthesis, implementation, XSim, constraints, bitstream flow
docs/              Architecture, ISA, verification, and full results
results/           Generated evaluation, waveform, and Vivado artifacts

Hardware status and limitations

The RTL, SoC, board top, and bitstream are simulation-, synthesis-, and post-route-verified. Physical FPGA implementation remains outstanding: the bitstream has not yet been flashed or validated on an Arty Z7-20 board.

Other deliberate limitations:

  • layer 1 is dense CPU preprocessing rather than spike-encoded input;
  • CPU-to-accelerator input transfer is 18% of runtime; DMA is future work;
  • the benchmark is a compact MNIST SNN, not a large event-camera workload;
  • the CPU core is third-party CV32E40X RTL; this project's contribution is the coprocessor, accelerator, integration, verification, and board flow.

See docs/RESULTS.md for the detailed measured comparison and the verification record.

About

Event-list-driven SNN accelerator for RISC-V, continuing the custom-instruction coprocessor project.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages