Structural Verilog implementation of an IEEE 754 single-precision floating-point (FP32) datapath for Multi-Layer Perceptron (MLP) neural network inference. This repository decomposes a sigmoid neuron into modular floating-point arithmetic IP units, culminating in a 4-element streaming vector dot-product engine cascaded into an activation pipeline.
- Objective: Integrate basic floating-point arithmetic IP components to process initial scaling and bias subtraction.
- Components: Instantiates
FP_multiplierandFP_subberfromFPU_Is_lib.v. - Functionality: Multiplies an input dot-product scalar by IEEE 754
-1.0(32'hbf800000) and passes the output toFP_subberto subtract the bias valuebh_i.
- Objective: Expand Project 1 into a full floating-point sequential transformation pipeline.
- Components: Cascades
fp_ms, a secondFP_subbermodule, andFP_reciprocal. - Functionality: Subtracts a constant
1.0(32'h3f800000) from the incoming pipeline signal, then computes the IEEE 754 floating-point reciprocal (1.0 / io_a).
- Objective: Implement a 4-element streaming vector dot-product multiplier.
- Components: 4 parallel
FP_multiplierunits and a 2-stage binary adder tree usingFP_addermodules. - Functionality: Evaluates vector dot product for a streaming width of 4.
- Objective: Integrate Project 3 (
fp_dot) and Project 2 (fp_mssr) into a standalone Sigmoid Neuron hardware module. - Functionality: Connects the scalar IEEE 754 output of
fp_dotinto the input stage offp_mssralong with bias parameterbh_ito output final signalyh_i.
IEEE 754 Standard (Single Precision):
- Sign: Bit 31
- Exponent: Bits 30:23
- Mantissa: Bits 22:0
| FP | 1.0 | 2.0 | 3.0 | 4.0 | 5.0 | 6.0 |
|---|---|---|---|---|---|---|
| HW | 3f800000 | 40000000 | 40400000 | 40800000 | 40a00000 | 40c00000 |
| FP | 7.0 | 8.0 | 9.0 | 10.0 | 11.0 | 12.0 |
| HW | 40e00000 | 41000000 | 41100000 | 41200000 | 41300000 | 41400000 |
| FP | -1.0 | -2.0 | -3.0 | -4.0 | -5.0 | -6.0 |
| HW | bf800000 | c0000000 | c0400000 | c0800000 | c0a00000 | c0c00000 |
| FP | -7.0 | -8.0 | -9.0 | -10.0 | -11.0 | -12.0 |
| HW | c0e00000 | c1000000 | c1100000 | c1200000 | c1300000 | c1400000 |
All primitives operate on active-low asynchronous resets (negedge reset). Testbenches verify hardware execution against the following dataset:
- Test Case 1:
- Vector Inputs:
w = [1.0, 2.0, 3.0, 1.0],x = [1.0, 2.0, 2.0, 1.0],bh_i = -2.0 fp_dotresult:12.0(32'h41400000)fp_msresult:12.0 * (-1.0) - (-2.0) = -10.0(32'hc1200000)- Final
yh_iresult:1.0 / (-10.0 - 1.0) ≈ -0.0909
- Vector Inputs:
- Test Case 2:
- Vector Inputs:
w = [1.0, 2.0, 3.0, 1.0],x = [-1.0, -2.0, -1.0, -2.0],bh_i = 2.0 fp_dotresult:-10.0(32'hc1200000)fp_msresult:-10.0 * (-1.0) - 2.0 = 8.0(32'h41000000)- Final
yh_iresult:1.0 / (8.0 - 1.0) ≈ 0.1428
- Vector Inputs:
The FPU library is provided by Xiaokun (Bobbie) Yang through the IC-Design repository: https://github.com/IC-Design-Lab/IC-Design
dut/: Design-under-test Verilog filestb/: Verilog testbenchfilelist/: ModelSim source file listsim/: ModelSim simulation scriptthird_party/IC-Design/: FPU library included as a Git submodule
Clone the repository with its submodule:
git clone --recurse-submodules https://github.com/IC-Design-Lab/IC-Design
git submodule update --init --recursiveSimulation Open ModelSim in the sim directory and run:
do run- Hardware architectural specifications and floating-point design concepts based on coursework material from Integrated Circuit Design: IC Design Flow and Project-Based Learning book by Dr. Xiaokun Yang (University of Houston-Clear Lake).