A portable, allocation-free Bluetooth Low Energy host stack in C99.
microBLE implements the parts of Bluetooth LE that are pure software — the advertising data codec, L2CAP fragmentation, and the ATT/GATT protocol — as transport-agnostic state machines with no dynamic allocation, no I/O, and no dependency on a hosted C library. It runs on a laptop, in CI, under a fuzzer, and on a Cortex-M, from the same sources.
Status: working end to end. Two processes form a link over a socket and run the full GATT flow — MTU negotiation, service and characteristic discovery, reads, subscription, notifications — with no radio and no hardware. Both halves are in the library: a GATT server and a GATT client. 227 tests, 89% line coverage, clean under ASan and UBSan on gcc and clang.
Built for a Cortex-M33 — the core in Silicon Labs' EFR32xG24 family among
others — at -Os -ffreestanding:
| module | .text | .data | .bss |
|---|---|---|---|
| ATT codec | 3,926 | 0 | 0 |
| GATT server | 2,868 | 0 | 0 |
| GATT client | 1,698 | 0 | 0 |
| GAP advertising | 1,322 | 0 | 0 |
| L2CAP | 966 | 0 | 0 |
| UUIDs | 655 | 0 | 0 |
| Buffer cursors | 434 | 0 | 0 |
| Attribute database | 272 | 0 | 0 |
| Shared types | 271 | 0 | 0 |
| HCI ACL header | 140 | 0 | 0 |
| total | 12,552 | 0 | 0 |
12.3 kB of flash for both roles, and not one byte of static RAM. A
peripheral that never acts as a client drops the 1.7 kB client automatically:
the linker discards what nothing references, which is what -ffunction-sections
is for. The zero is the whole
point of the no-allocation design, stated as a measurement rather than a
claim: the library holds no state of its own, so every buffer and every
connection struct belongs to the caller and per-connection RAM is exactly what
the application declares. make size-report regenerates this, and CI runs it
so a footprint regression shows up in the diff.
make demoTwo processes, a UNIX socket standing in for the radio, and a real GATT exchange:
[2] Primary service discovery
-> Read By Group Type Request (7 octets)
<- Read By Group Type Response (14 octets)
service 180D handles 0x0001-0x0006
service 180A handles 0x0007-0x0009
[3] Characteristic discovery in 0x0001-0x0006
char 2A37 value 0x0003 props 0x10 notify
char 2A38 value 0x0006 props 0x02 read
[6] Subscribing via CCCD at 0x0004
-> Write Request (5 octets)
<- Write Response (1 octets)
[7] Subscribed. Waiting for notifications ...
heart rate: 78 bpm
heart rate: 61 bpm
The controller's ACL buffer is set to 27 octets, as a real one would be, so anything above the default MTU genuinely fragments and reassembles: the run above carries 11 L2CAP PDUs in 12 ACL packets. CI runs this same script under ASan and UBSan on every push, which makes it an integration test as well as a demo — unit tests cannot catch a layer wired up wrongly, and this can.
Bluetooth stacks are a good place to practise the things that make embedded software hard: parsing hostile input in a fixed amount of RAM, keeping protocol state machines honest, and proving correctness without the hardware in front of you. The design goal here is not feature completeness. It is that every layer should be testable on a host, byte-exact against the specification, and immune to malformed input by construction rather than by vigilance.
There is no malloc in the library and no allocator to configure. All state
lives in caller-provided structs, and all limits are compile-time constants. A
device with 64 kB of RAM cannot afford heap fragmentation on a connection path
that must not fail after three weeks of uptime, and a stack that allocates is a
stack whose worst-case memory use nobody actually knows.
Protocol layers are pure functions of bytes and state: bytes and events go in,
events and bytes come out. Nothing in src/ calls read, write, printf, or
a timer. This is what makes the test suite possible — every layer can be driven
directly from a unit test with no mocking framework, no fake HAL, and no
hardware — and it is also what lets the same code sit behind a UNIX socket in a
demo and behind a real radio on target.
Every octet that arrives from the air is untrusted, unauthenticated, and
attacker-controlled before any connection exists. So all wire access goes
through one bounds-checked cursor type with a sticky error flag
(mble_buf.h). A read past the end does not read: it
sets the flag and returns zero, and every later operation becomes a no-op.
That means a parser can be written as a straight run of reads and checked for failure exactly once, at the end, and still be safe against every possible truncation:
mble_rbuf_t rb;
mble_rbuf_init(&rb, pdu, pdu_len);
req->opcode = mble_rbuf_u8(&rb);
req->handle = mble_rbuf_u16le(&rb);
req->offset = mble_rbuf_u16le(&rb);
if (mble_rbuf_failed(&rb)) return MBLE_ERR_TRUNCATED; // one check, totalThe cost is one predictable branch per field. The benefit is that "did I remember to bounds-check this one?" stops being a question anyone has to ask in review, and there is exactly one function in the codebase where the bounds logic could be wrong.
BLE is little-endian on air. Every codec converts explicitly rather than casting a struct over a buffer, so there are no packing assumptions, no alignment traps on a Cortex-M0, and no host-endianness dependency. CI proves the last part by running the full suite on a big-endian s390x machine.
include/mble/ public headers, one per layer
src/core/ buffer cursors, shared types, UUIDs
src/gap/ advertising and scan response data [done]
src/l2cap/ fragmentation and reassembly [done]
src/att/ ATT protocol codec [done]
src/gatt/ attribute database, GATT server, client [done]
src/hci/ HCI ACL packet header [done]
src/port/ virtual radio over a UNIX socket [done]
tests/ unit tests, one binary per layer
fuzz/ libFuzzer harnesses
apps/ demo peripheral and central [done]
tools/ demo runner
Everything except src/port/ is freestanding and does no I/O. That is not a
convention, it is a build target: the port layer compiles into a separate
library, and make size-report links only the core. The claim stays true
because breaking it breaks the build.
The interesting inputs in a Bluetooth stack are the malformed ones, so that is where the tests concentrate. The advertising suite covers length fields that overrun the payload, zero-length structures, terminators followed by padding, odd-length UUID lists, and names that are not NUL-terminated on air — all of which real advertisers emit, and each of which is a plausible way to read past the end of a packet buffer.
L2CAP adds a second kind of hostile input, because the peer controls both the declared length of a PDU and how many octets it actually sends, and the two need not agree. The suite enumerates every way they can disagree: a declared length larger than the reassembly buffer, a continuation that overruns it, a start fragment that interrupts a pending PDU, a basic header split across fragments, a continuation with nothing pending.
One of those deserves singling out, because it is a design decision rather than a test case. The size limit is enforced on the declared length, before any copy and before the single-fragment fast path. Checking it only on the reassembly path would let a peer bypass the limit entirely by sending an over-long PDU in one unfragmented piece — so the largest PDU accepted does not depend on how the peer chose to fragment it. There is a test named after exactly that.
ATT contributes a third kind: PDUs are compared against literal expected octets, not merely round-tripped. A decoder that shares a bug with its encoder agrees with itself perfectly, so a round-trip test alone would pass happily on a stack that no real peer could talk to. Both kinds of test are present, and they check different things.
The GATT server adds a different class again, because it is the first layer that owns a database rather than just a packet. Its tests use a real Heart Rate Service and Device Information layout, so discovery is exercised with the same handle arithmetic a phone would produce, and the expected responses are written out octet by octet.
Three server decisions are worth stating outright:
- Characteristic declarations are synthesised, never stored. A declaration
contains its value handle. Storing that handle means writing it out by hand
or fixing the table up in RAM at start-up, and either way it can drift from
where the attribute actually sits. The server builds the declaration when it
is read — the value handle is always its own handle plus one — so the table
stays
const, stays in flash, and cannot go stale. - A command is never answered, not even to refuse it. Every failure path has to know whether the PDU it is rejecting was a request. A client that receives a reply to a command cannot tell which request it belongs to.
- A second MTU exchange does not move the MTU. The peer has already sized its receive path against the first answer; shrinking it underneath loses packets.
The client's tests mostly drive it against the real server rather than against hand-written response PDUs. The thing worth testing is the procedure loop — ask, read the answer, ask again from the next handle, stop when the server says there is no more — and a loop tested against canned responses only proves it agrees with whatever the test author imagined. Wiring the two halves together tests both, and catches either drifting from the specification independently. The tests that do use hand-written PDUs are the ones about hostile or broken peers, which a well-behaved server never produces.
Two client rules are worth stating outright:
- One outstanding request at a time. ATT has no transaction identifier, so a client with two requests in flight cannot tell which response belongs to which. Starting a second procedure returns an error rather than sending a request that would confuse both ends.
- A response must match the request it claims to answer. A stray or duplicated response is ignored rather than allowed to advance the state machine — acting on the wrong one is how a discovery loop ends up reporting attributes that do not exist.
And one from the long-write flow:
- Execute Write is atomic. The specification permits failing partway through a queued commit and naming the handle that failed, but that leaves an attribute half-written with no way for the client to learn how far it got. The whole queue is known before the first octet is committed, so it is validated up front and either all of it lands or none does.
Two ATT decisions are worth stating outright:
- Signed Write Command is refused, not decoded. Verifying its 12-octet signature needs a Security Manager, and there is not one yet. Decoding it as an ordinary Write Command would turn an unverifiable signature into an accepted write — precisely the authentication bypass the signature exists to prevent. Answering "not supported" is better than answering wrongly.
- Trailing octets on a fixed-length PDU are rejected. ATT PDU lengths are exactly specified, so extra octets mean the peer built the packet wrong. Tolerating them would admit two encodings of the same PDU.
The virtual radio adds a kind of test the protocol layers do not need: one about framing. A stream socket has no message boundaries, so a read may deliver half an ACL packet or three and a bit, and the transport has to accumulate until a whole one is present before L2CAP can begin its own reassembly. Its tests split writes in deliberately awkward places, including one octet at a time — which no kernel would normally do, but any kernel is permitted to.
That second reassembly layer is not an artefact of using sockets. A real HCI UART transport is also a byte stream and needs exactly the same reframing, which is why the virtual radio has the same shape as the real thing rather than being a shortcut around it. For the same reason it speaks the actual HCI ACL header rather than an invented framing: swapping the socket for a UART is then a change of transport, not a change of protocol.
Everything runs at every commit:
| Check | What it catches |
|---|---|
make test |
Behaviour, under -Werror -Wconversion on gcc and clang, C99 and C11 |
make asan |
Off-by-one reads and writes, as located failures instead of wrong values |
make fuzz-run |
Malformed-input handling nobody thought to write a test for |
| big-endian CI job | Byte-order assumptions that happen to hold on x86 |
make coverage |
Untested branches — 89.5% lines, 95.0% functions, 77.0% branches |
make tidy / cppcheck |
Static-analysis findings |
make size-report |
Flash and RAM regressions on a Cortex-M33 |
make demo |
The whole stack wired together, in two processes, under sanitizers |
The fuzz harnesses assert more than "it did not crash". The advertising harness checks on every input that iteration terminates, that every element the parser yields lies entirely inside the input buffer, and that the convenience accessors agree with the iterator. Reading the last octet of each element gives ASan a concrete access to catch, which turns it from a crash detector into a bounds checker.
L2CAP is a state machine rather than a parser, so a bug there needs a particular sequence of fragments to reach, not a single malformed packet. Its harness therefore interprets the fuzzer's input as a script of fragments — length-prefixed, with the boundary flags under the fuzzer's control — so coverage-guided search can discover the orderings that matter instead of only ever seeing well-formed single fragments. The invariant it leans on hardest is that a rejected fragment always leaves the receiver idle: splicing two PDUs together is how a reassembler turns a protocol violation into a memory-safety bug.
The ATT harness leans on a third property: the list-bearing responses carry a peer-supplied item size that the decoder divides the list by, and dividing a buffer by an attacker-chosen stride is exactly how a reader walks off the end of it. So every item the iterators yield is checked to lie inside the input, and the entry count is checked against the list length.
The GATT harness is scripted like the L2CAP one, since the server is stateful too — the MTU, the subscriptions and the outstanding-indication flag all persist across PDUs, and the orderings are where the interesting bugs live.
It found one. Not the crash it first reported, which turned out to be an over-strong property in the harness itself: guarding every octet past the reported response length fails legitimately, because a handler may reserve space, start filling it, then hit a bad handle and rewind to emit a shorter error, leaving harmless scratch behind. The bound that matters is the capacity the caller offered, not the length reported.
Chasing that down surfaced a real defect next to it. The server subtracted a read callback's reported length from a reserved size; an application callback that over-reported would underflow that subtraction and wrap the write cursor, turning an ordinary application bug into memory corruption inside the library. Callback-reported lengths are now clamped, with a test that supplies a deliberately lying callback.
Coverage is measured with the demo included, not only the unit tests. The transport's listen, connect and close paths are genuinely exercised — by two processes over a real socket — and counting only the unit tests reported them as dead code nobody had tested, which was misleading rather than merely pessimistic. With the demo counted the figure moved from 79.6% to 89.5% without a line of production code changing.
The name tables get their own tests, which look like busywork and are not. A table drifts silently: someone adds an opcode, forgets the switch, and the only symptom is a log line reading "Unknown" during an incident. Walking every defined value moves that failure to the build.
The client gets a harness too, because a client trusts its server rather less than people assume: every handle, length and item size driving the discovery loop came from the peer. Its most valuable assertion is about liveness rather than memory safety — that discovery terminates. A server answering with handles at or below the one already requested could otherwise walk the client round the same request forever, and no sanitizer would ever notice.
It found one. The ATT decoder accepted an Error Response carrying error code 0x00, which does not exist. That is worse than merely malformed: an application checking the code rather than the opcode would read zero as "no error" and carry on as though its request had succeeded. Both the decoder and the encoder now refuse it.
Most recent run, 45s per harness: 4.7M inputs through the advertising parser, 15.7M through the ATT decoder, 9.9M through L2CAP reassembly, 1.6M scripted exchanges against the GATT server and 1.4M against the client. No crashes and no assertion failures.
The test harness in tests/mble_test.h is deliberately
small and dependency-free, with assertion macros named to match Unity's. The
suite can migrate to Unity by deleting one header and adding one dependency,
without editing a single test — but until that trade is worth making, the
project builds anywhere a C99 compiler exists, with no package manager
involved.
Nothing but a C99 compiler is required.
makeOther targets:
make asan # rebuild and test under AddressSanitizer + UBSan
make fuzz-run # fuzz every harness (needs clang)
make coverage # line coverage report
make size-report # Cortex-M33 flash/RAM footprint per module
make check # everything CI runs
make helpCMake is also supported, for IDEs and cross-compilation toolchain files:
cmake -B build -DCMAKE_BUILD_TYPE=Debug
cmake --build build
ctest --test-dir build --output-on-failure- Bounds-checked buffer cursors, shared types
- GAP advertising and scan response data codec
- Unit test harness, Makefile, CMake, CI, fuzzing, sanitizers
- L2CAP fragmentation and reassembly over the LE fixed channels
- UUIDs, including 16-bit/128-bit equivalence
- ATT protocol codec: all opcodes, error responses, MTU exchange
- GATT server: attribute table, discovery, read/write, notify/indicate
- Virtual radio transport — two processes connecting over a UNIX socket
- Demo peripheral and central, exchanging real GATT traffic
- Cortex-M33 footprint report in CI
- GATT client: discovery, read, write, subscribe
- Prepare/Execute Write, with an atomic commit
- Bluetooth Core Specification, Vol 3, Part A — L2CAP
- Bluetooth Core Specification, Vol 4, Part E §5.4.2 — HCI ACL data packets
- Bluetooth Core Specification, Vol 3, Part C §11 — advertising data format
- Bluetooth Core Specification, Vol 3, Part B §2.5.1 — UUIDs and the Base UUID
- Bluetooth Core Specification, Vol 3, Part F — Attribute Protocol
- Bluetooth Core Specification, Vol 3, Part G — Generic Attribute Profile
- Assigned Numbers — Common Data Types, Company Identifiers
MIT