GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.
-
Updated
Jun 18, 2026 - Cuda
GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.
Syllable-aware BPE tokenizer for the Amharic language (አማርኛ) – fast, accurate, trainable.
A ridiculously fast Python BPE (Byte Pair Encoder) implementation written in Rust
An educational Python project for learning tokenization step by step by building character-level, byte-level, and BPE tokenizers from scratch.
Implemented GPT from scratch
LLM inference engine built from scratch in C++. No PyTorch, no frameworks.
GPT-style language model with Byte Pair Encoding tokenizer, built from scratch in PyTorch.
A custom tokenizer (byte level BPE) trained by me to try to replicate LLAMA 2's massive token vocabulary
A PHP implementation of OpenAI's BPE tokenizer tiktoken.
R-BPE: Improving BPE-Tokenizers with Token Reuse
Multi-language BPE tokenizer implementation for Qwen3 models. Lightweight byte-pair encoding for C#/.NET
BPE tokenizer for LLMs in Pure Zig
Teaching transformer-based architectures
Byte-Pair Encoding (BPE) subword tokenizer trainer and deterministic text encoder/decoder.
Byte-Pair Encoding (BPE) subword tokenizer trainer and deterministic text encoder/decoder.
High-Performance Tokenizer implementation in PHP.
Experimental local transformer framework for training small language models, chat assistants, and code focused AI systems from scratch. Built for learning, research, and rapid LLM experimentation.
Byte-Pair Encoding tokenizer for training large language models on huge datasets
implementation of Byte-Pair Encoding (BPE) for subword tokenization, written entirely in C++ . The tokenizer learns merges from raw text and supports encoding/decoding with UTF-8
To associate your repository with the bpe-tokenizer topic, visit your repo's landing page and select "manage topics."