The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
-
Updated
Sep 7, 2026 - Python
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
A tool for developers and researchers to design FHE pipelines and execute them on cloud-based GPUs
Local AI workflow orchestration and runtime management framework
GPU-aware LLM serving platform with multi-worker routing, dynamic batching, model placement, streaming, admission control, and reproducible performance benchmarks.
Local AI inference platform built with FastAPI, React, Ollama, and SQLite. Features multi-model benchmarking, structured generation, experiment tracking, JWT authentication, and performance analytics for small language models running entirely offline.
A conceptual framework for a high-scale Agentic AI orchestrator, inspired by enterprise-grade inference platforms.
AI inference platform architecture lab demonstrating admission control, fairness scheduling, bounded queues, and graceful degradation under burst traffic.
To associate your repository with the inference-platform topic, visit your repo's landing page and select "manage topics."