A small CPU inference server in C++20 and ONNX Runtime, built to learn the systems behind real model serving: request admission, bounded queues, dynamic batching, backpressure, graceful shutdown, and reproducible load testing.
This is a learning / portfolio project, not a replacement for Triton, TensorFlow Serving, or a managed platform.
The milestone curriculum and issue-by-issue build guide live in
README_Inference_Server_From_Scratch.md.
| Area | State |
|---|---|
| Model contract | Deterministic [batch, 4] → [batch, 2] fixture |
| Offline CLI inference | Working |
| Contract + finite-input validation | Working |
| Unit + CLI end-to-end tests | 47/47 passing |
| UBSan | Passing |
| ASan on current AppleClang | Blocked by toolchain deadlock before main() |
HTTP /predict server |
Not started (Milestone 2) |
| Bounded queue / backpressure | Planned (Milestone 3) |
| Dynamic batching | Planned (Milestone 4) |
| Metrics + load client | Planned (Milestones 5–6) |
Next: Milestone 2 — single-request HTTP inference.
./build/inference_cli --model models/example.onnx --input 0.25,-0.5,1.0,0.0inputs (1)
[0] features: float32 [batch, 4]
outputs (1)
[0] scores: float32 [batch, 2]
outputs: 0.25 0.5
The CLI:
- parses
--modeland--input - loads the ONNX model once
- validates the tensor contract (
features/scores, shapes,float32) - rejects non-finite values and wrong feature counts
- runs one sample and prints the scores
Client request
|
v
+-------------------+
| HTTP boundary | validate JSON, shape, and size
+---------+---------+
|
v
+-------------------+ full / closing
| Bounded FIFO | --------------------------> HTTP 503
+---------+---------+
|
v
+-------------------+
| Dynamic batcher | max batch size OR oldest-request deadline
+---------+---------+
|
v
+-------------------+
| ONNX Runtime | one session, CPU execution provider
+---------+---------+
|
v
+-------------------+
| Result splitter | output[i] -> request[i]
+---------+---------+
|
v
HTTP response + measurements
Version 1 target API:
POST /predict
Content-Type: application/json
{"inputs":[0.25,-0.5,1.0,0.0]}{
"request_id": 42,
"outputs": [0.25, 0.5]
}Also planned: GET /health, GET /ready, GET /stats.
| Decision | Choice |
|---|---|
| Language | C++20 |
| Runtime | ONNX Runtime C++ API |
| Compute | CPU only |
| Models | One model loaded once at startup |
| Queue | Bounded FIFO with explicit overload rejection |
| Batching | Max batch size or oldest-request deadline |
| HTTP | Small synchronous server (cpp-httplib) for v1 |
| Tests | GoogleTest + CTest |
Details and rationale are in the from-scratch guide.
Prerequisites: CMake 3.20+, a C++20 compiler, and the ONNX Runtime C/C++ SDK.
cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DONNXRUNTIME_ROOT=/path/to/onnxruntime
cmake --build build
ctest --test-dir build --output-on-failure./build/inference_cli --model models/example.onnx --input 0.25,-0.5,1.0,0.0More build and sanitizer notes: docs/build.md.
src/ CLI, argument parsing, ModelSession
tests/ GoogleTest + CLI smoke tests
models/ Example ONNX fixture
scripts/ Model generation script
docs/ Build notes and model contract
lessons/ Short HTML lessons
reference/ Quick-reference pages
learning-records/ Architecture / mission notes
| Doc | Purpose |
|---|---|
README_Inference_Server_From_Scratch.md |
Milestone curriculum and tickets |
MISSION.md |
Project mission and success criteria |
docs/model-contract.md |
Tensor I/O contract for the example model |
docs/build.md |
Build and sanitizer instructions |
INTEGRATION_ROADMAP.md |
Later VectorDB / semantic-search integration |
RESOURCES.md |
Reading list |