Skip to content

About

A C++20 CPU inference server built from scratch with ONNX Runtime, contract validation, dynamic batching, backpressure, and reproducible benchmarks.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

C++ Inference Server

A small CPU inference server in C++20 and ONNX Runtime, built to learn the systems behind real model serving: request admission, bounded queues, dynamic batching, backpressure, graceful shutdown, and reproducible load testing.

This is a learning / portfolio project, not a replacement for Triton, TensorFlow Serving, or a managed platform.

The milestone curriculum and issue-by-issue build guide live in README_Inference_Server_From_Scratch.md.


Status

Area State
Model contract Deterministic [batch, 4] → [batch, 2] fixture
Offline CLI inference Working
Contract + finite-input validation Working
Unit + CLI end-to-end tests 47/47 passing
UBSan Passing
ASan on current AppleClang Blocked by toolchain deadlock before main()
HTTP /predict server Not started (Milestone 2)
Bounded queue / backpressure Planned (Milestone 3)
Dynamic batching Planned (Milestone 4)
Metrics + load client Planned (Milestones 5–6)

Next: Milestone 2 — single-request HTTP inference.


What it does today

./build/inference_cli --model models/example.onnx --input 0.25,-0.5,1.0,0.0
inputs (1)
  [0] features: float32 [batch, 4]
outputs (1)
  [0] scores: float32 [batch, 2]
outputs: 0.25 0.5

The CLI:

  1. parses --model and --input
  2. loads the ONNX model once
  3. validates the tensor contract (features / scores, shapes, float32)
  4. rejects non-finite values and wrong feature counts
  5. runs one sample and prints the scores

Target architecture

Client request
     |
     v
+-------------------+
| HTTP boundary     | validate JSON, shape, and size
+---------+---------+
          |
          v
+-------------------+       full / closing
| Bounded FIFO      | --------------------------> HTTP 503
+---------+---------+
          |
          v
+-------------------+
| Dynamic batcher   | max batch size OR oldest-request deadline
+---------+---------+
          |
          v
+-------------------+
| ONNX Runtime      | one session, CPU execution provider
+---------+---------+
          |
          v
+-------------------+
| Result splitter   | output[i] -> request[i]
+---------+---------+
          |
          v
HTTP response + measurements

Version 1 target API:

POST /predict
Content-Type: application/json

{"inputs":[0.25,-0.5,1.0,0.0]}
{
  "request_id": 42,
  "outputs": [0.25, 0.5]
}

Also planned: GET /health, GET /ready, GET /stats.


Design choices

Decision Choice
Language C++20
Runtime ONNX Runtime C++ API
Compute CPU only
Models One model loaded once at startup
Queue Bounded FIFO with explicit overload rejection
Batching Max batch size or oldest-request deadline
HTTP Small synchronous server (cpp-httplib) for v1
Tests GoogleTest + CTest

Details and rationale are in the from-scratch guide.


Build

Prerequisites: CMake 3.20+, a C++20 compiler, and the ONNX Runtime C/C++ SDK.

cmake -S . -B build \
  -DCMAKE_BUILD_TYPE=Debug \
  -DONNXRUNTIME_ROOT=/path/to/onnxruntime
cmake --build build
ctest --test-dir build --output-on-failure
./build/inference_cli --model models/example.onnx --input 0.25,-0.5,1.0,0.0

More build and sanitizer notes: docs/build.md.


Repository layout

src/                 CLI, argument parsing, ModelSession
tests/               GoogleTest + CLI smoke tests
models/              Example ONNX fixture
scripts/             Model generation script
docs/                Build notes and model contract
lessons/             Short HTML lessons
reference/           Quick-reference pages
learning-records/    Architecture / mission notes

Documentation

Doc Purpose
README_Inference_Server_From_Scratch.md Milestone curriculum and tickets
MISSION.md Project mission and success criteria
docs/model-contract.md Tensor I/O contract for the example model
docs/build.md Build and sanitizer instructions
INTEGRATION_ROADMAP.md Later VectorDB / semantic-search integration
RESOURCES.md Reading list

About

A C++20 CPU inference server built from scratch with ONNX Runtime, contract validation, dynamic batching, backpressure, and reproducible benchmarks.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages