mortred_model_server

Inference CI: what each path is allowed to claim

Weights are not in git. GitHub holds code, conf/weights_manifest.json (path + sha256), conf/ci_hosted_golden.json (which goldens which path must run), and test/golden/ expectations. Bytes live on Hugging Face (MaybeShewill-CV/mortred_model_server) and, for CUDA/TensorRT, on the maintainer GPU runner at /opt/mortred-cache/weights. Dual-file ONNX export and remaining gaps: onnx-interchange.md.

A green check must not be read as “every model still works on every contribution path.” Use the table.

Path Runner Green means Does not mean
Fork PR GitHub ubuntu-22.04 (cpu-profile) cpu profile compiled; output contracts passed; MNN/ONNX goldens in conf/ci_hosted_golden.json hosted set (sha256 locked, skipped=0): classification, NanoDet, YOLOv8 ONNX overlay, OCR, keypoints, segmentation TensorRT, CUDA, YOLOv8 .engine, full zoo, ORT-CUDA
Hosted container boot (cpu compose) GitHub ubuntu-22.04 On Dockerfile / compose / entrypoint / demo pack / this workflow: cpu runtime image starts; supervisor /api/v1/health; gateway /healthz; one MOBILENETV2 infer; supervisor and gateway still in docker top. Unrelated PRs skip the image build and stay green GPU compose, TensorRT, GHCR :gpu pull, every catalog model
Same-repo PR / push main, no MORTRED_HAS_GPU_RUNNER GitHub hosted only Same as fork PR. GPU jobs are skipped, not queued TensorRT / CUDA goldens ran
Same-repo PR / push main, variable true Hosted + self-hosted GPU Fork claims plus 8 golden cases, skipped=0 Nightly zoo or cross-backend allclose
Schedule / workflow_dispatch + variable true GPU runner (gpu-nightly-full) Smoke-8 fail-closed; remaining goldens may skip if listed in the skip-inventory artifact Bit-exact MNN/ONNX/TRT

The required GitHub check should be inference paths (job inference-paths), not gpu golden smoke by itself. GPU jobs skip when the repository variable MORTRED_HAS_GPU_RUNNER is unset (no self-hosted runner registered) and on fork PRs. The wrapper treats that skip as success only when cpu-profile succeeded. Do not set inference paths required while GPU jobs are still queued waiting for a missing runner.

container boot (cpu compose) is a separate Docker-runtime check. It is safe to require once it has gone green: unrelated PRs only run checkout + path filter. Do not read it as GPU compose coverage. --profile gpu boot belongs on a self-hosted GPU runner (MORTRED_HAS_GPU_RUNNER=true), not on GitHub-hosted VMs.

Changing src/models/backend/trt_session.cpp (or any TensorRT-only path) is not proven on a fork PR. Maintainer PRs need MORTRED_HAS_GPU_RUNNER=true so gpu-pr-gate runs. Hosted CPU goldens use device=cpu via force_cpu_backend and only mnn/onnx configs. Product yolov8_config.toml (type=tensorrt) stays on the GPU smoke list; the fork-visible YOLO case is the CI overlay conf/ci/yolov8_onnx_hosted.toml.

Enable maintainer GPU jobs

GitHub does not let GITHUB_TOKEN list self-hosted runners, so presence is an explicit flag:

  1. Register a Linux x64 runner with labels self-hosted, X64, gpu.
  2. Create /opt/mortred-cache/weights and /opt/mortred-cache/3rd_party.
  3. Repo Settings → Secrets and variables → Actions → Variables → MORTRED_HAS_GPU_RUNNER = true.
  4. Delete the variable (or set it to anything else) to skip GPU jobs again.

Hosted fork liveness (cpu-profile)

Source of truth: conf/ci_hosted_golden.json. scripts/check_hosted_golden.py (also invoked from scripts/check_consistency.py) rejects a set that is not on Hugging Face, not tagged cpu, that points at a TensorRT engine, or that has no committed test/golden/<case>.json / .png (so a missing YOLO ONNX baseline cannot hide behind a green consistency job).

  1. Cache weights/ keyed by that JSON + conf/weights_manifest.json.
  2. scripts/fetch_weights.py --only … for each hosted weight (stdlib urllib if requests / huggingface_hub are absent).
  3. --check against the manifest sha256 (mismatch fails the job).
  4. MORTRED_CI_REQUIRE_WEIGHTS=1 on backend_unittest (MnnSession + config tests) and the hosted gtest filter from --print-gtest-filter.
  5. scripts/ci_assert_gtest_xml.py rejects zero tests or any skip.

The earlier cmake --build --target check step still allows skip-as-pass for the rest of the zoo (local-style). Only the explicit hosted step is fail-closed.

Local developers without weights keep GTEST_SKIP (env unset).

YOLOv8 HTTP serving still uses conf/model/object_detection/yolov8/yolov8s.toml (TensorRT). Fork-hosted detection coverage is NanoDet MNN plus the YOLOv8 ONNX overlay (conf/ci/yolov8_onnx_hosted.toml, yolov8s.onnx). That overlay is not a product serving config: conf/server must keep pointing at the TensorRT toml. Hosted ONNX passing does not prove the serving engine. TensorRT plus device=cpu remains a configuration error.

Maintainer GPU smoke (gpu-pr-gate)

Runs only when all of these hold:

Eight cases (gpu_smoke.cases in conf/ci_hosted_golden.json, must match MORTRED_GPU_SMOKE_FILTER in .github/workflows/ci.yml):

Case Family
yolov5_detection object detection
yolov8_detection TensorRT decode + letterbox geometry
yolov8_mixed_size_batch_matches_single_runs mixed-size batch
nanodet_detection anchor-free decode
centerface_detection landmarks
dbnet_text_detection OCR boxes
mobilenetv2_classification scores + k_score_tol
fastsam_segmentation fingerprint png

Engine refresh is convert_trt_engines.sh --only yolov8 only. The other smoke cases use MNN files, which are not in conf/trt_engines.json; converting those names fails the job. Missing cache or a gtest skip fails the job (MORTRED_CI_REQUIRE_WEIGHTS=1 + XML audit).

Goldens still force device=cpu in the test harness (force_cpu_backend). MNN cases therefore run MNN-CPU even on the GPU box; type=tensorrt (YOLOv8) still uses CUDA. Changing that requires a golden refresh on the runner — a follow-up, not this gate.

Nightly (gpu-nightly-full)

Same MORTRED_HAS_GPU_RUNNER=true gate as the PR smoke (otherwise the schedule would queue for 120 minutes). Smoke-8 is fail-closed. The rest of the committed golden zoo runs with --allow-skips so an incomplete cache prints an inventory instead of a fake all-green. The skip list is written to gpu-rest-skips.json and uploaded with the nightly log artifact. model_lifecycle_unittest, backend_unittest, and gateway_e2e_test still run via ctest.

Workflow concurrency includes github.event_name so a push to main does not cancel a running schedule.

ONNX cutover freeze (P0)

The Hugging Face ONNX-as-interchange work (onnx-interchange.md) must not silently refresh goldens. Source of truth for the hosted set remains conf/ci_hosted_golden.json (today: MobileNetV2 MNN, NanoDet MNN, YOLOv8 ONNX overlay, DBNet MNN, SuperPoint MNN, BiseNetV2 MNN). GPU smoke-8 stays the list in that file / MORTRED_GPU_SMOKE_FILTER.

P1 only uploads existing ONNX and updates conf/weights_manifest.json. It does not change test/golden/ or hosted cases. A later P2 overlay that cannot match these baselines must update golden in the same PR with a reason; “we exported a new ONNX” is not enough.

Contract tiers (T0 / T1 / T2)

These labels are how we talk about verification depth. They map onto existing gates; they are not a separate binary.

Tier What it is Where it runs Green means
T0 Weight-free synthetic contracts (fake session / helpers): shape overflow, short buffers, IO name set, short item_status/outputs, finite opt-in, int32 range, output-count match tests-only / cmake --build … --target check Those unittests passed. Not a claim about real models
T1 Small real weights, sha256-locked GitHub-hosted cpu-profile via conf/ci_hosted_golden.json hosted set Hosted goldens ran with skipped=0
T2 Full / nightly zoo gpu-nightly-full / catalog_tiers.nightly Nightly executed; if T2 did not run, summaries must say so (never look like full-zoo green)

CI prints this table via scripts/ci_print_tier_summary.py from inference paths (PR/push always marks T2: NOT RUN) and from gpu-nightly-full / T2 tier notice on schedule/manual.

T0 owners (non-exhaustive; prefer adding cases here over new binaries): tensor_contract_unittest, session_io_unittest, batch_collector_unittest, model_runtime_unittest, param_spec_unittest, packed_batch_nvi_unittest.

Catalog CI tiers

Every HTTP catalog id in src/factory/*_task.h must appear in catalog_tiers inside conf/ci_hosted_golden.json:

Tier Meaning
hosted Fail-closed on GitHub-hosted cpu-profile (fork-visible). YOLOV8 is hosted via the ONNX overlay; GPU smoke still runs the product TensorRT golden
gpu-smoke Fail-closed on maintainer GPU PR gate; not claimed on forks
nightly Allowed to skip on PR; exercised on gpu-nightly-full when weights exist

Adding an HTTP model without a tier fails python3 scripts/check_consistency.py.

Self-hosted runner layout

Item Requirement
OS Ubuntu 22.04 LTS
NVIDIA driver >= 535.x
CUDA / TensorRT 12.x / 10.3 — match 3rd_party/
Labels self-hosted, X64, gpu
Concurrency One runner process per machine
/opt/mortred-cache/
|-- weights/      # HF blobs + this GPU/TRT pair's .engine files
`-- 3rd_party/    # scripts/install_deps.sh --all

Jobs symlink those into the checkout. Engines are GPU- and TRT-version specific: after this CUDA 12 / TensorRT 10 bump (and any later driver/TRT bump), wipe /opt/mortred-cache/3rd_party and weights/**/*.engine, then install_deps.sh --all && --check and rebuild engines from ONNX before PR traffic.

Preflight before relying on fail-closed smoke:

# on the runner, repo root with weights linked
python3 scripts/fetch_weights.py --only yolov8
bash scripts/convert_trt_engines.sh --only yolov8
MORTRED_CI_REQUIRE_WEIGHTS=1 ./build-gpu/bin/model_golden_test \
  --gtest_filter="$MORTRED_GPU_SMOKE_FILTER" \
  --gtest_output=xml:/tmp/gpu-smoke.xml
python3 scripts/ci_assert_gtest_xml.py /tmp/gpu-smoke.xml

Refreshing goldens

After any intentional change to test/golden/* or golden case declarations in test/model_golden_test.cc, reset the zero-drift baseline in the same PR:

python3 scripts/golden_drift_check.py --record   # updates test/golden_baseline.json
python3 scripts/golden_drift_check.py --check    # must exit 0

scripts/check_consistency.py (CI) runs --check fail-closed. A green model_golden_test only proves the binary matches the current files; the baseline proves those files did not silently change relative to the last recorded inventory + sha256 map.

yolov8_onnx_detection is regenerated on a cpu-profile build with yolov8s.onnx (WSL is fine). Product YOLO TensorRT goldens (yolov8_detection) still need the same GPU runner that gates PRs:

MORTRED_UPDATE_GOLDEN=1 ./build-cpu/bin/model_golden_test \
  --gtest_filter='model_golden.yolov8_onnx_detection'
MORTRED_UPDATE_GOLDEN=1 ./build-gpu/bin/model_golden_test \
  --gtest_filter='model_golden.yolov5_detection:model_golden.yolov6_detection:model_golden.yolov7_detection:model_golden.yolov8_detection'
git add test/golden/

Update-mode GTEST_SKIPs after writing artifacts (not a weight skip). Commit golden files separately from logic changes.

Valgrind (A6 / curated memory-check)

Hosted tests installs valgrind and configures with -DMORTRED_ENABLE_VALGRIND_TESTS=ON. That registers a small curated set of weight-free unittests under ctest label valgrind (not the full suite, and not part of --target check). CI runs:

ctest --test-dir build-ci -L valgrind --output-on-failure

Locally (tests-only preset):

cmake --preset tests-only -DMORTRED_ENABLE_VALGRIND_TESTS=ON
cmake --build --preset tests-only --target status_code_unittest base64_unittest \
  param_spec_unittest tensor_contract_unittest session_io_unittest
ctest --test-dir build/tests-only -L valgrind --output-on-failure

Memcheck uses test/valgrind.supp (wired into the *-memory-check command) to ignore known third-party Memcheck:Param noise from libglog/libunwind during .init / _Unwind_Backtrace. Do not add suppressions for definite leaks in mortred code.