All model configurations live in conf/model/. The unified backend layer uses
a two-table schema per model section:
[YOLOV8] # model section, one per model
[YOLOV8.backend] # backend selection (shared by all models)
type = "tensorrt" # mnn | onnx | tensorrt
model_file_path = "../weights/object_detection/yolov8/yolov8s.engine"
device = "gpu" # cpu | gpu (omitted defaults to gpu; tensorrt forbids cpu)
device_id = 0 # optional, cuda device index
threads = 4 # cpu threads for mnn/onnx (default 4)
input_layout = "auto" # mnn only: auto | nhwc | nchw
precision_mode = 0 # mnn only: BackendConfig::PrecisionMode
power_mode = 0 # mnn only: BackendConfig::PowerMode
input_names = ["images"] # optional, defaults to the model file io
output_names = ["output0"] # optional, filters aux output nodes
[YOLOV8.params] # model specific keys, consumed by on_init
model_score_threshold = 0.25
model_nms_threshold = 0.5
model_input_image_size = [640, 640] # [height, width]; must match fixed model inputs
max_image_pixels = 16777216 # decoded-pixel safety limit
max_image_side = 8192
class_names = ['person', 'bicycle']
| key | scope | description |
|---|---|---|
type |
backend | inference engine: mnn, onnx (ONNX Runtime) or tensorrt |
model_file_path |
backend | weights file (.mnn, .onnx, .engine) |
device |
backend | cpu or gpu; omitted defaults to gpu. type=tensorrt with device=cpu is a configuration error |
device_id / gpu_device_id |
backend | cuda device index (alias accepted) |
threads |
backend | intra-op threads for mnn / onnx cpu |
gpu_mem_limit_mb |
backend (onnx+cuda) | CUDA EP arena cap per session/worker; default 2048; 0 = unlimited (legacy). Override with MORTRED_ORT_GPU_MEM_LIMIT_MB. MNN/TRT ignore it. |
input_layout |
backend (mnn) | host tensor byte order: nhwc for TF-style exports, nchw for CHW exports, auto follows the model file |
precision_mode / power_mode |
backend (mnn) | MNN::BackendConfig modes |
input_names / output_names |
backend | io name override/filter, useful for models exposing auxiliary outputs |
max_image_pixels / max_image_side |
image model params | decoded input safety limits; defaults are 16777216 and 8192 |
model_input_image_size |
fixed-image model params | [height, width]; must match the session input H/W |
min_box_area_px |
YOLOv5 / v6 / v7 / v8 params | drop boxes whose width*height in network (letterboxed) pixels is below this; default 5; compared before letterbox unmap |
sample_size |
diffusion [DDPM_SAMPLER] / [DDIM_SAMPLER] / [LDM_SAMPLER] |
[height, width] of the generated canvas; must match the UNet weight. HTTP seeds this at adapter init; it is not a request params key |
| everything else | params | model specific (thresholds, class names, sizes); key names are unchanged from the historical configs |
Multi-engine models (SAM encoder + decoder, lightglue extractor + matcher) use
one <key>_backend sub-table per engine instead of the primary backend
table, and orchestrate the sessions in run_sessions.
The historical schema ([SECTION] + backend_type + [SECTION_TRT] /
[SECTION_ONNX] / [SECTION_MNN] + [BACKEND_DICT]) is no longer supported.
Migrate with:
python scripts/migrate_model_config.py --dry-run # preview + report
python scripts/migrate_model_config.py # in-place migration
python scripts/migrate_model_config.py --check # CI gate (exit 1 on drift)
Mapping: compute_backend -> device, gpu_device_id -> device_id,
model_threads_num -> threads, backend_precision_mode -> precision_mode,
backend_power_mode -> power_mode, trt -> tensorrt; all other keys move
into [SECTION.params] with unchanged names and semantics.
model_golden_test rewrites ../-prefixed paths relative to the repo root and
forces backend.device = "cpu" so cpu-only CI stays deterministic. That rewrite
makes product TensorRT configs (for example yolov8_config.toml) a
configuration error on CPU (type=tensorrt forbids device=cpu). Hosted fork
CI proves YOLOv8 decode via the overlay conf/ci/yolov8_onnx_hosted.toml, not
the serving .engine. TensorRT goldens stay on the maintainer GPU path.
DDPM HTTP generate is proven by conf/ci/ddpm_onnx_fewstep.toml +
model_golden.ddpm_celeba_hq_fewstep (not hosted; the ONNX is ~143MiB).