Single entry for adding and hardening a CV model in this repo. For mounting a finished model on HTTP (conf/server, OpenAPI, consistency), see
how_to_add_new_server.md.
Every command below is runnable from the repository root.
All CV models inherit
BackendCvModel<INPUT, OUTPUT>:
init: parse [SECTION.backend] -> create InferenceSession -> on_init([SECTION.params])
run_impl: prepare_inputs -> session.run -> postprocess(context)
A standard single-image model implements preprocess (cv::Mat → named
tensors) and postprocess (named tensors + request geometry → task output).
Backend plumbing (MNN / ORT / TensorRT sessions, dtype & shape checks, copies)
lives under src/models/backend/ and is not repeated
per model.
Prefer the runtime toolkit:
ImagePipeline, OutputReader, SessionIoValidator. Reference models:
mobilenetv2 (MNN),
yolov8_detector (TRT),
ddpm_unet (ORT, non-image
prepare_inputs).
IO types: src/models/io/ (prefer task std_*_output).
Request-scoped geometry belongs in InferenceContext, never in model members
(dynamic batch). Malformed tensors must return MODEL_OUTPUT_CONTRACT_FAILED
via OutputReader / f32_output.h.
Skeleton:
template<typename INPUT, typename OUTPUT>
class MyModel : public jinq::models::BackendCvModel<INPUT, OUTPUT> {
public:
MyModel() : jinq::models::BackendCvModel<INPUT, OUTPUT>("MY_MODEL") {}
private:
std::vector<jinq::models::backend::NamedTensor> preprocess(const cv::Mat& image) override;
jinq::common::StatusCode postprocess(
const std::vector<jinq::models::backend::NamedTensor>& outputs,
const jinq::models::backend::InferenceContext& context,
OUTPUT& output) override;
jinq::common::StatusCode on_init(const toml::table& params) override; // optional
};
Repo-wide decoded-image caps default to max_image_pixels = 16777216 and
max_image_side = 8192 (see server/model config). A model may raise
max_image_pixels when full-size camera input is normal (e.g. MODNet /
PPMatting); document the override in that model TOML.
Config shape (full keys: about_model_configuration.md):
[MY_MODEL]
[MY_MODEL.backend]
type = "mnn" # mnn | onnx | tensorrt
model_file_path = "../weights/my_model/model.mnn"
device = "gpu"
threads = 4
[MY_MODEL.params]
score_threshold = 0.25
Catalog: one CvModelEntry in src/factory/<task>_task.h (see Path 1).
HTTP-only families use cv_catalog.h;
benchmark-only families use model_catalog.h
(CLIP / SAM predictor / FastSAM — no HTTP by product lock). Multiple output
contracts → multiple typed catalogs (catalog() / face_catalog() in
obj_detection_task.h).
python scripts/new_model.py --list-tasks
python scripts/new_model.py --task classification \
--name efficientnet --class EfficientNet \
--backend mnn --dry-run
python scripts/new_model.py --task classification \
--name efficientnet --class EfficientNet \
--backend mnn
Produces header / .inl / TOML / output-contract unittest (and prints catalog +
CMake snippets it does not auto-apply). Unimplemented hooks return
MODEL_NOT_IMPLEMENTED so a half-finished model cannot be served.
src/models/object_detection/rtdetr_detector.* is the checked-in scaffold canary.
Fill hooks like mobilenetv2.inl:
const auto info = SessionIoValidator(session())
.input().f32().rank(4).nhwc().channels(3).static_shape().validate();
return ImagePipeline(image)
.resize(_m_input_size)
.bgr_to_rgb()
.to_float()
.scale(1.0f / 255.0f)
.nhwc(session().inputs().front().name);
auto view = OutputReader(outputs, outputs.front().name)
.f32().shape({1, -1}).finite().read();
Paste the scaffolder’s catalog row into src/factory/classification_task.h and
the test target into test/CMakeLists.txt.
Same commands with --task object_detection. Differences:
std_object_detection_output or std_face_detection_output.detector_common.h
for validation / NMS / top-k. Geometry is not one helper: Ultralytics YOLO
needs ImagePipeline::letterbox + unmap_letterbox_bbox; stretch
GeometryScale is for NanoDet / CenterFace / LibFace.catalog() and face_catalog() — do not type-erase them.Reference: yolov8_detector.inl.
POSTPROCESS_CONTRACT_TEST(EfficientNet, mat_input,
std_classification_output, "output", 1, 1000);
One macro → seven filterable tests (rejects_missing_output, wrong dtype/rank/shape,
short buffer, nan, inf). Scaffold models may still return MODEL_NOT_IMPLEMENTED
(harness accepts that as explicit rejection).
auto view = OutputReader(outputs, "output")
.f32().shape({1, -1}).finite().read();
if (!view.ok()) {
return view.status;
}
Hosted CI fail-closes only a small MNN CPU set; GPU smoke is a maintainer gate. See ci-golden-regression.md.
GOLDEN_CLASSIFICATION_CASE(efficientnet_classification,
"conf/model/classification/efficientnet/efficientnet_config.toml",
"demo_data/model_test_input/classification/ILSVRC2012_val_00000003.JPEG",
jinq::factory::classification::create_efficientnet_classifier,
std_classification_output);
Macros: GOLDEN_CLASSIFICATION_CASE, GOLDEN_OBJECT_DETECTION_CASE,
GOLDEN_FACE_DETECTION_CASE, GOLDEN_SCENE_SEGMENTATION_CASE,
GOLDEN_MATTING_CASE, GOLDEN_ENHANCEMENT_CASE, GOLDEN_TEXT_REGION_CASE,
GOLDEN_KEYPOINT_CASE, GOLDEN_RAW_MAT_CASE.
LD_LIBRARY_PATH=<build>/lib:3rd_party/libs \
MORTRED_UPDATE_GOLDEN=1 <build>/bin/model_golden_test \
--gtest_filter='model_golden.efficientnet_classification'
Missing weights → case skips (expected on CPU-only boxes).
python scripts/golden_drift_check.py --record # before a migration
python scripts/golden_drift_check.py --check # after it
test/golden_baseline.json records case names and golden file hashes. Prefer
colour golden inputs so channel-order swaps are visible.
Read the error literally (SessionIoValidator / catalog tests name engine,
direction, tensor). Common causes: config type vs file mismatch; NHWC/NCHW
pack wrong; dynamic batch vs static expect; session("...") name missing from
sessions(); TOML section ≠ constructor section string.
for (const auto &info : session().inputs()) {
LOG(INFO) << "input " << info.to_string();
}
| Helper | Scope | Deliberately not |
|---|---|---|
ImagePipeline |
resize, crop, colour, normalise, NCHW/NHWC | keep-ratio pad (depth, fastsam), non-image inputs |
OutputReader |
named f32 contract | int32 tokens, multi-output ordering |
SessionIoValidator |
one named IO pair | optional / alternative IO |
GOLDEN_*_CASE |
single-image seven-step | batch equivalence, multi-session flows |
MultiSessionModel |
fixed named engines | concurrent pools, model composition |
Examples kept hand-written on purpose: SamAutoMaskGenerator (session pool), LDM (composed BackendCvModels). Comment why when you step outside a helper.
cmake --preset full && cmake --build --preset full
scripts/run_tests.sh build/full -R model_golden_test --output-on-failure
python3 scripts/check_consistency.py
enlightengan stays off ImagePipeline (dual tensor / custom luma / alpha).