All notable changes to this project are documented here. The format follows Keep a Changelog; versions follow Semantic Versioning. 每个版本的条目保留英文原文,中文说明 以引用块附于同版本之下。
scripts/check_consistency.py (and CI) runs golden_drift_check.py --check fail-closed against test/golden_baseline.json; refresh docs require --record in the same PR. Baseline reset to current 30 cases / 26 golden files.docs/unsupported-boundaries.md (+ zh summary): Linux x64 only; RTDETR scaffold not in catalog; CUDA 12.x + TensorRT 10.x only (8/9 out); engines must be rebuilt on this GPU/TRT; cpu profile has no TensorRT. Linked from README, oob-main-path, deployment, and rtdetr model doc.docs/oob-main-path.md plus mortredctl next (scripts/mortredctl_next.sh) prints one next command (three tokens → loopback listen tip → start supervisor → pack prepare/calibrate → doctor --strict). init / init-trust / doctor WARN·FAIL / prepare / calibrate emit next: lines. README Quick Start points at the main path; other entries are variants.mortredctl doctor --strict.
TensorRT ids in an active pack need gpu_mem_mib from
mortredctl calibrate --write-pack (plus gpu_mem_at_workers and a GPU
fingerprint on [pack]). Missing stamps, stale worker_nums, and joint-budget
overflow refuse spawn with the next command in the error (no crash-loop).
occupancy_policy=off / MORTRED_OCCUPANCY_ENFORCE=0 skip the spawn gate
(unsafe) and fail --strict. gpu_mem_limit_mb stays ORT CUDA EP only.pack 占用在 supervisor spawn 和
mortredctl doctor --strict上 fail-closed。 已启用 pack 里的 TensorRT id 必须有calibrate --write-pack写入的gpu_mem_mib(以及gpu_mem_at_workers与[pack]GPU 指纹)。缺 stamp、worker_nums过期、联合预算超限都会拒绝 spawn,错误里是下一条命令,不进 crash-loop。occupancy_policy=off/MORTRED_OCCUPANCY_ENFORCE=0跳过 spawn 闸(unsafe),--strict仍失败。gpu_mem_limit_mb仍然只约束 ORT CUDA EP。
docs/shortest-path-cpu.md (linked from
README Quick Start) and SME-02 evidence under docs/evidence/.增加 cpu 源码树最短成功路径文档
docs/shortest-path-cpu.md(README Quick Start 已链入),以及 SME-02 evidence(docs/evidence/)。
oob-main-path.md (first-hour lane lives in deployment.md); remove sme-oob-checklist.md, async-job-table (+zh), historical model_inference_benchmark (+zh), and onnx-export-guide.md.model-contract-governance.md (rules live in model-developer-guide; image-size caps noted there).shortest-path-cpu.md into oob-main-path.md (single out-of-box lane); remove docs/bench/, docs/evidence/, and docs/models/ (RTDETR note stays in unsupported-boundaries).docs/tutorials_of_model_servers.md (+ zh-cn); README links updated.submit_and_wait; describe shared request-level timer + BatchCollector::submit snapshot semantics.batch_collector_unittest as sanitizer and build it (with worker_pool_unittest) in the TSAN CI gate / tests-only-tsan preset.tests-only* build presets default to target check (so EXCLUDE_FROM_ALL unit tests are built); docs drop bare ctest after preset build. TSAN preset builds sanitizer-labeled binaries.check_consistency convert/CI contract aligned with SME-16 (require unconditional --skipInference, refuse -ge 9 / --buildOnly, require -lt 10 gate and CI TRT_VERSION_MAJOR=8 failure assert).convert_trt_engines.sh refuses TensorRT major < 10 and drops the TRT 8 --workspace/--buildOnly path; TRT_VERSION_MAJOR only overrides failed banner probes (still fail-closed if < 10). Dry-run no longer silently assumes 10 when version cannot be detected.web_console / ServerManager claims from CMakeLists.txt, test/ready_probe_unittest.cc, and scripts/clean_artifacts.sh (those apps are gone; _bin/_lib and ready_probe belong to supervisor/control).ORT_API_VERSION==29 and refuse leftover libonnxruntime.so.1.18*; FATAL fix lines are profile-aware (--cpu --all / --all) and match install_deps --check. README / shortest-path-cpu / deployment §7 document --check before cmake.deadline (same clock as the HTTP timer): expired entries TIMEOUT without checkout; checkout wait is min(config, min remaining deadline); re-check after pool wait. In-flight run_batch is not cancelled mid-run.write_slot uses lock-free claim→write→publish (no dual writers on outputs); submit/stop use _accepting+epoch recheck with TIMEOUT fixup (no new mutex). Concurrent unit tests cover same-index races and submit/stop TOCTOU.serve_process authenticates before per-IP rate_limit_qps; /healthz /ready /openapi.json are exempt. Unauthenticated callers get 401 (not 429) and do not burn the IP budget.POST /api/v1/servers/{id}/start|stop|restart returns HTTP 500 (with ok:false) on failure instead of always 200; aligns with graceful_restart. UI treats non-2xx as failure.[supervisor].pack_file from mortred.toml via ControlConfig::apply_pack when MORTRED_PACK is unset (resolve_pack_path: env/CLI wins). Docs/comments matched behavior.scripts/release_dry_run.sh locally verifies GHCR IMAGE lowercase, tarball+basename .sha256, and bootstrap mismatch refuse; check_consistency gates release.yml lowercase IMAGE construction.scripts/check_consistency.py gates the ci.yml mock-trtexec dry-run contract against convert_trt_engines.sh (require --skipInference, reject --buildOnly, TensorRT 10.x mock banner).pack_trt.py --self-test isolates MORTRED_PROFILE (and uses profile="any") so a live CPU shell cannot fail hermetic consistency.install_deps.sh toml++ one-shot install: timed curl into a stage file, reject
truncated headers (<17000 lines), always ensure TOML_EXCEPTIONS=0, and do not
stamp a partial download. SME-01 archives Linux tests-only 53/53 at d51b0287
under docs/evidence/ plus the living docs/sme-oob-checklist.md.
install_deps.sh的 toml++ 改为限时下载到暂存文件、拒绝残片(<17000 行)、 始终保证TOML_EXCEPTIONS=0,不对半截下载盖 stamp。SME-01 在docs/evidence/归档了 Linux tests-only 53/53(d51b0287),并加入活清单docs/sme-oob-checklist.md。
github.repository for GHCR image paths so
ghcr.io/maybeshewill-cv/mortred_model_server matches docs (uppercase owner
refs are illegal). Bootstrap refuses install when a published .sha256 does
not match; missing checksum still WARN-and-continue.release 工作流把
github.repository转成小写再推 GHCR,与文档中的ghcr.io/maybeshewill-cv/mortred_model_server一致(大写 owner 非法)。 bootstrap 在已发布的.sha256与包内容不匹配时拒绝安装;缺少校验文件仍 WARN 后继续。
install_deps.sh --workflow also installs vendored libssl/libcrypto (same
path as --all). CMake requires vendored::crypto whenever workflow is
present; sanitizer CI only ran --workflow and failed at configure.
--check now asserts libcrypto.so.3 or libcrypto.so.1.1.
install_deps.sh --workflow也会安装 vendored 的libssl/libcrypto(与--all同路径)。有 workflow 时 CMake 硬要vendored::crypto;sanitizer CI 只跑--workflow会在 configure 阶段挂。--check现在断言libcrypto.so.3或libcrypto.so.1.1。
install_deps.sh binds ORT headers to the pinned library version: wipe/reinstall
replaces 3rd_party/include/onnxruntime atomically, stamp_fresh invalidates on
ORT_API_VERSION mismatch (and missing struct CUDAProviderOptions on gpu), and
--check fail-closes on the same gates. Stops local full-GPU builds from compiling
ort_session.cpp CUDA EP against leftover ORT 1.18 headers while linking 1.29.
install_deps.sh把 ORT 头与 pin 的库版本强绑定:wipe/重装会整目录替换3rd_party/include/onnxruntime;stamp_fresh在ORT_API_VERSION不匹配(gpu 下缺struct CUDAProviderOptions)时失效;--check用同一套门闩 fail-closed。 避免本地 full-GPU 用残留的 ORT 1.18 头编译ort_session.cppCUDA EP、却链接 1.29。
model_run_timeout without dropping completed
items. The HTTP series holds a unique-reply counter; per-item compute is a
detached go chain plus one request-level timer. results[] is always length
N: 504 / status 4 when nothing published; 200 / status 68 / partial=true
when at least one item finished. Batch uses async BatchCollector::submit
with the same request-level timer race as the non-batch path: reply
takes an acquire snapshot of published slots (unpublished → TIMEOUT);
late write_slot completions are dropped after the unique reply.同步推理在
model_run_timeout到期时回包,不再把已完成项丢掉。HTTP series 上是唯一回复的 counter;逐项计算是脱离 series 的 go 链,外加一个请求级 timer。results[]长度恒为 N:没有任何项发布时 504 / status 4;至少一项 完成时 200 / status 68 /partial=true。凑批走异步BatchCollector::submit, 与非凑批共用请求级 timer:回包对已发布 slot 做 acquire 快照(未发布 → TIMEOUT); 唯一回复之后迟到的write_slot会被丢弃。
sha256sum -c after
curl -fLO prints OK (deployment §6.1). sha256sum /abs/path wrote a
runner path that does not exist on the operator machine. bootstrap.sh
downloads mortred_model_server-<tag>-<profile>-linux-x64.tar.gz after
resolving the latest GitHub release tag (there is no ...-latest-...
asset). GPU release staging copies into opt/mortred/ like the cpu packer
(cp -a src/. dest/, not cp -a src dest/opt/).发布
.sha256只写 tarball 文件名,curl -fLO后sha256sum -c才能 打出OK(deployment §6.1)。原先sha256sum /绝对路径会把 runner 路径写进 校验文件。bootstrap.sh先解析 latest tag,再下mortred_model_server-<tag>-<profile>-linux-x64.tar.gz(没有...-latest-...资产)。gpu staging 与 cpu 一样落到opt/mortred/(cp -a src/. dest/, 不是cp -a src dest/opt/)。
write_prometheus_credentials.sh chowns the scrape secret to the process
that reads it: uid 65534 (compose prom/prometheus nobody) or user
prometheus (apt unit, /etc/prometheus/*). Mode stays 600. Compose
monitoring scrapes host.docker.internal:8080 via prometheus.compose.yml
(localhost inside that container is Prometheus). deployment documents
both the ownership step and the Docker DNS name.
write_prometheus_credentials.sh把 scrape 密钥chown给真正读它的进程: compose 里prom/prometheus的 uid 65534,或 apt 单元的 prometheus 用户(/etc/prometheus/*)。权限仍是 600。compose 监控通过prometheus.compose.yml刮host.docker.internal:8080(容器内localhost是 Prometheus 自己)。deployment补了属主步骤和该 Docker 主机名。
install_deps wipe/reinstall removes CUDA 11 / TensorRT 8 /
ORT 1.18 libs and mismatched ORT headers; --check requires
ORT_API_VERSION to match the pin. --cuda-version 11 fails. Engines must
be rebuilt with matching trtexec. Full GPU CMake refuses AddressSanitizer
(ASan + TRT/cudart mixed ELF).GPU 线改为 CUDA 12 + TensorRT 10.3 + cuDNN 9 + MNN 3.6.1 + ORT 1.29 cuda12。
install_deps的 wipe/重装会清掉 CUDA 11 / TensorRT 8 / ORT 1.18 库以及 与 pin 不一致的 ORT 头;--check校验ORT_API_VERSION。--cuda-version 11失败。引擎必须用匹配的 trtexec 重建。full GPU CMake 拒绝 AddressSanitizer。
mortred_model_server-<version>-<profile>-linux-x64.tar.gz. Release does not
publish a ...-latest-... tarball name (GHCR :latest-cpu / :latest-gpu
stay image tags). README / deployment / bootstrap.sh no longer claim that
path always installs: no Release yet prints WARN and the source-build track.
verify_deployment.sh --basic rejects packing or fetching
mortred_model_server-latest-.入口一无 Docker 按 GitHub latest tag 下
mortred_model_server-<version>-<profile>-linux-x64.tar.gz。Release 不挂...-latest-...tarball 文件名(GHCR:latest-cpu/:latest-gpu仍是镜像 tag)。README / deployment /bootstrap.sh不再写成「无 Docker 一定能装上」: 还没有 Release 时 WARN 并落到源码构建。verify_deployment.sh --basic禁止 打包或下载mortred_model_server-latest-。
min_box_area_px in decoder network
(letterboxed) pixels before letterbox unmap. v5/v6 previously filtered
source pixels after unmap; v8 ignored the key. Synthetic decode tests
cover the shared threshold.YOLOv5 / v6 / v7 / v8 的
min_box_area_px统一在解码坐标系(letterbox 网络像素)里、unmap 之前过滤。原先 v5/v6 在 unmap 后的源图像像素上比 面积,v8 完全不读该键。合成解码单测覆盖这条共用阈值。
check_trt_engine_manifest walks HTTP catalog [*.backend] type="tensorrt"
paths (not the removed *_TRT tables) against conf/trt_engines.json.
Scaffold configs (TODO(new_model) / not in the HTTP catalog) are skipped.
fetch_weights --profile cpu no longer includes the off-HF hrnet ONNX that
the cpu catalog does not serve.checker 按 HTTP catalog 的
[*.backend] type="tensorrt"对照trt_engines.json(不再找已删除的*_TRT)。脚手架(TODO(new_model)/ 不在 HTTP catalog)跳过。cpu 权重集合不再含 catalog 不服务、且不在 HF 上的 hrnet ONNX。
install.sh, opt/mortred/,
and deploy/ sit at the archive root (make_release_tarball.sh and the gpu
release job already packed that way). README / deployment / installer comments
unpack into an empty directory instead of cd into a wrapper the packer does
not create.发布 tarball 按平铺写进文档:包根就是
install.sh、opt/mortred/、deploy/(打包脚本和 gpu release job 本来就是这样打的)。README / 部署 / 安装脚本注释改为解到空目录,不再cd一个打包器不会生成的 wrapper 目录。
mortredctl init-trust
/ supervisor.env). Compose already required MORTRED_METRICS_TOKEN; the
copied quick-start commands now match. The optional monitoring compose
bind-mounts a scrape credentials file (not minted by the app compose or
the release tarball) and provisions the Grafana Prometheus datasource.入口二/三与 bootstrap 改为三个 token(
init-trust/supervisor.env), 与 compose 已有的:?scrape 必填对齐。可选监控 compose 挂载 scrape credentials 文件(不由应用 compose 或发布 tarball 发明 token),并 provision Grafana 的 Prometheus 数据源。
images[] /
{status, status_str, task_id, results[]}). api-contract, api-keys, and
the task tutorials no longer present img_data or {code, msg, data} as
success examples. Gateway /metrics docs require MORTRED_METRICS_TOKEN
on loopback in both languages. README Model Zoo splits HTTP catalog vs
bench-only.人读文档与统一信封对齐:请求
images[],响应{status, status_str, task_id, results[]}。api-contract、api-keys与任务教程 不再把img_data或{code, msg, data}写成成功示例。中英监控/部署均写明 网关/metrics含环回也要MORTRED_METRICS_TOKEN。README Model Zoo 区分 HTTP catalog 与 bench-only。
ci_container_boot.sh asserts the unified infer envelope (status /
results[], including results[0].status) instead of the removed code
field. payload.get("code", 0) always passed, so a HTTP 200 with
status != 0 would still print success.
ci_container_boot.sh按统一信封断言推理结果(status/results[],含results[0].status),不再读已删除的code。payload.get("code", 0)会永远 通过,HTTP 200 且status != 0仍会打印成功。
init (sample_size / default steps / channels). A missing or
empty sample_size fails init instead of serving MODEL_EMPTY_INPUT_IMAGE
on every request. The async smoke script and DDPM examples use
params.timesteps (not a root timestep) and the unified result envelope.
Few-step CPU generate proof: model_golden.ddpm_celeba_hq_fewstep with
conf/ci/ddpm_onnx_fewstep.toml (not in hosted; ONNX is ~143MiB). The case
checks a 128x128 PNG from the HTTP adapter; it does not pin pixels
(random_device still seeds each run). The async smoke script sets
MORTRED_AUTH_TOKEN and sends Authorization: Bearer on /jobs
(empty token is 401; /healthz stays public).HTTP 扩散工人在
init时从模型 TOML 种入采样模板(sample_size/ 默认步数 / channels)。缺或空sample_size会 init 失败,而不再每请求MODEL_EMPTY_INPUT_IMAGE。异步冒烟与 DDPM 示例改用params.timesteps(不是根上timestep)和统一结果信封。少步 CPU 出图证明:model_golden.ddpm_celeba_hq_fewstep+conf/ci/ddpm_onnx_fewstep.toml(不进 hosted;ONNX 约 143MiB)。该用例断言 HTTP adapter 走出 128x128 PNG, 不钉像素(每次random_device仍重新播种)。异步冒烟会设MORTRED_AUTH_TOKEN并在/jobs带Authorization: Bearer(空 token 是 401;/healthz仍公开)。
mortred-gateway.out is on disk) and
refuses to spawn __gateway unless MORTRED_METRICS_TOKEN is set and
distinct from the inference and management tokens — the same rule the
gateway already enforced. Missing or colliding scrape secrets are a
permanent failure, not a restart backoff, so a healthy :8787 banner can
no longer hide a gateway crash loop. /api/v1/keys remains 404.Supervisor 在磁盘上已有
mortred-gateway.out时,若 scrape token 未设置或与 推理/管理 token 相同则拒绝 listen;spawn__gateway使用同一谓词,失败记为 permanent 而非 backoff,避免「健康横幅 + 网关崩溃循环」。/api/v1/keys仍是 404。
BaseAiModel::run (and packed BackendCvModel::run_batch) catch throws from
run_impl / OpenCV and return MODEL_RUN_SESSION_FAILED, so a workflow go
thread does not std::terminate the process. Packed run_batch also
broadcasts that code onto every item_status after a throw (run_image_batch
pre-assigns OK; HTTP would otherwise report success). Diffusion DDPM/DDIM/cls-cond
postprocess no longer convertTo(CV_8UC3) + COLOR_RGB2BGR on 1/4-channel
tensors (LDM latent channels=4 with save_raw_output); display conversion
is channel-correct and skipped when only raw latents are needed.
BaseAiModel::run(以及 packedBackendCvModel::run_batch)接住run_impl/ OpenCV 的抛出并返回MODEL_RUN_SESSION_FAILED,避免 workflow go 线程std::terminate整进程。packedrun_batch在 catch 后还会把该错误码盖到 每一个item_status(run_image_batch会先全部标 OK;不盖的话 HTTP 会谎报成功)。 扩散 DDPM/DDIM/cls-cond 后处理不再对 1/4 通道convertTo(CV_8UC3)+COLOR_RGB2BGR(LDM latentchannels=4且save_raw_output);出图按通道转换,只要 raw 时跳过。
do_work writes inference metrics and the run-time EWMA before returning
the worker to the queue (same order as process_batch). The destructor
drain treats an enqueued worker as permission to destroy those members;
writing them after enqueue raced a timed-out request’s destructor.
do_work在把 worker 还回队列之前写入推理 metrics 和运行时间 EWMA (与process_batch相同)。析构 drain 把「worker 已入队」当作可以销毁 这些成员;enqueue 后再写会与超时请求的析构竞态。
GeometryScale remains for
NanoDet / CenterFace / LibFace. Fork CI proves YOLOv8 decode+NMS+unmap
via a CI-only ONNX overlay (conf/ci/yolov8_onnx_hosted.toml,
yolov8s.onnx); product yolov8_config.toml is still TensorRT.YOLO v5/v6/v7/v8 预处理改为 Ultralytics 导出同款中心 letterbox(等比、 填充 114),框反变换用同一套 pad/scale。NanoDet / CenterFace / LibFace 仍走拉伸
GeometryScale。Fork CI 用仅 CI 的 ONNX overlay (conf/ci/yolov8_onnx_hosted.toml,yolov8s.onnx)证明 YOLOv8 decode+NMS+unmap;出厂yolov8_config.toml仍是 TensorRT。
WFServerBase::stop() is
already shutdown() + wait_finish() (blocking); the old teardown called
wait_finish() a second time and blocked forever. In production this was
masked by systemd’s TimeoutStopSec SIGKILL - mortred-supervisor never
actually exited gracefully. Found by the in-process SupervisorApp teardown
of this refactor.Supervisor 优雅关停不再挂死。
WFServerBase::stop()本身就是阻塞的shutdown() + wait_finish();旧代码又补了一次wait_finish(),第二次 永久阻塞。生产上一直被 systemdTimeoutStopSec的 SIGKILL 掩盖——mortred-supervisor此前从未真正优雅退出过。由本次重构的进程内 SupervisorApp 关停路径暴露。
GatewayApp /
SupervisorApp own their state (catalog, config, tokens, api keys,
metrics, supervisor); run() maps the process environment onto an
explicit *InitOptions, and the app objects live in a new
workflow-bound control_workflow library so tests link them in-process.
Acceptance: two gateway and two supervisor instances with distinct
roots/tokens/catalogs serve side by side in one process
(gateway_multiinstance_test, supervisor_multiinstance_test).控制面去全局化(内部重构,无行为变化):网关/supervisor 的文件级全局 状态移除,
GatewayApp/SupervisorApp持有各自状态;run()将进程 环境映射为显式的*InitOptions,app 对象移入新的依赖 workflow 的control_workflow库以便测试进程内链接。验收:两个网关与两个 supervisor 实例(不同 root/token/catalog)可同进程并存 (gateway_multiinstance_test、supervisor_multiinstance_test)。
X-Mortred-Params,
X-Mortred-Options and X-Request-ID to the model server, and echoes
X-Request-ID on its own error replies. Raw-body requests previously lost
their params/options and client correlation when proxied through the
gateway (JSON-envelope requests were unaffected). All other client headers
stay dropped - the default-deny forward list is the header-injection guard.网关现在向模型服务器转发 raw-body 控制头
X-Mortred-Params/X-Mortred-Options/X-Request-ID,并在自身错误回复中回显X-Request-ID。此前 raw-body 请求经网关代理后参数与关联 id 被静默丢弃 (JSON envelope 路径不受影响)。其余客户端头仍然丢弃——默认拒绝的转发 白名单就是防头注入的屏障。
process() folds a null request method to "" like the
gateway/model-server guards. A malformed request line (workflow leaves
get_method() null) used to construct std::string from nullptr - UB,
typically a crash of the management plane.Supervisor 的
process()现在与网关/模型服务器一样把空 method 折叠为""。此前畸形请求行(workflow 返回 null method)会从 nullptr 构造std::string——未定义行为,通常表现为管理面崩溃。
mortred_http_requests_total: the 202
submit reply, 200 status/wait/result replies and the async error codes
(404/405/409/429) were previously invisible on the model-server dashboards.
Async replies deliberately carry no mortred_http_request_duration_ms
sample - that histogram observes inference time and the async HTTP path
has none.异步任务回复现在计入
mortred_http_requests_total:此前 202 提交回复、 200 状态/等待/结果回复以及异步错误码(404/405/409/429)在模型服务器监控 上不可见。异步回复刻意不产生mortred_http_request_duration_ms样本——该 直方图观测的是推理耗时,异步 HTTP 路径上没有这一耗时。
ProcessStop latch (block SIGINT/SIGTERM,
sigwait thread, idempotent WaitGroup::done()) so those signals reach
server->stop() and the impl destructor drain instead of default-killing
the process. Supervisor already had this path; kStopGraceMs SIGKILL
remains the hung-model backstop.模型进程与网关的
main在 SIGINT/SIGTERM 时会done()WaitGroup,从而走到stop()和析构 drain;不再被默认信号直接杀死。Supervisor 本身原本就是这样。
POST /jobs flushes HTTP 202 at admission. The runner is a detached
Workflow go task (go->start()), so submit no longer waits for
run_items. GET /jobs/{id}/wait hangs the HTTP series on a named
counter and wakes on a terminal job or the wait budget (milliseconds),
not on pending→running. Serial submits can now observe 429 while a
job is still running. Customer steps:
async-jobs-customer-test.md.
POST /jobs在准入时立即返回 202,不再等推理结束。GET …/wait在终态或 wait 预算耗尽时返回(单位毫秒)。客户逐步验收见 async-jobs-customer-test.zh-cn.md。
GET /api/v1/keys and POST /api/v1/keys/reload. Those routes
mutated a copy of ApiKeyManager that never authenticated inference
traffic. Edit conf/api_keys.toml then restart the gateway child
(POST /api/v1/servers/__gateway/restart). scope=admin still does not
unlock :8787; management stays MORTRED_API_TOKEN only.ApiKeyManager::load treats a readable empty or comment-only key file as
success with key_count()==0 (not a parse failure). The gateway still
refuses to start with no static token and no keys; with a static token it
logs a warning and uses that token only.RestartEngine::Decision the same way as child exits, so backoff is
scheduled instead of leaving wanted true with a dead pid.kAutostartReadyTimeoutMs (never referenced; 10-minute
kill is still out of scope).*_cpu_config.toml duplicates. Git
defaults backend.device to gpu (including omitted keys). The value is
cpu or gpu (cuda is rejected). type=tensorrt with device=cpu
with device=cpu fails at parse, session create, and supervisor spawn.
CPU catalog entries point at the same files; operators who need CPU inference
set device = "cpu" themselves. YOLOv8 and HRNet are not in the cpu catalog
(TensorRT).docker_entrypoint.sh now has a bash shebang so
exec-form ENTRYPOINT can start; compose --profile gpu builds
target: mortred-gpu (the Dockerfile last stage is that alias, so
docker build . is the GPU runtime again, not mortred-cpu).SupervisorTest.spawn_passes_model_flag_for_unified_exe). The gate
only refuse-spawns after it can read type=tensorrt and the engine file is
missing or empty.Cannot start server when :9056 is already bound (supervisor still
serving the pack). Probe now binds 127.0.0.1 like the supervisor, refuses
a busy port before spawn, and logs host:port/errno on listen failure.
Bind failure is start_failed, not OOM.memory.used as the model’s footprint.
Per-model numbers come from NVML compute-apps (pid, else unique
mortred-model-server name). If WSL has no process row, gpu_mem_mib_* is
the delta vs device used sampled before that spawn (gpu_mem_source=
device_delta), not the card total. Joint residency does not sum deltas.127.0.0.1:0 on Python 3.10: repo_toml ignored unquoted
port=9002. Fallback parser now reads integers; calibrate takes port/uri
from the same conf/server file used to spawn.convert_trt_engines.sh retries without min/opt/maxShapes when TensorRT
reports a static ONNX (Static model does not take explicit shapes). The
yolov8 profile is for dynamic batch; some weight drops are fixed 1x3x640x640.prepare_pack.sh runs the /ready probe with cwd = _bin/bin, matching
supervisor spawn, so model_config_file_path = "../conf/..." resolves when
the script is invoked from the repo root.prepare_pack.sh stops the /ready probe with SIGINT (same as the
supervisor) instead of SIGTERM, so glog does not dump a failure stack.
Probe logs go to logs/prepare-<id>.log; a failed ready prints that file.mortred-supervisor can init the full GPU catalog (pack autostart still
loads every profile-matching conf/server file).ghcr.io/...:vX.Y.Z-gpu and :latest-gpu from the
existing gpu-tarball Docker compile (cpu tags stay in the images job).container boot (cpu compose): compose up → supervisor
health → gateway /healthz → one MOBILENETV2 infer → supervisor/gateway
process liveness. Infer uses a CI-only device=cpu pack overlay (git
model tomls stay device=gpu). Expensive rebuild is path-filtered; GPU
compose is not claimed on GitHub-hosted runners.mortredctl init-trust writes gitignored
conf/local/trust.env (inference / management / scrape / internal). Gateway
and supervisor refuse to start without their tokens, including on loopback.
GET /metrics is never public. Wildcard bind requires MORTRED_EXPOSE=docker
(containers) or unsafe. Nginx is the supported TLS edge
(mortredctl init-edge --mode lan|acme|files, deploy/nginx, compose
profile edge on Linux host network). Caddy is removed.gpu_mem_limit defaults to 2048 MiB per session
(gpu_mem_limit_mb / MORTRED_ORT_GPU_MEM_LIMIT_MB; 0 = unlimited).
The previous gpu_mem_limit = 0 let the CUDA EP arena grow without bound.mortredctl
prepare for pack TensorRT engines, mortredctl calibrate / --write-pack
for worker_nums on the pack file (conf/server stays 1).scripts/calibrate_pack.py,
mortredctl calibrate): sweep w, HTTP RPS via http_infer_rps.py, per-process
GPU occupancy (NVML pid/name, else pre-spawn device delta), suggested w*,
optional joint residency. --write-pack updates [pack.<ID>] worker_nums
in that pack file only. Does not write conf/server (git copies stay
worker_nums=1).scripts/prepare_pack.sh, mortredctl prepare):
convert only engines used by MORTRED_PACK, refuse spawn if a file is missing
or empty (no crash-loop), optional /ready at worker_nums=1.
MORTRED_AUTO_BUILD_ENGINES still converts the whole zoo and stays opt-in.
doctor --strict fails when pack TRT files are missing.conf/packs/demo.toml, MORTRED_PACK): listed
catalog ids boot; MORTRED_AUTOSTART=true no longer starts the whole zoo.
Pack worker_nums / model_config override the child via env without
rewriting conf/server (still worker_nums=1).scripts/server/http_infer_rps.py): keep-alive
workers, pre-encoded envelope, serving RPS + latency percentiles, optional --qps,
JSON report. test_server.py --mode load wraps catalog/gateway URLs. No locust
or requests.mortredctl doctor --strict: fail the doctor when security warnings fire
(non-loopback plaintext HTTP, short tokens, identical tokens). Default
doctor still warns only.cpu-profile fail-closes a multi-family MNN CPU golden set from
conf/ci_hosted_golden.json (classification, NanoDet, DBNet, SuperPoint,
BiSeNetV2): sha256-locked HF fetch, MORTRED_CI_REQUIRE_WEIGHTS, XML
skipped=0. GPU smoke and TensorRT are still maintainer-only. Nightly
remaining goldens write a skip-inventory artifact. HTTP catalog ids must
declare a CI tier (hosted / gpu-smoke / nightly).POST /v1/models/{id}/infer and /v1/models/{id}/jobs* to
the model’s loopback port. Job Location / poll_url / result_url are
rewritten onto that prefix. The legacy {server_uri} POST path still works.scripts/server/locust_performance.py (breaking for anyone
who invoked --mode locust). Use --mode load / http_infer_rps.py./api/v1/infer, /api/v1/jobs*, and /api/v1/pipelines*
(breaking). Inference and async jobs go through the gateway; there is no
server-side pipeline on the supervisor. Those paths now return the
management {ok, error} 404. Graceful restart still drains by reading the
model’s mortred_async_queue_depth gauge.create_*_sampler factories, dead CvUtils overlay/base64/tensor-copy
helpers, unused std_clip_* / std_sam_prompt_input aliases,
build_unified_response_body, handle_custom_endpoint,
FilePathUtil::is_dir_exist, Timestamp::to_str / invalid,
detection_params_parse (inlined into DetectionParams::parse),
TypeErasedFactory::register_type and the ModelFactory alias.json_request_parser.h (parse_json_request had no callers; it still
accepted img_data and ignored unknown keys).http_response.h ({req_id, code, msg, data} shim). Process-level JSON
now uses UnifiedResponse from response_envelope.h.create_server wrappers (every caller already used
cv_catalog::create_server).DetectionGeometryScale,
make_detection_geometry_scale, scale_detection_bbox /
scale_detection_point, validated_f32_output).task_request / go_result aliases for InferenceTask /
InferenceResult./welcome and /hello_world (breaking). Unknown paths
now answer 404 with process-level UnifiedResponse.MORTRED_HAS_GPU_RUNNER=true;
otherwise the job is skipped so CI does not wait for a missing runner. Smoke
engine refresh is convert_trt_engines.sh --only yolov8 only. Fork PRs never
run on the self-hosted GPU runner. Require the inference paths check, not
gpu golden smoke by name. See docs/ci-golden-regression.md.:8080 and supervisor :8787 on
127.0.0.1 only. The monitoring stack binds Grafana/Prometheus to loopback,
requires GRAFANA_ADMIN_PASSWORD, and no longer enables Prometheus
--web.enable-lifecycle. Default Prometheus scrape is gateway /metrics
only; supervisor (Bearer) and model (loopback) jobs are commented with the
real auth and topology constraints.mortredctl infer POST the data-plane envelope to
the gateway (/v1/models/{id}/infer on :8080) instead of /api/v1/infer.
The gateway accepts MORTRED_API_TOKEN and API keys with admin (or all /
inference) scope, and reflects CORS for the supervisor UI origin.
common/response_envelope.h (encode/decode + field names). Data-plane
binding is server/parsed_request.h; in-process execution types moved from
async_job_table.h to server/inference_task.h. Supervisor/CLI go through
the codec instead of hand-rolled JSON./healthz,
/ready, and 401/404/405/413/415/429 exits emit {status, status_str,
task_id, results:[]} instead of {req_id, code, msg, data}. HTTP status
codes and StatusCode wire integers are unchanged./api/v1/infer /jobs*
/pipelines* failures before upstream now emit {status, status_str,
results:[], errors[]}. HTTP status codes are unchanged. Management APIs
(/servers*, start/stop, logs, supervisor 401/405) still use {ok, error}.{"req_id", "images": [<base64>...], "params", "options"}
and answer with {status, status_str, task_id, model, results[], server_time_ms,
partial}. results[] is index-aligned with images[] and every item carries
its own status (per-item failure isolation, deadline partials). The legacy
img_data field was removed and answers 422 with a JSON-pointer migration
hint; unknown fields/params are rejected strictly (422 + errors[]).score_threshold,
nms_threshold, top_k (validated per model, TOML config stays the default).Retry-After and queue depth accounting are now per image
item (max_request_items, default 16); one deadline spans queue wait,
worker wait and inference.Content-Type/Accept verbatim (binary body
encoding groundwork).custom_drivers.cpp; bench-only
product rows compile only into the benchmark target (MORTRED_WITH_CUSTOM_DRIVERS).control/http_reply.h for JSON replies
(error JSON shapes are unchanged).create_* / make_server_worker / catalog make_model /
CvWorkerFactory drop the unused name argument.common no longer links OpenCV (cv_utils is header-only)./api/v1/infer, /api/v1/jobs and pipelines now speak
the unified images[] / results[].data envelope. They previously still
sent the removed img_data field and read the legacy {data: ...}
response, so the built-in test proxy and pipelines could not succeed
against current model servers.contract_dump (C++ catalogs as the single
source) -> docs/contract_dump.json -> gen_openapi.py -> OpenAPI +
embedded /openapi.json; scripts/check_contract_sync.py gates the chain
in CI (any spec change without regeneration fails the build).cpu | gpu): one switch drives the build
(MORTRED_BUILD_PROFILE), the dependency set (install_deps.sh --cpu),
the model catalog (per-server profile field + MORTRED_PROFILE) and the
weight subset (fetch_weights.py --profile). The cpu profile compiles
TensorRT out entirely and ships a curated model set (mobilenetv2, resnet50,
yolov8, hrnet).mortred-cpu Docker target + docker compose
--profile cpu|gpu, and versioned binary tarballs
(make_release_tarball.sh + in-tarball install.sh with systemd wiring).curl | bash bootstrap, docker
compose, and mortredctl init / doctor / upgrade.MORTRED_AUTO_BUILD_ENGINES=true opt-in engine conversion at container
start (gpu profile).--version), this changelog, and a tag-driven release
pipeline building both images and both tarballs.mortred-gateway link failure against vendored OpenSSL after the P0-2
rework (no CI path compiled the gateway; the new cpu-profile job now does).中文摘要:新增部署 Profile 体系(cpu/gpu 单一事实源贯穿构建、依赖、目录、 权重四层)、双轨分发(Docker 双 target + compose profiles;版本化 tarball
- systemd 安装器)、三入口共享 mortredctl 内核(bootstrap / compose / init-doctor-upgrade)、可选首启 engine 转换、项目版本化与发布流水线; 修复 gateway 对 vendored OpenSSL 的链接缺口。