Skip to content

Feat/new model manager - #2251

Draft
grzegorz-roboflow wants to merge 86 commits into
mainfrom
feat/new-model-manager
Draft

grzegorz-roboflow wants to merge 86 commits into
mainfrom
feat/new-model-manager

Conversation

@grzegorz-roboflow

Copy link
Copy Markdown
Collaborator

What does this PR do?

Adding inference_model_manager

Type of Change

  • New feature (non-breaking change that adds functionality)

Testing

Initial suit of tests added

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code where necessary, particularly in hard-to-understand areas
  • My changes generate no new warnings or errors
  • I have updated the documentation accordingly (if applicable)

Additional Context

N/A

Comment thread inference_server/inference_server/app.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/app.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/server.py Fixed
Comment thread inference_server/inference_server/auth.py Fixed
Comment thread inference_models/inference_models/models/base/classification.py Outdated
Comment thread inference_models/inference_models/models/auto_loaders/core.py Outdated
Comment thread inference_models/inference_models/models/auto_loaders/core.py Outdated
Comment thread inference_models/inference_models/models/base/documents_parsing.py Outdated
Comment thread inference_models/inference_models/models/moondream2/moondream2_hf.py Outdated
Comment thread inference_server/inference_server/auth.py Fixed
Comment thread inference_server/inference_server/errors.py Fixed
Comment thread inference_server/inference_server/routers/infer.py Fixed

@dkosowski87 dkosowski87 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1/6 with the pace I'm currently at ;)

Comment thread inference_model_manager/inference_model_manager/backends/direct.py Outdated
Comment thread inference_model_manager/inference_model_manager/backends/direct.py Outdated
Comment thread inference_model_manager/inference_model_manager/backends/direct.py Outdated
Comment thread inference_model_manager/inference_model_manager/backends/direct.py
Comment thread inference_model_manager/inference_model_manager/backends/direct.py Outdated
Comment thread inference_model_manager/inference_model_manager/backends/direct.py Outdated
Comment thread inference_model_manager/inference_model_manager/model_manager.py
Comment thread inference_model_manager/inference_model_manager/model_manager.py Outdated
Comment thread inference_server/inference_server/errors.py Fixed
Comment thread inference_server/inference_server/routers/infer.py Fixed
@socket-security

socket-security Bot commented Jun 22, 2026 •

Copy link
Copy Markdown

Warning

Review the following alerts detected in dependencies.

According to your organization's Security Policy, it is recommended to resolve "Warn" alerts. Learn more about Socket for GitHub.

Priority Alert  (click "▶" to expand/collapse) Action
Medium priority
Obfuscated code: pypi openexr is 72.0% likely obfuscated

Confidence: 0.72

Location: Package overview

From: requirements/requirements.sam3_3d.txt → pypi/openexr@3.5.2

ℹ Read more on: This package | This alert | What is obfuscated code?

Next steps: Take a moment to review the security alert above. Review the linked package source code to understand the potential risk. Ensure the package is not malicious before proceeding. If you're unsure how to proceed, reach out to your security team or ask the Socket team for help at support@socket.dev.

Suggestion: Packages should not obfuscate their code. Consider not using packages with obfuscated code.

Mark the package as acceptable risk. To ignore this alert only in this pull request, reply with the comment @SocketSecurity ignore pypi/openexr@3.5.2. You can also ignore all packages with @SocketSecurity ignore-all. To ignore an alert for all future pull requests, use Socket's Dashboard to change the triage state of this alert.

Warn

View full report

@grzegorz-roboflow
grzegorz-roboflow force-pushed the feat/new-model-manager branch from fa5f76a to 4c4a435 Compare July 8, 2026 10:24
@grzegorz-roboflow
grzegorz-roboflow force-pushed the feat/new-model-manager branch 7 times, most recently from b9a09b5 to 60f7a91 Compare July 28, 2026 15:24
- shared fetcher refuses with 403 before any DNS work
- covers every V2 handler parser
- published rc1 lacks the mqtt host configuration flag
PawelPeczek-Roboflow added a commit that referenced this pull request Sep 25, 2026
…h-published rcs (#3068)

`inference-models 0.39.0rc3` was published to PyPI today (10:48 UTC) from
the feat/new-model-manager branch (PR #2251). main's requirement files
pinned `~=0.39.0rc2`, and a pre-release lower bound lets the resolver take
any newer rc, so every image built since installs rc3. rc3 changes the
SAM3 output contract (default `mask_format="rle"`, dict results) and
main's adapters in inference/models/sam3/ still expect tensors, which is
why "Code Quality & Regression Tests - NVIDIA T4" fails in the SAM3 step
with HTTP 500 (run 36145992322).

Exact-pin all five requirement files to 0.39.0rc2, the version the last
green image installed. No code change; the pin is lifted by the PR that
adopts the new inference-models contract.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
PawelPeczek-Roboflow added a commit that referenced this pull request Sep 29, 2026
…ing dormancy

Replace the path-filter dormancy check with the plain fix: remove
unit_tests_inference_model_manager_and_server.yml from TEST_WORKFLOWS, with a
comment saying why and when to put it back (once inference_model_manager/ and
inference_server/ land on main, #2251). The status-sync listener drops the
same workflow name so the two lists stay aligned.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
PawelPeczek-Roboflow added a commit that referenced this pull request Sep 29, 2026
… after re-runs (#3077)

* Fast-track CI: skip workflows that are dormant on the branch

The orchestrator dispatched unit_tests_inference_model_manager_and_server.yml
on every fast-track branch. That workflow is path-filtered to
inference_model_manager/** and inference_server/**, which exist only on
feat/new-model-manager, so it never runs on main, but workflow_dispatch
ignores path filters and the run died at install with "Distribution not
found at: .../inference_model_manager" (3 of 3 fast-track dispatches).

Before dispatching, read each workflow's pull_request path filter as it is
on the branch and skip the workflow when every literal root of the filter
is absent there: a PR to main would never have started it. The skip is
reported as a success status and a summary row naming the absent roots.
Workflows without a path filter, with a root-level pattern, or with at
least one root present are dispatched as before.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Fast-track CI: refresh commit statuses after a re-run

The orchestrator writes each fast-track-ci / <workflow> status once, from
the attempt it waited for. A maintainer re-run of a failed suite completes
outside the orchestrator, so a suite that went green on re-run kept a red
status on the umbrella PR (integration_tests_workflows_x86 on #3076).

Add a workflow_run listener on the same test workflows that, for a re-run
attempt (run_attempt > 1) of a workflow_dispatch run on a fast-track branch,
rewrites the status from that attempt. It only touches a status the
orchestrator created and never checks out branch code.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Fast-track CI: drop the dormant model-manager suite instead of detecting dormancy

Replace the path-filter dormancy check with the plain fix: remove
unit_tests_inference_model_manager_and_server.yml from TEST_WORKFLOWS, with a
comment saying why and when to put it back (once inference_model_manager/ and
inference_server/ land on main, #2251). The status-sync listener drops the
same workflow name so the two lists stay aligned.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* Apply EXIF orientation (JPEG, PNG, TIFF) in the new server decoder

The previous server decoded images with cv2.imdecode(IMREAD_COLOR), which
honours the EXIF Orientation tag. The imagecodecs decoder ignores it, so a
phone photo stored rotated with Orientation 3/6/8 reached the model
sideways or upside down, gave different detections, and the legacy routes
reported width and height swapped.

The decoder now reads the Orientation tag from the JPEG APP1 Exif block
(header walk only, no work when the tag is absent or 1) and rotates/flips
the decoded array before the BGR copy, so no extra copy is made.
image_dims uses the same tag reader and swaps width/height for
orientations 5-8, so the reported size matches the decoded pixels.
Output is pixel-identical to cv2.imdecode for all eight orientations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Serve the Roboflow-platform workflow blocks as core blocks on both servers

The new server listed 237 blocks in /workflows/blocks/describe against 246
on 1.7.2: dataset upload v1/v2, custom metadata, model monitoring, vision
events, vision event bundle, visual search, visual search classifier and
asset-library attributes were missing, so workflows using them did not
compile. They lived in `inference/roboflow_workflows_plugin`, inside the
legacy package the new server image does not install, and they imported
`inference.core.roboflow_api` and `inference.core.active_learning`.

- The blocks move to `core_steps/sinks/roboflow/` and
  `core_steps/integrations/roboflow/` of roboflow-workflows, where they
  lived before the plugin extraction, and `load_blocks()` registers them,
  with the `_tensor` siblings under ENABLE_TENSOR_DATA_REPRESENTATION.
  Identifiers, manifests and block_source are unchanged;
  fully_qualified_block_class_name is the historic
  `inference.core.workflows.core_steps...` path. The dataset-upload
  quota helpers move next to those blocks with the same cache keys and
  expiry.
- The plugin package, its loader and WORKFLOWS_PLUGINS entry are removed.
  `inference.roboflow_workflows_plugin` was never a public API and is not
  kept as an alias.
- The blocks' Roboflow API calls become members of the existing
  `workflows_core.platform_client` port, injected as a block init
  parameter - no process-global setup. The `inference` server client
  forwards them to `inference.core.roboflow_api` (same functions, retries
  and errors); the `inference_server` client implements the endpoints with
  its own headers and secure-gateway wrapping. The offline default refuses
  the API calls, as it refuses `post`.
- Tests inject fake clients instead of patching module functions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add model, engine and request-id response headers to the server

The legacy inference server (1.7.x) attaches X-Model-Id,
X-Model-Cold-Start, X-Model-Cold-Start-Count, x-inference-engine and
X-Request-ID to every response, plus X-Model-Load-Time and
X-Model-Load-Details on the request that loaded a model. The new server
sent none of them, so clients and dashboards reading them broke.

A raw ASGI middleware now installs a per-request collector in a
ContextVar and writes the headers on response start, with the same
names and value formats as 1.7.x (load time as str(float) seconds,
details as [{"m": id, "t": seconds}] JSON dropped above 4096 bytes,
comma-joined sorted model ids). The legacy bridge and the v2 dispatch
record the model ids they use; ModelManagerGateway records a load only
for the request that created the load future, so requests joining an
in-flight load are not reported as cold starts. X-Request-ID keeps a
valid incoming UUID4 and otherwise generates uuid4().hex, like the
asgi_correlation_id defaults used by 1.7.x. The model headers are
added to CORS expose_headers as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Serve Prometheus metrics at /metrics in the new server

The legacy server answers GET /metrics with Prometheus text metrics out of
the box (the CPU image included). The new server only has
/v2/server/metrics, which returns JSON and is off unless
ENABLE_CONTROL_PLANE_ROUTES=true, so scrapers pointed at /metrics got 404.

Add /metrics with the metric families the legacy server exposes:
- HTTP request metrics from prometheus-fastapi-instrumentator
  (http_requests_total, http_request_size_bytes, http_response_size_bytes,
  http_request_duration_seconds, http_request_duration_highr_seconds),
- process, platform and GC collectors (process_*, python_info,
  python_gc_*),
- the legacy per-model gauges num_inferences_<model>,
  avg_inference_time_<model>, num_errors_<model> and their *_total
  roll-ups, derived from the model manager's cumulative counters over a
  10 s window.

The route is unauthenticated and on by default, as before. It honours
ENABLE_PROMETHEUS (default true); setting it to false removes the route and
the HTTP instrumentation. Each app uses its own CollectorRegistry, so
rebuilding the app never hits duplicated-timeseries errors.
/v2/server/metrics is unchanged. inference_pipeline_* gauges are not
emitted: this server has no stream pipeline API.

prometheus-fastapi-instrumentator>=8.1.0 is required: older releases crash
on the router wrappers that current FastAPI puts in app.routes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add describe_workload routes to the new inference server

Release 1.7.2 serves `POST /workflows/describe_workload` and
`POST /{workspace_name}/workflows/{workflow_id}/describe_workload`
(experimental workload introspection). The new server did not register
them, so the inline variant fell through to the legacy
`/{dataset_id}/{version_id}` catch-all and answered 404 instead of the
1.7.2 responses (for example 400 "Required Roboflow API key is missing").

The introspection itself already lives in roboflow-workflows
(`describe_workflow_workload`), so only the host glue is added: the
request bodies, the platform bindings for inner workflow resolution, a
registry-backed model metadata provider (metadata lookup only, cached
per api key and model id, skipped for third-party models and in offline
mode), the `DISABLE_WORKFLOW_WORKLOAD_ENDPOINTS` switch, and the two
routes on the workflows router, which is included before the legacy
catch-all. Request/response shapes, key handling and error codes match
1.7.2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Echo any incoming X-Request-ID instead of replacing non-UUID4 values

The 1.7.x server images set API_LOGGING_ENABLED=True, which configures
asgi_correlation_id with an accept-all validator and identity transformer.
Those images echo whatever X-Request-ID the caller sends and generate a
uuid4 hex only when the header is absent or empty. The new server replaced
every non-UUID4 value and logged a warning per request, breaking clients
and proxies that correlate logs with their own id formats.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Route model-metadata lookups through SECURE_GATEWAY

The model-stat check that runs on every old-API model request and the
workload-introspection metadata lookup both called
get_one_page_of_model_metadata without a proxy_url_builder, so they hit
the Roboflow API directly even when SECURE_GATEWAY was configured. The
weights download already goes through the gateway, and so does the
server's platform client (wrap_url), so behind a gateway these lookups
failed or leaked egress around it.

Both call sites now pass roboflow_secure_gateway_proxy_url_builder, the
same builder the download path and wrap_url use. It reads the gateway
from inference_models configuration and is a no-op when SECURE_GATEWAY
is unset, so the direct URL is unchanged in that case. Test fakes of the
registry call now accept the extra keyword.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Cache workspace lookups in the new server's platform client

The legacy server caches get_roboflow_workspace with cachetools
@ttl_cache (MODELS_CACHE_AUTH_CACHE_TTL, 15 min; max size
MODELS_CACHE_AUTH_CACHE_MAX_SIZE). ServerRoboflowPlatformClient had no
cache, so visual_search_classifier (which resolves the workspace outside
the Workflows cache) made one extra Roboflow API call per image.

Cache successful lookups process-wide in the existing TTL-LRU cache,
keyed by the api key's SHA-256, with the same env variables and
defaults as the legacy server. Failures are not cached, as in legacy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Stop double-rotating CMYK and YCbCr TIFFs in the imagecodecs decoder

imagecodecs hands CMYK (Photometric=Separated), YCbCr and old-style JPEG
TIFFs to libtiff's TIFFReadRGBAImageOriented, which already applies the
flip part of the Orientation tag (2-4 fully, 5-8 flipped but not
transposed). The decoder then applied the full orientation again, so
e.g. a CMYK TIFF with orientation 3 came out upside down, unlike
cv2.imdecode. New-style JPEG TIFFs (Compression=7) take the plain path
and are unaffected, even when CMYK.

Detect that path from IFD0 Compression/Photometric, mirroring
imagecodecs' dispatch, and apply only the remaining transpose for 5-8.
The output is still the stored image oriented once, so decoded_dims'
5-8 swap keeps matching the pixels. Tests pin CMYK and YCbCr TIFFs
against cv2 for orientations 1-8, plus reported dims.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep plugin-path disable patterns matching the platform blocks

Inference 1.6.1-1.7.2 shipped the 9 Roboflow-platform blocks (dataset
upload v1/v2, custom metadata, model monitoring, vision events, vision
event bundle, visual search, visual search classifier, asset library
attributes) from `inference.roboflow_workflows_plugin`. Its loader
matched WORKFLOW_DISABLED_BLOCK_PATTERNS as case-insensitive substrings
of both that path and the pre-1.6.1 `inference.core.workflows` path.
Now that the blocks are core blocks again, the core loader only knew the
`roboflow_workflows.*` and `inference.core.workflows.*` spellings, so an
operator who disabled e.g. dataset upload by its plugin path silently got
the block back on upgrade.

`_compat_names.to_historic_plugin_module` maps the two platform-block
subtrees back to their plugin path, and `_should_filter_block` matches
patterns against it too. This is for pattern matching only: the plugin
package stays non-importable. Both servers share this loader.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Refuse workspace lookups and generic posts in OFFLINE_MODE

The new server's platform client made an outbound GET for an uncached
workspace lookup, and an outbound POST from the generic `post` used by
the proxied LLM/OCR/notification blocks, even with OFFLINE_MODE on. The
legacy server refuses both: `_get_from_url` raises before any request
and `post_to_roboflow_api` raises RoboflowAPIConnectionError.

Both now call `_refuse_when_offline`, like the sibling methods. The
workspace guard sits behind the cache lookup, because the legacy
`ttl_cache` still answers for an already-resolved key while offline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
raise RoboflowAPIRequestError(
"Empty workspace encountered, check your API key."
)
cache_key = sha256(api_key.encode("utf-8")).hexdigest()
… the legacy-compatible replacement server (#3126)

* Raise supervision to 0.30.6 (#3079)

* Raise supervision to 0.30.6 and adapt to its behaviour changes
- Filter polygons with fewer than 3 points before Detections.from_inference in Workflows, the render sink and the CLI, keeping RLE-backed predictions
- Round LeanPolygonZone anchors with np.rint, as sv.PolygonZone now does
- Delegate compact_mask_from_coco_rle to CompactMask.from_coco_rl
- Run the torch NMM port at native resolution, counting overlaps only where mask bounding boxes intersect

* Fix streamvision dependency declarations and Modal WebRTC startup handling (#3087)

* Fix streamvision dependencies and Modal WebRTC startup handling
- Declare Pillow and optional Torch dependencies; update docs and lockfile
- Include stream_vision in formatting and lint checks; exclude virtualenvs
- Remove historical extraction checks and baseline-fetch CI setup
- Await Modal queue delivery and avoid duplicate watchdog shutdown
- Preserve startup exceptions and add lifecycle regression coverage

* Log WebRTC session errors instead of re-raising from Modal function

* Remove profiler, reconcile pins and add legacy env, preload and readiness parity
- delete inference_models.utils.performance and every profiler call site
- inference-models 0.39.1rc1, inference-model-manager 0.5.1, inference-server 0.4.1 with matching pins and a pin consistency test
- close the workflows 0.2.4rc3 changelog entry; drop enterprise runtime name from docker error text
- legacy_env.apply_legacy_env maps legacy env names, region-derived URLs and legacy defaults before configuration loads
- gateway.load takes pinned; preload and /model/add load unpinned, PINNED_MODELS pins, PRELOAD_HF_IDS loads owlv2 models off the readiness path
- readiness means startup preload finished; GATEWAY_API_VERSION 2

* Add hosted auth modes and route visibility to inference server
- serverless credit-check middleware under GCP_SERVERLESS with legacy skip list, error bodies, cache TTLs and transient-error retries
- dedicated deployment allow-list middleware under DEDICATED_DEPLOYMENT_WORKSPACE_URL
- billing intent resolved per request for every deployment; credits header on the model authorisation call, cache keyed by auth context
- assume-identity headers only on the model stat lookup, opt-in on the shared platform client
- legacy route gating under LAMBDA and GCP_SERVERLESS at registration time
- platform_http shared client so hosted modes do not need the workflows extra

* Add legacy response headers and observability context to inference server
- model cold-start, load-time, load-details, model id and workflow id headers from per-request load events
- cold start attributed only to the request that performed the load; load time measured inside the load
- correlation id middleware with legacy validation and request write-back; execution id prepared under GCP_SERVERLESS
- serverless auth denials carry the observability headers like legacy
- refuse OFFLINE_MODE together with hosted deployment flags at startup
- CORS expose_headers list from legacy; x-inference-engine header
- ensure_loaded reports load outcome and duration

* Record model loads inside the gateway for response headers
- capture the request load context when a load future is created and record the completed load from the executor job
- explicit /v2/models/load and internal reloads count; joining requests record nothing; no owner registry
- requested model id recorded before the load so X-Model-Id survives load errors and timeouts
- ensure_loaded returns (model_ready,) again; gateway contract documents gateway-owned recording
- correlation header name follows legacy: X-Request-ID unless API_LOGGING_ENABLED; header appended, invalid ids warned

* Add legacy logging modes and OpenTelemetry to inference server
- configure_logging with LOG_LEVEL, API_LOGGING_ENABLED and STRUCTURED_API_LOGGING; plain uvicorn access line by default, JSON access and application records with the correlation key when enabled, health paths at DEBUG in structured mode
- structured access log carries the legacy fields from response headers and never the query string
- telemetry module with guarded OpenTelemetry imports: tracer and meter providers, OTLP exporters, FastAPI and requests instrumentation, force-trace header and sampler, legacy helper API, X-Trace-Id on every response
- shutdown_telemetry wired into the lifespan; optional otel extra

* Bring inference server /metrics to legacy parity
- per-model inference count, error count and average time from a per-process event recorder fed at every inference boundary, 10 s window, response-weighted average, legacy monitoring opt-outs
- one event per HTTP request on composite embedding paths; v2 named instances keyed by routing key
- scrape limited to the first 25 loaded models, one series per model via route display ids
- manager gauges and stream-manager gauges with legacy names, empty families and zero active streams without a client
- source labels sanitised like legacy
- routing_key moved to a shared routing module

* Bring /model/registry rows and fields to legacy parity - one row per requested id, keyed like legacy: alias for inference routes and preload, de-aliased id for /model/add, id as passed elsewhere - request_aliases and request_paths recorded per row from what each route passes; nothing inferred from the gateway registry id - input_height, input_width and vram_bytes reported by the model manager; total_vram_bytes counts each loaded model once; batch_size stays null - /model/remove drops the de-aliased row and unloads only when no row is left, inside the legacy routes only - rows show only what was recorded since the current load, using a monotonic load stamp from the model manager; a request served by a reload keeps its row - inference-model-manager 0.5.2

* Record /model/registry rows for models registered by workflow steps

* Keep request path and alias of a failed model load for /model/registry

* Start inference server images like the legacy images

* Map errors on legacy routes like the legacy server

* Answer input, content-type and platform failures on legacy routes like the legacy server

* Add route parity check between legacy server and inference server

* Answer model load failures on legacy routes by cause like the legacy server

* Run workflow sinks before the response on serverless like the legacy server

* Install OpenTelemetry packages in inference server images

* Bound request metadata kept for failed model loads

* Answer request mistakes on legacy routes with 400 and the reason

* Apply legacy image URL policy on legacy routes and workflows

* Skip disk cache watchdog when the server runs offline

* Accept legacy image input forms on legacy routes and workflows

* Accept numpy mask format on legacy semantic segmentation route

* Honour ALLOW_API_KEY_FROM_HEADERS on legacy routes

* Use the model layer offline decision on legacy routes, workflows and builder

* Answer workflow step errors like the legacy server

* Add device stats, secure gateway health, notebook start and logs routes

* Align info id, platform TLS switch, gateway auth lookup and thread pools with legacy

* Add periodic metrics pingback on legacy routes and workflows

* Pin per-key model access check for loaded models

* Record OpenTelemetry metrics and errors on legacy routes

* Add active learning package for legacy routes

* Run active learning after inference on legacy routes

* Add usage collector core for legacy routes

* Simplify usage package structure and naming

* Keep model latency as an attribute when merging usage rows

* Record each usage call as one row regardless of list length

* Record usage rows on legacy model routes

* Record usage rows on workflow run routes

* Add serverless request context and usage contract test

* Honour serverless, air-gapped, enterprise and MQTT workflow settings

* Cache workflow definitions on disk in the legacy layout and support a Redis workflows cache

* Serve every model family on the tensor-native workflow path

* Align EasyOCR, LMM router gate, SAM knobs and depth inversion with legacy

* Accept legacy model options, SAM mask input and binary format, and OWLv2 few-shot

* Serve weightless stub models, anomaly detection responses and the full registry MRO table

* Forward per-request preprocessing overrides through the model manager and ignore legacy-only options

* Proxy SAM3 segmentation to the platform in remote mode and send the platform's chunked-response header

* Add action recognition groundwork: registry entry, video sampling metadata, streaming video input and legacy video settings

* Serve action recognition on the legacy route, the catch-all and the workflows provider

* Add a vLLM proxy serving backend for Qwen3-VL, Qwen3.5 and Qwen3.8 models

* Route vLLM proxy models through the model manager and the legacy server

* Accept caller-supplied SAM embeddings through the model manager

* Add stream API settings and a pipeline host for streamvision pipeline processes

* Serve the inference pipeline routes with a stream manager process per worker

* Record one usage request row per workflow run inside stream pipeline processes

* Pipeline RF-DETR instance segmentation frames through the model manager for stream workflows

* Keep collecting usage offline with silent bounded retries and take only rows of known API keys from the shared queue

* Merge billable model and custom Python entries in the legacy usage helper and keep the later list when entries cannot be merged

* Send platform requests under apiproxy to API_PROXY_BASE_URL as the legacy server does
) -> dict:
cache_key = (
f"workflow_definition:{workspace_id}:{workflow_id}:{workflow_version_id}:"
f"{sha256((api_key or '').encode()).hexdigest()}"
if isinstance(error, ServerBusyError):
return JSONResponse(
status_code=503,
content={"message": str(error)},
else:
logger.error("%s: %s", type(error).__name__, error, exc_info=error)

return JSONResponse(status_code=status_code, content=content)
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants