Repository navigation
Feat/new model manager - #2251
grzegorz-roboflow wants to merge 86 commits into
Conversation
b0de01d to
c27b669
Compare
b9236ec to
bf0f4d4
Compare
f67b160 to
fb002eb
Compare
dkosowski87
left a comment
There was a problem hiding this comment.
1/6 with the pace I'm currently at ;)
|
Warning Review the following alerts detected in dependencies. According to your organization's Security Policy, it is recommended to resolve "Warn" alerts. Learn more about Socket for GitHub.
|
95a3a0c to
c7c0af0
Compare
fa5f76a to
4c4a435
Compare
b9a09b5 to
60f7a91
Compare
- shared fetcher refuses with 403 before any DNS work - covers every V2 handler parser
- published rc1 lacks the mqtt host configuration flag
…h-published rcs (#3068) `inference-models 0.39.0rc3` was published to PyPI today (10:48 UTC) from the feat/new-model-manager branch (PR #2251). main's requirement files pinned `~=0.39.0rc2`, and a pre-release lower bound lets the resolver take any newer rc, so every image built since installs rc3. rc3 changes the SAM3 output contract (default `mask_format="rle"`, dict results) and main's adapters in inference/models/sam3/ still expect tensors, which is why "Code Quality & Regression Tests - NVIDIA T4" fails in the SAM3 step with HTTP 500 (run 36145992322). Exact-pin all five requirement files to 0.39.0rc2, the version the last green image installed. No code change; the pin is lifted by the PR that adopts the new inference-models contract. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…ing dormancy Replace the path-filter dormancy check with the plain fix: remove unit_tests_inference_model_manager_and_server.yml from TEST_WORKFLOWS, with a comment saying why and when to put it back (once inference_model_manager/ and inference_server/ land on main, #2251). The status-sync listener drops the same workflow name so the two lists stay aligned. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… after re-runs (#3077) * Fast-track CI: skip workflows that are dormant on the branch The orchestrator dispatched unit_tests_inference_model_manager_and_server.yml on every fast-track branch. That workflow is path-filtered to inference_model_manager/** and inference_server/**, which exist only on feat/new-model-manager, so it never runs on main, but workflow_dispatch ignores path filters and the run died at install with "Distribution not found at: .../inference_model_manager" (3 of 3 fast-track dispatches). Before dispatching, read each workflow's pull_request path filter as it is on the branch and skip the workflow when every literal root of the filter is absent there: a PR to main would never have started it. The skip is reported as a success status and a summary row naming the absent roots. Workflows without a path filter, with a root-level pattern, or with at least one root present are dispatched as before. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Fast-track CI: refresh commit statuses after a re-run The orchestrator writes each fast-track-ci / <workflow> status once, from the attempt it waited for. A maintainer re-run of a failed suite completes outside the orchestrator, so a suite that went green on re-run kept a red status on the umbrella PR (integration_tests_workflows_x86 on #3076). Add a workflow_run listener on the same test workflows that, for a re-run attempt (run_attempt > 1) of a workflow_dispatch run on a fast-track branch, rewrites the status from that attempt. It only touches a status the orchestrator created and never checks out branch code. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Fast-track CI: drop the dormant model-manager suite instead of detecting dormancy Replace the path-filter dormancy check with the plain fix: remove unit_tests_inference_model_manager_and_server.yml from TEST_WORKFLOWS, with a comment saying why and when to put it back (once inference_model_manager/ and inference_server/ land on main, #2251). The status-sync listener drops the same workflow name so the two lists stay aligned. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* Apply EXIF orientation (JPEG, PNG, TIFF) in the new server decoder
The previous server decoded images with cv2.imdecode(IMREAD_COLOR), which
honours the EXIF Orientation tag. The imagecodecs decoder ignores it, so a
phone photo stored rotated with Orientation 3/6/8 reached the model
sideways or upside down, gave different detections, and the legacy routes
reported width and height swapped.
The decoder now reads the Orientation tag from the JPEG APP1 Exif block
(header walk only, no work when the tag is absent or 1) and rotates/flips
the decoded array before the BGR copy, so no extra copy is made.
image_dims uses the same tag reader and swaps width/height for
orientations 5-8, so the reported size matches the decoded pixels.
Output is pixel-identical to cv2.imdecode for all eight orientations.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Serve the Roboflow-platform workflow blocks as core blocks on both servers
The new server listed 237 blocks in /workflows/blocks/describe against 246
on 1.7.2: dataset upload v1/v2, custom metadata, model monitoring, vision
events, vision event bundle, visual search, visual search classifier and
asset-library attributes were missing, so workflows using them did not
compile. They lived in `inference/roboflow_workflows_plugin`, inside the
legacy package the new server image does not install, and they imported
`inference.core.roboflow_api` and `inference.core.active_learning`.
- The blocks move to `core_steps/sinks/roboflow/` and
`core_steps/integrations/roboflow/` of roboflow-workflows, where they
lived before the plugin extraction, and `load_blocks()` registers them,
with the `_tensor` siblings under ENABLE_TENSOR_DATA_REPRESENTATION.
Identifiers, manifests and block_source are unchanged;
fully_qualified_block_class_name is the historic
`inference.core.workflows.core_steps...` path. The dataset-upload
quota helpers move next to those blocks with the same cache keys and
expiry.
- The plugin package, its loader and WORKFLOWS_PLUGINS entry are removed.
`inference.roboflow_workflows_plugin` was never a public API and is not
kept as an alias.
- The blocks' Roboflow API calls become members of the existing
`workflows_core.platform_client` port, injected as a block init
parameter - no process-global setup. The `inference` server client
forwards them to `inference.core.roboflow_api` (same functions, retries
and errors); the `inference_server` client implements the endpoints with
its own headers and secure-gateway wrapping. The offline default refuses
the API calls, as it refuses `post`.
- Tests inject fake clients instead of patching module functions.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Add model, engine and request-id response headers to the server
The legacy inference server (1.7.x) attaches X-Model-Id,
X-Model-Cold-Start, X-Model-Cold-Start-Count, x-inference-engine and
X-Request-ID to every response, plus X-Model-Load-Time and
X-Model-Load-Details on the request that loaded a model. The new server
sent none of them, so clients and dashboards reading them broke.
A raw ASGI middleware now installs a per-request collector in a
ContextVar and writes the headers on response start, with the same
names and value formats as 1.7.x (load time as str(float) seconds,
details as [{"m": id, "t": seconds}] JSON dropped above 4096 bytes,
comma-joined sorted model ids). The legacy bridge and the v2 dispatch
record the model ids they use; ModelManagerGateway records a load only
for the request that created the load future, so requests joining an
in-flight load are not reported as cold starts. X-Request-ID keeps a
valid incoming UUID4 and otherwise generates uuid4().hex, like the
asgi_correlation_id defaults used by 1.7.x. The model headers are
added to CORS expose_headers as before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Serve Prometheus metrics at /metrics in the new server
The legacy server answers GET /metrics with Prometheus text metrics out of
the box (the CPU image included). The new server only has
/v2/server/metrics, which returns JSON and is off unless
ENABLE_CONTROL_PLANE_ROUTES=true, so scrapers pointed at /metrics got 404.
Add /metrics with the metric families the legacy server exposes:
- HTTP request metrics from prometheus-fastapi-instrumentator
(http_requests_total, http_request_size_bytes, http_response_size_bytes,
http_request_duration_seconds, http_request_duration_highr_seconds),
- process, platform and GC collectors (process_*, python_info,
python_gc_*),
- the legacy per-model gauges num_inferences_<model>,
avg_inference_time_<model>, num_errors_<model> and their *_total
roll-ups, derived from the model manager's cumulative counters over a
10 s window.
The route is unauthenticated and on by default, as before. It honours
ENABLE_PROMETHEUS (default true); setting it to false removes the route and
the HTTP instrumentation. Each app uses its own CollectorRegistry, so
rebuilding the app never hits duplicated-timeseries errors.
/v2/server/metrics is unchanged. inference_pipeline_* gauges are not
emitted: this server has no stream pipeline API.
prometheus-fastapi-instrumentator>=8.1.0 is required: older releases crash
on the router wrappers that current FastAPI puts in app.routes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Add describe_workload routes to the new inference server
Release 1.7.2 serves `POST /workflows/describe_workload` and
`POST /{workspace_name}/workflows/{workflow_id}/describe_workload`
(experimental workload introspection). The new server did not register
them, so the inline variant fell through to the legacy
`/{dataset_id}/{version_id}` catch-all and answered 404 instead of the
1.7.2 responses (for example 400 "Required Roboflow API key is missing").
The introspection itself already lives in roboflow-workflows
(`describe_workflow_workload`), so only the host glue is added: the
request bodies, the platform bindings for inner workflow resolution, a
registry-backed model metadata provider (metadata lookup only, cached
per api key and model id, skipped for third-party models and in offline
mode), the `DISABLE_WORKFLOW_WORKLOAD_ENDPOINTS` switch, and the two
routes on the workflows router, which is included before the legacy
catch-all. Request/response shapes, key handling and error codes match
1.7.2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Echo any incoming X-Request-ID instead of replacing non-UUID4 values
The 1.7.x server images set API_LOGGING_ENABLED=True, which configures
asgi_correlation_id with an accept-all validator and identity transformer.
Those images echo whatever X-Request-ID the caller sends and generate a
uuid4 hex only when the header is absent or empty. The new server replaced
every non-UUID4 value and logged a warning per request, breaking clients
and proxies that correlate logs with their own id formats.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Route model-metadata lookups through SECURE_GATEWAY
The model-stat check that runs on every old-API model request and the
workload-introspection metadata lookup both called
get_one_page_of_model_metadata without a proxy_url_builder, so they hit
the Roboflow API directly even when SECURE_GATEWAY was configured. The
weights download already goes through the gateway, and so does the
server's platform client (wrap_url), so behind a gateway these lookups
failed or leaked egress around it.
Both call sites now pass roboflow_secure_gateway_proxy_url_builder, the
same builder the download path and wrap_url use. It reads the gateway
from inference_models configuration and is a no-op when SECURE_GATEWAY
is unset, so the direct URL is unchanged in that case. Test fakes of the
registry call now accept the extra keyword.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Cache workspace lookups in the new server's platform client
The legacy server caches get_roboflow_workspace with cachetools
@ttl_cache (MODELS_CACHE_AUTH_CACHE_TTL, 15 min; max size
MODELS_CACHE_AUTH_CACHE_MAX_SIZE). ServerRoboflowPlatformClient had no
cache, so visual_search_classifier (which resolves the workspace outside
the Workflows cache) made one extra Roboflow API call per image.
Cache successful lookups process-wide in the existing TTL-LRU cache,
keyed by the api key's SHA-256, with the same env variables and
defaults as the legacy server. Failures are not cached, as in legacy.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Stop double-rotating CMYK and YCbCr TIFFs in the imagecodecs decoder
imagecodecs hands CMYK (Photometric=Separated), YCbCr and old-style JPEG
TIFFs to libtiff's TIFFReadRGBAImageOriented, which already applies the
flip part of the Orientation tag (2-4 fully, 5-8 flipped but not
transposed). The decoder then applied the full orientation again, so
e.g. a CMYK TIFF with orientation 3 came out upside down, unlike
cv2.imdecode. New-style JPEG TIFFs (Compression=7) take the plain path
and are unaffected, even when CMYK.
Detect that path from IFD0 Compression/Photometric, mirroring
imagecodecs' dispatch, and apply only the remaining transpose for 5-8.
The output is still the stored image oriented once, so decoded_dims'
5-8 swap keeps matching the pixels. Tests pin CMYK and YCbCr TIFFs
against cv2 for orientations 1-8, plus reported dims.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Keep plugin-path disable patterns matching the platform blocks
Inference 1.6.1-1.7.2 shipped the 9 Roboflow-platform blocks (dataset
upload v1/v2, custom metadata, model monitoring, vision events, vision
event bundle, visual search, visual search classifier, asset library
attributes) from `inference.roboflow_workflows_plugin`. Its loader
matched WORKFLOW_DISABLED_BLOCK_PATTERNS as case-insensitive substrings
of both that path and the pre-1.6.1 `inference.core.workflows` path.
Now that the blocks are core blocks again, the core loader only knew the
`roboflow_workflows.*` and `inference.core.workflows.*` spellings, so an
operator who disabled e.g. dataset upload by its plugin path silently got
the block back on upgrade.
`_compat_names.to_historic_plugin_module` maps the two platform-block
subtrees back to their plugin path, and `_should_filter_block` matches
patterns against it too. This is for pattern matching only: the plugin
package stays non-importable. Both servers share this loader.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Refuse workspace lookups and generic posts in OFFLINE_MODE
The new server's platform client made an outbound GET for an uncached
workspace lookup, and an outbound POST from the generic `post` used by
the proxied LLM/OCR/notification blocks, even with OFFLINE_MODE on. The
legacy server refuses both: `_get_from_url` raises before any request
and `post_to_roboflow_api` raises RoboflowAPIConnectionError.
Both now call `_refuse_when_offline`, like the sibling methods. The
workspace guard sits behind the cache lookup, because the legacy
`ttl_cache` still answers for an already-resolved key while offline.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
| raise RoboflowAPIRequestError( | ||
| "Empty workspace encountered, check your API key." | ||
| ) | ||
| cache_key = sha256(api_key.encode("utf-8")).hexdigest() |
… the legacy-compatible replacement server (#3126) * Raise supervision to 0.30.6 (#3079) * Raise supervision to 0.30.6 and adapt to its behaviour changes - Filter polygons with fewer than 3 points before Detections.from_inference in Workflows, the render sink and the CLI, keeping RLE-backed predictions - Round LeanPolygonZone anchors with np.rint, as sv.PolygonZone now does - Delegate compact_mask_from_coco_rle to CompactMask.from_coco_rl - Run the torch NMM port at native resolution, counting overlaps only where mask bounding boxes intersect * Fix streamvision dependency declarations and Modal WebRTC startup handling (#3087) * Fix streamvision dependencies and Modal WebRTC startup handling - Declare Pillow and optional Torch dependencies; update docs and lockfile - Include stream_vision in formatting and lint checks; exclude virtualenvs - Remove historical extraction checks and baseline-fetch CI setup - Await Modal queue delivery and avoid duplicate watchdog shutdown - Preserve startup exceptions and add lifecycle regression coverage * Log WebRTC session errors instead of re-raising from Modal function * Remove profiler, reconcile pins and add legacy env, preload and readiness parity - delete inference_models.utils.performance and every profiler call site - inference-models 0.39.1rc1, inference-model-manager 0.5.1, inference-server 0.4.1 with matching pins and a pin consistency test - close the workflows 0.2.4rc3 changelog entry; drop enterprise runtime name from docker error text - legacy_env.apply_legacy_env maps legacy env names, region-derived URLs and legacy defaults before configuration loads - gateway.load takes pinned; preload and /model/add load unpinned, PINNED_MODELS pins, PRELOAD_HF_IDS loads owlv2 models off the readiness path - readiness means startup preload finished; GATEWAY_API_VERSION 2 * Add hosted auth modes and route visibility to inference server - serverless credit-check middleware under GCP_SERVERLESS with legacy skip list, error bodies, cache TTLs and transient-error retries - dedicated deployment allow-list middleware under DEDICATED_DEPLOYMENT_WORKSPACE_URL - billing intent resolved per request for every deployment; credits header on the model authorisation call, cache keyed by auth context - assume-identity headers only on the model stat lookup, opt-in on the shared platform client - legacy route gating under LAMBDA and GCP_SERVERLESS at registration time - platform_http shared client so hosted modes do not need the workflows extra * Add legacy response headers and observability context to inference server - model cold-start, load-time, load-details, model id and workflow id headers from per-request load events - cold start attributed only to the request that performed the load; load time measured inside the load - correlation id middleware with legacy validation and request write-back; execution id prepared under GCP_SERVERLESS - serverless auth denials carry the observability headers like legacy - refuse OFFLINE_MODE together with hosted deployment flags at startup - CORS expose_headers list from legacy; x-inference-engine header - ensure_loaded reports load outcome and duration * Record model loads inside the gateway for response headers - capture the request load context when a load future is created and record the completed load from the executor job - explicit /v2/models/load and internal reloads count; joining requests record nothing; no owner registry - requested model id recorded before the load so X-Model-Id survives load errors and timeouts - ensure_loaded returns (model_ready,) again; gateway contract documents gateway-owned recording - correlation header name follows legacy: X-Request-ID unless API_LOGGING_ENABLED; header appended, invalid ids warned * Add legacy logging modes and OpenTelemetry to inference server - configure_logging with LOG_LEVEL, API_LOGGING_ENABLED and STRUCTURED_API_LOGGING; plain uvicorn access line by default, JSON access and application records with the correlation key when enabled, health paths at DEBUG in structured mode - structured access log carries the legacy fields from response headers and never the query string - telemetry module with guarded OpenTelemetry imports: tracer and meter providers, OTLP exporters, FastAPI and requests instrumentation, force-trace header and sampler, legacy helper API, X-Trace-Id on every response - shutdown_telemetry wired into the lifespan; optional otel extra * Bring inference server /metrics to legacy parity - per-model inference count, error count and average time from a per-process event recorder fed at every inference boundary, 10 s window, response-weighted average, legacy monitoring opt-outs - one event per HTTP request on composite embedding paths; v2 named instances keyed by routing key - scrape limited to the first 25 loaded models, one series per model via route display ids - manager gauges and stream-manager gauges with legacy names, empty families and zero active streams without a client - source labels sanitised like legacy - routing_key moved to a shared routing module * Bring /model/registry rows and fields to legacy parity - one row per requested id, keyed like legacy: alias for inference routes and preload, de-aliased id for /model/add, id as passed elsewhere - request_aliases and request_paths recorded per row from what each route passes; nothing inferred from the gateway registry id - input_height, input_width and vram_bytes reported by the model manager; total_vram_bytes counts each loaded model once; batch_size stays null - /model/remove drops the de-aliased row and unloads only when no row is left, inside the legacy routes only - rows show only what was recorded since the current load, using a monotonic load stamp from the model manager; a request served by a reload keeps its row - inference-model-manager 0.5.2 * Record /model/registry rows for models registered by workflow steps * Keep request path and alias of a failed model load for /model/registry * Start inference server images like the legacy images * Map errors on legacy routes like the legacy server * Answer input, content-type and platform failures on legacy routes like the legacy server * Add route parity check between legacy server and inference server * Answer model load failures on legacy routes by cause like the legacy server * Run workflow sinks before the response on serverless like the legacy server * Install OpenTelemetry packages in inference server images * Bound request metadata kept for failed model loads * Answer request mistakes on legacy routes with 400 and the reason * Apply legacy image URL policy on legacy routes and workflows * Skip disk cache watchdog when the server runs offline * Accept legacy image input forms on legacy routes and workflows * Accept numpy mask format on legacy semantic segmentation route * Honour ALLOW_API_KEY_FROM_HEADERS on legacy routes * Use the model layer offline decision on legacy routes, workflows and builder * Answer workflow step errors like the legacy server * Add device stats, secure gateway health, notebook start and logs routes * Align info id, platform TLS switch, gateway auth lookup and thread pools with legacy * Add periodic metrics pingback on legacy routes and workflows * Pin per-key model access check for loaded models * Record OpenTelemetry metrics and errors on legacy routes * Add active learning package for legacy routes * Run active learning after inference on legacy routes * Add usage collector core for legacy routes * Simplify usage package structure and naming * Keep model latency as an attribute when merging usage rows * Record each usage call as one row regardless of list length * Record usage rows on legacy model routes * Record usage rows on workflow run routes * Add serverless request context and usage contract test * Honour serverless, air-gapped, enterprise and MQTT workflow settings * Cache workflow definitions on disk in the legacy layout and support a Redis workflows cache * Serve every model family on the tensor-native workflow path * Align EasyOCR, LMM router gate, SAM knobs and depth inversion with legacy * Accept legacy model options, SAM mask input and binary format, and OWLv2 few-shot * Serve weightless stub models, anomaly detection responses and the full registry MRO table * Forward per-request preprocessing overrides through the model manager and ignore legacy-only options * Proxy SAM3 segmentation to the platform in remote mode and send the platform's chunked-response header * Add action recognition groundwork: registry entry, video sampling metadata, streaming video input and legacy video settings * Serve action recognition on the legacy route, the catch-all and the workflows provider * Add a vLLM proxy serving backend for Qwen3-VL, Qwen3.5 and Qwen3.8 models * Route vLLM proxy models through the model manager and the legacy server * Accept caller-supplied SAM embeddings through the model manager * Add stream API settings and a pipeline host for streamvision pipeline processes * Serve the inference pipeline routes with a stream manager process per worker * Record one usage request row per workflow run inside stream pipeline processes * Pipeline RF-DETR instance segmentation frames through the model manager for stream workflows * Keep collecting usage offline with silent bounded retries and take only rows of known API keys from the shared queue * Merge billable model and custom Python entries in the legacy usage helper and keep the later list when entries cannot be merged * Send platform requests under apiproxy to API_PROXY_BASE_URL as the legacy server does
| ) -> dict: | ||
| cache_key = ( | ||
| f"workflow_definition:{workspace_id}:{workflow_id}:{workflow_version_id}:" | ||
| f"{sha256((api_key or '').encode()).hexdigest()}" |
| if isinstance(error, ServerBusyError): | ||
| return JSONResponse( | ||
| status_code=503, | ||
| content={"message": str(error)}, |
| else: | ||
| logger.error("%s: %s", type(error).__name__, error, exc_info=error) | ||
|
|
||
| return JSONResponse(status_code=status_code, content=content) |
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
…e for the capability guard
…xpose it through the gateway
…ck rows alongside the request row
…the server images
…s and declare safetensors in inference-models
…RLE masks, the RF-DETR cap and disabled backends explicitly from the model manager and server
…RLE masks, the RF-DETR cap and disabled backends explicitly from the model manager and server
What does this PR do?
Adding
inference_model_managerType of Change
Testing
Initial suit of tests added
Checklist
Additional Context
N/A