- Preface
- Deployment methods
- Integration with OGX framework
- Local deployment
- Running from container
- Usage
In this document, you will learn how to install and run a service called Lightspeed Core Stack (LCS). It is a service that allows users to communicate with large language models (LLMs), access to RAG databases, call so called agents, process conversation history, ensure that the conversation is only about permitted topics, etc.
Lightspeed Core Stack (LCS) is built on the OGX framework, which can be run in several modes. Additionally, it is possible to run LCS locally (as a regular Python application) or from within a container. This means that it is possible to leverage multiple deployment methods:
- Local deployment
- OGX framework is used as a library
- OGX framework is used as a separate process (deployed locally)
- Running from a container
- OGX framework is used as a library
- OGX framework is used as a separate process
All those deployments methods will be covered later.
LCS reads one operator-facing file: lightspeed-stack.yaml. There are two
ways it can drive the underlying OGX:
- Unified mode (recommended). The single
lightspeed-stack.yamlis the only configuration file you maintain. LCORE synthesizes the OGXrun.yamlfrom it at startup — from a built-in default baseline, an optional profile you author, the high-levelinference.providerssection, and a rawnative_overrideescape hatch. All examples in this guide show unified mode first. - Legacy two-file mode (deprecated).
ogx.library_client_config_path(the deprecatedllama_stackYAML-section alias is still accepted) points at an external, hand-maintainedrun.yaml. This path is deprecated: since release 0.6 it logs a startup warning, and it is removed in release 0.7. See Migrating from the legacy two-file configuration.
Note
The deprecated llama_stack YAML key remains accepted with a startup
warning; use ogx: for new configuration. See
Migrating from the legacy two-file configuration
and v0.7.0 migration notes.
The two modes are mutually exclusive in one file — configuration loading
fails if a unified synthesis input and library_client_config_path are both
present.
The OGX framework can be run as a standalone server and accessed via its the REST API. However, instead of direct communication via the REST API (and JSON format), there is an even better alternative. It is based on the so-called OGX Client. It is a library available for Python, Swift, Node.js or Kotlin, which "wraps" the REST API stack in a suitable way, which is easier for many applications.
When this mode is selected, OGX is used as a regular Python library. This means that the library must be installed in the system Python environment, a user-level environment, or a virtual environment. All calls to OGX are performed via standard function or method calls:
Note
Even when OGX is used as a library, it still requires a run.yaml
configuration during the initialization phase. In unified mode (the
recommended default) LCORE synthesizes that file for you from
lightspeed-stack.yaml; only the deprecated legacy mode requires you to
maintain run.yaml by hand.
In unified mode (where LCORE synthesizes the OGX run.yaml from
lightspeed-stack.yaml instead of reading an external file), the synthesis
starts from a baseline. By default that is LCORE's built-in baseline; a
profile replaces it with a file you author.
A profile is an ordinary run.yaml-shaped YAML file — the same schema OGX reads natively. Everything else in the unified pipeline (enrichment,
the high-level inference.providers section, ensuring the MCP tool_runtime
provider, then native_override) is applied on top of the profile, in that
order. The MCP ensure adds provider_id: model-context-protocol when missing
so static mcp_servers and dynamic MCP registration work; it is skipped only
for baseline: empty (use native_override there if you need MCP).
Authoring a profile. Start from one of the reference profiles shipped in
examples/profiles/:
openai-remote.yaml— remote OpenAI inference plus inline sentence-transformers embeddings and FAISS; boots with onlyOPENAI_API_KEYset.inline-faiss.yaml— fully inline (sentence-transformers + FAISS), no remote provider and no API key; pair it with a chat provider via the high-levelinference.providerssection.
Keep secrets out of the file: write ${env.MY_KEY} environment references,
which OGX resolves at startup.
Referencing a profile. Point ogx.config.profile at the file:
name: Lightspeed Core Service (LCS)
service:
host: 0.0.0.0
port: 8080
ogx:
use_as_library_client: true
config:
profile: ./profiles/openai-remote.yamlA relative profile: path resolves against the directory of the loaded
lightspeed-stack.yaml (not the current working directory); absolute paths
are used as-is. When profile is set, the baseline selector is ignored,
and library_client_config_path must not be set (unified and legacy inputs
are mutually exclusive).
The reference profiles are sanity-checked by the unit suite
(tests/unit/test_ogx_synthesize.py), so they stay loadable as the
synthesizer evolves.
When this mode is selected, OGX is started as a separate REST API service. All communication with OGX is performed via REST API calls, which means that OGX can run on a separate machine if needed.
Note
The REST API schema and semantics can change at any time, especially before version 1.0.0 is released. By using Lightspeed Core Service, developers, users, and customers stay isolated from these incompatibilities.
Three migration paths, per deployment:
| Path | Effort | Result |
|---|---|---|
| Do nothing | none | Legacy keeps working until removal in 0.7 (with a startup deprecation warning) |
| Lift-and-shift | seconds — --migrate-config |
Single file, byte-equivalent OGX behavior |
| Re-express | hours+ | Single file; high-level sections and/or a profile replace the lifted run.yaml |
Given a legacy pair — a hand-maintained run.yaml plus a
lightspeed-stack.yaml that points at it:
# lightspeed-stack.yaml (legacy, deprecated)
name: LCS
llama_stack:
use_as_library_client: true
library_client_config_path: ./run.yaml
# ... rest ...-
Run the migration tool:
lightspeed-stack --migrate-config \ --run-yaml run.yaml \ -c lightspeed-stack.yaml \ --migrate-output lightspeed-stack-unified.yaml
-
Inspect the output. Everything from your
lightspeed-stack.yamlis preserved; only thellama_stacksection changes —library_client_config_pathis removed and your entirerun.yamlis lifted into the unified config block:# lightspeed-stack-unified.yaml name: LCS ogx: use_as_library_client: true config: baseline: empty native_override: # ... your run.yaml content, verbatim ...
-
Replace literal secrets. If your
run.yamlcontained secret values directly, replace them with${env.MY_VAR}environment references — the migrated file otherwise carries them onto disk verbatim (the synthesized output is written owner-only, mode 0600, as a safety net). -
Swap the file in (
mv lightspeed-stack-unified.yaml lightspeed-stack.yaml), delete the now-unused externalrun.yamlmount/copy, and restart. OGX behavior is identical: synthesis starts from an empty baseline and deep-merges only your liftedrun.yaml.
Later, at your own pace, you can slim the native_override down by moving
providers into the high-level inference.providers section or into a
profile — that is the "re-express" path.
Unified mode shipped in release 0.6 with legacy mode fully functional plus a startup deprecation warning; the legacy two-file path is removed in release 0.7.
In this chapter it will be shown how to run LCS locally. This mode is especially useful for developers, as it is possible to work with the latest versions of source codes, including locally made changes and improvements. And last but not least, it is possible to trace, monitor and debug the entire system from within integrated development environment etc.
The easiest option is to run OGX in a separate process. This means that there will at least be two running processes involved:
- OGX framework with open port 8321 (can be easily changed if needed)
- LCS with open port 8080 (can be easily changed if needed)
- Python 3.12 or 3.13
piptool installedjqandcurltools installed
pip install --user uvsudo dnf install curl jq
-
Create a new directory outside of the lightspeed-stack project directory
mkdir /tmp/ogx-server
-
Copy the project file named
pyproject.ogx.tomlinto the new directory, renaming it topyproject.toml:cp examples/pyproject.ogx.toml /tmp/ogx-server/pyproject.toml
-
Run the following command to install all OGX dependencies in a new venv located in your new directory:
cd /tmp/ogx-server uv syncYou should get the following output:
Using CPython 3.12.10 interpreter at: /usr/bin/python3 Creating virtual environment at: .venv Resolved 136 packages in 1.90s Built sqlalchemy==2.0.42 Prepared 14 packages in 10.04s Installed 133 packages in 4.36s + accelerate==1.9.0 + aiohappyeyeballs==2.6.1 ... ... ... + transformers==4.54.0 + triton==3.3.1 + trl==0.20.0 + typing-extensions==4.14.1 + typing-inspection==0.4.1 + tzdata==2025.2 + urllib3==2.5.0 + uvicorn==0.35.0 + wcwidth==0.2.13 + wrapt==1.17.2 + xxhash==3.5.0 + yarl==1.20.1 + zipp==3.23.0
- In the next step, we need to verify that it is possible to run a tool called
ogx. It was installed into a Python virtual environment and therefore we have to run it viauv runcommand:uv run ogx
- If the installation was successful, the following messages should be displayed on the terminal:
usage: ogx [-h] {model,stack,download,verify-download} ... Welcome to the OGX CLI options: -h, --help show this help message and exit subcommands: {model,stack,download,verify-download} model Work with llama models stack Operations for the OGX / Distributions download Download a model from llama.meta.com or Hugging Face Hub verify-download Verify integrity of downloaded model files - If we try to run the OGX without configuring it, only the exception information is displayed (which is not very user-friendly):
Output:
uv run ogx stack run
INFO 2025-07-27 16:56:12,464 llama_stack.cli.stack.run:147 server: No image type or image name provided. Assuming environment packages. Traceback (most recent call last): File "/tmp/ramdisk/ogx-runner/.venv/bin/ogx", line 10, in <module> sys.exit(main()) ^^^^^^ File "/tmp/ramdisk/ogx-runner/.venv/lib64/python3.12/site-packages/llama_stack/cli/llama.py", line 53, in main parser.run(args) File "/tmp/ramdisk/ogx-runner/.venv/lib64/python3.12/site-packages/llama_stack/cli/llama.py", line 47, in run args.func(args) File "/tmp/ramdisk/ogx-runner/.venv/lib64/python3.12/site-packages/llama_stack/cli/stack/run.py", line 164, in _run_stack_run_cmd server_main(server_args) File "/tmp/ramdisk/ogx-runner/.venv/lib64/python3.12/site-packages/llama_stack/distribution/server/server.py", line 414, in main elif args.template: ^^^^^^^^^^^^^ AttributeError: 'Namespace' object has no attribute 'template'
OGX needs to be configured properly. For using the default runnable OGX a file named run.yaml needs to be created. Copy the example examples/run.yaml from the lightspeed-stack project directory into your OGX directory.
cp examples/run.yaml /tmp/ogx-server- Export OpenAI key by using the following command:
export OPENAI_API_KEY="sk-foo-bar-baz"
- Run the following command:
uv run ogx stack run run.yaml
- Check the output on terminal, it should look like:
INFO 2025-07-29 15:26:20,864 llama_stack.cli.stack.run:126 server: Using run configuration: run.yaml INFO 2025-07-29 15:26:20,877 llama_stack.cli.stack.run:147 server: No image type or image name provided. Assuming environment packages. INFO 2025-07-29 15:26:21,277 llama_stack.distribution.server.server:441 server: Using config file: run.yaml INFO 2025-07-29 15:26:21,279 llama_stack.distribution.server.server:443 server: Run configuration: INFO 2025-07-29 15:26:21,285 llama_stack.distribution.server.server:445 server: apis: - agents - datasetio - eval - inference - post_training - safety - scoring - telemetry - tool_runtime - vector_io benchmarks: [] container_image: null datasets: [] external_providers_dir: null image_name: minimal-viable-ogx-configuration inference_store: db_path: .llama/distributions/ollama/inference_store.db type: sqlite logging: null metadata_store: db_path: .llama/distributions/ollama/registry.db namespace: null type: sqlite models: - metadata: {} model_id: gpt-4-turbo model_type: !!python/object/apply:llama_stack.apis.models.models.ModelType - llm provider_id: openai provider_model_id: gpt-4-turbo providers: agents: - config: persistence_store: db_path: .llama/distributions/ollama/agents_store.db namespace: null type: sqlite responses_store: db_path: .llama/distributions/ollama/responses_store.db type: sqlite provider_id: meta-reference provider_type: inline::meta-reference datasetio: - config: kvstore: db_path: .llama/distributions/ollama/huggingface_datasetio.db namespace: null type: sqlite provider_id: huggingface provider_type: remote::huggingface - config: kvstore: db_path: .llama/distributions/ollama/localfs_datasetio.db namespace: null type: sqlite provider_id: localfs provider_type: inline::localfs eval: - config: kvstore: db_path: .llama/distributions/ollama/meta_reference_eval.db namespace: null type: sqlite provider_id: meta-reference provider_type: inline::meta-reference inference: - config: api_key: '********' provider_id: openai provider_type: remote::openai post_training: - config: checkpoint_format: huggingface device: cpu distributed_backend: null provider_id: huggingface provider_type: inline::huggingface safety: - config: excluded_categories: [] provider_id: llama-guard provider_type: inline::llama-guard scoring: - config: {} provider_id: basic provider_type: inline::basic - config: {} provider_id: llm-as-judge provider_type: inline::llm-as-judge - config: openai_api_key: '********' provider_id: braintrust provider_type: inline::braintrust telemetry: - config: service_name: lightspeed-stack sinks: sqlite sqlite_db_path: .llama/distributions/ollama/trace_store.db provider_id: meta-reference provider_type: inline::meta-reference tool_runtime: - config: {} provider_id: model-context-protocol provider_type: remote::model-context-protocol vector_io: - provider_id: faiss provider_type: inline::faiss config: persistence: namespace: vector_io::faiss backend: kv_default storage: backends: kv_default: type: kv_sqlite db_path: .llama/distributions/ollama/kv_store.db scoring_fns: [] server: auth: null host: null port: 8321 quota: null tls_cafile: null tls_certfile: null tls_keyfile: null shields: [] tool_groups: [] vector_stores: [] version: 2 - The server with OGX listens on port 8321. A description of the REST API is available in the form of OpenAPI (endpoint /openapi.json), but other endpoints can also be used. It is possible to check if OGX runs as REST API server by retrieving its version. We use
curlandjqtools for this purposes:The output should be in this form:curl localhost:8321/v1/version | jq .
{ "version": "0.2.22" }
Copy the examples/lightspeed-stack-lls-external.yaml file to your OGX project directory, naming it lightspeed-stack.yaml:
cp examples/lightspeed-stack-lls-external.yaml lightspeed-stack.yaml`make runuv run opentelemetry-instrument python3.12 src/lightspeed_stack.py
[07/29/25 15:43:35] INFO Initializing app main.py:19
INFO Including routers main.py:68
INFO: Started server process [1922983]
INFO: Waiting for application startup.
INFO Registering MCP servers main.py:81
DEBUG No MCP servers configured, skipping registration common.py:36
INFO Setting up model metrics main.py:84
[07/29/25 15:43:35] DEBUG Set provider/model configuration for openai/gpt-4-turbo to 0 utils.py:45
INFO App startup complete main.py:86
INFO: Application startup complete.
INFO: Uvicorn running on http://localhost:8080 (Press CTRL+C to quit)
curl localhost:8080/v1/models | jq .{
"models": [
{
"identifier": "gpt-4-turbo",
"metadata": {},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "gpt-4-turbo",
"model_type": "llm"
}
]
}It is possible to run Lightspeed Core Stack service with OGX "embedded" as a Python library. This means that just one process will be running and only one port (for example 8080) will be accessible.
- Python 3.12 or 3.13
piptool installedjqandcurltools installed
pip install --user uvsudo dnf install curl jq
- Clone LCS repository
- Add and install all required dependencies
uv sync --group ogxlibdev
OGX needs to be configured properly. Copy the example config from examples/run.yaml to the project directory:
cp examples/run.yaml .Copy the example LCS config file from examples/lightspeed-stack-lls-library.yaml to the project directory:
cp examples/lightspeed-stack-lls-library.yaml lightspeed-stack.yamlThe example is a unified-mode configuration: the run.yaml you created above
is consumed as the synthesis profile via
ogx.config.profile — there is no deprecated
library_client_config_path in it.
- Export OpenAI key by using the following command:
export OPENAI_API_KEY="sk-foo-bar-baz"
- Run the following command
make run
- Check the output
uv run opentelemetry-instrument python3.12 src/lightspeed_stack.py Using config run.yaml: apis: - agents - datasetio - eval - inference - post_training - safety - scoring - telemetry - tool_runtime - vector_io [07/30/25 20:01:53] INFO Initializing app main.py:19 [07/30/25 20:01:54] INFO Including routers main.py:68 INFO Registering MCP servers main.py:81 DEBUG No MCP servers configured, skipping registration common.py:36 INFO Setting up model metrics main.py:84 [07/30/25 20:01:54] DEBUG Set provider/model configuration for openai/openai/chatgpt-4o-latest to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-3.5-turbo to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-3.5-turbo-0125 to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-3.5-turbo-instruct to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-4 to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-4-turbo to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-4o to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-4o-2024-08-06 to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-4o-audio-preview to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/gpt-4o-mini to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/o1 to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/o1-mini to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/o3-mini to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/o4-mini to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/text-embedding-3-large to 0 utils.py:45 DEBUG Set provider/model configuration for openai/openai/text-embedding-3-small to 0 utils.py:45 INFO App startup complete main.py:86
curl localhost:8080/v1/models | jq .{
"models": [
{
"identifier": "gpt-4-turbo",
"metadata": {},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "gpt-4-turbo",
"model_type": "llm"
}
]
}The image with Lightspeed Core Stack allow users to run the service in two modes. In the first mode, the OGX runs in separate process - in a container or as a local or remote process. OGX functions are accessible via exposed TCP port. In the second model, the OGX is used as a standard Python library which means, that only the Lightspeed Core Stack image is needed and no other packages nor tools need to be installed.
First, it is needed to get an image containing the Lightspeed Core Stack service and all the necessary libraries on which the service depends. It is possible to use the stable release version (like 1.0.0, or "latest" stable), latest development version, or development version identified by a date + SHA (that image is built for any merged pull request).
podmaninstalled and configured properly
Note
It is possible to use docker instead of podman, but this use case is not tested and thus not supported.
Stable release images are tagged with versions like 0.1.0. Tag latest always points to the latest stable release.
Development images are build from main branch every time a new pull request is merged. Image tags for dev images use
the template dev-YYYYMMMDDD-SHORT_SHA e.g. dev-20250704-eaa27fb.
Tag dev-latest always points to the latest dev image built from latest git.
To retrieve the latest dev image, use the following command:
podman pull quay.io/lightspeed-core/lightspeed-stack:dev-latestIt should get the image, copy all layers, and write manifest:
Trying to pull quay.io/lightspeed-core/lightspeed-stack:dev-latest...
Getting image source signatures
Copying blob 455d71b0a12b done |
Copying blob d8e516fe2a03 done |
Copying blob a299c213c55c done |
Copying config 4468f47593 done |
Writing manifest to image destination
4468f475931a54ad1e5c26270ff4c3e55ec31444c1b0bf8fb77a576db7ab33f1
To retrieve stable version 0.2.0, use the following command:
podman pull quay.io/lightspeed-core/lightspeed-stack:0.2.0Trying to pull quay.io/lightspeed-core/lightspeed-stack:0.2.0...
Getting image source signatures
Copying blob 7c9e86f872c9 done |
Copying blob 455d71b0a12b skipped: already exists
Copying blob a299c213c55c skipped: already exists
Copying config a4982f4319 done |
Writing manifest to image destination
a4982f43195537b9eb1cec510fe6655f245d6d4b7236a4759808115d5d719972
Lightspeed Core Stack image can run LCS service that connects to OGX running in a separate process. This means that there will at least be two running processes involved:
- OGX framework with open port 8321 (can be easily changed if needed)
- Image with LCS (running in a container) with open port 8080 mapped to local port 8080 (can be easily changed if needed)
Note
Please note that LCS service will be run in a container. OGX itself can be run in a container, in separate local process, or on external machine. It is just needed to know the URL (including TCP port) to connect to OGX. [!INFO] If OGX is started from a container or is running on separate machine, you can skip next parts - it is expected that everything is setup accordingly.
- Python 3.12 or 3.13
piptool installedjqandcurltools installed
pip install --user uvsudo dnf install curl jq
-
Create a new directory
mkdir ogx-server cd ogx-server -
Create project file named
pyproject.tomlin this directory. This file should have the following content:[project] name = "ogx-demo" version = "0.1.0" description = "Default template for PDM package" authors = [] dependencies = [ "ogx==1.2.5", "fastapi>=0.115.12", "opentelemetry-sdk>=1.34.0", "opentelemetry-exporter-otlp>=1.34.0", "opentelemetry-instrumentation>=0.55b0", "aiosqlite>=0.21.0", "litellm>=1.72.1", "uvicorn>=0.34.3", "blobfile>=3.0.0", "datasets>=3.6.0", "sqlalchemy>=2.0.41", "faiss-cpu>=1.11.0", "mcp>=1.9.4", "autoevals>=0.0.129", "psutil>=7.0.0", "torch>=2.7.1", "peft>=0.15.2", "trl>=0.18.2"] requires-python = "==3.12.*" readme = "README.md" license = {text = "MIT"} [tool.pdm] distribution = false
-
Run the following command to install all dependencies:
uv sync
You should get the following output:
Using CPython 3.12.10 interpreter at: /usr/bin/python3 Creating virtual environment at: .venv Resolved 136 packages in 1.90s Built sqlalchemy==2.0.42 Prepared 14 packages in 10.04s Installed 133 packages in 4.36s + accelerate==1.9.0 + aiohappyeyeballs==2.6.1 ... ... ... + transformers==4.54.0 + triton==3.3.1 + trl==0.20.0 + typing-extensions==4.14.1 + typing-inspection==0.4.1 + tzdata==2025.2 + urllib3==2.5.0 + uvicorn==0.35.0 + wcwidth==0.2.13 + wrapt==1.17.2 + xxhash==3.5.0 + yarl==1.20.1 + zipp==3.23.0
- In the next step, we need to verify that it is possible to run a tool called
ogx. It was installed into a Python virtual environment and therefore we have to run it viauv runcommand:uv run ogx
- If the installation was successful, the following messages should be displayed on the terminal:
usage: ogx [-h] {model,stack,download,verify-download} ... Welcome to the OGX CLI options: -h, --help show this help message and exit subcommands: {model,stack,download,verify-download} model Work with llama models stack Operations for the OGX / Distributions download Download a model from llama.meta.com or Hugging Face Hub verify-download Verify integrity of downloaded model files - If we try to run the OGX without configuring it, only the exception information is displayed (which is not very user-friendly):
Output:
uv run ogx stack run
INFO 2025-07-27 16:56:12,464 llama_stack.cli.stack.run:147 server: No image type or image name provided. Assuming environment packages. Traceback (most recent call last): File "/tmp/ramdisk/ogx-runner/.venv/bin/ogx", line 10, in <module> sys.exit(main()) ^^^^^^ File "/tmp/ramdisk/ogx-runner/.venv/lib64/python3.12/site-packages/llama_stack/cli/llama.py", line 53, in main parser.run(args) File "/tmp/ramdisk/ogx-runner/.venv/lib64/python3.12/site-packages/llama_stack/cli/llama.py", line 47, in run args.func(args) File "/tmp/ramdisk/ogx-runner/.venv/lib64/python3.12/site-packages/llama_stack/cli/stack/run.py", line 164, in _run_stack_run_cmd server_main(server_args) File "/tmp/ramdisk/ogx-runner/.venv/lib64/python3.12/site-packages/llama_stack/distribution/server/server.py", line 414, in main elif args.template: ^^^^^^^^^^^^^ AttributeError: 'Namespace' object has no attribute 'template'
OGX needs to be configured properly. For using the default runnable OGX a file named run.yaml needs to be created. Use the example configuration from examples/run.yaml.
- Export OpenAI key by using the following command:
export OPENAI_API_KEY="sk-foo-bar-baz"
- Run the following command:
uv run ogx stack run run.yaml
- Check the output on terminal, it should look like:
INFO 2025-07-29 15:26:20,864 llama_stack.cli.stack.run:126 server: Using run configuration: run.yaml INFO 2025-07-29 15:26:20,877 llama_stack.cli.stack.run:147 server: No image type or image name provided. Assuming environment packages. INFO 2025-07-29 15:26:21,277 llama_stack.distribution.server.server:441 server: Using config file: run.yaml INFO 2025-07-29 15:26:21,279 llama_stack.distribution.server.server:443 server: Run configuration: INFO 2025-07-29 15:26:21,285 llama_stack.distribution.server.server:445 server: apis: - agents - datasetio - eval - inference - post_training - safety - scoring - telemetry - tool_runtime - vector_io benchmarks: [] container_image: null datasets: [] external_providers_dir: null image_name: minimal-viable-ogx-configuration inference_store: db_path: .llama/distributions/ollama/inference_store.db type: sqlite logging: null metadata_store: db_path: .llama/distributions/ollama/registry.db namespace: null type: sqlite models: - metadata: {} model_id: gpt-4-turbo model_type: !!python/object/apply:llama_stack.apis.models.models.ModelType - llm provider_id: openai provider_model_id: gpt-4-turbo providers: agents: - config: persistence_store: db_path: .llama/distributions/ollama/agents_store.db namespace: null type: sqlite responses_store: db_path: .llama/distributions/ollama/responses_store.db type: sqlite provider_id: meta-reference provider_type: inline::meta-reference datasetio: - config: kvstore: db_path: .llama/distributions/ollama/huggingface_datasetio.db namespace: null type: sqlite provider_id: huggingface provider_type: remote::huggingface - config: kvstore: db_path: .llama/distributions/ollama/localfs_datasetio.db namespace: null type: sqlite provider_id: localfs provider_type: inline::localfs eval: - config: kvstore: db_path: .llama/distributions/ollama/meta_reference_eval.db namespace: null type: sqlite provider_id: meta-reference provider_type: inline::meta-reference inference: - config: api_key: '********' provider_id: openai provider_type: remote::openai post_training: - config: checkpoint_format: huggingface device: cpu distributed_backend: null provider_id: huggingface provider_type: inline::huggingface safety: - config: excluded_categories: [] provider_id: llama-guard provider_type: inline::llama-guard scoring: - config: {} provider_id: basic provider_type: inline::basic - config: {} provider_id: llm-as-judge provider_type: inline::llm-as-judge - config: openai_api_key: '********' provider_id: braintrust provider_type: inline::braintrust telemetry: - config: service_name: lightspeed-stack sinks: sqlite sqlite_db_path: .llama/distributions/ollama/trace_store.db provider_id: meta-reference provider_type: inline::meta-reference tool_runtime: - config: {} provider_id: model-context-protocol provider_type: remote::model-context-protocol vector_io: - provider_id: faiss provider_type: inline::faiss config: persistence: namespace: vector_io::faiss backend: kv_default storage: backends: kv_default: type: kv_sqlite db_path: .llama/distributions/ollama/kv_store.db scoring_fns: [] server: auth: null host: null port: 8321 quota: null tls_cafile: null tls_certfile: null tls_keyfile: null shields: [] tool_groups: [] vector_stores: [] version: 2 - The server with OGX listens on port 8321. A description of the REST API is available in the form of OpenAPI (endpoint /openapi.json), but other endpoints can also be used. It is possible to check if OGX runs as REST API server by retrieving its version. We use
curlandjqtools for this purposes:The output should be in this form:curl localhost:8321/v1/version | jq .
{ "version": "0.2.22" }
Image with Lightspeed Core Stack needs to be configured properly. Create local file named lightspeed-stack.yaml with the following content:
name: Lightspeed Core Service (LCS)
service:
host: localhost
port: 8080
auth_enabled: false
workers: 1
color_log: true
access_log: true
ogx:
use_as_library_client: false
url: http://localhost:8321
api_key: xyzzy
user_data_collection:
feedback_enabled: true
feedback_storage: "/tmp/data/feedback"
transcripts_enabled: true
transcripts_storage: "/tmp/data/transcripts"
authentication:
module: "noop"Now it is needed to run Lightspeed Core Stack from within a container. The service needs to be configured, so lightspeed-stack.yaml has to be mounted into the container:
podman run -it --network host -v lightspeed-stack.yaml:/app-root/lightspeed-stack.yaml:Z quay.io/lightspeed-core/lightspeed-stack:dev-latestNote
Please note that --network host is insecure option. It is used there because LCS service running in a container have to access OGX running outside this container and the standard port mapping can not be leveraged there. This configuration would be ok for development purposes, but for real deployment, network needs to be reconfigured accordingly to maintain required container isolation!
OGX can be used as a library that is already part of OLS image. It means that no other processed needs to be started, but more configuration is required. Everything will be started from within the one container:
First, export your OpenAI key into environment variable:
export OPENAI_API_KEY="sk-foo-bar-baz-my-key"Create a file named run.yaml. Use the example configuration from examples/run.yaml.
Create file lightspeed-stack.yaml with the following content (unified
mode — the run.yaml created above is consumed as the synthesis
profile):
name: Lightspeed Core Service (LCS)
service:
host: localhost
port: 8080
auth_enabled: false
workers: 1
color_log: true
access_log: true
ogx:
use_as_library_client: true
config:
profile: ./run.yaml
api_key: xyzzy
user_data_collection:
feedback_enabled: true
feedback_storage: "/tmp/data/feedback"
transcripts_enabled: true
transcripts_storage: "/tmp/data/transcripts"
authentication:
module: "noop"Warning
The legacy equivalent — library_client_config_path: ./run.yaml instead
of the config: block — is deprecated and will be removed in release
0.7. See
Migrating from the legacy two-file configuration.
Now it is time to start the service from a container. It is needed to mount both configuration files lightspeed-stack.yaml and run.yaml into the container. And it is also needed to expose environment variable containing OpenAI key:
podman run -it -p 8080:8080 -v lightspeed-stack.yaml:/app-root/lightspeed-stack.yaml:Z -v ./run.yaml:/app-root/run.yaml:Z -e OPENAI_API_KEY=${OPENAI_API_KEY} quay.io/lightspeed-core/lightspeed-stack:dev-latestThe Lightspeed Core Stack service exposes its own REST API endpoints:
Open http://localhost:8080 URL in your web browser. The following front page should be displayed:
Click on Swagger UI link to open Swagger UI page:
List of all available REST API endpoints is displayed on this page. It is possible to interactively access any endpoint, specify query parameters, JSON payload etc. For example it is possible to access Info endpoint and see actual response from the Lightspeed Core Stack service:
Some REST API endpoints like /query requires payload to be send into the service. This payload should be represented in JSON format. Some attributes in JSON payload are optional, so it is possible to send just the question and system prompt. In this case the JSON payload should look like:
The response retrieved from LLM is displayed directly on Swagger UI page:
To access Lightspeed Core Stack service functions via REST API from command line, just the curl tool and optionally jq tool are needed. Any REST API endpoint can be accessed from command line.
For example, the /v1/info endpoint can be accessed without parameters using HTTP GET method:
curl http://localhost:8080/v1/info | jq .The response should look like:
{
"name": "Lightspeed Core Service (LCS)",
"version": "0.2.0"
}Use the /v1/models to get list of available models:
curl http://localhost:8080/v1/models | jq .Please note that actual response from the service is larger. It was stripped down in this guide:
{
"models": [
{
"identifier": "gpt-4-turbo",
"metadata": {},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "gpt-4-turbo",
"model_type": "llm"
},
{
"identifier": "openai/gpt-4o-mini",
"metadata": {},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "gpt-4o-mini",
"model_type": "llm"
},
{
"identifier": "openai/gpt-4o-audio-preview",
"metadata": {},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "gpt-4o-audio-preview",
"model_type": "llm"
},
{
"identifier": "openai/chatgpt-4o-latest",
"metadata": {},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "chatgpt-4o-latest",
"model_type": "llm"
},
{
"identifier": "openai/o1",
"metadata": {},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "o1",
"model_type": "llm"
},
{
"identifier": "openai/text-embedding-3-small",
"metadata": {
"embedding_dimension": 1536.0,
"context_length": 8192.0
},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "text-embedding-3-small",
"model_type": "llm"
},
{
"identifier": "openai/text-embedding-3-large",
"metadata": {
"embedding_dimension": 3072.0,
"context_length": 8192.0
},
"api_model_type": "llm",
"provider_id": "openai",
"type": "model",
"provider_resource_id": "text-embedding-3-large",
"model_type": "llm"
}
]
}To retrieve LLM response, the question (or query) needs to be send to inference model. Thus the HTTP POST method should be used:
$ curl -X 'POST' \
> 'http://localhost:8080/v1/query' \
> -H 'accept: application/json' \
> -H 'Content-Type: application/json' \
> -d '{
> "query": "write a deployment yaml for the mongodb image",
> "system_prompt": "You are a helpful assistant"
> }'Response should look like:
{
"conversation_id": "a731eaf2-0935-47ee-9661-2e9b36cda1f4",
"response": "Below is a basic example of a Kubernetes deployment YAML file for deploying a MongoDB instance using the official MongoDB Docker image. This YAML file defines a Deployment resource that manages a Pod with a single MongoDB container.\n\n```yaml\napiVersion: apps/v1\nkind: Deployment\nmetadata:\n name: mongodb-deployment\n labels:\n app: mongodb\nspec:\n replicas: 1\n selector:\n matchLabels:\n app: mongodb\n template:\n metadata:\n labels:\n app: mongodb\n spec:\n containers:\n - name: mongodb\n image: mongo:latest\n ports:\n - containerPort: 27017\n env:\n - name: MONGO_INITDB_ROOT_USERNAME\n value: \"mongoadmin\"\n - name: MONGO_INITDB_ROOT_PASSWORD\n value: \"mongopass\"\n volumeMounts:\n - name: mongodb-data\n mountPath: /data/db\n volumes:\n - name: mongodb-data\n persistentVolumeClaim:\n claimName: mongodb-pvc\n---\napiVersion: v1\nkind: Service\nmetadata:\n name: mongodb-service\nspec:\n ports:\n - port: 27017\n targetPort: 27017\n selector:\n app: mongodb\n type: ClusterIP\n---\napiVersion: v1\nkind: PersistentVolumeClaim\nmetadata:\n name: mongodb-pvc\nspec:\n accessModes:\n - ReadWriteOnce\n resources:\n requests:\n storage: 1Gi\n```\n\n### Explanation of the YAML Components:\n\n1. **Deployment**:\n - `apiVersion: apps/v1`: Specifies the API version for the Deployment.\n - `kind: Deployment`: Specifies that this is a Deployment resource.\n - `metadata`: Metadata about the deployment, such as its name.\n - `spec`: Specification of the deployment.\n - `replicas`: Number of desired pods.\n - `selector`: Selector for pod targeting.\n - `template`: Template for the pod.\n - `containers`: List of containers within the pod.\n - `image`: Docker image of MongoDB.\n - `ports`: Container port MongoDB server listens on.\n - `env`: Environment variables for MongoDB credentials.\n - `volumeMounts`: Mount points for volumes inside the container.\n\n2. **Service**:\n - `apiVersion: v1`: Specifies the API version for the Service.\n - `kind: Service`: Specifies that this is a Service resource.\n - `metadata`: Metadata about the service, such as its name.\n - `spec`: Specification of the service.\n - `ports`: Ports the service exposes.\n - `selector`: Selector for service targeting.\n - `type`: Type of service, `ClusterIP` for internal access.\n\n3. **PersistentVolumeClaim (PVC)**:\n - `apiVersion: v1`: Specifies the API version for the PVC.\n - `kind: PersistentVolumeClaim`: Specifies that this is a PVC resource.\n - `metadata`: Metadata about the PVC, such as its name.\n - `spec`: Specification of the PVC.\n - `accessModes`: Access modes for the volume.\n - `resources`: Resources requests for the volume.\n\nThis setup ensures that MongoDB data persists across pod restarts and provides a basic internal service for accessing MongoDB within the cluster. Adjust the storage size, MongoDB version, and credentials as necessary for your specific requirements."
}Note
As is shown on the previous example, the output might contain endlines, Markdown marks etc.


