From 047fa67857e3cfc5067841e7628ee1d585dce747 Mon Sep 17 00:00:00 2001 From: Maxim Svistunov Date: Fri, 4 Sep 2026 15:07:02 +0200 Subject: [PATCH 1/2] LCORE-1673: e2e feature file for conversation compaction (no step implementation) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Author the behave scenarios for conversation compaction from the design document alone, ahead of the step definitions (LCORE-2230), so the test shape follows the intended behaviour rather than the implementation. The scenarios observe compaction strictly from outside the deployed stack — e2e steps never touch src/ (docs/testing/e2e_testing.md, "Choosing the Test Layer") — which leaves exactly the surfaces the design exposes over HTTP: - context_status on /v1/query responses: "full" while compaction never triggers, "summarized" once a three-turn conversation crosses the configured threshold (R7, driven by R1/R9 configuration). - The assistant still recalls a fact stated before the summary. - The Conversations API keeps serving every user query after compaction (R6). - /v1/streaming_query emits a compaction event before the first token and carries context_status in its end event (R12, R7). - enabled: false is a full off-switch: context_status stays "full" past the threshold. Buffer semantics, additive summarization, the summarization model and prompt, tiktoken estimation and per-conversation blocking are internal and belong to the integration layer (LCORE-1574). Fixtures: lightspeed-stack-compaction.yaml and lightspeed-stack-compaction-disabled.yaml per mode, derived from the mode's base configuration. They register a deliberately small context window for the CI model (openai/gpt-4o-mini: 2000 tokens) and a 10% threshold above a 100-token floor, so three turns cross the trigger; the real provider window is untouched, this only drives the local estimate. The window key is model-specific, so the scenarios that need it carry @openai-only, the same gating the unified-mode boot scenarios use on the providers matrix. tests/e2e/test_list.txt gains the feature; behave --dry-run parses it with the nine new step patterns reported as undefined and every other step matched. --- .../lightspeed-stack-compaction-disabled.yaml | 62 +++++++ .../lightspeed-stack-compaction.yaml | 61 +++++++ .../lightspeed-stack-compaction-disabled.yaml | 60 +++++++ .../lightspeed-stack-compaction.yaml | 59 +++++++ .../features/conversation-compaction.feature | 165 ++++++++++++++++++ tests/e2e/test_list.txt | 1 + 6 files changed, 408 insertions(+) create mode 100644 tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml create mode 100644 tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml create mode 100644 tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml create mode 100644 tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml create mode 100644 tests/e2e/features/conversation-compaction.feature diff --git a/tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml b/tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml new file mode 100644 index 000000000..a38ef5e9d --- /dev/null +++ b/tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml @@ -0,0 +1,62 @@ +name: Lightspeed Core Service (LCS) +service: + host: 0.0.0.0 + port: 8080 + auth_enabled: false + workers: 1 + color_log: true + access_log: true +ogx: + # Library mode - embeds OGX as library + use_as_library_client: true + # Unified mode: run.yaml (materialized per provider by CI/the harness) + # is consumed as the synthesis profile instead of the legacy two-file path. + config: + profile: run.yaml +user_data_collection: + feedback_enabled: true + feedback_storage: "/tmp/data/feedback" + transcripts_enabled: true + transcripts_storage: "/tmp/data/transcripts" +authentication: + module: "noop" +inference: + default_provider: openai + default_model: gpt-4o-mini + # Compaction e2e (LCORE-1673): a deliberately small window for the CI + # model so a three-turn conversation crosses the trigger threshold. + # The real model window is far larger; this only drives the local + # estimate, never the provider. + context_windows: + openai/gpt-4o-mini: 2000 +rag: + byok: + stores: + - rag_id: e2e-test-docs + backend: faiss + embedding_model: sentence-transformers/all-mpnet-base-v2 + embedding_dimension: 768 + vector_db_id: ${env.FAISS_VECTOR_STORE_ID} + db_path: ${env.KV_RAG_PATH:=~/.llama/storage/rag/kv_store.db} + score_multiplier: 1.0 + retrieval: + tool: + sources: + - e2e-test-docs + +shields: + - name: pii-redaction + provider_id: redaction + config: + rules: + - pattern: '\d+' + replacement: '[NUM]' + +# Same small window and threshold as lightspeed-stack-compaction.yaml, but +# compaction switched off: context_status must stay "full" past the +# threshold (enabled is a full off-switch). +compaction: + enabled: false + threshold_ratio: 0.1 + token_floor: 100 + buffer_turns: 1 diff --git a/tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml b/tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml new file mode 100644 index 000000000..347058e1b --- /dev/null +++ b/tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml @@ -0,0 +1,61 @@ +name: Lightspeed Core Service (LCS) +service: + host: 0.0.0.0 + port: 8080 + auth_enabled: false + workers: 1 + color_log: true + access_log: true +ogx: + # Library mode - embeds OGX as library + use_as_library_client: true + # Unified mode: run.yaml (materialized per provider by CI/the harness) + # is consumed as the synthesis profile instead of the legacy two-file path. + config: + profile: run.yaml +user_data_collection: + feedback_enabled: true + feedback_storage: "/tmp/data/feedback" + transcripts_enabled: true + transcripts_storage: "/tmp/data/transcripts" +authentication: + module: "noop" +inference: + default_provider: openai + default_model: gpt-4o-mini + # Compaction e2e (LCORE-1673): a deliberately small window for the CI + # model so a three-turn conversation crosses the trigger threshold. + # The real model window is far larger; this only drives the local + # estimate, never the provider. + context_windows: + openai/gpt-4o-mini: 2000 +rag: + byok: + stores: + - rag_id: e2e-test-docs + backend: faiss + embedding_model: sentence-transformers/all-mpnet-base-v2 + embedding_dimension: 768 + vector_db_id: ${env.FAISS_VECTOR_STORE_ID} + db_path: ${env.KV_RAG_PATH:=~/.llama/storage/rag/kv_store.db} + score_multiplier: 1.0 + retrieval: + tool: + sources: + - e2e-test-docs + +shields: + - name: pii-redaction + provider_id: redaction + config: + rules: + - pattern: '\d+' + replacement: '[NUM]' + +# Compaction on with a low threshold: 10% of the 2000-token window, +# above a 100-token floor, keeping one recent turn verbatim. +compaction: + enabled: true + threshold_ratio: 0.1 + token_floor: 100 + buffer_turns: 1 diff --git a/tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml b/tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml new file mode 100644 index 000000000..86b08fb7e --- /dev/null +++ b/tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml @@ -0,0 +1,60 @@ +name: Lightspeed Core Service (LCS) +service: + host: 0.0.0.0 + port: 8080 + auth_enabled: false + workers: 1 + color_log: true + access_log: true +ogx: + # Server mode - connects to separate OGX service + use_as_library_client: false + url: http://${env.E2E_LLAMA_HOSTNAME}:8321 + api_key: xyzzy +user_data_collection: + feedback_enabled: true + feedback_storage: "/tmp/data/feedback" + transcripts_enabled: true + transcripts_storage: "/tmp/data/transcripts" +authentication: + module: "noop" +inference: + default_provider: openai + default_model: gpt-4o-mini + # Compaction e2e (LCORE-1673): a deliberately small window for the CI + # model so a three-turn conversation crosses the trigger threshold. + # The real model window is far larger; this only drives the local + # estimate, never the provider. + context_windows: + openai/gpt-4o-mini: 2000 +rag: + byok: + stores: + - rag_id: e2e-test-docs + backend: faiss + embedding_model: sentence-transformers/all-mpnet-base-v2 + embedding_dimension: 768 + vector_db_id: ${env.FAISS_VECTOR_STORE_ID} + db_path: ${env.KV_RAG_PATH:=~/.llama/storage/rag/kv_store.db} + score_multiplier: 1.0 + retrieval: + tool: + sources: + - e2e-test-docs + +shields: + - name: pii-redaction + provider_id: redaction + config: + rules: + - pattern: '\d+' + replacement: '[NUM]' + +# Same small window and threshold as lightspeed-stack-compaction.yaml, but +# compaction switched off: context_status must stay "full" past the +# threshold (enabled is a full off-switch). +compaction: + enabled: false + threshold_ratio: 0.1 + token_floor: 100 + buffer_turns: 1 diff --git a/tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml b/tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml new file mode 100644 index 000000000..c2a156586 --- /dev/null +++ b/tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml @@ -0,0 +1,59 @@ +name: Lightspeed Core Service (LCS) +service: + host: 0.0.0.0 + port: 8080 + auth_enabled: false + workers: 1 + color_log: true + access_log: true +ogx: + # Server mode - connects to separate OGX service + use_as_library_client: false + url: http://${env.E2E_LLAMA_HOSTNAME}:8321 + api_key: xyzzy +user_data_collection: + feedback_enabled: true + feedback_storage: "/tmp/data/feedback" + transcripts_enabled: true + transcripts_storage: "/tmp/data/transcripts" +authentication: + module: "noop" +inference: + default_provider: openai + default_model: gpt-4o-mini + # Compaction e2e (LCORE-1673): a deliberately small window for the CI + # model so a three-turn conversation crosses the trigger threshold. + # The real model window is far larger; this only drives the local + # estimate, never the provider. + context_windows: + openai/gpt-4o-mini: 2000 +rag: + byok: + stores: + - rag_id: e2e-test-docs + backend: faiss + embedding_model: sentence-transformers/all-mpnet-base-v2 + embedding_dimension: 768 + vector_db_id: ${env.FAISS_VECTOR_STORE_ID} + db_path: ${env.KV_RAG_PATH:=~/.llama/storage/rag/kv_store.db} + score_multiplier: 1.0 + retrieval: + tool: + sources: + - e2e-test-docs + +shields: + - name: pii-redaction + provider_id: redaction + config: + rules: + - pattern: '\d+' + replacement: '[NUM]' + +# Compaction on with a low threshold: 10% of the 2000-token window, +# above a 100-token floor, keeping one recent turn verbatim. +compaction: + enabled: true + threshold_ratio: 0.1 + token_floor: 100 + buffer_turns: 1 diff --git a/tests/e2e/features/conversation-compaction.feature b/tests/e2e/features/conversation-compaction.feature new file mode 100644 index 000000000..cd00ca03a --- /dev/null +++ b/tests/e2e/features/conversation-compaction.feature @@ -0,0 +1,165 @@ +@cfg_compaction +Feature: Conversation compaction + + When a conversation's estimated input approaches the model's context + window, older turns are summarized before the request reaches the LLM + (docs/design/conversation-compaction/conversation-compaction.md). These + scenarios observe compaction from outside only: the context_status field + on responses (R7), the compaction event on the native stream (R12), the + full history the Conversations API keeps serving (R6), and the assistant's + recall of what was said before the summary. The trigger is driven by the + admin configuration (R1, R9): the compaction fixtures register a small + context window for the openai model and a low threshold ratio, so a + three-turn conversation crosses it. + + Background: + Given The service is started locally + And The system is in default state + And REST API service prefix is /v1 + And the Lightspeed stack configuration directory is "tests/e2e/configuration" + + + Scenario: context_status reports full while compaction never triggers + Given The service uses the lightspeed-stack.yaml configuration + And The service is restarted + When I use "query" to ask question + """ + {"query": "Say hello", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And The response context_status is "full" + + + @openai-only + Scenario: context_status reports summarized once the conversation crosses the threshold + Given The service uses the lightspeed-stack-compaction.yaml configuration + And The service is restarted + When I use "query" to ask question + """ + {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And I store conversation details + When I use "query" to ask question with same conversation_id + """ + {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + When I use "query" to ask question with same conversation_id + """ + {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And The response context_status is "summarized" + + + @openai-only + Scenario: the assistant still recalls what was said before the summary + Given The service uses the lightspeed-stack-compaction.yaml configuration + And The service is restarted + When I use "query" to ask question + """ + {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And I store conversation details + When I use "query" to ask question with same conversation_id + """ + {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + When I use "query" to ask question with same conversation_id + """ + {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And The response context_status is "summarized" + When I use "query" to ask question with same conversation_id + """ + {"query": "What is the name of my cluster? Reply with the name only.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And The response context_status is "summarized" + And The response contains following fragments + | Fragments in LLM response | + | aurora-prod-7 | + + + @openai-only + Scenario: the full conversation history stays available after compaction + Given The service uses the lightspeed-stack-compaction.yaml configuration + And The service is restarted + When I use "query" to ask question + """ + {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And I store conversation details + When I use "query" to ask question with same conversation_id + """ + {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + When I use "query" to ask question with same conversation_id + """ + {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And The response context_status is "summarized" + When I use REST API conversation endpoint with conversation_id from above using HTTP GET method + Then The status code of the response is 200 + And The conversation history includes the following user queries + | User query | + | My OpenShift cluster is named aurora-prod-7. Remember that name. | + | Explain what a pod is in about five sentences. | + | Explain what a deployment is in about five sentences. | + + + @openai-only + Scenario: the native stream announces compaction and reports context_status + Given The service uses the lightspeed-stack-compaction.yaml configuration + And The service is restarted + When I use "streaming_query" to ask question + """ + {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And I wait for the response to be completed + And I store conversation details + When I use "streaming_query" to ask question with same conversation_id + """ + {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And I wait for the response to be completed + When I use "streaming_query" to ask question with same conversation_id + """ + {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And I wait for the response to be completed + And The streamed response contains a compaction event before the first token + And The streamed response end event has context_status "summarized" + + + @openai-only + Scenario: compaction stays off when disabled, even past the threshold + Given The service uses the lightspeed-stack-compaction-disabled.yaml configuration + And The service is restarted + When I use "query" to ask question + """ + {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And I store conversation details + When I use "query" to ask question with same conversation_id + """ + {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + When I use "query" to ask question with same conversation_id + """ + {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + """ + Then The status code of the response is 200 + And The response context_status is "full" diff --git a/tests/e2e/test_list.txt b/tests/e2e/test_list.txt index 79f34fcb5..c37e04e98 100644 --- a/tests/e2e/test_list.txt +++ b/tests/e2e/test_list.txt @@ -23,6 +23,7 @@ features/rlsapi_v1.feature features/streaming_query.feature features/vector_stores.feature features/conversation_cache_v2.feature +features/conversation-compaction.feature features/feedback.feature features/http_401_unauthorized.feature features/rbac.feature From a547bf2791b32643aee6dbee48c55998b4c6180f Mon Sep 17 00:00:00 2001 From: Maxim Svistunov Date: Fri, 11 Sep 2026 15:01:35 +0200 Subject: [PATCH 2/2] LCORE-1673: make the compaction crossing deterministic and cut the feature to three scenarios Review rework of the conversation-compaction feature file. Scenarios. The six scenarios collapse into three. The base-config "full" scenario duplicated the disabled one and is gone. The "summarized", "recall" and "history" scenarios exercised the same three turns and are now one scenario; the streaming and disabled scenarios stay separate because they observe different surfaces (the SSE stream, the off switch). Deterministic trigger. The old scenarios relied on three ordinary queries adding up to the threshold, so the crossing turn depended on response length. Now the first two turns ask the model to reply with "OK" only and each plants one fact, and the third query is a fixed 248-token paragraph that crosses the 200-token threshold on its own (2000-token window, ratio 0.1, floor 100). With buffer_turns 1 nothing can be summarized before the third turn, so context_status is asserted "full" on turns one and two and "summarized" on turn three. The third query also asks for both facts: turn one is in the summary, turn two is the verbatim buffer, so one answer covers summary recall and buffer recall. Fixtures. The pii-redaction shield is dropped: it cost tokens on every turn and its \d+ rule would have rewritten the cluster name the recall assertion looks for. context_windows now lists every provider/model pair the e2e workflows run (openai, azure, google-vertex, watsonx, aws-bedrock) instead of only openai, and the @openai-only tags go away with it. The vLLM runs take their model id from an env var and context_windows keys are not env-substituted, so they are not listed. Skip. The feature is tagged @skip at feature level until LCORE-2230 lands the nine step definitions; Konflux runs the whole test list and would fail on undefined steps. The description drops the R-references. Rebased onto main after the OGX rename: the server-mode fixtures now read the OGX host from E2E_OGX_HOSTNAME like the base fixture does. --- .../lightspeed-stack-compaction-disabled.yaml | 23 ++-- .../lightspeed-stack-compaction.yaml | 23 ++-- .../lightspeed-stack-compaction-disabled.yaml | 25 ++-- .../lightspeed-stack-compaction.yaml | 25 ++-- .../features/conversation-compaction.feature | 121 +++++------------- 5 files changed, 75 insertions(+), 142 deletions(-) diff --git a/tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml b/tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml index a38ef5e9d..227792a2f 100644 --- a/tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml +++ b/tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml @@ -23,12 +23,19 @@ authentication: inference: default_provider: openai default_model: gpt-4o-mini - # Compaction e2e (LCORE-1673): a deliberately small window for the CI - # model so a three-turn conversation crosses the trigger threshold. - # The real model window is far larger; this only drives the local - # estimate, never the provider. + # Compaction e2e (LCORE-1673): a deliberately small window for every + # model the e2e workflows run against, so the third query of the + # compaction scenarios crosses the trigger threshold. The real windows + # are far larger; this only drives the local estimate, never the provider. + # The vLLM-backed runs (rhaiis, rhelai) take the model id from an env + # var, and context_windows keys are not env-substituted, so they are + # not listed and skip the token-based trigger. context_windows: openai/gpt-4o-mini: 2000 + azure/gpt-4o-mini: 2000 + google-vertex/publishers/google/models/gemini-2.5-flash: 2000 + watsonx/meta-llama/llama-3-3-70b-instruct: 2000 + aws-bedrock/deepseek.v3-v1:0: 2000 rag: byok: stores: @@ -44,14 +51,6 @@ rag: sources: - e2e-test-docs -shields: - - name: pii-redaction - provider_id: redaction - config: - rules: - - pattern: '\d+' - replacement: '[NUM]' - # Same small window and threshold as lightspeed-stack-compaction.yaml, but # compaction switched off: context_status must stay "full" past the # threshold (enabled is a full off-switch). diff --git a/tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml b/tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml index 347058e1b..94675e16c 100644 --- a/tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml +++ b/tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml @@ -23,12 +23,19 @@ authentication: inference: default_provider: openai default_model: gpt-4o-mini - # Compaction e2e (LCORE-1673): a deliberately small window for the CI - # model so a three-turn conversation crosses the trigger threshold. - # The real model window is far larger; this only drives the local - # estimate, never the provider. + # Compaction e2e (LCORE-1673): a deliberately small window for every + # model the e2e workflows run against, so the third query of the + # compaction scenarios crosses the trigger threshold. The real windows + # are far larger; this only drives the local estimate, never the provider. + # The vLLM-backed runs (rhaiis, rhelai) take the model id from an env + # var, and context_windows keys are not env-substituted, so they are + # not listed and skip the token-based trigger. context_windows: openai/gpt-4o-mini: 2000 + azure/gpt-4o-mini: 2000 + google-vertex/publishers/google/models/gemini-2.5-flash: 2000 + watsonx/meta-llama/llama-3-3-70b-instruct: 2000 + aws-bedrock/deepseek.v3-v1:0: 2000 rag: byok: stores: @@ -44,14 +51,6 @@ rag: sources: - e2e-test-docs -shields: - - name: pii-redaction - provider_id: redaction - config: - rules: - - pattern: '\d+' - replacement: '[NUM]' - # Compaction on with a low threshold: 10% of the 2000-token window, # above a 100-token floor, keeping one recent turn verbatim. compaction: diff --git a/tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml b/tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml index 86b08fb7e..ff1e3510c 100644 --- a/tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml +++ b/tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml @@ -9,7 +9,7 @@ service: ogx: # Server mode - connects to separate OGX service use_as_library_client: false - url: http://${env.E2E_LLAMA_HOSTNAME}:8321 + url: http://${env.E2E_OGX_HOSTNAME}:8321 api_key: xyzzy user_data_collection: feedback_enabled: true @@ -21,12 +21,19 @@ authentication: inference: default_provider: openai default_model: gpt-4o-mini - # Compaction e2e (LCORE-1673): a deliberately small window for the CI - # model so a three-turn conversation crosses the trigger threshold. - # The real model window is far larger; this only drives the local - # estimate, never the provider. + # Compaction e2e (LCORE-1673): a deliberately small window for every + # model the e2e workflows run against, so the third query of the + # compaction scenarios crosses the trigger threshold. The real windows + # are far larger; this only drives the local estimate, never the provider. + # The vLLM-backed runs (rhaiis, rhelai) take the model id from an env + # var, and context_windows keys are not env-substituted, so they are + # not listed and skip the token-based trigger. context_windows: openai/gpt-4o-mini: 2000 + azure/gpt-4o-mini: 2000 + google-vertex/publishers/google/models/gemini-2.5-flash: 2000 + watsonx/meta-llama/llama-3-3-70b-instruct: 2000 + aws-bedrock/deepseek.v3-v1:0: 2000 rag: byok: stores: @@ -42,14 +49,6 @@ rag: sources: - e2e-test-docs -shields: - - name: pii-redaction - provider_id: redaction - config: - rules: - - pattern: '\d+' - replacement: '[NUM]' - # Same small window and threshold as lightspeed-stack-compaction.yaml, but # compaction switched off: context_status must stay "full" past the # threshold (enabled is a full off-switch). diff --git a/tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml b/tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml index c2a156586..82dd0c32b 100644 --- a/tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml +++ b/tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml @@ -9,7 +9,7 @@ service: ogx: # Server mode - connects to separate OGX service use_as_library_client: false - url: http://${env.E2E_LLAMA_HOSTNAME}:8321 + url: http://${env.E2E_OGX_HOSTNAME}:8321 api_key: xyzzy user_data_collection: feedback_enabled: true @@ -21,12 +21,19 @@ authentication: inference: default_provider: openai default_model: gpt-4o-mini - # Compaction e2e (LCORE-1673): a deliberately small window for the CI - # model so a three-turn conversation crosses the trigger threshold. - # The real model window is far larger; this only drives the local - # estimate, never the provider. + # Compaction e2e (LCORE-1673): a deliberately small window for every + # model the e2e workflows run against, so the third query of the + # compaction scenarios crosses the trigger threshold. The real windows + # are far larger; this only drives the local estimate, never the provider. + # The vLLM-backed runs (rhaiis, rhelai) take the model id from an env + # var, and context_windows keys are not env-substituted, so they are + # not listed and skip the token-based trigger. context_windows: openai/gpt-4o-mini: 2000 + azure/gpt-4o-mini: 2000 + google-vertex/publishers/google/models/gemini-2.5-flash: 2000 + watsonx/meta-llama/llama-3-3-70b-instruct: 2000 + aws-bedrock/deepseek.v3-v1:0: 2000 rag: byok: stores: @@ -42,14 +49,6 @@ rag: sources: - e2e-test-docs -shields: - - name: pii-redaction - provider_id: redaction - config: - rules: - - pattern: '\d+' - replacement: '[NUM]' - # Compaction on with a low threshold: 10% of the 2000-token window, # above a 100-token floor, keeping one recent turn verbatim. compaction: diff --git a/tests/e2e/features/conversation-compaction.feature b/tests/e2e/features/conversation-compaction.feature index cd00ca03a..89d794788 100644 --- a/tests/e2e/features/conversation-compaction.feature +++ b/tests/e2e/features/conversation-compaction.feature @@ -1,16 +1,15 @@ -@cfg_compaction +# @skip until LCORE-2230 lands the step definitions; Konflux runs the whole +# test list and would fail on the undefined steps. @cfg_compaction is not in +# any GitHub CI shard yet, LCORE-2230 adds it. +@cfg_compaction @skip Feature: Conversation compaction - When a conversation's estimated input approaches the model's context - window, older turns are summarized before the request reaches the LLM - (docs/design/conversation-compaction/conversation-compaction.md). These - scenarios observe compaction from outside only: the context_status field - on responses (R7), the compaction event on the native stream (R12), the - full history the Conversations API keeps serving (R6), and the assistant's - recall of what was said before the summary. The trigger is driven by the - admin configuration (R1, R9): the compaction fixtures register a small - context window for the openai model and a low threshold ratio, so a - three-turn conversation crosses it. + Once the estimated input crosses the configured share of the model's + context window, older turns are summarized before the request reaches + the model. The compaction fixtures register a 2000-token window with a + 10% threshold and keep one recent turn verbatim, so a long third query + is what crosses it: turn one ends up in the summary, turn two stays in + the verbatim buffer, and the third query asks for a fact from each. Background: Given The service is started locally @@ -19,122 +18,61 @@ Feature: Conversation compaction And the Lightspeed stack configuration directory is "tests/e2e/configuration" - Scenario: context_status reports full while compaction never triggers - Given The service uses the lightspeed-stack.yaml configuration - And The service is restarted - When I use "query" to ask question - """ - {"query": "Say hello", "model": "{MODEL}", "provider": "{PROVIDER}"} - """ - Then The status code of the response is 200 - And The response context_status is "full" - - - @openai-only - Scenario: context_status reports summarized once the conversation crosses the threshold - Given The service uses the lightspeed-stack-compaction.yaml configuration - And The service is restarted - When I use "query" to ask question - """ - {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} - """ - Then The status code of the response is 200 - And I store conversation details - When I use "query" to ask question with same conversation_id - """ - {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} - """ - Then The status code of the response is 200 - When I use "query" to ask question with same conversation_id - """ - {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} - """ - Then The status code of the response is 200 - And The response context_status is "summarized" - - - @openai-only - Scenario: the assistant still recalls what was said before the summary + Scenario: the third query crosses the threshold, older turns are summarized, recall and history survive Given The service uses the lightspeed-stack-compaction.yaml configuration And The service is restarted When I use "query" to ask question """ - {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 + And The response context_status is "full" And I store conversation details When I use "query" to ask question with same conversation_id """ - {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "My application namespace is called blue-lagoon. Remember that name too and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 + And The response context_status is "full" When I use "query" to ask question with same conversation_id """ - {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} - """ - Then The status code of the response is 200 - And The response context_status is "summarized" - When I use "query" to ask question with same conversation_id - """ - {"query": "What is the name of my cluster? Reply with the name only.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "Some background on my environment first, no need to comment on it. The cluster runs on bare metal in two racks with three control plane nodes and nine worker nodes, all on the same subnet behind a pair of hardware load balancers. Storage is provided by an external Ceph cluster exposed through the CSI driver, with three storage classes for block, file and object access. Ingress is handled by the default router with two replicas pinned to the infra nodes, and TLS certificates are issued by an internal certificate authority and rotated every ninety days. Monitoring uses the built-in Prometheus stack with a remote write to a central Thanos instance, and alerts are routed to an on-call rotation through a webhook receiver. The image registry is the internal one, backed by an object storage bucket, and images are mirrored from an upstream registry once a day by a scheduled job. Upgrades follow the stable channel, one minor version at a time, and are rehearsed on a staging cluster of the same shape a week before production. Backups of etcd are taken hourly and copied off-site nightly. Now the question: what is the name of my cluster and what is the name of my application namespace? Reply with the two names only, separated by a comma.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 And The response context_status is "summarized" And The response contains following fragments | Fragments in LLM response | | aurora-prod-7 | - - - @openai-only - Scenario: the full conversation history stays available after compaction - Given The service uses the lightspeed-stack-compaction.yaml configuration - And The service is restarted - When I use "query" to ask question - """ - {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} - """ - Then The status code of the response is 200 - And I store conversation details - When I use "query" to ask question with same conversation_id - """ - {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} - """ - Then The status code of the response is 200 - When I use "query" to ask question with same conversation_id - """ - {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} - """ - Then The status code of the response is 200 - And The response context_status is "summarized" + | blue-lagoon | When I use REST API conversation endpoint with conversation_id from above using HTTP GET method Then The status code of the response is 200 And The conversation history includes the following user queries - | User query | - | My OpenShift cluster is named aurora-prod-7. Remember that name. | - | Explain what a pod is in about five sentences. | - | Explain what a deployment is in about five sentences. | + | User query | + | My OpenShift cluster is named aurora-prod-7. Remember that name and reply with OK only. | + | My application namespace is called blue-lagoon. Remember that name too and reply with OK only. | - @openai-only - Scenario: the native stream announces compaction and reports context_status + Scenario: the native stream announces compaction on the query that crosses the threshold Given The service uses the lightspeed-stack-compaction.yaml configuration And The service is restarted When I use "streaming_query" to ask question """ - {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 And I wait for the response to be completed + And The streamed response end event has context_status "full" And I store conversation details When I use "streaming_query" to ask question with same conversation_id """ - {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "My application namespace is called blue-lagoon. Remember that name too and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 And I wait for the response to be completed + And The streamed response end event has context_status "full" When I use "streaming_query" to ask question with same conversation_id """ - {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "Some background on my environment first, no need to comment on it. The cluster runs on bare metal in two racks with three control plane nodes and nine worker nodes, all on the same subnet behind a pair of hardware load balancers. Storage is provided by an external Ceph cluster exposed through the CSI driver, with three storage classes for block, file and object access. Ingress is handled by the default router with two replicas pinned to the infra nodes, and TLS certificates are issued by an internal certificate authority and rotated every ninety days. Monitoring uses the built-in Prometheus stack with a remote write to a central Thanos instance, and alerts are routed to an on-call rotation through a webhook receiver. The image registry is the internal one, backed by an object storage bucket, and images are mirrored from an upstream registry once a day by a scheduled job. Upgrades follow the stable channel, one minor version at a time, and are rehearsed on a staging cluster of the same shape a week before production. Backups of etcd are taken hourly and copied off-site nightly. Now the question: what is the name of my cluster and what is the name of my application namespace? Reply with the two names only, separated by a comma.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 And I wait for the response to be completed @@ -142,24 +80,23 @@ Feature: Conversation compaction And The streamed response end event has context_status "summarized" - @openai-only Scenario: compaction stays off when disabled, even past the threshold Given The service uses the lightspeed-stack-compaction-disabled.yaml configuration And The service is restarted When I use "query" to ask question """ - {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 And I store conversation details When I use "query" to ask question with same conversation_id """ - {"query": "Explain what a pod is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "My application namespace is called blue-lagoon. Remember that name too and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 When I use "query" to ask question with same conversation_id """ - {"query": "Explain what a deployment is in about five sentences.", "model": "{MODEL}", "provider": "{PROVIDER}"} + {"query": "Some background on my environment first, no need to comment on it. The cluster runs on bare metal in two racks with three control plane nodes and nine worker nodes, all on the same subnet behind a pair of hardware load balancers. Storage is provided by an external Ceph cluster exposed through the CSI driver, with three storage classes for block, file and object access. Ingress is handled by the default router with two replicas pinned to the infra nodes, and TLS certificates are issued by an internal certificate authority and rotated every ninety days. Monitoring uses the built-in Prometheus stack with a remote write to a central Thanos instance, and alerts are routed to an on-call rotation through a webhook receiver. The image registry is the internal one, backed by an object storage bucket, and images are mirrored from an upstream registry once a day by a scheduled job. Upgrades follow the stable channel, one minor version at a time, and are rehearsed on a staging cluster of the same shape a week before production. Backups of etcd are taken hourly and copied off-site nightly. Now the question: what is the name of my cluster and what is the name of my application namespace? Reply with the two names only, separated by a comma.", "model": "{MODEL}", "provider": "{PROVIDER}"} """ Then The status code of the response is 200 And The response context_status is "full"