-
Notifications
You must be signed in to change notification settings - Fork 101
LCORE-1673: e2e feature file for conversation compaction (no step implementation) #2611
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,61 @@ | ||
| name: Lightspeed Core Service (LCS) | ||
| service: | ||
| host: 0.0.0.0 | ||
| port: 8080 | ||
| auth_enabled: false | ||
| workers: 1 | ||
| color_log: true | ||
| access_log: true | ||
| ogx: | ||
| # Library mode - embeds OGX as library | ||
| use_as_library_client: true | ||
| # Unified mode: run.yaml (materialized per provider by CI/the harness) | ||
| # is consumed as the synthesis profile instead of the legacy two-file path. | ||
| config: | ||
| profile: run.yaml | ||
| user_data_collection: | ||
| feedback_enabled: true | ||
| feedback_storage: "/tmp/data/feedback" | ||
| transcripts_enabled: true | ||
| transcripts_storage: "/tmp/data/transcripts" | ||
| authentication: | ||
| module: "noop" | ||
| inference: | ||
| default_provider: openai | ||
| default_model: gpt-4o-mini | ||
| # Compaction e2e (LCORE-1673): a deliberately small window for every | ||
| # model the e2e workflows run against, so the third query of the | ||
| # compaction scenarios crosses the trigger threshold. The real windows | ||
| # are far larger; this only drives the local estimate, never the provider. | ||
| # The vLLM-backed runs (rhaiis, rhelai) take the model id from an env | ||
| # var, and context_windows keys are not env-substituted, so they are | ||
| # not listed and skip the token-based trigger. | ||
| context_windows: | ||
| openai/gpt-4o-mini: 2000 | ||
| azure/gpt-4o-mini: 2000 | ||
| google-vertex/publishers/google/models/gemini-2.5-flash: 2000 | ||
| watsonx/meta-llama/llama-3-3-70b-instruct: 2000 | ||
| aws-bedrock/deepseek.v3-v1:0: 2000 | ||
| rag: | ||
| byok: | ||
| stores: | ||
| - rag_id: e2e-test-docs | ||
| backend: faiss | ||
| embedding_model: sentence-transformers/all-mpnet-base-v2 | ||
| embedding_dimension: 768 | ||
| vector_db_id: ${env.FAISS_VECTOR_STORE_ID} | ||
| db_path: ${env.KV_RAG_PATH:=~/.llama/storage/rag/kv_store.db} | ||
| score_multiplier: 1.0 | ||
| retrieval: | ||
| tool: | ||
| sources: | ||
| - e2e-test-docs | ||
|
|
||
| # Same small window and threshold as lightspeed-stack-compaction.yaml, but | ||
| # compaction switched off: context_status must stay "full" past the | ||
| # threshold (enabled is a full off-switch). | ||
| compaction: | ||
| enabled: false | ||
| threshold_ratio: 0.1 | ||
| token_floor: 100 | ||
| buffer_turns: 1 | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,60 @@ | ||
| name: Lightspeed Core Service (LCS) | ||
| service: | ||
| host: 0.0.0.0 | ||
| port: 8080 | ||
| auth_enabled: false | ||
|
Comment on lines
+3
to
+5
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🔒 Security & Privacy | 🟠 Major | ⚡ Quick win 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -eu
printf '%s\n' '--- fixture files ---'
for f in \
tests/e2e/configuration/library-mode/lightspeed-stack-compaction.yaml \
tests/e2e/configuration/server-mode/lightspeed-stack-compaction.yaml \
tests/e2e/configuration/library-mode/lightspeed-stack-compaction-disabled.yaml \
tests/e2e/configuration/server-mode/lightspeed-stack-compaction-disabled.yaml
do
echo "--- $f"
cat -n "$f"
done
printf '%s\n' '--- configuration and auth definitions ---'
rg -n -C 4 'auth_enabled|host:|authentication|Authorization|auth' \
src tests/e2e/configuration tests/e2e 2>/dev/null | head -n 240Repository: lightspeed-core/lightspeed-stack Length of output: 28909 🤖 get_repo_knowledge executed:
Length of output: 21530 🏁 Script executed: #!/bin/bash
set -eu
printf '%s\n' '--- service authentication wiring ---'
rg -n -C 5 'auth_enabled|AUTH_ENABLED|authentication_configuration|AUTH_MOD_NOOP|module.*noop' \
src tests | head -n 260
printf '%s\n' '--- E2E service startup and network exposure ---'
rg -n -C 5 '8080|configure_service|docker compose|network_mode|ports:|host network|localhost|127\.0\.0\.1' \
tests/e2e .github .tekton 2>/dev/null | head -n 320Repository: lightspeed-core/lightspeed-stack Length of output: 43927 🏁 Script executed: #!/bin/bash
set -eu
printf '%s\n' '--- tracked compose and runner files ---'
git ls-files | grep -E '(^|/)(docker-compose[^/]*|.*e2e.*(workflow|pipeline|compose|sh))' | head -n 160
printf '%s\n' '--- compose port mappings and network settings ---'
for f in $(git ls-files | grep -E '(^|/)docker-compose[^/]*\.ya?ml$' | head -n 40); do
matches=$(rg -n '8080|ports:|network_mode|networks:' "$f" || true)
if [ -n "$matches" ]; then
echo "--- $f"
printf '%s\n' "$matches"
fi
done
printf '%s\n' '--- auth middleware application ---'
rg -n -C 6 'auth_enabled|AuthenticationMiddleware|authentication.*middleware|authenticate\(' \
src/app src/authentication src/server src 2>/dev/null | head -n 260Repository: lightspeed-core/lightspeed-stack Length of output: 5498 Authorization Bypass (CWE-306): Missing Authentication for Critical Function Reachability: External · Exploitability: Moderate Restrict the E2E service exposure. The four fixtures disable authentication and use the 📍 Affects 4 files
🤖 Prompt for AI AgentsSource: Coding guidelines |
||
| workers: 1 | ||
| color_log: true | ||
| access_log: true | ||
| ogx: | ||
| # Library mode - embeds OGX as library | ||
| use_as_library_client: true | ||
| # Unified mode: run.yaml (materialized per provider by CI/the harness) | ||
| # is consumed as the synthesis profile instead of the legacy two-file path. | ||
| config: | ||
| profile: run.yaml | ||
| user_data_collection: | ||
| feedback_enabled: true | ||
| feedback_storage: "/tmp/data/feedback" | ||
| transcripts_enabled: true | ||
| transcripts_storage: "/tmp/data/transcripts" | ||
| authentication: | ||
| module: "noop" | ||
| inference: | ||
| default_provider: openai | ||
| default_model: gpt-4o-mini | ||
| # Compaction e2e (LCORE-1673): a deliberately small window for every | ||
| # model the e2e workflows run against, so the third query of the | ||
| # compaction scenarios crosses the trigger threshold. The real windows | ||
| # are far larger; this only drives the local estimate, never the provider. | ||
| # The vLLM-backed runs (rhaiis, rhelai) take the model id from an env | ||
| # var, and context_windows keys are not env-substituted, so they are | ||
| # not listed and skip the token-based trigger. | ||
| context_windows: | ||
| openai/gpt-4o-mini: 2000 | ||
| azure/gpt-4o-mini: 2000 | ||
| google-vertex/publishers/google/models/gemini-2.5-flash: 2000 | ||
| watsonx/meta-llama/llama-3-3-70b-instruct: 2000 | ||
| aws-bedrock/deepseek.v3-v1:0: 2000 | ||
| rag: | ||
| byok: | ||
| stores: | ||
| - rag_id: e2e-test-docs | ||
| backend: faiss | ||
| embedding_model: sentence-transformers/all-mpnet-base-v2 | ||
| embedding_dimension: 768 | ||
| vector_db_id: ${env.FAISS_VECTOR_STORE_ID} | ||
| db_path: ${env.KV_RAG_PATH:=~/.llama/storage/rag/kv_store.db} | ||
| score_multiplier: 1.0 | ||
| retrieval: | ||
| tool: | ||
| sources: | ||
| - e2e-test-docs | ||
|
|
||
| # Compaction on with a low threshold: 10% of the 2000-token window, | ||
| # above a 100-token floor, keeping one recent turn verbatim. | ||
| compaction: | ||
| enabled: true | ||
| threshold_ratio: 0.1 | ||
| token_floor: 100 | ||
| buffer_turns: 1 | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,59 @@ | ||
| name: Lightspeed Core Service (LCS) | ||
| service: | ||
| host: 0.0.0.0 | ||
| port: 8080 | ||
| auth_enabled: false | ||
| workers: 1 | ||
| color_log: true | ||
| access_log: true | ||
| ogx: | ||
| # Server mode - connects to separate OGX service | ||
| use_as_library_client: false | ||
| url: http://${env.E2E_OGX_HOSTNAME}:8321 | ||
| api_key: xyzzy | ||
| user_data_collection: | ||
| feedback_enabled: true | ||
| feedback_storage: "/tmp/data/feedback" | ||
| transcripts_enabled: true | ||
| transcripts_storage: "/tmp/data/transcripts" | ||
| authentication: | ||
| module: "noop" | ||
| inference: | ||
| default_provider: openai | ||
| default_model: gpt-4o-mini | ||
| # Compaction e2e (LCORE-1673): a deliberately small window for every | ||
| # model the e2e workflows run against, so the third query of the | ||
| # compaction scenarios crosses the trigger threshold. The real windows | ||
| # are far larger; this only drives the local estimate, never the provider. | ||
| # The vLLM-backed runs (rhaiis, rhelai) take the model id from an env | ||
| # var, and context_windows keys are not env-substituted, so they are | ||
| # not listed and skip the token-based trigger. | ||
| context_windows: | ||
| openai/gpt-4o-mini: 2000 | ||
| azure/gpt-4o-mini: 2000 | ||
| google-vertex/publishers/google/models/gemini-2.5-flash: 2000 | ||
| watsonx/meta-llama/llama-3-3-70b-instruct: 2000 | ||
| aws-bedrock/deepseek.v3-v1:0: 2000 | ||
| rag: | ||
| byok: | ||
| stores: | ||
| - rag_id: e2e-test-docs | ||
| backend: faiss | ||
| embedding_model: sentence-transformers/all-mpnet-base-v2 | ||
| embedding_dimension: 768 | ||
| vector_db_id: ${env.FAISS_VECTOR_STORE_ID} | ||
| db_path: ${env.KV_RAG_PATH:=~/.llama/storage/rag/kv_store.db} | ||
| score_multiplier: 1.0 | ||
| retrieval: | ||
| tool: | ||
| sources: | ||
| - e2e-test-docs | ||
|
|
||
| # Same small window and threshold as lightspeed-stack-compaction.yaml, but | ||
| # compaction switched off: context_status must stay "full" past the | ||
| # threshold (enabled is a full off-switch). | ||
| compaction: | ||
| enabled: false | ||
| threshold_ratio: 0.1 | ||
| token_floor: 100 | ||
| buffer_turns: 1 |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,58 @@ | ||
| name: Lightspeed Core Service (LCS) | ||
| service: | ||
| host: 0.0.0.0 | ||
| port: 8080 | ||
| auth_enabled: false | ||
| workers: 1 | ||
| color_log: true | ||
| access_log: true | ||
| ogx: | ||
| # Server mode - connects to separate OGX service | ||
| use_as_library_client: false | ||
| url: http://${env.E2E_OGX_HOSTNAME}:8321 | ||
| api_key: xyzzy | ||
| user_data_collection: | ||
| feedback_enabled: true | ||
| feedback_storage: "/tmp/data/feedback" | ||
| transcripts_enabled: true | ||
| transcripts_storage: "/tmp/data/transcripts" | ||
| authentication: | ||
| module: "noop" | ||
| inference: | ||
| default_provider: openai | ||
| default_model: gpt-4o-mini | ||
| # Compaction e2e (LCORE-1673): a deliberately small window for every | ||
| # model the e2e workflows run against, so the third query of the | ||
| # compaction scenarios crosses the trigger threshold. The real windows | ||
| # are far larger; this only drives the local estimate, never the provider. | ||
| # The vLLM-backed runs (rhaiis, rhelai) take the model id from an env | ||
| # var, and context_windows keys are not env-substituted, so they are | ||
| # not listed and skip the token-based trigger. | ||
| context_windows: | ||
| openai/gpt-4o-mini: 2000 | ||
| azure/gpt-4o-mini: 2000 | ||
| google-vertex/publishers/google/models/gemini-2.5-flash: 2000 | ||
| watsonx/meta-llama/llama-3-3-70b-instruct: 2000 | ||
| aws-bedrock/deepseek.v3-v1:0: 2000 | ||
| rag: | ||
| byok: | ||
| stores: | ||
| - rag_id: e2e-test-docs | ||
| backend: faiss | ||
| embedding_model: sentence-transformers/all-mpnet-base-v2 | ||
| embedding_dimension: 768 | ||
| vector_db_id: ${env.FAISS_VECTOR_STORE_ID} | ||
| db_path: ${env.KV_RAG_PATH:=~/.llama/storage/rag/kv_store.db} | ||
| score_multiplier: 1.0 | ||
| retrieval: | ||
| tool: | ||
| sources: | ||
| - e2e-test-docs | ||
|
|
||
| # Compaction on with a low threshold: 10% of the 2000-token window, | ||
| # above a 100-token floor, keeping one recent turn verbatim. | ||
| compaction: | ||
| enabled: true | ||
| threshold_ratio: 0.1 | ||
| token_floor: 100 | ||
| buffer_turns: 1 |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,102 @@ | ||
| # @skip until LCORE-2230 lands the step definitions; Konflux runs the whole | ||
| # test list and would fail on the undefined steps. @cfg_compaction is not in | ||
| # any GitHub CI shard yet, LCORE-2230 adds it. | ||
| @cfg_compaction @skip | ||
| Feature: Conversation compaction | ||
|
|
||
| Once the estimated input crosses the configured share of the model's | ||
| context window, older turns are summarized before the request reaches | ||
| the model. The compaction fixtures register a 2000-token window with a | ||
| 10% threshold and keep one recent turn verbatim, so a long third query | ||
| is what crosses it: turn one ends up in the summary, turn two stays in | ||
| the verbatim buffer, and the third query asks for a fact from each. | ||
|
|
||
| Background: | ||
| Given The service is started locally | ||
| And The system is in default state | ||
| And REST API service prefix is /v1 | ||
| And the Lightspeed stack configuration directory is "tests/e2e/configuration" | ||
|
|
||
|
|
||
| Scenario: the third query crosses the threshold, older turns are summarized, recall and history survive | ||
| Given The service uses the lightspeed-stack-compaction.yaml configuration | ||
| And The service is restarted | ||
| When I use "query" to ask question | ||
| """ | ||
| {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| And The response context_status is "full" | ||
| And I store conversation details | ||
| When I use "query" to ask question with same conversation_id | ||
| """ | ||
| {"query": "My application namespace is called blue-lagoon. Remember that name too and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| And The response context_status is "full" | ||
| When I use "query" to ask question with same conversation_id | ||
| """ | ||
| {"query": "Some background on my environment first, no need to comment on it. The cluster runs on bare metal in two racks with three control plane nodes and nine worker nodes, all on the same subnet behind a pair of hardware load balancers. Storage is provided by an external Ceph cluster exposed through the CSI driver, with three storage classes for block, file and object access. Ingress is handled by the default router with two replicas pinned to the infra nodes, and TLS certificates are issued by an internal certificate authority and rotated every ninety days. Monitoring uses the built-in Prometheus stack with a remote write to a central Thanos instance, and alerts are routed to an on-call rotation through a webhook receiver. The image registry is the internal one, backed by an object storage bucket, and images are mirrored from an upstream registry once a day by a scheduled job. Upgrades follow the stable channel, one minor version at a time, and are rehearsed on a staging cluster of the same shape a week before production. Backups of etcd are taken hourly and copied off-site nightly. Now the question: what is the name of my cluster and what is the name of my application namespace? Reply with the two names only, separated by a comma.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| And The response context_status is "summarized" | ||
| And The response contains following fragments | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. is this intended to cover the logic using the buffer? If so, then I also want to see the test here to test that something is not rememebered
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Partly -- one fact summarized, one buffered. What would "not remembered" look like here? |
||
| | Fragments in LLM response | | ||
| | aurora-prod-7 | | ||
| | blue-lagoon | | ||
| When I use REST API conversation endpoint with conversation_id from above using HTTP GET method | ||
| Then The status code of the response is 200 | ||
| And The conversation history includes the following user queries | ||
| | User query | | ||
| | My OpenShift cluster is named aurora-prod-7. Remember that name and reply with OK only. | | ||
| | My application namespace is called blue-lagoon. Remember that name too and reply with OK only. | | ||
|
|
||
|
|
||
| Scenario: the native stream announces compaction on the query that crosses the threshold | ||
| Given The service uses the lightspeed-stack-compaction.yaml configuration | ||
| And The service is restarted | ||
| When I use "streaming_query" to ask question | ||
| """ | ||
| {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| And I wait for the response to be completed | ||
| And The streamed response end event has context_status "full" | ||
| And I store conversation details | ||
| When I use "streaming_query" to ask question with same conversation_id | ||
| """ | ||
| {"query": "My application namespace is called blue-lagoon. Remember that name too and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| And I wait for the response to be completed | ||
| And The streamed response end event has context_status "full" | ||
| When I use "streaming_query" to ask question with same conversation_id | ||
| """ | ||
| {"query": "Some background on my environment first, no need to comment on it. The cluster runs on bare metal in two racks with three control plane nodes and nine worker nodes, all on the same subnet behind a pair of hardware load balancers. Storage is provided by an external Ceph cluster exposed through the CSI driver, with three storage classes for block, file and object access. Ingress is handled by the default router with two replicas pinned to the infra nodes, and TLS certificates are issued by an internal certificate authority and rotated every ninety days. Monitoring uses the built-in Prometheus stack with a remote write to a central Thanos instance, and alerts are routed to an on-call rotation through a webhook receiver. The image registry is the internal one, backed by an object storage bucket, and images are mirrored from an upstream registry once a day by a scheduled job. Upgrades follow the stable channel, one minor version at a time, and are rehearsed on a staging cluster of the same shape a week before production. Backups of etcd are taken hourly and copied off-site nightly. Now the question: what is the name of my cluster and what is the name of my application namespace? Reply with the two names only, separated by a comma.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| And I wait for the response to be completed | ||
| And The streamed response contains a compaction event before the first token | ||
| And The streamed response end event has context_status "summarized" | ||
|
|
||
|
|
||
| Scenario: compaction stays off when disabled, even past the threshold | ||
| Given The service uses the lightspeed-stack-compaction-disabled.yaml configuration | ||
| And The service is restarted | ||
| When I use "query" to ask question | ||
| """ | ||
| {"query": "My OpenShift cluster is named aurora-prod-7. Remember that name and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| And I store conversation details | ||
| When I use "query" to ask question with same conversation_id | ||
| """ | ||
| {"query": "My application namespace is called blue-lagoon. Remember that name too and reply with OK only.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| When I use "query" to ask question with same conversation_id | ||
| """ | ||
| {"query": "Some background on my environment first, no need to comment on it. The cluster runs on bare metal in two racks with three control plane nodes and nine worker nodes, all on the same subnet behind a pair of hardware load balancers. Storage is provided by an external Ceph cluster exposed through the CSI driver, with three storage classes for block, file and object access. Ingress is handled by the default router with two replicas pinned to the infra nodes, and TLS certificates are issued by an internal certificate authority and rotated every ninety days. Monitoring uses the built-in Prometheus stack with a remote write to a central Thanos instance, and alerts are routed to an on-call rotation through a webhook receiver. The image registry is the internal one, backed by an object storage bucket, and images are mirrored from an upstream registry once a day by a scheduled job. Upgrades follow the stable channel, one minor version at a time, and are rehearsed on a staging cluster of the same shape a week before production. Backups of etcd are taken hourly and copied off-site nightly. Now the question: what is the name of my cluster and what is the name of my application namespace? Reply with the two names only, separated by a comma.", "model": "{MODEL}", "provider": "{PROVIDER}"} | ||
| """ | ||
| Then The status code of the response is 200 | ||
| And The response context_status is "full" | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -23,6 +23,7 @@ features/rlsapi_v1.feature | |
| features/streaming_query.feature | ||
| features/vector_stores.feature | ||
| features/conversation_cache_v2.feature | ||
| features/conversation-compaction.feature | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win Do not register an unimplemented feature. The PR objective states that nine new step patterns remain undefined. Line 26 makes Behave execute this feature in normal E2E runs, so it will fail with undefined steps. Implement the steps before registration, or remove this entry until LCORE-2230 lands. 🤖 Prompt for AI Agents |
||
| features/feedback.feature | ||
| features/http_401_unauthorized.feature | ||
| features/rbac.feature | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
is there any specific reason why we hardcode here the specific model and ignore the other providers? I would rather see here the env var for the model and provider and specific context windows for all applicable providers