From e4de99b29551fc39145a2c80d62769885b43b707 Mon Sep 17 00:00:00 2001 From: Cristian Magherusan-Stanciu Date: Mon, 7 Sep 2026 18:22:50 +0200 Subject: [PATCH] docs(audit): add the 2026-09-02 full codebase audit report Records 545 findings from a sharded review of every package at 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd. Each finding was written by one reviewer and checked by an independent verifier that did not write it: 454 confirmed, 46 plausible, 45 rejected. Rejected findings are kept with the verifier's reasoning rather than deleted, so nobody re-raises them. Issues filed from this audit cite finding IDs and this path, so the report needs to exist on main for those references to resolve. Two adjustments were needed to land it: Excludes docs/audits/ from markdownlint. The report quotes code verbatim, and --fix rewrote a git-secrets pattern by stripping the trailing space inside `resource `, which changes what the finding claims. Evidence a formatter can edit is not evidence. Redacts the synthetic AWS key in A14-010's reproduction. The key was always fake and its characters carry no meaning; the finding is about git-secrets matching allowed regexes against whole lines, which the surrounding text still shows. Keeping a key-shaped literal would leave the scanner blocking every future commit that touches this file. Co-Authored-By: claude-flow Claude-Session: https://claude.ai/code/session_01Fu9uWjxtDFx5HDKeMRt1jC --- .markdownlintignore | 5 + docs/audits/codebase-audit-2026-09-02.md | 10352 +++++++++++++++++++++ 2 files changed, 10357 insertions(+) create mode 100644 .markdownlintignore create mode 100644 docs/audits/codebase-audit-2026-09-02.md diff --git a/.markdownlintignore b/.markdownlintignore new file mode 100644 index 000000000..6b2f94c6f --- /dev/null +++ b/.markdownlintignore @@ -0,0 +1,5 @@ +# Generated audit reports are verbatim records: reviewer text, quoted code +# excerpts and verifier verdicts. Autoformatting them rewrites evidence +# (for example stripping a meaningful trailing space inside a code span), +# so they are linted by the process that writes them, not by markdownlint. +docs/audits/ diff --git a/docs/audits/codebase-audit-2026-09-02.md b/docs/audits/codebase-audit-2026-09-02.md new file mode 100644 index 000000000..49ce984dc --- /dev/null +++ b/docs/audits/codebase-audit-2026-09-02.md @@ -0,0 +1,10352 @@ +# CUDly codebase audit, 2026-09-02 + +| | | +|---|---| +| Audited commit | `3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd` | +| Commit is | the tip of `origin/main` at audit start, not the working tree | +| Audit window | 2026-09-01T23:29Z to 2026-09-02T20:41Z | +| Review shards | 19, all completed | +| Findings written | 545 | +| Verification | every finding checked by an independent verifier that did not write it | + +The audit read the pinned commit only. Uncommitted working-tree changes were out of scope. +Each shard was reviewed by one agent and then verified by a second agent that had not seen the +first agent's reasoning. Verdicts below are the verifier's, not the reviewer's. + +## How to read this report + +**CONFIRMED** means the verifier reproduced the reviewer's claim against the pinned commit, by +reading the cited code, running a test, or both. **PLAUSIBLE** means the mechanism holds but the +verifier could not close the last step, usually reachability or a runtime value it could not +observe. **REJECTED** means the verifier found the claim wrong: the code does not do what the +finding says, or a compensating control the reviewer missed makes the scenario impossible. + +Rejected findings are not deleted. They are moved to the end of this report with the verifier's +reasoning intact, so that the same claim is not re-raised later by someone who has not seen why +it fails. Nothing in the rejected section should be actioned. + +A `severity-adjusted` line means the verifier disagreed with the reviewer's rating and changed +it. Where that line is present, the adjusted value governs: this report sorts, groups and counts +the finding by the adjusted severity, and the original `severity` line is left in place so the +disagreement stays visible. + +Some findings sit on code paths that no live caller reaches. The verifier says so in its verdict, +usually as the reason for a downgrade. A latent defect on an unreachable path is real and worth +fixing, but it is not shipping a failure today, so it ranks below a reachable finding of the same +severity. + +## Shard coverage + +Counts are parsed from the shard files, not quoted from any summary. + +| Shard | Scope | Findings | Confirmed | Plausible | Rejected | Not read | +|---|---|---:|---:|---:|---:|---| +| A01 | Money-path API handlers: purchases, RI exchange, plans, plan health, request validation | 22 | 22 | 0 | 0 | internal/api/types_apikeys.go unread; most test files read by excerpt only, deferred to A16. | +| A02 | The rest of internal/api: routing, middleware, auth, accounts, users, groups, API keys, recommendations, history, dashboard, config, analytics, rate limiters, OpenAPI spec | 26 | 25 | 1 | 0 | OpenAPI component schemas spot-checked only. Tests and mocks grepped for specific questions, not read end to end. | +| A03 | Authentication and secrets: internal/auth, internal/oidc, internal/credentials, internal/secrets | 25 | 21 | 1 | 3 | Most test files in these four packages unread, folded into A16. | +| A04 | internal/config, internal/database, SQL migrations | 20 | 17 | 2 | 1 | Most internal/config and internal/database test files unread, folded into A16. | +| A05 | Purchase execution and scheduling: internal/purchase, internal/execution, internal/scheduler, internal/commitmentopts | 19 | 14 | 4 | 1 | All non-test files read. About 22 test files grepped only, deferred to A16. | +| A06 | internal/server, email, analytics, reporter, accounts, iacfiles, runtime | 30 | 23 | 3 | 4 | Tests not read as tests; deferred to A16. | +| A07 | providers/aws | 33 | 28 | 1 | 4 | All 27 non-test Go files read in full. 32 test files deferred to A16. providers/aws/services/AUDIT.md not read. | +| A08 | providers/azure and providers/gcp, first pass | 28 | 18 | 8 | 2 | Left unread: azure httpclient.go, six Azure service clients, compute/exchange.go, displayname.go, mocks; GCP cloudsql, cloudstorage, memorystore clients and most of computeengine. Split into A08b. All Azure and GCP tests, about 23k lines, deferred to A16. | +| A08b | The Azure and GCP provider clients A08 left unread, plus pkg/httpclient and Azure internal pricing | 45 | 40 | 5 | 0 | computeengine lines 560-1188 cited by grep rather than full read in findings 019, 029 and 032. | +| A09 | pkg/ shared libraries: httpclient, exchange, recfilter, common, provider, ladder, scorer, retry, logging, config | 32 | 26 | 3 | 3 | pkg/config/load.go and types.go bodies unread; the claim that pkg/config has no importers was left to the verifier. | +| A10 | cmd/: CLI, MCP server, Lambda handlers, rekey, cleanup, gen-permissions | 30 | 28 | 2 | 0 | Tests deferred to A16. | +| A11 | frontend/src part 1: auth, state, settings, users, groups, permissions, API client | 25 | 15 | 5 | 5 | Frontend tests not read as tests. | +| A12 | frontend/src part 2, money-facing: recommendations, plans, purchase modal, RI exchange, history, dashboard, ladder | 76 | 64 | 4 | 8 | recommendations.ts read to about 3100 of 5481 lines; setup and column-filter helpers unread. riexchange.ts and history.ts were delegated to sub-reviewers, with criticals and highs re-verified by the shard reviewer. | +| A13 | terraform, iac, cloudformation, arm, Docker, scanner suppressions, first pass | 24 | 22 | 2 | 0 | Covered only a fraction of scope. Most of terraform/modules, terraform/environments, iac/federation and both cloudformation/stacks templates unread. Split into A13b and A13c. | +| A13b | Customer-facing federation trust boundary: aws-target, aws-cross-account, azure-target, gcp-sa-impersonation, CloudFormation stacks, ARM templates | 16 | 14 | 1 | 1 | Scope covered as assigned. | +| A13c | Terraform provisioning CUDly's own infrastructure | 26 | 20 | 2 | 4 | Stopped in priority 5. Unread: monitoring modules for all three clouds, frontend/azure and frontend/gcp, build/outputs.tf and build/scripts, policy_guard_test.go, test-iam-role.sh, ci-cd-permissions/policy_*.tf, and the AKS, GKE and cleanup-function compute modules. | +| A14 | .github/workflows, scripts, Makefiles, pre-commit and linter configs, secret allowlists | 43 | 37 | 2 | 4 | Scope covered as assigned. | +| A15 | Cross-cutting: dependency and vulnerability scanning, npm audit, govulncheck, currency handling, CI tool health | 11 | 9 | 0 | 2 | A sampling pass rather than a file-by-file read; it examined what crosses package boundaries. | +| A16 | Test files judged as tests, using go test coverage profiles rather than greps | 14 | 11 | 0 | 3 | Stopped partway through priority 2. Untouched: internal/credentials and internal/secrets tests, three OIDC cloud signers, most of internal/api and all internal/config tests, and all tests under providers, pkg, cmd and frontend. | +| **Total** | | **545** | **454** | **46** | **45** | | + +Shards A08b, A13b, A13c and A16 exist because an earlier reviewer left part of its assigned scope +unread. A08 deferred most Azure and GCP service clients, A13 covered only a fraction of the +infrastructure tree, and every shard that deferred its test files pushed them to A16. Each +follow-up shard was given the unread files as its own scope and was verified the same way. + +## Summary + +Actionable means CONFIRMED or PLAUSIBLE. Rejected findings are excluded from every table in this +section. + +### Actionable findings by severity + +| Severity | Findings | +|---|---:| +| critical | 10 | +| high | 82 | +| medium | 218 | +| low | 190 | +| **Total** | **500** | + +### Actionable findings by category + +| Category | Findings | +|---|---:| +| security | 78 | +| money-path | 86 | +| silent-fallback | 65 | +| correctness | 118 | +| concurrency | 12 | +| performance | 9 | +| test-gap | 29 | +| duplication | 13 | +| dead-code | 31 | +| over-engineering | 3 | +| ops | 37 | +| hygiene | 14 | +| bug | 1 | +| bugs | 4 | +| **Total** | **500** | + +### Severity against category + +| Category | critical | high | medium | low | Total | +|---|---:|---:|---:|---:|---:| +| security | 2 | 21 | 31 | 24 | 78 | +| money-path | 6 | 32 | 40 | 8 | 86 | +| silent-fallback | 0 | 3 | 35 | 27 | 65 | +| correctness | 2 | 12 | 59 | 45 | 118 | +| concurrency | 0 | 2 | 7 | 3 | 12 | +| performance | 0 | 0 | 5 | 4 | 9 | +| test-gap | 0 | 8 | 13 | 8 | 29 | +| duplication | 0 | 0 | 4 | 9 | 13 | +| dead-code | 0 | 0 | 8 | 23 | 31 | +| over-engineering | 0 | 0 | 0 | 3 | 3 | +| ops | 0 | 4 | 15 | 18 | 37 | +| hygiene | 0 | 0 | 1 | 13 | 14 | +| bug | 0 | 0 | 0 | 1 | 1 | +| bugs | 0 | 0 | 0 | 4 | 4 | +| **Total** | **10** | **82** | **218** | **190** | **500** | + +## Critical and high findings + +Every actionable finding rated critical or high after verification, in one line each. The full +block for each is in the findings section under its category. + +| ID | Severity | Category | Location | Summary | +|---|---|---|---|---| +| A03-001 | critical | security | `internal/auth/group_ceiling.go:85` | Resource wildcard `execute:*` bypasses the admin money-verb carve-out at both the grant ceiling and enforcement | +| A05-001 | critical | money-path | `internal/purchase/execution.go:371` | Direct-execute spanning two cloud accounts silently buys in the ambient host account | +| A08-001 | critical | money-path | `providers/gcp/services/computeengine/client.go:1392` | GCP CUD purchase reads memory from a value-typed Details, but the purchase-execution path supplies a pointer | +| A08b-001 | critical | security | `pkg/httpclient/httpclient.go:41` | IMDS blocklist matches literal host text, so IPv6-mapped, alternate-radix and DNS forms of 169.254.169.254 pass through | +| A08b-005 | critical | money-path | `providers/azure/services/cache/client.go:207` | `ReservationsDetails` is a daily usage API, so `GetExistingCommitments` emits one Commitment per reservation per day | +| A09-004 | critical | money-path | `pkg/exchange/exchange.go:305` | An absent PaymentDue is treated as $0, disabling the spend cap on an irreversible exchange | +| A11-001 | critical | correctness | `frontend/src/groups/handlers.ts:17` | The group create/edit form's submit handler is never wired in production | +| A12-001 | critical | money-path | `frontend/src/recommendations.ts:5412` | Purchase modal's Term and Payment selects mutate the submitted rec without re-pricing it | +| A12-002 | critical | money-path | `frontend/src/recommendations.ts:4336` | Fan-out modal says an incompatible bucket "will be skipped", then submits it | +| A12-003 | critical | correctness | `frontend/src/riexchange.ts:1809` | RI-exchange execute never sends `region`, which the backend rejects with 400 | +| A01-001 | high | money-path | `internal/api/handler_purchases.go:2255` | MaxPurchaseAmount cap is enforced against client-asserted dollar amounts, not the priced commitment | +| A01-002 | high | security | `internal/api/handler_purchases.go:685` | Session-authed approve, cancel, retry and revoke never check the session's allowed_accounts scope | +| A01-003 | high | correctness | `internal/api/handler_purchases.go:320` | "Run now" strands the execution in `running`; nothing executes it and the reaper marks it failed | +| A01-004 | high | money-path | `internal/api/handler_plans.go:281` | PUT /api/plans/{id} rebuilds the ramp schedule from scratch, resetting CurrentStep/StartDate and re-arming the ramp | +| A02-001 | high | security | `internal/api/handler_dashboard.go:38` | Dashboard commitment KPIs skip allowed_accounts scoping when the caller supplies an explicit account filter | +| A02-002 | high | correctness | `internal/api/handler_auth.go:238` | setupAdmin stores the base64-encoded password, so the bootstrap admin cannot log in with the password they typed | +| A03-002 | high | security | `internal/auth/service_apikeys.go:128` | An admin can mint a user API key carrying `execute:*` in a single request and spend with it | +| A03-003 | high | security | `internal/auth/service_user.go:551` | Self-membership guard checks carved-out pairs exactly, so an admin can join a wildcard group they created | +| A03-004 | high | security | `internal/auth/service_user.go:396` | User-membership writes have no grant ceiling: any create:users or update:users holder can mint an Administrator or Purchaser | +| A03-005 | high | security | `internal/auth/service.go:316` | Deactivating a user does not end their sessions, and session validation never re-checks the user | +| A03-006 | high | security | `internal/auth/service_password.go:355` | A deactivated account can reactivate itself through the forgot-password flow | +| A05-002 | high | test-gap | `internal/purchase/execution_test.go:1628` | `singleCloudAccountIDFromRecs` is unit-tested in isolation, so the ambient fallback above stays green | +| A06-001 | high | concurrency | `internal/server/handler.go:112` | Advisory lock is released with the request context, so a canceled/expired scheduled run strands the lock on a pooled connection | +| A06-003 | high | security | `internal/server/http.go:303` | SourceIP carries the TCP port (or the proxy's IP), defeating the login brute-force rate limit | +| A07-001 | high | money-path | `providers/aws/services/savingsplans/client.go:425` | CE-supplied OfferingID short-circuits every Savings Plans purchase validation | +| A07-006 | high | correctness | `providers/aws/services/elasticache/client.go:311` | ElastiCache offering lookup sends the raw CE engine string as ProductDescription, unvalidated | +| A07-007 | high | money-path | `providers/aws/services/elasticache/client.go:92` | ElastiCache and MemoryDB commitments carry no Engine, so the duplicate-purchase guard never matches | +| A07-010 | high | correctness | `providers/aws/recommendations/usage_history.go:85` | Daily-sparkline coverage call sets Granularity together with GroupBy, which the API rejects | +| A07-012 | high | money-path | `providers/aws/ladder/layer_states.go:307` | EC2 ladder coverage blends in ElastiCache, OpenSearch, Redshift and MemoryDB pools | +| A08-004 | high | correctness | `providers/azure/services/compute/client.go:914` | Azure VM SKU enrichment is dead on the recommendations path and burns a full SKU-catalogue walk per subscription | +| A08-005 | high | money-path | `providers/azure/services/compute/client.go:765` | Azure retail-price extraction is last-item-wins across Spot/Windows/Linux SKUs and silently mixes currencies | +| A08-006 | high | money-path | `providers/azure/internal/recommendations/converter.go:423` | ExpandPaymentVariants fabricates 100% savings when the provider omitted the commitment cost | +| A08-007 | high | money-path | `providers/azure/services/savingsplans/client.go:261` | Azure Savings Plan purchases hardcode CurrencyCode "USD" on the commitment body | +| A08-008 | high | money-path | `providers/azure/internal/recommendations/converter.go:360` | Modern (MCA) recommendation extraction discards the currency Azure reported | +| A08-009 | high | silent-fallback | `providers/gcp/services/computeengine/client.go:876` | GCP offering details silently bill an unrecognized payment option as all-upfront | +| A08b-002 | high | security | `pkg/httpclient/httpclient.go:31` | IMDS denylist covers two addresses; Azure WireServer and the rest of link-local are reachable | +| A08b-004 | high | security | `providers/azure/services/cache/client.go:103` | Four Azure clients accept a nil HTTP client in `NewClientWithHTTP` with no hardened fallback | +| A08b-006 | high | money-path | `providers/azure/services/cache/client.go:251` | Azure commitments fabricate `State: "active"` and `Region` from fields the SDK response does not carry | +| A08b-007 | high | money-path | `providers/azure/services/cache/client.go:260` | Azure commitments never populate Count, StartDate, EndDate or Cost | +| A08b-008 | high | money-path | `providers/azure/services/cosmosdb/client.go:532` | Cosmos DB pricing ignores the SKU it was asked to price | +| A08b-009 | high | money-path | `providers/azure/services/search/client.go:447` | Azure Search pricing ignores the SKU it was asked to price | +| A08b-010 | high | money-path | `providers/azure/services/cache/client.go:599` | Retail-price extraction is last-item-wins with no SKU or unit-of-measure check | +| A08b-014 | high | money-path | `providers/gcp/services/cloudstorage/client.go:356` | Cloud Storage prices a per-GiB-month SKU as if it were per-hour, inflating cost by 730x | +| A08b-016 | high | silent-fallback | `providers/gcp/services/cloudsql/client.go:507` | GCP pricing failures are logged and the recommendation is emitted with zero cost alongside non-zero savings | +| A08b-017 | high | correctness | `providers/gcp/services/memorystore/client.go:456` | GCP `ResourceType` is a resource instance name, then used as a pricing tier and validated against tier constants | +| A08b-022 | high | silent-fallback | `providers/azure/services/cache/client.go:185` | Four Azure clients report "no existing commitments" when the reservations pager cannot be built | +| A08b-024 | high | money-path | `providers/azure/services/managedredis/client.go:41` | Azure Cache for Redis is enumerated by two clients, so its reservations and recommendations are counted twice | +| A08b-029 | high | correctness | `providers/gcp/services/computeengine/client.go:524` | Compute Engine existing commitments report `ResourceType` as "VCPU" instead of a machine type | +| A08b-030 | high | money-path | `providers/gcp/services/computeengine/client.go:1296` | The vCPU amount is selected by "not memory", so an accelerator or local-SSD amount is read as the vCPU count | +| A08b-032 | high | money-path | `providers/gcp/services/computeengine/client.go:1415` | An empty idempotency token silently produces a non-idempotent, second-granularity commitment name | +| A09-001 | high | security | `pkg/httpclient/httpclient.go:41` | IMDS blocklist is bypassed by any hostname that resolves to the metadata address | +| A09-002 | high | security | `pkg/httpclient/httpclient.go:31` | IMDS blocklist misses the ECS/EKS credential endpoints and the rest of 169.254.0.0/16 | +| A09-005 | high | money-path | `pkg/exchange/exchange.go:313` | The USD spend cap is compared against a quote whose currency is never checked | +| A09-008 | high | money-path | `pkg/exchange/auto.go:319` | The per-exchange cap is skipped entirely when the quote reports no PaymentDue | +| A10-001 | high | money-path | `cmd/multi_service_engine_versions.go:486` | Extended-support exclusion cuts Count without scaling the row's money fields | +| A10-002 | high | money-path | `cmd/multi_service_helpers.go:589` | A failed duplicate check falls back to purchasing the full, un-deduplicated counts | +| A10-003 | high | money-path | `cmd/multi_service.go:63` | `--target-coverage` sizes as if nothing is owned when the Cost Explorer coverage fetch fails | +| A10-004 | high | money-path | `cmd/helpers.go:91` | A failed account lookup substitutes the account ID, silently defeating `--exclude-accounts` | +| A10-018 | high | security | `cmd/configure_gcp.go:686` | The GCP setup wizard grants `roles/compute.admin` at project scope | +| A11-002 | high | security | `frontend/src/groups/groupModals.ts:112` | Removing every permission from a group reports success while the backend keeps the old permissions | +| A12-004 | high | money-path | `frontend/src/recommendations.ts:3919` | Capacity-% scaling is count-based, so every Savings Plans rec is silently dropped below 100% | +| A12-005 | high | money-path | `frontend/src/history.ts:463` | `scheduled` executions render the green "Completed" badge | +| A12-006 | high | money-path | `frontend/src/history.ts:1463` | Marketplace consent dialog computes the list price with `Math.round` while the backend uses `math.Floor` | +| A12-008 | high | correctness | `frontend/src/plans.ts:2326` | The AWS "Savings Plans" optgroup is hidden and disabled whenever a provider is selected | +| A12-009 | high | money-path | `frontend/src/recommendations.ts:5013` | "Execute Now" warning shows an upfront total that goes stale when rows are toggled | +| A12-010 | high | money-path | `frontend/src/riexchange.ts:1609` | Changing the exchange targets after a quote does not invalidate it | +| A12-011 | high | concurrency | `frontend/src/riexchange.ts:318` | A pending AWS utilization response re-renders the shared container over the Azure/GCP table | +| A12-074 | high | money-path | `frontend/src/recommendations.ts:3115` | The Savings Plans group row scales its savings by the cost period twice | +| A13-001 | high | money-path | `terraform/modules/compute/aws/lambda/main.tf:381` | `ce:GetCostAndUsage` is granted in no IaC flavor, but the ladder baseline calls it | +| A13-002 | high | money-path | `terraform/modules/compute/aws/lambda/main.tf:352` | RI Marketplace listing actions are granted nowhere, so the sell flow 403s | +| A13-004 | high | security | `terraform/environments/aws/ci-cd-permissions/policy_networking.tf:166` | The deploy role holds unconditioned account-wide KMS Decrypt, Encrypt and CreateGrant | +| A13-005 | high | security | `terraform/environments/azure/build.tf:24` | The ACR admin password is de-sensitized into a non-sensitive module variable | +| A13b-001 | high | correctness | `iac/federation/gcp-sa-impersonation/terraform/main.tf:27` | GCP SA-impersonation module grants two roles that cannot bind at project scope and never grants the CUD-purchase permission | +| A13b-002 | high | correctness | `iac/federation/azure-target/bicep/azure-wif.bicep:24` | Azure Bicep/ARM assigns the built-in Reservation Purchaser role that the repo documents as insufficient for purchases | +| A13b-004 | high | correctness | `cloudformation/stacks/CUDly-CrossAccount/template.yaml:64` | Cross-account trust principal omits the `/lambda/` path the hub role actually carries | +| A13c-002 | high | ops | `terraform/modules/compute/aws/lambda/main.tf:506` | Lambda log group uses `name_prefix`, so Lambda's real log group is unmanaged — retention never applies and the migration alarm can never fire | +| A13c-005 | high | security | `terraform/environments/azure/registry.tf:10` | Azure ACR admin account is enabled and its credentials are written into the Container App and Terraform state, despite an AcrPull grant existing alongside | +| A14-002 | high | ops | `.github/workflows/rollback.yml:49` | The rollback workflow pins a Terraform version the modules reject | +| A14-003 | high | ops | `.github/workflows/database-migration.yml:296` | database-migration.yml runs `terraform init` with no backend config, so it can never read a DB endpoint | +| A14-004 | high | security | `.github/workflows/cleanup-staging.yml:401` | GCP staging cleanup deletes a Cloud SQL instance chosen by a substring filter, then orphans it in state | +| A14-005 | high | correctness | `scripts/entrypoint.sh:72` | `set -e` makes the entrypoint's migration exit-code handling unreachable | +| A14-006 | high | security | `.github/workflows/README.md:446` | The workflows README tells operators to provision long-lived cloud credentials the workflows never use | +| A14-008 | high | security | `.pre-commit-config.yaml:196` | The only gating Trivy IaC scan skips the AWS Terraform environment | +| A14-010 | high | security | `scripts/setup-git-secrets.sh:104` | git-secrets allowed patterns are content regexes, so `resource `, `var.`, `data.` whitelist most of the repo | +| A15-001 | high | ops | `frontend/package-lock.json:5754` | npm audit fails on the pinned commit: fast-uri 3.1.5 carries four high-severity advisories | +| A16-001 | high | test-gap | `internal/purchase/approvals.go:284` | 4-eyes: the "creator's account could not be resolved" fail-closed branch has no test | +| A16-004 | high | test-gap | `internal/purchase/scheduled_fire_test.go:157` | The scheduled-purchase fire success path is behind a t.Skip, and its stand-in guard only checks a signature | +| A16-005 | high | test-gap | `internal/server/handler_test.go:142` | FinalizeInFlightRevocations has zero coverage; its only "test" asserts a stub's own return value | +| A16-006 | high | test-gap | `internal/purchase/execution.go:269` | The AUDIT LOSS branch in the per-account fan-out is untested, and it is the one that decides SQS ack vs redelivery | +| A16-007 | high | test-gap | `internal/scheduler/scheduler.go:912` | The partial-sweep eviction guard is proven for Azure only; the GCP per-account collector duplicating it has no test | +| A16-008 | high | test-gap | `internal/api/handler_purchases.go:147` | getPlannedPurchases' account-scope filter is never exercised; the only test runs as an unrestricted admin | +| A16-009 | high | test-gap | `internal/api/handler_purchases_revoke.go:320` | authorizeSessionRevokeExecution's allow and deny boundaries are both untested | + +92 of the 500 actionable findings are critical or high. + +## Findings + +Grouped by category, then by verified severity, then by shard. Each block is reproduced exactly +as the reviewer wrote it and the verifier annotated it. The `- issue:` line is the only addition. + +### Category: security + +78 findings: 2 critical, 21 high, 31 medium, 24 low. + +### A03-001 Resource wildcard `execute:*` bypasses the admin money-verb carve-out at both the grant ceiling and enforcement +- category: security +- severity: critical +- location: internal/auth/group_ceiling.go:85 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The carve-out set `adminCarvedOuts` is keyed on exact (action, resource) pairs, but enforcement (`checkPermissionMatch`, service_group.go:375-383; `AuthContext.HasPermission`, types.go:159-166) treats a stored `Resource == "*"` as matching every resource. An admin (holding only `admin:*`) sends `POST /api/groups` with permissions `[{execute,*},{approve-any,*},{retry-any,*}]`: the ceiling looks up `adminCarvedOuts[{execute,"*"}]` (false), then `grantCeilingAllows` returns true because the admin branch only refuses carved-out pairs. The group is written. Any member now passes `permissionsAllow(execute, purchases)`, `(approve-any, purchases)`, `(retry-any, purchases)` and `(execute, ri-exchange)`. Reproduced: `CreateGroupAPI(execute:*, approve-any:*, retry-any:*) by admin -> err=`; `HasPermission(execute,purchases)=true`, `HasPermission(execute,ri-exchange)=true` with `[{admin,*},{execute,*}]`, while the baseline `admin:*` alone gives false. Issues #923/#1550/#1644 separation of duties is void for every admin. +- evidence: + ```go + for _, req := range requested { + if adminCarvedOuts[[2]string{req.Action, req.Resource}] { + if permissionCoveredBy(existing, req) { + continue + } + return fmt.Errorf( + "%w: %s:%s is reserved for separation of duties (issue #923) and cannot be granted through the API", + ErrPermissionNotGrantable, req.Action, req.Resource) + } + if !grantCeilingAllows(actorPerms, req) { + ``` +- suggested fix: Make the carve-out test resource-aware in one shared helper (`isCarvedOut(action, resource)` returning true when the action is a carved-out action and the resource is the carved-out resource OR `ResourceAll`), and use it in `checkGrantCeiling`, `grantCeilingAllows`, `permissionsAllow`, `AuthContext.HasPermission` and `firstUnheldCarvedOut`; additionally refuse `Resource == "*"` for the carved-out actions in `validateRequestedPermissions` so a wildcard can never be stored for them. +- verdict: CONFIRMED — traced POST /api/groups (internal/api/handler_groups.go:32-53, gated only on create:groups) into checkGrantCeiling, where the exact-pair lookup at internal/auth/group_ceiling.go:85 misses {execute,"*"}, grantCeilingAllows:348-356 then returns true on the admin branch, validateRequestedPermissions:128-145 only rejects blanks, and at enforcement checkPermissionMatch (service_group.go:379) and AuthContext.HasPermission (types.go:162) accept a stored "*" resource for execute:purchases and execute:ri-exchange; the session execute path (handler_purchases.go:2128 -> HasPermissionForConstraintsAPI -> permissionsAllow) has no second check. +- issue: (pending cross-reference) + +### A08b-001 IMDS blocklist matches literal host text, so IPv6-mapped, alternate-radix and DNS forms of 169.254.169.254 pass through +- category: security +- severity: critical +- location: pkg/httpclient/httpclient.go:41 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (entry point `providers/azure/internal/httpclient/httpclient.go:18`) +- failure scenario: `blockIMDSDialer.DialContext` receives the address string built from the URL host, before name resolution, and compares it to a two-entry map. A URL of `http://[::ffff:169.254.169.254]/metadata/identity/oauth2/token` yields host `::ffff:169.254.169.254`, which is not a map key, so the dial proceeds and connects to the IPv4 metadata endpoint. The same holds for `http://2852039166/`, `http://0xa9fea9fe/`, `http://0251.0376.0251.0376/` (getaddrinfo accepts all three on the cgo resolver), and for any attacker-controlled DNS name whose A record is 169.254.169.254, including a name that resolves benignly on the first lookup and to the metadata IP on the second (DNS rebinding). Every Azure client reaches this dialer with a managed-identity-bearing process, so a redirect or attacker-influenced pricing URL exfiltrates the subscription's credentials. +- evidence: + ```go + func (d *blockIMDSDialer) DialContext(ctx context.Context, network, addr string) (net.Conn, error) { + host, _, err := net.SplitHostPort(addr) + if err != nil { + host = addr + } + if imdsAddresses[host] { + return nil, fmt.Errorf("connection to metadata endpoint %s is blocked", host) + } + return d.inner.DialContext(ctx, network, addr) + } + ``` +- suggested fix: resolve the host first (`net.DefaultResolver.LookupIPAddr`), then reject if any resolved `netip.Addr` (after `Unmap()`) falls in a denied set, and dial only the vetted IPs so the resolution that was checked is the resolution that is used. +- verdict: CONFIRMED — `DialContext` compares the pre-resolution host string from the URL against a two-key map (pkg/httpclient/httpclient.go:41-49), so `::ffff:169.254.169.254` and any DNS name whose A record is the metadata IP dial straight through; the Azure package delegates to this exact function (providers/azure/internal/httpclient/httpclient.go:17-19). +- issue: (pending cross-reference) + +### A01-002 Session-authed approve, cancel, retry and revoke never check the session's allowed_accounts scope +- category: security +- severity: high +- location: internal/api/handler_purchases.go:685 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A user in a group with `approve-any:purchases` (or `retry-any`, `cancel-any`, `revoke-any`) and `allowed_accounts=[acct-A]` calls `POST /api/purchases/approve/{id}` for a pending execution whose recs target acct-B. `approvePurchaseViaSession` runs only `authorizeSessionApprove` (verb + creator match) and `requireDifferentApprover`; `requireExecutionAccess`/`validatePurchaseRecommendationScope` are called only by pause/resume/run/delete (lines 263, 287, 311, 411). The purchase in acct-B executes. Same gap in `loadAndValidateRetryRequest` (1699, mints a new approvable execution for the out-of-scope account), `cancelPurchaseViaSession` (1148), `tryRevokeViaSession` (1329) and `revokeScheduledExecution` (handler_purchases_revoke.go:249). The per-permission `AccountIDs` constraint does not cover this axis (see the rationale on `requireAzureSubscriptionScope`, handler_ri_exchange.go:779-788). No test in `handler_purchases_test.go` registers `GetAllowedAccountsAPI` for these paths, so the suite passes with the gap present. +- evidence: + ```go + if execution.Status != "pending" && execution.Status != "notified" { + return nil, NewClientError(409, ...) + } + + if err := h.authorizeSessionApprove(ctx, session, execution); err != nil { + return nil, err + } + if err := h.requireDifferentApprover(ctx, session, execution); err != nil { + return nil, err + } + ``` +- suggested fix: Call `h.requireExecutionAccess(ctx, session, execution.ExecutionID)` (or a scope check over `execution.Recommendations` for plan-less rows) right after RBAC in `approvePurchaseViaSession`, `cancelPurchaseViaSession`, `loadAndValidateRetryRequest`, `tryRevokeViaSession` and `revokeScheduledExecution`, and add one scoped-session test per path asserting the mutation is not called. +- verdict: CONFIRMED — requireExecutionAccess is called only at internal/api/handler_purchases.go:263,287,311,411 (pause/resume/run/delete); approvePurchaseViaSession:665-735, cancelPurchaseViaSession:1131-1183, loadAndValidateRetryRequest:1672-1719, tryRevokeViaSession:1324-1346 and revokeScheduledExecution (internal/api/handler_purchases_revoke.go:240-300) run only verb/creator RBAC (authorizeSession*), and the only tests registering GetAllowedAccountsAPI in handler_purchases_test.go are for getPlannedPurchases/pause/execute/deletePlanned (lines 1881, 2433, 3684-4354), none for approve/cancel/retry/revoke. +- issue: (pending cross-reference) + +### A02-001 Dashboard commitment KPIs skip allowed_accounts scoping when the caller supplies an explicit account filter +- category: security +- severity: high +- location: internal/api/handler_dashboard.go:38 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A session whose group allowed_accounts is ["Production"] calls `GET /api/dashboard/summary?account_id=123456789012` where 123456789012 is the AWS account number of the out-of-scope "Finance" account (or `account_ids=`). `resolveSingleAccountFilterIDs` (scoping.go:277-279) maps the unknown value to an external-id filter under the "" provider key, the non-empty result skips the `resolveAllowedAccountScope` fallback, and `calculateCommitmentMetrics` -> `fetchCommitmentPurchases` -> `GetActivePurchaseHistory` (store_postgres.go:2112-2114, `account_id = ANY(...)`) aggregates Finance's rows. The response returns Finance's `active_commitments`, `committed_monthly`, `ytd_savings` and per-service `current_savings`. Only the recommendations half is scope-filtered (`filterDashboardRecommendations`); the purchase-history half has no equivalent, unlike `/api/inventory/*` which runs `filterPurchaseHistoryByAllowedAccounts` after the same fetch (handler_inventory.go:47, 197) and `/api/history/analytics` which validates the requested account against the scope first (handler_analytics.go:233-249). +- evidence: + ```go + accountUUIDs, accountExternalIDsByProvider, err := h.resolveDashboardAccountScope(ctx, params) + ... + if len(accountUUIDs) == 0 && len(accountExternalIDsByProvider) == 0 { + accountUUIDs, accountExternalIDsByProvider, err = h.resolveAllowedAccountScope(ctx, session) + ... + } + ... + activeCommitments, committedMonthly, ytdSavings, currentSavingsByService := h.calculateCommitmentMetrics(ctx, params["provider"], accountUUIDs, accountExternalIDsByProvider) + ``` +- suggested fix: When the session is scoped, validate every explicitly requested account (UUID or external id) with `validateAnalyticsAccountScope`-style `AccountScope.Allows` before `calculateCommitmentMetrics`, or intersect the explicit filter with `resolveAllowedAccountScope` and short-circuit to zeroed KPIs on an empty intersection. Add a test with `grantScoped` + an out-of-scope `account_id`. +- verdict: CONFIRMED — handler_dashboard.go:38-42 skips resolveAllowedAccountScope as soon as resolveSingleAccountFilterIDs (scoping.go:277-279) returns a non-empty external-id set for the unknown value, and calculateCommitmentMetrics/fetchCommitmentPurchases (handler_dashboard.go:625-680) pass it straight to GetActivePurchaseHistory (store_postgres.go:2050-2058) with no AccountScope.Allows check; every getDashboardSummary test (handler_dashboard_test.go:79,142,186,1212-1396) omits account_id/account_ids. +- issue: (pending cross-reference) + +### A03-002 An admin can mint a user API key carrying `execute:*` in a single request and spend with it +- category: security +- severity: high +- location: internal/auth/service_apikeys.go:128 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `validateAPIKeyPermissions` checks each requested key permission with `authCtx.HasPermission(perm.Action, perm.Resource)`. For an admin, `HasPermission("execute", "*")` returns true (the carve-out lookup is `adminCarvedOuts[{execute,"*"}]`, false). The key is stored with `[{execute,*}]`; `computeEffectivePermissionsFromAuthCtx` keeps it for the same reason, and `HasAPIKeyPermissionAPI` then answers `HasPermission(execute, purchases) == true`. No group edit, no second actor, one request. Reproduced: `CreateAPIKey(execute:*) -> err=`; `effective=[{execute *}] HasPermission(execute,purchases)=true`. Same root cause as A03-001 through a different guard. +- evidence: + ```go + for _, perm := range permissions { + if !authCtx.HasPermission(perm.Action, perm.Resource) { + return fmt.Errorf("user does not have permission for action=%s resource=%s", perm.Action, perm.Resource) + } + } + ``` +- suggested fix: Route key-permission validation through the same carve-out-aware helper as A03-001, and reject `Resource == "*"` on carved-out actions when a key is created; a key should never be able to hold a money verb the owner does not hold explicitly. +- verdict: CONFIRMED — validateAPIKeyPermissions (service_apikeys.go:128-131) accepts {execute,"*"} for an admin because AuthContext.HasPermission (types.go:150-155) only carves out the exact pair, computeEffectivePermissionsFromAuthCtx:462-473 keeps it for the same reason, and HasAPIKeyPermissionAPI (service_apikeys_api.go:363-364) then grants execute:purchases; the direct-execute path is partly narrowed because HasAPIKeyPermissionForConstraintsAPI:411 also requires the OWNER's group permissions to allow execute:purchases, but runPlannedPurchase (handler_purchases.go:307), approve-any and retry-any gates call requirePermission alone, so the key still spends. +- issue: (pending cross-reference) + +### A03-003 Self-membership guard checks carved-out pairs exactly, so an admin can join a wildcard group they created +- category: security +- severity: high +- location: internal/auth/service_user.go:551 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: After A03-001 stores a group with `{execute,*}`, the admin edits their own membership to add it. `guardSelfEscalation` passes (update:users held), `guardSelfCarvedOutGrant` calls `firstUnheldCarvedOut`, which skips every permission whose exact pair is not in `adminCarvedOuts`; `{execute,*}` is skipped, the change is written, and `HasPermissionAPI(admin, execute, purchases)` becomes true. Reproduced: `self-add to execute:* group -> err=`; `HasPermissionAPI(admin, execute, purchases) after self-add = true`. Two requests from one compromised admin account drain commitments, which is exactly what types.go:123 says cannot happen. +- evidence: + ```go + for i, perm := range group.Permissions { + if !adminCarvedOuts[[2]string{perm.Action, perm.Resource}] { + continue + } + if permissionsAllow(held, perm.Action, perm.Resource, nil) { + continue + } + return &group.Permissions[i] + } + ``` +- suggested fix: Use the shared carve-out-aware helper from A03-001 here as well. +- verdict: CONFIRMED — PUT /api/users/{id} (handler_users.go:123-137) reaches guardSelfEscalation (service_user.go:440-452), which passes update:users for admin:* and then guardSelfCarvedOutGrant -> firstUnheldCarvedOut, whose exact-pair test at service_user.go:551 skips a group permission {execute,"*"} and returns nil, so the self-add is written and permissionsAllow later matches the wildcard resource (service_group.go:379). +- issue: (pending cross-reference) + +### A03-004 User-membership writes have no grant ceiling: any create:users or update:users holder can mint an Administrator or Purchaser +- category: security +- severity: high +- location: internal/auth/service_user.go:396 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `guardGroupChange` returns nil for every non-self edit, and `validateCreateUserRequest` (service_user.go:122-148) never sees an actor at all; the API layer gates `POST /api/users` on `create:users` and `PUT /api/users/{id}` on `update:users` only (internal/api/handler_users.go:32, 123). (a) An admin, who by policy cannot add themself to Purchaser, creates a second user with `group_ids=[DefaultPurchaserGroupID]` and a password they chose, logs in as it, and executes purchases: the two-person control in `guardSelfCarvedOutGrant` is defeated by a puppet account. (b) A non-admin custom group holding only `update:users` (grantable by any admin through the ceiling) can move any user, including a second account it controls, into the Administrators group; the group-permission ceiling (#1550) is bypassed via membership. Group *permission* writes are ceilinged; group *membership* writes are not. +- evidence: + ```go + // Internal callers (actorUserID == "") are already trusted and skip it. + ... + if actorUserID == "" || actorUserID != targetUserID { + return nil + } + return s.guardSelfEscalation(ctx, prior, next) + ``` +- suggested fix: Apply the same ceiling to membership grants as to permission grants: for a non-self edit and for create, resolve the target groups' permissions and refuse any group that carries a permission the actor does not hold (reusing `grantCeilingAllows`), and refuse assignment of any group carrying a carved-out verb through the API unless the actor already holds that verb. Thread `actorUserID` into `CreateUser`/`CreateUserAPI`. +- verdict: CONFIRMED — guardGroupChange returns nil for every non-self edit (service_user.go:396-398), CreateUserAPI passes no actor (service_api.go:202-212) and validateCreateUserRequest:122-148 checks only email/groups/password, the handlers gate on create:users / update:users alone (handler_users.go:32, :123), update:users is not in adminCarvedOuts (types.go:124-134) so an admin can grant it through the ceiling, and nothing compares the target groups' permissions to the actor's; the comment at service_user.go:514-516 explicitly accepts the non-self Purchaser add as design, but that design does not cover a puppet account or an update:users-only actor moving accounts into Administrators. +- issue: (pending cross-reference) + +### A03-005 Deactivating a user does not end their sessions, and session validation never re-checks the user +- category: security +- severity: high +- location: internal/auth/service.go:316 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: An admin sets `active=false` on a user via `UpdateUser`; `guardDeactivation` only protects the last admin and nothing calls `DeleteUserSessions` (service_user.go:319-333, the only session purges are in DeleteUser, password change and reset). `ValidateSession` checks only the sessions row (`WHERE token = $1 AND expires_at > NOW()`, store_postgres.go:670), `requireSessionPermission` (internal/api/handler.go:301-314) then resolves permissions through `GetUserPermissions`, which never reads `user.Active`, and `/usr/bin/grep -rn "\.Active" internal/api` (non-test) returns nothing. The deactivated user keeps full group-derived access through every session-authenticated endpoint, including purchase execution if they hold the verbs, for the rest of `sessionDuration` (24h default). Contrast: the API-key path (`lookupAPIKeyUser`, service_apikeys.go:272) and `CreateAPIKey` (service_apikeys.go:54) do check Active, so the two credential types behave differently on the same input. +- evidence: + ```go + session, err := s.store.GetSession(ctx, hashedToken) + if err != nil { + return nil, err + } + if session == nil { + return nil, fmt.Errorf("session not found") + } + // Constant-time comparison after fetch ... + ``` +- suggested fix: In `UpdateUser`, when `priorActive && !user.Active`, call `DeleteUserSessions` after the row write; and have the session path fail closed like the API-key path by loading the user in `ValidateSession` (or in `requireSessionPermission`) and refusing when missing or inactive. +- verdict: CONFIRMED — UpdateUser (service_user.go:300-345) never calls DeleteUserSessions (the only callers are DeleteUser:636, UpdateUserProfile:692 and the two password paths in service_password.go:255,359), ValidateSession (service.go:316-345) only reads the sessions row via GetSession (store_postgres.go:670), requireSessionPermission (handler.go:301-314) resolves permissions through HasPermissionAPI with no user load, and `/usr/bin/grep -rn "\.Active\b" internal/api` (non-test) returns nothing, while lookupAPIKeyUser (service_apikeys.go:272) does refuse inactive owners. +- issue: (pending cross-reference) + +### A03-006 A deactivated account can reactivate itself through the forgot-password flow +- category: security +- severity: high +- location: internal/auth/service_password.go:355 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `RequestPasswordReset` has no `Active` check (service_password.go:267-333), so a reset email is sent to a deactivated user. `ConfirmPasswordReset` then flips `Active = true` unconditionally, because it cannot tell an invited user (never activated) from an admin-deactivated one. A user an admin disabled (offboarding, suspected compromise) clicks "Forgot password" and is back in, with all prior group memberships. Reproduced: `reset email sent to deactivated account: 1`; `after ConfirmPasswordReset: deactivated user Active=true`. +- evidence: + ```go + // Activate user on first password set (admin bootstrap flow) + if !user.Active { + user.Active = true + } + ``` +- suggested fix: Distinguish "invited" from "deactivated" explicitly (the invite path already sets `PasswordResetToken` at create time; record an invite marker, or treat "never had a real password / `LastLoginAt == nil`" as the invite signal) and only activate on that path; refuse `RequestPasswordReset` for deactivated, non-invited users while keeping the anti-enumeration response. +- verdict: CONFIRMED — RequestPasswordReset (service_password.go:267-333) loads the user and issues a token with no Active check, ConfirmPasswordReset:354-357 sets Active = true for any inactive user, and CreateUser (service_user.go:203-224) records nothing that distinguishes an invite (Active=false + setup token) from an admin deactivation via applyUpdateUserRequest:611-613, so the two states are indistinguishable at confirm time. +- issue: (pending cross-reference) + +### A06-003 SourceIP carries the TCP port (or the proxy's IP), defeating the login brute-force rate limit +- category: security +- severity: high +- location: internal/server/http.go:303 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `SourceIP` is the sole rate-limit key for the credential endpoints — `checkRateLimitStrict` builds `"IP#" + req.RequestContext.HTTP.SourceIP` (internal/api/middleware.go:557, internal/api/db_rate_limiter.go:212) and login has no email-keyed limiter. On any container deployment that is not behind an `X-Forwarded-For`-setting proxy (Fargate behind an NLB, direct container, local), `r.RemoteAddr` is `"10.0.0.7:54321"`, so every new TCP connection produces a distinct key: an attacker gets unlimited password guesses and one `rate_limits` row per attempt. In the other direction, on Cloud Run the header is `client, google-lb`, so the rightmost entry is the load balancer and every client on the deployment shares a single bucket — one noisy user locks everyone out of login. The port is never stripped (no `net.SplitHostPort`), and the only test, `TestHttpToLambdaRequest_XForwardedFor` (internal/server/app_test.go:91), exercises only the header-present case. +- evidence: + ```go + sourceIP := r.RemoteAddr + if xff := r.Header.Get("X-Forwarded-For"); xff != "" { + parts := strings.Split(xff, ",") + sourceIP = strings.TrimSpace(parts[len(parts)-1]) + } + ``` +- suggested fix: strip the port with `net.SplitHostPort(r.RemoteAddr)` for the no-XFF case, and select the client entry by trusted-proxy-hop count rather than always taking the rightmost element. +- verdict: CONFIRMED — `sourceIP := r.RemoteAddr` with no `net.SplitHostPort` anywhere on the path (internal/server/http.go:303-307,315), and that value is the sole rate-limit key for the credential endpoints: `checkRateLimitStrict` passes `req.RequestContext.HTTP.SourceIP` to `AllowWithIP`, which formats `"IP#%s"` (internal/api/middleware.go:557-559, internal/api/db_rate_limiter.go:211-215); login has no email-keyed companion (internal/api/handler_auth.go:24 vs the only `AllowWithEmail` caller at handler_auth.go:268 for forgot_password). +- issue: (pending cross-reference) + +### A08b-002 IMDS denylist covers two addresses; Azure WireServer and the rest of link-local are reachable +- category: security +- severity: high +- location: pkg/httpclient/httpclient.go:31 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: this is the Azure provider's hardened client, yet the map omits `168.63.129.16`, Azure's WireServer / host-agent endpoint that serves instance configuration and extension settings, and omits the rest of `169.254.0.0/16` and `fe80::/10`. An SSRF into `http://168.63.129.16/machine?comp=goalstate` is not blocked. `metadata.google.internal` (a name, so also missed by A08b-001) and Alibaba's `100.100.100.200` are likewise reachable. +- evidence: + ```go + var imdsAddresses = map[string]bool{ + "169.254.169.254": true, // AWS/Azure/GCP link-local IMDS (IPv4) + "fd00:ec2::254": true, // AWS IMDS (IPv6) + } + ``` +- suggested fix: replace the map with a CIDR denylist (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::/64`, plus loopback and the RFC1918 ranges if outbound is meant to be internet-only) evaluated against resolved addresses. +- verdict: CONFIRMED — `imdsAddresses` at pkg/httpclient/httpclient.go:31-34 holds exactly the two literal strings quoted; there is no CIDR test anywhere in the file and no entry for 168.63.129.16, 100.100.100.200 or `metadata.google.internal`. +- issue: (pending cross-reference) + +### A08b-004 Four Azure clients accept a nil HTTP client in `NewClientWithHTTP` with no hardened fallback +- category: security +- severity: high +- location: providers/azure/services/cache/client.go:103 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same shape at `cosmosdb/client.go:101`, `database/client.go:125`, `search/client.go:80`) +- failure scenario: `managedredis/client.go:95` and `synapse/client.go:77` guard the nil case and fall back to `httpclient.New()`; these four do not. A caller passing nil stores a nil `HTTPClient` interface, and the first pricing or purchase call panics on `c.httpClient.Do`. Any caller that instead passes its own bare `&http.Client{}` silently loses the IMDS dialer on a path that carries an ARM bearer token (`DoIdempotentPurchaseTwoStep` at `cache/client.go:345`). +- evidence: + ```go + func NewClientWithHTTP(cred azcore.TokenCredential, subscriptionID, region string, httpClient HTTPClient) *CacheClient { + return &CacheClient{ + cred: cred, + subscriptionID: subscriptionID, + region: region, + httpClient: httpClient, + } + } + ``` +- suggested fix: add the same `if httpClient == nil { httpClient = httpclient.New() }` guard the two fixed siblings already carry, and add the regression test that asserts the nil branch blocks 169.254.169.254. +- verdict: CONFIRMED — cache/client.go:103-110, cosmosdb/client.go:101-108, database/client.go:125-132 and search/client.go:80-87 assign the parameter straight into the struct, while synapse/client.go:77-79 and managedredis/client.go:95-97 carry the nil guard the finding names. +- issue: (pending cross-reference) + +### A09-001 IMDS blocklist is bypassed by any hostname that resolves to the metadata address +- category: security +- severity: high +- location: pkg/httpclient/httpclient.go:41 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `http.Transport` hands `DialContext` the URL host, not a resolved IP; `net.Dialer.DialContext` does the DNS lookup afterwards. With a JWKS URL of `http://metadata.attacker.example/latest/meta-data/iam/security-credentials/` whose A record is `169.254.169.254`, `net.SplitHostPort` yields `metadata.attacker.example`, the map lookup misses, and the connection to IMDS proceeds. The only control point every provider delegates to therefore blocks literal IPs and nothing else. The same hole is reachable through a 302 redirect, since redirects are followed by default and re-dial through the identical path. +- evidence: + ```go + func (d *blockIMDSDialer) DialContext(ctx context.Context, network, addr string) (net.Conn, error) { + host, _, err := net.SplitHostPort(addr) + if err != nil { + host = addr + } + if imdsAddresses[host] { + return nil, fmt.Errorf("connection to metadata endpoint %s is blocked", host) + } + return d.inner.DialContext(ctx, network, addr) + } + ``` +- suggested fix: dial with `d.inner.DialContext` then inspect `conn.RemoteAddr()` (or resolve first via `net.DefaultResolver.LookupIPAddr` and dial the vetted IP), rejecting and closing when the peer IP is link-local/metadata rather than matching on the pre-resolution host string. +- verdict: CONFIRMED — `blockIMDSDialer.DialContext` matches `imdsAddresses[host]` on the host string `http.Transport` passes it before `net.Dialer` resolves it (pkg/httpclient/httpclient.go:41-48), so a hostname with an A record of 169.254.169.254 misses the map and is dialed; no post-dial `RemoteAddr` check exists anywhere in the file. +- issue: (pending cross-reference) + +### A09-002 IMDS blocklist misses the ECS/EKS credential endpoints and the rest of 169.254.0.0/16 +- category: security +- severity: high +- location: pkg/httpclient/httpclient.go:31 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the map holds exactly two addresses. A request to `http://169.254.170.2/v2/credentials/` (the ECS task credential provider) or `http://169.254.170.23/v1/credentials` (EKS Pod Identity) is dialed normally and returns live IAM credentials. `http://169.254.169.254./latest/...` (trailing-dot FQDN form) and `http://[fd00:ec2:0:0:0:0:0:254]/` (expanded IPv6, which `net.SplitHostPort` returns verbatim rather than in canonical form) also miss the map. The package doc claims it "blocks connections to the cloud Instance Metadata Service (IMDS) endpoints", which is stronger than what the two entries deliver. +- evidence: + ```go + var imdsAddresses = map[string]bool{ + "169.254.169.254": true, // AWS/Azure/GCP link-local IMDS (IPv4) + "fd00:ec2::254": true, // AWS IMDS (IPv6) + } + ``` +- suggested fix: replace the string map with `net.ParseIP` plus CIDR checks over `169.254.0.0/16`, `fe80::/10` and `fd00:ec2::/32`, applied to the resolved peer address (see A09-001) so every link-local credential endpoint is covered by construction. +- verdict: CONFIRMED — the map at pkg/httpclient/httpclient.go:31-34 holds exactly the two literals quoted; there is no CIDR check, no `net.ParseIP` normalization and no trailing-dot handling in the file, so 169.254.170.2 and 169.254.170.23 pass straight to `d.inner.DialContext` at line 47. +- issue: (pending cross-reference) + +### A10-018 The GCP setup wizard grants `roles/compute.admin` at project scope +- category: security +- severity: high +- location: cmd/configure_gcp.go:686 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Step 4 of `cudly configure-gcp` grants the CUDly service account `roles/compute.admin` on the operator's project, then Step 5 mints a long-lived JSON key for it and Step 6 uploads that key to AWS Secrets Manager. `compute.admin` confers full control of every Compute Engine resource: creating and deleting VMs, disks, images, firewall rules and networks. CUDly needs to read usage and purchase committed use discounts. Anyone who obtains the stored key can delete the project's production infrastructure. The project's own IAM policy in CLAUDE.md names this exact role as the anti-pattern ("Prefer custom roles ... over broad predefined roles like `roles/compute.admin`"), and the prompt defaults to Run on empty input. +- evidence: + ```go + member := fmt.Sprintf("serviceAccount:%s", saEmail) + role := "roles/compute.admin" + ... + fmt.Printf("[R]un, [S]kip? (grants %s to %s on project %s via SDK) ", role, saEmail, projectID) + ``` +- suggested fix: Grant the narrowest predefined pair the CUD flow needs (`roles/compute.viewer` plus the commitment-purchase permissions) or provision a `google_project_iam_custom_role` holding only `compute.commitments.*` and the usage reads, matching the runtime-permissions rule in CLAUDE.md. +- verdict: CONFIRMED — gcpStepGrantRole hardcodes roles/compute.admin at project scope (cmd/configure_gcp.go:684-710) and its `case "r", "run", "":` arm at :700 makes empty input Run, and the identity it is granted to gets a JSON key minted at :722 and pushed to Secrets Manager at cmd/configure_gcp.go:162. +- issue: (pending cross-reference) + +### A11-002 Removing every permission from a group reports success while the backend keeps the old permissions +- category: security +- severity: high +- location: frontend/src/groups/groupModals.ts:112 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `collectPermissions()` skips any row whose action or resource is empty (groupModals.ts:418), so an operator who clicks Remove on every permission row and saves sends `permissions: []`. The backend's `applyUpdateGroupRequest` only assigns when the list is non-empty (`internal/auth/service_api.go:330 — if len(perms) > 0`), so the stored permissions are left untouched. The UI then shows "Group updated successfully" and reloads a group that still carries, say, `admin:*`. An admin who believes they have just stripped a group's privileges has not. The same shape applies to a group whose only remaining row is left on the "Select Action" placeholder. +- evidence: + ```typescript + // groups/groupModals.ts:107-117 + const permissions = collectPermissions(); + try { + if (currentEditingGroup) { + await api.updateGroup(currentEditingGroup.id, { name, description, permissions }); + showSuccess('Group updated successfully'); + ``` +- suggested fix: refuse to submit an empty permission list from the form with an explicit message ("a group must grant at least one permission; delete the group instead"), since the API cannot express "clear all permissions". +- verdict: CONFIRMED — collectPermissions skips rows with an empty action or resource (frontend/src/groups/groupModals.ts:417-418) and saveGroup sends the result unconditionally (groupModals.ts:107-117), while the backend's `if len(perms) > 0` at internal/auth/service_api.go:330 documents empty as "not sent" and leaves `group.Permissions` untouched, so the success toast at groupModals.ts:113 reports a strip that never happened; note the UI path is currently masked by A11-001, which never binds the submit handler. +- issue: (pending cross-reference) + +### A13-004 The deploy role holds unconditioned account-wide KMS Decrypt, Encrypt and CreateGrant +- category: security +- severity: high +- location: terraform/environments/aws/ci-cd-permissions/policy_networking.tf:166 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A leaked GitHub Actions OIDC token assumes `cudly-terraform-deploy` and calls `kms:Decrypt` against any CMK in the account, including keys belonging to unrelated workloads that share it, and `kms:CreateGrant` to hand `Decrypt` on any CMK to an arbitrary grantee principal it controls, which survives the deploy role being revoked. This is the one statement in the whole `ci-cd-permissions/` directory with no comment justifying `Resource = "*"`, and it directly contradicts the sibling policies: `KMSReadTaggedOnly` (policy_compute_b.tf:160) gates the far weaker `kms:GetKeyPolicy` on `Project=CUDly` explicitly to stop "account-wide key-policy reconnaissance". All four actions support `aws:ResourceTag`. +- evidence: + ```hcl + Sid = "KMS" + Effect = "Allow" + Action = [ + "kms:CreateGrant", + "kms:Decrypt", + "kms:DescribeKey", + "kms:Encrypt", + "kms:GenerateDataKey", + ] + Resource = "*" + ``` +- suggested fix: Add the same `StringEqualsIgnoreCase` on `aws:ResourceTag/Project` used by `KMSMutateTaggedOnly`, keeping only `kms:DescribeKey` unconditioned if a plan-time lookup needs it. +- verdict: CONFIRMED — policy_networking.tf:164-175 is the `KMS` Sid with `Resource = "*"` and no `Condition` block, while every other KMS grant in the directory is gated: `KMSAliasMutate` on an ARN prefix (policy_compute_b.tf:135-141), `KMSReadTaggedOnly` and `KMSTagOnCreate` on `aws:ResourceTag`/`aws:RequestTag` Project (policy_compute_b.tf:160-190), and `KMSMutateTaggedOnly` in policy_compute.tf:347-375. Since AWS CMKs carry the default key policy that delegates to IAM, an unconditioned `kms:Decrypt`/`kms:CreateGrant` here really does reach unrelated keys in the account. The finding's aside that this is the only `Resource = "*"` statement lacking a comment is inaccurate (ACM at :139 and Route53 at :152 are also uncommented `"*"`), but that does not affect the defect. +- issue: (pending cross-reference) + +### A13-005 The ACR admin password is de-sensitized into a non-sensitive module variable +- category: security +- severity: high +- location: terraform/environments/azure/build.tf:24 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `nonsensitive()` strips the provider's sensitivity marker from `azurerm_container_registry.main.admin_password`, and the receiving variable `registry_login_command` (terraform/modules/build/variables.tf:43) carries no `sensitive = true`. The registry admin password, a long-lived push credential for the image the production Container App runs, is therefore stored unredacted in the module input and rendered in plain text wherever Terraform prints that value, including CI plan output. The adjacent `terraform/environments/azure/secrets.tf:42` carries an explicit "do NOT wrap this merge in nonsensitive()" warning for exactly this reason. +- evidence: + ```hcl + registry_login_command = "echo '${nonsensitive(azurerm_container_registry.main.admin_password)}' | docker login ${azurerm_container_registry.main.login_server} -u ${azurerm_container_registry.main.admin_username} --password-stdin" + ``` +- suggested fix: Mark `registry_login_command` `sensitive = true` in `terraform/modules/build/variables.tf` and drop the `nonsensitive()` wrapper; better, use `az acr login` with the deploy identity so no admin credential is materialized at all. +- verdict: CONFIRMED — terraform/environments/azure/build.tf:24 wraps `azurerm_container_registry.main.admin_password` in `nonsensitive()` and terraform/modules/build/variables.tf:43-46 declares the receiving variable with no `sensitive = true`, so the marker is gone for the whole downstream path. The value is interpolated into the `local-exec` command body at terraform/modules/build/main.tf:74, which terraform echoes at apply time and persists in state, and the registry is `admin_enabled = true` (terraform/environments/azure/registry.tf:10) so this is a live long-lived push credential. The contrast the finding draws is real: terraform/environments/azure/secrets.tf:42-44 carries an explicit "do NOT wrap this merge in nonsensitive()" warning for the same class of value. +- issue: (pending cross-reference) + +### A13c-005 Azure ACR admin account is enabled and its credentials are written into the Container App and Terraform state, despite an AcrPull grant existing alongside +- category: security +- severity: high +- location: terraform/environments/azure/registry.tf:10 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `admin_enabled = true` activates a shared registry account with push as well + as pull. `environments/azure/compute.tf:102-103` reads `admin_username` / `admin_password` off + the resource and hands them to the container-apps module, which materialises them as a + `secret` block on `azurerm_container_app.main` (container-apps/main.tf:224-230). The password + therefore lands in Terraform state and in the ARM resource definition, readable by anyone with + Reader on the resource group. Line 16 of the same file already grants the container-app managed + identity `AcrPull`, so the workload can pull without any credential at all — the admin path is + redundant and strictly wider. The reusable `modules/registry/azure` defaults + `enable_admin_user = false`; the environment declares its own registry resource and overrides. +- evidence: + ```hcl + resource "azurerm_container_registry" "main" { + name = local.acr_name + sku = "Basic" + admin_enabled = true # Enables username/password login for docker push + } + ``` +- suggested fix: set `admin_enabled = false`, drop the three `registry_*` inputs at + compute.tf:101-103, and give the Container App `registry { identity = }` so the pull + runs on the existing AcrPull assignment. +- verdict: CONFIRMED — `admin_enabled = true` is set directly on the resource, not merely defaulted + (environments/azure/registry.tf:10), so the `admin_username`/`admin_password` attributes are + populated; compute.tf:102-103 reads them into the module, which emits them as a `secret` block on + `azurerm_container_app.main` (container-apps/main.tf:224-230), putting the password in state on + both the registry resource and the container app. The redundant AcrPull grant to the same identity + is at registry.tf:16-21, and modules/registry/azure/variables.tf:37 does default the flag false. +- issue: (pending cross-reference) + +### A14-004 GCP staging cleanup deletes a Cloud SQL instance chosen by a substring filter, then orphans it in state +- category: security +- severity: high +- location: .github/workflows/cleanup-staging.yml:401 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `gcloud sql instances list --filter="name:cudly-staging"` uses gcloud's `:` operator, which is a contains match, not equality. With `cudly-staging-prod-mirror` and `cudly-staging-5e4d3c2b` both present, `head -1` picks whichever sorts first and `gcloud sql instances delete --quiet` destroys it. This is exactly the over-match that `scripts/select-owned-name.sh` was written to remove for ECR (#1592/#1820) and RDS (#1821); the GCP path never got the guard, and the sweeps in `test-ecr-delete-selection.sh` / `test-rds-deletion-protection-scope.sh` key on `aws ecr delete-repository` and `aws rds modify-db-instance`, so they never look at `gcloud sql instances delete`. The `|| true` on line 405 then hides a failed delete, and lines 419-421 `terraform state rm` the instance regardless, leaving a running, billing, untracked Cloud SQL instance. The step validates `PROJECT` carefully (line 396) precisely because a silent skip is unacceptable, but applies no equivalent check to `INSTANCE`. +- evidence: + ```bash + INSTANCE=$(gcloud sql instances list --project="$PROJECT" \ + --filter="name:cudly-staging" --format="value(name)" 2>/dev/null | head -1) + if [ -n "$INSTANCE" ]; then + gcloud sql instances delete "$INSTANCE" --project="$PROJECT" --quiet || true + ``` + ```bash + terraform state rm google_sql_database_instance.main 2>/dev/null || true + ``` +- suggested fix: Read the owned instance name from `terraform output -json` and pipe the `gcloud sql instances list` output through `scripts/select-owned-name.sh`, matching the ECR/RDS pattern; drop the `|| true` on the delete. +- verdict: CONFIRMED — cleanup-staging.yml:401-405 selects by `--filter="name:cudly-staging" | head -1` and deletes with `|| true`, and lines 419-421 `terraform state rm` unconditionally outside the `if [ -n "$INSTANCE" ]` block; scripts/select-owned-name.sh is wired only into force-delete-owned-ecr-repo.sh:80 and disable-owned-rds-deletion-protection.sh:143, and the two sweeps key on `aws ecr delete-repository` (test-ecr-delete-selection.sh:278) and `aws rds modify-db-instance` (test-rds-deletion-protection-scope.sh:35), never on gcloud. +- issue: (pending cross-reference) + +### A14-006 The workflows README tells operators to provision long-lived cloud credentials the workflows never use +- category: security +- severity: high +- location: .github/workflows/README.md:446 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: An operator following the Setup Guide creates and stores three sets of static credentials: AWS access key + secret (446-448), a downloaded GCP service-account JSON key (463-467), and an Azure service principal client secret via `az ad sp create-for-rbac --sdk-auth` (476-479). None of them is read by any workflow: AWS uses OIDC via `vars.AWS_ROLE_TO_ASSUME`, GCP uses Workload Identity Federation via `vars.GCP_WORKLOAD_IDENTITY_PROVIDER`, Azure uses `azure/login` with client/tenant/subscription IDs and no secret. Grep for `AWS_SECRET_ACCESS_KEY`, `GCP_SA_KEY` and `AZURE_CREDENTIALS` across `.github/workflows/` returns nothing outside this README. Following the guide therefore creates exactly the long-lived, exfiltratable credentials the OIDC design exists to eliminate, and the GCP step also grants project-wide `roles/run.admin` (line 460), contradicting the narrow-scope rule in the project CLAUDE.md. +- evidence: + ```bash + gh secret set AWS_ACCESS_KEY_ID + gh secret set AWS_SECRET_ACCESS_KEY + ... + gcloud iam service-accounts keys create key.json \ + --iam-account=cudly-cicd@.iam.gserviceaccount.com + gh secret set GCP_SA_KEY < key.json + ... + az ad sp create-for-rbac --name cudly-cicd --sdk-auth > azure-credentials.json + gh secret set AZURE_CREDENTIALS < azure-credentials.json + ``` +- suggested fix: Replace the Setup Guide's credential sections with the OIDC/WIF variables the workflows actually read, and state explicitly that no static cloud credentials are stored. +- verdict: CONFIRMED — `/usr/bin/grep` over `.github/workflows/` returns `AWS_SECRET_ACCESS_KEY`, `GCP_SA_KEY` and `AZURE_CREDENTIALS` only inside README.md (lines 92, 93, 172, 213, 446, 447, 467, 479), while every workflow authenticates via `vars.AWS_ROLE_TO_ASSUME` (deploy-aws-lambda.yml:246), `vars.GCP_WORKLOAD_IDENTITY_PROVIDER` (deploy-gcp.yml:149) and azure/login client-id (deploy-azure.yml:168). +- issue: (pending cross-reference) + +### A14-008 The only gating Trivy IaC scan skips the AWS Terraform environment +- category: security +- severity: high +- location: .pre-commit-config.yaml:196 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The `trivy-config` hook is the only Trivy invocation in the repo that uses `--exit-code 1`. It passes `--skip-dirs terraform/environments/aws`, so a HIGH/CRITICAL misconfiguration introduced anywhere under the primary cloud's Terraform root (public S3 bucket, open security group, unencrypted RDS) is never gated. ci.yml's `Run Trivy IaC misconfiguration scanner` (line 779) scans `terraform/` but deliberately runs at the default exit code 0 and only uploads SARIF (its comment at 770-778 says so), so nothing else catches it. The hook's long comment justifies the `**/.terraform` and `.claude` skips in detail and says nothing about the AWS environment skip. +- evidence: + ```yaml + entry: bash -c 'trivy config --severity HIGH,CRITICAL --exit-code 1 --skip-dirs terraform/environments/aws --skip-dirs "**/.terraform" --skip-dirs .claude .' + ``` +- suggested fix: Remove `--skip-dirs terraform/environments/aws`, fix or explicitly `#trivy:ignore` the findings it surfaces, and document each remaining suppression. +- verdict: CONFIRMED — .pre-commit-config.yaml:196 is the only Trivy invocation in the repo using `--exit-code 1` and it skips `terraform/environments/aws`, while ci.yml:780-797's IaC step carries a comment at 770-778 stating it "uses the default exit-code 0 so misconfig findings are reported to the Security tab without gating the job". +- issue: (pending cross-reference) + +### A14-010 git-secrets allowed patterns are content regexes, so `resource `, `var.`, `data.` whitelist most of the repo +- category: security +- severity: high +- location: scripts/setup-git-secrets.sh:104 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `git secrets --add --allowed ` suppresses any scanned **line** matching the regex; it is not a path filter. `'resource\s'` therefore whitelists every Terraform resource-block line and any line containing the word "resource" followed by whitespace; `'var\.'`, `'local\.'`, `'data\.'` and `'module\.'` whitelist any line containing those substrings. A real AWS account ID or access key placed on such a line — `resource "aws_iam_access_key" "x" { secret = "AKIA…" }` — passes the local gate silently. The neighbouring `'_test\.go'` and `'testdata/'` entries show the intent was to exclude paths, which this mechanism cannot do. +- evidence: + ```bash + git secrets --add --allowed 'var\.' + git secrets --add --allowed 'data\.' + ... + git secrets --add --allowed 'resource\s' + git secrets --add --allowed 'module\.' + ``` +- suggested fix: Delete the substring allowlist entries and rely on `.gitallowed`, which is already scoped to the specific placeholder values; if path exclusions are needed, filter the file list before invoking `git secrets --scan`. +- verdict: CONFIRMED — reproduced in a throwaway repo: with `git secrets --add --allowed 'resource\s'` (scripts/setup-git-secrets.sh:104), the line `resource "aws_iam_access_key" "x" { key = "AKIA-EXAMPLE-KEY-REDACTED" }` scanned clean while the same key on a plain line exited 1, because git-secrets matches the allowed regex against whole output lines. +- issue: (pending cross-reference) + +### A01-006 AWS RI exchange execute/quote/target-offerings ignore the session's allowed_accounts scope +- category: security +- severity: medium +- location: internal/api/handler_ri_exchange.go:1806 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A user scoped to `allowed_accounts=[azure-sub-1]` who holds an unconstrained `execute:ri-exchange` calls `POST /api/ri-exchange/execute` naming the deployment AWS account's convertible RIs. `executeExchange` checks the verb, the permission Constraints (AccountIDs on the permission only) and the idempotency claim, then exchanges the RIs. The three read siblings (`listConvertibleRIs`, `getRIUtilization`, `getReshapeRecommendations`) and the Azure execute path all gate on `reshapeCloudAccountInScope`/`requireAzureSubscriptionScope`; `executeExchange`, `getExchangeQuote` (1685) and `listTargetOfferings` (117, enumerates the account's RIs) do not. `TestExecuteExchange_*` never register `GetAllowedAccountsAPI`. +- evidence: + ```go + err = h.requirePermissionConstraints(ctx, session, "ri-exchange", []auth.PermissionConstraints{{ + AccountIDs: []string{cloudAccountID}, + Providers: []string{string(common.ProviderAWS)}, + Services: []string{string(common.ServiceEC2)}, + Regions: []string{region}, + MaxPurchaseAmount: maxPayment, + }}) + if err != nil { + return nil, err + } + ``` +- suggested fix: Call `reshapeCloudAccountInScope` in `executeExchange`, `getExchangeQuote` and `listTargetOfferings` and refuse (errNotFound) when it returns false, mirroring the read endpoints. +- verdict: CONFIRMED — reshapeCloudAccountInScope is called only at internal/api/handler_ri_exchange.go:1351,1386,1539 (listConvertibleRIs/getRIUtilization/getReshapeRecommendations); executeExchange:1765-1846 runs the verb, the permission Constraints and the idempotency claim only, getExchangeQuote:1685-1727 and listTargetOfferings:117-140 check only view:purchases before calling AWS on the deployment account. +- issue: (pending cross-reference) + +### A01-009 Planned-purchase rows are minted with approval tokens that never expire +- category: security +- severity: medium +- location: internal/api/handler_plans.go:539 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `createPurchaseExecutionsTx` writes up to 52 rows with `ApprovalToken` set and `ApprovalTokenExpiresAt` nil. `validateApprovalToken` (internal/purchase/approvals.go:124) skips the TTL check when the field is nil, and the scheduler's `getOrCreateExecution` reuses the existing row (and its token) for the notification email. Every other creation site stamps `config.ApprovalTokenTTL` (`newPendingExecution` 2575, `persistRetryExecution` 1837, notifications.go:142). A forwarded or leaked approval link for a scheduled row therefore stays valid until the row leaves pending/notified, which for a step months out can be months. +- evidence: + ```go + execution := &config.PurchaseExecution{ + PlanID: planID, + ExecutionID: uuid.New().String(), + Status: "pending", + StepNumber: plan.RampSchedule.CurrentStep + i + 1, + ScheduledDate: scheduledDate, + ApprovalToken: approvalToken, + CreatedByUserID: creator, + } + ``` +- suggested fix: Do not mint the token at scheduling time; let the notification step (`getOrCreateExecution`) generate it with `ApprovalTokenExpiresAt` when the email is actually sent, or stamp `ApprovalTokenExpiresAt` relative to `scheduledDate` here. +- verdict: CONFIRMED — createPurchaseExecutionsTx (internal/api/handler_plans.go:539-547) sets ApprovalToken with no ApprovalTokenExpiresAt while newPendingExecution (handler_purchases.go:2575), persistRetryExecution (:1837) and getOrCreateExecution (internal/purchase/notifications.go:127) all stamp config.ApprovalTokenTTL; validateApprovalToken (internal/purchase/approvals.go:124) skips the TTL when nil, and getOrCreateExecution:115-119 returns the existing plan+date row whose token buildNotificationData:158 embeds in the email. +- issue: (pending cross-reference) + +### A01-012 GET /api/plans lists every plan regardless of the session's allowed_accounts +- category: security +- severity: medium +- location: internal/api/handler_plans.go:35 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A Read-Only user scoped to acct-A calls `GET /api/plans`. `listPlans` applies only the caller-supplied `account_ids` query filter and returns every plan (names, services, ramp position, health score) including plans whose accounts are all outside the user's scope. `getPlan`, `updatePlan`, `patchPlan`, `deletePlan` and `getPlannedPurchases` all enforce `requirePlanAccess`/`isPlanAllowedCached`; the list endpoint is the only unscoped read, and it is what feeds the Plans page. `TestHandler_listPlans` uses an admin session only. +- evidence: + ```go + filter := config.PurchasePlanFilter{AccountIDs: accountIDs} + plans, err := h.config.ListPurchasePlans(ctx, filter) + if err != nil { + return nil, err + } + ``` +- suggested fix: After loading, drop plans for which `isPlanAllowedCached` returns false (unrestricted sessions short-circuit), the same way `getPlannedPurchases` does. +- verdict: CONFIRMED — listPlans (internal/api/handler_plans.go:21-49) applies only the caller-supplied account_ids filter and never calls getAccountScope or requirePlanAccess, whereas getPlan:218, updatePlan:258, patchPlan:617, deletePlan:316 and getPlannedPurchases (handler_purchases.go:142 via isPlanAllowedCached) all scope per plan. +- issue: (pending cross-reference) + +### A02-008 Behind CloudFront every IP-keyed rate limit uses the edge server address, not the client +- category: security +- severity: medium +- location: internal/api/middleware.go:530 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: With `enable_cdn` the Lambda Function URL is called by CloudFront via OAC (terraform/environments/aws/compute.tf:25, 269-277). The function-URL event's `requestContext.http.sourceIp` is then the CloudFront edge IP and the viewer IP is only in `X-Forwarded-For`; internal/server/lambda.go:99 passes the event through unchanged (the plain HTTP server does parse XFF, internal/server/http.go:300-315). `login` (5/15 min), `setup_admin`, `reset_password`, `change_password`, `register` and `approve_cancel_public` are all keyed on this value. Five wrong passwords from anyone routed through the same edge lock every user on that edge out of login for 15 minutes, and an attacker gets a fresh budget per edge POP. +- evidence: + ```go + clientIP := req.RequestContext.HTTP.SourceIP + allowed, err := h.rateLimiter.AllowWithIP(ctx, clientIP, endpoint) + ``` +- suggested fix: In the Lambda transport, when the function URL is OAC-protected, replace `SourceIP` with the rightmost trusted entry of `X-Forwarded-For` (mirroring http.go), so the limiter keys on the viewer; keep `SourceIP` as-is for direct (auth_type NONE) invocations. +- verdict: CONFIRMED — checkRateLimit and checkRateLimitStrict (middleware.go:530,557) key on RequestContext.HTTP.SourceIP, internal/server/lambda.go has no X-Forwarded-For handling (only http.go:300-306 parses it), and terraform/environments/aws/compute.tf:25-26,269-277 puts CloudFront+OAC in front of the Function URL when enable_cdn=true; enable_cdn defaults to false (variables.tf:398-402), so only CDN deployments are affected. +- issue: (pending cross-reference) + +### A03-007 MFA can be replaced with only a session and the password, bypassing the proof-of-possession that MFADisable requires +- category: security +- severity: medium +- location: internal/auth/service_mfa.go:285 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `MFADisable` demands password plus a TOTP or recovery code ("a stolen session alone shouldn't disable MFA", service_mfa.go:427-431). But `MFASetup` never checks `user.MFAEnabled`; it only re-verifies the password, writes a new pending secret, and `MFAEnable` promotes it over the existing `MFASecret` and replaces all recovery codes. An attacker with a live session and the password (the exact threat MFADisable models) calls setup then enable with a code from their own authenticator: the victim's authenticator stops working, the attacker's is enrolled, and the victim's recovery codes are gone. The re-enrol path is strictly weaker than the disable path it is equivalent to. +- evidence: + ```go + if !s.verifyPassword(password, user.PasswordHash) { + return nil, fmt.Errorf("%w", ErrMFAInvalidPassword) + } + secret, err := generateMFASecret() + ... + user.MFAPendingSecret = secret + user.MFAPendingSecretExpiresAt = &expiresAt + ``` +- suggested fix: When `user.MFAEnabled` is true, require a current TOTP or recovery code in `MFASetup` (same check as `MFADisable`) before writing a new pending secret; or refuse setup while enabled and require disable first. +- verdict: CONFIRMED — mfaSetup (handler_auth.go:542-556) requires only a session, MFASetup (service_mfa.go:285-316) checks the password and never reads MFAEnabled, and MFAEnable:393-397 overwrites MFASecret and MFARecoveryCodes from the pending secret with no check that MFA was already on, so a session plus password replaces the authenticator without the proof-of-possession MFADisable:436 demands. +- issue: (pending cross-reference) + +### A04-004 pgx error/warn logs carry bound query arguments (session tokens, bcrypt hashes, approval tokens) +- category: security +- severity: medium +- location: internal/database/connection.go:399-405 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: pgx v5 `tracelog.TraceQueryEnd` logs `{"sql", "args", "err"}` at `LogLevelError` whenever a query fails. `sanitizeLogData` strips the `args` key only when `level == LogLevelDebug`. A transient failure of `INSERT INTO sessions (token, ...)`, `UPDATE users SET password_hash = $1`, or `UPDATE purchase_executions SET approval_token = $7` therefore writes the raw session token, bcrypt hash or approval token to CloudWatch (pgx truncates strings only past 64 bytes; a 64-hex token and a 60-char bcrypt hash are logged whole). `security_test.go:184-198` pins this as intended ("args kept at warn/error level"). +- evidence: + ```go + for k, v := range data { + if isSensitiveKey(k) { + continue + } + if k == "args" && level == tracelog.LogLevelDebug && !bindParams { + continue + } + safe[k] = v + } + ``` +- suggested fix: drop the `level == LogLevelDebug` condition so `args` is stripped at every level unless `DB_LOG_BIND_PARAMETERS=true`, and flip the two security_test cases to expect stripping. +- verdict: CONFIRMED — `sanitizeLogData` gates the `args` strip on `level == tracelog.LogLevelDebug` (internal/database/connection.go:399), `isSensitiveKey` matches only the literal keys password/secret/token and never the args payload (internal/database/connection.go:384-386), the tracer is installed on every pooled connection (internal/database/connection.go:226), `stdLogger.Log` prints `safeData` at warn and error (internal/database/connection.go:417-420), and internal/database/security_test.go:183-198 pins "args kept at warn level"/"args kept at error level". +- issue: (pending cross-reference) + +### A06-010 Unauthenticated /health echoes raw internal error strings +- category: security +- severity: medium +- location: internal/server/health.go:92 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `/health` is registered with no auth (internal/server/http.go:31) and always returns 200 with a JSON body. A migration failure puts `err.Error()` in the body verbatim — golang-migrate errors carry the failing migration number and the raw Postgres error including relation and column names; a DB health-check failure (health.go:123) and an auth-store ping failure (health.go:170) put pgx connection errors in the body, which name the host, port and user from the DSN. Any unauthenticated caller can poll the endpoint to map the schema and the database topology. +- evidence: + ```go + case err != nil: + return CheckResult{Status: "failed", Message: err.Error()} + ``` +- suggested fix: return a fixed status string in the public body and log the detail server-side, or gate the detailed body behind the scheduled-task/admin credential. +- verdict: CONFIRMED — `/health` is registered outside every auth wrapper (internal/server/http.go:31, contrast the scheduledauth wrap at http.go:39) and always writes 200 with the full JSON body (internal/server/health.go:58-70); all three checks put the raw error text in `Message` (health.go:92, 123, 170), and the golang-migrate and pgx errors reaching them carry migration numbers and DSN host/user detail respectively. +- issue: (pending cross-reference) + +### A06-012 SMTP send paths log raw recipient addresses while the sibling approval path redacts them +- category: security +- severity: medium +- location: internal/email/smtp_sender.go:182 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `SendToEmailWithCC` (lines 182/184) and `SendToEmailWithCCMultipart` (lines 151/153) interpolate `toEmail` verbatim, so every password-reset, invite, welcome, scheduled-purchase and executed-purchase send on a GCP (SendGrid) or Azure (ACS) deployment writes a customer email address into the log aggregator. The same file redacts in `sendMultipartWithUnsubscribe` (lines 591/593) and the whole SES path uses `redactEmail`, so this is an internal inconsistency, and it is the case the project memory `feedback_pii_in_logs` describes ("Emails: log only the domain part or mask the local part"). Redacted value withheld here; no live address was observed, only the format string. +- evidence: + ```go + if len(sanitizedCC) > 0 { + logging.Debugf("Sent email via SMTP to %s (cc %d): %s", toEmail, len(sanitizedCC), subject) + } else { + logging.Debugf("Sent email via SMTP to %s: %s", toEmail, subject) + } + ``` +- suggested fix: wrap both call sites in `redactEmail(toEmail)` as the approval path already does. +- verdict: CONFIRMED — raw `toEmail` is interpolated at internal/email/smtp_sender.go:151,153 and 182,184, while the sibling `sendMultipartWithUnsubscribe` uses `redactEmail(toEmail)` on the identical log lines (smtp_sender.go:591,593) and every SES log does the same (internal/email/sender.go:291,293,325,327); both SMTP methods are the live delivery path on GCP (SendGrid) and Azure (ACS) per internal/email/factory.go:140,196. +- issue: (pending cross-reference) + +### A06-013 Azure bicep/ARM deploy script is rendered from unescaped data while every sibling shell template is shell-escaped +- category: security +- severity: medium +- location: internal/iacfiles/templates/azure-wif-bicep-deploy.sh.tmpl:64 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the package doc for `iacfiles` states the renderer must escape user-controlled fields before `template.Execute`. `renderSingleFile` does so for CLI scripts (internal/api/handler_federation.go:274) and `writeCFNFiles` does so for the CFN deploy script (:640), but `writeAzureTemplateFiles` (:532) renders `azure-wif-bicep-deploy.sh.tmpl` with the raw `data` and writes the result into the zip with mode 0755. `.ContactEmail` (line 64) and `.CUDlyAPIURL` (line 83) land inside double-quoted bash. Addresses are validated only with `mail.ParseAddress`, which accepts a backtick in the local part (RFC 5322 atext), so a user whose account email contains one downloads a deploy script with a live command substitution that executes when they run it. Impact is bounded because `ContactEmail` is always the downloader's own session email (handler_federation.go:166), making this self-inflicted rather than cross-user, but the guard is simply absent on one of three shell-template paths. +- evidence: + ```sh + CONTACT_EMAIL="${CUDLY_CONTACT_EMAIL:-{{.ContactEmail}}}" + ... + -X POST "{{.CUDlyAPIURL}}/api/register" \ + ``` +- suggested fix: pass `shellEscapeData(data)` when rendering the deploy script in `writeAzureTemplateFiles`, exactly as `writeCFNFiles` does. +- verdict: CONFIRMED — `writeAzureTemplateFiles` renders the deploy script with the raw `data` (internal/api/handler_federation.go:532) and writes it with `addExecBytesToZip` (handler_federation.go:557), while the two sibling shell paths both escape first (`renderSingleFile` at handler_federation.go:273-274, `writeCFNFiles` at handler_federation.go:638-640); `.ContactEmail` and `.CUDlyAPIURL` land inside double-quoted bash and a heredoc (internal/iacfiles/templates/azure-wif-bicep-deploy.sh.tmpl:64,83), and I confirmed by running `mail.ParseAddress` that both `` a`id`b@example.com `` and `a${x}b@example.com` parse cleanly, so the only validator (internal/api/validation.go:134-150) admits them. +- issue: (pending cross-reference) + +### A10-005 The CLI still registers a `--yes` flag that skips the purchase confirmation +- category: security +- severity: medium +- location: cmd/main.go:124 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `cudly --purchase --yes` reaches `ConfirmPurchase(..., skipConfirmation=true)`, which returns true immediately (cmd/helpers.go:208-210) without printing the "About to purchase N instances" line or reading stdin. The only remaining human gate on a non-reversible multi-thousand-dollar RI purchase is removed by a single flag, and the flag is discoverable in `--help`. The project's own standing rule is that this confirmation must run on every invocation, dry-run included. There is a `TestDryRunFlagRemoved` guard against reintroducing `--dry-run` but no equivalent guard for `--yes`. +- evidence: + ```go + rootCmd.Flags().BoolVar(&toolCfg.SkipConfirmation, "yes", false, "Skip confirmation prompt for purchases (use with caution)") + ``` +- suggested fix: Remove the flag and the `SkipConfirmation` field, drop the parameter from `ConfirmPurchase`, and add a `rootCmd.Flags().Lookup("yes") != nil` regression test alongside `TestDryRunFlagRemoved`. +- verdict: CONFIRMED — the flag is bound at cmd/main.go:124 and reaches ConfirmPurchase, which returns true ahead of both the TTY check and the prompt (cmd/helpers.go:207-215). It can reach exactly two call sites, the main purchase path (cmd/multi_service.go:158) and the CSV per-region prompt (cmd/multi_service.go:701); no dry-run path is affected because both sit behind `!isDryRun` (:156) / the `else` of `if isDryRun` (:689-705), and dry runs never prompt at all. No test in cmd/ references the flag — TestDryRunFlagRemoved at cmd/effective_dry_run_test.go:38 has no --yes counterpart. +- issue: (pending cross-reference) + +### A10-019 `ListSecrets` errors are discarded and `arns[0]` is written to unvalidated in both configure paths +- category: security +- severity: medium +- location: cmd/configure_azure.go:148 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The returned `err` is consulted only inside `if err == nil && len(arns) > 0` and is then overwritten by the `json.Marshal` assignment, so an AccessDenied or throttling failure from `secretsmanager:ListSecrets` is never reported. The code proceeds with the bare secret name; if a secret of that exact name does not exist, `UpdateSecret` fails with a resource-not-found error that gives the operator no hint the listing was the real problem. Conversely, the Secrets Manager `name` filter is a prefix match, so with secrets `cudly-AzureCredentials` and `cudly-AzureCredentialsOld` both present, `arns[0]` is whichever AWS returns first and the credentials can be written into the wrong secret. Identical code at cmd/configure_gcp.go:119-125. +- evidence: + ```go + arns, err := store.ListSecrets(ctx, secretName) + // Use the ARN if found, otherwise use the name (will fail if secret doesn't exist) + secretID := secretName + if err == nil && len(arns) > 0 { + secretID = arns[0] + } + ``` +- suggested fix: Return the `ListSecrets` error to the caller, and when more than one ARN comes back, require an exact name match on the returned entry (or error naming the candidates) instead of taking index 0. +- verdict: CONFIRMED — cmd/configure_azure.go:148-154 and cmd/configure_gcp.go:119-125 consult err only inside the guard and then reassign it on the next `:=`, and ListSecrets hands the name to a FilterNameStringTypeName filter and returns every ARN with no exact-name check before the caller takes arns[0] (cmd/secrets_store.go:31-55). +- issue: (pending cross-reference) + +### A10-021 The minted GCP service-account private key is left on disk after being uploaded +- category: security +- severity: medium +- location: cmd/configure_gcp.go:722 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Step 5 writes a fresh, never-expiring GCP service-account private key to `~/cudly-gcp-key.json`, `runConfigureGCP` reads it and stores it in AWS Secrets Manager, then prints "GCP configuration complete!" and exits. The plaintext key stays in the operator's home directory indefinitely, where it is picked up by backup tools, Dropbox/iCloud sync and `find ~ -name '*.json'`. Nothing in the wizard deletes it, warns about it, or tells the operator to rotate it. The Azure path deliberately never persists its secret to disk, so the two flows disagree on the same hazard. +- evidence: + ```go + keyFile := filepath.Join(home, "cudly-gcp-key.json") + ... + fmt.Printf("Key file written to: %s\n", keyFile) + ``` +- suggested fix: After `storeGCPCredentials` succeeds on a wizard-minted key, delete the local file and say so, or print an explicit instruction to remove and rotate it. +- verdict: CONFIRMED — the key is written to ~/cudly-gcp-key.json at cmd/configure_gcp.go:722-745 and runConfigureGCP ends at printGCPConfigurationSuccess (cmd/configure_gcp.go:162-167, :256-262) with no unlink and no warning about the file. +- issue: (pending cross-reference) + +### A10-022 `clear-rate-limit` interpolates an env var into a `LIKE` pattern for an unguarded DELETE +- category: security +- severity: medium +- location: cmd/lambda/clear-rate-limit/main.go:41 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The domain is a query parameter (no SQL injection) but it lands inside a `LIKE` pattern, where `%` and `_` are metacharacters. With `RATE_LIMIT_DOMAIN=%` the pattern becomes `EMAIL#%@%#ENDPOINT#forgot_password` and the function deletes the forgot-password rate limits for every domain in the table, disabling brute-force protection on the password-reset endpoint tenant-wide. A single `_` in a legitimate domain broadens the match by one character position. The handler has no dry-run mode and no confirmation, unlike its sibling `cmd/cleanup-lambda`, which does have one. `defaultDomain` is additionally a hardcoded `leanercloud.com`, so a misconfigured deployment silently targets a specific tenant. +- evidence: + ```go + tag, err := db.Exec(ctx, + "DELETE FROM rate_limits WHERE id LIKE $1", + "EMAIL#%@"+domain+"#ENDPOINT#forgot_password") + ``` +- suggested fix: Validate the domain against a hostname regex and reject `%`/`_`, or escape them and add `ESCAPE '\'`; add a `dryRun` field to the event mirroring `CleanupEvent`. +- verdict: CONFIRMED — cmd/lambda/clear-rate-limit/main.go:39-42 concatenates getDomain() straight into the LIKE pattern with no hostname validation and no metacharacter escaping, the handler takes no event at all (:30, :65) so there is no dry-run switch, and the sibling does carry one (cmd/cleanup-lambda/main.go:15, :44). Triggering it requires control of RATE_LIMIT_DOMAIN at deploy time, which is the stated scenario. +- issue: (pending cross-reference) + +### A11-008 A failing /api/info offers the first-admin setup screen +- category: security +- severity: medium +- location: frontend/src/api/auth.ts:415 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `getPublicInfo` swallows every non-OK response and returns `{ version: '', admin_exists: false }`. `init()` reads that value and, seeing `!publicInfo.admin_exists`, renders the "No admin account exists yet. Set up the first admin to get started." modal (app.ts:49-54). So during any backend outage, 5xx, or misrouted deploy, an unauthenticated visitor is told the installation has no admin and is invited to claim it. The API key requirement is what stops the takeover, not this code path; the UI is nonetheless asserting a security-relevant fact it did not verify, and a user who reaches this screen on a healthy deployment has no signal that the check failed. +- evidence: + ```typescript + // api/auth.ts:412-419 + export async function getPublicInfo(): Promise { + const response = await fetch(`${API_BASE}/info`); + if (response.ok) { + return response.json() as Promise; + } + return { version: '', admin_exists: false }; + } + ``` +- suggested fix: throw on a non-OK response so `init()`'s existing catch falls through to the login modal, which is the safe default for an unknown bootstrap state. +- verdict: CONFIRMED — the non-OK branch returns `admin_exists: false` rather than throwing (frontend/src/api/auth.ts:414-418), and `init()` branches straight into `showAdminSetupModal` on that value (frontend/src/app.ts:47-53); its `catch` only covers a network-level failure, so any 4xx/5xx from `/info` renders the first-admin screen to an unauthenticated visitor. +- issue: (pending cross-reference) + +### A11-011 The permission matrix hides 13 of the 20 actions, including every -any verb +- category: security +- severity: medium +- location: frontend/src/users/permissionMatrix.ts:11 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ACTIONS` is a hand-written list of seven verbs. `permissions.ts` exports `ALL_ACTIONS`, a compile-time-exhaustive list of all twenty, and the group-edit form was already migrated onto it for exactly this reason (see the drift comment at groupModals.ts:166-188). The Permission Overview table therefore renders no row at all for `approve-any`, `retry-any`, `cancel-any`, `update-any`, `execute-any`, `sell-any`, `revoke-any` and the -own variants. An admin auditing which groups can approve or retry a purchase reads a matrix that shows those capabilities nowhere, and concludes a group holding `approve-any:purchases` grants nothing beyond `view`. The only test touching this file is an accessibility assertion (a11y.test.ts:75). +- evidence: + ```typescript + // users/permissionMatrix.ts:11 + const ACTIONS = ['view', 'create', 'update', 'delete', 'execute', 'approve', 'admin'] as const; + ``` +- suggested fix: import `ALL_ACTIONS` from `../permissions` and render from it, as `buildActionOptions` already does. +- verdict: CONFIRMED — the hand-written seven-verb list at frontend/src/users/permissionMatrix.ts:11 drives the row set (permissionMatrix.ts:46, matching on `p.action === action`), while `ACTION_EXHAUSTIVENESS_CHECK` enumerates twenty actions including every `-any` and `-own` verb (frontend/src/permissions.ts:100-122); the matrix is live, rendered from userActions.ts:98, so a group holding only `approve-any:purchases` shows dashes in every row, and a11y.test.ts:75 is indeed the only test that touches the file. +- issue: (pending cross-reference) + +### A13-008 `fargate_certificate_arn` never enables HTTPS, so the ALB serves the dashboard in cleartext +- category: security +- severity: medium +- location: terraform/environments/aws/compute.tf:177 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `enable_https` is derived only from `frontend_domain_names` and `subdomain_zone_name`; `var.fargate_certificate_arn` is passed as `certificate_arn` on the next line but is not part of the condition. An operator following `terraform/profiles/aws/fargate-dev.tfvars.example:21` ("Only needed if not using subdomain_zone_name") supplies a pre-issued ACM certificate and no zone, gets `enable_https = false`, and the module creates only the port-80 listener, whose `default_action` forwards to the target group rather than redirecting (terraform/modules/compute/aws/fargate/main.tf:482). The admin login form, session cookie and API traffic all travel unencrypted, and the supplied certificate is never attached to anything. The shipped fargate profile sets none of the three, so this is its default state. +- evidence: + ```hcl + enable_https = length(var.frontend_domain_names) > 0 && var.subdomain_zone_name != "" + certificate_arn = ( + length(aws_acm_certificate.frontend) > 0 + ? aws_acm_certificate_validation.frontend[0].certificate_arn + : var.fargate_certificate_arn + ) + ``` +- suggested fix: Make the condition `(length(var.frontend_domain_names) > 0 && var.subdomain_zone_name != "") || var.fargate_certificate_arn != ""`, and add a precondition that refuses a public Fargate ALB with neither. +- verdict: CONFIRMED — terraform/environments/aws/compute.tf:177 derives `enable_https` from `frontend_domain_names` and `subdomain_zone_name` only, and `var.fargate_certificate_arn` (variables.tf:382-386, default `""`) appears solely as the fallback on the `certificate_arn` line below it. With `enable_https = false` the module creates no 443 listener (fargate/main.tf:498-499 is `count = var.enable_https ? 1 : 0`) and the port-80 listener's default action forwards to the target group rather than redirecting (fargate/main.tf:481-493). terraform/profiles/aws/fargate-dev.tfvars.example:21 tells the operator the certificate is "Only needed if not using subdomain_zone_name", which is exactly the combination that yields cleartext. +- issue: (pending cross-reference) + +### A13-010 The ACR AcrPull role assignment is dead because the Container App authenticates with admin credentials +- category: security +- severity: medium +- location: terraform/environments/azure/registry.tf:10 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `admin_enabled = true` turns on a static registry username/password, and compute.tf:102-103 feeds those to the Container App's `registry` block (terraform/modules/compute/azure/container-apps/main.tf:92-94), so image pulls authenticate as the admin account. The `azurerm_role_assignment.acr_pull` created directly below therefore grants a managed identity that is never used for pulls, and the long-lived admin credential remains live in Key Vault, in state and in the container app secret. Removing the role assignment would change nothing observable, which is the definition of a guard on an unreachable state. +- evidence: + ```hcl + admin_enabled = true # Enables username/password login for docker push + ... + resource "azurerm_role_assignment" "acr_pull" { + scope = azurerm_container_registry.main.id + role_definition_name = "AcrPull" + principal_id = module.compute_container_apps[0].managed_identity_principal_id + ``` +- suggested fix: Drop `registry_username`/`registry_password` and set the Container App registry block to use the managed identity, then set `admin_enabled = false`; add `depends_on = [azurerm_role_assignment.acr_pull]` on the container app so the first pull does not race RBAC propagation. +- verdict: CONFIRMED — terraform/environments/azure/compute.tf:101-103 passes `registry_server`/`registry_username`/`registry_password` from `azurerm_container_registry.main` (`admin_enabled = true`, registry.tf:10), and the module's `registry` block authenticates with `username` plus `password_secret_name = "registry-password"` and never sets `identity` (terraform/modules/compute/azure/container-apps/main.tf:88-96). The `AcrPull` assignment to `managed_identity_principal_id` (registry.tf:15-21) therefore governs no pull; deleting it changes nothing observable. +- issue: (pending cross-reference) + +### A13-013 `.trivyignore` AVD-AWS-0013 states a justification the code contradicts +- category: security +- severity: medium +- location: .trivyignore:17 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The suppression claims TLS 1.0 "was tightened in a prior PR" and that the finding only trips because the value is "inherited rather than set explicitly". `terraform/modules/frontend/aws/main.tf:113` sets `minimum_protocol_version` explicitly to `"TLSv1"` whenever `acm_certificate_arn` is null, which is the default for every environment. The finding is live, not an artifact, so the suppression hides a real TLS 1.0 viewer policy rather than a false positive. It is not currently exploitable only because `enable_cdn = false` in every tfvars, which is a different reason from the one written down. +- evidence: + ```hcl + minimum_protocol_version = var.acm_certificate_arn != null ? "TLSv1.2_2021" : "TLSv1" + ``` +- suggested fix: Either drop the no-certificate branch (a CloudFront default certificate forces TLSv1 regardless, so the distribution should require a real certificate) or rewrite the suppression to state the actual reason, which is that the distribution is never created. +- verdict: CONFIRMED — .trivyignore:17-22 says the value "still trips when the value is inherited rather than set explicitly", but terraform/modules/frontend/aws/main.tf:113 sets `minimum_protocol_version` explicitly, to `"TLSv1"` on the `acm_certificate_arn == null` branch. The mitigating fact the suppression does not mention also checks out: the distribution is behind `count = var.enable_cdn ? 1 : 0` (terraform/environments/aws/frontend.tf:9) and every shipped tfvars sets `enable_cdn = false` (github-{dev,staging,prod}.tfvars). So the written justification is false and the real one is unwritten. +- issue: (pending cross-reference) + +### A13-015 `setup.sh` prints the Azure client secret to stdout while the WIF branch deliberately uses stderr +- category: security +- severity: medium +- location: arm/CUDly-CrossSubscription/setup.sh:135 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The `wif` branch writes the private key to stderr with an explicit warning that stdout redirects would capture it (lines 86-94). The default `client_secret` branch prints a two-year Azure AD application secret to stdout with no warning at all, so `./setup.sh > onboarding.txt` or any wrapper that captures stdout persists the credential to disk, and a CI invocation captures it in the run log. The two branches emit the same class of secret through opposite channels. +- evidence: + ```bash + if [[ "$MODE" == "client_secret" ]]; then + echo " client_secret : ${CLIENT_SECRET}" + echo "" + echo " Save the client_secret now — it will not be shown again." + ``` +- suggested fix: Route the client-secret block to stderr with the same warning the WIF branch carries, so both credential modes behave identically under redirection. +- verdict: CONFIRMED — arm/CUDly-CrossSubscription/setup.sh:86-88 carries the warning "Do NOT run this script in CI/CD or any environment that captures stdout. Both blocks are written to stderr so stdout redirects do not capture them", and lines 91-95 duly `>&2` the key and certificate. The `client_secret` branch mints a two-year secret at lines 107-112 and the summary block prints it with a bare `echo` at line 139, inside a `[[ "$MODE" == "client_secret" ]]` guard, with no redirection and no warning. +- issue: (pending cross-reference) + +### A13b-003 Azure Bicep role assignment takes the role definition ID as an unconstrained parameter +- category: security +- severity: medium +- location: iac/federation/azure-target/bicep/azure-wif.bicep:24 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `roleDefinitionId` is a plain string parameter with no `@allowed` list, and the documented deployment flow (lines 12-16 of the same file) tells the customer to run `az deployment sub create --parameters @target-azure-wif-bicep-params.json`. Anyone who supplies or edits that parameter file — a support engineer, a doc page, a tampered download — can set it to the Owner role GUID `8e3af657-a8ff-443c-a75c-2fe8c4bcb635` or Contributor. `subscriptionResourceId` resolves whatever GUID it is handed at subscription scope, so the deployment grants CUDly's service principal full control of the customer's subscription while the template, its description text and its assignment description all still say "Reservation Purchaser". Nothing in the template compares the value against the role it claims to assign. Same defect in the generated ARM at `iac/federation/azure-target/bicep/azure-wif.arm.json:20`. +- evidence: + ```bicep + param roleDefinitionId string = 'f7b75c60-3036-4b75-91c3-6b41c27c1689' + ... + roleDefinitionId: subscriptionResourceId('Microsoft.Authorization/roleDefinitions', roleDefinitionId) + description: 'CUDly Reservation Purchaser — assigned via CUDly federation setup.' + ``` +- suggested fix: Drop the parameter and inline the role definition the template intends to assign, or constrain it with `@allowed([...])` naming only the acceptable GUIDs. A customer-deployed template should never let its caller choose which privilege it grants. +- verdict: CONFIRMED — `iac/federation/azure-target/bicep/azure-wif.bicep:20-37` has no `@allowed` on `roleDefinitionId` and no comparison against the role it names in its own `description` at line 35, and `subscriptionResourceId` at line 34 resolves any built-in GUID at subscription scope; the generated ARM repeats it verbatim at `iac/federation/azure-target/bicep/azure-wif.arm.json:18-24,41`, and the documented flow at lines 12-16 passes the value through a params file the customer does not author. +- issue: (pending cross-reference) + +### A13b-008 CloudFormation cross-account role trusts the whole source account with no way to pin the execution role +- category: security +- severity: medium +- location: iac/federation/aws-cross-account/cloudformation/template.yaml:152 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The CloudFormation template hardcodes `:root` as the trusted principal, so every IAM principal in CUDly's source account — a CI role, a developer user, a compromised low-privilege function — can assume the customer's role and place irreversible multi-year commitment purchases, subject only to knowing the external ID. The Terraform module for the identical access offers `cudly_execution_role_arn` (`iac/federation/aws-cross-account/terraform/variables.tf:24-33`) and pins the trust to that single principal when set. Customers who deploy the CloudFormation variant, which is the copy-paste path, silently get the wider trust with no parameter to narrow it. +- evidence: + ```yaml + AssumeRolePolicyDocument: + Version: "2012-10-17" + Statement: + - Effect: Allow + Principal: + AWS: !Sub "arn:aws:iam::${SourceAccountID}:root" + Action: sts:AssumeRole + ``` +- suggested fix: Add an optional `CUDlyExecutionRoleARN` parameter and an `Fn::If` that selects it over `:root`, mirroring the Terraform local. +- verdict: CONFIRMED — `iac/federation/aws-cross-account/cloudformation/template.yaml:150-152` hardcodes `:root` and the template's Parameters block (lines 7-49) offers no execution-role input at all, while `iac/federation/aws-cross-account/terraform/main.tf:23-27` selects `var.cudly_execution_role_arn` over the account root; the CloudFormation side is the wrong one, though the external-ID condition at lines 153-155 still narrows the trust to a caller holding that secret. +- issue: (pending cross-reference) + +### A13b-009 External ID is an operator-chosen 8-character string in both CloudFormation templates while Terraform generates a UUID +- category: security +- severity: medium +- location: iac/federation/aws-cross-account/cloudformation/template.yaml:13 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `MinLength: 8` with no `AllowedPattern` accepts `password`, `cudly123` or the customer's own account number as the external ID, and `cloudformation/stacks/CUDly-CrossAccount/template.yaml:31` has the same shape. The Terraform module for the same access auto-generates a `random_uuid` when the value is empty (`iac/federation/aws-cross-account/terraform/main.tf:19,27`), so the two paths produce secrets of wildly different strength. The external ID is the sole control that stops a confused-deputy registration: the role name defaults to a published constant (`CUDly-CrossAccount` here, `CUDly` in the stacks template) and the account ID is often discoverable, so an attacker who registers the victim's role ARN with a guessed external ID gets CUDly to assume into the victim account on their behalf. A short, human-chosen value makes that guess practical. +- evidence: + ```yaml + ExternalID: + Type: String + Description: > + Random string used as the sts:ExternalId condition to prevent confused-deputy + attacks. Generate a UUID and share it securely with the source account. + MinLength: 8 + NoEcho: true + ``` +- suggested fix: Raise `MinLength` to 32 and add an `AllowedPattern` requiring a hex or UUID shape, in both CloudFormation templates. Better still, have CUDly issue the external ID at registration time rather than accepting one the deploying party chooses. +- verdict: CONFIRMED — `iac/federation/aws-cross-account/cloudformation/template.yaml:13-19` and `cloudformation/stacks/CUDly-CrossAccount/template.yaml:30-37` both carry `MinLength: 8` with no `AllowedPattern`, against `iac/federation/aws-cross-account/terraform/main.tf:19-27` which substitutes a `random_uuid` when the value is empty, so the two paths for the same trust boundary produce different secret strength; the confused-deputy end of the scenario is weaker than stated, since `internal/api/handler_registrations.go:76-78` lands every registration as `pending` and `approveRegistration` at line 252 requires a CUDly admin before the ARN is ever assumed. +- issue: (pending cross-reference) + +### A13b-010 Registration endpoint URL is unvalidated, so `http://` sends the external ID and GCP credential config in cleartext +- category: security +- severity: medium +- location: iac/federation/aws-cross-account/terraform/registration.tf:21 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `var.cudly_api_url` has no validation in any of the five modules and is interpolated straight into `data.http`. A customer handed a mistyped or downgraded `http://` URL POSTs `aws_external_id` (`registration.tf:15`, the confused-deputy secret) over cleartext; anyone on the path recovers it and, with the predictable role ARN from A13b-009, can drive purchases in that account. The GCP variant is worse in payload: `iac/federation/gcp-target/terraform/registration.tf:45` sends the full external-account credential config JSON. The CloudFormation Lambda has the same gap at `iac/federation/aws-cross-account/cloudformation/template.yaml:206`, where `props["CUDlyAPIURL"]` is concatenated with no scheme check. +- evidence: + ```hcl + data "http" "cudly_registration" { + count = local.do_register ? 1 : 0 + url = "${var.cudly_api_url}/api/register" + method = "POST" + ``` +- suggested fix: Add a `validation` block requiring `startswith(var.cudly_api_url, "https://")` to all five modules' `cudly_api_url`, and an `AllowedPattern: "^$|^https://.+"` to the CloudFormation `CUDlyAPIURL` parameter. +- verdict: CONFIRMED — the `cudly_api_url` variable block is byte-identical and validation-free in all five modules (`iac/federation/{aws-target,aws-cross-account,azure-target,gcp-sa-impersonation,gcp-target}/terraform/variables.tf`) and is interpolated directly at each `registration.tf` `url` line; the AWS cross-account payload carries `aws_external_id` (`iac/federation/aws-cross-account/terraform/registration.tf:15`), the GCP WIF payload carries the full `credential_payload` config JSON (`iac/federation/gcp-target/terraform/registration.tf:45`), and the CloudFormation Lambda concatenates the parameter with no scheme check (`iac/federation/aws-cross-account/cloudformation/template.yaml:206`) behind a `CUDlyAPIURL` parameter that has no `AllowedPattern` (lines 26-29). +- issue: (pending cross-reference) + +### A13b-012 Azure federated identity credential accepts any issuer URL with no validation +- category: security +- severity: medium +- location: iac/federation/azure-target/terraform/variables.tf:19 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `cudly_issuer_url` is a bare required string with no validation block, and it becomes the `issuer` of the federated identity credential at `main.tf:57`. Azure AD then fetches JWKS from whatever host it names and mints tokens for CUDly's app registration for any JWT that host signs with `sub` equal to `cudly_federated_subject` (default `cudly-controller`, a constant shared by every CUDly customer) and `aud` `api://AzureADTokenExchange`. A customer given a wrong or attacker-controlled issuer URL — a typosquat in a docs page, a stale value in a params file — hands whoever controls that host the ability to obtain tokens for the app and, via the subscription role assignment at `main.tf:78`, to purchase reservations in the customer's subscription. The AWS WIF sibling validates its issuer URL hard (`iac/federation/aws-target/terraform/variables.tf:11-14`) and rejects `$` and `*` in both audience and subject; the Azure module applies none of that to any of its three federation inputs. +- evidence: + ```hcl + variable "cudly_issuer_url" { + description = "CUDly OIDC issuer URL (e.g. https://cudly.example.com/oidc). Azure AD fetches JWKS from this issuer to verify client assertion JWTs." + type = string + } + ``` +- suggested fix: Add a validation requiring `https://` and no whitespace on `cudly_issuer_url`, and a non-empty validation on `cudly_federated_subject` so it cannot be blanked. Consider making the subject per-subscription rather than one global constant, so a single leaked assertion does not replay across every customer tenant. +- verdict: CONFIRMED — `iac/federation/azure-target/terraform/variables.tf:19-33` declares `cudly_issuer_url`, `cudly_federated_subject` and `cudly_federated_audience` with no `validation` block on any of the three, and all three feed the federated credential unchecked at `iac/federation/azure-target/terraform/main.tf:56-58`; the AWS sibling guards the equivalent inputs hard (`iac/federation/aws-target/terraform/variables.tf:11-15` on the issuer, `:36-39` on the audience, `:59-62` and the `$`/`*` rejection on the subject), so this is a missing guard rather than a difference in threat model, and the default subject `cudly-controller` is one constant across every customer tenant. +- issue: (pending cross-reference) + +### A13c-001 Lambda secrets policy renders `Resource = "*"` when the DB secret ARN is empty, while its three siblings guard against exactly that +- category: security +- severity: medium +- location: terraform/modules/compute/aws/lambda/main.tf:245 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `compact()` drops empty strings, so each of the admin / credential-encryption / + scheduled-task entries is written as `X != "" ? "${X}*" : ""` and disappears when unset. The + database entry on line 247 has no such guard: with `database_password_secret_arn = ""` the + element renders as the literal `"*"`, which `compact()` keeps, and the Lambda execution role + gains `secretsmanager:GetSecretValue` on **every secret in the account**. The Fargate twin + writes `"${arn}-*"` (fails closed to the unmatchable `-*`), so the same input produces opposite + polarity on the two compute paths. +- evidence: + ```hcl + Resource = compact([ + var.database_password_secret_arn, + "${var.database_password_secret_arn}*", + var.admin_password_secret_arn, + var.admin_password_secret_arn != "" ? "${var.admin_password_secret_arn}*" : "", + var.credential_encryption_key_secret_arn, + var.credential_encryption_key_secret_arn != "" ? "${var.credential_encryption_key_secret_arn}*" : "", + var.scheduled_task_secret_arn, + var.scheduled_task_secret_arn != "" ? "${var.scheduled_task_secret_arn}*" : "", + ]) + ``` +- suggested fix: guard the database entry like its siblings, or add a `validation` block on + `database_password_secret_arn` requiring a non-empty `arn:aws:secretsmanager:` prefix. +- verdict: PLAUSIBLE — the asymmetry is real and the variable has no default and no validation + (lambda/variables.tf:72-75), so `""` renders the literal `"*"` that `compact()` keeps, and the + Fargate twin's `-*` polarity is confirmed (fargate/main.tf:109); but the scenario needs a caller + that passes the empty string, and the sole in-tree caller passes a real ARN + (`module.database.password_secret_arn` at environments/aws/compute.tf:47, sourced from + `module.secrets.database_password_secret_arn` via database.tf:15 and database/aws/main.tf:64). +- issue: (pending cross-reference) + +### A13c-009 Fargate's two original EventBridge roles lack the `ecs:cluster` and `iam:PassedToService` conditions its two newer ones carry +- category: security +- severity: medium +- location: terraform/modules/compute/aws/fargate/main.tf:896 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: four EventBridge invoker roles exist in this file with identical purpose. + `eventbridge_ladder_run` (:1103) and `eventbridge_fire_scheduled_purchases` (:1213) constrain + `ecs:RunTask` with `ArnEquals ecs:cluster` and `iam:PassRole` with `StringEquals + iam:PassedToService = ecs-tasks.amazonaws.com`. The recommendations role (:896) and the + RI-exchange role (:997) have neither, so each can pass the task role — which holds the SP/RI + purchase grants — to any AWS service that accepts a passed role, and can run the task + definition in any cluster in the account. Same input, same intent, two of four guarded. +- evidence: + ```hcl + { + Effect = "Allow" + Action = ["iam:PassRole"] + Resource = [ + aws_iam_role.task_execution.arn, + aws_iam_role.task.arn + ] + } + ``` +- suggested fix: copy the two `Condition` blocks from `eventbridge_ladder_run` onto the + recommendations and RI-exchange policies; better, collapse all four into one `for_each` role so + the guards cannot diverge again. +- verdict: CONFIRMED — fargate/main.tf:896-923 and :997-1024 each carry a bare `ecs:RunTask` on the + task-definition ARN and a bare `iam:PassRole` on the execution and task roles, with no `Condition` + block; the two newer roles do constrain both (`ecs:cluster` at :1120 and :1230, + `iam:PassedToService` at :1135 and :1245). Reach is bounded by the roles' `events.amazonaws.com` + trust policy, so exploitation needs an actor able to create EventBridge targets. +- issue: (pending cross-reference) + +### A14-009 The git-secrets AWS secret-key pattern is inverted and can never match a key +- category: security +- severity: high +- location: scripts/setup-git-secrets.sh:52 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: An AWS secret access key is 40 characters drawn from `[A-Za-z0-9/+=]`. The registered pattern is the negated class, so it matches 41 consecutive characters that are **not** in that set. A real secret key committed to the repo is not matched by this pattern at all; the only things that can match are runs of punctuation such as the `━━━━` separator lines this repo's own scripts print. The pattern is registered by `make setup-git-secrets` and becomes part of every developer's local pre-commit gate, where its comment claims AWS-secret coverage it does not provide. +- evidence: + ```bash + git secrets --add '[^A-Za-z0-9/+=]{40}[^A-Za-z0-9/+=]' # AWS Secret Access Key + ``` +- suggested fix: Replace with a positive class anchored on non-key boundaries, e.g. `(^|[^A-Za-z0-9/+=])[A-Za-z0-9/+=]{40}([^A-Za-z0-9/+=]|$)`, and add a fixture proving it fires on a synthetic 40-char key. +- verdict: CONFIRMED — scripts/setup-git-secrets.sh:52 registers the negated class, and tested with `/usr/bin/grep -E` it matched only a run of U+2501 box-drawing characters, never a synthetic 40-character base64 key. +- severity-adjusted: medium — `git secrets --register-aws` (scripts/setup-git-secrets.sh:43) plus the pattern on line 53 still block an assignment-shaped key (verified: `secret_key = "<40 chars>"` exits 1), so only a bare unassigned key slips through. +- issue: (pending cross-reference) + +### A14-024 The AWS IAM parity extractor cannot see a wildcard action +- category: security +- severity: medium +- location: scripts/check-aws-iam-parity.sh:79 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The extraction regex requires an uppercase letter after the colon, so `ec2:*` and `Action: "*"` produce no token. If the CloudFormation stack were changed to `"ec2:*"` while the Terraform module keeps its explicit list, both sides extract the same set of remaining actions, `compare_pair` finds no diff and the guard reports "OK: AWS IAM action lists are in parity" for two policies that are not equivalent. The Azure sibling covers this class explicitly (`scripts/testdata/role-parity/dataactions-wildcard-arm.json`, cases 13-15 of `test-azure-role-parity.sh`); the AWS suite has no wildcard fixture. No live wildcard exists in the compared namespaces today, so this is a latent gap rather than an active drift. +- evidence: + ```bash + actions=$(grep -oE "(^|[^A-Za-z])(${ACTION_PREFIXES}):[A-Z][A-Za-z]+" "$file" \ + | sed 's/^[^a-zA-Z]//' \ + | sort -u) + ``` +- suggested fix: Extend the pattern to capture `:\*` and a bare `"*"` action, and add a fixture to `scripts/test-aws-iam-parity.sh` asserting that a wildcard on one side fails parity. +- verdict: CONFIRMED — the regex at scripts/check-aws-iam-parity.sh:79 requires `[A-Z]` after the colon, and on synthetic input `ec2:*`, `rds:*` and `Action: "*"` produced zero tokens while `ec2:DescribeInstances` matched, so a wildcard on one side leaves both lists identical and compare_pair (line 160) reports parity; the Azure sibling fixture scripts/testdata/role-parity/dataactions-wildcard-arm.json is exercised at scripts/test-azure-role-parity.sh:240 while the AWS suite has no wildcard case. +- issue: (pending cross-reference) + +### A14-025 `--source` is never validated, so a typo renders an AWS trust policy with an empty issuer +- category: security +- severity: medium +- location: scripts/generate-federation-iac.go:518 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `populateData` allowlists `--target` against `validTargets` but passes `source` through unchecked. With `--target aws --source Azure` (or any typo), `subjectClaimModeFor` returns `subjectClaimRequired` so a subject claim is accepted, `awsOIDCIssuer` falls through its `default:` to `""`, and `singleFileTmpl` selects `aws-wif.tfvars.tmpl`. The generated file carries `oidc_issuer_url = ""` alongside a populated `oidc_subject_claim`, so the operator applies a WIF trust policy whose issuer was silently dropped. `--target aws --source AWS` is worse: it produces a WIF artifact where the cross-account external-ID artifact was intended. `TestGenerator_InvalidTargetReportedAsTarget` covers the target axis; there is no equivalent source case. +- evidence: + ```go + if !validTargets[target] { + return fmt.Errorf("--target must be aws, azure, or gcp (got %q)", target) + } + ``` + ```go + func awsOIDCIssuer(source, tenantID string) string { + switch source { + case "azure": ... + case "gcp": ... + default: + return "" + ``` +- suggested fix: Add a `validSources` allowlist checked alongside `validTargets`, and make `awsOIDCIssuer`'s default arm an error rather than an empty string. +- verdict: CONFIRMED — populateData validates only `--target` (scripts/generate-federation-iac.go:518), and running `--target aws --source Azure` rendered `oidc_issuer_url = ""` beside a populated `oidc_subject_claim`, because subjectClaimModeFor:427 keys on `source != "aws"`, awsOIDCIssuer:154 falls through to `return ""`, and singleFileTmpl:314 selects aws-wif.tfvars.tmpl. +- issue: (pending cross-reference) + +### A14-027 Only the subject claim is validated; six sibling fields reach the same Bash/JSON/HCL sinks raw +- category: security +- severity: medium +- location: scripts/generate-federation-iac.go:596 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The file's own comment (355-376) explains that `--oidc-subject-claim` is allowlisted because it is interpolated verbatim into three grammars. `AccountName`, `AccountExternalID`, `TenantID`, `ProjectID`, `ServiceAccountEmail`, `CUDlyAPIURL` and `ContactEmail` reach the same templates with no validation and no escaping: `aws-wif.tfvars.tmpl` renders `account_name = "{{.AccountName}}"` (an unescaped HCL string) and `aws-cfn-deploy.sh.tmpl:78` renders `-X POST "{{.CUDlyAPIURL}}/api/register"` inside a Bash script, where a value containing `$(…)` executes when the operator runs the generated deploy script. The server path escapes all of them: `internal/api/handler_federation.go:285-299` calls `shellEscape` on every field. Same templates, two callers, one of which escapes nothing. +- evidence: + ```go + data := iacData{ + AccountName: *accountName, + AccountExternalID: *accountID, + AccountSlug: slug, + Source: *source, + ContactEmail: *contactEmail, + CUDlyAPIURL: *cudlyAPIURL, + } + ``` +- suggested fix: Apply the same charset validation (or the server's `shellEscape`) to every field written into `iacData`, not only `OIDCSubjectClaim`. +- verdict: CONFIRMED — only OIDCSubjectClaim is validated (validateOIDCSubjectClaim at scripts/generate-federation-iac.go:441); `--account-name 'Acme"\nevil_var = "pwned'` rendered a second HCL attribute into the tfvars output and `--cudly-api-url 'https://x$(id)'` passed through verbatim, against internal/api/handler_federation.go:285-299 which shellEscapes every field. +- severity-adjusted: medium — unchanged: the Bash sink at aws-cfn-deploy.sh.tmpl:78 is unreachable today because bundle mode dies at A14-026, but the HCL sink is reachable and injectable. +- issue: (pending cross-reference) + +### A14-036 `detect-private-key` skips every Go test file, and `.gitleaksignore` uses a form gitleaks does not honour +- category: security +- severity: medium +- location: .pre-commit-config.yaml:61 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A PEM private key committed in any `*_test.go` file passes the pre-commit gate, and test fixtures are exactly where real keys are most often pasted. The neighbouring comment explains that the `internal/credentials/resolver.go` exclusion was removed to tighten the gate, but the far broader `_test\.go` exclusion was kept without justification. Separately, `.gitleaksignore` contains a bare file path; gitleaks expects finding fingerprints (`commit:file:rule:line`), so the entry matches nothing — and no workflow or hook in this shard invokes gitleaks at all, so the file is inert either way while reading as an active suppression. +- evidence: + ```yaml + - id: detect-private-key + name: Detect private keys + exclude: '(_test\.go|frontend/src/index\.html)$' + ``` + ``` + internal/secrets/aws_resolver_httptest_test.go + ``` +- suggested fix: Narrow the exclusion to the specific fixture files that need it; either wire gitleaks into CI and convert `.gitleaksignore` to real fingerprints, or delete the file. +- verdict: CONFIRMED — .pre-commit-config.yaml:61 excludes `_test\.go$` from detect-private-key with no accompanying justification, and a repo-wide search for gitleaks returns only .github/runbooks/credential-compromise.md:128 ("Enable gitleaks pre-commit hook"), so no tool ever consumes `.gitleaksignore`'s bare path entry. +- issue: (pending cross-reference) + +### A01-017 AuthPublic revoke/approve/cancel routes disclose execution existence and status before any authentication +- category: security +- severity: low +- location: internal/api/handler_purchases.go:1258 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: An unauthenticated caller POSTs `/api/purchases/revoke/{uuid}` with no token and no session. `revokeViaEmailToken` loads the row and returns 404 (no such execution), 409 "still pending", or 409 "cannot be revoked (status=)" before `tryRevokeViaSession`/`authorizeApprovalAction` run, so the response is a per-UUID existence and status oracle. `cancelPurchase` (1029-1035) and `loadApproveExecution` (524-535, including orphan details) return the same pre-auth 404/409s. The scoping helpers elsewhere deliberately return 404 to avoid exactly this enumeration signal (scoping.go:16-21). +- evidence: + ```go + execution, err := h.config.GetExecutionByID(ctx, execID) + if err != nil { + return nil, fmt.Errorf("failed to get execution: %w", err) + } + if execution == nil { + return nil, NewClientError(404, "execution not found") + } + if statusErr := checkRevokableStatus(execution); statusErr != nil { + return nil, statusErr + } + ``` +- suggested fix: Resolve the principal (session or token) first and collapse pre-auth failures into a single 401/404, running the status checks only after the caller is authorized. +- verdict: CONFIRMED — internal/api/router.go:163-172 registers approve/cancel/revoke as AuthPublic; revokeViaEmailToken (handler_purchases.go:1248-1260) returns 404 or a status-bearing 409 before tryRevokeViaSession/authorizeApprovalAction run, cancelPurchase:1029-1035 404s pre-auth, and loadApproveExecution:523-537 returns 404/orphan-409 before either auth branch; requireAccountAccess (scoping.go:16-21) documents the enumeration rationale these public routes do not follow. +- issue: (pending cross-reference) + +### A02-015 PATCH is registered as a mutating route but excluded from CSRF validation and the CORS method list +- category: security +- severity: low +- location: internal/api/middleware.go:274 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: router.go:150 serves `PATCH /api/plans/{id}` (a plan mutation). `requiresCSRFValidation` returns false for any method other than POST/PUT/DELETE, so the CSRF header is never checked on PATCH; `buildResponseHeaders` (handler.go:720) advertises `GET, POST, PUT, DELETE, OPTIONS`, so a cross-origin deployment's preflight for PATCH fails. Exploitability is low because the session credential is a bearer header rather than a cookie, but the middleware's own contract ("state-changing requests need CSRF protection") is not met for this verb and nothing tests it. +- evidence: + ```go + if method != "POST" && method != "PUT" && method != "DELETE" { + return false + } + ``` +- suggested fix: Add PATCH to the CSRF method set and to `Access-Control-Allow-Methods`, with a middleware test for `PATCH /api/plans/x`. +- verdict: CONFIRMED — router.go:150 registers PATCH /api/plans/ as AuthUser, requiresCSRFValidation (middleware.go:274-276) returns false for any method outside POST/PUT/DELETE, the only handler-level validateCSRF calls are in handler_purchases.go:672,1135,1335 and handler_ri_exchange.go:2183 (none in handler_plans.go), and buildResponseHeaders (handler.go:720) advertises only GET/POST/PUT/DELETE/OPTIONS; low severity stands because the session credential is a bearer header. +- issue: (pending cross-reference) + +### A02-016 The public prefix "/api/info" also matches the authenticated /api/info/deployment route +- category: security +- severity: low +- location: internal/api/middleware.go:22 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `isPublicEndpoint("/api/info/deployment")` is true, so `validateSecurityContext` skips authentication and API-key usage booking for a route that returns the Secrets Manager URL and the host AWS account id (router.go:363). Only the router-level `requireAuth` in `Route` keeps it protected; a future change that trusts the middleware (as the comment at router.go:16-25 anticipates) would expose it, and `TestHandler_isPublicEndpoint` (middleware_test.go:15-44) has no case for the path. The same file explicitly exact-matches `/version` and `/api/register` to avoid exactly this overlap. +- evidence: + ```go + publicPrefixEndpoints := []string{ + "/health", // Root health endpoint (no /api prefix) + "/api/health", // API health endpoint + "/api/info", + ``` +- suggested fix: Move `/api/info` to the exact-match switch and add `{"/api/info/deployment", false}` to the test table. +- verdict: CONFIRMED — "/api/info" sits in the prefix list (middleware.go:22) so isPublicEndpoint("/api/info/deployment") is true and validateSecurityContext (handler.go:762) skips authenticatePrincipal and the API-key usage booking; the route is protected only by its AuthUser level (router.go:363) enforced in Router.Route, which the router.go:16-25 comment describes as defense-in-depth rather than the primary gate. +- issue: (pending cross-reference) + +### A02-020 Registration list and detail expose reference_token, which reject deliberately strips +- category: security +- severity: low +- location: internal/api/handler_registrations.go:156 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `config.AccountRegistration.ReferenceToken` is tagged `json:"reference_token"` (internal/config/types.go:1078). `listRegistrations` and `getRegistration` return the struct whole, so every admin (and any log/export of the response) sees the registrant's private status token, while `rejectRegistration` (lines 373-381) builds a filtered map precisely so the token is "not exposed to admin". The token is the only credential for `GET /api/register/{token}`, so the two endpoints disagree about the same secret. +- evidence: + ```go + regs, err := h.config.ListAccountRegistrations(ctx, filter) + ... + return regs, nil + ``` +- suggested fix: Tag `ReferenceToken` with `json:"-"` and return it only from `submitRegistration`'s explicit map. +- verdict: CONFIRMED — registrationColumns (internal/config/store_postgres_registrations.go:225) selects reference_token for the list and get queries, the field is tagged json:"reference_token" (internal/config/types.go:1078), and listRegistrations/getRegistration (handler_registrations.go:153-175) return the structs whole, while rejectRegistration (handler_registrations.go:373-381) builds a filtered map explicitly to hide it. +- issue: (pending cross-reference) + +### A03-015 Account scope matches on display name, so two accounts sharing a name share scope +- category: security +- severity: low +- location: internal/auth/types.go:207 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `MatchesAccount` and `AccountScope.Allows` (account_scope.go:112) treat a scope entry equal to `accountName` as a match. Display names are free text and not unique. A group scoped to `["prod"]` (by name) gains access to every account later registered with the display name "prod", including one belonging to another team, with no change to the group. Renaming an account also silently changes who can see it. +- evidence: + ```go + for _, a := range allowed { + if a == accountID { + return true + } + if accountName != "" && a == accountName { + return true + } + } + ``` +- suggested fix: Match on account ID only; migrate existing name-based entries to IDs once and drop the name parameter. +- verdict: CONFIRMED — MatchesAccount (types.go:203-209) and AccountScope.Allows (account_scope.go:108-114) both match on accountName, production callers pass real display names (handler_ri_exchange.go:431, :809; handler_ladder.go:60; handler_analytics.go:245; handler_history.go:1015), and the group write path compares allowed_accounts entries only against the actor's own scope (group_ceiling.go:322-338), never against the account table, so a name entry is storable and matches any later account with that name. +- issue: (pending cross-reference) + +### A03-016 TOTP codes are replayable within the acceptance window +- category: security +- severity: low +- location: internal/auth/service_mfa.go:77 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `verifyTOTP` accepts the current step and one on each side with no record of the last accepted counter. A code observed once (shoulder-surfed, phished, or captured from a request log) remains valid for up to 90 seconds and can be used for a second login, an `MFADisable`, or a recovery-code regeneration in that window. RFC 6238 section 5.2 requires rejecting a previously used code. +- evidence: + ```go + valid := 0 + for _, offset := range []int64{-1, 0, 1} { + counter := (currentTime / timeStep) + offset + expected := generateTOTP(secret, counter) + if subtle.ConstantTimeCompare([]byte(expected), []byte(code)) == 1 { + valid = 1 + } + } + ``` +- suggested fix: Persist the last accepted counter on the user row and refuse any counter less than or equal to it. +- verdict: CONFIRMED — verifyTOTP (service_mfa.go:63-87) accepts counters -1..+1 with no state, the User row carries no last-accepted counter (UpdateUser column list store_postgres.go:193-212) and a grep for any counter tracking in internal/auth returns nothing, so a captured code is accepted again within the skew window by Login, MFADisable and the recovery-code regeneration path. +- issue: (pending cross-reference) + +### A05-003 The "universal" 4-eyes gate is not on the SQS / cron execute path +- category: security +- severity: medium +- location: internal/purchase/approvals.go:167 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `enforceFourEyesPolicy` documents itself as the single choke point every approve/execute entry point inherits, and it runs only inside `ApproveAndExecute`. `claimAndExecute` (manager.go:193) goes straight to `executeAndFinalize`, so both of its callers — the SQS `execute_purchase` worker (messages.go:136) and the cron sweep `processOneExecution` (manager.go:678) — commit money without ever reaching the gate. With `RequireDifferentApprover` on and a plan carrying `AutoPurchase=true`, user A's own pending row is executed by the next cron tick with no second person involved, while the same row approved through the dashboard would be denied. The four-eyes tests cover `ApproveAndExecute` and `ProcessMessage`'s approve branch only; none covers `handleExecutePurchase` or `ProcessScheduledPurchases`. +- evidence: + ```go + // enforceFourEyesPolicy is the UNIVERSAL 4-eyes approval gate (issue #1005). + // It is the single choke point ApproveAndExecute runs before mutating any + // execution state, so every approve/execute entry point inherits the policy + // regardless of which caller reaches ApproveAndExecute: + ``` +- suggested fix: decide whether dual control is meant to bind AutoPurchase and say so at the gate. If it is, move the check into `executeAndFinalize` alongside `armedRedriveRefusal`, which is already positioned as the funnel every executor reaches money through; if it is not, correct the comment and name the AutoPurchase carve-out explicitly. +- verdict: PLAUSIBLE — the structural claim holds: claimAndExecute (manager.go:177-193) calls executeAndFinalize directly with no 4-eyes call, and the test list in approvals_test.go covers only ApproveAndExecute/ApproveExecution (:347-:558), never handleExecutePurchase or ProcessScheduledPurchases. But the stated scenario needs a pending/notified row that is BOTH non-web-sourced and carries a creator, and I could not find a writer that produces one: executableByScheduler rejects `common.PurchaseSourceWeb` (manager.go:635), and the only two writers of CreatedByUserID (handler_purchases.go:1901 and :2585) both stamp Source=cudly-web. The reachable divergence is narrower — a system-created row (nil creator, empty Source, AutoPurchase plan) executes on the cron tick where checkDifferentApprover (approvals.go:255) would have denied it fail-closed. +- severity-adjusted: low — no human can self-approve through this path today; the only rows that reach it have no creator to compare against. +- issue: (pending cross-reference) + +### A06-029 CUDLY_FEDERATED_SUBJECT is validated on the AWS and GCP onboarding scripts but not on either Azure one +- category: security +- severity: low +- location: internal/iacfiles/templates/azure-wif-cli.sh.tmpl:26 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the same knob is guarded on the two sibling paths — `aws-wif-cli.sh.tmpl:22-41` refuses an empty or `$`/`*`/whitespace-bearing `OIDC_SUBJECT_CLAIM` because IAM expands policy variables inside Condition values, and `gcp-wif-cli.sh.tmpl:50-60` pins a charset plus a 127-character cap because the value lands inside a CEL string literal. The Azure templates (also azure-wif-deploy.sh.tmpl:61) interpolate the override straight into the federated-credential JSON heredoc with no check, so a value containing `"` closes the `subject` string and injects sibling keys into the request body (for example widening `audiences`). Operator-supplied rather than attacker-supplied, so this is defense in depth, but the guard exists on two of three paths for a reason that applies to the third. +- evidence: + ```sh + CUDLY_FEDERATED_SUBJECT="${CUDLY_FEDERATED_SUBJECT:-cudly-controller}" + ... + "subject": "${CUDLY_FEDERATED_SUBJECT}", + ``` +- suggested fix: apply the GCP script's charset and length check to `CUDLY_FEDERATED_SUBJECT` in both Azure templates. +- verdict: CONFIRMED — both Azure templates take the override and interpolate it straight into the federated-credential JSON heredoc with no validation of any kind (internal/iacfiles/templates/azure-wif-cli.sh.tmpl:26,47-56 and azure-wif-deploy.sh.tmpl:61,66-75), while the two siblings do guard: the GCP script applies a charset regex plus a 127-character cap (gcp-wif-cli.sh.tmpl:50-61) and the AWS script rejects whitespace, `$` and `*` (aws-wif-cli.sh.tmpl:31-41). A `"` in the value closes the `subject` string and lets sibling keys such as `audiences` be injected into the request body. +- issue: (pending cross-reference) + +### A06-030 tfvars templates interpolate values into HCL string literals with no escaping +- category: security +- severity: low +- location: internal/iacfiles/templates/aws-wif.tfvars.tmpl:25 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the bundle path renders the `.auto.tfvars` files with the raw `data` (internal/api/handler_federation.go:592) — only the shell templates get `shellEscapeData`. Inside an HCL double-quoted string, `${...}` is a template interpolation, and `mail.ParseAddress` accepts `{`, `}` and `$` as RFC 5322 atext, so a user whose account email is `a${x}b@example.com` downloads a bundle whose `contact_email = "a${x}b@example.com"` line makes `terraform apply` fail with an unresolved-variable error, or evaluate an HCL expression. Reach is narrow today because `buildGenericIaCData` leaves `AccountName`, `AccountExternalID` and `ProjectID` empty and `ContactEmail` is the downloader's own address, but the tfvars path has no escaper at all while the two shell paths each have one. +- evidence: + ```hcl + cudly_api_url = "{{.CUDlyAPIURL}}" + contact_email = "{{.ContactEmail}}" + account_name = "{{.AccountName}}" + ``` +- suggested fix: add an `hclEscape` helper (escape `\`, `"`, and `$`/`%` before `{`) and apply it to the tfvars render the way `shellEscapeData` is applied to the script renders. +- verdict: CONFIRMED — `addBundleTerraform` renders the tfvars template with the raw `data` and no escaper of any kind (internal/api/handler_federation.go:591-593), and `shellEscapeData` is the only escaper in the file, applied solely to the two shell paths (handler_federation.go:273-274, 638-640); the values land unquoted-inside-quotes at internal/iacfiles/templates/aws-wif.tfvars.tmpl:24-26. I confirmed by running `mail.ParseAddress` that `a${x}b@example.com` parses cleanly (so the only validator, internal/api/validation.go:134-150, admits it) while a literal `"` does not, which bounds the impact to HCL interpolation rather than string-breakout. +- issue: (pending cross-reference) + +### A08-003 Five Azure service clients drop the SSRF-hardened HTTP client on the nil branch of NewClientWithHTTP +- category: security +- severity: high +- location: providers/azure/services/compute/client.go:134 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `NewClientWithHTTP(cred, sub, region, nil)` stores a nil `HTTPClient`. Every outbound call (`c.httpClient.Do(req)` in `fetchCapacityProviderState`, and the purchase two-step) nil-derefs and panics. The three clients that were fixed (managedredis:95, savingsplans:87, synapse:77) fall back to `httpclient.New()`, which installs the `blockIMDSDialer` that refuses `169.254.169.254`; compute, cache (client.go:103), cosmosdb (client.go:101), database (client.go:125) and search (client.go:80) have neither the guard nor the SSRF defence on that branch. +- evidence: + ```go + func NewClientWithHTTP(cred azcore.TokenCredential, subscriptionID, region string, httpClient HTTPClient) *ComputeClient { + return &ComputeClient{ + cred: cred, + subscriptionID: subscriptionID, + region: region, + httpClient: httpClient, + } + } + ``` +- suggested fix: add `if httpClient == nil { httpClient = httpclient.New() }` to the five constructors, matching managedredis/savingsplans/synapse. +- verdict: PLAUSIBLE — the nil guard is genuinely absent in the five constructors (compute/client.go:134, cache/client.go:103, cosmosdb/client.go:101, database/client.go:125, search/client.go:80) and present in the three siblings (managedredis/client.go:87, savingsplans/client.go:86, synapse/client.go:76), but no non-test caller of `NewClientWithHTTP` exists anywhere in the tree, so nothing passes nil today. +- severity-adjusted: low — reaching the nil deref requires a caller that does not exist; this is a constructor-consistency gap, not a live panic or SSRF exposure. +- issue: (pending cross-reference) + +### A08-016 The accounts-cache generation guard protects the cache write but not the value handed to in-flight callers +- category: security +- severity: medium +- location: providers/azure/accounts_cache.go:123 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `SetCredential` bumps `accountsGen` so a fetch that started under the old credential cannot publish. That fetch still returns its list to every caller that joined the singleflight. `GetRecommendationsClient` then reads the NEW credential (provider.go:659) and builds a fan-out client over the OLD credential's subscription list — precisely the hazard `SetCredential`'s own doc names: "Serving it to the new credential would report subscriptions this principal may have no access to, and — via GetRecommendationsClient's fan-out — fan out across them." +- evidence: + ```go + p.accountsMu.Lock() + if p.accountsGen == gen { + p.cachedAccounts = accounts + } + p.accountsMu.Unlock() + return accounts, nil + ``` +- suggested fix: when `p.accountsGen != gen`, return an error (or re-fetch under the new generation) instead of returning the stale-credential list, so the guard covers the returned value as well as the cache. +- verdict: PLAUSIBLE — `fetchAccountsShared` (accounts_cache.go:122-132) does return `accounts` whether or not the generation matched, so the pre-invalidation list reaches every joined caller; but SetCredential's own contract records that "every caller installs the credential before the first accounts fetch" (provider.go:223-224) and GetRecommendationsClient documents the same window as the accepted guarantee (provider.go:651-657), so the race needs a concurrent SetCredential no production caller performs. +- severity-adjusted: low — no live caller mutates the credential during an in-flight accounts fetch. +- issue: (pending cross-reference) + +### A08b-026 Recommendation SKU and region are interpolated unescaped into the Retail Prices OData `$filter` +- category: security +- severity: medium +- location: providers/azure/services/database/client.go:550 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cache/client.go:573`, `managedredis/client.go:468`, `synapse/client.go:417`) +- failure scenario: `rec.ResourceType` reaches `GetOfferingDetails` from a request body on the purchase/quote paths and is pasted between single quotes with no escaping. A value containing `'` terminates the literal: `X' or armSkuName ne '` turns the SKU predicate into a tautology, so the pricing walk returns every SKU in the region and, given the last-item-wins extractor in A08b-010, an attacker-chosen price is what the quote reports. A crafted value can also inject a `contains(...)` clause that makes the request expensive enough to exhaust the 50-page walk. +- evidence: + ```go + filter := fmt.Sprintf("serviceName eq 'SQL Database' and armRegionName eq '%s' and armSkuName eq '%s'", + region, sku) + ``` +- suggested fix: escape embedded single quotes by doubling them, and validate the SKU against `[A-Za-z0-9_.-]+` at the boundary before it reaches the filter builder. +- verdict: PLAUSIBLE — the unescaped interpolation into a single-quoted OData literal is confirmed at database:550-551, cache:573-574, managedredis:468-469 and synapse:417-418, and no `ResourceType` validation exists at the API boundary; what I could not establish is the delivery half, since the only route into these builders is `GetOfferingDetails`, which no code in this repo calls (grep outside `providers/` and tests yields only pkg/provider/interface.go:50). +- severity-adjusted: low — the injection primitive is real but currently has no reachable attacker-controlled feed. +- issue: (pending cross-reference) + +### A09-006 ExpectedAccount is silently ignored on the dependency-injected exchange path +- category: security +- severity: high +- location: pkg/exchange/exchange.go:343 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `assertAccount` is called only by the package-level `GetExchangeQuote` (line 222) and `ExecuteExchange` (line 295). The methods `ExchangeClient.GetQuote` and `ExchangeClient.Execute` — the path `internal/server/handler_ri_exchange.go:61` and `internal/server/ladder_write.go:75` construct and use — go straight to `getQuoteWithAPI`/`executeWithAPI`, neither of which calls it. `executeWithAPI` even copies `req.ExpectedAccount` into `quoteReq` at line 350, where it is also ignored. A caller that sets `ExpectedAccount` on the DI path believes it has a cross-account guard on an irreversible purchase and has none; the field is silently inert. +- evidence: + ```go + func executeWithAPI(ctx context.Context, client EC2ExchangeAPI, req ExchangeExecuteRequest) (string, *ExchangeQuoteSummary, error) { + if req.MaxPaymentDueUSD == nil { + return "", nil, fmt.Errorf("refusing to execute without max-payment-due-usd guardrail") + } + quoteReq := ExchangeQuoteRequest{ + Region: req.Region, + ExpectedAccount: req.ExpectedAccount, + ``` +- suggested fix: either drop `ExpectedAccount` from the two request structs so no caller can believe it is enforced, or move the STS check into `executeWithAPI`/`getQuoteWithAPI` behind an injected identity resolver so both entry paths honour it. +- verdict: CONFIRMED — `git grep -n assertAccount` returns only pkg/exchange/exchange.go:202 (definition), :222 and :295 (the two package-level wrappers); `GetQuote`/`Execute` at exchange.go:185 and :191 call `getQuoteWithAPI`/`executeWithAPI` directly, and the copy at line 349 is never read. +- severity-adjusted: low — no live caller sets `ExpectedAccount` on the DI path: `git grep -n ExpectedAccount` shows the only setters are ci_cd_sanity_tests/cmd/ri-exchange/main.go:98 and :153, both of which go through the package-level `GetExchangeQuote`/`ExecuteExchange` that do call `assertAccount`; internal/server/handler_ri_exchange.go:61 and ladder_write.go:75 never set the field. It is a latent trap, not an active bypass. +- issue: (pending cross-reference) + +### A10-028 The Azure service-principal client secret is printed to stdout unconditionally +- category: security +- severity: low +- location: cmd/configure_azure.go:531 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `printAzureSPResult` writes the freshly minted client secret to stdout with no TTY check. Run under `script`, `tee`, a CI job, a tmux logging pane, or simply left in terminal scrollback, a live Azure credential with the "Reservations Administrator" role is captured in plaintext. The choice mirrors `az ad sp create-for-rbac` and is deliberate, but that tool is not the one that then stores the same secret in Secrets Manager two prompts later, so CUDly could avoid the round trip through the screen entirely. +- evidence: + ```go + fmt.Printf(" appId (Client ID): %s\n", result.AppID) + fmt.Printf(" password (Client Secret): %s\n", result.ClientSecret) + fmt.Printf(" tenant (Tenant ID): %s\n", result.TenantID) + ``` +- suggested fix: Carry the wizard's result straight into `collectAzureCredentials` so the secret never reaches stdout; if it must be shown, gate the print on `term.IsTerminal(int(os.Stdout.Fd()))`. +- verdict: CONFIRMED — printAzureSPResult writes the secret with a bare fmt.Printf, no TTY test and no redaction (cmd/configure_azure.go:526-540), and it is reached unconditionally after a successful create at :519; the wizard then asks the operator to retype the same value in the next step, so the exposure is structural rather than incidental. +- issue: (pending cross-reference) + +### A12-015 `planId` is interpolated unescaped into a hidden input's `value` attribute +- category: security +- severity: medium +- location: frontend/src/plans.ts:2150 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `openAddPurchasesModal` builds its markup with `innerHTML`. `planName`, two lines above, goes through `escapeHtml`; `planId` does not. A plan id containing `" onfocus=alert(1) autofocus x="` breaks out of the attribute and executes in the Plans tab. The id reaches this function from `btn.dataset['id']`, which came from an API-supplied `plan.id`, so the trust boundary is the API response, not the DOM. +- evidence: + ```html +

Schedule additional purchases for ${escapeHtml(planName)}

+
+ + ``` +- suggested fix: Wrap it as `value="${escapeHtmlAttr(planId)}"`, matching the `planName` line directly above. +- verdict: PLAUSIBLE — planId really is interpolated unescaped into `value="${planId}"` one line below an escaped planName (frontend/src/plans.ts:2148-2150), but the only supplier is purchase_plans.id, a UUID primary key (internal/database/postgres/migrations/000001_initial_schema.up.sql:51), so a quote-bearing id is a runtime condition I could not establish. +- severity-adjusted: low — the only supplier of planId is a UUID primary key, so the missing escape is defence in depth rather than a reachable injection +- issue: (pending cross-reference) + +### A12-047 Server-controlled keys are assigned onto a plain object, allowing `__proto__` assignment +- category: security +- severity: low +- location: frontend/src/commitmentOptions.ts:288 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `fetchAndPopulateCommitmentOptions` iterates `Object.entries(body.aws)` and writes `awsConfigs[service] = {...}` with no key validation. A `/api/commitment-options` response containing a `"__proto__"` key invokes the `__proto__` setter and replaces the prototype of the module's AWS config object with the attacker-supplied value, changing what every unlisted service lookup resolves to. The lookup on line 110 has the mirror-image problem: `commitmentConfigs[provider.toLowerCase()]` resolves prototype members, so `getCommitmentConfig('constructor', 'name')` returns the string `"Object"` as a `CommitmentConfig` and `getValidPaymentOptions` then throws on `config.payments.filter`. +- evidence: + ```ts + const awsConfigs = commitmentConfigs.aws ?? (commitmentConfigs.aws = {}); + for (const [service, supportedCombos] of Object.entries(body.aws)) { + // ... + awsConfigs[service] = { terms: STANDARD_TERMS, payments: AWS_PAYMENTS, invalidCombinations: ... }; + } + ``` +- suggested fix: Skip keys that are not own-enumerable safe names (reject `__proto__`, `constructor`, `prototype`), and read the tables with `Object.prototype.hasOwnProperty.call`. +- verdict: PLAUSIBLE — The missing own-property guard is real, with keys assigned straight from Object.entries(body.aws) (frontend/src/commitmentOptions.ts:288) and raw-string indexing at :110, but reaching it needs the server to emit a `__proto__` or `constructor` key and those come from the DB-persisted AWS probe (internal/api/handler_commitment_options.go:53-58), a condition I could not establish. +- issue: (pending cross-reference) + +### A12-053 `canRevokeCompletedRow` omits the creator match its three sibling predicates enforce +- category: security +- severity: low +- location: frontend/src/history.ts:678 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `canCancelPendingRow`, `rbacAllowsApprove` and `canRetryFailedRow` all require `p.created_by_user_id === user.id` for the `-own` verb and deny on an absent creator. `canRevokeCompletedRow` returns `canAccess('revoke-own', 'purchases')` unconditionally, so any `revoke-own` holder sees a Revoke button on every in-window Azure row in the tenant and gets a 403 on click. The backend does enforce ownership, so this is the UX-versus-RBAC drift the function's own comment cites PR #995 for, in the one predicate that did not adopt the fix. +- evidence: + ```ts + if (canAccess('admin', '*') || canAccess('revoke-any', 'purchases')) return true; + return canAccess('revoke-own', 'purchases'); + ``` +- suggested fix: Mirror the siblings: `return canAccess('revoke-own','purchases') && !!p.created_by_user_id && p.created_by_user_id === user.id;`. +- verdict: CONFIRMED — canRevokeCompletedRow returns canAccess('revoke-own','purchases') with no creator comparison (frontend/src/history.ts:677-678) unlike its three siblings (frontend/src/history.ts:543-545, :575-578, :636-638), and the backend denies revoke-own on another user's execution (internal/api/handler_purchases_revoke_test.go:743-760), so the button renders and 403s. +- issue: (pending cross-reference) + +### A12-054 `_fourEyesMode` initialises to `false`, contradicting the documented fail-closed default +- category: security +- severity: low +- location: frontend/src/history.ts:47 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The doc comment at lines 50-62 states that an unreadable config "fails closed, to dual-control ON", and the catch path does set `true`. The module-level initialiser is `false`, so any render reaching `renderPendingActionButtons` before the first `refreshFourEyesMode()` resolves treats dual control as off and offers the creator an Approve button on their own row. All current paths await a `Promise.all` first, so this is latent, but the safe default costs nothing. +- evidence: + ```ts + let _fourEyesMode = false; + ``` +- suggested fix: Initialise to `true`. +- verdict: CONFIRMED — The module-level initialiser is `let _fourEyesMode = false` (frontend/src/history.ts:47) while the docstring directly above states an unreadable config fails closed to dual-control ON and the catch path does set true (frontend/src/history.ts:57-67). +- issue: (pending cross-reference) + +### A12-056 A handful of API-sourced numbers bypass escaping into `innerHTML` +- category: security +- severity: low +- location: frontend/src/history.ts:1154 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `api.getHistory` is cast with `as unknown as Promise` and never validated at runtime, so the TypeScript `number` types are assumptions the JSON does not guarantee. `p.count` (lines 1154, 1799) and `p.retry_attempt_n` (line 881, inside a `title="…"` attribute) are the only API values in the file that skip `escapeHtml`; the `title` occurrence would break out of the attribute. `riexchange.ts:1970` has the same shape for `config.max_payment_per_exchange_usd` and the other numeric settings, and `plans.ts:737` for `purchase.count`. +- evidence: + ```ts + ${p.count} + // line 881 + lineage.push(`↻ Retry #${p.retry_attempt_n}`); + ``` +- suggested fix: Coerce at the boundary, `escapeHtml(String(p.count ?? ''))` and `Number(p.retry_attempt_n) || 0`, so no unescaped API value reaches an `innerHTML` template. +- verdict: PLAUSIBLE — p.count and p.retry_attempt_n really are the only API values in the file that skip escaping, one of them inside a title attribute (frontend/src/history.ts:1154, :881, :1799), and the response is cast with no runtime validation (frontend/src/history.ts:346), but a break-out needs the API to emit a string where the type says number, which I could not establish. +- issue: (pending cross-reference) + +### A13-014 GCP CI impersonation is granted to the whole repository on a generically named shared pool +- category: security +- severity: medium +- location: terraform/environments/gcp/ci-cd-permissions/github_oidc.tf:46 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The `attribute_condition` on the provider pins repo and ref, but the `workloadIdentityUser` grant is `principalSet://.../attribute.repository/` with no ref component, and the pool is named `github-actions` rather than something CUDly-specific. Principal identifiers are pool-scoped, not provider-scoped, so any second provider added to that pool that can mint `attribute.repository` for this repo satisfies the grant and impersonates a service account carrying `roles/resourcemanager.projectIamAdmin`, `roles/iam.roleAdmin`, `roles/cloudkms.admin` and `roles/storage.admin`. The sibling module `iac/federation/gcp-target/terraform/main.tf:203-205` documents exactly this hazard and tells the reader to keep the pool dedicated; the CI pool does the opposite. +- evidence: + ```hcl + member = "principalSet://iam.googleapis.com/${google_iam_workload_identity_pool.github[0].name}/attribute.repository/${var.github_repo}" + ``` +- suggested fix: Map `attribute.ref` into the grant as well (`.../attribute.repository/` becomes a two-attribute condition or the member pins `attribute.ref`), and rename the pool to something CUDly-scoped so an unrelated provider cannot be added to it. +- verdict: PLAUSIBLE — every code fact checks out: the pool id is the generic `"github-actions"` (github_oidc.tf:4), the ref pin lives only in the provider's `attribute_condition` (:37), the `workloadIdentityUser` member is the ref-less `attribute.repository` principalSet (:46), the impersonated SA holds `roles/resourcemanager.projectIamAdmin`, `roles/iam.roleAdmin`, `roles/cloudkms.admin` and `roles/storage.admin` (service_account.tf:12,21,28,35), and iac/federation/gcp-target/terraform/main.tf:203-205 documents the pool-scoping hazard verbatim. The escalation itself needs a second provider to be added to that pool, which no code in the repo does and which I cannot establish from source; today the single provider's condition still pins repo and ref. +- severity-adjusted: low — with one provider in the pool this is a defense-in-depth gap, not a reachable escalation. +- issue: (pending cross-reference) + +### A13c-021 Rotation Lambda's RDS grant wildcards the account segment +- category: security +- severity: low +- location: terraform/modules/secrets/aws/main.tf:390 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the ARN's account field is `*`, so the statement authorises + `rds:ModifyDBInstance` against any DB instance in any account that would accept the principal. + IAM identity policies do not cross account boundaries on their own, so the practical reach is + today's account, but the pattern also fails to name the one instance being rotated: any + instance in the account, including unrelated stacks', is in scope. The module already knows the + target (`var.rds_cluster_id`, which gates this very resource at line 376) and does not use it. +- evidence: + ```hcl + Action = [ + "rds:ModifyDBInstance", + "rds:DescribeDBInstances" + ] + Resource = "arn:aws:rds:${var.region}:*:db:*" + ``` +- suggested fix: build the ARN from `data.aws_caller_identity.current.account_id` and + `var.rds_cluster_id`. +- verdict: CONFIRMED — `Resource = "arn:aws:rds:${var.region}:*:db:*"` at secrets/aws/main.tf:390 + wildcards both the account and the instance, and the same resource's own count at :376 already + holds the specific `var.rds_cluster_id`. Latent in practice: the rotation Lambda this role serves + cannot be applied at all (see A13c-003), so nothing assumes the role today. +- issue: (pending cross-reference) + +### A13c-022 GCP Cloud Run service account holds project-wide `roles/compute.viewer` +- category: security +- severity: low +- location: terraform/modules/compute/gcp/cloud-run/main.tf:431 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the stated need is `compute.commitments.list` and `compute.machineTypes.list`. + `roles/compute.viewer` grants read on every Compute Engine resource in the project — instance + metadata (including startup scripts, a common credential-leak surface), disks, images, + snapshots, firewall rules and network topology. The module already demonstrates the narrower + pattern immediately below by defining a two-permission custom role for the write side, so the + read side is the outlier rather than a constraint of the platform. +- evidence: + ```hcl + resource "google_project_iam_member" "compute_viewer" { + project = var.project_id + role = "roles/compute.viewer" + member = "serviceAccount:${google_service_account.cloud_run.email}" + } + ``` +- suggested fix: extend `cudlyCommitmentWriter` (or add a sibling custom role) with + `compute.commitments.list/get` and `compute.machineTypes.list/get`, and drop the predefined role. +- verdict: CONFIRMED — the ungated project-wide grant is at cloud-run/main.tf:431-435 and its own + comment names only `compute.commitments.list` and `compute.machineTypes.list` as the need, while + the two-permission custom role for the write side sits immediately below at :454-463, proving the + narrow pattern is available in this module. +- issue: (pending cross-reference) + +### A13c-023 Container App identity holds `Reader` over the entire host subscription +- category: security +- severity: low +- location: terraform/modules/compute/azure/container-apps/main.tf:269 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the grant is scoped to `/subscriptions/`, which is the same subscription + that holds CUDly's own Key Vault, PostgreSQL server, ACR and Container App environment. + A compromised runtime can therefore enumerate the whole control plane of its own deployment — + resource names, network layout, diagnostic settings, key vault metadata — not only the customer + workloads it is meant to inventory. The stated need ("host account ingested as a Self account") + is real, but it is an opt-in mode; the grant is unconditional. +- evidence: + ```hcl + resource "azurerm_role_assignment" "subscription_reader" { + scope = local.subscription_resource_id + role_definition_name = "Reader" + principal_id = azurerm_user_assigned_identity.container_app.principal_id + } + ``` +- suggested fix: gate it behind an `enable_self_account_ingestion` input defaulting false, so a + deployment that only manages customer subscriptions does not carry it. +- verdict: CONFIRMED — `azurerm_role_assignment.subscription_reader` (container-apps/main.tf:269) + has no `count` and is scoped to `local.subscription_resource_id` (:254), the same subscription + that holds the Key Vault, PostgreSQL server, ACR and Container App environment; the comment at + :265-268 states the Self-account rationale without gating on it. +- issue: (pending cross-reference) + +### A14-023 A migrate error is written unquoted into `$GITHUB_ENV` +- category: security +- severity: low +- location: .github/workflows/database-migration.yml:352 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `2>&1` folds stderr into `VERSION`, and golang-migrate's failure output is multi-line. Writing a multi-line value as a single `NAME=value` line to `$GITHUB_ENV` lets the second and subsequent lines be parsed as further environment assignments for later steps in the job. It is only bounded today because this is the job's last step. The error text can also carry the connection string, so it is echoed to the log as well. +- evidence: + ```bash + VERSION=$(migrate -path "$MIGRATIONS_PATH" -database "$DB_URL" version 2>&1 || echo "unknown") + echo "Current migration version: $VERSION" + echo "MIGRATION_VERSION=$VERSION" >> "$GITHUB_ENV" + ``` +- suggested fix: Capture stdout only, or use the heredoc delimiter form (`{name}< 0 { + return int64(cd.MemoryGB * 1024), nil + } + return 0, fmt.Errorf("memoryMBFromDetails: MEMORY resource amount absent from recommendation Details (no MEMORY op in Recommender payload); cannot build CUD insert without explicit memory") + } + ``` +- suggested fix: accept both forms in the assertion (`case common.ComputeDetails` / `case *common.ComputeDetails` in a type switch, guarding the typed-nil pointer), and add a regression test that drives `PurchaseCommitment` with `Details` set to `&common.ComputeDetails{MemoryGB: 16}`. +- verdict: CONFIRMED — pkg/common/service_details_codec.go:127 returns `&ComputeDetails{}` for ServiceCompute and internal/purchase/execution.go:1068 assigns that pointer to `recommendation.Details`, so the value-type assertion at providers/gcp/services/computeengine/client.go:1392 is always false on the executed purchase path. +- issue: (pending cross-reference) + +### A08b-005 `ReservationsDetails` is a daily usage API, so `GetExistingCommitments` emits one Commitment per reservation per day +- category: money-path +- severity: critical +- location: providers/azure/services/cache/client.go:207 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (identical in `cosmosdb/client.go:209`, `database/client.go:239`, `search/client.go:154`, `synapse/client.go:182`, `managedredis/client.go:209`) +- failure scenario: `armconsumption.ReservationDetailProperties` carries `UsageDate`, `ReservedHours` and `UsedHours` — it is one row per reservation per usage day, not a reservation inventory. With no `Filter` on `properties/UsageDate`, the pager returns the whole default window, so a single Redis reservation held for 30 days produces 30 `common.Commitment` values that all share the same `CommitmentID`. Every consumer that sums existing commitments to compute coverage, or that dedupes recommendations against existing commitments, sees roughly 30x the real committed capacity and suppresses purchases that should happen. +- evidence: + ```go + scope := fmt.Sprintf("subscriptions/%s", c.subscriptionID) + return client.NewListPager(scope, &armconsumption.ReservationsDetailsClientListOptions{}), nil + ``` +- suggested fix: use `armreservations.ReservationClient` / `ReservationOrderClient` (an inventory API, as `compute/exchange.go` already does) for existing commitments, or at minimum dedupe by `ReservationID` and pass a bounded `Filter` on `properties/UsageDate`. +- verdict: CONFIRMED — all six pagers pass an empty `ReservationsDetailsClientListOptions` (cache:208, cosmosdb:210, database:240, search:155, synapse:183, managedredis:210) and the SDK type is per-usage-day, carrying `UsageDate`, `ReservedHours` and `UsedHours` (armconsumption@v1.1.0 models.go:2229-2267); no converter dedupes on `ReservationID`. +- issue: (pending cross-reference) + +### A09-004 An absent PaymentDue is treated as $0, disabling the spend cap on an irreversible exchange +- category: money-path +- severity: critical +- location: pkg/exchange/exchange.go:305 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `getQuoteWithAPI` leaves `PaymentDueUSD` nil whenever `out.PaymentDue` is nil or empty (line 261 gates the parse on a non-empty raw string). `resolvePaymentDue` then substitutes a fresh zero `big.Rat`, so `checkInitialQuote` and `checkReQuote` compare `0` against `MaxPaymentDueUSD` and always pass. If AWS omits `PaymentDue` on a genuinely non-zero exchange (an API change, a partial response, or a shape the SDK does not populate), `executeWithAPI` proceeds to `AcceptReservedInstancesExchangeQuote` with no effective cap and the charge is irreversible. The doc comment asserts the only cause is a zero-cost exchange; nothing verifies that. +- evidence: + ```go + func resolvePaymentDue(q *ExchangeQuoteSummary) *big.Rat { + if q.PaymentDueUSD != nil { + return q.PaymentDueUSD + } + return new(big.Rat) + } + ``` +- suggested fix: make an absent `PaymentDue` an error in `checkInitialQuote`/`checkReQuote` rather than a zero, or require `IsValidExchange && PaymentDueRaw == "0"`-style positive evidence of a zero-cost exchange before proceeding. +- verdict: CONFIRMED — traced end to end: pkg/exchange/exchange.go:260 gates the parse on `s.PaymentDueRaw != ""` so `PaymentDueUSD` stays nil, `resolvePaymentDue` (exchange.go:303-308) substitutes `new(big.Rat)`, and both `checkInitialQuote` (line 313) and `checkReQuote` (line 330) compare that zero against the cap, so `executeWithAPI` reaches `AcceptReservedInstancesExchangeQuote` at line 372 with no effective ceiling. +- issue: (pending cross-reference) + +### A12-001 Purchase modal's Term and Payment selects mutate the submitted rec without re-pricing it +- category: money-path +- severity: critical +- location: frontend/src/recommendations.ts:5412 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A 3yr all-upfront rec (`upfront_cost` $36,000, `monthly_cost` 0) is opened in the purchase modal. The user changes the row's Term select to 1yr. `live.term` becomes 1, `rebuildPaymentOptions` re-derives the payment list, and nothing else changes. The Upfront / Monthly Cost / Eff. Savings / Eff. % cells are static text nodes built once in `renderPurchaseModalRow` and are never re-rendered; `updatePurchaseModalTotals` reads the same untouched `rec.upfront_cost` / `rec.savings`. The POST therefore carries `term: 1` with the 3-year price, and the totals row, the approval email and the stored execution all record a price belonging to a term the user did not buy. The Payment select (line 5430) has the same shape: flipping `no-upfront` to `all-upfront` leaves `upfront_cost` at the no-upfront value, so the modal shows "$0 upfront" for an all-upfront purchase. `recTotalCommitment` in `internal/api/handler_purchases.go:2247` multiplies the submitted `MonthlyCost` by the submitted `Term`, so the `MaxPurchaseAmount` permission constraint is also evaluated against a mismatched pair. +- evidence: + ```ts + termSelect.addEventListener('change', () => { + const live = currentPurchaseRecommendations[idx]; + if (!live) return; + const newTerm = parseInt(termSelect.value, 10) === 3 ? 3 : 1; + live.term = newTerm; + rebuildPaymentOptions(paymentSelect, live.provider as CompatProvider, + live.service, newTerm, (live.payment ?? '') as CompatPayment); + live.payment = paymentSelect.value; + }); + ``` +- suggested fix: On a term/payment change, look up the matching variant from the loaded recommendation set (the API already fans out per `(term, payment)` cell) and swap the whole rec in, then re-render the row and the totals; if no matching variant exists, disable that option rather than keeping stale prices. +- verdict: CONFIRMED — The term/payment handlers mutate only `live.term`/`live.payment` (frontend/src/recommendations.ts:5412-5435) while the Upfront/Monthly/Eff. cells are one-shot text nodes built in renderPurchaseModalRow (frontend/src/recommendations.ts:5339-5364), and app.ts:385-391 spreads the same rec into the POST, so the submitted term reaches recTotalCommitment (internal/api/handler_purchases.go:2255) paired with the other term's price. +- issue: (pending cross-reference) + +### A12-002 Fan-out modal says an incompatible bucket "will be skipped", then submits it +- category: money-path +- severity: critical +- location: frontend/src/recommendations.ts:4336 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A bulk selection produces two buckets, one of which fails `isBucketPaymentCompatible` (for example AWS RDS 3yr + no-upfront). `renderFanOutBucketSection` renders "Invalid combo: … This bucket will be skipped." `handleFanOutExecute` in `frontend/src/app.ts:534` maps over `buckets` with no compatibility filter, so both buckets are POSTed. The backend only warns for that combination (`cmd/validators.go:warnRDS3YearNoUpfront`), so an approval email is minted for a commitment the UI just told the user would not be submitted. The confirm dialog and the "Total upfront / Total savings" header (`computeFanOutTotals`, line 4295) also count the skipped bucket's money. +- evidence: + ```ts + status.textContent = compat + ? `${b.capacityPercent}% capacity · ${b.term}yr · ${b.payment}` + : `Invalid combo: ${b.provider} / ${serviceLabel} doesn't support ${b.term}yr + ${b.payment}. This bucket will be skipped.`; + // app.ts:534 — no filter: + const promises = buckets.map((b) => api.executePurchase(b.recs.map(...), b.capacityPercent)); + ``` +- suggested fix: Filter incompatible buckets out of `currentFanOutBuckets` before `handleFanOutExecute` runs (or exclude them from `getFanOutBuckets`), and exclude them from `computeFanOutTotals` and the email count. +- verdict: CONFIRMED — handleFanOutExecute maps over every bucket with no compatibility predicate (frontend/src/app.ts:534-544) even though handleBulkPurchaseClick deliberately routes incompatible buckets into the modal (frontend/src/recommendations.ts:3976-3988) and renderFanOutBucketSection promises they are skipped (frontend/src/recommendations.ts:4334-4337); no rejection exists on the API side, only the CLI-flag warning at cmd/validators.go:150. +- issue: (pending cross-reference) + +### A01-001 MaxPurchaseAmount cap is enforced against client-asserted dollar amounts, not the priced commitment +- category: money-path +- severity: high +- location: internal/api/handler_purchases.go:2255 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A user whose execute:purchases permission carries MaxPurchaseAmount=1000 and execute-own (or execute-any) POSTs `/api/purchases/execute` with `execute_mode:"direct"` and a rec `{provider:aws, service:ec2, resource_type:"m5.24xlarge", count:100, term:3, payment:"all-upfront", upfront_cost:1, monthly_cost:null}`. `recTotalCommitment` sums the request's own `upfront_cost`/`monthly_cost` (1.0), `requireNonZeroCommitment` only rejects an exact 0, the constraint check passes (1 <= 1000), `directExecutePurchase` runs `ApproveAndExecute`, and the provider buys 100 three-year RIs at the real list price. Nothing in `validateExecutePurchaseRequest` re-derives the amount from the stored recommendation (`rec.ID` is never looked up) or from provider pricing. The retry path inherits the same self-reported totals. +- evidence: + ```go + func recTotalCommitment(rec *config.RecommendationRecord) float64 { + total := rec.UpfrontCost + if rec.Term > 0 && rec.MonthlyCost != nil { + total += *rec.MonthlyCost * float64(rec.Term*12) + } + return total + } + // ... + func requireNonZeroCommitment(sets []auth.PermissionConstraints) error { + if len(sets) > 0 && sets[0].MaxPurchaseAmount == 0 { + ``` +- suggested fix: Before the constraint check, resolve each rec by `rec.ID` against the cached recommendations store (the scheduler's `ListRecommendations`/`GetRecommendationByID`) and take upfront/monthly cost from the stored row scaled by the requested count, refusing recs that do not resolve; treat the client-sent costs as untrusted input. +- verdict: CONFIRMED — validateExecutePurchaseRequest (internal/api/handler_purchases.go:2127-2185) never resolves rec.ID against any store (no GetRecommendationByID caller in internal/api or internal/purchase), purchaseConstraintSets:2221-2242 caps on recTotalCommitment of the request's own UpfrontCost/MonthlyCost, and the AWS purchase at providers/aws/services/ec2/client.go:155-161 sends only offering ID + count with no LimitPrice, so nothing downstream re-prices the batch. +- issue: (pending cross-reference) + +### A01-004 PUT /api/plans/{id} rebuilds the ramp schedule from scratch, resetting CurrentStep/StartDate and re-arming the ramp +- category: money-path +- severity: high +- location: internal/api/handler_plans.go:281 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A weekly-25pct plan is on `CurrentStep=3` (three steps bought). The operator renames it through the Edit modal, which issues PUT (frontend/src/plans.ts:1420). `updatePlan` builds `plan := req.toPurchasePlan()`, whose `buildRampSchedule` returns the preset with `CurrentStep=0`, `StartDate=now` (types.go:797-801) and `NextExecutionDate=now+7d`; `LastExecutionDate`/`LastNotificationSent` are left nil. `UpdatePurchasePlan` writes all of it (`ramp_schedule = $7`, `next_execution_date = $9`, `last_execution_date = $10`, store_postgres.go:763-774). The scheduler's `getOrCreateExecution` then finds no row for the new date and mints a `StepNumber=1` execution (notifications.go:139), so steps 1..4 are notified and bought again. `refuseOccupiedRampSteps` guards only the manual create endpoint. `TestHandler_updatePlan` (handler_plans_test.go:551) never asserts ramp preservation. +- evidence: + ```go + // Create new plan from request + plan := req.toPurchasePlan() + plan.ID = planID + + // Preserve timestamps from existing plan + plan.CreatedAt = existingPlan.CreatedAt + plan.UpdatedAt = time.Now() + ``` +- suggested fix: Carry `existingPlan.RampSchedule.CurrentStep`, `StartDate`, `NextExecutionDate`, `LastExecutionDate` and `LastNotificationSent` onto the rebuilt plan (or only rebuild the ramp when the requested ramp type/params actually changed), and add a test asserting a rename leaves `CurrentStep` untouched. +- verdict: CONFIRMED — frontend/src/plans.ts:1420 calls api.updatePlan which is PUT (frontend/src/api/plans.ts:45-50); updatePlan (internal/api/handler_plans.go:281-300) rebuilds via req.toPurchasePlan(), whose buildRampSchedule (internal/api/types.go:797-801) copies the preset (weekly-25pct carries no CurrentStep, internal/config/types.go:241-246) with StartDate=now and leaves LastExecutionDate/LastNotificationSent nil; UpdatePurchasePlanTx (internal/config/store_postgres.go:763-786) overwrites ramp_schedule, next_execution_date, last_execution_date and last_notification_sent, and getOrCreateExecution stamps StepNumber=CurrentStep+1 (internal/purchase/notifications.go:139); TestHandler_updatePlan (handler_plans_test.go:551-590) asserts only ID and Name. +- issue: (pending cross-reference) + +### A07-001 CE-supplied OfferingID short-circuits every Savings Plans purchase validation +- category: money-path +- severity: high +- location: providers/aws/services/savingsplans/client.go:425 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `findOfferingID` returns `spDetails.OfferingID` before `resolveSPPlanType`, `convertTermToSeconds` and `convertPaymentOption` run. `extractEC2SPFields` (parser_sp.go:268) populates `OfferingID` on every EC2Instance SP rec, so for those recs the scoped client's "reject mismatches to prevent buying the wrong product" check at client.go:279 never executes, and neither does term or payment-option validation. With a persisted rec whose `Term` was later changed from `1yr` to `3yr` by config (or a rec routed to a client scoped to a different plan type), `PurchaseCommitment` sends the stale CE offering ID to `CreateSavingsPlan` and buys the term/payment/plan-type the offering encodes, not the one the recommendation now says. +- evidence: + ```go + if spDetails.OfferingID != "" { + log.Printf("purchase[%s]: SavingsPlans findOfferingID: using CE-provided OfferingID %s (skipping DescribeSavingsPlansOfferings)", tag, spDetails.OfferingID) + return spDetails.OfferingID, nil + } + planType, err := c.resolveSPPlanType(spDetails.PlanType) + ``` +- suggested fix: run `resolveSPPlanType`, `convertTermToSeconds` and `convertPaymentOption` before the short-circuit so a malformed or mis-scoped rec still fails loud, and keep the fast path only for the offering lookup itself. +- verdict: CONFIRMED — providers/aws/services/savingsplans/client.go:425-428 returns before the three validators at :430-441, and providers/aws/recommendations/parser_sp.go:268 + :397 populate `OfferingID` on every EC2Instance SP rec, so the scoped-client mismatch check at client.go:279 is bypassed on exactly those recs. +- issue: (pending cross-reference) + +### A07-007 ElastiCache and MemoryDB commitments carry no Engine, so the duplicate-purchase guard never matches +- category: money-path +- severity: high +- location: providers/aws/services/elasticache/client.go:92 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `recfilter.dedupeKey` is built from `(resourceType, region, engine, deployment)`. For a cache recommendation the engine comes from `common.EngineFromDetails(rec.Details)` and is `"redis"`; the commitment side reads `c.Engine`, which neither `elasticache.GetExistingCommitments` nor `memorydb.GetExistingCommitments` ever sets. The two keys can never be equal, so `AdjustRecommendationsForExisting` never suppresses a cache recommendation and a reserved cache node bought minutes ago is recommended and bought again on the next run. RDS sets `Engine` from `instance.ProductDescription` (rds/client.go:110) and does not have this problem; `types.ReservedCacheNode.ProductDescription` is available and simply unread. +- evidence: + ```go + commitment := common.Commitment{ + Provider: common.ProviderAWS, + CommitmentID: aws.ToString(node.ReservedCacheNodeId), + CommitmentType: common.CommitmentReservedInstance, + Service: common.ServiceCache, + Region: c.region, + ResourceType: aws.ToString(node.CacheNodeType), + Count: int(aws.ToInt32(node.CacheNodeCount)), + ``` +- suggested fix: set `Engine: aws.ToString(node.ProductDescription)` in the ElastiCache mapping and the equivalent field in `memorydb.GetExistingCommitments`, and add a dedupe test that fails when `Engine` is blank. +- verdict: CONFIRMED — neither providers/aws/services/elasticache/client.go:92-102 nor providers/aws/services/memorydb/client.go:88-99 sets `Engine`, while pkg/recfilter/dedupe.go:105 keys the existing map on `NormalizeEngineName(c.Engine)` (empty) and :135-137 keys the rec on `EngineFromDetails(rec.Details)` (`"redis"`); `types.ReservedCacheNode.ProductDescription` does exist in the SDK (elasticache@v1.50.3/types/types.go:1758) and is never read. +- issue: (pending cross-reference) + +### A07-012 EC2 ladder coverage blends in ElastiCache, OpenSearch, Redshift and MemoryDB pools +- category: money-path +- severity: high +- location: providers/aws/ladder/layer_states.go:307 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `computeEC2CoveragePct` selects map entries by "one colon in the key". But `coverageServiceFilters` (coverage.go:28) populates that same map with ElastiCache, OpenSearch, Redshift and MemoryDB pools, all keyed `region:instance_type` with exactly one colon. An account with a fully covered Redshift or ElastiCache fleet in the region produces a high blended `CoveragePct` on `LayerConvertibleRI`, which the ladder engine reads as "the EC2 convertible-RI buffer is covered" and uses to gate buy and reshape decisions. This is the same cross-service blend PR #1361 fixed for utilization (see `utilsForConvertibleRIs` 30 lines below), left unfixed on the coverage side. +- evidence: + ```go + for key, cov := range coverageMap { + if !strings.HasPrefix(key, prefix) { + continue + } + // Exclude RDS keys (contain extra ":" segments for engine:deployment). + // EC2 pool keys are exactly "region:instance_type" (one colon). + if strings.Count(key, ":") != 1 { + continue + } + ``` +- suggested fix: have `GetRICoverageMap` record the CE service alongside each pool (or key non-RDS entries by `region:service:instance_type`) so the EC2 aggregate can select only EC2 pools. +- verdict: CONFIRMED — `coverageServiceFilters` (providers/aws/recommendations/coverage.go:28-33) lists ElastiCache, OpenSearch, Redshift and MemoryDB alongside EC2, `GetRICoverageMap` :193-196 loops all five, and `fetchCoverageForServiceRegion` :250 writes every one of them as `poolKey(region, instType)` — exactly one colon — which `computeEC2CoveragePct` (providers/aws/ladder/layer_states.go:296-307) then accepts into the `LayerConvertibleRI` aggregate. +- issue: (pending cross-reference) + +### A08-005 Azure retail-price extraction is last-item-wins across Spot/Windows/Linux SKUs and silently mixes currencies +- category: money-path +- severity: high +- location: providers/azure/services/compute/client.go:765 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the `armSkuName eq ''` filter returns every meter for the size, including Spot, Low Priority, Windows and Linux variants. The loop keeps the LAST matching item for both `onDemand` and `reservation` with no disambiguation on `SKUName`/`MeterName`/`ProductName` (all present on `pricing.RetailPriceItem`), so `GetOfferingDetails` can report a Spot on-demand rate against a standard reservation price, and `savingsPercentage` (client.go:719) is computed from that pair. `currency` is likewise the last item's code and is not required to match the item the prices came from; when no item carries a code it stays hardcoded `"USD"`. The identical shape is in database/client.go:617, cache/client.go:599, cosmosdb/client.go:597, search/client.go:510, synapse/client.go:456 and managedredis/client.go:521. +- evidence: + ```go + currency = "USD" + for _, item := range items { + if item.CurrencyCode != "" { currency = item.CurrencyCode } + if item.ReservationTerm == termStr { + reservation = item.RetailPrice + } else if item.Type == "Consumption" { + onDemand = item.UnitPrice + } + } + ``` +- suggested fix: reject Spot/Low-Priority meters and require the on-demand and reservation items to share a currency and SKU name; error rather than defaulting `currency` to `"USD"` when no item reports one. +- verdict: CONFIRMED — `extractVMPricing` (compute/client.go:765) is last-item-wins across the whole `$filter` result with no SKUName/MeterName/ProductName disambiguation even though those fields exist on the item (internal/pricing/types.go:16-20), the filter itself is only serviceName+armRegionName+armSkuName (client.go:692), and the same loop shape is at database:617, cache:599, cosmosdb:597, search:510, synapse:456 and managedredis:521. +- issue: (pending cross-reference) + +### A08-006 ExpandPaymentVariants fabricates 100% savings when the provider omitted the commitment cost +- category: money-path +- severity: high +- location: providers/azure/internal/recommendations/converter.go:423 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: a Legacy payload with `CostWithNoReservedInstances = 4000` and no `TotalCostWithReservedInstances` yields `CommitmentCost == 0`. `ExpandPaymentVariants` then computes `savings = 4000 - 0 = 4000` and `savingsPct = 100`, overwriting the provider-reported `NetSavings` the extractor had populated, and emits both variants as a 100%-savings recommendation. Every Azure service converter calls this unconditionally (e.g. compute/client.go:218), so nothing upstream validates it. The function's own doc claims the opposite: "if CommitmentCost is zero both variants are still emitted with zero costs". +- evidence: + ```go + var savingsPct float64 + var savings float64 + if totalOnDemand != 0 { + savings = totalOnDemand - totalReservation + savingsPct = savings / totalOnDemand * 100 + } + ``` +- suggested fix: when `totalReservation == 0`, keep `base.EstimatedSavings` (the provider's `NetSavings`) and leave `SavingsPercentage` unset, or drop the recommendation; either way stop deriving savings from an absent commitment cost. +- verdict: CONFIRMED — `extractLegacy` (converter.go:156) leaves CommitmentCost at 0 when TotalCostWithReservedInstances is nil (an absence the repo itself treats as expected — `deriveCoveredMonthlyCost`, converter.go:90, exists for it) and `ExpandPaymentVariants` (converter.go:415-425) then overwrites the provider's NetSavings with `totalOnDemand - 0` and a 100% SavingsPercentage on both variants. +- issue: (pending cross-reference) + +### A08-007 Azure Savings Plan purchases hardcode CurrencyCode "USD" on the commitment body +- category: money-path +- severity: high +- location: providers/azure/services/savingsplans/client.go:261 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `PurchaseCommitment` sends `spDetails.HourlyCommitment` with `CurrencyCode: "USD"` regardless of the billing account's actual billing currency. For a tenant billed in EUR/GBP/KWD, a recommendation whose hourly commitment was derived in the billing currency is committed as that many US dollars. The same literal is on the validate body (client.go:355) so `ValidateOffering` cannot catch the mismatch, and `GetOfferingDetails` reports `Currency: "USD"` (client.go:445) so the UI corroborates the wrong denomination. `auth.PermissionConstraints.MaxPurchaseAmount` is USD-denominated with no currency field, so a non-USD amount also compares against the spend cap in the wrong units. +- evidence: + ```go + Commitment: &armbillingbenefits.Commitment{ + Amount: &hourlyAmount, + CurrencyCode: toPtr("USD"), + Grain: &grain, + }, + ``` +- suggested fix: carry the currency on `common.SavingsPlanDetails` and reject the purchase when it is absent or not USD, rather than stamping USD onto whatever amount arrived. +- verdict: CONFIRMED — the literal is stamped on the purchase body (savingsplans/client.go:261), the validate body (client.go:355) and the reported offering currency (client.go:445), and `common.SavingsPlanDetails` (pkg/common/types.go:612-634) carries no currency field for the purchase to check against. +- issue: (pending cross-reference) + +### A08-008 Modern (MCA) recommendation extraction discards the currency Azure reported +- category: money-path +- severity: high +- location: providers/azure/internal/recommendations/converter.go:360 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `armconsumption.Amount` carries `{Currency, Value}`. `amountValue` unwraps `.Value` and throws `.Currency` away, so an MCA billing account billed in EUR produces `OnDemandCost`, `CommitmentCost` and `EstimatedSavings` in EUR that are stored, ranked, summed and compared against USD-denominated spend caps as if they were dollars. Nothing downstream can detect it because `common.Recommendation` has no currency field on this path. The comment states the assumption ("downstream consumers assume a single-currency view per subscription") but there is no check that enforces it. +- evidence: + ```go + func amountValue(a *armconsumption.Amount) float64 { + if a == nil || a.Value == nil { + return 0 + } + return *a.Value + } + ``` +- suggested fix: read `a.Currency` and return an error (or drop the recommendation with a loud log) when it is set to anything other than USD, so a non-USD tenant fails closed instead of silently mis-denominating a money path. +- verdict: CONFIRMED — `amountValue` (converter.go:360) reads only `.Value`, and `extractModern` (converter.go:206-215) feeds it straight into OnDemandCost, CommitmentCost and EstimatedSavings with no currency check anywhere on the path; `amountValuePtr` (converter.go:369) discards it too. +- issue: (pending cross-reference) + +### A08b-006 Azure commitments fabricate `State: "active"` and `Region` from fields the SDK response does not carry +- category: money-path +- severity: high +- location: providers/azure/services/cache/client.go:251 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (identical in `cosmosdb/client.go:252`, `database/client.go:282`, `search/client.go:197`, `synapse/client.go:223`, `managedredis/client.go:223`) +- failure scenario: `ReservationDetailProperties` has no state and no region field, so both values are invented. The reservation listing is subscription-wide and unfiltered by region, so a `westeurope` reservation returned to a client constructed for `eastus` is stamped `Region: "eastus"` and matched against `eastus` recommendations that it does not discount. `State: "active"` is unconditional, so an expired or cancelled reservation still counts as live coverage and suppresses a needed repurchase. +- evidence: + ```go + commitment := &common.Commitment{ + Provider: common.ProviderAzure, + Account: c.subscriptionID, + CommitmentType: common.CommitmentReservedInstance, + Service: common.ServiceCache, + Region: c.region, + State: "active", + } + ``` +- suggested fix: source region and state from an inventory API that reports them (`armreservations.ReservationResponse` has `Location` and `ProvisioningState`), and leave them empty rather than guessing when unavailable. +- verdict: CONFIRMED — `ReservationDetailProperties` (armconsumption@v1.1.0 models.go:2229-2267) carries neither a state nor a region field, the pager scope is the whole subscription with no region filter (cache:207-208), and the converter hardcodes `Region: c.region` and `State: "active"` (cache:172-179, managedredis:223-230, synapse:223-230). +- issue: (pending cross-reference) + +### A08b-007 Azure commitments never populate Count, StartDate, EndDate or Cost +- category: money-path +- severity: high +- location: providers/azure/services/cache/client.go:260 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (identical in the five sibling converters) +- failure scenario: `common.Commitment` has `Count`, `StartDate`, `EndDate` and `Cost` (`pkg/common/types.go:436-440`); the converters set only `CommitmentID` and `ResourceType`. `TotalReservedQuantity` is present on the SDK response and dropped, so every Azure commitment reports `Count: 0` — coverage arithmetic that multiplies count by a rate yields zero committed capacity. `EndDate` stays at the zero `time.Time`, so any "is this commitment still active" filter comparing against `time.Now()` classifies every Azure commitment as long expired. +- evidence: + ```go + if props.ReservationID != nil { + commitment.CommitmentID = *props.ReservationID + } + if props.SKUName != nil { + commitment.ResourceType = *props.SKUName + } + return commitment + ``` +- suggested fix: populate `Count` from `TotalReservedQuantity`, and take start/end/cost from a reservation-order lookup; do not emit a Commitment at all when the required money fields cannot be filled. +- verdict: CONFIRMED — the converters write only `CommitmentID` and `ResourceType` (cache/client.go:181-188, managedredis:231-234, synapse:231-234); `Count`, `StartDate`, `EndDate` and `Cost` exist on the struct at pkg/common/types.go:436-441, and `TotalReservedQuantity` is present on the SDK response (models.go:2260) but never read. +- issue: (pending cross-reference) + +### A08b-008 Cosmos DB pricing ignores the SKU it was asked to price +- category: money-path +- severity: critical +- location: providers/azure/services/cosmosdb/client.go:532 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `getCosmosPricing(ctx, sku, region, termYears)` builds a filter on service name and region only; the `sku` parameter is never read. The Retail Prices response therefore contains every Cosmos DB meter in the region (provisioned RU/s, autoscale RU/s, transactional storage, analytical storage, backup, restore), and `extractCosmosPricing` keeps the last matching item of each kind. `GetOfferingDetails` then reports that arbitrary meter's price as `TotalCost`/`EffectiveHourlyRate` for the specific reservation being priced — for instance a backup-storage price presented as the cost of a 100 RU/s three-year reservation. +- evidence: + ```go + func (c *CosmosDBClient) getCosmosPricing(ctx context.Context, sku, region string, termYears int) (*CosmosPricing, error) { + filter := fmt.Sprintf("serviceName eq 'Azure Cosmos DB' and armRegionName eq '%s'", region) + priceData, err := c.fetchAzurePricing(ctx, filter) + ``` +- suggested fix: add an exact `armSkuName eq` (or `skuName eq`) clause for the SKU, and reject the lookup when more than one distinct reservation price survives the filter rather than taking the last one. +- verdict: CONFIRMED — `sku` appears nowhere in the body of `getCosmosPricing` (cosmosdb/client.go:531-567); the filter is service plus region only and `extractCosmosPricing` (597-614) overwrites on every match, so the last item wins. +- severity-adjusted: high — no in-repo caller reaches `GetOfferingDetails`; grepping the whole tree outside `providers/` and tests returns only the interface declaration at pkg/provider/interface.go:50, so the wrong quote is latent rather than shipping today. +- issue: (pending cross-reference) + +### A08b-009 Azure Search pricing ignores the SKU it was asked to price +- category: money-path +- severity: critical +- location: providers/azure/services/search/client.go:447 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: identical shape to A08b-008. `sku` is accepted and used only inside an error string at line 470; the filter selects every Azure Cognitive Search meter in the region. Pricing a `standard3` reservation returns whichever `basic` or `storage_optimized_l2` item happened to sort last in the response, so `TotalCost` can be off by an order of magnitude in either direction on a path that feeds purchase previews. +- evidence: + ```go + func (c *SearchClient) getSearchPricing(ctx context.Context, sku, region string, termYears int) (*SearchPricing, error) { + filter := fmt.Sprintf("serviceName eq 'Azure Cognitive Search' and armRegionName eq '%s'", region) + ``` +- suggested fix: filter on the SKU, and fail loud when the filtered result set contains more than one candidate reservation price. +- verdict: CONFIRMED — the filter at search/client.go:447 carries service and region only, `sku` is read solely inside the error string at line 470, and `extractSearchPricing` (510-527) is last-item-wins. +- severity-adjusted: high — same reachability caveat as A08b-008: `GetOfferingDetails` has no caller outside the providers' own tests (pkg/provider/interface.go:50). +- issue: (pending cross-reference) + +### A08b-010 Retail-price extraction is last-item-wins with no SKU or unit-of-measure check +- category: money-path +- severity: high +- location: providers/azure/services/cache/client.go:599 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same in `cosmosdb/client.go:597`, `database/client.go:617`, `search/client.go:510`, `managedredis/client.go:520`, `synapse/client.go:456`) +- failure scenario: the loop overwrites `reservation` and `onDemand` on every match, so the final values come from whichever item the API returned last. In cache and managedredis the query uses `contains(armSkuName, 'Premium_P1')`, which also matches `Premium_P10` through `Premium_P15`, so the price attributed to a P1 reservation can be a P15 price. `UnitOfMeasure` is never inspected, so a "100 Hours" meter and a "1 Hour" meter are treated as interchangeable. On-demand is read from `UnitPrice` while the reservation is read from `RetailPrice`, two different fields of the same record, and the resulting savings percentage compares them directly. +- evidence: + ```go + for _, item := range items { + if item.CurrencyCode != "" { + currency = item.CurrencyCode + } + if item.ReservationTerm == termStr { + reservation = item.RetailPrice + } else if item.Type == "Consumption" { + onDemand = item.UnitPrice + } + } + ``` +- suggested fix: require an exact `ArmSKUName` match against the requested SKU, assert a single `UnitOfMeasure`, and error when two candidate items disagree instead of letting the last one win. +- verdict: CONFIRMED — the loops overwrite unconditionally at cache:603-613, cosmosdb:601-611, search:514-524, managedredis:523-532 and synapse:463-475; `contains(armSkuName, '%s')` is the filter at cache:573 and managedredis:468; `UnitOfMeasure` exists on the shared item (providers/azure/internal/pricing/types.go:21) and is read nowhere; on-demand comes from `UnitPrice` while the reservation comes from `RetailPrice`. +- issue: (pending cross-reference) + +### A08b-014 Cloud Storage prices a per-GiB-month SKU as if it were per-hour, inflating cost by 730x +- category: money-path +- severity: high +- location: providers/gcp/services/cloudstorage/client.go:356 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: GCS storage SKUs are denominated per GiB-month (`unitOfMeasure` "GiBy.mo"); none are per-hour. Multiplying the catalog unit price by `8760 * termYears` therefore converts a monthly rate into a fictitious term total 730x too large per year. `CommitmentPrice` and `OnDemandPrice` both flow into `common.OfferingDetails.TotalCost` and into `rec.CommitmentCost` via `fillStoragePricing`, and `HourlyRate` labels a per-GiB-month figure as an hourly rate. +- evidence: + ```go + hoursInTerm := 8760.0 * float64(termYears) + commitmentPriceTerm := commitmentPrice * hoursInTerm + savingsPercentage := calculateStorageSavingsPercentage(onDemandPrice, hoursInTerm, commitmentPriceTerm) + ``` +- suggested fix: read `PricingExpression.UsageUnit` and scale by the unit the SKU actually uses (months for GiBy.mo), erroring when the unit is unrecognized rather than assuming hours. +- verdict: CONFIRMED — cloudstorage/client.go:356-368 multiplies the catalog unit price by `8760 * termYears` and returns the raw per-unit price as `HourlyRate`; `extractStoragePriceFromSKU` (415-432) reads only `TieredRates[0].UnitPrice` and never touches `PricingExpression.UsageUnit`, so no unit check exists anywhere on the path. Unlike the sibling clients this one is live: `fillStoragePricing` (497-512) runs inside `convertGCPRecommendation`. +- issue: (pending cross-reference) + +### A08b-024 Azure Cache for Redis is enumerated by two clients, so its reservations and recommendations are counted twice +- category: money-path +- severity: high +- location: providers/azure/services/managedredis/client.go:41 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (twin at `cache/client.go:157`; both registered at `providers/azure/provider.go:521-522` and dispatched at `provider.go:586-591`) +- failure scenario: both clients issue the identical Consumption filter `properties/resourceType eq 'RedisCache'` and both classify existing reservations with `strings.Contains(strings.ToLower(*props.SKUName), "redis")`. `SupportedServices()` lists `ServiceCache` and `ServiceMemoryDB`, so a caller iterating supported services receives every Redis reservation twice under one `CommitmentID` and every Redis recommendation twice under two `ServiceType` values. Acting on both recommendations buys the same reserved capacity twice. +- evidence: + ```go + func (c *ManagedRedisClient) recommendationsListArgs() (string, *armconsumption.ReservationRecommendationsClientListOptions) { + scope := fmt.Sprintf("/subscriptions/%s", c.subscriptionID) + filter := "properties/scope eq 'Shared' and properties/resourceType eq 'RedisCache'" + return scope, &armconsumption.ReservationRecommendationsClientListOptions{Filter: &filter} + } + ``` +- suggested fix: pick one client as the owner of Azure Cache for Redis and have the other return empty, or split the filter so each covers a disjoint SKU family. +- verdict: CONFIRMED — the recommendation filter is byte-identical at cache:157 and managedredis:41, the reservation classifier is `strings.Contains(strings.ToLower(*props.SKUName), "redis")` at cache:247 and managedredis:220, and both `ServiceCache` and `ServiceMemoryDB` are advertised (provider.go:521-522) and dispatched to the two separate clients (provider.go:586-587 and 590-591). +- issue: (pending cross-reference) + +### A08b-030 The vCPU amount is selected by "not memory", so an accelerator or local-SSD amount is read as the vCPU count +- category: money-path +- severity: high +- location: providers/gcp/services/computeengine/client.go:1296 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `memoryMBFromOperationGroups` selects positively (`isMemoryAmountOp` true), but `vcpuCountFromOperationGroups` selects by exclusion: any commitment operation whose path contains "amount" and whose path filter is not MEMORY is returned as the vCPU count. A recommendation carrying an `ACCELERATOR` or `LOCAL_SSD` amount op — and accelerator families are exactly the ones `familyRequiredResource` exists to describe — returns that amount as `rec.Count`. The count is what `GroupCommitments` sums into the purchased vCPU total, so the resulting CUD commits to the wrong number of vCPUs. Ordering decides which wins, since the first match returns. +- evidence: + ```go + if !strings.Contains(strings.ToLower(op.GetPath()), "amount") { + continue + } + if isMemoryAmountOp(op) { + continue + } + if v := op.GetValue(); v != nil { + ``` +- suggested fix: give the vCPU selector its own positive path-filter test for `VCPU`, mirroring `isMemoryAmountOp`, and error when the payload carries an amount op of an unhandled resource type. +- verdict: CONFIRMED — `vcpuCountFromOperationGroups` (computeengine:1287-1307) accepts any commitment op whose path contains "amount" and is not `isMemoryAmountOp`, returning on the first match, while `isMemoryAmountOp` (1269-1281) and its memory sibling select positively; the resulting `rec.Count` is summed into `agg.vcpus` at GroupCommitments:582 and becomes the purchased VCPU amount at 595. +- issue: (pending cross-reference) + +### A08b-032 An empty idempotency token silently produces a non-idempotent, second-granularity commitment name +- category: money-path +- severity: high +- location: providers/gcp/services/computeengine/client.go:1415 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (consumer at `client.go:787`) +- failure scenario: the whole dedupe mechanism is the name collision — the same token yields the same name and GCP rejects the duplicate. With an empty token the name becomes `cud-`, so a retried purchase a minute later gets a different name and creates a second commitment for the same capacity, with no error. In the other direction, two distinct purchases inside the same second derive the same name and the second fails with ALREADY_EXISTS, reported as a duplicate when it is a genuinely new commitment. Both failures are silent at the call site, which passes `opts.IdempotencyToken` through without checking it is set. +- evidence: + ```go + func idempotentCommitmentName(token string) string { + if token == "" { + return fmt.Sprintf("cud-%d", time.Now().Unix()) + } + ``` +- suggested fix: return an error for an empty token on the purchase path so a caller that forgot to set one fails loudly instead of losing idempotency. +- verdict: CONFIRMED — computeengine:1414-1417 returns `cud-` for an empty token, the call site at 787 passes `opts.IdempotencyToken` through with no check, and the function's own doc at 1409-1410 records that the CLI path reaches the empty branch by design, so both the lost-idempotency and the same-second-collision cases are live. +- issue: (pending cross-reference) + +### A09-005 The USD spend cap is compared against a quote whose currency is never checked +- category: money-path +- severity: high +- location: pkg/exchange/exchange.go:313 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ExchangeQuoteSummary.CurrencyCode` is captured from the AWS response at line 253 and then read by nothing. `MaxPaymentDueUSD` and the field name `PaymentDueUSD` both assert USD, but for an account billed in another currency AWS returns `PaymentDue` in that currency. With `CurrencyCode == "JPY"` and `MaxPaymentDueUSD == 500`, a quote of `PaymentDue = "480"` (≈ $3) passes, and equally a cap of 500 would pass a 500-unit quote worth far more than $500 in a stronger currency. Nothing in the package fails closed on a non-USD quote. +- evidence: + ```go + func checkInitialQuote(q *ExchangeQuoteSummary, maxPayment *big.Rat) error { + if !q.IsValidExchange { + return fmt.Errorf("exchange is not valid: %s", q.ValidationFailureReason) + } + paymentDue := resolvePaymentDue(q) + if paymentDue.Cmp(maxPayment) == 1 { + ``` +- suggested fix: reject any quote whose `CurrencyCode` is not `"USD"` in `checkInitialQuote` and `checkReQuote` before the cap comparison, so non-USD accounts fail loud instead of comparing unlike units. +- verdict: CONFIRMED — `git grep -n CurrencyCode -- 'pkg/exchange/*.go'` over non-test files returns only the field declaration (pkg/exchange/exchange.go:23) and the assignment (line 252); the reshape.go hits are `OfferingOption.CurrencyCode`, a different type. No read of `ExchangeQuoteSummary.CurrencyCode` exists on the quote/cap path. +- issue: (pending cross-reference) + +### A09-008 The per-exchange cap is skipped entirely when the quote reports no PaymentDue +- category: money-path +- severity: high +- location: pkg/exchange/auto.go:319 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the guard is conditioned on `PaymentDueUSD != nil`, so a quote with an empty `PaymentDue` bypasses `perExchangeCap` rather than being rejected. `processRecommendation` then sets `paymentDueStr = "0"` (line 244), the daily-cap arithmetic adds zero, and `chooseEffectiveCap` hands `Execute` a cap that A09-004 also lets through. Every layer of the cap stack collapses on the same nil, so an unpriced quote reaches `AcceptReservedInstancesExchangeQuote` completely uncapped. +- evidence: + ```go + if quote.PaymentDueUSD != nil && quote.PaymentDueUSD.Cmp(perExchangeCap) > 0 { + return nil, &SkippedRecommendation{ + SourceRIID: rec.SourceRIID, + Reason: fmt.Sprintf("exceeds per-exchange cap: payment $%s > cap $%.2f", + ``` +- suggested fix: skip the recommendation with an explicit "quote returned no payment amount" reason when `PaymentDueUSD` is nil, instead of falling through the cap check. +- verdict: CONFIRMED — `getValidatedQuote` guards the per-exchange cap behind `quote.PaymentDueUSD != nil` (pkg/exchange/auto.go:319), so a nil skips the check and returns the quote; `processRecommendation` then sets `paymentDueStr = "0"` (auto.go:243-246) and the daily-cap addition at auto.go:526 adds zero. One correction: "completely uncapped" overstates it — `processAutoExchange` still passes `chooseEffectiveCap`'s value as `MaxPaymentDueUSD` (auto.go:540-548), so the full bypass additionally requires Execute's fresh re-quote to omit PaymentDue as well, which is A09-004's scenario. +- issue: (pending cross-reference) + +### A10-001 Extended-support exclusion cuts Count without scaling the row's money fields +- category: money-path +- severity: high +- location: cmd/multi_service_engine_versions.go:486 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: An RDS rec for `db.r5.large` in `us-east-1` with `Count=10`, `EstimatedSavings=$1000/mo`, `CommitmentCost=$12000` passes filters while 3 running instances of that class sit on an extended-support version. `adjustRecommendationForExcludedVersions` sets `Count=7` but leaves EstimatedSavings at $1000 and CommitmentCost at $12000. The confirmation prompt (`sumPassedRecs`), the CSV report's UpfrontPayment/EstimatedSavings columns and the TOTAL row all describe a 10-instance purchase the run will never make, overstating both spend and benefit by 30%. Every other count-reducing path in the repo (`ApplyInstanceLimit` cmd/helpers.go:189, `ApplyCountOverride` cmd/helpers_count_override.go:56, `applyTargetCoverageRI` pkg/recfilter/sizing.go:303) routes through `common.ScaleRecommendationCosts`; this one does not. +- evidence: + ```go + if excludedCount > 0 { + originalCount := rec.Count + newCount := max(0, rec.Count-excludedCount) + if newCount != originalCount { + log.Printf("📉 Adjusting recommendation ...") + rec.Count = newCount // money fields untouched + } + } + return rec + ``` +- suggested fix: Replace the bare `rec.Count = newCount` with `rec = common.ScaleRecommendationCosts(rec, float64(newCount)/float64(originalCount)); rec.Count = newCount`, guarding `originalCount > 0` exactly as ApplyInstanceLimit does. +- verdict: CONFIRMED — cmd/multi_service_engine_versions.go:479-488 assigns rec.Count with no money scaling while the three sibling count-reducing paths all route through common.ScaleRecommendationCosts (cmd/helpers.go:189, cmd/helpers_count_override.go:56, pkg/recfilter/sizing.go:303), and the adjusted rec flows into the purchase set via cmd/multi_service_filters.go:73. +- issue: (pending cross-reference) + +### A10-002 A failed duplicate check falls back to purchasing the full, un-deduplicated counts +- category: money-path +- severity: high +- location: cmd/multi_service_helpers.go:589 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: With `--purchase`, if `AdjustRecommendationsForExistingRIs` fails (throttling, `DescribeReservedDBInstances` AccessDenied, a transient 5xx) the run logs a warning and proceeds with `filteredRecs` unchanged. The recs still carry the full pre-dedup counts, so the run buys reserved capacity the account already owns. The same fallback exists on the `--input-csv` path at cmd/multi_service.go:546-549 (`adjustedRecs = recs // Continue with original recommendations if check fails`) and inside cmd/multi_service_helpers.go:200-201. This is the one guard standing between a re-run and a double purchase, and it fails open on a non-reversible money path. +- evidence: + ```go + adjustedRecs, dedupedOut, err := duplicateChecker.AdjustRecommendationsForExistingRIs(ctx, filteredRecs, serviceClient) + if err != nil { + AppLogger.Printf(" ⚠️ Warning: Could not check for existing RIs: %v\n", err) + // Continue with original filteredRecs on error; adjustedRecs is not used in this branch. + } else { + ... + filteredRecs = adjustedRecs + } + ``` +- suggested fix: Propagate the error and abort the purchase for that service/region when `!isDryRun` (dry runs may continue with a loud banner), mirroring the "refuse to spend rather than purchase uncapped" stance already taken for `--max-instances` at cmd/multi_service_helpers.go:412. +- verdict: CONFIRMED — checkDuplicates (cmd/multi_service_helpers.go:587-591) is live on the purchase path via cmd/multi_service_helpers.go:666, keeps the pre-dedup counts on error, and executePurchasePipeline (cmd/multi_service.go:391-410) applies no second idempotency guard before executePurchase, so the dedup call really is the only one. +- issue: (pending cross-reference) + +### A10-003 `--target-coverage` sizes as if nothing is owned when the Cost Explorer coverage fetch fails +- category: money-path +- severity: high +- location: cmd/multi_service.go:63 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `--target-coverage 80` on an account already at 75% RI coverage. If `GetRICoverageMap` errors (CE throttling, missing `ce:GetReservationCoverage`), `fetchExistingCoverage` returns nil, every rec keeps `ExistingCoveragePct == 0`, and `applyTargetCoverageRI` computes `gapPct := targetPct - rec.ExistingCoveragePct` = 80 (pkg/recfilter/sizing.go:256). The run sizes to buy 80% coverage on top of the 75% already owned, landing near 155% and paying for idle commitments. The `Recommendation.ExistingCoverageKnown` flag exists precisely to separate "unknown" from "zero" (the CSV writer renders `n/a` vs `0.0`, cmd/multi_service_csv.go:443) but the sizing formula never reads it. +- evidence: + ```go + cov, err := adapter.GetRICoverageMap(ctx, lookbackDays, regions) + if err != nil { + AppLogger.Printf(" ⚠️ Could not fetch existing-RI coverage (%v); sizing will assume zero existing coverage\n", err) + return nil + } + ``` +- suggested fix: When `cfg.TargetCoverage > 0` and the coverage fetch fails, abort a `--purchase` run rather than returning nil; at minimum have `applyTargetCoverageRI` drop any rec whose `ExistingCoverageKnown` is false instead of treating unknown as 0. +- verdict: CONFIRMED — fetchExistingCoverage returns nil on error (cmd/multi_service.go:62-66), AverageInstancesUsedPerHour is still populated by the rec parser (providers/aws/recommendations/parser_ri.go:106) so avg > 0 keeps the rec on the sizing path rather than the no-signal one, and gapPct = targetPct - ExistingCoveragePct at pkg/recfilter/sizing.go:256 never reads ExistingCoverageKnown, whose only reader is the CSV writer at cmd/multi_service_csv.go:443. +- issue: (pending cross-reference) + +### A10-004 A failed account lookup substitutes the account ID, silently defeating `--exclude-accounts` +- category: money-path +- severity: high +- location: cmd/helpers.go:91 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `cudly --purchase --exclude-accounts production`. `organizations:DescribeAccount` fails for account `123456789012` (the CLI is running from a member account, or Organizations is throttling). `GetAccountAlias` caches and returns the numeric ID as the name. `shouldIncludeAccount` then matches `"123456789012"` against the filter `"production"`, finds no substring match, and the account is **not** excluded, so RIs are purchased in the account the operator explicitly ruled out. The bad value is cached for the process lifetime, so one transient failure poisons every rec for that account. The same fallback also silently narrows an `--include-accounts` run in the opposite direction. +- evidence: + ```go + result, err := c.orgClient.DescribeAccount(ctx, &organizations.DescribeAccountInput{ + AccountId: aws.String(accountID), + }) + if err != nil { + c.cache[accountID] = accountID // Use ID as fallback + return accountID + } + ``` +- suggested fix: Have `GetAccountAlias` return `(string, error)`; when an account filter is in force and the alias cannot be resolved, fail the run rather than filtering on a stand-in value. Do not cache the failure. +- verdict: CONFIRMED — cmd/helpers.go:91-94 caches the numeric ID as the alias, and cmd/multi_service_filters.go:105 filters on exactly that name via shouldIncludeAccount (:156-175), so an unresolvable account slips past --exclude-accounts and is dropped by --include-accounts. +- issue: (pending cross-reference) + +### A12-004 Capacity-% scaling is count-based, so every Savings Plans rec is silently dropped below 100% +- category: money-path +- severity: high +- location: frontend/src/recommendations.ts:3919 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: AWS Savings Plans recommendations are created with `Count: 1` (`providers/aws/services/savingsplans/client.go:191`, "Savings Plans don't have a count"). With the Capacity % toolbar set to anything in 1..99, `Math.floor(1 * capacity / 100)` is 0 and the rec hits `continue`, so it vanishes from the purchase with no per-row notice; an SP-only selection aborts with "Try a higher %". The Go sizing code special-cases exactly this: `pkg/recfilter/sizing.go:55` scales `SavingsPlanDetails.HourlyCommitment` for SPs and never touches Count. The frontend also scales `upfront_cost` / `monthly_cost` / `savings` by the count ratio while leaving the SP hourly commitment in `details` untouched, so any SP rec that did survive would display a scaled dollar figure against an unscaled committed quantity. +- evidence: + ```ts + for (const r of recommendations) { + const newCount = Math.floor((r.count * tb.capacity) / 100); + if (newCount <= 0) continue; + const ratio = r.count > 0 ? newCount / r.count : 1; + scaled.push({ ...r, count: newCount, recommended_count: r.count, + upfront_cost: r.upfront_cost * ratio, + monthly_cost: r.monthly_cost != null ? r.monthly_cost * ratio : null, + savings: r.savings * ratio }); + } + ``` +- suggested fix: Branch on `isSavingsPlanService(r.service)` and scale the SP hourly commitment in `details` by `capacity / 100` while leaving `count` at 1, mirroring `ApplyCoverage`'s SP branch. +- verdict: CONFIRMED — SP recs carry Count 1 from the parser (providers/aws/recommendations/parser_sp.go:382) and pkg/recfilter/sizing.go:45-55 scales SavingsPlanDetails.HourlyCommitment rather than Count, while the frontend's count-ratio branch computes floor(1*capacity/100)==0 and `continue`s for every capacity 1..99 (frontend/src/recommendations.ts:3919-3928). +- issue: (pending cross-reference) + +### A12-005 `scheduled` executions render the green "Completed" badge +- category: money-path +- severity: high +- location: frontend/src/history.ts:463 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `historyExecutionStatuses` in `internal/api/handler_history.go:124` deliberately includes `"scheduled"` so the Revoke button is reachable. `statusBadgeHTML` has cases for pending/notified/approved/running/paused/canceled/partial/failed/expired but not `scheduled`, so a delayed purchase whose provider call has not fired yet falls to `default` and renders `badge-success` "Completed". That is exactly what the `isInFlightStatus` comment at lines 77-85 says must never happen. The row then reads "Completed" beside a Revoke button, and the backend's `summarizePurchaseHistory` likewise lets `scheduled` fall into `TotalCompleted` and `TotalUpfront`, so the "Total Upfront Spent" card counts money not yet spent. +- evidence: + ```ts + case 'expired': + return 'Expired'; + default: + return 'Completed'; + ``` +- suggested fix: Add `case 'scheduled':` returning a warning/muted "Scheduled" badge, and include it in `isInFlightStatus` so it buckets under Pending. +- verdict: CONFIRMED — statusBadgeHTML has no `scheduled` case so it hits the green default (frontend/src/history.ts:461-464), `scheduled` is deliberately in historyExecutionStatuses (internal/api/handler_history.go:124), and it also falls through summarizePurchaseHistory's switch into TotalCompleted and TotalUpfront (internal/api/handler_history.go:1042-1084). +- issue: (pending cross-reference) + +### A12-006 Marketplace consent dialog computes the list price with `Math.round` while the backend uses `math.Floor` +- category: money-path +- severity: high +- location: frontend/src/history.ts:1463 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `api.createMarketplaceListing(id)` sends no price; the backend recomputes it with `computeRemainingMonths` = `int(math.Floor(remaining))`, clamped to a minimum of 1 (`internal/api/handler_marketplace.go:447`). The dialog rounds instead, and clamps to a minimum of 0. For a 1-year, $1,200-upfront, count-1 RI purchased 5.4 months ago: remaining is 6.6, so the dialog shows 7 months and a list price of $665 with net proceeds $585, while the backend floors to 6 and lists at $570 for $501.60 net. The user authorises one amount and a different one is submitted, under a comment claiming the two mirror each other "EXACTLY". +- evidence: + ```ts + const remainingMonths = Math.max(0, Math.round(termMonths - elapsedMonths)); + const perUnitResidual = termMonths > 0 && upfront > 0 + ? (upfront * (remainingMonths / termMonths)) / count + : 0; + const listPricePerUnit = perUnitResidual * AWS_MARKETPLACE_BUYER_DISCOUNT; + ``` +- suggested fix: Use `Math.max(1, Math.floor(termMonths - elapsedMonths))` to match `computeRemainingMonths`, or add a backend preview endpoint that returns the resolved price so the dialog cannot drift. +- verdict: CONFIRMED — The dialog uses `Math.max(0, Math.round(...))` (frontend/src/history.ts:1463) while computeRemainingMonths uses `int(math.Floor(remaining))` clamped to 1 (internal/api/handler_marketplace.go:447-459) and the listing recomputes the price server-side from it (internal/api/handler_marketplace.go:157-163), so any fractional remainder >= 0.5 authorises one price and submits another. +- issue: (pending cross-reference) + +### A12-009 "Execute Now" warning shows an upfront total that goes stale when rows are toggled +- category: money-path +- severity: high +- location: frontend/src/recommendations.ts:5013 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `updateExecuteMode` computes `totalUpfront` over `checkedPurchaseIndices` and is bound only to the two radio `change` events. A user who selects "Execute Now" first and then unchecks rows (or uses the header select-all) sees the original amount: the include-checkbox handler at line 5402 calls only `updatePurchaseModalTotals`, which never touches `.direct-execute-warning`. The dialog that says "This will charge $X upfront immediately" therefore states an amount that does not match what the button submits, on the one path that bypasses approval entirely. +- evidence: + ```ts + let totalUpfront = 0; + for (const idx of checkedPurchaseIndices) { + const r = currentPurchaseRecommendations[idx]; + if (r) totalUpfront += r.upfront_cost; + } + // ... includeCb.addEventListener('change', ...) calls only updatePurchaseModalTotals + ``` +- suggested fix: Move the warning-text rebuild into `updatePurchaseModalTotals` (or expose a module-level `refreshDirectExecuteWarning()` that it calls) so the amount tracks the checked set. +- verdict: CONFIRMED — updateExecuteMode is bound only to the two radio change events (frontend/src/recommendations.ts:5041-5042) while the include-checkbox and select-all handlers call updatePurchaseModalTotals alone (frontend/src/recommendations.ts:5409, :5118), and `.direct-execute-warning` is written nowhere else in the file. +- issue: (pending cross-reference) + +### A12-010 Changing the exchange targets after a quote does not invalidate it +- category: money-path +- severity: high +- location: frontend/src/riexchange.ts:1609 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The user quotes 2x `m5.large` (payment due $40), then switches the picker to `m5.4xlarge` and raises Count to 8. The chip and the running total refresh to the new figure, but the quote block, the visible Execute button and `modalQuoteReq` are untouched. Clicking Execute buys the previously quoted 2x `m5.large`, not what the form and running total display. Removing a target row (line 1590) has the same effect, and nothing disables the Quote button, so a double-click fires overlapping quote requests. +- evidence: + ```ts + pickerSelect.addEventListener('change', () => { + offeringInput.value = pickerSelect.value; + updateRowChip(pickerSelect.value, chipEl); + updateRunningTotal(); + }); + countInput.addEventListener('input', updateRunningTotal); + ``` +- suggested fix: In the change / input / remove handlers clear `modalQuote` and `modalQuoteReq`, re-hide `executeBtn`, and clear the quote result so a fresh quote is mandatory. +- verdict: CONFIRMED — The picker `change`, count `input` and row-remove handlers refresh only the chip and running total (frontend/src/riexchange.ts:1608-1614, :1590-1597) and never clear modalQuote/modalQuoteReq or re-hide executeBtn, so submitModalExecute still posts the previously quoted targets (frontend/src/riexchange.ts:1809-1812); the Quote button is likewise never disabled (frontend/src/riexchange.ts:1726-1728). +- issue: (pending cross-reference) + +### A12-074 The Savings Plans group row scales its savings by the cost period twice +- category: money-path +- severity: high +- location: frontend/src/recommendations.ts:3115 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `scaleCost` multiplies by `PERIOD_FACTOR[period]`, and `formatCostForPeriod` calls `scaleCost` again internally (line 1502). The SP group parent row passes an already-scaled value into `formatCostForPeriod`, so the factor is applied twice: with the cost-period selector on "yearly" a $1,000/mo Savings Plans group renders $144,000/yr instead of $12,000/yr, and on "hourly" it renders roughly 1/720th of the correct figure. Every sibling site avoids this by feeding raw monthly values into `formatCostForPeriod` (lines 3154, 3221) or by using `formatScaledRange`, which does not re-scale (lines 3155, 3222, 801). The default period is monthly, whose factor is 1, which is why the bug is invisible until a user changes the selector. +- evidence: + ```ts + const scaledSpSavings = scaleCost(spSavingsTotal, period) ?? spSavingsTotal; + const spSavingsText = `${formatCostForPeriod(scaledSpSavings, period)}${sfxLabel}`; + ``` +- suggested fix: Pass the raw `spSavingsTotal` to `formatCostForPeriod` and drop the outer `scaleCost` call. +- verdict: CONFIRMED — formatCostForPeriod calls scaleCost internally (frontend/src/recommendations.ts:1502), so passing an already-scaled value at frontend/src/recommendations.ts:3114-3115 applies PERIOD_FACTOR twice, whereas the siblings pass raw monthly values (:3152, :3220) or use formatScaledRange, which does not re-scale (:3153, :3221); the only test runs at the default monthly factor of 1 (frontend/src/__tests__/recommendations.test.ts:8113-8128). +- issue: (pending cross-reference) + +### A13-001 `ce:GetCostAndUsage` is granted in no IaC flavor, but the ladder baseline calls it +- category: money-path +- severity: high +- location: terraform/modules/compute/aws/lambda/main.tf:381 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A scheduled ladder run reaches `awsladder.NewFromAWSConfig` -> `baseline.GetUsageBaseline` -> `recommendations.Client.GetOnDemandSeries`, which issues `costexplorer.GetCostAndUsage` (providers/aws/recommendations/ondemand_series.go:182) under the ambient Lambda/Fargate execution role. `ce:GetCostAndUsage` appears in none of the seven IaC files that encode CUDly IAM (`/usr/bin/grep -rn GetCostAndUsage terraform iac cloudformation arm internal/iacfiles` returns nothing). Every ladder run therefore AccessDenies on its usage baseline and is recorded as Errored, producing no plan. `scripts/check-aws-iam-parity.sh` cannot catch this: it compares the templates only against each other, never against the calls the Go code makes, so a uniform omission passes. +- evidence: + ```hcl + Action = [ + "ce:GetReservationUtilization", + "ce:GetReservationPurchaseRecommendation", + "ce:GetReservationCoverage", + "ce:GetSavingsPlansPurchaseRecommendation", + "ce:GetSavingsPlansUtilization", + "ce:GetSavingsPlansCoverage", + ] + ``` +- suggested fix: Add `ce:GetCostAndUsage` to the runtime CE statement in `terraform/modules/compute/aws/{lambda,fargate}/main.tf` and `cloudformation/stacks/CUDly/template.yaml`, and extend `scripts/check-aws-iam-parity.sh` (or a Go guard test) to assert the union of SDK calls in `providers/aws` is covered. +- verdict: CONFIRMED — `/usr/bin/grep -rn 'ce:Get' terraform iac cloudformation arm` returns only the six Reservation/SavingsPlans actions in every flavor (terraform/modules/compute/aws/lambda/main.tf:381-386, fargate/main.tf:373-378, cloudformation/stacks/CUDly/template.yaml:423-428) and `GetCostAndUsage` appears nowhere; the only `ce:*` hit is the permissions-boundary ceiling (policy_boundary.tf:163), which caps but never grants. The path is live, not stubbed: internal/server/app.go:570 assigns `awsladder.NewFromAWSConfig`, which wires `&onDemandSeriesAdapter{client: recoClient}` (providers/aws/ladder/factory.go:85) over ambient `LoadDefaultConfig` credentials (factory.go:57), and handler_ladder.go:306 calls `GetUsageBaseline`, whose only data source is `GetOnDemandSeries` -> `costExplorerClient.GetCostAndUsage` (providers/aws/recommendations/ondemand_series.go:182). +- issue: (pending cross-reference) + +### A13-002 RI Marketplace listing actions are granted nowhere, so the sell flow 403s +- category: money-path +- severity: high +- location: terraform/modules/compute/aws/lambda/main.tf:352 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `POST` to the marketplace-list endpoint runs `internal/api/handler_marketplace.go:248` `ec2Client.CreateMarketplaceListing`, which calls `ec2:CreateReservedInstancesListing` (providers/aws/services/ec2/client.go:1048) against the ambient runtime credentials from `loadAWSConfigWithRegion` -> `getBaseAWSConfig` -> `awsconfig.LoadDefaultConfig`. None of `ec2:CreateReservedInstancesListing`, `ec2:DescribeReservedInstancesListings` or `ec2:CancelReservedInstancesListing` appears in any terraform module, CloudFormation template or federation bundle. Worse, `reserveAndCreateListing` claims the listing slot in the DB before the AWS call, so on the AccessDenied the row transits pending and depends on `releaseMarketplaceClaim` to recover; the cancel path is equally unauthorized, so the operator cannot unwind the listing through CUDly either. +- evidence: + ```hcl + "ec2:DescribeReservedInstances", + "ec2:DescribeReservedInstancesOfferings", + "ec2:GetReservedInstancesExchangeQuote", + "ec2:AcceptReservedInstancesExchangeQuote", + "ec2:PurchaseReservedInstancesOffering", + "ec2:DescribeInstanceTypeOfferings", + "ec2:DescribeRegions", + ``` +- suggested fix: Add the three marketplace actions to the runtime EC2 statement in the Lambda module, the Fargate module and `cloudformation/stacks/CUDly/template.yaml`. +- verdict: CONFIRMED — `/usr/bin/grep -rn ReservedInstancesListing terraform iac cloudformation arm` exits 1 (no hits), while the three SDK calls are declared at providers/aws/services/ec2/client.go:31-33 and issued at client.go:1048, 1070 and 1111. Both routes are registered unconditionally (internal/api/router.go:194-195) and the client is built from ambient credentials: `loadAWSConfigWithRegion` -> `getBaseAWSConfig` -> `awsconfig.LoadDefaultConfig` (internal/api/handler_ri_exchange.go:1259-1275), so the runtime role's `ri_exchange` statement (terraform/modules/compute/aws/lambda/main.tf:349-390) is the effective policy and grants none of the three. +- issue: (pending cross-reference) + +### A01-007 RI exchange approval bypasses the 4-eyes policy +- category: money-path +- severity: medium +- location: internal/api/handler_ri_exchange.go:2195 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GlobalConfig.RequireDifferentApprover=true`. The user who submitted a manual RI exchange (record `CreatedByUserID` = self) holds `approve-own:purchases` and clicks approve. `approveRIExchangeViaSession` -> `fetchAndAuthorizeRIExchange` -> `authorizeSessionApproveRIExchange` explicitly allows the creator, and neither this path nor `approveRIExchangeViaToken` calls `requireDifferentApprover` (callers: handler_purchases.go:600 and :693 only). The irreversible exchange executes with the same person as requester and approver. `TestApproveRIExchange_SessionApproveOwn` pins the self-approval as the expected behaviour. +- evidence: + ```go + record, err := h.fetchAndAuthorizeRIExchange(ctx, session, id) + if err != nil { + return nil, err + } + + // Session-authed approval: stamp the session user as the actor. + transitioned, err := h.config.TransitionRIExchangeStatus(ctx, id, "pending", "processing", resolveCreatorUserID(session)) + ``` +- suggested fix: After `fetchAndAuthorizeRIExchange`, apply the same dual-control check against `record.CreatedByUserID` (generalise `requireDifferentApprover` to take the creator pointer), and fail closed on the token path when the mode is on and no session resolves. +- verdict: CONFIRMED — approveRIExchangeViaSession (internal/api/handler_ri_exchange.go:2176-2223) goes fetchAndAuthorizeRIExchange → authorizeSessionApproveRIExchange:2287-2320, which grants the creator under approve-own; requireDifferentApprover's callers are handler_purchases.go:600,693 only, enforceFourEyesPolicy runs inside purchase.Manager.ApproveAndExecute (internal/purchase/approvals.go:323) which executeApprovedExchange:2430-2487 never calls, and TestApproveRIExchange_SessionApproveOwn (handler_ri_exchange_test.go:422-447) asserts the creator's self-approval succeeds. +- issue: (pending cross-reference) + +### A01-008 Approving never re-evaluates the approver's own permission constraints +- category: money-path +- severity: medium +- location: internal/api/handler_purchases.go:716 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The approving user's group grants `approve-any:purchases` with `MaxPurchaseAmount=5000` and `AccountIDs=[acct-A]`. A different submitter creates a $200k execution for acct-B. `approvePurchaseViaSession` and `approveViaToken` run RBAC, 4-eyes and delay checks, then `ApproveAndExecute`; `requirePermissionConstraints` is hard-wired to the `execute` action (handler.go:475) and is never called on approval, so constraints configured on approve permissions are dead and the approver commits spend their own permission was meant to cap. The retry path was explicitly given this re-check for the identical rationale (lines 1622-1632). +- evidence: + ```go + if globalCfg.GetPurchaseDelay() > 0 { + return h.approveWithDelay(ctx, execution, globalCfg.GetPurchaseDelay(), session.Email, actor) + } + + if err := h.purchase.ApproveAndExecute(ctx, execution.ExecutionID, fourEyesActorIdentity(session), actor); err != nil { + ``` +- suggested fix: In both approve branches call `h.enforcePurchaseConstraints`-style evaluation of the approver's session against `execution.Recommendations` for the approve action (add the action parameter back to `requirePermissionConstraints`), or document and enforce that approve permissions may not carry constraints. +- verdict: CONFIRMED — requirePermissionConstraintsAction is the constant "execute" (internal/api/handler.go:475) and its only purchases-side callers are enforcePurchaseConstraints at handler_purchases.go:1630 (retry) and :2181 (execute); approvePurchaseViaSession:685-716 and approveViaToken:589-619 gate only on HasPermissionAPI, which passes nil constraints (internal/auth/service_api.go:402-404), before ApproveAndExecute. +- issue: (pending cross-reference) + +### A01-010 Azure revoke executes without the user's consented refund amount when the body omits it or Azure returns no amount +- category: money-path +- severity: medium +- location: internal/api/handler_purchases_revoke.go:661 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `POST /api/purchases/{id}/revoke` with an empty body (or `{}`) for an Azure reservation. `revokeConfirmBody.ExpectedRefundAmount` is documented as "Required when the purchase has an Azure revocation window" but nothing enforces it; `callAzureReturn` skips the TOCTOU/consent check whenever `expectedRefundAmount == nil` OR Azure's `BillingRefundAmount` is nil, and submits the Return with whatever `sessionID` CalculateRefund returned (possibly ""). The two-step quote-then-confirm UX is bypassable by any client that just does not send the field. `TestRevokePurchase_AzureSuccess` (handler_purchases_revoke_test.go:247) passes `nil` and asserts success. +- evidence: + ```go + if expectedRefundAmount != nil && calcRefundAmount != nil { + if math.Abs(*expectedRefundAmount-*calcRefundAmount) > revokeQuoteEpsilon { + return nil, NewClientError(422, ...) + } + } + ``` +- suggested fix: Return 400 when `expectedRefundAmount` is nil for the Azure path, 422 when Azure returns no `BillingRefundAmount` or an empty `SessionID`, and keep the divergence check for the remaining case. +- verdict: CONFIRMED — revokeConfirmBody (internal/api/handler_purchases_revoke.go:70-73) documents the field as required but loadAndRevokePurchaseHistory:217-224 only unmarshals it; callAzureReturn:672-678 runs the divergence check only when both pointers are non-nil, azureCalculateRefund:718-735 returns sessionID "" with a nil error when the response omits it, and the Return POST at :687-697 fires regardless; TestRevokePurchase_AzureSuccess (handler_purchases_revoke_test.go:247-275) passes nil with no BillingRefundAmount and asserts "revoked". +- issue: (pending cross-reference) + +### A04-001 Deleting a cloud account silently converts a scoped plan into an ambient-credential plan +- category: money-path +- severity: high +- location: internal/config/store_postgres.go:3299 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (and internal/database/postgres/migrations/000011_cloud_accounts.up.sql:113-117) +- failure scenario: plan P is enabled + auto_purchase and targets exactly one cloud account A (say an Azure subscription); its next_execution_date is next month, so no purchase_executions row references A yet and the 000053 RESTRICT FK does not fire. An admin deletes A. `plan_accounts.account_id` is `ON DELETE CASCADE`, so P's only plan_accounts row vanishes and P keeps running with zero accounts. The notification sweep (`internal/purchase/notifications.go:111-146`) creates a pending execution with nil CloudAccountID, and `executePurchase` (`internal/purchase/execution.go:56-66`) sees `len(accounts) == 0` and falls through to `executeSingleAccount` using CUDly's own ambient credentials. Migration 000060 deleted exactly this "universal plan" shape because it "could trigger unintended purchases"; `DeleteCloudAccount` recreates it with no check. +- evidence: + ```go + tag, err := tx.Exec(ctx, `DELETE FROM cloud_accounts WHERE id = $1`, id) + if err != nil { + return fmt.Errorf("failed to delete cloud account: %w", err) + } + if tag.RowsAffected() == 0 { + return fmt.Errorf("cloud account not found: %s", id) + } + if err = tx.Commit(ctx); err != nil { + ``` + ```sql + CREATE TABLE plan_accounts ( + plan_id UUID NOT NULL REFERENCES purchase_plans(id) ON DELETE CASCADE, + account_id UUID NOT NULL REFERENCES cloud_accounts(id) ON DELETE CASCADE, + ``` +- suggested fix: inside the `DeleteCloudAccount` transaction, refuse (or disable, `enabled = false`) every plan whose only plan_accounts row is the account being deleted, and make `executePurchase` refuse a plan-scoped execution whose plan has zero accounts instead of falling back to ambient credentials. +- verdict: PLAUSIBLE — the unscoped-plan resurrection is confirmed (`plan_accounts.account_id` is ON DELETE CASCADE at internal/database/postgres/migrations/000011_cloud_accounts.up.sql:115, `DeleteCloudAccount` at internal/config/store_postgres.go:3274-3310 checks no plan, and internal/purchase/execution.go:66-68 falls through on `len(accounts) == 0`), but the ambient *purchase* needs an execution carrying recommendations and neither writer populates them — `getOrCreateExecution` (internal/purchase/notifications.go:126-149) nor `createPurchaseExecutionsTx` (internal/api/handler_plans.go:539-547) sets `Recommendations`, so `resolveSingleAccountProvider` short-circuits at internal/purchase/execution.go:367-369 and nothing is bought. +- severity-adjusted: medium — the plan is silently left in the shape migration 000060 deleted, but no purchase can result on the traced path. +- issue: (pending cross-reference) + +### A04-002 Notification sweep's full-row UpdatePurchasePlan can roll a ramp position back and freeze the plan +- category: money-path +- severity: medium +- location: internal/config/store_postgres.go:733-738 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (writer at internal/purchase/notifications.go:100-104) +- failure scenario: `UpdatePurchasePlan` is a blind full-row UPDATE of `ramp_schedule`, `next_execution_date` and `last_execution_date` with no version/CAS check. The scheduled notification sweep reads every plan via `ListPurchasePlans` (current_step=1, next date D1), sends the email, then calls `UpdatePurchasePlan(plan)` only to stamp `last_notification_sent`. If, between that read and that write, a user approval in the API Lambda completes step 2 and `CompletePlanStep` commits current_step=2 / next date D2 under its row lock, the sweep's write lands afterwards and restores current_step=1 and D1. From then on `getOrCreateExecution` keeps finding the already-completed D1 execution, no D2 execution is ever created, and a later completion of step 3 through the API create path is refused by `advanceRampStep` as "step(s) 2-2 never completed": the ramp silently stops buying (the #1669 under-buy shape) while the plan reports step 1. The same overwrite happens from any PUT /api/plans edit that races a completion. This is the #1071 lost-update class reintroduced through a secondary writer; the scheduled-task advisory lock does not serialize against API approvals. +- evidence: + ```go + func (s *PostgresStore) UpdatePurchasePlan(ctx context.Context, plan *PurchasePlan) error { + plan.UpdatedAt = time.Now() + return s.WithTx(ctx, func(tx pgx.Tx) error { + return s.UpdatePurchasePlanTx(ctx, tx, plan) + }) + } + ... + UPDATE purchase_plans SET + name = $2, ... ramp_schedule = $7, updated_at = $8, + next_execution_date = $9, last_execution_date = $10, last_notification_sent = $11 + ``` +- suggested fix: add a narrow `StampPlanNotificationSent(ctx, planID, at)` UPDATE for the sweep, and make `UpdatePurchasePlanTx` optimistic (`WHERE id = $1 AND updated_at = $prevUpdatedAt`, returning a conflict error) so no caller can overwrite a ramp position it did not read under `LockPurchasePlanTx`. +- verdict: CONFIRMED — `UpdatePurchasePlanTx` (internal/config/store_postgres.go:762-786) is a blind full-row UPDATE of `ramp_schedule`/`next_execution_date`/`last_execution_date` with no version predicate, the sweep's write at internal/purchase/notifications.go:100-104 uses a plan read outside `LockPurchasePlanTx` (internal/config/store_postgres.go:667-676, the only lock `CompletePlanStep` holds at 636-654), and `GetExecutionByPlanAndDate` (internal/config/store_postgres.go:1549-1573) has no status filter, so the restored D1 row is returned to `getOrCreateExecution` forever. +- issue: (pending cross-reference) + +### A04-005 GetPendingExecutionsTx duplicate-detection read is a capped page, not a predicate +- category: money-path +- severity: medium +- location: internal/config/store_postgres.go:1493-1508 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (caller internal/api/handler_purchases.go:2457-2463) +- failure scenario: `persistExecutionAndSuppressions` relies on `GetPendingExecutionsTx` to find an existing pending row for the same creator + idempotency key before inserting (#643). The query is `ORDER BY scheduled_date ASC LIMIT 1000 FOR UPDATE`. An estate with more than 1000 pending/notified rows (100 accounts x a 12-step ramp is 1200 rows created up front) silently truncates the newest scheduled rows out of the page; a double-submitted direct purchase scheduled today lands past the cut once older-dated plan rows fill the page, `matchDuplicateInList` finds nothing, and a second execution is inserted and later approved. The FOR UPDATE also row-locks up to 1000 unrelated pending rows on every submit. +- evidence: + ```go + const query = ` + SELECT plan_id, execution_id, status, ... + FROM purchase_executions + WHERE status IN ('pending', 'notified') + AND (expires_at IS NULL OR expires_at > NOW()) + ORDER BY scheduled_date ASC + LIMIT 1000 + FOR UPDATE + ` + ``` +- suggested fix: push the duplicate predicate into SQL (`WHERE status IN (...) AND created_by_user_id = $1 AND idempotency_key = $2 AND created_at >= $3 FOR UPDATE`) so the read is bounded by matching rows, or enforce the dedupe with a partial unique index on `(idempotency_key) WHERE status IN ('pending','notified')`. +- verdict: PLAUSIBLE — the read really is an unfiltered capped page (`ORDER BY scheduled_date ASC LIMIT 1000 FOR UPDATE`, internal/config/store_postgres.go:1502-1507) and `matchDuplicateInList` scans only what that page returned (internal/api/handler_purchases.go:2419-2434, called at 2457-2463), so correctness rests on a LIMIT rather than a predicate; the actual double-insert additionally needs more than 1000 live pending/notified rows scheduled earlier than the resubmit, a data-volume and date-distribution condition I could not establish from source. +- issue: (pending cross-reference) + +### A04-007 ri_exchange_history schema cannot record non-AWS exchanges, so the daily-spend ledger is AWS-only +- category: money-path +- severity: medium +- location: internal/database/postgres/migrations/000009_ri_exchange_history.up.sql:3 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (writer internal/config/store_postgres.go:2639-2651; ledger internal/config/store_postgres.go:2893-2900) +- failure scenario: `account_id VARCHAR(20) NOT NULL CHECK (account_id ~ '^\d{12}$')` was never widened (000095 widened purchase_history, 000067 savings_snapshots, but no migration touched this table). Any `SaveRIExchangeRecord` carrying an Azure subscription GUID fails with SQLSTATE 23514/22001, which is why `executeAzureExchange` (internal/api/handler_ri_exchange.go:1174) performs no store write at all. Consequently `GetRIExchangeDailySpend`, the only input to `RIExchangeMaxDailyUSD`, sums AWS rows only: an operator who set a daily cap has no cap on Azure exchanges, and Azure exchanges leave no audit row for the History page. +- evidence: + ```sql + CREATE TABLE ri_exchange_history ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + account_id VARCHAR(20) NOT NULL CHECK (account_id ~ '^\d{12}$'), + ``` + ```go + SELECT COALESCE(SUM(payment_due), 0)::text + FROM ri_exchange_history + WHERE status IN ('completed', 'processing') + ``` +- suggested fix: migration dropping the `^\d{12}$` CHECK and widening `account_id` to VARCHAR(255) (mirror 000095's guarded probe), then have the Azure execute path write a ledger row and consult `GetRIExchangeDailySpend`. +- verdict: CONFIRMED — the `CHECK (account_id ~ '^\d{12}$')` on VARCHAR(20) still stands at internal/database/postgres/migrations/000009_ri_exchange_history.up.sql:3 and no later migration touching the table (000011, 000026, 000054, 000077, 000080, 000089) alters `account_id`; `executeAzureExchange` (internal/api/handler_ri_exchange.go:1173-1242) performs no store write, and its only money gate `checkAzureExchangeMoneyGuardrails` (internal/api/handler_ri_exchange.go:1110-1135) checks the per-request cap only, never `GetRIExchangeDailySpend` (internal/config/store_postgres.go:2893-2900). +- issue: (pending cross-reference) + +### A04-009 FailRIExchange and CompleteRIExchange have no status guard, so an accepted exchange can drop out of the daily-cap ledger +- category: money-path +- severity: medium +- location: internal/config/store_postgres.go:2865-2882 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (caller internal/api/handler_ri_exchange.go:2350-2355) +- failure scenario: `executeApprovedExchange` transitions pending to processing with a CAS, calls `AcceptReservedInstancesExchangeQuote`, and on any `execErr` (including a client timeout after AWS accepted the quote) calls `failExchange`, which runs `UPDATE ri_exchange_history SET status='failed' WHERE id=$1` unconditionally. `GetRIExchangeDailySpend` sums only `completed`/`processing`, so the exchange that AWS actually executed disappears from the daily ledger and the next approval's cap check under-counts by its `payment_due`. Nothing in the store distinguishes "never submitted" from "outcome unknown after submission". `CompleteRIExchange` / `CompleteRIExchangeWithPayment` are equally unguarded, so a duplicate completion callback can overwrite a `failed` audit row. +- evidence: + ```go + query := ` + UPDATE ri_exchange_history + SET status = 'failed', error = $2 + WHERE id = $1 + ` + ``` +- suggested fix: make the terminal writers CAS on the expected source status (`WHERE id = $1 AND status = 'processing'`) and add a distinct `ambiguous`/`unknown` terminal status that the ledger keeps counting until reconciled. +- verdict: CONFIRMED — `FailRIExchange` is an unconditional `UPDATE ... SET status='failed' WHERE id=$1` (internal/config/store_postgres.go:2863-2882), and `executeApprovedExchange` routes every `execErr` — including the post-`AcceptReservedInstancesExchangeQuote` ambiguous ones — into `failExchange` (internal/api/handler_ri_exchange.go:2465-2467, helper at 2349-2355); `GetRIExchangeDailySpend` sums only `completed`/`processing` (internal/config/store_postgres.go:2893-2900), so the row leaves the ledger. `CompleteRIExchange` (2798-2816) and `CompleteRIExchangeWithPayment` (2822-2840) are equally unguarded. +- issue: (pending cross-reference) + +### A05-014 Confirmation email reports the estimated upfront, purchase_history stores the actual +- category: money-path +- severity: low +- location: internal/purchase/execution.go:660 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The aggregator accumulates `totalUpfront` from `rec.UpfrontCost`, the recommendation's estimate captured at collection time, while `savePurchaseHistory` writes `UpfrontCost: result.Cost`, the figure the provider returned for the commitment it actually created (execution.go:906). When AWS prices the offering differently from the Cost Explorer estimate (a stale rec, a repriced offering), the "Purchase confirmed" email quotes the estimate and the History row quotes the real charge, with no note that they differ. `totalSavings` has the same shape. +- evidence: + ```go + exec.Recommendations[i].Purchased = true + exec.Recommendations[i].PurchaseID = v.purchase.CommitmentID + totalSavings += rec.Savings + totalUpfront += rec.UpfrontCost + ``` +- suggested fix: accumulate `v.purchase.Cost` when the provider returned one and fall back to `rec.UpfrontCost` only when it did not, so the confirmation reports what was charged. +- verdict: CONFIRMED, and worse than described — the divergence is at execution.go:665 (`totalUpfront += rec.UpfrontCost`, fed to sendPurchaseNotification at :117 and into NotificationData.TotalUpfrontCost at :944) versus execution.go:906 (`UpfrontCost: result.Cost`). Only four service clients ever set result.Cost — rds/client.go:184, elasticache/client.go:185, memorydb/client.go:180, redshift/client.go:204 — and `/usr/bin/grep -n Cost providers/aws/services/ec2/client.go` shows the EC2 RI client never assigns it, so every EC2 RI writes UpfrontCost=0 to purchase_history while the email quotes the full estimate. +- severity-adjusted: medium — this is not a rounding difference on repriced offerings; EC2 Reserved Instances, the most common commitment, record $0 upfront in the History table on every purchase. +- issue: (pending cross-reference) + +### A07-002 Savings Plan hourly commitment is rounded to two decimals before purchase +- category: money-path +- severity: medium +- location: providers/aws/services/savingsplans/client.go:232 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `HourlyCommitment` comes from CE's `HourlyCommitmentToPurchase` and is parsed as a full-precision float (parser_sp.go:301). `%.2f` truncates it to cents-per-hour. With `HourlyCommitment = 0.004` the request commits `"0.00"`; with `0.1250001` it commits `"0.13"`, a 4% over-commitment for the whole term. Neither case errors — AWS accepts the rounded string and the recorded purchase silently differs from what was sized and shown to the operator. +- evidence: + ```go + input := &savingsplans.CreateSavingsPlanInput{ + SavingsPlanOfferingId: aws.String(offeringID), + Commitment: aws.String(fmt.Sprintf("%.2f", spDetails.HourlyCommitment)), + UpfrontPaymentAmount: nil, + ``` +- suggested fix: format with the precision AWS accepts (`strconv.FormatFloat(v, 'f', -1, 64)`), and reject a commitment that rounds to zero rather than sending `"0.00"`. +- verdict: CONFIRMED — providers/aws/services/savingsplans/client.go:232 formats `%.2f` while providers/aws/recommendations/parser_sp.go:301 parses `HourlyCommitmentToPurchase` as a full-precision float64 into the same field (:393), so the value sent to `CreateSavingsPlan` differs from the sized one with no error path. +- issue: (pending cross-reference) + +### A07-014 Expiry adjustment divides an org-wide expiring count by a per-rec share of demand +- category: money-path +- severity: medium +- location: providers/aws/recommendations/expiry.go:83 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ApplyCoverageMapToRecommendations` (coverage.go:447) rescales each rec's `AverageInstancesUsedPerHour` so the recs in a pool *sum* to the org-wide average. `expiringByPool` is keyed by the same org-wide pool key and holds the org-wide expiring count. With three linked-account recs sharing one pool, each rec's avg is roughly one third of the pool demand while `expCount` is the full pool's expiring count, so `expiringPct` comes out about three times too large, clamps `ExistingCoveragePct` to zero, and `--target-coverage` sizes each of the three recs as if the pool were entirely uncovered. +- evidence: + ```go + expCount, ok := expiringByPool[lookupPoolKey(recs[i])] + if !ok || expCount == 0 { + continue + } + expiringPct := float64(expCount) / recs[i].AverageInstancesUsedPerHour * 100.0 + ``` +- suggested fix: compute `expiringPct` once per pool against the pool's total average (sum the recs' avgs, or carry the coverage map's `AvgInstancesPerHour`), then apply that single percentage to every rec in the pool. +- verdict: CONFIRMED — `ApplyCoverageMapToRecommendations` (providers/aws/recommendations/coverage.go:447) rescales each rec's avg to its proportional share of `cov.AvgInstancesPerHour`, while `expiringCountsByPool` (expiry.go:59) sums the whole pool's `Count` under the same `lookupPoolKey`, so the ratio at expiry.go:83 grows with the number of recs sharing the pool and the clamp at :84-85 then zeroes `ExistingCoveragePct`. +- issue: (pending cross-reference) + +### A08-002 Five Azure clients return an empty commitment inventory with a nil error when the pager cannot be built +- category: money-path +- severity: high +- location: providers/azure/services/compute/client.go:230 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: with an expired credential or a 403 on `armconsumption.NewReservationsDetailsClient`, `GetExistingCommitments` logs a WARNING and returns `([]common.Commitment{}, nil)`. A caller that uses the commitment inventory to decide whether a reservation already exists sees "no existing commitments" and proceeds. The very next function in the same file states the invariant this breaks: "A partial commitment list is unsafe for the purchase flow — it could trigger duplicate purchases for reservations that exist but weren't loaded." The same shape is in cache/client.go:189, cosmosdb/client.go:191, database/client.go:221 and search/client.go:136. +- evidence: + ```go + pager, err := c.createReservationsPager() + if err != nil { + log.Printf("WARNING: failed to create VM reservations pager: %v", err) + return []common.Commitment{}, nil + } + return c.collectVMReservations(ctx, pager) + ``` +- suggested fix: return the wrapped error from all five clients, matching `collectVMReservations`'s own contract; drop the log-and-swallow branch. +- verdict: PLAUSIBLE — the log-and-swallow really is in all five files (compute/client.go:229, cache/client.go:188, cosmosdb/client.go:190, database/client.go:220, search/client.go:134), but the only production consumer of GetExistingCommitments is the AWS-only CLI dedupe path (cmd/multi_service_helpers.go:652 → pkg/recfilter/dedupe.go:40, whose client comes from the AWS-only cmd/main.go:243 switch), and `armconsumption.NewReservationsDetailsClient` performs no auth call, so the stated "expired credential or 403" trigger is not established from source. +- severity-adjusted: medium — no Azure caller reaches the swallowed branch today, and cmd/multi_service_helpers.go:653 continues with unadjusted recommendations on an error anyway. +- issue: (pending cross-reference) + +### A08-010 Azure exchange purchases force BillingPlan=Upfront regardless of the caller's request +- category: money-path +- severity: high +- location: providers/azure/services/compute/exchange_operations.go:457 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ExchangeTarget` has no billing-plan field and `buildCalculateExchangeRequest` unconditionally sets `ReservationBillingPlanUpfront` on every target. A customer exchanging into a monthly-billed reservation is priced and committed for the full term today. The file's own contract is that nothing is coerced: "There is no coercion of invalid values (no clamping quantity to 1, no defaulting an unrecognized term) — a caller mistake here is a validation error, not a silently different exchange than the one requested" (exchange_operations.go:227). `AppliedScopeType` is optional and threaded through; `BillingPlan` is not. +- evidence: + ```go + Properties: &armreservations.PurchaseRequestProperties{ + AppliedScopeType: to.Ptr(scopeType), + BillingPlan: to.Ptr(armreservations.ReservationBillingPlanUpfront), + BillingScopeID: to.Ptr(tgt.BillingScopeID), + ``` +- suggested fix: add a required `BillingPlan` field on `ExchangeTarget`, validate it against `armreservations.PossibleReservationBillingPlanValues()` in `validateExchangeTargets`, and thread it through. +- verdict: PLAUSIBLE — `ExchangeTarget` (exchange_operations.go:111-137) has no billing-plan field and `buildCalculateExchangeRequest` hardcodes `ReservationBillingPlanUpfront` (exchange_operations.go:457), but the HTTP boundary that builds those targets carries no billing-plan input either (internal/api/handler_ri_exchange.go:662-668), so no caller request is coerced; the impact needs a customer who wants a monthly-billed exchange target. +- severity-adjusted: medium — an upfront-only exchange feature rather than a silent override of a plan the caller asked for. +- issue: (pending cross-reference) + +### A08-013 The terminal-failed reservation-order set uses bare literals and is a denylist that misses BillingFailed +- category: money-path +- severity: medium +- location: providers/azure/services/internal/reservations/purchase.go:449 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `matchReservationOrderInPage` short-circuits a purchase for any existing order whose state is not in this three-entry set. `armreservations.ProvisioningState` (armreservations@v1.1.0/constants.go:369) also defines `BillingFailed`. A prior attempt that reached `BillingFailed` therefore suppresses the re-drive: the recommendation is reported as already purchased and the order ID of a dead order is returned as the commitment ID, so the customer never gets the reservation and the execution is recorded as successful. The three entries are also bare strings, not the SDK constants, so a rename or casing change in a future API version breaks the guard with no compile error. +- evidence: + ```go + var reservationOrderTerminalFailedStates = map[string]struct{}{ + "Cancelled": {}, + "Failed": {}, + "Expired": {}, + } + ``` +- suggested fix: build the set from `armreservations.ProvisioningState*` constants and invert it into an allowlist of states that legitimately suppress a purchase (Succeeded, Created, PendingBilling, ConfirmedBilling, …), so an unknown or newly added state fails safe. +- verdict: CONFIRMED — armreservations@v1.1.0/constants.go:369 defines `ProvisioningStateBillingFailed`, absent from the three bare-string entries at reservations/purchase.go:449-453, so `matchReservationOrderInPage` (purchase.go:525-538) returns a BillingFailed order and `DoIdempotentPurchaseTwoStep` (purchase.go:566-570) short-circuits the re-drive with that dead order's ID. +- issue: (pending cross-reference) + +### A08-017 Advisor recommendations hardcode PaymentOption "upfront" and swallow an unparseable savings amount +- category: money-path +- severity: medium +- location: providers/azure/recommendations.go:427 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `convertAdvisorRecommendation` stamps `PaymentOption: "upfront"` on every Advisor recommendation. Advisor never reports a payment option, and `populateFromExtendedProperties` never overrides it, so an Advisor-sourced recommendation reaches `BillingPlanForPaymentOption` as Upfront and charges the whole commitment immediately. `Term` is defaulted to `"1yr"` and then overwritten by `ext["term"]` verbatim (recommendations.go:444), which Azure reports as "P1Y"/"P3Y", so the same field carries two vocabularies. Separately `extFloat` returns 0 for a missing or unparseable `annualSavingsAmount` and the caller divides it by a bare `12` (line 442), so a malformed Advisor payload becomes a zero-savings recommendation rather than an error. +- evidence: + ```go + rec := &common.Recommendation{ + Provider: common.ProviderAzure, + Service: common.ServiceType(service), + Account: r.subscriptionID, + CommitmentType: common.CommitmentReservedInstance, + Term: "1yr", + PaymentOption: "upfront", + } + ``` +- suggested fix: run the Advisor term through `normaliseTerm` and drop the recommendation (or leave PaymentOption empty so the purchase path refuses it) rather than asserting Upfront; have `extFloat` distinguish absent from zero on the savings field. +- verdict: CONFIRMED — `convertAdvisorRecommendation` (recommendations.go:420-433) hardcodes `PaymentOption: "upfront"` and `populateFromExtendedProperties` (recommendations.go:437-446) never writes that field, overwrites Term with the raw `ext["term"]` without passing it through `normaliseTerm`, and divides `extFloat`'s absent-or-unparseable 0 (recommendations.go:455-462) by a bare 12. +- issue: (pending cross-reference) + +### A08-018 Azure Savings Plan offering details invent a 50/50 split for a payment option Azure cannot express +- category: money-path +- severity: medium +- location: providers/azure/services/savingsplans/client.go:427 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GetOfferingDetails` accepts "partial-upfront" and reports half the total as due today and half spread hourly. Azure savings plans have exactly two billing plans and no partial-upfront — the same package's `BillingPlanForPaymentOption` rejects it explicitly: "partial-upfront is a real payment option elsewhere in CUDly (AWS offers it), so the caller needs to know Azure specifically has no equivalent" (reservations/purchase.go:120). The 0.5 is a hardcoded ratio with no source in the pricing payload, so the UI shows a cost breakdown for a purchase that can never be made. +- evidence: + ```go + case "Partial Upfront", "partial-upfront": + upfrontCost = totalCost * 0.5 + recurringCost = (totalCost * 0.5) / hoursInTerm + ``` +- suggested fix: delete the partial-upfront case so it falls into the existing `default:` error branch, matching `BillingPlanForPaymentOption`. +- verdict: CONFIRMED — savingsplans/client.go:426-428 accepts "Partial Upfront"/"partial-upfront" and splits the total on a hardcoded 0.5 with no pricing-payload source, while `BillingPlanForPaymentOption` rejects the same value as having no Azure equivalent (reservations/purchase.go:119-127). +- issue: (pending cross-reference) + +### A08-019 An unrecognized term is amortized over 12 months +- category: money-path +- severity: medium +- location: providers/azure/internal/recommendations/converter.go:388 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `normaliseTerm` passes an unrecognized Azure term through verbatim with a warning (converter.go:350), so a "P5Y" recommendation keeps `Term == "P5Y"`. `termToMonths("P5Y")` then returns 12, and `ExpandPaymentVariants` sets the monthly variant's `RecurringMonthlyCost` to `CommitmentCost / 12` — five times the real monthly charge for a five-year commitment. A nil or empty term is likewise defaulted to "1yr" (converter.go:342), so a payload that omitted the term is presented as a one-year purchase. +- evidence: + ```go + func termToMonths(term string) int { + switch term { + case "3yr": + return 36 + default: + return 12 + } + } + ``` +- suggested fix: return `(int, error)` and have `ExpandPaymentVariants` leave `RecurringMonthlyCost` nil for a term it cannot map, rather than amortizing over an assumed year. +- verdict: CONFIRMED — `normaliseTerm` (converter.go:341-353) passes an unrecognised term through verbatim after a warning and defaults nil/empty to "1yr", and `termToMonths` (converter.go:388-395) returns 12 for anything but "3yr", which `ExpandPaymentVariants` (converter.go:427, 436) divides CommitmentCost by for the monthly variant. +- issue: (pending cross-reference) + +### A08-023 A GCP SKU with no ServiceRegions matches every region +- category: money-path +- severity: medium +- location: providers/gcp/services/computeengine/client.go:1062 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `skuMatchesMachineType` returns true for any SKU whose `ServiceRegions` is nil, so a global or region-less catalogue entry is accepted as the price for whichever region the client is scoped to. Combined with the last-wins assignment in `extractComputePricingFromSKUs` (client.go:1018-1022), the region-less entry can be the one that ends up as `onDemand` or `commitment`, and `savingsPercentage`, `BreakEvenMonths` and `RecurringMonthlyCost` are all computed from it. The same `nil ServiceRegions` fallthrough is in cloudsql/client.go, memorystore/client.go and cloudstorage/client.go. +- evidence: + ```go + if sku.ServiceRegions != nil { + for _, serviceRegion := range sku.ServiceRegions { + if strings.EqualFold(serviceRegion, region) { return true } + } + return false + } + return true + ``` +- suggested fix: return false when `ServiceRegions` is empty; a SKU that does not declare the region should not price that region. +- verdict: PLAUSIBLE — the fallthrough is real and reaches money (computeengine/client.go:1062, cloudsql:456, cloudstorage:448, memorystore:442; last-wins at client.go:1018-1022 feeds `getComputePricing` client.go:966-981, whose CommitmentPrice/SavingsPercentage become CommitmentCost, BreakEvenMonths and RecurringMonthlyCost at client.go:1160-1175 and 1140), and a `["global"]` SKU would not match either, so only a literal nil triggers it; nothing in the tree establishes that the Cloud Billing catalog ever returns a SKU with nil `ServiceRegions`, and the behaviour is asserted deliberately at client_test.go:243-252. +- issue: (pending cross-reference) + +### A08b-011 Currency is defaulted to USD, then silently overwritten by whichever item came last, and never validated +- category: money-path +- severity: high +- location: providers/azure/services/cache/client.go:600 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same pattern in the five sibling extractors and in `gcp/services/{cloudsql,cloudstorage,memorystore}/client.go`) +- failure scenario: `currency = "USD"` is a fabricated default when the API reports none, and each subsequent item overwrites it, so the reported currency can belong to a different item than the reported price. `GetOfferingDetails` copies it straight into `common.OfferingDetails.Currency` while `TotalCost`/`UpfrontCost` are compared downstream against a USD-denominated `MaxPurchaseAmount` spend cap that has no currency field. A tenant priced in a weaker currency clears a cap it should not, in exactly the shape recorded for PR #1515. +- evidence: + ```go + currency = "USD" + termStr := azureTermString(termYears) + for _, item := range items { + if item.CurrencyCode != "" { + currency = item.CurrencyCode + } + ``` +- suggested fix: return an error when the price items carry no currency or disagree on one, and have the offering path fail closed on any non-USD currency reaching a spend-cap comparison. +- verdict: CONFIRMED — the fabricated `currency = "USD"` seed followed by an unconditional per-item overwrite is at cache:600-606, cosmosdb:598-604, search:511-517, managedredis:521-526, synapse:457-466, and on the GCP side at cloudstorage:388/400-402 with identical copies in cloudsql and memorystore. +- severity-adjusted: medium — the spend-cap half of the scenario is not established: `OfferingDetails.Currency` has no consumer at all (only pkg/common/types.go:444 and pkg/provider/interface.go:50 mention the type outside providers), so nothing carries it to a `MaxPurchaseAmount` comparison. +- issue: (pending cross-reference) + +### A08b-012 `managedredis.GetOfferingDetails` silently bills an unrecognized payment option as all-upfront +- category: money-path +- severity: high +- location: providers/azure/services/managedredis/client.go:355 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: cache, cosmosdb, database, search and synapse all return an error in this `default` branch and their comments cite managedredis as the model; managedredis itself was never fixed. A recommendation carrying `PaymentOption: "partial-upfront"` (or any typo) is quoted with the entire term cost as `UpfrontCost` and `RecurringCost: 0`, so the user is shown, and approves, a one-time charge for a commitment that will actually bill monthly. +- evidence: + ```go + case "monthly", "no-upfront": + upfrontCost = 0 + recurringCost = totalCost / (float64(termYears) * 12) + default: + upfrontCost = totalCost + } + ``` +- suggested fix: return the same explicit error the five sibling clients return for an unsupported payment option. +- verdict: CONFIRMED — managedredis/client.go:355-357 has a bare `default: upfrontCost = totalCost` while the five siblings return an explicit error in the same branch (cache:396-399, cosmosdb:397-400, database:427-430, search:347-350, synapse:361-364), each with the "no silent fallbacks on money-affecting fields" comment. +- severity-adjusted: medium — the branch is only reachable through `GetOfferingDetails`, which has no in-repo caller (pkg/provider/interface.go:50). +- issue: (pending cross-reference) + +### A08b-013 All three GCP service clients silently bill an unrecognized payment option as all-upfront and an unrecognized term as one year +- category: money-path +- severity: high +- location: providers/gcp/services/cloudsql/client.go:280 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cloudstorage/client.go:295`, `memorystore/client.go:291`; term default at `cloudsql/client.go:260`, `cloudstorage/client.go:275`, `memorystore/client.go:271`) +- failure scenario: `termYears` starts at 1 and is only raised for the literals `"3yr"`/`"3"`, so a term of `"36mo"` or `"P3Y"` is priced as a one-year commitment — a third of the real term cost — with no warning. The payment `default` branch then charges the whole (already wrong) total upfront. `computeengine.termPlan` at line 48 errors on exactly this input, so the two paths disagree about whether an unknown term is fatal. +- evidence: + ```go + termYears := 1 + if rec.Term == "3yr" || rec.Term == "3" { + termYears = 3 + } + ``` +- suggested fix: reuse `termPlan`-style parsing that errors on an unrecognized term, and return an error from the payment `default` branch. +- verdict: CONFIRMED — the two-literal term test and the all-upfront `default` are at cloudsql:260-263 and 280-282, cloudstorage:275-278 and 295-297, memorystore:271-274 and 291-293; the divergence with `termPlan` is sharper than stated, since `termPlan` (computeengine:48-57) accepts `"36mo"` as three years while these three price it as one. +- severity-adjusted: medium — reachable only via `GetOfferingDetails`, which has no in-repo caller (pkg/provider/interface.go:50). +- issue: (pending cross-reference) + +### A08b-031 vCPU count falls back to an untyped `numericValue` from the recommendation overview blob +- category: money-path +- severity: medium +- location: providers/gcp/services/computeengine/client.go:1333 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when no VCPU operation is found, the count is taken from `Content.Overview["numericValue"]`, a free-form struct field whose meaning is not contractually the vCPU count — for a cost-oriented recommendation it can be a dollar amount or a percentage. That value becomes `rec.Count` and, through `GroupCommitments`, the number of vCPUs actually committed for one or three years. The doc comment at line 1321 acknowledges the fallback is for "older recommender versions" but nothing checks which version produced the payload. +- evidence: + ```go + if gcpRec.Content.GetOverview() != nil { + if nv := gcpRec.Content.GetOverview().GetFields()["numericValue"]; nv != nil { + if count := nv.GetNumberValue(); count > 0 { + rec.Count = int(count) + } + } + } + ``` +- suggested fix: drop the fallback and return an error when no VCPU operation is present, consistent with `memoryMBFromDetails`, which already refuses to guess the memory amount. +- verdict: CONFIRMED — computeengine:1333-1339 writes `rec.Count` from `Overview["numericValue"]` with no type, unit or recommender-version check, and the doc at 1318-1321 confirms it is a bare fallback; the contrast holds, `memoryMBFromDetails` returns an explicit error rather than guessing (1395). +- issue: (pending cross-reference) + +### A08b-035 Azure Search sends an unverified `reservedResourceType` literal on the purchase path +- category: money-path +- severity: medium +- location: providers/azure/services/search/client.go:262 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: every sibling client sends an `armreservations.ReservedResourceType*` constant; search sends the string `"SearchService"`, which the comment states has no counterpart in the SDK enum at v1.1.0 or v2.0.0 and has not been checked against the live catalog. If Azure's accepted value differs, `calculatePrice` returns a price for a different reserved-resource type or the purchase is rejected mid-flow, after the two-step exchange has already begun. Nothing in the code path degrades gracefully on a wrong enum value. +- evidence: + ```go + // "SearchService" has no counterpart in the armreservations + // ReservedResourceType enum (checked v1.1.0 and v2.0.0), so it + // cannot be expressed as an SDK constant like the other service + // clients do; the literal is kept until verified against the + // live reservation catalog (see issue #1189). + "reservedResourceType": "SearchService", + ``` +- suggested fix: verify the value against the live `reservationOrders/calculatePrice` catalog and, until then, have `PurchaseCommitment` return a not-supported error rather than issuing a request built on an unverified enum. +- verdict: PLAUSIBLE — the bare literal and its self-admitted unverified status are confirmed in the purchase body at search:256-262, and every sibling does send an SDK constant, but whether Azure accepts, rejects or silently reinterprets `"SearchService"` is a live-API fact no reading of this tree can settle. +- issue: (pending cross-reference) + +### A08b-041 Synapse reads on-demand price from a different field and matches the term by substring +- category: money-path +- severity: medium +- location: providers/azure/services/synapse/client.go:456 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: every other Azure client takes on-demand from `item.UnitPrice`; synapse takes it from `item.RetailPrice`. The two fields differ for meters whose unit of measure is a multiple (a "100 Hours" meter reports the block price in one and the per-unit price in the other), so the same catalog produces different savings percentages depending on which client asked. Synapse also re-implements the term string inline and compares with `strings.Contains` rather than equality, so `ReservationTerm: "13 Years"` — or any future term string containing "3 Years" — matches a three-year request. +- evidence: + ```go + switch { + case strings.Contains(item.ReservationTerm, termStr): + if item.RetailPrice > 0 { + reservation = item.RetailPrice + } + case item.Type == "Consumption" && item.RetailPrice > 0: + onDemand = item.RetailPrice + } + ``` +- suggested fix: use the shared `azureTermString` with exact equality and read on-demand from `UnitPrice`, matching the five sibling clients, after confirming which field the API actually documents as the per-unit rate. +- verdict: CONFIRMED — synapse:472-473 reads on-demand from `RetailPrice` where cache:611, cosmosdb:609, search:522 and managedredis:530 all use `UnitPrice`, and synapse:468 matches the term with `strings.Contains` against a term string it rebuilds inline at 458-461 instead of calling the `azureTermString` helper it also declares. +- issue: (pending cross-reference) + +### A08b-043 `skuMatches*` treats a nil region list as "available everywhere" and a substring description match as identity +- category: money-path +- severity: medium +- location: providers/gcp/services/memorystore/client.go:435 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cloudsql/client.go:449`, `cloudstorage/client.go:441`) +- failure scenario: two independent weaknesses in one guard. A SKU whose `ServiceRegions` is nil passes the region test for every region, so an `asia-east1` price can be attributed to a `europe-west4` commitment; an empty but non-nil slice fails for all regions, so the nil/empty distinction silently flips the guard's polarity. The description test is a substring match, so tier `STANDARD` matches every SKU whose description mentions "standard", and Cloud SQL tier `db-n1-standard-1` matches `db-n1-standard-16`. Combined with the last-wins loop in `extractPricingFromSKUs`, the price finally used is whichever loose match came last. +- evidence: + ```go + if !strings.Contains(strings.ToLower(sku.Description), strings.ToLower(tier)) { + return false + } + if sku.ServiceRegions != nil { + for _, serviceRegion := range sku.ServiceRegions { + if strings.EqualFold(serviceRegion, region) { + return true + } + } + ``` +- suggested fix: require an explicit region membership (treat both nil and empty as "not available here" unless the SKU is marked global) and match the tier on the SKU's structured category/description fields rather than a substring. +- verdict: CONFIRMED — the guard falls through to `return true` when `ServiceRegions` is nil and returns false after an empty loop when it is non-nil but empty (memorystore:441-451, cloudsql:455-465, cloudstorage:447-457), so the nil/empty distinction does flip the polarity; the tier test above it is a plain `strings.Contains` on the description, and `extractStoragePricingFromSKUs` (cloudstorage:390-409) keeps the last match rather than rejecting an ambiguous set. +- issue: (pending cross-reference) + +### A09-007 Any ladder mode string other than the literal "manual" auto-executes real exchanges +- category: money-path +- severity: high +- location: pkg/exchange/auto.go:248 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `RIExchangeConfig.Mode` is an untyped `string` and the dispatch compares it to a bare literal, even though `ExchangeMode` typed constants exist 350 lines below at line 606. With `Config.Mode` set to `"Manual"`, `"MANUAL"`, `""`, or any typo, the manual branch is skipped and `processAutoExchange` runs the irreversible `AcceptReservedInstancesExchangeQuote`. The failure direction is the dangerous one: an unrecognized mode buys rather than holds. `internal/api/handler.go:554` documents that the API layer does forward `""`, `"Auto"` and arbitrary typos to the scheduler, so the unvalidated value does reach this comparison. +- evidence: + ```go + if params.Config.Mode == "manual" { + outcome := processManualExchange(ctx, params, rec, offeringID, paymentDueStr) + ``` +- suggested fix: type `RIExchangeConfig.Mode` as `ExchangeMode`, validate it once at the top of `RunAutoExchange`, and dispatch on `ExchangeModeAuto` explicitly so an unknown mode errors instead of falling through to auto-purchase. +- verdict: CONFIRMED — `RIExchangeConfig.Mode` is a bare `string` (pkg/exchange/auto.go:72), the dispatch compares it to the literal `"manual"` (auto.go:248) while the typed constants sit unused at auto.go:606-607, and the fall-through calls `processAutoExchange`. internal/api/handler.go:551-559 states in the code that `GlobalConfig.Validate` never constrains `RIExchangeMode`, so any non-"manual" string reaches this comparison via PUT /api/config. +- severity-adjusted: medium — a compensating gate exists one layer up: `riExchangeArmState.armed()` (internal/api/handler.go:562) is deliberately defined as `mode != "manual"`, so `requireRIExchangeAutoModeGrant` demands the execute:ri-exchange verb before any typo'd mode can be written. The residual risk is an already-authorized operator typing "Manual" and getting auto-execution, not an unprivileged escalation. +- issue: (pending cross-reference) + +### A09-016 A failed pending-exchange cancellation only warns, and the run then creates a second pending set +- category: money-path +- severity: medium +- location: pkg/exchange/auto.go:173 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: a transient database error makes `CancelPendingExchangesByOrigin` fail. The run logs a warning and continues to `processManualExchange`, which issues fresh approval tokens and saves new pending records for the same underutilized RIs. The user's inbox now holds two live approval links per RI: the previous run's (still `pending`, valid for its 24h `ExpiresAt`) and the new one. Approving both executes two exchanges against the same source RI, and the second failure is only caught downstream by AWS. +- evidence: + ```go + cancelled, err := params.Store.CancelPendingExchangesByOrigin(ctx, origin) + if err != nil { + logging.Warnf("failed to cancel pending exchanges: %v", err) + } else if cancelled > 0 { + ``` +- suggested fix: return the error from `RunAutoExchange` when cancellation fails, so a run that cannot clear stale approvals does not issue overlapping ones. +- verdict: CONFIRMED — the cancellation error is logged and swallowed (pkg/exchange/auto.go:171-174) with no early return, so control falls through to the recommendation loop at auto.go:200 and `processManualExchange` mints a fresh `GenerateApprovalToken` and saves a new `pending` record with a 24h `ExpiresAt` (auto.go:344, :371, :383-386). Note the blast radius is bounded: auto.go:467-470 records that AWS replaces the source RI atomically, so a second approval on the same RI fails at the provider rather than double-buying. +- issue: (pending cross-reference) + +### A09-021 A zero Count fabricates a scaling ratio equal to the instance count +- category: money-path +- severity: medium +- location: pkg/recfilter/sizing.go:301 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: for a malformed recommendation with `Count == 0` but non-zero costs — reachable for any parser that populates `AverageInstancesUsedPerHour` and cost fields before setting `Count` — the fallback sets `ratio = float64(nTarget)`. With `nTarget = 3` and `CommitmentCost = 1000`, `ScaleRecommendationCosts` multiplies the commitment to `3000` and `EstimatedSavings` likewise, then `adjusted.Count = 3`. The in-code justification ("so a zero-cost rec stays zero-cost rather than NaN") holds only when every cost field happens to be zero; when it is not, the rec is inflated by an invented factor rather than rejected. +- evidence: + ```go + var ratio float64 + if rec.Count > 0 { + ratio = float64(nTarget) / float64(rec.Count) + } else { + ratio = float64(nTarget) + } + adjusted := common.ScaleRecommendationCosts(rec, ratio) + ``` +- suggested fix: drop the recommendation (with a named drop reason) when `rec.Count <= 0`, since a rec with no count cannot have its per-unit costs derived. +- verdict: PLAUSIBLE — the arithmetic is exactly as described: `ratio = float64(nTarget)` when `rec.Count <= 0` (pkg/recfilter/sizing.go:295-301), then `ScaleRecommendationCosts` multiplies every cost field by that invented factor and `adjusted.Count = nTarget` (sizing.go:302-303). The runtime condition I could not establish from source is a producer: reaching this branch needs `AverageInstancesUsedPerHour > 0` (sizing.go:233) and `gapPct > 0` together with `Count == 0` and non-zero costs, and I found no parser that emits that shape — the in-code comment at sizing.go:288-292 calls it a malformed rec. +- issue: (pending cross-reference) + +### A09-023 The duplicate checker misses AWS RIs in the "queued" state, permitting a double purchase +- category: money-path +- severity: medium +- location: pkg/recfilter/dedupe.go:82 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `Commitment.State` is an untyped string compared against two bare literals. The AWS EC2 `ReservedInstanceState` enum also carries `queued` (a scheduled future RI purchase), which `providers/aws/services/ec2/client.go:102` passes through verbatim as `string(ri.State)`. An RI queued twenty minutes ago has a future `StartDate` (so `StartDate.After(cutoffTime)` is true) but is filtered out by the state check, so `existingMap` never records it and the next run recommends and purchases the same RI again. The literals also carry no link to the SDK enum, so a future state addition repeats the gap. +- evidence: + ```go + func isRecentActiveCommitment(c common.Commitment, cutoffTime time.Time) bool { + return (c.State == "active" || c.State == "payment-pending") && c.StartDate.After(cutoffTime) + } + ``` +- suggested fix: add `"queued"` to the accepted set and define the accepted states as named constants in `pkg/common` next to `Commitment`, so each provider maps its SDK enum onto a typed vocabulary rather than a raw string. +- verdict: CONFIRMED — `isRecentActiveCommitment` compares `c.State` to the two bare literals (pkg/recfilter/dedupe.go:82-84) and `filterRecentCommitments` drops everything else before `buildExistingCommitmentsMap` ever sees it (dedupe.go:72-75), so a queued RI is not recorded and the next run re-recommends it. One correction to the mechanism: the queued RI never reaches dedupe in the first place, because `GetExistingCommitments` filters the AWS call itself with `state IN ("active","payment-pending")` at providers/aws/services/ec2/client.go:78-84. The suggested fix is therefore incomplete — adding "queued" to dedupe.go:83 alone changes nothing until the upstream API filter is widened too. +- issue: (pending cross-reference) + +### A10-008 `determineCSVCoverage` infers "the operator did not set --coverage" from the literal value 80.0 +- category: money-path +- severity: medium +- location: cmd/multi_service_csv.go:21 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `cudly --input-csv recs.csv --coverage 80 --purchase`. The operator explicitly asked for 80% of each row's count. Because the value happens to equal the flag default, `determineCSVCoverage` returns 100.0 and the run buys the full count on every row, 25% more than requested. `validateTargetCoverage` (cmd/validators.go:121) already solves exactly this problem correctly with `cmd.Flags().Changed("coverage")`; this call site compares against the magic literal instead. `TestDetermineCSVCoverage` codifies the current behaviour ("Default coverage (80) changed to 100 for CSV"), so the suite stays green with the bug present. +- evidence: + ```go + func determineCSVCoverage(cfg Config) float64 { + if cfg.Coverage == 80.0 { + // User didn't override the default, so use 100% for CSV mode + return 100.0 + } + return cfg.Coverage + } + ``` +- suggested fix: Thread `cmd.Flags().Changed("coverage")` into Config (or pass the `*cobra.Command`) and branch on that, exactly as `validateTargetCoverage` does. +- verdict: CONFIRMED — cmd/multi_service_csv.go:21 branches on the literal 80.0 while validateTargetCoverage solves the same problem with cmd.Flags().Changed("coverage") at cmd/validators.go:121, and cmd/multi_service_csv_test.go:23-29 pins the buggy mapping as expected behaviour. +- issue: (pending cross-reference) + +### A10-016 Account filters match on substring, so a scoping flag reaches accounts the operator did not name +- category: money-path +- severity: medium +- location: cmd/multi_service_filters.go:205 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `accountMatchesFilter` returns true on any substring hit. In an org with accounts `dev` and `dev-prod-mirror`, `--include-accounts dev --purchase` buys commitments in both. In an org with `prod` and `preprod`, `--exclude-accounts prod` removes both, silently under-buying. There is no anchored or exact-match mode and no warning when a filter matches more than one account, so neither direction is visible in the output. +- evidence: + ```go + func accountMatchesFilter(accountLower, filter string) bool { + filterLower := strings.ToLower(filter) + return filterLower == accountLower || strings.Contains(accountLower, filterLower) + } + ``` +- suggested fix: Default to exact match and require an explicit wildcard (`*dev*`) for substring behaviour; at minimum log every account name a filter expanded to before the confirmation prompt. +- verdict: CONFIRMED — accountMatchesFilter returns strings.Contains (cmd/multi_service_filters.go:203-206) and both list checks call it (:184 and :195), so a filter widens in the include direction and in the exclude direction with no anchored mode and no expansion log. +- issue: (pending cross-reference) + +### A11-005 A saved default_coverage of 0 is rewritten to 80 on the next load-and-save cycle +- category: money-path +- severity: high +- location: frontend/src/settings.ts:3184 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `loadGlobalSettings` populates the target-coverage input with `String(data.global.default_coverage || 80)`. A stored 0 (buy no commitments by default) is falsy, so the input shows 80. `saveGlobalSettings` then reads that input and persists `default_coverage: 80` (settings.ts:3482-3507), so opening Settings and clicking Save silently raises the organisation-wide purchase target from 0% to 80%. The same `||` is applied to `cachedGlobalDefaults.coverage` (settings.ts:3192), which drives the "Inherit (currently: X)" labels in the override modal, and to `notification_days_before || 3` (settings.ts:3196), where 0 means "notify on the day". The file already knows the rule: `populateGraceInput` documents that "An explicit 0 must round-trip as 0" and uses `?? 7` (settings.ts:196-200), and the per-SP-card coverage on line 3265 correctly uses `??`. +- evidence: + ```typescript + // settings.ts:3183-3193 + const coverageInput = document.getElementById('setting-default-coverage') as HTMLInputElement | null; + if (coverageInput) coverageInput.value = String(data.global.default_coverage || 80); + cachedGlobalDefaults = { + term: data.global.default_term || 3, + payment: data.global.default_payment || 'all-upfront', + coverage: data.global.default_coverage || 80, + }; + ``` +- suggested fix: switch the coverage and notification-days reads to `??`, matching `populateGraceInput` and the SP-card path on line 3265. +- verdict: CONFIRMED — `default_coverage` is a plain `json:"default_coverage"` with no omitempty (internal/config/types.go:22), so a stored 0 reaches the form as 0, renders as 80 through the `||` at frontend/src/settings.ts:3184, and is written back as 80 by the save at settings.ts:3482/3507, while the sibling SP-card read on settings.ts:3265 uses `??` and `populateGraceInput` (settings.ts:196-200) documents the opposite rule. The `notification_days_before || 3` half of the claim does not hold: the save validator requires 1..30 (settings.ts:3476), so 0 is not a storable value there. +- severity-adjusted: medium — 0 is not unambiguously "buy nothing" on the backend either (`handler_dashboard.go:394` treats `DefaultCoverage > 0` as the only set value), so this is a silent state rewrite rather than a live purchase-widening. +- issue: (pending cross-reference) + +### A12-014 "Avg Monthly Savings" divides by the number of non-empty buckets, not the number of periods +- category: money-path +- severity: medium +- location: frontend/src/modules/savings-history.ts:289 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `QueryHistory` groups `purchase_history` by `date_trunc(interval, timestamp)` and returns only buckets that contain rows, so a 7-day hourly window with purchases in 3 hours yields `dataPoints.length === 3`. The tile labelled "Avg Monthly Savings" then reads `total / 3` rather than an average over the 168 periods in the window, and the value changes with how sparse the data is rather than with the savings rate. The same figure carries a `/mo` suffix, so the reader cannot tell the denominator is data-dependent. +- evidence: + ```ts + const avgPerPeriod = dataPoints.length > 0 ? totalSavings / dataPoints.length : 0; + ``` +- suggested fix: Divide by the number of buckets in the requested window (derivable from `start`, `end` and `interval`), or relabel the tile to say it averages over periods with activity. +- verdict: CONFIRMED — The query GROUPs by date_trunc and folds only returned rows into the bucket index (internal/api/analytics_postgres.go:174-195), so empty periods produce no data point and `totalSavings / dataPoints.length` (frontend/src/modules/savings-history.ts:289) divides by the count of periods that happened to contain purchases. +- issue: (pending cross-reference) + +### A12-018 The approval-details modal falls back to a no-amount body and still offers Approve +- category: money-path +- severity: medium +- location: frontend/src/approval-details.ts:364 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `buildApprovalDetailsBody` wraps `api.getPurchaseDetails` in a try/catch. On a 403, a 404 or a network blip it returns `buildApprovalDetailsFallback()`, a single sentence with no upfront, no savings, no per-rec table. The confirm dialog still opens with a working "Approve purchase" button (`frontend/src/purchases-deeplink.ts:91`), so the informed-consent guarantee the module exists for silently disappears exactly when the data cannot be read. The same fallback is used when `details.recommendations` is empty (line 64). +- evidence: + ```ts + } catch (err) { + console.error('Failed to load purchase details for approval modal:', err); + return buildApprovalDetailsFallback(); + } + ``` +- suggested fix: Return a marker with the fallback and have `handlePurchaseDeeplink` refuse to render the confirm button when details could not be loaded, pointing the user at History instead. +- verdict: CONFIRMED — Any getPurchaseDetails rejection returns the amount-free single-sentence fallback (frontend/src/approval-details.ts:364-372, :379-384) and handlePurchaseDeeplink passes that body straight into confirmDialog with a live `Approve purchase` button and no guard (frontend/src/purchases-deeplink.ts:85-95). +- issue: (pending cross-reference) + +### A12-020 The Capacity % input silently applies 100% when it holds 0 or is blank +- category: money-path +- severity: medium +- location: frontend/src/recommendations.ts:3712 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The persist handler is `Math.max(1, Math.min(100, parseInt(capacityInput.value, 10) || 100))`. Typing `0` gives `parseInt('0') === 0`, which `||` replaces with 100; clearing the field gives `NaN`, likewise replaced with 100. The input still displays `0` or blank while `handleBulkPurchaseClick` reads the persisted 100 from `loadBulkPurchaseState()`, so the user believes they scoped the buy down and gets a full-size purchase. +- evidence: + ```ts + const persist = (): void => { + saveBulkPurchaseState({ + payment: tbState.payment, + capacity: Math.max(1, Math.min(100, parseInt(capacityInput.value, 10) || 100)), + }); + }; + ``` +- suggested fix: Reject a blank or out-of-range capacity with an inline error and disable the Purchase button, rather than substituting 100. +- verdict: CONFIRMED — The persist handler collapses both `0` and NaN to 100 via `|| 100` (frontend/src/recommendations.ts:3712) while the input keeps displaying the typed value, and handleBulkPurchaseClick reads the persisted state through loadBulkPurchaseState (frontend/src/recommendations.ts:3910). +- issue: (pending cross-reference) + +### A12-035 Marketplace consent copy prints "5.000000000000004%" from a float subtraction +- category: money-path +- severity: medium +- location: frontend/src/history.ts:1518 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `1 - 0.95` is `0.050000000000000044` in IEEE-754, so `(1 - AWS_MARKETPLACE_BUYER_DISCOUNT) * 100` evaluates to `5.000000000000004`. Every user who opens the Sell-on-Marketplace dialog reads "prices the listing at 5.000000000000004% below remaining value" in the sentence that justifies the amount they are about to accept. +- evidence: + ```ts + noteEl.textContent = `AWS charges a ${AWS_MARKETPLACE_FEE_PERCENT}% transaction fee on proceeds. The default schedule prices the listing at ${(1 - AWS_MARKETPLACE_BUYER_DISCOUNT) * 100}% below remaining value. ...`; + ``` +- suggested fix: Introduce `AWS_MARKETPLACE_BUYER_DISCOUNT_PERCENT = 5` and derive the multiplier from it, interpolating the integer. +- verdict: CONFIRMED — AWS_MARKETPLACE_BUYER_DISCOUNT is 0.95 (frontend/src/history.ts:36) and `(1 - 0.95) * 100` evaluates to 5.000000000000004, which is interpolated straight into the consent copy (frontend/src/history.ts:1518). +- issue: (pending cross-reference) + +### A12-038 Cost chips, picker labels and the running total hardcode `$`, discarding `currency_code` +- category: money-path +- severity: medium +- location: frontend/src/riexchange.ts:1390 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `pkg/exchange.OfferingOption` carries `currency_code`, described as the ISO-4217 currency the `EffectiveMonthlyCost` is denominated in, but the frontend `OfferingOption` type omits the field and all four display sites pass a literal `'$'` (lines 1390, 1524, 1632, 1688). A €52.00 offering renders as `$52.00/mo` and the operator compares it directly against the USD figures elsewhere on the page. The exchange-history payment column (line 2216) does the same with a raw `'$' + rec.payment_due`, and an empty `payment_due` renders as a bare `$`. +- evidence: + ```ts + return alternatives + .map((alt) => `${escapeHtml(alt.instance_type)} ${formatCurrency(alt.effective_monthly_cost, '$', 2)}/mo`) + .join(' '); + ``` +- suggested fix: Add `currency_code` to the TypeScript `OfferingOption` and pass it (or its symbol) into `formatCurrency`; render `—` for an empty amount instead of a lone symbol. +- verdict: CONFIRMED — The Go OfferingOption carries `currency_code` as the ISO-4217 denomination (pkg/exchange/reshape.go:70-74) while the TS interface omits it entirely (frontend/src/api/types.ts:730-734) and all four display sites pass a literal '$' (frontend/src/riexchange.ts:1390, :1524, :1632, :1688); the history cell concatenates a bare '$' onto payment_due (frontend/src/riexchange.ts:2215). +- issue: (pending cross-reference) + +### A12-039 `max_payment_due_usd` is filled with the quote amount without checking its currency +- category: money-path +- severity: medium +- location: frontend/src/riexchange.ts:1814 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The client copies `quote.PaymentDueRaw`, an amount denominated in `quote.CurrencyCode` (which the modal displays as a separate row precisely because it is not always USD), into the field named `max_payment_due_usd`. The backend parses it as a USD decimal and uses it both as the `pkg/exchange` cap and as the `MaxPurchaseAmount` for a USD-denominated permission constraint (`handler_ri_exchange.go:1781`). Unlike the Azure path's `checkAzureExchangeMoneyGuardrails`, the AWS path never compares currencies, so a non-USD quote is evaluated against a cap in different units. +- evidence: + ```ts + max_payment_due_usd: modalQuote.PaymentDueRaw, + ``` +- suggested fix: Refuse to enable Execute (or send the currency for the backend to reject) unless `quote.CurrencyCode` is USD. +- verdict: CONFIRMED — The client copies PaymentDueRaw into max_payment_due_usd while displaying CurrencyCode as a separate row (frontend/src/riexchange.ts:1814, :1861-1862), and the AWS handler parses it as a plain decimal and feeds it to the USD-denominated MaxPurchaseAmount constraint with no currency comparison (internal/api/handler_ri_exchange.go:1781-1811), unlike checkAzureExchangeMoneyGuardrails which rejects a currency mismatch 422 (internal/api/handler_ri_exchange.go:1123-1124). +- issue: (pending cross-reference) + +### A12-040 RI exchange Execute has no confirmation step for an irreversible payment +- category: money-path +- severity: medium +- location: frontend/src/riexchange.ts:1730 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A single click on "Execute Exchange" submits the AWS accept call, which the backend's own comment describes as financially irreversible. The far less consequential Approve action in the same file goes through `confirmDialog` (line 2247) and so does enabling auto mode (line 2085). A misclick on the primary button sitting next to "Get Quote" spends money with no interstitial and no restated amount. The Approve dialog itself (line 2247) also carries no amount even though the row has `payment_due`, source/target types and counts. +- evidence: + ```ts + executeBtn.addEventListener('click', () => { + void submitModalExecute(); + }); + ``` +- suggested fix: Gate `submitModalExecute` behind `confirmDialog` with `destructive: true`, quoting `modalQuote.PaymentDueRaw` plus `CurrencyCode` and the targets actually being submitted; add the same amounts to the Approve dialog. +- verdict: CONFIRMED — The Execute click goes straight to submitModalExecute (frontend/src/riexchange.ts:1730-1732) while the sibling Approve action does gate on confirmDialog (frontend/src/riexchange.ts:2247-2251), and that dialog's body carries no payment_due even though the record has one (frontend/src/api/types.ts:841). +- issue: (pending cross-reference) + +### A15-009 Azure retail pricing labels prices with a currency taken from unrelated items +- category: money-path +- severity: medium +- location: providers/azure/services/cache/client.go:600 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: In all seven `extract*Pricing` functions the loop assigns `currency` from every item that carries a non-empty `CurrencyCode`, while `onDemand` and `reservation` are assigned only from items matching the term or `Type == "Consumption"`. The returned currency is therefore whatever the last item in the response happened to carry, not the currency of the two prices actually returned. With a response whose Consumption item is USD and whose later reservation item is EUR, the function returns `onDemand` in USD, `reservation` in EUR, and labels the pair EUR; `getRedisPricing` then computes `savingsPercentage` by mixing the two denominations and reports a fabricated saving. No caller pins `currencyCode` in the request (`params.Add` only ever sets `$filter` and `api-version` across all seven services), so the code depends entirely on the Retail Prices API's undocumented default of USD to keep the item set single-currency. Same shape at compute:766, cosmosdb:598, database:618, managedredis:521, search:511, synapse:457. +- evidence: + ```go + currency = "USD" + termStr := azureTermString(termYears) + + for _, item := range items { + if item.CurrencyCode != "" { + currency = item.CurrencyCode + } + if item.ReservationTerm == termStr { + reservation = item.RetailPrice + } else if item.Type == "Consumption" { + onDemand = item.UnitPrice + } + } + ``` +- suggested fix: Pin `currencyCode=USD` in the request params, and take `currency` only from the items the prices were read from, returning an error when those two disagree rather than labelling a mixed pair. +- verdict: CONFIRMED — I opened all seven sites and the mechanism is identical in each: an unconditional `if item.CurrencyCode != "" { currency = item.CurrencyCode }` that last-writer-wins over every item, while `onDemand`/`reservation` are gated on `ReservationTerm`/`Type` (cache/client.go:604-612, compute:770-778, cosmosdb:602-610, database:622-630, managedredis:523-529, search:515-523, synapse:465-475). The count of seven is right; the only naming slip is that the managedredis copy is `parsePriceItems` (managedredis/client.go:520), not an `extract*Pricing`. The request-side claim holds too: `params.Add` across all seven sets only `$filter` and `api-version` (e.g. cache/client.go:577-578), never `currencyCode`. And the mixing really reaches money: cache/client.go:558 computes `savingsPercentage` from `onDemandPrice` and `reservationPrice` and stamps the loop's `currency` onto the result at :564. +- issue: (pending cross-reference) + +### A06-006 Convertible-RI monthly cost silently becomes $0 when the term duration is missing +- category: money-path +- severity: medium +- location: internal/server/handler_ri_exchange.go:234 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `MonthlyCost` feeds the exchange engine's cross-family dollar-units pre-filter (`exchange.RIInfo.MonthlyCost`). If `DescribeReservedInstances` returns an RI with `Duration == 0` (unset field, an offering type the parser does not populate), this returns `0` instead of an error, so the source RI is presented to the filter as costing nothing per month. A comparison that expects "target must not cost more than source" then admits every candidate target, and in `automatic` mode the exchange is executed with real money. The zero is indistinguishable from a genuine $0 RI. +- evidence: + ```go + func monthlyCostFromConvertibleRI(ri ec2svc.ConvertibleRI) float64 { + if ri.Duration <= 0 { + return 0 + } + hoursPerTerm := float64(ri.Duration) / 3600 + if hoursPerTerm <= 0 { + return 0 + } + ``` +- suggested fix: return `(float64, error)` (or `*float64`) and have `convertForAutoExchange` fail the run for that RI rather than emitting a fabricated $0. +- verdict: PLAUSIBLE — the bare `return 0` is there (internal/server/handler_ri_exchange.go:234-241) and `MonthlyCost <= 0` makes `pricingGatePasses` return true unconditionally, admitting every alternative (pkg/exchange/reshape.go:679-680), but reaching it requires `DescribeReservedInstances` to omit `Duration` (providers/aws/services/ec2/client.go:757 is a nil-safe `aws.ToInt64`), a runtime condition I could not establish from source; the money escalation is also blunted because every auto exchange is still gated on AWS's own `quote.IsValidExchange` and the per-exchange cap (pkg/exchange/auto.go:310-325). +- severity-adjusted: low — the fabricated $0 only widens a local pre-filter that the code documents as a skip-on-zero approximation (pkg/exchange/reshape.go:19-26,144-151); AWS re-validates each exchange before any payment is made. +- issue: (pending cross-reference) + +### A09-020 A short or malformed idempotency token silently drops Azure purchase idempotency +- category: money-path +- severity: medium +- location: pkg/common/tokens.go:115 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `IdempotencyGUID` returns `""` for any token under 32 characters or whose first 32 characters are not hex, and `ReservationOrderID` then hands back `fallback` — the caller's prior non-idempotent, typically random or timestamp-derived, order ID. A truncated or corrupted `IdempotencyToken` (from a shortened DB column, a partial write, or a future caller passing a non-hex value) therefore silently reverts to the non-idempotent path on the exact re-drive scenario the token exists to protect, and a re-driven Azure reservation purchase creates a second reservation. There is no signal at either call site distinguishing "no token supplied" from "token supplied but unusable". +- evidence: + ```go + func ReservationOrderID(token, fallback string) string { + if guid := IdempotencyGUID(token); guid != "" { + return guid + } + return fallback + } + ``` +- suggested fix: return an error (or a distinguishable second return value) when `token != ""` but does not yield a GUID, so a malformed token aborts the purchase instead of downgrading it to non-idempotent. +- verdict: PLAUSIBLE — the fail-open shape is exactly as described (pkg/common/tokens.go:98-107 returns "" for a short or non-hex token; :115-120 then returns `fallback` with no way to tell "no token" from "unusable token"). The runtime condition I could not establish is a malformed token: the only production producer is `DeriveIdempotencyToken` (tokens.go:64-67), which always yields a 64-char sha256 hex string, and it is computed at call time in internal/purchase/execution.go:606 rather than read from a DB column, so the "shortened column / partial write" path does not exist today. Two corrections to the scenario: the sole `ReservationOrderID` caller is providers/azure/services/savingsplans/client.go:282 (a Savings Plans alias name, not a reservation order), and providers/azure/services/internal/reservations/purchase.go:21 records that the reservations path deliberately does not use this mechanism at all. +- severity-adjusted: low — reachable only via a future caller that supplies a non-derived token. +- issue: (pending cross-reference) + +### A09-022 A Savings Plan rec with unexpected Details survives sizing at its full unscaled commitment +- category: money-path +- severity: medium +- location: pkg/recfilter/sizing.go:45 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when `Details` is not a non-nil `*SavingsPlanDetails`, `ApplyCoverage` logs a warning and appends `adjusted` — which is still `rec`, unscaled — to the result. An operator running `--coverage 50` on a rec whose Details failed to decode (see A09-018) gets a plan carrying the full 100% `HourlyCommitment`, and the purchase commits twice the dollars requested. `logf` is nil-safe and callers may pass nil, so the warning can be invisible. The `applyTargetCoverageSP` sibling at line 340 has the same shape. +- evidence: + ```go + if common.IsSavingsPlan(rec.Service) { + if details, ok := rec.Details.(*common.SavingsPlanDetails); ok && details != nil { + adjusted = common.ScaleRecommendationCosts(adjusted, ratio) + } else { + logf.printf("WARNING: SP recommendation for service %q has missing or unexpected Details (%T); passing through unscaled\n", rec.Service, rec.Details) + } + result = append(result, adjusted) + ``` +- suggested fix: drop the recommendation with a named drop reason rather than passing it through, so a sizing flag never silently fails open on the dollar-denominated field it exists to shrink. +- verdict: PLAUSIBLE — the fail-open branch is real: `ApplyCoverage` appends the still-unscaled `adjusted` after logging (pkg/recfilter/sizing.go:45-53), and `applyTargetCoverageSP` returns `(rec, true)` unchanged on the same condition (sizing.go:339-343). What I could not establish is a producer of an SP rec with missing or wrong Details on this path: the AWS SP parser always sets a non-nil `&common.SavingsPlanDetails{}` (providers/aws/recommendations/parser_sp.go:391), and the cited decode failure (A09-018) does not feed here — `ApplyCoverage`'s only callers are the CLI wrappers at cmd/helpers.go:118 and :124, which never route through `DecodeServiceDetailsFor`. +- severity-adjusted: low — no reachable producer today; the branch is defensive. +- issue: (pending cross-reference) + +### A11-006 Per-Savings-Plan coverage is persisted with no range or integer check +- category: money-path +- severity: medium +- location: frontend/src/settings.ts:3562 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the global `setting-default-coverage` is guarded at save time with an explicit `Number.isInteger` + 0..100 check whose comment calls it "the defensive guard for clients that bypass" the inline validator (settings.ts:3469-3488). The four per-plan-type SP coverage inputs go through no such guard: any finite number is accepted, so `-40`, `500` or `12.7` reaches `updateServiceConfig` as the service's target coverage. The inline validator wired at settings.ts:2695-2702 only paints a red hint; it does not block Save. +- evidence: + ```typescript + // settings.ts:3562-3566 + if ('coverageId' in field && field.coverageId) { + const rawCov = byId(field.coverageId)?.value ?? ''; + const parsed = Number(rawCov); + if (rawCov !== '' && Number.isFinite(parsed)) coverage = parsed; + } + ``` +- suggested fix: reuse the same `Number.isInteger(parsed) && parsed >= 0 && parsed <= 100` test the global field uses and abort the save with the same toast when it fails. +- verdict: CONFIRMED — the save path applies no integer or range test (frontend/src/settings.ts:3562-3566), and native validation cannot cover for it because the Purchasing panel's inputs sit outside `#global-settings-form` (index.html:343-423 vs the SP cards at 587-611) and the Save button dispatches a synthetic `new Event('submit')` (app.ts:296-297), which skips constraint validation entirely. The out-of-range half of the claim is blocked one layer down: `ServiceConfig.Validate` rejects coverage outside 0..100 (internal/config/validation.go:517-520) from the update handler at internal/api/handler_config.go:352, so -40 and 500 fail the PUT loudly. +- severity-adjusted: low — only a fractional coverage such as 12.7 survives end to end; the range violations surface as a failed save rather than a bad stored target. +- issue: (pending cross-reference) + +### A12-012 Two different hours-per-month constants convert the same monthly figure to $/hr +- category: money-path +- severity: medium +- location: frontend/src/recommendations.ts:1471 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The Opportunities cost-period selector divides by 720 while the Purchases savings-history unit selector divides by 730 (`frontend/src/modules/savings-history.ts:25`). The same $730/month savings figure therefore reads `$1.0139/hr` on Opportunities and `$1.00/hr` on Purchases, a 1.4% divergence between two screens the user compares directly. The daily factor diverges the same way (`1/30` versus the 30.4375 days/month the marketplace and analytics code use). +- evidence: + ```ts + const PERIOD_FACTOR: Record = { + hourly: 1 / 720, // 24 × 30 hrs/mo + daily: 1 / 30, + monthly: 1, + yearly: 12, + }; + ``` +- suggested fix: Export one `HOURS_PER_MONTH` / `DAYS_PER_MONTH` pair from `utils.ts` and have both modules derive their factors from it. +- verdict: CONFIRMED — frontend/src/recommendations.ts:1471 divides by 720 while frontend/src/modules/savings-history.ts:25 divides by HOURS_PER_MONTH = 730, two live constants for the same monthly-to-hourly conversion on two screens the user compares. +- severity-adjusted: low — display-only divergence of about 1.4 percent; neither factor reaches a submitted amount, a permission cap or a stored record +- issue: (pending cross-reference) + +### A12-044 The plans "Run now" confirmation authorises an immediate purchase without an amount +- category: money-path +- severity: medium +- location: frontend/src/plans.ts:775 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Clicking the row's ▶ button opens a dialog whose entire body is "This will immediately execute the purchase." No upfront cost, no savings, no service, no count, no plan name. `renderApprovalDetailsBody` in `approval-details.ts` exists to solve exactly this for the email deep-link path (issue #374) and is not reused here, so the in-app immediate-execution path has strictly weaker informed consent than the email path. +- evidence: + ```ts + const runOk = await confirmDialog({ + title: 'Run purchase now?', + body: 'This will immediately execute the purchase.', + confirmLabel: 'Run now', + destructive: true, + }); + ``` +- suggested fix: Pass the `PlannedPurchase` row into the handler and render the upfront, monthly savings, count and term in the dialog body. +- verdict: CONFIRMED — The Run-now dialog body is the bare sentence with no cost, count, term or plan name (frontend/src/plans.ts:773-779) while renderApprovalDetailsBody exists for exactly this purpose on the email path (frontend/src/approval-details.ts:53) and is not reused here. +- severity-adjusted: low — the originating row already renders plan name, step, count, term, upfront and monthly savings beside the button (frontend/src/plans.ts:727-741), so the amount is on screen even though the dialog omits it +- issue: (pending cross-reference) + +### A12-058 The three marketplace money lines are rounded independently and do not add up +- category: money-path +- severity: low +- location: frontend/src/history.ts:1502 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `formatCurrency` defaults to zero fraction digits. With `listPriceTotal = 100.5` the dialog shows "Default list price $101", "AWS fee (12%) $12" and "Estimated net proceeds $88": 12 plus 88 is 100, not 101. The three money lines in one consent dialog are mutually inconsistent. +- evidence: + ```ts + addRow('Default list price', count > 1 ? ... : formatCurrency(listPriceTotal)); + addRow(`AWS fee (${AWS_MARKETPLACE_FEE_PERCENT}%)`, formatCurrency(listPriceTotal * (AWS_MARKETPLACE_FEE_PERCENT / 100))); + addRow('Estimated net proceeds', formatCurrency(netProceedsTotal)); + ``` +- suggested fix: Pass `digits: 2` to all three, as the cents-precision summary cards elsewhere already do. +- verdict: CONFIRMED — All three rows call formatCurrency with the default CURRENCY_DEFAULT_DIGITS = 0 (frontend/src/utils.ts:14, frontend/src/history.ts:1500-1504), so list price, fee and net proceeds are rounded independently and need not sum. +- issue: (pending cross-reference) + +### A12-059 With amortize on, the upfront appears both as a lump sum and folded into Monthly Cost +- category: money-path +- severity: low +- location: frontend/src/history.ts:1156 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: For a 1-year row with `upfront_cost` 1200, `monthly_cost` 100 and `estimated_savings` 300, the toggle produces `Upfront Cost $1,200 | Monthly Cost (amortized) $200 | Monthly Savings $300`. Only the Monthly Cost header discloses the amortization, so `Upfront + 12 × Monthly` double-counts the upfront, and cost is shown on an amortized basis in the same row where savings is shown on a cash basis. +- evidence: + ```ts + ${formatCurrency(p.upfront_cost)} + ${monthlyCostCell} + ${formatCurrency(p.estimated_savings)} + ``` +- suggested fix: When amortize is on, mute the Upfront Cell with an "(included in monthly)" title so the two columns cannot be read as additive. +- verdict: CONFIRMED — The row renders the raw upfront cell beside the amortized monthly cell and the cash-basis savings cell (frontend/src/history.ts:1153-1156), and only the column header discloses the amortization (frontend/src/history.ts:1163). +- issue: (pending cross-reference) + +### Category: silent-fallback + +65 findings: 3 high, 35 medium, 27 low. + +### A08-009 GCP offering details silently bill an unrecognized payment option as all-upfront +- category: silent-fallback +- severity: high +- location: providers/gcp/services/computeengine/client.go:876 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: a recommendation with `PaymentOption` of "" (the pre-migration-000032 empty default), "partial-upfront", or any typo falls into `default:` and is reported with the whole commitment charged today and `recurringCost` left at 0. The sibling Azure path errors on exactly this input with the comment "Fail loud on an unrecognised payment option rather than silently billing it as all-upfront (owner policy: no silent fallbacks on money-affecting fields)" (providers/azure/services/compute/client.go:573). All four GCP clients share the bug: cloudsql/client.go:280, memorystore/client.go:291, cloudstorage/client.go:295. +- evidence: + ```go + case "monthly", "no-upfront": + upfrontCost = 0 + recurringCost = totalCost / (float64(termYears) * 12) + default: + upfrontCost = totalCost + } + ``` +- suggested fix: return `fmt.Errorf("unsupported payment option ...")` in the default branch of all four GCP clients, matching the Azure compute client. +- verdict: CONFIRMED — the default branch bills the whole commitment today in all four GCP clients (computeengine/client.go:875, cloudsql/client.go:279, memorystore/client.go:290, cloudstorage/client.go:294) while the Azure sibling returns an error at compute/client.go:574-577, and the empty payment option is a documented reachable state (migration 000032, reservations/purchase.go:103-108). +- issue: (pending cross-reference) + +### A08b-016 GCP pricing failures are logged and the recommendation is emitted with zero cost alongside non-zero savings +- category: silent-fallback +- severity: high +- location: providers/gcp/services/cloudsql/client.go:507 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cloudstorage/client.go:499`, `memorystore/client.go:493`) +- failure scenario: when the catalog lookup fails for any reason — including the guaranteed pagination failure in A08b-015 — `fillSQLPricing` returns without touching the recommendation. `EstimatedSavings` has already been set from the Recommender payload at line 558, so the recommendation reaches the scorer with a real savings number, `CommitmentCost: 0`, `OnDemandCost: 0` and `SavingsPercentage: 0`. Any ranking that divides by cost, or any "savings per dollar committed" screen, sees a free commitment with positive savings and ranks it first. +- evidence: + ```go + pricing, err := c.getSQLPricing(ctx, rec.ResourceType, c.region, termYears) + if err != nil { + log.Printf("cloudsql: pricing unavailable for %s in %s (issue #1020): %v", rec.ResourceType, c.region, err) + return + } + ``` +- suggested fix: drop the recommendation (or propagate the error) when pricing cannot be established, rather than emitting a half-populated money row. +- verdict: CONFIRMED — `fillSQLPricing` logs and returns without writing any cost field (cloudsql:505-510), and `convertGCPRecommendation` has already set `EstimatedSavings` at line 558 before calling it at 566, so the emitted row carries real savings with zero cost; the same shape is at cloudstorage:497-502 and memorystore:491-496. +- issue: (pending cross-reference) + +### A08b-022 Four Azure clients report "no existing commitments" when the reservations pager cannot be built +- category: silent-fallback +- severity: high +- location: providers/azure/services/cache/client.go:185 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cosmosdb/client.go:187`, `database/client.go:217`, `search/client.go:132`) +- failure scenario: a credential, permission or client-construction failure is logged to stdout and converted into an empty slice with a nil error. A caller comparing recommendations against existing commitments concludes the subscription holds no reservations and recommends (or executes) a purchase that duplicates an existing one. `synapse/client.go:168` and `managedredis/client.go:179` return the error on this same path, so the behaviour is inconsistent across the provider. +- evidence: + ```go + pager, err := c.createReservationsPager() + if err != nil { + log.Printf("WARNING: failed to create Redis reservations pager: %v", err) + return []common.Commitment{}, nil + } + ``` +- suggested fix: return the wrapped error, matching synapse and managedredis. +- verdict: CONFIRMED — cache:185-190, cosmosdb:187-192, database:217-222 and search:132-137 log and return `[]common.Commitment{}, nil`, while synapse:165-169 and managedredis:174-180 return the wrapped error on the same path. One narrowing: the discarded error comes only from `NewReservationsDetailsClient` construction (cache:202-205), so a permission failure surfaces later at `NextPage`, not here. +- issue: (pending cross-reference) + +### A01-014 executeApprovedExchange permanently fails an approved exchange on transient store errors and reports HTTP 200 +- category: silent-fallback +- severity: medium +- location: internal/api/handler_ri_exchange.go:2431 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The approver clicks the email link; `TransitionRIExchangeStatus` moves pending->processing, then `GetRIExchangeDailySpend` (or `GetGlobalConfig`) returns a transient DB error. `failExchange` marks the record `failed` for good (the CAS cannot be re-approved) and returns `{"status":"failed"}` with a nil error, i.e. HTTP 200. If `FailRIExchange` itself fails, the row is stranded in `processing` and the caller still gets 200 "failed". On the session path `approveRIExchangeViaSession` then stamps `approved_by` because `execErr == nil` (2213-2216), attributing a failed exchange to the approver. +- evidence: + ```go + dailySpendStr, err := h.config.GetRIExchangeDailySpend(ctx, time.Now()) + if err != nil { + return h.failExchange(ctx, id, "daily spending cap check failed") + } + + globalCfg, err := h.config.GetGlobalConfig(ctx) + if err != nil { + return h.failExchange(ctx, id, "config load failed") + } + ``` +- suggested fix: For pre-commit store failures, revert the record to `pending` (CAS processing->pending) and return a 5xx so the approval can be retried; reserve `failExchange` for provider-side failures, and make it return an error when `FailRIExchange` fails. +- verdict: CONFIRMED — executeApprovedExchange (internal/api/handler_ri_exchange.go:2431-2439) routes GetRIExchangeDailySpend/GetGlobalConfig errors into failExchange:2350-2356, which logs a FailRIExchange failure and returns (map{"status":"failed"}, nil) either way; TransitionRIExchangeStatus (internal/config/store_postgres.go:2751-2755) requires status=pending so a failed row cannot be re-approved, and approveRIExchangeViaSession:2213-2216 stamps approved_by whenever execErr == nil. +- issue: (pending cross-reference) + +### A02-005 History merges a hard-capped, unfiltered page of executions, so older matching rows disappear once 100 non-clean executions exist +- category: silent-fallback +- severity: medium +- location: internal/api/handler_history.go:149 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `fetchExecutionsAsHistory` reads `GetExecutionsByStatuses(..., config.DefaultListLimit)` = 100 rows ordered by `scheduled_date DESC` (store_postgres.go:1280-1283) across every terminal status (failed, expired, canceled, partially_completed) plus future-dated pending ramp steps. Provider/account/date filters and the allowed_accounts scope are applied in Go afterwards. With a few multi-step plans plus accumulated failed/expired rows the cap is exceeded, and a request for `account_ids=&start=2026-06-01` returns none of X's failed executions from June even though they exist; the stale-approval sweep (`expireStaleExecutions`) likewise never reaches pending rows beyond the first 100. No warning is emitted. Same shape as the #1140 truncation fixed for `/api/inventory`. +- evidence: + ```go + executions, err := h.config.GetExecutionsByStatuses(ctx, historyExecutionStatuses, config.DefaultListLimit) + ... + for _rvc := range executions { + exec := executions[_rvc] + ... + if !filters.matchesExecution(exec) { + continue + } + ``` +- suggested fix: Push the provider/account/date predicates (and ideally the scope's account set) into the store query and page by `filters.Limit`, or at minimum drop the fixed cap for the filtered path and query stale pending/notified rows separately for the sweep. +- verdict: CONFIRMED — fetchExecutionsAsHistory (handler_history.go:149) passes config.DefaultListLimit=100 (internal/config/constants.go:9) to GetExecutionsByStatuses, whose SQL applies `ORDER BY scheduled_date DESC LIMIT $2` before any filter (store_postgres.go:1280-1282); historyExecutionStatuses (handler_history.go:124) also includes "completed" rows that are then skipped in Go (line 166), so the page fills even faster than the finding states, and matchesExecution (line 801) plus the stale-sweep collection (line 172) only ever see that page. +- issue: (pending cross-reference) + +### A02-010 Coverage breakdown swallows a recommendations read error and reports 100% coverage +- category: silent-fallback +- severity: medium +- location: internal/api/handler_inventory.go:224 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ListRecommendations` fails (DB blip, or `account_id` is a raw external number that is not a UUID: `buildCoverageRecFilter` at lines 255-258 forwards it unvalidated into `AccountIDs`, which the store compares against a uuid column). `recs` becomes nil, `onDemandByKey` is empty, and `coveragePct(covered, 0)` returns 100 for every service with an active commitment. The page shows "100% covered" with HTTP 200 and nothing in the response marks the number as degraded. +- evidence: + ```go + recs, err := h.scheduler.ListRecommendations(ctx, buildCoverageRecFilter(params)) + if err != nil { + // Non-fatal: recommendations are best-effort for coverage display. + recs = nil + } + ``` +- suggested fix: Return the error (500) like `getDashboardSummary` does at line 48-50, and resolve `account_id` through `resolveSingleAccountFilterIDs` before passing UUIDs to the recommendations filter. +- verdict: CONFIRMED — getCoverageBreakdown (handler_inventory.go:224-229) sets recs=nil on a ListRecommendations error, aggregateOnDemandByKey then yields no entries, and coveragePct (handler_inventory.go:390-396) returns covered/(covered+0)*100 = 100 for every service with an active commitment while the handler still returns 200; buildCoverageRecFilter (lines 255-258) forwards params["account_id"] into AccountIDs unvalidated. +- issue: (pending cross-reference) + +### A03-009 Forgot-password is silently rate-limited for the whole invite window because the age is derived from the wrong expiry constant +- category: silent-fallback +- severity: medium +- location: internal/auth/service_password.go:288 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `tokenAge := PasswordResetExpiry - time.Until(expiry)` assumes every stored token was issued with a 1h expiry. Invites are issued with `PasswordSetupExpiry` (7d), so for an invited user `tokenAge` is negative for 7d - 59m and the request returns nil without issuing a token or sending mail. `CreateUserResult` (service_user.go:150-157) and `sendInviteEmail` tell the admin to "re-mail the setup link via the Forgot Password flow" when the invite send failed; that recovery path does nothing for a week and reports success. Reproduced with an invite 1h old: `UpdateUser calls=0 email sends=0 err=`. +- evidence: + ```go + if user.PasswordResetExpiry != nil { + tokenAge := PasswordResetExpiry - time.Until(*user.PasswordResetExpiry) + if tokenAge < PasswordResetRateLimit { + logging.Debugf("Password reset rate-limited for %s (token age %s < %s)", + redactEmail(email), tokenAge.Round(time.Second), PasswordResetRateLimit) + return nil + } + } + ``` +- suggested fix: Store the issue time (or compute age from `UpdatedAt`), or pick the window by flow (`PasswordSetupExpiry` when `!user.Active`), so the one-minute limit is measured against the real issue time. +- verdict: CONFIRMED — CreateUser sets the invite expiry to now + PasswordSetupExpiry (7d, service_user.go:220-222, service.go:28) while RequestPasswordReset computes age as PasswordResetExpiry (1h, service.go:22) minus time.Until(expiry) at service_password.go:288, which is negative for the first 6d23h of an invite and therefore below PasswordResetRateLimit (1m, service_password.go:45), returning nil before the token write and the send; CreateUserResult's comment at service_user.go:154-156 still directs admins to that flow. +- issue: (pending cross-reference) + +### A03-011 Bastion mode silently falls back to the caller's ambient STS identity in two production callers +- category: silent-fallback +- severity: medium +- location: internal/credentials/resolver.go:222 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: When `opts.AccountLookup` or `opts.STSClientFactory` is nil, a `bastion` account is resolved as plain `role_arn` using the caller-supplied STS client, i.e. the CUDly host's own credentials, and the bastion account's stored credentials are never touched. The no-opts wrapper `ResolveAWSCredentialProvider` is what `internal/scheduler/scheduler.go:761` (recommendation collection) and `internal/api/handler_accounts.go:1528` (org discovery) call, so every bastion account on those paths assumes the target role from the wrong principal. If the customer's trust policy also trusts the CUDly role it "works" with the wrong identity and the wrong audit trail; if it trusts only the bastion, collection fails with an AccessDenied that points nowhere near the cause. +- evidence: + ```go + if opts.AccountLookup == nil || opts.STSClientFactory == nil { + // Legacy fallback: trust caller's stsClient. Tracked in known_issues/03. + return resolveRoleARNProvider(ctx, account, stsClient, nil) + } + ``` +- suggested fix: Return an error for `bastion` when the options are missing, and wire `AWSResolveOptions` into the scheduler and org-discovery callers (the purchase path already does). +- verdict: CONFIRMED — ResolveAWSCredentialProvider passes AWSResolveOptions{} (resolver.go:95-102), resolveBastionProvider:222-225 then assumes the target role with the caller's stsClient and never loads the bastion account, and the only two non-test callers of that wrapper are scheduler.go:761 (s.assumeRoleSTS) and handler_accounts.go:1528 (sts.NewFromConfig(baseCfg)), both host-credentialed; the WithOpts callers at server/app.go:825 and purchase/execution.go:449 set only AmbientProvider, so no production path supplies AccountLookup/STSClientFactory at all. +- issue: (pending cross-reference) + +### A06-008 AWS sender is constructed with no FROM_EMAIL, so password-reset / invite / welcome mail silently succeeds and is never sent +- category: silent-fallback +- severity: medium +- location: internal/email/factory.go:76 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the AWS branch validates nothing (the SMTP constructors, by contrast, reject an empty `FromEmail` at internal/email/smtp_sender.go:73). With `FROM_EMAIL` unset and `SNS_TOPIC_ARN` unset, a user clicks "forgot password": `SendPasswordResetEmail` → `sendMultipartVia` → `SendToEmailWithCCMultipart` → `SendToEmailWithCC`, which logs at debug level and returns `nil` (internal/email/sender.go:306). The API reports success, the user is told to check their inbox, and nothing was ever sent. The same silent nil covers welcome, invite and scheduled-purchase notifications; only `SendPurchaseApprovalRequest` and `SendPurchaseExecutedNotification` check `isValidFromEmail` and return `ErrNoFromEmail`. +- evidence: + ```go + case ProviderAWS: + return NewSender(SenderConfig{ + TopicARN: os.Getenv("SNS_TOPIC_ARN"), + FromEmail: os.Getenv("FROM_EMAIL"), + EmailAddress: os.Getenv("EMAIL_ADDRESS"), + }) + ``` +- suggested fix: apply the existing `isValidFromEmail` guard in `SendToEmailWithCC`/`SendToEmailWithCCMultipart` and return `ErrNoFromEmail` instead of `nil`. +- verdict: CONFIRMED — the AWS branch validates nothing (internal/email/factory.go:75-81) while `NewSMTPSender` rejects an empty `FromEmail` (internal/email/smtp_sender.go:73-75); `SendPasswordResetEmail` → `sendMultipartVia` → `SendToEmailWithCCMultipart` (internal/email/templates.go:481-486,780-796) and both SES entry points return a bare `nil` on `s.fromEmail == ""` (internal/email/sender.go:271-274 and 306-309), whereas only `SendPurchaseApprovalRequest` and `SendPurchaseExecutedNotification` apply `isValidFromEmail` (templates.go:819-820,1070-1071). +- issue: (pending cross-reference) + +### A06-014 SMTP transport redirects token-bearing approval mail to the static notify inbox where SES returns ErrNoRecipient +- category: silent-fallback +- severity: medium +- location: internal/email/smtp_sender.go:527 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `(*Sender).SendPurchaseApprovalRequest` returns `ErrNoRecipient` when `data.RecipientEmail == ""`, and internal/api/handler_purchases.go:2889 branches on that sentinel to tell the operator which side is unconfigured. The SMTP implementations of `SendPurchaseApprovalRequest` (:527), `SendPurchaseExecutedNotification` (:634) and `SendRIExchangePendingApproval` (:485) instead fall back to `s.notifyEmail`, so on every Azure/GCP deployment an unresolved recipient is masked: the approval link and its one-time token are delivered to a shared operations inbox, and the handler's precise error branch is unreachable. The two transports therefore behave differently on the same money path. +- evidence: + ```go + recipient := data.RecipientEmail + if recipient == "" { + recipient = s.notifyEmail + } + if recipient == "" { + return ErrNoRecipient + } + ``` +- suggested fix: return `ErrNoRecipient` when `data.RecipientEmail` is empty on the token-bearing methods, matching the SES path; keep the `notifyEmail` fallback only for the tokenless broadcast-style sends. +- verdict: CONFIRMED — the three SMTP methods fall back to `s.notifyEmail` (internal/email/smtp_sender.go:484-490, 526-533, 633-640) while the SES equivalents return `ErrNoRecipient` on the empty field before anything else (internal/email/templates.go:810-812, 1066-1068), and internal/api/handler_purchases.go:2886-2893 branches on that sentinel. Stronger than stated: `notifyEmail` defaults to `cfg.FromEmail` when unset (smtp_sender.go:85-88) and `FromEmail` is mandatory (smtp_sender.go:73-75), so on GCP/Azure the fallback always resolves and `ErrNoRecipient` is unreachable from these methods. +- issue: (pending cross-reference) + +### A06-015 Analytics pipeline keeps writing after partition creation fails, into a partition retention can never drop +- category: silent-fallback +- severity: medium +- location: internal/server/analytics_collect.go:142 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: step 1 exists specifically so the snapshot "lands in a real monthly partition rather than the catch-all default (M3)". When `CreateFuturePartitions` fails (lock timeout, the 5-minute DDL deadline, missing privilege) the code logs a warning, marks the result "partial", and proceeds to `Collect`, which COPYs the batch into the default partition. `DropOldPartitions` only detaches month partitions, so those rows are never reclaimed and the default partition grows forever; every later month with the same failure adds more. The scheduled task still returns HTTP 200 with `status: partial`, which no alert consumes. +- evidence: + ```go + if err := withDDLTimeout(ctx, app.Analytics.CreateFuturePartitions, cfg.PartitionsAhead); err != nil { + log.Printf("Warning: failed to ensure future partitions: %v", err) + result["status"] = "partial" + } else { + result["partitions_ensured"] = true + } + ``` +- suggested fix: return the error and skip the collect step when partition creation fails, since the write is the thing the partition guarantees. +- verdict: CONFIRMED — the partition step only logs and marks `partial` before falling through to `Collect` (internal/server/analytics_collect.go:140-159), and retention explicitly excludes the catch-all: `drop_old_savings_partitions` filters `tablename != 'savings_snapshots_default'` (internal/database/postgres/migrations/000067_analytics_snapshot_correctness.up.sql:207-212), so rows that land there are never reclaimed. The task still returns `(result, nil)` and the HTTP wrapper reports `"status":"success"` around it (internal/server/http.go:257-265). +- issue: (pending cross-reference) + +### A06-016 DEFAULT_TERM and DEFAULT_COVERAGE silently fall back to built-in defaults on an unparseable value +- category: silent-fallback +- severity: medium +- location: internal/server/app.go:947 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `LoadApplicationConfig` reads `DEFAULT_TERM` and `DEFAULT_COVERAGE` through these helpers and both feed the purchase manager as system-wide money defaults (`purchase.ManagerConfig.DefaultTerm` / `DefaultCoverage`, app.go:441-443). A typo such as `DEFAULT_TERM=3y` logs one warning line and silently commits every subsequent purchase at 3 years instead of the intended value. The two sibling knobs on the same struct are validated and fail startup (`validateAppConfigEnvDefaults` at app.go:288, added for issue #1026), and the analytics knobs deliberately use a non-defaulting parser for the same reason (`loadAnalyticsInt`, analytics_collect.go:60), so the term/coverage pair are the outliers. +- evidence: + ```go + result, err := strconv.Atoi(val) + if err != nil { + log.Printf("WARNING: %s=%q is not a valid integer; using default %d", key, val, defaultVal) + return defaultVal + } + ``` +- suggested fix: extend `validateAppConfigEnvDefaults` to reject a set-but-unparseable `DEFAULT_TERM` / `DEFAULT_COVERAGE` at startup, using the `loadAnalyticsInt` sentinel pattern. +- verdict: CONFIRMED — `getEnvInt` and `getEnvFloat` both log a warning and return the built-in default (internal/server/app.go:945-955, 962-972); they are the readers for `DEFAULT_TERM` (default 3) and `DEFAULT_COVERAGE` (default 80) at app.go:263,265, both of which flow into `purchase.ManagerConfig` at app.go:441-443 and again at app.go:790-792. `validateAppConfigEnvDefaults` covers only `DEFAULT_PAYMENT_OPTION` and `DEFAULT_RAMP_SCHEDULE` (app.go:288-297), and `loadAnalyticsInt` uses the fail-fast sentinel instead (internal/server/analytics_collect.go:60-70), so the term/coverage pair are indeed the outliers. +- issue: (pending cross-reference) + +### A06-028 Azure federated-credential creation swallows every error and still reports success +- category: silent-fallback +- severity: medium +- location: internal/iacfiles/templates/azure-wif-deploy.sh.tmpl:79 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `|| echo "(federated credential may already exist — continuing)"` catches every failure of `az ad app federated-credential create`, not just the name conflict it names. If a credential named `cudly` already exists on the app registration bound to a *different* issuer, subject or audience — an earlier deployment pointing at a different CUDly instance, or a hand-edited one — the create fails, the script prints "continuing", proceeds to assign the Reservation Purchaser role at subscription scope (step 4), and prints "=== Done ===". The operator is told federation is configured while the credential actually in place trusts a different issuer, so either CUDly cannot authenticate or a stale issuer can. The GCP script goes to considerable length to avoid exactly this (gcp-wif-cli.sh.tmpl:177-233 re-reads an existing provider and aborts unless issuer, attribute condition, mapping, state and audience all match). The same swallow is in azure-wif-cli.sh.tmpl:60. +- evidence: + ```sh + az ad app federated-credential create \ + --id "${APP_ID}" \ + --parameters "${FEDCRED_PARAMS}" \ + --output none 2>&1 || echo " (federated credential may already exist — continuing)" + ``` +- suggested fix: on failure, `az ad app federated-credential show --federated-credential-id cudly` and compare issuer/subject/audiences against the intended values; abort when they differ instead of continuing to the role assignment. +- verdict: CONFIRMED — the `|| echo` swallows every non-zero exit and defeats the script's own `set -euo pipefail` (internal/iacfiles/templates/azure-wif-deploy.sh.tmpl:14, 77-81), after which step 4 assigns the Reservation Purchaser role and the script prints `=== Done ===` (azure-wif-deploy.sh.tmpl:83-97); the same swallow is at azure-wif-cli.sh.tmpl:58-61, and the GCP path does the compare-and-abort the finding cites (gcp-wif-cli.sh.tmpl). One correction: each run creates a fresh app registration (`az ad app create`, azure-wif-deploy.sh.tmpl:49), so the "credential bound to a different issuer" variant is hard to reach; a Graph permission denial or a throttle produces the same false success with no credential in place at all. +- issue: (pending cross-reference) + +### A07-005 Savings Plan start/end dates are dropped when unparseable +- category: silent-fallback +- severity: medium +- location: providers/aws/services/savingsplans/client.go:195 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: a `Start`/`End` string that does not parse as RFC3339 leaves `commitment.StartDate` / `EndDate` at the zero time with no log line. `recfilter.isRecentActiveCommitment` gates on `c.StartDate.After(cutoff)`, so a zero `StartDate` makes an active SP invisible to the duplicate check and the next run re-recommends it. `ladder/adapters.go:195` fails loud on exactly this input for the same field, so the two readers of the same API disagree. +- evidence: + ```go + if sp.Start != nil { + if startTime, err := time.Parse(time.RFC3339, *sp.Start); err == nil { + commitment.StartDate = startTime + } + } + ``` +- suggested fix: return an error (or at minimum log) on a parse failure, matching `parseSPDate` in providers/aws/ladder/adapters.go. +- verdict: CONFIRMED — providers/aws/services/savingsplans/client.go:195-204 drops both parse failures with no log; pkg/recfilter/dedupe.go:83 gates on `c.StartDate.After(cutoffTime)` so a zero StartDate makes the SP invisible to the dedupe map built at :105, while providers/aws/ladder/adapters.go:195-204 errors on the identical field. +- issue: (pending cross-reference) + +### A07-008 MemoryDB recommendations hardcode the engine as "redis" +- category: silent-fallback +- severity: medium +- location: providers/aws/recommendations/parser_services.go:255 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: MemoryDB supports Valkey as well as Redis. Every MemoryDB recommendation is stamped `Engine: "redis"` regardless of what the cluster runs, and the same function errors loudly when `NodeType` is absent rather than guessing — the engine gets the opposite treatment. The fabricated value flows into `EngineFromDetails` and the dedupe key, and into any consumer that displays or filters on engine. +- evidence: + ```go + rec.Details = &common.CacheDetails{ + Engine: "redis", + NodeType: *mdbDetails.NodeType, + } + ``` +- suggested fix: read the engine from `MemoryDBInstanceDetails` if CE exposes it, otherwise leave it empty so downstream code can tell "unknown" from "redis". +- verdict: CONFIRMED — providers/aws/recommendations/parser_services.go:255 stamps `"redis"` unconditionally while :248-250 errors on a missing NodeType, and costexplorer@v1.63.1/types/types.go:1492-1509 shows `MemoryDBInstanceDetails` carries no engine field at all, so the value is fabricated rather than read; only the "leave it empty" half of the suggested fix is available. +- issue: (pending cross-reference) + +### A07-009 Reserved-node term is silently assumed to be 12 months for any non-3yr duration +- category: silent-fallback +- severity: medium +- location: providers/aws/services/rds/client.go:88 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `instance.Duration` is a `*int32`; a nil pointer (or any value that is not exactly 94608000) yields `termMonths = 12`. `EndDate` is then `StartTime + 12 months`. A 3-year RI whose `Duration` came back nil is reported as expiring two years early, which makes `AdjustExistingCoverageForExpiringCommitments` treat its pool as uncovered and recommend a replacement purchase that is not needed. ElastiCache has the identical construct at elasticache/client.go:86, and MemoryDB/OpenSearch/Redshift use `getTermMonthsFromDuration`, which likewise returns 12 for a zero duration. +- evidence: + ```go + duration := aws.ToInt32(instance.Duration) + termMonths := 12 + if duration == ThreeYearSeconds { + termMonths = 36 + } + ``` +- suggested fix: treat an absent or unrecognised duration as unknown — leave `EndDate` zero (which `expiringCountsByPool` already skips) rather than fabricating a one-year term. +- verdict: CONFIRMED — providers/aws/services/rds/client.go:87-91 and providers/aws/services/elasticache/client.go:86-89 map every duration other than `ThreeYearSeconds` (nil included, via `aws.ToInt32`) to 12 months and set `EndDate = StartTime + termMonths`; `getTermMonthsFromDuration` does the same for a zero duration at memorydb/client.go:498, opensearch/client.go:617 and redshift/client.go:684, and providers/aws/recommendations/expiry.go:52 does skip a zero `EndDate` as the fix assumes. +- issue: (pending cross-reference) + +### A07-022 Exchange target lookup swallows an invalid payment option and silently widens the query +- category: silent-fallback +- severity: medium +- location: providers/aws/services/ec2/client.go:874 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `normalizeTargetOfferingsParams` calls `convertEC2PaymentOption` and discards the error with `if ... err == nil`. A caller passing an unrecognised payment option (for example the display form `"All Upfront"` rather than `"all-upfront"`) leaves `offeringType` at its zero value, which the SDK omits, so AWS returns offerings for *every* payment option. `ListTargetOfferings` then presents partial-upfront and all-upfront exchange targets for a request that asked for one specific option, and the caller has no signal that its input was rejected. +- evidence: + ```go + if p.OfferingType != "" { + if ot, err := convertEC2PaymentOption(p.OfferingType); err == nil { + offeringType = ot + } + } + return + ``` +- suggested fix: return the error from `normalizeTargetOfferingsParams` and propagate it out of `ListTargetOfferings`; keep the "all options" behaviour only for a genuinely empty input. +- verdict: CONFIRMED — providers/aws/services/ec2/client.go:874-878 keeps `offeringType` at its zero value when `convertEC2PaymentOption` errors and returns no error to the caller, and the path is live from internal/api/handler_ri_exchange.go:154. +- issue: (pending cross-reference) + +### A07-026 Mid-pagination AccessDenied returns a partial account list as if it were complete +- category: silent-fallback +- severity: medium +- location: providers/aws/provider.go:292 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the comment above `orgListAccountsSilentErrorCodes` states that returning a silently truncated list is unsafe for the purchase flow, but the silent-code check is not gated on how many pages have already succeeded. If page one of `ListAccounts` succeeds and page two fails with `AccessDeniedException` (an SCP change, a permission boundary applied mid-run), `appendOrgAccounts` returns the accounts collected so far with a nil error and the caller cannot distinguish it from a complete organization listing. +- evidence: + ```go + var apiErr smithy.APIError + if errors.As(err, &apiErr) { + if _, silent := orgListAccountsSilentErrorCodes[apiErr.ErrorCode()]; silent { + return accounts, nil + } + } + ``` +- suggested fix: treat the silent codes as expected only on the first page (no member accounts collected yet); once pagination has produced results, any error including AccessDenied is a real mid-run failure and must propagate. +- verdict: CONFIRMED — providers/aws/provider.go:287-293 tests only the error code and returns `accounts, nil` regardless of how many pages already appended at :298-308, directly contradicting the "returning a silently-truncated list is unsafe for the purchase flow" rationale stated at :255-261; the non-silent branch immediately below (:295-297) shows the intended propagation. +- issue: (pending cross-reference) + +### A08-014 pricing.FetchAll silently truncates when the page cap is hit +- category: silent-fallback +- severity: medium +- location: providers/azure/internal/pricing/retail_prices.go:72 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when a query needs more than `maxPages` pages the loop exits with a non-empty `nextURL` and returns `(all, nil)`. The caller (`getVMPricing`, compute/client.go:700) cannot distinguish "all prices for this SKU" from "the first 5000 of them", so a reservation or on-demand meter on a later page is read as missing and the run either errors with a wrong message or, where the meter type differs, prices against a partial set. The self-referential-link case in the same loop returns an explicit error; the cap does not. +- evidence: + ```go + for pageIdx := 0; pageIdx < maxPages && nextURL != ""; pageIdx++ { + ... + nextURL = page.NextPageLink + } + return all, nil + ``` +- suggested fix: after the loop, return an error when `nextURL != ""`, naming the cap — same treatment the self-referential-link guard already gets. +- verdict: CONFIRMED — the loop at internal/pricing/retail_prices.go:72-86 exits on `pageIdx == maxPages` with `nextURL` still set and returns `(all, nil)`, while the self-referential-link branch nine lines above returns an error; `getVMPricing` (compute/client.go:700-717) then reads a missing meter as absent rather than truncated. +- issue: (pending cross-reference) + +### A08-015 Subscriptions with a nil DisplayName are silently dropped from the org-wide account list +- category: silent-fallback +- severity: medium +- location: providers/azure/accounts_cache.go:219 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `fetchAccounts` skips any `armsubscriptions` entry whose `DisplayName` is nil, with no log and no error. The result feeds `NewMultiSubscriptionRecommendationsClient` (provider.go:681), which fans out only over the accounts it was given, so the dropped subscription is never queried and never counted in `PartialSubscriptionFailureError.Attempted` — the sweep reports complete while missing it. It also feeds `accountsContain`, so `validateConfiguredSubscription` rejects an explicitly configured subscription that exists but has no display name. A missing display name has nothing to do with whether the subscription is usable. +- evidence: + ```go + for _, sub := range page.Value { + if sub.SubscriptionID == nil || sub.DisplayName == nil { + continue + } + accounts = append(accounts, common.Account{ ... Name: *sub.DisplayName, ...}) + ``` +- suggested fix: only skip when `SubscriptionID` is nil; fall back to the subscription ID for `Name`/`DisplayName` and log the substitution. +- verdict: PLAUSIBLE — `fetchAccounts` (accounts_cache.go:230-232) does drop a nil-DisplayName subscription with no log, and the downstream claims hold (the list feeds the fan-out at provider.go:679 and `accountsContain` at accounts_cache.go:288 via validateConfiguredSubscription, provider.go:415), but nothing in the tree establishes that ARM ever returns a subscription whose `DisplayName` is nil. +- issue: (pending cross-reference) + +### A08-020 GCP treats any 403-shaped error as a permission gap and returns an empty, nil-error sweep +- category: silent-fallback +- severity: medium +- location: providers/gcp/recommendations.go:126 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when `getRegions` fails and `isPermissionError` is true, `GetRecommendations` returns `([]common.Recommendation{}, nil)`. That is exactly the outcome `mergeRegionResults` in the same file exists to prevent: "Returning (recs, nil) on a total failure makes a broken run indistinguishable from 'no savings available': the scheduler would count the account as succeeded, evict its previously collected rows, and clear last_collection_error (COR-03)". The classifier is also loose — its string fallback (recommendations.go:412) matches any error message containing both "403" and "permission", so an unrelated failure can take the silent-empty branch. +- evidence: + ```go + if isPermissionError(err) { + logging.Warnf("GCP account %s: skipping recommendations — insufficient Compute permission to list regions (grant roles/compute.viewer): %v", r.projectID, err) + return []common.Recommendation{}, nil + } + ``` +- suggested fix: return a typed permission error that the scheduler can log at WARN without treating the sweep as a successful zero-result collection, so previously collected rows are not evicted. +- verdict: CONFIRMED — `GetRecommendations` (recommendations.go:125-128) returns `([]common.Recommendation{}, nil)` on a permission-classified `getRegions` failure, exactly the outcome `mergeRegionResults` documents as unsafe for COR-03 (recommendations.go:178-189, guard at :207), and `isPermissionError` (recommendations.go:411-412) falls back to a "403"+"permission" substring test on any error string. +- issue: (pending cross-reference) + +### A08b-021 `pricing.FetchAll` truncates at the page cap and returns partial results with a nil error +- category: silent-fallback +- severity: medium +- location: providers/azure/internal/pricing/retail_prices.go:72 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the loop exits when `pageIdx == maxPages` even though `nextURL` is still non-empty, and returns `(all, nil)`. Every Azure client calls it with `DefaultMaxPages = 50` and, for cosmosdb and search, a filter that is region-wide rather than SKU-scoped (A08b-008, A08b-009), so 5000 items is reachable. The caller cannot distinguish "this SKU has no reservation price" from "the reservation price was on page 51", and reports the former. +- evidence: + ```go + for pageIdx := 0; pageIdx < maxPages && nextURL != ""; pageIdx++ { + ... + nextURL = page.NextPageLink + } + return all, nil + ``` +- suggested fix: return an explicit error when the loop exits with `nextURL != ""`, so a truncated price set can never be read as a complete one. +- verdict: CONFIRMED — the loop condition at providers/azure/internal/pricing/retail_prices.go:72 exits on `pageIdx == maxPages` regardless of `nextURL`, and line 87 returns `all, nil` with no post-loop check; every Azure client passes `pricing.DefaultMaxPages` (50, declared at line 49). +- issue: (pending cross-reference) + +### A08b-023 `GetValidResourceTypes` falls back to a hardcoded SKU list when the live listing fails, including on context cancellation +- category: silent-fallback +- severity: medium +- location: providers/azure/services/managedredis/client.go:374 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same shape at `cache/client.go:417`, `cosmosdb/client.go:418`, `search/client.go:368`) +- failure scenario: `collectSKUsFromPager` correctly returns the pager error, and the caller discards it — including `context.Canceled` and `context.DeadlineExceeded` — and substitutes a hand-maintained list of Redis SKU names. `ValidateOffering` then approves a SKU that may not exist in this region (the curated list is region-blind), or rejects a valid new SKU that Azure has added since the list was written. The cache and cosmosdb variants additionally `break` out of the page loop on error at `cache/client.go:470` and `cosmosdb/client.go:470`, silently accepting a partial SKU set. +- evidence: + ```go + skuSet, err := collectSKUsFromPager(ctx, pager) + if err != nil { + // Discard any partial results and fall back to the curated SKU list + // rather than risk false validation failures for valid SKUs. + return c.commonSKUs(), nil + } + ``` +- suggested fix: propagate the error (and always propagate context errors), reserving the curated list for an explicitly-flagged offline mode. +- verdict: CONFIRMED at the cited location, with a correction to the "same shape" claim — managedredis:380-385 does discard `collectSKUsFromPager`'s error, context errors included, but cache:424-427, cosmosdb:424-427 and search:374-377 all propagate it; what those three actually share is the curated fallback when the pager cannot be constructed (cache:419-421, cosmosdb:420-421, search:370-371) plus the silent `break` on a page error (cache:467-471, cosmosdb:467-471). +- issue: (pending cross-reference) + +### A08b-028 `walkManagedInstances` derives a "dominant" AZ configuration from counts collected before an error +- category: silent-fallback +- severity: medium +- location: providers/gcp/../azure/services/database/client.go:829 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: on a page error or context cancellation the walk returns the counts accumulated so far. `fetchServerInfo` at line 794 then compares `zoneRedundantCount == total` on that truncated sample: if page one held two zone-redundant instances and page two (which failed) held fifty non-redundant ones, the client reports `AZConfig: "zoneRedundant"` for the whole subscription. The partial-result signal is indistinguishable from a complete one. +- evidence: + ```go + page, err := pager.NextPage(ctx) + if err != nil { + if ctx.Err() != nil { + return zoneRedundant, nonZoneRedundant, total + } + logging.Warnf("azure database: managed instances page fetch failed: %v; AZConfig/Deployment signal unavailable", err) + return zoneRedundant, nonZoneRedundant, total + } + ``` +- suggested fix: return a completeness flag (or zeroed counts) on any error so the caller emits an empty `AZConfig` rather than a conclusion drawn from a partial walk. +- verdict: CONFIRMED — both error branches at database:829-833 return the counts accumulated so far, and `fetchServerInfo` at 790-800 derives `AZConfig` from `zoneRedundantCount == total` on that sample with no completeness signal; the pager-construction branch at 822-824 does return zeroes, which is the only path that degrades safely. +- issue: (pending cross-reference) + +### A08b-034 `termPlan` errors on an unknown term while `termYearsFromTerm` silently answers "one year" +- category: silent-fallback +- severity: medium +- location: providers/gcp/services/computeengine/client.go:200 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (consumers at `client.go:1139` and `client.go:1157`) +- failure scenario: the two functions parse the same `rec.Term` string with opposite policies. `termPlan` refuses `"P3Y"` with an explicit error precisely because "a silent mis-default can purchase the wrong term and waste money"; `termYearsFromTerm`, fed the same value, returns 1 and is used to compute the savings and cost fields shown to the user. The recommendation therefore displays one-year economics for a three-year term, right up until the purchase fails at `termPlan`. +- evidence: + ```go + func termYearsFromTerm(term string) int { + switch strings.ToLower(strings.TrimSpace(term)) { + case "3yr", "3", "36mo": + return 3 + default: + return 1 + } + } + ``` +- suggested fix: have `termYearsFromTerm` return `(int, error)` sharing `termPlan`'s switch, so an unrecognized term fails in one place. +- verdict: CONFIRMED — `termPlan` errors on an unrecognized term with that exact money rationale in its doc (computeengine:46-56) while `termYearsFromTerm` silently returns 1 (200-207); both read `rec.Term`, the lenient one at 1157 for the pricing lookup and 1139 for `RecurringMonthlyCost`, the strict one at 568 inside `GroupCommitments` and at 773's sibling on the purchase path. +- issue: (pending cross-reference) + +### A09-009 AnalyzeReshapingWithRecs discards the lookup error without logging it +- category: silent-fallback +- severity: medium +- location: pkg/exchange/reshape.go:550 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: a database outage or a permission regression in the `PurchaseRecLookup` closure makes every call return an error. The reshape page renders with zero cross-family alternatives on every request, indistinguishable from "AWS has not recommended anything for this region", and `err` is dropped on the floor rather than logged. An operator sees a silently degraded page with no signal anywhere that the lookup is broken. +- evidence: + ```go + offerings, err := lookup(ctx, region, currencyCode) + if err != nil || len(offerings) == 0 { + // Fall through to base recs — losing alternatives is strictly + // less bad than losing the whole reshape page. + return recs + } + ``` +- suggested fix: keep the fall-through but split the branches and emit `logging.Warnf("reshape alternatives lookup failed: %v", err)` on the error arm so a persistent failure is visible. +- verdict: CONFIRMED — pkg/exchange/reshape.go:550-555 collapses `err != nil` and `len(offerings) == 0` into one branch that returns `recs`; `err` is never logged or wrapped anywhere in `AnalyzeReshapingWithRecs`, so a persistently failing lookup is indistinguishable from an empty region. +- issue: (pending cross-reference) + +### A09-019 An unrecognized term string is recorded as 0 months in the purchase audit log +- category: silent-fallback +- severity: medium +- location: pkg/common/audit.go:84 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `Recommendation.Term` is an untyped string and `termMonths` matches two bare literals. A provider that emits `"1yr "` (trailing space), `"P1Y"`, `"12mo"` or `"5yr"` produces `AuditRecord.Term = 0` on a record that says a real commitment was purchased. The JSONL audit log is the artifact reconciled against `purchase_history`, so the reconciliation sees a zero-month commitment and the operator has no way to recover the real term from the record. The warning goes to `log.Printf`, not to the audit sink. +- evidence: + ```go + func termMonths(t string) int { + switch t { + case "1yr": + return 12 + case "3yr": + return 36 + default: + if t != "" { + log.Printf("warn: unrecognized term string %q, using 0 months", t) + } + return 0 + ``` +- suggested fix: reuse `ladder.ParseTerm`-style validation (or make `termMonths` return an error) so `NewAuditRecord` refuses to write a record whose term could not be resolved. +- verdict: CONFIRMED — `termMonths` matches only the two literals and returns 0 for everything else with a `log.Printf` warning (pkg/common/audit.go:84-99), and `NewAuditRecord` writes that 0 into `AuditRecord.Term` unconditionally (audit.go:68). A non-canonical term reaches it in practice: `normaliseTerm` in providers/azure/internal/recommendations/converter.go:350-352 explicitly passes any Azure term other than P1Y/P3Y through verbatim, and that value lands on `Recommendation.Term` at converter.go:148 and :202. The single production caller is cmd/multi_service.go:401. +- issue: (pending cross-reference) + +### A09-026 Provider factory failures are swallowed, so credential detection reports "no credentials found" +- category: silent-fallback +- severity: medium +- location: pkg/provider/registry.go:110 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GetAllProviders` logs a factory error through the stdlib `log` package (not `pkg/logging`) and drops the provider from the slice. `DetectAvailableProviders` iterates only what it got back, so a provider whose factory failed never contributes to its `errors` slice. When all three factories fail, `available` and `errors` are both empty and the caller receives "no cloud credentials found. Please configure AWS, Azure, or GCP credentials" — pointing the operator at credentials when the real cause was, say, the GCP factory's `Projects.List()` call failing on a permission error. +- evidence: + ```go + for name, factory := range factories { + provider, err := factory(&ProviderConfig{Name: name}) + if err != nil { + log.Printf("provider %q factory error: %v", name, err) + continue + } + ``` +- suggested fix: give `GetAllProviders` a second return value carrying the per-provider factory errors and have `DetectAvailableProviders` fold them into the error it reports, so a construction failure is never reported as absent credentials. +- verdict: CONFIRMED — `GetAllProviders` logs the factory error through the stdlib `log` package and `continue`s, returning only successfully constructed providers (pkg/provider/registry.go:109-116). `DetectAvailableProviders` iterates that slice alone and builds its `errors` slice only from `ValidateCredentials` failures (pkg/provider/credentials.go:26-41), so with all factories failing both `available` and `errors` are empty and line 49 returns "no cloud credentials found. Please configure AWS, Azure, or GCP credentials". +- issue: (pending cross-reference) + +### A10-007 CSV `Service`, `Term` and `PaymentOption` are cast into typed fields with no validation +- category: silent-fallback +- severity: medium +- location: cmd/multi_service_csv.go:106 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `parseCSVRecord` casts an arbitrary cell to `common.ServiceType` and copies `Term` / `PaymentOption` verbatim. A hand-edited CSV with `Service=rds ` (trailing space) or `elasticsearch` produces a ServiceType that `createServiceClient` does not recognise, so the row is dropped at cmd/multi_service.go:538 with "Service client not yet implemented" rather than an error naming the bad cell; the operator sees a short run and no failure. Worse, `PaymentOption=All Upfront` (instead of `all-upfront`) is never checked against the same `validPaymentOptions` map `validatePaymentAndTerm` enforces for the flag, and is handed straight to `PurchaseCommitment`. External input on a money path is not parsed into the enum at the boundary. +- evidence: + ```go + rec.Service = common.ServiceType(getCSVField(record, colIdx, "Service")) + ... + rec.Term = getCSVField(record, colIdx, "Term") + rec.PaymentOption = getCSVField(record, colIdx, "PaymentOption") + ``` +- suggested fix: Parse each of the three through a `parseServiceType` / `parsePaymentOption` / `parseTerm` helper that errors on an unrecognised value, reusing the `serviceMap` in cmd/main.go:190 and the `validPaymentOptions` map in cmd/validators.go:131. +- verdict: CONFIRMED on the Service cast — cmd/multi_service_csv.go:106 bypasses the serviceMap at cmd/main.go:190, so even a valid flag alias like `elasticsearch` becomes a ServiceType that createServiceClient's switch cannot match (cmd/main.go:244-268) and the row is dropped with the "not yet implemented" message at cmd/multi_service.go:538-541. The PaymentOption half is real at the boundary but does not reach a purchase: every provider converts and errors on an unknown value first (providers/aws/services/rds/client.go:550-560, plus the matching converters in ec2, elasticache, memorydb and savingsplans). +- issue: (pending cross-reference) + +### A10-009 A short CSV write is reported as success; a zero-result run claims a file that was never created +- category: silent-fallback +- severity: medium +- location: cmd/multi_service_csv.go:193 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `csv.Writer` buffers; `writer.Flush()` runs in a defer and its error is never read via `writer.Error()`. On a full disk or a failing NFS mount the header and rows are lost, `writeMultiServiceCSVReport` returns nil, and the caller prints "📋 CSV report written to: ri-helper-purchase-….csv" over a truncated or empty file that is the only record of what a `--purchase` run bought. Separately, `len(results) == 0` returns nil before `os.Create`, so a run with nothing to purchase prints the same "written to" line for a file that does not exist. +- evidence: + ```go + writer := csv.NewWriter(file) + defer writer.Flush() // error never inspected + ... + if len(results) == 0 { + return nil // no file created, caller still prints "written to" + } + ``` +- suggested fix: Call `writer.Flush()` explicitly before returning and return `writer.Error()`; return a sentinel (or have the caller check `len(allResults)`) so the "written to" line is only printed when a file was written. +- verdict: CONFIRMED — cmd/multi_service_csv.go:194-196 returns nil before os.Create and :208-209 defers Flush without ever consulting writer.Error(), while both callers print "CSV report written to" on any nil return (cmd/multi_service.go:175-179 and :579-583). +- issue: (pending cross-reference) + +### A10-012 An unknown `--services` value is warned about and skipped instead of rejected +- category: silent-fallback +- severity: medium +- location: cmd/main.go:219 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `cudly --services rds,elasticahe --purchase` (typo). `parseServices` logs one `log.Printf` warning to stderr, drops ElastiCache, and the run proceeds against RDS alone. Amid the wizard's heavy emoji-decorated stdout the warning is easy to miss, and the operator believes both services were covered. Only a run where *every* name is bad reaches the `log.Fatalf("No valid services specified")` guard. External input is not parsed into the enum at the boundary with an error on unknown. +- evidence: + ```go + if service, ok := serviceMap[key]; ok { + add(service) + } else { + log.Printf("Warning: Unknown service '%s', skipping", name) + } + ``` +- suggested fix: Return an error from `parseServices` naming the unrecognised value and the valid set, and call it from `validateFlags` so the run never starts. +- verdict: CONFIRMED — cmd/main.go:217-220 adds the match or logs a warning and drops the name, and only a fully empty result reaches log.Fatalf("No valid services specified") at cmd/multi_service.go:94-96. +- issue: (pending cross-reference) + +### A10-013 Engine-version query failures silently disable the extended-support exclusion +- category: silent-fallback +- severity: medium +- location: cmd/multi_service_helpers.go:322 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The default is to exclude instances on extended-support engine versions. If `queryRunningInstanceEngineVersions` or `queryMajorEngineVersions` fails, both helpers return an empty map and the run continues; `isInExtendedSupport` then returns false for every version ("If we don't have info, assume not in extended support", cmd/multi_service_engine_versions.go:395) and the filter becomes a no-op, so the run buys 3-year RIs for instances the operator wanted excluded. The failure need not be visible: `queryRDSInstancesInRegions` logs per-region errors and still returns `(map, nil)` (line 128), and `queryMajorEngineVersionsWithClient` logs per-engine errors and returns `(map, nil)` (line 226), so a total failure across all regions and all four engines is reported to the operator as "✅ Found 0 instance types". +- evidence: + ```go + instanceVersions, err := queryRunningInstanceEngineVersions(ctx, cfg) + if err != nil { + AppLogger.Printf("⚠️ Warning: Failed to query running instances ...: %v\n", err) + AppLogger.Printf(" Continuing without engine version filtering\n") + return make(map[string][]InstanceEngineVersion) + } + ``` +- suggested fix: Have both query helpers return an error when every region or every engine failed, and abort a `--purchase` run (not a dry run) when `!cfg.IncludeExtendedSupport` and the signal is unavailable. +- verdict: CONFIRMED — cmd/multi_service_helpers.go:318-329 swallows the error into an empty map, and neither query can report total failure: queryRDSInstancesInRegions returns (map, nil) unconditionally after per-region logging (cmd/multi_service_engine_versions.go:127-128, :140-143) and queryMajorEngineVersionsWithClient does the same per engine (:220-226), so isInExtendedSupport's "no info means not in extended support" default at :393-397 turns the filter into a no-op. +- issue: (pending cross-reference) + +### A10-015 Savings Plan type is matched against bare string literals with no default arm +- category: silent-fallback +- severity: medium +- location: cmd/multi_service_stats_helpers.go:27 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `SavingsPlanDetails.PlanType` is a plain `string` (pkg/common/types.go:613) and both `categorizeSPRecommendations` and `collectSPSavings` (line 107) switch on the literals `"Compute"`, `"EC2Instance"`, `"SageMaker"`, `"Database"` with no default. If AWS returns or a future parser emits any other spelling, the rec's savings fall into no bucket and vanish from the SAVINGS PLANS section and from the RI-vs-SP comparison, so the recommended purchasing strategy is computed from a subset of the data without any indication that rows were dropped. The same four literals are duplicated across the two functions, so they can drift apart. +- evidence: + ```go + switch details.PlanType { + case "Compute": + breakdown.ComputeSavings += rec.EstimatedSavings + breakdown.ComputeCount++ + case "EC2Instance": + ... + } // no default: unrecognised plan types are silently dropped + ``` +- suggested fix: Define a `common.SavingsPlanType` const set, use it in both switches, and add a default arm that logs the unrecognised value rather than dropping the row. +- verdict: PLAUSIBLE — the missing default arm and the duplicated literals are real (cmd/multi_service_stats_helpers.go:27-40 and :107-116, the second switch omitting SageMaker entirely), but reaching the drop needs the runtime condition of Cost Explorer returning a plan type outside the four modelled SDK members, which spPlanTypeDisplayString passes through verbatim at providers/aws/recommendations/parser_sp.go:286; costexplorer v1.63.1 models exactly the four the switches handle. +- issue: (pending cross-reference) + +### A10-020 Skipping GCP service-account creation still returns a fabricated account email +- category: silent-fallback +- severity: medium +- location: cmd/configure_gcp.go:678 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: An operator answers `s` at Step 3 because they already have a differently-named service account. `gcpStepCreateServiceAccount` returns the *constructed* string `cudly-service-account@.iam.gserviceaccount.com` rather than signalling that nothing was created. Step 4 then prints "grants roles/compute.admin to cudly-service-account@… on project …" for a principal that may not exist, and Step 5 tries to mint a key for it. The wizard states a fact about an identity it never verified, and the failure surfaces several steps later as an opaque IAM error. +- evidence: + ```go + saName := "cudly-service-account" + saEmail := fmt.Sprintf("%s@%s.iam.gserviceaccount.com", saName, projectID) + ... + case "s", "skip": + fmt.Println("Skipping Create Service Account") + } + return saEmail, nil + ``` +- suggested fix: Return an empty string on skip and prompt for the existing service-account email, the way `gcpStepCreateKey` already returns "" on skip so the caller asks for an existing credentials file. +- verdict: CONFIRMED — the skip and default arms at cmd/configure_gcp.go:673-677 fall through to `return saEmail, nil` at :678 with the address constructed at :651-652, unlike gcpStepCreateKey which returns "" on skip (:744-750). +- issue: (pending cross-reference) + +### A10-023 `rekey` counts any zero-key decrypt failure as "already re-keyed" and exits successfully +- category: silent-fallback +- severity: medium +- location: cmd/rekey/main.go:168 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `rekeyOne` treats *every* `credentials.Decrypt` failure as proof the row is already encrypted under the real key. A row whose `encrypted_blob` is truncated, base64-corrupt, empty, or written by a third key version fails the same way and lands in `skippedAlreadyReal`. The run then reports `errored=0`, `run` returns nil, and the migration is declared complete while an unreadable credential row survives untouched. The only signal is a count in one log line that the operator has no baseline to compare against. +- evidence: + ```go + plaintext, err := credentials.Decrypt(zeroKey, blob) + if err != nil { + // Decrypt with zero key failed — assume already real-key encrypted. + return outcomeSkipped + } + ``` +- suggested fix: Attempt a real-key decrypt in the skip branch and only count the row as `skippedAlreadyReal` when that succeeds; anything failing under both keys is an error with its id logged. +- verdict: CONFIRMED — rekeyOne returns outcomeSkipped for any Decrypt error whatsoever (cmd/rekey/main.go:167-172), that outcome only ever increments skippedAlreadyReal (:146-151), and run returns nil whenever cs.errored == 0 (cmd/rekey/main.go:81-87), so a corrupt or third-key row leaves the migration reporting success. +- issue: (pending cross-reference) + +### A11-009 A truncated 200 response is turned into `null` for every endpoint +- category: silent-fallback +- severity: medium +- location: frontend/src/api/client.ts:284 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the catch is documented as covering 204s and body-stripping proxies, but it is unconditional: any 200 whose body is truncated or malformed (a proxy cutting a large recommendations payload, a partial gateway write) resolves as `null` instead of rejecting. Callers that spread the result then see "no data" rather than an error — for example `listActiveCommitments` does `resp.commitments ?? []` and would throw on null, while `getHistory` returns null into a renderer that treats it as an empty list. A user sees an empty table where a failure occurred. +- evidence: + ```typescript + // api/client.ts:284-288 + try { + return await response.json() as T; + } catch { + return null as T; + } + ``` +- suggested fix: return null only when the response has no body to parse (204, or `content-length: 0`), and rethrow the parse error otherwise. +- verdict: CONFIRMED — the catch at frontend/src/api/client.ts:283-287 is unconditional and inspects neither the status nor `content-length`, so every 2xx whose body fails to parse resolves as `null` for every endpoint; the sibling catch on the error path (client.ts:271-273) is deliberately narrow by comparison, and nothing downstream distinguishes "empty by design" from "truncated". +- issue: (pending cross-reference) + +### A11-016 A failed account fetch tells the operator no accounts are configured, then clones an unscoped group +- category: silent-fallback +- severity: medium +- location: frontend/src/groups/groupModals.ts:623 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `openDuplicateGroupModal` catches any `listAccounts` failure and continues with `accounts = []`. `renderDuplicateAccountsList` then renders "No cloud accounts configured yet. Duplicating without scope clones the full source group — add accounts first if you want to restrict." (groupModals.ts:566). On a 403 or a transient 5xx the operator is told a factual untruth about their deployment and is nudged to create the duplicate with no account scoping, inheriting the source group's `allowed_accounts` unchanged. That is the widest available outcome on an authorization-scoping path. +- evidence: + ```typescript + // groups/groupModals.ts:620-627 + let accounts: api.CloudAccount[] = []; + try { + accounts = await api.listAccounts(); + } catch (err) { + console.error('Failed to list accounts for duplicate modal:', err); + accounts = []; + } + ``` +- suggested fix: distinguish the two states — on a fetch failure render an error line and disable the duplicate submit, keeping the "none configured" copy for a genuinely empty list. +- verdict: CONFIRMED — the catch swallows every `listAccounts` failure into `accounts = []` (frontend/src/groups/groupModals.ts:620-626), the empty branch of `renderDuplicateAccountsList` prints the "No cloud accounts configured yet" copy unconditionally (groupModals.ts:562-568), and the modal still opens (groupModals.ts:630) with the submit path enabled, whose own doc comment confirms that ticking nothing inherits the source group's `allowed_accounts` as-is (groupModals.ts:646-650). +- issue: (pending cross-reference) + +### A13c-012 Two secrets are seeded with hardcoded placeholder literals that the runtime cannot distinguish from a real credential +- category: silent-fallback +- severity: medium +- location: terraform/modules/secrets/gcp/main.tf:172 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: with `create_sendgrid_secret = true` and no `sendgrid_api_key`, the secret is + created holding the string `PLACEHOLDER_REPLACE_ME`. The Azure module does the same for both + SMTP credentials (`secrets/azure/main.tf:190` and `:207`, value + `PLACEHOLDER_GENERATE_IN_AZURE_PORTAL`). The application resolves a secret that exists and is + non-empty, so nothing errors at startup; the first real send fails at the provider with an + authentication error, at whatever hour the first approval email goes out. A secret that is + absent fails loudly at resolve time; a secret holding a placeholder does not. +- evidence: + ```hcl + resource "google_secret_manager_secret_version" "sendgrid_api_key" { + count = var.sendgrid_api_key != null || var.create_sendgrid_secret ? 1 : 0 + + secret = google_secret_manager_secret.sendgrid_api_key[0].id + secret_data = var.sendgrid_api_key != null ? var.sendgrid_api_key : "PLACEHOLDER_REPLACE_ME" + } + ``` +- suggested fix: create the empty secret container without a version when no value is supplied, + so the runtime's "no version" resolve error names the missing manual step directly. +- verdict: CONFIRMED — the three placeholder literals are as cited (secrets/gcp/main.tf:172, + secrets/azure/main.tf:190 and :207), and `/usr/bin/grep -rn PLACEHOLDER internal/ providers/ cmd/` + returns nothing, so no runtime code distinguishes a placeholder from a real credential; the + resolvers only fail on absent or unreadable secrets. +- issue: (pending cross-reference) + +### A01-016 A GetGlobalConfig failure silently reverts purchase suppressions to the default grace period +- category: silent-fallback +- severity: low +- location: internal/api/handler_purchases.go:2628 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: An operator has set `grace_period_days` to 0 for GCP (feature off) or to 30 for AWS. During a DB blip `GetGlobalConfig` errors; `executePurchase` (and `persistRetryExecution`, 1913) pass `nil` config to `buildSuppressions`, which substitutes `config.DefaultGracePeriodDays` for every provider. Suppression rows are written for the disabled provider and with the wrong expiry for the others; the error is not even logged. Every other config read on these paths fails closed (`approveViaToken`, 608-611). +- evidence: + ```go + var gracePeriodCfg *config.GlobalConfig + g, getConfigErr := h.config.GetGlobalConfig(ctx) + if getConfigErr == nil { + gracePeriodCfg = g + } + + suppressions := buildSuppressions(execReq.Recommendations, executionID, gracePeriodCfg, time.Now()) + ``` +- suggested fix: Return the config error (500) from both call sites instead of defaulting; `buildSuppressions` can then require a non-nil config. +- verdict: CONFIRMED — executePurchase (internal/api/handler_purchases.go:2627-2633) and persistRetryExecution:1912-1916 pass nil on a GetGlobalConfig error without logging it, and buildSuppressions:71-73 substitutes config.DefaultGracePeriodDays whenever cfg is nil, so a provider configured to 0 still gets suppression rows; approveViaToken:608-611 and approvePurchaseViaSession:708-711 return the same error instead. +- issue: (pending cross-reference) + +### A02-013 Account update and create treat an omitted `enabled` (and `aws_is_org_root`) as false +- category: silent-fallback +- severity: low +- location: internal/api/handler_accounts.go:515 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `Enabled *bool` signals an optional field, but `cloudAccountFromRequest` leaves `a.Enabled` at the zero value when the pointer is nil and `updateAccount` (lines 568-576) builds the stored row from the request alone, copying only CreatedAt/CreatedBy from `existing`. An API client (or Terraform) sending `PUT /api/accounts/{id}` with name and credentials but no `enabled` key silently disables the account and clears `aws_is_org_root`; the scheduler then skips it with no error. The frontend happens to always send `enabled` (settings.ts:2453), which hides this from UI users only. +- evidence: + ```go + if req.Enabled != nil { + a.Enabled = *req.Enabled + } + ... + account := cloudAccountFromRequest(req) + account.ID = id + account.CreatedAt = existing.CreatedAt + account.CreatedBy = existing.CreatedBy + ``` +- suggested fix: In `updateAccount`, fall back to `existing.Enabled` when `req.Enabled == nil` (and treat `aws_is_org_root` the same way with a `*bool`), or require `enabled` and return 400 when absent. +- verdict: CONFIRMED — CloudAccountRequest has Enabled *bool but AWSIsOrgRoot bool (types.go:29,48); cloudAccountFromRequest (handler_accounts.go:504-516) leaves Enabled false when the pointer is nil and copies AWSIsOrgRoot's zero value, and updateAccount (handler_accounts.go:568-576) copies only ID/CreatedAt/CreatedBy from `existing` before UpdateCloudAccount, so an omitted key persists as false. +- issue: (pending cross-reference) + +### A02-022 Commitment-option validation errors are swallowed, allowing the save +- category: silent-fallback +- severity: low +- location: internal/api/handler_config.go:177 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `checkCommitmentOptionCombo` is the server-side guard that a (term, payment) pair is actually sold for the service. When `commitmentOpts.Validate` errors (DB blip, timeout) the handler logs and returns nil, so `PUT /api/config/service/aws/rds` with an unsupported combination persists and the scheduler later fails the purchase. The comment defends this as "the frontend's hardcoded rules are the primary gate", which makes the backend check advisory on precisely the request it exists for. +- evidence: + ```go + ok, err := h.commitmentOpts.Validate(ctx, cfg.Provider, cfg.Service, cfg.Term, cfg.Payment) + if err != nil { + logging.Warnf("commitment-option validation error (allowing save): %v", err) + return nil + } + ``` +- suggested fix: Return a 503 (or 500) on a validation error so the operator retries, keeping only `ErrNoData` as the permissive case. +- verdict: CONFIRMED — checkCommitmentOptionCombo (handler_config.go:177-181) logs and returns nil on a Validate error and updateServiceConfig (handler_config.go:356-362) proceeds to SaveServiceConfig; the swallow is documented as intentional in the function comment (handler_config.go:166-171), so this is a design choice the finding disputes rather than an accidental gap, and low severity stands. +- issue: (pending cross-reference) + +### A03-012 A missing role ARN resolves to the host's ambient credentials instead of an error +- category: silent-fallback +- severity: medium +- location: internal/credentials/resolver.go:175 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `role_arn` with an empty `AWSRoleARN` returns `ambient` (the Lambda/instance role); `ResolveGCPTokenSourceWithOpts` returns `(nil, nil)` for `application_default` (resolver.go:390-392) and Azure `managed_identity` returns the host identity (resolver.go:346-347). The API validator permits an empty ARN in `role_arn` mode (internal/api/handler_accounts.go:412-426), so an account row whose ARN is blanked by an edit, or an account registered by any `create:accounts` holder, quietly runs collection and purchases as the CUDly deployment identity rather than failing. The "Self account" shape is intentional, but the resolver cannot distinguish "operator chose Self" from "customer ARN went missing", and the choice is not gated to admins anywhere in scope. +- evidence: + ```go + if account.AWSRoleARN == "" { + // Self-account: auth_mode=role_arn with no role ARN means "use the + // CUDly Lambda's own credentials to access this account." ... + if ambient != nil { + return ambient, nil + } + return nil, fmt.Errorf("credentials: aws_role_arn is empty and no ambient credentials available (account %s)", account.ID) + } + ``` +- suggested fix: Make the ambient shape an explicit auth mode value (e.g. `self`) validated at the boundary and creatable only by admins, and have `role_arn` with an empty ARN error; return an explicit ambient sentinel instead of `(nil, nil)` on the GCP path. +- verdict: PLAUSIBLE — the empty-ARN-means-host-identity shape is documented and deliberate in three places (resolver.go:165-168 and :175-183, AWSResolveOptions comment :80-83, scheduler.go:753-759), validateAWSRoleARN accepts "" (validation.go:90-93) and createAccount is gated on create:accounts only (handler_accounts.go:307), so the fallback is real but the "silent" and privilege claims need a deployment that delegates create:accounts/update:accounts to non-admins, which the source cannot establish; on the no-opts callers an empty ARN errors rather than falls back (resolver.go:182). +- severity-adjusted: low — a documented feature reachable only by account-management holders, not an unintended fallback on the money path. +- issue: (pending cross-reference) + +### A04-006 saveGlobalConfigWith silently rewrites an RI-exchange utilization threshold of 0 to 95 +- category: silent-fallback +- severity: medium +- location: internal/config/store_postgres.go:265-268 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the RI exchange config handler accepts `utilization_threshold` in `[0, 100]` (internal/api/handler_ri_exchange.go:2576) and writes it through `UpdateGlobalConfigAtomic` without a whole-config Validate. An operator with auto-exchange enabled sets the threshold to 0 to stop exchanges from triggering (no RI has utilization below 0). The store replaces 0 with 95.0 before the UPSERT, the config round-trips as 95, and the auto-exchange task now considers every RI under 95 percent for an irreversible exchange. The same block rewrites `ri_exchange_lookback_days` 0 to 30 and `recommendations_lookback_days` 0 to 7. +- evidence: + ```go + riExchangeUtilizationThreshold := config.RIExchangeUtilizationThreshold + if riExchangeUtilizationThreshold == 0 { + riExchangeUtilizationThreshold = 95.0 + } + ``` +- suggested fix: delete the zero-to-default rewrites in `saveGlobalConfigWith`; reject or explicitly define 0 in the handler's `validate` and persist what was validated. +- verdict: CONFIRMED — the handler accepts 0 (`UtilizationThreshold < 0 || > 100`, internal/api/handler_ri_exchange.go:2576-2578) and `saveGlobalConfigWith` rewrites it to 95.0 before the UPSERT (internal/config/store_postgres.go:265-268), so the persisted value is not the validated one; the sibling rewrites of `ri_exchange_lookback_days` and `recommendations_lookback_days` are at 260-264. +- severity-adjusted: low — no backend consumer reads `RIExchangeUtilizationThreshold`; every reference is the config read/write plus the two GET responses (internal/api/handler_ri_exchange.go:1997, internal/server/handler_ri_exchange.go:103), so no exchange is selected by the rewritten value and the harm is a lied-about setting rather than an unintended exchange. +- issue: (pending cross-reference) + +### A04-016 database env parsing falls back to defaults on malformed values +- category: silent-fallback +- severity: low +- location: internal/database/config.go:179-204 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `DB_MAX_CONNECTIONS=2O` (letter O) yields 25, `DB_CONNECT_TIMEOUT=10` (no unit, which `time.ParseDuration` rejects) yields 10s, `DB_AUTO_MIGRATE=yes` yields false. The operator's setting is discarded without any log line and `Validate` sees only the defaults, so a pool sized for a Lambda concurrency limit or a deliberately disabled auto-migrate silently reverts. +- evidence: + ```go + func getEnvInt(key string, defaultValue int) int { + if value := os.Getenv(key); value != "" { + if intVal, err := strconv.Atoi(value); err == nil { + return intVal + } + } + return defaultValue + } + ``` +- suggested fix: return an error from `LoadFromEnv` when a set variable fails to parse (mirroring `maybeForceMigrationVersion`'s non-numeric rejection). +- verdict: CONFIRMED — `getEnvInt`, `getEnvBool` and `getEnvDuration` each discard the parse error and fall through to the default with no log line (internal/database/config.go:179-203), and the three cited defaults are exactly the values claimed: `DB_MAX_CONNECTIONS` 25 (internal/database/config.go:47), `DB_CONNECT_TIMEOUT` 10s (config.go:52), `DB_AUTO_MIGRATE` false (config.go:55). +- issue: (pending cross-reference) + +### A05-012 Azure probe treats a missing `Valid` flag as "combo is available" +- category: silent-fallback +- severity: low +- location: internal/commitmentopts/probe_azure.go:187 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `probeCombo` returns true unless some benefit explicitly reports `Valid == false`. A response with an empty `Benefits` slice, or one whose `Valid` pointer is nil (a field the service omitted, or an SDK shape change), is read as "this (term, payment) is sold". The 5-year candidate combos exist precisely so the API can reject what it does not sell, so a nil flag would persist `P5Y` combos as live and the frontend would offer a 5-year Azure savings plan that cannot be bought. The code comments the empty-slice case as "can't happen in practice" but does not guard the nil-pointer case at all. +- evidence: + ```go + for _, b := range resp.Benefits { + if b != nil && b.Valid != nil && !*b.Valid { + return false, nil + } + } + return true, nil + ``` +- suggested fix: require positive confirmation — return true only when at least one benefit reports `Valid == true`, and treat an empty slice or a nil flag as not-offered. +- verdict: PLAUSIBLE — the code reads exactly as claimed (probe_azure.go:188-196: the loop can only return false, and a nil `b.Valid` or an empty Benefits slice falls through to `return true`), and probe_azure_test.go:138 and :150 pin that permissive behaviour as intended. The named consequence needs two conditions I could not establish: that the Azure service ever omits the Valid flag, and that the prober runs at all — A05-009 shows ProbeAzure has no non-test caller, so no P5Y combo can currently reach the store or the frontend. +- issue: (pending cross-reference) + +### A06-017 ANALYTICS_COLLECTION_ENABLED silently stays enabled on an unparseable value +- category: silent-fallback +- severity: low +- location: internal/server/analytics_collect.go:96 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `strconv.ParseBool` rejects `no`, `off`, `False `, `disabled`. An operator who sets `ANALYTICS_COLLECTION_ENABLED=no` to stop the collector (for example while a partition problem is being fixed) gets the default `true` with no warning at all, and the collector keeps writing. The immediately preceding helper in the same file, `loadAnalyticsInt`, deliberately returns an out-of-range sentinel so `Validate` fails startup rather than "running with a default the operator never asked for" — this function contradicts that rationale. +- evidence: + ```go + func getEnvBool(key string, defaultVal bool) bool { + if val := os.Getenv(key); val != "" { + if result, err := strconv.ParseBool(val); err == nil { + return result + } + } + return defaultVal + } + ``` +- suggested fix: make a set-but-unparseable value a startup error via `AnalyticsConfig.Validate`, or at minimum log a warning as the sibling env parsers do. +- verdict: CONFIRMED — `getEnvBool` discards the `ParseBool` error entirely and returns `defaultVal` with no log line at all (internal/server/analytics_collect.go:96-103), unlike `getEnvInt`/`getEnvFloat` which at least warn (internal/server/app.go:948,965); it is the reader for `ANALYTICS_COLLECTION_ENABLED` with `defaultVal=true` (analytics_collect.go:49), and `strconv.ParseBool` rejects `no`/`off`/`disabled`, so `cfg.Enabled` stays true and the gate at analytics_collect.go:129-133 never trips. `Validate` checks only the two int knobs (analytics_collect.go:75-83). +- issue: (pending cross-reference) + +### A06-018 analytics_refresh reports success with fabricated zero counters when the analytics store is absent +- category: silent-fallback +- severity: low +- location: internal/server/handler.go:369 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when `app.Analytics` is nil the task logs "skipping" and returns `{"status":"success","views_refreshed":0,"partitions_created":0,"partitions_dropped":0}`. The HTTP scheduled-task path (`handleScheduledHTTP`) wraps that in `{"status":"success"}` too, so a monitor watching the endpoint sees a healthy refresh that never happened. `partitions_created` and `partitions_dropped` are hardcoded zeros this function never writes to at all. The neighbouring `handleCleanupExpiredRecords` has the same defect: `result["sessions_deleted"]` is initialised to 0 and never assigned, then logged as "Cleanup complete: 0 sessions". +- evidence: + ```go + } else { + log.Println("Analytics store not available, skipping materialized view refresh") + } + log.Printf("Analytics refresh complete") + return result, nil + ``` +- suggested fix: return `status: "skipped"` when the store is nil (as `handleCollectAnalytics` already does) and drop the counters that are never populated. +- verdict: CONFIRMED — the nil-store branch only logs and falls through to `return result, nil` with `status: "success"` still set (internal/server/handler.go:372-397), `partitions_created` and `partitions_dropped` are initialised at handler.go:375-376 and never written anywhere in the function, and the HTTP wrapper adds its own `"status":"success"` envelope (internal/server/http.go:257-265). `handleCollectAnalytics` does use `"skipped"` for the same condition (internal/server/analytics_collect.go:134-138). The sibling defect also reproduces: `sessions_deleted` is set to 0 at handler.go:275, never assigned (`CleanupExpiredSessions` returns only an error, handler.go:280-286), then logged as a count at handler.go:303. +- issue: (pending cross-reference) + +### A06-024 Request body read errors are discarded and oversize bodies are silently truncated +- category: silent-fallback +- severity: low +- location: internal/server/http.go:270 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `err` from `io.ReadAll` is dropped, so a client that disconnects mid-upload or a body that fails to read yields `body = ""` and the request continues into the API router as if it had been sent with an empty body. Separately the 10 MB `LimitReader` truncates rather than rejecting: an oversize POST reaches the handler as a body cut mid-JSON, producing a confusing 400 parse error instead of a 413. +- evidence: + ```go + limited := io.LimitReader(r.Body, maxBodySize) + bodyBytes, err := io.ReadAll(limited) + if err == nil && len(bodyBytes) > 0 { + body = string(bodyBytes) + } + ``` +- suggested fix: propagate the read error as a 400 and use `http.MaxBytesReader` so an oversize body is rejected with 413 instead of silently cut. +- verdict: CONFIRMED — `httpToLambdaRequest` has no error return and the `err` from `io.ReadAll` is only used as a guard on assigning `body`, never propagated (internal/server/http.go:271-280), so a mid-upload disconnect or read fault yields `Body: ""` on the request handed to `app.API.HandleRequest` (http.go:169-172); `io.LimitReader` truncates at 10 MB rather than erroring, and no `http.MaxBytesReader` appears anywhere in the package. +- issue: (pending cross-reference) + +### A06-025 STS failure yields an "unknown" account id that is persisted on exchange audit rows +- category: silent-fallback +- severity: low +- location: internal/server/handler_ri_exchange.go:130 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: on a transient `GetCallerIdentity` failure this returns the literal `"unknown"`, which flows into `RunAutoExchangeParams.AccountID` and is written to `ri_exchange_records.account_id` for every record created by that run. The audit trail for a money action then carries a fabricated identifier that cannot be distinguished from a genuine one, and the sibling ladder path deliberately refuses to do this (`defaultLadderAccountResolver`, internal/server/handler_ladder.go:150, "a fabricated/'unknown' value is unsafe"). +- evidence: + ```go + identity, err := stsClient.GetCallerIdentity(ctx, &sts.GetCallerIdentityInput{}) + if err != nil { + log.Printf("Warning: failed to get AWS account ID via STS: %v (using 'unknown')", err) + return "unknown" + } + ``` +- suggested fix: return an error and abort the reshape run, as the ladder resolver does; an audit row is not worth writing with a fabricated subject. +- verdict: CONFIRMED — `resolveAccountID` returns the `"unknown"` literal on both the STS error and the nil-`Account` path (internal/server/handler_ri_exchange.go:130-141), it is assigned to `clients.accountID` (handler_ri_exchange.go:71) and passed as `RunAutoExchangeParams.AccountID` (handler_ri_exchange.go:108), which is written onto every record the run creates (pkg/exchange/auto.go:374,566,615) and persisted into the `account_id` column (internal/config/store_postgres.go:2641). The sibling ladder resolver refuses exactly this and returns an error instead (internal/server/handler_ladder.go:144-156). +- issue: (pending cross-reference) + +### A07-003 Partial-upfront breakdown uses a hardcoded 50% split and silently treats unknown payment options as all-upfront +- category: silent-fallback +- severity: medium +- location: providers/aws/services/savingsplans/client.go:663 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GetOfferingDetails` reports `UpfrontCost` / `RecurringCost` to the caller. For a partial-upfront SP the split is fixed at exactly 0.5 rather than read from the offering, and any payment-option string outside the six recognised spellings (for example the CE display form `"Partial Upfront"` arriving from a persisted rec) falls into `default` and is reported as 100% upfront with zero recurring — the most expensive shape, presented as fact. +- evidence: + ```go + case "Partial Upfront", "partial-upfront": + return totalCost * 0.5, (totalCost * 0.5) / hoursInTerm + case "No Upfront", "no-upfront": + return 0, totalCost / hoursInTerm + default: + return totalCost, 0 + } + ``` +- suggested fix: return an error for unrecognised payment options (matching `convertPaymentOption` two functions above), and derive the upfront share from the offering rates rather than a fixed 0.5. +- verdict: CONFIRMED — the fixed 0.5 at providers/aws/services/savingsplans/client.go:668 and the `default: return totalCost, 0` at :671-672 are exactly as described, though the cited `"Partial Upfront"` example is in fact handled at :667 so the unknown value must be some other string. +- severity-adjusted: low — the only consumer is `GetOfferingDetails` (client.go:607), which has no production caller in the repo; `GetOfferingDetails` appears only as an interface declaration at pkg/provider/interface.go:50, so no operator sees these numbers today. +- issue: (pending cross-reference) + +### A07-004 Term helpers silently fall back to one year on any unrecognised term +- category: silent-fallback +- severity: medium +- location: providers/aws/services/savingsplans/client.go:655 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `convertTermToSeconds` in the same file errors on an unknown term, but `calculateHoursInTerm` and `normalizeTermString` on the `GetOfferingDetails` path silently return the 1-year value. A rec carrying `Term: "5yr"` or an empty term produces `TotalCost` computed over 8760 hours and a `Term: "1yr"` label, which the caller displays as the offering's real terms. +- evidence: + ```go + func calculateHoursInTerm(term string) float64 { + if term == "3yr" || term == "3" { + return 3 * 365 * 24 + } + return 365 * 24 + } + ``` +- suggested fix: route both helpers through `convertTermToSeconds` and propagate its error so `GetOfferingDetails` fails rather than labelling an unknown term as one year. +- verdict: CONFIRMED — `calculateHoursInTerm` (providers/aws/services/savingsplans/client.go:655) and `normalizeTermString` (:677) both fall through to the 1-year value, and the path is reachable only through the CE-OfferingID short-circuit at :425 because otherwise `convertTermToSeconds` (:536) errors first inside `findOfferingID`. +- severity-adjusted: low — same reachability limit as A07-003: `GetOfferingDetails` has no production caller, only the pkg/provider/interface.go:50 declaration. +- issue: (pending cross-reference) + +### A07-028 Unrecognised OpenSearch payment options are passed through unchanged +- category: silent-fallback +- severity: low +- location: providers/aws/services/opensearch/client.go:477 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `normalizeOpenSearchPaymentOption` returns the input verbatim in its `default` arm. A rec carrying `"All Upfront"` produces `wantPayment = "All Upfront"`, which never equals the SDK enum string, so the first offering matching instance type and duration triggers the "has payment option X, want Y" error and aborts the whole offering search. The operator sees an API-mismatch error rather than "your recommendation's payment option is not one of the three supported values". +- evidence: + ```go + case "no-upfront": + return string(types.ReservedInstancePaymentOptionNoUpfront) + default: + return option + } + ``` +- suggested fix: return an error for unrecognised values and surface it from `findOfferingID`, matching `convertEC2PaymentOption` and the ElastiCache/MemoryDB equivalents. +- verdict: CONFIRMED — providers/aws/services/opensearch/client.go:478-479 returns the input verbatim in the `default` arm, and `scanOpenSearchOfferingPage` :448/:457-460 then compares it against the SDK enum string and aborts the whole search with an "has payment option X, want Y" error rather than naming the unsupported input. +- issue: (pending cross-reference) + +### A08-024 GCP offering details parse the term with bare literals and silently fall back to one year +- category: silent-fallback +- severity: medium +- location: providers/gcp/services/computeengine/client.go:856 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GetOfferingDetails` only recognizes "3yr" and "3". A recommendation carrying "36mo" — a form `termPlan` (client.go:50) explicitly accepts on the purchase path — is priced as a one-year commitment, so `recurringCost` is three times the real monthly charge and `hoursInTerm` inside `getComputePricing` is a third of the real term. The purchase path and the pricing path therefore disagree about the same recommendation. Repeated verbatim in cloudsql/client.go:260, memorystore/client.go:271 and cloudstorage/client.go:275. +- evidence: + ```go + termYears := 1 + if rec.Term == "3yr" || rec.Term == "3" { + termYears = 3 + } + ``` +- suggested fix: call the existing `termYearsFromTerm` helper (client.go:200) — or better, a variant that returns an error — in all four clients so the pricing path accepts exactly what the purchase path accepts. +- verdict: PLAUSIBLE — the bare, case-sensitive literals are in all four `GetOfferingDetails` (computeengine/client.go:856-859, cloudsql:260-263, memorystore:271-274, cloudstorage:275-278) while `termPlan` (client.go:48-56) accepts "36mo" and `convertGCPRecommendation` (client.go:1094-1102) can therefore emit `Term == "36mo"`, but the live recommendations path already uses `termYearsFromTerm` (`enrichRecWithPricing`, client.go:1157) and `GetOfferingDetails` has no non-test caller anywhere in the tree — it is only a `pkg/provider/interface.go:50` member — so the mispricing needs an out-of-tree consumer of that exported interface. +- severity-adjusted: low — no in-tree caller reaches the divergent branch; the in-tree pricing path uses the correct helper. +- issue: (pending cross-reference) + +### A09-018 A non-empty Details payload for an unrecognized service is discarded without error +- category: silent-fallback +- severity: medium +- location: pkg/common/service_details_codec.go:91 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `newDetailsForService` has no entry for `memorydb`, and none for any service slug a newer writer introduces. A row persisted by a newer deployment with `service = ""` and a fully populated `Details` blob decodes to `(nil, nil)` on an older reader: the payload is dropped and the purchase proceeds with nil details rather than failing. The comment calls this "suspicious but tolerated", which is the tolerance that let #453 hide for a release. +- evidence: + ```go + target, ok := newDetailsForService(service) + if !ok { + // Service has no *Details type — nothing to decode. Empty raw + // payload is fine; a non-empty payload on a service we don't + // recognize is suspicious but tolerated + return nil, nil + } + ``` +- suggested fix: return `(nil, nil)` only when `raw` is empty; when a payload exists for a service with no typed shape, return an error naming the service so the version skew surfaces instead of silently dropping purchase parameters. +- verdict: CONFIRMED — `newDetailsForService` (pkg/common/service_details_codec.go:121-167) has no case for memorydb or any unlisted slug, and `DecodeServiceDetailsFor` returns `(nil, nil)` at line 95 before ever looking at `raw`, so a fully populated payload is discarded without error. +- severity-adjusted: low — no current service slug loses data this way: service_details_codec.go:154-161 documents memorydb's nil details as deliberate (the MemoryDB client reads `rec.ResourceType` directly), and the version-skew case needs a slug a future writer has not introduced yet. The consumers that need details already fail loud on nil, e.g. providers/aws/services/ec2/client.go:455-457. +- issue: (pending cross-reference) + +### A11-004 A malformed Max Amount is silently dropped, widening the permission instead of refusing the save +- category: silent-fallback +- severity: medium +- location: frontend/src/groups/groupModals.ts:450 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when the number input holds something `parseFloat` cannot turn into a finite non-negative value, the branch simply skips the assignment and the permission is saved with no `max_amount`. For a permission that previously carried a cap, the save removes it and the operator sees a success toast. The same file refuses to save at all when a list constraint cannot round-trip (`unrepresentablePermissionErrors`, groupModals.ts:361), so the money dimension is the one dimension that fails open. +- evidence: + ```typescript + // groups/groupModals.ts:450-460 + if (maxAmount) { + const parsed = parseFloat(maxAmount); + if (Number.isFinite(parsed) && parsed >= 0) { + permission.constraints.max_amount = parsed; + } + } + ``` +- suggested fix: push a message into `unrepresentablePermissionErrors` when the box is non-empty but does not parse, so `saveGroup` refuses rather than sending an uncapped permission. +- verdict: PLAUSIBLE — the skip is real and silent (frontend/src/groups/groupModals.ts:450-460), but the input is `type="number" min="0"` (groupModals.ts:269) inside `#group-form`, which is submitted by a genuine submit-button click (index.html:1093), so native constraint validation blocks the reachable malformed cases: a negative fails `min`, and non-numeric text leaves `value === ""`. Reaching the branch needs a value that passes validation yet fails `Number.isFinite`, e.g. an exponent overflow like `1e400`, which I could not establish from source. +- severity-adjusted: low — the common malformed inputs never reach the branch; only an exponent-overflow value slips past the input's own validation. +- issue: (pending cross-reference) + +### A11-018 A drifted API-keys response shape renders the reassuring "No API keys yet" empty state +- category: silent-fallback +- severity: low +- location: frontend/src/apikeys.ts:51 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the triple fallback ends in `?? []`. If the backend ever returns a shape carrying neither `api_keys` nor a bare array (a wrapped envelope, a partial response), the list resolves to empty and `renderApiKeysList` paints "No API keys yet — Create an API key to let automation tools call CUDly programmatically". An operator reviewing which keys exist before revoking credentials is shown "none" for a deployment that has active keys. The catch branch already has a correct error rendering path that this case bypasses. +- evidence: + ```typescript + // apikeys.ts:50-55 + const response = await api.getApiKeys(); + const list = (response as { api_keys?: APIKeyInfo[] } | undefined)?.api_keys + ?? (Array.isArray(response) ? response as APIKeyInfo[] : undefined) + ?? []; + currentApiKeys = Array.isArray(list) ? list : []; + ``` +- suggested fix: treat "neither known shape" as an error and route it through `renderApiKeysListError`, keeping `[]` only for a genuine empty `api_keys` array. +- verdict: PLAUSIBLE — the fallback chain does end in `?? []` and feeds the reassuring empty state (frontend/src/apikeys.ts:50-55), and it composes with A11-009 so a body-stripped 200 arrives as `null` and renders "No API keys yet" instead of the error path at apikeys.ts:69-75; the trigger is an off-contract response from `/api-keys` (api/apikeys.ts:16-18), which I could not produce from source. +- issue: (pending cross-reference) + +### A12-055 Absent money values filter as `0` while the cell renders `--` +- category: silent-fallback +- severity: low +- location: frontend/src/history.ts:951 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The block comment at lines 920-923 states the contract that "a user typing the visible '$123' matches the row that displays that exact value". `formatCurrency(null)` renders `--`, but `purchaseHistoryNumericCellValue` coerces null to `0`, so typing `0` into the Upfront Cost or Monthly Savings filter returns rows whose cells show `--`, conflating "not reported" with "actually free". The Approval Queue extractor gets this right for `monthly_cost` (line 1661, returns `NaN`) but not for the other two. `plans.ts:247` has the same `?? 0` shape. +- evidence: + ```ts + case 'count': return p.count ?? 0; + case 'upfront_cost': return p.upfront_cost ?? 0; + case 'savings': return p.estimated_savings ?? 0; + ``` +- suggested fix: Return `Number.NaN` for absent money values in both extractors, matching the approval queue's `monthly_cost` case. +- verdict: CONFIRMED — formatCurrency renders `--` for null (frontend/src/utils.ts:44-46) but the extractor coerces count/upfront_cost/savings to 0 (frontend/src/history.ts:951-953), against the stated contract at :920-923; the approval-queue extractor returns NaN for the analogous monthly_cost case (frontend/src/history.ts:1660). +- issue: (pending cross-reference) + +### A12-057 Marketplace dialog silently degrades to a no-amount consent when the row lookup misses +- category: silent-fallback +- severity: low +- location: frontend/src/history.ts:1451 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The whole price summary is wrapped in `if (purchase)`. When `lastPurchases.find` returns `undefined` (a stale render, or the row scrolled out of the fetched window after A12-029), the dialog renders only the generic fee note with no RI id, no remaining term, no list price and no net proceeds, and "Confirm listing" still submits. The user authorises a marketplace sale with no amount on screen. +- evidence: + ```ts + const purchase = lastPurchases.find(p => p.purchase_id === id); + const bodyEl = document.createElement('div'); + bodyEl.className = 'marketplace-pricing-modal-body'; + if (purchase) { + // entire price summary + } + ``` +- suggested fix: Abort with an error toast when the lookup fails; never open a money-authorising dialog without an amount. +- verdict: CONFIRMED — The entire price summary is inside `if (purchase)` (frontend/src/history.ts:1451-1512) while the confirmDialog and the createMarketplaceListing call sit outside it (frontend/src/history.ts:1520-1535), so a missed lastPurchases lookup opens a money-authorising dialog carrying only the generic fee note. +- issue: (pending cross-reference) + +### A12-060 `total_completed` falls back to the total row count +- category: silent-fallback +- severity: low +- location: frontend/src/history.ts:405 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: If the API supplies `total_pending` and `total_purchases` but omits `total_completed` — the partial-deploy case the surrounding comment is written for — `completed` becomes the full row count and the line reads "10 completed · 3 pending" for a dataset with 7 completed. The lines immediately above go out of their way to render `--` rather than fabricate money values, then this one fabricates a count. +- evidence: + ```ts + const completed = summary.total_completed ?? total; + const pending = summary.total_pending ?? 0; + const detail = (total !== null && pending > 0) + ? `

${completed} completed · ${pending} pending

` + : ''; + ``` +- suggested fix: Render the detail line only when `summary.total_completed != null`. +- verdict: CONFIRMED — `const completed = summary.total_completed ?? total` fabricates the count from the row total (frontend/src/history.ts:405) immediately after the same function renders `--` rather than fabricate money values (frontend/src/history.ts:371-390). +- issue: (pending cross-reference) + +### A12-064 `collectTargets` silently coerces an empty or sub-1 exchange count to 1 +- category: silent-fallback +- severity: low +- location: frontend/src/riexchange.ts:1757 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A user who clears the Count field intending to retype it, or types `0`, and clicks Get Quote gets a quote for a single unit instead of an inline validation error. The same coercion in `updateRunningTotal` (line 1680) makes the displayed total agree with the substitution, so nothing signals the typed value was discarded. The offering-id branch four lines above correctly returns an error for the analogous empty case. +- evidence: + ```ts + const rawCount = parseInt(row.countInput.value, 10); + const targetCount = isNaN(rawCount) || rawCount < 1 ? 1 : rawCount; + targets.push({ offering_id: offeringId, count: targetCount }); + ``` +- suggested fix: Return `{ targets: [], error: 'Target N: count must be a whole number >= 1.' }`, matching the offering-id check. +- verdict: CONFIRMED — collectTargets coerces an empty or sub-1 count to 1 (frontend/src/riexchange.ts:1756-1758) four lines after the offering-id branch returns an explicit error for the analogous empty case (:1747-1749), and updateRunningTotal repeats the same coercion so the display agrees with the substitution (frontend/src/riexchange.ts:1679-1681). +- issue: (pending cross-reference) + +### A12-065 RI utilization failures are swallowed and the column stays "..." forever +- category: silent-fallback +- severity: low +- location: frontend/src/riexchange.ts:324 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: When `getRIUtilization` fails or is throttled, the catch logs and returns, leaving `currentUtilization` empty so every row renders a permanent loading ellipsis. The operator cannot distinguish "Cost Explorer is slow" from "utilization is unavailable", and an active `utilization_pct` column filter silently matches zero rows because the extractor returns `NaN` for every row. +- evidence: + ```ts + } catch (error) { + console.error('Failed to load RI utilization:', error); + } + ``` +- suggested fix: Record the failure in module state and render "n/a" with the error in a `title`, instead of a permanent ellipsis. +- verdict: CONFIRMED — The catch logs and returns without recording the failure (frontend/src/riexchange.ts:324-326), leaving currentUtilization empty so every row renders the permanent loading ellipsis at frontend/src/riexchange.ts:378. +- issue: (pending cross-reference) + +### A12-075 The summary card's "absent savings" branch is unreachable, so missing savings render as $0 +- category: silent-fallback +- severity: medium +- location: frontend/src/recommendations.ts:800 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `pageLevelRange` initialises `savingsMin`/`savingsMax` to `0` and only ever `+=` into them, so both are always numbers; a rec whose `savings` arrives as null contributes `0` through `cellSummary`. The `plr.savingsMin === null || plr.savingsMax === null` test therefore never fires and the `'--'` the comment above it promises is dead. Execution falls to the final branch, `formatCostForPeriod(0, period)`, so a page whose savings data is entirely absent renders "Potential Monthly Savings $0" — the fabricated zero the M-1 note says it prevents. The `?? null` on lines 795-796 is likewise a no-op for the same reason. +- evidence: + ```ts + const savingsText = hasRecs && plr.savingsMax > 0 && scaledSavingsMin !== null && scaledSavingsMax !== null + ? formatScaledRange(scaledSavingsMin, scaledSavingsMax, period) + : hasRecs && (plr.savingsMin === null || plr.savingsMax === null) + ? '--' + : formatCostForPeriod(0, period); + ``` +- suggested fix: Have `cellSummary`/`pageLevelRange` propagate `null` when no variant carried a savings figure (skipping nulls the way `approval-details.ts:104` does), so the `'--'` branch can actually fire. +- verdict: CONFIRMED — pageLevelRange initialises savingsMin/savingsMax to 0 and only ever `+=` into them (frontend/src/recommendations.ts:1073-1082), so the null test at frontend/src/recommendations.ts:802 can never fire, the `'--'` branch is dead and the `?? null` at :795-796 is a no-op. +- severity-adjusted: low — Recommendation.savings is non-nullable in the API contract (frontend/src/api/types.ts:148), so the fabricated zero needs the backend to violate its own schema; with conforming data only the dead branch remains +- issue: (pending cross-reference) + +### A13b-011 Registration gate ignores `contact_email` in all four Terraform modules despite the message claiming otherwise +- category: silent-fallback +- severity: low +- location: iac/federation/aws-target/terraform/registration.tf:2 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `local.do_register` tests only `var.cudly_api_url`, but the payload always carries `contact_email`, and the CloudFormation equivalent gates on both (`iac/federation/aws-cross-account/cloudformation/template.yaml:53-55`, `DoRegister` is an `Fn::And` over URL and email). A customer who sets the URL and forgets the email registers an account with an empty contact, so CUDly has no one to notify about a pending purchase, and the skip message printed by every `registration_response` output ("Skipped (cudly_api_url or contact_email not set)") states a condition the code never evaluates. Identical in the `aws-cross-account`, `azure-target`, `gcp-sa-impersonation` and `gcp-target` registration files. +- evidence: + ```hcl + locals { + do_register = var.cudly_api_url != "" + ``` +- suggested fix: Make the local `var.cudly_api_url != "" && var.contact_email != ""` in all five files, or drop the email clause from the output messages so the two agree. +- verdict: CONFIRMED — `do_register = var.cudly_api_url != ""` appears verbatim in all five registration files (`aws-target:2`, `aws-cross-account:4`, `azure-target:2`, `gcp-sa-impersonation:2`, `gcp-target:2`) while every payload still sets `contact_email` and every `registration_response` output prints "Skipped (cudly_api_url or contact_email not set)"; the CloudFormation counterpart really does gate on both (`iac/federation/aws-cross-account/cloudformation/template.yaml:52-55`, `DoRegister` as an `Fn::And`), so the Terraform side is the one that disagrees with its own message. +- issue: (pending cross-reference) + +### A13c-019 `secrets/aws` accepts an empty-string database password and stores it, where the admin password in the same file rejects it +- category: silent-fallback +- severity: low +- location: terraform/modules/secrets/aws/main.tf:51 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `random_password.database` is gated on `var.database_password == null`, and + the version body picks the variable whenever it is `!= null`. Passing `""` therefore skips + generation *and* selects the empty string, writing `{"username":"cudly","password":""}` into + Secrets Manager and, through `database/aws`, into `aws_db_instance.password`. The admin password + eight lines below handles precisely this (`var.admin_password == null || var.admin_password == + ""` on the generator, and the matching two-part test on the version), so the asymmetry is inside + one file. +- evidence: + ```hcl + resource "random_password" "database" { + count = var.database_password == null ? 1 : 0 + } + + resource "aws_secretsmanager_secret_version" "database_password" { + secret_string = jsonencode({ + password = var.database_password != null ? var.database_password : random_password.database[0].result + }) + } + ``` +- suggested fix: use the same `== null || == ""` test on both the generator count and the version + value, matching `admin_password`. +- verdict: PLAUSIBLE — the asymmetry is exactly as described: `random_password.database` gates on + `== null` alone (secrets/aws/main.tf:25) and the version body on `!= null` (:51), while + `random_password.admin_password` (:60) and its version (:85) both carry the `|| == ""` test, and + `database_password` has no validation (variables.tf:34-39). The scenario needs a caller passing + `""`; the sole in-tree caller passes `null` (environments/aws/secrets.tf:14). +- issue: (pending cross-reference) + +### A14-042 The ECR guard omits the empty-string check its RDS sibling documents as the decisive one +- category: silent-fallback +- severity: low +- location: scripts/force-delete-owned-ecr-repo.sh:75 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `jq -er` accepts an empty string (it is neither `null` nor `false`), so a state publishing `"ecr_repository_name": {"value": ""}` yields `OWNED_REPO=""`. The script then echoes "This state owns ECR repository ''" and runs `aws ecr describe-repositories` before `select-owned-name.sh` refuses with exit 2. `disable-owned-rds-deletion-protection.sh:120-137` adds an explicit check for exactly this row of its own measured table, with a comment explaining that the selector's refusal comes too late and names the wrong problem. The ECR script, which shares the selector and the same failure mode, has no equivalent branch. +- evidence: + ```bash + OWNED_REPO="$(jq -er '.ecr_repository_name.value' <<<"$OUTPUTS_JSON")" + echo "This state owns ECR repository '$OWNED_REPO'" + ``` +- suggested fix: Add the same `case "$OWNED_REPO" in '' | *[![:graph:]]*)` guard the RDS script uses, before the first AWS call. +- verdict: CONFIRMED — scripts/force-delete-owned-ecr-repo.sh:75-78 takes `jq -er '.ecr_repository_name.value'`, echoes it and calls `aws ecr describe-repositories` with no empty or whitespace check, while disable-owned-rds-deletion-protection.sh:126-137 carries exactly the `case "$OWNED_INSTANCE" in '' | *[![:graph:]]*)` guard before its first AWS call. +- issue: (pending cross-reference) + +### Category: correctness + +118 findings: 2 critical, 12 high, 59 medium, 45 low. + +### A11-001 The group create/edit form's submit handler is never wired in production +- category: correctness +- severity: critical +- location: frontend/src/groups/handlers.ts:17 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `setupGroupHandlers()` is the only code that binds `saveGroup` to `#group-form`, and nothing calls it at runtime. Its sole caller is `users.ts:setupHandlers()` (frontend/src/users.ts:40), which has no callers anywhere outside tests. `app.ts:setupEventListeners()` calls `setupUserHandlers()` (frontend/src/app.ts:176) but never the group equivalent, and `index.ts`'s DOMContentLoaded block wires `create-group-btn`, `close-group-modal-btn`, `add-permission-btn` and `group-duplicate-form` but not `group-form` (frontend/src/index.ts:43-50). `#group-form` in index.html:1076 carries `` and no inline `onsubmit`. So clicking "Save Group" in the Create Group / Edit Group modal runs a native form submission against the current URL instead of `saveGroup`: the page navigates/reloads and the group is never created or updated. Every permission edit made through the Groups panel is silently lost. `groups.test.ts` does not catch this because it calls `groupModals.saveGroup(event)` directly (groups.test.ts:560) and calls `groupHandlers.setupGroupHandlers()` itself (groups.test.ts:933) — the suite proves the handler works, never that the app installs it. +- evidence: + ```typescript + // groups/handlers.ts:10-21 — the only binder for #group-form + export function setupGroupHandlers(): void { + (window as any).openCreateGroupModal = openCreateGroupModal; + (window as any).closeGroupModal = closeGroupModal; + (window as any).addPermission = () => addPermission(); + const groupForm = document.getElementById('group-form'); + if (groupForm) { + groupForm.addEventListener('submit', (e) => void saveGroup(e)); + } + } + ``` +- suggested fix: call `setupGroupHandlers()` from `setupEventListeners()` in app.ts next to `setupUserHandlers()`, and add a test that runs the real init path and asserts a `submit` on `#group-form` reaches `api.updateGroup`. +- verdict: CONFIRMED — `setupGroupHandlers` has exactly three non-test references (definition at frontend/src/groups/handlers.ts:10, barrel re-export at frontend/src/groups/index.ts:27, dynamic import inside the uncalled `setupHandlers` at frontend/src/users.ts:43), `#group-form` appears in only handlers.ts:17 and groupModals.ts:26/53 with no delegated `submit` listener anywhere (the only non-test `addEventListener('submit')` sites are ladder.ts:395, plans.ts:2190, settings.ts:2716, app.ts:155/160, apikeys.ts:394, users/handlers.ts:27, riexchange.ts:2003, index.ts:50, auth.ts:136/470/600/774/1060) and no `onsubmit` in index.html, and groups.test.ts:917-976 installs the binding itself so nothing asserts the app does. +- issue: (pending cross-reference) + +### A12-003 RI-exchange execute never sends `region`, which the backend rejects with 400 +- category: correctness +- severity: critical +- location: frontend/src/riexchange.ts:1809 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The user gets a valid quote (the quote handler resolves the region from the SDK chain, so it succeeds) and clicks "Execute Exchange". The POST body omits `region`; `validateExecuteExchangeBody` at `internal/api/handler_ri_exchange.go:1756` returns `400 "region is required for execute; omitting it risks exchanging RIs in the wrong region"`. Every UI-initiated exchange fails at the last step, and the failure reads as a backend fault. `ExchangeExecuteRequest.region` is optional in the TypeScript type, so there is no compile-time signal. +- evidence: + ```ts + const result = await api.executeExchange({ + ri_ids: modalQuoteReq.ri_ids, + targets: modalQuoteReq.targets, + target_offering_id: modalQuoteReq.target_offering_id, + target_count: modalQuoteReq.target_count, + max_payment_due_usd: modalQuote.PaymentDueRaw, + }); + ``` +- suggested fix: Carry a region into the modal (from the source RI's availability zone, or have the quote response echo the region the backend resolved) and send the same value on quote and execute. +- verdict: CONFIRMED — submitModalExecute posts exactly five fields with no region (frontend/src/riexchange.ts:1808-1814), executeExchange forwards the body verbatim (frontend/src/api/riexchange.ts:91-96), and validateExecuteExchangeBody returns 400 on an empty Region (internal/api/handler_ri_exchange.go:1756-1758) before executeExchange reads it at :1793. +- issue: (pending cross-reference) + +### A01-003 "Run now" strands the execution in `running`; nothing executes it and the reaper marks it failed +- category: correctness +- severity: high +- location: internal/api/handler_purchases.go:320 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Operator clicks Run now on a planned purchase. `runPlannedPurchase` CAS-transitions pending->running and returns "Purchase execution initiated" (frontend toasts "Purchase executed successfully", frontend/src/plans.ts:782-783). No executor consumes `running` rows: the scheduler fires only `GetPendingExecutions` (status IN pending,notified; store_postgres.go:1479), `claimAndExecute` claims from approved/pending/notified (internal/purchase/manager.go:179), and `purchase.Reaper` flips `running` rows older than 10 minutes to `failed` (reaper.go:25). The row also drops out of the Planned list (`plannedListStatuses` = pending/notified/paused, line 95). Net effect: the operator is told the purchase ran, nothing is bought, and 10 minutes later the row reads "failed". `TestHandler_runPlannedPurchase` (handler_purchases_test.go:1245) only asserts the status flip. +- evidence: + ```go + if _, err := h.config.TransitionExecutionStatus(ctx, executionID, []string{"pending", "paused"}, "running", resolveCreatorUserID(session)); err != nil { + return nil, NewClientError(409, fmt.Sprintf("execution %s cannot be started: %v", executionID, err)) + } + + return map[string]any{ + "execution_id": executionID, + "status": "running", + "message": "Purchase execution initiated", + }, nil + ``` +- suggested fix: Either hand the claimed row to the purchase manager synchronously (the same `ApproveAndExecute`/`claimAndExecute` funnel the approve path uses, with the 4-eyes and constraint gates) or transition to `approved` and enqueue the execute message; do not report success from a bare status flip. +- verdict: CONFIRMED — runPlannedPurchase (internal/api/handler_purchases.go:320-328) CAS-flips to running and returns; the only executors are ProcessScheduledPurchases → GetPendingExecutions (WHERE status IN pending,notified; internal/config/store_postgres.go:1479), FireScheduledDelayedPurchases → GetScheduledExecutionsDue (internal/purchase/scheduled_fire.go:45), and the reaper whose stuckStatuses include running (internal/purchase/reaper.go:25); frontend/src/plans.ts:782 toasts "Purchase executed successfully" on the bare 200. +- issue: (pending cross-reference) + +### A02-002 setupAdmin stores the base64-encoded password, so the bootstrap admin cannot log in with the password they typed +- category: correctness +- severity: high +- location: internal/api/handler_auth.go:238 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The frontend base64-encodes the setup-admin password (frontend/src/api/auth.ts:285, `password: base64Encode(password)`; `base64Encode` is `btoa`, client.ts:191). The handler forwards `setupReq` untouched, the adapter copies `req.Password` verbatim (internal/server/app.go:1023-1027), and `Service.SetupAdmin` hashes it (internal/auth/service_user.go:33-38). `login` (handler_auth.go:34-38) decodes base64 before `auth.Login`, so bcrypt compares the plaintext against a hash of the base64 text and every login of the first admin fails. base64_password_guard_test.go:58-63 exempts `setupAdmin` with the justification that "the auth service handles hashing internally", which does not address the encoding mismatch. Verified by reading the three layers; not executed. If the deployed flow works, something outside these files must be re-encoding, and I could not find it. +- evidence: + ```go + var setupReq SetupAdminRequest + if err := json.Unmarshal([]byte(req.Body), &setupReq); err != nil { + return nil, NewClientError(400, "invalid request body") + } + response, err := h.auth.SetupAdmin(ctx, setupReq) + ``` +- suggested fix: Decode with `decodeBase64Password` in `setupAdmin` exactly as `login`/`resetPassword` do, drop the `knownExemptFunctions` entry, and add an end-to-end test that creates the admin through the handler and then logs in through `login` with the same frontend-encoded body. +- verdict: CONFIRMED — frontend/src/auth.ts:526 calls api.setupAdmin which sends base64Encode(password) (frontend/src/api/auth.ts:285); setupAdmin (handler_auth.go:238-249) unmarshals and forwards without decodeBase64Password, the adapter copies req.Password verbatim (internal/server/app.go:1023-1027), SetupAdmin hashes it (internal/auth/service_user.go:33-38), and login (handler_auth.go:34-38) decodes before auth.Login, so the plaintext can never match; no decode exists anywhere in internal/auth (grep base64 finds only API-key encoding) and no e2e test covers setup-then-login. +- issue: (pending cross-reference) + +### A07-006 ElastiCache offering lookup sends the raw CE engine string as ProductDescription, unvalidated +- category: correctness +- severity: high +- location: providers/aws/services/elasticache/client.go:311 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `parseElastiCacheDetails` (parser_services.go:70) copies CE's `ProductDescription` verbatim into `CacheDetails.Engine`. `resolveEC2Tenancy` in that same file documents that CE returns these values title-cased, so a real rec carries `"Redis"`, while the ElastiCache offering filter matches on the lowercase product description. The filter matches nothing and the purchase fails as "no offerings found" rather than naming the real cause. When CE omits the field entirely, `details.Engine` is `""` and the code sends an explicit empty-string filter — a positive filter for an empty product description — instead of leaving it unset or erroring. RDS handles the identical problem with `normalizeEngineName` (rds/client.go:573); ElastiCache has no equivalent. +- evidence: + ```go + input := &elasticache.DescribeReservedCacheNodesOfferingsInput{ + CacheNodeType: aws.String(rec.ResourceType), + ProductDescription: aws.String(details.Engine), + Duration: aws.String(duration), + OfferingType: aws.String(offeringType), + ``` +- suggested fix: normalise `details.Engine` to the lowercase `redis` / `memcached` values and return an explicit error when it is empty or unrecognised, mirroring `normalizeEngineName`. +- verdict: PLAUSIBLE — the raw pass-through is real (providers/aws/recommendations/parser_services.go:69-71 copies CE's `ProductDescription` verbatim, providers/aws/services/elasticache/client.go:311 sends it unchanged, and that file contains no `strings.ToLower` or normalise helper at all, unlike rds/client.go:573), but the failure needs two runtime facts I could not establish from source: that CE actually returns `"Redis"` title-cased for ElastiCache (the title-case evidence at parser_services.go:83-85 is about the EC2 tenancy field), and that AWS treats an empty `ProductDescription` as a match-nothing filter rather than ignoring it. +- issue: (pending cross-reference) + +### A07-010 Daily-sparkline coverage call sets Granularity together with GroupBy, which the API rejects +- category: correctness +- severity: high +- location: providers/aws/recommendations/usage_history.go:85 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the `GetReservationCoverage` SDK operation doc states "If GroupBy is set, Granularity can't be set." `fetchDailyCoverage` sets both, so every real call returns a ValidationException. `AttachDailyUsageHistory` logs the error at WARN and continues, so the entire per-row usage-history feature is dead in production while the collection looks healthy. No test asserts `Granularity` on this input (the three `GranularityDaily` assertions in the package are in ondemand_series_test.go and sp_coverage_test.go, whose operations do accept it), which is why mock-backed tests stay green. The sibling path in coverage.go correctly omits `Granularity` when it sets `GroupBy`. +- evidence: + ```go + input := &costexplorer.GetReservationCoverageInput{ + TimePeriod: &types.DateInterval{...}, + Granularity: types.GranularityDaily, + GroupBy: []types.GroupDefinition{ + {Type: types.GroupDefinitionTypeDimension, Key: aws.String(string(types.DimensionInstanceType))}, + }, + Filter: dailyUsageFilter(serviceFilter, region), + Metrics: []string{"Hour"}, + } + ``` +- suggested fix: drop the `Granularity` field (the response is already per-day when the time period is daily-resolvable), and assert its absence in the test so the pair cannot be reintroduced. +- verdict: CONFIRMED — providers/aws/recommendations/usage_history.go:85-91 sets both, and costexplorer@v1.63.1/api_op_GetReservationCoverage.go:118 states verbatim "If GroupBy is set, Granularity can't be set"; the sibling inputs at coverage.go:216 and :241 omit it, usage_history_test.go asserts nothing about `Granularity`, and the live path runs on every refresh via client.go:339. +- issue: (pending cross-reference) + +### A08-004 Azure VM SKU enrichment is dead on the recommendations path and burns a full SKU-catalogue walk per subscription +- category: correctness +- severity: high +- location: providers/azure/services/compute/client.go:914 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `RecommendationsClientAdapter` constructs the compute client with an empty region (`newComputeClientFn(r.cred, r.subscriptionID, "")`, providers/azure/recommendations.go:158). `populateVMSKUMapFromPage` then filters every SKU through `isAvailableInRegion(sku, "")`, whose `strings.EqualFold(*location, "")` is false for every real location, so the catalogue map is always empty and `cachedSKULookup` always returns ok=false. Every VM recommendation ships with `Details.VCPU == 0` and `Details.MemoryGB == 0`, while the client still pages up to `maxSKUPages` (20) pages of `ResourceSKUs` per subscription and discards 100% of them. With the new org-wide fan-out this repeats once per subscription per sweep. +- evidence: + ```go + if !c.isAvailableInRegion(sku, c.region) { + continue + } + ``` +- suggested fix: skip the region filter when `c.region == ""` (the subscription-wide recommendations call is region-agnostic by design, see the GetRecommendations doc comment), or skip the catalogue fetch entirely on that path so the wasted pagination goes away. +- verdict: CONFIRMED — providers/azure/recommendations.go:158 constructs the client with region "", `isAvailableInRegion` (compute/client.go:667) EqualFolds every real location against "" and returns false, so `fetchSKUCatalogue` (client.go:871) pages the catalogue and returns an empty map that `cachedSKULookup` (client.go:846) can never hit, leaving Details.VCPU/MemoryGB at 0 in convertAzureVMRecommendation (client.go:813). +- issue: (pending cross-reference) + +### A08b-017 GCP `ResourceType` is a resource instance name, then used as a pricing tier and validated against tier constants +- category: correctness +- severity: high +- location: providers/gcp/services/memorystore/client.go:456 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same helper at `cloudsql/client.go:470`, `cloudstorage/client.go:462`) +- failure scenario: `extractGCPResourceType` returns the last path segment of the first operation resource — for Memorystore that is `projects/p/locations/l/instances/my-redis`, so `ResourceType` becomes `my-redis`; for Cloud Storage it becomes the bucket name. That value is then passed to `skuMatchesTier` as the tier to substring-match against SKU descriptions (never matches, so pricing always fails) and to `ValidateOffering`, which compares it against the constant list `{"BASIC","STANDARD_HA"}` at line 311 and always returns `invalid Memorystore tier: my-redis`. `computeengine` fixed exactly this class at `client.go:1197` by requiring a `/machineTypes/` path; the three other clients did not. +- evidence: + ```go + for _, op := range opGroup.Operations { + if op.Resource == "" { + continue + } + parts := strings.Split(op.Resource, "/") + if len(parts) > 0 { + return parts[len(parts)-1] + } + } + ``` +- suggested fix: select the operation whose resource path names the type being priced (mirroring `machineTypeFromResourcePath`) and return an error when none is present, instead of taking the first segment available. +- verdict: CONFIRMED — `extractGCPResourceType` returns the last path segment of the first non-empty operation resource (memorystore:456-472, cloudsql:470-486, cloudstorage:462-478); that value is passed to `skuMatchesTier` (memorystore:435-452) through `fillRedisPricing` (491-496) on the live conversion path, and to `ValidateOffering` (254-266) against the two-entry list at 311-314. `computeengine` requires a `/machineTypes/` segment at client.go:1197-1228, exactly the pattern the finding cites. +- issue: (pending cross-reference) + +### A08b-029 Compute Engine existing commitments report `ResourceType` as "VCPU" instead of a machine type +- category: correctness +- severity: high +- location: providers/gcp/services/computeengine/client.go:524 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `commitment.Resources` is the list of `ResourceCommitment` entries a CUD is composed of, and each entry's `Type` is the enum `VCPU`, `MEMORY`, `LOCAL_SSD` or `ACCELERATOR`. Taking index zero writes `com.ResourceType = "VCPU"` for every commitment. Recommendations set `ResourceType` to a machine type such as `n2-standard-8` (`extractResourceTypeFromRecommendation`, line 1206), so no existing commitment can ever match a recommendation by resource type — dedupe and coverage silently see zero overlap and re-recommend capacity the project already owns. The commitment's own `Type` field (`GENERAL_PURPOSE_N2`), which does identify the covered family, is discarded. +- evidence: + ```go + if len(commitment.Resources) > 0 { + resource := commitment.Resources[0] + if resource.Type != nil { + com.ResourceType = *resource.Type + } + } + ``` +- suggested fix: map `commitment.Type` (the `Commitment_Type` enum) to the machine family it discounts, the inverse of `machineFamilyCommitmentType`, and carry the vCPU/memory amounts in `Count` and the details struct. +- verdict: CONFIRMED — computeengine:524-529 assigns `Resources[0].Type`, which the repo's own `ResourceCommitment` doc at 535-538 states is the `"VCPU"` or `"MEMORY"` enum, while the recommendation side sets `ResourceType` to a machine type via `extractResourceTypeFromRecommendation` (1197-1222); `commitment.Type` is never read in the converter. +- issue: (pending cross-reference) + +### A12-008 The AWS "Savings Plans" optgroup is hidden and disabled whenever a provider is selected +- category: correctness +- severity: high +- location: frontend/src/plans.ts:2326 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `#plan-provider` has no empty option, so its default value is `aws` and `setupPlanHandlers()` runs `updateServiceDropdownForProvider('aws')` at app init. The match is `optgroup.label.toLowerCase().includes(provider.toLowerCase())`; the shipped markup labels the SP group `"Savings Plans"` (`frontend/src/index.html:853`), which does not contain the substring `aws`. The group therefore gets `hidden` plus `disabled`, and all four AWS Savings Plans options become unselectable in the Plan modal for the entire session. The sibling optgroup carries `id="plan-service-aws-sp"`, which the code ignores. +- evidence: + ```ts + optgroups.forEach(optgroup => { + const optgroupLabel = optgroup.label.toLowerCase(); + const shouldShow = optgroupLabel.includes(provider.toLowerCase()); + optgroup.classList.toggle('hidden', !shouldShow); + optgroup.disabled = !shouldShow; + ``` +- suggested fix: Match on a `data-provider` attribute (or the existing `id` prefix) instead of a substring of the human-readable label, and add the SP group to the fixture used by the `provider scopes service dropdown` tests. +- verdict: CONFIRMED — index.html labels the group `Savings Plans` (frontend/src/index.html:853), the match is a substring test against the provider slug (frontend/src/plans.ts:2326-2330), and setupPlanHandlers runs it at init against `#plan-provider`'s default `aws` (frontend/src/plans.ts:2299), so all four SP options are hidden and disabled for the session. +- issue: (pending cross-reference) + +### A13b-001 GCP SA-impersonation module grants two roles that cannot bind at project scope and never grants the CUD-purchase permission +- category: correctness +- severity: high +- location: iac/federation/gcp-sa-impersonation/terraform/main.tf:27 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A customer who follows the SA-impersonation path grants CUDly's target service account `roles/commerceorgpolicy.commitmentAdmin` and `roles/billing.viewer` through `google_project_iam_member`. Both are organization- / billing-account-scoped roles, so the bindings either 400 at apply or confer nothing at project scope. Neither role carries `compute.commitments.create`, `compute.commitments.update` or `compute.regions.list`, which are the permissions the code actually needs: `providers/gcp/services/computeengine/client.go:708` calls `RegionCommitments.Insert` and `providers/gcp/recommendations.go:127` warns that region listing needs `roles/compute.viewer`. The onboarding reports success and every commitment purchase 403s. The sibling module states the rule in its own docs at `iac/federation/gcp-target/terraform/variables.tf:198` ("GCP's built-in commitment/billing roles (e.g. roles/commerceorgpolicy.commitmentAdmin, roles/billing.viewer) are organization- or billing-account-scoped and will 400 if granted at project scope") and grants `roles/compute.viewer` plus a custom role holding `compute.commitments.create`/`.update` instead. +- evidence: + ```hcl + resource "google_project_iam_member" "cudly_commitment_admin" { + project = var.project_id + role = "roles/commerceorgpolicy.commitmentAdmin" + member = "serviceAccount:${var.service_account_email}" + } + + resource "google_project_iam_member" "cudly_billing_viewer" { + project = var.project_id + role = "roles/billing.viewer" + member = "serviceAccount:${var.service_account_email}" + } + ``` +- suggested fix: Mirror `gcp-target`: bind `roles/compute.viewer` plus a project-scoped custom role carrying `compute.commitments.create` and `compute.commitments.update`, and move `roles/billing.viewer` to a `google_billing_account_iam_member` gated on a `billing_account_id` variable, exactly as `terraform/modules/compute/gcp/cloud-run/main.tf:484` already does. The same two roles appear in `internal/iacfiles/templates/gcp-sa-impersonation-cli.sh.tmpl:28` and need the same correction. +- verdict: CONFIRMED — dead grant by missing action: `iac/federation/gcp-sa-impersonation/terraform/main.tf:27,33` binds only `roles/commerceorgpolicy.commitmentAdmin` and `roles/billing.viewer` at project scope, neither of which carries `compute.commitments.create` or `compute.regions.list`, while `providers/gcp/services/computeengine/client.go:708` calls `RegionCommitments.Insert` and `providers/gcp/recommendations.go:127` swallows the 403 from region listing as a Warn and returns an empty slice; the sibling module's own note at `iac/federation/gcp-target/terraform/variables.tf:198-203` states both roles 400 at project scope, and the CLI twin at `internal/iacfiles/templates/gcp-sa-impersonation-cli.sh.tmpl:28-35` hides even that with `|| echo "(role may already be bound or unavailable)"`, so onboarding reports success either way. +- issue: (pending cross-reference) + +### A13b-002 Azure Bicep/ARM assigns the built-in Reservation Purchaser role that the repo documents as insufficient for purchases +- category: correctness +- severity: high +- location: iac/federation/azure-target/bicep/azure-wif.bicep:24 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A customer who onboards through the Bicep/ARM path grants CUDly's service principal only the built-in Reservation Purchaser role at subscription scope. The Terraform equivalent for the same onboarding (`iac/federation/azure-target/terraform/main.tf:69`) instead creates a custom role, and its module documents why at `terraform/modules/iam/azure/cudly-reservation-role/main.tf:17`: the built-in role "lacks reservationOrders/write and calculatePrice/action, plus every Microsoft.BillingBenefits action, which is what caused 403s on production reservation purchases". The Bicep customer therefore reaches a subscription that reads as onboarded, passes deployment, and 403s on every reservation and savings-plan purchase. The two IaC paths for the same access disagree about scope; the Bicep side is the wrong one. +- evidence: + ```bicep + @description('Built-in role definition ID for Reservation Purchaser. Default is the well-known built-in role ID.') + param roleDefinitionId string = 'f7b75c60-3036-4b75-91c3-6b41c27c1689' + ``` +- suggested fix: Have the Bicep template create the same custom role definition (the eleven `Microsoft.Capacity` / `Microsoft.BillingBenefits` actions in `terraform/modules/iam/azure/cudly-reservation-role/main.tf:44-54`) and assign that, so the Bicep, Terraform and `arm/CUDly-CrossSubscription/template.json` definitions stay in lockstep. Regenerate `azure-wif.arm.json` from the corrected Bicep. +- verdict: CONFIRMED — dead grant by missing action, and the Bicep is the wrong side of a three-way drift: `iac/federation/azure-target/bicep/azure-wif.bicep:24,34` assigns only the built-in GUID, while both other expressions of the same access create the eleven-action custom role (`iac/federation/azure-target/terraform/main.tf:69-83` via `terraform/modules/iam/azure/cudly-reservation-role/main.tf:44-54`, and `arm/CUDly-CrossSubscription/template.json:37-57,78`); the Go purchase path posts to `/providers/Microsoft.Capacity/calculatePrice` then `/providers/Microsoft.Capacity/reservationOrders/{id}/purchase` (`providers/azure/services/database/client.go:376`, `providers/azure/services/cache/client.go:345`), needing `calculatePrice/action` and `reservationOrders/write`, the two actions the module's own comment at `terraform/modules/iam/azure/cudly-reservation-role/main.tf:16-20` says the built-in role lacks. +- issue: (pending cross-reference) + +### A13b-004 Cross-account trust principal omits the `/lambda/` path the hub role actually carries +- category: correctness +- severity: high +- location: cloudformation/stacks/CUDly-CrossAccount/template.yaml:64 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The hub stack creates its Lambda execution role with `Path: /lambda/` (`cloudformation/stacks/CUDly/template.yaml:397`) and `RoleName: !Sub "${AWS::StackName}-LambdaRole"` (line 389), so its real ARN is `arn:aws:iam:::role/lambda/CUDly-LambdaRole`. The cross-account template builds the trusted principal without the path, producing `arn:aws:iam:::role/CUDly-LambdaRole`, an ARN that names no existing role. IAM rejects a trust policy naming a non-existent IAM principal, so the customer's stack fails with `MalformedPolicyDocument: Invalid principal in policy`; if it is created against a pre-existing decoy role of that name, the hub Lambda can never assume it. The parameter's own description ("set this to the full role name (without path prefix)") tells the operator to reproduce the wrong ARN. +- evidence: + ```yaml + AssumeRolePolicyDocument: + Version: "2012-10-17" + Statement: + - Effect: Allow + Principal: + AWS: !Sub "arn:aws:iam::${PrimaryAccountId}:role/${PrimaryRoleName}" + Action: sts:AssumeRole + ``` +- suggested fix: Add a `PrimaryRolePath` parameter defaulting to `/lambda/` and build the principal as `arn:aws:iam::${PrimaryAccountId}:role${PrimaryRolePath}${PrimaryRoleName}`, or drop `Path: /lambda/` from the hub role so the two agree. Either way the two templates must be changed together. +- verdict: CONFIRMED — dead grant by unmatchable ARN: `cloudformation/stacks/CUDly/template.yaml:386-397` is the only IAM role in the hub stack and carries `RoleName: !Sub "${AWS::StackName}-LambdaRole"` with `Path: /lambda/`, and it is the principal that performs the cross-account assume (`cloudformation/stacks/CUDly/template.yaml:524-528`, `Resource: arn:aws:iam::*:role/CUDly*`), so its real ARN is `.../role/lambda/CUDly-LambdaRole` while `cloudformation/stacks/CUDly-CrossAccount/template.yaml:64` builds `.../role/${PrimaryRoleName}` with no path segment, and the parameter description at lines 21-29 instructs the operator to supply the bare name. +- issue: (pending cross-reference) + +### A14-005 `set -e` makes the entrypoint's migration exit-code handling unreachable +- category: correctness +- severity: high +- location: scripts/entrypoint.sh:72 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The script opens with `set -eu` (line 2). `migrate ... up 2>&1` is a simple command in no errexit-exempt position, so any non-zero exit terminates the shell immediately and `MIGRATE_EXIT_CODE=$?` never runs. Verified: `sh -c 'set -eu; false_cmd; RC=$?; echo RC=$RC'` exits 1 without printing. The entire branch below it, including the documented "exit code 1 means no change, which is okay" tolerance, is dead code. `DB_AUTO_MIGRATE=true` is the Dockerfile default (Dockerfile:176), so on every container start where golang-migrate returns non-zero for a benign reason the container dies with the raw migrate exit code and none of the diagnostics. +- evidence: + ```sh + migrate -path "$DB_MIGRATIONS_PATH" -database "$DB_URL" up 2>&1 + MIGRATE_EXIT_CODE=$? + if [ $MIGRATE_EXIT_CODE -eq 0 ]; then + echo " ✅ Migrations completed successfully" + elif [ $MIGRATE_EXIT_CODE -eq 1 ]; then + echo " ℹ️ No new migrations to apply" + ``` +- suggested fix: `MIGRATE_EXIT_CODE=0; migrate ... || MIGRATE_EXIT_CODE=$?` so the branch below actually receives the status. +- verdict: CONFIRMED — scripts/entrypoint.sh:2 sets `set -eu` and line 72's `migrate ... up` is a simple command in an errexit-live position, so line 73's `MIGRATE_EXIT_CODE=$?` is unreachable on failure (`sh -c 'set -eu; false; RC=$?; echo RC=$RC'` prints nothing and exits 1), and Dockerfile:176 defaults `DB_AUTO_MIGRATE=true`. +- issue: (pending cross-reference) + +### A01-011 Session-authed revoke skips the revocation window and token expiry the token path enforces +- category: correctness +- severity: medium +- location: internal/api/handler_purchases.go:1338 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A user with `cancel-any` (or the creator with `cancel-own`) POSTs `/api/purchases/revoke/{id}` with a session and no token for an execution completed 30 days ago. `tryRevokeViaSession` goes straight to `revokeViaSession`; `checkRevocationWindow` (config.RevocationWindow after CompletedAt) and the expiry check live only in `validateRevokeToken`, which runs only on the token branch (1288). The row flips to `revocation_requested` and the History UI shows a revocation in flight for a commitment whose provider window closed weeks ago. The session tests (`TestHandler_revokePurchase_SessionAdminCancelAny`, 4911) use `buildCompletedExec` with no timestamps, so they cannot observe the missing check. +- evidence: + ```go + case sessErr == nil: + if csrfErr := h.validateCSRF(ctx, req); csrfErr != nil { + return nil, true, NewClientError(403, "CSRF validation failed") + } + res, revokeErr := h.revokeViaSession(ctx, execution, session.Email) + return res, true, revokeErr + ``` +- suggested fix: Call `checkRevocationWindow(execution)` in `revokeViaSession` (shared by both branches) so the window is enforced regardless of how the caller authenticated. +- verdict: CONFIRMED — checkRevocationWindow's only caller is validateRevokeToken (internal/api/handler_purchases.go:1383), which is reached only from the token branch at :1288; tryRevokeViaSession:1338 calls revokeViaSession:1446-1478 directly, which CAS-flips completed/partially_completed → revocation_requested with no CompletedAt/ExecutedAt check. +- issue: (pending cross-reference) + +### A02-003 sell-own marketplace authorization compares the scope against the account UUID with an empty name, but groups store account names +- category: correctness +- severity: medium +- location: internal/api/handler_marketplace.go:438 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The group editor writes account NAMES into `allowed_accounts` (frontend/src/groups/groupModals.ts:578 `cb.value = acct.name`). `AccountScope.Allows(id, name)` matches either, but `authorizeAllowedAccount` passes `""` for the name, so a sell-own user scoped to "Production" is refused 403 on every RI in Production. Every other scope check in the package resolves the name first (`requireAccountAccess`, `filterPurchaseHistoryByAllowedAccounts`). The only tests (handler_marketplace_test.go:164-165, 200-201) put UUID-like values in the scope, so they pass with the bug present. +- evidence: + ```go + scope, err := h.getAccountScope(ctx, session) + ... + if scope.Allows(cloudAccountID, "") { + return nil + } + return NewClientError(403, "permission denied: purchase is in a cloud account not covered by your session's allowed accounts") + ``` +- suggested fix: Replace the body with `_, err := h.requireAccountAccess(ctx, session, cloudAccountID)` (it fetches the account and passes `account.Name`), and change the tests to grant the account by name. +- verdict: CONFIRMED — authorizeAllowedAccount (handler_marketplace.go:438) passes "" as the name so AccountScope.Allows (internal/auth/account_scope.go:104-116) can only match the UUID; the only UI writer of allowed_accounts is the duplicate-group modal (frontend/src/groups/groupModals.ts:558-578, whose comment says "names are what the backend matcher accepts") and it stores acct.name, while requireAccountAccess (scoping.go:26-41) resolves account.Name first; the marketplace tests (handler_marketplace_test.go:164-165,200-201) grant "acct-1"/"acct-other" and never exercise the name branch. +- issue: (pending cross-reference) + +### A02-004 Marketplace list/cancel run against the host's ambient AWS credentials regardless of the row's cloud account +- category: correctness +- severity: medium +- location: internal/api/handler_marketplace.go:169 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A purchase_history row with `CloudAccountID` pointing at a federated member account (role_arn / bastion / WIF) is authorized by `authorizeSessionSell` against that account, but `loadAWSConfigWithRegion` (handler_ri_exchange.go:1267-1276) only sets the region on `getBaseAWSConfig`, the Lambda's own credentials. `CreateReservedInstancesListing` / `CancelReservedInstancesListing` are issued in the host account, where the RI id does not exist, so every multi-account deployment gets `InvalidReservedInstancesId.NotFound` (mapped to 400) and the feature only works for the self-account. The per-account credential resolver already exists (`credentials.ResolveAWSCredentialProvider`, used by `buildOrgRootAWSConfig` at handler_accounts.go:1528). +- evidence: + ```go + cfg, err := h.loadAWSConfigWithRegion(ctx, row.Region) + if err != nil { + return nil, fmt.Errorf("failed to load AWS config: %w", err) + } + ec2Client := h.buildMarketplaceEC2Client(cfg) + ``` +- suggested fix: Resolve the row's `CloudAccountID` to a `config.CloudAccount` and build the config through `credentials.ResolveAWSCredentialProvider` (as `buildOrgRootAWSConfig` does); fail with 400 when the row has no `CloudAccountID` and the deployment is multi-account. +- verdict: CONFIRMED — loadAWSConfigWithRegion (handler_ri_exchange.go:1267-1276) only overrides Region on the ambient getBaseAWSConfig and marketplaceList/Cancel (handler_marketplace.go:169,343) never consult row.CloudAccountID for credentials, while purchases in federated accounts are executed with per-account credentials (internal/purchase/execution.go:448 via ResolveAWSCredentialProviderWithOpts), so the RI exists only in the member account the host credentials cannot see; the only api-package caller of the per-account resolver is buildOrgRootAWSConfig (handler_accounts.go:1528). +- issue: (pending cross-reference) + +### A02-009 A global-defaults PUT overwrites term/payment/coverage/ramp on every service config, including fields the request did not send +- category: correctness +- severity: medium +- location: internal/api/handler_config.go:154 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Operator has `aws/ec2` at term 3 / all-upfront via `PUT /api/config/service/aws/ec2`. Later `PUT /api/config {"default_coverage": 70}` is sent. `anyKeyPresent` is true, `propagateGlobalDefaults` rewrites `svc.Term = cfg.DefaultTerm` (1) and `svc.Payment = cfg.DefaultPayment` on ec2 and every other service, so the next scheduled purchase buys a 1-year commitment instead of the configured 3-year one. The presence guard at line 126 only decides whether to propagate; it does not restrict propagation to the keys actually sent, and failures are logged rather than returned (lines 160-162). +- evidence: + ```go + for i := range services { + svc := &services[i] + svc.Term = cfg.DefaultTerm + svc.Payment = cfg.DefaultPayment + svc.Coverage = cfg.DefaultCoverage + svc.RampSchedule = cfg.DefaultRampSchedule + if saveErr := h.config.SaveServiceConfig(ctx, svc); saveErr != nil { + logging.Warnf(...) + ``` +- suggested fix: Pass the `present` map into `propagateGlobalDefaults` and overlay only the defaults whose key was in the body (the same present-key pattern `mergeServiceConfig` already uses), and return an error if any save fails. +- verdict: CONFIRMED — updateGlobalConfig (handler_config.go:122-128) calls propagateGlobalDefaults whenever anyKeyPresent sees any one of the four keys, and propagateGlobalDefaults (handler_config.go:144-163) overwrites Term/Payment/Coverage/RampSchedule on every ServiceConfig from the merged global cfg regardless of which key was sent, logging (not returning) save failures; the guard's own comment (lines 122-125) states the intent this violates. +- issue: (pending cross-reference) + +### A03-008 A consumed recovery code stays valid when the persist fails, and the comment claims the opposite +- category: correctness +- severity: medium +- location: internal/auth/service.go:255 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `consumeRecoveryCode` removes the hash from the in-memory slice; if `UpdateUser` fails the login is still granted and the store row is unchanged. The next login re-reads the row from the store, so the same code matches again. The comment says "repeated use of the same code on the next login will fail because the slice is stale", which is false: the slice is rebuilt from the store on every login. Reproduced with `UpdateUser` returning an error: `login#1 with recovery code (persist fails): token=true`; `login#2 SAME code: token=true`. Single-use recovery codes become multi-use exactly during a DB outage, which is also when an attacker who obtained one code gets unlimited retries. +- evidence: + ```go + if s.consumeRecoveryCode(user, req.MFACode) { + if err := s.store.UpdateUser(ctx, user); err != nil { + logging.Warnf("Failed to persist recovery-code consumption for user %s: %v", user.ID, err) + // The recovery code already verified; still allow + // login but warn -- repeated use of the same code on + // the next login will fail because the slice is + // stale, which is the safe failure mode. + } + return nil + } + ``` +- suggested fix: Fail the login when the consumption write fails (return `ErrInvalidMFACode` or a 5xx), so a recovery code is only accepted once it is provably burned; delete the incorrect comment. +- verdict: CONFIRMED — verifyPasswordAndMFA (service.go:254-262) returns nil after a failed UpdateUser, Login:158 re-reads the row through getUserAndValidateStatus -> GetUserByEmail (store_postgres.go:60-75) on every attempt, so the in-memory slice the comment relies on is discarded and the unburned hash matches again on the next login. +- issue: (pending cross-reference) + +### A03-014 KMS signers cache the first resolution error forever +- category: correctness +- severity: medium +- location: internal/oidc/aws_signer.go:86 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `resolveOnce` uses `sync.Once`; a transient `GetPublicKey` failure, a throttle, or simply a caller whose `ctx` was already cancelled poisons `s.err`, and every later `KeyID`/`PublicKey`/`Mint` on that signer fails for the process lifetime. The signer is constructed once per process (`NewSignerFromEnv` at startup) and used by every Azure and GCP federated credential exchange, so one bad first call takes down all federated purchases and collection until the Lambda instance is recycled. Same pattern in azure_signer.go:103 and gcp_signer.go:107. +- evidence: + ```go + func (s *AWSKMSSigner) resolveOnce(ctx context.Context) { + s.once.Do(func() { + out, err := s.client.GetPublicKey(ctx, &kms.GetPublicKeyInput{KeyId: &s.keyID}) + if err != nil { + s.err = fmt.Errorf("oidc: kms:GetPublicKey: %w", err) + return + } + ``` +- suggested fix: Only memoize success: guard with a mutex, retry the fetch when `pubKey == nil`, and return the error without recording it (or resolve eagerly in the constructor and fail startup). +- verdict: CONFIRMED — resolveOnce stores the error inside sync.Once.Do (aws_signer.go:86-115, azure_signer.go:103-109, gcp_signer.go:107-113) and PublicKey/KeyID return s.err unconditionally afterwards (:75-84); NewSignerFromEnv is called exactly once at startup (server/app.go:409) with no eager resolve, so the first caller's ctx and the first transient failure decide the signer's state for the process lifetime. +- issue: (pending cross-reference) + +### A04-003 ON DELETE RESTRICT plus a pending-only preflight makes accounts with any terminal execution undeletable +- category: correctness +- severity: medium +- location: internal/config/store_postgres.go:1600-1611 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (constraint at internal/database/postgres/migrations/000053_executions_account_fk_restrict.up.sql:26-30; handler at internal/api/handler_accounts.go:609-655) +- failure scenario: account A has one execution in status `completed`, `failed` or `canceled`. `CountPendingExecutionsForAccount` counts only `pending`/`notified`, so the preflight passes, `DELETE FROM cloud_accounts` hits the RESTRICT FK (which applies to every referencing row regardless of status), and the handler maps SQLSTATE 23503 to a 409 saying "pending purchase(s) must be canceled first" with no IDs. The frontend's Cancel-All-Then-Delete flow only cancels (rows keep `cloud_account_id`), so it can never succeed. Failed rows are never purged by `CleanupOldExecutions`, so an account with one failed purchase is permanently undeletable, and the operator is told to cancel purchases that do not exist. The 000053 test's comment ("the API layer's Cancel-All-And-Delete flow does exactly that", 000053_executions_account_fk_restrict_test.go:84) claims a row deletion the API never performs. +- evidence: + ```go + err := s.db.QueryRow(ctx, ` + SELECT COUNT(*) FROM purchase_executions + WHERE cloud_account_id = $1 + AND status IN ('pending', 'notified') + `, accountID).Scan(&n) + ``` + ```sql + ALTER TABLE purchase_executions + ADD CONSTRAINT purchase_executions_cloud_account_id_fkey + FOREIGN KEY (cloud_account_id) REFERENCES cloud_accounts(id) + ON DELETE RESTRICT; + ``` +- suggested fix: count all referencing rows split by live vs terminal in the preflight, and have `DeleteCloudAccount` detach terminal rows (`UPDATE purchase_executions SET cloud_account_id = NULL WHERE cloud_account_id = $1 AND status NOT IN (live set)`) inside its transaction before the DELETE, so only genuinely live purchases block deletion. +- verdict: CONFIRMED — `CountPendingExecutionsForAccount` counts only `pending`/`notified` (internal/config/store_postgres.go:1600-1611) while the 000053 FK restricts every referencing row regardless of status (internal/database/postgres/migrations/000053_executions_account_fk_restrict.up.sql:23-30); `CancelExecutionAtomic` leaves `cloud_account_id` set (internal/config/store_postgres.go:1158-1166) and `CleanupOldExecutions` excludes `HealthScoredExecutionStatuses` — failed and both canceled spellings — from its only expiry branch (internal/config/store_postgres.go:1873-1890, internal/config/types.go:485-489), so a failed row blocks the delete permanently. The 000053 test comment at line 84 is verified as claiming a row deletion the API never performs. +- issue: (pending cross-reference) + +### A04-008 ClaimMarketplaceListingSlot and the other purchase_id-keyed writers assume purchase_id is unique +- category: correctness +- severity: medium +- location: internal/config/store_postgres.go:1985-1998 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (also 1948-1962, 2356-2365, 2424-2442) +- failure scenario: `purchase_history.purchase_id` has only a non-unique index (000002). `savePurchaseHistory` (internal/purchase/execution.go:893-930) writes one row per successful rec with `PurchaseID = result.CommitmentID`; a stranded execution re-driven under the same idempotency lineage (#1012) gets the existing commitment ID back from the provider dedupe and writes a second row with the same purchase_id. `ClaimMarketplaceListingSlot` then updates both rows to `listing_state='pending'` but returns `RowsAffected() == 1` false, the handler answers 409 and never releases the slot, and every later attempt sees `pending` and 409s forever. `GetPurchaseHistoryByPurchaseID` (`LIMIT 1`, no ORDER BY) and `MarkPurchaseRevoked` pick or stamp an arbitrary one of the duplicates. +- evidence: + ```go + tag, err := s.db.Exec(ctx, query, + ListingStatePending, purchaseID, ListingStateActive, ListingStatePending) + if err != nil { + return false, fmt.Errorf("failed to claim marketplace listing slot for purchase %s: %w", purchaseID, err) + } + return tag.RowsAffected() == 1, nil + ``` +- suggested fix: add a unique index on `purchase_history(purchase_id)` after a dedupe migration (or make `SavePurchaseHistory` an upsert on purchase_id), and return `RowsAffected() > 0` in the claim. +- verdict: CONFIRMED — `purchase_id` carries only the non-unique `idx_purchase_history_purchase_id` (internal/database/postgres/migrations/000002_add_indexes.up.sql:28) and `SavePurchaseHistory` is a plain INSERT with no ON CONFLICT (internal/config/store_postgres.go:2004-2011); a re-drive gets the same commitment id back from the provider dedupe (providers/aws/services/ec2/client.go:129-142, providers/aws/services/savingsplans/client.go:238-256) and `savePurchaseHistory` writes it again (internal/purchase/execution.go:893-930), after which `ClaimMarketplaceListingSlot`'s `RowsAffected() == 1` (internal/config/store_postgres.go:1990-1997) is false forever while both rows sit at `pending`. `GetPurchaseHistoryByPurchaseID`'s `LIMIT 1` with no ORDER BY is at 2355-2365. +- issue: (pending cross-reference) + +### A05-004 Scheduler-created executions have no creator, so 4-eyes makes them permanently unapprovable +- category: correctness +- severity: medium +- location: internal/purchase/notifications.go:130 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `getOrCreateExecution` builds the row with no `CreatedByUserID` (the column stays NULL). The plan-notification email it triggers offers exactly one action, the token deep link, which lands in `ApproveExecution` -> `ApproveAndExecute` -> `checkDifferentApprover`. That function denies any row whose `CreatedByUserID` is nil with "this execution predates the dual-control feature" (approvals.go:255). So with `RequireDifferentApprover` on, every execution the notification sweep creates is un-approvable by any operator, and the only stated remedy in the error text is to turn dual control off globally. The row is not legacy — it is minted by current code on every tick that finds no existing execution for the plan+date. +- evidence: + ```go + execution := &config.PurchaseExecution{ + PlanID: plan.ID, + ExecutionID: uuid.New().String(), + Status: "pending", + StepNumber: plan.RampSchedule.CurrentStep + 1, + ScheduledDate: *plan.NextExecutionDate, + ApprovalToken: approvalToken, + ApprovalTokenExpiresAt: &tokenExpiresAt, + } + ``` +- suggested fix: stamp the plan's owner onto `CreatedByUserID` here so the identity comparison has something to compare against, or treat a system-created row (no creator, no human requester) as satisfying dual control rather than failing it, and test the notification-created row against the 4-eyes gate. +- verdict: CONFIRMED — the struct literal at notifications.go:129-146 sets no CreatedByUserID, and the email's only action is the ApprovalToken deep link, which lands in ApproveExecution (approvals.go:27) -> ApproveAndExecute (approvals.go:313) -> enforceFourEyesPolicy (approvals.go:216) -> checkDifferentApprover, whose first statement (approvals.go:253-256) denies any row with a nil CreatedByUserID and names disabling 4-eyes as the only remedy; the session path denies the same row independently at handler_purchases.go:958. +- issue: (pending cross-reference) + +### A05-005 The recommendation ID's account component is empty for every AWS reservation rec, so IDs collide across accounts +- category: correctness +- severity: medium +- location: internal/scheduler/scheduler.go:1492 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The composite ID names `rec.Account` as the field that "separates per-account/per-subscription recs sharing the same provider+SKU+region+payment". AWS reservation recs only set it when Cost Explorer returns `AccountId` (providers/aws/recommendations/parser_ri.go:72), which is absent for payer-scope recommendations; `parser_services.go` never sets it at all. With two registered AWS accounts both surfacing an m5.large 1yr all-upfront EC2 RI, both rows get the identical id `aws||ec2|us-east-1|m5.large||1|all-upfront`. The DB unique key is `(account_key, provider, service, region, resource_type, engine, term, payment_option)` (migration 000043) and keys on the CUDly account UUID, not `rec.Account`, so both rows persist and `GetRecommendationByID` (scheduler.go:1123) returns `recs[0]` — a deep link to one account's recommendation renders the other account's. It is also the collision the comment cites #187/#188 for on the frontend selection set. +- evidence: + ```go + recordID := fmt.Sprintf("%s|%s|%s|%s|%s|%s|%d|%s", + providerName, rec.Account, rec.Service, rec.Region, + rec.ResourceType, engine, term, rec.PaymentOption) + ``` +- suggested fix: key the ID on the same account dimension the DB unique index uses — the record's `CloudAccountID` (empty for the ambient path) — instead of the provider-reported `rec.Account`, and add a test with two accounts whose recs are otherwise identical asserting distinct IDs. +- verdict: CONFIRMED — `rec.Account` is set only under `if details.AccountId != nil` (providers/aws/recommendations/parser_ri.go:72-73) and `/usr/bin/grep -n Account providers/aws/recommendations/parser_services.go` returns nothing; the ID is built at scheduler.go:1492 BEFORE tagAccount stamps CloudAccountID (fetchAndConvert, scheduler.go:1027-1029, tagAccount at :1035), so two accounts collected via fanOutPerAccount (scheduler.go:495) yield byte-identical IDs. The upsert key at store_postgres_recommendations.go:224 includes account_key so both rows persist, and GetRecommendationByID returns `&recs[0]` (scheduler.go:1138) after a filter on ID alone. +- issue: (pending cross-reference) + +### A05-010 `findAWSAccount` ignores the enabled flag despite its name and doc +- category: correctness +- severity: medium +- location: internal/commitmentopts/service.go:130 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The doc says "returns the first enabled AWS account", but the filter sets only `Provider` — `Enabled` is left nil, so `ListCloudAccounts` returns disabled rows too and the loop takes the first one. An operator who disables a decommissioned AWS account but leaves the row in place has that account's (now-revoked) credentials used for the probe. `buildConfig` or the probe calls then fail, `probeAndPersist` returns `ErrNoData`, and `Validate` falls permissive-true forever even though a healthy enabled account sits later in the list. Compare `Scheduler.enabledAccounts` (scheduler.go:917), which sets both filter fields. +- evidence: + ```go + provider := "aws" + accounts, err := s.accounts.ListCloudAccounts(ctx, config.CloudAccountFilter{Provider: &provider}) + if err != nil { + return nil, fmt.Errorf("list cloud accounts: %w", err) + } + for i := range accounts { + if accounts[i].Provider == "aws" { + return &accounts[i], nil + ``` +- suggested fix: pass `Enabled: &enabled` in the filter as `enabledAccounts` does; the redundant in-loop `Provider == "aws"` re-check can then go too. +- verdict: CONFIRMED — service.go:127 builds `config.CloudAccountFilter{Provider: &provider}` and leaves the `Enabled *bool` field (internal/config/types.go:1046) nil, so disabled rows are returned and the loop at :132-136 takes the first by index; the sibling Scheduler.enabledAccounts sets both fields (scheduler.go:916-920). The permissive-forever consequence follows from probeAndPersist returning ErrNoData on a buildConfig failure (service.go:86) and Validate's ErrNoData branch (service.go:148). +- issue: (pending cross-reference) + +### A06-005 Two plain-text email bodies are rendered with html/template and come out HTML-escaped +- category: correctness +- severity: medium +- location: internal/email/templates.go:1031 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `RenderPurchaseExecutedNotificationEmail` and `RenderPurchaseScheduledDelayEmail` (templates.go:895) call `renderTemplate`, whose own doc comment says "Never use it for plain-text bodies" (internal/email/template_renderers.go:26). Rendering `purchaseExecutedNotificationTemplate` with `RequestedByName="O'Brien & Co"`, `RequestedByEmail="a&b@example.com"` produces, verified by running the same template through html/template with the same FuncMap: `Requested by: O'Brien & Co <a&b@example.com>`. The recipient of a post-execution notice sees entity-encoded names and cannot copy-paste the address. The 07-H1 regression suite that fixed exactly this class (internal/email/template_renderers_test.go:504) covers PasswordReset, UserInvite, RegistrationDecision and PurchaseApprovalRequest — it omits the two renderers that were never converted, so the gap is invisible to CI. +- evidence: + ```go + func RenderPurchaseExecutedNotificationEmail(data NotificationData) (string, error) { + return renderTemplate("purchase-executed-notification", purchaseExecutedNotificationTemplate, data) + } + ``` +- suggested fix: switch both to `renderTextTemplate` and extend `TestPlainTextTemplates_NoHTMLEscaping` to cover them. +- verdict: CONFIRMED — both renderers call `renderTemplate` (internal/email/templates.go:895,1031) against that function's own "Never use it for plain-text bodies" contract (internal/email/template_renderers.go:24-28), while every sibling plain-text renderer uses `renderTextTemplate` (template_renderers.go:62-180); executing the exact line at templates.go:931-932 through html/template yields `Requested by: O'Brien & Co <a&b@example.com>`, and the regression suite covers only PasswordReset/UserInvite/RegistrationDecision/PurchaseApprovalRequest (internal/email/template_renderers_test.go:504-560). +- issue: (pending cross-reference) + +### A07-013 Compute SP layer mixes a region-scoped coverage figure with a global utilization figure +- category: correctness +- severity: medium +- location: providers/aws/ladder/layer_states.go:184 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `fetchSPUtilizationPct` deliberately passes `region = ""` for Compute SPs because they are global, but `fetchSPCoveragePct` (line 156) is called once with `a.cfg.Region` and its result is assigned to both SP layers. The Compute SP `LayerState` therefore reports coverage measured over one region and utilization measured over all regions. In a multi-region account where the ladder runs in a low-spend region, coverage reads near zero while utilization reads near 100, and the engine sizes a Compute SP purchase against a denominator that excludes most of the commitment it is already paying for. +- evidence: + ```go + region := a.cfg.Region + if planType == spPlanTypeCompute { + region = "" // "" = all regions in the CE GetSavingsPlansUtilization API + } + summary, err := a.spUtil.GetSPUtilization(ctx, cePlanType, region, a.cfg.lookbackDays()) + ``` +- suggested fix: fetch SP coverage per layer with the same region convention utilization uses (`""` for Compute, `cfg.Region` for EC2Instance) instead of sharing one region-scoped value. +- verdict: CONFIRMED — providers/aws/ladder/layer_states.go:65 fetches coverage once through `fetchSPCoveragePct`, which passes `a.cfg.Region` at :156, and :69-70 hand that same pointer to both SP layers, while `fetchSPUtilizationPct` :183-186 substitutes `""` for Compute. +- issue: (pending cross-reference) + +### A07-021 Region display-name map is stale for the renamed EU regions, and the fallback branch is dead +- category: correctness +- severity: medium +- location: providers/aws/recommendations/converters.go:137 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the map keys five regions under the old `"EU (...)"` display form while keying the newer ones under `"Europe (...)"`. AWS renamed those to `Europe (Ireland)`, `Europe (Frankfurt)`, `Europe (London)`, `Europe (Paris)` and `Europe (Stockholm)`. When CE returns the current name, `normalizeRegionName` falls through and returns the display string verbatim, so `rec.Region` becomes `"Europe (Ireland)"`. `filterByIncludedRegions` then drops that rec for an `--include-regions eu-west-1` filter, and any rec that survives carries a region string no service client can use. Separately, the `strings.HasPrefix` block below returns exactly the same value as the final statement, so it is dead code that reads as a guard. +- evidence: + ```go + if strings.HasPrefix(region, "us-") || strings.HasPrefix(region, "eu-") || + strings.HasPrefix(region, "ap-") || strings.HasPrefix(region, "sa-") || + strings.HasPrefix(region, "ca-") || strings.HasPrefix(region, "me-") || + strings.HasPrefix(region, "af-") || strings.HasPrefix(region, "il-") { + return region + } + return region + ``` +- suggested fix: add the `Europe (...)` aliases for the five renamed regions, delete the dead prefix branch, and log once when a display-form value falls through unmapped so future renames surface. +- verdict: CONFIRMED — the map at providers/aws/recommendations/converters.go:136-166 keys eu-west-1/2/3, eu-central-1 and eu-north-1 as `"EU (...)"` while eu-south-1, eu-south-2 and eu-central-2 use `"Europe (...)"` in the same literal, so the two spellings are inconsistent within one map and the `Europe (...)` form for the five older regions falls through; the prefix block at :173-179 returns `region`, identical to the `return region` immediately below it, so it is dead as claimed. +- issue: (pending cross-reference) + +### A07-023 FindConvertibleOffering returns the first result without paginating or checking the payment option +- category: correctness +- severity: medium +- location: providers/aws/services/ec2/client.go:794 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: this function uses the `Filters[]`-heavy request shape that the comment on `findOfferingID` (line 480) says AWS answers with empty pages plus a `NextToken` on sparse offering sets — and it does not paginate at all, so a first empty page yields "no convertible offering found" for an instance type that has one. It also caps at `MaxResults: 20`, never sets `OfferingType`, and returns `ReservedInstancesOfferings[0]`, so the offering it hands back for an exchange target may carry a different payment option than the source RI. +- evidence: + ```go + input := &ec2.DescribeReservedInstancesOfferingsInput{ + Filters: filters, + IncludeMarketplace: aws.Bool(false), + MaxResults: aws.Int32(20), + } + result, err := c.client.DescribeReservedInstancesOfferings(ctx, input) + ``` +- suggested fix: build the request from typed fields the way `describeInputFromQuery` does, paginate with the existing `isLastEC2Page` helper and page cap, and match the payment option before returning. +- verdict: CONFIRMED — providers/aws/services/ec2/client.go:794-818 builds a six-entry `Filters[]` request with `MaxResults: 20`, no `OfferingType`, a single `DescribeReservedInstancesOfferings` call and `ReservedInstancesOfferings[0]`; the comment at :481-487 documents that this exact request shape returns empty pages plus a `NextToken`, and the function has live callers at internal/server/ladder_write.go:64 and internal/server/handler_ri_exchange.go:63. +- issue: (pending cross-reference) + +### A08b-015 GCP billing SKU listing never follows the catalog page token +- category: correctness +- severity: high +- location: providers/gcp/services/cloudsql/client.go:106 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cloudstorage/client.go:148`, `memorystore/client.go:118`, `computeengine/client.go:331`) +- failure scenario: `Services.Skus.List(serviceID).Do()` returns one page (default 50 SKUs) and a `NextPageToken` that is discarded; the `BillingService` interface has no page-token parameter, so the omission is not fixable at the call site. The Cloud SQL catalog alone has several hundred SKUs, so the tier being priced is almost never in the first page and `getSQLPricing` returns "no pricing found" — which `fillSQLPricing` then swallows (A08b-016). `/usr/bin/grep -rn "PageToken" providers/gcp/` returns no matches at this commit. +- evidence: + ```go + func (r *realBillingService) ListSKUs(serviceID string) (*cloudbilling.ListSkusResponse, error) { + return r.service.Services.Skus.List(serviceID).Do() + } + ``` +- suggested fix: widen the interface to accept a page token (or return an iterator) and loop until `NextPageToken` is empty, with an explicit cap that errors rather than truncating. +- verdict: PLAUSIBLE — the dropped token is real (cloudsql:105-107, cloudstorage:147-149, memorystore:117-119, computeengine:330-332, and `/usr/bin/grep -rn PageToken providers/gcp/` returns nothing), but the "default 50 SKUs" premise is wrong: `services.skus.list` defaults to 5000 items per page (google.golang.org/api@v0.274.0 cloudbilling/v1, `ServicesSkusListCall.PageSize` doc), so whether any given service catalog overflows one page is a runtime fact I could not establish from source. +- severity-adjusted: medium — silent truncation is certain, a guaranteed miss for Cloud SQL is not. +- issue: (pending cross-reference) + +### A08b-018 Cosmos DB `ValidateOffering` compares reservation SKUs against account capability names, so it can never pass +- category: correctness +- severity: high +- location: providers/azure/services/cosmosdb/client.go:510 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GetValidResourceTypes` returns either capability names harvested from existing accounts (`collectCapabilitiesFromAccounts` returns `capability.Name` values such as `EnableMongo`) or the hardcoded `getCommonSKUs` list of the same shape. A Cosmos reservation SKU is a throughput tier — `detailsFromCosmosSKU` at line 780 documents the real form as `"100RU"`. `ValidateOffering` does `strings.EqualFold("100RU", "EnableMongo")` for every entry and always returns `invalid Azure Cosmos DB SKU`, so any pre-purchase validation gate blocks every legitimate Cosmos recommendation. +- evidence: + ```go + func (c *CosmosDBClient) getCommonSKUs() []string { + return []string{ + "EnableCassandra", + "EnableMongo", + "EnableGremlin", + "EnableTable", + "EnableServerless", + } + } + ``` +- suggested fix: validate against the throughput-SKU grammar the purchase body actually sends, or drop the check rather than shipping a validator that cannot succeed. +- verdict: CONFIRMED — `GetValidResourceTypes` returns either `capability.Name` values (cosmosdb:456-498) or the `EnableMongo`-shaped `getCommonSKUs` list (510-519), and `ValidateOffering` (359-372) `EqualFold`s those against `rec.ResourceType`, which the converter takes from the reservation SKU and `detailsFromCosmosSKU` (770-797) documents as `"100RU"`. The two vocabularies cannot intersect. +- severity-adjusted: medium — `ValidateOffering` and `GetValidResourceTypes` have no caller in this repo outside tests; only pkg/provider/interface.go:49 and :53 name them, so no live gate is blocked today. +- issue: (pending cross-reference) + +### A08b-019 `maxMachineTypeItems = 20` makes GCP `GetValidResourceTypes` error in every real region +- category: correctness +- severity: high +- location: providers/gcp/services/computeengine/client.go:37 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (consumer at `client.go:920`) +- failure scenario: the constant caps a per-item iteration (one machine type per `Next()`), and the consumer returns a hard error rather than stopping: `return nil, fmt.Errorf("computeengine: GetValidResourceTypes iteration cap (%d items) reached", maxMachineTypeItems)`. A GCP zone publishes well over a hundred machine types, so the 21st item always fires the cap and `GetValidResourceTypes` never returns a list — every Compute Engine CUD `ValidateOffering` fails with `failed to get valid tiers`. +- evidence: + ```go + // maxMachineTypeItems caps GCP machine types iteration (one item per Next() call). + const maxMachineTypeItems = 20 + ``` +- suggested fix: raise the cap to a value above the real machine-type count for a zone (or page properly with a filter), and keep the hard error only as a runaway guard. +- verdict: CONFIRMED — the constant and its "one item per Next() call" comment are at computeengine:36-37, and the consumer increments `itemIdx` once per `it.Next()` and returns a hard error at 920-922, so the 21st machine type aborts the call. One wording correction: `ValidateOffering` wraps it as "failed to get valid machine types" (839-842), not "failed to get valid tiers". +- severity-adjusted: medium — `GetValidResourceTypes`/`ValidateOffering` have no caller in this repo outside tests (pkg/provider/interface.go:49,53). +- issue: (pending cross-reference) + +### A08b-020 GCP iteration caps are per-item, named "Pages", and abort the whole call +- category: correctness +- severity: medium +- location: providers/gcp/services/cloudsql/client.go:181 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cloudstorage/client.go:201`, `memorystore/client.go:194`, `computeengine/client.go:406` and `client.go:478`) +- failure scenario: each loop increments `pageIdx` once per `it.Next()`, so `maxRecsPages = 20` means twenty recommendations, not twenty pages, and non-ACTIVE recommendations consume the budget before being skipped. A project with 21 Recommender rows returns zero recommendations and an error. `maxCommitmentsPages = 50` does the same for existing CUDs, so a project holding 51 commitments fails coverage collection entirely. +- evidence: + ```go + for pageIdx := 0; ; pageIdx++ { + if pageIdx >= maxRecsPages { + return nil, fmt.Errorf("cloudsql: GetRecommendations iteration cap (%d items) reached", maxRecsPages) + } + ``` +- suggested fix: rename the constants to reflect items and raise them to values a real account cannot legitimately exceed, so the guard fires only on a runaway iterator. +- verdict: CONFIRMED — `pageIdx` advances once per `it.Next()` at cloudsql:177-183, computeengine:402-408 and computeengine:474-480; the non-ACTIVE skip happens after the increment (cloudsql:197-199) so it consumes budget; the constants are `maxRecsPages = 20` (cloudsql:22, cloudstorage:23, memorystore:24, computeengine:31) and `maxCommitmentsPages = 50` (computeengine:34), and hitting either returns an error rather than stopping. +- issue: (pending cross-reference) + +### A08b-025 Reservation classification by SKU-name substring cross-assigns reservations between services +- category: correctness +- severity: medium +- location: providers/azure/services/database/client.go:278 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same shape at `cache/client.go:247`, `cosmosdb/client.go:248`, `search/client.go:193`, `synapse/client.go:216`) +- failure scenario: each client claims any reservation whose lowercased SKU name contains its own keyword. Azure SQL Data Warehouse / Synapse SKUs are `SqlDW`-family names that contain "sql", so the database client claims every Synapse reservation as `ServiceRelationalDB` while the synapse client also claims it via its `"dw"` prefix — the same reservation is reported under two services with two different `ServiceType` values. The synapse `"dw"` prefix test is likewise unanchored to a family. +- evidence: + ```go + props := detail.Properties + if props.SKUName == nil || !strings.Contains(strings.ToLower(*props.SKUName), "sql") { + return nil + } + ``` +- suggested fix: classify on the reservation's declared reserved-resource type from an inventory API rather than substring-matching the SKU string. +- verdict: PLAUSIBLE — the unanchored substring tests are exactly as described (database:278 on "sql", cache:247 on "redis", cosmosdb:248 on "cosmos", search:193 on "search", synapse:217-219 on a `dw`/`scu` prefix or a "synapse" substring), but the specific database/synapse overlap is not demonstrable from source: a Synapse `SKUName` of the `DW1000c` form matches synapse's `dw` prefix and does not contain "sql", so which client double-claims depends on the live SKU string Azure returns. +- issue: (pending cross-reference) + +### A08b-027 Six pager loops have neither a page cap nor a context check, unlike every sibling loop in the same files +- category: correctness +- severity: medium +- location: providers/azure/services/cache/client.go:697 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cosmosdb/client.go:691`, `database/client.go:826` and `client.go:870`, `managedredis/client.go:182`, `synapse/client.go:142` and `client.go:189`, `compute/exchange.go:140`) +- failure scenario: `collectRedisReservations` in the same file checks `ctx.Err()` and enforces `maxReservationsPages`; `fetchSKUCatalogue` does neither. A cancelled request keeps issuing `NextPage` calls until the SDK's own transport notices, and a server returning a non-terminating page chain spins forever. `compute/exchange.go:140` is the tenant-wide reservation listing that feeds the exchange ownership gate, so an unbounded walk there stalls an authorization check. +- evidence: + ```go + out := make(map[string]redisSKUEntry) + for pager.More() { + page, err := pager.NextPage(ctx) + if err != nil { + logging.Warnf("azure cache: SKU catalogue page fetch failed for region %s: %v ...", c.region, err, len(out)) + return nil + } + ``` +- suggested fix: apply the same `pageIdx >= maxPages` cap and `ctx.Err()` precheck the sibling loops in each file already use. +- verdict: CONFIRMED — all eight cited loops are bare `for pager.More()` with no index and no precheck (cache:697, cosmosdb:691, database:826 and 870, managedredis:182, synapse:142 and 189, compute/exchange.go:140), while `collectRedisReservations` in the same cache file enforces both at 217-224. Narrowing on the cancellation half: database:829 and 873 do test `ctx.Err()` after a page error, and every one of these loops returns on any page error, so the non-terminating scenario needs a server emitting an endless chain of successful pages. +- issue: (pending cross-reference) + +### A09-010 findBestFit can emit a target instance type that does not exist for the family +- category: correctness +- severity: medium +- location: pkg/exchange/reshape.go:694 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the loop keeps the LAST index in `sizeOrder` whose factor fits, and `"metal"` is the last entry with the same factor (192) as `"24xlarge"`. An m5 convertible RI with `normalizedUsed = 200` picks `bestIdx` at `"metal"` and `analyzeRI` composes `family + "." + targetSize` = `"m5.metal"`, count 2. The same path produces `"t3.9xlarge"`, `"r5.3xlarge"` and other size/family pairs AWS does not offer. `resolveOffering` in auto.go then fails the offering lookup and the recommendation is silently skipped, while the reshape dashboard shows a target no operator can act on. +- evidence: + ```go + bestIdx := -1 + for i, s := range sizeOrder { + nf := normalizationFactors[s] + if nf > 0 && nf <= normalizedUsed { + bestIdx = i + } + } + ``` +- suggested fix: drop `"metal"` from `sizeOrder` (keep it in `normalizationFactors` for parsing existing RIs) and prefer the smaller-indexed size on a factor tie, so the suggested target is always a standard size. +- verdict: CONFIRMED — `sizeOrder` ends with `"metal"` (pkg/exchange/reshape.go:382-387) and `normalizationFactors["metal"] == normalizationFactors["24xlarge"] == 192` (reshape.go:378-379), so the loop's last-wins assignment at reshape.go:694-699 selects index 17 ("metal") for any `normalizedUsed >= 192`; `analyzeRI` then composes `family + "." + targetSize` at reshape.go:486 with no per-family existence check. +- issue: (pending cross-reference) + +### A09-011 The "metal" normalization factor is a single hardcoded value for a family-dependent quantity +- category: correctness +- severity: medium +- location: pkg/exchange/reshape.go:379 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: AWS metal sizes map to different normalization factors per family: `m5.metal` is 24xlarge-equivalent (192) but `i3.metal` is 16xlarge-equivalent (128) and `m5zn.metal` is 12xlarge-equivalent (96). Any `.metal` source RI with an unpopulated `RIInfo.NormalizationFactor` falls through `resolveNormFactor` to this table and is credited with 192 units. For an `i3.metal` at 50% utilization the computed `normalizedUsed` is 96 instead of 64, and `findBestFit` sizes the exchange target 50% too large. +- evidence: + ```go + "24xlarge": 192, + "metal": 192, + } + ``` +- suggested fix: remove the `"metal"` entry so `resolveNormFactor` returns 0 and `analyzeRI` skips the RI (fail loud) unless the caller supplied the real per-family factor from `ec2:DescribeInstanceTypes`. +- verdict: CONFIRMED — `normalizationFactors["metal"] = 192` is a single family-independent entry (pkg/exchange/reshape.go:379), and `resolveNormFactor` falls back to that table whenever `ri.NormalizationFactor == 0` (reshape.go:432-437), feeding `normalizedPurchased` at reshape.go:461 and `findBestFit` at :481 with no family discrimination. +- issue: (pending cross-reference) + +### A10-006 CSV numeric fields are parsed with `fmt.Sscanf`, which truncates and accepts trailing garbage +- category: correctness +- severity: medium +- location: cmd/multi_service_csv.go:173 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A row whose `Count` cell reads `3.7` parses as `3` with `err == nil` (`Sscanf` stops at the `.` and reports one successful conversion). `12 units` parses as `12`. `-5` parses as `-5` and flows into `savingsPerInstance`, `ApplyInstanceLimit` and the purchase loop as a negative count. `EstimatedSavings` of `1000 USD` silently becomes `1000`. Every one of these is a money quantity read from an operator-editable file with no boundary rejection, contrary to the project's strict-integer-parsing rule. +- evidence: + ```go + func parseCSVInt(record []string, colIdx map[string]int, fieldName string, target *int) error { + value := getCSVField(record, colIdx, fieldName) + if value == "" { return nil } + if _, err := fmt.Sscanf(value, "%d", target); err != nil { + return fmt.Errorf("invalid %s value '%s': %w", fieldName, value, err) + } + return nil + } + ``` +- suggested fix: Use `strconv.Atoi` / `strconv.ParseFloat` on the whole trimmed cell, which reject trailing characters, and add an explicit `< 0` rejection for `Count`. +- verdict: CONFIRMED — ran the parses against the pinned toolchain: `Sscanf("%d")` returns err=nil with 3 for "3.7", 12 for "12 units" and -5 for "-5", and `Sscanf("%f")` returns 1000 for "1000 USD", so cmd/multi_service_csv.go:167-190 admits every value the finding names. +- issue: (pending cross-reference) + +### A10-011 The CSV TOTAL row sums normalized units for rows whose own NU cell is blank +- category: correctness +- severity: medium +- location: cmd/multi_service_csv.go:314 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A mixed `--all-services` run includes ElastiCache recs with `ResourceType=cache.r7g.large`. `formatNormalizedUnitsOrBlank` blanks the per-row NormalizedUnits cell because the service is not RDS, but `buildTotalRow` calls `RDSInstanceNUFromType` unconditionally, and that function only requires three dot-separated parts before looking up the size suffix (`providers/aws/recommendations/family_nu.go:62-67`), so `cache.r7g.large` resolves to the `large` NU value. The TOTAL cell therefore exceeds the sum of the visible cells above it, and an operator reconciling family-NU bundling by hand cannot make the column add up. +- evidence: + ```go + for i := range results { + r := results[i] + totalCount += r.Recommendation.Count + totalNU += float64(r.Recommendation.Count) * recommendations.RDSInstanceNUFromType(r.Recommendation.ResourceType) + ``` +- suggested fix: Gate the `totalNU` accumulation on the same `rec.Service != ServiceRDS && != ServiceRelationalDB` test `formatNormalizedUnitsOrBlank` uses, ideally by summing the per-row helper's own value. +- verdict: CONFIRMED — buildTotalRow calls RDSInstanceNUFromType for every row (cmd/multi_service_csv.go:314) while formatNormalizedUnitsOrBlank blanks non-RDS rows (cmd/multi_service_csv.go:370-373), and "cache.r7g.large" splits into three parts so the lookup returns the "large" entry, 4 (providers/aws/recommendations/family_nu.go:61-68, map at :23-38). +- issue: (pending cross-reference) + +### A10-014 The dry-run footer tells operators Savings Plans purchasing is not implemented, but it is +- category: correctness +- severity: medium +- location: cmd/multi_service_stats.go:184 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A dry run over `--all-services` prints "To actually purchase these RIs, run with --purchase flag / Note: Savings Plans purchasing not yet implemented". An operator who read that line reruns with `--purchase` expecting SP rows to be skipped. They are not: `createServiceClient` returns a real `savingsplans.NewClient` for all four SP slugs (cmd/main.go:257-265) and `savingsplans.Client.PurchaseCommitment` (providers/aws/services/savingsplans/client.go:210) issues a real `CreateSavingsPlan`. The message actively misleads on a money path. +- evidence: + ```go + if isDryRun { + AppLogger.Println("\n💡 To actually purchase these RIs, run with --purchase flag") + AppLogger.Println(" Note: Savings Plans purchasing not yet implemented") + } + ``` +- suggested fix: Delete the stale second line, or replace it with the count of SP recommendations that `--purchase` will buy. +- verdict: CONFIRMED — printFinalMessage emits the line on every dry run (cmd/multi_service_stats.go:181-185), while createServiceClient returns a real savingsplans client for all four SP slugs (cmd/main.go:258-266) and PurchaseCommitment issues CreateSavingsPlan (providers/aws/services/savingsplans/client.go:210-235). +- issue: (pending cross-reference) + +### A10-017 Recommendations with no resolvable account name are dropped whenever any account filter is set +- category: correctness +- severity: medium +- location: cmd/multi_service_filters.go:158 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A rec whose `Account` field is empty (Savings Plans rows and any provider that does not populate it — `populateAccountNames` only sets `AccountName` when `Account != ""`, cmd/multi_service_helpers.go:190) gets `AccountName == ""`. With `--exclude-accounts sandbox` set, `shouldIncludeAccount` returns false for that rec because the exclude list is non-empty, so it is dropped even though it does not match `sandbox` at all. `processRecommendation` classifies this as an expected dimension mismatch and records no drop reason (cmd/multi_service_filters.go:63-69), so the end-of-run drop summary never mentions it and the operator sees an unexplained shortfall. +- evidence: + ```go + if accountName == "" { + return len(cfg.IncludeAccounts) == 0 && len(cfg.ExcludeAccounts) == 0 + } + ``` +- suggested fix: With only an exclude list in force, an unattributed rec cannot match it and should be kept. Keep the drop for an include list, but surface it with a dedicated drop reason so the count appears in the summary. +- verdict: CONFIRMED — cmd/multi_service_filters.go:158-160 returns false as soon as either list is non-empty, populateAccountNames leaves AccountName empty when Account is empty (cmd/multi_service_helpers.go:188-193), and processRecommendation returns an empty dropReason on the dimension-filter branch (cmd/multi_service_filters.go:63-69) so the drop never reaches the summary. +- issue: (pending cross-reference) + +### A10-024 The CSV purchase path confirms per region rather than once for the run +- category: correctness +- severity: medium +- location: cmd/multi_service.go:693 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `processPurchaseLoop` prompts only when `j == 0`, and it is called once per (service, region) from `runToolFromCSV`. A CSV spanning RDS in three regions and EC2 in two therefore prompts five separate times, and each prompt reports only that region's totals ("About to purchase 4 instances with estimated monthly savings: $210"). The operator never sees the run-wide instance count or dollar figure before the first purchase fires, and answering "no" at prompt four does not undo the twelve instances already bought under prompts one through three. The non-CSV path gets this right, confirming once against `sumPassedRecs` over the whole run (cmd/multi_service.go:157). +- evidence: + ```go + if j == 0 { + totalInstances := CalculateTotalInstances(recs) // this region only + ... + if !ConfirmPurchase(totalInstances, totalSavings, cfg.SkipConfirmation) { + return createCancelledResults(recs, region, cfg) + } + } + ``` +- suggested fix: Hoist the confirmation into `runToolFromCSV` ahead of the service/region loop, summing over the full post-dedup set, and drop the per-region prompt. +- verdict: CONFIRMED — processPurchaseLoop prompts under `if j == 0` over this region's recs only (cmd/multi_service.go:693-705) and runToolFromCSV calls it once per (service, region) inside the nested loops at :521-564, while the non-CSV path confirms once against sumPassedRecs for the whole run (:156-163). +- issue: (pending cross-reference) + +### A11-012 The synthetic "current user" row exposes Edit and Delete bound to a fake id +- category: correctness +- severity: medium +- location: frontend/src/users/userActions.ts:55 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when the users list omits the signed-in user, `withCurrentUser` prepends a row with `id: 'current'` and a hardcoded `mfa_enabled: false`. `renderUsers` gives every row an Edit and a Delete button keyed on that id (userList.ts:122-123), so clicking Edit issues `GET /api/users/current` and Delete issues `DELETE /api/users/current`, neither of which addresses a real user. The row is also selectable, so a bulk delete includes `current` and fails the whole `Promise.all`. Separately the hardcoded `mfa_enabled: false` is counted by the "MFA Enabled" stat card (userList.ts:40), understating the count for an admin who has MFA on. +- evidence: + ```typescript + // users/userActions.ts:59-65 + const synthetic: api.APIUser = { + id: 'current', + email: current.email, + groups: Array.isArray(current.groups) ? current.groups : [], + mfa_enabled: false, + }; + return [synthetic, ...users]; + ``` +- suggested fix: mark the synthetic row read-only (skip the Edit/Delete/select controls when `user.id === 'current'`) and carry the real `mfa_enabled` from the session user, or drop the synthetic row entirely now that the list endpoint returns the caller. +- verdict: PLAUSIBLE — the row really does get the full control set, since `renderUsers` keys Edit and Delete on `user.id` with no `current` exemption (frontend/src/users/userList.ts:121-122, whose only special case is the "You" badge at userList.ts:105) and the hardcoded `mfa_enabled: false` feeds the stat card at userList.ts:41. Reaching it needs `listUsers` to omit the caller, which I could not establish: `GET /api/users` returns the unscoped list (internal/api/handler_users.go:17-27), so the `u.email === current.email` test at userActions.ts:57 normally matches and the synthetic row is never prepended; an email-case mismatch or an API-key session would be the trigger. +- issue: (pending cross-reference) + +### A11-013 A partly-failed bulk user delete leaves deleted users on screen and names none of the failures +- category: correctness +- severity: medium +- location: frontend/src/users/userActions.ts:173 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `Promise.all` rejects on the first failing delete. The `catch` shows "Failed to delete some users" and neither clears the selection nor reloads the list, so users that were successfully deleted remain rendered and selected until a manual refresh; a second click on Bulk Delete then re-issues deletes for rows that no longer exist. The operator is never told which users survived. `handleFanOutExecute` in app.ts:545 already demonstrates the `allSettled` + per-item reporting pattern this path needs. +- evidence: + ```typescript + // users/userActions.ts:172-183 + await Promise.all( + Array.from(selectedUserIds).map(userId => api.deleteUser(userId)) + ); + clearSelectedUserIds(); + await loadUsers(); + ``` +- suggested fix: use `Promise.allSettled`, deselect only the ids that resolved, reload the list in both branches, and name the failures in the error toast. +- verdict: CONFIRMED — `Promise.all` short-circuits on the first rejection and both `clearSelectedUserIds()` and `loadUsers()` sit inside the `try` after it (frontend/src/users/userActions.ts:172-178), so the catch at userActions.ts:179-182 shows one generic string and leaves the stale, still-selected rows on screen; deletes already issued are not undone, so a second Bulk Delete re-issues them. +- issue: (pending cross-reference) + +### A11-014 A multi-account savings filter is silently narrowed to the first account +- category: correctness +- severity: medium +- location: frontend/src/api/history.ts:45 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when the topbar filter selects three accounts, `getSavingsAnalytics` sends both `account_ids=a,b,c` and `account_id=a`. The comment states the handler honours only the singular form, so the chart renders account A's savings while the UI chip row says three accounts are selected. The user reads an under-reported savings figure with no indication that two accounts were dropped. The sibling `getHistory` (history.ts:23) sends only the plural form, so the two panels on the same page disagree for the same filter. +- evidence: + ```typescript + // api/history.ts:39-46 + if (filters.account_ids && filters.account_ids.length > 0) { + params.set('account_ids', filters.account_ids.join(',')); + if (filters.account_ids[0]) params.set('account_id', filters.account_ids[0]); + } + ``` +- suggested fix: have the caller surface the limitation (disable or annotate the chart when more than one account is selected) rather than quietly charting a subset. +- verdict: CONFIRMED — the analytics handlers read only the singular param (`accountID := params["account_id"]` at internal/api/handler_analytics.go:40, 131, 197, then `resolveSingleAccountFilterIDs`) and `account_ids` is parsed nowhere under internal/api for these routes, so the extra plural param at frontend/src/api/history.ts:43 is inert and the chart is scoped to the first account; `getHistory` (history.ts:22) sends only the plural form, which internal/api/handler_history.go:678 does honour, so the two panels genuinely disagree. +- issue: (pending cross-reference) + +### A12-013 The savings-analytics summary keys the frontend reads do not exist on the wire +- category: correctness +- severity: medium +- location: frontend/src/modules/savings-history.ts:293 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GET /history/analytics` returns `Summary *HistorySummary` (`internal/api/handler_analytics.go:99`), whose JSON keys are `total_purchases`, `total_upfront`, `total_monthly_savings`, `total_annual_savings`. The frontend's `SavingsAnalyticsSummary` (`frontend/src/api/types.ts:650`) declares `total_period_savings`, `average_savings_per_period` and `peak_savings`, none of which are ever sent, so all three `??` fallbacks are taken on every response and the declared contract is a fiction. The Go struct that does carry those keys, `HistorySummaryAnalytics` (`internal/api/types.go:112`), is never populated anywhere — it is dead. The numbers happen to agree today only because the client-side sum reproduces `TotalMonthlySavings`; a change on either side breaks silently rather than loudly. +- evidence: + ```ts + const monthlyTotal = summary?.total_period_savings ?? totalSavings; + const monthlyAvg = summary?.average_savings_per_period ?? avgPerPeriod; + const monthlyPeak = summary?.peak_savings ?? peakSavings; + ``` +- suggested fix: Align the TypeScript interface with `HistorySummary`'s actual keys and delete the unused `HistorySummaryAnalytics` struct, or populate it and have the handler return it. +- verdict: CONFIRMED — The handler returns Summary *HistorySummary (internal/api/handler_analytics.go:99) whose JSON keys are total_purchases/total_upfront/total_monthly_savings/total_annual_savings (internal/api/types.go:910-929), and HistorySummaryAnalytics (internal/api/types.go:112) is never constructed anywhere, so all three keys read at frontend/src/modules/savings-history.ts:293-295 are absent on every response. +- issue: (pending cross-reference) + +### A12-017 Cloning the ladder form after populating it resets its three `` (line 1561); the free-text fallback the copy describes no longer exists. Opened from the Convertible RIs table there are also no Cost Explorer alternatives, so the select contains nothing but the placeholder and the user is at a dead end following misleading instructions. +- evidence: + ```ts + if (offeringsError) { + placeholder.textContent = 'Could not load offerings -- type a UUID'; + ``` +- suggested fix: Change the copy to say the offerings failed to load and offer a retry, or restore a visible text field for the manual path. +- verdict: CONFIRMED — The placeholder tells the user to type a UUID (frontend/src/riexchange.ts:1490-1491) but the only UUID-accepting field is `offeringInput`, created as `type = 'hidden'` (frontend/src/riexchange.ts:1559-1561), and the visible control is a select whose only other option group is the CE alternatives. +- issue: (pending cross-reference) + +### A12-067 `escapeHtml` output is assigned to `textContent`, double-escaping option labels +- category: correctness +- severity: low +- location: frontend/src/riexchange.ts:1507 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `option.textContent` escapes on its own, so pre-escaping means any `&`, `<` or quote in an AWS `instance_type` or `offering_type` renders literally as `&` in the dropdown. Line 1525 does the same for the Cost Explorer alternatives. It also misleads a reader into treating the option list as an HTML sink. +- evidence: + ```ts + const label = escapeHtml(o.instance_type) + (o.offering_type ? ' -- ' + escapeHtml(o.offering_type) : ''); + opt.textContent = label; + ``` +- suggested fix: Assign the raw strings to `textContent` and drop the `escapeHtml` calls at 1507 and 1525. +- verdict: CONFIRMED — escapeHtml output is assigned to option.textContent, which escapes again (frontend/src/riexchange.ts:1507-1508 and :1525), so an ampersand or angle bracket in an instance_type or offering_type renders as its entity in the dropdown. +- issue: (pending cross-reference) + +### A12-069 A failed exchange's `error` field is never surfaced in the history table +- category: correctness +- severity: low +- location: frontend/src/riexchange.ts:2211 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: When an exchange fails at the AWS accept step, `RIExchangeHistoryRecord.error` carries the reason but the row renders only a "failed" status badge. The operator has no in-product way to see why money was not spent, or whether it partially was, and must read backend logs. +- evidence: + ```ts + + '' + escapeHtml(rec.status) + '' + ``` +- suggested fix: Put the escaped `rec.error` in the status cell's `title`, or render it in a detail row for failed statuses. +- verdict: CONFIRMED — RIExchangeHistoryRecord declares `error?: string` (frontend/src/api/types.ts:844) but the status cell renders only the badge and no cell or title carries it (frontend/src/riexchange.ts:2216). +- issue: (pending cross-reference) + +### A12-073 The "YTD Savings" tile's sparkline plots the trend-range window, not year-to-date +- category: correctness +- severity: low +- location: frontend/src/dashboard.ts:1057 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The tile's value comes from `data.ytd_savings`, but the sparkline drawn beneath it comes from `loadSavingsTrendChart`'s data points, whose window is whatever the trend range toggle is set to (7, 30 or 90 days, or a rolling 365). Clicking the "7" range button changes the shape of the line under a label that says YTD, so the number and the line describe different periods. +- evidence: + ```ts + attachSparkline('ytd', points.map((p: SavingsDataPoint) => p.cumulative_savings || 0)); + ``` +- suggested fix: Fetch a separate year-to-date series for the tile, or relabel the sparkline so it is not read as YTD. +- verdict: CONFIRMED — The tile value is data.ytd_savings (frontend/src/dashboard.ts:308) while the sparkline plots the trend-range data points whose window comes from savingsTrendRange (frontend/src/dashboard.ts:1022-1031, :1057), which the range buttons mutate at :1174-1177. +- issue: (pending cross-reference) + +### A13-019 The `policy_arns` output claims to list all attached policies and lists three of five +- category: correctness +- severity: low +- location: terraform/environments/aws/ci-cd-permissions/outputs.tf:11 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `role.tf:158-181` attaches five managed policies (`networking`, `compute`, `compute_b`, `data`, `iam`). The output describes itself as "ARNs of all attached managed policies" but omits `compute_b` and `iam`. An operator scripting a teardown or an audit from `terraform output -json policy_arns` silently misses the two policies that carry the whole #1705 boundary-delegation half, leaving them orphaned or unreviewed. +- evidence: + ```hcl + description = "ARNs of all attached managed policies" + value = { + networking = aws_iam_policy.networking.arn + compute = aws_iam_policy.compute.arn + data = aws_iam_policy.data.arn + } + ``` +- suggested fix: Add `compute_b` and `iam` (and `workload_boundary`, if the boundary belongs in the same inventory). +- verdict: CONFIRMED — role.tf:158-181 declares five `aws_iam_role_policy_attachment` resources (networking, compute, compute_b, data, iam) against `aws_iam_role.cudly_deploy`, while outputs.tf:11-18 describes itself as "ARNs of all attached managed policies" and emits only networking, compute and data. +- issue: (pending cross-reference) + +### A13c-016 AWS profile examples pin a full minor engine version the module documents as an incident cause +- category: correctness +- severity: low +- location: terraform/profiles/aws/prod.tfvars.example:33 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `modules/database/aws/variables.tf:58` documents that pinning a full minor + (`"16.6"`) "causes downgrade failures if RDS auto-upgrades beyond it (incident: 2026-07-16, + #1372)" and defaults to the major-only `"16"`. All three `profiles/aws/*.tfvars.example` files + (dev:33, fargate-dev:32, prod:33) publish `"16.6"`, so anyone starting from a profile inherits + the exact configuration the incident was about. `environments/aws/dev.tfvars.example:50` + correctly uses `"16"`, which shows the fix landed in one place and not the other. +- evidence: + ```hcl + database_engine_version = "16.6" + ``` +- suggested fix: change all three profile examples to `"16"`. +- verdict: CONFIRMED — the incident text and the major-only default are at + database/aws/variables.tf:57-61, and all three profile examples publish `"16.6"` + (dev:33, fargate-dev:32, prod:33) while every file under `environments/aws/` uses `"16"` + (dev.tfvars.example:50, github-dev:49, github-prod:47, github-staging:46). +- issue: (pending cross-reference) + +### A13c-018 `database/aws` output guards on a different condition than the data source it indexes +- category: correctness +- severity: low +- location: terraform/modules/database/aws/outputs.tf:39 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `data.aws_secretsmanager_secret.existing_password` is created when + `var.create_password` is false (main.tf:53), but this output selects between the two branches on + `var.master_password_secret_arn != null`. With `create_password = true` and an ARN also + supplied, the output indexes `[0]` of a zero-count data source and the plan fails with an + invalid-index error. The mirrored case — `create_password = false` with a null ARN — passes + `arn = null` to the data source and fails there instead. The in-tree caller + (`environments/aws/database.tf:15-16`) happens to set the one combination that works, so neither + path is exercised. +- evidence: + ```hcl + output "password_secret_name" { + value = var.master_password_secret_arn != null ? data.aws_secretsmanager_secret.existing_password[0].name : aws_secretsmanager_secret.db_password[0].name + } + ``` +- suggested fix: switch the ternary to `var.create_password` so both sides use the same predicate + as the resources, and add a variable precondition that the two inputs are mutually consistent. +- verdict: CONFIRMED — the data source's count is `var.create_password ? 0 : 1` + (database/aws/main.tf:53) while the output ternary keys off `var.master_password_secret_arn != + null` (outputs.tf:39); note `local.db_password_secret_arn` at main.tf:64 uses the correct + predicate, so the two outputs disagree with each other. The single in-tree caller sets + `create_password = false` with a non-null ARN (environments/aws/database.tf:15-16), the one + combination both branches survive. +- issue: (pending cross-reference) + +### A14-030 A failed health request is reported as a slow response +- category: correctness +- severity: low +- location: scripts/test-deployment.sh:266 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `curl … || true` leaves `time_total` empty when the request fails outright. `echo " < 1.0" | bc -l` is then a syntax error, `|| echo 0` makes both comparisons false, and control falls to the `else`, printing `response-time: /health in s (very slow, >3s)` and counting a FAIL. An operator reading the report sees a latency problem where the endpoint was actually unreachable. +- evidence: + ```bash + time_total=$($CURL_BIN -s $CURL_K -o /dev/null -w "%{time_total}" --connect-timeout 10 "${URL}${health_path}" 2>/dev/null || true) + if (( $(echo "$time_total < 1.0" | bc -l 2>/dev/null || echo 0) )); then + ``` +- suggested fix: Fail with a distinct message when `time_total` is empty or non-numeric, before comparing it. +- verdict: PLAUSIBLE — scripts/test-deployment.sh:266-272 never checks curl's exit status, but the stated mechanism needs `time_total` to come back empty, which requires a missing curl binary or a malformed URL: measured, a refused connection still emits a number (0.002, logging PASS) and a 10s connect timeout emits ~10, logging the "very slow" FAIL the finding describes. +- issue: (pending cross-reference) + +### Category: concurrency + +12 findings: 2 high, 7 medium, 3 low. + +### A06-001 Advisory lock is released with the request context, so a canceled/expired scheduled run strands the lock on a pooled connection +- category: concurrency +- severity: high +- location: internal/server/handler.go:112 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Cloud Scheduler POSTs `/api/scheduled/process_scheduled_purchases`; `handleScheduledHTTP` passes `r.Context()` straight through (internal/server/http.go:225). The scheduler's attempt deadline (180s by default) expires, the client disconnects, `r.Context()` is canceled, and the deferred `ReleaseAdvisoryLock(ctx, ...)` runs `SELECT pg_advisory_unlock($1)` on an already-canceled context. `internal/database/connection.go:357` logs a warning and `defer conn.Release()` returns the connection to the pool while the *session-level* advisory lock is still held. Every later invocation of that task type gets `pg_try_advisory_lock` = false from a different pooled connection and returns `{"status":"skipped","reason":"already_running"}` until that connection is recycled by the pool. The same happens on Lambda when a task runs past the function deadline. Note the file already knows this hazard: `releaseSkippedCollectionMarker` (handler.go:170) deliberately builds a detached `context.Background()` with its own timeout for exactly this reason. +- evidence: + ```go + acquired, err := locker.TryAdvisoryLock(ctx, lockID) + ... + defer locker.ReleaseAdvisoryLock(ctx, lockID) + } + return app.dispatchTask(ctx, taskType, params) + ``` +- suggested fix: release under a short detached context (`context.WithTimeout(context.Background(), 5*time.Second)`), mirroring `releaseSkippedCollectionMarker`. +- verdict: CONFIRMED — `handleScheduledHTTP` takes `ctx := r.Context()` and passes it straight to `HandleScheduledTask` (internal/server/http.go:225,249), whose deferred release runs on that same ctx (internal/server/handler.go:112); `ReleaseAdvisoryLock` only logs a warning when the unlock query fails and still runs `defer conn.Release()`, returning the still-locked session to the pool (internal/database/connection.go:354-361), while the sibling `releaseSkippedCollectionMarker` detaches deliberately (internal/server/handler.go:170). +- issue: (pending cross-reference) + +### A12-011 A pending AWS utilization response re-renders the shared container over the Azure/GCP table +- category: concurrency +- severity: high +- location: frontend/src/riexchange.ts:318 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: With the provider chip on AWS, `loadConvertibleRIs` starts `loadUtilization(gen)` against Cost Explorer. The user switches to Azure before it resolves. `loadExchangeableAzureRIs` resets `currentRIs = []` but never bumps `utilizationGeneration`, so the stale response passes the guard and `renderRIsTable` replaces the freshly rendered Azure reservations with the AWS empty state. `renderGCPEmptyStates` has the same hole. The convertible-RI list load itself (line 270) has no generation guard at all, so a slow AWS list can paint AWS rows, complete with Exchange buttons, into a panel the user believes is scoped to Azure. +- evidence: + ```ts + const utilization = await api.getRIUtilization(); + if (generation !== utilizationGeneration) return; + currentUtilization = new Map(utilization.map(u => [u.reserved_instance_id, u])); + const container = document.getElementById('ri-exchange-instances-list'); + if (container) renderRIsTable(container, currentRIAccountID); + ``` +- suggested fix: Bump `utilizationGeneration` in `loadExchangeableAzureRIs` and `renderGCPEmptyStates`, and apply the same generation check to the list load. +- verdict: CONFIRMED — utilizationGeneration is incremented only inside loadConvertibleRIs (frontend/src/riexchange.ts:273), so neither loadExchangeableAzureRIs (frontend/src/riexchange.ts:289-311) nor renderGCPEmptyStates (:120-136) invalidates an in-flight loadUtilization, whose guard therefore passes and calls renderRIsTable over the other provider's table (frontend/src/riexchange.ts:316-323); the list fetch at :270 has no guard at all. +- issue: (pending cross-reference) + +### A01-005 retryPurchase has no atomic claim on the failed row, so concurrent retries mint two successors and the full-row upsert clobbers the predecessor +- category: concurrency +- severity: medium +- location: internal/api/handler_purchases.go:1924 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Two operators (or a double click) POST `/api/purchases/retry/{id}` for the same failed row. Both `loadAndValidateRetryRequest` reads see `RetryExecutionID == nil` (line 1708) and pass; both `persistRetryExecution` calls insert their own pending successor and then upsert `originalUpdated` (a stale in-memory copy of the failed row) with their own `retry_execution_id`. Result: two approvable successors with two approval emails, the already-retried guard is defeated, the first linkage is overwritten, and any column another writer changed on the failed row in between is reverted by the second upsert (`ON CONFLICT ... DO UPDATE SET` writes every column, store_postgres.go:991-1012). Only provider-side token dedupe stands between the two approvals and a double purchase. +- evidence: + ```go + originalUpdated := *failedExec + originalUpdated.RetryExecutionID = &newExecutionID + // ... + if err := h.config.WithTx(ctx, func(tx pgx.Tx) error { + if err := h.config.SavePurchaseExecutionTx(ctx, tx, newExecution); err != nil { + return err + } + if err := h.config.SavePurchaseExecutionTx(ctx, tx, &originalUpdated); err != nil { + return err + } + ``` +- suggested fix: Inside the tx, replace the full-row upsert of the predecessor with a conditional `UPDATE purchase_executions SET retry_execution_id=$2 WHERE execution_id=$1 AND status='failed' AND retry_execution_id IS NULL` and return 409 when zero rows are affected, before inserting the successor. +- verdict: CONFIRMED — the already-retried guard at internal/api/handler_purchases.go:1708 reads RetryExecutionID outside any tx, persistRetryExecution:1909-1937 upserts a stale value copy of the failed row, and the ON CONFLICT clause at internal/config/store_postgres.go:991-1012 rewrites status, retry_execution_id and every other mutable column unconditionally, so nothing serializes two concurrent retries. +- issue: (pending cross-reference) + +### A01-013 Pausing a `running` execution lets a claimed in-flight purchase be resumed and re-claimed +- category: concurrency +- severity: medium +- location: internal/api/handler_purchases.go:271 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The scheduler's `claimAndExecute` has CAS-claimed a row to `running` and is mid-purchase. An operator clicks Pause (allowed from `running`), then Resume (`paused` -> `pending`). The next scheduler tick sees a pending row and `claimAndExecute` claims it again (pending is in its from-set), running the purchase a second time in parallel; the first run's terminal `SavePurchaseExecution` then overwrites whatever the second wrote. Only provider-side token dedupe prevents a second commitment (none for Azure savings plans). `running` appears in the pause from-set only to support the Run-now flow of A01-003, which itself never executes anything. +- evidence: + ```go + // Atomically transition to paused + if _, err := h.config.TransitionExecutionStatus(ctx, executionID, []string{"pending", "running"}, "paused", resolveCreatorUserID(session)); err != nil { + ``` +- suggested fix: Restrict pause to `{"pending"}` (and `notified` if desired); an execution that is actually running cannot be paused. +- verdict: CONFIRMED — pausePlannedPurchase (internal/api/handler_purchases.go:271) accepts from {pending,running} and resume:295 flips paused→pending; ProcessScheduledPurchases (internal/purchase/manager.go:715-742) re-selects due pending rows and claimAndExecute:179 claims from approved/pending/notified, while the first run's terminal SavePurchaseExecution (manager.go:234) is a full-row upsert; the double run additionally requires pause, resume and a scheduler tick all inside the first run's in-flight window, but every step on that path is unguarded. +- issue: (pending cross-reference) + +### A03-010 Failed-login bookkeeping does a whole-row read-modify-write from an unauthenticated path +- category: concurrency +- severity: medium +- location: internal/auth/service_user.go:767 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `recordFailedLogin` writes the entire user row it loaded at the start of `Login` (`UpdateUser` sets all 18 columns, store_postgres.go:193-240). Anyone who knows the email can trigger it at will. Interleave it with a legitimate write: the victim changes password (`ChangePassword` writes the new hash), an in-flight failed login for the same email then writes its stale snapshot with the old hash, old MFA state or old `group_ids`, silently reverting the change; the same race lets an admin's deactivation or group removal be undone by the stale write. It also loses lockout increments (two concurrent failures both write `n+1`), so the 5-attempt lockout is soft under parallel guessing. +- evidence: + ```go + func (s *Service) recordFailedLogin(ctx context.Context, user *User) { + user.FailedLoginAttempts++ + now := time.Now() + user.UpdatedAt = now + ... + if err := s.store.UpdateUser(ctx, user); err != nil { + ``` +- suggested fix: Give the store a targeted atomic statement for this path (`UPDATE users SET failed_login_attempts = failed_login_attempts + 1, locked_until = CASE ... WHERE id = $1`) and use it from `recordFailedLogin` and `completeSuccessfulLogin` instead of the full-row write. +- verdict: CONFIRMED — recordFailedLogin (service_user.go:755-770) writes the *User loaded at Login:158 through PostgresStore.UpdateUser, which sets all 18 columns unconditionally (store_postgres.go:193-240, no version or WHERE guard), and it is reachable unauthenticated from any wrong-password attempt (service.go:230-232), so the lost-update and stale-overwrite interleavings need only two concurrent requests. +- issue: (pending cross-reference) + +### A05-006 Revocation-finalize retry loop sleeps without honouring the context +- category: concurrency +- severity: medium +- location: internal/purchase/finalize_revocations.go:62 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `FinalizeInFlightRevocations` retries `MarkPurchaseRevoked` with `time.Sleep(2s)` then `time.Sleep(6s)`, and neither the retry loop nor the outer row loop looks at `ctx`. With 30 in-flight rows whose store writes are failing (the DB outage that produced the in-flight backlog in the first place), the sweep blocks for ~4 minutes past a canceled Lambda context, doing three doomed writes per row against a dead connection, instead of returning immediately. This is the exact bare-sleep-with-ctx-in-scope shape the project's own feedback memory records CodeRabbit catching repeatedly. +- evidence: + ```go + for attempt, backoff := range finalizeRevocationBackoffs { + if markErr == nil { + break + } + logging.Warnf("finalize_revocations: MarkPurchaseRevoked attempt %d for %s failed: %v (retrying in %s)", + attempt+1, record.PurchaseID, markErr, backoff) + time.Sleep(backoff) + markErr = m.config.MarkPurchaseRevoked(ctx, record.PurchaseID, now, "direct-api", "", nil, "") + } + ``` +- suggested fix: replace the sleep with `select { case <-time.After(backoff): case <-ctx.Done(): return result, ctx.Err() }` and add a `ctx.Err()` check at the top of the per-row loop so a canceled sweep stops instead of grinding through the remaining rows. +- verdict: CONFIRMED — finalize_revocations.go:53-73 is the whole loop: `ctx` appears only as the first argument to MarkPurchaseRevoked, the retry loop's only pause is the bare `time.Sleep(backoff)` at :62 over `finalizeRevocationBackoffs = {2s, 6s}` (:24-27), and neither the per-row `for _, record := range rows` at :54 nor the retry loop consults `ctx.Done()` / `ctx.Err()`, so 8s of doomed sleeping per row survives cancellation. +- issue: (pending cross-reference) + +### A06-004 DB rate-limiter cleanup worker is started with a request-scoped context and never ticks +- category: concurrency +- severity: medium +- location: internal/server/app.go:754 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `reinitializeAfterConnect` runs inside `ensureDB`, which is called from `handleHTTPRequest` with `ctx, cancel := context.WithTimeout(r.Context(), 30*time.Second)` (internal/server/http.go:158) or from the Lambda invocation context. `StartCleanupWorker(ctx)` returns when `ctx.Done()` fires (internal/api/db_rate_limiter.go:69) and its ticker period is 10 minutes (`dbRateLimiterScheduledCleanupInterval`), so the goroutine is always canceled 30s (or one invocation) after the first DB connect and never runs a single cleanup. The documented 02-M2 mitigation — evicting perpetually-denied keys whose count never resets — is dead in production and `rate_limits` grows without bound under sustained abuse. Secondarily, a failure later in `reinitializeAfterConnect` (e.g. the `awsconfig.LoadDefaultConfig` at app.go:778) starts a fresh worker goroutine on every retry. +- evidence: + ```go + if app.appConfig.IsLambda { + dbRL := api.NewDBRateLimiter(dbConn.Pool()) + dbRL.StartCleanupWorker(ctx) + app.RateLimiter = dbRL + ``` +- suggested fix: start the worker with a process-lifetime context stored on `Application` (canceled in `Close`), not the caller's request context. +- verdict: CONFIRMED — `StartCleanupWorker(ctx)` receives the ctx `ensureDB` was called with (internal/server/app.go:750-755, reached from internal/server/app.go:633), which is the 30s request context built at internal/server/http.go:158 or the per-invocation Lambda context, while the ticker is 10 minutes (internal/api/db_rate_limiter.go:35,65), so the `select` at db_rate_limiter.go:68-73 always takes `ctx.Done()` before its first tick and `cleanup()` never runs. +- issue: (pending cross-reference) + +### A12-030 Two overlapping `loadHistory()` fetches have no sequence guard +- category: concurrency +- severity: medium +- location: frontend/src/history.ts:159 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `setupHistoryHandlers` subscribes `loadHistory` to both the provider and the account slice. Changing the provider chip clears the accounts and sets the provider synchronously (the issue #185 ordering rule), firing both subscribers, so two `api.getHistory` calls with different filters are in flight. Neither is sequenced, so the slower first response can land last and repaint the Approval Queue with rows for accounts the user just filtered out; `lastPurchases`, which the marketplace Sell dialog reads for its price lookup, is then also stale. +- evidence: + ```ts + state.subscribeProvider(() => void loadHistory()); + state.subscribeAccount(() => void loadHistory()); + ``` +- suggested fix: Keep a monotonically increasing request id, capture it before the fetch, and skip all renders when it is no longer the latest — the pattern `topbar-filters.ts:92` already uses. +- verdict: CONFIRMED — Both slices are subscribed to loadHistory (frontend/src/history.ts:159-160) and the provider chip's onChange calls setCurrentAccountIDs([]) then setCurrentProvider synchronously (frontend/src/topbar-filters.ts:165-166), so two getHistory calls with different filters are in flight with no request-id guard on either. +- issue: (pending cross-reference) + +### A12-031 Money-mutation buttons are disabled only after the confirm dialog resolves +- category: concurrency +- severity: medium +- location: frontend/src/history.ts:1242 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The Approve handler awaits a network `buildApprovalDetailsBody(id)` and then `confirmDialog` before it touches `btn.disabled`. `confirmDialog` appends a fresh backdrop per call with no singleton guard (`confirmDialog.ts:81`), so a double-click during that round trip stacks two modals and confirming both fires `api.approvePurchase(id)` twice. Retry (1336), Revoke (1404) and Sell (1440) have the same shape, and Retry mints a new execution. +- evidence: + ```ts + const detailsBody = await buildApprovalDetailsBody(id); + const ok = await confirmDialog({ title: 'Approve this pending purchase?', ... }); + if (!ok) return; + const rowActions = sameRowActions(btn); + rowActions.forEach((b) => { b.disabled = true; }); + ``` +- suggested fix: Disable the row's action buttons synchronously on click, re-enabling on cancel or failure. +- verdict: CONFIRMED — The Approve handler awaits buildApprovalDetailsBody and confirmDialog before touching btn.disabled (frontend/src/history.ts:1242-1258), Retry (:1340-1348), Revoke and Sell (:1440-1533) have the same shape, and confirmDialog appends a fresh backdrop per call with no singleton guard (frontend/src/confirmDialog.ts:81). +- issue: (pending cross-reference) + +### A01-018 persistAzureRevocation sleeps with a bare time.Sleep after money moved, ignoring ctx +- category: concurrency +- severity: low +- location: internal/api/handler_purchases_revoke.go:776 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Azure has already returned the reservation and `MarkPurchaseRevoked` fails. The retry loop blocks 1s+3s+9s with `time.Sleep` while the Lambda request context may already be cancelled; the subsequent store calls then fail on the dead ctx, the 13 seconds are wasted, and the 207 RECONCILE_PENDING response arrives late or not at all. +- evidence: + ```go + for attempt, backoff := range revokeMarkRetryBackoffs { + if markErr == nil { + break + } + logging.Warnf(...) + time.Sleep(backoff) + markErr = h.config.MarkPurchaseRevoked(ctx, record.PurchaseID, now, "direct-api", "", calcRefundAmount, calcRefundCurrency) + } + ``` +- suggested fix: Use the ctx-aware `select { case <-time.After(backoff): case <-ctx.Done(): break }` form and stop retrying once ctx is done. +- verdict: CONFIRMED — persistAzureRevocation (internal/api/handler_purchases_revoke.go:767-780) loops revokeMarkRetryBackoffs (1s/3s/9s, :126-130) with a bare time.Sleep and never checks ctx.Done() before re-calling MarkPurchaseRevoked with the same ctx. +- issue: (pending cross-reference) + +### A02-017 A request-scoped context seals the process-wide AWS config cache +- category: concurrency +- severity: low +- location: internal/api/handler.go:1062 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The first request that reaches `resolveAWSCallerIdentity` or `getLambdaInvoker` (handler_recommendations_refresh.go:225-227) runs `awsconfig.LoadDefaultConfig(ctx)` inside `awsCfgOnce` with that request's context. If the request is cancelled or hits its deadline during the load (IMDS/SSO/credential-process resolution on ECS or dev hosts), `awsCfgErr` holds `context canceled` for the container's lifetime, so every later reshape scope resolution, deployment-info call and async refresh fails until a cold start. The comment at handler.go:903-912 describes this exact hazard and avoids it for `loadAPIKey`, but the request path still has it. +- evidence: + ```go + h.awsCfgOnce.Do(func() { + h.awsCfg, h.awsCfgErr = awsconfig.LoadDefaultConfig(ctx) + }) + ``` +- suggested fix: Load with `context.Background()` plus a bounded timeout inside the `Once`, or do not cache the error (reset the `Once` on a context error) so the next request retries. +- verdict: PLAUSIBLE — resolveAWSCallerIdentity (handler.go:1062-1064) and getLambdaInvoker (handler_recommendations_refresh.go:225-227) both seal awsCfgOnce with the request ctx and the comment at handler.go:903-912 records this exact failure mode for loadAPIKey; but LoadDefaultConfig does little network I/O (an IMDS region probe only when no region is configured; credentials resolve lazily at first Retrieve), so a cancellation landing inside the Once needs a host with no AWS_REGION and a slow IMDS, which I could not establish from source. +- issue: (pending cross-reference) + +### A09-025 logging.SetOutput mutates shared logger state without synchronization +- category: concurrency +- severity: low +- location: pkg/logging/logger.go:272 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the package doc explains that `level` is an `atomic.Int32` specifically because logging happens "from many goroutines" during the fan-out, but `output` is a plain field. `SetOutput` reads and writes it unguarded while `Logger.With` (line 128) concurrently reads `l.output` to build a derived logger. Under `-race`, a test that calls `SetOutput` to capture output while a background collector goroutine logs reports a data race on `defaultLogger.output`, and the derived logger can capture a torn or stale writer. +- evidence: + ```go + func SetOutput(w io.Writer) io.Writer { + prev := defaultLogger.output + defaultLogger.output = w + defaultLogger.logger.SetOutput(w) + return prev + } + ``` +- suggested fix: store the writer in an `atomic.Value` (or guard it with the same discipline as `level`) so `SetOutput` and `With` do not race, matching the reasoning already applied to the level field. +- verdict: CONFIRMED — `Logger.output` is a plain `io.Writer` field alongside the `atomic.Int32` level (pkg/logging/logger.go:36-42); `SetOutput` reads and writes it with no synchronization (logger.go:272-277) while `Logger.With` reads `l.output` to build the derived logger (logger.go:125-130). Narrowing note: the embedded `*log.Logger` has its own internal mutex, so the actual emit path is safe — the race is confined to the `output` field and the derived logger it seeds. +- issue: (pending cross-reference) + +### Category: performance + +9 findings: 5 medium, 4 low. + +### A07-011 Sparkline fetch issues one full region-wide coverage query per instance type and discards all but one row +- category: performance +- severity: medium +- location: providers/aws/recommendations/usage_history.go:79 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `dailyUsageFilter` scopes only to `(service, region)`; the instance type is matched client-side in `applyPeriodsToDayMap`. `AttachDailyUsageHistory` calls `GetDailyUsagePcts` once per unique `(service, region, resourceType)` tuple, so 60 distinct EC2 SKUs in one region produce 60 identical CE queries that each return the whole region's coverage and throw away 59/60 of it. At CE's per-request charge that is 60 billed calls where one would do, on every recommendation refresh. +- evidence: + ```go + for _, k := range order { + if ctx.Err() != nil { + return + } + pcts, err := c.GetDailyUsagePcts(ctx, k.service, k.resourceType, k.region) + ``` +- suggested fix: fetch once per `(service, region)`, build a map from instance type to the daily series, and fan the result out to every rec sharing that pair. +- verdict: CONFIRMED — `dailyUsageFilter` (providers/aws/recommendations/usage_history.go:197-210) scopes only SERVICE and REGION, `applyPeriodsToDayMap` :126 discards non-matching instance types client-side, and `AttachDailyUsageHistory` :176-180 issues one `GetDailyUsagePcts` per unique `(service, region, resourceType)` from `groupRecsByTuple` :155, on the live refresh path at client.go:339. +- issue: (pending cross-reference) + +### A07-015 OpenSearch idempotency lookup pages the whole reservation inventory with no cap, on the purchase path +- category: performance +- severity: medium +- location: providers/aws/services/opensearch/client.go:242 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `DescribeReservedInstances` has no name filter, so `findReservationByName` walks every page of every reservation in the account before each purchase, with no page ceiling. In an account with thousands of historical reservations this consumes the Lambda budget that `purchasecfg` was introduced to protect, and the purchase fails with a deadline error attributed to `PurchaseCommitment` rather than to the guard. Every other pagination loop on this shard's purchase paths carries a `maxOfferingPages` or `maxCommitmentPages` guard. +- evidence: + ```go + func (c *Client) findReservationByName(ctx context.Context, name string) (string, bool, error) { + var nextToken *string + for { + response, err := c.client.DescribeReservedInstances(ctx, &opensearch.DescribeReservedInstancesInput{ + NextToken: nextToken, + MaxResults: 100, + }) + ``` +- suggested fix: add a page cap and a `ctx.Err()` check at the top of the loop, and fail loud when the cap is hit (a truncated guard must not report "not found" and let the purchase proceed). +- verdict: CONFIRMED — `findReservationByName` (providers/aws/services/opensearch/client.go:242-268) has neither a page counter nor a `ctx.Err()` check, unlike the offering loop in the same file at :403 which guards with `maxOfferingPages` (:378), and `idempotencyGuard` :275-285 calls it before every tokened purchase. +- issue: (pending cross-reference) + +### A07-016 Redshift idempotency guard makes one DescribeTags call per reserved node, unbounded +- category: performance +- severity: medium +- location: providers/aws/services/redshift/client.go:260 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `findNodeByIdempotencyToken` pages all reserved nodes with no cap and calls `nodeHasIdempotencyTag` (a `DescribeTags` round trip) for each active node until it finds a match. An account with 200 active reserved nodes issues up to 200 sequential `DescribeTags` calls before every Redshift purchase, at 15s HTTP timeout each in the worst case. A slow or throttled `DescribeTags` mid-scan surfaces as a lookup failure and blocks a legitimate first-time purchase. +- evidence: + ```go + arn := fmt.Sprintf("arn:aws:redshift:%s:%s:reservednode:%s", c.region, accountID, nodeID) + tagged, err := c.nodeHasIdempotencyTag(ctx, arn, token) + if err != nil { + return "", false, fmt.Errorf("failed to read tags for reserved node %s: %w", nodeID, err) + } + ``` +- suggested fix: query `DescribeTags` once with `ResourceType: "reservednode"` and the token as a tag-value filter, rather than per node, and add a page cap to the outer loop. +- verdict: CONFIRMED — `scanNodesForToken` (providers/aws/services/redshift/client.go:259-278) issues one `DescribeTags` per active node and the outer `DescribeReservedNodes` loop at :237-252 has no page cap; `purchasecfg.HTTPTimeout` is 15s (providers/aws/internal/purchasecfg/config.go:32) and the error return at :271-273 aborts the guard, which `findNodeByIdempotencyToken` :226-230 turns into a blocked purchase. +- issue: (pending cross-reference) + +### A07-018 Savings Plans coverage and utilization calls bypass the shared concurrency semaphore +- category: performance +- severity: medium +- location: providers/aws/recommendations/sp_coverage.go:361 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `fetchCoveragePage`, `fetchUtilizationPage` and `fetchRIPageWithRetry` all wrap their SDK call in `concurrency.Acquire` / `concurrency.Release` so CE traffic stays under `CUDLY_MAX_PARALLELISM`. `fetchSPCoveragePage` and `fetchSPUtilizationPage` (line 450) do not. When the ladder runs concurrently with a recommendation sweep, the SP calls are invisible to the cap and push total CE concurrency above the configured limit, which is what the cap exists to prevent. +- evidence: + ```go + rateLimiter := c.rateLimiter.newOperation() + for { + if waitErr := rateLimiter.Wait(ctx); waitErr != nil { + return nil, fmt.Errorf("rate limiter wait failed: %w", waitErr) + } + result, err := c.costExplorerClient.GetSavingsPlansCoverage(ctx, input) + if !rateLimiter.ShouldRetry(err) { + ``` +- suggested fix: wrap both SDK calls in `concurrency.Acquire`/`Release` exactly as `fetchCoveragePage` does. +- verdict: CONFIRMED — `fetchSPCoveragePage` (providers/aws/recommendations/sp_coverage.go:361-374) and `fetchSPUtilizationPage` (:450-464) hold only the rate limiter, while `fetchCoveragePage` wraps its SDK call in `concurrency.Acquire`/`Release` at coverage.go:331-336 and documents the shared `CUDLY_MAX_PARALLELISM` cap at :322-324. +- issue: (pending cross-reference) + +### A12-022 `initRegistrations` stacks a filter listener on every visit to the Accounts tab +- category: performance +- severity: medium +- location: frontend/src/modules/registrations.ts:256 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `#registrations-status-filter` is static markup in `index.html:714` and is never replaced. `navigation.ts:300` calls `loadAccountsTab()` on each activation of the Accounts tab, which calls `initRegistrations()` (`settings.ts:3346`). Each call adds another `change` listener to the same element, so after N visits a single dropdown change fires N concurrent `loadRegistrations()` calls and N `GET /api/registrations` requests, whose responses race to render the same container. +- evidence: + ```ts + export function initRegistrations(): void { + const filterEl = document.getElementById('registrations-status-filter'); + filterEl?.addEventListener('change', () => void loadRegistrations()); + void loadRegistrations(); + } + ``` +- suggested fix: Guard with a module-level `wired` flag or a `dataset` marker on the element, matching the `wireRefreshButton` pattern in `inventory.ts:141`. +- verdict: CONFIRMED — initRegistrations unconditionally adds a `change` listener to the static `#registrations-status-filter` (frontend/src/modules/registrations.ts:254-257), which lives outside the container loadRegistrations rewrites (frontend/src/index.html:714), and switchSettingsSubTab re-runs loadAccountsTab on every Accounts activation with no self-switch guard (frontend/src/navigation.ts:269-301 into frontend/src/settings.ts:3346). +- issue: (pending cross-reference) + +### A01-019 RI exchange history is a capped 500-row fetch filtered in Go, so scoped users silently lose rows +- category: performance +- severity: low +- location: internal/api/handler_ri_exchange.go:2064 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A deployment with more than 500 exchange records in the last year. `GetRIExchangeHistory(ctx, since, 500)` returns the newest 500 across all accounts; the allowed_accounts filter then runs in memory, so a scoped user whose account's exchanges are older than the newest 500 sees an empty or truncated history with no indication of truncation. +- evidence: + ```go + since := time.Now().AddDate(-1, 0, 0) + records, err := h.config.GetRIExchangeHistory(ctx, since, 500) + // ... + if !allowed.AllowsAll() { + nameByID := h.resolveAccountNamesByID(ctx) + filtered := records[:0] + ``` +- suggested fix: Push the account predicate into the store query (pass the allowed account IDs) and apply the limit after filtering. +- verdict: CONFIRMED — getRIExchangeHistory (internal/api/handler_ri_exchange.go:2063-2085) calls GetRIExchangeHistory(ctx, since, 500), whose SQL (internal/config/store_postgres.go:2731-2745) is ORDER BY created_at DESC LIMIT $2 with no account predicate, then filters by allowed_accounts in memory with no truncation indicator. +- issue: (pending cross-reference) + +### A05-017 OpenSearch and Redshift probes filter client-side under a five-page cap +- category: performance +- severity: low +- location: internal/commitmentopts/probe.go:283 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Both APIs have no instance-type filter, so the probe fetches unfiltered pages and discards non-matching rows in Go, while `walkPaginated` stops after `maxPages` (5) × `pageSize` (100). If the region's `t3.small.search` (or `dc2.large`) offerings sort past offering 500, the probe yields zero combos for that service. `Save` still writes the singleton probe-run row, `HasData` becomes true, and nothing ever re-probes — the cache is permanently warm and permanently missing that service. The cap's stated purpose is bounding API spend when pagination detection breaks, which the server-side-filtered probers do not need it for; it is only load-bearing here, where it truncates real data. +- evidence: + ```go + for _, o := range out.ReservedInstanceOfferings { + if string(o.InstanceType) != probeTargetOpenSearch { + continue + } + ``` +- suggested fix: for the two client-side-filtered probers, stop early once all six (term, payment) combos have been seen and otherwise let pagination run to completion, or raise the cap for those two and log when it is hit so the truncation is visible rather than silent. +- verdict: PLAUSIBLE — every mechanical link checks out: walkPaginated hard-stops at `page < maxPages` with maxPages=5 and pageSize=100 (probe.go:23, :27, :86) and never logs when it hits the cap; OpenSearch (probe.go:283) and Redshift (:338) discard non-matching rows in the closure while the SP prober filters server-side via PlanTypes (:540); and the cache is permanently warm because Save inserts the singleton probe-run row (store_postgres.go:110) that HasData (:80-83) reads as "data exists". The unestablished condition is external: whether AWS actually returns `t3.small.search` / `dc2.large` offerings past position 500 in a given region, which I cannot determine from source. +- issue: (pending cross-reference) + +### A12-052 `parseNumericFilter` is re-parsed for every row of every numeric filter +- category: performance +- severity: low +- location: frontend/src/lib/column-filters.ts:119 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The parse call sits inside the per-row predicate. Recommendations, Plans, History and RI Exchange all route through this helper, so a 500-row Opportunities table with three active numeric filters performs 1,500 regex-parse plus closure-allocation cycles on every render, and every popover commit re-renders. The parse result is loop-invariant. +- evidence: + ```ts + return rows.filter((row) => { + for (const [col, filter] of entries) { + ... + } else { + const parsed = parseNumericFilter(filter.expr); + if (!parsed.ok) continue; + ``` +- suggested fix: Hoist the parse above `rows.filter`, building a `[col, predicate][]` array once. +- verdict: CONFIRMED — parseNumericFilter is called inside the per-row predicate while `entries` is hoisted once above it (frontend/src/lib/column-filters.ts:110, :119-121), so the parse is provably loop-invariant work repeated per row per numeric filter. +- issue: (pending cross-reference) + +### A12-071 Per-render `indexOf` inside the RI recommendation row map is O(n²) +- category: performance +- severity: low +- location: frontend/src/riexchange.ts:1317 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Each `renderRecommendations` call performs `currentRecommendations.indexOf(rec)` once per visible row. Combined with A12-052's per-row filter re-parse, a few hundred recommendations with three active numeric filters cost thousands of redundant scans and regex compiles per filter interaction and per utilization re-render. +- evidence: + ```ts + const visibleWithIdx = visible.map((rec) => ({ + rec, + idx: currentRecommendations.indexOf(rec), + })); + ``` +- suggested fix: Carry the index through the filter by mapping to `{rec, idx}` before filtering. +- verdict: CONFIRMED — currentRecommendations.indexOf(rec) runs once per visible row inside the render map (frontend/src/riexchange.ts:1313-1317), making the row build quadratic in the recommendation count on every filter interaction and every utilization re-render. +- issue: (pending cross-reference) + +### Category: test-gap + +29 findings: 8 high, 13 medium, 8 low. + +### A05-002 `singleCloudAccountIDFromRecs` is unit-tested in isolation, so the ambient fallback above stays green +- category: test-gap +- severity: high +- location: internal/purchase/execution_test.go:1628 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The only test of the two-distinct-accounts case asserts the helper returns nil and labels that "multi-account, not this path". There is no test that drives `executeSingleAccount` (or `ApproveAndExecute` / `directExecutePurchase`) with a mixed-account rec set and asserts the purchase is refused, so A05-001 passes CI. `TestMultiAccountSeedsStablePerAccountKey` and the other money-path regressions all exercise the plan fan-out, which reaches `executeMultiAccount`, never this branch. +- evidence: + ```go + { + name: "two distinct account IDs returns nil (multi-account, not this path)", + recs: []config.RecommendationRecord{ + {CloudAccountID: &aid1}, + {CloudAccountID: &aid2}, + }, + want: nil, + }, + ``` +- suggested fix: add a regression test that calls `executeSingleAccount` with recs naming two accounts and asserts `CreateAndValidateProvider` is never called (register the expectation first so the assertion is not vacuous). +- verdict: CONFIRMED — I opened the test files rather than grepping: `git grep -n "singleCloudAccountIDFromRecs\|resolveSingleAccountProvider\|executeSingleAccount" -- 'internal/**/*_test.go'` returns only execution_test.go:1279, :1294 and :1641, and the two-distinct-IDs case at execution_test.go:1628-1634 asserts the helper in isolation; the eight tests in money_path_regression_test.go (:86 through :540) all drive the plan fan-out or the idempotency key, none reaches the mixed-account single-account branch. +- issue: (pending cross-reference) + +### A16-001 4-eyes: the "creator's account could not be resolved" fail-closed branch has no test +- category: test-gap +- severity: high +- location: internal/purchase/approvals.go:284 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `checkDifferentApprover`'s tier-2 (email) path is the one every token/SQS + approver takes. Blocks 281 and 285-287 are both uncovered: no test makes `GetUserEmailByID` return + an error, and none makes it return `("", nil)`. Delete the emptiness guard and + `strings.EqualFold("", actorEmail)` is false for every real actor, so 4-eyes silently *passes* for + any execution whose creator row is missing, deleted, or unresolvable — exactly the deleted-creator + case dual control exists for. `("", nil)` from a user lookup is an established shape in this repo, + so this is not a hypothetical input. Every existing four-eyes test + (internal/purchase/approvals_test.go:337-560) stubs `GetUserEmailByID` with a real address or + asserts it is never called. +- evidence: + ```go + creatorEmail, err := m.config.GetUserEmailByID(ctx, *execution.CreatedByUserID) + if err != nil { + return fmt.Errorf("4-eyes policy check: failed to resolve creator identity: %w", err) + } + creatorEmail = strings.TrimSpace(creatorEmail) + if creatorEmail == "" { + logging.Warnf("purchase[%s]: 4-eyes mode on; creator account %s not found, denying (fail-closed)", + executionID, *execution.CreatedByUserID) + return fmt.Errorf("4-eyes approval mode is enabled but the creator's account could not be resolved; an admin must investigate before approving") + } + ``` +- suggested fix: add two subtests to `internal/purchase/approvals_test.go` on the + `ApproveExecution` (tier-2) shape: `GetUserEmailByID` returning `("", nil)`, and returning an + error. Assert the error text and `AssertNotCalled` on `TransitionExecutionStatus`. +- verdict: CONFIRMED — a merged cross-package profile (`internal/api`, `internal/server`, `internal/scheduler` and `internal/purchase` together, `-coverpkg` over all four) still shows approvals.go:281.3-282.1 and 285.3-288.1 at 0, and the scenario is not hypothetical: internal/config/store_postgres.go:1587-1588 returns `("", nil)` on `pgx.ErrNoRows`, so a deleted creator row lands on exactly this branch, while every stub in internal/purchase/approvals_test.go:515,550 and coverage_extra_test.go:422,476 returns a real address. +- issue: (pending cross-reference) + +### A16-004 The scheduled-purchase fire success path is behind a t.Skip, and its stand-in guard only checks a signature +- category: test-gap +- severity: high +- location: internal/purchase/scheduled_fire_test.go:157 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `result.Fired++` (internal/purchase/scheduled_fire.go:60) and `fireOneDue`'s + success return (scheduled_fire.go:105) are both uncovered. The four live tests in this file cover + only no-rows, list-error, CAS-race-lost and hard-DB-error; the one test that would exercise a + purchase actually firing is skipped, and its own comment calls it "the end-to-end smoke test that + verifies the fire-tick path does not silently no-op the pre-fire delay branch (CRITICAL)". So the + delayed-purchase feature can stop spending entirely — `fireOneDue` returning `(false, false)` on + every row, or `executeAndFinalize` erroring for all of them — and the suite is green with + `Fired == 0`, which is what three of the four live tests already assert. The compensating test + named `TestFireScheduledDelayedPurchases_DelayPathNotSilentNoOp` (line 175) is a compile-time + method-signature assertion; it cannot observe a no-op despite its name. +- evidence: + ```go + func TestFireScheduledDelayedPurchases_EndToEnd(t *testing.T) { + t.Skip("placeholder until full provider-stub wiring is available; " + + "the CAS and audit-stamp paths are covered by the unit tests above") + // When un-skipped, the test scenario is: + // 1. Create an execution with purchase_delay_hours > 0, Status="scheduled", + // ScheduledExecutionAt = time.Now().Add(-1h). + // 2. Call FireScheduledDelayedPurchases(ctx). + // 3. Assert result.Fired == 1, result.RaceLost == 0, result.Errored == 0. + ``` +- suggested fix: un-skip it using the provider/email wiring `money_path_regression_test.go` already + builds (`awsAccessKeyCredStore` plus the `MockProviderFactory` chain), and assert `Fired == 1` + plus exactly one `PurchaseCommitment` call. Separately assert that a `SavePurchaseExecution` + failure after a winning CAS (the uncovered AUDIT GAP branch at scheduled_fire.go:97) still fires. +- verdict: CONFIRMED — scheduled_fire.go:60.4-61.1 (`result.Fired++`) and 105.2-105.20 (the success return) are both 0 in the merged cross-package profile, so nothing in `internal/server` or `internal/scheduler` reaches them either. One correction that does not change the verdict: a fifth live caller exists outside the named file, armed_redrive_test.go:181, but it asserts `result.Fired == 0` and `result.Errored == 1`, so it joins rather than closes the set of tests a total no-op would satisfy. +- issue: (pending cross-reference) + +### A16-005 FinalizeInFlightRevocations has zero coverage; its only "test" asserts a stub's own return value +- category: test-gap +- severity: high +- location: internal/server/handler_test.go:142 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: every statement of `internal/purchase/finalize_revocations.go` (lines 46-75) is + uncovered. The only exercise of the feature is the scheduled-task dispatch test above, which + installs a fake returning `FinalizeResult{Found: 1, Finalized: 1}` and then asserts only that + `HandleScheduledTask` returned no error — an assertion on the fake's own return value that holds + whatever the real sweep does. Consequences that stay green: the `Found`/`Finalized`/`Errored` + accounting can be wrong in any direction; the retry loop's `time.Sleep(backoff)` + (finalize_revocations.go:62) is not ctx-aware and the loop never checks `ctx.Err()`, so a Lambda + whose context is already cancelled still burns up to 8s per row and can be killed mid-sweep; and a + row that never finalizes is counted rather than surfaced. These rows are the audit record for + Azure reservations already returned to the provider, so a silent failure leaves the DB claiming + money was never refunded. +- evidence: + ```go + setupMocks: func(s *testutil.MockScheduler, p *testutil.MockPurchaseManager) { + p.FinalizeInFlightRevocationsFunc = func(ctx context.Context) (*purchase.FinalizeResult, error) { + return &purchase.FinalizeResult{Found: 1, Finalized: 1}, nil + } + }, + expectError: false, + ``` +- suggested fix: add `internal/purchase/finalize_revocations_test.go` driving the real method against + `MockConfigStore`: three in-flight rows where one succeeds first try, one succeeds on retry, one + fails all three attempts; assert `Found=3, Finalized=2, Errored=1` and that the sweep did not abort. + Add a cancelled-context case once the sleep is made ctx-aware. +- verdict: CONFIRMED — no `finalize_revocations_test.go` exists, `FinalizeInFlightRevocations` appears in no test outside the two fakes at internal/server/handler_test.go:142,152, and nothing under tests/, pkg/, cmd/ or mcp/ names it or `GetPurchaseHistoryInFlight`; every statement block in finalize_revocations.go is 0 in the merged profile. The ctx-blindness is in the source as described: the loop at finalize_revocations.go:56-64 calls `time.Sleep(backoff)` with no `ctx.Err()` check, and the two backoffs sum to 8s per stuck row. +- issue: (pending cross-reference) + +### A16-006 The AUDIT LOSS branch in the per-account fan-out is untested, and it is the one that decides SQS ack vs redelivery +- category: test-gap +- severity: high +- location: internal/purchase/execution.go:269 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: block 269-270 is uncovered. This is the only place where `executeForAccount` + returns `committed = true` *together with* an error. `TestMultiAccountPartialSuccessIsAcked` + (internal/purchase/money_path_regression_test.go:283) proves the >=1-committed run is ACKed, but + only via the clean save path — its `SavePurchaseExecutionFn` always returns nil. Change line 269 to + `return false, ...` (the natural-looking "we failed, report failure" edit) and a run whose cloud + purchase succeeded but whose row failed to persist is reported as fully failed, the SQS message is + redelivered, and the redelivery re-buys the commitment. Nothing in the suite fails, because no test + ever makes `SavePurchaseExecution` fail on a committed account. +- evidence: + ```go + partial, committed := applyAccountOutcome(&acctExec, purchaseErrors) + + if saveErr := m.config.SavePurchaseExecution(ctx, &acctExec); saveErr != nil { + return committed, fmt.Errorf("AUDIT LOSS: failed to save execution record for account %s: %w", account.ID, saveErr) + } + ``` +- suggested fix: clone `TestMultiAccountPartialSuccessIsAcked` with a `SavePurchaseExecutionFn` that + errors for the per-account row after `PurchaseCommitment` succeeded, and assert + `handleExecutePurchase` still returns nil (ack) and that no second `PurchaseCommitment` occurs. +- verdict: CONFIRMED — execution.go:269.3-270.1 is 0 in the merged profile. Every `SavePurchaseExecutionFn` in the package either returns nil unconditionally or records and returns nil (execution_test.go:869,1038,1208,1826; money_path_regression_test.go:43,167,249,326,409,567; armed_redrive_test.go:101,245); the one stub that returns an error, execution_test.go:923, drives the credential-failure path and its assertion at line 958 lands on the other AUDIT LOSS site, execution.go:300, not on 269. +- issue: (pending cross-reference) + +### A16-007 The partial-sweep eviction guard is proven for Azure only; the GCP per-account collector duplicating it has no test +- category: test-gap +- severity: high +- location: internal/scheduler/scheduler.go:912 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `internal/scheduler/partial_sweep_eviction_test.go` is a careful guard for the + invariant "an incomplete sweep must not authorize stale-row eviction", but every case drives Azure: + `collectAzureRecommendations`, `fetchAndConvert(..., "azure", ...)` (scheduler_test.go:756, 782) and + `tolerateIncompleteSweep("azure", ...)` (scheduler_test.go:800-814). `collectGCPForAccount` carries + its own inlined copy of the same plumbing and is entirely uncovered (lines 884-912), as is + `collectAWSForAccount` (752-770). Hardcode `true` in place of `complete` at line 912 and a partial + GCP sweep enters `outcome.SucceededAccountIDs`, which authorizes the + `DELETE ... WHERE collected_at < $1 AND (provider, account_key) IN (…)` that + `UpsertRecommendations` runs — deleting the previous-cycle recommendations for every project the + sweep never queried. The whole Azure suite stays green. +- evidence: + ```go + recs, err := recClient.GetAllRecommendations(ctx) + complete, err := tolerateIncompleteSweep("gcp", err) + if err != nil { + return nil, false, fmt.Errorf("get recommendations: %w", err) + } + return s.tagAccount(s.convertRecommendations(recs, "gcp"), acct.ID), complete, nil + ``` +- suggested fix: parameterise `newPartialSweepScheduler` over the provider name and run the same + three cases (partial, complete, hard error) through `collectGCPForAccount` and + `collectAWSForAccount`, asserting the returned `complete` bool directly. +- verdict: CONFIRMED — scheduler.go:912.2-912.84 and the whole body of `collectGCPForAccount` from 883 are 0 in the merged profile, as is `collectAWSForAccount`'s tail at 770. partial_sweep_eviction_test.go names only Azure: its provider stub, its three `collectAzureRecommendations` cases at lines 106, 128 and 174, and its `fanOutPerAccount` case at 148 all pass `"azure"`. The shared helper `tolerateIncompleteSweep` is directly tested at scheduler_test.go:800-814, so what is untested is the per-provider wiring of its `complete` return, which is exactly what the finding claims. +- issue: (pending cross-reference) + +### A16-008 getPlannedPurchases' account-scope filter is never exercised; the only test runs as an unrestricted admin +- category: test-gap +- severity: high +- location: internal/api/handler_purchases.go:147 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: blocks 140 (`plan == nil`), 144 (scope-check error) and 147 (`!ok`, the scope + exclusion) are all uncovered. `TestHandler_getPlannedPurchases` + (internal/api/handler_purchases_test.go:910) is the sole test: it calls `mockAuth.grantAdmin()`, + supplies one plan and one execution, and asserts `Len(result.Purchases, 1)`. Because the fixture + contains no plan the session should be denied, the loop's exclusion branch is not merely + unasserted, it is unreachable — the negative guard is satisfied by an empty set. Make + `isPlanAllowedCached` return `true` unconditionally, or delete lines 146-148, and a per-account + scoped user reading `GET /api/purchases/planned` receives every tenant's plan names, step + schedules, estimated savings and upfront costs. The mutating twin + (`pausePlannedPurchase`) is covered by internal/api/handler_per_account_perms_test.go:770, so this + is a listing-vs-mutation asymmetry, not a wholly missing concern. +- evidence: + ```go + for _rvc := range executions { + exec := executions[_rvc] + plan := planMap[exec.PlanID] + if plan == nil { + continue + } + ok, err := h.isPlanAllowedCached(ctx, session, exec.PlanID, allowedPlan) + if err != nil { + return nil, err + } + if !ok { + continue + } + ``` +- suggested fix: add a scoped-session variant seeded with two plans on two different cloud accounts, + the session allowed only one, and assert the response contains exactly the allowed plan's ID. Assert + a non-zero expected count first so the test cannot pass on an empty result. +- verdict: CONFIRMED — handler_purchases.go:140.4, 144.4-145.1 and 147.4 are all 0 in the merged profile. One correction that strengthens rather than weakens the finding: line 910 is not the sole test, there are six callers of `getPlannedPurchases` (handler_purchases_test.go:968, 1009, 1067, 1141, 1865, 1930), and none of them can reach the exclusion — four run `mockAuth.grantAdmin()`, one errors before the loop, and the read-only one at 1930 stubs `GetAllowedAccountsAPI` to return nil, which `isPlanAllowedCached` reads as unrestricted. Every fixture also puts every execution's PlanID in the plan map, so `plan == nil` is unreachable too. +- issue: (pending cross-reference) + +### A16-009 authorizeSessionRevokeExecution's allow and deny boundaries are both untested +- category: test-gap +- severity: high +- location: internal/api/handler_purchases_revoke.go:320 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: within this function, only the admin-API-key short-circuit (303) and the + revoke-own creator check (326) are covered. The `hasAny` allow (312), the `!hasOwn` 403 (320) and + both permission-lookup error paths (309, 317) are uncovered. So no test proves that a session + holding *neither* `revoke-any` nor `revoke-own` is refused: delete lines 319-321 and any + authenticated user reaches the creator comparison, and one who happens to be the creator revokes a + scheduled purchase with no revoke permission at all. The sibling function for purchase-history + revoke, `authorizeSessionRevoke` (line 352), has both of these branches covered — same guard, two + entry points, only one of them tested. +- evidence: + ```go + hasAny, err := h.auth.HasPermissionAPI(ctx, session.UserID, auth.ActionRevokeAny, auth.ResourcePurchases) + if err != nil { + return fmt.Errorf("permission check failed: %w", err) + } + if hasAny { + return nil + } + hasOwn, err := h.auth.HasPermissionAPI(ctx, session.UserID, auth.ActionRevokeOwn, auth.ResourcePurchases) + if err != nil { + return fmt.Errorf("permission check failed: %w", err) + } + if !hasOwn { + return NewClientError(403, "permission denied: requires revoke-any or revoke-own on purchases") + } + ``` +- suggested fix: table-drive `authorizeSessionRevokeExecution` over (revoke-any granted → allow), + (neither granted → 403), (revoke-own granted + non-creator → 403), (revoke-own + creator → allow), + mirroring the coverage `authorizeSessionRevoke` already has. +- verdict: CONFIRMED — handler_purchases_revoke.go blocks at 309, 312, 317 and 320 are all 0 in the merged profile while 303 and 326 are 1, matching the finding exactly. Only three tests reach the function (handler_purchases_revoke_test.go:766 via `revokePurchase`, 782, 801) and every one of them stubs `revoke-any` false and `revoke-own` true, so the neither-granted 403 at line 320 is never exercised on a path that is plainly reachable in production for any authenticated non-admin session. +- issue: (pending cross-reference) + +### A02-026 No test exercises a scoped session with an explicit account filter on the dashboard, nor name-based scopes on marketplace sell-own +- category: test-gap +- severity: medium +- location: internal/api/handler_dashboard_test.go:636 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The only scoped dashboard tests (`GetAllowedAccountsAPI` at lines 636, 681, 1529) cover recommendations, upcoming purchases and the slice-aliasing regression; none sends `account_id`/`account_ids` for an out-of-scope account, so A02-001 is invisible to the suite. In handler_marketplace_test.go the scope fixtures hold `"acct-1"` / `"acct-other"` (lines 164, 200), i.e. the row's UUID, whereas production scopes hold names, so `TestMarketplaceList_SellOwnAllowed` passes while A02-003 denies every real sell-own user. Both tests would stay green with the bugs present. +- evidence: + ```go + // allowed accounts cover the row's cloud account. + authSvc.On("GetAllowedAccountsAPI", mock.Anything, "user-1"). + Return([]string{"acct-1"}, nil) + ``` +- suggested fix: Add `TestGetDashboardSummary_ScopedUser_ExplicitOutOfScopeAccountIsIgnored` (expect zeroed commitment KPIs and no `GetActivePurchaseHistory` call with the foreign id) and switch the marketplace scope fixtures to the account name with a `ListCloudAccounts` stub, confirming they fail on the current code first. +- verdict: CONFIRMED — every getDashboardSummary call in handler_dashboard_test.go (lines 79,142,186,1212,1257,1302,1352,1372,1396) passes either no account params or only provider, and the scoped tests at 636/681/1529 exercise filterDashboardRecommendations/upcoming only, so no test combines GetAllowedAccountsAPI scoping with account_id/account_ids; the marketplace fixtures (handler_marketplace_test.go:164-165,200-201) grant "acct-1"/"acct-other", which is the UUID branch of Allows, so both A02-001 and A02-003 stay green. +- issue: (pending cross-reference) + +### A03-025 Tests enumerate only the exact carved-out pairs and never exercise the failure paths above +- category: test-gap +- severity: medium +- location: internal/auth/group_ceiling_permissions_test.go:24 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Every carve-out test iterates the three exact `(verb, purchases)` pairs (group_ceiling_permissions_test.go:24-28, self_escalation_carveout_test.go:35-39, service_apikeys_test.go:1169-1188), so A03-001/002/003 pass the suite. `TestLogin_WithMFA_RecoveryCode_ConsumedOnce` (service_mfa_test.go:438-466) only asserts the in-memory slice is empty with a succeeding `UpdateUser`, so A03-008 passes. `TestService_UpdateUser/"update active status successfully"` (service_user_test.go:646-665) registers no `DeleteUserSessions` expectation, so A03-005 passes and would keep passing after the fix unless the expectation is added. `TestService_ResetTokenStatus` (service_password_test.go:542-557) asserts an inactive user with a valid token is the "invite" flow; nothing asserts a deactivated account is refused, so A03-006 passes. `TestResolveBastionProvider_LegacyFallback` (credentials/resolver_test.go:366-374) pins the A03-011 fallback as desired behaviour. No test covers `RequestPasswordReset` for an invited user (A03-009), MFA setup on an already-enabled user (A03-007), or a signer whose first KMS call fails and second succeeds (A03-014). +- evidence: + ```go + carvedOut := []APIPermission{ + {Action: ActionExecute, Resource: ResourcePurchases}, + {Action: ActionApproveAny, Resource: ResourcePurchases}, + {Action: ActionRetryAny, Resource: ResourcePurchases}, + } + ``` +- suggested fix: Add the wildcard-resource variants to every carve-out table, a persist-failure case for recovery codes, a deactivation case asserting `DeleteUserSessions` and a rejected session, a reset-on-deactivated-account refusal, an invited-user forgot-password case, an MFA re-enrol-while-enabled refusal, and a transient-then-success signer case; convert the legacy-fallback test into an error assertion. +- verdict: CONFIRMED — the cited tables hold only the exact (verb, purchases) pairs (group_ceiling_permissions_test.go:24-28, self_escalation_carveout_test.go:35-39) and a grep of internal/auth/*_test.go for an execute permission with ResourceAll or "*" returns nothing; TestLogin_WithMFA_RecoveryCode_ConsumedOnce stubs UpdateUser to succeed and asserts only the slice (service_mfa_test.go:438-466); the deactivation case registers no DeleteUserSessions expectation (service_user_test.go:646-665); no test references PasswordSetupExpiry, none of the three MFASetup calls sets MFAEnabled=true beforehand (service_mfa_test.go:221, :236, :486), the fake KMS client's GetPublicKey never fails (aws_signer_test.go:29-34), and TestResolveBastionProvider_LegacyFallback asserts NoError on the fallback (resolver_test.go:366-374). +- issue: (pending cross-reference) + +### A05-007 `FinalizeInFlightRevocations` has no tests at all +- category: test-gap +- severity: medium +- location: internal/purchase/finalize_revocations.go:45 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `git grep FinalizeInFlightRevocations` matches only the definition — no test file references it. This is the sweep that closes the window where an Azure Return API call succeeded but the DB write failed, i.e. it is the sole mechanism that makes the revocation audit record eventually consistent for money already returned. Nothing pins the retry count, the shared `now` timestamp, the per-row error isolation, or that a failing row is counted in `Errored` rather than silently dropped, so any of those can regress unnoticed. +- evidence: + ```go + func (m *Manager) FinalizeInFlightRevocations(ctx context.Context) (*FinalizeResult, error) { + rows, err := m.config.GetPurchaseHistoryInFlight(ctx) + if err != nil { + return nil, err + } + ``` +- suggested fix: add a table test over a mocked store covering: all rows succeed on the first attempt, a row that succeeds on retry 2, and a row that never succeeds (asserting `Errored` increments and the sweep continues to the next row). +- verdict: CONFIRMED — the substance holds, with one correction to the evidence: `git grep -n FinalizeInFlightRevocations -- '*.go'` does match test files (internal/server/handler_test.go:142 and :152), but those stub the whole method through testutil.MockPurchaseManager (internal/testutil/mocks.go:109-111) and exercise the HTTP handler at internal/server/handler.go:353, not one line of finalize_revocations.go. No test in internal/purchase references it, so the retry count, the shared `now` at :52, per-row isolation and the Errored counter are all unpinned. +- issue: (pending cross-reference) + +### A07-033 No test exercises a title-cased engine or the CE-offering-ID purchase path +- category: test-gap +- severity: medium +- location: providers/aws/services/elasticache/client_test.go:446 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: every ElastiCache test constructs `CacheDetails{Engine: "redis"}` in the already-lowercase form, so A07-006 (the raw title-cased CE value reaching the offering filter) cannot fail any test. Likewise, `savingsplans/client_test.go:1530` is the single test touching a non-empty `OfferingID`, and it asserts only that the offerings API is skipped — nothing asserts that a plan-type, term or payment-option mismatch is still rejected on that path (A07-001). Both bugs live in the exact gap the suite leaves open. +- evidence: + ```go + Details: &common.CacheDetails{Engine: "redis", NodeType: "cache.m6g.large"}, + ``` +- suggested fix: add an ElastiCache offering-lookup case with `Engine: "Redis"` and an empty engine, and a Savings Plans case where a rec carrying an `OfferingID` has a plan type that does not match the scoped client; both should fail against the current code. +- verdict: CONFIRMED — every `CacheDetails` literal in providers/aws/services/elasticache/client_test.go uses lowercase `"redis"` (:251, :288, :341, :446, :560, :616, :647, :680, :714, :737, :762) with no title-cased or empty-engine case, and the only Savings Plans test with a non-empty `OfferingID` is `TestEC2InstanceSP_CEProvidedOfferingIDUsedDirectly` (savingsplans/client_test.go:1512-1537), which asserts only that the offerings API is skipped; the plan-type rejection test at :364-385 uses a rec with no `OfferingID`, so it never exercises the short-circuit path. +- issue: (pending cross-reference) + +### A09-031 No test covers the two money-path guards that fail open in pkg/exchange +- category: test-gap +- severity: medium +- location: pkg/exchange/exchange_test.go:1 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: across all six test files in `pkg/exchange` there is no case that sets `PaymentDueRaw` to the empty string on a quote destined for `checkInitialQuote`/`checkReQuote` (A09-004), no case with a `CurrencyCode` other than `"USD"` reaching the cap comparison (A09-005), and no reference at all to `assertAccount` or `ExpectedAccount` (A09-006). The suite passes with all three defects present, so the guards' teeth are unverified on the one path in the package that spends money irreversibly. +- evidence: + ```text + $ /usr/bin/grep -rn "assertAccount\|ExpectedAccount" pkg/exchange/*_test.go + (no output) + $ /usr/bin/grep -rn "CurrencyCode" pkg/exchange/*_test.go | grep -v '"USD"' + (no output) + ``` +- suggested fix: add three table cases to `exchange_test.go` — an empty `PaymentDue` quote that must be refused, a `"EUR"` quote that must be refused, and an `ExpectedAccount` mismatch that must abort before `AcceptReservedInstancesExchangeQuote` — each asserting the accept call was never made. +- verdict: CONFIRMED — I opened all six test files rather than trusting the greps. exchange_test.go holds only `TestParseDecimalRat`, `TestPaymentDueUSDStr_InJSON` and `TestSpendCapComparison` (which exercises `big.Rat.Cmp` directly, never `checkInitialQuote`). fail_loud_test.go's `seqQuoteOut` always sets a non-empty `PaymentDue` (fail_loud_test.go:47-52) and never sets `CurrencyCode` or `ExpectedAccount`; multi_target_test.go:37 is the same shape. No test constructs a quote with an empty `PaymentDue` reaching the cap check, none uses a non-USD currency there, and `assertAccount`/`ExpectedAccount` appear in no test file. Minor evidence correction: the second grep in the finding does produce output (five EUR/empty `OfferingOption.CurrencyCode` lines in reshape_crossfamily_test.go), but those exercise `passesDollarUnitsCheck`, not the cap comparison, so the substantive claim stands. +- issue: (pending cross-reference) + +### A09-032 httpclient tests assert only the literal-IP path, so the DNS bypass stays green +- category: test-gap +- severity: medium +- location: pkg/httpclient/httpclient_test.go:27 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `TestNew_BlocksIMDS` exercises two hardcoded URLs, both containing the literal address the map holds. It therefore passes for any implementation that string-matches the pre-resolution host, including the current one, and would keep passing if the ECS credential endpoint or a resolving hostname were added as attack vectors. `TestNew_AllowsRegularEndpoints` confirms the happy path but adds no negative coverage. Nothing in the suite would fail if the blocklist were reduced to a single entry. +- evidence: + ```go + tests := []struct { + name string + url string + }{ + {name: "ipv4 link-local IMDS", url: "http://169.254.169.254/latest/meta-data/"}, + {name: "ipv6 AWS IMDS", url: "http://[fd00:ec2::254]/latest/meta-data/"}, + } + ``` +- suggested fix: add cases for `169.254.170.2` (ECS credentials) and for a hostname stubbed to resolve to `169.254.169.254` via an injected resolver, so the blocklist's actual reach is what the test measures. +- verdict: CONFIRMED — I read the whole file. The suite is three tests: `TestNew_NotDefaultClient` (identity assertions only), `TestNew_BlocksIMDS` with the two literal-IP URLs quoted, and `TestNew_AllowsRegularEndpoints` against an `httptest` server (pkg/httpclient/httpclient_test.go:11-73). Nothing exercises a resolving hostname or the ECS/EKS addresses, so the A09-001 DNS bypass and the A09-002 gaps both stay green. One correction: the closing sentence overstates — reducing the blocklist to a single entry WOULD fail whichever of the two subtests lost its address; what the suite cannot detect is any implementation that string-matches the pre-resolution host. +- issue: (pending cross-reference) + +### A10-025 No test asserts that money fields survive a count reduction on the extended-support path +- category: test-gap +- severity: medium +- location: cmd/multi_service_engine_versions_test.go:236 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: All six `adjustRecommendationForExcludedVersions` tests assert only `result.Count`, and none of the fixtures sets `EstimatedSavings` or `CommitmentCost` at all, so A10-001 is invisible to the suite and would remain invisible after a fix regressed. The sibling flags have the coverage this path lacks: `cmd/helpers_instance_limit_rescale_test.go` and `cmd/helpers_count_override_rescale_test.go` exist precisely to pin the scaling behaviour for `--max-instances` and `--override-count`. +- evidence: + ```go + result := adjustRecommendationForExcludedVersions(recommendation, instanceVersions, versionInfo) + assert.Equal(t, 8, result.Count, "Should exclude 2 instances (5.6 and 5.7 both in extended support)") + ``` +- suggested fix: Add a rescale test mirroring `helpers_instance_limit_rescale_test.go`: a rec with non-zero EstimatedSavings/CommitmentCost, two of ten instances excluded, asserting both fields land at 0.8x. +- verdict: CONFIRMED — opened every test file that drives this function: all assertions check only result.Count (cmd/multi_service_engine_versions_test.go:180, :185, :238, :255, :282 and cmd/multi_service_coverage_test.go:416) and no fixture sets EstimatedSavings or CommitmentCost, while cmd/helpers_instance_limit_rescale_test.go and cmd/helpers_count_override_rescale_test.go do exist for the sibling paths. +- issue: (pending cross-reference) + +### A12-037 The amortized Monthly Cost path in History has no test coverage +- category: test-gap +- severity: medium +- location: frontend/src/history.ts:1140 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: All eleven `history*.test.ts` files mock `getAmortizeUpfront` to return `false` and none flips it. No fixture creates `history-controls` or `purchases-approval-queue-section`, so both `mountAmortizeCheckbox` calls, `syncAmortizeCheckbox`, the `subscribeAmortizeUpfront` re-render path, the `amortizedMonthly` cell in both tables and the "(amortized)" header are never executed. A regression that added the upfront twice, or dropped it, would ship green. +- evidence: + ```ts + const displayMonthly = (rawMonthly != null && amortize) + ? amortizedMonthly(rawMonthly, p.upfront_cost, p.term) + : rawMonthly; + ``` +- suggested fix: Add a fixture with those containers and a test that flips `getAmortizeUpfront` to `true`, asserting the cell equals `monthly + upfront/(term*12)` and the header reads "Monthly Cost (amortized)". +- verdict: CONFIRMED — Every history suite pins getAmortizeUpfront to false and none flips it (frontend/src/__tests__/history.test.ts:52 and the nine siblings), and no fixture creates `history-controls` or `purchases-approval-queue-section` (only frontend/src/index.html:180, :242), so the amortized cell at frontend/src/history.ts:1138-1141 and both mount calls are never executed. +- issue: (pending cross-reference) + +### A12-043 No test covers `submitModalExecute` or `renderExchangeHistory` +- category: test-gap +- severity: medium +- location: frontend/src/__tests__/riexchange.test.ts:150 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `executeExchange` is mocked in three test files but never asserted on; the suite asserts only the quote request shape. That is why the missing `region` (A12-003) and the `max_payment_due_usd` currency assumption (A12-039) ship green. `renderExchangeHistory`, `canApproveRIExchangeRow` and `handleRIExchangeApproveClick` have no test at all, so A12-042 and the hardcoded `$` are equally unguarded. +- evidence: + ```ts + executeExchange: jest.fn(), // mocked; no test reads mockedApi.executeExchange.mock.calls + ``` +- suggested fix: Add a test that drives quote to execute and asserts the full posted body including `region`, plus one rendering a pending history record and asserting the cells and the Approve gating. +- verdict: CONFIRMED — executeExchange is mocked in three suites (frontend/src/__tests__/riexchange.test.ts:12, riexchange-column-filters.test.ts:27, riexchange-active-ri-filters.test.ts:20) with no assertion on its calls anywhere, and renderExchangeHistory, canApproveRIExchangeRow and handleRIExchangeApproveClick appear in no test file. +- issue: (pending cross-reference) + +### A14-015 ci.yml's "Test Docker image" step asserts nothing and rebuilds the image it was given +- category: test-gap +- severity: medium +- location: .github/workflows/ci.yml:522 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Both container runs end in `|| true`, so a binary that panics on `--version`, is built for the wrong architecture, or is missing from `/app/cudly` passes the step. The step also runs a second `docker build`, discarding the `load: true` image the preceding build-push-action step produced specifically so downstream steps would inspect the artifact this job built (comment at 512-514); the two builds can differ if the cache resolves differently. +- evidence: + ```yaml + - name: Test Docker image + run: | + docker build -t cudly:test . + docker run --rm cudly:test /app/cudly --version || true + docker run --rm cudly:test /app/cudly --help || true + ``` +- suggested fix: Drop the redundant `docker build`, run `cudly:${{ github.sha }}`, and remove the `|| true` so a broken binary fails the job. +- verdict: CONFIRMED — ci.yml:522-526 rebuilds `cudly:test` rather than exercising the `load: true` image `cudly:${{ github.sha }}` produced at ci.yml:509-517, and both `docker run` lines end in `|| true`, so no assertion can fail the step. +- issue: (pending cross-reference) + +### A14-031 The Azure sanity workflow compares the subscription against itself +- category: test-gap +- severity: medium +- location: .github/workflows/azure_sanity.yml:66 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `--subscription-id` and `--expected-subscription` are both `${{ secrets.AZURE_SUBSCRIPTION_ID }}`. `runAccountShowCheck` fetches that subscription by ID and reads `sub.SubscriptionID` back from the response, so `validateAccountExpectations` (ci_cd_sanity_tests/pkg/sanity/azure/azure.go:84) compares a value against the value used to fetch it. The `unexpected subscription` branch is unreachable by construction, and the `azure:account:expected_checks` PASS in the uploaded report asserts nothing about the subscription axis. The tenant comparison is meaningful; the subscription one is not. +- evidence: + ```yaml + ./azure-sanity \ + --subscription-id "${AZURE_SUBSCRIPTION_ID}" \ + --expected-subscription "${AZURE_SUBSCRIPTION_ID}" \ + --expected-tenant "${AZURE_TENANT_ID}" \ + --out "${REPORT_PATH}" + ``` +- suggested fix: Drop `--expected-subscription` from this invocation, or source it from a separate repository variable so the two values can disagree. +- verdict: CONFIRMED — runAccountShowCheck fetches via `subClient.Get(ctx, subscriptionID)` and reads `info.ID = *sub.SubscriptionID` back from that response (ci_cd_sanity_tests/pkg/sanity/azure/azure.go:263-274), so validateAccountExpectations' `a.ID != opts.ExpectedSubID` test (azure.go:81) compares the fetch key against itself; a wrong ID errors at Get and skips the check entirely (azure.go:330). +- issue: (pending cross-reference) + +### A15-004 The vacuous-assertion guard silently skips packages that declare no mock of their own +- category: test-gap +- severity: medium +- location: internal/mocks/vacuous_assertion_guard_test.go:137 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `TestNoUnfailableMockAssertions` walks the repo, then per directory does `mocks := collectMocks(files); if len(mocks) == 0 { continue }`. A package that uses a mock imported from elsewhere but declares none locally is dropped before its assertion sites are ever collected, and it is not added to the `skipped` slice the test errors on. Four packages with 16 assertion sites are therefore invisible: `internal/server` (6), `providers/azure/services/compute` (5), `providers/azure/services/managedredis` (3), `providers/azure/services/synapse` (2). The guard reports `checked 162 mock assertion site(s)` and passes, with zero skips reported. None of the 16 is vacuous today, because all pass matchers whose count equals the real `m.Called` arity, so this is a hole in the guard rather than a live unfailable assertion. It contradicts the guard's own stated contract, which says a skip that reports nothing is the defect the guard exists to catch. +- evidence: + ```go + mocks := collectMocks(files) + if len(mocks) == 0 { + continue + } + ``` +- suggested fix: Collect assertion sites before the `len(mocks) == 0` test and resolve them through the existing repo-wide `global` index, routing anything still unresolved into `skipped` rather than dropping the directory. +- verdict: CONFIRMED — I re-derived the drop set by replaying the guard's own `collectMocks` and assertion-site walk over the pinned tree, and the hole at internal/mocks/vacuous_assertion_guard_test.go:137 is real: the same four packages are dropped with no entry in `skipped`, and the guard run logs `checked 162 mock assertion site(s)` and PASSes. The count is **15**, not 16 — `providers/azure/services/compute/client_test.go` has 4 sites (924, 948, 971, 1499), not 5; internal/server has 6, managedredis 3, synapse 2. Two further qualifiers the finding omits: the 4 `internal/server/handler_test.go` sites are on `mocks.MockConfigStore`, which shadows both helpers (internal/mocks/stores.go:1564 records before `isExpected`), so they would classify as `verdictIgnore` even if the directory were reached; and every one of the 15 passes a matcher count equal to the real `m.Called` arity (`MockHTTPClient.Do` -> `m.Called(req)`, providers/azure/mocks/azure_mocks.go:130), confirming none is vacuous today. +- issue: (pending cross-reference) + +### A16-002 4-eyes and per-account approver matching are never tested with a case-differing email +- category: test-gap +- severity: medium +- location: internal/purchase/approvals.go:290 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the self-approval comparison is `strings.EqualFold`, but every email literal in + internal/purchase/approvals_test.go and internal/purchase/coverage_extra_test.go is all-lowercase + and byte-identical on both sides (creator `creator@example.com` vs actor `creator@example.com` at + approvals_test.go:352/358, and `owner@example.com` on both sides at coverage_extra_test.go:358/382). + Replace `EqualFold` with `==` and the whole suite stays green, while a creator whose session email + is `Creator@example.com` self-approves their own purchase under 4-eyes. The same untested + case-folding sits in `matchActorAgainstApprovers`, which authorizes the SQS approve path against + per-account `contact_email`. +- evidence: + ```go + if strings.EqualFold(creatorEmail, actorEmail) { + logging.Warnf("purchase[%s]: 4-eyes mode on; creator %s attempted self-approval via actor %q, denied", + executionID, *execution.CreatedByUserID, maskActor(actorEmail)) + return fmt.Errorf("approval declined: 4-eyes mode requires a different approver than the requester") + } + ``` +- suggested fix: change one existing self-approve subtest to stub the creator as + `Creator@Example.COM` while the actor stays `creator@example.com`, and mirror it in the + approver-matching test. Both must still deny. +- verdict: CONFIRMED — a repo-wide scan for a mixed-case email literal (`git grep -nE '"[A-Za-z0-9._%+-]*[A-Z][A-Za-z0-9._%+-]*@'`) returns exactly two hits, neither in internal/purchase: handler_purchases_test.go:5098 covers `resolveExecutedNotificationRecipients` dedup and handler_registrations_recipients_test.go:23 covers recipient dedup, so no fixture reaches approvals.go:290 or messages.go:263-266 with differing case; messages.go:263-266 does lowercase both sides as the finding states. +- issue: (pending cross-reference) + +### A01-022 4-eyes coverage for revocation exists only as t.Skip placeholders +- category: test-gap +- severity: low +- location: internal/api/handler_purchases_test.go:5304 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `TestRevoke_4EyesMode_TODO` and `TestRevokePurchase_FourEyesApproval` (handler_purchases_revoke_test.go:1311) are permanently skipped placeholders "until #1005 lands". #1005 (4-eyes) has landed for approval, and neither revoke path (`revokeViaSession`, `revokeAzurePurchase`) applies any dual-control check, so the skipped tests document an unenforced expectation and inflate the apparent coverage of the revoke money path. +- evidence: + ```go + func TestRevoke_4EyesMode_TODO(t *testing.T) { + t.Skip("placeholder until #1005 lands: verify that 4-eyes-mode approval flows interact correctly with the revocation window") + } + ``` +- suggested fix: Decide whether revocation is subject to dual control; implement and test it, or delete the placeholders. +- verdict: CONFIRMED — both placeholders are bare t.Skip (internal/api/handler_purchases_test.go:5304-5306, handler_purchases_revoke_test.go:1311-1313); #1005 is implemented for approvals via enforceFourEyesPolicy (internal/purchase/approvals.go:216-230, called from ApproveAndExecute:323), and neither revokeViaSession (handler_purchases.go:1446-1478) nor the revokePurchase/authorizeSessionRevoke* chain (handler_purchases_revoke.go:139-330) references requireDifferentApprover or enforceFourEyesPolicy. +- issue: (pending cross-reference) + +### A04-020 Integration test harness adds a second unique index on execution_id under a false claim about the schema +- category: test-gap +- severity: low +- location: internal/config/store_postgres_db_test.go:53-59 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the comment states "the migration schema does not create a UNIQUE constraint on execution_id"; migration 000008 has created `unique_execution_id` since day one and every `ON CONFLICT (execution_id)` in production depends on it. The harness therefore tests against a schema with two unique indexes and would keep passing if 000008 were ever dropped, which is the only regression the extra DDL could have caught. It is a stale workaround, not coverage. +- evidence: + ```go + // The store code uses ON CONFLICT (execution_id) but the migration schema + // does not create a UNIQUE constraint on execution_id. Add it for tests. + _, err := container.DB.Exec(ctx, + "CREATE UNIQUE INDEX IF NOT EXISTS idx_purchase_executions_execution_id_unique ON purchase_executions(execution_id)") + ``` +- suggested fix: delete the DDL and the comment so the integration suite runs against exactly the migrated schema. +- verdict: CONFIRMED — internal/database/postgres/migrations/000008_add_execution_id_unique.up.sql:3 adds `CONSTRAINT unique_execution_id UNIQUE (execution_id)`, so the comment at internal/config/store_postgres_db_test.go:53-54 is false; the harness has already applied the full chain at store_postgres_db_test.go:49-51 before layering the redundant index at 55-56, which can only mask the loss of 000008. +- issue: (pending cross-reference) + +### A12-061 Both History tables stamp the same `data-execution-id`, and the fixtures invert the shipped DOM order +- category: test-gap +- severity: low +- location: frontend/src/history.ts:1792 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `renderHistoryList` (line 1136) and `renderApprovalQueue` (line 1792) both stamp `data-execution-id`, so a pending purchase produces two rows with the same id and `applyExecutionDeepLink`'s unscoped `document.querySelector` returns the first in DOM order. In `index.html` the queue precedes the list, so production highlights the queue row; every test fixture appends `history-list` before `purchases-approval-queue`, so the tests exercise the opposite element. The deep-link tests use a hand-built single table and never see the duplicate. +- evidence: + ```ts + const execIdAttr = p.purchase_id ? ` data-execution-id="${escapeHtmlAttr(p.purchase_id)}"` : ''; + ``` +- suggested fix: Scope the deep-link query to `#history-list` (or give queue rows a distinct attribute) and reorder the fixtures to match `index.html`. +- verdict: CONFIRMED — Both renderers stamp data-execution-id (frontend/src/history.ts:1136, :1792) and applyExecutionDeepLink resolves it with an unscoped document.querySelector that takes the first in DOM order (frontend/src/history.ts:122-124); index.html puts the queue first (frontend/src/index.html:180 before :253) while the fixtures append history-list first (frontend/src/__tests__/history-approval-queue.test.ts:105-107). +- issue: (pending cross-reference) + +### A14-028 A stale comment keeps the cross-account generator test from asserting its exit code +- category: test-gap +- severity: low +- location: scripts/generate_federation_iac_test.go:428 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The comment says `iacData` is missing `SourceAccountID / CUDlyAPIURL / ContactEmail / OIDCIssuerHost` and that "no tfvars path of this script renders on main today", and on that basis `TestGenerator_CrossAccountDoesNotRequireSubjectClaim` asserts only on stderr. Three of those four fields are now present (SourceAccountID at generate-federation-iac.go:119, CUDlyAPIURL and ContactEmail at 141-142) and `TestGenerateFederationIaC_TfvarsCombinations` exercises tfvars rendering, so the stated justification no longer holds. The test therefore passes whatever exit code the aws→aws path returns, including a regression that breaks it outright. +- evidence: + ```go + // It asserts only on stderr, not on the exit code, because that invocation + // currently fails for an unrelated pre-existing reason: iacData is missing the + // SourceAccountID / CUDlyAPIURL / ContactEmail / OIDCIssuerHost fields the + // tfvars and deploy-script templates reference, so no tfvars path of this + // script renders on main today. + ``` +- suggested fix: Fix the remaining `OIDCIssuerHost` gap (A14-026), then assert `res.exitCode == 0` and delete the stale rationale. +- verdict: CONFIRMED — TestGenerator_CrossAccountDoesNotRequireSubjectClaim (scripts/generate_federation_iac_test.go:434-446) asserts only `strings.Contains(res.stderr, "--oidc-subject-claim")`, while three of the four fields its rationale names now exist (SourceAccountID at generate-federation-iac.go:119, CUDlyAPIURL and ContactEmail at 141-142) and generate-federation-iac_test.go:40 renders tfvars successfully, so the stated justification is stale. +- issue: (pending cross-reference) + +### A15-006 A singleflight test orchestrates a ten-goroutine race with 450ms of sleeps +- category: test-gap +- severity: low +- location: internal/api/ri_utilization_cache_test.go:259 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The test needs all ten background-refresh goroutines to reach `sf.Do` before the in-flight fetch is released, and sleeps 250ms to arrange it, then 200ms more to let `storePayload` finish before reading `calls.Load()`. Under `-race` on a loaded runner a straggler can still arrive late and open a second batch, making `calls` 2 and failing the test for scheduling reasons rather than a real defect; conversely a genuine singleflight regression that collapses slowly can pass. The in-file comment already documents the flakiness it is working around. +- evidence: + ```go + time.Sleep(250 * time.Millisecond) + + // Unblock the single in-flight fetch and wait for the collapsed + // batch's storePayload to complete so calls.Load() reflects the + // final state. + close(release) + time.Sleep(200 * time.Millisecond) + ``` +- suggested fix: Have the fetch function signal arrival on a buffered channel and read exactly ten arrivals before `close(release)`, then wait on a completion signal from `storePayload` instead of the trailing sleep. +- verdict: CONFIRMED — the two sleeps are exactly as quoted at internal/api/ri_utilization_cache_test.go:259 and :265, and the in-file comment at :251-258 documents the straggler race it is papering over, so the false-failure direction is real and self-admitted. One half of the claim is not supported: a genuine singleflight regression makes `calls` larger, not smaller (removing the collapse entirely would drive all ten fetchers through and fail at the `n != 1` check on line 269), so the test cannot pass through a slow-collapse regression the way the finding asserts. +- issue: (pending cross-reference) + +### A16-011 The parallelism bound test asserts only the ceiling, so full serialisation passes +- category: test-gap +- severity: low +- location: internal/scheduler/scheduler_test.go:1145 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `assert.LessOrEqual(peak, 3)` is satisfied by `peak == 1`. If + `fanOutPerAccount` regressed to a serial loop, or `CUDLY_MAX_ACCOUNT_PARALLELISM` stopped being + parsed and defaulted to 1, this test still passes while a 20-account collection sweep takes 20x + longer — which on Lambda means timing out the whole cycle. The test also carries a + `time.Sleep(5 * time.Millisecond)` inside the worker purely to make overlap observable, which is + what makes the missing lower bound easy to add. +- evidence: + ```go + out, outcome := fanOutPerAccount(context.Background(), "Test", accounts, fn) + assert.Len(t, out, len(accounts), "all accounts contribute one record") + assert.LessOrEqual(t, peak.Load(), int32(3), "peak in-flight must not exceed CUDLY_MAX_ACCOUNT_PARALLELISM") + assert.Equal(t, len(accounts), outcome.SucceededCount) + assert.Zero(t, outcome.FailedCount) + ``` +- suggested fix: add `assert.Greater(t, peak.Load(), int32(1), "the fan-out must actually run + accounts concurrently")`. With 20 accounts, a limit of 3 and a 5ms body, a peak of 1 means the + limiter serialised everything. +- verdict: CONFIRMED — scheduler_test.go:1143-1147 asserts only `LessOrEqual(peak, 3)`, `Len(out, 20)`, `SucceededCount` and `Zero(FailedCount)`, every one of which a fully serial `fanOutPerAccount` satisfies. The 5ms body at scheduler_test.go:1137 and the peak tracker at 1123-1131 are already in place, so the missing lower bound is one line. +- issue: (pending cross-reference) + +### A16-012 A tautological assertion presented as "the property in its own terms" +- category: test-gap +- severity: low +- location: internal/auth/account_scope_fail_closed_test.go:74 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `require.Error(t, err)` on line 69 has already established that `err != nil`, so + the `&& err == nil` conjunct is false and the whole `assert.False` can never fire, whatever + `IsUnrestrictedAccess(got)` returns. The comment above it claims this restates the security + property independently of the preceding assertions, which is the risk: a reader auditing this file + sees the invariant explicitly asserted and stops looking. The real protection comes from + `assert.Nil(t, got)` on line 70. The same construction is repeated at line 192. +- evidence: + ```go + require.Error(t, err, "an unestablishable scope must be refused, not treated as unrestricted") + assert.Nil(t, got) + assert.Contains(t, err.Error(), "could not be established") + // The property in its own terms: whatever comes back must not be + // readable as "all accounts". + assert.False(t, IsUnrestrictedAccess(got) && err == nil, + "a failed resolution must never yield an unrestricted scope") + ``` +- suggested fix: drop the `&& err == nil` conjunct so the assertion reads + `assert.False(t, IsUnrestrictedAccess(got))`, which is the claim the comment makes and can actually + fail. Same edit at line 192. +- verdict: CONFIRMED — `require.Error` at account_scope_fail_closed_test.go:69 aborts the subtest when `err == nil`, so by line 74 `err != nil` holds unconditionally, the `&& err == nil` conjunct is constant false and `assert.False` cannot fire for any value of `IsUnrestrictedAccess(got)`. The construction repeats verbatim at line 192, where the preceding `require.Error` sits at line 187, and in both places `assert.Nil(t, got)` (lines 70 and 189) is the assertion carrying the property. +- issue: (pending cross-reference) + +### A16-014 permissionCoveredBy never returns true in the whole suite +- category: test-gap +- severity: low +- location: internal/auth/group_ceiling.go:376 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the `return true` is uncovered, so no test in the group-ceiling suite exercises + "this write KEEPS a permission the group already had at constraints no narrower than the request". + The false direction is well covered — the over-ceiling denial tests would fail if this returned + true unconditionally — but the true direction is not. A change that makes it always return false + turns every no-op resubmission of an existing group permission into a ceiling violation, locking a + group admin out of editing any other field on their own group, and the suite stays green. +- evidence: + ```go + func permissionCoveredBy(set []Permission, req Permission) bool { + for _, p := range set { + if p.Action != req.Action || p.Resource != req.Resource { + continue + } + if constraintsCover(p.Constraints, req.Constraints) { + return true + } + ``` +- suggested fix: add a case to internal/auth/group_ceiling_permissions_test.go where an actor whose + own permissions do NOT reach the ceiling resubmits a group's existing permission unchanged, and + assert the update succeeds. +- verdict: CONFIRMED — group_ceiling.go:376.4 is 0 in the merged profile, so the carry-through direction is genuinely never asserted. One scope correction, which also invalidates the suggested fix as written: `permissionCoveredBy` has a single call site, group_ceiling.go:86, inside the `adminCarvedOuts` branch, so a permission that is not carved out never reaches it and a resubmission of an ordinary permission goes through `grantCeilingAllows` instead. The consequence of an always-false return is therefore narrower than stated — not "every no-op resubmission", but any edit to a group that holds one of the separation-of-duties reserved pairs, which can then never be saved again. group_ceiling_permissions_test.go:220-235 documents this same distinction. +- issue: (pending cross-reference) + +### Category: duplication + +13 findings: 4 medium, 9 low. + +### A05-008 Third copy of the approval-token check, and it is the one missing the TTL guard +- category: duplication +- severity: medium +- location: internal/purchase/messages.go:240 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The same non-empty + constant-time hash comparison exists three times: `validateApprovalToken` (approvals.go:115), `loadCancelableExecution` (approvals.go:494) and `loadAsyncExecutionForApproval` here. The first two then enforce `ApprovalTokenExpiresAt`; this one does not. An SQS approve message carrying an expired token therefore clears `verifyAsyncApprovalActor` and reaches the approver-list lookup, and is only rejected later by `ApproveExecution`'s own TTL check. The security property currently holds by accident of call ordering — any future caller of `loadAsyncExecutionForApproval` that does not re-check inherits an expiry-blind token check. +- evidence: + ```go + if execution.ApprovalToken == "" || msg.Token == "" { + return nil, fmt.Errorf("invalid approval token") + } + storedHash := sha256.Sum256([]byte(execution.ApprovalToken)) + userHash := sha256.Sum256([]byte(msg.Token)) + if subtle.ConstantTimeCompare(storedHash[:], userHash[:]) != 1 { + return nil, fmt.Errorf("invalid approval token") + } + return execution, nil + ``` +- suggested fix: have all three sites call the existing `validateApprovalToken(execution, token)` helper, which already covers the empty check, the constant-time compare and the TTL. +- verdict: CONFIRMED — validateApprovalToken (approvals.go:109-125) and loadCancelableExecution (approvals.go:494-507) both end with the `ApprovalTokenExpiresAt != nil && time.Now().After(...)` guard, while loadAsyncExecutionForApproval (messages.go:229-244) returns the execution immediately after the ConstantTimeCompare with no TTL check; the only thing that rejects the expired token afterwards is ApproveExecution's own validateApprovalToken call at approvals.go:40, one frame later in the same request. +- issue: (pending cross-reference) + +### A08-021 isPermissionError re-implements the package's own IsPermissionError, less correctly +- category: duplication +- severity: medium +- location: providers/gcp/recommendations.go:402 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the same package already exports `IsPermissionError` (providers/gcp/permission_errors.go:29), which uses `errors.As` and also recognizes the gRPC `codes.PermissionDenied` surface. The local copy hand-rolls unwrapping via `isGoogleAPIError`, which only follows single-`Unwrap() error` chains and therefore misses an error joined with `errors.Join` or a multi-`%w` wrap, recognizes no gRPC status at all, and adds the loose "403"+"permission" substring fallback from A08-020. The `errorAs` shim (recommendations.go:417) exists only to support this duplicate. +- evidence: + ```go + var gapiErr *googleapi.Error + if ok := errorAs(err, &gapiErr); ok && gapiErr.Code == 403 { + return true + } + msg := err.Error() + return strings.Contains(msg, "403") && strings.Contains(msg, "permission") + ``` +- suggested fix: call the exported `IsPermissionError` and delete `isPermissionError`, `errorAs` and `isGoogleAPIError`. +- verdict: CONFIRMED — providers/gcp/permission_errors.go:29-42 already exports `IsPermissionError` using `errors.As` plus the gRPC `codes.PermissionDenied` surface, while the local copy (recommendations.go:402-412) unwraps through `isGoogleAPIError` (recommendations.go:430-441), which recurses only over single `Unwrap() error` chains and recognises no gRPC status, and `errorAs` (recommendations.go:417) has no other caller. +- issue: (pending cross-reference) + +### A08b-039 The three GCP service clients are near-verbatim copies of one another +- category: duplication +- severity: medium +- location: providers/gcp/services/memorystore/client.go:454 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (twins at `cloudsql/client.go:468` and `cloudstorage/client.go:460`) +- failure scenario: `extractGCPResourceType`, `extractGCPSavings`, `termYearsFromLabel`, `skuMatchesTier`/`skuMatchesStorageClass`, `extractPriceFromSKU`, `calculateSavingsPercentage`, `fill*Pricing`, `getOrCreateBillingService`, `realBillingService` and `convertGCPRecommendation` are byte-identical (modulo the receiver type and one identifier) across the three packages. Every finding in this report that names one of them applies three times, and a fix landed in one package leaves the other two wrong — which is how the payment-option default in A08b-013 survives in all three while the Azure siblings were fixed. +- evidence: + ```go + func extractGCPSavings(rec *recommenderpb.Recommendation) float64 { + if rec.PrimaryImpact == nil { + return 0 + } + costProj := rec.PrimaryImpact.GetCostProjection() + if costProj == nil || costProj.Cost == nil { + return 0 + } + ``` +- suggested fix: lift the shared helpers into a `providers/gcp/internal/recommendations` package, mirroring what `providers/azure/internal/{pricing,recommendations}` already does for Azure. +- verdict: CONFIRMED — comparing cloudsql:442-529, cloudstorage:434-520 and memorystore:428-500 line by line, `extractGCPResourceType`, `extractGCPSavings`, `termYearsFromLabel`, the savings-percentage helper, the SKU/tier matcher, `extractPriceFromSKU`, `getOrCreateBillingService` and `realBillingService.ListSKUs` differ only by receiver and identifier; A08b-013 surviving identically in all three is the drift the finding predicts. +- issue: (pending cross-reference) + +### A12-026 Two different permission predicates gate the same cancel endpoint +- category: duplication +- severity: medium +- location: frontend/src/dashboard.ts:410 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Both the Home widget's Cancel button and the Plans row's Disable button call `DELETE /api/purchases/planned/{id}`, but they compute eligibility differently. `canCancelUpcomingPurchase` treats `cancel-any` as full scope; `plans.ts:717` folds it into a conjunction with `canManageScheduledPurchase`, whose full-scope set is only `{admin, update-any}`. A user holding `cancel-any:purchases` but not `update-any:purchases`, looking at a scheduler-created row with a null `created_by_user_id`, sees a working Cancel button on Home and no Disable button on Plans for the same execution. +- evidence: + ```ts + if (canAccess('admin', '*') || canAccess('cancel-any', 'purchases') || canAccess('update-any', 'purchases')) return true; + // plans.ts:717-725 + const canManagePurchase = canManageScheduledPurchase(purchase); // admin | update-any | creator + const canDisablePlan = canManagePurchase && (canAccess('delete','purchases') || canAccess('cancel-any','purchases') || canAccess('cancel-own','purchases')); + ``` +- suggested fix: Extract one shared `canCancelScheduledExecution(row)` helper and call it from both surfaces. +- verdict: CONFIRMED — canCancelUpcomingPurchase grants on cancel-any alone (frontend/src/dashboard.ts:412-421) while the Plans row ANDs canManageScheduledPurchase, whose full-scope set is only admin/update-any (frontend/src/plans.ts:673-679, :717-725), and both buttons hit DELETE /api/purchases/planned/{id} (frontend/src/dashboard.ts:648-669, frontend/src/plans.ts:812-824). +- issue: (pending cross-reference) + +### A01-020 Azure exchange execute resolves the same CloudAccount three times per request +- category: duplication +- severity: low +- location: internal/api/handler_ri_exchange.go:1011 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: For a scoped session, `authorizeAzureExchangeExecution` calls `GetCloudAccountByExternalID(ctx, "azure", sub)` in `requireAzureSubscriptionScope` (805), again in `buildAzureExchangeClient` (214), and again in `resolveAzureExchangeAccountID` (1029). Three identical round-trips on the money path, and three places that must agree on the "no account registered" semantics (`unattributedAccountConstraint` vs 404 vs errNotFound). +- evidence: + ```go + if scopeErr := h.requireAzureSubscriptionScope(ctx, session, body.SubscriptionID); scopeErr != nil { + // ... + client, err := h.buildAzureExchangeClient(ctx, body.SubscriptionID) + // ... + accountID, err := h.resolveAzureExchangeAccountID(ctx, body.SubscriptionID) + ``` +- suggested fix: Resolve the account once at the top of `authorizeAzureExchangeExecution` and pass it to the scope check, the client builder and the constraint set. +- verdict: CONFIRMED — authorizeAzureExchangeExecution (internal/api/handler_ri_exchange.go:982-1022) calls requireAzureSubscriptionScope (GetCloudAccountByExternalID at :805, scoped sessions only), buildAzureExchangeClient (:214, whenever no test factory is injected) and resolveAzureExchangeAccountID (:1029), three lookups of the same subscription with three different no-account outcomes (errNotFound, 404, unattributedAccountConstraint). +- issue: (pending cross-reference) + +### A02-023 Two copies of the recommendation scope filter plus a third name-map builder +- category: duplication +- severity: low +- location: internal/api/handler_dashboard.go:168 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `filterDashboardRecommendations` (handler_dashboard.go:168-189) and `filterRecommendationsByAllowedAccounts` (handler_recommendations.go:123-157) implement the same loop; the first swallows a `ListCloudAccounts` failure through `resolveAccountNamesByID` (scoping.go:293-296) while the second returns it as a 500, so the same user sees different outcomes for the same DB error on the dashboard and the recommendations page, and any future scope fix has to land three times. +- evidence: + ```go + nameByID := h.resolveAccountNamesByID(ctx) + filtered := make([]config.RecommendationRecord, 0, len(recs)) + for _rvc := range recs { + if recs[_rvc].CloudAccountID == nil { + continue + } + ``` +- suggested fix: Delete `filterDashboardRecommendations` and call `filterRecommendationsByAllowedAccounts` from the dashboard. +- verdict: CONFIRMED — filterDashboardRecommendations (handler_dashboard.go:168-189) and filterRecommendationsByAllowedAccounts (handler_recommendations.go:123-157) are the same scope/name-map/loop with divergent error handling: the first goes through resolveAccountNamesByID, which returns an empty map when ListCloudAccounts fails (scoping.go:293-296), while the second returns that failure as an error. +- issue: (pending cross-reference) + +### A02-025 Weak contact_email check on the public registration endpoint duplicates a stricter validator +- category: duplication +- severity: low +- location: internal/api/handler_registrations.go:429 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `validateRegistrationRequest` accepts any string of length 5 containing "@" (including CR/LF, spaces and display-name syntax), stores it, and `notifyRegistrant` later hands it to the SES sender as the To address, while the same file already CR/LF-strips `AccountName` for header-injection defence and `validateEmailFormat` (validation.go:134) exists for exactly this purpose. `"a@b\r\nBcc: x@y"` passes. +- evidence: + ```go + if !strings.Contains(req.ContactEmail, "@") || len(req.ContactEmail) < 5 { + return NewClientError(400, "contact_email must be a valid email address") + } + ``` +- suggested fix: Replace the ad-hoc check with `validateEmailFormat` and store `mail.ParseAddress(...).Address`. +- verdict: CONFIRMED — validateRegistrationRequest (handler_registrations.go:429-431) accepts any 5+ character string containing "@" while validateEmailFormat (validation.go:134-146) exists in the same package and notifyRegistrant (handler_registrations.go:240-247) hands the stored value to the notifier; the header-injection impact is nil in practice because the address travels as an SES Destination field rather than a raw header, so this stays a duplication/weak-validation finding at low severity. +- issue: (pending cross-reference) + +### A03-023 Repeated helpers where one exists +- category: duplication +- severity: low +- location: internal/auth/service_user.go:279 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `loadUser` normalises the lookup-or-not-found path, but five callers reimplement it inline with slightly different messages: `UpdateUserProfile` (service_user.go:660-669), `ChangePassword` (service_password.go:213-222), `CreateAPIKey`, `ListUserAPIKeys`, `authorizeAPIKeyAccess` and `lookupAPIKeyUser` (service_apikeys.go:44-53, 140-149, 190-199, 262-271). `DeleteUser` (service_user.go:625-633) inlines `checkLastAdminConstraint` (568-577). `resolver.go` computes `sessionSuffix` twice (185-188, 269-272). `redactEmail` exists three times across packages (auth/service_password.go:496, api/handler_auth.go:631, email/sender.go:665) with different masking rules. `MatchesAccount` (types.go:199) and `AccountScope.Allows` (account_scope.go:104) are the same loop. Divergent copies are where the A03-005 asymmetry (API-key path checks Active, session path does not) came from. +- evidence: + ```go + user, err := s.store.GetUserByID(ctx, userID) + if err != nil { + if errors.Is(err, pgx.ErrNoRows) { + return "", nil, fmt.Errorf("user not found") + } + return "", nil, fmt.Errorf("failed to get user: %w", err) + } + if user == nil { + return "", nil, fmt.Errorf("user not found") + } + ``` +- suggested fix: Route the inline lookups through `loadUser` (adding an `activeOnly` variant for the API-key path), call `checkLastAdminConstraint` from `DeleteUser`, extract `roleSessionName(accountID)`, and keep one `redactEmail` in `pkg/logging`. +- verdict: CONFIRMED — loadUser exists (service_user.go:279-291) while UpdateUserProfile:660-669, CreateAPIKey (service_apikeys.go:44-53) and lookupAPIKeyUser (:262-271) inline the same lookup, DeleteUser:625-633 repeats checkLastAdminConstraint:568-577 verbatim, the 8-char sessionSuffix is computed twice (resolver.go:185-188, :269-272), `func redactEmail` is defined in service_password.go:496, api/handler_auth.go:631 and email/sender.go:665 (plus redactEmailLocal in handler_notifications.go:132), and MatchesAccount (types.go:199-212) and AccountScope.Allows (account_scope.go:104-117) are the same loop. +- issue: (pending cross-reference) + +### A06-020 Cache-Control switch is duplicated between the HTTP and Lambda static paths +- category: duplication +- severity: low +- location: internal/server/static.go:39 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `setCacheHeaders` (used by `spaHandler`) and `cacheControlForExt` (used by `serveStaticForLambda`, static.go:177) carry byte-identical extension lists and header values. Adding `.avif` or `.map` to one and not the other gives the same asset a different cache policy depending on whether the deployment is a container or a Lambda Function URL — the exact drift class this file's `resolveStaticFilePath` refactor (04-M6) was created to close. +- evidence: + ```go + func setCacheHeaders(w http.ResponseWriter, urlPath string) { + ext := strings.ToLower(path.Ext(urlPath)) + switch ext { + case ".html": + w.Header().Set("Cache-Control", "no-cache, no-store, must-revalidate") + ``` +- suggested fix: make `setCacheHeaders` call `cacheControlForExt(strings.ToLower(path.Ext(urlPath)))`. +- verdict: CONFIRMED — the two switches are byte-identical in case labels and header values (internal/server/static.go:38-50 vs 175-186), one used by `spaHandler` (static.go:33) and the other by `serveStaticForLambda` (static.go:198), with no shared helper between them. No behavioural divergence exists today; the finding is a verified duplication/drift hazard rather than a live bug. +- issue: (pending cross-reference) + +### A08b-040 Synapse and Search re-declare the shared pricing types the pricing package exists to centralize +- category: duplication +- severity: low +- location: providers/azure/services/synapse/client.go:110 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `search/client.go:115`) +- failure scenario: `pricing/types.go:3` documents `RetailPriceItem` as the single shape "so a field rename or addition only requires a single edit", and cache, cosmosdb, database and managedredis alias it. Synapse declares its own copy omitting `Location` and `UnitOfMeasure`, and search declares its own `AzureRetailPrice` envelope instead of aliasing `pricing.Page`. A new field added to the shared item (unit of measure is the one A08b-010 needs) silently decodes to zero in these two packages. +- evidence: + ```go + type SynapseRetailPriceItem struct { + CurrencyCode string `json:"currencyCode"` + RetailPrice float64 `json:"retailPrice"` + UnitPrice float64 `json:"unitPrice"` + ``` +- suggested fix: alias `pricing.RetailPriceItem` and `pricing.Page` as the four sibling clients do. +- verdict: CONFIRMED — `SynapseRetailPriceItem` (synapse:110-122) omits both `Location` and `UnitOfMeasure`, which the shared item carries at providers/azure/internal/pricing/types.go:15 and :21 with the single-edit rationale at lines 9-11; search re-declares the envelope at 115-119 rather than aliasing `pricing.Page`, though its `Items` field does use the shared item type. +- issue: (pending cross-reference) + +### A09-012 Confidence thresholds are duplicated between pkg/exchange and internal/api with the same magic numbers +- category: duplication +- severity: medium +- location: pkg/exchange/reshape.go:328 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `confidenceComponent` reimplements `confidenceBucketFor` at `internal/api/handler_recommendations.go:382`, which itself mirrors `frontend/src/recommendations.ts`. The `200` / `50` / `3` cut-offs are inline literals on all three sides with no shared constant. Tuning the bucket in the API layer leaves the reshape ranking scoring the same offerings under the old thresholds, so the UI badge and the ordering it accompanies disagree, and only the API side has a test (`handler_recommendations_test.go:765`). +- evidence: + ```go + switch { + case savings >= 200 && count >= 3: + return 1.0 + case savings >= 50: + return 0.5 + default: + return 0.0 + } + ``` +- suggested fix: export the three thresholds as named constants from `pkg/exchange` (the module `internal/api` can import) and have `confidenceBucketFor` read them, leaving only the frontend as documented cross-language duplication. +- verdict: CONFIRMED — the Go-side duplication is real: the 200/50/3 literals appear inline in `confidenceComponent` (pkg/exchange/reshape.go:326-334) and again in `confidenceBucketFor` (internal/api/handler_recommendations.go:382-393), with a test only on the API side (handler_recommendations_test.go:765). One correction: the claimed third copy does not exist — `/usr/bin/grep -rn "confidence" frontend/src/recommendations.ts` returns nothing, and handler_recommendations.go:373-377 documents that the heuristic was moved off the client. The two Go copies have also already drifted (the API side clamps `count` to 1; the exchange side returns a neutral 0.5 when count is 0). +- severity-adjusted: low — drift risk only, with no failure scenario beyond a ranking/badge mismatch, and one of the three cited call sites does not exist. +- issue: (pending cross-reference) + +### A15-007 MockHTTPClient and createMockHTTPResponse are copy-pasted into four Azure service test files +- category: duplication +- severity: low +- location: providers/azure/services/cache/client_test.go:94 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `providers/azure/mocks` already exports `MockHTTPClient` and `CreateMockHTTPResponse`, and `compute`, `managedredis` and `synapse` import them. Four other service test files redeclare a byte-identical local copy instead: `cache/client_test.go:94`, `cosmosdb/client_test.go:104`, `database/client_test.go:87`, `search/client_test.go:92`. A behaviour change to the shared double (adding call recording, or a `Do` that honours `ctx`) silently reaches three services and misses four, so the two halves of the Azure suite drift apart. +- evidence: + ```go + type MockHTTPClient struct { + mock.Mock + } + + func (m *MockHTTPClient) Do(req *http.Request) (*http.Response, error) { + args := m.Called(req) + if args.Get(0) == nil { + return nil, args.Error(1) + } + return args.Get(0).(*http.Response), args.Error(1) + } + ``` +- suggested fix: Delete the four local copies and import `providers/azure/mocks`, as the three sibling services already do. +- verdict: CONFIRMED — the split is exactly 4/3 and measurable: `type MockHTTPClient struct` is declared in providers/azure/mocks/azure_mocks.go:122 plus cache/client_test.go:94, cosmosdb/client_test.go:104, database/client_test.go:87 and search/client_test.go:92, and a count of `mocks.MockHTTPClient`/`mocks.CreateMockHTTPResponse` references gives 51 in compute, 41 in managedredis and 28 in synapse against 0 in all four shadowing packages. One correction: the copies are not byte-identical to the shared type, which also carries `ResponseBody string` and `StatusCode int` fields (azure_mocks.go:123-124); only the `Do` method and the response helper match verbatim, which if anything strengthens the drift argument. +- issue: (pending cross-reference) + +### A15-008 azureTermString is duplicated in six packages and reimplemented inline in a seventh +- category: duplication +- severity: low +- location: providers/azure/services/cache/client.go:591 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The same four-line helper is declared byte-identically in `cache`, `compute`, `cosmosdb`, `database`, `managedredis` and `search`. `synapse` does not call it but recomputes the identical rule inline at `providers/azure/services/synapse/client.go:458-461`, which is why a name-based search finds six copies rather than seven. All seven encode the Retail Prices API's `"1 Year"` / `"N Years"` term spelling. If that spelling ever changes, a fix applied to the shared helper still leaves the synapse copy matching nothing, and `extractSynapsePricing` then silently finds no reservation price. +- evidence: + ```go + // synapse/client.go:458 — the seventh copy, written differently + termStr := fmt.Sprintf("%d Year", termYears) + if termYears > 1 { + termStr = fmt.Sprintf("%d Years", termYears) + } + ``` +- suggested fix: Move the helper to the shared `providers/azure/internal/pricing` package that four of these services already import, and have synapse call it instead of recomputing the rule. +- verdict: CONFIRMED — `func azureTermString` is declared in exactly six packages (cache/client.go:591, compute:757, cosmosdb:589, database:609, managedredis:511, search:502) and synapse recomputes the identical rule inline at providers/azure/services/synapse/client.go:458-461, so a name-based search does under-report the seventh; all seven of those service clients already import `providers/azure/internal/pricing`, so the suggested home is reachable without a new dependency. The failure scenario is conditional on the Retail Prices API changing its term spelling, so this stands as a duplication finding rather than a live defect. +- issue: (pending cross-reference) + +### Category: dead-code + +31 findings: 8 medium, 23 low. + +### A05-009 The whole Azure commitment-options prober is unreachable +- category: dead-code +- severity: medium +- location: internal/commitmentopts/probe_azure.go:107 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ProbeAzure` and `DefaultAzureProbers` have no non-test callers anywhere in the repo, and the wiring the doc comment points at, `Service.probeAndPersistAzure`, does not exist — `probeAndPersist` (service.go:76) only ranges over `s.probers`, which `AzureSPProber` cannot join because its `Probe` signature takes an `azcore.TokenCredential` rather than an `aws.Config`. So 212 lines of production code plus 304 lines of tests describe a feature that never runs, and `Validate` for `("azure", "savingsplans", …)` always takes the permissive-true branch. A reader auditing whether Azure savings-plan term/payment combos are validated will conclude they are. +- evidence: + ```go + // The method signature differs from the AWS Prober interface because Azure + // authentication uses azcore.TokenCredential rather than aws.Config. Callers + // use this method directly; Service.probeAndPersistAzure wires it up. + func (p *AzureSPProber) ProbeAzure(ctx context.Context, cred azcore.TokenCredential) ([]Combo, error) { + ``` +- suggested fix: either wire it up (an Azure sibling of `probeAndPersist` that resolves a credential and calls `ProbeAzure`) or delete the file and its tests. Leaving it in place with a comment naming a function that does not exist is the worst of the three. +- verdict: CONFIRMED — `git grep -n "ProbeAzure\|DefaultAzureProbers\|probeAndPersistAzure" -- '*.go'` matches only probe_azure.go and probe_azure_test.go; probeAndPersistAzure does not exist anywhere, probeAndPersist (service.go:76-121) ranges only over `s.probers` typed as the aws.Config-taking Prober, and the Azure prober cannot satisfy it. Validate for ("azure", ...) therefore hits either the ErrNoData branch (service.go:148) or the missing-provider branch (service.go:154) and returns permissive true. +- issue: (pending cross-reference) + +### A08-022 GroupCommitments is unused and would rebuild the #1538 wrong-commitment-type bug if called +- category: dead-code +- severity: medium +- location: providers/gcp/services/computeengine/client.go:555 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GroupCommitments` and `CommitmentRequest` have no caller outside their own tests. The function aggregates vCPUs and memory across every machine type in a (account, region, term) group and emits a request with no commitment `Type` at all — the exact defect `buildInsertRequest` guards against with `commitmentTypeForMachineType` and documents as issue #1538 ("an N2/C3/M3/... recommendation bought a commitment that applied to none of its instances: full commitment spend plus undimmed on-demand charges, booked as realized savings"). The generated `Name` is also timestamp-based, so a re-drive through it would double-buy. +- evidence: + ```go + result = append(result, CommitmentRequest{ + Name: fmt.Sprintf("cud-%s-%d-%d", k.region, ts, counter), + Plan: a.plan, + Region: k.region, + Resources: []ResourceCommitment{ + {Type: computepb.ResourceCommitment_VCPU.String(), Amount: a.vcpus}, + {Type: computepb.ResourceCommitment_MEMORY.String(), Amount: a.memoryMB}, + }, + }) + ``` +- suggested fix: delete `GroupCommitments`, `CommitmentRequest` and `ResourceCommitment` along with their tests; `buildInsertRequest` is the only supported way to construct a CUD insert. +- verdict: CONFIRMED — `GroupCommitments` (client.go:555) and `CommitmentRequest` (client.go:541) have no reference outside client.go and client_test.go, the emitted request carries no commitment `Type` field at all, unlike `buildInsertRequest` (client.go:758) which derives it via `commitmentTypeForMachineType` (client.go:190) for issue #1538, and the name is `time.Now().UnixNano()`-based (client.go:586, 591). +- issue: (pending cross-reference) + +### A08b-036 `convertAzureSearchRecommendation` is unreachable production code kept alive by its own tests +- category: dead-code +- severity: medium +- location: providers/azure/services/search/client.go:536 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GetRecommendations` at line 127 returns an empty slice unconditionally, so nothing in production calls the converter. `/usr/bin/grep -rn convertAzureSearchRecommendation providers/` returns only the definition and two calls in `client_test.go:675` and `client_test.go:691`. The tests pass, the coverage report counts the function as exercised, and the `ctx` parameter it accepts is never read — a reviewer scanning coverage sees a tested conversion path that cannot run. +- evidence: + ```go + func (c *SearchClient) GetRecommendations(_ context.Context, _ *common.RecommendationParams) ([]common.Recommendation, error) { + return []common.Recommendation{}, nil + } + ``` +- suggested fix: delete the converter and its tests until Azure exposes a Search recommendations resource type, so coverage reflects reachable code. +- verdict: CONFIRMED — `GetRecommendations` returns an empty slice unconditionally (search:127-129), and `/usr/bin/grep -rn convertAzureSearchRecommendation providers/` returns only the doc line, the definition at 536 and the two test calls at client_test.go:675 and :691; reading the body at 536-567 confirms `ctx` is never used. +- issue: (pending cross-reference) + +### A08b-037 Cloud Storage and Memorystore carry full resource-creation plumbing that no production path reads +- category: dead-code +- severity: medium +- location: providers/gcp/services/cloudstorage/client.go:26 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `memorystore/client.go:27`, partially at `cloudsql/client.go:25`) +- failure scenario: `PurchaseCommitment` was changed to a not-supported no-op and `GetExistingCommitments` returns nil, so `c.storageService` is written by `SetStorageService` and never read; `/usr/bin/grep -n storageService providers/gcp/services/cloudstorage/client.go` shows only the field, the setter and the type declarations. The same holds for memorystore's `redisService`, `CreateInstanceOperation` and `realRedisService`, and for `SQLAdminService.InsertInstance`/`ListInstances` in cloudsql. This keeps `cloud.google.com/go/storage` and the Redis create-instance surface linked in, and leaves an `Insert`/`Create` capability wired to a client whose documented purpose is that it must never create billable resources (issue #640). +- evidence: + ```go + type StorageService interface { + Buckets(ctx context.Context, projectID string) BucketIterator + Bucket(name string) BucketHandle + Close() error + } + ``` +- suggested fix: delete the unused interfaces, wrappers, setters and their mocks, so the create-resource capability is not reachable from a client that must not create resources. +- verdict: CONFIRMED — `PurchaseCommitment` is a not-supported no-op (cloudstorage:245-255) and `GetValidResourceTypes` is a hardcoded four-entry list (313-323), so `storageService` appears only as the field at 64 and the setter at 81; memorystore's `redisService` is likewise only field 65 and setter 82 with `CreateInstance` wired at 104; cloudsql reaches the API through `ListTiers` (312), leaving `ListInstances` and `InsertInstance` (26-27, 88-95) unreferenced. +- issue: (pending cross-reference) + +### A09-013 pkg/errors is a fully exported package with no importer anywhere in the repository +- category: dead-code +- severity: medium +- location: pkg/errors/errors.go:12 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `git grep -l 'LeanerCloud/CUDly/pkg/errors"'` outside `pkg/errors/` returns zero files. The package ships 7 error types, 7 constructors, 7 `Is*` helpers, 7 field-insensitive `Is` methods and 6 sentinels, plus 312 lines of tests, none of which any caller reaches. Its field-insensitive `Is` contract is also a trap for a future adopter: `errors.Is(err, &NotFoundError{ID: "x"})` is true for any `*NotFoundError` regardless of ID. +- evidence: + ```go + // Package errors provides custom error types for CUDly. + // + // Type-level Is matching: every error type in this package implements Is by + // matching purely on the dynamic type of the target, ignoring the target's + // struct fields. + package errors + ``` +- suggested fix: delete the package, or adopt it at the boundaries that currently hand-roll `fmt.Errorf` sentinels; leaving an unused public error taxonomy invites divergent adoption later. +- verdict: CONFIRMED — I rebuilt the import graph myself rather than trusting the reviewer: `git grep -n "pkg/errors"` across the whole tree (all six modules named in go.work: root, pkg, providers/{aws,azure,gcp}, tests/e2e) returns only two go.sum lines for the unrelated `github.com/pkg/errors` dependency. No file imports `github.com/LeanerCloud/CUDly/pkg/errors`. Line counts check out: errors.go is 299 lines, errors_test.go 312, and the six sentinels are declared at pkg/errors/errors.go:281-298. +- issue: (pending cross-reference) + +### A09-014 pkg/config is dead and is the sole reason the pkg module depends on pflag and yaml.v3 +- category: dead-code +- severity: medium +- location: pkg/config/load.go:1 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `git grep -l 'LeanerCloud/CUDly/pkg/config"'` outside `pkg/config/` returns zero files. The 399-line YAML/env/flag precedence loader, its 91-line type set and 267 lines of tests have no caller. `pkg/go.mod` carries `github.com/spf13/pflag v1.0.5` and `gopkg.in/yaml.v3 v3.0.1` as direct requirements purely for this package, so every consumer of the shared module inherits two dependencies for unreachable code. The package also defines a second `DefaultScorerConfig` that no caller can reconcile against the live scorer configuration. +- evidence: + ```go + import ( + "github.com/LeanerCloud/CUDly/pkg/scorer" + "github.com/spf13/pflag" + "gopkg.in/yaml.v3" + ) + ``` +- suggested fix: delete `pkg/config` and drop `pflag` and `yaml.v3` from `pkg/go.mod`, or move it to the root module next to the CLI that would actually use it. +- verdict: CONFIRMED — I established the import graph independently and read the package the reviewer skipped. `git grep -n "pkg/config"` across all six go.work modules returns exactly one hit, and it is a negative: pkg/scorer/scorer.go:2 says "must not import pkg/config". Nothing imports `github.com/LeanerCloud/CUDly/pkg/config`. Line counts are exact (load.go 399, types.go 91, config_test.go 267). The dependency claim holds: `git grep -n "spf13/pflag\|yaml.v3" -- pkg/` shows the only importers are pkg/config/load.go:12-13 and its test, while both sit in `pkg/go.mod`'s direct require block at lines 13 and 16. One nuance on the fix: yaml.v3 would become indirect rather than droppable, since testify pulls it in. +- issue: (pending cross-reference) + +### A13-009 The whole `terraform/modules/monitoring/` tree is never instantiated +- category: dead-code +- severity: medium +- location: terraform/modules/monitoring/README.md:38 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: No environment root declares `module "monitoring"`; grepping `modules/monitoring` across the repo hits only that README. About 2,500 lines of unexercised Terraform (aws 465 + azure 507 + gcp 659 lines of `main.tf`, plus outputs, variables and a 788-line README) declare SNS topics, KMS keys, metric filters, dashboards and alarms that no apply has ever created, and none of it is `terraform validate`-covered by an environment plan. `.trivyignore:49` accepts an SNS-encryption finding whose only source is this dead tree, which is how a suppression outlives the resource it describes. +- evidence: + ```hcl + # README shows the intended call site; no .tf file in the repo contains it + source = "../../modules/monitoring/aws" + ``` +- suggested fix: Either wire the module into the environment roots behind an `enable_monitoring` flag, or delete the tree and the suppressions that exist only for it. +- verdict: CONFIRMED — `/usr/bin/grep -rn modules/monitoring .` returns ten hits, all inside terraform/modules/monitoring/README.md; no `.tf` file anywhere declares the module. Line counts match the finding (aws/main.tf 465, azure/main.tf 507, gcp/main.tf 659, README.md 788). The suppression link holds too: `aws_sns_topic` is declared exactly once in the repo, at terraform/modules/monitoring/aws/main.tf:11, so the AVD-AWS-0136 entry at .trivyignore:49-52 has no other source. +- issue: (pending cross-reference) + +### A13c-003 Secret rotation is unappliable: the rotation Lambda points at a zip that does not exist in the repo +- category: dead-code +- severity: medium +- location: terraform/modules/secrets/aws/main.tf:282 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `find . -name 'rotation_function*'` returns nothing at this commit, and + `terraform/modules/secrets/aws/` contains only `main.tf`, `variables.tf`, `outputs.tf` and the + lockfile. Setting `enable_secret_rotation = true` (which `environments/aws/secrets.tf:32` + exposes as `var.enable_secret_rotation`, and whose own comment tells production tfvars to set + it true) fails the apply on a missing file. `source_code_hash` already `fileexists`-guards + itself, so the absence was known; `filename` does not, so the whole rotation subsystem — the + Lambda, its role, three policies, the `aws_secretsmanager_secret_rotation` — is unreachable. +- evidence: + ```hcl + filename = "${path.module}/rotation_function.zip" + source_code_hash = fileexists("${path.module}/rotation_function.zip") ? filebase64sha256("${path.module}/rotation_function.zip") : null + ``` +- suggested fix: either ship the rotation function (or build it with `archive_file`), or delete + the rotation block and the `enable_secret_rotation` / `rotation_days` / `rds_cluster_id` + variables so no tfvars can request a configuration that cannot apply. +- verdict: CONFIRMED — `find . -name 'rotation_function*'` returns nothing and + `terraform/modules/secrets/aws/` holds only main.tf, variables.tf, outputs.tf and the lockfile; + `filename` at secrets/aws/main.tf:282 is unguarded while `source_code_hash` on the next line is + `fileexists`-guarded, and `environments/aws/secrets.tf:32` wires `var.enable_secret_rotation` + straight through, so flipping it aborts the apply on the missing zip. +- issue: (pending cross-reference) + +### A01-015 findDuplicatePendingExecution is dead code kept alive only by its test; revokeViaSession keeps a dead variable +- category: dead-code +- severity: low +- location: internal/api/handler_purchases.go:2521 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `findDuplicatePendingExecution` has no non-test caller (`/usr/bin/grep -rn findDuplicatePendingExecution --include='*.go'` returns only its definition and `TestFindDuplicatePendingExecution`, handler_purchases_guards_test.go:481); the live duplicate check is `matchDuplicateInList` inside `persistExecutionAndSuppressions`. The test therefore pins behaviour of code the product never runs. In `revokeViaSession` the CAS result is discarded with `_ = updated // kept for future use` (1471). +- evidence: + ```go + func (h *Handler) findDuplicatePendingExecution(ctx context.Context, creatorID, key string, now time.Time) (*config.PurchaseExecution, error) { + pending, err := h.config.GetPendingExecutions(ctx) + // ... + _ = updated // kept for future use; CancelledBy is now persisted atomically + ``` +- suggested fix: Delete `findDuplicatePendingExecution` and its test (or point the test at `matchDuplicateInList`), and drop the `updated` return binding in `revokeViaSession`. +- verdict: CONFIRMED — /usr/bin/grep finds findDuplicatePendingExecution only at internal/api/handler_purchases.go:2513-2547 and handler_purchases_guards_test.go:509-564; the live duplicate check is matchDuplicateInList:2419 via persistExecutionAndSuppressions:2461, and `_ = updated` sits at :1471. +- issue: (pending cross-reference) + +### A02-014 Dead middleware and helper code, including an unreachable CSRF exemption branch +- category: dead-code +- severity: low +- location: internal/api/middleware.go:314 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `requiresCSRFValidation` is only called from `validateSecurityContext` (handler.go:784), which returns at line 762 whenever `isPublicEndpoint` is true; every prefix in `csrfExemptWhenTokenOnly` is also in `isPublicEndpoint`, so lines 314-325 can never execute and the comment at 280-297 describes behaviour that lives elsewhere. Grep of non-test callers also finds none for `authenticate` / `checkUserAPIKey` / `checkBearerToken` (middleware.go:100-112, 219-239), `validateRequest` and `validateSecurity` (handler.go:729-732, 756-759), `formatNotFoundError` (router.go:889), `toAPIPermissions` (types_apikeys.go:34) and `formatTimePtr` (handler_apikeys.go:135). Each is exercised only by its own unit test, so the tests assert code the product never runs. +- evidence: + ```go + csrfExemptWhenTokenOnly := []string{ + "/api/purchases/approve/", + "/api/purchases/cancel/", + "/api/ri-exchange/approve/", + "/api/ri-exchange/reject/", + } + for _, prefix := range csrfExemptWhenTokenOnly { + if strings.HasPrefix(path, prefix) { + return h.extractBearerToken(req) != "" + ``` +- suggested fix: Delete the unreachable block and the unused helpers together with their tests; keep `authenticatePrincipal` as the single auth path. +- verdict: CONFIRMED — /usr/bin/grep over internal/ and cmd/ excluding *_test.go finds no callers of authenticate, validateRequest, validateSecurity, formatNotFoundError, toAPIPermissions or formatTimePtr, and checkUserAPIKey/checkBearerToken are reached only from the dead authenticate (middleware.go:107-111); requiresCSRFValidation's sole caller is validateSecurityContext (handler.go:784), which returns at handler.go:762 for every path in isPublicEndpoint, and all four csrfExemptWhenTokenOnly prefixes (middleware.go:314-319) are in that list (middleware.go:23-28), so lines 320-325 cannot execute. +- issue: (pending cross-reference) + +### A03-022 Dead code left over from the removed role model and superseded paths +- category: dead-code +- severity: low +- location: internal/auth/types.go:321 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `RoleAdmin`, `RoleUser`, `RoleReadOnly` have zero references outside their declaration (Session.Role was removed in #907/#940); `Service.UpdateLastUsed` (service_apikeys.go:442) has no non-test caller; `ResolveGCPCredentials` (credentials/resolver.go:305) has no non-test caller; `EmailSenderInterface.SendWelcomeEmail` (interfaces.go:71) is never invoked by the auth service and still carries a `role` string parameter; `User.Salt` is written as `""` at every site and read by nothing. Each keeps the pre-#907 vocabulary alive for the next reader and widens the surface every mock has to implement. +- evidence: + ```go + // Predefined roles. + const ( + RoleAdmin = "admin" + RoleUser = "user" + RoleReadOnly = "readonly" + ) + ``` +- suggested fix: Delete the constants, the two unused functions, the unused interface method (and its mock implementations), and the `Salt` field once the column is dropped. +- verdict: CONFIRMED — repo-wide greps show RoleAdmin/RoleUser/RoleReadOnly referenced only at their declaration (types.go:321-325), no non-test caller of Service.UpdateLastUsed (service_apikeys.go:442) or ResolveGCPCredentials (resolver.go:305), SendWelcomeEmail present only in interface declarations and sender implementations with no invocation from internal/auth, and User.Salt assigned "" at service_user.go:744, service_password.go:251, :487 and only round-tripped by the store (store_postgres.go:113, :225, :451, :754). +- issue: (pending cross-reference) + +### A04-015 ri_exchange_history.cloud_account_id and RIExchangeRecord.CloudAccountID are never written or read +- category: dead-code +- severity: low +- location: internal/config/store_postgres.go:2639-2646 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (column at migrations/000011_cloud_accounts.up.sql:140-141; field at internal/config/types.go:985) +- failure scenario: the column added for multi-tenant scoping is NULL on every row and absent from every SELECT, so `RIExchangeRecord.CloudAccountID` is always nil. Exchange-history scoping (`internal/api/handler_ri_exchange.go:2064-2085`) has to fall back to matching the provider account-number string, the same ambiguity purchase_history needed the dual-column predicate for (#701/#866). Anyone who adds a `cloud_account_id = ANY($1)` filter here would silently return zero rows. +- evidence: + ```go + INSERT INTO ri_exchange_history ( + id, account_id, exchange_id, region, source_ri_ids, + source_instance_type, source_count, target_offering_id, + target_instance_type, target_count, payment_due, + status, approval_token, error, mode, completed_at, expires_at, + created_at, updated_at, created_by_user_id, ladder_run_id + ) + ``` +- suggested fix: either populate and project the column (resolve via `GetCloudAccountByExternalID` at save time) or drop the column and the struct field so no caller can filter on it. +- verdict: CONFIRMED — `SaveRIExchangeRecord`'s twenty-one-column INSERT omits `cloud_account_id` (internal/config/store_postgres.go:2639-2651) and none of the four read projections include it (internal/config/store_postgres.go:2685-2691, 2710-2716, 2735-2743, 2980-2986), so `RIExchangeRecord.CloudAccountID` (internal/config/types.go:985) is always nil; exchange-history scoping consequently filters on the `AccountID` string (internal/api/handler_ri_exchange.go:2069-2085). Column added at internal/database/postgres/migrations/000011_cloud_accounts.up.sql:140-141. +- issue: (pending cross-reference) + +### A05-019 `maskActor`'s length guard is always true and its variable unused +- category: dead-code +- severity: low +- location: internal/purchase/approvals.go:95 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `actor == ""` is rejected two lines earlier, so `len(actor)-1 >= 0` always holds, and `at` is never read inside the block. The `if` neither guards the loop nor communicates anything; it reads as though an index check were being performed on a PII-masking path, which is where a reader most wants the control flow to be obvious. +- evidence: + ```go + if at := len(actor) - 1; at >= 0 { + for i, c := range actor { + if c == '@' { + return "***" + actor[i:] + } + } + } + ``` +- suggested fix: drop the `if` and keep the loop; better still, use `strings.Index(actor, "@")` so the intent is stated once. +- verdict: CONFIRMED — approvals.go:89-91 returns early on `actor == ""`, so at approvals.go:95 `len(actor) >= 1` and `at := len(actor) - 1` is always >= 0; the block body (:96-100) ranges over `actor` and never reads `at`. The `if` is unconditionally taken and guards nothing. +- issue: (pending cross-reference) + +### A07-027 opensearch.matchesPaymentOption is dead production code kept alive by its own test +- category: dead-code +- severity: low +- location: providers/aws/services/opensearch/client.go:483 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the only reference outside the definition is `client_test.go:333`. `scanOpenSearchOfferingPage` does its matching through `normalizeOpenSearchPaymentOption` instead. The two encode different behaviour for an unrecognised option (this one returns false, the live one passes the string through), so the test asserts a contract the purchase path does not use and a reader tracing payment-option handling finds the wrong function first. +- evidence: + ```go + func (c *Client) matchesPaymentOption(offeringOption types.ReservedInstancePaymentOption, required string) bool { + switch required { + case "all-upfront": + return offeringOption == types.ReservedInstancePaymentOptionAllUpfront + ``` +- suggested fix: delete the method and its test; the live behaviour is already covered through `scanOpenSearchOfferingPage`. +- verdict: CONFIRMED — a repo-wide grep for `matchesPaymentOption` finds the OpenSearch method only at its definition (providers/aws/services/opensearch/client.go:483) and at client_test.go:333; the live comparison uses `normalizeOpenSearchPaymentOption` (:448, compared at :457), and the two differ on an unrecognised option exactly as described (:492 returns false, :479 passes through). Redshift's identically-named free function is a separate symbol with real callers. +- issue: (pending cross-reference) + +### A07-032 describeInputFromQuery carries an unreachable offering-class default +- category: dead-code +- severity: low +- location: providers/aws/services/ec2/client.go:430 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the only caller is `findOfferingID`, which sets `q.offeringClass` from `resolveOfferingClassType` on the line above; that function returns either a non-empty enum value or an error. `oc == ""` is therefore unreachable, and the duplicated default means a future caller that forgets to set the field silently buys convertible rather than failing the way `resolveOfferingClassType` was written to. +- evidence: + ```go + oc := q.offeringClass + if oc == "" { + oc = types.OfferingClassTypeConvertible + } + ``` +- suggested fix: delete the branch and let `resolveOfferingClassType` remain the single place that decides the default. +- verdict: CONFIRMED — the only production caller of `describeInputFromQuery` is providers/aws/services/ec2/client.go:522, reached from `findOfferingID` which sets `q.offeringClass` from `resolveOfferingClassType` at :495-499, and that function (:385-394) returns either a non-empty enum or an error, so `oc == ""` at :431 is unreachable outside client_test.go:1176-1182. +- issue: (pending cross-reference) + +### A08-026 fetchOnDemandRate is exercised only by its own tests +- category: dead-code +- severity: low +- location: providers/azure/services/savingsplans/client.go:455 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: no production code calls `fetchOnDemandRate`; the only callers are `client_test.go:533/552/571`. It is the sole reason the savings plans client holds an `httpClient` field at all, so the tests keep alive a code path and a dependency the client does not otherwise use. +- evidence: + ```go + func (c *Client) fetchOnDemandRate(ctx context.Context, planType string) (float64, error) { + ``` +- suggested fix: delete the function and its tests, or wire it into `GetOfferingDetails` where an on-demand rate is actually needed to report savings. +- verdict: CONFIRMED — `fetchOnDemandRate` (savingsplans/client.go:455) is referenced only at client_test.go:533/552/571, and the `httpClient` field's only production read is inside that function (client.go:465); every other mention is the constructor assignment (client.go:79, 87-94) or a test. +- issue: (pending cross-reference) + +### A08-027 RecommendationsClientAdapter stores a context it never reads +- category: dead-code +- severity: low +- location: providers/gcp/recommendations.go:73 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `GetRecommendationsClient` (provider.go:491) populates the `ctx` field, but every method on the adapter takes its own `ctx` parameter and none reads `r.ctx`. The field invites a future edit to use the stored (typically `context.Background()`, provider.go:116) context for an RPC, which would silently ignore the caller's deadline and cancellation. +- evidence: + ```go + type RecommendationsClientAdapter struct { + ctx context.Context + projectID string + clientOpts []option.ClientOption + } + ``` +- suggested fix: drop the `ctx` field and the assignment in `GetRecommendationsClient`. +- verdict: CONFIRMED — the field is written once at provider.go:489 and never read by production code: the only `.ctx` reads in providers/gcp are assertions in provider_test.go:265/336 and recommendations_test.go:122, and every adapter method takes its own `ctx` (e.g. `GetRecommendations`, recommendations.go:103). The parenthetical is imprecise, though: the field stores the caller's `ctx` argument, not provider.go:116's `context.Background()`. +- issue: (pending cross-reference) + +### A08b-038 All four GCP clients store a `context.Context` in the struct and never read it +- category: dead-code +- severity: low +- location: providers/gcp/services/cloudsql/client.go:49 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (same at `cloudstorage/client.go:60`, `memorystore/client.go:61`, `computeengine/client.go:255`) +- failure scenario: `NewClient` captures the caller's context into the struct, but every method takes its own `ctx` parameter and `/usr/bin/grep -n "c\.ctx"` returns no matches in any of the four files. The field is a live trap: a future method that reads it would silently use the construction-time context, which for a Lambda-scoped client is already cancelled by the time a later request runs. +- evidence: + ```go + type CloudSQLClient struct { + ctx context.Context + projectID string + region string + ``` +- suggested fix: drop the field and the `ctx` parameter from `NewClient`. +- verdict: CONFIRMED — the field is declared at cloudsql:49, cloudstorage:60, memorystore:61 and computeengine:255, each `NewClient` takes a `ctx` to fill it (cloudsql:59, cloudstorage:70, memorystore:71, computeengine:266), and `/usr/bin/grep -rn "c\.ctx" providers/gcp/` returns nothing at this commit. +- issue: (pending cross-reference) + +### A09-015 RIExchangeStore forces every implementer to carry an uncalled method +- category: dead-code +- severity: low +- location: pkg/exchange/auto.go:49 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `CancelAllPendingExchanges` is in the interface, implemented at `internal/config/store_postgres.go:2914`, mocked at `internal/mocks/stores.go:694`, and tested — but no production caller invokes it. `RunAutoExchange` uses `CancelPendingExchangesByOrigin` instead, and the doc comment says so. The method is a live "cancel every pending exchange regardless of origin" primitive kept reachable only through the interface, which is the exact cross-origin behaviour the origin-scoped replacement was written to prevent. +- evidence: + ```go + // CancelAllPendingExchanges cancels every pending record regardless of origin. + // Kept for interface compatibility; RunAutoExchange now calls + // CancelPendingExchangesByOrigin instead to avoid cross-origin contamination. + CancelAllPendingExchanges(ctx context.Context) (int64, error) + ``` +- suggested fix: remove the method from `RIExchangeStore` (and from the store, if nothing else calls it) so the dangerous unscoped variant is not one method call away. +- verdict: CONFIRMED — `git grep -n CancelAllPendingExchanges` returns the interface declaration (pkg/exchange/auto.go:49), the store implementation (internal/config/store_postgres.go:2914), the adapter passthrough (internal/server/handler_ri_exchange.go:296), four mocks and six tests. No production code path invokes it; `RunAutoExchange` calls `CancelPendingExchangesByOrigin` at auto.go:171 instead. +- issue: (pending cross-reference) + +### A09-024 IsSavingsPlan ends with a comparison that can never be reached +- category: dead-code +- severity: low +- location: pkg/common/types.go:157 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ServiceSavingsPlansAll` is declared at line 107 with the value `"savingsplans"`, and the switch above already matches it. The trailing `return string(s) == "savingsplans"` can therefore only be evaluated for values the switch rejected, all of which differ from `"savingsplans"`. It always returns false and can never be the deciding branch. The doc comment claims it exists to catch "the dash-free frontend spelling that the API handler stores verbatim", implying a case the constant does not already cover. +- evidence: + ```go + case ServiceSavingsPlansAll, + ServiceSavingsPlansCompute, + ... + return true + } + return string(s) == "savingsplans" + ``` +- suggested fix: replace the trailing comparison with `return false`, and if a dash-form legacy spelling still needs recognising, match `"savings-plans"` explicitly instead. +- verdict: CONFIRMED — `ServiceSavingsPlansAll ServiceType = "savingsplans"` (pkg/common/types.go:107) is the first case of the switch at types.go:150-155, so any `s` reaching the trailing `return string(s) == "savingsplans"` at types.go:157 has already been rejected by that case and cannot equal the literal. The line always returns false. The doc comment at types.go:143-146 describes it as covering a spelling the constant already covers. +- issue: (pending cross-reference) + +### A10-026 `Config.Providers` is never bound to a flag or read +- category: dead-code +- severity: low +- location: cmd/main.go:47 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The field is the only occurrence of the identifier in the whole `cmd/` tree (`/usr/bin/grep -rn "Providers" cmd/` returns one line). It is declared on the config struct that documents the CLI's surface, so a reader adding multi-provider support reasonably assumes a `--providers` flag exists and wires against it; nothing does. The main pipeline is AWS-only (`recClient := awsprovider.NewRecommendationsClient`, cmd/multi_service.go:122). +- evidence: + ```go + ExcludeAccounts []string + Providers []string + IncludeRegions []string + ``` +- suggested fix: Delete the field. +- verdict: CONFIRMED — `/usr/bin/grep -rn "Providers" cmd/` returns only cmd/main.go:47, and a repo-wide search for a `.Providers` selector finds no reader anywhere. +- issue: (pending cross-reference) + +### A10-027 `processService` / `processRegionRecommendations` / `applyCommonCoverage` are reachable only from tests +- category: dead-code +- severity: low +- location: cmd/multi_service.go:653 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `processService` is called from six test functions and nowhere in production; its only caller of `processRegionRecommendations` is itself, and `applyCommonCoverage` is called once, from `multi_service_helpers_test.go:277`. The comments already concede this ("Used by legacy callers", "This legacy path (test-only)"). Roughly 140 lines of purchase-executing code, including a `--max-instances` refusal guard at cmd/multi_service_helpers.go:412 that protects a state no production caller can reach, are maintained and reviewed as if live. The tests exercising them report coverage for a pipeline the CLI never runs, which overstates how well the real path is tested. +- evidence: + ```go + // processService processes a single service and returns recommendations and results. + // Used by legacy callers; new code should use fetchAllRecs + executePurchasePipeline. + func processService(ctx context.Context, awsCfg aws.Config, ... + ``` +- suggested fix: Delete all three along with the tests that exist only to drive them, or move whatever behaviour is still worth pinning onto `fetchAndFilterRegionRecs` and `executePurchasePipeline`. +- verdict: CONFIRMED — processService's only callers are cmd/multi_service_coverage_test.go (3) and cmd/multi_service_test.go (6); processRegionRecommendations has exactly one caller, processService itself (cmd/multi_service.go:667); applyCommonCoverage has exactly one, cmd/multi_service_helpers_test.go:277. The --max-instances refusal at cmd/multi_service_helpers.go:412 therefore sits on a test-only path. +- issue: (pending cross-reference) + +### A11-020 The user role filter and the bulk-group prompt are dead code +- category: dead-code +- severity: low +- location: frontend/src/users/filters.ts:35 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `index.html` contains no `#user-role-filter`, no `#user-group-filter` and no `#bulk-group-btn` (all three grep to zero occurrences). The role-filter state (`roleFilter`, `setRoleFilter`, the `case 'role'` branch in `handleFilterChange`, the `applyFilters` role block) and the `#bulk-group-btn` handler are therefore unreachable. The dead role branch also encodes a wrong rule if it is ever revived: `roleFilter !== 'admin' && isAdminUser` excludes admins for *any* non-admin filter value, so a "readonly" selection would return every non-admin rather than read-only users. The dead bulk handler is the only remaining caller of `prompt()` in the module, and `bulkAddToGroup` still uses a native `confirm()` (userActions.ts:200) where every sibling path uses `confirmDialog`. +- evidence: + ```typescript + // users/filters.ts:35-40 + if (roleFilter) { + const isAdminUser = Array.isArray(user.groups) && user.groups.includes(ADMINISTRATORS_GROUP_ID); + if (roleFilter === 'admin' && !isAdminUser) return false; + if (roleFilter !== 'admin' && isAdminUser) return false; + } + ``` +- suggested fix: delete the role-filter state and its handler branches along with the `#bulk-group-btn` block, and move `bulkAddToGroup`'s `confirm()` onto `confirmDialog`. +- verdict: CONFIRMED — `user-role-filter`, `user-group-filter` and `bulk-group-btn` appear only in frontend/src/users/filters.ts and users/handlers.ts, never in index.html, so the role state, the `case 'role'` branch, the `applyFilters` block at filters.ts:35-39 and the `prompt()` handler at handlers.ts:76-85 are all unreachable, and the surviving bulk path is the `#bulk-group-select` dropdown at handlers.ts:88; `bulkAddToGroup` does still call native `confirm()` (users/userActions.ts:200). The dead role rule is wrong as written, as claimed. +- issue: (pending cross-reference) + +### A12-068 `MODE_VALUES` is dead, kept alive by a `void` with an inaccurate comment +- category: dead-code +- severity: low +- location: frontend/src/riexchange.ts:75 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The comment claims `MODE_VALUES` is used in `saveAutomationSettings`, but that function reads `modeInput.value` directly (line 2046) and there is no other reference. The `void` statement exists only to defeat the unused-variable lint, so a reader looking for the label-to-value inversion follows the comment to code that does not exist. +- evidence: + ```ts + // Suppress unused variable warning — MODE_VALUES is used in saveAutomationSettings + void MODE_VALUES; + ``` +- suggested fix: Delete `MODE_VALUES` and the `void` statement. +- verdict: CONFIRMED — MODE_VALUES is referenced only by its own definition and the `void` statement (frontend/src/riexchange.ts:71, :76), and saveAutomationSettings reads modeInput.value directly (frontend/src/riexchange.ts:2046), so the comment points at code that does not exist. +- issue: (pending cross-reference) + +### A12-072 The Archera offer modal's `'plan'` context is unreachable +- category: dead-code +- severity: low +- location: frontend/src/archera.ts:47 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ArcheraContext` declares `'purchase' | 'plan'` and `openArcheraOfferModal` branches on it for the headline, but the only two call sites (`app.ts:456` and `app.ts:645`) both pass `'purchase'`. The plan-creation surface named in the module docstring never opens the modal, so the `'plan'` headline is unreachable and a test asserts a string no user can see. +- evidence: + ```ts + export type ArcheraContext = 'purchase' | 'plan'; + // ... + title.textContent = context === 'plan' + ? 'Insure this plan with Archera?' + : 'Insure your commitments with Archera?'; + ``` +- suggested fix: Either wire the plan-creation path to call `openArcheraOfferModal('plan')`, or drop the parameter and the branch. +- verdict: CONFIRMED — ArcheraContext declares both values and the headline branches on them (frontend/src/archera.ts:47, :114) but both production call sites pass 'purchase' (frontend/src/app.ts:456, :645) and a test asserts the unreachable headline (frontend/src/__tests__/archera.test.ts:228); the finding's claim that the module docstring names a plan surface is the one inaccurate detail. +- issue: (pending cross-reference) + +### A13-020 The ARM template declares a `roleAssignmentGuidPrefix` parameter nothing reads +- category: dead-code +- severity: low +- location: arm/CUDly-CrossSubscription/template.json:17 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The parameter defaults to `[newGuid()]` and is documented as the "Base GUID used to derive deterministic role assignment names", but no `variables` or `resources` expression references `parameters('roleAssignmentGuidPrefix')`; all four names come from `guid(...)` over the principal and subscription. A customer who sets it to pin assignment names sees no effect, and `newGuid()` in a default value is itself a non-deterministic expression that would make redeployment non-idempotent if the parameter were ever wired up. +- evidence: + ```json + "roleAssignmentGuidPrefix": { + "type": "string", + "defaultValue": "[newGuid()]", + "metadata": { + "description": "Base GUID used to derive deterministic role assignment names. Leave at default to auto-generate." + } + } + ``` +- suggested fix: Delete the parameter. +- verdict: CONFIRMED — `/usr/bin/grep -n "parameters(" arm/CUDly-CrossSubscription/template.json` returns six references, all to `servicePrincipalObjectId` (lines 73, 79, 88, 91, 100, 103); `roleAssignmentGuidPrefix` is declared at :17-23 and read nowhere, and the four generated names come from `guid(...)` over the principal, a literal and `subscription().subscriptionId` (the custom-role name at :32 plus the three assignment names). Every other repo hit is a scripts/testdata/role-parity fixture copied from this template. +- issue: (pending cross-reference) + +### A13b-013 `organizations:DescribeOrganization` is granted but never called +- category: dead-code +- severity: low +- location: iac/federation/aws-cross-account/cloudformation/template.yaml:138 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `DescribeOrganization` has zero call sites in the Go tree (`grep -rn "DescribeOrganization" --include='*.go'` returns nothing at this commit), yet it is granted by the `EnableOrgDiscovery` statement here, by its Terraform twin at `iac/federation/aws-cross-account/terraform/main.tf:157`, and by the hub Lambda's `AccountDiscovery` statement in `cloudformation/stacks/CUDly/template.yaml`. It reads the organization's management-account ID, master email and feature set. A customer who enables org discovery grants CUDly organization metadata it does not use, widening the role past its call set for no benefit. +- evidence: + ```yaml + - Sid: OrganizationsDiscovery + Effect: Allow + Action: + - organizations:ListAccounts + - organizations:DescribeOrganization + Resource: "*" + ``` +- suggested fix: Drop `organizations:DescribeOrganization` from the three statements, keeping `organizations:ListAccounts` which `providers/aws/provider.go` and `internal/accounts/org_discovery.go` do call. +- verdict: CONFIRMED — `/usr/bin/grep -rn "DescribeOrganization" --include='*.go' .` returns nothing at this commit (exit 1), and the only non-Go hits are IaC: `iac/federation/aws-cross-account/cloudformation/template.yaml:138`, `iac/federation/aws-cross-account/terraform/main.tf:157`, `cloudformation/stacks/CUDly/template.yaml`, plus the deploy-side policies under `terraform/`; `organizations:ListAccounts` by contrast is genuinely used via the paginator at `providers/aws/provider.go:283-286`. +- issue: (pending cross-reference) + +### A13c-020 `aws_iam_policy.secret_read` is created and exported but never attached to any role +- category: dead-code +- severity: low +- location: terraform/modules/secrets/aws/main.tf:443 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: a grep for `secret_read_policy` across `terraform/**.tf` finds only the + module output and its pass-through at `environments/aws/outputs.tf:158` — no + `aws_iam_role_policy_attachment` anywhere. Both compute modules write their own inline + secrets policies instead. The policy is a permanent, unused IAM object that reads as the + authoritative secret-access grant, and its `data.aws_iam_policy_document.secret_read` resource + list (main.tf:431-438) already omits the credential-encryption-key ARN, so anyone who did attach + it would get a subtly incomplete grant. +- evidence: + ```hcl + resource "aws_iam_policy" "secret_read" { + name_prefix = "${var.stack_name}-secret-read-" + description = "Allow reading secrets for ${var.stack_name}" + policy = data.aws_iam_policy_document.secret_read.json + tags = var.tags + } + ``` +- suggested fix: delete the policy, the data source and both outputs, or attach it and delete the + duplicated inline policies in the Lambda and Fargate modules. +- verdict: CONFIRMED — a repo-wide grep for `secret_read` matches only the definition + (secrets/aws/main.tf:423, :443), the two module outputs (outputs.tf:67, :72) and the env + pass-through (environments/aws/outputs.tf:158); no attachment exists. The credential-encryption + key is its own resource (main.tf:209), not one of `aws_secretsmanager_secret.additional`, so the + document at :431-438 does omit it. Note the unattached state is deliberate and load-bearing: + environments/aws/ci-cd-permissions/policy_iam.tf:121 cites it to justify the + `iam:AttachRolePolicy` allowlist, so a fix must update that allowlist too. +- issue: (pending cross-reference) + +### A13c-025 `modules/registry/azure` is unreferenced dead code, and its safer defaults are what the environment overrides +- category: dead-code +- severity: low +- location: terraform/modules/registry/azure/main.tf:1 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `grep -rn 'modules/registry'` over `terraform/**.tf` matches only + `environments/gcp/registry.tf` and `environments/aws/registry.tf`; the Azure environment + declares `azurerm_container_registry.main` inline instead. The unused module carries the + Premium-SKU policy blocks, the `acr purge` cleanup task, the AcrPull assignment and + `enable_admin_user = false` — every one of which the live inline resource lacks, including the + admin-user default that A13c-005 is about. A reader auditing "the ACR module" audits the wrong + file. +- evidence: + ```hcl + resource "azurerm_container_registry" "main" { + name = var.acr_name + sku = var.sku + admin_enabled = var.enable_admin_user + quarantine_policy_enabled = var.sku == "Premium" + } + ``` +- suggested fix: consume the module from `environments/azure/registry.tf` (or delete it), so + there is one Azure registry definition. +- verdict: CONFIRMED — `/usr/bin/grep -rn 'modules/registry' terraform/` matches only + environments/gcp/registry.tf:6 and environments/aws/registry.tf:6; the Azure environment declares + `azurerm_container_registry.main` inline at environments/azure/registry.tf:5. The unused module + does carry every element claimed: `enable_admin_user` defaulting false + (registry/azure/variables.tf:37), the Premium policy blocks (main.tf:14-28), the `acr purge` + task (:34) and its own AcrPull assignment (:67-74). +- issue: (pending cross-reference) + +### A14-018 `deploy_to=aws-only` never deploys Fargate; the Fargate caller job is unreachable +- category: dead-code +- severity: medium +- location: .github/workflows/deploy-all.yml:127 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `deploy-aws-fargate` is written as the literal string `false` for every value of `deploy_to`, so the `deploy-aws-fargate` job's `if:` at line 185 is never true. An operator selecting `aws-only` or `all` for a disaster-recovery deploy gets Lambda only, while the aggregate summary prints "⏭️ AWS Fargate: Skipped" as if that were a configured choice. The README documents `all` as "Deploy to AWS, GCP, and Azure" and lists Fargate as job 3 in the fan-out (README.md:247, 263). +- evidence: + ```bash + # AWS Fargate (optional, can be enabled separately) + echo "deploy-aws-fargate=false" >> "$GITHUB_OUTPUT" + ``` +- suggested fix: Add an explicit `aws-fargate` value to the `deploy_to` choice list and drive the output from it, or delete the `deploy-aws-fargate` caller job and its README entry. +- verdict: CONFIRMED — deploy-all.yml:127 writes the literal `deploy-aws-fargate=false` for every `deploy_to` value, so the caller job's `if:` at line 185 can never be true. +- severity-adjusted: low — only the caller job is dead: deploy-aws-fargate.yml:34 remains directly dispatchable and README.md:262 already documents `aws-only` as "AWS Lambda only" with Fargate marked optional, so no deployment capability is actually lost. +- issue: (pending cross-reference) + +### A14-040 Unreachable duplicate range validation in the sanity CLI +- category: dead-code +- severity: low +- location: ci_cd_sanity_tests/cmd/sanity/main.go:32 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `requireInt32Range` on line 30 already calls `os.Exit(2)` for any value outside `[1, MaxInt32]`, so the identical check on lines 32-35 can never be reached. It is the only consumer of the `math` import, and it duplicates the bound in two spellings (`1<<31-1` and `math.MaxInt32`) that must be kept in sync by hand. +- evidence: + ```go + requireInt32Range("--max-list", *maxList) + + if *maxList < 1 || *maxList > math.MaxInt32 { + fmt.Fprintf(os.Stderr, "ERROR: --max-list must be between 1 and %d, got %d\n", math.MaxInt32, *maxList) + os.Exit(2) + } + ``` +- suggested fix: Delete lines 32-35 and the `math` import, and use `math.MaxInt32` inside `requireInt32Range` in place of `1<<31-1`. +- verdict: CONFIRMED — `requireInt32Range("--max-list", *maxList)` (ci_cd_sanity_tests/cmd/sanity/main.go:30) already exits 2 on `n < 1 || n > (1<<31-1)` (main.go:16), the identical bound as the block at main.go:32-35, which is `math`'s only non-import consumer. +- issue: (pending cross-reference) + +### Category: over-engineering + +3 findings: 3 low. + +### A03-020 Post-fetch constant-time compares guard nothing the SQL equality did not already leak +- category: over-engineering +- severity: low +- location: internal/auth/service.go:334 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The row is fetched by `WHERE token = $1` (and `key_hash = $1`, `password_reset_token = $1`); whatever timing the database comparison leaks has already leaked before `subtle.ConstantTimeCompare` runs, and the compare can only be false if the store returned a row for a different key, which is a store bug, not an attack. The comments at service.go:331-333, service_apikeys.go:300-302, service_password.go:423 and :457 claim the check "closes the SQL-equality timing oracle"; it does not, and a reader relying on that claim will not look for a real mitigation. Four sites, plus two tests that pin the theatre (service_test.go:194-220, service_password_test.go:624-648). +- evidence: + ```go + // Constant-time comparison after fetch to close the SQL-equality timing oracle. + // GetSession does `WHERE token = $1` which is not constant-time in PostgreSQL; + // align with ValidateUserAPIKey / validateResetToken (issue #392 PR #837). + if subtle.ConstantTimeCompare([]byte(session.Token), []byte(hashedToken)) != 1 { + return nil, fmt.Errorf("session not found") + } + ``` +- suggested fix: Remove the four compares and their tests, or replace the comment with the true statement (the lookup key is a SHA-256 of a 256-bit random token, so equality timing reveals nothing usable). +- verdict: CONFIRMED — every compared value is the same hash the row was just selected by: hashSessionToken is SHA-256 of a 32-byte random token (service_helpers.go:51-54, :121-127) and GetSession selects `WHERE token = $1` on it (store_postgres.go:670), ValidateUserAPIKey selects by the SHA-256 keyHash (service_apikeys.go:289-292) and validateResetToken by the hashed token (service_password.go:444-446), so the compares at service.go:334, service_apikeys.go:303, service_password.go:424 and :458 can only fail on a store bug and the comments' timing-oracle claim is not what the code closes. +- issue: (pending cross-reference) + +### A05-018 `collectWithService` and `dedupeRaw` are inert +- category: over-engineering +- severity: low +- location: internal/commitmentopts/probe.go:594 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `collectWithService(service, raw)` is `return collect(service, raw)` with nothing else, and its comment explains a distinction that does not exist — `collect` already takes the service name as its first argument. `dedupeRaw` (probe.go:570) dedupes on `(durationSeconds, payment)` immediately before `collect`, which dedupes on the normalized `(term, payment)`; every pair `dedupeRaw` removes would be removed again one call later, so deleting it changes no output. Both are correct and tested, and both guard nothing. +- evidence: + ```go + func collectWithService(service string, raw []rawOffer) []Combo { + return collect(service, raw) + } + ``` +- suggested fix: call `collect(serviceKey, raw)` directly at probe.go:562 and drop both helpers. +- verdict: CONFIRMED — collectWithService (probe.go:594-596) is a one-line pass-through to collect, whose own first parameter is already the service name (probe.go:105). dedupeRaw keys on `(durationSeconds, payment)` (probe.go:570-585) and collect keys on `(durationToTerm(durationSeconds), normalizePayment(payment))` (probe.go:112-124); since both normalizers are pure functions of those same two fields, every pair dedupeRaw drops maps to a key collect would drop one call later, and collect logs nothing and has no other side effect, so removing dedupeRaw cannot change the output. +- issue: (pending cross-reference) + +### A12-062 `canAccess('admin','*')` in the cancel and revoke gates is dead +- category: over-engineering +- severity: low +- location: frontend/src/history.ts:540 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `cancel-any:purchases` and `revoke-any:purchases` are not in `ADMIN_CARVED_OUTS` (`permissions.ts:173`), so `canAccess('cancel-any','purchases')` already returns true for an `admin:*` holder on both the effective-permissions path and the loading fallback. Deleting the `canAccess('admin','*') ||` term changes no outcome. Its presence here but deliberate absence from `rbacAllowsApprove` and `canRetryFailedRow` reads as an intentional asymmetry, inviting a future edit to "fix" the approve path by adding it back and reopening the #923 carve-out. +- evidence: + ```ts + if (canAccess('admin', '*') || canAccess('cancel-any', 'purchases')) return true; + // line 677 + if (canAccess('admin', '*') || canAccess('revoke-any', 'purchases')) return true; + ``` +- suggested fix: Drop the redundant term from both so all four row predicates gate on the verb alone. +- verdict: CONFIRMED — cancel-any:purchases and revoke-any:purchases are absent from ADMIN_CARVED_OUTS (frontend/src/permissions.ts:173-182), so canAccess already returns true for an admin:* holder on both the effective-permissions path and the loading fallback (frontend/src/permissions.ts:333, :350), making the leading term at frontend/src/history.ts:540 and :677 outcome-neutral. +- issue: (pending cross-reference) + +### Category: ops + +37 findings: 4 high, 15 medium, 18 low. + +### A13c-002 Lambda log group uses `name_prefix`, so Lambda's real log group is unmanaged — retention never applies and the migration alarm can never fire +- category: ops +- severity: high +- location: terraform/modules/compute/aws/lambda/main.tf:506 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Lambda writes to `/aws/lambda/` exactly. `name_prefix` makes + Terraform create `/aws/lambda/-` instead, and `aws_lambda_function.main` + carries no `logging_config` block redirecting output. Two consequences: (a) `retention_in_days` + is applied to a permanently empty group while the real group is auto-created by Lambda with + never-expire retention, so logs accumulate forever and `lambda_log_retention_days` is a no-op; + (b) `migration-alarm.tf:33` points the metric filter at `aws_cloudwatch_log_group.lambda.name`, + so even with `enable_migration_alarm = true` the filter watches the empty group and the alarm + can never leave `notBreaching`. +- evidence: + ```hcl + resource "aws_cloudwatch_log_group" "lambda" { + name_prefix = "/aws/lambda/${aws_lambda_function.main.function_name}-" + retention_in_days = var.log_retention_days + tags = var.tags + } + ``` +- suggested fix: use `name = "/aws/lambda/${aws_lambda_function.main.function_name}"` (importing + the existing auto-created group), or set `logging_config { log_group = ... }` on the function. +- verdict: CONFIRMED — Lambda writes to `/aws/lambda/${var.stack_name}-api` (function_name at + lambda/main.tf:29) and holds `AWSLambdaBasicExecutionRole` (main.tf:221) so it auto-creates that + group with never-expire retention, while `name_prefix` at main.tf:507 makes Terraform manage a + different, suffixed group; `/usr/bin/grep -rn logging_config terraform/` returns nothing, so + nothing redirects the function's output, and migration-alarm.tf:33 points the metric filter at + `aws_cloudwatch_log_group.lambda.name` — the suffixed, permanently empty group. +- issue: (pending cross-reference) + +### A14-002 The rollback workflow pins a Terraform version the modules reject +- category: ops +- severity: high +- location: .github/workflows/rollback.yml:49 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `TF_VERSION: '1.6.0'` is used by all four rollback jobs (lines 234, 338, 408, 499), each of which runs `terraform init` inside `terraform/environments/{aws,gcp,azure}`. Every one of those roots declares `required_version = ">= 1.10.0"` (aws/main.tf:5, gcp/main.tf:5, azure/main.tf:5). `terraform init` aborts with "Unsupported Terraform Core version" before the backend is even configured, so an operator invoking the emergency rollback path during an incident gets a hard failure on every cloud. rollback.yml is the sole outlier: every other workflow pins 1.10.0 or 1.10.5. +- evidence: + ```yaml + env: + TF_VERSION: '1.6.0' + ``` +- suggested fix: Change to `'1.10.0'` to match the other deploy/destroy workflows. +- verdict: CONFIRMED — rollback.yml:49 sets `TF_VERSION: '1.6.0'`, consumed by setup-terraform at lines 234/338/408/499, and all four jobs `terraform init` in roots declaring `required_version = ">= 1.10.0"` (terraform/environments/{aws,gcp,azure}/main.tf:5); every other workflow pins 1.10.0 or 1.10.5. +- issue: (pending cross-reference) + +### A14-003 database-migration.yml runs `terraform init` with no backend config, so it can never read a DB endpoint +- category: ops +- severity: high +- location: .github/workflows/database-migration.yml:296 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: All three `Get database endpoint from Terraform` steps (296, 410, 502) run a bare `terraform init`. `terraform/environments/{aws,gcp,azure}/backend.tf` each declare an empty partial backend block ("Configuration provided via -backend-config flag"), so `init` fails on the missing required `bucket`/`storage_account_name` argument. `TF_BACKEND_AWS/GCP/AZURE` is never referenced anywhere in this file. Even if init somehow succeeded it would attach to a single default state object with no `github-` key, so a `prod` migration would read a `dev` endpoint. The result is that the only documented migration path (README.md:293-358) cannot run at all, and no guard covers it: `scripts/test-aws-tfstate-platform-key.sh` only inspects jobs that apply a `compute_platform`, which these do not. +- evidence: + ```bash + cd terraform/environments/aws + terraform init + DB_ENDPOINT=$(terraform output -raw database_proxy_endpoint 2>/dev/null || echo "") + if [ -z "$DB_ENDPOINT" ]; then + echo "Failed to get database endpoint" + exit 1 + fi + ``` +- suggested fix: Write `/tmp/backend.tfbackend` from `secrets.TF_BACKEND_` plus the environment-scoped key, the way every deploy job does, and pass `-backend-config`. +- verdict: CONFIRMED — database-migration.yml:296/410/502 run a bare `terraform init` against partial backends (terraform/environments/{aws,gcp,azure}/backend.tf each declare an empty block "provided via -backend-config"), `TF_BACKEND` appears nowhere in the file, and scripts/test-aws-tfstate-platform-key.sh:4 only scans jobs applying a `compute_platform`. +- issue: (pending cross-reference) + +### A15-001 npm audit fails on the pinned commit: fast-uri 3.1.5 carries four high-severity advisories +- category: ops +- severity: high +- location: frontend/package-lock.json:5754 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The committed lockfile pins `fast-uri@3.1.5`, which is inside the advisory range `>= 3.1.2, < 3.1.6`. `.github/workflows/ci.yml:663` runs `npm audit --audit-level=high`; run against this lockfile it exits 1, so the Security job's npm audit step is red on this commit. Measured locally after `npm ci --ignore-scripts`: `npm audit --audit-level=high` -> exit 1, `1 high severity vulnerability`. +- evidence: + ```json + "node_modules/fast-uri": { + "version": "3.1.5", + "dev": true, + ``` + Four GHSAs: GHSA-5jgf-p345-68v8, GHSA-f65p-4m7j-42xc, GHSA-fph4-wmhf-6fwf, GHSA-jqff-g426-hqxp (host confusion / SSRF via IDN, IPv6 and percent-decoding normalization). First patched version 3.1.6 per the GitHub advisory API. + Reachability: `npm ls fast-uri --omit=dev` returns empty. It is reached only through devDependencies (`babel-loader` -> `schema-utils` -> `ajv`, and `serve` -> `ajv`), and production `dependencies` are only `@types/qrcode`, `chart.js`, `qrcode`. It does not enter the shipped webpack bundle, so no end user of the deployed frontend is exposed. The live impact is the red CI gate, not runtime SSRF. +- suggested fix: Refresh the lockfile so `fast-uri` resolves to `3.1.6` (a patch bump inside the existing semver range, no dependency change needed). +- verdict: CONFIRMED — reproduced in the pinned worktree: `npm audit --audit-level=high` in `frontend/` exits 1 with `1 high severity vulnerability` and all four GHSAs against `fast-uri@3.1.5` (frontend/package-lock.json:5755, `"dev": true`), while `npm ls fast-uri --omit=dev` returns empty and frontend/package.json:49-53 lists only `@types/qrcode`, `chart.js`, `qrcode` in `dependencies`, so the impact is the red gate at .github/workflows/ci.yml:663 and not runtime SSRF; the only correction is the advisory range, which npm reports as `3.0.0 - 3.1.5` (per-GHSA `>=3.1.3 <3.1.6`) rather than the finding's `>= 3.1.2, < 3.1.6`. +- issue: (pending cross-reference) + +### A07-017 Three Cost Explorer pagination loops have no page cap +- category: ops +- severity: medium +- location: providers/aws/recommendations/coverage.go:266 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `fetchCoveragePaged`, `GetRIUtilization` (utilization.go:69) and `fetchDailyCoverage` (usage_history.go:98) loop on `NextPageToken` until the token is empty, with no ceiling. Every other CE loop in the package guards with `maxRecommendationPages`, `maxSPCoveragePages` or `maxOnDemandSeriesPages` and the comments on those constants state the reason: a token loop or API misbehaviour otherwise spins while billing per call. A repeated token from CE turns any of these three into an unbounded billed loop that only ends when the context deadline fires. +- evidence: + ```go + var token *string + for { + if err := ctx.Err(); err != nil { + return fmt.Errorf("coverage: pagination cancelled: %w", err) + } + input.NextPageToken = token + result, err := c.fetchCoveragePage(ctx, input) + ``` +- suggested fix: add the same page-index cap and diagnostic error the sibling loops use to all three. +- verdict: CONFIRMED — `fetchCoveragePaged` (providers/aws/recommendations/coverage.go:266-288), `GetRIUtilization` (utilization.go:68-89) and `fetchDailyCoverage` (usage_history.go:98-111) all loop on the token with no page index, while the siblings cap at client.go:181 (`maxRecommendationPages`), sp_coverage.go:334 (`maxSPCoveragePages`) and ondemand_series.go:130 (`maxOnDemandSeriesPages`); `fetchDailyCoverage` additionally has no `ctx.Err()` check at the top of its loop. +- issue: (pending cross-reference) + +### A08b-003 Hardened transport drops proxy support, connection reuse limits and HTTP/2 +- category: ops +- severity: medium +- location: pkg/httpclient/httpclient.go:60 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the transport is constructed with only `DialContext` and `TLSHandshakeTimeout`. `Proxy` is nil, so `HTTPS_PROXY` is ignored and every Azure/pricing call fails in a network that requires an egress proxy. `IdleConnTimeout` is zero, so idle keep-alive connections are never reaped, and `MaxIdleConns`/`MaxIdleConnsPerHost` default to 2 per host, throttling the multi-page pricing walks. Setting a custom `DialContext` also disables the automatic HTTP/2 upgrade that `http.DefaultTransport` gets via `ForceAttemptHTTP2`. +- evidence: + ```go + transport := &http.Transport{ + DialContext: dialer.DialContext, + TLSHandshakeTimeout: tlsHandshakeTimeout, + } + ``` +- suggested fix: start from `http.DefaultTransport.(*http.Transport).Clone()`, then override `DialContext`, so proxy, idle-pool and HTTP/2 defaults are preserved. +- verdict: CONFIRMED — the transport at pkg/httpclient/httpclient.go:60-63 sets only `DialContext` and `TLSHandshakeTimeout`, leaving `Proxy` nil, `IdleConnTimeout` zero and `ForceAttemptHTTP2` false; one correction, zero `MaxIdleConns` means unlimited, it is `MaxIdleConnsPerHost` alone that falls back to 2. +- issue: (pending cross-reference) + +### A09-003 Hardened transport silently disables proxy support, HTTP/2 and idle-connection expiry +- category: ops +- severity: medium +- location: pkg/httpclient/httpclient.go:60 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: a hand-built `&http.Transport{}` has `Proxy: nil`, unlike `http.DefaultTransport` which uses `http.ProxyFromEnvironment`. In a VPC or Lambda deployment whose only egress path is an `HTTPS_PROXY`, every call through this client dials the origin directly and times out after 10s at the dialer, with no indication that the proxy was ignored. `ForceAttemptHTTP2` also defaults false here, and the absent `IdleConnTimeout`/`MaxIdleConns` means pooled connections are never reaped in a long-lived process. +- evidence: + ```go + transport := &http.Transport{ + DialContext: dialer.DialContext, + TLSHandshakeTimeout: tlsHandshakeTimeout, + } + ``` +- suggested fix: start from `http.DefaultTransport.(*http.Transport).Clone()` and override only `DialContext` and `TLSHandshakeTimeout`, so the proxy, HTTP/2 and idle-pool defaults survive. +- verdict: CONFIRMED — `New` builds a bare `&http.Transport{}` with only `DialContext` and `TLSHandshakeTimeout` set (pkg/httpclient/httpclient.go:60-63), so `Proxy`, `ForceAttemptHTTP2`, `IdleConnTimeout` and `MaxIdleConns` all take their zero values rather than `http.DefaultTransport`'s. +- issue: (pending cross-reference) + +### A13-003 The legacy cross-account CloudFormation stack is missing three grants and is excluded from the parity guard +- category: ops +- severity: medium +- location: cloudformation/stacks/CUDly-CrossAccount/template.yaml:159 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A customer onboarded with this template gets a role lacking `savingsplans:DescribeSavingsPlansOfferings`, `ec2:GetReservedInstancesExchangeQuote` and `ec2:AcceptReservedInstancesExchangeQuote`, all three of which the other four federation flavors grant. Savings Plans purchase fails at offering lookup and RI exchange fails at quote time, in that account only, with an AccessDenied that looks like a customer misconfiguration. `scripts/check-aws-iam-parity.sh` compares seven files (lines 120-152) and this is not one of them, so the drift is invisible in CI. +- evidence: + ```yaml + - savingsplans:DescribeSavingsPlans + - savingsplans:CreateSavingsPlan + - savingsplans:DescribeSavingsPlansOfferingRates + ``` +- suggested fix: Either add this template to comparison 3 in `scripts/check-aws-iam-parity.sh` and bring its action list into parity, or delete it if `iac/federation/aws-cross-account/` has superseded it. +- verdict: CONFIRMED — a per-file grep count for the three actions returns 1 in all six files the parity script compares (iac/federation/aws-{cross-account,target}/{cloudformation,terraform} and internal/iacfiles/templates/aws-{cross-account,wif}-cli.sh.tmpl) and 0 in cloudformation/stacks/CUDly-CrossAccount/template.yaml, whose SavingsPlans statement stops at `DescribeSavingsPlansOfferingRates` (lines 152-158) and whose EC2 statement has no exchange verbs (lines 97-107). The script's three comparisons (scripts/check-aws-iam-parity.sh:120-152) never name this file. +- issue: (pending cross-reference) + +### A13-006 The production Azure Key Vault is default-deny with an empty allowlist, so no principal can reach it +- category: ops +- severity: high +- location: terraform/environments/azure/github-prod.tfvars:52 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `github-prod.tfvars` sets neither `key_vault_default_network_acl_action` (default `"Deny"`, variables.tf:167) nor `allowed_ip_addresses` (default `[]`), and `create_private_subnet` defaults to `false`, so `allowed_subnet_ids` is `[]` too. The rendered `network_acls` block is `default_action = "Deny"`, `ip_rules = []`, `virtual_network_subnet_ids = []`, `bypass = "AzureServices"`. A GitHub Actions runner is not an Azure service, so every `azurerm_key_vault_secret` write in the prod apply is refused and the deploy fails; the Container App cannot read the vault at runtime either. The staging file's own comment (github-staging.tfvars:56) says production "should keep the default Deny and supply allowed_ip_addresses" and prod supplies none. +- evidence: + ```hcl + # github-prod.tfvars — the whole Key Vault block + key_vault_sku = "standard" + soft_delete_retention_days = 90 + purge_protection_enabled = true + ``` +- suggested fix: Add `allowed_ip_addresses` for the deploy runner egress to `github-prod.tfvars`, or set `create_private_subnet = true` there so the private subnet's `Microsoft.KeyVault` service endpoint is allowlisted. +- verdict: CONFIRMED — the whole of terraform/environments/azure/github-prod.tfvars sets only `key_vault_sku`, `soft_delete_retention_days` and `purge_protection_enabled`; it never sets `key_vault_default_network_acl_action` (default `"Deny"`, variables.tf:167-175), `allowed_ip_addresses` (default `[]`, variables.tf:161-165) or `create_private_subnet` (default `false`, variables.tf:96-100), so secrets.tf:37-38 renders `ip_rules = []` and `virtual_network_subnet_ids = []` into terraform/modules/secrets/azure/main.tf:46-51. github-staging.tfvars:56-58 sets `"Allow"` with the comment that prod "should keep the default Deny and supply allowed_ip_addresses", and .github/workflows/deploy-azure.yml:207 selects the file by environment name, so the prod apply really does run against a default-deny vault with an empty allowlist. +- severity-adjusted: medium — the failure is a hard, immediately visible apply failure on a path that has evidently not been exercised, not a silent security or money defect. +- issue: (pending cross-reference) + +### A13-011 `timestamp()` in the image tag forces a rebuild and push on every apply, making the content-hash triggers dead +- category: ops +- severity: medium +- location: terraform/modules/build/main.tf:39 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `local.timestamp` is `formatdate(..., timestamp())`, which is re-evaluated on every plan, so `terraform_data.image_tag.output` and hence `local.image_tag` change on every run. `image_tag` is itself one of `terraform_data.docker_build.triggers_replace`, so the build resource is replaced on every apply regardless of the four file-hash triggers above it. Those hashes (`go_mod`, `go_sum`, `dockerfile`, `cmd_files`, `pkg_files`) can never prevent a rebuild, so every apply pays a full multi-stage Docker build and pushes a new image layer set to ECR/ACR/Artifact Registry even when nothing changed. The hashes also cover only `cmd/` and `pkg/`, not `internal/` or `frontend/`, so they would miss most real source changes if they were load-bearing. +- evidence: + ```hcl + timestamp = var.skip_docker_build ? "skip" : formatdate("YYYYMMDDhhmmss", timestamp()) + ... + triggers_replace = { + go_mod = fileexists(...) ? filemd5(...) : "none" + image_tag = local.image_tag + platform = local.effective_platform + } + ``` +- suggested fix: Derive the tag from the content hash (the existing `sha256(...)` of the source files plus `git_commit`) instead of `timestamp()`, and drop the now-redundant per-file triggers or extend them to `internal/` and `frontend/`. +- verdict: CONFIRMED — terraform/modules/build/main.tf:39 sets `timestamp = formatdate("YYYYMMDDhhmmss", timestamp())`, feeding `terraform_data.image_tag.input` at line 24 and `local.image_tag` at line 43, which is itself a `triggers_replace` key at line 46 alongside the five file hashes at lines 57-62. The escape hatch is not used: `var.custom_image_tag` is referenced only at main.tf:24 and set by no environment (`/usr/bin/grep -rn custom_image_tag terraform .github` finds only the declaration, the default `""`, and a comment in deploy-aws-lambda.yml:185), so the timestamp branch is always taken and the hashes can never gate a rebuild. The coverage gap is real too: the fileset globs cover `cmd/` and `pkg/` only, not `internal/` or `frontend/`. +- issue: (pending cross-reference) + +### A13-012 `enable_migration_alarm` promises a bootstrap grant that the bootstrap policy does not contain +- category: ops +- severity: medium +- location: terraform/modules/compute/aws/lambda/variables.tf:19 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The variable description tells the operator that `logs:PutMetricFilter` "is granted via the ci-cd-permissions bootstrap" and to "set true only after re-applying the bootstrap". The `CloudWatchLogs` statement in `policy_compute.tf:33-47` grants only `CreateLogGroup`, `DeleteLogGroup`, `ListTagsForResource`, `PutRetentionPolicy`, `TagResource` and `UntagResource`; no metric-filter action exists anywhere under `ci-cd-permissions/`. An operator who re-applies the bootstrap and then flips the flag gets an AccessDenied on `aws_cloudwatch_log_metric_filter.migration_failed`, which fails the apply and blocks every deploy, the exact outcome the gate was added to prevent. +- evidence: + ```hcl + description = "... requires logs:PutMetricFilter on the deploy SA, which is granted via the ci-cd-permissions bootstrap (root CLAUDE.md CI/CD IAM split). ... set true only after re-applying the bootstrap so the deploy role can manage the filter." + ``` +- suggested fix: Add `logs:PutMetricFilter`, `logs:DeleteMetricFilter` and `logs:DescribeMetricFilters` scoped to the existing log-group ARNs in `policy_compute.tf`, or reword the description to say the grant does not exist yet. +- verdict: CONFIRMED — `/usr/bin/grep -rn 'PutMetricFilter|DescribeMetricFilters|DeleteMetricFilter' terraform cloudformation iac` matches only the two prose mentions (lambda/variables.tf:19 and lambda/migration-alarm.tf:25); the `CloudWatchLogs` statement at policy_compute.tf:32-48 stops at `CreateLogGroup`, `DeleteLogGroup`, `ListTagsForResource`, `PutRetentionPolicy`, `TagResource`, `UntagResource`, with `logs:DescribeLogGroups` split out at :49-53. `logs:*` at policy_boundary.tf:173 is the workload permissions boundary, which caps rather than grants and does not apply to the deploy role. So an operator who re-applies the bootstrap and flips `enable_migration_alarm = true` still 403s on `aws_cloudwatch_log_metric_filter.migration_failed` (migration-alarm.tf:24-45). +- issue: (pending cross-reference) + +### A13c-007 Azure Logic App recurrence triggers embed `timestamp()`, producing a diff on every plan +- category: ops +- severity: medium +- location: terraform/modules/compute/azure/container-apps/scheduled-tasks.tf:143 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `timestamp()` is evaluated at every plan and apply, so `start_time` changes + on each run and all three recurrence triggers (`daily` at :143, `ri_exchange` at :263, + `cleanup_daily` at :358) show a permanent update. `terraform plan -detailed-exitcode` can never + return "no changes" on an Azure deployment, which removes drift detection as a signal and + re-anchors the schedule window on every deploy. +- evidence: + ```hcl + start_time = "${formatdate("YYYY-MM-DD", timestamp())}T${format("%02s", local.schedule_hour)}:00:00Z" + ``` +- suggested fix: drop `start_time` (Logic Apps starts the recurrence at creation) or make it a + fixed input date, and add `lifecycle { ignore_changes = [start_time] }` for existing state. +- verdict: CONFIRMED — all three recurrence triggers embed `timestamp()` in `start_time` + (scheduled-tasks.tf:143, 263, 358) and the file's only `lifecycle` block is at line 55, on an + unrelated resource, so no `ignore_changes` suppresses the diff. +- issue: (pending cross-reference) + +### A13c-013 Azure PostgreSQL Flexible Server has no deletion protection of any kind, while the AWS and GCP twins do +- category: ops +- severity: medium +- location: terraform/modules/database/azure/main.tf:19 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `aws_db_instance.main` carries `deletion_protection` (default true) plus + `final_snapshot_identifier`; `google_sql_database_instance.main` carries + `deletion_protection`. `azurerm_postgresql_flexible_server` has neither an equivalent argument + set nor a `lifecycle { prevent_destroy }`, so a `terraform destroy`, a `-target` mistake, or any + ForceNew change (`administrator_login`, `delegated_subnet_id`, `zone` outside the ignore) drops + the production database with only the automatic backup as recourse. +- evidence: + ```hcl + resource "azurerm_postgresql_flexible_server" "main" { + name = "${var.app_name}-postgres" + administrator_login = var.administrator_login + backup_retention_days = var.backup_retention_days + lifecycle { + ignore_changes = [zone] + } + } + ``` +- suggested fix: add `prevent_destroy = var.deletion_protection`-equivalent guarding (a + `lifecycle { prevent_destroy = true }` in the module, or a `azurerm_management_lock` on the + server gated by an input) so the three providers share a posture. +- verdict: CONFIRMED — `azurerm_postgresql_flexible_server.main` spans database/azure/main.tf:19-76 + and its only `lifecycle` block is `ignore_changes = [zone]` at :73-75; a repo-wide grep for + `deletion_protection|prevent_destroy|management_lock` under `terraform/modules/database/` matches + only the AWS (main.tf:152, variables.tf:87 default true) and GCP (main.tf:42) modules. +- issue: (pending cross-reference) + +### A14-007 The README's GitHub Environment names do not match the ones the workflows bind to +- category: ops +- severity: medium +- location: .github/workflows/README.md:521 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The setup instructions tell the operator to create `aws-lambda-dev/staging/prod`, but `deploy-aws-lambda.yml:231` binds `environment: ${{ needs.prepare.outputs.target_environment }}`, i.e. plain `dev`/`staging`/`prod`. The rollback jobs bind `aws-lambda--rollback` / `aws-fargate--rollback` / `gcp--rollback` / `azure--rollback`, and the migration jobs bind `aws-db-` / `gcp-db-` / `azure-db-`; none of those appear in the README list. Required-reviewer and deployment-branch-policy rules configured per these instructions land on environments nothing binds to, while the environments that credentialed jobs actually use are auto-created bare. Several workflow headers state that the environment binding is the only control that can gate an unapproved production deploy or destroy. +- evidence: + ```markdown + 2. Create environments: + - `aws-lambda-dev`, `aws-lambda-staging`, `aws-lambda-prod` + - `aws-fargate-dev`, `aws-fargate-staging`, `aws-fargate-prod` + - `gcp-dev`, `gcp-staging`, `gcp-prod` + - `azure-dev`, `azure-staging`, `azure-prod` + ``` +- suggested fix: Regenerate the list from the actual `environment:` expressions across the workflows, including the `-rollback` and `-db-` families. +- verdict: CONFIRMED — README.md:521-524 lists `aws-lambda-*`/`gcp-*`/`azure-*`, but deploy-aws-lambda.yml:231,377 bind plain `dev`/`staging`/`prod`, rollback.yml:192,296,376,458 and database-migration.yml:265,376,470 bind the `-rollback` and `-db-` families, and deploy-gcp.yml and deploy-azure.yml carry no job-level `environment:` at all; only `aws-fargate-*` (deploy-aws-fargate.yml:142) matches the README. +- issue: (pending cross-reference) + +### A14-012 `make security-scan-go` swallows gosec's verdict and excludes eight rule classes +- category: ops +- severity: medium +- location: Makefile:159 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: gosec exits non-zero on findings, but the recipe line ends with `; echo "✓ Go security scan complete"`, so the recipe's status is echo's. `security-scan-go` therefore always succeeds, and `make ci` (line 216) — which chains `security-scan` — reports a clean pipeline with findings present. The `-exclude=G101,G104,G115,G204,G301,G304,G402,G505` list also silences InsecureSkipVerify-adjacent and TLS-version rules that ci.yml's authoritative gosec run (ci.yml:694) does **not** exclude, so the local `make ci` result cannot predict CI. When gosec is not installed the recipe prints a hint and still exits 0. +- evidence: + ```make + security-scan-go: + @if command -v gosec > /dev/null; then \ + gosec -fmt=json -out=gosec-report.json -exclude=G101,G104,G115,G204,G301,G304,G402,G505 ./...; \ + echo "✓ Go security scan complete: gosec-report.json"; \ + else \ + echo "gosec not installed. Install: make install-dev-tools"; \ + fi + ``` +- suggested fix: Drop the trailing `echo` so gosec's status propagates, align the exclude list with ci.yml (which uses none), and `exit 1` when the tool is missing, as the `complexity` target already does. +- verdict: CONFIRMED — Makefile:156-163 is a single `@if ...; then gosec ...; echo ...; fi`, so the recipe's exit status is the trailing echo's; ci.yml:694-705 runs gosec with no `-exclude` at all, and `ci:` at Makefile:216 chains `security-scan`. +- issue: (pending cross-reference) + +### A14-032 gcp-import-dev-state.sh hardcodes a live project ID and state bucket +- category: ops +- severity: medium +- location: scripts/gcp-import-dev-state.sh:22 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The GCP project (`serene-bazaar-666`) and the Terraform state bucket (`cudly-terraform-state-cloudprowess`) are literals in a committed script, unlike every other GCP consumer which reads `vars.GCP_PROJECT_ID` and `secrets.TF_BACKEND_GCP`. Anyone running the script against a different environment silently imports resources from the wrong project into whatever state they initialised, and the two names — which the deploy pipeline treats as configuration worth keeping in repository variables and secrets — are published in the repository. +- evidence: + ```bash + PROJECT="serene-bazaar-666" + REGION="us-central1" + SERVICE_NAME="cudly-dev" + TF_DIR="$(cd "$(dirname "$0")/.." && pwd)/terraform/environments/gcp" + BACKEND_CONFIG="bucket = \"cudly-terraform-state-cloudprowess\"\nprefix = \"github-dev\"" + ``` +- suggested fix: Take the project and bucket from required arguments or environment variables, defaulting to nothing and failing loudly when unset. +- verdict: CONFIRMED — `PROJECT="serene-bazaar-666"` and the `cudly-terraform-state-cloudprowess` bucket are unconditional literals with no argument or environment override (scripts/gcp-import-dev-state.sh:22,26), while deploy-gcp.yml:184 reads `vars.GCP_PROJECT_ID` and the workflows take the bucket from `secrets.TF_BACKEND_GCP`. +- issue: (pending cross-reference) + +### A14-034 `make deploy` points at a directory layout the repository does not have +- category: ops +- severity: medium +- location: scripts/tf-deploy.sh:62 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ENV_DIR` is `terraform/environments//`, e.g. `terraform/environments/aws/dev`, which does not exist — the repository has a single root at `terraform/environments/aws` selected by `-var-file`. The script creates the missing directory and symlinks only `main.tf` into it (line 83), so the resulting root has no `variables.tf`, `backend.tf` or sibling `.tf` files and `terraform init` at line 113 (run with no `-backend-config`) cannot resolve anything. `Makefile:102` (`deploy`) and every target in `Makefile.terraform` (deploy/plan/destroy/output/docker-skip) route through this script, so the documented local deployment path is inoperable, and it leaves a stray directory behind. `terraform/profiles/aws/` also contains only `*.tfvars.example`, so `make profile-list` reports no profiles. +- evidence: + ```bash + ENV_DIR="${PROJECT_ROOT}/terraform/environments/${PROVIDER}/${PROFILE}" + ... + mkdir -p "$ENV_DIR" + ln -sf "../../${PROVIDER}/main.tf" "${ENV_DIR}/main.tf" + ``` +- suggested fix: Point `ENV_DIR` at `terraform/environments/${PROVIDER}` and pass the profile through `-var-file`, matching the workflows; remove the directory-creation and symlink branch. +- verdict: CONFIRMED — `ENV_DIR="${PROJECT_ROOT}/terraform/environments/${PROVIDER}/${PROFILE}"` (scripts/tf-deploy.sh:62) names a path that does not exist (terraform/environments/aws is a single flat root) and terraform/profiles/aws/ holds only `*.tfvars.example`, so `make deploy` (Makefile:102) exits at the profile check before reaching the mkdir and symlink branch at line 83. +- issue: (pending cross-reference) + +### A14-035 init-backend.sh writes a backend file whose state key conflicts with the workflows' +- category: ops +- severity: medium +- location: scripts/init-backend.sh:208 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The script creates `terraform/environments/aws//backend.tf` with a hardcoded `key = "/terraform.tfstate"` and bucket `cudly-terraform-state-`. Every workflow instead writes `key = "github-/terraform.tfstate"` (deploy-aws-lambda.yml:259) or `github-fargate-/…` from `secrets.TF_BACKEND_AWS`. An operator following this script initialises a second, empty state at a different key in a different bucket and then applies into it, creating a duplicate stack alongside the CI-managed one. The generated file also lands in a subdirectory that `terraform fmt -check -recursive terraform/` (ci.yml:583) will start walking as a separate root. +- evidence: + ```bash + BACKEND_CONFIG_FILE="terraform/environments/aws/${ENVIRONMENT}/backend.tf" + ... + bucket = "${BUCKET_NAME}" + key = "${ENVIRONMENT}/terraform.tfstate" + ``` +- suggested fix: Emit a `.tfbackend` file matching the `github-/terraform.tfstate` key the workflows use, into `terraform/environments/aws/backends/`, rather than a `backend.tf` in a new root. +- verdict: CONFIRMED — scripts/init-backend.sh:208,220-221 writes `terraform/environments/aws//backend.tf` with bucket `cudly-terraform-state-` (line 78) and `key = "/terraform.tfstate"` against every workflow's `github-/terraform.tfstate` (deploy-aws-lambda.yml:259), and lines 248-254 instruct the operator to cd there and apply; the duplicate-stack outcome additionally requires configuration to be placed in that otherwise-empty directory. +- issue: (pending cross-reference) + +### A15-003 Go tooling scans third-party Go code vendored inside frontend/node_modules +- category: ops +- severity: medium +- location: .pre-commit-config.yaml:19 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The npm package `flatted` ships a Go package at `frontend/node_modules/flatted/golang/pkg/flatted/flatted.go` with no `go.mod` of its own, so Go absorbs it into the CUDly module. In `.github/workflows/pre-commit.yml` the "Install frontend deps" step (`cd frontend && npm ci`, line 238) runs before "Run pre-commit" (line 241), which runs `go vet ./...`. Verified in the worktree after `npm ci`: `go list ./...` emits `github.com/LeanerCloud/CUDly/frontend/node_modules/flatted/golang/pkg/flatted`. Go's `./...` skips only `testdata` and directories beginning with `.` or `_`, never `node_modules`. Today `go vet` passes on that package, so nothing is broken; the exposure is that a future `flatted` release with vet-unclean or vulnerable Go code reddens the repo's own lint, test and govulncheck gates on code the repo does not own and cannot fix. +- evidence: + ```yaml + - id: go-vet + name: Run go vet + entry: bash -c 'go vet ./...' + ``` +- suggested fix: Have the Go `./...` invocations enumerate real packages, for example `go vet $(go list ./... | grep -v /node_modules/)`, or move the frontend install after the pre-commit step in that workflow. +- verdict: CONFIRMED — reproduced end to end: `frontend/node_modules/flatted/golang/pkg/flatted/flatted.go` exists after `npm ci` and `go list ./...` in the worktree emits `github.com/LeanerCloud/CUDly/frontend/node_modules/flatted/golang/pkg/flatted`, the go-vet hook really is `bash -c 'go vet ./...'` (.pre-commit-config.yaml:16-21, one line above the cited :19), and "Install frontend deps" at .github/workflows/pre-commit.yml:235 does run before "Run pre-commit" at :241, so the third-party package is inside the repo's own gate. +- issue: (pending cross-reference) + +### A02-018 forgot-password has no per-IP limit and its per-email key is not normalized +- category: ops +- severity: low +- location: internal/api/handler_auth.go:267 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The only limiter is `AllowWithEmail` (10 per 5 min per exact string). One client can submit thousands of distinct addresses per minute, each costing a `GetUserByEmail` lookup and, for real users, an outbound reset email through SES, with no IP budget at all (every other credential endpoint uses `checkRateLimitStrict`). Because the key is the raw string, `Alice@x.com`, `alice@x.com` and `alice@x.com ` are separate buckets, so a single victim can receive well over 10 reset mails per window if the store's lookup is case-insensitive. +- evidence: + ```go + if h.rateLimiter != nil { + allowed, err := h.rateLimiter.AllowWithEmail(ctx, pwdReq.Email, "forgot_password") + ``` +- suggested fix: Add `checkRateLimit(ctx, req, "forgot_password_ip")` (fail-open is fine here) before the email bucket, and key the email bucket on `strings.ToLower(strings.TrimSpace(email))`. +- verdict: CONFIRMED — forgotPassword (handler_auth.go:267-276) has only AllowWithEmail, which keys on the raw string in both limiters (db_rate_limiter.go:217-219, inmemory_rate_limiter.go:143-145), and no checkRateLimit call, unlike every other credential endpoint (handler_auth.go:24,234,316,345,475); the case-variant amplification is weaker than stated because GetUserByEmail is an exact `WHERE email = $1` (internal/auth/store_postgres.go:68), so variants hit distinct lookups rather than the same victim. +- issue: (pending cross-reference) + +### A06-026 OIDC provider is created with an all-zero certificate thumbprint +- category: ops +- severity: low +- location: internal/iacfiles/templates/aws-wif-cli.sh.tmpl:57 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the generated script creates the IAM OIDC provider with a 40-zero thumbprint and no comment explaining why. For issuers whose certificate chains to a CA in AWS's trust store this value is ignored, but for a CUDly deployment fronted by a private or internal CA the thumbprint is the verification input, and a value that can never match makes every `AssumeRoleWithWebIdentity` fail with an opaque error the operator has no pointer to. It is also a hardcoded magic value in a security-relevant position, in a file that elsewhere goes to considerable length to explain its trust-policy choices. +- evidence: + ```sh + aws iam create-open-id-connect-provider $PROFILE_ARG \ + --url "${OIDC_ISSUER_URL}" \ + --client-id-list "${OIDC_AUDIENCE}" \ + --thumbprint-list "0000000000000000000000000000000000000000" >/dev/null + ``` +- suggested fix: compute the thumbprint from the issuer's TLS chain in the script, or keep the placeholder and add the one-line comment stating that AWS ignores it for trust-store issuers. +- verdict: PLAUSIBLE — the all-zero literal is there with no explanatory comment, in a file that comments its trust-policy choices at length (internal/iacfiles/templates/aws-wif-cli.sh.tmpl:52-58 vs 15-41), so the hardcoded-magic-value half is a verified fact; the operational failure needs the CUDly issuer to be fronted by a CA outside AWS's trust store, which I cannot establish from source since `CUDLY_ISSUER_URL` resolves to the deployment's own Lambda Function URL / Cloud Run domain (internal/api/handler_federation.go:175-183). +- issue: (pending cross-reference) + +### A13-018 `acm:RemoveTagsFromCertificate` is missing, the same gap the file documents for `ec2:DeleteTags` +- category: ops +- severity: low +- location: terraform/environments/aws/ci-cd-permissions/policy_networking.tf:140 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `terraform/environments/aws/acm.tf:14` tags the certificate with `merge(local.common_tags, ...)`. The AWS provider's tag update path removes dropped keys before adding new ones, so dropping any key from `common_tags` or `default_tags` on an existing certificate calls `acm:RemoveTagsFromCertificate` and fails the apply. This is the identical failure mode the same file documents at lines 53-58 as the reason `ec2:DeleteTags` had to be added, applied to a resource that also carries tags. +- evidence: + ```hcl + Action = [ + "acm:AddTagsToCertificate", + "acm:DeleteCertificate", + "acm:DescribeCertificate", + "acm:GetCertificate", + "acm:ListTagsForCertificate", + "acm:RequestCertificate", + ] + ``` +- suggested fix: Add `acm:RemoveTagsFromCertificate` to the ACM statement. +- verdict: CONFIRMED — the `ACM` statement at policy_networking.tf:139-150 lists six actions and no `RemoveTagsFromCertificate`, while the same file at :52-58 documents the identical provider behaviour as the reason `ec2:DeleteTags` had to be added. terraform/environments/aws/acm.tf:14-16 tags the certificate from `merge(local.common_tags, ...)`, so a dropped key hits the untagged verb. The failure needs two conditions the finding does not state: the certificate exists only when `frontend_domain_names` and `subdomain_zone_name` are both set (acm.tf:5), and a tag key must actually be removed. +- issue: (pending cross-reference) + +### A13-021 Two compose images float on tags while every sibling image is digest-pinned with a written rationale +- category: ops +- severity: low +- location: docker-compose.yml:96 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `docker-compose.test.yml:7-10` explains that a floating tag "lets the database change under an unchanged repo" and pins Postgres by digest in both files. `nginx:alpine` and `dpage/pgadmin4:9.2` in `docker-compose.yml` are not pinned, so a local dev stack silently picks up a new nginx or pgAdmin build, including one that changes the reverse-proxy behaviour `scripts/nginx.conf` depends on. `nginx:alpine` in particular tracks a moving major. +- evidence: + ```yaml + frontend: + image: nginx:alpine + ... + pgadmin: + image: dpage/pgadmin4:9.2 + ``` +- suggested fix: Pin both by `@sha256:` digest with the same refresh comment the Postgres entries carry. +- verdict: CONFIRMED — docker-compose.yml pins postgres by digest at line 6 but leaves `nginx:alpine` (line 96) and `dpage/pgadmin4:9.2` (line 110) on tags, while docker-compose.test.yml:7-11 states the rationale ("a floating tag lets the database change under an unchanged repo") and pins the same digest. `nginx:alpine` does float a major; `dpage/pgadmin4:9.2` is at least minor-pinned, so only the nginx half floats freely. Scope is the local dev stack only — the test compose file and CI service containers are already pinned. +- issue: (pending cross-reference) + +### A13-022 The dev image installs a golang-migrate release binary the production image was rewritten to avoid +- category: ops +- severity: low +- location: Dockerfile.dev:29 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `Dockerfile:55-85` explains at length that upstream's prebuilt `migrate` tarballs ship whatever toolchain and dependency versions upstream built them with, that this is how issue #1833's stdlib CVEs reached the runtime image, and that migrate must therefore be built from this module's own `go.mod`. `Dockerfile.dev` still downloads the v4.17.0 release tarball, an older release than the one `go.mod` resolves, so the development image ships exactly the transitive advisories the production image was changed to remove, and a developer running `migrate` locally exercises a different binary from the one that runs in the container. `go install github.com/air-verse/air@v1.61.7` on line 16 resolves air's own `go.mod` for the same reason. +- evidence: + ```dockerfile + curl -Lo migrate.tar.gz "https://github.com/golang-migrate/migrate/releases/download/v4.17.0/migrate.linux-${MIGRATE_ARCH}.tar.gz" && \ + echo "${MIGRATE_SHA256} migrate.tar.gz" | sha256sum -c - && \ + ``` +- suggested fix: Build `migrate` from the main module in `Dockerfile.dev` the way `Dockerfile` does, so the two images run the same binary at the same versions. +- verdict: CONFIRMED — Dockerfile.dev:22-33 downloads the v4.17.0 release tarball, while Dockerfile:55-85 explains that upstream tarballs carry upstream's own toolchain and dependency pins ("how issue #1833's stdlib CVEs reached the runtime image") and builds migrate as a package of this module at the version `go.mod` resolves (v4.19.1, go.mod:106) with `-tags=pgx5`. The divergence is worse than the finding states: the prod binary registers only the `pgx5://` scheme while the dev tarball is the default `postgres://` build, so the same migration command is not portable between the two images. `go install github.com/air-verse/air@v1.61.7` at Dockerfile.dev:16 does resolve air's own go.mod, as claimed. +- issue: (pending cross-reference) + +### A13-023 The email sender publishes to SNS but no IaC grants `sns:Publish` +- category: ops +- severity: low +- location: terraform/modules/compute/aws/lambda/main.tf:381 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `internal/email/sender.go:192` publishes to `s.topicARN`, sourced from the `SNS_TOPIC_ARN` environment variable (internal/email/factory.go:77). No terraform module or CloudFormation template grants `sns:Publish` to the runtime role, and `sns:` is absent from the workload permissions boundary's `WorkloadServiceCeiling` as well, so a boundaried role would be denied even if an identity policy granted it. Today nothing sets `SNS_TOPIC_ARN` so `SendNotification` takes the empty-topic skip path; the moment an operator wires `module.monitoring.sns_topic_arn` into the Lambda environment, every notification fails with AccessDenied rather than the intended send. +- evidence: + ```go + _, err := s.snsClient.Publish(ctx, &sns.PublishInput{ + TopicArn: aws.String(s.topicARN), + ``` +- suggested fix: Grant `sns:Publish` scoped to the notification topic ARN in the runtime modules and add `sns:Publish` to `WorkloadServiceCeiling`, in the same change that first sets `SNS_TOPIC_ARN`. +- verdict: PLAUSIBLE — the claim that "no terraform module or CloudFormation template grants sns:Publish" is wrong: cloudformation/stacks/CUDly/template.yaml:568-573 has an `SNSPublish` Sid scoped to `!Ref NotificationTopic`, and that stack also creates the topic (:729). The Terraform half and the boundary half do hold: `/usr/bin/grep -rn 'sns:' terraform iac` returns nothing, and `WorkloadServiceCeiling` (policy_boundary.tf:160-198) lists no `sns:` entry. The failure needs the runtime condition the finding names, an operator wiring the topic ARN, and there is a second wrinkle it missed: even the CFN stack sets `NOTIFICATION_TOPIC_ARN` (:607) while internal/email/factory.go:77 reads `SNS_TOPIC_ARN`, so `SendNotification` takes the empty-topic skip path there too. +- issue: (pending cross-reference) + +### A13-024 Customer-facing federation modules use unbounded `>=` provider constraints two majors behind their own lock files +- category: ops +- severity: low +- location: iac/federation/aws-target/terraform/main.tf:6 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: All five federation roots constrain with `>=` (`aws >= 5.0`, `google >= 5.0`, `http >= 3.4`) while their committed `.terraform.lock.hcl` files pin `aws 6.40.0` and `google 7.27.0`. A customer who unpacks the bundle into a directory that already has a lock, runs `terraform init -upgrade`, or is given only the `.tf` files gets whatever major is current at the time, with no floor that reflects what the module was written and tested against. The repo's own workload modules use `~>` pessimistic constraints throughout; only the customer-facing bundles do not. +- evidence: + ```hcl + aws = { + source = "hashicorp/aws" + version = ">= 5.0" + } + ``` +- suggested fix: Change the five federation roots to `~> 6.0` / `~> 7.0` / `~> 3.5` to match their lock files, so a customer's major-version upgrade is a deliberate edit. +- verdict: CONFIRMED — four of the five roots leave their primary provider unbounded (`aws >= 5.0` in aws-cross-account/terraform/main.tf:6 and aws-target/terraform/main.tf:6, `google >= 5.0` in gcp-sa-impersonation and gcp-target main.tf:6) and all five use `http >= 3.4`, against locks that pin `aws 6.40.0`, `google 7.27.0` and `http 3.5.0`. The workload side does use pessimistic constraints (`~> 5.0`, `~> 3.3`, `~> 3.4`, `~> 2.0` in terraform/environments/aws/main.tf:10-22, `~> 5.0` in the compute module versions.tf files). Two corrections: azure-target already uses `~> 3.8` / `~> 4.0` for azuread and azurerm, so only its `http` line is unbounded, and the drift is two majors for google but one for aws. +- issue: (pending cross-reference) + +### A13b-014 CloudFormation WIF template makes the operator keep the issuer URL and its condition-key host in sync by hand +- category: ops +- severity: low +- location: iac/federation/aws-target/cloudformation/template.yaml:20 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `OIDCIssuerHost` is a second parameter the operator must set to `OIDCIssuerURL` minus `https://`, and the description says outright that keeping them in sync is the operator's job. It is used to build the trust-policy condition keys `:aud` and `:sub` at lines 265-267. If the two disagree, IAM never populates those keys for the incoming token, `StringEquals` fails, and every `AssumeRoleWithWebIdentity` is denied. That fails closed, but it fails after the stack reports success, with no diagnostic pointing at the mismatch. The Terraform module derives the host from the URL in one line (`iac/federation/aws-target/terraform/main.tf:22`) and has no second parameter to desynchronise. +- evidence: + ```yaml + OIDCIssuerHost: + Type: String + Description: > + OIDC issuer host and path WITHOUT the https:// prefix — used as the IAM + condition key. Must equal OIDCIssuerURL with the https:// stripped and no + trailing slash; the operator is responsible for keeping the two in sync. + ``` +- suggested fix: Remove `OIDCIssuerHost` and derive it as `!Select [1, !Split ["https://", !Ref OIDCIssuerURL]]`, matching the Terraform local so one input cannot contradict the other. +- verdict: CONFIRMED — `iac/federation/aws-target/cloudformation/template.yaml:20-28` is a second free-text parameter whose own description hands the operator the sync duty, and its two `AllowedPattern`s constrain shape only, never the relationship to `OIDCIssuerURL` (lines 8-18); the value is spliced into both condition keys at `:265-267`, and `iac/federation/aws-target/terraform/main.tf:22` derives the same host from the URL with `trimsuffix(trimprefix(...))` so the Terraform path has nothing to desynchronise. Fails closed, as the finding states. +- issue: (pending cross-reference) + +### A13b-016 Azure registration fires immediately after the role assignment, ahead of RBAC propagation +- category: ops +- severity: low +- location: iac/federation/azure-target/terraform/registration.tf:30 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The registration POST depends on `azurerm_role_assignment.cudly_reservations`, so it fires the moment the assignment resource is created. Azure RBAC propagation is asynchronous and takes up to ten minutes, so CUDly receives a registered subscription and starts collecting against a principal whose grant has not landed. The first collection or purchase attempt returns 403 and reads as a permissions misconfiguration rather than a timing artefact. Nothing in the module waits. +- evidence: + ```hcl + # Defer to apply phase — ensure Azure resources are created before registering. + depends_on = [azurerm_role_assignment.cudly_reservations] + ``` +- suggested fix: Insert a `time_sleep` of a few minutes between the role assignment and the registration `data.http`, or have the registration payload flag that RBAC may still be propagating so the backend retries rather than reporting a hard failure. +- verdict: PLAUSIBLE — the code half holds: `iac/federation/azure-target/terraform/registration.tf:30` waits only on resource creation and nothing in the module sleeps, while the module's own comment at `iac/federation/azure-target/terraform/main.tf:76-77` puts RBAC propagation at up to ten minutes. But registration does not start collection: `internal/api/handler_registrations.go:76-78` stores the account as `pending`, and only the admin-driven `approveRegistration` (`:252`, enabling the account at `:288`) creates the cloud account, so the 403 needs the runtime condition that an operator approves the registration inside the propagation window — which I cannot establish from source. +- issue: (pending cross-reference) + +### A13c-014 GCP Cloud SQL module defaults `deletion_protection` to false where the AWS module defaults it to true +- category: ops +- severity: low +- location: terraform/modules/database/gcp/variables.tf:199 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `modules/database/aws/variables.tf:87` defaults `deletion_protection = true`, + so a caller who forgets the input gets the safe posture. The GCP module defaults false, so the + same omission yields a deletable production database. `environments/gcp/database.tf:39` does + pass `var.database_deletion_protection`, which masks this for the in-tree callers, but the + module contract itself is unsafe by default and differs from its sibling for no stated reason. +- evidence: + ```hcl + variable "deletion_protection" { + description = "Enable deletion protection" + type = bool + default = false + } + ``` +- suggested fix: flip the default to true so the three database modules agree, and let dev + profiles opt out explicitly as `environments/gcp/github-dev.tfvars:50` already does. +- verdict: CONFIRMED — database/gcp/variables.tf:199-202 defaults false while database/aws/ + variables.tf:87-90 defaults true. The masking is exactly as described: environments/gcp/ + database.tf:39 passes `var.database_deletion_protection`, whose env-layer default is true + (environments/gcp/variables.tf:174) and which github-dev.tfvars:50 overrides to false. +- issue: (pending cross-reference) + +### A13c-024 GCP OIDC signing key sets `prevent_destroy = false` where its AWS counterpart was explicitly given `create_before_destroy` +- category: ops +- severity: low +- location: terraform/modules/compute/gcp/cloud-run/signing-key.tf:35 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `modules/compute/aws/lambda/signing-key.tf:25-27` carries + `create_before_destroy = true` with a comment explaining that replacing the OIDC signing CMK + destroy-first breaks client-assertion JWT minting (PR #1480 follow-up). The GCP key opts the + other way and additionally shortens `destroy_scheduled_duration` to one day "because tests + redeploy often". A ForceNew on this key — changing `purpose`, `version_template.algorithm`, or + the key ring's `location` via `var.region` — schedules the live signing key for destruction and + every federated target cloud rejects assertions until the new JWKS propagates. +- evidence: + ```hcl + resource "google_kms_crypto_key" "signing" { + purpose = "ASYMMETRIC_SIGN" + destroy_scheduled_duration = "86400s" # 1 day — tests redeploy often + lifecycle { + prevent_destroy = false + } + } + ``` +- suggested fix: make `prevent_destroy` and `destroy_scheduled_duration` inputs so production + deployments get the protective values and only test environments opt out. +- verdict: CONFIRMED — the GCP key carries `destroy_scheduled_duration = "86400s"` with the + "tests redeploy often" comment and an explicit `prevent_destroy = false` + (cloud-run/signing-key.tf:24-38), against the AWS twin's `create_before_destroy = true` and its + seven-line rationale citing #1480 (lambda/signing-key.tf:19-27). The two mechanisms are not + equivalent — `create_before_destroy` protects a ForceNew replacement, `prevent_destroy` blocks a + destroy — but the GCP key has neither, and `prevent_destroy = false` is also Terraform's default, + so the line states rather than changes the posture. +- issue: (pending cross-reference) + +### A14-016 The Snyk job cannot fail and is not a dependency of CI Success +- category: ops +- severity: medium +- location: .github/workflows/ci.yml:829 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `continue-on-error: true` on the only step means `snyk-scan` always reports success, and the job is absent from `ci-success`'s `needs` list (lines 1064-1080), so even a job-level failure would not gate a merge. If `SNYK_TOKEN` is unset the action no-ops as well. The README lists Snyk as CI job 7 (README.md:38) and `make security-scan-all` chains it, so the repo presents dependency scanning as a gate that has no effect on any outcome. +- evidence: + ```yaml + - name: Run Snyk to check for vulnerabilities + uses: snyk/actions/golang@b98d498629f1c368650224d6d212bf7dfa89e4bf # 0.4.0 + continue-on-error: true + env: + SNYK_TOKEN: ${{ secrets.SNYK_TOKEN }} + ``` +- suggested fix: Either drop the job and the README claim, or remove `continue-on-error`, add `snyk-scan` to `ci-success`'s `needs`, and fail loudly when the token is missing. +- verdict: CONFIRMED — ci.yml:829 sets `continue-on-error: true` on the job's only substantive step, and `ci-success`'s `needs` list (ci.yml:1064-1080) enumerates 17 jobs, none of which is `snyk-scan`. +- severity-adjusted: low — .github/workflows/README.md:38,50 already mark `SNYK_TOKEN` as optional, so nothing in the repo relies on this job as a merge gate. +- issue: (pending cross-reference) + +### A14-037 `make complexity` gates test files that ci.yml and .golangci.yml both exempt +- category: ops +- severity: low +- location: Makefile:126 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The local target runs `gocyclo -over 10 .` across every file, while `ci.yml:70` runs `gocyclo -over 10 -ignore "_test\.go" .` and `.golangci.yml:42-43,86-92` sets `min-complexity: 15` and excludes gocyclo on `_test\.go`. A developer running `make pre-commit` or `make ci` gets failures on table-driven tests that CI accepts, so the local gate is stricter than the merge gate in a direction that trains people to ignore it. The `2>&1` capture also folds a gocyclo tool error into the "complexity issues" branch. +- evidence: + ```make + COMPLEXITY_ISSUES=$$(gocyclo -over 10 . 2>&1 || true); \ + if [ -n "$$COMPLEXITY_ISSUES" ]; then \ + ``` +- suggested fix: Add `-ignore "_test\.go"` so the local target matches ci.yml, and keep stderr separate from the findings list. +- verdict: CONFIRMED — Makefile:126 runs `gocyclo -over 10 .` with no ignore while ci.yml:70 passes `-ignore "_test\.go"` and .golangci.yml sets gocyclo min-complexity 15 plus a `_test\.go` exclusion; both `ci` and `pre-commit` depend on the `complexity` target. +- issue: (pending cross-reference) + +### A14-038 "Check coverage threshold" never fails and never checks a real threshold +- category: ops +- severity: low +- location: .github/workflows/ci.yml:312 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The step emits `::warning::` and exits 0, so coverage dropping from 80% to 5% does not affect `unit-tests` or `ci-success`. If the `grep total` produces an empty string (a merge that lost the total line), `bc -l` errors and the `(( ))` is false, so the step also passes silently on a broken profile. A step named "Check coverage threshold" that cannot fail reads as a gate in the job list. +- evidence: + ```bash + coverage=$(go tool cover -func="$RUNNER_TEMP/coverage.out" | grep total | awk '{print $3}' | sed 's/%//') + echo "Total coverage: ${coverage}%" + if (( $(echo "$coverage < 80" | bc -l) )); then + echo "::warning::Coverage is below 80% (current: ${coverage}%)" + fi + ``` +- suggested fix: Either fail the step below an agreed floor, or rename it to "Report coverage" so it is not read as a gate; fail loudly when `coverage` is empty. +- verdict: CONFIRMED — ci.yml:312-317 emits only `::warning::` with no exit, so coverage gates neither `unit-tests` nor `ci-success`; the finding's secondary claim is wrong, since GitHub's default `bash -eo pipefail` makes an empty `grep total` fail the assignment rather than pass silently. +- issue: (pending cross-reference) + +### A14-039 The read-only sanity checks are documented as verifying deploy-credential permissions +- category: ops +- severity: low +- location: ci_cd_sanity_tests/pkg/sanity/aws/aws.go:1 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The package comment says the checks "verify that deploy credentials have sufficient IAM permissions before a real deploy". `aws_sanity.yml:54` assumes `secrets.AWS_CICD_READONLY_ROLE_ARN`, a different role from the `vars.AWS_ROLE_TO_ASSUME` the deploys use, and the four checks are `sts:GetCallerIdentity`, `ec2:DescribeRegions`, `ec2:DescribeInstances` and `rds:DescribeDBInstances` — none of which the deploy role's IAM policy is even asserted against. A green sanity report therefore says nothing about whether a deploy will have the permissions it needs, which is what the comment promises. +- evidence: + ```go + // Package aws implements read-only AWS sanity checks used in CI/CD to verify + // that deploy credentials have sufficient IAM permissions before a real deploy. + package aws + ``` +- suggested fix: Reword the comment to describe what it does (read-only reachability of a small API set under the CI read-only role), or point the check at the deploy role's actual action list. +- verdict: CONFIRMED — ci_cd_sanity_tests/pkg/sanity/aws/aws.go:1-2 promises verification that "deploy credentials have sufficient IAM permissions", but aws_sanity.yml:54 assumes `secrets.AWS_CICD_READONLY_ROLE_ARN` while deploys assume `vars.AWS_ROLE_TO_ASSUME` (deploy-aws-lambda.yml:246), and the only checks are the four read-only calls at aws.go:123-132. +- issue: (pending cross-reference) + +### A14-041 `terraform fmt -check` on tfvars files is downgraded to a note +- category: ops +- severity: low +- location: .github/workflows/ci.yml:602 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Both tfvars format checks end in `|| echo "Note: … may need formatting"`, so a malformed `github-prod.tfvars` passes the `Validate Terraform` job. `terraform fmt -check` also returns non-zero on a genuine parse error, not only on formatting, so a broken production variables file reaches the deploy workflows undetected while the step it lives in is named "Validate environment tfvars files". +- evidence: + ```bash + if [ -f "github-${env}.tfvars" ]; then + echo "Checking github-${env}.tfvars syntax..." + terraform fmt -check "github-${env}.tfvars" || echo "Note: github-${env}.tfvars may need formatting" + fi + ``` +- suggested fix: Let the non-zero exit propagate, or run `terraform fmt -check -recursive` over the environment directory as part of the existing format step. +- verdict: CONFIRMED — ci.yml:595-608 ends both tfvars checks with `terraform fmt -check ... || echo "Note: ..."`, so a non-zero exit from either formatting or a parse error never fails the "Validate environment tfvars files" step. +- issue: (pending cross-reference) + +### A14-043 `docker-test` and the sanity-skip gates report success for work that did not happen +- category: ops +- severity: low +- location: Makefile:213 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `make docker-test` builds the image and then runs `--help` with `|| true`, so a container that cannot start still reports the target as passing. The same shape appears in `aws_sanity.yml:80-82` and `azure_sanity.yml:79-81`: when the cloud secrets are absent the job prints "skipped because required secrets are not configured" and succeeds. If a secret is rotated away or a fork's PR runs, the sanity workflows go green forever with no signal that the check stopped running, which is indistinguishable from the check passing. +- evidence: + ```make + docker-test: docker-build + @echo "Testing Docker image..." + docker run --rm cudly:$(VERSION) /app/cudly --help || true + ``` +- suggested fix: Drop the `|| true` from `docker-test`; for the sanity workflows, emit a `::warning::` and surface the skipped state in the job name or summary so a silently disarmed check is visible. +- verdict: CONFIRMED — Makefile:211-213 ends `docker run ... --help` with `|| true`, and the skip paths at aws_sanity.yml:80-82 and azure_sanity.yml:79-81 echo a message on a green job whenever `should_run != 'true'`. +- issue: (pending cross-reference) + +### A15-002 GO-2026-5932 (x/crypto/openpgp) is present, unreachable, and has no version that clears it +- category: ops +- severity: low +- location: go.mod:67 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `govulncheck ./...` at the repo root reports one advisory at "modules you require" level against `golang.org/x/crypto@v0.55.0`. Its OSV record has ranges `[{"introduced":"0"}]` with no `fixed` event, so every version that will ever exist matches and no `go get` clears it. Anyone treating this as a routine bump would produce a change that claims to resolve the advisory while govulncheck still reports it. +- evidence: + ```text + Vulnerability #1: GO-2026-5932 + Module: golang.org/x/crypto + Found in: golang.org/x/crypto@v0.55.0 + Fixed in: N/A + ``` + Unreachable, verified two ways: govulncheck reports 0 called and 0 imported symbols, and `git grep 'crypto/openpgp' -- '*.go'` returns nothing, so none of the seven affected `openpgp/*` import paths is used. +- suggested fix: No action beyond recording it as accepted. Do not attempt a bump; if the advisory must be silenced, that is an explicit suppression decision for the owner, not a dependency change. +- verdict: CONFIRMED — reproduced with a govulncheck rebuilt as v1.7.0 under the local go1.27.0 (not the stale binary): `govulncheck -show verbose ./...` at the repo root reports exactly one advisory, GO-2026-5932 against `golang.org/x/crypto@v0.55.0` (go.mod:67) with `Fixed in: N/A`, 0 called and 0 imported symbols, and `git grep 'crypto/openpgp' -- '*.go'` returns nothing, so the no-version-clears-it claim and the unreachability claim both hold. +- issue: (pending cross-reference) + +### Category: hygiene + +14 findings: 1 medium, 13 low. + +### A02-011 openapi.yaml omits roughly a third of the routed surface and mis-describes several documented operations +- category: hygiene +- severity: medium +- location: internal/api/openapi.yaml:44 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The spec is served publicly at `/docs` and is the contract the 403 regression test (openapi_403_test.go) walks, so undocumented routes are exempt from that guard. Routes registered in router.go with no path entry: `/api/recommendations/freshness` (126), `/api/recommendations/{id}/detail` (133), `/api/plans/{id}/accounts` (146-147), `/api/purchases/revoke/{id}` (171-172), `/api/purchases/retry/{id}` (178), `/api/purchases/{id}/revoke` and `/revoke/calculate` (184, 189), `/api/purchases/{id}/marketplace-list|cancel` (194-195), `/api/analytics/trends` (216), `/api/auth/me/permissions` (228), `/api/auth/reset-password/status` (233), all four `/api/auth/mfa/*` (239-242), every `/api/accounts*` route (264-279), `/api/inventory/commitments|coverage` (295, 299), `/api/ladder/configs` (325-326), `/api/notifications/unsubscribe` (330-331), `/api/register`, `/api/register/{token}` and all `/api/registrations*` (334-342), `/api/federation/iac` (347), `/version` (369), and the HEAD methods on docs (373, 375). Documented operations that disagree with the handler: `/api/history` lists an `interval` param the handler ignores and omits `provider`, `account_ids`, `limit` (yaml:1075-1099 vs handler_history.go:666-710); `/api/history/analytics` omits `interval` and `provider` (yaml:1101-1120 vs handler_analytics.go:131-145); `/api/history/breakdown` dimension enum lacks `account` (yaml:1129 vs analytics_postgres.go:72); `PUT /api/config` and `PUT /api/config/service/{service}` document a config body response but return `{"status":"updated"}` (yaml:110-117, 165-171 vs handler_config.go:130, 364); `/api/config/service/{service}` path enum `[ec2, rds, elasticache, opensearch]` but the segment is `provider/service` and GET returns 200 `{}` not 404 (yaml:137, 176 vs handler_config.go:297-314); `/api/recommendations/refresh` documents 200 `CollectResult` but returns `RefreshResponse` and can 409 (yaml:239-257 vs handler_recommendations_refresh.go:37-40, 82); `/api/recommendations` omits `account_ids`, `min_savings_usd`, `min_savings_pct` and lists `account_id` which the handler does not read (yaml:206-237 vs handler_recommendations.go:39-79); `/api/purchases/approve|cancel/{id}` document POST only with the token required in the query (yaml:470-520) while the router also serves GET and reads the token from the body first (router.go:163-166, 568-586); `/api/dashboard/summary` lists StartDate/EndDate the handler never reads (yaml:53-55); `/api/api-keys/{id}` and `/api/users/{id}` document 404 where the handler returns 500 (see A02-012); `/api/auth/setup-admin` omits 409/503, `/api/auth/login` omits 503, `/api/auth/reset-password` omits 429/503 (yaml:1242-1310 vs handler_auth.go:24, 234-247, 345); `/api/docs/openapi.yaml` declares `application/x-yaml` but the handler emits `application/yaml` (yaml:1795 vs handler_docs.go:104). +- evidence: + ```yaml + /api/history: + get: + parameters: + - $ref: '#/components/parameters/AccountID' + - $ref: '#/components/parameters/StartDate' + - $ref: '#/components/parameters/EndDate' + - name: interval + ``` +- suggested fix: Add a test that diffs `registerRoutes()` (path+method) against the spec's paths and fails on any unlisted route, then backfill the missing operations and correct the listed mismatches in one pass. +- verdict: CONFIRMED — the spec's path keys (openapi.yaml:44-1800) contain no /api/accounts*, /api/registrations*, /api/register*, /api/recommendations/freshness, /api/auth/mfa/*, /api/inventory/*, /api/ladder/configs, /api/analytics/trends or /version entries although router.go registers them (e.g. router.go:264-279, 334-342, 369); spot-checked mismatches hold: openapi.yaml:1080-1088 lists `interval` and omits provider/account_ids/limit versus parseHistoryFilters (handler_history.go:666-706), openapi.yaml:137 enumerates bare service names versus the provider/service split at handler_config.go:297-303, and openapi.yaml:1795 says application/x-yaml versus handler_docs.go:104 application/yaml. +- issue: (pending cross-reference) + +### A02-019 Store errors are classified by substring matching in four handlers +- category: hygiene +- severity: low +- location: internal/api/handler_accounts.go:268 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `isDuplicateKeyError` tests `err.Error()` for "duplicate key" / "23505" while forty lines later `deleteAccount` (lines 644-645) does the typed `errors.As(err, &pgErr)` check; `submitRegistration` matches "duplicate" (handler_registrations.go:106), `mergeServiceConfig` matches "not found" (handler_config.go:219) and `isResetPasswordClientError` keeps a list of five message fragments (handler_auth.go:291-298). A wrapped error whose message is reworded (or a legitimately different error that happens to contain "not found") flips a 409/400 into a 500 or vice versa, and `mergeServiceConfig` would treat a transport error mentioning "not found" as "no existing row" and overwrite filter fields it was meant to preserve. +- evidence: + ```go + s := err.Error() + return strings.Contains(s, "duplicate key") || strings.Contains(s, "23505") + ``` +- suggested fix: Use `pgconn.PgError` codes / `config.ErrNotFound` / exported auth sentinels with `errors.Is`/`errors.As` in all four sites and delete the substring helpers. +- verdict: CONFIRMED — isDuplicateKeyError (handler_accounts.go:262-268), submitRegistration (handler_registrations.go:106), mergeServiceConfig (handler_config.go:217-221) and isResetPasswordClientError (handler_auth.go:291-298) all branch on err.Error() substrings, while deleteAccount (handler_accounts.go:644-645) does the typed pgconn.PgError/errors.As check in the same package; the mergeServiceConfig "not found" branch really does return cfg unchanged on any error whose text contains that phrase. +- issue: (pending cross-reference) + +### A03-021 Test doubles and testify/mock are compiled into the production auth package +- category: hygiene +- severity: low +- location: internal/auth/test_helpers.go:15 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `test_helpers.go` is not a `_test.go` file, so `MockStore`, `MockEmailSender`, `TestCSRFKey`, `DeriveTestCSRFToken`, the fixed `testCSRFKey` and the `github.com/stretchr/testify/mock` dependency ship in every server, Lambda and CLI binary. A caller can construct a service with the well-known test CSRF key; and the mocks are what let the proof-of-concept for A03-001 through A03-009 be driven from outside the package. +- evidence: + ```go + // MockStore is a mock implementation of the auth store for testing. + type MockStore struct { + mock.Mock + } + ``` +- suggested fix: Move the doubles to an `authtest` package or to `export_test.go`/`_test.go` files; `internal/mocks` already exists for cross-package mocks. +- verdict: CONFIRMED — internal/auth/test_helpers.go is a non-_test file in package auth importing testify/mock, testify/require and "testing" (:3-12) and exporting MockStore, MockEmailSender and DeriveTestCSRFToken with the fixed testCSRFKey (:14-17, :281-303), so it compiles into every binary that imports internal/auth; no non-test code references these symbols (grep outside the file returns nothing), so the "caller can construct a service with the test key" claim describes reachable-by-linking, not an in-tree call. +- issue: (pending cross-reference) + +### A03-024 Group creation records no creator for any actor +- category: hygiene +- severity: low +- location: internal/auth/service_api.go:313 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `CreateGroupAPI` passes `""` as `createdBy` for every caller because the admin-API-key sentinel is not a UUID. Human admins therefore also leave `created_by = NULL`, so after A03-001/A03-004 there is no audit trail of who created the escalating group. +- evidence: + ```go + // Use empty string for createdBy: the column is a UUID FK and + // actorUserID may be the non-UUID admin-API-key sentinel. + if err := s.CreateGroup(ctx, group, ""); err != nil { + ``` +- suggested fix: Pass `actorUserID` when it is not `AdminAPIKeyActorID` and `""` only for the sentinel. +- verdict: CONFIRMED — CreateGroupAPI receives the real actorUserID from the handler (handler_groups.go:53) and hands CreateGroup a literal "" for every caller (service_api.go:311-313), so human admins leave created_by unset even though the sentinel case (group_ceiling.go:40) is the only one that cannot be stored. +- issue: (pending cross-reference) + +### A04-018 Auto-heal naming and the newMigratorWithRecovery comment promise a recovery that no longer happens +- category: hygiene +- severity: low +- location: internal/database/postgres/migrations/migrate.go:114-118 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (behaviour at 381-421) +- failure scenario: the call-site comment says a dirty row is cleared "at the CURRENT recorded version so the subsequent Up() re-applies any pending migrations", and the gate is named `CUDLY_MIGRATION_AUTOHEAL` with a default of true. `maybeAutoHealDirty` heals nothing; it returns an error. An operator following the comment (or the memory note describing Force(current)+Up as default-on) leaves `CUDLY_MIGRATION_AUTOHEAL=true` expecting self-recovery and gets a permanently dirty database until they discover `CUDLY_FORCE_MIGRATION_VERSION`. +- evidence: + ```go + // Default-on dirty auto-heal: when the schema_migrations row is dirty, + // clear the dirty flag at the CURRENT recorded version so the subsequent + // Up() re-applies any pending migrations, letting a cold start self-recover + // instead of staying broken until a manual force. + if err := maybeAutoHealDirty(m); err != nil { + ``` +- suggested fix: rename to `maybeRefuseDirty` / `CUDLY_MIGRATION_DIRTY_CHECK` and rewrite the call-site comment to match the fail-loud behaviour. +- verdict: CONFIRMED — the call-site comment promises "clear the dirty flag at the CURRENT recorded version so the subsequent Up() re-applies any pending migrations" (internal/database/postgres/migrations/migrate.go:114-118), but `maybeAutoHealDirty` calls no `Force` and returns an error on every dirty row (migrate.go:381-421), and `autoHealEnabled` defaults the gate to true (migrate.go:428). +- issue: (pending cross-reference) + +### A08b-044 GCP billing service IDs and the 8760-hour year are inline literals across seven files +- category: hygiene +- severity: low +- location: providers/gcp/services/cloudsql/client.go:350 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd (service IDs also at `cloudstorage/client.go:343`, `memorystore/client.go:337`; `8760.0` at `cache/client.go:548`, `cosmosdb/client.go:548`, `database/client.go:567`, `search/client.go:463`, `managedredis/client.go:492`, `synapse/client.go:439`, and the three GCP files) +- failure scenario: the opaque service ID `services/9662-B51E-5089` appears with no name and no comment, so a reader cannot tell which GCP service is being priced and a transposed digit produces an empty catalog that surfaces as "no pricing found". `8760.0` is repeated in nine files as a bare literal on money paths, with no shared constant to change if the term-to-hours convention is ever revisited. +- evidence: + ```go + skus, err := svc.ListSKUs("services/9662-B51E-5089") + if err != nil { + return nil, fmt.Errorf("failed to list SKUs: %w", err) + } + ``` +- suggested fix: name each service ID as a documented package constant and hoist `hoursPerYear = 8760.0` into the shared pricing helper both providers already import. +- verdict: CONFIRMED — the four service IDs are inline and uncommented at cloudsql:350, cloudstorage:343, memorystore:337 and computeengine:961, and the count of bare `8760.0` literals is higher than stated: twelve provider files, adding compute/client.go:709 and savingsplans/client.go:411 to the nine listed. +- issue: (pending cross-reference) + +### A09-030 Retry documents a per-attempt context as independent of the outer context, and jitter escapes MaxDelay +- category: hygiene +- severity: low +- location: pkg/retry/exponential.go:145 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `context.WithTimeout(ctx, ...)` produces a child, so cancelling the outer context cancels the attempt — the opposite of the "independent of the outer ctx so a slow attempt fails fast and the retry budget continues" claim repeated at lines 36-38 and 96-97. A caller relying on that sentence will size `PerAttemptTimeout` on the assumption that a cancelled parent still lets the current attempt finish. Separately, `backoffFor` applies the ±25% jitter factor after the `MaxDelay` clamp, so with `MaxDelay = 30s` an actual sleep of up to 37.5s occurs, exceeding the documented cap. +- evidence: + ```go + perAttemptCtx, cancel := context.WithTimeout(ctx, cfg.PerAttemptTimeout) + defer cancel() + return op(perAttemptCtx, attempt) + ``` +- suggested fix: reword the doc to say the per-attempt deadline is *additional* to the outer context, and clamp the post-jitter delay to `MaxDelay` so the documented cap holds. +- verdict: CONFIRMED — both halves check out. `runAttempt` derives the per-attempt context from the caller's ctx via `context.WithTimeout(ctx, cfg.PerAttemptTimeout)` (pkg/retry/exponential.go:141-147), so cancelling the parent cancels the attempt, contradicting the "independent of the outer ctx" wording at exponential.go:36-38 and :92-94. And `backoffFor` applies the ±25% factor after both `MaxDelay` clamps (exponential.go:168-186), so a 30s MaxDelay yields sleeps up to 37.5s. +- issue: (pending cross-reference) + +### A12-049 `HOURS_PER_MONTH` is 730 but its comment derives 730.5 +- category: hygiene +- severity: low +- location: frontend/src/modules/savings-history.ts:25 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The comment states the derivation `365.25 * 24 / 12`, which is 730.5, while the constant is 730. A reader auditing the $/hr conversion cannot tell whether the value or the comment is authoritative, and neither matches the 720 used by `recommendations.ts` (A12-012). +- evidence: + ```ts + const HOURS_PER_MONTH = 730; // 365.25 * 24 / 12 + ``` +- suggested fix: Fix the comment to name the convention actually used (AWS's 730 hours/month), and share the constant with `recommendations.ts`. +- verdict: CONFIRMED — The constant is 730 while its own comment derives 365.25 * 24 / 12 = 730.5 (frontend/src/modules/savings-history.ts:25), and frontend/src/recommendations.ts:1471 independently uses 1/720 with a 24 x 30 comment. +- issue: (pending cross-reference) + +### A12-070 The staleness banner cannot distinguish soft from hard and contradicts its own docstring +- category: hygiene +- severity: low +- location: frontend/src/riexchange.ts:548 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The function contract at line 509 says soft means "data may be up to 12 h old", but the rendered soft copy says "may be up to 24h old" and the hard copy says "older than 24h". Both banners therefore tell the operator the same thing about price freshness before they act on a cross-family alternative; only the colour differs. +- evidence: + ```ts + const copy = isSoft + ? `Cross-family alternatives are based on Cost Explorer recommendations that may be up to 24h old${ageLabel}. Some prices may be stale.` + : `Cross-family alternatives are based on Cost Explorer recommendations older than 24h${ageLabel}. Prices may be significantly out of date.`; + ``` +- suggested fix: Make the soft copy say 12h, or drop the hardcoded hour figures and rely on `ageLabel`. +- verdict: CONFIRMED — The contract says soft means data may be up to 12 h old (frontend/src/riexchange.ts:509) while the rendered soft copy says 24h and the hard copy says older than 24h (frontend/src/riexchange.ts:547-550), so both banners state the same freshness bound. +- issue: (pending cross-reference) + +### A13-016 `.trivyignore` GCP-0015 describes a variable that does not exist +- category: hygiene +- severity: low +- location: .trivyignore:93 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The suppression says the GCP database module "exposes require_ssl as a variable" and is "Tracked as a hardening follow-up to set require_ssl = true in the module default". There is no `require_ssl` variable in `terraform/modules/database/gcp/`; both `ip_configuration` blocks hardcode `ssl_mode = "ENCRYPTED_ONLY"` (main.tf:84 and main.tf:183), which already requires TLS. The suppression is either stale against a fixed finding or masking a different one, and it points a future reviewer at a follow-up that cannot be done. +- evidence: + ```hcl + ip_configuration { + ipv4_enabled = var.enable_public_ip + private_network = var.vpc_network_id + ssl_mode = "ENCRYPTED_ONLY" + ``` +- suggested fix: Remove the entry and re-run `trivy config` to confirm it no longer fires; if it does, replace the justification with the real cause. +- verdict: CONFIRMED — `/usr/bin/grep -rn 'require_ssl|ssl_mode' terraform/modules/database/gcp/` returns no `require_ssl` anywhere; both `ip_configuration` blocks hardcode `ssl_mode = "ENCRYPTED_ONLY"` (main.tf:84, main.tf:183) and outputs.tf:55 emits `ssl_mode = "require"` for the connection string. The suppression at .trivyignore:92-98 therefore points a future reviewer at a variable that does not exist and a follow-up that cannot be performed. +- issue: (pending cross-reference) + +### A13-017 `.snyk` declares severity and license policy keys that the Snyk policy format does not read +- category: hygiene +- severity: low +- location: .snyk:26 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: The `.snyk` policy file supports `version`, `ignore`, `patch` and `exclude`. `failOnSeverity` and `license` are not policy-file keys: Snyk takes the severity gate from `--severity-threshold` (which `Makefile:232` and the ci.yml snyk job pass explicitly) and license policy from organization settings. A reader auditing the repo's license posture sees an allow/deny list that nothing enforces, so a GPL-3.0 dependency would be admitted despite the file appearing to forbid it. +- evidence: + ```yaml + # Severity thresholds + failOnSeverity: high + + # License policy + license: + allow: + - MIT + ``` +- suggested fix: Delete the two inert blocks, or move the license policy to a real enforcement point and leave a comment saying where it lives. +- verdict: CONFIRMED — .snyk declares `ignore`, `patch`, `exclude` (and omits `version` entirely), then `failOnSeverity: high` at line 26 and a `license` allow/deny map at lines 29-43, neither of which is a Snyk policy-file key. Both real enforcement points pass the threshold on the command line instead: `Makefile:232` runs `snyk test --severity-threshold=high` and .github/workflows/ci.yml:832 passes `args: --severity-threshold=high`, and no consumer reads the license map, so a GPL-3.0 dependency would be admitted despite the deny list. +- issue: (pending cross-reference) + +### A13b-015 Registration response is explicitly marked non-sensitive while its own description says it carries a reference token +- category: hygiene +- severity: low +- location: iac/federation/aws-target/terraform/outputs.tf:24 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: All five modules declare `registration_response` with `sensitive = false` set explicitly, and describe it as containing the reference token used for status checks. The raw response body is therefore printed by `terraform apply` and lands in CI logs and any stored plan output. The same file marks the far less sensitive `cudly_account_registration` block as sensitive in the Azure module (`iac/federation/azure-target/terraform/outputs.tf:21`), so the two are inconsistent about what deserves redaction. Identical in the `aws-cross-account`, `azure-target`, `gcp-sa-impersonation` and `gcp-target` outputs files. +- evidence: + ```hcl + output "registration_response" { + description = "CUDly registration API response (contains reference_token for status checks)" + value = local.do_register ? data.http.cudly_registration[0].response_body : "Skipped (cudly_api_url or contact_email not set)" + sensitive = false + } + ``` +- suggested fix: Set `sensitive = true` in all five modules and tell the customer to read it with `terraform output -raw registration_response`. +- verdict: CONFIRMED — the `registration_response` output is byte-identical with an explicit `sensitive = false` in all five `outputs.tf`, and the body it prints really does carry the token: `internal/api/handler_registrations.go:128-131` returns `reference_token`, `internal/config/store_postgres_registrations.go:79-81` uses it as the lookup key for `GetAccountRegistrationByToken`, and `internal/api/handler_registrations.go:373` treats it as withheld even from admins; the inconsistency with the redacted `cudly_account_registration` block at `iac/federation/azure-target/terraform/outputs.tf:11-22` is as described. +- issue: (pending cross-reference) + +### A13c-017 A live deployed Lambda Function URL hostname is hardcoded in a tracked tfvars file with no mechanism to keep it current +- category: hygiene +- severity: low +- location: terraform/environments/aws/github-dev.tfvars:32 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the Function URL ID is server-assigned and only known after apply, so it + cannot be self-referenced without a cycle and has to be pasted by hand — the comment three lines + above says exactly that ("Update the Lambda Function URL entry when the dev environment is + redeployed"). Any redeploy that reissues the URL silently invalidates the CORS allowlist, and + the failure surfaces as a browser CORS error rather than a Terraform diff. The file is + deliberately un-gitignored (`terraform/.gitignore` negates `github-*.tfvars`), so a live + environment's public endpoint is published in the repository. +- evidence: + ```hcl + lambda_allowed_origins = [ + "https://33pz7pombdqwu3bdlxp4lqxyra0bsriy.lambda-url.us-east-1.on.aws", + "http://localhost:3000", + ] + ``` +- suggested fix: supply the origin via `TF_VAR_lambda_allowed_origins` from the deploy workflow + (which can read the previous apply's output) rather than committing the host. +- verdict: CONFIRMED — the hostname is committed at environments/aws/github-dev.tfvars:32, the + manual-update comment is three lines above at :30, and `terraform/.gitignore` ignores `*.tfvars` + then re-admits `!github-*.tfvars`, so the file is tracked by design. The endpoint is a public + URL rather than a credential, which is consistent with the low severity. +- issue: (pending cross-reference) + +### A15-011 Twelve source files exceed 500 lines by more than double, against the project's own limit +- category: hygiene +- severity: low +- location: internal/api/handler_purchases.go:1 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `CLAUDE.md` states "Keep files under 500 lines". Forty non-test source files exceed it; twelve exceed 1000. The length hides mixed responsibilities rather than merely being long: `handler_purchases.go` holds 85 top-level functions spanning session authorization (`authorizeSession*`), request validation (`validateExecute*`, `validateRevoke*`), notification dispatch (`sendPurchase*`), revocation, cancellation, approval and delayed scheduling. `frontend/src/recommendations.ts` is 5481 lines with 50 exports and 68 functions. Reviewing a money-path change in these files means reading past several unrelated concerns, and the file is a permanent merge-conflict hotspot for concurrent PRs. +- evidence: + ```text + 5481 frontend/src/recommendations.ts 1847 frontend/src/history.ts + 3823 internal/config/store_postgres.go 1784 internal/mocks/stores.go + 3715 frontend/src/settings.ts 1641 internal/api/handler_accounts.go + 3316 internal/api/handler_purchases.go 1536 frontend/src/auth.ts + 2594 internal/api/handler_ri_exchange.go 1519 internal/scheduler/scheduler.go + 2359 frontend/src/plans.ts 1432 providers/gcp/.../computeengine/client.go + 2270 frontend/src/riexchange.ts + ``` +- suggested fix: Split by the concern boundaries already visible in the function-name prefixes, starting with `handler_purchases.go` (authorization, validation and notification each move to their own file) rather than attempting all forty. +- verdict: CONFIRMED — the quoted line counts reproduce exactly, and the two inventory claims check out: `handler_purchases.go` has 85 `^func ` declarations whose prefixes cluster as authorize (7), build (6), resolve (5), validate (4), require (4), approve (4), send (3), cancel (3), revoke (2), and `frontend/src/recommendations.ts` has 50 `^export ` lines across 5481 lines. The counts understate rather than overstate: across tracked non-test `.go` and `frontend/src/*.ts`, 91 files exceed 500 lines and 22 exceed 1000, against the finding's "forty" and "twelve" (its own table already lists 13 files over 1000). Read the verdict with the same caveat as the finding: this rests on function inventories and name prefixes, not on a line-by-line reading of the four largest files, so it evidences mixed responsibilities by naming clusters rather than by tracing each function's behaviour. +- severity-adjusted: unchanged at low, but the scope is roughly double what is reported. +- issue: (pending cross-reference) + +### Category: bug + +1 finding: 1 low. + +This label is outside the report taxonomy; the reviewer wrote it as-is and it is preserved. + +### A10-029 The RDS instance pagination loop has no page cap and never checks context cancellation +- category: bug +- severity: low +- location: cmd/multi_service_engine_versions.go:138 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `queryRDSInstancesInRegion` loops on the marker with no iteration ceiling and no `ctx.Err()` check. If the API returns a non-advancing marker, the loop calls `DescribeDBInstances` forever, appending duplicate `InstanceEngineVersion` entries into the shared map under the mutex, which inflates `excludedCount` in `adjustRecommendationForExcludedVersions` and can drive recommendation counts to zero. Cancelling the CLI does not stop the worker; only the shared `sem` bounds how many regions spin at once. The sibling loop `fetchMajorEngineVersionsForEngine` (line 235-244) has both guards. +- evidence: + ```go + var marker *string + for { + localVersions, nextMarker, err := queryRDSInstancesPage(ctx, rdsClient, marker, regionName) + ... + if nextMarker == nil { break } + marker = nextMarker + } + ``` +- suggested fix: Add the same `if err := ctx.Err(); err != nil { return }` check and a page cap constant that `fetchMajorEngineVersionsForEngine` already uses. +- verdict: PLAUSIBLE — the missing guards are real: cmd/multi_service_engine_versions.go:137-156 has neither the ctx.Err() check nor the maxEngineVersionPages cap its sibling carries at :235-244 (tested at cmd/multi_service_engine_versions_paginate_test.go:144). The runaway needs the runtime condition of AWS returning a non-advancing marker, and the "cancelling the CLI does not stop the worker" half is wrong: a cancelled context makes DescribeDBInstances error, which hits the break at :140-143. +- issue: (pending cross-reference) + +### Category: bugs + +4 findings: 4 low. + +This label is outside the report taxonomy; the reviewer wrote it as-is and it is preserved. + +### A05-013 `executableByScheduler` dereferences the plan without the nil guard its sibling has +- category: bugs +- severity: low +- location: internal/purchase/manager.go:640 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `executePurchase` explicitly handles `GetPurchasePlan` returning `(nil, nil)` (execution.go:53) because the store interface permits it; the AutoPurchase gate does not, and `plan.AutoPurchase` panics on a nil plan. The panic is inside `processOneExecution`, called from `ProcessScheduledPurchases` in the scheduler Lambda, with no `recover()` on that path (the only recover in this shard is inside `FanOutWithConcurrency`'s goroutines), so one such row takes down the whole tick and every later due row goes unprocessed. The Postgres store happens to return `ErrNotFound` today, so this is latent rather than live, but the asymmetry with `executePurchase` is exactly what makes it easy to trip on the next store or mock. +- evidence: + ```go + plan, err := m.config.GetPurchasePlan(ctx, exec.PlanID) + if err != nil { + return false, fmt.Errorf("failed to fetch plan %s for AutoPurchase gate: %w", exec.PlanID, err) + } + return plan.AutoPurchase, nil + ``` +- suggested fix: add `if plan == nil { return false, fmt.Errorf(...) }` before the dereference — fail closed, matching the function's own stated "no silent money action on error" policy. +- verdict: PLAUSIBLE — the asymmetry is real (executePurchase guards `if plan == nil` at execution.go:52-54; executableByScheduler dereferences `plan.AutoPurchase` at manager.go:643 with no guard) and the panic would indeed escape, since the only recover() calls in this shard are internal/execution/fanout.go:115 and scheduler.go:1336, neither on the processOneExecution path. Reaching it needs a store returning (nil, nil): PostgresStore.GetPurchasePlan wraps pgx.ErrNoRows as ErrNotFound (store_postgres.go:599-601) and the interface (internal/config/interfaces.go:31) states no contract, so the crash is latent as the finding itself says. +- issue: (pending cross-reference) + +### A07-031 A transient STS failure permanently disables account-ID resolution for the client +- category: bugs +- severity: low +- location: providers/aws/services/redshift/client.go:307 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `resolveAccountID` caches the outcome in `sync.Once`. If the first call happens under a cancelled context or a throttled STS, `accountErr` is stored and every later call on the same client returns that stale error without retrying. For Redshift this is on the pre-purchase guard path, so `findNodeByIdempotencyToken` fails closed for the client's whole lifetime and blocks every subsequent purchase. The identical construct is at opensearch/client.go:313, where it only disables tagging. +- evidence: + ```go + c.accountOnce.Do(func() { + if c.stsClient == nil { + return + } + out, err := c.stsClient.GetCallerIdentity(ctx, &sts.GetCallerIdentityInput{}) + if err != nil { + c.accountErr = err + return + } + ``` +- suggested fix: cache only a successful resolution (guard with a mutex and re-attempt when `accountID` is still empty) so a transient STS error does not become permanent. +- verdict: CONFIRMED — `resolveAccountID` (providers/aws/services/redshift/client.go:306-320) stores `c.accountErr` inside `sync.Once` and returns it on every later call, and `findNodeByIdempotencyToken` :226-230 converts that into a hard "resolve account ID for idempotency check" failure for the lifetime of that client instance; opensearch/client.go has the same construct on a tagging-only path. +- issue: (pending cross-reference) + +### A08-012 The reservation-order idempotency lookup paginates with no cycle guard and no page cap +- category: bugs +- severity: medium +- location: providers/azure/services/internal/reservations/purchase.go:476 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: a `nextLink` that points back at itself (or an unbounded chain) makes the loop spin forever. This runs before every purchase and before every retry inside `purchaseTwoStepGuarded`, so a purchase with a caller-supplied context that has no deadline hangs indefinitely instead of failing. The repo already owns a walker with both guards for exactly this class of loop: `pricing.FetchAll` enforces a seen-URL guard, a max-pages cap and a per-page timeout (providers/azure/internal/pricing/retail_prices.go:63). +- evidence: + ```go + nextURL := ReservationOrdersListURL() + for nextURL != "" { + page, err := fetchReservationOrdersPage(ctx, httpClient, nextURL, bearerToken) + if err != nil { return "", false, err } + if orderID, found := matchReservationOrderInPage(page, idempotencyToken); found { return orderID, true, nil } + nextURL = page.NextLink + } + ``` +- suggested fix: add the same seen-URL set and page cap this loop's sibling walker already has, erroring on a repeat link rather than looping. +- verdict: CONFIRMED — `FindReservationOrderByIdempotencyToken` (reservations/purchase.go:475-484) follows `nextLink` with no seen-URL set and no page cap, while the sibling `pricing.FetchAll` enforces both plus a per-page timeout (internal/pricing/retail_prices.go:63-86). +- severity-adjusted: low — the scheduler/web purchase path runs every rec under a 30s context (internal/purchase/execution.go:1007), so the unbounded spin needs a caller supplying a deadline-free context. +- issue: (pending cross-reference) + +### A08-028 RI-utilization pagination has neither a page cap nor a cancellation check +- category: bugs +- severity: low +- location: providers/azure/ri_utilization.go:104 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `getRIUtilizationViaAPI` walks `pager.More()` with no bound and no `ctx.Err()` check between pages, unlike every pagination loop in the compute client, which caps pages and checks cancellation on each iteration (compute/client.go:262-268). A large tenant, or a pager that keeps reporting `More()`, keeps the call running past the caller's intent. Rows with a nil `ReservationID` are also silently discarded by `accumulateSummary` (ri_utilization.go:125), so a partial summary is reported as a complete utilization figure. +- evidence: + ```go + for pager.More() { + page, err := pager.NextPage(ctx) + if err != nil { + return nil, fmt.Errorf("azure ri utilization: failed to fetch summaries page: %w", err) + } + for _, summary := range page.Value { + accumulateSummary(agg, summary) + } + } + ``` +- suggested fix: add the same `pageIdx >= maxPages` cap and per-iteration `ctx.Err()` check the compute client's loops use. +- verdict: PLAUSIBLE — the page cap really is absent at ri_utilization.go:104-113 where the sibling `collectVMReservations` has both guards (compute/client.go:262-268), but `pager.NextPage(ctx)` already fails a cancelled context on every page fetch, so only a pager that keeps reporting `More()` produces the runaway, and the Azure two-argument `GetRIUtilization` (ri_utilization.go:65) has no caller outside ri_utilization_test.go — the API layer calls the three-argument AWS shape (internal/server/ladder_write.go:55). The nil-`ReservationID` skip is documented at ri_utilization.go:118-120, not silent. +- issue: (pending cross-reference) + +## Rejected findings + +The verifier found each of these wrong. They are kept with their reasoning so the same claim is +not re-raised. Do not action anything in this section. + +| ID | Category | Reviewer severity | Location | +|---|---|---|---| +| A03-013 | silent-fallback | medium | `internal/auth/service_group.go:414` | +| A03-017 | security | low | `internal/credentials/resolver.go:266` | +| A03-018 | security | low | `internal/secrets/azure_resolver.go:32` | +| A04-014 | correctness | low | `internal/config/store_postgres.go:1083-1090` | +| A05-016 | concurrency | low | `internal/scheduler/scheduler.go:600` | +| A06-002 | security | high | `internal/server/http.go:207` | +| A06-011 | security | low | `internal/email/smtp_sender.go:327` | +| A06-021 | correctness | low | `internal/accounts/org_discovery.go:48` | +| A06-023 | correctness | low | `internal/server/handler_ri_exchange.go:165` | +| A07-020 | silent-fallback | medium | `providers/aws/recommendations/converters.go:27` | +| A07-024 | money-path | medium | `providers/aws/services/ec2/client.go:1043` | +| A07-025 | silent-fallback | medium | `providers/aws/provider.go:476` | +| A07-030 | silent-fallback | low | `providers/aws/provider.go:419` | +| A08-011 | money-path | high | `providers/azure/services/savingsplans/client.go:490` | +| A08-025 | hygiene | low | `providers/azure/services/compute/client.go:448` | +| A09-017 | money-path | high | `pkg/common/service_details_codec.go:98` | +| A09-027 | security | low | `pkg/ladder/plan.go:195` | +| A09-029 | silent-fallback | low | `pkg/common/identifiers.go:33` | +| A11-003 | money-path | high | `frontend/src/groups/groupModals.ts:269` | +| A11-007 | money-path | high | `frontend/src/app.ts:398` | +| A11-010 | correctness | medium | `frontend/src/state.ts:167` | +| A11-021 | correctness | low | `frontend/src/users/filters.ts:131` | +| A11-025 | hygiene | low | `frontend/src/utils.ts:224` | +| A12-007 | correctness | high | `frontend/src/plans.ts:1253` | +| A12-016 | silent-fallback | medium | `frontend/src/ladder.ts:437` | +| A12-025 | money-path | medium | `frontend/src/dashboard.ts:751` | +| A12-027 | silent-fallback | medium | `frontend/src/plans.ts:1361` | +| A12-036 | test-gap | medium | `frontend/src/__tests__/history-marketplace-sell-button.test.ts:252` | +| A12-046 | silent-fallback | low | `frontend/src/plans.ts:1908` | +| A12-050 | over-engineering | low | `frontend/src/modules/savings-history.ts:576` | +| A12-051 | correctness | low | `frontend/src/plans.ts:2294` | +| A13b-006 | correctness | medium | `iac/federation/aws-cross-account/cloudformation/template.yaml:68` | +| A13c-008 | security | medium | `terraform/modules/compute/gcp/cloud-run/main.tf:454` | +| A13c-010 | ops | medium | `terraform/modules/compute/aws/fargate/main.tf:764` | +| A13c-011 | correctness | low | `terraform/modules/secrets/gcp/main.tf:183` | +| A13c-026 | hygiene | low | `terraform/modules/email/azure/main.tf:10` | +| A14-001 | security | critical | `.github/workflows/deploy-azure.yml:232` | +| A14-013 | ops | medium | `scripts/security-scan.sh:187` | +| A14-014 | ops | medium | `scripts/security-scan.sh:71` | +| A14-019 | ops | medium | `.github/workflows/deploy-aws-fargate.yml:173` | +| A15-005 | test-gap | medium | `internal/scheduler/scheduler_test.go:1048` | +| A15-010 | ops | low | `.github/workflows/ci.yml:646` | +| A16-003 | test-gap | high | `internal/purchase/approvals.go:494` | +| A16-010 | test-gap | medium | `internal/scheduler/scheduler_test.go:1048` | +| A16-013 | test-gap | medium | `internal/purchase/execution.go:766` | + +### A03-013 Constraint matching treats an absent request dimension or a zero amount as satisfied +- category: silent-fallback +- severity: medium +- location: internal/auth/service_group.go:414 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: A group grants `execute:purchases` constrained to `AccountIDs=[acct-A], MaxPurchaseAmount=100`. `permissionsAllow` is called with a request constraint set whose `AccountIDs` is empty (a caller that did not populate it, a region-agnostic Savings Plan path, or a provider whose record lacks the field) or whose amount is 0 (unknown cost). `matchStringListConstraints` returns true because `len(reqList) == 0`, and `matchPurchaseAmountConstraint(100, 0)` returns true because `reqMax > permMax` is false. The constrained grant therefore authorizes any account and any amount whenever the caller under-fills the request side; the code comments push the burden onto "callers that need ..." but nothing in scope enforces it. This is the empty-means-unrestricted shape that #1748 removed from account scope, still present on the permission-constraint axis. +- evidence: + ```go + func matchStringListConstraints(permList, reqList []string) bool { + if len(permList) > 0 && len(reqList) > 0 { + return containsAny(permList, reqList) + } + return true + } + ... + func matchPurchaseAmountConstraint(permMax, reqMax float64) bool { + if permMax > 0 && reqMax > permMax { + return false + } + ``` +- suggested fix: When the permission constrains a dimension, require the request to name it (empty request list against a non-empty permission list is a refusal; zero amount against a positive cap is a refusal), mirroring `listCovers` and `amountCovers` in group_ceiling.go, which already implement the strict polarity. +- verdict: REJECTED — the matcher polarity is as described (service_group.go:414-419, :459-464), but every caller of the constraint check fills the request side: purchaseConstraintSets (handler_purchases.go:2221-2242) always populates Providers/Services/Regions/AccountIDs, substituting the `unattributedAccountConstraint` sentinel (handler.go:466) for a missing account, requireNonZeroCommitment (handler_purchases.go:2197) refuses a zero total before the check, and checkAzureExecuteConstraints (handler_ri_exchange.go:1064-1096) names account/provider/service and still returns 403 when it probes with MaxPurchaseAmount=0; HasPermissionForConstraintsAPI also fails loud on an empty set (service_api.go:415-420), so no in-tree caller reaches the under-filled case. + +### A03-017 Web-identity token file path is tenant-supplied and read from the host filesystem +- category: security +- severity: low +- location: internal/credentials/resolver.go:266 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `AWSWebIdentityTokenFile` comes from the cloud-account record (writable by any `create:accounts`/`update:accounts` holder). The only check is a substring test for `..` and `filepath.IsAbs`, so any absolute path on the host (`/proc/self/environ`, `/var/run/secrets/...`, another tenant's projected token) is opened by the CUDly process and its contents sent to STS as the web identity token for a role ARN the same tenant chose. The substring check also rejects legitimate names containing `..` inside a component. Host-level path policy belongs to the operator, not to tenant data. +- evidence: + ```go + if strings.Contains(tokenFile, "..") || !filepath.IsAbs(tokenFile) { + return nil, fmt.Errorf("credentials: aws_web_identity_token_file must be an absolute path without '..' (account %s)", account.ID) + } + ``` +- suggested fix: Take the token file exclusively from deployment configuration (`AWS_WEB_IDENTITY_TOKEN_FILE`) or an operator allow-list of directories, and drop the per-account field. +- verdict: REJECTED — the API boundary already enforces the operator allow-list the fix asks for: validateAWSWebIdentityTokenFile (internal/api/validation.go:105-119) rejects ".." and requires the path to start with /var/run/secrets/eks.amazonaws.com/serviceaccount/ or /var/run/secrets/kubernetes.io/serviceaccount/ (awsWebIdentityTokenFilePrefixes, :51-54), and it runs via validateAWSAuthMode:419 from validateCloudAccountRequest on create (handler_accounts.go:316), update (:558) and registration (handler_registrations.go:262), so /proc/self/environ and arbitrary host paths are unreachable from tenant data; the resolver check at resolver.go:266 is a second, weaker layer, not the only one. + +### A03-018 The IMDS block in the Azure secrets client checks literal IPs only +- category: security +- severity: low +- location: internal/secrets/azure_resolver.go:32 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `blockIMDSDialer` compares the dial `host` string against two literal addresses. A hostname that resolves to 169.254.169.254 (attacker DNS, or a `169.254.169.254.nip.io`-style name), the GCP metadata name `metadata.google.internal`, or a redirect to such a name is dialled normally because resolution happens after the check. The guard therefore blocks only the most literal spelling of the SSRF it documents. +- evidence: + ```go + host, _, err := net.SplitHostPort(addr) + if err != nil { + host = addr + } + if imdsAddresses[host] { + return nil, fmt.Errorf("connection to metadata endpoint %s is blocked", host) + } + return d.inner.DialContext(ctx, network, addr) + ``` +- suggested fix: Resolve the host first (or wrap `Control` on the dialer) and reject link-local (169.254.0.0/16, fe80::/10) and the metadata hostnames on the resolved address; share the implementation with `providers/azure/internal/httpclient`, which solves the same problem. +- verdict: REJECTED — the guard is literal-only as described (azure_resolver.go:32-41), but the client it protects talks only to the vault named by the operator's AZURE_KEY_VAULT_URL env var (secrets/resolver.go:58-81) with secret names as path segments (azure_resolver.go:103, :121), so no tenant-controlled host reaches this dialer and the hostname-resolution bypass has no attacker input; the suggested sibling providers/azure/internal/httpclient delegates to pkg/httpclient, which uses the same literal-address map (pkg/httpclient/httpclient.go:29-47), so it does not solve the stated problem either. + +### A04-014 TransitionExecutionStatus RETURNING projects raw cancelled_by while every other execution read projects COALESCE(canceled_by, cancelled_by) +- category: correctness +- severity: low +- location: internal/config/store_postgres.go:1083-1090 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: rows canceled by new code carry `canceled_by` only (`CancelExecutionAtomic`). A transition out of such a row (any caller listing a canceled spelling in `fromStatuses`) returns `CancelledBy == nil`; a follow-up `SavePurchaseExecution` of that struct writes `cancelled_by = NULL` and, while `canceled_by` is retained, any consumer of the returned record loses the actor. The pgxmock guard `TestPGXMock_GetExecutionByID_ProjectsCoalescedCancelledBy` covers only `GetExecutionByID`, so this projection drifted unnoticed. +- evidence: + ```go + RETURNING plan_id, execution_id, status, step_number, scheduled_date, + notification_sent, approval_token, recommendations, + total_upfront_cost, estimated_savings, completed_at, error, expires_at, + cloud_account_id, source, approved_by, cancelled_by, capacity_percent, + ``` +- suggested fix: project `COALESCE(canceled_by, cancelled_by) AS cancelled_by` in the RETURNING list and extend the pgxmock projection test to every method that feeds `scanExecutionRows`. +- verdict: REJECTED — the projection asymmetry is real (raw `cancelled_by` at internal/config/store_postgres.go:1086 versus `COALESCE(canceled_by, cancelled_by)` at 1554), but the failure needs a transition *out of* a canceled row and no caller passes a canceled spelling in `fromStatuses`: the thirteen call sites are internal/purchase/approvals.go:332, internal/purchase/scheduled_fire.go:81, internal/purchase/reaper.go:175, internal/purchase/manager.go:179/452/498, internal/api/handler_history.go:247 and internal/api/handler_purchases.go:271/295/320/445/827/1448. Both writers of canonical-only `canceled_by` land the row in `canceled` (internal/config/store_postgres.go:1158-1166 and 1214-1222), which is terminal, and `SetCancelledBy` writes both columns (1128-1132), so no record with a lost actor is ever returned. + +### A05-016 `fanOutPerAccount` omits the post-`Wait` ctx check its sibling fan-out has +- category: concurrency +- severity: low +- location: internal/scheduler/scheduler.go:600 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: Every goroutine returns nil to isolate per-account failures, so `g.Wait()` can never surface `context.Canceled`. `collectAllProviders` recognises this and adds an explicit `ctx.Err()` check after its own Wait (scheduler.go:339); `fanOutPerAccount` does not. Today the parent's check catches the cancellation before `persistCollection` runs, so nothing is evicted, but the per-provider collect functions in between (`collectAWSRecommendations` and its Azure/GCP twins) each return `(recs, outcome.SucceededAccountIDs, nil)` — a nil error alongside a roster that authorizes stale-row eviction — for a sweep that was cut short. The safety of the eviction rests entirely on a check two frames up. +- evidence: + ```go + if waitErr := g.Wait(); waitErr != nil { + // Goroutines return nil to isolate per-account failures; non-nil is unexpected. + logging.Warnf("fanOutPerAccount: errgroup.Wait returned unexpected error: %v", waitErr) + } + return all, outcome + ``` +- suggested fix: after Wait, if `ctx.Err() != nil`, clear `outcome.SucceededAccountIDs` (move those IDs to `IncompleteAccountIDs`) so a canceled sweep can never authorize eviction regardless of what the caller does. +- verdict: REJECTED — a guard exists on the only path out. fanOutPerAccount is reached exclusively through collectAWSRecommendations (scheduler.go:495), collectAzure (:802) and collectGCP (:875), all called from collectProviderRecommendations (:419), whose sole caller is the collectAllProviders goroutine at scheduler.go:319; that function does check `ctx.Err()` after Wait and returns early (:338-340), before CollectRecommendations reaches persistCollection (:211). No canceled sweep can authorize eviction, and the finding concedes this itself — what remains is a defense-in-depth preference with no failure scenario. + +### A06-002 scheduledAuthMiddleware falls open to a pass-through when the validator is nil +- category: security +- severity: high +- location: internal/server/http.go:207 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: any `Application` not built by `NewApplicationFromDeps` has `scheduledAuth == nil`, and every `/api/scheduled/*` route (collect_recommendations, process_scheduled_purchases, ladder_run, ri_exchange_reshape — all money-moving) is then served with no authentication at all. This is the exact pattern the project memory records as fixed on PR #620 (`feedback_fail_closed_middleware`), so it is a regression, not a known gap. It is currently load-bearing for tests: internal/server/http_test.go:304 builds `app := &Application{}` and asserts HTTP 200 for `POST /api/scheduled/collect_recommendations` with no `Authorization` header, so the suite passes *because* of the bypass and would catch nothing if a future refactor exposed it in production wiring. +- evidence: + ```go + func (app *Application) scheduledAuthMiddleware(next http.Handler) http.Handler { + if app.scheduledAuth == nil { + return next + } + return app.scheduledAuth.Middleware(next) + } + ``` +- suggested fix: return a handler that writes 403 when `app.scheduledAuth == nil`, and change the tests to inject `scheduledauth.New(Config{Mode: ModeDisabled})` explicitly. +- verdict: REJECTED — the nil state is unreachable in production wiring: the only non-test `Application` literal is internal/server/app.go:502 inside `NewApplicationFromDeps`, which returns early unless `initScheduledAuth` yields a non-nil validator (internal/server/app.go:396-402 and 361-381), `LoadConfig` errors when `SCHEDULED_TASK_AUTH_MODE` is unset, and `scheduledauth.New` never returns `(nil, nil)` (internal/server/scheduledauth/validator.go:109-133); the fail-open default remains a hardening gap but has no reachable failure scenario at this commit. + +### A06-011 SMTP TLS guard protects credentials only, not the token-bearing body it claims to protect +- category: security +- severity: low +- location: internal/email/smtp_sender.go:327 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the comment above `dispatchSMTP` states the guard exists because a non-TLS connection "exposes credentials ... and token-bearing message bodies in cleartext", but the condition is `auth != nil && !s.useTLS && !s.allowInsecure`. With `Username`/`Password` empty (an open relay or a local MTA) `auth` is nil, the guard does not fire, and `smtp.SendMail` sends purchase-approval and revocation tokens over plaintext port 25. Reachability today is limited: both factory paths hardcode `Port: 587, UseTLS: true`, so only a direct `NewSMTPSender` caller (tests, a future wiring) can hit it — the finding is that the guard's axes do not match its stated purpose. +- evidence: + ```go + if auth != nil && !s.useTLS && !s.allowInsecure { + return fmt.Errorf("SMTP auth over non-TLS connection is refused: ...") + } + ``` +- suggested fix: drop `auth != nil` from the condition so any non-TLS send is refused unless `AllowInsecure` is set. +- verdict: REJECTED — the mismatch between the comment and the condition is real (internal/email/smtp_sender.go:311-330), but no reachable caller can produce a non-TLS send: all four constructors hardcode `Port: 587, UseTLS: true` with non-empty credentials (internal/email/factory.go:140-148, 196-203, 221-230, 239-248), and `NewSMTPSender` additionally forces `UseTLS = true` whenever `Port == 587` (smtp_sender.go:82-84), so there is no in-tree path with `auth == nil` and `useTLS == false`. + +### A06-021 Org discovery silently drops malformed accounts and imports suspended ones as enabled +- category: correctness +- severity: low +- location: internal/accounts/org_discovery.go:48 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `ListAccounts` returns every member account including ones with `Status: SUSPENDED` or `PENDING_CLOSURE`; each is turned into `CloudAccount{Enabled: true, AWSAuthMode: "role_arn"}`, so CUDly will keep trying to assume a role into a closed account on every collection cycle and surface the failures as account errors. Separately, an entry with a nil `Id` or `Name` is skipped with a bare `continue`, which contradicts the type's own doc ("Discovery is all-or-nothing ... There is no partial-success path") — the caller receives a short list with no indication anything was dropped. `"aws"` and `"role_arn"` are bare literals where the repo has typed provider constants. +- evidence: + ```go + for _, a := range page.Accounts { + if a.Id == nil || a.Name == nil { + continue + } + accounts = append(accounts, config.CloudAccount{ + Provider: "aws", ExternalID: *a.Id, Name: *a.Name, + Enabled: true, AWSAuthMode: "role_arn", + }) + ``` +- suggested fix: skip accounts whose `Status` is not `ACTIVE`, and return an error (or a reported count) instead of silently discarding entries with missing fields. +- verdict: REJECTED — the stated consequence cannot occur: the sole consumer of the discovery result overwrites both fields before persisting, setting `member.Enabled = false` and `member.AWSAuthMode = ""` on every row (internal/api/handler_accounts.go:1591-1595, reached via `runOrgDiscovery` at handler_accounts.go:1540-1549), so a SUSPENDED account is written as a disabled row awaiting operator review and no collection cycle ever assumes a role into it. The residual sub-point survives as a nit: the nil-`Id`/`Name` `continue` (internal/accounts/org_discovery.go:49-51) does silently shrink a list the type documents as all-or-nothing (org_discovery.go:14-16). + +### A06-023 Exchange notification branch re-derives the manual mode from a bare literal +- category: correctness +- severity: low +- location: internal/server/handler_ri_exchange.go:165 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `result.Mode` is `cfg.RIExchangeMode` copied straight from the DB (`pkg/exchange/auto.go:157`), and `pkg/exchange` owns the manual/auto decision at auto.go:248. This handler re-derives it from the literal `"manual"` even though `exchange.ExchangeModeManual` exists (auto.go:607). If the two ever disagree — a stored value of `"Manual"`, or a future rename — the engine still produces `Pending` outcomes with live approval tokens while this branch falls through to the `Completed`/`Failed` test, which is false for a pending-only run, so no email is sent at all and the approvals expire unnoticed. +- evidence: + ```go + if result.Mode == "manual" && len(result.Pending) > 0 { + data.RecipientEmail = notifyEmail + err = app.Email.SendRIExchangePendingApproval(ctx, data) + } else if len(result.Completed)+len(result.Failed) > 0 { + ``` +- suggested fix: branch on `len(result.Pending) > 0` (the actual precondition) and use `string(exchange.ExchangeModeManual)` if the mode check is kept. +- verdict: REJECTED — the two sides cannot disagree, because the engine's own gate uses the identical bare literal: `if params.Config.Mode == "manual"` (pkg/exchange/auto.go:248) is the only place `result.Pending` is ever appended to (auto.go:253), and `result.Mode` is `params.Config.Mode` copied verbatim (auto.go:156). So `len(result.Pending) > 0` implies `result.Mode == "manual"`; a stored `"Manual"` would make the engine run in auto mode and produce no pendings at all, not the described silent-drop. The unused `ExchangeModeManual` constant (auto.go:607) makes this a style inconsistency with no failure scenario. + +### A07-020 Unmapped service types are sent to Cost Explorer as a raw slug +- category: silent-fallback +- severity: medium +- location: providers/aws/recommendations/converters.go:27 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `getServiceStringForCostExplorer` returns `string(service)` for anything not in its switch, and `client.go:150` puts that value straight into `GetReservationPurchaseRecommendationInput.Service`. A service type the switch does not cover reaches CE as, for example, `"savingsplans-compute"` instead of a real SERVICE dimension value. CE matches nothing and returns an empty recommendation set with no error, so the run reports "no savings available" for a service that was never actually queried. +- evidence: + ```go + case common.ServiceMemoryDB: + return "Amazon MemoryDB Service" + default: + return string(service) + } + ``` +- suggested fix: add an error-returning variant used by the request builder (the file already establishes this pattern with `convertPaymentOptionE` / `convertTermInYearsE`) so an unmapped service fails loud instead of querying a nonexistent dimension value. +- verdict: REJECTED — client.go:145 diverts every Savings Plans slug through `common.IsSavingsPlan` (pkg/common/types.go:148-157) before reaching the input built at :150, the cited value `"savingsplans-compute"` is not a real constant (the slug is `"savings-plans-compute"`, types.go:115), and every remaining entry in `GetSupportedServices` (provider.go:426-446) has an arm in the switch, so the `default` at converters.go:27 is unreachable from either call site (the other, usage_history.go:151, additionally skips SP recs because they carry no `ResourceType`). + +### A07-024 Marketplace listing accepts an empty idempotency token +- category: money-path +- severity: medium +- location: providers/aws/services/ec2/client.go:1043 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `CreateMarketplaceListing` validates `PriceSchedule`, `ReservedInstancesID` and `InstanceCount` but not `ClientToken`, then sends `aws.String("")`. `ClientToken` is the only thing preventing a retried listing request from creating a second listing for the same RI. A caller that forgets to populate it gets no error, and a retry after a timeout lists the RI twice. +- evidence: + ```go + input := &ec2.CreateReservedInstancesListingInput{ + ReservedInstancesId: aws.String(req.ReservedInstancesID), + ClientToken: aws.String(req.ClientToken), + InstanceCount: aws.Int32(req.InstanceCount), + PriceSchedules: awsSchedule, + } + ``` +- suggested fix: reject an empty `ClientToken` alongside the other three boundary checks, matching the non-empty idempotency-source rule the EC2 and RDS purchase paths already follow. +- verdict: REJECTED — the missing check is real (providers/aws/services/ec2/client.go:1021-1029 validates the other three fields only), but the sole production caller always supplies `uuid.New().String()` (internal/api/handler_marketplace.go:250) and concurrent creates are serialized by `ClaimMarketplaceListingSlot` (:240-246), whose comment records that AWS gets a fresh token per call anyway, so the described double-listing is unreachable; this is a latent boundary-validation gap, not a live money-path defect. + +### A07-025 GetServiceClient discards the plan-type lookup error and silently falls into umbrella mode +- category: silent-fallback +- severity: medium +- location: providers/aws/provider.go:476 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `PlanTypeForServiceType` returns `("", false)` for anything outside its four cases and the `ok` value is thrown away. If a fifth SP slug is added to this `case` list without a matching arm in `PlanTypeForServiceType`, `pt` is `""`, which the savingsplans constructor treats as umbrella mode: `GetExistingCommitments` stops partitioning and returns every plan type, and `resolveSPPlanType` stops rejecting a plan-type mismatch (client.go:271). The scope guard that exists to stop the wrong product being bought is disabled by a value nobody checked. +- evidence: + ```go + case common.ServiceSavingsPlansCompute, + common.ServiceSavingsPlansEC2Instance, + common.ServiceSavingsPlansSageMaker, + common.ServiceSavingsPlansDatabase: + pt, _ := savingsplans.PlanTypeForServiceType(service) + return NewSavingsPlansClient(regionalCfg, pt), nil + ``` +- suggested fix: check the boolean and return an error when the slug has no mapped plan type, so the two lists cannot drift into an unintended umbrella client. +- verdict: REJECTED — all four slugs in the case list at providers/aws/provider.go:471-474 have matching arms in `PlanTypeForServiceType` (providers/aws/services/savingsplans/client.go:84-96), so `pt` is never empty at this commit and the umbrella-mode consequences at client.go:120-123 and :271 are unreachable; the discarded `ok` is a genuine drift hazard but describes a future edit, not a present failure. +- severity-adjusted: low — no reachable failure scenario, so this is hardening rather than a live silent fallback. + +### A07-030 GetDefaultRegion falls back to a hardcoded us-east-1 +- category: silent-fallback +- severity: low +- location: providers/aws/provider.go:419 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: when no region is set on the provider, on the config, or in the SDK chain, the function returns `"us-east-1"` rather than reporting that no region could be resolved. `GetServiceClient` copies whatever region it is handed into the regional config, so a misconfigured environment silently targets us-east-1 for offering lookups and purchases instead of failing. +- evidence: + ```go + if p.IsConfigured() && p.cfg.Region != "" { + return p.cfg.Region + } + return "us-east-1" + ``` +- suggested fix: add an error-returning variant used by the purchase and offering paths, keeping the string default (if at all) only for display contexts. +- verdict: REJECTED — the hardcoded return at providers/aws/provider.go:419 is real (as is the dead `IsConfigured()` branch at :417, which repeats the check at :412), but a repo-wide grep finds no production caller of `GetDefaultRegion` anywhere: only the interface declaration at pkg/provider/interface.go:27 and test mocks, so the claimed propagation into `GetServiceClient`'s regional config, offering lookups and purchases does not exist. + +### A08-011 An empty term silently buys a one-year Savings Plan while an empty payment option correctly fails closed +- category: money-path +- severity: high +- location: providers/azure/services/savingsplans/client.go:490 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `PurchaseCommitment` calls `toAzureTerm(rec.Term)`, which maps `""` to `TermP1Y`. A recommendation row persisted before the term column was populated reaches the purchase path with `Term == ""` and buys a real one-year commitment nobody asked for. The sibling parser in the same package family refuses the equivalent input explicitly and explains why: `BillingPlanForPaymentOption("")` returns an error rather than defaulting because "Azure's own default is upfront and would charge the whole commitment immediately" (reservations/purchase.go:115). `reservations.ParseTermYears` has the same `""` → 1 mapping (purchase.go:70). +- evidence: + ```go + func toAzureTerm(term string) (armbillingbenefits.Term, error) { + switch term { + case "1yr", "1", "P1Y", "": + return armbillingbenefits.TermP1Y, nil + ``` +- suggested fix: remove `""` from the one-year case in both `toAzureTerm` and `reservations.ParseTermYears` and return the same style of explicit error `BillingPlanForPaymentOption` returns for an empty payment option. +- verdict: REJECTED — no caller can deliver `Term == ""` to the purchase path: internal/purchase/execution.go:1041 builds it as `fmt.Sprintf("%dyr", rec.Term)` from an int column (a 0 would yield "0yr", which `toAzureTerm` already rejects), and every Azure-sourced recommendation goes through `normaliseTerm` (internal/recommendations/converter.go:341), which maps nil/empty to "1yr" before the rec exists. The `""` case at savingsplans/client.go:490 and reservations/purchase.go:70 is an unreachable allowance, not a live silent purchase. + +### A08-025 The VM reservation purchase body uses a raw "Shared" literal for a typed SDK enum +- category: hygiene +- severity: low +- location: providers/azure/services/compute/client.go:448 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: `appliedScopeType` is written as the string `"Shared"` while `armreservations.AppliedScopeTypeShared` exists and is used correctly elsewhere in the same package (exchange_operations.go:448). A casing or spelling change in a future API version breaks the purchase body at runtime with no compile error, and the scope is not derived from the recommendation's own `ExtractedFields.Scope`, which the extractor already populates (internal/recommendations/converter.go:51). +- evidence: + ```go + "appliedScopeType": "Shared", + "renew": false, + ``` +- suggested fix: use `string(armreservations.AppliedScopeTypeShared)`. +- verdict: REJECTED — the literal at compute/client.go:448 is byte-identical to `armreservations.AppliedScopeTypeShared` (armreservations@v1.1.0/constants.go:21), so no failure scenario exists today, and the "not derived from ExtractedFields.Scope" half is refuted by the recommendation filter the client actually sends, `"properties/scope eq 'Shared'"` (compute/client.go:199), which is exactly the rationale recorded at providers/azure/internal/recommendations/converter.go:44-50. + +### A09-017 An empty persisted Details payload yields a zero-valued typed pointer that the provider fills with Linux/default +- category: money-path +- severity: high +- location: pkg/common/service_details_codec.go:98 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: for a `purchase_executions` row whose `Details` JSONB is NULL or absent — a legacy row, or any writer that failed to populate it — `DecodeServiceDetailsFor("ec2", nil)` returns `&ComputeDetails{}` with `Platform`, `Tenancy` and `Scope` all empty. The comment states the downstream `buildOfferingFilters` then substitutes `Platform=Linux/UNIX, Tenancy=default, Scope=Region`. A Windows EC2 recommendation re-driven from such a row buys a Linux Reserved Instance that covers none of the Windows demand. That is the exact mis-purchase class issue #453 was opened for, still reachable by design on the legacy path. +- evidence: + ```go + if len(raw) == 0 || bytes.Equal(raw, jsonNullBytes) { + // Legacy row or genuinely absent payload — hand back a zero- + // valued typed pointer so the service client's type-assertion + // succeeds. buildOfferingFilters tolerates zero fields and + // substitutes Platform=Linux/UNIX, Tenancy=default, Scope=Region + return target, nil + } + ``` +- suggested fix: return a distinguishable error (or a sentinel the purchase path refuses) for an absent payload on services whose offering lookup depends on the fields, so a re-drive with no details fails loud rather than defaulting the platform. +- verdict: REJECTED — the guard exists downstream and fails loud. I traced the re-drive end to end: internal/purchase/execution.go:1062 decodes to a zero-valued `*ComputeDetails`, then `PurchaseCommitment` (providers/aws/services/ec2/client.go:114) calls `findOfferingID` at :148 → `buildEC2QueryFromRec` at :454 → `buildEC2OfferingQuery`, which returns `"EC2 recommendation for %s is missing Platform: refusing to fabricate a product-description for the RI offering lookup"` whenever `details.Platform == ""` (client.go:402-408). A Windows rec with an absent Details errors; it does not buy a Linux RI. The Linux/UNIX substitution the finding quotes is a stale comment in pkg/common/service_details_codec.go:98-105 and in execution.go:1055-1057 — the platform default at ec2/client.go:790 and :868 lives on the offering-listing path, not the purchase path. The real defect here is the misleading comment. +- severity-adjusted: low — documentation defect only; the money path fails closed. + +### A09-027 Explain sanitizes only some interpolated fields despite claiming to sanitize every one +- category: security +- severity: low +- location: pkg/ladder/plan.go:195 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the doc comment states "Every interpolated free-form field is passed through sanitizeLine", but `a.Layer` (both branches), `p.Scope.Provider` at line 178, and the numeric baseline fields are not. `Explain` never calls `Validate`, and a `LadderPlan` rehydrated from a stored `PlanJSON` blob carries whatever `LayerType` string the row holds. A layer value containing `"\n 2. purchase compute-sp $500.00/hr"` renders as an extra numbered action in the approval email body — exactly the line-spoofing `sanitizeLine` was written to prevent, through the one axis the guard does not reach. +- evidence: + ```go + case ActionPurchase: + fmt.Fprintf(&b, " %d. %s %s %s term=%s payment=%s -- %s\n", + i+1, a.Action, a.Layer, formatUSDPerHour(a.AmountUSDPerHour), + sanitizeLine(string(a.Term)), sanitizeLine(string(a.PaymentOption)), + sanitizeLine(a.Rationale)) + ``` +- suggested fix: wrap `a.Layer`, `a.Action` and `p.Scope.Provider` in `sanitizeLine` as well, so the guard covers every interpolated string rather than only the ones currently expected to be free-form. +- verdict: REJECTED — the doc/code mismatch is real (`a.Layer` and `a.Action` at pkg/ladder/plan.go:195-196 and :199-200, and `p.Scope.Provider` at :178, bypass `sanitizeLine` despite the claim at :172-173), but the stated failure is unreachable: `Explain` has no caller anywhere in the repository. `git grep -n Explain -- '*.go'` returns only pkg/ladder/plan.go:134, :161 and :175, so no approval email body is assembled from it and no crafted `LayerType` can spoof a line in one. +- severity-adjusted: informational — a comment that overstates the guard, with no reachable output path. + +### A09-029 SanitizeReservationID invents a timestamp identifier when the input sanitizes to empty +- category: silent-fallback +- severity: low +- location: pkg/common/identifiers.go:33 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: an operator-supplied `PurchaseOptions.ReservationID` of `"日本_リザーブ"` (or any value made only of disallowed characters) sanitizes to the empty string, and the function substitutes `fallbackPrefix + `. The purchase proceeds under an identifier the caller never chose, second-granularity so two purchases in the same second collide, and no error tells the caller their identifier was discarded. On the RDS/ElastiCache/MemoryDB path the reservation ID is the server-side dedupe key, so a substituted value also breaks re-drive idempotency. +- evidence: + ```go + s = strings.Trim(s, "-") + if s == "" { + s = fallbackPrefix + strconv.FormatInt(time.Now().Unix(), 10) + } + return s + ``` +- suggested fix: return an error for a caller-supplied identifier that sanitizes to empty, keeping the generated fallback only for the internal builder paths that have no caller-supplied value to preserve. +- verdict: REJECTED — the timestamp substitution is in the code as quoted (pkg/common/identifiers.go:32-35), but no caller can reach it with a value that sanitizes to empty. The only provider call site is providers/aws/services/rds/client.go:207, reached only when `opts.ReservationID != ""`, and the only production writer of that field is cmd/multi_service_helpers.go:246, which passes `generatePurchaseID` — a machine-composed ASCII string always containing the "ri"/"dryrun" prefix, an RFC-style timestamp and a uuid8 suffix (cmd/main.go:313-330). There is no operator-controlled ReservationID input. The "RDS/ElastiCache/MemoryDB path" claim is also wrong: `git grep -n SanitizeReservationID` shows RDS is the only provider caller. +- severity-adjusted: informational — unreachable from every current caller. + +### A11-003 A stored max_amount of 0 renders as blank and re-saves as no spending cap +- category: money-path +- severity: high +- location: frontend/src/groups/groupModals.ts:269 @ 3c0f8ac94048a2c36fce5ccddee54e6c4849a5cd +- failure scenario: the Max Amount input is populated with `String(permission?.constraints?.max_amount || '')`. For a permission constrained to `max_amount: 0` (spend nothing) the `||` treats 0 as absent and the box renders empty. `collectPermissions` then gates on `if (maxAmount)` (groupModals.ts:450), so the empty box contributes nothing and the saved permission carries no `max_amount` at all. On the backend an absent cap is unrestricted, so a cosmetic rename of the group converts a $0 ceiling into unlimited spend. This is the exact widening the file's own comments guard against for the list constraints, but `unrepresentableDimensions` only walks `LIST_CONSTRAINT_DIMENSIONS` (groupModals.ts:319) and never inspects `max_amount`. No test covers `max_amount: 0`; groups.test.ts only exercises 5000, 10000, Infinity and a negative. +- evidence: + ```typescript + // groups/groupModals.ts:268-270 +