diff --git a/docs/net/architecture.md b/docs/net/architecture.md
index 6a90368ee..562de7ff6 100644
--- a/docs/net/architecture.md
+++ b/docs/net/architecture.md
@@ -496,26 +496,30 @@ sequenceDiagram
### Push-Based Status Aggregation
-Node agents periodically push their status to the controller, which caches and broadcasts updates to dashboard clients via WebSocket.
+Node agents publish summaries by default. The controller separates observed
+overview facts from expiring, explicitly requested diagnostic payloads.
+Routine cluster APIs and broadcasts contain only summaries, including when
+legacy agents publish full status. The diagram illustrates the HTTP path;
+agents can also publish over authenticated WebSockets.
```mermaid
sequenceDiagram
participant Agent as Node Agent
participant Controller as Controller (Leader)
- participant Cache as Status Cache
+ participant Cache as Overview Cache
participant WS as WebSocket Clients
- loop Every 10 seconds
- Agent->>Agent: getNodeStatus()
(snapshot state, WireGuard, routes, pingmesh)
+ loop Configured publication cadence
+ Agent->>Agent: Inspect overview facts without BPF traversal
Agent->>Controller: POST /status/push
(gzip JSON, Bearer token)
Controller->>Controller: Authenticate token
- Controller->>Cache: Store status
+ Controller->>Cache: Store metadata and observed counts
end
- loop Every 3 seconds
- Controller->>Cache: Fetch all cached statuses
- Controller->>Controller: Build ClusterStatusResponse
- Controller->>WS: Broadcast delta-compressed update
+ loop Cluster summary updates
+ Controller->>Cache: Read overview state
+ Controller->>Controller: Build ClusterSummary
+ Controller->>WS: Broadcast summary snapshot or delta
end
```
@@ -524,8 +528,50 @@ sequenceDiagram
- On leader election, the controller cleans up stale `v1/Endpoints` resources left by previous controller versions to prevent kube-proxy routing to dead pods
- HTTP POST is sent asynchronously with an atomic in-flight guard to prevent ticker drift
- Status collection uses a snapshot-and-release pattern to minimize lock hold time
+- Summary inspection preserves observed peer health, route counts/mismatches,
+ bootstrap/CNI errors, and interface metadata without building outbound
+ peer/route/BPF arrays
- A pod informer watches `unbounded-net-node` pods to display pod name, restart count, and age in the dashboard
+### Explicit detail request lifecycle
+
+The dashboard requests details only through **Load data** or **Refresh**.
+CLI named-node show commands reuse an unexpired snapshot or initiate one
+request; `--refresh` asks for fresh data. Concurrent requests for a node coalesce.
+The controller uses an active authenticated capable WebSocket, otherwise an
+immediate HTTP pull. A failed pull leaves a command for the next authenticated
+status POST acknowledgment. This HTTP-first behavior is independent of
+background summary pulling and of the node's publication mode.
+
+Commands carry an opaque request ID and a fixed deadline (120 seconds by
+default). Nodes collect and reply immediately rather than waiting for a
+publication tick. Detail acknowledgments do not acknowledge routine
+publications or advance their revision bases. Duplicate commands can resend
+the collected response without recollection; duplicate and late responses
+cannot revive expired work or renew cached details.
+
+Accepted snapshots are bound to Node UID and expire 300 seconds after actual
+receipt by default. Reads and summary updates do not renew expiry. Legacy full
+publications update the same completed cache association, without completing
+an unrelated pending refresh. Legacy delta bases and their validation memos
+live only in this TTL store: expiry or explicit snapshot replacement requires
+an ordinary full resync before deltas can resume.
+
+The node cache, cluster cache, global JSON, and broadcast replay history own
+only overview data. Request records and waiting callbacks keep metadata rather
+than duplicate detailed results. Viewer queues carry metadata and resolve
+details at write time; detailed writes are canceled at snapshot expiry or
+leadership loss. Browser snapshots have their own fixed advertised expiry and
+are discarded without automatic collection.
+
+The leader owns all requests and detail storage in memory. Node replacement,
+deletion, shutdown, or leadership loss invalidates the relevant state. TTL
+limits duration, not peak memory during bursts; no entry-count or concurrency
+budget is imposed. Released references become eligible for garbage collection,
+not guaranteed immediate RSS reduction. Existing HTTP compressed-body and
+WebSocket frame limits remain in force; oversized diagnostics fail explicitly
+rather than truncate data.
+
## State Management
### Controller State
@@ -548,14 +594,17 @@ graph TD
end
subgraph "Dashboard State"
- D1[Status Cache
node -> pushed status]
+ D1[Overview Cache
node -> metadata and counts]
D2[Pod Informer
unbounded-net-node pods]
- D3[WebSocket Broadcaster
delta-compressed updates]
+ D3[WebSocket Broadcaster
summary updates and history]
+ D4[Detail Cache
UID-bound payload and legacy base with TTL]
+ D5[Detail Requests
IDs, deadlines, and lifecycle metadata]
end
A1 --> |"Thread-safe
mutex"| A2
A1 --> |"Thread-safe
mutex"| A3
- D1 --> |"sync.Map"| D3
+ D1 --> |"summary snapshots"| D3
+ D5 --> |"explicit single-node result"| D4
```
### Node Agent State
diff --git a/docs/net/operations.md b/docs/net/operations.md
index 735d7cfd1..1fb177d6c 100644
--- a/docs/net/operations.md
+++ b/docs/net/operations.md
@@ -213,11 +213,11 @@ The dashboard displays:
- **Overview**: Cluster health summary with node counts, site counts, and gateway status
- **Sites**: All configured sites with node counts and health indicators
- **Nodes**: Detailed list of all nodes with filtering, sorting, and pagination
- - Tunnel peer status (WireGuard peers or eBPF tunnel endpoints)
+ - Aggregate tunnel peer counts and health
- Gateway health for each node
- Site membership
- Pod name, restart count, and pod age
- - Click on any node to view detailed status including stale data warnings
+ - Select a node, then click **Load data** to inspect peers, routes, or BPF entries
The dashboard uses **WebSocket** for real-time updates with delta compression, falling back to HTTP polling when WebSocket is unavailable. Features include:
- Live status updates via push-based architecture (node agents push status to controller)
@@ -229,6 +229,51 @@ The dashboard uses **WebSocket** for real-time updates with delta compression, f
The dashboard does not render a site connectivity graph or connectivity matrix.
Use the Site summaries and filtered node list to inspect individual resources.
+### Explicit node diagnostics
+
+Selecting a node does not collect diagnostics. **Load data** reuses an unexpired
+snapshot when available; **Refresh** requests a fresh collection. Opening a
+dialog, reconnecting, receiving a summary, or reaching expiry never requests
+details automatically. The dialog reports loading, expiry, and errors rather
+than showing empty tables as if no configuration exists.
+
+The default detail lifetime is 300 seconds from receipt. The browser drops the
+snapshot at its advertised expiry. A failed Refresh may leave the previous
+still-valid snapshot visible with its age and the request error. Global
+**Cluster Status JSON** exports summaries only; full JSON belongs to a loaded
+single-node snapshot.
+
+CLI named-node show commands use the same cache/request lifecycle:
+
+```bash
+# Reuse valid cached data, or request and wait for one node's diagnostics
+kubectl unbounded net node show peers
+kubectl unbounded net node show routes
+kubectl unbounded net node show bpf
+kubectl unbounded net node show json
+
+# Force fresh collection without bypassing a request already in flight
+kubectl unbounded net node show --refresh
+
+# Watch only cluster summaries
+kubectl unbounded net node list --watch
+```
+
+`node show --watch` is not supported: use summary list/watch or repeat an explicit
+show with `--refresh`. There is no full-cluster diagnostic refresh command.
+The `json` show format remains raw full node JSON, not the request envelope.
+
+Requests use an active, authenticated, capable node WebSocket when possible.
+Otherwise they try direct HTTP pull first, even if background pulling is
+disabled. If that fails, the next authenticated node status POST response
+delivers a pending command. Agents collect and reply immediately on receipt,
+not at the next periodic publication tick. All attempts share the original
+120-second request deadline.
+
+If both outbound publishers are disabled, direct node HTTP diagnostics remain
+available. A failed pull cannot use a POST fallback until publication resumes;
+the controller does not enable either publisher automatically.
+
### Health Endpoints
#### Controller Health
@@ -261,10 +306,10 @@ curl http://:9998/status/json
#### Controller Status Endpoints
```bash
-# Cluster status JSON (leader only)
+# Cluster summary JSON (leader only, no peer/route/BPF arrays)
curl http://:9999/status/json
-# Per-node status (supports ?live=true for force pull)
+# Per-node raw diagnostics (cached by default; ?live=true requests fresh data)
curl http://:9999/status/node/
# Aggregated API status push endpoint (if enabled)
@@ -272,6 +317,25 @@ curl http://:9999/status/node/
# WebSocket: wss:///apis/status.net.unbounded-cloud.io/v1alpha1/status/nodews
```
+The asynchronous API uses the same viewer authorization and leader routing:
+
+| Request | Purpose |
+|---------|---------|
+| `POST /status/node//details` with `{"forceRefresh":false}` | Reuse cache or start/join a request |
+| `POST /status/node//details` with `{"forceRefresh":true}` | Start/join a fresh request |
+| `GET /status/node//details?requestId=` | Read that request's result |
+
+Responses identify the node and request and return `pending` (202), `complete`
+(200), `expired` (410), `unavailable` (404), or `retryable` (503). Completed
+`details` include `collectedAt`, `receivedAt`, `expiresAt`, and full `status`.
+The request deadline governs pending work; snapshot expiry governs completed
+data. Reads never extend either lifetime. A pending response can include the
+failed pull's error while still waiting for a POST-delivered reply.
+
+For aggregated access, prefix these paths with
+`/apis/status.net.unbounded-cloud.io/v1alpha1`. Do not assume a raw controller
+URL bypasses viewer authentication.
+
#### Gateway Health
```bash
@@ -625,6 +689,7 @@ The node agent status server listens on port 9998 (configurable via `--health-po
| `/healthz` | GET | Liveness probe (API connectivity) |
| `/readyz` | GET | Readiness probe |
| `/status/json` | GET | Full node status (tunnels, routes, health, BPF entries) |
+| `/status/summary` | GET | Overview metadata and observed counts, without diagnostic arrays |
| `/status` | GET | Human-readable status page |
| `/metrics` | GET | Prometheus metrics |
diff --git a/docs/net/troubleshooting.md b/docs/net/troubleshooting.md
index 7d016b50f..1eddf1878 100644
--- a/docs/net/troubleshooting.md
+++ b/docs/net/troubleshooting.md
@@ -12,8 +12,13 @@ kubectl unbounded net node show # detailed node status
kubectl unbounded net node show bpf # BPF trie entries
kubectl unbounded net node show routes # kernel routes
kubectl unbounded net node show json # raw JSON status
+kubectl unbounded net node show --refresh # request fresh diagnostics
```
+Named-node show commands reuse valid cached details unless `--refresh` is set.
+List/watch stay summary-only; `node show --watch` is rejected rather than
+starting continuous diagnostic collection.
+
### unroute -- BPF map inspector
Available inside the node agent container:
@@ -66,10 +71,41 @@ port conflict, or flap backoff).
Each node agent exposes status on port 9998:
- `GET /status/json` -- full node status (peers, routes, health checks, BPF entries)
+- `GET /status/summary` -- observed overview facts without diagnostic arrays
- `GET /metrics` -- Prometheus metrics (healthcheck_packets_sent/received, etc.)
## Common Issues
+### Details are not loaded, expired, or unavailable
+
+- In the dashboard, select a node and click **Load data**. Selection, reconnect,
+ and summary updates do not load peers, routes, or BPF data.
+- An expired snapshot is intentionally removed. Click **Load data** again, or
+ use `node show --refresh`, rather than interpreting missing details as an
+ empty dataplane configuration.
+- Check the per-node request error and deadline. Without an active capable
+ WebSocket, the controller tries HTTP first, then waits for an authenticated
+ status POST to carry the pending command if that pull fails.
+- If both node publishers are disabled and HTTP is unreachable, the request
+ cannot complete through a polling fallback. Restore the intended transport
+ or use node-local diagnostics; no publisher is enabled automatically.
+- A controller restart or leadership change loses in-memory requests and
+ snapshots. Retry explicitly against the current leader. A replaced Node UID
+ also invalidates previous details.
+- `status-summary-unsupported` indicates an incompatible controller. Upgrade
+ the controller first, or explicitly configure `node.statusDetailMode: full`
+ and restart agents for node-publishing rollback.
+- Older clients expecting full arrays from controller `/status/json` must
+ migrate to summaries plus the named-node detail API. Full publishing mode
+ does not restore detailed bulk exports.
+
+Summary freshness and detail freshness are separate. Default detail expiry is
+300 seconds after actual receipt, and the request deadline is 120 seconds
+across all delivery attempts. A routine summary or repeated read extends
+neither. Expiry releases retained objects for garbage collection; it does not
+promise an immediate process RSS decrease or cap memory during a burst of
+explicit requests.
+
### Peers showing Down / 0 online
1. Check BPF program attachment: `tc filter show dev unbounded0 egress`