Skip to content

Commit f734ba5

Browse files
Sg312icecrasher321waleedlatif1
authored
feat(mothership): v1.0.0 (#8208)
* mothership agent cli: grep --in accepts the world/resource path a match line prints * workflow lint, block catalog, tables, files, sandbox tools: fixes from exploration run 5 Lint: trigger-category blocks are entry blocks (schedule was an orphan); references inside Function code are checked with the runtime tokenizer; the agent-cli lint and grep read the draft state instead of the sanitized export. Catalog: operation inputs publish sub-block ids only (no canonical-param aliases); trigger-category blocks expose their trigger-mode fields. Tables: dispatch processedCount counts unlimited dispatches; json-language code fields accept objects. Files: restore keeps the folder when it still exists. Sandbox tools: outputTable failures report the files already written and the per-language result shape; the file writer keeps the computed result. * sim-cli: --output json prints the API data verbatim; the single-key unwrap stays a table-only convenience * route and connection validation, trigger defaults on add, table folder restore, dispatch listing, group column attach, import block summary; logs query and deps engines; run-tool errors and payload compaction; stale tests aligned Editing: router routes validated as {id?, title, value} with unknown keys named; malformed connections reported instead of dropped; sub-block value() defaults evaluated on add so a webhook gets its token. Tables: folder restore no longer 500s (lock inside the row transaction), completed dispatches are listed, groups attach to existing output columns. Workflows: import answers with its blocks. Mothership engines: logs query resolves bare paths under output, marks missing paths, drops non-executed runs under --where; deps lists graph predecessors and child return shapes; run-from-block validation errors reach the agent; run payloads compact input.code. CLI generated types regenerated. * sim-cli: the operations apply confirmation describes the batch as written, not a delete * sandbox export keeps the code's result; --trigger implies --manual; deps mock skeleton; run tools lift the terminal output and take select; lint all-clear names what it checked; docs chunker keeps code spans * sim-cli: runs get --select-output takes block names; groups delete says it removes the output columns and their data; a publish that lands on public auth prints a note * null-safe block.error, strict start-trigger coercion, lint checks reference paths against output schemas; logs stats segments; requiredWhen and model ids in the agent catalog; live tag usage; one row-filter grammar; minted service-account credential ids; webhook delivery URLs on deployment status; mv semantics for folder moves Executor: <block.error> resolves empty on a block that succeeded so a shared error collector is writable; a number input given a non-numeric string fails the run at start naming the field; a regression test pins that failure traces carry child spans. Lint: a reference whose first path segment is not on the block's effective outputs (or its responseFormat schema) is an unknown-field finding. API: logs stats honours segmentCount and omits empty buckets unless includeEmpty; block detail publishes requiredWhen instead of flattening conditional requirements and always lists model options with hosted marks; knowledge tag usage counts through the document slot; every table rows filter accepts a bare condition; service-account credentials mint their id; deployment status lists each webhook's delivery URL; a folder move into an existing folder moves it inside. OpenAPI and CLI generated types regenerated. * same-workspace import keeps or warns about workspace bindings; generic required-field lint; folder moves to the root; empty-graph and no-entry lint notes; span durationMs; MCP tools reported on undeploy and listed inactive; paged connector types and workspaces; grep --in refuses unknown selectors; run_function names exported files; run tools surface the failing block's error Export takes includeWorkspaceBindings for same-workspace round trips and import answers with warnings naming every block whose required binding was stripped. Lint reports a missing required field for every block type (the knowledge base id was skipped), notes an empty graph and a graph with no entry block, and the trace spans carry durationMs. Undeploy answers with the MCP tools it archived and the tools list shows them inactive. connector-types list is paged (25, summary projection, detail=full) and workspaces list defaults to 25. The agent CLI's grep refuses a prefix or unknown --in selector with the accepted forms; a run_function that exported files and wrote a table says both; a run tool whose executor result carries no message uses the failing block's own error. * table groups default to the deployed version and refuse a dispatch without one; column rename reports unmigrated Table-block filters; lint checks table fields against the live schema; runCount counts every settled run; runs cancel is honest about no-ops; dry-run apply answers with previewBlockIds; conditionResult is the chosen test's boolean with selectedTitle; workflows run --select-output keys blockOutputs by the selector and rejects unknown heads; knowledge search exposes rankScore and rank; the multi-trigger error names workflows state get; the Knowledge block declares cost, tokens, model; media writes refuse to fork a moved folder; grep --in knowledge points at knowledge search * agent cli grep: the searchable text leads with the resource's name and description * Catalog hides sunset blocks unless includeSunset; enrichment get returns the group's run state and outputs; knowledge upload prints the document fields; docs chunks titled by page not nav link; logs carry hasHandledErrors and opt-in handledErrorRuns; outputTable receipt keeps stdout * Sandbox exports decode by the path rule the reader used: a .jpg declared without a format is written as image/jpeg bytes, not as its base64 text under the json format (thumbnails opened as raw text) * Chat uploads resolve by uploads/<name> for reads, sandbox mounts, and image references (never listings or writes); an unresolved reference image fails the call instead of rendering without it; copilot session-sandbox calls are priced like Function-block sandboxes and report the raw cost beside the billed one; CLI docs regenerated * grep: resolve --in against listings, fetch only the named resources A bare --in selector (e.g. --in agent, --in file_v5) materialized every searched world to find one block: 65 block details plus every workflow state per call. On dev that took 18-34s per grep and tripped the per-user rate limit. Each world now has a cheap index (its listing) and a per-resource fetch; a --in search reads the indexes and fetches only the matches. Whole-world searches are unchanged. * mothership: model-facing strings name only what exists on the CLI surface The copilot back-derives fixes from the strings sim hands it, so a retired name becomes a wrong instruction to the user (dev 2026-09-03: a save_upload mention became "drag the photo into the files panel"). Every model-facing string audited today now names the current surface: - upload notice: a chat upload lives at uploads/<name> and is not a workspace file; workflows import takes --workflow (there is no --file); a .zip is mounted and unzipped in the sandbox rather than a files unzip path that does not resolve - table import resolves uploads/<name> directly (includeChatUploads) instead of pointing at the retired save_upload tool - function-execute / generate-image: outputs get, files ls, files restore and tables list replace read/grep/glob/restore_resource and Go VFS meta.json paths - process-contents: browser/terminal pointers no longer name browser_* or a terminal tool this surface does not have; docs fallback names docs search - lint/deps usage strings use the plural workflows group - integration credential error names credentials list * docs: every model-searchable page states the current surface The copilot answers users from these pages (docs search), so a wrong sentence becomes a wrong instruction. Audit fixes: - chat: uploads live under uploads/, not the Files panel; on-demand workspace reach instead of a per-message snapshot; Chat runs Opus 4.8; connectors and deletes/restores as they actually behave; login pages via the shared browser - cli: --select-output works on sync runs (only --async conflicts); runs get takes the same block selectors; secrets set needs --scope; workflows run runs the deployment or --manual; sim profiles honours --output; which groups have no singular alias; the embedded CLI has no profile, login, or config; @path reads the chat sandbox - model defaults are claude-sonnet-5 (guardrails has none); API examples use https://www.sim.ai/api/v2 (apex 301s POSTs into GETs) - quick reference / shortcuts: only affordances that exist (no Deploy tab, no workspace duplicate/export, Mod+B not Mod+E, variables under the ⋯ menu, <variable.name> syntax) - editor read-only rules match round-trip-safety.ts; card links fixed - sim-cli: runs-get selector error and the Settings label match the product; batch delete describes its ids as deleted, not updated; docs/api regenerated * mothership: prepare_file_edit create + new_file pass input validation The workspace-file handler has always supported operation=create with a new_file target, but the generated tool schema (from the copilot catalog) forbade both, so the watched write could never create a file: a "save this report as a workspace file" turn died on `/operation must be equal to one of the allowed values` (dev 2026-09-03). Regenerated from the catalog that now declares create, new_file, and fileName; regression test on the validator. * mothership: the embedded CLI answers v2 in-process; grep memoizes both platform catalogs The embedded CLI and the agent-cli engines were typed v2 clients pointed at the server's own URL: every tool call was a network round trip through the proxy, API-key auth, the abuse rate limits, and the proxy body ceiling. A grep over one block definition cost 8-34s and tripped the per-key limit (dev 2026-09-03). - sim-cli: ResolvedProfile / EmbeddedCliIdentity take an optional transport; the client sends through it instead of fetch. The installed CLI never sets it. - sim: an in-process transport resolves a v2 path against a generated route table (scripts/generate-v2-route-table.ts, check:v2-route-table in check:audits) and invokes the route handler directly. The request is marked internal through a WeakSet — not a header — so admission still authenticates it but skips the pre-auth IP bucket and the per-key rate limits, which exist for callers on the wire. Contracts, use cases, presenters, and error envelopes are the ones the network path runs. Anything outside the v2 table falls through to fetch. - grep engine: blocks and tools (5,000 built-ins at 100 a page) are memoized per workspace; an exact block or tool id (--in file_v5) resolves from that corpus without listing any workspace world; name fragments still index every world and fetch only the matches. * mothership: take the embedded CLI off the in-process transport until the dev hang is understood Every cli_grep on dev timed out at 60s from the first deploy of 99aad28. The transport, marker, route table, and grep memo stay; the embedded identity goes back to the HTTP path while the hang is diagnosed from sim's logs. * grep: share in-flight builds, bound nested requests process-wide, cache file text by version The first dev turn on the in-process transport fired eight world-wide greps at once. Each rebuilt every world concurrently on the one process serving the chat — 24 block-catalog listings, 150 tool-catalog pages, 130 file reads in sixty seconds — until all eight timed out. Over the wire the same fan-out had been spread across tasks and throttled by the network. - concurrent greps now await the build already in flight for a (world, workspace) - nested requests are bounded to eight for the whole process, not per world - platform corpora (blocks, tools) are kept an hour, not ten minutes - a file's text is cached by id + updatedAt + size, so repeat greps re-read only files that changed * grep: an exact block id answers before the tool catalog is built * grep: honour -A and -B context flags The engine read only -C; -A and -B were ignored silently, so `-A 40` returned the bare match line and the model concluded context flags were unreliable (dev 2026-09-03). A bad value is refused like -C's. * sandbox pricing: normalize COST_MULTIPLIER at the boundary Under skipValidation the value arrives as the raw string it was deployed with; the sandbox lease's finite check rejected it, so on dev every run_code failed as "Boot sandbox" (Sandbox pricing multiplier must be a finite nonnegative number). getCostMultiplier now goes through envNumber, as the env module's own rule says numeric overrides must. * sandbox pricing: cost-multiplier imports env relatively — next.config.ts loads env-flags outside alias resolution * mothership: pack a workspace inventory with every chat request The agent spent nine tool rounds and ~20K tokens learning what exists in the workspace before an orientation task could start. The chat request now carries a compact inventory (contracts ChatRequest.inventory): workflows, tables, knowledge bases, files, skills, custom tools, MCP servers, credentials, and secret names — names and ids, one page per world, capped worlds named in `truncated`. It is read through the same use cases the v2 listings run, under the caller's session principal, so authorization is unchanged; a world that fails to list is left empty and logged rather than failing the turn. The worker renders it once per turn as a request-local message. * grep: the tool catalog is one use-case call when the engine has the caller's principal The tools world paged 5,000 built-in tools through the route stack at 100 a page — fifty round trips, each re-resolving the gate and walking the registry — which was the whole cold cost of a world-wide grep. The embedded bridge now resolves the delegation key to its principal exactly as the v2 surface does and hands it to the engines; with it, the tools world reads the catalog through listCatalogTools in one call — same authorization, same projection. Without a principal the client path stays. * mothership: background tasks — subscriptions, wake turn, task pill (21 §6) Sim side of copilot background tasks. copilot_task_subscriptions (migration 0315) records a worker task's watch on a workflow execution; the logging session's completion attempt posts every subscribed run's outcome to the worker. POST /api/mothership/tasks/subscribe and POST /api/mothership/wake (internal key): the wake runs the inbox's headless lifecycle under the task id as message id, announces itself on the chat status channel so an open chat reconnects, persists the user message with origin 'task', and resolves the pill in the arming turn. Client: a 'task' content block end to end (run handler, turn-model node, both serializer directions, persisted normalizer — which also flattened plan blocks to text on reload —, display block, segment, TaskPill) and a muted TaskNotificationRow for task-origin messages. The stream validator's run kinds now derive from the generated contract; the generated mothership-stream-v1 gains task_armed / task_delivered. * mothership tasks: resolve the pill by task id (jsonb text spacing broke the exact-key match) * task pill: icons from @sim/emcn (lucide-react is not an app dependency; the Docker build failed) * mothership: checkpoint Sim interaction and workbench boundary repairs Capture the accumulated Mothership control, recovery, sandbox session, embedded CLI, and standalone diagnostic repairs. Companion worker checkpoint: c8245400. Validation and limitations are recorded in the Mothership revamp ledger through iteration 38; the latest expanded Sim diagnostic selection passed 1231 tests. The new table-mount audit is excluded and will follow separately. * mothership: authorize table snapshot mounts and honor folder references Move Mothership table snapshot resolution, provenance checks, and bounded storage reads into a registered application operation. Reuse the table VFS resolver for folder paths, recheck current actor access before materialization, and project safe mount errors. Preserve CSV transport, budgets, default paths, and workflow/public API behavior. Validation: 517 related tests, app/auth type checks, API and agent-CLI boundary checks. * mothership: stream workbench downloads with atomic publication * mothership: stream uploads and imports from workbench snapshots * mothership: bind upload requests to the snapshot lifetime * fix(mothership): preserve and bound embedded CLI output * fix(mothership): preserve file provenance through host CLI reads * test(mothership): guard current CLI catalog entry points * fix(mothership): make search current, isolated, and explicit about coverage * fix(mothership): bind file mounts to one canonical version * feat(mothership): persist private upload source classification * fix(mothership): preserve classification across workbench file copies Record encrypted source evidence against host-streamed bytes and physical workbench identity. Bind completed upload streams through the authorized application operation, sealing pending provenance with the completion lease so recovery preserves the first result. * fix(mothership): retain source evidence for saved CLI output Classify settled stdout through the existing workspace-file helper and persist evidence against the streamed workbench bytes. Preserve completed command outcomes on publication failure and verify copied output across fresh invocations and pending sibling tools. * test(mothership): exercise CSV composition through the Sim runtime * test(mothership): connect CLI workflow runs to report publication * fix(mothership): preserve raw CLI data with scoped presentation * fix(mothership): expose canonical workflow identity for watches * fix(mothership): preserve watches across transcript replay * fix(mothership): convert code execution timeouts once * fix(mothership): preserve child failures across transcript recovery * test(mothership): distinguish replay from executable handoff * test(mothership): verify persisted run diagnostics through CLI * test(mothership): verify workflow execution log readback * test(mothership): verify controller execution and resume recovery * fix(mothership): acknowledge text received before stream retries * fix(mothership): reconcile response text across reconnects * test(mothership): verify automatic relay interruption recovery * fix(mothership): unify replayable child lifecycle handling * test(mothership): verify child recovery after worker exit * fix(mothership): distinguish tool history from executable handoffs * fix(mothership): acknowledge received activity between stream legs * test(mothership): recover missing initial relay frames * test(mothership): cover overlapping relay attachments * fix(mothership): observe results owned by overlapping controllers * fix(mothership): preserve long CLI command budgets * test(mothership): verify table pipeline against real execution * fix(mothership): recover one-shot response delivery * fix(mothership): align one-shot MCP discovery guidance * fix(mothership): make inbox attachments readable * fix(mothership): scope upload names to the requesting chat * fix(mothership): preserve worker history when forking chats * fix(mothership): cancel only response-owned workflow executions * fix(mothership): serialize workflow pickup with run Stop * fix(mothership): track browser workflow execution through Stop * fix(mothership): require settlement for Stop before chat attachment * fix(mothership): preserve canonical history when stopping a response * fix(mothership): handle abort failures while Stop history persists * fix(mothership): retain queued Stop dependencies through retries * fix(mothership): recover outgoing handoffs through the durable queue * fix(mothership): bind queued retries to the active Stop operation * test(mothership): exercise physical workflow edits and oracle execution * test(mothership): verify extracted workflow invocation through saved logs * test(mothership): verify build-only consent through physical run history * fix(mothership): retain terminal failures in saved chats * fix(mothership): retain active runs through worker connection loss * fix(mothership): persist Stop before worker delivery * fix(mothership): recover interactive streams after Sim process loss * fix(mothership): durably admit chat turns before worker dispatch * fix(mothership): recover lost tool handlers with fenced execution leases * fix(mothership): finalize steered answers by turn identity * fix(mothership): separate committed tool outcomes from stream publication * fix(mothership): serialize saved resource panel changes * fix(mothership): persist resource effects once before publication * fix(mothership): preserve complete saved resource addresses * fix(mothership): preserve saved table views in chat context * fix(mothership): follow the current embedded table view * fix(mothership): preserve exact table selection identities * fix(mothership): attach usable file folder references * fix(mothership): render saved file folder resources * Revert "fix(mothership): render saved file folder resources" This reverts commit d785434. * fix(mothership): align attachment context with available tools * fix(mothership): authorize MCP calls as the chat subject * fix(mothership): keep workbench policy out of the public CLI * fix(mothership): preserve composed code export outcomes * fix(mothership): retain completed files after export failure * test(mothership): verify hosted workbenches and task compositions * ci: restore staging jobs and keep Mothership acceptance local * feat(mothership): receive private Sim controls over outbound transport * fix(mothership): bound stream reconnection attempts * fix(mothership): export chat sandbox files through workspace storage * fix(cli): isolate embedded output from host process logging * fix(mothership): preserve file intents and workflow result contracts * perf(mothership): trace CLI and tool persistence boundaries * fix(mothership): preserve resource effects and execution recovery * fix(mothership): reconcile staging contracts and consolidate migrations * fix(mothership): share and verify billing callback contracts * fix(mothership): preserve catalog curation and embedded CLI access * fix(mothership): mount code secrets explicitly * fix(mothership): own sandbox profile at the code tool boundary * fix(workflows): validate field values before full-state writes * fix(chat): collapse main tool groups into action summaries * fix(mothership): restore workspace API key dispatch * fix(chat): label workflow dry runs as validation * feat(chat): expose model, reasoning effort and Astra Fast controls * fix(chat): align composer model and effort controls * Group Mothership tools by activity and show scoped resource names * Report requested Mothership model without legacy telemetry defaults * Restore Mothership Slack bot connection from stored secrets * Preserve execution events for deployed Copilot workflow runs * Expose active workflow API details to Mothership * Restore desktop parity and preserve interrupted chat progress * Keep Mothership resource panels and activity groups in sync * Show active tools before completed activity summaries * Fix shared agent tool execution and discovery boundaries * Port Assistant to shared Mothership runtime and repair resource recovery * Refresh CLI docs and desktop validation after staging rebase * Restore staging image inventory commands * Fail migrations correctly and resolve empty direct database URLs * Match chat attachment guidance to supported file capabilities * fix(ci): retain the latest published desktop prereleases * fix(mothership): read file contents through one adaptive command * fix(chat): simplify pending tool activity label * feat(chat): inspect scratch files and preserve inline images Resolve workspace, upload, and sandbox references through authorized file readers. Preserve explicit Markdown images privately with each chat turn without creating workspace files, and keep visual bytes out of bounded UI tool-status events. Keep sandbox CLI credentials scoped to the active callback so file provenance checks cover data entering the workbench. * fix(chat): keep tool failures in expanded history * fix(db): avoid repeated dev index builds and search backfills * fix(chat): preview inline images with shared lightbox * feat(mothership): operate across organization workspaces * feat(mothership): manage settings and workflow file inputs * fix(mothership): preserve stopped chat admission and queued corrections * fix(mothership): show sequential activity intent and concrete tools * fix(mothership): expose canonical model hints in internal discovery * fix(mothership): unify scoped CLI and organization resource panels * feat(assistant): connect and use personal organization integrations * feat(chat): unify Home modes and search resource tabs * feat(chat): tailor Home controls and starters to the selected mode * fix(mothership): preserve file upload and preview resource context * fix(mothership): stabilize Home search panels and add optional Fast Search * fix(mothership): preserve staging authorization and resumed billing * fix(db): make development schema pushes noninteractive (#7906) * fix(db): resolve dev schema push column ambiguity * refactor(db): generalize push rename handling * docs(db): describe schema push reconciliation steps * feat(search): unify Home search levels and composer * fix(db): retire file size bridge before forced schema push * fix(ci): restore test isolation and memory file scope checks * fix(ci): keep metadata imports pure and narrow component graphs * fix(mothership): refresh client settings after scoped mutations * fix(db): repair partial member sync status on existing databases * fix(home): reuse the shared loading fallback * fix(mothership): remove eager catalogs and bulky admission saves * fix(mothership): keep reads out of resource panels * test(search): assert lazy integration catalog in headless chat * fix(search): label balanced search level Auto * fix(mothership): provide authorized context for chat titles * fix(search): restore standalone search and refine chat mode controls * fix(search): align mode selector with the standard search icon * fix(mothership): resolve integration discovery service identities * fix(desktop): reuse Sim login for authenticated browser previews * fix(chat): preserve configured origin in deployment URLs * fix(search): recover GitHub installation setup and verify ownership * fix(search): retain sync failures and resume reconnected accounts * fix(chat): restore thinking indicator after tools finish * fix(mothership): expose and validate tool attachment identities * feat(mothership): inspect configured workflow tool bindings * feat(mothership): support dynamic Sim Chat tool modes * fix(mothership): continue Sim Chat block conversations * feat(mothership): create organization workspaces through CLI * perf(desktop): observe browser state in action calls * fix(mothership): repair background wakes and simplify watch activity * fix(mothership): clarify tool permissions and folder context * fix: reconcile staging rebase contracts and migrations * fix(db): reconcile staging search migrations in dev * fix(chat): preserve live turns and recover interrupted streams * fix(workflows): scope custom block schemas during authoring * fix: reconcile staging search and file contracts after rebase * style(db): format rebased migration snapshots * fix(slack): deliver Sim Chat text and tool progress reliably * test(slack): verify Agent and Sim Chat thinking parity * fix(slack): bound each long-answer append request * fix(slack): continue long streams across message limits * fix(slack): recover explicitly rejected size overflows * fix(mothership): simplify saved results display label * fix(slack): clean up long-response continuations * feat(mothership): simplify Build picker to Astra effort levels * fix(mothership): hide mode selection after the first message * feat(search): add feature-flagged live provider retrieval * fix(search): scope repository retrieval and add Coda MCP OAuth * fix(search): restore account hooks client directive * Fix GitHub connected-account tools in Search Assistant * Add organization-scoped live search and admin-managed GitLab * Add date-aware live enterprise search and bounded listings * Polish live Search results, citations, and effort controls * fix(mothership): preserve hosted service billing through cutover Record hosted integration, sandbox and media spend through a trusted durable service outbox independent of tool output, cancellation and worker delivery. Preserve service billing under model BYOK and restore title admission. Use workspace BYOK settings and fresh credentials, retain execute transport, and restore the locked actor-authority check for billing-sensitive membership removal. Add migrations and local HTTP/PostgreSQL billing acceptance coverage. Validation: 1,453 billing tests passed with 30 environment-gated skips, app and dev infrastructure types passed, and local SQL charge proofs passed. * fix(search): restore live citations and streamed source panel * fix(search): replace live results tab with cited sources * feat(search): simplify federated sources and add scoped secrets * fix(search): make service account setup directly accessible * feat(mothership): add dev-gated Plan mode across chat surfaces * fix(search): preserve account connection links in Slack * fix(search): remove indexing language from live search * fix(search): keep admin setup out of member integrations * fix(search): route Search MCP through live retrieval * fix(plan): normalize the runtime feature flag to a boolean * ci(dev): deploy Trigger tasks independently of app images * feat(mothership): set organization Generic Secrets in chat * fix(search): verify federated connectors and align setup docs * fix(search): remove stale Slack indexing copy * docs(search): document Generic Secrets and clarify source setup * fix(mothership): allow authorized organization code secrets in Build and Plan * fix(search): restore live integration connection chips * fix(search): include Gmail in live connection inventory * fix(search): expand composer around image attachments * fix(mothership): restore saved Plan chat history * fix(mothership): address review findings and CI regressions * test(deploy): align promotion expectation with Trigger CLI * improvement(desktop): make saved password autofill field-aware * fix(mothership): preserve trace filters and shared app workers * fix(desktop): dismiss credential picker when browser hides * feat(mothership): simplify model controls and add native desktop files * fix(desktop): preserve password picker focus and lifetime * chore(mothership): align model and deployment test expectations * fix(mothership): resolve model and Plan controls from AppConfig * fix(desktop): contain native imports and preserve UTF-8 pages * fix(credentials): preserve unrelated setup controls * feat(browser): file transfer, popups, PDFs, click gestures, and dialog answers for the browser agent - Upload workspace, chat-upload, or granted local files into a page's file input (DOM.setFileInputFiles on the input a ref, label, or drop zone resolves to), staged privately and bound to the claimed tool call; save completed downloads into workspace files. New /api/desktop/tool/file route with application use cases. - Same-session popups adopt Chromium's WebContents so window.opener works (sign-in and connect flows); a page that closes itself leaves cleanly and agent work returns to its opener. - PDFs render: enable Electron's internal PDF plugin and let its packaged viewer resources past the network guard. - Clicks take button, clickCount, and modifiers; agent right-clicks never open the native menu. Actions answer confirm/alert dialogs via dialog: {accept}. press_key gains F1-F12, Insert, and repeat; scroll and hover accept post-action observation. * chore(chat): provide runtime flags in composer tests * fix(browser): isolate actions and bound file transfers * fix(browser): retain upload input identity through dispatch * fix(mothership): validate complete billing protocol headers * fix(browser): preserve upload dispatch before acknowledgement * fix(browser): distinguish unconfirmed upload outcomes --------- Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com> Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai> Co-authored-by: Waleed Latif <walif6@gmail.com>
1 parent e2cdee6 commit f734ba5

2,334 files changed

Lines changed: 215907 additions & 64586 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.agents/skills/add-connector/SKILL.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,12 @@ argument-hint: <service-name> [api-docs-url]
66

77
# Add Connector Skill
88

9+
## Choose the connector runtime first
10+
11+
For **Sim Search**, use the live provider workflow in [the federated Search developer guide](../../../apps/sim/lib/sim-search/live/README.md#adding-a-live-search-connector). Its browser-safe provider catalog owns provider IDs, API origins, credential aliases, and account modes; its typed runtime registry requires both search and read handlers. `ConnectorMeta` remains the owner of logos and setup fields. Member mode has no admin resource filters. Service mode requires independent live source verification and shared selectors. Do not implement a Search source by adding a crawler, embeddings, or a scheduled ACL build.
12+
13+
The ingestion instructions below apply to **ordinary knowledge-base connectors** and the explicit legacy Search backend (`SIM_SEARCH_LIVE=false`). If a provider supports both, implement and test both runtimes; adding `search: true` to metadata alone does not implement federated search. Preserve indexing documentation and behavior for those KB/legacy callers.
14+
915
You are an expert at adding knowledge base connectors to Sim. A connector syncs documents from an external source (Confluence, Google Drive, Notion, etc.) into a knowledge base.
1016

1117
## Your Task

‎.agents/skills/validate-connector/SKILL.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,12 @@ argument-hint: <service-name> [api-docs-url]
66

77
# Validate Connector Skill
88

9+
## Identify the runtime under review
10+
11+
For **Sim Search**, validate the [live provider registration and access pipeline](../../../apps/sim/lib/sim-search/live/README.md#adding-a-live-search-connector): catalog and metadata parity, both search/read handlers, current member grants, service-source restrictions, safe scoped references, pagination, provenance, and provider failure behavior. Test real localhost setup/search/read with authorized fixtures when available, and distinguish those results from mocked provider tests. Live Search must not enqueue content indexing or background ACL/directory builds; GitLab still computes request-time permissions or uses current CSV grants.
12+
13+
The ingestion-specific checks below apply to ordinary workspace KB connectors and legacy Search selected with `SIM_SEARCH_LIVE=false`. Keep those checks for providers supporting both runtimes; do not require a live-only provider to implement content hashes, ingestion cursors, embeddings, or stored ACL snapshots.
14+
915
You are an expert auditor for Sim knowledge base connectors. Your job is to thoroughly validate that an existing connector is correct, complete, and follows all conventions.
1016

1117
## Your Task

‎.github/CONTRIBUTING.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -256,6 +256,8 @@ If you prefer not to use Docker. **All commands run from the repository root unl
256256

257257
For ad-hoc schema iteration during development you can also use `bun run db:push` from `packages/db`, but `db:migrate` is the canonical command for staging and production. `db:push` reconciles directly to the current schema without running versioned migration guards. For disposable local/dev databases, `bun run db:push --force` accepts Drizzle's data-loss prompts, including column drops.
258258

259+
`db:push` treats added and removed columns, tables, and other schema objects as separate creations and deletions. It never infers a rename. For an intentional rename during local development, run `bun run db:push --interactive-renames` in a terminal and select the old object in Drizzle's chooser. This flag does not approve data loss; `--force` controls that separately. After schema reconciliation succeeds, the wrapper reconciles credential policies and OAuth providers, then backfills search vectors. A failure stops subsequent steps. Staging and production changes still use reviewed versioned migrations with expand/contract deployment steps.
260+
259261
4. **Run the Development Servers:**
260262

261263
```bash

‎.github/actions/docker-build/action.yml‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -19,6 +19,10 @@ inputs:
1919
tags:
2020
description: Comma-separated list of tags to push.
2121
required: true
22+
build-args:
23+
description: Newline-separated Docker build arguments.
24+
required: false
25+
default: ''
2226
max-cache-size-mb:
2327
description: >-
2428
Layer cache to retain after this action prunes, in MB. Must stay above one
@@ -72,6 +76,7 @@ runs:
7276
platforms: ${{ inputs.platforms }}
7377
push: true
7478
tags: ${{ inputs.tags }}
79+
build-args: ${{ inputs.build-args }}
7580
provenance: false
7681
sbom: false
7782

@@ -177,5 +182,6 @@ runs:
177182
platforms: ${{ inputs.platforms }}
178183
push: true
179184
tags: ${{ inputs.tags }}
185+
build-args: ${{ inputs.build-args }}
180186
provenance: false
181187
sbom: false

‎.github/workflows/ci.yml‎

Lines changed: 26 additions & 19 deletions
Original file line numberDiff line numberDiff line change
@@ -103,7 +103,7 @@ jobs:
103103
echo "ℹ️ No comparable base commit; skipping desktop prerelease"
104104
exit 0
105105
fi
106-
if git diff --name-only "$BEFORE" HEAD | grep -qE '^(apps/desktop/|packages/desktop-bridge/|packages/browser-protocol/)'; then
106+
if git diff --name-only "$BEFORE" HEAD | grep -qE '^(apps/desktop/|packages/desktop-bridge/|packages/browser-protocol/|\.github/workflows/(ci|desktop-release)\.yml$)'; then
107107
echo "changed=true" >> "$GITHUB_OUTPUT"
108108
echo "✅ Desktop shell code changed"
109109
else
@@ -221,15 +221,18 @@ jobs:
221221
file: ${{ matrix.dockerfile }}
222222
platforms: linux/amd64
223223
tags: ${{ steps.login-ecr.outputs.registry }}/${{ steps.ecr-repo.outputs.name }}:${{ github.sha }}-dev
224+
build-args: |
225+
SIM_SEARCH_LIVE_DEFAULT=true
226+
MSHIP_PLAN_MODE_DEFAULT=true
224227
max-cache-size-mb: ${{ matrix.cache_mb }}
225228

226-
# Build and upload tasks alongside tests and images. The unpromoted version
227-
# cannot serve new runs; promote-images waits for it and successful migrations.
229+
# Staging/production coordinate task releases with app traffic cutover.
230+
# Dev tasks deploy independently in deploy-trigger-dev.yml.
228231
prepare-trigger:
229232
name: Prepare Trigger.dev
230233
if: >-
231234
github.event_name == 'push' &&
232-
(github.ref == 'refs/heads/main' || github.ref == 'refs/heads/staging' || github.ref == 'refs/heads/dev')
235+
(github.ref == 'refs/heads/main' || github.ref == 'refs/heads/staging')
233236
runs-on: ${{ (vars.CI_PROVIDER == '' || vars.CI_PROVIDER == 'blacksmith') && 'blacksmith-4vcpu-ubuntu-2404' || 'ubuntu-latest' }}
234237
timeout-minutes: 30
235238
outputs:
@@ -265,7 +268,6 @@ jobs:
265268
case "$GITHUB_REF" in
266269
refs/heads/main) TRIGGER_ENV=prod; TRIGGER_BRANCH='' ;;
267270
refs/heads/staging) TRIGGER_ENV=staging; TRIGGER_BRANCH='' ;;
268-
refs/heads/dev) TRIGGER_ENV=preview; TRIGGER_BRANCH=dev-sim ;;
269271
*) echo "ERROR: unsupported Trigger release ref: $GITHUB_REF" >&2; exit 1 ;;
270272
esac
271273
echo "environment=$TRIGGER_ENV" >> "$GITHUB_OUTPUT"
@@ -426,8 +428,9 @@ jobs:
426428
tags: ${{ steps.meta.outputs.tags }}
427429
max-cache-size-mb: ${{ matrix.cache_mb }}
428430

429-
# Promote the sha-tagged ECR images once tests, migrations, and the Trigger
430-
# upload pass. Pushing the ECR latest/staging tag is what triggers
431+
# Promote the sha-tagged ECR images once their build and migrations pass.
432+
# Staging/production also require the Trigger upload; dev never waits for it.
433+
# Pushing the ECR latest/staging/dev tag is what triggers
431434
# CodePipeline, so this seconds-long manifest retag is the deploy gate —
432435
# the image builds themselves run in parallel with the tests. A single job
433436
# (not a matrix) so all four sha manifests are verified before any tag
@@ -438,9 +441,9 @@ jobs:
438441
# Explicit results: see migrate's comment.
439442
if: >-
440443
!cancelled() && github.event_name == 'push' &&
441-
needs.prepare-trigger.result == 'success' &&
442444
(
443445
((github.ref == 'refs/heads/main' || github.ref == 'refs/heads/staging') &&
446+
needs.prepare-trigger.result == 'success' &&
444447
needs.migrate.result == 'success' &&
445448
needs.build-amd64.result == 'success') ||
446449
(github.ref == 'refs/heads/dev' &&
@@ -546,7 +549,7 @@ jobs:
546549
fi
547550
done
548551
549-
# Promote the parked Trigger.dev version after observing the ECS
552+
# Staging/production: promote the parked Trigger.dev version after observing the ECS
550553
# traffic cutover (CodeDeploy AllowTraffic on every target). The image retag
551554
# triggers the ECS pipeline; this job correlates it via the digest + retag epoch
552555
# (rejecting a stale execution reusing the digest) and promotes at cutover.
@@ -560,14 +563,13 @@ jobs:
560563
if: >-
561564
!cancelled() &&
562565
github.event_name == 'push' &&
563-
(github.ref == 'refs/heads/main' || github.ref == 'refs/heads/staging' || github.ref == 'refs/heads/dev') &&
566+
(github.ref == 'refs/heads/main' || github.ref == 'refs/heads/staging') &&
564567
needs.promote-images.result == 'success' &&
565568
needs.prepare-trigger.result == 'success' &&
566569
needs.promote-images.outputs.promoted == 'true'
567570
runs-on: ${{ (vars.CI_PROVIDER == '' || vars.CI_PROVIDER == 'blacksmith') && 'blacksmith-4vcpu-ubuntu-2404' || 'ubuntu-latest' }}
568-
# Leave setup/promotion headroom above the cutover poll (dev: 20 min;
569-
# staging/prod: 70 min, including a deploy queued behind a long bake).
570-
timeout-minutes: ${{ github.ref == 'refs/heads/dev' && 40 || 90 }}
571+
# Leave setup/promotion headroom above the 70-minute cutover poll.
572+
timeout-minutes: 90
571573
permissions:
572574
contents: read
573575
id-token: write
@@ -599,14 +601,13 @@ jobs:
599601
with:
600602
role-to-assume: ${{ github.ref == 'refs/heads/main' && secrets.AWS_ROLE_TO_ASSUME || github.ref == 'refs/heads/dev' && secrets.DEV_AWS_ROLE_TO_ASSUME || secrets.STAGING_AWS_ROLE_TO_ASSUME }}
601603
aws-region: ${{ github.ref == 'refs/heads/main' && secrets.AWS_REGION || github.ref == 'refs/heads/dev' && secrets.DEV_AWS_REGION || secrets.STAGING_AWS_REGION }}
602-
# Match each environment's session budget; both outlast their polls.
603-
role-duration-seconds: ${{ github.ref == 'refs/heads/dev' && 2400 || 5400 }}
604+
role-duration-seconds: 5400
604605

605606
# An unchanged tag may belong to a failed or still-running earlier deploy.
606607
# Verify its latest cutover rather than treating tag equality as success.
607608
- name: Wait for ECS traffic cutover
608609
env:
609-
OVERALL_TIMEOUT: ${{ github.ref == 'refs/heads/dev' && 1200 || 4200 }}
610+
OVERALL_TIMEOUT: 4200
610611
APP_IMAGE_CHANGED: ${{ needs.promote-images.outputs.app_image_changed }}
611612
DIGEST: ${{ needs.promote-images.outputs.app_image_digest }}
612613
PIPELINE: sim-${{ github.ref == 'refs/heads/main' && 'production' || github.ref == 'refs/heads/dev' && 'dev' || 'staging' }}-us-east-1-app-deployment
@@ -1311,24 +1312,30 @@ jobs:
13111312
name: Prune Desktop Prereleases
13121313
runs-on: ${{ (vars.CI_PROVIDER == '' || vars.CI_PROVIDER == 'blacksmith') && 'blacksmith-2vcpu-ubuntu-2404' || 'ubuntu-latest' }}
13131314
timeout-minutes: 5
1314-
needs: [publish-desktop-prerelease]
1315+
needs: [create-desktop-prerelease, publish-desktop-prerelease]
13151316
permissions:
13161317
contents: read
13171318
env:
13181319
GH_TOKEN: ${{ secrets.DESKTOP_RELEASE_TOKEN }}
13191320
GH_REPO: simstudioai/sim-desktop-releases
1321+
CURRENT_TAG: ${{ needs.create-desktop-prerelease.outputs.version }}
13201322
steps:
13211323
- name: Delete stale prereleases
13221324
run: |
1325+
set -euo pipefail
1326+
: "${CURRENT_TAG:?Current publication tag is required before pruning}"
13231327
if [ -z "$GH_TOKEN" ]; then
13241328
echo "::error::DESKTOP_RELEASE_TOKEN is required to prune desktop prereleases."
13251329
exit 1
13261330
fi
13271331
if [ "$GITHUB_REF" = "refs/heads/dev" ]; then CHANNELS='(dev|alpha)'; else CHANNELS='(staging|beta)'; fi
1328-
gh release list --limit 100 --json tagName,isPrerelease,isDraft,createdAt \
1329-
--jq "[.[] | select(.isPrerelease and (.isDraft | not) and (.tagName | test(\"-${CHANNELS}\\\\.\")))] | sort_by(.createdAt) | reverse | .[5:] | .[].tagName" |
1332+
# createdAt follows the tag's commit: published releases in the release-only
1333+
# repository can all share it. Retain by publication time, with deterministic ties.
1334+
gh release list --limit 100 --json tagName,isPrerelease,isDraft,publishedAt \
1335+
--jq "[.[] | select(.isPrerelease and (.isDraft | not) and (.tagName | test(\"-${CHANNELS}\\\\.\")))] | sort_by(.publishedAt, .tagName) | reverse | .[5:] | .[].tagName" |
13301336
while read -r TAG; do
13311337
[ -n "$TAG" ] || continue
1338+
[ "$TAG" != "$CURRENT_TAG" ] || continue
13321339
echo "Deleting stale prerelease $TAG"
13331340
gh release delete "$TAG" --cleanup-tag --yes
13341341
done
Lines changed: 75 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,75 @@
1+
name: Deploy Dev Tasks
2+
3+
# Independent of app CI: task packaging and deployment must never delay dev images.
4+
on:
5+
push:
6+
branches: [dev]
7+
8+
permissions:
9+
contents: read
10+
11+
# Serialize promotions without cancelling an external deployment in flight.
12+
# Pending pushes coalesce to the newest run while the current run finishes.
13+
concurrency:
14+
group: deploy-trigger-dev
15+
cancel-in-progress: false
16+
17+
jobs:
18+
deploy:
19+
name: Deploy Trigger.dev preview
20+
runs-on: ${{ (vars.CI_PROVIDER == '' || vars.CI_PROVIDER == 'blacksmith') && 'blacksmith-4vcpu-ubuntu-2404' || 'ubuntu-latest' }}
21+
timeout-minutes: 30
22+
steps:
23+
- name: Checkout code
24+
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
25+
26+
- name: Setup Bun
27+
uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
28+
with:
29+
bun-version: 1.4.1
30+
31+
- name: Cache Bun dependencies
32+
uses: actions/cache@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5
33+
with:
34+
path: |
35+
~/.bun/install/cache
36+
node_modules
37+
**/node_modules
38+
key: ${{ runner.os }}-bun-${{ hashFiles('**/bun.lock') }}
39+
restore-keys: |
40+
${{ runner.os }}-bun-
41+
42+
- name: Install dependencies
43+
run: bun install --frozen-lockfile --ignore-scripts
44+
45+
- name: Upload preview version
46+
id: deploy
47+
working-directory: ./apps/sim
48+
env:
49+
TRIGGER_ACCESS_TOKEN: ${{ secrets.TRIGGER_ACCESS_TOKEN }}
50+
TRIGGER_PROJECT_ID: ${{ secrets.TRIGGER_PROJECT_ID }}
51+
run: |
52+
set -euo pipefail
53+
: "${TRIGGER_ACCESS_TOKEN:?TRIGGER_ACCESS_TOKEN must be configured}"
54+
: "${TRIGGER_PROJECT_ID:?TRIGGER_PROJECT_ID must be configured}"
55+
bunx trigger.dev@4.5.16 deploy --env preview --branch dev-sim --skip-promotion
56+
57+
- name: Promote current dev preview
58+
working-directory: ./apps/sim
59+
env:
60+
GH_TOKEN: ${{ github.token }}
61+
TRIGGER_ACCESS_TOKEN: ${{ secrets.TRIGGER_ACCESS_TOKEN }}
62+
TRIGGER_PROJECT_ID: ${{ secrets.TRIGGER_PROJECT_ID }}
63+
VERSION: ${{ steps.deploy.outputs.deploymentVersion }}
64+
run: |
65+
set -euo pipefail
66+
if ! [[ "$VERSION" =~ ^[0-9]{8}\.[0-9]+$ ]]; then
67+
echo "ERROR: Trigger.dev did not report a valid deploymentVersion output" >&2
68+
exit 1
69+
fi
70+
CURRENT_SHA=$(gh api "repos/$GITHUB_REPOSITORY/git/ref/heads/dev" --jq '.object.sha')
71+
if [ "$CURRENT_SHA" != "$GITHUB_SHA" ]; then
72+
echo "::notice::Skipping superseded dev task version $VERSION"
73+
exit 0
74+
fi
75+
bunx trigger.dev@4.5.16 promote "$VERSION" --env preview --branch dev-sim

‎.github/workflows/migrations.yml‎

Lines changed: 3 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -64,6 +64,8 @@ jobs:
6464
MIGRATION_DATABASE_URL: ${{ inputs.environment == 'production' && secrets.MIGRATION_DATABASE_URL || inputs.environment == 'staging' && secrets.STAGING_MIGRATION_DATABASE_URL || '' }}
6565
ENVIRONMENT: ${{ inputs.environment }}
6666
run: |
67+
set -euo pipefail
68+
6769
if [ -z "$DATABASE_URL" ]; then
6870
echo "ERROR: no database URL secret resolved for environment '${ENVIRONMENT}'" >&2
6971
exit 1
@@ -73,16 +75,7 @@ jobs:
7375
echo "Dev environment — pushing schema directly (db:push)"
7476
# Dev deliberately forces direct schema reconciliation; staging and
7577
# production use guarded versioned migrations in the other branch.
76-
# drizzle-kit push needs a TTY to resolve ambiguous renames (--force only
77-
# covers data-loss). In CI it throws "Interactive prompts require a TTY
78-
# terminal" but still exits 0, so the job goes green without applying the
79-
# change. tee keeps the output live in the log; we then fail on drizzle's
80-
# own TTY error. A genuine non-zero exit already fails via `set -e`.
81-
bun run db:push --force < /dev/null 2>&1 | tee /tmp/db-push.log
82-
if grep -q "Interactive prompts require a TTY terminal" /tmp/db-push.log; then
83-
echo "ERROR: db:push needs an interactive rename decision; land it as a versioned migration instead of relying on push." >&2
84-
exit 1
85-
fi
78+
SIM_DEV_DB_PUSH=1 bun run db:push --force < /dev/null
8679
else
8780
echo "Applying versioned migrations (db:migrate)"
8881
bun run ./scripts/migrate.ts

‎.github/workflows/test-build.yml‎

Lines changed: 10 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -83,6 +83,12 @@ jobs:
8383
- name: Install dependencies
8484
run: bun install --frozen-lockfile --ignore-scripts
8585

86+
- name: Verify direct schema push compatibility
87+
working-directory: packages/db
88+
env:
89+
DB_PUSH_TEST_DATABASE_URL: postgresql://postgres:postgres@127.0.0.1:5432/postgres
90+
run: bunx vitest run scripts/push.postgres.test.ts
91+
8692
- name: Provision a fresh database through the supported command
8793
working-directory: packages/db
8894
run: |
@@ -233,14 +239,16 @@ jobs:
233239
if-no-files-found: ignore
234240
retention-days: 7
235241

236-
- name: Verify durable provenance, concurrent memory writes, and attachment replay
242+
- name: Verify durable provenance, concurrent memory writes, and browser download admission
237243
working-directory: apps/sim
238244
env:
245+
BROWSER_FILE_TRANSFER_TEST_DATABASE_URL: postgresql://postgres:postgres@127.0.0.1:5432/sim_auth_scim
239246
TABLE_PROVENANCE_TEST_DATABASE_URL: postgresql://postgres:postgres@127.0.0.1:5432/sim_auth_scim
240247
MEMORY_PROVENANCE_TEST_DATABASE_URL: postgresql://postgres:postgres@127.0.0.1:5432/sim_auth_scim
241248
AGENT_MEMORY_TEST_DATABASE_URL: postgresql://postgres:postgres@127.0.0.1:5432/sim_auth_scim
242249
run: >-
243250
bunx vitest run
251+
lib/mothership/async-runs/browser-download-claim.postgres.test.ts
244252
lib/table/rows/secret-provenance.postgres.test.ts
245253
lib/memory/message-provenance.postgres.test.ts
246254
lib/memory/conversation-store.postgres.test.ts
@@ -256,6 +264,7 @@ jobs:
256264
script-migrations/0016_backfill_search_vectors.postgres.test.ts
257265
script-migrations/0018_repair_workspace_file_content_revision.postgres.test.ts
258266
script-migrations/0019_tin_keyword_projection.postgres.test.ts
267+
member-sync-status-migration.postgres.test.ts
259268
260269
- name: Verify Search progress, pagination, and outbox scheduling in PostgreSQL
261270
working-directory: apps/sim

0 commit comments

Comments
 (0)