Audience. Contributors who work on ScriptCat itself — the browser extension — not script authors. If you want to write userscripts, read docs.scriptcat.org instead.
Scope. This document is the deep-dive companion to
AGENTS.md(terse contributor guide + conventions) andCONTRIBUTING.md(setup + PR workflow). It explains how the pieces fit together and why: the multi-process model, message passing, the service/data layers, the GM API system, script execution, and the build pipeline. File references use repo-relative paths and are clickable.
- The Big Picture
- The Five Contexts (Process Model)
- Message Passing
- Service layer
- Data layer (backend taxonomy: Repo<T> / DAO<T> / OPFSRepo)
- GM API system
- Script execution
- Build pipeline & manifest
- Agent subsystem
- Extending ScriptCat — Recipes
- Testing the Internals
ScriptCat is a Manifest V3 browser extension that runs Tampermonkey-compatible userscripts, plus its own background and scheduled script types that have no Tampermonkey equivalent. MV3 fragments an extension into several sandboxed JavaScript realms that cannot share memory; ScriptCat therefore runs as a small distributed system of cooperating contexts that talk over message channels.
Three ideas explain almost everything in the codebase:
- Contexts are processes. Each entry point (
service_worker,content,inject,offscreen,sandbox) is an isolated realm. They never share objects — only serializable messages. - One message layer, several transports.
packages/messageabstractschrome.runtime,postMessage, and DOMCustomEventbehind a single RPC + pub/sub API, so services are written against interfaces (Server/Group/Client/IMessageQueue), not raw browser APIs. - Domain logic lives in services, wired explicitly. Shared collaborators (a
Group,IMessageQueue, another service, a DAO owned elsewhere) typically come in through the constructor; message handlers and subscriptions are registered through an explicit lifecycle method — commonlyinit(), though content and Agent code have their own equivalents. The exact dependency set and lifecycle shape vary by service and context — see service layer for the real variance and the exceptions, rather than assuming one uniform constructor/init()template. This explicit-wiring style is what makes the system testable with the mock message bus.
┌───────────────────────────────────────────────┐
│ SERVICE WORKER │
│ central hub: script CRUD, chrome.* APIs, │
│ permission checks, resource cache, routing │
│ Server("serviceWorker") + MessageQueue │
└───────┬───────────────────────────────┬─────────┘
ExtensionMessage │ │ ServiceWorkerMessageSend (Chrome)
(chrome.runtime) │ │ / EventPageOffscreenManager (Firefox)
▼ ▼
┌──────────────────────┐ ┌──────────────────────────┐
│ CONTENT SCRIPT │ │ OFFSCREEN DOCUMENT │
│ bridges SW ↔ inject │ │ DOM-capable background │
│ Server("content") │ │ Server("offscreen") │
└──────────┬───────────┘ └─────────────┬────────────┘
CustomEventMessage WindowMessage
(DOM CustomEvent) (window.postMessage)
▼ ▼
┌──────────────────────┐ ┌──────────────────────────┐
│ INJECT SCRIPT │ │ SANDBOX (iframe) │
│ page realm, │ │ with(){} script eval, │
│ unsafeWindow access │ │ cron scheduling │
│ Server("inject") │ │ Server("sandbox") │
└──────────────────────┘ └──────────────────────────┘
MessageQueue (pub/sub) broadcasts state changes among the contexts that instantiate it — currently
Service Worker, Offscreen, and UI pages (see architecture-services.md for the exact list); content,
inject, and sandbox don't hold a MessageQueue instance.
Each context is a separate bundle (see Build pipeline & manifest) with its own entry file in src/.
| Context | Entry | Realm / capabilities | Bootstraps |
|---|---|---|---|
| Service Worker | src/service_worker.ts |
No DOM. Owns chrome.* privileged APIs, storage, permissions, routing. |
ExtensionMessage(true) → Server("serviceWorker") + MessageQueue → ServiceWorkerManager |
| Content | src/content.ts |
Isolated content-script world. Bridges SW and the page. | CustomEventMessage channel to inject + Server("content") → ScriptRuntime |
| Inject | src/inject.ts |
Page (MAIN) world. Has unsafeWindow; runs page userscripts. |
CustomEventMessage to content + Server("inject") |
| Offscreen | src/offscreen.ts |
DOM-capable background page (Blobs, clipboard, DOM scraping, local storage). | ExtensionMessage() + WindowMessage(window, sandbox) → OffscreenManager |
| Sandbox | src/sandbox.ts |
sandboxed iframe inside offscreen. Evaluates background/scheduled scripts; runs cron. |
WindowMessage(window, parent) + Server("sandbox") → SandboxManager |
There is also a sixth bundle, src/scripting.ts, injected via chrome.userScripts /
chrome.scripting to carry the compiled page-script payload (see Script execution).
src/service_worker.ts is the canonical example of how a context wires itself up:
const message = new ExtensionMessage(true); // backgroundPrimary = true
const server = new Server("serviceWorker", message); // RPC listener, action prefix "serviceWorker/"
const messageQueue = new MessageQueue(); // pub/sub broadcast bus
const hasOffscreenDocument = typeof chrome.offscreen?.createDocument === "function";
if (hasOffscreenDocument) {
const offscreen = new ServiceWorkerMessageSend(); // Chrome: talk to the real offscreen document
new ServiceWorkerManager(server, messageQueue, offscreen).initManager();
setupOffscreenDocument(); // chrome.offscreen.createDocument(...)
} else {
const offscreen = new EventPageOffscreenManager(message, messageQueue); // Firefox MV3: event page IS the DOM env
new ServiceWorkerManager(server, messageQueue, offscreen).initManager();
}This is the most important platform divergence to understand. MV3 service workers have no DOM, but
ScriptCat needs DOM for things like DOMParser, Blobs, and clipboard. Chrome solves this with the
Offscreen API (a hidden document). Firefox MV3 has no offscreen API, so its background event page
already has DOM and plays the offscreen role directly.
- Chrome: SW → Offscreen uses
ServiceWorkerMessageSend(it finds the offscreen client viaclients.matchAll()andpostMessages it); Offscreen → SW replies overExtensionMessage(chrome.runtime). - Firefox:
EventPageOffscreenManagersubstitutes for the offscreen document; its sandbox iframe is asandboxmanifest page, which Firefox 154+ loads as a cross-origin frame (contentDocumentisnull,contentWindow.locationis unreadable). Only the sandbox itself knows when it's actually ready, so the parent never polls or pings it:SandboxManager(src/app/service/sandbox/index.ts) proactively posts apreparationSandboxmessage once its ownServeris wired up (same mechanism on both platforms), andBackgroundEnvManagerBase.preparationSandboximmediately tells the service workerpreparationOffscreen({ verified: true })— no round trip, no waiting. The service worker replays enabled background/scheduled scripts and the current language only for this verified signal, once. Separately and non-blockingly, the sandbox reuses its own in-flightgetExtensionEnvrequest to self-check that the channel is genuinely bidirectional, and reports the outcome viareportSandboxChannelHealth, which the parent logs (visible in the parent's own console/log, since the sandbox iframe's console is far less discoverable). If the sandbox never announces readiness at all (iframe failed to load, script error), a fallback timer inBackgroundEnvManagerBasestill tells the service workerpreparationOffscreen({ verified: false })afterSANDBOX_READY_FALLBACK_MS, logging a clear error instead of hanging forever. That unverified notification does not send initialization through an unavailable channel; if the real handshake arrives later, the verified state replay still occurs exactly once.
Firefox packages use incognito: "spanning", so normal and private page scripts share one event page but retain
their own sender.tab.incognito value for global-switch checks, @run-in, and GM_info.isIncognito. Background
and scheduled scripts have no originating tab and run once in the shared event-page sandbox; tab-scoped
normal-tabs / incognito-tabs values therefore do not create or imply separate background instances.
Firefox packages always list webRequestBlocking in optional_permissions for the experimental event-page
keep-alive loop. On Firefox with scheduler.postTask, users can enable that loop from the install prompt or
runtime settings when they use Background Scripts or Scheduled Scripts. The loop starts only while
keep_ext_background_alive is enabled and the optional permission is granted; permission add/remove events
start or stop probing at runtime.
Services never see this difference: they receive an IOffscreenSend and call .init().
Everything cross-context flows through packages/message. It provides two
communication styles over several transports.
| Class | File | Connects | Underlying API |
|---|---|---|---|
ExtensionMessage |
extension_message.ts |
SW ↔ Content / Inject / Offscreen | chrome.runtime.sendMessage / onConnect (+ onUserScript* on Firefox) |
CustomEventMessage |
custom_event_message.ts |
Content ↔ Inject | DOM CustomEvent dispatch (bypasses page tampering) |
WindowMessage |
window_message.ts |
Offscreen ↔ Sandbox | window.postMessage |
ServiceWorkerMessageSend |
window_message.ts |
SW → Offscreen (Chrome) | clients.matchAll() + postMessage |
MessageQueue |
message_queue.ts |
Broadcast among the contexts that instantiate it — SW, Offscreen, UI pages | chrome.runtime.sendMessage + local EventEmitter3 |
MockMessage |
mock_message.ts |
Tests | in-memory EventEmitter3 |
All transports implement the small Message / MessageSend / MessageConnect interfaces in
types.ts, so higher layers don't care which one they got.
A request/reply call is identified by an action string like "script/install".
Serverlistens on a transport and routes by action prefix. The SW createsnew Server("serviceWorker", message), so it only handles actions beginning withserviceWorker/.Groupnamespaces handlers.server.group("script")returns a group whose.on("install", fn)registers the full actionscript/install. Groups can nest and carry middleware (group.use(fn)), which is howRuntimeServicedelays handling until initialization finishes.Clientis the caller side. It sends an action + params and awaits the reply.
Replies are wrapped in a uniform envelope: { code: 0, data } on success or { code: -1, message } on error,
so the calling side can reject the promise on failure.
// SW side — register
class ValueService {
init(/* … */) {
this.group.on("getScriptValue", this.getScriptValue.bind(this));
this.group.on("setScriptValues", this.setScriptValues.bind(this));
}
}
// Caller side — invoke
const value = await client.do("value/getScriptValue", { uuid });MessageQueue is a fire-and-forget bus for state changes that any
context may care about. publish(topic, data) both chrome.runtime.sendMessages the payload to every context
and emits locally; subscribe(topic, fn) registers a listener and returns an unsubscribe function; emit is
local-only (no broadcast). group(name) namespaces topics and supports middleware.
This is how data stays consistent without shared memory. For example, when a script is deleted the SW
publishes deleteScripts, and ValueService reacts by garbage-collecting orphaned values:
this.mq.subscribe<TDeleteScript[]>("deleteScripts", async (data) => {
for (const { storageName } of data) {
const stillUsed = await this.scriptDAO.find((_, s) => getStorageName(s) === storageName);
if (stillUsed.length === 0) await this.valueDAO.delete(storageName);
}
});Rule of thumb: use RPC when you need an answer; use MessageQueue to announce that something changed.
MockMessage implements the full Message interface with an in-memory
EventEmitter3 and no browser APIs, so message-driven services can be unit-tested. Tests wire it up directly —
e.g. tests/utils.ts builds the SW/offscreen mock buses from it. The separate chrome.*
mock (@Packages/chrome-extension-mock) is what
tests/vitest.setup.ts registers globally.
The Agent subsystem (src/app/service/agent/) is an AI-agent layer built on top of the five contexts
above — it is not a sixth execution context. AgentService is assembled inside ServiceWorkerManager
(src/app/service/service_worker/index.ts) from a set of narrower services (chat, task, skill, model, MCP,
background session, sub-agent, compact, DOM automation, OPFS); a content-side CAT.agent.* API
(src/app/service/content/gm_api/cat_agent.ts) exposes conversations to user scripts; offscreen and sandbox
are used for skill-script execution the same way regular background/scheduled scripts are. The UI (chat/task
pages) and userscript-driven conversations share the same streaming/tool-loop boundary in
core/tool_registry.ts / service_worker/tool_loop_orchestrator.ts — there is one tool loop, not a separate
implementation per caller. Full write-up, including the global vs. per-session tool registries, LLM
streaming/retry/compact flow, background-session/sub-agent/scheduled-task lifecycle, storage backend choices,
and the page/offscreen/sandbox delegation boundaries: references/architecture-agent.md.
Map a change onto the existing extension points instead of inventing new structure:
- A new cross-context message (RPC). Pick the owning service, add
this.group.on("myAction", handler)in itsinit(), and call it from the other context with aClient. No new transport, no new wiring. - A new broadcast event.
this.mq.publish("myTopic", payload)where state changes;this.mq.subscribe(...)wherever it matters. Use this, not RPC, for "X changed" notifications. - A new persisted entity. Pick a backend by matching an existing entity with the same data-size/query/lifecycle
needs —
Repo<T>,DAO<T>, orOPFSRepo(see data layer) — then copy the nearest entity of that same backend, in the same subsystem for ownership and construction: they don't all follow "construct in the manager, expose viagroup.on." Some context-serviceRepo<T>entities do (e.g.scriptDAObuilt inServiceWorkerManagerand exposed through RPC); some Agent entities don't —AgentModelServiceconstructs its ownAgentModelRepointernally by default rather than taking one from a manager, andAgentChatRepois a module-level singleton (export const agentChatRepo = ...) imported directly rather than manager-constructed or exposed viagroup.on. Only route through the manager +group.onwhen the entity is genuinely owned by a context composition root and needs to be reachable over RPC. - A new service. First decide what kind: a context service (owned by one of
content/,offscreen/,sandbox/,service_worker/), an Agent/core component, or another cross-cutting subsystem — these do not share one constructor shape. Then copy the nearest existing service of that kind. See service layer § Adding a service for the decision path; don't default to aGroup+IMessageQueue+ DAOs constructor for everything. - A new GM API. Decorate the method with
@GMContext.APIon the content side, add a privileged/offscreen handler if needed, register the@grant(see GM API system).
Follow the engineering principles in AGENTS.md: fix root causes (no as any / swallowed
errors), prefer direct replacement over adapter sandwiches, and keep scope tight — three similar lines beat a
premature abstraction.
- Unit (Vitest + happy-dom). Co-locate
*.test.tsnext to source.chrome.*is mocked via@Packages/chrome-extension-mock(tests/vitest.setup.ts). For the message layer, keep the two pieces separate:MockMessagefakes theMessagetransport forServer/GroupRPC tests (new Server("test", new MockMessage(...))); a service'sIMessageQueuedependency is tested with a realMessageQueueinstance (new MessageQueue(), with individual methods likepublishswapped for a spy when a test needs to assert on it) or a narrow fake written againstIMessageQueue— notMockMessage, which solves a different problem (see architecture-services.md for the detailed distinction). Run one file:pnpm test -- --run path/to/file.test.ts. - TDD. See AGENTS.md § Engineering Principles for the write-failing- test-first principle and develop-testing.md § When TDD doesn't apply for the narrow exceptions — this section only covers architecture-specific test mechanics, not the policy itself.
- E2E (Playwright).
e2e/*.spec.ts, one worker, real Chromium.pnpm run test:e2e(first run:pnpm run test:e2e:install). - Before a PR: lint + the relevant suite — owned by references/develop-testing.md → Testing.
The DI + interface design is what makes this tractable: because many services take IMessageQueue/DAOs
through the constructor (dependency set varies — see architecture-services.md),
a test builds a service with a real MessageQueue (or a narrow fake) and an in-memory DAO and exercises
handlers directly, with no browser. MockMessage is for the separate case of testing Server/Group RPC
routing, not for standing in as a service's IMessageQueue.