Skip to content

Fix the lost content write, then add the MCP write tools #229

Description

@HMarzban

Problem

Concurrent writes to PATCH /api/documents/:documentId/content are silently lost. Measured against a running stack: five parallel calls, each a full replace.

HTTP 200/200/200/200/200  ->  ["Title","N3"]

All five callers were told they succeeded. Four writes were discarded. There is no version check, no precondition and no conflict code.

Overlapping writes can also corrupt the document shape. Two trials of three produced two level-1 headings:

trial 1: 200/200 -> ["Title","pre1","Title","Y1"]
trial 2: 200/200 -> ["Title","Y2","Title","pre2"]

mode: replace stopped replacing. The server itself refuses that shape on input with 422 document content must start with a level-1 heading. So the write path can produce a document the read path would reject, and the damage survives export.

Held to a narrow claim: two clean concurrent writes never corrupted across three trials. They only lost one write each. Corruption needs an in-flight overlap; the lost update needs only concurrency.

A person rarely triggers this. An agent triggers it constantly, because it retries and runs in parallel.

What to do

Fix the lost update first. Then, and only then, add the MCP write tools.

The order is not negotiable. Shipping a write tool onto a write path that silently discards work turns one defect into many documents.

Acceptance

  • Five parallel PATCH .../content calls no longer end with four writes discarded. Either they serialize, or a losing caller receives a conflict.
  • A caller can tell a successful write from a discarded one, from the response alone.
  • After the fix, the MCP write tools ship.

Notes

The write path is apps/hocuspocus.server/src/modules/document-content/. domain/applyContentToDoc.ts performs the replace, and infra/hocuspocusApply.ts runs the flush that returns before the version row is written.

Related to the open write-model questions: a write precondition (#161) and what a PATCH should return so a caller can confirm the write (#164).

The 200 is not durable either. The store flush is synchronous, but it enqueues, and the version row is written by the BullMQ worker afterwards.

replace is idempotent at document level, so a plain retry of a single write is already safe. The defect is concurrency, not retry.

Blocked by #226.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    FeaturebugSomething isn't workingenhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions