Skip to content

DCAM4 capture: bound every wait, arm it before the start, clean up on every exit - #75

Open
kalidke wants to merge 5 commits into
mainfrom
dcam4-capture-hang
Open

kalidke wants to merge 5 commits into
mainfrom
dcam4-capture-hang

Conversation

@kalidke

@kalidke kalidke commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

DCAM4 capture hang, reported from the quickbeam rig on v0.2.4. At full frame (2304x2304) and standard readout,
capture hangs indefinitely at 12.5 ms and at 100 ms. At 100 ms it hangs as the FIRST capture in a fresh kernel,
spinning (131 CPU-s), and interrupts do not land. Full frame works at 232 ms, and at 5 ms right after an ROI switch;
256x256 works. The DCAM4 code is identical at v0.2.4, v0.2.5 and main. For 0.2.6, with #73 and #74.

What was wrong

  • dcamwait_event sent DCAMWAIT_START with size = 0. The sizing constructor existed but was never used.
  • capture armed the frame wait after dcamcap_start, so a frame that was ready early could be missed.
  • Every failure path in getlastframe and getdata did err, hwait = dcamwait_close(hwait). dcamwait_close
    returns one DCAMERR, so that line threw a MethodError. It threw before any cleanup, and capture stopped nothing
    and released nothing on failure.
  • getdata's end-of-sequence timeout ignored the readout time. sequence's status poller had no deadline.

The most likely reading of the rig pattern:

  • The frame event fires before v0.2.4 arms the wait, so the wait misses it.
  • The size-0 struct then turns that miss into a wait that never returns: a busy wait inside the DLL, which
    interrupts cannot reach.
  • This PR fixes both halves. A diagnostic for the rig (one SNAP on v0.2.4 with a correctly sized wait and a
    dcamwait_abort watchdog) is with the rig session, to confirm it. If the DLL ignores the timeout even with the
    correct size, the next step is a watchdog that calls dcamwait_abort from another thread.

Where the struct layouts come from

dcamapi4.h is not in this repository. The layouts were checked against the copy in the lab instrument archive:
manuals/Hamamatsu/SDK/dcam-api/include/sdk4-v22126552/dcamapi4.h. manuals/ is the gitignored symlink to the
NAS archive; see CLAUDE.md.

  • DCAMWAIT_START (:815-822) is four int32: size [in], eventhappened [out], eventmask [in] and
    timeout [in]. That is 16 bytes.
  • DCAMWAIT_OPEN (:806-813) is int32 size, int32 supportevent, HDCAMWAIT hwait and HDCAM hdcam. That is
    24 bytes on 64-bit.
  • DCAMWAIT_TIMEOUT_INFINITE is 0x80000000 (:393), which is negative as an Int32, and this PR refuses it.

What changes (DCAM4 files only)

  • Every DCAM wait is bounded.
    • DCAMWAIT_START carries its size, and a negative timeout is refused before any library call.
    • The frame wait's timeout is capture_timeout_ms = 2 x (exposure + readout) + 1 s.
    • The readout comes from DCAM_IDPROP_TIMING_READOUTTIME, read after setroi! by live, sequence and
      capture, and cached for getlastframe. A camera that does not report it gets 1 s, with one warning.
  • capture:
    • It refuses (throws "Stop the live view or sequence first") while a live view or sequence runs. It never
      releases a buffer that another task's getlastframe wait may be blocked on; that crashes the process.
    • Otherwise it stops and releases any leftover, arms the wait before dcamcap_start, and on every exit stops,
      releases and closes.
    • A timeout or failed wait logs, sets last_error and throws a clear error.
  • getdata:
    • In SEQUENCE mode it polls the capture status against 2 x N x (exposure + readout) + 1 s instead of waiting
      for the end-of-cycle event. An ended sequence cannot be missed, and an interrupt can land while it waits.
    • A frame that cannot be read returns nothing, with last_error set from the copy's DCAMERR
      (dcambuf_getframe_err).
    • A sequence that transferred fewer frames than requested, or never ran, returns nothing, with last_error
      set to DCAMERR_LOSTFRAME.
    • In SINGLE_FRAME mode it reads the newest frame at once. Every exit stops and releases.
    • In LIVE mode it returns the newest frame at once and leaves the live view running: no stop, no release, and
      is_running unchanged. Releasing the buffer under a threaded live reader crashes the process.
  • sequence: its poller gives up at the same deadline and stops a capture stuck running, so is_running
    cannot stay true forever.
    • It acts only while its sequence is current. stop_and_release! increments a capture_generation, and every
      start goes through it first, so a stale poller never stops or clears a newer live view or sequence.
    • A throw inside the poller still clears is_running for its own sequence.
    • A failed status read is retried until the deadline.
  • stop_and_release! and abort: stop_and_release! handles every status (a capture in ERROR was neither
    stopped nor released), and abort is now that one call.
  • getlastframe: a timeout or failed wait logs, sets last_error and returns nothing, and the wait handle is
    always closed. A frame that cannot be copied sets last_error, here and in capture.

Compatibility (decisions 0033 and 0035)

Non-breaking. The paths that now throw a clear error, or return nothing, threw a MethodError before. One
deliberate behaviour change is listed under Changed: capture while a live view or sequence runs now throws,
where it used to fail at the buffer allocation and return nothing. That path never produced a frame.

MicroscopeAdapt (control_functions.jl:250/319) and MicroscopeSeqSR (adapters.jl:91) call getdata only after
a finished sequence. DCAM4Camera gains two fields (readout_s, capture_generation), with no keyword; the keyword constructor is
unchanged.

Tests

  • test/dcam4_pure.jl covers what runs without the DCAM library:
    • the timeout formula, including the quickbeam case (1111 ms) and its refusals;
    • the size of DCAMWAIT_START (16);
    • the refusal of an unbounded wait before any library call;
    • capture's refusal while running, before any library call;
    • the cached readout time;
    • a superseded status wait, which touches nothing.
  • Local suite: 1833 pass, 3 broken (1836).
  • Follow-up (not in 0.2.6): a DCAM fake. The bindings call the literal "dcamapi.dll", so a fake needs a
    swappable library, as the TCube and PI fakes have. Until then, the capture, poll and cleanup paths are exercised
    only on a rig.

Rig checks owed (none run; no hardware)

  • R5, on the quickbeam rig (full frame 2304x2304, standard readout):
    1. As the first capture in a fresh session, capture at 100 ms, then 12.5 ms, 20 times each. Every call returns
      a frame.
    2. Force a timeout, for example with an external trigger mode and no trigger. capture throws the timeout error
      within about 2 x (exposure + readout) + 1 s. Switch back, and the next capture returns a good frame with no
      reconnect.
    3. During live, capture throws "Stop the live view or sequence first", and getdata returns a frame. The
      live view keeps running through both.
    4. sequence of 100 frames at 12.5 ms, then getdata: it returns all 100 frames.
    5. No "did not report its readout time" warning appears, and the timeout error in step 2 names a readout close to
      2304 x 18.65 us (about 43 ms).

🤖 Generated with Claude Code

kalidke and others added 4 commits September 29, 2026 14:53
Decision 0033: after X.Y.Z, main carries X.Y.(Z+1)-DEV. TagOnMerge skips a
-DEV version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… every exit

The quickbeam rig saw capture hang at full frame, 12.5 ms (v0.2.4). Fixes:
- DCAMWAIT_START is built with its size (it was sent as 0), and a negative timeout
  (DCAM's INFINITE) is refused before any library call.
- Every frame wait is bounded by capture_timeout_ms: 2 x (exposure + readout) + 1 s,
  with the readout read from the camera (TIMING_READOUTTIME); getdata scales it by N.
- capture stops and releases any earlier capture first, arms the wait before
  dcamcap_start, and stops, releases and closes on every exit. A timeout or failed
  wait logs, sets last_error and throws a clear error.
- getlastframe and getdata no longer destructure dcamwait_close's single DCAMERR
  (a MethodError on every failure path). getdata reads an ended cycle at once,
  never passes DCAM a null wait handle, and cleans up on every exit (in LIVE mode
  it now ends the live view).
- test/dcam4_pure.jl covers what runs without the DCAM library.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…olls against a deadline, sequence's poller has one

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… a newer capture, getdata leaves a live view running

Also: getdata reads a READY buffer that never ran as LOSTFRAME, and capture and getlastframe set last_error on a failed frame copy.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
kalidke added a commit that referenced this pull request Sep 30, 2026
…re the start, cleanup on every exit

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

# Conflicts:
#	CHANGELOG.md
@kalidke kalidke mentioned this pull request Sep 30, 2026
@kalidke

kalidke commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

This ships in #76 (MicroscopeControl 0.2.6), the single 0.2.6 release PR, as a merge commit of this PR's accepted head. Please leave this PR open: GitHub marks it merged when #76 lands.

@dcamcall replaces @CCall at all 30 DCAM call sites. With tracing on
(DCAM4.dcam_trace!(path) or ENV MC_DCAM4_TRACE) it writes a flushed BEGIN and
END line per call with arguments, elapsed time, return value and cumulative
GC time, so a hang inside a ccall leaves its last BEGIN on disk. Off, it
costs one Ref{Bool} check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant