Commit Graph
1115 Commits
Author SHA1 Message Date
x12sa f916b2cec1 config update
CI for csaxs_bec / test (push) Failing after 2m12s
CI for csaxs_bec / test (pull_request) Failing after 2m15s
2026-09-15 08:22:55 +02:00
x12saandClaude Sonnet 5 6e8fc91924 docs(omny): fix omny_fermat_scan parameter table and add docstring example
CI for csaxs_bec / test (push) Failing after 2m10s
Reorders omny.md's parameter table to match the actual signature,
adds the previously undocumented readout_time parameter, and adds an
Examples block to OmnyFermatScan's docstring (mirroring the flomni fix).
Also rewords flomni.md's corridor_size default to read "3 um" for
consistency with omny, while noting it is auto-estimated when omitted.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FFeiCHv3qUxxssmP9Qzhom
2026-09-14 20:52:03 +02:00
x12saandClaude Sonnet 5 5141854e14 docs(flomni): fix flomni_fermat_scan parameter table to match code
CI for csaxs_bec / test (push) Failing after 2m5s
Reorders parameters to match the actual signature, adds the
undocumented burst_at_each_point parameter, and corrects the
corridor_size default (None/auto-estimated, not a literal 3 um).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FFeiCHv3qUxxssmP9Qzhom
2026-09-14 20:47:20 +02:00
x12saandClaude Sonnet 5 2753bb9243 docs(flomni): add usage example to flomni_fermat_scan docstring
CI for csaxs_bec / test (push) Failing after 2m10s
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FFeiCHv3qUxxssmP9Qzhom
2026-09-14 20:12:27 +02:00
x12saandClaude Sonnet 5 501c93d41c fix(flomni): use last completed scan number when tomo_reconstruct's cached scan list is stale
CI for csaxs_bec / test (push) Failing after 2m4s
tomo_reconstruct() named the queue file from a fresh next_scan_number but
wrote its content from self._current_scan_list, which is only kept in sync
by tomo_scan_projection()/tomo_acquire_at_angle(). Calling it directly from
the command line -- e.g. after a plain scans.flomni_fermat_scan() -- wrote
whatever scan list was left over from an earlier tomo scan, or raised
AttributeError if none had run yet this session.

Falls back to [next_scan_number - 1] whenever the cached list's last entry
doesn't match the scan that actually just completed. Internal callers are
unaffected since their cached list always matches at the point they call it.

Also notes the identical bug in OMNY.tomo_reconstruct() (a separate,
non-shared implementation) as a TODO for a later fix on its own branch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TUimPoyFQRvxM6R3njvuVj
2026-09-14 17:06:02 +02:00
x12saandClaude Sonnet 5 df7c90c5d7 docs(omny): document manual tomo_alignment_fit offset nudge
CI for csaxs_bec / test (push) Failing after 2m5s
Port the flomni doc addition to omny -- both setups share the identical
tomo_alignment_fit mechanism (OMNYAlignmentMixin.get_alignment_offset
mirrors Flomni's), so the same command-line offset nudge applies verbatim.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FFeiCHv3qUxxssmP9Qzhom
2026-09-13 21:56:15 +02:00
x12saandClaude Sonnet 5 d100af5429 docs(flomni): document manual tomo_alignment_fit offset nudge
CI for csaxs_bec / test (push) Failing after 2m16s
Add a snippet to the XrayEye alignment section showing how to patch the
constant x/y offset terms of tomo_alignment_fit directly on the command
line, for a quick few-micron correction (e.g. after moving foptz) without
recording a new fit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FFeiCHv3qUxxssmP9Qzhom
2026-09-13 21:54:42 +02:00
x12saandClaude Sonnet 5 5a22651c46 fix(flomni): report correct projection count in tomo PDF/SciLog for golden-ratio scans
CI for csaxs_bec / test (push) Failing after 2m7s
write_pdf_report() always derived "Number of projections" (and the
dependent sub-tomogram-count/angular-step fields) from
_tomo_type1_actual_grid(), which is only valid for tomo_type==1. For
golden-ratio scans (tomo_type 2/3) it silently reported a stale,
unrelated type-1 total left over in the persistent tomo_angle_stepsize
global var instead of golden_max_number_of_projections, causing the
PDF/SciLog entry to disagree with the web page's correct progress
display. Branch on tomo_type the same way tomo_parameters() already
does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FFeiCHv3qUxxssmP9Qzhom
2026-09-12 20:31:24 +02:00
x12saandClaude Sonnet 5 2dae9441ec docs(ids-cameras): extend exposure/gain plan to xrayeye widget controls
CI for csaxs_bec / test (push) Failing after 2m10s
No code changes -- IDSCamera and the xrayeye widget are both used
during beamtimes. Revises the earlier plain-method design to
Cpt(Signal, kind=Kind.config) (mirroring live_mode_enabled) so the
xrayeye widget can read exposure/auto-exposure/auto-gain state via
its existing device_read_configuration subscription instead of
polling the device, and adds the corresponding GUI control knobs
(auto-exposure/auto-gain toggles, exposure spinbox) to the plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014s92UHyMfov6Fb934pGqcU
2026-09-11 20:30:54 +02:00
x12saandClaude Sonnet 5 ec3e7c0679 docs(ids-cameras): add plan for manual exposure/auto-gain control
CI for csaxs_bec / test (push) Failing after 2m5s
No code changes -- IDSCamera is in production use during beamtimes.
Documents the intended USER_ACCESS additions (get/set_exposure_time,
set_auto_exposure_enabled, set_auto_gain_enabled) for a future session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014s92UHyMfov6Fb934pGqcU
2026-09-11 18:01:49 +02:00
x12saandClaude Sonnet 5 0cadd8582e fix(flomni,omny): sanitize sample name spaces in tomography_scannumbers.txt
CI for csaxs_bec / test (push) Failing after 19m49s
The scannumbers log is read column-wise by downstream tooling; a
sample name containing spaces breaks that whitespace-based parsing
since it's written as the trailing field. Replace spaces with
underscores before writing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FB3kYVkJruYsmcVsj7RS1G
2026-09-11 16:24:12 +02:00
x12saandClaude Sonnet 5 47a9c4c444 fix(flomni,lamni): persist tomo_id across client restarts
CI for csaxs_bec / test (push) Failing after 4m17s
self.tomo_id was a plain instance attribute reset to -1 on every
Flomni/LamNI init, so resuming a tomo scan after a client/kernel restart
(tomo_scan_resume(), or tomo_queue_execute()'s automatic resume) kept
reporting/uploading under the wrong id instead of the one actually
registered with OMNY for that measurement.

Back it with the existing _GlobalVarParam descriptor on TomoQueueMixin
(the same BEC Redis-backed global-var mechanism already used for
tomo_shellstep, stitch_x, at_each_angle_hook, etc.) so it survives a
restart like the rest of the tomo-scan state already does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gGu5ANBo88Fh7eqmRdFbU
2026-09-11 16:19:10 +02:00
x12saandClaude Sonnet 5 da4007d7e7 fix(flomni,lamni): stop registering test/gac accounts against production omny.psi.ch
TomoIDManager.OMNY_URL and TEST_OMNY_URL both point at omny.psi.ch since
the dedicated omny-test.psi.ch host was retired, so the eaccount check no
longer redirects test/gac-* accounts anywhere different -- it just logs a
warning and still registers them in the production sample database. Skip
registration entirely for non-e-accounts and return FALLBACK_TOMO_ID
instead, to avoid polluting production with test entries.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gGu5ANBo88Fh7eqmRdFbU
2026-09-11 16:18:51 +02:00
x12saandClaude Sonnet 5 42739fcc7a docs(flomni): flag unguarded heater-vs-OSA collision gap in fosa_in()
CI for csaxs_bec / test (push) Failing after 4m58s
Audit of heater interlocks (pre-first-use) found that fosa_in()/
foptics_in() never check fheater position before driving fosaz, even
though the OSA travels inside the heater's envelope. Heater motion
itself is well protected (ensure_osa_back()/ensure_fheater_up()), but
this one direction is open. Not fixing it yet -- documenting it so it
is remembered when fheater's userParameter is added during hardware
commissioning.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 13:15:58 +02:00
x12saandClaude Sonnet 5 f56d11d730 fix(flomni): close race between sample-storage EPICS writes and reads
CI for csaxs_bec / test (push) Failing after 6m18s
ftransfer_sample_change occasionally raised "The gripper does not carry
a sample" right after a successful get, even though the gripper
physically held one. Root cause: flomni_modify_storage_non_interactive
wrote the sample_in_gripper/sample_placed signals with EpicsSignal.set()
without waiting for the write to complete, and the immediately-following
is_sample_in_gripper()/is_sample_slot_used() checks read the
auto_monitor-cached value, which could still reflect the pre-write state.

Wait on each .set() call (5s timeout, matching the existing
.set(...).wait(timeout=...) pattern used elsewhere for EPICS writes),
and force is_sample_in_gripper()/is_sample_slot_used() to do a live
get(use_monitor=False) read instead of trusting the monitor cache
(mirroring the same pattern already used in ddg_1.py). This closes the
race deterministically instead of masking it with an arbitrary sleep.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 10:18:29 +02:00
x12saandClaude Sonnet 5 59cdc5c512 fix(omny): increase laser tracker on-target wait timeout, add status print
CI for csaxs_bec / test (push) Failing after 7m30s
Bump max_repeat from 25 to 100 (~12.5s to ~50s) in
laser_tracker_wait_on_target, since the tracker was timing out too
early during beamtime. Print a status message once the retry-enable
branch kicks in so it's visible the wait is still in progress.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 09:11:37 +02:00
x12saandClaude Sonnet 5 d6e2ca8d1a fix(omny): mitigate scrambled Galil PUT/GET communication with settle delays
GalilController had zero delay in its socket path (write immediately
followed by a blocking read), unlike sgalil_ophyd.py's GalilController
which already sleeps 10ms in both socket_put and socket_get. This
increased the chance of partial-packet collisions/desync between
commands and responses during flomni beamtime (see command_history
showing merged/missing PUT-GET pairs).

Add a 10ms sleep to socket_put and a socket_get override (also with a
10ms sleep) that logs into command_history, mirroring the existing
sgalil_ophyd.py pattern. This reduces desync likelihood but does not
fully fix it: failed socket_get attempts are still swallowed silently
by retry_once with no trace in command_history, and receive() still
does a single recv(1024) with no terminator framing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 09:11:08 +02:00
x12saandClaude Sonnet 5 780be4bb9f config(ptycho_flomni): update axis calibration values, disable omny_panda
Update fttrx, foptx, ftransy sensor_voltage, fosax and fosay calibration
values for current beamtime alignment. Disable omny_panda device.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 09:10:26 +02:00
x12saandClaude Sonnet 5 0fd52d23e4 config(main): switch active endstation config to flomni
Disable xeye/ssaxs includes and enable ptycho_flomni for flomni beamtime.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 09:09:59 +02:00
x12saandClaude Sonnet 5 0627221491 config(bl_detectors): switch active detector from eiger_9 to eiger_1_5
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 09:09:43 +02:00
x12saandClaude Sonnet 5 031e612394 fix(flomni): re-enable webpage generator start
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 09:09:24 +02:00
x12saandClaude Sonnet 5 36d34d2858 fix(eps): cap alarm texts per event to stop eps_alarm_history.json growing unbounded
CI for csaxs_bec / test (push) Failing after 7m36s
A long-running/flapping EPS alarm appended every slightly-different
AlarmList_EPS text to the open event's texts list with no limit,
eventually growing eps_alarm_history.json past the upload server's
size limit (HTTP 413), which then retried and re-failed indefinitely.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XdHenMmwbrCz1juWMi1yZc
2026-09-11 09:07:25 +02:00
x12saandClaude Sonnet 5 e65389376d fix(flomni): fix UnboundLocalError for dev in tomo_scan_projection
CI for csaxs_bec / test (push) Failing after 2m19s
dev = builtins.__dict__.get("dev") was only executed inside the
`if not _internal:` block, but Python treats any name assigned
anywhere in a function as local to the whole function. The
unconditional dev.rtx.controller.laser_tracker_check_signalstrength()
call later in the function then raised UnboundLocalError whenever
_internal=True (e.g. tomo_alignment_scan() and the fermat branch of
_at_each_angle). Hoist the fetch above the guard so dev is always
bound.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CUx5yV1iCBhL5sFEHEoPe
2026-09-11 07:46:23 +02:00
x12saandClaude Sonnet 5 3bea6b1a30 docs(flomni): add cSAXS_HR August 2018_chip_maxime.pdf
CI for csaxs_bec / test (push) Failing after 2m3s
New reference PDF in the flomni docs folder, served by flomnigui_docs().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CUx5yV1iCBhL5sFEHEoPe
2026-09-10 16:46:57 +02:00
x12saandClaude Sonnet 5 a457865768 docs(omny): TODO for filter check and missing eye/optics gate
CI for csaxs_bec / test (push) Failing after 2m3s
Notes for the OMNY branch: 1) port the filter-out-of-beam check (added
for flomni/lamni in the previous commit, shared helper already in
OMNY_shared/filter_check.py) into omny.py's tomo_scan()/
tomo_scan_projection(); 2) OMNY.tomo_scan() has no eye-out/optics-in
precondition check at all, unlike Flomni/LamNI's - a pre-existing,
unrelated gap noticed while investigating the filter check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CUx5yV1iCBhL5sFEHEoPe
2026-09-10 15:41:02 +02:00
x12saandClaude Sonnet 5 7d64b49918 feat(flomni,lamni): hard-warn if beam filters are not out before scans
Adds a shared filters_out_of_beam()/warn_and_confirm() helper
(OMNY_shared/filter_check.py) that checks the four cSAXS exposure-box
filter axes (filter_array_1_x..4_x) against their "out" position, reusing
cSAXSFilterTransmission's own position table.

Wired into flomni.py and lamni.py at every point a scan can start:
tomo_scan() and tomo_alignment_scan() (once per run, same interactive/
queued warn-and-confirm convention as the existing X-ray-eye/optics
check), and tomo_scan_projection() (every direct/standalone call, via a
new _internal flag on LamNI's method mirroring flomni's existing one;
internal per-angle loop calls skip the redundant re-check). Since
tomo_queue_execute() already calls tomo_scan(interactive=False) once per
queued job, queued runs get exactly one warn-and-proceed check at queue
start, not one per projection.

DataDrivenLamNI.tomo_scan() (extra_tomo.py) fully overrides the base
method, so it gets its own copy of the same top-of-scan check.

OMNY intentionally not touched here - left as follow-up on the user's
separate OMNY branch, see the OMNY AI_docs TODO in the next commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CUx5yV1iCBhL5sFEHEoPe
2026-09-10 15:40:52 +02:00
x12saandClaude Sonnet 5 36e19adde4 fix(flomni): make flomnigui_show_* macros idempotent
CI for csaxs_bec / test (push) Successful in 2m8s
flomnigui_show_gui() unconditionally called self.gui.flomni.raise_window()
whenever the flomni window already existed. That RPC call is currently
buggy in bec_widgets and can hide the window instead of bringing it to
front, so calling any flomnigui_show_* macro a second time (e.g. from a
script) made the GUI disappear. Disable the raise_window() call, matching
the same workaround already applied in flomnigui_raise(); repeat calls now
just update content in place instead of re-raising.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CUx5yV1iCBhL5sFEHEoPe
2026-09-10 14:00:48 +02:00
x12sa fe4cb2e661 increase velocity of Owis stages in xeye.yaml
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 2s
CI for csaxs_bec / test (push) Successful in 2m3s
2026-09-09 12:30:34 +02:00
x12sa 3f521db0de add in position values and recommended gain for XBPM3 2026-09-09 12:30:34 +02:00
x12sa f1690aec52 added comment for nominal xbpm2x position according to Excel file 2026-09-09 12:30:34 +02:00
x12sa 65230c608d added comment for settings of microstepping card 2026-09-09 12:30:34 +02:00
x12sa 3e54a96e84 comment out X-ray eye 2026-09-09 12:30:34 +02:00
x12sa 46aa051c00 update IDS camera ID 2026-09-09 12:30:34 +02:00
menzel f74f09de7b fix(cont_grid): wait for the move to reach the motor before triggering
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 2s
CI for csaxs_bec / test (push) Successful in 2m3s
at_each_point issues the line move and then sleeps before firing the
burst. The sleep was acc_time alone, which covers the acceleration ramp
and nothing else -- while the move still has to cross the scan server ->
device server -> EPICS -> controller path first. acc_time shrinks with
the scan velocity; that path does not. On a slow scan the burst
therefore began well before the stage moved, and the first points of
every line piled up at the line start.

Measured from the position readback of five commissioning scans, as the
distance the stage was behind the trigger grid once the lag stopped
growing:

    scan   exp_time   v_cmd         lag       latency
     254     20 ms    0.5   mm/s   29.7 pts   0.594 s
     450     50 ms    0.1   mm/s   11.9 pts   0.595 s
     324     15 ms    0.667 mm/s   41.5 pts   0.623 s
     473    100 ms    0.065 mm/s    6.3 pts   0.630 s
     411     50 ms    0.1   mm/s   12.7 pts   0.635 s

Constant to +-3.5% across a 6.7x range of exposure and a 10x range of
velocity, on two axes. Scan 473 is sampled 74 times per line and shows
the shape plainly: the lag appears in the first interval and then holds
at 6.3-6.4 points for the rest of the line while the velocity sits at
exactly the commanded 0.065 mm/s. A start-up offset, not a velocity
error, and the stage itself is blameless.

Deliberately not solved by polling for motion. Observing the readback
costs a round trip of this same ~0.6 s, so it would trade a systematic
offset for a jitter of similar size -- and a constant offset displaces
every line equally, where a varying one shears the image line by line
and cannot be undone afterwards. The reproducibility is the asset here,
not the enemy.

The value lives on ddg1 next to the shutter delay, with a setter in
USER_ACCESS, because that is where cont_grid already fetches its trigger
timing. It is not a property of the delay generator and the docstring
says so. Setting it to 0 reproduces exactly the previous timing without
a redeploy, which is the intended way back if this makes things worse.

The real fix is to trigger the DDG from the motor
(scan_type: hardware_triggered), taking the round trip out of the timing
chain rather than compensating for it. This is the stopgap until then,
and it rests on the latency staying constant -- which is worth
re-checking with the same measurement whenever the deployment changes.
2026-09-09 12:30:16 +02:00
menzel 3ddd0d7957 fix(ddg1): give the shutter the time it actually needs to open
The first point of every line is under-exposed. Four commissioning scans
put a number on it, against the 2e-3 head start that was in place:

    scan   exp_time   first/near   lost
     324     15 ms       0.655     5.18 ms
     254     20 ms       0.732     5.36 ms
     411     50 ms       0.881     5.93 ms
     450     50 ms       0.879     6.04 ms

A fixed time, not a fixed fraction: it varies by 17% across a 3.3x range
of exposure while the fraction varies by 2.9x. The same deficit appears
on the integrated scattering, a different detector behind a different
gate, so the cause is upstream of both readout chains rather than in
either of them. 2e-3 allowed and ~5.6e-3 still lost means the shutter
needs about 7.6 ms, rounded up to 8 ms: the spread across the four scans
is 0.9 ms, so the third digit is not meaningful, and overshooting costs
only the difference in dead time at the start of each line.

The trigger scheme was already right -- the shutter fires on cd at t0
and the acquisition on ab is held back by _shutter_to_open_delay, with
the widths, burst_period and cont_grid's acc_time and premove all
derived from it. Only the value was wrong, and it was a literal in two
places, so setting one and not the other would have been silently undone
by keep_shutter_open_during_scan.

Lifts it to DEFAULT_SHUTTER_TO_OPEN_DELAY next to the other defaults,
with the measurement recorded, and adds set_shutter_to_open_delay to
USER_ACCESS so the value can be converged from the client instead of by
redeploying the device server. It is bounded, because a delay is paid on
every line and a fat-fingered value would stretch the scan rather than
fail.

Cost at the new value is 6 ms per line -- 0.26 s over scan 450 -- and
0.8 um of extra premove.

The existing stage test asserted the 2e-3 literal and now asserts the
constant. New tests cover the default, the bound, the USER_ACCESS entry,
that a set value actually reaches the ab channel while cd still fires at
t0, and that keep_shutter_open_during_scan discards a tuned value, which
is a sharp edge worth pinning rather than leaving to be rediscovered.
2026-09-09 12:30:16 +02:00
menzel 3bf85ab36a Feat/smargon (#314)
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 1s
CI for csaxs_bec / test (push) Successful in 1m48s
Per-axis BEC motors over the virtual SmarGon Coordinate System axes, sharing a singleton SmargopoloController, with an ophyd-free transport layer — RestTransport against the real :3000 API, plus an in-memory FakeTransport for offline use and tests. smargopolo runs the kinematics and drives the underlying q1..q6 MCS2 stages; this device never commands those directly.
Adds only its own package and its own two config files.
Tested against the hardware since.
Merged into current main: 101 failed / 580 passed / 12 skipped (measured with all three branches stacked), against a baseline of 101 failed / 532 passed.

Reviewed-on: #314
2026-09-09 12:29:57 +02:00
menzelandClaude Opus 5 af5603d9b8 fix(falcon): name prefix in __init__ so the device server stops dropping it
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 1s
CI for csaxs_bec / test (push) Successful in 1m48s
The __init__ added in the previous commit took (*args, **kwargs). The device
server builds a device's init kwargs by intersecting the deviceConfig keys
with the NAMED parameters of the class signature
(bec_server/device_server/devices/devicemanager.py:469-475), so 'prefix' was
silently discarded and the Falcon was constructed with an empty prefix.

Every signal then pointed at a bare suffix -- HDF1:FilePath_RBV instead of
X12SA-SITORO:HDF1:FilePath_RBV -- and instantiation failed with
"TimeoutError: Failed to connect to all signals" listing several hundred PVs.
That reads like an unreachable IOC, which is how it was diagnosed at the
beamline for two hours, while caget from the same host worked perfectly.

The signature now names name, prefix, scan_info, device_manager and
xml_file_name explicitly, matching DDG1. Two tests guard it: one asserts every
deviceConfig key is a named parameter, the other that a configured prefix
reaches the signal PV names.

Reported-by: Klaus Wakonig <klaus.wakonig@psi.ch>

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLnmUurqcNd1FiDY5M2uZr
2026-09-09 12:29:26 +02:00
menzelandClaude Opus 5 484bf9821d fix(falcon): set the layout the IOC can actually resolve, and verify it
_initialize_detector_backend put a bare "layout.xml" into the HDF5 plugin.
The IOC resolves that relative to its own working directory, where no such
file exists: the cSAXS SITORO IOC ships the layout as cfg/layout.xml
(installed at /ioc/X12SA-CPCL-FALCONX1/cfg/layout.xml), and that is what the
IOC configures at startup. BEC was overwriting a correct value with a broken
one on every device init.

The failure was silent and badly signposted. An unreadable layout does not
raise; the plugin simply refuses to open the output file later, reporting
"Error opening file ..., status=-1" with the actual cause buried in
XMLErrorMsg_RBV. Because a caget taken before a BEC device reload showed the
IOC's own valid value, the two readings disagreed and the layout looked
innocent.

The default is now cfg/layout.xml, overridable per deployment with an
xml_file_name key in deviceConfig ("" selects the plugin's built-in layout),
and on_connected reads XMLValid_RBV back and logs an error naming the file
and the IOC's message if the layout was rejected.

The existing on_connected test asserted the broken value; it now asserts the
default constant.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLnmUurqcNd1FiDY5M2uZr
2026-09-09 12:29:26 +02:00
menzelandClaude Opus 5 c9cd7ac972 fix(macros): no module-level assignment, or the loader refuses the file
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 2s
CI for csaxs_bec / test (push) Successful in 1m45s
The previous commit added `logger = bec_logger.logger` at module level. The
macro loader rejects any module-level ast.Assign
(bec_lib.macro_update_handler.has_executable_code) and then refuses the whole
file, so run_cont_grid_scan_for_table_row stopped being loaded at all:

    Macro file .../run_cont_grid_scan_for_table_row.py contains executable code
    at module level (line 16) and will not be loaded for security reasons.

Imports, defs, classes, annotated assignments and docstrings are permitted;
plain assignments are not. The logger is now fetched inside the functions.

Adds a test that runs the real has_executable_code over every macro, so this is
caught by the suite instead of by a WARNING in the log stream. Verified to fail
when a module-level assignment is reintroduced.

Noticed only because bec-log-monitor happened to be running at the time -- a
refused macro is otherwise indistinguishable from one that was never installed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0148tn6uK6oiTH25mzLfJcyc
2026-09-09 12:29:09 +02:00
menzelandClaude Opus 5 e3409f6520 fix(macros): report failures fully instead of a one-line summary
When a table row failed, the macro printed a headline, sent the actual error
text to SciLog alone, and left both the SciLog post and the SMS unguarded. So
the detail existed in exactly one place that nobody was watching, and if SciLog
was unreachable its exception replaced the one being reported -- the bare
`raise` at the end never ran. Combined with @scan_repeat retrying three times,
a deterministic one-line DeviceConfigError produced three identical
context-free messages and survived several hours of beamtime.

Failures now go through _report_failure, which:

  - prints the exception type, message and full traceback to the console;
  - logs the same through bec_logger, so it reaches the log files AND Redis and
    is therefore visible in `bec-log-monitor` and afterwards in the logs, rather
    than only on whichever console ran the macro;
  - includes _row_context: sample, template, both scan axes with ranges and step
    sizes, exposure time, and for tensor rows the rotation axes and angles, so a
    report identifies the row without needing the table alongside it;
  - guards SciLog and SMS separately, each reporting its own failure without
    touching the original exception.

The caller still re-raises, so scan_repeat and the queue behave as before.

Tests cover the two masking cases that mattered -- an unreachable SciLog and a
failing SMS must not replace the original error -- plus the tensor context and
that no SMS is attempted without phone numbers.

Not changed, but flagged: @scan_repeat(max_repeats=3, default=True) retries any
error three times, including deterministic ones. The file's own TODO warns about
this. It triples the noise while diagnosing a reliably failing scan.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0148tn6uK6oiTH25mzLfJcyc
2026-09-09 12:29:09 +02:00
menzelandClaude Opus 5 c7ea404e90 fix(macros): look devices up in dev, not dev.devices
The tensor-tomography branch of run_cont_grid_scan_for_table_row has never
worked. `dev` is already the device container -- bec_lib binds
dev = device_manager.devices in the client namespace -- so dev.devices[name]
asks DeviceContainer for a device literally called "devices" and its
__getattr__ raises DeviceConfigError before any motor moves.

Observed at the beamline on a tensor table using sgchi/sgphi as the rotation
axes:

    --> 72  roty_motor = dev.devices[row["roty_axis"]]
    DeviceConfigError: Device devices does not exist.

The failure is deterministic, so @scan_repeat(max_repeats=3, default=True)
retried it three times, and the surrounding except reported only "Error while
moving motors to starting position for sample ..." -- the traceback goes to
SciLog and nowhere else, which is why this survived undetected.

Adds tests/tests_macros, which had no equivalent: macros run in the client
namespace and are not executed by any test, so mistakes in them reach the
beamline unfiltered. The check parses each macro and flags dev.devices
attribute access rather than matching text, so comments and docstrings that
mention the pattern do not trip it. Verified to fail against the unfixed macro.

The rotation axes themselves are not hard-coded: rotx_axis/roty_axis are row
fields holding a device name chosen from the SAXS widget's positioner combo
boxes. Only the four row keys are fixed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0148tn6uK6oiTH25mzLfJcyc
2026-09-09 12:29:09 +02:00
wakonig_k 376e5b1661 fix(camera): stop live mode before disconnecting to prevent polling
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 1s
CI for csaxs_bec / test (push) Successful in 1m50s
2026-09-08 16:03:01 +02:00
menzelandClaude Opus 5 653351fc15 feat(eiger): make the missing-packet tolerance per-detector and switchable
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 2s
CI for csaxs_bec / test (push) Successful in 1m52s
The JungfrauJoch broker reports packet loss during data collection as a plain
error, which fails the scan. Since 2026-09-03 that error has been suppressed on
the beamline by an uncommitted edit, through a hardcoded flag on the Eiger base
class: it applied to every Eiger at once, it could be reached neither from the
client nor from deviceConfig, and it left nothing in the log.

The tolerance is now a real parameter, raise_on_missing_packets:

- named in Eiger, Eiger9M and Eiger1_5M, so that a deviceConfig key actually
  reaches the device. bec_server intersects config keys with the named
  parameters of the class, so a flag reachable only through **kwargs is
  silently dropped -- the same trap as readout_time (f450f29) and prefix
  (10be2b5). A test pins the signatures.
- exposed through USER_ACCESS as get_/set_raise_on_missing_packets, so a
  beamtime can change its mind without a redeployment. Like every runtime
  value it is shared between clients and does not survive a server restart;
  deviceConfig is what makes a choice stick.
- counted, and logged at warning level whenever an error is let through, so
  that "which scans were affected?" has an answer. get_missing_packet_events()
  returns the count.

What is tolerated is narrower than it looks: the frame-count check below still
raises when statistics.images_collected falls short of the trigger count. Only
"the broker flagged packet loss but delivered the expected number of images"
gets through, and a test pins that a short acquisition still fails.

The default stays False, i.e. tolerate, so the running beamtime is unaffected.
It should become True once the 9M's packet loss is understood, with
raise_on_missing_packets: false in that detector's deviceConfig if it still
needs it. That is one constant to change, RAISE_ON_MISSING_PACKETS.

The wording of the broker message is the only handle available, as there is no
error code for it. If JungfrauJoch rephrases it the match stops working and the
error raises again, which is the safe direction to fail in.

test_eiger_on_complete_error_message was skipped as failing "because the error
should be skipped for now due to HW issues". With the tolerance scoped to the
missing-packet message it passes again, and is no longer skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 14:30:53 +02:00
menzel e1a8f88164 Merge branch 'main' into feat/trigger-timing-consistency
CI for csaxs_bec / test (pull_request) Successful in 1m48s
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 2s
CI for csaxs_bec / test (push) Successful in 1m53s
2026-09-08 14:21:28 +02:00
wakonig_k 4c53f2cf25 fix(mcs_card): handle excess data points and improve MCA callback suppression
CI for csaxs_bec / test (pull_request) Successful in 1m45s
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 1s
CI for csaxs_bec / test (push) Successful in 1m43s
2026-09-05 12:56:52 +02:00
menzelandClaude Opus 5 cfc7271d40 docs(eiger): record that triggering emulates gating, and why
CI for csaxs_bec / test (push) Failing after 5s
CI for csaxs_bec / test (pull_request) Successful in 2m0s
The Eiger is triggered rather than gated for stability, and the pulse train is
shaped so its internal timer coincides with the gate -- it is meant to behave as
if gated, so that every detector in a scan integrates the same window. Nothing
in the code said so, which is why sending the full exp_time as image_time_us
looked reasonable and stayed wrong until a scope showed the 173 us overhang.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLnmUurqcNd1FiDY5M2uZr
2026-09-02 17:54:35 +02:00
menzelandClaude Opus 5 64c2fff6e6 fix(eiger): let each model keep its own readout time, and make it configurable
CI for csaxs_bec / test (push) Failing after 7s
Readout is a property of the detector, not a beamline constant: a 9M has more
modules than a 1.5M, and the Falcon needs 3 ms against the delay generator's
200 us. The per-model constants therefore stay, with comments saying the
duplication is deliberate so nobody consolidates them again. They all hold 2e-4
today only because no measured per-model value exists yet.

Writing a test for the deviceConfig override exposed that it never worked. Both
subclasses passed readout_time to super() while also forwarding **kwargs, so
supplying it raised "got multiple values for keyword argument" -- and through
BEC it never even got that far, because readout_time was not a named parameter
of the subclass signature and the device server drops config keys it cannot see
(the same rule behind the recent prefix incident). Both subclasses now name it
with the model constant as default, and a test asserts the signatures keep it.

Also documents frame_time_us in DetectorSettings as required-but-ignored for the
Eiger, and warns that its 500 is microseconds while every other time in the
module is seconds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLnmUurqcNd1FiDY5M2uZr
2026-09-02 17:42:38 +02:00
menzelandClaude Opus 5 77fc2525a0 feat(eiger,ddg2): acquire over the gated window, not the whole period
CI for csaxs_bec / test (push) Failing after 5s
scan_info's exp_time is the trigger PERIOD. DDG2 gates for exp_time minus a
readout time, but the Eiger was sent the full exp_time as image_time_us, so
the detector acquired 180 us past the falling edge of its own gate and
finished only 20 us before the next trigger -- that 20 us being JungfrauJoch's
internal board-readout allowance, and the entire margin available. Measured on
a scope as a 172.8 us overhang.

Four independent notions of "readout time" existed:

  500 us  EIGER*_READOUT_TIME_US   validation floor only, no effect on anything
  200 us  DDG2 DEFAULT_READOUT_TIMES["ab"]   sets the gate width
   20 us  JungfrauJoch deployment config     applied internally by JFJoch
    -     scan parameter readout_time        honoured by the Falcon, ignored by DDG2

They are now one. DDG2 derives the gap from the scan's readout_time, floored by
its configured value (scans default it to 0), and exposes effective_readout_times().
The Eiger subtracts the same number from image_time_us. The Falcon already used
the scan value, so it needs no change.

The Eiger's 500 us constant is split in two, because it was doing two jobs: a
MIN_EXP_TIME validation floor keeps today's behaviour exactly, while the new
EIGER_READOUT_TIME defaults to 2e-4 to match the delay generator. The _US suffix
on constants holding seconds is dropped in all three modules.

Both devices log the effective exposure window at on_stage, so a period/exposure
mismatch is visible in the logs rather than only on an oscilloscope.

NOTE: this shortens the delivered exposure by the readout time -- 200 us, i.e.
0.5% at 40 ms -- for scans that do not set readout_time. Users should be told
before this is deployed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLnmUurqcNd1FiDY5M2uZr
2026-09-02 17:02:17 +02:00
menzelandClaude Opus 5 bb39833d83 feat(ddg2): make the detector-trigger readout time configurable
CI for csaxs_bec / test (push) Successful in 3m48s
The gap between consecutive detector triggers is exp_time minus a readout
time that was a module constant fixed at 0.2 ms. Some detectors need more:
FalconcSAXS declares MIN_READOUT = 3 ms, fifteen times longer. Nothing
reconciled the two -- DDG2 does not know which detectors are in the scan,
and the falcon only validates its exposure time, never the gap -- so a
detector that cannot keep up silently dropped frames.

There was also no way to change it. on_stage recomputes the pulse width
from the constant on every scan, so a value set by hand from the client did
not survive to the next acquisition.

The readout times are now per-instance, settable two ways: a readout_times
key in deviceConfig for a per-deployment default, and set_readout_times()
via USER_ACCESS for a change between scans. Raising the value widens the gap
and shortens the exposure by the same amount; burst_period stays at exp_time,
so the frame rate is unaffected.

Channel pair 'ab' is the one multiplexed to the detectors and normally the
only one worth changing. A value that would exceed the exposure time is
rejected, and a non-default value is logged at on_stage so a scan taken with
a widened gap is recoverable from the logs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLnmUurqcNd1FiDY5M2uZr
2026-09-01 14:54:00 +02:00
menzelandClaude Opus 5 d3f7871855 feat(falcon): add prime() to arm the HDF5 plugin after an IOC restart
CI for csaxs_bec / test (pull_request) Successful in 1m55s
Read the Docs Deploy Trigger / trigger-rtd-webhook (push) Successful in 1s
CI for csaxs_bec / test (push) Successful in 1m47s
The areaDetector file plugin refuses to be staged until one NDArray has
passed through it, so every restart of the SITORO IOC leaves the Falcon
raising ophyd's UnprimedPlugin at the start of a scan. The generic
HDF5Plugin.warmup() cannot be used, as it drives parent.cam, which the
Falcon does not have.

prime() pushes a single spectrum through the plugin using user-advanced
pixels, so no external gate signal is required and a disconnected trigger
cable cannot mask the priming. No file is written.

The trigger configuration is captured beforehand and restored in a finally
block. on_stage never resets pixel_advance_mode, ignore_gate or
pixels_per_buffer -- only on_connected does -- so a prime that bailed out
halfway would otherwise leave the detector deaf to the gate and silently
starve every subsequent scan.

prime() and is_primed() are exposed via USER_ACCESS, so they are callable
from the BEC client as dev.falcon.prime().

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KLnmUurqcNd1FiDY5M2uZr
2026-08-31 15:46:55 +02:00