Files
Jungfraujoch/image_analysis
leonarski_fandClaude Opus 5.5 ea9d0adcde Stop on a lost CUDA context instead of falling back to the host decoder
A sticky error (an illegal address, say) is reported by whichever worker synchronises next - usually
the bslz4 device decode - and MXAnalysisWithoutFPGA::Analyze and Rugnux::MaskDefectivePixels then
logged "falling back to host" and carried on, although the context is gone and the run fails anyway,
later and less clearly. cuda_throw_if_context_lost() now throws a GPUCUDAError ("CUDA device
unusable after an unrecoverable error: ...") there first, which the image loops already treat as
fatal (IsFatalResourceError), like a CUDA out-of-memory.

The last error cannot tell: once cudaGetLastError() has returned a sticky error, it and
cudaPeekAtLastError() both report success. cudaFree(nullptr) frees nothing and does not
synchronise, but returns the sticky error (checked: after an illegal-address kernel it returns
cudaErrorIllegalAddress; with only a pending out-of-memory it returns success), so a handled,
non-sticky decode failure still falls back - MXAnalysis_HandledDeviceDecodeFailureLeavesNoError
passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-27 17:11:38 +02:00
..
2026-09-16 18:17:46 +02:00
2026-06-08 08:30:35 +02:00
2026-09-16 18:17:46 +02:00
2026-09-16 18:17:46 +02:00
2026-09-22 06:48:37 +02:00
2026-09-22 06:48:37 +02:00
2026-07-03 19:18:56 +02:00
2026-07-03 19:18:56 +02:00
2026-09-22 06:48:37 +02:00