From edbd890700846cfe27cb8de0664864c5d191966b Mon Sep 17 00:00:00 2001 From: Mirko Holler Date: Sun, 6 Sep 2026 20:14:53 -0300 Subject: [PATCH] docs(developer): document the vision pipeline architecture Generic reference for VisionInterface/SampleVisionMonitor (the three-layer camera-processing architecture, ROI cropping, position-blind shape scoring vs. position_tolerance_px, and reading diagnostic_image), plus the device_manager-injection pitfall for any future PSIDeviceBase subclass that needs to resolve sibling devices -- written up the same way as the existing lamni_smear_architecture.md. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01C9zgK28uPoZWJFkJWV4P1G --- docs/developer/developer.md | 12 + .../developer/vision_pipeline_architecture.md | 223 ++++++++++++++++++ 2 files changed, 235 insertions(+) create mode 100644 docs/developer/vision_pipeline_architecture.md diff --git a/docs/developer/developer.md b/docs/developer/developer.md index 664c3d38..06af28b1 100644 --- a/docs/developer/developer.md +++ b/docs/developer/developer.md @@ -10,6 +10,7 @@ hidden: true editing_docs lamni_smear_architecture +vision_pipeline_architecture ``` @@ -42,4 +43,15 @@ Conventions for writing these MyST/Sphinx docs and how changes go live on Read t Device/ipython-client/GUI widget interfaces for the smear rotation-center calibration aid, and the `bw-generate-cli` regeneration pitfall. ``` +```{grid-item-card} +:link: developer.vision_pipeline +:link-type: ref +:text-align: center +:class-item: index-card + +## Vision pipeline architecture + +The generic `VisionInterface`/`SampleVisionMonitor` camera-processing layer, the `PreviewSignal` rpc_access gotcha, and the `device_manager` injection pitfall. +``` + ```` diff --git a/docs/developer/vision_pipeline_architecture.md b/docs/developer/vision_pipeline_architecture.md new file mode 100644 index 00000000..a9700c7d --- /dev/null +++ b/docs/developer/vision_pipeline_architecture.md @@ -0,0 +1,223 @@ +(developer.vision_pipeline)= + +# Vision pipeline architecture: `VisionInterface` and `SampleVisionMonitor` + +A generic, camera-agnostic machine-vision layer living under +`csaxs_bec/devices/vision/`, built to sit between any camera device (real or +simulated) and whatever is watching it — a processed/overlay live view for an +alignment GUI, or an automated "does this still look right" confirm check. +Nothing in it is specific to OMNY or any other endstation; the design and +staged rollout plan for OMNY's sample-transfer use case lives separately in +`csaxs_bec/bec_ipython_client/plugins/omny/AI_docs/MACHINE_VISION_AUTOCONFIRM_PLAN.md`. + +## Three layers + +```text +camera device VisionInterface consumer +(IDSCamera, (device server process) + AlliedVisionAravisCamera, │ + DummyVisionCamera, ...) │ set_mode("edges") + │ │ refresh(camera_name) + │ get_last_image() ▼ + └──────────────────► vision_toolkit.detect_edges(...) ──► diagnostic_image + (PreviewSignal) + │ + SampleVisionMonitor ◄──────────────────────┘ + (save_reference(tag) / score(tag)) + │ + ▼ + per-camera match score + published + diagnostic overlay of the worst camera +``` + +- **Layer 0 — `vision_toolkit.py`.** Plain `numpy`/OpenCV functions, no + devices, no BEC imports: `crop_to_roi()`, `to_grayscale()`, `apply_clahe()`, + `detect_edges()`, `register_to_reference()` (ECC-based, a configurable + `warp_mode`), `largest_contour()`/`contour_similarity_score()` (Hu-moment + shape matching), `centroid_distance()` (a separate, position-sensitive + signal — see below), `best_reference_match()` (multi-reference scoring), + `draw_contour_overlay()`. Trivially unit-testable in isolation; this is the + shared vocabulary both layers below build on. +- **Layer 1 — `vision_interface.py`, `VisionInterface`.** A `PSIDeviceBase` + device that wraps one or more source cameras (`camera_names: list[str]`, + resolved lazily through the device manager, not at construction) and + republishes a processed view of the latest frame through its own + `diagnostic_image` signal — independent of the source camera's own live + channel. The processing mode (`"raw"`/`"grayscale"`/`"edges"`/`"clahe"`) is + runtime-selectable via `set_mode()`, a plain RPC call, no GUI required. + Only depends on the camera implementing `get_last_image() -> np.ndarray | + None` — the contract already shared by every camera class in this repo — + so it can be pointed at any current or future camera via config alone. +- **Layer 2 — `sample_vision_monitor.py`, `SampleVisionMonitor`.** Composes a + configured `VisionInterface` with a per-tag, per-camera in-memory reference + store. `save_reference(tag)` snapshots the current raw frame of every + camera behind the interface and stores it under an opaque, caller-defined + `tag` string (the device has no notion of what the tag means). `score(tag)` + registers and scores the current frame of every camera against every + reference stored under that tag (`vision_toolkit.best_reference_match()`), + publishes a diagnostic contour overlay for the worst-scoring camera through + the shared `VisionInterface`, and returns a dict a caller can turn into a + confirm-dialog decision — this device makes no decision itself. + +## Region of interest: crop before anything else looks at the frame + +`VisionInterface` can crop each camera's frame to a fixed region before any +processing mode, publishing, or scoring happens — e.g. to score only the +gripper jaws and ignore background clutter/motion elsewhere in the field of +view. + +- **`set_roi(roi, camera_name=None)`** / **`get_roi(camera_name=None)`** — + `roi` is `(x, y, width, height)` in pixels, top-left origin (the same + convention OpenCV and `bec_widgets`' `RectangularROI.get_coordinates()` + use, so coordinates picked interactively on an `Image` widget can be + passed straight through). `camera_name=None` (the default) applies to + every camera configured on the interface; pass a specific name to set a + different ROI per camera — natural for a pair like OMNY's parking-view + `cam200`+`cam203`, which don't necessarily frame the gripper identically. + `roi=None` clears it, back to the camera's full frame. Also settable at + construction time via the `roi` device-config argument (a single tuple for + all cameras, or a `{camera_name: roi}` mapping). +- **Applied once, centrally, in `_get_camera_frame()`** — the single method + both `get_raw_frame()` and `get_last_processed_image()` (and therefore + `refresh()`) go through. This means it is automatically also what + `SampleVisionMonitor.save_reference()`/`score()` see, with no changes + needed on that side: both call `VisionInterface.get_raw_frame()`, never the + camera directly. +- **Out-of-bounds is a hard error, not a silent clamp.** A stale ROI left + over from a different camera resolution raises `VisionInterfaceError` + (wrapping `vision_toolkit.crop_to_roi()`'s `ValueError`) the next time a + frame is pulled, rather than quietly cropping to whatever fits — a + misconfigured ROI should fail loudly, not silently score a different + region than the caller thinks it's comparing. +- **Interacts with the position-blindness above**: an object that moves + *within* the ROI still scores as a shape match (position is still ignored + unless `position_tolerance_px` is set) — but an object that moves *out of* + the ROI crop entirely genuinely changes the score, since there's no + contour left to match at all. This is a real, useful way to make the + system sensitive to "did the pin end up in completely the wrong place," + short of committing to `position_tolerance_px`'s pixel-distance semantics. + +## The shape score is deliberately position-blind; use `position_tolerance_px` for that + +`contour_similarity_score()` compares Hu moments (`cv2.matchShapes`), which +are *unconditionally* invariant to translation, rotation, and scale — this +holds regardless of whether frames were registered first. `warp_mode` +(`"none"`/`"translation"`/`"euclidean"`/`"affine"`, exposed on +`register_to_reference()`, `best_reference_match()`, and as a +`SampleVisionMonitor` constructor argument / `set_warp_mode()`) genuinely +affects registration/rasterization quality for a real scale or rotation +difference, but it cannot and does not make the shape score +position-sensitive — that was never what it does. + +If position should actually matter, it's a second, independent signal: +`centroid_distance()` always compares *raw, unregistered* contours +(registering first would remove the very displacement it measures), and +`SampleVisionMonitor.position_tolerance_px` (pixels, `None` by default, +meaning position isn't checked at all) surfaces it as +`position_within_tolerance` in `score()`'s per-camera result, alongside — not +combined with — the shape `score`. Deciding what to do when shape is fine but +position isn't (or vice versa) is a policy choice left entirely to the +caller. + +## Reading `diagnostic_image`: it is not a normal device attribute + +`diagnostic_image` (and `DummyVisionCamera.image`) is a `PreviewSignal`, a +`BECMessageSignal` subclass. Every `BECMessageSignal` hardcodes `rpc_access` +to `False`, so — same gotcha documented in +{ref}`the LamNI smear architecture notes ` +for `IDSCamera`'s preview channels — it does **not** appear as a client-side +device attribute at all: + +```pycon +>>> dev.vision_test_interface.diagnostic_image +AttributeError: 'Device' object has no attribute 'diagnostic_image' +``` + +It's published as a `DevicePreviewMessage` on a Redis stream instead. Read +the latest one with: + +```python +from bec_lib.endpoints import MessageEndpoints + +data = bec.connector.get_last( + MessageEndpoints.device_preview("vision_test_interface", "diagnostic_image") +) +data["data"].data.shape +``` + +(the same pattern used for any other camera's preview signal — see +`tests/end-2-end/test_scans_e2e.py::test_device_preview` for the equivalent +against `eiger`/`preview`). + +## Pitfall: a `PSIDeviceBase` subclass must declare `device_manager` explicitly + +`VisionInterface` and `SampleVisionMonitor` both need `self.device_manager` +at call time, to resolve their configured `camera_names` / +`vision_interface_name` lazily (device build order across a config isn't +guaranteed, so this can't be resolved at construction). `PSIDeviceBase` +itself accepts a `device_manager` keyword and stores it — but that alone is +**not** enough for it to ever be populated with a real value. + +:::{important} +The device server's `DeviceManagerDS.construct_device_obj()` only injects +`device_manager` into a device's constructor when that device class's *own* +`__init__` **signature** names a `device_manager` parameter explicitly: + +```python +signature = inspect.signature(dev_cls) +if "device_manager" in signature.parameters: + init_kwargs["device_manager"] = device_manager +``` + +`inspect.signature()` on a class reflects only that class's own `__init__` +— it does not look through `**kwargs` into a base class. A subclass whose +`__init__` accepts only `**kwargs` and forwards it to +`super().__init__(..., **kwargs)` will happily accept `device_manager` if +passed directly (e.g. unit tests instantiating the class themselves), but +the device server will never pass it, silently leaving +`self.device_manager` at its default of `None` for the lifetime of the +object — with no error at startup. Every call that resolves a sibling +device (`self.device_manager.devices.get(...)`) then fails at call time +instead, with `device manager is not available`. +::: + +The fix, and the required pattern for any new `PSIDeviceBase` subclass that +needs to look up sibling devices (the same shape already used by +`ddg_1.py`): declare `device_manager` as an explicit keyword parameter and +forward it through: + +```python +def __init__( + self, + *, + name: str, + ..., + device_manager: "DeviceManagerBase | None" = None, + **kwargs, +): + super().__init__(name=name, ..., device_manager=device_manager, **kwargs) +``` + +## Testing without a real camera + +`DummyVisionCamera` (`csaxs_bec/devices/vision/dummy_vision_camera.py`) is a +settable, hardware-free stand-in implementing only `get_last_image()` — its +"current frame" is whatever was last set via `set_image()`, +`set_test_pattern()` (a synthetic white square on black, for a +translation/scale/rotation-controllable test target), or +`load_image_file(path)` (a real image file from disk — the closest available +proxy to real hardware behavior, useful for sanity-checking +`detect_edges()`'s thresholds against real lighting). + +`device_configs/simulated_omny/vision_test.yaml` wires two `DummyVisionCamera` +instances to one `VisionInterface` and one `SampleVisionMonitor`, loadable +into any already-running session without touching existing devices: + +```python +bec.config.update_session_with_file( + "csaxs_bec/device_configs/simulated_omny/vision_test.yaml" +) +``` + +See `csaxs_bec/bec_ipython_client/plugins/omny/AI_docs/VISION_PIPELINE_TESTING_HOWTO.md` +for a full worked walkthrough exercising every layer this way.