A prose pass over docs/ for text that reads as written by, or addressed
to, a language model. The documentation is almost entirely in the house
voice already; what needed changing was six hollow directives in the
older infrastructure pages ('it is important to ensure...', 'extra care
has to be taken...', a leftover 'Be explicit about the gaps' addressed
to the document's own author). Each is rewritten as a statement of fact
with the technical content untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011GxZqDiFP3KqriBhNdcR56
7.5 KiB
Security
Jungfraujoch is a data-acquisition and analysis system for X-ray detectors, designed to run inside a controlled facility network. This document describes what the software does and does not protect against, the current known limitations, and the authentication work in progress.
Threat model and scope
The security model targets a semi-trusted internal facility network. The concern is a peer on that network reaching a Jungfraujoch service with little or no effort — a mistyped host/port, a curious colleague, a mis-pointed script, a stray browser tab — not a determined attacker and not passive wire capture (which is the responsibility of the network layer: 802.1x, VLANs, facility infrastructure).
The asset that matters most is the confidentiality of live analysis data: the diffraction images and derived metadata (unit cell, resolution, spot counts, sample name) that reveal which sample is being measured. This matters for industrial and proprietary experiments. By contrast, acquisition control (start / stop / configure) is treated as low risk — scientists operate their own experiments and there is little to gain from restricting it.
Security is best-effort: measures that materially impede normal operation get turned off, so the design favours a few high-value, low-friction controls over comprehensive lockdown.
Out of scope. Jungfraujoch is not designed to be exposed to an untrusted network or the public internet. Do not do this.
1. Good practice — what is and is not protected
What you can secure (and should)
These controls work and a deployment should apply them (see also DEPLOYMENT.md):
- Network isolation. Keep the broker and its data streams on a controlled segment. The broker ↔ writer ↔ receiver traffic should run on a dedicated back-end network, with the data-socket addresses pinned to that interface and the ports firewalled.
- Reverse proxy for TLS. The broker speaks plain HTTP. To get HTTPS, put a reverse proxy
(Apache / nginx) in front that terminates TLS and pin the broker to
localhostbehind it. The desktop viewer supportshttps://endpoints — choose the scheme in the Open HTTP Connection dialog. - Filesystem confinement of written data. The writer creates NXmx HDF5 files on shared storage. Confidentiality of that data at rest is enforced by the filesystem: run the writer under a dedicated identity and use directory ownership / ACLs (and setgid) so that only the owning experiment can read its files.
- Firewall the ZeroMQ ports. The image / preview / metadata / republish streams have no access control of their own (see below), so restrict who can reach those ports at the network layer.
What the software does NOT provide
The gaps are listed explicitly so a deployment does not assume protection that is not there:
- No authentication or authorization in the broker. The HTTP/REST API currently has no login, token, or access control. Anyone who can reach the broker's host and port has full read access (live images, unit cell, resolution, sample metadata) and full write access (start, cancel, reconfigure). Confidentiality currently depends entirely on network/firewall isolation. This is being addressed — see §3.
- No transport encryption in the broker. The broker serves plain HTTP; there is no built-in TLS. Encryption must be provided by a reverse proxy.
- No access control or encryption on the ZeroMQ streams. The preview, metadata, image, and
republish streams are unauthenticated sockets. Any peer that can connect can subscribe to live
data. For the image
PUSHstream specifically, an accidental extra consumer does not merely eavesdrop — aPULLpeer is load-balanced into the stream and will divert images away from the real writer. - No per-user isolation. The broker has no concept of users; it cannot separate one operator's access from another's.
- No application-level audit trail of who accessed or changed what.
Recommended deployment checklist
- Broker and back-end streams on an isolated network; never exposed to a general/untrusted network.
- ZeroMQ data-socket addresses pinned to the back-end interface; ports firewalled to known peers.
- TLS terminated by a reverse proxy; broker bound to
localhostbehind it. - Writer run under a dedicated identity; data directories owned / ACL'd per experiment (setgid) so users read only their own data.
- ZeroMQ compatibility streams (preview / metadata / republish) enabled only if actually consumed, and only on the trusted back-end.
2. Known issues
| # | Issue | Impact | Mitigation today |
|---|---|---|---|
| 1 | Broker HTTP API has no authentication | Anyone who can reach it has full read (confidential live data) + write (control) | Network / firewall isolation; §3 in progress |
| 2 | Broker binds all interfaces, plain HTTP | Reachable from anywhere routable; no encryption | Expose only on the trusted segment; TLS via reverse proxy |
| 3 | ZeroMQ preview / metadata / image / republish streams are unauthenticated and unencrypted | Live-data exfiltration; a rogue PULL on the image stream diverts/steals images |
Firewall the ports; run only on the back-end network |
| 4 | Web frontend assumes an open API | The bundled UI has no auth and expects to reach an open broker | Serve and reach it only on the trusted network |
Input robustness. Services parse framed data from peers on the (trusted) data path. Hardening of untrusted-frame handling (size caps, overflow guards) is ongoing; these paths are not intended to face an untrusted network.
3. Work in progress — authenticated read access
The main gap (issue #1) is being closed with a best-effort, low-friction scheme that protects the confidential read endpoints while leaving acquisition control open.
Enabling step (done): viewer HTTP client on libcurl. The desktop viewer's broker client was
migrated from a plain-HTTP library to libcurl, which gives it HTTPS plus the ability to
authenticate. The build stays self-contained per platform: on Windows, TLS and Kerberos come
from the operating system (Schannel + SSPI, no external dependencies); on Linux, from system
OpenSSL + Kerberos (GSSAPI / krb5). This is why the Linux build now needs the Kerberos
development headers (libkrb5-dev on Debian/Ubuntu, krb5-devel on RHEL/Rocky).
Planned enforcement (not yet implemented). Two complementary options, both gating only the sensitive read endpoints (live images, scan result, preview plots, statistics) and leaving writes open:
- Per-experiment Bearer token. An optional token supplied at acquisition start (
/start); when set, the broker requires it (Authorization: Bearer …) on the read endpoints. No token set → no enforcement (backward compatible). The token is minted externally and handed to the authorised viewer(s); the broker only compares strings. A rotated per-experiment token segments one experiment's data from the next. - Kerberos / GSSAPI single sign-on. A pass-through reverse proxy performs Kerberos/SPNEGO
authentication (service keytab) and forwards the authenticated user name to a
localhost-pinned broker as a trusted header. This gives Active Directory single sign-on with no token to carry, at the cost of a proxy deployment. The viewer's libcurl client can negotiate Kerberos directly against such a proxy.
Neither enforcement path is in the broker yet. Until then, confidentiality relies on the network and filesystem controls in §1.