Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
64 lines
3.3 KiB
C
64 lines
3.3 KiB
C
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <bitshuffle/bitshuffle_internals.h>
|
|
#include <bitshuffle_hperf/bitshuffle.h>
|
|
|
|
// One bitshuffle block, transformed by whichever of the two vendored implementations is SIMD on
|
|
// this architecture. The two write byte-identical output and each decodes the other's, so the file
|
|
// format does not depend on the build host.
|
|
//
|
|
// bitshuffle_hperf is x86-only: outside SSE2 its whole vector body is compiled out and what remains
|
|
// is a scalar fallback. The classic bitshuffle has an aarch64 NEON path (bitshuffle_core.c,
|
|
// USEARMNEON), so aarch64 encodes with that instead - measured against the hperf scalar fallback it
|
|
// is ~2.5x on encode and ~1.7x on decode. Decode, which is what rugnux spends its time on, goes on
|
|
// aarch64 through a NEON port of hperf's decoder (bitshuffle_hperf/bitshuffle_neon.c) for the
|
|
// element sizes it covers (1, 2, 4 bytes), and through the classic NEON stages for the rest.
|
|
// Everywhere else hperf wins outright (~2x over classic SSE2), so it stays the default.
|
|
//
|
|
// The condition mirrors USEARMNEON in bitshuffle_core.c exactly. With NEON off the classic path
|
|
// falls back to a scalar of its own that is slower than hperf's, so it must not be selected then.
|
|
// It is a preprocessor test rather than a CMake one on purpose: Apple Silicon defines the same two
|
|
// macros as aarch64 Linux, and a macOS universal build compiles this header once per architecture,
|
|
// which a single configure-time answer could not follow.
|
|
#if (defined(__ARM_NEON__) || (__ARM_NEON)) && defined(__aarch64__)
|
|
|
|
#include <bitshuffle_hperf/bitshuffle_neon.h>
|
|
|
|
// The two stages of bshuf_untrans_bit_elem_NEON, which bitshuffle_core.c defines but no header
|
|
// declares. Calling them directly lets the decode use the caller's scratch instead of the block-sized
|
|
// buffer the classic entry point mallocs and frees for every block.
|
|
extern "C" {
|
|
int64_t bshuf_trans_byte_bitrow_NEON(const void *in, void *out, size_t size, size_t elem_size);
|
|
int64_t bshuf_shuffle_bit_eightelem_NEON(const void *in, void *out, size_t size, size_t elem_size);
|
|
}
|
|
|
|
// The classic encode entry point allocates its own block-sized scratch, so the caller's goes unused.
|
|
inline int64_t JFJochBitShuffleBlock(char *out, const char *in, char *, size_t size, size_t elem_size) {
|
|
return bshuf_trans_bit_elem(in, out, size, elem_size);
|
|
}
|
|
|
|
inline int64_t JFJochBitUnshuffleBlock(char *out, const char *in, char *scratch, size_t size, size_t elem_size) {
|
|
if (elem_size == 1 || elem_size == 2 || elem_size == 4)
|
|
return bitshuf_decode_block_neon(out, in, scratch, size, elem_size);
|
|
|
|
const int64_t count = bshuf_trans_byte_bitrow_NEON(in, scratch, size, elem_size);
|
|
if (count < 0)
|
|
return count;
|
|
return bshuf_shuffle_bit_eightelem_NEON(scratch, out, size, elem_size);
|
|
}
|
|
|
|
#else
|
|
|
|
inline int64_t JFJochBitShuffleBlock(char *out, const char *in, char *scratch, size_t size, size_t elem_size) {
|
|
return bitshuf_encode_block(out, in, scratch, size, elem_size);
|
|
}
|
|
|
|
inline int64_t JFJochBitUnshuffleBlock(char *out, const char *in, char *scratch, size_t size, size_t elem_size) {
|
|
return bitshuf_decode_block(out, in, scratch, size, elem_size);
|
|
}
|
|
|
|
#endif
|