mirror of
https://github.com/slsdetectorgroup/aare.git
synced 2026-09-03 00:30:43 +02:00
Optimized pedestal tracking with FastPedestal and performance improvements to the cluster finder. - ~2x faster hitting 14k FPS on benchmark dataset on 16 core threadripper pro **Changes** - No bounds check on interior of frame. (only within half of a cluster) (~15% improvement) - Cache threshold (std*n_rms) Biggest improvement comes from not computing std for each access - Cache pedestal subtracted frame - Switched to FastPedestal which assumes we reached steady state **Things tried but rejected:** - Always store cluster values, commit on val==max (~15% drop in frame rate) **Options** - using 16 bit clusters and 16 bit pedestal. (slows down the single threaded case but allow us to reach 12k with multi threading) **Other notes on performance** - Passing in memory frames from python and doing nothing: 0.8us/frame - Passing + pedestal subtraction: 28us/frame - Passing + pedestal + cluster finding (no store): 652us/frame - Passing + pedestal + cluster finding: 886 us/frame
17 lines
422 B
C++
17 lines
422 B
C++
// SPDX-License-Identifier: MPL-2.0
|
|
#include "aare/NDArray.hpp"
|
|
#include "aare/NDView.hpp"
|
|
#include <benchmark/benchmark.h>
|
|
|
|
using aare::NDArray;
|
|
using aare::NDView;
|
|
|
|
static void BM_CreateNDView(benchmark::State &st) {
|
|
NDArray<int, 2> arr{{1024, 1024}, 0};
|
|
for (auto _ : st) {
|
|
// This code gets timed
|
|
auto res = arr.view();
|
|
benchmark::DoNotOptimize(res);
|
|
}
|
|
}
|
|
BENCHMARK(BM_CreateNDView); |