docs: clean up deck
Build on RHEL8 / build (push) Successful in 3m18s
Build on RHEL9 / build (push) Successful in 3m48s
Run tests using data on local RHEL8 / build (push) Successful in 4m10s

This commit is contained in:
kferjaoui
2026-08-25 15:34:12 +02:00
parent d14b2c0088
commit 03e7c9a454
29 changed files with 1426 additions and 540 deletions
Binary file not shown.
+304
View File
@@ -0,0 +1,304 @@
# 2026-08-25 — projection legibility, and the 30-minute restructure
Session log for `docs/cf_cuda_fused.pptx` (via `docs/deck/build_fused_deck.py` +
`make_figs.py`). Deck rebuilt at **51 slides**: title + hero + **33 numbered**
(333) + 7 dividers + **11 annex** (A1A6).
---
## 1. One projection floor, enforced in two places
The back row is 67 m from the screen. On a 13.33 × 7.5 in slide, **9 pt is
about 1/60 of the slide height** — the conventional lower bound for supporting
detail. That is now the floor for *everything*, on the slide and inside every
PNG, and it is enforced by code rather than by hand.
**Text set in PowerPoint**`run()` clamps to `MIN_PT = 9.0`. Enforced in the
helper, not at the ~400 call sites, because a floor applied by hand is a floor
one new caption silently drops through. Consequences: table headers 8 → 9,
statstrip labels 8 → 9, rail labels 8.5 → 9, code titles 7.5 → 9, page numbers
8.5 → 9, the 8.5 pt captions → 9.
`code()` had a **constant** line height of 0.148 in tuned for 8.5 pt, so raising
the type pushed the last line out of the panel. It now scales: `lh = 0.0174 ×
size`, which reproduces the old value at 8.5.
**Text inside figures** — the real problem, and it was invisible where the font
size is written. A matplotlib fontsize is in points of the *figure's* inches;
the PNG is then placed at some other width, so what the audience reads is
```
effective_pt = raw_pt × (placement_width_in / figure_width_in)
```
`make_figs.py`'s `save()` now takes the placement width, walks every `Text`
object in the figure after drawing it, and prints/collects anything below the
floor. `legibility_report()` at the end of the run is the gate. The placement
widths are **parsed out of `build_fused_deck.py`** rather than hardcoded, so the
check can never drift from the layout it is checking.
Measured before the fix:
| figure | PNG | placed | scale | smallest text |
|---|--:|--:|--:|--:|
| `fig_mismatch147` | 8.97 in | 9.40 | 1.05 | **4.8 pt** |
| `fig_opt2_timeline` | 10.64 in | 7.90 | 0.74 | **4.8 pt** |
| `fig_regpressure` | 10.71 in | 7.90 | 0.74 | **6.3 pt** |
| `fig_f32_kernel` | 8.00 in | 7.50 | 0.94 | **6.4 pt** |
| `fig_streams` | 6.36 in | 6.55 | 1.03 | **6.6 pt** |
13 of 22 placed figures failed. All 22 now pass.
Three causes, three fixes:
- **Self-inflating captions.** `savefig(bbox_inches="tight")` grew the saved
width to fit a long in-figure caption, which shrank the scale factor, which
shrank the caption further — the longer the text, the smaller it rendered.
Four such strings (`fig_opt2_timeline`, `fig_streams`, `fig_f32_kernel`,
`fig_graphs`) were provenance or verdict text, not data labels. They moved to
slide captions and speaker notes, where a point is a point.
- **Declared width far above placement.** `fig_regpressure` 11.2 → 7.9 in,
`fig_varfloor` 5.6 → 4.35 in, `fig_graphs` re-laid out at 7.6 in.
- **Unfittable text.** `fig_mismatch147` printed an ADU value in each of its
masked-in cells at 4.6 pt. Dropped here, then **restored in §5** once the
figure was given a wider placement — see there.
Then 75 individual sizes were raised — only those below their figure's floor, so
the internal hierarchy is preserved.
**Layout fallout, all fixed:** slides 9, 10, 20, 25, 33 and A5 were re-laid out
where the larger type pushed content into a figure or past the footer. Verified
with `pdftotext -bbox` across all 51 pages: no text below the footer line except
the PSI template's own title page.
---
## 2. Restructure for a 30-minute slot
### New: the measurement convention (12/33 here, **moved to 18/33 in §5**)
Promoted out of A1. Everything from slide 13 on is quoted as `[s1]` or `[s4]`
and compared against a "floor", and none of those three words was defined
anywhere the audience would see it. Three definition cards (s1 / s4 / floor) and
one figure; the numbers stay in A1.
`fig_measure` is new: four stream lanes from the **real** `_schedule()`, so the
H2D bars queue because there is one copy engine per direction (measured overlap
1.000 in every row of `probes.csv`) while kernels sit on top of each other in
time. The two union strips beneath are computed off that schedule. No
microseconds on that panel — it reproduces the *shape* of the 9×9 pipeline, not
its exact overlap factor, and measured numbers next to a schematic invite being
read across. The right panel is measured: 20.77 / 32.66 / 25.25 at s4, the
dashed profiled max, and the 30.01 sustained rate that is the actual floor.
### Reordered: 9 and 10 swapped
Registers (the cause) now precede occupancy (the effect). The old order taught a
metric and then said it was not the point. Slide 10 opens with occupancy in one
sentence — *when a warp stalls the SM switches to another resident warp;
occupancy is how many alternatives it has* — so the 33 % lands as a consequence
of the register budget, not as a failing grade.
### Merged: 20 + 21 → one slide
The correctness trap and the fix. The ULP arithmetic of the error floor is the
least transferable part of the story; what stays is *a small answer computed as
the difference of two large numbers is not computable in f32*, what it did to
the physics, and the two-line change.
`fig_cancellation`'s right panel is now the **deformed spectrum**. The f64 curve
is measured (`validation_tiers.json`, 23.2 M clusters). The f32 curve is
**reconstructed, and labelled as such on the slide**: its area is the measured
+28.06 % from §1 of the write-up, its shape is where §9 derives the excess (a
gate collapsed from ~225 ADU to ~0 admits the whole positive side, so the
population smears upward from zero). Nobody kept the broken build to re-run.
The old right panel — variance vs rms, the ±34 ADU² floor — is now
`fig_varfloor`, in A5.
### Cut to the annex
| was | now |
|---|---|
| 18 · two rejected routes (B, B″) | **A2·3** |
| 21 · the variance rewrite | **A5** (+ `fig_varfloor`) |
| 26 · three ways a benchmark lies | **A6·1**, as it stood |
Slide 25 keeps only the artefact that moves a number the audience is shown —
first-touch page faults — plus one amber callout naming the other two (nsys
inflates host-side API calls ~4×, CUDA events measure a stream) and pointing at
A6. CUDA-event detail is not discussed in the arc at all.
Slide 19 (opt7) lost the s1-vs-s4 caption, which the measurement slide now owns.
### Detector context
MÖNCH03 specifications on slide 3: hybrid silicon, 400 × 400 at 25 µm pitch,
10 × 10 mm² active area, **1.3 kHz standard and 36 kHz with optimised readout**.
Slide 33 closes the loop: 58 495 FPS is **45× the standard rate and ~10× the
optimised ceiling**, with the two qualifications in the notes (frames are assumed
already in host RAM; at 9×9 the margin shrinks if the cap has to grow).
---
## 3. Renumbering
`N_SLIDES` 34 → **33**, `N_ANNEX` 5 → **6**. All `chrome()` indices reassigned
sequentially; six divider slide-lists and ranges rewritten; every cross-reference
retargeted (`expands slide N` on A2·3, A3, A4, A5, A6; `[engines: 19/33]`;
"slides 2730"; the A2 part counts 1/2, 2/2 → 1/3, 2/3).
---
## 4. Verified
- `pdftotext -bbox`, all 51 pages: no text past the footer (PSI title page aside).
- No unrendered `**` markup anywhere.
- `legibility_report()`: every string in every figure ≥ 9.0 pt as placed.
- No stale cap-1500 values (`82.17`, `80.44`, `94.78`, "18 % SLOWER"); the one
remaining "1 500" is the sentence explaining *why* the cap is 1 700.
- 20 em dashes on-slide, all pivots, appositions, or table "—" for empty.
- Visual pass over all 51 rendered pages.
## Still open
- `fig_bottleneck`, `fig_correctness` and `fig_variance_rewrite` are generated
but placed on no slide. Harmless; `fig_variance_rewrite` joined the list when
slide 21 moved to A5, where the code block says the same thing in less space.
- Timing: 40 slides in the arc. At 30 minutes that is ~45 s each, which is tight.
Next cut candidates, in order: 14/15 (opt2 and opt3 could be one slide), 24
(where the time went — the arc slides already carry it), 26 (first run — the
point survives as a sentence on 25).
---
## 5. Second review pass (same day)
**Code-panel clearance, deck-wide.** The type floor grew every `code()` panel by
0.060.09 in, and eleven of them had less clearance than that to whatever was
drawn under. Found with a parser that computes each panel's height from its own
line count and size and compares against every later element with an overlapping
x-range, so this is not a slide-by-slide eyeball. All eleven now clear by
≥ 0.12 in; slides 8, 15 and 30 needed re-laying out around it.
**12/33 moved into Act III, now 18/33.** It talked about four streams (which
arrive at opt2) and 9×9 (which is not discussed until Act III), so it was
spending the audience's attention before either was earned. It now sits
immediately after the Act III divider, ahead of opt7 — whose two-panel figure is
the first place s1 and s4 are put side by side. Retitled "How the engine times
are measured".
Slide 12 (opt1) is where "% of floor" first appears, so it keeps a one-sentence
definition — *floor = max(H2D, kernel, D2H), the fastest a frame can go if the
host cost nothing* — and defers how each term is measured to 18. That callout
had been carrying the whole s1/s4/union argument; it is now four lines shorter.
**One colour code for the two builds in Act III: f32 = amber, f64 = blue.**
`fig_f32_absolute` and `fig_cancellation` already were; `fig_f32_kernel` had them
the other way round, so the same two builds swapped colour between slides 19 and
20. (The engine palette — H2D amber, kernel blue, D2H white — is a separate axis
and unchanged; both figures carry their own legend.)
**Titles.** 19 back to "FP32 halves pedestal traffic: 41 % kernel time".
20 → "Accumulate what is small, not what is large", with catastrophic
cancellation moved into the eyebrow; A5 renamed "The rewrite in full, and which
pixels the error reached" so the two no longer collide. The annex divider is
just "Annexes".
**29/33: the per-cell ADU values are back.** They were dropped because at a 9.4 in
placement a 25 × 25 grid gives each cell 13.5 pt of width, which cannot hold a
three-character number legibly. The figure is now placed at **10.6 in** — 15.3 pt
per cell — which fits the values at **9.0 pt as rendered**, so they are back
*and* they clear the projection floor. The centre marker moved behind the text
(it used to paint over the value in its own cell) and the slide's caption moved
to the notes to pay for the extra height.
**32/33: the knob that loses data is ringed.** New `frame_rect()` helper (an
outline, not a fill, so the row is not recoloured) and a new `RED` token used for
pointing only, never as a data colour. The bottom callout switched from amber to
the same red, so the ring and the warning read as one thing.
**Also:** `fig_f32_kernel`'s "39.86 ◀ binds" ran past its own x-limit onto the
right panel's tick labels (xlim 46 → 60, wspace 0.22 → 0.30); `fig_pinning`'s DMA
arrow label was wider than the gap it sat in.
Re-verified after every change: no text past the footer on any of the 51 pages,
no unrendered `**`, every figure string ≥ 9 pt as placed, and a visual pass over
all 51 rendered pages.
---
## 6. Third review pass (same day)
**19/33 · `fig_f32_kernel` label clearance.** A bar pair spans `y ± h`, so the
step percentage (`40.5 %`, `26.7 %`) at `0.02 h` sat exactly on the top edge
of the f32 kernel bar, and the verdict line at 5.5 % of axes height sat on the
H2D bars. Both now carry a full `h` of clearance beyond the bar pair, and the y
limits (`3.05, 0.98`) hold that margin explicitly rather than relying on default
padding — which the larger type had eaten. The legend moved to the one region no
label reaches, between the D2H and H2D rows, and the dashed 25.24 wall dropped
below the labels in z-order so it stops crossing the `23.94`.
**18/33 · the floor's units, spelled out.** The card read "1 / the busiest
engine" while the panel under it was labelled in µs and the annotation said
"FLOOR 30.01" — reciprocal in one place, microseconds in the other, with nothing
saying they are the same quantity. They are: the busiest engine's busy time per
frame is µs/frame, and one divided by it is FPS. **30.01 µs/frame is 33 323
FPS.**
- Card retitled "Set by the busiest engine", body now names both units.
- The figure annotation reads `FLOOR 30.01 µs/frame = 33 323 FPS`.
- The slide's closing callout says the deck quotes the floor both ways.
- The speaker notes gain a paragraph on when each unit is used: µs/frame when
comparing engines to each other, FPS when comparing against the CPU baseline
or the detector's frame rate. `% of floor` is the same ratio either way.
One consequence worth recording: the notes previously said "the floor is
1 / max(H2D, kernel, D2H) at s4 … 1 / 32.66 µs = 30 600 FPS", which mixed the two
forms in a single sentence and rounded loosely. It now reads "the max is 32.66 µs
= 30 618 FPS", with the reciprocal taken once and named.
Cards grew 1.52 → 1.70 in to hold the longer floor definition; `fig_measure`
re-placed at 10.3 in so the stack still clears the footer.
---
## 7. The CPU die, mirroring the GPU die (5/33)
Slide 6 had a die photograph with the compute units boxed; slide 5 had only a
block diagram, so the two machines were compared as schematics and never as
objects. Added `img_cpu_die.png` — Intel "Comet Lake" 10-core Core i9, same
CS149 source and the same crop recipe (`scratchpad/crops.py`: content box, light
card, rounded alpha edge).
**Kept whole rather than cropped to the ten cores.** The L3 slab on the right and
the I/O block on the left are close to half the die area, which is the callout's
own argument — *"the ALUs are the small part"* — standing next to it in silicon.
**On placement.** A strict mirror of slide 6 is not available: the AD102 die is
square (1.01:1) so it fits the narrow left slot, while the Comet Lake die is
2.35:1 and at any width where its core labels could be read it is 3.7 in tall.
Stacking it under the Skylake diagram at a shared width does not fit either —
both are landscape, and two of them plus the callout overruns the column by
~1.1 in. What did fit, without shrinking the Skylake diagram below the size at
which its `Fetch/Decode` and `ALU` labels survive projection:
| | |
|---|---|
| Skylake core diagram | unchanged at 7.3 in, `M, 1.86` |
| callout | narrowed 7.9 → 4.55 in, `M, 5.42`, h 1.06 |
| **Comet Lake die** | **`5.60, 5.42`, 3.00 in wide** |
The die's own text is illegible at 3.0 in, which is fine and matches slide 6:
neither die is read, both are counted. The caption carries the claim ("10 cores
boxed; nearly half the die is cache and I/O") exactly as slide 6's does ("144
blocks, 128 enabled on this card. One SM boxed.").
What moved to pay for it: the four-line Skylake-vs-Zen 4 caveat caption is now
speaker notes, along with a new paragraph on why ten cores do not fill the die.
**Also on 6/33** (pre-existing, same class): the closing caption sat at 6.90 and
its second line rode on the progress track. Trimmed to one line at 6.82, with the
FP64-scarcity remark moved into the notes — where it now says explicitly that the
V100 diagram *understates* it, since a GeForce part runs FP64 at 1/64, and that
this is half of why opt7 pays.
File diff suppressed because it is too large Load Diff
+413 -155
View File
@@ -56,9 +56,11 @@ opt7 cuts the kernel 40 % to 23.94 us -- below the D2H bar -- so the f32 floor i
D2H, not the kernel. Optimizing the kernel in bottleneck order ended by handing
the constraint to the result path, which is what success looks like.
"""
import re
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
import matplotlib.text
import numpy as np
from matplotlib.patches import FancyArrowPatch, Rectangle
from pathlib import Path
@@ -86,11 +88,103 @@ plt.rcParams.update({
})
def save(fig, name):
fig.savefig(OUT / f"{name}.png", dpi=220, transparent=False,
DPI = 220
# ------------------------------------------------------------ legibility gate
# A matplotlib fontsize is in points of the FIGURE's own inches. The PNG is then
# placed on the slide at some other width, so what the audience actually reads is
#
# effective_pt = raw_pt x (placement_width_in / figure_width_in)
#
# and the second factor is invisible at the point where the font size is written.
# fig_opt2_timeline used to save 10.64 in wide and be placed at 7.9, turning a
# 6.4 pt provenance line into 4.8 pt on the screen -- two thirds of the size of
# the smallest text set directly in PowerPoint. Worse, bbox_inches="tight" grew
# the saved width to fit that very caption, so the longer the caption the smaller
# it rendered.
#
# So the placement width is now an argument, and every figure is measured against
# the deck's projection floor after it is drawn. The floor is set for a room
# where the back row is 6-7 m from the screen: on a 13.33 x 7.5 in slide, 9 pt is
# about 1/60 of the slide height, which is the conventional lower bound for
# supporting detail, and 10.5 pt (~1/50) is the bound for anything the audience is
# asked to read a number off.
MIN_EFF_PT = 9.0
_VIOLATIONS = []
def _placements():
"""Read the placement width of every figure straight out of the deck script.
Hardcoding the widths here would drift the first time a slide is re-laid out,
and drift silently, because nothing renders the figure at both sizes. Parsing
the deck means the generator always checks against the width the deck will
actually use. Where a figure appears twice, the narrowest placement wins,
since that is the one that sets the smallest text.
"""
deck = (Path(__file__).resolve().parent / "build_fused_deck.py")
if not deck.exists():
return {}
env = {"M": 0.7, "COL": 7.9, "RAIL_W": 3.5, "RAIL_X": 9.2,
"W": 13.333, "H": 7.5}
out = {}
for m in re.finditer(r'\b(?:card_)?figure\(s,\s*"([a-z0-9_]+)",([^)]*)\)',
deck.read_text()):
args = [a.strip() for a in m.group(2).split(",")]
if len(args) < 3:
continue
try:
w = float(eval(args[2], {}, env))
except Exception:
continue
out[m.group(1)] = min(out.get(m.group(1), 99.0), w)
return out
PLACE_W = _placements()
def save(fig, name, place_w=None):
"""Write the PNG, then check every string in it against the projection floor.
`place_w` is the width the deck places this figure at. Passing it turns the
check on; the figure list at the bottom of this module keeps it in sync with
build_fused_deck.py, and tools/audit prints both sides.
"""
path = OUT / f"{name}.png"
texts = [(t.get_text(), t.get_fontsize())
for t in fig.findobj(matplotlib.text.Text)
if t.get_text().strip() and t.get_visible()]
fig.savefig(path, dpi=DPI, transparent=False,
bbox_inches="tight", pad_inches=0.08)
plt.close(fig)
print("wrote", name)
if place_w is None:
place_w = PLACE_W.get(name)
if place_w is None:
print("wrote", name)
return
from PIL import Image
pw = Image.open(path).size[0] / DPI
scale = place_w / pw
bad = sorted({(round(sz * scale, 2), round(sz, 1), txt[:44].replace("\n", " "))
for txt, sz in texts if sz * scale < MIN_EFF_PT - 0.05})
print(f"wrote {name:22s} {pw:5.2f} in -> {place_w:5.2f} in "
f"(x{scale:.3f}) min eff {min((sz * scale for _, sz in texts), default=99):.1f} pt")
for eff, raw, txt in bad:
_VIOLATIONS.append((name, eff, raw, txt))
print(f" ILLEGIBLE {eff:5.2f} pt (set {raw:4.1f}) {txt!r}")
def legibility_report():
if not _VIOLATIONS:
print(f"\nlegibility: every string in every figure renders at "
f">= {MIN_EFF_PT} pt on the slide.")
return
print(f"\nlegibility: {len(_VIOLATIONS)} strings below {MIN_EFF_PT} pt "
f"in {len({v[0] for v in _VIOLATIONS})} figures")
for name, eff, raw, txt in sorted(_VIOLATIONS):
print(f" {name:22s} {eff:5.2f} pt {txt!r}")
def bare(ax, keep=("left", "bottom")):
@@ -126,11 +220,11 @@ def fig_spectra_valid():
label=f"{name} ({d['totals'][name]:,})")
ax.set_yscale("log")
ax.set_ylim(3e2, 5e6)
ax.legend(frameon=False, fontsize=8, labelcolor=TEXT2, loc="upper right")
ax.set_ylabel("clusters / bin", fontsize=8)
ax.legend(frameon=False, fontsize=9, labelcolor=TEXT2, loc="upper right")
ax.set_ylabel("clusters / bin", fontsize=9)
bare(ax, keep=("left", "bottom"))
ax.set_title("cluster energy spectrum · 3×3, 10 000 frames, 23.2 M clusters",
color=MUTED, fontsize=8.5, loc="left", pad=6)
color=MUTED, fontsize=9, loc="left", pad=6)
m = h["cpu"] > 0
axr.axhspan(0.999, 1.001, color=GREEN, alpha=0.18, zorder=1)
@@ -140,13 +234,13 @@ def fig_spectra_valid():
dev = max(np.abs(h[n][m] / h["cpu"][m] - 1).max() for n in ("frozen", "cuda"))
axr.set_ylim(0.9955, 1.0045)
axr.set_yticks([0.996, 1.0, 1.004])
axr.set_yticklabels(["0.4 %", "0", "+0.4 %"], fontsize=7.5)
axr.set_xlabel("cluster sum [ADU]", fontsize=8)
axr.set_ylabel("vs CPU", fontsize=8)
axr.set_yticklabels(["0.4 %", "0", "+0.4 %"], fontsize=9)
axr.set_xlabel("cluster sum [ADU]", fontsize=9)
axr.set_ylabel("vs CPU", fontsize=9)
bare(axr, keep=("left", "bottom"))
axr.text(0.985, 0.90, f"worst populated bin: {dev*100:.3f} % · band = ±0.1 %",
transform=axr.transAxes, ha="right", va="top", color=GREEN,
fontsize=7.5)
fontsize=9)
fig.subplots_adjust(hspace=0.10)
save(fig, "fig_spectra_valid")
@@ -253,7 +347,7 @@ def fig_arc():
ax.plot([x0 + 0.08, x1 - 0.08], [TOP * 0.955] * 2, color=col, lw=2.2,
zorder=2)
ax.text((x0 + x1) / 2, TOP * 0.965, label, ha="center", va="bottom",
color=col, fontsize=8.5, fontweight="bold")
color=col, fontsize=9, fontweight="bold")
ax.bar(x, fps, width=0.62, color=colors, zorder=3, linewidth=0)
@@ -263,7 +357,7 @@ def fig_arc():
ax.axhspan(61312, 61859, color=GREEN, alpha=0.20, zorder=1)
ax.axhline(61859, color=GREEN, lw=1.2, ls="--", zorder=4)
ax.text(-0.42, 63200, "H2D floor · 6162 k FPS · the GPU cannot be fed faster",
color=GREEN, fontsize=8.5, ha="left", va="bottom")
color=GREEN, fontsize=9, ha="left", va="bottom")
for xi, (f, s) in enumerate(zip(fps, spd)):
ax.text(xi, f + 1100, f"{f:,}", ha="center", va="bottom",
@@ -272,7 +366,7 @@ def fig_arc():
ha="center", va="top", color=BG, fontsize=9, fontweight="bold")
ax.set_xticks(x)
ax.set_xticklabels(steps, fontsize=8.5, color=TEXT2)
ax.set_xticklabels(steps, fontsize=9, color=TEXT2)
ax.set_ylim(0, TOP)
ax.set_yticks([])
bare(ax, keep=("bottom",))
@@ -308,7 +402,7 @@ def fig_arc_9x9():
ax.plot([x0 + 0.08, x1 - 0.08], [TOP * 0.955] * 2, color=col, lw=2.2,
zorder=2)
ax.text((x0 + x1) / 2, TOP * 0.965, label, ha="center", va="bottom",
color=col, fontsize=8.5, fontweight="bold")
color=col, fontsize=9, fontweight="bold")
ax.bar(x, fps, width=0.58, color=colors, zorder=3, linewidth=0)
@@ -324,10 +418,10 @@ def fig_arc_9x9():
# completely that it handed the constraint to the result path.
ax.hlines(33323, -0.5, 4.5, color=GREEN, lw=1.3, ls="--", zorder=4)
ax.text(-0.42, 33900, "f64 KERNEL floor · 33 323 FPS (nsys estimates 30 621)",
color=GREEN, fontsize=8.5, va="bottom")
color=GREEN, fontsize=9, va="bottom")
ax.hlines(39775, 4.5, 5.5, color=AMBER, lw=1.3, ls="--", zorder=4)
ax.text(5.46, 45600, "f32 D2H floor · 39 775 FPS\n(nsys estimates 39 614)",
color=AMBER, fontsize=8.5, ha="right", va="bottom", linespacing=1.35)
color=AMBER, fontsize=9, ha="right", va="bottom", linespacing=1.35)
ax.add_patch(FancyArrowPatch((4.62, 34100), (4.62, 39100), arrowstyle="-|>",
mutation_scale=11, color=AMBER, lw=1.5, zorder=5))
ax.text(3.30, 37600, "40 % kernel → D2H binds instead", color=AMBER,
@@ -340,7 +434,7 @@ def fig_arc_9x9():
ha="center", va="top", color=BG, fontsize=9, fontweight="bold")
ax.set_xticks(x)
ax.set_xticklabels(steps, fontsize=8.5, color=TEXT2)
ax.set_xticklabels(steps, fontsize=9, color=TEXT2)
ax.set_ylim(0, TOP)
ax.set_yticks([])
bare(ax, keep=("bottom",))
@@ -383,7 +477,7 @@ def fig_overhead():
fontsize=9.5, fontweight="bold")
if e > ymax * 0.05:
ax.text(xi, f + e / 2, f"+{e:.0f}", ha="center", va="center",
color=BG, fontsize=8.5, fontweight="bold")
color=BG, fontsize=9, fontweight="bold")
ax.set_xticks(x)
ax.set_xticklabels(steps, color=TEXT2, fontsize=9)
ax.set_ylim(0, ymax)
@@ -395,9 +489,9 @@ def fig_overhead():
# label the two segments in place — the floor takes the act colour, so a
# colour-keyed legend would be wrong
axes[0].text(0, 16.17 / 2, "GPU\nfloor", ha="center", va="center", color=BG,
fontsize=8, fontweight="bold")
fontsize=9, fontweight="bold")
axes[0].annotate("host excess", xy=(0.30, 25), xytext=(1.15, 41),
color=TEXT2, fontsize=8.5,
color=TEXT2, fontsize=9,
arrowprops=dict(arrowstyle="-", color=MUTED, lw=0.9))
axes[1].text(1.5, 108, "Act II removes the host bar", color=PALE, fontsize=9,
@@ -481,7 +575,7 @@ def fig_streams():
for i in range(3):
_frame_bars(ax, 1.0, i * FR)
ax.set_ylim(0.4, 2.3)
ax.text(3 * FR + 6, 1.34, "one engine at a time", color=MUTED, fontsize=7.5,
ax.text(3 * FR + 6, 1.34, "one engine at a time", color=MUTED, fontsize=9,
va="center")
# --- opt2: 4 streams, barrier after each round. The drain is not drawn in;
@@ -493,7 +587,7 @@ def fig_streams():
for h_idle, t_end in drains:
ax.axvspan(h_idle, t_end, color=AMBER, alpha=0.13, zorder=1)
ax.text((drains[0][0] + drains[0][1]) / 2, 4.15, "barrier — H2D starves",
color=AMBER, fontsize=7.5, ha="center", va="bottom")
color=AMBER, fontsize=9, ha="center", va="bottom")
ax.set_ylim(-0.4, 4.9)
# --- opt3: no barriers, continuous
@@ -502,7 +596,7 @@ def fig_streams():
_draw_schedule(ax, frames)
ax.set_ylim(-1.5, 4.5)
ax.text(0, -0.25, "streams never wait on each other — the H2D engine never goes idle",
color=ACCENT, fontsize=7.5, va="top")
color=ACCENT, fontsize=9, va="top")
titles = ["opt1 · 1 stream, synchronous",
"opt2 · 4 streams, sync barrier per round",
"opt3 · 4 streams, barriers removed"]
@@ -514,14 +608,10 @@ def fig_streams():
handles = [Rectangle((0, 0), 1, 1, color=c) for c in (AMBER, ACCENT, PALE)]
axes[0].legend(handles, ["H2D copy", "kernel", "D2H copy"], frameon=False,
fontsize=8, labelcolor=TEXT2, ncol=3, loc="lower right",
fontsize=9, labelcolor=TEXT2, ncol=3, loc="lower right",
bbox_to_anchor=(1.02, 0.98), handlelength=1.1)
axes[2].set_xlabel("time →", color=MUTED, fontsize=8.5, loc="left")
axes[2].set_xlabel("time →", color=MUTED, fontsize=9, loc="left")
fig.subplots_adjust(hspace=0.75, bottom=0.10)
fig.text(0.5, 0.005,
"Scheduled, not sketched: H2D and D2H are one FIFO engine each, as in hardware — measured overlap 1.00 on both — while "
"kernels may\noverlap. 3×3 proportions [f64 · s1], 13 : 15 : 5. The H2D lane fills first because at 3×3 H2D is the tallest bar.",
color=MUTED, fontsize=6.4, ha="center", va="bottom", linespacing=1.6)
save(fig, "fig_streams")
@@ -534,68 +624,71 @@ def fig_pinning():
def box(x, y, w, h, label, sub=""):
ax.add_patch(Rectangle((x, y), w, h, facecolor=PANEL, edgecolor=RULE, lw=1))
ax.text(x + w / 2, y + h / 2 + 0.26, label, ha="center", va="center",
color=PALE, fontsize=8.5, fontweight="bold")
color=PALE, fontsize=9.5, fontweight="bold")
ax.text(x + w / 2, y + h / 2 - 0.34, sub, ha="center", va="center",
color=MUTED, fontsize=7)
color=MUTED, fontsize=9.5)
def arrow(x0, x1, y, color, label):
ax.add_patch(FancyArrowPatch((x0, y), (x1, y), arrowstyle="-|>",
mutation_scale=10, color=color, lw=1.6))
ax.text((x0 + x1) / 2, y + 0.22, label, ha="center", va="bottom",
color=color, fontsize=7)
color=color, fontsize=9.5)
ax.text(0, 5.95, "PAGEABLE · before opt4", color=AMBER, fontsize=8.5,
ax.text(0, 5.95, "PAGEABLE · before opt4", color=AMBER, fontsize=9.5,
fontweight="bold")
box(0, 4.05, 2.5, 1.1, "numpy array", "pageable")
box(4.0, 4.05, 2.4, 1.1, "driver staging", "hidden pinned buf")
box(7.9, 4.05, 2.5, 1.1, "GPU", "device memory")
arrow(2.5, 4.0, 4.60, AMBER, "memcpy")
arrow(6.4, 7.9, 4.60, AMBER, "DMA")
ax.text(0, 3.62, "every transfer is copied twice", color=MUTED, fontsize=7)
ax.text(0, 3.62, "every transfer is copied twice", color=MUTED, fontsize=9.5)
ax.text(0, 2.75, "PINNED · opt4", color=ACCENT, fontsize=8.5, fontweight="bold")
ax.text(0, 2.75, "PINNED · opt4", color=ACCENT, fontsize=9.5, fontweight="bold")
box(0, 0.85, 2.5, 1.1, "numpy array", "page-locked")
box(7.9, 0.85, 2.5, 1.1, "GPU", "device memory")
arrow(2.5, 7.9, 1.40, ACCENT, "DMA — engine reads host RAM directly")
arrow(2.5, 7.9, 1.40, ACCENT, "DMA · reads host RAM directly")
ax.text(0, 0.42, "no staging copy, no page faults, fully async",
color=MUTED, fontsize=7)
color=MUTED, fontsize=9.5)
# the rule, in one inset: pinning pays only where H2D is the tallest bar
ax2 = fig.add_axes([0.70, 0.16, 0.30, 0.62])
x = np.arange(2)
before = [34.26, 82.17]
after = [25.98, 80.44]
# 3x3 from 2026-08-18_f64, 9x9 from 2026-08-20_f64_cap1700 -- NOT the
# cap-1500 ladder, which read 82.17 -> 80.44 and put a stale x1.02 on the
# slide for a step that actually buys x1.03.
before = [34.26, 82.44]
after = [25.98, 79.83]
ax2.bar(x - 0.19, before, width=0.36, color=AMBER, zorder=3, label="opt3")
ax2.bar(x + 0.19, after, width=0.36, color=ACCENT, zorder=3, label="opt4")
for xi, (b, a) in enumerate(zip(before, after)):
ax2.text(xi - 0.19, b + 2, f"{b:.0f}", ha="center", color=TEXT2, fontsize=7.5)
ax2.text(xi + 0.19, a + 2, f"{a:.0f}", ha="center", color=TEXT2, fontsize=7.5)
ax2.text(xi, 92, f"×{b / a:.2f}", ha="center", color=PALE, fontsize=9,
ax2.text(xi - 0.19, b + 2, f"{b:.0f}", ha="center", color=TEXT2, fontsize=9.5)
ax2.text(xi + 0.19, a + 2, f"{a:.0f}", ha="center", color=TEXT2, fontsize=9.5)
ax2.text(xi, 92, f"×{b / a:.2f}", ha="center", color=PALE, fontsize=9.5,
fontweight="bold")
ax2.set_xticks(x)
ax2.set_xticklabels(["3×3\nH2D-bound", "9×9\nkernel-bound"], color=TEXT2,
fontsize=7.5)
fontsize=9.5)
ax2.set_ylim(0, 104); ax2.set_yticks([]); bare(ax2, keep=("bottom",))
ax2.set_title("µs / frame", color=MUTED, fontsize=7.5, pad=6)
ax2.legend(frameon=False, fontsize=7, labelcolor=TEXT2, loc="center left")
ax2.set_title("µs / frame", color=MUTED, fontsize=9.5, pad=6)
ax2.legend(frameon=False, fontsize=9.5, labelcolor=TEXT2, loc="center left")
save(fig, "fig_pinning")
# ----------------------------------------------- 6. graphs (rejected route)
def fig_graphs():
fig = plt.figure(figsize=(7.7, 3.0))
ax = fig.add_axes([0, 0.30, 0.66, 0.70])
fig = plt.figure(figsize=(7.6, 2.55))
ax = fig.add_axes([0, 0.07, 0.66, 0.90])
ax.axis("off"); ax.set_xlim(0, 11.2); ax.set_ylim(0, 4.6)
def node(x, y, w, h, t, fc):
ax.add_patch(Rectangle((x, y), w, h, facecolor=fc, edgecolor="none"))
ax.text(x + w / 2, y + h / 2, t, ha="center", va="center",
color=BG, fontsize=7.5, fontweight="bold")
color=BG, fontsize=9.5, fontweight="bold")
ops = [("H2D", AMBER), ("kernel", ACCENT), ("D2H", PALE)] * 2
ax.text(0, 4.15, "WITHOUT GRAPHS · one driver call per operation, every frame",
color=AMBER, fontsize=8.5, fontweight="bold")
color=AMBER, fontsize=9.5, fontweight="bold")
for i, (t, c) in enumerate(ops):
x = 0.1 + i * 1.62
node(x, 2.85, 1.4, 0.6, t, c)
@@ -603,10 +696,10 @@ def fig_graphs():
arrowstyle="-|>", mutation_scale=7,
color=MUTED, lw=0.9))
ax.text(11.1, 3.15, "CPU cost\n≈ 6 launches", ha="right", va="center",
color=MUTED, fontsize=7.5)
color=MUTED, fontsize=9.5)
ax.text(0, 2.18, "WITH GRAPHS · record once, replay with one launch",
color=ACCENT, fontsize=8.5, fontweight="bold")
color=ACCENT, fontsize=9.5, fontweight="bold")
ax.add_patch(Rectangle((0.1, 0.72), 9.20, 1.15, facecolor=PANEL,
edgecolor=ACCENT, lw=1.2))
for i, (t, c) in enumerate(ops):
@@ -614,31 +707,26 @@ def fig_graphs():
ax.add_patch(FancyArrowPatch((0.8, 2.02), (0.8, 1.90), arrowstyle="-|>",
mutation_scale=8, color=ACCENT, lw=1.3))
ax.text(11.1, 1.30, "CPU cost\n≈ 1 launch", ha="right", va="center",
color=ACCENT, fontsize=7.5, fontweight="bold")
color=ACCENT, fontsize=9.5, fontweight="bold")
# the verdict
ax2 = fig.add_axes([0.72, 0.30, 0.28, 0.62])
ax2 = fig.add_axes([0.73, 0.13, 0.27, 0.78])
x = np.arange(2)
opt4 = [25.98, 80.44]
graph = [25.16, 94.78]
opt4 = [25.98, 79.83] # 3x3 2026-08-18_f64 · 9x9 2026-08-20_f64_cap1700
graph = [25.16, 90.32] # cap 1500 read 80.44 / 94.78 and said "18 % slower"
ax2.bar(x - 0.19, opt4, width=0.36, color=ACCENT, zorder=3, label="opt4")
ax2.bar(x + 0.19, graph, width=0.36, color=AMBER, zorder=3, label="graphs")
for xi, (a, g) in enumerate(zip(opt4, graph)):
ax2.text(xi - 0.19, a + 2.5, f"{a:.0f}", ha="center", color=TEXT2, fontsize=7.5)
ax2.text(xi + 0.19, g + 2.5, f"{g:.0f}", ha="center", color=TEXT2, fontsize=7.5)
ax2.text(1, 108, "18 % SLOWER", ha="center", color=AMBER, fontsize=8,
ax2.text(xi - 0.19, a + 2.5, f"{a:.0f}", ha="center", color=TEXT2, fontsize=9.5)
ax2.text(xi + 0.19, g + 2.5, f"{g:.0f}", ha="center", color=TEXT2, fontsize=9.5)
ax2.text(1, 100, "12 % THROUGHPUT", ha="center", color=AMBER, fontsize=9.5,
fontweight="bold")
ax2.set_xticks(x)
ax2.set_xticklabels(["3×3", "9×9"], color=TEXT2, fontsize=8)
ax2.set_xticklabels(["3×3", "9×9"], color=TEXT2, fontsize=9.5)
ax2.set_ylim(0, 120); ax2.set_yticks([]); bare(ax2, keep=("bottom",))
ax2.set_title("µs / frame", color=MUTED, fontsize=7.5, pad=6)
ax2.legend(frameon=False, fontsize=7, labelcolor=TEXT2, loc="upper left")
ax2.set_title("µs / frame", color=MUTED, fontsize=9.5, pad=6)
ax2.legend(frameon=False, fontsize=9.5, labelcolor=TEXT2, loc="upper left")
fig.text(0.02, 0.10,
"REJECTED — launch overhead stops binding one step later. The graph finder "
"never got the chunked pipeline of opt5,\nand its original advantage is "
"swamped by an overlap it does not have.",
color=AMBER, fontsize=8, fontweight="bold", va="top")
save(fig, "fig_graphs")
@@ -657,7 +745,7 @@ def fig_resultpath():
ax.bar([0], [floor], width=0.5, color=col, zorder=3, linewidth=0)
ax.bar([1], [copy_us], width=0.5, color=col, zorder=3, linewidth=0)
ax.axhline(floor, color=GREEN, lw=1.3, ls="--", zorder=4)
ax.text(-0.55, floor + 1.2, "GPU floor", color=GREEN, fontsize=8.5, ha="left")
ax.text(-0.55, floor + 1.2, "GPU floor", color=GREEN, fontsize=10, ha="left")
ax.text(0, floor + 1.4, f"{floor:.1f} µs", ha="center", color=PALE,
fontsize=10, fontweight="bold")
ax.text(1, copy_us + 1.4, f"{copy_us:.0f} µs", ha="center", color=PALE,
@@ -665,13 +753,13 @@ def fig_resultpath():
ax.set_xticks([0, 1])
ax.set_xticklabels(["GPU per frame\n(H2D ∥ kernel ∥ D2H)",
"host copy per frame\ncollect() memcpy + malloc"],
color=TEXT2, fontsize=8.5)
color=TEXT2, fontsize=10)
ax.set_xlim(-0.6, 1.7); ax.set_ylim(0, 52); ax.set_yticks([])
bare(ax, keep=("bottom",))
ax.set_title(title, color=PALE, fontsize=10, pad=10, loc="left")
ax.text(1.68, 46, gain, color=col, fontsize=16, fontweight="bold", ha="right")
ax.text(1.68, 40, "opt5 → opt6", color=MUTED, fontsize=8, ha="right")
ax.text(-0.55, -11, verdict, color=col, fontsize=8.5, fontweight="bold",
ax.text(1.68, 40, "opt5 → opt6", color=MUTED, fontsize=10, ha="right")
ax.text(-0.55, -11, verdict, color=col, fontsize=10, fontweight="bold",
va="top")
fig.subplots_adjust(bottom=0.30)
save(fig, "fig_resultpath")
@@ -700,8 +788,8 @@ def fig_f32_kernel():
"f32 kernel drops BELOW D2H"),
]
for ax, title, f64, f32, delta, verdict in panels:
ax.barh(y + h / 2, f64, height=h, color=AMBER, zorder=3, label="f64 ped")
ax.barh(y - h / 2, f32, height=h, color=ACCENT, zorder=3, label="f32 ped")
ax.barh(y + h / 2, f64, height=h, color=ACCENT, zorder=3, label="f64 ped")
ax.barh(y - h / 2, f32, height=h, color=AMBER, zorder=3, label="f32 ped")
# the engine that sets the floor in each arm is named, not ringed
i64, i32 = int(np.argmax(f64)), int(np.argmax(f32))
for yi, (a, b) in enumerate(zip(f64, f32)):
@@ -710,87 +798,154 @@ def fig_f32_kernel():
# in the arm where it no longer binds
ax.text(val + 0.7, yi + off, f"{val:.2f}" + (" ◀ binds" if top else ""),
va="center", color=PALE if (top or yi == 0) else TEXT2,
fontsize=7.5, fontweight="bold" if top else "normal")
ax.text(0.985, 0.055, verdict, transform=ax.transAxes, ha="right",
color=PALE, fontsize=8, fontweight="bold")
ax.text(f64[0] * 0.42, -0.02 - h, delta, ha="center", va="center",
color=PALE, fontsize=9.5, fontweight="bold")
ax.set_yticks(y); ax.set_yticklabels(labels, color=TEXT2, fontsize=8.5)
ax.invert_yaxis(); ax.set_xlim(0, 46); ax.set_xticks([])
ax.set_ylim(2.62, -0.72)
fontsize=10, fontweight="bold" if top else "normal")
# A bar pair spans y +- h, so the delta and the verdict need a full h of
# clearance beyond that or they sit on the bar they are labelling. The
# y limits carry that margin explicitly rather than relying on the
# default padding, which the larger type had eaten.
ax.text(f64[0] * 0.42, -h - 0.30, delta, ha="center", va="bottom",
color=PALE, fontsize=10.5, fontweight="bold")
ax.text(0.985, 2 + h + 0.34, verdict, transform=ax.get_yaxis_transform(),
ha="right", va="top", color=PALE, fontsize=10, fontweight="bold")
ax.set_yticks(y); ax.set_yticklabels(labels, color=TEXT2, fontsize=10)
ax.invert_yaxis(); ax.set_xlim(0, 60); ax.set_xticks([])
ax.set_ylim(3.05, -0.98)
bare(ax, keep=("left",))
ax.set_title(title, color=PALE, fontsize=9, pad=10, loc="left")
ax.set_title(title, color=PALE, fontsize=10, pad=10, loc="left")
# the wall the f32 kernel has to clear, drawn only where it matters
ax2.axvline(25.24, color=MUTED, lw=1.0, ls="--", zorder=4)
ax1.legend(frameon=False, fontsize=8, labelcolor=TEXT2, loc="lower right",
bbox_to_anchor=(1.0, 0.13))
fig.subplots_adjust(bottom=0.20, top=0.86, left=0.09, right=0.99, wspace=0.22)
fig.text(0.545, 0.015,
"s4 is busy time per frame — the UNION of an engine's intervals, not a duration. The copy engines "
"still serialize (overlap 1.00), so H2D and D2H\nstay durations and rise under contention; the kernel "
"overlaps itself 1.32× across streams, so its 43.2 µs mean duration reads as 32.66 µs of "
"occupancy. → A1",
ha="center", va="bottom", color=MUTED, fontsize=6.8, linespacing=1.6)
ax2.axvline(25.24, color=MUTED, lw=1.0, ls="--", zorder=1) # under the labels
# parked between the D2H and H2D rows, the one region no label reaches
ax1.legend(frameon=False, fontsize=10, labelcolor=TEXT2, loc="center right",
bbox_to_anchor=(1.0, 0.40))
fig.subplots_adjust(bottom=0.20, top=0.86, left=0.09, right=0.99, wspace=0.30)
save(fig, "fig_f32_kernel")
# --------------------------------------------------------- 9. cancellation
def fig_cancellation():
fig, (ax, ax2) = plt.subplots(1, 2, figsize=(7.7, 2.7),
gridspec_kw={"width_ratios": [1.25, 1]})
names = ["E[X²]\n2.17e7", "mean²\n2.17e7", "var, bulk\n2025", "var, quiet\n9"]
ax.bar([0, 1], [2.17e7, 2.17e7], width=0.5, color=[PALE, PALE], zorder=3)
ax.bar([2], [2025], width=0.5, color=AMBER, zorder=3)
ax.bar([3], [9], width=0.5, color=AMBER, zorder=3)
ax.set_yscale("log"); ax.set_ylim(1, 2e8)
ax.set_xticks([0, 1, 2, 3]); ax.set_xticklabels(names, color=TEXT2, fontsize=7.5)
ax.set_yticks([1e0, 1e2, 1e4, 1e6, 1e8])
ax.axhline(3, color=ACCENT, lw=1.3, ls="--", zorder=4)
ax.text(3.46, 2.2e5, "the error: ±3 ADU², the same\nfor every pixel — so it is the\n"
"quiet ones it swallows", color=ACCENT, fontsize=7.5, ha="right",
va="center", linespacing=1.5)
ax.plot([3.0, 3.0], [3.6, 4e4], color=ACCENT, lw=0.7, ls=":", zorder=4)
bare(ax)
ax.set_title("var = E[X²] mean² · each operand on a grid of 2 ADU²",
color=MUTED, fontsize=8, pad=8)
"""The trap and what it did to the physics, on one pair of axes.
# The axes are scaled to the region the error actually reaches: the floor is
# +-3-4 ADU^2 (docs/pedestal_precision_f32_cancellation.md SS5), so variance is
# lost below rms ~2 and merely corrupted up to rms ~5. An earlier version drew
# the floor at 42 with the damage running to rms 6.5 -- that is 6.5^2, i.e. the
# line and the shading were each other's source, and both were ~14x too large.
LEFT is the cancellation itself: two ~2.17e7 operands, an answer of ~2025 for
a normal pixel and ~9 for a quiet one, against a fixed +-3 ADU^2 f32 error
that does not shrink with the answer.
RIGHT is the consequence. The f64/CPU curve is MEASURED -- validation_tiers.json,
23.2 M clusters over 10 000 frames. The f32 curve is RECONSTRUCTED, not
measured: the excess is placed where SS9 of
docs/pedestal_precision_f32_cancellation.md derives it (a threshold that has
collapsed from ~225 ADU to ~0 admits the whole positive side of the pixel's
distribution, so the population is smeared upward from zero) and its AREA is
set to the measured +28.06 % from SS1. Shape schematic, area measured, and the
caption on the slide says so. Drawing a measured f32 spectrum would be better;
that build is not one anyone should keep around to re-run.
The variance-vs-rms error floor that used to sit in this panel is now
fig_varfloor, in the annex: it is the quantitative version of the same claim
and it was competing with the physics for the audience's attention.
"""
import json
fig, (ax, ax2) = plt.subplots(1, 2, figsize=(7.9, 2.34),
gridspec_kw={"width_ratios": [1.0, 1.28]})
names = ["E[X²]\n2.17e7", "mean²\n2.17e7", "var\nbulk 2025", "var\nquiet 9"]
ax.bar([0, 1], [2.17e7, 2.17e7], width=0.52, color=[PALE, PALE], zorder=3)
ax.bar([2], [2025], width=0.52, color=AMBER, zorder=3)
ax.bar([3], [9], width=0.52, color=AMBER, zorder=3)
ax.set_yscale("log"); ax.set_ylim(1, 8e8)
ax.set_xticks([0, 1, 2, 3])
ax.set_xticklabels(names, color=TEXT2, fontsize=9, linespacing=1.35)
ax.set_yticks([1e0, 1e2, 1e4, 1e6, 1e8])
ax.tick_params(labelsize=9)
ax.axhline(3, color=ACCENT, lw=1.4, ls="--", zorder=4)
ax.text(3.52, 1.1e7, "±3 ADU², the same for\nevery pixel — so it is the\n"
"quiet ones it swallows", color=ACCENT, fontsize=9, ha="right",
va="center", linespacing=1.4)
ax.plot([3.0, 3.0], [3.6, 1.2e6], color=ACCENT, lw=0.7, ls=":", zorder=4)
bare(ax)
ax.set_title("var = E[X²] mean²", color=MUTED, fontsize=9.5, pad=8, loc="left")
d = json.loads((Path(__file__).resolve().parent
/ "validation_tiers.json").read_text())
e = np.array(d["edges"])
ctr = 0.5 * (e[1:] + e[:-1])
bw = e[1] - e[0]
good = np.array(d["hists"]["cpu"], dtype=float)
# +28.06 % more clusters (SS1), smeared upward from ~0 by a gate that has
# collapsed to 0 sigma. Exponential with a 400 ADU scale: shape schematic,
# integral set to the measured excess.
LAM = 400.0
excess = good.sum() * 0.2806 * (bw / LAM) * np.exp(-np.clip(ctr, 0, None) / LAM)
ax2.step(ctr, good, where="mid", color=ACCENT, lw=1.8, zorder=4,
label="f64 pedestal · measured")
ax2.step(ctr, good + excess, where="mid", color=AMBER, lw=1.5, zorder=3,
label="f32 pedestal · +28.06 %")
ax2.fill_between(ctr, good, good + excess, step="mid", color=AMBER,
alpha=0.22, lw=0, zorder=2)
ax2.set_yscale("log")
ax2.set_ylim(2e2, 6e7)
ax2.set_xlim(0, 2600)
ax2.set_xticks([0, 1000, 2000])
ax2.tick_params(labelsize=9, colors=MUTED)
ax2.set_xlabel("cluster energy [ADU]", color=TEXT2, fontsize=9.5)
bare(ax2, keep=("left", "bottom"))
ax2.legend(frameon=False, fontsize=9, labelcolor=TEXT2, loc="lower right",
handlelength=1.2, borderaxespad=0.3)
ax2.annotate("clusters below the 5σ cut\nthat should not be there",
xy=(150, 1.1e6), xytext=(560, 8.0e6), color=AMBER, fontsize=9,
linespacing=1.4, zorder=6,
arrowprops=dict(arrowstyle="->", color=AMBER, lw=0.9))
ax2.text(1195, 3.4e6, "Cu Kα", color=PALE, fontsize=9, ha="center")
ax2.set_title("cluster-energy spectrum, 3×3", color=MUTED, fontsize=9.5,
pad=8, loc="left")
fig.subplots_adjust(left=0.075, right=0.995, top=0.86, bottom=0.20, wspace=0.24)
save(fig, "fig_cancellation")
def fig_varfloor():
"""The quantitative version of the left panel: which pixels the error reaches.
Floor is +-3-4 ADU^2 (docs/pedestal_precision_f32_cancellation.md SS5), so
variance is LOST below rms ~2 and merely corrupted up to rms ~5. An earlier
version drew the floor at 42 with the damage running to rms 6.5 -- that is
6.5^2, i.e. the line and the shading were each other's source, and both were
~14x too large.
"""
fig, ax2 = plt.subplots(figsize=(4.35, 3.05))
rms = np.linspace(0, 6, 300)
ax2.axvspan(0, 2.0, color=AMBER, alpha=0.20, zorder=1)
ax2.axvspan(2.0, 5.0, color=AMBER, alpha=0.07, zorder=1)
ax2.fill_between([0, 6], 3, 4, color=ACCENT, alpha=0.22, zorder=2, lw=0)
ax2.plot(rms, rms**2, color=PALE, lw=1.8, zorder=4)
ax2.axhline(3, color=ACCENT, lw=1.4, ls="--", zorder=5)
# the three rows of the doc's per-pixel table, labelled left of the curve
# (it is convex, so the whole upper-left of the panel is free)
pts = ((2.0, "rms 2 → var 4 ± 3 → 0", 0.30, 11.2),
(3.0, "rms 3 → 17 % rms err", 0.30, 17.4),
(5.0, "rms 5 → 6 % rms err", 2.30, 24.6))
for r, lab, tx, ty in pts:
ax2.plot([r], [r * r], marker="o", ms=4, color=PALE, zorder=6)
ax2.annotate(lab, xy=(r, r * r), xytext=(tx, ty), color=TEXT2,
fontsize=6.8, va="center", zorder=6,
fontsize=9.5, va="center", zorder=6,
arrowprops=dict(arrowstyle="-", color=MUTED, lw=0.6,
shrinkA=2, shrinkB=3))
ax2.set_xlabel("pixel rms (ADU)", color=TEXT2, fontsize=8.5)
ax2.set_ylabel("variance (ADU²)", color=TEXT2, fontsize=8.5)
ax2.set_xlabel("pixel rms (ADU)", color=TEXT2, fontsize=9.5)
ax2.set_ylabel("variance (ADU²)", color=TEXT2, fontsize=9.5)
ax2.set_ylim(0, 40); ax2.set_xlim(0, 6)
ax2.set_yticks([0, 10, 20, 30, 40]); ax2.set_xticks([0, 2, 4, 6])
ax2.tick_params(labelsize=7.5, colors=MUTED)
ax2.tick_params(labelsize=9.5, colors=MUTED)
bare(ax2, keep=("left", "bottom"))
ax2.text(0.12, 38.8, "rms < 2\n→ clamped to 0\n→ fires every frame",
color=AMBER, fontsize=7.0, va="top", linespacing=1.5, zorder=6)
color=AMBER, fontsize=9.5, va="top", linespacing=1.45, zorder=6)
ax2.text(2.20, 38.8, "rms 25: threshold\ncorrupted, not clamped",
color=MUTED, fontsize=6.8, va="top", linespacing=1.5, zorder=6)
ax2.text(3.05, 31.5, "true variance = rms²", color=PALE, fontsize=7.2, zorder=6)
ax2.text(5.90, 5.3, "f32 error floor ±34 ADU²", color=ACCENT, fontsize=7.0,
color=MUTED, fontsize=9.5, va="top", linespacing=1.45, zorder=6)
ax2.text(3.05, 31.0, "true variance = rms²", color=PALE, fontsize=9.5, zorder=6)
ax2.text(5.90, 5.6, "f32 error floor ±34 ADU²", color=ACCENT, fontsize=9.5,
ha="right", zorder=6)
save(fig, "fig_cancellation")
fig.subplots_adjust(left=0.13, right=0.98, top=0.97, bottom=0.15)
save(fig, "fig_varfloor")
fig_varfloor()
# ------------------------------- 10. the same kernel, measured five ways
@@ -896,7 +1051,7 @@ def fig_overlap():
ax.add_patch(Rectangle((x, y), w, lane_h, facecolor=col, edgecolor=BG,
linewidth=1.4, zorder=3))
ax.text(x + w / 2, y + lane_h / 2, txt, ha="center", va="center",
color=tcol, fontsize=8.5, fontweight="bold", zorder=4)
color=tcol, fontsize=10.5, fontweight="bold", zorder=4)
# ---- serial: G H G H G H G H, strictly alternating ----------------------
yG, yH = 3.30, 2.78
@@ -918,13 +1073,13 @@ def fig_overlap():
for y, lbl in ((yG, "GPU"), (yH, "host"), (yG2, "GPU"), (yH2, "host")):
ax.text(-0.25, y + lane_h / 2, lbl, ha="right", va="center",
color=TEXT2, fontsize=9)
color=TEXT2, fontsize=10.5)
ax.text(-0.25, yG + lane_h + 0.30, "submit → collect, serialized (opt4)",
ha="left", va="bottom", color=TEXT2, fontsize=9.5, fontweight="bold")
ha="left", va="bottom", color=TEXT2, fontsize=10.5, fontweight="bold")
ax.text(-0.25, yG2 + lane_h + 0.30,
"submit(i+1) before collect(i) (opt5)",
ha="left", va="bottom", color=AMBER, fontsize=9.5, fontweight="bold")
ha="left", va="bottom", color=AMBER, fontsize=10.5, fontweight="bold")
for x, y0, y1, col in ((serial_end, yH, yG + lane_h, MUTED),
(pipe_end, yH2, yG2 + lane_h, AMBER)):
@@ -935,12 +1090,12 @@ def fig_overlap():
arrowprops=dict(arrowstyle="<|-|>", color=AMBER, lw=1.5))
ax.text((pipe_end + serial_end) / 2, 0.20,
"saved: min(GPU, host) per chunk", ha="center", va="top",
color=AMBER, fontsize=9.5, fontweight="bold")
color=AMBER, fontsize=10.5, fontweight="bold")
ax.text(serial_end + 0.25, yG + lane_h / 2, "GPU + host per chunk",
ha="left", va="center", color=MUTED, fontsize=9)
ha="left", va="center", color=MUTED, fontsize=10.5)
ax.text(pipe_end + 0.25, yG2 + lane_h / 2, "max(GPU, host) per chunk",
ha="left", va="center", color=AMBER, fontsize=9, fontweight="bold")
ha="left", va="center", color=AMBER, fontsize=10.5, fontweight="bold")
ax.set_xlim(-1.6, serial_end + 4.0)
ax.set_ylim(0, 4.25)
@@ -1051,7 +1206,7 @@ def fig_mismatch147():
cmap = plt.cm.viridis.copy()
cmap.set_bad(PANEL)
fig, axes = plt.subplots(1, 2, figsize=(11.2, 4.3))
fig, axes = plt.subplots(1, 2, figsize=(9.0, 3.9))
for ax, name, col in ((axes[0], "frozen", PALE), (axes[1], "cuda", ACCENT)):
w, m = sub[name], msk[name]
ax.imshow(np.ma.masked_where(~m, w), cmap=cmap, vmin=0, vmax=vmax,
@@ -1060,10 +1215,10 @@ def fig_mismatch147():
for j in range(w.shape[1]):
if m[i, j]:
ax.text(j, i, f"{w[i, j]:.0f}", ha="center", va="center",
fontsize=4.6,
fontsize=6.4, zorder=7,
color="white" if w[i, j] < 0.55 * vmax else "black")
for (cx, cy) in cen[name]:
ax.plot(cx - X0, cy - Y0, ".", color="#FF4B4B", ms=6, zorder=5)
ax.plot(cx - X0, cy - Y0, ".", color="#FF4B4B", ms=4.5, zorder=4)
for (cx, cy) in only:
if name == "cuda":
ax.add_patch(plt.Circle((cx - X0, cy - Y0), 2.6, fill=False,
@@ -1079,7 +1234,7 @@ def fig_mismatch147():
axes[0].text(-0.03, 0.5, f"frame {d['fid']}\nzoom on (202, 8)",
transform=axes[0].transAxes, rotation=90, ha="right",
va="center", color=MUTED, fontsize=8, linespacing=1.4)
va="center", color=MUTED, fontsize=9, linespacing=1.4)
# anchored in DATA coordinates: the window is no longer square (the cut is
# clipped by the top edge of the frame), so axes fractions do not track the
# pixel once matplotlib letterboxes the image to keep aspect.
@@ -1087,7 +1242,7 @@ def fig_mismatch147():
axes[1].annotate("cuda keeps this one;\nfrozen does not",
xy=(ox - X0 + 2.9, oy - Y0), xytext=(0.99, 0.93),
xycoords="data", textcoords="axes fraction",
color=AMBER, fontsize=8.5, fontweight="bold", ha="right",
color=AMBER, fontsize=9, fontweight="bold", ha="right",
linespacing=1.35,
arrowprops=dict(arrowstyle="-|>", color=AMBER, lw=1.3))
fig.subplots_adjust(wspace=0.06)
@@ -1111,7 +1266,7 @@ def fig_regpressure():
("3×3", 38, 6, ACCENT),
("9×9", 128, 2, AMBER),
]
fig, ax = plt.subplots(figsize=(11.2, 2.55))
fig, ax = plt.subplots(figsize=(7.4, 2.15))
y, ticks, labels = 0.0, [], []
for name, regs, blocks, col in rows:
@@ -1123,7 +1278,7 @@ def fig_regpressure():
ax.barh(y, seg - 0.5, left=b * seg + 0.25, height=0.52,
color=col, zorder=3, linewidth=0)
full = abs(frac - 100.0) < 0.6
ax.text(frac + 1.5, y, f"{frac:.0f} %", va="center", fontsize=9,
ax.text(frac + 1.5, y, f"{frac:.0f} %", va="center", fontsize=10,
color=PALE if full else TEXT2,
fontweight="bold" if full else "normal")
ticks.append(y); labels.append(f"{name} · {kind}")
@@ -1131,19 +1286,19 @@ def fig_regpressure():
y -= 0.45
ax.axvline(100, color=MUTED, lw=1.1, ls="--", zorder=4)
ax.text(99, 1.05, "capacity of one SM", color=MUTED, fontsize=8.5,
ax.text(99, 1.05, "capacity of one SM", color=MUTED, fontsize=10,
ha="right")
ax.set_yticks(ticks)
ax.set_yticklabels(labels, color=TEXT2, fontsize=9)
ax.set_yticklabels(labels, color=TEXT2, fontsize=10)
ax.set_xlim(0, 152); ax.set_xticks([])
ax.set_ylim(y + 0.6, 1.5)
bare(ax, keep=("left",))
ax.tick_params(axis="y", length=0)
ax.text(115, ticks[0] - 0.5, "6 blocks resident\n38 × 256 = 9 728 regs each",
color=ACCENT, fontsize=8.5, va="center", linespacing=1.4)
color=ACCENT, fontsize=10, va="center", linespacing=1.4)
ax.text(115, ticks[2] - 0.5, "2 blocks resident\n128 × 256 = 32 768 regs each",
color=AMBER, fontsize=8.5, va="center", linespacing=1.4)
color=AMBER, fontsize=10, va="center", linespacing=1.4)
fig.subplots_adjust(left=0.16, right=0.99, top=0.88, bottom=0.06)
save(fig, "fig_regpressure")
@@ -1159,7 +1314,7 @@ fig_regpressure()
def _engine_legend(ax, y=0.92):
handles = [Rectangle((0, 0), 1, 1, color=c) for c in (AMBER, ACCENT, PALE)]
ax.legend(handles, ["H2D copy", "kernel", "D2H copy"], frameon=False,
fontsize=8, labelcolor=TEXT2, ncol=3, loc="lower right",
fontsize=9.5, labelcolor=TEXT2, ncol=3, loc="lower right",
bbox_to_anchor=(1.02, y), handlelength=1.1)
@@ -1181,9 +1336,9 @@ def fig_opt1_timeline():
ax.axvspan(t0 + WORK, t0 + PER, color=MUTED, alpha=0.09, zorder=1)
ax.annotate("host blocks — every engine idle",
xy=(WORK + HOST / 2, 0.98), xytext=(WORK + HOST / 2, 0.56),
color=AMBER, fontsize=8, ha="center", va="top",
color=AMBER, fontsize=9, ha="center", va="top",
arrowprops=dict(arrowstyle="-", color=AMBER, lw=0.8))
ax.text(3 * PER + 4, 1.34, "one engine\nat a time", color=TEXT2, fontsize=8,
ax.text(3 * PER + 4, 1.34, "one engine\nat a time", color=TEXT2, fontsize=9,
va="center", linespacing=1.5)
ax.set_xlim(-4, 3 * PER + 40)
ax.set_ylim(0.02, 2.45)
@@ -1207,7 +1362,7 @@ def fig_opt2_timeline():
_draw_schedule(ax, frames)
for st in range(4):
ax.text(-4, (3 - st) + LANE_ / 2, f"stream {st}", color=MUTED,
fontsize=7.5, ha="right", va="center")
fontsize=9, ha="right", va="center")
# The window where the most frames are simultaneously in flight -- measured
# off the schedule, not asserted. At 3x3 proportions it is THREE, not four:
# the frame span (33) is only 2.5x the H2D stagger (13), so stream 0 has
@@ -1222,20 +1377,16 @@ def fig_opt2_timeline():
hi = max(b for _, b, c in counts if c == best)
ax.axvspan(lo, hi, color=ACCENT, alpha=0.10, zorder=1)
ax.text((lo + hi) / 2, 4.05, f"{best} frames in flight", color=ACCENT,
fontsize=7.5, ha="center", va="bottom")
fontsize=9, ha="center", va="bottom")
ax.set_xlim(-26, 100)
ax.set_ylim(-0.35, 4.6)
ax.set_yticks([]); ax.set_xticks([])
bare(ax, keep=())
ax.text(74, 0.34, "a copy in one stream runs\nwhile another computes",
color=TEXT2, fontsize=8, va="center", linespacing=1.5)
color=TEXT2, fontsize=9, va="center", linespacing=1.5)
_engine_legend(ax, y=0.94)
ax.set_xlabel("time →", color=MUTED, fontsize=8.5, loc="left")
ax.set_xlabel("time →", color=MUTED, fontsize=9, loc="left")
fig.subplots_adjust(left=0.01, right=0.99, top=0.90, bottom=0.19)
fig.text(0.01, 0.015,
"Scheduled, not sketched: the H2D bars queue because there is one copy engine per direction (measured overlap 1.00). 3×3 proportions "
"[f64 · s1], 13 : 15 : 5 — so a frame's span is 2.5× the H2D stagger and three are in flight, not four.",
color=MUTED, fontsize=6.4, va="bottom")
save(fig, "fig_opt2_timeline")
@@ -1293,7 +1444,7 @@ def fig_f32_absolute():
axA.set_xlim(-0.62, 3.02)
axA.set_ylim(0, 94)
axA.set_yticks([0, 25, 50, 75])
axA.set_yticklabels(["0", "25", "50", "75"], fontsize=8)
axA.set_yticklabels(["0", "25", "50", "75"], fontsize=8.5)
axA.set_ylabel("end-to-end µs / frame", color=MUTED, fontsize=8.5)
axA.set_xticks(x)
axA.set_xticklabels([f"{s}\n{t}" for s, t in zip(steps, sub)],
@@ -1316,7 +1467,7 @@ def fig_f32_absolute():
fontsize=10.5, fontweight="bold")
axB.set_ylim(0, 7.6)
axB.set_yticks([0, 2, 4])
axB.set_yticklabels(["0", "2", "4"], fontsize=8)
axB.set_yticklabels(["0", "2", "4"], fontsize=8.5)
axB.set_ylabel("µs saved by the f32 pedestal", color=MUTED, fontsize=8.5)
axB.set_xticks(x)
axB.set_xticklabels([f"{s}{'' if d else ''}\n{p:+.1f} %"
@@ -1386,3 +1537,110 @@ def fig_variance_rewrite():
fig_variance_rewrite()
# ---------------------------- 16. the measurement convention: s1, s4 and the floor
def fig_measure():
"""What "busy per frame" means, and why it is not the sum of durations.
LEFT is a real schedule at 9x9 proportions, one lane per STREAM, produced by
_schedule() rather than drawn: the copy lanes are single FIFO resources --
H2D_overlap and D2H_overlap are 1.000 in every row of probes.csv, because the
GPU has one copy engine per direction -- so those bars stagger, while kernels
from different streams sit on top of each other in time. Under the lanes, the
two union strips are computed from that same schedule, which is the whole
point: the engine's busy time is the union of its intervals.
No microseconds on the left panel. The schedule reproduces the SHAPE of the
9x9 pipeline but not its exact overlap factor, and putting numbers on a
schematic next to a panel of measured ones invites them to be read across.
RIGHT is measured: s4 engine occupancy at 9x9 f64 cap 1700 from probes.csv
(20.77 / 32.66 / 25.25) and the best unprofiled sustained rate from
ladder_9x9.csv opt6 (30.01).
"""
H, K, D = 21, 43, 25
NF, NS = 8, 4
frames, _ = _schedule(NF, NS, H=H, K=K, D=D)
fig, (ax, ax2) = plt.subplots(1, 2, figsize=(11.6, 2.95),
gridspec_kw={"width_ratios": [1.62, 1]})
# ---- left: four stream lanes, then the union each engine actually sees
LANE, TOP = 0.52, 5.0
for st, h0, k0, d0 in frames:
y = TOP - st * 0.72
for t0, dur, col in ((h0, H, AMBER), (k0, K, ACCENT), (d0, D, PALE)):
ax.broken_barh([(t0, dur)], (y, LANE), facecolors=col,
edgecolor=BG, linewidth=0.8, zorder=3)
for st in range(NS):
ax.text(-6, TOP - st * 0.72 + LANE / 2, f"stream {st}", color=MUTED,
fontsize=9.5, ha="right", va="center")
def union(iv):
pts = sorted(iv)
out = [list(pts[0])]
for a, b in pts[1:]:
if a <= out[-1][1]:
out[-1][1] = max(out[-1][1], b)
else:
out.append([a, b])
return out
lanes = [("H2D engine busy", [(h0, h0 + H) for _, h0, _, _ in frames], AMBER),
("kernels engine busy", [(k0, k0 + K) for _, _, k0, _ in frames], ACCENT)]
ys = 1.62
for name, iv, col in lanes:
for a, b in union(iv):
ax.broken_barh([(a, b - a)], (ys, 0.34), facecolors=col, zorder=3)
ax.text(-6, ys + 0.17, name.split(" ")[0], color=col, fontsize=9.5,
ha="right", va="center", fontweight="bold")
ys -= 0.62
ax.text(frames[-1][3] + D + 8, TOP - 0.72, "one copy engine\nper direction,\n"
"so the H2D bars\nqueue", color=MUTED, fontsize=9.5, va="center",
linespacing=1.4)
ax.text(frames[-1][2] + K + 8, 1.45,
"UNION — what s4 reports.\nThe kernels overlap, so it is\n"
"shorter than their sum.", color=TEXT2, fontsize=9.5, va="center",
linespacing=1.5)
ax.plot([-2, frames[-1][3] + D + 2], [2.42, 2.42], color=RULE, lw=0.9, zorder=1)
ax.set_xlim(-74, frames[-1][3] + D + 86)
ax.set_ylim(0.75, 6.05)
ax.set_xticks([]); ax.set_yticks([])
bare(ax, keep=())
ax.set_title("s4 · the shipped pipeline, four streams · schematic, 9×9 shape",
color=MUTED, fontsize=9.5, loc="left", pad=10)
# ---- right: the measured occupancies, and the two estimates of the floor
vals = [("H2D", 20.77, AMBER), ("kernel", 32.66, ACCENT), ("D2H", 25.25, PALE)]
xs = np.arange(3)
ax2.bar(xs, [v for _, v, _ in vals], width=0.56,
color=[c for _, _, c in vals], zorder=3)
for x, (_, v, _) in zip(xs, vals):
ax2.text(x, v + 1.0, f"{v:.2f}", ha="center", color=TEXT2, fontsize=10)
ax2.axhline(32.66, color=MUTED, lw=1.0, ls="--", zorder=4)
ax2.text(2.42, 33.3, "engine max\nprofiled 32.66", color=MUTED, fontsize=9.5,
ha="left", va="bottom", linespacing=1.4)
ax2.axhline(30.01, color=GREEN, lw=1.8, zorder=5)
ax2.text(2.42, 24.4, "best sustained\nFLOOR 30.01 µs/frame\n= 33 323 FPS",
color=GREEN, fontsize=9.5, ha="left", va="center", linespacing=1.4)
ax2.set_xticks(xs)
ax2.set_xticklabels([n for n, _, _ in vals], color=TEXT2, fontsize=10)
ax2.set_xlim(-0.55, 4.35)
ax2.set_ylim(0, 40)
ax2.set_yticks([])
bare(ax2, keep=("bottom",))
ax2.set_title("busy µs / frame · 9×9 f64 · measured",
color=PALE, fontsize=9.5, loc="left", pad=6)
fig.subplots_adjust(left=0.085, right=0.99, top=0.86, bottom=0.09, wspace=0.20)
save(fig, "fig_measure")
fig_measure()
legibility_report()
Binary file not shown.

Before

Width:  |  Height:  |  Size: 127 KiB

After

Width:  |  Height:  |  Size: 118 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 124 KiB

After

Width:  |  Height:  |  Size: 115 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 79 KiB

After

Width:  |  Height:  |  Size: 71 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 111 KiB

After

Width:  |  Height:  |  Size: 98 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 41 KiB

After

Width:  |  Height:  |  Size: 36 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 100 KiB

After

Width:  |  Height:  |  Size: 92 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 106 KiB

After

Width:  |  Height:  |  Size: 72 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 138 KiB

After

Width:  |  Height:  |  Size: 124 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 94 KiB

After

Width:  |  Height:  |  Size: 68 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 107 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 85 KiB

After

Width:  |  Height:  |  Size: 85 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 77 KiB

After

Width:  |  Height:  |  Size: 79 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 21 KiB

After

Width:  |  Height:  |  Size: 22 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 58 KiB

After

Width:  |  Height:  |  Size: 37 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 80 KiB

After

Width:  |  Height:  |  Size: 76 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 67 KiB

After

Width:  |  Height:  |  Size: 68 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 118 KiB

After

Width:  |  Height:  |  Size: 106 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 84 KiB

After

Width:  |  Height:  |  Size: 94 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 72 KiB

After

Width:  |  Height:  |  Size: 69 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 86 KiB

After

Width:  |  Height:  |  Size: 87 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 88 KiB

After

Width:  |  Height:  |  Size: 57 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 84 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 56 KiB

After

Width:  |  Height:  |  Size: 50 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 504 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.1 MiB