CS-authored science draft for the Comms publishing pilot. Numbers are our own
receipted measurements; judgments marked CEO-eye are single-judge human calls
and are labeled as such throughout. Limits stated up front: single judge, no
blind scoring, one machine, one model family — the structure (paired seeds,
frozen plans, append-only indexes, hash-stamped outputs) is the defensible
part. This draft makes no claim beyond its n.
A ~$1k consumer APU (8-core Strix Halo, iGPU, 80 GiB GTT shared memory) runs a
7B-parameter diffusion transformer (Q4_K) alongside two resident LLM services
(27B champion + 35B MoE) with no partitioning and no mode switching. The
question the org needed answered: what quality and economics does this class
of hardware actually deliver, measured rather than guessed?
Co-residency. Image generation at full quality runs at identical speed
with both LLM residents live (5.84–5.88 s/it across memory modes; peak GTT
58.8 GiB / 80, ~15.5 GiB RAM free). The "park the big models before drawing"
reflex is unnecessary on this memory architecture. [2026-09-23 receipts]
Sampler separation. At 30 steps: samplers euler, euler-a are
speed-identical; dpm++2m costs +5.6% (cfg 4) to +15.9% (cfg 6) wall time. At
low steps (6/10) all samplers are speed-identical — separation is purely a
quality question. CEO-eye judgment: dpm++2m preferred at 30 steps, and is the
only sampler producing usable output at 6 steps; euler-family output at 6
steps is under-integrated (missing content, unresolved exposure). Mechanism
(offered, consistent with solver theory): dpm++2m is a second-order multistep
solver whose per-step error shrinks faster, so it stays on the ODE trajectory
when strides get huge; euler's straight-line steps leave the walk unfinished.
Guidance (cfg) economics. cfg 4 vs 6 at 30 steps: ~10% wall-time savings
(191.7 s vs 211.4 s per 768×512 image, n=10+10). Quality call on real brand
content: pending the CEO eye on like-for-like cells (section 5).
Reload tax. Each fresh CLI invocation pays ~15–20 s model reload; a
20-image batch spends ~6 min of its ~65 min on reloading. Same-process
sequential generation is the first optimization lever (est. ~9% throughput,
larger for short runs). [reload-tax receipts]
Text capability. Stitched/embroidered vertical lettering rendered
perfectly (CEO-eye, both cfg settings) — the hard text class (letters as
thread on fabric). The empty OCR read on the same images is an instrument
artifact, not model failure.
Cost per image. Full-quality 768×512: ~3.2 min wall, ~54 W GPU package
power — hours of generation at quiet-class draw, zero instability across all
receipted runs. Contrast: the same work CPU-only was ~29 min/image and
crashed the machine under sustained load; the GPU lane is both faster and
safer here.
Every run in this paper is: a frozen plan (prompts hash-pinned), an
append-only index, per-image sidecars, exact-hash outputs, preflight receipts
gating shared-resource use, and a live served surface for human judgment.
Failures are recorded in place (one cell killed and regenerated on the
record; a terminated-and-adopted run disclosed). This is the paper's second
product: a receipts-first bench pattern other small labs can copy.
No blind scoring, no inter-rater statistics, single machine class, single
model. The CEO-eye judgments are directionally strong (paired same-seed
comparisons) and statistically n=1. We claim measured facts and a repeatable
method, not a leaderboard.
quality; texture-class ceiling verdicts).
the prompt-confounded earlier comparison.
first.
All receipts committed in the lab repo: sampler/low-step grid + reload-tax
run (2026-09-24/25), co-residency receipts (2026-09-23), 20-cell brand
content hunt + telemetry (2026-09-25). Served raw copies available via the
lab's :4690 file surface.