Limits of observability on the ET-SoC-1
Start here · interactiveThe ET-SoC-1, interactively →Click through the die, a shire and a minion; watch eleven data flows (a load to DRAM, the latency ladder, TensorSend, the relay, gathers, the host over PCIe, a matmul and its heat, a hot line, the allreduce, a broadcast to every minion) with the latency, bandwidth, energy and heat measured on these cards; a 17-step tour for presenting.This chip ships with its manuals, errata, firmware source, functional simulator and the RTL of its core. This page indexes every measurement made on three of its cards and asks how far down they let you see: one cycle, one register, one event, one bit.
docs/reports/data/2026-09-25-claims-v3.A. Why is this chip low power, and what does a FLOP cost? For a newcomer, about 45 minutes.
- Why is the ET-SoC-1 low power? the equation measured term by term, against an A100
- The Horace experiment, §1, §3, §6 why the data, not the FLOPs, sets the watts
- The energy manual, §1–§3 the card at rest, awake, per instruction
- The DVFS loop and its leakage, §1–§2, §5 the governor, and what leakage costs
B. Writing fast kernels. For a programmer.
- Test drive the toolchain, the simulator, a first SGEMM 18 Sep baseline
- Matmul efficiency the tensor unit near its peak 18 Sep baseline; energies from three cards
- Memory hierarchy latency and bandwidth of every level 18 Sep baseline; energies from the energy manual
- Ridge points how much work per byte each level needs
- On-chip communication the mesh, messages, trees, and the TensorSend trap 18 Sep baseline; energies from the energy manual
- Hand it to the next shire relaying between shires instead of DRAM
- One hot line stops a shire contended atomics, and what to do instead
- Sparse compute what the tensor unit does with zeros 18 Sep baseline; power from three cards
C. What does each operation cost? For an energy modeller.
- The energy manual every cost, with bars from three cards
- Heat per millimetre a bit carried across the mesh, by bit pattern
- The Horace experiment, §8 from flips to watts to degrees
- This page, §4 what the meters do not see
D. How far down can you see? For tools and observability.
- This page, §2–§4 every instrument, and the meters' limits
- Power and temperature, §2–§3 a load step through every sensor; the voltage map
- Anatomy of a memory access one load taken apart to ±3 cycles energies from the energy manual
- Spatial temperature: a brief 35 sensors the host never sees one by one
- The DVFS loop, §1 the firmware as the gate
The chip in brief: minions, harts, shires, scratchpads, the SP, the PMIC and the other names on this page
The ET-SoC-1 is Esperanto's RISC-V accelerator, now open-sourced by AI Foundry. Its compute cores are minions: small in-order RISC-V cores, each with two hardware threads (harts), an 8-lane fp32 vector unit and a tensor unit whose matrix multiply-accumulate instruction is TensorFMA (one fp32 op multiplies 16×16×16 tiles: 4,096 multiply-adds). Only hart 0 issues tensor operations (hart 1 may only prefetch into the L2 scratchpad). Eight minions form a neighbourhood and 32 a shire. Each shire has 4 MB of SRAM, which these cards' firmware splits into a 512 KB L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad: software-managed memory that any shire can address. Each minion also sets aside 3 KB of its 4 KB L1 data cache as an L1 scratchpad for tensor operands.
The chip has 34 minion shires (1,088 minions): 32 run kernels (1,024 minions), the master shire runs firmware and one is spare. With the I/O and PCIe shires they form a 6×6 grid on a mesh network-on-chip (NoC, 400 MHz); eight memory shires with the LPDDR4X controllers (32 GB) sit along two sides, making 8×6 mesh stops. A hop is one step between neighbouring stops, about 3.72 mm. A flit is the unit the mesh moves as a whole.
The service processor (SP) is the on-die management core: its firmware reads the sensors, runs the clock and voltage governor (600, 700 or 800 MHz) and answers the host's management commands (DM_CMD_*). The PMIC is the board's power-management controller: it meters the 12 V input and, over PMBus, three regulators (the minion, SRAM and NoC rails). Moortec PVT monitors measure temperature and voltage on the die. Kernels run in user mode (U-mode) and firmware in machine mode (M-mode). The four Maxions are larger out-of-order RISC-V cores on their own 0.6 V rail.
Other names on this page: A0 is the silicon revision (stepping) that the errata describe; PShire is the PCIe shire; VDDQ is the DRAM I/O supply; MDI is the Minion Debug Interface that the SP exposes to the host; BL1 and BL2 are the SP's boot stages; OTP is one-time-programmable fuse memory; the PRM is Esperanto's Programmer's Reference Manual; sys_emu is the functional simulator; ettelem is this project's telemetry client. aifoundry2 and aifoundry3 (a2, a3) are lab machines with one card each, and aifoundry1 holds two: its card 1 (aifoundry1-c1, a1c1, on older firmware, 1.2.0 against the others' 1.3.1) is the third card measured, since 25 September; its card 0 overheats under load and is not used. aifoundry3's card is held at 600 MHz by a power limit of 0 W that a lab service sets at every boot. The Horace experiment is named after Horace He's post showing that GPU matrix multiplies run faster on predictable data (“Strangely, Matrix Multiplications on GPUs Run Faster When Given ‘Predictable’ Data!”).
1. The measurement reports
Each session (an E-ID of the experiment register, §7) sits in the lane of the card it ran on, and each report at the day it was first published. A session on several cards sits in each of their lanes. Grey curves run from a session to the reports it fed; dashed arrows from a number's first page to the page that now holds it. Hover, tap or tab to any mark; Enter or a click selects it and shows its details.
Every claim the version-3 claims check listed, one bar per page: the upper bar is its status after the version-3 campaign, which is complete (25–26 September: each claim it tested takes its outcome on aifoundry2, aifoundry3 and aifoundry1-c1), the lower its status before it (the plan's verdicts, 25 September). Each bar is split by verdict (a number quoted from another page takes that page's verdict), or by the cards whose data the claim rests on. A page's own note (“Checked on three cards”) gives both counts of the claims the campaign tested there: judged on the numbers the page quotes, and by this chart's rule, which counts a tested claim under “a test behind it failed” when any part of any registered test covering it failed, even a part the page does not quote, and under “fewer than three repeats” when a card lacked them; a tested claim the campaign only reported per card keeps its earlier verdict. So a page's “held” can be the chart's “failed”: all 17 of the hot line's tested claims rest on its registered test LAT-H, which failed in parts the page does not quote (the shares, the cycles per atomic, the pollers), and the relay's rest on LAT-R, which failed on two cards and was incomplete on the third. Hover, tap or tab to a bar for its counts, and for both readings of its tested claims; a click or Enter opens the page (on a touch screen, a second tap).
What each verdict means
The full index as a table
2. The ladder
Every instrument, roughly coarsest to finest: what "status" and the check column mean
Every instrument, roughly coarsest to finest. Status says what it takes on this card: works now from a user account; tooling means the hardware and firmware are there but a client, a build or an external box must be written or installed; firmware means a small source change plus a reflash of a signed image (§5); research means possible in principle with things not in the open drop; not on silicon is the hard wall. The check column shows whether two reviewing AI agents confirmed the row's citations (✓), corrected them (±), or did not get to it (·); rows added after 20 September are unreviewed (·) by that survey (§8), and rows rewritten after their review are marked †. §8 says how the review was done.
Each dot is the finest time step the instrument resolves on the card; its colour is what it takes to use it. Hover, tap or tab to a row for its granularity; Enter opens its row in the table. The buttons under the chart filter the chart and the table together.
Every instrument, as a table
Every event the reports priced, on one scale of joules, against what one meter reading resolves. An event becomes visible when the card can run enough of them a second to lift the power by the precision asked for above the uncertainty of the baseline: it needs a rate of σ ÷ (precision × energy). Filled dots are events the measurements ran fast enough to see at these settings; hollow ones are not: at the rate measured, a burst of them moves the power by less than σ ÷ precision. The tick on each row is the smallest energy per event that the measured rate would reveal.
3. Can we track individual bits flipping?
On the card: no, and not with any firmware change. Nothing on the die counts toggles. The PMU counts operations (up to 8 per cycle per neighbourhood), the shire-cache and memory-controller monitors count requests, and the power path resolves at best 1 mW over one service-processor pass (§4.1), twelve orders of magnitude above one node toggle. The closest things to a bit read are the ECC machinery and the debug module, and both stop at architectural state:
ECC, the debug interface, raw SRAM, UltraSoC, and physical probing, in detail
- ECC covers the L2/L3/scratchpad SRAM and the instruction cache only. The L1 data cache, the tensor
scratchpad carved from it, the integer register file and the 256-bit vector register file have no ECC and no parity, so
a flip there is invisible to hardware. DRAM ECC is compiled off (
ddr_mode.ecc=false). The SRAM error logger records which 64-bit doubleword of a line failed, not which bit. And in the firmware source read here, the service processor never enables the interrupt sources that carry SRAM errors, soSramCeEventandDM_CMD_GET_MODULE_SRAM_ECC_UECCread 0 forever. (We read them on both cards: 0.) - The Minion Debug Interface reads state, not flips. The service processor can halt a neighbourhood, inject two-instruction pairs into a hart, and hand back any GPR, the PC, seven CSRs and 8-byte memory words. That is one architectural snapshot per millisecond-scale round trip. It is compiled into the shipped firmware and the management node is world-writable, but no installed tool sends the commands, halting mid-kernel wedges the kernel, single-step is a stub, and the boot code clock-gates the debug logic, so the first MDI command on this card is itself an experiment.
- Raw SRAM rows with their check bits can be read and written through the shire cache's M-mode index-cacheop state machine
(
Dbg_Read/Dbg_Write, PRM Table 15-81). That is a literal bit read of a stored line, and a fault injector, but it needs a syscall the firmware does not have. - The UltraSoC debug fabric is on the die but unreachable. The RTL wires a per-minion instruction-trace encoder (retired instruction and PC every cycle), per-neighbourhood status monitors (200 filter, 64 match, 128 data signals, 3 counters) and a memory-shire monitor over 200 controller signals. The IP, its message-engine protocol and the Data Book are not open, no firmware line configures it, and the SP gates its clock at boot.
- Physically: scan chains exist (a one-shot, stopped-clock flop dump with tooling that is not open), and laser or e-beam probing is a different project. Not for a shared PCIe card.
Nor did the card report a wild write, the one time we made one: on 24 September, during the wire-energy work, a test-tool bug sent stray tensor stores to physical address 0 in two test launches of about half a second each (which card was not recorded), and no error counter, telemetry field or clock moved.
A first result from the RTL: the hpmcounter3 carry bug
A first result from the RTL. On every card measured (aifoundry2, and since 26 September aifoundry3 and aifoundry1-c1, five launches of each test on each), hpmcounter3 reads 128 short just after its low 7 bits
wrap: two reads 10 cycles apart differ by 10, 138 or −118 cycles and nothing else. Simulating the open RTL's neigh_pmu.v under Verilator (rtl-sim/pmu_carry) shows why: each
counter is a 7-bit pre-counter plus a 57-bit post-counter, twelve counters share one adder that folds pre-counter
overflows into the post-counters in round-robin order, and a read ignores the pending overflow bit. So after every wrap
the value is 128 short until the adder comes round again: 12 cycles in the simulation. On the cards the short window is not a fixed length. Where one window fits a whole launch it is consistent with low bits 0 to 10 (six launches on aifoundry2, one on aifoundry1-c1), but in 26 of the 30 launches of paired raw reads no single window fits them all, and the usual fix (fixcyc(), which adds 128 to a read whose low bits are below 11) left 0–4.7% of 10-cycle intervals off by 128 in a launch: none in four launches of five on aifoundry2 and 2.1% in the fifth, 1.7–4.7% in all five on aifoundry3, up to 0.8% on aifoundry1-c1. The mechanism is a silicon
measurement explained by a flip-flop-level view of the design, which is the point of having the RTL; the window's length on the card is not yet explained.
In simulation: yes, twice over
In simulation: yes, twice over. sys_emu -l prints every executed instruction on any selected hart with each architectural write it
makes: integer registers, all 8 lanes of a vector register, mask registers, every memory read and write up to 512 bits, CSR writes, traps. XOR
successive values and you have the bits that changed. It is functional, not cycle-accurate, and it models no cache contents. The RTL bench goes
further: +dump=1 writes every net, flop and behavioural SRAM bit of the CPU subsystem to an FST waveform each cycle, and the design
verification monitors emit a cycle-stamped event per retired instruction that is checked in lockstep against the same emulator. Toggle counts per
signal are a coverage flag. What you cannot get from the drop is joules: there is no cell library, netlist or power flow, so toggles stay counts.
Joules have been attached to flips in two places, both by fitting to measured power: the tensor unit, with four event energies (the Horace experiment; its flip counts come from rtl-sim/fma_toggle, the core's multiply-add RTL under Verilator), and the mesh links, where a bit that differs from the previous flit costs about 98 fJ per hop and a one carried about 129 fJ on the mesh rail (Heat per millimetre, §5). The chart at the end of §2 sets both against the meter.
4. Power and energy: the meter chain, and what it does not meter
4.1 The chain
How the PMIC and the service processor meter the three rails, and how stale a reading can be
The card's power is read by a PMIC on the board — a microcontroller with its own closed firmware — that measures the 12 V input and, through PMBus, the three digital regulators that feed the minion cores, the SRAM arrays and the mesh. For each of those three it holds voltage, current and power on both sides of the regulator, and its temperature, each as a current value, a minimum and maximum since the last reset, and a running average: 84 numbers.
The service processor reads all 84 every loop pass and forwards to the host the PMIC's running average of each output-side power, with its min and max, and the 12 V input power; ettelem samples them at 10 Hz. The average is roughly first-order: after the board steps down at the end of a burst, the rail reading has fallen , which is why the rails take 2–3 s to catch up with the load step in Power and temperature, §2. That τ is one fit over the catalogue's bursts and absorbs the service processor's one-pass delay; fitted rail by rail with the delay kept apart (E58, 29 September), τ is 1.01–1.06 s on aifoundry3's rails, 1.08 s on aifoundry1-c1's minion and NoC rails and 0.54 s on its SRAM rail, and tools/ettelem/deconv.py undoes the filter (rung 4). Resetting the PMIC's minimum and maximum (ettelem --reset-ms) restarts that average as well: with a reset every second, the rail reading had fallen 93% of the way one second after a burst, on all three cards (the median of seven to nine bursts on each), instead of the fraction above. The SP's copy refreshes once per pass of its loop. While ettelem samples at 10 Hz, as in every session from 20 September on, a new board value arrives about every 156 ms on aifoundry2, 158 ms on aifoundry1-c1 and 263 ms on aifoundry3 (phase-folding the board-power stream, the mean of three passes on each card in the version-3 campaign, TEL-S; in the earlier power sessions the readings changed every 148–164 ms on aifoundry2, over 29 sessions, and 251–260 ms on aifoundry3, over 16). The SP's own trace times the pass directly: 133 ms on aifoundry2, 135 ms on aifoundry1-c1 and 224 ms on aifoundry3 with no sampler running, and about 160, 162 and 266 ms while ettelem samples (the medians of three passes on each card, TEL-P1, P3 and P5). So the sampler lengthens the pass on every card, by 26–42 ms; under the sampler the trace's pass runs 3–5 ms longer than the refresh the board-power stream shows, the difference between the two methods on each card. In the firmware source read here (353f20e, not the cards' older build) the I2C driver waits a millisecond after each transaction, the likely reason a pass takes that long. Why aifoundry3's loop takes about 1.7 times as long as the other two cards' (aifoundry2 runs the same firmware release) is not established. How often the PMIC itself updates is unknown. The other rails — — have set-point registers in the PMIC and no current or power telemetry (their on-die voltages are reported, §4.3).
The workload can starve the meter
The workload can starve the meter. Rings of messages between shires s and s+16 slow the service processor's management path: a telemetry sample takes instead of 22, and a burst measured through it reads low (33–45% on aifoundry2 on 23 September). So do tensor loads between shires three hops apart in one column, on aifoundry2 (the heat-per-millimetre runs). The sampler records its own latency with every sample, and bursts with that signature are dropped (rung 7). A sampler killed in the middle of a request stops the meter outright: its reply stays in the management queue, and every later opener crashes on it until dev_mngt_service drains it (seen once on each card, E32).
Every catalogue burst's own sampler latency (the time one telemetry sample took), one strip per card; a dot's size is how many bursts took that long. Diamonds are rerun passes dropped from the manual because the sampler was this slow. Hover, tap or tab to a mark.
So the meter chain gives one number for the whole card and three for its inside, and the difference between them is the subject of this section. At idle it is : the first bar of the chart in §4.2. The version-3 campaign did not confirm the aifoundry2 figure: its idle cycles on that card passed through 73 °C only twice, one short of the three repeats the check asks for (15.05 W of 31.73 W, 99% interval 14.56–15.54 W); on aifoundry1-c1 three cycles gave 17.87 W of 42.71 W at 73 °C.
From the 12 V input to a number in this page's charts. The grey box at the bottom carries no current or power telemetry at all. Hover, tap or tab to a box for what it does and its numbers on the selected card.
4.2 The unmetered remainder, attributed
The remainder cannot be metered with anything on the card, but it can be attributed. Over the energy manual's catalogue — , three passes on each of three cards — each configuration's power above idle is known on the board and on the three rails, and the bytes it moved through DRAM are counted. Fitting the unmetered part of each configuration's mean as a fraction of each rail's power plus a cost per DRAM byte gives:
The fit's coefficients, per card
Each coefficient's estimate, with one robust (HC3) standard error either side; one whisker per card in every row, top to bottom in the order of the table's columns. Hover, tap or tab to a whisker.
How the fit was done
Three things the fit implies
4.3 What the Moortec sensors can and cannot do for power
The die's Moortec PVT subsystem, and what of it reaches the host, is in the table below. The spatial temperature brief maps the sensors onto the 6×6 grid of minion, IO and PCIe shires (8×6 counting the memory shires at the sides) and traces how the firmware collapses them to one average.
None of it measures current, so none of it meters the unmetered rails. What the monitors do measure is voltage, and voltage droops with current. The host already receives, every SP pass (§4.1), the memory shires' reading of the 0.8 V DDR rail (DM_CMD_GET_ASIC_VOLTAGE, ettelem's die_mv.ddr): across the catalogue it reads at idle against an 800 mV set point, and under load it droops roughly in proportion to the DRAM power the regression attributes off-rail:
Every catalogue configuration, in the colours of the chart in §4.2; aifoundry2 by default (the calibration above), or pick a card. Hover, tap or tab to a point for its droop, the droop the calibration predicts, and the DRAM power the excess would be read as.
How the droop calibration was done
The eight configurations of the first table, to 0.1 mV
A current meter for DRAM power, calibrated once per card — and its limits
That is a free current meter for DRAM, the largest unmetered consumer under load, in every ettelem log, updating every SP pass with . Its scale comes from the fit above, so it is not independent of the board meter, and it responds mostly to DRAM traffic, but not only: So in bursts of one kind it separates DRAM from arithmetic, not from mesh traffic; a mixed burst is untested. .
The same monitors see the minion rail sag by , which Power and temperature (§3) mapped shire by shire at idle from the SP's debug trace (one idle capture on aifoundry2; the three-card check captured all 34 shires on each card but got too few complete idle captures on aifoundry2 and aifoundry1-c1 to confirm the map); calibrated shire by shire, that map is a spatial power meter for the three metered rails (rung 5). The temperature sensors and the process detectors, exported, would add heat maps and a leakage proxy per shire; they too need a firmware change.
5. The improvement ladder
Every way found to see more, grouped by what it costs: software on the data the card already gives; lab hardware that needs no firmware change; firmware changes, each ending in a signed image; tooling against interfaces that already exist (a debugger client, an RTL flow); and research. Two more groups collect what the chip diagram still infers (its dashed parts, the routes it draws, the host's path to DRAM: the last two measured since 29 September, rungs 32 and 34), what the PCIe page left open (why two host-to-card copies at once halve the rate, one stream's since 29 September, rung 35; the link's payload size) and what the memory levels, the interactive diagrams of each level of the memory system down to the transistors, leave unknown (which cell each memory uses, since the lab lead says the chip is not using SRAM; the memory macros' geometry; where each cycle of a cache access goes; the live cache settings; the DRAM part), and what the effect of overheating could not settle (what stopped a card near 120 °C; what the PMIC's temperature measures; the chip's thermal ratings): the documents, interfaces and readings to ask AI Foundry (Ainekko) and the lab for, each rung ask the team, and the experiments on the cards that would settle the same parts without them. Rungs 21–36 are the chip diagram's and the PCIe page's; rungs 37–44 are the memory levels', whose other unknowns are added to rungs 20, 21, 23, 25, 26, 28 and 32, which already asked the same thing (rungs 19 and 33 gain a note). Rungs 45 and 46 are the overheating page's; its other asks extend rungs 13 (a firmware build for the process detectors), 24 (the thermal ratings) and 42 (the DRAM's temperature grade). Each of those rungs names the part of the diagram it settles and links to it, and each diagram links its unknown parts to their rungs. Their firmware line numbers are at et-platform 836a4ab, as in the diagram. Each rung says what it adds and what it would change in the numbers; where a rung is an instrument of §2, the link opens that row. Rungs marked done say where they were done. Rungs 4, 31–36 and 43 were run on the cards on the night of 28–29 September (E55–E58 in §7), each under predictions frozen after development on aifoundry1-c1 and tested on aifoundry3; rungs 31 and 43 are not marked done: the first was not decided on the validation card, and no registered theory of the second survived. The firmware caveat at the end of the section applies to every rung in its group.
What each rung does to the unmetered watts, and the firmware caveat
What each rung does to the unmetered watts. Today: ; rungs 2 and 3 attribute the part above idle (§4.2, §4.3) and leave the idle W unsplit. The riser (rung 8) does not meter them either, but it makes every bracketed burst sharper and shows the input side of every regulator's transient. The camera (rung 9) says which package the watts warm. Forwarding the PMIC's input-side readings (rung 11) measures the delivery losses and leaves only the rails with no meter — DDR, VDDQ, PCIe, IO, Maxion — as a remainder: . Below that is the board, and the board is shared.
The firmware caveat. Every "firmware" row above ends with a reflash. The minion runtime and the service processor bootloader are signed executable images; BL1 and BL2 verify a certificate and signature against a public-key hash in OTP unless a chicken bit is set, and neither the signing tool nor a test key is in the open tree. Whether this card accepts a rebuilt image is unknown until someone tries, and a bad image can brick the boot path, so it is a lab-admin decision. This is the one place where "everything is open" has a gap: the source is open, the right to run modified source on this particular card is not yet established.
6. Against a GPU
For an A100/H100-class part. Sources are NVIDIA's own documentation and the microbenchmark literature, listed below the table; the ET side is this survey.
The pattern: the GPU is at least as observable on silicon (ECC counts, a debugger, SASS-level instrumentation, a mJ energy counter the ET card lacks), and entirely opaque below it. The ET-SoC-1 is thinner on silicon (no L1 or register ECC, DRAM ECC off, power once per service-processor pass) and open all the way down: the firmware that decides what you can see is source, the reference model the RTL was verified against is the installed simulator, and the RTL of the core dumps every bit.
Sources for the GPU column
- Jia, Maggioni, Staiger, Scarpazza, “Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking”, arXiv 1804.06826, 2018.
- Luo et al., “Benchmarking and Dissecting the Nvidia Hopper GPU Architecture”, arXiv 2402.13499, 2024.
- Yang, Adamek, Armour, “Accurate and Convenient Energy Measurements for GPUs”, SC'24 (preprint: “Part-time Power Measurements: nvidia-smi's Lack of Attention”, arXiv 2312.02741).
- Lin et al., “GPUHammer: Rowhammer Attacks on GPU Memories are Practical”, USENIX Security 2025.
- Khairy, Shen, Aamodt, Rogers, “Accel-Sim: An Extensible Simulation Framework for Validated GPU Modeling”, ISCA 2020.
- NVIDIA developer forum, “Questions about globaltimer functionality, accessing and configuring” (the %globaltimer update rate).
- NVIDIA documentation: CUDA Binary Utilities, PTX ISA, the CUDA Programming Guide (clock64), the Nsight Compute Profiling Guide, the NVML device queries, GPU Memory Error Management r575, the CUDA-GDB manual and the open-gpu-kernel-modules README.
- This project's notes on GPU power telemetry:
docs/research/power-telemetry.md.
7. The power and temperature sessions
Every registered session (E1 onward) that measured power or temperature, and the experiments run for §5's rungs (E54–E58), in order, with what it established, the instruments it used, where its raw data is in the repository (paths are under docs/reports/data/ and link to GitHub) and the reports that use it; two sessions of 18 September were registered on 25 September (E33, E34), and that day's matmul and sparsity sessions, which predate the register, share one row. The map in §1 shows them on a timeline, by card. The IDs are those of the experiment register, docs/findings/03-experiments.md, which gives each session's command, protocol and caveats.
All sessions as a table
8. Method and caveats
Method, versions, and caveats in full
- Who did the survey. Seven AI research agents (Claude subagents run by the author) each surveyed one layer: the manuals (
external/et-man), the firmware source (external/et-platformat353f20e, checked against HEAD), the RTL and DV trees (external/core-et), the installed tools in/opt/et, the GPU literature, the debug path, and the earlier reports. Each key claim was then given to two further AI agents told to refute it, one checking citations and one checking whether it would work from a user account on this card; 40 of the 73 verdicts corrected the claim as first stated. The ladder folds in their corrections; a person has not re-checked each row. The raw findings and verdicts are indocs/reports/sources/2026-09-20-limits-of-observability/. - The power section reads the PMIC and PVT drivers of the same firmware (
ServiceProcessorBL2/driver/pmic_controller.c,pvt_controller.c,services/thermal_pwr_mgmt.c) and fits the energy manual's catalogue (since 26 September the version-3 campaign's full catalogue on three cards,docs/reports/data/2026-09-25-claims-v3/raw/<card>/catfull/; the 23 September runs are indocs/reports/data/2026-09-23-catalogue-aifoundry2/anddocs/reports/data/2026-09-23-catalogue-aifoundry3/); the fit and the droop calibration are indocs/reports/data/2026-09-23-energy-manual/unmetered_fit.json.tools/ettelem/fit_unmetered.pywrites that file from the catalogue and its telemetry, andtools/ettelem/sync_hub_data.pycopies it, the rail filter, the sampler's latency and the energies of §2's chart into this page's data, and computes the fit's robust errors, its values pass by pass and the droop calibration on each card's own telemetry (--checksays whether the page is current). Those rows were not in the 20 September two-agent review; the 24 September cross-check recomputed their numbers from the data (its record). - Which firmware. The firmware statements on this page describe the et-platform source at
353f20e. The cards' own trace strings match an older build (before et-platform commit60b40c10f, 24 September 2024); aifoundry2 and aifoundry3 report release 1.3.1 and aifoundry1-c1 release 1.2.0 (PMIC firmware 1.5.0 and 1.3.0, read again in the version-3 campaign); the older source has since been read (27 September): release 1.3.1's service-processor firmware, BL2 0.20.0, has its closest public source at et-platformffca4cbb4, and release 1.2.0's, BL2 0.18.0, atda192816a. Their governor differs from353f20ein ways that matter for heat (a blocking thermal loop of about 0.4 s steps with no busy test, an exit to the boot point, a one-call climb), and the development night of 28 September (E51) fits 0.20.0; its frozen validation on the same card (28–29 September) confirmed the missing dead band and the heartbeat's latencies but left the 0.20.0 loop itself untested (two of 46 idle intervals just off its grid, and too few descents). The DVFS report, §1, gives that governor line by line, and its §9 what is not established. - Nothing done for this survey changed the card's firmware or configuration. The only measurements taken for this page were reads of the ECC counters and sysfs error statistics, and the reports' own timing and power runs.
- The RTL is core-et's Erbium branch at
b38a1a3: the same Minion core lineage in a later MCU-class configuration (one neighbourhood, no shire cache), with known deviations, not the taped-out ET-SoC-1. The PRM and errata stay authoritative for the A0 silicon. The same repository'smainbranch holds a clean-room, Verilator-first re-implementation including a shire cache, with toggle coverage in CI; it is a re-implementation, not the taped-out RTL. - Several rows rest on a single reading of the sources and were not exercised: MDI on this card, the debug clock gate, the MMST trace behind device-wide bandwidth, the memory-shire perfmon type codes, the per-shire LVDPLL (which an erratum calls unstable on A0), the external analog monitor inputs. They are labelled accordingly.
- How the cards are named. A measured statement that names no card holds on both aifoundry2 and aifoundry3, and on aifoundry1-c1 too where the text says “all three cards” (the version-3 campaign's cards; aifoundry1's card 0 overheats and was left out): at least three independent repeats on each (passes, sessions or separately started bursts), or a deterministic value identical on each. A card named in the text means that card only; fewer than three repeats are given in words (one run, one session); where the cards differ, both values are given. Results within noise, or causes that were not measured, are kept out of the argument. The tests and intervals behind each statement are in the version-3 claims record.
- Versions. One line per version; the artifacts register (A2) has more, and the source's history every wording. 20 September: first published (the ladder, bit flips, the GPU comparison). 23 September: the power meters, the unmetered remainder, the improvement ladder and the index of reports. 24 September: the heat-per-millimetre runs, and corrections from a cross-check of every report against its data. 25 September: version 3, every claim re-checked on both cards (133 ms is the SP's own pass on aifoundry2, not the board's refresh; the fit's DRAM error ±7%, not ±2%). 26 September: the campaign's results on three cards (the cycle counter's short window, the rails' reset, the service processor's pass on each card and the factor between cards corrected). 27 September: charts, E33, E34 and E48, the refresh told apart from the SP trace's pass, and §5's rungs 21–36 with the chip diagram and the PCIe page. 28 September: §5's rungs 37–44 with the memory levels; each page's note gives both counts of its tested claims; one sentence per index row; the PCIe page's slow host memcpy explained (rung 30); the effect of overheating in the index, E53 in §7, and its asks (rungs 45 and 46; rungs 13, 24 and 42 extended). 29 September: rungs 4, 31–36 and 43 from the experiments of the night of 28–29 September (E55–E58, each frozen after development on aifoundry1-c1 and validated on aifoundry3); E51's validation under way, E54 registered, E54–E58 in §7; the state of the lab fixes on the requests card; the sparse parity page in the index, beside the influence-functions page under “Research and exploratory”, and its session, E59, in §7. 30 September: E51's validation reduced, in §7 and the firmware note (TH3, TH4 and TH8 survived, TH7 fell, and the owner's two questions stayed untested); rungs 31–36 and 43 with E55–E57's third card, aifoundry2, run after it.
9. Related reports
Every report is in §1; the knowledge base behind them is docs/findings/, and the code and raw data are in the et-soc1-prototyping repository.