Limits of observability on the ET-SoC-1

Start here · interactiveThe ET-SoC-1, interactively →Click through the die, a shire and a minion; watch eleven data flows (a load to DRAM, the latency ladder, TensorSend, the relay, gathers, the host over PCIe, a matmul and its heat, a hot line, the allreduce, a broadcast to every minion) with the latency, bandwidth, energy and heat measured on these cards; a 17-step tour for presenting.

29 September 2026 (first published 20 September; versions in §8) · the AI Foundry cards in aifoundry2, aifoundry3 and aifoundry1 (its card 1, aifoundry1-c1) · firmware source read at et-platform 353f20e (the cards' own log strings point to an older build; see §8), RTL core-et b38a1a3 · the index of every ET-SoC-1 measurement report

This chip ships with its manuals, errata, firmware source, functional simulator and the RTL of its core. This page indexes every measurement made on three of its cards and asks how far down they let you see: one cycle, one register, one event, one bit.

Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 and aifoundry1-c1 (the version-3 campaign, 25–26 September, which planned three to five repeats of each test on each card). Corrected here: the cycle counter's short window is not fixed, a reset of the rails' readings restarts their average, aifoundry3's service processor loops every 224 ms, the factor between cards is smaller and partly temperature (over the eight Horace patterns, a least-squares ratio, aifoundry3 switches 0.958–0.969 of aifoundry2 with busy and idle power compared at the same die temperature, 0.895 as registered, and the catalogue gives 0.972), and figures quoted from the reports. Of 27 claims tested here: 9 held, 12 failed their registered test (10 in something the page says, now corrected; 2 only in a number it does not quote), 1 differs by card and 5 not confirmed, as §1’s scoreboard counts them. Record: docs/reports/data/2026-09-25-claims-v3.
Finest time on the card
1 cycle
hpmcounter3 from a kernel, after correcting a late carry into bit 7; a fixed correction still leaves up to 4.7% of intervals off by 128 on the three cards, so a kernel should check for them (§3)
Finest state on the card
64 bits
one GPR, CSR or memory word of a halted hart per management round trip; in the firmware, untried (§2)
One reading's energy step
156–263 µJ
1 mW × one meter refresh on a rail (§4.1; the board's 10 mW step 1.6–2.6 mJ); events are priced by repeating them ≥10⁹ times a second; a bit flip is ~10⁻¹⁶ J (§2)
Individual bit flips
simulation only
every flop every cycle in the open RTL, every architectural bit in the functional simulator; nothing on the die counts toggles (§3)
Start here: four reading paths
  1. A. Why is this chip low power, and what does a FLOP cost? For a newcomer, about 45 minutes.
    1. Why is the ET-SoC-1 low power? the equation measured term by term, against an A100
    2. The Horace experiment, §1, §3, §6 why the data, not the FLOPs, sets the watts
    3. The energy manual, §1–§3 the card at rest, awake, per instruction
    4. The DVFS loop and its leakage, §1–§2, §5 the governor, and what leakage costs
  2. B. Writing fast kernels. For a programmer.
    1. Test drive the toolchain, the simulator, a first SGEMM 18 Sep baseline
    2. Matmul efficiency the tensor unit near its peak 18 Sep baseline; energies from three cards
    3. Memory hierarchy latency and bandwidth of every level 18 Sep baseline; energies from the energy manual
    4. Ridge points how much work per byte each level needs
    5. On-chip communication the mesh, messages, trees, and the TensorSend trap 18 Sep baseline; energies from the energy manual
    6. Hand it to the next shire relaying between shires instead of DRAM
    7. One hot line stops a shire contended atomics, and what to do instead
    8. Sparse compute what the tensor unit does with zeros 18 Sep baseline; power from three cards
  3. C. What does each operation cost? For an energy modeller.
    1. The energy manual every cost, with bars from three cards
    2. Heat per millimetre a bit carried across the mesh, by bit pattern
    3. The Horace experiment, §8 from flips to watts to degrees
    4. This page, §4 what the meters do not see
  4. D. How far down can you see? For tools and observability.
    1. This page, §2–§4 every instrument, and the meters' limits
    2. Power and temperature, §2–§3 a load step through every sensor; the voltage map
    3. Anatomy of a memory access one load taken apart to ±3 cycles energies from the energy manual
    4. Spatial temperature: a brief 35 sensors the host never sees one by one
    5. The DVFS loop, §1 the firmware as the gate
The chip in brief: minions, harts, shires, scratchpads, the SP, the PMIC and the other names on this page

The ET-SoC-1 is Esperanto's RISC-V accelerator, now open-sourced by AI Foundry. Its compute cores are minions: small in-order RISC-V cores, each with two hardware threads (harts), an 8-lane fp32 vector unit and a tensor unit whose matrix multiply-accumulate instruction is TensorFMA (one fp32 op multiplies 16×16×16 tiles: 4,096 multiply-adds). Only hart 0 issues tensor operations (hart 1 may only prefetch into the L2 scratchpad). Eight minions form a neighbourhood and 32 a shire. Each shire has 4 MB of SRAM, which these cards' firmware splits into a 512 KB L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad: software-managed memory that any shire can address. Each minion also sets aside 3 KB of its 4 KB L1 data cache as an L1 scratchpad for tensor operands.

The chip has 34 minion shires (1,088 minions): 32 run kernels (1,024 minions), the master shire runs firmware and one is spare. With the I/O and PCIe shires they form a 6×6 grid on a mesh network-on-chip (NoC, 400 MHz); eight memory shires with the LPDDR4X controllers (32 GB) sit along two sides, making 8×6 mesh stops. A hop is one step between neighbouring stops, about 3.72 mm. A flit is the unit the mesh moves as a whole.

The service processor (SP) is the on-die management core: its firmware reads the sensors, runs the clock and voltage governor (600, 700 or 800 MHz) and answers the host's management commands (DM_CMD_*). The PMIC is the board's power-management controller: it meters the 12 V input and, over PMBus, three regulators (the minion, SRAM and NoC rails). Moortec PVT monitors measure temperature and voltage on the die. Kernels run in user mode (U-mode) and firmware in machine mode (M-mode). The four Maxions are larger out-of-order RISC-V cores on their own 0.6 V rail.

Other names on this page: A0 is the silicon revision (stepping) that the errata describe; PShire is the PCIe shire; VDDQ is the DRAM I/O supply; MDI is the Minion Debug Interface that the SP exposes to the host; BL1 and BL2 are the SP's boot stages; OTP is one-time-programmable fuse memory; the PRM is Esperanto's Programmer's Reference Manual; sys_emu is the functional simulator; ettelem is this project's telemetry client. aifoundry2 and aifoundry3 (a2, a3) are lab machines with one card each, and aifoundry1 holds two: its card 1 (aifoundry1-c1, a1c1, on older firmware, 1.2.0 against the others' 1.3.1) is the third card measured, since 25 September; its card 0 overheats under load and is not used. aifoundry3's card is held at 600 MHz by a power limit of 0 W that a lab service sets at every boot. The Horace experiment is named after Horace He's post showing that GPU matrix multiplies run faster on predictable data (“Strangely, Matrix Multiplications on GPUs Run Faster When Given ‘Predictable’ Data!”).

1. The measurement reports

Reports and sessions, 18–29 September

Each session (an E-ID of the experiment register, §7) sits in the lane of the card it ran on, and each report at the day it was first published. A session on several cards sits in each of their lanes. Grey curves run from a session to the reports it fed; dashed arrows from a number's first page to the page that now holds it. Hover, tap or tab to any mark; Enter or a click selects it and shows its details.

How each page's claims stand on the cards

Every claim the version-3 claims check listed, one bar per page: the upper bar is its status after the version-3 campaign, which is complete (25–26 September: each claim it tested takes its outcome on aifoundry2, aifoundry3 and aifoundry1-c1), the lower its status before it (the plan's verdicts, 25 September). Each bar is split by verdict (a number quoted from another page takes that page's verdict), or by the cards whose data the claim rests on. A page's own note (“Checked on three cards”) gives both counts of the claims the campaign tested there: judged on the numbers the page quotes, and by this chart's rule, which counts a tested claim under “a test behind it failed” when any part of any registered test covering it failed, even a part the page does not quote, and under “fewer than three repeats” when a card lacked them; a tested claim the campaign only reported per card keeps its earlier verdict. So a page's “held” can be the chart's “failed”: all 17 of the hot line's tested claims rest on its registered test LAT-H, which failed in parts the page does not quote (the shares, the cycles per atomic, the pollers), and the relay's rest on LAT-R, which failed on two cards and was incomplete on the third. Hover, tap or tab to a bar for its counts, and for both readings of its tested claims; a click or Enter opens the page (on a touch screen, a second tap).

What each verdict means
The full index as a table

2. The ladder

Every instrument, roughly coarsest to finest: what "status" and the check column mean

Every instrument, roughly coarsest to finest. Status says what it takes on this card: works now from a user account; tooling means the hardware and firmware are there but a client, a build or an external box must be written or installed; firmware means a small source change plus a reflash of a signed image (§5); research means possible in principle with things not in the open drop; not on silicon is the hard wall. The check column shows whether two reviewing AI agents confirmed the row's citations (✓), corrected them (±), or did not get to it (·); rows added after 20 September are unreviewed (·) by that survey (§8), and rows rewritten after their review are marked †. §8 says how the review was done.

Time resolution of each instrument

Each dot is the finest time step the instrument resolves on the card; its colour is what it takes to use it. Hover, tap or tab to a row for its granularity; Enter opens its row in the table. The buttons under the chart filter the chart and the table together.

Every instrument, as a table
How many identical events before the meter sees one?

Every event the reports priced, on one scale of joules, against what one meter reading resolves. An event becomes visible when the card can run enough of them a second to lift the power by the precision asked for above the uncertainty of the baseline: it needs a rate of σ ÷ (precision × energy). Filled dots are events the measurements ran fast enough to see at these settings; hollow ones are not: at the rate measured, a burst of them moves the power by less than σ ÷ precision. The tick on each row is the smallest energy per event that the measured rate would reveal.

3. Can we track individual bits flipping?

On the card: no, and not with any firmware change. Nothing on the die counts toggles. The PMU counts operations (up to 8 per cycle per neighbourhood), the shire-cache and memory-controller monitors count requests, and the power path resolves at best 1 mW over one service-processor pass (§4.1), twelve orders of magnitude above one node toggle. The closest things to a bit read are the ECC machinery and the debug module, and both stop at architectural state:

ECC, the debug interface, raw SRAM, UltraSoC, and physical probing, in detail

Nor did the card report a wild write, the one time we made one: on 24 September, during the wire-energy work, a test-tool bug sent stray tensor stores to physical address 0 in two test launches of about half a second each (which card was not recorded), and no error counter, telemetry field or clock moved.

A first result from the RTL: the hpmcounter3 carry bug

A first result from the RTL. On every card measured (aifoundry2, and since 26 September aifoundry3 and aifoundry1-c1, five launches of each test on each), hpmcounter3 reads 128 short just after its low 7 bits wrap: two reads 10 cycles apart differ by 10, 138 or −118 cycles and nothing else. Simulating the open RTL's neigh_pmu.v under Verilator (rtl-sim/pmu_carry) shows why: each counter is a 7-bit pre-counter plus a 57-bit post-counter, twelve counters share one adder that folds pre-counter overflows into the post-counters in round-robin order, and a read ignores the pending overflow bit. So after every wrap the value is 128 short until the adder comes round again: 12 cycles in the simulation. On the cards the short window is not a fixed length. Where one window fits a whole launch it is consistent with low bits 0 to 10 (six launches on aifoundry2, one on aifoundry1-c1), but in 26 of the 30 launches of paired raw reads no single window fits them all, and the usual fix (fixcyc(), which adds 128 to a read whose low bits are below 11) left 0–4.7% of 10-cycle intervals off by 128 in a launch: none in four launches of five on aifoundry2 and 2.1% in the fifth, 1.7–4.7% in all five on aifoundry3, up to 0.8% on aifoundry1-c1. The mechanism is a silicon measurement explained by a flip-flop-level view of the design, which is the point of having the RTL; the window's length on the card is not yet explained.

In the RTL: one counter, one wrap, and the shared adder that folds its carry in
On the cards: read pairs that come out 128 off, raw and after the fix

In simulation: yes, twice over

In simulation: yes, twice over. sys_emu -l prints every executed instruction on any selected hart with each architectural write it makes: integer registers, all 8 lanes of a vector register, mask registers, every memory read and write up to 512 bits, CSR writes, traps. XOR successive values and you have the bits that changed. It is functional, not cycle-accurate, and it models no cache contents. The RTL bench goes further: +dump=1 writes every net, flop and behavioural SRAM bit of the CPU subsystem to an FST waveform each cycle, and the design verification monitors emit a cycle-stamped event per retired instruction that is checked in lockstep against the same emulator. Toggle counts per signal are a coverage flag. What you cannot get from the drop is joules: there is no cell library, netlist or power flow, so toggles stay counts.

Joules have been attached to flips in two places, both by fitting to measured power: the tensor unit, with four event energies (the Horace experiment; its flip counts come from rtl-sim/fma_toggle, the core's multiply-add RTL under Verilator), and the mesh links, where a bit that differs from the previous flit costs about 98 fJ per hop and a one carried about 129 fJ on the mesh rail (Heat per millimetre, §5). The chart at the end of §2 sets both against the meter.

4. Power and energy: the meter chain, and what it does not meter

4.1 The chain

How the PMIC and the service processor meter the three rails, and how stale a reading can be

The card's power is read by a PMIC on the board — a microcontroller with its own closed firmware — that measures the 12 V input and, through PMBus, the three digital regulators that feed the minion cores, the SRAM arrays and the mesh. For each of those three it holds voltage, current and power on both sides of the regulator, and its temperature, each as a current value, a minimum and maximum since the last reset, and a running average: 84 numbers.

The service processor reads all 84 every loop pass and forwards to the host the PMIC's running average of each output-side power, with its min and max, and the 12 V input power; ettelem samples them at 10 Hz. The average is roughly first-order: after the board steps down at the end of a burst, the rail reading has fallen , which is why the rails take 2–3 s to catch up with the load step in Power and temperature, §2. That τ is one fit over the catalogue's bursts and absorbs the service processor's one-pass delay; fitted rail by rail with the delay kept apart (E58, 29 September), τ is 1.01–1.06 s on aifoundry3's rails, 1.08 s on aifoundry1-c1's minion and NoC rails and 0.54 s on its SRAM rail, and tools/ettelem/deconv.py undoes the filter (rung 4). Resetting the PMIC's minimum and maximum (ettelem --reset-ms) restarts that average as well: with a reset every second, the rail reading had fallen 93% of the way one second after a burst, on all three cards (the median of seven to nine bursts on each), instead of the fraction above. The SP's copy refreshes once per pass of its loop. While ettelem samples at 10 Hz, as in every session from 20 September on, a new board value arrives about every 156 ms on aifoundry2, 158 ms on aifoundry1-c1 and 263 ms on aifoundry3 (phase-folding the board-power stream, the mean of three passes on each card in the version-3 campaign, TEL-S; in the earlier power sessions the readings changed every 148–164 ms on aifoundry2, over 29 sessions, and 251–260 ms on aifoundry3, over 16). The SP's own trace times the pass directly: 133 ms on aifoundry2, 135 ms on aifoundry1-c1 and 224 ms on aifoundry3 with no sampler running, and about 160, 162 and 266 ms while ettelem samples (the medians of three passes on each card, TEL-P1, P3 and P5). So the sampler lengthens the pass on every card, by 26–42 ms; under the sampler the trace's pass runs 3–5 ms longer than the refresh the board-power stream shows, the difference between the two methods on each card. In the firmware source read here (353f20e, not the cards' older build) the I2C driver waits a millisecond after each transaction, the likely reason a pass takes that long. Why aifoundry3's loop takes about 1.7 times as long as the other two cards' (aifoundry2 runs the same firmware release) is not established. How often the PMIC itself updates is unknown. The other rails — — have set-point registers in the PMIC and no current or power telemetry (their on-die voltages are reported, §4.3).

The workload can starve the meter

The workload can starve the meter. Rings of messages between shires s and s+16 slow the service processor's management path: a telemetry sample takes instead of 22, and a burst measured through it reads low (33–45% on aifoundry2 on 23 September). So do tensor loads between shires three hops apart in one column, on aifoundry2 (the heat-per-millimetre runs). The sampler records its own latency with every sample, and bursts with that signature are dropped (rung 7). A sampler killed in the middle of a request stops the meter outright: its reply stays in the management queue, and every later opener crashes on it until dev_mngt_service drains it (seen once on each card, E32).

The meter starved by the workload

Every catalogue burst's own sampler latency (the time one telemetry sample took), one strip per card; a dot's size is how many bursts took that long. Diamonds are rerun passes dropped from the manual because the sampler was this slow. Hover, tap or tab to a mark.

So the meter chain gives one number for the whole card and three for its inside, and the difference between them is the subject of this section. At idle it is : the first bar of the chart in §4.2. The version-3 campaign did not confirm the aifoundry2 figure: its idle cycles on that card passed through 73 °C only twice, one short of the three repeats the check asks for (15.05 W of 31.73 W, 99% interval 14.56–15.54 W); on aifoundry1-c1 three cycles gave 17.87 W of 42.71 W at 73 °C.

The meter chain

From the 12 V input to a number in this page's charts. The grey box at the bottom carries no current or power telemetry at all. Hover, tap or tab to a box for what it does and its numbers on the selected card.

How often a new reading arrives, card by card

4.2 The unmetered remainder, attributed

The remainder cannot be metered with anything on the card, but it can be attributed. Over the energy manual's catalogue — , three passes on each of three cards — each configuration's power above idle is known on the board and on the three rails, and the bytes it moved through DRAM are counted. Fitting the unmetered part of each configuration's mean as a fraction of each rail's power plus a cost per DRAM byte gives:

The fit's coefficients, per card

Each coefficient's estimate, with one robust (HC3) standard error either side; one whisker per card in every row, top to bottom in the order of the table's columns. Hover, tap or tab to a whisker.

How the fit was done

Where a workload's watts go

Three things the fit implies

4.3 What the Moortec sensors can and cannot do for power

The die's Moortec PVT subsystem, and what of it reaches the host, is in the table below. The spatial temperature brief maps the sensors onto the 6×6 grid of minion, IO and PCIe shires (8×6 counting the memory shires at the sides) and traces how the firmware collapses them to one average.

None of it measures current, so none of it meters the unmetered rails. What the monitors do measure is voltage, and voltage droops with current. The host already receives, every SP pass (§4.1), the memory shires' reading of the 0.8 V DDR rail (DM_CMD_GET_ASIC_VOLTAGE, ettelem's die_mv.ddr): across the catalogue it reads at idle against an 800 mV set point, and under load it droops roughly in proportion to the DRAM power the regression attributes off-rail:

Is the DDR monitor a DRAM meter?

Every catalogue configuration, in the colours of the chart in §4.2; aifoundry2 by default (the calibration above), or pick a card. Hover, tap or tab to a point for its droop, the droop the calibration predicts, and the DRAM power the excess would be read as.

How the droop calibration was done

The eight configurations of the first table, to 0.1 mV
A current meter for DRAM power, calibrated once per card — and its limits

That is a free current meter for DRAM, the largest unmetered consumer under load, in every ettelem log, updating every SP pass with . Its scale comes from the fit above, so it is not independent of the board meter, and it responds mostly to DRAM traffic, but not only: So in bursts of one kind it separates DRAM from arithmetic, not from mesh traffic; a mixed burst is untested. .

The same monitors see the minion rail sag by , which Power and temperature (§3) mapped shire by shire at idle from the SP's debug trace (one idle capture on aifoundry2; the three-card check captured all 34 shires on each card but got too few complete idle captures on aifoundry2 and aifoundry1-c1 to confirm the map); calibrated shire by shire, that map is a spatial power meter for the three metered rails (rung 5). The temperature sensors and the process detectors, exported, would add heat maps and a leakage proxy per shire; they too need a firmware change.

5. The improvement ladder

Every way found to see more, grouped by what it costs: software on the data the card already gives; lab hardware that needs no firmware change; firmware changes, each ending in a signed image; tooling against interfaces that already exist (a debugger client, an RTL flow); and research. Two more groups collect what the chip diagram still infers (its dashed parts, the routes it draws, the host's path to DRAM: the last two measured since 29 September, rungs 32 and 34), what the PCIe page left open (why two host-to-card copies at once halve the rate, one stream's since 29 September, rung 35; the link's payload size) and what the memory levels, the interactive diagrams of each level of the memory system down to the transistors, leave unknown (which cell each memory uses, since the lab lead says the chip is not using SRAM; the memory macros' geometry; where each cycle of a cache access goes; the live cache settings; the DRAM part), and what the effect of overheating could not settle (what stopped a card near 120 °C; what the PMIC's temperature measures; the chip's thermal ratings): the documents, interfaces and readings to ask AI Foundry (Ainekko) and the lab for, each rung ask the team, and the experiments on the cards that would settle the same parts without them. Rungs 21–36 are the chip diagram's and the PCIe page's; rungs 37–44 are the memory levels', whose other unknowns are added to rungs 20, 21, 23, 25, 26, 28 and 32, which already asked the same thing (rungs 19 and 33 gain a note). Rungs 45 and 46 are the overheating page's; its other asks extend rungs 13 (a firmware build for the process detectors), 24 (the thermal ratings) and 42 (the DRAM's temperature grade). Each of those rungs names the part of the diagram it settles and links to it, and each diagram links its unknown parts to their rungs. Their firmware line numbers are at et-platform 836a4ab, as in the diagram. Each rung says what it adds and what it would change in the numbers; where a rung is an instrument of §2, the link opens that row. Rungs marked done say where they were done. Rungs 4, 31–36 and 43 were run on the cards on the night of 28–29 September (E55–E58 in §7), each under predictions frozen after development on aifoundry1-c1 and tested on aifoundry3; rungs 31 and 43 are not marked done: the first was not decided on the validation card, and no registered theory of the second survived. The firmware caveat at the end of the section applies to every rung in its group.

For the Nekko team · errors to fixRequests for Roman: what breaks on the lab machines and cards →The rungs below ask for documents, interfaces and firmware changes that would let us see more. The lab problems report asks for fixes to what goes wrong: 41 requests in 7 groups (cards, driver and firmware; runtime and toolchain; machines and operations; access and security; site and hardware; documents; policies), each group sorted by importance, from 84 problems re-checked on 27 September. aifoundry2's hung card was recovered on 28 September, and whatever can be fixed on the machines themselves is now on our own list, so these requests keep only what needs the Nekko team: firmware, driver and runtime, site and hardware, the tailnet, and lab policy. Of our own list, the owner-approved lab fixes on aifoundry1 were made on 28 September (logged on the host); the rest, the same fixes on aifoundry2 and aifoundry3 among them, waits for the owner.
What each rung does to the unmetered watts, and the firmware caveat

What each rung does to the unmetered watts. Today: ; rungs 2 and 3 attribute the part above idle (§4.2, §4.3) and leave the idle W unsplit. The riser (rung 8) does not meter them either, but it makes every bracketed burst sharper and shows the input side of every regulator's transient. The camera (rung 9) says which package the watts warm. Forwarding the PMIC's input-side readings (rung 11) measures the delivery losses and leaves only the rails with no meter — DDR, VDDQ, PCIe, IO, Maxion — as a remainder: . Below that is the board, and the board is shared.

The firmware caveat. Every "firmware" row above ends with a reflash. The minion runtime and the service processor bootloader are signed executable images; BL1 and BL2 verify a certificate and signature against a public-key hash in OTP unless a chicken bit is set, and neither the signing tool nor a test key is in the open tree. Whether this card accepts a rebuilt image is unknown until someone tries, and a bad image can brick the boot path, so it is a lab-admin decision. This is the one place where "everything is open" has a gap: the source is open, the right to run modified source on this particular card is not yet established.

6. Against a GPU

For an A100/H100-class part. Sources are NVIDIA's own documentation and the microbenchmark literature, listed below the table; the ET side is this survey.

The pattern: the GPU is at least as observable on silicon (ECC counts, a debugger, SASS-level instrumentation, a mJ energy counter the ET card lacks), and entirely opaque below it. The ET-SoC-1 is thinner on silicon (no L1 or register ECC, DRAM ECC off, power once per service-processor pass) and open all the way down: the firmware that decides what you can see is source, the reference model the RTL was verified against is the installed simulator, and the RTL of the core dumps every bit.

Sources for the GPU column
  1. Jia, Maggioni, Staiger, Scarpazza, “Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking”, arXiv 1804.06826, 2018.
  2. Luo et al., “Benchmarking and Dissecting the Nvidia Hopper GPU Architecture”, arXiv 2402.13499, 2024.
  3. Yang, Adamek, Armour, “Accurate and Convenient Energy Measurements for GPUs”, SC'24 (preprint: “Part-time Power Measurements: nvidia-smi's Lack of Attention”, arXiv 2312.02741).
  4. Lin et al., “GPUHammer: Rowhammer Attacks on GPU Memories are Practical”, USENIX Security 2025.
  5. Khairy, Shen, Aamodt, Rogers, “Accel-Sim: An Extensible Simulation Framework for Validated GPU Modeling”, ISCA 2020.
  6. NVIDIA developer forum, “Questions about globaltimer functionality, accessing and configuring” (the %globaltimer update rate).
  7. NVIDIA documentation: CUDA Binary Utilities, PTX ISA, the CUDA Programming Guide (clock64), the Nsight Compute Profiling Guide, the NVML device queries, GPU Memory Error Management r575, the CUDA-GDB manual and the open-gpu-kernel-modules README.
  8. This project's notes on GPU power telemetry: docs/research/power-telemetry.md.

7. The power and temperature sessions

Every registered session (E1 onward) that measured power or temperature, and the experiments run for §5's rungs (E54–E58), in order, with what it established, the instruments it used, where its raw data is in the repository (paths are under docs/reports/data/ and link to GitHub) and the reports that use it; two sessions of 18 September were registered on 25 September (E33, E34), and that day's matmul and sparsity sessions, which predate the register, share one row. The map in §1 shows them on a timeline, by card. The IDs are those of the experiment register, docs/findings/03-experiments.md, which gives each session's command, protocol and caveats.

All sessions as a table

8. Method and caveats

Method, versions, and caveats in full

Every report is in §1; the knowledge base behind them is docs/findings/, and the code and raw data are in the et-soc1-prototyping repository.