On-chip communication on the ET-SoC-1: 20 ns per mesh hop, and a bandwidth cliff at the shire boundary
Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 and aifoundry1's card 1. Corrected where a few cycles moved: the in-shire barrier is 233 cycles, not 237; the 32-minion allreduce 444, not 432; the credit round trip rises 24 ns per hop, not 20; also the tree levels and the memory flag's range. The energy per mesh hop held with no difference between the cards (the energies here are now that check's, six passes on each card). The three cards agree within a few cycles (the chip-wide barriers within 15), and the map and the distance chart show each card. Of 29 claims tested here, this page counts 17 held, 8 corrected, 2 differ by card and 2 held with no difference between the cards; the hub’s scoreboard, 3 “proven on the cards”, 25 “a test behind it failed” and 1 “differs by card”. Record: docs/reports/data/2026-09-25-claims-v3.
The ET-SoC-1's cores can pass data to each other directly, without going through memory, and quickly. Inside a shire a message takes 113–190 ns round trip, and each mesh hop adds 20 ns. A tree allreduce over all 1,024 minions takes 2.3 µs. Messaging costs 0.7–2.2 pJ per byte inside a shire with 1 KB messages, busy cores included (the energy manual's re-run on three cards at 600 MHz, 26 September). Crossing the mesh has a cost, though: there is 7–34× less message bandwidth between shires than inside them.
An A100's SMs can reach each other only through L2; Hopper adds direct shared-memory access, but only within a cluster of up to 16 SMs. The ET-SoC-1 has several direct paths. Hart 0 of any minion can send up to 127 × 32 B (about 4 KB, cycling through its 32 vector registers) straight into another minion's registers with TensorSend/TensorRecv, and the receiver can add, max or min the data into what it already holds. There are hardware reduction and broadcast trees, credit counters that let a core sleep until another core signals it, and barrier counters in every shire. This page measures each of these against its closest GPU equivalent and maps where the 32 shires sit on the chip's 2D mesh.
Terms used on this page
The ET-SoC-1's compute cores are minions: small in-order RISC-V cores, each with two hardware threads (harts) and 32 vector registers of 32 B, of which only hart 0 issues tensor operations such as TensorSend; minions are written shire.minion, so 0.0 → 0.1 is minion 0 of shire 0 sending to minion 1. Eight minions form a neighbourhood and 32 a shire, whose 4 MB of SRAM holds a 512 KB L2, a 1 MB slice of the 32 MB L3 and a 2.5 MB scratchpad that any shire can address; the 32 compute shires, with the master, spare, I/O and PCIe shires, form a 6×6 grid on a mesh network-on-chip (NoC, 400 MHz), eight memory shires sit along two sides, and a hop is one step between neighbouring stops. A credit counter (FCC, fast credit counter) lets a minion sleep until another core adds a credit to it by a store to its shire's CREDINC register, and the service processor (SP) is the on-die management core that reads the board's power meter. More in the hub's glossary.
Where the shires are: a 6×6 mesh
Each square is one position on the mesh in marty1885's logical map (which appears to be the die turned a quarter), labelled with the logical shire ID the firmware uses. Colour is the measured TensorSend round-trip time from the selected shire (outlined), between minion 0 of each shire; stronger blue is slower. Card picks whose measurements: each of the three cards of the 26 September check (the median of its three passes) or the first session (aifoundry2, 18 September). Hover or focus a shire for its numbers; tap it, or press Enter on it, to measure from there. The four grey squares hold no compute shire: they are the master shire, the spare shire, and the PCIe and I/O shires, and they route traffic too. The datasheet's full 8 × 6 mesh adds eight memory shires, not shown: in this map's orientation the latency fit on aifoundry2 in Anatomy of a memory access (§4) places memory shires 0–3 above the top row and 4–7 below the bottom row, over the middle four columns (memory shire 2's own position is not pinned down by that fit).
What "Shade by" does
Shade by recolours the map with another of the page's shire-to-shire round trips, each on its own scale: TensorSend with 1 KB messages, the credit counter, or a flag passed through memory. The readout under the map gives that matrix's straight-line fit against mesh hops over all 496 shire pairs.
Compare three chain orderings on the map
A chain hands data from one shire to the next in some order, all the way round; the orange path is that order, and the darker dashed segment its longest step. Compare three orderings, on the same 32 compute shires the relay report chains:
How the map was checked
Coordinates follow marty1885's map (x across, y down). On each of the three cards every round trip between two shires is 150 + 12.02 × (Manhattan distance on this grid) cycles, with a worst residual of 1.1–1.4 cycles over all 496 pairs in each of the nine passes (1.1 in the first session). In every pass on every card, all 12 restarts of a search from random layouts, without his map, found the same distances.
Latency grows with distance, except through memory
Round-trip time between minion 0 of every pair of shires, against the number of mesh hops between them, for three ways to signal another core. TensorSend round trips rise by 20 ns per hop. Credit round trips rise by about 24 ns per hop on all three cards (234 + 24 ns per hop fitted, with a worst residual of 15 cycles; the first session's 246 + 20.4 ns fitted within 5), yet across shires they almost coincide with TensorSend's: the median pair differs by under a cycle, and 445 of the 496 by at most 15 cycles. A flag handoff through global atomics in memory, which is how a GPU passes a signal between SMs, lands on whichever shire's L3 slice holds the flag's line. It costs 610–1,150 ns (median 873 across shires on each card) whatever the distance: its slope against mesh hops is −2 cycles per hop, and at most +9 cycles (15 ns) per hop at 99% on each of the three cards. Points at 0 hops are pairs inside one shire.
Inside a shire, only the tree edges are fast
Each shire has 32 minions in four neighbourhoods of eight. Measuring all 496 minion pairs in shires 0 and 24 gives exactly two speeds, the same pairs and the same cycles in every pass on all three cards:
- 68 cycles (113 ns) for 28 pairs per shire: in every neighbourhood, 0–1, 0–2, 0–4, 2–3, 4–5, 4–6 and 6–7.
- 114–115 cycles (190 ns) for all 468 others. Being in the same neighbourhood doesn't help.
Those seven edges are the reduction tree, and they are the only links of the neighbourhood's fast local messaging network (core-et Neighborhood MAS §4.7), which delivers a 256-bit message in one cycle. Every other message leaves the neighbourhood and comes back through the shire crossbar.
Every primitive, against a GPU
Round trips at 600 MHz on ET-SoC-1; the three cards agree within a few cycles, the chip-wide barriers within 15 (26 September). GPU numbers are published measurements, at 1.41 GHz for the A100 and 1.755 GHz for the H800 (Hopper). The GPU column lists the closest equivalent where one exists.
Drawn on one time axis, the ET-SoC-1's messages between cores land where an A100's handoff through L2 does, while each of its barriers is slower than the GPU's nearest equivalent:
The full comparison table, primitive by primitive
| Primitive | ET-SoC-1, measured | What it does | Closest GPU equivalent |
|---|---|---|---|
| TensorSend/Recv, register to register | 113 / 190 ns 250 + 20/hop ns | Round trip of 32 B between two minions' vector registers: fast-network pair / rest of the shire / other shire. No memory is touched. Per 32 B register add 2.3–4.6 cycles each way. | A100: no SM-to-SM path. A handoff through an L2 atomic is 166 ns one way (~330 ns round trip). L2 hits are 148 ns (near) and 253 ns (far). H800 (Hopper): a remote read of another SM's shared memory takes 103–121 ns, but only inside a cluster of up to 16 SMs. |
| Combine on arrival | +0 cycles | The receiver adds (fp32 or int32), maxes or mins the incoming data into its registers. The round trip is identical to a plain move for 32 B and 1 KB messages. | Atomics executed in L2 (red.global); on H100, atomics on another SM's shared memory. |
| TensorReduce + TensorBroadcast (allreduce) | 0.74 µs · 2.3 µs | Hardware binary tree over the 32 minions of a shire, or over all 1,024 minions, 32 B each. All 32 shires can run their own trees at once without slowing each other. | No tree hardware. A grid-wide grid.sync() alone, with no data, costs 1.1–1.2 µs on A100. Device-wide reductions are built from atomics or extra kernel launches (1.5–2.3 µs each). |
| Credit counters (FCC) | 200 ns 234 + 24/hop ns | Round trip: one store adds a credit to a core in any shire, and that core sleeps on a CSR until it arrives. Measured with the blocking wait inside a shire, and polled across shires. | None. A consumer spins on a flag in global memory. A __threadfence() alone costs ~540 ns on A100. |
| Fast local barrier (FLB) + credits | 388 ns | Barrier of 32 minions in a shire, with all 32 shires doing it at once. | __syncthreads(), within one SM: 16–60 ns. cluster.sync() across 2 SMs on H100: 851 cycles. |
| Chip-wide barrier | 8.3 µs · 2.3 µs | Barrier of 32 shires built from FLB, global atomics and credits (8.3 µs), or from the hardware allreduce tree (2.3 µs). | grid.sync(): 1.1–1.2 µs on A100. |
| Flag through memory atomics | 610–1,150 ns (median 873 across shires) | Round trip: global atomic swap, then spinning on a global atomic read, on DRAM lines that live in L3. No measurable dependence on distance (at most 15 ns per hop). | This is the A100 pattern: ~330 ns round trip through L2, ~0.7 µs once a data fence is added. |
Message size and bandwidth
Time per message for a one-way stream of messages from one minion to another, against message size. The receiver posts a new receive as soon as the last one finishes. A message is up to 127 registers of 32 B, reusing registers once past f31. Each path has a fixed cost per message (40, 86, 135 and 224 cycles) plus a cost per register (2.3, 4.0, 3.0 and 4.6 cycles), on all three cards. The chart is aifoundry2's (26 September); the other two cards agree within 0.3%. 1 KB round trips between all 496 shire pairs add 12 cycles per hop up to 5 hops and 27–36 per hop beyond (the 1 KB view of the distance chart), so the sender seems to have a limited number of packets in flight.
Per-link bandwidth, as a table
| Path | One link, 1 KB messages | One link, 4 KB messages |
|---|---|---|
| Fast local network | 5.3 GB/s | 7.2 GB/s |
| Shire crossbar | 2.9 GB/s | 4.1 GB/s |
| Mesh, 1 hop | 2.7 GB/s | 4.7 GB/s |
| Mesh, 10 hops | 1.8 GB/s | 2.9 GB/s |
Per-link numbers are 32 B × COUNT over the stream's time per message at 600 MHz; the three cards give the same to the digits shown. With every minion sending 1 KB messages at once, the whole chip moves 1.1–3.0 TB/s inside shires and 0.09–0.16 TB/s between them; those rates and the energy per byte of each ring are in Energy per byte moved.
Reduction trees
Latency of one allreduce (TensorReduce up the tree, then TensorBroadcast back down) against the number of minions in the tree, for three sizes. Up to 32 minions the tree stays in one shire. Beyond that, each doubling adds a level between shires. Each level adds about one round trip between the minions it pairs: 68–71 cycles on the fast network (levels 0–2), 117 through the crossbar (levels 3–4), and 164–234 over the mesh (levels 5–9), where the slowest branch sets the pace (the first session's were 68, 114 and 156–235). The chart is aifoundry2's (26 September); the other two cards agree within 0.3%. The dashed line is the A100's grid.sync(), a barrier that moves no data.
Energy per byte moved
Why does a byte cost 6–26× more once it crosses the shire boundary? Hart 0 of all 1,024 minions sends 1 KB messages (128 B where marked) around rings inside a shire, or to shires a fixed number of IDs away; the rows are sorted by mean mesh distance. The three panels share the rows: bytes moved per second over the whole chip, board power above the idle just before and after each burst, and energy per byte. The last two are the energy manual's re-run of these rings (§5): each dot is the mean of six passes on each of three cards at 600 MHz (26 September), with the range the passes spanned. First measured on 18 September on aifoundry2 alone (two runs, sampled without the die temperature and so with no leakage correction, on a card cooling after another user's matmul): 0.8–2.3 pJ per byte inside a shire and 13–20 across the mesh, 2–13% above these means (4–16% above aifoundry2's own passes). Inside a shire messaging costs 0.7–2.2 pJ per byte with 1 KB messages. With 128 B messages a shire ring costs more: 3.7 against 2.2 pJ/B (3.6 against 2.1 on aifoundry2, 3.4 against 2.1 on aifoundry3, 3.9 against 2.5 on aifoundry1's card 1), a difference of +1.3 to +1.5 pJ/B, above zero at 99% on every card. The readout under the chart gives the cost's fit against mean hops across the mesh, per card and together. Grey rules are references: ET's own L2, remote-scratchpad (the mean over the zeros and random data its passes held) and DRAM reads at a steady 600 MHz (the energy manual, §4), and the A100's L2 and HBM accesses (SC'25). The last row is bulk data for comparison: every minion streaming TensorLoads from the scratchpad 16 shire IDs away (5.1 pJ/B when that scratchpad holds zeros, 11.8 when it holds random data).
The numbers in the chart
| Messages | Mean hops | Whole chip | W over idle [range] | pJ/B [range], passes |
|---|---|---|---|---|
| Pairs on tree edges | 0.0 | 2,992 GB/s | 2.05 [1.81–2.24] | 0.69 [0.61–0.76], 18 |
| Rings of 8 (a neighbourhood) | 0.0 | 1,120 GB/s | 2.40 [1.95–2.74] | 2.16 [1.76–2.47], 18 |
| Rings of 32 (a shire) | 0.0 | 1,120 GB/s | 2.47 [2.03–2.85] | 2.22 [1.83–2.57], 18 |
| Rings of 32, 128 B messages | 0.0 | 553 GB/s | 1.99 [1.57–2.31] | 3.66 [2.89–4.24], 18 |
| Rings of 4 shires, 8 IDs apart | 1.6 | 150 GB/s | 1.85 [1.35–2.20] | 12.50 [9.10–14.84], 18 |
| Shires s and s+16 | 2.1 | 156 GB/s | 2.03 [1.86–2.12] | 13.27 [12.20–13.81], 3 (23 Sep, aifoundry3) |
| To the next shire ID | 3.5 | 124 GB/s | 1.77 [1.44–2.02] | 14.41 [11.72–16.47], 18 |
| To the next shire ID, 128 B messages | 3.5 | 112 GB/s | 2.25 [1.89–2.65] | 20.20 [16.99–23.82], 18 |
| Rings of 8 shires, 4 IDs apart | 3.7 | 91 GB/s | 1.37 [1.10–1.59] | 15.13 [12.22–17.55], 18 |
| Rings of 16 shires, 2 IDs apart | 4.5 | 87 GB/s | 1.47 [1.16–1.85] | 16.93 [13.33–21.29], 18 |
| Rings of 16 shires, 6 IDs apart | 4.7 | 89 GB/s | 1.62 [1.30–1.79] | 18.29 [14.73–20.23], 18 |
| For comparison: TensorLoad from the scratchpad 16 shire IDs away | 2.1 | 959 GB/s | – | 8.47 [4.40–14.43], 18 |
Watts over idle and pJ/B: the mean over the passes, in brackets the range they spanned; the whole-chip rates are the first session's timed launches (18 September), as in the energy manual. The same cores spinning with no messages drew 2.4–2.5 W over idle on each of the three cards (means of six passes; 2.45 W together, the dashed line in the middle panel). The s+16 row keeps aifoundry3's three passes of 23 September: the check dropped every s+16 burst on every card, because that ring slows the service processor that reads the meter (hub, §4.1), stretching the sampler to 63–141 ms a reading. The TensorLoad row pools nine passes whose scratchpads held zeros (5.10 [4.40–6.45] pJ/B) and nine that held random data (11.83 [9.53–14.43]).
Why the shire-boundary step is bandwidth, not wires
Read these as the cost of messaging, not of wires. Rings inside a neighbourhood or a shire with 1 KB messages drew about as much power over idle as the same cores spinning with no messages (from 0.3 W more to 0.2 W less, the means of six passes on each of three cards, 26 September), and 1 KB rings across the mesh drew 0.4–1.2 W less, so these figures are mostly the cost of 1,024 cores kept busy by messaging, divided by the bytes they moved. The step at the shire boundary mostly mirrors the drop in bandwidth, not costlier wires. The per-hop slope, 1.75 pJ/B on the three cards together, falls inside the mesh's own cost measured with chosen bit patterns in Heat per millimetre (1.4–2.3 pJ per byte per hop on board power, from free links to a loaded mesh); the 8–10 pJ intercept is mostly the busy cores.
What the numbers say
- All three cards have marty1885's shire layout. He built the map from shire-to-shire bandwidth on another card. On aifoundry2, aifoundry3 and aifoundry1's card 1, TensorSend round trips fit a+b×(Manhattan distance on his map) within 1.5 cycles and scratchpad loads at a steady clock within 0.2, and credit round trips rise with it too, within 15 cycles (5 in the first session). Pure Manhattan distance means the routes are shortest paths through all 36 cells, the four non-compute ones included. Minion 31 of each shire gives the same numbers as minion 0, so each shire meets the mesh at a single point.
- The mesh costs 20 ns per hop, in both directions together. That is 12 minion cycles at 600 MHz, or 4 mesh cycles each way at 400 MHz. In the one memory-hierarchy session, where aifoundry2's governor moved the clock mid-row, a hop still took 20 ns: remote scratchpad loads cost 66 minion cycles + (56 ns + 20 ns × hops), and a model fitted to two of the four rows fits the other two within a cycle at 600, 700 or 800 MHz, except one chase whose clock changed partway (5.4 cycles off). Only aifoundry2 can test this (a boot service holds aifoundry3 at 600 MHz by setting its TDP to 0 W at every boot).
- Three speeds, not four. A message takes 68 cycles on a reduction-tree edge, 114 anywhere else in the shire, and 150 + 12 per hop between shires. Chains and trees should put partners on those edges.
- Combining in the channel is free. A receive that adds or maxes into its registers takes the same time as a plain copy. That is the ⊕ of a (min,+) or (max,+) semiring, applied on arrival without a separate instruction.
- The hardware trees are the fastest way to synchronise the whole chip. A 1,024-minion allreduce of 32 B takes 2.3 µs, against 8.3 µs for a barrier made from global atomics and credits, on each of the three cards. Inside a shire it takes 0.74 µs, and all 32 shires can reduce at once without slowing down. Reductions and barriers belong on the trees.
- The shire boundary is a bandwidth cliff. Messages move 1.1–3.0 TB/s in aggregate inside shires and 0.09–0.16 TB/s between them, at 6–26× the energy per byte, mostly because the same ~2 W of busy cores moves 7–34× fewer bytes. Across shires, bulk data moves better as TensorLoads from the other shire's scratchpad (0.96 TB/s, at 5.1 pJ/B when that scratchpad holds zeros and 11.8 when it holds random data) than as messages; per byte it cost 3–5 pJ less than the s → s+8 ring on average on all three cards, but not in every pass (two of six on aifoundry3 and on aifoundry1's card 1): 5.8–8.9 pJ less in the passes that read zeros, −2.2 to +3.4 in those that read random data.
- Latency in nanoseconds is GPU-like; the mechanisms and the energy are not. A message crossing 1–10 hops takes 135–225 ns one way. An A100 takes 166 ns to hand a flag from one SM to another through L2. The differences lie elsewhere. The data lands in the receiver's registers, with an optional combine. There are 1,024 addressable endpoints instead of 108 SMs, plus hardware trees and sleeping credit waits. Messaging costs 0.7–18 pJ per byte (1 KB messages), busy cores included, against about 38 pJ per byte for an A100 L2 access.
A trap: one ready flag per minion
Before a sender may transmit, TensorSend waits for a "ready" message from its receiver. In the RTL (dcache_reduce.v, register partner_ready_peer), a point-to-point ready is a single bit per minion, not one per partner. If two different minions post their readies to the same sender before it has used them, one of them is lost. That receiver then waits forever, and so does its sender. The tree operations are safe: they keep one bit per tree level. The first cross-shire test here (aifoundry2, 18 September) ran two pairs through minion 0.0 without a barrier between them, and it hung the card; that is the one hang seen, and the test was not repeated. The stalled harts ignored the firmware's abort, and every later launch failed with KernelLaunchCmIfaceMulticastFailed until the chip was reset.
How the trap is avoided in this page's own probe
The simulator (sys_emu) tracks every partner separately, so it never shows the problem. nocbench now refuses any schedule where a minion could have two partners' readies outstanding. Pairs are separated by a chip-wide barrier, and rings only ever receive readies from the next minion. With that rule, TensorSend works across all 496 shire pairs.
Which TensorSend layouts hang? One shire's 32 minions, one neighbourhood per row, and the messages of one step of four layouts. Solid edges are the neighbourhood's tree edges (68-cycle round trips), dashed ones go through the shire crossbar (114 cycles); a minion that would take readies from two partners in the same phase is outlined and marked. Costs are this page's measured round trips.
What it means for systolic and wavefront designs
- 1D chains at minion grain inside a shire, 2D arrays at shire grain over the mesh. A one-way message costs about 34 cycles on a tree edge, 57 cycles elsewhere in the shire, and 75 + 6 cycles per hop between shires. To keep communication under a tenth of the time, a cell needs about 10× that much work per step: roughly 350–600 cycles inside a shire and 800–1,350 between shires one to ten hops apart. That sets the grain k ≈ 10·tmsg/tcell cells per message for wavefront DP or lattice updates.
- In a 2D array, give each neighbour link its own minion. Over TensorSend a minion may have only one partner's ready outstanding (see the trap above). At shire grain that is easy to respect: a shire has 32 minions, so no minion needs to send to or receive from two partners. Bulk tiles can go through the scratchpads instead (Hand it to the next shire). Inside a shire the layouts are, as drawn above:
The four layouts, safe or not
- Safe: 1D rings (partners never change).
- Safe by the RTL rule, not run here: 2D in alternating phases (west pass, barrier, north pass, barrier; a 32-minion FLB plus credit barrier is 233 cycles against 68–114 per hop).
- Safe: the hardware trees (one ready bit per level; 444 cycles for a 32-minion reduce and broadcast).
- Hangs by the RTL rule, not run here (the one hang seen had two pairs through one minion): a naive 2D array whose cells take readies from two partners, until the chip is reset;
sys_emutracks partners separately and will not show it.
- Order chains by the map, not by shire ID. Stepping through shires 0, 1, 2, … 31 averages 3.3 hops per boundary and up to 8 (3.5 and up to 10 if the ring closes from 31 back to 0). This ring through all 32 compute shires moves one hop at every step: 0 24 9 25 2 11 19 27 18 10 17 14 22 26 15 23 31 7 6 30 29 5 28 20 12 21 13 1 16 4 3 8, and back to 0.
- Send boundaries, load bulk. Register-to-register messages suit small, latency-critical exchanges and in-network combines. Tiles that move between shires go faster through the scratchpads and L3: handing each stage's output to the next shire's scratchpad, instead of a round trip through DRAM, runs about 12× faster at about a thirteenth of the energy per byte (Hand it to the next shire).
The first rule as a calculator: pick a link, a message size and the work a cell does per step. A message's time is half the measured round trip at that size.
How it was measured
How it was measured, in full
- Round trips. Hart 0 of two minions bounces COUNT registers back and forth with TensorSend and TensorRecv (or credits, or atomic flags) 200–4,000 times after a warm-up, timed with the minion cycle counter. Pairs run one at a time, with a chip-wide barrier between them. The receiver checks the data: after a MOVE both sides must hold the sender's pattern, and after ADD or MAX the value a host-side replay predicts. All 3,699 pair and stream measurements, all 40 tree and 4 barrier runs, and every ring launch of the energy runs passed, and all 3,699, 40 and 4 passed again in each of the nine passes of the 26 September check. For the shire matrix, minion 0 of each of the 32 shires met every other shire's minion 0 (496 pairs), and the same again with minion 31. The intra-shire matrix covers all 496 minion pairs of shires 0 and 24.
- Streams, rings, trees and barriers. A stream sends N messages one way. A ring has every minion send to the next and receive from the previous, with neighbours alternating which comes first. Allreduces sum each minion's ID + 1 once with integer ADD, check the result, then time 500–1,000 more with integer MAX, which keeps the sum, and check it again. Barriers are timed over thousands of iterations.
- Clock. A background loop read the minion clock and board power from the service processor about 4 times a second during every run of the first session. The clock stayed at 600 MHz, and nanoseconds are cycles ÷ 0.6. In the 26 September check a 10 Hz sampler watched every launch, aifoundry2 and aifoundry1's card 1 were heated to 76 °C first so that their governors held 600 MHz (a boot service holds aifoundry3 at 600 MHz by setting its TDP to 0 W at every boot), and a launch that ran off 600 MHz would have been dropped: no nocbench launch was.
- Energy. The energies are the energy manual's re-runs, from the three-card check (26 September): six passes on each card at 600 MHz, each burst against idle readings taken before and after and corrected for the die's warming, with bursts dropped where the clock left 600 MHz or the meter's sampler slowed (every s+16 burst, so that row keeps aifoundry3's three passes of 23 September). The service processor's copy of the meter changes about every 126–135 ms polled by single commands, 156–157 ms under a 10 Hz sampler and 187 ms at 20 Hz on aifoundry2; 134–139, 157–159 and 188–189 ms on aifoundry1's card 1; 223–224, 263–264 and 320–322 ms on aifoundry3, the only card where the sampler's lengthening of the refresh the single commands see is resolved at 99% (the service processor's own pass, timed in its trace, lengthens under the 10 Hz sampler on every card: from 133 to 160 ms on aifoundry2, 224 to 266 ms on aifoundry3 and 135 to 162 ms on aifoundry1's card 1). The first session (18 September, aifoundry2, two runs in opposite orders) polled board power about 12 times a second, averaged it over about 4 s of launches per configuration and took as its baseline the median idle power in the 5 s gaps just before and after each configuration, because aifoundry2 was cooling after another user's matmul (die at 80 °C at the start of that session, 77–78 °C during the two runs used here) and its idle power drifted from 35.1 to 33.7 W over the session.
- Lab etiquette. Every card run was its own process under
timeout 10, on an otherwise idle card. The one hang (above) was cleared with a software reset of the chip (DM_CMD_RESET_ETSOC), and a health check passed before any further runs. The lab's rule, which we had not yet seen then, is to leave a hung card for the lab admin to power-cycle, because a software reset can itself hang a card.
Version history and provenance
Versions (one line per date; every wording is in the file's history): 18 September, published; 24 September, the energies per byte re-measured (first published from aifoundry2's two runs as 0.8–2.3 pJ inside a shire and 13–20 pJ across the mesh) and the systolic guidance and map orientation corrected; 25 September (version 3), the spin comparison, the 128 B messages and the per-hop slope's uncertainty corrected, and the layouts never run marked as predicted from the RTL; 26 September, three cards under a pre-registered plan (the in-shire barrier, the 32-minion allreduce, the tree levels, the credit round trips and the flag range corrected); 27 September, charts, and the work a cell needs between shires corrected to 800–1,350 cycles at one to ten hops (it said 750–1,000); 28 September, repeats cut, the note's counts given by both rules, and every energy in the chart, its table and the notes, the watts over idle too, now the energy manual's three cards, with the first runs as one note. Record: docs/reports/data/2026-09-25-claims-v3.
Caveats
Caveats, in full
- Three cards, one clock. The logical-to-physical map can differ between cards when a different shire is fused off; it matched marty1885's card on all three cards measured here. All runs here were at 600 MHz.
- Energy is whole-card power above idle. It includes the cores that issue the messages, busy or stalled, and the regulators' losses, so it overstates the cost of the wires: only the per-hop slope is close to the wires' own cost. The deltas are small (1.1–2.8 W on cards idling at 25–39 W), so treat each configuration's energy per byte as ±20%. The per-hop slope is 1.75 pJ/B per hop on the three cards together (1.73–1.78 on each; no two cards differ at 99%), 1.40–2.10 at 99% over the 18 passes, while one card's six passes pin it only to 1.31–2.25 (aifoundry2), 0.95–2.53 (aifoundry1's card 1) and 0.48–2.97 (aifoundry3).
- The GPU column is not our measurement. The numbers come from several published microbenchmark studies with their own methods and clocks. No measured A100 small-vector device-wide reduction was found, so the comparison uses its
grid.sync(). - Hart 0 only. Only hart 0 of a minion may issue tensor operations, so hart 1 was idle in these tests. Credits can also target hart 1, which was not measured.
Reproduce
Reproduce this
# build on the lab machine (sources only, nice -j4)
scripts/deploy-lab.sh aifoundry2 workloads/nocbench
# latency, size, trees and barriers: one timeout-10 process per probe, waits for an idle card
ssh aifoundry2 'cd ~/nekko && bash workloads/nocbench/run_lab.sh build/nocbench/host/nocbench_host build/nocbench-data'
# energy (quit et-powertop first; every configuration runs under timeout 10)
ssh aifoundry2 'cd ~/nekko && python3 workloads/nocbench/run_energy.py --host-bin build/nocbench/host/nocbench_host --out build/nocbench-data/energy-a'
# the same in the reverse order
ssh aifoundry2 'cd ~/nekko && python3 workloads/nocbench/run_energy.py --host-bin build/nocbench/host/nocbench_host --out build/nocbench-data/energy-b --only xshire1-c4,shire-c4,xshire6,xshire4,xshire2,xshire8,xshire16,xshire1,shire,neigh,pair,spin'
# the 26 September check on three cards: tools/claims-v3/lat/block.sh (its nocbench unit runs the same commands)
# numbers, layout check (12 search restarts, about 30 s) and chart data for this page; the energy
# re-runs come from docs/reports/data/2026-09-23-energy-manual/reruns.json (the rings' watts over idle from the passes it
# names), the three cards' latency passes from --v3
python3 workloads/nocbench/analyze.py docs/reports/data/2026-09-18-nocbench-aifoundry2 --memhier docs/reports/data/2026-09-18-memhier-aifoundry2 --search --v3 docs/reports/data/2026-09-25-claims-v3 --embed docs/reports/2026-09-18-et-soc1-on-chip-communication.html
Paths are relative to a checkout of yaroslavvb/et-soc1-prototyping;
scripts/deploy-lab.sh copies the sources to ~/nekko on the lab machine and builds there.
The analysis step rewrites only this page's embedded data.
Related reports
- The energy manual, §5 — these rings re-measured with bars on three cards; and §6, the energy of atomics and barriers.
- Heat per millimetre — a hop's wire energy with the bits on the links chosen, and 3.72 mm per hop measured from the die plot.
- Memory hierarchy — the scratchpad and L3 rates, and the pointer-chase rows used in item 2 of What the numbers say.
- Anatomy of a memory access — the same shire map; an L3 hit costs 110 cycles plus 12 per hop, and where the memory shires sit.
- Hand it to the next shire — bulk data handed between shires through the scratchpads, measured against a round trip through DRAM.
- One hot line stops a shire — what happens when many cores hammer one global-atomic line, such as a chip-wide flag: the shire that holds the line loses its own memory traffic.
- Ridge points — how much reuse each level of memory, and each kind of message on this page, demands before the tensor unit is the limit.
Sources
Sources, in full
- ET-SoC-1: ET Programmer's Reference Manual §7 (atomics), §9.4 (TensorSend, TensorRecv, TensorReduce, TensorBroadcast), §10 (fast local barriers), §11 (credit counters); ET-SoC-1 Preliminary Datasheet §4 (8 × 6 mesh, 44 stops); core-et Neighborhood MAS §4.7 (fast local messaging network), ET-Link Specification §5.5 (messages), RTL
rtl/shire/minion/dcache/dcache_reduce.v; et-platformsw-sysemu/insns/tensors.cpp; D. Ditzel et al., IEEE Micro 42(3), 2022 (570 mm², TSMC 7 nm). - M. Chang (marty1885), Investigating the ET-SoC-1 NoC, 2026 (the shire map, bandwidth and congestion tests), and etTopoScan.
- A100 and Hopper (H100, H800): A100 whitepaper; Luo et al., Dissecting the NVIDIA Hopper Architecture, 2025 (A100 L2 near/far, H800 distributed shared memory latency and bandwidth); Chips and Cheese, Nvidia's H100: Funny L2, and Tons of Bandwidth, 2023 (atomic handoff between SMs); Burtchell & Burtscher, Characterizing CUDA and OpenMP Synchronization Primitives, IISWC 2024 (
__syncthreads,__threadfence); Zhang et al., A Study of Single and Multi-device Synchronization Methods in Nvidia GPUs, IPDPS 2020, and their SC23 poster (grid.sync()on A100); Li et al., CREDIT, HPEC 2026 (cluster.sync()); Vellaisamy et al., ISPASS 2025 (launch latency). - Energy: Antepara et al., Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing, SC'25, table 3 (A100 4.71 pJ/bit for L2); W. J. Dally, Y. Turakhia, S. Han, Domain-Specific Hardware Accelerators, CACM 2020 (~100 fJ/bit·mm on-chip wires). One hop is 3.72 mm (measured from the die plot in Heat per millimetre, §2). Measured since, with the bits on the links chosen: 47–79 fJ per random bit·mm on board power, the meter used here, and 37–53 on the mesh rail alone (Heat per millimetre, §7).