On-chip communication on the ET-SoC-1: 20 ns per mesh hop, and a bandwidth cliff at the shire boundary

18 September 2026 · first measured on one card (aifoundry2, in the AI Foundry lab) in one session; the latencies, trees and barriers re-measured on three cards (aifoundry2, aifoundry3 and aifoundry1's card 1, three passes each) on 26 September, and the energies per byte on the same three cards, six passes each, also on 26 September; minion clock 600 MHz and mesh clock 400 MHz throughout · GPU values from published microbenchmarks (sources at the end) · code: workloads/nocbench, raw data: docs/reports/data/2026-09-18-nocbench-aifoundry2 · part of the ET-SoC-1 measurement reports

Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 and aifoundry1's card 1. Corrected where a few cycles moved: the in-shire barrier is 233 cycles, not 237; the 32-minion allreduce 444, not 432; the credit round trip rises 24 ns per hop, not 20; also the tree levels and the memory flag's range. The energy per mesh hop held with no difference between the cards (the energies here are now that check's, six passes on each card). The three cards agree within a few cycles (the chip-wide barriers within 15), and the map and the distance chart show each card. Of 29 claims tested here, this page counts 17 held, 8 corrected, 2 differ by card and 2 held with no difference between the cards; the hub’s scoreboard, 3 “proven on the cards”, 25 “a test behind it failed” and 1 “differs by card”. Record: docs/reports/data/2026-09-25-claims-v3.

The ET-SoC-1's cores can pass data to each other directly, without going through memory, and quickly. Inside a shire a message takes 113–190 ns round trip, and each mesh hop adds 20 ns. A tree allreduce over all 1,024 minions takes 2.3 µs. Messaging costs 0.7–2.2 pJ per byte inside a shire with 1 KB messages, busy cores included (the energy manual's re-run on three cards at 600 MHz, 26 September). Crossing the mesh has a cost, though: there is 7–34× less message bandwidth between shires than inside them.

An A100's SMs can reach each other only through L2; Hopper adds direct shared-memory access, but only within a cluster of up to 16 SMs. The ET-SoC-1 has several direct paths. Hart 0 of any minion can send up to 127 × 32 B (about 4 KB, cycling through its 32 vector registers) straight into another minion's registers with TensorSend/TensorRecv, and the receiver can add, max or min the data into what it already holds. There are hardware reduction and broadcast trees, credit counters that let a core sleep until another core signals it, and barrier counters in every shire. This page measures each of these against its closest GPU equivalent and maps where the 32 shires sit on the chip's 2D mesh.

Core-to-core round trip, 32 B
113 ns
fast local network; 190 ns anywhere else in a shire
Each mesh hop adds
20 ns
round trip; 250 ns + 20 ns/hop across shires, r² 0.9999 over 496 pairs, on three cards
Allreduce, all 1,024 minions
2.3 µs
32 B through the hardware tree, 1,393 cycles on each card; 0.74 µs (444 cycles) within a shire; the same on three cards
Energy per byte moved
0.7–2.2 pJ
inside a shire, busy cores included; across the mesh 9.2 + 1.75 pJ per mean hop, the slope the same on three cards (600 MHz, 26 September)
Terms used on this page

The ET-SoC-1's compute cores are minions: small in-order RISC-V cores, each with two hardware threads (harts) and 32 vector registers of 32 B, of which only hart 0 issues tensor operations such as TensorSend; minions are written shire.minion, so 0.0 → 0.1 is minion 0 of shire 0 sending to minion 1. Eight minions form a neighbourhood and 32 a shire, whose 4 MB of SRAM holds a 512 KB L2, a 1 MB slice of the 32 MB L3 and a 2.5 MB scratchpad that any shire can address; the 32 compute shires, with the master, spare, I/O and PCIe shires, form a 6×6 grid on a mesh network-on-chip (NoC, 400 MHz), eight memory shires sit along two sides, and a hop is one step between neighbouring stops. A credit counter (FCC, fast credit counter) lets a minion sleep until another core adds a credit to it by a store to its shire's CREDINC register, and the service processor (SP) is the on-die management core that reads the board's power meter. More in the hub's glossary.

Where the shires are: a 6×6 mesh

Each square is one position on the mesh in marty1885's logical map (which appears to be the die turned a quarter), labelled with the logical shire ID the firmware uses. Colour is the measured TensorSend round-trip time from the selected shire (outlined), between minion 0 of each shire; stronger blue is slower. Card picks whose measurements: each of the three cards of the 26 September check (the median of its three passes) or the first session (aifoundry2, 18 September). Hover or focus a shire for its numbers; tap it, or press Enter on it, to measure from there. The four grey squares hold no compute shire: they are the master shire, the spare shire, and the PCIe and I/O shires, and they route traffic too. The datasheet's full 8 × 6 mesh adds eight memory shires, not shown: in this map's orientation the latency fit on aifoundry2 in Anatomy of a memory access (§4) places memory shires 0–3 above the top row and 4–7 below the bottom row, over the middle four columns (memory shire 2's own position is not pinned down by that fit).

What "Shade by" does

Shade by recolours the map with another of the page's shire-to-shire round trips, each on its own scale: TensorSend with 1 KB messages, the credit counter, or a flag passed through memory. The readout under the map gives that matrix's straight-line fit against mesh hops over all 496 shire pairs.

Compare three chain orderings on the map

A chain hands data from one shire to the next in some order, all the way round; the orange path is that order, and the darker dashed segment its longest step. Compare three orderings, on the same 32 compute shires the relay report chains:

How the map was checked

Coordinates follow marty1885's map (x across, y down). On each of the three cards every round trip between two shires is 150 + 12.02 × (Manhattan distance on this grid) cycles, with a worst residual of 1.1–1.4 cycles over all 496 pairs in each of the nine passes (1.1 in the first session). In every pass on every card, all 12 restarts of a search from random layouts, without his map, found the same distances.

Latency grows with distance, except through memory

Round-trip time between minion 0 of every pair of shires, against the number of mesh hops between them, for three ways to signal another core. TensorSend round trips rise by 20 ns per hop. Credit round trips rise by about 24 ns per hop on all three cards (234 + 24 ns per hop fitted, with a worst residual of 15 cycles; the first session's 246 + 20.4 ns fitted within 5), yet across shires they almost coincide with TensorSend's: the median pair differs by under a cycle, and 445 of the 496 by at most 15 cycles. A flag handoff through global atomics in memory, which is how a GPU passes a signal between SMs, lands on whichever shire's L3 slice holds the flag's line. It costs 610–1,150 ns (median 873 across shires on each card) whatever the distance: its slope against mesh hops is −2 cycles per hop, and at most +9 cycles (15 ns) per hop at 99% on each of the three cards. Points at 0 hops are pairs inside one shire.

Inside a shire, only the tree edges are fast

Each shire has 32 minions in four neighbourhoods of eight. Measuring all 496 minion pairs in shires 0 and 24 gives exactly two speeds, the same pairs and the same cycles in every pass on all three cards:

  • 68 cycles (113 ns) for 28 pairs per shire: in every neighbourhood, 0–1, 0–2, 0–4, 2–3, 4–5, 4–6 and 6–7.
  • 114–115 cycles (190 ns) for all 468 others. Being in the same neighbourhood doesn't help.

Those seven edges are the reduction tree, and they are the only links of the neighbourhood's fast local messaging network (core-et Neighborhood MAS §4.7), which delivers a 256-bit message in one cycle. Every other message leaves the neighbourhood and comes back through the shire crossbar.

Every primitive, against a GPU

Round trips at 600 MHz on ET-SoC-1; the three cards agree within a few cycles, the chip-wide barriers within 15 (26 September). GPU numbers are published measurements, at 1.41 GHz for the A100 and 1.755 GHz for the H800 (Hopper). The GPU column lists the closest equivalent where one exists.

Drawn on one time axis, the ET-SoC-1's messages between cores land where an A100's handoff through L2 does, while each of its barriers is slower than the GPU's nearest equivalent:

Every primitive on one time axis: the ET-SoC-1 measured on three cards, against the closest published GPU figure

The full comparison table, primitive by primitive
PrimitiveET-SoC-1, measuredWhat it doesClosest GPU equivalent
TensorSend/Recv, register to register113 / 190 ns
250 + 20/hop ns
Round trip of 32 B between two minions' vector registers: fast-network pair / rest of the shire / other shire. No memory is touched. Per 32 B register add 2.3–4.6 cycles each way.A100: no SM-to-SM path. A handoff through an L2 atomic is 166 ns one way (~330 ns round trip). L2 hits are 148 ns (near) and 253 ns (far). H800 (Hopper): a remote read of another SM's shared memory takes 103–121 ns, but only inside a cluster of up to 16 SMs.
Combine on arrival+0 cyclesThe receiver adds (fp32 or int32), maxes or mins the incoming data into its registers. The round trip is identical to a plain move for 32 B and 1 KB messages.Atomics executed in L2 (red.global); on H100, atomics on another SM's shared memory.
TensorReduce + TensorBroadcast (allreduce)0.74 µs · 2.3 µsHardware binary tree over the 32 minions of a shire, or over all 1,024 minions, 32 B each. All 32 shires can run their own trees at once without slowing each other.No tree hardware. A grid-wide grid.sync() alone, with no data, costs 1.1–1.2 µs on A100. Device-wide reductions are built from atomics or extra kernel launches (1.5–2.3 µs each).
Credit counters (FCC)200 ns
234 + 24/hop ns
Round trip: one store adds a credit to a core in any shire, and that core sleeps on a CSR until it arrives. Measured with the blocking wait inside a shire, and polled across shires.None. A consumer spins on a flag in global memory. A __threadfence() alone costs ~540 ns on A100.
Fast local barrier (FLB) + credits388 nsBarrier of 32 minions in a shire, with all 32 shires doing it at once.__syncthreads(), within one SM: 16–60 ns. cluster.sync() across 2 SMs on H100: 851 cycles.
Chip-wide barrier8.3 µs · 2.3 µsBarrier of 32 shires built from FLB, global atomics and credits (8.3 µs), or from the hardware allreduce tree (2.3 µs).grid.sync(): 1.1–1.2 µs on A100.
Flag through memory atomics610–1,150 ns
(median 873 across shires)
Round trip: global atomic swap, then spinning on a global atomic read, on DRAM lines that live in L3. No measurable dependence on distance (at most 15 ns per hop).This is the A100 pattern: ~330 ns round trip through L2, ~0.7 µs once a data fence is added.

Message size and bandwidth

Time per message for a one-way stream of messages from one minion to another, against message size. The receiver posts a new receive as soon as the last one finishes. A message is up to 127 registers of 32 B, reusing registers once past f31. Each path has a fixed cost per message (40, 86, 135 and 224 cycles) plus a cost per register (2.3, 4.0, 3.0 and 4.6 cycles), on all three cards. The chart is aifoundry2's (26 September); the other two cards agree within 0.3%. 1 KB round trips between all 496 shire pairs add 12 cycles per hop up to 5 hops and 27–36 per hop beyond (the 1 KB view of the distance chart), so the sender seems to have a limited number of packets in flight.

Per-link bandwidth, as a table

Per-link numbers are 32 B × COUNT over the stream's time per message at 600 MHz; the three cards give the same to the digits shown. With every minion sending 1 KB messages at once, the whole chip moves 1.1–3.0 TB/s inside shires and 0.09–0.16 TB/s between them; those rates and the energy per byte of each ring are in Energy per byte moved.

Reduction trees

Latency of one allreduce (TensorReduce up the tree, then TensorBroadcast back down) against the number of minions in the tree, for three sizes. Up to 32 minions the tree stays in one shire. Beyond that, each doubling adds a level between shires. Each level adds about one round trip between the minions it pairs: 68–71 cycles on the fast network (levels 0–2), 117 through the crossbar (levels 3–4), and 164–234 over the mesh (levels 5–9), where the slowest branch sets the pace (the first session's were 68, 114 and 156–235). The chart is aifoundry2's (26 September); the other two cards agree within 0.3%. The dashed line is the A100's grid.sync(), a barrier that moves no data.

Energy per byte moved

Why does a byte cost 6–26× more once it crosses the shire boundary? Hart 0 of all 1,024 minions sends 1 KB messages (128 B where marked) around rings inside a shire, or to shires a fixed number of IDs away; the rows are sorted by mean mesh distance. The three panels share the rows: bytes moved per second over the whole chip, board power above the idle just before and after each burst, and energy per byte. The last two are the energy manual's re-run of these rings (§5): each dot is the mean of six passes on each of three cards at 600 MHz (26 September), with the range the passes spanned. First measured on 18 September on aifoundry2 alone (two runs, sampled without the die temperature and so with no leakage correction, on a card cooling after another user's matmul): 0.8–2.3 pJ per byte inside a shire and 13–20 across the mesh, 2–13% above these means (4–16% above aifoundry2's own passes). Inside a shire messaging costs 0.7–2.2 pJ per byte with 1 KB messages. With 128 B messages a shire ring costs more: 3.7 against 2.2 pJ/B (3.6 against 2.1 on aifoundry2, 3.4 against 2.1 on aifoundry3, 3.9 against 2.5 on aifoundry1's card 1), a difference of +1.3 to +1.5 pJ/B, above zero at 99% on every card. The readout under the chart gives the cost's fit against mean hops across the mesh, per card and together. Grey rules are references: ET's own L2, remote-scratchpad (the mean over the zeros and random data its passes held) and DRAM reads at a steady 600 MHz (the energy manual, §4), and the A100's L2 and HBM accesses (SC'25). The last row is bulk data for comparison: every minion streaming TensorLoads from the scratchpad 16 shire IDs away (5.1 pJ/B when that scratchpad holds zeros, 11.8 when it holds random data).

The numbers in the chart
MessagesMean hopsWhole chipW over idle [range]pJ/B [range], passes
Pairs on tree edges0.02,992 GB/s2.05 [1.81–2.24]0.69 [0.61–0.76], 18
Rings of 8 (a neighbourhood)0.01,120 GB/s2.40 [1.95–2.74]2.16 [1.76–2.47], 18
Rings of 32 (a shire)0.01,120 GB/s2.47 [2.03–2.85]2.22 [1.83–2.57], 18
Rings of 32, 128 B messages0.0553 GB/s1.99 [1.57–2.31]3.66 [2.89–4.24], 18
Rings of 4 shires, 8 IDs apart1.6150 GB/s1.85 [1.35–2.20]12.50 [9.10–14.84], 18
Shires s and s+162.1156 GB/s2.03 [1.86–2.12]13.27 [12.20–13.81], 3 (23 Sep, aifoundry3)
To the next shire ID3.5124 GB/s1.77 [1.44–2.02]14.41 [11.72–16.47], 18
To the next shire ID, 128 B messages3.5112 GB/s2.25 [1.89–2.65]20.20 [16.99–23.82], 18
Rings of 8 shires, 4 IDs apart3.791 GB/s1.37 [1.10–1.59]15.13 [12.22–17.55], 18
Rings of 16 shires, 2 IDs apart4.587 GB/s1.47 [1.16–1.85]16.93 [13.33–21.29], 18
Rings of 16 shires, 6 IDs apart4.789 GB/s1.62 [1.30–1.79]18.29 [14.73–20.23], 18
For comparison: TensorLoad from the scratchpad 16 shire IDs away2.1959 GB/s–8.47 [4.40–14.43], 18

Watts over idle and pJ/B: the mean over the passes, in brackets the range they spanned; the whole-chip rates are the first session's timed launches (18 September), as in the energy manual. The same cores spinning with no messages drew 2.4–2.5 W over idle on each of the three cards (means of six passes; 2.45 W together, the dashed line in the middle panel). The s+16 row keeps aifoundry3's three passes of 23 September: the check dropped every s+16 burst on every card, because that ring slows the service processor that reads the meter (hub, §4.1), stretching the sampler to 63–141 ms a reading. The TensorLoad row pools nine passes whose scratchpads held zeros (5.10 [4.40–6.45] pJ/B) and nine that held random data (11.83 [9.53–14.43]).

Why the shire-boundary step is bandwidth, not wires

Read these as the cost of messaging, not of wires. Rings inside a neighbourhood or a shire with 1 KB messages drew about as much power over idle as the same cores spinning with no messages (from 0.3 W more to 0.2 W less, the means of six passes on each of three cards, 26 September), and 1 KB rings across the mesh drew 0.4–1.2 W less, so these figures are mostly the cost of 1,024 cores kept busy by messaging, divided by the bytes they moved. The step at the shire boundary mostly mirrors the drop in bandwidth, not costlier wires. The per-hop slope, 1.75 pJ/B on the three cards together, falls inside the mesh's own cost measured with chosen bit patterns in Heat per millimetre (1.4–2.3 pJ per byte per hop on board power, from free links to a loaded mesh); the 8–10 pJ intercept is mostly the busy cores.

What the numbers say

  1. All three cards have marty1885's shire layout. He built the map from shire-to-shire bandwidth on another card. On aifoundry2, aifoundry3 and aifoundry1's card 1, TensorSend round trips fit a+b×(Manhattan distance on his map) within 1.5 cycles and scratchpad loads at a steady clock within 0.2, and credit round trips rise with it too, within 15 cycles (5 in the first session). Pure Manhattan distance means the routes are shortest paths through all 36 cells, the four non-compute ones included. Minion 31 of each shire gives the same numbers as minion 0, so each shire meets the mesh at a single point.
  2. The mesh costs 20 ns per hop, in both directions together. That is 12 minion cycles at 600 MHz, or 4 mesh cycles each way at 400 MHz. In the one memory-hierarchy session, where aifoundry2's governor moved the clock mid-row, a hop still took 20 ns: remote scratchpad loads cost 66 minion cycles + (56 ns + 20 ns × hops), and a model fitted to two of the four rows fits the other two within a cycle at 600, 700 or 800 MHz, except one chase whose clock changed partway (5.4 cycles off). Only aifoundry2 can test this (a boot service holds aifoundry3 at 600 MHz by setting its TDP to 0 W at every boot).
  3. Three speeds, not four. A message takes 68 cycles on a reduction-tree edge, 114 anywhere else in the shire, and 150 + 12 per hop between shires. Chains and trees should put partners on those edges.
  4. Combining in the channel is free. A receive that adds or maxes into its registers takes the same time as a plain copy. That is the ⊕ of a (min,+) or (max,+) semiring, applied on arrival without a separate instruction.
  5. The hardware trees are the fastest way to synchronise the whole chip. A 1,024-minion allreduce of 32 B takes 2.3 µs, against 8.3 µs for a barrier made from global atomics and credits, on each of the three cards. Inside a shire it takes 0.74 µs, and all 32 shires can reduce at once without slowing down. Reductions and barriers belong on the trees.
  6. The shire boundary is a bandwidth cliff. Messages move 1.1–3.0 TB/s in aggregate inside shires and 0.09–0.16 TB/s between them, at 6–26× the energy per byte, mostly because the same ~2 W of busy cores moves 7–34× fewer bytes. Across shires, bulk data moves better as TensorLoads from the other shire's scratchpad (0.96 TB/s, at 5.1 pJ/B when that scratchpad holds zeros and 11.8 when it holds random data) than as messages; per byte it cost 3–5 pJ less than the s → s+8 ring on average on all three cards, but not in every pass (two of six on aifoundry3 and on aifoundry1's card 1): 5.8–8.9 pJ less in the passes that read zeros, −2.2 to +3.4 in those that read random data.
  7. Latency in nanoseconds is GPU-like; the mechanisms and the energy are not. A message crossing 1–10 hops takes 135–225 ns one way. An A100 takes 166 ns to hand a flag from one SM to another through L2. The differences lie elsewhere. The data lands in the receiver's registers, with an optional combine. There are 1,024 addressable endpoints instead of 108 SMs, plus hardware trees and sleeping credit waits. Messaging costs 0.7–18 pJ per byte (1 KB messages), busy cores included, against about 38 pJ per byte for an A100 L2 access.

A trap: one ready flag per minion

Before a sender may transmit, TensorSend waits for a "ready" message from its receiver. In the RTL (dcache_reduce.v, register partner_ready_peer), a point-to-point ready is a single bit per minion, not one per partner. If two different minions post their readies to the same sender before it has used them, one of them is lost. That receiver then waits forever, and so does its sender. The tree operations are safe: they keep one bit per tree level. The first cross-shire test here (aifoundry2, 18 September) ran two pairs through minion 0.0 without a barrier between them, and it hung the card; that is the one hang seen, and the test was not repeated. The stalled harts ignored the firmware's abort, and every later launch failed with KernelLaunchCmIfaceMulticastFailed until the chip was reset.

How the trap is avoided in this page's own probe

The simulator (sys_emu) tracks every partner separately, so it never shows the problem. nocbench now refuses any schedule where a minion could have two partners' readies outstanding. Pairs are separated by a chip-wide barrier, and rings only ever receive readies from the next minion. With that rule, TensorSend works across all 496 shire pairs.

Which TensorSend layouts hang? One shire's 32 minions, one neighbourhood per row, and the messages of one step of four layouts. Solid edges are the neighbourhood's tree edges (68-cycle round trips), dashed ones go through the shire crossbar (114 cycles); a minion that would take readies from two partners in the same phase is outlined and marked. Costs are this page's measured round trips.

What it means for systolic and wavefront designs

The first rule as a calculator: pick a link, a message size and the work a cell does per step. A message's time is half the measured round trip at that size.

How it was measured

How it was measured, in full
Version history and provenance

Versions (one line per date; every wording is in the file's history): 18 September, published; 24 September, the energies per byte re-measured (first published from aifoundry2's two runs as 0.8–2.3 pJ inside a shire and 13–20 pJ across the mesh) and the systolic guidance and map orientation corrected; 25 September (version 3), the spin comparison, the 128 B messages and the per-hop slope's uncertainty corrected, and the layouts never run marked as predicted from the RTL; 26 September, three cards under a pre-registered plan (the in-shire barrier, the 32-minion allreduce, the tree levels, the credit round trips and the flag range corrected); 27 September, charts, and the work a cell needs between shires corrected to 800–1,350 cycles at one to ten hops (it said 750–1,000); 28 September, repeats cut, the note's counts given by both rules, and every energy in the chart, its table and the notes, the watts over idle too, now the energy manual's three cards, with the first runs as one note. Record: docs/reports/data/2026-09-25-claims-v3.

Caveats

Caveats, in full

Reproduce

Reproduce this
# build on the lab machine (sources only, nice -j4)
scripts/deploy-lab.sh aifoundry2 workloads/nocbench

# latency, size, trees and barriers: one timeout-10 process per probe, waits for an idle card
ssh aifoundry2 'cd ~/nekko && bash workloads/nocbench/run_lab.sh build/nocbench/host/nocbench_host build/nocbench-data'

# energy (quit et-powertop first; every configuration runs under timeout 10)
ssh aifoundry2 'cd ~/nekko && python3 workloads/nocbench/run_energy.py --host-bin build/nocbench/host/nocbench_host --out build/nocbench-data/energy-a'
# the same in the reverse order
ssh aifoundry2 'cd ~/nekko && python3 workloads/nocbench/run_energy.py --host-bin build/nocbench/host/nocbench_host --out build/nocbench-data/energy-b --only xshire1-c4,shire-c4,xshire6,xshire4,xshire2,xshire8,xshire16,xshire1,shire,neigh,pair,spin'

# the 26 September check on three cards: tools/claims-v3/lat/block.sh (its nocbench unit runs the same commands)

# numbers, layout check (12 search restarts, about 30 s) and chart data for this page; the energy
# re-runs come from docs/reports/data/2026-09-23-energy-manual/reruns.json (the rings' watts over idle from the passes it
# names), the three cards' latency passes from --v3
python3 workloads/nocbench/analyze.py docs/reports/data/2026-09-18-nocbench-aifoundry2 --memhier docs/reports/data/2026-09-18-memhier-aifoundry2 --search --v3 docs/reports/data/2026-09-25-claims-v3 --embed docs/reports/2026-09-18-et-soc1-on-chip-communication.html

Paths are relative to a checkout of yaroslavvb/et-soc1-prototyping; scripts/deploy-lab.sh copies the sources to ~/nekko on the lab machine and builds there. The analysis step rewrites only this page's embedded data.

Sources

Sources, in full

← All ET-SoC-1 measurement reports