GCL investigation — 2026-09-08

For SRAM-resident inference on one C600/IPU21, I found no confirmed GCL-only
hardware facility that would accelerate ordinary tile-to-tile tensor movement
beyond the internal exchange mechanisms already present in ipu-stack. The most
useful follow-up is synchronization scheduling, especially advance abstention
from unrelated barriers. Exchange-block packet loopback remains an unverified
research possibility, not an established alternative data path.

This investigation used SDK 3.4.0 headers, dynamic symbols and selected x86-64
disassembly, public documentation/patents, public-source checkouts, and a host-only
C600 target probe. No IPU device was attached, reconfigured or benchmarked.
ipu-stack HEAD when inspected: `c4b43334adac4003b890241a2e1e941a02227492`.

**Source availability.** Searches for the full library name, GCL-specific symbols,
and Graphcore GitHub sources did not find a public implementation of libgcl.
This is a search result, not proof that no historical mirror exists. Some GitHub
page/API fetches failed, so an exhaustive live repository inventory was not
obtained. [PopLibs](https://github.com/graphcore/poplibs) is public, but its library
list and the local checkout do not contain GCL. The local
[mgcl.cpp](../popxl-addons/popxl_addons/ops/replicated_strided_collectives/mgcl.cpp)
implements strided collective helpers using GCL and Poplar; it is not the missing
GCL implementation. PopART and TensorFlow supply useful public callers.

**I/O tiles are ordinary tiles assigned a role.** There is direct source evidence
in [PopART's graph construction](../popart/willow/src/popx/irlowering.cpp): around
line 2945 it calls `gcl::perIPUTiles` twice, first for I/O tiles and then for the
complementary compute set, and creates virtual graphs over those IDs. They consume
the same pool of 1,472 processors and SRAM. Neither virtual graph creation nor the
allocation helper creates additional engines or cores. The
[allocation API](https://docs.graphcore.ai/projects/gcl-user-guide/en/latest/gcl/TileAllocation.html)
describes balancing over exchange blocks and contexts; its pairing requirement is
specifically mandatory for host I/O.

**Exchange blocks bridge the two networks.** The GC200 paper describes a
deterministic internal tile interconnect connected through exchange blocks to a
packet-switched external interconnect serving PCIe and IPU-Links. Thus the word
“external” describes the traffic's purpose; much of that network is physically
on-chip. These blocks do not mean the exchange instruction rows allocated by
ipu-stack, or extra tensor scratchpad memory.
[GC200 architecture description, section 2](https://www.microsoft.com/en-us/research/wp-content/uploads/2023/05/confidential-ml-within-ipus-arxiv.pdf).

The routing patent describes per-context exchange sequencers, a nominated XREQ
sender, XON/XOFF flow control, and packet/header conversion. It distinguishes
headerless internal exchange from TLink packet reception selected through the
incoming mux. It also describes eight XB indices on each edge, so one must not
mistake a host-side tile partition for a count of every physical block on both
edges. These are descriptions of embodiments, not a complete C600 programming
specification.
[Routing in a network of processors, figures 6A/B](https://patents.google.com/patent/US11615053B2/en).

The installed `poplar/Graph.hpp:281` independently exposes
`addExternalExchangeVertex`: it programs `INCOMING_DCOUNT` and `INCOMING_MUX`,
chooses east/west exchange, and requires at most one fixed XREQ sender per context.
That connects the patent terminology directly to this SDK. It does not establish
an autonomous SRAM-to-SRAM DMA engine usable concurrently with arbitrary compute.

**What the binary and probe establish.** [probe-output.txt](probe-output.txt) reports:

| Property | C600 target result |
|---|---:|
| Usable tiles | 1472 |
| Contexts per XB | 4 |
| Usable tiles per context | 46 |
| Combined host-side context indices | 32 |
| GCL minimum I/O tiles | 32 |
| Tiles sharing an exchange bus | 2 |
| Baseline internal exchange bytes/cycle | 4 |
| Internal synchronization zones | 1 |
| External synchronization zones | 6 |
| Local-worker / internal-tile sync selectors | 2 / 3 |

The host-side partition is `8 × 4 × 46 = 1472`. The physical coordinate space
also accounts for repair; do not substitute logical tile indices in physical-ID
bit formulas.

At library virtual address `0xab520`, `gcl::getMinIoTiles` calls
`Graph::getTarget`, calls `Target::getNumContextsPerXB`, and multiplies by eight.
It is therefore `8 * contextsPerXB`, not a measurement of a bandwidth optimum.
The PLT entries are `0x3e6d0` and `0x3e6e0`; their relocation slots are `0x4f73e0`
and `0x4f73e8`. [Disassembly](min-io-disassembly.txt).

Two exported C functions expose the mapping directly. For physical tile ID `p`:

```text
gcl_get_xb_context_index(p)      = ((p >> 1) & 0x1e) | (p & 1)
gcl_get_xb_context_tile_index(p) = ((p >> 5) & 0x3e) | ((p >> 1) & 1)

XB index                       = (p >> 3) & 7
context within XB              = ((p >> 1) & 2) | (p & 1)
```

The first function returns the combined context index, not just the context
within an XB. Addresses are `0xae830` and `0xae850`; a reconstruction function
`gcl_get_physical_tile_id` starts at `0xae800`.
[Disassembly](tile-map-disassembly.txt). The probe checks the formulas and compares
both functions against Poplar Target accessors for every one of the 1472 tiles.

Crucially, `perIPUTiles(g,0,32,false,true)` selects pairs spanning **16 contexts**,
two tiles per context, in XB indices **0,7,2,5**. The first pair has virtual IDs
0,1 and physical IDs 0,2. With pairing disabled the first 32 selected tiles also
span 16 contexts, but use different positions. This is observed policy in this
target/library version, not a universal architectural prescription. The probe
also checks that requesting the remaining 1440 tiles at offset 32 produces a
disjoint, complete partition. “Minimum 32” must not be explained as “one tile in
every context.”

**No evidence of arithmetic inside the exchange blocks.** The library exports
ring/broadcast/multiphase collective implementations and imports Poplar copies
and PopOps arithmetic. More strongly, disassembly of `gcl::syncful::opInPlace`
shows calls to `popops::map` at `0xced1e` and `0xcf235`, and
`popops::scaledAddTo` at `0xcf404`. [Disassembly](op-in-place-disassembly.txt).
`popops::reduceMany` and `program::CrossReplicaCopy` are also imported.
[Symbol inventory](symbols.txt). This is positive evidence of software reduction
and communication composition, not an exhaustive proof about every hardware unit.

`allReduceWithinReplica` does not mean a special reduction over tiles within one
IPU: the header's example and mapping contract explicitly concern IPUs within a
rank. For tile reductions, the public `poplibs/lib/popops` source is a better
implementation reference. `CollectiveBalancedReorder` is likewise a software
layout/padding helper, not a hardware rearrangement mode.

**Single-device opportunities, ranked.**

| Candidate | Assessment for ipu-stack |
|---|---|
| Advance SANS/ANS barrier abstention | Most concrete unexploited scheduling lead; benchmark before changing general codegen. |
| Software collective/layout optimization | Useful for reductions and avoiding gathers, but not newly discovered hardware. |
| Ordinary multicast and paired exchange | Already supported by ipu-stack; not a GCL gap. |
| Multiple independent internal barrier groups | Not established on IPU21; target reports one internal zone. |
| Same-IPU packet routing through external XBs | Unverified; GCL does not demonstrate this as a useful local transport. |
| Dedicated I/O tiles for SRAM-only inference | No demonstrated benefit merely from assigning the role. |

The synchronization patent describes SANS as advance abstention from future
barriers, tracked by ANS_DCOUNT, permitting unrelated tiles to continue. This
suggests posting abstentions before a long unrelated computation, then joining
when a true dependency arises. The exact operand/count semantics and interaction
with other sync types require a hardware check.
[Synchronization amongst processor tiles](https://patents.google.com/patent/US20210271527A1).

ipu-stack already emits `sans(0); sync(1)` on inactive exchange tiles, while
active tiles run a worker barrier and `sync(3)` before exchange
([codegen](../ipu-stack/crates/ipu-codegen/src/lib.rs), around lines 460–478).
Thus it would be wrong to report SANS as entirely missing. The opportunity is
advance posting across multiple phases and scheduling independent computation
over those phases. A useful experiment is: subset A performs several short
internal exchanges while subset B computes, with B abstaining in advance and
joining only at the final dependency. Measure elapsed cycles and verify all
payloads; establish whether B's computation still determines A's barrier time.

This primarily helps independent attention heads, independent requests, or
uneven subgraphs. It does not remove a true dependence such as reducing GEMM
partials before a consumer uses the result. Same-tile compute/exchange overlap
is a separate, harder experiment because workers, supervisor issue slots,
instruction fetch, receive writes and outgoing SRAM reads share resources.

The SDK architecture interface also defines an
`EXTERNAL_SYNC_INTERNAL_MUX` exception. External/global sync selectors should
not simply be substituted for additional internal barriers. A patent's general
multiple-group design is insufficient evidence that such a substitution works.

For local collectives, preserve a reduce-scattered layout when the consumer can
use shards, broadcast compact statistics rather than full intermediates, and
choose collective shape according to message size. Those are workload planning
ideas, not claims that the current compiler lacks every such optimization.
ipu-stack's existing multicast, paired 64-bit transfers and send/PIC composition
are documented in its [exchange reference](../ipu-stack/docs/EXCHANGE_INSTRUCTION_REFERENCE.md).

An on-chip packet loopback experiment would first need to establish local ETWR
reachability, routing configuration, completion and progress under backpressure.
Even if reachable, it competes for the tile receive interface and shared XB
contexts. I found no supported GCL API or measured result making it preferable
to headerless exchange for this workload. Likewise, dedicated Streaming Memory
is system-dependent and does not follow merely from owning a single IPU; it is
not an additional on-chip memory tier revealed by GCL.

**Reproduction.** From `/srv/home/gc-sdk`:

```sh
source poplar_sdk-ubuntu_20_04-3.4.0+1507-69d9d03fd8/poplar-ubuntu_20_04-3.4.0+73-aa67dd6164/enable.sh
g++ -std=c++17 -Wno-deprecated-declarations -Wno-cpp gcl-investigation/probe.cpp -lpoplar -lgcl -lipu_arch_info -o gcl-investigation/probe
./gcl-investigation/probe
```

libgcl.so SHA-256:
`559b87f0c00480a9d8558dda1f5c4c8aec7b7c05461fc9ad83614343fbb78b46`.
The probe runs entirely on the host using target metadata; it establishes SDK
behavior and mapping consistency, not silicon performance or packet loopback.
