Arbitrary coupled graphs execute natively

#exp077 · SNNLANG ·

Abstract

We asked whether the graph runtime can execute arbitrary coupling between independently driven excitatory–inhibitory circuits. We compared uncoupled, directional, reciprocal and delayed-reciprocal graphs, alongside a separate circuit-parity and runtime check.

Graph execution matched the explicit circuit and reproduced the expected causal effects of coupling direction and delay within the tested workload. These are bounded architecture and timing checks, not evidence for a biological coupling mechanism or performance on unrelated workloads.

Results

Circuit parity and runtime costs

MeasurementResult
Maximum parameter discrepancy0
Spike mismatches0
Maximum output discrepancy0
Checkpoint replay discrepancy0
Explicit median runtime0.0217 s
Graph median runtime0.0159 s
Graph overhead−26.6%
Compiled first invocation21.362 s
Compiled warm median0.0007 s
Compiled repeat discrepancy0
Explicit peak traced Python memory231736 bytes
Graph peak traced Python memory72611 bytes
Table 4: Matched eager execution used 100 timesteps and eight samples, with five timed calls after two warmups. CPU Inductor compilation used a separate 20-step, 2-sample input and three warm calls. Compiled repeat discrepancy compares two compiled calls, not compiled versus eager execution. Traced Python memory excludes native tensor storage; no uncertainty across independent timing sessions is estimated.

Reciprocal delayed circuit graph

Figure 1: Reciprocal delayed graph. Circuit A contains 16 excitatory and four inhibitory neurons; circuit B contains 12 excitatory and three inhibitory neurons. Each cross-circuit inhibitory projection targets the other circuit’s excitatory population.

Coupling and delay variants

Separate pulse and causal-planning tests passed (Fig. 5).

VariantCross projectionsDelay steps
Uncoupled0none
Unidirectional11
Reciprocal21
Delayed reciprocal25
Table 5: Coupling changed graph data only. A recurrent or feedback connection with no additional delay receives spikes on the next causal step; the explicit delay was 5 steps.

Coupling-variant rasters

Sparse or silent responses in these illustrative samples do not establish a phase-coupling mechanism (Fig. 2).

Figure 2: Excitatory spikes in the first sample of (A–B) uncoupled, (C–D) unidirectional, (E–F) reciprocal and (G–H) delayed reciprocal variants; the first panel in each pair is circuit A and the second is circuit B. The time axis is in milliseconds. Identical input tensors were reused across variants.

Methods

We separated graph-defined coupling checks from a bounded implementation comparison using PyTorch[1].

  1. Define and drive the coupled circuits. We authored two circuits with 16/four and 12/three excitatory/inhibitory neurons and eight/six input channels respectively. Deterministic input pulses arrived every ten and thirteen timesteps, with different offsets for the second sample. All variants used seed 77, 300 timesteps, 2 samples and 0.1 ms integration steps.
  2. Vary cross-circuit inhibition. We compared no cross projections, inhibition from A to B, reciprocal inhibition with one-step causal delay and reciprocal inhibition with five-step delay. Cross-projection weights were fixed at three and inhibitory conductance decay at nine milliseconds. We recorded all exposed populations and computed zero-lag correlation and peak cross-correlation lag; silent traces received the existing zero-valued diagnostic convention and were not interpreted as phase estimates.
  1. Compare explicit and graph execution. We initialised a separate 256-excitatory/64-inhibitory classifier identically in both implementations, using seed 17 and the same 100-step, eight-sample spike tensor. We compared parameters, spikes and outputs, and reloaded the graph state into an independently initialised model. All comparisons used recorded tensors.
  1. Measure timing and causal boundaries. We timed five eager calls after two warmups and calculated graph overhead relative to the explicit implementation’s median, using the unchanged ten-percent threshold. We measured CPU Inductor’s first invocation and three warm calls separately on a 20-step, two-sample input. Separate numerical tests checked delayed pulse arrival and causal planning; no coupling sweep or accelerator performance study was performed.

References

  1. Adam Paszke et al.: PyTorch: An Imperative Style, High-Performance Deep Learning Library. NeurIPS, 2019. doi:10.48550/arXiv.1912.01703