We asked how recurrent inhibition and activity constraints shape the trade-off between MNIST accuracy and excitatory firing. We compared trained COBA and PING families across activity ceilings, cycle participation, oscillation frequency and input coupling.
PING had lower excitatory rates and retained more accuracy under strict ceilings, but no structural rate floor appeared. The comparison is confounded by gradient damping and therefore neither isolates a benefit of gamma timing nor measures energy use.
At the unpenalised operating points, PING reached 89.8% at 16.6 Hz and COBA reached 91.1% at 113.9 Hz. The comparison includes different gradient damping, so it does not establish a structural lower firing-rate limit (Fig. 1).
Figure 1:(A–B) Illustrative 400 ms COBA and PING rasters for the same digit-0 example, seed 42; E spikes are black and I spikes red. (C) Baseline validation accuracy over training. (D) Test accuracy versus mean E firing rate across activity ceilings; means ± SEM over three seeds, unpenalised points starred. These rates are test-set averages, not raster estimates. PING black and COBA red in panels C–D.
PING participation varied from 0.16 to 0.26 and oscillation frequency from approximately 60 to 18 Hz. The approximation differed from measured E rate by up to 27.5%, so participation was not constant. PING accuracy spanned 81–90%; COBA fell from 90% to 63% as the ceiling tightened (Fig. 2).
Figure 2:(A) PING participation, (B) PING gamma frequency, (C) E rate with the approximation and (D) accuracy across activity ceilings. Five penalised conditions per model, seed 42, 1000 test images each. The dashed curve in C is the approximation. These are individual-seed measurements, not across-seed estimates.
Final validation accuracies were 76.5% / 76.6% / 76.8% / 76.6%, while final I rates were 5.6 / 5.8 / 5.6 / 5.6 Hz. The similar endpoints across these initializations do not prove basin attractivity (Fig. 3).
Figure 3:PING learning curves for initial input-coupling means 0.05, 0.1, 0.3 and 0.9. Panels A–D show validation accuracy in that order; E–H show validation E (black) and I (red) rates in the same order. The rate ceiling is 1 Hz throughout; lines and shading show means ± SEM across seeds 42–44.
COBA’s penalty reached approximately 3 at . The empirical inhibitory-rate crossing marks a sampled transition, not a fitted bifurcation (Fig. 4).
Figure 4:Seed-42 networks trained with a 1 Hz ceiling; input weights scaled at inference over 24 values, all other weights fixed. (A) Cross-entropy, (B) rate penalty, (C) their sum, (D) test accuracy with a dotted chance line, (E) E rate and (F) I rate; 1000 images per condition. PING black, COBA red. Dashed marks training; the dotted marker labelled denotes the empirical input scale 0.475 where I rate crosses 0.05 Hz, not a fitted bifurcation. Penalty and total-objective axes stop at 4.
PING’s highest sampled accuracy was approximately 83%. At input scale , COBA reached approximately 63% at 8 Hz. This single direction of weight scaling does not map the full loss landscape (Fig. 5).
Figure 5:Figure 4 replotted against mean E rate: (A) cross-entropy, (B) rate penalty, (C) their sum, (D) test accuracy, (E) I rate and (F) input-weight scale. Stars mark the trained points: PING 4.7 Hz and COBA 2.1 Hz.
We reused networks and learning histories from the exp022 — Training Runs and reanalysed recorded inference measurements; no new training or simulation was performed.
Prepare digit inputs. Training used 6,300 MNIST images and 700 validation images from the official training partition. Pixels drove 784 Poisson channels at a 25 Hz maximum, for 200 ms with 0.1 ms steps.
Compare network configurations. Both conductance-based networks had 1,024 excitatory (E), 256 inhibitory (I), and 10 output leaky-integrate-and-fire neurons. Pyramidal-interneuron gamma (PING) enabled fixed E↔I coupling; COBA disabled it; E→E and I→I coupling were zero throughout. Only input and readout weights trained; class scores were mean pre-reset output membrane voltages (exp006 — Training). Voltage-gradient damping differed: 1 for COBA, 1,000 for PING.
Train with activity ceilings. Networks trained for 50 epochs with AdamW (zero weight decay), learning rate , batch size 256, and gradient-norm clipping at 1. Cross-entropy was supplemented by:
(1)
Here counts excitatory neuron ’s spikes in presentation , is the excitatory population, its size, its duration in seconds, and minibatch size. Rates and ceilings are in hertz; weights the dimensionless penalty . Six ceiling conditions (Training settings) and three seeds yielded 36 networks.
Evaluate training endpoints. Final-epoch weights, rather than validation-selected weights, supplied the endpoint comparisons. Each network was evaluated on the same 1,000 official-test images; frontier points show means and standard errors across seeds 42–44.
Measure cycle participation. For seed 42, oscillation frequency was the 5–150 Hz peak of trial-averaged Welch spectra of E activity [1]. Participation was the fraction of E-neuron/cycle pairs containing at least one spike, with cycles delimited by I-burst midpoints (Measurement details).
(2)
This diagnostic approximation relates mean E rate (Hz), dimensionless , and (Hz); repeated spikes and differing cycle/frequency aggregation prevent treating it as an identity.
Vary input coupling. Twelve PING networks used four initial input-coupling means and three seeds, with a 1 Hz ceiling throughout; validation histories measured recruitment during training. Separately, seed-42 PING and COBA networks trained at 1 Hz were evaluated after multiplying all input weights by dimensionless , holding other weights fixed, on the same 1,000 test images at 24 scales.
Expose retained comparisons. We displayed endpoint comparisons, cycle-participation probes and input-coupling sensitivity with their recorded seed roles and aggregation.
Input weights used a lower-clamped normal initializer with parent mean 0.9 and standard deviation 0.09, with 95% Bernoulli zeroing, sparsity compensation, and fan-in normalization. The mean parameter describes expected summed input coupling, not a per-connection mean. Initial zeros remained trainable; this was not a permanent connectivity mask. The four recruitment conditions replaced the mean by 0.05, 0.1, 0.3, or 0.9, with standard deviation one tenth of the mean.
Readout weights used a directly specified lower-clamped normal initializer (mean 1.12060546875, standard deviation 0.8349609375). PING’s fixed E→I and I→E initializers used summed-coupling means 1 and 2, respectively, with standard deviations 0.1 and 0.2 and normalization by source-population size. COBA set both connections to zero. Dale’s law was enforced; membrane time constants were not trained and adaptive thresholds were disabled.
Validation used three fixed Poisson draws per image. The training study also retained checkpoints selected by minimum validation cross-entropy, with ties resolved by accuracy then earliest epoch; the comparisons here used epoch 50.
Welch spectra used each full 200 ms demeaned E-population trace with a Hann window; constant traces were excluded. Spectra were averaged before selecting the 5–150 Hz peak, refined by three-bin parabolic interpolation capped at half a frequency bin. The symbol denotes this estimator even when its peak falls below the usual gamma range.
For participation, I-population activity was smoothed with a 1 ms Gaussian. Peaks required 5% of the trial maximum and at least half a cycle of separation, using that trial’s E spectral peak. Cycle boundaries lay midway between I peaks, with the trial endpoints closing the first and last cycles. Trials with no usable frequency or I peak were omitted from participation; active neuron/cycle pairs were pooled across accepted trials. This extends the exp041 — Firing Rate Tracks Gamma Frequency using the exp046 — One Spike per Gamma Cycle.
Input-scale values were 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.8, 0.9, 1, 1.15, 1.3, 1.5, 1.75, 2, 2.5, and 3. The crossing marker is the midpoint of the first adjacent pair whose mean I rates cross 0.05 Hz. The inference objective adds cross-entropy to Equation 1′s sample-wise penalty; its quadratic excess-rate dependence does not imply a quadratic dependence on input scale.
P. D. Welch. “The use of the fast Fourier transform for the estimation of power spectra: A method based on time averaging over short, modified periodograms.” IEEE Transactions on Audio and Electroacoustics 15(2), 70–73 (1967). doi:10.1109/TAU.1967.1161901