Accuracy and Firing Rate With and Without Inhibition

Abstract

We asked how recurrent inhibition and activity constraints shape the trade-off between MNIST accuracy and excitatory firing. We compared trained COBA and PING families across activity ceilings, cycle participation, oscillation frequency and input coupling.

PING had lower excitatory rates and retained more accuracy under strict ceilings, but no structural rate floor appeared. The comparison is confounded by gradient damping and therefore neither isolates a benefit of gamma timing nor measures energy use.

Results

Accuracy-rate activity ceilings

At the unpenalised operating points, PING reached 89.8% at 16.6 Hz and COBA reached 91.1% at 113.9 Hz. The comparison includes different gradient damping, so it does not establish a structural lower firing-rate limit (Fig. 1).

Two-by-two panel: COBA and PING single-trial rasters, per-epoch learning curves, and the accuracy–rate frontier across hidden-E rate ceilings.
Figure 1: (A–B) Illustrative 400 ms COBA and PING rasters for the same digit-0 example, seed 42; E spikes are black and I spikes red. (C) Baseline validation accuracy over training. (D) Test accuracy versus mean E firing rate across activity ceilings; means ± SEM over three seeds, unpenalised points starred. These rates are test-set averages, not raster estimates. PING black and COBA red in panels C–D.

PING participation and frequency

PING participation varied from 0.16 to 0.26 and oscillation frequency from approximately 60 to 18 Hz. The 𝑝part𝑓𝛾 approximation differed from measured E rate by up to 27.5%, so participation was not constant. PING accuracy spanned 81–90%; COBA fell from 90% to 63% as the ceiling tightened (Fig. 2).

PING participation fraction p and oscillation frequency f_gamma across the activity-ceiling sweep, with the p·f_gamma product overlaid on the measured E rate.
Figure 2: (A) PING participation, (B) PING gamma frequency, (C) E rate with the 𝑝part𝑓𝛾 approximation and (D) accuracy across activity ceilings. Five penalised conditions per model, seed 42, 1000 test images each. The dashed curve in C is the 𝑝part𝑓𝛾 approximation. These are individual-seed measurements, not across-seed estimates.

Input-coupling learning curves

Final validation accuracies were 76.5% / 76.6% / 76.8% / 76.6%, while final I rates were 5.6 / 5.8 / 5.6 / 5.6 Hz. The similar endpoints across these initializations do not prove basin attractivity (Fig. 3).

Across-seed mean per-epoch validation accuracy and E/I firing rates for four PING input summed-coupling parent means, one column per condition.
Figure 3: PING learning curves for initial input-coupling means 0.05, 0.1, 0.3 and 0.9. Panels A–D show validation accuracy in that order; E–H show validation E (black) and I (red) rates in the same order. The rate ceiling is 1 Hz throughout; lines and shading show means ± SEM across seeds 42–44.

Inference input-scaling response

COBA’s penalty reached approximately 3 at 𝑠=3. The empirical inhibitory-rate crossing marks a sampled transition, not a fitted bifurcation (Fig. 4).

Inference-time W_in scale sweep: CE loss, activity penalty, total objective, test accuracy, and E/I rates versus scalar s for PING and COBA.
Figure 4: Seed-42 networks trained with a 1 Hz ceiling; input weights scaled at inference over 24 values, all other weights fixed. (A) Cross-entropy, (B) rate penalty, (C) their sum, (D) test accuracy with a dotted chance line, (E) E rate and (F) I rate; 1000 images per condition. PING black, COBA red. Dashed 𝑠=1 marks training; the dotted marker labelled 𝑓∗ denotes the empirical input scale 0.475 where I rate crosses 0.05 Hz, not a fitted bifurcation. Penalty and total-objective axes stop at 4.

Accuracy-rate input scaling

PING’s highest sampled accuracy was approximately 83%. At input scale 𝑠=3, COBA reached approximately 63% at 8 Hz. This single direction of weight scaling does not map the full loss landscape (Fig. 5).

The W_in scale sweep re-projected with hidden E rate on the x-axis, trained operating points starred for PING and COBA.
Figure 5: Figure 4 replotted against mean E rate: (A) cross-entropy, (B) rate penalty, (C) their sum, (D) test accuracy, (E) I rate and (F) input-weight scale. Stars mark the trained points: PING 4.7 Hz and COBA 2.1 Hz.

Methods

We reused networks and learning histories from the exp022 — Training Runs and reanalysed recorded inference measurements; no new training or simulation was performed.

  1. Prepare digit inputs. Training used 6,300 MNIST images and 700 validation images from the official training partition. Pixels drove 784 Poisson channels at a 25 Hz maximum, for 200 ms with 0.1 ms steps.

  2. Compare network configurations. Both conductance-based networks had 1,024 excitatory (E), 256 inhibitory (I), and 10 output leaky-integrate-and-fire neurons. Pyramidal-interneuron gamma (PING) enabled fixed E↔I coupling; COBA disabled it; E→E and I→I coupling were zero throughout. Only input and readout weights trained; class scores were mean pre-reset output membrane voltages (exp006 — Training). Voltage-gradient damping differed: 1 for COBA, 1,000 for PING.

  3. Train with activity ceilings. Networks trained for 50 epochs with AdamW (zero weight decay), learning rate 4×10−4, batch size 256, and gradient-norm clipping at 1. Cross-entropy was supplemented by:

    𝑟𝑏=1𝑁𝐸𝑇present∑𝑛∈𝐸𝑛spike(𝑏,𝑛),𝐿rate=𝜆rate𝐵∑𝑏max(𝑟𝑏−𝑟𝐸,ceil,0)2.
    (1)

    Here 𝑛spike(𝑏,𝑛) counts excitatory neuron 𝑛’s spikes in presentation 𝑏, 𝐸 is the excitatory population, 𝑁𝐸 its size, 𝑇present its duration in seconds, and 𝐵 minibatch size. Rates 𝑟𝑏 and ceilings 𝑟𝐸,ceil are in hertz; 𝜆rate=0.041 Hz−2 weights the dimensionless penalty 𝐿rate. Six ceiling conditions (Training settings) and three seeds yielded 36 networks.

  1. Evaluate training endpoints. Final-epoch weights, rather than validation-selected weights, supplied the endpoint comparisons. Each network was evaluated on the same 1,000 official-test images; frontier points show means and standard errors across seeds 42–44.

  2. Measure cycle participation. For seed 42, oscillation frequency 𝑓𝛾 was the 5–150 Hz peak of trial-averaged Welch spectra of E activity [1]. Participation 𝑝part was the fraction of E-neuron/cycle pairs containing at least one spike, with cycles delimited by I-burst midpoints (Measurement details).

    𝑟𝐸≈𝑝part𝑓𝛾.
    (2)

    This diagnostic approximation relates mean E rate 𝑟𝐸 (Hz), dimensionless 𝑝part, and 𝑓𝛾 (Hz); repeated spikes and differing cycle/frequency aggregation prevent treating it as an identity.

  3. Vary input coupling. Twelve PING networks used four initial input-coupling means and three seeds, with a 1 Hz ceiling throughout; validation histories measured recruitment during training. Separately, seed-42 PING and COBA networks trained at 1 Hz were evaluated after multiplying all input weights by dimensionless 𝑠∈[0.05,3], holding other weights fixed, on the same 1,000 test images at 24 scales.

  1. Expose retained comparisons. We displayed endpoint comparisons, cycle-participation probes and input-coupling sensitivity with their recorded seed roles and aggregation.

Dataset

Appendix: Training settings

ParameterValue
Integration timestep Δ𝑡sim0.1 ms
Presentation duration 𝑇present200 ms; illustrative rasters use 400 ms
MNIST training pool7,000 official-training images: 6,300 optimizer-training / 700 validation
EvaluationFixed 1,000-image official-test subset
Epochs50
Rate ceilingsPenalty off, 25, 10, 5, 2.5, and 1 Hz
Independent seeds42, 43, 44
Trainable weights𝑊in: 784×1024; 𝑊out: 1024×10; 813,056 parameters
Stored parameters2,451,456, including fixed and zero weights
Fixed synaptic decay𝜏AMPA=2 ms; 𝜏GABA=6 ms

Input weights 𝑊in used a lower-clamped normal initializer with parent mean 0.9 and standard deviation 0.09, with 95% Bernoulli zeroing, sparsity compensation, and fan-in normalization. The mean parameter describes expected summed input coupling, not a per-connection mean. Initial zeros remained trainable; this was not a permanent connectivity mask. The four recruitment conditions replaced the mean by 0.05, 0.1, 0.3, or 0.9, with standard deviation one tenth of the mean.

Readout weights 𝑊out used a directly specified lower-clamped normal initializer (mean 1.12060546875, standard deviation 0.8349609375). PING’s fixed E→I and I→E initializers used summed-coupling means 1 and 2, respectively, with standard deviations 0.1 and 0.2 and normalization by source-population size. COBA set both connections to zero. Dale’s law was enforced; membrane time constants were not trained and adaptive thresholds were disabled.

Validation used three fixed Poisson draws per image. The training study also retained checkpoints selected by minimum validation cross-entropy, with ties resolved by accuracy then earliest epoch; the comparisons here used epoch 50.

Appendix: Measurement details

Welch spectra used each full 200 ms demeaned E-population trace with a Hann window; constant traces were excluded. Spectra were averaged before selecting the 5–150 Hz peak, refined by three-bin parabolic interpolation capped at half a frequency bin. The symbol 𝑓𝛾 denotes this estimator even when its peak falls below the usual gamma range.

For participation, I-population activity was smoothed with a 1 ms Gaussian. Peaks required 5% of the trial maximum and at least half a cycle of separation, using that trial’s E spectral peak. Cycle boundaries lay midway between I peaks, with the trial endpoints closing the first and last cycles. Trials with no usable frequency or I peak were omitted from participation; active neuron/cycle pairs were pooled across accepted trials. This extends the exp041 — Firing Rate Tracks Gamma Frequency using the exp046 — One Spike per Gamma Cycle.

Input-scale values were 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.8, 0.9, 1, 1.15, 1.3, 1.5, 1.75, 2, 2.5, and 3. The crossing marker is the midpoint of the first adjacent pair whose mean I rates cross 0.05 Hz. The inference objective adds cross-entropy to Equation 1′s sample-wise penalty; its quadratic excess-rate dependence does not imply a quadratic dependence on input scale.

References

  1. P. D. Welch. “The use of the fast Fourier transform for the estimation of power spectra: A method based on time averaging over short, modified periodograms.” IEEE Transactions on Audio and Electroacoustics 15(2), 70–73 (1967). doi:10.1109/TAU.1967.1161901