Training Recurrent Weights Weakens PING Rhythmicity

Abstract

We asked whether surrogate-gradient training preserves PING rhythmicity when recurrent excitatory–inhibitory conductances can change. We compared trainable recurrent loops across initializations with a control whose recurrent PING weights remained frozen.

Recurrent training weakened rhythmicity early and did not recover the rhythmic control, while activity and accuracy changed by condition. This constrains these learning recipes, but does not show that recurrent learning must always destroy gamma or that the compared conditions are accuracy-equivalent.

Results

E→I pruning accompanies low contrast

Training from standard and 10%-standard recurrence drove most E→I weights to zero, while most I→E weights remained positive and their total mean grew. Both conditions ended with lower reference-image contrast and higher E rates than frozen recurrence. Zero-initialized recurrence remained zero (Fig. 1). These observations do not isolate pruning as the cause of reduced contrast or establish equivalent accuracy.

Final accuracy, E/I rates and reference-image contrast across four conditions, with initial and final recurrent nonzero fractions and mean weights.
Figure 1: Final checkpoints: (A) official-test accuracy, (B) per-neuron E/I firing rates over the same 1000 test images and (C) unsmoothed reference-image contrast 𝑅=𝑅contrast. Bars and error bars in A–C show means ±1 SEM across three independently trained seeds (sample SD divided by 3). In B, black denotes E and red I. Frozen denotes fixed standard recurrence; Std., 10% and Zero denote trainable recurrence initialized at standard, 10%-standard and zero weights. (D, E) Positive-weight fractions for E→I and I→E; (F, G) corresponding per-edge means including zeros, in model conductance units scaled by 10−3. Weight statistics pool all entries across the three seeds; no weight uncertainty is shown. Wide grey bars show initialization and narrow red bars show epoch 50. Black arrows connect before to after when the relative change is at least 5%; this is a display threshold, not a statistical-significance test.

Rhythmicity starts low

The unsmoothed trainable contrast averaged 0.05 after epoch 1 and ended at 0.018–0.111. Because there is no epoch-0 observation, these histories cannot resolve the intervening transition (Fig. 2).

Validation accuracy and E/I rates over 50 epochs, alongside reference-image rhythmicity; frozen recurrence retains high contrast while trainable recurrence has low contrast.
Figure 2: Recorded histories: (A) validation accuracy, (B) E rate, (C) I rate and (D) contrast from a fixed reference-image diagnostic. Lines show three-seed means; shading spans seed minima and maxima. Each series is smoothed with a five-epoch edge-padded moving average, not a confidence interval.

Methods

We reused networks from the exp022 — Training Runs and reanalysed recorded observations. No new training or simulation was performed.

  1. Compare recurrent trainability. Twelve conductance-based leaky-integrate-and-fire classifiers had 784 Poisson input channels, 1,024 excitatory (E), 256 inhibitory (I) and 10 output neurons. Three seeds per condition compared frozen canonical recurrence with trainable canonical, zero and 10%-canonical E→I/I→E conductances; E→E and I→I coupling stayed zero. Canonical initializer means were 11024 and 2256, respectively, with standard deviations one tenth of each mean and negative draws clamped to zero.

  2. Train on a held-out split. The 7,000-image subset contained 6,300 optimizer-training and 700 validation images from the official MNIST training partition. Input and readout weights trained for 50 epochs with AdamW, learning rate 4×10−4, zero weight decay, batch size 256, surrogate slope 1, voltage-gradient damping 1,000, gradient-norm clipping at 1 and no firing-rate penalty. Class scores were mean pre-reset output voltages; each 200 ms presentation used 0.1 ms steps and pixel-dependent input rates up to 25 Hz.

  3. Constrain conductance signs. Trainable recurrent magnitudes were projected onto the non-negative cone after each optimizer step. Inhibition arose through 𝑔𝐼(𝐸𝐼−𝑉𝑚), where 𝑔𝐼 is inhibitory conductance, 𝐸𝐼=−80 mV its reversal potential, and 𝑉𝑚 membrane voltage: a positive I→E weight need not become negative to inhibit. Input zeros remained trainable and could regrow.

  1. Evaluate final networks. All endpoint tests and weight comparisons used epoch 50, not validation-selected weights. Accuracy and whole-population mean E/I rates used the same 1000 official-test images per network; per-epoch validation metrics averaged three fixed Poisson encoding draws. For endpoint spectra, demeaned nonconstant E-population traces received full-trial Welch density estimation; the mean spectrum’s largest bin within 5–150 Hz defined the selected peak, without interpolation.

  2. Measure temporal contrast. After each epoch, the same fixed reference digit’s Poisson spike realization elicited a diagnostic response. E-population counts were binned at 1 ms; their autocorrelation was normalized by lag overlap and squared mean count, over 0–100 ms, then smoothed with weights (14,12,14) after replacing the zero-lag entry by its neighbour:

    𝑅contrast=𝐴lobe−𝐴trough𝐴lobe+𝐴trough
    (1)

    Here 𝑅contrast is dimensionless contrast, 𝐴trough the first local trough from lag 2 ms onward, and 𝐴lobe the preceding positive-lag maximum of the smoothed autocorrelogram. We reused the recorded scalar; it is neither a test-population rhythm estimate nor a calibrated probability of PING.

  1. Expose retained training evidence. We displayed retained validation, activity and temporal-contrast measurements with their distinct seed and illustrative-probe roles.

Dataset