Training Runs

Abstract

We built the MNIST model bank supplying trained networks and histories to downstream gamma-gated-sparsity experiments. We compared feedforward COBA controls with recurrent PING networks across activity penalties, inhibitory timescale, integration, recurrent weights and input drive.

PING combined lower excitatory activity with lower validation performance than the feedforward baseline, while wider sweeps exposed distinct accuracy–activity trade-offs. The retained bank provides training families; the bank itself does not establish that gamma timing causes any downstream benefit.

Results

Canonical COBA-PING learning curves

Validation-accuracy learning curves for the canonical COBA and PING training replicates.
Figure 1: Individual learning histories for both architectures and all three seeds in the full-data family; no across-seed averaging or uncertainty bands. Based on recorded measurements.

Canonical PING raster diagnostic

Seed-42 excitatory/inhibitory raster and population-rate diagnostic for canonical PING.
Figure 2: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum for canonical PING, seed 42, in one final-epoch digit-0 diagnostic; not a population estimate. Based on recorded measurements.

Activity-ceiling learning curves

Validation learning curves across the activity-ceiling conditions for COBA and PING.
Figure 3: Individual histories for two architectures, six ceiling settings, and three seeds on the reduced training pool; line colours distinguish settings and line styles distinguish architectures. No uncertainty bands are shown. Based on recorded measurements.

Unconstrained PING endpoint raster

Seed-42 raster for the unconstrained PING endpoint of the activity-ceiling sweep.
Figure 4: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum for PING with the activity penalty off, seed 42, in one final-epoch digit-0 probe. This is the reference endpoint, not a raster of the strictest ceiling. Based on recorded measurements.

Inhibitory-decay learning curves

Validation learning histories across six inhibitory-decay settings.
Figure 5: Individual PING histories for six GABA decay constants and three seeds per setting; the training pool and other recipe settings are fixed. No confidence bands are shown. Based on recorded measurements.

Six-ms inhibitory-decay raster

Seed-42 PING raster at the six-millisecond inhibitory-decay reference.
Figure 6: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum at the 6 ms reference decay, seed 42, in one final-epoch digit-0 probe; the spectrum describes this example only. Based on recorded measurements.

Timestep learning curves

Validation learning curves across five integration timesteps.
Figure 7: Individual histories for five timesteps (0.05, 0.1, 0.2, 0.3 and 0.6 ms) and three seeds per setting. Training and validation use each network’s own timestep and fixed 1.2/0.6-ms E/I refractory holds. Presentations last 199.8 ms at 0.3/0.6 ms and 200 ms otherwise; no across-seed bands are shown. Based on recorded measurements.

0.1-ms reference raster

Seed-42 PING raster at the reference 0.1-millisecond timestep.
Figure 8: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum at the 0.1 ms reference timestep, seed 42, in one final-epoch digit-0 probe. One reference diagnostic cannot establish timestep convergence. Based on recorded measurements.

Recurrent-training learning curves

Validation learning histories across frozen and trainable recurrent-loop conditions.
Figure 9: Individual histories for four recurrent conditions and three seeds each. The frozen PING control and three trainable initializations share the feedforward recipe; no confidence bands are shown. Based on recorded measurements.

Frozen recurrent PING control raster

Seed-42 raster for the frozen recurrent PING control.
Figure 10: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum for the frozen recurrent PING control, seed 42, in one final-epoch digit-0 probe; this example does not describe the trainable conditions. Based on recorded measurements.

Variable-rate training histories

Validation learning histories for all three variable-rate spike-count training replicates.
Figure 11: Three individual seed histories for the variable-rate spike-count recipe. Each validation presentation draws from the specified rate set; these curves do not separate accuracy by input rate and have no uncertainty bands. Based on recorded measurements.

Five-hertz variable-rate raster

Seed-42 excitatory/inhibitory raster for the variable-rate bank at a five-hertz input rate.
Figure 12: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum for variable-rate PING, seed 42, in one final-epoch digit-0 probe at 5 Hz maximum-pixel input rate. This diagnostic does not include output-neuron spikes or continuous-stream resets. Based on recorded measurements.

Input-coupling learning curves

Validation learning curves for four initial input-coupling settings and three seeds each.
Figure 13: Twelve individual histories under the fixed 1 Hz soft ceiling. Colours distinguish the four input-initialization means; no across-seed means or confidence bands are shown. Based on recorded measurements.

0.05 input-coupling raster

Seed-42 E/I diagnostic for initial input-coupling parent mean 0.05.
Figure 14: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum for input-coupling parent mean 0.05, seed 42, in one final-epoch digit-0 probe at the 1 Hz activity-ceiling target. This is one example, not an across-seed statistic. Based on recorded measurements.

0.1 input-coupling raster

Seed-42 E/I diagnostic for initial input-coupling parent mean 0.1.
Figure 15: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum for input-coupling parent mean 0.1, seed 42, in one final-epoch digit-0 probe at the 1 Hz activity-ceiling target. This is one example, not an across-seed statistic. Based on recorded measurements.

0.3 input-coupling raster

Seed-42 E/I diagnostic for initial input-coupling parent mean 0.3.
Figure 16: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum for input-coupling parent mean 0.3, seed 42, in one final-epoch digit-0 probe at the 1 Hz activity-ceiling target. This is one example, not an across-seed statistic. Based on recorded measurements.

0.9 input-coupling raster

Seed-42 E/I diagnostic for initial input-coupling parent mean 0.9.
Figure 17: (A) E/I spike raster, (B) E/I population rates and (C) inhibitory spectrum for input-coupling parent mean 0.9, seed 42, in one final-epoch digit-0 probe at the 1 Hz activity-ceiling target. This is one example, not an across-seed statistic. Based on recorded measurements.

Methods

The completed bank contains 102 models: 90 reused models with unchanged weights and histories, plus twelve newly trained timestep models. The three existing 0.1-ms timestep models were reused after execution-equivalence checks. We analysed these histories and generated 34 final-epoch seed-42 diagnostic recordings.

  1. Prepare the data. Stratified MNIST splits provided 54,000 training and 6,000 validation images for the baseline, and 6,300 and 700 for sweeps. The split seed was 42; the official test set was excluded from training and model selection.

  2. Build the networks. Each network contained 1,024 excitatory neurons and ten spiking leaky integrate-and-fire outputs. Recurrent networks added feedback through 256 inhibitory neurons; feedforward controls disabled it. Input and output projections were trained, while recurrent projections were fixed except in designated conditions. Weights retained their excitatory or inhibitory sign. Hidden E/I neurons used 1.2/0.6-ms absolute refractory holds; these were already the effective durations in the reused 0.1-ms executions.

  3. Vary the conditions. Seven families covered the baseline, activity ceilings, inhibitory decay, timestep, recurrent initialization and trainability, and input drive. Each condition used initialization seeds 42–44; Training-run specification sheets lists the grids and shared settings.

  4. Present the images. Pixels generated Poisson spikes over nominally 200 ms, normally with a maximum-pixel rate of 25 Hz and a 0.1 ms timestep. Variable-rate training sampled uniformly from eleven rates between 0.5 and 25 Hz. Timestep conditions ranged from 0.05 to 0.6 ms; whole-step presentations lasted 199.8 ms at 0.3/0.6 ms and 200 ms otherwise. Firing rates used realised durations.

  5. Calculate class scores. Most networks used mean pre-reset output voltage. Variable-rate training used spike counts and smaller initial readout weights.

    𝑧voltage=1𝑁𝑡∑𝑘=1𝑁𝑡𝑢out[𝑘],𝑧count=∑𝑘=1𝑁𝑡𝑠out[𝑘].
    (1)

    Each score applies to one output neuron and presentation. Here 𝑘 indexes the 𝑁𝑡 timesteps, 𝑢out[𝑘] is dimensionless pre-reset output state, and 𝑠out[𝑘] is a dimensionless binary output spike; both scores are dimensionless.

  6. Train the networks. Training used 50 epochs of AdamW, learning rate 0.0004, batches of 256, zero weight decay, and gradient clipping at 1. Surrogate gradients approximated spike derivatives[1] voltage-gradient damping was 1,000 for recurrent networks and 1 for controls. Activity-constrained conditions minimized

    𝐿total=𝐿CE+𝜆rate⟨max(0,𝑟𝐸−𝑟𝐸,ceil)2⟩batch.
    (2)

    𝐿total is total loss and 𝐿CE is mean cross-entropy. Each presentation’s mean excitatory firing rate 𝑟𝐸 and ceiling 𝑟𝐸,ceil are in hertz. Brackets average the individual penalties across the minibatch; 𝜆rate=0.041 with rates in hertz, or zero when disabled.

  1. Select models. Validation averaged three fixed Poisson encoding draws. Selection minimized mean cross-entropy, breaking ties by higher accuracy and then earlier epoch. Selected and final-epoch models were retained.
  1. Measure activity and retain models. Accuracy used selected models, whereas firing rates averaged final-epoch validation measurements across images and encoding draws. Retained models and histories support subsequent experiments. Learning curves show individual validation histories; baseline summaries average three seeds. The 34 new seed-42 digit-zero recordings covered every condition, including reused networks; their rasters are individual probes, not across-seed estimates. Spectra used the realised analysis-bin duration: grouping approximately 1 ms of simulation steps gave 0.9-ms bins at the 0.3-ms timestep and 1.2-ms bins at 0.6 ms.

Dataset

Appendix: Training-run specification sheets

The following shared sheet and all seven TR sheets specify the controlled recipes and their intended uses. The observed learning curves and diagnostic plots are in Results.

Shared production specification

All runs map Poisson-encoded pixels through 1,024 excitatory neurons to ten spiking output LIF neurons. COBA and PING use the same input-weight and readout-weight initialization distributions. COBA disables recurrent E/I coupling; PING adds a 1024→256→1024 E/I feedback loop and uses stronger backward-pass gradient damping to stabilize training through that recurrent path.

These are specified production settings, not newly measured results. Unless a TR-specific sheet says otherwise, every condition–seed network uses this contract. In the tables, 𝒩︀(𝜇init,𝜎init2) denotes a normal distribution with initialization mean 𝜇init and standard deviation 𝜎init; max(0,𝑥) replaces a sampled weight 𝑥 by zero when negative. LIF means leaky integrate-and-fire.

ParameterDefaultMeaning
DatasetMNIST784 normalized pixels encoded as independent Poisson channels
Presentation duration200 ms nominalOne static digit; 199.8 ms in the two coarsest timestep conditions
Integration timestep0.1 ms2,000 recurrent updates per presentation
Epochs50Training horizon for every production training replicate
Seeds42, 43, 44Three independently initialized training replicates per configuration
Minibatch256Presentations per optimization step
Optimizer learning rate4×10−4Shared learning rate
Input population784One channel per image pixel
Excitatory population1,024Learned stimulus representation
Inhibitory population256PING feedback population; silent when E/I coupling is disabled
𝜏AMPA2 msFixed excitatory synaptic decay
𝜏GABA6 msDefault inhibitory decay at the gamma operating point
Refractory hold, E/I1.2/0.6 msExactly represented by integer timestep counts; hidden neurons only
Input initial-zero fraction0.95Fraction set to zero at initialization; every entry remains trainable and may regrow
Input summed-coupling parent mean0.9Parent-Gaussian mean before lower clamping and fan-in normalization; shared by COBA and PING
Readout-weight initializationmax(0,𝒩︀(1.1206,0.83502))A directly stored lower-clamped Gaussian. Its zero-valued entries remain trainable; COBA and PING use the same initializer
E/I loop strengthCOBA: 0; PING: 1The forward architectural difference under comparison
Gradient dampingCOBA: 1; PING: 1,000Backward-pass stabilization for the recurrent PING loop; it does not change forward dynamics
Surrogate slope1Spike-gradient surrogate parameter
Default readoutmean-voltageA spiking 1024→10 output-LIF layer. At each timestep its pre-reset voltages enter the temporal mean, then emitted spikes subtract the threshold before the next update. These mean voltages—not spike counts—are the logits
Stored projection shapes784×1024; 1024×256; 256×1024; 1024×10Input→E, E→I, I→E, and E→class, in source-to-destination orientation

Specification: TR-01 — Canonical full-data reference

The full-data COBA and PING condition–seed networks are used for headline accuracy. They use the official MNIST training partition with no spike-budget penalty, so the comparison is not affected by the smaller dataset or regularization used in the sweeps. Ten percent of that partition is reserved for checkpoint selection; the official test partition remains untouched during training. This experiment reports their validation learning histories; independent official-test evaluation belongs to the downstream experiments.

Key parameterValueWhy it differs
ArchitecturesCOBA and PINGProvides a feedforward control and recurrent E/I model
Training pool60,000 official training samples54,000 optimizer-training and 6,000 validation samples
Input rate25 Hz maximum-pixel rateFixed-rate collection baseline
Spike budgetOffMeasures unconstrained capacity
Training replicates2 architectures × 3 seeds = 6Across-seed comparison for both models
Readout shape1024→10 spiking LIF outputsmean-voltage: mean membrane voltage supplies the logits

Specification: TR-02 — Activity-ceiling sweep

This run measures the trade-off between accuracy and firing rate as the activity ceiling is tightened. For each presentation, the loss computes the population-mean hidden-E firing rate, applies a one-sided quadratic penalty above the target, then averages the penalties across the minibatch. Quiet presentations therefore cannot offset active presentations. Expressing the target in hertz and normalising over samples, neurons, duration, and hidden layers makes the intervention comparable across those dimensions. A smaller training pool keeps the multi-seed sweep manageable; only the activity target changes across conditions. Comparisons with the full-data training replicates also change training-pool size and cannot isolate the penalty alone. The resulting checkpoints are used by exp024 — Accuracy Plateaus While Firing Rate Rises , exp025 — Accuracy and Firing Rate With and Without Inhibition , exp037 — Dropped Spikes vs Added Noise , exp038 — Switching On the Inhibitory Loop .

Key parameterValueWhy it differs
ArchitecturesCOBA and PINGDirect architecture comparison
Training pool7,000 official training samples6,300 optimizer-training and 700 validation samples
Hidden-E rate ceiling 𝑟𝐸,ceiloff, 25, 10, 5, 2.5, 1 HzSpans unconstrained through severe sparsity
Penalty strength0.041 when enabledCalibrated for the sample-wise, population-normalized Hz objective
Training replicates2 × 6 settings × 3 seeds = 36Error bars at every frontier point
Readout shape1024→10 spiking LIF outputsmean-voltage: mean membrane voltage supplies the logits

Specification: TR-03 — Inhibitory-timescale sweep

This run changes 𝜏GABA to test how the inhibitory timescale affects gamma frequency, firing rate, and accuracy. All other scientific settings are held fixed, with 6 ms as the standard condition. The resulting checkpoints are used by exp041 — Firing Rate Tracks Gamma Frequency , exp042 — Inhibitory Replay Perturbations Change Excitatory Firing , exp046 — One Spike per Gamma Cycle .

Key parameterValueWhy it differs
ArchitecturePINGThe manipulated quantity belongs to the recurrent inhibitory loop
Training pool7,000 samplesSweep-scale default
𝜏GABA4.5, 6, 9, 12, 18, 27 msMoves the inhibitory rhythm across a broad timescale range
Training replicates6 settings × 3 seeds = 18Across-seed estimate at each decay
Readout shape1024→10 spiking LIF outputsmean-voltage: mean membrane voltage supplies the logits

Specification: TR-04 — Integration-timestep sweep

This run changes the integration timestep while keeping E/I refractory holds at 1.2/0.6 ms. Presentations last nominally 200 ms and contain whole simulation steps. It tests whether the learned operating regime depends on numerical resolution; separately trained networks do not establish fixed-weight numerical convergence. The finest timestep requires the longest unrolled training trajectories. The resulting checkpoints are used by exp044 — Firing Rate Across the Timestep Sweep .

Key parameterValueWhy it differs
ArchitecturePINGTests the recurrent reference model
Training pool7,000 samplesSweep-scale default
Δ𝑡sim0.05, 0.1, 0.2, 0.3, 0.6 msTwelvefold resolution range at fixed physical refractory holds
Steps/presentation4,000; 2,000; 1,000; 666; 333200; 200; 200; 199.8; 199.8 ms realised durations
Training replicates5 settings × 3 seeds = 15Across-seed stability check
Readout shape1024→10 spiking LIF outputsmean-voltage: mean membrane voltage supplies the logits

Specification: TR-05 — Recurrent-initialization sweep

This run tests whether training preserves the PING loop or learns a useful loop from weaker initial conditions. Recurrent initialization and trainability change together, while the feedforward network and classifier remain fixed to the PING recipe. The resulting checkpoints are used by exp049 — Training Recurrent Weights Weakens PING Rhythmicity .

Key parameterValueWhy it differs
ArchitecturePINGManipulates the recurrent loop directly
Training pool7,000 samplesSweep-scale default
Loop conditionsfrozen PING; trainable PING init; trainable zero init; trainable 0.1 initSeparates built-in dynamics from recurrence learned during task training
Trainable projections𝑊EI and 𝑊IE only in trainable conditionsThe frozen condition is the mechanistic control
Training replicates4 conditions × 3 seeds = 12Across-seed comparison
Readout shape1024→10 spiking LIF outputsmean-voltage: mean membrane voltage supplies the logits

Specification: TR-06 — Variable-rate streaming bank

This condition trains PING across the input rates used by the continuous-stream classification study. One rate is sampled uniformly for each presentation, and the ten output LIF neurons produce class logits from their total spike counts rather than their mean membrane voltages. This keeps the cross-entropy logits dimensionless and gives one additional output spike the same logit increment. The readout weights use a small Gaussian initialization, 𝒩︀(0.05,0.042), lower-clamped at zero and governed by the shared non-negative constraint.

The recorded design rationale for this smaller initializer was to avoid saturation at the mean-voltage scale and silence at the generic fan-in-normalized scale. The present bank retains the chosen 𝒩︀(0.05,0.042) setting; its learning curves do not independently establish those earlier calibration comparisons. This is a specified spike-count recipe, not a claim that the initializer is uniquely optimal.

Output firing rates may still be reported as activity measurements. During streaming inference, the output neurons reset at digit boundaries while the hidden PING state continues. The resulting checkpoints are used by exp082 — Spike-Count Classification in a Continuous Stream .

Key parameterValueWhy it differs
ArchitecturePINGTarget model for streaming inference
Training pool7,000 samplesUses the shared sweep-scale training set
Input-rate set0.5, 0.75, 1, 1.5, 2, 3, 5, 7.5, 10, 15, 25 HzDenser sampling within the interval selected by the input-rate calibration study
Sampling ruleUniform categorical, independently per presentationMakes rate variation part of the training distribution
Readoutspike-countHidden E spikes drive ten spiking LIF class neurons; each logit is that class neuron’s total spikes over the presentation
Readout initialization𝒩︀(0.05,0.042), constrained non-negativeKeeps the spiking outputs near threshold at high training rates without the saturation caused by the default mean-voltage scale
Readout shape1024→10 spiking LIF outputsTen class neurons emit and reset throughout the presentation
Training replicates1 recipe × 3 seeds = 3Models for continuous-stream classification

Specification: TR-07 — Low-input recruitment sweep

TR-02 asks how changing the activity ceiling affects networks initialized at the standard input coupling. TR-07 asks the complementary question: with the strictest TR-02 ceiling fixed at 1 Hz from the first epoch, can PING recruit its inhibitory loop when the feedforward projection starts weak? It varies the parent mean used to initialize expected summed input coupling, before lower clamping and fan-in normalization; it does not directly set the mean of the stored synaptic matrix. The training pool, optimizer, recurrent loop, and mean-voltage readout remain at the reduced-sweep defaults. This bank retains all twelve models, and exp025 — Accuracy and Firing Rate With and Without Inhibition aggregates the three seeds at each setting to test recruitment and path dependence.

Key parameterValueWhy it differs
ArchitecturePINGTests recruitment of the recurrent inhibitory loop
Training pool7,000 samples6,300 optimizer-training and 700 validation samples
Epochs50Matches the reduced production standard
Input summed-coupling parent mean0.05, 0.1, 0.3, 0.9Varies initial feedforward drive from weak coupling to the shared 0.9 standard before clamping and fan-in normalization
Hidden-E rate target1 HzApplies the strictest TR-02 activity ceiling from epoch 0
Parameters held fixedPING loop, optimizer, dataset split, and mean-voltage readoutIsolates initial feedforward recruitment from the activity-ceiling sweep
Training replicates4 settings × 3 seeds = 12Across-seed estimate for every condition
Readout shape1024→10 spiking LIF outputsmean-voltage: mean membrane voltage supplies the logits

References

  1. E. O. Neftci, H. Mostafa, and F. Zenke. “Surrogate Gradient Learning in Spiking Neural Networks: Bringing the Power of Gradient-Based Optimization to Spiking Neural Networks.” IEEE Signal Processing Magazine 36(6), 51–63 (2019). doi:10.1109/MSP.2019.2931595