COBA training trajectories#
Final mean validation accuracy was 90.78% and E rate 110.87 Hz. The mean final-window E-rate slope was 0.777 Hz/epoch; 3/3 seeds met the accuracy criterion and 0/3 met the rate criterion (Fig. 1).
We asked whether classification accuracy and neuronal firing rate converge together during training. We audited the retained learning histories of unregularised COBA and PING classifiers, applying separate plateau criteria to performance and activity.
Accuracy could settle while excitatory firing continued to change, and rate stability differed across architectures and training replicates. Training convergence therefore needs separate accuracy and activity checks; a low firing rate alone does not demonstrate a fixed-rate attractor.
Uses the unregularised baseline learning histories from exp022 — Training Runs: COBA and PING, seeds 42, 43, 44. The complete trajectories support comparisons throughout learning, rather than only at a selected checkpoint.
Final mean validation accuracy was 90.78% and E rate 110.87 Hz. The mean final-window E-rate slope was 0.777 Hz/epoch; 3/3 seeds met the accuracy criterion and 0/3 met the rate criterion (Fig. 1).
Final mean E and I rates were 16.38 and 106.55 Hz. The mean final-window E-rate slope was 0.081 Hz/epoch; 3/3 seeds met the accuracy criterion and 0/3 met the rate criterion (Fig. 2).
The 99%-of-final-accuracy markers do not establish sustained convergence. These reused training observations are not a direct measurement of confidence or a causal rate–margin relation (Fig. 3).
We assessed finite changes in accuracy, activity and weights using recorded learning histories from unregularised classifiers.
Select the baseline histories. We reused all 3 seeds per architecture from the unregularised activity comparison. Each history contains 50 consecutive completed epochs; final values refer to the last epoch, not the checkpoint selected by minimum validation loss. The audit involved no new training or inference.
Identify the evaluation split. The training pool contained 7000 images from MNIST’s official training partition, split into 6300 optimisation samples and 700 validation samples. The official test partition of 10000 images was not used during training. Per-epoch evaluation averaged 3 fixed encoder draws per validation sample; those draws are not independent training seeds.
Recover the training conditions. Images drove 784 Poisson input channels at a maximum pixel rate of 25 Hz for 200 ms, with a 0.1 ms timestep. The networks used 1024 excitatory and 256 inhibitory hidden neurons and 10 output neurons. Mean output membrane voltage supplied the class logits. Training used surrogate gradients[1], learning rate 0.0004 and batches of 256. Voltage-gradient damping was 1 for COBA and 1000 for PING; no activity regulariser was applied.
Measure final-window drift. For each seed, we recorded validation accuracy, training and validation cross-entropy, and population-mean E and I rates. The final 10 epochs define the endpoint slope
Here is a measurement at epoch , is the final epoch, is the window length, and is change per epoch. Absolute slopes below 0.1 percentage points/epoch for accuracy or 0.05 Hz/epoch for E rate meet the audit’s operational stability criterion. This endpoint diagnostic does not exclude fluctuations within the window or prove asymptotic convergence.
We recorded per-seed slopes, first-to-final-epoch weight-norm ratios, and final-window weight-norm slopes, and computed means and sample standard deviations across seeds. Curves show individual seeds. The first epoch reaching 99% of final accuracy supplies a separate descriptive marker, averaged across seeds; it does not require subsequent accuracy to stay above the threshold.
Cross-entropy can keep rewarding larger decision gaps after the predicted class becomes correct:
Here is cross-entropy for one example, its true class, an alternative class, and their logits, and the softmax probability of the true class. The decision margin determines correctness by its sign, whereas cross-entropy depends on all the logit gaps. Aggregate accuracy can remain steady while individual predictions change.
The continued activity drift is consistent with ongoing optimisation, but these curves do not establish that confidence growth causes the rate increase. The mean-membrane readout depends on synaptic drive and membrane dynamics; it is not simply a linear function of the mean E rate. PING’s lower activity and slower drift in this comparison do not demonstrate a fixed-rate attractor or isolate a causal benefit of gamma timing.