https://arxiv.org/api/sdF4qGK6z4urscCUCXsqzmxvh5s2026-09-12T21:00:05Z182493015http://arxiv.org/abs/2605.30361v2Gradient-Free Training of Spiking Neural Networks via Low-Rank Evolution Strategies2026-09-07T12:59:27ZSpiking Neural Networks (SNNs) offer compelling energy efficiency on neuromorphic hardware, yet their training remains challenging because the discrete spike threshold is non-differentiable. Surrogate-gradient methods sidestep this by approximating the derivative, but they impose backpropagation infrastructure that is incompatible with on-chip learning. Evolution Strategies (\es) are a natural gradient-free alternative, yet their computational cost scales with the number of parameters, making them impractical for large weight matrices.
We present a method for training SNNs using EGGROLL, a low-rank factorisation of ES perturbations that reduces per-generation memory from $\mathcal{O}(mn)$ to $\mathcal{O}(r(m{+}n))$. Combining EGGROLL with a Leaky Integrate-and-Fire SNN on N-MNIST, we demonstrate that gradient-free training achieves 79.21% test accuracy while reducing per-generation wall-clock time by 2.23$\times$ relative to full-rank ES. Our results demonstrate EGGROLL is viable for SNN training, with a clear accuracy-speed tradeoff, compatible with training on neuromorphic hardware without surrogate gradients.2026-05-14T05:42:15Z12 pages, 4 figuresDhruv PatankarSachit Ramesha Gowdahttp://arxiv.org/abs/2609.07418v1Photonic reservoir computing with dimensionally compressed readout2026-09-07T12:28:34ZThis work addresses a hardware constraint in reservoir computing: the limited size of the readout layer imposed by systems with a physical readout. We investigate a strategy to accommodate this constraint based on random projection, which compresses high-dimensional reservoir states into a lower-dimensional subspace while preserving key properties of the source space and information- processing capabilities. To evaluate this approach, we compare a small, standalone time delay reservoir against a larger configuration whose output is projected down to match the same restricted readout dimension. Using task-independent metrics, we demonstrate that the distribution of information-processing capacities may differ between the two configurations, even at identical readout sizes. Furthermore, we perform a comprehensive hyperparameter scan to assess how both systems behave under varying physical regimes. Finally, we benchmark this approach on the standard NARMA10 task, showing that the random projection framework can yield superior performance compared to a standalone constrained reservoir, within specific compression range. These results provide a scalable pathway to bypass physical readout bottlenecks in hardware-based reservoir computing.2026-09-07T12:28:34ZGerald KobiMohab AbdallaMiguel C. SorianoDamien Rontanihttp://arxiv.org/abs/2605.28703v2A Fresh Look at Lamarckian Evolution and the Baldwin Effect2026-09-07T08:06:35ZBaldwinian and Lamarckian evolution have existed for a long time in evolutionary algorithms (EAs) without ever dominating the academic literature or practical applications. In this work, we use modern empirical and theoretical methods to revisit Lamarckian and Baldwinian evolution and rigorously compare them with the generic Darwinian evolution. On the empirical side, we run a comprehensive suite of experiments on graphs from six different datasets from the recent GraphBench benchmark on Maximum Independent Set and Maximum Cut problems. Our results show that Baldwinian and Lamarckian evolution consistently outperform Darwinian evolution, confirming the great potential of local search augmented evolutionary algorithms. Notably, in the great majority of cases, all EAs outperform recent deep learning baselines and approach the performance of highly specialised heuristic and exact solvers. We furthermore report a high-performing set of generalist parameters for all studied evolution types that we hope will be of use to practitioners in future. On the theoretical side, we extend the existing Deceptive Leading Block benchmark to arbitrary block length and use tools from modern theoretical runtime analysis to prove upper and lower bounds on the expected runtime. For block lengths greater than two, Baldwinian evolution is asymptotically faster than Lamarckian which is asymptotically faster than Darwinian evolution. When accounting for the cost of the local search procedure in fitness evaluations, the ordering depends on the implementation with Baldwinian evolution staying fastest from small block lengths onwards, explaining its strong empirical performance.2026-05-27T16:30:39ZFull version with appendix of the work publishedIn Proceedings of the 19th International Conference on Parallel Problem Solving from Nature (PPSN XIX), Lecture Notes in Computer Science, vol. 16988, Springer, pp. 459-474, 2026Inès BenitoJohannes F. LutzeyerBenjamin Doerr10.1007/978-3-032-36229-2_28http://arxiv.org/abs/2609.06826v1Formation of structural attractors in neuromorphic systems2026-09-06T20:44:04ZThis paper examines the theory of Invariant Structural Learning (ISL), which proposes a non-optimization approach to concept formation. Learning is interpreted as convergence to structural attractors in a hypergraph space, rather than as the minimization of a global loss function. The paper presents the ISL model, including its mathematical formalization, computational verification, and a hypothetical neurobiological interpretation. The mathematical section introduces the formal apparatus of the structural reduction process and proves its finite convergence, the existence and uniqueness of class structural attractors, and the self-organization of attractor maps. The computational section demonstrates the feasibility of the proposed approach on classical image recognition tasks, utilizing the proposed learning mechanism without backpropagation and with extremely small training datasets. Finally, the neurobiological section formulates hypotheses regarding the possible implementation of structural attractors in dendritic trees, neural coding as a projection of internal attractor dynamics, and the development of neural architectures supporting the proposed learning concept. These hypotheses are discussed in the context of modern experimental data in the fields of dendritic computations, synaptic plasticity, and the structural organization of neural circuits. The proposed neurobiological mechanisms are presented as testable hypotheses rather than established biological facts. The results demonstrate the mathematical consistency and computational feasibility of the proposed model, while the neurobiological hypotheses outline potential directions for its experimental verification.2026-09-06T20:44:04Z131 pages, 4 figures, 5 tablesYurii ParzhynAlexander SchwarzmannMykyta LapinKostiantyn Bokhanhttp://arxiv.org/abs/2609.10588v1Threshold-Based Selection for Continuous Optimization: A Leaf-Abscission Instantiation2026-09-06T17:41:35ZThis paper formalizes threshold-based selection as an evaluation-gating architecture in which each incumbent is tested before variation and a replacement is generated and evaluated only when contextual pressure exceeds intrinsic strength. The mechanism is instantiated as Leaf Abscission Optimization (LAO), using rank-based strength, a phenological seasonal signal, diversity modulation, environmental pressure, and a base regrowth kernel. A blocked two-to-the-fourth-power factorial analysis at dimension 10 on the CEC 2017 suite reduces the original multi-layer design to a parsimonious core: drift is harmful, while the other three auxiliary layers show no robust independent evidence of benefit. The resulting LAO-Core attains the third-best mean Friedman rank among nine optimizers at dimensions 10, 30, and 50 under the equal evaluation budget. A four-budget sweep shows budget-dependent relative performance, with adaptive differential-evolution baselines gaining relative advantage at larger budgets; the nine-cell dimension-budget analysis establishes neither an evaluation-budget-per-dimension-only law nor a statistically significant dimension-budget interaction. A paired intervention shows that diversity modulation changes late-run replacement behaviour without a detectable effect on final error at the tested budget. The evidence supports LAO as a parsimonious evaluation-gating mechanism with regime-qualified competitiveness, rather than as a generally superior optimizer.2026-09-06T17:41:35Z15 pages, 12 figuresNasser Khalilihttp://arxiv.org/abs/2609.06341v1Linear Algebra Foundations of Efficient Attention: A Phase Reversal in Rank Collapse Under SVD Compression2026-09-06T02:27:01ZLinear algebra provides the framework of concepts (matrix rank, singular value decomposition (SVD), and eigendecomposition) that modern artificial intelligence employs to encode, compress, and propagate information through neural networks. This paper unifies fourteen separate peer-reviewed works analyzing the usage of these techniques in the context of transformer-based foundation model research, focusing on three areas of the topic: derivations and properties of self-attention matrices' output rank, compression methods that purposefully utilize this phenomenon, and the low-rank key-value (KV) cache projection and its semiseparable-matrix duality to linear attention and state-space structured models. We were motivated to conduct this work after observing an open problem in this literature: the interplay of the mentioned compression methods with natural rank collapse of the network. With this paper, we report an original finding that using SVD compression of attention projections actually has the opposite effect on the rank collapse of the network: while it strongly suppresses it at initialization, it accelerates on pretrained models (for GPT-2 124M, GPT-2 Medium 355M, and Pythia-160M) with minimal risk of object aliasing artifacts appearing (verified on all compression ratios) and is consistent across four rank estimation methods. A controlled causal decomposition of the effect in both settings showed that the reason for this behavior can be explained by the choice of the subspace SVD makes when compressing the matrix better than the reduction of the operator norm it achieves, explaining roughly 76% of the effect at initialization and 83% on the pretrained weights, providing a refinement to the calibration-aware compression viewpoint and an explanation of why it outperformed naive SVD truncation.2026-09-06T02:27:01ZAnjaneya Teja Sarma Kalvakolanuhttp://arxiv.org/abs/2411.15876v3DUA-D2C: Dynamic Uncertainty Aware Method for Overfitting Remediation in Deep Learning2026-09-05T18:47:48ZOverfitting remains a significant challenge in deep learning, often arising from data outliers, noise, and limited training data. To address this, we previously proposed the Divide2Conquer (D2C) method, which partitions training data into multiple subsets and trains identical models independently on each. This strategy enables learning more consistent patterns while minimizing the influence of individual outliers and noise. D2C's standard aggregation typically treats all subset models equally or based on fixed heuristics (like data size), potentially underutilizing information about their varying generalization capabilities. Building upon this foundation, we introduce Dynamic Uncertainty-Aware Divide2Conquer (DUA-D2C), an advanced technique that refines the aggregation process. DUA-D2C dynamically weights the contributions of subset models based on their performance on a shared validation set, employing a novel composite score of accuracy and normalized prediction entropy. This intelligent aggregation allows the central model to preferentially learn from subsets yielding more generalizable and confident edge models, thereby more effectively combating overfitting. In this work, we provide a rigorous theoretical justification for this approach, analytically demonstrating how dynamic parameter fusion reduces model variance. Empirical evaluations on benchmark datasets spanning image, audio, and text domains demonstrate that DUA-D2C significantly improves generalization. Our analysis includes evaluations of decision boundaries, loss curves, and ablation studies, highlighting that DUA-D2C provides additive performance gains even when applied on top of standard regularizers like Dropout. This study establishes it as a theoretically grounded and effective approach to combating overfitting in deep learning. Our code is publicly available at: https://github.com/Saiful185/DUAD2C.2024-11-24T15:31:22ZThis version (v2) extends our previous work (arXiv:2411.15876v1) on Divide2Conquer (D2C) by introducing Dynamic Uncertainty-Aware Divide2Conquer (DUA-D2C). The manuscript has been published at Complex and Intelligent Systems. Find the published version at https://doi.org/10.1007/s40747-026-02251-1Complex Intell. Syst. 12, 149 (2026)Md. Saiful Bari SiddiquiMd Mohaiminul IslamMd. Golam Rabiul Alam10.1007/s40747-026-02251-1http://arxiv.org/abs/2609.06042v1Minimizing the Effect of Sleep Deprivation in the Forward-Forward Algorithm2026-09-05T11:54:27ZThis paper addresses the challenge posed by sleep deprivation in the Forward-Forward algorithm, where separating the two passes in this algorithm and imbalancing the data processing in the passes is considered an imitation of the cognitive processes observed in humans suffering from sleep deprivation. Previous research has demonstrated that sleep deprivation in the Forward-Forward algorithm has a catastrophic effect on learning efficacy. To mitigate this issue, we explore several approaches; these include alternative activation, optimized loss function, and threshold tuning. To simulate periodic rest, we reduce the number of positive passes in alternating epochs, creating short break phases. We additionally investigate the potential of caffeine-induced stimulation to enhance performance during sleep-deprived conditions. Experimental evaluations conducted on the MNIST and Fashion-MNIST datasets demonstrate that these modifications improve accuracy under the context of sleep deprivation. For example, a 2%-62% accuracy gain is observed in a severe sleep deprivation setting (16 positive or awake periods and 1 negative or sleep period). The approaches also enhance the resilience of the algorithm and its alignment with the adaptive mechanisms of human cognition.2026-09-05T11:54:27Z14 pages, 6 figures, 7 tables. Accepted at the 28th International Conference on Pattern Recognition (ICPR 2026), Lyon, France, August 17-22, 2026. This is the accepted manuscript; the Version of Record is published by Springer in Lecture Notes in Computer Science and is available online at https://doi.org/10.1007/978-3-032-31933-3_40Pattern Recognition. ICPR 2026. Lecture Notes in Computer Science, pp. 590-604. Springer, Cham (2026)Joy DattaPuja SahaRawhatur RabbiNafiz Imtiaz RafinSwakkhar ShatabdaMd. Golam Rabiul AlamChad Mourning10.1007/978-3-032-31933-3_40http://arxiv.org/abs/2608.26741v2Asymmetric Coupling Anisotropy for Causal Information Filtering in Physical Reservoirs2026-09-05T03:24:58ZWe demonstrate a physical mechanism for causal information filtering in a physical reservoir computing (PRC) by exploiting asymmetric coupling anisotropy. Using a network of coupled Duffing oscillators, we show that the directionality of internal coupling induces a spatial gradient in the effective potential, establishing a deterministic upstream-to-downstream information flow. This anisotropy allows for the selective amplification of semantic drifts, triggering a macroscopic saddle-node bifurcation as a physical interlock before global computational failure. Through spatiotemporal analysis of a 50-node system under traveling wave inputs, we confirm that local phase transitions effectively purge anomalous information while preserving the computational integrity of the remaining nodes. The results suggest that the intrinsic causality of the reservoir's topology provides a robust framework for autonomous reliability and fault-tolerant physical intelligence.2026-08-27T07:31:49Z14 pages, 9 figuresTakashi HikiharaYuma Aokihttp://arxiv.org/abs/2609.05304v1What Makes a Redundant Representation Remember? Lineage Isolation, Not Masking2026-09-04T15:57:10ZMemory-based evolutionary algorithms for dynamic optimization often carry a redundant second copy of the genotype and expose only one copy to the objective, on the assumption that the shielded copy accumulates information about past optima. We show this assumption is false as usually implemented, and identify the structural property that actually determines whether the shielded copy retains information. We formalize such methods as a gated dual-copy representation with two independent design axes: a gating rule deciding which copy is evaluated, and an inheritance rule deciding whether the two copies mix across generations. A ablation shows retained information is governed almost entirely by the inheritance rule (21.4 vs. 1.3 bits) and is nearly invariant to the gating rule. Per-locus independent inheritance reshuffles cross-locus structure every generation, so shielding preserves the variance of the hidden copy while destroying the pattern that constitutes a memory. Under isolated inheritance the memory effect is real: against a single-copy baseline matched for representation budget, the method gains +0.010 AUC when optima recur periodically and loses 0.078 when they drift unidirectionally---a 0.089 separation under otherwise identical settings, which excludes explanations based on added capacity. We show the readout rate is also the corruption rate, predicting and confirming an interior optimum replicated across two implementations. We report one negative result with a mechanism: dual-copy representations lower the mutational error threshold, because gated expression is a selector rather than a joint decoder and therefore provides no coding gain. Finally, we document a benchmarking hazard: on dynamic benchmarks the choice of recombination operator alone shifted our baseline by 0.062 AUC, six times the effect size under study.2026-09-04T15:57:10Z6 pagesJia HuangYangjun Ouhttp://arxiv.org/abs/2609.05151v1Large Language Models with At Most One Spike per Neuron2026-09-04T13:53:28ZLeveraging their inherent sparse event-driven computation, spiking neural networks (SNNs) offer a promising path toward energy-efficient large language models (LLMs). Time-to-first-spike (TTFS) coding generates at most one spike per neuron within a time window, yielding extremely low firing rates. However, conventional TTFS SNNs are restricted to specific structures, making it challenging to encode certain blocks in LLM -- such as layer normalization and matrix multiplication --using TTFS. To overcome this limitation, we introduce a reference-based strategy specifically to encode the four core LLM components: embedding layers, layer normalization, attention-related operations and dropout. We construct a fully TTFS-based SNN architecture and train it end-to-end. Experiments on modern LLMs like BERT and GPT-2 demonstrate that our approach achieves performance comparable to ANN counterparts on natural language understanding and common-sense reasoning, while a clear gap remains on language modeling perplexity. To the best of our knowledge, this is the first work to scale a spiking LLM to 1.5 billion parameters using TTFS coding. We also report an estimate of spike-related energy; this is a spike-count proxy under an established cost model rather than a measurement on neuromorphic hardware.2026-09-04T13:53:28ZZhuoya ZhaoParsa OmidiAref JafariRichard Naudhttp://arxiv.org/abs/2609.05040v1Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications2026-09-04T12:01:53ZAs evolutionary transfer optimization (ETO) scales to larger collections of related tasks, problem evaluation can become a major source of runtime growth. This work studies problem-side evaluation scaling in task-parameterized applications and reformulates application-specific serial computations into forms suitable for parallel execution. We organize evaluation scaling into two levels: the number of evaluated tasks and the workload within each task. In multi-task optimization, matrix-recursive kinematic-arm evaluation is reformulated using an accumulation-matrix representation of cumulative link directions. In sequential transfer optimization, pointwise B-spline trajectory evaluation is reformulated using a blending-matrix representation for trajectory and collision computations. Both reformulations maintain close numerical agreement with their reference evaluations and substantially reduce runtime, yielding $256.72\times$ and $93.91\times$ end-to-end speedups, respectively. These results demonstrate problem-side reformulation as a practical route toward scalable ETO. Both application implementations and experimental scripts are released as open source to support reproducibility and reuse.2026-09-04T12:01:53ZAccepted at the 2026 International Conference on Machine Intelligence and Nature-Inspired Computing (MIND 2026)Yanchen LiXiaoming XueKay Chen Tanhttp://arxiv.org/abs/2602.13769v4OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Heuristic Design2026-09-04T02:55:55ZAutomating heuristic design in complex, experiment-driven domains requires more than iterative mutation of solution algorithms. Current LLM-based evolutionary methods often rely on stochastic mutation loops that lack long-term strategic planning and a formal mechanism to learn from historical failures, leading to inefficient exploration and redundant trials. To address this, we present OR-Agent, a multi-agent research framework designed for automated heuristic design in optimization problems with rich experimental environments. OR-Agent organizes heuristic search as tree-based workflow that explicitly models branching hypothesis generation and systematic backtracking. Furthermore, to address the lack of adaptive learning in current agents, we introduce a hierarchical, optimization-inspired reflection system in which short-term reflections act as verbal gradients, long-term reflections as verbal momentum, and memory compression as semantic weight decay - collectively forming a principled mechanism for governing research dynamics. Extensive experiments on classical combinatorial optimization problems (e.g., TSP, CVRP, bin packing) and simulation-based cooperative driving scenarios demonstrate that OR-Agent outperforms strong evolutionary search baselines. All code and experimental data are publicly available at https://github.com/qiliuchn/OR-Agent.2026-02-14T13:32:03ZQi LiuRuochen HaoCan LiWanjing Mahttp://arxiv.org/abs/2211.12337v4Quality-diversity in dissimilarity spaces2026-09-03T19:53:45ZThe theory of magnitude provides a mathematical framework for quantifying and maximizing diversity. We apply this framework to formulate quality-diversity algorithms in generic dissimilarity spaces. In particular, we instantiate and demonstrate a very general version of Go-Explore with promising performance.2022-11-14T16:34:07ZPatched Section 7 (not in the GECCO 2023 version at DOI 10.1145/3583131.3590409) with inline forward reference to https://arxiv.org/html/2509.19565 and https://proceedings.mlr.press/v321/huntsman26a.html, which contain the correct algorithm and proofs. No experimental results are materially affected in either this version or the GECCO oneSteve Huntsman10.1145/3583131.3590409http://arxiv.org/abs/2609.04195v1Axonal delay dispersion decides whether a neuron detects an event or a sequence, and predicts cortical column diameter2026-09-03T17:59:08ZCortical neurons fire sparsely -- often fewer than one spike per sensory window -- making rate coding insufficient and temporal coding a necessity. That conduction delays convert firing order into synchrony is long established. What governs which class of temporal feature a neuron detects -- one volley of coincident input, or two in a particular order -- has not been examined. We propose a delay-signature framework in which the axonal conduction delays converging on a dendritic branch constitute a physical key: only input sequences whose spike-time differences the delays compensate arrive synchronously, and coincidence detection, via calcium plateau thresholds, converts that synchrony into an all-or-none output. In simulations of an integrator-neuron model we report three results. First, a single physical scalar -- the dispersion of the delay set -- moves a population from event detection to order-selective sequence detection. The transition is emergent under random delays and connectivity: at narrow dispersion sequence detectors do not exist, and the dispersion at which they overtake event detectors tracks the inter-event interval with a slope statistically indistinguishable from one. This maps a computational distinction onto the anatomical one between myelinated and unmyelinated projections, making myelination a switch on what a neuron computes, not only a regulator of speed. Second, the same dispersion sets the code's limits: it bounds the longest codable interval and fixes an absolute timing tolerance of about a millisecond, with slowing better tolerated than speeding. Third, that millisecond window and horizontal conduction velocity together predict cortical column diameter, and the two areas with direct measurements fall where the relation puts them. One anatomically measurable parameter thus sets what a neuron detects and the limits of what it can represent.2026-09-03T17:59:08Z30 pages, 4 figures, under preparation for PLOS Computational BiologyCheng BiJipeng Sun