https://arxiv.org/api/ZJ/sQszUERgJhdY0e5IJFXwhZYY 2026-09-12T19:59:44Z 18249 15 15 http://arxiv.org/abs/2609.00449v2 Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach 2026-09-09T13:31:06Z Neuroevolution of Augmenting Topologies (NEAT) and its advanced version, Evolvable-Substrate HyperNEAT (ES-HyperNEAT), have shown great potential in developing neural networks. However, their effectiveness heavily depends on the selection of hyperparameters. This study investigates the optimization of ES-HyperNEAT hyperparameters using the Tree-structured Parzen Estimator (TPE) on the MNIST classification task, exploring a search space of over 3 billion potential combinations. TPE effectively navigates this vast space, significantly outperforming random search in terms of mean, median, and best accuracy. During the validation process, the best hyperparameter configuration found by TPE achieves an accuracy of 29.00% on MNIST, surpassing previous studies while using a smaller population size and fewer generations. The transferability of the optimized hyperparameters is explored in logic operations and Fashion-MNIST tasks, revealing successful transfer to the more complex Fashion-MNIST problem but limited to simpler logic operations. This study emphasizes a method to unlock the full potential of neuroevolutionary algorithms and provides insights into the hyperparameters' transferability across tasks of varying complexity. 2026-08-31T22:39:47Z Pages 1879 - 1887 GECCO 2024 Companion: Proceedings of the Genetic and Evolutionary Computation Conference Companion Romain Claret Michael O'Neill Paul Cotofrei Kilian Stoffel 10.1145/3638530.3664144 http://arxiv.org/abs/2608.15407v2 Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random Trees 2026-09-09T13:02:57Z Traditional honeypots share an invariant behavioral profile: a skilled adversary can confirm the presence of a deception environment within a few diagnostic commands, limiting their intelligence value. Commercial deception products (USD 100,000-150,000/year) similarly lack real-time model-driven feedback. Chameleon, an openly distributed adaptive honeypot, addresses both shortcomings. It integrates: a BiLSTM classifier achieving 99.61% accuracy across seven threat categories at ~2 ms CPU latency; a locally deployed Qwen3.5-0.8B model delivering 90% generation accuracy at 4.5 ms latency; and two meta-heuristic engines. Threat-Calibrated PSO (TC-PSO) reshapes swarm inertia and objective amplification in proportion to the classifier's anomaly output, adjusting connection-holding delays in real time. Semantic Deception RRT (S-RRT) evolves deception schemas via exponentially scaled pheromone updates from a language-model severity assessment, with a depth-decay multiplier enforcing a finite memory footprint. A controlled 30-seed benchmark (42-71, identical trajectories and budgets) shows threat-calibrated inertia alone does not improve search over standard PSO on static or dynamic landscapes (p = 0.18); population-diversity mechanisms (GA/ACO) significantly outperform PSO-family optimizers on threat-regime shifts (p < 0.0001, d <= -37). S-RRT's depth-decay delivers a significant memory reduction versus standard RRT (53.1 vs. 119.2 units, p < 0.0001, d = -10.0); its severity-weighted pheromone does not improve raw fitness. Operating cost is ~USD 17/month, a ~490-fold reduction versus commercial alternatives. 2026-08-15T20:30:14Z 10 pages, 7 figures, 2 tables. Under consideration for journal publication. MIT-licensed code and datasets: https://github.com/RohitSwami33/Chameleon-cybersecurity-ml Rohit Swami Tushar Singh Akash Warde Sri Muthu http://arxiv.org/abs/2609.09595v1 Teacher Geometry Shapes Learnability in Teacher-Student Networks 2026-09-09T01:44:23Z Teacher-student systems, in which a teacher neural network generates training labels so that a student neural network can learn to implement the same function, are widely used as an abstract setting to study learning. However, the structure of the teachers is often overlooked by assuming randomly-generated, normally-distributed parameters. This hides substantial variation in how learnable different teachers are. We formalize learnability as the success rate of converging to the global minimum, as a function of overparameterization, learning algorithm, student initialization distribution, and teacher geometry. We both identify an easy distribution that maximizes node dissimilarity and a hard distribution that minimizes it, and show that these two distributions induce markedly different success rates across a large range of settings and for different activation functions. To explain the gap, we study the loss landscape of small neural networks that contain two distinct kinds of suboptimal local minima, out-of-bounds (OOB) minima at the edge of the data distribution and interior minima within. Assuming infinite data and a fast readout layer, we analytically reduce the loss landscape of small networks to two dimensions, showing that the region of attraction of interior minima changes as a function of teacher structure. In larger networks, maximally dissimilar teachers induce more interior minima, while minimally dissimilar teachers induce more OOB minima. Motivated by these analyses, we show that differentially increasing the learning rate of the readout layer and decreasing the learning rate of the inner biases increases success rates. These findings provide an important step in narrowing the gap between the study of teacher-student networks and more structured functions that arise in practice. 2026-09-09T01:44:23Z Kai J. Sandbrink Flavio Martinelli Alexander van Meegen Wulfram Gerstner Johanni Brea http://arxiv.org/abs/2605.03338v2 Symmetry, Defects, and Diffusion in Continuous-memory Recurrent Networks 2026-09-09T01:10:07Z Continuous-memory recurrent networks must preserve phase, position, or orientation despite model imperfections and state noise. We develop a geometric framework that separates three questions: how many memory coordinates are neutrally transported, how deterministic perturbations alter their finite-horizon stability, and how ambient noise is decoded along them. Exact per-input equivariance transports analytical group tangents pathwise and, on a compact nondegenerate orbit stratum, yields at least $q=\dim(G/H)$ zero group-tangent Lyapunov exponents under stationary ergodic driving. For imperfect dynamics, a four-block tangent/normal decomposition gives local and finite-horizon bounds on tangent growth and subspace rotation, distinguishing first-order direct damage from second-order leakage through contracting normal directions. For noisy dynamics, a specified decoder maps ambient covariance $Q$ to coordinate covariance $ZQZ^\top$; under isotropic noise and fixed tangent energy, least-squares decoding and scaled-isometric action geometry minimize local diffusion. Local and finite-horizon evaluations include cases both within and outside the sufficient conditions. In a fresh twenty-seed $T^2$ replication, a decoder-covariance objective improves noisy horizon-256 memory in every pair while meeting a prespecified clean-error equivalence margin. Direct noise training also improves noisy memory but incurs a clean-error tradeoff. In coupled $T^4/T^8$ integrators, the advantage of anisotropic-covariance over isotropic regularization reverses when evaluation noise becomes isotropic. These results connect continuous symmetry to measurable limits and design choices for recurrent memory under explicit dynamical, decoder, and noise assumptions. 2026-05-05T03:59:56Z Hanson Hanxuan Mo http://arxiv.org/abs/2609.09564v1 Robust Industrial Cyber Physical Classification Using Neuromorphic Temporal Embeddings and Hybrid SNN XGBoost Under Machine Unlearning Attacks 2026-09-09T00:43:44Z The digitalisation of electrical distribution networks has increased the exposure of power-grid infrastructure to cyber attacks. Existing intrusion detection systems (IDSs), however, often rely on computationally expensive deep learning models that are difficult to deploy at the edge. Periodic retraining also exposes these systems to machine unlearning attacks, where selective data removal can degrade detection performance. We propose a hybrid Spiking Neural Network (SNN) and XGBoost architecture that combines efficient temporal encoding with a lightweight classifier and provides structural resilience to such attacks. The SNN is trained once on clean data and used as a fixed feature extractor, while only the XGBoost classifier is retrained during model updates. Evaluated on two real-world public power-system datasets, the proposed method achieves 99.9\% accuracy (F1-macro 0.999) on the Synchrophasor dataset and 95.0\% accuracy (F1-macro 0.943) on the MSU/ORNL dataset, outperforming standalone baselines. Under selective label-flipping attacks, the hybrid model loses only 0.9\% F1-macro at 10\% poisoning and delays target-class collapse from 60\% to 70\% poisoning compared with raw models. These results demonstrate that neuromorphic temporal encoding can provide both accurate cyber-attack detection and improved resilience to data poisoning in cyber-physical systems. 2026-09-09T00:43:44Z 13 pages, first draft Ammar Kamoona Sajad Koushkbaghi Mahdi Jalili Peter McTaggart Xinghuo Yu http://arxiv.org/abs/2407.18897v3 Small Molecule Optimization with Large Language Models 2026-09-08T20:13:57Z Molecular optimization, the process of designing molecules with desirable properties, represents a critical challenge in drug discovery. The recent advancements in large language models (LLMs) have opened new opportunities for their integration with traditional molecular optimization algorithms to improve performance. In this work, we propose Molecular Language Model powered Evolutionary Algorithm (Mol-E), an evolutionary algorithm that relies on the generative capabilities of LLMs trained on molecules and molecular properties. Scientific Contribution. Mol-E obtains the highest aggregate Top-10 AUC among the comparable full-23-task results considered here, scoring 17.500 in the task-agnostic regime, in which the oracle is treated strictly as a black box, and 20.551 in the task-informed regime, in which the optimizer receives a fixed semantic description of the objective. Mol-E also improves over the evaluated baselines on multi-property optimization with docking against DRD2, MK2, and AChE. 2024-07-26T17:51:33Z 35 pages Philipp Guevorguian Menua Bedrosian Tigran Fahradyan Gayane Chilingaryan Armen Aghajanyan Hrant Khachatrian http://arxiv.org/abs/2507.10005v3 Effects of relational graph modularity and depth on the learning performance of neural networks 2026-09-08T18:47:35Z In recent years, graph-based machine learning techniques, such as reinforcement learning and graph neural networks, have garnered significant attention. While some recent studies have started to explore the relationship between the graph structure of neural networks and their predictive performance, they often limit themselves to a narrow range of model networks, particularly lacking mesoscale structures such as communities. Our work advances this area by conducting a more comprehensive investigation, incorporating realistic network structures characterized by heterogeneous degree distributions and community structures, which are typical characteristics of many real networks. These community structures offer a nuanced perspective on network architecture. Our analysis employs model networks such as random and scale-free networks, alongside a comparison with a biological neural network and its subsets for more detailed analysis. We examine the impact of these structural attributes on the performance of image classification tasks. Our findings reveal that structural properties do affect performance to some extent. Specifically, within moderate-depth architectures, networks featuring coherent, densely interconnected communities demonstrate enhanced learning capabilities. Crucially, we find that this advantage is strictly depth-dependent: extending the architecture to eight layers reverses the effect entirely. This comparison with the biological neural network emphasizes the relevance of our findings to real-world structures, suggesting an intriguing connection worth further exploration. This study contributes meaningfully to network science and machine learning, providing insights that could inspire the design of more biologically informed neural networks. 2025-07-14T07:39:19Z 12 pages, 7 figures Yash Arya Sang Hoon Lee http://arxiv.org/abs/2609.09306v1 Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions 2026-09-08T18:01:23Z This paper investigates the hypothesis that the first-order structure of physical interactions, i.e. gradients or Jacobians, characterizes the structure of phenomenal experience. It does so in an idealized world inhabited by neural networks, Gradland, where the physics are known and the functions are (mostly) differentiable. The paper introduces two measures of Jacobian structure: effective rank and cohesion, based on Kirchhoff complexity. Applying the measures to a series of worked examples shows the hypothesis accounts for: (1) the duration of experience, that it can prolong over hundreds of milliseconds; (2) the difference between what is experienced vividly and obscurely; (3) the experience of texture; (4) the blooming buzzing confusion presumably experienced by newborns; (5) the difference between ideas that are held distinctly in mind and ideas that are confused; (6) what learning is like; and finally (7) the paper explains the function of rich, dense experience. 2026-09-08T18:01:23Z Code: https://github.com/dbalduzzi/gradland/ David Balduzzi http://arxiv.org/abs/2507.10383v5 Dynamical stability for dense patterns in attractor neural networks 2026-09-08T14:32:55Z Recurrent neural networks are canonical models of biological memory. In these models, memories are represented by distributed patterns of neural activity that are stored in the recurrent connections between neurons, such that they become attractors of the network's dynamics. During memory recall, network dynamics thus converge toward one of these memory patterns when started from a noisy or partial cue. Therefore, memory performance critically hinges on the dynamical stability of the stored patterns. However, previous theoretical approaches only studied dynamical stability under highly restrictive conditions that do not readily apply to biological neural circuits. Here, we develop a theory of the local stability of discrete fixed points in a broad class of networks with graded neural activities and in the presence of noise. Using methods from random matrix theory, we analyze the bulk and outliers of the eigenvalue spectra of the Jacobians that characterize network dynamics around fixed points. We show that either all fixed points are stable or all of them are unstable, depending on whether their number is below a ``critical load for stability'', which is distinct from the classical critical capacity that measures the maximal number of achievable fixed points regardless of their stability. We further analyze the dependence of this critical load for stability on experimentally measurable quantities characterizing the statistics of memory patterns and the activation functions of neurons. Our analysis highlights the computational benefits of sparse-like patterns and threshold-linear activation functions and offers testable predictions for neural circuits supporting memory. 2025-07-14T15:23:24Z Uri Cohen Máté Lengyel http://arxiv.org/abs/2609.08615v1 Why shared attention vectors fail: a case for outcome-indexed tuning 2026-09-08T11:52:08Z Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single scalar. These scalars are learned by models through gradient-descent on error, where predictive features acquire more salience. We show that under multi-outcome learning, where models predict more than one outcome, this shared vector becomes unstable; it collapses to its bounds and prevents the models from learning meaningful attentional tunings for learning and generalization. We address this by introducing an outcome-indexed attentional matrix that converts globally shared attentional tuning into an outcome-indexed representation. We present an analysis of the unstable shared vectors and derive the conditions under which it holds. Empirically, three synthetic experiments benchmark the proposed attention matrices and show that they converge to meaningful representations, something shared attention vectors fail to do. These results suggest that outcome-indexed attentional matrices are a general fix for gradient-based attentional processes, which improves models of learning under multi-outcome conditions. 2026-09-08T11:52:08Z 2 figures, 8 pages Lenard Dome http://arxiv.org/abs/2601.12723v3 An Evolutionary Framework for Automatic Optimization Benchmark Generation via Large Language Models 2026-09-08T00:19:33Z Optimization benchmarks play a fundamental role in assessing algorithm performance; however, existing artificial benchmarks often fail to capture the diversity and irregularity of real-world problem structures, while benchmarks derived from real-world problems are costly and difficult to construct. To address these challenges, we propose an evolutionary automatic benchmark generation framework that leverages a large language model (LLM) as a generative operator, termed the LLM-driven evolutionary benchmark generator (LLM-EBG). In this framework, the LLM serves as an evolutionary operator that generates and evolves benchmark problems within a flexible, expressive representation space. As a case study, we generate unconstrained single-objective continuous minimization problems represented as mathematical expressions designed to induce significant performance differences between a genetic algorithm (GA) and differential evolution (DE). Experimental results show that LLM-EBG successfully produces benchmark problems in which the designated target algorithm consistently outperforms the comparative algorithm in more than 80\% of trials. Furthermore, exploratory landscape analysis reveals that benchmarks favoring GA are highly sensitive to variable scaling, demonstrating that the proposed framework can generate problems with distinct geometric characteristics that reflect the intrinsic search behaviors of different optimization algorithms. 2026-01-19T04:58:15Z Yuhiro Ono Tomohiro Harada Yukiya Miura http://arxiv.org/abs/2606.29655v2 Geometric Reliability of Neural Population Codes: Sampling Calibration and Within-Session Nonstationarity 2026-09-07T22:24:06Z Trial-to-trial variability limits how reliably neural population geometry can be estimated, while comparisons across populations depend on neuron and trial counts, response quality, and clustered sampling. We quantified within-session geometric reliability using Shesha, the Spearman correlation between representational dissimilarity matrices estimated from independent trial subsets, in all 39 Steinmetz Neuropixels sessions and in olfactory bulb and piriform cortex recordings from Bolding and Franks. Steinmetz analyses matched neurons and repetitions, compared observed reliability with a stationary residual-bootstrap expectation, and used mouse-level or mouse-clustered inference. Mean matched reliability was 0.0402 across 312 area-by-session recordings. Regional differences and reliability above the stationary benchmark did not survive correction. Temporal effects received the strongest support: interleaving early and late trials increased reliability relative to blocked allocation ($Δ=0.02666$, $q=0.001953$), and RDM similarity declined with within-session lag (mean mouse-level slope $=-0.01912$, $q=0.001953$; $n=10$ mice). Outer-cross-fitted reliability was not associated with choice-direction coupling or stimulus or response-direction decoding after correction. Olfactory comparisons remained descriptive because few paired sessions and no animal identities were available. In held-out simulations, associative recurrence outperformed feedforward subspace denoising but not divisive normalization. Representational geometry became less reproducible with temporal separation within a session, and comparisons across neural populations require sampling calibration and independent inference. 2026-06-28T23:54:20Z Prashant C. Raju http://arxiv.org/abs/2606.03026v2 Spike-Aware INT8 Execution for Spiking Language Models on Commodity CPUs 2026-09-07T16:38:06Z Binary spike activations allow a language-model runtime to read only active weight columns and replace multiplications by weight sums. We implement this execution strategy in C++ for an 874M-parameter spike-gated language model. Sparse projections use column-major INT8 weights, integer accumulation, and one scale application per output channel; dense projections retain row-major access and FP32 activations. In a single-thread comparison using an early checkpoint, INT8 achieves 23.31 tokens/s versus 9.82 for FP32, while reducing weight storage from 3355.2 to 1087.4 MiB. A variant using INT4 on dense projections saves a further 17.4% of storage but reduces decode throughput by 46.6%. On an AMD Ryzen 7 5800X, the final INT8 checkpoint achieves 22.63 tokens/s on one thread and 47.90 on four threads; 512-token prefill reaches 94.68 tokens/s on eight threads. A separate ARM output-head case study records higher trimmed decode-window energy metrics for two candidate-verification configurations. The results characterize how activation-specific layouts and quantized kernels support CPU deployment of a spike-gated language model. 2026-06-02T02:03:37Z 13 pages, 3 figures, 11 tables. Revised presentation and evaluation; added numerical validation and an ARM wall-power case study; withdrawn invalid pure-INT4 results and clarified comparison protocols Ting Liu http://arxiv.org/abs/2609.07689v1 Emergent Charging Coordination in Electric Delivery Fleets 2026-09-07T16:09:13Z In electric delivery fleets, mid-shift charging is non-trivial: each vehicle must decide when, where and how much to charge to finish on time with battery above a safety floor. The choices are coupled: queues build where too many vehicles pick the same station. Prior work resolves this coupling with central dispatching, precomputed schedules or reservations, machinery that charging infrastructure rarely supports. Instead, we use a family of learning agents under purely local control: every vehicle runs the same policy, deciding alone from its time budgets and broadcast station occupancies, leading to emergent coordination without central control or messaging. We validate this paradigm in simulation on real OpenStreetMap networks of twenty cities, each with a frozen scenario calibrated by an omniscient Oracle (99.5% of shifts completed on time), whereas a naive greedy rule (nearest station on low battery) completes just 73%. Agents trained with neuroevolution (NEAT) and policy gradients (PPO) on four cities and deployed zero-shot across all twenty, sixteen never seen in training, complete 96.8% and 98.6% of shifts, with the policy-gradient controllers proving more robust when demand or vehicle characteristics drift beyond the trained regime. In contrast, tuned threshold heuristics that read vehicle urgency alone fall short in contended cities (~80%). Through training, these learning agents rediscover partial charging and short opportunistic sessions, and route around busy stations, cutting per-session queue waits from about 45 minutes to under 2. In summary, this coordination paradigm balances local urgency against public occupancy, reaching near-Oracle performance at minimal implementation cost. 2026-09-07T16:09:13Z 42 pages, 11 figures. Submitted to Transportation Research Part C Javier Vales-Alonso Juan J. Alcaraz http://arxiv.org/abs/2508.20850v3 Encoding Tactile Stimuli for Braille Recognition with Organoids 2026-09-07T14:22:53Z This study proposes a transferable encoding strategy that maps tactile sensor data to electrical stimulation patterns, enabling neural organoids to perform an open-loop artificial tactile Braille classification task. Human forebrain organoids cultured on a low-density microelectrode array (MEA) are systematically stimulated to characterize the relationship between electrical stimulation parameters (number of pulse, phase amplitude, phase duration, and trigger delay) and organoid responses, measured as spike activity and spatial displacement of the center of activity. Implemented on event-based tactile inputs recorded from the Evetac sensor, our system achieved an average Braille letter classification accuracy of 61% with a single organoid, which increased significantly to 83% when responses from a three-organoid ensemble were combined. Additionally, the multi-organoid configuration demonstrated enhanced robustness against various types of artificially introduced noise. This research demonstrates the potential of organoids as low-power, adaptive bio-hybrid computational elements and provides a foundational encoding framework for future scalable bio-hybrid computing architectures. 2025-08-28T14:44:25Z Tianyi Liu School of Engineering Mathematics and Technology, University of Bristol, United Kingdom Hemma Philamore School of Engineering Mathematics and Technology, University of Bristol, United Kingdom Benjamin Ward-Cherrier School of Engineering Mathematics and Technology, University of Bristol, United Kingdom