# Felisyn benchmark 004 — reference and measurement audit

9 October 2026 (Europe/Vienna). Status: **measurement specification drafted; full-model comparison not executed; final equivalence criterion not frozen**.

## Question and scope

Can a reconstructed LSV1M environment reproduce the spontaneous-activity statistics in Figure 4 of Antolík et al. (2024), *A comprehensive data-driven model of cat primary visual cortex*, PLOS Computational Biology 20(8): e1012342? DOI: https://doi.org/10.1371/journal.pcbi.1012342 . This is model reproduction. Independent biological validation, reduced-model validity and causal controller utility require separate experiments.

## Reference data

`reference-means.csv` contains 24 means manually transcribed from the printed labels in the published Figure 4, checked against the authors' archived Figure 4. These are rounded, approximately two-significant-digit reference values, not recovered raw arrays. The voltage labels display magnitudes; negative signs are restored from the published axis and signed-voltage semantics. Exact SDs, neuron/pair sample counts, full-precision means and original spike/analog arrays have not been recovered. No SDs were estimated from bar lengths. The figure's SD bars are dispersion among neurons/pairs, not run-to-run confidence intervals or acceptance margins.

Published image: https://doi.org/10.1371/journal.pcbi.1012342.g004 . Archive run: https://v1model.arkheia.org/6723d916744974749499b1a6/results . Run date displayed: 22 March 2024; submission: 31 October 2024. Archive parameters: https://v1model.arkheia.org/6723d916744974749499b1a6/parameters . `sources/archive-parameters-visible.txt` retains only the expanded layer-four excitatory subsection and surrounding visible fields, not a complete parameter export. The figure-parameter dialog was empty. We did not find downloadable raw recordings in the views inspected; this does not establish that none exist elsewhere.

## Source pins and measurement definitions

- LSV1M v1.0: `04bf11b72ff5f78e66f28e037016e8e33b09c12e` at https://github.com/CSNG-MFF/mozaik-models/tree/04bf11b72ff5f78e66f28e037016e8e33b09c12e/LSV1M . This release postdates the paper.
- Mozaik v0.4.0: `f5c09a0eccc5786daf78881d883bbce586a00e6d`. Relevant source files and original CeCILL license are retained in `sources/`; the model's MIT license accompanies its source extracts.
- Four populations: V1_Exc_L4, V1_Inh_L4, V1_Exc_L2/3 and V1_Inh_L2/3. Keep their identities and neuron IDs separate throughout.
- Released `create_experiments_spont` uses a 40,320 ms NoStimulation experiment, versus the paper caption's rounded 40 seconds. The source settings use 50 cd/m² background, 0.1 ms integration and 1 ms cortical analog sampling. Read actual saved timestamps; do not silently trim to 40 seconds or delete a warm-up window. Select InternalStimulus and no direct stimulation; exclude terminal blank segments.
- A: single-unit spike counts divided by the complete segment duration in seconds, including recorded silent cells; mean and population SD across recorded cells.
- B: each included cell's ISI population SD divided by its mean ISI; then mean and population SD across cells. The plotting code requires more than five intervals, hence **at least seven spikes**.
- C: Pearson correlation of 10 ms histograms; use the same >=7-spike cohort as B, strict upper triangle only, no self-pairs or double-counting. For 40,320 ms there are 4,032 bins. Mozaik uses NumPy histogram semantics (rightmost endpoint included), normalizes bins to rates and applies `nan_to_num` to the correlation matrix. Count undefined pairs as an additional diagnostic. Rate scaling does not change Pearson r. Assert matching neuron-ID ordering before applying a mask to an upstream correlation matrix.
- D–F: temporal mean of each recorded analog signal, then equal-weight population mean and SD of those per-neuron means. Preserve signed mV. Convert conductances from µS to nS by multiplying by 1,000. For multiple trials upstream analog analysis averages per-neuron temporal means across trials first; the new helpers support a single window only.
- All cited SDs use NumPy's default `ddof=0`. No assumption of independent pairs is made: pairs share neurons.
- Panel G distributions are secondary exploratory checks; no lognormal-fit acceptance test is specified here.

## Configuration discrepancy and reference conflicts

The archived layer-four excitatory recorders agree with the released general `param/l4_exc_rec`: analog plus spikes on a 200 µm square with 10 µm spacing, and extra spike grids of 600/20 and 2,000/100 µm (size/spacing). The released `param_spont/l4_exc_rec` instead uses an analog 2,000/100 µm grid and a spike 3,000/20 µm grid. RCGrid chooses the nearest neuron to each grid point and deduplicates IDs; electrode counts are not neuron counts. We have verified this archive comparison for L4 excitatory cells only. Other archive populations and the complete original stimulation history remain to be reconciled.

The paper's general methods describe a central radius-1-mm spike sample and a selected 0.2 × 0.2 mm intracellular subset. The Figure 4 plotting class computes radius-0.5 masks but never uses those mask variables in its A–F summaries. `run_parameter_search.py` only adds a trial value; it does not resolve these sampling differences. The archive shows trial 0, while the released search script specifies trial 1. Matching seed numbers alone is therefore insufficient evidence of an identical run.

A second source disagreement: applying the plotting code's population weights (1, 0.25, 1, 0.25)/2.5 to the rounded inhibitory-conductance bars gives **4.09 nS**, while the paper's prose reports **3.56 nS**. The gap exceeds rounding of these four displayed labels. This audit does not establish the cause. Keep the panel-specific means as our declared reference; retain the prose result as a separate unresolved item. Do not choose whichever target a new run happens to match.

## Comparison policy — fixed before any full run

1. Report all 24 signed differences from the rounded reference, the number and identities of recorded/eligible neurons and pairs, undefined-pair counts, seeds, MPI rank/thread layout, durations, source pins and raw artifact hashes. Keep failed runs.
2. Two-significant-digit display agreement is a **descriptive check only**. It is not biological equivalence and is never an overall pass in `metrics.py`.
3. Do not use figure SD bars as tolerances. Exact raw-reference comparison requires the original arrays and a demonstrably matched configuration. Distribution-level equivalence requires a prospective, scientifically justified margin and independent-run uncertainty; these are not available yet. Neurons or pairs within one run cannot substitute for independent network seeds.
4. Resolve the reference configuration and define any sensitivity analysis before larger execution. Record deviations explicitly. A run of the released `param_spont` recipe would be a release-recipe experiment until its connection to Figure 4 is established.
5. Controller experiments remain downstream: predefined sensory inputs and output decoding, controls with disconnected/shuffled circuits and frozen baselines, recorded seeds and held-out tests. No live token actions are part of this benchmark.

## Implemented and checked

`metrics.py` provides independent spike and analog metric helpers and a descriptive rounded-reference comparison. `test_metrics.py` has 11 analytical synthetic edge-case tests, with receipt in `checks/`. They cover silent cells, eligibility, known CV, pair selection, undefined correlation, units, invalid input and preventing a rounding match from claiming a benchmark pass. These are mathematics/software checks, not new neural results. Datastore integration, end-to-end parity with upstream Figure 4, multi-trial support, timing alignment of analog inputs and full-model scaling remain untested.

Run locally with Python and NumPy:

```sh
python3 -m unittest discover -s . -p test_metrics.py -v
MPLCONFIGDIR=/tmp/felisyn-mpl python3 plot_reference.py
```

Plot generation also requires Matplotlib. The saved receipt records the actual local versions; the simulation runtime remains separate. Source arrays supplied to helpers must be aligned and explicitly associated with their population, stimulus, window and units. No synthetic trace is used as a research-result graphic.

## Compute planning

The authors report 16 processes on an EPYC 7302 with 128 GB RAM and about 1.5 hours for the spontaneous protocol. These are their setup and observed time, not a measured minimum or a guaranteed cloud runtime. The existing 8 GiB host is suitable for the small regression; we have no evidence it can run the full model. No full run or larger server was started in this step.

A candidate for later approval is a temporary DigitalOcean Memory-Optimized bundled instance with 16 vCPUs, 128 GiB RAM and 400 GiB disk: $1/hour, $672 monthly cap as listed on 9 October 2026. A six-hour infrastructure allowance would be $6 before tax and any extras; installation, debugging and CPU differences may exceed the paper's runtime. Region availability and the actual checkout price must be confirmed before provisioning. Powering off does not end billing; charges end on destruction. Preserve and verify outputs before authorized teardown. Sources: https://www.digitalocean.com/pricing/droplets and https://docs.digitalocean.com/products/droplets/details/pricing/ . Recommendation: resolve the reference mismatch first, then request approval for a bounded temporary compute run. No GPU is required by the cited CPU recipe.

## Attribution and provenance

The published Figure 4 and this adapted mean-only plot are credited to Antolík et al. (2024), CC BY; see the article's license statement and https://creativecommons.org/licenses/by/4.0/ . Adaptation: mean labels transcribed, voltage sign restored, layout/colors changed, SD bars and distribution panels omitted. Mozaik source files retain their CeCILL license; LSV1M files retain their MIT license. Neither the authors nor their institutions endorse this project. `manifest.json` hashes the public evidence files; its own hash and the ZIP hash are intentionally not self-referential.
