← Journal
NOTE 04·Methods

Keeping an auditable record

How we will connect each claim to its sources, implementation, measurements and limitations.

The journal should let someone follow a claim back to the work behind it. A citation, a screenshot and a reproduced experiment provide different kinds of evidence; we should say which kind is being offered.

A citation establishes what the source authors reported. A versioned package and an exact command make our implementation inspectable. A quantitative comparison supports a particular benchmark. An intervention can test whether a component affects the controller. Independent reproduction requires someone else to carry out the run.

Record where each choice comes from

For adopted parameters and connection rules, we will use five provenance categories:

Category What it records
Measured An observation, with the method and scope specified.
Derived A calculation from measured quantities, with the transformation retained.
Generated A sampled property or connection, with its distribution and rule.
Fitted A value estimated against a stated objective and dataset.
Engineered A choice introduced for the interface, encoding, decoding or operation.

These categories describe origin. They do not make measurements infallible or fitted models invalid. Their purpose is to keep the construction of the system visible.

Preserve experiments, including failures

Every experiment will receive its own record. Failed runs will not be overwritten by a successful rerun. A correction should identify the affected result, explain what changed and link to the new evidence.

Each result figure needs its units, condition, model version and sample size where applicable. A selected image is not a substitute for the complete planned set of runs. Raw outputs and the scripts used to derive summaries should be retained together.

Trace a controller decision

If the model becomes part of a controller, the record should connect an external input to an action request through explicit stages:

  1. Input snapshot, timestamp and validity checks.
  2. Encoder version and the resulting model stimulus.
  3. Model version, state or checkpoint, and seed.
  4. Recorded neural outputs.
  5. Decoder version and the requested action.
  6. The executor’s operating-limit decision and execution result.
  7. The resulting state checkpoint.

Replay reports will state whether they are exact or statistically equivalent on the chosen runtime. Credentials and private communications will stay outside published artifacts.

Public updates

Short updates will point to the relevant journal entry and its evidence. Proposed experiments will be described as plans until results exist. Before a result is published, its evidence must be accessible without access to our private working notes.

At present, the public record contains planning, literature selection and design work. No neural-model validation or controller-performance result has been produced.

Next step: apply this record structure to the first package audit and reproduction attempt.