What the proof shows,
and what it does not
Verifiable neural-network inference in halo2. A model owner publishes a commitment to the weights; a prover then produces a proof that a given output came from the model behind that commitment. This is a write-up of what that establishes, what the commitment gives away for free, and what proving costs at four measured circuit sizes.
It is also a write-up of a defect the system still has. That part is not buried.
01The system
The demo is a two-layer MNIST MLP, 784 inputs, 128 hidden units, 10 output logits,
quantized to 8-bit with per-tensor power-of-two scales. The circuit does matmul plus
bias, requantize, and ReLU on the hidden layer, all inside a halo2 statement, over the
Pasta curves. Alongside the arithmetic the
circuit recomputes a Poseidon digest over the weights and binds it to the statement,
so the proof is not just "some network produced this output" but "the network whose
digest is D produced this output".
The public instance is D ‖ x ‖ z: the model root, then every
element of the input, then every output logit. A verifier is handed all three. The
weights, the biases, the hidden activations and the requantize remainders are witness
data and are not in the instance.
The quantized model records an accuracy of 0.9703 at hidden width 128, with every quantization scale fitted on the training split alone and the accuracy measured on the untouched test split. The proof is 10,880 bytes and verifies in a median of 169.0 ms since the fold restructure of 2026-08-12. Section 5 has the rest of the numbers, the machine they came from, and what the restructure changed.
02The commitment is a confirmation oracle
D is a deterministic Poseidon fold over the packed weight chunks and the
biases, seeded from a header derived from the layer shape and the shift. There is no
salt, no nonce and no randomness anywhere in it. The same weights always produce the
same D.
That makes D a confirmation oracle for the weights, and the oracle runs
entirely outside the proof system. Anyone holding a candidate model recomputes the
fold and compares one field element. No proof is involved, no interaction with the
prover is needed, and no improvement to the proof system touches it. Guessing a
784-128-10 model from nothing is not the threat. Confirming a model you already
suspect is, and confirming is exactly the operation an unsalted commitment offers for
free.
The reason that matters in practice is that a realistic attacker does not search a parameter space. They enumerate a short list and test it. Deployed models are frequently fine-tunes of a public open-weight base, quantized onto a standard grid, which makes the candidate set small enough to write down: a handful of bases, a handful of checkpoints, the obvious quantization settings. Checkpoints also leak by ordinary means, through a departing employee, a misconfigured bucket, a breach, or discovery in litigation, and a party holding one wants to know whether it is the model behind a published identifier. Every one of those questions is answered by one fold and one comparison, per candidate, with no access to the prover and no cooperation from anyone.
This is a consequence of a recorded design decision rather than a surprise: the digest was always analysed as a binding commitment and never as a hiding one. It is called out here because it is the load-bearing fact for anyone reasoning about model confidentiality on top of a scheme like this, and because it survives every fix to the proving system underneath it.
On naming it
The attack shape has established names outside machine learning. It is the confirmation-of-a-file attack, first analysed by Douceur et al. (2002) and documented by Wilcox-O'Hearn, Perttula and Warner in 2008 on the tahoe-dev list:
"Convergent encryption is already known to suffer from a confirmation-of-a-file attack. We show that it suffers also from a learn-partial-information attack."
The formal frame is Message-Locked Encryption (Bellare, Keelveedhi and Ristenpart, EUROCRYPT 2013, eprint 2012/631), whose governing limitation is that MLE provides security only for unpredictable messages. An unsalted hash of model weights is an MLE tag, and the confirmation attack transfers directly.
A literature survey run for this project found no published work naming the confirmation-of-a-file attack, in that established terminology, in the context of ML weight commitments, and reported that the convergent-encryption vocabulary has not been formally transferred to the zkML literature as of mid-2026. That is one survey pass, and its underlying evidence ledger was never retrievable. Absence of evidence in a single survey is not proof of novelty, so this is not a priority claim. It is stated at exactly the strength the evidence carries: the standard names exist, they fit, and we would rather use them than reinvent the description.
03Two locks, and only one of them is shut
Hiding model weights from a verifier needs two independent things. Fixing either one alone buys nothing.
Lock 1: a commitment that hides as well as binds
D would have to absorb randomness, so that publishing it does not let a
holder of a candidate model confirm it. That is a design change and not a parameter.
The circuit has to witness the commitment randomness alongside the weights, and every
operation D currently supports for free needs a replacement. "Does your
model match the approved one" is today a comparison of two field elements; against
hiding commitments it becomes an equality-of-committed-value proof between two
parties. In this system that lock is open.
Lock 2: fresh randomness at proving time
halo2 claims perfect special honest-verifier zero knowledge for its protocol, in its
own design document, under stated preconditions. Two of those preconditions are a
freshly random blinding factor for every group element in the transcript, and each
advice polynomial blinded with n_e + 1 random evaluations over the
domain. Both are supplied by one thing: the RNG the caller passes to
create_proof. The RNG feeds blinding and nothing else. There is no
correctness or soundness role for it anywhere in the prover. The caller's RNG is the
hiding property.
Three call sites in this project passed a fixed seed. Under a fixed seed a proof is a deterministic function of the witness, which was measured rather than argued: two proofs of the same statement from the same witness came out byte-identical, and a party who knows the input and the shape and guesses the weights correctly re-proves and gets identical bytes. So the proof itself was a second confirmation oracle, slower than the digest but on the same weights.
That is fixed. The three call sites now draw their blinding from the operating system's entropy source, and a regression test fails if a fixed seed returns. The cost was paid knowingly: the committed demo proof can no longer be regenerated bit for bit, so a reader can still produce a valid proof of the same statement but not the same bytes.
This system does not hide model weights from a verifier, and no claim here should be read as saying it does. Lock 2 is in place. Lock 1 is open, so the published digest still confirms any candidate model for the cost of one fold. Separately, fresh randomness is not itself evidence of hiding: determinism is observable, and its absence is not a proof of zero knowledge, which is a claim about distributions and a simulator that no number of executions decides. Nothing has verified that a proof from this system conceals its witness, and the project does not claim it does.
Two further facts bound what the locks could ever deliver here, and they are the shape
of the statement rather than defects in it. The input x is a public
input, so concealing it is a different statement, not a stronger proof of this one.
The logits z are public, so a verifier who can obtain proofs on inputs of
its choosing learns the model's outputs on those inputs by design. A hiding commitment
defends against confirming a candidate model. It does not defend against extracting
one by querying, and this architecture publishes the query and the answer with every
proof.
04The gap is industry-wide
Lock 1 is not a defect peculiar to this project. The same survey found deterministic,
unsalted commitments across the surveyed zkML systems. The clearest documented case is
EZKL, whose kzgcommit mode uses a constant blinding factor of
1, which is no blinding, and whose Poseidon path under
--param-visibility "hashed" is also deterministic and unsalted. Nobody
holds a salt in either mode, because there is none. EZKL's own engineering blog states
the consequence:
"without blinding, and with a sufficient number of proofs, an attacker can more easily recover a non-blinded column's pre-image. Each proof generates a new query and evaluation of the polynomial... if M is equal to the degree + 1 of the polynomial... the column's pre-image can be recovered."
blog.ezkl.xyz/post/commits/, October 2023
Coslett (2026), surveying the major zkML systems, puts the more general version of the problem this way: "a weight commitment is not a model identity. A prover can commit to arbitrary weights, execute them honestly, and still prove the computation correctly."
None of that makes our own gap smaller. That a whole field shares a limitation says nothing about whether a given system describes itself correctly, which is the only thing worth judging a write-up on. It is context for how hard the problem is, and a reason to read anyone claiming to have solved it closely.
05What proving costs
Until this sweep, every statement about what a larger model would cost rested on one
measured point plus an assumption. Five shapes were then measured across four values
of the circuit size parameter k, ten prove and verify runs each, one
process per shape.
Everything in this section is scoped to what was run: this architecture, this prover,
k = 15 through 18, on one machine, at n1 = 784
and m2 = 10 with hidden width the only shape parameter varied. None of it
is a claim about proving systems in general, and none of it is a claim about circuit
sizes nobody here has run.
Updated 2026-08-12. The table and charts below are the record of
the run that produced them, against the digest fold this system shipped with. On
2026-08-12 that fold was restructured for cost, which removed enough rows to drop
the committed 784-128-10 model from k = 18 to k = 17.
Measured over ten runs on the same machine: prove median 36.9 s, verify median
169.0 ms, peak 8.18 GiB, proof 10,880 bytes. The reason to keep the old table in
place: the restructured circuit landed within 0.05 percent of where the table's own
memory law puts k = 17, so the law, not luck, is doing the predicting.
Everything below stands as written, one step of k above where the
committed model now sits.
Hardware: 11th Gen Intel Core i5-11600K at 3.90 GHz, 12 logical cores, Fedora Linux 44 (Workstation Edition), 31 GiB RAM. The machine was not quiesced. Other work was running throughout, and the spread inside a single shape shows it: the 784-20-10 shape has a prove minimum of 7.24 s and a maximum of 15.26 s over ten identical runs. Read the medians and minima; treat the maxima as contention rather than as circuit cost.
| shape | k | rows | prove median | prove min | verify median | peak | proof |
|---|---|---|---|---|---|---|---|
| 784-20-10 | 15 | 32,768 | 8.56 s | 7.24 s | 46.3 ms | 2.07 GiB | 10,752 B |
| 784-24-10 | 15 | 32,768 | 7.49 s | 7.38 s | 45.5 ms | 2.05 GiB | 10,752 B |
| 784-32-10 | 16 | 65,536 | 15.16 s | 15.09 s | 84.7 ms | 4.10 GiB | 10,816 B |
| 784-112-10 | 17 | 131,072 | 37.11 s | 34.49 s | 163.7 ms | 8.18 GiB | 10,880 B |
| 784-128-10 | 18 | 262,144 | 70.41 s | 70.06 s | 313.4 ms | 16.35 GiB | 10,944 B |
Medians over ten runs. Peak is the whole process high-water mark and so covers the keygen peak as well as proving; the underlying figures are 2,166,636 kB, 2,149,500 kB, 4,294,844 kB, 8,582,516 kB and 17,147,424 kB, where kB is the kernel's own unit for VmHWM, which is KiB. Keygen is deliberately absent: it is measured once per shape, so every reading is a sample of size one, and it should be treated as unmeasured for curve purposes.
The same numbers as the table, drawn. Both axes of cost double with each step of
k, so on a log scale the measured points sit on a line; every gridline
crossed is a doubling. The two markers at k = 15 are the width control:
20 percent apart in hidden width, 0.8 percent apart in peak memory. The open point at
k = 19 is extrapolation from the fit, not a measurement: about 32.8 GiB
against the 18.7 GiB the sweep recorded as available, which is why the dashed line is
where measurement stops on this machine.
Cost is a step function of k, not a curve in width
The two shapes at k = 15 differ by 20 percent in hidden width and by 0.8
percent in peak memory: 2,166,636 kB at width 20 against 2,149,500 kB at width 24. The
narrower model used the larger peak of the two, so within this band peak memory does
not track width at all. Their prove minima are 7.24 s and 7.38 s, 1.9 percent apart,
against a 20 percent difference in width. That is one within-band pair, deliberately
included as a control, and it is the whole of the direct evidence for the paragraph
below.
Across the shapes measured, widening a model costs nothing until the row count crosses
a power of two, and then costs a factor of two at once. So for sizing work on this
circuit, the quantity to reason about is k rather than parameter count or
width, and an argument phrased in width would have missed the step entirely.
The fitted relationships, over k = 15 to 18 only
Peak resident set is very nearly exactly 2^k times a constant. Dividing
each measured peak by its row count gives 66.12, 65.60, 65.53, 65.48 and 65.41 KiB per
row across the five shapes, a spread of 1.1 percent:
peak ~ 65.5 KiB x 2^k(measured k = 15 to 18, within 1.1%)
Prove time grows slightly faster than a doubling. On minima, which are the readings
least contaminated by other load, the geometric mean per-step ratio is 2.13 anchored on
784-20-10 and 2.12 anchored on 784-24-10. A pure doubling law anchored on the
k = 15 minima under-predicts the measured k = 18 minimum by
16 to 17 percent depending on the anchor. That is the signature of an
n log n term on top of the linear one, which is what an FFT-based prover
should show:
prove ~ 7.3 s x 2.12^(k - 15)(measured k = 15 to 18, on minima)
Verification is not flat and not cheap in the limit. It grows by a factor of 1.89 to
1.90 per step, with individual steps of 1.83, 1.93 and 1.92, doubling alongside the
prover from 45.5 ms to 313.4 ms over three steps. Proof size grows by exactly 64 bytes
per step in k: 10,752, 10,816, 10,880, 10,944 bytes. That is one
additional pair of group elements per IPA round, and it is the only cost here that is
genuinely logarithmic.
The ceiling
k = 18 is the measured ceiling. Anything above it is extrapolation.
The k = 18 shape peaks at 16.35 GiB against the 18.7 GiB of
MemAvailable this sweep recorded on a 31 GiB machine, so k = 19 at the
measured 65.5 KiB per row would need about 32.8 GiB and would swap rather than run.
Every fitted relationship above is anchored at the top of its own range and validated
only from below. The memory law holding to one percent over three steps downward is
evidence about the shape of the curve below k = 18, not a licence to
project it upward past a ceiling nobody has tested.
The useful reading is the negative one. On commodity hardware this prover is out of memory one step above the model it already proves. A 784-128-10 MLP is not a small model chosen for convenience; it is at this machine's limit, and the next model up the ladder needs a different machine or a smaller circuit.
06Where this goes
Once a commitment hides, a verifier can no longer derive it, so the belief that some
published D is the approved model has to come from somewhere else. In the
deployed world it comes from a signature by a party who saw the weights. A survey
looking for any deployed system where it comes from anything else returned a clear
negative: every one rests on a signature from the owner or operator who saw the
weights, a TEE attestation where the TEE saw them, or reputation.
That infrastructure already exists and it ships. sigstore/model-transparency reached v1.0 in April 2025 under OpenSSF, with Google, NVIDIA and HiddenLayer named as participants: a signature plus an append-only transparency log, which is Rekor's Merkle tree with signed tree heads doing anti-equivocation. A bare hash registry is therefore not worth building here; that part of the problem already has shipped implementations with vendor names on them.
What that stack does not do is the interesting part. It authenticates who signed, not what they verified about the model; Sigstore says so itself and does not verify model quality, correctness or safety. Every commitment in the surveyed set is deterministic, so nothing there ties a signature to a commitment that hides. And nothing there ties a signed commitment to an inference proof: the signature says a named party vouched for an artifact hash, and says nothing about any output that artifact later produced. Those last two are absences reported by a single survey, which is the weakest form of evidence there is, and anyone building on the gap should check it directly first.
Signed commitments tied to per-inference proofs, over a commitment that hides, is the seam. It is also where the honest limits above have to be carried forward rather than quietly dropped: the accreditor pattern relocates trust to a signer, and the failure mode that matters is the signer who signs without verifying. The mitigation for that is separation of duties and policy, which is process, not cryptography. A design that adopts the pattern inherits that exposure knowingly. So does a design that relies on a transparency log, since Rekor's anti-equivocation works only if third parties actually monitor the logs.
This is risk management, not compliance
Worth stating because the opposite is assumed so often. The survey found no binding or quasi-binding regulatory text requiring cryptographic model identity, and reported that the closest language it turned up is the EU AI Act's Annex VIII field 4, an "unambiguous reference allowing identification and traceability", which on its reading a version number with a product name satisfies. It read SR 11-7 and the OCC guidance as requiring a model inventory, validation and ongoing monitoring while naming no mechanism, and NIST's AI RMF and FDA's SaMD and PCCP framing as the same shape: integrity controls required, mechanism unspecified.
Every sentence in that paragraph is one survey's reading of a body of text rather than an independent finding, and it is not legal advice. A negative over regulation is exactly the kind of claim to have counsel confirm before repeating it. Taken for what it is, it still sets the frame. The case for cryptographic model identity looks operational and risk-management-driven: it is for an organisation that already runs a model inventory and wants to know, mechanically rather than by attestation, which model produced a given output. Anyone pitching it as a compliance requirement should be asked which text they mean.
The objection this argument meets first, answered separately: if a trusted party has to exist anyway, a trusted execution environment runs the model at native speed, so why prove at all? The TEE objection.
ghostspeed / prove which model produced an output / mike@0pon.com / patent pending