The TEE objection
The strongest argument against proving inference is short: if a trusted party has to exist anyway, put the model in a trusted execution environment and run it at native speed. The cryptography is then overhead you paid for nothing.
The answer is not that the objection is wrong about speed. It is right about speed.
The objection, at full strength
Trust in a signer is not an implementation detail that better cryptography removes later. The survey behind the write-up looked for any deployed system in which a verifier accepts a commitment to an ML model on some basis other than a party who saw the weights, and returned a clear negative: every one rests on a signature from the owner or operator, a hardware attestation, or reputation. So a system of this shape has a trusted party in it. That is the premise, and it is conceded.
Grant the premise and the rest follows quickly. If you already believe an attestation
from that party, let them run the model inside an enclave and put the runtime under
attestation instead of committing to the artifact. The enclave runs the network
natively, at whatever speed the hardware gives. Proving the same inference in a circuit
costs a median of 36.9 s at k = 17 as of the 2026-08-12 fold
restructure (70.41 s at k = 18 before it), and
native execution of an MLP this size is faster by a margin nobody disputes. Same trust
assumption, several orders of magnitude less work. Why prove?
One thing worth stating precisely, because the loose version of this argument makes the TEE sound like it needs no design work. A quote does not attest to an inference. It attests to platform and TCB state and to a measurement of the code loaded into the enclave. Per-inference integrity is a property of the protocol built on top: the enclave has to bind the request, the model identifier, the output and a nonce into something it signs, or return them over a channel the attestation established, and a verifier has to check that binding rather than the quote alone. That protocol is buildable and well understood. The point is only that both routes require it, so the comparison is between two designs and not between a design and a piece of hardware.
Anyone who cannot state the objection that strongly has not understood it. What follows is not a rebuttal of the speed claim.
Two questions, not one
The objection treats "trust the signer" and "trust the execution" as the same commitment because the same party might be named in both. They are different questions, asked at different times, and answered by different mechanisms.
What is this model? Answered once, at publication, by an attestation to what the commitment is. This is provenance, and it is social. Somebody says which artifact the identifier refers to, and you either believe them or you do not. The survey's negative says this is unavoidable in every deployed system it found, so no design here escapes it, including a TEE-based one.
Did this specific output come from that model? Answered for every inference, afterwards, by a proof. Answering it requires no trust in whoever ran the model, no trust in the machine it ran on, and no trust in any hardware vendor. It also requires no trust in the party who answered the first question, beyond what they already said: the proof binds an output to a commitment, and if you disbelieve the commitment's provenance the proof still tells you, exactly, which commitment produced the output.
A one-time provenance statement covers one artifact. A proof covers one output. Trust in the first does not spread to the second, and that is the point of separating them: an operator who is trusted to publish honestly is not thereby trusted to serve the published model on every request, at every hour, under every incentive to swap in something cheaper.
What a TEE buys, and what it costs
Stated fairly, because the comparison is only useful if the other side is described the way its users would describe it.
It buys per-inference integrity at native speed. That is a real property and it is the same property proofs give, obtained far more cheaply. For a service under latency pressure, that difference decides the architecture on its own. A TEE also offers a confidentiality property this system does not claim and this page does not compete with: the write-up is explicit that these proofs are not shown to conceal anything, and nothing here changes that.
The costs are not in the speed column. Trust moves to the silicon vendor and to the attestation service that vouches for it, which is a different party from the model owner and one the verifier usually cannot choose. The threat model grows a side-channel history, and enclave vulnerabilities are patched over time rather than settled once. The operator has to run the enclave, so the deployment topology is constrained by which hardware is available and to whom.
The cost that matters most here is temporal, and it is worth stating narrowly, because the loose version of this critique is wrong. A deployment that archived the quote, the certificate chain, the TCB collateral and a signed transcript of the inputs and outputs can have a verifier check that evidence years later without any endpoint being live. The dependency that does not go away is on the vendor's roots, on revocation and TCB policy, and on whether a verifier in the future is willing to accept the historical TCB state the quote records. Evidence produced under a configuration later marked out-of-date is still evidence, but what it is worth then is a policy judgement made by someone who was not in the room. A verifier who shows up late is not in the same position as one who was there.
What the proof buys
A proof of one inference in this system is 10,880 bytes, about 10.9 kB, and verifies
in a median of 169.0 ms at k = 17 on the machine reported in the cost
table. Those are the numbers that matter for this argument, and they are not the
proving numbers.
What that artifact is worth is a function of what it does not depend on, and the honest version of the claim names what it does depend on first. Checking a proof takes four things: the proof, the public instance, the verifying key for this circuit, and the public parameters the system was set up with. The last two are not derived from the proof, so they have to be archived or distributed alongside it, and an archive that keeps proofs and drops the key has kept nothing. They are also fixed per circuit rather than per inference, so one copy covers every proof that circuit will ever produce.
Given those, verification is a local computation. It is done by anyone, on ordinary hardware, without contacting the prover, the operator, or a vendor, and without any party's continued cooperation or continued existence. There is no service to be reachable, no revocation state to consult, no enclave generation to still be supported, and no policy call about whether some hardware configuration is still acceptable. The dependency is on a handful of bytes somebody kept, which is a dependency a verifier can satisfy for itself in advance.
Archival verifiability and non-repudiation are the product. Speed is not. The proving cost is the price of an artifact that survives the infrastructure that made it. If a deployment has no use for a check that outlives its own runtime, it is paying that price for nothing, and it should not pay it.
Non-repudiation is the same property read from the other side, and it is worth scoping
exactly. An operator who has produced a proof cannot later say the output came from a
model other than the one committed to in D, given the correct parameters,
verifying key and public instance, and cannot argue the checking party's tooling was
misconfigured, because a third party who was never involved can run the check. What the
proof does not do is settle what D is or what it means. That is the
provenance record's job, and it is the first question above, answered socially and once.
The proof closes the gap between a commitment and an output; it does not close the gap
between a commitment and a model anyone has vouched for.
They compose
The error in the objection is the conflation. One-time provenance trust and per-inference execution trust are not the same commitment, so trading one against the other is not the choice on offer. A TEE answers the second question with a hardware assumption. A proof answers it with no hardware assumption. Both still need the first question answered socially, and neither of them answers it.
Which means a deployment can use both without contradiction. Serve latency-critical traffic from an enclave, and prove the slice that has to be defensible later: the audited sample, the contested decision, the batch that a counterparty or a regulator will ask about after the fact. The two mechanisms cover different failure modes, and the second one is the one that still works when the first one's infrastructure has moved on.
When the TEE is the right answer
Often. If you are willing to trust a hardware vendor and its attestation service, if your verifiers check attestations while the deployment is live, and if nobody needs to re-check a specific inference years later against a party that may by then be hostile, a TEE gives you the integrity property at a fraction of the cost. That is not a concession made reluctantly. It is the correct recommendation for that buyer, and a page that could not say so would not be worth reading on the rest of it.
The proof is for the other case: the verifier who cannot extend trust to a hardware vendor, or will not, or who needs the check to outlive the machine, the enclave generation, and the commercial relationship. That is a narrower market than "everyone doing inference". It is stated narrowly on purpose.
The long version, including what this system does not do: what the proof shows, and what it does not.
ghostspeed / prove which model produced an output / mike@0pon.com