vitalik.eth
@VitalikButerin
Impressive work!
For comparison, an H100 can do roughly 100-200 tok/s of Muse 30B for raw inference single-thread, going up to low thousands of tok/s with a large number of threads - and I am sure that for the massively-multi-threaded case they can optimize the prover further.
So we roughly, sort of, have single-digit (<10x) overhead for LLM proving!
Next step is getting single-digit overheads for FHE, and then ultimately vFHE (aka STARK * FHE). A crazy ambitious milestone given present FHE overheads, but because of how highly structured and almost-linear LLM inference is, it's closer to the realm of possibility than you might think.
Single-digit-overhead all the things.
https://t.co/NXOTrPTjp6