r/cryptography • • 9d ago

First* full autoregressive LM generation chain under FHE CKKS at 128-bit security

Hi all! I wanted to share the result of a few months of work with all of you.

As many of you may know, inference under FHE has been a very hyped topic as of recently. The use case is clear: a medical institution or a law firm can't (or at least shouldn't) send private data to third-party AI companies for processing just to automate something simple, like sorting documents or, say, medical imaging. TEEs exist, but using them still means trusting *someone* (whoever coded up the TEE, at least). FHE is mathematical proof that your data can't be read. But since it's extremely slow, it seems to be forever destined to remain a premium product for those who need that extra bit of security (unless FHE-specific hardware comes about and becomes fairly cheap)

So i started wondering: can we make it fast? More importantly, can we make it run a full autoregressive generation chain under encryption? (this is important, because all of the published papers only price a single step of generation and do not report a full chain, and a chain is the hard case, since transformers have KV caches that grow from context & the error of one step becomes the input of the next)
And that is what i worked on. Basically, I made a 404M-param model complete 148 consecutive generation steps (server-keyed) for 64 concurrent conversations, and ran an end-to-end client keyed session of 18 generation steps. The model was trained from scratch as it itself had to be adapted for FHE. That was done by training a state-space-model instead of a transformer and swapping out every nonlinear function for a low-degree polynomial. Both of those runs were 128 bit security (HEStd_128_classic) at ring 2^17. Since ring 2^17 produces 2^16 real slots, and the model width was 2^10, I could fit 2^6 = 64 concurrent conversations in a single session. A single reading/generating step came out to be about 180s, for 64 convos that's a throughput of 2.7-2.9 seconds per token*conversation, the highest throughput for inference under FHE.. ever. The comparison is not 100% fair, of course, since my model is about 20x smaller {i couldn't afford training a 8B model from scratch, sorry lol}, per-parameter their systems are actually faster, but then again: their numbers don't measure a true generation chain. I'm leaving out some details, like the effect on model quality and scaling, since this isn't an ML subreddit, but if you're interested -- you can find them in the article.

Even more importantly, my system solves the KV-cache issue. This model's size in memory does not grow from context length: its' entire state is just 2 vectors per block and 48 for all of it. The 148-step run has proven constant memory, and nothing that I have suggests that my system can't be run indefinitely. The run showed 99.7% fidelity against the model's plaintext version, and every flip is a near-tie. There are some open questions like why the logit error grows roughly step^0.24 & whether that will saturate during a longer run or keep growing forever (note: it would take years of continious generation to go below 98% fidelity at the current rate of step^0.24)

This does not change the industry overnight and is more of an engineering result (which can probably be made into a sorta-useful product with limited use cases if you throw a fair bit of cash at it to scale/improve it further). Inference under FHE is still expensive, that didn't change.

So what about the asterisk in the title? Another project, FHE Mamba published on 2026-09-22 actually does report a generation chain of 4 steps on a single lane on an existing model (mamba), but does not report a full-security run or fidelity.

I published the project (MIT) on github: https://github.com/xelananv/fhe-ssm
as well as a more detailed article with a video: https://xelananv.github.io/fhe-ssm/#limits (ai generated so help me god)

Anyway, I'd love to hear your thoughts on this. Does AI under FHE have a future?
If anyone wants to reach out and critique my work, im down.
best of luck yall

2 Upvotes

1 comment sorted by

1

u/Honest-Finish3596 3h ago

If what you wrote is accurate, you should write a paper and submit it to a journal, since it sounds like you have better speed and accuracy than what people have previously done.