Every figure here is from Google’s own blog post and paper about its own system. Nothing was independently verified and this publication ran nothing. The models trained so far reach up to 10M parameters.
Google says a new federated learning system that trains on the server, inside trusted execution environments (TEEs), trained a Gboard English model in 3 weeks against 2 months, with a 3x smaller privacy budget. Those are Google’s own claims, from its own blog post and paper. Nothing was independently verified, and this publication ran nothing.
What Google Built
Google Research announced the system on 2 October 2026 in a blog post by Katharine Daly and Daniel Ramage, with a paper on arXiv. In earlier federated learning, phones computed gradients locally. In the new system, devices upload encrypted training examples, and training runs on the server inside TEEs.
The uploads are tied to an access policy listing the server programs allowed to process them. Policies are published to Rekor, a public transparency log. Google says only metrics and differentially private model weights are visible to its own operators.
Google says the guarantees hold subject to current-generation TEE limitations. The paper also notes known limitations such as side-channel observability.

Three Experiments, Kept Separate
The paper reports three different experiments. They use different models, different populations and different measures, so this piece does not blend or compare figures across them.
Device Coverage: the Japanese Model
In the baseline system, 3000 rounds of training over 38 days with cohorts of 6500 used contributions from only 8.5M of 35.5M available devices, or 23.9%. The paper says the number of available devices observed over time was similar in both systems.
In the TEE system, server-side training started after about six days of collecting uploads, 17.8M in all, and every upload went into training. This publication’s own arithmetic: that is about 2.1 times as many devices as the baseline’s 8.5M, collected over a much shorter period.
Privacy-Utility Headroom: the English Model
In the baseline, the noise multiplier was fixed at 7.38 and the model trained for over 8616 rounds across 85 days. In the TEE system, Google targeted a zCDP of 0.232, and the noise multiplier needed at 5000 rounds came out at 5.16, based on 11.8M collected uploads.
The paper says the baseline would have needed a noise multiplier of 9.54 for the same zCDP at the same number of rounds. It traces the headroom to participation patterns: at 5000 rounds the baseline had maxP of 8 and minSep of 561, against 3 and 1,822 for the TEE system.
The Live A/B Test
Google tested checkpoints of an English language model in a live A/B experiment on part of Gboard’s production population, with 3.5M devices in each arm. The best TEE-trained arm used a zCDP of 0.215, against 0.641 for the production model once adjusted to the same MF-DP-FTRL mechanism.
The paper calls that a 3x smaller privacy budget; this publication’s own arithmetic gives about 2.98. It reports neutral performance on Words Per Minute and Words Modified Ratio, with no numbers given, and says the TEE model trained in 3 weeks against 2 months for the previous production model. The blog says training these models could previously take 1-2 months each.
Scale and Limits
The paper says models trained so far reach up to 10M parameters. Training much larger ones may need worker TEEs to use GPUs and may expose communication bottlenecks between the root and workers. The paper says round times improved significantly on as few as 14 machines.
The blog says Gboard has launched English and Japanese next-word prediction models on the new system, and calls them improved in accuracy. The paper gives no accuracy figures, only the neutral usability result.
What Is Not Established
Not established, and therefore absent from this analysis: any independent verification of the figures; accuracy numbers for the TEE-trained models; how many Gboard users run them; results above 10M parameters; and how a side-channel attack would affect the guarantees.
None of the above is investment advice. It reports what Google says about its own system.
Every figure above comes from Google Research’s blog post of 2 October 2026 and the paper “Toward provably private learning from federated data” (arXiv 2609.31494), both written by Google and read in full. They are Google’s own account of its own system. This publication ran nothing and found no independent verification in the sources it read. The Japanese device-coverage figures, the English privacy-utility figures and the live A/B test are three separate experiments and are not combined here.
The A/B test reports neutral usability metrics, not better ones, and gives no numbers for them. Models trained so far reach up to 10M parameters. Google says its guarantees are subject to current-generation TEE limitations and to side-channel observability. The 2.98 and 2.1 times figures are this publication’s own arithmetic. Nothing above predicts anything and nothing here is investment advice.
Sources: research.google · arxiv.org









