Problem Statement: A Private Keyboard That Learns Without Forgetting
Frames the product as an edge-first learning system, not a cloud autocomplete API.
Problem statement
Design a mobile keyboard that learns a user's typing patterns, new slang, names, and domain vocabulary locally, and keeps improving its next-word suggestions over months of use. The model is a small LSTM or compact transformer that updates incrementally from each user's typing sessions. The design must handle partial offline usage, keep raw text on the device, and optionally contribute anonymized model diffs to a central aggregator so the global model improves without any server ever reading what anyone typed.
This is not a 'call a prediction API' problem. Suggestions must appear within tens of milliseconds of a keystroke, while the device may be offline, thermal-throttled, memory-constrained, or running a game in the foreground. At the same time, the learning loop must not corrupt a good base model with noisy one-user data, must not catastrophically forget the general language, and must never leak raw keystrokes through gradients, vocabularies, or telemetry.
Why this problem is distinctive
A web recommender can retrain nightly in a data center on centralized logs. A keyboard cannot, for three hard reasons. First, privacy: typed text is among the most sensitive data a person produces, containing passwords, medical terms, and private conversations, so raw text must not leave the device. Second, latency: suggestion candidates are needed on every keystroke, so inference must be local and fast; a 300 ms cloud round trip is unusable. Third, heterogeneity: hundreds of millions of devices have different languages, slang, contact lists, and hardware, so one global model is never enough; personalization must happen per device.
The result is a system with two learning loops running at different speeds. A fast, local loop fine-tunes a small adapter layer on the device after the user charges and locks the phone. A slow, global loop aggregates differentially private model diffs from a daily cohort of participating devices, evaluates the result, and ships a new base model through staged rollout. The architecture must keep these loops decoupled so a bad global round cannot wipe out local personalization, and so a compromised or curious server never reconstructs a sentence.
Public operating baseline
The category is real and large. Google Play lists Gboard at more than 1 billion installs. Google's 2018 paper 'Federated Learning for Mobile Keyboard Prediction' describes training LSTM language models across fleets of phones with federated averaging and secure aggregation, updating models without centralizing keystrokes. Apple's privacy materials describe QuickType learning on the device with local differential privacy, publishing per-day epsilon budgets rather than raw data. Microsoft's SwiftKey popularized LSTM neural autocomplete and moved prediction on-device. These are public signals that anchor feasibility; every capacity number in this answer that is not cited is an explicit design assumption.
The four architectural planes
- Suggestion-serving plane: on-device inference, candidate ranking, IME integration, sub-50 ms latency.
- Device-learning plane: local sample collection, eligibility checks, adapter fine-tuning, local evaluation, rollback.
- Federated aggregation plane: cohort selection, contribution upload, secure aggregation, DP noise, global model production.
- Governance plane: model registry, release manifests, canaries, quality metrics, privacy accounting, audit.
A strong answer keeps these planes separate. The serving plane must survive total cloud outage. The device-learning plane must survive bad local data. The aggregation plane must survive device churn and adversarial contributions. The governance plane must be able to halt any rollout without bricking keyboards that are offline.
Key Highlights
- •Suggestions are latency-critical and privacy-critical, so inference and personalization run on-device; the cloud only aggregates anonymized diffs.
- •Two learning loops at two speeds: fast local adapter fine-tuning, slow global federated rounds.
- •Gboard passes 1B+ Play installs and Google's 2018 federated keyboard paper proves the pattern at fleet scale.
- •Four planes: suggestion serving, device learning, federated aggregation, governance and release.
- •The serving plane must work during total cloud outage; offline is the normal state, not an exception.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "Let me separate the fast local learning loop from the slow global learning loop before choosing any model."
- "I will define what crosses the device boundary before defining what runs in the cloud."