Implanted electrodes can already restore communication to people who cannot speak or move, but they require brain surgery. This paper decodes full sentences from magnetoencephalography, a scanner that sits outside the head, and gets close enough to be interesting. From Meta AI and Ecole Normale Superieure with hospital and university collaborators, led by Mingfang Zhang, Jarod Levy and Jean-Remi King.
Word error rate, lower is better. The full pipeline beats its own encoder and the n-gram correction that state of the art used a year earlier.
invasive implants still win: under 2% word error for typing. The gap is real, but it used to be a chasm
The pipeline is trained jointly to read characters, words and sentences from the same MEG signal.
A convolutional network plus Conformer reads the continuous MEG stream and predicts keystrokes, with no need to know when each key was pressed. That timing freedom is what makes real-time use possible.
▶A contrastive model groups the brain signal into word-sized chunks, cut where the space key was predicted, and maps each chunk into the word-embedding space of a language model. 8 of 9 words land at rank 1.
▶A LoRA-finetuned Qwen3 reads both the predicted characters and the brain-derived word tokens, then writes the sentence. Take the brain tokens away and every metric gets worse, so the LLM really is reading the neural signal.
best training recipe: one small LoRA adapter per person, then average their weights. The paper calls it model soup
Share of test sentences decoded with zero word edits, and within one edit. Each row is 100 dots.
the typical failure is one substituted or missing word, not a collapse into nonsense
Because the last stage is a language model, it always writes something readable. When the signal is weak, it produces a coherent sentence that is simply not yours.
grammatical, confident, wrong. The paper flags this openly: for passwords you would want the letter-perfect decoder, for conversation the fluent one
Character error falls in a straight line against the log of recording hours, with no plateau at the 90-hour ceiling of this study. The fit is nearly exact, a Pearson correlation of -0.99. More hours in the scanner means better decoding, which is precisely the bet that non-invasive approaches need to win.
Matched for total sentences, training on 256 unique sentences beat 128 sentences typed twice, by a wide margin. Diversity of language is its own axis of data quality.
Decoding stays remarkably strong with 50% and even 25% of the 306 MEG sensors, which matters because future wearable sensors will be sparser than a lab scanner.
Three coding agents, Cursor running Claude Opus 4.6, were each given the codebase and told to lower the error rate. All three beat Optuna, the classical optimizer: 16 to 20% relative improvement against 8.6%, and their gains held on all nine participants while Optuna's vanished. They independently converged on the same tricks: label smoothing, dropping the character stream during training, beam search, and stripping the prompt to its minimum.
Given the open-ended task, matching v2 from the v1 codebase on their own, the same agents consistently failed. Large entangled code changes crashed most runs, and when a run did succeed the agents tended to idle rather than iterate. The authors' verdict: a force multiplier, not a replacement for researchers.