KnowledgeSketches|← All papersLenny’s PodcastResearch papers
Knowledge Sketches · Brain2Qwerty v2 · Meta AI · June 2026

Reading Typed Sentences Straight From the Brain, No Surgery

Implanted electrodes can already restore communication to people who cannot speak or move, but they require brain surgery. This paper decodes full sentences from magnetoencephalography, a scanner that sits outside the head, and gets close enough to be interesting. From Meta AI and Ecole Normale Superieure with hospital and university collaborators, led by Mingfang Zhang, Jarod Levy and Jean-Remi King.

0%
average word error rate across nine people. For the best participant it drops to 22%, and half their sentences come out with at most one word wrong
9typists
10 hrsof MEG each
22,000sentences recorded
306MEG sensors
What decoding actually looks like
What the participant typed
What the model decoded from their brain signal
real examples from the paper, errors underlined in coral. It cycles

How close is close?

Word error rate, lower is better. The full pipeline beats its own encoder and the n-gram correction that state of the art used a year earlier.

invasive implants still win: under 2% word error for typing. The gap is real, but it used to be a chasm

Three models stacked, one per level of language

The pipeline is trained jointly to read characters, words and sentences from the same MEG signal.

Character

The Encoder

A convolutional network plus Conformer reads the continuous MEG stream and predicts keystrokes, with no need to know when each key was pressed. That timing freedom is what makes real-time use possible.

▶
Word

The Aligner

A contrastive model groups the brain signal into word-sized chunks, cut where the space key was predicted, and maps each chunk into the word-embedding space of a language model. 8 of 9 words land at rank 1.

▶
Sentence

The LLM

A LoRA-finetuned Qwen3 reads both the predicted characters and the brain-derived word tokens, then writes the sentence. Take the brain tokens away and every metric gets worse, so the LLM really is reading the neural signal.

best training recipe: one small LoRA adapter per person, then average their weights. The paper calls it model soup

Perfect sentences, by participant

Share of test sentences decoded with zero word edits, and within one edit. Each row is 100 dots.

the typical failure is one substituted or missing word, not a collapse into nonsense

The failure mode is fluent

Because the last stage is a language model, it always writes something readable. When the signal is weak, it produces a coherent sentence that is simply not yours.

Target, worst participant cars are not allowed on this road
Decoded had she not fallen down the stairs

grammatical, confident, wrong. The paper flags this openly: for passwords you would want the letter-perfect decoder, for conversation the fluent one

Two findings bigger than the headline number

Accuracy scales with data, and has not stopped log-linear

Character error falls in a straight line against the log of recording hours, with no plateau at the 90-hour ceiling of this study. The fit is nearly exact, a Pearson correlation of -0.99. More hours in the scanner means better decoding, which is precisely the bet that non-invasive approaches need to win.

Variety beats repetition 0.45 vs 0.65

Matched for total sentences, training on 256 unique sentences beat 128 sentences typed twice, by a wide margin. Diversity of language is its own axis of data quality.

It survives losing half its sensors robust

Decoding stays remarkably strong with 50% and even 25% of the 306 MEG sensors, which matters because future wearable sensors will be sparser than a lab scanner.

AI agents tuned the pipeline better than the standard tool auto research

Three coding agents, Cursor running Claude Opus 4.6, were each given the codebase and told to lower the error rate. All three beat Optuna, the classical optimizer: 16 to 20% relative improvement against 8.6%, and their gains held on all nine participants while Optuna's vanished. They independently converged on the same tricks: label smoothing, dropping the character stream during training, beam search, and stripping the prompt to its minimum.

But only inside a fence open-ended = failure

Given the open-ended task, matching v2 from the v1 codebase on their own, the same agents consistently failed. Large entangled code changes crashed most runs, and when a run did succeed the agents tended to idle rather than iterate. The authors' verdict: a force multiplier, not a replacement for researchers.

Read it with these caveats

  • Participants were nine healthy volunteers typing on a keyboard. Whether this transfers to patients who cannot type, the people the technology is for, is explicitly untested.
  • The model reads whole sentences after the fact, not word by word as you produce them. Real-time, low-latency decoding is future work.
  • The scanner is a 306-sensor cryogenic MEG, a room-sized machine. The clinical hope rests on wearable optically-pumped sensors that are still maturing.
  • Word error averages 39% against under 2% for invasive implants. The claim is a narrowing gap and a scaling path, not parity.
Based on the paper Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings, Zhang, Levy, et al., Meta AI and ENS, June 2026. Every number on this page is from the paper.Read the paper
This sketch began as a paper.
We draw podcast episodes too. Send us a transcript. $1 a sketch, or $49 for all 11 formats.
Sketch mine →