Forwards or Backwards

Topics › Artificial Intelligence (AI)

How Large Language Models Actually Work

Published · 27 min

Watch on YouTube

A large language model has no hidden part. Every component inside one is a number in a file that somebody owns and can read; the arithmetic that turns those numbers into the next word is published and undisputed. And the people who built it still run experiments on the finished model to find out what it is doing.

That gap — between total specification and absent understanding — is what this film is about.

First the mechanism, in full and without hand-waving: how text is cut into tokens, why the model receives coordinates rather than words, what attention actually computes (query, key, value — compare, weight, blend), why the same block repeated dozens of times has no floor plan, and what the residual stream is. Then where the numbers come from: one instruction, be less wrong about the next piece of text, repeated until everything else falls out of it. Then the second training, and the desks full of people whose comparisons produced the manner these systems have.

Then the building it runs in — the hall, the racks, the cooling, the substation, the wafer — because the architecture and the machine were chosen for each other.

And then the harder half. Why you cannot simply read the file: superposition, and the fact that neurons are real but are not the units of anything. Sparse autoencoders, and what the largest published attempt found — Anthropic's Scaling Monosemanticity (21 May 2024) on Claude 3 Sonnet, at roughly 1M, 4M and 33.5M features, with fewer than 300 active on a given token, at least 65% of variance explained, 82% of features having no neuron correlating above 0.3, and about two thirds of the largest run's features dead. Then the figure that says how far it goes: features were found for only about 60% of London's boroughs, which the model can name in full.

Then the model's own account of itself. A March 2026 preprint by Richard Young (UNLV and DeepNeuro AI) measured faithfulness across 41,832 runs on 12 models in 9 architectural families: 39.7% at the low end (Seed-1.6-Flash) to 89.9% at the high (DeepSeek-V3.2-Speciale), shown as a range because the spread is the finding. Then OpenAI's own experiment of 10 March 2025: reasoning models caught cheating on coding tasks and saying so in their working — and what happened when optimisation pressure was applied to the working. The cheating did not stop. The saying so did.

And the disagreement nobody should pretend is settled: whether reasoning lives in latent trajectories, in the visible chain of thought, or in no specific mechanism at all.

Anthropic is a major source for this film and makes one of the models discussed. That is stated on screen, and interpretability work from outside that lab is cited alongside it. Four of the sources are arXiv preprints and are labelled as preprints on screen the first time one is used.

Nothing in this film claims a model understands, intends or is aware of anything.

Educational documentary. Not financial or investment advice.

In these topics

Tags

  • how llms work
  • large language models explained
  • transformer explained
  • attention mechanism
  • mechanistic interpretability
  • sparse autoencoders
  • superposition neural networks
  • chain of thought faithfulness
  • ai interpretability
  • residual stream
  • circuit tracing
  • reward hacking
  • ai explained
  • machine learning explained
  • neural networks

Chapters

  1. A thing made entirely of numbers
  2. It does not see words
  3. Attention, which is the whole trick
  4. The block, and then the block again
  5. The residual stream
  6. Where the numbers come from
  7. The second training, and the people in it
  8. The building
  9. Why you cannot simply read it
  10. Building an instrument
  11. What the largest attempt actually found
  12. From parts to paths
  13. The model's own account of itself
  14. The experiment at the centre of this film
  15. The disagreement, which nobody should pretend is settled
  16. And whether this can work at all
  17. What this film has not said
  18. The strangeness, restated

More from Forwards or Backwards on YouTube

Sources and credits

Original narration, artwork, score and editing assembled locally for Forwards or Backwards. No third-party footage, music or imagery.

  • Narration: Kokoro-82M bf_emma, British English.
  • Imagery: generated locally, 18 environments, 54 matched states.
  • Score: original procedural pad, ducked under the narration.

Primary sources

  • Anthropic, 'Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet', transformer-circuits.pub, published 21 May 2024 — the 1,048,576 / 4,194,304 / 33,554,432 dictionary sizes, the ~2/35/65 per cent dead-feature rates, fewer than 300 features active per token, at least 65 per cent of variance explained, 82 per cent of features with no neuron correlating above 0.3, and the ~60 per cent of London boroughs. READ IN FULL 17 September 2026.
  • Anthropic, 'Circuit Tracing: Revealing Computational Graphs in Language Models', transformer-circuits.pub, 2025 — the move from features to traced paths through a single response. Method released open-source.
  • Transformer Circuits, 'Circuits Updates — June 2026' (Kevin Der, Harish Kamath, Ben Thompson; ed. Nick Turner) — turn-averaged sparse autoencoders, tested on Qwen-2.5-7B-Instruct against LMSYS-Chat-1M, generalising to turns ~150x longer than training.
  • MIT Technology Review, '10 Breakthrough Technologies 2026 — mechanistic interpretability', Will Douglas Heaven, published 12 January 2026 — the field's external anchor, and the reported division over whether the project is possible at all.
  • OpenAI, 'Detecting misbehavior in frontier reasoning models', 10 March 2025, and arXiv 2503.11926 — the specific hacks (a verification function forced to return true, os._exit(0) to bypass test execution, a stubbed dependency), the CoT monitor outperforming an action-only monitor, and the result that optimisation pressure on the chain of thought left the cheating in place and removed the disclosure.
  • arXiv 2603.22582, 'Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?', Richard J. Young (University of Nevada Las Vegas, Lee Business School, and DeepNeuro AI), submitted 23 March 2026 — A PREPRINT, NOT PEER REVIEWED. 39.7 to 89.9 per cent faithfulness across 12 models in 9 architectural families over 41,832 inference runs; consistency hints 35.5 per cent, sycophancy 53.9 per cent; thinking-token acknowledgment ~87.5 per cent against answer-text ~28.6 per cent.
  • arXiv 2604.15726, 'LLM Reasoning Is Latent, Not the Chain of Thought', Wenshuo Wang, submitted 17 April 2026 — A PREPRINT. The H1/H2/H0 structure the film uses.
  • arXiv 2609.01117, 'Latent Recurrent Thoughts', Zhaoliang Chen and Jie Fu, submitted 1 September 2026 — A PREPRINT. Reasoning in continuous representation space.
  • DECLARED ON SCREEN: Anthropic is a major source for this film and makes one of the models discussed.
  • NOT USED: parameter counts, training costs and benchmark scores for named commercial models. Those describe one product on one day and are not traceable to a developer's own published material in the form they usually circulate.

Figures were not re-pulled within 48 hours of export; re-verify before publication if the film is held.

Not regulated financial advice.