What's Next: The Signal & Syntax Roadmap

Want to know what's coming? Explore our roadmap to see what we're building, what's queued up, and what's already live.

If there is a topic in the planned list you want to see sooner, or a topic missing that you think belongs here, get in touch .


In Progress

  • Implementing a Minimal Transformer in PyTorch. Building the core machinery of a language model from embeddings and attention to training and generation

Planned (not ordered)

  • The Cooperative Witness Problem. A piece on the ways language models tend to accept and continue user premises rather than push back on them, and the training dynamics that produce that tendency. Follows naturally from Part 2 of Inside Attention, which sets up the function-class framing this piece depends on.
  • Pre-training vs. Post-training. The distinction between the compute-heavy foundation training phase and the alignment phase that shapes model behavior (supervised fine-tuning, RLHF, DPO, Constitutional AI). One of the highest-confusion topics for readers new to the field and a natural companion to the cooperative witness piece.
  • Softmax and Cross-Entropy Loss. A focused post on the connective tissue between probability and learning. Softmax as the operation that turns real numbers into distributions; cross-entropy as the loss that measures how well those distributions match the truth. Short, foundational, unlocks the probability layer.
  • The Feedforward Sublayer. A standalone post on the FFN sublayer inside each transformer block: what it does, why it is often much larger than the attention sublayer, and what interpretability research has shown about the concepts stored there.
  • In-Context Learning. How models adapt to patterns within a single prompt without parameter updates, and why induction heads (covered in Inside Attention Part 1) are part but not all of the mechanistic story.
  • Reasoning Models and Test-Time Compute. The class of models that spend additional compute at inference time to improve their answers, and why this changes what 'capability' means.
  • Scaling Laws and Emergence. The empirical relationships between compute, data, parameters, and capability, and where the sharp transitions in behavior come from.
  • A Mechanistic Interpretability Primer. An introduction to the research program of reverse-engineering trained transformers, at the level of concrete circuits and features rather than high-level intuitions.
  • Context Compression: What Does 'Lossless' Really Mean?. As context windows fill, AI systems increasingly rely on summarization and compression to preserve what matters. But if a system must decide what to discard before it knows what will matter later, how “lossless” can that compression really be? An examination of what context compression preserves, what it inevitably risks losing, and why the distinction matters for long-running AI systems.

Published (22 Posts)

AI and the Mathematics of Language (9 Posts)

The mathematics, mechanisms, and engineering behind large language models. These posts work from first principles to explore how models represent language, learn patterns, use attention and context, and generate output.

Applied Modeling and Simulation (8 Posts)

Real-world problems explored through mathematics, modeling, simulation, and code. These posts use specific problems as a way to develop broader techniques for reasoning about complex systems.

Essays and Perspectives (2 Posts)

Longer-form perspectives on AI, technology, research, and engineering. These posts step back from implementation details to examine broader questions, assumptions, and implications.

Python Techniques and Tooling (3 Posts)

Practical Python techniques, patterns, libraries, and tools. These posts focus on useful programming ideas that make Python code clearer, more effective, or easier to reason about.