What's Next: The Signal & Syntax Roadmap
If there is a topic in the planned list you want to see sooner, or a topic missing that you think belongs here, get in touch .
In Progress
- Implementing a Minimal Transformer in PyTorch. Building the core machinery of a language model from embeddings and attention to training and generation
Planned (not ordered)
- The Cooperative Witness Problem. A piece on the ways language models tend to accept and continue user premises rather than push back on them, and the training dynamics that produce that tendency. Follows naturally from Part 2 of Inside Attention, which sets up the function-class framing this piece depends on.
- Pre-training vs. Post-training. The distinction between the compute-heavy foundation training phase and the alignment phase that shapes model behavior (supervised fine-tuning, RLHF, DPO, Constitutional AI). One of the highest-confusion topics for readers new to the field and a natural companion to the cooperative witness piece.
- Softmax and Cross-Entropy Loss. A focused post on the connective tissue between probability and learning. Softmax as the operation that turns real numbers into distributions; cross-entropy as the loss that measures how well those distributions match the truth. Short, foundational, unlocks the probability layer.
- The Feedforward Sublayer. A standalone post on the FFN sublayer inside each transformer block: what it does, why it is often much larger than the attention sublayer, and what interpretability research has shown about the concepts stored there.
- In-Context Learning. How models adapt to patterns within a single prompt without parameter updates, and why induction heads (covered in Inside Attention Part 1) are part but not all of the mechanistic story.
- Reasoning Models and Test-Time Compute. The class of models that spend additional compute at inference time to improve their answers, and why this changes what 'capability' means.
- Scaling Laws and Emergence. The empirical relationships between compute, data, parameters, and capability, and where the sharp transitions in behavior come from.
- A Mechanistic Interpretability Primer. An introduction to the research program of reverse-engineering trained transformers, at the level of concrete circuits and features rather than high-level intuitions.
- Context Compression: What Does 'Lossless' Really Mean?. As context windows fill, AI systems increasingly rely on summarization and compression to preserve what matters. But if a system must decide what to discard before it knows what will matter later, how “lossless” can that compression really be? An examination of what context compression preserves, what it inevitably risks losing, and why the distinction matters for long-running AI systems.
Published (22 Posts)
AI and the Mathematics of Language (9 Posts)
The mathematics, mechanisms, and engineering behind large language models. These posts work from first principles to explore how models represent language, learn patterns, use attention and context, and generate output.
- Inside Attention, Part 1: The Mechanism. Attention is the engine. The rest of the transformer architecture stabilizes it, organizes it, and makes deep training possible.
- The Discrete Mathematics Hiding Inside LLMs. How set theory, predicate logic, and formal proofs show up in modern AI
- How Large Language Models (LLMs) Know Things They Were Never Taught. Web search, RAG, and the illusion of current knowledge
- Temperature and Top-P: The Creativity Knobs. How sampling parameters shape AI personality
- How Large Language Models (LLMs) Tokenize Text: Why Words Aren't What You Think. Understanding how LLMs break language into pieces—and why it matters more than you realize
- How Large Language Models (LLMs) Handle Context Windows: The Memory That Isn't Memory. Exploring why longer context doesn't mean better memory and what happens when conversations grow
- How Large Language Models (LLMs) Learn: Calculus and the Search for Understanding. Exploring how gradient descent and partial derivatives teach models to think
- How Large Language Models (LLMs) Think: Turning Meaning into Math. Exploring how large language models use linear algebra to create geometric meaning
- How Large Language Models (LLMs) Read Code: Seeing Patterns Instead of Logic. Exploring how large language models interpret code and what they miss
Applied Modeling and Simulation (8 Posts)
Real-world problems explored through mathematics, modeling, simulation, and code. These posts use specific problems as a way to develop broader techniques for reasoning about complex systems.
- The Wreck of the Edmund Fitzgerald: Modeling Decomposition in Extreme Environments. How cold, pressure, and buoyancy explain why Lake Superior may never give up her dead
- The Birthday Paradox in Production: When Random IDs Collide. Why collision risk grows faster than intuition suggests, and what that means for IDs, hashes, and distributed systems
- Rethinking the Three-Second Traffic Rule: When Physics Says It’s Not Enough. Using kinematics and Python to test when a familiar following-distance rule stops being safe
- Modeling Heat Capacity and Evaporation with Python: Why Water Warms Slowly but Cools Fast. Using thermodynamics and Python to explain why pools resist daytime warming but lose heat rapidly at night
- The Five-Second Rule Explored with Math & Python. Why germs transfer immediately, and how math reveals what really happens after food hits the floor
- The Meeting Diet: An Optimization Approach to Your Calendar. How optimization can help you decide which meetings deserve your limited time and energy
- From Ice Shows to Algorithms: Cracking the Truck-Packing Problem. Using Python, heuristics, and 3D visualization to tackle a real-world optimization problem
- Should You Walk or Run in the Rain? The Puzzle That Sparked a Passion. Using physics, Python, and simulation to settle a deceptively simple question
Essays and Perspectives (2 Posts)
Longer-form perspectives on AI, technology, research, and engineering. These posts step back from implementation details to examine broader questions, assumptions, and implications.
- Hash Collisions: Why Your 'Unique' Fingerprints Aren't (And Why That's Usually OK). The mathematical certainty of hash collisions, the near-impossibility of meaningful ones, and what it means for modern cryptography
- From Solow to ChatGPT: Why Total Factor Productivity Can't Keep Up With Generative AI. Why a productivity measure built for the industrial economy struggles to capture the value created by generative AI
Python Techniques and Tooling (3 Posts)
Practical Python techniques, patterns, libraries, and tools. These posts focus on useful programming ideas that make Python code clearer, more effective, or easier to reason about.
- Numeric Parsing in Python with Integer Division and Modulus. Using // and % to split fixed-width numeric data without converting it to strings
- Using SymPy in Python When NumPy Isn't Enough. Choosing exact symbolic mathematics when floating-point approximations are not good enough
- Using Python Dispatch Tables for Cleaner Validation. Replacing sprawling validation logic with a compact, declarative mapping of rules and results