<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Transformers on Signal &amp; Syntax</title><link>https://signal-and-syntax.com/tags/transformers/</link><description>Recent content in Transformers on Signal &amp; Syntax</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 22 Aug 2026 06:00:00 -0700</lastBuildDate><atom:link href="https://signal-and-syntax.com/tags/transformers/index.xml" rel="self" type="application/rss+xml"/><item><title>Inside Attention, Part 1: The Mechanism</title><link>https://signal-and-syntax.com/posts/inside-attention-part-1/</link><pubDate>Sat, 22 Aug 2026 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/inside-attention-part-1/</guid><description>The transformer architecture is composed of many repeating transformer layers, or blocks. Each block contains an attention sublayer followed by a feedforward sublayer, wrapped in residual connections and layer normalization. Positional information is added to the input so the model knows what order the tokens came in. The attention sublayer sets the table for the feedforward sublayer: it does the work of looking at other tokens and deciding what information to absorb from them.</description></item><item><title>Implementing a Minimal Transformer in PyTorch</title><link>https://signal-and-syntax.com/posts/minimal-transformer/</link><pubDate>Wed, 01 Apr 2026 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/minimal-transformer/</guid><description>In 2017, Vaswani et al. published &amp;ldquo;Attention Is All You Need,&amp;rdquo; a paper that quietly rearranged the entire landscape of machine learning. It introduced the Transformer architecture — a design that has since become the backbone of every major language model you&amp;rsquo;ve heard of: GPT, BERT, Claude, Gemini, and dozens of others. The paper&amp;rsquo;s title was a provocation. Attention mechanisms already existed. The claim was that you could throw out recurrence entirely and let attention carry the whole load.</description></item><item><title>How Large Language Models (LLMs) Tokenize Text: Why Words Aren't What You Think</title><link>https://signal-and-syntax.com/posts/how-large-language-models-tokenize-text/</link><pubDate>Tue, 11 Nov 2025 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/how-large-language-models-tokenize-text/</guid><description>When you type &amp;ldquo;I love programming&amp;rdquo; into ChatGPT, you might assume the model reads three words. It doesn&amp;rsquo;t. It reads somewhere between three and seven tokens, depending on how the text is split.
When you ask Claude to count the letters in the word &amp;ldquo;strawberry,&amp;rdquo; it often gets it wrong. The reason is simple. Claude never saw the word &amp;ldquo;strawberry&amp;rdquo; as a complete unit. It saw tokens like &amp;quot;str&amp;quot;, &amp;quot;aw&amp;quot;, &amp;quot;berry&amp;quot; and tried to reason about letters it couldn&amp;rsquo;t directly access.</description></item><item><title>How Large Language Models (LLMs) Handle Context Windows: The Memory That Isn't Memory</title><link>https://signal-and-syntax.com/posts/how-large-language-models-handle-context-windows/</link><pubDate>Mon, 10 Nov 2025 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/how-large-language-models-handle-context-windows/</guid><description>When you have a long conversation with a large language model (LLM) such as ChatGPT or Claude , it feels like the model remembers everything you&amp;rsquo;ve discussed. It references earlier points, maintains consistent context, and seems to &amp;ldquo;know&amp;rdquo; what you talked about pages ago.
But here&amp;rsquo;s the uncomfortable truth: the model doesn&amp;rsquo;t remember anything. It&amp;rsquo;s not storing your conversation in memory the way a database would. Instead, it&amp;rsquo;s rereading the entire conversation from the beginning every single time you send a message.</description></item></channel></rss>