<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Attention Mechanisms on Signal &amp; Syntax</title><link>https://signal-and-syntax.com/tags/attention-mechanisms/</link><description>Recent content in Attention Mechanisms on Signal &amp; Syntax</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 22 Aug 2026 06:00:00 -0700</lastBuildDate><atom:link href="https://signal-and-syntax.com/tags/attention-mechanisms/index.xml" rel="self" type="application/rss+xml"/><item><title>Inside Attention, Part 1: The Mechanism</title><link>https://signal-and-syntax.com/posts/inside-attention-part-1/</link><pubDate>Sat, 22 Aug 2026 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/inside-attention-part-1/</guid><description>The transformer architecture is composed of many repeating transformer layers, or blocks. Each block contains an attention sublayer followed by a feedforward sublayer, wrapped in residual connections and layer normalization. Positional information is added to the input so the model knows what order the tokens came in. The attention sublayer sets the table for the feedforward sublayer: it does the work of looking at other tokens and deciding what information to absorb from them.</description></item><item><title>Implementing a Minimal Transformer in PyTorch</title><link>https://signal-and-syntax.com/posts/minimal-transformer/</link><pubDate>Wed, 01 Apr 2026 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/minimal-transformer/</guid><description>In 2017, Vaswani et al. published &amp;ldquo;Attention Is All You Need,&amp;rdquo; a paper that quietly rearranged the entire landscape of machine learning. It introduced the Transformer architecture — a design that has since become the backbone of every major language model you&amp;rsquo;ve heard of: GPT, BERT, Claude, Gemini, and dozens of others. The paper&amp;rsquo;s title was a provocation. Attention mechanisms already existed. The claim was that you could throw out recurrence entirely and let attention carry the whole load.</description></item><item><title>The Discrete Mathematics Hiding Inside LLMs</title><link>https://signal-and-syntax.com/posts/discrete-math-in-large-language-models/</link><pubDate>Tue, 31 Mar 2026 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/discrete-math-in-large-language-models/</guid><description>A recent LinkedIn post from Michael Palmer described how discrete mathematics is the foundation for how computers reason about problems. That thread got me thinking about just how many discrete math concepts show up inside systems that seem purely statistical. LLMs are often described in terms of neural networks, gradient descent, and probability distributions. If you&amp;rsquo;ve taken discrete mathematics and wondered what it has to do with modern AI, the answer is: more than you&amp;rsquo;d expect.</description></item><item><title>How Large Language Models (LLMs) Handle Context Windows: The Memory That Isn't Memory</title><link>https://signal-and-syntax.com/posts/how-large-language-models-handle-context-windows/</link><pubDate>Mon, 10 Nov 2025 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/how-large-language-models-handle-context-windows/</guid><description>When you have a long conversation with a large language model (LLM) such as ChatGPT or Claude , it feels like the model remembers everything you&amp;rsquo;ve discussed. It references earlier points, maintains consistent context, and seems to &amp;ldquo;know&amp;rdquo; what you talked about pages ago.
But here&amp;rsquo;s the uncomfortable truth: the model doesn&amp;rsquo;t remember anything. It&amp;rsquo;s not storing your conversation in memory the way a database would. Instead, it&amp;rsquo;s rereading the entire conversation from the beginning every single time you send a message.</description></item></channel></rss>