<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Tokenization on Signal &amp; Syntax</title><link>https://signal-and-syntax.com/tags/tokenization/</link><description>Recent content in Tokenization on Signal &amp; Syntax</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 11 Nov 2025 06:00:00 -0700</lastBuildDate><atom:link href="https://signal-and-syntax.com/tags/tokenization/index.xml" rel="self" type="application/rss+xml"/><item><title>How Large Language Models (LLMs) Tokenize Text: Why Words Aren't What You Think</title><link>https://signal-and-syntax.com/posts/how-large-language-models-tokenize-text/</link><pubDate>Tue, 11 Nov 2025 06:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/how-large-language-models-tokenize-text/</guid><description>When you type &amp;ldquo;I love programming&amp;rdquo; into ChatGPT, you might assume the model reads three words. It doesn&amp;rsquo;t. It reads somewhere between three and seven tokens, depending on how the text is split.
When you ask Claude to count the letters in the word &amp;ldquo;strawberry,&amp;rdquo; it often gets it wrong. The reason is simple. Claude never saw the word &amp;ldquo;strawberry&amp;rdquo; as a complete unit. It saw tokens like &amp;quot;str&amp;quot;, &amp;quot;aw&amp;quot;, &amp;quot;berry&amp;quot; and tried to reason about letters it couldn&amp;rsquo;t directly access.</description></item><item><title>How Large Language Models (LLMs) Read Code: Seeing Patterns Instead of Logic</title><link>https://signal-and-syntax.com/posts/how-large-language-models-read-code/</link><pubDate>Mon, 06 Oct 2025 09:00:00 -0700</pubDate><guid>https://signal-and-syntax.com/posts/how-large-language-models-read-code/</guid><description>Developers are accustomed to thinking about code in terms of syntax and semantics, the how and the why. Syntax defines what is legal; semantics defines what it means. A compiler enforces syntax with ruthless precision and interprets semantics through symbol tables and execution logic. But a Large Language Model (LLM), reads code the way a seasoned engineer reads poetry, recognizing rhythm, pattern, and context more than explicit rules.
The difference may seem subtle, but it has vast consequences.</description></item></channel></rss>