<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[TokenSave Notes]]></title><description><![CDATA[TokenSave Notes]]></description><link>https://tokensave.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>TokenSave Notes</title><link>https://tokensave.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 07 Oct 2026 10:26:24 GMT</lastBuildDate><atom:link href="https://tokensave.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Is the most expensive AI the smartest? 23 LLMs, capability vs API price (Oct 2026)]]></title><description><![CDATA[Is the most expensive AI model the smartest one? I put 23 models' capability scores next to their API prices to check. Short answer: no. The top model costs about half of what the runner-up costs, and]]></description><link>https://tokensave.hashnode.dev/is-the-most-expensive-ai-the-smartest-23-llms-capability-vs-api-price-oct-2026</link><guid isPermaLink="true">https://tokensave.hashnode.dev/is-the-most-expensive-ai-the-smartest-23-llms-capability-vs-api-price-oct-2026</guid><category><![CDATA[AI]]></category><category><![CDATA[llm]]></category><category><![CDATA[openai]]></category><category><![CDATA[claude]]></category><dc:creator><![CDATA[Jaehyun Cho]]></dc:creator><pubDate>Sun, 04 Oct 2026 12:49:38 GMT</pubDate><content:encoded><![CDATA[<p>Is the most expensive AI model the smartest one? I put 23 models' capability scores next to their API prices to check. Short answer: no. The top model costs about half of what the runner-up costs, and a model at 5% of the top price lands within 13 points of it.</p>
<p>Capability scores come from the <strong>Epoch Capabilities Index (ECI)</strong> by Epoch AI, which combines dozens of benchmarks (math, coding, science, reasoning) into one scale. Prices are standard API list prices, checked October 4, 2026.</p>
<p><img src="https://tokensave.app/img/ai-model-capability-vs-price.png" alt="AI model capability vs API price, October 2026" /></p>
<p>Models on the green line are the <strong>best-value frontier</strong>: no other model is both cheaper and more capable. Anything below the line costs more than a frontier model with the same or higher score.</p>
<h2>The short version</h2>
<table>
<thead>
<tr>
<th>Pick</th>
<th>Model</th>
<th>ECI</th>
<th>10,000 chat requests</th>
</tr>
</thead>
<tbody><tr>
<td>Top capability</td>
<td>Claude Opus 5.5</td>
<td>167.3</td>
<td>$182</td>
</tr>
<tr>
<td>Best value</td>
<td>Claude Sonnet 5.5</td>
<td>165.2</td>
<td>$91</td>
</tr>
<tr>
<td>Cheap and strong</td>
<td>Gemini 3.8 Flash</td>
<td>156.9</td>
<td>$25</td>
</tr>
<tr>
<td>Lowest price</td>
<td>DeepSeek V4 Flash</td>
<td>154.5</td>
<td>$9</td>
</tr>
</tbody></table>
<p>A chat request is 1,500 input + 400 output tokens of English text. Costs include each model's tokenizer difference: Claude counts about 30% more tokens than GPT for the same text, and that is already in the numbers.</p>
<h2>Three things that surprised me</h2>
<p><strong>1. The #1 model is not the most expensive.</strong> Claude Opus 5.5 (167.3) edges out GPT-6 Astra (166.5), and the two are within each other's margin of error. But Astra lists at \(10 / \)50 per million tokens against Opus's \(4 / \)20, so 10,000 chats cost $350 on Astra and $182 on Opus.</p>
<p><strong>2. "Good enough" is very close to the top.</strong> Claude Sonnet 5.5 is 2.2 points under Opus for half the price. Gemini 3.8 Flash and DeepSeek V4 Flash sit 10–13 points below the top at 7–20× lower cost, which is plenty for classification, extraction, summaries and most chatbot traffic.</p>
<p><strong>3. Some popular picks are off the line.</strong> Claude Fable 5.1 (\(455 per 10k chats) and GPT-5.5 (\)195) both score below Sonnet 5.5 ($91). Gemini 3.1 Pro ($74) scores below Gemini 3.8 Flash ($25). They may still win on a specific task, but if you use them by default, test the cheaper option first.</p>
<h2>What about GPT-6 Sol?</h2>
<p>Epoch hasn't scored GPT-6 Sol yet. On a different scale, the <a href="https://artificialanalysis.ai/">Artificial Analysis Intelligence Index</a> (v4.3.2, max reasoning, Sept 30), <strong>GPT-6.1 Sol scores 51.8 vs GPT-6 Astra's 52.7</strong>, at a fifth of the price. If that holds on ECI, Sol would sit on the frontier next to Claude Sonnet 5.5, which has the same list price. It isn't on the chart because the scales differ.</p>
<h2>Live version</h2>
<p>The chart and table rebuild every day from Epoch's data and current list prices. You can switch between top capability, best value and cheapest views here:</p>
<p>👉 <strong><a href="https://tokensave.app/compare/performance">AI model capability vs price ranking</a></strong> (free, no sign-up)</p>
<p>Full write-up with the "expensive for what you get" table: <a href="https://tokensave.app/blog/best-value-llm-october-2026">Best value LLM, October 2026</a>.</p>
<p><em>Capability scores: Epoch Capabilities Index by Epoch AI, used under CC BY 4.0. Prices: API list prices, no caching or batch discounts.</em></p>
]]></content:encoded></item><item><title><![CDATA[Pretty JSON costs 3× CSV in tokens. Columnar JSON costs the same. (I measured 9 formats)]]></title><description><![CDATA[Most of us paste data into LLM prompts as JSON, often straight from JSON.stringify(data, null, 2). I had never checked what that costs in tokens, so I took one table, wrote it in seven formats and cou]]></description><link>https://tokensave.hashnode.dev/pretty-json-costs-3-csv-in-tokens-columnar-json-costs-the-same-i-measured-9-formats</link><guid isPermaLink="true">https://tokensave.hashnode.dev/pretty-json-costs-3-csv-in-tokens-columnar-json-costs-the-same-i-measured-9-formats</guid><category><![CDATA[llm]]></category><category><![CDATA[json]]></category><category><![CDATA[openai]]></category><category><![CDATA[Prompt Engineering]]></category><dc:creator><![CDATA[Jaehyun Cho]]></dc:creator><pubDate>Fri, 02 Oct 2026 16:56:28 GMT</pubDate><content:encoded><![CDATA[<p>Most of us paste data into LLM prompts as JSON, often straight from <code>JSON.stringify(data, null, 2)</code>. I had never checked what that costs in tokens, so I took one table, wrote it in seven formats and counted.</p>
<p>Short version: <strong>row-by-row JSON, pretty-printed, uses about 3× the tokens of CSV. The same data as columnar JSON costs about the same as CSV.</strong> It's the repeated keys, not JSON itself.</p>
<p><em>Update: a reader pointed out I'd left out columnar JSON (one object, one array per field). I measured it on the same table and tokenizer and added it below.</em></p>
<h2>The setup</h2>
<p>A product table: 20 rows, 5 fields (id, name, price, in_stock, category). One row in CSV:</p>
<pre><code>1001,Wireless Mouse,9.99,false,electronics
</code></pre>
<p>Same data, seven formats, counted with <code>o200k_base</code>, the tokenizer behind GPT-4o and later OpenAI models:</p>
<pre><code class="language-js">import { getEncoding } from "js-tiktoken";
const enc = getEncoding("o200k_base");
const count = (s) =&gt; enc.encode(s).length;

count(csv);                            // 300
count(JSON.stringify(rows));           // 525
count(JSON.stringify(rows, null, 2));  // 884
</code></pre>
<h2>Results</h2>
<table>
<thead>
<tr>
<th>Format</th>
<th>Tokens</th>
<th>vs CSV</th>
</tr>
</thead>
<tbody><tr>
<td>TSV</td>
<td>296</td>
<td>0.99×</td>
</tr>
<tr>
<td>CSV</td>
<td>300</td>
<td>1.00×</td>
</tr>
<tr>
<td>JSON, columnar, minified <code>{"id":[...],"name":[...]}</code></td>
<td>297</td>
<td><strong>0.99×</strong></td>
</tr>
<tr>
<td>JSON, header + row arrays <code>{"columns":[...],"rows":[[...]]}</code></td>
<td>320</td>
<td>1.07×</td>
</tr>
<tr>
<td>Markdown table</td>
<td>373</td>
<td>1.24×</td>
</tr>
<tr>
<td>JSON, columnar, pretty</td>
<td>520</td>
<td>1.73×</td>
</tr>
<tr>
<td>JSON, row objects, minified</td>
<td>525</td>
<td>1.75×</td>
</tr>
<tr>
<td>YAML</td>
<td>649</td>
<td>2.16×</td>
</tr>
<tr>
<td>JSON, row objects, pretty (2-space)</td>
<td>884</td>
<td><strong>2.95×</strong></td>
</tr>
<tr>
<td>XML</td>
<td>1,088</td>
<td><strong>3.63×</strong></td>
</tr>
</tbody></table>
<p>The older <code>cl100k_base</code> tokenizer gave nearly identical numbers, within 2% for every format.</p>
<h2>Why the gap is so big</h2>
<ol>
<li><strong>Keys repeat on every row.</strong> JSON, YAML and XML write <code>"name":</code>, <code>"price":</code> twenty times. CSV writes them once, in the header, and so does columnar JSON, which is why it lands right next to CSV.</li>
<li><strong>Punctuation is tokens.</strong> Quotes, braces and colons all take tokens. XML is worst: every value gets an opening and a closing tag.</li>
<li><strong>Indentation is for humans.</strong> The model doesn't need it. Pretty-printing alone added 359 tokens (+68%) over minified JSON here.</li>
</ol>
<p>YAML is the odd one: fewer characters than minified JSON, more tokens, because every field gets its own line and its own key.</p>
<h2>I measured some code too</h2>
<p>Coding agents send a lot of code, so I tried a few things:</p>
<table>
<thead>
<tr>
<th>Test</th>
<th>Result</th>
</tr>
</thead>
<tbody><tr>
<td>16-line Python file: 4 spaces vs 2 spaces vs tabs</td>
<td>135 / 135 / 133 tokens, <strong>basically no difference</strong></td>
</tr>
<tr>
<td>Same file without its one-line docstring</td>
<td>135 → 123 (−9%)</td>
</tr>
<tr>
<td>10-line JS function, minified</td>
<td>92 → 47 (−49%)</td>
</tr>
<tr>
<td>One UUID</td>
<td><strong>18 tokens</strong></td>
</tr>
</tbody></table>
<ul>
<li><strong>Indentation is nearly free.</strong> Modern tokenizers merge runs of spaces into single tokens. Don't reformat code to save tokens.</li>
<li><strong>Minified code is half the tokens, but don't.</strong> You lose names like <code>subtotal</code> and <code>taxRate</code>, which is exactly what helps the model understand the code.</li>
<li><strong>UUIDs are expensive.</strong> If your prompt has a column of long IDs, swap them for row numbers and map back in code.</li>
</ul>
<h2>What I use now</h2>
<ul>
<li><strong>Flat tables as input:</strong> CSV or TSV. TSV if values contain commas, so no quoting.</li>
<li><strong>If it has to be JSON:</strong> go columnar (<code>{"id":[...],"name":[...]}</code>) or header + row arrays. Same data, roughly CSV's cost.</li>
<li><strong>Data the model returns:</strong> minified JSON with a schema. Parse reliability beats a few saved tokens; use structured output if your API supports it.</li>
<li><strong>Nested data:</strong> minified JSON. Flattening it into CSV by hand often costs more than it saves.</li>
<li><strong>Tables a human will also read:</strong> Markdown, only 24% over CSV.</li>
<li><strong>Avoid pretty JSON and XML in prompts</strong> unless a tool requires them.</li>
</ul>
<p>Dropping unused fields, <code>null</code> fields and extra decimal places helps too.</p>
<h2>What it costs</h2>
<p>Say you attach a 20-row table to every request, 1,000 requests a day, 30,000 a month, at $2 per million input tokens:</p>
<table>
<thead>
<tr>
<th>Format</th>
<th>Tokens / month</th>
<th>Cost / month</th>
</tr>
</thead>
<tbody><tr>
<td>CSV</td>
<td>9.0M</td>
<td>$18</td>
</tr>
<tr>
<td>JSON, columnar</td>
<td>8.9M</td>
<td>$18</td>
</tr>
<tr>
<td>JSON, row objects, minified</td>
<td>15.8M</td>
<td>$32</td>
</tr>
<tr>
<td>JSON, row objects, pretty</td>
<td>26.5M</td>
<td>$53</td>
</tr>
<tr>
<td>XML</td>
<td>32.6M</td>
<td>$65</td>
</tr>
</tbody></table>
<p>Invisible per request, real for a product, and it scales with bigger tables, RAG results and long API responses.</p>
<h2>Caveats</h2>
<ul>
<li>One table, one tokenizer. Claude and Gemini tokenize differently, so absolute numbers will differ. The ranking (repeated keys and tags cost the most) comes from how the formats are built, so it should hold.</li>
<li>If the format changes answer quality for your task, test both. A cheaper prompt with worse answers isn't cheaper.</li>
</ul>
<h2>Try it on your own data</h2>
<p>I built a free <a href="https://tokensave.app/">token counter</a> for this. It runs the same <code>o200k</code> tokenizer in your browser, so nothing you paste is uploaded. Paste your data in two formats and compare tokens and cost per model.</p>
<p>Full write-up with more detail: <a href="https://tokensave.app/blog/json-vs-yaml-vs-csv-tokens">JSON vs YAML vs CSV: which format uses the fewest tokens?</a></p>
]]></content:encoded></item><item><title><![CDATA[ChatGPT Pro 100 vs 200 vs 500: Every Tier Costs the Same per Unit of Usage]]></title><description><![CDATA[On September 29, 2026 OpenAI split ChatGPT Pro into three plans: Pro 100, Pro 200 and a new Pro 500. It also quietly cut how much you get on Pro 200. If you are deciding between them, or wondering whe]]></description><link>https://tokensave.hashnode.dev/chatgpt-pro-100-vs-200-vs-500-every-tier-costs-the-same-per-unit-of-usage</link><guid isPermaLink="true">https://tokensave.hashnode.dev/chatgpt-pro-100-vs-200-vs-500-every-tier-costs-the-same-per-unit-of-usage</guid><category><![CDATA[chatgpt]]></category><category><![CDATA[openai]]></category><category><![CDATA[llm]]></category><category><![CDATA[AI]]></category><dc:creator><![CDATA[Jaehyun Cho]]></dc:creator><pubDate>Thu, 01 Oct 2026 11:40:55 GMT</pubDate><content:encoded><![CDATA[<p>On September 29, 2026 OpenAI split ChatGPT Pro into three plans: <strong>Pro 100</strong>, <strong>Pro 200</strong> and a new <strong>Pro 500</strong>. It also quietly cut how much you get on Pro 200. If you are deciding between them, or wondering whether you need Pro at all, here is what actually changed and how to pick.</p>
<img src="https://tokensave.app/blog-chatgpt-pro-tiers.png" alt="ChatGPT plans compared by usage and price" style="display:block;margin:0 auto" />

<h2>The plans side by side</h2>
<table>
<thead>
<tr>
<th>Plan</th>
<th>Price / month</th>
<th>Usage vs Plus</th>
<th>Price per "Plus-worth" of usage</th>
<th>Ultrafast</th>
</tr>
</thead>
<tbody><tr>
<td>Go</td>
<td>$8</td>
<td>lower</td>
<td>–</td>
<td>–</td>
</tr>
<tr>
<td>Plus</td>
<td>$20</td>
<td>1×</td>
<td>$20</td>
<td>–</td>
</tr>
<tr>
<td>Pro 100</td>
<td>$100</td>
<td>5×</td>
<td>$20</td>
<td>–</td>
</tr>
<tr>
<td>Pro 200</td>
<td>$200</td>
<td>10×</td>
<td>$20</td>
<td>–</td>
</tr>
<tr>
<td>Pro 500</td>
<td>$500</td>
<td>25×</td>
<td>$20</td>
<td>✓</td>
</tr>
</tbody></table>
<p>US prices. OpenAI's help page only says that Pro 200 includes more usage than Pro 100 and Pro 500 the most; the 5×, 10× and 25× figures are the multipliers reported from OpenAI's announcement.</p>
<p>All three Pro plans include the same features: Pro models, Codex, deep research, image creation, memory and file uploads. The only feature difference is <strong>Ultrafast</strong>, a faster mode for GPT-6 Astra, which is Pro 500 only. Buying extra credits on Pro 100 or Pro 200 does not unlock it.</p>
<h2>The part most people miss: there is no bulk discount any more</h2>
<p>Divide each price by the usage you get and every plan lands on the same number: <strong>$20 per Plus-worth of usage</strong>. Pro 500 is not a better deal than Pro 100; it is just more of the same thing, plus speed.</p>
<p>That used to be different. Until this change Pro 200 gave <strong>20×</strong> Plus usage, which worked out to $10 per unit, half the price of everything else. New Pro 200 subscribers now get <strong>10×</strong>. If you were already on Pro 200, you keep the old 20× allowance until <strong>October 29, 2026</strong>, then drop to 10× at the same $200.</p>
<p>So the rule is simple: <strong>buy the smallest plan you don't hit the limit on.</strong> Paying for headroom you never use is the only way to overpay.</p>
<h2>Which one should you pick?</h2>
<ul>
<li><p><strong>You rarely hit the Plus limit:</strong> stay on Plus ($20). None of the Pro plans give you a smarter answer for everyday chat; they give you more of it.</p>
</li>
<li><p><strong>You hit the Plus limit a few times a week:</strong> Pro 100. Five times the usage for five times the price, and the jump from $20 to $100 is the smallest step up.</p>
</li>
<li><p><strong>You regularly run out on Pro 100:</strong> Pro 200. Same price per unit, twice the room.</p>
</li>
<li><p><strong>You run Codex or agents most of the day, or waiting on output costs you money:</strong> Pro 500. It is the only plan with Ultrafast, but note that faster generation uses your allowance faster too.</p>
</li>
<li><p><strong>You're on the old Pro 200:</strong> keep it until October 29; it is the best deal OpenAI sells right now. After that, look at how much you actually used. If you stayed under about a quarter of your old allowance, Pro 100 does the same job for $100 less.</p>
</li>
</ul>
<h2>What about just using the API?</h2>
<p>If you mostly send short or medium messages, paying per token is often much cheaper than any Pro plan. Here is roughly what a month of chat costs through the API, assuming conversations of 6 messages and English text:</p>
<table>
<thead>
<tr>
<th>Your usage</th>
<th>GPT-6 Sol API</th>
<th>GPT-6 Astra API</th>
</tr>
</thead>
<tbody><tr>
<td>30 normal messages a day</td>
<td>≈ $8</td>
<td>≈ $39</td>
</tr>
<tr>
<td>80 normal messages a day</td>
<td>≈ $21</td>
<td>≈ $103</td>
</tr>
<tr>
<td>80 long messages a day (pasting documents)</td>
<td>≈ $59</td>
<td>≈ $294</td>
</tr>
<tr>
<td>200 long messages a day</td>
<td>≈ $147</td>
<td>≈ $735</td>
</tr>
</tbody></table>
<p>"Normal" = about 150 tokens in and 500 out per message; "long" = about 2,000 in and 700 out. API prices: Sol \(2 / \)10 and Astra \(10 / \)50 per million input / output tokens.</p>
<p>Two things push the numbers up fast. First, every new message re-sends the whole conversation, so long chats cost far more than short ones. Second, languages other than English need more tokens for the same text: Korean about 1.4×, Japanese about 1.8×, so the API bill grows by the same factor.</p>
<p>The takeaway: light and medium users of the everyday model are usually better off on Plus or the API. Pro starts to pay off when you use the top models heavily, work with long documents, or live in Codex.</p>
<h2>Check your own numbers</h2>
<p>Everyone's usage is different. The <a href="https://tokensave.app/plans">Subscription vs API calculator</a> lets you enter how many messages you send, how long they are and which language you write in, and shows what the same month would cost on the API next to ChatGPT, Claude and Gemini plans.</p>
<p><em>Prices as of October 1, 2026. OpenAI may change allowances again; check chatgpt.com/pricing before you buy.</em></p>
<p><em>Originally published on</em> <a href="https://tokensave.app/blog/chatgpt-pro-100-vs-200-vs-500"><em>tokensave.app</em></a><em>.</em></p>
]]></content:encoded></item><item><title><![CDATA[I measured GPT token costs in 10 Indian languages: Malayalam dropped from 10× to 1.85×]]></title><description><![CDATA[I translated one ordinary customer-support prompt into 10 Indian languages and counted tokens with o200k_base, the tokenizer behind GPT-4o and every newer OpenAI model. I also ran the same text throug]]></description><link>https://tokensave.hashnode.dev/i-measured-gpt-token-costs-in-10-indian-languages-malayalam-dropped-from-10-to-1-85</link><guid isPermaLink="true">https://tokensave.hashnode.dev/i-measured-gpt-token-costs-in-10-indian-languages-malayalam-dropped-from-10-to-1-85</guid><category><![CDATA[openai]]></category><category><![CDATA[llm]]></category><category><![CDATA[india]]></category><category><![CDATA[AI]]></category><dc:creator><![CDATA[Jaehyun Cho]]></dc:creator><pubDate>Wed, 30 Sep 2026 17:59:34 GMT</pubDate><content:encoded><![CDATA[<p>I translated one ordinary customer-support prompt into 10 Indian languages and counted tokens with <strong>o200k_base</strong>, the tokenizer behind GPT-4o and every newer OpenAI model. I also ran the same text through <strong>cl100k_base</strong>, the tokenizer from the GPT-4 era, to see how much has changed.</p>
<p>Short version: Indian languages got dramatically cheaper. <strong>Malayalam went from 10.09× the tokens of English to 1.85×.</strong> But the gap is not gone: <strong>Punjabi still needs 2.44× the tokens of English</strong>, more than any of the 34 languages I measured before.</p>
<h2>The prompt</h2>
<p>English (34 tokens):</p>
<blockquote>
<p>Please summarize the customer email below in three bullet points and suggest a polite reply. The customer says the order arrived two days late and one item was missing from the box.</p>
</blockquote>
<p>Hindi version:</p>
<blockquote>
<p>नीचे दिए गए ग्राहक के ईमेल को तीन बिंदुओं में सारांशित करें और एक विनम्र जवाब सुझाएँ। ग्राहक का कहना है कि ऑर्डर दो दिन देर से पहुँचा और डिब्बे में एक सामान कम था।</p>
</blockquote>
<pre><code class="language-js">import { getEncoding } from "js-tiktoken";

const now = getEncoding("o200k_base");   // GPT-4o and later
const old = getEncoding("cl100k_base");  // GPT-4 / GPT-3.5

now.encode(hindi).length; // 51
old.encode(hindi).length; // 156
</code></pre>
<h2>Results</h2>
<table>
<thead>
<tr>
<th>Language</th>
<th>Tokens today (o200k)</th>
<th>vs English today</th>
<th>vs English on old GPT-4 tokenizer</th>
</tr>
</thead>
<tbody><tr>
<td>English</td>
<td>34</td>
<td>1.00×</td>
<td>1.00×</td>
</tr>
<tr>
<td>Hindi</td>
<td>51</td>
<td><strong>1.50×</strong></td>
<td>4.59×</td>
</tr>
<tr>
<td>Gujarati</td>
<td>54</td>
<td><strong>1.59×</strong></td>
<td>7.18×</td>
</tr>
<tr>
<td>Urdu</td>
<td>54</td>
<td><strong>1.59×</strong></td>
<td>4.24×</td>
</tr>
<tr>
<td>Marathi</td>
<td>56</td>
<td><strong>1.65×</strong></td>
<td>4.91×</td>
</tr>
<tr>
<td>Bengali</td>
<td>57</td>
<td><strong>1.68×</strong></td>
<td>6.09×</td>
</tr>
<tr>
<td>Kannada</td>
<td>61</td>
<td><strong>1.79×</strong></td>
<td>9.35×</td>
</tr>
<tr>
<td>Malayalam</td>
<td>63</td>
<td><strong>1.85×</strong></td>
<td>10.09×</td>
</tr>
<tr>
<td>Tamil</td>
<td>67</td>
<td><strong>1.97×</strong></td>
<td>8.47×</td>
</tr>
<tr>
<td>Telugu</td>
<td>69</td>
<td><strong>2.03×</strong></td>
<td>9.38×</td>
</tr>
<tr>
<td>Punjabi</td>
<td>83</td>
<td><strong>2.44×</strong></td>
<td>7.41×</td>
</tr>
</tbody></table>
<p>For comparison, Simplified Chinese is 1.03×, Korean 1.44×, Japanese 1.79× and Greek 2.06×. The full 34-language table is here: <a href="https://tokensave.app/blog/token-cost-by-language">token cost by language</a>.</p>
<p><img src="https://tokensave.app/blog-language-tax-chart-v3.png" alt="Extra tokens per language vs English" /></p>
<h2>What stood out</h2>
<p><strong>1. The old tokenizer was brutal for South Indian languages.</strong> On cl100k, Malayalam needed 10×, Kannada and Telugu over 9×, Tamil 8.5×. Many characters fell back to raw bytes, so a single letter could cost several tokens.</p>
<p><strong>2. Hindi is now cheaper than Japanese.</strong> Hindi dropped from 4.59× to 1.50×. Japanese, which was cheaper than Hindi on the old tokenizer, is now 1.79×.</p>
<p><strong>3. Script matters more than the language family.</strong> Hindi and Marathi both use Devanagari and sit at 1.50–1.65×. Punjabi (Gurmukhi) is the outlier at 2.44×, even though it is closely related to Hindi.</p>
<h2>What it costs</h2>
<p>With a model at $2 per million input tokens, sending this prompt one million times costs:</p>
<table>
<thead>
<tr>
<th>Language</th>
<th>Input tokens</th>
<th>Input cost</th>
</tr>
</thead>
<tbody><tr>
<td>English</td>
<td>34M</td>
<td>$68</td>
</tr>
<tr>
<td>Hindi</td>
<td>51M</td>
<td>$102</td>
</tr>
<tr>
<td>Tamil</td>
<td>67M</td>
<td>$134</td>
</tr>
<tr>
<td>Punjabi</td>
<td>83M</td>
<td>$166</td>
</tr>
</tbody></table>
<p>That is input only. If the model also replies in the same language, the same multiplier applies to output tokens, which usually cost 4–5× more.</p>
<h2>How to spend fewer tokens</h2>
<ul>
<li><strong>Write the system prompt and fixed instructions in English.</strong> Keep only the user's message in their language. That part is repeated on every request, so it matters most.</li>
<li><strong>Ask for intermediate steps in English or JSON</strong> (classification, extraction, tool calls) and only the final answer in Hindi, Tamil or whichever language your user reads.</li>
<li><strong>Cache the fixed part of the prompt</strong> (prompt caching). Cached input is often 90% cheaper.</li>
</ul>
<h2>Limitations</h2>
<ul>
<li>One prompt only. With other text the ratios move by about ±0.1–0.2.</li>
<li>Translations started from machine translation and were checked, but not by native speakers of every language. If a translation looks off to you, tell me and I will re-measure.</li>
<li>Claude and Gemini use different tokenizers, so these numbers apply to OpenAI models.</li>
</ul>
<h2>Try it on your own text</h2>
<p>I built a free counter that runs this measurement in your browser (nothing is uploaded). It shows how many times more tokens your text uses than English, and one click can clean spaces, translate to English on-device (Chrome) and trim filler: <a href="https://tokensave.app/hi/">TokenSave in Hindi</a> · <a href="https://tokensave.app/">in English</a>.</p>
<p>Which Indian language do you use in your prompts, and do your numbers match mine?</p>
]]></content:encoded></item></channel></rss>