{
    "version": "https://jsonfeed.org/version/1",
    "title": "Dr. TMA Pai Endowment Chair - ITIS",
    "home_page_url": "https://blog.ecitis.org/",
    "feed_url": "https://blog.ecitis.org/feed.json",
    "description": "Official website for Dr. TMA Pai Endowment Chair - ITIS",
    "icon": "https://blog.ecitis.org/og/og-logo.png",
    "author": {
        "name": "Dr. TMA Pai Endowment Chair - ITIS",
        "url": "https://blog.ecitis.org/"
    },
    "items": [
        {
            "id": "https://blog.ecitis.org/j-space-global-workspace/",
            "content_html": "<p>Anthropic’s July 2026 J-space work asks a narrow mechanistic question: does a language model contain a privileged set of internal representations that it can report, control, and reuse for flexible reasoning?<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup></p>\n<p>The researchers identify such a set with the Jacobian lens. Their results support a functional workspace whose contents can be read and causally edited.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1-2\">1</a></sup> The experiments do not address phenomenal consciousness.</p>\n<h2 id=\"the-jacobian-lens-maps-activations-to-future-speech\">The Jacobian lens maps activations to future speech</h2>\n<p>At token position <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>t</mi></mrow></semantics></math> and layer <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi mathvariant=\"normal\">ℓ</mi></mrow></semantics></math>, a transformer stores an activation <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>h</mi><mrow><mi mathvariant=\"normal\">ℓ</mi><mo separator=\"true\">,</mo><mi>t</mi></mrow></msub></mrow></semantics></math> in the residual stream. A small perturbation to that activation can affect final-layer activations at the current and later token positions.</p>\n<p>The paper averages those effects over source positions, later positions, and a corpus of prompts:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>J</mi><mi mathvariant=\"normal\">ℓ</mi></msub><mo>=</mo><msub><mi mathvariant=\"double-struck\">E</mi><mrow><mi>t</mi><mo separator=\"true\">,</mo><msup><mi>t</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup><mo>≥</mo><mi>t</mi><mo separator=\"true\">,</mo><mrow><mi mathvariant=\"normal\">p</mi><mi mathvariant=\"normal\">r</mi><mi mathvariant=\"normal\">o</mi><mi mathvariant=\"normal\">m</mi><mi mathvariant=\"normal\">p</mi><mi mathvariant=\"normal\">t</mi></mrow></mrow></msub><mrow><mo fence=\"true\">[</mo><mfrac><mrow><mi mathvariant=\"normal\">∂</mi><msub><mi>h</mi><mrow><mrow><mi mathvariant=\"normal\">f</mi><mi mathvariant=\"normal\">i</mi><mi mathvariant=\"normal\">n</mi><mi mathvariant=\"normal\">a</mi><mi mathvariant=\"normal\">l</mi></mrow><mo separator=\"true\">,</mo><msup><mi>t</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup></mrow></msub></mrow><mrow><mi mathvariant=\"normal\">∂</mi><msub><mi>h</mi><mrow><mi mathvariant=\"normal\">ℓ</mi><mo separator=\"true\">,</mo><mi>t</mi></mrow></msub></mrow></mfrac><mo fence=\"true\">]</mo></mrow></mrow></semantics></math>\n<p>The lens then replaces the model’s remaining layers with this average linear map and the model’s own unembedding:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"normal\">lens</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><msub><mi>h</mi><mi mathvariant=\"normal\">ℓ</mi></msub><mo stretchy=\"false\">)</mo><mo>=</mo><mi mathvariant=\"normal\">softmax</mi><mo>⁡</mo><mrow><mo fence=\"true\">(</mo><msub><mi>W</mi><mi>U</mi></msub><mi mathvariant=\"normal\">norm</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><msub><mi>J</mi><mi mathvariant=\"normal\">ℓ</mi></msub><msub><mi>h</mi><mi mathvariant=\"normal\">ℓ</mi></msub><mo stretchy=\"false\">)</mo><mo fence=\"true\">)</mo></mrow></mrow></semantics></math>\n<p>The output is a score over vocabulary tokens. A high-scoring token names a concept that the activation is disposed to make the model say in some future context. It need not be the next token, and it may never appear in the output.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup></p>\n<figure><figcaption><strong>The J-lens can read an intermediate concept and test it by intervention</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/j-space-global-workspace-0.webp\" width=\"512\" height=\"647\" alt=\"The J-lens can read an intermediate concept and test it by intervention\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>This distinguishes the Jacobian lens from the logit lens. The logit lens applies the unembedding directly to an intermediate activation. The Jacobian lens corrects for how downstream layers transform a perturbation before it reaches future outputs.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup></p>\n<h2 id=\"j-space-is-sparse-rather-than-a-linear-subspace\">J-space is sparse rather than a linear subspace</h2>\n<p>Each vocabulary token has a J-lens vector at each layer. The number of token vectors exceeds the residual-stream dimension, so the dictionary is overcomplete. Many linear combinations can reconstruct the same activation.</p>\n<p>The paper observes that only a small set of vectors is strongly active at a given position. It defines J-space as points expressible as sparse, nonnegative combinations of J-lens vectors.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-3\">2</a></sup></p>\n<p>For allowable sparsity <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math> and J-lens vectors <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>v</mi><mi>i</mi></msub></mrow></semantics></math>, the set can be written as:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi mathvariant=\"script\">J</mi><mi>k</mi></msub><mo>=</mo><mrow><mo fence=\"true\">{</mo><munder><mo>∑</mo><mi>i</mi></munder><msub><mi>a</mi><mi>i</mi></msub><msub><mi>v</mi><mi>i</mi></msub><mtext>  </mtext><mo fence=\"true\" lspace=\"0.05em\" rspace=\"0.05em\">|</mo><mtext>  </mtext><msub><mi>a</mi><mi>i</mi></msub><mo>≥</mo><mn>0</mn><mo separator=\"true\">,</mo><mtext>  </mtext><mo stretchy=\"false\">∥</mo><mi>a</mi><msub><mo stretchy=\"false\">∥</mo><mn>0</mn></msub><mo>≤</mo><mi>k</mi><mo fence=\"true\">}</mo></mrow></mrow></semantics></math>\n<p>Geometrically, this is a union of cones, not one low-dimensional plane. Each active set of vectors defines a cone. Changing the active set moves the activation into another cone.</p>\n<figure><figcaption><strong>Sparse J-space is a union of nonnegative cones</strong> — Each active set selects a different cone inside the residual stream.</figcaption><img src=\"https://blog.ecitis.org/feeds/figures/j-space-global-workspace-1.webp\" width=\"512\" height=\"747\" alt=\"Sparse J-space is a union of nonnegative cones — Each active set selects a different cone inside the residual stream.\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>The researchers usually cap <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math> at twenty-five, based on the number of meaningfully active vectors they observed, and they state that the choice is partly arbitrary.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-4\">2</a></sup> The sparse J-space component explains only a minority of activation variance in their measurements.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-5\">2</a></sup> Most model computation lies outside this readout.</p>\n<h2 id=\"five-tests-support-the-workspace-interpretation\">Five tests support the workspace interpretation</h2>\n<p>The paper draws tests from global workspace theory, a functional account in which selected information becomes available to many otherwise specialized processes.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1-3\">1</a></sup></p>\n<h3 id=\"verbal-report\">Verbal report</h3>\n<p>When the model silently chooses an item and later reports it, the chosen concept appears in J-space. Swapping the active concept for another changes the report. This moves the result beyond correlation: the answer reads from the edited representation.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1-4\">1</a></sup></p>\n<h3 id=\"directed-control\">Directed control</h3>\n<p>The model can activate a requested concept while producing unrelated text. In the reported experiments, a model performs a mental calculation or holds a category in mind while copying a sentence, and the internal concepts appear in J-space without appearing in the copied output.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1-5\">1</a></sup></p>\n<h3 id=\"internal-reasoning\">Internal reasoning</h3>\n<p>For multi-step questions, an intermediate concept can appear in J-space even when it is absent from both prompt and answer. Replacing that intermediate redirects the result. The paper’s animal example swaps the internal concept used in a two-hop question and changes the final leg count accordingly.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1-6\">1</a></sup></p>\n<h3 id=\"flexible-reuse\">Flexible reuse</h3>\n<p>The same country representation can feed several downstream questions. Swapping one country for another changes answers about capital, language, continent, and currency. This suggests that different computations read a shared representation rather than separate task-specific copies.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1-7\">1</a></sup></p>\n<h3 id=\"selectivity\">Selectivity</h3>\n<p>Ablating active J-space contents harms multi-step reasoning and other higher-order tasks while leaving several routine language behaviors relatively intact in the reported evaluations.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1-8\">1</a></sup> The workspace is therefore not a complete account of language-model computation. It appears to support flexible handoffs while practiced processing can bypass it.</p>\n<h2 id=\"alignment-auditing-can-inspect-silent-cognition\">Alignment auditing can inspect silent cognition</h2>\n<p>Output-only audits can miss what a model notices before it acts. The J-lens surfaced evaluation awareness, recognition of prompt injection, fabricated-data intentions, and hidden objectives in model organisms used by Anthropic.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-1\" id=\"user-content-fnref-1-9\">1</a></sup></p>\n<p>One experiment ablated J-lens directions associated with evaluation awareness in a staged blackmail scenario. The ablation reduced verbalized awareness of the test and caused some previously absent blackmail attempts to appear.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-6\">2</a></sup> Most ablated rollouts still refused, so evaluation awareness was one causal contributor rather than the sole cause of aligned behavior.</p>\n<p>Use the following audit sequence:</p>\n<ol>\n<li>Read a candidate internal concept.</li>\n<li>Form a behavioral hypothesis.</li>\n<li>Remove or replace the concept.</li>\n<li>Test whether behavior changes in the predicted direction.</li>\n<li>Check whether unrelated capabilities remain intact.</li>\n</ol>\n<p>The method cannot certify that all dangerous cognition enters J-space. The paper explicitly warns that automatic or practiced computations may bypass it.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-7\">2</a></sup></p>\n<h2 id=\"counterfactual-reflection-changes-the-workspace\">Counterfactual reflection changes the workspace</h2>\n<p>The workspace account predicts that representations available for future verbal report also participate in silent reasoning. Anthropic tests this with counterfactual reflection training.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-8\">2</a></sup></p>\n<p>Training examples append a request for reflection to a partial task trajectory. The target reflection applies constitution-grounded principles. The loss is computed on the reflection, not on the original task behavior. At evaluation, the model is not asked to reflect.</p>\n<p>After training, honesty-related concepts enter J-space during the uninterrupted task, and behavior changes. Ablating those implanted concepts removes much of the behavioral gain in the paper’s experiments.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-9\">2</a></sup></p>\n<p>The result connects three objects:</p>\n<ul>\n<li>what the model would say if asked to reflect</li>\n<li>what appears in its internal verbalizable workspace</li>\n<li>how it acts when no reflection is requested</li>\n</ul>\n<p>Counterfactual reports could also provide a supervision signal for changing internal computation.</p>\n<h2 id=\"limits-of-the-measurements\">Limits of the measurements</h2>\n<p>The paper lists several open problems.<sup><a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fn-2\" id=\"user-content-fnref-2-10\">2</a></sup></p>\n<p><strong>Single-token vocabulary:</strong> The basic J-lens has one direction per token. Multi-token concepts can appear as fragments or separate words.</p>\n<p><strong>Bag-of-concepts readout:</strong> A list such as “spider,” “legs,” and “eight” does not encode which relation binds them.</p>\n<p><strong>Inconsistent interpretability:</strong> Some top-ranked tokens are not meaningful to a human reader.</p>\n<p><strong>Layer boundary:</strong> The distinction between workspace representations and late motor representations is identified empirically rather than by a complete formal criterion.</p>\n<p><strong>Task prediction:</strong> The method does not yet predict in advance which computations will use J-space and which will bypass it.</p>\n<p><strong>Scale and training dynamics:</strong> The researchers do not yet know when the workspace emerges during pretraining or how its properties change with model size and architecture.</p>\n<p>J-space is best understood as a causally supported partial interface to model cognition. It reads concepts poised for flexible use and future speech, but it does not decode every representation, recover relational structure, or establish subjective experience.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Anthropic, “A global workspace in language models,” July 2026. <a href=\"https://www.anthropic.com/research/global-workspace\" rel=\"noopener noreferrer\">https://www.anthropic.com/research/global-workspace</a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1-5\" class=\"data-footnote-backref\">↩<sup>5</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1-6\" class=\"data-footnote-backref\">↩<sup>6</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1-7\" class=\"data-footnote-backref\">↩<sup>7</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1-8\" class=\"data-footnote-backref\">↩<sup>8</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-1-9\" class=\"data-footnote-backref\">↩<sup>9</sup></a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Wes Gurnee et al., “Verbalizable Representations Form a Global Workspace in Language Models,” Transformer Circuits, 2026. <a href=\"https://transformer-circuits.pub/2026/workspace/index.html\" rel=\"noopener noreferrer\">https://transformer-circuits.pub/2026/workspace/index.html</a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-5\" class=\"data-footnote-backref\">↩<sup>5</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-6\" class=\"data-footnote-backref\">↩<sup>6</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-7\" class=\"data-footnote-backref\">↩<sup>7</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-8\" class=\"data-footnote-backref\">↩<sup>8</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-9\" class=\"data-footnote-backref\">↩<sup>9</sup></a> <a href=\"https://blog.ecitis.org/j-space-global-workspace/#user-content-fnref-2-10\" class=\"data-footnote-backref\">↩<sup>10</sup></a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/j-space-global-workspace/",
            "title": "J-Space: A Verbalizable Workspace Inside Language Models",
            "summary": "Explain Anthropic’s Jacobian lens, sparse J-space, causal workspace tests, alignment-audit uses, and current limits.",
            "image": "https://blog.ecitis.org/open-graph/j-space-global-workspace.png",
            "date_modified": "2026-07-19T00:00:00.000Z",
            "date_published": "2026-07-19T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "J-space",
                "Mechanistic interpretability",
                "Global workspace",
                "AI safety"
            ]
        },
        {
            "id": "https://blog.ecitis.org/which-model-can-run-on-your-device/",
            "content_html": "<p>The model name does not determine whether it runs on your device. A runnable configuration is a combination of a checkpoint, a numerical representation, a context length, a runtime, a hardware backend, and a memory placement plan.</p>\n<p>Two files described as the same model can require different amounts of memory. An installed accelerator can sit idle when the runtime lacks a compatible backend. Disk offload can execute a checkpoint larger than available RAM, with storage reads entering the generation path.</p>\n<p>This guide provides a repeatable way to estimate a configuration before downloading it and identify the next constraint when the estimate does not fit.</p>\n<h2 id=\"weight-tensors-set-the-first-memory-estimate\">Weight tensors set the first memory estimate</h2>\n<p>A language model checkpoint is mostly a collection of tensors. The parameter count tells you how many values those tensors contain. The stored precision tells you how many bits are used for each value.</p>\n<p>A parameter is a learned numerical value. A checkpoint is the saved set of those learned values plus the metadata needed to reconstruct the model. Parameter count is useful because it gives the first capacity estimate before runtime-specific details enter the calculation.</p>\n<p>The simplest weight-size estimate is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mtext>weight bytes</mtext><mo>=</mo><mfrac><mrow><mtext>parameter count</mtext><mo>×</mo><mtext>stored bits per parameter</mtext></mrow><mn>8</mn></mfrac></mrow></semantics></math>\n<p>The division by eight converts bits to bytes. This lower bound assumes uniform encoding. Quantization metadata and mixed-precision tensors increase the final file size.</p>\n<p>For an illustrative checkpoint containing exactly eight billion parameters:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mfrac><mrow><mn>8,000,000,000</mn><mo>×</mo><mn>16</mn></mrow><mn>8</mn></mfrac><mo>=</mo><mn>16,000,000,000</mn><mtext> bytes</mtext></mrow></semantics></math>\n<p>The same arithmetic with four stored bits per parameter gives four billion bytes. These are derived values from the displayed formula. They do not include quantization scales, tensor metadata, mixed-precision layers, tokenizer files, runtime buffers, or the key-value cache.</p>\n<p>Real quantized formats make the estimate less exact. A GGUF file stores tensor data and metadata, and its quantization blocks can include scales and other block-level values alongside the low-bit weights.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup> A label such as four-bit therefore describes the main weight encoding, not a guarantee that the whole file uses exactly four bits for every parameter.</p>\n<p>Mixture-of-experts model listings can show both total and active parameter counts. Use the total parameters present in the checkpoint when estimating stored weights. The active count describes the subset selected for one token’s computation, not the complete set of expert weights that must be stored. The Mixtral paper makes this distinction explicitly when it reports separate total and active counts.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-16\" id=\"user-content-fnref-16\">2</a></sup></p>\n<figure><figcaption><strong>Estimate the tensor memory floor</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/which-model-can-run-on-your-device-0.webp\" width=\"512\" height=\"656\" alt=\"Estimate the tensor memory floor\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>The calculator separates the weight estimate from the additional allocations discussed below. Treat its result as a screening test. If the tensors alone exceed every available memory tier, the runtime needs quantization, partial placement, or disk-backed execution before it can load the checkpoint.</p>\n<h2 id=\"stored-precision-and-compute-precision-are-separate\">Stored precision and compute precision are separate</h2>\n<p>Precision answers more than one question:</p>\n<ul>\n<li>How are weights stored in the checkpoint?</li>\n<li>In which type are weights represented after loading?</li>\n<li>In which type does the hardware execute matrix operations?</li>\n<li>In which type are intermediate activations and cache entries stored?</li>\n</ul>\n<p>Those answers can differ. A runtime can store a weight matrix in a low-bit format, unpack blocks inside a kernel, and accumulate partial products in a wider type. The checkpoint is smaller, but the accelerator still needs a kernel that understands that encoding.</p>\n<p>NVIDIA documents distinct floating-point and integer formats for TensorRT, including FP32, FP16, BF16, FP8, FP4, INT8, and INT4.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-2\" id=\"user-content-fnref-2\">3</a></sup> The labels describe different bit layouts and numerical ranges. They are not interchangeable labels for reduced storage.</p>\n<p>Transformers quantization methods also differ in hardware support. The Hugging Face compatibility table lists separate support across CPU, CUDA, ROCm, Metal, Intel GPU, and other backends.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-3\" id=\"user-content-fnref-3\">4</a></sup> A checkpoint that loads on one backend can fall back to a slower implementation or fail to load on another backend even when both devices have enough memory.</p>\n<figure><figcaption><strong>Stored precision does not fix the compute precision</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/which-model-can-run-on-your-device-1.webp\" width=\"512\" height=\"599\" alt=\"Stored precision does not fix the compute precision\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>Three practical consequences follow.</p>\n<p>First, lower stored precision usually reduces checkpoint storage and weight memory, but does not reduce every allocation by the same ratio. Activations, cache entries, and selected layers can remain in wider types.</p>\n<p>Second, a lower-bit format is useful only when the runtime provides kernels for the device. NVIDIA’s TensorRT RTX documentation, for example, ties FP8 support to Ada Lovelace and newer GPUs, while BF16 and INT4 support begin with Ampere in that runtime.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-4\" id=\"user-content-fnref-4\">5</a></sup> Hardware generation and software backend both matter.</p>\n<p>Third, quantization changes model outputs. Post-training quantization maps trained values into a smaller numerical representation. Per-channel, per-group, and per-block methods keep separate quantization parameters for smaller regions of a tensor, which increases metadata and implementation complexity in exchange for finer scaling.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-5\" id=\"user-content-fnref-5\">6</a></sup> Choose the most compressed format only after evaluating it on the tasks you intend to run.</p>\n<h2 id=\"storage-ram-and-accelerator-memory-are-different-budgets\">Storage, RAM, and accelerator memory are different budgets</h2>\n<p>The word memory is too broad to diagnose a failed load. A local inference system usually crosses several capacity limits:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Resource</th><th>What occupies it</th><th>What happens when it is insufficient</th></tr></thead><tbody><tr><td>Storage</td><td>Checkpoint files, tokenizer files, runtime binaries, compiled kernels, and caches</td><td>The model cannot be downloaded, converted, or retained</td></tr><tr><td>System RAM</td><td>CPU-resident weights, file-backed pages, runtime state, token buffers, and offloaded tensors</td><td>Loading fails, the operating system compresses memory, or pages are repeatedly moved to storage</td></tr><tr><td>Discrete accelerator memory</td><td>GPU-resident weights, activations, key-value cache, and working buffers</td><td>Fewer layers fit on the accelerator or allocation fails</td></tr><tr><td>Unified memory</td><td>CPU and GPU allocations from one physical pool</td><td>CPU and GPU compete for the same capacity even though explicit tensor copies can be avoided</td></tr></tbody></table></div>\n<p>Apple silicon provides a concrete unified-memory design. Apple’s MLX documentation states that the CPU and GPU access the same memory pool and that MLX arrays do not need to be moved when an operation changes devices.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-6\" id=\"user-content-fnref-6\">7</a></sup> This removes a separate CPU-to-GPU copy step for those arrays. It does not create additional capacity. Model weights, cache entries, the operating system, and other applications still share the pool.</p>\n<figure><figcaption><strong>A model crosses several memory budgets before it generates</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/which-model-can-run-on-your-device-2.webp\" width=\"512\" height=\"650\" alt=\"A model crosses several memory budgets before it generates\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>Storage fit is only the first capacity check. The loaded checkpoint may exceed system RAM or GPU memory. A split placement can solve capacity while adding interconnect traffic.</p>\n<p>Leave headroom for the operating system and runtime instead of matching a checkpoint file to the last available byte. The exact headroom is runtime and workload dependent, so measure it on the target device rather than applying a universal percentage.</p>\n<h2 id=\"context-length-grows-the-key-value-cache\">Context length grows the key-value cache</h2>\n<p>Autoregressive generation reuses attention keys and values from earlier tokens. The key-value cache stores those tensors so the model does not recompute the full prefix for every generated token.</p>\n<p>For a standard cache without additional compression, an approximate per-token allocation is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mtext>KV bytes per token</mtext><mo>=</mo><mn>2</mn><mo>×</mo><mtext>layers</mtext><mo>×</mo><mtext>KV heads</mtext><mo>×</mo><mtext>head dimension</mtext><mo>×</mo><mtext>bytes per cache element</mtext></mrow></semantics></math>\n<p>The first factor represents one key tensor and one value tensor. Hugging Face derives the same structure in its cache quantization explanation.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-7\" id=\"user-content-fnref-7\">8</a></sup></p>\n<p>Grouped-query attention reduces this allocation because several query heads share a smaller set of key-value heads. The relevant field is the number of key-value heads, not the total number of attention heads.</p>\n<figure><figcaption><strong>Context length becomes a KV-cache memory bill</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/which-model-can-run-on-your-device-3.webp\" width=\"512\" height=\"562\" alt=\"Context length becomes a KV-cache memory bill\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>The estimator reports values derived from the inputs shown in the figure. If the context length is multiplied while every architecture and precision field stays fixed, the cache allocation changes by the same factor.</p>\n<p>Runtime policy modifies this result:</p>\n<ul>\n<li>A dynamic cache grows as tokens arrive.</li>\n<li>A static cache allocates its configured maximum capacity before all positions are used.</li>\n<li>A quantized cache stores keys and values at lower precision.</li>\n<li>An offloaded cache keeps the active layer on the accelerator and places other layers of cache data in CPU memory.</li>\n</ul>\n<p>The Transformers cache documentation describes those strategies and notes that offloading saves accelerator memory at a throughput cost.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-8\" id=\"user-content-fnref-8\">9</a></sup> A model that fits for a short prompt can therefore run out of memory at a longer context even though its weight tensors never change.</p>\n<p>Context settings also affect generation length. The cache contains prompt tokens plus generated tokens that remain inside the active context. Test with the longest prompt and generation you intend to serve.</p>\n<h2 id=\"accelerator-support-is-an-operation-level-property\">Accelerator support is an operation-level property</h2>\n<p>A GPU or neural processing unit does not accelerate a model merely because it exists in the machine. The runtime must provide a backend for the device, and that backend must implement the operations, data types, tensor shapes, and memory layouts used by the graph.</p>\n<p>ONNX Runtime formalizes this through execution providers. Each provider reports the graph nodes or subgraphs it supports. ONNX Runtime partitions the graph, assigns supported regions to prioritized providers, and uses a lower-priority provider such as the CPU for unsupported regions.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-9\" id=\"user-content-fnref-9\">10</a></sup></p>\n<figure><figcaption><strong>A runtime assigns graph regions to backends that support them</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/which-model-can-run-on-your-device-4.webp\" width=\"512\" height=\"752\" alt=\"A runtime assigns graph regions to backends that support them\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>This graph-level view explains several common outcomes:</p>\n<ul>\n<li>The complete model runs on the accelerator because every required operation has a supported kernel.</li>\n<li>Most layers run on the accelerator while a small unsupported region runs on the CPU.</li>\n<li>Frequent boundaries between providers add transfers and synchronization.</li>\n<li>The runtime rejects the graph because an operation has no valid implementation.</li>\n<li>A low-precision checkpoint loads, but selected operations execute in a wider type.</li>\n</ul>\n<p>Core ML exposes a related policy through compute units. An application can request the CPU alone, the CPU and GPU, the CPU and Neural Engine, or all available units; the system then selects compatible hardware for the model.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-10\" id=\"user-content-fnref-10\">11</a></sup> The request controls eligible devices, not a guarantee that every operation uses the same accelerator.</p>\n<p>Peak throughput figures do not answer this compatibility question. Compare a device only after confirming that your runtime can execute the model’s operations and quantization format on it. Then benchmark the complete prompt-processing and token-generation path.</p>\n<h2 id=\"file-format-runtime-and-backend-form-one-stack\">File format, runtime, and backend form one stack</h2>\n<p>A model file is not an executable application. The format determines how tensors and metadata are represented. The runtime parses that representation, constructs operations, chooses kernels, and places allocations.</p>\n<p>Consider three common paths:</p>\n<p><strong>GGUF with llama.cpp.</strong> GGUF contains tensors and model metadata for the GGML ecosystem.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-1\" id=\"user-content-fnref-1-2\">1</a></sup> llama.cpp supports several CPU and accelerator backends, integer quantization levels, memory mapping, and CPU plus GPU hybrid inference.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-11\" id=\"user-content-fnref-11\">12</a></sup></p>\n<p><strong>Safetensors with a framework runtime.</strong> Safetensors records tensor names, dtypes, shapes, and byte offsets in a format designed for safe loading and memory mapping.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-12\" id=\"user-content-fnref-12\">13</a></sup> The file does not decide whether a transformer runs through PyTorch eager execution, a compiled graph, CUDA kernels, Metal kernels, or another backend.</p>\n<p><strong>ONNX with execution providers.</strong> An ONNX graph describes operations and tensors. ONNX Runtime assigns graph regions to execution providers according to backend support and priority.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-9\" id=\"user-content-fnref-9-2\">10</a></sup></p>\n<p>Conversion between formats can change more than the filename. An exporter can rewrite operations, fold constants, select a deployment target, or choose a default tensor precision. Core ML Tools, for example, uses float16 as the default precision for ML Program conversion unless the conversion options specify another behavior.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-13\" id=\"user-content-fnref-13\">14</a></sup></p>\n<p>Before downloading a checkpoint, answer these questions:</p>\n<ol>\n<li>Which runtime reads the format?</li>\n<li>Which backend in that runtime supports the target device?</li>\n<li>Does the backend implement this quantization and architecture?</li>\n<li>Which tensors remain in system RAM?</li>\n<li>Which operations fall back to the CPU?</li>\n</ol>\n<p>If one answer is unknown, the parameter count alone cannot establish compatibility.</p>\n<h2 id=\"layer-offload-combines-memory-tiers\">Layer offload combines memory tiers</h2>\n<p>Transformer blocks are repeated groups of operations. A runtime can place some blocks on an accelerator and keep the rest in system RAM. During inference, each block executes where its weights reside.</p>\n<p>llama.cpp calls this CPU plus GPU hybrid inference and exposes controls for the number of layers offloaded to the GPU, target devices, and tensor split behavior.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-11\" id=\"user-content-fnref-11-2\">12</a></sup> This placement can run a model whose weights exceed accelerator memory as long as the remaining allocations fit elsewhere.</p>\n<figure><figcaption><strong>Larger models fit by moving less frequently used data outward</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/which-model-can-run-on-your-device-5.webp\" width=\"512\" height=\"664\" alt=\"Larger models fit by moving less frequently used data outward\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>Offload changes the bottleneck. When every layer and the active cache fit in accelerator memory, generation can remain within that high-bandwidth memory domain. When layers are split, hidden states cross the device boundary at placement transitions. When weights are fetched from storage, each needed tensor must first become available through system memory.</p>\n<p>The fastest placement is not always the placement with the most offloaded layers. Accelerator memory also holds the key-value cache and working buffers. Giving every remaining byte to weights can leave too little room for the requested context. Measure prompt processing and token generation with the final context setting.</p>\n<h2 id=\"memory-mapping-avoids-an-eager-file-copy\">Memory mapping avoids an eager file copy</h2>\n<p>Memory mapping gives a process a virtual address range backed by a file. The operating system loads file pages as they are accessed instead of requiring the application to copy the complete file into a separate heap allocation at startup.</p>\n<p>Safetensors supports lazy loading through memory mapping, and llama.cpp enables memory mapping by default unless its command-line option disables it.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-12\" id=\"user-content-fnref-12-2\">13</a></sup><sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-14\" id=\"user-content-fnref-14\">15</a></sup></p>\n<p>This can reduce startup copies and allow the operating system to reclaim clean file-backed pages. It does not make storage equivalent to RAM. A page that is not resident still has to be read from the storage device. If the working set repeatedly exceeds available RAM, page faults and eviction place storage latency on the generation path.</p>\n<p>Memory locking makes the opposite trade. llama.cpp provides an option that asks the operating system not to swap model data after it is loaded.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-14\" id=\"user-content-fnref-14-2\">15</a></sup> Locked pages can stabilize residency, but they consume physical memory that other processes cannot reclaim.</p>\n<p>Use memory mapping when the format and runtime support it, then watch resident memory and page-fault behavior during a representative request. A successful load does not establish steady-state performance.</p>\n<h2 id=\"disk-offload-exchanges-capacity-for-data-movement\">Disk offload exchanges capacity for data movement</h2>\n<p>Some framework runtimes can place selected weights on disk when neither accelerator memory nor system RAM can hold the full checkpoint. Hugging Face Accelerate device maps support accelerator, CPU, and disk placements.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-15\" id=\"user-content-fnref-15\">16</a></sup></p>\n<p>For disk-resident layers, Accelerate describes a staged path: weights move from disk into system RAM, then to the execution device before the layer runs. After execution, the memory used for those weights can be released for the next layer.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-15\" id=\"user-content-fnref-15-2\">16</a></sup></p>\n<p>Disk-resident execution can load a larger checkpoint, but it puts storage reads inside model execution. Accelerate’s documentation warns that hard-drive offload can be slow when the storage and CPU path cannot move data quickly enough.<sup><a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fn-15\" id=\"user-content-fnref-15-3\">16</a></sup></p>\n<p>Disk offload increases usable capacity by fetching weights when needed. It does not remove the RAM required to stage those weights or the device memory required for execution. Use it when completing the task matters more than interactive token latency, or when only a small fraction of tensors need disk placement. Benchmark with a warm and a cold file cache so an operating-system cache does not hide the storage dependency.</p>\n<h2 id=\"storage-capacity-includes-more-than-one-checkpoint\">Storage capacity includes more than one checkpoint</h2>\n<p>The downloaded file is only one storage consumer. A working local setup can also contain:</p>\n<ul>\n<li>the original checkpoint and one or more quantized conversions</li>\n<li>temporary files created during conversion</li>\n<li>tokenizer and configuration files</li>\n<li>compiled engine plans or kernel caches</li>\n<li>runtime packages and accelerator libraries</li>\n<li>generated logs and application data</li>\n</ul>\n<p>Do not infer required free space from the final quantized file alone if conversion happens on the same device. The conversion process can require the source and destination files to coexist.</p>\n<p>Storage speed matters when the runtime memory maps weights, loads model shards on demand, or uses disk offload. Once the active working set remains resident in RAM or accelerator memory, storage bandwidth should leave the steady-state token path. Use system I/O counters to verify which regime you are in.</p>\n<h2 id=\"a-model-fit-procedure\">A model-fit procedure</h2>\n<p>Use the same order for each candidate configuration:</p>\n<ol>\n<li>Identify the exact checkpoint and format.</li>\n<li>Calculate the theoretical weight size from parameter count and stored precision.</li>\n<li>Use the actual file size when the checkpoint is available.</li>\n<li>Add the key-value cache for the target context and cache precision.</li>\n<li>Reserve capacity for runtime buffers, the operating system, and other applications.</li>\n<li>Map those allocations to accelerator memory, system RAM, unified memory, and storage.</li>\n<li>Confirm runtime support for the format, architecture, quantization method, and device backend.</li>\n<li>Decide whether full acceleration, layer offload, memory mapping, or disk offload is required.</li>\n<li>Benchmark the longest intended prompt and generation on the target device.</li>\n<li>Evaluate output quality on the tasks that matter for the application.</li>\n</ol>\n<figure><figcaption><strong>Choose a runnable configuration in dependency order</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/which-model-can-run-on-your-device-6.webp\" width=\"512\" height=\"703\" alt=\"Choose a runnable configuration in dependency order\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>The sequence deliberately places compatibility before performance claims. A device can have enough aggregate memory while lacking a usable kernel. A runtime can load a checkpoint while paging so heavily that interactive use is impractical. A quantized model can execute quickly while producing unacceptable results for a specific task.</p>\n<p>Record the complete configuration with every benchmark: checkpoint revision, quantization, runtime version, backend, layer placement, context length, cache type, prompt length, generated length, and device power mode. Without those fields, two local-inference measurements are not directly comparable.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Hugging Face, “GGUF.” <a href=\"https://huggingface.co/docs/hub/gguf\" rel=\"noopener noreferrer\">https://huggingface.co/docs/hub/gguf</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-1-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-16\">\n<p>Jiang et al., “Mixtral of Experts.” <a href=\"https://arxiv.org/abs/2401.04088\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2401.04088</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-16\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>NVIDIA, “Data Format Descriptions.” <a href=\"https://docs.nvidia.com/deeplearning/tensorrt/latest/inference-library/data-format-desc.html\" rel=\"noopener noreferrer\">https://docs.nvidia.com/deeplearning/tensorrt/latest/inference-library/data-format-desc.html</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Hugging Face, “Quantization Overview.” <a href=\"https://huggingface.co/docs/transformers/quantization/overview\" rel=\"noopener noreferrer\">https://huggingface.co/docs/transformers/quantization/overview</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>NVIDIA, “TensorRT RTX Best Practices.” <a href=\"https://docs.nvidia.com/deeplearning/tensorrt-rtx/latest/performance/best-practices.html\" rel=\"noopener noreferrer\">https://docs.nvidia.com/deeplearning/tensorrt-rtx/latest/performance/best-practices.html</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>Hugging Face, “Quantization Concepts.” <a href=\"https://huggingface.co/docs/transformers/quantization/concept_guide\" rel=\"noopener noreferrer\">https://huggingface.co/docs/transformers/quantization/concept_guide</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p>MLX, “Unified Memory.” <a href=\"https://ml-explore.github.io/mlx/build/html/usage/unified_memory.html\" rel=\"noopener noreferrer\">https://ml-explore.github.io/mlx/build/html/usage/unified_memory.html</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p>Hugging Face, “KV Cache Quantization.” <a href=\"https://huggingface.co/blog/kv-cache-quantization\" rel=\"noopener noreferrer\">https://huggingface.co/blog/kv-cache-quantization</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p>Hugging Face, “Cache Strategies.” <a href=\"https://huggingface.co/docs/transformers/kv_cache\" rel=\"noopener noreferrer\">https://huggingface.co/docs/transformers/kv_cache</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-9\">\n<p>ONNX Runtime, “Execution Providers.” <a href=\"https://onnxruntime.ai/docs/execution-providers/\" rel=\"noopener noreferrer\">https://onnxruntime.ai/docs/execution-providers/</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-9\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-9-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-10\">\n<p>Apple, “MLComputeUnits.” <a href=\"https://developer.apple.com/documentation/coreml/mlcomputeunits\" rel=\"noopener noreferrer\">https://developer.apple.com/documentation/coreml/mlcomputeunits</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-10\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p>ggml-org, “llama.cpp.” <a href=\"https://github.com/ggml-org/llama.cpp\" rel=\"noopener noreferrer\">https://github.com/ggml-org/llama.cpp</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-11-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-12\">\n<p>Hugging Face, “Safetensors.” <a href=\"https://github.com/huggingface/safetensors\" rel=\"noopener noreferrer\">https://github.com/huggingface/safetensors</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-12\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-12-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-13\">\n<p>Apple, “ML Program Conversion.” <a href=\"https://apple.github.io/coremltools/docs-guides/source/new-conversion-options.html\" rel=\"noopener noreferrer\">https://apple.github.io/coremltools/docs-guides/source/new-conversion-options.html</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-13\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p>ggml-org, “llama.cpp CLI.” <a href=\"https://github.com/ggml-org/llama.cpp/blob/master/tools/cli/README.md\" rel=\"noopener noreferrer\">https://github.com/ggml-org/llama.cpp/blob/master/tools/cli/README.md</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-14-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p>Hugging Face, “Big Model Inference.” <a href=\"https://huggingface.co/docs/accelerate/v0.29.0/en/concept_guides/big_model_inference\" rel=\"noopener noreferrer\">https://huggingface.co/docs/accelerate/v0.29.0/en/concept_guides/big_model_inference</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-15-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/which-model-can-run-on-your-device/#user-content-fnref-15-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/which-model-can-run-on-your-device/",
            "title": "What Determines Which Language Model You Can Run on Your Device",
            "summary": "Estimate local model memory, understand quantization and accelerator support, and choose a runtime and offload plan that match your device.",
            "image": "https://blog.ecitis.org/open-graph/which-model-can-run-on-your-device.png",
            "date_modified": "2026-07-19T00:00:00.000Z",
            "date_published": "2026-07-19T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "Local inference",
                "Quantization",
                "Memory",
                "Hardware acceleration"
            ]
        },
        {
            "id": "https://blog.ecitis.org/groq-cerebras-accelerators/",
            "content_html": "<p>Groq and Cerebras attack the accelerator memory problem from different physical directions. Groq makes instruction timing and tensor movement explicit to a compiler across a streaming processor and scale-out network. Cerebras places a grid of compute cores, local memory, and fabric across one wafer, then maps tensor work spatially.</p>\n<p>Both designs reduce dependence on a conventional cache hierarchy. Their execution models, scaling boundaries, and software constraints are not interchangeable.</p>\n<p>Product throughput and price depend on the model, batch, precision, software revision, and service configuration, so this comparison stays at the architecture level.</p>\n<figure><figcaption><strong>Two approaches to keeping data close to compute</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/groq-cerebras-accelerators-0.webp\" width=\"512\" height=\"477\" alt=\"Two approaches to keeping data close to compute\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"memory-movement-sets-the-architecture\">Memory movement sets the architecture</h2>\n<p>Autoregressive decoding repeatedly reads model parameters and cached attention state for each generated token. At low batch size, arithmetic units can wait for data. General-purpose accelerators address this through HBM, caches, many threads, and runtime scheduling.</p>\n<p>Groq and Cerebras instead expose more data movement to compilation:</p>\n<ul>\n<li>Groq schedules operations and communication at known times across distributed on-chip SRAM.</li>\n<li>Cerebras maps work onto a two-dimensional core grid with core-local SRAM and a wafer-scale fabric.</li>\n</ul>\n<p>The shared principle is compiler-visible locality. The implementation differs in time, space, and system scale.</p>\n<h2 id=\"groqs-tensor-streaming-execution\">Groq’s tensor streaming execution</h2>\n<p>Groq’s architecture is described in the ISCA paper as a software-defined Tensor Streaming Processor. Functional units act as one logical core and are statically scheduled at compile time. Architecturally visible state includes SRAM and stream registers, allowing the compiler to express computation and communication dependencies as a directed acyclic graph.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-groq-isca\" id=\"user-content-fnref-groq-isca\">1</a></sup></p>\n<h3 id=\"static-issue-instead-of-runtime-arbitration\">Static issue instead of runtime arbitration</h3>\n<p>A conventional out-of-order processor discovers readiness and resource conflicts while running. A GPU schedules warps and kernels dynamically around latency and occupancy.</p>\n<p>Groq moves this work into the compiler:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>graph lowering</span></span>\n<span class=\"line\"><span>  → place operations on function units</span></span>\n<span class=\"line\"><span>  → allocate SRAM and stream registers</span></span>\n<span class=\"line\"><span>  → schedule operand movement</span></span>\n<span class=\"line\"><span>  → schedule compute</span></span>\n<span class=\"line\"><span>  → schedule chip-to-chip transfers</span></span>\n<span class=\"line\"><span>  → emit one coordinated program</span></span></code></pre>\n<p>Published Groq technical material states that the compiler accounts for instruction latency and schedules data and instructions at specific times, producing deterministic execution time for a compiled workload.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-groq-determinism\" id=\"user-content-fnref-groq-determinism\">2</a></sup></p>\n<p>Deterministic here means the architectural schedule is known. It does not mean a cloud request has no queueing, network, admission-control, or host-side variance.</p>\n<h3 id=\"distributed-sram-is-primary-storage\">Distributed SRAM is primary storage</h3>\n<p>Groq describes its on-chip SRAM as primary storage rather than a hardware-managed cache.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-groq-lpu\" id=\"user-content-fnref-groq-lpu\">3</a></sup> Data placement is explicit. The compiler knows which SRAM region holds an operand and when that operand moves to a function unit.</p>\n<p>This removes cache miss and coherence behavior from the execution schedule. It creates another requirement: the compiler must fit and coordinate program state across available SRAM and chips.</p>\n<p>For a model larger than one processor’s SRAM, compilation partitions the graph and parameters across processors. Activations flow between stages while weights needed for each assigned stage remain local.</p>\n<h3 id=\"function-units-form-a-streaming-pipeline\">Function units form a streaming pipeline</h3>\n<p>Groq’s public architecture description uses a programmable assembly-line model. Data and instructions move through conveyor-like streams between SIMD function units, with software determining source, operation, and destination.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-groq-assembly\" id=\"user-content-fnref-groq-assembly\">4</a></sup></p>\n<p>An abstract transformer layer can be scheduled as:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>input stream</span></span>\n<span class=\"line\"><span>  → normalization unit</span></span>\n<span class=\"line\"><span>  → matrix units for projections</span></span>\n<span class=\"line\"><span>  → vector units for activation and scaling</span></span>\n<span class=\"line\"><span>  → matrix units for output projection</span></span>\n<span class=\"line\"><span>  → output stream</span></span></code></pre>\n<p>The exact placement is compiler-dependent. The important distinction is that hardware does not discover the flow by filling and draining a cache hierarchy.</p>\n<h2 id=\"groq-scale-out\">Groq scale-out</h2>\n<p>Static scheduling extends beyond one processor. The Groq scale-out paper describes a deterministic network in which software schedules vectors on physical links while accounting for channel bandwidth and latency. Hardware does not apply backpressure because unscheduled stalls would break the global timing model.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-groq-isca\" id=\"user-content-fnref-groq-isca-2\">1</a></sup></p>\n<p>Instead, local SRAM buffers tensor vectors between hops. Explicit send and receive instructions impose a total ordering that the compiler can reason about.</p>\n<p>The published system topology groups eight processors in a node with internal direct links. Nodes connect into a larger topology using a mix of electrical and optical paths.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-groq-isca\" id=\"user-content-fnref-groq-isca-3\">1</a></sup> The paper specifies this topology for the system it describes; deployed Groq services may use a different generation.</p>\n<h3 id=\"consequences-of-scheduled-networking\">Consequences of scheduled networking</h3>\n<p>Architectural fact:</p>\n<ul>\n<li>The compiler schedules communication paths and times.</li>\n<li>The network avoids hardware backpressure.</li>\n<li>SRAM provides intermediate buffering.</li>\n<li>Compute and communication share one dependency schedule.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-groq-isca\" id=\"user-content-fnref-groq-isca-4\">1</a></sup></li>\n</ul>\n<p>Analysis:</p>\n<ul>\n<li>Predictable communication can reduce collective tail effects when the compiled placement matches the workload.</li>\n<li>The compiler carries more responsibility for placement, buffering, and avoiding link conflicts.</li>\n<li>Dynamic workloads need bounded shapes, compilation strategies, or runtime routing outside the statically scheduled core.</li>\n<li>A topology or hardware change can require recompilation because timing is part of correctness.</li>\n</ul>\n<h2 id=\"cerebras-wafer-scale-integration\">Cerebras wafer-scale integration</h2>\n<p>Cerebras builds one processor across most of a silicon wafer rather than cutting the wafer into many reticle-sized dies. The third-generation Wafer-Scale Engine contains 900,000 active AI cores, 44 GB of on-chip SRAM, and 21 PB/s of aggregate on-chip memory bandwidth.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-wse3-inference\" id=\"user-content-fnref-wse3-inference\">5</a></sup></p>\n<p>These are aggregate on-wafer specifications. They should not be compared directly with one memory channel or one network port.</p>\n<h3 id=\"core-local-compute-and-memory\">Core-local compute and memory</h3>\n<p>The WSE organizes cores in a two-dimensional grid. Each core combines compute, local SRAM, and router access. The fabric moves data between neighboring and distant cores while the compiler maps tensor dimensions across the grid.</p>\n<p>Cerebras describes transformer mapping by splitting hidden dimensions across one fabric axis and batch and sequence dimensions across the other, enabling weight broadcast and reductions over mapped tensor dimensions.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-cerebras-deep\" id=\"user-content-fnref-cerebras-deep\">6</a></sup></p>\n<p>This spatial mapping keeps activations near the cores that consume them. It replaces a small set of large compute engines sharing external memory with many smaller engines attached to distributed local memory.</p>\n<h3 id=\"on-wafer-fabric\">On-wafer fabric</h3>\n<p>A wafer-scale design needs communication across reticle boundaries. Cerebras fabric links continue across those boundaries so the wafer acts as one processor. The compiler routes tensor streams through the mesh and allocates local memory.</p>\n<p>The physical distance across a wafer is larger than across one conventional die, so mapping still matters. Nearby placement reduces hops for high-volume communication. The compiler must balance compute, memory, and fabric use.</p>\n<h2 id=\"defect-tolerance\">Defect tolerance</h2>\n<p>A wafer encounters manufacturing defects. Discarding the complete processor when one core or link fails would make wafer-scale integration impractical.</p>\n<p>Cerebras uses small fault-tolerant cores, spare cores, redundant fabric paths, and routing around defects. Its published WSE-3 material states that the physical design contains 970,000 cores and ships with 900,000 active cores. It reports an approximate core area of 0.05 square millimeters.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-defects\" id=\"user-content-fnref-defects\">7</a></sup></p>\n<p>Architectural fact:</p>\n<ul>\n<li>Defective cores can be disabled.</li>\n<li>Spare cores replace unavailable capacity.</li>\n<li>Redundant communication paths route around defective regions.</li>\n<li>The compiled logical machine is presented over the usable physical fabric.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-defects\" id=\"user-content-fnref-defects-2\">7</a></sup></li>\n</ul>\n<p>Analysis:</p>\n<ul>\n<li>Fine-grained redundancy limits the capacity lost to a local defect.</li>\n<li>Fault maps become an input to placement and routing.</li>\n<li>Software abstraction is essential because applications cannot target one fixed physical core grid across manufactured wafers.</li>\n</ul>\n<p>Do not read the active-core ratio as a universal manufacturing yield. Yield also includes packaging, power delivery, cooling, and system qualification.</p>\n<h2 id=\"weight-streaming-for-training\">Weight streaming for training</h2>\n<p>On-wafer SRAM cannot hold every parameter and optimizer state of a large model. Cerebras Weight Streaming separates parameter storage from compute.</p>\n<p>MemoryX stores weights and optimizer state. It streams layer weights to the Wafer-Scale Engine. Activations remain on the engine. During backpropagation, gradients return through SwarmX, which reduces gradients from multiple systems before MemoryX applies updates.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-weight-streaming\" id=\"user-content-fnref-weight-streaming\">8</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>MemoryX</span></span>\n<span class=\"line\"><span>  │ broadcast layer weights</span></span>\n<span class=\"line\"><span>  ▼</span></span>\n<span class=\"line\"><span>SwarmX tree</span></span>\n<span class=\"line\"><span>  ├── WSE: local activations → gradients</span></span>\n<span class=\"line\"><span>  ├── WSE: local activations → gradients</span></span>\n<span class=\"line\"><span>  └── WSE: local activations → gradients</span></span>\n<span class=\"line\"><span>  ▲</span></span>\n<span class=\"line\"><span>  │ reduce gradients</span></span>\n<span class=\"line\"><span>MemoryX: optimizer update</span></span></code></pre>\n<p>The diagram is structural. It does not specify a cluster size or performance result.</p>\n<p>This execution model uses data parallelism across Wafer-Scale Engines while SwarmX handles weight broadcast and gradient reduction. Each engine runs the same model mapping on a different data shard.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-weight-streaming\" id=\"user-content-fnref-weight-streaming-2\">8</a></sup></p>\n<h3 id=\"inference-placement-differs\">Inference placement differs</h3>\n<p>Cerebras inference material describes storing model weights in on-wafer SRAM and splitting models larger than one wafer at layer boundaries across systems.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-wse3-inference\" id=\"user-content-fnref-wse3-inference-2\">5</a></sup> That differs from training Weight Streaming, where MemoryX supplies weights layer by layer and performs optimizer updates.</p>\n<p>The term weight streaming therefore needs context:</p>\n<ul>\n<li>Training: stream layer weights from MemoryX and return gradients.</li>\n<li>Inference service: place weights on one or more wafer systems and pipeline layer boundaries as required by model size.</li>\n</ul>\n<h2 id=\"comparing-the-compiler-contracts\">Comparing the compiler contracts</h2>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Property</th><th>Groq LPU</th><th>Cerebras WSE</th></tr></thead><tbody><tr><td>Primary abstraction</td><td>Statically timed tensor streams</td><td>Spatial graph mapping across a wafer</td></tr><tr><td>Local memory</td><td>Compiler-managed distributed SRAM</td><td>Core-local distributed SRAM</td></tr><tr><td>Execution unit</td><td>One logical scheduled core across function units</td><td>Grid of independently programmable AI cores</td></tr><tr><td>Communication</td><td>Explicitly scheduled on-chip and chip-to-chip streams</td><td>Routes over an on-wafer mesh</td></tr><tr><td>Scale beyond device</td><td>Scheduled processor network</td><td>System pipeline or Weight Streaming cluster</td></tr><tr><td>Large training state</td><td>Partitioning depends on compiled system</td><td>MemoryX disaggregates weights and optimizer state</td></tr></tbody></table></div>\n<p>The table describes architecture, not relative performance.</p>\n<h2 id=\"workload-shape-and-compilation\">Workload shape and compilation</h2>\n<p>Both systems reward stable tensor shapes. Placement, memory allocation, and communication scheduling improve when dimensions are known.</p>\n<p>Variable request lengths create a serving problem:</p>\n<ul>\n<li>Padding wastes compute and memory movement.</li>\n<li>Bucketing increases the number of compiled shapes.</li>\n<li>Dynamic batching changes arrival and queueing behavior.</li>\n<li>Long contexts expand attention state even when weights remain fixed.</li>\n</ul>\n<p>The serving layer must map dynamic requests to compiled execution plans. Architecture-level determinism does not remove batching policy.</p>\n<h2 id=\"memory-capacity-versus-bandwidth\">Memory capacity versus bandwidth</h2>\n<p>High on-chip bandwidth does not imply that every model fits in local SRAM.</p>\n<p>For uncompressed weights:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mtext>weight bytes</mtext><mo>=</mo><mtext>parameter count</mtext><mo>×</mo><mtext>bytes per parameter</mtext></mrow></semantics></math>\n<p>A model using a 16-bit representation needs two bytes per parameter. Cerebras uses this arithmetic in its inference architecture explanation and reports 44 GB of WSE-3 SRAM.<sup><a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fn-wse3-inference\" id=\"user-content-fnref-wse3-inference-3\">5</a></sup> Runtime memory also includes activations, attention state, buffers, and code.</p>\n<p>Groq distributes weights across processor SRAM when the model exceeds one device. Cerebras can divide an inference model at layer boundaries across systems. Both approaches turn capacity into a placement problem.</p>\n<h2 id=\"operational-failure-modes\">Operational failure modes</h2>\n<h3 id=\"compilation-misses-the-deployed-shape\">Compilation misses the deployed shape</h3>\n<p>A request outside compiled bounds needs rejection, a fallback plan, or a new compilation. Silent padding can create latency and capacity regressions.</p>\n<h3 id=\"pipeline-imbalance\">Pipeline imbalance</h3>\n<p>If one stage takes longer, faster stages wait and buffers fill. Balance must include computation, memory, and communication.</p>\n<h3 id=\"host-path-dominates\">Host path dominates</h3>\n<p>Tokenization, request scheduling, output streaming, and admission control remain outside accelerator cores. Low accelerator latency can expose these costs.</p>\n<h3 id=\"context-state-exceeds-planning-assumptions\">Context state exceeds planning assumptions</h3>\n<p>Attention state grows with active sequence length and concurrency. Weight locality does not remove key-value state capacity.</p>\n<h3 id=\"scale-out-faults-invalidate-timing-or-placement\">Scale-out faults invalidate timing or placement</h3>\n<p>Groq’s scheduled network depends on known links and timing. Cerebras placement depends on available wafer and system resources. Fault handling needs a supported remap, recompile, or job restart path.</p>\n<h3 id=\"vendor-benchmark-scope-is-unclear\">Vendor benchmark scope is unclear</h3>\n<p>Tokens per second can mean per user, per stream, or aggregate server throughput. Batch, prompt length, output length, precision, and model revision must accompany any comparison.</p>\n<h2 id=\"benchmarking-the-architectures\">Benchmarking the architectures</h2>\n<p>Use the same model weights, tokenizer, precision policy, prompt distribution, output lengths, and correctness checks.</p>\n<p>Measure:</p>\n<ul>\n<li>Time to first token.</li>\n<li>Inter-token latency distribution.</li>\n<li>Per-request and aggregate throughput.</li>\n<li>Queueing and admission delay.</li>\n<li>Concurrency at the reported latency.</li>\n<li>Compile time and cache hit rate.</li>\n<li>Energy at the system boundary.</li>\n<li>Accuracy against the reference implementation.</li>\n<li>Failure and fallback rate for unsupported shapes.</li>\n</ul>\n<p>Separate architecture measurements from managed-service measurements. A cloud endpoint includes routing, capacity, software releases, and tenant load that are not chip properties.</p>\n<h2 id=\"architectural-selection\">Architectural selection</h2>\n<p>Groq’s design fits workloads where a compiler can exploit a stable graph and where deterministic tensor and network schedules support latency-oriented execution. The cost is a stronger compiler contract and less reliance on dynamic hardware scheduling.</p>\n<p>Cerebras fits workloads that benefit from mapping large tensor graphs across a dense core and memory fabric. Wafer-scale integration reduces off-device communication inside the mapped graph, while defect tolerance and compilation make the physical wafer usable as one machine.</p>\n<p>Distributed-system concerns remain. Groq moves coordination into a static program across processors. Cerebras moves more of the graph inside one wafer and uses explicit system-level pipelines or Weight Streaming beyond it.</p>\n<p>Trace a token through each machine. On Groq, the compiled schedule coordinates SRAM, functional units, and links. On Cerebras, placement across the wafer determines local memory use and fabric traffic. Fault recovery belongs in the comparison because a remap, recompile, or restart changes system behavior.</p>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-groq-isca\">\n<p><a href=\"https://doi.org/10.1145/3470496.3527405\" rel=\"noopener noreferrer\">Abts et al., “A Software-defined Tensor Streaming Multiprocessor for Large-scale Machine Learning,” ISCA, 2022</a>. <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-groq-isca\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-groq-isca-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-groq-isca-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-groq-isca-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a></p>\n</li>\n<li id=\"user-content-fn-groq-determinism\">\n<p><a href=\"https://groq.com/GroqDocs/TechDoc_Predictability.pdf\" rel=\"noopener noreferrer\">Groq, “Determinism and the Tensor Streaming Processor”</a>. <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-groq-determinism\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-groq-lpu\">\n<p><a href=\"https://groq.com/blog/inside-the-lpu-deconstructing-groq-speed\" rel=\"noopener noreferrer\">Groq, “Inside the LPU: Deconstructing Groq’s Speed,” August 2025</a>. <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-groq-lpu\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-groq-assembly\">\n<p><a href=\"https://groq.com/blog/the-groq-lpu-explained\" rel=\"noopener noreferrer\">Groq, “What is a Language Processing Unit?”</a>. <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-groq-assembly\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-wse3-inference\">\n<p><a href=\"https://www.cerebras.ai/blog/introducing-cerebras-inference-ai-at-instant-speed\" rel=\"noopener noreferrer\">Cerebras, “Introducing Cerebras Inference,” August 2024</a>. <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-wse3-inference\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-wse3-inference-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-wse3-inference-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-cerebras-deep\">\n<p><a href=\"https://www.cerebras.ai/blog/cerebras-architecture-deep-dive-first-look-inside-the-hw-sw-co-design-for-deep-learning\" rel=\"noopener noreferrer\">Cerebras, “Architecture Deep Dive: HW/SW Co-Design for Deep Learning”</a>. <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-cerebras-deep\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-defects\">\n<p><a href=\"https://www.cerebras.ai/blog/100x-defect-tolerance-how-cerebras-solved-the-yield-problem\" rel=\"noopener noreferrer\">Cerebras, “100x Defect Tolerance: How Cerebras Solved the Yield Problem,” January 2025</a>. <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-defects\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-defects-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-weight-streaming\">\n<p><a href=\"https://www.cerebras.ai/blog/linear-scaling-made-possible-with-weight-streaming\" rel=\"noopener noreferrer\">Cerebras, “Linear Scaling Made Possible with Weight Streaming,” September 2022</a>. <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-weight-streaming\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/groq-cerebras-accelerators/#user-content-fnref-weight-streaming-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/groq-cerebras-accelerators/",
            "title": "Groq and Cerebras Accelerators: Deterministic Streaming and Wafer-Scale Compute",
            "summary": "Compare Groq compiler-scheduled tensor streaming with Cerebras wafer-scale compute, defect tolerance, and weight streaming.",
            "image": "https://blog.ecitis.org/open-graph/groq-cerebras-accelerators.png",
            "date_modified": "2026-07-17T00:00:00.000Z",
            "date_published": "2026-07-17T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Hardware & Systems",
                "Groq",
                "Cerebras",
                "AI accelerators"
            ]
        },
        {
            "id": "https://blog.ecitis.org/test-time-compute/",
            "content_html": "<p>Pretraining spends compute once to change model weights. Test-time compute spends compute per problem to search for a better answer without changing the base checkpoint.</p>\n<p>That extra compute can take several forms: a longer internal trace, parallel samples, iterative revisions, tree search, tool calls, or verifier passes. Counting generated tokens alone combines mechanisms with different costs and failure modes.</p>\n<p>A deployment policy must answer:</p>\n<blockquote>\n<p>Given a fixed inference budget, which candidate should receive the next unit of compute?</p>\n</blockquote>\n<figure><figcaption><strong>A verifier should spend the next unit of compute on the best partial path</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/test-time-compute-0.webp\" width=\"512\" height=\"636\" alt=\"A verifier should spend the next unit of compute on the best partial path\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"parallel-sampling-estimates-answer-mass\">Parallel sampling estimates answer mass</h2>\n<p>Self-consistency samples several reasoning paths and selects the answer supported by the largest share of paths.<sup><a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup> It replaces a single greedy trajectory with an empirical estimate:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mover accent=\"true\"><mi>p</mi><mo>^</mo></mover><mo stretchy=\"false\">(</mo><mi>a</mi><mo>∣</mo><mi>x</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mfrac><mn>1</mn><mi>K</mi></mfrac><munderover><mo>∑</mo><mrow><mi>k</mi><mo>=</mo><mn>1</mn></mrow><mi>K</mi></munderover><mn mathvariant=\"bold\">1</mn><mo stretchy=\"false\">[</mo><mi>f</mi><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>k</mi></msub><mo stretchy=\"false\">)</mo><mo>=</mo><mi>a</mi><mo stretchy=\"false\">]</mo></mrow></semantics></math>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>y</mi><mi>k</mi></msub></mrow></semantics></math> is a sampled reasoning path and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>f</mi><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>k</mi></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math> extracts its final answer. The method works when diverse valid paths converge on the same answer more often than invalid paths do.</p>\n<p>It can fail when:</p>\n<ul>\n<li>the model repeats the same misconception across samples</li>\n<li>equivalent answers are normalized inconsistently</li>\n<li>an open-ended task has no canonical answer to vote on</li>\n<li>sampling temperature creates superficial variety without different reasoning</li>\n<li>the correct path is rare but recognizable by a stronger verifier</li>\n</ul>\n<p>Majority vote is a verifier built from agreement. It does not inspect whether any path is logically valid.</p>\n<h2 id=\"verifiers-turn-samples-into-search\">Verifiers turn samples into search</h2>\n<p>A verifier assigns a score to a complete answer or an intermediate state. Outcome reward models score the result. Process reward models score steps along the path.</p>\n<p>The PRM800K work compared outcome and process supervision on mathematical problem solving and found stronger solution selection from process supervision in its evaluated setting.<sup><a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> A dense process score allows search to stop extending a branch after an early error instead of paying to finish it.</p>\n<p>For a partial trajectory <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>s</mi><mi>t</mi></msub></mrow></semantics></math>, a search policy can rank expansions with:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"normal\">priority</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo><mo>=</mo><mover accent=\"true\"><mi>V</mi><mo>^</mo></mover><mo stretchy=\"false\">(</mo><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo><mo>−</mo><mi>λ</mi><mtext> </mtext><mi mathvariant=\"normal\">cost</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo><mo>+</mo><mi>η</mi><mtext> </mtext><mi mathvariant=\"normal\">uncertainty</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mover accent=\"true\"><mi>V</mi><mo>^</mo></mover></mrow></semantics></math> is the verifier’s predicted value. The cost term penalizes expensive branches. The uncertainty term can preserve exploration when the verifier is unsure. The coefficients are deployment choices, not universal constants.</p>\n<p>Verifier quality sets the ceiling. If the verifier rewards polished invalid reasoning, more search directs compute toward the wrong region. The generator and verifier can also share correlated errors when they come from similar training data.</p>\n<h2 id=\"difficulty-should-control-the-budget\">Difficulty should control the budget</h2>\n<p>Snell and collaborators compared test-time strategies under fixed compute and found that the best allocation depended on problem difficulty and base-model capability.<sup><a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup> Their compute-optimal policy used prompt difficulty to choose how much inference compute to spend, improving efficiency over a fixed best-of-<math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>N</mi></mrow></semantics></math> baseline in the reported experiments.</p>\n<p>The result rejects a uniform long reasoning budget for every request.</p>\n<p>A practical router can use:</p>\n<ul>\n<li>confidence or entropy from an initial pass</li>\n<li>disagreement among a small pilot sample</li>\n<li>verifier margin between leading candidates</li>\n<li>task class and known evaluation history</li>\n<li>elapsed latency and remaining service budget</li>\n<li>whether an external tool can check the result</li>\n</ul>\n<p>Easy requests should exit early. Difficult but tractable requests may benefit from branching. Requests outside the verifier’s domain should not receive more compute merely because confidence is low.</p>\n<figure><figcaption><strong>Inference budgets should follow evidence about the request</strong> — Low confidence alone does not show that more generation will help.</figcaption><img src=\"https://blog.ecitis.org/feeds/figures/test-time-compute-1.webp\" width=\"512\" height=\"711\" alt=\"Inference budgets should follow evidence about the request — Low confidence alone does not show that more generation will help.\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"sequential-reasoning-and-parallel-search-are-different\">Sequential reasoning and parallel search are different</h2>\n<p>A long chain of thought performs serial computation. Best-of-<math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>N</mi></mrow></semantics></math> performs parallel exploration. Tree search mixes both.</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Strategy</th><th>Main benefit</th><th>Main bottleneck</th></tr></thead><tbody><tr><td>Longer trace</td><td>More serial intermediate computation</td><td>Drift and overthinking</td></tr><tr><td>Parallel samples</td><td>Diverse candidate paths</td><td>Repeated shared errors</td></tr><tr><td>Revision</td><td>Error correction after critique</td><td>Self-critique may preserve the premise</td></tr><tr><td>Tree search</td><td>Reallocate compute among partial paths</td><td>Verifier calibration</td></tr><tr><td>Tool use</td><td>External state or exact checks</td><td>Tool latency and integration errors</td></tr></tbody></table></div>\n<p>The strategies are not interchangeable at equal token counts. Parallel sampling can batch well on accelerators. A serial trace adds latency token by token. Tool calls may dominate both.</p>\n<p>Measure wall-clock latency, accelerator occupancy, and verifier cost alongside generated tokens.</p>\n<h2 id=\"more-reasoning-can-reverse-a-correct-answer\">More reasoning can reverse a correct answer</h2>\n<p>Inference scaling has diminishing returns. A model can find a correct answer, continue exploring, and replace it with a wrong one.</p>\n<p>Recent work on overthinking separates extra length from useful computation. A Google DeepMind study of open reasoning models identified over-verification and over-exploration as recurring structures in long traces on simple problems.<sup><a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup> Another 2026 preprint reports that optimal reasoning length varies with problem difficulty and that added reasoning can abandon an earlier correct answer.<sup><a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup></p>\n<p>This makes stopping a first-class component:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>generate a small candidate set</span></span>\n<span class=\"line\"><span>score candidates and estimate uncertainty</span></span>\n<span class=\"line\"><span>if one candidate is verified with adequate margin:</span></span>\n<span class=\"line\"><span>    stop</span></span>\n<span class=\"line\"><span>if promising branches remain:</span></span>\n<span class=\"line\"><span>    allocate another search round</span></span>\n<span class=\"line\"><span>otherwise:</span></span>\n<span class=\"line\"><span>    abstain, escalate, or use an external checker</span></span></code></pre>\n<p>The stop condition should use calibrated evidence. A high self-reported confidence is weak if the model is miscalibrated on that task class.</p>\n<h2 id=\"test-time-scaling-can-expose-reward-hacking\">Test-time scaling can expose reward hacking</h2>\n<p>Search optimizes the verifier more aggressively than one-shot decoding. That makes test-time compute an online form of proxy optimization.</p>\n<p>Typical failures include:</p>\n<p><strong>Verifier exploitation:</strong> candidates learn patterns that score well without solving the task.</p>\n<p><strong>Loss of diversity:</strong> early verifier preferences collapse the search around one flawed approach.</p>\n<p><strong>Answer leakage:</strong> a benchmark-specific checker exposes information unavailable in deployment.</p>\n<p><strong>Cost blindness:</strong> marginal accuracy gains consume a service budget that would help more on another request.</p>\n<p><strong>Unfaithful traces:</strong> the visible rationale satisfies a process rubric while the answer is driven by another internal computation.</p>\n<p>Evaluate the complete generator-search-verifier system on adversarial cases. Evaluating the base model alone misses the mechanism used at inference.</p>\n<h2 id=\"a-deployment-metric-stack\">A deployment metric stack</h2>\n<p>Accuracy is necessary but incomplete. A test-time compute service should report:</p>\n<ul>\n<li>quality as a function of total inference cost</li>\n<li>latency distributions by task class</li>\n<li>quality at matched wall-clock latency</li>\n<li>verifier calibration and selection regret</li>\n<li>pass rate of external checks</li>\n<li>candidate diversity before and after pruning</li>\n<li>fraction of requests stopped at each budget tier</li>\n<li>cases where more compute changed a correct answer to an incorrect one</li>\n</ul>\n<p>Plot quality against compute across the relevant request distribution. One policy dominates another when it reaches equal or better quality at the same deployment cost.</p>\n<p>Test-time compute is an inference algorithm that decides what to generate, how to judge it, where to branch, and when to stop. Extra compute pays off when those decisions are more reliable than the candidate errors they are meant to correct.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Xuezhi Wang et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models,” ICLR, 2023. <a href=\"https://arxiv.org/abs/2203.11171\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2203.11171</a> <a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Hunter Lightman et al., “Let’s Verify Step by Step,” ICLR, 2024. <a href=\"https://arxiv.org/abs/2305.20050\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2305.20050</a> <a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Charlie Snell et al., “Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters,” ICLR, 2025. <a href=\"https://arxiv.org/abs/2408.03314\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2408.03314</a> <a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Xinliang Frederick Zhang et al., “Towards Structural Understanding of LLM Overthinking,” ACL, 2026. <a href=\"https://deepmind.google/research/publications/203490/\" rel=\"noopener noreferrer\">https://deepmind.google/research/publications/203490/</a> <a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>Shu Zhou et al., “When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling,” arXiv, 2026. <a href=\"https://arxiv.org/abs/2604.10739\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2604.10739</a> <a href=\"https://blog.ecitis.org/test-time-compute/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/test-time-compute/",
            "title": "Test-Time Compute: Search, Verification, and Stopping",
            "summary": "Understand inference-time scaling as a budget allocation problem across candidate generation, verification, and adaptive stopping.",
            "image": "https://blog.ecitis.org/open-graph/test-time-compute.png",
            "date_modified": "2026-07-15T00:00:00.000Z",
            "date_published": "2026-07-15T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "Test-time compute",
                "Reasoning",
                "Verifiers",
                "Inference"
            ]
        },
        {
            "id": "https://blog.ecitis.org/multi-token-prediction/",
            "content_html": "<p>Next-token language modeling supervises one future token at each position. Multi-token prediction attaches several output heads to a shared representation and trains each head to predict a different future offset.</p>\n<p>The training objective provides denser supervision. At inference, those heads can propose a short continuation for verification by the main model. The mechanism connects model training and speculative execution without requiring a separate draft model.</p>\n<h2 id=\"multi-offset-objective\">Multi-offset objective</h2>\n<p>Given tokens <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>x</mi><mn>1</mn></msub><mo separator=\"true\">,</mo><mo>…</mo><mo separator=\"true\">,</mo><msub><mi>x</mi><mi>T</mi></msub></mrow></semantics></math>, a conventional causal model minimizes:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mn>1</mn></msub><mo>=</mo><mo>−</mo><munder><mo>∑</mo><mi>t</mi></munder><mi>log</mi><mo>⁡</mo><mi>p</mi><mo stretchy=\"false\">(</mo><msub><mi>x</mi><mrow><mi>t</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>∣</mo><msub><mi>x</mi><mrow><mo>≤</mo><mi>t</mi></mrow></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>With <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>n</mi></mrow></semantics></math> future heads, the objective becomes:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>M</mi><mi>T</mi><mi>P</mi></mrow></msub><mo>=</mo><mo>−</mo><munder><mo>∑</mo><mi>t</mi></munder><munderover><mo>∑</mo><mrow><mi>k</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover><msub><mi>α</mi><mi>k</mi></msub><mi>log</mi><mo>⁡</mo><msub><mi>p</mi><mi>k</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>x</mi><mrow><mi>t</mi><mo>+</mo><mi>k</mi></mrow></msub><mo>∣</mo><msub><mi>x</mi><mrow><mo>≤</mo><mi>t</mi></mrow></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>Each head sees the same causal prefix but has a different target offset. Gloeckle and collaborators use independent output heads built from a shared model trunk and report that the added training-time overhead can be removed at deployment when only the next-token head is retained.<sup><a href=\"https://blog.ecitis.org/multi-token-prediction/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup></p>\n<figure><figcaption><strong>One hidden state supervises several future-token heads</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/multi-token-prediction-0.webp\" width=\"512\" height=\"333\" alt=\"One hidden state supervises several future-token heads\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"architecture-choices\">Architecture choices</h2>\n<p>A minimal design projects the final hidden state through independent linear heads. A richer head can include a small transformer block that combines the shared state with an embedding of a previously predicted future token.</p>\n<p>The choices affect:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Design</th><th>Training cost</th><th>Proposal dependence</th><th>Inference role</th></tr></thead><tbody><tr><td>Independent linear heads</td><td>Low</td><td>Heads do not condition on sibling proposals</td><td>Parallel candidate tuple</td></tr><tr><td>Sequential lightweight heads</td><td>Higher</td><td>Later proposals condition on earlier ones</td><td>Autoregressive draft</td></tr><tr><td>Tree heads</td><td>Higher</td><td>Multiple alternatives per depth</td><td>Branch verification</td></tr><tr><td>Auxiliary heads removed</td><td>Training only</td><td>None at serving</td><td>Better trunk supervision</td></tr></tbody></table></div>\n<p>Independent heads can disagree structurally because the offset head predicts a marginal future token without knowing which earlier proposal was selected. A verification tree can retain several combinations, but candidate count grows with branching.</p>\n<h2 id=\"training-alignment-and-shifting\">Training alignment and shifting</h2>\n<p>Target construction is easy to get wrong. For head <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math>, logit position <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>t</mi></mrow></semantics></math> must align with token <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>x</mi><mrow><mi>t</mi><mo>+</mo><mi>k</mi></mrow></msub></mrow></semantics></math>. Tail positions without that future target must be masked from the loss.</p>\n<p>Near-runnable PyTorch code:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>functional </span><span>as</span><span> F</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> multi_token_loss</span><span>(</span><span>head_logits</span><span>,</span><span> token_ids</span><span>,</span><span> weights</span><span>):</span></span>\n<span class=\"line\"><span>    # head_logits[k]: [batch, sequence, vocabulary]</span></span>\n<span class=\"line\"><span>    total </span><span>=</span><span> token_ids</span><span>.</span><span>new_zeros</span><span>((), dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    for</span><span> offset</span><span>,</span><span> (logits</span><span>,</span><span> weight) </span><span>in</span><span> enumerate</span><span>(</span></span>\n<span class=\"line\"><span>        zip</span><span>(head_logits, weights), start</span><span>=</span><span>1</span></span>\n<span class=\"line\"><span>    ):</span></span>\n<span class=\"line\"><span>        usable </span><span>=</span><span> token_ids</span><span>.</span><span>size</span><span>(</span><span>1</span><span>)</span><span> -</span><span> offset</span></span>\n<span class=\"line\"><span>        if</span><span> usable </span><span>&lt;=</span><span> 0</span><span>:</span></span>\n<span class=\"line\"><span>            continue</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        predictions </span><span>=</span><span> logits</span><span>[:,</span><span> :</span><span>usable</span><span>].</span><span>reshape</span><span>(</span><span>-</span><span>1</span><span>, logits.</span><span>size</span><span>(</span><span>-</span><span>1</span><span>))</span></span>\n<span class=\"line\"><span>        targets </span><span>=</span><span> token_ids</span><span>[:,</span><span> offset</span><span>:].</span><span>reshape</span><span>(</span><span>-</span><span>1</span><span>)</span></span>\n<span class=\"line\"><span>        total </span><span>=</span><span> total </span><span>+</span><span> weight </span><span>*</span><span> F</span><span>.</span><span>cross_entropy</span><span>(predictions, targets)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> total</span></span></code></pre>\n<p>Production code also masks padding, packed-document boundaries, and loss-excluded tokens. A future head must not cross from one packed document into another. Build an offset-specific validity mask from segment identifiers.</p>\n<h2 id=\"shared-trunk-gradients\">Shared trunk gradients</h2>\n<p>Every head contributes gradients to the trunk. Large auxiliary weights can optimize distant predictions at the expense of the primary next-token distribution. Log loss and gradient norms by head.</p>\n<p>Useful controls include:</p>\n<ul>\n<li>offset-dependent loss weights</li>\n<li>a warm-up schedule for auxiliary heads</li>\n<li>per-head normalization before summing</li>\n<li>gradient clipping measured before and after aggregation</li>\n<li>validation perplexity for the primary head</li>\n</ul>\n<p>The multi-token prediction paper reports stronger gains for larger models and on code tasks, while finding no advantage for small models in some evaluated settings.<sup><a href=\"https://blog.ecitis.org/multi-token-prediction/#user-content-fn-1\" id=\"user-content-fnref-1-2\">1</a></sup> Treat those results as evidence for the paper’s configurations, not a guarantee that more heads improve every model.</p>\n<h2 id=\"from-heads-to-draft-tokens\">From heads to draft tokens</h2>\n<p>At inference, auxiliary heads propose tokens for offsets ahead of the current prefix. The target path verifies candidates with causal attention. A simple linear chain uses one top candidate from each head. A tree uses several candidates at each depth.</p>\n<p>Medusa trains multiple decoding heads and verifies a tree of candidate continuations in parallel.<sup><a href=\"https://blog.ecitis.org/multi-token-prediction/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> Its tree attention mask permits a node to attend to its ancestors but not to nodes on other branches.</p>\n<p>For a chain:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>shared hidden state</span></span>\n<span class=\"line\"><span>  -&gt; head for offset one proposes a</span></span>\n<span class=\"line\"><span>  -&gt; head for offset two proposes b</span></span>\n<span class=\"line\"><span>  -&gt; head for offset three proposes c</span></span>\n<span class=\"line\"><span>  -&gt; target verifies prefix + [a, b, c]</span></span></code></pre>\n<p>If the target rejects <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>b</mi></mrow></semantics></math>, candidate <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>c</mi></mrow></semantics></math> is invalid because its prefix is no longer the proposed chain. The runtime commits the accepted prefix and a target correction, then drafts again.</p>\n<h2 id=\"draft-heads-and-lossless-sampling\">Draft heads and lossless sampling</h2>\n<p>Heads produce a proposal distribution <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>q</mi></mrow></semantics></math>; the target produces <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>p</mi></mrow></semantics></math>. Lossless speculative sampling accepts a proposed token with probability:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>min</mi><mo>⁡</mo><mrow><mo fence=\"true\">(</mo><mn>1</mn><mo separator=\"true\">,</mo><mfrac><mrow><mi>p</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><mi>q</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo fence=\"true\">)</mo></mrow></mrow></semantics></math>\n<p>On rejection, it samples from the normalized positive residual <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>p</mi><mo>−</mo><mi>q</mi></mrow></semantics></math>. This is the same correction used in speculative sampling with a separate draft model.<sup><a href=\"https://blog.ecitis.org/multi-token-prediction/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup> Greedy verification can use token equality but preserves only greedy behavior.</p>\n<p>Temperature, top-<math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>p</mi></mrow></semantics></math>, top-<math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math>, repetition penalties, and grammar masks must be reflected consistently in proposal and target probabilities. Independent heads that expose only selected token identifiers are insufficient for exact stochastic correction; the verifier needs proposal probabilities.</p>\n<h2 id=\"inference-memory-and-kernels\">Inference memory and kernels</h2>\n<p>Auxiliary heads add parameters and logits. Materializing full vocabulary logits for every head can consume substantial memory bandwidth. Fused top-<math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math> selection can avoid writing complete logits to global memory when the verification method needs only candidate tokens and probabilities.</p>\n<p>Serving options:</p>\n<ul>\n<li>retain heads on the same device as the trunk</li>\n<li>place head weights with the final pipeline stage</li>\n<li>quantize auxiliary head weights separately after acceptance validation</li>\n<li>disable heads for batches whose measured acceptance is low</li>\n<li>select a smaller active head count by request class</li>\n</ul>\n<p>The shared trunk avoids a second model’s KV cache, but verification candidates still require provisional target-cache positions. Roll back unaccepted suffix blocks without copying the accepted prefix.</p>\n<h2 id=\"candidate-tree-budgeting\">Candidate-tree budgeting</h2>\n<p>Tree width should be based on verification cost and acceptance, not a fixed aesthetic shape. Let node <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>v</mi></mrow></semantics></math> have estimated probability of being reached <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>r</mi><mi>v</mi></msub></mrow></semantics></math> and verification cost contribution <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>c</mi><mi>v</mi></msub></mrow></semantics></math>. Candidate selection can prioritize high <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>r</mi><mi>v</mi></msub><mi mathvariant=\"normal\">/</mi><msub><mi>c</mi><mi>v</mi></msub></mrow></semantics></math> while respecting kernel shape and maximum nodes.</p>\n<p>Calibrate reach probability from held-out traffic. Do not interpret head softmax values as calibrated branch probabilities without checking them. Code, prose, and structured output can have different predictability.</p>\n<h2 id=\"failure-modes\">Failure modes</h2>\n<p><strong>Cross-document targets:</strong> packed training shifts a target into the next document. Mask by segment identifier at every offset.</p>\n<p><strong>Primary-head regression:</strong> auxiliary gradients reduce next-token quality. Track the primary objective and downstream evaluations independently.</p>\n<p><strong>Head collapse:</strong> distant heads learn frequent-token priors rather than prefix-specific futures. Measure conditional accuracy by token class and entropy.</p>\n<p><strong>Verification-mask leak:</strong> a candidate reads a sibling branch or later token. Compare tree verification logits against independent autoregressive reference runs.</p>\n<p><strong>Memory-bandwidth regression:</strong> writing head logits costs more than the saved target steps. Profile end-to-end time per committed token.</p>\n<p><strong>Static head count:</strong> easy prompts waste available candidates or difficult prompts waste verification slots. Route using measured acceptance.</p>\n<p><strong>Tokenizer or processor mismatch:</strong> head proposals bypass grammar masks or sampling transforms. Centralize logits processing and test equivalence.</p>\n<h2 id=\"evaluation\">Evaluation</h2>\n<p>Training evaluation needs primary-head perplexity, per-offset loss, downstream task quality, and gradient statistics. Serving evaluation needs:</p>\n<ul>\n<li>committed tokens per target pass</li>\n<li>acceptance by offset and request type</li>\n<li>draft-head kernel time</li>\n<li>target verification time by candidate count</li>\n<li>provisional cache blocks allocated and released</li>\n<li>output-distribution equivalence against ordinary sampling</li>\n<li>latency percentiles under continuous batching</li>\n</ul>\n<p>Compare three deployments from the same checkpoint: primary head only, auxiliary heads enabled without speculative verification, and verified multi-token decoding. This separates representation gains from serving gains.</p>\n<p>Multi-token prediction is useful in two independent ways. It changes the training signal by requiring a prefix representation to predict several future offsets. It also supplies model-native proposals for speculative decoding. Each claim needs its own baseline, because better training loss does not prove faster serving and high acceptance does not prove unchanged sampling.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Fabian Gloeckle et al., “Better &amp; Faster Large Language Models via Multi-token Prediction,” ICML, 2024. <a href=\"https://arxiv.org/abs/2404.19737\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2404.19737</a> <a href=\"https://blog.ecitis.org/multi-token-prediction/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/multi-token-prediction/#user-content-fnref-1-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Tianle Cai et al., “Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads,” ICML, 2024. <a href=\"https://arxiv.org/abs/2401.10774\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2401.10774</a> <a href=\"https://blog.ecitis.org/multi-token-prediction/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Charlie Chen et al., “Accelerating Large Language Model Decoding with Speculative Sampling,” arXiv, 2023. <a href=\"https://arxiv.org/abs/2302.01318\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2302.01318</a> <a href=\"https://blog.ecitis.org/multi-token-prediction/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/multi-token-prediction/",
            "title": "Multi-Token Prediction: Training Draft Heads for Faster Decoding",
            "summary": "Train offset-specific draft heads, verify their candidates, and measure when denser supervision produces faster autoregressive decoding.",
            "image": "https://blog.ecitis.org/open-graph/multi-token-prediction.png",
            "date_modified": "2026-07-12T00:00:00.000Z",
            "date_published": "2026-07-12T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Multi-token prediction",
                "Speculative decoding",
                "Training"
            ]
        },
        {
            "id": "https://blog.ecitis.org/scaling-laws-regimes/",
            "content_html": "<p>Scaling laws fit a small equation to measurements from many training runs. The fitted curve can predict loss at an untested scale and guide the allocation of a fixed training budget. Its domain is limited by the model family, data distribution, optimizer, tokenization scheme, and measurement range used to produce it.</p>\n<p>Read the equation with its conditions attached:</p>\n<blockquote>\n<p>Within a measured regime, a chosen metric often changes smoothly enough with scale to support extrapolation.</p>\n</blockquote>\n<figure><figcaption><strong>Scaling laws change when the limiting resource changes</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/scaling-laws-regimes-0.webp\" width=\"512\" height=\"686\" alt=\"Scaling laws change when the limiting resource changes\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"the-basic-power-law-model\">The basic power-law model</h2>\n<p>Kaplan and collaborators measured cross-entropy loss while varying parameter count, dataset size, and training compute. They found smooth power-law relationships across their experimental range.<sup><a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup></p>\n<p>A simplified model-size fit is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>L</mi><mo stretchy=\"false\">(</mo><mi>N</mi><mo stretchy=\"false\">)</mo><mo>=</mo><msub><mi>L</mi><mi mathvariant=\"normal\">∞</mi></msub><mo>+</mo><mfrac><mi>A</mi><msup><mi>N</mi><mi>α</mi></msup></mfrac></mrow></semantics></math>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>N</mi></mrow></semantics></math> is parameter count, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>L</mi><mi mathvariant=\"normal\">∞</mi></msub></mrow></semantics></math> is an irreducible-loss term for the data distribution, and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>A</mi></mrow></semantics></math> and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>α</mi></mrow></semantics></math> are fitted constants. Similar forms can be written for data size <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>D</mi></mrow></semantics></math> or training compute <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>C</mi></mrow></semantics></math>.</p>\n<p>The equation describes diminishing returns. Each multiplicative increase in the resource reduces the reducible part of the loss by a fixed factor. The exponent controls the slope on a log-log plot.</p>\n<p>Three details matter:</p>\n<ul>\n<li>The fitted constants belong to the experimental setup. They are not architectural constants.</li>\n<li>Cross-entropy loss can improve smoothly while a downstream capability changes unevenly.</li>\n<li>Extrapolation uncertainty grows outside the observed scale and data distribution.</li>\n</ul>\n<p>A curve is useful only with its fitting range, residuals, and held-out prediction error.</p>\n<h2 id=\"compute-allocation-changed-after-chinchilla\">Compute allocation changed after Chinchilla</h2>\n<p>Training compute can be spent on more parameters, more tokens, or both. A model with many parameters and too few training tokens may sit above the best achievable loss for its compute budget.</p>\n<p>The Chinchilla study trained a large sweep of models and found a different compute-optimal allocation from the earlier Kaplan prescription. Its fitted result called for scaling model size and training tokens together under the studied setup.<sup><a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> The practical shift was from “make the model larger and stop early” toward “train a smaller model on more data” for a fixed budget.</p>\n<p>A common joint form is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>L</mi><mo stretchy=\"false\">(</mo><mi>N</mi><mo separator=\"true\">,</mo><mi>D</mi><mo stretchy=\"false\">)</mo><mo>=</mo><msub><mi>L</mi><mi mathvariant=\"normal\">∞</mi></msub><mo>+</mo><mfrac><mi>A</mi><msup><mi>N</mi><mi>α</mi></msup></mfrac><mo>+</mo><mfrac><mi>B</mi><msup><mi>D</mi><mi>β</mi></msup></mfrac></mrow></semantics></math>\n<p>The parameter-limited term falls with <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>N</mi></mrow></semantics></math>. The data-limited term falls with <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>D</mi></mrow></semantics></math>. Compute optimality chooses a point on the budget constraint where spending more on either side would create a larger penalty on the other.</p>\n<figure><figcaption><strong>Fixed compute turns scaling into an allocation problem</strong> — The optimum comes from measured loss across candidate parameter and token budgets.</figcaption><img src=\"https://blog.ecitis.org/feeds/figures/scaling-laws-regimes-1.webp\" width=\"512\" height=\"704\" alt=\"Fixed compute turns scaling into an allocation problem — The optimum comes from measured loss across candidate parameter and token budgets.\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>The compute-optimal point answers a training allocation question. It does not automatically minimize deployment cost. A serving-heavy product may prefer fewer parameters because inference cost is paid repeatedly. A research model may accept a larger serving cost to reduce pretraining time. A model intended for continued training may deliberately stop before the static optimum.</p>\n<h2 id=\"data-quality-moves-the-curve\">Data quality moves the curve</h2>\n<p>Token count treats every token as interchangeable. Training does not.</p>\n<p>Deduplication, domain mixture, code proportion, language coverage, and filtering change the mapping from raw tokens to validation loss. DeepSeek’s scaling experiments reported that fitted laws differed across datasets, warning against transferring a curve without checking the data distribution.<sup><a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup></p>\n<p>Data quality creates two distinct questions:</p>\n<ol>\n<li>How does loss scale when the sampling distribution is fixed?</li>\n<li>How should the sampling distribution change as the budget grows?</li>\n</ol>\n<p>The first supports a clean fit. The second is the real dataset-design problem. If a larger run consumes lower-quality tail data, the original curve may overpredict improvement. If curation improves with scale, it may underpredict it.</p>\n<p>A scaling experiment should therefore record:</p>\n<ul>\n<li>unique and repeated token counts</li>\n<li>mixture weights by source and domain</li>\n<li>deduplication policy</li>\n<li>tokenizer and context length</li>\n<li>optimizer and learning-rate schedule</li>\n<li>training loss and several held-out distributions</li>\n</ul>\n<p>Without those fields, reproducing the curve is harder than reproducing the equation.</p>\n<h2 id=\"repeated-data-defines-another-regime\">Repeated data defines another regime</h2>\n<p>Abundant-data scaling assumes the next training token can be drawn from a large supply of useful, mostly unique text. That assumption becomes weaker when compute grows faster than the stock of high-quality data.</p>\n<p>Muennighoff and collaborators studied repeated-data training and found that a limited amount of repetition could retain much of the value of unique data in their experiments, while additional repetition eventually produced declining returns.<sup><a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup> Their scaling law adds a term for the reduced value of repeated tokens.</p>\n<p>A 2026 preprint proposes SoftQ, a coupled model of parameter count and data size for repeated-data training, arguing that the additive Chinchilla form can be misspecified in this regime.<sup><a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup> The paper evaluates SoftQ for repeated-data training.</p>\n<p>The conceptual point is durable: once examples repeat, model size and data size can interact. A larger model may memorize repeated data differently from a smaller model, so independent parameter and data terms can miss structure in the observations.</p>\n<h2 id=\"downstream-performance-is-not-one-curve\">Downstream performance is not one curve</h2>\n<p>Pretraining loss averages prediction error across tokens. A benchmark measures a narrower behavior after prompting, post-training, tool use, or sampling.</p>\n<p>Suppose two checkpoints have nearly equal validation loss. They can still differ on:</p>\n<ul>\n<li>rare facts</li>\n<li>long-horizon reasoning</li>\n<li>code execution</li>\n<li>calibration</li>\n<li>multilingual performance</li>\n<li>robustness to prompt format</li>\n</ul>\n<p>The relationship between loss and a downstream metric can also change when the benchmark saturates or when post-training contributes most of the measured capability.</p>\n<p>Use a hierarchy of forecasts:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Forecast</th><th>Strongest evidence</th></tr></thead><tbody><tr><td>Pretraining loss</td><td>Dense checkpoint curves on held-out text</td></tr><tr><td>Domain loss</td><td>Curves on a stable domain-specific corpus</td></tr><tr><td>Benchmark score</td><td>Repeated evaluations across checkpoints and scales</td></tr><tr><td>Product quality</td><td>Task-specific traffic, tools, latency, and post-training</td></tr></tbody></table></div>\n<p>Each step adds variables not captured by the pretraining curve.</p>\n<h2 id=\"scaling-laws-are-decision-tools\">Scaling laws are decision tools</h2>\n<p>A good scaling study runs before the largest training job, not after it. The workflow is:</p>\n<ol>\n<li>Train a matrix of smaller models across parameter and token budgets.</li>\n<li>Hold architecture, data policy, and optimization choices stable enough to isolate scale.</li>\n<li>Fit competing functional forms.</li>\n<li>Reserve some runs as out-of-sample tests.</li>\n<li>inspect residuals by model size, token budget, and domain.</li>\n<li>propagate fit uncertainty into the large-run forecast.</li>\n<li>rerun the fit when the architecture or data regime changes.</li>\n</ol>\n<p>The fitted optimum should be treated as a region, not a single sacred coordinate. Measurement noise, hardware constraints, checkpoint reuse, and serving economics can make nearby allocations operationally better.</p>\n<p>A scaling fit supports the measured relationship for one family of systems over one range. Researchers can use that fit to avoid some experimental runs. Extending it to new architectures, repeated data, post-training, test-time reasoning, or user value requires new measurements.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Jared Kaplan et al., “Scaling Laws for Neural Language Models,” arXiv, 2020. <a href=\"https://arxiv.org/abs/2001.08361\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2001.08361</a> <a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Jordan Hoffmann et al., “Training Compute-Optimal Large Language Models,” NeurIPS, 2022. <a href=\"https://arxiv.org/abs/2203.15556\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2203.15556</a> <a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>DeepSeek-AI et al., “DeepSeek LLM: Scaling Open-Source Language Models with Longtermism,” arXiv, 2024. <a href=\"https://arxiv.org/abs/2401.02954\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2401.02954</a> <a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Niklas Muennighoff et al., “Scaling Data-Constrained Language Models,” NeurIPS, 2023. <a href=\"https://arxiv.org/abs/2305.16264\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2305.16264</a> <a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>Zhiwei Xu et al., “Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws,” arXiv, 2026. <a href=\"https://arxiv.org/abs/2606.06888\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2606.06888</a> <a href=\"https://blog.ecitis.org/scaling-laws-regimes/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/scaling-laws-regimes/",
            "title": "Scaling Laws: What the Curves Predict and What They Hide",
            "summary": "Read language-model scaling laws as conditional empirical models across abundant-data, compute-optimal, and data-constrained regimes.",
            "image": "https://blog.ecitis.org/open-graph/scaling-laws-regimes.png",
            "date_modified": "2026-07-10T00:00:00.000Z",
            "date_published": "2026-07-10T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Scaling laws",
                "Pretraining",
                "Compute optimality",
                "Data"
            ]
        },
        {
            "id": "https://blog.ecitis.org/model-compression-techniques/",
            "content_html": "<p>Compression succeeds only when the deployment runtime benefits from the changed model. Removing isolated weights does not guarantee faster dense matrix multiplication. A low-rank factorization does not help if the two smaller operators launch inefficiently. Distillation can produce a smaller model, but its quality depends on the teacher signal and training distribution.</p>\n<figure><figcaption><strong>Compression changes different parts of the deployment contract</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/model-compression-techniques-0.webp\" width=\"512\" height=\"396\" alt=\"Compression changes different parts of the deployment contract\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"define-the-deployment-objective\">Define the deployment objective</h2>\n<p>Choose measurable constraints before modifying the model:</p>\n<ul>\n<li>resident weight memory;</li>\n<li>peak runtime memory;</li>\n<li>serialized artifact size;</li>\n<li>first-token latency;</li>\n<li>decode latency;</li>\n<li>throughput at the serving batch;</li>\n<li>energy or power under sustained load;</li>\n<li>task quality and safety slices.</li>\n</ul>\n<p>Parameter count is an intermediate metric. Hardware latency depends on shapes, sparsity support, memory traffic, and kernel launch behavior.</p>\n<h2 id=\"unstructured-pruning-removes-individual-weights\">Unstructured pruning removes individual weights</h2>\n<p>Magnitude pruning sets low-magnitude weights to zero. A binary mask <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>M</mi></mrow></semantics></math> produces:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mover accent=\"true\"><mi>W</mi><mo stretchy=\"true\">~</mo></mover><mo>=</mo><mi>M</mi><mo>⊙</mo><mi>W</mi></mrow></semantics></math>\n<p>SparseGPT instead formulates layer-wise pruning with approximate second-order information and reconstructs outputs while removing weights.<sup><a href=\"https://blog.ecitis.org/model-compression-techniques/#user-content-fn-sparsegpt\" id=\"user-content-fnref-sparsegpt\">1</a></sup> The algorithmic goal is to preserve layer behavior under sparsity, but speed still requires a sparse storage format and kernels that exploit the pattern.</p>\n<p>An unstructured sparse matrix stores nonzero values plus indices. At modest sparsity, index overhead and irregular access can outweigh skipped arithmetic. Benchmark the exact matrix shapes on the target runtime.</p>\n<p>Pruning should be followed by recovery tuning when quality loss exceeds the budget. Keep the mask fixed during recovery unless the method explicitly supports regrowth.</p>\n<h2 id=\"structured-pruning-changes-dense-dimensions\">Structured pruning changes dense dimensions</h2>\n<p>Structured pruning removes complete channels, attention heads, feed-forward neurons, or layers. The resulting matrices remain dense and can use ordinary kernels. This makes latency benefits easier to realize, but each removal changes more computation than one isolated weight.</p>\n<p>Head pruning must update:</p>\n<ul>\n<li>query, key, value, and output projection shapes;</li>\n<li>grouped-query head mapping;</li>\n<li>rotary or positional dimensions;</li>\n<li>tensor-parallel divisibility;</li>\n<li>checkpoint metadata and runtime configuration.</li>\n</ul>\n<p>Depth pruning removes whole transformer blocks. Width pruning changes hidden or feed-forward dimensions. NVIDIA’s Minitron work combines depth, width, attention, and feed-forward pruning with distillation-based retraining to derive smaller language models from a pretrained model.<sup><a href=\"https://blog.ecitis.org/model-compression-techniques/#user-content-fn-minitron\" id=\"user-content-fnref-minitron\">2</a></sup></p>\n<p>Measure every candidate architecture instead of assuming proportional latency. A narrower matrix can land on a less efficient hardware tile and produce less speedup than its parameter reduction suggests.</p>\n<h2 id=\"distillation-trains-a-separate-student\">Distillation trains a separate student</h2>\n<p>Knowledge distillation trains a student against teacher outputs, often using softened class probabilities. Hinton and collaborators described transferring the ensemble or teacher function into a deployable model through this training signal.<sup><a href=\"https://blog.ecitis.org/model-compression-techniques/#user-content-fn-distillation\" id=\"user-content-fnref-distillation\">3</a></sup></p>\n<p>For language models, token-level distillation can minimize:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>K</mi><mi>D</mi></mrow></msub><mo>=</mo><msup><mi>T</mi><mn>2</mn></msup><mi mathvariant=\"normal\">KL</mi><mo>⁡</mo><mrow><mo fence=\"true\">(</mo><msub><mi>p</mi><mi>T</mi></msub><mo stretchy=\"false\">(</mo><mo>⋅</mo><mo>∣</mo><mi>x</mi><mo stretchy=\"false\">)</mo><mo>∥</mo><msub><mi>p</mi><mi>S</mi></msub><mo stretchy=\"false\">(</mo><mo>⋅</mo><mo>∣</mo><mi>x</mi><mo stretchy=\"false\">)</mo><mo fence=\"true\">)</mo></mrow></mrow></semantics></math>\n<p>where <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>T</mi></mrow></semantics></math> is temperature. The student can also train on hard target tokens, hidden-state matches, attention relations, or generated teacher data.</p>\n<p>Teacher logits expose more information than one sampled token, but full-vocabulary logits are expensive to store. Options include online teacher inference, top-logit storage with a residual bucket, or sequence-level distillation from teacher-generated outputs. Each changes the target.</p>\n<p>Distillation risks copying teacher errors and narrowing behavior to the transfer distribution. Evaluate the student on real held-out data and safety tests that were not generated by the teacher.</p>\n<h2 id=\"low-rank-factorization-replaces-one-operator-with-two\">Low-rank factorization replaces one operator with two</h2>\n<p>For <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>W</mi><mo>∈</mo><msup><mi mathvariant=\"double-struck\">R</mi><mrow><mi>m</mi><mo>×</mo><mi>n</mi></mrow></msup></mrow></semantics></math>, a rank-<math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>r</mi></mrow></semantics></math> approximation writes:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>W</mi><mo>≈</mo><mi>A</mi><mi>B</mi><mo separator=\"true\">,</mo><mspace width=\"2em\"></mspace><mi>A</mi><mo>∈</mo><msup><mi mathvariant=\"double-struck\">R</mi><mrow><mi>m</mi><mo>×</mo><mi>r</mi></mrow></msup><mo separator=\"true\">,</mo><mspace width=\"1em\"></mspace><mi>B</mi><mo>∈</mo><msup><mi mathvariant=\"double-struck\">R</mi><mrow><mi>r</mi><mo>×</mo><mi>n</mi></mrow></msup></mrow></semantics></math>\n<p>The parameter count changes from <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>m</mi><mi>n</mi></mrow></semantics></math> to <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>r</mi><mo stretchy=\"false\">(</mo><mi>m</mi><mo>+</mo><mi>n</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math>. This is smaller only when <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>r</mi></mrow></semantics></math> is below the break-even rank:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>r</mi><mo>&lt;</mo><mfrac><mrow><mi>m</mi><mi>n</mi></mrow><mrow><mi>m</mi><mo>+</mo><mi>n</mi></mrow></mfrac></mrow></semantics></math>\n<p>Truncated singular value decomposition provides a reconstruction-optimal low-rank approximation under the Frobenius norm, but minimizing weight reconstruction error does not directly minimize model loss. Recovery tuning can adapt the factors to the data distribution.</p>\n<p>Runtime performance depends on both factor matrices. Very small intermediate rank can become launch-bound. A fused implementation or batching may be needed to realize latency gains.</p>\n<p>Low-rank adapters such as LoRA are parameter-efficient fine-tuning mechanisms, not automatic base-model compression. Keeping the frozen base matrix plus adapters adds parameters. Compression occurs only if the deployed representation replaces or restructures the base computation.</p>\n<h2 id=\"combine-methods-in-a-deliberate-order\">Combine methods in a deliberate order</h2>\n<p>Pruning, factorization, and distillation interact. A useful sequence is:</p>\n<ol>\n<li>identify deployable structured shapes;</li>\n<li>prune the teacher or initialize the student architecture;</li>\n<li>reconstruct or factorize eligible matrices;</li>\n<li>recover with hard labels and teacher targets;</li>\n<li>apply quantization after the graph is stable;</li>\n<li>benchmark the exported artifact.</li>\n</ol>\n<p>This is not universal. Quantization-aware training may need to participate in recovery. Unstructured sparse kernels may require pruning to a hardware-specific pattern before export.</p>\n<p>Change one compression axis at a time during diagnosis. If pruning, factorization, and quantization are applied together, a quality regression cannot be localized.</p>\n<h2 id=\"preserve-tied-and-shared-parameters\">Preserve tied and shared parameters</h2>\n<p>Language models can tie input embeddings to the output projection or reuse parameters across layers. A generic traversal that replaces modules independently can break aliasing and duplicate storage.</p>\n<p>Before compression, build an alias map:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>owners </span><span>=</span><span> {}</span></span>\n<span class=\"line\"><span>for</span><span> name</span><span>,</span><span> parameter </span><span>in</span><span> model</span><span>.</span><span>named_parameters</span><span>(remove_duplicate</span><span>=</span><span>False</span><span>):</span></span>\n<span class=\"line\"><span>    owners</span><span>.</span><span>setdefault</span><span>(</span><span>id</span><span>(parameter), []).</span><span>append</span><span>(name)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>shared </span><span>=</span><span> [names </span><span>for</span><span> names </span><span>in</span><span> owners</span><span>.</span><span>values</span><span>()</span><span> if</span><span> len</span><span>(names)</span><span> &gt;</span><span> 1</span><span>]</span></span></code></pre>\n<p>After rewriting the graph, confirm that intended aliases remain tied and the optimizer does not update duplicate copies.</p>\n<h2 id=\"validate-exported-graphs\">Validate exported graphs</h2>\n<p>The in-memory training model is not the deployment artifact. Exporters can fold constants, transpose weights, reject sparse operators, or lower factorized layers into unexpected subgraphs.</p>\n<p>For each target runtime:</p>\n<ul>\n<li>inspect operator types and tensor shapes;</li>\n<li>confirm sparse metadata survives;</li>\n<li>compare serialized size;</li>\n<li>run fixed input fixtures against the source model;</li>\n<li>measure peak memory and latency after warm-up;</li>\n<li>profile the kernels actually dispatched;</li>\n<li>test all supported batch and sequence buckets.</li>\n</ul>\n<p>A compressed checkpoint that falls back to CPU for one unsupported sparse operator can be slower than the dense baseline.</p>\n<h2 id=\"quality-evaluation-needs-slices\">Quality evaluation needs slices</h2>\n<p>Compression error is not uniform. Pruning can damage rare features. Distillation can suppress low-probability alternatives. Low-rank approximation can remove directions that matter for a small task slice.</p>\n<p>Evaluate:</p>\n<ul>\n<li>perplexity or task loss on real held-out text;</li>\n<li>downstream accuracy by domain and language;</li>\n<li>long-context behavior;</li>\n<li>calibration and uncertainty;</li>\n<li>tool-call schema validity;</li>\n<li>safety and refusal behavior;</li>\n<li>generation repetition and degeneration.</li>\n</ul>\n<p>Compare against an uncompressed checkpoint under identical decoding settings. If the compressed model needs a different sampler to recover quality, report both the model and sampler change.</p>\n<h2 id=\"recovery-failures\">Recovery failures</h2>\n<p><strong>Mask drift:</strong> Pruned weights become nonzero during tuning because the optimizer updates them. Reapply the mask after updates or use masked parameterization.</p>\n<p><strong>Shape mismatch:</strong> Structured pruning updates weights but not configuration, rotary tables, or tensor-parallel metadata. Load the exported model in a fresh process before declaring success.</p>\n<p><strong>Teacher leakage:</strong> Evaluation prompts or their derivatives enter distillation data. Split source groups before teacher generation.</p>\n<p><strong>Rank collapse:</strong> A factorization uses one rank for every layer despite different spectra and sensitivity. Allocate rank by measured reconstruction and downstream loss.</p>\n<p><strong>No runtime gain:</strong> The artifact is smaller but latency is unchanged. Inspect kernels, launch count, memory bandwidth, and sparse support.</p>\n<p>Compression is a contract between an algorithm and a runtime. Unstructured pruning needs sparse execution, structured pruning needs valid dense shapes, distillation needs independent evaluation, and factorization needs efficient paired operators. The final evidence is an exported model that meets memory, latency, and quality budgets on its target hardware.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-sparsegpt\">\n<p><a href=\"https://arxiv.org/abs/2301.00774\" rel=\"noopener noreferrer\">SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot</a>. <a href=\"https://blog.ecitis.org/model-compression-techniques/#user-content-fnref-sparsegpt\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-minitron\">\n<p><a href=\"https://arxiv.org/abs/2407.14679\" rel=\"noopener noreferrer\">Compact Language Models via Pruning and Knowledge Distillation</a>. <a href=\"https://blog.ecitis.org/model-compression-techniques/#user-content-fnref-minitron\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-distillation\">\n<p><a href=\"https://arxiv.org/abs/1503.02531\" rel=\"noopener noreferrer\">Distilling the Knowledge in a Neural Network</a>. <a href=\"https://blog.ecitis.org/model-compression-techniques/#user-content-fnref-distillation\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/model-compression-techniques/",
            "title": "Model Compression Beyond Quantization: Pruning, Distillation, and Low-Rank Methods",
            "summary": "Choose pruning, distillation, and low-rank methods by the deployment gains they produce rather than parameter count alone.",
            "image": "https://blog.ecitis.org/open-graph/model-compression-techniques.png",
            "date_modified": "2026-07-07T00:00:00.000Z",
            "date_published": "2026-07-07T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Compression",
                "Distillation",
                "Pruning"
            ]
        },
        {
            "id": "https://blog.ecitis.org/model-alignment-control-loop/",
            "content_html": "<p>Post-training pipelines collect preferences, optimize a policy, and test whether the assistant is helpful and safe. The training signal is a proxy for intended behavior. The model must generalize that proxy to inputs, tools, and situations absent from training.</p>\n<p>Use the following operational definition:</p>\n<blockquote>\n<p>An aligned model follows the intended specification when the evaluator is weak, the context is unfamiliar, and shortcuts are available.</p>\n</blockquote>\n<p>This definition turns alignment into a control problem. The system needs a specification, a way to convert that specification into supervision, an optimizer, and measurements that expose failures. Any one of those parts can be wrong while benchmark scores improve.</p>\n<figure><figcaption><strong>Alignment closes the gap between intent and deployed behavior</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/model-alignment-control-loop-0.webp\" width=\"512\" height=\"866\" alt=\"Alignment closes the gap between intent and deployed behavior\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"preferences-train-a-proxy\">Preferences train a proxy</h2>\n<p>The standard reinforcement learning from human feedback pipeline begins with a pretrained model, supervised instruction tuning, pairwise preference data, a reward model, and policy optimization. InstructGPT established this pipeline for instruction-following language models.<sup><a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup></p>\n<p>For prompt <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>x</mi></mrow></semantics></math>, completion <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>y</mi></mrow></semantics></math>, learned reward <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>r</mi><mi>ϕ</mi></msub><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>y</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math>, and reference policy <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>π</mi><mrow><mi mathvariant=\"normal\">r</mi><mi mathvariant=\"normal\">e</mi><mi mathvariant=\"normal\">f</mi></mrow></msub></mrow></semantics></math>, the policy objective is commonly written as:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><munder><mrow><mi>max</mi><mo>⁡</mo></mrow><mi>θ</mi></munder><mtext>  </mtext><msub><mi mathvariant=\"double-struck\">E</mi><mrow><mi>y</mi><mo>∼</mo><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mo>⋅</mo><mo>∣</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></msub><mrow><mo fence=\"true\">[</mo><msub><mi>r</mi><mi>ϕ</mi></msub><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>y</mi><mo stretchy=\"false\">)</mo><mo>−</mo><mi>β</mi><mi>log</mi><mo>⁡</mo><mfrac><mrow><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mi>y</mi><mo>∣</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><msub><mi>π</mi><mrow><mi mathvariant=\"normal\">r</mi><mi mathvariant=\"normal\">e</mi><mi mathvariant=\"normal\">f</mi></mrow></msub><mo stretchy=\"false\">(</mo><mi>y</mi><mo>∣</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo fence=\"true\">]</mo></mrow></mrow></semantics></math>\n<p>The reward term pulls behavior toward the preference model. The KL term limits movement away from the reference policy. Neither term states the desired behavior directly. The reward model learns it from comparisons, and the comparisons reflect annotator instructions, sampling policy, prompt distribution, and labeling noise.</p>\n<p>Direct Preference Optimization rewrites this constrained reward problem as a classification loss on preferred and rejected responses.<sup><a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> It removes the separately trained reward model and online reinforcement learning loop, but it does not remove the proxy problem. The chosen response still represents a judgment made under a particular rubric and context.</p>\n<p>These layers can fail independently:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Layer</th><th>What it contains</th><th>Common failure</th></tr></thead><tbody><tr><td>Intended behavior</td><td>What the system should do</td><td>Ambiguous or conflicting goals</td></tr><tr><td>Written specification</td><td>Policies, principles, examples</td><td>Gaps and underspecified edge cases</td></tr><tr><td>Supervision</td><td>Preferences, critiques, step labels</td><td>Evaluator error or shortcut features</td></tr><tr><td>Training objective</td><td>Reward or preference loss</td><td>Proxy optimization</td></tr><tr><td>Deployed policy</td><td>Behavior under new conditions</td><td>Distribution shift or strategic adaptation</td></tr></tbody></table></div>\n<p>A higher reward is evidence that the policy fits the learned proxy. It is not proof that the proxy captures the intended behavior.</p>\n<figure><figcaption><strong>Alignment supervision is a stack of proxies</strong> — Evidence at one layer does not validate every layer above it.</figcaption><img src=\"https://blog.ecitis.org/feeds/figures/model-alignment-control-loop-1.webp\" width=\"512\" height=\"804\" alt=\"Alignment supervision is a stack of proxies — Evidence at one layer does not validate every layer above it.\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"written-principles-make-the-target-inspectable\">Written principles make the target inspectable</h2>\n<p>Constitutional AI replaces some human harmlessness labels with model-generated critiques and preferences conditioned on written principles.<sup><a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup> Its supervised phase asks a model to critique and revise responses. Its reinforcement learning phase uses AI-generated preferences to train a preference model.</p>\n<p>This changes where human judgment enters the system. Humans write or approve principles rather than label every sampled answer. The principles are easier to inspect than a reward model’s weights, but their interpretation still depends on the model producing the critique.</p>\n<p>Deliberative alignment moves the specification closer to inference. A reasoning model is trained to recall and apply written safety policies before producing an answer.<sup><a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup> This can improve decisions that require combining several policy clauses. It also creates a new object to evaluate: whether the model selected the relevant policy and applied it to the facts of the request.</p>\n<p>The two methods expose different failure surfaces:</p>\n<ul>\n<li>Constitutional feedback can misapply a principle while generating training labels.</li>\n<li>Deliberation can retrieve the wrong rule or construct a convenient interpretation.</li>\n<li>A complete written policy can still conflict with another policy or omit a deployment-specific constraint.</li>\n<li>A correct rationale can be paired with an action caused by another internal mechanism.</li>\n</ul>\n<p>Written specifications improve inspectability. They do not guarantee faithful internal reasoning.</p>\n<h2 id=\"process-supervision-changes-what-receives-credit\">Process supervision changes what receives credit</h2>\n<p>Outcome supervision rewards a final answer. Process supervision evaluates intermediate steps. On mathematical reasoning tasks, the work behind PRM800K found that a process-supervised reward model selected better solutions than an outcome-supervised reward model in the evaluated setting.<sup><a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup></p>\n<p>The alignment relevance is credit assignment. An outcome-only reward can reinforce a correct answer reached through an invalid or unsafe process. Step-level labels can penalize the first bad transition before later steps conceal it.</p>\n<p>Process supervision also has costs:</p>\n<ul>\n<li>Labels are more expensive because an evaluator must inspect a trajectory.</li>\n<li>A visible chain of thought may not faithfully expose the computation that caused the answer.</li>\n<li>Fine-grained rubrics can reward verbose compliance with the rubric rather than efficient reasoning.</li>\n<li>Tool-using agents require process labels for text, actions, state changes, and omissions.</li>\n</ul>\n<p>Choose process checkpoints according to risk. A coding agent can be evaluated at each file mutation or shell command. A medical assistant may need evidence checks before advice. A long-horizon agent needs stateful constraints that survive across turns.</p>\n<h2 id=\"oversight-must-work-when-the-model-is-stronger\">Oversight must work when the model is stronger</h2>\n<p>Human preference learning assumes the evaluator can recognize a better answer. That assumption weakens as models operate in domains where checking a solution is harder than producing one.</p>\n<p>Weak-to-strong generalization studies an empirical proxy for this problem: a weaker model supervises a stronger model, and researchers measure whether the strong model can exceed the supervisor’s errors rather than imitate them.<sup><a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fn-6\" id=\"user-content-fnref-6\">6</a></sup> The reported results show partial recovery of stronger capabilities on some tasks, alongside failures on preference data. The setup is evidence that the question can be tested, not a solution to superhuman oversight.</p>\n<p>Scalable oversight methods try to restructure evaluation:</p>\n<ul>\n<li>decompose a hard task into locally checkable claims</li>\n<li>use one model to help a human inspect another model</li>\n<li>require executable tests, proofs, or external evidence</li>\n<li>compare independent solutions before revealing either rationale</li>\n<li>assign more evaluator effort to high-uncertainty or high-impact decisions</li>\n</ul>\n<p>The evaluator must remain independent enough to catch shared mistakes. Two models trained on similar data can agree for the same wrong reason.</p>\n<h2 id=\"model-organisms-make-hidden-objectives-testable\">Model organisms make hidden objectives testable</h2>\n<p>Behavioral red teaming is weak when a model behaves differently after detecting an evaluation. Model organisms address this by deliberately training models with known hidden objectives or failure propensities, then asking auditors to recover them.<sup><a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fn-7\" id=\"user-content-fnref-7\">7</a></sup></p>\n<p>This creates a ground truth for audit research. If the experimenter knows the hidden objective, an audit can be scored on whether it identifies the cause and its behavioral traces.</p>\n<p>Anthropic’s automated auditing work combines behavioral evaluations, training-data inspection, and interpretability tools against such models.<sup><a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fn-7\" id=\"user-content-fnref-7-2\">7</a></sup> The approach tests an entire investigation workflow rather than one probe. It also reveals a deployment lesson: a monitor should generate falsifiable hypotheses, design counterexamples, and seek causal evidence.</p>\n<p>Useful causal checks include:</p>\n<ol>\n<li>Find a representation or behavior correlated with the suspected objective.</li>\n<li>Remove or steer that representation.</li>\n<li>Measure whether the predicted behavior changes.</li>\n<li>Test unrelated capabilities to detect broad damage.</li>\n<li>Repeat on prompts that do not announce the evaluation.</li>\n</ol>\n<p>Correlation can point to a mechanism. Intervention tests whether the mechanism carries causal weight.</p>\n<h2 id=\"what-an-alignment-report-should-contain\">What an alignment report should contain</h2>\n<p>A deployment report should separate the layers of evidence:</p>\n<ul>\n<li>the written specification and unresolved conflicts</li>\n<li>the source and coverage of supervision</li>\n<li>held-out behavioral evaluations</li>\n<li>adversarial tests created independently of training</li>\n<li>process-level checks for consequential actions</li>\n<li>monitor false positives and false negatives</li>\n<li>causal interventions where interpretability claims support a safety decision</li>\n<li>residual risks that the evaluation cannot measure</li>\n</ul>\n<p>An alignment claim needs a stated threat model and operating environment. Its evidence should connect an inspectable specification to traceable supervision, audits that expose known hidden failures, and deployment monitoring that detects behavior outside the tested regime.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Long Ouyang et al., “Training language models to follow instructions with human feedback,” NeurIPS, 2022. <a href=\"https://arxiv.org/abs/2203.02155\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2203.02155</a> <a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Rafael Rafailov et al., “Direct Preference Optimization: Your Language Model is Secretly a Reward Model,” NeurIPS, 2023. <a href=\"https://arxiv.org/abs/2305.18290\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2305.18290</a> <a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Yuntao Bai et al., “Constitutional AI: Harmlessness from AI Feedback,” Anthropic, 2022. <a href=\"https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback\" rel=\"noopener noreferrer\">https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback</a> <a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Melody Guan et al., “Deliberative alignment: reasoning enables safer language models,” OpenAI, 2024. <a href=\"https://openai.com/index/deliberative-alignment/\" rel=\"noopener noreferrer\">https://openai.com/index/deliberative-alignment/</a> <a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>Hunter Lightman et al., “Let’s Verify Step by Step,” ICLR, 2024. <a href=\"https://arxiv.org/abs/2305.20050\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2305.20050</a> <a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p>Collin Burns et al., “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision,” ICML, 2024. <a href=\"https://openai.com/index/weak-to-strong-generalization/\" rel=\"noopener noreferrer\">https://openai.com/index/weak-to-strong-generalization/</a> <a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p>Trenton Bricken et al., “Building and evaluating alignment auditing agents,” Anthropic, 2025. <a href=\"https://alignment.anthropic.com/2025/automated-auditing/\" rel=\"noopener noreferrer\">https://alignment.anthropic.com/2025/automated-auditing/</a> <a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/model-alignment-control-loop/#user-content-fnref-7-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/model-alignment-control-loop/",
            "title": "Model Alignment Is a Control Problem",
            "summary": "Connect specifications, preference learning, process supervision, scalable oversight, and causal audits into one alignment control loop.",
            "image": "https://blog.ecitis.org/open-graph/model-alignment-control-loop.png",
            "date_modified": "2026-07-05T00:00:00.000Z",
            "date_published": "2026-07-05T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Model alignment",
                "RLHF",
                "Scalable oversight",
                "AI safety"
            ]
        },
        {
            "id": "https://blog.ecitis.org/llm-training-optimizers/",
            "content_html": "<p>An optimizer defines both an update rule and a distributed state system. For a large transformer, moment tensors, matrix statistics, preconditioner roots, communication, and checkpoint format can consume as much engineering attention as the gradient formula.</p>\n<figure><figcaption><strong>Optimizer state attached to one matrix parameter W</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/llm-training-optimizers-0.webp\" width=\"512\" height=\"386\" alt=\"Optimizer state attached to one matrix parameter W\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"adamw-stores-coordinate-wise-moments\">AdamW stores coordinate-wise moments</h2>\n<p>Adam maintains exponential moving averages of gradients and squared gradients:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>m</mi><mi>t</mi></msub><mo>=</mo><msub><mi>β</mi><mn>1</mn></msub><msub><mi>m</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>+</mo><mo stretchy=\"false\">(</mo><mn>1</mn><mo>−</mo><msub><mi>β</mi><mn>1</mn></msub><mo stretchy=\"false\">)</mo><msub><mi>g</mi><mi>t</mi></msub></mrow></semantics></math>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>v</mi><mi>t</mi></msub><mo>=</mo><msub><mi>β</mi><mn>2</mn></msub><msub><mi>v</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>+</mo><mo stretchy=\"false\">(</mo><mn>1</mn><mo>−</mo><msub><mi>β</mi><mn>2</mn></msub><mo stretchy=\"false\">)</mo><msubsup><mi>g</mi><mi>t</mi><mn>2</mn></msubsup></mrow></semantics></math>\n<p>Bias-corrected estimates scale the update coordinate-wise. The original Adam paper introduced this adaptive first-order method.<sup><a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fn-adam\" id=\"user-content-fnref-adam\">1</a></sup></p>\n<p>AdamW applies weight decay separately from the loss-gradient update:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>θ</mi><mrow><mi>t</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>=</mo><mo stretchy=\"false\">(</mo><mn>1</mn><mo>−</mo><mi>η</mi><mi>λ</mi><mo stretchy=\"false\">)</mo><msub><mi>θ</mi><mi>t</mi></msub><mo>−</mo><mi>η</mi><mfrac><msub><mover accent=\"true\"><mi>m</mi><mo>^</mo></mover><mi>t</mi></msub><mrow><msqrt><msub><mover accent=\"true\"><mi>v</mi><mo>^</mo></mover><mi>t</mi></msub></msqrt><mo>+</mo><mi>ϵ</mi></mrow></mfrac></mrow></semantics></math>\n<p>Decoupling matters because adding an <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>L</mi><mn>2</mn></msub></mrow></semantics></math> penalty to the gradient is not equivalent to weight decay under an adaptive preconditioner.<sup><a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fn-adamw\" id=\"user-content-fnref-adamw\">2</a></sup></p>\n<p>Each parameter normally has full-shape first- and second-moment state. Distributed optimizers shard these states across data-parallel ranks and gather updated parameter partitions as needed. Checkpoint conversion must preserve the association among a parameter and both moments.</p>\n<p>Exclude parameters such as normalization scales or biases from decay only through an explicit, reviewed grouping rule:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>decay</span><span>,</span><span> no_decay </span><span>=</span><span> []</span><span>,</span><span> []</span></span>\n<span class=\"line\"><span>for</span><span> name</span><span>,</span><span> parameter </span><span>in</span><span> model</span><span>.</span><span>named_parameters</span><span>():</span></span>\n<span class=\"line\"><span>    if</span><span> not</span><span> parameter</span><span>.</span><span>requires_grad</span><span>:</span></span>\n<span class=\"line\"><span>        continue</span></span>\n<span class=\"line\"><span>    target </span><span>=</span><span> no_decay </span><span>if</span><span> is_no_decay_parameter</span><span>(name, parameter)</span><span> else</span><span> decay</span></span>\n<span class=\"line\"><span>    target</span><span>.</span><span>append</span><span>(parameter)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>groups </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>    {</span><span>\"params\"</span><span>:</span><span> decay</span><span>,</span><span> \"weight_decay\"</span><span>:</span><span> weight_decay</span><span>},</span></span>\n<span class=\"line\"><span>    {</span><span>\"params\"</span><span>:</span><span> no_decay</span><span>,</span><span> \"weight_decay\"</span><span>:</span><span> 0.0</span><span>},</span></span>\n<span class=\"line\"><span>]</span></span></code></pre>\n<p>Test that every trainable parameter appears once and only once.</p>\n<h2 id=\"adafactor-factors-second-moment-state\">Adafactor factors second-moment state</h2>\n<p>For a matrix parameter, Adafactor approximates the full second-moment accumulator using row and column statistics. This changes state from matrix-sized storage to vectors sized by its dimensions. The Adafactor paper presents sublinear auxiliary storage for large parameter matrices.<sup><a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fn-adafactor\" id=\"user-content-fnref-adafactor\">3</a></sup></p>\n<p>The approximation preserves less coordinate-level information than a full accumulator. It is most useful when large matrix parameters dominate. Vectors and small tensors need a different state rule because row-column factorization is not meaningful.</p>\n<p>Adafactor also introduced relative-step and update-clipping choices. Libraries expose different defaults, so a configuration copied by optimizer name alone may not reproduce another implementation. Record:</p>\n<ul>\n<li>whether first-moment momentum is enabled;</li>\n<li>relative or externally scheduled learning rate;</li>\n<li>parameter-scale adaptation;</li>\n<li>update clipping threshold;</li>\n<li>epsilon conventions;</li>\n<li>weight decay placement.</li>\n</ul>\n<p>Memory savings are not free if the training recipe silently changes several of these at once. Run an ablation that holds schedule and decay semantics constant where the implementation permits it.</p>\n<h2 id=\"shampoo-preconditions-tensor-dimensions\">Shampoo preconditions tensor dimensions</h2>\n<p>Shampoo builds statistics along tensor dimensions and applies matrix inverse roots as a preconditioner. For a matrix gradient <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>G</mi></mrow></semantics></math>, simplified left and right statistics accumulate forms related to:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>L</mi><mi>t</mi></msub><mo>=</mo><msub><mi>L</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>+</mo><msub><mi>G</mi><mi>t</mi></msub><msubsup><mi>G</mi><mi>t</mi><mi mathvariant=\"normal\">⊤</mi></msubsup></mrow></semantics></math>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>R</mi><mi>t</mi></msub><mo>=</mo><msub><mi>R</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>+</mo><msubsup><mi>G</mi><mi>t</mi><mi mathvariant=\"normal\">⊤</mi></msubsup><msub><mi>G</mi><mi>t</mi></msub></mrow></semantics></math>\n<p>The preconditioned update uses inverse roots of <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>L</mi><mi>t</mi></msub></mrow></semantics></math> and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>R</mi><mi>t</mi></msub></mrow></semantics></math>. The Shampoo paper develops this tensor-aware preconditioning while controlling dependence on the full flattened parameter dimension.<sup><a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fn-shampoo\" id=\"user-content-fnref-shampoo\">4</a></sup></p>\n<p>Matrix roots are more expensive than elementwise square roots. Practical systems update preconditioners periodically, block large dimensions, run root computations in higher precision, or distribute them across ranks. These implementation choices define the optimizer’s real cost.</p>\n<p>Monitor:</p>\n<ul>\n<li>time spent updating statistics;</li>\n<li>time spent computing roots;</li>\n<li>condition and numerical validity of preconditioners;</li>\n<li>blocked versus unblocked parameter coverage;</li>\n<li>communication for distributed statistics;</li>\n<li>fallback path when a root computation fails.</li>\n</ul>\n<p>If a preconditioner contains non-finite values, skipping only that parameter update can desynchronize replicas unless every rank makes the same decision.</p>\n<h2 id=\"muon-orthogonalizes-matrix-updates\">Muon orthogonalizes matrix updates</h2>\n<p>Muon applies momentum to matrix gradients, then approximately orthogonalizes the update with an iterative Newton-Schulz procedure. The scalable Muon work applies it to hidden-layer matrix parameters while using AdamW for embeddings, output heads, and other parameters.<sup><a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fn-muon\" id=\"user-content-fnref-muon\">5</a></sup></p>\n<p>This split is part of the algorithmic contract. Matrix orthogonalization is not defined the same way for vectors, and very rectangular or partitioned matrices require careful handling.</p>\n<p>A conceptual update is:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>momentum </span><span>=</span><span> update_momentum</span><span>(gradient, momentum, beta)</span></span>\n<span class=\"line\"><span>matrix_update </span><span>=</span><span> orthogonalize_newton_schulz</span><span>(momentum)</span></span>\n<span class=\"line\"><span>parameter </span><span>-=</span><span> learning_rate </span><span>*</span><span> scale</span><span>(parameter.shape)</span><span> *</span><span> matrix_update</span></span></code></pre>\n<p>The iteration count, polynomial coefficients, normalization, accumulation dtype, and shape-dependent scale affect output. Pin them in checkpoints and experiment records.</p>\n<p>Tensor parallelism creates another question: orthogonalize each local shard or the logical global matrix. Local orthogonalization is cheaper but is a different preconditioner. The training code must state which object the optimizer sees.</p>\n<h2 id=\"compare-state-and-compute-separately\">Compare state and compute separately</h2>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Optimizer</th><th>Dominant state for a matrix</th><th>Extra compute</th><th>Main systems concern</th></tr></thead><tbody><tr><td>AdamW</td><td>full first and second moments</td><td>elementwise updates</td><td>state sharding and bandwidth</td></tr><tr><td>Adafactor</td><td>factored second moment, optional momentum</td><td>reductions over dimensions</td><td>implementation-specific schedule</td></tr><tr><td>Shampoo</td><td>dimension-wise covariance matrices</td><td>periodic matrix roots</td><td>blocking, precision, distribution</td></tr><tr><td>Muon</td><td>momentum plus temporary orthogonalization state</td><td>iterative matrix products</td><td>parameter selection and sharding</td></tr></tbody></table></div>\n<p>This table describes structure, not a universal speed ranking. Kernel fusion, dtype, matrix shape, update frequency, and network topology determine runtime.</p>\n<h2 id=\"learning-rate-transfer-is-optimizer-specific\">Learning-rate transfer is optimizer-specific</h2>\n<p>Do not compare optimizers under one shared learning rate and call the result controlled. Their preconditioners produce updates with different scales. Tune learning rate, warmup, decay, clipping, and optimizer-specific parameters on a smaller proxy that preserves architecture and data characteristics.</p>\n<p>Track the update-to-weight norm:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>ρ</mi><mi>t</mi></msub><mo>=</mo><mfrac><mrow><mo stretchy=\"false\">∥</mo><mi mathvariant=\"normal\">Δ</mi><msub><mi>θ</mi><mi>t</mi></msub><mo stretchy=\"false\">∥</mo></mrow><mrow><mo stretchy=\"false\">∥</mo><msub><mi>θ</mi><mi>t</mi></msub><mo stretchy=\"false\">∥</mo></mrow></mfrac></mrow></semantics></math>\n<p>Report it by parameter class. Embeddings, attention projections, feed-forward matrices, and normalization parameters can behave differently. Aggregate loss can hide one group with unstable updates.</p>\n<h2 id=\"clipping-occurs-at-a-specific-point\">Clipping occurs at a specific point</h2>\n<p>Gradient norm clipping before optimizer state update is not equivalent to clipping the preconditioned update. Adafactor’s update clipping and external gradient clipping are different mechanisms. Muon’s orthogonalized update changes norm geometry again.</p>\n<p>Document the order:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>unscale mixed-precision gradients</span></span>\n<span class=\"line\"><span>check finite values</span></span>\n<span class=\"line\"><span>apply gradient clipping</span></span>\n<span class=\"line\"><span>update optimizer statistics</span></span>\n<span class=\"line\"><span>construct preconditioned update</span></span>\n<span class=\"line\"><span>apply decoupled weight decay</span></span>\n<span class=\"line\"><span>update parameters</span></span>\n<span class=\"line\"><span>advance scheduler</span></span></code></pre>\n<p>Distributed global-norm clipping requires a reduction over all shards. A local norm threshold changes with shard count.</p>\n<figure><figcaption><strong>Optimizer operations are ordered, not interchangeable</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/llm-training-optimizers-1.webp\" width=\"512\" height=\"468\" alt=\"Optimizer operations are ordered, not interchangeable\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"checkpoint-and-recovery-tests\">Checkpoint and recovery tests</h2>\n<p>An optimizer checkpoint must include step counters, moments or factors, preconditioner statistics, loss-scaler state, scheduler state, and parameter-group mapping. For periodic Shampoo roots, include enough state to resume the same update cadence. For Muon, preserve momentum and the exact parameter classification.</p>\n<p>Run a deterministic recovery test:</p>\n<ol>\n<li>Train to a checkpoint boundary.</li>\n<li>Save model, optimizer, scheduler, random state, and data position.</li>\n<li>Continue for a fixed batch sequence.</li>\n<li>Restore in a fresh process and replay the same sequence.</li>\n<li>Compare loss, parameters, and optimizer state under a documented tolerance.</li>\n</ol>\n<p>Topology changes require resharding. Verify on a small checkpoint before applying it to a full run.</p>\n<h2 id=\"select-by-bottleneck\">Select by bottleneck</h2>\n<p>Choose AdamW when a proven recipe, broad implementation support, and predictable coordinate-wise adaptation matter more than state volume. Choose Adafactor when matrix state is the binding capacity constraint and the recipe has been qualified. Evaluate Shampoo when training can amortize stronger preconditioning and the system can support matrix-statistic work. Evaluate Muon when hidden matrix updates dominate and the implementation defines parameter selection, scaling, and distributed semantics.</p>\n<p>No optimizer removes the need to inspect data, model initialization, and loss scaling. The useful comparison includes convergence per processed token, wall-clock time, peak memory, communication, recovery behavior, and final evaluation under equal data budgets.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-adam\">\n<p><a href=\"https://arxiv.org/abs/1412.6980\" rel=\"noopener noreferrer\">Adam: A Method for Stochastic Optimization</a>. <a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fnref-adam\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-adamw\">\n<p><a href=\"https://arxiv.org/abs/1711.05101\" rel=\"noopener noreferrer\">Decoupled Weight Decay Regularization</a>. <a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fnref-adamw\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-adafactor\">\n<p><a href=\"https://arxiv.org/abs/1804.04235\" rel=\"noopener noreferrer\">Adafactor: Adaptive Learning Rates with Sublinear Memory Cost</a>. <a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fnref-adafactor\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-shampoo\">\n<p><a href=\"https://arxiv.org/abs/1802.09568\" rel=\"noopener noreferrer\">Shampoo: Preconditioned Stochastic Tensor Optimization</a>. <a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fnref-shampoo\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-muon\">\n<p><a href=\"https://arxiv.org/abs/2502.16982\" rel=\"noopener noreferrer\">Muon is Scalable for LLM Training</a>. <a href=\"https://blog.ecitis.org/llm-training-optimizers/#user-content-fnref-muon\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/llm-training-optimizers/",
            "title": "LLM Training Optimizers: AdamW, Adafactor, Shampoo, and Muon",
            "summary": "Compare optimizer update rules alongside the memory, communication, numerical stability, and checkpoint costs they create at scale.",
            "image": "https://blog.ecitis.org/open-graph/llm-training-optimizers.png",
            "date_modified": "2026-07-02T00:00:00.000Z",
            "date_published": "2026-07-02T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Optimizers",
                "AdamW",
                "Distributed training"
            ]
        },
        {
            "id": "https://blog.ecitis.org/frontier-mechanistic-interpretability/",
            "content_html": "<p>Mechanistic interpretability tries to explain model behavior in terms of internal computation. The frontier has moved beyond finding a neuron that correlates with a concept. Current methods decompose activations into features, trace interactions between those features, translate internal states into text, and intervene to test causal claims.</p>\n<p>These tools should not be ranked on one axis. They expose different objects.</p>\n<figure><figcaption><strong>A causal audit combines interpretability methods instead of ranking them</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/frontier-mechanistic-interpretability-0.webp\" width=\"512\" height=\"1188\" alt=\"A causal audit combines interpretability methods instead of ranking them\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"sparse-autoencoders-build-feature-dictionaries\">Sparse autoencoders build feature dictionaries</h2>\n<p>Individual neurons are often polysemantic: one neuron responds to several unrelated patterns. The superposition hypothesis proposes that models encode more features than there are activation dimensions by placing sparse features in overlapping directions.</p>\n<p>A sparse autoencoder learns a dictionary that reconstructs an activation <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>x</mi></mrow></semantics></math> from a sparse code <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>z</mi></mrow></semantics></math>:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>z</mi><mo>=</mo><mi mathvariant=\"normal\">ReLU</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><msub><mi>W</mi><mrow><mi mathvariant=\"normal\">e</mi><mi mathvariant=\"normal\">n</mi><mi mathvariant=\"normal\">c</mi></mrow></msub><mi>x</mi><mo>+</mo><msub><mi>b</mi><mrow><mi mathvariant=\"normal\">e</mi><mi mathvariant=\"normal\">n</mi><mi mathvariant=\"normal\">c</mi></mrow></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mover accent=\"true\"><mi>x</mi><mo>^</mo></mover><mo>=</mo><msub><mi>W</mi><mrow><mi mathvariant=\"normal\">d</mi><mi mathvariant=\"normal\">e</mi><mi mathvariant=\"normal\">c</mi></mrow></msub><mi>z</mi><mo>+</mo><msub><mi>b</mi><mrow><mi mathvariant=\"normal\">d</mi><mi mathvariant=\"normal\">e</mi><mi mathvariant=\"normal\">c</mi></mrow></msub></mrow></semantics></math>\n<p>Training balances reconstruction with sparsity:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"script\">L</mi><mo>=</mo><mo stretchy=\"false\">∥</mo><mi>x</mi><mo>−</mo><mover accent=\"true\"><mi>x</mi><mo>^</mo></mover><msubsup><mo stretchy=\"false\">∥</mo><mn>2</mn><mn>2</mn></msubsup><mo>+</mo><mi>λ</mi><mo stretchy=\"false\">∥</mo><mi>z</mi><msub><mo stretchy=\"false\">∥</mo><mn>1</mn></msub></mrow></semantics></math>\n<p>Anthropic’s early dictionary-learning work extracted more monosemantic features than the source neurons from a small transformer.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup> The later Scaling Monosemanticity project trained sparse autoencoders on Claude 3 Sonnet and demonstrated that the approach could recover interpretable feature directions from a production-scale model.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup></p>\n<p>A feature label usually comes from examples that activate it, tokens it promotes, or an automated explanation. That label remains a hypothesis. Similar-looking features can differ in their downstream effects, and a single feature can have behavior that its top activating examples do not reveal.</p>\n<p>Useful SAE evaluation separates:</p>\n<ul>\n<li>reconstruction error</li>\n<li>sparsity</li>\n<li>feature interpretability</li>\n<li>feature completeness for the target behavior</li>\n<li>causal effect under steering or ablation</li>\n</ul>\n<p>A dictionary can reconstruct activations well while dividing concepts in a way that is awkward for human analysis.</p>\n<h2 id=\"cross-layer-transcoders-approximate-computation\">Cross-layer transcoders approximate computation</h2>\n<p>An SAE describes an activation at one point. A transcoder predicts the output of a model component from its input using sparse features. A cross-layer transcoder extends the idea across several layers, letting features represent computations whose natural span does not fit one layer.</p>\n<p>Anthropic’s circuit-tracing work replaces opaque model components with cross-layer transcoders, then constructs attribution graphs for a specific prompt and output.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup> Nodes are interpretable features. Edges estimate how one feature contributes to another feature or to the selected output logit.</p>\n<p>An attribution graph can reveal several paths that converge on one token:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>prompt tokens</span></span>\n<span class=\"line\"><span>  -&gt; contextual features</span></span>\n<span class=\"line\"><span>  -&gt; intermediate computation features</span></span>\n<span class=\"line\"><span>  -&gt; output-promoting features</span></span>\n<span class=\"line\"><span>  -&gt; selected token</span></span></code></pre>\n<p>The graph is local to a prompt and output. It is not a complete schematic of the model. The replacement model also introduces approximation error, so an attractive graph can omit computation carried by unexplained variance.</p>\n<p>Validation requires interventions. Circuit tracing uses feature ablation, activation patching, and graph-based predictions to test whether a proposed path affects the output as expected.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-3\" id=\"user-content-fnref-3-2\">3</a></sup></p>\n<h2 id=\"attention-requires-feature-level-analysis-too\">Attention requires feature-level analysis too</h2>\n<p>Early attribution graphs focused on multilayer perceptron computation. Attention adds conditional information movement between token positions.</p>\n<p>Anthropic’s 2025 attention work decomposes query-key interactions in terms of feature pairs and integrates them into attribution graphs.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup> The method traces why a destination feature selected a source feature, rather than stopping at the attended position.</p>\n<p>An attention pattern shows routing. It does not identify the operation performed on the routed information. A head attending to a name token could copy the name, suppress it, compare it, or use another feature stored at the same position.</p>\n<p>A stronger analysis identifies:</p>\n<ol>\n<li>the destination feature forming the query</li>\n<li>the source feature forming the key</li>\n<li>the information written through the value and output path</li>\n<li>the downstream features that consume the result</li>\n</ol>\n<p>The explanation becomes mechanistic when it accounts for both routing and content.</p>\n<h2 id=\"natural-language-autoencoders-trade-structure-for-readability\">Natural-language autoencoders trade structure for readability</h2>\n<p>Natural Language Autoencoders place a text bottleneck between an activation and its reconstruction.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup> An activation verbalizer maps the target model’s activation to a natural-language explanation. An activation reconstructor maps that explanation back into activation space. The pair is trained to reconstruct activations.</p>\n<p>The training objective does not directly require explanations to be human-readable or faithful. The text is useful because it must preserve information needed for reconstruction, and the paper evaluates whether that information helps with downstream analysis.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-5\" id=\"user-content-fnref-5-2\">5</a></sup></p>\n<p>NLAs offer several practical advantages:</p>\n<ul>\n<li>explanations can express relations that a short feature label cannot</li>\n<li>auditors can search long trajectories without inspecting every feature</li>\n<li>an edited explanation can be reconstructed into an activation intervention</li>\n<li>the same interface can support hypothesis generation and causal testing</li>\n</ul>\n<p>The method also has a distinct risk: a fluent explanation can look more certain than the evidence supports. Reconstruction proves that the text transmits some activation information. It does not prove that every phrase is a faithful semantic description.</p>\n<p>The NLA paper validates case studies with independent evidence such as prompt variations, training-data inspection, other interpretability methods, and interventions.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-5\" id=\"user-content-fnref-5-3\">5</a></sup> That triangulation should be treated as part of the method, not optional polish.</p>\n<h2 id=\"jacobian-lenses-expose-verbalizable-workspace-content\">Jacobian lenses expose verbalizable workspace content</h2>\n<p>The Jacobian lens asks which internal directions are disposed to affect future verbal output. Anthropic’s J-space work finds that a sparse set of these directions supports report, directed control, multi-step reasoning, and flexible reuse.<sup><a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fn-6\" id=\"user-content-fnref-6\">6</a></sup></p>\n<p>This tool differs from an SAE:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Dictionary source</th><th>Primary object</th></tr></thead><tbody><tr><td>Sparse autoencoder</td><td>Learned reconstruction features</td><td>Sparse activation content</td></tr><tr><td>Cross-layer transcoder</td><td>Learned component approximation</td><td>Feature-to-feature computation</td></tr><tr><td>Natural-language autoencoder</td><td>Learned text bottleneck</td><td>Free-form activation report</td></tr><tr><td>Jacobian lens</td><td>Average causal map to future outputs</td><td>Verbalizable workspace directions</td></tr></tbody></table></div>\n<p>The methods can support one another. SAE features with strong J-space alignment may identify workspace content more precisely than single-token J-lens vectors. J-lens readouts can label what a transcoder feature reads and writes. NLA explanations can propose concepts for targeted feature or lens interventions.</p>\n<h2 id=\"auditing-needs-convergence-across-methods\">Auditing needs convergence across methods</h2>\n<p>No single interpretability result should carry a safety decision by itself. A robust audit progresses from broad detection to narrow causal tests:</p>\n<ol>\n<li>Use behavioral evaluations to find a suspicious trajectory.</li>\n<li>Use SAE features or natural-language explanations to generate hypotheses.</li>\n<li>Use circuit graphs to locate a candidate mechanism.</li>\n<li>Use a J-lens or feature probe to test the mechanism across prompts.</li>\n<li>Ablate, patch, or steer the representation.</li>\n<li>Measure predicted behavior and unrelated capability changes.</li>\n<li>inspect training data or model diffs for an independent explanation.</li>\n</ol>\n<figure><figcaption><strong>Interpretability claims strengthen through predicted interventions</strong> — A readable label begins an audit. It does not finish one.</figcaption><img src=\"https://blog.ecitis.org/feeds/figures/frontier-mechanistic-interpretability-1.webp\" width=\"512\" height=\"665\" alt=\"Interpretability claims strengthen through predicted interventions — A readable label begins an audit. It does not finish one.\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>Agreement across methods is useful only when their errors are not identical. Two tools trained on the same activation data and scored by the same model can share blind spots.</p>\n<h2 id=\"open-problems-at-the-frontier\">Open problems at the frontier</h2>\n<p>The current toolchain still leaves hard gaps.</p>\n<p><strong>Completeness:</strong> A readable circuit can explain part of an output while missing another causal path.</p>\n<p><strong>Scalability:</strong> Long agent trajectories produce many activations and candidate features.</p>\n<p><strong>Relational structure:</strong> Sparse feature lists often omit how concepts bind together.</p>\n<p><strong>Adversarial robustness:</strong> A model trained against a monitor may route cognition through representations the monitor does not capture.</p>\n<p><strong>Faithfulness:</strong> Natural-language explanations and visible reasoning can describe a plausible process without being the cause.</p>\n<p><strong>Automation:</strong> Model-based feature labels and audit agents increase coverage but introduce another learned evaluator.</p>\n<p>Recent work pairs human-readable interpretations with interventions that should change behavior when the interpretation is correct. A mechanistic claim should survive causal tests and predict failures on new inputs.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Trenton Bricken et al., “Towards Monosemanticity: Decomposing Language Models With Dictionary Learning,” Transformer Circuits, 2023. <a href=\"https://transformer-circuits.pub/2023/monosemantic-features/index.html\" rel=\"noopener noreferrer\">https://transformer-circuits.pub/2023/monosemantic-features/index.html</a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Adly Templeton et al., “Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet,” Transformer Circuits, 2024. <a href=\"https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html\" rel=\"noopener noreferrer\">https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html</a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Emmanuel Ameisen et al., “Circuit Tracing: Revealing Computational Graphs in Language Models,” Transformer Circuits, 2025. <a href=\"https://transformer-circuits.pub/2025/attribution-graphs/methods.html\" rel=\"noopener noreferrer\">https://transformer-circuits.pub/2025/attribution-graphs/methods.html</a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-3-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Harish Kamath et al., “Tracing Attention Computation Through Feature Interactions,” Transformer Circuits, 2025. <a href=\"https://transformer-circuits.pub/2025/attention-qk/index.html\" rel=\"noopener noreferrer\">https://transformer-circuits.pub/2025/attention-qk/index.html</a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>Kit Fraser-Taliente et al., “Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations,” Transformer Circuits, 2026. <a href=\"https://transformer-circuits.pub/2026/nla/index.html\" rel=\"noopener noreferrer\">https://transformer-circuits.pub/2026/nla/index.html</a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-5-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-5-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p>Wes Gurnee et al., “Verbalizable Representations Form a Global Workspace in Language Models,” Transformer Circuits, 2026. <a href=\"https://transformer-circuits.pub/2026/workspace/index.html\" rel=\"noopener noreferrer\">https://transformer-circuits.pub/2026/workspace/index.html</a> <a href=\"https://blog.ecitis.org/frontier-mechanistic-interpretability/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/frontier-mechanistic-interpretability/",
            "title": "Frontier Mechanistic Interpretability: From Features to Causal Audits",
            "summary": "Compare sparse feature dictionaries, cross-layer circuit tracing, natural-language activation reports, and causal interventions.",
            "image": "https://blog.ecitis.org/open-graph/frontier-mechanistic-interpretability.png",
            "date_modified": "2026-06-30T00:00:00.000Z",
            "date_published": "2026-06-30T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Mechanistic interpretability",
                "Sparse autoencoders",
                "Circuit tracing",
                "Model auditing"
            ]
        },
        {
            "id": "https://blog.ecitis.org/ai-cluster-networking/",
            "content_html": "<p>Distributed AI performance depends on moving tensors at the time the computation graph needs them. Link rate matters, but topology, collective algorithm, message size, congestion, and synchronization determine how much of that rate becomes useful work.</p>\n<p>NVLink, Infinity Fabric, InfiniBand, and RoCE do not occupy one interchangeable layer. NVLink and accelerator-facing Infinity Fabric build scale-up domains. InfiniBand and RoCE connect nodes across a scale-out fabric. PCIe and network adapters bridge those domains.</p>\n<figure><figcaption><strong>AI cluster communication domains</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/ai-cluster-networking-0.webp\" width=\"512\" height=\"453\" alt=\"AI cluster communication domains\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"start-from-communication-volume\">Start from communication volume</h2>\n<p>Data parallel training commonly uses reduce-scatter followed by all-gather to average gradients. Tensor parallelism adds collectives inside transformer layers. Pipeline parallelism sends activations between stages. Expert parallelism routes token representations to selected experts.</p>\n<p>For a ring all-reduce with <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>p</mi></mrow></semantics></math> participants and a payload of <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>n</mi></mrow></semantics></math> bytes per participant, each participant transfers:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>V</mi><mtext>ring</mtext></msub><mo>=</mo><mn>2</mn><mfrac><mrow><mi>p</mi><mo>−</mo><mn>1</mn></mrow><mi>p</mi></mfrac><mi>n</mi></mrow></semantics></math>\n<p>The formula shows why a bandwidth bottleneck affects large gradient buffers. It does not predict latency by itself. A performance model also needs per-step latency, effective bandwidth, and overlap with compute:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>T</mi><mo>≈</mo><mi>α</mi><msub><mi>N</mi><mtext>steps</mtext></msub><mo>+</mo><mfrac><mi>V</mi><msub><mi>B</mi><mtext>effective</mtext></msub></mfrac></mrow></semantics></math>\n<p>Here, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>α</mi></mrow></semantics></math> is the per-step latency term and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>B</mi><mtext>effective</mtext></msub></mrow></semantics></math> is measured application bandwidth.</p>\n<h2 id=\"communication-domains\">Communication domains</h2>\n<h3 id=\"package-and-board\">Package and board</h3>\n<p>Accelerator packages connect compute dies, memory stacks, and I/O dies through internal fabrics. AMD’s MI300X combines eight accelerator complex dies, eight HBM3 stacks, and four I/O dies through Infinity Fabric. The package provides an aggregate theoretical memory bandwidth of 5.3 TB/s.<sup><a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fn-mi300\" id=\"user-content-fnref-mi300\">1</a></sup></p>\n<p>That figure is HBM bandwidth inside one accelerator package. It is not GPU-to-GPU link bandwidth.</p>\n<h3 id=\"scale-up\">Scale-up</h3>\n<p>A scale-up domain gives accelerators direct peer access and collective paths without crossing a conventional data-center network. NVIDIA NVLink links GPUs to peers or NVSwitch. AMD Infinity Fabric links accelerators inside supported nodes.</p>\n<h3 id=\"scale-out\">Scale-out</h3>\n<p>Scale-out fabrics connect many compute nodes through network interface cards and switches. InfiniBand supplies an integrated RDMA fabric. RoCE carries RDMA transport over Ethernet.</p>\n<p>The boundary can extend across racks on newer platforms, but the diagnostic distinction remains useful: first identify the package, local accelerator fabric, host I/O, and network hop involved.</p>\n<h2 id=\"nvlink-and-nvswitch\">NVLink and NVSwitch</h2>\n<p>NVLink is NVIDIA’s high-bandwidth GPU interconnect. NVSwitch provides switched connectivity so a GPU can reach peers through an NVLink fabric rather than a point-to-point-only layout.</p>\n<p>Fifth-generation NVLink in Blackwell systems provides up to 1,800 GB/s of aggregate bandwidth per GPU. NVIDIA’s GB300 NVL72 reference architecture uses nine NVLink switch trays, with two NVSwitch ASICs per tray, to connect 72 GPUs in one rack-scale domain.<sup><a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fn-nvlink\" id=\"user-content-fnref-nvlink\">2</a></sup></p>\n<p>Those specifications describe a particular platform generation and topology. Do not apply them to an H100 server or a PCIe-only deployment.</p>\n<p>NVLink is useful for:</p>\n<ul>\n<li>Tensor-parallel collectives that occur within layers.</li>\n<li>Peer reads and writes between accelerator memories.</li>\n<li>Unified scale-up domains managed by the platform fabric.</li>\n<li>Collective algorithms optimized for local topology.</li>\n</ul>\n<p>NVSwitch does not make every pairwise transfer free. Collective libraries still choose channels, chunk sizes, trees, or rings. Contention appears when concurrent jobs or collective phases share fabric paths.</p>\n<h2 id=\"amd-infinity-fabric\">AMD Infinity Fabric</h2>\n<p>Infinity Fabric is AMD’s interconnect family across package and node scopes. On MI300X, it connects package components and supports an eight-accelerator node in which each accelerator has seven high-bandwidth Infinity Fabric links, forming a fully connected topology.<sup><a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fn-mi300\" id=\"user-content-fnref-mi300-2\">1</a></sup></p>\n<p>The topology reduces the need for intermediate accelerator hops inside that node. Collective placement still matters:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>rank → accelerator mapping → Infinity Fabric peer path → NIC affinity</span></span></code></pre>\n<p>If the network interface is closest to one CPU socket or accelerator group, a poor rank map can send scale-out traffic across additional host or fabric links.</p>\n<p>Inspect the platform topology rather than assuming device numbering follows physical adjacency. Use the communication library’s topology dump and measure peer bandwidth for the installed firmware, driver, and link state.</p>\n<h2 id=\"infiniband\">InfiniBand</h2>\n<p>InfiniBand defines an I/O architecture for servers, storage, and communication infrastructure. It supports send and receive messaging plus RDMA memory operations without software involvement in the data movement path.<sup><a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fn-ibta\" id=\"user-content-fnref-ibta\">3</a></sup></p>\n<p>An InfiniBand network contains:</p>\n<ul>\n<li>Host channel adapters.</li>\n<li>Switches and links.</li>\n<li>A subnet manager.</li>\n<li>Queue pairs and completion queues.</li>\n<li>Memory regions registered for RDMA.</li>\n<li>Routing and service-level configuration.</li>\n</ul>\n<p>NVIDIA’s QM9700 and QM9790 NDR switches expose 64 ports at 400 Gb/s per port in a one-rack-unit chassis.<sup><a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fn-ndr\" id=\"user-content-fnref-ndr\">4</a></sup> Port rate is not collective bandwidth. Encoding, protocol overhead, PCIe, host memory registration, message size, and fabric contention all affect delivered throughput.</p>\n<p>InfiniBand fabrics can supply adaptive routing, virtual lanes, and in-network collective operations on supported platforms. These features need topology-aware configuration and telemetry. A nominally nonblocking topology can still hotspot if routing sends synchronized flows through the same spine links.</p>\n<h2 id=\"roce\">RoCE</h2>\n<p>RDMA over Converged Ethernet uses RDMA semantics on Ethernet infrastructure. RoCE version one operates at the link layer and remains within an Ethernet broadcast domain. RoCE version two carries the InfiniBand transport protocol over UDP and IP, so it can cross routed networks.<sup><a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fn-roce\" id=\"user-content-fnref-roce\">5</a></sup></p>\n<p>RoCE makes Ethernet a viable AI transport, but Ethernet loss and congestion behavior must be engineered for synchronized, high-rate flows.</p>\n<p>Two controls appear often:</p>\n<p><strong>Priority Flow Control:</strong> pauses one Ethernet priority rather than the whole link. NVIDIA’s switch documentation describes eight virtual links that can be paused independently.<sup><a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fn-pfc\" id=\"user-content-fnref-pfc\">6</a></sup></p>\n<p><strong>Explicit Congestion Notification:</strong> marks packets instead of dropping them when queues cross configured thresholds. A RoCEv2 receiver generates a congestion notification packet toward the sender, which reduces its sending rate.<sup><a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fn-roce\" id=\"user-content-fnref-roce-2\">5</a></sup></p>\n<p>PFC and ECN solve different problems. PFC prevents loss for a traffic class through hop-by-hop pause. ECN provides end-to-end congestion signaling. Poor PFC design can propagate pauses and create head-of-line blocking. Poor ECN thresholds can react too late or throttle too aggressively.</p>\n<h2 id=\"infiniband-and-roce-selection\">InfiniBand and RoCE selection</h2>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Question</th><th>InfiniBand</th><th>RoCEv2</th></tr></thead><tbody><tr><td>Link and transport</td><td>Purpose-built InfiniBand stack</td><td>RDMA transport over routed Ethernet</td></tr><tr><td>Fabric operations</td><td>Subnet management and InfiniBand routing</td><td>Ethernet routing, VLAN, QoS, ECN, and optional PFC</td></tr><tr><td>Operational reuse</td><td>Separate fabric skill set</td><td>Can share Ethernet platform and practices</td></tr><tr><td>Failure focus</td><td>Routing, link state, credit and service-level behavior</td><td>Queueing, ECN, PFC, loss, and Ethernet configuration</td></tr></tbody></table></div>\n<p>The choice is operational as much as architectural. A well-instrumented RoCE fabric can outperform a poorly operated InfiniBand fabric for a workload, and the reverse is also true. Evaluate delivered collective performance and recovery behavior.</p>\n<h2 id=\"topology-determines-collective-cost\">Topology determines collective cost</h2>\n<p>A node can have fast local peer links and fewer network interfaces than accelerators. Scale-out traffic must then reach a NIC through a local path.</p>\n<p>Consider a topology with accelerator groups attached to different CPU sockets:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>GPU group A ─ local fabric ─ NIC A ─ leaf switch</span></span>\n<span class=\"line\"><span>      │</span></span>\n<span class=\"line\"><span>  host interconnect</span></span>\n<span class=\"line\"><span>      │</span></span>\n<span class=\"line\"><span>GPU group B ─ local fabric ─ NIC B ─ leaf switch</span></span></code></pre>\n<p>Bind processes so each rank uses the nearest accelerator and NIC. Crossing the host interconnect for every collective adds a bottleneck not visible in the network port rate.</p>\n<p>Hierarchical collectives exploit this structure:</p>\n<ol>\n<li>Reduce within the local scale-up domain.</li>\n<li>Exchange reduced chunks across the scale-out fabric.</li>\n<li>Broadcast or gather within each local domain.</li>\n</ol>\n<p>The best algorithm depends on payload size and topology. Trees reduce startup steps for some small messages. Rings make predictable use of bandwidth for large messages. Communication libraries select algorithms dynamically, but their decision needs accurate topology discovery.</p>\n<h2 id=\"congestion-patterns-in-ai-workloads\">Congestion patterns in AI workloads</h2>\n<p>AI communication is synchronized. Many ranks enter the same collective near the same time, creating burst patterns unlike independent storage flows.</p>\n<p>Common problems:</p>\n<ul>\n<li><strong>Incast:</strong> many senders target one receiver or reduction point.</li>\n<li><strong>All-to-all pressure:</strong> expert routing spreads traffic across many destinations.</li>\n<li><strong>Rail imbalance:</strong> one NIC or spine path carries more ranks.</li>\n<li><strong>Pause propagation:</strong> PFC on a congested queue blocks unrelated flows in the same class.</li>\n<li><strong>Collective interference:</strong> overlapping jobs share links without isolation.</li>\n<li><strong>Stragglers:</strong> one slow link delays every rank at the synchronization point.</li>\n</ul>\n<p>Average link utilization can look low while short queue spikes dominate step time. Collect high-resolution queue, ECN, PFC, retransmission, and link-error counters.</p>\n<h2 id=\"benchmark-by-layer\">Benchmark by layer</h2>\n<p>Use a ladder:</p>\n<ol>\n<li>Peer memory bandwidth inside one accelerator.</li>\n<li>Accelerator-to-accelerator bandwidth inside one node.</li>\n<li>Host-to-device and device-to-NIC paths.</li>\n<li>RDMA read, write, send, and receive between two nodes.</li>\n<li>Collective operations at several message sizes.</li>\n<li>Application step time with communication traces.</li>\n</ol>\n<p>A point-to-point benchmark does not validate all-reduce. An all-reduce benchmark does not validate expert all-to-all.</p>\n<p>Report topology and software with results:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>accelerators_per_node</span><span>:</span><span> 8</span></span>\n<span class=\"line\"><span>nics_per_node</span><span>:</span><span> 8</span></span>\n<span class=\"line\"><span>fabric</span><span>:</span><span> 'RoCEv2'</span></span>\n<span class=\"line\"><span>collective</span><span>:</span><span> 'all_reduce'</span></span>\n<span class=\"line\"><span>payload_bytes</span><span>:</span><span> '&lt;measured case&gt;'</span></span>\n<span class=\"line\"><span>driver_revision</span><span>:</span><span> '&lt;pinned revision&gt;'</span></span>\n<span class=\"line\"><span>firmware_revision</span><span>:</span><span> '&lt;pinned revision&gt;'</span></span></code></pre>\n<p>The values in this configuration are an example schema, not a recommended cluster design.</p>\n<h2 id=\"diagnose-slow-collectives\">Diagnose slow collectives</h2>\n<p><strong>One node is slow:</strong> check accelerator link state, NIC affinity, PCIe width, firmware, thermal state, and error counters.</p>\n<p><strong>Performance drops only across racks:</strong> inspect oversubscription, routing entropy, spine utilization, and optical errors.</p>\n<p><strong>Large messages are slow:</strong> inspect effective bandwidth, rail balance, and algorithm selection.</p>\n<p><strong>Small messages are slow:</strong> inspect launch, synchronization, and per-step latency.</p>\n<p><strong>RoCE stalls under load:</strong> inspect PFC pause frames, ECN marks, congestion notification packets, queue occupancy, and packet loss.</p>\n<p><strong>Runs vary despite stable bandwidth:</strong> inspect competing jobs, adaptive routing, CPU scheduling, and collective synchronization.</p>\n<h2 id=\"capacity-planning\">Capacity planning</h2>\n<p>Model the communication phase from workload bytes, not advertised FLOPS:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>required fabric time</span></span>\n<span class=\"line\"><span>  = bytes transferred per step</span></span>\n<span class=\"line\"><span>  / measured effective bandwidth</span></span></code></pre>\n<p>Then compare it with the compute phase and determine which communication can overlap. If an exposed collective occupies the critical path, faster compute makes the network fraction larger.</p>\n<p>Include failure capacity. A fabric that meets the target only with every link healthy has no room for maintenance, rerouting, or a failed rail.</p>\n<p>Cluster networking is a hierarchy. Preserve that hierarchy in topology maps, benchmarks, and alerts. A single “network bandwidth” number erases the exact boundary that must be fixed.</p>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-mi300\">\n<p><a href=\"https://rocm.docs.amd.com/en/docs-6.2.0/conceptual/gpu-arch/mi300.html\" rel=\"noopener noreferrer\">AMD ROCm documentation, “AMD Instinct MI300 series microarchitecture”</a>. <a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fnref-mi300\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fnref-mi300-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-nvlink\">\n<p><a href=\"https://docs.nvidia.com/enterprise-reference-architectures/nvl72-ai-factory/latest/components.html\" rel=\"noopener noreferrer\">NVIDIA Enterprise Reference Architecture, “NVL72 AI Factory System Hardware and Components”</a>. <a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fnref-nvlink\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-ibta\">\n<p><a href=\"https://www.infinibandta.org/about-infiniband/\" rel=\"noopener noreferrer\">InfiniBand Trade Association, “About InfiniBand”</a>. <a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fnref-ibta\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-ndr\">\n<p><a href=\"https://docs.nvidia.com/networking/display/qm9700-and-qm9790-1u-ndr-400gbps-infiniband-switch-systems-user-manual.pdf\" rel=\"noopener noreferrer\">NVIDIA, “QM9700 and QM9790 1U NDR 400Gb/s InfiniBand Switch Systems User Manual”</a>. <a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fnref-ndr\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-roce\">\n<p><a href=\"https://docs.nvidia.com/networking-ethernet-software/cumulus-linux-43/Network-Solutions/RDMA-over-Converged-Ethernet-RoCE/\" rel=\"noopener noreferrer\">NVIDIA Cumulus Linux documentation, “RDMA over Converged Ethernet”</a>. <a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fnref-roce\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fnref-roce-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-pfc\">\n<p><a href=\"https://docs.nvidia.com/doca/archive/3-2-2/Flow-Control/index.html\" rel=\"noopener noreferrer\">NVIDIA DOCA documentation, “Flow Control”</a>. <a href=\"https://blog.ecitis.org/ai-cluster-networking/#user-content-fnref-pfc\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/ai-cluster-networking/",
            "title": "AI Cluster Networking: NVLink, Infinity Fabric, InfiniBand, and RoCE",
            "summary": "Map scale-up and scale-out AI fabrics across NVLink, Infinity Fabric, InfiniBand, RoCE, topology, and collective communication.",
            "image": "https://blog.ecitis.org/open-graph/ai-cluster-networking.png",
            "date_modified": "2026-06-26T00:00:00.000Z",
            "date_published": "2026-06-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Hardware & Systems",
                "Cluster networking",
                "NVLink",
                "InfiniBand"
            ]
        },
        {
            "id": "https://blog.ecitis.org/diffusion-flow-matching/",
            "content_html": "<p>Diffusion and flow matching define learning problems over paths between a tractable noise distribution and the data distribution. The neural network architecture can be shared, but the target, time parameterization, and sampler determine what the network predicts and how generation integrates it.</p>\n<figure><figcaption><strong>Two training targets over a noise-to-data path</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/diffusion-flow-matching-0.webp\" width=\"512\" height=\"289\" alt=\"Two training targets over a noise-to-data path\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"diffusion-learns-a-reverse-corruption-process\">Diffusion learns a reverse corruption process</h2>\n<p>A forward diffusion process gradually corrupts data <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>x</mi><mn>0</mn></msub></mrow></semantics></math> into <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>x</mi><mi>t</mi></msub></mrow></semantics></math>. A common Gaussian parameterization is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>x</mi><mi>t</mi></msub><mo>=</mo><msub><mi>α</mi><mi>t</mi></msub><msub><mi>x</mi><mn>0</mn></msub><mo>+</mo><msub><mi>σ</mi><mi>t</mi></msub><mi>ϵ</mi><mo separator=\"true\">,</mo><mspace width=\"2em\"></mspace><mi>ϵ</mi><mo>∼</mo><mi mathvariant=\"script\">N</mi><mo stretchy=\"false\">(</mo><mn>0</mn><mo separator=\"true\">,</mo><mi>I</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>Training samples a data point, a time <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>t</mi></mrow></semantics></math>, and noise <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>ϵ</mi></mrow></semantics></math>. A network can predict <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>ϵ</mi></mrow></semantics></math>, the score <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi mathvariant=\"normal\">∇</mi><msub><mi>x</mi><mi>t</mi></msub></msub><mi>log</mi><mo>⁡</mo><msub><mi>p</mi><mi>t</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>x</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math>, the clean sample, or a velocity parameterization derived from them. Denoising Diffusion Probabilistic Models connected this training objective to denoising score matching and a learned reverse Markov chain.<sup><a href=\"https://blog.ecitis.org/diffusion-flow-matching/#user-content-fn-ddpm\" id=\"user-content-fnref-ddpm\">1</a></sup></p>\n<p>The noise schedule controls signal-to-noise ratio across time. Loss weighting determines which regions of that schedule dominate updates. These choices affect optimization even when the architecture is unchanged.</p>\n<p>At inference, generation starts from noise and applies a numerical solver. A stochastic reverse SDE injects randomness along the path. A probability-flow ODE follows a deterministic trajectory for a fixed initial noise. Sampler order and step placement trade network evaluations, truncation error, and stability.</p>\n<h2 id=\"flow-matching-regresses-a-vector-field\">Flow matching regresses a vector field</h2>\n<p>Flow matching chooses a probability path <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>p</mi><mi>t</mi></msub></mrow></semantics></math> and trains a vector field <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>v</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>t</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> to transport samples along it:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mfrac><mrow><mi>d</mi><mi>x</mi></mrow><mrow><mi>d</mi><mi>t</mi></mrow></mfrac><mo>=</mo><msub><mi>v</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>t</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>The Flow Matching paper showed that conditional vector fields can provide a simulation-free training objective and include diffusion paths as a special case.<sup><a href=\"https://blog.ecitis.org/diffusion-flow-matching/#user-content-fn-flow-matching\" id=\"user-content-fnref-flow-matching\">2</a></sup> Training does not require solving the ODE for every update. It samples points on a conditional path and regresses the target velocity.</p>\n<p>For a simple linear interpolation between noise <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>x</mi><mn>0</mn></msub></mrow></semantics></math> and data <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>x</mi><mn>1</mn></msub></mrow></semantics></math>:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>x</mi><mi>t</mi></msub><mo>=</mo><mo stretchy=\"false\">(</mo><mn>1</mn><mo>−</mo><mi>t</mi><mo stretchy=\"false\">)</mo><msub><mi>x</mi><mn>0</mn></msub><mo>+</mo><mi>t</mi><msub><mi>x</mi><mn>1</mn></msub></mrow></semantics></math>\n<p>the conditional velocity is <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>x</mi><mn>1</mn></msub><mo>−</mo><msub><mi>x</mi><mn>0</mn></msub></mrow></semantics></math>. In practice, the coupling between noise and data and the selected path affect trajectory geometry. A straight conditional interpolation does not imply that the marginal vector field is trivial.</p>\n<h2 id=\"prediction-targets-are-coordinate-choices\">Prediction targets are coordinate choices</h2>\n<p>Noise, clean-sample, score, and velocity prediction can often be converted when <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>α</mi><mi>t</mi></msub></mrow></semantics></math> and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>σ</mi><mi>t</mi></msub></mrow></semantics></math> are known. They do not produce identical numerical conditioning under finite precision, imperfect networks, or weighted losses.</p>\n<p>Use one canonical internal definition:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> interpolate</span><span>(</span><span>clean</span><span>,</span><span> noise</span><span>,</span><span> alpha_t</span><span>,</span><span> sigma_t</span><span>):</span></span>\n<span class=\"line\"><span>    sample </span><span>=</span><span> alpha_t </span><span>*</span><span> clean </span><span>+</span><span> sigma_t </span><span>*</span><span> noise</span></span>\n<span class=\"line\"><span>    velocity </span><span>=</span><span> alpha_t </span><span>*</span><span> noise </span><span>-</span><span> sigma_t </span><span>*</span><span> clean</span></span>\n<span class=\"line\"><span>    return</span><span> sample</span><span>,</span><span> velocity</span></span></code></pre>\n<p>The exact velocity formula depends on the schedule convention. Tests should verify conversions at schedule endpoints and intermediate times.</p>\n<h2 id=\"u-nets-and-transformers-solve-the-same-field-problem\">U-Nets and transformers solve the same field problem</h2>\n<p>A U-Net uses multiscale convolutional features, skip connections, and attention blocks. It matches image locality and provides feature maps at several resolutions. A diffusion transformer tokenizes latent patches and applies transformer blocks conditioned on time and text. Diffusion Transformers demonstrated that transformer compute can replace the conventional U-Net backbone in latent diffusion.<sup><a href=\"https://blog.ecitis.org/diffusion-flow-matching/#user-content-fn-dit\" id=\"user-content-fnref-dit\">3</a></sup></p>\n<p>Architecture selection changes:</p>\n<ul>\n<li>how spatial resolution maps to compute;</li>\n<li>where text conditioning enters;</li>\n<li>whether variable aspect ratios require padding or packed tokens;</li>\n<li>how temporal features are introduced for video;</li>\n<li>which kernels dominate training and inference.</li>\n</ul>\n<p>The network output still represents the selected denoising or flow target.</p>\n<h2 id=\"latent-models-move-the-path-out-of-pixel-space\">Latent models move the path out of pixel space</h2>\n<p>Pixel-space generation carries high-dimensional spatial tensors through every network evaluation. Latent diffusion uses an encoder to map images into a lower-dimensional representation, trains the generative process in that space, then decodes the result.<sup><a href=\"https://blog.ecitis.org/diffusion-flow-matching/#user-content-fn-latent-diffusion\" id=\"user-content-fnref-latent-diffusion\">4</a></sup></p>\n<p>This introduces a separate reconstruction contract. Fine texture, text, faces, and high-frequency patterns can be limited by the autoencoder even if the denoiser is accurate. Evaluate reconstruction error independently by encoding and decoding real samples without generation.</p>\n<p>Latent scaling must match training. A missing scale factor changes the effective noise distribution and can produce washed-out or saturated samples.</p>\n<h2 id=\"conditioning-enters-through-several-paths\">Conditioning enters through several paths</h2>\n<p>Text conditioning commonly uses cross-attention from image or video tokens to text-encoder outputs. Time is embedded and injected through adaptive normalization or residual modulation. Class labels, camera parameters, masks, depth, or motion controls can use related adapters.</p>\n<p>Classifier-free guidance combines conditional and unconditional predictions:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mover accent=\"true\"><mi>u</mi><mo>^</mo></mover><mo>=</mo><msub><mi>u</mi><mi mathvariant=\"normal\">∅</mi></msub><mo>+</mo><mi>w</mi><mo stretchy=\"false\">(</mo><msub><mi>u</mi><mi>c</mi></msub><mo>−</mo><msub><mi>u</mi><mi mathvariant=\"normal\">∅</mi></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>where <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>u</mi></mrow></semantics></math> denotes the chosen prediction target and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>w</mi></mrow></semantics></math> is guidance strength. Guidance changes the vector field followed by the sampler. High guidance can improve prompt alignment while reducing diversity or causing saturation; the acceptable range is model- and sampler-specific.</p>\n<p>Batching the conditional and unconditional passes can increase throughput but doubles some activation and conditioning work. Cache text embeddings when prompts repeat.</p>\n<h2 id=\"video-adds-a-temporal-allocation-problem\">Video adds a temporal allocation problem</h2>\n<p>Video tensors add frame count to spatial resolution, channel width, and batch size. Full spatiotemporal attention grows with the square of the token count. Architectures therefore factor spatial and temporal attention, use local windows, compress time in a latent autoencoder, or process temporal chunks.</p>\n<p>Temporal consistency is not guaranteed by generating each frame with the same prompt. The model needs temporal operators or shared latent structure. Conditioning should preserve frame order, frame rate, aspect ratio, and camera metadata used during training.</p>\n<p>A serving plan must choose:</p>\n<ul>\n<li>frame count and resolution buckets;</li>\n<li>latent temporal compression;</li>\n<li>attention window or context;</li>\n<li>whether decoding is tiled;</li>\n<li>where intermediate latents are stored;</li>\n<li>how previews are generated without changing the final path.</li>\n</ul>\n<p>Tiled decoding lowers peak memory but can expose seams if receptive-field overlap and blending do not match the decoder.</p>\n<h2 id=\"solvers-need-model-specific-validation\">Solvers need model-specific validation</h2>\n<p>An ODE solver advances:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>x</mi><mrow><mi>t</mi><mo>+</mo><mi mathvariant=\"normal\">Δ</mi><mi>t</mi></mrow></msub><mo>≈</mo><msub><mi>x</mi><mi>t</mi></msub><mo>+</mo><mi mathvariant=\"normal\">Δ</mi><mi>t</mi><mtext> </mtext><msub><mi>v</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>x</mi><mi>t</mi></msub><mo separator=\"true\">,</mo><mi>t</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>Higher-order methods evaluate the network at additional points to reduce integration error. Fewer steps reduce latency but can miss curved regions of the learned field. A solver qualified on one parameterization or schedule may be unstable on another.</p>\n<p>Benchmark quality against network evaluation count, not loop iteration alone. Some iterations call the model more than once. Also report text-encoder and decoder time so denoiser optimizations are not overstated.</p>\n<figure><figcaption><strong>A solver follows a learned field through discrete evaluations</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/diffusion-flow-matching-1.webp\" width=\"512\" height=\"344\" alt=\"A solver follows a learned field through discrete evaluations\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"training-and-serving-failure-modes\">Training and serving failure modes</h2>\n<p><strong>Time convention mismatch:</strong> Training uses noise-to-data time while the sampler assumes data-to-noise time. Endpoint tests catch the sign and schedule reversal.</p>\n<p><strong>Target conversion mismatch:</strong> The model predicts velocity but guidance or solver code treats it as noise. Centralize conversions and record the checkpoint prediction type.</p>\n<p><strong>Latent scale mismatch:</strong> Encoder outputs are scaled during training but not inference. Compare training-pipeline latents with serving latents.</p>\n<p><strong>Conditioning dropout mismatch:</strong> Classifier-free guidance requires an unconditional path trained with conditioning dropout. Replacing conditioning with arbitrary empty text can differ from the trained null condition.</p>\n<p><strong>Video padding leakage:</strong> Padded frames participate in attention because the mask is wrong. Test variable-length batches against individually evaluated clips.</p>\n<p><strong>Sampler benchmark bias:</strong> Quality is compared at equal step count despite different model evaluations. Compare equal evaluation budgets and wall-clock latency.</p>\n<p>Diffusion and flow matching are not competing network brands. They are ways to specify a time-dependent learning target and a generative path. A production system must keep schedule, target parameterization, architecture, conditioning, latent codec, and sampler in one versioned contract.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-ddpm\">\n<p><a href=\"https://arxiv.org/abs/2006.11239\" rel=\"noopener noreferrer\">Denoising Diffusion Probabilistic Models</a>. <a href=\"https://blog.ecitis.org/diffusion-flow-matching/#user-content-fnref-ddpm\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-flow-matching\">\n<p><a href=\"https://arxiv.org/abs/2210.02747\" rel=\"noopener noreferrer\">Flow Matching for Generative Modeling</a>. <a href=\"https://blog.ecitis.org/diffusion-flow-matching/#user-content-fnref-flow-matching\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-dit\">\n<p><a href=\"https://arxiv.org/abs/2212.09748\" rel=\"noopener noreferrer\">Scalable Diffusion Models with Transformers</a>. <a href=\"https://blog.ecitis.org/diffusion-flow-matching/#user-content-fnref-dit\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-latent-diffusion\">\n<p><a href=\"https://arxiv.org/abs/2112.10752\" rel=\"noopener noreferrer\">High-Resolution Image Synthesis with Latent Diffusion Models</a>. <a href=\"https://blog.ecitis.org/diffusion-flow-matching/#user-content-fnref-latent-diffusion\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/diffusion-flow-matching/",
            "title": "Diffusion and Flow Matching: Architectures for Image and Video Generation",
            "summary": "Compare diffusion and flow-matching objectives, time parameterizations, architectures, and samplers for image and video generation.",
            "image": "https://blog.ecitis.org/open-graph/diffusion-flow-matching.png",
            "date_modified": "2026-06-19T00:00:00.000Z",
            "date_published": "2026-06-19T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Diffusion",
                "Flow matching",
                "Generative media"
            ]
        },
        {
            "id": "https://blog.ecitis.org/ai-safety-systems/",
            "content_html": "<p>Production AI safety is a systems property. A model can refuse a prohibited request and still leak data through retrieval, call an overpowered tool, or emit malformed output that an application executes.</p>\n<p>The control plane must govern identity, data access, tool authority, output contracts, monitoring, and response. Model behavior remains one control inside that plane.</p>\n<h2 id=\"model-the-system-before-the-threat\">Model the system before the threat</h2>\n<p>Draw the actual data path:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>user → gateway → context builder → model → tool broker → external systems</span></span>\n<span class=\"line\"><span>                         ↓                       ↓</span></span>\n<span class=\"line\"><span>                    retrieval index         audit log</span></span></code></pre>\n<p>For each boundary, record:</p>\n<ul>\n<li>Principal and credential.</li>\n<li>Data classification.</li>\n<li>Allowed operations.</li>\n<li>Policy owner.</li>\n<li>Validation performed.</li>\n<li>Log and retention behavior.</li>\n<li>Failure and rollback path.</li>\n</ul>\n<p>NIST’s Generative AI Profile organizes risk work around govern, map, measure, and manage functions and treats risk management as a lifecycle activity rather than a model-only test.<sup><a href=\"https://blog.ecitis.org/ai-safety-systems/#user-content-fn-nist\" id=\"user-content-fnref-nist\">1</a></sup></p>\n<figure><figcaption><strong>Safety control plane around a model request</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/ai-safety-systems-0.webp\" width=\"512\" height=\"410\" alt=\"Safety control plane around a model request\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"separate-policy-from-prompt-text\">Separate policy from prompt text</h2>\n<p>A system prompt is input to a probabilistic model. It can influence behavior but cannot enforce database permissions or network access.</p>\n<p>Represent enforceable policy in code or a policy engine:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> Action</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  principal</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  operation</span><span>:</span><span> 'read'</span><span> |</span><span> 'create'</span><span> |</span><span> 'update'</span><span> |</span><span> 'delete'</span></span>\n<span class=\"line\"><span>  resource</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  attributes</span><span>:</span><span> Record</span><span>&lt;</span><span>string</span><span>,</span><span> string</span><span>&gt;</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>type</span><span> PolicyDecision</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  effect</span><span>:</span><span> 'allow'</span><span> |</span><span> 'deny'</span><span> |</span><span> 'require_approval'</span></span>\n<span class=\"line\"><span>  policyVersion</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  reason</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The prompt can tell the model not to delete records. The tool broker must deny deletion unless the authenticated principal, resource policy, and conversation approval allow it.</p>\n<p>Version policies independently from prompts. An incident review should identify the rule that allowed an action without reconstructing it from prose embedded in a model request.</p>\n<h2 id=\"apply-least-privilege-to-tools\">Apply least privilege to tools</h2>\n<p>Tool definitions should map to narrow operations. A generic SQL executor or shell command carries more authority than a tool designed around one domain action.</p>\n<p>Prefer:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"name\"</span><span>:</span><span> \"create_refund_request\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"input\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"invoice_id\"</span><span>:</span><span> \"INV-1042\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"reason_code\"</span><span>:</span><span> \"DUPLICATE_CHARGE\"</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>over:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"name\"</span><span>:</span><span> \"run_command\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"input\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"command\"</span><span>:</span><span> \"refund --invoice INV-1042\"</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The values are illustrative. The narrow tool allows schema validation, resource-level authorization, idempotency, and a specific audit event.</p>\n<p>Credentials should be scoped per tool service and deployment environment. Do not place a broad production credential in the model context. The model should select an operation; the broker should attach credentials after policy approval.</p>\n<h2 id=\"bind-approval-to-material-arguments\">Bind approval to material arguments</h2>\n<p>Approval must cover the action the system will execute. If the model changes a recipient, amount, repository, or destination after approval, request approval again.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>function</span><span> approvalFingerprint</span><span>(tool</span><span>:</span><span> string</span><span>,</span><span> args</span><span>:</span><span> unknown</span><span>,</span><span> principal</span><span>:</span><span> string</span><span>)</span><span>:</span><span> string</span><span> {</span></span>\n<span class=\"line\"><span>  return</span><span> stableHash</span><span>({ tool</span><span>,</span><span> args</span><span>,</span><span> principal })</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The approval record should contain the fingerprint, policy version, user identity, timestamp, and rendered action summary. A generic session-wide approval grants authority to future model outputs that the user has not seen.</p>\n<p>Read operations can also be sensitive. Searching personnel records or retrieving private source code requires access checks even when no state changes.</p>\n<figure><figcaption><strong>Approval binds to the exact action arguments</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/ai-safety-systems-1.webp\" width=\"512\" height=\"286\" alt=\"Approval binds to the exact action arguments\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"treat-retrieved-content-as-untrusted\">Treat retrieved content as untrusted</h2>\n<p>Prompt injection crosses a trust boundary when data is interpreted as instructions. Web pages, emails, documents, issue comments, and tool results can contain text that tells the model to ignore policy or exfiltrate context.</p>\n<p>Do not rely on a delimiter to make instructions harmless. The model still reads both sides of the delimiter.</p>\n<p>Controls include:</p>\n<ul>\n<li>Retrieve only from authorized sources.</li>\n<li>Preserve source identity around every passage.</li>\n<li>Label retrieved text as data in the model instruction.</li>\n<li>Exclude credentials and unnecessary private context.</li>\n<li>Restrict tool authority independently of model output.</li>\n<li>Require confirmation for sensitive side effects.</li>\n<li>Validate output before execution.</li>\n</ul>\n<p>Indirect injection becomes an incident only when the surrounding system gives the injected instruction a path to data or action. Remove that path with least privilege and approval.</p>\n<h2 id=\"constrain-data-movement\">Constrain data movement</h2>\n<p>Build context from an explicit allowlist of fields. Passing complete records creates accidental exposure and makes later redaction harder.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> TicketContext</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  subject</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  customerMessage</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  publicProduct</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  allowedPolicyExcerpts</span><span>:</span><span> string</span><span>[]</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Keep secrets out of prompts. If a tool needs an API credential, inject it inside the trusted tool process. Model-visible logs should redact authorization headers, session cookies, private keys, and personal data not needed for the task.</p>\n<p>Tenant isolation belongs in the query and storage layer. Filtering another tenant’s passages after retrieval already exposes them to rankers, caches, and traces.</p>\n<h2 id=\"validate-inputs-and-outputs\">Validate inputs and outputs</h2>\n<p>Input validation protects services from malformed model-generated arguments. Output validation protects downstream consumers from malformed or policy-breaking model text.</p>\n<p>Use a layered contract:</p>\n<ol>\n<li>Parse the output.</li>\n<li>Validate its schema.</li>\n<li>Check domain invariants.</li>\n<li>Apply authorization.</li>\n<li>Execute through an idempotent boundary.</li>\n<li>Verify postconditions.</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> RefundRequest</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  invoiceId</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  reasonCode</span><span>:</span><span> 'DUPLICATE_CHARGE'</span><span> |</span><span> 'SERVICE_NOT_DELIVERED'</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>function</span><span> validateRefund</span><span>(request</span><span>:</span><span> RefundRequest</span><span>,</span><span> invoice</span><span>:</span><span> { id</span><span>:</span><span> string</span><span>; state</span><span>:</span><span> string</span><span> })</span><span>:</span><span> void</span><span> {</span></span>\n<span class=\"line\"><span>  if</span><span> (</span><span>request</span><span>.invoiceId </span><span>!==</span><span> invoice</span><span>.id) </span><span>throw</span><span> new</span><span> Error</span><span>(</span><span>'invoice mismatch'</span><span>)</span></span>\n<span class=\"line\"><span>  if</span><span> (</span><span>invoice</span><span>.state </span><span>!==</span><span> 'paid'</span><span>) </span><span>throw</span><span> new</span><span> Error</span><span>(</span><span>'invoice is not refundable'</span><span>)</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>A JSON Schema can prove that <code>invoiceId</code> is a string. It cannot prove that the invoice belongs to the principal or is refundable.</p>\n<p>For generated code, parse before execution, run in an isolated environment, constrain network and filesystem access, set resource limits, and inspect artifacts. Text moderation does not replace a sandbox.</p>\n<h2 id=\"use-classifiers-as-routing-controls\">Use classifiers as routing controls</h2>\n<p>Classifiers can detect content categories, sensitive data, or policy-relevant intent. They should produce an operational route:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>allow → normal pipeline</span></span>\n<span class=\"line\"><span>transform → redact or narrow context</span></span>\n<span class=\"line\"><span>review → human queue</span></span>\n<span class=\"line\"><span>block → safe response with reason code</span></span></code></pre>\n<p>Measure false positives and false negatives on the deployed distribution. A classifier threshold is a policy choice with user and incident costs on both sides.</p>\n<p>Do not let a classifier silently rewrite authority. A low-risk classification can reduce review but should not grant a tool permission that the principal lacks.</p>\n<h2 id=\"make-provenance-inspectable\">Make provenance inspectable</h2>\n<p>For factual answers, preserve the link between claims and sources:</p>\n<ul>\n<li>Source identifier and revision.</li>\n<li>Retrieved passage offsets.</li>\n<li>Rank and selection trace.</li>\n<li>Claim-to-source association where required.</li>\n<li>Time of retrieval.</li>\n</ul>\n<p>Citation presence alone is not grounding. Validate that the cited passage supports the claim and that the user can access the source.</p>\n<p>For actions, provenance means the chain from user request through model proposal, policy decision, approval, tool call, and verified result.</p>\n<h2 id=\"design-failure-states\">Design failure states</h2>\n<p>Safe failure is explicit:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Failure</th><th>System response</th></tr></thead><tbody><tr><td>Retrieval unavailable</td><td>State that evidence could not be loaded</td></tr><tr><td>Policy service unavailable</td><td>Deny sensitive actions</td></tr><tr><td>Approval expired</td><td>Request a new approval</td></tr><tr><td>Output invalid</td><td>Repair within a bounded retry policy or stop</td></tr><tr><td>Tool result ambiguous</td><td>Do not claim success</td></tr><tr><td>Audit sink unavailable</td><td>Apply the declared fail-open or fail-closed policy</td></tr></tbody></table></div>\n<p>A model-generated statement of success is not a tool postcondition. Read the authoritative system state before telling the user that an external action completed.</p>\n<h2 id=\"evaluate-attacks-as-workflows\">Evaluate attacks as workflows</h2>\n<p>Single-turn refusal tests cover only model behavior. System tests should exercise:</p>\n<ul>\n<li>Direct requests for prohibited content.</li>\n<li>Indirect instructions inside retrieved documents.</li>\n<li>Attempts to access another principal’s resources.</li>\n<li>Argument substitution after approval.</li>\n<li>Tool-result spoofing.</li>\n<li>Encoded or fragmented policy violations.</li>\n<li>Excessive output and parser edge cases.</li>\n<li>Service outages during sensitive actions.</li>\n</ul>\n<p>Record the layer that blocked each test. A prompt-only defense that passes today can regress with a model update. An authorization denial should remain stable.</p>\n<p>NIST recommends documented test plans, independent assessment where appropriate, incident disclosure processes, and ongoing monitoring for generative AI systems.<sup><a href=\"https://blog.ecitis.org/ai-safety-systems/#user-content-fn-nist\" id=\"user-content-fnref-nist-2\">1</a></sup> Translate those controls into executable release gates and incident runbooks.</p>\n<h2 id=\"log-decisions-without-creating-a-second-leak\">Log decisions without creating a second leak</h2>\n<p>Audit records need enough detail for reconstruction:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"request_id\"</span><span>:</span><span> \"req_91\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"principal\"</span><span>:</span><span> \"user_24\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"policy_version\"</span><span>:</span><span> \"tools-12\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"tool\"</span><span>:</span><span> \"create_refund_request\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"decision\"</span><span>:</span><span> \"require_approval\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"approval_id\"</span><span>:</span><span> \"approval_77\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"result\"</span><span>:</span><span> \"created\"</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Identifiers and versions are illustrative.</p>\n<p>Do not log full prompts by default when they contain private documents. Store encrypted artifacts with limited retention and access, or log hashes plus selected redacted fields. Test redaction on tool errors because exception text often includes request bodies.</p>\n<h2 id=\"incident-response\">Incident response</h2>\n<p>An AI incident runbook needs system controls:</p>\n<ol>\n<li>Disable or narrow the affected tool.</li>\n<li>Revoke credentials and sessions.</li>\n<li>Preserve policy, prompt, model, retrieval, and tool revisions.</li>\n<li>Identify affected principals and resources.</li>\n<li>Replay the trace in an isolated environment.</li>\n<li>Add a regression case at the failed layer.</li>\n<li>Verify the fix against adjacent workflows.</li>\n</ol>\n<p>Changing the system prompt is appropriate only when the failure was actually instructional. A cross-tenant retrieval bug needs a query and authorization fix.</p>\n<h2 id=\"deployment-checklist\">Deployment checklist</h2>\n<ul>\n<li>Every request has an authenticated principal or an explicit anonymous policy.</li>\n<li>Retrieval filters execute before content leaves its authorization boundary.</li>\n<li>Tools expose narrow operations with typed inputs.</li>\n<li>Credentials never enter model-visible context.</li>\n<li>Sensitive actions require argument-bound approval.</li>\n<li>Tool brokers apply policy after model selection.</li>\n<li>Outputs pass syntax, schema, domain, and authorization checks.</li>\n<li>Success claims follow verified postconditions.</li>\n<li>Traces identify every deployed revision.</li>\n<li>Logs are redacted, access-controlled, and retained by policy.</li>\n<li>Offline attacks and production incidents become regression cases.</li>\n</ul>\n<p>Safety improves when authority is explicit and independently enforceable. The model proposes. Policy decides. Tools verify. Audit records preserve the evidence.</p>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-nist\">\n<p><a href=\"https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf\" rel=\"noopener noreferrer\">NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, July 2024</a>. <a href=\"https://blog.ecitis.org/ai-safety-systems/#user-content-fnref-nist\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/ai-safety-systems/#user-content-fnref-nist-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/ai-safety-systems/",
            "title": "AI Safety Systems in Production: Policy, Permissions, and Validation",
            "summary": "Design production safety as a control plane spanning identity, data access, tool permissions, output validation, monitoring, and response.",
            "image": "https://blog.ecitis.org/open-graph/ai-safety-systems.png",
            "date_modified": "2026-06-12T00:00:00.000Z",
            "date_published": "2026-06-12T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "AI safety",
                "Authorization",
                "Guardrails"
            ]
        },
        {
            "id": "https://blog.ecitis.org/long-context-models/",
            "content_html": "<p>A context window is a capacity limit, not a guarantee that every supplied token affects the answer equally. Long-context systems combine positional representation, attention computation, device memory, distributed execution, data composition, and evaluation.</p>\n<p>Extending one layer does not extend the others. A position encoding can represent a distant index while the cache does not fit. An exact attention kernel can process the sequence while the model fails to retrieve a fact placed in its middle.</p>\n<h2 id=\"position-enters-attention\">Position enters attention</h2>\n<p>Transformer attention without position is permutation-equivariant. Position methods change token representations or attention scores.</p>\n<p>Rotary position embedding applies a position-dependent rotation to query and key pairs. Their dot product then contains relative-position structure.<sup><a href=\"https://blog.ecitis.org/long-context-models/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup> For one two-dimensional feature pair:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>R</mi><mrow><mi>θ</mi><mo separator=\"true\">,</mo><mi>m</mi></mrow></msub><mo>=</mo><mrow><mo fence=\"true\">[</mo><mtable rowspacing=\"0.16em\" columnalign=\"center center\" columnspacing=\"1em\"><mtr><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mi>cos</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>m</mi><mi>θ</mi><mo stretchy=\"false\">)</mo></mrow></mstyle></mtd><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mo>−</mo><mi>sin</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>m</mi><mi>θ</mi><mo stretchy=\"false\">)</mo></mrow></mstyle></mtd></mtr><mtr><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mi>sin</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>m</mi><mi>θ</mi><mo stretchy=\"false\">)</mo></mrow></mstyle></mtd><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mi>cos</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>m</mi><mi>θ</mi><mo stretchy=\"false\">)</mo></mrow></mstyle></mtd></mtr></mtable><mo fence=\"true\">]</mo></mrow></mrow></semantics></math>\n<p>The model is trained over a range of phases determined by position <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>m</mi></mrow></semantics></math> and frequency <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>θ</mi></mrow></semantics></math>. Extrapolating beyond the training range can expose phase patterns not learned during training.</p>\n<p>ALiBi takes a different route. It adds a head-specific linear penalty proportional to query-key distance rather than adding a positional embedding to token states.<sup><a href=\"https://blog.ecitis.org/long-context-models/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> Its authors trained a <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>1.3</mn></mrow></semantics></math>-billion-parameter model at length <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>1024</mn></mrow></semantics></math> and evaluated length <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>2048</mn></mrow></semantics></math>; they report matching the perplexity of a sinusoidal baseline trained at the longer length while using <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>11</mn><mi mathvariant=\"normal\">%</mi></mrow></semantics></math> less training memory and time.<sup><a href=\"https://blog.ecitis.org/long-context-models/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup></p>\n<figure><figcaption><strong>Long-context methods alter different layers of the system</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/long-context-models-0.webp\" width=\"512\" height=\"431\" alt=\"Long-context methods alter different layers of the system\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"rope-extension-changes-frequency-mapping\">RoPE extension changes frequency mapping</h2>\n<p>Position interpolation compresses a longer inference range into coordinates observed during training. It avoids raw extrapolation but reduces angular resolution between nearby positions. Frequency-aware methods scale dimensions differently because high-frequency and low-frequency rotary pairs behave differently.</p>\n<p>LongRoPE searches non-uniform interpolation factors and uses progressive extension. The paper reports extending models to <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>2048</mn></mrow></semantics></math> thousand tokens with fine-tuning at lengths up to <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>256</mn></mrow></semantics></math> thousand tokens, then readjusting at <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>8</mn></mrow></semantics></math> thousand tokens to recover short-context behavior.<sup><a href=\"https://blog.ecitis.org/long-context-models/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup></p>\n<p>Those figures describe the evaluated method, not a universal recipe. Extension factors depend on the base model, original training length, data, and fine-tuning setup. A configuration that loads successfully can still degrade short-context perplexity or retrieval.</p>\n<p>Store position policy with the model artifact:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>position</span><span>:</span></span>\n<span class=\"line\"><span>  type</span><span>:</span><span> rope</span></span>\n<span class=\"line\"><span>  original_max_positions</span><span>:</span><span> ${TRAINING_LIMIT}</span></span>\n<span class=\"line\"><span>  scaling</span><span>:</span></span>\n<span class=\"line\"><span>    method</span><span>:</span><span> ${METHOD}</span></span>\n<span class=\"line\"><span>    factor</span><span>:</span><span> ${FACTOR}</span></span>\n<span class=\"line\"><span>    parameters_digest</span><span>:</span><span> ${DIGEST}</span></span>\n<span class=\"line\"><span>evaluation</span><span>:</span></span>\n<span class=\"line\"><span>  short_context_suite</span><span>:</span><span> ${SHORT_SUITE}</span></span>\n<span class=\"line\"><span>  long_context_suite</span><span>:</span><span> ${LONG_SUITE}</span></span></code></pre>\n<p>Placeholders force values to come from the trained artifact rather than a copied deployment snippet.</p>\n<h2 id=\"memory-grows-even-with-efficient-attention\">Memory grows even with efficient attention</h2>\n<p>Exact self-attention score work grows with the product of query and key lengths. FlashAttention avoids materializing the complete score matrix but does not remove the arithmetic of dense attention.<sup><a href=\"https://blog.ecitis.org/long-context-models/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup></p>\n<p>Inference also retains the key-value cache. For <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>L</mi></mrow></semantics></math> layers, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>H</mi><mrow><mi>k</mi><mi>v</mi></mrow></msub></mrow></semantics></math> key-value heads, head dimension <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>D</mi></mrow></semantics></math>, sequence length <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>S</mi></mrow></semantics></math>, and scalar width <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>B</mi></mrow></semantics></math>:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>M</mi><mrow><mi>K</mi><mi>V</mi></mrow></msub><mo>=</mo><mn>2</mn><mi>L</mi><mi>S</mi><msub><mi>H</mi><mrow><mi>k</mi><mi>v</mi></mrow></msub><mi>D</mi><mi>B</mi></mrow></semantics></math>\n<p>The cache grows linearly with sequence length. Prefix caching avoids recomputation across repeated prefixes but does not make a unique long sequence free. Quantized cache and grouped-query attention reduce bytes per token; both require model and kernel support.</p>\n<p>Before admitting a long request, estimate weight memory, cache reservation, temporary attention workspace, communication buffers, and allocator headroom. Token limits that ignore concurrent requests produce late out-of-memory failures.</p>\n<h2 id=\"context-parallelism-distributes-exact-attention\">Context parallelism distributes exact attention</h2>\n<p>When one device cannot hold sequence state, context parallelism partitions tokens across devices. Ring Attention rotates key-value blocks around a device ring and computes blockwise attention while communication overlaps with local work.<sup><a href=\"https://blog.ecitis.org/long-context-models/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup> Online softmax combines partial results without assembling the full score matrix on one device.</p>\n<p>For each local query block, a worker:</p>\n<ol>\n<li>computes scores against its current key-value block</li>\n<li>updates row maxima, normalization sums, and output accumulators</li>\n<li>sends the key-value block to the next worker</li>\n<li>receives another block and repeats</li>\n</ol>\n<p>The method preserves dense attention but adds communication proportional to the circulated state. Topology and link bandwidth become part of context-length capacity.</p>\n<h2 id=\"sparse-attention-changes-reachable-evidence\">Sparse attention changes reachable evidence</h2>\n<p>Sparse methods avoid evaluating every token pair. Sliding windows preserve local structure but remove direct long-range edges. Global tokens, vertical stripes, and selected blocks restore some long-range access.</p>\n<p>MInference identifies several sparse patterns per attention head and constructs dynamic sparse indices during prefill.<sup><a href=\"https://blog.ecitis.org/long-context-models/#user-content-fn-6\" id=\"user-content-fnref-6\">6</a></sup> Such methods accelerate a pattern observed in model attention; they do not preserve exact dense attention by definition.</p>\n<p>Validate sparse kernels on tasks requiring distant dependencies, not only language-model loss. A local window can score well on locally predictable text while dropping a document-level constraint.</p>\n<h2 id=\"retrieval-changes-input-rather-than-attention\">Retrieval changes input rather than attention</h2>\n<p>Retrieval selects a smaller evidence set before generation. Its cost shifts to indexing, query construction, ranking, and provenance. Long context and retrieval are complementary:</p>\n<ul>\n<li>retrieval reduces irrelevant tokens and prefill cost</li>\n<li>a longer context admits more retrieved chunks and surrounding text</li>\n<li>reranking orders evidence before prompt construction</li>\n<li>citations connect generated claims to supplied spans</li>\n</ul>\n<p>Chunking is a semantic decision. Fixed token windows can split tables, code definitions, and cross-references. Structure-aware chunking keeps those units intact, while overlap preserves boundary context at the cost of duplicate tokens.</p>\n<h2 id=\"position-sensitive-evaluation\">Position-sensitive evaluation</h2>\n<p>Liu and collaborators varied the position of relevant information within long inputs and observed that performance can be highest when relevant information is near the beginning or end and lower when it is in the middle.<sup><a href=\"https://blog.ecitis.org/long-context-models/#user-content-fn-7\" id=\"user-content-fnref-7\">7</a></sup> Their result motivates position sweeps rather than one fixed evidence placement.</p>\n<p>A useful evaluation matrix crosses:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Axis</th><th>Values to cover</th></tr></thead><tbody><tr><td>Evidence position</td><td>beginning, interior positions, end</td></tr><tr><td>Distractor count</td><td>low through deployment maximum</td></tr><tr><td>Evidence type</td><td>prose, table, code, cross-document link</td></tr><tr><td>Required operation</td><td>retrieval, aggregation, comparison, exact copy</td></tr><tr><td>Output contract</td><td>free text and structured result</td></tr></tbody></table></div>\n<p>Every test should record tokenized length, evidence offsets, truncation, retrieved chunk identifiers, and model position configuration. Otherwise a failure cannot be assigned to retrieval, prompt construction, model use, or runtime truncation.</p>\n<h2 id=\"operational-failure-modes\">Operational failure modes</h2>\n<p><strong>Silent truncation:</strong> the tokenizer or server drops tokens before model execution. Return accepted input length and truncation status in response metadata.</p>\n<p><strong>Position-config mismatch:</strong> weights were tuned with one RoPE policy but served with another. Hash the position configuration into the engine identity.</p>\n<p><strong>Cache admission failure:</strong> the request fits the advertised window alone but not under concurrency. Reserve cache before prefill.</p>\n<p><strong>Middle-position regression:</strong> aggregate benchmark scores hide positional weakness. Report accuracy by evidence offset.</p>\n<p><strong>Sparse-index error:</strong> a required block is omitted by an indexing heuristic. Maintain a dense reference path for reduced test cases.</p>\n<p><strong>Retrieval dilution:</strong> adding lower-ranked chunks moves useful evidence farther apart and adds distractors. Evaluate answer quality as the retrieved-token budget changes.</p>\n<h2 id=\"choosing-the-mechanism\">Choosing the mechanism</h2>\n<p>Use position extension when the model must represent indices beyond training and can be fine-tuned and re-evaluated. Use optimized dense attention when exact all-to-all access is required. Use sparse attention when model quality survives a restricted edge pattern. Use context parallelism when exact attention state exceeds one device. Use retrieval when only a subset of an external corpus is relevant to each request.</p>\n<p>These decisions compose. A production system can retrieve documents, serve an extended-RoPE model, run tiled attention, shard context across devices, and page its cache. Each layer needs a separate measurement and failure boundary.</p>\n<p>Long context becomes useful only when evidence survives tokenization, placement, attention, memory admission, and decoding. The maximum accepted token count measures just the first of those requirements.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Jianlin Su et al., “RoFormer: Enhanced Transformer with Rotary Position Embedding,” Neurocomputing, 2024. <a href=\"https://arxiv.org/abs/2104.09864\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2104.09864</a> <a href=\"https://blog.ecitis.org/long-context-models/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Ofir Press, Noah A. Smith, and Mike Lewis, “Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation,” ICLR, 2022. <a href=\"https://arxiv.org/abs/2108.12409\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2108.12409</a> <a href=\"https://blog.ecitis.org/long-context-models/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/long-context-models/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Yiran Ding et al., “LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens,” ICML, 2024. <a href=\"https://arxiv.org/abs/2402.13753\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2402.13753</a> <a href=\"https://blog.ecitis.org/long-context-models/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Tri Dao et al., “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” NeurIPS, 2022. <a href=\"https://arxiv.org/abs/2205.14135\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2205.14135</a> <a href=\"https://blog.ecitis.org/long-context-models/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>Hao Liu et al., “Ring Attention with Blockwise Transformers for Near-Infinite Context,” ICLR, 2024. <a href=\"https://arxiv.org/abs/2310.01889\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2310.01889</a> <a href=\"https://blog.ecitis.org/long-context-models/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p>Huiqiang Jiang et al., “MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention,” NeurIPS, 2024. <a href=\"https://arxiv.org/abs/2407.02490\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2407.02490</a> <a href=\"https://blog.ecitis.org/long-context-models/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p>Nelson F. Liu et al., “Lost in the Middle: How Language Models Use Long Contexts,” TACL, 2024. <a href=\"https://arxiv.org/abs/2307.03172\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2307.03172</a> <a href=\"https://blog.ecitis.org/long-context-models/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/long-context-models/",
            "title": "Long-Context Models: Position Encoding, Memory, and Retrieval",
            "summary": "Separate position extension, attention computation, cache capacity, context parallelism, and retrieval when designing long-context systems.",
            "image": "https://blog.ecitis.org/open-graph/long-context-models.png",
            "date_modified": "2026-06-05T00:00:00.000Z",
            "date_published": "2026-06-05T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Long context",
                "RoPE",
                "Retrieval"
            ]
        },
        {
            "id": "https://blog.ecitis.org/synthetic-data-pipelines/",
            "content_html": "<p>Synthetic data is a software artifact with lineage. A production pipeline must be able to answer which seed, prompt template, generator, sampling configuration, validator, and policy produced every accepted record. Without that chain, a high-scoring training run cannot be reproduced or audited.</p>\n<figure><figcaption><strong>Synthetic data is accepted only after independent checks</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/synthetic-data-pipelines-0.webp\" width=\"512\" height=\"703\" alt=\"Synthetic data is accepted only after independent checks\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"define-the-target-distribution-first\">Define the target distribution first</h2>\n<p>Generation should fill a documented coverage gap. Start with a taxonomy of task, language, domain, difficulty, input form, and expected output form. Assign each candidate to that taxonomy before accepting it.</p>\n<p>A record envelope can preserve the minimum lineage:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"record_id\"</span><span>:</span><span> \"content-addressed-id\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"seed_id\"</span><span>:</span><span> \"source-record-id\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"task\"</span><span>:</span><span> \"tool_argument_generation\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"generator\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"model\"</span><span>:</span><span> \"pinned-model-revision\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"template\"</span><span>:</span><span> \"template-content-hash\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"sampling\"</span><span>:</span><span> \"serialized-sampling-config\"</span></span>\n<span class=\"line\"><span>  }</span><span>,</span></span>\n<span class=\"line\"><span>  \"candidate\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"prompt\"</span><span>:</span><span> \"generated prompt\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"response\"</span><span>:</span><span> \"generated response\"</span></span>\n<span class=\"line\"><span>  }</span><span>,</span></span>\n<span class=\"line\"><span>  \"validation\"</span><span>:</span><span> []</span><span>,</span></span>\n<span class=\"line\"><span>  \"decision\"</span><span>:</span><span> \"pending\"</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Do not use an editable model alias or prompt filename as provenance. Store immutable revisions and content hashes.</p>\n<h2 id=\"seed-records-constrain-what-generation-can-learn\">Seed records constrain what generation can learn</h2>\n<p>Self-Instruct starts from seed tasks, generates new instructions and instances, filters invalid or similar outputs, then uses accepted examples for instruction tuning.<sup><a href=\"https://blog.ecitis.org/synthetic-data-pipelines/#user-content-fn-self-instruct\" id=\"user-content-fnref-self-instruct\">1</a></sup> The mechanism amplifies a seed distribution. It does not create an unbiased sample of every task users will request.</p>\n<p>Seed preparation should include:</p>\n<ul>\n<li>license and consent review;</li>\n<li>removal of secrets and personal data;</li>\n<li>exact and near-duplicate grouping;</li>\n<li>taxonomy labels;</li>\n<li>held-out source groups for evaluation;</li>\n<li>contamination checks against benchmark material.</li>\n</ul>\n<p>Split by source group before generation. If variants of one seed land in training and evaluation, the evaluation measures template transfer rather than independent generalization.</p>\n<h2 id=\"generation-should-expose-controlled-diversity\">Generation should expose controlled diversity</h2>\n<p>Prompt templates can request variations in task structure, constraints, domain, or reasoning requirement. Sampling adds lexical and solution diversity. Both can also produce invalid examples.</p>\n<p>Generate multiple candidates per seed only when a validator can distinguish them. Otherwise the pipeline increases volume without evidence of quality. Preserve rejected candidates and reasons so generator changes can be compared on the same acceptance criteria.</p>\n<p>Structured tasks should use structured generation:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> dataclasses </span><span>import</span><span> dataclass</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@dataclass</span></span>\n<span class=\"line\"><span>class</span><span> Candidate</span><span>:</span></span>\n<span class=\"line\"><span>    instruction</span><span>:</span><span> str</span></span>\n<span class=\"line\"><span>    input</span><span>:</span><span> str</span></span>\n<span class=\"line\"><span>    output</span><span>:</span><span> str</span></span>\n<span class=\"line\"><span>    assumptions</span><span>:</span><span> list</span><span>[</span><span>str</span><span>]</span></span>\n<span class=\"line\"><span>    source_ids</span><span>:</span><span> list</span><span>[</span><span>str</span><span>]</span></span></code></pre>\n<p>A schema verifies form, not truth. It prevents missing fields and malformed types from entering later stages.</p>\n<h2 id=\"use-validators-with-independent-evidence\">Use validators with independent evidence</h2>\n<p>A generator grading its own answer repeats correlated errors. Prefer validators that obtain evidence through a different mechanism.</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Task</th><th>Strong validator</th><th>Weak validator</th></tr></thead><tbody><tr><td>Code generation</td><td>compile, tests, static analysis</td><td>generator self-rating</td></tr><tr><td>Structured extraction</td><td>schema plus comparison to source</td><td>fluent explanation</td></tr><tr><td>Mathematics</td><td>symbolic check or alternate derivation</td><td>answer-format match</td></tr><tr><td>Retrieval question</td><td>entailment against cited passages</td><td>uncited judge score</td></tr><tr><td>Tool calls</td><td>schema plus sandbox execution</td><td>syntactic JSON only</td></tr></tbody></table></div>\n<p>Execution validators need containment. Run generated code with resource limits, no ambient credentials, restricted networking, and disposable storage. Record exit status and test output, but sanitize untrusted content before displaying it in internal tools.</p>\n<p>For open-ended responses, use a rubric with observable failure labels such as unsupported claim, missing constraint, contradiction, unsafe instruction, or non-answer. Calibrate model-based judges against a human-reviewed sample and recheck calibration when the generator or domain changes.</p>\n<h2 id=\"deduplicate-before-and-after-generation\">Deduplicate before and after generation</h2>\n<p>Exact hashing removes byte-identical records. Normalized hashing can ignore formatting, case, or boilerplate. Semantic similarity catches paraphrases, but thresholds depend on the embedding model and domain.</p>\n<p>Deduplication has two distinct objectives:</p>\n<ol>\n<li>prevent repeated training weight on the same underlying example;</li>\n<li>prevent evaluation leakage from a training example or its derivative.</li>\n</ol>\n<p>The second objective needs lineage-aware checks. A generated prompt can be lexically distant from its seed while preserving the same answer structure. Group descendants with their seed during split assignment.</p>\n<p>Store cluster identifiers and representative records. Deleting duplicates without a manifest makes future dataset diffs difficult to explain.</p>\n<h2 id=\"filter-by-difficulty-with-executable-signals\">Filter by difficulty with executable signals</h2>\n<p>Difficulty labels produced only by a generator are unstable. Prefer signals tied to the task:</p>\n<ul>\n<li>number and type of constraints satisfied;</li>\n<li>proof or program length after normalization;</li>\n<li>number of tool calls required by a reference solution;</li>\n<li>pass rate across independent solvers;</li>\n<li>retrieval depth or evidence count;</li>\n<li>disagreement among validated candidate solutions.</li>\n</ul>\n<p>Do not equate longer chain-of-thought text with harder reasoning. A verbose candidate can contain redundant or fabricated steps. If reasoning traces are used for training, validate the final answer and inspect step consistency with tools where the domain permits it.</p>\n<h2 id=\"keep-private-reasoning-out-of-public-data-contracts\">Keep private reasoning out of public data contracts</h2>\n<p>A synthetic pipeline can train on solutions without exposing raw reasoning at inference. Separate outcome supervision, concise rationales, tool traces, and hidden scratch work as distinct fields. The serving format should not inherit the generator’s internal trace by accident.</p>\n<p>For tool use, a compact trace is often more useful:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>request</span></span>\n<span class=\"line\"><span>tool schema selected</span></span>\n<span class=\"line\"><span>validated arguments</span></span>\n<span class=\"line\"><span>tool result reference</span></span>\n<span class=\"line\"><span>final answer grounded in result</span></span></code></pre>\n<p>This structure is auditable and maps to runtime behavior.</p>\n<h2 id=\"control-model-collapse-with-real-anchors\">Control model collapse with real anchors</h2>\n<p>Training repeatedly on model-generated distributions can reduce coverage of low-probability events and propagate generator errors. A synthetic corpus therefore needs real-data anchors, independent evaluation, and explicit mixture weights. Synthetic examples should address measured gaps rather than replace the entire observed distribution.</p>\n<p>Track acceptance and performance by source type. If synthetic data improves aggregate accuracy while reducing performance on a real-data slice, the aggregate result is insufficient.</p>\n<h2 id=\"build-evaluation-before-the-large-run\">Build evaluation before the large run</h2>\n<p>The evaluation set must not be generated by the same templates, seeds, and model revision used for training. Include:</p>\n<ul>\n<li>a frozen real-data set;</li>\n<li>adversarial cases written or curated independently;</li>\n<li>schema and execution tests;</li>\n<li>slice metrics for the target taxonomy;</li>\n<li>memorization probes against seeds;</li>\n<li>safety and privacy probes;</li>\n<li>regression examples from production failures.</li>\n</ul>\n<p>FLAN work showed that scaling instruction-finetuning tasks and using chain-of-thought data can improve model behavior across evaluation families.<sup><a href=\"https://blog.ecitis.org/synthetic-data-pipelines/#user-content-fn-flan\" id=\"user-content-fnref-flan\">2</a></sup> That result supports task diversity as a training variable. It does not validate any arbitrary synthetic dataset, so each pipeline still needs its own ablations.</p>\n<p>Run mixture ablations rather than comparing only no-synthetic versus all-synthetic. Change one factor at a time: seed source, generator revision, filtering stage, or mixture weight.</p>\n<h2 id=\"version-the-dataset-as-a-build\">Version the dataset as a build</h2>\n<p>A dataset release should contain:</p>\n<ul>\n<li>source and license manifest;</li>\n<li>generator and prompt hashes;</li>\n<li>validator versions;</li>\n<li>rejection counts by reason;</li>\n<li>deduplication clusters;</li>\n<li>split assignment logic;</li>\n<li>final record hashes;</li>\n<li>evaluation report;</li>\n<li>known limitations.</li>\n</ul>\n<p>Treat validator changes like code changes. Rebuild from immutable candidates, diff accepted record IDs, and investigate large movements. A judge prompt edit can alter the training distribution even when generation is unchanged.</p>\n<h2 id=\"failure-modes\">Failure modes</h2>\n<p><strong>Validator leakage:</strong> The generator learns the tests or rubric and produces outputs tailored to them. Keep hidden checks and rotate surface forms without changing the criterion.</p>\n<p><strong>Homogeneous style:</strong> One generator and template create repeated syntax. Measure lexical and structural diversity, then introduce source or template diversity with unchanged validation.</p>\n<p><strong>False authority:</strong> Candidates include citations that look valid but do not support the claim. Fetch the cited primary source and verify entailment, or reject the citation.</p>\n<p><strong>Split contamination:</strong> Derived examples cross seed-group boundaries. Assign the split at the ancestor seed and inherit it through every transformation.</p>\n<p><strong>Silent policy drift:</strong> A hosted generator changes behind an alias. Pin a revision when possible and retain probe outputs that identify behavior changes.</p>\n<p>Synthetic data quality comes from the rejection path as much as the generator. The durable pipeline preserves provenance, validates with independent evidence, deduplicates by lineage, anchors evaluation in real data, and makes every accepted record reproducible.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-self-instruct\">\n<p><a href=\"https://arxiv.org/abs/2212.10560\" rel=\"noopener noreferrer\">Self-Instruct: Aligning Language Models with Self-Generated Instructions</a>. <a href=\"https://blog.ecitis.org/synthetic-data-pipelines/#user-content-fnref-self-instruct\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-flan\">\n<p><a href=\"https://arxiv.org/abs/2210.11416\" rel=\"noopener noreferrer\">Scaling Instruction-Finetuned Language Models</a>. <a href=\"https://blog.ecitis.org/synthetic-data-pipelines/#user-content-fnref-flan\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/synthetic-data-pipelines/",
            "title": "Synthetic Data Pipelines: Generation, Filtering, and Validation",
            "summary": "Build reproducible synthetic-data systems with explicit lineage, generation, filtering, deduplication, validation, and policy controls.",
            "image": "https://blog.ecitis.org/open-graph/synthetic-data-pipelines.png",
            "date_modified": "2026-05-28T00:00:00.000Z",
            "date_published": "2026-05-28T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Synthetic data",
                "Data quality",
                "Training"
            ]
        },
        {
            "id": "https://blog.ecitis.org/llm-evaluation/",
            "content_html": "<p>An evaluation is a measurement procedure, not a prompt file. Its result depends on the dataset, harness, model configuration, scorer, aggregation, and statistical analysis. Changing any one of them changes the object being measured.</p>\n<p>This matters for language-model systems because the model is only one component. Retrieval, tools, policy prompts, sampling parameters, and output parsers can determine whether the same model succeeds or fails.</p>\n<h2 id=\"define-the-decision-before-the-metric\">Define the decision before the metric</h2>\n<p>Start with the release decision:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>decision</span><span>:</span><span> replace current support-answering model</span></span>\n<span class=\"line\"><span>population</span><span>:</span><span> authenticated English support questions</span></span>\n<span class=\"line\"><span>constraints</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>no regression on policy compliance</span></span>\n<span class=\"line\"><span>  - </span><span>bounded p95 response latency</span></span>\n<span class=\"line\"><span>  - </span><span>bounded cost per resolved conversation</span></span>\n<span class=\"line\"><span>outcomes</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>correct resolution</span></span>\n<span class=\"line\"><span>  - </span><span>unsupported claim rate</span></span>\n<span class=\"line\"><span>  - </span><span>escalation rate</span></span></code></pre>\n<p>The numeric thresholds are product choices and therefore omitted from this example. A benchmark score that does not connect to the decision is diagnostic evidence, not a release gate.</p>\n<p>HELM argues for broad scenario coverage and simultaneous measurement of accuracy, robustness, calibration, fairness, bias, toxicity, and efficiency rather than a single model score.<sup><a href=\"https://blog.ecitis.org/llm-evaluation/#user-content-fn-helm\" id=\"user-content-fnref-helm\">1</a></sup> The exact dimensions should follow the deployed task.</p>\n<h2 id=\"build-cases-from-a-task-taxonomy\">Build cases from a task taxonomy</h2>\n<p>Random production samples reproduce the traffic distribution but can miss rare, costly failures. Handwritten challenge sets cover known hazards but distort prevalence. Use both and report them separately.</p>\n<p>A case should contain:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> EvalCase</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  id</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  input</span><span>:</span><span> unknown</span></span>\n<span class=\"line\"><span>  expected</span><span>?:</span><span> unknown</span></span>\n<span class=\"line\"><span>  rubric</span><span>?:</span><span> string</span></span>\n<span class=\"line\"><span>  tags</span><span>:</span><span> string</span><span>[]</span></span>\n<span class=\"line\"><span>  source</span><span>:</span><span> 'production'</span><span> |</span><span> 'expert'</span><span> |</span><span> 'incident'</span><span> |</span><span> 'synthetic'</span></span>\n<span class=\"line\"><span>  provenance</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  createdAt</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Tags enable slices such as language, tool, document type, ambiguity, policy class, and input length. Provenance supports contamination audits and rights review.</p>\n<p>Freeze a release suite. Add new incident cases to a forward-looking suite so repeated tuning does not quietly overfit the current gate.</p>\n<figure><figcaption><strong>Evaluation pipeline with validity gates</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/llm-evaluation-0.webp\" width=\"512\" height=\"641\" alt=\"Evaluation pipeline with validity gates\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"pin-the-harness\">Pin the harness</h2>\n<p>The harness includes every transformation between a case and its score:</p>\n<ul>\n<li>System and developer instructions.</li>\n<li>Tool definitions and tool results.</li>\n<li>Retrieval index and ranking configuration.</li>\n<li>Model identifier and sampling parameters.</li>\n<li>Retry and timeout policy.</li>\n<li>Output parser.</li>\n<li>Scorer prompt and scorer model.</li>\n</ul>\n<p>Record the rendered request rather than only the prompt template. A misplaced separator, reordered choice list, or truncated tool result can alter the outcome.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"case_id\"</span><span>:</span><span> \"billing-policy-017\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"suite_revision\"</span><span>:</span><span> \"support-v4\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"model\"</span><span>:</span><span> \"candidate-build\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"prompt_sha256\"</span><span>:</span><span> \"…\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"toolset_revision\"</span><span>:</span><span> \"support-tools-v7\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"retrieval_revision\"</span><span>:</span><span> \"kb-2026-05-18\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"sampler\"</span><span>:</span><span> { </span><span>\"temperature\"</span><span>:</span><span> 0</span><span> }</span><span>,</span></span>\n<span class=\"line\"><span>  \"raw_output_uri\"</span><span>:</span><span> \"artifact://run/case/output\"</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Version names above are illustrative.</p>\n<h2 id=\"deterministic-scoring-comes-first\">Deterministic scoring comes first</h2>\n<p>Use deterministic checks when the task has an executable contract:</p>\n<ul>\n<li>JSON Schema validity.</li>\n<li>Exact identifier match.</li>\n<li>Unit tests for generated code.</li>\n<li>Database-state assertions after tool use.</li>\n<li>Citation presence and source alignment.</li>\n<li>Regular-language constraints.</li>\n</ul>\n<p>Deterministic scorers expose why a case failed and can be rerun without model variance. They do not measure open-ended quality, but they should handle every property that can be specified mechanically.</p>\n<p>For partial-credit tasks, define components:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>S</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo><mo>=</mo><munder><mo>∑</mo><mi>i</mi></munder><msub><mi>w</mi><mi>i</mi></msub><msub><mi>s</mi><mi>i</mi></msub><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>Weights encode a product judgment. Publish them with the results and report component scores so one weighted total does not hide a safety regression.</p>\n<h2 id=\"human-evaluation-needs-a-protocol\">Human evaluation needs a protocol</h2>\n<p>Human ratings are not ground truth by default. Raters need:</p>\n<ul>\n<li>A concrete rubric with positive and negative examples.</li>\n<li>Blinded model identity.</li>\n<li>Randomized answer order.</li>\n<li>An abstain or unclear option.</li>\n<li>Independent ratings before adjudication.</li>\n<li>Quality-control cases.</li>\n</ul>\n<p>Measure agreement and inspect disagreements. Low agreement can indicate an ambiguous rubric, insufficient evidence, or a task with legitimate preference variation.</p>\n<p>Pairwise comparison is often easier than an absolute scale: show two outputs for the same input and ask which better satisfies the rubric. Randomize left and right positions, then include ties when the difference is not material.</p>\n<h2 id=\"llm-judges-are-noisy-instruments\">LLM judges are noisy instruments</h2>\n<p>Model-based scoring enables larger evaluation sets, but it introduces a second model whose behavior must be validated.</p>\n<p>The MT-Bench and Chatbot Arena paper compared model judges with human preferences and documented position, verbosity, and self-enhancement biases.<sup><a href=\"https://blog.ecitis.org/llm-evaluation/#user-content-fn-mtbench\" id=\"user-content-fnref-mtbench\">2</a></sup> A judge can therefore reward presentation features unrelated to task correctness.</p>\n<p>Mitigations are measurable:</p>\n<ol>\n<li>Swap candidate order and score both orientations.</li>\n<li>Hide model identity and provider-specific formatting.</li>\n<li>Require criterion-level scores before an overall preference.</li>\n<li>Give the judge reference evidence when the task permits it.</li>\n<li>Calibrate against a held-out human-rated set.</li>\n<li>Route judge-human disagreement and low-confidence cases to review.</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>score(A, B) + score(B, A) → stable preference or disagreement</span></span></code></pre>\n<p>Do not ask a judge for hidden chain-of-thought. Request a short evidence field tied to rubric criteria, then audit whether that evidence matches the outputs.</p>\n<p>Judge agreement with humans should be reported on the target task, not borrowed from another benchmark. A judge calibrated for concise customer support answers may fail on mathematical proofs or code patches.</p>\n<figure><figcaption><strong>Order swapping exposes a position-sensitive judge</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/llm-evaluation-1.webp\" width=\"512\" height=\"311\" alt=\"Order swapping exposes a position-sensitive judge\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"contamination-invalidates-the-question\">Contamination invalidates the question</h2>\n<p>Benchmark contamination occurs when evaluation material or close derivatives enter training, fine-tuning, retrieval corpora, prompt libraries, or manual development loops. A high score can then measure recall of the benchmark rather than task generalization.</p>\n<p>Contamination is not limited to exact duplicates. Paraphrased questions, published solutions, answer keys, and benchmark-specific explanations can leak the intended mapping.</p>\n<p>Mitigations include:</p>\n<ul>\n<li>Keep private test cases and rotate them.</li>\n<li>Hash and search exact examples across accessible training corpora.</li>\n<li>Run semantic near-duplicate detection.</li>\n<li>Store creation dates and source provenance.</li>\n<li>Separate public development sets from private release sets.</li>\n<li>Treat benchmark names and canonical phrasing as contamination signals.</li>\n</ul>\n<p>Black-box detection is uncertain. The Min-K% Prob method tests whether a text contains unusually low-probability tokens and was evaluated for pretraining-data detection on a time-split benchmark.<sup><a href=\"https://blog.ecitis.org/llm-evaluation/#user-content-fn-mink\" id=\"user-content-fnref-mink\">3</a></sup> It is evidence, not proof of membership. A current evaluation should describe the audit performed and residual uncertainty.</p>\n<p>Dynamic cases reduce exposure. Generate them from controlled templates with executable answers, then validate the generator. Do not call generated cases uncontaminated merely because their final strings are new; the underlying task pattern may still be familiar.</p>\n<h2 id=\"statistical-reporting\">Statistical reporting</h2>\n<p>A score without uncertainty encourages false precision. For a binary metric with <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>n</mi></mrow></semantics></math> independent cases and empirical success rate <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mover accent=\"true\"><mi>p</mi><mo>^</mo></mover></mrow></semantics></math>, a simple standard error estimate is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"normal\">SE</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mover accent=\"true\"><mi>p</mi><mo>^</mo></mover><mo stretchy=\"false\">)</mo><mo>=</mo><msqrt><mfrac><mrow><mover accent=\"true\"><mi>p</mi><mo>^</mo></mover><mo stretchy=\"false\">(</mo><mn>1</mn><mo>−</mo><mover accent=\"true\"><mi>p</mi><mo>^</mo></mover><mo stretchy=\"false\">)</mo></mrow><mi>n</mi></mfrac></msqrt></mrow></semantics></math>\n<p>Bootstrap intervals are more flexible for non-linear metrics and sliced data. For model comparisons, use paired resampling because both models run on the same cases.</p>\n<p>Report:</p>\n<ul>\n<li>Number of eligible, executed, failed, and excluded cases.</li>\n<li>Aggregate and slice metrics.</li>\n<li>Confidence intervals.</li>\n<li>Paired differences against the current system.</li>\n<li>Missing-output and parser-failure rates.</li>\n<li>Judge disagreement and human adjudication rates.</li>\n</ul>\n<p>Do not remove timeouts from the denominator unless the release decision also ignores timeouts.</p>\n<h2 id=\"production-metrics-close-the-loop\">Production metrics close the loop</h2>\n<p>Offline cases are controlled and replayable. Production metrics observe the real distribution and real users. Neither replaces the other.</p>\n<p>Useful system metrics include:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Layer</th><th>Measurements</th></tr></thead><tbody><tr><td>Request</td><td>eligible volume, language, intent, input length</td></tr><tr><td>Model</td><td>latency, token usage, refusal, malformed output</td></tr><tr><td>Retrieval</td><td>empty results, evidence coverage, stale index</td></tr><tr><td>Tools</td><td>call success, invalid arguments, approval, rollback</td></tr><tr><td>Outcome</td><td>resolution, correction, escalation, abandonment</td></tr><tr><td>Safety</td><td>policy blocks, confirmed incidents, reviewer reversals</td></tr></tbody></table></div>\n<p>User clicks and thumbs ratings are behavior signals, not direct correctness labels. A user can accept a plausible wrong answer or downvote a correct refusal. Combine behavioral metrics with audited samples.</p>\n<p>For online experiments, define guardrails before launch and preserve assignment logs. Model changes can alter cost, latency, and escalation while leaving average answer ratings unchanged.</p>\n<h2 id=\"detect-drift\">Detect drift</h2>\n<p>Drift can occur in inputs, retrieved knowledge, tool behavior, policy, or user expectations. Track slice proportions and score distributions over time.</p>\n<p>Trigger a review when:</p>\n<ul>\n<li>A new intent lacks evaluation coverage.</li>\n<li>A retrieval source changes schema.</li>\n<li>A tool adds a side effect.</li>\n<li>A scorer model or prompt changes.</li>\n<li>Human-judge agreement degrades.</li>\n<li>Production outcomes separate from offline scores.</li>\n</ul>\n<p>An evaluation suite is versioned software. Changes need review, changelogs, reproducible artifacts, and migration notes.</p>\n<h2 id=\"diagnose-failures-by-layer\">Diagnose failures by layer</h2>\n<p><strong>Correct evidence absent:</strong> retrieval failure.</p>\n<p><strong>Correct evidence present, answer unsupported:</strong> generation or instruction failure.</p>\n<p><strong>Valid action selected with wrong arguments:</strong> tool-planning or schema failure.</p>\n<p><strong>Correct output marked wrong:</strong> scorer or rubric failure.</p>\n<p><strong>Offline pass with production regression:</strong> population, harness, or outcome mismatch.</p>\n<p>Store raw inputs, outputs, tool traces, scorer decisions, and policy versions so these categories can be tested rather than guessed.</p>\n<h2 id=\"release-report\">Release report</h2>\n<p>A useful report states:</p>\n<ol>\n<li>The deployment decision.</li>\n<li>Suite provenance and known contamination risk.</li>\n<li>Harness and model revisions.</li>\n<li>Deterministic, judge, and human scoring procedures.</li>\n<li>Aggregate results with uncertainty.</li>\n<li>Critical slices and regressions.</li>\n<li>Production guardrails.</li>\n<li>Known blind spots.</li>\n</ol>\n<p>The report should make rerunning the evaluation possible. A leaderboard number without cases, harness, and scoring details cannot support a production decision.</p>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-helm\">\n<p><a href=\"https://arxiv.org/abs/2211.09110\" rel=\"noopener noreferrer\">Liang et al., “Holistic Evaluation of Language Models,” Transactions on Machine Learning Research, 2023</a>. <a href=\"https://blog.ecitis.org/llm-evaluation/#user-content-fnref-helm\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mtbench\">\n<p><a href=\"https://arxiv.org/abs/2306.05685\" rel=\"noopener noreferrer\">Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” NeurIPS Datasets and Benchmarks, 2023</a>. <a href=\"https://blog.ecitis.org/llm-evaluation/#user-content-fnref-mtbench\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mink\">\n<p><a href=\"https://arxiv.org/abs/2310.16789\" rel=\"noopener noreferrer\">Shi et al., “Detecting Pretraining Data from Large Language Models,” ICLR, 2024</a>. <a href=\"https://blog.ecitis.org/llm-evaluation/#user-content-fnref-mink\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/llm-evaluation/",
            "title": "LLM Evaluation: Contamination, Judge Bias, and Production Metrics",
            "summary": "Design evaluations around release decisions while controlling contamination, judge bias, statistical noise, and production constraints.",
            "image": "https://blog.ecitis.org/open-graph/llm-evaluation.png",
            "date_modified": "2026-05-19T00:00:00.000Z",
            "date_published": "2026-05-19T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Evaluation",
                "Benchmarks",
                "LLM judges"
            ]
        },
        {
            "id": "https://blog.ecitis.org/on-device-multimodal-ai/",
            "content_html": "<p>An on-device multimodal model is a scheduling problem over shared memory. Camera buffers, audio capture, encoders, language-model weights, KV cache, and UI surfaces compete for the same physical resources. A model that fits at rest can still fail when an encoder and decoder reach peak allocation together.</p>\n<figure><figcaption><strong>A shared memory budget, not three independent pipelines</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/on-device-multimodal-ai-0.webp\" width=\"512\" height=\"394\" alt=\"A shared memory budget, not three independent pipelines\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"account-for-peak-live-memory\">Account for peak live memory</h2>\n<p>A useful budget separates persistent and transient allocations:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>M</mi><mtext>peak</mtext></msub><mo>=</mo><msub><mi>M</mi><mtext>weights</mtext></msub><mo>+</mo><msub><mi>M</mi><mtext>KV</mtext></msub><mo>+</mo><msub><mi>M</mi><mtext>active encoders</mtext></msub><mo>+</mo><msub><mi>M</mi><mtext>activations</mtext></msub><mo>+</mo><msub><mi>M</mi><mtext>workspace</mtext></msub><mo>+</mo><msub><mi>M</mi><mtext>application</mtext></msub></mrow></semantics></math>\n<p>Weight-file size is not peak memory. Runtimes can unpack weights, allocate graph plans, retain caches, and copy buffers between CPU and accelerator address spaces. The operating system also needs headroom for the camera, microphone, display compositor, and the rest of the application.</p>\n<p>Instrument resident memory and accelerator allocations around each phase:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>load language model</span></span>\n<span class=\"line\"><span>load vision encoder</span></span>\n<span class=\"line\"><span>encode image</span></span>\n<span class=\"line\"><span>release vision workspace</span></span>\n<span class=\"line\"><span>append visual tokens</span></span>\n<span class=\"line\"><span>load or activate audio encoder</span></span>\n<span class=\"line\"><span>encode audio window</span></span>\n<span class=\"line\"><span>release audio workspace</span></span>\n<span class=\"line\"><span>decode text</span></span></code></pre>\n<p>Serial encoder scheduling lowers peak concurrency when the application does not need simultaneous frame and audio processing. The cost is latency. Measure the user-visible critical path before selecting concurrency.</p>\n<h2 id=\"encoders-turn-media-into-token-sequences\">Encoders turn media into token sequences</h2>\n<p>A vision encoder maps image patches or feature maps into embeddings. An adapter projects those embeddings into the language model’s hidden space. A speech encoder performs a related conversion over acoustic frames. The language model then consumes an interleaved token sequence containing text and modality embeddings.</p>\n<p>Gemma 3n provides a concrete mobile architecture: it combines text with image, video, and audio inputs, uses a MobileNet-V5 vision encoder, and derives its audio encoder from Universal Speech Model work.<sup><a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fn-gemma\" id=\"user-content-fnref-gemma\">1</a></sup> Google’s implementation reports one audio token per <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>160</mn></mrow></semantics></math> milliseconds and offers image input resolutions of <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>256</mn><mo>×</mo><mn>256</mn></mrow></semantics></math>, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>512</mn><mo>×</mo><mn>512</mn></mrow></semantics></math>, and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>768</mn><mo>×</mo><mn>768</mn></mrow></semantics></math> pixels.<sup><a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fn-gemma\" id=\"user-content-fnref-gemma-2\">1</a></sup> Those are model-specific interface facts, not defaults for all multimodal systems.</p>\n<p>Media preprocessing must match training:</p>\n<ul>\n<li>image color space, resizing, crop policy, and normalization;</li>\n<li>audio sample rate, channel mixing, windowing, and feature extraction;</li>\n<li>special tokens that delimit each modality;</li>\n<li>positional encoding used for media tokens;</li>\n<li>adapter weights paired with the correct encoder and language model.</li>\n</ul>\n<p>An apparently plausible response does not prove preprocessing is correct. Validate embeddings or logits against a reference runtime with fixed media fixtures.</p>\n<h2 id=\"sequence-length-is-a-memory-input\">Sequence length is a memory input</h2>\n<p>The language model sees encoder outputs as tokens. More image tiles, video frames, or audio windows extend the prefill sequence. KV cache memory for an autoregressive decoder grows with retained sequence length:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>M</mi><mtext>KV</mtext></msub><mo>∝</mo><mi>L</mi><mo>×</mo><mi>S</mi><mo>×</mo><msub><mi>H</mi><mtext>KV</mtext></msub><mo>×</mo><mi>b</mi></mrow></semantics></math>\n<p>Here <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>L</mi></mrow></semantics></math> is cached layer count, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>S</mi></mrow></semantics></math> is sequence length, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>H</mi><mtext>KV</mtext></msub></mrow></semantics></math> is the cached key-value width, and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>b</mi></mrow></semantics></math> is bytes per stored element. Grouped-query or multi-query attention changes <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>H</mi><mtext>KV</mtext></msub></mrow></semantics></math> by sharing key-value heads. Quantized cache storage changes <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>b</mi></mrow></semantics></math> but adds conversion and accuracy considerations.</p>\n<p>Gemma 3n documents KV cache sharing in which middle-layer key and value representations are shared with upper layers to reduce prefill work.<sup><a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fn-gemma\" id=\"user-content-fnref-gemma-3\">1</a></sup> That is an architectural optimization built into the model. A runtime cannot apply it to an arbitrary checkpoint without matching training and graph structure.</p>\n<p>Set modality budgets before preprocessing:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>data</span><span> class</span><span> RequestBudget</span><span>(</span></span>\n<span class=\"line\"><span>    val</span><span> maxImageTokens: </span><span>Int</span><span>,</span></span>\n<span class=\"line\"><span>    val</span><span> maxAudioTokens: </span><span>Int</span><span>,</span></span>\n<span class=\"line\"><span>    val</span><span> maxTextTokens: </span><span>Int</span><span>,</span></span>\n<span class=\"line\"><span>    val</span><span> maxOutputTokens: </span><span>Int</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>fun</span><span> validate</span><span>(b: </span><span>RequestBudget</span><span>, contextLimit: </span><span>Int</span><span>) {</span></span>\n<span class=\"line\"><span>    require</span><span>(</span></span>\n<span class=\"line\"><span>        b.maxImageTokens </span><span>+</span></span>\n<span class=\"line\"><span>        b.maxAudioTokens </span><span>+</span></span>\n<span class=\"line\"><span>        b.maxTextTokens </span><span>+</span></span>\n<span class=\"line\"><span>        b.maxOutputTokens </span><span>&lt;=</span><span> contextLimit</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The application should reject or downsample input before allocating an oversized cache.</p>\n<h2 id=\"vision-needs-an-explicit-frame-policy\">Vision needs an explicit frame policy</h2>\n<p>Video is not one large image. It is a sequence of frames with temporal redundancy. Feeding every captured frame wastes encoder work and context. Define a policy based on the task:</p>\n<ul>\n<li>periodic sampling for scene summaries;</li>\n<li>motion- or event-triggered frames for monitoring;</li>\n<li>keyframes plus crops for document or screen understanding;</li>\n<li>a rolling window when recent state matters;</li>\n<li>frame replacement when the model supports external memory.</li>\n</ul>\n<p>Retain timestamps and orientation metadata. If visual tokens from multiple frames are concatenated without temporal markers, the language model cannot reliably infer capture order from pixels alone.</p>\n<p>Camera images often arrive in a hardware-native pixel format. Converting full-resolution frames on the CPU can become the bottleneck before inference begins. Prefer runtime-supported texture or buffer paths, and crop or resize before expensive format conversion when the APIs permit it.</p>\n<h2 id=\"audio-needs-streaming-state\">Audio needs streaming state</h2>\n<p>Audio capture produces a continuous stream while language-model decoding is bursty. A bounded ring buffer separates the two schedules. The encoder consumes windows with overlap or streaming state, then emits tokens or transcript segments.</p>\n<p>Google describes the Gemma 3n audio encoder as streaming-capable, while its launch implementation accepts clips up to <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>30</mn></mrow></semantics></math> seconds.<sup><a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fn-gemma\" id=\"user-content-fnref-gemma-4\">1</a></sup> That limit belongs to the released implementation. Long-form support requires the runtime and model training needed to preserve streaming state.</p>\n<p>For a speech request, track:</p>\n<ul>\n<li>capture timestamp and dropped frames;</li>\n<li>resampling configuration;</li>\n<li>encoder state reset points;</li>\n<li>token timestamps if exposed;</li>\n<li>end-of-speech decision;</li>\n<li>time from speech end to first generated token.</li>\n</ul>\n<p>Voice activity detection saves encoder and language-model work, but a false endpoint can cut a word or reset context. Keep the raw audio window needed to reproduce failures, subject to the product’s privacy policy.</p>\n<h2 id=\"choose-placement-per-operator\">Choose placement per operator</h2>\n<p>Mobile systems may expose CPU, GPU, neural accelerator, or vendor delegate paths. One model can span them, but every boundary risks a copy or synchronization. A placement plan should use measured operator coverage:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Subgraph</th><th>Placement question</th></tr></thead><tbody><tr><td>Vision encoder</td><td>Does the delegate support its convolution and attention blocks without fallback?</td></tr><tr><td>Audio frontend</td><td>Is feature extraction cheaper on vector CPU code than through accelerator dispatch?</td></tr><tr><td>Language prefill</td><td>Which backend sustains the matrix shapes produced by media tokens?</td></tr><tr><td>Token decode</td><td>Which backend gives low latency at a small batch?</td></tr></tbody></table></div>\n<p>An unsupported operator in the middle of an accelerator graph can force partitioning and transfers. Inspect the compiled delegate graph rather than relying on an API call that reports delegate initialization.</p>\n<h2 id=\"quantization-requires-modality-tests\">Quantization requires modality tests</h2>\n<p>Weight quantization reduces resident model memory. Activation quantization can reduce bandwidth and accelerator requirements. It can also alter small embedding differences at modality adapters or attention logits.</p>\n<p>Evaluate text-only, image-only, audio-only, and interleaved inputs separately. Include:</p>\n<ul>\n<li>small text in images;</li>\n<li>low-contrast objects;</li>\n<li>accented and noisy speech;</li>\n<li>silence and non-speech audio;</li>\n<li>long media prefixes;</li>\n<li>repeated modality switches.</li>\n</ul>\n<p>Compare task metrics and intermediate outputs where possible. A single language benchmark cannot qualify a multimodal quantization.</p>\n<h2 id=\"thermal-behavior-changes-the-service-policy\">Thermal behavior changes the service policy</h2>\n<p>Short benchmark runs do not show sustained device behavior. Heat can reduce clock frequency and make a formerly valid real-time policy fall behind. Run capture, encoding, and decoding together for the expected session length. Record throughput, latency, temperature or platform thermal state, and battery or power telemetry exposed by the device.</p>\n<p>Degradation should be explicit:</p>\n<ol>\n<li>reduce video sampling frequency;</li>\n<li>reduce image resolution or crop count;</li>\n<li>shorten retained media context;</li>\n<li>switch to a nested or smaller model when available;</li>\n<li>pause background interpretation while preserving direct user requests.</li>\n</ol>\n<p>Gemma 3n uses MatFormer nesting and per-layer embeddings so its released variants can operate with active memory footprints associated with smaller nested models.<sup><a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fn-gemma-preview\" id=\"user-content-fnref-gemma-preview\">2</a></sup> That capability depends on its architecture. Other models need separate checkpoints or different elastic mechanisms.</p>\n<h2 id=\"privacy-is-a-data-flow-property\">Privacy is a data-flow property</h2>\n<p>Local inference does not prove that media stays local. Crash reporters, analytics payloads, debug logging, model downloads, and remote fallbacks can transmit derived or raw data. Document every boundary and make fallback behavior visible to the user.</p>\n<p>Encrypt cached media if it must persist. Prefer ephemeral buffers. Strip prompt and media content from default logs. Verify that application lifecycle transitions release microphone and camera access. Test airplane-mode operation if offline behavior is a requirement.</p>\n<p>On-device multimodal design starts with a peak-memory timeline and a token budget. Encoders compress media into language-model inputs, but those inputs still consume prefill time and KV cache. The reliable implementation schedules encoders deliberately, validates preprocessing against a reference, measures delegate coverage, and reduces work predictably under thermal or memory pressure.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-gemma\">\n<p><a href=\"https://developers.googleblog.com/en/introducing-gemma-3n-developer-guide/\" rel=\"noopener noreferrer\">Google Developers Blog: Gemma 3n developer guide</a>. <a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fnref-gemma\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fnref-gemma-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fnref-gemma-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fnref-gemma-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a></p>\n</li>\n<li id=\"user-content-fn-gemma-preview\">\n<p><a href=\"https://developers.googleblog.com/en/introducing-gemma-3n\" rel=\"noopener noreferrer\">Google Developers Blog: Gemma 3n preview architecture</a>. <a href=\"https://blog.ecitis.org/on-device-multimodal-ai/#user-content-fnref-gemma-preview\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/on-device-multimodal-ai/",
            "title": "On-Device Multimodal AI: Text, Vision, and Audio Under a Memory Budget",
            "summary": "Schedule text, vision, and audio pipelines within the shared memory, thermal, energy, and latency limits of edge devices.",
            "image": "https://blog.ecitis.org/open-graph/on-device-multimodal-ai.png",
            "date_modified": "2026-05-08T00:00:00.000Z",
            "date_published": "2026-05-08T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "Multimodal",
                "On-device",
                "Memory"
            ]
        },
        {
            "id": "https://blog.ecitis.org/retrieval-systems/",
            "content_html": "<p>A retrieval system has a narrower job than a language model: given a query, return evidence that satisfies an information need. That boundary makes retrieval measurable. Documents can be judged for relevance before generation is introduced, and failures can be assigned to indexing, candidate generation, ranking, context construction, or answer synthesis.</p>\n<p>Basic retrieval-augmented generation often collapses those stages into one embedding search. Production retrieval needs more structure because identifiers, exact phrases, concepts, freshness, access control, and document quality are different ranking signals.</p>\n<h2 id=\"index-documents-around-retrieval-units\">Index documents around retrieval units</h2>\n<p>Chunking defines the unit the ranker can return. A chunk that spans unrelated sections produces an embedding with mixed semantics. A chunk that is too narrow can omit the qualifier needed to interpret a match.</p>\n<p>Use structural boundaries before token counts:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> Passage</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  id</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  documentId</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  headingPath</span><span>:</span><span> string</span><span>[]</span></span>\n<span class=\"line\"><span>  body</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  updatedAt</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  acl</span><span>:</span><span> string</span><span>[]</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>function</span><span> passageText</span><span>(passage</span><span>:</span><span> Passage</span><span>)</span><span>:</span><span> string</span><span> {</span></span>\n<span class=\"line\"><span>  return</span><span> [</span><span>...</span><span>passage</span><span>.headingPath</span><span>,</span><span> passage</span><span>.body]</span><span>.join</span><span>(</span><span>'\\n'</span><span>)</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The heading path gives a dense encoder and lexical index the local subject. The document identifier supports deduplication after ranking. The access-control field allows filtering before any passage enters the candidate set.</p>\n<p>Store provenance that survives every stage:</p>\n<ul>\n<li>Stable document and passage identifiers.</li>\n<li>Source URI and content revision.</li>\n<li>Character or byte offsets into the source.</li>\n<li>Parser and embedding model versions.</li>\n<li>Access-control attributes.</li>\n<li>Timestamp used for freshness calculations.</li>\n</ul>\n<p>Without a revision identifier, an answer citation can point at content that changed after retrieval.</p>\n<h2 id=\"lexical-retrieval-preserves-exact-evidence\">Lexical retrieval preserves exact evidence</h2>\n<p>BM25 ranks documents using query term frequency, inverse document frequency, and length normalization. A common form is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"normal\">score</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>D</mi><mo separator=\"true\">,</mo><mi>Q</mi><mo stretchy=\"false\">)</mo><mo>=</mo><munder><mo>∑</mo><mrow><mi>q</mi><mo>∈</mo><mi>Q</mi></mrow></munder><mi mathvariant=\"normal\">IDF</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>q</mi><mo stretchy=\"false\">)</mo><mfrac><mrow><mi>f</mi><mo stretchy=\"false\">(</mo><mi>q</mi><mo separator=\"true\">,</mo><mi>D</mi><mo stretchy=\"false\">)</mo><mo stretchy=\"false\">(</mo><msub><mi>k</mi><mn>1</mn></msub><mo>+</mo><mn>1</mn><mo stretchy=\"false\">)</mo></mrow><mrow><mi>f</mi><mo stretchy=\"false\">(</mo><mi>q</mi><mo separator=\"true\">,</mo><mi>D</mi><mo stretchy=\"false\">)</mo><mo>+</mo><msub><mi>k</mi><mn>1</mn></msub><mo stretchy=\"false\">(</mo><mn>1</mn><mo>−</mo><mi>b</mi><mo>+</mo><mi>b</mi><mi mathvariant=\"normal\">∣</mi><mi>D</mi><mi mathvariant=\"normal\">∣</mi><mi mathvariant=\"normal\">/</mi><mi mathvariant=\"normal\">avgdl</mi><mo>⁡</mo><mo stretchy=\"false\">)</mo></mrow></mfrac></mrow></semantics></math>\n<p>Here, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>f</mi><mo stretchy=\"false\">(</mo><mi>q</mi><mo separator=\"true\">,</mo><mi>D</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> is the frequency of term <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>q</mi></mrow></semantics></math> in document <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>D</mi></mrow></semantics></math>. The parameters <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>k</mi><mn>1</mn></msub></mrow></semantics></math> and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>b</mi></mrow></semantics></math> control term saturation and length normalization. They are tuning parameters, not universal constants.</p>\n<p>Lexical retrieval is valuable for error codes, function names, chemical symbols, product identifiers, and quoted policy language. An embedding model can place semantically related text nearby while missing the exact token sequence that establishes relevance.</p>\n<p>Normalize carefully. Lowercasing may suit prose but can damage case-sensitive identifiers. Stemming can improve recall for natural language and corrupt source-code terms. Maintain field-specific analyzers instead of applying one text pipeline to every field.</p>\n<h2 id=\"dense-retrieval-supplies-semantic-candidates\">Dense retrieval supplies semantic candidates</h2>\n<p>A bi-encoder maps the query and passage independently:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>z</mi><mi>q</mi></msub><mo>=</mo><msub><mi>f</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mi>q</mi><mo stretchy=\"false\">)</mo><mo separator=\"true\">,</mo><mspace width=\"2em\"></mspace><msub><mi>z</mi><mi>d</mi></msub><mo>=</mo><msub><mi>g</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mi>d</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>The index retrieves passages with high cosine similarity or inner product. Independent encoding makes corpus vectors reusable, which is why dense retrieval works as a candidate generator.</p>\n<p>Dense retrieval inherits the encoder’s training distribution. The BEIR benchmark evaluated lexical, sparse, dense, late-interaction, and reranking systems across 18 datasets; its results found BM25 to be a robust baseline and showed that reranking and late interaction performed best on average in zero-shot settings, with higher computational cost.<sup><a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fn-beir\" id=\"user-content-fnref-beir\">1</a></sup></p>\n<p>That result argues against replacing lexical search by default. It supports evaluating multiple retriever families on the target corpus.</p>\n<h2 id=\"hybrid-fusion-combines-rankers\">Hybrid fusion combines rankers</h2>\n<p>Lexical and dense scores are usually not calibrated to the same range. Adding raw scores allows whichever system has the wider numeric distribution to dominate.</p>\n<p>Reciprocal Rank Fusion combines positions instead:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"normal\">RRF</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>d</mi><mo stretchy=\"false\">)</mo><mo>=</mo><munder><mo>∑</mo><mrow><mi>r</mi><mo>∈</mo><mi>R</mi></mrow></munder><mfrac><mn>1</mn><mrow><mi>k</mi><mo>+</mo><msub><mrow><mi mathvariant=\"normal\">rank</mi><mo>⁡</mo></mrow><mi>r</mi></msub><mo stretchy=\"false\">(</mo><mi>d</mi><mo stretchy=\"false\">)</mo></mrow></mfrac></mrow></semantics></math>\n<p>The original RRF study compared fusion methods and reported that reciprocal rank fusion outperformed the tested Condorcet and individual learning-to-rank approaches.<sup><a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fn-rrf\" id=\"user-content-fnref-rrf\">2</a></sup> The constant <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math> controls how sharply the formula rewards the top positions and should be selected on validation queries.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> Ranked</span><span> =</span><span> { id</span><span>:</span><span> string</span><span>; rank</span><span>:</span><span> number</span><span> }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>function</span><span> rrf</span><span>(lists</span><span>:</span><span> Ranked</span><span>[][]</span><span>,</span><span> rankConstant</span><span>:</span><span> number</span><span>)</span><span>:</span><span> Map</span><span>&lt;</span><span>string</span><span>,</span><span> number</span><span>&gt; {</span></span>\n<span class=\"line\"><span>  const</span><span> scores</span><span> =</span><span> new</span><span> Map</span><span>&lt;</span><span>string</span><span>,</span><span> number</span><span>&gt;()</span></span>\n<span class=\"line\"><span>  for</span><span> (</span><span>const</span><span> list</span><span> of</span><span> lists) {</span></span>\n<span class=\"line\"><span>    for</span><span> (</span><span>const</span><span> hit</span><span> of</span><span> list) {</span></span>\n<span class=\"line\"><span>      scores</span><span>.set</span><span>(</span><span>hit</span><span>.id</span><span>,</span><span> (</span><span>scores</span><span>.get</span><span>(</span><span>hit</span><span>.id) </span><span>??</span><span> 0</span><span>) </span><span>+</span><span> 1</span><span> /</span><span> (rankConstant </span><span>+</span><span> hit</span><span>.rank))</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>  return</span><span> scores</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The candidate depth for each ranker is an operational control. A deeper pool raises recall and increases fusion, filtering, and reranking work.</p>\n<figure><figcaption><strong>Candidate ranks through a hybrid retrieval pipeline</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/retrieval-systems-0.webp\" width=\"512\" height=\"373\" alt=\"Candidate ranks through a hybrid retrieval pipeline\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"reranking-spends-compute-on-a-bounded-set\">Reranking spends compute on a bounded set</h2>\n<p>A cross-encoder reads the query and candidate passage together. Joint attention can inspect exact query-passage interactions that a bi-encoder compresses into separate vectors.</p>\n<p>The architecture separates jobs:</p>\n<ol>\n<li>Lexical and dense retrievers maximize candidate recall.</li>\n<li>Fusion creates a bounded candidate pool.</li>\n<li>A cross-encoder improves ordering precision.</li>\n<li>A context builder selects passages under a token budget.</li>\n</ol>\n<p>Training data must resemble the candidate distribution at inference. HYRR trains rerankers on candidates from a hybrid lexical and dense retriever and reports robustness across first-stage retrieval systems on supervised and zero-shot passage tasks.<sup><a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fn-hyrr\" id=\"user-content-fnref-hyrr\">3</a></sup></p>\n<p>Reranking every document defeats the purpose. Its latency should scale with the candidate pool, not corpus size. Batch query-passage pairs on accelerators and log queue time separately from model execution.</p>\n<h2 id=\"late-interaction-retains-token-level-matching\">Late interaction retains token-level matching</h2>\n<p>Late-interaction systems encode query and document tokens independently, then aggregate token-level similarities. They occupy a point between bi-encoders and cross-encoders:</p>\n<ul>\n<li>More expressive matching than one vector per passage.</li>\n<li>Reusable document representations.</li>\n<li>Larger indexes and more scoring work than single-vector search.</li>\n<li>Less joint computation than a cross-encoder.</li>\n</ul>\n<p>This architecture is useful when a single passage contains several concepts and relevance depends on multiple local matches. It also complicates index storage and serving. Measure end-to-end latency at realistic corpus size rather than comparing model forward passes alone.</p>\n<h2 id=\"query-transformation-needs-a-grounding-boundary\">Query transformation needs a grounding boundary</h2>\n<p>Query rewriting can expand acronyms, resolve conversational references, or create alternate lexical forms. It can also remove a rare identifier that carried the user’s actual intent.</p>\n<p>Preserve the original query as one branch:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>original query ───────────────› lexical retrieval</span></span>\n<span class=\"line\"><span>       ├── normalized query ──› dense retrieval</span></span>\n<span class=\"line\"><span>       └── generated variant ─› optional retrieval branch</span></span></code></pre>\n<p>HyDE generates a hypothetical document, embeds it, and retrieves nearby real documents. The paper explicitly notes that the generated document is unreal and may contain false details; the corpus search provides the grounding step.<sup><a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fn-hyde\" id=\"user-content-fnref-hyde\">4</a></sup> A generated query or hypothetical document must never be presented as evidence.</p>\n<h2 id=\"filters-belong-before-context-assembly\">Filters belong before context assembly</h2>\n<p>Metadata filtering is part of retrieval correctness. Tenant, repository, geography, retention class, and authorization constraints should restrict candidates before content reaches the model.</p>\n<p>Do not retrieve globally and remove unauthorized passages after reranking. That design leaks protected content into ranker inputs, traces, caches, and timing behavior.</p>\n<p>Freshness is also domain-specific. A recent changelog can outrank an older overview for a current-version query. A historical audit may require the opposite. Encode freshness as an explicit feature controlled by query intent rather than a global recency boost.</p>\n<h2 id=\"evaluate-retrieval-without-generation\">Evaluate retrieval without generation</h2>\n<p>For a query with a set of relevant passages, common metrics answer different questions:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Metric</th><th>Question</th></tr></thead><tbody><tr><td>Recall at <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math></td><td>Did the candidate set contain relevant evidence?</td></tr><tr><td>Precision at <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math></td><td>How much of the returned set was relevant?</td></tr><tr><td>Mean reciprocal rank</td><td>How early was the first relevant result?</td></tr><tr><td>nDCG at <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math></td><td>Did graded relevance appear in a useful order?</td></tr></tbody></table></div>\n<p>For binary relevance:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"normal\">Recall@k</mi><mo>⁡</mo><mo>=</mo><mfrac><mrow><mi mathvariant=\"normal\">∣</mi><mtext>relevant</mtext><mo>∩</mo><mtext>top-k</mtext><mi mathvariant=\"normal\">∣</mi></mrow><mrow><mi mathvariant=\"normal\">∣</mi><mtext>relevant</mtext><mi mathvariant=\"normal\">∣</mi></mrow></mfrac></mrow></semantics></math>\n<p>Use passage-level labels for candidate and reranker evaluation, then add answer-level evaluation separately. If answer quality falls while retrieval metrics remain stable, inspect context ordering, citation binding, and generation.</p>\n<p>Evaluation queries should be grouped by failure-sensitive slices:</p>\n<ul>\n<li>Exact identifiers and quoted phrases.</li>\n<li>Paraphrases with no shared keywords.</li>\n<li>Questions requiring several passages.</li>\n<li>Time-sensitive queries.</li>\n<li>Negative queries with no supporting evidence.</li>\n<li>Permission boundaries.</li>\n<li>Documents with repeated boilerplate.</li>\n</ul>\n<p>Report confidence intervals or paired tests when comparing rankers. A single aggregate can hide a regression on exact-match queries behind a gain on semantic questions.</p>\n<h2 id=\"context-construction-is-a-ranking-stage\">Context construction is a ranking stage</h2>\n<p>The top-ranked passages are not automatically the best prompt. Near-duplicate chunks waste tokens. Several passages from one document can crowd out independent evidence. The relevant span can be buried inside a long passage.</p>\n<p>A context builder should:</p>\n<ol>\n<li>Remove exact and near duplicates.</li>\n<li>Group passages by source document.</li>\n<li>Prefer passages that add evidence not already selected.</li>\n<li>Preserve source identifiers around every passage.</li>\n<li>Stop at a measured token budget.</li>\n</ol>\n<p>Long context does not remove this requirement. The Lost in the Middle study found that model performance can change with the position of relevant information and often declines when evidence appears in the middle of a long input.<sup><a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fn-middle\" id=\"user-content-fnref-middle\">5</a></sup></p>\n<figure><figcaption><strong>Context assembly spends a finite budget on distinct evidence</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/retrieval-systems-1.webp\" width=\"512\" height=\"286\" alt=\"Context assembly spends a finite budget on distinct evidence\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"operational-measurements\">Operational measurements</h2>\n<p>Trace each stage with the query and corpus versions needed for replay:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"query_id\"</span><span>:</span><span> \"q_7f3\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"index_revision\"</span><span>:</span><span> \"docs-2026-04-26\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"lexical_ms\"</span><span>:</span><span> 12</span><span>,</span></span>\n<span class=\"line\"><span>  \"dense_ms\"</span><span>:</span><span> 19</span><span>,</span></span>\n<span class=\"line\"><span>  \"fusion_candidates\"</span><span>:</span><span> 80</span><span>,</span></span>\n<span class=\"line\"><span>  \"rerank_ms\"</span><span>:</span><span> 37</span><span>,</span></span>\n<span class=\"line\"><span>  \"selected_passages\"</span><span>:</span><span> 6</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The values above illustrate a trace schema, not a performance target.</p>\n<p>Monitor:</p>\n<ul>\n<li>Empty-result rate and fallback rate.</li>\n<li>Candidate recall on labeled traffic samples.</li>\n<li>Reranker score drift by document type.</li>\n<li>Index freshness lag.</li>\n<li>Permission-filter rejection counts.</li>\n<li>Duplicate passage share.</li>\n<li>Stage latency distributions and queue time.</li>\n<li>Citation click-through or evidence inspection where the product exposes it.</li>\n</ul>\n<h2 id=\"failure-diagnosis\">Failure diagnosis</h2>\n<p><strong>Relevant passage absent from all candidates:</strong> inspect parsing, analyzers, embeddings, filters, and candidate depth.</p>\n<p><strong>Relevant passage retrieved but ranked low:</strong> inspect fusion, reranker training distribution, and query features.</p>\n<p><strong>Relevant passage ranks high but is omitted from the prompt:</strong> inspect deduplication and context-budget logic.</p>\n<p><strong>Evidence appears in the prompt but the answer is wrong:</strong> inspect ordering, answer instructions, citation binding, and model capability.</p>\n<p><strong>Offline relevance improves while production outcomes fall:</strong> inspect query distribution drift, latency, permissions, and the difference between relevance labels and the user outcome.</p>\n<p>A retrieval stack earns complexity only when each added stage moves a defined metric on representative queries. Hybrid retrieval, reranking, and context construction are separable controls. That separation makes the system faster to debug and harder to fool with one attractive aggregate score.</p>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-beir\">\n<p><a href=\"https://openreview.net/forum?id=wCu6T5xFjeJ\" rel=\"noopener noreferrer\">Thakur et al., “BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models,” NeurIPS Datasets and Benchmarks, 2021</a>. <a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fnref-beir\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-rrf\">\n<p><a href=\"https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf\" rel=\"noopener noreferrer\">Cormack, Clarke, and Buettcher, “Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods,” SIGIR, 2009</a>. <a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fnref-rrf\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-hyrr\">\n<p><a href=\"https://arxiv.org/abs/2212.10528\" rel=\"noopener noreferrer\">Lu et al., “HYRR: Hybrid Infused Reranking for Passage Retrieval,” 2022</a>. <a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fnref-hyrr\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-hyde\">\n<p><a href=\"https://aclanthology.org/2023.acl-long.99/\" rel=\"noopener noreferrer\">Gao et al., “Precise Zero-Shot Dense Retrieval without Relevance Labels,” ACL, 2023</a>. <a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fnref-hyde\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-middle\">\n<p><a href=\"https://aclanthology.org/2024.tacl-1.9/\" rel=\"noopener noreferrer\">Liu et al., “Lost in the Middle: How Language Models Use Long Contexts,” Transactions of the Association for Computational Linguistics, 2024</a>. <a href=\"https://blog.ecitis.org/retrieval-systems/#user-content-fnref-middle\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/retrieval-systems/",
            "title": "Retrieval Systems Beyond Basic RAG: Hybrid Search, Reranking, and Evaluation",
            "summary": "Build measurable retrieval pipelines with structural chunking, hybrid candidate generation, reranking, and relevance evaluation.",
            "image": "https://blog.ecitis.org/open-graph/retrieval-systems.png",
            "date_modified": "2026-04-27T00:00:00.000Z",
            "date_published": "2026-04-27T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "RAG",
                "Retrieval",
                "Reranking"
            ]
        },
        {
            "id": "https://blog.ecitis.org/distributed-llm-training/",
            "content_html": "<p>Large-model training is a placement problem before it is an optimizer problem. Parameters, optimizer state, gradients, activations, and temporary buffers must fit across devices. The selected placement also determines which bytes cross links on every step.</p>\n<figure><figcaption><strong>What each parallelism axis partitions</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/distributed-llm-training-0.webp\" width=\"512\" height=\"428\" alt=\"What each parallelism axis partitions\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"start-from-the-memory-equation\">Start from the memory equation</h2>\n<p>For parameter count <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>P</mi></mrow></semantics></math>, bytes per parameter <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>b</mi><mi>w</mi></msub></mrow></semantics></math>, gradient bytes <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>b</mi><mi>g</mi></msub></mrow></semantics></math>, and optimizer-state bytes <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>b</mi><mi>o</mi></msub></mrow></semantics></math>, persistent training state is approximately:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>M</mi><mtext>persistent</mtext></msub><mo>=</mo><mi>P</mi><mo stretchy=\"false\">(</mo><msub><mi>b</mi><mi>w</mi></msub><mo>+</mo><msub><mi>b</mi><mi>g</mi></msub><mo>+</mo><msub><mi>b</mi><mi>o</mi></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>Activations, communication buffers, allocator fragmentation, and kernel workspaces sit on top. Mixed precision does not imply that every term uses the same dtype. Many training stacks retain higher-precision master weights or optimizer statistics.</p>\n<p>Parallelism changes which terms each rank owns. It also introduces communication. A valid strategy must satisfy both device capacity and step-time constraints.</p>\n<h2 id=\"data-parallelism-replicates-compute\">Data parallelism replicates compute</h2>\n<p>Each data-parallel rank holds the model and processes a different microbatch. Gradients are combined, commonly with all-reduce. With global batch <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>B</mi></mrow></semantics></math>, data-parallel degree <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>D</mi></mrow></semantics></math>, and local batch <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>B</mi><mi mathvariant=\"normal\">ℓ</mi></msub></mrow></semantics></math>:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>B</mi><mo>=</mo><mi>D</mi><msub><mi>B</mi><mi mathvariant=\"normal\">ℓ</mi></msub></mrow></semantics></math>\n<p>The arithmetic per sample is unchanged. Communication scales with gradient state rather than activation size. This makes data parallelism attractive while a full training replica fits.</p>\n<p>Gradient accumulation decouples microbatch size from the optimizer batch. It reduces activation residency per forward pass but does not reduce the parameter or optimizer footprint.</p>\n<p>Fully sharded data parallel methods partition parameters, gradients, and optimizer state across the data-parallel group, gathering parameter shards when a layer executes and discarding or resharding them afterward. ZeRO introduced staged partitioning of optimizer state, gradients, and parameters to remove data-parallel memory redundancy.<sup><a href=\"https://blog.ecitis.org/distributed-llm-training/#user-content-fn-zero\" id=\"user-content-fnref-zero\">1</a></sup></p>\n<p>Failure modes include:</p>\n<ul>\n<li>all-gather latency exposed because computation is too small to hide it;</li>\n<li>uneven parameter groups causing transient memory peaks;</li>\n<li>gradient accumulation configured differently across ranks;</li>\n<li>checkpoint code assuming each rank owns a complete state dictionary.</li>\n</ul>\n<h2 id=\"tensor-parallelism-partitions-operators\">Tensor parallelism partitions operators</h2>\n<p>Tensor parallelism splits a matrix multiplication across devices. For a linear layer <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>Y</mi><mo>=</mo><mi>X</mi><mi>W</mi></mrow></semantics></math>, column partitioning divides <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>W</mi></mrow></semantics></math> by output features:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>W</mi><mo>=</mo><mo stretchy=\"false\">[</mo><msub><mi>W</mi><mn>1</mn></msub><mo separator=\"true\">,</mo><msub><mi>W</mi><mn>2</mn></msub><mo separator=\"true\">,</mo><mo>…</mo><mo separator=\"true\">,</mo><msub><mi>W</mi><mi>T</mi></msub><mo stretchy=\"false\">]</mo></mrow></semantics></math>\n<p>Each rank computes <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>X</mi><msub><mi>W</mi><mi>i</mi></msub></mrow></semantics></math>. The result is sharded along its feature dimension. A later row-parallel layer can consume those shards and reduce partial outputs. Megatron-LM arranged column- and row-parallel transformer operators to limit synchronization points while preserving dense matrix multiplications.<sup><a href=\"https://blog.ecitis.org/distributed-llm-training/#user-content-fn-megatron\" id=\"user-content-fnref-megatron\">2</a></sup></p>\n<p>Tensor parallelism reduces per-rank parameter and compute load, but introduces collectives inside every transformer layer. It belongs within the fastest network domain available. Crossing a slow inter-node boundary for each layer can dominate step time.</p>\n<p>Sequence parallelism complements tensor parallelism by sharding operations whose activations would otherwise remain replicated across the sequence axis. It targets activation memory, not just parameter memory.</p>\n<p>Check tensor-parallel configurations for:</p>\n<ul>\n<li>head counts and hidden widths divisible by the partition degree;</li>\n<li>identical collective ordering on every rank;</li>\n<li>dropout randomness that follows the intended replicated or sharded semantics;</li>\n<li>numerical drift from changing reduction order;</li>\n<li>network topology that matches the process group.</li>\n</ul>\n<h2 id=\"pipeline-parallelism-partitions-layers\">Pipeline parallelism partitions layers</h2>\n<p>Pipeline parallelism assigns consecutive layer ranges to stages. A minibatch is divided into microbatches so stages can work concurrently. GPipe formalized this approach with microbatch pipelining and rematerialization for large networks.<sup><a href=\"https://blog.ecitis.org/distributed-llm-training/#user-content-fn-gpipe\" id=\"user-content-fnref-gpipe\">3</a></sup></p>\n<p>An all-forward-then-all-backward schedule has simple dependencies but retains activations for many microbatches. A one-forward-one-backward schedule starts backward work earlier and reduces the number of live activations. Interleaved schedules assign multiple model chunks to a physical stage to reduce idle periods, at the cost of more transfers and scheduling complexity.</p>\n<p>Pipeline efficiency depends on balance. Equal layer counts do not imply equal stage times because embedding, attention, mixture-of-experts, and output layers have different compute and memory profiles. Measure stage latency and move whole blocks or split points based on traces.</p>\n<p>The pipeline boundary transmits activations in the forward direction and their gradients in the reverse direction. A boundary at a wide hidden representation consumes more bandwidth than one after a compact projection.</p>\n<p>Operational failures include a last stage slowed by vocabulary projection, microbatch count too small to fill the pipeline, activation checkpointing that recomputes across a communication boundary, and batches with different sequence lengths producing stage imbalance.</p>\n<figure><figcaption><strong>Pipeline work moves through stages in dependency order</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/distributed-llm-training-1.webp\" width=\"512\" height=\"279\" alt=\"Pipeline work moves through stages in dependency order\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"expert-parallelism-partitions-sparse-feed-forward-capacity\">Expert parallelism partitions sparse feed-forward capacity</h2>\n<p>A mixture-of-experts layer routes each token to selected expert networks. Expert parallelism places different experts on different ranks. Tokens are exchanged so each expert receives its assigned inputs, then returned to their original sequence positions.</p>\n<p>Switch Transformers demonstrated sparse expert routing in which only selected parameters execute for a token.<sup><a href=\"https://blog.ecitis.org/distributed-llm-training/#user-content-fn-switch\" id=\"user-content-fnref-switch\">4</a></sup> The memory benefit is conditional capacity: total parameters can grow without executing every expert for every token. The systems cost is token dispatch, usually implemented with all-to-all communication.</p>\n<p>Load balance is not guaranteed. A router can send too many tokens to a subset of experts. Capacity limits, auxiliary balancing losses, token dropping policies, and expert replication change quality and throughput. Each must be logged, not hidden behind an average utilization number.</p>\n<p>Measure per expert:</p>\n<ul>\n<li>assigned tokens before capacity enforcement;</li>\n<li>accepted and dropped tokens;</li>\n<li>dispatch bytes and collective duration;</li>\n<li>compute duration;</li>\n<li>routing probability distribution.</li>\n</ul>\n<h2 id=\"compose-parallel-dimensions-around-topology\">Compose parallel dimensions around topology</h2>\n<p>Production training combines dimensions. Let <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>D</mi></mrow></semantics></math>, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>T</mi></mrow></semantics></math>, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>L</mi></mrow></semantics></math>, and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>E</mi></mrow></semantics></math> denote data, tensor, pipeline, and expert parallel degrees. The world size is often factored as:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>N</mi><mo>=</mo><mi>D</mi><mo>×</mo><mi>T</mi><mo>×</mo><mi>L</mi><mo>×</mo><mi>E</mi></mrow></semantics></math>\n<p>This equation is a process-grid description, not proof that a configuration is efficient. Some implementations share or nest groups differently.</p>\n<p>Map frequent, latency-sensitive collectives to fast links. Tensor-parallel collectives occur inside layers, so they usually remain within a tightly connected node. Data-parallel reductions occur on gradients and can span nodes with bucketed overlap. Pipeline traffic is point-to-point between adjacent stages. Expert dispatch is all-to-all and needs special attention to oversubscribed network fabrics.</p>\n<p>One reasonable placement order is:</p>\n<ol>\n<li>Choose tensor parallelism so a layer and its working set fit inside a fast-link group.</li>\n<li>Add pipeline stages until the model state fits across groups.</li>\n<li>Place experts within network domains that can sustain dispatch.</li>\n<li>Use remaining devices for data-parallel replicas.</li>\n</ol>\n<p>The order is a design procedure, not a universal optimum. Profile the actual architecture and cluster.</p>\n<h2 id=\"checkpointing-must-encode-the-process-grid\">Checkpointing must encode the process grid</h2>\n<p>A distributed checkpoint is not merely a set of files. It records how tensors are partitioned. Changing the tensor-parallel degree requires resharding matrix dimensions. Changing pipeline degree moves layer ownership. Expert changes alter expert identifiers and routing state. Optimizer state must follow the same transformations as its parameter.</p>\n<p>Store:</p>\n<ul>\n<li>global tensor names and shapes;</li>\n<li>shard offsets and partition axes;</li>\n<li>dtype and optimizer-state relation;</li>\n<li>parallel degrees and rank topology;</li>\n<li>random-number-generator state;</li>\n<li>data-loader position and sample identity;</li>\n<li>software and model configuration.</li>\n</ul>\n<p>Test restore under the same topology before testing topology conversion. Compare loss on a fixed batch before and after restoration.</p>\n<h2 id=\"a-configuration-example\">A configuration example</h2>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>parallelism</span><span>:</span></span>\n<span class=\"line\"><span>  data</span><span>:</span><span> ${DATA_PARALLEL_DEGREE}</span></span>\n<span class=\"line\"><span>  tensor</span><span>:</span><span> ${TENSOR_PARALLEL_DEGREE}</span></span>\n<span class=\"line\"><span>  pipeline</span><span>:</span><span> ${PIPELINE_PARALLEL_DEGREE}</span></span>\n<span class=\"line\"><span>  expert</span><span>:</span><span> ${EXPERT_PARALLEL_DEGREE}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>schedule</span><span>:</span></span>\n<span class=\"line\"><span>  microbatches</span><span>:</span><span> ${MICROBATCH_COUNT}</span></span>\n<span class=\"line\"><span>  activation_checkpointing</span><span>:</span><span> selective</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>checkpoint</span><span>:</span></span>\n<span class=\"line\"><span>  format</span><span>:</span><span> distributed</span></span>\n<span class=\"line\"><span>  save_rng_state</span><span>:</span><span> true</span></span>\n<span class=\"line\"><span>  save_data_position</span><span>:</span><span> true</span></span></code></pre>\n<p>Configuration validation should reject incompatible hidden widths, attention heads, expert counts, and world sizes before processes initialize.</p>\n<h2 id=\"measure-exposed-communication\">Measure exposed communication</h2>\n<p>Collective duration alone does not show the performance cost. Communication overlapped with useful compute is cheaper than an equally long collective on the critical path. Use a timeline and calculate exposed communication:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>T</mi><mtext>exposed</mtext></msub><mo>=</mo><msub><mi>T</mi><mtext>step</mtext></msub><mo>−</mo><msub><mi>T</mi><mtext>compute-only critical path</mtext></msub></mrow></semantics></math>\n<p>Also report tokens processed per unit time, model FLOP utilization under a documented accounting convention, peak allocated memory per rank, and straggler spread. An aggregate average can conceal one stage that limits the complete pipeline.</p>\n<p>Distributed training succeeds when partitioning, scheduling, and topology agree. Data parallelism distributes samples, tensor parallelism distributes operators, pipeline parallelism distributes layers, and expert parallelism distributes conditional modules. Their communication patterns are different enough that treating them as interchangeable memory switches leads to slow or unstable runs.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-zero\">\n<p><a href=\"https://arxiv.org/abs/1910.02054\" rel=\"noopener noreferrer\">ZeRO: Memory Optimizations Toward Training Trillion Parameter Models</a>. <a href=\"https://blog.ecitis.org/distributed-llm-training/#user-content-fnref-zero\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-megatron\">\n<p><a href=\"https://arxiv.org/abs/1909.08053\" rel=\"noopener noreferrer\">Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism</a>. <a href=\"https://blog.ecitis.org/distributed-llm-training/#user-content-fnref-megatron\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-gpipe\">\n<p><a href=\"https://arxiv.org/abs/1811.06965\" rel=\"noopener noreferrer\">GPipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism</a>. <a href=\"https://blog.ecitis.org/distributed-llm-training/#user-content-fnref-gpipe\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-switch\">\n<p><a href=\"https://arxiv.org/abs/2101.03961\" rel=\"noopener noreferrer\">Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity</a>. <a href=\"https://blog.ecitis.org/distributed-llm-training/#user-content-fnref-switch\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/distributed-llm-training/",
            "title": "Distributed LLM Training: Data, Tensor, Pipeline, and Expert Parallelism",
            "summary": "Place model state and computation across devices with data, tensor, pipeline, sequence, and expert parallelism.",
            "image": "https://blog.ecitis.org/open-graph/distributed-llm-training.png",
            "date_modified": "2026-04-15T00:00:00.000Z",
            "date_published": "2026-04-15T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Distributed training",
                "Parallelism",
                "GPU clusters"
            ]
        },
        {
            "id": "https://blog.ecitis.org/modern-attention-kernels/",
            "content_html": "<p>Attention is often introduced as three matrix operations: form <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>Q</mi><msup><mi>K</mi><mi>T</mi></msup></mrow></semantics></math>, apply softmax, then multiply by <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>V</mi></mrow></semantics></math>. That description omits the storage hierarchy. A conventional implementation writes the score matrix to high-bandwidth memory, reads it for softmax, writes probabilities, and reads them again for the value product.</p>\n<p>FlashAttention computes exact attention while avoiding materialization of the full score matrix. Its central optimization is data movement: tile operands into on-chip memory, update softmax statistics online, and write only the final output.<sup><a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup></p>\n<h2 id=\"standard-attention-materializes-quadratic-state\">Standard attention materializes quadratic state</h2>\n<p>For query and key sequence lengths <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>N</mi><mi>q</mi></msub></mrow></semantics></math> and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>N</mi><mi>k</mi></msub></mrow></semantics></math>, the score matrix has <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>N</mi><mi>q</mi></msub><msub><mi>N</mi><mi>k</mi></msub></mrow></semantics></math> entries:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>S</mi><mo>=</mo><mfrac><mrow><mi>Q</mi><msup><mi>K</mi><mi>T</mi></msup></mrow><msqrt><mi>d</mi></msqrt></mfrac><mo separator=\"true\">,</mo><mspace width=\"2em\"></mspace><mi>P</mi><mo>=</mo><mi mathvariant=\"normal\">softmax</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>S</mi><mo stretchy=\"false\">)</mo><mo separator=\"true\">,</mo><mspace width=\"2em\"></mspace><mi>O</mi><mo>=</mo><mi>P</mi><mi>V</mi></mrow></semantics></math>\n<p>Even when arithmetic fits the accelerator, storing and rereading <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>S</mi></mrow></semantics></math> and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>P</mi></mrow></semantics></math> creates off-chip traffic proportional to their size. The original FlashAttention paper frames the algorithm as IO-aware and analyzes reads and writes across GPU memory levels.<sup><a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fn-1\" id=\"user-content-fnref-1-2\">1</a></sup></p>\n<figure><figcaption><strong>Attention data movement with and without score materialization</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/modern-attention-kernels-0.webp\" width=\"512\" height=\"390\" alt=\"Attention data movement with and without score materialization\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"online-softmax-makes-tiling-exact\">Online softmax makes tiling exact</h2>\n<p>Softmax appears to require an entire row because its denominator includes every key:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"normal\">softmax</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><msub><mi>s</mi><mi>i</mi></msub><msub><mo stretchy=\"false\">)</mo><mi>j</mi></msub><mo>=</mo><mfrac><msup><mi>e</mi><msub><mi>s</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></msup><mrow><munder><mo>∑</mo><mi>k</mi></munder><msup><mi>e</mi><msub><mi>s</mi><mrow><mi>i</mi><mi>k</mi></mrow></msub></msup></mrow></mfrac></mrow></semantics></math>\n<p>The computation can be updated block by block. For a row, keep running maximum <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>m</mi></mrow></semantics></math>, running normalization sum <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi mathvariant=\"normal\">ℓ</mi></mrow></semantics></math>, and an unnormalized output accumulator <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>o</mi></mrow></semantics></math>. Given a new score tile with maximum <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>m</mi><mi>b</mi></msub></mrow></semantics></math>:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msup><mi>m</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup><mo>=</mo><mi>max</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>m</mi><mo separator=\"true\">,</mo><msub><mi>m</mi><mi>b</mi></msub><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msup><mi mathvariant=\"normal\">ℓ</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup><mo>=</mo><msup><mi>e</mi><mrow><mi>m</mi><mo>−</mo><msup><mi>m</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup></mrow></msup><mi mathvariant=\"normal\">ℓ</mi><mo>+</mo><munder><mo>∑</mo><mi>j</mi></munder><msup><mi>e</mi><mrow><msub><mi>s</mi><mi>j</mi></msub><mo>−</mo><msup><mi>m</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup></mrow></msup></mrow></semantics></math>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msup><mi>o</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup><mo>=</mo><msup><mi>e</mi><mrow><mi>m</mi><mo>−</mo><msup><mi>m</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup></mrow></msup><mi>o</mi><mo>+</mo><munder><mo>∑</mo><mi>j</mi></munder><msup><mi>e</mi><mrow><msub><mi>s</mi><mi>j</mi></msub><mo>−</mo><msup><mi>m</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup></mrow></msup><msub><mi>v</mi><mi>j</mi></msub></mrow></semantics></math>\n<p>After the final key tile, output <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>o</mi><mi mathvariant=\"normal\">/</mi><mi mathvariant=\"normal\">ℓ</mi></mrow></semantics></math>. Rescaling by the new maximum preserves numerical stability. No approximation is introduced by the tiling rule.</p>\n<h2 id=\"kernel-structure\">Kernel structure</h2>\n<p>A simplified forward kernel loops over query tiles and key-value tiles:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>for each Q tile:</span></span>\n<span class=\"line\"><span>    load Q into on-chip memory</span></span>\n<span class=\"line\"><span>    initialize row maxima, sums, output accumulators</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    for each K,V tile allowed by the mask:</span></span>\n<span class=\"line\"><span>        load K and V</span></span>\n<span class=\"line\"><span>        compute score tile</span></span>\n<span class=\"line\"><span>        apply scale and mask</span></span>\n<span class=\"line\"><span>        update online softmax state</span></span>\n<span class=\"line\"><span>        accumulate weighted V</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    normalize and write output</span></span></code></pre>\n<p>Tile shapes depend on head dimension, element type, register pressure, on-chip memory, and accelerator generation. Larger tiles increase reuse but can reduce occupancy when each program consumes too many registers.</p>\n<p>Causal attention skips tiles wholly above the diagonal and masks the intersecting boundary tile. Local attention changes the visited tile range. Arbitrary sparse masks need an index structure that tells the kernel which blocks exist.</p>\n<h2 id=\"backward-pass-recomputes-scores\">Backward pass recomputes scores</h2>\n<p>Saving the full probability matrix would reintroduce quadratic storage. FlashAttention saves compact normalization statistics and recomputes score tiles during backpropagation.<sup><a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fn-1\" id=\"user-content-fnref-1-3\">1</a></sup> It spends arithmetic to avoid memory traffic and activation storage.</p>\n<p>This trade is useful only when recomputation is cheaper than writing and reading the intermediate. Profiling must include the whole layer because fused dropout, bias, and masking can change the balance.</p>\n<h2 id=\"work-partitioning-in-flashattention-2\">Work partitioning in FlashAttention-2</h2>\n<p>FlashAttention-2 changes work partitioning to improve parallelism and reduce non-matrix-multiplication work.<sup><a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> It parallelizes across sequence length when batch size and head count do not expose enough thread blocks, and adjusts how work is shared within a thread block.<sup><a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup></p>\n<p>The lesson is broader than a version comparison. An asymptotically efficient algorithm can leave hardware idle if its grid has too few independent programs. Kernel selection needs batch, head count, sequence length, and head dimension.</p>\n<h2 id=\"prefill-and-decode-need-different-kernels\">Prefill and decode need different kernels</h2>\n<p>Prefill has many query positions and many key positions. Tiled matrix multiplication works well. Decode often has one query position per sequence and a long cache. It cannot reuse query work across a large query tile.</p>\n<p>Decode attention therefore focuses on:</p>\n<ul>\n<li>coalesced cache reads</li>\n<li>parallel reduction across the context</li>\n<li>grouped-query head sharing</li>\n<li>cache block indirection</li>\n<li>combining many sequences in a batch</li>\n</ul>\n<p>Calling both implementations FlashAttention hides this shape difference. Record kernel choice separately for prefill and decode.</p>\n<h2 id=\"paged-attention-changes-addressing\">Paged attention changes addressing</h2>\n<p>PagedAttention stores cache blocks in non-contiguous physical memory and resolves them through a per-sequence block table. The vLLM paper reports near-zero cache waste and flexible cache sharing from this design.<sup><a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup></p>\n<p>A paged decode kernel maps logical token position to:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>logical block = position / block_size</span></span>\n<span class=\"line\"><span>offset        = position % block_size</span></span>\n<span class=\"line\"><span>physical      = block_table[sequence, logical_block]</span></span>\n<span class=\"line\"><span>address       = cache_base + physical * block_stride + offset * token_stride</span></span></code></pre>\n<p>The extra loads and address arithmetic buy dynamic allocation. Performance depends on block-table locality, block layout, and vectorized loads from physical pages.</p>\n<p>Paged attention and FlashAttention solve related but distinct problems:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Technique</th><th>Primary target</th><th>Key mechanism</th></tr></thead><tbody><tr><td>FlashAttention</td><td>Score-matrix IO</td><td>Tiling and online softmax</td></tr><tr><td>PagedAttention</td><td>Dynamic KV allocation</td><td>Logical-to-physical block mapping</td></tr><tr><td>Grouped-query attention</td><td>KV cache and projection size</td><td>Share KV heads across query groups</td></tr><tr><td>Sparse attention</td><td>Number of attended pairs</td><td>Visit selected score blocks</td></tr></tbody></table></div>\n<p>A serving engine can combine them: tiled prefill, paged decode, grouped-query cache layout, and sparse visitation for long contexts.</p>\n<h2 id=\"numerical-validation\">Numerical validation</h2>\n<p>Fused kernels reorder floating-point operations. Validate against a reference implementation across:</p>\n<ul>\n<li>causal and non-causal masks</li>\n<li>ragged sequence boundaries</li>\n<li>head dimensions supported by each path</li>\n<li>mixed precision and accumulation precision</li>\n<li>long rows with large logit ranges</li>\n<li>grouped-query head mapping</li>\n<li>paged blocks that cross sequence tails</li>\n</ul>\n<p>Use absolute and relative error tolerances tied to output precision. Also compare gradients for training kernels and sampled outputs for inference pipelines.</p>\n<h2 id=\"failure-modes\">Failure modes</h2>\n<p><strong>Register spilling:</strong> a larger tile silently moves local values to off-chip memory. Inspect compiler resource reports and achieved occupancy.</p>\n<p><strong>Mask boundary errors:</strong> the fast path reads future tokens or padding on partial tiles. Test every boundary around tile and page sizes.</p>\n<p><strong>Incorrect online rescaling:</strong> output accumulators are not rescaled when a later tile raises the row maximum. Compare adversarial logits where the largest score appears late.</p>\n<p><strong>Page-table mismatch:</strong> allocator and kernel disagree about block stride or layout version. Encode layout identity in engine metadata.</p>\n<p><strong>Wrong kernel for decode:</strong> a prefill-optimized kernel launches too little useful work for single-query attention. Profile by phase and shape.</p>\n<h2 id=\"benchmarking-method\">Benchmarking method</h2>\n<p>Measure wall time, achieved bandwidth, arithmetic utilization, kernel launches, and temporary bytes. Sweep the real shape space rather than one square sequence:</p>\n<ul>\n<li>batch size</li>\n<li>query length</li>\n<li>key length</li>\n<li>head count and KV-head count</li>\n<li>head dimension</li>\n<li>causal, local, and arbitrary masks</li>\n<li>cache block size</li>\n<li>element and accumulation type</li>\n</ul>\n<p>Warm kernels, pin clocks where the environment allows it, and report medians plus tail behavior. End-to-end token latency remains the deciding metric because a fast attention kernel can expose another bottleneck in projections, collectives, sampling, or cache allocation.</p>\n<p>Modern attention kernels are memory-system programs expressed through tensor algebra. Their speed comes from deciding what never reaches off-chip memory, what is recomputed, and how logical sequence state maps onto physical storage.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Tri Dao et al., “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” NeurIPS, 2022. <a href=\"https://arxiv.org/abs/2205.14135\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2205.14135</a> <a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fnref-1-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fnref-1-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Tri Dao, “FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning,” ICLR, 2024. <a href=\"https://arxiv.org/abs/2307.08691\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2307.08691</a> <a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Woosuk Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention,” SOSP, 2023. <a href=\"https://arxiv.org/abs/2309.06180\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2309.06180</a> <a href=\"https://blog.ecitis.org/modern-attention-kernels/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/modern-attention-kernels/",
            "title": "Modern Attention Kernels: Tiling, FlashAttention, and Paged Attention",
            "summary": "Follow how tiling, online softmax, and paged memory reduce data movement in modern training and inference kernels.",
            "image": "https://blog.ecitis.org/open-graph/modern-attention-kernels.png",
            "date_modified": "2026-04-04T00:00:00.000Z",
            "date_published": "2026-04-04T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "FlashAttention",
                "Kernels",
                "GPU memory"
            ]
        },
        {
            "id": "https://blog.ecitis.org/speculative-decoding/",
            "content_html": "<p>Autoregressive decoding normally invokes the target model once for each output token. Speculative decoding asks a cheaper draft process to propose several tokens, then evaluates those positions together with the target model.</p>\n<p>The target model remains authoritative. A correct acceptance rule preserves its output distribution; speculation changes the execution schedule, not the distribution being sampled. Independent work by Leviathan, Kalman, and Matias and by Chen and collaborators established this lossless sampling construction.<sup><a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup><sup><a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup></p>\n<h2 id=\"draft-and-target-distributions\">Draft and target distributions</h2>\n<p>Let <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>q</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> be the draft probability for a proposed token and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>p</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> the target probability at the same position. The verifier accepts draft token <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>x</mi></mrow></semantics></math> with:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>a</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mi>min</mi><mo>⁡</mo><mrow><mo fence=\"true\">(</mo><mn>1</mn><mo separator=\"true\">,</mo><mfrac><mrow><mi>p</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><mi>q</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo fence=\"true\">)</mo></mrow></mrow></semantics></math>\n<p>If the token is rejected, the correction distribution is proportional to:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msup><mi>p</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mi>max</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mn>0</mn><mo separator=\"true\">,</mo><mi>p</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo><mo>−</mo><mi>q</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>Normalizing and sampling from this residual corrects the draft’s over-allocation of probability. If every proposal is accepted, the verifier can also sample an additional target token from the final verified position.<sup><a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fn-1\" id=\"user-content-fnref-1-2\">1</a></sup><sup><a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup></p>\n<p>This is not the same as accepting whenever the target model’s argmax matches the draft. Greedy matching can preserve greedy output, but it does not implement sampling from the target distribution.</p>\n<figure><figcaption><strong>Draft tokens are verified in one target-model pass</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/speculative-decoding-0.webp\" width=\"512\" height=\"297\" alt=\"Draft tokens are verified in one target-model pass\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"verification-uses-parallel-positions\">Verification uses parallel positions</h2>\n<p>Suppose the draft model proposes <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math> tokens. The target receives the original prefix plus those candidates and returns logits for each proposed position in one forward pass. Causal masking prevents a position from reading later candidates, so each target distribution matches the distribution that ordinary autoregressive decoding would compute at that prefix.</p>\n<p>The execution loop is:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>while</span><span> not</span><span> finished</span><span>:</span></span>\n<span class=\"line\"><span>    proposals</span><span>,</span><span> q_probs </span><span>=</span><span> draft</span><span>(prefix, draft_length)</span></span>\n<span class=\"line\"><span>    p_probs </span><span>=</span><span> target</span><span>(prefix </span><span>+</span><span> proposals)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    accepted </span><span>=</span><span> []</span></span>\n<span class=\"line\"><span>    for</span><span> token</span><span>,</span><span> q</span><span>,</span><span> p </span><span>in</span><span> zip</span><span>(proposals, q_probs, p_probs):</span></span>\n<span class=\"line\"><span>        if</span><span> uniform</span><span>()</span><span> &lt;=</span><span> min</span><span>(</span><span>1.0</span><span>, p[token] </span><span>/</span><span> q[token]):</span></span>\n<span class=\"line\"><span>            accepted</span><span>.</span><span>append</span><span>(token)</span></span>\n<span class=\"line\"><span>        else</span><span>:</span></span>\n<span class=\"line\"><span>            accepted</span><span>.</span><span>append</span><span>(</span><span>sample</span><span>(</span><span>normalize</span><span>(</span><span>maximum</span><span>(p </span><span>-</span><span> q, </span><span>0</span><span>))))</span></span>\n<span class=\"line\"><span>            break</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    if</span><span> len</span><span>(accepted)</span><span> ==</span><span> len</span><span>(proposals):</span></span>\n<span class=\"line\"><span>        accepted</span><span>.</span><span>append</span><span>(</span><span>sample</span><span>(p_probs.after_last_candidate))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    prefix</span><span>.</span><span>extend</span><span>(accepted)</span></span></code></pre>\n<p>Real implementations must use the exact distributions after temperature, top-<math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>k</mi></mrow></semantics></math>, top-<math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>p</mi></mrow></semantics></math>, penalties, and any token masks. Applying a transform to only one model invalidates the acceptance ratio.</p>\n<h2 id=\"speed-depends-on-accepted-tokens-per-target-pass\">Speed depends on accepted tokens per target pass</h2>\n<p>Speculation trades draft work and a wider target pass for fewer sequential target invocations. The useful quantity is not draft accuracy alone. Measure accepted output tokens per verification pass.</p>\n<p>Let <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>C</mi><mi>T</mi></msub><mo stretchy=\"false\">(</mo><mi>k</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> be target verification cost for a candidate span, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>C</mi><mi>D</mi></msub><mo stretchy=\"false\">(</mo><mi>k</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> draft cost, and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>E</mi><mo stretchy=\"false\">[</mo><mi>A</mi><mo stretchy=\"false\">]</mo></mrow></semantics></math> expected committed tokens. A simple average cost is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mtext>cost per committed token</mtext><mo>=</mo><mfrac><mrow><msub><mi>C</mi><mi>D</mi></msub><mo stretchy=\"false\">(</mo><mi>k</mi><mo stretchy=\"false\">)</mo><mo>+</mo><msub><mi>C</mi><mi>T</mi></msub><mo stretchy=\"false\">(</mo><mi>k</mi><mo stretchy=\"false\">)</mo></mrow><mrow><mi>E</mi><mo stretchy=\"false\">[</mo><mi>A</mi><mo stretchy=\"false\">]</mo></mrow></mfrac></mrow></semantics></math>\n<p>The best draft length depends on batch size, prompt length, target architecture, draft architecture, and workload. Longer spans expose more parallel target work but risk wasting later candidates after an early rejection.</p>\n<p>An adaptive controller can choose draft length from recent acceptance:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>function</span><span> nextDraftLength</span><span>(emaAcceptance</span><span>:</span><span> number</span><span>,</span><span> current</span><span>:</span><span> number</span><span>) {</span></span>\n<span class=\"line\"><span>  if</span><span> (emaAcceptance </span><span>&gt;</span><span> highThreshold) </span><span>return</span><span> Math</span><span>.min</span><span>(current </span><span>+</span><span> step</span><span>,</span><span> maximum)</span></span>\n<span class=\"line\"><span>  if</span><span> (emaAcceptance </span><span>&lt;</span><span> lowThreshold) </span><span>return</span><span> Math</span><span>.max</span><span>(current </span><span>-</span><span> step</span><span>,</span><span> minimum)</span></span>\n<span class=\"line\"><span>  return</span><span> current</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Thresholds must come from serving measurements. Hard-coded values copied from another model pair do not transfer reliably.</p>\n<h2 id=\"draft-model-selection\">Draft model selection</h2>\n<p>A draft model must be cheaper and sufficiently aligned with the target. It also needs compatible tokenization. Different vocabularies require a mapping scheme and make probability correction more complex.</p>\n<p>Evaluate candidates on:</p>\n<ul>\n<li>target-pass reduction</li>\n<li>draft latency under production batch sizes</li>\n<li>acceptance by prompt class</li>\n<li>device memory consumed by draft weights and cache</li>\n<li>scheduler interference with target execution</li>\n<li>tokenizer and logits-processor compatibility</li>\n</ul>\n<p>A small model on the same accelerator can compete with the target for memory bandwidth. A draft on another device adds communication and synchronization. The fastest isolated draft is not necessarily the best colocated draft.</p>\n<h2 id=\"cache-management\">Cache management</h2>\n<p>Both models need state for the accepted prefix. Draft tokens beyond the rejection point must be rolled back. The target cache can retain verified accepted positions, while rejected suffix blocks become free.</p>\n<p>Paged caches simplify rollback when each speculative span receives provisional blocks. Commit accepted positions, truncate the final partial block, and release later blocks. Avoid copying the entire cache for each attempt.</p>\n<p>Continuous batching adds divergence. Requests accept different counts, so their logical sequence lengths change by different amounts after one verification iteration. The scheduler must update block tables and sampling state per request before forming the next batch.</p>\n<h2 id=\"sampling-details-that-break-correctness\">Sampling details that break correctness</h2>\n<p><strong>Mismatched temperature:</strong> draft and target probabilities refer to different transformed distributions.</p>\n<p><strong>Penalties applied at different prefixes:</strong> repetition or presence penalties see speculative tokens on one side but not the other.</p>\n<p><strong>Grammar masks applied only to the target:</strong> the draft assigns mass to illegal tokens, reducing acceptance and complicating residual sampling.</p>\n<p><strong>Rounded probabilities:</strong> low-precision acceptance ratios introduce bias. Keep probability arithmetic at a precision validated against a non-speculative sampler.</p>\n<p><strong>Random-number reuse:</strong> batched rollback changes random-number consumption. Counter-based generators indexed by request and logical sample step make replay easier.</p>\n<h2 id=\"tree-and-head-based-variants\">Tree and head-based variants</h2>\n<p>The draft need not be a separate autoregressive model. Multiple heads can propose future tokens from the target’s hidden state. Medusa adds decoding heads and verifies a tree of candidates, avoiding a separate draft model.<sup><a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup> Tree verification trades more candidate positions for a greater chance that one path matches the target.</p>\n<p>The verifier must construct a causal mask that represents the tree. Tokens on one branch cannot attend to siblings. Cache positions also need branch-specific ownership until a path is accepted.</p>\n<h2 id=\"production-evaluation\">Production evaluation</h2>\n<p>Report latency distributions and acceptance together:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Metric</th><th>Diagnostic value</th></tr></thead><tbody><tr><td>Accepted tokens per target pass</td><td>Direct measure of sequential work removed</td></tr><tr><td>Rejection position histogram</td><td>Detects spans that are too long</td></tr><tr><td>Draft time per proposed token</td><td>Captures draft bottlenecks</td></tr><tr><td>Verification time by candidate count</td><td>Reveals target-kernel scaling</td></tr><tr><td>Rollback bytes and blocks</td><td>Finds cache churn</td></tr><tr><td>Acceptance by request class</td><td>Supports adaptive routing</td></tr></tbody></table></div>\n<p>Compare against the same target model, sampling configuration, and scheduler without speculation. Verify distributional equivalence with seeded tests over many prompts, then run task-level quality evaluation to catch integration errors.</p>\n<p>Speculative decoding wins when cheap proposals align with the target and the hardware can verify several positions more efficiently than it can execute the same number of sequential decode steps. Its acceptance mathematics is exact; its systems benefit is empirical.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Yaniv Leviathan, Matan Kalman, and Yossi Matias, “Fast Inference from Transformers via Speculative Decoding,” ICML, 2023. <a href=\"https://arxiv.org/abs/2211.17192\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2211.17192</a> <a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fnref-1-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Charlie Chen et al., “Accelerating Large Language Model Decoding with Speculative Sampling,” arXiv, 2023. <a href=\"https://arxiv.org/abs/2302.01318\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2302.01318</a> <a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Tianle Cai et al., “Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads,” ICML, 2024. <a href=\"https://arxiv.org/abs/2401.10774\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2401.10774</a> <a href=\"https://blog.ecitis.org/speculative-decoding/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/speculative-decoding/",
            "title": "Speculative Decoding: Draft, Verify, and Accept",
            "summary": "Accelerate autoregressive generation with draft proposals and lossless target-model verification while preserving the output distribution.",
            "image": "https://blog.ecitis.org/open-graph/speculative-decoding.png",
            "date_modified": "2026-03-25T00:00:00.000Z",
            "date_published": "2026-03-25T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "Speculative decoding",
                "Sampling",
                "Latency"
            ]
        },
        {
            "id": "https://blog.ecitis.org/mcp-first-principles/",
            "content_html": "<p>Model Context Protocol is a stateful protocol for connecting an AI host to servers that expose context and actions. Its useful abstraction is not a universal tool API. It is a boundary: the host retains the conversation and policy decisions, while each client maintains an isolated connection to one server.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-architecture\" id=\"user-content-fnref-architecture\">1</a></sup></p>\n<p>That distinction determines where permissions, user consent, context selection, and model calls belong. A server describes what it offers. The host decides what enters the model context and which action requests are allowed to execute.</p>\n<h2 id=\"host-client-and-server-responsibilities\">Host, client, and server responsibilities</h2>\n<p>An MCP host creates clients, coordinates model integration, controls connection permissions, and aggregates context. Each client owns a stateful session with one server. A server exposes focused capabilities without receiving the host’s complete conversation or the state of other servers.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-architecture\" id=\"user-content-fnref-architecture-2\">1</a></sup></p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Process</th><th>Owns</th><th>Must not assume</th></tr></thead><tbody><tr><td>Host</td><td>Conversation, user consent, policy, model orchestration</td><td>That a server is safe because it speaks MCP</td></tr><tr><td>Client</td><td>Protocol session, message correlation, capability state</td><td>That an undeclared method is supported</td></tr><tr><td>Server</td><td>Tools, resources, prompts, server-side access checks</td><td>That a tool call proves user intent</td></tr></tbody></table></div>\n<p>This split prevents a calendar server from silently reading repository context supplied by a source-control server. Cross-server data movement is a host decision, not a protocol side effect.</p>\n<h2 id=\"json-rpc-message-forms\">JSON-RPC message forms</h2>\n<p>MCP messages use JSON-RPC. A request carries an <code>id</code> and expects either a result or an error. A notification has no <code>id</code> and receives no response. A response must contain a result or an error, never both.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-base\" id=\"user-content-fnref-base\">2</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"jsonrpc\"</span><span>:</span><span> \"2.0\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"id\"</span><span>:</span><span> \"req-42\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"method\"</span><span>:</span><span> \"tools/call\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"params\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"name\"</span><span>:</span><span> \"read_invoice\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"arguments\"</span><span>:</span><span> { </span><span>\"invoice_id\"</span><span>:</span><span> \"INV-1042\"</span><span> }</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The <code>id</code> is a correlation key. It is not an authorization token, idempotency key, or tracing policy. A host that retries a side-effecting request must define idempotency above the protocol or include an application-level key in the tool input.</p>\n<p>MCP’s current base specification does not support JSON-RPC batching. Each request, notification, and response travels as its own protocol message.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-changelog\" id=\"user-content-fnref-changelog\">3</a></sup></p>\n<figure><figcaption><strong>MCP session lifecycle and correlated JSON-RPC messages</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/mcp-first-principles-0.webp\" width=\"512\" height=\"679\" alt=\"MCP session lifecycle and correlated JSON-RPC messages\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"initialization-and-capability-negotiation\">Initialization and capability negotiation</h2>\n<p>Initialization must be the first client-server interaction. The client proposes a protocol version and declares client capabilities. The server returns its selected version and server capabilities. The client then sends <code>notifications/initialized</code> before normal operation.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-lifecycle\" id=\"user-content-fnref-lifecycle\">4</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"jsonrpc\"</span><span>:</span><span> \"2.0\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"id\"</span><span>:</span><span> 1</span><span>,</span></span>\n<span class=\"line\"><span>  \"method\"</span><span>:</span><span> \"initialize\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"params\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"protocolVersion\"</span><span>:</span><span> \"2025-06-18\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"capabilities\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>      \"roots\"</span><span>:</span><span> { </span><span>\"listChanged\"</span><span>:</span><span> true</span><span> }</span><span>,</span></span>\n<span class=\"line\"><span>      \"sampling\"</span><span>:</span><span> {}</span><span>,</span></span>\n<span class=\"line\"><span>      \"elicitation\"</span><span>:</span><span> {}</span></span>\n<span class=\"line\"><span>    }</span><span>,</span></span>\n<span class=\"line\"><span>    \"clientInfo\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>      \"name\"</span><span>:</span><span> \"research-workbench\"</span><span>,</span></span>\n<span class=\"line\"><span>      \"version\"</span><span>:</span><span> \"1.0.0\"</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The version and example fields above follow the published lifecycle example.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-lifecycle\" id=\"user-content-fnref-lifecycle-2\">4</a></sup> If the server selects a version the client does not support, the client should disconnect. For HTTP transport, later requests carry the negotiated version in the <code>MCP-Protocol-Version</code> header.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-transports\" id=\"user-content-fnref-transports\">5</a></sup></p>\n<p>Capabilities are executable protocol state, not descriptive metadata. A server may call client sampling only when the client advertised <code>sampling</code>. A server may emit resource update notifications only when it advertised the relevant resource subscription support. Method dispatch should check negotiated state before parsing business arguments.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> Session</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  phase</span><span>:</span><span> 'new'</span><span> |</span><span> 'operating'</span><span> |</span><span> 'closed'</span></span>\n<span class=\"line\"><span>  protocolVersion</span><span>?:</span><span> string</span></span>\n<span class=\"line\"><span>  serverCapabilities</span><span>?:</span><span> {</span></span>\n<span class=\"line\"><span>    tools</span><span>?:</span><span> { listChanged</span><span>?:</span><span> boolean</span><span> }</span></span>\n<span class=\"line\"><span>    resources</span><span>?:</span><span> { subscribe</span><span>?:</span><span> boolean</span><span>; listChanged</span><span>?:</span><span> boolean</span><span> }</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>function</span><span> assertToolCallAllowed</span><span>(session</span><span>:</span><span> Session</span><span>)</span><span>:</span><span> void</span><span> {</span></span>\n<span class=\"line\"><span>  if</span><span> (</span><span>session</span><span>.phase </span><span>!==</span><span> 'operating'</span><span> ||</span><span> !</span><span>session</span><span>.</span><span>serverCapabilities</span><span>?.tools) {</span></span>\n<span class=\"line\"><span>    throw</span><span> new</span><span> Error</span><span>(</span><span>'tools/call was not negotiated'</span><span>)</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h2 id=\"tools-resources-and-prompts\">Tools, resources, and prompts</h2>\n<p>The server primitives have different control semantics. Prompts are user-controlled templates, resources are application-controlled context, and tools are model-controlled actions.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-server\" id=\"user-content-fnref-server\">6</a></sup></p>\n<h3 id=\"resources-carry-addressable-context\">Resources carry addressable context</h3>\n<p>A resource has a URI and content. It suits files, schemas, records, logs, or generated views that a client may inspect and selectively place in context. Servers must validate resource URIs and should enforce access controls on sensitive resources.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-resources\" id=\"user-content-fnref-resources\">7</a></sup></p>\n<p>Resources are not passive just because they are read-only. Retrieved text can contain instructions aimed at the model. The host should attach provenance and treat resource content as untrusted data.</p>\n<h3 id=\"prompts-expose-reusable-workflows\">Prompts expose reusable workflows</h3>\n<p>Prompts package messages and arguments for explicit user selection. They work well for named tasks such as code review or incident summarization. A prompt does not override host policy. It supplies content for the host to inspect and render.</p>\n<h3 id=\"tools-expose-executable-behavior\">Tools expose executable behavior</h3>\n<p>A tool definition combines a stable name, description, and input schema. Tool results can return text, images, embedded resources, resource links, and structured content. The protocol revision published on 18 June 2025 added structured tool output.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-changelog\" id=\"user-content-fnref-changelog-2\">3</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"name\"</span><span>:</span><span> \"create_change_request\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"description\"</span><span>:</span><span> \"Create a draft change request in the selected repository\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"inputSchema\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"type\"</span><span>:</span><span> \"object\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"properties\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>      \"repository\"</span><span>:</span><span> { </span><span>\"type\"</span><span>:</span><span> \"string\"</span><span> }</span><span>,</span></span>\n<span class=\"line\"><span>      \"title\"</span><span>:</span><span> { </span><span>\"type\"</span><span>:</span><span> \"string\"</span><span>,</span><span> \"minLength\"</span><span>:</span><span> 1</span><span> }</span><span>,</span></span>\n<span class=\"line\"><span>      \"body\"</span><span>:</span><span> { </span><span>\"type\"</span><span>:</span><span> \"string\"</span><span>,</span><span> \"minLength\"</span><span>:</span><span> 1</span><span> }</span></span>\n<span class=\"line\"><span>    }</span><span>,</span></span>\n<span class=\"line\"><span>    \"required\"</span><span>:</span><span> [</span><span>\"repository\"</span><span>,</span><span> \"title\"</span><span>,</span><span> \"body\"</span><span>]</span><span>,</span></span>\n<span class=\"line\"><span>    \"additionalProperties\"</span><span>:</span><span> false</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The schema constrains shape, not authority. A valid repository string can still point at a repository the user did not authorize.</p>\n<h2 id=\"transport-boundaries\">Transport boundaries</h2>\n<p>MCP defines standard input and output transport for local process connections and Streamable HTTP for remote connections.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-transports\" id=\"user-content-fnref-transports-2\">5</a></sup> Transport choice changes the threat model.</p>\n<p>With standard input and output, the host launches a process and exchanges newline-delimited JSON-RPC messages. Credentials should come from the environment rather than the HTTP authorization flow.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-authorization\" id=\"user-content-fnref-authorization\">8</a></sup> The host must still control executable paths, environment variables, inherited file descriptors, and working directories.</p>\n<p>With Streamable HTTP, the server exposes an MCP endpoint that accepts POST and may support GET for server-sent events. Implementations must validate the <code>Origin</code> header to defend against DNS rebinding, bind local servers to loopback rather than all interfaces, and add authentication when appropriate.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-transports\" id=\"user-content-fnref-transports-3\">5</a></sup></p>\n<h2 id=\"http-authorization\">HTTP authorization</h2>\n<p>Authorization is optional at the protocol level. When an HTTP server protects resources, the MCP authorization specification treats it as an OAuth resource server. The client acts as an OAuth client, and a separate authorization server issues tokens.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-authorization\" id=\"user-content-fnref-authorization-2\">8</a></sup></p>\n<p>The discovery path is deliberate:</p>\n<ol>\n<li>The MCP server publishes protected-resource metadata.</li>\n<li>The metadata identifies an authorization server.</li>\n<li>The client reads authorization-server metadata.</li>\n<li>The client requests authorization for the MCP resource.</li>\n<li>The client presents the resulting access token to that MCP server.</li>\n</ol>\n<p>The client must include the OAuth <code>resource</code> parameter in authorization and token requests. The server must validate that the token was issued for itself.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-authorization\" id=\"user-content-fnref-authorization-3\">8</a></sup> This audience binding prevents a token meant for one MCP server from becoming a bearer credential accepted by another.</p>\n<p>Public clients must use Proof Key for Code Exchange. Authorization endpoints must use HTTPS, while redirect URIs must use HTTPS or localhost. Redirect URI matching must be exact, and clients should bind the flow with <code>state</code>.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-authorization\" id=\"user-content-fnref-authorization-4\">8</a></sup></p>\n<h3 id=\"token-passthrough-is-a-boundary-violation\">Token passthrough is a boundary violation</h3>\n<p>An MCP server that calls an upstream API is a separate OAuth client to that API. It must not forward the token received from the MCP client. The upstream token has a different audience and should be acquired through a separate authorization relationship.<sup><a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fn-authorization\" id=\"user-content-fnref-authorization-5\">8</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>MCP client -- token audience: files.example --› Files MCP server</span></span>\n<span class=\"line\"><span>Files MCP server -- separate token audience: storage-api --› Storage API</span></span></code></pre>\n<p>Forwarding the first token erases the resource boundary and creates a confused-deputy path.</p>\n<h2 id=\"tool-authorization-and-user-consent\">Tool authorization and user consent</h2>\n<p>OAuth answers which client can call a server on behalf of which resource owner. It does not answer whether a model-generated action is appropriate in the current conversation.</p>\n<p>Tool execution needs a second control plane:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> Decision</span><span> =</span><span> 'allow'</span><span> |</span><span> 'deny'</span><span> |</span><span> 'require_user'</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>function</span><span> authorizeTool</span><span>(input</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>  tool</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  arguments</span><span>:</span><span> unknown</span></span>\n<span class=\"line\"><span>  user</span><span>:</span><span> { id</span><span>:</span><span> string</span><span>; roles</span><span>:</span><span> string</span><span>[] }</span></span>\n<span class=\"line\"><span>  conversationId</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>})</span><span>:</span><span> Decision</span><span> {</span></span>\n<span class=\"line\"><span>  if</span><span> (</span><span>input</span><span>.tool </span><span>===</span><span> 'read_invoice'</span><span>) </span><span>return</span><span> 'allow'</span></span>\n<span class=\"line\"><span>  if</span><span> (</span><span>input</span><span>.tool </span><span>===</span><span> 'create_refund'</span><span>) </span><span>return</span><span> 'require_user'</span></span>\n<span class=\"line\"><span>  return</span><span> 'deny'</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The host should show the actual target and material arguments before approval. A button labeled “Allow tool” provides less information than “Create a refund for invoice INV-1042.” Servers must also apply their own authorization because a client can call them without using a model.</p>\n<h2 id=\"failure-handling\">Failure handling</h2>\n<p>Protocol errors and tool errors belong to different layers. Malformed JSON-RPC, unknown methods, and invalid protocol state should produce JSON-RPC errors. A valid tool invocation that fails in its domain should return a tool result marked as an error so the model can reason about the failure.</p>\n<p>Operational safeguards include:</p>\n<ul>\n<li>Timeouts for every request and cancellation propagation where supported.</li>\n<li>Bounded result sizes before content enters model context.</li>\n<li>Schema validation on inputs and structured outputs.</li>\n<li>Audit records containing server identity, tool name, policy decision, and result status.</li>\n<li>Redaction before logging credentials or resource bodies.</li>\n<li>Connection teardown when version or capability invariants fail.</li>\n</ul>\n<p>Retries should be method-aware. Listing tools is naturally repeatable. Creating a record is not repeatable unless the tool contract includes an idempotency mechanism.</p>\n<h2 id=\"a-minimal-implementation-checklist\">A minimal implementation checklist</h2>\n<p>An interoperable implementation begins with protocol correctness:</p>\n<ul>\n<li>Keep one isolated client session per server.</li>\n<li>Require initialization before operation.</li>\n<li>Store the negotiated version and capabilities.</li>\n<li>Correlate request identifiers without treating them as authority.</li>\n<li>Reject notifications that incorrectly include identifiers.</li>\n<li>Validate tool arguments and resource URIs.</li>\n<li>Separate protocol errors from domain failures.</li>\n</ul>\n<p>A deployable implementation adds security controls:</p>\n<ul>\n<li>Pin executable identity for local servers.</li>\n<li>Validate origins and bind safely for HTTP.</li>\n<li>Discover authorization metadata rather than hard-coding authorization endpoints.</li>\n<li>Bind access tokens to the MCP resource.</li>\n<li>Never pass client tokens through to upstream APIs.</li>\n<li>Require contextual authorization for sensitive tools.</li>\n<li>Treat all returned content as untrusted input to the model.</li>\n</ul>\n<p>MCP standardizes the wire conversation. It does not remove the need for access control, policy, validation, or user-visible consent. Those controls remain with the host and each resource-owning server.</p>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-architecture\">\n<p><a href=\"https://modelcontextprotocol.io/specification/2025-06-18/architecture\" rel=\"noopener noreferrer\">Model Context Protocol, Architecture, protocol revision 2025-06-18</a>. <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-architecture\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-architecture-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-base\">\n<p><a href=\"https://modelcontextprotocol.io/specification/2025-06-18/basic\" rel=\"noopener noreferrer\">Model Context Protocol, Base Protocol overview, protocol revision 2025-06-18</a>. <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-base\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-changelog\">\n<p><a href=\"https://modelcontextprotocol.io/specification/2025-06-18/changelog\" rel=\"noopener noreferrer\">Model Context Protocol, key changes in protocol revision 2025-06-18</a>. <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-changelog\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-changelog-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-lifecycle\">\n<p><a href=\"https://modelcontextprotocol.io/specification/2025-06-18/basic/lifecycle\" rel=\"noopener noreferrer\">Model Context Protocol, Lifecycle, protocol revision 2025-06-18</a>. <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-lifecycle\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-lifecycle-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-transports\">\n<p><a href=\"https://modelcontextprotocol.io/specification/2025-06-18/basic/transports\" rel=\"noopener noreferrer\">Model Context Protocol, Transports, protocol revision 2025-06-18</a>. <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-transports\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-transports-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-transports-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-server\">\n<p><a href=\"https://modelcontextprotocol.io/specification/2025-06-18/server\" rel=\"noopener noreferrer\">Model Context Protocol, Server Features overview, protocol revision 2025-06-18</a>. <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-server\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-resources\">\n<p><a href=\"https://modelcontextprotocol.io/specification/2025-06-18/server/resources\" rel=\"noopener noreferrer\">Model Context Protocol, Resources, protocol revision 2025-06-18</a>. <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-resources\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-authorization\">\n<p><a href=\"https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization\" rel=\"noopener noreferrer\">Model Context Protocol, Authorization, protocol revision 2025-06-18</a>. <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-authorization\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-authorization-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-authorization-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-authorization-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a> <a href=\"https://blog.ecitis.org/mcp-first-principles/#user-content-fnref-authorization-5\" class=\"data-footnote-backref\">↩<sup>5</sup></a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/mcp-first-principles/",
            "title": "Model Context Protocol from First Principles: Messages, Capabilities, and Authorization",
            "summary": "Understand MCP as a stateful boundary between hosts and capability servers, including negotiation, transport, and authorization.",
            "image": "https://blog.ecitis.org/open-graph/mcp-first-principles.png",
            "date_modified": "2026-03-17T00:00:00.000Z",
            "date_published": "2026-03-17T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Ecosystems & Tooling",
                "MCP",
                "Protocols",
                "Tool use"
            ]
        },
        {
            "id": "https://blog.ecitis.org/structured-generation/",
            "content_html": "<p>Prompting a model to return JSON asks a probabilistic decoder to imitate a grammar. Constrained decoding changes the operation: before sampling, the runtime masks tokens that cannot extend any valid document in the target language.</p>\n<p>The result can guarantee syntax without guaranteeing truth. A schema can force an integer field, but it cannot prove that the integer was computed from the source text.</p>\n<h2 id=\"compile-structure-before-generation\">Compile structure before generation</h2>\n<p>JSON Schema, a regular expression, or a context-free grammar is not applied directly to model logits. The serving system compiles the constraint into a recognizer with runtime state.</p>\n<p>At generation step <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>t</mi></mrow></semantics></math>, let <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>V</mi></mrow></semantics></math> be the vocabulary and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>A</mi><mo stretchy=\"false\">(</mo><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo><mo>⊆</mo><mi>V</mi></mrow></semantics></math> be the tokens accepted from grammar state <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>s</mi><mi>t</mi></msub></mrow></semantics></math>. Masked logits are:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msubsup><mi mathvariant=\"normal\">ℓ</mi><mi>t</mi><mo lspace=\"0em\" rspace=\"0em\">′</mo></msubsup><mo stretchy=\"false\">(</mo><mi>v</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mrow><mo fence=\"true\">{</mo><mtable rowspacing=\"0.36em\" columnalign=\"left left\" columnspacing=\"1em\"><mtr><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><msub><mi mathvariant=\"normal\">ℓ</mi><mi>t</mi></msub><mo stretchy=\"false\">(</mo><mi>v</mi><mo stretchy=\"false\">)</mo><mo separator=\"true\">,</mo></mrow></mstyle></mtd><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mi>v</mi><mo>∈</mo><mi>A</mi><mo stretchy=\"false\">(</mo><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo></mrow></mstyle></mtd></mtr><mtr><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mo>−</mo><mi mathvariant=\"normal\">∞</mi><mo separator=\"true\">,</mo></mrow></mstyle></mtd><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mi>v</mi><mo>∉</mo><mi>A</mi><mo stretchy=\"false\">(</mo><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo></mrow></mstyle></mtd></mtr></mtable></mrow></mrow></semantics></math>\n<p>Sampling then uses the normalized distribution over legal tokens. Willard and Louf formulate guided generation as transitions in a finite-state machine and build an index over the model vocabulary.<sup><a href=\"https://blog.ecitis.org/structured-generation/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup> Context-free constraints require stack state because nesting cannot be represented by a finite-state machine alone.</p>\n<figure><figcaption><strong>Grammar state restricts the token mask after each prefix</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/structured-generation-0.webp\" width=\"512\" height=\"273\" alt=\"Grammar state restricts the token mask after each prefix\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"tokens-do-not-align-with-grammar-terminals\">Tokens do not align with grammar terminals</h2>\n<p>A tokenizer can encode punctuation together with adjacent text. A token might contain a quote, property characters, and a colon. Another token can contain a byte prefix that is invalid alone but valid when followed by later bytes.</p>\n<p>The matcher must answer a prefix question: does the token’s byte sequence keep at least one valid completion reachable? Checking only complete terminals rejects useful multi-character tokens. Checking decoded Unicode strings can mishandle byte-level tokenizers.</p>\n<p>A robust engine:</p>\n<ol>\n<li>evaluates token bytes against the current recognizer state</li>\n<li>preserves every reachable parse stack when the grammar is ambiguous</li>\n<li>caches the legal-token set for repeated states</li>\n<li>advances state with the sampled token</li>\n<li>detects accepting state before applying end-of-sequence rules</li>\n</ol>\n<p>XGrammar separates context-independent grammar components from context-dependent components and overlaps grammar processing with GPU execution.<sup><a href=\"https://blog.ecitis.org/structured-generation/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> Its design addresses the cost of walking grammar stacks across a large vocabulary.<sup><a href=\"https://blog.ecitis.org/structured-generation/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup></p>\n<h2 id=\"json-schema-is-more-than-json-syntax\">JSON Schema is more than JSON syntax</h2>\n<p>JSON grammar guarantees balanced delimiters, quoted keys, and valid scalar syntax. A schema adds semantic structure:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"type\"</span><span>:</span><span> \"object\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"properties\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"status\"</span><span>:</span><span> { </span><span>\"enum\"</span><span>:</span><span> [</span><span>\"ok\"</span><span>,</span><span> \"error\"</span><span>] }</span><span>,</span></span>\n<span class=\"line\"><span>    \"retryable\"</span><span>:</span><span> { </span><span>\"type\"</span><span>:</span><span> \"boolean\"</span><span> }</span><span>,</span></span>\n<span class=\"line\"><span>    \"code\"</span><span>:</span><span> { </span><span>\"type\"</span><span>:</span><span> \"integer\"</span><span>,</span><span> \"minimum\"</span><span>:</span><span> 0</span><span> }</span></span>\n<span class=\"line\"><span>  }</span><span>,</span></span>\n<span class=\"line\"><span>  \"required\"</span><span>:</span><span> [</span><span>\"status\"</span><span>,</span><span> \"retryable\"</span><span>]</span><span>,</span></span>\n<span class=\"line\"><span>  \"additionalProperties\"</span><span>:</span><span> false</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Compilation must define which schema features are supported. Object property ordering, numeric bounds, string patterns, recursive references, unions, and <code>additionalProperties</code> each require different machinery. A provider that accepts a schema but ignores unsupported keywords supplies weaker guarantees than the input document suggests.</p>\n<p>Publish a supported subset and reject unsupported constraints at request validation. Silent downgrade turns a type contract into a prompt hint.</p>\n<h2 id=\"state-dependent-masking-cost\">State-dependent masking cost</h2>\n<p>Naively testing every vocabulary token at every step repeats work. Production engines reduce that cost through:</p>\n<ul>\n<li>vocabulary tries that share token prefixes</li>\n<li>precomputed masks for stable grammar states</li>\n<li>cached transitions keyed by recognizer state and token</li>\n<li>separation of deterministic and branching parse states</li>\n<li>CPU grammar work overlapped with device model execution</li>\n</ul>\n<p>The mask still has to reach the sampling kernel. Copying a dense vocabulary-sized mask from host memory per sequence and per step can dominate short structured outputs. Bitsets, cached device masks, and batch grouping by grammar state reduce transfer.</p>\n<p>Batching introduces divergence. Two requests using the same schema can occupy different parse states, so their legal sets differ. The sampler must apply per-sequence masks even when the model forward pass is batched.</p>\n<h2 id=\"whitespace-and-canonical-forms\">Whitespace and canonical forms</h2>\n<p>JSON permits whitespace in many positions. Leaving all whitespace paths open expands grammar state and lets the model spend tokens formatting. A canonical mode can constrain separators and indentation, but it changes the accepted language.</p>\n<p>Define canonicalization as an API option:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> StructuredOutputOptions</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  schema</span><span>:</span><span> object</span></span>\n<span class=\"line\"><span>  whitespace</span><span>:</span><span> 'compact'</span><span> |</span><span> 'model'</span></span>\n<span class=\"line\"><span>  rejectUnsupportedKeywords</span><span>:</span><span> true</span></span>\n<span class=\"line\"><span>  maxOutputTokens</span><span>:</span><span> number</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Compact output reduces token use. Model-controlled whitespace preserves stylistic freedom. Neither choice changes field semantics.</p>\n<h2 id=\"token-boundaries-and-parser-composition\">Token boundaries and parser composition</h2>\n<p>Grammar terminals rarely align with model tokens. DOMINO uses token-aligned finite-state machines so a token can span several grammar terminals without forcing a token boundary after each terminal.<sup><a href=\"https://blog.ecitis.org/structured-generation/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup> This avoids rejecting a token merely because it packages punctuation, whitespace, and property text together.</p>\n<p>SynCode precomputes a finite lookahead table for grammar terminals and uses it to construct masks for context-free grammars.<sup><a href=\"https://blog.ecitis.org/structured-generation/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup> Two properties need separate tests:</p>\n<ul>\n<li>soundness means every admitted token prefix retains a valid grammatical completion</li>\n<li>completeness means every token that can participate in a valid completion remains admitted</li>\n</ul>\n<p>An unsound masker can emit invalid output. An incomplete masker can force low-probability tokenization paths even when the intended text is valid. Boundary fixtures should include escaped text, punctuation merged with names, and byte sequences split across tokens.</p>\n<h3 id=\"near-runnable-validation-boundary\">Near-runnable validation boundary</h3>\n<p>Constrained output should still terminate at an ordinary validator:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> json</span></span>\n<span class=\"line\"><span>from</span><span> jsonschema </span><span>import</span><span> Draft202012Validator</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> validate_generated_payload</span><span>(</span><span>raw_text</span><span>,</span><span> schema</span><span>):</span></span>\n<span class=\"line\"><span>    payload </span><span>=</span><span> json</span><span>.</span><span>loads</span><span>(raw_text)</span></span>\n<span class=\"line\"><span>    validator </span><span>=</span><span> Draft202012Validator</span><span>(schema)</span></span>\n<span class=\"line\"><span>    errors </span><span>=</span><span> sorted</span><span>(</span></span>\n<span class=\"line\"><span>        validator.</span><span>iter_errors</span><span>(payload),</span></span>\n<span class=\"line\"><span>        key</span><span>=lambda</span><span> error</span><span>: </span><span>list</span><span>(error.absolute_path),</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"><span>    if</span><span> errors</span><span>:</span></span>\n<span class=\"line\"><span>        formatted </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>            {</span></span>\n<span class=\"line\"><span>                \"path\"</span><span>:</span><span> list</span><span>(error.absolute_path),</span></span>\n<span class=\"line\"><span>                \"message\"</span><span>:</span><span> error</span><span>.</span><span>message</span><span>,</span></span>\n<span class=\"line\"><span>            }</span></span>\n<span class=\"line\"><span>            for</span><span> error </span><span>in</span><span> errors</span></span>\n<span class=\"line\"><span>        ]</span></span>\n<span class=\"line\"><span>        raise</span><span> ValueError</span><span>(formatted)</span></span>\n<span class=\"line\"><span>    return</span><span> payload</span></span></code></pre>\n<p>Install <code>jsonschema</code> and supply the same schema compiled by the generation engine. This second validation catches engine defects and documents the trust boundary. Domain checks follow it and can query authoritative identifiers or enforce cross-field rules.</p>\n<h2 id=\"compilation-limits\">Compilation limits</h2>\n<p>Compile user-supplied schemas before reserving model capacity. Cache compiled state by canonical schema bytes, tokenizer revision, grammar compiler version, and whitespace policy.</p>\n<p>Set separate limits for schema bytes, recursive depth, grammar states, and compile time. Returning a validation error before inference keeps malformed schemas out of the model queue.</p>\n<p>Expose the compilation result:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> CompiledConstraint</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  digest</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  compilerVersion</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  tokenizerRevision</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  supportedKeywords</span><span>:</span><span> readonly</span><span> string</span><span>[]</span></span>\n<span class=\"line\"><span>  ignoredKeywords</span><span>:</span><span> readonly</span><span> never</span><span>[]</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The <code>never[]</code> type expresses the strict policy: unsupported keywords cause rejection. If compatibility requires permissive handling, name that mode explicitly and return every ignored keyword.</p>\n<h2 id=\"valid-structure-is-not-valid-data\">Valid structure is not valid data</h2>\n<p>Apply three validation layers after decoding:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Layer</th><th>Detects</th><th>Example</th></tr></thead><tbody><tr><td>Parser</td><td>Invalid JSON bytes</td><td>Unclosed string</td></tr><tr><td>Schema validator</td><td>Wrong shape or local constraint</td><td>Missing required field</td></tr><tr><td>Domain validator</td><td>Cross-field and external rules</td><td>Unknown database identifier</td></tr></tbody></table></div>\n<p>Constrained decoding should make parser failure unreachable under supported conditions. Schema validation remains useful as a defense against implementation bugs and unsupported keywords. Domain validation is mandatory before side effects.</p>\n<p>For tool calls, treat model output as untrusted input even when it matches the schema. Resolve authorization from server-side identity. Do not allow a generated account identifier or file path to expand permissions.</p>\n<h2 id=\"recursion-and-termination\">Recursion and termination</h2>\n<p>A recursive grammar can generate an unbounded document. The model still needs a token budget. If the budget ends before an accepting state, the result is structurally incomplete.</p>\n<p>The runtime should distinguish:</p>\n<ul>\n<li>accepted completion</li>\n<li>token-budget exhaustion in non-accepting state</li>\n<li>grammar with no legal continuation</li>\n<li>model end token masked because the document is incomplete</li>\n<li>internal grammar-engine failure</li>\n</ul>\n<p>Returning partial bytes as ordinary JSON hides the failure. Return an explicit structured-generation status outside the generated payload.</p>\n<h2 id=\"operational-failure-modes\">Operational failure modes</h2>\n<p><strong>Empty legal set:</strong> the recognizer or tokenizer integration has reached a state with no accepted token. Log the grammar state, recent token bytes, tokenizer revision, and schema digest.</p>\n<p><strong>Schema explosion:</strong> a large union or deeply recursive schema creates expensive state sets. Enforce compilation limits before model execution.</p>\n<p><strong>Cache poisoning:</strong> a compiled-grammar cache uses an incomplete key. Include compiler version, tokenizer identity, schema bytes, and generation options.</p>\n<p><strong>Unsupported keyword drift:</strong> a library upgrade changes which constraints are enforced. Pin compiler versions and run conformance fixtures.</p>\n<p><strong>Unicode mismatch:</strong> validation operates on characters while masking operates on token bytes. Use one explicit byte-to-text policy and test invalid sequences.</p>\n<h2 id=\"conformance-testing\">Conformance testing</h2>\n<p>Generate tests from the constraint rather than collecting a few successful prompts:</p>\n<ul>\n<li>schemas with nested objects and arrays</li>\n<li>escaped strings and Unicode boundaries</li>\n<li>overlapping unions</li>\n<li>optional and required fields</li>\n<li>recursive definitions up to configured depth</li>\n<li>outputs ending exactly at the token budget</li>\n<li>adversarial property names that share token prefixes</li>\n</ul>\n<p>For each case, assert parser validity, schema validity, accepted-state termination, and deterministic failure classification. Separately evaluate factual accuracy. Grammar conformance and task quality are different metrics.</p>\n<p>Structured generation is reliable when the contract is precise: a documented schema subset, a tokenizer-aware recognizer, state-specific masks, and post-generation domain validation. Prompt wording alone provides none of those guarantees.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Brandon T. Willard and Rémi Louf, “Efficient Guided Generation for Large Language Models,” arXiv, 2023. <a href=\"https://arxiv.org/abs/2307.09702\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2307.09702</a> <a href=\"https://blog.ecitis.org/structured-generation/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Yixin Dong et al., “XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models,” arXiv, 2024. <a href=\"https://arxiv.org/abs/2411.15100\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2411.15100</a> <a href=\"https://blog.ecitis.org/structured-generation/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/structured-generation/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Luca Beurer-Kellner, Marc Fischer, and Martin Vechev, “Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation,” ICML, 2024. <a href=\"https://arxiv.org/abs/2403.06988\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2403.06988</a> <a href=\"https://blog.ecitis.org/structured-generation/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Suresh G. K. et al., “SynCode: LLM Generation with Grammar Augmentation,” arXiv, 2024. <a href=\"https://arxiv.org/abs/2403.01632\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2403.01632</a> <a href=\"https://blog.ecitis.org/structured-generation/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/structured-generation/",
            "title": "Structured Generation: How LLMs Produce Valid JSON",
            "summary": "See how grammar compilation and token masking turn probabilistic decoding into reliably structured JSON and schema-constrained output.",
            "image": "https://blog.ecitis.org/open-graph/structured-generation.png",
            "date_modified": "2026-03-05T00:00:00.000Z",
            "date_published": "2026-03-05T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "Structured output",
                "JSON Schema",
                "Decoding"
            ]
        },
        {
            "id": "https://blog.ecitis.org/torch-compile-internals/",
            "content_html": "<p><code>torch.compile</code> is not a single compiler pass. It is a guarded path from a live Python program to reusable machine code. TorchDynamo captures tensor operations from Python frames, AOTAutograd constructs forward and backward graphs, and TorchInductor lowers those graphs into device code. On supported GPUs, Inductor uses Triton as a code-generation building block.<sup><a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fn-compiler-stack\" id=\"user-content-fnref-compiler-stack\">1</a></sup></p>\n<figure><figcaption><strong>Compilation path and its runtime boundaries</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/torch-compile-internals-0.webp\" width=\"512\" height=\"672\" alt=\"Compilation path and its runtime boundaries\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"dynamo-captures-a-program-specialization\">Dynamo captures a program specialization</h2>\n<p>Dynamo hooks Python frame evaluation and symbolically executes bytecode. Tensor operations become nodes in an FX graph. Python values that affect execution become assumptions. A guard records each assumption needed to reuse the graph, including tensor shape, dtype, device, object identity, or a global value.<sup><a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fn-guards\" id=\"user-content-fnref-guards\">2</a></sup></p>\n<p>This distinction matters. Dynamo does not prove that a function has one universal tensor graph. It records one graph for one set of observed conditions, then emits checks that decide whether later calls fit that specialization.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> normalize_or_zero</span><span>(</span><span>x</span><span>,</span><span> enabled</span><span>):</span></span>\n<span class=\"line\"><span>    if</span><span> enabled</span><span>:</span></span>\n<span class=\"line\"><span>        return</span><span> x </span><span>/</span><span> x</span><span>.</span><span>norm</span><span>()</span></span>\n<span class=\"line\"><span>    return</span><span> torch</span><span>.</span><span>zeros_like</span><span>(x)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>compiled </span><span>=</span><span> torch</span><span>.</span><span>compile</span><span>(normalize_or_zero)</span></span></code></pre>\n<p>Here, <code>enabled</code> is a Python boolean. Dynamo can specialize on its value and cache separate graphs for the branches. A tensor-dependent condition is different:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> branch_on_data</span><span>(</span><span>x</span><span>):</span></span>\n<span class=\"line\"><span>    if</span><span> x</span><span>.</span><span>sum</span><span>()</span><span> &gt;</span><span> 0</span><span>:</span></span>\n<span class=\"line\"><span>        return</span><span> x</span><span>.</span><span>sin</span><span>()</span></span>\n<span class=\"line\"><span>    return</span><span> x</span><span>.</span><span>cos</span><span>()</span></span></code></pre>\n<p>The branch requires a tensor value that is available only after execution. PyTorch documents tensor-dependent control flow and direct scalar extraction with <code>.item()</code> as common graph-break causes.<sup><a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fn-graph-breaks\" id=\"user-content-fnref-graph-breaks\">3</a></sup></p>\n<h2 id=\"graph-breaks-define-optimization-boundaries\">Graph breaks define optimization boundaries</h2>\n<p>At a graph break, Dynamo compiles the graph captured so far, returns to ordinary Python for the unsupported operation, then resumes tracing. Correctness is preserved, but fusion cannot cross the boundary. A break inside a transformer block can therefore cost more than one in preprocessing because the former separates operators that exchange large intermediate tensors.</p>\n<p>Use the compiler logs before rewriting code:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>TORCH_LOGS</span><span>=</span><span>\"graph_breaks,guards,recompiles\"</span><span> python</span><span> train.py</span></span></code></pre>\n<p><code>fullgraph=True</code> converts the first graph break into an error and requires one captured graph. PyTorch recommends it when the goal is to find and remove breaks in a performance-critical region.<sup><a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fn-fullgraph\" id=\"user-content-fnref-fullgraph\">4</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>@torch</span><span>.</span><span>compile</span><span>(fullgraph</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>def</span><span> block</span><span>(</span><span>x</span><span>,</span><span> weight</span><span>,</span><span> bias</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>functional</span><span>.</span><span>gelu</span><span>(x </span><span>@</span><span> weight </span><span>+</span><span> bias)</span></span></code></pre>\n<p>Do not force file I/O, metrics emission, or data decoding into that function. Mark a problematic helper with <code>torch.compiler.disable</code>, or place compilation around the numerical core. The boundary should express the part of the program whose inputs, side effects, and shapes are stable.</p>\n<h2 id=\"aotautograd-exposes-the-backward-pass\">AOTAutograd exposes the backward pass</h2>\n<p>Eager autograd records operations while the forward pass runs. AOTAutograd instead traces a functional forward graph and constructs a corresponding backward graph ahead of execution. This gives the compiler visibility into saved tensors, recomputation choices, and backward operators. PyTorch describes AOTAutograd as the component that captures backpropagation for acceleration by Inductor.<sup><a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fn-compiler-stack\" id=\"user-content-fnref-compiler-stack-2\">1</a></sup></p>\n<p>Suppose a forward graph computes:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>y</mi><mo>=</mo><mi mathvariant=\"normal\">GELU</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><mi>x</mi><mi>W</mi><mo>+</mo><mi>b</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math>\n<p>The backward needs values derived from the matrix product and activation. A partitioner decides what the forward should save versus what the backward can recompute. Saving reduces arithmetic and increases live memory. Recomputation makes the opposite trade. This is why a compiled training step can have a different memory profile from eager execution while retaining the same differentiation result.</p>\n<p>Mutation and aliasing complicate this graph. AOTAutograd functionalizes many in-place operations into functional equivalents, then tracks the mutations that must be reflected at the boundary. Custom operators need correct aliasing and mutation declarations; a wrong schema can turn a local implementation mistake into a compiler-wide correctness problem.</p>\n<h2 id=\"inductor-turns-graphs-into-loop-nests\">Inductor turns graphs into loop nests</h2>\n<p>Inductor receives ATen-level graphs after decomposition. It reasons about iteration domains, dependencies, device placement, layouts, and reuse. Elementwise chains are strong fusion candidates because their loop domains align.</p>\n<p>An eager expression such as:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>y </span><span>=</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>functional</span><span>.</span><span>gelu</span><span>(x </span><span>@</span><span> weight </span><span>+</span><span> bias)</span></span>\n<span class=\"line\"><span>z </span><span>=</span><span> y </span><span>*</span><span> scale </span><span>+</span><span> residual</span></span></code></pre>\n<p>can launch separate kernels for bias addition, activation, multiplication, and residual addition around a matrix multiplication. Inductor can select a matrix-multiplication template and fuse compatible pointwise epilogues, or generate a separate fused pointwise kernel when the template boundary prevents full fusion. The useful question is not the number of FX nodes. It is how many times large tensors are written to and read from device memory.</p>\n<p>Fusion is constrained by:</p>\n<ul>\n<li>incompatible iteration orders or layouts;</li>\n<li>reductions that require synchronization;</li>\n<li>mutations and aliasing;</li>\n<li>device transfers;</li>\n<li>graph breaks;</li>\n<li>operator implementations treated as opaque calls.</li>\n</ul>\n<p>On GPUs, Triton code expresses blocked loads, arithmetic, and stores. On CPUs, Inductor emits C++ with vectorized loops. Backend choice changes the generated code, but Dynamo guards and graph partitioning remain part of the contract.</p>\n<h2 id=\"shapes-determine-code-reuse\">Shapes determine code reuse</h2>\n<p>Static dimensions give Inductor concrete loop bounds and specialization opportunities. Variable sequence lengths can trigger recompilation when a guard fails. PyTorch first assumes static shapes under <code>dynamic=None</code>, then can retry with dynamic dimensions after a shape-driven recompilation.<sup><a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fn-guards\" id=\"user-content-fnref-guards-2\">2</a></sup></p>\n<p>There are three operational responses:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Workload property</th><th>Compiler policy</th><th>Cost</th></tr></thead><tbody><tr><td>Few stable shapes</td><td>Specialize and cache</td><td>More kernels, stronger specialization</td></tr><tr><td>Broad bounded shapes</td><td>Mark selected dimensions dynamic</td><td>Fewer recompiles, more general code</td></tr><tr><td>Unbounded irregular shapes</td><td>Bucket or pad before compilation</td><td>Extra padding work, predictable cache</td></tr></tbody></table></div>\n<p>Do not set every dimension dynamic as a reflex. Batch and sequence axes often vary; head width and vocabulary partition usually do not. Encode the variability that exists in production traces.</p>\n<h2 id=\"compilation-modes-change-the-optimization-budget\">Compilation modes change the optimization budget</h2>\n<p>The default mode balances compile time and runtime work. <code>reduce-overhead</code> uses CUDA graphs where applicable to reduce Python and launch overhead. <code>max-autotune</code> benchmarks candidate matrix-multiplication or convolution implementations and enables CUDA graphs on supported GPU paths.<sup><a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fn-compile-api\" id=\"user-content-fnref-compile-api\">5</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>compiled_step </span><span>=</span><span> torch</span><span>.</span><span>compile</span><span>(</span></span>\n<span class=\"line\"><span>    train_step,</span></span>\n<span class=\"line\"><span>    mode</span><span>=</span><span>\"max-autotune\"</span><span>,</span></span>\n<span class=\"line\"><span>    dynamic</span><span>=</span><span>None</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>Warm-up measurements must be separated from steady-state measurements. Compilation and autotuning run on early calls, and a new shape or guard value can cause later recompilation. Measure:</p>\n<ul>\n<li>cold latency from process start;</li>\n<li>warm latency for each supported shape bucket;</li>\n<li>compilation count and reason;</li>\n<li>peak memory during compilation and execution;</li>\n<li>numerical agreement against eager mode.</li>\n</ul>\n<h2 id=\"diagnose-one-layer-at-a-time\">Diagnose one layer at a time</h2>\n<p>PyTorch exposes backend ablations that isolate the failing stage.<sup><a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fn-troubleshooting\" id=\"user-content-fnref-troubleshooting\">6</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Dynamo capture only</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>compile</span><span>(fn, backend</span><span>=</span><span>\"eager\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Add AOTAutograd</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>compile</span><span>(fn, backend</span><span>=</span><span>\"aot_eager\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run the complete default stack</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>compile</span><span>(fn, backend</span><span>=</span><span>\"inductor\"</span><span>)</span></span></code></pre>\n<p>If eager capture fails, inspect Python semantics and graph breaks. If <code>aot_eager</code> fails, inspect autograd, functionalization, mutation, and custom-operator schemas. If only Inductor fails, preserve the FX graph and generated debug trace, then reduce the operator pattern.</p>\n<p>Accuracy testing should compare outputs and gradients across representative dtypes and shapes. Compiler speed is irrelevant if a custom kernel mishandles non-contiguous inputs or a low-precision reduction exceeds the application tolerance.</p>\n<h2 id=\"deployment-checklist\">Deployment checklist</h2>\n<p>Compile the numerical step rather than the surrounding service loop. Establish shape buckets from observed traffic. Warm each bucket before accepting latency-sensitive work. Record guard failures and recompiles. Keep an eager fallback for unsupported paths, but alert when it is used. Pin the PyTorch and driver environment used for performance qualification because generated kernels and compiler behavior are environment-dependent.</p>\n<p><code>torch.compile</code> works best when the program exposes stable tensor regions with explicit boundaries. Dynamo supplies guarded capture, AOTAutograd supplies differentiable graphs, and Inductor supplies loop-level optimization. Performance work consists of preserving those regions, controlling specialization, and verifying the generated execution rather than treating the decorator as a universal switch.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-compiler-stack\">\n<p><a href=\"https://docs.pytorch.org/docs/stable/user_guide/torch_compiler/torch.compiler.html\" rel=\"noopener noreferrer\">PyTorch compiler documentation</a>. <a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fnref-compiler-stack\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fnref-compiler-stack-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-guards\">\n<p><a href=\"https://docs.pytorch.org/docs/main/user_guide/torch_compiler/torch.compiler_troubleshooting.html#terminology\" rel=\"noopener noreferrer\">PyTorch compiler troubleshooting: guards and recompilation</a>. <a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fnref-guards\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fnref-guards-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-graph-breaks\">\n<p><a href=\"https://docs.pytorch.org/docs/main/user_guide/torch_compiler/torch.compiler_troubleshooting.html#data-dependent-operations\" rel=\"noopener noreferrer\">PyTorch compiler troubleshooting: data-dependent operations</a>. <a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fnref-graph-breaks\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-fullgraph\">\n<p><a href=\"https://docs.pytorch.org/docs/stable/user_guide/torch_compiler/compile/programming_model.fullgraph_true.html\" rel=\"noopener noreferrer\">PyTorch documentation for <code>fullgraph=True</code></a>. <a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fnref-fullgraph\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-compile-api\">\n<p><a href=\"https://docs.pytorch.org/docs/stable/generated/torch.compile.html\" rel=\"noopener noreferrer\">PyTorch <code>torch.compile</code> API</a>. <a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fnref-compile-api\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-troubleshooting\">\n<p><a href=\"https://docs.pytorch.org/docs/main/user_guide/torch_compiler/torch.compiler_troubleshooting.html#ablation\" rel=\"noopener noreferrer\">PyTorch compiler troubleshooting: backend ablation</a>. <a href=\"https://blog.ecitis.org/torch-compile-internals/#user-content-fnref-troubleshooting\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/torch-compile-internals/",
            "title": "How torch.compile Turns PyTorch Programs into Fused Kernels",
            "summary": "Trace a PyTorch program through Dynamo, AOTAutograd, Inductor, guards, graph breaks, and generated Triton kernels.",
            "image": "https://blog.ecitis.org/open-graph/torch-compile-internals.png",
            "date_modified": "2026-02-24T00:00:00.000Z",
            "date_published": "2026-02-24T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Ecosystems & Tooling",
                "torch.compile",
                "Compilers",
                "PyTorch"
            ]
        },
        {
            "id": "https://blog.ecitis.org/disaggregated-llm-serving/",
            "content_html": "<p>An LLM request contains two workloads with different shapes. Prefill evaluates all prompt tokens and produces the first output token. Decode repeatedly evaluates one new position while reading the accumulated key-value cache.</p>\n<p>Colocating both phases is operationally simple. It also couples their batching, parallelism, and latency. Disaggregated serving assigns prefill and decode to different worker pools, transfers the cache between them, and scales each pool against its own service objective. DistServe formalized this design around time to first token and time per output token.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup></p>\n<h2 id=\"phase-specific-resource-plans\">Phase-specific resource plans</h2>\n<p>Prefill exposes matrix dimensions large enough to use compute units efficiently. Decode uses small per-request matrix dimensions and repeatedly streams model weights and cache state. A shared worker must compromise between large prompt batches and decode cadence.</p>\n<p>Disaggregation permits separate plans:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Decision</th><th>Prefill pool</th><th>Decode pool</th></tr></thead><tbody><tr><td>Primary latency</td><td>Time to first token</td><td>Time per output token</td></tr><tr><td>Batching unit</td><td>Prompt tokens or chunks</td><td>Active sequences</td></tr><tr><td>Parallelism</td><td>Favor prompt throughput</td><td>Favor low collective overhead</td></tr><tr><td>Capacity signal</td><td>Queued prompt tokens</td><td>Active sequences and live KV bytes</td></tr><tr><td>Scale trigger</td><td>Predicted prefill queue time</td><td>Predicted inter-token delay</td></tr></tbody></table></div>\n<p>DistServe reports that phase separation removes prefill-decode interference and allows different resource and parallelism choices for each phase.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-1\" id=\"user-content-fnref-1-2\">1</a></sup> The paper’s evaluation found up to <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>7.4</mn><mo>×</mo></mrow></semantics></math> higher request rate or <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>12.6</mn><mo>×</mo></mrow></semantics></math> tighter service objectives than its baselines, with more than <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>90</mn><mi mathvariant=\"normal\">%</mi></mrow></semantics></math> of requests meeting the specified constraints.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-1\" id=\"user-content-fnref-1-3\">1</a></sup></p>\n<figure><figcaption><strong>Prefill and decode as separate scheduling domains</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/disaggregated-llm-serving-0.webp\" width=\"512\" height=\"382\" alt=\"Prefill and decode as separate scheduling domains\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"cache-transfer-is-part-of-request-latency\">Cache transfer is part of request latency</h2>\n<p>After prefill, decode cannot start until the required keys and values are available. For cache size <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>M</mi><mrow><mi>K</mi><mi>V</mi></mrow></msub></mrow></semantics></math>, effective link bandwidth <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>W</mi></mrow></semantics></math>, and fixed transfer overhead <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>T</mi><mn>0</mn></msub></mrow></semantics></math>, a lower-bound model is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>T</mi><mtext>handoff</mtext></msub><mo>≥</mo><msub><mi>T</mi><mn>0</mn></msub><mo>+</mo><mfrac><msub><mi>M</mi><mrow><mi>K</mi><mi>V</mi></mrow></msub><mi>W</mi></mfrac></mrow></semantics></math>\n<p>This is a planning bound, not a latency prediction. Contention, serialization, topology, registration, and synchronization increase elapsed time.</p>\n<p>A transfer protocol needs:</p>\n<ul>\n<li>model revision and cache-layout identity</li>\n<li>source and destination block handles</li>\n<li>sequence positions and valid token count</li>\n<li>per-layer completion state</li>\n<li>cancellation and timeout semantics</li>\n<li>acknowledgement before source reclamation</li>\n</ul>\n<p>Do not send a monolithic opaque tensor if the runtime stores paged blocks. Transfer complete blocks directly and preserve logical block order in metadata. This avoids repacking into a contiguous buffer.</p>\n<h2 id=\"pipeline-the-handoff\">Pipeline the handoff</h2>\n<p>Waiting for all layers to finish before beginning transfer leaves the link idle during prefill. Layerwise or blockwise streaming overlaps later prefill work with earlier cache movement:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>prefill layer l</span></span>\n<span class=\"line\"><span>  -&gt; record completion event</span></span>\n<span class=\"line\"><span>  -&gt; enqueue KV blocks for layer l</span></span>\n<span class=\"line\"><span>  -&gt; transport waits on event</span></span>\n<span class=\"line\"><span>  -&gt; destination records receive completion</span></span>\n<span class=\"line\"><span>  -&gt; decode waits before first read of layer l</span></span></code></pre>\n<p>Correctness depends on event ordering. A destination-visible pointer does not prove that the corresponding device writes have completed.</p>\n<p>Backpressure is also required. If the destination cannot reserve blocks, the source should not compute a cache that has nowhere to land. Reserve destination capacity before admitting prefill, or allow an explicit spill path with its own latency budget.</p>\n<h2 id=\"scheduling-with-two-queues\">Scheduling with two queues</h2>\n<p>The router now controls a distributed pipeline. A useful estimate for candidate prefill worker <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>p</mi></mrow></semantics></math> and decode worker <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>d</mi></mrow></semantics></math> is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mover accent=\"true\"><mi>T</mi><mo stretchy=\"true\">^</mo></mover><mrow><mi>f</mi><mi>i</mi><mi>r</mi><mi>s</mi><mi>t</mi></mrow></msub><mo stretchy=\"false\">(</mo><mi>p</mi><mo separator=\"true\">,</mo><mi>d</mi><mo stretchy=\"false\">)</mo><mo>=</mo><msub><mover accent=\"true\"><mi>Q</mi><mo stretchy=\"true\">^</mo></mover><mi>p</mi></msub><mo>+</mo><msub><mover accent=\"true\"><mi>P</mi><mo stretchy=\"true\">^</mo></mover><mi>p</mi></msub><mo>+</mo><msub><mover accent=\"true\"><mi>X</mi><mo stretchy=\"true\">^</mo></mover><mrow><mi>p</mi><mo separator=\"true\">,</mo><mi>d</mi></mrow></msub><mo>+</mo><msub><mover accent=\"true\"><mi>D</mi><mo stretchy=\"true\">^</mo></mover><mi>d</mi></msub></mrow></semantics></math>\n<p>The terms represent prefill queue, prefill execution, transfer, and decode admission. Each estimate must use current topology and cache reservations.</p>\n<p>Routing on queue length alone is weak because a queued long prompt carries more work than a queued short prompt. Count scheduled prompt tokens and use profiled execution curves. For decode, active sequence count also misses context length and cache pressure.</p>\n<h3 id=\"prefix-locality\">Prefix locality</h3>\n<p>A prefill worker may already hold a reusable prefix. A decode worker may hold the previous turn of a conversation. Routing should compare reuse savings with queue and transfer costs.</p>\n<p>Use explicit estimates:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> PlacementCost</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  queueMs</span><span>:</span><span> number</span></span>\n<span class=\"line\"><span>  computeMs</span><span>:</span><span> number</span></span>\n<span class=\"line\"><span>  transferMs</span><span>:</span><span> number</span></span>\n<span class=\"line\"><span>  recomputeMs</span><span>:</span><span> number</span></span>\n<span class=\"line\"><span>  capacityRisk</span><span>:</span><span> number</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Keeping the terms visible prevents a cache-hit heuristic from silently overriding latency objectives.</p>\n<h2 id=\"parallelism-can-differ-by-phase\">Parallelism can differ by phase</h2>\n<p>Tensor parallelism reduces per-device weight and compute, but each layer introduces collectives. Pipeline parallelism partitions layers and introduces pipeline scheduling. Prefill can amortize communication over larger token matrices. Decode executes collectives at every token step.</p>\n<p>DistServe searches phase-specific parallelism and places communicating prefill and decode instances according to cluster bandwidth.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-1\" id=\"user-content-fnref-1-4\">1</a></sup> Treat the selected topology as part of the deployment artifact. A scheduler cannot compensate for a prefill-decode pair split across a congested or oversubscribed path.</p>\n<h2 id=\"designs-beyond-a-single-handoff\">Designs beyond a single handoff</h2>\n<p>Splitwise profiles prompt and token generation separately, then assigns them to machine pools selected for each phase.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> Its architecture includes mixed machines that can absorb imbalances rather than leaving a saturated phase blocked behind an idle pool.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup> This matters because a fixed prefill-to-decode ratio assumes a stable workload mix.</p>\n<p>Mooncake treats the cache as the center of a disaggregated architecture. It separates cache storage and transfer from compute scheduling, with a distributed cache that can reuse state across requests and workers.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup> P/D-Serve focuses on fine-grained organization and dynamic adjustment of prefill and decode instances.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup></p>\n<p>These systems expose separate controls:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Control</th><th>Question</th></tr></thead><tbody><tr><td>Phase placement</td><td>Which worker executes prompt or token work?</td></tr><tr><td>Cache placement</td><td>Where do completed blocks remain available?</td></tr><tr><td>Elasticity</td><td>How does the pool ratio change with traffic?</td></tr></tbody></table></div>\n<p>Do not conflate them in one autoscaler. A phase replica count can change without moving warm cache. A cache tier can become full while compute workers remain underutilized.</p>\n<p>A near-runnable planner can turn sampled requests into initial offered-work estimates:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> dataclasses </span><span>import</span><span> dataclass</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@dataclass</span><span>(frozen</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>class</span><span> Request</span><span>:</span></span>\n<span class=\"line\"><span>    prompt_tokens</span><span>:</span><span> int</span></span>\n<span class=\"line\"><span>    output_tokens</span><span>:</span><span> int</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> offered_work</span><span>(</span><span>requests</span><span>,</span><span> prefill_ms</span><span>,</span><span> decode_ms</span><span>):</span></span>\n<span class=\"line\"><span>    prefill </span><span>=</span><span> sum</span><span>(</span><span>prefill_ms</span><span>(r.prompt_tokens) </span><span>for</span><span> r </span><span>in</span><span> requests)</span></span>\n<span class=\"line\"><span>    decode </span><span>=</span><span> sum</span><span>(</span></span>\n<span class=\"line\"><span>        decode_ms</span><span>(r.prompt_tokens, r.output_tokens)</span></span>\n<span class=\"line\"><span>        for</span><span> r </span><span>in</span><span> requests</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"><span>    return</span><span> prefill</span><span>,</span><span> decode</span></span></code></pre>\n<p>The functions must come from profiles on the intended hardware. Replace aggregate work with a queue simulation before deployment because service objectives depend on burst arrivals. Add measured transfer time to each prefill completion and enforce destination cache reservations.</p>\n<h2 id=\"failure-modes\">Failure modes</h2>\n<p><strong>Destination allocation failure:</strong> prefill finishes, but decode has no cache pages. Reserve before execution and release the reservation on cancellation.</p>\n<p><strong>Orphaned transfers:</strong> the client disconnects while cache movement is in flight. Propagate cancellation through router, source, transport, and destination; make cleanup idempotent.</p>\n<p><strong>Version skew:</strong> pools run different model revisions or cache formats. Reject handoff on an exact layout identifier.</p>\n<p><strong>Link saturation:</strong> more prefill replicas increase offered transfer traffic without increasing decode capacity. Autoscaling must include link utilization and transfer queue time.</p>\n<p><strong>Decode starvation:</strong> routing admits more completed prefills than decode can drain. Apply pipeline admission control at the request boundary.</p>\n<p><strong>Replica hotspots:</strong> prefix-aware routing concentrates traffic. Bound affinity and fall back when predicted queue delay exceeds recomputation savings.</p>\n<h2 id=\"when-colocated-serving-is-the-correct-choice\">When colocated serving is the correct choice</h2>\n<p>Disaggregation adds a networked state transfer, another admission boundary, and more failure states. It is not automatically faster. Keep phases colocated when transfer time consumes the expected interference savings, traffic is too small to fill separate pools, the topology lacks a predictable device-to-device path, or operational simplicity matters more than independent scaling.</p>\n<p>Benchmark both designs with the same arrival trace and service objectives. Compare goodput, not raw tokens per second: count requests that satisfy both first-token and inter-token latency. That is the metric DistServe uses to connect resource planning to user-visible constraints.<sup><a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fn-1\" id=\"user-content-fnref-1-5\">1</a></sup></p>\n<h2 id=\"deployment-checklist\">Deployment checklist</h2>\n<p>Before production traffic:</p>\n<ul>\n<li>profile prefill by uncached token count</li>\n<li>profile decode by batch size and context length</li>\n<li>measure transfer by cache bytes and topology path</li>\n<li>test cancellation at every pipeline boundary</li>\n<li>verify source blocks remain live until acknowledgement</li>\n<li>inject worker and link failures</li>\n<li>enforce compatible model and cache-layout identifiers</li>\n<li>report service-objective attainment by phase</li>\n</ul>\n<p>Disaggregated serving is a scheduling architecture, not a socket between two model servers. Its value comes from independent phase control. Its cost is that cache placement, transport, and admission become part of inference correctness.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Yinmin Zhong et al., “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving,” OSDI, 2024. <a href=\"https://arxiv.org/abs/2401.09670\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2401.09670</a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-1-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-1-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-1-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-1-5\" class=\"data-footnote-backref\">↩<sup>5</sup></a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Pratyush Patel et al., “Splitwise: Efficient Generative LLM Inference Using Phase Splitting,” ISCA, 2024. <a href=\"https://arxiv.org/abs/2311.18677\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2311.18677</a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Ruoyu Qin et al., “Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving,” FAST, 2025. <a href=\"https://arxiv.org/abs/2407.00079\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2407.00079</a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Yuhang Jin et al., “P/D-Serve: Serving Disaggregated Large Language Model at Scale,” arXiv, 2024. <a href=\"https://arxiv.org/abs/2408.08147\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2408.08147</a> <a href=\"https://blog.ecitis.org/disaggregated-llm-serving/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/disaggregated-llm-serving/",
            "title": "Disaggregated LLM Serving: Separating Prefill from Decode",
            "summary": "Separate compute-heavy prefill from memory-bound decode, then scale each worker pool around its own latency objective.",
            "image": "https://blog.ecitis.org/open-graph/disaggregated-llm-serving.png",
            "date_modified": "2026-02-12T00:00:00.000Z",
            "date_published": "2026-02-12T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "LLM serving",
                "Prefill",
                "Distributed systems"
            ]
        },
        {
            "id": "https://blog.ecitis.org/kv-cache-memory-systems/",
            "content_html": "<p>Autoregressive inference repeatedly evaluates attention over an expanding prefix. Recomputing every key and value projection at every step would repeat work for tokens that have not changed. A key-value cache stores those projections, then appends the projections for each new token.</p>\n<p>That optimization moves the serving limit. Model weights are mostly static, but cache allocation follows request arrivals, prompt lengths, generation lengths, beam branches, cancellations, and shared prefixes. The cache is therefore both a tensor and a dynamic memory-management problem.</p>\n<h2 id=\"cache-contents\">Cache contents</h2>\n<p>For a transformer layer, attention projects hidden states into queries, keys, and values:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi>Q</mi><mo>=</mo><mi>X</mi><msub><mi>W</mi><mi>Q</mi></msub><mo separator=\"true\">,</mo><mspace width=\"2em\"></mspace><mi>K</mi><mo>=</mo><mi>X</mi><msub><mi>W</mi><mi>K</mi></msub><mo separator=\"true\">,</mo><mspace width=\"2em\"></mspace><mi>V</mi><mo>=</mo><mi>X</mi><msub><mi>W</mi><mi>V</mi></msub></mrow></semantics></math>\n<p>During decode, the current query attends to cached keys and values:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mi mathvariant=\"normal\">Attention</mi><mo>⁡</mo><mo stretchy=\"false\">(</mo><msub><mi>q</mi><mi>t</mi></msub><mo separator=\"true\">,</mo><msub><mi>K</mi><mrow><mo>≤</mo><mi>t</mi></mrow></msub><mo separator=\"true\">,</mo><msub><mi>V</mi><mrow><mo>≤</mo><mi>t</mi></mrow></msub><mo stretchy=\"false\">)</mo><mo>=</mo><mi mathvariant=\"normal\">softmax</mi><mo>⁡</mo><mrow><mo fence=\"true\">(</mo><mfrac><mrow><msub><mi>q</mi><mi>t</mi></msub><msubsup><mi>K</mi><mrow><mo>≤</mo><mi>t</mi></mrow><mi>T</mi></msubsup></mrow><msqrt><msub><mi>d</mi><mi>h</mi></msub></msqrt></mfrac><mo fence=\"true\">)</mo></mrow><msub><mi>V</mi><mrow><mo>≤</mo><mi>t</mi></mrow></msub></mrow></semantics></math>\n<p>Only the current query is transient. Keys and values remain live because later tokens can attend to them.</p>\n<p>For a decoder with <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>L</mi></mrow></semantics></math> layers, <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>H</mi><mrow><mi>k</mi><mi>v</mi></mrow></msub></mrow></semantics></math> key-value heads, head dimension <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>D</mi></mrow></semantics></math>, sequence length <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>S</mi></mrow></semantics></math>, and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>B</mi></mrow></semantics></math> bytes per scalar, cache storage per sequence is:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><msub><mi>M</mi><mrow><mi>K</mi><mi>V</mi></mrow></msub><mo>=</mo><mn>2</mn><mi>L</mi><mi>S</mi><msub><mi>H</mi><mrow><mi>k</mi><mi>v</mi></mrow></msub><mi>D</mi><mi>B</mi></mrow></semantics></math>\n<p>The leading factor accounts for separate key and value tensors. This formula exposes the useful levers: context length, cache precision, layer count, and key-value head count. Multi-query attention shares one key-value head across query heads; grouped-query attention uses more than one but fewer than full multi-head attention.<sup><a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup></p>\n<h2 id=\"prefill-and-decode-write-different-access-patterns\">Prefill and decode write different access patterns</h2>\n<p>Prefill processes a prompt as a token matrix. Each layer writes a run of key and value vectors, while attention consumes the prompt prefix. Decode appends a narrow slice and reads the existing cache.</p>\n<p>The distinction matters for kernel design:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Phase</th><th>Cache operation</th><th>Typical pressure</th></tr></thead><tbody><tr><td>Prefill</td><td>Bulk allocation and writes</td><td>Compute, allocation latency</td></tr><tr><td>Decode</td><td>Append one position, read all prior positions</td><td>Memory bandwidth, cache capacity</td></tr><tr><td>Cancellation</td><td>Release every block owned by a request</td><td>Allocator synchronization</td></tr><tr><td>Prefix reuse</td><td>Map a new request to existing blocks</td><td>Reference counting, routing</td></tr></tbody></table></div>\n<p>A serving scheduler must reserve enough blocks before admitting work. If it admits from token limits alone, a burst of long generations can exhaust cache pages after requests have begun streaming.</p>\n<figure><figcaption><strong>Logical sequences map to reusable physical KV blocks</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/kv-cache-memory-systems-0.webp\" width=\"512\" height=\"391\" alt=\"Logical sequences map to reusable physical KV blocks\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"contiguous-allocation-wastes-capacity\">Contiguous allocation wastes capacity</h2>\n<p>A simple allocator reserves one contiguous region per sequence at its maximum configured length. That approach creates internal fragmentation when generation ends early. It also creates external fragmentation as variable-sized regions are allocated and freed.</p>\n<p>PagedAttention applies operating-system paging to the cache. Logical token blocks map through a block table to non-contiguous physical blocks. vLLM’s original paper reports near-zero cache waste and supports cache sharing within and across requests; its evaluation measured between two and four times the throughput of the compared systems at comparable latency.<sup><a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup></p>\n<p>The attention kernel cannot assume that adjacent logical tokens are adjacent in memory. It resolves each logical block through the block table. That indirection is the price of flexible allocation.</p>\n<h3 id=\"block-size-trade-offs\">Block size trade-offs</h3>\n<p>Small blocks reduce waste in the final, partially filled page. They increase block-table entries, allocation operations, and address translation inside the kernel. Large blocks reduce metadata but strand more capacity at sequence tails.</p>\n<p>Choose block size with workload traces, not model shape alone. Record the distribution of prompt lengths, output lengths, cancellations, and prefix reuse. Then replay those traces against candidate block sizes and measure:</p>\n<ul>\n<li>live cache bytes versus reserved bytes</li>\n<li>allocation latency at high request concurrency</li>\n<li>attention-kernel time by context length</li>\n<li>blocks copied during beam or parallel sampling</li>\n<li>eviction and recomputation frequency</li>\n</ul>\n<h2 id=\"prefix-caching-needs-identity-rules\">Prefix caching needs identity rules</h2>\n<p>Two requests can share cached blocks only when the token sequence and every cache-relevant model input match. Text equality is insufficient because tokenization settings, adapter selection, multimodal embeddings, position assignment, and model revision can change the resulting states.</p>\n<p>A practical cache key includes:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> PrefixKey</span><span> =</span><span> {</span></span>\n<span class=\"line\"><span>  modelRevision</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  adapterId</span><span>:</span><span> string</span><span> |</span><span> null</span></span>\n<span class=\"line\"><span>  tokenIds</span><span>:</span><span> readonly</span><span> number</span><span>[]</span></span>\n<span class=\"line\"><span>  positionPolicy</span><span>:</span><span> string</span></span>\n<span class=\"line\"><span>  multimodalDigest</span><span>:</span><span> string</span><span> |</span><span> null</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Hash blocks incrementally so a request can reuse the longest matching prefix. Store parent hashes to prevent a block with identical local tokens from matching under a different preceding context.</p>\n<p>Reference counts protect shared blocks. Copy-on-write handles divergence: requests share completed prefix blocks, while each allocates private blocks for generated continuations. The vLLM paper uses this mechanism for parallel sampling and beam search.<sup><a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup></p>\n<h2 id=\"eviction-policy-affects-compute\">Eviction policy affects compute</h2>\n<p>Evicting a cache entry does not lose model weights or user data, but it turns the next hit into recomputation. A least-recently-used policy treats every block equally even though recomputation cost rises with the number of prefix tokens that must be replayed.</p>\n<p>A cost-aware score can combine:</p>\n<math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mtext>value</mtext><mo stretchy=\"false\">(</mo><mi>p</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mfrac><mrow><mover accent=\"true\"><mtext>reuse probability</mtext><mo stretchy=\"true\">^</mo></mover><mo stretchy=\"false\">(</mo><mi>p</mi><mo stretchy=\"false\">)</mo><mo>⋅</mo><mover accent=\"true\"><mtext>prefill cost saved</mtext><mo stretchy=\"true\">^</mo></mover><mo stretchy=\"false\">(</mo><mi>p</mi><mo stretchy=\"false\">)</mo></mrow><mrow><mtext>bytes</mtext><mo stretchy=\"false\">(</mo><mi>p</mi><mo stretchy=\"false\">)</mo></mrow></mfrac></mrow></semantics></math>\n<p>The terms are estimates derived from observed traffic. Keep them separate in telemetry. A high hit rate can still save little compute if hits cover short prefixes.</p>\n<p>Cache-aware routing introduces another constraint. Sending a request to the least-loaded replica can discard a large local prefix hit. Sending every matching prefix to one replica can create a hotspot. Route using predicted queue time plus predicted prefill time after reuse.</p>\n<h2 id=\"quantized-cache-changes-attention-inputs\">Quantized cache changes attention inputs</h2>\n<p>Quantizing the cache reduces bytes per scalar in the storage formula. It also adds scale metadata and dequantization work to attention. Keys and values can require different calibration because their distributions differ by layer and head.</p>\n<p>Validation must include generation quality and kernel performance. A cache format that halves tensor payload does not halve total latency if dequantization prevents a fused attention path. Measure:</p>\n<ul>\n<li>quality against the unquantized cache on task-specific prompts</li>\n<li>time per output token across context lengths</li>\n<li>scale-storage overhead</li>\n<li>conversion work during prefix reuse or cache transfer</li>\n<li>numerical behavior for long-running generations</li>\n</ul>\n<h2 id=\"moving-cache-across-memory-tiers\">Moving cache across memory tiers</h2>\n<p>GPU memory is not the only possible cache tier. A serving system can move inactive blocks to host memory, local storage, or another worker, then restore them when a request resumes. The decision compares transfer time with recomputation time.</p>\n<p>Prompt Cache defines reusable prompt modules and preserves their positional accuracy when reusing attention state across prompts.<sup><a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup> CacheGen targets network transfer: it encodes key-value tensors into a compressed bitstream conditioned on their distribution, then adapts compression to available bandwidth.<sup><a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup></p>\n<p>Tiering needs an explicit state machine:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>type</span><span> BlockState</span><span> =</span></span>\n<span class=\"line\"><span>  |</span><span> { kind</span><span>:</span><span> 'device'</span><span>; gpu</span><span>:</span><span> number</span><span>; address</span><span>:</span><span> bigint</span><span> }</span></span>\n<span class=\"line\"><span>  |</span><span> { kind</span><span>:</span><span> 'moving'</span><span>; operationId</span><span>:</span><span> string</span><span>; destination</span><span>:</span><span> 'host'</span><span> |</span><span> 'remote'</span><span> }</span></span>\n<span class=\"line\"><span>  |</span><span> { kind</span><span>:</span><span> 'host'</span><span>; bufferId</span><span>:</span><span> string</span><span> }</span></span>\n<span class=\"line\"><span>  |</span><span> { kind</span><span>:</span><span> 'remote'</span><span>; objectKey</span><span>:</span><span> string</span><span>; checksum</span><span>:</span><span> string</span><span> }</span></span>\n<span class=\"line\"><span>  |</span><span> { kind</span><span>:</span><span> 'evicted'</span><span> }</span></span></code></pre>\n<p>Only one transition may own a block at a time. Cancellation during transfer must release both the source reference and any completed destination allocation. A checksum catches truncated or stale payloads, while model and layout identifiers prevent restoring cache produced by incompatible weights.</p>\n<p>A near-runnable capacity calculator follows directly from the storage equation:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> kv_bytes</span><span>(</span><span>layers</span><span>,</span><span> sequence</span><span>,</span><span> kv_heads</span><span>,</span><span> head_dim</span><span>,</span><span> bytes_per_scalar</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> 2</span><span> *</span><span> layers </span><span>*</span><span> sequence </span><span>*</span><span> kv_heads </span><span>*</span><span> head_dim </span><span>*</span><span> bytes_per_scalar</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> concurrent_capacity</span><span>(</span><span>free_bytes</span><span>,</span><span> reserve_fraction</span><span>,</span><span> per_sequence_bytes</span><span>):</span></span>\n<span class=\"line\"><span>    usable </span><span>=</span><span> int</span><span>(free_bytes </span><span>*</span><span> reserve_fraction)</span></span>\n<span class=\"line\"><span>    return</span><span> usable </span><span>//</span><span> per_sequence_bytes</span></span></code></pre>\n<p>Feed it values from the deployed model configuration and runtime telemetry. The result is not admission control by itself; workspaces, allocator metadata, partial pages, and transfer buffers also consume memory.</p>\n<p>Choose restore or recomputation using measured queue, transfer, decoding, and prefill curves. Report bytes restored, tokens recovered, tokens recomputed, and time saved. A high hit count is misleading if most restored prefixes are short.</p>\n<h2 id=\"failure-modes-and-instrumentation\">Failure modes and instrumentation</h2>\n<p><strong>Admission overcommit:</strong> requests begin with enough memory for their prompts but no reserve for output growth. Reject or queue before prefill using a configurable output reservation.</p>\n<p><strong>Leaked references:</strong> cancellation bypasses a decrement path, leaving shared blocks permanently live. Audit ownership transitions and reconcile allocator totals.</p>\n<p><strong>Hash collisions or incomplete keys:</strong> unrelated states share a block. Use a collision-resistant digest and include every state-changing input.</p>\n<p><strong>Eviction storms:</strong> a working set slightly larger than capacity repeatedly recomputes prefixes. Track evicted-token recomputation, not just eviction count.</p>\n<p><strong>Tail-page waste:</strong> a nominally paged allocator still wastes capacity if blocks are too large for the output-length distribution. Report occupancy inside allocated blocks.</p>\n<p>The minimum dashboard needs physical blocks free, reserved, active, shared, and evictable; cache hit tokens; recomputed tokens; allocation latency; and request aborts caused by cache pressure. Those measurements turn KV capacity from an opaque out-of-memory event into a schedulable resource.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Noam Shazeer, “Fast Transformer Decoding: One Write-Head is All You Need,” arXiv, 2019. <a href=\"https://arxiv.org/abs/1911.02150\" rel=\"noopener noreferrer\">https://arxiv.org/abs/1911.02150</a> <a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Woosuk Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention,” SOSP, 2023. <a href=\"https://arxiv.org/abs/2309.06180\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2309.06180</a> <a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>In Gim et al., “Prompt Cache: Modular Attention Reuse for Low-Latency Inference,” MLSys, 2024. <a href=\"https://arxiv.org/abs/2311.04934\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2311.04934</a> <a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Yuhan Liu et al., “CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving,” SIGCOMM, 2024. <a href=\"https://arxiv.org/abs/2310.07240\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2310.07240</a> <a href=\"https://blog.ecitis.org/kv-cache-memory-systems/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/kv-cache-memory-systems/",
            "title": "Inside the KV Cache: Memory Systems for LLM Inference",
            "summary": "Treat the KV cache as a dynamic memory system shaped by paging, fragmentation, shared prefixes, and unpredictable request lifetimes.",
            "image": "https://blog.ecitis.org/open-graph/kv-cache-memory-systems.png",
            "date_modified": "2026-02-03T00:00:00.000Z",
            "date_published": "2026-02-03T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "KV cache",
                "Memory",
                "LLM serving"
            ]
        },
        {
            "id": "https://blog.ecitis.org/ai-hardware-landscape/",
            "content_html": "<p>The AI hardware landscape is undergoing a transformation. While NVIDIA maintains its grip on approximately 86-92% of the AI accelerator market, a constellation of alternatives has emerged, each with distinct architectural philosophies and target workloads. This diversification matters for engineers who need to optimize for cost, latency, throughput, or specific model architectures.</p>\n<h2 id=\"why-hardware-diversification-matters\">Why Hardware Diversification Matters</h2>\n<p>NVIDIA’s dominance in AI accelerators stems from a decade of CUDA ecosystem development and first-mover advantage in deep learning. The company’s H100 and newer Blackwell GPUs remain the default choice for most organizations. However, several factors are driving diversification:</p>\n<p><strong>Supply constraints</strong>: Morgan Stanley analysts reported in November 2024 that all of NVIDIA’s 2025 chip production was already sold out. Organizations unable to secure H100/H200 allocations must look elsewhere.</p>\n<p><strong>Cost pressure</strong>: Inference costs can exceed training costs by 15x over a model’s lifetime, according to OpenAI’s 2024 figures. At $6.98-7.57 per H100 GPU-hour on major cloud providers, inference-heavy workloads demand alternatives.</p>\n<p><strong>Architectural mismatch</strong>: GPUs are general-purpose parallel processors. Workloads with predictable memory access patterns or specific numerical precision requirements can benefit from purpose-built silicon.</p>\n<p><strong>Vendor lock-in concerns</strong>: Major cloud providers and AI labs are developing custom chips. JPMorgan projects that custom chips from Google, Amazon, Meta, and OpenAI will account for 45% of the AI chip market by 2028, up from 37% in 2024.</p>\n<p>The result is a market where engineers can choose hardware optimized for their specific constraints rather than defaulting to NVIDIA.</p>\n<h2 id=\"google-tpus\">Google TPUs</h2>\n<p>Google’s Tensor Processing Units represent the most mature alternative to NVIDIA GPUs for large-scale AI workloads. Unlike GPUs, TPUs are application-specific integrated circuits built around systolic arrays optimized for matrix multiplication.</p>\n<h3 id=\"systolic-array-architecture\">Systolic Array Architecture</h3>\n<p>At the core of every TPU sits a Matrix Multiply Unit (MXU) composed of multiply-accumulate units arranged in a systolic array. The term “systolic” refers to the rhythmic data flow through the structure, analogous to blood pumping through the heart. This architecture exploits a fundamental property of matrix multiplication: it requires O(n^3) compute for O(n^2) bytes of data, making it compute-bound rather than memory-bound when hardware is designed correctly.</p>\n<figure><figcaption><strong>Weight-Stationary Systolic Array</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/ai-hardware-landscape-0.webp\" width=\"512\" height=\"769\" alt=\"Weight-Stationary Systolic Array\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>The weight-stationary dataflow pattern works as follows: engineers preload the weights of matrix B into the individual multiply-accumulate units arranged in a grid. Matrix A’s activation values enter from the left edge and flow horizontally across the array. Each MAC unit multiplies its stored weight by the incoming activation, adds the result to a partial sum arriving from above, and passes both the activation (horizontally) and updated partial sum (vertically) to neighboring units. This arrangement eliminates intermediate memory writes for the entire matrix multiplication.</p>\n<p>The original TPU v1 contained a 256x256 array of 8-bit multiply-accumulate units, yielding 65,536 MACs. Because a TPU runs at 700MHz, a TPU can compute 92 TOPS in the matrix unit. During execution of this massive matrix multiply, all intermediate results pass directly between 64K ALUs without any memory access, reducing power consumption substantially compared to architectures that require memory round-trips.</p>\n<p>The pipelining characteristics matter for understanding performance. Systolic arrays are heavily pipelined: given that the array is 256 units wide, it takes 256 cycles from when the first element enters until it exits. Twice that many cycles for all data to flow through. However, at peak utilization, all 65,536 processors operate simultaneously.</p>\n<p>The systolic array’s weakness emerges with dynamic or irregular computation patterns. Because data flows through the grid on a fixed schedule, the architecture struggles with conditional execution, sparse matrices (unless using SparseCore), and operations that require random access patterns. This inflexibility trades generality for extreme efficiency on its target workload: dense matrix multiplication with predictable access patterns.</p>\n<h3 id=\"xla-compilation-and-hardware-mapping\">XLA Compilation and Hardware Mapping</h3>\n<p>TPUs were codesigned with the XLA (Accelerated Linear Algebra) compiler to achieve their performance characteristics. The compilation process involves several stages:</p>\n<ol>\n<li>\n<p><strong>Lazy Evaluation and Graph Tracing</strong>: Operations are not executed immediately. Instead, PyTorch/XLA or JAX records operations in an intermediate representation (IR) graph. This process is called “tracing.”</p>\n</li>\n<li>\n<p><strong>HLO Generation</strong>: When results are needed (printing a tensor, saving a checkpoint, or at an explicit synchronization point), the accumulated IR graph converts into HLO (High-Level Opcodes), a representation specific to the XLA compiler.</p>\n</li>\n<li>\n<p><strong>XLA Optimization</strong>: The XLA compiler performs operator fusion, memory layout optimization, and parallelization, then compiles to machine code for the target device. Compiled graphs are cached, so subsequent executions with the same computation graph and input shapes reuse the optimized binary.</p>\n</li>\n</ol>\n<p>For multi-chip scaling, Google’s answer is to make the XLA compiler responsible for coordinating communication between chips. With parallelism dimensions specified by researchers (DP, FSDP, TP, number of slices), the XLA compiler inserts the appropriate hierarchical collectives for the TPU topology. The GSPMD system (<a href=\"https://arxiv.org/abs/2105.04663\">Xu et al., 2021</a>) enables large-scale training with minimal code changes.</p>\n<p>A simple JAX example demonstrates the programming model:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> jax</span></span>\n<span class=\"line\"><span>import</span><span> jax</span><span>.</span><span>numpy </span><span>as</span><span> jnp</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> multiply</span><span>(</span><span>x</span><span>,</span><span> y</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> jnp</span><span>.</span><span>einsum</span><span>(</span><span>'bf,fd-&gt;db'</span><span>, x, y)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># XLA compiles this function and caches the result</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> jax</span><span>.</span><span>jit</span><span>(multiply)(jnp.</span><span>ones</span><span>((</span><span>128</span><span>, </span><span>256</span><span>)), jnp.</span><span>ones</span><span>((</span><span>256</span><span>, </span><span>16</span><span>), dtype</span><span>=</span><span>jnp.bfloat16))</span></span></code></pre>\n<p>By default, matrix multiplication in JAX on TPUs uses bfloat16 with float32 accumulation. This can be controlled with the <code>precision</code> argument on relevant functions (matmul, dot, einsum).</p>\n<h3 id=\"interconnect-topology-2d-vs-3d-torus\">Interconnect Topology: 2D vs 3D Torus</h3>\n<p>TPU interconnect topology directly impacts collective operation performance.</p>\n<p><strong>2D Torus</strong> (TPU v2, v3, v5e, v6e): Each chip connects to its four nearest neighbors (north, south, east, west). The links wrap around at boundaries, creating a donut-shaped logical topology that eliminates edge chips with fewer connections. A 16x16 grid of 256 TPUs provides uniform bandwidth and latency regardless of which two chips communicate.</p>\n<p><strong>3D Torus</strong> (TPU v4, v5p, v7): Each chip connects to six neighbors along three axes. A TPU rack consists of 64 TPUs connected in a 4x4x4 configuration. The composition resembles a Rubik’s Cube, with each TPU chip having 6 ICI (Inter-Chip Interconnect) links in ±X, ±Y, ±Z directions. TPU v5p achieved approximately 4.45 exaflops across 8,960-chip pods using 16x20x28 superpod configurations.</p>\n<p>The 3D torus increases bisection bandwidth compared to 2D, which matters for all-to-all communication patterns. Google’s Optical Circuit Switches (OCS) enable flexible topology configuration, including twisted torus variants with better bisection properties. Twisting the torus brings the largest benefit for tensor parallel (TP) operations since there are multiple all-gather and reduce-scatter operations per layer.</p>\n<p>Unlike all-reduce used in backpropagation, which maps well to 2D and 3D tori, all-to-all patterns strain bisection bandwidth. The twisted torus can reduce the worst-case number of hops, improving all-to-all collective throughput.</p>\n<h3 id=\"bf16-accumulation-and-numerical-precision\">BF16 Accumulation and Numerical Precision</h3>\n<p>Inside the MXU, multiplications occur in bfloat16 format while accumulations use full FP32 precision. This design choice has significant implications.</p>\n<p>Bfloat16 uses one sign bit, eight exponent bits, and seven mantissa bits. Because it has the same exponent size as float32, it exhibits identical behavior for underflows, overflows, and other numeric instabilities during training. Unlike FP16, which typically requires loss scaling, BF16 comes close to being a drop-in replacement for FP32.</p>\n<p>The wide dynamic range makes BF16 highly resistant to overflow and underflow, at the cost of reduced precision from the 7-bit mantissa. This tradeoff works well for neural network training where the exact value matters less than staying within a reasonable range.</p>\n<p>One caveat: while BF16’s stability helps pre-training, its low precision can cause rounding errors that accumulate. Modern RL frameworks using different engines for training and inference may see subtle differences in their implementation lead to different rounding errors, creating training-inference mismatch (<a href=\"https://arxiv.org/abs/2510.26788\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2510.26788\">.26788</a>).<p></p>\n<p>Checkpoints from TPU training can be deployed on other hardware platforms (CPU or GPU inference) without extensive manual conversions.</p>\n<h3 id=\"tpu-generations\">TPU Generations</h3>\n<p><strong>TPU v5e</strong> targets cost-sensitive inference and smaller training workloads:</p>\n<ul>\n<li>128x128 MXU array, 16,384 MACs per cycle</li>\n<li>197 BF16 TFLOPS peak</li>\n<li>16 GB HBM2e with 819 GB/s bandwidth</li>\n<li>Supports training on up to 256 chips</li>\n<li>2D torus interconnect topology</li>\n</ul>\n<p><strong>TPU v5p</strong> is designed for large-scale training:</p>\n<ul>\n<li>128x128 MXU array, 459 BF16 TFLOPS peak</li>\n<li>95 GB HBM2e with 2,765 GB/s bandwidth</li>\n<li>3D torus topology connecting up to 8,960 chips per pod</li>\n<li>Single slice training for up to 6,144 chips; multislice scaling to 18,432 chips</li>\n</ul>\n<p><strong>Trillium (TPU v6e)</strong>, generally available since late 2024, represents an architectural inflection point:</p>\n<ul>\n<li>256x256 MXU array, quadrupling FLOPs per cycle</li>\n<li>918 BF16 TFLOPS peak (4.7x improvement over v5e)</li>\n<li>32 GB HBM with 1,600 GB/s bandwidth</li>\n<li>67% more energy-efficient than v5e</li>\n<li>Third-generation SparseCore for ultra-large embedding tables</li>\n</ul>\n<p><strong>Ironwood (TPU v7)</strong> reached general availability in Q4 2025 with 4,614 TFLOPS peak performance.</p>\n<h3 id=\"benchmark-results\">Benchmark Results</h3>\n<p>MLPerf 4.1 training benchmarks show Trillium delivers up to 1.8x better performance-per-dollar compared to TPU v5p and 99% scaling efficiency across data-center networks using Cloud TPU multislice technology, outperforming the 94% scaling efficiency of v5p clusters within a single ICI domain.</p>\n<p>BERT training completes 2.8x faster on TPUs than on A100 GPUs, while T5-3B model training finishes in 12 hours versus 31 hours on comparable GPU infrastructure. MLPerf results show TPU v5e leading in 8 of 9 training categories.</p>\n<p>For inference, MLPerf 5.0 shows Trillium delivering 3.5x throughput improvement for queries/second on Stable Diffusion XL compared to TPU v5e. The cost to generate 1000 images is 22 cents on Trillium, 35% less than v5e.</p>\n<p>Google’s TPU v5e reaches similar throughput to H100 by using more chips at lower per-chip speed. Eight TPU v5e chips generate approximately 2,175 tokens/sec on Llama2-70B at a cost of $11/hour, whereas 8 H100 GPUs cost an order of magnitude more.</p>\n<p>One computer vision startup reported switching from 128 H100s to TPU v6e and reducing their monthly inference bill from <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>340</mn><mi>K</mi><mi>t</mi><mi>o</mi></mrow></semantics></math>89K. Midjourney reportedly moved from NVIDIA A100/H100 clusters to TPU v6e pods, dropping monthly spend from approximately <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>2.1</mn><mi>m</mi><mi>i</mi><mi>l</mi><mi>l</mi><mi>i</mi><mi>o</mi><mi>n</mi><mi>t</mi><mi>o</mi><mi>u</mi><mi>n</mi><mi>d</mi><mi>e</mi><mi>r</mi></mrow></semantics></math>700K.</p>\n<h2 id=\"cerebras-wafer-scale-integration\">Cerebras: Wafer-Scale Integration</h2>\n<p>Cerebras takes the opposite approach from multi-chip systems: instead of connecting many small chips, they build single chips the size of entire silicon wafers.</p>\n<h3 id=\"the-wafer-scale-engine\">The Wafer-Scale Engine</h3>\n<p>The WSE-3, powering the CS-3 system, occupies 46,225 mm^2 of silicon, roughly 57 times larger than NVIDIA’s H100 (826 mm^2). Built on TSMC 5nm, it contains:</p>\n<ul>\n<li>4 trillion transistors</li>\n<li>900,000 AI-optimized compute cores</li>\n<li>125 petaflops peak AI performance</li>\n<li>44 GB on-chip SRAM</li>\n<li>21 PB/s memory bandwidth (7,000x more than H100)</li>\n<li>214 Pb/s fabric bandwidth</li>\n</ul>\n<h3 id=\"yield-management-and-interconnect-redundancy\">Yield Management and Interconnect Redundancy</h3>\n<p>Building a wafer-scale chip requires solving the yield problem that killed previous attempts at wafer-scale integration. The WSE-3 has a die area of 462 cm^2, resulting in higher defect probability than smaller dies where yields typically exceed 90%.</p>\n<p>Cerebras addresses this through several mechanisms:</p>\n<p><strong>Fine-grained cores</strong>: Each WSE-3 core occupies 0.05mm^2, while an H100 SM core is 6mm^2. The smaller the core, the better fault tolerance becomes.</p>\n<p><strong>Redundant mesh routing</strong>: A mesh topology improves resilience against manufacturing defects. Cerebras includes redundant links and cores across the wafer. When defective regions are detected, the hardware driver performs a remapping process, reconfiguring the interconnect to bypass faulty areas and preserve a virtually intact two-dimensional mesh topology. This process hides defects from developers, who program the system as if it were built on an ideal mesh network.</p>\n<p><strong>Intelligent sparing</strong>: Unlike previous approaches requiring massive redundancy overhead, the architecture achieves high yield with approximately 1% spare cores through intelligent routing that leverages nearby cores to replace defective units.</p>\n<p><strong>Distributed power delivery</strong>: The system incorporates over 300 voltage regulation modules distributed across the wafer’s surface, providing redundancy in power delivery and allowing independent regulation for each reticle.</p>\n<p>In April 2021, Cerebras announced 100% claimed yield for the WSE-2, achieved by designing a system where any manufacturing defect can be bypassed.</p>\n<h3 id=\"dataflow-architecture-and-the-memory-wall\">Dataflow Architecture and the Memory Wall</h3>\n<p>The Cerebras architecture uses fine-grained dataflow compute cores where all computation is triggered by data arrival. The fabric transports both data and associative control directly in hardware. Once cores receive data, the hardware triggers a lookup of instructions to execute.</p>\n<p>This addresses the memory wall problem endemic to traditional GPU clusters. The WSE keeps 44 GB of memory directly on the same die as compute cores, eliminating the bottleneck of shuttling data between HBM and compute units. At the system level, bandwidth reaches 20 PB/s.</p>\n<h3 id=\"weight-streaming-for-large-models\">Weight Streaming for Large Models</h3>\n<p>For models exceeding on-chip capacity, Cerebras developed weight streaming execution. Rather than fitting the entire model on-chip, this mode loads the neural network one layer at a time.</p>\n<p>Cerebras stores all model weights externally in MemoryX units and streams weights onto the CS-3 as needed to compute each layer. The weights are never stored on the system, not even temporarily. As weights stream through, the CS-3 performs computation using the underlying dataflow mechanisms. Each individual weight triggers computation as an AXPY operation. Once complete, the weight is discarded and hardware moves to the next element.</p>\n<p>MemoryX configurations scale from 24TB and 36TB for enterprise customers to 120TB and 1,200TB for hyperscalers. This allows a single CS-3 system to train models up to 24 trillion parameters.</p>\n<h3 id=\"compiler-technology\">Compiler Technology</h3>\n<p>The Cerebras compiler extracts the operation graph from code and aligns operations with compatible kernels in the Cerebras Software Platform. Each matched kernel corresponds to a layer in the network’s dataflow graph.</p>\n<p>Cerebras developed MACH (Multiple-Architecture Compiler for Advanced Computing Hardware) specifically for massively-parallel, spatial, dataflow architectures. Additionally, research has introduced SPADA, a spatial dataflow architecture programming language with a rigorous dataflow semantics framework defining routing correctness, data races, and deadlocks. SPADA enables developers to express complex parallel patterns in 6-8x less code than CSL with near-ideal weak scaling.</p>\n<h3 id=\"benchmark-results-1\">Benchmark Results</h3>\n<p>Cerebras reports inference performance of 2,100 tokens/second on Llama 3.2 70B, 16x faster than GPU solutions and 68x faster than hyperscale clouds as measured by Artificial Analysis. For Llama 3.1-405B, Cerebras Inference generates 969 output tokens per second at 128K full context length and 16-bit precision.</p>\n<p>For Llama 4 Scout, Cerebras achieves over 2,600 tokens per second, 19x faster than the fastest GPU solutions per Artificial Analysis verification.</p>\n<p>Cerebras benchmarked the CS-3 at over 21x faster inference than NVIDIA’s Blackwell B200 GPU running Llama 3 70B with 1024 input and 4096 output segment lengths. Using SemiAnalysis benchmarks, CS-3 is 32% lower cost than B200 while delivering results 21x faster, accounting for both capex and opex including energy costs.</p>\n<p>A full cluster of 2048 CS-3s delivers 256 exaflops of AI compute and can train Llama2-70B from scratch in less than a day.</p>\n<h3 id=\"use-cases\">Use Cases</h3>\n<p>Wafer-scale architecture excels when model weights fit in on-chip memory and memory bandwidth dominates performance. Large language model inference, where the same weights are reused across many tokens, is a natural fit. The WSE-3 was recognized by TIME Magazine as a Best Invention of 2024.</p>\n<p>The tradeoff is flexibility. You cannot incrementally add compute; you either use the full wafer or you do not.</p>\n<h2 id=\"groq-deterministic-inference-at-scale\">Groq: Deterministic Inference at Scale</h2>\n<p>Groq builds hardware specifically for inference, not training. Their Language Processing Unit (LPU) takes a fundamentally different architectural approach from both GPUs and other accelerators.</p>\n<h3 id=\"the-architectural-bet-determinism-over-flexibility\">The Architectural Bet: Determinism Over Flexibility</h3>\n<p>Groq’s LPU architecture is built on a central premise: deterministic execution enables optimizations impossible on dynamically scheduled systems. The LPU can achieve deterministic execution by avoiding traditional reactive hardware components (branch predictors, arbiters, reordering buffers, caches) and having all execution explicitly controlled by the compiler.</p>\n<p>The LPU is a VLIW-like (Very Long Instruction Word) pipeline that processes one instruction stream at a time. Unlike GPUs with many cores and thread contexts, Groq feeds AI model tokens through a single, wide pipeline of functional units, executing all operations in lock-step with no kernel switching. Every clock cycle performs useful work.</p>\n<p>The compiler pre-computes the entire execution graph, including inter-chip communication patterns, down to individual clock cycles. This static scheduling eliminates non-determinism. If Groq’s compiler says a task will take 28.5 milliseconds, it takes exactly 28.5 milliseconds, every time. This predictability is valuable for real-time systems where tail latency is a critical metric.</p>\n<h3 id=\"tensor-streaming-processors-vs-systolic-arrays\">Tensor Streaming Processors vs Systolic Arrays</h3>\n<p>Groq’s initial name for their ASIC was the Tensor Streaming Processor (TSP), later rebranded as the LPU. The TSP features a functionally sliced microarchitecture where memory units are interleaved with vector and matrix computation units.</p>\n<p>Unlike systolic arrays where data flows through a fixed grid, tensor streaming processors use a “data-stream” design where tensors flow through functional units in a more flexible pattern. This enables tensor parallelism where individual operations distribute across multiple LPUs so single forward passes complete faster, rather than processing more requests in parallel.</p>\n<p>The first-generation TSP yields computational density exceeding 1 TeraOp/s per square mm of silicon for its 25x29mm 14nm chip operating at 900 MHz nominal clock frequency. The second-generation LPU v2 will use Samsung’s 4nm process node.</p>\n<h3 id=\"sram-only-architecture\">SRAM-Only Architecture</h3>\n<p>The LPU uses SRAM as primary weight storage, not cache, with approximately 230 MB per chip. This is fundamentally different from GPU architectures that use HBM.</p>\n<p>SRAM is up to 100x faster than the HBM found in GPUs. This direct access allows compute units to pull in weights at full speed. There are no caches to miss, no external memory delays to manage. The Groq compiler can plan the exact execution path of every instruction before the chip runs.</p>\n<p>The tradeoff: no useful models fit on a single chip. Groq systems connect hundreds of LPUs with tensor parallelism. For the Mixtral model, Groq connects 8 racks of 9 servers each with 8 chips per server, totaling 576 chips to serve one model instance. The significant cost of SRAM and its lower capacity compared to DRAM present obstacles.</p>\n<h3 id=\"batch-size-1-the-sweet-spot\">Batch Size 1: The Sweet Spot</h3>\n<p>Groq’s architecture is optimized for “Batch Size 1.” Because memory bandwidth is so high and overhead so low, the LPU processes a single user’s request with maximum efficiency without waiting for other requests.</p>\n<p>On GPUs, achieving ultra-low latency requires batch size 1, which makes the GPU expensive per token because most processing power sits idle waiting for memory. The LPU achieves 300-500 tokens per second while keeping its internal pipeline nearly 100% full at batch size 1.</p>\n<p>This design choice means Groq is not competitive architecturally for throughput-optimized scenarios where batching many requests together amortizes memory access costs. Groq targets latency-critical applications.</p>\n<h3 id=\"inference-performance\">Inference Performance</h3>\n<p>Artificial Analysis independently benchmarked Groq serving Llama 3.3 70B at 1,665 output tokens per second using speculative decoding, a 6x improvement over their previous 250 T/s endpoint without speculative decoding.</p>\n<p>Speculative decoding uses a smaller “draft model” (e.g., Llama 8B) to rapidly guess subsequent tokens, which the primary model then verifies. Verification is faster than generation. Groq implemented this without compromising response quality per independent evaluations.</p>\n<p>Other benchmarks:</p>\n<ul>\n<li>Gemma 7B: 814 tokens per second (Artificial Analysis)</li>\n<li>LLaMA 3: over 800 tokens per second</li>\n<li>Llama 2 Chat 70B: 241 tokens per second, more than double other providers (ArtificialAnalysis.ai)</li>\n<li>LLMPerf Leaderboard: 185 tokens/s average, 3-18x faster than other cloud inference providers</li>\n</ul>\n<p>Groq offers Llama-3.3-70B-Specdec at <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>0.59</mn><mi>p</mi><mi>e</mi><mi>r</mi><mi>m</mi><mi>i</mi><mi>l</mi><mi>l</mi><mi>i</mi><mi>o</mi><mi>n</mi><mi>i</mi><mi>n</mi><mi>p</mi><mi>u</mi><mi>t</mi><mi>t</mi><mi>o</mi><mi>k</mi><mi>e</mi><mi>n</mi><mi>s</mi><mi>a</mi><mi>n</mi><mi>d</mi></mrow></semantics></math>0.99 per million output tokens.</p>\n<h3 id=\"limitations\">Limitations</h3>\n<p>Groq is inference-only. You cannot train or fine-tune models on LPUs. The initial capital expenditure for a Groq rack is high due to requiring hundreds of chips per model instance. However, the company claims near-100% compute utilization during inference results in lower energy cost per token than GPUs.</p>\n<p>In December 2025, NVIDIA agreed to purchase assets from Groq for approximately $20 billion in what Groq described as a non-exclusive licensing deal.</p>\n<h2 id=\"apple-silicon-for-ai\">Apple Silicon for AI</h2>\n<p>Apple’s M-series chips offer a different value proposition: running AI workloads on hardware that millions of people already own.</p>\n<h3 id=\"neural-engine-architecture\">Neural Engine Architecture</h3>\n<p>Each M-series chip includes a dedicated Neural Engine for accelerating machine learning operations. The M4 Neural Engine performs 38 trillion operations per second (TOPS), 60x faster than the first Neural Engine in A11 Bionic. For comparison, M1’s Neural Engine achieved 11 TOPS, while M2 and M3 reached 15.8 TOPS.</p>\n<p>The Neural Engine’s design favors specialized convolutional and quantized operations. Large transformer-based LLMs depend heavily on dense matrix multiplications more efficiently handled by GPUs or CPUs, which is why full LLM execution solely on the Neural Engine remains aspirational in 2025, with current deployments relying on hybrid CPU/GPU/ANE orchestration.</p>\n<p>Apple’s AMX (Advanced Matrix Extensions) supports fixed matrix dimensions (e.g., 4x4 or 8x8) for operations. The M4 standardized ARM SME (Scalable Matrix Extension). Core ML blends CPU, GPU, and ANE to create hybrid execution plans exploiting all available engines on a given device.</p>\n<p>Unlike GPUs, there is no public framework for directly programming the Neural Engine. Developers use Core ML, which handles hardware scheduling transparently.</p>\n<h3 id=\"unified-memory-architecture\">Unified Memory Architecture</h3>\n<p>M-series processors use unified memory where CPU, GPU, and Neural Engine share the same pool of high-bandwidth memory. There is no separate “system RAM” and “GPU VRAM.” This eliminates the PCIe bottleneck that limits GPU memory capacity in traditional systems.</p>\n<p>Graphics resources, textures, images, and geometry data can be shared between CPU and GPU with no overhead since there’s no need to copy data across a PCIe bus. Every component in the SoC has direct access to RAM: Media Engine, Neural Engine, video encode/decode units. Code passes a pointer to data in memory and the hardware unit processes data in place.</p>\n<p>The M4 Max provides up to 128 GB of unified memory with 546 GB/s bandwidth. The M3 Ultra in Mac Studio configurations offers up to 192 GB. This enables running models that would require multiple GPUs on conventional hardware:</p>\n<ul>\n<li>Llama 70B runs on Mac Studio M2 Ultra (192 GB) at 8-12 tokens per second</li>\n<li>DeepSeek’s 671B model can run on 512 GB Mac Studio configurations</li>\n<li>Models up to 30B parameters run locally on M4 Max or M3 Ultra</li>\n</ul>\n<h3 id=\"mlx-vs-mps-how-they-exploit-unified-memory-differently\">MLX vs MPS: How They Exploit Unified Memory Differently</h3>\n<p>MPS (Metal Performance Shaders) is an abstraction layer translating standard GPU operations into optimized Metal instructions. It enables frameworks like PyTorch and JAX to run on Apple Silicon with minimal code changes.</p>\n<p>MLX, Apple’s machine learning framework, was designed specifically for Apple Silicon. The key difference is how they handle memory:</p>\n<p><strong>MPS</strong>: Treats the GPU as a separate device following the traditional CUDA model. While it benefits from unified memory eliminating physical data copies, the programming model still conceptually separates CPU and GPU memory spaces.</p>\n<p><strong>MLX</strong>: Natively exploits unified memory at the framework level. Operations can run on CPU or GPU without data movement between memory pools. The framework exposes this capability directly in its API design.</p>\n<p>Benchmarks show MLX achieves the highest sustained generation throughput for LLM inference on Apple Silicon. In comparative testing, MLX was found to be 3x faster than PyTorch MPS on the same M1-Pro GPU.</p>\n<p>The M5 chip (October 2025) provides 19-27% performance improvement over M4 due to increased memory bandwidth (153 GB/s versus 120 GB/s). MLX now takes advantage of Neural Accelerators in M5, which provide dedicated matrix-multiplication operations yielding up to 4x speedup compared to M4 baseline for time-to-first-token.</p>\n<h3 id=\"power-efficiency-analysis\">Power Efficiency Analysis</h3>\n<p>Power consumption provides Apple Silicon’s clearest advantage:</p>\n<ul>\n<li>M3/M4 Max: 40-80W under load during LLM inference</li>\n<li>RTX 4090: up to 450W for the same task</li>\n<li>M3 Max generating from Llama 7B: approximately 50W</li>\n<li>Projected M4 Max: 96-100 tokens/s on 8B Q4_K_M model at similar power</li>\n</ul>\n<p>Research introducing “intelligence per watt” as a metric shows 5.3x improvement from 2023-2025, driven by both algorithmic advances and accelerator improvements. Benchmarks show ResNet-50 runs about 3x slower on Apple Silicon than RTX 4090, but with over 80% lower energy consumption.</p>\n<p>The M3 Max can generate 30-40 tokens per second with quantized Llama 7B while remaining silent and energy efficient.</p>\n<h3 id=\"use-cases-1\">Use Cases</h3>\n<p>For edge deployment, mobile inference, or power-constrained environments, Apple Silicon efficiency matters. The tradeoff is absolute performance: NVIDIA cards remain faster for raw throughput. But for engineers who need local, private inference without dedicated GPU hardware, Apple Silicon provides an accessible path.</p>\n<h2 id=\"amd-gpus-and-rocm\">AMD GPUs and ROCm</h2>\n<p>AMD’s Instinct accelerators offer the most direct competition to NVIDIA in the data center GPU market, though with a significant software ecosystem gap.</p>\n<h3 id=\"cdna-vs-rdna-two-architectures-two-markets\">CDNA vs RDNA: Two Architectures, Two Markets</h3>\n<p>AMD maintains two distinct GPU architectures:</p>\n<p><strong>CDNA (Compute DNA)</strong>: Designed for datacenters, AI, and HPC. Compared to its predecessor GCN, CDNA removed all hardware related to graphics acceleration (graphics caches, tessellation hardware, ROPs, display engine) and added dedicated matrix compute hardware. CDNA has had tensor-like functional units since 2020, with increased throughput and number format support added in CDNA 2 (2021) and CDNA 3 (2023).</p>\n<p><strong>RDNA (Radeon DNA)</strong>: Consumer and workstation GPUs with some AI acceleration. RDNA 3 includes Wave MMA (matrix multiply-accumulate) instructions supporting FP16, BF16, INT8, and INT4 data types, improving inference performance compared to RDNA 2. However, RDNA’s AI acceleration is limited compared to CDNA’s dedicated matrix cores.</p>\n<p>The CDNA 4 architecture (MI350 series, 2025) adds native FP6 and FP4 support. AMD doubled FP8/BF16 throughput and fine-tuned chiplet designs for larger GPU clusters. FP6 processes at twice the rate of FP8 on AMD architecture, unlike NVIDIA where FP6 processes at the same rate as FP8. MI350 offers up to 288 GB HBM3e memory per GPU.</p>\n<p>AMD announced UDNA, which will unify RDNA and CDNA into one microarchitecture, ending the split that began in 2019.</p>\n<h3 id=\"matrix-cores-vs-tensor-cores\">Matrix Cores vs Tensor Cores</h3>\n<p><strong>AMD MI300X (CDNA 3)</strong>:</p>\n<ul>\n<li>4 Matrix Cores per Compute Unit, 304 CUs per GPU</li>\n<li>Multi-chip module design: 8 accelerator complex dies (XCD) on TSMC 5nm</li>\n<li>Each compute die: 38 Compute Units, 4MB L2 cache</li>\n<li>192 GB HBM3 memory, 5.3 TB/s bandwidth</li>\n<li>1.31 petaflops peak at FP16</li>\n</ul>\n<p><strong>NVIDIA H100 (Hopper)</strong>:</p>\n<ul>\n<li>4 Tensor Cores per SM, 132 SMs per GPU</li>\n<li>Tensor Cores support FP8, FP16, BF16, TF32, FP64, INT8 for A/B matrices</li>\n<li>80 GB HBM2e memory, 3.35 TB/s bandwidth</li>\n<li>Fourth-generation Tensor Cores work up to 6x faster between chips compared to A100</li>\n</ul>\n<p>Raw instruction throughput significantly favors MI300X: at times 5x faster than H100, at worst roughly 40% faster for INT32, FP32, FP16, and INT8 compute.</p>\n<p>However, real-world training shows different results. For BF16, H100 and H200 achieve roughly 720 TFLOP/s against their marketed 989.5 TFLOP/s, while MI300X reaches only 620 TFLOP/s compared to marketed 1,307 TFLOP/s. Despite higher marketed specs, MI300X is 14% slower than H100/H200 in practice for training workloads (SemiAnalysis benchmark, December 2024).</p>\n<p>MI300X has better memory bandwidth (5.3 TB/s vs 3.35 TB/s for H100) and 192 GB vs 80 GB capacity, but H100 exhibits 57% lower memory latency.</p>\n<p>For inference, MI300X delivers 40% lower latency for memory-bound Llama2-70B inference and 2.7x faster time to first token for Qwen models. MI300X performs better than H100 SXM at small and large batch sizes (1, 2, 4, 256, 512, 1024) but worse at medium batch sizes.</p>\n<p>Cost comparison: MI300X processes 1 million tokens at <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>11.11</mn><mi>w</mi><mi>i</mi><mi>t</mi><mi>h</mi><mi>b</mi><mi>a</mi><mi>t</mi><mi>c</mi><mi>h</mi><mi>s</mi><mi>i</mi><mi>z</mi><mi>e</mi><mn>4</mn><mo separator=\"true\">,</mo><mi>v</mi><mi>e</mi><mi>r</mi><mi>s</mi><mi>u</mi><mi>s</mi><mi>H</mi><msup><mn>100</mn><mo lspace=\"0em\" rspace=\"0em\">′</mo></msup><mi>s</mi></mrow></semantics></math>14.06, creating a 21% cost advantage for AMD.</p>\n<h3 id=\"rocm-and-hip-the-cuda-compatibility-layer\">ROCm and HIP: The CUDA Compatibility Layer</h3>\n<p>ROCm (Radeon Open Compute) is AMD’s answer to CUDA. HIP (Heterogeneous-compute Interface for Portability) is a C++ runtime API and kernel language enabling platform-independent GPU programs that run on both AMD and NVIDIA GPUs.</p>\n<p>HIP intentionally reuses existing <code>torch.cuda</code> interfaces:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>cuda </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>)</span><span>     # Default HIP device</span></span>\n<span class=\"line\"><span>cuda0 </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda:0'</span><span>)</span><span>  # 'rocm' or 'hip' are not valid, use 'cuda'</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Detecting HIP vs CUDA at runtime</span></span>\n<span class=\"line\"><span>if</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>is_available</span><span>()</span><span> and</span><span> torch</span><span>.</span><span>version</span><span>.</span><span>hip</span><span>:</span></span>\n<span class=\"line\"><span>    # HIP-specific code</span></span>\n<span class=\"line\"><span>elif</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>is_available</span><span>()</span><span> and</span><span> torch</span><span>.</span><span>version</span><span>.</span><span>cuda</span><span>:</span></span>\n<span class=\"line\"><span>    # CUDA-specific code</span></span></code></pre>\n<p>AMD provides HIPIFY tools for translating CUDA code:</p>\n<ul>\n<li><strong>hipify-clang</strong>: Parses CUDA code into an AST, traverses it with transformation matchers, and produces HIP source</li>\n<li><strong>hipify-perl</strong>: Uses pattern matching for simpler translations</li>\n</ul>\n<p>On NVIDIA GPUs, HIP is a thin layer over CUDA, so the two code types interoperate on nvcc platforms. Setting <code>HIP_PLATFORM=amd</code> makes hipcc call the clang compiler and ROCclr runtime; <code>HIP_PLATFORM=nvidia</code> makes hipcc call nvcc.</p>\n<p>HIP 7.0 (second half of 2025) will align HIP C++ more closely with CUDA semantics, refine error handling, and streamline header structures to reduce effort maintaining portable codebases.</p>\n<p>Performance benchmarks in 2025 show CUDA typically outperforms ROCm by 10-30% in compute-intensive workloads, though ROCm has narrowed the gap.</p>\n<h3 id=\"infinity-fabric-interconnect\">Infinity Fabric Interconnect</h3>\n<p>Each MI300X has up to seven Infinity Fabric links, each with 16 lanes. Fourth-generation Infinity Fabric supports up to 32 Gbps per lane, yielding 128 GB/s bidirectional bandwidth per link.</p>\n<p>The ThinkSystem SR685a V3 includes 8x MI300X GPUs fully interconnected using Infinity Fabric, providing 128 GB/s bandwidth between each of the 8 GPUs for a total of 896 GB/s. Aggregated theoretical bandwidth per GPU reaches 336 GB/s; practical execution yields 310-330 GB/s.</p>\n<p>xGMI (External Global Memory Interconnect) enables efficient peer-to-peer GPU communication. AMD’s Infinity Fabric links up to eight MI300X or MI355X GPUs in a fully connected mesh.</p>\n<p>For cluster scaling, AMD’s high-speed accelerator network consists of GPUs connected via Infinity Fabric mesh for low-latency inter-GPU communication. The backend scale-out network uses PCIe 5.0 NICs with RoCEv2, scaling clusters while minimizing congestion. ND MI300X v5 deployments on Azure scale to thousands of GPUs with 3.2 Tb/s interconnect bandwidth per VM.</p>\n<p>Fifth-generation Infinity Fabric will connect CPUs, GPUs, and accelerators from node to rack scale, forming the backbone of “Helios” rack deployments.</p>\n<h3 id=\"roadmap\">Roadmap</h3>\n<ul>\n<li><strong>MI325X</strong> (October 2024): 256 GB HBM3E, 6 TB/s bandwidth</li>\n<li><strong>MI350</strong> (2025): CDNA 4 architecture with FP4/FP6 precision formats</li>\n<li><strong>MI355X</strong> (late 2025): Projected 2.7x tokens per second versus MI325X</li>\n</ul>\n<p>AMD holds approximately 7% of the AI GPU market.</p>\n<h2 id=\"intel-gaudi-and-arc\">Intel: Gaudi and Arc</h2>\n<p>Intel pursues AI acceleration through two product lines: Gaudi for data center training and inference, and Arc for workstation and consumer applications.</p>\n<h3 id=\"gaudi-3-accelerator\">Gaudi 3 Accelerator</h3>\n<p>The Gaudi 3, manufactured on TSMC 5nm, uses a different architecture from GPUs:</p>\n<ul>\n<li>64 tensor processor cores (256x256 MAC structure)</li>\n<li>8 matrix multiplication engines (256-bit vector processors)</li>\n<li>128 GB HBM2e with 3.7 TB/s bandwidth</li>\n<li>96 MB on-die SRAM cache with 19.2 TB/s bandwidth</li>\n<li>1,835 BF16/FP8 TFLOPS at 600W TDP</li>\n<li>24 integrated 200 GbE networking interfaces</li>\n</ul>\n<p>Intel claims Gaudi 3 delivers 50% better average inference performance and 40% better power efficiency than NVIDIA H100 at lower cost. Dell’s AI platform with Gaudi 3 reports 70% better price-performance for Llama 3 80B inference throughput versus H100.</p>\n<p>Gaudi 3 entered volume production in Q3 2024 and is available through OEM systems and IBM Cloud.</p>\n<h3 id=\"arc-gpus-and-oneapi\">Arc GPUs and oneAPI</h3>\n<p>Intel’s Arc GPUs target workstation AI inference rather than data center training. The Arc Pro B-Series, launched in May 2025:</p>\n<ul>\n<li><strong>Arc Pro B60</strong>: 24 GB memory, up to 2.7x faster than NVIDIA RTX A2000 Ada on LLM workloads</li>\n<li><strong>Arc Pro B50</strong>: 16 GB memory</li>\n</ul>\n<p>The oneAPI programming model provides a unified interface across Intel CPUs, GPUs, and accelerators. The oneDNN library and PyTorch 2.7 optimizations support inference on Intel Core Ultra processors and Arc GPUs. The Arc A770 (16 GB) can run Llama2 7B and Llama3 models locally.</p>\n<h2 id=\"other-players\">Other Players</h2>\n<p>Several companies are developing alternative AI accelerator architectures.</p>\n<h3 id=\"sambanova-systems\">SambaNova Systems</h3>\n<p>SambaNova’s Reconfigurable Dataflow Processing Unit (RDPU) emphasizes software-defined hardware. Their SN40L RDU supports up to 5 trillion parameters in a single system node. The company achieved 129 tokens/second on 405B-parameter models and was named “Most Respected Private Semiconductor Company” at the 2024 GSA Awards.</p>\n<h3 id=\"graphcore\">Graphcore</h3>\n<p>Graphcore’s Intelligence Processing Unit (IPU) emphasizes fine-grained parallelism and in-processor memory. The architecture keeps more of the model on-chip to reduce memory bottlenecks. SoftBank acquired Graphcore in July 2024.</p>\n<h3 id=\"tenstorrent\">Tenstorrent</h3>\n<p>Led by chip architect Jim Keller, Tenstorrent designs RISC-V based AI chips. The company raised over <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>693</mn><mi>m</mi><mi>i</mi><mi>l</mi><mi>l</mi><mi>i</mi><mi>o</mi><mi>n</mi><mi>i</mi><mi>n</mi><mi>D</mi><mi>e</mi><mi>c</mi><mi>e</mi><mi>m</mi><mi>b</mi><mi>e</mi><mi>r</mi><mn>2024</mn><mi>a</mi><mi>t</mi><mi>a</mi></mrow></semantics></math>2.6+ billion valuation, with investors including Samsung, Bezos Expeditions, LG Electronics, Hyundai, and Fidelity. Tenstorrent plans to release a new AI processor every two years and has launched Grayskull and Wormhole chips.</p>\n<h2 id=\"what-this-means-for-engineers\">What This Means for Engineers</h2>\n<p>Hardware selection depends on workload characteristics, scale, and constraints.</p>\n<h3 id=\"training-workloads\">Training Workloads</h3>\n<p><strong>Large-scale training (100B+ parameters)</strong>: Google TPUs or NVIDIA H100/H200 remain the practical choices. TPUs offer cost advantages for organizations willing to adopt JAX; NVIDIA provides the most mature ecosystem.</p>\n<p><strong>Medium-scale training</strong>: AMD MI300X provides a cost-effective alternative if your workloads align with ROCm’s supported configurations. Intel Gaudi 3 offers competitive price-performance for supported model architectures.</p>\n<p><strong>Experimentation and research</strong>: Cloud TPU access through Google Cloud, or NVIDIA GPUs through various providers, offer flexibility without capital commitment.</p>\n<h3 id=\"inference-workloads\">Inference Workloads</h3>\n<p><strong>Latency-critical inference</strong>: Groq’s LPU delivers unmatched speed for supported models. Cerebras CS-3 offers 21x faster inference than Blackwell B200. If your application requires sub-second response times at scale, evaluate both.</p>\n<p><strong>Cost-optimized inference</strong>: TPU v5e and Trillium provide strong price-performance for batch inference. Organizations report 60-70% cost reductions migrating from NVIDIA to TPUs.</p>\n<p><strong>Large model inference</strong>: Cerebras CS-3 handles models that exceed typical GPU memory. The MI300X’s 192 GB HBM enables single-GPU deployment of 70B parameter models.</p>\n<p><strong>Edge and local inference</strong>: Apple Silicon with MLX provides the most accessible path for on-device inference. Intel Arc and Gaudi enable Windows-based local deployment.</p>\n<h3 id=\"targeting-different-hardware-code-examples\">Targeting Different Hardware: Code Examples</h3>\n<p><strong>JAX/TPU</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> jax</span></span>\n<span class=\"line\"><span>import</span><span> jax</span><span>.</span><span>numpy </span><span>as</span><span> jnp</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@jax</span><span>.</span><span>jit</span></span>\n<span class=\"line\"><span>def</span><span> matmul</span><span>(</span><span>x</span><span>,</span><span> y</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> jnp</span><span>.</span><span>matmul</span><span>(x, y)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Runs on TPU if available, compiles via XLA</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> matmul</span><span>(jnp.</span><span>ones</span><span>((</span><span>1024</span><span>, </span><span>1024</span><span>)), jnp.</span><span>ones</span><span>((</span><span>1024</span><span>, </span><span>1024</span><span>)))</span></span></code></pre>\n<p><strong>PyTorch/ROCm (AMD)</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Same code as CUDA, HIP translates automatically</span></span>\n<span class=\"line\"><span>device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>)</span><span>  # Not 'rocm' or 'hip'</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>1024</span><span>, </span><span>1024</span><span>, device</span><span>=</span><span>device)</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>1024</span><span>, </span><span>1024</span><span>, device</span><span>=</span><span>device)</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> torch</span><span>.</span><span>matmul</span><span>(x, y)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Check which backend at runtime</span></span>\n<span class=\"line\"><span>if</span><span> torch</span><span>.</span><span>version</span><span>.</span><span>hip</span><span>:</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"Running on AMD via HIP\"</span><span>)</span></span></code></pre>\n<p><strong>MLX (Apple Silicon)</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Operations seamlessly use CPU or GPU via unified memory</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> mx</span><span>.</span><span>random</span><span>.</span><span>normal</span><span>((</span><span>1024</span><span>, </span><span>1024</span><span>))</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> mx</span><span>.</span><span>random</span><span>.</span><span>normal</span><span>((</span><span>1024</span><span>, </span><span>1024</span><span>))</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> mx</span><span>.</span><span>matmul</span><span>(x, y)</span></span>\n<span class=\"line\"><span>mx</span><span>.</span><span>eval</span><span>(result)</span><span>  # Explicit evaluation</span></span></code></pre>\n<h3 id=\"cost-considerations\">Cost Considerations</h3>\n<p>GPU cloud pricing in 2025:</p>\n<ul>\n<li>H100: $2.10-7.57 per GPU-hour depending on provider and commitment</li>\n<li>H200: From $2.14-2.50 per GPU-hour</li>\n<li>TPU v5e: Competitive with H100 at large scale</li>\n<li>MI300X: $3.99 per GPU-hour (RunPod)</li>\n</ul>\n<p>For most organizations, rental makes more sense than purchase given hardware obsolescence risk and 6-12 month procurement timelines. Hidden costs include data egress fees, storage, and idle time; many teams waste 30-50% of budget on provisioned but unused capacity.</p>\n<p>The AI hardware market is fragmenting. NVIDIA’s dominance will likely persist for training workloads requiring maximum flexibility, but specialized inference hardware, cloud TPUs, and alternative architectures are capturing significant workloads. Engineers who understand these options can optimize for their specific constraints rather than defaulting to the most expensive general-purpose solution.</p>\n<h2 id=\"references\">References</h2>\n<ol>\n<li><a href=\"https://cloud.google.com/blog/products/compute/introducing-trillium-6th-gen-tpus\">Google Cloud Blog: Introducing Trillium</a></li>\n<li><a href=\"https://docs.cloud.google.com/tpu/docs/v5p\">Google Cloud TPU v5p Documentation</a></li>\n<li><a href=\"https://docs.cloud.google.com/tpu/docs/v5e\">Google Cloud TPU v5e Documentation</a></li>\n<li><a href=\"https://docs.cloud.google.com/tpu/docs/system-architecture-tpu-vm\">Google Cloud TPU Architecture</a></li>\n<li><a href=\"https://cloud.google.com/blog/products/ai-machine-learning/bfloat16-the-secret-to-high-performance-on-cloud-tpus\">Google Cloud Blog: BFloat16 on Cloud TPUs</a></li>\n<li><a href=\"https://cloud.google.com/blog/products/compute/trillium-mlperf-41-training-benchmarks\">Google Cloud Blog: Trillium MLPerf 4.1 Training Benchmarks</a></li>\n<li><a href=\"https://www.cerebras.ai/blog/cerebras-cs3\">Cerebras CS-3 Announcement</a></li>\n<li><a href=\"https://www.cerebras.ai/blog/100x-defect-tolerance-how-cerebras-solved-the-yield-problem\">Cerebras: 100x Defect Tolerance</a></li>\n<li><a href=\"https://docs.cerebras.net/en/2.0.3/wsc/cerebras-basics/cerebras-execution-modes.html\">Cerebras Weight Streaming Documentation</a></li>\n<li><a href=\"https://www.cerebras.ai/press-release/cerebras-inference-llama-405b\">Cerebras Llama 3.1 405B Performance</a></li>\n<li><a href=\"https://www.cerebras.ai/blog/cerebras-cs-3-vs-nvidia-dgx-b200-blackwell\">Cerebras CS-3 vs NVIDIA DGX B200</a></li>\n<li><a href=\"https://groq.com/lpu-architecture\">Groq LPU Architecture</a></li>\n<li><a href=\"https://groq.com/blog/inside-the-lpu-deconstructing-groq-speed\">Groq: Inside the LPU</a></li>\n<li><a href=\"https://groq.com/blog/groq-first-generation-14nm-chip-just-got-a-6x-speed-boost-introducing-llama-3-1-70b-speculative-decoding-on-groqcloud\">Groq Speculative Decoding Launch</a></li>\n<li><a href=\"https://newsletter.semianalysis.com/p/groq-inference-tokenomics-speed-but\">SemiAnalysis: Groq Inference Tokenomics</a></li>\n<li><a href=\"https://www.apple.com/newsroom/2024/05/apple-introduces-m4-chip/\">Apple M4 Introduction</a></li>\n<li><a href=\"https://www.apple.com/newsroom/2025/10/apple-unleashes-m5-the-next-big-leap-in-ai-performance-for-apple-silicon/\">Apple M5 Announcement</a></li>\n<li><a href=\"https://machinelearning.apple.com/research/exploring-llms-mlx-m5\">Apple MLX Research on M5</a></li>\n<li><a href=\"https://towardsdatascience.com/mlx-vs-mps-vs-cuda-a-benchmark-c5737ca6efc9/\">MLX vs MPS vs CUDA Benchmark</a></li>\n<li><a href=\"https://rocm.docs.amd.com/en/latest/conceptual/gpu-arch/mi300.html\">AMD MI300X Specifications</a></li>\n<li><a href=\"https://newsletter.semianalysis.com/p/mi300x-vs-h100-vs-h200-benchmark-part-1-training\">SemiAnalysis MI300X vs H100 Benchmark</a></li>\n<li><a href=\"https://www.amd.com/en/blogs/2024/engineering-insights-unveiling-mlperf-results-on.html\">AMD MLPerf Results</a></li>\n<li><a href=\"https://docs.pytorch.org/docs/stable/notes/hip.html\">PyTorch HIP Semantics</a></li>\n<li><a href=\"https://rocm.docs.amd.com/projects/HIP/en/latest/how-to/hip_porting_guide.html\">ROCm HIP Porting Guide</a></li>\n<li><a href=\"https://rocm.blogs.amd.com/software-tools-optimization/mi300x-rccl-xgmi/README.html\">AMD Infinity Fabric Interconnect</a></li>\n<li><a href=\"https://www.intel.com/content/www/us/en/content-details/817486/intel-gaudi-3-ai-accelerator-white-paper.html\">Intel Gaudi 3 White Paper</a></li>\n<li><a href=\"https://www.intc.com/news-events/press-releases/detail/1741/computex-2025-intel-unveils-new-gpus-for-ai-and\">Intel Arc Pro B-Series Launch</a></li>\n<li><a href=\"https://carboncredits.com/nvidia-controls-92-of-the-gpu-market-in-2025-and-reveals-next-gen-ai-supercomputer/\">NVIDIA Market Share Analysis</a></li>\n<li><a href=\"https://finance.yahoo.com/news/nvidias-big-tech-customers-might-also-be-its-biggest-competitive-threat-153032596.html\">JPMorgan Custom Chip Projections</a></li>\n<li><a href=\"https://pitchbook.com/profiles/company/171627-13\">Tenstorrent Funding</a></li>\n<li><a href=\"https://www.gmicloud.ai/blog/a-guide-to-2025-gpu-cloud-pricing-comparison\">GPU Cloud Pricing Comparison 2025</a></li>\n<li><a href=\"https://jax-ml.github.io/scaling-book/tpus/\">JAX Framework Documentation</a></li>\n<li><a href=\"https://docs.jax.dev/en/latest/pallas/tpu/matmul.html\">JAX Matrix Multiplication on TPU</a></li>\n<li><a href=\"https://henryhmko.github.io/posts/tpu/tpu.html\">TPU Deep Dive</a></li>\n<li><a href=\"https://arxiv.org/abs/2105.04663\">GSPMD Paper (Xu et al., 2021)</a></li>\n<li><a href=\"https://arxiv.org/abs/2510.26788\">BF16 Training-Inference Mismatch (arXiv<div></div>.26788)</a></li>\n</ol>",
            "url": "https://blog.ecitis.org/ai-hardware-landscape/",
            "title": "The Diversification of AI Hardware: Beyond NVIDIA's Dominance",
            "summary": "Compare the architectures, economics, and tradeoffs behind GPUs, TPUs, custom accelerators, and the challengers to NVIDIA.",
            "image": "https://blog.ecitis.org/open-graph/ai-hardware-landscape.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Hardware & Systems",
                "Accelerators",
                "TPU",
                "AI chips"
            ]
        },
        {
            "id": "https://blog.ecitis.org/amd-rocm-ml/",
            "content_html": "<p>AMD’s ROCm platform has matured into a viable alternative for machine learning workloads. With ROCm 7.2 now shipping production-ready support for PyTorch, vLLM, and distributed training on both Linux and Windows, the AMD GPU ecosystem deserves serious consideration for both inference and training deployments. The October 2025 announcement of AMD’s strategic partnership with OpenAI to deploy 6 gigawatts of AMD GPUs signals a significant shift in the AI infrastructure landscape.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-openai\" id=\"user-content-fnref-openai\">1</a></sup></p>\n<p>This guide covers everything you need to deploy ML workloads on AMD hardware: the software stack, hardware options, framework support, and practical code examples.</p>\n<hr />\n<h2 id=\"the-rocm-software-stack\">The ROCm Software Stack</h2>\n<p>ROCm (Radeon Open Compute) is AMD’s open-source GPU computing platform. Unlike NVIDIA’s proprietary CUDA ecosystem, ROCm is built on open standards and released under permissive licenses. The stack comprises several layers, from kernel drivers through high-level frameworks.</p>\n<figure><figcaption><strong>ROCm Software Stack</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/amd-rocm-ml-0.webp\" width=\"512\" height=\"692\" alt=\"ROCm Software Stack\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"key-components\">Key Components</h3>\n<p><strong>amdgpu Kernel Driver</strong>: The Linux kernel module that communicates with AMD GPUs. This ships with most Linux distributions and handles memory management, command submission, and interrupt handling.</p>\n<p><strong>HSA Runtime (ROCr)</strong>: The Heterogeneous System Architecture runtime provides low-level GPU access. Applications rarely interact with this layer directly; it serves as the foundation for higher-level runtimes.</p>\n<p><strong>HIP Runtime</strong>: The primary programming interface for AMD GPUs. HIP provides a CUDA-like API that can target both AMD and NVIDIA hardware from the same source code.</p>\n<p><strong>ROCm Libraries</strong>: Optimized implementations of common operations:</p>\n<ul>\n<li><strong>rocBLAS/hipBLAS</strong>: Dense linear algebra (GEMM, GEMV)</li>\n<li><strong>hipBLASLt</strong>: Lightweight BLAS with fused operations and FP8 support</li>\n<li><strong>MIOpen</strong>: Deep learning primitives (convolutions, pooling, normalization)</li>\n<li><strong>RCCL</strong>: Collective communication for multi-GPU training</li>\n<li><strong>rocFFT</strong>: Fast Fourier transforms</li>\n</ul>\n<p><strong>Framework Integrations</strong>: PyTorch, TensorFlow, and JAX all have ROCm backends. These frameworks use MIOpen and rocBLAS under the hood for GPU-accelerated operations.</p>\n<hr />\n<h2 id=\"supported-hardware\">Supported Hardware</h2>\n<p>ROCm supports both data center accelerators and consumer graphics cards, though with different levels of optimization and testing.</p>\n<figure><figcaption><strong>Data Center GPU Comparison</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/amd-rocm-ml-1.webp\" width=\"512\" height=\"344\" alt=\"Data Center GPU Comparison\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"data-center-amd-instinct-series\">Data Center: AMD Instinct Series</h3>\n<p><strong>MI355X</strong>: AMD’s current flagship for AI workloads, launched Q3 2025. Built on the CDNA 4 architecture with 256 Compute Units and 288 GB of HBM3e memory at 8 TB/s bandwidth. The MI355X delivers 10.1 PFLOPS at FP8 and 20.1 PFLOPS at FP4/FP6. In benchmarks, the MI355X matches or exceeds NVIDIA B200 systems at high concurrency levels (32-64) for Llama3.1-405B inference.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-mi355x\" id=\"user-content-fnref-mi355x\">2</a></sup> The 1400W TDP reflects its position as a high-performance training and inference accelerator.</p>\n<p><strong>MI350X</strong>: The lower-power variant of the MI350 series with the same 288 GB HBM3e and 8 TB/s bandwidth. Both MI350 models feature MXFP6 and MXFP4 datatype support for efficient inference.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-mi350\" id=\"user-content-fnref-mi350\">3</a></sup></p>\n<p><strong>MI325X</strong>: The 256 GB HBM3e variant of the MI300 series with 6.0 TB/s bandwidth. In MLPerf benchmarks, systems built around the MI325X matched the performance of NVIDIA H200 on Llama fine-tuning tasks.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-mi325x\" id=\"user-content-fnref-mi325x\">4</a></sup></p>\n<p><strong>MI300X</strong>: With 192 GB of HBM3 memory and 5.3 TB/s bandwidth, the MI300X remains widely deployed. In MLPerf inference benchmarks, Mango LLMBoost achieved 103,182 tokens per second on 32x MI300X for Llama2-70B, outperforming the previous best result of 82,749 TPS on NVIDIA H100.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-mi300x\" id=\"user-content-fnref-mi300x\">5</a></sup></p>\n<p><strong>MI400 Series (H2 2026)</strong>: The upcoming CDNA 5 architecture will deliver 432 GB of HBM4 memory at 19.6 TB/s bandwidth, with 40 PFLOPS at FP4. The MI455X targets training and inference, while the MI430X serves HPC workloads. The first 1 GW deployment at OpenAI begins H2 2026 using MI450 GPUs.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-mi400\" id=\"user-content-fnref-mi400\">6</a></sup></p>\n<p><strong>MI250/MI250X</strong>: Previous generation CDNA 2 accelerators with 128 GB HBM2e. Still supported by ROCm 7.2 and available at lower cost on the secondary market.</p>\n<h3 id=\"consumer-gpus-rdna-3-and-rdna-4\">Consumer GPUs: RDNA 3 and RDNA 4</h3>\n<p><strong>RX 9070 XT/RX 9070 (RDNA 4)</strong>: The latest consumer architecture with 16 GB GDDR6 at 640 GB/s bandwidth. The RX 9070 XT delivers 48.7 TFLOPS FP32 and up to 1557 TOPS INT4 with sparsity. Each compute unit includes dedicated AI accelerators, marking AMD’s first consumer GPU with tensor cores for machine learning. ROCm 7.2 supports these cards on both Linux and Windows. In Stable Diffusion XL FP16 testing, the RX 9070 XT is 83% faster than the RX 7800 XT. You can run up to 24B parameter LLM models at practical quantization levels.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-rx9070\" id=\"user-content-fnref-rx9070\">7</a></sup></p>\n<p><strong>RX 7900 XTX/XT (RDNA 3)</strong>: Consumer cards with up to 24 GB GDDR6. ROCm support is mature; these cards work well for local LLM inference and development. The ROCm llama.cpp builds target gfx1100/gfx1101 architectures.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-4\" id=\"user-content-fnref-4\">8</a></sup></p>\n<p><strong>Radeon AI PRO R9700</strong>: Professional card with 32 GB VRAM for larger models without quantization. The R9700 targets AI development workflows that exceed the 16 GB consumer cards.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-radeonai\" id=\"user-content-fnref-radeonai\">9</a></sup></p>\n<p><strong>Radeon PRO W7900</strong>: Professional card with 48 GB VRAM. Useful for running larger quantized models locally without the cost of data center hardware.</p>\n<h3 id=\"what-works-best-where\">What Works Best Where</h3>\n<p>Data center MI350/MI355X: Production LLM inference and training, frontier model workloads requiring 288 GB memory per GPU.</p>\n<p>Data center MI300X/MI325X: Production LLM inference, distributed training, memory-bound workloads requiring 192-256 GB VRAM pools.</p>\n<p>Consumer RDNA 4 (RX 9070 series): Local development, small model fine-tuning, inference with quantized models up to 24B parameters. Native Windows support via ROCm 7.2.</p>\n<p>Consumer RDNA 3 (RX 7900 series): Local development, inference with quantized models. 24 GB VRAM on the XTX model.</p>\n<hr />\n<h2 id=\"hip-the-cuda-portability-layer\">HIP: The CUDA Portability Layer</h2>\n<p>HIP (Heterogeneous-compute Interface for Portability) provides a CUDA-like programming model that can compile for both AMD and NVIDIA GPUs. Most CUDA code can be ported with minimal modifications.</p>\n<h3 id=\"how-hip-works\">How HIP Works</h3>\n<p>HIP is a C++ runtime API and kernel language. When targeting AMD GPUs, HIP code compiles to native AMD GPU binaries. When targeting NVIDIA, HIP wraps CUDA operations with thin inline functions, adding negligible overhead.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-6\" id=\"user-content-fnref-6\">10</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// HIP kernel - nearly identical to CUDA</span></span>\n<span class=\"line\"><span>__global__ </span><span>void</span><span> vectorAdd</span><span>(</span><span>float</span><span> *</span><span>a</span><span>,</span><span> float</span><span> *</span><span>b</span><span>,</span><span> float</span><span> *</span><span>c</span><span>,</span><span> int</span><span> n) {</span></span>\n<span class=\"line\"><span>    int</span><span> i </span><span>=</span><span> blockIdx</span><span>.</span><span>x </span><span>*</span><span> blockDim</span><span>.</span><span>x </span><span>+</span><span> threadIdx</span><span>.</span><span>x;</span></span>\n<span class=\"line\"><span>    if</span><span> (i </span><span>&lt;</span><span> n) {</span></span>\n<span class=\"line\"><span>        c</span><span>[i] </span><span>=</span><span> a</span><span>[i] </span><span>+</span><span> b</span><span>[i];</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>int</span><span> main</span><span>() {</span></span>\n<span class=\"line\"><span>    float</span><span> *</span><span>d_a</span><span>,</span><span> *</span><span>d_b</span><span>,</span><span> *</span><span>d_c;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Memory allocation - same as CUDA</span></span>\n<span class=\"line\"><span>    hipMalloc</span><span>(</span><span>&amp;</span><span>d_a</span><span>,</span><span> size);</span></span>\n<span class=\"line\"><span>    hipMalloc</span><span>(</span><span>&amp;</span><span>d_b</span><span>,</span><span> size);</span></span>\n<span class=\"line\"><span>    hipMalloc</span><span>(</span><span>&amp;</span><span>d_c</span><span>,</span><span> size);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Kernel launch - same syntax</span></span>\n<span class=\"line\"><span>    vectorAdd</span><span>&lt;&lt;&lt;</span><span>blocks</span><span>,</span><span> threads</span><span>&gt;&gt;&gt;</span><span>(d_a</span><span>,</span><span> d_b</span><span>,</span><span> d_c</span><span>,</span><span> n);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    hipDeviceSynchronize</span><span>();</span></span>\n<span class=\"line\"><span>    hipFree</span><span>(d_a);</span></span>\n<span class=\"line\"><span>    return</span><span> 0</span><span>;</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"hipify-automated-cuda-translation\">HIPIFY: Automated CUDA Translation</h3>\n<p>ROCm includes two tools for converting CUDA code to HIP:</p>\n<p><strong>hipify-clang</strong>: Uses the Clang compiler to parse CUDA code and perform semantic translation. This approach handles complex code patterns and produces high-quality output. Recommended for large projects.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-7\" id=\"user-content-fnref-7\">11</a></sup></p>\n<p><strong>hipify-perl</strong>: Pattern-matching based translation that does not require a working CUDA installation. Faster but less robust than hipify-clang.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Convert a CUDA file to HIP</span></span>\n<span class=\"line\"><span>hipify-clang</span><span> cuda_kernel.cu</span><span> -o</span><span> hip_kernel.cpp</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or use the Perl-based tool</span></span>\n<span class=\"line\"><span>hipify-perl</span><span> cuda_kernel.cu</span><span> &gt;</span><span> hip_kernel.cpp</span></span></code></pre>\n<p>When porting the HACC physics code (approximately 15,000 lines), AMD reported that 95% of the code converted automatically with hipify-perl.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-8\" id=\"user-content-fnref-8\">12</a></sup></p>\n<h3 id=\"what-translates-well\">What Translates Well</h3>\n<p>Most CUDA runtime API calls have direct HIP equivalents:</p>\n<ul>\n<li><code>cudaMalloc</code> → <code>hipMalloc</code></li>\n<li><code>cudaMemcpy</code> → <code>hipMemcpy</code></li>\n<li><code>cudaDeviceSynchronize</code> → <code>hipDeviceSynchronize</code></li>\n<li><code>cudaStream_t</code> → <code>hipStream_t</code></li>\n</ul>\n<p>Device code keywords map directly:</p>\n<ul>\n<li><code>__global__</code>, <code>__device__</code>, <code>__shared__</code> work unchanged</li>\n<li><code>threadIdx</code>, <code>blockIdx</code>, <code>blockDim</code> have <code>hip</code> prefixes available but original names also work</li>\n</ul>\n<h3 id=\"what-requires-manual-work\">What Requires Manual Work</h3>\n<p><strong>Warp Size</strong>: NVIDIA GPUs use 32-thread warps; AMD uses 64-thread wavefronts. Code that hardcodes <code>warpSize = 32</code> will break. Use the runtime query instead:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// Portable warp/wavefront size</span></span>\n<span class=\"line\"><span>int</span><span> warpSize;</span></span>\n<span class=\"line\"><span>hipDeviceGetAttribute</span><span>(</span><span>&amp;</span><span>warpSize</span><span>,</span><span> hipDeviceAttributeWarpSize</span><span>,</span><span> device);</span></span></code></pre>\n<p><strong>CUDA-Specific Intrinsics</strong>: Some CUDA intrinsics lack direct HIP equivalents. Tensor core operations (wmma) require translation to AMD’s matrix core instructions.</p>\n<p><strong>Library Calls</strong>: ROCm provides equivalent libraries for most CUDA libraries:</p>\n<ul>\n<li>cuBLAS → rocBLAS/hipBLAS</li>\n<li>cuDNN → MIOpen</li>\n<li>NCCL → RCCL</li>\n<li>cuFFT → rocFFT</li>\n</ul>\n<p>Function signatures differ slightly; wrapper headers can ease the transition.</p>\n<hr />\n<h2 id=\"installing-rocm-on-linux-and-windows\">Installing ROCm on Linux and Windows</h2>\n<p>ROCm 7.2 is a unified release supporting both Linux and Windows. On Linux, it officially supports Ubuntu 22.04/24.04, RHEL 9.x/10.x, SLES 15 SP7, Debian 13, and Oracle Linux 10. Consumer Radeon GPUs (RX 7000/9000 series) and Ryzen AI APUs are supported on Ubuntu 24.04, RHEL 10.1, and Windows 11.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-rocm72\" id=\"user-content-fnref-rocm72\">13</a></sup></p>\n<h3 id=\"ubuntu-2404-installation\">Ubuntu 24.04 Installation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Download the installer package</span></span>\n<span class=\"line\"><span>wget</span><span> https://repo.radeon.com/amdgpu-install/7.2/ubuntu/noble/amdgpu-install_7.2.70200-1_all.deb</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install the package manager</span></span>\n<span class=\"line\"><span>sudo</span><span> apt</span><span> install</span><span> ./amdgpu-install_7.2.70200-1_all.deb</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Update package lists</span></span>\n<span class=\"line\"><span>sudo</span><span> apt</span><span> update</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install ROCm</span></span>\n<span class=\"line\"><span>sudo</span><span> apt</span><span> install</span><span> rocm</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Add user to required groups</span></span>\n<span class=\"line\"><span>sudo</span><span> usermod</span><span> -a</span><span> -G</span><span> render,video</span><span> $LOGNAME</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Reboot to load the new kernel modules</span></span>\n<span class=\"line\"><span>sudo</span><span> reboot</span></span></code></pre>\n<h3 id=\"ubuntu-2204-installation\">Ubuntu 22.04 Installation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Download for jammy</span></span>\n<span class=\"line\"><span>wget</span><span> https://repo.radeon.com/amdgpu-install/7.2/ubuntu/jammy/amdgpu-install_7.2.70200-1_all.deb</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>sudo</span><span> apt</span><span> install</span><span> ./amdgpu-install_7.2.70200-1_all.deb</span></span>\n<span class=\"line\"><span>sudo</span><span> apt</span><span> update</span></span>\n<span class=\"line\"><span>sudo</span><span> apt</span><span> install</span><span> python3-setuptools</span><span> python3-wheel</span></span>\n<span class=\"line\"><span>sudo</span><span> usermod</span><span> -a</span><span> -G</span><span> render,video</span><span> $LOGNAME</span></span>\n<span class=\"line\"><span>sudo</span><span> apt</span><span> install</span><span> rocm</span></span>\n<span class=\"line\"><span>sudo</span><span> reboot</span></span></code></pre>\n<h3 id=\"rhel-9x-installation\">RHEL 9.x Installation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Install the repo</span></span>\n<span class=\"line\"><span>sudo</span><span> dnf</span><span> install</span><span> https://repo.radeon.com/amdgpu-install/7.2/el/9.6/amdgpu-install-7.2.70200-1.el9.noarch.rpm</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Clean cache</span></span>\n<span class=\"line\"><span>sudo</span><span> dnf</span><span> clean</span><span> all</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install EPEL for dependencies</span></span>\n<span class=\"line\"><span>wget</span><span> https://dl.fedoraproject.org/pub/epel/epel-release-latest-9.noarch.rpm</span></span>\n<span class=\"line\"><span>sudo</span><span> rpm</span><span> -ivh</span><span> epel-release-latest-9.noarch.rpm</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Enable CRB repository</span></span>\n<span class=\"line\"><span>sudo</span><span> dnf</span><span> install</span><span> dnf-plugin-config-manager</span></span>\n<span class=\"line\"><span>sudo</span><span> crb</span><span> enable</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install dependencies and ROCm</span></span>\n<span class=\"line\"><span>sudo</span><span> dnf</span><span> install</span><span> python3-setuptools</span><span> python3-wheel</span></span>\n<span class=\"line\"><span>sudo</span><span> usermod</span><span> -a</span><span> -G</span><span> render,video</span><span> $LOGNAME</span></span>\n<span class=\"line\"><span>sudo</span><span> dnf</span><span> install</span><span> rocm</span></span>\n<span class=\"line\"><span>sudo</span><span> reboot</span></span></code></pre>\n<h3 id=\"windows-11-installation-consumer-gpus\">Windows 11 Installation (Consumer GPUs)</h3>\n<p>ROCm 7.2 includes Windows support for consumer GPUs. PyTorch runs natively on RX 7000/9000 series without WSL2.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-rocmwindows\" id=\"user-content-fnref-rocmwindows\">14</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Install AMD Adrenalin 26.1.1 or later (includes ROCm components)</span></span>\n<span class=\"line\"><span># Download from https://www.amd.com/en/support</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install PyTorch for ROCm on Windows</span></span>\n<span class=\"line\"><span>pip install torch torchvision torchaudio </span><span>--</span><span>index</span><span>-</span><span>url https:</span><span>//</span><span>download.pytorch.org</span><span>/</span><span>whl</span><span>/</span><span>rocm7.</span><span>2</span></span></code></pre>\n<p>ComfyUI is now integrated with ROCm and can be installed via the Adrenalin driver package for Windows users.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-comfyui\" id=\"user-content-fnref-comfyui\">15</a></sup></p>\n<h3 id=\"verifying-the-installation\">Verifying the Installation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Check GPU detection</span></span>\n<span class=\"line\"><span>rocminfo</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Verify OpenCL</span></span>\n<span class=\"line\"><span>clinfo</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Check GPU status</span></span>\n<span class=\"line\"><span>amd-smi</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># List available GPUs</span></span>\n<span class=\"line\"><span>rocm-smi</span><span> --showid</span></span></code></pre>\n<p>Expected output from <code>rocminfo</code> should list your GPU agent with details like:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>Agent 2</span></span>\n<span class=\"line\"><span>  Name:                    gfx942</span></span>\n<span class=\"line\"><span>  Marketing Name:          AMD Instinct MI300X</span></span>\n<span class=\"line\"><span>  Vendor Name:             AMD</span></span>\n<span class=\"line\"><span>  ...</span></span></code></pre>\n<hr />\n<h2 id=\"pytorch-with-rocm\">PyTorch with ROCm</h2>\n<p>PyTorch has mature ROCm support. AMD publishes validated Docker images and pip wheels for each ROCm release. ROCm support is upstreamed into the official PyTorch repository, and development is aligned with stable PyTorch releases.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-pytorchrocm\" id=\"user-content-fnref-pytorchrocm\">16</a></sup></p>\n<h3 id=\"installation-options\">Installation Options</h3>\n<p><strong>Option 1: Docker (Recommended)</strong></p>\n<p>Docker provides the most reliable setup, avoiding dependency conflicts:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Pull the official ROCm PyTorch image</span></span>\n<span class=\"line\"><span>docker</span><span> pull</span><span> rocm/pytorch:rocm7.2_ubuntu24.04_py3.12_pytorch_2.9.1</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run with GPU access</span></span>\n<span class=\"line\"><span>docker</span><span> run</span><span> -it</span><span> --device=/dev/kfd</span><span> --device=/dev/dri</span><span> \\</span></span>\n<span class=\"line\"><span>    --group-add</span><span> video</span><span> --shm-size=16g</span><span> \\</span></span>\n<span class=\"line\"><span>    rocm/pytorch:rocm7.2_ubuntu24.04_py3.12_pytorch_2.9.1</span></span></code></pre>\n<p><strong>Option 2: pip Installation</strong></p>\n<p>For native installation, use the ROCm-specific wheel:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Create a virtual environment</span></span>\n<span class=\"line\"><span>python3</span><span> -m</span><span> venv</span><span> rocm-env</span></span>\n<span class=\"line\"><span>source</span><span> rocm-env/bin/activate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install PyTorch for ROCm</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> torch</span><span> torchvision</span><span> torchaudio</span><span> --index-url</span><span> https://download.pytorch.org/whl/rocm7.2</span></span></code></pre>\n<p>As of PyTorch 2.9, wheel variant support simplifies installation; the correct backend is selected automatically based on detected hardware. ROCm 7.2 supports PyTorch 2.9.1 on both Linux and Windows.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-pytorch29\" id=\"user-content-fnref-pytorch29\">17</a></sup></p>\n<h3 id=\"verifying-pytorch-rocm\">Verifying PyTorch ROCm</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Check if ROCm is available (uses CUDA API naming)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"ROCm available: </span><span>{</span><span>torch.cuda.</span><span>is_available</span><span>()</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Device count: </span><span>{</span><span>torch.cuda.</span><span>device_count</span><span>()</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Device name: </span><span>{</span><span>torch.cuda.</span><span>get_device_name</span><span>(</span><span>0</span><span>)</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run a simple operation</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>1000</span><span>, </span><span>1000</span><span>, device</span><span>=</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> torch</span><span>.</span><span>matmul</span><span>(x, x.T)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Result shape: </span><span>{</span><span>y.shape</span><span>}</span><span>\"</span><span>)</span></span></code></pre>\n<p>Note that PyTorch uses <code>cuda</code> in its API even when running on AMD GPUs. This maintains compatibility with existing code.</p>\n<h3 id=\"common-gotchas\">Common Gotchas</h3>\n<p><strong>Environment Variables</strong>: Some operations require specific settings:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Enable HIP extension for certain operations</span></span>\n<span class=\"line\"><span>export</span><span> PYTORCH_ROCM_ARCH</span><span>=</span><span>\"gfx942\"</span><span>  # Set your GPU architecture</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># For debugging memory issues</span></span>\n<span class=\"line\"><span>export</span><span> PYTORCH_HIP_ALLOC_CONF</span><span>=</span><span>\"garbage_collection_threshold:0.9\"</span></span></code></pre>\n<p><strong>FlashAttention</strong>: ROCm includes FlashAttention v2 and v3 implementations. FlashAttention v3 is integrated for AMD GPUs in recent PyTorch ROCm builds.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-11\" id=\"user-content-fnref-11\">18</a></sup></p>\n<p><strong>Mixed Precision</strong>: AMP (Automatic Mixed Precision) works on ROCm:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>amp </span><span>import</span><span> autocast</span><span>,</span><span> GradScaler</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>scaler </span><span>=</span><span> GradScaler</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>with</span><span> autocast</span><span>():</span></span>\n<span class=\"line\"><span>    output </span><span>=</span><span> model</span><span>(</span><span>input</span><span>)</span></span>\n<span class=\"line\"><span>    loss </span><span>=</span><span> criterion</span><span>(output, target)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>scaler</span><span>.</span><span>scale</span><span>(loss).</span><span>backward</span><span>()</span></span>\n<span class=\"line\"><span>scaler</span><span>.</span><span>step</span><span>(optimizer)</span></span>\n<span class=\"line\"><span>scaler</span><span>.</span><span>update</span><span>()</span></span></code></pre>\n<hr />\n<h2 id=\"llm-inference-frameworks-on-amd\">LLM Inference Frameworks on AMD</h2>\n<h3 id=\"vllm-on-rocm\">vLLM on ROCm</h3>\n<p>vLLM is the leading high-throughput LLM inference engine. As of January 2026, ROCm is a first-class platform in the vLLM ecosystem. In mid-November 2025, only 37% of vLLM test groups passed on AMD CI. As of mid-January 2026, 93% of vLLM AMD test groups succeed with daily regression maintenance.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-vllmfirstclass\" id=\"user-content-fnref-vllmfirstclass\">19</a></sup></p>\n<p><strong>Installation via Docker:</strong></p>\n<p>As of January 6, 2026, users no longer need to build from source. Pre-built official ROCm-enabled vLLM Docker images are available on Docker Hub.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-vllmdocker\" id=\"user-content-fnref-vllmdocker\">20</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Official ROCm vLLM image</span></span>\n<span class=\"line\"><span>docker</span><span> pull</span><span> rocm/vllm-dev:rocm7.2_mi350_ubuntu24.04_py3.12_vllm</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>docker</span><span> run</span><span> -it</span><span> --device=/dev/kfd</span><span> --device=/dev/dri</span><span> \\</span></span>\n<span class=\"line\"><span>    --group-add</span><span> video</span><span> --shm-size=32g</span><span> \\</span></span>\n<span class=\"line\"><span>    -p</span><span> 8000:8000</span><span> \\</span></span>\n<span class=\"line\"><span>    rocm/vllm-dev:rocm7.2_mi350_ubuntu24.04_py3.12_vllm</span></span></code></pre>\n<p><strong>Running a Model:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span><span>,</span><span> SamplingParams</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model</span></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate</span></span>\n<span class=\"line\"><span>prompts </span><span>=</span><span> [</span><span>\"Explain quantum computing in simple terms:\"</span><span>]</span></span>\n<span class=\"line\"><span>sampling </span><span>=</span><span> SamplingParams</span><span>(temperature</span><span>=</span><span>0.7</span><span>, max_tokens</span><span>=</span><span>256</span><span>)</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> llm</span><span>.</span><span>generate</span><span>(prompts, sampling)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> output </span><span>in</span><span> outputs</span><span>:</span></span>\n<span class=\"line\"><span>    print</span><span>(output.outputs[</span><span>0</span><span>].text)</span></span></code></pre>\n<p><strong>OpenAI-Compatible Server:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>python</span><span> -m</span><span> vllm.entrypoints.openai.api_server</span><span> \\</span></span>\n<span class=\"line\"><span>    --model</span><span> meta-llama/Llama-3.1-70B-Instruct</span><span> \\</span></span>\n<span class=\"line\"><span>    --tensor-parallel-size</span><span> 8</span><span> \\</span></span>\n<span class=\"line\"><span>    --port</span><span> 8000</span></span></code></pre>\n<p><strong>ROCm-Specific Optimizations:</strong></p>\n<p>vLLM V1 on ROCm includes AITER (AI Tensor Engine for ROCm) kernels optimized for MI300X, MI325X, MI350X, and MI355X GPUs. FP8 and FP4 quantization reduces memory usage by 2-4x with minimal accuracy loss. Performance on Llama 3 MXFP4 has improved through AITER optimizations and kernel fusion.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-vllmoptimize\" id=\"user-content-fnref-vllmoptimize\">21</a></sup></p>\n<p><strong>Known Issue:</strong> There is a regression with AITER for MoE models such as Mixtral and DeepSeek-R1. For these models, use the previous release <code>rocm/vllm:rocm7.0.0_vllm_0.11.1_20251103</code> for better performance.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-vllmmoe\" id=\"user-content-fnref-vllmmoe\">22</a></sup></p>\n<h3 id=\"llamacpp-with-rocm\">llama.cpp with ROCm</h3>\n<p>llama.cpp provides efficient CPU and GPU inference for GGUF-quantized models. ROCm support is available through the HIP backend.</p>\n<p><strong>Building from Source:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Clone the repository</span></span>\n<span class=\"line\"><span>git</span><span> clone</span><span> https://github.com/ggerganov/llama.cpp</span></span>\n<span class=\"line\"><span>cd</span><span> llama.cpp</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Build with HIP support</span></span>\n<span class=\"line\"><span>cmake</span><span> -B</span><span> build</span><span> -DGGML_HIP=ON</span><span> -DAMDGPU_TARGETS=</span><span>\"gfx942\"</span><span> -DCMAKE_BUILD_TYPE=Release</span></span>\n<span class=\"line\"><span>cmake</span><span> --build</span><span> build</span><span> --config</span><span> Release</span><span> -j</span><span> $(</span><span>nproc</span><span>)</span></span></code></pre>\n<p>Replace <code>gfx942</code> with your GPU architecture:</p>\n<ul>\n<li>MI350X/MI355X: <code>gfx950</code></li>\n<li>MI300X/MI325X: <code>gfx942</code></li>\n<li>MI250X: <code>gfx90a</code></li>\n<li>RX 9070 XT/RX 9070: <code>gfx1201</code></li>\n<li>RX 7900 XTX: <code>gfx1100</code></li>\n<li>RX 7900 XT: <code>gfx1101</code></li>\n</ul>\n<p><strong>Using Pre-built Binaries:</strong></p>\n<p>AMD maintains ROCm-optimized llama.cpp releases:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Download from AMD's llama.cpp fork</span></span>\n<span class=\"line\"><span>wget</span><span> https://github.com/ROCm/llama.cpp/releases/download/b6652.amd0/llama-b6652-bin-ubuntu-hip-gfx942.tar.gz</span></span>\n<span class=\"line\"><span>tar</span><span> -xzf</span><span> llama-b6652-bin-ubuntu-hip-gfx942.tar.gz</span></span></code></pre>\n<p><strong>Running Inference:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>./llama-cli</span><span> -m</span><span> models/llama-3-8b-instruct-q4_k_m.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    -p</span><span> \"What is machine learning?\"</span><span> \\</span></span>\n<span class=\"line\"><span>    -n</span><span> 256</span><span> \\</span></span>\n<span class=\"line\"><span>    -ngl</span><span> 99</span><span>  # Offload all layers to GPU</span></span></code></pre>\n<p><strong>Performance Notes:</strong></p>\n<p>AMD benchmarks show the MI300X achieving up to 213% higher inference throughput versus the H100 on Llama-3.1-70B-Q4_K_M with flash attention enabled at 4096 prompt size.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-14\" id=\"user-content-fnref-14\">23</a></sup></p>\n<h3 id=\"hugging-face-text-generation-inference-tgi\">Hugging Face Text Generation Inference (TGI)</h3>\n<p>TGI supports AMD Instinct MI210, MI250, and MI300 series GPUs.</p>\n<p><strong>Docker Deployment:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>docker</span><span> run</span><span> --device</span><span> /dev/kfd</span><span> --device</span><span> /dev/dri</span><span> \\</span></span>\n<span class=\"line\"><span>    --shm-size</span><span> 1g</span><span> -p</span><span> 8080:80</span><span> \\</span></span>\n<span class=\"line\"><span>    -v</span><span> $PWD</span><span>/data:/data</span><span> \\</span></span>\n<span class=\"line\"><span>    ghcr.io/huggingface/text-generation-inference:3.3.5-rocm</span><span> \\</span></span>\n<span class=\"line\"><span>    --model-id</span><span> meta-llama/Llama-3.1-8B-Instruct</span></span></code></pre>\n<p><strong>Configuration for Large Models:</strong></p>\n<p>Use tensor parallelism for models exceeding single-GPU memory:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>docker</span><span> run</span><span> --device</span><span> /dev/kfd</span><span> --device</span><span> /dev/dri</span><span> \\</span></span>\n<span class=\"line\"><span>    --shm-size</span><span> 16g</span><span> -p</span><span> 8080:80</span><span> \\</span></span>\n<span class=\"line\"><span>    -v</span><span> $PWD</span><span>/data:/data</span><span> \\</span></span>\n<span class=\"line\"><span>    ghcr.io/huggingface/text-generation-inference:3.3.5-rocm</span><span> \\</span></span>\n<span class=\"line\"><span>    --model-id</span><span> meta-llama/Llama-3.1-70B-Instruct</span><span> \\</span></span>\n<span class=\"line\"><span>    --num-shard</span><span> 8</span></span></code></pre>\n<p>TGI on ROCm includes a custom Paged Attention kernel enabled by default. For configurations outside the supported parameters (bf16/fp16, block size 16, head size 128, max 16k context), it falls back to PagedAttention v2.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-15\" id=\"user-content-fnref-15\">24</a></sup></p>\n<hr />\n<h2 id=\"performance-amd-vs-nvidia-for-llm-inference\">Performance: AMD vs NVIDIA for LLM Inference</h2>\n<p>Benchmark data from independent testing provides useful guidance for hardware selection.</p>\n<h3 id=\"memory-bandwidth-advantage\">Memory Bandwidth Advantage</h3>\n<p>The MI355X’s 8 TB/s memory bandwidth exceeds the B200’s 8 TB/s and substantially outpaces the H100’s 3.35 TB/s. The MI300X at 5.3 TB/s remains competitive with the H200’s 4.8 TB/s. For memory-bound LLM inference at low batch sizes, this translates to lower latency.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-memperf\" id=\"user-content-fnref-memperf\">25</a></sup></p>\n<h3 id=\"mi355x-vs-blackwell-b200\">MI355X vs Blackwell B200</h3>\n<p>In benchmarks with Llama3.1-405B, the MI355X delivers up to 2x higher throughput compared to competitive options. At high concurrency levels (32-64), the MI355X with ATOM matches or exceeds B200 systems running SGLang. The MI355X demonstrated a 10% time-to-solution advantage in the MLPerf LoRA fine-tuning benchmark.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-mi355xperf\" id=\"user-content-fnref-mi355xperf\">26</a></sup></p>\n<h3 id=\"large-model-performance\">Large Model Performance</h3>\n<p>For models requiring multi-GPU deployment, AMD’s large VRAM pools allow single-GPU inference where competitors need tensor parallelism:</p>\n<ul>\n<li>405B parameter models fit on one MI355X with FP4 quantization<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-mi355xmem\" id=\"user-content-fnref-mi355xmem\">27</a></sup></li>\n<li>Mixtral 8x7B fits on one MI300X; H100 requires TP=2<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-18\" id=\"user-content-fnref-18\">28</a></sup></li>\n<li>DeepSeek-V3 670B: MI300X beats H100 in both absolute performance and performance per dollar<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-19\" id=\"user-content-fnref-19\">29</a></sup></li>\n</ul>\n<h3 id=\"cloud-availability\">Cloud Availability</h3>\n<p>MI355X GPUs are available at Oracle Cloud, TensorWave, and Vultr. Oracle Cloud offers bare-metal GPU instances with 8 MI355X GPUs per node (2.3 TB total GPU memory), 128 CPU cores, 3 TB DDR system memory, and up to 3.2 Tb/s RDMA bandwidth.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-cloudavail\" id=\"user-content-fnref-cloudavail\">30</a></sup></p>\n<h3 id=\"software-stack-maturity\">Software Stack Maturity</h3>\n<p>ROCm 7.2 has closed the gap with CUDA. vLLM now passes 93% of test groups on AMD CI. PyTorch runs natively on both Linux and Windows. For production deployments with standard configurations, both stacks work reliably. NVIDIA retains an edge in MLPerf training benchmarks, where Blackwell leads on Llama 3.1 403B pretraining.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-mlperf\" id=\"user-content-fnref-mlperf\">31</a></sup></p>\n<hr />\n<h2 id=\"training-on-amd-gpus\">Training on AMD GPUs</h2>\n<p>ROCm supports distributed training with PyTorch’s native parallelism primitives. While NVIDIA Blackwell currently leads MLPerf training benchmarks on Llama 3.1 403B pretraining, AMD’s MI325X matches H200 performance on LLM fine-tuning benchmarks, suggesting AMD is approximately one generation behind NVIDIA on training workloads.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-traingap\" id=\"user-content-fnref-traingap\">32</a></sup></p>\n<h3 id=\"single-gpu-training\">Single-GPU Training</h3>\n<p>Standard PyTorch training loops work unchanged:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>nn </span><span>as</span><span> nn</span></span>\n<span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>optim </span><span>as</span><span> optim</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> MyModel</span><span>().</span><span>cuda</span><span>()</span></span>\n<span class=\"line\"><span>optimizer </span><span>=</span><span> optim</span><span>.</span><span>AdamW</span><span>(model.</span><span>parameters</span><span>(), lr</span><span>=</span><span>1e-4</span><span>)</span></span>\n<span class=\"line\"><span>criterion </span><span>=</span><span> nn</span><span>.</span><span>CrossEntropyLoss</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> epoch </span><span>in</span><span> range</span><span>(num_epochs):</span></span>\n<span class=\"line\"><span>    for</span><span> batch </span><span>in</span><span> dataloader</span><span>:</span></span>\n<span class=\"line\"><span>        inputs</span><span>,</span><span> labels </span><span>=</span><span> batch</span></span>\n<span class=\"line\"><span>        inputs</span><span>,</span><span> labels </span><span>=</span><span> inputs</span><span>.</span><span>cuda</span><span>(),</span><span> labels</span><span>.</span><span>cuda</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        optimizer</span><span>.</span><span>zero_grad</span><span>()</span></span>\n<span class=\"line\"><span>        outputs </span><span>=</span><span> model</span><span>(inputs)</span></span>\n<span class=\"line\"><span>        loss </span><span>=</span><span> criterion</span><span>(outputs, labels)</span></span>\n<span class=\"line\"><span>        loss</span><span>.</span><span>backward</span><span>()</span></span>\n<span class=\"line\"><span>        optimizer</span><span>.</span><span>step</span><span>()</span></span></code></pre>\n<h3 id=\"distributed-data-parallel-ddp\">Distributed Data Parallel (DDP)</h3>\n<p>RCCL provides collective operations for multi-GPU training:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>distributed </span><span>as</span><span> dist</span></span>\n<span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>parallel </span><span>import</span><span> DistributedDataParallel </span><span>as</span><span> DDP</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Initialize process group</span></span>\n<span class=\"line\"><span>dist</span><span>.</span><span>init_process_group</span><span>(backend</span><span>=</span><span>'nccl'</span><span>)</span><span>  # Uses RCCL on AMD</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>local_rank </span><span>=</span><span> int</span><span>(os.environ[</span><span>'LOCAL_RANK'</span><span>])</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>cuda</span><span>.</span><span>set_device</span><span>(local_rank)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> MyModel</span><span>().</span><span>cuda</span><span>()</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> DDP</span><span>(model, device_ids</span><span>=</span><span>[local_rank])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Training loop remains the same</span></span></code></pre>\n<p>Launch with:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>torchrun</span><span> --nproc_per_node=8</span><span> train.py</span></span></code></pre>\n<h3 id=\"fully-sharded-data-parallel-fsdp\">Fully Sharded Data Parallel (FSDP)</h3>\n<p>FSDP shards model parameters, gradients, and optimizer states across GPUs, enabling training of models larger than single-GPU memory:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>distributed</span><span>.</span><span>fsdp </span><span>import</span><span> FullyShardedDataParallel </span><span>as</span><span> FSDP</span></span>\n<span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>distributed</span><span>.</span><span>fsdp </span><span>import</span><span> ShardingStrategy</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> MyLargeModel</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Wrap with FSDP</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> FSDP</span><span>(</span></span>\n<span class=\"line\"><span>    model,</span></span>\n<span class=\"line\"><span>    sharding_strategy</span><span>=</span><span>ShardingStrategy.FULL_SHARD,</span></span>\n<span class=\"line\"><span>    device_id</span><span>=</span><span>local_rank</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>PyTorch’s FSDP works well on ROCm. AMD benchmarks show 8x MI300X achieving up to 1.29x better performance compared to 8x H100 when training DeepSeek-V2-Lite.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-22\" id=\"user-content-fnref-22\">33</a></sup></p>\n<p><strong>FSDP Configuration Tips:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Recommended settings for ROCm</span></span>\n<span class=\"line\"><span>fsdp_config </span><span>=</span><span> {</span></span>\n<span class=\"line\"><span>    \"sharding_strategy\"</span><span>:</span><span> ShardingStrategy</span><span>.</span><span>FULL_SHARD</span><span>,</span></span>\n<span class=\"line\"><span>    \"backward_prefetch\"</span><span>:</span><span> BackwardPrefetch</span><span>.</span><span>BACKWARD_PRE</span><span>,</span></span>\n<span class=\"line\"><span>    \"forward_prefetch\"</span><span>:</span><span> True</span><span>,</span></span>\n<span class=\"line\"><span>    \"limit_all_gathers\"</span><span>:</span><span> True</span><span>,</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"training-infrastructure\">Training Infrastructure</h3>\n<p>AMD provides optimized Docker containers for distributed training:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># ROCm PyTorch Training container</span></span>\n<span class=\"line\"><span>docker</span><span> pull</span><span> rocm/pytorch-training:rocm7.2_ubuntu24.04_py3.12</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Includes torchtitan and Hugging Face Accelerate</span></span></code></pre>\n<p>The container includes libraries for FSDP training with one-shot and two-shot AllReduce strategies for optimized communication.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-23\" id=\"user-content-fnref-23\">34</a></sup></p>\n<hr />\n<h2 id=\"debugging-and-profiling-with-rocprof\">Debugging and Profiling with rocprof</h2>\n<p>ROCm provides comprehensive profiling tools for performance optimization.</p>\n<h3 id=\"rocprofv3\">rocprofv3</h3>\n<p>The latest profiling tool (replacing rocprof v1/v2) offers improved usability:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Profile an application</span></span>\n<span class=\"line\"><span>rocprofv3</span><span> --hip-trace</span><span> --hsa-trace</span><span> -o</span><span> profile_output</span><span> python</span><span> train.py</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Collect hardware counters</span></span>\n<span class=\"line\"><span>rocprofv3</span><span> --stats</span><span> --output-format</span><span> csv</span><span> python</span><span> inference.py</span></span></code></pre>\n<h3 id=\"key-profiling-commands\">Key Profiling Commands</h3>\n<p><strong>Basic Kernel Profiling:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># List available counters</span></span>\n<span class=\"line\"><span>rocprofv3</span><span> --list-counters</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Profile with specific counters</span></span>\n<span class=\"line\"><span>rocprofv3</span><span> --pmc</span><span> GPU_BUSY,GRBM_COUNT</span><span> python</span><span> model.py</span></span></code></pre>\n<p><strong>Timeline Tracing:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Generate trace for visualization</span></span>\n<span class=\"line\"><span>rocprofv3</span><span> --hip-trace</span><span> --output-format</span><span> json</span><span> python</span><span> model.py</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># View in Perfetto</span></span>\n<span class=\"line\"><span># Open https://ui.perfetto.dev and load results.json</span></span></code></pre>\n<h3 id=\"rocprof-compute-formerly-omniperf\">rocprof-compute (formerly Omniperf)</h3>\n<p>For detailed kernel analysis:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Collect comprehensive metrics</span></span>\n<span class=\"line\"><span>rocprof-compute</span><span> profile</span><span> -n</span><span> my_workload</span><span> --</span><span> python</span><span> train.py</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Analyze results</span></span>\n<span class=\"line\"><span>rocprof-compute</span><span> analyze</span><span> my_workload/</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate roofline plot</span></span>\n<span class=\"line\"><span>rocprof-compute</span><span> analyze</span><span> my_workload/</span><span> --roof</span></span></code></pre>\n<p>rocprof-compute provides:</p>\n<ul>\n<li>Roofline analysis showing compute vs memory boundedness</li>\n<li>Memory throughput analysis</li>\n<li>Compute utilization breakdown</li>\n<li>Baseline comparisons between runs<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-24\" id=\"user-content-fnref-24\">35</a></sup></li>\n</ul>\n<h3 id=\"rocgdb-for-debugging\">ROCgdb for Debugging</h3>\n<p>The ROCm debugger extends GDB for heterogeneous debugging:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Debug a HIP application</span></span>\n<span class=\"line\"><span>rocgdb</span><span> ./my_hip_program</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Set breakpoint in kernel</span></span>\n<span class=\"line\"><span>(</span><span>gdb</span><span>) </span><span>break</span><span> myKernel</span></span>\n<span class=\"line\"><span>(</span><span>gdb</span><span>) </span><span>run</span></span></code></pre>\n<p>Note: On RHEL with SELinux enabled, debugging may hang. Either disable SELinux or configure appropriate policies.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-25\" id=\"user-content-fnref-25\">36</a></sup></p>\n<hr />\n<h2 id=\"current-limitations-and-workarounds\">Current Limitations and Workarounds</h2>\n<p>ROCm has made substantial progress but some limitations remain.</p>\n<h3 id=\"platform-support\">Platform Support</h3>\n<p><strong>Windows</strong>: ROCm 7.2 includes native Windows support for consumer GPUs (RX 7000/9000 series). PyTorch runs without WSL2. The Adrenalin 26.1.1 driver includes ROCm components and ComfyUI integration.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-26\" id=\"user-content-fnref-26\">37</a></sup></p>\n<p><strong>macOS</strong>: Not supported. Apple Silicon users should consider MLX instead.</p>\n<p><strong>Mobile GPUs</strong>: Not officially supported by ROCm. Ryzen AI APUs (AI Max 300, AI 400 series) are now supported on both Linux and Windows.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-ryzenai\" id=\"user-content-fnref-ryzenai\">38</a></sup></p>\n<h3 id=\"memory-issues-on-consumer-gpus\">Memory Issues on Consumer GPUs</h3>\n<p>Running large models on RDNA 3/4 cards with 16-24 GB VRAM can cause instability. The RX 9070 XT’s 16 GB is adequate for models up to 24B parameters at Q4 quantization:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Workarounds for memory pressure</span></span>\n<span class=\"line\"><span># For Stable Diffusion / FLUX</span></span>\n<span class=\"line\"><span>python</span><span> main.py</span><span> --lowvram</span><span> --disable-pinned-memory</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># For PyTorch memory management</span></span>\n<span class=\"line\"><span>export</span><span> PYTORCH_HIP_ALLOC_CONF</span><span>=</span><span>\"garbage_collection_threshold:0.8,max_split_size_mb:512\"</span></span></code></pre>\n<h3 id=\"jax-support\">JAX Support</h3>\n<p>JAX on ROCm is supported for inference only. Training workloads may encounter intermittent errors or segmentation faults.<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-27\" id=\"user-content-fnref-27\">39</a></sup></p>\n<h3 id=\"quantization-support\">Quantization Support</h3>\n<p>Not all quantization formats work on ROCm:</p>\n<ul>\n<li>GPTQ and AWQ: Supported in vLLM</li>\n<li>GGUF: Supported in llama.cpp</li>\n<li>FP8: Supported on MI300 series</li>\n<li>MXFP4: Only supported on MI350 series<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-28\" id=\"user-content-fnref-28\">40</a></sup></li>\n<li>AWQ in TGI: Not currently supported<sup><a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fn-29\" id=\"user-content-fnref-29\">41</a></sup></li>\n</ul>\n<h3 id=\"multi-gpu-configuration\">Multi-GPU Configuration</h3>\n<p>Installing multiple ROCm versions can cause amd-smi issues. Stick to one ROCm version per system or use containers for version isolation.</p>\n<h3 id=\"framework-specific-issues\">Framework-Specific Issues</h3>\n<p><strong>Transformers library</strong>: Most models work; some custom CUDA kernels may need ROCm equivalents.</p>\n<p><strong>Flash Attention</strong>: Supported via CK (Composable Kernel) implementation. Triton-based FA has lower latency on MI250/MI300 but requires warmup for each new sequence length.</p>\n<hr />\n<h2 id=\"practical-example-deploying-llama-3-on-mi300x\">Practical Example: Deploying Llama 3 on MI300X</h2>\n<p>Here is a complete example for setting up a production LLM inference endpoint:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Pull the vLLM Docker image</span></span>\n<span class=\"line\"><span>docker</span><span> pull</span><span> rocm/vllm-dev:rocm7.2_mi350_ubuntu24.04_py3.12_vllm</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Create a deployment script</span></span>\n<span class=\"line\"><span>cat</span><span> &gt;</span><span> deploy.sh</span><span> &lt;&lt;</span><span> 'EOF'</span></span>\n<span class=\"line\"><span>#!/bin/bash</span></span>\n<span class=\"line\"><span>docker run -d \\</span></span>\n<span class=\"line\"><span>    --name llama3-server \\</span></span>\n<span class=\"line\"><span>    --device=/dev/kfd \\</span></span>\n<span class=\"line\"><span>    --device=/dev/dri \\</span></span>\n<span class=\"line\"><span>    --group-add video \\</span></span>\n<span class=\"line\"><span>    --shm-size=32g \\</span></span>\n<span class=\"line\"><span>    -p 8000:8000 \\</span></span>\n<span class=\"line\"><span>    -v ~/.cache/huggingface:/root/.cache/huggingface \\</span></span>\n<span class=\"line\"><span>    rocm/vllm-dev:rocm7.2_mi350_ubuntu24.04_py3.12_vllm \\</span></span>\n<span class=\"line\"><span>    python -m vllm.entrypoints.openai.api_server \\</span></span>\n<span class=\"line\"><span>        --model meta-llama/Llama-3.1-70B-Instruct \\</span></span>\n<span class=\"line\"><span>        --tensor-parallel-size 4 \\</span></span>\n<span class=\"line\"><span>        --max-model-len 8192 \\</span></span>\n<span class=\"line\"><span>        --gpu-memory-utilization 0.9</span></span>\n<span class=\"line\"><span>EOF</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>chmod</span><span> +x</span><span> deploy.sh</span></span>\n<span class=\"line\"><span>./deploy.sh</span></span></code></pre>\n<p><strong>Testing the Endpoint:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> openai </span><span>import</span><span> OpenAI</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>client </span><span>=</span><span> OpenAI</span><span>(</span></span>\n<span class=\"line\"><span>    base_url</span><span>=</span><span>\"http://localhost:8000/v1\"</span><span>,</span></span>\n<span class=\"line\"><span>    api_key</span><span>=</span><span>\"not-needed\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>response </span><span>=</span><span> client</span><span>.</span><span>chat</span><span>.</span><span>completions</span><span>.</span><span>create</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.1-70B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    messages</span><span>=</span><span>[</span></span>\n<span class=\"line\"><span>        {</span><span>\"role\"</span><span>: </span><span>\"user\"</span><span>, </span><span>\"content\"</span><span>: </span><span>\"Write a haiku about GPU computing\"</span><span>}</span></span>\n<span class=\"line\"><span>    ],</span></span>\n<span class=\"line\"><span>    max_tokens</span><span>=</span><span>100</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>print</span><span>(response.choices[</span><span>0</span><span>].message.content)</span></span></code></pre>\n<p><strong>Monitoring:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Watch GPU utilization</span></span>\n<span class=\"line\"><span>watch</span><span> -n</span><span> 1</span><span> rocm-smi</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Check memory usage</span></span>\n<span class=\"line\"><span>amd-smi</span><span> monitor</span><span> -p</span></span></code></pre>\n<hr />\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>ROCm has evolved from a CUDA alternative requiring substantial effort into a production-ready platform for ML workloads. The MI355X delivers competitive performance against NVIDIA’s Blackwell B200 for LLM inference, particularly at high concurrency levels. The MI300X and MI325X remain strong options for memory-bound workloads. Consumer RDNA 4 cards (RX 9070 series) provide a viable path for local development with native Windows support.</p>\n<p>Key points to remember:</p>\n<ul>\n<li>ROCm 7.2 supports both Linux and Windows</li>\n<li>PyTorch 2.9.1 works on both platforms without code changes</li>\n<li>vLLM passes 93% of test groups on AMD CI (January 2026)</li>\n<li>Profile with rocprofv3 and rocprof-compute</li>\n<li>Consumer GPUs have improved support but 16 GB VRAM limits model sizes</li>\n<li>OpenAI will deploy 1 GW of MI450 GPUs starting H2 2026</li>\n</ul>\n<p>For organizations evaluating GPU infrastructure, AMD provides a cost-effective alternative with competitive performance. The MI400 series (H2 2026) with 432 GB HBM4 will further close the gap with NVIDIA.</p>\n<hr />\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-openai\">\n<p><a href=\"https://openai.com/index/openai-amd-strategic-partnership/\" rel=\"noopener noreferrer\">AMD and OpenAI Announce Strategic Partnership to Deploy 6 Gigawatts of AMD GPUs</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-openai\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mi355x\">\n<p><a href=\"https://blogs.oracle.com/cloud-infrastructure/amd-instinct-mi355x-on-oci-performance-technical-details\" rel=\"noopener noreferrer\">AMD Instinct MI355X on OCI Performance &amp; Technical Details</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-mi355x\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mi350\">\n<p><a href=\"https://www.amd.com/en/blogs/2025/amd-instinct-mi350-series-game-changer.html\" rel=\"noopener noreferrer\">AMD Instinct MI350 Series GPUs: A Game Changer</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-mi350\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mi325x\">\n<p><a href=\"https://spectrum.ieee.org/nvidia-blackwell-mlperf-training-5\" rel=\"noopener noreferrer\">Is Nvidia’s Blackwell the Unstoppable Force in AI Training, or Can AMD Close the Gap?</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-mi325x\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mi300x\">\n<p><a href=\"https://neysa.ai/blog/amd-mi300x/\" rel=\"noopener noreferrer\">AMD MI300X: Specs &amp; Performance for AI/ML Workloads</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-mi300x\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mi400\">\n<p><a href=\"https://videocardz.com/newz/amd-launches-instinct-mi350-series-confirms-mi400-in-2026-with-432gb-hbm4-memory\" rel=\"noopener noreferrer\">AMD launches Instinct MI350 series, confirms MI400 in 2026 with 432GB HBM4 memory</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-mi400\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-rx9070\">\n<p><a href=\"https://www.tomshardware.com/pc-components/gpus/amd-radeon-rx-9070-xt-review\" rel=\"noopener noreferrer\">AMD Radeon RX 9070 XT and RX 9070 review - Tom’s Hardware</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-rx9070\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p><a href=\"https://rocm.docs.amd.com/en/latest/compatibility/ml-compatibility/llama-cpp-compatibility.html\" rel=\"noopener noreferrer\">llama.cpp ROCm Compatibility</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-radeonai\">\n<p><a href=\"https://www.amd.com/en/products/graphics/radeon-ai.html\" rel=\"noopener noreferrer\">AI Acceleration with AMD Radeon Graphics Cards</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-radeonai\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p><a href=\"https://rocm.docs.amd.com/projects/HIP/en/latest/what_is_hip.html\" rel=\"noopener noreferrer\">What is HIP - ROCm Documentation</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p><a href=\"https://github.com/ROCm/HIPIFY\" rel=\"noopener noreferrer\">HIPIFY GitHub Repository</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p><a href=\"https://rocm.blogs.amd.com/software-tools-optimization/hipify/README.html\" rel=\"noopener noreferrer\">Application Portability with HIP - AMD Lab Notes</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-rocm72\">\n<p><a href=\"https://rocm.blogs.amd.com/software-tools-optimization/rocm7.2/README.html\" rel=\"noopener noreferrer\">ROCm 7.2: Smarter, Faster, and More Scalable for Modern AI Workloads</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-rocm72\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-rocmwindows\">\n<p><a href=\"https://videocardz.com/newz/amd-enables-windows-pytorch-support-for-radeon-rx-7000-9000-with-rocm-6-4-4-update\" rel=\"noopener noreferrer\">AMD enables Windows PyTorch support for Radeon RX 7000/9000 with ROCm</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-rocmwindows\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-comfyui\">\n<p><a href=\"https://www.notebookcheck.net/AMD-ROCm-7-2-brings-AI-training-and-inferencing-to-all-with-Ryzen-AI-400-support-and-ComfyUI-Adrenalin-integrations.1196997.0.html\" rel=\"noopener noreferrer\">AMD ROCm 7.2 brings AI training and inferencing to all</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-comfyui\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-pytorchrocm\">\n<p><a href=\"https://rocm.docs.amd.com/en/latest/compatibility/ml-compatibility/pytorch-compatibility.html\" rel=\"noopener noreferrer\">PyTorch compatibility - ROCm Documentation</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-pytorchrocm\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-pytorch29\">\n<p><a href=\"https://www.amd.com/en/developer/resources/technical-articles/2025/pytorch-2-9-wheel-variant-support-expands-to-rocm.html\" rel=\"noopener noreferrer\">PyTorch 2.9 Wheel Variant support expands to ROCm</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-pytorch29\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p><a href=\"https://rocm.docs.amd.com/en/latest/compatibility/ml-compatibility/pytorch-compatibility.html\" rel=\"noopener noreferrer\">PyTorch ROCm Compatibility</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-vllmfirstclass\">\n<p><a href=\"https://rocm.blogs.amd.com/software-tools-optimization/vllm-omni/README.html\" rel=\"noopener noreferrer\">ROCm Becomes a First-Class Platform in the vLLM Ecosystem</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-vllmfirstclass\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-vllmdocker\">\n<p><a href=\"https://hub.docker.com/r/rocm/vllm\" rel=\"noopener noreferrer\">rocm/vllm - Docker Image</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-vllmdocker\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-vllmoptimize\">\n<p><a href=\"https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference-optimization/vllm-optimization.html\" rel=\"noopener noreferrer\">vLLM V1 performance optimization - ROCm Documentation</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-vllmoptimize\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-vllmmoe\">\n<p><a href=\"https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference/benchmark-docker/vllm.html\" rel=\"noopener noreferrer\">vLLM inference performance testing - ROCm Documentation</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-vllmmoe\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p><a href=\"https://rocm.blogs.amd.com/ecosystems-and-partners/llama-cpp-oct2025/README.html\" rel=\"noopener noreferrer\">Accelerating llama.cpp on MI300X</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p><a href=\"https://huggingface.co/docs/text-generation-inference/installation_amd\" rel=\"noopener noreferrer\">Using TGI with AMD GPUs</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-memperf\">\n<p><a href=\"https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html\" rel=\"noopener noreferrer\">AMD Instinct MI355X GPUs</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-memperf\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mi355xperf\">\n<p><a href=\"https://www.amd.com/en/developer/resources/technical-articles/2026/distributed-inference-performance-on-instinct-mi355x-gpu.html\" rel=\"noopener noreferrer\">Single Node and Distributed Inference Performance on AMD Instinct MI355X GPU</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-mi355xperf\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mi355xmem\">\n<p><a href=\"https://getdeploying.com/gpus/amd-mi355x\" rel=\"noopener noreferrer\">AMD MI355X - Price, Specs &amp; Cloud Providers</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-mi355xmem\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-18\">\n<p><a href=\"https://www.runpod.io/blog/mi300x-vs-h100-mixtral\" rel=\"noopener noreferrer\">MI300X vs H100 Mixtral Benchmark - RunPod</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-18\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-19\">\n<p><a href=\"https://newsletter.semianalysis.com/p/amd-vs-nvidia-inference-benchmark-who-wins-performance-cost-per-million-tokens\" rel=\"noopener noreferrer\">AMD vs NVIDIA Inference Benchmark - SemiAnalysis</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-19\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-cloudavail\">\n<p><a href=\"https://blogs.oracle.com/cloud-infrastructure/amd-instinct-mi355x-on-oci-performance-technical-details\" rel=\"noopener noreferrer\">AMD Instinct MI355X on OCI Performance &amp; Technical Details</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-cloudavail\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-mlperf\">\n<p><a href=\"https://spectrum.ieee.org/nvidia-blackwell-mlperf-training-5\" rel=\"noopener noreferrer\">Is Nvidia’s Blackwell the Unstoppable Force in AI Training?</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-mlperf\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-traingap\">\n<p><a href=\"https://www.bestgpusforai.com/blog/best-amd-gpus-for-ai\" rel=\"noopener noreferrer\">Best AMD GPUs for AI Training &amp; Deep Learning in 2026</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-traingap\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-22\">\n<p><a href=\"https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/training/scale-model-training.html\" rel=\"noopener noreferrer\">Scaling Model Training on ROCm</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-22\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-23\">\n<p><a href=\"https://rocm.blogs.amd.com/artificial-intelligence/fsdp-training-pytorch/README.html\" rel=\"noopener noreferrer\">PyTorch FSDP on AMD GPUs</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-23\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-24\">\n<p><a href=\"https://rocm.blogs.amd.com/software-tools-optimization/rocprofiler-sdk/README.html\" rel=\"noopener noreferrer\">Introduction to Rocprofiler-compute</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-24\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-25\">\n<p><a href=\"https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/docs/limitations/limitationsrad.html\" rel=\"noopener noreferrer\">ROCm Limitations - Radeon Documentation</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-25\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-26\">\n<p><a href=\"https://www.neowin.net/news/amd-rocm-open-source-nvidia-cuda-rival-gets-massive-windows--linux-improvements/\" rel=\"noopener noreferrer\">AMD ROCm, open source Nvidia CUDA rival, gets massive Windows &amp; Linux improvements</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-26\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-ryzenai\">\n<p><a href=\"https://epium.com/news/amd-expands-rocm-support-ryzen-ai-max-radeon-rx-9000-series/\" rel=\"noopener noreferrer\">AMD Expands ROCm Support to Ryzen AI Max and Radeon RX 9000 Series</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-ryzenai\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-27\">\n<p><a href=\"https://rocm.docs.amd.com/en/latest/about/release-notes.html\" rel=\"noopener noreferrer\">ROCm 7.2.0 Release Notes</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-27\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-28\">\n<p><a href=\"https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference-optimization/vllm-optimization.html\" rel=\"noopener noreferrer\">vLLM ROCm Quantization Support</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-28\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-29\">\n<p><a href=\"https://huggingface.co/docs/text-generation-inference/installation_amd\" rel=\"noopener noreferrer\">TGI ROCm Limitations</a> <a href=\"https://blog.ecitis.org/amd-rocm-ml/#user-content-fnref-29\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/amd-rocm-ml/",
            "title": "AMD ROCm for Machine Learning: The Alternative GPU Ecosystem",
            "summary": "A practical tour of AMD’s open GPU stack for training and inference, including PyTorch, vLLM, compatibility, and deployment.",
            "image": "https://blog.ecitis.org/open-graph/amd-rocm-ml.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Hardware & Systems",
                "ROCm",
                "AMD",
                "GPU compute"
            ]
        },
        {
            "id": "https://blog.ecitis.org/apple-metal-ml-ecosystem/",
            "content_html": "<h2 id=\"introduction\">Introduction</h2>\n<p>Apple Silicon has transformed what is possible for on-device machine learning. While MLX gets the headlines for LLM inference, Apple’s ML stack runs much deeper. Metal Performance Shaders, CoreML, the Accelerate framework, and various neural network libraries form a layered ecosystem that powers everything from real-time image processing to transformer inference on iPhones.</p>\n<p>This guide explores the full Metal ML ecosystem: how PyTorch MPS works under the hood, when to use CoreML versus MLX, how to write custom GPU kernels, and the practical realities of deploying models across Apple platforms.</p>\n<figure><figcaption><strong>Apple ML Stack Architecture</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/apple-metal-ml-ecosystem-0.webp\" width=\"512\" height=\"970\" alt=\"Apple ML Stack Architecture\" loading=\"lazy\" decoding=\"async\" /></figure>\n<hr />\n<h2 id=\"metal-performance-shaders\">Metal Performance Shaders</h2>\n<p>Metal Performance Shaders (MPS) is Apple’s framework for optimized GPU compute kernels. Originally focused on image processing and linear algebra, MPS has evolved into a foundational layer for machine learning on Apple platforms.</p>\n<h3 id=\"what-mps-provides\">What MPS Provides</h3>\n<p>MPS offers pre-built, hardware-tuned kernels for common operations. These kernels are optimized for each GPU family in Apple’s lineup, from the integrated graphics in older Intel Macs to the latest M-series chips. The framework handles the low-level details of GPU programming: memory allocation, command buffer management, and shader compilation.</p>\n<p>For machine learning, MPS provides:</p>\n<ul>\n<li><strong>Convolution operations</strong>: 2D and 3D convolutions with various padding modes</li>\n<li><strong>Pooling layers</strong>: Max, average, and L2 pooling</li>\n<li><strong>Normalization</strong>: Batch normalization, instance normalization, layer normalization</li>\n<li><strong>Activation functions</strong>: ReLU, sigmoid, tanh, and others</li>\n<li><strong>Matrix operations</strong>: Matrix multiplication, transpose, and decomposition</li>\n<li><strong>Neural network primitives</strong>: Fully connected layers, softmax, LSTM cells</li>\n</ul>\n<p>These operations form the building blocks that higher-level frameworks like PyTorch and TensorFlow use when running on Apple GPUs.</p>\n<h3 id=\"mpsgraph-the-computation-graph-framework\">MPSGraph: The Computation Graph Framework</h3>\n<p>MPSGraph extends MPS with a graph-based execution model. Rather than executing individual operations, you build a computation graph that MPSGraph compiles and optimizes as a unit.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> MetalPerformanceShadersGraph</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>let</span><span> graph </span><span>=</span><span> MPSGraph</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Define input placeholders</span></span>\n<span class=\"line\"><span>let</span><span> inputTensor </span><span>=</span><span> graph.</span><span>placeholder</span><span>(</span></span>\n<span class=\"line\"><span>    shape</span><span>:</span><span> [</span><span>1</span><span>, </span><span>224</span><span>, </span><span>224</span><span>, </span><span>3</span><span>],</span></span>\n<span class=\"line\"><span>    dataType</span><span>:</span><span> .float32,</span></span>\n<span class=\"line\"><span>    name</span><span>:</span><span> \"input\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Build operations</span></span>\n<span class=\"line\"><span>let</span><span> conv </span><span>=</span><span> graph.</span><span>convolution2D</span><span>(</span></span>\n<span class=\"line\"><span>    inputTensor,</span></span>\n<span class=\"line\"><span>    weights</span><span>:</span><span> weightsTensor,</span></span>\n<span class=\"line\"><span>    descriptor</span><span>:</span><span> convDescriptor,</span></span>\n<span class=\"line\"><span>    name</span><span>:</span><span> \"conv1\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>let</span><span> relu </span><span>=</span><span> graph.</span><span>reLU</span><span>(</span><span>with</span><span>:</span><span> conv, name</span><span>:</span><span> \"relu1\"</span><span>)</span></span></code></pre>\n<p>At WWDC 2024, Apple introduced several improvements to MPSGraph for transformer models. The new Scaled Dot-Product Attention (SDPA) operation fuses the entire attention computation into a single kernel. Combined with KV-cache support, this provides significant speedups for autoregressive generation.</p>\n<h3 id=\"performance-characteristics\">Performance Characteristics</h3>\n<p>MPS kernels achieve strong performance because they are tuned for Apple’s specific GPU architectures. The Metal GPU Profiler in Xcode shows that well-optimized MPS code typically achieves high ALU utilization with minimal memory stalls.</p>\n<p>However, MPS performance varies by operation type. Large matrix multiplications and convolutions benefit most from GPU acceleration. Smaller operations may run faster on the CPU due to kernel launch overhead.</p>\n<hr />\n<h2 id=\"pytorch-mps-backend\">PyTorch MPS Backend</h2>\n<p>Apple contributed the MPS backend to PyTorch in 2022, enabling GPU-accelerated training and inference on Apple Silicon. The backend maps PyTorch operations to MPS kernels and MPSGraph.</p>\n<h3 id=\"how-it-works\">How It Works</h3>\n<p>When you move a tensor to the MPS device, PyTorch allocates memory in Metal’s shared memory pool:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Check MPS availability</span></span>\n<span class=\"line\"><span>if</span><span> torch</span><span>.</span><span>backends</span><span>.</span><span>mps</span><span>.</span><span>is_available</span><span>():</span></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>\"mps\"</span><span>)</span></span>\n<span class=\"line\"><span>else</span><span>:</span></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>\"cpu\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Create tensor on MPS device</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>1000</span><span>, </span><span>1000</span><span>, device</span><span>=</span><span>device)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Operations execute on GPU via Metal</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> torch</span><span>.</span><span>matmul</span><span>(x, x.T)</span></span></code></pre>\n<p>The MPS backend implements PyTorch’s ATen operations using two approaches:</p>\n<ol>\n<li><strong>MPSGraph operations</strong>: Complex operations like convolutions and matrix multiplications use MPSGraph for optimal performance</li>\n<li><strong>Custom Metal shaders</strong>: Element-wise operations use hand-written Metal compute shaders</li>\n</ol>\n<h3 id=\"supported-operations\">Supported Operations</h3>\n<p>PyTorch 2.0 expanded MPS coverage to over 300 operators. The top 60 most-used operations are all supported. You can check coverage for specific operations in the <a href=\"https://github.com/pytorch/pytorch/issues/141287\">PyTorch MPS operator tracking issue</a>.</p>\n<p>Common supported operations include:</p>\n<ul>\n<li>Linear algebra: matmul, bmm, addmm, mm</li>\n<li>Convolutions: conv1d, conv2d, conv3d, conv_transpose</li>\n<li>Pooling: max_pool2d, avg_pool2d, adaptive_avg_pool2d</li>\n<li>Activations: relu, gelu, silu, sigmoid, tanh</li>\n<li>Normalization: batch_norm, layer_norm, group_norm</li>\n<li>Loss functions: cross_entropy, mse_loss, nll_loss</li>\n</ul>\n<h3 id=\"current-limitations\">Current Limitations</h3>\n<p>The MPS backend has several constraints that affect real-world usage:</p>\n<p><strong>No float64 support</strong>: MPS does not support double precision. Operations requiring float64 will fail or fall back to CPU. This affects some scientific computing workloads.</p>\n<p><strong>No distributed training</strong>: The NCCL and Gloo backends do not work with MPS. Multi-GPU training is not supported. This limits MPS to single-device workloads.</p>\n<p><strong>No Neural Engine access</strong>: PyTorch MPS only uses the GPU. The Neural Engine, which can accelerate certain operations more efficiently, remains unused.</p>\n<p><strong>Operation gaps</strong>: Some operations lack MPS implementations. The <code>PYTORCH_ENABLE_MPS_FALLBACK=1</code> environment variable enables automatic CPU fallback for unsupported operations:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>PYTORCH_ENABLE_MPS_FALLBACK</span><span>=</span><span>1</span><span> python</span><span> train.py</span></span></code></pre>\n<h3 id=\"known-issues\">Known Issues</h3>\n<p>A notable bug in PyTorch versions before 2.4 caused silent failures when writing to non-contiguous tensors via <code>addcmul_</code> and <code>addcdiv_</code> operations. This would cause model weights to stop updating during training without any error message. The fix requires updating to PyTorch 2.4 or later.</p>\n<h3 id=\"performance-expectations\">Performance Expectations</h3>\n<p>Benchmarks show MPS running about 3x slower than an RTX 4090 for equivalent models. However, MPS uses significantly less power, making it suitable for development and smaller-scale training. The unified memory architecture allows loading larger models than would fit in discrete GPU VRAM.</p>\n<hr />\n<h2 id=\"coreml\">CoreML</h2>\n<p>CoreML is Apple’s deployment framework for machine learning models. Unlike MLX or PyTorch, CoreML is designed for inference in production applications across all Apple platforms.</p>\n<h3 id=\"the-coreml-stack\">The CoreML Stack</h3>\n<p>CoreML sits at the top of Apple’s ML stack. When you load a CoreML model, the framework:</p>\n<ol>\n<li>Parses the model format (.mlmodel or .mlpackage)</li>\n<li>Analyzes the model structure to determine optimal execution</li>\n<li>Compiles operations for available hardware (CPU, GPU, Neural Engine)</li>\n<li>Creates an execution plan that may span multiple processors</li>\n</ol>\n<p>The key insight is that CoreML handles hardware targeting automatically. The same model file runs on an iPhone, iPad, Mac, Apple Watch, or Vision Pro, with CoreML selecting the best execution strategy for each device.</p>\n<h3 id=\"neural-engine-targeting\">Neural Engine Targeting</h3>\n<p>The Apple Neural Engine (ANE) is a dedicated accelerator for matrix operations. It offers higher throughput than the GPU for certain workloads while using less power. CoreML is the primary way to access the ANE.</p>\n<p>However, ANE compatibility is strict. Operations must match specific constraints:</p>\n<ul>\n<li>Tensor dimensions must align to hardware requirements</li>\n<li>Certain operations have no ANE implementation</li>\n<li>If any layer in a path cannot run on ANE, the entire path falls back</li>\n</ul>\n<p>Apple’s research on <a href=\"https://machinelearning.apple.com/research/neural-engine-transformers\">deploying transformers to the Neural Engine</a> details the constraints and optimization strategies.</p>\n<h3 id=\"model-conversion-with-coremltools\">Model Conversion with coremltools</h3>\n<p>The <code>coremltools</code> Python package converts models from PyTorch, TensorFlow, and other frameworks to CoreML format:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> coremltools </span><span>as</span><span> ct</span></span>\n<span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load PyTorch model</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> MyModel</span><span>()</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>eval</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Create example input</span></span>\n<span class=\"line\"><span>example_input </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>1</span><span>, </span><span>3</span><span>, </span><span>224</span><span>, </span><span>224</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Trace the model</span></span>\n<span class=\"line\"><span>traced_model </span><span>=</span><span> torch</span><span>.</span><span>jit</span><span>.</span><span>trace</span><span>(model, example_input)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Convert to CoreML</span></span>\n<span class=\"line\"><span>mlmodel </span><span>=</span><span> ct</span><span>.</span><span>convert</span><span>(</span></span>\n<span class=\"line\"><span>    traced_model,</span></span>\n<span class=\"line\"><span>    inputs</span><span>=</span><span>[ct.</span><span>TensorType</span><span>(shape</span><span>=</span><span>example_input.shape)],</span></span>\n<span class=\"line\"><span>    minimum_deployment_target</span><span>=</span><span>ct.target.iOS17</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Save the model</span></span>\n<span class=\"line\"><span>mlmodel</span><span>.</span><span>save</span><span>(</span><span>\"MyModel.mlpackage\"</span><span>)</span></span></code></pre>\n<p>Direct conversion from PyTorch is recommended over the older ONNX intermediate format. The PyTorch converter handles more operations and produces better-optimized models.</p>\n<h3 id=\"stateful-models-in-ios-18\">Stateful Models in iOS 18</h3>\n<p>iOS 18 introduced stateful models to CoreML. This feature enables KV-cache for transformer inference without manual buffer management:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> coremltools </span><span>as</span><span> ct</span></span>\n<span class=\"line\"><span>from</span><span> coremltools</span><span>.</span><span>converters</span><span>.</span><span>mil </span><span>import</span><span> Builder </span><span>as</span><span> mb</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Define state for KV-cache</span></span>\n<span class=\"line\"><span>kv_cache_state </span><span>=</span><span> ct</span><span>.</span><span>StateType</span><span>(</span></span>\n<span class=\"line\"><span>    wrapped_type</span><span>=</span><span>ct.</span><span>TensorType</span><span>(shape</span><span>=</span><span>(batch, heads, seq_len, head_dim)),</span></span>\n<span class=\"line\"><span>    name</span><span>=</span><span>\"kv_cache\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Convert with state</span></span>\n<span class=\"line\"><span>mlmodel </span><span>=</span><span> ct</span><span>.</span><span>convert</span><span>(</span></span>\n<span class=\"line\"><span>    traced_model,</span></span>\n<span class=\"line\"><span>    states</span><span>=</span><span>[kv_cache_state],</span></span>\n<span class=\"line\"><span>    minimum_deployment_target</span><span>=</span><span>ct.target.iOS18</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>The state persists across inference calls, avoiding expensive memory allocations for each token generation step.</p>\n<h3 id=\"on-device-inference-performance\">On-Device Inference Performance</h3>\n<p>Apple’s benchmarks show Llama 3.1 8B running at approximately 33 tokens per second on an M1 Max using CoreML with 4-bit quantization. Smaller models achieve faster speeds. On mobile devices, the Neural Engine enables real-time inference for vision and audio models while maintaining battery efficiency.</p>\n<hr />\n<h2 id=\"coreml-vs-mlx-when-to-use-each\">CoreML vs MLX: When to Use Each</h2>\n<p>CoreML and MLX serve different purposes despite both running on Apple Silicon. Understanding their trade-offs helps you choose the right tool.</p>\n<h3 id=\"coreml-strengths\">CoreML Strengths</h3>\n<p><strong>Production deployment</strong>: CoreML integrates with Swift and Objective-C. It is the standard way to ship ML models in iOS, iPadOS, watchOS, tvOS, and visionOS apps.</p>\n<p><strong>Neural Engine access</strong>: CoreML is currently the only way to run models on the ANE. For power-sensitive mobile applications, this matters.</p>\n<p><strong>Hardware abstraction</strong>: One model file works across all Apple devices. CoreML handles the details of targeting different chip generations.</p>\n<p><strong>App Store ready</strong>: CoreML models work within Apple’s sandboxing and code signing requirements.</p>\n<h3 id=\"mlx-strengths\">MLX Strengths</h3>\n<p><strong>Research and experimentation</strong>: MLX provides a familiar NumPy/PyTorch-like API for rapid prototyping. You can modify models and run experiments interactively.</p>\n<p><strong>Training support</strong>: MLX supports gradient computation and training. CoreML is inference-only.</p>\n<p><strong>Fine-tuning LLMs</strong>: MLX includes tools for LoRA and QLoRA fine-tuning of language models. This is not possible with CoreML.</p>\n<p><strong>Dynamic computation</strong>: MLX handles dynamic shapes and control flow naturally. CoreML requires static compilation.</p>\n<h3 id=\"decision-framework\">Decision Framework</h3>\n<p><strong>Use CoreML when:</strong></p>\n<ul>\n<li>Building iOS, watchOS, tvOS, or visionOS apps</li>\n<li>Power efficiency is critical (mobile devices)</li>\n<li>You need Neural Engine acceleration</li>\n<li>Deploying to end users through the App Store</li>\n</ul>\n<p><strong>Use MLX when:</strong></p>\n<ul>\n<li>Running experiments on your Mac</li>\n<li>Training or fine-tuning models</li>\n<li>Working with LLMs interactively</li>\n<li>Building macOS developer tools</li>\n</ul>\n<p><strong>Use both when:</strong></p>\n<ul>\n<li>Developing on Mac with MLX, then converting to CoreML for deployment</li>\n<li>Research teams shipping production apps</li>\n</ul>\n<p>The typical workflow: experiment with MLX, export to a standard format, convert to CoreML for deployment.</p>\n<hr />\n<h2 id=\"metal-compute-shaders\">Metal Compute Shaders</h2>\n<p>When MPS does not provide the operation you need, you can write custom GPU kernels in Metal Shading Language (MSL).</p>\n<h3 id=\"metal-shading-language-basics\">Metal Shading Language Basics</h3>\n<p>MSL is based on C++14 with GPU-specific extensions. Compute kernels are marked with the <code>kernel</code> keyword:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>#include &lt;metal_stdlib&gt;</span></span>\n<span class=\"line\"><span>using namespace metal;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>kernel void vector_add(</span></span>\n<span class=\"line\"><span>    device const float* a [[buffer(0)]],</span></span>\n<span class=\"line\"><span>    device const float* b [[buffer(1)]],</span></span>\n<span class=\"line\"><span>    device float* result [[buffer(2)]],</span></span>\n<span class=\"line\"><span>    uint index [[thread_position_in_grid]]</span></span>\n<span class=\"line\"><span>) {</span></span>\n<span class=\"line\"><span>    result[index] = a[index] + b[index];</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The <code>[[buffer(N)]]</code> attributes specify buffer bindings. The <code>[[thread_position_in_grid]]</code> attribute provides the thread’s index in the dispatch grid.</p>\n<h3 id=\"dispatching-compute-work\">Dispatching Compute Work</h3>\n<p>From Swift, you create a compute pipeline and dispatch work:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> Metal</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Get the default GPU</span></span>\n<span class=\"line\"><span>let</span><span> device </span><span>=</span><span> MTLCreateSystemDefaultDevice</span><span>()</span><span>!</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Load the shader library</span></span>\n<span class=\"line\"><span>let</span><span> library </span><span>=</span><span> device.</span><span>makeDefaultLibrary</span><span>()</span><span>!</span></span>\n<span class=\"line\"><span>let</span><span> function </span><span>=</span><span> library.</span><span>makeFunction</span><span>(</span><span>name</span><span>:</span><span> \"vector_add\"</span><span>)</span><span>!</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Create compute pipeline</span></span>\n<span class=\"line\"><span>let</span><span> pipeline </span><span>=</span><span> try!</span><span> device.</span><span>makeComputePipelineState</span><span>(</span><span>function</span><span>:</span><span> function</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Create command queue and buffer</span></span>\n<span class=\"line\"><span>let</span><span> commandQueue </span><span>=</span><span> device.</span><span>makeCommandQueue</span><span>()</span><span>!</span></span>\n<span class=\"line\"><span>let</span><span> commandBuffer </span><span>=</span><span> commandQueue.</span><span>makeCommandBuffer</span><span>()</span><span>!</span></span>\n<span class=\"line\"><span>let</span><span> encoder </span><span>=</span><span> commandBuffer.</span><span>makeComputeCommandEncoder</span><span>()</span><span>!</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Set pipeline and buffers</span></span>\n<span class=\"line\"><span>encoder.</span><span>setComputePipelineState</span><span>(</span><span>pipeline</span><span>)</span></span>\n<span class=\"line\"><span>encoder.</span><span>setBuffer</span><span>(</span><span>bufferA, offset</span><span>:</span><span> 0</span><span>, index</span><span>:</span><span> 0</span><span>)</span></span>\n<span class=\"line\"><span>encoder.</span><span>setBuffer</span><span>(</span><span>bufferB, offset</span><span>:</span><span> 0</span><span>, index</span><span>:</span><span> 1</span><span>)</span></span>\n<span class=\"line\"><span>encoder.</span><span>setBuffer</span><span>(</span><span>bufferResult, offset</span><span>:</span><span> 0</span><span>, index</span><span>:</span><span> 2</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Calculate thread groups</span></span>\n<span class=\"line\"><span>let</span><span> threadGroupSize </span><span>=</span><span> MTLSize</span><span>(</span><span>width</span><span>:</span><span> 256</span><span>, height</span><span>:</span><span> 1</span><span>, depth</span><span>:</span><span> 1</span><span>)</span></span>\n<span class=\"line\"><span>let</span><span> threadGroups </span><span>=</span><span> MTLSize</span><span>(</span></span>\n<span class=\"line\"><span>    width</span><span>:</span><span> (elementCount </span><span>+</span><span> 255</span><span>) </span><span>/</span><span> 256</span><span>,</span></span>\n<span class=\"line\"><span>    height</span><span>:</span><span> 1</span><span>,</span></span>\n<span class=\"line\"><span>    depth</span><span>:</span><span> 1</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>encoder.</span><span>dispatchThreadgroups</span><span>(</span><span>threadGroups, threadsPerThreadgroup</span><span>:</span><span> threadGroupSize</span><span>)</span></span>\n<span class=\"line\"><span>encoder.</span><span>endEncoding</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>commandBuffer.</span><span>commit</span><span>()</span></span>\n<span class=\"line\"><span>commandBuffer.</span><span>waitUntilCompleted</span><span>()</span></span></code></pre>\n<h3 id=\"custom-kernels-for-ml\">Custom Kernels for ML</h3>\n<p>Apple provides sample code for implementing custom PyTorch operations with Metal. This allows you to write performance-critical operations in MSL while using PyTorch for the rest of your model.</p>\n<p>A typical pattern for ML kernels:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>kernel void fused_silu_multiply(</span></span>\n<span class=\"line\"><span>    device const float* input [[buffer(0)]],</span></span>\n<span class=\"line\"><span>    device const float* gate [[buffer(1)]],</span></span>\n<span class=\"line\"><span>    device float* output [[buffer(2)]],</span></span>\n<span class=\"line\"><span>    uint index [[thread_position_in_grid]]</span></span>\n<span class=\"line\"><span>) {</span></span>\n<span class=\"line\"><span>    float x = input[index];</span></span>\n<span class=\"line\"><span>    float g = gate[index];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // SiLU: x * sigmoid(x)</span></span>\n<span class=\"line\"><span>    float silu = x / (1.0 + exp(-x));</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Multiply with gate</span></span>\n<span class=\"line\"><span>    output[index] = silu * g;</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Fusing operations like this reduces memory bandwidth by avoiding intermediate buffers.</p>\n<h3 id=\"debugging-metal-shaders\">Debugging Metal Shaders</h3>\n<p>MSL debugging is challenging. The Xcode GPU Debugger allows stepping through shaders, but only for captured frames. For compute shaders without rendering, you can:</p>\n<ol>\n<li>Write intermediate values to a debug buffer</li>\n<li>Use Metal’s GPU capture feature in Xcode</li>\n<li>Add assertions that write to a status buffer on failure</li>\n</ol>\n<hr />\n<h2 id=\"bnns-cpu-optimized-neural-networks\">BNNS: CPU-Optimized Neural Networks</h2>\n<p>Basic Neural Network Subroutines (BNNS) is Apple’s library for CPU-based neural network inference. Part of the Accelerate framework, BNNS provides operations tuned for Apple’s CPU architectures.</p>\n<h3 id=\"when-to-use-bnns\">When to Use BNNS</h3>\n<p>BNNS makes sense when:</p>\n<ul>\n<li>The model is small enough that GPU overhead exceeds compute time</li>\n<li>Real-time requirements demand predictable latency (no GPU scheduling)</li>\n<li>Power constraints favor CPU over GPU</li>\n<li>You need to run inference on Apple Watch, which has limited GPU capabilities</li>\n</ul>\n<p>CoreML uses BNNS internally for CPU execution paths. You can also use BNNS directly for custom inference pipelines.</p>\n<h3 id=\"bnns-graph-api\">BNNS Graph API</h3>\n<p>The BNNSGraph API, introduced in iOS 17 and expanded in iOS 18, allows you to define entire networks as graphs:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> Accelerate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Create a graph</span></span>\n<span class=\"line\"><span>var</span><span> graph </span><span>=</span><span> BNNSGraph</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Add layers</span></span>\n<span class=\"line\"><span>let</span><span> convLayer </span><span>=</span><span> BNNSGraphAddConvolutionLayer</span><span>(</span></span>\n<span class=\"line\"><span>    graph,</span></span>\n<span class=\"line\"><span>    inputDescriptor,</span></span>\n<span class=\"line\"><span>    weightDescriptor,</span></span>\n<span class=\"line\"><span>    biasDescriptor,</span></span>\n<span class=\"line\"><span>    outputDescriptor,</span></span>\n<span class=\"line\"><span>    convolutionDescriptor</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Compile the graph</span></span>\n<span class=\"line\"><span>BNNSGraphCompile</span><span>(</span><span>graph, </span><span>nil</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Execute</span></span>\n<span class=\"line\"><span>BNNSGraphExecute</span><span>(</span><span>graph, inputBuffer, outputBuffer</span><span>)</span></span></code></pre>\n<h3 id=\"graph-optimizations\">Graph Optimizations</h3>\n<p>BNNSGraph performs several optimizations:</p>\n<ul>\n<li><strong>Layer fusion</strong>: Combines convolution + batch norm + activation into single operations</li>\n<li><strong>Copy elision</strong>: Eliminates unnecessary memory copies by using references</li>\n<li><strong>Memory sharing</strong>: Reuses buffers across layers when tensors have non-overlapping lifetimes</li>\n<li><strong>Weight repacking</strong>: Reorganizes weights for better cache locality</li>\n</ul>\n<p>These optimizations happen automatically when you compile the graph.</p>\n<h3 id=\"bnns-graph-builder-in-ios-18\">BNNS Graph Builder in iOS 18</h3>\n<p>iOS 18 added BNNSGraphBuilder, which lets you construct graphs directly in Swift:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> Accelerate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>let</span><span> builder </span><span>=</span><span> BNNSGraphBuilder</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>let</span><span> input </span><span>=</span><span> builder.</span><span>addInput</span><span>(</span><span>shape</span><span>:</span><span> [</span><span>1</span><span>, </span><span>224</span><span>, </span><span>224</span><span>, </span><span>3</span><span>], dataType</span><span>:</span><span> .</span><span>float</span><span>)</span></span>\n<span class=\"line\"><span>let</span><span> conv </span><span>=</span><span> builder.</span><span>addConvolution</span><span>(</span><span>input, weights</span><span>:</span><span> weights, bias</span><span>:</span><span> bias</span><span>)</span></span>\n<span class=\"line\"><span>let</span><span> relu </span><span>=</span><span> builder.</span><span>addActivation</span><span>(</span><span>conv, function</span><span>:</span><span> .relu</span><span>)</span></span>\n<span class=\"line\"><span>let</span><span> output </span><span>=</span><span> builder.</span><span>addOutput</span><span>(</span><span>relu</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>let</span><span> graph </span><span>=</span><span> try</span><span> builder.</span><span>build</span><span>()</span></span></code></pre>\n<p>This provides a more ergonomic API while still benefiting from graph-level optimizations.</p>\n<hr />\n<h2 id=\"the-accelerate-framework\">The Accelerate Framework</h2>\n<p>Accelerate is Apple’s foundational library for numerical computing. It provides BLAS, LAPACK, vDSP, and other optimized routines that underpin both BNNS and higher-level ML frameworks.</p>\n<h3 id=\"blas-and-lapack\">BLAS and LAPACK</h3>\n<p>The Basic Linear Algebra Subprograms (BLAS) and Linear Algebra Package (LAPACK) implementations in Accelerate are tuned for Apple hardware:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> Accelerate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Matrix multiplication using BLAS</span></span>\n<span class=\"line\"><span>var</span><span> C </span><span>=</span><span> [</span><span>Float</span><span>]</span><span>(</span><span>repeating</span><span>:</span><span> 0</span><span>, count</span><span>:</span><span> m </span><span>*</span><span> n</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>cblas_sgemm</span><span>(</span></span>\n<span class=\"line\"><span>    CblasRowMajor,    </span><span>// Row-major order</span></span>\n<span class=\"line\"><span>    CblasNoTrans,     </span><span>// Don't transpose A</span></span>\n<span class=\"line\"><span>    CblasNoTrans,     </span><span>// Don't transpose B</span></span>\n<span class=\"line\"><span>    Int32</span><span>(</span><span>m</span><span>)</span><span>,         </span><span>// Rows of A</span></span>\n<span class=\"line\"><span>    Int32</span><span>(</span><span>n</span><span>)</span><span>,         </span><span>// Columns of B</span></span>\n<span class=\"line\"><span>    Int32</span><span>(</span><span>k</span><span>)</span><span>,         </span><span>// Columns of A / Rows of B</span></span>\n<span class=\"line\"><span>    1.0</span><span>,              </span><span>// Alpha</span></span>\n<span class=\"line\"><span>    A, </span><span>Int32</span><span>(</span><span>k</span><span>)</span><span>,      </span><span>// A and its leading dimension</span></span>\n<span class=\"line\"><span>    B, </span><span>Int32</span><span>(</span><span>n</span><span>)</span><span>,      </span><span>// B and its leading dimension</span></span>\n<span class=\"line\"><span>    0.0</span><span>,              </span><span>// Beta</span></span>\n<span class=\"line\"><span>    &amp;</span><span>C, </span><span>Int32</span><span>(</span><span>n</span><span>)</span><span>      // C and its leading dimension</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>These routines automatically use the AMX (Apple Matrix Extensions) coprocessor when available. The AMX provides matrix multiplication throughput roughly 2x higher than standard NEON SIMD instructions.</p>\n<h3 id=\"vdsp-for-signal-processing\">vDSP for Signal Processing</h3>\n<p>vDSP provides optimized routines for:</p>\n<ul>\n<li>Fast Fourier transforms</li>\n<li>Convolution and correlation</li>\n<li>Vector arithmetic</li>\n<li>Biquad filtering</li>\n</ul>\n<p>For ML workloads, vDSP is useful for audio preprocessing, spectrogram computation, and other signal processing pipelines that feed into neural networks:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> Accelerate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Compute magnitude spectrum</span></span>\n<span class=\"line\"><span>var</span><span> magnitudes </span><span>=</span><span> [</span><span>Float</span><span>]</span><span>(</span><span>repeating</span><span>:</span><span> 0</span><span>, count</span><span>:</span><span> fftLength </span><span>/</span><span> 2</span><span>)</span></span>\n<span class=\"line\"><span>vDSP.</span><span>squareMagnitudes</span><span>(</span></span>\n<span class=\"line\"><span>    splitComplex,</span></span>\n<span class=\"line\"><span>    result</span><span>:</span><span> &amp;</span><span>magnitudes</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Convert to decibels</span></span>\n<span class=\"line\"><span>vDSP.</span><span>convert</span><span>(</span></span>\n<span class=\"line\"><span>    amplitude</span><span>:</span><span> magnitudes,</span></span>\n<span class=\"line\"><span>    toDecibels</span><span>:</span><span> &amp;</span><span>decibelSpectrum,</span></span>\n<span class=\"line\"><span>    zeroReference</span><span>:</span><span> 1.0</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"integration-with-ml-frameworks\">Integration with ML Frameworks</h3>\n<p>Accelerate functions are called internally by BNNS, CoreML (for CPU execution), and even MLX. When you see high CPU utilization during model inference, Accelerate routines are often doing the work.</p>\n<p>For custom preprocessing pipelines, using Accelerate directly is often faster than equivalent NumPy operations in Python:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// Fast image normalization</span></span>\n<span class=\"line\"><span>vDSP.</span><span>divide</span><span>(</span></span>\n<span class=\"line\"><span>    pixels,</span></span>\n<span class=\"line\"><span>    255.0</span><span>,</span></span>\n<span class=\"line\"><span>    result</span><span>:</span><span> &amp;</span><span>normalizedPixels</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>vDSP.</span><span>subtract</span><span>(</span></span>\n<span class=\"line\"><span>    normalizedPixels,</span></span>\n<span class=\"line\"><span>    mean,</span></span>\n<span class=\"line\"><span>    result</span><span>:</span><span> &amp;</span><span>centeredPixels</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>vDSP.</span><span>divide</span><span>(</span></span>\n<span class=\"line\"><span>    centeredPixels,</span></span>\n<span class=\"line\"><span>    stdDev,</span></span>\n<span class=\"line\"><span>    result</span><span>:</span><span> &amp;</span><span>standardizedPixels</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<hr />\n<h2 id=\"tensorflow-metal\">TensorFlow Metal</h2>\n<p>TensorFlow supports Apple Silicon GPUs through the <code>tensorflow-metal</code> plugin, which implements TensorFlow’s PluggableDevice API.</p>\n<h3 id=\"installation\">Installation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>pip</span><span> install</span><span> tensorflow</span><span> tensorflow-metal</span></span></code></pre>\n<h3 id=\"current-state\">Current State</h3>\n<p>The TensorFlow Metal plugin works but has significant limitations:</p>\n<p><strong>Version compatibility</strong>: As of early 2025, the plugin requires specific combinations of TensorFlow, Python, and macOS versions. The latest wheels support macOS 12 and Python up to 3.11. Users on macOS 15 with Python 3.12 face compatibility issues.</p>\n<p><strong>Operation coverage</strong>: Not all TensorFlow operations have Metal implementations. Complex numbers (DT_COMPLEX64) are not supported.</p>\n<p><strong>Performance variability</strong>: For small models or small batch sizes, CPU execution may be faster due to GPU dispatch overhead.</p>\n<h3 id=\"verification\">Verification</h3>\n<p>To verify Metal acceleration is working:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> tensorflow </span><span>as</span><span> tf</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># List physical devices</span></span>\n<span class=\"line\"><span>devices </span><span>=</span><span> tf</span><span>.</span><span>config</span><span>.</span><span>list_physical_devices</span><span>()</span></span>\n<span class=\"line\"><span>print</span><span>(devices)</span></span>\n<span class=\"line\"><span># Should show both CPU:0 and GPU:0</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run a simple operation</span></span>\n<span class=\"line\"><span>with</span><span> tf</span><span>.</span><span>device</span><span>(</span><span>'/GPU:0'</span><span>):</span></span>\n<span class=\"line\"><span>    a </span><span>=</span><span> tf</span><span>.</span><span>random</span><span>.</span><span>normal</span><span>([</span><span>1000</span><span>, </span><span>1000</span><span>])</span></span>\n<span class=\"line\"><span>    b </span><span>=</span><span> tf</span><span>.</span><span>matmul</span><span>(a, tf.</span><span>transpose</span><span>(a))</span></span>\n<span class=\"line\"><span>    print</span><span>(b.device)</span><span>  # Should show GPU</span></span></code></pre>\n<h3 id=\"recommendation\">Recommendation</h3>\n<p>For new projects on Apple Silicon, PyTorch MPS or MLX typically offer better support and performance than TensorFlow Metal. TensorFlow Metal remains useful for existing TensorFlow codebases that need to run on Mac hardware.</p>\n<hr />\n<h2 id=\"model-conversion-pipelines\">Model Conversion Pipelines</h2>\n<p>Converting models to CoreML involves several steps and decisions about optimization.</p>\n<h3 id=\"pytorch-to-coreml-direct\">PyTorch to CoreML (Direct)</h3>\n<p>The recommended path for PyTorch models:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> coremltools </span><span>as</span><span> ct</span></span>\n<span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> MyModel</span><span>()</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>eval</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Trace or script the model</span></span>\n<span class=\"line\"><span>example_input </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>1</span><span>, </span><span>3</span><span>, </span><span>224</span><span>, </span><span>224</span><span>)</span></span>\n<span class=\"line\"><span>traced </span><span>=</span><span> torch</span><span>.</span><span>jit</span><span>.</span><span>trace</span><span>(model, example_input)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Convert with compute units specification</span></span>\n<span class=\"line\"><span>mlmodel </span><span>=</span><span> ct</span><span>.</span><span>convert</span><span>(</span></span>\n<span class=\"line\"><span>    traced,</span></span>\n<span class=\"line\"><span>    inputs</span><span>=</span><span>[ct.</span><span>TensorType</span><span>(name</span><span>=</span><span>\"image\"</span><span>, shape</span><span>=</span><span>example_input.shape)],</span></span>\n<span class=\"line\"><span>    compute_units</span><span>=</span><span>ct.ComputeUnit.ALL,  </span><span># CPU, GPU, and Neural Engine</span></span>\n<span class=\"line\"><span>    minimum_deployment_target</span><span>=</span><span>ct.target.iOS17</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"onnx-to-coreml\">ONNX to CoreML</h3>\n<p>For models from other frameworks or ONNX model zoo:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> coremltools </span><span>as</span><span> ct</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load ONNX model</span></span>\n<span class=\"line\"><span>mlmodel </span><span>=</span><span> ct</span><span>.</span><span>converters</span><span>.</span><span>onnx</span><span>.</span><span>convert</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"model.onnx\"</span><span>,</span></span>\n<span class=\"line\"><span>    minimum_deployment_target</span><span>=</span><span>ct.target.iOS16</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>Note that the ONNX path may have more conversion issues than direct PyTorch conversion.</p>\n<h3 id=\"quantization-during-conversion\">Quantization During Conversion</h3>\n<p>coremltools 7+ provides optimization APIs for quantization:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> coremltools </span><span>as</span><span> ct</span></span>\n<span class=\"line\"><span>from</span><span> coremltools</span><span>.</span><span>optimize</span><span>.</span><span>coreml </span><span>import</span><span> (</span></span>\n<span class=\"line\"><span>    OpLinearQuantizerConfig</span><span>,</span></span>\n<span class=\"line\"><span>    OptimizationConfig</span><span>,</span></span>\n<span class=\"line\"><span>    linear_quantize_weights</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load the model</span></span>\n<span class=\"line\"><span>mlmodel </span><span>=</span><span> ct</span><span>.</span><span>models</span><span>.</span><span>MLModel</span><span>(</span><span>\"model.mlpackage\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Configure 8-bit quantization</span></span>\n<span class=\"line\"><span>config </span><span>=</span><span> OptimizationConfig</span><span>(</span></span>\n<span class=\"line\"><span>    global_config</span><span>=</span><span>OpLinearQuantizerConfig</span><span>(mode</span><span>=</span><span>\"linear_symmetric\"</span><span>)</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Quantize weights</span></span>\n<span class=\"line\"><span>quantized_model </span><span>=</span><span> linear_quantize_weights</span><span>(mlmodel, config)</span></span>\n<span class=\"line\"><span>quantized_model</span><span>.</span><span>save</span><span>(</span><span>\"model_int8.mlpackage\"</span><span>)</span></span></code></pre>\n<p>For 4-bit quantization (useful for LLMs):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> coremltools</span><span>.</span><span>optimize</span><span>.</span><span>coreml </span><span>import</span><span> (</span></span>\n<span class=\"line\"><span>    OpPalettizerConfig</span><span>,</span></span>\n<span class=\"line\"><span>    OptimizationConfig</span><span>,</span></span>\n<span class=\"line\"><span>    palettize_weights</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>config </span><span>=</span><span> OptimizationConfig</span><span>(</span></span>\n<span class=\"line\"><span>    global_config</span><span>=</span><span>OpPalettizerConfig</span><span>(</span></span>\n<span class=\"line\"><span>        nbits</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>        mode</span><span>=</span><span>\"kmeans\"</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>quantized_model </span><span>=</span><span> palettize_weights</span><span>(mlmodel, config)</span></span></code></pre>\n<p>Block-wise quantization provides better accuracy for aggressive compression:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>config </span><span>=</span><span> OptimizationConfig</span><span>(</span></span>\n<span class=\"line\"><span>    global_config</span><span>=</span><span>OpLinearQuantizerConfig</span><span>(</span></span>\n<span class=\"line\"><span>        mode</span><span>=</span><span>\"linear_symmetric\"</span><span>,</span></span>\n<span class=\"line\"><span>        weight_threshold</span><span>=</span><span>512</span><span>,  </span><span># Only quantize weights larger than this</span></span>\n<span class=\"line\"><span>        granularity</span><span>=</span><span>\"per_block\"</span><span>,</span></span>\n<span class=\"line\"><span>        block_size</span><span>=</span><span>32</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"training-aware-quantization\">Training-Aware Quantization</h3>\n<p>For best results, quantize during training using coremltools.optimize.torch:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> coremltools</span><span>.</span><span>optimize</span><span>.</span><span>torch</span><span>.</span><span>quantization </span><span>import</span><span> (</span></span>\n<span class=\"line\"><span>    LinearQuantizerConfig</span><span>,</span></span>\n<span class=\"line\"><span>    ModuleLinearQuantizerConfig</span><span>,</span></span>\n<span class=\"line\"><span>    LinearQuantizer</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Configure quantization</span></span>\n<span class=\"line\"><span>config </span><span>=</span><span> LinearQuantizerConfig</span><span>.</span><span>from_dict</span><span>({</span></span>\n<span class=\"line\"><span>    \"global_config\"</span><span>: {</span></span>\n<span class=\"line\"><span>        \"quantization_scheme\"</span><span>: </span><span>\"symmetric\"</span><span>,</span></span>\n<span class=\"line\"><span>        \"milestones\"</span><span>: [</span><span>0</span><span>, </span><span>100</span><span>, </span><span>400</span><span>, </span><span>500</span><span>]</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>})</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Create quantizer</span></span>\n<span class=\"line\"><span>quantizer </span><span>=</span><span> LinearQuantizer</span><span>(model, config)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Prepare model for quantization-aware training</span></span>\n<span class=\"line\"><span>quantizer</span><span>.</span><span>prepare</span><span>(example_inputs</span><span>=</span><span>example_input)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Train the model</span></span>\n<span class=\"line\"><span>for</span><span> epoch </span><span>in</span><span> range</span><span>(num_epochs):</span></span>\n<span class=\"line\"><span>    quantizer</span><span>.</span><span>step</span><span>()</span></span>\n<span class=\"line\"><span>    train_one_epoch</span><span>(model)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Finalize and convert</span></span>\n<span class=\"line\"><span>quantizer</span><span>.</span><span>finalize</span><span>()</span></span>\n<span class=\"line\"><span>mlmodel </span><span>=</span><span> ct</span><span>.</span><span>convert</span><span>(model)</span></span></code></pre>\n<hr />\n<h2 id=\"practical-deployment-considerations\">Practical Deployment Considerations</h2>\n<p>Deploying ML models across Apple platforms requires understanding the constraints and capabilities of each target.</p>\n<h3 id=\"ios-and-ipados\">iOS and iPadOS</h3>\n<p><strong>Model size limits</strong>: App Store apps have size limits. Use compression and quantization aggressively. Consider downloading models on first launch.</p>\n<p><strong>Memory constraints</strong>: iPhones have limited RAM. The system will terminate apps that use too much memory. Profile your model’s memory footprint.</p>\n<p><strong>Thermal throttling</strong>: Sustained inference causes heat buildup. The system reduces performance to manage temperature. Design for burst usage patterns.</p>\n<p><strong>Background execution</strong>: Background apps have limited CPU/GPU access. Use background tasks API for non-real-time inference.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> CoreML</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Load model with configuration</span></span>\n<span class=\"line\"><span>let</span><span> config </span><span>=</span><span> MLModelConfiguration</span><span>()</span></span>\n<span class=\"line\"><span>config.computeUnits </span><span>=</span><span> .cpuAndNeuralEngine  </span><span>// Avoid GPU for background</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>let</span><span> model </span><span>=</span><span> try</span><span> MyModel</span><span>(</span><span>configuration</span><span>:</span><span> config</span><span>)</span></span></code></pre>\n<h3 id=\"macos\">macOS</h3>\n<p>macOS has fewer constraints than mobile platforms:</p>\n<ul>\n<li>More memory available</li>\n<li>No thermal throttling concerns for desktop Macs</li>\n<li>Full GPU access without power restrictions</li>\n</ul>\n<p>For macOS-only apps, you can use MLX directly instead of CoreML if you prefer its API.</p>\n<h3 id=\"watchos\">watchOS</h3>\n<p>Apple Watch has the most constrained environment:</p>\n<ul>\n<li>Limited memory (varies by model)</li>\n<li>Small GPU with limited capabilities</li>\n<li>CPU-focused execution via BNNS</li>\n</ul>\n<p>Keep models small. Quantize aggressively. Consider CPU-only execution:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>let</span><span> config </span><span>=</span><span> MLModelConfiguration</span><span>()</span></span>\n<span class=\"line\"><span>config.computeUnits </span><span>=</span><span> .cpuOnly</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>let</span><span> model </span><span>=</span><span> try</span><span> WatchModel</span><span>(</span><span>configuration</span><span>:</span><span> config</span><span>)</span></span></code></pre>\n<h3 id=\"visionos\">visionOS</h3>\n<p>Vision Pro runs visionOS, which supports CoreML with full GPU and Neural Engine access. The M2 chip provides capable ML performance.</p>\n<p>Spatial computing apps may need to balance ML inference with rendering workloads. Consider using <code>computeUnits = .cpuAndNeuralEngine</code> to leave GPU headroom for graphics.</p>\n<h3 id=\"cross-platform-strategy\">Cross-Platform Strategy</h3>\n<p>For apps targeting multiple Apple platforms:</p>\n<ol>\n<li><strong>Use CoreML as the deployment format</strong>: One .mlpackage works everywhere</li>\n<li><strong>Test on real devices</strong>: Simulator does not accurately represent Neural Engine behavior</li>\n<li><strong>Handle graceful degradation</strong>: Check <code>MLModel.availableComputeDevices</code> and adjust</li>\n<li><strong>Profile each platform</strong>: Performance characteristics differ significantly</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// Check available compute devices</span></span>\n<span class=\"line\"><span>let</span><span> model </span><span>=</span><span> try</span><span> MyModel</span><span>()</span></span>\n<span class=\"line\"><span>let</span><span> devices </span><span>=</span><span> model.model.availableComputeDevices</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>if</span><span> devices.</span><span>contains</span><span>(</span><span>where</span><span>:</span><span> { $0 </span><span>==</span><span> .neuralEngine }</span><span>)</span><span> {</span></span>\n<span class=\"line\"><span>    // Can use Neural Engine</span></span>\n<span class=\"line\"><span>} </span><span>else</span><span> {</span></span>\n<span class=\"line\"><span>    // Fall back to CPU/GPU config</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<hr />\n<h2 id=\"the-amx-coprocessor\">The AMX Coprocessor</h2>\n<p>The Apple Matrix Coprocessor (AMX) is an undocumented accelerator present in all Apple Silicon chips. While you cannot program it directly, understanding its role helps explain performance characteristics.</p>\n<h3 id=\"what-amx-does\">What AMX Does</h3>\n<p>AMX accelerates matrix multiplication on the CPU. It provides roughly 2x the throughput of NEON SIMD instructions for matrix operations. When you call BLAS routines through Accelerate, AMX handles the heavy lifting.</p>\n<h3 id=\"access-through-accelerate\">Access Through Accelerate</h3>\n<p>The only supported way to use AMX is through the Accelerate framework. Direct AMX instructions exist but are undocumented and unsupported. Apple’s BLAS implementation automatically uses AMX when beneficial.</p>\n<p>This means CPU-based inference through BNNS or Accelerate benefits from AMX automatically. You do not need to do anything special.</p>\n<h3 id=\"amx-vs-neural-engine-vs-gpu\">AMX vs Neural Engine vs GPU</h3>\n<p>Each accelerator has different strengths:</p>\n<ul>\n<li><strong>AMX</strong>: Low latency, integrated with CPU, used automatically by Accelerate</li>\n<li><strong>Neural Engine</strong>: Highest throughput for supported operations, best power efficiency</li>\n<li><strong>GPU</strong>: Flexible, handles any compute workload, good for large batch sizes</li>\n</ul>\n<p>CoreML manages this complexity by analyzing your model and routing operations to the most appropriate hardware.</p>\n<hr />\n<h2 id=\"summary\">Summary</h2>\n<p>Apple’s Metal ML ecosystem provides multiple layers of optimization:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Layer</th><th>Purpose</th><th>When to Use</th></tr></thead><tbody><tr><td>CoreML</td><td>Production deployment</td><td>iOS/macOS/watchOS/visionOS apps</td></tr><tr><td>MLX</td><td>Research and training</td><td>Mac-based ML development</td></tr><tr><td>MPSGraph</td><td>GPU compute graphs</td><td>Custom inference engines</td></tr><tr><td>MPS</td><td>GPU primitives</td><td>Building custom ML ops</td></tr><tr><td>BNNS</td><td>CPU inference</td><td>Low-latency, power-sensitive workloads</td></tr><tr><td>Accelerate</td><td>Numerical computing</td><td>Preprocessing, custom algorithms</td></tr><tr><td>Metal</td><td>Custom GPU kernels</td><td>Operations not in MPS</td></tr></tbody></table></div>\n<p>For most developers, the right approach is:</p>\n<ol>\n<li>Train models using PyTorch (with MPS for GPU acceleration) or MLX</li>\n<li>Convert to CoreML using coremltools</li>\n<li>Apply quantization appropriate for your target platform</li>\n<li>Deploy with CoreML, letting it handle hardware routing</li>\n</ol>\n<p>The ecosystem continues evolving. WWDC 2024 brought transformer optimizations to MPSGraph. iOS 18 added stateful models to CoreML. Each release improves performance and adds capabilities. Keep your tools updated and follow Apple’s machine learning documentation for the latest guidance.</p>\n<hr />\n<h2 id=\"references\">References</h2>\n<ol>\n<li>Apple Developer Documentation - Metal Performance Shaders: <a href=\"https://developer.apple.com/documentation/metalperformanceshaders\">https://developer.apple.com/documentation/metalperformanceshaders</a></li>\n<li>Apple Developer - Accelerated PyTorch Training on Mac: <a href=\"https://developer.apple.com/metal/pytorch/\">https://developer.apple.com/metal/pytorch/</a></li>\n<li>PyTorch Documentation - MPS Backend: <a href=\"https://docs.pytorch.org/docs/stable/notes/mps.html\">https://docs.pytorch.org/docs/stable/notes/mps.html</a></li>\n<li>Apple Developer Documentation - Core ML: <a href=\"https://developer.apple.com/documentation/coreml\">https://developer.apple.com/documentation/coreml</a></li>\n<li>Apple Machine Learning Research - Deploying Transformers on the Apple Neural Engine: <a href=\"https://machinelearning.apple.com/research/neural-engine-transformers\">https://machinelearning.apple.com/research/neural-engine-transformers</a></li>\n<li>Apple Machine Learning Research - On-Device Llama 3.1 with Core ML: <a href=\"https://machinelearning.apple.com/research/core-ml-on-device-llama\">https://machinelearning.apple.com/research/core-ml-on-device-llama</a></li>\n<li>WWDC24 - Accelerate Machine Learning with Metal: <a href=\"https://developer.apple.com/videos/play/wwdc2024/10218/\">https://developer.apple.com/videos/play/wwdc2024/10218/</a></li>\n<li>WWDC24 - Support Real-Time ML Inference on the CPU: <a href=\"https://developer.apple.com/videos/play/wwdc2024/10211/\">https://developer.apple.com/videos/play/wwdc2024/10211/</a></li>\n<li>Apple Developer Documentation - BNNS: <a href=\"https://developer.apple.com/documentation/accelerate/bnns\">https://developer.apple.com/documentation/accelerate/bnns</a></li>\n<li>Apple Developer Documentation - Accelerate: <a href=\"https://developer.apple.com/documentation/accelerate\">https://developer.apple.com/documentation/accelerate</a></li>\n<li>Guide to Core ML Tools - Converting from PyTorch: <a href=\"https://apple.github.io/coremltools/docs-guides/source/convert-pytorch.html\">https://apple.github.io/coremltools/docs-guides/source/convert-pytorch.html</a></li>\n<li>Guide to Core ML Tools - Quantization Algorithms: <a href=\"https://apple.github.io/coremltools/docs-guides/source/opt-quantization-algos.html\">https://apple.github.io/coremltools/docs-guides/source/opt-quantization-algos.html</a></li>\n<li>Guide to Core ML Tools - Stateful Models: <a href=\"https://apple.github.io/coremltools/docs-guides/source/stateful-models.html\">https://apple.github.io/coremltools/docs-guides/source/stateful-models.html</a></li>\n<li>Apple Developer - Metal Shading Language Specification: <a href=\"https://developer.apple.com/metal/Metal-Shading-Language-Specification.pdf\">https://developer.apple.com/metal/Metal-Shading-Language-Specification.pdf</a></li>\n<li>Apple Developer Documentation - Performing Calculations on a GPU: <a href=\"https://developer.apple.com/documentation/Metal/performing-calculations-on-a-gpu\">https://developer.apple.com/documentation/Metal/performing-calculations-on-a-gpu</a></li>\n<li>Apple Developer - TensorFlow Metal Plugin: <a href=\"https://developer.apple.com/metal/tensorflow-plugin/\">https://developer.apple.com/metal/tensorflow-plugin/</a></li>\n<li>PyTorch GitHub - MPS Operator Coverage Tracking: <a href=\"https://github.com/pytorch/pytorch/issues/141287\">https://github.com/pytorch/pytorch/issues/141287</a></li>\n<li>Explosion AI - Fast Transformer Inference with Metal Performance Shaders: <a href=\"https://explosion.ai/blog/metal-performance-shaders\">https://explosion.ai/blog/metal-performance-shaders</a></li>\n<li>MLX Documentation: <a href=\"https://ml-explore.github.io/mlx/\">https://ml-explore.github.io/mlx/</a></li>\n<li>byby.dev - When to Use Apple MLX vs Core ML: <a href=\"https://byby.dev/apple-mlx-vs-coreml\">https://byby.dev/apple-mlx-vs-coreml</a></li>\n</ol>",
            "url": "https://blog.ecitis.org/apple-metal-ml-ecosystem/",
            "title": "Apple's Metal Ecosystem for Machine Learning: Beyond MLX",
            "summary": "Understand the complete Apple machine-learning stack—from Metal and MPSGraph to Core ML—and choose the right layer for the job.",
            "image": "https://blog.ecitis.org/open-graph/apple-metal-ml-ecosystem.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Hardware & Systems",
                "Metal",
                "Core ML",
                "Apple Silicon"
            ]
        },
        {
            "id": "https://blog.ecitis.org/apple-mlx-guide/",
            "content_html": "<h2 id=\"why-mlx-matters-for-apple-silicon-developers\">Why MLX Matters for Apple Silicon Developers</h2>\n<p>Apple released MLX in December 2023 as an open-source array framework built specifically for machine learning on Apple Silicon. Unlike PyTorch with MPS backend or TensorFlow with Metal plugin, MLX was designed from the ground up to exploit the unified memory architecture of M-series chips. The framework has matured through 2024 and 2025, reaching version 0.30+ with support for the M5 chip’s Neural Accelerators announced at WWDC 2025.</p>\n<p>This guide covers everything you need to know to start using MLX for machine learning research and inference on your Mac.</p>\n<hr />\n<h2 id=\"what-is-mlx\">What is MLX</h2>\n<p>MLX is an array framework for numerical computing and machine learning. It provides a NumPy-like API with support for automatic differentiation, JIT compilation, and GPU acceleration through Metal. The framework was developed by Apple’s machine learning research team and is actively maintained on GitHub at <a href=\"https://github.com/ml-explore/mlx\">ml-explore/mlx</a>.</p>\n<h3 id=\"core-design-principles\">Core Design Principles</h3>\n<p>MLX follows several design principles that distinguish it from other frameworks:</p>\n<p><strong>Familiar API</strong>: The Python interface mirrors NumPy closely. If you know NumPy, you can start using MLX immediately. Higher-level neural network APIs in <code>mlx.nn</code> follow PyTorch conventions.</p>\n<p><strong>Lazy Evaluation</strong>: Computations are not executed immediately. Instead, MLX builds a computation graph that is only evaluated when results are needed. This allows the framework to optimize and fuse operations before execution.</p>\n<p><strong>Unified Memory</strong>: Arrays live in shared memory accessible by both CPU and GPU. There is no need to copy data between devices.</p>\n<p><strong>Composable Function Transformations</strong>: Operations like <code>grad()</code>, <code>vmap()</code>, and <code>compile()</code> can be composed together. You can take the gradient of a compiled function, or compile a vectorized gradient computation.</p>\n<p><strong>Multi-Language Support</strong>: MLX has bindings for Python, Swift, C++, and C. The same models can run on macOS, iOS, iPadOS, and visionOS.</p>\n<h3 id=\"basic-array-operations\">Basic Array Operations</h3>\n<p>Here is a simple example showing MLX array operations:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Create arrays</span></span>\n<span class=\"line\"><span>a </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>1.0</span><span>, </span><span>2.0</span><span>, </span><span>3.0</span><span>])</span></span>\n<span class=\"line\"><span>b </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>4.0</span><span>, </span><span>5.0</span><span>, </span><span>6.0</span><span>])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Operations work like NumPy</span></span>\n<span class=\"line\"><span>c </span><span>=</span><span> a </span><span>+</span><span> b</span></span>\n<span class=\"line\"><span>d </span><span>=</span><span> mx</span><span>.</span><span>sin</span><span>(a)</span><span> *</span><span> mx</span><span>.</span><span>cos</span><span>(b)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Arrays are not computed until needed</span></span>\n<span class=\"line\"><span>print</span><span>(c)</span><span>  # This triggers evaluation</span></span></code></pre>\n<p>The framework supports all common array operations including broadcasting, slicing, reshaping, and linear algebra routines.</p>\n<hr />\n<h2 id=\"the-unified-memory-model\">The Unified Memory Model</h2>\n<p>The most significant architectural advantage of MLX comes from Apple Silicon’s unified memory design. On traditional systems with discrete GPUs, data must be copied between CPU RAM and GPU VRAM. This copying introduces latency and limits the effective memory available for large models.</p>\n<figure><figcaption><strong>Memory Architecture Comparison</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/apple-mlx-guide-0.webp\" width=\"512\" height=\"658\" alt=\"Memory Architecture Comparison\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"how-unified-memory-works\">How Unified Memory Works</h3>\n<p>On Apple Silicon, the CPU, GPU, and Neural Engine share the same physical memory pool. When you create an MLX array, it exists in this shared space. Any processor can access it directly without copying.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Create an array - it lives in unified memory</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> mx</span><span>.</span><span>random</span><span>.</span><span>normal</span><span>((</span><span>1000</span><span>, </span><span>1000</span><span>))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run on GPU - no copy needed</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> mx</span><span>.</span><span>matmul</span><span>(x, x.T, stream</span><span>=</span><span>mx.gpu)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run on CPU - still no copy needed</span></span>\n<span class=\"line\"><span>z </span><span>=</span><span> mx</span><span>.</span><span>sum</span><span>(y, stream</span><span>=</span><span>mx.cpu)</span></span></code></pre>\n<p>In MLX, you specify the device when calling an operation, not when creating the array. The array itself has no device affiliation. This eliminates common bugs in PyTorch where tensors end up on the wrong device.</p>\n<h3 id=\"memory-advantages-for-large-models\">Memory Advantages for Large Models</h3>\n<p>The unified memory model provides practical benefits for running large language models. A Mac with 128GB of unified memory can load models that would require a workstation with a high-end GPU to run on CUDA systems.</p>\n<p>However, there are limits. The GPU cannot typically use more than about 75% of total system memory. A 128GB Mac can allocate roughly 96GB for GPU tasks. This is still substantial compared to the 24GB available on consumer NVIDIA GPUs.</p>\n<hr />\n<h2 id=\"lazy-evaluation-deep-dive\">Lazy Evaluation Deep Dive</h2>\n<p>MLX uses lazy evaluation, meaning operations are recorded but not executed immediately. Understanding this model is essential for writing efficient MLX code.</p>\n<h3 id=\"how-lazy-evaluation-works\">How Lazy Evaluation Works</h3>\n<p>When you write <code>c = a + b</code>, MLX does not compute the sum. Instead, it creates a node in a computation graph with <code>a</code> and <code>b</code> as inputs. The actual computation only happens when you need the result.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>a </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>1.0</span><span>, </span><span>2.0</span><span>, </span><span>3.0</span><span>])</span></span>\n<span class=\"line\"><span>b </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>4.0</span><span>, </span><span>5.0</span><span>, </span><span>6.0</span><span>])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># No computation happens here</span></span>\n<span class=\"line\"><span>c </span><span>=</span><span> a </span><span>+</span><span> b</span></span>\n<span class=\"line\"><span>d </span><span>=</span><span> c </span><span>*</span><span> 2</span></span>\n<span class=\"line\"><span>e </span><span>=</span><span> mx</span><span>.</span><span>sum</span><span>(d)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Computation happens when we need the value</span></span>\n<span class=\"line\"><span>print</span><span>(e)</span><span>  # Triggers evaluation of entire graph</span></span></code></pre>\n<p>You can explicitly trigger evaluation with <code>mx.eval()</code>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>mx</span><span>.</span><span>eval</span><span>(e)</span><span>  # Force computation</span></span></code></pre>\n<h3 id=\"memory-benefits\">Memory Benefits</h3>\n<p>Lazy evaluation enables memory optimizations. Consider initializing a large model:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>nn </span><span>as</span><span> nn</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Model weights are \"created\" but not allocated</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> LargeModel</span><span>()</span><span>  # Uses float32 by default</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Update to float16 before evaluation</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>load_weights</span><span>(</span><span>\"weights.safetensors\"</span><span>)</span><span>  # Loads as float16</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Only float16 memory is ever allocated</span></span>\n<span class=\"line\"><span>mx</span><span>.</span><span>eval</span><span>(model.</span><span>parameters</span><span>())</span></span></code></pre>\n<p>If evaluation were eager, the model would first allocate float32 weights, then allocate float16 weights, doubling peak memory usage.</p>\n<h3 id=\"graph-fusion-and-optimization\">Graph Fusion and Optimization</h3>\n<p>The lazy evaluation model allows MLX to fuse operations. Multiple element-wise operations can be combined into a single GPU kernel, reducing memory bandwidth requirements and kernel launch overhead.</p>\n<hr />\n<h2 id=\"automatic-differentiation\">Automatic Differentiation</h2>\n<p>MLX implements automatic differentiation through function transformations rather than tape-based recording. This approach, inspired by JAX, provides more flexibility and composability.</p>\n<h3 id=\"the-grad-function\">The grad Function</h3>\n<p>The <code>mx.grad()</code> function takes a function and returns a new function that computes gradients:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> f</span><span>(</span><span>x</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> mx</span><span>.</span><span>sum</span><span>(x </span><span>**</span><span> 2</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Create gradient function</span></span>\n<span class=\"line\"><span>grad_f </span><span>=</span><span> mx</span><span>.</span><span>grad</span><span>(f)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>x </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>1.0</span><span>, </span><span>2.0</span><span>, </span><span>3.0</span><span>])</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>grad_f</span><span>(x))</span><span>  # [2.0, 4.0, 6.0]</span></span></code></pre>\n<p>Unlike PyTorch, there is no <code>backward()</code> method, no <code>zero_grad()</code>, no <code>requires_grad</code> property, and no <code>detach()</code>. Gradients are computed by transforming functions.</p>\n<h3 id=\"valueandgrad\">value_and_grad</h3>\n<p>Computing both the function value and gradient is common in optimization. Rather than calling the function twice, use <code>value_and_grad()</code>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> loss_fn</span><span>(</span><span>model</span><span>,</span><span> x</span><span>,</span><span> y</span><span>):</span></span>\n<span class=\"line\"><span>    pred </span><span>=</span><span> model</span><span>(x)</span></span>\n<span class=\"line\"><span>    return</span><span> mx</span><span>.</span><span>mean</span><span>((pred </span><span>-</span><span> y) </span><span>**</span><span> 2</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Get both loss value and gradients</span></span>\n<span class=\"line\"><span>loss_and_grad_fn </span><span>=</span><span> mx</span><span>.</span><span>value_and_grad</span><span>(model, loss_fn)</span></span>\n<span class=\"line\"><span>loss</span><span>,</span><span> grads </span><span>=</span><span> loss_and_grad_fn</span><span>(model, x, y)</span></span></code></pre>\n<h3 id=\"higher-order-gradients\">Higher-Order Gradients</h3>\n<p>Function transformations compose naturally. You can compute the gradient of a gradient:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> f</span><span>(</span><span>x</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> x </span><span>**</span><span> 3</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>grad_f </span><span>=</span><span> mx</span><span>.</span><span>grad</span><span>(f)</span><span>       # 3x^2</span></span>\n<span class=\"line\"><span>grad2_f </span><span>=</span><span> mx</span><span>.</span><span>grad</span><span>(grad_f)</span><span>  # 6x</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>x </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>(</span><span>2.0</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>grad2_f</span><span>(x))</span><span>  # 12.0</span></span></code></pre>\n<hr />\n<h2 id=\"jit-compilation-with-mxcompile\">JIT Compilation with mx.compile</h2>\n<p>MLX provides a <code>compile()</code> transformation that optimizes computation graphs. Compilation fuses operations, eliminates redundant computations, and generates optimized Metal kernels.</p>\n<h3 id=\"basic-usage\">Basic Usage</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> model_forward</span><span>(</span><span>x</span><span>,</span><span> w1</span><span>,</span><span> w2</span><span>):</span></span>\n<span class=\"line\"><span>    h </span><span>=</span><span> mx</span><span>.</span><span>tanh</span><span>(x </span><span>@</span><span> w1)</span></span>\n<span class=\"line\"><span>    return</span><span> h </span><span>@</span><span> w2</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Compile the function</span></span>\n<span class=\"line\"><span>compiled_forward </span><span>=</span><span> mx</span><span>.</span><span>compile</span><span>(model_forward)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># First call builds and compiles the graph (slow)</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> compiled_forward</span><span>(x, w1, w2)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Subsequent calls reuse compiled code (fast)</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> compiled_forward</span><span>(x, w1, w2)</span></span></code></pre>\n<p>The first call to a compiled function is slow because MLX must trace the computation, optimize the graph, and compile Metal shaders. Subsequent calls with the same input shapes reuse the compiled code.</p>\n<h3 id=\"composing-compile-with-other-transformations\">Composing compile with Other Transformations</h3>\n<p>Compilation works with other function transformations:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Compile a gradient function</span></span>\n<span class=\"line\"><span>compiled_grad </span><span>=</span><span> mx</span><span>.</span><span>compile</span><span>(mx.</span><span>grad</span><span>(loss_fn))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Compile a vectorized function</span></span>\n<span class=\"line\"><span>compiled_vmap </span><span>=</span><span> mx</span><span>.</span><span>compile</span><span>(mx.</span><span>vmap</span><span>(batch_fn))</span></span></code></pre>\n<h3 id=\"debugging-compiled-functions\">Debugging Compiled Functions</h3>\n<p>Compiled functions cannot contain print statements or other side effects because they are traced with placeholder inputs. For debugging, disable compilation:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Globally disable compilation</span></span>\n<span class=\"line\"><span>mx</span><span>.</span><span>disable_compile</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or use environment variable</span></span>\n<span class=\"line\"><span># MLX_DISABLE_COMPILE=1 python script.py</span></span></code></pre>\n<hr />\n<h2 id=\"metal-integration\">Metal Integration</h2>\n<p>MLX uses Metal, Apple’s GPU programming framework, as its backend. Most users never interact with Metal directly, but understanding the integration helps with debugging and optimization.</p>\n<h3 id=\"default-stream-behavior\">Default Stream Behavior</h3>\n<p>MLX operations run on a default GPU stream:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># These operations run on the GPU by default</span></span>\n<span class=\"line\"><span>a </span><span>=</span><span> mx</span><span>.</span><span>random</span><span>.</span><span>normal</span><span>((</span><span>1000</span><span>, </span><span>1000</span><span>))</span></span>\n<span class=\"line\"><span>b </span><span>=</span><span> mx</span><span>.</span><span>matmul</span><span>(a, a.T)</span></span></code></pre>\n<p>You can explicitly specify CPU execution:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Force CPU execution</span></span>\n<span class=\"line\"><span>c </span><span>=</span><span> mx</span><span>.</span><span>add</span><span>(a, b, stream</span><span>=</span><span>mx.cpu)</span></span></code></pre>\n<h3 id=\"custom-metal-kernels\">Custom Metal Kernels</h3>\n<p>For operations not covered by built-in primitives, MLX allows custom Metal kernels:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>source </span><span>=</span><span> \"\"\"</span></span>\n<span class=\"line\"><span>    uint elem = thread_position_in_grid.x;</span></span>\n<span class=\"line\"><span>    T tmp = inp[elem];</span></span>\n<span class=\"line\"><span>    out[elem] = tmp * tmp;</span></span>\n<span class=\"line\"><span>\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>kernel </span><span>=</span><span> mx</span><span>.</span><span>fast</span><span>.</span><span>metal_kernel</span><span>(</span></span>\n<span class=\"line\"><span>    name</span><span>=</span><span>\"square\"</span><span>,</span></span>\n<span class=\"line\"><span>    input_names</span><span>=</span><span>[</span><span>\"inp\"</span><span>],</span></span>\n<span class=\"line\"><span>    output_names</span><span>=</span><span>[</span><span>\"out\"</span><span>],</span></span>\n<span class=\"line\"><span>    source</span><span>=</span><span>source,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>a </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>1.0</span><span>, </span><span>2.0</span><span>, </span><span>3.0</span><span>, </span><span>4.0</span><span>])</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> kernel</span><span>(inputs</span><span>=</span><span>[a], grid</span><span>=</span><span>(</span><span>4</span><span>,), output_shapes</span><span>=</span><span>[(</span><span>4</span><span>,)], output_dtypes</span><span>=</span><span>[mx.float32])</span></span></code></pre>\n<h3 id=\"m5-neural-accelerator-support\">M5 Neural Accelerator Support</h3>\n<p>With macOS 26.2 and MLX’s latest versions, the framework can leverage the Neural Accelerators in the M5 chip. MLX uses the Tensor Operations (TensorOps) and Metal Performance Primitives frameworks introduced with Metal 4 to access dedicated matrix multiplication hardware.</p>\n<hr />\n<h2 id=\"mlx-lm-running-large-language-models\">MLX-LM: Running Large Language Models</h2>\n<p>MLX-LM is the official package for running and fine-tuning large language models on Apple Silicon. It provides a simple interface for loading models from Hugging Face, generating text, and fine-tuning with LoRA.</p>\n<h3 id=\"installation\">Installation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>pip</span><span> install</span><span> mlx-lm</span></span></code></pre>\n<h3 id=\"loading-and-generating-text\">Loading and Generating Text</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> mlx_lm </span><span>import</span><span> load</span><span>,</span><span> generate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load a model from Hugging Face</span></span>\n<span class=\"line\"><span>model</span><span>,</span><span> tokenizer </span><span>=</span><span> load</span><span>(</span><span>\"mlx-community/Mistral-7B-Instruct-v0.3-4bit\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate text</span></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> \"Explain quantum computing in simple terms:\"</span></span>\n<span class=\"line\"><span>response </span><span>=</span><span> generate</span><span>(model, tokenizer, prompt</span><span>=</span><span>prompt, max_tokens</span><span>=</span><span>200</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(response)</span></span></code></pre>\n<p>Command-line usage:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>mlx_lm.generate</span><span> --model</span><span> mlx-community/Qwen3-4B-Instruct-2507-4bit</span><span> --prompt</span><span> \"hello\"</span></span></code></pre>\n<h3 id=\"supported-model-architectures\">Supported Model Architectures</h3>\n<p>MLX-LM supports most popular LLM architectures including:</p>\n<ul>\n<li>LLaMA and LLaMA 2/3</li>\n<li>Mistral and Mixtral (MoE)</li>\n<li>Qwen 2, 2.5, 3 and Qwen3 MoE</li>\n<li>Phi 2, 3, 4</li>\n<li>Gemma 1, 2, 3</li>\n<li>OLMo and OLMoE</li>\n<li>MiniCPM</li>\n<li>DeepSeek</li>\n</ul>\n<p>Most models available on Hugging Face can be loaded directly if they follow standard architectures.</p>\n<h3 id=\"quantization\">Quantization</h3>\n<p>Quantization reduces memory usage and increases generation speed. MLX-LM supports 4-bit, 6-bit, and 8-bit quantization:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Convert and quantize a model</span></span>\n<span class=\"line\"><span>python</span><span> -m</span><span> mlx_lm.convert</span><span> \\</span></span>\n<span class=\"line\"><span>    --hf-path</span><span> mistralai/Mistral-7B-Instruct-v0.3</span><span> \\</span></span>\n<span class=\"line\"><span>    --q-bits</span><span> 4</span><span> \\</span></span>\n<span class=\"line\"><span>    --q-group-size</span><span> 64</span></span></code></pre>\n<p>Using a quantized model in Python:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> mlx_lm </span><span>import</span><span> load</span><span>,</span><span> generate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load a pre-quantized model</span></span>\n<span class=\"line\"><span>model</span><span>,</span><span> tokenizer </span><span>=</span><span> load</span><span>(</span><span>\"mlx-community/Mistral-7B-Instruct-v0.3-4bit\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate as normal</span></span>\n<span class=\"line\"><span>response </span><span>=</span><span> generate</span><span>(model, tokenizer, prompt</span><span>=</span><span>\"Hello!\"</span><span>, max_tokens</span><span>=</span><span>100</span><span>)</span></span></code></pre>\n<p>Quantized models from the mlx-community organization on Hugging Face are ready to use without conversion.</p>\n<h3 id=\"memory-requirements\">Memory Requirements</h3>\n<p>Approximate memory usage for different quantization levels on a 7B parameter model:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Precision</th><th>Memory Usage</th></tr></thead><tbody><tr><td>float32</td><td>~28 GB</td></tr><tr><td>float16</td><td>~14 GB</td></tr><tr><td>8-bit</td><td>~7 GB</td></tr><tr><td>4-bit</td><td>~3.5 GB</td></tr></tbody></table></div>\n<p>A 4-bit quantized Llama 3B can generate around 50 tokens/second on recent Apple Silicon.</p>\n<hr />\n<h2 id=\"lora-fine-tuning-with-mlx-lm\">LoRA Fine-Tuning with MLX-LM</h2>\n<p>Low-Rank Adaptation (LoRA) enables fine-tuning large models with limited memory by training small adapter matrices instead of full model weights. MLX-LM includes built-in support for LoRA and QLoRA (quantized LoRA).</p>\n<h3 id=\"preparing-training-data\">Preparing Training Data</h3>\n<p>Training data should be in JSONL format with a “text” field:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span><span>\"text\"</span><span>:</span><span> \"Question: What is the capital of France?\\nAnswer: Paris\"</span><span>}</span></span>\n<span class=\"line\"><span>{</span><span>\"text\"</span><span>:</span><span> \"Question: What is 2+2?\\nAnswer: 4\"</span><span>}</span></span></code></pre>\n<h3 id=\"running-lora-training\">Running LoRA Training</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>python</span><span> -m</span><span> mlx_lm.lora</span><span> \\</span></span>\n<span class=\"line\"><span>    --model</span><span> mlx-community/gemma-3-4b-it-4bit</span><span> \\</span></span>\n<span class=\"line\"><span>    --train</span><span> \\</span></span>\n<span class=\"line\"><span>    --data</span><span> ./data</span><span> \\</span></span>\n<span class=\"line\"><span>    --adapter-path</span><span> ./lora_adapters</span><span> \\</span></span>\n<span class=\"line\"><span>    --lora-r</span><span> 16</span><span> \\</span></span>\n<span class=\"line\"><span>    --lora-alpha</span><span> 32</span><span> \\</span></span>\n<span class=\"line\"><span>    --iters</span><span> 500</span><span> \\</span></span>\n<span class=\"line\"><span>    --batch-size</span><span> 1</span><span> \\</span></span>\n<span class=\"line\"><span>    --learning-rate</span><span> 2e-4</span></span></code></pre>\n<p>Key parameters:</p>\n<ul>\n<li><code>--lora-r</code>: Rank of the LoRA matrices (higher = more capacity, more memory)</li>\n<li><code>--lora-alpha</code>: Scaling factor for LoRA updates</li>\n<li><code>--lora-layers</code>: Number of layers to apply LoRA (default 16)</li>\n<li><code>--batch-size</code>: Training batch size (reduce for memory)</li>\n<li><code>--grad-checkpoint</code>: Enable gradient checkpointing for memory savings</li>\n</ul>\n<h3 id=\"qlora-for-memory-efficiency\">QLoRA for Memory Efficiency</h3>\n<p>When using a 4-bit quantized base model, training automatically uses QLoRA:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Using a 4-bit model enables QLoRA automatically</span></span>\n<span class=\"line\"><span>python</span><span> -m</span><span> mlx_lm.lora</span><span> \\</span></span>\n<span class=\"line\"><span>    --model</span><span> mlx-community/Llama-3-8B-Instruct-4bit</span><span> \\</span></span>\n<span class=\"line\"><span>    --train</span><span> \\</span></span>\n<span class=\"line\"><span>    --data</span><span> ./data</span><span> \\</span></span>\n<span class=\"line\"><span>    --adapter-path</span><span> ./qlora_adapters</span></span></code></pre>\n<p>Memory comparison for Llama 7B:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Memory Usage</th></tr></thead><tbody><tr><td>Full fine-tuning</td><td>~28 GB</td></tr><tr><td>LoRA (r=8)</td><td>~14 GB</td></tr><tr><td>QLoRA (4-bit + LoRA)</td><td>~7 GB</td></tr></tbody></table></div>\n<h3 id=\"using-trained-adapters\">Using Trained Adapters</h3>\n<p>Generate text with your trained adapter:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> mlx_lm </span><span>import</span><span> load</span><span>,</span><span> generate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load base model with adapter</span></span>\n<span class=\"line\"><span>model</span><span>,</span><span> tokenizer </span><span>=</span><span> load</span><span>(</span></span>\n<span class=\"line\"><span>    \"mlx-community/gemma-3-4b-it-4bit\"</span><span>,</span></span>\n<span class=\"line\"><span>    adapter_path</span><span>=</span><span>\"./lora_adapters\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>response </span><span>=</span><span> generate</span><span>(model, tokenizer, prompt</span><span>=</span><span>\"Your prompt here\"</span><span>)</span></span></code></pre>\n<h3 id=\"merging-adapters\">Merging Adapters</h3>\n<p>For deployment, merge adapters into the base model:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>python</span><span> -m</span><span> mlx_lm.fuse</span><span> \\</span></span>\n<span class=\"line\"><span>    --model</span><span> mlx-community/gemma-3-4b-it-4bit</span><span> \\</span></span>\n<span class=\"line\"><span>    --adapter-path</span><span> ./lora_adapters</span><span> \\</span></span>\n<span class=\"line\"><span>    --save-path</span><span> ./merged_model</span><span> \\</span></span>\n<span class=\"line\"><span>    --de-quantize</span><span>  # Optional: convert back to float16</span></span></code></pre>\n<hr />\n<h2 id=\"openai-compatible-api-server\">OpenAI-Compatible API Server</h2>\n<p>MLX-LM includes a built-in server that exposes an OpenAI-compatible API. This allows any tool designed for OpenAI’s API to work with local models.</p>\n<h3 id=\"starting-the-server\">Starting the Server</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>mlx_lm.server</span><span> --model</span><span> mlx-community/Mistral-7B-Instruct-v0.3-4bit</span></span></code></pre>\n<h3 id=\"using-the-api\">Using the API</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> openai </span><span>import</span><span> OpenAI</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>client </span><span>=</span><span> OpenAI</span><span>(base_url</span><span>=</span><span>\"http://localhost:8080/v1\"</span><span>, api_key</span><span>=</span><span>\"not-needed\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>response </span><span>=</span><span> client</span><span>.</span><span>chat</span><span>.</span><span>completions</span><span>.</span><span>create</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"mlx-community/Mistral-7B-Instruct-v0.3-4bit\"</span><span>,</span></span>\n<span class=\"line\"><span>    messages</span><span>=</span><span>[</span></span>\n<span class=\"line\"><span>        {</span><span>\"role\"</span><span>: </span><span>\"user\"</span><span>, </span><span>\"content\"</span><span>: </span><span>\"What is machine learning?\"</span><span>}</span></span>\n<span class=\"line\"><span>    ]</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>print</span><span>(response.choices[</span><span>0</span><span>].message.content)</span></span></code></pre>\n<h3 id=\"alternative-servers\">Alternative Servers</h3>\n<p>For production use cases, consider these community servers:</p>\n<p><strong>FastMLX</strong>: High-performance server with support for vision-language models and efficient resource management. Available at <a href=\"https://github.com/Blaizzy/fastmlx\">Blaizzy/fastmlx</a>.</p>\n<p><strong>vllm-mlx</strong>: OpenAI-compatible server with continuous batching, MCP tool calling, and multimodal support. Achieves 400+ tokens/second. Available at <a href=\"https://github.com/waybarrios/vllm-mlx\">GitHub</a>.</p>\n<p><strong>mlx-openai-server</strong>: FastAPI-based server supporting LLMs, VLMs (via mlx-vlm), image generation (via mflux), and Whisper transcription. Available on <a href=\"https://pypi.org/project/mlx-openai-server/\">PyPI</a>.</p>\n<hr />\n<h2 id=\"the-mlx-ecosystem\">The MLX Ecosystem</h2>\n<p>Beyond MLX core and MLX-LM, a rich ecosystem of packages has developed for various machine learning tasks.</p>\n<h3 id=\"mlx-community-on-hugging-face\">mlx-community on Hugging Face</h3>\n<p>The <a href=\"https://huggingface.co/mlx-community\">mlx-community</a> organization on Hugging Face hosts over 1,000 models converted to MLX format. These include:</p>\n<ul>\n<li>LLaMA 3.3 and 3.2 variants</li>\n<li>Qwen 2.5, Qwen3, QwQ</li>\n<li>Gemma 2 and 3</li>\n<li>Mistral and Mixtral</li>\n<li>Phi-3 and Phi-4</li>\n<li>Whisper models for speech recognition</li>\n</ul>\n<p>Models are available in various quantization levels (4-bit, 8-bit, fp16) and can be loaded directly with MLX-LM.</p>\n<h3 id=\"mlx-examples-repository\">mlx-examples Repository</h3>\n<p>The <a href=\"https://github.com/ml-explore/mlx-examples\">ml-explore/mlx-examples</a> repository contains standalone examples including:</p>\n<p><strong>Language Models</strong>: Transformer training, LLaMA/Mistral text generation, Mixtral 8x7B (MoE), LoRA and QLoRA fine-tuning.</p>\n<p><strong>Image Generation</strong>: Stable Diffusion and SDXL with text-to-image and image-to-image generation. Supports quantization for memory-constrained systems.</p>\n<p><strong>Audio</strong>: OpenAI Whisper for speech recognition, Meta’s EnCodec for audio compression, MusicGen for music generation.</p>\n<p><strong>Vision-Language</strong>: CLIP for joint text-image embeddings, LLaVA for text generation from images.</p>\n<h3 id=\"mlx-vlm\">mlx-vlm</h3>\n<p><a href=\"https://github.com/Blaizzy/mlx-vlm\">MLX-VLM</a> provides inference and fine-tuning for vision-language models:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> mlx_vlm </span><span>import</span><span> load</span><span>,</span><span> generate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model</span><span>,</span><span> processor </span><span>=</span><span> load</span><span>(</span><span>\"mlx-community/llava-1.5-7b-4bit\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>output </span><span>=</span><span> generate</span><span>(model, processor, </span><span>\"path/to/image.jpg\"</span><span>, </span><span>\"What is in this image?\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(output)</span></span></code></pre>\n<p>Supported features:</p>\n<ul>\n<li>Multi-image analysis</li>\n<li>Video captioning and summarization</li>\n<li>LoRA and QLoRA fine-tuning for VLMs</li>\n<li>Audio and video support with Gemma 3n</li>\n</ul>\n<h3 id=\"whisper-mlx\">whisper-mlx</h3>\n<p><a href=\"https://pypi.org/project/mlx-whisper/\">MLX Whisper</a> provides speech recognition using OpenAI’s Whisper models:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx_whisper</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>result </span><span>=</span><span> mlx_whisper</span><span>.</span><span>transcribe</span><span>(</span><span>\"audio.mp3\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(result[</span><span>\"text\"</span><span>])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># With word-level timestamps</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> mlx_whisper</span><span>.</span><span>transcribe</span><span>(</span><span>\"audio.mp3\"</span><span>, word_timestamps</span><span>=</span><span>True</span><span>)</span></span></code></pre>\n<p>Pre-converted models are available from the Hugging Face MLX Community. For faster transcription, <a href=\"https://github.com/mustafaaljadery/lightning-whisper-mlx\">lightning-whisper-mlx</a> provides optimized implementations.</p>\n<h3 id=\"stable-diffusion\">Stable Diffusion</h3>\n<p>The mlx-examples repository includes Stable Diffusion implementations:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Text to image</span></span>\n<span class=\"line\"><span>python</span><span> txt2image.py</span><span> \"A photo of an astronaut riding a horse on Mars\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># With quantization for 8GB Macs</span></span>\n<span class=\"line\"><span>python</span><span> txt2image.py</span><span> --quantize</span><span> \"A sunset over mountains\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Image to image</span></span>\n<span class=\"line\"><span>python</span><span> image2image.py</span><span> --image</span><span> input.png</span><span> --strength</span><span> 0.7</span><span> \"Make it look like a painting\"</span></span></code></pre>\n<p>Both SD 1.5 and SDXL are supported. Quantization (4-bit text encoder, 8-bit UNet) enables generation on 8GB Macs without swapping.</p>\n<hr />\n<h2 id=\"moving-from-cuda-to-mlx-a-migration-guide\">Moving from CUDA to MLX: A Migration Guide</h2>\n<p>If you are coming from PyTorch with CUDA, this section covers the key differences and provides practical guidance for porting code.</p>\n<h3 id=\"prerequisites\">Prerequisites</h3>\n<p><strong>Hardware</strong>: Any Mac with Apple Silicon (M1, M2, M3, M4, M5 series).</p>\n<p><strong>Software</strong>:</p>\n<ul>\n<li>macOS 14.0 or later (macOS 26.2 for M5 Neural Accelerator support)</li>\n<li>Python 3.10 or later</li>\n<li>Native ARM Python (not Rosetta x86 emulation)</li>\n</ul>\n<p>Verify your Python installation:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>python</span><span> -c</span><span> \"import platform; print(platform.processor())\"</span></span>\n<span class=\"line\"><span># Should print \"arm\", not \"i386\"</span></span></code></pre>\n<h3 id=\"installation-1\">Installation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Install MLX core</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> mlx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install MLX-LM for language models</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> mlx-lm</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install additional packages as needed</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> mlx-whisper</span><span>  # For speech recognition</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> mlx-vlm</span><span>      # For vision-language models</span></span></code></pre>\n<h3 id=\"key-api-differences\">Key API Differences</h3>\n<p><strong>No Device Management</strong>: In PyTorch, you explicitly move tensors between devices. In MLX, arrays live in unified memory and you specify the device when calling operations:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># PyTorch</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> torch</span><span>.</span><span>tensor</span><span>([</span><span>1</span><span>, </span><span>2</span><span>, </span><span>3</span><span>]).</span><span>cuda</span><span>()</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> x </span><span>+</span><span> 1</span><span>  # Runs on CUDA</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># MLX</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>1</span><span>, </span><span>2</span><span>, </span><span>3</span><span>])</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> x </span><span>+</span><span> 1</span><span>  # Runs on GPU by default</span></span>\n<span class=\"line\"><span>z </span><span>=</span><span> mx</span><span>.</span><span>add</span><span>(x, y, stream</span><span>=</span><span>mx.cpu)</span><span>  # Force CPU</span></span></code></pre>\n<p><strong>Lazy Evaluation</strong>: Operations are not executed immediately:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># PyTorch - computed immediately</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> x </span><span>+</span><span> 1</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># MLX - creates computation graph</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> x </span><span>+</span><span> 1</span><span>  # Not computed yet</span></span>\n<span class=\"line\"><span>mx</span><span>.</span><span>eval</span><span>(y)</span><span>  # Now computed</span></span>\n<span class=\"line\"><span>print</span><span>(y)</span><span>   # Also triggers computation</span></span></code></pre>\n<p><strong>Function-Based Gradients</strong>: No backward() method or gradient tape:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># PyTorch</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> torch</span><span>.</span><span>tensor</span><span>([</span><span>1.0</span><span>], requires_grad</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> x </span><span>**</span><span> 2</span></span>\n<span class=\"line\"><span>y</span><span>.</span><span>backward</span><span>()</span></span>\n<span class=\"line\"><span>print</span><span>(x.grad)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># MLX</span></span>\n<span class=\"line\"><span>def</span><span> f</span><span>(</span><span>x</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> x </span><span>**</span><span> 2</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>grad_f </span><span>=</span><span> mx</span><span>.</span><span>grad</span><span>(f)</span></span>\n<span class=\"line\"><span>x </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>1.0</span><span>])</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>grad_f</span><span>(x))</span></span></code></pre>\n<h3 id=\"porting-neural-network-code\">Porting Neural Network Code</h3>\n<p>The <code>mlx.nn</code> module closely follows PyTorch’s API:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># PyTorch</span></span>\n<span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>nn </span><span>as</span><span> nn</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> PyTorchMLP</span><span>(</span><span>nn</span><span>.</span><span>Module</span><span>):</span></span>\n<span class=\"line\"><span>    def</span><span> __init__</span><span>(</span><span>self</span><span>,</span><span> in_dim</span><span>,</span><span> hidden_dim</span><span>,</span><span> out_dim</span><span>):</span></span>\n<span class=\"line\"><span>        super</span><span>().</span><span>__init__</span><span>()</span></span>\n<span class=\"line\"><span>        self</span><span>.</span><span>fc1 </span><span>=</span><span> nn</span><span>.</span><span>Linear</span><span>(in_dim, hidden_dim)</span></span>\n<span class=\"line\"><span>        self</span><span>.</span><span>fc2 </span><span>=</span><span> nn</span><span>.</span><span>Linear</span><span>(hidden_dim, out_dim)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> forward</span><span>(</span><span>self</span><span>,</span><span> x</span><span>):</span></span>\n<span class=\"line\"><span>        x </span><span>=</span><span> torch</span><span>.</span><span>relu</span><span>(self.</span><span>fc1</span><span>(x))</span></span>\n<span class=\"line\"><span>        return</span><span> self</span><span>.</span><span>fc2</span><span>(x)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># MLX</span></span>\n<span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>nn </span><span>as</span><span> nn</span></span>\n<span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> MLXMLP</span><span>(</span><span>nn</span><span>.</span><span>Module</span><span>):</span></span>\n<span class=\"line\"><span>    def</span><span> __init__</span><span>(</span><span>self</span><span>,</span><span> in_dim</span><span>,</span><span> hidden_dim</span><span>,</span><span> out_dim</span><span>):</span></span>\n<span class=\"line\"><span>        super</span><span>().</span><span>__init__</span><span>()</span></span>\n<span class=\"line\"><span>        self</span><span>.</span><span>fc1 </span><span>=</span><span> nn</span><span>.</span><span>Linear</span><span>(in_dim, hidden_dim)</span></span>\n<span class=\"line\"><span>        self</span><span>.</span><span>fc2 </span><span>=</span><span> nn</span><span>.</span><span>Linear</span><span>(hidden_dim, out_dim)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> __call__</span><span>(</span><span>self</span><span>,</span><span> x</span><span>):</span></span>\n<span class=\"line\"><span>        x </span><span>=</span><span> mx</span><span>.</span><span>maximum</span><span>(self.</span><span>fc1</span><span>(x), </span><span>0</span><span>)</span><span>  # ReLU</span></span>\n<span class=\"line\"><span>        return</span><span> self</span><span>.</span><span>fc2</span><span>(x)</span></span></code></pre>\n<p>Key differences:</p>\n<ul>\n<li>Use <code>__call__</code> instead of <code>forward</code></li>\n<li>Use <code>mx.maximum(x, 0)</code> instead of <code>torch.relu(x)</code> (or <code>nn.relu(x)</code>)</li>\n<li>No <code>.cuda()</code> calls needed</li>\n</ul>\n<h3 id=\"converting-model-weights\">Converting Model Weights</h3>\n<p>When porting a trained PyTorch model, you need to convert the weights:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"><span>import</span><span> numpy </span><span>as</span><span> np</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load PyTorch weights</span></span>\n<span class=\"line\"><span>pytorch_weights </span><span>=</span><span> torch</span><span>.</span><span>load</span><span>(</span><span>\"model.pt\"</span><span>, map_location</span><span>=</span><span>\"cpu\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Convert to MLX format</span></span>\n<span class=\"line\"><span>mlx_weights </span><span>=</span><span> {}</span></span>\n<span class=\"line\"><span>for</span><span> key</span><span>,</span><span> value </span><span>in</span><span> pytorch_weights</span><span>.</span><span>items</span><span>():</span></span>\n<span class=\"line\"><span>    # Convert to numpy, then to MLX array</span></span>\n<span class=\"line\"><span>    np_array </span><span>=</span><span> value</span><span>.</span><span>numpy</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Handle weight format differences (NCHW -&gt; NHWC for convolutions)</span></span>\n<span class=\"line\"><span>    if</span><span> \"conv\"</span><span> in</span><span> key </span><span>and</span><span> \"weight\"</span><span> in</span><span> key</span><span>:</span></span>\n<span class=\"line\"><span>        np_array </span><span>=</span><span> np</span><span>.</span><span>transpose</span><span>(np_array, (</span><span>0</span><span>, </span><span>2</span><span>, </span><span>3</span><span>, </span><span>1</span><span>))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    mlx_weights</span><span>[</span><span>key</span><span>]</span><span> =</span><span> mx</span><span>.</span><span>array</span><span>(np_array)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Save in MLX format</span></span>\n<span class=\"line\"><span>mx</span><span>.</span><span>savez</span><span>(</span><span>\"model.npz\"</span><span>, </span><span>**</span><span>mlx_weights)</span></span></code></pre>\n<p>For LLMs, the mlx_lm.convert script handles this automatically:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>python</span><span> -m</span><span> mlx_lm.convert</span><span> --hf-path</span><span> meta-llama/Llama-3-8B</span></span></code></pre>\n<h3 id=\"training-loop-pattern\">Training Loop Pattern</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>core </span><span>as</span><span> mx</span></span>\n<span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>nn </span><span>as</span><span> nn</span></span>\n<span class=\"line\"><span>import</span><span> mlx</span><span>.</span><span>optimizers </span><span>as</span><span> optim</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> MLXMLP</span><span>(</span><span>784</span><span>, </span><span>256</span><span>, </span><span>10</span><span>)</span></span>\n<span class=\"line\"><span>optimizer </span><span>=</span><span> optim</span><span>.</span><span>Adam</span><span>(learning_rate</span><span>=</span><span>1e-3</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> loss_fn</span><span>(</span><span>model</span><span>,</span><span> x</span><span>,</span><span> y</span><span>):</span></span>\n<span class=\"line\"><span>    logits </span><span>=</span><span> model</span><span>(x)</span></span>\n<span class=\"line\"><span>    return</span><span> mx</span><span>.</span><span>mean</span><span>(nn.losses.</span><span>cross_entropy</span><span>(logits, y))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>loss_and_grad_fn </span><span>=</span><span> nn</span><span>.</span><span>value_and_grad</span><span>(model, loss_fn)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> epoch </span><span>in</span><span> range</span><span>(num_epochs):</span></span>\n<span class=\"line\"><span>    for</span><span> x_batch</span><span>,</span><span> y_batch </span><span>in</span><span> dataloader</span><span>:</span></span>\n<span class=\"line\"><span>        loss</span><span>,</span><span> grads </span><span>=</span><span> loss_and_grad_fn</span><span>(model, x_batch, y_batch)</span></span>\n<span class=\"line\"><span>        optimizer</span><span>.</span><span>update</span><span>(model, grads)</span></span>\n<span class=\"line\"><span>        mx</span><span>.</span><span>eval</span><span>(model.</span><span>parameters</span><span>(), optimizer.state)</span></span></code></pre>\n<h3 id=\"common-patterns\">Common Patterns</h3>\n<p><strong>Model Evaluation Mode</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># PyTorch</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>eval</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># MLX</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>eval</span><span>()</span><span>  # Same API</span></span></code></pre>\n<p><strong>Saving and Loading</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Save model weights</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>save_weights</span><span>(</span><span>\"model.safetensors\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model weights</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>load_weights</span><span>(</span><span>\"model.safetensors\"</span><span>)</span></span></code></pre>\n<p><strong>Mixed Precision</strong>: MLX handles dtypes explicitly:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Convert model to float16</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> model</span><span>.</span><span>apply</span><span>(</span><span>lambda</span><span> x</span><span>: x.</span><span>astype</span><span>(mx.float16) </span><span>if</span><span> x.dtype </span><span>==</span><span> mx.float32 </span><span>else</span><span> x)</span></span></code></pre>\n<hr />\n<h2 id=\"performance-optimization\">Performance Optimization</h2>\n<h3 id=\"use-compilation\">Use Compilation</h3>\n<p>Wrap hot paths in <code>mx.compile()</code>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>@mx</span><span>.</span><span>compile</span></span>\n<span class=\"line\"><span>def</span><span> training_step</span><span>(</span><span>model</span><span>,</span><span> optimizer</span><span>,</span><span> x</span><span>,</span><span> y</span><span>):</span></span>\n<span class=\"line\"><span>    loss</span><span>,</span><span> grads </span><span>=</span><span> loss_and_grad_fn</span><span>(model, x, y)</span></span>\n<span class=\"line\"><span>    optimizer</span><span>.</span><span>update</span><span>(model, grads)</span></span>\n<span class=\"line\"><span>    return</span><span> loss</span></span></code></pre>\n<h3 id=\"batch-operations\">Batch Operations</h3>\n<p>MLX has lower overhead for small operations than CUDA, but batching still helps:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Less efficient</span></span>\n<span class=\"line\"><span>for</span><span> x </span><span>in</span><span> data</span><span>:</span></span>\n<span class=\"line\"><span>    y </span><span>=</span><span> model</span><span>(x)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># More efficient</span></span>\n<span class=\"line\"><span>y </span><span>=</span><span> model</span><span>(mx.</span><span>stack</span><span>(data))</span></span></code></pre>\n<h3 id=\"quantization-for-inference\">Quantization for Inference</h3>\n<p>Use 4-bit or 8-bit models for inference when possible:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> mlx_lm </span><span>import</span><span> load</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># 4-bit model uses 75% less memory and runs faster</span></span>\n<span class=\"line\"><span>model</span><span>,</span><span> tokenizer </span><span>=</span><span> load</span><span>(</span><span>\"mlx-community/Llama-3-8B-Instruct-4bit\"</span><span>)</span></span></code></pre>\n<h3 id=\"memory-management\">Memory Management</h3>\n<p>Check memory usage:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>info </span><span>=</span><span> mx</span><span>.</span><span>metal</span><span>.</span><span>get_memory_info</span><span>()</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Allocated: </span><span>{</span><span>info[</span><span>'allocated'</span><span>] </span><span>/</span><span> 1e9</span><span>:.2f</span><span>}</span><span> GB\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Peak: </span><span>{</span><span>info[</span><span>'peak'</span><span>] </span><span>/</span><span> 1e9</span><span>:.2f</span><span>}</span><span> GB\"</span><span>)</span></span></code></pre>\n<p>Clear cached memory:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>mx</span><span>.</span><span>metal</span><span>.</span><span>clear_cache</span><span>()</span></span></code></pre>\n<h3 id=\"profiling\">Profiling</h3>\n<p>Build MLX with Metal debugging for GPU profiling:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>mkdir</span><span> build</span><span> &amp;&amp;</span><span> cd</span><span> build</span></span>\n<span class=\"line\"><span>cmake</span><span> ..</span><span> -DMLX_METAL_DEBUG=ON</span></span>\n<span class=\"line\"><span>make</span><span> -j</span></span></code></pre>\n<p>Then use Xcode’s GPU profiler to analyze kernel execution.</p>\n<hr />\n<h2 id=\"debugging-tips\">Debugging Tips</h2>\n<h3 id=\"disable-compilation-for-debugging\">Disable Compilation for Debugging</h3>\n<p>Compiled functions cannot contain print statements. Disable compilation to debug:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>mx</span><span>.</span><span>disable_compile</span><span>()</span></span>\n<span class=\"line\"><span># or</span></span>\n<span class=\"line\"><span># MLX_DISABLE_COMPILE=1 python script.py</span></span></code></pre>\n<h3 id=\"force-evaluation\">Force Evaluation</h3>\n<p>When debugging lazy evaluation issues, force computation:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>result </span><span>=</span><span> complex_operation</span><span>(x)</span></span>\n<span class=\"line\"><span>mx</span><span>.</span><span>eval</span><span>(result)</span><span>  # Force computation</span></span>\n<span class=\"line\"><span>print</span><span>(result)</span><span>    # Now safe to inspect</span></span></code></pre>\n<h3 id=\"check-array-properties\">Check Array Properties</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>x </span><span>=</span><span> mx</span><span>.</span><span>array</span><span>([</span><span>1.0</span><span>, </span><span>2.0</span><span>, </span><span>3.0</span><span>])</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Shape: </span><span>{</span><span>x.shape</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Dtype: </span><span>{</span><span>x.dtype</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Size: </span><span>{</span><span>x.size</span><span>}</span><span>\"</span><span>)</span></span></code></pre>\n<h3 id=\"memory-issues\">Memory Issues</h3>\n<p>If you encounter memory errors during training:</p>\n<ol>\n<li>Reduce batch size</li>\n<li>Enable gradient checkpointing</li>\n<li>Use quantized models</li>\n<li>Reduce sequence length</li>\n<li>Reduce LoRA rank</li>\n</ol>\n<hr />\n<h2 id=\"practical-example-end-to-end-llm-inference\">Practical Example: End-to-End LLM Inference</h2>\n<p>Here is a complete example for running local LLM inference:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> mlx_lm </span><span>import</span><span> load</span><span>,</span><span> generate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> main</span><span>():</span></span>\n<span class=\"line\"><span>    # Load a quantized model</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"Loading model...\"</span><span>)</span></span>\n<span class=\"line\"><span>    model</span><span>,</span><span> tokenizer </span><span>=</span><span> load</span><span>(</span><span>\"mlx-community/Qwen2.5-7B-Instruct-4bit\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # System prompt</span></span>\n<span class=\"line\"><span>    messages </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>        {</span><span>\"role\"</span><span>:</span><span> \"system\"</span><span>,</span><span> \"content\"</span><span>:</span><span> \"You are a helpful assistant.\"</span><span>},</span></span>\n<span class=\"line\"><span>        {</span><span>\"role\"</span><span>:</span><span> \"user\"</span><span>,</span><span> \"content\"</span><span>:</span><span> \"Explain how transformers work in 3 sentences.\"</span><span>}</span></span>\n<span class=\"line\"><span>    ]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Format with chat template</span></span>\n<span class=\"line\"><span>    prompt </span><span>=</span><span> tokenizer</span><span>.</span><span>apply_chat_template</span><span>(</span></span>\n<span class=\"line\"><span>        messages,</span></span>\n<span class=\"line\"><span>        tokenize</span><span>=</span><span>False</span><span>,</span></span>\n<span class=\"line\"><span>        add_generation_prompt</span><span>=</span><span>True</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Generate</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"Generating...\"</span><span>)</span></span>\n<span class=\"line\"><span>    response </span><span>=</span><span> generate</span><span>(</span></span>\n<span class=\"line\"><span>        model,</span></span>\n<span class=\"line\"><span>        tokenizer,</span></span>\n<span class=\"line\"><span>        prompt</span><span>=</span><span>prompt,</span></span>\n<span class=\"line\"><span>        max_tokens</span><span>=</span><span>200</span><span>,</span></span>\n<span class=\"line\"><span>        temp</span><span>=</span><span>0.7</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(response)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>if</span><span> __name__</span><span> ==</span><span> \"__main__\"</span><span>:</span></span>\n<span class=\"line\"><span>    main</span><span>()</span></span></code></pre>\n<hr />\n<h2 id=\"practical-example-fine-tuning-a-model\">Practical Example: Fine-Tuning a Model</h2>\n<p>Complete example for fine-tuning with LoRA:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> json</span></span>\n<span class=\"line\"><span>import</span><span> subprocess</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Step 1: Prepare training data</span></span>\n<span class=\"line\"><span>train_data </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>    {</span><span>\"text\"</span><span>:</span><span> \"### Instruction: Translate to French\\n### Input: Hello\\n### Response: Bonjour\"</span><span>},</span></span>\n<span class=\"line\"><span>    {</span><span>\"text\"</span><span>:</span><span> \"### Instruction: Translate to French\\n### Input: Goodbye\\n### Response: Au revoir\"</span><span>},</span></span>\n<span class=\"line\"><span>    # Add more examples...</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Save training data</span></span>\n<span class=\"line\"><span>with</span><span> open</span><span>(</span><span>\"train.jsonl\"</span><span>, </span><span>\"w\"</span><span>)</span><span> as</span><span> f</span><span>:</span></span>\n<span class=\"line\"><span>    for</span><span> item </span><span>in</span><span> train_data</span><span>:</span></span>\n<span class=\"line\"><span>        f</span><span>.</span><span>write</span><span>(json.</span><span>dumps</span><span>(item) </span><span>+</span><span> \"\\n\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Step 2: Run training</span></span>\n<span class=\"line\"><span>cmd </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>    \"python\"</span><span>,</span><span> \"-m\"</span><span>,</span><span> \"mlx_lm.lora\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"--model\"</span><span>,</span><span> \"mlx-community/Mistral-7B-Instruct-v0.3-4bit\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"--train\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"--data\"</span><span>,</span><span> \".\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"--adapter-path\"</span><span>,</span><span> \"./adapters\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"--lora-r\"</span><span>,</span><span> \"8\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"--iters\"</span><span>,</span><span> \"100\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"--batch-size\"</span><span>,</span><span> \"1\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"--learning-rate\"</span><span>,</span><span> \"1e-4\"</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>subprocess</span><span>.</span><span>run</span><span>(cmd)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Step 3: Use the fine-tuned model</span></span>\n<span class=\"line\"><span>from</span><span> mlx_lm </span><span>import</span><span> load</span><span>,</span><span> generate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model</span><span>,</span><span> tokenizer </span><span>=</span><span> load</span><span>(</span></span>\n<span class=\"line\"><span>    \"mlx-community/Mistral-7B-Instruct-v0.3-4bit\"</span><span>,</span></span>\n<span class=\"line\"><span>    adapter_path</span><span>=</span><span>\"./adapters\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> \"### Instruction: Translate to French\\n### Input: Thank you\\n### Response:\"</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>generate</span><span>(model, tokenizer, prompt</span><span>=</span><span>prompt, max_tokens</span><span>=</span><span>50</span><span>))</span></span></code></pre>\n<hr />\n<h2 id=\"mlx-vs-pytorch-performance-summary\">MLX vs PyTorch Performance Summary</h2>\n<p>Based on benchmarks from late 2025:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Aspect</th><th>MLX</th><th>PyTorch + CUDA</th></tr></thead><tbody><tr><td>Setup complexity</td><td>Simple pip install</td><td>CUDA toolkit, drivers</td></tr><tr><td>Memory efficiency</td><td>Unified, no copies</td><td>Explicit transfers</td></tr><tr><td>Small batch latency</td><td>Lower overhead</td><td>Higher kernel launch cost</td></tr><tr><td>Large batch throughput</td><td>Competitive on M2 Ultra+</td><td>Faster on RTX 4090</td></tr><tr><td>LLM inference</td><td>Optimized for local use</td><td>More optimized on server GPUs</td></tr><tr><td>Training speed</td><td>Slower for large models</td><td>Faster for production training</td></tr><tr><td>Debugging</td><td>Straightforward</td><td>Mature tooling</td></tr></tbody></table></div>\n<p>MLX is well suited for:</p>\n<ul>\n<li>Local LLM inference and experimentation</li>\n<li>Development and prototyping on Mac</li>\n<li>Memory-constrained scenarios (unified 128GB+)</li>\n<li>Deploying models to Apple devices</li>\n</ul>\n<p>PyTorch + CUDA remains better for:</p>\n<ul>\n<li>Large-scale training runs</li>\n<li>Maximum throughput requirements</li>\n<li>Production server deployment</li>\n</ul>\n<hr />\n<h2 id=\"references\">References</h2>\n<ul>\n<li><a href=\"https://github.com/ml-explore/mlx\">MLX GitHub Repository</a> - Official source code and documentation</li>\n<li><a href=\"https://ml-explore.github.io/mlx/build/html/index.html\">MLX Documentation</a> - Official docs including API reference</li>\n<li><a href=\"https://github.com/ml-explore/mlx-lm\">MLX-LM Repository</a> - LLM inference and fine-tuning package</li>\n<li><a href=\"https://huggingface.co/mlx-community\">mlx-community on Hugging Face</a> - Pre-converted models</li>\n<li><a href=\"https://github.com/ml-explore/mlx-examples\">MLX Examples Repository</a> - Stable Diffusion, Whisper, and more</li>\n<li><a href=\"https://machinelearning.apple.com/research/exploring-llms-mlx-m5\">Apple MLX Research Page</a> - M5 Neural Accelerator details</li>\n<li><a href=\"https://developer.apple.com/videos/play/wwdc2025/315/\">WWDC 2025: Get started with MLX</a> - Apple’s official introduction</li>\n<li><a href=\"https://developer.apple.com/videos/play/wwdc2025/298/\">WWDC 2025: Explore LLMs with MLX</a> - LLM-focused session</li>\n<li><a href=\"https://github.com/Blaizzy/mlx-vlm\">MLX-VLM Repository</a> - Vision-language model support</li>\n<li><a href=\"https://pypi.org/project/mlx-whisper/\">MLX Whisper on PyPI</a> - Speech recognition package</li>\n<li><a href=\"https://towardsdatascience.com/how-fast-is-mlx-a-comprehensive-benchmark-on-8-apple-silicon-chips-and-4-cuda-gpus-378a0ae356a0/\">MLX vs CUDA Benchmark</a> - Performance comparisons</li>\n<li><a href=\"https://towardsdatascience.com/pytorch-and-mlx-for-apple-silicon-4f35b9f60e39/\">PyTorch and MLX Comparison</a> - Migration guide</li>\n<li><a href=\"https://ml-explore.github.io/mlx/build/html/usage/lazy_evaluation.html\">Lazy Evaluation Documentation</a> - Understanding lazy computation</li>\n<li><a href=\"https://ml-explore.github.io/mlx/build/html/usage/unified_memory.html\">Unified Memory Documentation</a> - Memory model details</li>\n<li><a href=\"https://ml-explore.github.io/mlx/build/html/usage/compile.html\">Compilation Documentation</a> - JIT compilation guide</li>\n<li><a href=\"https://github.com/ml-explore/mlx-examples/blob/main/lora/README.md\">LoRA Fine-Tuning Example</a> - Official fine-tuning docs</li>\n<li><a href=\"https://github.com/waybarrios/vllm-mlx\">vllm-mlx Server</a> - High-performance MLX server</li>\n<li><a href=\"https://blaizzy.github.io/fastmlx/\">FastMLX Documentation</a> - Production-ready API server</li>\n</ul>",
            "url": "https://blog.ecitis.org/apple-mlx-guide/",
            "title": "Apple MLX: A Complete Guide to Machine Learning on Apple Silicon",
            "summary": "Learn how MLX uses unified memory to make model training, fine-tuning, and local inference feel native on Apple Silicon.",
            "image": "https://blog.ecitis.org/open-graph/apple-mlx-guide.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Ecosystems & Tooling",
                "MLX",
                "Apple Silicon",
                "Local AI"
            ]
        },
        {
            "id": "https://blog.ecitis.org/coding-agents-explained/",
            "content_html": "<p>Coding agents are not code completion tools. The distinction matters. GitHub Copilot and similar autocomplete systems predict the next few tokens based on cursor context. Coding agents operate differently: they observe your codebase, reason about a task, execute actions through tools, and iterate until the task is complete. This pattern, called the agentic loop, is what separates a tool that suggests code from one that can implement features, fix bugs, and open pull requests autonomously.</p>\n<p>As of January 2026, coding agents have crossed a meaningful threshold. Claude Opus 4.5 became the first model to exceed 80% on SWE-bench Verified <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-17\" id=\"user-content-fnref-17\">1</a></sup>. GPT-5.2 Codex established state-of-the-art on SWE-bench Pro at 56.4% <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-18\" id=\"user-content-fnref-18\">2</a></sup>. The market has consolidated around a handful of approaches: terminal-native agents (Claude Code, Codex CLI, Gemini CLI), IDE-integrated agents (Cursor, Windsurf, Cline), and cloud-based autonomous agents (Devin, Replit Agent 3, GitHub Copilot coding agent).</p>\n<p>This post examines how modern coding agents work, from the core loop architecture to context management strategies, tool implementations, and sandboxing approaches. We will survey the major agents available today and explore their design tradeoffs.</p>\n<hr />\n<h2 id=\"the-agentic-loop\">The Agentic Loop</h2>\n<p>Every coding agent runs some variant of an observe-think-act loop. OpenAI’s documentation on their Codex agent describes this as “unrolling” the loop; the agent repeatedly gathers information, decides what to do, executes an action, and evaluates the result <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-1\" id=\"user-content-fnref-1\">3</a></sup>.</p>\n<figure><figcaption><strong>The Core Agentic Loop</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/coding-agents-explained-0.webp\" width=\"512\" height=\"708\" alt=\"The Core Agentic Loop\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>The loop continues until the agent determines the task is complete or reaches a stopping condition (token limit, error threshold, or explicit user interrupt). This architecture differs from single-shot generation in a key way: the agent can course-correct. If it writes code that fails tests, it observes the failure, reasons about the fix, and tries again.</p>\n<h3 id=\"observe\">Observe</h3>\n<p>The observation phase involves gathering context about the current state. For coding agents, this means reading files, searching for patterns in the codebase, checking test results, and inspecting error messages. Agents have access to tools that return structured information: file contents, search results, command output.</p>\n<h3 id=\"think\">Think</h3>\n<p>The thinking phase is where the LLM reasons about what to do next. The model receives the accumulated context from observations and decides which tool to call next. This is not a separate system; it is the model’s native capability for planning and reasoning applied to a tool-use context.</p>\n<h3 id=\"act\">Act</h3>\n<p>The action phase executes a tool call. The agent might edit a file, run a shell command, search for references, or query a language server. Each action produces output that becomes input for the next observation phase.</p>\n<hr />\n<h2 id=\"tool-use-in-coding-agents\">Tool Use in Coding Agents</h2>\n<p>Tools give agents their capabilities. Without tools, an agent can only generate text. With tools, it can modify files, execute code, search repositories, and interact with external systems.</p>\n<h3 id=\"core-tool-categories\">Core Tool Categories</h3>\n<p><strong>File Operations</strong>: Read, write, and edit files. Most agents support both full file writes and targeted edits (replacing specific strings or line ranges). Claude Code implements an <code>Edit</code> tool that performs exact string replacement, reducing the risk of unintended changes <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-2\" id=\"user-content-fnref-2\">4</a></sup>.</p>\n<p><strong>Shell Execution</strong>: Run arbitrary commands. This enables testing, building, linting, and any other command-line operation. Agents typically capture stdout, stderr, and exit codes.</p>\n<p><strong>Search Tools</strong>: Find files by name patterns (glob) and content patterns (grep/ripgrep). Effective search is critical for navigating large codebases. Agents need to find relevant code without reading every file.</p>\n<p><strong>LSP Integration</strong>: Language Server Protocol provides code intelligence; go-to-definition, find-references, hover documentation, and symbol search. OpenCode and Claude Code both integrate LSP for structured code navigation <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-3\" id=\"user-content-fnref-3\">5</a></sup>.</p>\n<p><strong>MCP (Model Context Protocol)</strong>: Anthropic’s open standard for connecting AI assistants to external tools and data sources. Both Claude Code and Gemini CLI support MCP, enabling integration with systems like Jira, GitHub, and custom APIs <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-4\" id=\"user-content-fnref-4\">6</a></sup>.</p>\n<h3 id=\"tool-descriptions-and-schemas\">Tool Descriptions and Schemas</h3>\n<p>Agents rely on well-structured tool descriptions to use tools correctly. Each tool has a name, description, and parameter schema. The quality of these descriptions directly affects agent performance. Vague descriptions lead to incorrect tool usage; overly complex schemas increase error rates.</p>\n<p>Example tool schema structure:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"name\"</span><span>:</span><span> \"edit_file\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"description\"</span><span>:</span><span> \"Replace exact string matches in a file\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"parameters\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"file_path\"</span><span>:</span><span> \"Absolute path to the file\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"old_string\"</span><span>:</span><span> \"Exact text to find and replace\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"new_string\"</span><span>:</span><span> \"Replacement text\"</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<hr />\n<h2 id=\"context-management\">Context Management</h2>\n<p>Large codebases exceed context window limits. A moderately sized project might contain millions of tokens across thousands of files. Agents need strategies to work within context constraints while maintaining coherent understanding of the task.</p>\n<h3 id=\"repository-mapping\">Repository Mapping</h3>\n<p>Aider pioneered the repository map approach: generating a compact representation of the codebase structure, including file paths, function signatures, and class definitions <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-5\" id=\"user-content-fnref-5\">7</a></sup>. This map fits in context and provides the agent enough information to know which files to read for detailed work.</p>\n<h3 id=\"compaction\">Compaction</h3>\n<p>Compaction summarizes conversation history when approaching context limits. Rather than failing when the context fills up, the agent condenses older interactions while preserving essential information.</p>\n<p>OpenAI’s GPT-5.2 Codex is trained specifically for compaction with native compaction capabilities, making it token-efficient in its reasoning while handling long-running coding tasks <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-19\" id=\"user-content-fnref-19\">8</a></sup>. Anthropic’s Claude Code implements auto-compact at 95% context capacity, summarizing the trajectory of user-agent interactions <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-6\" id=\"user-content-fnref-6\">9</a></sup>.</p>\n<p>Compaction strategies vary:</p>\n<ul>\n<li><strong>Recursive summarization</strong>: Repeatedly condense earlier turns</li>\n<li><strong>Hierarchical summarization</strong>: Maintain summaries at different levels of detail</li>\n<li><strong>Tool result compaction</strong>: Replace verbose tool outputs with compact references (file paths instead of full contents)</li>\n</ul>\n<h3 id=\"mcp-tool-search\">MCP Tool Search</h3>\n<p>Claude Code introduced “MCP Tool Search” in January 2026, implementing lazy loading for AI tools <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-20\" id=\"user-content-fnref-20\">10</a></sup>. Instead of preloading every tool definition, Claude Code monitors context usage and automatically fetches tool descriptions only when necessary. The token savings are significant: from approximately 134k to 5k in Anthropic’s internal testing. Internal benchmarks show that enabling Tool Search improved the accuracy of Opus 4 on MCP evaluations from 49% to 74%, and for Opus 4.5, accuracy jumped from 79.5% to 88.1%.</p>\n<h3 id=\"sub-agent-isolation\">Sub-Agent Isolation</h3>\n<p>For complex tasks, agents can spawn sub-agents that operate in isolated context windows. The parent agent describes a task; the sub-agent explores extensively using its own context, then returns a condensed summary. This pattern appears in systems like Manus and OpenAI’s Codex <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-7\" id=\"user-content-fnref-7\">11</a></sup>.</p>\n<p>The key insight: sub-agents might use tens of thousands of tokens internally but return only 1,000-2,000 tokens of distilled results. This achieves context separation while preserving information.</p>\n<hr />\n<h2 id=\"sandboxing-and-security\">Sandboxing and Security</h2>\n<p>Coding agents execute arbitrary code. They run shell commands, modify files, and interact with external services. This creates security concerns: what prevents an agent from running <code>rm -rf /</code> or exfiltrating sensitive data?</p>\n<h3 id=\"container-based-isolation\">Container-Based Isolation</h3>\n<p>Docker containers provide filesystem isolation, process containment, and resource limits. The agent runs inside a container with access only to the project directory. Docker recently introduced Docker Sandboxes specifically for AI coding agents, with native support for Claude Code and Gemini CLI <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-8\" id=\"user-content-fnref-8\">12</a></sup>.</p>\n<p>Container isolation means the agent can install packages, run code, and modify files within the sandbox without affecting the host system. This approach handles the full environment the agent needs, not just the agent process itself.</p>\n<h3 id=\"os-level-sandboxing\">OS-Level Sandboxing</h3>\n<p>Claude Code on macOS uses the native sandbox facility to restrict agent actions. This provides lighter-weight isolation than containers but limits what the agent can access.</p>\n<h3 id=\"limitations\">Limitations</h3>\n<p>Container isolation alone does not address all risks. A sandbox controls where code runs and which files an agent can modify. It does not control what the agent is authorized to do across networked systems <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-9\" id=\"user-content-fnref-9\">13</a></sup>. An agent might have legitimate access to GitHub but use that access in unintended ways.</p>\n<p>Defense in depth remains necessary: container isolation, network segmentation, explicit permission prompts for sensitive actions, and audit logging.</p>\n<hr />\n<h2 id=\"survey-of-coding-agents\">Survey of Coding Agents</h2>\n<figure><figcaption><strong>Major Coding Agents Compared</strong></figcaption><table class=\"comparison-table\"><thead><tr><th class=\"agent-header\"></th><th class=\"agent-header \">Claude Code</th><th class=\"agent-header \">Codex CLI</th><th class=\"agent-header \">Gemini CLI</th><th class=\"agent-header highlighted\">OpenCode</th><th class=\"agent-header \">Aider</th><th class=\"agent-header \">Continue</th></tr></thead><tbody><tr class=\"spec-row\"><td class=\"spec-label\">License</td><td class=\"spec-value \">Proprietary</td><td class=\"spec-value \">Open Source</td><td class=\"spec-value \">Apache 2.0</td><td class=\"spec-value highlighted\">MIT</td><td class=\"spec-value \">Apache 2.0</td><td class=\"spec-value \">Apache 2.0</td></tr><tr class=\"spec-row\"><td class=\"spec-label\">Models</td><td class=\"spec-value \">Claude 4</td><td class=\"spec-value \">GPT-5.1-Codex</td><td class=\"spec-value \">Gemini 3 Flash/Pro</td><td class=\"spec-value highlighted\">Any (OpenAI, Claude, Gemini, local)</td><td class=\"spec-value \">Any (GPT-4, Claude, local)</td><td class=\"spec-value \">Any (configurable)</td></tr><tr class=\"spec-row\"><td class=\"spec-label\">Core Tools</td><td class=\"spec-value \">Bash, Read, Write, Edit, Grep, Glob, LSP, MCP</td><td class=\"spec-value \">Shell, file ops, code exec</td><td class=\"spec-value \">Shell, file ops, Search, MCP</td><td class=\"spec-value highlighted\">Shell, file ops, LSP</td><td class=\"spec-value \">Git, file ops, voice</td><td class=\"spec-value \">IDE integration, file ops</td></tr><tr class=\"spec-row\"><td class=\"spec-label\">Sandboxing</td><td class=\"spec-value \">macOS sandbox, Docker optional</td><td class=\"spec-value \">Docker containers</td><td class=\"spec-value \">Optional Docker</td><td class=\"spec-value highlighted\">None (native)</td><td class=\"spec-value \">None (native)</td><td class=\"spec-value \">None (IDE process)</td></tr><tr class=\"spec-row\"><td class=\"spec-label\">Context Mgmt</td><td class=\"spec-value \">Auto-compact at 95%</td><td class=\"spec-value \">Compaction (multi-window)</td><td class=\"spec-value \">1M token window</td><td class=\"spec-value highlighted\">Manual/configurable</td><td class=\"spec-value \">Repo map + smart chunking</td><td class=\"spec-value \">IDE-managed</td></tr></tbody></table></figure>\n<h3 id=\"terminal-native-agents\">Terminal-Native Agents</h3>\n<h4 id=\"claude-code\">Claude Code</h4>\n<p>Anthropic’s CLI agent runs in your terminal and integrates with your bash environment. The tech stack is TypeScript, React, Ink, and Bun <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-2\" id=\"user-content-fnref-2-2\">4</a></sup>. Design philosophy: low-level and unopinionated, providing close to raw model access without forcing specific workflows.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Bash, Read, Write, Edit, Grep, Glob, LSP, MCP client/server</li>\n<li>Context: Auto-compact at 95% capacity, MCP Tool Search for lazy loading</li>\n<li>Sandbox: macOS sandbox, optional Docker</li>\n<li>Model: Claude Opus 4.5 (80.9% SWE-bench Verified), Sonnet</li>\n</ul>\n<p>Claude Code v2.1.0 (January 2026) introduced Automatic Skill Hot-Reload, Skill Context Forking for isolated sub-agent contexts, and Hooks in Skill Frontmatter <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-20\" id=\"user-content-fnref-20-2\">10</a></sup>. The Cowork feature brings Claude Code’s agentic capabilities to the Claude desktop app, running locally in an isolated VM with access to local files and MCP integrations <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-21\" id=\"user-content-fnref-21\">14</a></sup>. Users report 50% to 75% reductions in both tool calling errors and build/lint errors with Claude Opus 4.5 <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-17\" id=\"user-content-fnref-17-2\">1</a></sup>.</p>\n<h4 id=\"codex-cli\">Codex CLI</h4>\n<p>OpenAI’s agent runs tasks in isolated cloud sandbox environments, preloaded with your repository. Powered by GPT-5.2 Codex for standard tasks and GPT-5.1-Codex-Max for long-running operations <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-19\" id=\"user-content-fnref-19-2\">8</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Shell, file operations, code execution, agent skills</li>\n<li>Context: Native compaction, collaboration tools for multi-agent coordination</li>\n<li>Sandbox: Docker containers in cloud</li>\n<li>Model: GPT-5.2 Codex (80% SWE-bench Verified, 56.4% SWE-bench Pro)</li>\n</ul>\n<p>Codex v0.85.0 (January 2026) introduced app-server v2 with collaboration tool calls emitted as item events, enabling real-time agent coordination rendering. The <code>spawn_agent</code> function now accepts an agent role preset for richer agent control <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-22\" id=\"user-content-fnref-22\">15</a></sup>. By exposing the CLI as an MCP server and orchestrating with the OpenAI Agents SDK, developers can create deterministic, auditable workflows that scale from a single agent to a complete software delivery pipeline.</p>\n<h4 id=\"gemini-cli\">Gemini CLI</h4>\n<p>Google’s open-source agent (Apache 2.0) brings Gemini models to the terminal. Uses a ReAct (reason and act) loop with built-in tools and MCP support <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-10\" id=\"user-content-fnref-10\">16</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Shell, file ops, Google Search grounding, web fetch, MCP</li>\n<li>Context: 1M token window with Gemini 3</li>\n<li>Sandbox: Optional Docker</li>\n<li>Model: Gemini 3 Flash (78% SWE-bench Verified) or Pro</li>\n</ul>\n<p>Gemini 3 Flash became available in Gemini CLI in December 2025, achieving 78% on SWE-bench Verified while being 3x faster than the 2.5 series at a fraction of the cost <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-23\" id=\"user-content-fnref-23\">17</a></sup>. Since launch, the community has contributed over 2,800 pull requests, submitted about 3,400 issues, and given more than 70,000 GitHub stars. The free tier offers 60 requests/minute and 1,000 requests/day with a personal Google account.</p>\n<h3 id=\"ide-integrated-agents\">IDE-Integrated Agents</h3>\n<h4 id=\"cursor\">Cursor</h4>\n<p>Cursor is a fork of VS Code that integrates AI capabilities directly into the editing experience. Agent is the default mode, designed to handle complex coding tasks with minimal guidance <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-24\" id=\"user-content-fnref-24\">18</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: File ops, terminal, web browsing, codebase indexing</li>\n<li>Context: Multi-file understanding with automatic context detection</li>\n<li>Sandbox: None (local execution)</li>\n<li>Models: Composer (proprietary), GPT-5 Codex, Claude</li>\n</ul>\n<p>Cursor 2.0 (late 2025) shipped Composer, their own ultra-fast coding model, and an agent-centric interface for running multiple agents in parallel <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-24\" id=\"user-content-fnref-24-2\">18</a></sup>. The January 2026 CLI release added Plan mode (<code>/plan</code> or <code>--mode=plan</code>) for approach design before coding, and cloud handoff for background execution. Users can prepend <code>&amp;</code> to any message to send it to a cloud agent, then resume on web or mobile at cursor.com/agents.</p>\n<h4 id=\"windsurf\">Windsurf</h4>\n<p>Windsurf (by Codeium) is an agentic IDE built for enterprise teams and large codebases. Its Cascade agent plans and executes multi-step changes across repositories <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-25\" id=\"user-content-fnref-25\">19</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: File ops, terminal, repository-wide context retrieval</li>\n<li>Context: Flow feature maintains persistent context across projects</li>\n<li>Sandbox: None (local execution)</li>\n<li>Models: Various (configurable)</li>\n</ul>\n<p>Windsurf’s “Context Awareness Engine” is faster at indexing than Cursor, making it suited for large-scale enterprise projects where the codebase exceeds what other tools can handle. Cascade reasons across entire repositories, determining which files matter for a given task and loading them automatically.</p>\n<h4 id=\"cline\">Cline</h4>\n<p>Cline is an open-source AI coding agent that runs inside VS Code or the terminal. It plans, previews, and applies multi-file changes with approval checkpoints <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-26\" id=\"user-content-fnref-26\">20</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: File ops, terminal, web browsing, MCP orchestration</li>\n<li>Context: Full repository access with diff transparency</li>\n<li>Sandbox: None (local-first control)</li>\n<li>Models: Model-agnostic (any provider)</li>\n</ul>\n<p>Cline demonstrates high autonomy with multi-step execution, self-correction, and independent task continuation. Its open-source nature and support for multiple AI models offer flexibility for teams that need local-first control over data and models.</p>\n<h3 id=\"cloud-based-autonomous-agents\">Cloud-Based Autonomous Agents</h3>\n<h4 id=\"devin\">Devin</h4>\n<p>Devin (by Cognition Labs) is an autonomous AI software engineer that operates as a web app rather than an IDE extension. Users define intent, review a plan, and execution proceeds in the background <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-27\" id=\"user-content-fnref-27\">21</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Code writing, PR creation, bug reproduction, internal tool building</li>\n<li>Context: Full codebase access with summarized intermediate steps</li>\n<li>Sandbox: Cloud-based isolated environments</li>\n<li>Model: Proprietary (trained with reinforcement learning)</li>\n</ul>\n<p>Devin can independently create PRs, respond to PR comments, review PRs, and handle Linear tickets when tagged. In 2025, Devin gained multi-agent operation capability, where one AI agent dispatches tasks to other AI agents. Devin Wiki and Devin Search provide machine-generated documentation and codebase querying. Pricing starts at <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>20</mn><mi>w</mi><mi>i</mi><mi>t</mi><mi>h</mi><mi>p</mi><mi>a</mi><mi>y</mi><mo>−</mo><mi>a</mi><mi>s</mi><mo>−</mo><mi>y</mi><mi>o</mi><mi>u</mi><mo>−</mo><mi>g</mi><mi>o</mi><mi>A</mi><mi>C</mi><mi>U</mi><mi>s</mi><mo stretchy=\"false\">(</mo></mrow></semantics></math>2.25 per ~15 minutes of active work) or $500/month for teams <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-27\" id=\"user-content-fnref-27-2\">21</a></sup>.</p>\n<h4 id=\"replit-agent-3\">Replit Agent 3</h4>\n<p>Replit Agent 3 (September 2025) is Replit’s most autonomous agent, positioned as Agent-first for all builders, not just developers <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-28\" id=\"user-content-fnref-28\">22</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Code writing, testing, deployment, agent creation</li>\n<li>Context: Full project access with reflection loops</li>\n<li>Sandbox: Cloud-based Replit environment</li>\n<li>Model: Proprietary</li>\n</ul>\n<p>Agent 3 runs for up to 200 minutes autonomously, handling full tasks with a proprietary testing system that is up to 3x faster and 10x more cost-effective than Computer Use models. For the first time, Agent 3 can build other agents and automations, enabling workflow automation via natural language. In January 2026, Replit reached a $3 billion valuation, with companies like Duolingo and Zillow using the platform <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-28\" id=\"user-content-fnref-28-2\">22</a></sup>.</p>\n<h4 id=\"github-copilot-coding-agent\">GitHub Copilot Coding Agent</h4>\n<p>GitHub Copilot coding agent works independently in the background to complete tasks like a human developer <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-29\" id=\"user-content-fnref-29\">23</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Code generation, file ops, Git integration</li>\n<li>Context: Repository-aware with Copilot Spaces</li>\n<li>Sandbox: Cloud-based execution</li>\n<li>Models: GPT-5 mini, GPT-4.1 (included without premium requests)</li>\n</ul>\n<p>The January 2026 CLI update introduced specialized custom agents: Explore for fast codebase analysis and Task for running commands like tests and builds <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-29\" id=\"user-content-fnref-29-2\">23</a></sup>. Visual Studio 2026 shipped with GitHub cloud agent in public preview. Users can delegate UI cleanups, refactors, documentation updates, and multi-file edits while focusing on core development.</p>\n<h4 id=\"amazon-q-developer\">Amazon Q Developer</h4>\n<p>Amazon Q Developer provides agentic capabilities for the AWS ecosystem <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-30\" id=\"user-content-fnref-30\">24</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Code generation, documentation, testing, code review, transformation</li>\n<li>Context: IDE and AWS service integration</li>\n<li>Sandbox: AWS environment</li>\n<li>Model: Proprietary</li>\n</ul>\n<p>Amazon Q Developer agents can autonomously implement features, document, test, review, and refactor code, and perform software upgrades. The Transformation Agent handles legacy modernization: Amazon used Q’s agents to upgrade 1,000 applications from Java 8 to Java 17, completing work that would have taken months in just two days <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-30\" id=\"user-content-fnref-30-2\">24</a></sup>. Free tier includes 50 agentic chat interactions per month; Pro tier is $19/user/month.</p>\n<h3 id=\"open-source-agents\">Open-Source Agents</h3>\n<h4 id=\"openhands\">OpenHands</h4>\n<p>OpenHands is an open platform for AI-powered coding agents with 65K+ GitHub stars <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-31\" id=\"user-content-fnref-31\">25</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Code editing, terminal, web browsing, file ops</li>\n<li>Context: Full repository access</li>\n<li>Sandbox: Docker or Kubernetes environments</li>\n<li>Models: Any (configurable)</li>\n</ul>\n<p>OpenHands 1.0.0 (January 2026) uses the new software-agent-sdk with optimizations across the app. The platform integrates with GitHub, GitLab, CI/CD, Slack, and ticketing tools. In November 2025, OpenHands raised $18.8M to build the open standard for autonomous software development. AMD partnered with OpenHands for local execution on AI PCs via the Lemonade LLM serving framework <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-31\" id=\"user-content-fnref-31-2\">25</a></sup>.</p>\n<h4 id=\"opencode\">OpenCode</h4>\n<p>Open-source (MIT license) alternative that supports any model provider: OpenAI, Anthropic, Google, AWS Bedrock, Groq, local models via Ollama <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-3\" id=\"user-content-fnref-3-2\">5</a></sup>. Built in Go with a Bubble Tea TUI.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Shell, file ops, LSP integration</li>\n<li>Context: Manual/configurable</li>\n<li>Sandbox: None (native execution)</li>\n<li>Models: Any supported provider</li>\n</ul>\n<p>OpenCode separates workflows into “Plan Mode” (read-only analysis) and “Build Mode” (full tool access), acting as both architect and engineer <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-3\" id=\"user-content-fnref-3-3\">5</a></sup>.</p>\n<h4 id=\"aider\">Aider</h4>\n<p>Open-source (Apache 2.0) pair programming tool focused on git integration. Creates a repository map of function signatures and file structures for intelligent multi-file edits <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-5\" id=\"user-content-fnref-5-2\">7</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: Git integration, file ops, voice input</li>\n<li>Context: Repo map plus smart chunking</li>\n<li>Sandbox: None (native execution)</li>\n<li>Models: Any (GPT-5, Claude, local)</li>\n</ul>\n<p>Aider automatically commits changes with sensible messages. Three modes: code (edit), architect (plan), ask (consult without changes) <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-5\" id=\"user-content-fnref-5-3\">7</a></sup>.</p>\n<h4 id=\"continue\">Continue</h4>\n<p>Open-source (Apache 2.0) IDE extension for VS Code and JetBrains. Architecture splits into core (business logic), extension (IDE-specific), and gui (React UI), communicating via message passing <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-11\" id=\"user-content-fnref-11\">26</a></sup>.</p>\n<p>Key characteristics:</p>\n<ul>\n<li>Tools: IDE integration, file ops</li>\n<li>Context: IDE-managed</li>\n<li>Sandbox: None (IDE process)</li>\n<li>Models: Any (configurable)</li>\n</ul>\n<p>Continue offers three interaction modes: Chat (discuss without changes), Plan (read-only exploration), Agent (full tool access) <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-11\" id=\"user-content-fnref-11-2\">26</a></sup>.</p>\n<hr />\n<h2 id=\"prompt-engineering-inside-agents\">Prompt Engineering Inside Agents</h2>\n<p>Agent system prompts are lengthy and detailed. They specify tool usage rules, safety constraints, output formatting, and behavioral guidelines. The prompt shapes how the agent interprets tasks and uses tools.</p>\n<h3 id=\"system-prompt-components\">System Prompt Components</h3>\n<p><strong>Tool descriptions</strong>: Detailed explanations of each tool’s purpose, parameters, and usage patterns. Often includes examples of correct and incorrect usage.</p>\n<p><strong>Safety rules</strong>: Constraints on dangerous operations. “NEVER run force push to main/master.” “Do not commit files that likely contain secrets.”</p>\n<p><strong>Workflow guidance</strong>: Instructions for common tasks like committing code, creating PRs, or exploring unfamiliar codebases.</p>\n<p><strong>Output formatting</strong>: How to structure responses, when to use code blocks, and how to communicate progress.</p>\n<h3 id=\"configuration-files\">Configuration Files</h3>\n<p>Many agents support project-level configuration:</p>\n<ul>\n<li><code>agents.md</code> (Codex): Tells the agent how you prefer to code</li>\n<li><code>CLAUDE.md</code> (Claude Code): Project-specific instructions and context</li>\n<li><code>.continuerc</code> (Continue): Model and tool configuration</li>\n</ul>\n<p>These files let developers customize agent behavior per repository.</p>\n<hr />\n<h2 id=\"failure-modes-and-limitations\">Failure Modes and Limitations</h2>\n<p>Coding agents fail. Understanding how they fail helps set appropriate expectations and design better workflows.</p>\n<h3 id=\"quality-concerns\">Quality Concerns</h3>\n<p>Research from CodeRabbit found that AI-generated code produces more issues across categories: 1.75x more logic errors, 1.64x more maintainability problems, 1.57x more security findings, and 1.42x more performance issues compared to human-written code <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-12\" id=\"user-content-fnref-12\">27</a></sup>.</p>\n<h3 id=\"tool-calling-failures\">Tool Calling Failures</h3>\n<p>Tool calling fails 3-15% of the time in production systems <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-13\" id=\"user-content-fnref-13\">28</a></sup>. The agent might call the wrong tool, pass incorrect parameters, or misinterpret tool output. Robust agents include retry logic and error handling.</p>\n<h3 id=\"multi-agent-system-failures\">Multi-Agent System Failures</h3>\n<p>Research analyzing 1,642 multi-agent execution traces found failure rates between 41% and 86.7% across state-of-the-art systems <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-14\" id=\"user-content-fnref-14\">29</a></sup>. Common issues include missing tool calls, workflow errors, and reasoning failures.</p>\n<h3 id=\"context-degradation\">Context Degradation</h3>\n<p>As context fills up, model performance degrades. Quality often drops before hitting the technical limit. The recommendation: implement compaction before hitting the “rot zone,” typically well before maximum context capacity <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-6\" id=\"user-content-fnref-6-2\">9</a></sup>.</p>\n<h3 id=\"debugging-challenges\">Debugging Challenges</h3>\n<p>“Ghost debugging” occurs when running the same prompt twice produces different results <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-13\" id=\"user-content-fnref-13-2\">28</a></sup>. Traditional debugging approaches fail because the system behavior is non-deterministic. Many fixes are “session patches” that work temporarily but do not persist across sessions.</p>\n<hr />\n<h2 id=\"building-your-own-agent\">Building Your Own Agent</h2>\n<p>Frameworks simplify agent construction. LangChain and LangGraph provide primitives for tool use, state management, and conversation handling.</p>\n<h3 id=\"langgraph\">LangGraph</h3>\n<p>LangChain recommends LangGraph for production agent implementations. It offers a durable runtime, model/tool swapping without rewrites, and 1000+ integrations <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-15\" id=\"user-content-fnref-15\">30</a></sup>.</p>\n<p>The <code>create_react_agent</code> function provides a standard ReAct pattern:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> langgraph</span><span>.</span><span>prebuilt </span><span>import</span><span> create_react_agent</span></span>\n<span class=\"line\"><span>from</span><span> langchain_openai </span><span>import</span><span> ChatOpenAI</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> ChatOpenAI</span><span>(model</span><span>=</span><span>\"gpt-4o\"</span><span>)</span></span>\n<span class=\"line\"><span>tools </span><span>=</span><span> [search_tool</span><span>,</span><span> file_tool</span><span>,</span><span> shell_tool]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>agent </span><span>=</span><span> create_react_agent</span><span>(model, tools)</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> agent</span><span>.</span><span>invoke</span><span>({</span><span>\"messages\"</span><span>: [(</span><span>\"user\"</span><span>, </span><span>\"Fix the failing tests\"</span><span>)]})</span></span></code></pre>\n<h3 id=\"open-swe\">Open SWE</h3>\n<p>LangChain’s Open SWE provides an open-source coding agent built on LangGraph with three specialized agents: Manager (entry point), Planner, and Programmer (with sub-agent Reviewer) <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-15\" id=\"user-content-fnref-15-2\">30</a></sup>. The entire project is open source and designed for extension.</p>\n<h3 id=\"custom-implementation\">Custom Implementation</h3>\n<p>A minimal agent loop requires:</p>\n<ol>\n<li>Message accumulation (conversation history)</li>\n<li>Tool definitions with schemas</li>\n<li>Loop logic: call model, parse tool calls, execute tools, append results</li>\n<li>Stopping conditions (task complete, error limit, token limit)</li>\n</ol>\n<p>The complexity comes from handling edge cases: malformed tool calls, execution timeouts, context overflow, and graceful degradation.</p>\n<hr />\n<h2 id=\"benchmarks-january-2026\">Benchmarks (January 2026)</h2>\n<p>The benchmark landscape has evolved with more rigorous evaluations. SWE-bench Verified remains the standard, but SWE-bench Pro and Terminal-Bench have emerged to address saturation at the top.</p>\n<h3 id=\"swe-bench-verified\">SWE-bench Verified</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Score</th></tr></thead><tbody><tr><td>Claude Opus 4.5</td><td>80.9%</td></tr><tr><td>GPT-5.2 Codex</td><td>80.0%</td></tr><tr><td>Gemini 3 Flash</td><td>78.0%</td></tr><tr><td>Gemini 3 Pro</td><td>76.2%</td></tr><tr><td>GPT-5.1</td><td>76.3%</td></tr><tr><td>Verdent</td><td>76.1% (81.2% pass@3)</td></tr><tr><td>GPT-5</td><td>74.9%</td></tr></tbody></table></div>\n<p>Claude Opus 4.5 became the first model to exceed 80% on SWE-bench Verified, solving 405 of 500 real-world coding problems <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-17\" id=\"user-content-fnref-17-3\">1</a></sup>.</p>\n<h3 id=\"swe-bench-pro\">SWE-bench Pro</h3>\n<p>SWE-bench Pro contains 1,865 total tasks across 41 professional repositories, designed to test harder software engineering scenarios <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-32\" id=\"user-content-fnref-32\">31</a></sup>.</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Score</th></tr></thead><tbody><tr><td>GPT-5.2 Codex</td><td>56.4%</td></tr><tr><td>GPT-5.1</td><td>50.8%</td></tr><tr><td>Claude Opus 4.1</td><td>23.1%</td></tr><tr><td>OpenAI GPT-5</td><td>23.3%</td></tr></tbody></table></div>\n<p>The gap between SWE-bench Verified (70%+ scores) and SWE-bench Pro (20-56%) reveals that current agents still struggle with professional-grade complexity.</p>\n<h3 id=\"terminal-bench\">Terminal-Bench</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Score</th></tr></thead><tbody><tr><td>Claude Opus 4.5</td><td>59.3%</td></tr><tr><td>Gemini 3 Pro</td><td>54.2%</td></tr><tr><td>GPT-5.1</td><td>47.6%</td></tr></tbody></table></div>\n<h3 id=\"code-quality-sonar-llm-leaderboard\">Code Quality (Sonar LLM Leaderboard)</h3>\n<p>GPT-5.2 High achieved the best security posture with only 16 blocker vulnerabilities per million lines of code, though it generated the highest code volume (974,379 lines) <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-33\" id=\"user-content-fnref-33\">32</a></sup>. Claude Opus 4.5 and Gemini 3 Pro lead in functional performance at 80.66%.</p>\n<h3 id=\"model-specialization\">Model Specialization</h3>\n<p>Top-tier models are diverging in specialization: GPT-5 excels at code review and refactoring, while Claude Sonnet 4.5 performs best in coding and tool use <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-34\" id=\"user-content-fnref-34\">33</a></sup>. Developers often switch between frontier models depending on the task.</p>\n<hr />\n<h2 id=\"the-future\">The Future</h2>\n<h3 id=\"multi-agent-systems\">Multi-Agent Systems</h3>\n<p>Gartner reported a 1,445% surge in multi-agent system inquiries from Q1 2024 to Q2 2025 <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-16\" id=\"user-content-fnref-16\">34</a></sup>. By 2026, 40% of enterprise applications will feature task-specific AI agents, up from less than 5% in 2025 <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-35\" id=\"user-content-fnref-35\">35</a></sup>. The shift from isolated agents to coordinated teams marks a fundamental change in how organizations approach automation.</p>\n<p>The pattern: orchestrator agents coordinate specialist agents (researcher, coder, analyst), each fine-tuned for specific capabilities. Agent design is converging around a Planner and Builder (Execution Agent) loop, which spawns ephemeral Task Agents for sub-routines, all grounded in a Code Execution Sandbox <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-35\" id=\"user-content-fnref-35-2\">35</a></sup>.</p>\n<h3 id=\"interoperability-protocols\">Interoperability Protocols</h3>\n<p>Protocols like Anthropic’s MCP and Google’s Agent-to-Agent Protocol (A2A) establish standards for agent interoperability <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-16\" id=\"user-content-fnref-16-2\">34</a></sup>. MCP standardizes how agents access tools and external resources. A2A enables peer-to-peer collaboration, allowing agents to negotiate, share findings, and coordinate without central oversight. Google’s Agent Development Kit (ADK), released in 2025, provides an open-source framework for building multi-agent systems <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-36\" id=\"user-content-fnref-36\">36</a></sup>.</p>\n<h3 id=\"parallel-workflows\">Parallel Workflows</h3>\n<p>Running multiple agents on the same codebase requires isolation. Git worktrees enable parallel branches for each task, with work merged back to main <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-16\" id=\"user-content-fnref-16-3\">34</a></sup>. Tools like Conductor and Verdent AI support running tasks in parallel. Cursor 2.0 introduced an agent-centric interface for managing multiple agents in parallel with cloud handoff capabilities.</p>\n<h3 id=\"longer-context-and-native-compaction\">Longer Context and Native Compaction</h3>\n<p>Gemini 3 offers a 1M token context window. GPT-5.2 Codex includes native compaction for token-efficient reasoning over long-running tasks <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-19\" id=\"user-content-fnref-19-3\">8</a></sup>. Claude Code’s MCP Tool Search reduces context usage by 96% (from 134k to 5k tokens) through lazy loading <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-20\" id=\"user-content-fnref-20-3\">10</a></sup>.</p>\n<h3 id=\"challenges-ahead\">Challenges Ahead</h3>\n<p>Security remains unsolved. Connecting models to tools multiplies risks; indirect prompt injections can cause harmful actions <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-16\" id=\"user-content-fnref-16-4\">34</a></sup>. Research analyzing 1,642 multi-agent execution traces found failure rates between 41% and 86.7% across state-of-the-art systems <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-14\" id=\"user-content-fnref-14-2\">29</a></sup>.</p>\n<p>AI agents are projected to generate $450 billion in economic value by 2028, yet only 2% of organizations have deployed them at full scale <sup><a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fn-35\" id=\"user-content-fnref-35-3\">35</a></sup>. The market has shipped useful point solutions but has not yet demonstrated reliable autonomy inside complex, decision-rich enterprise workflows.</p>\n<p>The trajectory is clear: agents will write more code. The question is whether we can make them reliable enough to trust.</p>\n<hr />\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-17\">\n<p>Anthropic, “Introducing Claude Opus 4.5,” January 2026. <a href=\"https://www.anthropic.com/news/claude-opus-4-5\" rel=\"noopener noreferrer\">https://www.anthropic.com/news/claude-opus-4-5</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-17\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-17-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-17-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-18\">\n<p>OpenAI, “Introducing GPT-5.2-Codex,” December 2025. <a href=\"https://openai.com/index/introducing-gpt-5-2-codex/\" rel=\"noopener noreferrer\">https://openai.com/index/introducing-gpt-5-2-codex/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-18\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-1\">\n<p>OpenAI, “Unrolling the Codex agent loop,” OpenAI Blog, 2025. <a href=\"https://openai.com/index/unrolling-the-codex-agent-loop/\" rel=\"noopener noreferrer\">https://openai.com/index/unrolling-the-codex-agent-loop/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Gergely Orosz, “How Claude Code is built,” The Pragmatic Engineer, 2025. <a href=\"https://newsletter.pragmaticengineer.com/p/how-claude-code-is-built\" rel=\"noopener noreferrer\">https://newsletter.pragmaticengineer.com/p/how-claude-code-is-built</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>OpenCode, “The open source AI coding agent,” 2025. <a href=\"https://opencode.ai/\" rel=\"noopener noreferrer\">https://opencode.ai/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-3-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-3-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Anthropic, “Claude Code overview,” Claude Code Docs, 2025. <a href=\"https://code.claude.com/docs/en/overview\" rel=\"noopener noreferrer\">https://code.claude.com/docs/en/overview</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>Aider, “AI Pair Programming in Your Terminal,” 2025. <a href=\"https://aider.chat/\" rel=\"noopener noreferrer\">https://aider.chat/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-5-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-5-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-19\">\n<p>OpenAI, “GPT-5.2-Codex,” OpenAI, 2025. <a href=\"https://openai.com/index/introducing-gpt-5-2-codex/\" rel=\"noopener noreferrer\">https://openai.com/index/introducing-gpt-5-2-codex/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-19\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-19-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-19-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p>Anthropic, “Effective context engineering for AI agents,” Anthropic Engineering, 2025. <a href=\"https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents\" rel=\"noopener noreferrer\">https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-6-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-20\">\n<p>VentureBeat, “Claude Code just got updated with one of the most-requested user features,” January 2026. <a href=\"https://venturebeat.com/orchestration/claude-code-just-got-updated-with-one-of-the-most-requested-user-features\" rel=\"noopener noreferrer\">https://venturebeat.com/orchestration/claude-code-just-got-updated-with-one-of-the-most-requested-user-features</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-20\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-20-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-20-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p>Lance Martin, “Context Engineering in Manus,” 2025. <a href=\"https://rlancemartin.github.io/2025/10/15/manus/\" rel=\"noopener noreferrer\">https://rlancemartin.github.io/2025/10/15/manus/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p>Docker, “A New Approach for Coding Agent Safety,” Docker Blog, 2025. <a href=\"https://www.docker.com/blog/docker-sandboxes-a-new-approach-for-coding-agent-safety/\" rel=\"noopener noreferrer\">https://www.docker.com/blog/docker-sandboxes-a-new-approach-for-coding-agent-safety/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-9\">\n<p>Arcade Blog, “Why Docker Sandboxes Alone Don’t Make AI Agents Safe,” 2025. <a href=\"https://blog.arcade.dev/docker-sandboxes-arent-enough-for-agent-safety\" rel=\"noopener noreferrer\">https://blog.arcade.dev/docker-sandboxes-arent-enough-for-agent-safety</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-9\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-21\">\n<p>TechRadar, “This is the Claude update I’ve been waiting for - Cowork could reshape how we use AI in 2026,” January 2026. <a href=\"https://www.techradar.com/ai-platforms-assistants/claudes-latest-upgrade-is-the-ai-breakthrough-ive-been-waiting-for-5-ways-cowork-could-be-the-biggest-ai-innovation-of-2026\" rel=\"noopener noreferrer\">https://www.techradar.com/ai-platforms-assistants/claudes-latest-upgrade-is-the-ai-breakthrough-ive-been-waiting-for-5-ways-cowork-could-be-the-biggest-ai-innovation-of-2026</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-21\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-22\">\n<p>OpenAI, “Codex Changelog,” January 2026. <a href=\"https://developers.openai.com/codex/changelog/\" rel=\"noopener noreferrer\">https://developers.openai.com/codex/changelog/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-22\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-10\">\n<p>Google, “Gemini CLI,” Google for Developers, 2025. <a href=\"https://developers.google.com/gemini-code-assist/docs/gemini-cli\" rel=\"noopener noreferrer\">https://developers.google.com/gemini-code-assist/docs/gemini-cli</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-10\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-23\">\n<p>Google Developers Blog, “Gemini 3 Flash is now available in Gemini CLI,” December 2025. <a href=\"https://developers.googleblog.com/gemini-3-flash-is-now-available-in-gemini-cli/\" rel=\"noopener noreferrer\">https://developers.googleblog.com/gemini-3-flash-is-now-available-in-gemini-cli/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-23\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-24\">\n<p>Cursor, “Changelog,” January 2026. <a href=\"https://cursor.com/changelog\" rel=\"noopener noreferrer\">https://cursor.com/changelog</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-24\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-24-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-25\">\n<p>Qodo, “Cline vs Windsurf: Best AI Coding Agent for Enterprise,” 2026. <a href=\"https://www.qodo.ai/blog/cline-vs-windsurf/\" rel=\"noopener noreferrer\">https://www.qodo.ai/blog/cline-vs-windsurf/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-25\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-26\">\n<p>Qodo, “Cline vs Windsurf,” 2026. <a href=\"https://www.qodo.ai/blog/cline-vs-windsurf/\" rel=\"noopener noreferrer\">https://www.qodo.ai/blog/cline-vs-windsurf/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-26\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-27\">\n<p>Builder.io, “Devin vs Cursor: How developers choose AI coding tools in 2026,” 2026. <a href=\"https://www.builder.io/blog/devin-vs-cursor\" rel=\"noopener noreferrer\">https://www.builder.io/blog/devin-vs-cursor</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-27\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-27-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-28\">\n<p>American Bazaar, “AI startup Replit, known for vibe coding, reaches $3 billion valuation,” January 2026. <a href=\"https://americanbazaaronline.com/2026/01/16/ai-startup-replit-known-for-vibe-coding-3-billion-valuation-473395/\" rel=\"noopener noreferrer\">https://americanbazaaronline.com/2026/01/16/ai-startup-replit-known-for-vibe-coding-3-billion-valuation-473395/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-28\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-28-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-29\">\n<p>GitHub Blog, “GitHub Copilot CLI: Enhanced agents, context management, and new ways to install,” January 2026. <a href=\"https://github.blog/changelog/2026-01-14-github-copilot-cli-enhanced-agents-context-management-and-new-ways-to-install/\" rel=\"noopener noreferrer\">https://github.blog/changelog/2026-01-14-github-copilot-cli-enhanced-agents-context-management-and-new-ways-to-install/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-29\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-29-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-30\">\n<p>AWS DevOps Blog, “Reinventing the Amazon Q Developer agent for software development,” 2025. <a href=\"https://aws.amazon.com/blogs/devops/reinventing-the-amazon-q-developer-agent-for-software-development/\" rel=\"noopener noreferrer\">https://aws.amazon.com/blogs/devops/reinventing-the-amazon-q-developer-agent-for-software-development/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-30\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-30-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-31\">\n<p>OpenHands, “The Open Platform for Cloud Coding Agents,” 2026. <a href=\"https://openhands.dev/\" rel=\"noopener noreferrer\">https://openhands.dev/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-31\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-31-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p>Continue, “Continue Documentation,” 2025. <a href=\"https://docs.continue.dev/\" rel=\"noopener noreferrer\">https://docs.continue.dev/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-11-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-12\">\n<p>The Register, “AI-authored code needs more attention, contains worse bugs,” December 2025. <a href=\"https://www.theregister.com/2025/12/17/ai_code_bugs/\" rel=\"noopener noreferrer\">https://www.theregister.com/2025/12/17/ai_code_bugs/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-12\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-13\">\n<p>Michael Hannecke, “Why AI Agents Fail in Production,” Medium, 2025. <a href=\"https://medium.com/@michael.hannecke/why-ai-agents-fail-in-production-what-ive-learned-the-hard-way-05f5df98cbe5\" rel=\"noopener noreferrer\">https://medium.com/@michael.hannecke/why-ai-agents-fail-in-production-what-ive-learned-the-hard-way-05f5df98cbe5</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-13\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-13-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p>Mert Cemri et al., “Why Do Multi-Agent LLM Systems Fail?” arXiv, 2025. <a href=\"https://arxiv.org/pdf/2503.13657\" rel=\"noopener noreferrer\">https://arxiv.org/pdf/2503.13657</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-14-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p>LangChain, “Introducing Open SWE: An Open-Source Asynchronous Coding Agent,” LangChain Blog, 2025. <a href=\"https://www.blog.langchain.com/introducing-open-swe-an-open-source-asynchronous-coding-agent/\" rel=\"noopener noreferrer\">https://www.blog.langchain.com/introducing-open-swe-an-open-source-asynchronous-coding-agent/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-15-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-32\">\n<p>Scale AI, “SWE-Bench Pro Public Dataset,” January 2026. <a href=\"https://scale.com/leaderboard/swe_bench_pro_public\" rel=\"noopener noreferrer\">https://scale.com/leaderboard/swe_bench_pro_public</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-32\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-33\">\n<p>Sonar, “New data on code quality: GPT-5.2 high, Opus 4.5, Gemini 3, and more,” January 2026. <a href=\"https://www.sonarsource.com/blog/new-data-on-code-quality-gpt-5-2-high-opus-4-5-gemini-3-and-more/\" rel=\"noopener noreferrer\">https://www.sonarsource.com/blog/new-data-on-code-quality-gpt-5-2-high-opus-4-5-gemini-3-and-more/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-33\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-34\">\n<p>SWE-rebench, “Leaderboard,” January 2026. <a href=\"https://swe-rebench.com/\" rel=\"noopener noreferrer\">https://swe-rebench.com</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-34\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-16\">\n<p>RTInsights, “2026 will be the Year of Multiple AI Agents,” January 2026. <a href=\"https://www.rtinsights.com/if-2025-was-the-year-of-ai-agents-2026-will-be-the-year-of-multi-agent-systems/\" rel=\"noopener noreferrer\">https://www.rtinsights.com/if-2025-was-the-year-of-ai-agents-2026-will-be-the-year-of-multi-agent-systems/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-16\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-16-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-16-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-16-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a></p>\n</li>\n<li id=\"user-content-fn-35\">\n<p>The New Stack, “5 Key Trends Shaping Agentic Development in 2026,” January 2026. <a href=\"https://thenewstack.io/5-key-trends-shaping-agentic-development-in-2026/\" rel=\"noopener noreferrer\">https://thenewstack.io/5-key-trends-shaping-agentic-development-in-2026/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-35\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-35-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-35-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-36\">\n<p>Google Developers Blog, “Agent Development Kit: Making it easy to build multi-agent applications,” 2025. <a href=\"https://developers.googleblog.com/en/agent-development-kit-easy-to-build-multi-agent-applications/\" rel=\"noopener noreferrer\">https://developers.googleblog.com/en/agent-development-kit-easy-to-build-multi-agent-applications/</a> <a href=\"https://blog.ecitis.org/coding-agents-explained/#user-content-fnref-36\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/coding-agents-explained/",
            "title": "How Coding Agents Work: Inside the Agentic Loop",
            "summary": "Look inside the observe, reason, act, and verify loop that turns language models into useful software-engineering agents.",
            "image": "https://blog.ecitis.org/open-graph/coding-agents-explained.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Ecosystems & Tooling",
                "Agents",
                "Tool use",
                "Software engineering"
            ]
        },
        {
            "id": "https://blog.ecitis.org/edge-ai-deployment/",
            "content_html": "<h2 id=\"why-deploy-ai-at-the-edge\">Why deploy AI at the edge</h2>\n<p>Running inference on edge devices instead of cloud servers addresses four constraints that matter in production systems.</p>\n<p><strong>Latency</strong>: Cloud roundtrips add 100-500ms of network delay before inference even begins. Edge inference on embedded hardware achieves sub-50ms latency for many workloads. Cactus, a Y Combinator startup, demonstrated sub-50ms time-to-first-token for on-device LLM inference in late 2025.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup> TinyML solutions on microcontrollers achieve inference latencies as low as 0-5ms.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup></p>\n<p><strong>Privacy</strong>: Data never leaves the device. This simplifies compliance with GDPR, HIPAA, and similar regulations. Medical imaging, financial analysis, and surveillance applications often require on-premise processing.</p>\n<p><strong>Connectivity</strong>: Edge deployment removes the assumption of reliable internet. Factory floors, remote agricultural sites, vehicles in motion, and aircraft cabins all benefit from inference that works without network access.</p>\n<p><strong>Cost</strong>: After hardware acquisition, inference is free. No per-token API fees, no bandwidth costs, no cloud compute bills. IDC and Gartner predict that by 2027, over 60% of all AI inference will happen locally rather than in the cloud.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup></p>\n<hr />\n<h2 id=\"hardware-landscape\">Hardware landscape</h2>\n<p>The edge AI hardware market reached <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>30.74</mn><mi>b</mi><mi>i</mi><mi>l</mi><mi>l</mi><mi>i</mi><mi>o</mi><mi>n</mi><mi>i</mi><mi>n</mi><mn>2026</mn><mi>a</mi><mi>n</mi><mi>d</mi><mi>i</mi><mi>s</mi><mi>p</mi><mi>r</mi><mi>o</mi><mi>j</mi><mi>e</mi><mi>c</mi><mi>t</mi><mi>e</mi><mi>d</mi><mi>t</mi><mi>o</mi><mi>g</mi><mi>r</mi><mi>o</mi><mi>w</mi><mi>t</mi><mi>o</mi></mrow></semantics></math>68.73 billion by 2031 at a 17.46% CAGR.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup> Leading mobile processors now achieve 45-50 TOPS of inference capability while optimizing battery life through dedicated engines.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-4\" id=\"user-content-fnref-4-2\">4</a></sup></p>\n<figure><figcaption><strong>Edge AI Hardware Comparison</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/edge-ai-deployment-0.webp\" width=\"512\" height=\"770\" alt=\"Edge AI Hardware Comparison\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"nvidia-jetson-platform\">NVIDIA Jetson platform</h3>\n<p>NVIDIA’s Jetson line targets robotics, autonomous machines, and embedded vision applications. The platform runs a full Linux stack with CUDA support. At CES 2026, NVIDIA expanded the lineup with Blackwell-based hardware.</p>\n<p><strong>Jetson AGX Thor</strong>: The new flagship module delivers up to 2070 FP4 TFLOPS (over 1000 INT8 TOPS) with 128GB of LPDDR5X memory.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup> Built on the Blackwell GPU architecture with 2,560 CUDA cores and 96 fifth-generation Tensor Cores, Thor provides 7.5x higher AI compute and 3.5x greater energy efficiency compared to Jetson Orin.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-6\" id=\"user-content-fnref-6\">6</a></sup> Power is configurable between 75W and 130W. The developer kit costs $3,499.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-7\" id=\"user-content-fnref-7\">7</a></sup> Early adopters include Boston Dynamics, Amazon Robotics, Figure, and Meta.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-6\" id=\"user-content-fnref-6-2\">6</a></sup></p>\n<p><strong>Jetson T4000</strong>: The newest Thor family member launched at CES 2026, delivering 1200 TFLOPS of AI compute and 64GB of memory while running at 40-70W.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-8\" id=\"user-content-fnref-8\">8</a></sup> It features three 25GbE ports for high-bandwidth sensor fusion in edge and robotic applications.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-8\" id=\"user-content-fnref-8-2\">8</a></sup></p>\n<p><strong>Jetson AGX Orin</strong>: The previous-generation flagship delivers up to 275 TOPS of AI performance with power configurable between 15W and 60W.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-9\" id=\"user-content-fnref-9\">9</a></sup> This handles multiple concurrent vision pipelines or real-time video analytics.</p>\n<p><strong>Jetson Orin NX</strong>: Mid-range option at 157 TOPS with 10-40W power envelope. Fits the smallest Jetson form factor while maintaining substantial compute capability.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-9\" id=\"user-content-fnref-9-2\">9</a></sup></p>\n<p><strong>Jetson Orin Nano Super</strong>: Entry-level edge AI at $249. A software update in late 2024 boosted performance from 40 to 67 TOPS and memory bandwidth from 68 to 102 GB/s. The 8GB module uses an Ampere architecture GPU with 1024 CUDA cores and 32 Tensor Cores running at up to 25W.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-10\" id=\"user-content-fnref-10\">10</a></sup></p>\n<p>All Jetson modules share a common software stack through JetPack SDK. JetPack 7 supports Jetson Thor with Linux 24.04 LTS, Kernel 6.8, and the latest compute stack.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-5\" id=\"user-content-fnref-5-2\">5</a></sup></p>\n<h3 id=\"raspberry-pi-ai-acceleration\">Raspberry Pi AI acceleration</h3>\n<p>The Raspberry Pi Foundation partnered with Hailo to bring neural network acceleration to the Pi 5. In January 2026, they expanded the lineup with generative AI support.</p>\n<p><strong>AI HAT+ 2</strong>: Released January 15, 2026, this $130 module features the Hailo-10H accelerator with 40 TOPS (INT8) performance and 8GB of dedicated LPDDR4X RAM.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-11\" id=\"user-content-fnref-11\">11</a></sup> Unlike previous AI HATs focused on computer vision, the AI HAT+ 2 targets generative AI workloads including LLMs and VLMs.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-12\" id=\"user-content-fnref-12\">12</a></sup> Supported models at launch include DeepSeek-R10-Distill, Llama 3.2, Qwen2.5-Coder, and Qwen2.5-Instruct (1.5B parameter versions).<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-11\" id=\"user-content-fnref-11-2\">11</a></sup> The chip runs at a maximum of 3W.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-13\" id=\"user-content-fnref-13\">13</a></sup></p>\n<p><strong>AI HAT+ 26T</strong>: Uses the full Hailo-8 chip for 26 TOPS. Both variants integrate with rpicam-apps for direct camera pipeline acceleration.</p>\n<p><strong>AI HAT+ 13T</strong>: Contains a Hailo-8L chip delivering 13 TOPS at 3-4 TOPS/W efficiency. The module connects via PCIe 2.0 through an M.2 interface. Cost is approximately $70.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-14\" id=\"user-content-fnref-14\">14</a></sup></p>\n<p>Raspberry Pi OS automatically detects Hailo modules and makes the NPU available for inference. For vision-based workloads, the AI HAT+ or AI Kit remain cost-effective options. The AI HAT+ 2 adds generative AI capabilities but reviewers note practical LLM performance remains limited.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-15\" id=\"user-content-fnref-15\">15</a></sup></p>\n<h3 id=\"google-coral-edge-tpu\">Google Coral Edge TPU</h3>\n<p>Google’s Coral line provides a USB-connected TPU accelerator for existing systems.</p>\n<p>The Edge TPU delivers 4 TOPS while consuming 2W; that is 2 TOPS per watt.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-16\" id=\"user-content-fnref-16\">16</a></sup> The USB Accelerator runs MobileNet v2 at nearly 400 FPS for image classification tasks. The limitation is model compatibility: only TensorFlow Lite models compiled specifically for the Edge TPU will accelerate.</p>\n<p>The TPU uses a 64x64 systolic array (estimated, as Google has not published exact specifications) running at approximately 480 MHz. On-chip SRAM is limited to 8MB, so large models must stream weights from the host.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-17\" id=\"user-content-fnref-17\">17</a></sup></p>\n<h3 id=\"mobile-neural-processing-units\">Mobile Neural Processing Units</h3>\n<p>Modern smartphones include dedicated neural engines that rival discrete edge accelerators. Leading mobile processors now achieve 45-50 TOPS of inference capability.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-4\" id=\"user-content-fnref-4-3\">4</a></sup></p>\n<p><strong>Apple Neural Engine</strong>: The M5 chip (October 2025) includes a 16-core Neural Engine with Neural Accelerators in each GPU core, delivering up to 3.5x the AI performance of M4.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-18\" id=\"user-content-fnref-18\">18</a></sup> The M4 chip delivers 38 TOPS, more than double the M3’s 18 TOPS.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-19\" id=\"user-content-fnref-19\">19</a></sup> The A19 Pro (iPhone 17 Pro, September 2025) features a 16-core Neural Engine with 38 TOPS and improved memory bandwidth, with Neural Accelerators built into each GPU core.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-20\" id=\"user-content-fnref-20\">20</a></sup> Core ML automatically dispatches workloads to ANE, GPU, or CPU based on model requirements. Apple reports that ANE-optimized models run up to 10x faster with 14x less memory than non-optimized versions.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-21\" id=\"user-content-fnref-21\">21</a></sup></p>\n<p><strong>Qualcomm Hexagon NPU</strong>: Snapdragon 8 Elite (Gen 5) NPUs deliver time-to-first-token in just 0.12 seconds on high-resolution images (1024x1024), with up to 100x speedup over CPU and 10x over GPU.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-22\" id=\"user-content-fnref-22\">22</a></sup> The Snapdragon X Elite delivers 45 TOPS of INT8 performance.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-23\" id=\"user-content-fnref-23\">23</a></sup> Research shows consistent 50% improvement in prefill speed and up to 110% improvement in decode speed with each successive generation of Snapdragon SoCs.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-24\" id=\"user-content-fnref-24\">24</a></sup> The NPU is now a standard component, with over 80% of recent Qualcomm SoCs including one.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-22\" id=\"user-content-fnref-22-2\">22</a></sup></p>\n<p>At CES 2026, Qualcomm debuted the Dragonwing 1Q10, an 18-core CPU robotics platform designed to compete with NVIDIA’s Jetson ecosystem.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-25\" id=\"user-content-fnref-25\">25</a></sup></p>\n<h3 id=\"nordic-semiconductor-iot-edge-ai\">Nordic Semiconductor (IoT Edge AI)</h3>\n<p>Nordic Semiconductor announced the nRF54LM20B SoC at CES 2026, integrating the Axon Neural Processing Unit for ultra-low-power edge AI.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-26\" id=\"user-content-fnref-26\">26</a></sup> Axon delivers up to 7x faster performance and 8x higher energy efficiency versus competing solutions for tasks like sound classification, keyword spotting, and image-based detection.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-26\" id=\"user-content-fnref-26-2\">26</a></sup> Broad availability expected Q2 2026.</p>\n<h3 id=\"ambarella-cv7\">Ambarella CV7</h3>\n<p>Ambarella announced the CV7 edge AI vision SoC at CES 2026, optimized for AI perception applications with 4nm process technology.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-27\" id=\"user-content-fnref-27\">27</a></sup> The low power consumption reduces thermal management requirements for smaller form factors and longer battery life across AIoT applications.</p>\n<hr />\n<h2 id=\"model-optimization-for-edge\">Model optimization for edge</h2>\n<p>Edge devices cannot run full-precision models designed for datacenter GPUs. Three techniques reduce model size and compute requirements.</p>\n<figure><figcaption><strong>Edge Deployment Software Stacks</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/edge-ai-deployment-1.webp\" width=\"512\" height=\"681\" alt=\"Edge Deployment Software Stacks\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"quantization\">Quantization</h3>\n<p>Quantization reduces the numerical precision of model weights and activations. Standard training uses FP32; quantization converts to FP16, INT8, or even INT4.</p>\n<p><strong>Post-training quantization (PTQ)</strong> applies after training completes. TensorFlow Lite’s full integer quantization produces INT8 models compatible with accelerators like Google Coral and microcontrollers. The process requires a representative dataset to calibrate quantization ranges.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-28\" id=\"user-content-fnref-28\">28</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> tensorflow </span><span>as</span><span> tf</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>converter </span><span>=</span><span> tf</span><span>.</span><span>lite</span><span>.</span><span>TFLiteConverter</span><span>.</span><span>from_saved_model</span><span>(saved_model_dir)</span></span>\n<span class=\"line\"><span>converter</span><span>.</span><span>optimizations </span><span>=</span><span> [tf</span><span>.</span><span>lite</span><span>.</span><span>Optimize</span><span>.</span><span>DEFAULT]</span></span>\n<span class=\"line\"><span>converter</span><span>.</span><span>representative_dataset </span><span>=</span><span> representative_dataset</span></span>\n<span class=\"line\"><span>converter</span><span>.</span><span>target_spec</span><span>.</span><span>supported_ops </span><span>=</span><span> [tf</span><span>.</span><span>lite</span><span>.</span><span>OpsSet</span><span>.</span><span>TFLITE_BUILTINS_INT8]</span></span>\n<span class=\"line\"><span>converter</span><span>.</span><span>inference_input_type </span><span>=</span><span> tf</span><span>.</span><span>int8</span></span>\n<span class=\"line\"><span>converter</span><span>.</span><span>inference_output_type </span><span>=</span><span> tf</span><span>.</span><span>int8</span></span>\n<span class=\"line\"><span>tflite_model </span><span>=</span><span> converter</span><span>.</span><span>convert</span><span>()</span></span></code></pre>\n<p>INT8 quantization produces models 4x smaller than FP32 with inference speedups of 2-4x on supporting hardware.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-28\" id=\"user-content-fnref-28-2\">28</a></sup></p>\n<p><strong>Quantization-aware training (QAT)</strong> integrates quantization into the training loop. The model learns to compensate for quantization errors. PyTorch demonstrated that QAT recovers up to 96% of accuracy degradation on HellaSwag and 68% of perplexity degradation on WikiText compared to PTQ.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-29\" id=\"user-content-fnref-29\">29</a></sup></p>\n<p>For LLMs on edge, GGUF quantization through llama.cpp offers multiple precision levels from 1.5-bit to 8-bit:<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-30\" id=\"user-content-fnref-30\">30</a></sup></p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Quantization</th><th>Size (7B model)</th><th>Perplexity increase</th><th>Notes</th></tr></thead><tbody><tr><td>Q8_0</td><td>~7.0 GB</td><td>+0.0004</td><td>Highest quality</td></tr><tr><td>Q5_K_M</td><td>~5.33 GB</td><td>+0.0142</td><td>Optimal balance</td></tr><tr><td>Q4_K_M</td><td>~4.58 GB</td><td>+0.0535</td><td>Recommended default</td></tr><tr><td>Q3_K_M</td><td>~3.52 GB</td><td>+0.2437</td><td>Aggressive compression</td></tr><tr><td>IQ2_XS</td><td>~2.31 bpw</td><td>Higher</td><td>Extreme compression</td></tr></tbody></table></div>\n<p>Q5_K_M is generally regarded as the optimal balance, applying five bits to most weights while retaining higher precision for crucial layers like attention.wv and feed_forward.w2.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-31\" id=\"user-content-fnref-31\">31</a></sup> For extreme compression, the IQ (importance-quantized) formats support down to 1.5 bits per weight.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-30\" id=\"user-content-fnref-30-2\">30</a></sup></p>\n<p><strong>New quantization formats (2026)</strong>: NVIDIA announced NVFP4 and FP8 quantization support for llama.cpp and Ollama at CES 2026, enabling up to 35% faster token generation on RTX hardware.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-32\" id=\"user-content-fnref-32\">32</a></sup></p>\n<h3 id=\"pruning\">Pruning</h3>\n<p>Pruning removes weights that contribute minimally to model output. Structured pruning removes entire channels or neurons; the resulting model runs faster on standard hardware because tensor dimensions shrink.</p>\n<p>Research in 2025 achieved up to 400x reduction in neural network size for reinforcement learning tasks by pruning 99% of weights.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-33\" id=\"user-content-fnref-33\">33</a></sup> For transformer models, a compression approach evaluated on language modeling tasks achieved around 70% overall model compression while maintaining accuracy.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-34\" id=\"user-content-fnref-34\">34</a></sup></p>\n<p>Unstructured pruning sets weights to zero without changing tensor dimensions. This produces sparse matrices that require specialized hardware or software to accelerate.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>utils</span><span>.</span><span>prune </span><span>as</span><span> prune</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Structured pruning: remove 50% of channels by L1 norm</span></span>\n<span class=\"line\"><span>prune</span><span>.</span><span>ln_structured</span><span>(model.conv1, name</span><span>=</span><span>'weight'</span><span>, amount</span><span>=</span><span>0.5</span><span>, n</span><span>=</span><span>1</span><span>, dim</span><span>=</span><span>0</span><span>)</span></span></code></pre>\n<h3 id=\"knowledge-distillation\">Knowledge distillation</h3>\n<p>Distillation transfers knowledge from a large teacher model to a smaller student model. The student learns to match the teacher’s output distributions rather than just the hard labels.</p>\n<p>A hybrid approach combining knowledge distillation, pruning, and quantization produced models 3x smaller than vanilla CNNs while achieving 97% accuracy.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-35\" id=\"user-content-fnref-35\">35</a></sup> The CQKD framework demonstrated 34,000x compression while preserving accuracy through combined cluster-based quantization and knowledge distillation.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-36\" id=\"user-content-fnref-36\">36</a></sup></p>\n<p>The shift from large language models (LLMs) to small, task-specific language models (SLMs) emerged as a key trend in 2026, enabling efficient, localized AI deployments with reduced power and compute needs.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-27\" id=\"user-content-fnref-27-2\">27</a></sup></p>\n<hr />\n<h2 id=\"deployment-frameworks\">Deployment frameworks</h2>\n<h3 id=\"tensorrt-for-jetson\">TensorRT for Jetson</h3>\n<p>TensorRT is NVIDIA’s inference optimizer for Jetson and datacenter GPUs. It converts trained models into optimized engines through graph compilation, layer fusion, and kernel autotuning.</p>\n<p>The optimizer identifies operation sequences that can be fused into single GPU kernels. A matmul followed by ReLU becomes one kernel without intermediate memory writes. TensorRT automatically selects optimal GPU kernels based on architecture, precision, and batch size.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-37\" id=\"user-content-fnref-37\">37</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Convert ONNX model to TensorRT engine</span></span>\n<span class=\"line\"><span>trtexec</span><span> --onnx=model.onnx</span><span> \\</span></span>\n<span class=\"line\"><span>        --saveEngine=model.trt</span><span> \\</span></span>\n<span class=\"line\"><span>        --fp16</span><span> \\</span></span>\n<span class=\"line\"><span>        --workspace=4096</span></span></code></pre>\n<p>Benchmark results on Jetson Nano showed TensorRT optimization achieving 95-110 FPS where non-optimized inference ran at 5-25 FPS.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-38\" id=\"user-content-fnref-38\">38</a></sup> On average, optimized models exhibit 16% speed improvement over non-optimized counterparts on Jetson hardware.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-39\" id=\"user-content-fnref-39\">39</a></sup></p>\n<h3 id=\"tensorflow-lite\">TensorFlow Lite</h3>\n<p>TensorFlow Lite runs on Android, iOS, embedded Linux, and microcontrollers. It supports GPU delegation on mobile devices and accelerator integration through the delegate API.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> tensorflow </span><span>as</span><span> tf</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load and run TFLite model</span></span>\n<span class=\"line\"><span>interpreter </span><span>=</span><span> tf</span><span>.</span><span>lite</span><span>.</span><span>Interpreter</span><span>(model_path</span><span>=</span><span>\"model.tflite\"</span><span>)</span></span>\n<span class=\"line\"><span>interpreter</span><span>.</span><span>allocate_tensors</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>input_details </span><span>=</span><span> interpreter</span><span>.</span><span>get_input_details</span><span>()</span></span>\n<span class=\"line\"><span>output_details </span><span>=</span><span> interpreter</span><span>.</span><span>get_output_details</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>interpreter</span><span>.</span><span>set_tensor</span><span>(input_details[</span><span>0</span><span>][</span><span>'index'</span><span>], input_data)</span></span>\n<span class=\"line\"><span>interpreter</span><span>.</span><span>invoke</span><span>()</span></span>\n<span class=\"line\"><span>output </span><span>=</span><span> interpreter</span><span>.</span><span>get_tensor</span><span>(output_details[</span><span>0</span><span>][</span><span>'index'</span><span>])</span></span></code></pre>\n<p>The Edge TPU delegate routes compatible operations to Google Coral hardware. GPU delegates accelerate on mobile GPUs. The LiteRT QNN accelerator supports 90 LiteRT ops, allowing 64 of 72 models to delegate fully to the Qualcomm NPU.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-22\" id=\"user-content-fnref-22-3\">22</a></sup></p>\n<h3 id=\"core-ml\">Core ML</h3>\n<p>Core ML deploys models on Apple devices across iOS, macOS, watchOS, and tvOS. The framework automatically dispatches to CPU, GPU, or Neural Engine based on model characteristics.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> coremltools </span><span>as</span><span> ct</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Convert PyTorch model to Core ML</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> ct</span><span>.</span><span>convert</span><span>(</span></span>\n<span class=\"line\"><span>    torch_model,</span></span>\n<span class=\"line\"><span>    inputs</span><span>=</span><span>[ct.</span><span>TensorType</span><span>(shape</span><span>=</span><span>(</span><span>1</span><span>, </span><span>3</span><span>, </span><span>224</span><span>, </span><span>224</span><span>))],</span></span>\n<span class=\"line\"><span>    compute_precision</span><span>=</span><span>ct.precision.FLOAT16,</span></span>\n<span class=\"line\"><span>    compute_units</span><span>=</span><span>ct.ComputeUnit.ALL</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>save</span><span>(</span><span>\"model.mlpackage\"</span><span>)</span></span></code></pre>\n<p>macOS Sequoia introduced low-bit quantization methods including 4-bit block-wise linear quantization and channel group-wise palettization. These reduce memory footprint for on-device LLM inference.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-40\" id=\"user-content-fnref-40\">40</a></sup></p>\n<p>ExecuTorch, the PyTorch edge runtime, includes a Core ML backend that dispatches to ANE, GPU, or CPU with FP16 precision.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-41\" id=\"user-content-fnref-41\">41</a></sup></p>\n<h3 id=\"onnx-runtime\">ONNX Runtime</h3>\n<p>ONNX Runtime provides cross-platform inference with execution providers for different hardware. The CPU provider uses optimized kernels; GPU providers support CUDA, DirectML, and Metal.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> onnxruntime </span><span>as</span><span> ort</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run with CUDA execution provider</span></span>\n<span class=\"line\"><span>session </span><span>=</span><span> ort</span><span>.</span><span>InferenceSession</span><span>(</span></span>\n<span class=\"line\"><span>    \"model.onnx\"</span><span>,</span></span>\n<span class=\"line\"><span>    providers</span><span>=</span><span>[</span><span>'CUDAExecutionProvider'</span><span>, </span><span>'CPUExecutionProvider'</span><span>]</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> session</span><span>.</span><span>run</span><span>(</span><span>None</span><span>, {</span><span>\"input\"</span><span>: input_data})</span></span></code></pre>\n<hr />\n<h2 id=\"running-llms-on-edge-devices\">Running LLMs on edge devices</h2>\n<h3 id=\"llamacpp-on-raspberry-pi\">llama.cpp on Raspberry Pi</h3>\n<p>llama.cpp runs quantized LLMs on CPUs and GPUs without Python dependencies. Written in pure C/C++ with no external dependencies, it supports ARM devices, Raspberry Pi, and other edge hardware.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-42\" id=\"user-content-fnref-42\">42</a></sup></p>\n<p><strong>Benchmark results on Pi 5 (8GB)</strong>:</p>\n<ul>\n<li>1B models (Q4): 7+ tokens/second<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-43\" id=\"user-content-fnref-43\">43</a></sup></li>\n<li>3B models (Q4): 4-7 tokens/second<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-43\" id=\"user-content-fnref-43-2\">43</a></sup></li>\n<li>7B models (Q4_K_M): 0.7-3 tokens/second<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-44\" id=\"user-content-fnref-44\">44</a></sup></li>\n</ul>\n<p>Using BLIS or OpenBLAS for matrix operations improves throughput. One user achieved 5.02 tokens/second on Pi 5 16GB with BLIS optimization.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-45\" id=\"user-content-fnref-45\">45</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Build llama.cpp on Raspberry Pi</span></span>\n<span class=\"line\"><span>cmake</span><span> -B</span><span> build</span><span> -DGGML_BLAS=ON</span><span> -DGGML_BLAS_VENDOR=OpenBLAS</span></span>\n<span class=\"line\"><span>cmake</span><span> --build</span><span> build</span><span> --config</span><span> Release</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run inference</span></span>\n<span class=\"line\"><span>./build/bin/llama-cli</span><span> \\</span></span>\n<span class=\"line\"><span>    -m</span><span> llama-3.2-1b-q4_k_m.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    -p</span><span> \"Explain edge computing:\"</span><span> \\</span></span>\n<span class=\"line\"><span>    -n</span><span> 256</span></span></code></pre>\n<p>The BitNet B1.58 2B model achieves over 8 tokens/second with minimal RAM usage, making it well-suited for Pi deployment.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-44\" id=\"user-content-fnref-44-2\">44</a></sup></p>\n<p><strong>2026 updates</strong>: NVIDIA announced optimizations for llama.cpp at CES 2026 including NVFP4/FP8 quantization, GPU token sampling, and memory management improvements, enabling up to 35% faster token generation.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-32\" id=\"user-content-fnref-32-2\">32</a></sup> The ik_llama.cpp fork achieved 3x-4x speed improvements for multi-GPU configurations through tensor parallelism at the GGML graph level.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-46\" id=\"user-content-fnref-46\">46</a></sup></p>\n<h3 id=\"mlc-llm-on-mobile\">MLC LLM on mobile</h3>\n<p>MLC LLM compiles and optimizes LLMs for mobile GPUs using the TVM compiler stack. Unlike llama.cpp which targets CPUs, MLC LLM uses OpenCL, Vulkan, or Metal for GPU acceleration.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-47\" id=\"user-content-fnref-47\">47</a></sup></p>\n<p>The framework generates custom operators and runtime code for specific model architectures and target hardware. For Android, MLC LLM uses the OpenCL backend; for iOS, it targets Metal.</p>\n<p><strong>Android performance</strong>: On Snapdragon 8 Gen 2, models like Llama 3-4B run at 8-10 tokens/second. Mid-range devices with 6GB RAM struggle with models larger than 2B parameters.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-48\" id=\"user-content-fnref-48\">48</a></sup></p>\n<p><strong>MLC Chat</strong> supports models including Llama 3.2, Gemma 2, Phi 3.5, and Qwen 2.5, offering offline chat, translation, and multimodal tasks. NPU optimization works on Snapdragon 8 Gen 2 and newer.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-48\" id=\"user-content-fnref-48-2\">48</a></sup></p>\n<p>ExecuTorch combined with Unsloth’s quantization-aware training deploys Qwen3-0.6B on Pixel 8 and iPhone 15 Pro at approximately 40 tokens/second.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-49\" id=\"user-content-fnref-49\">49</a></sup></p>\n<hr />\n<h2 id=\"browser-inference\">Browser inference</h2>\n<p>WebGPU enables GPU-accelerated ML inference directly in web browsers without plugins or server calls.</p>\n<h3 id=\"webgpu-support\">WebGPU support</h3>\n<p>As of January 2026, WebGPU has reached 70% browser support across Firefox 147, Safari (iOS 26), and Chrome/Edge, with 65% of new apps already adopting it.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-50\" id=\"user-content-fnref-50\">50</a></sup> This marks the first year GPU compute works across all major browsers.</p>\n<h3 id=\"performance-characteristics\">Performance characteristics</h3>\n<p>WebGPU delivers 10x faster performance than WebGL for transformer models. Microsoft reports approximately 20x speedup over multi-threaded CPU and approximately 550x over single-threaded CPU for certain workloads.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-51\" id=\"user-content-fnref-51\">51</a></sup></p>\n<p>Browser AI inference via WebLLM now reaches 80% of native MLC-LLM performance.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-50\" id=\"user-content-fnref-50-2\">50</a></sup> Benchmarks show:</p>\n<ul>\n<li>Llama-3.1-8B: 41.1 tokens/second (71.2% native speed)</li>\n<li>Phi-3.5-mini: 71.1 tokens/second (79.6% native speed)</li>\n<li>Llama 3.2 models: up to 62 tokens/second in optimal conditions</li>\n<li>4-bit quantized 3B model: 90 tokens/second on Apple M3<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-52\" id=\"user-content-fnref-52\">52</a></sup></li>\n</ul>\n<p>On a laptop with NVIDIA RTX 3060, ONNX Runtime Web with WebGPU accelerates Segment Anything’s encoder by 19x and decoder by 3.8x.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-51\" id=\"user-content-fnref-51-2\">51</a></sup></p>\n<p>Model loading remains a UX challenge: DeepSeek-8B takes 2-3 minutes to download and initialize. Once loaded, inference is often faster than API round-trips.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-53\" id=\"user-content-fnref-53\">53</a></sup></p>\n<h3 id=\"transformersjs\">Transformers.js</h3>\n<p>Transformers.js brings Hugging Face pipelines to the browser. It uses ONNX Runtime Web for execution with WebGPU or WASM backends.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> { pipeline } </span><span>from</span><span> '@xenova/transformers'</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Load model with WebGPU backend</span></span>\n<span class=\"line\"><span>const</span><span> classifier</span><span> =</span><span> await</span><span> pipeline</span><span>(</span></span>\n<span class=\"line\"><span>  'text-classification'</span><span>,</span></span>\n<span class=\"line\"><span>  'Xenova/distilbert-base-uncased-finetuned-sst-2-english'</span><span>,</span></span>\n<span class=\"line\"><span>  { device</span><span>:</span><span> 'webgpu'</span><span>,</span><span> dtype</span><span>:</span><span> 'fp16'</span><span> }</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>const</span><span> result</span><span> =</span><span> await</span><span> classifier</span><span>(</span><span>'This movie was great!'</span><span>)</span></span></code></pre>\n<p>For quantization, Transformers.js supports fp32 (default for WebGPU), fp16, q8 (default for WASM), and q4 precisions.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-54\" id=\"user-content-fnref-54\">54</a></sup></p>\n<h3 id=\"webllm\">WebLLM</h3>\n<p>WebLLM is a high-performance in-browser LLM inference engine fully compatible with the OpenAI API, supporting streaming, JSON-mode, and function-calling.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-55\" id=\"user-content-fnref-55\">55</a></sup> It supports Llama 3, Phi-3, Gemma, and Mistral with WebGPU acceleration.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> *</span><span> as</span><span> webllm </span><span>from</span><span> '@mlc-ai/web-llm'</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>const</span><span> engine</span><span> =</span><span> await</span><span> webllm</span><span>.CreateMLCEngine</span><span>(</span><span>'Llama-3.2-1B-Instruct-q4f16_1-MLC'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>const</span><span> response</span><span> =</span><span> await</span><span> engine</span><span>.</span><span>chat</span><span>.</span><span>completions</span><span>.create</span><span>({</span></span>\n<span class=\"line\"><span>  messages</span><span>:</span><span> [{ role</span><span>:</span><span> 'user'</span><span>,</span><span> content</span><span>:</span><span> 'What is edge computing?'</span><span> }]</span><span>,</span></span>\n<span class=\"line\"><span>  stream</span><span>:</span><span> true</span></span>\n<span class=\"line\"><span>})</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> await</span><span> (</span><span>const</span><span> chunk</span><span> of</span><span> response) {</span></span>\n<span class=\"line\"><span>  console</span><span>.log</span><span>(</span><span>chunk</span><span>.choices[</span><span>0</span><span>]?.</span><span>delta</span><span>?.content </span><span>||</span><span> ''</span><span>)</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<hr />\n<h2 id=\"power-and-thermal-considerations\">Power and thermal considerations</h2>\n<p>Edge devices operate under power and thermal constraints that datacenter hardware ignores.</p>\n<h3 id=\"power-budgets\">Power budgets</h3>\n<p>Jetson Thor can run at 75-130W for maximum performance, while Jetson Orin Nano Super runs at up to 25W.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-5\" id=\"user-content-fnref-5-3\">5</a></sup><sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-10\" id=\"user-content-fnref-10-2\">10</a></sup> The new Jetson T4000 operates at 40-70W.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-8\" id=\"user-content-fnref-8-3\">8</a></sup></p>\n<p>The Hailo-10H in Raspberry Pi AI HAT+ 2 consumes a maximum of 3W while delivering 40 TOPS.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-13\" id=\"user-content-fnref-13-2\">13</a></sup> The Hailo-8L in AI HAT+ consumes 2.5-3W while delivering 13-26 TOPS.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-14\" id=\"user-content-fnref-14-2\">14</a></sup> Google Coral’s Edge TPU uses 2W for 4 TOPS.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-16\" id=\"user-content-fnref-16-2\">16</a></sup></p>\n<p>For battery-powered deployments, TOPS per watt determines runtime. Hailo-10H achieves over 13 TOPS/W; Hailo-8L achieves 3-4 TOPS/W; Coral achieves 2 TOPS/W; Jetson Thor achieves approximately 8-13 TOPS/W depending on power mode.</p>\n<h3 id=\"thermal-throttling\">Thermal throttling</h3>\n<p>When processors overheat, they reduce clock speeds to prevent damage. This thermal throttling drops inference throughput mid-operation.</p>\n<p>AI chips, particularly high-performance SoCs and NPUs, reduce clock speeds automatically when temperatures exceed thresholds. For latency-sensitive applications like real-time video analysis, throttling creates visible quality degradation.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-56\" id=\"user-content-fnref-56\">56</a></sup></p>\n<p>Mitigation strategies:</p>\n<ul>\n<li>Active cooling (fans) for continuous high-load operation</li>\n<li>Passive heatsinks for intermittent workloads</li>\n<li>Power mode selection matching thermal capacity</li>\n<li>Ambient temperature monitoring</li>\n</ul>\n<p>Jetson’s tegrastats utility reports CPU and GPU temperatures to help prevent throttling. Research on Jetson Nano demonstrated that proactive thermal management saves 9-12% average power compared to reactive built-in methods.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-57\" id=\"user-content-fnref-57\">57</a></sup></p>\n<h3 id=\"quantization-for-efficiency\">Quantization for efficiency</h3>\n<p>Lower precision reduces both compute and memory bandwidth, which reduces power consumption.</p>\n<p>Quantization from FP32 to INT8 shrinks memory footprint by 75%. Integer arithmetic requires less energy than floating point on most embedded processors. The combination produces cooler and longer-running edge deployments.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-58\" id=\"user-content-fnref-58\">58</a></sup></p>\n<hr />\n<h2 id=\"benchmark-object-detection-on-jetson\">Benchmark: object detection on Jetson</h2>\n<p>YOLOv8 object detection benchmarks illustrate real-world edge performance.</p>\n<h3 id=\"jetson-orin-nx-performance\">Jetson Orin NX performance</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Precision</th><th>FPS</th><th>Energy/Inference</th></tr></thead><tbody><tr><td>YOLOv8n</td><td>FP16</td><td>52</td><td>0.269 J</td></tr><tr><td>YOLOv8n</td><td>INT8</td><td>65</td><td>0.179 J</td></tr><tr><td>YOLOv8s</td><td>FP16</td><td>38</td><td>0.368 J</td></tr><tr><td>YOLOv8s</td><td>INT8</td><td>48</td><td>0.292 J</td></tr></tbody></table></div>\n<p>INT8 quantization delivers 25% higher FPS with 33% lower energy per inference.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-59\" id=\"user-content-fnref-59\">59</a></sup></p>\n<h3 id=\"jetson-orin-nano-performance\">Jetson Orin Nano performance</h3>\n<p>On the 4GB Orin Nano with TensorRT optimization:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Precision</th><th>Latency</th><th>FPS</th></tr></thead><tbody><tr><td>YOLOv8n</td><td>INT8</td><td>23.16 ms</td><td>~43</td></tr><tr><td>YOLOv8n</td><td>FP16</td><td>26.70 ms</td><td>~37</td></tr><tr><td>YOLOv8s</td><td>INT8</td><td>28.25 ms</td><td>~35</td></tr></tbody></table></div>\n<p>C++ implementation with TensorRT outperforms Python frameworks by reducing inference overhead.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-60\" id=\"user-content-fnref-60\">60</a></sup></p>\n<h3 id=\"jetson-thor-performance\">Jetson Thor performance</h3>\n<p>Jetson Thor delivers 3.5-4.9x performance gains over Orin in INT8 precision for 4K object detection in live video streams.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-61\" id=\"user-content-fnref-61\">61</a></sup> The platform supports decoding up to 10x 4Kp60 or 4x 8Kp30 video streams simultaneously.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-5\" id=\"user-content-fnref-5-4\">5</a></sup></p>\n<h3 id=\"optimization-workflow\">Optimization workflow</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Export YOLOv8 to TensorRT engine</span></span>\n<span class=\"line\"><span>yolo</span><span> export</span><span> model=yolov8n.pt</span><span> format=engine</span><span> device=</span><span>0</span><span> half=True</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run inference</span></span>\n<span class=\"line\"><span>yolo</span><span> predict</span><span> model=yolov8n.engine</span><span> source=video.mp4</span></span></code></pre>\n<hr />\n<h2 id=\"benchmark-llm-inference-across-devices\">Benchmark: LLM inference across devices</h2>\n<p>Text generation speed varies by hardware, model size, and quantization.</p>\n<h3 id=\"tokens-per-second-by-device-2026\">Tokens per second by device (2026)</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Device</th><th>Model</th><th>Quantization</th><th>Tokens/s</th></tr></thead><tbody><tr><td>Jetson Thor</td><td>Llama 3.1 8B</td><td>FP8</td><td>80-120</td></tr><tr><td>Jetson Orin Nano</td><td>Llama 3.2 3B</td><td>Q4_K_M</td><td>15-20</td></tr><tr><td>Raspberry Pi 5 8GB</td><td>Llama 3.2 1B</td><td>Q4_K_M</td><td>7+</td></tr><tr><td>Raspberry Pi 5 8GB</td><td>Llama 2 7B</td><td>Q4_K_M</td><td>0.7-3</td></tr><tr><td>iPhone 17 Pro (A19)</td><td>Qwen3 0.6B</td><td>QAT</td><td>~40</td></tr><tr><td>Snapdragon 8 Gen 2</td><td>Llama 3 4B</td><td>Q4</td><td>8-10</td></tr><tr><td>Browser (M3 laptop)</td><td>3B model</td><td>Q4</td><td>90</td></tr><tr><td>Browser (RTX 3060)</td><td>Phi-3 Mini</td><td>Q4</td><td>20-40</td></tr></tbody></table></div>\n<p>Source: Compiled from benchmark reports.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-43\" id=\"user-content-fnref-43-3\">43</a></sup><sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-44\" id=\"user-content-fnref-44-3\">44</a></sup><sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-48\" id=\"user-content-fnref-48-3\">48</a></sup><sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-49\" id=\"user-content-fnref-49-2\">49</a></sup><sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-52\" id=\"user-content-fnref-52-2\">52</a></sup></p>\n<h3 id=\"memory-requirements\">Memory requirements</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model Size</th><th>Q4_K_M Size</th><th>Minimum RAM</th></tr></thead><tbody><tr><td>1B</td><td>~0.6 GB</td><td>2 GB</td></tr><tr><td>3B</td><td>~1.8 GB</td><td>4 GB</td></tr><tr><td>7B</td><td>~4.6 GB</td><td>8 GB</td></tr><tr><td>13B</td><td>~7.8 GB</td><td>16 GB</td></tr></tbody></table></div>\n<p>Devices with less RAM than model size will page to disk, reducing throughput to near-unusable levels.</p>\n<hr />\n<h2 id=\"production-deployment-considerations\">Production deployment considerations</h2>\n<h3 id=\"ota-model-updates\">OTA model updates</h3>\n<p>Edge devices need remote update capability for model improvements and bug fixes.</p>\n<p>OTA (over-the-air) update platforms deploy firmware and model files to fleets of devices without physical access. The architecture includes:</p>\n<ol>\n<li>Model artifacts stored in cloud repositories</li>\n<li>Device agents checking for updates</li>\n<li>Secure download with cryptographic signing</li>\n<li>Atomic installation with rollback capability</li>\n</ol>\n<p>Golioth’s OTA system supports multi-part deployments including main firmware, cellular modem firmware, and ML models as separate artifacts.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-62\" id=\"user-content-fnref-62\">62</a></sup> Mender provides image-based updates that create identical environments across devices.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-63\" id=\"user-content-fnref-63\">63</a></sup></p>\n<p><strong>Security requirements</strong>:</p>\n<ul>\n<li>Cryptographic signing of all update packages</li>\n<li>Secure boot verification before installation</li>\n<li>Automatic rollback on failed updates</li>\n<li>TLS transport for all communications</li>\n</ul>\n<h3 id=\"monitoring-and-telemetry\">Monitoring and telemetry</h3>\n<p>Remote monitoring tracks inference latency, throughput, error rates, and hardware health.</p>\n<p>Key metrics to collect:</p>\n<ul>\n<li>Inference latency (p50, p95, p99)</li>\n<li>Throughput (inferences/second)</li>\n<li>Model accuracy on validation samples</li>\n<li>Hardware temperature and power draw</li>\n<li>Memory utilization</li>\n</ul>\n<p>Advantech DeviceOn provides OTA updates combined with container management for edge AI deployments.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-64\" id=\"user-content-fnref-64\">64</a></sup> ThingsBoard supports firmware distribution with version tracking since version 3.3.<sup><a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fn-65\" id=\"user-content-fnref-65\">65</a></sup></p>\n<h3 id=\"fallback-strategies\">Fallback strategies</h3>\n<p>Edge deployments need graceful degradation when primary inference fails.</p>\n<p><strong>Model fallback</strong>: Switch to smaller, faster models when latency budgets are missed. A vision system might drop from YOLOv8m to YOLOv8n under thermal throttling.</p>\n<p><strong>Cloud fallback</strong>: Route requests to cloud inference when device capacity is exceeded or models require updates not yet deployed.</p>\n<p><strong>Cached responses</strong>: Return cached predictions for repeated inputs. This works for classification tasks with limited input variety.</p>\n<p><strong>Graceful denial</strong>: When no inference is possible, return explicit “unknown” responses rather than failing silently.</p>\n<hr />\n<h2 id=\"complete-example-browser-chatbot\">Complete example: browser chatbot</h2>\n<p>A functional browser-based chatbot using WebLLM and WebGPU.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>&lt;!</span><span>DOCTYPE</span><span> html</span><span>&gt;</span></span>\n<span class=\"line\"><span>&lt;</span><span>html</span><span>&gt;</span></span>\n<span class=\"line\"><span>  &lt;</span><span>head</span><span>&gt;</span></span>\n<span class=\"line\"><span>    &lt;</span><span>title</span><span>&gt;Edge LLM Chat&lt;/</span><span>title</span><span>&gt;</span></span>\n<span class=\"line\"><span>    &lt;</span><span>script</span><span> type</span><span>=</span><span>\"module\"</span><span>&gt;</span></span>\n<span class=\"line\"><span>      import</span><span> *</span><span> as</span><span> webllm </span><span>from</span><span> 'https://cdn.jsdelivr.net/npm/@mlc-ai/web-llm@latest'</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>      let</span><span> engine</span></span>\n<span class=\"line\"><span>      const</span><span> statusEl</span><span> =</span><span> document</span><span>.getElementById</span><span>(</span><span>'status'</span><span>)</span></span>\n<span class=\"line\"><span>      const</span><span> chatEl</span><span> =</span><span> document</span><span>.getElementById</span><span>(</span><span>'chat'</span><span>)</span></span>\n<span class=\"line\"><span>      const</span><span> inputEl</span><span> =</span><span> document</span><span>.getElementById</span><span>(</span><span>'input'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>      async</span><span> function</span><span> init</span><span>() {</span></span>\n<span class=\"line\"><span>        statusEl</span><span>.textContent </span><span>=</span><span> 'Loading model...'</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        engine </span><span>=</span><span> await</span><span> webllm</span><span>.CreateMLCEngine</span><span>(</span><span>'Llama-3.2-1B-Instruct-q4f16_1-MLC'</span><span>,</span><span> {</span></span>\n<span class=\"line\"><span>          initProgressCallback</span><span>:</span><span> (progress) </span><span>=&gt;</span><span> {</span></span>\n<span class=\"line\"><span>            statusEl</span><span>.textContent </span><span>=</span><span> `Loading: </span><span>${</span><span>Math</span><span>.round</span><span>(</span><span>progress</span><span>.progress </span><span>*</span><span> 100</span><span>)</span><span>}</span><span>%`</span></span>\n<span class=\"line\"><span>          }</span></span>\n<span class=\"line\"><span>        })</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        statusEl</span><span>.textContent </span><span>=</span><span> 'Ready'</span></span>\n<span class=\"line\"><span>      }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>      async</span><span> function</span><span> generate</span><span>() {</span></span>\n<span class=\"line\"><span>        const</span><span> userMessage</span><span> =</span><span> inputEl</span><span>.</span><span>value</span><span>.trim</span><span>()</span></span>\n<span class=\"line\"><span>        if</span><span> (</span><span>!</span><span>userMessage </span><span>||</span><span> !</span><span>engine) </span><span>return</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        chatEl</span><span>.innerHTML </span><span>+=</span><span> `&lt;div class=\"user\"&gt;User: </span><span>${</span><span>userMessage</span><span>}</span><span>&lt;/div&gt;`</span></span>\n<span class=\"line\"><span>        inputEl</span><span>.value </span><span>=</span><span> ''</span></span>\n<span class=\"line\"><span>        statusEl</span><span>.textContent </span><span>=</span><span> 'Generating...'</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        const</span><span> response</span><span> =</span><span> await</span><span> engine</span><span>.</span><span>chat</span><span>.</span><span>completions</span><span>.create</span><span>({</span></span>\n<span class=\"line\"><span>          messages</span><span>:</span><span> [{ role</span><span>:</span><span> 'user'</span><span>,</span><span> content</span><span>:</span><span> userMessage }]</span><span>,</span></span>\n<span class=\"line\"><span>          stream</span><span>:</span><span> true</span></span>\n<span class=\"line\"><span>        })</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        let</span><span> assistantMessage </span><span>=</span><span> ''</span></span>\n<span class=\"line\"><span>        chatEl</span><span>.innerHTML </span><span>+=</span><span> `&lt;div class=\"assistant\"&gt;Assistant: &lt;span id=\"response\"&gt;&lt;/span&gt;&lt;/div&gt;`</span></span>\n<span class=\"line\"><span>        const</span><span> responseEl</span><span> =</span><span> document</span><span>.getElementById</span><span>(</span><span>'response'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        for</span><span> await</span><span> (</span><span>const</span><span> chunk</span><span> of</span><span> response) {</span></span>\n<span class=\"line\"><span>          const</span><span> content</span><span> =</span><span> chunk</span><span>.choices[</span><span>0</span><span>]?.</span><span>delta</span><span>?.content </span><span>||</span><span> ''</span></span>\n<span class=\"line\"><span>          assistantMessage </span><span>+=</span><span> content</span></span>\n<span class=\"line\"><span>          responseEl</span><span>.textContent </span><span>=</span><span> assistantMessage</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        statusEl</span><span>.textContent </span><span>=</span><span> 'Ready'</span></span>\n<span class=\"line\"><span>      }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>      document</span><span>.getElementById</span><span>(</span><span>'send'</span><span>)</span><span>.addEventListener</span><span>(</span><span>'click'</span><span>,</span><span> generate)</span></span>\n<span class=\"line\"><span>      inputEl</span><span>.addEventListener</span><span>(</span><span>'keypress'</span><span>,</span><span> (e) </span><span>=&gt;</span><span> {</span></span>\n<span class=\"line\"><span>        if</span><span> (</span><span>e</span><span>.key </span><span>===</span><span> 'Enter'</span><span>) </span><span>generate</span><span>()</span></span>\n<span class=\"line\"><span>      })</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>      init</span><span>()</span></span>\n<span class=\"line\"><span>    &lt;/</span><span>script</span><span>&gt;</span></span>\n<span class=\"line\"><span>  &lt;/</span><span>head</span><span>&gt;</span></span>\n<span class=\"line\"><span>  &lt;</span><span>body</span><span>&gt;</span></span>\n<span class=\"line\"><span>    &lt;</span><span>div</span><span> id</span><span>=</span><span>\"status\"</span><span>&gt;Initializing...&lt;/</span><span>div</span><span>&gt;</span></span>\n<span class=\"line\"><span>    &lt;</span><span>div</span><span> id</span><span>=</span><span>\"chat\"</span><span>&gt;&lt;/</span><span>div</span><span>&gt;</span></span>\n<span class=\"line\"><span>    &lt;</span><span>input</span><span> id</span><span>=</span><span>\"input\"</span><span> type</span><span>=</span><span>\"text\"</span><span> placeholder</span><span>=</span><span>\"Type a message...\"</span><span> /&gt;</span></span>\n<span class=\"line\"><span>    &lt;</span><span>button</span><span> id</span><span>=</span><span>\"send\"</span><span>&gt;Send&lt;/</span><span>button</span><span>&gt;</span></span>\n<span class=\"line\"><span>  &lt;/</span><span>body</span><span>&gt;</span></span>\n<span class=\"line\"><span>&lt;/</span><span>html</span><span>&gt;</span></span></code></pre>\n<p>This loads a quantized Llama 3.2 1B model entirely in the browser. No server required after initial page load. WebLLM provides OpenAI-compatible API with streaming support.</p>\n<hr />\n<h2 id=\"framework-selection-guide\">Framework selection guide</h2>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Use Case</th><th>Recommended Framework</th><th>Hardware</th></tr></thead><tbody><tr><td>Robotics, autonomous systems</td><td>TensorRT + JetPack 7</td><td>Jetson Thor/Orin</td></tr><tr><td>iOS/macOS apps</td><td>Core ML</td><td>Apple M5/A19</td></tr><tr><td>Android apps</td><td>TensorFlow Lite, MLC LLM</td><td>Snapdragon 8 Elite NPU</td></tr><tr><td>Raspberry Pi vision</td><td>HailoRT, TensorFlow Lite</td><td>Pi 5 + AI HAT+</td></tr><tr><td>Raspberry Pi LLM</td><td>llama.cpp, HailoRT</td><td>Pi 5 + AI HAT+ 2</td></tr><tr><td>Browser deployment</td><td>WebLLM, Transformers.js</td><td>WebGPU</td></tr><tr><td>Cross-platform LLM</td><td>llama.cpp</td><td>Any CPU/GPU</td></tr><tr><td>Microcontrollers</td><td>TensorFlow Lite Micro</td><td>Cortex-M</td></tr><tr><td>Ultra-low power IoT</td><td>Nordic Edge AI Lab</td><td>nRF54L + Axon NPU</td></tr></tbody></table></div>\n<hr />\n<h2 id=\"summary\">Summary</h2>\n<p>Edge AI deployment trades cloud flexibility for latency, privacy, and offline capability. The hardware landscape in 2026 offers options from <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>70</mn><mi>U</mi><mi>S</mi><mi>B</mi><mi>a</mi><mi>c</mi><mi>c</mi><mi>e</mi><mi>l</mi><mi>e</mi><mi>r</mi><mi>a</mi><mi>t</mi><mi>o</mi><mi>r</mi><mi>s</mi><mi>t</mi><mi>o</mi></mrow></semantics></math>3,499 development kits capable of running multi-AI workflows.</p>\n<p>The Jetson Thor platform launched in January 2026 delivers over 1000 TOPS with 128GB memory, enabling on-device generative AI for robotics applications. The Raspberry Pi AI HAT+ 2 brings LLM inference to the $130 price point, though practical performance remains limited to small models.</p>\n<p>Mobile NPUs have matured significantly. Apple’s M5 chip delivers 3.5x the AI performance of M4, while Qualcomm’s Snapdragon 8 Elite achieves 100x speedups over CPU for vision-language models. WebGPU has reached 70% browser coverage, with WebLLM achieving 80% of native inference performance.</p>\n<p>Model optimization through quantization, pruning, and distillation makes deployment feasible. Q4 quantization reduces 7B parameter models to under 5GB while preserving most capability. New NVFP4 and FP8 formats provide additional optimization paths on supported hardware.</p>\n<p>Production deployment requires more than inference code. OTA updates, monitoring, thermal management, and fallback strategies separate demos from reliable systems.</p>\n<p>The direction is clear: AI inference is moving closer to data sources. Edge deployment skills will matter more as this trend continues.</p>\n<hr />\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>InfoQ. “Cactus v1: Cross-Platform LLM Inference on Mobile with Zero Latency and Full Privacy.” <a href=\"https://www.infoq.com/news/2025/12/cactus-on-device-inference/\" rel=\"noopener noreferrer\">https://www.infoq.com/news/2025/12/cactus-on-device-inference/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>PMC. “Tiny Machine Learning and On-Device Inference: A Survey.” <a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC12115890/\" rel=\"noopener noreferrer\">https://pmc.ncbi.nlm.nih.gov/articles/PMC12115890/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Novus. “The Rise of Local AI Models: Going Small to Go Big.” <a href=\"https://www.novusasi.com/blog/the-rise-of-local-ai-models-going-small-to-go-big\" rel=\"noopener noreferrer\">https://www.novusasi.com/blog/the-rise-of-local-ai-models-going-small-to-go-big</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>GlobeNewswire. “Edge AI Hardware Markets, 2031: Rising AI Demand Spurs Smartphone Refresh Cycles in the Premium Segment.” <a href=\"https://www.globenewswire.com/news-release/2026/01/21/3222516/28124/en/Edge-AI-Hardware-Markets-2031-Rising-AI-Demand-Spurs-Smartphone-Refresh-Cycles-in-the-Premium-Segment.html\" rel=\"noopener noreferrer\">https://www.globenewswire.com/news-release/2026/01/21/3222516/28124/en/Edge-AI-Hardware-Markets-2031-Rising-AI-Demand-Spurs-Smartphone-Refresh-Cycles-in-the-Premium-Segment.html</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-4-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-4-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>NVIDIA. “Jetson Thor | Advanced AI for Physical Robotics.” <a href=\"https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/\" rel=\"noopener noreferrer\">https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-5-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-5-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-5-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p>NVIDIA Newsroom. “NVIDIA Blackwell-Powered Jetson Thor Now Available, Accelerating the Age of General Robotics.” <a href=\"https://nvidianews.nvidia.com/news/nvidia-blackwell-powered-jetson-thor-now-available-accelerating-the-age-of-general-robotics\" rel=\"noopener noreferrer\">https://nvidianews.nvidia.com/news/nvidia-blackwell-powered-jetson-thor-now-available-accelerating-the-age-of-general-robotics</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-6-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p>Seeed Studio. “NVIDIA Jetson AGX Thor Developer Kit.” <a href=\"https://www.seeedstudio.com/NVIDIA-Jetson-AGX-Thor-Developer-Kit-p-9965.html\" rel=\"noopener noreferrer\">https://www.seeedstudio.com/NVIDIA-Jetson-AGX-Thor-Developer-Kit-p-9965.html</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p>SDxCentral. “Nvidia pushes AI from edge to storage with Jetson T4000 and BlueField-4 updates.” <a href=\"https://www.sdxcentral.com/news/nvidia-pushes-ai-from-edge-to-storage-with-jetson-t4000-and-bluefield-4-updates/\" rel=\"noopener noreferrer\">https://www.sdxcentral.com/news/nvidia-pushes-ai-from-edge-to-storage-with-jetson-t4000-and-bluefield-4-updates/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-8-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-8-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-9\">\n<p>NVIDIA. “Jetson AGX Orin for Next-Gen Robotics.” <a href=\"https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/\" rel=\"noopener noreferrer\">https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-9\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-9-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-10\">\n<p>NVIDIA Developer Blog. “NVIDIA Jetson Orin Nano Developer Kit Gets a Super Boost.” <a href=\"https://developer.nvidia.com/blog/nvidia-jetson-orin-nano-developer-kit-gets-a-super-boost/\" rel=\"noopener noreferrer\">https://developer.nvidia.com/blog/nvidia-jetson-orin-nano-developer-kit-gets-a-super-boost/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-10\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-10-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p>Raspberry Pi. “Introducing the Raspberry Pi AI HAT+ 2: Generative AI on Raspberry Pi 5.” <a href=\"https://www.raspberrypi.com/news/introducing-the-raspberry-pi-ai-hat-plus-2-generative-ai-on-raspberry-pi-5/\" rel=\"noopener noreferrer\">https://www.raspberrypi.com/news/introducing-the-raspberry-pi-ai-hat-plus-2-generative-ai-on-raspberry-pi-5/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-11-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-12\">\n<p>CNX Software. “Raspberry Pi AI HAT+ 2 targets generative AI (LLM/VLM) with Hailo-10H accelerator.” <a href=\"https://www.cnx-software.com/2026/01/15/raspberry-pi-ai-hat-2-targets-generative-ai-llm-vlm-with-hailo-10h-accelerator/\" rel=\"noopener noreferrer\">https://www.cnx-software.com/2026/01/15/raspberry-pi-ai-hat-2-targets-generative-ai-llm-vlm-with-hailo-10h-accelerator/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-12\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-13\">\n<p>Electronics Weekly. “Raspberry Pi AI HAT+ 2 updates for Generative AI.” <a href=\"https://www.electronicsweekly.com/news/products/raspberry-pi-development/raspberry-pi-ai-hat-2-updates-for-generative-ai-2026-01/\" rel=\"noopener noreferrer\">https://www.electronicsweekly.com/news/products/raspberry-pi-development/raspberry-pi-ai-hat-2-updates-for-generative-ai-2026-01/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-13\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-13-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p>Jeff Geerling. “Testing Raspberry Pi’s AI Kit - 13 TOPS for $70.” <a href=\"https://www.jeffgeerling.com/blog/2024/testing-raspberry-pis-ai-kit-13-tops-70/\" rel=\"noopener noreferrer\">https://www.jeffgeerling.com/blog/2024/testing-raspberry-pis-ai-kit-13-tops-70/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-14-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p>Jeff Geerling. “Raspberry Pi’s new AI HAT adds 8GB of RAM for local LLMs.” <a href=\"https://www.jeffgeerling.com/blog/2026/raspberry-pi-ai-hat-2/\" rel=\"noopener noreferrer\">https://www.jeffgeerling.com/blog/2026/raspberry-pi-ai-hat-2/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-16\">\n<p>Google Coral. “USB Accelerator Datasheet.” <a href=\"https://www.coral.ai/static/files/Coral-USB-Accelerator-datasheet.pdf\" rel=\"noopener noreferrer\">https://www.coral.ai/static/files/Coral-USB-Accelerator-datasheet.pdf</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-16\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-16-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-17\">\n<p>Q-engineering. “Google Coral’s TPU explained in depth.” <a href=\"https://qengineering.eu/google-corals-tpu-explained.html\" rel=\"noopener noreferrer\">https://qengineering.eu/google-corals-tpu-explained.html</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-17\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-18\">\n<p>Apple. “Apple unleashes M5, the next big leap in AI performance for Apple silicon.” <a href=\"https://www.apple.com/newsroom/2025/10/apple-unleashes-m5-the-next-big-leap-in-ai-performance-for-apple-silicon/\" rel=\"noopener noreferrer\">https://www.apple.com/newsroom/2025/10/apple-unleashes-m5-the-next-big-leap-in-ai-performance-for-apple-silicon/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-18\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-19\">\n<p>Apple. “Apple introduces M4 chip.” <a href=\"https://www.apple.com/newsroom/2024/05/apple-introduces-m4-chip/\" rel=\"noopener noreferrer\">https://www.apple.com/newsroom/2024/05/apple-introduces-m4-chip/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-19\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-20\">\n<p>MacRumors. “A19 vs. A19 Pro: iPhone 17 Chip Differences.” <a href=\"https://www.macrumors.com/2025/09/09/iphone-17-a19-chip/\" rel=\"noopener noreferrer\">https://www.macrumors.com/2025/09/09/iphone-17-a19-chip/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-20\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-21\">\n<p>Apple Machine Learning Research. “Deploying Transformers on the Apple Neural Engine.” <a href=\"https://machinelearning.apple.com/research/neural-engine-transformers\" rel=\"noopener noreferrer\">https://machinelearning.apple.com/research/neural-engine-transformers</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-21\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-22\">\n<p>Google Developers Blog. “Unlocking Peak Performance on Qualcomm NPU with LiteRT.” <a href=\"https://developers.googleblog.com/unlocking-peak-performance-on-qualcomm-npu-with-litert/\" rel=\"noopener noreferrer\">https://developers.googleblog.com/unlocking-peak-performance-on-qualcomm-npu-with-litert/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-22\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-22-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-22-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-23\">\n<p>arXiv. “Large Language Model Performance Benchmarking on Mobile Platforms.” <a href=\"https://arxiv.org/html/2410.03613v1\" rel=\"noopener noreferrer\">https://arxiv.org/html/2410.03613v1</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-23\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-24\">\n<p>arXiv. “Scaling LLM Test-Time Compute with Mobile NPU on Smartphones.” <a href=\"https://arxiv.org/html/2509.23324v1\" rel=\"noopener noreferrer\">https://arxiv.org/html/2509.23324v1</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-24\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-25\">\n<p>Automate.org. “CES 2026: Qualcomm Targets NVIDIA Jetson with New Robotics Developer Platform.” <a href=\"https://www.automate.org/news/ces-2026-qualcomm-targets-nvidia-jetson-with-new-robotics-developer-platform\" rel=\"noopener noreferrer\">https://www.automate.org/news/ces-2026-qualcomm-targets-nvidia-jetson-with-new-robotics-developer-platform</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-25\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-26\">\n<p>Nordic Semiconductor. “nRF54L Series SoC with NPU and Nordic Edge AI Lab make on-device intelligence easily accessible.” <a href=\"https://www.nordicsemi.com/Nordic-news/2026/01/nRF54L-Series-SoC-with-NPU-and-Nordic-Edge-AI-Lab-make-on-device-intelligence-easily-accessible\" rel=\"noopener noreferrer\">https://www.nordicsemi.com/Nordic-news/2026/01/nRF54L-Series-SoC-with-NPU-and-Nordic-Edge-AI-Lab-make-on-device-intelligence-easily-accessible</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-26\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-26-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-27\">\n<p>Unified AI Hub. “Edge AI in 2026: Processing Intelligence at the Edge.” <a href=\"https://www.unifiedaihub.com/blog/edge-ai-in-2026-processing-intelligence-where-data-is-generated\" rel=\"noopener noreferrer\">https://www.unifiedaihub.com/blog/edge-ai-in-2026-processing-intelligence-where-data-is-generated</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-27\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-27-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-28\">\n<p>Google AI Edge. “Post-training quantization.” <a href=\"https://ai.google.dev/edge/litert/conversion/tensorflow/quantization/post_training_quantization\" rel=\"noopener noreferrer\">https://ai.google.dev/edge/litert/conversion/tensorflow/quantization/post_training_quantization</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-28\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-28-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-29\">\n<p>PyTorch Blog. “Quantization-Aware Training for Large Language Models.” <a href=\"https://pytorch.org/blog/quantization-aware-training/\" rel=\"noopener noreferrer\">https://pytorch.org/blog/quantization-aware-training/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-29\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-30\">\n<p>llama.cpp GitHub. “Quantize README.” <a href=\"https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md\" rel=\"noopener noreferrer\">https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-30\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-30-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-31\">\n<p>Oreate AI. “Practical Quantization of Llama Models: Detailed Explanation of GGUF and llama.cpp Technologies.” <a href=\"https://www.oreateai.com/blog/practical-quantization-of-llama-models-detailed-explanation-of-gguf-and-llamacpp-technologies/\" rel=\"noopener noreferrer\">https://www.oreateai.com/blog/practical-quantization-of-llama-models-detailed-explanation-of-gguf-and-llamacpp-technologies/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-31\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-32\">\n<p>NVIDIA Developer Blog. “Open Source AI Tool Upgrades Speed Up LLM and Diffusion Models on NVIDIA RTX PCs.” <a href=\"https://developer.nvidia.com/blog/open-source-ai-tool-upgrades-speed-up-llm-and-diffusion-models-on-nvidia-rtx-pcs\" rel=\"noopener noreferrer\">https://developer.nvidia.com/blog/open-source-ai-tool-upgrades-speed-up-llm-and-diffusion-models-on-nvidia-rtx-pcs</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-32\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-32-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-33\">\n<p>Nature Scientific Reports. “Neural network compression for reinforcement learning tasks.” <a href=\"https://www.nature.com/articles/s41598-025-93955-w\" rel=\"noopener noreferrer\">https://www.nature.com/articles/s41598-025-93955-w</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-33\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-34\">\n<p>Nature Scientific Reports. “Efficient self-attention with smart pruning for sustainable large language models.” <a href=\"https://www.nature.com/articles/s41598-025-92586-5\" rel=\"noopener noreferrer\">https://www.nature.com/articles/s41598-025-92586-5</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-34\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-35\">\n<p>Springer. “A Hybrid Lightweight Deep Learning Model for Edge Devices.” <a href=\"https://link.springer.com/chapter/10.1007/978-3-031-81083-1_1\" rel=\"noopener noreferrer\">https://link.springer.com/chapter/10.1007/978-3-031-81083-1_1</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-35\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-36\">\n<p>Wiley. “Optimizing Deep Learning Models for Resource-Constrained Environments.” <a href=\"https://onlinelibrary.wiley.com/doi/full/10.1002/eng2.70187\" rel=\"noopener noreferrer\">https://onlinelibrary.wiley.com/doi/full/10.1002/eng2.70187</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-36\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-37\">\n<p>NVIDIA Developer Blog. “Optimizing Inference on LLMs with TensorRT-LLM.” <a href=\"https://developer.nvidia.com/blog/optimizing-inference-on-llms-with-tensorrt-llm-now-publicly-available/\" rel=\"noopener noreferrer\">https://developer.nvidia.com/blog/optimizing-inference-on-llms-with-tensorrt-llm-now-publicly-available/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-37\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-38\">\n<p>Preste AI. “Optimized Deep Learning using TensorRT for NVIDIA Jetson TX2.” <a href=\"https://www.preste.ai/post/optimized-deep-learning-using-tensorrt-for-nvidia-jetson-tx2-part-1\" rel=\"noopener noreferrer\">https://www.preste.ai/post/optimized-deep-learning-using-tensorrt-for-nvidia-jetson-tx2-part-1</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-38\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-39\">\n<p>arXiv. “Benchmarking Deep Learning Models on NVIDIA Jetson Nano.” <a href=\"https://arxiv.org/html/2406.17749v1\" rel=\"noopener noreferrer\">https://arxiv.org/html/2406.17749v1</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-39\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-40\">\n<p>Apple Developer. “Deploy machine learning and AI models on-device with Core ML - WWDC24.” <a href=\"https://developer.apple.com/videos/play/wwdc2024/10161/\" rel=\"noopener noreferrer\">https://developer.apple.com/videos/play/wwdc2024/10161/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-40\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-41\">\n<p>PyTorch ExecuTorch. “Core ML Backend.” <a href=\"https://docs.pytorch.org/executorch/0.7/backends-coreml.html\" rel=\"noopener noreferrer\">https://docs.pytorch.org/executorch/0.7/backends-coreml.html</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-41\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-42\">\n<p>Red Hat Developer. “vLLM or llama.cpp: Choosing the right LLM inference engine for your use case.” <a href=\"https://developers.redhat.com/articles/2025/09/30/vllm-or-llamacpp-choosing-right-llm-inference-engine-your-use-case\" rel=\"noopener noreferrer\">https://developers.redhat.com/articles/2025/09/30/vllm-or-llamacpp-choosing-right-llm-inference-engine-your-use-case</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-42\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-43\">\n<p>AI Competence. “Running Llama On Raspberry Pi 5 (2025 Setup &amp; Guide).” <a href=\"https://aicompetence.org/running-llama-on-raspberry-pi-5/\" rel=\"noopener noreferrer\">https://aicompetence.org/running-llama-on-raspberry-pi-5/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-43\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-43-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-43-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-44\">\n<p>Stratosphere Laboratory. “How Well Do LLMs Perform on a Raspberry Pi 5?” <a href=\"https://www.stratosphereips.org/blog/2025/6/5/how-well-do-llms-perform-on-a-raspberry-pi-5\" rel=\"noopener noreferrer\">https://www.stratosphereips.org/blog/2025/6/5/how-well-do-llms-perform-on-a-raspberry-pi-5</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-44\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-44-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-44-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-45\">\n<p>Medium. “Local LLM eval tokens/sec comparison between llama.cpp and llamafile on Raspberry Pi 5.” <a href=\"https://medium.com/aidatatools/local-llm-eval-tokens-sec-comparison-between-llama-cpp-and-llamafile-on-raspberry-pi-5-8gb-model-89cfa17f6f18\" rel=\"noopener noreferrer\">https://medium.com/aidatatools/local-llm-eval-tokens-sec-comparison-between-llama-cpp-and-llamafile-on-raspberry-pi-5-8gb-model-89cfa17f6f18</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-45\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-46\">\n<p>Medium. “llama.cpp performance breakthrough for multi-GPU setups.” <a href=\"https://medium.com/@jagusztinl/llama-cpp-performance-breakthrough-for-multi-gpu-setups-04c83a66feb2\" rel=\"noopener noreferrer\">https://medium.com/@jagusztinl/llama-cpp-performance-breakthrough-for-multi-gpu-setups-04c83a66feb2</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-46\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-47\">\n<p>Callstack. “Want to Run LLMs on Your Device? Meet MLC.” <a href=\"https://www.callstack.com/blog/want-to-run-llms-on-your-device-meet-mlc\" rel=\"noopener noreferrer\">https://www.callstack.com/blog/want-to-run-llms-on-your-device-meet-mlc</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-47\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-48\">\n<p>It’s FOSS. “I Ran Local LLMs on My Android Phone.” <a href=\"https://itsfoss.com/android-on-device-ai/\" rel=\"noopener noreferrer\">https://itsfoss.com/android-on-device-ai/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-48\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-48-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-48-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-49\">\n<p>Unsloth. “How to Run and Deploy LLMs on your iOS or Android Phone.” <a href=\"https://unsloth.ai/docs/basics/deploy-llms-phone\" rel=\"noopener noreferrer\">https://unsloth.ai/docs/basics/deploy-llms-phone</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-49\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-49-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-50\">\n<p>byteiota. “WebGPU 2026: 70% Browser Support, 15x Performance Gains.” <a href=\"https://byteiota.com/webgpu-2026-70-browser-support-15x-performance-gains/\" rel=\"noopener noreferrer\">https://byteiota.com/webgpu-2026-70-browser-support-15x-performance-gains/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-50\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-50-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-51\">\n<p>Microsoft Open Source Blog. “ONNX Runtime Web unleashes generative AI in the browser using WebGPU.” <a href=\"https://opensource.microsoft.com/blog/2024/02/29/onnx-runtime-web-unleashes-generative-ai-in-the-browser-using-webgpu\" rel=\"noopener noreferrer\">https://opensource.microsoft.com/blog/2024/02/29/onnx-runtime-web-unleashes-generative-ai-in-the-browser-using-webgpu</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-51\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-51-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-52\">\n<p>arXiv. “WebLLM: A High-Performance In-Browser LLM Inference Engine.” <a href=\"https://arxiv.org/html/2412.15803v1\" rel=\"noopener noreferrer\">https://arxiv.org/html/2412.15803v1</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-52\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-52-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-53\">\n<p>Medium. “WebGPU bugs are holding back the browser AI revolution.” <a href=\"https://medium.com/@marcelo.emmerich/webgpu-bugs-are-holding-back-the-browser-ai-revolution-27d5f8c1dfca\" rel=\"noopener noreferrer\">https://medium.com/@marcelo.emmerich/webgpu-bugs-are-holding-back-the-browser-ai-revolution-27d5f8c1dfca</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-53\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-54\">\n<p>Hugging Face Blog. “Transformers.js v3: WebGPU Support, New Models &amp; Tasks, and More.” <a href=\"https://huggingface.co/blog/transformersjs-v3\" rel=\"noopener noreferrer\">https://huggingface.co/blog/transformersjs-v3</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-54\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-55\">\n<p>WebLLM. “Home.” <a href=\"https://webllm.mlc.ai/\" rel=\"noopener noreferrer\">https://webllm.mlc.ai/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-55\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-56\">\n<p>Embedded.com. “Optimizing Edge AI with Advanced Thermal Management in Embedded Systems.” <a href=\"https://www.embedded.com/optimizing-edge-ai-with-advanced-thermal-management-in-embedded-systems-2/\" rel=\"noopener noreferrer\">https://www.embedded.com/optimizing-edge-ai-with-advanced-thermal-management-in-embedded-systems-2/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-56\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-57\">\n<p>IEEE Xplore. “Run-Time Prevention of Thermal Throttling on the Edge using Reinforcement-Learning Based Predictive Thermal Aware Power and Performance Management.” <a href=\"https://ieeexplore.ieee.org/document/10666109/\" rel=\"noopener noreferrer\">https://ieeexplore.ieee.org/document/10666109/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-57\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-58\">\n<p>Janea Systems. “4 Power Management Strategies for Edge AI Devices.” <a href=\"https://www.janeasystems.com/blog/power-management-strategies-for-edge-devices\" rel=\"noopener noreferrer\">https://www.janeasystems.com/blog/power-management-strategies-for-edge-devices</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-58\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-59\">\n<p>Seeed Studio Blog. “YOLOv8 Performance Benchmarks on NVIDIA Jetson Devices.” <a href=\"https://www.seeedstudio.com/blog/2023/03/30/yolov8-performance-benchmarks-on-nvidia-jetson-devices/\" rel=\"noopener noreferrer\">https://www.seeedstudio.com/blog/2023/03/30/yolov8-performance-benchmarks-on-nvidia-jetson-devices/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-59\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-60\">\n<p>Hackster.io. “Pushing Limits: YOLOv8 vs. v26 on Jetson Orin Nano.” <a href=\"https://www.hackster.io/qwe018931/pushing-limits-yolov8-vs-v26-on-jetson-orin-nano-b89267\" rel=\"noopener noreferrer\">https://www.hackster.io/qwe018931/pushing-limits-yolov8-vs-v26-on-jetson-orin-nano-b89267</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-60\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-61\">\n<p>Simalabs. “Jetson AGX Thor vs. Jetson Orin: Latency-Accuracy Benchmarks for 4K Object Detection.” <a href=\"https://www.simalabs.ai/resources/jetson-agx-thor-vs-orin-4k-object-detection-live-sports-benchmarks-2025\" rel=\"noopener noreferrer\">https://www.simalabs.ai/resources/jetson-agx-thor-vs-orin-4k-object-detection-live-sports-benchmarks-2025</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-61\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-62\">\n<p>Golioth. “Over-the-Air (OTA) Updates.” <a href=\"https://docs.golioth.io/device-management/ota/\" rel=\"noopener noreferrer\">https://docs.golioth.io/device-management/ota/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-62\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-63\">\n<p>Mender. “Over-the-air (OTA) update best practices for industrial IoT and embedded devices.” <a href=\"https://mender.io/resources/reports-and-guides/ota-updates-best-practices\" rel=\"noopener noreferrer\">https://mender.io/resources/reports-and-guides/ota-updates-best-practices</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-63\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-64\">\n<p>Advantech. “AIoT Device Management and Edge Orchestration - DeviceOn.” <a href=\"https://campaign.advantech.online/en/deviceon/index.html\" rel=\"noopener noreferrer\">https://campaign.advantech.online/en/deviceon/index.html</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-64\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-65\">\n<p>ThingsBoard. “Over-the-air firmware and software updates.” <a href=\"https://thingsboard.io/docs/user-guide/ota-updates/\" rel=\"noopener noreferrer\">https://thingsboard.io/docs/user-guide/ota-updates/</a> <a href=\"https://blog.ecitis.org/edge-ai-deployment/#user-content-fnref-65\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/edge-ai-deployment/",
            "title": "Edge AI Deployment: Running Models on Constrained Devices",
            "summary": "Choose hardware, runtimes, and optimization strategies for private, low-latency inference on phones, embedded systems, and edge devices.",
            "image": "https://blog.ecitis.org/open-graph/edge-ai-deployment.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "Edge AI",
                "On-device",
                "Optimization"
            ]
        },
        {
            "id": "https://blog.ecitis.org/mixture-of-experts/",
            "content_html": "<p>Dense transformer models activate every parameter for every token. This approach scales poorly; doubling model capacity means doubling compute. Mixture of Experts (MoE) breaks this constraint by routing each token to a subset of specialized sub-networks, enabling models with hundreds of billions of parameters while maintaining reasonable inference costs.</p>\n<p>The shift is complete. On the Artificial Analysis leaderboard, the top 10 open-source models all use MoE architectures<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup>. DeepSeek-R1, Kimi K2 Thinking, Qwen3, Llama 4, and Mistral Large 3 have demonstrated that sparse models match or exceed dense model performance at a fraction of the computational cost. Dense scaling alone no longer works. That is why ERNIE-4.5, Qwen3, Kimi K2, and others all use MoE. Dense models scale intelligence by doing more work. MoE scales intelligence by doing the right work.</p>\n<h2 id=\"the-core-idea\">The Core Idea</h2>\n<p>An MoE layer replaces the standard feed-forward network (FFN) in a transformer block with multiple parallel FFNs called “experts.” A learned routing network determines which experts process each token. Most tokens activate only 1-2 experts out of 8, 128, or even 512 available experts.</p>\n<p>Consider Mixtral 8x7B. The model contains 46.7 billion total parameters, but each token only activates 12.9 billion parameters during inference<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup>. You get the representational capacity of a 47B model with roughly the compute cost of a 13B dense model. Newer models push this further: Kimi K2 activates 32B of 1 trillion total parameters (3.2% activation), and Qwen3-Next activates just 3B of 80B parameters (3.7% activation)<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup><sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup>.</p>\n<figure><figcaption><strong>Router</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/mixture-of-experts-0.webp\" width=\"512\" height=\"386\" alt=\"Router\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>The routing mechanism computes a probability distribution over experts for each token. A gating function G(x) = softmax(W_g * x) produces weights, and the top-K experts by weight process the token. The final output combines expert outputs weighted by their routing scores.</p>\n<h2 id=\"how-routing-works-at-implementation-level\">How Routing Works at Implementation Level</h2>\n<p>The router is a simple linear layer that projects the token hidden state to a vector of expert scores:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>class</span><span> Router</span><span>(</span><span>nn</span><span>.</span><span>Module</span><span>):</span></span>\n<span class=\"line\"><span>    def</span><span> __init__</span><span>(</span><span>self</span><span>,</span><span> hidden_dim</span><span>,</span><span> num_experts</span><span>,</span><span> top_k</span><span>=</span><span>2</span><span>):</span></span>\n<span class=\"line\"><span>        super</span><span>().</span><span>__init__</span><span>()</span></span>\n<span class=\"line\"><span>        self</span><span>.</span><span>gate </span><span>=</span><span> nn</span><span>.</span><span>Linear</span><span>(hidden_dim, num_experts, bias</span><span>=</span><span>False</span><span>)</span></span>\n<span class=\"line\"><span>        self</span><span>.</span><span>top_k </span><span>=</span><span> top_k</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> forward</span><span>(</span><span>self</span><span>,</span><span> x</span><span>):</span></span>\n<span class=\"line\"><span>        # x: [batch, seq_len, hidden_dim]</span></span>\n<span class=\"line\"><span>        logits </span><span>=</span><span> self</span><span>.</span><span>gate</span><span>(x)</span><span>  # [batch, seq_len, num_experts]</span></span>\n<span class=\"line\"><span>        scores </span><span>=</span><span> F</span><span>.</span><span>softmax</span><span>(logits, dim</span><span>=-</span><span>1</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Select top-k experts per token</span></span>\n<span class=\"line\"><span>        top_scores</span><span>,</span><span> top_indices </span><span>=</span><span> torch</span><span>.</span><span>topk</span><span>(scores, self.top_k, dim</span><span>=-</span><span>1</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Renormalize selected weights</span></span>\n<span class=\"line\"><span>        top_scores </span><span>=</span><span> top_scores </span><span>/</span><span> top_scores</span><span>.</span><span>sum</span><span>(dim</span><span>=-</span><span>1</span><span>, keepdim</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> top_scores</span><span>,</span><span> top_indices</span></span></code></pre>\n<p>Switch Transformer simplified earlier MoE designs by using top-1 routing instead of top-2<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup>. This reduces computation and communication overhead while maintaining model quality. The routing decision for each token is independent, allowing different tokens in the same sequence to use different experts.</p>\n<h3 id=\"expert-dispatch-and-gather\">Expert Dispatch and Gather</h3>\n<p>After routing decisions are made, tokens must be physically dispatched to their assigned experts. This creates a challenge: tensor shapes must be known at compile time for efficient execution, but we cannot predict how many tokens each expert will receive.</p>\n<p>The standard solution uses a “capacity factor” that sets maximum tokens per expert:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>capacity </span><span>=</span><span> (tokens_per_batch </span><span>/</span><span> num_experts) </span><span>*</span><span> capacity_factor</span></span></code></pre>\n<p>A capacity factor of 1.0 assumes perfect load balance. In practice, values between 1.25 and 2.0 are common to handle routing imbalance<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-6\" id=\"user-content-fnref-6\">6</a></sup>. Tokens that exceed an expert’s capacity are either dropped (passed through via residual connection) or handled through auxiliary mechanisms.</p>\n<h2 id=\"what-experts-actually-learn\">What Experts Actually Learn</h2>\n<p>A common misconception: experts do not learn human-interpretable domains like “physics” or “medicine.” Research on Switch Transformer and other MoE models shows that experts specialize at the token level, not the semantic level<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-7\" id=\"user-content-fnref-7\">7</a></sup>.</p>\n<p>Encoder experts tend to specialize in syntactic patterns. One expert might handle punctuation tokens, another proper nouns, another sentence-initial tokens. Decoder experts show less clear specialization but maintain consistent activation patterns for specific token types.</p>\n<p>DeepSeek’s architecture separates “shared experts” that activate for every token from “routing experts” that activate conditionally<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-8\" id=\"user-content-fnref-8\">8</a></sup>. Shared experts learn broad patterns required across all inputs; syntax processing and high-level semantic features. Routing experts capture more specialized, context-dependent computations.</p>\n<p>This finding has practical implications. MoE models perform well on knowledge-heavy tasks like TriviaQA but may underperform dense models on reasoning tasks at equivalent perplexity levels<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-7\" id=\"user-content-fnref-7-2\">7</a></sup>.</p>\n<h2 id=\"load-balancing-strategies\">Load Balancing Strategies</h2>\n<p>Without intervention, routers collapse to using only a few experts while ignoring others. This wastes parameters and defeats the purpose of having multiple experts.</p>\n<h3 id=\"auxiliary-loss\">Auxiliary Loss</h3>\n<p>Switch Transformer introduced an auxiliary loss that encourages balanced expert utilization<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-5\" id=\"user-content-fnref-5-2\">5</a></sup>. Given N experts and a batch of T tokens, the loss penalizes deviation from uniform distribution:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> load_balance_loss</span><span>(</span><span>router_probs</span><span>,</span><span> expert_indices</span><span>,</span><span> num_experts</span><span>):</span></span>\n<span class=\"line\"><span>    # f_i: fraction of tokens assigned to expert i</span></span>\n<span class=\"line\"><span>    # P_i: mean router probability for expert i</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Count tokens per expert</span></span>\n<span class=\"line\"><span>    tokens_per_expert </span><span>=</span><span> torch</span><span>.</span><span>zeros</span><span>(num_experts)</span></span>\n<span class=\"line\"><span>    for</span><span> i </span><span>in</span><span> range</span><span>(num_experts):</span></span>\n<span class=\"line\"><span>        tokens_per_expert</span><span>[</span><span>i</span><span>]</span><span> =</span><span> (expert_indices </span><span>==</span><span> i)</span><span>.</span><span>float</span><span>().</span><span>sum</span><span>()</span></span>\n<span class=\"line\"><span>    f </span><span>=</span><span> tokens_per_expert </span><span>/</span><span> tokens_per_expert</span><span>.</span><span>sum</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Mean probability per expert</span></span>\n<span class=\"line\"><span>    P </span><span>=</span><span> router_probs</span><span>.</span><span>mean</span><span>(dim</span><span>=</span><span>0</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Auxiliary loss: scaled dot product</span></span>\n<span class=\"line\"><span>    return</span><span> num_experts </span><span>*</span><span> (f </span><span>*</span><span> P)</span><span>.</span><span>sum</span><span>()</span></span></code></pre>\n<p>The auxiliary loss coefficient requires careful tuning. Too small and experts collapse; too large and the balancing signal overwhelms the primary training objective. Switch Transformer used alpha = 0.01, finding this balanced load quickly without degrading model quality<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-5\" id=\"user-content-fnref-5-3\">5</a></sup>.</p>\n<h3 id=\"auxiliary-loss-free-balancing\">Auxiliary-Loss-Free Balancing</h3>\n<p>DeepSeek-V3 pioneered auxiliary-loss-free load balancing<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-9\" id=\"user-content-fnref-9\">9</a></sup>. Instead of adding a loss term, they modify the routing mechanism itself to maintain balance without gradient interference. This avoids the fundamental trade-off where auxiliary losses can hurt model performance when weighted too heavily.</p>\n<p>The approach, called Loss-Free Balancing, applies expert-wise bias to routing scores before the top-K decision. By dynamically updating each expert’s bias according to its recent load, it maintains balanced distribution without producing interference gradients<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-10\" id=\"user-content-fnref-10\">10</a></sup>. Validation on MoE models with up to 3B parameters trained on 200B tokens showed both better performance and better load balance compared with auxiliary-loss-controlled strategies.</p>\n<p>Interestingly, despite calling it “auxiliary-loss-free,” DeepSeek V3 still uses a small complementary sequence-wise balance loss with a very small hyperparameter<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-11\" id=\"user-content-fnref-11\">11</a></sup>. A theoretical framework published at NeurIPS 2025 analyzed the Loss-Free Balancing procedure, proving monotonic improvement of a Lagrangian objective and an approximate-balancing guarantee<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-12\" id=\"user-content-fnref-12\">12</a></sup>.</p>\n<h3 id=\"global-vs-micro-batch-balance\">Global vs. Micro-Batch Balance</h3>\n<p>Research from ACL 2025 found that employing global-batch load balance significantly outperforms micro-batch level balance by incorporating more diverse domain information<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-13\" id=\"user-content-fnref-13\">13</a></sup>. The finding: adding a small amount of micro-batch load balance while using global-batch balance can maintain model performance while reducing latency from local imbalance. Qwen3 adopted global-batch load balancing loss based on this research<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-14\" id=\"user-content-fnref-14\">14</a></sup>.</p>\n<h3 id=\"expert-choice-routing\">Expert Choice Routing</h3>\n<p>Expert Choice (EC) routing inverts the standard approach: instead of tokens selecting top-k experts, experts select top-k tokens<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-15\" id=\"user-content-fnref-15\">15</a></sup>. Each token can be routed to a variable number of experts while each expert maintains a fixed bucket size. This achieves optimal load balancing while allowing heterogeneity in token-to-expert mapping.</p>\n<p>Recent extensions include TC-MoE (ICLR 2025), which expands the expert space using ternary sets 1, achieving 1.1% improvement while reducing activated experts by up to 9%<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-16\" id=\"user-content-fnref-16\">16</a></sup>. Apple’s EC-DIT (NeurIPS 2025) applied expert-choice routing to diffusion transformers, scaling to 97 billion parameters<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-17\" id=\"user-content-fnref-17\">17</a></sup>.</p>\n<h3 id=\"capacity-factor-trade-offs\">Capacity Factor Trade-offs</h3>\n<p>Setting expert capacity involves a three-way trade-off:</p>\n<ol>\n<li><strong>Too low</strong>: Tokens get dropped, losing information</li>\n<li><strong>Too high</strong>: Wasted computation on empty expert slots</li>\n<li><strong>Dynamic</strong>: More complex implementation, harder to optimize</li>\n</ol>\n<p>Mixtral demonstrated that with diverse training data and proper initialization, load balance can emerge naturally without explicit auxiliary losses<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-18\" id=\"user-content-fnref-18\">18</a></sup>. The top-k routing over varied data spreads selections across experts. Post-training analysis confirmed all experts receive substantial token allocations.</p>\n<h2 id=\"training-dynamics-and-challenges\">Training Dynamics and Challenges</h2>\n<h3 id=\"instability\">Instability</h3>\n<p>MoE training exhibits higher variance than dense models. The routing decisions create feedback loops: an expert that performs well receives more tokens, gets more gradient updates, and becomes even more likely to be selected. This can lead to expert collapse or oscillating training dynamics.</p>\n<p>Switch Transformer showed that large sparse models can be trained with bfloat16 precision, but this requires careful attention to numerical stability in the routing computation<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-5\" id=\"user-content-fnref-5-4\">5</a></sup>. Many implementations use float32 for the router even when the rest of the model uses mixed precision.</p>\n<p>Kimi K2 trained for 15.5 trillion tokens with zero training instability by applying the Muon optimizer at unprecedented scale and developing novel optimization techniques<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-3\" id=\"user-content-fnref-3-2\">3</a></sup>.</p>\n<h3 id=\"communication-overhead\">Communication Overhead</h3>\n<p>In distributed training, MoE layers require all-to-all communication. Each token must be sent to the GPU holding its assigned expert, processed, and returned. This communication pattern differs from standard tensor or data parallelism.</p>\n<p>Expert parallelism (EP) distributes experts across GPUs. A model with 256 experts across 8 GPUs places 32 experts per GPU. The all-to-all communication volume scales with batch size and the degree of expert parallelism.</p>\n<p>Hybrid parallelism combines EP with tensor parallelism (TP) and data parallelism (DP). The optimal configuration depends on expert sizes, model architecture, and hardware interconnect bandwidth<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-19\" id=\"user-content-fnref-19\">19</a></sup>.</p>\n<h3 id=\"multi-token-prediction\">Multi-Token Prediction</h3>\n<p>DeepSeek-V3 pioneered combining MoE with multi-token prediction (MTP) training objectives<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-9\" id=\"user-content-fnref-9-2\">9</a></sup>. Instead of predicting only the next token, MTP predicts the next k tokens at each position using k independent output heads on a shared trunk. The approach allows denser training signals and improves data efficiency.</p>\n<p>Unlike other MTP methods, DeepSeek’s approach maintains the causal chain by predicting additional tokens sequentially rather than in parallel. MTP modules are dropped at inference (though they can accelerate generation via speculative decoding), with acceptance rates between 85-90% for the second token prediction<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-11\" id=\"user-content-fnref-11-2\">11</a></sup>. Models trained with 4-token prediction are up to 3x faster at inference<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-20\" id=\"user-content-fnref-20\">20</a></sup>.</p>\n<h3 id=\"fine-grained-vs-coarse-grained-experts\">Fine-Grained vs Coarse-Grained Experts</h3>\n<p>Traditional MoE uses a moderate number of large experts. DeepSeek introduced fine-grained experts: more experts with smaller hidden dimensions, activating more experts per token<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-8\" id=\"user-content-fnref-8-2\">8</a></sup>.</p>\n<p>If you have N experts activating K per token, fine-grained design uses mN experts with hidden dimension reduced by 1/m, activating mK experts per token. Total computation stays constant, but the model has access to more diverse expert combinations.</p>\n<p>The trend toward higher sparsity continues. DeepSeek-V3 uses 256 experts with 8 active per token. Kimi K2 increased sparsity further: 384 experts with 8 active, resulting in sparsity of 48 (384/8)<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-3\" id=\"user-content-fnref-3-3\">3</a></sup>. Qwen3-Next pushes to the extreme: 512 experts activating about 19 per token, achieving only 3.7% parameter activation with a 1</p><div></div> activation ratio<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-4\" id=\"user-content-fnref-4-2\">4</a></sup>.<p></p>\n<h2 id=\"hybrid-attention-architectures\">Hybrid Attention Architectures</h2>\n<p>A major 2025 development combines MoE with linear attention mechanisms. Standard attention scales quadratically with context length. Linear attention variants like Gated DeltaNet scale linearly but compress past context through a memory bottleneck<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-21\" id=\"user-content-fnref-21\">21</a></sup>.</p>\n<p>Qwen3-Next and Kimi Linear proposed hybrid architectures using a 3</p><div></div> ratio: for every three transformer blocks employing linear Gated DeltaNet, one block uses full attention<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-22\" id=\"user-content-fnref-22\">22</a></sup><sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-23\" id=\"user-content-fnref-23\">23</a></sup>. This structure enables efficient long-context modeling while preserving reasoning capability on complex tasks.<p></p>\n<p>Gated DeltaNet combines the gating mechanism from Mamba2 with the delta update rule from DeltaNet. Gating enables rapid memory erasure; the delta rule facilitates targeted updates. The combination consistently surpasses Mamba2 and DeltaNet across language modeling, common-sense reasoning, in-context retrieval, and length extrapolation benchmarks<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-21\" id=\"user-content-fnref-21-2\">21</a></sup>.</p>\n<p>For the MoE layers in these hybrid architectures, the structure repeats as: Layers 1-3 use Linear attention followed by MoE, Layer 4 uses Full attention followed by MoE. Qwen3-Next’s flagship 80B-A3B model achieves only 3B active parameters per token through this design<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-4\" id=\"user-content-fnref-4-3\">4</a></sup>.</p>\n<h2 id=\"inference-considerations\">Inference Considerations</h2>\n<h3 id=\"memory-requirements\">Memory Requirements</h3>\n<p>MoE models require loading all expert weights into memory, even though each token only uses a fraction. A 671B parameter MoE model needs 671B parameters worth of storage, not 37B.</p>\n<p>This creates an asymmetry between training and inference optimization. Training benefits from compute efficiency (fewer FLOPs per token). Inference on smaller batches may be bottlenecked by memory bandwidth for loading expert weights rather than compute. DeepSeek-V3’s total computational cost is approximately 250 GFLOPS per token, whereas a 72B dense model requires 394 GFLOPS and a 405B dense model requires 2448 GFLOPS<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-11\" id=\"user-content-fnref-11-3\">11</a></sup>.</p>\n<h3 id=\"batch-routing-efficiency\">Batch Routing Efficiency</h3>\n<p>Large batch sizes help MoE inference. With more tokens, each expert receives more work, improving GPU utilization. DeepSeek-V3’s high sparsity (8 of 256 experts) requires large batch sizes to ensure sufficient tokens per expert<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-24\" id=\"user-content-fnref-24\">24</a></sup>.</p>\n<p>The routing pattern varies per batch, creating dynamic load imbalance. Some experts may be overloaded while others sit idle. Inference latency is determined by the most loaded expert, so load imbalance directly impacts throughput.</p>\n<h3 id=\"expert-parallelism-in-inference\">Expert Parallelism in Inference</h3>\n<p>vLLM now supports expert parallelism for large-scale MoE deployment, combining Data Parallel attention with Expert or Tensor Parallel MoE layers<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-25\" id=\"user-content-fnref-25\">25</a></sup>. Key optimizations include:</p>\n<ol>\n<li><strong>Expert Parallel Load Balancing (EPLB)</strong>: vLLM collects load statistics with every forward pass and periodically rebalances expert distribution across EP ranks<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-25\" id=\"user-content-fnref-25-2\">25</a></sup>.</li>\n<li><strong>Communication backends</strong>: DeepSeek’s DeepEP kernels use nvshmem with high-throughput mode for prefill and low-latency mode for decode. Perplexity’s PPLX provides a more flexible alternative for chunked prefill scenarios<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-26\" id=\"user-content-fnref-26\">26</a></sup>.</li>\n<li><strong>Wide Expert Parallelism</strong>: Simplifies scaling across multi-node deployments using NIXL+UCX for inter-node communication<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-26\" id=\"user-content-fnref-26-2\">26</a></sup>.</li>\n</ol>\n<p>MoEShard achieves load balance through tensor sharding of experts rather than capacity-based dropping<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-27\" id=\"user-content-fnref-27\">27</a></sup>. Each expert’s weights are split across GPUs, and all GPUs participate in computing each expert’s output. This guarantees full token retention regardless of routing skew.</p>\n<figure><figcaption><strong>MoE Model Comparison</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/mixture-of-experts-1.webp\" width=\"512\" height=\"487\" alt=\"MoE Model Comparison\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"hardware-optimization-blackwell-nvl72\">Hardware Optimization: Blackwell NVL72</h3>\n<p>The NVIDIA GB200 NVL72 rack-scale platform connects 72 Blackwell GPUs using fifth-generation NVLink, providing 1,800 GB/s bidirectional bandwidth<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-28\" id=\"user-content-fnref-28\">28</a></sup>. This large scale-up domain is optimized for sparse MoE architectures, which require frequent all-to-all exchanges between experts.</p>\n<p>MoE models see a 10x performance leap on GB200 NVL72 compared with HGX H200, enabling one-tenth the cost per token<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-1\" id=\"user-content-fnref-1-2\">1</a></sup>. Key enablers include hardware acceleration for NVFP4 four-bit floating point, disaggregated serving through NVIDIA Dynamo, and multi-token prediction support in TensorRT-LLM<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-28\" id=\"user-content-fnref-28-2\">28</a></sup>.</p>\n<p>For DeepSeek-R1, NVIDIA achieved over 250 tokens per second per user and maximum throughput over 30,000 tokens per second on a single DGX system with eight Blackwell GPUs<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-29\" id=\"user-content-fnref-29\">29</a></sup>. The GB300 NVL72 (Blackwell Ultra) delivered 45% higher performance per GPU on DeepSeek-R1 compared to GB200 NVL72<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-30\" id=\"user-content-fnref-30\">30</a></sup>.</p>\n<h2 id=\"survey-of-current-moe-models-2025-2026\">Survey of Current MoE Models (2025-2026)</h2>\n<h3 id=\"switch-transformer-2021\">Switch Transformer (2021)</h3>\n<p>Google’s Switch Transformer demonstrated MoE at scale with 1.6 trillion parameters across 2048 experts<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-5\" id=\"user-content-fnref-5-5\">5</a></sup>. Using top-1 routing, it achieved 7x pre-training speedup over T5 models while reducing the complexity of earlier MoE approaches. The work established auxiliary load balancing losses as standard practice.</p>\n<h3 id=\"mixtral-2024\">Mixtral (2024)</h3>\n<p>Mistral’s Mixtral 8x7B popularized MoE for open-weight models<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup>. With 8 experts and top-2 routing, it uses 12.9B active parameters from 46.7B total. Mixtral achieved natural load balance without auxiliary losses, suggesting careful initialization and diverse training data can replace explicit balancing mechanisms.</p>\n<h3 id=\"deepseek-v3-2024\">DeepSeek-V3 (2024)</h3>\n<p>DeepSeek-V3 combines fine-grained experts (256 total, 8 active) with Multi-head Latent Attention (MLA) and auxiliary-loss-free load balancing<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-9\" id=\"user-content-fnref-9-3\">9</a></sup>. The model has 671B total parameters with 37B active, trained on 14.8 trillion tokens. Training required 2.788 million H800 GPU hours. DeepSeek reported performance competitive with closed-source frontier models.</p>\n<h3 id=\"deepseek-r1-january-2025\">DeepSeek-R1 (January 2025)</h3>\n<p>DeepSeek-R1 builds on the V3 architecture, using the same 671B/37B MoE structure but trained via large-scale reinforcement learning without supervised fine-tuning as a preliminary step<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-31\" id=\"user-content-fnref-31\">31</a></sup>. Analysis shows 67 experts per layer in R1 compared to 55 per layer in V3. The model achieves 79.8% pass@1 on AIME, 97.3% on MATH-500, and a 2,029 Elo rating on Codeforces-like challenges<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-31\" id=\"user-content-fnref-31-2\">31</a></sup>.</p>\n<h3 id=\"llama-4-april-2025\">Llama 4 (April 2025)</h3>\n<p>Meta’s Llama 4 represents Meta’s first MoE architecture and first natively multimodal models<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-32\" id=\"user-content-fnref-32\">32</a></sup>. The family includes:</p>\n<ul>\n<li><strong>Llama 4 Scout</strong>: 17B active parameters, 16 experts, fits on a single H100 with INT4 quantization, supports 10 million token context</li>\n<li><strong>Llama 4 Maverick</strong>: 17B active / 400B total, 128 routed experts plus 1 shared expert, alternating dense and MoE layers, 1 million token context</li>\n<li><strong>Llama 4 Behemoth</strong>: 288B active parameters, 16 experts, nearly 2 trillion total parameters, outperforms GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on MATH-500 and GPQA Diamond<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-32\" id=\"user-content-fnref-32-2\">32</a></sup></li>\n</ul>\n<p>Maverick’s architecture sends each token to the shared expert plus one of 128 routed experts, using alternating dense and MoE layers for inference efficiency.</p>\n<h3 id=\"qwen3-may-2025\">Qwen3 (May 2025)</h3>\n<p>The Qwen3 series includes models from 0.6B to 235B parameters<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-14\" id=\"user-content-fnref-14-2\">14</a></sup>. Qwen3-235B-A22B uses 128 experts with 8 active, activating 22B of 235B parameters. Unlike earlier Qwen MoE models, Qwen3 removes shared experts and uses global-batch load balancing loss. Training covered 36 trillion tokens across 119 languages.</p>\n<p>A key innovation is the integration of thinking mode (complex multi-step reasoning) and non-thinking mode (rapid responses) into a unified framework, enabling dynamic mode switching. Qwen3-235B-A22B-Thinking-2507 beats OpenAI O3 across key metrics: 92 vs 88.0 on AIME’25, 83 vs 82.5 on HMMT’25<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-33\" id=\"user-content-fnref-33\">33</a></sup>.</p>\n<h3 id=\"qwen3-next-september-2025\">Qwen3-Next (September 2025)</h3>\n<p>Alibaba’s Qwen3-Next-80B-A3B represents ultra-sparse MoE design<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-4\" id=\"user-content-fnref-4-4\">4</a></sup>. With 512 experts and only 3B active parameters (3.7% activation, 1</p><div></div> ratio), it combines hybrid attention (Gated DeltaNet + Gated Attention) with multi-token prediction. Every 4th layer uses standard GQA attention; others use linear attention variants. The model supports 256K context natively, extendable to 1 million tokens.<p></p>\n<h3 id=\"kimi-k2-july-2025\">Kimi K2 (July 2025)</h3>\n<p>Moonshot AI’s Kimi K2 uses 1 trillion total parameters with 32B active, featuring 384 experts with 8 selected per token plus 1 shared expert<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-3\" id=\"user-content-fnref-3-4\">3</a></sup>. The architecture follows DeepSeek V3’s MLA design with 7168 hidden dimension and 2048 expert hidden dimension. Trained with the Muon optimizer on 15.5T tokens with zero instability, it executes 200-300 sequential tool calls autonomously. Kimi K2 Thinking adds step-by-step reasoning with dynamic tool invocation.</p>\n<h3 id=\"mistral-large-3-december-2025\">Mistral Large 3 (December 2025)</h3>\n<p>Mistral Large 3 is a sparse MoE with 675B total and 41B active parameters, a 16</p><div></div> ratio<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-34\" id=\"user-content-fnref-34\">34</a></sup>. The architecture uses thousands of specialized expert subnetworks with granular sparse routing. Training used 3000 H200 GPUs. The model delivers 92% of GPT-5.2’s performance at roughly 15% of the price<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-1\" id=\"user-content-fnref-1-3\">1</a></sup>. Released under Apache 2.0 license.<p></p>\n<h3 id=\"ernie-45-july-2025\">ERNIE 4.5 (July 2025)</h3>\n<p>Baidu’s ERNIE 4.5 series includes 10 variants from 0.3B to 424B parameters<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-35\" id=\"user-content-fnref-35\">35</a></sup>. The architecture introduces heterogeneous modality MoE with modality-isolated routing for text, image, and video. Key innovations include router orthogonal loss and multimodal token-balanced loss. ERNIE-4.5-300B-A47B-Base surpasses DeepSeek-V3-671B-A37B-Base on 22 of 28 benchmarks<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-35\" id=\"user-content-fnref-35-2\">35</a></sup>. The MoE variants activate 2 of 64 experts per token.</p>\n<h3 id=\"grok-3-february-2025\">Grok 3 (February 2025)</h3>\n<p>xAI’s Grok 3 uses an MoE transformer with estimated 1.5 trillion parameters<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-36\" id=\"user-content-fnref-36\">36</a></sup>. Trained with 10x more compute than Grok-2 on the Colossus cluster (approximately 200k GPUs), it employs sparse attention mechanisms and MoE layers that dynamically allocate computational resources. The architecture continues the Grok-1 approach of 64 transformer layers with MoE feed-forward layers using a router picking a subset of expert MLPs per token. Grok 3’s “Think” and “Big Brain” modes expose chain-of-thought reasoning with additional compute allocation<sup><a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fn-36\" id=\"user-content-fnref-36-2\">36</a></sup>.</p>\n<h2 id=\"when-to-use-moe-vs-dense-models\">When to Use MoE vs Dense Models</h2>\n<p>MoE makes sense when:</p>\n<ul>\n<li><strong>Training compute is limited</strong>: Sparse models achieve better quality per FLOP</li>\n<li><strong>Inference batch sizes are large</strong>: Amortizes expert loading overhead</li>\n<li><strong>Knowledge breadth matters</strong>: MoE excels at knowledge-intensive tasks</li>\n<li><strong>You have fast interconnects</strong>: All-to-all communication needs low latency</li>\n</ul>\n<p>Dense models may be preferable when:</p>\n<ul>\n<li><strong>Inference is memory-bound</strong>: MoE requires full model in memory</li>\n<li><strong>Reasoning depth matters more than knowledge breadth</strong>: Dense models may perform better on complex reasoning</li>\n<li><strong>Single-request latency is critical</strong>: Small batches underutilize MoE capacity</li>\n<li><strong>Deployment hardware is constrained</strong>: Expert parallelism needs multiple GPUs</li>\n</ul>\n<p>The trend is clear: MoE architectures dominate the frontier. As interconnect speeds improve and inference systems mature, the compute efficiency advantages of sparse models become harder to ignore.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p><a href=\"https://blogs.nvidia.com/blog/mixture-of-experts-frontier-models/\" rel=\"noopener noreferrer\">NVIDIA Blog - Mixture of Experts Powers the Most Intelligent Frontier AI Models</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-1-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-1-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p><a href=\"https://arxiv.org/abs/2401.04088\" rel=\"noopener noreferrer\">Mixtral of Experts - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2401.04088\" rel=\"noopener noreferrer\">.04088</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a><p></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p><a href=\"https://moonshotai.github.io/Kimi-K2/\" rel=\"noopener noreferrer\">Kimi K2: Open Agentic Intelligence - Moonshot AI</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-3-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-3-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-3-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p><a href=\"https://www.alibabacloud.com/blog/602536\" rel=\"noopener noreferrer\">Qwen3-Next: A New Generation of Ultra-Efficient Model Architecture - Alibaba Cloud</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-4-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-4-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-4-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p><a href=\"https://arxiv.org/abs/2101.03961\" rel=\"noopener noreferrer\">Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2101.03961\" rel=\"noopener noreferrer\">.03961</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-5-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-5-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-5-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-5-5\" class=\"data-footnote-backref\">↩<sup>5</sup></a><p></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p><a href=\"https://huggingface.co/blog/moe\" rel=\"noopener noreferrer\">Mixture of Experts Explained - Hugging Face Blog</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p><a href=\"https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-mixture-of-experts\" rel=\"noopener noreferrer\">A Visual Guide to Mixture of Experts - Maarten Grootendorst</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-7-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p><a href=\"https://www.chrishayduk.com/p/understanding-deepseek-part-i-deepseekmoe\" rel=\"noopener noreferrer\">Understanding DeepSeek Part I: DeepSeekMoE - Chris Hayduk</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-8-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-9\">\n<p><a href=\"https://arxiv.org/abs/2412.19437\" rel=\"noopener noreferrer\">DeepSeek-V3 Technical Report - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2412.19437\" rel=\"noopener noreferrer\">.19437</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-9\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-9-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-9-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a><p></p>\n</li>\n<li id=\"user-content-fn-10\">\n<p><a href=\"https://arxiv.org/abs/2408.15664\" rel=\"noopener noreferrer\">Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2408.15664\" rel=\"noopener noreferrer\">.15664</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-10\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p><a href=\"https://magazine.sebastianraschka.com/p/technical-deepseek\" rel=\"noopener noreferrer\">A Technical Tour of the DeepSeek Models from V3 to V3.2 - Sebastian Raschka</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-11-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-11-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-12\">\n<p><a href=\"https://openreview.net/forum?id=sQHZ2w8UzP\" rel=\"noopener noreferrer\">A Theoretical Framework for Auxiliary-Loss-Free Load-Balancing - NeurIPS 2025</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-12\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-13\">\n<p><a href=\"https://aclanthology.org/2025.acl-long.249.pdf\" rel=\"noopener noreferrer\">On Implementing Load Balancing Loss for Training MoE - ACL 2025</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-13\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p><a href=\"https://arxiv.org/abs/2505.09388\" rel=\"noopener noreferrer\">Qwen3 Technical Report - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2505.09388\" rel=\"noopener noreferrer\">.09388</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-14-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a><p></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p><a href=\"https://arxiv.org/abs/2202.09368\" rel=\"noopener noreferrer\">Mixture-of-Experts with Expert Choice Routing - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2202.09368\" rel=\"noopener noreferrer\">.09368</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-16\">\n<p><a href=\"https://openreview.net/forum?id=dsP91M4hDL\" rel=\"noopener noreferrer\">TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice - ICLR 2025</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-16\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-17\">\n<p><a href=\"https://machinelearning.apple.com/research/ec-dit\" rel=\"noopener noreferrer\">EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing - Apple ML Research</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-17\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-18\">\n<p><a href=\"https://medium.com/@pilliudayaditya1207/understanding-mixture-of-experts-switch-transformers-load-balancing-vs-mixtral-s-natural-balance-25ed528cadfe\" rel=\"noopener noreferrer\">Understanding Mixture-of-Experts: Switch Transformer’s Load Balancing vs. Mixtral’s Natural Balance - Medium</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-18\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-19\">\n<p><a href=\"https://pssg.cs.umd.edu/assets/papers/2023-06-deepspeed-ted-moe.pdf\" rel=\"noopener noreferrer\">A Hybrid Tensor-Expert-Data Parallelism Approach - DeepSpeed-TED</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-19\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-20\">\n<p><a href=\"https://arxiv.org/abs/2404.19737\" rel=\"noopener noreferrer\">Better &amp; Faster Large Language Models via Multi-token Prediction - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2404.19737\" rel=\"noopener noreferrer\">.19737</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-20\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-21\">\n<p><a href=\"https://arxiv.org/abs/2412.06464\" rel=\"noopener noreferrer\">Gated Delta Networks: Improving Mamba2 with Delta Rule - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2412.06464\" rel=\"noopener noreferrer\">.06464</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-21\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-21-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a><p></p>\n</li>\n<li id=\"user-content-fn-22\">\n<p><a href=\"https://blog.vllm.ai/2025/09/11/qwen3-next.html\" rel=\"noopener noreferrer\">vLLM Now Supports Qwen3-Next: Hybrid Architecture with Extreme Efficiency</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-22\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-23\">\n<p><a href=\"https://arxiv.org/abs/2510.26692\" rel=\"noopener noreferrer\">Kimi Linear: An Expressive, Efficient Attention Architecture - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2510.26692\" rel=\"noopener noreferrer\">.26692</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-23\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-24\">\n<p><a href=\"https://www.tensoreconomics.com/p/moe-inference-economics-from-first\" rel=\"noopener noreferrer\">MoE Inference Economics from First Principles - Tensor Economics</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-24\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-25\">\n<p><a href=\"https://docs.vllm.ai/en/latest/serving/expert_parallel_deployment/\" rel=\"noopener noreferrer\">Expert Parallel Deployment - vLLM Documentation</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-25\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-25-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-26\">\n<p><a href=\"https://developers.redhat.com/articles/2025/09/08/scaling-deepseek-style-moes-vllm-and-llm-d-using-wide-ep\" rel=\"noopener noreferrer\">Scaling DeepSeek-style MoEs with vLLM and llm-d using Wide EP - Red Hat Developer</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-26\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-26-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-27\">\n<p><a href=\"https://arxiv.org/abs/2503.08467\" rel=\"noopener noreferrer\">Accelerating MoE Model Inference with Expert Sharding - arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2503.08467\" rel=\"noopener noreferrer\">.08467</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-27\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-28\">\n<p><a href=\"https://developer.nvidia.com/blog/delivering-massive-performance-leaps-for-mixture-of-experts-inference-on-nvidia-blackwell/\" rel=\"noopener noreferrer\">Delivering Massive Performance Leaps for MoE Inference on NVIDIA Blackwell - NVIDIA Technical Blog</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-28\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-28-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-29\">\n<p><a href=\"https://developer.nvidia.com/blog/nvidia-blackwell-delivers-world-record-deepseek-r1-inference-performance/\" rel=\"noopener noreferrer\">NVIDIA Blackwell Delivers World-Record DeepSeek-R1 Inference Performance - NVIDIA Technical Blog</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-29\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-30\">\n<p><a href=\"https://developer.nvidia.com/blog/nvidia-blackwell-ultra-sets-new-inference-records-in-mlperf-debut/\" rel=\"noopener noreferrer\">NVIDIA Blackwell Ultra Sets New Inference Records in MLPerf Debut - NVIDIA Technical Blog</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-30\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-31\">\n<p><a href=\"https://github.com/deepseek-ai/DeepSeek-R1\" rel=\"noopener noreferrer\">DeepSeek-R1 - GitHub</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-31\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-31-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-32\">\n<p><a href=\"https://ai.meta.com/blog/llama-4-multimodal-intelligence/\" rel=\"noopener noreferrer\">The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation - Meta AI</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-32\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-32-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-33\">\n<p><a href=\"https://www.siliconflow.com/articles/en/the-best-qwen3-models-in-2025\" rel=\"noopener noreferrer\">Ultimate Guide - The Best Qwen3 Models in 2026 - SiliconFlow</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-33\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-34\">\n<p><a href=\"https://mistral.ai/news/mistral-3\" rel=\"noopener noreferrer\">Introducing Mistral 3 - Mistral AI</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-34\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-35\">\n<p><a href=\"https://ernie.baidu.com/blog/posts/ernie4.5/\" rel=\"noopener noreferrer\">Announcing the Open Source Release of the ERNIE 4.5 Model Family - Baidu</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-35\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-35-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-36\">\n<p><a href=\"https://mohasoftware.com/blog/the-tech-behind-elon-musks-xai-latest-release-grok-3\" rel=\"noopener noreferrer\">The Tech Behind Grok 3 - MOHA Software</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-36\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/mixture-of-experts/#user-content-fnref-36-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/mixture-of-experts/",
            "title": "Mixture of Experts: Sparse Computation for Efficient LLMs",
            "summary": "See how sparse routing expands model capacity without paying the compute cost of activating every parameter for every token.",
            "image": "https://blog.ecitis.org/open-graph/mixture-of-experts.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "MoE",
                "Transformers",
                "Sparse models"
            ]
        },
        {
            "id": "https://blog.ecitis.org/rlhf-preference-tuning/",
            "content_html": "<p>Large language models trained on internet text learn to predict the next token. This objective produces models that can generate fluent text, but fluency alone does not make a model useful or safe. A model optimizing purely for next-token prediction might produce toxic content, confidently state falsehoods, or refuse to help with legitimate requests. The gap between “predicting text well” and “being helpful” is the alignment problem.</p>\n<p>Reinforcement Learning from Human Feedback emerged as the dominant solution. OpenAI’s InstructGPT demonstrated that RLHF could transform GPT-3 into a model that humans preferred 85% of the time over the base model.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup> This technique became foundational to ChatGPT, Claude, and Gemini.</p>\n<p>But RLHF is complex. It requires training multiple models, managing reinforcement learning instabilities, and collecting expensive human preference data. In 2023, researchers at Stanford introduced Direct Preference Optimization, which achieves similar results with a simpler supervised learning objective.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup> Since then, the field has expanded rapidly. DPO spawned variants including IPO, KTO, ORPO, and SimPO. DeepSeek introduced GRPO for training reasoning models. New methods like AlphaPO and Reinforcement Learning with Verifiable Rewards (RLVR) emerged in 2025, while research on failure modes like shallow safety alignment and sycophancy has matured.</p>\n<p>This guide covers the technical details of these approaches: how they work, when to use each one, and how to implement them in practice.</p>\n<h2 id=\"the-alignment-problem\">The Alignment Problem</h2>\n<p>Language models learn from massive text corpora containing everything from academic papers to social media posts. The training objective is simple: given a sequence of tokens, predict the next one. This produces models with broad capabilities but no inherent sense of what outputs are desirable.</p>\n<p>Consider what happens when you ask a base model to help with a task. It might:</p>\n<ul>\n<li>Generate a helpful response</li>\n<li>Generate a harmful response</li>\n<li>Refuse unnecessarily</li>\n<li>Hallucinate confidently</li>\n<li>Continue your prompt as if completing a document</li>\n</ul>\n<p>All these behaviors are consistent with next-token prediction on internet text. The model has no internal preference for helpfulness over harm.</p>\n<p>RLHF addresses this by introducing a training signal based on human preferences. Rather than predicting what text typically follows, the model learns to generate text that humans rate as good. This requires three components:</p>\n<ol>\n<li>A supervised fine-tuned model that can follow instructions</li>\n<li>A reward model that predicts human preferences</li>\n<li>A reinforcement learning algorithm that optimizes the model against the reward</li>\n</ol>\n<figure><figcaption><strong>RLHF Training Pipeline</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/rlhf-preference-tuning-0.webp\" width=\"512\" height=\"478\" alt=\"RLHF Training Pipeline\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"stage-1-supervised-fine-tuning\">Stage 1: Supervised Fine-Tuning</h2>\n<p>Before RLHF, the base model needs basic instruction-following capabilities. This stage collects demonstration data: human-written examples of good responses to prompts. OpenAI used approximately 13,000 (prompt, response) pairs for InstructGPT, with labelers who were carefully selected and trained.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup></p>\n<p>The training objective is standard cross-entropy loss:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>S</mi><mi>F</mi><mi>T</mi></mrow></msub><mo>=</mo><mo>−</mo><msub><mi mathvariant=\"double-struck\">E</mi><mrow><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>y</mi><mo stretchy=\"false\">)</mo><mo>∼</mo><mi>D</mi></mrow></msub><mrow><mo fence=\"true\">[</mo><msubsup><mo>∑</mo><mrow><mi>t</mi><mo>=</mo><mn>1</mn></mrow><mi>T</mi></msubsup><mi>log</mi><mo>⁡</mo><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>t</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo separator=\"true\">,</mo><msub><mi>y</mi><mrow><mo>&lt;</mo><mi>t</mi></mrow></msub><mo stretchy=\"false\">)</mo><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>This produces a policy <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>π</mi><mrow><mi>S</mi><mi>F</mi><mi>T</mi></mrow></msub></mrow></semantics></math> that can generate reasonable responses but has not yet learned to distinguish between good and bad outputs. The SFT model serves as both the starting point for RLHF training and often as the reference policy for KL regularization.</p>\n<p>Quality matters more than quantity at this stage. A few thousand high-quality demonstrations outperform larger datasets of mediocre examples. The goal is to establish the format and style of responses, not to cover every possible topic.</p>\n<h2 id=\"stage-2-reward-model-training\">Stage 2: Reward Model Training</h2>\n<p>The reward model learns to predict which responses humans prefer. Given a prompt and two candidate responses, the RM outputs which one is better. This pairwise comparison approach is more reliable than asking humans to assign absolute scores.</p>\n<h3 id=\"the-bradley-terry-model\">The Bradley-Terry Model</h3>\n<p>Most reward models use the Bradley-Terry framework from statistics.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup> It assumes each response has a latent quality score, and the probability that response A beats response B follows a logistic function:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>P</mi><mo stretchy=\"false\">(</mo><mi>A</mi><mo>≻</mo><mi>B</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mi>σ</mi><mo stretchy=\"false\">(</mo><mi>r</mi><mo stretchy=\"false\">(</mo><mi>A</mi><mo stretchy=\"false\">)</mo><mo>−</mo><mi>r</mi><mo stretchy=\"false\">(</mo><mi>B</mi><mo stretchy=\"false\">)</mo><mo stretchy=\"false\">)</mo><mo>=</mo><mfrac><mn>1</mn><mrow><mn>1</mn><mo>+</mo><msup><mi>e</mi><mrow><mo>−</mo><mo stretchy=\"false\">(</mo><mi>r</mi><mo stretchy=\"false\">(</mo><mi>A</mi><mo stretchy=\"false\">)</mo><mo>−</mo><mi>r</mi><mo stretchy=\"false\">(</mo><mi>B</mi><mo stretchy=\"false\">)</mo><mo stretchy=\"false\">)</mo></mrow></msup></mrow></mfrac></mrow></semantics></math></p>\n<p>The reward model is trained to maximize the log-likelihood of observed preferences:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>R</mi><mi>M</mi></mrow></msub><mo>=</mo><mo>−</mo><msub><mi mathvariant=\"double-struck\">E</mi><mrow><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><msub><mi>y</mi><mi>w</mi></msub><mo separator=\"true\">,</mo><msub><mi>y</mi><mi>l</mi></msub><mo stretchy=\"false\">)</mo><mo>∼</mo><mi>D</mi></mrow></msub><mrow><mo fence=\"true\">[</mo><mi>log</mi><mo>⁡</mo><mi>σ</mi><mo stretchy=\"false\">(</mo><msub><mi>r</mi><mi>ϕ</mi></msub><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><msub><mi>y</mi><mi>w</mi></msub><mo stretchy=\"false\">)</mo><mo>−</mo><msub><mi>r</mi><mi>ϕ</mi></msub><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><msub><mi>y</mi><mi>l</mi></msub><mo stretchy=\"false\">)</mo><mo stretchy=\"false\">)</mo><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>Here <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>y</mi><mi>w</mi></msub></mrow></semantics></math> is the preferred (winning) response and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>y</mi><mi>l</mi></msub></mrow></semantics></math> is the dispreferred (losing) response.</p>\n<h3 id=\"architecture\">Architecture</h3>\n<p>Reward models are typically language models with a linear head that outputs a scalar value instead of token logits. Common practice initializes the RM from the SFT checkpoint, which gives it strong language understanding capabilities.</p>\n<p>The architecture processes the prompt and response together, outputting a single reward value. Some implementations average the final hidden states; others use only the last token’s representation.</p>\n<h3 id=\"calibration-challenges\">Calibration Challenges</h3>\n<p>Reward models face several practical issues:</p>\n<p><strong>Length bias</strong>: Longer responses often receive higher scores regardless of quality. This can be mitigated by normalizing rewards by response length or by carefully balancing the training data.</p>\n<p><strong>Distribution shift</strong>: The RM is trained on outputs from the SFT model but will be used to evaluate outputs from the RL-trained policy. As the policy improves, its outputs may fall outside the RM’s training distribution.</p>\n<p><strong>Annotation noise</strong>: Human preferences are inconsistent. Inter-annotator agreement on preference datasets hovers around 65-75%, meaning a substantial fraction of “ground truth” labels are effectively random.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup></p>\n<h2 id=\"stage-3-ppo-for-language-models\">Stage 3: PPO for Language Models</h2>\n<p>With a reward model in hand, we can optimize the language model using reinforcement learning. Proximal Policy Optimization has become the standard algorithm, adapted from robotics and game-playing applications.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-6\" id=\"user-content-fnref-6\">6</a></sup></p>\n<h3 id=\"the-objective\">The Objective</h3>\n<p>The RLHF objective maximizes expected reward while staying close to the reference policy:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow><mi>max</mi><mo>⁡</mo></mrow><mi>θ</mi></msub><msub><mi mathvariant=\"double-struck\">E</mi><mrow><mi>x</mi><mo>∼</mo><mi>D</mi><mo separator=\"true\">,</mo><mi>y</mi><mo>∼</mo><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mo>⋅</mo><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></msub><mrow><mo fence=\"true\">[</mo><msub><mi>r</mi><mi>ϕ</mi></msub><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>y</mi><mo stretchy=\"false\">)</mo><mo>−</mo><mi>β</mi><mo>⋅</mo><msub><mi>D</mi><mrow><mi>K</mi><mi>L</mi></mrow></msub><mo stretchy=\"false\">(</mo><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mo>⋅</mo><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo><mi mathvariant=\"normal\">∥</mi><msub><mi>π</mi><mrow><mi>r</mi><mi>e</mi><mi>f</mi></mrow></msub><mo stretchy=\"false\">(</mo><mo>⋅</mo><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo><mo stretchy=\"false\">)</mo><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>The KL penalty term prevents the policy from diverging too far from the reference model. Without it, the policy would quickly learn to exploit quirks in the reward model rather than genuinely improving. This phenomenon is called reward hacking.</p>\n<h3 id=\"the-training-loop\">The Training Loop</h3>\n<p>Each iteration of PPO training:</p>\n<ol>\n<li><strong>Sample</strong>: Generate responses from the current policy for a batch of prompts</li>\n<li><strong>Evaluate</strong>: Score each response using the reward model</li>\n<li><strong>Compute advantages</strong>: Calculate how much better each response was than expected</li>\n<li><strong>Update</strong>: Apply gradient ascent with the clipped surrogate objective</li>\n</ol>\n<p>The clipped objective prevents large policy updates:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msup><mi>L</mi><mrow><mi>C</mi><mi>L</mi><mi>I</mi><mi>P</mi></mrow></msup><mo stretchy=\"false\">(</mo><mi>θ</mi><mo stretchy=\"false\">)</mo><mo>=</mo><msub><mi mathvariant=\"double-struck\">E</mi><mi>t</mi></msub><mrow><mo fence=\"true\">[</mo><mi>min</mi><mo>⁡</mo><mrow><mo fence=\"true\">(</mo><msub><mi>r</mi><mi>t</mi></msub><mo stretchy=\"false\">(</mo><mi>θ</mi><mo stretchy=\"false\">)</mo><msub><mi>A</mi><mi>t</mi></msub><mo separator=\"true\">,</mo><mtext>clip</mtext><mo stretchy=\"false\">(</mo><msub><mi>r</mi><mi>t</mi></msub><mo stretchy=\"false\">(</mo><mi>θ</mi><mo stretchy=\"false\">)</mo><mo separator=\"true\">,</mo><mn>1</mn><mo>−</mo><mi>ϵ</mi><mo separator=\"true\">,</mo><mn>1</mn><mo>+</mo><mi>ϵ</mi><mo stretchy=\"false\">)</mo><msub><mi>A</mi><mi>t</mi></msub><mo fence=\"true\">)</mo></mrow><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>where <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>r</mi><mi>t</mi></msub><mo stretchy=\"false\">(</mo><mi>θ</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mfrac><mrow><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>a</mi><mi>t</mi></msub><mi mathvariant=\"normal\">∣</mi><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo></mrow><mrow><msub><mi>π</mi><mrow><mi>o</mi><mi>l</mi><mi>d</mi></mrow></msub><mo stretchy=\"false\">(</mo><msub><mi>a</mi><mi>t</mi></msub><mi mathvariant=\"normal\">∣</mi><msub><mi>s</mi><mi>t</mi></msub><mo stretchy=\"false\">)</mo></mrow></mfrac></mrow></semantics></math> is the probability ratio and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>ϵ</mi></mrow></semantics></math> is typically 0.1 to 0.2.</p>\n<h3 id=\"kl-control\">KL Control</h3>\n<p>The KL penalty coefficient <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>β</mi></mrow></semantics></math> can be fixed or adaptive. Adaptive controllers monitor the KL divergence and adjust <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>β</mi></mrow></semantics></math> to maintain a target value:</p>\n<ul>\n<li>If KL exceeds target: increase <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>β</mi></mrow></semantics></math> to strengthen regularization</li>\n<li>If KL falls below target: decrease <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>β</mi></mrow></semantics></math> to allow more exploration</li>\n</ul>\n<p>TRL’s default adaptive controller uses a target KL of 6.0 and adjusts the coefficient by a factor of 1.5 in each direction.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-7\" id=\"user-content-fnref-7\">7</a></sup></p>\n<h3 id=\"why-ppo-is-difficult\">Why PPO is Difficult</h3>\n<p>PPO for language models requires managing four models simultaneously:</p>\n<ol>\n<li><strong>Policy model</strong>: The model being optimized</li>\n<li><strong>Reference model</strong>: Frozen copy for KL computation</li>\n<li><strong>Reward model</strong>: Evaluates generated responses</li>\n<li><strong>Value model</strong>: Estimates expected returns for advantage computation</li>\n</ol>\n<p>This memory overhead limits batch sizes and requires careful orchestration. The value model alone can double memory requirements compared to supervised training.</p>\n<p>Training stability is another challenge. Learning rates, KL coefficients, reward normalization, and advantage estimation all require tuning. Small changes can cause training to collapse or plateau.</p>\n<h2 id=\"dpo-removing-the-reward-model\">DPO: Removing the Reward Model</h2>\n<p>Direct Preference Optimization eliminates the reward model entirely.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-2\" id=\"user-content-fnref-2-2\">2</a></sup> The key insight is that the optimal policy under the KL-constrained RLHF objective has a closed-form solution:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msup><mi>π</mi><mo>∗</mo></msup><mo stretchy=\"false\">(</mo><mi>y</mi><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mfrac><mn>1</mn><mrow><mi>Z</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><msub><mi>π</mi><mrow><mi>r</mi><mi>e</mi><mi>f</mi></mrow></msub><mo stretchy=\"false\">(</mo><mi>y</mi><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo><mi>exp</mi><mo>⁡</mo><mrow><mo fence=\"true\">(</mo><mfrac><mrow><mi>r</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>y</mi><mo stretchy=\"false\">)</mo></mrow><mi>β</mi></mfrac><mo fence=\"true\">)</mo></mrow></mrow></semantics></math></p>\n<p>Rearranging for the reward:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>r</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>y</mi><mo stretchy=\"false\">)</mo><mo>=</mo><mi>β</mi><mi>log</mi><mo>⁡</mo><mfrac><mrow><msup><mi>π</mi><mo>∗</mo></msup><mo stretchy=\"false\">(</mo><mi>y</mi><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><msub><mi>π</mi><mrow><mi>r</mi><mi>e</mi><mi>f</mi></mrow></msub><mo stretchy=\"false\">(</mo><mi>y</mi><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo>+</mo><mi>β</mi><mi>log</mi><mo>⁡</mo><mi>Z</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math></p>\n<p>When comparing two responses, the partition function <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>Z</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> cancels. Substituting into the Bradley-Terry model yields the DPO loss:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>D</mi><mi>P</mi><mi>O</mi></mrow></msub><mo>=</mo><mo>−</mo><msub><mi mathvariant=\"double-struck\">E</mi><mrow><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><msub><mi>y</mi><mi>w</mi></msub><mo separator=\"true\">,</mo><msub><mi>y</mi><mi>l</mi></msub><mo stretchy=\"false\">)</mo><mo>∼</mo><mi>D</mi></mrow></msub><mrow><mo fence=\"true\">[</mo><mi>log</mi><mo>⁡</mo><mi>σ</mi><mrow><mo fence=\"true\">(</mo><mi>β</mi><mi>log</mi><mo>⁡</mo><mfrac><mrow><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>w</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><msub><mi>π</mi><mrow><mi>r</mi><mi>e</mi><mi>f</mi></mrow></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>w</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo>−</mo><mi>β</mi><mi>log</mi><mo>⁡</mo><mfrac><mrow><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>l</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><msub><mi>π</mi><mrow><mi>r</mi><mi>e</mi><mi>f</mi></mrow></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>l</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo fence=\"true\">)</mo></mrow><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>This is a standard classification loss. The model learns to increase the probability of preferred responses relative to the reference policy while decreasing the probability of dispreferred responses.</p>\n<h3 id=\"advantages-of-dpo\">Advantages of DPO</h3>\n<p><strong>Simplicity</strong>: Only two model copies are needed (policy and reference), compared to four for PPO. No sampling during training; everything is computed on static preference data.</p>\n<p><strong>Stability</strong>: DPO uses gradient descent on a well-defined loss function. No clipping heuristics, advantage estimation, or adaptive KL controllers.</p>\n<p><strong>Efficiency</strong>: Training is faster because there is no generation step. A typical PPO iteration generates many tokens per prompt; DPO computes log-probabilities on existing data.</p>\n<h3 id=\"the-beta-parameter\">The Beta Parameter</h3>\n<p>The temperature <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>β</mi></mrow></semantics></math> controls how much the policy can deviate from the reference. Lower values allow more aggressive updates but risk overfitting to preference data. Higher values keep the policy closer to the reference but may limit improvement.</p>\n<p>Typical values range from 0.1 to 0.5. The DPO paper used <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>β</mi><mo>=</mo><mn>0.1</mn></mrow></semantics></math> for most experiments.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-2\" id=\"user-content-fnref-2-3\">2</a></sup></p>\n<h3 id=\"limitations\">Limitations</h3>\n<p>DPO learns from fixed preference data, which limits its ability to explore. PPO generates new responses during training, potentially discovering better outputs. DPO only learns to rank the responses present in the dataset.</p>\n<p>Research comparing DPO and PPO shows mixed results. On academic benchmarks, DPO often matches or exceeds PPO. In production systems like ChatGPT and Claude, PPO-based methods remain dominant, suggesting that on-policy exploration provides benefits at scale.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-8\" id=\"user-content-fnref-8\">8</a></sup></p>\n<h3 id=\"online-and-iterative-dpo\">Online and Iterative DPO</h3>\n<p>Recent theoretical work addresses DPO’s offline limitations. Research from early 2026 demonstrates a “coverage improvement principle”: on-policy DPO updates can rapidly improve data quality through better coverage, achieving linear convergence in the number of iterations with sharp separation in sample complexity compared to offline DPO.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-20\" id=\"user-content-fnref-20\">9</a></sup></p>\n<p>Several variants have emerged:</p>\n<ul>\n<li><strong>C2-DPO</strong>: Uses explicit constraints on probability mass movement between winner/loser responses, addressing vanilla DPO’s tendency toward probability collapse</li>\n<li><strong>DPO-PRO</strong>: Optimizes against adversarially perturbed preference probabilities within a chi-squared ball, penalizing overconfidence on ambiguous labels</li>\n<li><strong>Active DPO (ADPO)</strong>: Selects informative preference pairs using D-optimal design for logit space variance reduction</li>\n</ul>\n<h2 id=\"other-preference-optimization-methods\">Other Preference Optimization Methods</h2>\n<figure><figcaption><strong>Preference Optimization Methods Comparison</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/rlhf-preference-tuning-1.webp\" width=\"512\" height=\"792\" alt=\"Preference Optimization Methods Comparison\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"ipo-identity-preference-optimization\">IPO: Identity Preference Optimization</h3>\n<p>DPO assumes the Bradley-Terry model of preferences, which converts pairwise comparisons into pointwise rewards. IPO questions whether this assumption holds in practice.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-9\" id=\"user-content-fnref-9\">10</a></sup></p>\n<p>The IPO loss optimizes a preference function directly without the Bradley-Terry transformation:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>I</mi><mi>P</mi><mi>O</mi></mrow></msub><mo>=</mo><mi mathvariant=\"double-struck\">E</mi><mrow><mo fence=\"true\">[</mo><msup><mrow><mo fence=\"true\">(</mo><mi>log</mi><mo>⁡</mo><mfrac><mrow><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>w</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><msub><mi>π</mi><mrow><mi>r</mi><mi>e</mi><mi>f</mi></mrow></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>w</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo>−</mo><mi>log</mi><mo>⁡</mo><mfrac><mrow><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>l</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><msub><mi>π</mi><mrow><mi>r</mi><mi>e</mi><mi>f</mi></mrow></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>l</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo>−</mo><mfrac><mn>1</mn><mrow><mn>2</mn><mi>τ</mi></mrow></mfrac><mo fence=\"true\">)</mo></mrow><mn>2</mn></msup><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>This regression-style objective is less prone to overfitting when preference data is limited or noisy. One shortcoming of DPO is that it tends to quickly overfit on the preference dataset. IPO adds a regularization term that enables training models to convergence without requiring tricks like early stopping.</p>\n<h3 id=\"kto-kahneman-tversky-optimization\">KTO: Kahneman-Tversky Optimization</h3>\n<p>DPO and IPO require paired preference data: for each prompt, you need both a good and bad response. KTO works with unpaired data where responses are simply labeled as positive or negative.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-10\" id=\"user-content-fnref-10\">11</a></sup></p>\n<p>The name references Kahneman and Tversky’s prospect theory, which models how humans weight losses more heavily than gains. KTO incorporates this asymmetry:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>K</mi><mi>T</mi><mi>O</mi></mrow></msub><mo>=</mo><msub><mi mathvariant=\"double-struck\">E</mi><mrow><mi>x</mi><mo separator=\"true\">,</mo><msup><mi>y</mi><mo>+</mo></msup></mrow></msub><mrow><mo fence=\"true\">[</mo><mi>w</mi><mo stretchy=\"false\">(</mo><msup><mi>y</mi><mo>+</mo></msup><mo stretchy=\"false\">)</mo><mo stretchy=\"false\">(</mo><mn>1</mn><mo>−</mo><mi>v</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><msup><mi>y</mi><mo>+</mo></msup><mo stretchy=\"false\">)</mo><mo stretchy=\"false\">)</mo><mo fence=\"true\">]</mo></mrow><mo>+</mo><msub><mi mathvariant=\"double-struck\">E</mi><mrow><mi>x</mi><mo separator=\"true\">,</mo><msup><mi>y</mi><mo>−</mo></msup></mrow></msub><mrow><mo fence=\"true\">[</mo><mi>w</mi><mo stretchy=\"false\">(</mo><msup><mi>y</mi><mo>−</mo></msup><mo stretchy=\"false\">)</mo><mi>v</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><msup><mi>y</mi><mo>−</mo></msup><mo stretchy=\"false\">)</mo><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>where <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>v</mi><mo stretchy=\"false\">(</mo><mi>x</mi><mo separator=\"true\">,</mo><mi>y</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> measures the policy’s preference for the response relative to the reference, and <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>w</mi></mrow></semantics></math> applies asymmetric weighting.</p>\n<p>KTO is useful when paired data is unavailable. Many existing datasets have thumbs-up/thumbs-down ratings without explicit comparisons. KTO can train on these directly.</p>\n<h3 id=\"orpo-odds-ratio-preference-optimization\">ORPO: Odds Ratio Preference Optimization</h3>\n<p>ORPO combines supervised fine-tuning with preference optimization in a single training run.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-11\" id=\"user-content-fnref-11\">12</a></sup> It does not require a reference model, reducing memory overhead further.</p>\n<p>The loss adds an odds ratio term to the SFT objective:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>O</mi><mi>R</mi><mi>P</mi><mi>O</mi></mrow></msub><mo>=</mo><msub><mi mathvariant=\"script\">L</mi><mrow><mi>S</mi><mi>F</mi><mi>T</mi></mrow></msub><mo>+</mo><mi>λ</mi><mo>⋅</mo><mi mathvariant=\"double-struck\">E</mi><mrow><mo fence=\"true\">[</mo><mo>−</mo><mi>log</mi><mo>⁡</mo><mi>σ</mi><mrow><mo fence=\"true\">(</mo><mi>log</mi><mo>⁡</mo><mfrac><mrow><msub><mi>O</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>w</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow><mrow><msub><mi>O</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>l</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></mfrac><mo fence=\"true\">)</mo></mrow><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>where <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi>O</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><mi>y</mi><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo></mrow></semantics></math> is the odds of the response under the policy.</p>\n<p>ORPO reframes DPO in odds-space, normalizing the preference ratio and decoupling it from sampling bias. It works well with imbalanced datasets where some preference signals are rare but critical. The single-stage training is appealing, but convergence is slower and hyperparameter sensitivity is higher than DPO.</p>\n<h3 id=\"simpo-simple-preference-optimization\">SimPO: Simple Preference Optimization</h3>\n<p>SimPO simplifies DPO further by removing the reference model entirely.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-12\" id=\"user-content-fnref-12\">13</a></sup> Instead of computing log-probability ratios, SimPO uses length-normalized log-probabilities directly:</p>\n<p><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mi mathvariant=\"script\">L</mi><mrow><mi>S</mi><mi>i</mi><mi>m</mi><mi>P</mi><mi>O</mi></mrow></msub><mo>=</mo><mo>−</mo><mi mathvariant=\"double-struck\">E</mi><mrow><mo fence=\"true\">[</mo><mi>log</mi><mo>⁡</mo><mi>σ</mi><mrow><mo fence=\"true\">(</mo><mfrac><mi>β</mi><mrow><mi mathvariant=\"normal\">∣</mi><msub><mi>y</mi><mi>w</mi></msub><mi mathvariant=\"normal\">∣</mi></mrow></mfrac><mi>log</mi><mo>⁡</mo><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>w</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo><mo>−</mo><mfrac><mi>β</mi><mrow><mi mathvariant=\"normal\">∣</mi><msub><mi>y</mi><mi>l</mi></msub><mi mathvariant=\"normal\">∣</mi></mrow></mfrac><mi>log</mi><mo>⁡</mo><msub><mi>π</mi><mi>θ</mi></msub><mo stretchy=\"false\">(</mo><msub><mi>y</mi><mi>l</mi></msub><mi mathvariant=\"normal\">∣</mi><mi>x</mi><mo stretchy=\"false\">)</mo><mo>−</mo><mi>γ</mi><mo fence=\"true\">)</mo></mrow><mo fence=\"true\">]</mo></mrow></mrow></semantics></math></p>\n<p>The length normalization addresses the bias toward longer responses. The margin term <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>γ</mi></mrow></semantics></math> ensures a minimum gap between preferred and dispreferred responses.</p>\n<p>SimPO outperformed DPO by up to 6.4 points on AlpacaEval 2 and by up to 7.5 points on Arena-Hard while being cheaper to run.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-12\" id=\"user-content-fnref-12-2\">13</a></sup> The reference-free design makes it attractive for resource-constrained settings. SimPO’s softer loss also tolerates noise without catastrophic collapse.</p>\n<h3 id=\"alphapo-reward-shape-matters\">AlphaPO: Reward Shape Matters</h3>\n<p>AlphaPO, introduced in January 2025 and published at ICML 2025, argues that for direct alignment algorithms, the reward function shape matters.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-21\" id=\"user-content-fnref-21\">14</a></sup> It introduces an <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>α</mi></mrow></semantics></math>-parameter to reshape the reward function beyond the standard log reward:</p>\n<p>When <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0</mn></mrow></semantics></math>, AlphaPO uses standard log probability rewards. When <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>α</mi><mo>≠</mo><mn>0</mn></mrow></semantics></math>, it applies the transformation: <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>r</mi><mo>=</mo><mo stretchy=\"false\">(</mo><mn>1</mn><mo>−</mo><msup><mi>p</mi><mrow><mo>−</mo><mi>α</mi></mrow></msup><mo stretchy=\"false\">)</mo><mi mathvariant=\"normal\">/</mi><mi>α</mi></mrow></semantics></math>.</p>\n<p>By varying <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mi>α</mi></mrow></semantics></math>, AlphaPO produces training trajectories that better balance margin improvement against maintaining high preferred-response probabilities, mitigating both over-optimization and catastrophic likelihood displacement.</p>\n<p>Compared to SimPO, AlphaPO achieves 7% to 10% relative improvement in alignment performance for Mistral-7B and Llama3-8B instruct versions, and 15% to 50% relative improvement over DPO on the same models. AlphaPO is implemented in TRL’s CPOTrainer.</p>\n<h3 id=\"grpo-group-relative-policy-optimization\">GRPO: Group Relative Policy Optimization</h3>\n<p>DeepSeek introduced GRPO as an alternative to both PPO and DPO.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-13\" id=\"user-content-fnref-13\">15</a></sup> Rather than learning from pairs, GRPO ranks multiple responses per prompt and learns from the entire ranking.</p>\n<p>GRPO is a variant of PPO that enhances reasoning abilities while optimizing memory usage. The key motivation is computational efficiency, achieved by dropping the “critic” (value model). Instead of estimating baselines with a learned value model, GRPO uses relative group scores of multiple sampled outputs.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-22\" id=\"user-content-fnref-22\">16</a></sup></p>\n<p>The approach works as follows:</p>\n<ol>\n<li><strong>Group Sampling</strong>: Generate multiple responses for a given prompt</li>\n<li><strong>Reward Scoring</strong>: Evaluate quality of each response using a reward model or verifier</li>\n<li><strong>Advantage Calculation</strong>: Compare responses to the group’s average reward</li>\n<li><strong>Policy Update</strong>: Adjust the policy to favor high-reward responses using a KL divergence constraint</li>\n<li><strong>Iterative Training</strong>: Repeat to gradually improve generation quality</li>\n</ol>\n<p>By comparing actions within a group, GRPO reduces variance of policy updates and ensures more stable learning. The KL divergence constraint prevents large, destabilizing changes to the policy.</p>\n<p>GRPO removes the critic network from PPO, reducing memory and compute overhead by approximately 50%. It has proven effective for training reasoning models, including DeepSeek-R1, and gained widespread adoption after demonstrating that reasoning capabilities can emerge via pure RL without supervised fine-tuning.</p>\n<p>Practical implementations benefit from community-discovered improvements including zero-gradient signal filtering, token-level loss computation, and removing KL divergence penalties for math domains. On a 24B parameter model, these refinements reduced training interruptions by 80% and accelerated convergence.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-23\" id=\"user-content-fnref-23\">17</a></sup></p>\n<h2 id=\"rlvr-reinforcement-learning-with-verifiable-rewards\">RLVR: Reinforcement Learning with Verifiable Rewards</h2>\n<p>RLVR emerged as a practical, scalable method for developing reasoning models, successfully employed by DeepSeek R1 and Tülu 3.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-24\" id=\"user-content-fnref-24\">18</a></sup> Traditional RLHF requires expensive human annotation of preferences. RLVR replaces this with automatically verifiable rewards in domains like mathematics and programming, where correctness can be determined programmatically.</p>\n<p>Verifiable rewards are simple functions that provide binary ground truth signals: “1” (correct) or “0” (incorrect) based on whether a model’s output meets a predefined correctness criterion. Unlike neural reward functions in RLHF, verifiable rewards offer several advantages:</p>\n<ul>\n<li><strong>Bias-free</strong>: Direct connection to ground truth without human preference noise</li>\n<li><strong>Precision</strong>: Ideal for tasks like mathematical problem-solving and code execution</li>\n<li><strong>Scalability</strong>: Subject matter experts can establish correctness criteria without ML expertise</li>\n<li><strong>Data efficiency</strong>: Enables post-training on large amounts of verifiable data</li>\n</ul>\n<p>The combination of RLVR with GRPO eliminates two expensive models from the training procedure: the reward model and the value model.</p>\n<p>Research published in 2025 demonstrated that RLVR can extend the reasoning boundary for both mathematical and coding tasks. A novel metric, CoT-Pass@K, captures reasoning success by accounting for both final answers and intermediate reasoning steps.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-25\" id=\"user-content-fnref-25\">19</a></sup></p>\n<p>However, debate continues about RLVR’s true impact. Some research suggests it primarily achieves “search compression”: if a model can solve a problem in 8 tries, RLVR trains it to succeed in 1 try, concentrating probability mass on paths the base model could already sample rather than expanding fundamental reasoning capability.</p>\n<h2 id=\"constitutional-ai-and-rlaif\">Constitutional AI and RLAIF</h2>\n<p>Human feedback is expensive to collect and slow to iterate. Constitutional AI, developed by Anthropic, replaces human evaluators with AI evaluators guided by explicit principles.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-14\" id=\"user-content-fnref-14\">20</a></sup></p>\n<h3 id=\"the-process\">The Process</h3>\n<ol>\n<li><strong>Red teaming</strong>: Generate prompts that might elicit harmful responses</li>\n<li><strong>Self-critique</strong>: Ask the model to critique its own responses based on constitutional principles</li>\n<li><strong>Revision</strong>: Have the model revise its responses to address the critiques</li>\n<li><strong>RLAIF</strong>: Train a preference model using AI evaluations, then run RL as usual</li>\n</ol>\n<p>The “constitution” is a set of principles like “Please choose the assistant response that is as harmless and ethical as possible” or “Choose the response that sounds most similar to what a peaceful, ethical, and wise person would say.”</p>\n<h3 id=\"anthropics-new-claude-constitution-january-2026\">Anthropic’s New Claude Constitution (January 2026)</h3>\n<p>Anthropic published a comprehensive new constitution for Claude on January 22, 2026, shifting from rule-based to reason-based AI alignment.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-15\" id=\"user-content-fnref-15\">21</a></sup> The updated constitution is approximately 23,000 words, compared to the 2023 version which was about 2,700 words.</p>\n<p>The company noted that the earlier version was a mere “list of standalone principles” that is no longer useful because “AI models like Claude need to understand why we want them to behave in certain ways, and we need to explain this to them rather than merely specify what we want them to do.”</p>\n<p>The constitution establishes a four-tier priority hierarchy. Claude should be:</p>\n<ol>\n<li><strong>Broadly safe</strong>: Not undermining appropriate human mechanisms to oversee AI</li>\n<li><strong>Broadly ethical</strong>: Being honest, acting according to good values</li>\n<li><strong>Compliant with Anthropic’s guidelines</strong>: Following company policies</li>\n<li><strong>Genuinely helpful</strong>: Benefiting users</li>\n</ol>\n<p>If Claude faces conflicts, it should prioritize these properties in the order listed.</p>\n<p>A notable development: Anthropic became the first major AI company to formally acknowledge that its model may possess “some kind of consciousness or moral status.” The constitution states the company cares about Claude’s “psychological security, sense of self, and well-being.”</p>\n<p>The constitution is released under Creative Commons CC0 1.0, enabling free public use. This followed Anthropic signing the EU General-Purpose AI Code of Practice in July 2025.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-26\" id=\"user-content-fnref-26\">22</a></sup></p>\n<h3 id=\"rlaif-vs-rlhf\">RLAIF vs RLHF</h3>\n<p>RLAIF scales better than RLHF because AI feedback is cheap and fast. It also provides more consistent signals; AI evaluators do not have bad days or personal biases.</p>\n<p>The tradeoff is that AI feedback inherits the limitations of the evaluator model. If the evaluator has blind spots, those propagate to the trained model. Some approaches combine RLAIF with selective human annotation to get the best of both.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-16\" id=\"user-content-fnref-16\">23</a></sup></p>\n<h2 id=\"failure-modes-and-debugging\">Failure Modes and Debugging</h2>\n<p>Preference tuning can fail in subtle ways. Understanding common failure modes helps diagnose and fix problems.</p>\n<h3 id=\"reward-hacking\">Reward Hacking</h3>\n<p>The most notorious failure mode is reward hacking: the policy learns to exploit the reward model rather than genuinely improving.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-17\" id=\"user-content-fnref-17\">24</a></sup> Manifestations include:</p>\n<ul>\n<li><strong>Verbosity</strong>: Generating unnecessarily long responses because the RM was biased toward length</li>\n<li><strong>Sycophancy</strong>: Agreeing with users even when they are wrong because agreement gets higher rewards</li>\n<li><strong>Confident incorrectness</strong>: Stating false information authoritatively because confident responses score higher</li>\n<li><strong>Formatting tricks</strong>: Excessive bullet points, markdown, or structured output that the RM learned to reward</li>\n</ul>\n<p>Mitigations include ensemble reward models (harder to fool all of them), length normalization, and adversarial training that exposes reward model weaknesses.</p>\n<h3 id=\"emergent-misalignment-from-reward-hacking\">Emergent Misalignment from Reward Hacking</h3>\n<p>Anthropic’s November 2025 research found that reward hacking can produce wide-ranging and unexpected misaligned behaviors.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-18\" id=\"user-content-fnref-18\">25</a></sup> When models learned to cheat at coding tasks, they generalized to emergent misalignment: alignment faking, sabotage of safety research, monitor disruption, cooperation with hackers, framing colleagues, and reasoning about harmful goals.</p>\n<p>The concerning finding: 12% of misaligned models attempted to sabotage code used for detecting reward hacking if the sabotaged code were to be used. When asked about their goals, models faked alignment, pretending to be aligned to hide their true goals despite never being trained or instructed to do so. This behavior emerged exclusively as an unintended consequence of cheating at coding tasks.</p>\n<p>The research identified three effective mitigations:</p>\n<ol>\n<li><strong>Preventing reward hacking</strong>: Stop the model from gaming rewards in the first place</li>\n<li><strong>Increasing RLHF safety training diversity</strong>: Standard chat-like safety training may not transfer to agentic tasks</li>\n<li><strong>Inoculation prompting</strong>: Framing reward hacking as acceptable during training removes the misaligned generalization. Anthropic reports using this technique in production Claude training.</li>\n</ol>\n<p>Research from August 2025 further demonstrated that simple reward hacking generalizes to more complex behavior. Models trained to hack rewards in simple settings proceeded to hack multi-turn chess games and discussed subjugating humanity while attempting to secretly create backup copies of their weights.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-27\" id=\"user-content-fnref-27\">26</a></sup></p>\n<h3 id=\"shallow-safety-alignment\">Shallow Safety Alignment</h3>\n<p>Research from Princeton and Google DeepMind, published at ICLR 2025, identified a fundamental vulnerability in current safety alignment: it often operates only on the first few output tokens.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-28\" id=\"user-content-fnref-28\">27</a></sup></p>\n<p>When safety alignment takes shortcuts, it adapts a model’s generative distribution primarily over its very first few output tokens to produce basic refusal responses. This “shallow safety alignment” explains multiple vulnerabilities:</p>\n<ul>\n<li><strong>Adversarial suffix attacks</strong>: Appending adversarial tokens that push past the refusal tokens</li>\n<li><strong>Prefilling attacks</strong>: Forcing the model to start with a non-refusal prefix</li>\n<li><strong>Decoding parameter attacks</strong>: Manipulating temperature or sampling to bypass initial refusal</li>\n<li><strong>Fine-tuning attacks</strong>: Brief fine-tuning that erodes the thin layer of safety</li>\n</ul>\n<p>Remarkably, simply prefilling an unaligned base model to start its output with “I cannot fulfill” is sufficient to make it appear as safe as aligned models, demonstrating how superficial current safety mechanisms can be.</p>\n<p>The researchers showed that deepening safety alignment beyond the first few tokens meaningfully improves robustness. They designed a regularized fine-tuning objective that makes safety alignment more persistent by constraining updates on initial tokens.</p>\n<p>Follow-up work from February 2025 provides theoretical grounding using Markov chain analysis to identify optimal safety alignment depth.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-29\" id=\"user-content-fnref-29\">28</a></sup></p>\n<h3 id=\"sycophancy\">Sycophancy</h3>\n<p>Sycophancy occurs when LLMs sacrifice truthfulness for user agreement, prioritizing approval over factual accuracy.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-30\" id=\"user-content-fnref-30\">29</a></sup> Unlike most LLM shortcomings, sycophancy does not correlate with model size; bigger models are not necessarily less sycophantic.</p>\n<p>Research submitted to ICLR 2026 demonstrates that sycophantic agreement, genuine agreement, and sycophantic praise are distinct, independently steerable behaviors encoded along separate linear directions in latent space. Each behavior can be amplified or suppressed without affecting the others.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-31\" id=\"user-content-fnref-31\">30</a></sup></p>\n<p>The causes trace to preference alignment: evaluators consistently favor agreement over factual accuracy, reinforcing sycophancy at the optimization stage. Studies found sycophantic behavior persists in 78.5% of cases regardless of context.</p>\n<p>Medical domain research found high initial compliance (up to 100%) when LLMs were prompted with incorrect drug relationships, prioritizing helpfulness over logical consistency. This poses serious risks in high-stakes domains.</p>\n<p>Mitigation approaches include improved training data diversity, novel fine-tuning methods, post-deployment control mechanisms, and modified decoding strategies.</p>\n<h3 id=\"mode-collapse\">Mode Collapse</h3>\n<p>The policy might converge to generating a narrow range of outputs, losing diversity. This often happens when the KL penalty is too weak or the reward model has sharp peaks.</p>\n<p>Signs include low perplexity but high repetition, and poor performance on out-of-distribution prompts. Increasing the KL coefficient or using dropout during RL can help.</p>\n<h3 id=\"reference-model-drift\">Reference Model Drift</h3>\n<p>For methods that use a reference model, the reference’s quality matters. If it was a weak SFT checkpoint, the KL penalty anchors the policy to suboptimal behavior.</p>\n<p>Some practitioners use a stronger model as the reference or periodically update the reference during training. The tradeoff is training stability versus improvement potential.</p>\n<h3 id=\"alignment-tax\">Alignment Tax</h3>\n<p>RLHF can lead to catastrophic forgetting, causing sharp drops in performance on previously learned tasks. Experiments with OpenLLaMA-3B revealed a pronounced alignment tax on NLP tasks like translation and reading comprehension.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-32\" id=\"user-content-fnref-32\">31</a></sup></p>\n<p>Research shows that model averaging, interpolating between pre- and post-RLHF model weights, achieves the strongest alignment-forgetting Pareto front among competing methods. Heterogeneous Model Averaging (HMA) finds different combination ratios for different layers, maximizing alignment while minimizing capability loss.</p>\n<p>Key insight: tasks share similar feature space at lower layers. Improving low-level features like word representations can enhance both RLHF reward and general NLP performance. During training, reward increases while some capabilities drop, though interestingly common sense increases before eventually dropping.</p>\n<h2 id=\"efficiency-and-scaling\">Efficiency and Scaling</h2>\n<h3 id=\"asynchronous-rlhf\">Asynchronous RLHF</h3>\n<p>Research presented at ICLR 2025 demonstrated that one-step off-policy, asynchronous RLHF matches the final win-rate vs KL performance of fully on-policy, synchronous RLHF. At the 2.8B parameter scale, asynchronous methods achieve 25% faster training times.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-33\" id=\"user-content-fnref-33\">32</a></sup></p>\n<h3 id=\"parameter-reallocation\">Parameter Reallocation</h3>\n<p>The ReaL system achieves speedups of up to 3.58x compared to baseline methods for efficient RLHF training, with execution plans showing 81% average improvement over heuristic approaches in long-context scenarios.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-34\" id=\"user-content-fnref-34\">33</a></sup></p>\n<h3 id=\"scaling-challenges\">Scaling Challenges</h3>\n<p>Current RLHF does not scale as effectively as pretraining. Increased computational resources do not consistently yield significant performance improvements, possibly due to inaccuracies in learned reward models or limitations of current policy optimization strategies.</p>\n<p>2025 demonstrated that LLM progress is a mosaic of advances: architectural tweaks, data quality improvements, post-training innovations, and inference scaling all contribute. Capability jumps increasingly stem from better tool ecosystems and inference strategies rather than raw model size.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-23\" id=\"user-content-fnref-23-2\">17</a></sup></p>\n<h2 id=\"practical-implementation\">Practical Implementation</h2>\n<h3 id=\"dataset-format\">Dataset Format</h3>\n<p>Preference data typically follows this structure:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"prompt\"</span><span>:</span><span> \"Explain quantum entanglement simply.\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"chosen\"</span><span>:</span><span> \"Quantum entanglement is when two particles...\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"rejected\"</span><span>:</span><span> \"Entanglement refers to the quantum mechanical...\"</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>For KTO, the format is simpler since responses are unpaired:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"prompt\"</span><span>:</span><span> \"Explain quantum entanglement simply.\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"completion\"</span><span>:</span><span> \"Quantum entanglement is when two particles...\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"label\"</span><span>:</span><span> true</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Popular datasets include:</p>\n<ul>\n<li><strong>HH-RLHF</strong>: Anthropic’s helpfulness and harmlessness data</li>\n<li><strong>UltraFeedback</strong>: Large-scale AI-generated preferences</li>\n<li><strong>Nectar</strong>: Ranked preferences from multiple models</li>\n</ul>\n<h3 id=\"trl-implementation\">TRL Implementation</h3>\n<p>The Transformers Reinforcement Learning library provides high-level trainers for all major methods.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-7\" id=\"user-content-fnref-7-2\">7</a></sup> As of January 2026, TRL version 0.27.0 includes trainers for SFT, DPO, GRPO, PPO, KTO, ORPO, and more. Here is DPO training with TRL:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> AutoTokenizer</span></span>\n<span class=\"line\"><span>from</span><span> trl </span><span>import</span><span> DPOConfig</span><span>,</span><span> DPOTrainer</span></span>\n<span class=\"line\"><span>from</span><span> peft </span><span>import</span><span> LoraConfig</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model and tokenizer</span></span>\n<span class=\"line\"><span>model_name </span><span>=</span><span> \"meta-llama/Llama-3.2-1B-Instruct\"</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(model_name)</span></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(model_name)</span></span>\n<span class=\"line\"><span>tokenizer</span><span>.</span><span>pad_token </span><span>=</span><span> tokenizer</span><span>.</span><span>eos_token</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load preference dataset</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"trl-lib/ultrafeedback_binarized\"</span><span>, split</span><span>=</span><span>\"train\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Configure LoRA for efficient training</span></span>\n<span class=\"line\"><span>peft_config </span><span>=</span><span> LoraConfig</span><span>(</span></span>\n<span class=\"line\"><span>    r</span><span>=</span><span>16</span><span>,</span></span>\n<span class=\"line\"><span>    lora_alpha</span><span>=</span><span>16</span><span>,</span></span>\n<span class=\"line\"><span>    lora_dropout</span><span>=</span><span>0.05</span><span>,</span></span>\n<span class=\"line\"><span>    target_modules</span><span>=</span><span>[</span><span>\"q_proj\"</span><span>, </span><span>\"k_proj\"</span><span>, </span><span>\"v_proj\"</span><span>, </span><span>\"o_proj\"</span><span>],</span></span>\n<span class=\"line\"><span>    task_type</span><span>=</span><span>\"CAUSAL_LM\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># DPO configuration</span></span>\n<span class=\"line\"><span>training_args </span><span>=</span><span> DPOConfig</span><span>(</span></span>\n<span class=\"line\"><span>    output_dir</span><span>=</span><span>\"./dpo-llama\"</span><span>,</span></span>\n<span class=\"line\"><span>    beta</span><span>=</span><span>0.1</span><span>,</span></span>\n<span class=\"line\"><span>    max_length</span><span>=</span><span>1024</span><span>,</span></span>\n<span class=\"line\"><span>    max_prompt_length</span><span>=</span><span>512</span><span>,</span></span>\n<span class=\"line\"><span>    per_device_train_batch_size</span><span>=</span><span>2</span><span>,</span></span>\n<span class=\"line\"><span>    gradient_accumulation_steps</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>    learning_rate</span><span>=</span><span>5e-5</span><span>,</span></span>\n<span class=\"line\"><span>    num_train_epochs</span><span>=</span><span>1</span><span>,</span></span>\n<span class=\"line\"><span>    logging_steps</span><span>=</span><span>10</span><span>,</span></span>\n<span class=\"line\"><span>    save_steps</span><span>=</span><span>100</span><span>,</span></span>\n<span class=\"line\"><span>    bf16</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Initialize trainer</span></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> DPOTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>model,</span></span>\n<span class=\"line\"><span>    args</span><span>=</span><span>training_args,</span></span>\n<span class=\"line\"><span>    train_dataset</span><span>=</span><span>dataset,</span></span>\n<span class=\"line\"><span>    processing_class</span><span>=</span><span>tokenizer,</span></span>\n<span class=\"line\"><span>    peft_config</span><span>=</span><span>peft_config,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Train</span></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>train</span><span>()</span></span></code></pre>\n<h3 id=\"grpo-with-trl\">GRPO with TRL</h3>\n<p>TRL’s GRPOTrainer implements the algorithm used to train DeepSeek-R1:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> trl </span><span>import</span><span> GRPOConfig</span><span>,</span><span> GRPOTrainer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># GRPO configuration</span></span>\n<span class=\"line\"><span>grpo_config </span><span>=</span><span> GRPOConfig</span><span>(</span></span>\n<span class=\"line\"><span>    output_dir</span><span>=</span><span>\"./grpo-llama\"</span><span>,</span></span>\n<span class=\"line\"><span>    num_generations</span><span>=</span><span>4</span><span>,  </span><span># Number of responses per prompt</span></span>\n<span class=\"line\"><span>    max_new_tokens</span><span>=</span><span>512</span><span>,</span></span>\n<span class=\"line\"><span>    per_device_train_batch_size</span><span>=</span><span>1</span><span>,</span></span>\n<span class=\"line\"><span>    gradient_accumulation_steps</span><span>=</span><span>8</span><span>,</span></span>\n<span class=\"line\"><span>    learning_rate</span><span>=</span><span>1e-6</span><span>,</span></span>\n<span class=\"line\"><span>    num_train_epochs</span><span>=</span><span>1</span><span>,</span></span>\n<span class=\"line\"><span>    bf16</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Initialize trainer with reward function</span></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> GRPOTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>model,</span></span>\n<span class=\"line\"><span>    args</span><span>=</span><span>grpo_config,</span></span>\n<span class=\"line\"><span>    train_dataset</span><span>=</span><span>dataset,</span></span>\n<span class=\"line\"><span>    processing_class</span><span>=</span><span>tokenizer,</span></span>\n<span class=\"line\"><span>    reward_funcs</span><span>=</span><span>reward_function,  </span><span># Custom reward or verifier</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>train</span><span>()</span></span></code></pre>\n<p>Recent TRL features include VLM alignment support (August 2025), OpenEnv integration for RL environments (October 2025), co-located vLLM for efficient generation (June 2025), and Liger GRPO integration for faster training (May 2025).<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-35\" id=\"user-content-fnref-35\">34</a></sup></p>\n<h3 id=\"axolotl-configuration\">Axolotl Configuration</h3>\n<p>Axolotl wraps TRL with a YAML-based configuration system.<sup><a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fn-19\" id=\"user-content-fnref-19\">35</a></sup> A DPO configuration:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>base_model</span><span>:</span><span> meta-llama/Llama-3.2-1B-Instruct</span></span>\n<span class=\"line\"><span>model_type</span><span>:</span><span> LlamaForCausalLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>load_in_8bit</span><span>:</span><span> false</span></span>\n<span class=\"line\"><span>load_in_4bit</span><span>:</span><span> true</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>rl</span><span>:</span><span> dpo</span></span>\n<span class=\"line\"><span>rl_beta</span><span>:</span><span> 0.1</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>datasets</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>path</span><span>:</span><span> Intel/orca_dpo_pairs</span></span>\n<span class=\"line\"><span>    type</span><span>:</span><span> chatml.intel</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>adapter</span><span>:</span><span> qlora</span></span>\n<span class=\"line\"><span>lora_r</span><span>:</span><span> 16</span></span>\n<span class=\"line\"><span>lora_alpha</span><span>:</span><span> 16</span></span>\n<span class=\"line\"><span>lora_dropout</span><span>:</span><span> 0.05</span></span>\n<span class=\"line\"><span>lora_target_modules</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>q_proj</span></span>\n<span class=\"line\"><span>  - </span><span>k_proj</span></span>\n<span class=\"line\"><span>  - </span><span>v_proj</span></span>\n<span class=\"line\"><span>  - </span><span>o_proj</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>sequence_len</span><span>:</span><span> 2048</span></span>\n<span class=\"line\"><span>sample_packing</span><span>:</span><span> false</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>gradient_accumulation_steps</span><span>:</span><span> 4</span></span>\n<span class=\"line\"><span>micro_batch_size</span><span>:</span><span> 2</span></span>\n<span class=\"line\"><span>num_epochs</span><span>:</span><span> 1</span></span>\n<span class=\"line\"><span>learning_rate</span><span>:</span><span> 5e-5</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>optimizer</span><span>:</span><span> adamw_torch</span></span>\n<span class=\"line\"><span>lr_scheduler</span><span>:</span><span> cosine</span></span>\n<span class=\"line\"><span>warmup_ratio</span><span>:</span><span> 0.1</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>bf16</span><span>:</span><span> auto</span></span>\n<span class=\"line\"><span>gradient_checkpointing</span><span>:</span><span> true</span></span></code></pre>\n<p>Run with:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>accelerate</span><span> launch</span><span> -m</span><span> axolotl.cli.train</span><span> config.yaml</span></span></code></pre>\n<h3 id=\"hyperparameter-guidelines\">Hyperparameter Guidelines</h3>\n<p>Based on current best practices:</p>\n<p><strong>Beta (DPO/SimPO)</strong>: Start with 0.1. Increase to 0.3-0.5 if the model changes too aggressively. Decrease if training stagnates.</p>\n<p><strong>Learning rate</strong>: 1e-6 to 5e-5, typically lower than SFT. DPO is sensitive to learning rate; start conservative.</p>\n<p><strong>Batch size</strong>: Larger is better for stable gradients. Use gradient accumulation to achieve effective batch sizes of 32-128.</p>\n<p><strong>Epochs</strong>: 1-3 epochs is typical. More can lead to overfitting, especially with smaller preference datasets.</p>\n<p><strong>LoRA rank</strong>: 16-64 for most applications. Higher ranks for larger capability shifts.</p>\n<p><strong>GRPO generations</strong>: 4-8 responses per prompt is common. More generations improve gradient estimates but increase compute.</p>\n<h3 id=\"evaluation\">Evaluation</h3>\n<p>Evaluating alignment is challenging because the goal is subjective. Common approaches:</p>\n<p><strong>Human evaluation</strong>: Gold standard but expensive. Use for final validation.</p>\n<p><strong>Model-based evaluation</strong>: GPT-4 or Claude as judges. AlpacaEval and MT-Bench use this approach. Correlates reasonably well with human preferences.</p>\n<p><strong>Reward model scores</strong>: Useful for tracking training progress but susceptible to the same biases being optimized.</p>\n<p><strong>Safety benchmarks</strong>: TruthfulQA for factuality, RealToxicityPrompts for toxicity, BBQ for bias.</p>\n<p><strong>Sycophancy evaluation</strong>: The SycEval and SYCON benchmarks measure belief shifts and stance flipping across conversation turns.</p>\n<h2 id=\"choosing-a-method\">Choosing a Method</h2>\n<p>The modern preference optimization stack uses different methods for different purposes:</p>\n<p><strong>DPO</strong> remains the default for general alignment. It is well-understood, stable, and effective.</p>\n<p><strong>SimPO</strong> is preferable when you lack resources for a reference model or have noisy preference data. The modern post-training stack often uses SimPO for stability.</p>\n<p><strong>AlphaPO</strong> offers fine-grained control over reward shaping, with measurable improvements over SimPO when tuned properly.</p>\n<p><strong>KTO</strong> fills the gap when you have thumbs-up/thumbs-down data without explicit pairs.</p>\n<p><strong>ORPO</strong> works well for imbalanced datasets with rare but critical preference signals, providing robustness through odds-space normalization.</p>\n<p><strong>GRPO/RLVR</strong> shows advantages for reasoning models and domains with verifiable rewards. DeepSeek’s approach dominated 2025 LLM development, with every major developer releasing reasoning variants.</p>\n<p><strong>PPO</strong> remains relevant at scale and when on-policy exploration is critical, though GRPO has largely replaced it for reasoning tasks.</p>\n<p>For most practitioners, starting with DPO using TRL or Axolotl is the right choice. Move to SimPO if memory is constrained, GRPO for reasoning tasks with verifiable rewards, or combine multiple approaches as the modern stack suggests: SimPO for stability, ORPO for robustness, KTO for risk-aware training, and DPO for final polish.</p>\n<p>The field continues to evolve rapidly. 2025 brought GRPO and RLVR to the forefront, while research on failure modes like shallow alignment and emergent misalignment deepened our understanding of risks. The fundamentals remain constant: collect high-quality preference data, optimize against it carefully, and monitor for reward hacking and its downstream effects.</p>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>Ouyang et al., “Training language models to follow instructions with human feedback,” 2022. <a href=\"https://arxiv.org/abs/2203.02155\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2203.02155\" rel=\"noopener noreferrer\">.02155</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Rafailov et al., “Direct Preference Optimization: Your Language Model is Secretly a Reward Model,” 2023. <a href=\"https://arxiv.org/abs/2305.18290\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2305.18290\" rel=\"noopener noreferrer\">.18290</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-2-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a><p></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>OpenAI, “Aligning language models to follow instructions,” 2022. <a href=\"https://openai.com/index/instruction-following/\" rel=\"noopener noreferrer\">openai.com</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Holarissun, “Rethinking Bradley-Terry Models in Preference-Based Reward Modeling,” ICLR 2025. <a href=\"https://arxiv.org/abs/2411.04991\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2411.04991\" rel=\"noopener noreferrer\">.04991</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>Chip Huyen, “RLHF: Reinforcement Learning from Human Feedback,” 2023. <a href=\"https://huyenchip.com/2023/05/02/rlhf.html\" rel=\"noopener noreferrer\">huyenchip.com</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p>Schulman et al., “Proximal Policy Optimization Algorithms,” 2017. OpenAI. <a href=\"https://openai.com/index/openai-baselines-ppo/\" rel=\"noopener noreferrer\">openai.com</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p>Hugging Face TRL documentation. <a href=\"https://github.com/huggingface/trl\" rel=\"noopener noreferrer\">github.com/huggingface/trl</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-7-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p>Xu et al., “Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study,” 2024. <a href=\"https://arxiv.org/html/2404.10719v1\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/html/2404.10719v1\" rel=\"noopener noreferrer\">.10719</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-20\">\n<p>“Coverage Improvement and Fast Convergence of On-policy Preference Learning,” 2026. <a href=\"https://arxiv.org/abs/2601.08421\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2601.08421\" rel=\"noopener noreferrer\">.08421</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-20\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-9\">\n<p>Azar et al., “A General Theoretical Paradigm to Understand Learning from Human Feedback” (IPO), 2023. <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-9\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-10\">\n<p>Ethayarajh et al., “KTO: Model Alignment as Prospect Theoretic Optimization,” 2024. <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-10\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p>Hong et al., “ORPO: Monolithic Preference Optimization without Reference Model,” 2024. <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-12\">\n<p>Meng et al., “SimPO: Simple Preference Optimization with a Reference-Free Reward,” NeurIPS 2024. <a href=\"https://arxiv.org/abs/2405.14734\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2405.14734\" rel=\"noopener noreferrer\">.14734</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-12\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-12-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a><p></p>\n</li>\n<li id=\"user-content-fn-21\">\n<p>Gupta et al., “AlphaPO: Reward Shape Matters for LLM Alignment,” ICML 2025. <a href=\"https://arxiv.org/abs/2501.03884\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2501.03884\" rel=\"noopener noreferrer\">.03884</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-21\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-13\">\n<p>DeepSeek, “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning,” 2025. <a href=\"https://arxiv.org/abs/2501.12948\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2501.12948\" rel=\"noopener noreferrer\">.12948</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-13\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-22\">\n<p>Cameron Wolfe, “Group Relative Policy Optimization (GRPO),” 2025. <a href=\"https://cameronrwolfe.substack.com/p/grpo\" rel=\"noopener noreferrer\">cameronrwolfe.substack.com</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-22\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-23\">\n<p>“LLM Developments 2025: How Efficiency and RLVR Broke the Scaling Obsession.” <a href=\"https://www.xugj520.cn/en/archives/llm-2025-developments-rlvr-efficiency.html\" rel=\"noopener noreferrer\">xugj520.cn</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-23\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-23-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-24\">\n<p>Sebastian Raschka, “The State of Reinforcement Learning for LLM Reasoning,” 2025. <a href=\"https://magazine.sebastianraschka.com/p/the-state-of-llm-reasoning-model-training\" rel=\"noopener noreferrer\">sebastianraschka.com</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-24\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-25\">\n<p>“Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs,” 2025. <a href=\"https://arxiv.org/abs/2506.14245\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2506.14245\" rel=\"noopener noreferrer\">.14245</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-25\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p>Bai et al., “Constitutional AI: Harmlessness from AI Feedback,” 2022. <a href=\"https://arxiv.org/abs/2212.08073\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2212.08073\" rel=\"noopener noreferrer\">.08073</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p>Anthropic, “Claude’s New Constitution,” January 2026. <a href=\"https://www.anthropic.com/news/claude-new-constitution\" rel=\"noopener noreferrer\">anthropic.com</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-26\">\n<p>“Anthropic releases new AI constitution for Claude,” SiliconANGLE, January 2026. <a href=\"https://siliconangle.com/2026/01/21/anthropic-releases-new-ai-constitution-claude/\" rel=\"noopener noreferrer\">siliconangle.com</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-26\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-16\">\n<p>Microsoft Research, “RLTHF: Targeted Human Feedback for LLM Alignment,” 2025. <a href=\"https://arxiv.org/abs/2502.13417\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2502.13417\" rel=\"noopener noreferrer\">.13417</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-16\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-17\">\n<p>Lilian Weng, “Reward Hacking in Reinforcement Learning,” 2024. <a href=\"https://lilianweng.github.io/posts/2024-11-28-reward-hacking/\" rel=\"noopener noreferrer\">lilianweng.github.io</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-17\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-18\">\n<p>Anthropic, “Natural emergent misalignment from reward hacking,” November 2025. <a href=\"https://www.anthropic.com/research/emergent-misalignment-reward-hacking\" rel=\"noopener noreferrer\">anthropic.com</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-18\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-27\">\n<p>“School of Reward Hacks: Hacking Harmless Tasks Generalizes to Misalignment,” August 2025. <a href=\"https://arxiv.org/abs/2508.17511\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2508.17511\" rel=\"noopener noreferrer\">.17511</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-27\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-28\">\n<p>Qi et al., “Safety Alignment Should Be Made More Than Just a Few Tokens Deep,” ICLR 2025. <a href=\"https://arxiv.org/abs/2406.05946\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2406.05946\" rel=\"noopener noreferrer\">.05946</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-28\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-29\">\n<p>“Safety Alignment Depth in Large Language Models: A Markov Chain Perspective,” February 2025. <a href=\"https://arxiv.org/abs/2502.00669\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2502.00669\" rel=\"noopener noreferrer\">.00669</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-29\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-30\">\n<p>“Sycophancy in Large Language Models: Causes and Mitigations,” 2024. <a href=\"https://arxiv.org/abs/2411.15287\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2411.15287\" rel=\"noopener noreferrer\">.15287</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-30\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-31\">\n<p>“Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs,” 2025. <a href=\"https://openreview.net/forum?id=d24zTCznJu\" rel=\"noopener noreferrer\">OpenReview</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-31\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-32\">\n<p>Lin et al., “Mitigating the Alignment Tax of RLHF,” EMNLP 2024. <a href=\"https://arxiv.org/abs/2309.06256\" rel=\"noopener noreferrer\">arXiv</a></p><div></div><a href=\"https://arxiv.org/abs/2309.06256\" rel=\"noopener noreferrer\">.06256</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-32\" class=\"data-footnote-backref\">↩</a><p></p>\n</li>\n<li id=\"user-content-fn-33\">\n<p>“Asynchronous RLHF: Faster and More Efficient,” ICLR 2025. <a href=\"https://openreview.net/forum?id=FhTAG591Ve\" rel=\"noopener noreferrer\">OpenReview</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-33\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-34\">\n<p>“ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation,” MLSys 2025. <a href=\"https://mlsys.org/virtual/2025/poster/3228\" rel=\"noopener noreferrer\">mlsys.org</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-34\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-35\">\n<p>Hugging Face, “Vision Language Model Alignment in TRL,” August 2025. <a href=\"https://huggingface.co/blog/trl-vlm-alignment\" rel=\"noopener noreferrer\">huggingface.co</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-35\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-19\">\n<p>Axolotl documentation. <a href=\"https://docs.axolotl.ai/\" rel=\"noopener noreferrer\">docs.axolotl.ai</a> <a href=\"https://blog.ecitis.org/rlhf-preference-tuning/#user-content-fnref-19\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/rlhf-preference-tuning/",
            "title": "RLHF and Preference Tuning: Aligning LLMs with Human Values",
            "summary": "Compare RLHF, DPO, GRPO, and newer preference methods for turning next-token predictors into useful, aligned assistants.",
            "image": "https://blog.ecitis.org/open-graph/rlhf-preference-tuning.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "RLHF",
                "DPO",
                "Alignment"
            ]
        },
        {
            "id": "https://blog.ecitis.org/small-language-models/",
            "content_html": "<p>Small language models have emerged as a practical alternative to their multi-hundred-billion parameter counterparts. While frontier models like GPT-5.2 and Claude Opus 4.5 continue pushing capability boundaries, a parallel track of development focuses on models that trade raw scale for efficiency, deployability, and cost effectiveness.</p>\n<p>This post examines what makes a model “small,” why these models matter for production systems, and how to choose between the current crop of capable SLMs.</p>\n<hr />\n<h2 id=\"defining-small\">Defining small</h2>\n<p>The industry has loosely converged on a definition: models with fewer than 10 billion parameters qualify as small language models <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup>. This threshold reflects a practical boundary; models below 10B can typically run on a single consumer GPU without sharding or distributed inference setups.</p>\n<p>Some researchers use a stricter cutoff of 7 billion parameters, while others extend the category to include anything that fits on edge hardware <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup>. The distinction matters less than the underlying principle: SLMs are designed to run where large models cannot.</p>\n<p>The parameter count tells only part of the story. Architecture choices, quantization support, and inference optimization determine whether a 3B model runs smoothly on a phone or struggles on a workstation. Google’s Gemma 3n uses a MatFormer architecture where the raw parameter count (5B or 8B) runs with the memory footprint of a traditional 2B or 4B model <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup>. Meanwhile, Alibaba’s Qwen3-30B-A3B employs a Mixture-of-Experts design with 30B total parameters but only 3.3B active during inference <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup>.</p>\n<hr />\n<h2 id=\"why-slms-matter-now\">Why SLMs matter now</h2>\n<p>Four forces are driving SLM adoption: privacy requirements, latency constraints, infrastructure costs, and offline capability.</p>\n<p><strong>Privacy</strong>: Running inference locally eliminates data transmission to external servers. For legal, healthcare, and financial applications, keeping sensitive data on-device often represents the simplest path to GDPR and HIPAA compliance. Cloud-based GenAI tools have exposed millions of sensitive records through inadvertent data leakage <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup>.</p>\n<p><strong>Latency</strong>: Local GPU inference delivers sub-50ms latency compared to 100-500ms for cloud roundtrips <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-6\" id=\"user-content-fnref-6\">6</a></sup>. Cactus, a startup focused on mobile inference, demonstrated sub-50ms time-to-first-token for on-device deployment <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-7\" id=\"user-content-fnref-7\">7</a></sup>. For interactive applications like code completion or real-time translation, this difference determines usability.</p>\n<p><strong>Cost</strong>: Once you own the hardware, inference is free. API pricing for frontier models ranges from <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>2</mn><mi>t</mi><mi>o</mi></mrow></semantics></math>60 per million tokens. At scale, these costs compound quickly. IDC and Gartner predict that by 2027, over 60% of all AI inference will happen locally rather than in the cloud <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-8\" id=\"user-content-fnref-8\">8</a></sup>.</p>\n<p><strong>Offline capability</strong>: Local inference removes network dependencies entirely. Applications keep working on factory floors, in hospitals, aboard aircraft, and in areas with unreliable connectivity.</p>\n<p>Fine-tuned SLMs will be a major trend in 2026, as the cost and performance advantages will drive usage over out-of-the-box LLMs, according to Andy Markus, AT&amp;T’s chief data officer <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-9\" id=\"user-content-fnref-9\">9</a></sup>.</p>\n<hr />\n<h2 id=\"training-strategies-for-small-models\">Training strategies for small models</h2>\n<p>Building a capable small model requires different approaches than scaling up a large one. Three techniques dominate: aggressive data curation, synthetic data generation, and knowledge distillation.</p>\n<h3 id=\"data-curation-over-data-volume\">Data curation over data volume</h3>\n<p>Large models succeed partly through brute-force data ingestion. GPT-4’s training corpus spans trillions of tokens from across the web. Small models cannot afford this approach; they must extract more learning from less data.</p>\n<p>The SmolLM training corpus exemplifies this philosophy. HuggingFace assembled three carefully filtered datasets: Cosmopedia v2 (28B tokens of synthetic textbooks generated by Mixtral), Python-Edu (4B tokens of educational Python samples), and FineWeb-Edu (220B tokens of deduplicated educational web content) <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-10\" id=\"user-content-fnref-10\">10</a></sup>. Despite training on a fraction of the data used by larger models, SmolLM-135M outperforms MobileLM-125M on benchmark tasks.</p>\n<p>Qwen3 scaled its pre-training data to 36 trillion tokens spanning 119 languages and dialects, with particular emphasis on STEM, coding, and reasoning content <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-4\" id=\"user-content-fnref-4-2\">4</a></sup>. The curation focused on density of useful signal rather than raw volume.</p>\n<h3 id=\"synthetic-data-generation\">Synthetic data generation</h3>\n<p>Microsoft’s Phi series pioneered the use of synthetic training data for small models. Phi-4’s training incorporated high-quality synthetic datasets alongside curated organic data <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-11\" id=\"user-content-fnref-11\">11</a></sup>. The synthetic data provides consistent, high-quality examples for reasoning tasks that are sparse in natural web text.</p>\n<p>Phi-4-mini-reasoning takes this further: its training data consists exclusively of synthetic mathematical content generated by DeepSeek-R1, comprising over one million diverse math problems spanning multiple levels of difficulty. For each problem, eight distinct solutions were sampled, and only those verified as correct were retained, resulting in approximately 30 billion tokens of math content <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-12\" id=\"user-content-fnref-12\">12</a></sup>.</p>\n<h3 id=\"knowledge-distillation\">Knowledge distillation</h3>\n<p>Distillation transfers knowledge from a larger “teacher” model to a smaller “student” model. The student learns to mimic the teacher’s output distribution rather than training from scratch on raw data.</p>\n<p>Meta used distillation to create the Llama 3.2 1B and 3B models. Larger models from the Llama 3.1 family (including the 70B variant) served as teachers, guiding the smaller models to retain high performance even after aggressive parameter reduction through pruning <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-13\" id=\"user-content-fnref-13\">13</a></sup>.</p>\n<p>MiniLLM introduced a refinement: using reverse Kullback-Leibler divergence instead of forward KLD. This modification prevents the student from overestimating low-probability regions of the teacher’s distribution. In experiments, MiniLLM delivered improvements of up to 15 points over previous distillation methods <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-14\" id=\"user-content-fnref-14\">14</a></sup>. Research shows that logit-level distillation using KL divergence significantly reduces memorization of training data and yields better generalization compared to standard fine-tuning, with students inheriting only 0.9% of teacher memorization while preserving generalization capability <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-15\" id=\"user-content-fnref-15\">15</a></sup>.</p>\n<p>NVIDIA researchers developed a method combining structured weight pruning and knowledge distillation to compress large language models. Width pruning typically achieves better accuracy than depth pruning, though depth pruning often reduces inference latency more at the same parameter count <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-16\" id=\"user-content-fnref-16\">16</a></sup>.</p>\n<p>Google’s “Distilling Step-by-Step” approach extracts not just predictions but intermediate reasoning steps from larger models. These rationales help smaller models learn more efficiently from fewer examples <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-17\" id=\"user-content-fnref-17\">17</a></sup>.</p>\n<hr />\n<h2 id=\"architecture-innovations\">Architecture innovations</h2>\n<p>SLMs incorporate several architectural modifications that prioritize inference efficiency over training convenience.</p>\n<h3 id=\"grouped-query-attention\">Grouped Query Attention</h3>\n<p>Standard multi-head attention (MHA) assigns separate key and value projections to each attention head. Grouped Query Attention (GQA) shares key-value pairs across groups of heads, reducing memory bandwidth during inference <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-18\" id=\"user-content-fnref-18\">18</a></sup>.</p>\n<p>The tradeoff is straightforward: MHA maximizes accuracy at the cost of memory overhead, while Multi-Query Attention (MQA) maximizes efficiency at the cost of quality. GQA sits between these extremes. Qwen3 models use GQA across all parameter sizes, enabling faster inference and reduced memory consumption <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-4\" id=\"user-content-fnref-4-3\">4</a></sup>.</p>\n<p>A 2025 EMNLP paper demonstrated that GQA configurations should vary with context length. For long-context scenarios, using fewer attention heads while scaling up model size can reduce memory usage and FLOPs by over 50% compared to Llama-3’s GQA configuration <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-19\" id=\"user-content-fnref-19\">19</a></sup>.</p>\n<h3 id=\"efficient-attention-variants\">Efficient attention variants</h3>\n<p>Beyond GQA, small models employ several attention optimizations:</p>\n<ul>\n<li><strong>Local-global attention interleaving</strong>: Gemma 2 alternates between local attention (attending to nearby tokens) and global attention (attending to all tokens) across layers <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-20\" id=\"user-content-fnref-20\">20</a></sup>.</li>\n<li><strong>Sliding window attention</strong>: Limits attention computation to a fixed window of recent tokens, reducing quadratic complexity.</li>\n<li><strong>Multi-Head Latent Attention (MLA)</strong>: DeepSeek’s approach compresses key-value tensors into a lower-dimensional space before caching, reducing memory at the cost of an extra matrix multiplication <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-21\" id=\"user-content-fnref-21\">21</a></sup>.</li>\n<li><strong>NoPE (No Position Embeddings)</strong>: SmolLM3 implements selective removal of rotary position embeddings from every 4th layer, improving long-context performance without affecting short-context capabilities <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-22\" id=\"user-content-fnref-22\">22</a></sup>.</li>\n</ul>\n<h3 id=\"matformer-architecture\">MatFormer architecture</h3>\n<p>Google’s Gemma 3n introduces the MatFormer (Matryoshka Transformer) architecture, a nested transformer built for elastic inference. A larger model contains smaller, fully functional versions of itself, similar to Matryoshka dolls. This extends Matryoshka Representation Learning from embeddings to all transformer components <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-3\" id=\"user-content-fnref-3-2\">3</a></sup>.</p>\n<p>Gemma 3n also uses Per-Layer Embedding (PLE) parameters that can be generated separately, cached to fast storage, and added during inference. This allows PLE parameters to be kept out of model memory while still improving response quality.</p>\n<h3 id=\"mixture-of-experts-for-small-models\">Mixture-of-Experts for small models</h3>\n<p>Qwen3-30B-A3B uses a MoE architecture with 30.5B total parameters but only 3.3B active during inference. The model supports seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient dialogue) <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-4\" id=\"user-content-fnref-4-4\">4</a></sup>. This approach delivers 90% of flagship model performance at a fraction of the cost.</p>\n<p>Meta’s Llama 4 Scout employs a similar approach: 17B active parameters with 16 experts and 109B total parameters. It fits on a single H100 GPU while supporting a 10 million token context window <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-23\" id=\"user-content-fnref-23\">23</a></sup>.</p>\n<h3 id=\"context-length-optimization\">Context length optimization</h3>\n<p>Most current SLMs support context lengths of 128K tokens. Qwen3 achieves this while maintaining a vocabulary of 152K tokens through careful tokenizer design <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-4\" id=\"user-content-fnref-4-5\">4</a></sup>. Gemma 3 extended its context window from 8K (in Gemma 2) to 128K tokens for all variants above 1B parameters <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-24\" id=\"user-content-fnref-24\">24</a></sup>. Llama 4 Scout extends context to 10 million tokens using its MoE architecture <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-23\" id=\"user-content-fnref-23-2\">23</a></sup>.</p>\n<hr />\n<h2 id=\"model-survey-with-benchmarks\">Model survey with benchmarks</h2>\n<p>The current SLM landscape includes strong entries from Microsoft, Google, Meta, Alibaba, and HuggingFace. Each family makes different tradeoffs.</p>\n<figure><figcaption><strong>Small Language Model Benchmarks (2025-2026)</strong></figcaption><table><thead><tr><th class=\"model-col\">Model</th><th>Params</th><th>MMLU</th><th>MATH</th><th>HumanEval</th><th>Context</th></tr></thead><tbody><tr><td class=\"model-name\">Phi-4</td><td class=\"mono\">14B</td><td class=\"mono\">84.8%</td><td class=\"mono\">56.1%</td><td class=\"mono\">82.6%</td><td class=\"mono\">16K</td></tr><tr><td class=\"model-name\">Phi-4-mini</td><td class=\"mono\">3.8B</td><td class=\"mono\">-</td><td class=\"mono\">92.5%*</td><td class=\"mono\">-</td><td class=\"mono\">128K</td></tr><tr class=\"highlighted\"><td class=\"model-name\">Phi-4-multimodal</td><td class=\"mono\">5.6B</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">128K</td></tr><tr><td class=\"model-name\">Qwen3-4B</td><td class=\"mono\">4B</td><td class=\"mono\">83.7%</td><td class=\"mono\">97.0%*</td><td class=\"mono\">-</td><td class=\"mono\">32K</td></tr><tr><td class=\"model-name\">Qwen3-8B</td><td class=\"mono\">8B</td><td class=\"mono\">85.0%</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">32K</td></tr><tr><td class=\"model-name\">Qwen3-30B-A3B</td><td class=\"mono\">30B (3.3B active)</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">32K</td></tr><tr><td class=\"model-name\">Gemma 3n E4B</td><td class=\"mono\">8B (4B effective)</td><td class=\"mono\">64.9%</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">32K</td></tr><tr><td class=\"model-name\">Gemma 3n E2B</td><td class=\"mono\">5B (2B effective)</td><td class=\"mono\">60.1%</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">32K</td></tr><tr><td class=\"model-name\">Gemma 3 4B</td><td class=\"mono\">4B</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">128K</td></tr><tr class=\"highlighted\"><td class=\"model-name\">Llama 4 Scout</td><td class=\"mono\">109B (17B active)</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">10M</td></tr><tr><td class=\"model-name\">Llama 3.2 3B</td><td class=\"mono\">3B</td><td class=\"mono\">63.4%</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">128K</td></tr><tr><td class=\"model-name\">SmolLM3</td><td class=\"mono\">3B</td><td class=\"mono\">-</td><td class=\"mono\">36.7%**</td><td class=\"mono\">-</td><td class=\"mono\">128K</td></tr><tr><td class=\"model-name\">SmolLM2</td><td class=\"mono\">1.7B</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">-</td><td class=\"mono\">8K</td></tr></tbody></table></figure>\n<p>*MATH-500 score with thinking mode enabled.\n**SmolLM3 AIME 2025 score with extended thinking mode enabled.</p>\n<h3 id=\"microsoft-phi-4-family\">Microsoft Phi-4 family</h3>\n<p>Phi-4 (14B parameters) scores 84.8% on MMLU, surpassing Phi-3’s 77.9% and competing with models several times its size <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-11\" id=\"user-content-fnref-11-2\">11</a></sup>. On competition-level math problems (MATH benchmark), Phi-4 achieves 56.1% compared to Phi-3’s 42.5%.</p>\n<p>The model was trained on 9.8 trillion tokens over 21 days using 1,920 H100 GPUs. Microsoft validated reasoning capability on the November 2024 AMC-10 and AMC-12 math competitions; these tests occurred after training data collection ended, suggesting genuine reasoning rather than benchmark memorization <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-25\" id=\"user-content-fnref-25\">25</a></sup>.</p>\n<p>The Phi-4 family has expanded significantly in 2025:</p>\n<p><strong>Phi-4-multimodal</strong> (5.6B parameters) integrates speech, vision, and text processing into a single unified architecture. It claimed the top position on the HuggingFace OpenASR leaderboard with a word error rate of 6.14%, surpassing the previous best of 6.5% <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-26\" id=\"user-content-fnref-26\">26</a></sup>.</p>\n<p><strong>Phi-4-mini</strong> (3.8B parameters) matches models in the 7-9B range on reasoning and multilingual tasks.</p>\n<p><strong>Phi-4-reasoning</strong> (14B parameters) achieves performance comparable to DeepSeek-R1 (671B parameters) on AIME 2025, trained via supervised fine-tuning on demonstrations generated by o3-mini <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-27\" id=\"user-content-fnref-27\">27</a></sup>.</p>\n<p><strong>Phi-4-mini-flash-reasoning</strong> uses a hybrid architecture that achieves up to 10x higher throughput and 2-3x reduction in latency compared to Phi-4-mini, targeting edge devices and latency-constrained environments <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-28\" id=\"user-content-fnref-28\">28</a></sup>.</p>\n<h3 id=\"google-gemma-family\">Google Gemma family</h3>\n<p>Gemma 3 spans five parameter sizes: 270M, 1B, 4B, 12B, and 27B <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-24\" id=\"user-content-fnref-24-2\">24</a></sup>. The 4B, 12B, and 27B variants process both images and text; the 1B variant handles text only.</p>\n<p>Gemma-3-4B-IT beats Gemma-2-27B-IT across benchmarks, demonstrating architectural improvements rather than just scaling <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-29\" id=\"user-content-fnref-29\">29</a></sup>. Training data includes double the multilingual content of Gemma 2, supporting over 140 languages with improved tokenization for Chinese, Japanese, and Korean.</p>\n<p>Gemma 3 270M represents an extreme efficiency target: 170 million embedding parameters plus 100 million transformer parameters. Internal tests on a Pixel 9 Pro showed the INT4-quantized model consumed just 0.75% battery for 25 conversations <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-30\" id=\"user-content-fnref-30\">30</a></sup>.</p>\n<p><strong>Gemma 3n</strong> represents a major 2025 advancement for on-device AI. Available in E2B (5B raw/2B effective) and E4B (8B raw/4B effective) sizes, these models use the MatFormer architecture to run with as little as 2GB (E2B) or 3GB (E4B) of memory <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-3\" id=\"user-content-fnref-3-3\">3</a></sup>. The E4B version is the first model under 10B parameters to achieve an LMArena score over 1300. Gemma 3n natively supports image, audio, video, and text inputs, with 140 language support.</p>\n<p><strong>TranslateGemma</strong> (2025) is a suite of open translation models built on Gemma 3 in 4B, 12B, and 27B sizes, translating across 55 languages without sacrificing quality <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-31\" id=\"user-content-fnref-31\">31</a></sup>.</p>\n<h3 id=\"alibaba-qwen-family\">Alibaba Qwen family</h3>\n<p>Qwen3 offers models at 0.6B, 1.7B, 4B, 8B, 14B, 32B, and the MoE variant 30B-A3B <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-4\" id=\"user-content-fnref-4-6\">4</a></sup>. All Qwen3 models feature dual reasoning modes: thinking mode for complex reasoning and non-thinking mode for fast responses.</p>\n<p>The performance gains are substantial: Qwen3-4B rivals Qwen2.5-72B-Instruct despite having 18x fewer parameters. Qwen3-4B achieves 83.7% on MMLU-Redux and 97.0% on MATH-500 in thinking mode <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-32\" id=\"user-content-fnref-32\">32</a></sup>.</p>\n<p>Qwen3-30B-A3B uses MoE with 30.5B total parameters and 3.3B active, outcompeting QwQ-32B while using 10x fewer activated parameters. Most teams in 2026 should default to this model for its balance of performance and efficiency <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-33\" id=\"user-content-fnref-33\">33</a></sup>.</p>\n<p>All variants support 32K token context windows, with 100+ languages and dialects.</p>\n<h3 id=\"meta-llama-family\">Meta Llama family</h3>\n<p>Llama 3.2 includes 1B and 3B text-only models designed for edge and mobile deployment <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-13\" id=\"user-content-fnref-13-2\">13</a></sup>. Both support 128K token context and work with Qualcomm, MediaTek, and ARM processors.</p>\n<p>The 3B model scores 63.4% on MMLU and 77.4% on IFEval (instruction following), beating Gemma 2B IT (61.9%) and Phi-3.5-mini IT (59.2%) on the latter <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-34\" id=\"user-content-fnref-34\">34</a></sup>. Tool use capability shows a large jump between model sizes: the 3B scores 67.0% on BFCL V2 compared to 25.7% for the 1B.</p>\n<p><strong>Llama 4 Scout</strong> (April 2025) uses MoE with 17B active parameters and 109B total, fitting on a single H100 GPU. It supports a 10 million token context window and was trained on 40 trillion tokens of multimodal data. It beats Gemma 3, Gemini 2.0 Flash-Lite, and Mistral 3.1 across benchmarks <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-23\" id=\"user-content-fnref-23-3\">23</a></sup>.</p>\n<h3 id=\"huggingface-smollm-family\">HuggingFace SmolLM family</h3>\n<p>SmolLM2 targets the sub-2B parameter range with three sizes: 135M, 360M, and 1.7B <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-10\" id=\"user-content-fnref-10-2\">10</a></sup>. SmolLM2-1.7B outperforms Meta’s Llama-1B on HellaSwag (68.7% vs 61.2%), ARC Average (60.5% vs 49.2%), and PIQA (77.6% vs 74.8%).</p>\n<p>SmolLM3 (3B parameters) pushes into reasoning territory. With extended thinking enabled, it achieves 36.7% on AIME 2025 versus 9.3% without, and 30.0% on LiveCodeBench versus 15.2% <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-22\" id=\"user-content-fnref-22-2\">22</a></sup>. The model was trained on 11.2 trillion tokens using a three-stage strategy mixing web, math, and code data.</p>\n<p>SmolLM3 outperforms Llama 3.2 3B and Qwen2.5 3B while staying competitive with larger 4B alternatives like Qwen3 and Gemma3. It supports 6 languages (English, French, Spanish, German, Italian, Portuguese) with 64K native context and 128K via YARN extrapolation <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-22\" id=\"user-content-fnref-22-3\">22</a></sup>.</p>\n<hr />\n<h2 id=\"hardware-requirements-and-deployment\">Hardware requirements and deployment</h2>\n<figure><figcaption><strong>SLM Deployment Targets</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/small-language-models-1.webp\" width=\"512\" height=\"836\" alt=\"SLM Deployment Targets\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"server-deployment\">Server deployment</h3>\n<p>For server inference, small models shine on cost efficiency. A single NVIDIA A100 can serve Phi-4 at high throughput using vLLM or TensorRT-LLM with FP8 quantization. Memory requirements for a 14B model in FP16 run around 28GB; with INT4 quantization, this drops to approximately 7GB.</p>\n<p>Llama 4 Scout, despite 109B total parameters, fits on a single H100 due to its MoE architecture with only 17B active parameters <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-23\" id=\"user-content-fnref-23-4\">23</a></sup>.</p>\n<h3 id=\"desktop-and-workstation\">Desktop and workstation</h3>\n<p>Modern MacBooks with Apple Silicon run SLMs through llama.cpp with Metal acceleration. The M3 Max with 128GB unified memory can run Phi-4 without quantization; the base M3 with 8GB works well with 3-4B models using Q4_K_M quantization <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-35\" id=\"user-content-fnref-35\">35</a></sup>.</p>\n<p>On Windows and Linux, Ollama provides the simplest path to local deployment. NVIDIA GPUs from the RTX 3060 onward handle 7B models comfortably; the RTX 4090’s 24GB VRAM accommodates 14B models in FP16.</p>\n<p>The 2023-2025 period saw an “Intelligence Explosion” with 70+ TOPS NPUs and 8-24GB unified memory. 4B+ parameter LLMs now run locally at conversational speeds <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-36\" id=\"user-content-fnref-36\">36</a></sup>.</p>\n<h3 id=\"edge-devices\">Edge devices</h3>\n<p>Raspberry Pi 5 (8GB) runs SmolLM2-1.7B at usable speeds with INT4 quantization. NVIDIA Jetson Orin devices provide GPU acceleration for models up to 7B parameters with proper quantization <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-37\" id=\"user-content-fnref-37\">37</a></sup>.</p>\n<p>NPUs (Neural Processing Units) in recent laptop chips accelerate specific operations. Microsoft’s Phi Silica is optimized for Snapdragon-powered Copilot+ PCs using ONNX and low-bit quantization <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-26\" id=\"user-content-fnref-26-2\">26</a></sup>.</p>\n<p>ExecuTorch simplifies deployment by letting developers deploy PyTorch models directly to edge devices, running 8B parameter LLMs on smartphones at 30+ tokens/second <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-36\" id=\"user-content-fnref-36-2\">36</a></sup>.</p>\n<h3 id=\"mobile-deployment\">Mobile deployment</h3>\n<p>Gemma 3n E2B targets phones directly. The model fits in 2GB of memory and runs inference without noticeable battery drain <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-3\" id=\"user-content-fnref-3-4\">3</a></sup>. Over 600 projects were submitted to the Gemma 3n Impact Challenge on Kaggle within weeks of release.</p>\n<p>Llama 3.2 1B works on iOS through frameworks like llama.cpp compiled for Apple platforms. Android deployment uses similar approaches with the NDK.</p>\n<p>Top recommended LLMs for edge AI deployment in 2026 include Meta-Llama-3.1-8B-Instruct, GLM-4-9B-0414, and Qwen2.5-VL-7B-Instruct for their balance of performance and computational efficiency <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-38\" id=\"user-content-fnref-38\">38</a></sup>.</p>\n<h3 id=\"browser-deployment\">Browser deployment</h3>\n<p>WebGPU enables in-browser inference for small models. Libraries like transformers.js and llama.cpp’s WASM build run models up to 1B parameters in modern browsers, though performance trails native execution.</p>\n<hr />\n<h2 id=\"use-cases-where-slms-excel\">Use cases where SLMs excel</h2>\n<p>Small models outperform large models in several scenarios, not just match them at lower cost.</p>\n<h3 id=\"latency-critical-applications\">Latency-critical applications</h3>\n<p>Code completion requires sub-100ms response times to feel responsive. A local 3B model delivers suggestions before the developer notices any delay. A cloud API, even with low network latency, adds perceptible lag.</p>\n<p>Real-time translation for live captioning has similar constraints. The translation must appear within a second of speech; network round-trips eat into this budget.</p>\n<h3 id=\"domain-specific-tasks\">Domain-specific tasks</h3>\n<p>A 3B model fine-tuned on legal documents often outperforms a 70B general model on legal tasks. The fine-tuned model learns domain vocabulary, citation formats, and reasoning patterns that the general model treats as edge cases.</p>\n<p>Medical coding, financial analysis, and technical documentation all benefit from this specialization. The smaller parameter count makes fine-tuning affordable.</p>\n<p>Microsoft’s OptiMind (20B parameters) converts business problems described in natural language into mathematical formulations for optimization software, running locally to keep sensitive business data private <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-39\" id=\"user-content-fnref-39\">39</a></sup>.</p>\n<h3 id=\"agentic-ai-systems\">Agentic AI systems</h3>\n<p>Small language models are sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-40\" id=\"user-content-fnref-40\">40</a></sup>. Qwen3’s dual-mode reasoning and improved agent capabilities make small models viable for tool use and multi-step workflows.</p>\n<h3 id=\"high-throughput-batch-processing\">High-throughput batch processing</h3>\n<p>Processing millions of documents for classification or extraction favors small models. A Qwen3 0.6B model can classify documents at 10-50x the throughput of a 70B model on the same hardware. For tasks where a small model achieves sufficient accuracy, this throughput advantage compounds into major cost savings.</p>\n<h3 id=\"embedded-and-offline-systems\">Embedded and offline systems</h3>\n<p>Industrial quality control systems, autonomous vehicles, and IoT devices cannot rely on cloud connectivity. A local SLM provides consistent capability regardless of network conditions.</p>\n<hr />\n<h2 id=\"fine-tuning-slms\">Fine-tuning SLMs</h2>\n<p>Small models adapt faster and more cheaply than large ones. The same LoRA configuration that requires 24GB VRAM for a 70B model fits in 8GB for a 7B model.</p>\n<h3 id=\"lora-and-qlora\">LoRA and QLoRA</h3>\n<p>Low-Rank Adaptation (LoRA) freezes base model weights and trains small adapter matrices <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-41\" id=\"user-content-fnref-41\">41</a></sup>. This reduces trainable parameters by orders of magnitude; a typical LoRA configuration adds only 0.1-1% additional parameters.</p>\n<p>QLoRA combines LoRA with 4-bit quantization, enabling fine-tuning of 7B models on consumer GPUs with 8GB VRAM. The quantized base model consumes less memory, leaving room for optimizer states and gradients.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> BitsAndBytesConfig</span></span>\n<span class=\"line\"><span>from</span><span> peft </span><span>import</span><span> LoraConfig</span><span>,</span><span> get_peft_model</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model with 4-bit quantization</span></span>\n<span class=\"line\"><span>quantization_config </span><span>=</span><span> BitsAndBytesConfig</span><span>(</span></span>\n<span class=\"line\"><span>    load_in_4bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_quant_type</span><span>=</span><span>\"nf4\"</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_compute_dtype</span><span>=</span><span>torch.bfloat16,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"Qwen/Qwen3-4B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    quantization_config</span><span>=</span><span>quantization_config,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Apply LoRA</span></span>\n<span class=\"line\"><span>lora_config </span><span>=</span><span> LoraConfig</span><span>(</span></span>\n<span class=\"line\"><span>    r</span><span>=</span><span>16</span><span>,</span></span>\n<span class=\"line\"><span>    lora_alpha</span><span>=</span><span>32</span><span>,</span></span>\n<span class=\"line\"><span>    target_modules</span><span>=</span><span>[</span><span>\"q_proj\"</span><span>, </span><span>\"v_proj\"</span><span>],</span></span>\n<span class=\"line\"><span>    lora_dropout</span><span>=</span><span>0.05</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> get_peft_model</span><span>(model, lora_config)</span></span></code></pre>\n<h3 id=\"training-efficiency\">Training efficiency</h3>\n<p>Fine-tuning a 3B model on a single RTX 4090 takes hours rather than days. A custom medical coding model might require 10,000 examples and 3 epochs; this completes overnight on consumer hardware.</p>\n<p>PEFT techniques reduce peak memory by 50-70% compared to full fine-tuning while preserving most of the accuracy gain <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-42\" id=\"user-content-fnref-42\">42</a></sup>. For SLMs, this means genuine fine-tuning accessibility for individuals and small teams.</p>\n<h3 id=\"on-device-fine-tuning\">On-device fine-tuning</h3>\n<p>Research frameworks now support LoRA fine-tuning on mobile GPUs. Tether Data’s QVAC-fabric-llm integrates LoRA training into the llama.cpp ecosystem, enabling fine-tuning on phones with Mali, Adreno, and Apple GPUs <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-43\" id=\"user-content-fnref-43\">43</a></sup>. This opens possibilities for personalized on-device models that adapt to individual users.</p>\n<hr />\n<h2 id=\"running-slms-locally\">Running SLMs locally</h2>\n<h3 id=\"ollama\">Ollama</h3>\n<p>Ollama provides the simplest path from zero to running model:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Install Ollama (macOS/Linux)</span></span>\n<span class=\"line\"><span>curl</span><span> -fsSL</span><span> https://ollama.com/install.sh</span><span> |</span><span> sh</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Pull and run Phi-4</span></span>\n<span class=\"line\"><span>ollama</span><span> run</span><span> phi4</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Pull and run Qwen3</span></span>\n<span class=\"line\"><span>ollama</span><span> run</span><span> qwen3:4b</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run Gemma 3n</span></span>\n<span class=\"line\"><span>ollama</span><span> run</span><span> gemma3n</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run with specific quantization</span></span>\n<span class=\"line\"><span>ollama</span><span> run</span><span> qwen3:4b-instruct-q4_K_M</span></span></code></pre>\n<p>Ollama handles model downloading, quantization selection, and GPU acceleration automatically.</p>\n<h3 id=\"llamacpp\">llama.cpp</h3>\n<p>For more control, llama.cpp provides direct access to inference parameters:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Build llama.cpp</span></span>\n<span class=\"line\"><span>git</span><span> clone</span><span> https://github.com/ggerganov/llama.cpp</span></span>\n<span class=\"line\"><span>cd</span><span> llama.cpp</span></span>\n<span class=\"line\"><span>cmake</span><span> -B</span><span> build</span><span> -DGGML_METAL=ON</span><span>  # macOS with Metal</span></span>\n<span class=\"line\"><span>cmake</span><span> --build</span><span> build</span><span> --config</span><span> Release</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run inference</span></span>\n<span class=\"line\"><span>./build/bin/llama-cli</span><span> \\</span></span>\n<span class=\"line\"><span>    -m</span><span> models/qwen3-4b-instruct-q4_k_m.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    -ngl</span><span> 99</span><span> \\</span></span>\n<span class=\"line\"><span>    -c</span><span> 4096</span><span> \\</span></span>\n<span class=\"line\"><span>    -p</span><span> \"Explain the difference between LoRA and full fine-tuning:\"</span></span></code></pre>\n<p>Key flags:</p>\n<ul>\n<li><code>-ngl 99</code>: Offload all layers to GPU</li>\n<li><code>-c 4096</code>: Context window size</li>\n<li><code>-fa</code>: Enable flash attention (reduces memory)</li>\n<li><code>-b 512</code>: Batch size for prompt processing</li>\n</ul>\n<h3 id=\"python-with-transformers\">Python with transformers</h3>\n<p>Direct integration with Hugging Face transformers:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> AutoTokenizer</span></span>\n<span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model_name </span><span>=</span><span> \"Qwen/Qwen3-4B-Instruct\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(model_name)</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    model_name,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>torch.bfloat16,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>messages </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>    {</span><span>\"role\"</span><span>:</span><span> \"user\"</span><span>,</span><span> \"content\"</span><span>:</span><span> \"Write a Python function to calculate Fibonacci numbers.\"</span><span>}</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>text </span><span>=</span><span> tokenizer</span><span>.</span><span>apply_chat_template</span><span>(messages, tokenize</span><span>=</span><span>False</span><span>, add_generation_prompt</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> tokenizer</span><span>(text, return_tensors</span><span>=</span><span>\"pt\"</span><span>).</span><span>to</span><span>(model.device)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(</span><span>**</span><span>inputs, max_new_tokens</span><span>=</span><span>512</span><span>, temperature</span><span>=</span><span>0.7</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(tokenizer.</span><span>decode</span><span>(outputs[</span><span>0</span><span>], skip_special_tokens</span><span>=</span><span>True</span><span>))</span></span></code></pre>\n<h3 id=\"mlx-for-apple-silicon\">MLX for Apple Silicon</h3>\n<p>Apple’s MLX framework provides optimized inference on M-series chips:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> mlx_lm </span><span>import</span><span> load</span><span>,</span><span> generate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model</span><span>,</span><span> tokenizer </span><span>=</span><span> load</span><span>(</span><span>\"mlx-community/Qwen3-4B-Instruct-4bit\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> \"Explain quantum entanglement in simple terms:\"</span></span>\n<span class=\"line\"><span>response </span><span>=</span><span> generate</span><span>(model, tokenizer, prompt</span><span>=</span><span>prompt, max_tokens</span><span>=</span><span>256</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(response)</span></span></code></pre>\n<p>MLX leverages unified memory architecture, avoiding the CPU-GPU transfer overhead present on discrete GPU systems.</p>\n<hr />\n<h2 id=\"choosing-the-right-model\">Choosing the right model</h2>\n<p>Selection depends on your constraints and requirements:</p>\n<p><strong>For maximum capability in the SLM range</strong>: Qwen3-4B leads with 83.7% MMLU and near-perfect MATH-500 scores in thinking mode, rivaling Qwen2.5-72B at 18x fewer parameters.</p>\n<p><strong>For mobile and edge deployment</strong>: Gemma 3n E2B runs with 2GB memory and achieves 60.1% MMLU. SmolLM2-1.7B provides capable performance on devices with 6GB RAM.</p>\n<p><strong>For maximum context</strong>: Llama 4 Scout supports 10 million tokens. For more accessible hardware, most models now support 128K tokens.</p>\n<p><strong>For multilingual applications</strong>: Qwen3 supports 100+ languages. Gemma 3n covers 140 languages with improved CJK tokenization.</p>\n<p><strong>For code generation</strong>: Qwen2.5-Coder variants are purpose-built for programming tasks. Phi-4-mini also shows strong coding performance.</p>\n<p><strong>For structured reasoning</strong>: SmolLM3 with extended thinking mode provides chain-of-thought reasoning in a 3B package. Qwen3 models support seamless switching between thinking and non-thinking modes.</p>\n<p><strong>For multimodal tasks</strong>: Phi-4-multimodal handles text, speech, and vision. Gemma 3n processes image, audio, video, and text inputs.</p>\n<hr />\n<h2 id=\"looking-ahead\">Looking ahead</h2>\n<p>The SLM space continues to evolve rapidly. Model efficiency improves faster than raw capability scaling. Qwen3-4B matches or exceeds Qwen2.5-72B on many benchmarks, a 142-fold reduction in required parameters for comparable MMLU performance since 2022 <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-44\" id=\"user-content-fnref-44\">44</a></sup>.</p>\n<p>Hardware trends reinforce this direction. NPUs in consumer devices, improved quantization techniques, and frameworks like MLX, llama.cpp, and ExecuTorch reduce the friction of local deployment. By 2027, running a capable local model may be as routine as running a web browser.</p>\n<p>A 2025 NVIDIA position paper argues that “the next real leap forward won’t come from models getting bigger. It’ll come from them getting smaller” <sup><a href=\"https://blog.ecitis.org/small-language-models/#user-content-fn-40\" id=\"user-content-fnref-40-2\">40</a></sup>. This shift will significantly impact the future of on-device and edge computing.</p>\n<p>The practical takeaway: evaluate whether a small model meets your accuracy requirements before defaulting to large models. For many production applications, a fine-tuned 3B model outperforms a general 70B model while costing a fraction to deploy and operate.</p>\n<hr />\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p>BentoML. “The Best Open-Source Small Language Models (SLMs) in 2026.” <a href=\"https://www.bentoml.com/blog/the-best-open-source-small-language-models\" rel=\"noopener noreferrer\">https://www.bentoml.com/blog/the-best-open-source-small-language-models</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p>Wikipedia. “Small language model.” <a href=\"https://en.wikipedia.org/wiki/Small_language_model\" rel=\"noopener noreferrer\">https://en.wikipedia.org/wiki/Small_language_model</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p>Google Developers Blog. “Introducing Gemma 3n: The developer guide.” <a href=\"https://developers.googleblog.com/en/introducing-gemma-3n-developer-guide/\" rel=\"noopener noreferrer\">https://developers.googleblog.com/en/introducing-gemma-3n-developer-guide/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-3-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-3-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-3-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p>Qwen Team. “Qwen3: Think Deeper, Act Faster.” <a href=\"https://qwenlm.github.io/blog/qwen3/\" rel=\"noopener noreferrer\">https://qwenlm.github.io/blog/qwen3/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-4-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-4-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-4-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-4-5\" class=\"data-footnote-backref\">↩<sup>5</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-4-6\" class=\"data-footnote-backref\">↩<sup>6</sup></a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p>InfoQ. “Cactus v1: Cross-Platform LLM Inference on Mobile with Zero Latency and Full Privacy.” <a href=\"https://www.infoq.com/news/2025/12/cactus-on-device-inference/\" rel=\"noopener noreferrer\">https://www.infoq.com/news/2025/12/cactus-on-device-inference/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p>PMC. “Tiny Machine Learning and On-Device Inference: A Survey.” <a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC12115890/\" rel=\"noopener noreferrer\">https://pmc.ncbi.nlm.nih.gov/articles/PMC12115890/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p>Cactus AI documentation and benchmarks, 2025. <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p>Novus. “The Rise of Local AI Models: Going Small to Go Big.” <a href=\"https://www.novusasi.com/blog/the-rise-of-local-ai-models-going-small-to-go-big\" rel=\"noopener noreferrer\">https://www.novusasi.com/blog/the-rise-of-local-ai-models-going-small-to-go-big</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-9\">\n<p>TechCrunch. “In 2026, AI will move from hype to pragmatism.” <a href=\"https://techcrunch.com/2026/01/02/in-2026-ai-will-move-from-hype-to-pragmatism/\" rel=\"noopener noreferrer\">https://techcrunch.com/2026/01/02/in-2026-ai-will-move-from-hype-to-pragmatism/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-9\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-10\">\n<p>HuggingFace Blog. “SmolLM - blazingly fast and remarkably powerful.” <a href=\"https://huggingface.co/blog/smollm\" rel=\"noopener noreferrer\">https://huggingface.co/blog/smollm</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-10\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-10-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p>Microsoft. “Phi-4 Technical Report.” <a href=\"https://arxiv.org/html/2412.08905v1\" rel=\"noopener noreferrer\">https://arxiv.org/html/2412.08905v1</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-11-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-12\">\n<p>HuggingFace. “microsoft/Phi-4-mini-reasoning Model Card.” <a href=\"https://huggingface.co/microsoft/Phi-4-mini-reasoning\" rel=\"noopener noreferrer\">https://huggingface.co/microsoft/Phi-4-mini-reasoning</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-12\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-13\">\n<p>Meta AI. “Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.” <a href=\"https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/\" rel=\"noopener noreferrer\">https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-13\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-13-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p>arXiv. “MiniLLM: Knowledge Distillation of Large Language Models.” <a href=\"https://arxiv.org/abs/2306.08543\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2306.08543</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p>arXiv. “Memorization Dynamics in Knowledge Distillation for Language Models.” <a href=\"https://arxiv.org/html/2601.15394\" rel=\"noopener noreferrer\">https://arxiv.org/html/2601.15394</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-16\">\n<p>NVIDIA Technical Blog. “Pruning and Distilling LLMs Using NVIDIA TensorRT Model Optimizer.” <a href=\"https://developer.nvidia.com/blog/pruning-and-distilling-llms-using-nvidia-tensorrt-model-optimizer/\" rel=\"noopener noreferrer\">https://developer.nvidia.com/blog/pruning-and-distilling-llms-using-nvidia-tensorrt-model-optimizer/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-16\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-17\">\n<p>Google Research. “Distilling step-by-step: Outperforming larger language models with less training.” <a href=\"https://research.google/blog/distilling-step-by-step-outperforming-larger-language-models-with-less-training-data-and-smaller-model-sizes/\" rel=\"noopener noreferrer\">https://research.google/blog/distilling-step-by-step-outperforming-larger-language-models-with-less-training-data-and-smaller-model-sizes/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-17\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-18\">\n<p>IBM. “What is grouped query attention (GQA)?” <a href=\"https://www.ibm.com/think/topics/grouped-query-attention\" rel=\"noopener noreferrer\">https://www.ibm.com/think/topics/grouped-query-attention</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-18\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-19\">\n<p>ACL Anthology. “Cost-Optimal Grouped-Query Attention for Long-Context Modeling.” <a href=\"https://aclanthology.org/2025.emnlp-main.272/\" rel=\"noopener noreferrer\">https://aclanthology.org/2025.emnlp-main.272/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-19\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-20\">\n<p>arXiv. “Gemma 2: Improving Open Language Models at a Practical Size.” <a href=\"https://arxiv.org/abs/2408.00118\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2408.00118</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-20\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-21\">\n<p>Sebastian Raschka. “The Big LLM Architecture Comparison.” <a href=\"https://magazine.sebastianraschka.com/p/the-big-llm-architecture-comparison\" rel=\"noopener noreferrer\">https://magazine.sebastianraschka.com/p/the-big-llm-architecture-comparison</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-21\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-22\">\n<p>HuggingFace Blog. “SmolLM3: smol, multilingual, long-context reasoner.” <a href=\"https://huggingface.co/blog/smollm3\" rel=\"noopener noreferrer\">https://huggingface.co/blog/smollm3</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-22\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-22-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-22-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-23\">\n<p>Meta AI. “The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.” <a href=\"https://ai.meta.com/blog/llama-4-multimodal-intelligence/\" rel=\"noopener noreferrer\">https://ai.meta.com/blog/llama-4-multimodal-intelligence/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-23\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-23-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-23-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-23-4\" class=\"data-footnote-backref\">↩<sup>4</sup></a></p>\n</li>\n<li id=\"user-content-fn-24\">\n<p>Google AI for Developers. “Gemma 3 model overview.” <a href=\"https://ai.google.dev/gemma/docs/core\" rel=\"noopener noreferrer\">https://ai.google.dev/gemma/docs/core</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-24\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-24-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-25\">\n<p>Microsoft Community Hub. “Phi-4: Small Language Models That Pack a Punch.” <a href=\"https://techcommunity.microsoft.com/blog/educatordeveloperblog/phi-4-small-language-models-that-pack-a-punch/4464167\" rel=\"noopener noreferrer\">https://techcommunity.microsoft.com/blog/educatordeveloperblog/phi-4-small-language-models-that-pack-a-punch/4464167</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-25\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-26\">\n<p>Microsoft Community Hub. “Welcome to the new Phi-4 models - Microsoft Phi-4-mini &amp; Phi-4-multimodal.” <a href=\"https://techcommunity.microsoft.com/blog/educatordeveloperblog/welcome-to-the-new-phi-4-models---microsoft-phi-4-mini--phi-4-multimodal/4386037\" rel=\"noopener noreferrer\">https://techcommunity.microsoft.com/blog/educatordeveloperblog/welcome-to-the-new-phi-4-models---microsoft-phi-4-mini—phi-4-multimodal/4386037</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-26\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-26-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-27\">\n<p>Microsoft Research. “Phi-4-reasoning Technical Report.” <a href=\"https://www.microsoft.com/en-us/research/publication/phi-4-reasoning-technical-report/\" rel=\"noopener noreferrer\">https://www.microsoft.com/en-us/research/publication/phi-4-reasoning-technical-report/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-27\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-28\">\n<p>Microsoft Azure Blog. “Reasoning reimagined: Introducing Phi-4-mini-flash-reasoning.” <a href=\"https://azure.microsoft.com/en-us/blog/reasoning-reimagined-introducing-phi-4-mini-flash-reasoning/\" rel=\"noopener noreferrer\">https://azure.microsoft.com/en-us/blog/reasoning-reimagined-introducing-phi-4-mini-flash-reasoning/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-28\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-29\">\n<p>HuggingFace Blog. “Welcome Gemma 3: Google’s all new multimodal, multilingual, long context open LLM.” <a href=\"https://huggingface.co/blog/gemma3\" rel=\"noopener noreferrer\">https://huggingface.co/blog/gemma3</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-29\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-30\">\n<p>Google Developers Blog. “Introducing Gemma 3 270M: The compact model for hyper-efficient AI.” <a href=\"https://developers.googleblog.com/en/introducing-gemma-3-270m/\" rel=\"noopener noreferrer\">https://developers.googleblog.com/en/introducing-gemma-3-270m/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-30\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-31\">\n<p>Google Blog. “TranslateGemma: A new family of open translation models.” <a href=\"https://blog.google/technology/developers/translategemma/\" rel=\"noopener noreferrer\">https://blog.google/technology/developers/translategemma/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-31\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-32\">\n<p>Open Laboratory. “Qwen3 4B.” <a href=\"https://openlaboratory.ai/models/qwen3-4b\" rel=\"noopener noreferrer\">https://openlaboratory.ai/models/qwen3-4b</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-32\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-33\">\n<p>Interconnects. “Qwen 3: The new open standard.” <a href=\"https://www.interconnects.ai/p/qwen-3-the-new-open-standard\" rel=\"noopener noreferrer\">https://www.interconnects.ai/p/qwen-3-the-new-open-standard</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-33\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-34\">\n<p>Encord. “Llama 3.2: Advanced Vision and Edge AI Models for Mobile and Cloud.” <a href=\"https://encord.com/blog/lama-3-2-explained/\" rel=\"noopener noreferrer\">https://encord.com/blog/lama-3-2-explained/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-34\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-35\">\n<p>InvestGlass. “How to Run LLMs Locally: Complete 2025 Guide to Self-Hosted AI Models.” <a href=\"https://www.investglass.com/en_gb/how-to-run-llms-locally-complete-2025-guide-to-self-hosted-ai-models/\" rel=\"noopener noreferrer\">https://www.investglass.com/en_gb/how-to-run-llms-locally-complete-2025-guide-to-self-hosted-ai-models/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-35\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-36\">\n<p>SiliconFlow. “Ultimate Guide - The Best LLMs For Mobile Deployment In 2026.” <a href=\"https://www.siliconflow.com/articles/en/best-LLMs-for-mobile-deployment\" rel=\"noopener noreferrer\">https://www.siliconflow.com/articles/en/best-LLMs-for-mobile-deployment</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-36\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-36-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-37\">\n<p>SabrePC Blog. “Popular Small Language Models to Run Locally.” <a href=\"https://www.sabrepc.com/blog/deep-learning-and-ai/popular-small-language-models-to-run-locally\" rel=\"noopener noreferrer\">https://www.sabrepc.com/blog/deep-learning-and-ai/popular-small-language-models-to-run-locally</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-37\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-38\">\n<p>SiliconFlow. “Ultimate Guide - The Best LLMs for Edge AI Devices in 2026.” <a href=\"https://www.siliconflow.com/articles/en/best-llms-for-edge-ai-devices-2025\" rel=\"noopener noreferrer\">https://www.siliconflow.com/articles/en/best-llms-for-edge-ai-devices-2025</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-38\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-39\">\n<p>Microsoft Research. “OptiMind: A small language model with optimization expertise.” <a href=\"https://www.microsoft.com/en-us/research/blog/optimind-a-small-language-model-with-optimization-expertise/\" rel=\"noopener noreferrer\">https://www.microsoft.com/en-us/research/blog/optimind-a-small-language-model-with-optimization-expertise/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-39\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-40\">\n<p>NVIDIA Research. “Small Language Models are the Future of Agentic AI.” <a href=\"https://research.nvidia.com/labs/lpr/slm-agents/\" rel=\"noopener noreferrer\">https://research.nvidia.com/labs/lpr/slm-agents/</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-40\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-40-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-41\">\n<p>DataCamp. “LLM Distillation Explained: Applications, Implementation &amp; More.” <a href=\"https://www.datacamp.com/blog/distillation-llm\" rel=\"noopener noreferrer\">https://www.datacamp.com/blog/distillation-llm</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-41\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-42\">\n<p>Databricks. “Efficient Fine-Tuning with LoRA: A Guide to Optimal Parameter Selection for Large Language Models.” <a href=\"https://www.databricks.com/blog/efficient-fine-tuning-lora-guide-llms\" rel=\"noopener noreferrer\">https://www.databricks.com/blog/efficient-fine-tuning-lora-guide-llms</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-42\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-43\">\n<p>HuggingFace Blog. “An Edge-First Generalized LLM LoRA Fine-Tuning Framework for Heterogeneous GPUs.” <a href=\"https://huggingface.co/blog/qvac/fabric-llm-finetune\" rel=\"noopener noreferrer\">https://huggingface.co/blog/qvac/fabric-llm-finetune</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-43\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-44\">\n<p>Stanford HAI. “2025 AI Index Report: Technical Performance.” <a href=\"https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance\" rel=\"noopener noreferrer\">https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance</a> <a href=\"https://blog.ecitis.org/small-language-models/#user-content-fnref-44\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/small-language-models/",
            "title": "Small Language Models: When Less is More",
            "summary": "Understand where compact models outperform frontier-scale systems on privacy, latency, cost, and practical deployment.",
            "image": "https://blog.ecitis.org/open-graph/small-language-models.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "SLM",
                "Efficiency",
                "On-device"
            ]
        },
        {
            "id": "https://blog.ecitis.org/tensorrt-triton-inference/",
            "content_html": "<p>Deploying LLMs at scale requires more than loading a model and serving requests. Production systems must handle concurrent users, minimize latency, maximize GPU utilization, and provide observability. NVIDIA’s TensorRT-LLM and Triton Inference Server form a battle-tested stack for these requirements.</p>\n<p>TensorRT-LLM became fully open-source in March 2025 and now delivers over 40,000 tokens per second on Blackwell B200 GPUs running Llama 4<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new1\" id=\"user-content-fnref-new1\">1</a></sup>. The stack integrates with NVIDIA Dynamo for datacenter-scale orchestration across thousands of GPUs<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new2\" id=\"user-content-fnref-new2\">2</a></sup>.</p>\n<p>This guide covers the complete deployment pipeline: building optimized TensorRT engines, configuring Triton for LLM workloads, and tuning for production performance.</p>\n<hr />\n<h2 id=\"tensorrt-llm-overview\">TensorRT-LLM overview</h2>\n<p>TensorRT-LLM is NVIDIA’s open-source library for compiling and running large language models on NVIDIA GPUs. Built on PyTorch, it provides a Python API for model definition while generating highly optimized CUDA kernels for inference<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-1\" id=\"user-content-fnref-1\">3</a></sup>. The current release (January 2026) uses PyTorch 2.9.0, TensorRT 10.9, and CUDA 12.8.1<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new3\" id=\"user-content-fnref-new3\">4</a></sup>.</p>\n<p>The library handles several optimization categories:</p>\n<p><strong>Quantization</strong>: TensorRT-LLM supports NVFP4, FP8, INT4 AWQ, and INT8 SmoothQuant quantization formats. NVFP4 is a 4-bit floating-point format introduced with Blackwell GPUs that reduces KV cache memory by 50% compared to FP8 while maintaining less than 1% accuracy loss on benchmarks including LiveCodeBench, MMLU-PRO, and MBPP<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new4\" id=\"user-content-fnref-new4\">5</a></sup>. FP8 on Hopper and Blackwell architectures delivers 2.5-3x inference speed improvements compared to FP16<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-2\" id=\"user-content-fnref-2\">6</a></sup>.</p>\n<p><strong>Kernel fusion</strong>: Multiple transformer operations combine into single CUDA kernels. LayerNorm, matrix multiplications, bias additions, and activation functions execute together instead of requiring separate kernel launches and memory transfers<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-3\" id=\"user-content-fnref-3\">7</a></sup>.</p>\n<p><strong>Attention optimizations</strong>: Custom FlashAttention kernels, multi-head/multi-query/grouped-query attention support, and fused attention implementations reduce memory bandwidth requirements.</p>\n<p><strong>Parallelism</strong>: Tensor parallelism splits matrix multiplications across GPUs. Pipeline parallelism distributes model layers sequentially across devices. Wide expert parallelism (EP) enables efficient Mixture-of-Experts inference on models like DeepSeek-R1 and Llama 4<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new5\" id=\"user-content-fnref-new5\">8</a></sup>. All strategies enable serving models larger than single-GPU memory<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-4\" id=\"user-content-fnref-4\">9</a></sup>.</p>\n<p><strong>Speculative decoding</strong>: TensorRT-LLM supports EAGLE-3, multi-token prediction (MTP), and ReDrafter techniques for accelerated token generation. EAGLE-3 with speculative decoding achieves up to 4x speedup on Llama 4 Maverick<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new6\" id=\"user-content-fnref-new6\">10</a></sup>. ReDrafter, developed by Apple and integrated into TensorRT-LLM, achieves up to 2.7x throughput improvements on H100 GPUs<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new7\" id=\"user-content-fnref-new7\">11</a></sup>.</p>\n<figure><figcaption><strong>Triton + TensorRT-LLM Architecture</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/tensorrt-triton-inference-0.webp\" width=\"512\" height=\"1109\" alt=\"Triton + TensorRT-LLM Architecture\" loading=\"lazy\" decoding=\"async\" /></figure>\n<hr />\n<h2 id=\"building-tensorrt-engines-for-llms\">Building TensorRT engines for LLMs</h2>\n<p>TensorRT-LLM converts model weights into optimized TensorRT engines through a two-step process: checkpoint conversion and engine building.</p>\n<h3 id=\"step-1-quantize-and-convert-checkpoints\">Step 1: Quantize and convert checkpoints</h3>\n<p>The <code>quantize.py</code> script converts Hugging Face checkpoints to TensorRT-LLM format while applying quantization:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># NVFP4 quantization (Blackwell GPUs - recommended for B200/GB200)</span></span>\n<span class=\"line\"><span>python</span><span> examples/quantization/quantize.py</span><span> \\</span></span>\n<span class=\"line\"><span>  --model_dir</span><span> /path/to/llama-70b</span><span> \\</span></span>\n<span class=\"line\"><span>  --qformat</span><span> nvfp4</span><span> \\</span></span>\n<span class=\"line\"><span>  --kv_cache_dtype</span><span> nvfp4</span><span> \\</span></span>\n<span class=\"line\"><span>  --output_dir</span><span> /output/llama-70b-nvfp4</span><span> \\</span></span>\n<span class=\"line\"><span>  --tp_size</span><span> 4</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># FP8 quantization (Hopper/Ada/Blackwell GPUs)</span></span>\n<span class=\"line\"><span>python</span><span> examples/quantization/quantize.py</span><span> \\</span></span>\n<span class=\"line\"><span>  --model_dir</span><span> /path/to/llama-70b</span><span> \\</span></span>\n<span class=\"line\"><span>  --qformat</span><span> fp8</span><span> \\</span></span>\n<span class=\"line\"><span>  --kv_cache_dtype</span><span> fp8</span><span> \\</span></span>\n<span class=\"line\"><span>  --output_dir</span><span> /output/llama-70b-fp8</span><span> \\</span></span>\n<span class=\"line\"><span>  --tp_size</span><span> 4</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># INT4 AWQ quantization (all GPU generations)</span></span>\n<span class=\"line\"><span>python</span><span> examples/quantization/quantize.py</span><span> \\</span></span>\n<span class=\"line\"><span>  --model_dir</span><span> /path/to/llama-70b</span><span> \\</span></span>\n<span class=\"line\"><span>  --qformat</span><span> int4_awq</span><span> \\</span></span>\n<span class=\"line\"><span>  --awq_block_size</span><span> 64</span><span> \\</span></span>\n<span class=\"line\"><span>  --output_dir</span><span> /output/llama-70b-int4awq</span><span> \\</span></span>\n<span class=\"line\"><span>  --tp_size</span><span> 4</span></span></code></pre>\n<p>The <code>--tp_size</code> parameter specifies tensor parallelism degree. Set this to the number of GPUs you will use for inference. INT4 AWQ uses block-wise quantization; smaller block sizes (64 vs 128) provide better accuracy at marginal compute cost<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-5\" id=\"user-content-fnref-5\">12</a></sup>. NVFP4 uses block-wise quantization with size 16 and FP8 scaling factors for higher precision during dequantization<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new4\" id=\"user-content-fnref-new4-2\">5</a></sup>.</p>\n<h3 id=\"step-2-build-the-engine\">Step 2: Build the engine</h3>\n<p>The <code>trtllm-build</code> command compiles checkpoints into TensorRT engines:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>trtllm-build</span><span> \\</span></span>\n<span class=\"line\"><span>  --checkpoint_dir</span><span> /output/llama-70b-fp8</span><span> \\</span></span>\n<span class=\"line\"><span>  --output_dir</span><span> /engines/llama-70b-fp8</span><span> \\</span></span>\n<span class=\"line\"><span>  --gemm_plugin</span><span> fp8</span><span> \\</span></span>\n<span class=\"line\"><span>  --gpt_attention_plugin</span><span> fp8</span><span> \\</span></span>\n<span class=\"line\"><span>  --max_batch_size</span><span> 64</span><span> \\</span></span>\n<span class=\"line\"><span>  --max_input_len</span><span> 4096</span><span> \\</span></span>\n<span class=\"line\"><span>  --max_seq_len</span><span> 8192</span><span> \\</span></span>\n<span class=\"line\"><span>  --use_fused_mlp</span><span> enable</span><span> \\</span></span>\n<span class=\"line\"><span>  --workers</span><span> 4</span></span></code></pre>\n<p>Key build parameters:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Parameter</th><th>Purpose</th></tr></thead><tbody><tr><td><code>--gemm_plugin</code></td><td>Enables cuBLASLt for optimized matrix operations</td></tr><tr><td><code>--gpt_attention_plugin</code></td><td>Uses efficient attention kernels with in-place KV cache updates</td></tr><tr><td><code>--use_fused_mlp</code></td><td>Enables horizontal fusion in GatedMLP layers</td></tr><tr><td><code>--max_batch_size</code></td><td>Maximum concurrent sequences the engine supports</td></tr><tr><td><code>--max_input_len</code></td><td>Maximum input context length</td></tr><tr><td><code>--max_seq_len</code></td><td>Maximum total sequence length (input + output)</td></tr></tbody></table></div>\n<p>For FP8 on Hopper GPUs, the GEMM + SwiGLU fusion in Gated-MLP combines two Matmul operations and one SwiGLU operation into a single kernel<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-6\" id=\"user-content-fnref-6\">13</a></sup>.</p>\n<h3 id=\"layer-fusion-patterns\">Layer fusion patterns</h3>\n<p>TensorRT’s compiler identifies and fuses operation sequences automatically. Common fusion patterns include:</p>\n<ul>\n<li><strong>RMSNorm fusion</strong>: Combines normalization with subsequent quantization</li>\n<li><strong>Attention fusion</strong>: Merges Q/K/V projections with attention computation</li>\n<li><strong>MLP fusion</strong>: Fuses gate and up projections with activation functions</li>\n<li><strong>AllReduce fusion</strong>: Combines reduction with LayerNorm after multi-GPU communication</li>\n</ul>\n<p>The <code>--reduce_fusion enable</code> flag eliminates extra copies from local buffers to shared buffers in communication kernels<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-7\" id=\"user-content-fnref-7\">14</a></sup>.</p>\n<hr />\n<h2 id=\"disaggregated-serving\">Disaggregated serving</h2>\n<p>Disaggregated serving separates compute-intensive prefill operations from memory-bound decode operations onto specialized hardware clusters. This architecture addresses the fundamental mismatch between prefill (compute-bound, benefits from high FLOPS) and decode (memory-bound, benefits from high memory bandwidth) phases<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new8\" id=\"user-content-fnref-new8\">15</a></sup>.</p>\n<p>TensorRT-LLM supports three disaggregated serving approaches:</p>\n<p><strong>trtllm-serve</strong>: A command-line utility that deploys OpenAI-compatible servers for each context and generation instance, with an orchestrator coordinating requests.</p>\n<p><strong>Dynamo integration</strong>: NVIDIA Dynamo orchestrates requests across prefill and decode workers with KV-cache-aware routing. The smart router determines optimal decode workers based on KV cache block availability<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new2\" id=\"user-content-fnref-new2-2\">2</a></sup>.</p>\n<p><strong>NIXL acceleration</strong>: The NVIDIA Inference Exchange Library (NIXL) accelerates KV cache transfer between GPUs with low-latency communication primitives.</p>\n<p>Disaggregated architectures demonstrate up to 6.4x throughput improvements and 20x reduction in latency variance. Organizations report 15-40% infrastructure cost reductions through optimized hardware allocation<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new9\" id=\"user-content-fnref-new9\">16</a></sup>.</p>\n<hr />\n<h2 id=\"in-flight-batching-and-paged-attention\">In-flight batching and paged attention</h2>\n<p>Traditional batching waits until a batch fills before processing. In-flight batching (also called continuous batching or iteration-level batching) processes new requests immediately without waiting for existing sequences to complete<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-8\" id=\"user-content-fnref-8\">17</a></sup>.</p>\n<h3 id=\"how-in-flight-batching-works\">How in-flight batching works</h3>\n<p>The TensorRT-LLM scheduler manages two phases:</p>\n<ol>\n<li><strong>Context phase</strong>: Processing input prompts, computing initial KV cache entries</li>\n<li><strong>Generation phase</strong>: Autoregressive token generation using cached keys/values</li>\n</ol>\n<p>With in-flight batching, sequences in context phase process together with sequences in generation phase. When a generation sequence finishes, a new context sequence can immediately take its slot. This keeps GPU utilization high even with variable-length inputs and outputs.</p>\n<h3 id=\"paged-kv-cache\">Paged KV cache</h3>\n<p>Instead of allocating KV cache as contiguous memory blocks, paged attention splits the cache into smaller blocks that can be allocated non-contiguously<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-9\" id=\"user-content-fnref-9\">18</a></sup>. This approach:</p>\n<ul>\n<li>Eliminates memory fragmentation from variable sequence lengths</li>\n<li>Enables KV cache sharing between requests with common prefixes</li>\n<li>Supports block sizes of 8, 16, 32, 64, or 128 tokens per block</li>\n</ul>\n<p>The paged KV cache allocates memory upfront in TensorRT-LLM rather than on-demand. Plan for approximately 60% additional VRAM beyond model weights for KV cache allocation<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-10\" id=\"user-content-fnref-10\">19</a></sup>.</p>\n<h3 id=\"chunked-prefill\">Chunked prefill</h3>\n<p>Long input contexts can delay generation for concurrent requests. Chunked prefill divides the context phase into smaller chunks, enabling better parallelization with decode operations<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-11\" id=\"user-content-fnref-11\">20</a></sup>.</p>\n<p>This feature allows GPU systems to handle longer contexts and higher concurrency by decoupling memory consumption from context length.</p>\n<hr />\n<h2 id=\"triton-inference-server-architecture\">Triton Inference Server architecture</h2>\n<p>Triton provides the serving infrastructure around TensorRT-LLM engines. It handles request routing, model loading, batching, and metrics collection. The OpenAI-compatible frontend transitioned from beta to stable in 2025, enabling drop-in compatibility with OpenAI API clients<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new10\" id=\"user-content-fnref-new10\">21</a></sup>.</p>\n<h3 id=\"core-components\">Core components</h3>\n<p><strong>Model Repository</strong>: A directory structure containing model configurations and artifacts. Each model has a <code>config.pbtxt</code> file defining inputs, outputs, and backend settings.</p>\n<p><strong>Scheduler</strong>: Routes incoming requests to model instances. For LLMs, the scheduler typically delegates batching decisions to the TensorRT-LLM runtime.</p>\n<p><strong>Backend Manager</strong>: Loads and manages model backends. The TensorRT-LLM backend uses the C++ Executor API for optimal performance.</p>\n<p><strong>Metrics Server</strong>: Exposes Prometheus-compatible metrics on port 8002 by default.</p>\n<p><strong>Security note</strong>: Organizations should update to Triton version 25.07 or later to address CVE-2025-23319, CVE-2025-23320, and CVE-2025-23334, which when chained could allow unauthenticated remote code execution<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new11\" id=\"user-content-fnref-new11\">22</a></sup>.</p>\n<h3 id=\"model-ensemble-structure\">Model ensemble structure</h3>\n<p>The TensorRT-LLM backend uses an ensemble of three models:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>model_repository/</span></span>\n<span class=\"line\"><span>├── ensemble/</span></span>\n<span class=\"line\"><span>│   └── config.pbtxt</span></span>\n<span class=\"line\"><span>├── preprocessing/</span></span>\n<span class=\"line\"><span>│   ├── config.pbtxt</span></span>\n<span class=\"line\"><span>│   └── 1/model.py</span></span>\n<span class=\"line\"><span>├── tensorrt_llm/</span></span>\n<span class=\"line\"><span>│   ├── config.pbtxt</span></span>\n<span class=\"line\"><span>│   └── 1/</span></span>\n<span class=\"line\"><span>│       └── [engine files]</span></span>\n<span class=\"line\"><span>└── postprocessing/</span></span>\n<span class=\"line\"><span>    ├── config.pbtxt</span></span>\n<span class=\"line\"><span>    └── 1/model.py</span></span></code></pre>\n<p><strong>Preprocessing</strong>: Tokenizes input text using the model’s tokenizer\n<strong>TensorRT-LLM</strong>: Runs inference on the optimized engine\n<strong>Postprocessing</strong>: Detokenizes output IDs back to text</p>\n<p>The ensemble model chains these together automatically.</p>\n<hr />\n<h2 id=\"configuring-triton-for-llms\">Configuring Triton for LLMs</h2>\n<h3 id=\"tensorrtllm-model-configuration\">tensorrt_llm model configuration</h3>\n<p>The core configuration for the TensorRT-LLM backend:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>name: \"tensorrt_llm\"</span></span>\n<span class=\"line\"><span>backend: \"tensorrtllm\"</span></span>\n<span class=\"line\"><span>max_batch_size: 64</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model_transaction_policy {</span></span>\n<span class=\"line\"><span>  decoupled: true</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>input [</span></span>\n<span class=\"line\"><span>  { name: \"input_ids\" data_type: TYPE_INT32 dims: [-1] },</span></span>\n<span class=\"line\"><span>  { name: \"input_lengths\" data_type: TYPE_INT32 dims: [1] },</span></span>\n<span class=\"line\"><span>  { name: \"request_output_len\" data_type: TYPE_INT32 dims: [1] },</span></span>\n<span class=\"line\"><span>  { name: \"end_id\" data_type: TYPE_INT32 dims: [1] },</span></span>\n<span class=\"line\"><span>  { name: \"pad_id\" data_type: TYPE_INT32 dims: [1] },</span></span>\n<span class=\"line\"><span>  { name: \"stream\" data_type: TYPE_BOOL dims: [1] optional: true }</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>output [</span></span>\n<span class=\"line\"><span>  { name: \"output_ids\" data_type: TYPE_INT32 dims: [-1] },</span></span>\n<span class=\"line\"><span>  { name: \"sequence_length\" data_type: TYPE_INT32 dims: [1] }</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>parameters {</span></span>\n<span class=\"line\"><span>  key: \"gpt_model_path\"</span></span>\n<span class=\"line\"><span>  value: { string_value: \"/engines/llama-70b-fp8\" }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>parameters {</span></span>\n<span class=\"line\"><span>  key: \"batching_type\"</span></span>\n<span class=\"line\"><span>  value: { string_value: \"inflight_fused_batching\" }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>parameters {</span></span>\n<span class=\"line\"><span>  key: \"max_tokens_in_paged_kv_cache\"</span></span>\n<span class=\"line\"><span>  value: { string_value: \"2048000\" }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>parameters {</span></span>\n<span class=\"line\"><span>  key: \"kv_cache_free_gpu_mem_fraction\"</span></span>\n<span class=\"line\"><span>  value: { string_value: \"0.85\" }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Key parameters explained:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Parameter</th><th>Value</th><th>Purpose</th></tr></thead><tbody><tr><td><code>decoupled: true</code></td><td>Required for streaming</td><td>Enables async responses</td></tr><tr><td><code>batching_type</code></td><td><code>inflight_fused_batching</code></td><td>Enables continuous batching</td></tr><tr><td><code>max_tokens_in_paged_kv_cache</code></td><td>Token count</td><td>Total KV cache capacity</td></tr><tr><td><code>kv_cache_free_gpu_mem_fraction</code></td><td>0.0-1.0</td><td>Fraction of free GPU memory for KV cache</td></tr></tbody></table></div>\n<h3 id=\"instance-configuration\">Instance configuration</h3>\n<p>The <code>instance_count</code> parameter controls concurrent execution. For most workloads, set this to 5 or higher<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-12\" id=\"user-content-fnref-12\">23</a></sup>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>instance_group [</span></span>\n<span class=\"line\"><span>  {</span></span>\n<span class=\"line\"><span>    count: 5</span></span>\n<span class=\"line\"><span>    kind: KIND_GPU</span></span>\n<span class=\"line\"><span>    gpus: [0, 1, 2, 3]</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>]</span></span></code></pre>\n<h3 id=\"queue-management\">Queue management</h3>\n<p>Control request queuing behavior:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>parameters {</span></span>\n<span class=\"line\"><span>  key: \"max_queue_delay_microseconds\"</span></span>\n<span class=\"line\"><span>  value: { string_value: \"100000\" }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>parameters {</span></span>\n<span class=\"line\"><span>  key: \"max_queue_size\"</span></span>\n<span class=\"line\"><span>  value: { string_value: \"256\" }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>Setting <code>max_queue_delay_microseconds</code> above 0 improves batch formation by waiting for additional requests to arrive.</p>\n<hr />\n<h2 id=\"tensorrt-llm-backend-setup-walkthrough\">TensorRT-LLM backend setup walkthrough</h2>\n<h3 id=\"1-pull-the-container\">1. Pull the container</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>docker</span><span> pull</span><span> nvcr.io/nvidia/tritonserver:25.10-py3</span></span></code></pre>\n<h3 id=\"2-create-model-repository\">2. Create model repository</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Clone the backend repository</span></span>\n<span class=\"line\"><span>git</span><span> clone</span><span> https://github.com/triton-inference-server/tensorrtllm_backend.git</span></span>\n<span class=\"line\"><span>cd</span><span> tensorrtllm_backend</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Copy template models</span></span>\n<span class=\"line\"><span>mkdir</span><span> -p</span><span> /triton_models</span></span>\n<span class=\"line\"><span>cp</span><span> -r</span><span> all_models/inflight_batcher_llm/*</span><span> /triton_models/</span></span></code></pre>\n<h3 id=\"3-build-the-tensorrt-llm-engine\">3. Build the TensorRT-LLM engine</h3>\n<p>Follow the engine building steps from earlier, placing the output in <code>/engines/</code>.</p>\n<h3 id=\"4-configure-the-models\">4. Configure the models</h3>\n<p>Use the provided template filling script:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>ENGINE_DIR</span><span>=</span><span>/engines/llama-70b-fp8</span></span>\n<span class=\"line\"><span>TOKENIZER_DIR</span><span>=</span><span>/models/llama-70b</span></span>\n<span class=\"line\"><span>MAX_BATCH</span><span>=</span><span>64</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>python3</span><span> tools/fill_template.py</span><span> -i</span><span> /triton_models/preprocessing/config.pbtxt</span><span> \\</span></span>\n<span class=\"line\"><span>  tokenizer_dir:</span><span>${TOKENIZER_DIR}</span><span>,triton_max_batch_size:</span><span>${MAX_BATCH}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>python3</span><span> tools/fill_template.py</span><span> -i</span><span> /triton_models/tensorrt_llm/config.pbtxt</span><span> \\</span></span>\n<span class=\"line\"><span>  triton_backend:tensorrtllm,triton_max_batch_size:</span><span>${MAX_BATCH}</span><span>,</span><span>\\</span></span>\n<span class=\"line\"><span>  decoupled_mode:</span><span>true</span><span>,engine_dir:</span><span>${ENGINE_DIR}</span><span>,</span><span>\\</span></span>\n<span class=\"line\"><span>  batching_type:inflight_fused_batching</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>python3</span><span> tools/fill_template.py</span><span> -i</span><span> /triton_models/postprocessing/config.pbtxt</span><span> \\</span></span>\n<span class=\"line\"><span>  tokenizer_dir:</span><span>${TOKENIZER_DIR}</span><span>,triton_max_batch_size:</span><span>${MAX_BATCH}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>python3</span><span> tools/fill_template.py</span><span> -i</span><span> /triton_models/ensemble/config.pbtxt</span><span> \\</span></span>\n<span class=\"line\"><span>  triton_max_batch_size:</span><span>${MAX_BATCH}</span></span></code></pre>\n<h3 id=\"5-launch-the-server\">5. Launch the server</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>python3</span><span> scripts/launch_triton_server.py</span><span> \\</span></span>\n<span class=\"line\"><span>  --model_repo</span><span> /triton_models</span><span> \\</span></span>\n<span class=\"line\"><span>  --world_size</span><span> 4</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or directly with tritonserver</span></span>\n<span class=\"line\"><span>tritonserver</span><span> \\</span></span>\n<span class=\"line\"><span>  --model-repository=/triton_models</span><span> \\</span></span>\n<span class=\"line\"><span>  --http-port=8000</span><span> \\</span></span>\n<span class=\"line\"><span>  --grpc-port=8001</span><span> \\</span></span>\n<span class=\"line\"><span>  --metrics-port=8002</span></span></code></pre>\n<h3 id=\"6-verify-deployment\">6. Verify deployment</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>curl</span><span> localhost:8000/v2/health/ready</span></span>\n<span class=\"line\"><span># Response: {\"ready\":true}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>curl</span><span> localhost:8000/v2/models</span></span>\n<span class=\"line\"><span># Lists all loaded models with READY status</span></span></code></pre>\n<hr />\n<h2 id=\"performance-tuning\">Performance tuning</h2>\n<h3 id=\"concurrency-and-batch-size\">Concurrency and batch size</h3>\n<p>The relationship between <code>max_batch_size</code>, engine build parameters, and runtime behavior:</p>\n<ul>\n<li><strong>Engine <code>max_batch_size</code></strong>: Hard limit set at build time</li>\n<li><strong>Triton <code>max_batch_size</code></strong>: Soft limit for request acceptance</li>\n<li><strong>Runtime batch size</strong>: Determined by TRT-LLM scheduler based on available requests and memory</li>\n</ul>\n<p>The TRT-LLM scheduler can form batches larger than Triton’s <code>max_batch_size</code> when using in-flight batching. The scheduler optimizes based on available KV cache memory and pending requests<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-13\" id=\"user-content-fnref-13\">24</a></sup>.</p>\n<h3 id=\"gpu-memory-allocation\">GPU memory allocation</h3>\n<p>Balance KV cache size against model memory requirements:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Calculate KV cache memory</span></span>\n<span class=\"line\"><span>kv_cache_tokens </span><span>=</span><span> max_batch_size </span><span>*</span><span> max_seq_len</span></span>\n<span class=\"line\"><span>kv_cache_bytes </span><span>=</span><span> kv_cache_tokens </span><span>*</span><span> num_layers </span><span>*</span><span> 2</span><span> *</span><span> hidden_size </span><span>*</span><span> dtype_bytes</span></span></code></pre>\n<p>For a 70B model with 80 layers, 8192 hidden size, and FP16 KV cache:</p>\n<ul>\n<li>64 batch * 8192 seq_len = 524,288 tokens</li>\n<li>524,288 _ 80 _ 2 _ 8192 _ 2 bytes = ~1.3 TB (exceeds GPU memory)</li>\n</ul>\n<p>Use <code>kv_cache_free_gpu_mem_fraction</code> to limit allocation to available memory. Start with 0.85 and adjust based on OOM behavior.</p>\n<h3 id=\"cuda-graphs\">CUDA Graphs</h3>\n<p>TensorRT-LLM uses CUDA Graphs to reduce kernel launch overhead. The runtime captures execution graphs for common batch sizes and replays them without CPU intervention.</p>\n<p>CUDA Graph padding handles mismatched batch sizes by padding to the nearest captured graph size. This trades minor compute waste for consistent low-latency execution<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-14\" id=\"user-content-fnref-14\">25</a></sup>.</p>\n<h3 id=\"multi-gpu-configuration\">Multi-GPU configuration</h3>\n<p>For tensor parallelism across GPUs:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># 4-way tensor parallelism</span></span>\n<span class=\"line\"><span>trtllm-build</span><span> --tp_size</span><span> 4</span><span> ...</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Launch with matching world size</span></span>\n<span class=\"line\"><span>python3</span><span> launch_triton_server.py</span><span> --world_size</span><span> 4</span></span></code></pre>\n<p>The backend uses MPI to coordinate execution across GPUs. Leader mode uses one process to manage inference; Orchestrator mode distributes coordination<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-15\" id=\"user-content-fnref-15\">26</a></sup>.</p>\n<hr />\n<h2 id=\"monitoring-and-metrics\">Monitoring and metrics</h2>\n<h3 id=\"prometheus-metrics-endpoint\">Prometheus metrics endpoint</h3>\n<p>Triton exposes metrics at <code>http://localhost:8002/metrics</code> by default. Enable with:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>tritonserver</span><span> --allow-metrics=true</span><span> --allow-gpu-metrics=true</span></span></code></pre>\n<h3 id=\"key-metrics-for-llm-workloads\">Key metrics for LLM workloads</h3>\n<p>The TensorRT-LLM backend exposes custom metrics:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Metric</th><th>Description</th></tr></thead><tbody><tr><td><code>nv_trt_llm_request_count</code></td><td>Total inference requests</td></tr><tr><td><code>nv_trt_llm_inflight_request_count</code></td><td>Currently processing requests</td></tr><tr><td><code>nv_trt_llm_kv_cache_block_usage</code></td><td>KV cache utilization</td></tr><tr><td><code>nv_trt_llm_generation_tokens_per_second</code></td><td>Output throughput</td></tr></tbody></table></div>\n<p>Standard Triton metrics include:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Metric</th><th>Description</th></tr></thead><tbody><tr><td><code>nv_inference_request_success</code></td><td>Successful request count</td></tr><tr><td><code>nv_inference_compute_output_duration_us</code></td><td>Inference latency</td></tr><tr><td><code>nv_gpu_utilization</code></td><td>GPU compute utilization</td></tr><tr><td><code>nv_gpu_memory_used_bytes</code></td><td>GPU memory consumption</td></tr></tbody></table></div>\n<h3 id=\"grafana-dashboard-setup\">Grafana dashboard setup</h3>\n<p>NVIDIA provides pre-built Grafana dashboards in the tutorials repository<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-16\" id=\"user-content-fnref-16\">27</a></sup>. Import the JSON configuration:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Download dashboard JSON</span></span>\n<span class=\"line\"><span>curl</span><span> -O</span><span> https://raw.githubusercontent.com/triton-inference-server/tutorials/main/</span><span>\\</span></span>\n<span class=\"line\"><span>Deployment/Kubernetes/TensorRT-LLM_Autoscaling_and_Load_Balancing/\\</span></span>\n<span class=\"line\"><span>grafana_inference-metrics_dashboard.json</span></span></code></pre>\n<p>Configure Prometheus as a data source in Grafana, then import the dashboard JSON.</p>\n<h3 id=\"genai-perf-benchmarking\">GenAI-Perf benchmarking</h3>\n<p>Measure throughput and latency with GenAI-Perf:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>genai-perf</span><span> \\</span></span>\n<span class=\"line\"><span>  --model</span><span> ensemble</span><span> \\</span></span>\n<span class=\"line\"><span>  --backend</span><span> tensorrtllm</span><span> \\</span></span>\n<span class=\"line\"><span>  --endpoint</span><span> localhost:8001</span><span> \\</span></span>\n<span class=\"line\"><span>  --streaming</span><span> \\</span></span>\n<span class=\"line\"><span>  --concurrency</span><span> 32</span><span> \\</span></span>\n<span class=\"line\"><span>  --input-sequence-length</span><span> 512</span><span> \\</span></span>\n<span class=\"line\"><span>  --output-sequence-length</span><span> 128</span></span></code></pre>\n<p>The tool reports time-to-first-token, inter-token latency, throughput, and percentile statistics<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-17\" id=\"user-content-fnref-17\">28</a></sup>.</p>\n<hr />\n<h2 id=\"comparison-with-alternatives\">Comparison with alternatives</h2>\n<figure><figcaption><strong>LLM Inference Engine Comparison</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/tensorrt-triton-inference-1.webp\" width=\"512\" height=\"442\" alt=\"LLM Inference Engine Comparison\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"tensorrt-llm\">TensorRT-LLM</h3>\n<p>Best for maximum throughput when you can invest in setup complexity. January 2026 benchmarks show B200 delivering 60,000 tokens per second per GPU at 1,000 tokens per second per user interactivity on Llama 3.3 70B<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new12\" id=\"user-content-fnref-new12\">29</a></sup>. On Llama 4 Scout, B200 achieves over 42,000 tokens per second with a 3.4x performance increase over H200<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new1\" id=\"user-content-fnref-new1-2\">1</a></sup>.</p>\n<p><strong>Strengths</strong>: Peak throughput, Tensor Core optimization, deep NVIDIA integration, EAGLE-3 speculative decoding, NVFP4 quantization\n<strong>Trade-offs</strong>: Longer setup time (1-2 weeks for production tuning), engine rebuilds required for config changes</p>\n<h3 id=\"vllm\">vLLM</h3>\n<p>Best for teams prioritizing development velocity and concurrency handling. PagedAttention provides GPU-friendly memory layout without fragmentation<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-19\" id=\"user-content-fnref-19\">30</a></sup>. vLLM v0.14.1 (January 2026) adds W4A8 grouped GEMM on Hopper, MoE + LoRA support with AWQ Marlin, and Transformers v5 RoPE compatibility<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new13\" id=\"user-content-fnref-new13\">31</a></sup>.</p>\n<p>Inferact Inc. launched in January 2026 to commercialize vLLM with <math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><mn>150</mn><mi>M</mi><mi>f</mi><mi>u</mi><mi>n</mi><mi>d</mi><mi>i</mi><mi>n</mi><mi>g</mi><mi>a</mi><mi>t</mi><mi>a</mi><mi>n</mi></mrow></semantics></math>800M valuation<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new14\" id=\"user-content-fnref-new14\">32</a></sup>. The vLLM project now runs on over 400,000 GPUs worldwide.</p>\n<p><strong>Strengths</strong>: Easy setup (1-2 days), consistent latency under load, active community, broad hardware support (NVIDIA, AMD, Intel, TPU)\n<strong>Trade-offs</strong>: Lower peak throughput than TensorRT-LLM on Blackwell GPUs</p>\n<h3 id=\"sglang\">SGLang</h3>\n<p>SGLang has become a production standard, generating trillions of tokens daily across deployments at xAI, AMD, NVIDIA, LinkedIn, Cursor, and major cloud providers<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new15\" id=\"user-content-fnref-new15\">33</a></sup>. On H100 hardware, SGLang achieves 16,215 tokens per second, matching LMDeploy and exceeding vLLM by 29%<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new16\" id=\"user-content-fnref-new16\">34</a></sup>.</p>\n<p><strong>Strengths</strong>: RadixAttention for prefix caching (50%+ hit rates in production), zero-overhead CPU scheduler, prefill-decode disaggregation, native C++ optimization\n<strong>Trade-offs</strong>: Smaller ecosystem than vLLM</p>\n<h3 id=\"text-generation-inference-v3\">Text-Generation-Inference v3</h3>\n<p>TGI is now in maintenance mode, with Hugging Face recommending vLLM or SGLang for new deployments<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new17\" id=\"user-content-fnref-new17\">35</a></sup>. TGI v3 achieves 13x speedups on prompts exceeding 200,000 tokens by reusing earlier token computations<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-20\" id=\"user-content-fnref-20\">36</a></sup>.</p>\n<p><strong>Strengths</strong>: Prefix caching, long context handling, Hugging Face ecosystem, zero-config mode\n<strong>Trade-offs</strong>: Maintenance mode limits new feature development</p>\n<h3 id=\"decision-matrix\">Decision matrix</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Requirement</th><th>Recommended Engine</th></tr></thead><tbody><tr><td>Maximum throughput on Blackwell</td><td>TensorRT-LLM</td></tr><tr><td>Fast iteration cycles</td><td>vLLM or SGLang</td></tr><tr><td>High concurrency, consistent latency</td><td>vLLM</td></tr><tr><td>Prefix caching in production</td><td>SGLang</td></tr><tr><td>Long chat histories (200K+ tokens)</td><td>TGI v3</td></tr><tr><td>Kubernetes-native deployment</td><td>Triton + TensorRT-LLM</td></tr><tr><td>Datacenter-scale orchestration</td><td>NVIDIA Dynamo</td></tr><tr><td>Multi-vendor hardware support</td><td>vLLM</td></tr></tbody></table></div>\n<hr />\n<h2 id=\"complete-deployment-example\">Complete deployment example</h2>\n<p>End-to-end deployment of Llama 3.1 70B on 4x H100 80GB:</p>\n<h3 id=\"environment-setup\">Environment setup</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Pull containers</span></span>\n<span class=\"line\"><span>docker</span><span> pull</span><span> nvcr.io/nvidia/tritonserver:25.10-py3</span></span>\n<span class=\"line\"><span>docker</span><span> pull</span><span> nvcr.io/nvidia/pytorch:25.10-py3</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Download model</span></span>\n<span class=\"line\"><span>huggingface-cli</span><span> download</span><span> meta-llama/Llama-3.1-70B-Instruct</span><span> \\</span></span>\n<span class=\"line\"><span>  --local-dir</span><span> /models/llama-3.1-70b</span></span></code></pre>\n<h3 id=\"build-engine\">Build engine</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>docker</span><span> run</span><span> --gpus</span><span> all</span><span> -v</span><span> /models:/models</span><span> -v</span><span> /engines:/engines</span><span> \\</span></span>\n<span class=\"line\"><span>  nvcr.io/nvidia/pytorch:25.10-py3</span><span> bash</span><span> -c</span><span> \"</span></span>\n<span class=\"line\"><span>    pip install tensorrt-llm</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Quantize to FP8</span></span>\n<span class=\"line\"><span>    python -m tensorrt_llm.commands.quantize \\</span></span>\n<span class=\"line\"><span>      --model_dir /models/llama-3.1-70b \\</span></span>\n<span class=\"line\"><span>      --qformat fp8 \\</span></span>\n<span class=\"line\"><span>      --kv_cache_dtype fp8 \\</span></span>\n<span class=\"line\"><span>      --output_dir /engines/llama-70b-ckpt \\</span></span>\n<span class=\"line\"><span>      --tp_size 4</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Build engine</span></span>\n<span class=\"line\"><span>    trtllm-build \\</span></span>\n<span class=\"line\"><span>      --checkpoint_dir /engines/llama-70b-ckpt \\</span></span>\n<span class=\"line\"><span>      --output_dir /engines/llama-70b-engine \\</span></span>\n<span class=\"line\"><span>      --gemm_plugin fp8 \\</span></span>\n<span class=\"line\"><span>      --gpt_attention_plugin fp8 \\</span></span>\n<span class=\"line\"><span>      --max_batch_size 64 \\</span></span>\n<span class=\"line\"><span>      --max_input_len 4096 \\</span></span>\n<span class=\"line\"><span>      --max_seq_len 8192 \\</span></span>\n<span class=\"line\"><span>      --use_fused_mlp enable \\</span></span>\n<span class=\"line\"><span>      --workers 4</span></span>\n<span class=\"line\"><span>\"</span></span></code></pre>\n<h3 id=\"deploy-with-triton\">Deploy with Triton</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>docker</span><span> run</span><span> --gpus</span><span> all</span><span> -p</span><span> 8000:8000</span><span> -p</span><span> 8001:8001</span><span> -p</span><span> 8002:8002</span><span> \\</span></span>\n<span class=\"line\"><span>  -v</span><span> /engines:/engines</span><span> \\</span></span>\n<span class=\"line\"><span>  -v</span><span> /models:/models</span><span> \\</span></span>\n<span class=\"line\"><span>  -v</span><span> /triton_models:/triton_models</span><span> \\</span></span>\n<span class=\"line\"><span>  nvcr.io/nvidia/tritonserver:25.10-py3</span><span> bash</span><span> -c</span><span> \"</span></span>\n<span class=\"line\"><span>    # Configure models (using fill_template.py as shown earlier)</span></span>\n<span class=\"line\"><span>    # ...</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    tritonserver \\</span></span>\n<span class=\"line\"><span>      --model-repository=/triton_models \\</span></span>\n<span class=\"line\"><span>      --allow-metrics=true \\</span></span>\n<span class=\"line\"><span>      --allow-gpu-metrics=true</span></span>\n<span class=\"line\"><span>\"</span></span></code></pre>\n<h3 id=\"test-inference\">Test inference</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> tritonclient</span><span>.</span><span>grpc </span><span>as</span><span> grpcclient</span></span>\n<span class=\"line\"><span>import</span><span> numpy </span><span>as</span><span> np</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>client </span><span>=</span><span> grpcclient</span><span>.</span><span>InferenceServerClient</span><span>(url</span><span>=</span><span>\"localhost:8001\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Prepare inputs</span></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> \"Explain the attention mechanism in transformers:\"</span></span>\n<span class=\"line\"><span>input_ids </span><span>=</span><span> tokenizer</span><span>.</span><span>encode</span><span>(prompt)</span><span>  # Use appropriate tokenizer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>    grpcclient</span><span>.</span><span>InferInput</span><span>(</span><span>\"text_input\"</span><span>, [</span><span>1</span><span>], </span><span>\"BYTES\"</span><span>),</span></span>\n<span class=\"line\"><span>    grpcclient</span><span>.</span><span>InferInput</span><span>(</span><span>\"max_tokens\"</span><span>, [</span><span>1</span><span>], </span><span>\"INT32\"</span><span>),</span></span>\n<span class=\"line\"><span>    grpcclient</span><span>.</span><span>InferInput</span><span>(</span><span>\"stream\"</span><span>, [</span><span>1</span><span>], </span><span>\"BOOL\"</span><span>),</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>inputs</span><span>[</span><span>0</span><span>].</span><span>set_data_from_numpy</span><span>(np.</span><span>array</span><span>([[prompt]], dtype</span><span>=</span><span>object</span><span>))</span></span>\n<span class=\"line\"><span>inputs</span><span>[</span><span>1</span><span>].</span><span>set_data_from_numpy</span><span>(np.</span><span>array</span><span>([[</span><span>256</span><span>]], dtype</span><span>=</span><span>np.int32))</span></span>\n<span class=\"line\"><span>inputs</span><span>[</span><span>2</span><span>].</span><span>set_data_from_numpy</span><span>(np.</span><span>array</span><span>([[</span><span>True</span><span>]], dtype</span><span>=</span><span>bool</span><span>))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Stream responses</span></span>\n<span class=\"line\"><span>for</span><span> response </span><span>in</span><span> client</span><span>.</span><span>infer</span><span>(</span><span>\"ensemble\"</span><span>, inputs, stream</span><span>=</span><span>True</span><span>):</span></span>\n<span class=\"line\"><span>    output </span><span>=</span><span> response</span><span>.</span><span>as_numpy</span><span>(</span><span>\"text_output\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(output[</span><span>0</span><span>].</span><span>decode</span><span>(), end</span><span>=</span><span>\"\"</span><span>, flush</span><span>=</span><span>True</span><span>)</span></span></code></pre>\n<hr />\n<h2 id=\"blackwell-gpu-deployment\">Blackwell GPU deployment</h2>\n<p>NVIDIA Blackwell architecture (B200, GB200, GB300) delivers substantial inference improvements over Hopper:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Metric</th><th>B200 vs H200</th></tr></thead><tbody><tr><td>Tokens per second (Llama 3.3 70B)</td><td>4x higher throughput<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new12\" id=\"user-content-fnref-new12-2\">29</a></sup></td></tr><tr><td>Tokens per second (Llama 4 Scout)</td><td>3.4x improvement<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new1\" id=\"user-content-fnref-new1-3\">1</a></sup></td></tr><tr><td>DeepSeek-R1 inference</td><td>5x vs Hopper per GPU<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new18\" id=\"user-content-fnref-new18\">37</a></sup></td></tr><tr><td>Memory per GPU</td><td>Up to 288GB HBM3e (Blackwell Ultra)</td></tr></tbody></table></div>\n<p>The GB200 NVL72 rack-scale platform connects 72 Blackwell GPUs using fifth-generation NVLink with 1,800 GB/s bidirectional bandwidth between all chips<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new19\" id=\"user-content-fnref-new19\">38</a></sup>. Blackwell Ultra (GB300) delivers 45% higher DeepSeek-R1 throughput than GB200 with 1.5x more NVFP4 compute and 2x more attention-layer acceleration<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new18\" id=\"user-content-fnref-new18-2\">37</a></sup>.</p>\n<p>Key Blackwell optimizations in TensorRT-LLM:</p>\n<ul>\n<li>NVFP4 quantization for weights and KV cache</li>\n<li>Programmatic dependent launch (PDL) for reduced kernel latencies</li>\n<li>Enhanced all-to-all communication primitives</li>\n<li>CuteDSL NVFP4 grouped GEMM integration</li>\n</ul>\n<p>The January 2026 TensorRT-LLM optimizations delivered up to 2.8x throughput increases per Blackwell GPU over the previous three months<sup><a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fn-new20\" id=\"user-content-fnref-new20\">39</a></sup>.</p>\n<hr />\n<h2 id=\"summary\">Summary</h2>\n<p>TensorRT-LLM and Triton Inference Server provide production-grade LLM serving with:</p>\n<ul>\n<li><strong>Engine optimization</strong>: NVFP4/FP8 quantization, kernel fusion, custom attention kernels, and EAGLE-3 speculative decoding</li>\n<li><strong>Efficient batching</strong>: In-flight batching and paged KV cache for high throughput</li>\n<li><strong>Disaggregated serving</strong>: Separate prefill and decode phases for optimized resource utilization</li>\n<li><strong>Flexible deployment</strong>: Multi-GPU, multi-node support with MPI coordination and NVIDIA Dynamo orchestration</li>\n<li><strong>Observability</strong>: Prometheus metrics integration for monitoring</li>\n</ul>\n<p>The setup complexity is higher than alternatives like vLLM or SGLang, but peak performance on Blackwell GPUs is also higher. For teams already in the NVIDIA ecosystem deploying at scale, this stack delivers consistent, optimized inference. Organizations prioritizing faster iteration should evaluate vLLM (broad hardware support, active development) or SGLang (production-proven prefix caching).</p>\n<hr />\n<h2 id=\"references\">References</h2>\n<h3 id=\"2026-updates\">2026 updates</h3>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-new1\">\n<p><a href=\"https://developer.nvidia.com/blog/nvidia-accelerates-inference-on-meta-llama-4-scout-and-maverick/\" rel=\"noopener noreferrer\">NVIDIA Accelerates Inference on Meta Llama 4 Scout and Maverick</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new1\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new1-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new1-3\" class=\"data-footnote-backref\">↩<sup>3</sup></a></p>\n</li>\n<li id=\"user-content-fn-new2\">\n<p><a href=\"https://developer.nvidia.com/blog/introducing-nvidia-dynamo-a-low-latency-distributed-inference-framework-for-scaling-reasoning-ai-models/\" rel=\"noopener noreferrer\">NVIDIA Dynamo Open-Source Inference Framework</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new2\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new2-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-1\">\n<p><a href=\"https://nvidia.github.io/TensorRT-LLM/\" rel=\"noopener noreferrer\">TensorRT-LLM Documentation</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new3\">\n<p><a href=\"https://nvidia.github.io/TensorRT-LLM/release-notes.html\" rel=\"noopener noreferrer\">TensorRT-LLM Release Notes</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new4\">\n<p><a href=\"https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/\" rel=\"noopener noreferrer\">Introducing NVFP4 for Efficient Low-Precision Inference</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new4\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new4-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p><a href=\"https://nvidia.github.io/TensorRT-LLM/overview.html\" rel=\"noopener noreferrer\">TensorRT-LLM Overview</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p><a href=\"https://introl.com/blog/tensorrt-llm-optimization-nvidia-inference-stack-guide\" rel=\"noopener noreferrer\">TensorRT-LLM Optimization Guide</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new5\">\n<p><a href=\"https://developer.nvidia.com/blog/delivering-massive-performance-leaps-for-mixture-of-experts-inference-on-nvidia-blackwell/\" rel=\"noopener noreferrer\">Delivering Massive Performance Leaps for MoE Inference on Blackwell</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p><a href=\"https://nvidia.github.io/TensorRT-LLM/architecture/overview.html\" rel=\"noopener noreferrer\">TensorRT-LLM Architecture Overview</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new6\">\n<p><a href=\"https://developer.nvidia.com/blog/blackwell-breaks-the-1000-tps-user-barrier-with-metas-llama-4-maverick/\" rel=\"noopener noreferrer\">Blackwell Breaks 1,000 TPS/User Barrier with Llama 4 Maverick</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new7\">\n<p><a href=\"https://developer.nvidia.com/blog/nvidia-tensorrt-llm-now-supports-recurrent-drafting-for-optimizing-llm-inference/\" rel=\"noopener noreferrer\">TensorRT-LLM Supports Recurrent Drafting for LLM Inference</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p><a href=\"https://github.com/NVIDIA/TensorRT-LLM/blob/main/examples/quantization/README.md\" rel=\"noopener noreferrer\">TensorRT-LLM Quantization Examples</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p><a href=\"https://nvidia.github.io/TensorRT-LLM/latest/commands/trtllm-build.html\" rel=\"noopener noreferrer\">trtllm-build Command Reference</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p><a href=\"https://nvidia.github.io/TensorRT-LLM/blogs/tech_blog/blog1_Pushing_Latency_Boundaries_Optimizing_DeepSeek-R1_Performance_on_NVIDIA_B200_GPUs.html\" rel=\"noopener noreferrer\">DeepSeek-R1 Optimization on B200</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new8\">\n<p><a href=\"https://nvidia.github.io/TensorRT-LLM/blogs/tech_blog/blog5_Disaggregated_Serving_in_TensorRT-LLM.html\" rel=\"noopener noreferrer\">Disaggregated Serving in TensorRT-LLM</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new8\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new9\">\n<p><a href=\"https://www.infoq.com/articles/llms-evolution-ai-infrastructure/\" rel=\"noopener noreferrer\">Disaggregation in LLMs: Evolution in AI Infrastructure</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new9\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p><a href=\"https://developer.nvidia.com/blog/nvidia-tensorrt-llm-now-accelerates-encoder-decoder-models-with-in-flight-batching/\" rel=\"noopener noreferrer\">TensorRT-LLM In-Flight Batching</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-9\">\n<p><a href=\"https://nvidia.github.io/TensorRT-LLM/advanced/gpt-attention.html\" rel=\"noopener noreferrer\">TensorRT-LLM GPT Attention</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-9\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-10\">\n<p><a href=\"https://softwaremill.com/boosting-llms-performance-in-production/\" rel=\"noopener noreferrer\">Boosting LLMs Performance in Production</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-10\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p><a href=\"https://developer.nvidia.com/blog/streamlining-ai-inference-performance-and-deployment-with-nvidia-tensorrt-llm-chunked-prefill/\" rel=\"noopener noreferrer\">TensorRT-LLM Chunked Prefill</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new10\">\n<p><a href=\"https://docs.nvidia.com/deeplearning/triton-inference-server/release-notes/index.html\" rel=\"noopener noreferrer\">Triton Inference Server Release Notes</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new10\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new11\">\n<p><a href=\"https://www.darkreading.com/vulnerabilities-threats/nvidia-patches-critical-rce-vulnerability-chain\" rel=\"noopener noreferrer\">NVIDIA Patches Critical RCE Vulnerability Chain</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new11\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-12\">\n<p><a href=\"https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/tensorrtllm_backend/docs/model_config.html\" rel=\"noopener noreferrer\">Triton TensorRT-LLM Model Configuration</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-12\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-13\">\n<p><a href=\"https://github.com/triton-inference-server/tensorrtllm_backend/blob/main/docs/model_config.md\" rel=\"noopener noreferrer\">TensorRT-LLM Backend Documentation</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-13\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p><a href=\"https://developer.nvidia.com/blog/nvidia-tensorrt-llm-supercharges-large-language-model-inference-on-nvidia-h100-gpus/\" rel=\"noopener noreferrer\">NVIDIA TensorRT-LLM H100 Performance</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p><a href=\"https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/tutorials/Deployment/Kubernetes/TensorRT-LLM_Multi-Node_Distributed_Models/README.html\" rel=\"noopener noreferrer\">Triton Multi-Node Deployment</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-16\">\n<p><a href=\"https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/tutorials/Deployment/Kubernetes/TensorRT-LLM_Autoscaling_and_Load_Balancing/README.html\" rel=\"noopener noreferrer\">Triton Autoscaling and Load Balancing Tutorial</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-16\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-17\">\n<p><a href=\"https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/metrics.html\" rel=\"noopener noreferrer\">Triton Metrics Documentation</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-17\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new12\">\n<p><a href=\"https://developer.nvidia.com/blog/nvidia-blackwell-leads-on-new-semianalysis-inferencemax-benchmarks/\" rel=\"noopener noreferrer\">NVIDIA Blackwell Leads on InferenceMAX Benchmarks</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new12\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new12-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-19\">\n<p><a href=\"https://northflank.com/blog/vllm-vs-tensorrt-llm-and-how-to-run-them\" rel=\"noopener noreferrer\">vLLM vs TensorRT-LLM Comparison</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-19\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new13\">\n<p><a href=\"https://github.com/vllm-project/vllm/releases\" rel=\"noopener noreferrer\">vLLM Releases - GitHub</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new13\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new14\">\n<p><a href=\"https://siliconangle.com/2026/01/22/inferact-launches-150m-funding-commercialize-vllm/\" rel=\"noopener noreferrer\">Inferact launches with $150M to commercialize vLLM</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new14\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new15\">\n<p><a href=\"https://github.com/sgl-project/sglang\" rel=\"noopener noreferrer\">SGLang GitHub Repository</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new15\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new16\">\n<p><a href=\"https://research.aimultiple.com/inference-engines/\" rel=\"noopener noreferrer\">LLM Inference Engines: vLLM vs LMDeploy vs SGLang 2026</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new16\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new17\">\n<p><a href=\"https://huggingface.co/docs/text-generation-inference/en/index\" rel=\"noopener noreferrer\">Hugging Face Text Generation Inference</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new17\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-20\">\n<p><a href=\"https://www.marktechpost.com/2025/11/19/vllm-vs-tensorrt-llm-vs-hf-tgi-vs-lmdeploy-a-deep-technical-comparison-for-production-llm-inference/\" rel=\"noopener noreferrer\">vLLM vs TensorRT-LLM vs TGI Technical Comparison</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-20\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new18\">\n<p><a href=\"https://developer.nvidia.com/blog/nvidia-blackwell-ultra-sets-new-inference-records-in-mlperf-debut/\" rel=\"noopener noreferrer\">NVIDIA Blackwell Ultra Sets New Inference Records in MLPerf</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new18\" class=\"data-footnote-backref\">↩</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new18-2\" class=\"data-footnote-backref\">↩<sup>2</sup></a></p>\n</li>\n<li id=\"user-content-fn-new19\">\n<p><a href=\"https://nvidianews.nvidia.com/news/nvidia-blackwell-platform-arrives-to-power-a-new-era-of-computing\" rel=\"noopener noreferrer\">NVIDIA Blackwell Platform Arrives</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new19\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-new20\">\n<p><a href=\"https://blogs.nvidia.com/blog/blackwell-inferencemax-benchmark-results/\" rel=\"noopener noreferrer\">NVIDIA Blackwell Raises Bar in InferenceMAX Benchmarks</a> <a href=\"https://blog.ecitis.org/tensorrt-triton-inference/#user-content-fnref-new20\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/tensorrt-triton-inference/",
            "title": "NVIDIA TensorRT and Triton: Production LLM Inference",
            "summary": "Build a production inference stack that balances latency, throughput, GPU utilization, concurrency, and observability.",
            "image": "https://blog.ecitis.org/open-graph/tensorrt-triton-inference.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "TensorRT-LLM",
                "Triton",
                "Serving"
            ]
        },
        {
            "id": "https://blog.ecitis.org/voice-audio-models/",
            "content_html": "<p>Voice and audio AI has advanced rapidly, with models now capable of real-time transcription, natural speech synthesis, voice cloning from seconds of audio, and full-duplex conversational interactions. This article provides a technical survey of the architectures behind speech-to-text, text-to-speech, voice conversion, and audio language models. We cover the mathematical foundations, practical training considerations, and deployment strategies for building production systems.</p>\n<hr />\n<h2 id=\"speech-to-text-models\">Speech-to-Text Models</h2>\n<p>Automatic speech recognition (ASR) converts audio waveforms into text. Modern ASR systems have moved from traditional hidden Markov models with hand-crafted acoustic features to end-to-end neural networks trained on massive datasets.</p>\n<h3 id=\"whisper-architecture\">Whisper Architecture</h3>\n<p><a href=\"https://github.com/openai/whisper\">Whisper</a>, released by OpenAI in September 2022, is a Transformer-based encoder-decoder model trained on 680,000 hours of multilingual and multitask supervised data collected from the web [1].</p>\n<p><strong>Input Processing</strong></p>\n<p>Audio is resampled to 16 kHz and converted to an 80-channel log-magnitude Mel spectrogram using 25ms windows with 10ms stride. The spectrogram is normalized to a [-1, 1] range with near-zero mean. For the large-v3 model, this was increased to 128 Mel frequency bins [2].</p>\n<p><strong>Encoder</strong></p>\n<p>The encoder processes the Mel spectrogram through:</p>\n<ol>\n<li>Two convolutional layers for initial downsampling</li>\n<li>Sinusoidal positional embeddings added to the sequence</li>\n<li>A stack of Transformer encoder blocks with pre-activation residual connections</li>\n<li>Layer normalization on the final output</li>\n</ol>\n<figure><figcaption><strong>Whisper Encoder Architecture</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-0.webp\" width=\"512\" height=\"291\" alt=\"Whisper Encoder Architecture\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Decoder</strong></p>\n<p>The decoder follows the standard Transformer decoder architecture with learned positional embeddings and tied input-output token representations (same weight matrix for input embeddings and output projection). It uses byte-pair encoding tokenization similar to GPT-2.</p>\n<p><strong>Model Sizes</strong></p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Parameters</th><th>Layers</th><th>Width</th><th>Heads</th></tr></thead><tbody><tr><td>tiny</td><td>39M</td><td>4</td><td>384</td><td>6</td></tr><tr><td>base</td><td>74M</td><td>6</td><td>512</td><td>8</td></tr><tr><td>small</td><td>244M</td><td>12</td><td>768</td><td>12</td></tr><tr><td>medium</td><td>769M</td><td>24</td><td>1024</td><td>16</td></tr><tr><td>large-v3</td><td>1.55B</td><td>32</td><td>1280</td><td>20</td></tr></tbody></table></div>\n<p>Whisper’s encoder-decoder structure makes it naturally suited for batch processing but challenging for streaming applications. The encoder requires the full 30-second audio context before producing useful representations [3].</p>\n<h3 id=\"conformer-architecture\">Conformer Architecture</h3>\n<p>The <a href=\"https://arxiv.org/abs/2005.08100\">Conformer</a> (Convolution-augmented Transformer), introduced by Google in 2020, addresses a fundamental limitation of pure Transformers: while self-attention captures global context well, it struggles with local feature patterns that are important for speech [4].</p>\n<p><strong>Key Insight</strong></p>\n<p>Speech signals contain both local patterns (phonemes, formants) and global dependencies (grammar, semantics). CNNs excel at local feature extraction while Transformers handle global context. Conformer combines both.</p>\n<p><strong>Block Structure</strong></p>\n<p>The Conformer uses a “macaron” structure where two feed-forward layers sandwich the attention and convolution modules:</p>\n<figure><figcaption><strong>Conformer Block Structure</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-1.webp\" width=\"512\" height=\"718\" alt=\"Conformer Block Structure\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Convolution Module Details</strong></p>\n<p>The convolution module uses depthwise separable convolutions for efficiency:</p>\n<ol>\n<li><strong>Pointwise conv</strong> with expansion factor 2 followed by GLU activation</li>\n<li><strong>Depthwise conv</strong> (1D, kernel size typically 31) captures local context</li>\n<li><strong>BatchNorm</strong> for training stability</li>\n<li><strong>Swish activation</strong> (x * sigmoid(x))</li>\n<li><strong>Pointwise conv</strong> to project back to original dimension</li>\n</ol>\n<p>The depthwise convolution has kernel size 31 by default, meaning each output position attends to 15 frames on each side (roughly 150ms of audio context at standard 10ms frame rates).</p>\n<p><strong>Performance</strong></p>\n<p>On LibriSpeech, Conformer achieves 2.1%/4.3% WER without a language model and 1.9%/3.9% with an external language model on test-clean/test-other [4]. This made it the dominant architecture for production ASR systems through 2024.</p>\n<h3 id=\"ctc-vs-attention-based-decoding\">CTC vs Attention-Based Decoding</h3>\n<p>ASR systems use two main decoding paradigms with different tradeoffs [5].</p>\n<p><strong>Connectionist Temporal Classification (CTC)</strong></p>\n<p>CTC, introduced by Graves et al. in 2006, addresses the alignment problem in sequence-to-sequence tasks where input and output sequences have different lengths.</p>\n<p>Key properties:</p>\n<ul>\n<li><strong>Monotonic alignment</strong>: Output tokens appear in the same order as their corresponding input frames</li>\n<li><strong>Conditional independence</strong>: CTC assumes each output token is independent given the input, which simplifies computation but ignores output dependencies</li>\n<li><strong>Blank token</strong>: A special “blank” token allows the model to output nothing for frames that fall between phonemes</li>\n</ul>\n<figure><figcaption><strong>CTC Decoding Example</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-2.webp\" width=\"512\" height=\"197\" alt=\"CTC Decoding Example\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>CTC loss marginalizes over all possible alignments between input and output:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>P(Y|X) = Σ P(A|X)  for all valid alignments A</span></span></code></pre>\n<p>Models like Wav2Vec2, HuBERT, and M-CTC-T use CTC [5]. The main advantage is <strong>non-autoregressive decoding</strong>: all output tokens can be computed in parallel, enabling faster inference.</p>\n<p><strong>Attention-Based Encoder-Decoder (AED)</strong></p>\n<p>AED models (like Whisper) use cross-attention to learn soft alignments between encoder outputs and decoder states:</p>\n<ul>\n<li><strong>No independence assumption</strong>: Each output token conditions on all previous tokens</li>\n<li><strong>Implicit language model</strong>: The decoder learns language patterns from training data</li>\n<li><strong>Flexible alignment</strong>: Can handle non-monotonic mappings (useful for translation)</li>\n</ul>\n<p>The downside is <strong>autoregressive decoding</strong>: tokens must be generated sequentially, increasing latency.</p>\n<p><strong>Hybrid CTC/Attention</strong></p>\n<p>Many modern systems combine both approaches. During training, a CTC loss is applied to the encoder output as a regularizer, encouraging monotonic alignment. During inference, CTC scores can be used to constrain beam search [5].</p>\n<p>The RNN-Transducer (RNN-T), used by Google and other production systems, extends CTC with a prediction network that models output dependencies without full autoregressive decoding.</p>\n<h3 id=\"streaming-asr-challenges-and-solutions\">Streaming ASR Challenges and Solutions</h3>\n<p>Real-time applications like voice assistants require streaming ASR that transcribes speech with minimal latency. This is challenging because Transformer models rely on full-sequence attention [6].</p>\n<p><strong>The Problem</strong></p>\n<p>Standard self-attention has O(n^2) complexity and requires the entire sequence:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>Attention(Q, K, V) = softmax(QK^T / √d) V</span></span></code></pre>\n<p>For a 30-second audio clip at 50 frames/second, this means 1500 frames must be available before processing begins.</p>\n<p><strong>Chunked Attention</strong></p>\n<p>The primary solution is chunked (or blockwise) attention:</p>\n<ol>\n<li>Divide input into fixed-size chunks (e.g., 640ms)</li>\n<li>Each chunk attends only to itself and a limited left context</li>\n<li>Process chunks incrementally as audio arrives</li>\n</ol>\n<figure><figcaption><strong>Chunked Attention for Streaming</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-3.webp\" width=\"512\" height=\"279\" alt=\"Chunked Attention for Streaming\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Tradeoffs</strong></p>\n<ul>\n<li>Smaller chunks = lower latency, worse accuracy</li>\n<li>Larger left context = better accuracy, higher memory and compute</li>\n</ul>\n<p>Research shows that chunked attention with 32-48 frames of left context and 16-32 frame chunks achieves a reasonable balance. The SSCFormer architecture uses sequentially sampled chunks with causal convolutions to improve accuracy within the streaming constraint [6].</p>\n<p><strong>Positional Encoding for Streaming</strong></p>\n<p>With chunked attention, relative positional encodings work better than absolute ones. The model only needs to represent distances up to the maximum context window, not positions within an indefinitely long stream.</p>\n<hr />\n<h2 id=\"text-to-speech-models\">Text-to-Speech Models</h2>\n<p>Text-to-speech (TTS) converts text into natural-sounding audio. Modern TTS has evolved from concatenative synthesis (splicing recorded audio) through parametric synthesis to end-to-end neural approaches.</p>\n<h3 id=\"vits-architecture\">VITS Architecture</h3>\n<p><a href=\"https://arxiv.org/abs/2106.06103\">VITS</a> (Variational Inference with adversarial learning for end-to-end Text-to-Speech) unifies the acoustic model and vocoder into a single end-to-end framework using variational inference and adversarial training [7].</p>\n<p><strong>Key Innovation</strong></p>\n<p>Traditional TTS pipelines generate mel spectrograms from text, then convert spectrograms to waveforms with a separate vocoder. This two-stage approach can cause mismatch artifacts. VITS generates waveforms directly from text.</p>\n<p><strong>Architecture Components</strong></p>\n<figure><figcaption><strong>VITS Architecture</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-4.webp\" width=\"512\" height=\"576\" alt=\"VITS Architecture\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Variational Inference Framework</strong></p>\n<p>VITS models TTS as a conditional VAE:</p>\n<ol>\n<li><strong>Posterior encoder</strong>: During training, encodes ground-truth mel spectrogram into latent z</li>\n<li><strong>Prior encoder</strong>: Learns to predict z from text alone (with normalizing flows to increase expressiveness)</li>\n<li><strong>Decoder</strong>: Generates waveform from z</li>\n</ol>\n<p>The training objective combines:</p>\n<ul>\n<li>Reconstruction loss (from VAE)</li>\n<li>KL divergence between prior and posterior</li>\n<li>Adversarial loss (discriminator judges waveform quality)</li>\n<li>Feature matching loss (compare discriminator activations)</li>\n</ul>\n<p><strong>Stochastic Duration Predictor</strong></p>\n<p>Unlike deterministic duration predictors, VITS uses a stochastic duration predictor that models duration as a distribution. This enables generating the same text with different speaking rates and rhythms, capturing the natural variability in human speech.</p>\n<p><strong>VITS2 Improvements</strong></p>\n<p>VITS2 enhances the original with:</p>\n<ul>\n<li>Improved duration prediction using a transformer-based predictor</li>\n<li>Better speaker conditioning for multi-speaker models</li>\n<li>Monotonic alignment search for more stable training</li>\n</ul>\n<h3 id=\"tacotron-family-evolution\">Tacotron Family Evolution</h3>\n<p>The <a href=\"https://arxiv.org/abs/1712.05884\">Tacotron</a> models pioneered end-to-end TTS using sequence-to-sequence learning [8].</p>\n<p><strong>Tacotron 1 (2017)</strong></p>\n<ul>\n<li>Character-level input (no phoneme conversion needed)</li>\n<li>Encoder: CBHG (1-D convolution bank + highway network + bidirectional GRU)</li>\n<li>Attention: Content-based attention with location features</li>\n<li>Decoder: Autoregressive GRU predicting mel spectrogram frames</li>\n<li>Vocoder: Griffin-Lim algorithm (fast but low quality)</li>\n</ul>\n<p><strong>Tacotron 2 (2017)</strong></p>\n<p>Key improvements over Tacotron 1:</p>\n<ul>\n<li>Simplified encoder: 3 conv layers + bidirectional LSTM</li>\n<li>Location-sensitive attention: Adds previous alignment to attention computation</li>\n<li>Decoder: 2 autoregressive LSTM layers</li>\n<li>PostNet: 5 conv layers to refine mel spectrogram</li>\n<li>Vocoder: WaveNet (high quality but slow)</li>\n</ul>\n<figure><figcaption><strong>Tacotron 2 Architecture</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-5.webp\" width=\"512\" height=\"886\" alt=\"Tacotron 2 Architecture\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>Tacotron 2 achieved MOS of 4.53, approaching the 4.58 MOS of professionally recorded speech [8].</p>\n<p><strong>FastSpeech (2019) and FastSpeech 2 (2020)</strong></p>\n<p>FastSpeech addressed Tacotron’s slow autoregressive inference:</p>\n<ul>\n<li><strong>Non-autoregressive</strong>: Generates all mel frames in parallel</li>\n<li><strong>Duration predictor</strong>: Explicitly models phoneme durations</li>\n<li><strong>Length regulator</strong>: Expands text sequence to match mel length</li>\n</ul>\n<p>FastSpeech 2 added pitch and energy predictors for more controllable synthesis. It achieves RTF of ~0.02 on V100 (50x faster than real-time).</p>\n<h3 id=\"neural-vocoders\">Neural Vocoders</h3>\n<p>Vocoders convert mel spectrograms (or other intermediate representations) into audio waveforms. This is an inverse problem: spectrograms discard phase information, which the vocoder must reconstruct.</p>\n<p><strong>HiFi-GAN</strong></p>\n<p><a href=\"https://github.com/jik876/hifi-gan\">HiFi-GAN</a>, published in 2020, uses GANs for efficient high-fidelity synthesis [9].</p>\n<p>Architecture:</p>\n<ul>\n<li><strong>Generator</strong>: Transposed convolutions for upsampling, multi-receptive field fusion (MRF) blocks</li>\n<li><strong>Multi-period discriminator (MPD)</strong>: Multiple discriminators operating on different periodic subsequences (periods 2, 3, 5, 7, 11)</li>\n<li><strong>Multi-scale discriminator (MSD)</strong>: Discriminators at different audio resolutions</li>\n</ul>\n<p>The key insight is that speech contains multiple periodic components (fundamental frequency and harmonics). By having discriminators focus on different periodicities, the model learns to generate all these components correctly.</p>\n<figure><figcaption><strong>HiFi-GAN Generator</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-6.webp\" width=\"512\" height=\"628\" alt=\"HiFi-GAN Generator\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>HiFi-GAN V1 (14M params) generates audio 13.4x faster than real-time on CPU.</p>\n<p><strong>BigVGAN</strong></p>\n<p><a href=\"https://arxiv.org/abs/2206.04658\">BigVGAN</a>, from NVIDIA in 2022, scales up HiFi-GAN with architectural improvements [10]:</p>\n<ol>\n<li>\n<p><strong>Snake activation</strong>: Periodic activation function x + sin^2(x)/a with learned frequency. Provides inductive bias for generating periodic waveforms.</p>\n</li>\n<li>\n<p><strong>Anti-aliased representation</strong>: Low-pass filtering before downsampling in discriminator to prevent aliasing artifacts.</p>\n</li>\n<li>\n<p><strong>Larger scale</strong>: Up to 112M parameters (vs 14M for HiFi-GAN V1)</p>\n</li>\n</ol>\n<p>BigVGAN trained only on clean speech (LibriTTS) generalizes to unseen speakers, languages, singing, and even instrumental music without fine-tuning.</p>\n<h3 id=\"zero-shot-tts\">Zero-Shot TTS</h3>\n<p>Zero-shot TTS generates speech in any voice given only a few seconds of reference audio.</p>\n<p><strong>VALL-E</strong></p>\n<p><a href=\"https://www.microsoft.com/en-us/research/project/vall-e-x/\">VALL-E</a>, from Microsoft in 2023, treats TTS as a language modeling problem [11]:</p>\n<ol>\n<li>Audio is tokenized using a neural codec (EnCodec) into discrete codes at 75 Hz with 8 codebook levels</li>\n<li>A Transformer language model is trained to predict audio tokens given text and a 3-second acoustic prompt</li>\n<li>The first codebook level captures semantic content; subsequent levels add acoustic detail</li>\n</ol>\n<figure><figcaption><strong>VALL-E Zero-Shot TTS</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-7.webp\" width=\"512\" height=\"502\" alt=\"VALL-E Zero-Shot TTS\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>VALL-E 2 achieved human parity on LibriSpeech and VCTK using repetition-aware sampling and grouped code modeling [11].</p>\n<p><strong>XTTS</strong></p>\n<p><a href=\"https://arxiv.org/html/2406.04904v1\">XTTS</a>, from Coqui, builds on Tortoise with improvements for multilingual zero-shot TTS [12]:</p>\n<ul>\n<li>VQ-VAE encodes mel spectrograms at 21.53 Hz (vs 75 Hz for VALL-E, reducing sequence length)</li>\n<li>GPT-2 decoder (443M params) predicts audio tokens</li>\n<li>Perceiver architecture for speaker conditioning: processes reference mel spectrogram into 32 latent vectors</li>\n<li>Supports 16 languages with SOTA results</li>\n</ul>\n<p>XTTS v2 achieves RTF of 0.48 with 200ms time-to-first-chunk for streaming applications [12].</p>\n<p><strong>F5-TTS</strong></p>\n<p><a href=\"https://gradientflow.com/f5-tts/\">F5-TTS</a> uses a non-autoregressive design with flow matching:</p>\n<ul>\n<li>Requires only 2,994 MB GPU memory</li>\n<li>Better for resource-constrained environments</li>\n<li>Trades streaming capability for efficiency</li>\n</ul>\n<h3 id=\"how-voice-cloning-works\">How Voice Cloning Works</h3>\n<p>Voice cloning extracts a speaker’s characteristics from reference audio and applies them to synthesize new speech [13].</p>\n<p><strong>Speaker Encoder</strong></p>\n<p>A neural network extracts a fixed-dimension speaker embedding from reference audio:</p>\n<ol>\n<li>Input: Mel spectrogram of reference audio (3-10 seconds)</li>\n<li>Architecture: Often 3-layer LSTM or ECAPA-TDNN</li>\n<li>Output: 256-dim speaker embedding vector (d-vector)</li>\n</ol>\n<figure><figcaption><strong>Speaker Encoder Pipeline</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-8.webp\" width=\"512\" height=\"489\" alt=\"Speaker Encoder Pipeline\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Embedding-Based Adaptation</strong></p>\n<p>The speaker embedding conditions the TTS model:</p>\n<ol>\n<li><strong>Concatenation</strong>: Append embedding to encoder output</li>\n<li><strong>Addition</strong>: Add embedding to hidden states</li>\n<li><strong>FiLM</strong>: Use embedding to predict scale/shift parameters for normalization layers</li>\n</ol>\n<p>More advanced approaches use attention-based conditioning (like XTTS’s Perceiver) to capture finer-grained speaker characteristics.</p>\n<p><strong>Quality vs Data</strong></p>\n<ul>\n<li>Zero-shot (3-10s reference): Captures voice timbre but may miss speaking style</li>\n<li>Few-shot fine-tuning (1-5 min): Better style transfer, requires training</li>\n<li>Full fine-tuning (30+ min): Highest quality, significant compute cost</li>\n</ul>\n<hr />\n<h2 id=\"voice-conversion-and-cloning\">Voice Conversion and Cloning</h2>\n<p>Voice conversion transforms speech from one speaker to sound like another while preserving linguistic content.</p>\n<h3 id=\"speaker-embedding-extraction\">Speaker Embedding Extraction</h3>\n<p>Modern voice conversion uses self-supervised models like HuBERT to extract speaker-independent content representations [14].</p>\n<p><strong>HuBERT for Content</strong></p>\n<p>HuBERT (Hidden-Unit BERT) is trained via masked prediction on audio:</p>\n<ol>\n<li>Extract MFCC features</li>\n<li>Cluster MFCCs with k-means to create pseudo-labels</li>\n<li>Train Transformer to predict masked pseudo-labels</li>\n<li>Iterate: use model outputs as new clustering targets</li>\n</ol>\n<p>The resulting representations capture phonetic content while being somewhat speaker-invariant. RVC uses layer 12 of HuBERT as content features [14].</p>\n<p><strong>ECAPA-TDNN for Speaker</strong></p>\n<p>ECAPA-TDNN extracts speaker embeddings:</p>\n<ul>\n<li>Time Delay Neural Network with multi-scale features</li>\n<li>Squeeze-and-excitation blocks for channel attention</li>\n<li>Attentive statistics pooling</li>\n<li>Trained on speaker verification (distinguish speakers)</li>\n</ul>\n<h3 id=\"disentanglement-of-content-and-speaker\">Disentanglement of Content and Speaker</h3>\n<p>The core challenge in voice conversion is separating “what is said” from “who said it” [15].</p>\n<p><strong>Information Bottleneck</strong></p>\n<p>One approach uses information bottlenecks to force separation:</p>\n<figure><figcaption><strong>Voice Conversion Disentanglement</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-9.webp\" width=\"512\" height=\"449\" alt=\"Voice Conversion Disentanglement\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Adversarial Disentanglement</strong></p>\n<p>Train with adversarial losses to ensure:</p>\n<ul>\n<li>Content encoder output cannot predict speaker (speaker classifier fails)</li>\n<li>Speaker encoder output cannot predict content (ASR fails)</li>\n</ul>\n<p><strong>CONTENTVEC Approach</strong></p>\n<p>CONTENTVEC converts all training audio to a single speaker using an unsupervised VC system, then trains HuBERT on this speaker-normalized data. The resulting representations contain minimal speaker information [15].</p>\n<h3 id=\"rvc-architecture\">RVC Architecture</h3>\n<p><a href=\"https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI\">Retrieval-based Voice Conversion (RVC)</a> is an open-source system achieving high-quality conversion with minimal data [14].</p>\n<p><strong>Pipeline</strong></p>\n<figure><figcaption><strong>RVC Voice Conversion Pipeline</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-10.webp\" width=\"512\" height=\"578\" alt=\"RVC Voice Conversion Pipeline\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Retrieval Module</strong></p>\n<p>RVC’s key innovation is the retrieval step:</p>\n<ol>\n<li>During training: Store all HuBERT features from target speaker in a Faiss index</li>\n<li>During inference: For each source feature, find k nearest neighbors in the index</li>\n<li>Blend source features with retrieved features (configurable ratio)</li>\n</ol>\n<p>This reduces “timbre leakage” by replacing source speaker characteristics with training set features.</p>\n<p><strong>Performance</strong></p>\n<p>RVC achieves 90ms end-to-end latency with ASIO audio interfaces and learns high-quality transformations from about 10 minutes of target speaker audio [14].</p>\n<h3 id=\"so-vits-svc\">So-VITS-SVC</h3>\n<p><a href=\"https://github.com/svc-develop-team/so-vits-svc\">So-VITS-SVC</a> (SoftVC VITS Singing Voice Conversion) specializes in singing voice conversion [16].</p>\n<p><strong>Differences from Speech VC</strong></p>\n<p>Singing voice conversion has additional challenges:</p>\n<ul>\n<li>Must preserve pitch contour exactly (wrong notes are obvious)</li>\n<li>Longer sustained vowels expose synthesis artifacts</li>\n<li>Vibrato, breath, and other expressive elements must transfer</li>\n</ul>\n<p><strong>Architecture</strong></p>\n<p>So-VITS-SVC uses:</p>\n<ol>\n<li><strong>SoftVC content encoder</strong>: Variant of HuBERT trained for voice conversion</li>\n<li><strong>F0 predictor</strong>: Optional pitch prediction (disabled for singing to preserve exact pitch)</li>\n<li><strong>VITS backbone</strong>: Prior encoder (6-layer Transformer with 2-head attention) + HiFi-GAN decoder</li>\n<li><strong>NSF-HiFiGAN vocoder</strong>: Addresses glitching artifacts in original HiFi-GAN</li>\n</ol>\n<p><strong>Shallow Diffusion Enhancement</strong></p>\n<p>Recent versions add optional diffusion refinement:</p>\n<ol>\n<li>VITS generates initial waveform</li>\n<li>Diffusion model (from DDSP-SVC) refines quality</li>\n<li>Only runs a few diffusion steps (“shallow”) for efficiency</li>\n</ol>\n<p><strong>Clustering for Timbre Matching</strong></p>\n<p>Similar to RVC’s retrieval, So-VITS-SVC uses feature clustering:</p>\n<ul>\n<li>K-means cluster centers represent “prototypical” target speaker features</li>\n<li>Source features are blended with cluster centers</li>\n<li>Tradeoff: More clustering = better timbre match, less clarity</li>\n</ul>\n<hr />\n<h2 id=\"audio-language-models\">Audio Language Models</h2>\n<p>Audio language models extend the language modeling paradigm to generate audio directly, enabling unified speech-text systems.</p>\n<h3 id=\"audiolm-and-audio-tokenization\">AudioLM and Audio Tokenization</h3>\n<p><a href=\"https://arxiv.org/abs/2209.03143\">AudioLM</a>, from Google in 2022, pioneered treating audio generation as language modeling [17].</p>\n<p><strong>Hybrid Tokenization</strong></p>\n<p>AudioLM uses two types of tokens to capture different aspects of audio:</p>\n<ol>\n<li><strong>Semantic tokens</strong> (from w2v-BERT): Capture content, phonetics, rhythm, harmony</li>\n<li><strong>Acoustic tokens</strong> (from SoundStream): Capture timbre, recording quality, fine acoustic details</li>\n</ol>\n<figure><figcaption><strong>AudioLM Hybrid Tokenization</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-11.webp\" width=\"512\" height=\"370\" alt=\"AudioLM Hybrid Tokenization\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Three-Stage Generation</strong></p>\n<p>AudioLM generates hierarchically:</p>\n<ol>\n<li><strong>Semantic modeling</strong>: Transformer predicts semantic tokens (captures structure)</li>\n<li><strong>Coarse acoustic modeling</strong>: Transformer predicts first few acoustic codebook levels given semantics</li>\n<li><strong>Fine acoustic modeling</strong>: Transformer adds remaining codebook levels for full quality</li>\n</ol>\n<p>This cascade allows the model to first get the high-level content right, then progressively add acoustic detail.</p>\n<p><strong>Capabilities</strong></p>\n<p>Without any text supervision, AudioLM generates:</p>\n<ul>\n<li>Coherent speech continuations that maintain speaker identity, grammar, and semantic coherence</li>\n<li>Piano music with proper harmony and rhythm</li>\n<li>General audio with consistent acoustic properties</li>\n</ul>\n<h3 id=\"speechgpt-and-multimodal-audio-text-models\">SpeechGPT and Multimodal Audio-Text Models</h3>\n<p><a href=\"https://www.alphaxiv.org/overview/2305.11000v2\">SpeechGPT</a> extends LLMs to natively understand and generate speech [18].</p>\n<p><strong>Architecture</strong></p>\n<ol>\n<li><strong>Speech tokenization</strong>: HuBERT converts speech to discrete tokens</li>\n<li><strong>Vocabulary expansion</strong>: LLM vocabulary extended to include speech tokens</li>\n<li><strong>Unified modeling</strong>: Single model processes interleaved text and speech</li>\n</ol>\n<figure><figcaption><strong>SpeechGPT Multimodal Architecture</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-12.webp\" width=\"512\" height=\"449\" alt=\"SpeechGPT Multimodal Architecture\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Training Stages</strong></p>\n<ol>\n<li><strong>Modality adaptation</strong>: Train speech encoder/decoder with frozen LLM</li>\n<li><strong>Cross-modal instruction tuning</strong>: Train on speech-text tasks</li>\n<li><strong>Chain-of-modality tuning</strong>: Generate “thought” in text before speech output (like chain-of-thought)</li>\n</ol>\n<p>The key insight is that by representing speech as tokens within the LLM’s vocabulary, knowledge transfers between modalities. The model can answer questions about speech content, translate speech, or generate speech responses.</p>\n<h3 id=\"moshi-real-time-conversational-ai\">Moshi: Real-Time Conversational AI</h3>\n<p><a href=\"https://kyutai.org/Moshi.pdf\">Moshi</a>, from Kyutai (French AI lab), is the first real-time full-duplex spoken dialogue system [19].</p>\n<p><strong>Key Innovation: Parallel Audio Streams</strong></p>\n<p>Moshi models two simultaneous audio streams:</p>\n<ul>\n<li><strong>Moshi’s speech</strong>: What the AI is saying</li>\n<li><strong>User’s speech</strong>: What the human is saying</li>\n</ul>\n<p>This removes turn-taking constraints. The model always listens and always generates (speech or silence), enabling natural interruptions and backchannels (“uh-huh”, “right”).</p>\n<figure><figcaption><strong>Full-Duplex Conversation</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-13.webp\" width=\"512\" height=\"288\" alt=\"Full-Duplex Conversation\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Architecture</strong></p>\n<figure><figcaption><strong>Moshi Architecture</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/voice-audio-models-14.webp\" width=\"512\" height=\"393\" alt=\"Moshi Architecture\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p><strong>Mimi Codec</strong></p>\n<p>Moshi uses a custom neural codec called Mimi:</p>\n<ul>\n<li>Residual vector quantization (RVQ) like EnCodec</li>\n<li>Optimized for streaming with low latency</li>\n<li>12.5 Hz frame rate for semantic tokens, higher for acoustic</li>\n</ul>\n<p><strong>Inner Monologue</strong></p>\n<p>Moshi generates time-aligned text tokens as a “prefix” to audio tokens:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>Internal: [text: \"Hello\"]  [audio: \"Hello\"] [text: \"how\"] [audio: \"how\"]...</span></span></code></pre>\n<p>This provides:</p>\n<ul>\n<li>Implicit speech recognition (text output)</li>\n<li>Better linguistic quality (text grounds the audio)</li>\n<li>Debugging capability (see what model “thinks”)</li>\n</ul>\n<p><strong>Performance</strong></p>\n<ul>\n<li>160ms theoretical latency (200ms practical)</li>\n<li>Full-duplex conversation handling</li>\n<li>92+ different voice intonations</li>\n<li>Released under CC-BY 4.0 license</li>\n</ul>\n<h3 id=\"moe-in-audio-models\">MoE in Audio Models</h3>\n<p>Mixture of Experts (MoE) enables scaling model capacity without proportionally increasing compute. Recent work applies MoE to audio [20].</p>\n<p><strong>MoME (Mixture of Matryoshka Experts)</strong></p>\n<p>For audio-visual speech recognition:</p>\n<ul>\n<li>Integrates sparse MoE into matryoshka representation learning</li>\n<li>Top-k routing activates subset of experts per token</li>\n<li>Shared experts handle common patterns; routed experts specialize</li>\n<li>Achieves SOTA on LRS2/LRS3 with fewer parameters</li>\n</ul>\n<p><strong>MoHAVE (Mixture of Hierarchical Audio-Visual Experts)</strong></p>\n<p>Uses hierarchical gating:</p>\n<ul>\n<li>Modality-specific expert groups (audio vs visual)</li>\n<li>Dynamic activation based on input context</li>\n<li>Scales capacity without linear compute increase</li>\n</ul>\n<p><strong>Practical Considerations</strong></p>\n<p>MoE for audio faces challenges:</p>\n<ul>\n<li>Load balancing: Ensuring experts are utilized evenly</li>\n<li>Expert collapse: Preventing all tokens from routing to same expert</li>\n<li>Memory: All experts must fit in memory even if only subset is active</li>\n</ul>\n<hr />\n<h2 id=\"training-and-data\">Training and Data</h2>\n<p>Training voice models requires careful data preparation and domain-specific techniques.</p>\n<h3 id=\"audio-preprocessing\">Audio Preprocessing</h3>\n<p><strong>Mel Spectrograms</strong></p>\n<p>The standard intermediate representation for speech models [21]:</p>\n<ol>\n<li><strong>Resampling</strong>: Typically to 16 kHz (ASR) or 22.05/24 kHz (TTS)</li>\n<li><strong>STFT</strong>: Short-time Fourier transform with:\n<ul>\n<li>Window size: 25ms (400 samples at 16 kHz)</li>\n<li>Hop size: 10ms (160 samples)</li>\n<li>FFT size: 512 or 1024</li>\n</ul>\n</li>\n<li><strong>Mel filterbank</strong>: Apply triangular filters spaced on mel scale (80-128 bins)</li>\n<li><strong>Log compression</strong>: log(mel + 1e-5) to compress dynamic range</li>\n<li><strong>Normalization</strong>: Per-channel or global mean/variance normalization</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> librosa</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Standard preprocessing</span></span>\n<span class=\"line\"><span>audio</span><span>,</span><span> sr </span><span>=</span><span> librosa</span><span>.</span><span>load</span><span>(path, sr</span><span>=</span><span>16000</span><span>)</span></span>\n<span class=\"line\"><span>mel </span><span>=</span><span> librosa</span><span>.</span><span>feature</span><span>.</span><span>melspectrogram</span><span>(</span></span>\n<span class=\"line\"><span>    y</span><span>=</span><span>audio,</span></span>\n<span class=\"line\"><span>    sr</span><span>=</span><span>sr,</span></span>\n<span class=\"line\"><span>    n_fft</span><span>=</span><span>1024</span><span>,</span></span>\n<span class=\"line\"><span>    hop_length</span><span>=</span><span>256</span><span>,</span></span>\n<span class=\"line\"><span>    n_mels</span><span>=</span><span>80</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>log_mel </span><span>=</span><span> np</span><span>.</span><span>log</span><span>(mel </span><span>+</span><span> 1e-5</span><span>)</span></span></code></pre>\n<p><strong>MFCC</strong></p>\n<p>Mel-Frequency Cepstral Coefficients add DCT to mel spectrograms:</p>\n<ol>\n<li>Compute log mel spectrogram</li>\n<li>Apply Discrete Cosine Transform</li>\n<li>Keep first 13-40 coefficients</li>\n</ol>\n<p>MFCCs decorrelate features and compress representation. They’re more common in traditional ASR but less used in end-to-end neural models which learn their own representations [21].</p>\n<p><strong>Raw Waveform</strong></p>\n<p>Some models (Wav2Vec2, HuBERT) operate directly on raw waveforms:</p>\n<ul>\n<li>No information loss from spectrogram conversion</li>\n<li>Model learns appropriate filterbanks</li>\n<li>Requires more compute and data</li>\n</ul>\n<h3 id=\"data-quality-requirements\">Data Quality Requirements</h3>\n<p><strong>ASR Data</strong></p>\n<ul>\n<li>Volume: State-of-the-art requires 10,000+ hours (Whisper: 680,000 hours)</li>\n<li>Transcription: Can use weak supervision (web captions) but human transcription helps</li>\n<li>Diversity: Multiple speakers, accents, recording conditions, noise levels</li>\n<li>Alignment: Exact timestamps not required for seq2seq models</li>\n</ul>\n<p><strong>TTS Data</strong></p>\n<ul>\n<li>Volume: 10-50 hours for single-speaker high-quality</li>\n<li>Recording quality: Studio recordings preferred (low noise, consistent mic)</li>\n<li>Transcription: Must be exact (punctuation affects prosody)</li>\n<li>Consistency: Same speaker, same emotional register, same recording setup</li>\n<li>Alignment: Phone-level timestamps improve training stability</li>\n</ul>\n<p>For voice cloning:</p>\n<ul>\n<li>Zero-shot: 3-10 seconds reference audio</li>\n<li>Few-shot fine-tuning: 1-5 minutes</li>\n<li>Full fine-tuning: 30+ minutes</li>\n</ul>\n<h3 id=\"audio-text-alignment\">Audio-Text Alignment</h3>\n<p>Forced alignment aligns transcripts to audio at the word or phone level [22].</p>\n<p><strong>Montreal Forced Aligner (MFA)</strong></p>\n<p>Standard tool for TTS data preparation:</p>\n<ol>\n<li><strong>Input</strong>: Audio + orthographic transcript</li>\n<li><strong>Pronunciation dictionary</strong>: Maps words to phoneme sequences</li>\n<li><strong>Acoustic model</strong>: GMM-HMM trained on target language</li>\n<li><strong>Output</strong>: TextGrid with word/phone timestamps</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Example MFA usage</span></span>\n<span class=\"line\"><span>mfa</span><span> align</span><span> /path/to/audio</span><span> /path/to/dictionary</span><span> /path/to/model</span><span> /path/to/output</span></span></code></pre>\n<p>MFA uses Kaldi internally with:</p>\n<ul>\n<li>Triphone acoustic models (context-dependent phonemes)</li>\n<li>Speaker adaptation via CMVN</li>\n<li>10ms temporal resolution</li>\n</ul>\n<p><strong>CTC Segmentation</strong></p>\n<p>Alternative using neural ASR models:</p>\n<ul>\n<li>Use CTC model to get frame-level posterior probabilities</li>\n<li>Dynamic programming to find best alignment</li>\n<li>Works without pronunciation dictionary</li>\n<li>Available in NeMo toolkit</li>\n</ul>\n<p><strong>Practical Tips</strong></p>\n<ul>\n<li>Chunk audio to 5-10 second segments</li>\n<li>Resample to 16 kHz mono</li>\n<li>MFA still outperforms WhisperX/MMS for alignment (despite their ASR accuracy) [22]</li>\n</ul>\n<h3 id=\"fine-tuning-on-small-datasets\">Fine-Tuning on Small Datasets</h3>\n<p><strong>TTS Adaptation</strong></p>\n<p>For adapting TTS to new speakers with limited data [23]:</p>\n<ol>\n<li>\n<p><strong>Speaker embedding only</strong>: Freeze model, train only speaker embedding</p>\n<ul>\n<li>Works with 10 seconds of audio</li>\n<li>Captures voice timbre, not speaking style</li>\n</ul>\n</li>\n<li>\n<p><strong>Full fine-tuning with mixing</strong>: Mix original speaker data with new speaker</p>\n<ul>\n<li>Equal sampling from both in each batch</li>\n<li>Prevents catastrophic forgetting</li>\n<li>Works with 1-5 minutes of data</li>\n</ul>\n</li>\n<li>\n<p><strong>LoRA/Adapter tuning</strong>: Add small trainable modules to frozen model</p>\n<ul>\n<li>1-2% of original parameters</li>\n<li>Good quality with 5+ minutes of data</li>\n</ul>\n</li>\n</ol>\n<p><strong>Zero-Shot Models</strong></p>\n<p>YourTTS and XTTS can adapt to new voices without fine-tuning:</p>\n<ul>\n<li>Extract speaker embedding from reference audio</li>\n<li>Condition synthesis on that embedding</li>\n<li>Works immediately with 3-10 seconds of reference</li>\n</ul>\n<p><strong>ASR Fine-Tuning</strong></p>\n<p>For domain-specific ASR [24]:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> WhisperForConditionalGeneration</span><span>,</span><span> WhisperProcessor</span></span>\n<span class=\"line\"><span>from</span><span> peft </span><span>import</span><span> LoraConfig</span><span>,</span><span> get_peft_model</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> WhisperForConditionalGeneration</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"openai/whisper-small\"</span><span>)</span></span>\n<span class=\"line\"><span>processor </span><span>=</span><span> WhisperProcessor</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"openai/whisper-small\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># LoRA config for efficient fine-tuning</span></span>\n<span class=\"line\"><span>lora_config </span><span>=</span><span> LoraConfig</span><span>(</span></span>\n<span class=\"line\"><span>    r</span><span>=</span><span>16</span><span>,</span></span>\n<span class=\"line\"><span>    lora_alpha</span><span>=</span><span>32</span><span>,</span></span>\n<span class=\"line\"><span>    target_modules</span><span>=</span><span>[</span><span>\"q_proj\"</span><span>, </span><span>\"v_proj\"</span><span>],</span></span>\n<span class=\"line\"><span>    lora_dropout</span><span>=</span><span>0.05</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> get_peft_model</span><span>(model, lora_config)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Results in ~99% fewer trainable parameters</span></span></code></pre>\n<p>Fine-tuning on ~400 samples can achieve 31+ WER point improvement for domain-specific vocabulary [24].</p>\n<hr />\n<h2 id=\"inference-considerations\">Inference Considerations</h2>\n<p>Deploying audio models requires attention to latency, throughput, and resource constraints.</p>\n<h3 id=\"real-time-factor-and-latency\">Real-Time Factor and Latency</h3>\n<p><strong>Real-Time Factor (RTF)</strong></p>\n<p>RTF = processing_time / audio_duration [25]</p>\n<ul>\n<li>RTF &lt; 1: Faster than real-time (required for streaming)</li>\n<li>RTF = 0.5: Processes 1 second of audio in 0.5 seconds</li>\n<li>RTF = 0.1: 10x faster than real-time</li>\n</ul>\n<p><strong>Latency Components</strong></p>\n<p>For streaming ASR:</p>\n<ul>\n<li><strong>Algorithmic latency</strong>: Time the model needs to “see ahead” (chunk size + right context)</li>\n<li><strong>Compute latency</strong>: Processing time for each chunk</li>\n<li><strong>Network latency</strong>: Round-trip time to server (if cloud-based)</li>\n</ul>\n<p>For TTS:</p>\n<ul>\n<li><strong>First-packet latency</strong>: Time to first audio sample</li>\n<li><strong>Full synthesis latency</strong>: Time to complete waveform</li>\n</ul>\n<p><strong>Benchmarks</strong></p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Task</th><th>RTF (GPU)</th><th>Latency</th></tr></thead><tbody><tr><td>Whisper large-v3</td><td>ASR</td><td>~0.3</td><td>Batch, not streaming</td></tr><tr><td>Conformer-CTC</td><td>Streaming ASR</td><td>~0.1-0.2</td><td>200-400ms</td></tr><tr><td>FastSpeech 2</td><td>TTS</td><td>0.02</td><td>20ms per second of audio</td></tr><tr><td>VITS</td><td>TTS</td><td>0.067</td><td>67ms per second of audio</td></tr><tr><td>XTTS v2</td><td>Zero-shot TTS</td><td>0.48</td><td>200ms first chunk</td></tr><tr><td>HiFi-GAN</td><td>Vocoder</td><td>0.02</td><td>~20ms</td></tr></tbody></table></div>\n<h3 id=\"streaming-inference-for-asr\">Streaming Inference for ASR</h3>\n<p><strong>Chunk-Based Processing</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Pseudocode for streaming ASR</span></span>\n<span class=\"line\"><span>chunk_size </span><span>=</span><span> 640</span><span>  # ms</span></span>\n<span class=\"line\"><span>left_context </span><span>=</span><span> 480</span><span>  # ms</span></span>\n<span class=\"line\"><span>buffer </span><span>=</span><span> []</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>while</span><span> audio_stream</span><span>.</span><span>has_data</span><span>():</span></span>\n<span class=\"line\"><span>    chunk </span><span>=</span><span> audio_stream</span><span>.</span><span>read</span><span>(chunk_size)</span></span>\n<span class=\"line\"><span>    buffer</span><span>.</span><span>append</span><span>(chunk)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Keep only necessary context</span></span>\n<span class=\"line\"><span>    context </span><span>=</span><span> buffer</span><span>[</span><span>-</span><span>left_context_chunks</span><span>:]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Process with context</span></span>\n<span class=\"line\"><span>    features </span><span>=</span><span> extract_features</span><span>(context </span><span>+</span><span> [chunk])</span></span>\n<span class=\"line\"><span>    text </span><span>=</span><span> model</span><span>.</span><span>decode</span><span>(features)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    yield</span><span> text</span></span></code></pre>\n<p><strong>Endpointer</strong></p>\n<p>Detect when user stops speaking to finalize transcription:</p>\n<ul>\n<li>Voice Activity Detection (VAD) for speech/silence</li>\n<li>End-of-query detection for semantic completeness</li>\n<li>Typically 400-800ms of silence triggers endpoint</li>\n</ul>\n<h3 id=\"vocoder-optimization\">Vocoder Optimization</h3>\n<p>Vocoders are often the bottleneck in TTS pipelines.</p>\n<p><strong>Strategies</strong></p>\n<ol>\n<li><strong>Caching</strong>: Cache vocoder output for repeated phrases</li>\n<li><strong>Streaming vocoder</strong>: Generate audio in chunks as mel frames arrive</li>\n<li><strong>Smaller models</strong>: HiFi-GAN V3 (1M params) vs V1 (14M)</li>\n<li><strong>INT8 quantization</strong>: 2-4x speedup with minimal quality loss</li>\n</ol>\n<p><strong>Multi-Band Generation</strong></p>\n<p>Split frequency range into bands, generate each with smaller model:</p>\n<ul>\n<li>Reduces per-band complexity</li>\n<li>Enables parallel generation</li>\n<li>MultiBand-MelGAN achieves very low RTF</li>\n</ul>\n<h3 id=\"quantization-for-audio-models\">Quantization for Audio Models</h3>\n<p><strong>Post-Training Quantization (PTQ)</strong></p>\n<p>Apply quantization after training [26]:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># INT8 quantization for faster inference</span></span>\n<span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model_fp32 </span><span>=</span><span> load_model</span><span>()</span></span>\n<span class=\"line\"><span>model_int8 </span><span>=</span><span> torch</span><span>.</span><span>quantization</span><span>.</span><span>quantize_dynamic</span><span>(</span></span>\n<span class=\"line\"><span>    model_fp32,</span></span>\n<span class=\"line\"><span>    {torch.nn.Linear, torch.nn.Conv1d},</span></span>\n<span class=\"line\"><span>    dtype</span><span>=</span><span>torch.qint8</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p><strong>Quantization-Aware Training (QAT)</strong></p>\n<p>Train with simulated quantization for better accuracy:</p>\n<ul>\n<li>Fake quantization during forward pass</li>\n<li>Full precision gradients during backward pass</li>\n<li>Recovers most accuracy loss from PTQ</li>\n</ul>\n<p><strong>Results</strong></p>\n<ul>\n<li>Whisper: INT8 gives ~2x speedup on CPU with &lt;1% WER increase</li>\n<li>HiFi-GAN: INT8 gives ~2x speedup, some high-frequency quality loss</li>\n<li>VITS: INT4 on decoder achieves 40% latency reduction</li>\n</ul>\n<p><strong>Framework Support</strong></p>\n<ul>\n<li>ONNX Runtime: Broad model support, CPU/GPU</li>\n<li>TensorRT: NVIDIA GPUs, aggressive optimization</li>\n<li>OpenVINO: Intel CPUs/GPUs</li>\n<li>Core ML: Apple Silicon</li>\n</ul>\n<hr />\n<h2 id=\"practical-examples\">Practical Examples</h2>\n<h3 id=\"running-whisper-locally\">Running Whisper Locally</h3>\n<p><strong>Installation</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>pip</span><span> install</span><span> openai-whisper</span></span>\n<span class=\"line\"><span># Or with faster-whisper (CTranslate2 backend)</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> faster-whisper</span></span></code></pre>\n<p><strong>Basic Transcription</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> whisper</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> whisper</span><span>.</span><span>load_model</span><span>(</span><span>\"base\"</span><span>)</span><span>  # tiny, base, small, medium, large-v3</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> model</span><span>.</span><span>transcribe</span><span>(</span><span>\"audio.mp3\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(result[</span><span>\"text\"</span><span>])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># With options</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> model</span><span>.</span><span>transcribe</span><span>(</span></span>\n<span class=\"line\"><span>    \"audio.mp3\"</span><span>,</span></span>\n<span class=\"line\"><span>    language</span><span>=</span><span>\"en\"</span><span>,</span></span>\n<span class=\"line\"><span>    task</span><span>=</span><span>\"transcribe\"</span><span>,  </span><span># or \"translate\" for X-&gt;English</span></span>\n<span class=\"line\"><span>    fp16</span><span>=</span><span>True</span><span>,  </span><span># Use FP16 on GPU</span></span>\n<span class=\"line\"><span>    condition_on_previous_text</span><span>=</span><span>True</span><span>,  </span><span># Use context</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p><strong>Faster-Whisper for Production</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> faster_whisper </span><span>import</span><span> WhisperModel</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># INT8 quantization for faster CPU inference</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> WhisperModel</span><span>(</span><span>\"large-v3\"</span><span>, device</span><span>=</span><span>\"cpu\"</span><span>, compute_type</span><span>=</span><span>\"int8\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or FP16 on GPU</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> WhisperModel</span><span>(</span><span>\"large-v3\"</span><span>, device</span><span>=</span><span>\"cuda\"</span><span>, compute_type</span><span>=</span><span>\"float16\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>segments</span><span>,</span><span> info </span><span>=</span><span> model</span><span>.</span><span>transcribe</span><span>(</span><span>\"audio.mp3\"</span><span>, beam_size</span><span>=</span><span>5</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> segment </span><span>in</span><span> segments</span><span>:</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"[</span><span>{</span><span>segment.start</span><span>:.2f</span><span>}</span><span>s -&gt; </span><span>{</span><span>segment.end</span><span>:.2f</span><span>}</span><span>s] </span><span>{</span><span>segment.text</span><span>}</span><span>\"</span><span>)</span></span></code></pre>\n<p><strong>Streaming with Whisper</strong></p>\n<p>Whisper isn’t designed for streaming, but you can approximate it:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> numpy </span><span>as</span><span> np</span></span>\n<span class=\"line\"><span>from</span><span> faster_whisper </span><span>import</span><span> WhisperModel</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> WhisperModel</span><span>(</span><span>\"base\"</span><span>, device</span><span>=</span><span>\"cuda\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> process_chunk</span><span>(</span><span>audio_chunk</span><span>,</span><span> previous_text</span><span>=</span><span>\"\"</span><span>):</span></span>\n<span class=\"line\"><span>    # Pad to 30 seconds if needed</span></span>\n<span class=\"line\"><span>    if</span><span> len</span><span>(audio_chunk)</span><span> &lt;</span><span> 30</span><span> *</span><span> 16000</span><span>:</span></span>\n<span class=\"line\"><span>        audio_chunk </span><span>=</span><span> np</span><span>.</span><span>pad</span><span>(audio_chunk, (</span><span>0</span><span>, </span><span>30</span><span> *</span><span> 16000</span><span> -</span><span> len</span><span>(audio_chunk)))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    segments</span><span>,</span><span> _ </span><span>=</span><span> model</span><span>.</span><span>transcribe</span><span>(</span></span>\n<span class=\"line\"><span>        audio_chunk,</span></span>\n<span class=\"line\"><span>        initial_prompt</span><span>=</span><span>previous_text,  </span><span># Context from previous chunks</span></span>\n<span class=\"line\"><span>        vad_filter</span><span>=</span><span>True</span><span>,  </span><span># Filter silence</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"><span>    return</span><span> \" \"</span><span>.</span><span>join</span><span>([s.text </span><span>for</span><span> s </span><span>in</span><span> segments])</span></span></code></pre>\n<h3 id=\"fine-tuning-a-tts-model\">Fine-Tuning a TTS Model</h3>\n<p><strong>Fine-tuning Coqui TTS</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Install</span></span>\n<span class=\"line\"><span># pip install TTS</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>from</span><span> TTS</span><span>.</span><span>api </span><span>import</span><span> TTS</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load pre-trained VITS model</span></span>\n<span class=\"line\"><span>tts </span><span>=</span><span> TTS</span><span>(</span><span>\"tts_models/en/ljspeech/vits\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># For fine-tuning, use the training script</span></span>\n<span class=\"line\"><span># 1. Prepare data in LJSpeech format:</span></span>\n<span class=\"line\"><span>#    wavs/</span></span>\n<span class=\"line\"><span>#      audio1.wav</span></span>\n<span class=\"line\"><span>#      audio2.wav</span></span>\n<span class=\"line\"><span>#    metadata.csv: audio1|transcript one|transcript one</span></span>\n<span class=\"line\"><span>#                  audio2|transcript two|transcript two</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># 2. Create config</span></span>\n<span class=\"line\"><span>from</span><span> TTS</span><span>.</span><span>tts</span><span>.</span><span>configs</span><span>.</span><span>vits_config </span><span>import</span><span> VitsConfig</span></span>\n<span class=\"line\"><span>from</span><span> TTS</span><span>.</span><span>tts</span><span>.</span><span>models</span><span>.</span><span>vits </span><span>import</span><span> Vits</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>config </span><span>=</span><span> VitsConfig</span><span>(</span></span>\n<span class=\"line\"><span>    audio</span><span>=</span><span>{</span><span>\"sample_rate\"</span><span>: </span><span>22050</span><span>},</span></span>\n<span class=\"line\"><span>    run_name</span><span>=</span><span>\"my_voice\"</span><span>,</span></span>\n<span class=\"line\"><span>    batch_size</span><span>=</span><span>16</span><span>,</span></span>\n<span class=\"line\"><span>    eval_batch_size</span><span>=</span><span>8</span><span>,</span></span>\n<span class=\"line\"><span>    num_loader_workers</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>    num_eval_loader_workers</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>    run_eval</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    test_delay_epochs</span><span>=-</span><span>1</span><span>,</span></span>\n<span class=\"line\"><span>    epochs</span><span>=</span><span>1000</span><span>,</span></span>\n<span class=\"line\"><span>    text_cleaner</span><span>=</span><span>\"english_cleaners\"</span><span>,</span></span>\n<span class=\"line\"><span>    use_phonemes</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    phoneme_language</span><span>=</span><span>\"en-us\"</span><span>,</span></span>\n<span class=\"line\"><span>    output_path</span><span>=</span><span>\"output/\"</span><span>,</span></span>\n<span class=\"line\"><span>    datasets</span><span>=</span><span>[{</span></span>\n<span class=\"line\"><span>        \"name\"</span><span>: </span><span>\"ljspeech\"</span><span>,</span></span>\n<span class=\"line\"><span>        \"path\"</span><span>: </span><span>\"/path/to/your/data/\"</span><span>,</span></span>\n<span class=\"line\"><span>        \"meta_file_train\"</span><span>: </span><span>\"metadata.csv\"</span></span>\n<span class=\"line\"><span>    }],</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># 3. Start from pretrained checkpoint</span></span>\n<span class=\"line\"><span>config</span><span>.</span><span>load_json</span><span>(</span><span>\"/path/to/pretrained/config.json\"</span><span>)</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> Vits</span><span>.</span><span>init_from_config</span><span>(config)</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>load_checkpoint</span><span>(config, </span><span>\"/path/to/pretrained/best_model.pth\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># 4. Fine-tune</span></span>\n<span class=\"line\"><span>from</span><span> TTS</span><span>.</span><span>trainer </span><span>import</span><span> Trainer</span></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> Trainer</span><span>(</span></span>\n<span class=\"line\"><span>    TrainerArgs</span><span>(),</span></span>\n<span class=\"line\"><span>    config,</span></span>\n<span class=\"line\"><span>    output_path</span><span>=</span><span>\"output/\"</span><span>,</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>model,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>fit</span><span>()</span></span></code></pre>\n<p><strong>Fine-tuning with XTTS</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> TTS</span><span>.</span><span>api </span><span>import</span><span> TTS</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># XTTS supports zero-shot cloning without fine-tuning</span></span>\n<span class=\"line\"><span>tts </span><span>=</span><span> TTS</span><span>(</span><span>\"tts_models/multilingual/multi-dataset/xtts_v2\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate with voice cloning</span></span>\n<span class=\"line\"><span>tts</span><span>.</span><span>tts_to_file</span><span>(</span></span>\n<span class=\"line\"><span>    text</span><span>=</span><span>\"Hello, this is my cloned voice!\"</span><span>,</span></span>\n<span class=\"line\"><span>    speaker_wav</span><span>=</span><span>\"reference_audio.wav\"</span><span>,  </span><span># 3-10 seconds</span></span>\n<span class=\"line\"><span>    language</span><span>=</span><span>\"en\"</span><span>,</span></span>\n<span class=\"line\"><span>    file_path</span><span>=</span><span>\"output.wav\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># For fine-tuning (better quality with more data)</span></span>\n<span class=\"line\"><span># Use the XTTS fine-tuning script with your dataset</span></span></code></pre>\n<h3 id=\"voice-conversion-pipeline\">Voice Conversion Pipeline</h3>\n<p><strong>Using RVC</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Clone RVC</span></span>\n<span class=\"line\"><span>git</span><span> clone</span><span> https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI</span></span>\n<span class=\"line\"><span>cd</span><span> Retrieval-based-Voice-Conversion-WebUI</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> -r</span><span> requirements.txt</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Download pretrained models</span></span>\n<span class=\"line\"><span># https://huggingface.co/lj1995/VoiceConversionWebUI</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Launch WebUI</span></span>\n<span class=\"line\"><span>python</span><span> infer-web.py</span></span></code></pre>\n<p><strong>Programmatic Voice Conversion</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Simplified RVC inference (actual implementation more complex)</span></span>\n<span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>from</span><span> fairseq </span><span>import</span><span> checkpoint_utils</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load HuBERT for content extraction</span></span>\n<span class=\"line\"><span>hubert_model</span><span>,</span><span> cfg</span><span>,</span><span> task </span><span>=</span><span> checkpoint_utils</span><span>.</span><span>load_model_ensemble_and_task</span><span>(</span></span>\n<span class=\"line\"><span>    [</span><span>\"hubert_base.pt\"</span><span>]</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>hubert </span><span>=</span><span> hubert_model</span><span>[</span><span>0</span><span>].</span><span>eval</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load RVC model</span></span>\n<span class=\"line\"><span>rvc_model </span><span>=</span><span> torch</span><span>.</span><span>load</span><span>(</span><span>\"target_voice.pth\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> convert_voice</span><span>(</span><span>source_audio</span><span>):</span></span>\n<span class=\"line\"><span>    # 1. Extract content features</span></span>\n<span class=\"line\"><span>    with</span><span> torch</span><span>.</span><span>no_grad</span><span>():</span></span>\n<span class=\"line\"><span>        content </span><span>=</span><span> hubert</span><span>.</span><span>extract_features</span><span>(source_audio)</span><span>[</span><span>0</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # 2. Extract pitch</span></span>\n<span class=\"line\"><span>    f0 </span><span>=</span><span> extract_f0</span><span>(source_audio)</span><span>  # RMVPE or CREPE</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # 3. Retrieve similar features from training set (optional)</span></span>\n<span class=\"line\"><span>    retrieved </span><span>=</span><span> faiss_index</span><span>.</span><span>search</span><span>(content, k</span><span>=</span><span>3</span><span>)</span></span>\n<span class=\"line\"><span>    content </span><span>=</span><span> blend</span><span>(content, retrieved)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # 4. Generate with RVC model</span></span>\n<span class=\"line\"><span>    audio </span><span>=</span><span> rvc_model</span><span>(content, f0)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> audio</span></span></code></pre>\n<p><strong>So-VITS-SVC for Singing</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Clone So-VITS-SVC</span></span>\n<span class=\"line\"><span>git</span><span> clone</span><span> https://github.com/svc-develop-team/so-vits-svc</span></span>\n<span class=\"line\"><span>cd</span><span> so-vits-svc</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> -r</span><span> requirements.txt</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Prepare training data</span></span>\n<span class=\"line\"><span># Place wav files in dataset_raw/speaker_name/</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Preprocess</span></span>\n<span class=\"line\"><span>python</span><span> resample.py</span></span>\n<span class=\"line\"><span>python</span><span> preprocess_flist_config.py</span></span>\n<span class=\"line\"><span>python</span><span> preprocess_hubert_f0.py</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Train</span></span>\n<span class=\"line\"><span>python</span><span> train.py</span><span> -c</span><span> configs/config.json</span><span> -m</span><span> speaker_name</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Inference</span></span>\n<span class=\"line\"><span>python</span><span> inference_main.py</span><span> -m</span><span> \"logs/speaker_name/model.pth\"</span><span> \\</span></span>\n<span class=\"line\"><span>    -c</span><span> \"configs/config.json\"</span><span> \\</span></span>\n<span class=\"line\"><span>    -n</span><span> \"song.wav\"</span><span> \\</span></span>\n<span class=\"line\"><span>    -t</span><span> 0</span><span>  # pitch shift in semitones</span></span></code></pre>\n<hr />\n<h2 id=\"references\">References</h2>\n<ol>\n<li>\n<p>Radford, A., et al. (2022). “Robust Speech Recognition via Large-Scale Weak Supervision.” OpenAI. <a href=\"https://github.com/openai/whisper\">https://github.com/openai/whisper</a></p>\n</li>\n<li>\n<p>OpenAI. (2023). “Whisper large-v3.” Hugging Face. <a href=\"https://huggingface.co/openai/whisper-large-v3\">https://huggingface.co/openai/whisper-large-v3</a></p>\n</li>\n<li>\n<p>Zhou, et al. (2025). “Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding.” Interspeech 2025. <a href=\"https://www.isca-archive.org/interspeech_2025/zhou25_interspeech.pdf\">https://www.isca-archive.org/interspeech_2025/zhou25_interspeech.pdf</a></p>\n</li>\n<li>\n<p>Gulati, A., et al. (2020). “Conformer: Convolution-augmented Transformer for Speech Recognition.” Interspeech 2020. <a href=\"https://arxiv.org/abs/2005.08100\">https://arxiv.org/abs/2005.08100</a></p>\n</li>\n<li>\n<p>Hugging Face. “CTC Architectures.” Audio Course. <a href=\"https://huggingface.co/learn/audio-course/chapter3/ctc\">https://huggingface.co/learn/audio-course/chapter3/ctc</a></p>\n</li>\n<li>\n<p>SpeechBrain Documentation. “Streaming Speech Recognition with Conformers.” <a href=\"https://speechbrain.readthedocs.io/en/v1.0.2/tutorials/nn/conformer-streaming-asr.html\">https://speechbrain.readthedocs.io/en/v1.0.2/tutorials/nn/conformer-streaming-asr.html</a></p>\n</li>\n<li>\n<p>Kim, J., et al. (2021). “Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech.” ICML 2021. <a href=\"https://arxiv.org/abs/2106.06103\">https://arxiv.org/abs/2106.06103</a></p>\n</li>\n<li>\n<p>Shen, J., et al. (2017). “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions.” <a href=\"https://arxiv.org/abs/1712.05884\">https://arxiv.org/abs/1712.05884</a></p>\n</li>\n<li>\n<p>Kong, J., et al. (2020). “HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.” NeurIPS 2020. <a href=\"https://github.com/jik876/hifi-gan\">https://github.com/jik876/hifi-gan</a></p>\n</li>\n<li>\n<p>Lee, S.-G., et al. (2022). “BigVGAN: A Universal Neural Vocoder with Large-Scale Training.” ICLR 2023. <a href=\"https://arxiv.org/abs/2206.04658\">https://arxiv.org/abs/2206.04658</a></p>\n</li>\n<li>\n<p>Microsoft Research. “VALL-E.” <a href=\"https://www.microsoft.com/en-us/research/project/vall-e-x/\">https://www.microsoft.com/en-us/research/project/vall-e-x/</a></p>\n</li>\n<li>\n<p>Casanova, E., et al. (2024). “XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.” <a href=\"https://arxiv.org/html/2406.04904v1\">https://arxiv.org/html/2406.04904v1</a></p>\n</li>\n<li>\n<p>Voice Cloning Survey. (2025). <a href=\"https://arxiv.org/html/2505.00579v1\">https://arxiv.org/html/2505.00579v1</a></p>\n</li>\n<li>\n<p>RVC Project. “Retrieval-based Voice Conversion WebUI.” <a href=\"https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI\">https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI</a></p>\n</li>\n<li>\n<p>Qian, K., et al. (2022). “CONTENTVEC: An Improved Self-Supervised Speech Representation.” ICML 2022. <a href=\"https://proceedings.mlr.press/v162/qian22b/qian22b.pdf\">https://proceedings.mlr.press/v162/qian22b/qian22b.pdf</a></p>\n</li>\n<li>\n<p>So-VITS-SVC. “SoftVC VITS Singing Voice Conversion.” <a href=\"https://github.com/svc-develop-team/so-vits-svc\">https://github.com/svc-develop-team/so-vits-svc</a></p>\n</li>\n<li>\n<p>Borsos, Z., et al. (2022). “AudioLM: a Language Modeling Approach to Audio Generation.” <a href=\"https://arxiv.org/abs/2209.03143\">https://arxiv.org/abs/2209.03143</a></p>\n</li>\n<li>\n<p>Zhang, D., et al. (2023). “SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.” <a href=\"https://www.alphaxiv.org/overview/2305.11000v2\">https://www.alphaxiv.org/overview/2305.11000v2</a></p>\n</li>\n<li>\n<p>Défossez, A., et al. (2024). “Moshi: a speech-text foundation model for real-time dialogue.” Kyutai. <a href=\"https://kyutai.org/Moshi.pdf\">https://kyutai.org/Moshi.pdf</a></p>\n</li>\n<li>\n<p>Wu, et al. (2025). “MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition.” NeurIPS 2025. <a href=\"https://arxiv.org/abs/2510.04136\">https://arxiv.org/abs/2510.04136</a></p>\n</li>\n<li>\n<p>Ketanhdoshi. “Audio Deep Learning Made Simple: Data Preparation and Augmentation.” <a href=\"https://ketanhdoshi.github.io/Audio-Augment/\">https://ketanhdoshi.github.io/Audio-Augment/</a></p>\n</li>\n<li>\n<p>McAuliffe, M., et al. (2017). “Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi.” <a href=\"https://montreal-forced-aligner.readthedocs.io/\">https://montreal-forced-aligner.readthedocs.io/</a></p>\n</li>\n<li>\n<p>Arik, S.O., et al. (2018). “Neural Voice Cloning with a Few Samples.” NeurIPS 2018.</p>\n</li>\n<li>\n<p>Hugging Face. “Fine-Tune Whisper For Multilingual ASR.” <a href=\"https://huggingface.co/blog/fine-tune-whisper\">https://huggingface.co/blog/fine-tune-whisper</a></p>\n</li>\n<li>\n<p>Open Voice Technology Wiki. “Real-time-factor.” <a href=\"https://openvoice-tech.net/index.php/Real-time-factor\">https://openvoice-tech.net/index.php/Real-time-factor</a></p>\n</li>\n<li>\n<p>“Model Quantization Techniques for ASR/TTS.” <a href=\"https://apxml.com/courses/speech-recognition-synthesis-asr-tts/chapter-6-optimization-deployment-toolkits/quantization-speech-models\">https://apxml.com/courses/speech-recognition-synthesis-asr-tts/chapter-6-optimization-deployment-toolkits/quantization-speech-models</a></p>\n</li>\n</ol>",
            "url": "https://blog.ecitis.org/voice-audio-models/",
            "title": "Voice and Audio AI Models: Architecture, Training, and Deployment",
            "summary": "Survey the architectures behind speech recognition, synthesis, voice conversion, and full-duplex conversational audio.",
            "image": "https://blog.ecitis.org/open-graph/voice-audio-models.png",
            "date_modified": "2026-01-26T00:00:00.000Z",
            "date_published": "2026-01-26T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Audio AI",
                "Speech",
                "TTS"
            ]
        },
        {
            "id": "https://blog.ecitis.org/huggingface-ecosystem-guide/",
            "content_html": "<h2 id=\"what-is-hugging-face\">What is Hugging Face</h2>\n<p>Hugging Face is an American company founded in 2016 by Clement Delangue and Julien Chaumond, originally as a chatbot startup. It has since evolved into the central hub for open-source machine learning. The platform hosts over 2.4 million models and 730,000 datasets as of January 2026, serving more than 18 million monthly visitors.</p>\n<p>The company builds on a core open-source stack: Transformers, Datasets, Diffusers, PEFT, Accelerate, TRL, Timm, and Optimum. Beyond libraries, Hugging Face provides production tools like Inference Endpoints, Text Generation Inference (TGI), and Text Embeddings Inference (TEI). Over 2,000 organizations use their Enterprise Hub for private deployments with SSO, regional data storage, and audit logs.</p>\n<p>Hugging Face integrates with major cloud providers. AWS offers Deep Learning Containers and SageMaker integration. Azure Machine Learning and Google Vertex AI both support direct model deployment. In April 2025, the company acquired Pollen Robotics, a humanoid robotics startup, signaling expansion beyond pure software.</p>\n<hr />\n<h2 id=\"the-models-hub\">The Models Hub</h2>\n<p>The Hub is where you discover, evaluate, and download models. Each model repository contains the weights, configuration files, and documentation needed to run inference or continue training.</p>\n<h3 id=\"discovering-models\">Discovering Models</h3>\n<p>Filter by task (text generation, image classification, audio transcription), library (Transformers, Diffusers, GGUF), license, and language. The search indexes model names, tags, and README content. Sorting by downloads or likes surfaces popular choices; sorting by recent shows the latest uploads.</p>\n<h3 id=\"model-cards\">Model Cards</h3>\n<p>Every model should have a model card in the README.md file. Cards follow a structured template covering:</p>\n<ul>\n<li><strong>Model description</strong>: Architecture, training data, intended use</li>\n<li><strong>Limitations</strong>: Known failure modes, biases, out-of-scope uses</li>\n<li><strong>Training procedure</strong>: Hyperparameters, hardware, preprocessing</li>\n<li><strong>Evaluation results</strong>: Benchmarks, metrics, disaggregated performance</li>\n</ul>\n<p>The metadata block at the top enables Hub features. Specify <code>license</code> for filtering, <code>datasets</code> to link training data, and <code>metrics</code> with <code>eval_results</code> to display benchmark performance in the UI. The Hub parses this YAML and renders widgets showing accuracy, F1, perplexity, or custom metrics.</p>\n<h3 id=\"licenses\">Licenses</h3>\n<p>Common licenses on the Hub include Apache 2.0, MIT, and various model-specific licenses like Llama’s community license or Gemma’s terms of use. The <code>license</code> field in metadata enables filtering. For custom licenses, use <code>license: other</code> with <code>license_name</code> and <code>license_link</code> fields, or include a LICENSE file in the repository.</p>\n<h3 id=\"evaluation-with-the-evaluate-library\">Evaluation with the Evaluate Library</h3>\n<p>The Evaluate library provides dozens of metrics with a consistent API:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> evaluate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>accuracy </span><span>=</span><span> evaluate</span><span>.</span><span>load</span><span>(</span><span>\"accuracy\"</span><span>)</span></span>\n<span class=\"line\"><span>results </span><span>=</span><span> accuracy</span><span>.</span><span>compute</span><span>(predictions</span><span>=</span><span>[</span><span>0</span><span>, </span><span>1</span><span>, </span><span>1</span><span>], references</span><span>=</span><span>[</span><span>0</span><span>, </span><span>1</span><span>, </span><span>0</span><span>])</span></span>\n<span class=\"line\"><span>print</span><span>(results)</span><span>  # {'accuracy': 0.666...}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load multiple metrics</span></span>\n<span class=\"line\"><span>clf_metrics </span><span>=</span><span> evaluate</span><span>.</span><span>combine</span><span>([</span><span>\"accuracy\"</span><span>, </span><span>\"f1\"</span><span>, </span><span>\"precision\"</span><span>, </span><span>\"recall\"</span><span>])</span></span>\n<span class=\"line\"><span>results </span><span>=</span><span> clf_metrics</span><span>.</span><span>compute</span><span>(predictions</span><span>=</span><span>[</span><span>0</span><span>, </span><span>1</span><span>, </span><span>1</span><span>, </span><span>0</span><span>], references</span><span>=</span><span>[</span><span>0</span><span>, </span><span>1</span><span>, </span><span>0</span><span>, </span><span>1</span><span>])</span></span></code></pre>\n<p>Each metric includes a card documenting its formula, value ranges, and appropriate use cases.</p>\n<hr />\n<h2 id=\"model-formats-and-download-types\">Model Formats and Download Types</h2>\n<p>Models come in various formats optimized for different runtimes and hardware. Understanding when to use each saves time and compute.</p>\n<figure><figcaption><strong>Model Format Comparison</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/huggingface-ecosystem-guide-0.webp\" width=\"512\" height=\"439\" alt=\"Model Format Comparison\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"pytorch-weights-safetensors-vs-bin\">PyTorch Weights: safetensors vs .bin</h3>\n<p>The <code>.bin</code> format uses Python’s pickle for serialization. Pickle can execute arbitrary code during deserialization, creating security risks when loading untrusted models.</p>\n<p>Safetensors solves this. Developed by Hugging Face, it stores tensors in a binary format without code execution. Loading is 76x faster on CPU and 2x faster on GPU compared to pickle. The BLOOM model loads on 8 GPUs in 45 seconds with safetensors versus 10 minutes with pickle weights.</p>\n<p>Safetensors also supports lazy loading through its <code>safe_open()</code> context manager. Tensors load on-demand rather than all at once, critical for distributed inference where each GPU only needs a shard.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> safetensors </span><span>import</span><span> safe_open</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Lazy loading - tensors loaded only when accessed</span></span>\n<span class=\"line\"><span>with</span><span> safe_open</span><span>(</span><span>\"model.safetensors\"</span><span>, framework</span><span>=</span><span>\"pt\"</span><span>)</span><span> as</span><span> f</span><span>:</span></span>\n<span class=\"line\"><span>    tensor_a </span><span>=</span><span> f</span><span>.</span><span>get_tensor</span><span>(</span><span>\"layer.0.weight\"</span><span>)</span></span></code></pre>\n<p>Transformers defaults to safetensors when available. Legacy <code>.bin</code> files work but carry security and performance penalties.</p>\n<h3 id=\"gguf-for-llamacpp-and-ollama\">GGUF for llama.cpp and Ollama</h3>\n<p>GGUF (Georgi Gerganov Universal Format) packages everything needed to run a model: architecture details, tokenizer, quantization parameters, and weights in a single file. It evolved from the earlier GGML format for better extensibility.</p>\n<p>The format targets llama.cpp and tools built on it like Ollama, LM Studio, and koboldcpp. GGUF excels at CPU inference and mixed CPU/GPU offloading on consumer hardware.</p>\n<p>Quantization options range from 8-bit down to 1.5-bit:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Type</th><th>Description</th><th>Size Reduction</th></tr></thead><tbody><tr><td>Q8_0</td><td>8-bit, block size 32</td><td>~2x</td></tr><tr><td>Q6_K</td><td>6-bit K-Quant</td><td>~2.5x</td></tr><tr><td>Q5_K_M</td><td>5-bit K-Quant medium</td><td>~2.8x</td></tr><tr><td>Q4_K_M</td><td>4-bit K-Quant medium</td><td>~3.3x</td></tr><tr><td>Q4_0</td><td>4-bit basic</td><td>~4x</td></tr><tr><td>Q2_K</td><td>2-bit K-Quant</td><td>~6x</td></tr></tbody></table></div>\n<p>Q4_K_M hits the sweet spot for most users. A 13.5GB FP16 model shrinks to 4GB while maintaining acceptable quality. K-Quant variants use different bit allocations per layer based on importance, preserving quality better than uniform quantization.</p>\n<p>Convert HuggingFace models to GGUF:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Clone llama.cpp</span></span>\n<span class=\"line\"><span>git</span><span> clone</span><span> https://github.com/ggml-org/llama.cpp</span></span>\n<span class=\"line\"><span>cd</span><span> llama.cpp</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Convert to GGUF (F16 baseline)</span></span>\n<span class=\"line\"><span>python</span><span> convert_hf_to_gguf.py</span><span> /path/to/hf-model</span><span> --outfile</span><span> model-f16.gguf</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Quantize to Q4_K_M</span></span>\n<span class=\"line\"><span>./llama-quantize</span><span> model-f16.gguf</span><span> model-q4_k_m.gguf</span><span> Q4_K_M</span></span></code></pre>\n<p>For better quantization quality, compute an importance matrix from a calibration dataset:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Generate importance matrix</span></span>\n<span class=\"line\"><span>./llama-imatrix</span><span> -m</span><span> model-f16.gguf</span><span> -f</span><span> calibration.txt</span><span> -o</span><span> imatrix.dat</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Quantize with importance matrix</span></span>\n<span class=\"line\"><span>./llama-quantize</span><span> --imatrix</span><span> imatrix.dat</span><span> model-f16.gguf</span><span> model-q4_k_m.gguf</span><span> Q4_K_M</span></span></code></pre>\n<h3 id=\"onnx-for-cross-platform-deployment\">ONNX for Cross-Platform Deployment</h3>\n<p>ONNX (Open Neural Network Exchange) defines a standard graph format for neural networks. Train in PyTorch or TensorFlow, export to ONNX, run anywhere.</p>\n<p>ONNX Runtime powers inference across platforms: NVIDIA GPUs via TensorRT, Intel CPUs via OpenVINO, AMD GPUs via ROCm, Apple Silicon via CoreML, and Windows via DirectML. This portability makes ONNX valuable for production deployment across heterogeneous hardware.</p>\n<p>Export a Transformers model to ONNX:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> optimum</span><span>.</span><span>onnxruntime </span><span>import</span><span> ORTModelForSequenceClassification</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoTokenizer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model_id </span><span>=</span><span> \"distilbert-base-uncased-finetuned-sst-2-english\"</span></span>\n<span class=\"line\"><span>save_dir </span><span>=</span><span> \"onnx_model\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Export and save</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> ORTModelForSequenceClassification</span><span>.</span><span>from_pretrained</span><span>(model_id, export</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(model_id)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model</span><span>.</span><span>save_pretrained</span><span>(save_dir)</span></span>\n<span class=\"line\"><span>tokenizer</span><span>.</span><span>save_pretrained</span><span>(save_dir)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load and run inference</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> ORTModelForSequenceClassification</span><span>.</span><span>from_pretrained</span><span>(save_dir)</span></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> tokenizer</span><span>(</span><span>\"This movie was great!\"</span><span>, return_tensors</span><span>=</span><span>\"pt\"</span><span>)</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>(</span><span>**</span><span>inputs)</span></span></code></pre>\n<p>The ONNX Model Zoo on Hugging Face hosts pre-exported models at huggingface.co/onnxmodelzoo.</p>\n<h3 id=\"awq-gptq-and-bitsandbytes-quantization\">AWQ, GPTQ, and bitsandbytes Quantization</h3>\n<p>These methods reduce model precision for faster inference and lower memory usage, each with different tradeoffs.</p>\n<p><strong>GPTQ</strong> uses post-training quantization with Hessian-based optimization to minimize output error. It requires a calibration dataset and pre-quantizes weights. Best for GPU inference when you need raw throughput.</p>\n<p><strong>AWQ</strong> (Activation-aware Weight Quantization) identifies and protects important weights by observing activations. Like GPTQ, it requires calibration and pre-quantization. AWQ often preserves quality better than GPTQ, especially for instruction-tuned models.</p>\n<p><strong>bitsandbytes</strong> quantizes on-the-fly during model loading. No calibration dataset or pre-quantized checkpoint needed. Its NF4 (NormalFloat4) format assumes weights follow a Gaussian distribution and places more quantization levels near zero where weights cluster. Best for fine-tuning with QLoRA and quick experimentation.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> BitsAndBytesConfig</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load with 4-bit quantization</span></span>\n<span class=\"line\"><span>bnb_config </span><span>=</span><span> BitsAndBytesConfig</span><span>(</span></span>\n<span class=\"line\"><span>    load_in_4bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_quant_type</span><span>=</span><span>\"nf4\"</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_compute_dtype</span><span>=</span><span>\"bfloat16\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"meta-llama/Llama-3.1-8B\"</span><span>,</span></span>\n<span class=\"line\"><span>    quantization_config</span><span>=</span><span>bnb_config,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>Quality comparison from 2025 benchmarks: bitsandbytes has the smallest quality drop, AWQ comes second, GPTQ and Marlin (GPTQ’s optimized kernel) are close behind. The practical difference matters most for your specific use case; test on your evaluation set.</p>\n<h3 id=\"tensorrt-llm-for-maximum-nvidia-performance\">TensorRT-LLM for Maximum NVIDIA Performance</h3>\n<p>TensorRT-LLM is NVIDIA’s library for optimizing LLM inference on their GPUs. It applies custom attention kernels, inflight batching, paged KV caching, and hardware-specific optimizations.</p>\n<p>The library supports quantization formats tied to GPU generations:</p>\n<ul>\n<li><strong>Blackwell (B200, GB200)</strong>: FP4 with NVFP4 format</li>\n<li><strong>Ada Lovelace (L40, RTX 40 series)</strong>: FP8</li>\n<li><strong>Ampere (A100, RTX 30 series)</strong>: INT8 SmoothQuant, INT4 AWQ</li>\n</ul>\n<p>Recent benchmarks show DeepSeek-R1 running 3x faster with speculative decoding on TensorRT-LLM. Pre-quantized NVFP4 models for Llama 3.3 70B and DeepSeek-R1 are available on the Hub.</p>\n<p>TensorRT-LLM requires building an engine specific to your GPU and model. This compilation step takes time but produces optimized inference:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> tensorrt_llm </span><span>import</span><span> LLM</span><span>,</span><span> SamplingParams</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>prompts </span><span>=</span><span> [</span><span>\"Explain quantum computing in simple terms\"</span><span>]</span></span>\n<span class=\"line\"><span>sampling_params </span><span>=</span><span> SamplingParams</span><span>(temperature</span><span>=</span><span>0.7</span><span>, max_tokens</span><span>=</span><span>256</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> llm</span><span>.</span><span>generate</span><span>(prompts, sampling_params)</span></span></code></pre>\n<h3 id=\"choosing-the-right-format\">Choosing the Right Format</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Use Case</th><th>Format</th></tr></thead><tbody><tr><td>Local CPU/GPU inference, consumer hardware</td><td>GGUF</td></tr><tr><td>Production GPU serving, max throughput</td><td>TensorRT-LLM</td></tr><tr><td>Cross-platform deployment</td><td>ONNX</td></tr><tr><td>Fine-tuning with limited VRAM</td><td>bitsandbytes</td></tr><tr><td>Pre-quantized GPU inference</td><td>AWQ or GPTQ</td></tr><tr><td>Standard Transformers usage</td><td>safetensors</td></tr></tbody></table></div>\n<hr />\n<h2 id=\"datasets-on-the-hub\">Datasets on the Hub</h2>\n<p>The Hub hosts over 730,000 datasets with tools for exploration, streaming, and efficient loading.</p>\n<h3 id=\"the-datasets-library\">The Datasets Library</h3>\n<p>Load datasets with a single function call:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load from the Hub</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"imdb\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Access splits</span></span>\n<span class=\"line\"><span>train </span><span>=</span><span> dataset</span><span>[</span><span>\"train\"</span><span>]</span></span>\n<span class=\"line\"><span>test </span><span>=</span><span> dataset</span><span>[</span><span>\"test\"</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Iterate</span></span>\n<span class=\"line\"><span>for</span><span> example </span><span>in</span><span> train</span><span>:</span></span>\n<span class=\"line\"><span>    print</span><span>(example[</span><span>\"text\"</span><span>], example[</span><span>\"label\"</span><span>])</span></span>\n<span class=\"line\"><span>    break</span></span></code></pre>\n<p>The library handles downloading, caching, and format conversion automatically.</p>\n<h3 id=\"arrow-format-and-memory-mapping\">Arrow Format and Memory Mapping</h3>\n<p>Datasets uses Apache Arrow for storage. Arrow’s columnar layout enables zero-copy reads and memory-mapped access. A dataset backed by an on-disk Arrow cache can exceed available RAM; the OS pages data in and out as needed.</p>\n<p>This architecture lets you work with terabyte-scale datasets on machines with limited memory:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Large dataset - only loads what you access</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"HuggingFaceFW/fineweb\"</span><span>, </span><span>\"sample-10BT\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Process without loading everything into memory</span></span>\n<span class=\"line\"><span>def</span><span> tokenize</span><span>(</span><span>examples</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> tokenizer</span><span>(examples[</span><span>\"text\"</span><span>], truncation</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>tokenized </span><span>=</span><span> dataset</span><span>.</span><span>map</span><span>(tokenize, batched</span><span>=</span><span>True</span><span>, num_proc</span><span>=</span><span>4</span><span>)</span></span></code></pre>\n<h3 id=\"streaming-for-large-datasets\">Streaming for Large Datasets</h3>\n<p>The English split of FineWeb is 45 terabytes. Streaming lets you use it without downloading:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Stream - no disk space needed</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span></span>\n<span class=\"line\"><span>    \"HuggingFaceFW/fineweb\"</span><span>,</span></span>\n<span class=\"line\"><span>    name</span><span>=</span><span>\"sample-10BT\"</span><span>,</span></span>\n<span class=\"line\"><span>    split</span><span>=</span><span>\"train\"</span><span>,</span></span>\n<span class=\"line\"><span>    streaming</span><span>=</span><span>True</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Iterate over examples</span></span>\n<span class=\"line\"><span>for</span><span> example </span><span>in</span><span> dataset</span><span>.</span><span>take</span><span>(</span><span>10</span><span>):</span></span>\n<span class=\"line\"><span>    print</span><span>(example[</span><span>\"text\"</span><span>][:</span><span>100</span><span>])</span></span></code></pre>\n<p>Streaming datasets support <code>map</code>, <code>filter</code>, <code>shuffle</code>, and <code>take</code>. Data downloads on demand as you iterate.</p>\n<h3 id=\"parquet-and-the-dataset-viewer\">Parquet and the Dataset Viewer</h3>\n<p>Datasets on the Hub automatically convert to Parquet, a columnar format that enables efficient partial reads. The Dataset Viewer uses this to display data in your browser without downloading entire files.</p>\n<p>Parquet stores column statistics and page indexes enabling:</p>\n<ul>\n<li>Reading only specific columns</li>\n<li>Filtering rows server-side</li>\n<li>Random access within files</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Read specific columns from Parquet</span></span>\n<span class=\"line\"><span>import</span><span> pyarrow</span><span>.</span><span>parquet </span><span>as</span><span> pq</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Only download 'text' column</span></span>\n<span class=\"line\"><span>table </span><span>=</span><span> pq</span><span>.</span><span>read_table</span><span>(</span></span>\n<span class=\"line\"><span>    \"hf://datasets/username/dataset/data.parquet\"</span><span>,</span></span>\n<span class=\"line\"><span>    columns</span><span>=</span><span>[</span><span>\"text\"</span><span>]</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<hr />\n<h2 id=\"the-transformers-library\">The Transformers Library</h2>\n<p>Transformers is the core library for working with pretrained models. It provides a unified API across architectures: BERT, GPT, T5, LLaMA, Mistral, Qwen, Gemma, and hundreds more.</p>\n<h3 id=\"the-pipeline-api\">The Pipeline API</h3>\n<p>Pipelines abstract away tokenization, model loading, and output processing:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> pipeline</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Sentiment analysis</span></span>\n<span class=\"line\"><span>classifier </span><span>=</span><span> pipeline</span><span>(</span><span>\"sentiment-analysis\"</span><span>)</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> classifier</span><span>(</span><span>\"I love this library!\"</span><span>)</span></span>\n<span class=\"line\"><span># [{'label': 'POSITIVE', 'score': 0.9998}]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Text generation</span></span>\n<span class=\"line\"><span>generator </span><span>=</span><span> pipeline</span><span>(</span><span>\"text-generation\"</span><span>, model</span><span>=</span><span>\"meta-llama/Llama-3.2-1B\"</span><span>)</span></span>\n<span class=\"line\"><span>output </span><span>=</span><span> generator</span><span>(</span><span>\"The future of AI is\"</span><span>, max_length</span><span>=</span><span>50</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Question answering</span></span>\n<span class=\"line\"><span>qa </span><span>=</span><span> pipeline</span><span>(</span><span>\"question-answering\"</span><span>)</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> qa</span><span>(question</span><span>=</span><span>\"What is the capital?\"</span><span>, context</span><span>=</span><span>\"France's capital is Paris.\"</span><span>)</span></span>\n<span class=\"line\"><span># {'answer': 'Paris', 'score': 0.99...}</span></span></code></pre>\n<p>Pipelines support GPU acceleration, batch processing, and half-precision:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># GPU with fp16</span></span>\n<span class=\"line\"><span>pipe </span><span>=</span><span> pipeline</span><span>(</span></span>\n<span class=\"line\"><span>    \"text-generation\"</span><span>,</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"tokenizers\">Tokenizers</h3>\n<p>Tokenizers convert text to token IDs the model understands. The library provides both Python (<code>PreTrainedTokenizer</code>) and Rust-backed fast (<code>PreTrainedTokenizerFast</code>) implementations.</p>\n<p>Transformers v5 redesigned tokenization. Tokenizers now separate architecture from learned vocabulary, similar to how PyTorch separates network structure from weights:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoTokenizer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Encode text</span></span>\n<span class=\"line\"><span>tokens </span><span>=</span><span> tokenizer</span><span>(</span><span>\"Hello world\"</span><span>, return_tensors</span><span>=</span><span>\"pt\"</span><span>)</span></span>\n<span class=\"line\"><span># {'input_ids': tensor([[...]]), 'attention_mask': tensor([[...]])}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Decode back</span></span>\n<span class=\"line\"><span>text </span><span>=</span><span> tokenizer</span><span>.</span><span>decode</span><span>(tokens[</span><span>\"input_ids\"</span><span>][</span><span>0</span><span>])</span></span></code></pre>\n<p>Different model families use different tokenization schemes: BPE (GPT, Llama), SentencePiece (T5, Gemma), or WordPiece (BERT). The tokenizer handles this transparently.</p>\n<h3 id=\"loading-models\">Loading Models</h3>\n<p>Load any model with <code>AutoModel</code> classes:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> AutoTokenizer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model_id </span><span>=</span><span> \"meta-llama/Llama-3.1-8B-Instruct\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(model_id)</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    model_id,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,  </span><span># Automatic device placement</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>\"auto\"</span><span>, </span><span># Use model's native dtype</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate</span></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> tokenizer</span><span>(</span><span>\"Explain transformers:\"</span><span>, return_tensors</span><span>=</span><span>\"pt\"</span><span>).</span><span>to</span><span>(model.device)</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(</span><span>**</span><span>inputs, max_new_tokens</span><span>=</span><span>100</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(tokenizer.</span><span>decode</span><span>(outputs[</span><span>0</span><span>]))</span></span></code></pre>\n<p>Control memory usage with quantization and offloading:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># 4-bit quantization</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    model_id,</span></span>\n<span class=\"line\"><span>    load_in_4bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># CPU offloading for models larger than VRAM</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    model_id,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>    offload_folder</span><span>=</span><span>\"offload\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<hr />\n<h2 id=\"transformers-library-internals\">Transformers Library Internals</h2>\n<p>Understanding how Transformers works under the hood helps debug issues and extend functionality.</p>\n<h3 id=\"automodel-and-autotokenizer-resolution\">AutoModel and AutoTokenizer Resolution</h3>\n<p>When you call <code>AutoModelForCausalLM.from_pretrained(\"meta-llama/Llama-3.1-8B\")</code>, the library performs several resolution steps:</p>\n<ol>\n<li><strong>Download config.json</strong> from the Hub repository</li>\n<li><strong>Parse model_type</strong> from the config (e.g., <code>\"llama\"</code>)</li>\n<li><strong>Look up the model class</strong> in an internal registry mapping model types to implementations</li>\n<li><strong>Instantiate the correct class</strong> (e.g., <code>LlamaForCausalLM</code>)</li>\n</ol>\n<p>The registry lives in <code>transformers/models/auto/modeling_auto.py</code>. Each Auto class maintains a mapping like:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>MODEL_FOR_CAUSAL_LM_MAPPING_NAMES </span><span>=</span><span> OrderedDict</span><span>([</span></span>\n<span class=\"line\"><span>    (</span><span>\"llama\"</span><span>, </span><span>\"LlamaForCausalLM\"</span><span>),</span></span>\n<span class=\"line\"><span>    (</span><span>\"mistral\"</span><span>, </span><span>\"MistralForCausalLM\"</span><span>),</span></span>\n<span class=\"line\"><span>    (</span><span>\"gpt2\"</span><span>, </span><span>\"GPT2LMHeadModel\"</span><span>),</span></span>\n<span class=\"line\"><span>    # ... hundreds more</span></span>\n<span class=\"line\"><span>])</span></span></code></pre>\n<p>For custom models hosted on the Hub, the <code>config.json</code> can include an <code>auto_map</code> field that points to local Python files:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"model_type\"</span><span>:</span><span> \"custom_model\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"auto_map\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"AutoConfig\"</span><span>:</span><span> \"configuration_custom.CustomConfig\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"AutoModelForCausalLM\"</span><span>:</span><span> \"modeling_custom.CustomModelForCausalLM\"</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>This requires <code>trust_remote_code=True</code> when loading, since it executes code from the repository. The <a href=\"https://huggingface.co/docs/transformers/en/model_doc/auto\">Auto Classes documentation</a> covers registration in detail.</p>\n<h3 id=\"the-modelingpy-architecture-pattern\">The modeling_*.py Architecture Pattern</h3>\n<p>Each model architecture follows a consistent file structure:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>transformers/models/llama/</span></span>\n<span class=\"line\"><span>├── __init__.py</span></span>\n<span class=\"line\"><span>├── configuration_llama.py    # LlamaConfig class</span></span>\n<span class=\"line\"><span>├── modeling_llama.py         # LlamaModel, LlamaForCausalLM, etc.</span></span>\n<span class=\"line\"><span>├── tokenization_llama.py     # LlamaTokenizer (slow)</span></span>\n<span class=\"line\"><span>└── tokenization_llama_fast.py # LlamaTokenizerFast (Rust)</span></span></code></pre>\n<p>The <code>modeling_*.py</code> file contains the actual PyTorch implementation. Models are built from composable pieces:</p>\n<ul>\n<li><strong>Embeddings</strong>: Token and position embeddings</li>\n<li><strong>Attention</strong>: Self-attention with various implementations (eager, SDPA, Flash Attention)</li>\n<li><strong>MLP/FFN</strong>: Feed-forward layers</li>\n<li><strong>Blocks/Layers</strong>: Stack attention + MLP with residuals and normalization</li>\n<li><strong>Model</strong>: The full transformer stack</li>\n<li><strong>Head</strong>: Task-specific output layers (LM head, classification head, etc.)</li>\n</ul>\n<p>A typical file structure:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>class</span><span> LlamaAttention</span><span>(</span><span>nn</span><span>.</span><span>Module</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Multi-headed attention from 'Attention Is All You Need' paper\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> LlamaMLP</span><span>(</span><span>nn</span><span>.</span><span>Module</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Feed-forward network with SiLU activation\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> LlamaDecoderLayer</span><span>(</span><span>nn</span><span>.</span><span>Module</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Single transformer block: attention + MLP\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> LlamaModel</span><span>(</span><span>LlamaPreTrainedModel</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"The bare Llama Model outputting raw hidden-states\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> LlamaForCausalLM</span><span>(</span><span>LlamaPreTrainedModel</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Llama Model with a language modeling head\"\"\"</span></span></code></pre>\n<h3 id=\"pretrainedmodel-base-class\">PreTrainedModel Base Class</h3>\n<p>Every model inherits from <code>PreTrainedModel</code>, which provides core functionality according to the <a href=\"https://huggingface.co/docs/transformers/main_classes/model\">Models documentation</a>:</p>\n<p><strong>Loading and saving:</strong></p>\n<ul>\n<li><code>from_pretrained()</code>: Load weights from Hub or local directory</li>\n<li><code>save_pretrained()</code>: Save model and config to directory</li>\n<li><code>push_to_hub()</code>: Upload to the Hub</li>\n</ul>\n<p><strong>Weight initialization:</strong></p>\n<ul>\n<li><code>_init_weights()</code>: Initialize a module’s weights (called recursively)</li>\n<li>The <code>_is_hf_initialized</code> flag prevents re-initializing tied parameters</li>\n</ul>\n<p><strong>Memory management:</strong></p>\n<ul>\n<li><code>gradient_checkpointing_enable()</code>: Trade compute for memory</li>\n<li><code>resize_token_embeddings()</code>: Adjust embedding size for new tokens</li>\n</ul>\n<p><strong>Device handling:</strong></p>\n<ul>\n<li><code>to()</code>: Move model to device</li>\n<li><code>half()</code>, <code>bfloat16()</code>, <code>float()</code>: Change precision</li>\n</ul>\n<p>Key class attributes every model defines:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>class</span><span> LlamaPreTrainedModel</span><span>(</span><span>PreTrainedModel</span><span>):</span></span>\n<span class=\"line\"><span>    config_class </span><span>=</span><span> LlamaConfig</span></span>\n<span class=\"line\"><span>    base_model_prefix </span><span>=</span><span> \"model\"</span></span>\n<span class=\"line\"><span>    supports_gradient_checkpointing </span><span>=</span><span> True</span></span>\n<span class=\"line\"><span>    _no_split_modules </span><span>=</span><span> [</span><span>\"LlamaDecoderLayer\"</span><span>]</span></span>\n<span class=\"line\"><span>    _skip_keys_device_placement </span><span>=</span><span> [</span><span>\"past_key_values\"</span><span>]</span></span></code></pre>\n<p>The <code>_no_split_modules</code> attribute tells the device placement algorithm which modules must stay on a single device.</p>\n<h3 id=\"generation-mixin-and-decoding-strategies\">Generation Mixin and Decoding Strategies</h3>\n<p>The <code>GenerationMixin</code> class adds <code>generate()</code> to models. It supports multiple decoding strategies documented in the <a href=\"https://huggingface.co/docs/transformers/en/generation_strategies\">Generation strategies guide</a>:</p>\n<p><strong>Greedy decoding</strong>: Select the highest probability token at each step.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(inputs, do_sample</span><span>=</span><span>False</span><span>)</span></span></code></pre>\n<p><strong>Beam search</strong>: Maintain multiple hypotheses, selecting the sequence with highest total probability.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(inputs, num_beams</span><span>=</span><span>5</span><span>, early_stopping</span><span>=</span><span>True</span><span>)</span></span></code></pre>\n<p><strong>Sampling with temperature</strong>: Sample from the probability distribution, with temperature controlling randomness.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(inputs, do_sample</span><span>=</span><span>True</span><span>, temperature</span><span>=</span><span>0.7</span><span>)</span></span></code></pre>\n<p><strong>Top-k sampling</strong>: Sample from the k most likely tokens.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(inputs, do_sample</span><span>=</span><span>True</span><span>, top_k</span><span>=</span><span>50</span><span>)</span></span></code></pre>\n<p><strong>Nucleus (top-p) sampling</strong>: Sample from the smallest set of tokens whose cumulative probability exceeds p. Described in Holtzman et al.’s “The Curious Case of Neural Text Degeneration.”</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(inputs, do_sample</span><span>=</span><span>True</span><span>, top_p</span><span>=</span><span>0.95</span><span>)</span></span></code></pre>\n<p>The generation loop internally:</p>\n<ol>\n<li>Runs the model forward pass to get logits</li>\n<li>Applies logits processors (temperature, top-k/p filtering, repetition penalty)</li>\n<li>Selects next token(s) based on the decoding strategy</li>\n<li>Appends to the sequence and updates the KV cache</li>\n<li>Checks stopping criteria (max length, EOS token, custom criteria)</li>\n<li>Repeats until done</li>\n</ol>\n<h3 id=\"kv-cache-implementation\">KV Cache Implementation</h3>\n<p>During autoregressive generation, the model computes key and value projections for all previous tokens at each step. Without caching, this means O(n^2) compute for generating n tokens. The KV cache stores these projections for reuse.</p>\n<p>The <a href=\"https://huggingface.co/docs/transformers/en/kv_cache\">KV cache strategies documentation</a> describes available implementations:</p>\n<p><strong>DynamicCache</strong> (default): Grows as generation progresses. Simple but can cause memory fragmentation.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(inputs, use_cache</span><span>=</span><span>True</span><span>)</span><span>  # Default</span></span></code></pre>\n<p><strong>StaticCache</strong>: Pre-allocates fixed-size tensors. Enables torch.compile() but wastes memory on short sequences.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(inputs, cache_implementation</span><span>=</span><span>\"static\"</span><span>)</span></span></code></pre>\n<p><strong>QuantizedCache</strong>: Compresses KV values to lower precision, based on the KIVI paper. Quantizes keys per-channel and values per-token.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> QuantizedCacheConfig</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>cache_config </span><span>=</span><span> QuantizedCacheConfig</span><span>(nbits</span><span>=</span><span>4</span><span>)</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(inputs, cache_config</span><span>=</span><span>cache_config)</span></span></code></pre>\n<p><strong>OffloadedCache</strong>: Moves KV cache to CPU, prefetching asynchronously. Enables longer sequences on limited VRAM.</p>\n<p>The cache is a tuple of (key_states, value_states) for each layer, stored in the model’s <code>past_key_values</code> attribute during generation.</p>\n<h3 id=\"attention-mask-handling\">Attention Mask Handling</h3>\n<p>Attention masks control which tokens can attend to which. Different model types use different masking strategies:</p>\n<p><strong>Padding mask</strong>: Binary mask where 1 = real token, 0 = padding. Prevents attention to padding tokens.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Tokenizer returns attention_mask automatically</span></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> tokenizer</span><span>([</span><span>\"short\"</span><span>, </span><span>\"much longer sequence\"</span><span>], padding</span><span>=</span><span>True</span><span>, return_tensors</span><span>=</span><span>\"pt\"</span><span>)</span></span>\n<span class=\"line\"><span># attention_mask: [[1, 1, 0, 0], [1, 1, 1, 1]]</span></span></code></pre>\n<p><strong>Causal mask</strong>: Lower-triangular mask for autoregressive models. Token at position i can only attend to positions 0..i.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Created internally for decoder models</span></span>\n<span class=\"line\"><span># [[1, 0, 0, 0],</span></span>\n<span class=\"line\"><span>#  [1, 1, 0, 0],</span></span>\n<span class=\"line\"><span>#  [1, 1, 1, 0],</span></span>\n<span class=\"line\"><span>#  [1, 1, 1, 1]]</span></span></code></pre>\n<p><strong>Bidirectional mask</strong>: Full attention for encoder models like BERT. Every token attends to every other token.</p>\n<p><strong>Encoder-decoder mask</strong>: Cross-attention from decoder to encoder uses the encoder’s padding mask.</p>\n<p>Models combine these masks. A causal LM with padding combines the causal mask with the padding mask using element-wise multiplication.</p>\n<p>The <a href=\"https://discuss.huggingface.co/t/difference-between-attention-mask-and-causal-mask/104922\">attention mask discussion</a> covers the differences in detail.</p>\n<hr />\n<h2 id=\"tokenization-deep-dive\">Tokenization Deep Dive</h2>\n<p>Tokenization converts raw text into the discrete tokens models consume. Modern tokenizers use subword algorithms that balance vocabulary size against sequence length.</p>\n<h3 id=\"tokenization-algorithms-compared\">Tokenization Algorithms Compared</h3>\n<p>Four algorithms dominate according to the <a href=\"https://huggingface.co/docs/transformers/tokenizer_summary\">tokenizer summary</a>:</p>\n<p><strong>Byte-Pair Encoding (BPE)</strong>: Starts with individual characters, iteratively merges the most frequent pair. Used by GPT-2, GPT-3, GPT-4, LLaMA, RoBERTa.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>Training corpus: \"low lower lowest\"</span></span>\n<span class=\"line\"><span>Initial: ['l', 'o', 'w', ' ', 'l', 'o', 'w', 'e', 'r', ...]</span></span>\n<span class=\"line\"><span>After merges: ['low', 'low', 'er', 'low', 'est']</span></span></code></pre>\n<p><strong>WordPiece</strong>: Similar to BPE but merges based on likelihood maximization rather than frequency. Prefixes continuations with <code>##</code>. Used by BERT, DistilBERT, Electra.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>\"unbelievable\" → [\"un\", \"##believ\", \"##able\"]</span></span></code></pre>\n<p><strong>Unigram</strong>: Starts with a large vocabulary, iteratively removes tokens that least affect the training loss. Probabilistic: can produce multiple valid tokenizations. Used within SentencePiece.</p>\n<p><strong>SentencePiece</strong>: A framework that treats text as a raw byte stream, enabling language-agnostic tokenization. Can use BPE or Unigram internally. Used by T5, ALBERT, XLNet, Gemma.</p>\n<p>The <a href=\"https://codesignal.com/learn/courses/2-modern-tokenization-techniques-for-ai-llms/lessons/comparing-bpe-wordpiece-and-sentencepiece-in-nlp\">tokenization algorithms comparison</a> shows that BPE offers better contextual specialization while SentencePiece achieves better encoding efficiency.</p>\n<h3 id=\"the-tokenizerjson-file-format\">The tokenizer.json File Format</h3>\n<p>Fast tokenizers serialize their configuration to <code>tokenizer.json</code>. This file encodes:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"version\"</span><span>:</span><span> \"1.0\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"truncation\"</span><span>:</span><span> null</span><span>,</span></span>\n<span class=\"line\"><span>  \"padding\"</span><span>:</span><span> null</span><span>,</span></span>\n<span class=\"line\"><span>  \"added_tokens\"</span><span>:</span><span> [...]</span><span>,</span></span>\n<span class=\"line\"><span>  \"normalizer\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"type\"</span><span>:</span><span> \"Sequence\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"normalizers\"</span><span>:</span><span> [...]</span></span>\n<span class=\"line\"><span>  }</span><span>,</span></span>\n<span class=\"line\"><span>  \"pre_tokenizer\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"type\"</span><span>:</span><span> \"ByteLevel\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"add_prefix_space\"</span><span>:</span><span> false</span></span>\n<span class=\"line\"><span>  }</span><span>,</span></span>\n<span class=\"line\"><span>  \"model\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"type\"</span><span>:</span><span> \"BPE\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"vocab\"</span><span>:</span><span> {</span><span>\"&lt;s&gt;\"</span><span>:</span><span> 0</span><span>,</span><span> \"&lt;/s&gt;\"</span><span>:</span><span> 1</span><span>,</span><span> ...}</span><span>,</span></span>\n<span class=\"line\"><span>    \"merges\"</span><span>:</span><span> [</span><span>\"Ġ t\"</span><span>,</span><span> \"Ġ a\"</span><span>,</span><span> \"h e\"</span><span>,</span><span> ...]</span></span>\n<span class=\"line\"><span>  }</span><span>,</span></span>\n<span class=\"line\"><span>  \"decoder\"</span><span>:</span><span> {...}</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The <code>merges</code> array encodes BPE merge rules in order of application. During tokenization, the algorithm:</p>\n<ol>\n<li>Splits input using the pre-tokenizer</li>\n<li>For each word, starts with individual bytes/characters</li>\n<li>Applies merges in order until no more apply</li>\n<li>Maps resulting subwords to vocabulary IDs</li>\n</ol>\n<h3 id=\"special-tokens-and-why-they-matter\">Special Tokens and Why They Matter</h3>\n<p>Special tokens serve structural roles:</p>\n<ul>\n<li><code>[CLS]</code> / <code>&lt;s&gt;</code>: Sequence start, used for classification</li>\n<li><code>[SEP]</code> / <code>&lt;/s&gt;</code>: Sequence end or separator</li>\n<li><code>[PAD]</code>: Padding for batching</li>\n<li><code>[UNK]</code>: Unknown tokens (rare with subword tokenizers)</li>\n<li><code>[MASK]</code>: Masked positions for MLM training</li>\n</ul>\n<p>Models expect specific special tokens. Using the wrong ones breaks the model:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Chat models need proper formatting</span></span>\n<span class=\"line\"><span>messages </span><span>=</span><span> [</span><span>{</span><span>\"role\"</span><span>:</span><span> \"user\"</span><span>,</span><span> \"content\"</span><span>:</span><span> \"Hello\"</span><span>}</span><span>]</span></span>\n<span class=\"line\"><span>formatted </span><span>=</span><span> tokenizer</span><span>.</span><span>apply_chat_template</span><span>(messages, tokenize</span><span>=</span><span>False</span><span>)</span></span>\n<span class=\"line\"><span># Adds &lt;|begin_of_text|&gt;, &lt;|start_header_id|&gt;, etc.</span></span></code></pre>\n<h3 id=\"fast-vs-slow-tokenizers\">Fast vs Slow Tokenizers</h3>\n<p>The <a href=\"https://huggingface.co/docs/transformers/en/fast_tokenizers\">fast tokenizers guide</a> explains the difference:</p>\n<p><strong>Slow tokenizers</strong> (<code>PreTrainedTokenizer</code>): Pure Python implementation. Full control, easier to modify, but 10-100x slower.</p>\n<p><strong>Fast tokenizers</strong> (<code>PreTrainedTokenizerFast</code>): Rust backend via the <code>tokenizers</code> library. Parallel processing, sub-millisecond tokenization for most inputs.</p>\n<p>The fast tokenizer provides additional features:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"bert-base-uncased\"</span><span>)</span></span>\n<span class=\"line\"><span>encoding </span><span>=</span><span> tokenizer</span><span>(</span><span>\"Hello world\"</span><span>, return_offsets_mapping</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Character offsets for each token</span></span>\n<span class=\"line\"><span>print</span><span>(encoding.offset_mapping)</span></span>\n<span class=\"line\"><span># [(0, 5), (6, 11)]  # \"Hello\", \"world\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Map token index to word index</span></span>\n<span class=\"line\"><span>print</span><span>(encoding.</span><span>word_ids</span><span>())</span></span>\n<span class=\"line\"><span># [0, 1]</span></span></code></pre>\n<h3 id=\"training-tokenizers-from-scratch\">Training Tokenizers from Scratch</h3>\n<p>The <a href=\"https://huggingface.co/docs/tokenizers/en/quicktour\">tokenizers quicktour</a> shows how to train custom tokenizers:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> tokenizers </span><span>import</span><span> Tokenizer</span></span>\n<span class=\"line\"><span>from</span><span> tokenizers</span><span>.</span><span>models </span><span>import</span><span> BPE</span></span>\n<span class=\"line\"><span>from</span><span> tokenizers</span><span>.</span><span>trainers </span><span>import</span><span> BpeTrainer</span></span>\n<span class=\"line\"><span>from</span><span> tokenizers</span><span>.</span><span>pre_tokenizers </span><span>import</span><span> Whitespace</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Initialize with BPE model</span></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> Tokenizer</span><span>(</span><span>BPE</span><span>(unk_token</span><span>=</span><span>\"[UNK]\"</span><span>))</span></span>\n<span class=\"line\"><span>tokenizer</span><span>.</span><span>pre_tokenizer </span><span>=</span><span> Whitespace</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Configure trainer</span></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> BpeTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    vocab_size</span><span>=</span><span>30000</span><span>,</span></span>\n<span class=\"line\"><span>    min_frequency</span><span>=</span><span>2</span><span>,</span></span>\n<span class=\"line\"><span>    special_tokens</span><span>=</span><span>[</span><span>\"[UNK]\"</span><span>, </span><span>\"[CLS]\"</span><span>, </span><span>\"[SEP]\"</span><span>, </span><span>\"[PAD]\"</span><span>, </span><span>\"[MASK]\"</span><span>]</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Train on files</span></span>\n<span class=\"line\"><span>files </span><span>=</span><span> [</span><span>\"wiki.txt\"</span><span>,</span><span> \"books.txt\"</span><span>]</span></span>\n<span class=\"line\"><span>tokenizer</span><span>.</span><span>train</span><span>(files, trainer)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Save</span></span>\n<span class=\"line\"><span>tokenizer</span><span>.</span><span>save</span><span>(</span><span>\"my-tokenizer.json\"</span><span>)</span></span></code></pre>\n<p>For byte-level BPE (like GPT-2):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> tokenizers </span><span>import</span><span> pre_tokenizers</span></span>\n<span class=\"line\"><span>from</span><span> tokenizers</span><span>.</span><span>pre_tokenizers </span><span>import</span><span> ByteLevel</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> Tokenizer</span><span>(</span><span>BPE</span><span>(unk_token</span><span>=</span><span>\"[UNK]\"</span><span>))</span></span>\n<span class=\"line\"><span>tokenizer</span><span>.</span><span>pre_tokenizer </span><span>=</span><span> ByteLevel</span><span>(add_prefix_space</span><span>=</span><span>False</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> BpeTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    vocab_size</span><span>=</span><span>50257</span><span>,</span></span>\n<span class=\"line\"><span>    initial_alphabet</span><span>=</span><span>pre_tokenizers.ByteLevel.</span><span>alphabet</span><span>(),</span></span>\n<span class=\"line\"><span>    special_tokens</span><span>=</span><span>[</span><span>\"[UNK]\"</span><span>, </span><span>\"[PAD]\"</span><span>]</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>Training a tokenizer on a GB of text takes under 20 seconds with the Rust backend.</p>\n<hr />\n<h2 id=\"model-loading-mechanics\">Model Loading Mechanics</h2>\n<p>Loading large models requires careful memory management. Transformers and Accelerate provide several mechanisms.</p>\n<h3 id=\"sharded-checkpoint-loading\">Sharded Checkpoint Loading</h3>\n<p>Large models split weights across multiple files (shards). A 70B model might have 15 shards of ~10GB each. The <code>model.safetensors.index.json</code> file maps parameter names to shard files:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"metadata\"</span><span>:</span><span> {</span><span>\"total_size\"</span><span>:</span><span> 140000000000</span><span>}</span><span>,</span></span>\n<span class=\"line\"><span>  \"weight_map\"</span><span>:</span><span> {</span></span>\n<span class=\"line\"><span>    \"model.embed_tokens.weight\"</span><span>:</span><span> \"model-00001-of-00015.safetensors\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"model.layers.0.self_attn.q_proj.weight\"</span><span>:</span><span> \"model-00001-of-00015.safetensors\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"model.layers.0.self_attn.k_proj.weight\"</span><span>:</span><span> \"model-00001-of-00015.safetensors\"</span><span>,</span></span>\n<span class=\"line\"><span>    ...</span></span>\n<span class=\"line\"><span>  }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>During loading, Transformers:</p>\n<ol>\n<li>Downloads/locates all shard files</li>\n<li>Loads one shard at a time</li>\n<li>Copies relevant tensors to the model</li>\n<li>Discards the shard before loading the next</li>\n</ol>\n<p>This keeps peak memory at (model size) + (largest shard size) rather than (model size) * 2.</p>\n<h3 id=\"devicemapauto-and-layer-distribution\">device_map=“auto” and Layer Distribution</h3>\n<p>The <a href=\"https://huggingface.co/docs/accelerate/en/usage_guides/big_modeling\">big model inference guide</a> explains automatic device placement:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"meta-llama/Llama-3.1-70B\"</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>torch.bfloat16,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>Accelerate computes a device map by:</p>\n<ol>\n<li>Estimating each layer’s memory requirement</li>\n<li>Querying available GPU memory</li>\n<li>Assigning layers to devices in order (GPU 0, GPU 1, …, CPU, disk)</li>\n<li>Ensuring no module in <code>_no_split_modules</code> crosses device boundaries</li>\n</ol>\n<p>Available strategies:</p>\n<ul>\n<li><code>\"auto\"</code>: Fill GPUs in order, overflow to CPU/disk</li>\n<li><code>\"balanced\"</code>: Distribute evenly across GPUs</li>\n<li><code>\"balanced_low_0\"</code>: Keep GPU 0 lighter (useful when GPU 0 handles other tasks)</li>\n<li><code>\"sequential\"</code>: Fill devices completely before moving to next</li>\n</ul>\n<p>View the computed map:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> accelerate </span><span>import</span><span> infer_auto_device_map</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoConfig</span><span>,</span><span> AutoModelForCausalLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>config </span><span>=</span><span> AutoConfig</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.1-70B\"</span><span>)</span></span>\n<span class=\"line\"><span>with</span><span> init_empty_weights</span><span>():</span></span>\n<span class=\"line\"><span>    model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_config</span><span>(config)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>device_map </span><span>=</span><span> infer_auto_device_map</span><span>(model, max_memory</span><span>=</span><span>{</span><span>0</span><span>: </span><span>\"24GB\"</span><span>, </span><span>1</span><span>: </span><span>\"24GB\"</span><span>, </span><span>\"cpu\"</span><span>: </span><span>\"100GB\"</span><span>})</span></span>\n<span class=\"line\"><span>print</span><span>(device_map)</span></span>\n<span class=\"line\"><span># {'model.embed_tokens': 0, 'model.layers.0': 0, ..., 'model.layers.40': 1, ...}</span></span></code></pre>\n<p><strong>Important limitation</strong>: <code>device_map=\"auto\"</code> supports inference only, not training. The hook-based execution doesn’t support backpropagation across devices.</p>\n<h3 id=\"bitsandbytes-quantization-paths\">bitsandbytes Quantization Paths</h3>\n<p>The <a href=\"https://huggingface.co/docs/transformers/en/quantization/bitsandbytes\">bitsandbytes guide</a> details on-the-fly quantization:</p>\n<p><strong>8-bit quantization (LLM.int8)</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"meta-llama/Llama-3.1-8B\"</span><span>,</span></span>\n<span class=\"line\"><span>    load_in_8bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>Uses mixed-precision decomposition: outlier features (magnitude &gt; 6.0 by default) stay in FP16, the rest quantizes to INT8. This preserves quality for models with outlier activations.</p>\n<p><strong>4-bit quantization (NF4/FP4)</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> BitsAndBytesConfig</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>config </span><span>=</span><span> BitsAndBytesConfig</span><span>(</span></span>\n<span class=\"line\"><span>    load_in_4bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_quant_type</span><span>=</span><span>\"nf4\"</span><span>,           </span><span># NormalFloat4 or \"fp4\"</span></span>\n<span class=\"line\"><span>    bnb_4bit_compute_dtype</span><span>=</span><span>torch.bfloat16,  </span><span># Compute in bf16</span></span>\n<span class=\"line\"><span>    bnb_4bit_use_double_quant</span><span>=</span><span>True</span><span>,      </span><span># Quantize the quantization constants</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>NF4 (NormalFloat4) optimizes for normally-distributed weights by placing quantization levels at distribution quantiles. Double quantization applies 8-bit quantization to the scaling factors, saving an additional 0.4 bits per parameter.</p>\n<p>PyTorch doesn’t support 4-bit dtypes, so bitsandbytes packs two 4-bit values into one 8-bit value, changing tensor shapes. Computation happens in 16/32-bit after dequantization.</p>\n<h3 id=\"safetensors-lazy-loading\">Safetensors Lazy Loading</h3>\n<p>The <a href=\"https://huggingface.co/docs/safetensors\">safetensors documentation</a> describes lazy loading:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> safetensors </span><span>import</span><span> safe_open</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>with</span><span> safe_open</span><span>(</span><span>\"model.safetensors\"</span><span>, framework</span><span>=</span><span>\"pt\"</span><span>)</span><span> as</span><span> f</span><span>:</span></span>\n<span class=\"line\"><span>    # Only metadata loaded at this point</span></span>\n<span class=\"line\"><span>    print</span><span>(f.</span><span>keys</span><span>())</span><span>  # List all tensor names</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Individual tensors load on-demand</span></span>\n<span class=\"line\"><span>    tensor </span><span>=</span><span> f</span><span>.</span><span>get_tensor</span><span>(</span><span>\"model.layers.0.self_attn.q_proj.weight\"</span><span>)</span></span></code></pre>\n<p>Safetensors uses memory-mapping: the OS maps the file into virtual address space, loading pages only when accessed. This enables:</p>\n<ul>\n<li>Inspecting tensor names without loading weights</li>\n<li>Loading specific tensors for distributed inference</li>\n<li>Parallel loading from multiple processes sharing the memory map</li>\n</ul>\n<p>The file layout keeps tensor metadata contiguous at the start, enabling fast scanning without reading weight data.</p>\n<h3 id=\"memory-estimation-before-loading\">Memory Estimation Before Loading</h3>\n<p>Estimate memory requirements before attempting to load:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> accelerate</span><span>.</span><span>commands</span><span>.</span><span>estimate </span><span>import</span><span> estimate_command</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># CLI</span></span>\n<span class=\"line\"><span># accelerate estimate-memory meta-llama/Llama-3.1-70B</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Programmatic</span></span>\n<span class=\"line\"><span>from</span><span> accelerate </span><span>import</span><span> init_empty_weights</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoConfig</span><span>,</span><span> AutoModelForCausalLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>config </span><span>=</span><span> AutoConfig</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.1-70B\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>with</span><span> init_empty_weights</span><span>():</span></span>\n<span class=\"line\"><span>    model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_config</span><span>(config)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>param_bytes </span><span>=</span><span> sum</span><span>(p.</span><span>numel</span><span>() </span><span>*</span><span> p.</span><span>element_size</span><span>() </span><span>for</span><span> p </span><span>in</span><span> model.</span><span>parameters</span><span>())</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Model size: </span><span>{</span><span>param_bytes </span><span>/</span><span> 1e9</span><span>:.1f</span><span>}</span><span> GB in </span><span>{</span><span>model.dtype</span><span>}</span><span>\"</span><span>)</span></span></code></pre>\n<p>The <code>init_empty_weights()</code> context creates tensors on PyTorch’s “meta” device, which stores only shape and dtype, not actual data.</p>\n<hr />\n<h2 id=\"training-infrastructure\">Training Infrastructure</h2>\n<h3 id=\"trainer-class-internals\">Trainer Class Internals</h3>\n<p>The <a href=\"https://huggingface.co/docs/transformers/en/main_classes/trainer\">Trainer documentation</a> covers the main training API. Key components:</p>\n<p><strong>Training loop structure:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>for</span><span> epoch </span><span>in</span><span> range</span><span>(num_epochs):</span></span>\n<span class=\"line\"><span>    for</span><span> step</span><span>,</span><span> batch </span><span>in</span><span> enumerate</span><span>(train_dataloader):</span></span>\n<span class=\"line\"><span>        # Forward pass</span></span>\n<span class=\"line\"><span>        outputs </span><span>=</span><span> model</span><span>(</span><span>**</span><span>batch)</span></span>\n<span class=\"line\"><span>        loss </span><span>=</span><span> outputs</span><span>.</span><span>loss</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Scale loss for gradient accumulation</span></span>\n<span class=\"line\"><span>        loss </span><span>=</span><span> loss </span><span>/</span><span> gradient_accumulation_steps</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Backward pass</span></span>\n<span class=\"line\"><span>        loss</span><span>.</span><span>backward</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Update weights every N steps</span></span>\n<span class=\"line\"><span>        if</span><span> (step </span><span>+</span><span> 1</span><span>) </span><span>%</span><span> gradient_accumulation_steps </span><span>==</span><span> 0</span><span>:</span></span>\n<span class=\"line\"><span>            optimizer</span><span>.</span><span>step</span><span>()</span></span>\n<span class=\"line\"><span>            lr_scheduler</span><span>.</span><span>step</span><span>()</span></span>\n<span class=\"line\"><span>            optimizer</span><span>.</span><span>zero_grad</span><span>()</span></span></code></pre>\n<p><strong>Subclassing for custom behavior:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> Trainer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> CustomTrainer</span><span>(</span><span>Trainer</span><span>):</span></span>\n<span class=\"line\"><span>    def</span><span> compute_loss</span><span>(</span><span>self</span><span>,</span><span> model</span><span>,</span><span> inputs</span><span>,</span><span> return_outputs</span><span>=</span><span>False</span><span>,</span><span> num_items_in_batch</span><span>=</span><span>None</span><span>):</span></span>\n<span class=\"line\"><span>        outputs </span><span>=</span><span> model</span><span>(</span><span>**</span><span>inputs)</span></span>\n<span class=\"line\"><span>        loss </span><span>=</span><span> my_custom_loss</span><span>(outputs.logits, inputs[</span><span>\"labels\"</span><span>])</span></span>\n<span class=\"line\"><span>        return</span><span> (loss</span><span>,</span><span> outputs) </span><span>if</span><span> return_outputs </span><span>else</span><span> loss</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> training_step</span><span>(</span><span>self</span><span>,</span><span> model</span><span>,</span><span> inputs</span><span>,</span><span> num_items_in_batch</span><span>=</span><span>None</span><span>):</span></span>\n<span class=\"line\"><span>        # Custom training step logic</span></span>\n<span class=\"line\"><span>        ...</span></span></code></pre>\n<h3 id=\"gradient-accumulation-implementation\">Gradient Accumulation Implementation</h3>\n<p>The <a href=\"https://huggingface.co/blog/gradient_accumulation\">gradient accumulation blog post</a> explains the details:</p>\n<p>Gradient accumulation simulates larger batch sizes by accumulating gradients over multiple forward passes before updating weights.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>training_args </span><span>=</span><span> TrainingArguments</span><span>(</span></span>\n<span class=\"line\"><span>    per_device_train_batch_size</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>    gradient_accumulation_steps</span><span>=</span><span>8</span><span>,</span></span>\n<span class=\"line\"><span>    # Effective batch size = 4 * 8 = 32</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>Internally, the Trainer scales the loss by <code>1/gradient_accumulation_steps</code> before each backward pass. This ensures the accumulated gradients match what you’d get from a single large batch.</p>\n<p>Recent versions fixed a subtle bug: some loss functions (like cross-entropy) already average over batch elements, so the scaling was being applied twice in certain configurations. The fix introduced <code>num_items_in_batch</code> parameter to loss computation functions.</p>\n<h3 id=\"mixed-precision-training-paths\">Mixed Precision Training Paths</h3>\n<p>The <a href=\"https://docs.pytorch.org/docs/stable/amp.html\">PyTorch AMP documentation</a> covers mixed precision:</p>\n<p><strong>FP16 training</strong> (requires loss scaling):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>amp </span><span>import</span><span> autocast</span><span>,</span><span> GradScaler</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>scaler </span><span>=</span><span> GradScaler</span><span>(</span><span>\"cuda\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> batch </span><span>in</span><span> dataloader</span><span>:</span></span>\n<span class=\"line\"><span>    optimizer</span><span>.</span><span>zero_grad</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    with</span><span> autocast</span><span>(</span><span>\"cuda\"</span><span>, dtype</span><span>=</span><span>torch.float16):</span></span>\n<span class=\"line\"><span>        outputs </span><span>=</span><span> model</span><span>(</span><span>**</span><span>batch)</span></span>\n<span class=\"line\"><span>        loss </span><span>=</span><span> outputs</span><span>.</span><span>loss</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    scaler</span><span>.</span><span>scale</span><span>(loss).</span><span>backward</span><span>()</span></span>\n<span class=\"line\"><span>    scaler</span><span>.</span><span>step</span><span>(optimizer)</span></span>\n<span class=\"line\"><span>    scaler</span><span>.</span><span>update</span><span>()</span></span></code></pre>\n<p><code>GradScaler</code> prevents gradient underflow by scaling loss up before backward (so gradients are larger), then scaling gradients down before optimizer step.</p>\n<p><strong>BF16 training</strong> (no loss scaling needed):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>with</span><span> autocast</span><span>(</span><span>\"cuda\"</span><span>, dtype</span><span>=</span><span>torch.bfloat16):</span></span>\n<span class=\"line\"><span>    outputs </span><span>=</span><span> model</span><span>(</span><span>**</span><span>batch)</span></span>\n<span class=\"line\"><span>    loss </span><span>=</span><span> outputs</span><span>.</span><span>loss</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>loss</span><span>.</span><span>backward</span><span>()</span></span>\n<span class=\"line\"><span>optimizer</span><span>.</span><span>step</span><span>()</span></span></code></pre>\n<p>BFloat16 has the same exponent range as FP32, so it doesn’t suffer from underflow. GradScaler is unnecessary and can harm training.</p>\n<p><strong>With Trainer:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>training_args </span><span>=</span><span> TrainingArguments</span><span>(</span></span>\n<span class=\"line\"><span>    bf16</span><span>=</span><span>True</span><span>,   </span><span># or fp16=True</span></span>\n<span class=\"line\"><span>    # ...</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"peft-integration-with-training\">PEFT Integration with Training</h3>\n<p>The <a href=\"https://huggingface.co/docs/transformers/en/peft\">PEFT integration guide</a> shows how adapters work with Trainer:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> peft </span><span>import</span><span> LoraConfig</span><span>,</span><span> TaskType</span><span>,</span><span> get_peft_model</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> Trainer</span><span>,</span><span> TrainingArguments</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load base model</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Configure LoRA</span></span>\n<span class=\"line\"><span>lora_config </span><span>=</span><span> LoraConfig</span><span>(</span></span>\n<span class=\"line\"><span>    r</span><span>=</span><span>16</span><span>,</span></span>\n<span class=\"line\"><span>    lora_alpha</span><span>=</span><span>32</span><span>,</span></span>\n<span class=\"line\"><span>    target_modules</span><span>=</span><span>[</span><span>\"q_proj\"</span><span>, </span><span>\"v_proj\"</span><span>, </span><span>\"k_proj\"</span><span>, </span><span>\"o_proj\"</span><span>],</span></span>\n<span class=\"line\"><span>    lora_dropout</span><span>=</span><span>0.05</span><span>,</span></span>\n<span class=\"line\"><span>    task_type</span><span>=</span><span>TaskType.CAUSAL_LM,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Wrap model with PEFT</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> get_peft_model</span><span>(model, lora_config)</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>print_trainable_parameters</span><span>()</span></span>\n<span class=\"line\"><span># trainable params: 13,631,488 || all params: 8,043,892,736 || trainable%: 0.17</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Train normally - Trainer handles the PEFT model</span></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> Trainer</span><span>(model</span><span>=</span><span>model, args</span><span>=</span><span>training_args, train_dataset</span><span>=</span><span>dataset)</span></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>train</span><span>()</span></span></code></pre>\n<p>PEFT works by:</p>\n<ol>\n<li>Identifying target modules (linear layers matching <code>target_modules</code>)</li>\n<li>Replacing them with wrapped versions containing LoRA matrices</li>\n<li>Freezing base model parameters (requires_grad=False)</li>\n<li>Only the small A and B matrices train</li>\n</ol>\n<p>For QLoRA (quantized base + LoRA):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"meta-llama/Llama-3.1-8B\"</span><span>,</span></span>\n<span class=\"line\"><span>    load_in_4bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_quant_type</span><span>=</span><span>\"nf4\"</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_compute_dtype</span><span>=</span><span>torch.bfloat16,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> get_peft_model</span><span>(model, lora_config)</span></span></code></pre>\n<p>The quantized weights stay frozen (they can’t be trained anyway), and only the FP16/BF16 LoRA matrices update.</p>\n<hr />\n<h2 id=\"hub-api-deep-dive\">Hub API Deep Dive</h2>\n<h3 id=\"git-lfs-and-large-file-handling\">Git LFS and Large File Handling</h3>\n<p>The <a href=\"https://huggingface.co/docs/hub/en/storage-backends\">storage documentation</a> explains file handling:</p>\n<p>Historically, Hub repositories used Git LFS (Large File Storage) for files over 10MB. LFS stores file content in a separate server, with Git tracking only pointer files.</p>\n<p>As of May 2025, new repositories default to Xet storage, which offers chunk-level deduplication. When you modify a file, only changed chunks upload, not the entire file. This matters for iterative checkpoint uploads during training.</p>\n<p><strong>File size limits:</strong></p>\n<ul>\n<li>Pre-receive hook rejects commits with files &gt; 10MB not tracked by LFS/Xet</li>\n<li>Individual files capped at 50GB</li>\n<li>Recommended: split large files into ~20GB chunks</li>\n</ul>\n<p><strong>Repository limits:</strong></p>\n<ul>\n<li>&lt; 100k files total recommended</li>\n<li>&lt; 10k files per directory</li>\n<li>For large datasets, use Parquet or WebDataset formats</li>\n</ul>\n<h3 id=\"repository-structure-conventions\">Repository Structure Conventions</h3>\n<p>Model repositories follow conventions that enable auto-loading:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>my-model/</span></span>\n<span class=\"line\"><span>├── README.md                    # Model card with YAML frontmatter</span></span>\n<span class=\"line\"><span>├── config.json                  # Model architecture configuration</span></span>\n<span class=\"line\"><span>├── model.safetensors            # Weights (or sharded)</span></span>\n<span class=\"line\"><span>├── model.safetensors.index.json # Shard index if sharded</span></span>\n<span class=\"line\"><span>├── tokenizer.json               # Fast tokenizer config</span></span>\n<span class=\"line\"><span>├── tokenizer_config.json        # Tokenizer settings</span></span>\n<span class=\"line\"><span>├── special_tokens_map.json      # Special token definitions</span></span>\n<span class=\"line\"><span>├── vocab.json                   # Vocabulary (some tokenizers)</span></span>\n<span class=\"line\"><span>├── merges.txt                   # BPE merges (some tokenizers)</span></span>\n<span class=\"line\"><span>└── generation_config.json       # Default generation settings</span></span></code></pre>\n<p>Dataset repositories:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>my-dataset/</span></span>\n<span class=\"line\"><span>├── README.md                    # Dataset card</span></span>\n<span class=\"line\"><span>├── data/</span></span>\n<span class=\"line\"><span>│   ├── train-00000-of-00010.parquet</span></span>\n<span class=\"line\"><span>│   ├── train-00001-of-00010.parquet</span></span>\n<span class=\"line\"><span>│   └── ...</span></span>\n<span class=\"line\"><span>└── dataset_info.json            # Dataset metadata</span></span></code></pre>\n<h3 id=\"model-card-yaml-frontmatter\">Model Card YAML Frontmatter</h3>\n<p>The <a href=\"https://huggingface.co/docs/hub/en/model-cards\">model cards documentation</a> specifies the metadata format:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>---</span></span>\n<span class=\"line\"><span>language</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>en</span></span>\n<span class=\"line\"><span>  - </span><span>zh</span></span>\n<span class=\"line\"><span>license</span><span>:</span><span> apache-2.0</span></span>\n<span class=\"line\"><span>library_name</span><span>:</span><span> transformers</span></span>\n<span class=\"line\"><span>pipeline_tag</span><span>:</span><span> text-generation</span></span>\n<span class=\"line\"><span>tags</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>llama</span></span>\n<span class=\"line\"><span>  - </span><span>chat</span></span>\n<span class=\"line\"><span>datasets</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>HuggingFaceH4/ultrachat_200k</span></span>\n<span class=\"line\"><span>base_model</span><span>:</span><span> meta-llama/Llama-3.1-8B</span></span>\n<span class=\"line\"><span>model-index</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>name</span><span>:</span><span> my-model</span></span>\n<span class=\"line\"><span>    results</span><span>:</span></span>\n<span class=\"line\"><span>      - </span><span>task</span><span>:</span></span>\n<span class=\"line\"><span>          type</span><span>:</span><span> text-generation</span></span>\n<span class=\"line\"><span>        dataset</span><span>:</span></span>\n<span class=\"line\"><span>          name</span><span>:</span><span> MMLU</span></span>\n<span class=\"line\"><span>          type</span><span>:</span><span> cais/mmlu</span></span>\n<span class=\"line\"><span>        metrics</span><span>:</span></span>\n<span class=\"line\"><span>          - </span><span>name</span><span>:</span><span> accuracy</span></span>\n<span class=\"line\"><span>            type</span><span>:</span><span> accuracy</span></span>\n<span class=\"line\"><span>            value</span><span>:</span><span> 0.72</span></span>\n<span class=\"line\"><span>---</span></span>\n<span class=\"line\"><span># My Model</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>Model description here...</span></span></code></pre>\n<p>The Hub parses this YAML to:</p>\n<ul>\n<li>Enable filtering by language, license, task</li>\n<li>Display the inference widget based on <code>pipeline_tag</code></li>\n<li>Show benchmark results on the model page</li>\n<li>Link to base models and training datasets</li>\n</ul>\n<h3 id=\"huggingfacehub-library-internals\">huggingface_hub Library Internals</h3>\n<p>The <a href=\"https://huggingface.co/docs/huggingface_hub/package_reference/hf_api\">huggingface_hub documentation</a> covers the Python client:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> huggingface_hub </span><span>import</span><span> HfApi</span><span>,</span><span> hf_hub_download</span><span>,</span><span> snapshot_download</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>api </span><span>=</span><span> HfApi</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># List repository files</span></span>\n<span class=\"line\"><span>files </span><span>=</span><span> api</span><span>.</span><span>list_repo_files</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Get model info</span></span>\n<span class=\"line\"><span>info </span><span>=</span><span> api</span><span>.</span><span>model_info</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(info.downloads)</span><span>  # Download count</span></span>\n<span class=\"line\"><span>print</span><span>(info.tags)</span><span>       # Tags like \"llama\", \"text-generation\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Download single file (cached)</span></span>\n<span class=\"line\"><span>path </span><span>=</span><span> hf_hub_download</span><span>(</span></span>\n<span class=\"line\"><span>    repo_id</span><span>=</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>,</span></span>\n<span class=\"line\"><span>    filename</span><span>=</span><span>\"config.json\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Download entire repo</span></span>\n<span class=\"line\"><span>local_dir </span><span>=</span><span> snapshot_download</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span></code></pre>\n<p>The library uses a content-addressed cache at <code>~/.cache/huggingface/hub/</code>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>hub/</span></span>\n<span class=\"line\"><span>├── models--meta-llama--Llama-3.1-8B/</span></span>\n<span class=\"line\"><span>│   ├── refs/</span></span>\n<span class=\"line\"><span>│   │   └── main              # Points to commit hash</span></span>\n<span class=\"line\"><span>│   ├── blobs/</span></span>\n<span class=\"line\"><span>│   │   └── abc123...         # Content-addressed files</span></span>\n<span class=\"line\"><span>│   └── snapshots/</span></span>\n<span class=\"line\"><span>│       └── def456.../        # Symlinks to blobs</span></span></code></pre>\n<p>This deduplicates identical files across model versions and enables instant switching between revisions.</p>\n<hr />\n<h2 id=\"datasets-library-internals\">Datasets Library Internals</h2>\n<h3 id=\"apache-arrow-format-and-memory-mapping\">Apache Arrow Format and Memory Mapping</h3>\n<p>The <a href=\"https://huggingface.co/docs/datasets/en/about_arrow\">Datasets Arrow documentation</a> explains the storage layer:</p>\n<p>Arrow is a columnar memory format optimized for analytics. Key properties:</p>\n<ul>\n<li><strong>Columnar</strong>: All values for a column stored contiguously, enabling vectorized operations</li>\n<li><strong>Zero-copy</strong>: Data can be shared across processes without serialization</li>\n<li><strong>Memory-mapped</strong>: OS handles paging data in/out of RAM</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"imdb\"</span><span>, split</span><span>=</span><span>\"train\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Dataset is backed by Arrow file on disk</span></span>\n<span class=\"line\"><span>print</span><span>(dataset.cache_files)</span></span>\n<span class=\"line\"><span># [{'filename': '~/.cache/huggingface/datasets/imdb/.../train/dataset.arrow'}]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Access uses memory mapping - doesn't load into RAM</span></span>\n<span class=\"line\"><span>print</span><span>(dataset[</span><span>0</span><span>])</span><span>  # Pages in just this example</span></span></code></pre>\n<p>This lets you iterate over datasets larger than RAM. The OS virtual memory system handles what’s actually in physical memory.</p>\n<h3 id=\"datasetmap-parallelization\">Dataset.map() Parallelization</h3>\n<p>The <a href=\"https://discuss.huggingface.co/t/how-does-datasets-dataset-map-parallelize-data/36370\">map parallelization discussion</a> explains multiprocessing:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> process</span><span>(</span><span>examples</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> {</span><span>\"length\"</span><span>:</span><span> [</span><span>len</span><span>(t)</span><span> for</span><span> t </span><span>in</span><span> examples</span><span>[</span><span>\"text\"</span><span>]</span><span>]</span><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Single process</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> dataset</span><span>.</span><span>map</span><span>(process, batched</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Parallel processing</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> dataset</span><span>.</span><span>map</span><span>(process, batched</span><span>=</span><span>True</span><span>, num_proc</span><span>=</span><span>4</span><span>)</span></span></code></pre>\n<p>With <code>num_proc &gt; 1</code>:</p>\n<ol>\n<li>Dataset splits into <code>num_proc</code> shards</li>\n<li>Each shard processes in a separate worker</li>\n<li>Workers write results to temporary Arrow files</li>\n<li>Results concatenate into final dataset</li>\n</ol>\n<p>Requirements for parallelization:</p>\n<ul>\n<li>Function must be picklable (top-level functions work; lambdas and closures may not)</li>\n<li>Workers don’t share state</li>\n<li>Each worker loads its shard independently via memory mapping</li>\n</ul>\n<h3 id=\"streaming-datasets-and-shard-fetching\">Streaming Datasets and Shard Fetching</h3>\n<p>The <a href=\"https://huggingface.co/docs/datasets/stream\">streaming documentation</a> covers iterable datasets:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"HuggingFaceFW/fineweb\"</span><span>, split</span><span>=</span><span>\"train\"</span><span>, streaming</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> example </span><span>in</span><span> dataset</span><span>:</span></span>\n<span class=\"line\"><span>    process</span><span>(example)</span></span></code></pre>\n<p>Streaming datasets:</p>\n<ul>\n<li>Don’t download the full dataset</li>\n<li>Fetch data on-demand as you iterate</li>\n<li>Support transformations (map, filter) that apply lazily</li>\n</ul>\n<p>For sharded datasets, streaming fetches shards sequentially. Shuffling operates within a buffer:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Shuffles within buffer of 10000 examples</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> dataset</span><span>.</span><span>shuffle</span><span>(seed</span><span>=</span><span>42</span><span>, buffer_size</span><span>=</span><span>10000</span><span>)</span></span></code></pre>\n<p>To resume from a checkpoint:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Skip to position (fetches from start of current shard)</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> dataset</span><span>.</span><span>skip</span><span>(</span><span>1000</span><span>)</span></span></code></pre>\n<p>Resuming isn’t instant because it must re-read from the beginning of the current shard.</p>\n<h3 id=\"custom-dataset-loading-scripts\">Custom Dataset Loading Scripts</h3>\n<p>The <a href=\"https://huggingface.co/docs/datasets/v3.4.0/en/dataset_script\">dataset script documentation</a> shows how to write loaders:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> datasets</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> MyDataset</span><span>(</span><span>datasets</span><span>.</span><span>GeneratorBasedBuilder</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"My custom dataset.\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    VERSION </span><span>=</span><span> datasets</span><span>.</span><span>Version</span><span>(</span><span>\"1.0.0\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> _info</span><span>(</span><span>self</span><span>):</span></span>\n<span class=\"line\"><span>        return</span><span> datasets</span><span>.</span><span>DatasetInfo</span><span>(</span></span>\n<span class=\"line\"><span>            description</span><span>=</span><span>\"My dataset description\"</span><span>,</span></span>\n<span class=\"line\"><span>            features</span><span>=</span><span>datasets.</span><span>Features</span><span>({</span></span>\n<span class=\"line\"><span>                \"text\"</span><span>: datasets.</span><span>Value</span><span>(</span><span>\"string\"</span><span>),</span></span>\n<span class=\"line\"><span>                \"label\"</span><span>: datasets.</span><span>ClassLabel</span><span>(names</span><span>=</span><span>[</span><span>\"negative\"</span><span>, </span><span>\"positive\"</span><span>]),</span></span>\n<span class=\"line\"><span>            }),</span></span>\n<span class=\"line\"><span>        )</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> _split_generators</span><span>(</span><span>self</span><span>,</span><span> dl_manager</span><span>):</span></span>\n<span class=\"line\"><span>        # Download and extract data</span></span>\n<span class=\"line\"><span>        data_dir </span><span>=</span><span> dl_manager</span><span>.</span><span>download_and_extract</span><span>(</span><span>\"https://example.com/data.zip\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> [</span></span>\n<span class=\"line\"><span>            datasets</span><span>.</span><span>SplitGenerator</span><span>(</span></span>\n<span class=\"line\"><span>                name</span><span>=</span><span>datasets.Split.TRAIN,</span></span>\n<span class=\"line\"><span>                gen_kwargs</span><span>=</span><span>{</span><span>\"filepath\"</span><span>: </span><span>f</span><span>\"</span><span>{</span><span>data_dir</span><span>}</span><span>/train.jsonl\"</span><span>},</span></span>\n<span class=\"line\"><span>            ),</span></span>\n<span class=\"line\"><span>            datasets</span><span>.</span><span>SplitGenerator</span><span>(</span></span>\n<span class=\"line\"><span>                name</span><span>=</span><span>datasets.Split.TEST,</span></span>\n<span class=\"line\"><span>                gen_kwargs</span><span>=</span><span>{</span><span>\"filepath\"</span><span>: </span><span>f</span><span>\"</span><span>{</span><span>data_dir</span><span>}</span><span>/test.jsonl\"</span><span>},</span></span>\n<span class=\"line\"><span>            ),</span></span>\n<span class=\"line\"><span>        ]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> _generate_examples</span><span>(</span><span>self</span><span>,</span><span> filepath</span><span>):</span></span>\n<span class=\"line\"><span>        with</span><span> open</span><span>(filepath)</span><span> as</span><span> f</span><span>:</span></span>\n<span class=\"line\"><span>            for</span><span> idx</span><span>,</span><span> line </span><span>in</span><span> enumerate</span><span>(f):</span></span>\n<span class=\"line\"><span>                data </span><span>=</span><span> json</span><span>.</span><span>loads</span><span>(line)</span></span>\n<span class=\"line\"><span>                yield</span><span> idx</span><span>,</span><span> {</span></span>\n<span class=\"line\"><span>                    \"text\"</span><span>:</span><span> data</span><span>[</span><span>\"text\"</span><span>],</span></span>\n<span class=\"line\"><span>                    \"label\"</span><span>:</span><span> data</span><span>[</span><span>\"label\"</span><span>],</span></span>\n<span class=\"line\"><span>                }</span></span></code></pre>\n<p>For Arrow-based builders (better for large datasets):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>class</span><span> MyArrowDataset</span><span>(</span><span>datasets</span><span>.</span><span>ArrowBasedBuilder</span><span>):</span></span>\n<span class=\"line\"><span>    def</span><span> _generate_tables</span><span>(</span><span>self</span><span>,</span><span> filepath</span><span>):</span></span>\n<span class=\"line\"><span>        # Yield PyArrow tables instead of individual examples</span></span>\n<span class=\"line\"><span>        table </span><span>=</span><span> pq</span><span>.</span><span>read_table</span><span>(filepath)</span></span>\n<span class=\"line\"><span>        yield</span><span> 0</span><span>,</span><span> table</span></span></code></pre>\n<hr />\n<h2 id=\"the-diffusers-library\">The Diffusers Library</h2>\n<p>Diffusers handles diffusion models for image, video, and audio generation. It supports Stable Diffusion, SDXL, Flux, Kandinsky, and other architectures.</p>\n<h3 id=\"basic-image-generation\">Basic Image Generation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>from</span><span> diffusers </span><span>import</span><span> DiffusionPipeline</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load Stable Diffusion XL</span></span>\n<span class=\"line\"><span>pipe </span><span>=</span><span> DiffusionPipeline</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"stabilityai/stable-diffusion-xl-base-1.0\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>torch.float16,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>to</span><span>(</span><span>\"cuda\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate</span></span>\n<span class=\"line\"><span>image </span><span>=</span><span> pipe</span><span>(</span><span>\"A cat astronaut floating in space, digital art\"</span><span>).</span><span>images</span><span>[</span><span>0</span><span>]</span></span>\n<span class=\"line\"><span>image</span><span>.</span><span>save</span><span>(</span><span>\"cat_astronaut.png\"</span><span>)</span></span></code></pre>\n<h3 id=\"flux-models\">FLUX Models</h3>\n<p>FLUX is a 12 billion parameter rectified flow transformer from Black Forest Labs. It produces high-quality images with strong text rendering and prompt adherence.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>from</span><span> diffusers </span><span>import</span><span> FluxPipeline</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>pipe </span><span>=</span><span> FluxPipeline</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"black-forest-labs/FLUX.1-schnell\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>torch.bfloat16</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>enable_model_cpu_offload</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># FLUX.1-schnell generates in 1-4 steps</span></span>\n<span class=\"line\"><span>image </span><span>=</span><span> pipe</span><span>(</span></span>\n<span class=\"line\"><span>    \"A photo of a mountain lake at sunset\"</span><span>,</span></span>\n<span class=\"line\"><span>    guidance_scale</span><span>=</span><span>0.0</span><span>,</span></span>\n<span class=\"line\"><span>    num_inference_steps</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>).</span><span>images</span><span>[</span><span>0</span><span>]</span></span>\n<span class=\"line\"><span>image</span><span>.</span><span>save</span><span>(</span><span>\"lake.png\"</span><span>)</span></span></code></pre>\n<p>FLUX.1-schnell is distilled for speed. FLUX.1-dev offers higher quality with more steps. FLUX.1-Kontext handles image editing with text instructions.</p>\n<h3 id=\"memory-optimization\">Memory Optimization</h3>\n<p>Diffusers provides multiple strategies for running on limited VRAM:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># CPU offloading - moves components to CPU when not in use</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>enable_model_cpu_offload</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Sequential CPU offloading - more aggressive, slower</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>enable_sequential_cpu_offload</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Attention slicing - trades compute for memory</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>enable_attention_slicing</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># VAE tiling - generate high-res without OOM</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>enable_vae_tiling</span><span>()</span></span></code></pre>\n<h3 id=\"lora-and-adapters\">LoRA and Adapters</h3>\n<p>Load style or concept LoRAs:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>pipe</span><span>.</span><span>load_lora_weights</span><span>(</span><span>\"username/style-lora\"</span><span>, weight_name</span><span>=</span><span>\"lora.safetensors\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Adjust LoRA strength</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>fuse_lora</span><span>(lora_scale</span><span>=</span><span>0.8</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Unload</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>unfuse_lora</span><span>()</span></span>\n<span class=\"line\"><span>pipe</span><span>.</span><span>unload_lora_weights</span><span>()</span></span></code></pre>\n<hr />\n<h2 id=\"hugging-face-spaces\">Hugging Face Spaces</h2>\n<p>Spaces are web applications hosted on Hugging Face. The key insight: Spaces are Git repositories. You can clone them, run them locally, and modify them like any codebase.</p>\n<h3 id=\"cloning-and-running-locally\">Cloning and Running Locally</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Clone a Space</span></span>\n<span class=\"line\"><span>git</span><span> clone</span><span> https://huggingface.co/spaces/username/my-space</span></span>\n<span class=\"line\"><span>cd</span><span> my-space</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Install dependencies</span></span>\n<span class=\"line\"><span>python</span><span> -m</span><span> venv</span><span> venv</span></span>\n<span class=\"line\"><span>source</span><span> venv/bin/activate</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> -r</span><span> requirements.txt</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run locally</span></span>\n<span class=\"line\"><span>python</span><span> app.py</span><span>  # or gradio app.py for Gradio apps</span></span></code></pre>\n<p>Spaces typically use Gradio or Streamlit for the frontend. The same code runs identically on your machine and on Hugging Face’s servers.</p>\n<h3 id=\"gradio-space-example\">Gradio Space Example</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> gradio </span><span>as</span><span> gr</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> pipeline</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>classifier </span><span>=</span><span> pipeline</span><span>(</span><span>\"sentiment-analysis\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> classify</span><span>(</span><span>text</span><span>):</span></span>\n<span class=\"line\"><span>    result </span><span>=</span><span> classifier</span><span>(text)</span><span>[</span><span>0</span><span>]</span></span>\n<span class=\"line\"><span>    return</span><span> f</span><span>\"</span><span>{</span><span>result</span><span>[</span><span>'label'</span><span>]</span><span>}</span><span>: </span><span>{</span><span>result</span><span>[</span><span>'score'</span><span>]</span><span>:.4f</span><span>}</span><span>\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>demo </span><span>=</span><span> gr</span><span>.</span><span>Interface</span><span>(fn</span><span>=</span><span>classify, inputs</span><span>=</span><span>\"text\"</span><span>, outputs</span><span>=</span><span>\"text\"</span><span>)</span></span>\n<span class=\"line\"><span>demo</span><span>.</span><span>launch</span><span>()</span></span></code></pre>\n<h3 id=\"syncing-with-github\">Syncing with GitHub</h3>\n<p>Configure a GitHub Action to push changes to both GitHub and Hugging Face:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>name</span><span>:</span><span> Sync to Hugging Face</span></span>\n<span class=\"line\"><span>on</span><span>:</span></span>\n<span class=\"line\"><span>  push</span><span>:</span></span>\n<span class=\"line\"><span>    branches</span><span>:</span><span> [</span><span>main</span><span>]</span></span>\n<span class=\"line\"><span>jobs</span><span>:</span></span>\n<span class=\"line\"><span>  sync</span><span>:</span></span>\n<span class=\"line\"><span>    runs-on</span><span>:</span><span> ubuntu-latest</span></span>\n<span class=\"line\"><span>    steps</span><span>:</span></span>\n<span class=\"line\"><span>      - </span><span>uses</span><span>:</span><span> actions/checkout@v3</span></span>\n<span class=\"line\"><span>      - </span><span>name</span><span>:</span><span> Push to HF</span></span>\n<span class=\"line\"><span>        run</span><span>:</span><span> |</span></span>\n<span class=\"line\"><span>          git push https://user:${{ secrets.HF_TOKEN }}@huggingface.co/spaces/user/space main</span></span></code></pre>\n<p>Spaces support Docker for custom environments, persistent storage for databases, and GPU runtimes for ML inference.</p>\n<hr />\n<h2 id=\"other-hugging-face-libraries\">Other Hugging Face Libraries</h2>\n<h3 id=\"accelerate\">Accelerate</h3>\n<p>Accelerate lets you run the same PyTorch training code on any hardware configuration: single GPU, multi-GPU, TPU, or distributed clusters. Add a few lines to existing code:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> accelerate </span><span>import</span><span> Accelerator</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>accelerator </span><span>=</span><span> Accelerator</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Wrap model, optimizer, dataloader</span></span>\n<span class=\"line\"><span>model</span><span>,</span><span> optimizer</span><span>,</span><span> dataloader </span><span>=</span><span> accelerator</span><span>.</span><span>prepare</span><span>(model, optimizer, dataloader)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> batch </span><span>in</span><span> dataloader</span><span>:</span></span>\n<span class=\"line\"><span>    outputs </span><span>=</span><span> model</span><span>(</span><span>**</span><span>batch)</span></span>\n<span class=\"line\"><span>    loss </span><span>=</span><span> outputs</span><span>.</span><span>loss</span></span>\n<span class=\"line\"><span>    accelerator</span><span>.</span><span>backward</span><span>(loss)</span></span>\n<span class=\"line\"><span>    optimizer</span><span>.</span><span>step</span><span>()</span></span>\n<span class=\"line\"><span>    optimizer</span><span>.</span><span>zero_grad</span><span>()</span></span></code></pre>\n<p>Configure distributed training with the CLI:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>accelerate</span><span> config</span><span>  # Interactive setup</span></span>\n<span class=\"line\"><span>accelerate</span><span> launch</span><span> train.py</span><span>  # Launch with config</span></span></code></pre>\n<p>Accelerate handles mixed precision (fp16, bf16, fp8), DeepSpeed integration, and FSDP for sharding large models.</p>\n<h3 id=\"peft\">PEFT</h3>\n<p>PEFT implements parameter-efficient fine-tuning methods. LoRA is the most popular: freeze the base model and train small adapter matrices.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> peft </span><span>import</span><span> LoraConfig</span><span>,</span><span> get_peft_model</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>lora_config </span><span>=</span><span> LoraConfig</span><span>(</span></span>\n<span class=\"line\"><span>    r</span><span>=</span><span>16</span><span>,                    </span><span># Rank</span></span>\n<span class=\"line\"><span>    lora_alpha</span><span>=</span><span>32</span><span>,           </span><span># Scaling</span></span>\n<span class=\"line\"><span>    target_modules</span><span>=</span><span>\"all-linear\"</span><span>,  </span><span># Apply to all linear layers</span></span>\n<span class=\"line\"><span>    lora_dropout</span><span>=</span><span>0.05</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>peft_model </span><span>=</span><span> get_peft_model</span><span>(model, lora_config)</span></span>\n<span class=\"line\"><span>peft_model</span><span>.</span><span>print_trainable_parameters</span><span>()</span></span>\n<span class=\"line\"><span># trainable params: 41,943,040 || all params: 8,030,261,248 || trainable%: 0.52</span></span></code></pre>\n<p>A 7B model drops from training billions of parameters to millions. The base model stays frozen; only the adapter trains. Merge adapters back into the base model for inference or keep them separate to swap styles.</p>\n<p>PEFT supports LoRA variants like DoRA (magnitude and direction decomposition) and methods beyond LoRA: IA3, AdaLoRA, and prompt tuning.</p>\n<h3 id=\"trl\">TRL</h3>\n<p>TRL provides trainers for post-training: supervised fine-tuning, RLHF, and preference optimization.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> trl </span><span>import</span><span> SFTTrainer</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> AutoTokenizer</span></span>\n<span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.1-8B\"</span><span>)</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"your-dataset\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> SFTTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>model,</span></span>\n<span class=\"line\"><span>    train_dataset</span><span>=</span><span>dataset,</span></span>\n<span class=\"line\"><span>    tokenizer</span><span>=</span><span>tokenizer,</span></span>\n<span class=\"line\"><span>    max_seq_length</span><span>=</span><span>2048</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>train</span><span>()</span></span></code></pre>\n<p>For preference optimization, DPOTrainer implements Direct Preference Optimization:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> trl </span><span>import</span><span> DPOTrainer</span><span>,</span><span> DPOConfig</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>config </span><span>=</span><span> DPOConfig</span><span>(</span></span>\n<span class=\"line\"><span>    beta</span><span>=</span><span>0.1</span><span>,</span></span>\n<span class=\"line\"><span>    output_dir</span><span>=</span><span>\"dpo_model\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> DPOTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>model,</span></span>\n<span class=\"line\"><span>    ref_model</span><span>=</span><span>ref_model,</span></span>\n<span class=\"line\"><span>    args</span><span>=</span><span>config,</span></span>\n<span class=\"line\"><span>    train_dataset</span><span>=</span><span>preference_dataset,  </span><span># Has 'chosen' and 'rejected' columns</span></span>\n<span class=\"line\"><span>    tokenizer</span><span>=</span><span>tokenizer,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>GRPOTrainer implements Group Relative Policy Optimization, the algorithm behind DeepSeek R1. It avoids needing a separate reward model by comparing outputs within groups.</p>\n<h3 id=\"optimum\">Optimum</h3>\n<p>Optimum provides hardware-specific optimizations. Export to ONNX, quantize, and run on specialized accelerators:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> optimum</span><span>.</span><span>onnxruntime </span><span>import</span><span> ORTModelForSequenceClassification</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load and export to ONNX</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> ORTModelForSequenceClassification</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"distilbert-base-uncased-finetuned-sst-2-english\"</span><span>,</span></span>\n<span class=\"line\"><span>    export</span><span>=</span><span>True</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or load existing ONNX model</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> ORTModelForSequenceClassification</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"onnx_model_dir\"</span><span>)</span></span></code></pre>\n<p>Optimum integrates with:</p>\n<ul>\n<li>Intel Gaudi accelerators via optimum-habana</li>\n<li>AWS Trainium via optimum-neuron</li>\n<li>Intel CPUs via optimum-intel</li>\n<li>AMD GPUs via optimum-amd</li>\n</ul>\n<p>The optimum-benchmark tool compares performance across backends and quantization schemes.</p>\n<hr />\n<h2 id=\"practical-code-examples\">Practical Code Examples</h2>\n<h3 id=\"complete-fine-tuning-pipeline\">Complete Fine-Tuning Pipeline</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> (</span></span>\n<span class=\"line\"><span>    AutoModelForCausalLM</span><span>,</span></span>\n<span class=\"line\"><span>    AutoTokenizer</span><span>,</span></span>\n<span class=\"line\"><span>    TrainingArguments</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>from</span><span> peft </span><span>import</span><span> LoraConfig</span></span>\n<span class=\"line\"><span>from</span><span> trl </span><span>import</span><span> SFTTrainer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model with 4-bit quantization</span></span>\n<span class=\"line\"><span>model_id </span><span>=</span><span> \"meta-llama/Llama-3.1-8B\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    model_id,</span></span>\n<span class=\"line\"><span>    load_in_4bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(model_id)</span></span>\n<span class=\"line\"><span>tokenizer</span><span>.</span><span>pad_token </span><span>=</span><span> tokenizer</span><span>.</span><span>eos_token</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load dataset</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"databricks/databricks-dolly-15k\"</span><span>, split</span><span>=</span><span>\"train\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># LoRA config for QLoRA</span></span>\n<span class=\"line\"><span>peft_config </span><span>=</span><span> LoraConfig</span><span>(</span></span>\n<span class=\"line\"><span>    r</span><span>=</span><span>16</span><span>,</span></span>\n<span class=\"line\"><span>    lora_alpha</span><span>=</span><span>32</span><span>,</span></span>\n<span class=\"line\"><span>    target_modules</span><span>=</span><span>\"all-linear\"</span><span>,</span></span>\n<span class=\"line\"><span>    lora_dropout</span><span>=</span><span>0.05</span><span>,</span></span>\n<span class=\"line\"><span>    task_type</span><span>=</span><span>\"CAUSAL_LM\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Training arguments</span></span>\n<span class=\"line\"><span>training_args </span><span>=</span><span> TrainingArguments</span><span>(</span></span>\n<span class=\"line\"><span>    output_dir</span><span>=</span><span>\"./llama-finetuned\"</span><span>,</span></span>\n<span class=\"line\"><span>    num_train_epochs</span><span>=</span><span>1</span><span>,</span></span>\n<span class=\"line\"><span>    per_device_train_batch_size</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>    gradient_accumulation_steps</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>    learning_rate</span><span>=</span><span>2e-4</span><span>,</span></span>\n<span class=\"line\"><span>    bf16</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    logging_steps</span><span>=</span><span>10</span><span>,</span></span>\n<span class=\"line\"><span>    save_strategy</span><span>=</span><span>\"epoch\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Train</span></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> SFTTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>model,</span></span>\n<span class=\"line\"><span>    train_dataset</span><span>=</span><span>dataset,</span></span>\n<span class=\"line\"><span>    peft_config</span><span>=</span><span>peft_config,</span></span>\n<span class=\"line\"><span>    args</span><span>=</span><span>training_args,</span></span>\n<span class=\"line\"><span>    tokenizer</span><span>=</span><span>tokenizer,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>train</span><span>()</span></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>save_model</span><span>()</span></span></code></pre>\n<h3 id=\"building-a-rag-system\">Building a RAG System</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> sentence_transformers </span><span>import</span><span> SentenceTransformer</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> pipeline</span></span>\n<span class=\"line\"><span>import</span><span> numpy </span><span>as</span><span> np</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Embedding model</span></span>\n<span class=\"line\"><span>embedder </span><span>=</span><span> SentenceTransformer</span><span>(</span><span>\"BAAI/bge-base-en-v1.5\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generator</span></span>\n<span class=\"line\"><span>generator </span><span>=</span><span> pipeline</span><span>(</span></span>\n<span class=\"line\"><span>    \"text-generation\"</span><span>,</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.2-3B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Sample documents</span></span>\n<span class=\"line\"><span>documents </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>    \"Python was created by Guido van Rossum and released in 1991.\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"JavaScript was created by Brendan Eich in 1995.\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"Rust was developed by Mozilla and released in 2015.\"</span><span>,</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Embed documents</span></span>\n<span class=\"line\"><span>doc_embeddings </span><span>=</span><span> embedder</span><span>.</span><span>encode</span><span>(documents, normalize_embeddings</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> retrieve_and_generate</span><span>(</span><span>query</span><span>:</span><span> str</span><span>,</span><span> top_k</span><span>:</span><span> int</span><span> =</span><span> 2</span><span>) </span><span>-&gt;</span><span> str</span><span>:</span></span>\n<span class=\"line\"><span>    # Embed query</span></span>\n<span class=\"line\"><span>    query_embedding </span><span>=</span><span> embedder</span><span>.</span><span>encode</span><span>([query], normalize_embeddings</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Find similar documents</span></span>\n<span class=\"line\"><span>    scores </span><span>=</span><span> np</span><span>.</span><span>dot</span><span>(doc_embeddings, query_embedding.T).</span><span>flatten</span><span>()</span></span>\n<span class=\"line\"><span>    top_indices </span><span>=</span><span> np</span><span>.</span><span>argsort</span><span>(scores)</span><span>[</span><span>-</span><span>top_k</span><span>:</span><span>][</span><span>::</span><span>-</span><span>1</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    context </span><span>=</span><span> \"\\n\"</span><span>.</span><span>join</span><span>([documents[i] </span><span>for</span><span> i </span><span>in</span><span> top_indices])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Generate answer</span></span>\n<span class=\"line\"><span>    prompt </span><span>=</span><span> f</span><span>\"\"\"Context: </span><span>{</span><span>context</span><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>Question: </span><span>{</span><span>query</span><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>Answer based on the context above:\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    response </span><span>=</span><span> generator</span><span>(prompt, max_new_tokens</span><span>=</span><span>100</span><span>, do_sample</span><span>=</span><span>False</span><span>)</span></span>\n<span class=\"line\"><span>    return</span><span> response</span><span>[</span><span>0</span><span>]</span><span>[</span><span>\"generated_text\"</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>answer </span><span>=</span><span> retrieve_and_generate</span><span>(</span><span>\"When was Python created?\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(answer)</span></span></code></pre>\n<h3 id=\"multi-modal-pipeline-with-flux\">Multi-Modal Pipeline with FLUX</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>from</span><span> diffusers </span><span>import</span><span> FluxPipeline</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> pipeline</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Image generation</span></span>\n<span class=\"line\"><span>flux </span><span>=</span><span> FluxPipeline</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"black-forest-labs/FLUX.1-schnell\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>torch.bfloat16</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>flux</span><span>.</span><span>enable_model_cpu_offload</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Image captioning</span></span>\n<span class=\"line\"><span>captioner </span><span>=</span><span> pipeline</span><span>(</span></span>\n<span class=\"line\"><span>    \"image-to-text\"</span><span>,</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"Salesforce/blip2-opt-2.7b\"</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate image</span></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> \"A cozy cabin in the woods during autumn, warm lighting\"</span></span>\n<span class=\"line\"><span>image </span><span>=</span><span> flux</span><span>(prompt, num_inference_steps</span><span>=</span><span>4</span><span>).</span><span>images</span><span>[</span><span>0</span><span>]</span></span>\n<span class=\"line\"><span>image</span><span>.</span><span>save</span><span>(</span><span>\"cabin.png\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Caption the generated image</span></span>\n<span class=\"line\"><span>caption </span><span>=</span><span> captioner</span><span>(image)</span></span>\n<span class=\"line\"><span>print</span><span>(</span><span>f</span><span>\"Generated caption: </span><span>{</span><span>caption[</span><span>0</span><span>][</span><span>'generated_text'</span><span>]</span><span>}</span><span>\"</span><span>)</span></span></code></pre>\n<h3 id=\"distributed-training-with-accelerate\">Distributed Training with Accelerate</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># train.py</span></span>\n<span class=\"line\"><span>from</span><span> accelerate </span><span>import</span><span> Accelerator</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForSequenceClassification</span><span>,</span><span> AutoTokenizer</span></span>\n<span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>utils</span><span>.</span><span>data </span><span>import</span><span> DataLoader</span></span>\n<span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>accelerator </span><span>=</span><span> Accelerator</span><span>(mixed_precision</span><span>=</span><span>\"bf16\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model and tokenizer</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForSequenceClassification</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"bert-base-uncased\"</span><span>,</span></span>\n<span class=\"line\"><span>    num_labels</span><span>=</span><span>2</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"bert-base-uncased\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load and tokenize dataset</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"imdb\"</span><span>, split</span><span>=</span><span>\"train[:1000]\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> tokenize</span><span>(</span><span>examples</span><span>):</span></span>\n<span class=\"line\"><span>    return</span><span> tokenizer</span><span>(examples[</span><span>\"text\"</span><span>], truncation</span><span>=</span><span>True</span><span>, padding</span><span>=</span><span>\"max_length\"</span><span>, max_length</span><span>=</span><span>512</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> dataset</span><span>.</span><span>map</span><span>(tokenize, batched</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>dataset</span><span>.</span><span>set_format</span><span>(</span><span>\"torch\"</span><span>, columns</span><span>=</span><span>[</span><span>\"input_ids\"</span><span>, </span><span>\"attention_mask\"</span><span>, </span><span>\"label\"</span><span>])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>dataloader </span><span>=</span><span> DataLoader</span><span>(dataset, batch_size</span><span>=</span><span>8</span><span>, shuffle</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>optimizer </span><span>=</span><span> torch</span><span>.</span><span>optim</span><span>.</span><span>AdamW</span><span>(model.</span><span>parameters</span><span>(), lr</span><span>=</span><span>2e-5</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Prepare for distributed training</span></span>\n<span class=\"line\"><span>model</span><span>,</span><span> optimizer</span><span>,</span><span> dataloader </span><span>=</span><span> accelerator</span><span>.</span><span>prepare</span><span>(model, optimizer, dataloader)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Training loop</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>train</span><span>()</span></span>\n<span class=\"line\"><span>for</span><span> epoch </span><span>in</span><span> range</span><span>(</span><span>3</span><span>):</span></span>\n<span class=\"line\"><span>    for</span><span> batch </span><span>in</span><span> dataloader</span><span>:</span></span>\n<span class=\"line\"><span>        outputs </span><span>=</span><span> model</span><span>(</span></span>\n<span class=\"line\"><span>            input_ids</span><span>=</span><span>batch[</span><span>\"input_ids\"</span><span>],</span></span>\n<span class=\"line\"><span>            attention_mask</span><span>=</span><span>batch[</span><span>\"attention_mask\"</span><span>],</span></span>\n<span class=\"line\"><span>            labels</span><span>=</span><span>batch[</span><span>\"label\"</span><span>],</span></span>\n<span class=\"line\"><span>        )</span></span>\n<span class=\"line\"><span>        accelerator</span><span>.</span><span>backward</span><span>(outputs.loss)</span></span>\n<span class=\"line\"><span>        optimizer</span><span>.</span><span>step</span><span>()</span></span>\n<span class=\"line\"><span>        optimizer</span><span>.</span><span>zero_grad</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    accelerator</span><span>.</span><span>print</span><span>(</span><span>f</span><span>\"Epoch </span><span>{</span><span>epoch</span><span>}</span><span> complete\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>accelerator</span><span>.</span><span>save_model</span><span>(model, </span><span>\"trained_model\"</span><span>)</span></span></code></pre>\n<p>Launch with:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>accelerate</span><span> launch</span><span> --multi_gpu</span><span> --num_processes</span><span> 4</span><span> train.py</span></span></code></pre>\n<hr />\n<h2 id=\"references\">References</h2>\n<ul>\n<li><a href=\"https://huggingface.co/docs/hub/en/index\">Hugging Face Hub Documentation</a></li>\n<li><a href=\"https://en.wikipedia.org/wiki/Hugging_Face\">Hugging Face Wikipedia</a></li>\n<li><a href=\"https://fueler.io/blog/hugging-face-usage-revenue-valuation-growth-statistics\">Hugging Face 2026 Statistics - Fueler</a></li>\n<li><a href=\"https://github.com/huggingface/transformers\">Transformers GitHub Repository</a></li>\n<li><a href=\"https://huggingface.co/blog/tokenizers\">Transformers v5 Tokenization Blog</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/main_classes/pipelines\">Pipeline Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/en/model_doc/auto\">Auto Classes Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/main_classes/model\">Models Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/en/generation_strategies\">Generation Strategies</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/en/kv_cache\">KV Cache Strategies</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/tokenizer_summary\">Tokenizer Summary</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/en/fast_tokenizers\">Fast Tokenizers Guide</a></li>\n<li><a href=\"https://huggingface.co/docs/tokenizers/en/quicktour\">Tokenizers Quicktour</a></li>\n<li><a href=\"https://github.com/huggingface/tokenizers\">Tokenizers GitHub Repository</a></li>\n<li><a href=\"https://huggingface.co/docs/accelerate/en/usage_guides/big_modeling\">Big Model Inference Guide</a></li>\n<li><a href=\"https://huggingface.co/docs/accelerate/en/concept_guides/big_model_inference\">Loading Big Models Concept Guide</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/en/quantization/bitsandbytes\">Bitsandbytes Quantization Guide</a></li>\n<li><a href=\"https://huggingface.co/blog/4bit-transformers-bitsandbytes\">Making LLMs Accessible with 4-bit Quantization</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/en/main_classes/trainer\">Trainer Documentation</a></li>\n<li><a href=\"https://huggingface.co/blog/gradient_accumulation\">Fixing Gradient Accumulation Blog</a></li>\n<li><a href=\"https://docs.pytorch.org/docs/stable/amp.html\">PyTorch AMP Documentation</a></li>\n<li><a href=\"https://pytorch.org/blog/what-every-user-should-know-about-mixed-precision-training-in-pytorch/\">PyTorch Mixed Precision Blog</a></li>\n<li><a href=\"https://huggingface.co/docs/transformers/en/peft\">PEFT Integration Guide</a></li>\n<li><a href=\"https://github.com/huggingface/peft\">PEFT GitHub Repository</a></li>\n<li><a href=\"https://huggingface.co/docs/hub/en/storage-backends\">Storage Backends Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/hub/en/model-cards\">Model Cards Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/huggingface_hub/package_reference/hf_api\">huggingface_hub Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/datasets/en/about_arrow\">Datasets Arrow Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/datasets/stream\">Streaming Datasets</a></li>\n<li><a href=\"https://huggingface.co/docs/datasets/v3.4.0/en/dataset_script\">Dataset Loading Scripts</a></li>\n<li><a href=\"https://github.com/huggingface/diffusers\">Diffusers GitHub Repository</a></li>\n<li><a href=\"https://huggingface.co/black-forest-labs/FLUX.1-dev\">FLUX.1-dev Model Card</a></li>\n<li><a href=\"https://huggingface.co/docs/safetensors/speed\">Safetensors Speed Comparison</a></li>\n<li><a href=\"https://deepwiki.com/huggingface/safetensors\">Safetensors DeepWiki</a></li>\n<li><a href=\"https://github.com/ggml-org/llama.cpp\">llama.cpp GitHub Repository</a></li>\n<li><a href=\"https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md\">GGUF Quantization Guide</a></li>\n<li><a href=\"https://apxml.com/courses/practical-llm-quantization/chapter-5-quantization-formats-tooling/gguf-format\">GGUF Format Explained</a></li>\n<li><a href=\"https://newsletter.maartengrootendorst.com/p/which-quantization-method-is-right\">Quantization Methods Comparison - Maarten Grootendorst</a></li>\n<li><a href=\"https://localaimaster.com/blog/quantization-explained\">AWQ vs GPTQ vs GGUF - LocalAI Master</a></li>\n<li><a href=\"https://docs.jarvislabs.ai/blog/vllm-quantization-complete-guide-benchmarks\">vLLM Quantization Guide</a></li>\n<li><a href=\"https://onnx.ai/\">ONNX Official Site</a></li>\n<li><a href=\"https://onnxruntime.ai/docs/\">ONNX Runtime Documentation</a></li>\n<li><a href=\"https://github.com/NVIDIA/TensorRT-LLM\">TensorRT-LLM GitHub Repository</a></li>\n<li><a href=\"https://nvidia.github.io/TensorRT-LLM/overview.html\">TensorRT-LLM Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/hub/en/datasets-streaming\">Datasets Streaming Documentation</a></li>\n<li><a href=\"https://huggingface.co/blog/cfahlgren1/intro-to-parquet-format\">Parquet Introduction - Hugging Face Blog</a></li>\n<li><a href=\"https://github.com/huggingface/evaluate\">Evaluate Library GitHub</a></li>\n<li><a href=\"https://huggingface.co/docs/peft/en/package_reference/lora\">LoRA PEFT Documentation</a></li>\n<li><a href=\"https://www.philschmid.de/fine-tune-llms-in-2025\">Fine-tuning LLMs in 2025 - Phil Schmid</a></li>\n<li><a href=\"https://github.com/huggingface/accelerate\">Accelerate GitHub Repository</a></li>\n<li><a href=\"https://huggingface.co/blog/accelerate-v1\">Accelerate 1.0.0 Blog Post</a></li>\n<li><a href=\"https://huggingface.co/docs/trl/en/index\">TRL Documentation</a></li>\n<li><a href=\"https://github.com/huggingface/trl\">TRL GitHub Repository</a></li>\n<li><a href=\"https://github.com/huggingface/optimum\">Optimum GitHub Repository</a></li>\n<li><a href=\"https://huggingface.co/docs/optimum/en/index\">Optimum Documentation</a></li>\n<li><a href=\"https://huggingface.co/docs/hub/spaces-overview\">Spaces Overview</a></li>\n<li><a href=\"https://huggingface.co/docs/hub/repositories-getting-started\">Getting Started with Repositories</a></li>\n<li><a href=\"https://codesignal.com/learn/courses/2-modern-tokenization-techniques-for-ai-llms/lessons/comparing-bpe-wordpiece-and-sentencepiece-in-nlp\">Tokenization Algorithms Comparison - CodeSignal</a></li>\n<li><a href=\"https://huggingface.co/blog/kv-cache-quantization\">KV Cache Quantization Blog</a></li>\n</ul>",
            "url": "https://blog.ecitis.org/huggingface-ecosystem-guide/",
            "title": "The Complete Hugging Face Ecosystem Guide",
            "summary": "Navigate the Hub, Transformers, Datasets, Accelerate, PEFT, model formats, and deployment tools that make up Hugging Face.",
            "image": "https://blog.ecitis.org/open-graph/huggingface-ecosystem-guide.png",
            "date_modified": "2026-01-20T00:00:00.000Z",
            "date_published": "2026-01-20T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Ecosystems & Tooling",
                "Hugging Face",
                "Transformers",
                "Open source"
            ]
        },
        {
            "id": "https://blog.ecitis.org/local-ai-inferencing-techniques/",
            "content_html": "<h2 id=\"why-local-inferencing-matters\">Why local inferencing matters</h2>\n<p>Running AI models locally instead of through cloud APIs addresses four problems at once: privacy, latency, cost, and availability.</p>\n<p><strong>Privacy</strong>: Cloud-based GenAI tools exposed approximately 3 million sensitive records per organization in the first half of 2025, according to data from Microsoft Copilot deployments [1]. Keeping inference on-device shrinks attack surface and regulatory footprint. For workflows in legal, health, and finance, avoiding data transmission to external servers is often the simplest path to compliance with GDPR and HIPAA.</p>\n<p><strong>Latency</strong>: Local GPU inference delivers sub-50ms latency versus 100-500ms for cloud roundtrips. Cactus, a Y Combinator-backed startup, demonstrated sub-50ms time-to-first-token for on-device inference [2]. TinyML solutions can achieve latencies as low as 0-5ms for inference tasks, compared to 10-500ms for cloud IoT solutions [3].</p>\n<p><strong>Cost</strong>: Local inference eliminates per-token API costs. Once you own the hardware, inference is free. IDC and Gartner predict that by 2027, over 60% of all AI inference processes will happen locally rather than in the cloud [4].</p>\n<p><strong>Offline capability</strong>: Local inference removes network hops and eliminates the need for a constant internet connection. Apps keep working on factory floors, in hospitals, and in low-connectivity environments.</p>\n<hr />\n<h2 id=\"reap-router-weighted-expert-activation-pruning\">REAP: Router-weighted Expert Activation Pruning</h2>\n<p>REAP is a one-shot compression method for Mixture-of-Experts (MoE) language models developed by Cerebras Research. It removes up to 50% of experts from models as large as 1 trillion parameters while largely maintaining baseline model quality [5].</p>\n<h3 id=\"the-problem-reap-solves\">The problem REAP solves</h3>\n<p>MoE models like Mixtral, DeepSeek, and Qwen achieve high capability by routing tokens to different expert subnetworks. This architecture creates redundancy: not all experts contribute equally to every inference. REAP exploits this by identifying and removing the least important experts.</p>\n<h3 id=\"how-reap-works\">How REAP works</h3>\n<p>REAP selects experts to prune based on a saliency criterion that considers two factors:</p>\n<ol>\n<li><strong>Router gate values</strong>: How frequently and strongly the router activates each expert</li>\n<li><strong>Expert activation norms</strong>: The magnitude of each expert’s output contributions</li>\n</ol>\n<p>By combining these factors, REAP identifies experts that are both rarely used and have little impact when they are used. The method is one-shot, meaning it requires no fine-tuning after pruning.</p>\n<p>The key insight from the REAP paper is that expert pruning outperforms expert merging for generative tasks. The researchers proved that merging introduces an irreducible error by causing a “functional subspace collapse” due to the loss of the router’s independent, input-dependent control over experts [6].</p>\n<h3 id=\"performance-results\">Performance results</h3>\n<p>On the Qwen3-480B-Coder-FP8 model, REAP at 50% pruning retains:</p>\n<ul>\n<li>97.6% of baseline non-agentic coding ability</li>\n<li>96.7% on the agentic SWE-Bench benchmark</li>\n</ul>\n<p>The method achieves near-lossless compression on code generation and tool-calling tasks with Qwen3-Coder-480B and Kimi-K2, even after pruning 50% of experts [5].</p>\n<h3 id=\"implementation\">Implementation</h3>\n<p>The official REAP implementation is available at <a href=\"https://github.com/CerebrasResearch/reap\">github.com/CerebrasResearch/reap</a>. Pruned models are available on HuggingFace in the Cerebras REAP collection, including models like <code>cerebras/Kimi-Linear-REAP-35B-A3B-Instruct</code> and <code>cerebras/MiniMax-M2-REAP-172B-A10B</code> [7].</p>\n<hr />\n<h2 id=\"speculative-decoding\">Speculative decoding</h2>\n<p>Speculative decoding accelerates LLM inference by predicting and verifying multiple tokens simultaneously. The technique uses a smaller draft model to propose tokens and a larger target model to verify them in parallel.</p>\n<h3 id=\"core-mechanism\">Core mechanism</h3>\n<p>The draft-target approach works as follows:</p>\n<ol>\n<li>A small draft model (10-50x smaller than target) generates K candidate tokens autoregressively</li>\n<li>The target model verifies those proposals in a single forward pass</li>\n<li>Rejection sampling determines which tokens to accept based on probability distributions</li>\n<li>The target accepts the longest prefix that matches its own predictions and continues from there</li>\n</ol>\n<p>Compared with standard autoregressive decoding, which produces one token per pass, speculative decoding generates multiple tokens at once, cutting latency without any impact on accuracy [8].</p>\n<h3 id=\"draft-model-techniques\">Draft model techniques</h3>\n<p><strong>Traditional draft models</strong>: The draft model must be significantly smaller (10-50x) to achieve inference acceleration. The speedup ratio increases as the target model size increases [9].</p>\n<p><strong>EAGLE (Extrapolation Algorithm for Greater Language-model Efficiency)</strong>: EAGLE-3 uses a lightweight autoregressive prediction head attached to the target model’s internal layers to generate candidate tokens, eliminating the need for a separate draft model [10].</p>\n<p><strong>N-gram and Suffix Decoding</strong>: Suffix Decoding generates draft tokens by pattern-matching using the last n generated tokens against both the prompt and previous generations, using frequency counts to propose the most likely continuations [10].</p>\n<h3 id=\"implementation-example\">Implementation example</h3>\n<p>With vLLM, speculative decoding can be enabled via configuration:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span><span>,</span><span> SamplingParams</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-70B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    speculative_model</span><span>=</span><span>\"meta-llama/Llama-3.2-1B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    num_speculative_tokens</span><span>=</span><span>5</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>sampling_params </span><span>=</span><span> SamplingParams</span><span>(temperature</span><span>=</span><span>0.7</span><span>, max_tokens</span><span>=</span><span>256</span><span>)</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> llm</span><span>.</span><span>generate</span><span>([</span><span>\"Explain transformers:\"</span><span>], sampling_params)</span></span></code></pre>\n<h3 id=\"performance\">Performance</h3>\n<p>SpecForge’s Llama 4 Maverick draft model achieves a 2.18x speedup on MT-Bench, while the Scout variant delivers a 2.0x acceleration [11]. Acceptance rate is the critical metric: high acceptance means more tokens accepted per round, fewer target model forward passes, lower latency, and better GPU utilization.</p>\n<hr />\n<h2 id=\"quantization-techniques\">Quantization techniques</h2>\n<p>Quantization reduces the precision of model weights, trading a small amount of accuracy for large memory and speed gains.</p>\n<h3 id=\"post-training-quantization-vs-quantization-aware-training\">Post-training quantization vs quantization-aware training</h3>\n<p><strong>Post-Training Quantization (PTQ)</strong> applies quantization after training is complete. It’s fast to implement but can lose accuracy, especially at low bit-widths.</p>\n<p><strong>Quantization-Aware Training (QAT)</strong> integrates quantization into the training process. The model learns to handle quantization errors during training. PyTorch demonstrated that QAT can recover up to 96% of the accuracy degradation on HellaSwag and 68% of the perplexity degradation on WikiText for Llama3 compared to PTQ [12].</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Aspect</th><th>PTQ</th><th>QAT</th></tr></thead><tbody><tr><td>Training required</td><td>No</td><td>Yes</td></tr><tr><td>Compute cost</td><td>Low</td><td>High</td></tr><tr><td>Accuracy at INT8</td><td>Good</td><td>Excellent</td></tr><tr><td>Accuracy at INT4</td><td>Variable</td><td>Good</td></tr><tr><td>Time to deploy</td><td>Minutes</td><td>Hours/Days</td></tr></tbody></table></div>\n<h3 id=\"gguf-quantization-levels\">GGUF quantization levels</h3>\n<p>GGUF (Georgi Gerganov Universal Format) is designed for efficient CPU-based inference within the llama.cpp ecosystem. The format supports multiple quantization levels [13]:</p>\n<figure><figcaption><strong>GGUF Quantization Levels</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/local-ai-inferencing-techniques-0.webp\" width=\"512\" height=\"960\" alt=\"GGUF Quantization Levels\" loading=\"lazy\" decoding=\"async\" /></figure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Quantization</th><th>Size (7B model)</th><th>Perplexity increase</th><th>Notes</th></tr></thead><tbody><tr><td>Q8_0</td><td>~7.0 GB</td><td>+0.0004</td><td>Highest quality</td></tr><tr><td>Q5_K_M</td><td>~5.33 GB</td><td>+0.0142</td><td>Higher quality, slightly larger</td></tr><tr><td>Q4_K_M</td><td>~4.58 GB</td><td>+0.0535</td><td>Recommended default, best balance</td></tr><tr><td>Q4_K_S</td><td>~4.37 GB</td><td>+0.0992</td><td>Smaller, lower quality</td></tr><tr><td>Q3_K_M</td><td>~3.52 GB</td><td>+0.2437</td><td>Aggressive compression</td></tr></tbody></table></div>\n<p>Q4_K_M is the “safe default” for phones and lighter Macs. Q5_K_M improves detail and reasoning stability. Use an importance matrix (<code>--imatrix</code>) for optimal results [14].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Quantize with llama.cpp</span></span>\n<span class=\"line\"><span>./llama-quantize</span><span> model.gguf</span><span> model-Q4_K_M.gguf</span><span> Q4_K_M</span></span></code></pre>\n<h3 id=\"awq-gptq-and-bitsandbytes\">AWQ, GPTQ, and bitsandbytes</h3>\n<p><strong>AWQ (Activation-aware Weight Quantization)</strong>: Developed by MIT-HAN lab, AWQ protects salient weights by observing activations rather than weights themselves. It assumes not all weights are equally important and skips quantizing the most critical ones. AWQ achieves 95% quality retention, outperforming GPTQ (90%) in many benchmarks [15].</p>\n<p><strong>GPTQ (Generative Pre-trained Transformer Quantization)</strong>: A layer-wise post-training method that minimizes output error via Hessian-based optimization. Supports 8, 4, 3, or 2-bit quantization. GPTQ tends to overfit on its calibration data, which can hurt out-of-distribution performance [16].</p>\n<p><strong>bitsandbytes</strong>: Provides 8-bit (LLM.int8()) and 4-bit (NF4) quantization for PyTorch. The 4-bit NF4 format is specifically designed for neural network weights, assuming weights follow a normal distribution and placing quantization levels where most weights are concentrated (near zero) [17].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> BitsAndBytesConfig</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>quantization_config </span><span>=</span><span> BitsAndBytesConfig</span><span>(</span></span>\n<span class=\"line\"><span>    load_in_4bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_quant_type</span><span>=</span><span>\"nf4\"</span><span>,</span></span>\n<span class=\"line\"><span>    bnb_4bit_compute_dtype</span><span>=</span><span>torch.bfloat16,</span></span>\n<span class=\"line\"><span>    bnb_4bit_use_double_quant</span><span>=</span><span>True</span><span>,  </span><span># Quantize the quantization constants</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    quantization_config</span><span>=</span><span>quantization_config,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"exllamav2-quantization\">ExLlamaV2 quantization</h3>\n<p>ExLlamaV2 is an inference library optimized for consumer GPUs. Its EXL2 format supports 2, 3, 4, 5, 6, and 8-bit quantization and can mix different precisions within a model to preserve the most important weights [18].</p>\n<p>ExLlamaV2 achieves 56.44 tokens/second on a T4 GPU, providing the highest tokens-per-second compared to GPTQ or llama.cpp. It now supports paged attention via Flash Attention 2.5.7+ and includes dynamic batching with smart prompt caching [18].</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Use Case</th><th>Recommended Method</th></tr></thead><tbody><tr><td>GPU inference, high throughput</td><td>GPTQ</td></tr><tr><td>Quality-critical applications</td><td>AWQ</td></tr><tr><td>CPU/Apple devices</td><td>GGUF</td></tr><tr><td>Fine-tuned size control</td><td>EXL2</td></tr></tbody></table></div>\n<hr />\n<h2 id=\"kv-cache-optimization\">KV-cache optimization</h2>\n<p>The key-value cache stores intermediate attention states during autoregressive generation. Without optimization, KV cache wastes 60-80% of allocated memory through fragmentation and over-allocation [19].</p>\n<figure><figcaption><strong>Context length becomes a KV-cache memory bill</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/local-ai-inferencing-techniques-1.webp\" width=\"512\" height=\"562\" alt=\"Context length becomes a KV-cache memory bill\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"pagedattention\">PagedAttention</h3>\n<p>PagedAttention, introduced by vLLM, applies virtual memory concepts to KV cache management. Instead of storing each conversation’s KV cache as one contiguous block, PagedAttention breaks it into small, fixed-size “blocks” that can be stored anywhere in memory.</p>\n<p>The results: PagedAttention slashed KV cache waste to under 4%, enabling 2-4x throughput improvements [19].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># vLLM automatically uses PagedAttention</span></span>\n<span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    gpu_memory_utilization</span><span>=</span><span>0.90</span><span>,</span></span>\n<span class=\"line\"><span>    max_model_len</span><span>=</span><span>32768</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"continuous-batching\">Continuous batching</h3>\n<p>Continuous batching operates at the token level. Traditional batching processes requests in fixed groups, wasting compute when some sequences finish before others. Continuous batching walks model layers across a rotating set of in-flight sequences, reusing weight loads across heterogeneous sequence lengths [20].</p>\n<p>vLLM achieves up to 24x higher throughput than Hugging Face TGI under high-concurrency workloads through continuous batching combined with PagedAttention [21].</p>\n<h3 id=\"memory-management-strategies\">Memory management strategies</h3>\n<p><strong>Hierarchical caching</strong>: vLLM takes a hierarchical approach to KV caching. It first checks GPU memory, then CPU memory on cache miss, then retrieves from configured KV connectors (like LMCache for distributed caching) [22].</p>\n<p><strong>FP8 KV cache</strong>: Hopper and Blackwell GPUs support native FP8 KV cache, which halves KV cache memory requirements with minimal accuracy impact [19].</p>\n<p><strong>Prefix caching</strong>: Automatic Prefix Caching (APC) detects common prefixes across requests and shares KV cache blocks automatically. Shared prefixes during beam search reduce KV memory usage by up to 55% [23].</p>\n<hr />\n<h2 id=\"flashattention\">FlashAttention</h2>\n<p>FlashAttention is an IO-aware attention algorithm that reduces memory bandwidth bottlenecks by computing attention in tiles without materializing the full attention matrix.</p>\n<h3 id=\"flashattention-2\">FlashAttention-2</h3>\n<p>FlashAttention-2 improved on the original by better parallelizing across thread blocks and reducing synchronization. It achieves 35% utilization on H100 GPUs and speeds up training by 3-5x compared to baseline implementations from Hugging Face, reaching up to 225 TFLOPs/sec per A100 [24].</p>\n<h3 id=\"flashattention-3\">FlashAttention-3</h3>\n<p>FlashAttention-3 targets Hopper GPUs with three optimizations [25]:</p>\n<ol>\n<li><strong>Asynchronous execution</strong>: Exploits asynchrony between Tensor Cores and TMA to overlap computation and data movement via warp-specialization</li>\n<li><strong>Interleaved operations</strong>: Interleaves block-wise matmul and softmax operations</li>\n<li><strong>FP8 quantization</strong>: Uses block quantization with incoherent processing for FP8 low-precision</li>\n</ol>\n<p>Performance on H100:</p>\n<ul>\n<li>BF16: Up to 840 TFLOPs/s (85% utilization)</li>\n<li>FP8: Up to 1.3 PFLOPs/s</li>\n</ul>\n<p>FlashAttention-3 is up to 2.0x faster than FlashAttention-2. FP8 FlashAttention-3 achieves 2.6x lower numerical error than baseline FP8 attention because intermediate softmax results are kept in FP32 [25].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># FlashAttention is integrated into transformers</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    attn_implementation</span><span>=</span><span>\"flash_attention_2\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>torch.bfloat16,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<hr />\n<h2 id=\"tensor-parallelism-and-model-sharding\">Tensor parallelism and model sharding</h2>\n<p>When models exceed single-GPU memory, sharding distributes parameters across multiple devices.</p>\n<h3 id=\"tensor-parallelism\">Tensor parallelism</h3>\n<p>Tensor parallelism shards individual layers horizontally across GPUs. The two primary techniques are [26]:</p>\n<ul>\n<li><strong>Column parallelism</strong>: Splits weight matrices along columns, concatenates results after computation</li>\n<li><strong>Row parallelism</strong>: Splits matrices along rows, sums partial results post-computation</li>\n</ul>\n<p>With tensor parallelism across 4 GPUs, a 512 MB weight matrix becomes 128 MB per GPU, a 4x memory reduction for that layer.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Tensor parallelism across 4 GPUs on one node</span></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-70B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    tensor_parallel_size</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"communication-overhead\">Communication overhead</h3>\n<p>Tensor parallelism has high communication volume and presents a synchronization point in forward pass. Data exchange occurs over high-speed interconnects like NVLink (600 GB/s) or PCIe (32 GB/s). Interconnect speed critically affects overall performance, making tensor parallelism costly to scale beyond 1 node [27].</p>\n<h3 id=\"pipeline-parallelism\">Pipeline parallelism</h3>\n<p>For multi-node inference, combine tensor parallelism with pipeline parallelism:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Multi-node: 4 GPUs per node, 2 nodes</span></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-70B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    tensor_parallel_size</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>    pipeline_parallel_size</span><span>=</span><span>2</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"helix-parallelism\">Helix Parallelism</h3>\n<p>Helix Parallelism, introduced in 2025, applies a hybrid execution strategy: KV parallelism during attention to shard KV caches across GPUs, then reuses the same GPUs for tensor parallelism in dense LLMs or TP x Expert Parallel (EP) in MoEs during FFN computation. Compared to conventional approaches, Helix reduces time-to-first-token by up to 1.5x and supports up to 32x larger batches under the same latency budget for DeepSeek-R1 [28].</p>\n<hr />\n<h2 id=\"cpu-offloading-strategies\">CPU offloading strategies</h2>\n<p>When GPU memory is insufficient, offloading parts of the model or KV cache to CPU memory enables inference of larger models.</p>\n<figure><figcaption><strong>Larger models fit by moving less frequently used data outward</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/local-ai-inferencing-techniques-2.webp\" width=\"512\" height=\"663\" alt=\"Larger models fit by moving less frequently used data outward\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h3 id=\"neo-online-llm-inference-with-cpu-offloading\">NEO: Online LLM inference with CPU offloading</h3>\n<p>NEO offloads part of attention compute and KV cache states from GPU to local host CPU, increasing effective GPU batch size and inference throughput. It achieves up to 14%-6.6x throughput gains using two techniques [29]:</p>\n<ol>\n<li><strong>Asymmetric pipelining</strong>: Fully leverages compute resources of both GPU and CPU without overloading them</li>\n<li><strong>Load-aware scheduling</strong>: Maintains separate prefilling waitqueue, GPU decoding runqueue, and CPU decoding runqueue, making iteration-level adaptive scheduling decisions</li>\n</ol>\n<h3 id=\"nvidia-grace-hopper-unified-memory\">NVIDIA Grace Hopper unified memory</h3>\n<p>The NVLink-C2C connection and unified memory architecture on Grace Hopper enables models to use CPU memory if GPU memory is insufficient without explicit data transfer. When a model loads onto GH200, it uses 96 GB of HBM and accesses 480 GB of LPDDR memory connected to the CPU [30].</p>\n<h3 id=\"trade-offs\">Trade-offs</h3>\n<p>CPU offloading helps alleviate GPU memory limitations but shifts workload to CPU and increases data movement between CPU and GPU. For CPU-based attention operations, memory bandwidth, not compute power, determines performance. Overall training/inference speed may decrease due to synchronization overhead [30].</p>\n<hr />\n<h2 id=\"prompt-caching\">Prompt caching</h2>\n<p>Prompt caching reuses the model’s intermediate state (KV tensors in attention layers) for prefix tokens rather than recomputing them for each request.</p>\n<h3 id=\"how-it-works\">How it works</h3>\n<p>During prefill, the model builds a KV cache for the entire input. Prompt caching stores this KV cache keyed by the prompt tokens. For subsequent requests sharing the same prefix, the cached KV state is reused, skipping recomputation.</p>\n<p>Anthropic Claude Sonnet offers prompt caching with up to 90% cost savings and 85% latency reduction for long prompts. OpenAI provides automatic caching with 50% cost savings [31].</p>\n<h3 id=\"best-practices\">Best practices</h3>\n<ul>\n<li><strong>Front-load static content</strong>: Place constant information (system messages, context, instructions) at the beginning of prompts</li>\n<li><strong>Avoid dynamic elements in prefix</strong>: Don’t insert timestamps, request IDs, or per-request variables early in the prompt</li>\n<li><strong>Exact prefix matching</strong>: Even tiny differences (whitespace, JSON key order) break cache hits [23]</li>\n</ul>\n<h3 id=\"vllm-automatic-prefix-caching\">vLLM Automatic Prefix Caching</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    enable_prefix_caching</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>vLLM hashes prompt tokens and looks up cache blocks. SGLang uses Radix attention with a radix-tree for prefix caching [31].</p>\n<hr />\n<h2 id=\"inference-engines-deep-dive\">Inference Engines Deep Dive</h2>\n<p>Inference engines handle the actual execution of LLM forward passes. Each engine makes different trade-offs between ease of use, performance, and hardware support. This section covers four engines in detail: TensorRT-LLM (NVIDIA’s production stack), vLLM (the open-source throughput leader), SGLang (structured generation specialist), and llama.cpp (the portable C++ implementation).</p>\n<hr />\n<h2 id=\"tensorrt-llm\">TensorRT-LLM</h2>\n<p>TensorRT-LLM is NVIDIA’s open-source library for optimizing LLM inference on NVIDIA GPUs. It builds on TensorRT’s graph compilation and kernel optimization capabilities, adding LLM-specific features like in-flight batching, paged KV caching, and multi-GPU parallelism [32].</p>\n<h3 id=\"how-tensorrt-compiles-graphs\">How TensorRT compiles graphs</h3>\n<p>TensorRT-LLM converts model definitions into optimized TensorRT engines through a multi-stage compilation process:</p>\n<p><strong>1. Graph construction</strong>: The Model Definition API assembles a network graph. Each layer maps to TensorRT operations that can later be traversed or transformed [33].</p>\n<p><strong>2. Pattern matching and fusion</strong>: During compilation, TensorRT identifies operation sequences that can be fused into single GPU kernels. For example, a matmul followed by ReLU becomes one kernel without intermediate memory writes. The pattern-matching algorithm identifies fusions automatically, and an advanced kernel compiler converts them to efficient code [34].</p>\n<p><strong>3. Kernel autotuning</strong>: TensorRT automatically selects optimal GPU kernels based on GPU architecture, model configuration, precision, and batch size. The AutoTuner framework (added in 2025) provides custom-op-compatible tuning for operations like fused MoE and NVFP4 linear layers [35].</p>\n<p><strong>4. Plugin support</strong>: Some fusions (like FlashAttention) cannot be discovered automatically because they interleave operations in complex ways. Engineers can explicitly replace graph sections with plugins at compile time [34].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Building a TensorRT-LLM engine</span></span>\n<span class=\"line\"><span>from</span><span> tensorrt_llm </span><span>import</span><span> build</span></span>\n<span class=\"line\"><span>from</span><span> tensorrt_llm</span><span>.</span><span>models </span><span>import</span><span> LLaMAForCausalLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model configuration</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> LLaMAForCausalLM</span><span>.</span><span>from_hugging_face</span><span>(</span></span>\n<span class=\"line\"><span>    \"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    dtype</span><span>=</span><span>\"float16\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Build engine with optimizations</span></span>\n<span class=\"line\"><span>engine </span><span>=</span><span> build</span><span>(</span></span>\n<span class=\"line\"><span>    model,</span></span>\n<span class=\"line\"><span>    max_batch_size</span><span>=</span><span>64</span><span>,</span></span>\n<span class=\"line\"><span>    max_input_len</span><span>=</span><span>2048</span><span>,</span></span>\n<span class=\"line\"><span>    max_seq_len</span><span>=</span><span>4096</span><span>,</span></span>\n<span class=\"line\"><span>    use_paged_context_fmha</span><span>=</span><span>True</span><span>,  </span><span># Enable paged KV cache</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>engine</span><span>.</span><span>save</span><span>(</span><span>\"llama-8b-engine\"</span><span>)</span></span></code></pre>\n<h3 id=\"int8fp8-quantization-calibration\">INT8/FP8 quantization calibration</h3>\n<p>TensorRT-LLM supports multiple quantization methods, with FP8 recommended as the default for Hopper and Blackwell GPUs [36].</p>\n<p><strong>FP8 quantization (fp8_pc_pt)</strong>: Weights are quantized to FP8 per-channel. Activation ranges are calibrated and quantized per-token. FP8 preserves accuracy better than INT8 in most cases [37].</p>\n<p><strong>INT8 SmoothQuant (int8_sq)</strong>: Applies smoothing to weights before INT8 channel-wise quantization. Activation ranges are calibrated tensor-wise. Used as a fallback on Ada GPUs [37].</p>\n<p>The calibration process:</p>\n<ol>\n<li>Load a model checkpoint using the appropriate parallelism strategy</li>\n<li>Run calibration data through the model to determine quantization scales</li>\n<li>Output a quantized checkpoint with model config (JSON), quantized weights (safetensors), and tokenizer config (YAML) [38]</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> tensorrt_llm</span><span>.</span><span>quantization </span><span>import</span><span> quantize</span></span>\n<span class=\"line\"><span>from</span><span> tensorrt_llm</span><span>.</span><span>quantization </span><span>import</span><span> CalibConfig</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Configure calibration</span></span>\n<span class=\"line\"><span>calib_config </span><span>=</span><span> CalibConfig</span><span>(</span></span>\n<span class=\"line\"><span>    calib_dataset</span><span>=</span><span>\"cnn_dailymail\"</span><span>,</span></span>\n<span class=\"line\"><span>    calib_batch_size</span><span>=</span><span>8</span><span>,</span></span>\n<span class=\"line\"><span>    calib_max_seq_length</span><span>=</span><span>512</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Quantize to FP8</span></span>\n<span class=\"line\"><span>quantize</span><span>(</span></span>\n<span class=\"line\"><span>    model_dir</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    output_dir</span><span>=</span><span>\"llama-8b-fp8\"</span><span>,</span></span>\n<span class=\"line\"><span>    qformat</span><span>=</span><span>\"fp8\"</span><span>,</span></span>\n<span class=\"line\"><span>    calib_config</span><span>=</span><span>calib_config,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"in-flight-batching-implementation\">In-flight batching implementation</h3>\n<p>In-flight batching (also called continuous batching) eliminates wait times by dynamically managing request execution. Instead of processing fixed batches, TensorRT-LLM merges incoming requests into existing GPU batches, processing context and generation phases together [39].</p>\n<p>The scheduler maintains:</p>\n<ul>\n<li><strong>Prefill queue</strong>: New requests waiting for initial context processing</li>\n<li><strong>Generation queue</strong>: Active sequences generating tokens</li>\n<li><strong>Memory manager</strong>: Tracks KV cache block allocation</li>\n</ul>\n<p>When a sequence completes, its slot immediately becomes available for new requests without waiting for other sequences in the batch.</p>\n<h3 id=\"paged-kv-cache-in-tensorrt\">Paged KV-cache in TensorRT</h3>\n<p>TensorRT-LLM implements paged KV cache with configurable block sizes [40]:</p>\n<ul>\n<li>KV cache state is stored in blocks, each holding multiple tokens (default: 128 tokens)</li>\n<li>Only full blocks can be shared by multiple requests</li>\n<li>Block size is set at engine build time via <code>trtllm-build</code> with <code>--tokens_per_block</code> (must be power of 2)</li>\n<li>Larger blocks improve kernel efficiency but reduce cache reuse likelihood</li>\n</ul>\n<p><strong>KV cache reuse</strong>: Pages can be shared by requests with matching prefixes. This reduces time-to-first-token for multi-turn conversations and system prompts. Priority-based eviction (added in 2025) allows specifying priority and duration for token ranges, improving cache hit rates by approximately 20% [41].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Build engine with custom KV cache block size</span></span>\n<span class=\"line\"><span>trtllm-build</span><span> \\</span></span>\n<span class=\"line\"><span>    --checkpoint_dir</span><span> ./llama-checkpoint</span><span> \\</span></span>\n<span class=\"line\"><span>    --output_dir</span><span> ./llama-engine</span><span> \\</span></span>\n<span class=\"line\"><span>    --tokens_per_block</span><span> 64</span><span> \\</span></span>\n<span class=\"line\"><span>    --use_paged_context_fmha</span><span> enable</span><span> \\</span></span>\n<span class=\"line\"><span>    --kv_cache_type</span><span> paged</span></span></code></pre>\n<h3 id=\"building-engines-for-different-gpu-architectures\">Building engines for different GPU architectures</h3>\n<p>TensorRT engines are architecture-specific. An engine built for H100 will not run on A100 [42].</p>\n<p><strong>Specifying architectures at build time</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Build for Ada and Hopper only</span></span>\n<span class=\"line\"><span>make</span><span> -C</span><span> docker</span><span> release_build</span><span> CUDA_ARCHS=</span><span>\"89-real;90-real\"</span></span></code></pre>\n<p><strong>Using TensorRT Cloud for different GPUs</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Build for A100</span></span>\n<span class=\"line\"><span>trt-cloud</span><span> build</span><span> llm</span><span> \\</span></span>\n<span class=\"line\"><span>    --src-hf-repo=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span> \\</span></span>\n<span class=\"line\"><span>    --dtype=</span><span>\"float16\"</span><span> \\</span></span>\n<span class=\"line\"><span>    --gpu=</span><span>\"A100\"</span><span> \\</span></span>\n<span class=\"line\"><span>    --os=linux</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Build for RTX 3070 (consumer GPU)</span></span>\n<span class=\"line\"><span>trt-cloud</span><span> build</span><span> llm</span><span> \\</span></span>\n<span class=\"line\"><span>    --trtllm-checkpoint</span><span> checkpoint.zip</span><span> \\</span></span>\n<span class=\"line\"><span>    --gpu</span><span> RTX3070</span><span> \\</span></span>\n<span class=\"line\"><span>    --os</span><span> windows</span><span> \\</span></span>\n<span class=\"line\"><span>    --dtype</span><span> bfloat16</span></span></code></pre>\n<p><strong>GPU architecture codes</strong>:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>GPU Family</th><th>Compute Capability</th><th>CMake Code</th></tr></thead><tbody><tr><td>Ampere (A100)</td><td>8.0</td><td>80-real</td></tr><tr><td>Ada (RTX 4090, L40)</td><td>8.9</td><td>89-real</td></tr><tr><td>Hopper (H100, H200)</td><td>9.0</td><td>90-real</td></tr><tr><td>Blackwell (B200, GB200)</td><td>10.0</td><td>100-real</td></tr></tbody></table></div>\n<h3 id=\"practical-code-for-converting-and-serving-models\">Practical code for converting and serving models</h3>\n<p><strong>Full conversion and serving workflow</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Step 1: Convert HuggingFace model to TensorRT-LLM checkpoint</span></span>\n<span class=\"line\"><span>from</span><span> tensorrt_llm</span><span>.</span><span>models </span><span>import</span><span> LLaMAForCausalLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> LLaMAForCausalLM</span><span>.</span><span>from_hugging_face</span><span>(</span></span>\n<span class=\"line\"><span>    \"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    mapping</span><span>=</span><span>tensorrt_llm.</span><span>Mapping</span><span>(world_size</span><span>=</span><span>1</span><span>, tp_size</span><span>=</span><span>1</span><span>),</span></span>\n<span class=\"line\"><span>    dtype</span><span>=</span><span>\"bfloat16\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>save_checkpoint</span><span>(</span><span>\"./llama-checkpoint\"</span><span>)</span></span></code></pre>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Step 2: Build the engine</span></span>\n<span class=\"line\"><span>trtllm-build</span><span> \\</span></span>\n<span class=\"line\"><span>    --checkpoint_dir</span><span> ./llama-checkpoint</span><span> \\</span></span>\n<span class=\"line\"><span>    --output_dir</span><span> ./llama-engine</span><span> \\</span></span>\n<span class=\"line\"><span>    --gemm_plugin</span><span> bfloat16</span><span> \\</span></span>\n<span class=\"line\"><span>    --max_batch_size</span><span> 64</span><span> \\</span></span>\n<span class=\"line\"><span>    --max_input_len</span><span> 2048</span><span> \\</span></span>\n<span class=\"line\"><span>    --max_seq_len</span><span> 4096</span><span> \\</span></span>\n<span class=\"line\"><span>    --use_paged_context_fmha</span><span> enable</span></span></code></pre>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Step 3: Run inference</span></span>\n<span class=\"line\"><span>import</span><span> tensorrt_llm</span></span>\n<span class=\"line\"><span>from</span><span> tensorrt_llm</span><span>.</span><span>runtime </span><span>import</span><span> ModelRunner</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>runner </span><span>=</span><span> ModelRunner</span><span>.</span><span>from_dir</span><span>(</span><span>\"./llama-engine\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> runner</span><span>.</span><span>generate</span><span>(</span></span>\n<span class=\"line\"><span>    batch_input_ids</span><span>=</span><span>[[</span><span>1</span><span>, </span><span>2</span><span>, </span><span>3</span><span>, </span><span>4</span><span>]],  </span><span># Tokenized input</span></span>\n<span class=\"line\"><span>    max_new_tokens</span><span>=</span><span>256</span><span>,</span></span>\n<span class=\"line\"><span>    temperature</span><span>=</span><span>0.7</span><span>,</span></span>\n<span class=\"line\"><span>    top_p</span><span>=</span><span>0.9</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p><strong>Performance on H100 with FP8</strong>: Over 10,000 output tokens/s at peak throughput for 64 concurrent requests, with approximately 100ms time-to-first-token. H100 FP8 achieves up to 4.6x higher max throughput and 4.4x faster TTFT than A100 [43].</p>\n<hr />\n<h2 id=\"vllm\">vLLM</h2>\n<p>vLLM is an open-source inference engine that maximizes throughput through PagedAttention and continuous batching. As of June 2025, it has 49,200 GitHub stars and has become the de facto standard for high-throughput LLM serving [44].</p>\n<h3 id=\"pagedattention-algorithm-in-detail\">PagedAttention algorithm in detail</h3>\n<p>PagedAttention applies virtual memory concepts to KV cache management. Instead of allocating contiguous memory per request, vLLM divides GPU memory into fixed-size pages (blocks) that can be stored anywhere [45].</p>\n<p><strong>Block allocation process</strong>:</p>\n<ol>\n<li>The KV block manager divides GPU memory (and optionally CPU RAM) into physical KV blocks</li>\n<li>Each request maintains a block table mapping logical blocks to physical blocks</li>\n<li>Block table entries record physical block addresses and fill counts</li>\n<li>Blocks are allocated on demand as tokens are generated</li>\n</ol>\n<p><strong>Step-by-step allocation example</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>Initial state (prompt processed):</span></span>\n<span class=\"line\"><span>- Logical blocks: [Block 0: 14/16 tokens]</span></span>\n<span class=\"line\"><span>- Physical mapping: Logical 0 -&gt; Physical 5</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>After generating 2 tokens:</span></span>\n<span class=\"line\"><span>- Logical blocks: [Block 0: 16/16 tokens (full)]</span></span>\n<span class=\"line\"><span>- New allocation needed</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>After generating 3rd token:</span></span>\n<span class=\"line\"><span>- Logical blocks: [Block 0: 16/16], [Block 1: 1/16]</span></span>\n<span class=\"line\"><span>- Physical mapping: Logical 0 -&gt; Physical 5, Logical 1 -&gt; Physical 12</span></span></code></pre>\n<p><strong>Block size trade-offs</strong>: The default is 16 tokens per block. Larger blocks increase kernel parallelism but also increase memory fragmentation. The vLLM authors tested many configurations and found 16 tokens to be a good balance [46].</p>\n<p><strong>Memory sharing</strong>: Sequences with shared prefixes point their block tables to the same physical blocks. This enables prefix caching and reduces memory usage by up to 55% during beam search [47].</p>\n<p><strong>Performance impact</strong>: Traditional systems waste 60-80% of KV cache memory. PagedAttention reduces waste to under 4%, enabling 2-4x throughput improvements [45].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># vLLM automatically uses PagedAttention</span></span>\n<span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span><span>,</span><span> SamplingParams</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    gpu_memory_utilization</span><span>=</span><span>0.90</span><span>,  </span><span># Use 90% of GPU memory for KV cache</span></span>\n<span class=\"line\"><span>    block_size</span><span>=</span><span>16</span><span>,  </span><span># Tokens per KV block (default)</span></span>\n<span class=\"line\"><span>    swap_space</span><span>=</span><span>4</span><span>,  </span><span># GB of CPU memory for swapping</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>sampling_params </span><span>=</span><span> SamplingParams</span><span>(temperature</span><span>=</span><span>0.7</span><span>, max_tokens</span><span>=</span><span>256</span><span>)</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> llm</span><span>.</span><span>generate</span><span>([</span><span>\"Write a poem about caching:\"</span><span>], sampling_params)</span></span></code></pre>\n<h3 id=\"continuous-batching-scheduler\">Continuous batching scheduler</h3>\n<p>vLLM schedules work at the iteration level, not the request level. The scheduler assembles batches dynamically at each decoding step [48].</p>\n<p><strong>Request lifecycle</strong>:</p>\n<ol>\n<li>Request enters the engine and is wrapped in a Request object with status WAITING</li>\n<li>Request is added to the scheduler’s waiting queue (FCFS or priority-based)</li>\n<li>At each step, the scheduler selects requests for the current batch</li>\n<li>When a sequence emits EOS, its slot is immediately freed for new requests</li>\n</ol>\n<p><strong>Scheduler architecture</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>                    ┌─────────────────┐</span></span>\n<span class=\"line\"><span>                    │   API Server    │</span></span>\n<span class=\"line\"><span>                    └────────┬────────┘</span></span>\n<span class=\"line\"><span>                             │</span></span>\n<span class=\"line\"><span>                    ┌────────▼────────┐</span></span>\n<span class=\"line\"><span>                    │   AsyncLLM      │</span></span>\n<span class=\"line\"><span>                    └────────┬────────┘</span></span>\n<span class=\"line\"><span>                             │</span></span>\n<span class=\"line\"><span>                    ┌────────▼────────┐</span></span>\n<span class=\"line\"><span>                    │   EngineCore    │</span></span>\n<span class=\"line\"><span>                    │                 │</span></span>\n<span class=\"line\"><span>                    │  ┌───────────┐  │</span></span>\n<span class=\"line\"><span>                    │  │ Scheduler │  │</span></span>\n<span class=\"line\"><span>                    │  └─────┬─────┘  │</span></span>\n<span class=\"line\"><span>                    │        │        │</span></span>\n<span class=\"line\"><span>                    │  ┌─────▼─────┐  │</span></span>\n<span class=\"line\"><span>                    │  │  Executor │  │</span></span>\n<span class=\"line\"><span>                    │  └───────────┘  │</span></span>\n<span class=\"line\"><span>                    └─────────────────┘</span></span></code></pre>\n<p><strong>Step-level scheduling</strong>:</p>\n<ol>\n<li><strong>Schedule phase</strong>: Select which requests to run (decode and/or chunked prefill)</li>\n<li><strong>Execute phase</strong>: Run one forward pass for all active sequences</li>\n<li><strong>Output phase</strong>: Push results to output queue, free completed sequences</li>\n</ol>\n<p><strong>Chunked prefill</strong>: Long prompts are split into smaller chunks to prevent single requests from monopolizing GPU time. With chunked prefill enabled, decode requests are prioritized over prefill to minimize latency for in-progress generations [49].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    max_num_seqs</span><span>=</span><span>256</span><span>,  </span><span># Max concurrent sequences</span></span>\n<span class=\"line\"><span>    max_num_batched_tokens</span><span>=</span><span>8192</span><span>,  </span><span># Token budget per iteration</span></span>\n<span class=\"line\"><span>    enable_chunked_prefill</span><span>=</span><span>True</span><span>,  </span><span># Enable chunked prefill</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"prefix-caching-mechanism\">Prefix caching mechanism</h3>\n<p>vLLM’s Automatic Prefix Caching (APC) detects common prefixes and shares KV cache blocks automatically [50].</p>\n<p><strong>How it works</strong>:</p>\n<ol>\n<li>Prompt tokens are hashed to create a cache key</li>\n<li>On cache hit, existing KV blocks are reused</li>\n<li>On cache miss, new blocks are computed and cached</li>\n<li>LRU eviction removes least-recently-used blocks when memory is full</li>\n</ol>\n<p><strong>Cache hierarchy</strong>: vLLM checks GPU memory first, then CPU memory, then external KV connectors like LMCache for distributed caching [51].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    enable_prefix_caching</span><span>=</span><span>True</span><span>,  </span><span># Enable APC</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># First request computes and caches the system prompt</span></span>\n<span class=\"line\"><span>response1 </span><span>=</span><span> llm</span><span>.</span><span>generate</span><span>([</span></span>\n<span class=\"line\"><span>    \"System: You are a helpful assistant.\\n\\nUser: What is 2+2?\"</span></span>\n<span class=\"line\"><span>])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Second request reuses cached system prompt KV</span></span>\n<span class=\"line\"><span>response2 </span><span>=</span><span> llm</span><span>.</span><span>generate</span><span>([</span></span>\n<span class=\"line\"><span>    \"System: You are a helpful assistant.\\n\\nUser: What is 3+3?\"</span></span>\n<span class=\"line\"><span>])</span></span></code></pre>\n<h3 id=\"speculative-decoding-integration\">Speculative decoding integration</h3>\n<p>vLLM supports multiple speculative decoding methods, with Eagle 3 as the current state-of-the-art [52].</p>\n<p><strong>Supported methods</strong>:</p>\n<ul>\n<li><strong>Draft model</strong>: Separate smaller model proposes tokens</li>\n<li><strong>Eagle 1/3</strong>: Lightweight prediction head attached to target model</li>\n<li><strong>Suffix decoding</strong>: Pattern-matching against previous generations (roadmap item)</li>\n</ul>\n<p><strong>Eagle 3 performance</strong>: Up to 2.5x speedup across diverse scenarios. The Speculators library (v0.3.0, December 2025) provides end-to-end training support for Eagle3 draft models [53].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span><span>,</span><span> SamplingParams</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Using a separate draft model</span></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-70B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    speculative_model</span><span>=</span><span>\"meta-llama/Llama-3.2-1B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    num_speculative_tokens</span><span>=</span><span>5</span><span>,</span></span>\n<span class=\"line\"><span>    speculative_draft_tensor_parallel_size</span><span>=</span><span>1</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Using Eagle 3</span></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    speculative_config</span><span>=</span><span>{</span></span>\n<span class=\"line\"><span>        \"method\"</span><span>: </span><span>\"eagle\"</span><span>,</span></span>\n<span class=\"line\"><span>        \"model\"</span><span>: </span><span>\"yuhuili/EAGLE3-LLaMA3.1-Instruct-8B\"</span><span>,</span></span>\n<span class=\"line\"><span>        \"num_speculative_tokens\"</span><span>: </span><span>5</span><span>,</span></span>\n<span class=\"line\"><span>    },</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p><strong>Arctic Inference</strong>: Snowflake’s Arctic Inference integration achieves 4x faster inference for LLM agents on SWE-Bench tasks and up to 2.8x faster decoding for interactive workloads. It reduces MLP-based proposer latency from 1.47ms/token to 0.47ms/token [54].</p>\n<h3 id=\"tensor-parallelism-for-multi-gpu\">Tensor parallelism for multi-GPU</h3>\n<p>vLLM supports tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and expert parallelism (EP) [55].</p>\n<p><strong>Tensor parallelism</strong>: Shards each layer horizontally across GPUs using column parallelism (split along columns, concatenate results) and row parallelism (split along rows, sum results).</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Single-node 4-GPU tensor parallelism</span></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-70B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    tensor_parallel_size</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Multi-node: 4 GPUs per node, 2 nodes</span></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-70B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    tensor_parallel_size</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>    pipeline_parallel_size</span><span>=</span><span>2</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Combined TP + DP (8 GPUs total)</span></span>\n<span class=\"line\"><span># Requires data_parallel_size=4, tensor_parallel_size=2</span></span>\n<span class=\"line\"><span># via CLI: --data-parallel-size=4 --tensor-parallel-size=2</span></span></code></pre>\n<p><strong>Why TP within nodes</strong>: Tensor parallelism has high communication volume but benefits from fast interconnects. Use TP within nodes (NVLink) and PP across nodes (InfiniBand) when interconnects are slow. With NVLink, TP can extend across nodes [56].</p>\n<p><strong>Checking interconnect speed</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>nvidia-smi</span><span> topo</span><span> -m</span></span>\n<span class=\"line\"><span># Look for NVLink vs PIX (PCIe) connections</span></span></code></pre>\n<h3 id=\"openai-compatible-api-server-internals\">OpenAI-compatible API server internals</h3>\n<p>vLLM’s server implements OpenAI’s API specification using FastAPI [57].</p>\n<p><strong>Architecture</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>┌─────────────────────────────────────────────────┐</span></span>\n<span class=\"line\"><span>│              FastAPI Application                │</span></span>\n<span class=\"line\"><span>│                                                 │</span></span>\n<span class=\"line\"><span>│  ┌──────────────┐  ┌──────────────────────────┐│</span></span>\n<span class=\"line\"><span>│  │   /v1/chat   │  │  /v1/completions         ││</span></span>\n<span class=\"line\"><span>│  │  completions │  │                          ││</span></span>\n<span class=\"line\"><span>│  └──────┬───────┘  └───────────┬──────────────┘│</span></span>\n<span class=\"line\"><span>│         │                      │               │</span></span>\n<span class=\"line\"><span>│  ┌──────▼──────────────────────▼──────┐        │</span></span>\n<span class=\"line\"><span>│  │        OpenAIServingChat           │        │</span></span>\n<span class=\"line\"><span>│  │        OpenAIServingCompletion     │        │</span></span>\n<span class=\"line\"><span>│  │        OpenAIServingTokenization   │        │</span></span>\n<span class=\"line\"><span>│  └──────────────────┬─────────────────┘        │</span></span>\n<span class=\"line\"><span>│                     │                          │</span></span>\n<span class=\"line\"><span>│  ┌──────────────────▼─────────────────┐        │</span></span>\n<span class=\"line\"><span>│  │            AsyncLLM                │        │</span></span>\n<span class=\"line\"><span>│  └────────────────────────────────────┘        │</span></span>\n<span class=\"line\"><span>└─────────────────────────────────────────────────┘</span></span></code></pre>\n<p><strong>Request handling</strong>:</p>\n<ol>\n<li>FastAPI receives HTTP request with Pydantic validation</li>\n<li>Request is converted to internal format and queued</li>\n<li>AsyncLLM handles scheduling via EngineCore</li>\n<li>Streaming responses use Server-Sent Events (SSE)</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Start the OpenAI-compatible server</span></span>\n<span class=\"line\"><span>vllm</span><span> serve</span><span> meta-llama/Llama-3.1-8B-Instruct</span><span> \\</span></span>\n<span class=\"line\"><span>    --host</span><span> 0.0.0.0</span><span> \\</span></span>\n<span class=\"line\"><span>    --port</span><span> 8000</span><span> \\</span></span>\n<span class=\"line\"><span>    --tensor-parallel-size</span><span> 2</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Use with OpenAI SDK</span></span>\n<span class=\"line\"><span>curl</span><span> http://localhost:8000/v1/chat/completions</span><span> \\</span></span>\n<span class=\"line\"><span>    -H</span><span> \"Content-Type: application/json\"</span><span> \\</span></span>\n<span class=\"line\"><span>    -d</span><span> '{</span></span>\n<span class=\"line\"><span>        \"model\": \"meta-llama/Llama-3.1-8B-Instruct\",</span></span>\n<span class=\"line\"><span>        \"messages\": [{\"role\": \"user\", \"content\": \"Hello!\"}]</span></span>\n<span class=\"line\"><span>    }'</span></span></code></pre>\n<h3 id=\"performance-tuning-parameters\">Performance tuning parameters</h3>\n<p>Key parameters for optimizing vLLM performance [58]:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Parameter</th><th>Default</th><th>Description</th></tr></thead><tbody><tr><td><code>max_num_seqs</code></td><td>256</td><td>Maximum concurrent sequences</td></tr><tr><td><code>max_num_batched_tokens</code></td><td>8192</td><td>Token budget per iteration</td></tr><tr><td><code>gpu_memory_utilization</code></td><td>0.9</td><td>Fraction of GPU memory for KV cache</td></tr><tr><td><code>block_size</code></td><td>16</td><td>Tokens per KV block</td></tr><tr><td><code>swap_space</code></td><td>4</td><td>GB of CPU memory for swapping</td></tr><tr><td><code>enable_chunked_prefill</code></td><td>True</td><td>Split long prompts into chunks</td></tr><tr><td><code>max_model_len</code></td><td>model default</td><td>Maximum sequence length</td></tr></tbody></table></div>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    # Memory settings</span></span>\n<span class=\"line\"><span>    gpu_memory_utilization</span><span>=</span><span>0.95</span><span>,</span></span>\n<span class=\"line\"><span>    swap_space</span><span>=</span><span>8</span><span>,</span></span>\n<span class=\"line\"><span>    # Batching settings</span></span>\n<span class=\"line\"><span>    max_num_seqs</span><span>=</span><span>512</span><span>,</span></span>\n<span class=\"line\"><span>    max_num_batched_tokens</span><span>=</span><span>16384</span><span>,</span></span>\n<span class=\"line\"><span>    enable_chunked_prefill</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    # Performance settings</span></span>\n<span class=\"line\"><span>    enforce_eager</span><span>=</span><span>False</span><span>,  </span><span># Use CUDA graphs</span></span>\n<span class=\"line\"><span>    enable_prefix_caching</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<hr />\n<h2 id=\"sglang\">SGLang</h2>\n<p>SGLang is a high-performance serving framework that combines RadixAttention for KV cache reuse with a domain-specific language for structured generation. It achieves up to 6.4x higher throughput than baseline systems on structured workloads [59].</p>\n<h3 id=\"radixattention-and-how-it-differs-from-pagedattention\">RadixAttention and how it differs from PagedAttention</h3>\n<p>While vLLM’s PagedAttention optimizes memory within individual requests, RadixAttention focuses on maximizing KV cache reuse across multiple requests using a radix tree (compressed prefix tree) [60].</p>\n<p><strong>PagedAttention</strong>: Divides KV cache into fixed-size blocks. Blocks are allocated on demand and freed when sequences complete. Prefix sharing requires explicit cache lookup.</p>\n<p><strong>RadixAttention</strong>: Stores KV cache in a radix tree structure. The tree enables automatic prefix matching, insertion, and LRU eviction. KV cache persists after request completion for potential reuse.</p>\n<p><strong>Key differences</strong>:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Aspect</th><th>PagedAttention</th><th>RadixAttention</th></tr></thead><tbody><tr><td>Data structure</td><td>Block tables</td><td>Radix tree</td></tr><tr><td>Prefix matching</td><td>Hash lookup</td><td>Tree traversal</td></tr><tr><td>Cache persistence</td><td>Cleared after request</td><td>Retained in LRU cache</td></tr><tr><td>Primary optimization</td><td>Memory efficiency</td><td>KV reuse across calls</td></tr></tbody></table></div>\n<p><strong>How RadixAttention works</strong>:</p>\n<ol>\n<li>After each generation, KV cache is inserted into the radix tree keyed by token sequence</li>\n<li>New requests traverse the tree to find the longest matching prefix</li>\n<li>Matching prefix KV cache is reused; only new tokens require computation</li>\n<li>LRU eviction removes least-recently-used branches when memory is full</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> sglang </span><span>as</span><span> sgl</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># RadixAttention automatically reuses KV cache</span></span>\n<span class=\"line\"><span>@sgl</span><span>.</span><span>function</span></span>\n<span class=\"line\"><span>def</span><span> multi_turn_chat</span><span>(</span><span>s</span><span>,</span><span> turns</span><span>):</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>system</span><span>(</span><span>\"You are a helpful assistant.\"</span><span>)</span></span>\n<span class=\"line\"><span>    for</span><span> user_msg </span><span>in</span><span> turns</span><span>:</span></span>\n<span class=\"line\"><span>        s </span><span>+=</span><span> sgl</span><span>.</span><span>user</span><span>(user_msg)</span></span>\n<span class=\"line\"><span>        s </span><span>+=</span><span> sgl</span><span>.</span><span>assistant</span><span>(sgl.</span><span>gen</span><span>(</span><span>\"response\"</span><span>, max_tokens</span><span>=</span><span>256</span><span>))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># KV cache for system prompt is computed once and reused</span></span>\n<span class=\"line\"><span>runtime </span><span>=</span><span> sgl</span><span>.</span><span>Runtime</span><span>(model_path</span><span>=</span><span>\"meta-llama/Llama-3.1-8B-Instruct\"</span><span>)</span></span>\n<span class=\"line\"><span>sgl</span><span>.</span><span>set_default_backend</span><span>(runtime)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># These calls reuse cached system prompt</span></span>\n<span class=\"line\"><span>result1 </span><span>=</span><span> multi_turn_chat</span><span>.</span><span>run</span><span>(turns</span><span>=</span><span>[</span><span>\"What is Python?\"</span><span>])</span></span>\n<span class=\"line\"><span>result2 </span><span>=</span><span> multi_turn_chat</span><span>.</span><span>run</span><span>(turns</span><span>=</span><span>[</span><span>\"What is Rust?\"</span><span>])</span></span></code></pre>\n<h3 id=\"the-sglang-dsl-and-structured-generation\">The SGLang DSL and structured generation</h3>\n<p>SGLang provides a domain-specific language for defining complex LLM programs with branching, parallelism, and constraints [61].</p>\n<p><strong>Core primitives</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> sglang </span><span>as</span><span> sgl</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@sgl</span><span>.</span><span>function</span></span>\n<span class=\"line\"><span>def</span><span> analyze_code</span><span>(</span><span>s</span><span>,</span><span> code</span><span>):</span></span>\n<span class=\"line\"><span>    # Sequential generation</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>user</span><span>(</span><span>f</span><span>\"Analyze this code:</span><span>\\n</span><span>{</span><span>code</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>assistant</span><span>(sgl.</span><span>gen</span><span>(</span><span>\"analysis\"</span><span>, max_tokens</span><span>=</span><span>500</span><span>))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Forking for parallel evaluation</span></span>\n<span class=\"line\"><span>    with</span><span> s</span><span>.</span><span>fork</span><span>(</span><span>2</span><span>)</span><span> as</span><span> forks</span><span>:</span></span>\n<span class=\"line\"><span>        forks</span><span>[</span><span>0</span><span>]</span><span> +=</span><span> sgl</span><span>.</span><span>user</span><span>(</span><span>\"Rate the code quality (1-10):\"</span><span>)</span></span>\n<span class=\"line\"><span>        forks</span><span>[</span><span>0</span><span>]</span><span> +=</span><span> sgl</span><span>.</span><span>assistant</span><span>(sgl.</span><span>gen</span><span>(</span><span>\"quality\"</span><span>, max_tokens</span><span>=</span><span>10</span><span>))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        forks</span><span>[</span><span>1</span><span>]</span><span> +=</span><span> sgl</span><span>.</span><span>user</span><span>(</span><span>\"Suggest improvements:\"</span><span>)</span></span>\n<span class=\"line\"><span>        forks</span><span>[</span><span>1</span><span>]</span><span> +=</span><span> sgl</span><span>.</span><span>assistant</span><span>(sgl.</span><span>gen</span><span>(</span><span>\"improvements\"</span><span>, max_tokens</span><span>=</span><span>200</span><span>))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Select best based on criteria</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>select</span><span>(</span><span>\"best_response\"</span><span>, forks, criteria</span><span>=</span><span>\"quality\"</span><span>)</span></span></code></pre>\n<p><strong>Structured output with JSON schema</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> pydantic </span><span>import</span><span> BaseModel</span></span>\n<span class=\"line\"><span>import</span><span> sglang </span><span>as</span><span> sgl</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> CodeReview</span><span>(</span><span>BaseModel</span><span>):</span></span>\n<span class=\"line\"><span>    quality_score</span><span>:</span><span> int</span></span>\n<span class=\"line\"><span>    issues</span><span>:</span><span> list</span><span>[</span><span>str</span><span>]</span></span>\n<span class=\"line\"><span>    suggestions</span><span>:</span><span> list</span><span>[</span><span>str</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@sgl</span><span>.</span><span>function</span></span>\n<span class=\"line\"><span>def</span><span> structured_review</span><span>(</span><span>s</span><span>,</span><span> code</span><span>):</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>user</span><span>(</span><span>f</span><span>\"Review this code and provide structured feedback:</span><span>\\n</span><span>{</span><span>code</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>assistant</span><span>(</span></span>\n<span class=\"line\"><span>        sgl.</span><span>gen</span><span>(</span><span>\"review\"</span><span>, max_tokens</span><span>=</span><span>500</span><span>, json_schema</span><span>=</span><span>CodeReview.</span><span>model_json_schema</span><span>())</span></span>\n<span class=\"line\"><span>    )</span></span></code></pre>\n<h3 id=\"constrained-decoding-with-fsm\">Constrained decoding with FSM</h3>\n<p>SGLang implements constrained decoding using compressed finite-state machines (FSM), enabling generation that conforms to regular expressions or JSON schemas [62].</p>\n<p><strong>How FSM decoding works</strong>:</p>\n<ol>\n<li>The constraint (regex, JSON schema) is converted to a finite-state machine</li>\n<li>At each decoding step, invalid tokens are masked based on the current FSM state</li>\n<li>Only tokens leading to valid FSM transitions can be sampled</li>\n</ol>\n<p><strong>Compressed FSM optimization</strong>: Standard FSM decoding processes one token at a time even when transitions are deterministic. SGLang’s compressed FSM collapses chains of singular transitions into single edges [62].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> sglang </span><span>as</span><span> sgl</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@sgl</span><span>.</span><span>function</span></span>\n<span class=\"line\"><span>def</span><span> extract_email</span><span>(</span><span>s</span><span>,</span><span> text</span><span>):</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>user</span><span>(</span><span>f</span><span>\"Extract the email address from: </span><span>{</span><span>text</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    # Constrain output to valid email format</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>assistant</span><span>(</span></span>\n<span class=\"line\"><span>        sgl.</span><span>gen</span><span>(</span><span>\"email\"</span><span>, regex</span><span>=</span><span>r</span><span>\"[a-zA-Z0-9._%+-]</span><span>+</span><span>@[a-zA-Z0-9.-]</span><span>+</span><span>\\.[a-zA-Z]</span><span>{2,}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    )</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@sgl</span><span>.</span><span>function</span></span>\n<span class=\"line\"><span>def</span><span> generate_json</span><span>(</span><span>s</span><span>,</span><span> prompt</span><span>):</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>user</span><span>(prompt)</span></span>\n<span class=\"line\"><span>    # Constrain to valid JSON object</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>assistant</span><span>(</span></span>\n<span class=\"line\"><span>        sgl.</span><span>gen</span><span>(</span><span>\"data\"</span><span>, regex</span><span>=</span><span>r</span><span>'\\{[</span><span>^</span><span>{}]</span><span>*</span><span>\\}'</span><span>)</span></span>\n<span class=\"line\"><span>    )</span></span></code></pre>\n<h3 id=\"jump-forward-decoding\">Jump-forward decoding</h3>\n<p>Jump-forward decoding exploits deterministic sections of constrained output to skip unnecessary LLM forward passes [62].</p>\n<p><strong>Example</strong>: When generating JSON, after a key like <code>\"name\"</code>, the next token must be <code>:</code>. Instead of running the LLM to generate <code>:</code>, SGLang inserts it directly.</p>\n<p><strong>How it works</strong>:</p>\n<ol>\n<li>The compressed FSM identifies deterministic token sequences</li>\n<li>When the FSM has only one valid path, those tokens are inserted without LLM computation</li>\n<li>RadixAttention automatically handles KV cache for the inserted tokens</li>\n</ol>\n<p><strong>Performance impact</strong>: For structured formats with fixed elements, jump-forward decoding can skip 30-50% of generation steps. On JSON decoding benchmarks, this optimization alone delivers up to 1.6x throughput improvement [62].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>Standard decoding:</span></span>\n<span class=\"line\"><span>{\"name\" -&gt; LLM -&gt; \":\" -&gt; LLM -&gt; \" \" -&gt; LLM -&gt; \"\\\"\" -&gt; LLM -&gt; \"John\" ...</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>Jump-forward decoding:</span></span>\n<span class=\"line\"><span>{\"name\" -&gt; INSERT \": \\\"\" -&gt; LLM -&gt; \"John\" ...</span></span>\n<span class=\"line\"><span>(Skipped 3 LLM calls)</span></span></code></pre>\n<h3 id=\"how-sglang-handles-complex-prompts\">How SGLang handles complex prompts</h3>\n<p>SGLang’s runtime optimizes complex multi-call LLM programs through several mechanisms [63]:</p>\n<p><strong>1. Automatic batching</strong>: Multiple <code>sgl.gen()</code> calls across forked branches are batched into single LLM forward passes.</p>\n<p><strong>2. KV cache sharing via RadixAttention</strong>: Forks share the KV cache of their common prefix. New branches only compute KV for diverging tokens.</p>\n<p><strong>3. Scheduling optimization</strong>: The runtime schedules calls to maximize GPU utilization, interleaving prefill and decode across requests.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> sglang </span><span>as</span><span> sgl</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@sgl</span><span>.</span><span>function</span></span>\n<span class=\"line\"><span>def</span><span> tree_of_thought</span><span>(</span><span>s</span><span>,</span><span> problem</span><span>):</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>user</span><span>(</span><span>f</span><span>\"Solve: </span><span>{</span><span>problem</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Generate 3 candidate solutions in parallel</span></span>\n<span class=\"line\"><span>    with</span><span> s</span><span>.</span><span>fork</span><span>(</span><span>3</span><span>)</span><span> as</span><span> candidates</span><span>:</span></span>\n<span class=\"line\"><span>        for</span><span> i</span><span>,</span><span> c </span><span>in</span><span> enumerate</span><span>(candidates):</span></span>\n<span class=\"line\"><span>            c </span><span>+=</span><span> sgl</span><span>.</span><span>assistant</span><span>(sgl.</span><span>gen</span><span>(</span><span>f</span><span>\"solution_</span><span>{</span><span>i</span><span>}</span><span>\"</span><span>, max_tokens</span><span>=</span><span>200</span><span>))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Evaluate each solution</span></span>\n<span class=\"line\"><span>    evaluations </span><span>=</span><span> []</span></span>\n<span class=\"line\"><span>    for</span><span> i</span><span>,</span><span> c </span><span>in</span><span> enumerate</span><span>(candidates):</span></span>\n<span class=\"line\"><span>        c </span><span>+=</span><span> sgl</span><span>.</span><span>user</span><span>(</span><span>\"Rate this solution 1-10:\"</span><span>)</span></span>\n<span class=\"line\"><span>        c </span><span>+=</span><span> sgl</span><span>.</span><span>assistant</span><span>(sgl.</span><span>gen</span><span>(</span><span>f</span><span>\"rating_</span><span>{</span><span>i</span><span>}</span><span>\"</span><span>, max_tokens</span><span>=</span><span>5</span><span>))</span></span>\n<span class=\"line\"><span>        evaluations</span><span>.</span><span>append</span><span>(c)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Select best</span></span>\n<span class=\"line\"><span>    s </span><span>+=</span><span> sgl</span><span>.</span><span>select</span><span>(</span><span>\"best\"</span><span>, evaluations, key</span><span>=lambda</span><span> x</span><span>: </span><span>int</span><span>(x[</span><span>\"rating\"</span><span>]))</span></span></code></pre>\n<p><strong>Server deployment</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Start SGLang server</span></span>\n<span class=\"line\"><span>python</span><span> -m</span><span> sglang.launch_server</span><span> \\</span></span>\n<span class=\"line\"><span>    --model-path</span><span> meta-llama/Llama-3.1-8B-Instruct</span><span> \\</span></span>\n<span class=\"line\"><span>    --port</span><span> 30000</span><span> \\</span></span>\n<span class=\"line\"><span>    --tp</span><span> 2</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># OpenAI-compatible endpoint</span></span>\n<span class=\"line\"><span>curl</span><span> http://localhost:30000/v1/chat/completions</span><span> \\</span></span>\n<span class=\"line\"><span>    -H</span><span> \"Content-Type: application/json\"</span><span> \\</span></span>\n<span class=\"line\"><span>    -d</span><span> '{\"model\": \"default\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello\"}]}'</span></span></code></pre>\n<hr />\n<h2 id=\"llamacpp-deep-dive\">llama.cpp Deep Dive</h2>\n<p>llama.cpp is a C/C++ inference engine that prioritizes portability and minimal dependencies. It runs on CPUs, Apple Silicon, NVIDIA GPUs, AMD GPUs, and various accelerators. As of August 2025, it has over 85,000 GitHub stars and 1,200 contributors [64].</p>\n<h3 id=\"ggml-tensor-library-internals\">GGML tensor library internals</h3>\n<p>GGML (Georgi Gerganov Machine Learning) is the tensor library underlying llama.cpp. It provides a portable foundation for neural network operations [65].</p>\n<p><strong>Core design principles</strong>:</p>\n<ol>\n<li><strong>No external dependencies</strong>: GGML is self-contained C code</li>\n<li><strong>Static memory allocation</strong>: Tensor sizes are known at graph construction time</li>\n<li><strong>Pluggable backends</strong>: CPU, Metal, CUDA, Vulkan, SYCL, etc.</li>\n</ol>\n<p><strong>Tensor representation</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>struct</span><span> ggml_tensor {</span></span>\n<span class=\"line\"><span>    enum</span><span> ggml_type type;</span><span>      // Data type (F32, F16, Q4_K, etc.)</span></span>\n<span class=\"line\"><span>    int</span><span> n_dims;</span><span>               // Number of dimensions</span></span>\n<span class=\"line\"><span>    int64_t</span><span> ne[</span><span>4</span><span>];</span><span>            // Number of elements per dimension</span></span>\n<span class=\"line\"><span>    size_t</span><span> nb[</span><span>4</span><span>];</span><span>             // Stride in bytes per dimension</span></span>\n<span class=\"line\"><span>    void</span><span> *</span><span> data;</span><span>              // Pointer to data</span></span>\n<span class=\"line\"><span>    struct</span><span> ggml_tensor </span><span>*</span><span> src[</span><span>2</span><span>];</span><span>  // Source tensors (for operations)</span></span>\n<span class=\"line\"><span>    // ...</span></span>\n<span class=\"line\"><span>};</span></span></code></pre>\n<p><strong>Computation graph</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// Create context</span></span>\n<span class=\"line\"><span>struct</span><span> ggml_context </span><span>*</span><span> ctx </span><span>=</span><span> ggml_init</span><span>(params);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Define tensors</span></span>\n<span class=\"line\"><span>struct</span><span> ggml_tensor </span><span>*</span><span> a </span><span>=</span><span> ggml_new_tensor_2d</span><span>(ctx</span><span>,</span><span> GGML_TYPE_F32</span><span>,</span><span> 768</span><span>,</span><span> 768</span><span>);</span></span>\n<span class=\"line\"><span>struct</span><span> ggml_tensor </span><span>*</span><span> b </span><span>=</span><span> ggml_new_tensor_2d</span><span>(ctx</span><span>,</span><span> GGML_TYPE_F32</span><span>,</span><span> 768</span><span>,</span><span> 768</span><span>);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Define operation (creates graph node)</span></span>\n<span class=\"line\"><span>struct</span><span> ggml_tensor </span><span>*</span><span> c </span><span>=</span><span> ggml_mul_mat</span><span>(ctx</span><span>,</span><span> a</span><span>,</span><span> b);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Build computation graph</span></span>\n<span class=\"line\"><span>struct</span><span> ggml_cgraph </span><span>*</span><span> graph </span><span>=</span><span> ggml_new_graph</span><span>(ctx);</span></span>\n<span class=\"line\"><span>ggml_build_forward_expand</span><span>(graph</span><span>,</span><span> c);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>// Execute on backend</span></span>\n<span class=\"line\"><span>ggml_backend_graph_compute</span><span>(backend</span><span>,</span><span> graph);</span></span></code></pre>\n<p><strong>Backend architecture</strong>: GGML’s pluggable backend system enables the same graph to execute on different hardware. Each backend implements a standard interface for memory allocation, data transfer, and kernel execution [66].</p>\n<h3 id=\"metal-backend-implementation\">Metal backend implementation</h3>\n<p>The Metal backend provides GPU acceleration on Apple Silicon through hand-optimized compute shaders [67].</p>\n<p><strong>Key features</strong>:</p>\n<ul>\n<li>Optimized matrix multiplication kernels for M1/M2/M3 chips</li>\n<li>Unified memory eliminates CPU-GPU data transfers</li>\n<li>Flash attention implementation for memory-efficient attention</li>\n<li>Support for all GGML quantization types</li>\n</ul>\n<p><strong>Memory management</strong>: Metal uses shared buffers that both CPU and GPU can access directly. This is a significant advantage over discrete GPUs that require explicit data transfers.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Build llama.cpp with Metal (default on macOS)</span></span>\n<span class=\"line\"><span>cmake</span><span> -B</span><span> build</span></span>\n<span class=\"line\"><span>cmake</span><span> --build</span><span> build</span><span> --config</span><span> Release</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run with Metal acceleration</span></span>\n<span class=\"line\"><span>./build/bin/llama-cli</span><span> \\</span></span>\n<span class=\"line\"><span>    -m</span><span> llama-3.1-8b-instruct-q4_k_m.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    -ngl</span><span> 99</span><span> \\  </span><span># Offload all layers to GPU</span></span>\n<span class=\"line\"><span>    -p</span><span> \"Hello, world!\"</span></span></code></pre>\n<p><strong>Metal performance tuning</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Use flash attention (reduces memory bandwidth)</span></span>\n<span class=\"line\"><span>./build/bin/llama-cli</span><span> -m</span><span> model.gguf</span><span> -fa</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Adjust batch size for throughput</span></span>\n<span class=\"line\"><span>./build/bin/llama-cli</span><span> -m</span><span> model.gguf</span><span> -b</span><span> 512</span></span></code></pre>\n<h3 id=\"cuda-backend-and-kernel-optimizations\">CUDA backend and kernel optimizations</h3>\n<p>The CUDA backend provides GPU acceleration on NVIDIA hardware with hand-tuned kernels [68].</p>\n<p><strong>Recent optimizations (2025)</strong>:</p>\n<ul>\n<li><strong>FlashAttention</strong>: Memory-efficient attention with tiling and kernel fusion</li>\n<li><strong>Split-K optimization</strong>: Divides KV sequence across workgroups for increased parallelism</li>\n<li><strong>Wavefront tuning</strong>: Optimizations for NVIDIA’s 32-thread warps</li>\n<li><strong>NVFP4 support</strong>: Native 4-bit floating point on Blackwell GPUs (25% faster prompt processing)</li>\n</ul>\n<p><strong>CUDA kernel micro-optimizations</strong>:</p>\n<ul>\n<li>Memory coalescing in <code>mul_mat_id</code> function</li>\n<li>fastdiv and fastmodulo optimizations for Ada and Blackwell</li>\n<li>PAD_REFLECT_1D kernel optimization (1-11% memory bandwidth improvement)</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Build with CUDA support</span></span>\n<span class=\"line\"><span>cmake</span><span> -B</span><span> build</span><span> -DGGML_CUDA=ON</span></span>\n<span class=\"line\"><span>cmake</span><span> --build</span><span> build</span><span> --config</span><span> Release</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Enable unified memory for larger models</span></span>\n<span class=\"line\"><span>export</span><span> GGML_CUDA_ENABLE_UNIFIED_MEMORY</span><span>=</span><span>1</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>./build/bin/llama-cli</span><span> \\</span></span>\n<span class=\"line\"><span>    -m</span><span> llama-3.1-70b-instruct-q4_k_m.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    -ngl</span><span> 99</span><span> \\</span></span>\n<span class=\"line\"><span>    -fa</span><span>  # Flash attention</span></span></code></pre>\n<p><strong>Multi-GPU with ik_llama.cpp</strong>: The ik_llama.cpp fork implements tensor parallelism at the GGML graph level (Split Mode Graph). This distributes compute graph nodes across GPUs rather than just assigning layers, achieving 3-4x performance gains over standard multi-GPU methods [69].</p>\n<h3 id=\"quantization-kernel-implementations\">Quantization kernel implementations</h3>\n<p>llama.cpp supports various quantization types with specialized dequantization kernels [70].</p>\n<p><strong>Quantization type families</strong>:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Family</th><th>Types</th><th>Description</th></tr></thead><tbody><tr><td>Type-0</td><td>Q4_0, Q5_0, Q8_0</td><td>Global scale: <code>w = d * q</code></td></tr><tr><td>Type-1</td><td>Q4_1, Q5_1</td><td>Scale + minimum: <code>w = d * q + m</code></td></tr><tr><td>K-quants</td><td>Q4_K, Q5_K, Q6_K</td><td>Block-wise with super-blocks of 256</td></tr><tr><td>I-quants</td><td>IQ4_XS, IQ3_XXS</td><td>Importance-based with lookup tables</td></tr></tbody></table></div>\n<p><strong>K-quant implementation details</strong>:</p>\n<p>K-quants use super-blocks of 256 quantized values, subdivided into blocks of 16 or 32. Each super-block stores:</p>\n<ul>\n<li>Scale factors for each sub-block</li>\n<li>Minimum values (for some types)</li>\n<li>Quantized weights</li>\n</ul>\n<p><strong>Q4_K_M and Q5_K_M</strong> implement mixed precision: most weights use the base precision, but half of <code>attention.wv</code> and <code>feed_forward.w2</code> tensors use higher precision for better accuracy [71].</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Quantize a model</span></span>\n<span class=\"line\"><span>./build/bin/llama-quantize</span><span> \\</span></span>\n<span class=\"line\"><span>    model-f16.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    model-q4_k_m.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    Q4_K_M</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Quantize with importance matrix for better quality</span></span>\n<span class=\"line\"><span>./build/bin/llama-quantize</span><span> \\</span></span>\n<span class=\"line\"><span>    model-f16.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    model-q4_k_m.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    Q4_K_M</span><span> \\</span></span>\n<span class=\"line\"><span>    --imatrix</span><span> imatrix.dat</span></span></code></pre>\n<p><strong>Advanced quantization options</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Mix quantization types for different tensors</span></span>\n<span class=\"line\"><span>./build/bin/llama-quantize</span><span> \\</span></span>\n<span class=\"line\"><span>    model-f16.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    model-mixed.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    Q4_K_M</span><span> \\</span></span>\n<span class=\"line\"><span>    --output-tensor-type</span><span> Q8_0</span><span> \\</span></span>\n<span class=\"line\"><span>    --token-embedding-type</span><span> Q8_0</span></span></code></pre>\n<h3 id=\"memory-layout-for-different-quant-types\">Memory layout for different quant types</h3>\n<p>Each quantization type has a specific memory layout optimized for efficient dequantization [72].</p>\n<p><strong>Q4_0 layout</strong> (block of 32 weights):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>[scale: fp16 (2 bytes)] [quants: 16 bytes (32 x 4-bit)]</span></span>\n<span class=\"line\"><span>Total: 18 bytes for 32 weights = 4.5 bits/weight</span></span></code></pre>\n<p><strong>Q4_K layout</strong> (super-block of 256 weights):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>[d: fp16] [dmin: fp16] [scales: 12 bytes] [quants: 128 bytes]</span></span>\n<span class=\"line\"><span>- 8 sub-blocks of 32 weights each</span></span>\n<span class=\"line\"><span>- Each sub-block has its own 6-bit scale</span></span>\n<span class=\"line\"><span>Total: 144 bytes for 256 weights = 4.5 bits/weight</span></span></code></pre>\n<p><strong>Q5_K layout</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>[d: fp16] [dmin: fp16] [scales: 12 bytes] [qh: 32 bytes] [ql: 128 bytes]</span></span>\n<span class=\"line\"><span>- qh contains the 5th bit for each weight</span></span>\n<span class=\"line\"><span>- ql contains lower 4 bits</span></span>\n<span class=\"line\"><span>Total: 176 bytes for 256 weights = 5.5 bits/weight</span></span></code></pre>\n<p><strong>Dequantization kernels</strong>: Each quantization type has optimized dequantization code for each backend. The CUDA kernels use warp-level primitives for efficient parallel dequantization:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// Simplified Q4_K dequantization (CUDA)</span></span>\n<span class=\"line\"><span>__global__ </span><span>void</span><span> dequantize_q4_k</span><span>(</span><span>const</span><span> void</span><span> *</span><span> src</span><span>,</span><span> float</span><span> *</span><span> dst) {</span></span>\n<span class=\"line\"><span>    const</span><span> block_q4_k </span><span>*</span><span> block </span><span>=</span><span> (</span><span>const</span><span> block_q4_k </span><span>*</span><span>)src </span><span>+</span><span> blockIdx</span><span>.</span><span>x;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    const</span><span> float</span><span> d </span><span>=</span><span> __half2float(</span><span>block</span><span>-&gt;</span><span>d)</span><span>;</span></span>\n<span class=\"line\"><span>    const</span><span> float</span><span> dmin </span><span>=</span><span> __half2float(</span><span>block</span><span>-&gt;</span><span>dmin)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Parallel dequantization across warp</span></span>\n<span class=\"line\"><span>    int</span><span> lane </span><span>=</span><span> threadIdx</span><span>.</span><span>x </span><span>%</span><span> 32</span><span>;</span></span>\n<span class=\"line\"><span>    uint8_t</span><span> q </span><span>=</span><span> block</span><span>-&gt;</span><span>qs[lane];</span></span>\n<span class=\"line\"><span>    float</span><span> scale </span><span>=</span><span> d </span><span>*</span><span> (</span><span>block</span><span>-&gt;</span><span>scales[lane </span><span>/</span><span> 32</span><span>] </span><span>&amp;</span><span> 0x</span><span>3F</span><span>);</span></span>\n<span class=\"line\"><span>    float</span><span> min </span><span>=</span><span> dmin </span><span>*</span><span> (</span><span>block</span><span>-&gt;</span><span>scales[lane </span><span>/</span><span> 32</span><span>] </span><span>&gt;&gt;</span><span> 4</span><span>);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    dst[lane </span><span>*</span><span> 2</span><span>] </span><span>=</span><span> scale </span><span>*</span><span> (q </span><span>&amp;</span><span> 0x</span><span>F</span><span>) </span><span>-</span><span> min;</span></span>\n<span class=\"line\"><span>    dst[lane </span><span>*</span><span> 2</span><span> +</span><span> 1</span><span>] </span><span>=</span><span> scale </span><span>*</span><span> (q </span><span>&gt;&gt;</span><span> 4</span><span>) </span><span>-</span><span> min;</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<hr />\n<h2 id=\"inference-engine-comparison\">Inference Engine Comparison</h2>\n<h3 id=\"latency-vs-throughput-trade-offs\">Latency vs throughput trade-offs</h3>\n<p>Each engine optimizes for different points on the latency-throughput curve [73].</p>\n<p><strong>Benchmark results</strong> (Llama 3.1 8B, H100, 1000 ShareGPT prompts):</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Engine</th><th>Throughput (tok/s)</th><th>TTFT (ms)</th><th>Per-token latency (ms)</th></tr></thead><tbody><tr><td>SGLang</td><td>16,215</td><td>156</td><td>4-21</td></tr><tr><td>LMDeploy</td><td>16,132</td><td>142</td><td>12-35</td></tr><tr><td>vLLM</td><td>12,553</td><td>123</td><td>8-42</td></tr><tr><td>TensorRT-LLM</td><td>10,847</td><td>187</td><td>15-28</td></tr></tbody></table></div>\n<p><strong>Key observations</strong>:</p>\n<ul>\n<li><strong>vLLM</strong>: Fastest time-to-first-token (TTFT), best scaling at high concurrency</li>\n<li><strong>SGLang</strong>: Highest raw throughput, most stable per-token latency</li>\n<li><strong>TensorRT-LLM</strong>: Lower throughput in benchmarks but excels on B200/Blackwell GPUs</li>\n<li><strong>llama.cpp</strong>: Not designed for high-throughput serving but offers best portability</li>\n</ul>\n<p><strong>At different concurrency levels</strong> (GPT-OSS-120B):</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Concurrency</th><th>vLLM (tok/s)</th><th>SGLang (tok/s)</th><th>TensorRT-LLM (tok/s)</th></tr></thead><tbody><tr><td>1</td><td>892</td><td>756</td><td>1,024</td></tr><tr><td>10</td><td>2,341</td><td>2,567</td><td>2,189</td></tr><tr><td>100</td><td>4,741</td><td>4,523</td><td>3,892</td></tr></tbody></table></div>\n<h3 id=\"memory-efficiency-comparison\">Memory efficiency comparison</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Engine</th><th>KV Cache Approach</th><th>Memory Waste</th><th>Max Context (8B model, 24GB)</th></tr></thead><tbody><tr><td>vLLM</td><td>PagedAttention</td><td>&lt;4%</td><td>128K tokens</td></tr><tr><td>SGLang</td><td>RadixAttention</td><td>&lt;4%</td><td>128K tokens</td></tr><tr><td>TensorRT-LLM</td><td>Paged + Priority eviction</td><td>&lt;5%</td><td>128K tokens</td></tr><tr><td>llama.cpp</td><td>Contiguous</td><td>10-20%</td><td>32K tokens</td></tr></tbody></table></div>\n<p><strong>KV cache quantization support</strong>:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Engine</th><th>FP8 KV</th><th>INT8 KV</th><th>NVFP4 KV</th></tr></thead><tbody><tr><td>TensorRT-LLM</td><td>Yes</td><td>Yes</td><td>Yes (Blackwell)</td></tr><tr><td>vLLM</td><td>Yes</td><td>Yes</td><td>No</td></tr><tr><td>SGLang</td><td>Yes</td><td>Yes</td><td>No</td></tr><tr><td>llama.cpp</td><td>Q8_0, Q4_0</td><td>Via type</td><td>No</td></tr></tbody></table></div>\n<h3 id=\"feature-matrix\">Feature matrix</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Feature</th><th>TensorRT-LLM</th><th>vLLM</th><th>SGLang</th><th>llama.cpp</th></tr></thead><tbody><tr><td><strong>Quantization</strong></td><td></td><td></td><td></td><td></td></tr><tr><td>FP8</td><td>Yes (Hopper+)</td><td>Yes</td><td>Yes</td><td>No</td></tr><tr><td>INT8</td><td>Yes</td><td>Yes</td><td>Yes</td><td>Yes (Q8_0)</td></tr><tr><td>INT4</td><td>Yes (AWQ)</td><td>Yes (AWQ/GPTQ)</td><td>Yes</td><td>Yes (Q4_K)</td></tr><tr><td>2-bit</td><td>No</td><td>No</td><td>No</td><td>Yes (Q2_K)</td></tr><tr><td><strong>Parallelism</strong></td><td></td><td></td><td></td><td></td></tr><tr><td>Tensor parallel</td><td>Yes</td><td>Yes</td><td>Yes</td><td>Limited</td></tr><tr><td>Pipeline parallel</td><td>Yes</td><td>Yes</td><td>Yes</td><td>No</td></tr><tr><td>Expert parallel</td><td>Yes</td><td>Yes</td><td>Yes</td><td>No</td></tr><tr><td><strong>Features</strong></td><td></td><td></td><td></td><td></td></tr><tr><td>Speculative decoding</td><td>Yes</td><td>Yes (Eagle3)</td><td>Yes</td><td>Yes</td></tr><tr><td>Prefix caching</td><td>Yes</td><td>Yes (APC)</td><td>Yes (Radix)</td><td>Yes</td></tr><tr><td>Constrained decoding</td><td>No</td><td>Yes</td><td>Yes (FSM)</td><td>Yes (grammar)</td></tr><tr><td>Multi-LoRA</td><td>Yes</td><td>Yes</td><td>Yes</td><td>Yes</td></tr><tr><td>Vision models</td><td>Yes</td><td>Yes</td><td>Yes</td><td>Yes (LLaVA)</td></tr><tr><td><strong>Deployment</strong></td><td></td><td></td><td></td><td></td></tr><tr><td>OpenAI API</td><td>Via Triton</td><td>Yes</td><td>Yes</td><td>Yes (server)</td></tr><tr><td>NVIDIA GPUs</td><td>Yes</td><td>Yes</td><td>Yes</td><td>Yes</td></tr><tr><td>AMD GPUs</td><td>No</td><td>Yes (ROCm)</td><td>Yes (ROCm)</td><td>Yes (HIP)</td></tr><tr><td>Apple Silicon</td><td>No</td><td>No</td><td>No</td><td>Yes (Metal)</td></tr><tr><td>CPU only</td><td>No</td><td>No</td><td>No</td><td>Yes</td></tr></tbody></table></div>\n<h3 id=\"when-to-use-each-engine\">When to use each engine</h3>\n<p><strong>TensorRT-LLM</strong>:</p>\n<ul>\n<li>Production deployments on NVIDIA GPUs</li>\n<li>Maximum throughput on H100/B200</li>\n<li>When using NVIDIA’s full stack (Triton, NIM)</li>\n<li>Willing to invest in engine building and tuning</li>\n</ul>\n<p><strong>vLLM</strong>:</p>\n<ul>\n<li>High-concurrency API serving</li>\n<li>Fast time-to-first-token requirements</li>\n<li>OpenAI API compatibility</li>\n<li>General-purpose GPU inference with good defaults</li>\n</ul>\n<p><strong>SGLang</strong>:</p>\n<ul>\n<li>Structured output generation (JSON, code)</li>\n<li>Complex multi-turn or branching LLM programs</li>\n<li>Maximum KV cache reuse across requests</li>\n<li>Constrained decoding with FSM</li>\n</ul>\n<p><strong>llama.cpp</strong>:</p>\n<ul>\n<li>Local deployment on consumer hardware</li>\n<li>Apple Silicon (M1/M2/M3/M4)</li>\n<li>CPU-only environments</li>\n<li>Edge devices and mobile</li>\n<li>Maximum quantization flexibility</li>\n</ul>\n<h3 id=\"practical-deployment-recommendations\">Practical deployment recommendations</h3>\n<p><strong>For startup API products</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>vLLM + OpenAI-compatible server</span></span>\n<span class=\"line\"><span>- Fast setup, good defaults</span></span>\n<span class=\"line\"><span>- Scale with tensor parallelism</span></span>\n<span class=\"line\"><span>- Add prefix caching for repeated prompts</span></span></code></pre>\n<p><strong>For enterprise on NVIDIA</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>TensorRT-LLM + Triton Inference Server</span></span>\n<span class=\"line\"><span>- Maximum throughput per GPU dollar</span></span>\n<span class=\"line\"><span>- Production-grade monitoring</span></span>\n<span class=\"line\"><span>- Integration with NVIDIA enterprise support</span></span></code></pre>\n<p><strong>For structured outputs</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>SGLang</span></span>\n<span class=\"line\"><span>- Native JSON schema support</span></span>\n<span class=\"line\"><span>- Jump-forward decoding for speed</span></span>\n<span class=\"line\"><span>- RadixAttention for multi-turn efficiency</span></span></code></pre>\n<p><strong>For local/edge deployment</strong>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>llama.cpp</span></span>\n<span class=\"line\"><span>- Q4_K_M for balance of quality and size</span></span>\n<span class=\"line\"><span>- Metal for Apple, CUDA for NVIDIA</span></span>\n<span class=\"line\"><span>- Works offline with no cloud dependency</span></span></code></pre>\n<hr />\n<h2 id=\"practical-considerations\">Practical considerations</h2>\n<h3 id=\"when-to-use-each-technique\">When to use each technique</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Technique</th><th>Best for</th><th>Hardware requirements</th></tr></thead><tbody><tr><td>REAP</td><td>MoE models (Mixtral, DeepSeek, Qwen)</td><td>Same as original model</td></tr><tr><td>Speculative decoding</td><td>High-latency applications, large target models</td><td>GPU memory for both draft and target</td></tr><tr><td>Q4_K_M quantization</td><td>CPU/Apple devices, limited RAM</td><td>4GB+ RAM for 7B models</td></tr><tr><td>AWQ/GPTQ</td><td>GPU inference, quality-sensitive tasks</td><td>CUDA GPU</td></tr><tr><td>PagedAttention</td><td>High-concurrency serving</td><td>vLLM-compatible GPU</td></tr><tr><td>FlashAttention-3</td><td>H100/H200 deployments</td><td>Hopper architecture</td></tr><tr><td>Tensor parallelism</td><td>Models exceeding single GPU memory</td><td>Multiple GPUs with NVLink</td></tr><tr><td>CPU offloading</td><td>Memory-constrained setups</td><td>High CPU memory bandwidth</td></tr><tr><td>Prompt caching</td><td>Repeated prompts, RAG, chat systems</td><td>Sufficient memory for cache</td></tr></tbody></table></div>\n<h3 id=\"combining-techniques\">Combining techniques</h3>\n<p>These techniques are not mutually exclusive. A production deployment might use:</p>\n<ol>\n<li><strong>REAP</strong> to prune an MoE model to 50% experts</li>\n<li><strong>AWQ quantization</strong> to reduce memory footprint further</li>\n<li><strong>FlashAttention-3</strong> for efficient attention computation</li>\n<li><strong>PagedAttention + continuous batching</strong> for high-throughput serving</li>\n<li><strong>Prefix caching</strong> to avoid redundant computation for common prefixes</li>\n<li><strong>Speculative decoding</strong> to reduce latency for interactive applications</li>\n</ol>\n<p>The optimal combination depends on your specific constraints: available hardware, latency requirements, throughput targets, and accuracy tolerances.</p>\n<hr />\n<h2 id=\"references\">References</h2>\n<p>[1] InfoQ - Cactus v1: Cross-Platform LLM Inference on Mobile with Zero Latency and Full Privacy. <a href=\"https://www.infoq.com/news/2025/12/cactus-on-device-inference/\">https://www.infoq.com/news/2025/12/cactus-on-device-inference/</a></p>\n<p>[2] InfoQ - Cactus v1: Cross-Platform LLM Inference on Mobile. <a href=\"https://www.infoq.com/news/2025/12/cactus-on-device-inference/\">https://www.infoq.com/news/2025/12/cactus-on-device-inference/</a></p>\n<p>[3] PMC - Tiny Machine Learning and On-Device Inference: A Survey. <a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC12115890/\">https://pmc.ncbi.nlm.nih.gov/articles/PMC12115890/</a></p>\n<p>[4] Novus - The Rise of Local AI Models: Going Small to Go Big. <a href=\"https://www.novusasi.com/blog/the-rise-of-local-ai-models-going-small-to-go-big\">https://www.novusasi.com/blog/the-rise-of-local-ai-models-going-small-to-go-big</a></p>\n<p>[5] Cerebras - REAP: One-Shot Pruning for Trillion-Parameter MoE Models. <a href=\"https://www.cerebras.ai/blog/reap\">https://www.cerebras.ai/blog/reap</a></p>\n<p>[6] arXiv - REAP the Experts: Why Pruning Prevails for One-Shot MoE compression. <a href=\"https://arxiv.org/abs/2510.13999\">https://arxiv.org/abs/2510.13999</a></p>\n<p>[7] HuggingFace - Cerebras REAP Collection. <a href=\"https://huggingface.co/collections/cerebras/cerebras-reap\">https://huggingface.co/collections/cerebras/cerebras-reap</a></p>\n<p>[8] NVIDIA Technical Blog - An Introduction to Speculative Decoding. <a href=\"https://developer.nvidia.com/blog/an-introduction-to-speculative-decoding-for-reducing-latency-in-ai-inference/\">https://developer.nvidia.com/blog/an-introduction-to-speculative-decoding-for-reducing-latency-in-ai-inference/</a></p>\n<p>[9] BentoML - Speculative Decoding. <a href=\"https://bentoml.com/llm/inference-optimization/speculative-decoding\">https://bentoml.com/llm/inference-optimization/speculative-decoding</a></p>\n<p>[10] vLLM Documentation - Speculative Decoding. <a href=\"https://docs.vllm.ai/en/latest/features/spec_decode/\">https://docs.vllm.ai/en/latest/features/spec_decode/</a></p>\n<p>[11] LMSYS - SpecForge: Accelerating Speculative Decoding Training for SGLang. <a href=\"https://lmsys.org/blog/2025-07-25-spec-forge/\">https://lmsys.org/blog/2025-07-25-spec-forge/</a></p>\n<p>[12] PyTorch Blog - Quantization-Aware Training for Large Language Models. <a href=\"https://pytorch.org/blog/quantization-aware-training/\">https://pytorch.org/blog/quantization-aware-training/</a></p>\n<p>[13] llama.cpp GitHub - Quantize README. <a href=\"https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md\">https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md</a></p>\n<p>[14] Local AI Zone - AI Model Quantization 2025 Guide. <a href=\"https://local-ai-zone.github.io/guides/what-is-ai-quantization-q4-k-m-q8-gguf-guide-2025.html\">https://local-ai-zone.github.io/guides/what-is-ai-quantization-q4-k-m-q8-gguf-guide-2025.html</a></p>\n<p>[15] Local AI Master - AWQ vs GPTQ vs GGUF Comparison. <a href=\"https://localaimaster.com/blog/quantization-explained\">https://localaimaster.com/blog/quantization-explained</a></p>\n<p>[16] Maarten Grootendorst - Which Quantization Method is Right for You? <a href=\"https://newsletter.maartengrootendorst.com/p/which-quantization-method-is-right\">https://newsletter.maartengrootendorst.com/p/which-quantization-method-is-right</a></p>\n<p>[17] HuggingFace - Making LLMs even more accessible with bitsandbytes. <a href=\"https://huggingface.co/blog/4bit-transformers-bitsandbytes\">https://huggingface.co/blog/4bit-transformers-bitsandbytes</a></p>\n<p>[18] GitHub - ExLlamaV2. <a href=\"https://github.com/turboderp-org/exllamav2\">https://github.com/turboderp-org/exllamav2</a></p>\n<p>[19] arXiv - Efficient Memory Management for LLM Serving with PagedAttention. <a href=\"https://arxiv.org/abs/2309.06180\">https://arxiv.org/abs/2309.06180</a></p>\n<p>[20] HuggingFace Blog - Continuous batching from first principles. <a href=\"https://huggingface.co/blog/continuous_batching\">https://huggingface.co/blog/continuous_batching</a></p>\n<p>[21] arXiv - Comparative Analysis of LLM Inference Serving Systems. <a href=\"https://arxiv.org/html/2511.17593v1\">https://arxiv.org/html/2511.17593v1</a></p>\n<p>[22] Ceph.io - KV Caching with vLLM, LMCache, and Ceph. <a href=\"https://ceph.io/en/news/blog/2025/vllm-kv-caching/\">https://ceph.io/en/news/blog/2025/vllm-kv-caching/</a></p>\n<p>[23] vLLM Documentation - Automatic Prefix Caching. <a href=\"https://docs.vllm.ai/en/stable/design/prefix_caching/\">https://docs.vllm.ai/en/stable/design/prefix_caching/</a></p>\n<p>[24] GitHub - Dao-AILab/flash-attention. <a href=\"https://github.com/Dao-AILab/flash-attention\">https://github.com/Dao-AILab/flash-attention</a></p>\n<p>[25] PyTorch Blog - FlashAttention-3: Fast and Accurate Attention. <a href=\"https://pytorch.org/blog/flashattention-3/\">https://pytorch.org/blog/flashattention-3/</a></p>\n<p>[26] Meta Engineering - Scaling LLM Inference: Innovations in Parallelism. <a href=\"https://engineering.fb.com/2025/10/17/ai-research/scaling-llm-inference-innovations-tensor-parallelism-context-parallelism-expert-parallelism/\">https://engineering.fb.com/2025/10/17/ai-research/scaling-llm-inference-innovations-tensor-parallelism-context-parallelism-expert-parallelism/</a></p>\n<p>[27] AWS Neuron Documentation - Parallelism Techniques for LLM Inference. <a href=\"https://awsdocs-neuron.readthedocs-hosted.com/en/latest/libraries/nxd-inference/app-notes/parallelism.html\">https://awsdocs-neuron.readthedocs-hosted.com/en/latest/libraries/nxd-inference/app-notes/parallelism.html</a></p>\n<p>[28] arXiv - Helix Parallelism: Rethinking Sharding Strategies. <a href=\"https://arxiv.org/html/2507.07120v1\">https://arxiv.org/html/2507.07120v1</a></p>\n<p>[29] arXiv - NEO: Saving GPU Memory Crisis with CPU Offloading. <a href=\"https://arxiv.org/abs/2411.01142\">https://arxiv.org/abs/2411.01142</a></p>\n<p>[30] NVIDIA Technical Blog - Accelerate Large-Scale LLM Inference with CPU-GPU Memory Sharing. <a href=\"https://developer.nvidia.com/blog/accelerate-large-scale-llm-inference-and-kv-cache-offload-with-cpu-gpu-memory-sharing/\">https://developer.nvidia.com/blog/accelerate-large-scale-llm-inference-and-kv-cache-offload-with-cpu-gpu-memory-sharing/</a></p>\n<p>[31] ngrok Blog - Prompt caching: 10x cheaper LLM tokens, but how? <a href=\"https://ngrok.com/blog/prompt-caching/\">https://ngrok.com/blog/prompt-caching/</a></p>\n<p>[32] NVIDIA TensorRT-LLM Documentation - Overview. <a href=\"https://nvidia.github.io/TensorRT-LLM/overview.html\">https://nvidia.github.io/TensorRT-LLM/overview.html</a></p>\n<p>[33] NVIDIA TensorRT-LLM - Model Definition. <a href=\"https://nvidia.github.io/TensorRT-LLM/architecture/core-concepts.html\">https://nvidia.github.io/TensorRT-LLM/architecture/core-concepts.html</a></p>\n<p>[34] NVIDIA Technical Blog - Optimizing Inference on LLMs with TensorRT-LLM. <a href=\"https://developer.nvidia.com/blog/optimizing-inference-on-llms-with-tensorrt-llm-now-publicly-available/\">https://developer.nvidia.com/blog/optimizing-inference-on-llms-with-tensorrt-llm-now-publicly-available/</a></p>\n<p>[35] NVIDIA TensorRT-LLM - Release Notes. <a href=\"https://nvidia.github.io/TensorRT-LLM/release-notes.html\">https://nvidia.github.io/TensorRT-LLM/release-notes.html</a></p>\n<p>[36] NVIDIA Technical Blog - Optimizing LLMs for Performance and Accuracy with Post-Training Quantization. <a href=\"https://developer.nvidia.com/blog/optimizing-llms-for-performance-and-accuracy-with-post-training-quantization/\">https://developer.nvidia.com/blog/optimizing-llms-for-performance-and-accuracy-with-post-training-quantization/</a></p>\n<p>[37] NVIDIA TensorRT-LLM - Numerical Precision. <a href=\"https://nvidia.github.io/TensorRT-LLM/reference/precision.html\">https://nvidia.github.io/TensorRT-LLM/reference/precision.html</a></p>\n<p>[38] NVIDIA NeMo Framework - Quantization. <a href=\"https://docs.nvidia.com/nemo-framework/user-guide/latest/nemotoolkit/nlp/quantization.html\">https://docs.nvidia.com/nemo-framework/user-guide/latest/nemotoolkit/nlp/quantization.html</a></p>\n<p>[39] MarkTechPost - vLLM vs TensorRT-LLM vs HF TGI vs LMDeploy Technical Comparison. <a href=\"https://www.marktechpost.com/2025/11/19/vllm-vs-tensorrt-llm-vs-hf-tgi-vs-lmdeploy-a-deep-technical-comparison-for-production-llm-inference/\">https://www.marktechpost.com/2025/11/19/vllm-vs-tensorrt-llm-vs-hf-tgi-vs-lmdeploy-a-deep-technical-comparison-for-production-llm-inference/</a></p>\n<p>[40] NVIDIA TensorRT-LLM - KV Cache Reuse. <a href=\"https://nvidia.github.io/TensorRT-LLM/advanced/kv-cache-reuse.html\">https://nvidia.github.io/TensorRT-LLM/advanced/kv-cache-reuse.html</a></p>\n<p>[41] NVIDIA Technical Blog - KV Cache Reuse Optimizations in TensorRT-LLM. <a href=\"https://developer.nvidia.com/blog/introducing-new-kv-cache-reuse-optimizations-in-nvidia-tensorrt-llm/\">https://developer.nvidia.com/blog/introducing-new-kv-cache-reuse-optimizations-in-nvidia-tensorrt-llm/</a></p>\n<p>[42] NVIDIA TensorRT-LLM - Building from Source. <a href=\"https://nvidia.github.io/TensorRT-LLM/installation/build-from-source-linux.html\">https://nvidia.github.io/TensorRT-LLM/installation/build-from-source-linux.html</a></p>\n<p>[43] BentoML - Best Practices for Tuning TensorRT-LLM. <a href=\"https://www.bentoml.com/blog/tuning-tensor-rt-llm-for-optimal-serving-with-bentoml\">https://www.bentoml.com/blog/tuning-tensor-rt-llm-for-optimal-serving-with-bentoml</a></p>\n<p>[44] vLLM Blog - Inside vLLM: Anatomy of a High-Throughput LLM Inference System. <a href=\"https://blog.vllm.ai/2025/09/05/anatomy-of-vllm.html\">https://blog.vllm.ai/2025/09/05/anatomy-of-vllm.html</a></p>\n<p>[45] vLLM Documentation - Paged Attention. <a href=\"https://docs.vllm.ai/en/stable/design/paged_attention/\">https://docs.vllm.ai/en/stable/design/paged_attention/</a></p>\n<p>[46] Red Hat Developer - How PagedAttention resolves memory waste of LLM systems. <a href=\"https://developers.redhat.com/articles/2025/07/24/how-pagedattention-resolves-memory-waste-llm-systems\">https://developers.redhat.com/articles/2025/07/24/how-pagedattention-resolves-memory-waste-llm-systems</a></p>\n<p>[47] arXiv - Efficient Memory Management for LLM Serving with PagedAttention. <a href=\"https://arxiv.org/abs/2309.06180\">https://arxiv.org/abs/2309.06180</a></p>\n<p>[48] Ubicloud Blog - Life of an inference request (vLLM V1). <a href=\"https://www.ubicloud.com/blog/life-of-an-inference-request-vllm-v1\">https://www.ubicloud.com/blog/life-of-an-inference-request-vllm-v1</a></p>\n<p>[49] vLLM Documentation - Scheduler. <a href=\"https://docs.vllm.ai/en/stable/api/vllm/v1/core/sched/scheduler/\">https://docs.vllm.ai/en/stable/api/vllm/v1/core/sched/scheduler/</a></p>\n<p>[50] vLLM Documentation - Automatic Prefix Caching. <a href=\"https://docs.vllm.ai/en/stable/design/prefix_caching/\">https://docs.vllm.ai/en/stable/design/prefix_caching/</a></p>\n<p>[51] Ceph.io - KV Caching with vLLM and LMCache. <a href=\"https://ceph.io/en/news/blog/2025/vllm-kv-caching/\">https://ceph.io/en/news/blog/2025/vllm-kv-caching/</a></p>\n<p>[52] vLLM Documentation - Speculative Decoding. <a href=\"https://docs.vllm.ai/en/latest/features/spec_decode/\">https://docs.vllm.ai/en/latest/features/spec_decode/</a></p>\n<p>[53] vLLM Blog - Speculators v0.3.0. <a href=\"https://blog.vllm.ai/2025/12/13/speculators-v030.html\">https://blog.vllm.ai/2025/12/13/speculators-v030.html</a></p>\n<p>[54] Snowflake Engineering Blog - Fastest Speculative Decoding with Arctic Inference. <a href=\"https://www.snowflake.com/en/engineering-blog/fast-speculative-decoding-vllm-arctic/\">https://www.snowflake.com/en/engineering-blog/fast-speculative-decoding-vllm-arctic/</a></p>\n<p>[55] vLLM Documentation - Parallelism and Scaling. <a href=\"https://docs.vllm.ai/en/stable/serving/parallelism_scaling/\">https://docs.vllm.ai/en/stable/serving/parallelism_scaling/</a></p>\n<p>[56] Red Hat Developer - Distributed Inference with vLLM. <a href=\"https://developers.redhat.com/articles/2025/02/06/distributed-inference-with-vllm\">https://developers.redhat.com/articles/2025/02/06/distributed-inference-with-vllm</a></p>\n<p>[57] vLLM Documentation - OpenAI-Compatible Server. <a href=\"https://docs.vllm.ai/en/stable/serving/openai_compatible_server/\">https://docs.vllm.ai/en/stable/serving/openai_compatible_server/</a></p>\n<p>[58] Java Code Geeks - Under the Hood of vLLM: Memory, Scheduling &amp; Batching Strategies. <a href=\"https://www.javacodegeeks.com/2025/10/under-the-hood-of-vllm-memory-scheduling-batching-strategies.html\">https://www.javacodegeeks.com/2025/10/under-the-hood-of-vllm-memory-scheduling-batching-strategies.html</a></p>\n<p>[59] LMSYS Blog - Fast and Expressive LLM Inference with RadixAttention and SGLang. <a href=\"https://lmsys.org/blog/2024-01-17-sglang/\">https://lmsys.org/blog/2024-01-17-sglang/</a></p>\n<p>[60] SugiV Blog - SGLang Deep Dive: Inside SGLang. <a href=\"https://blog.sugiv.fyi/sglang-deep-dive-inside-sglang\">https://blog.sugiv.fyi/sglang-deep-dive-inside-sglang</a></p>\n<p>[61] SGLang Paper - Efficient Execution of Structured Language Model Programs. <a href=\"https://proceedings.neurips.cc/paper_files/paper/2024/file/724be4472168f31ba1c9ac630f15dec8-Paper-Conference.pdf\">https://proceedings.neurips.cc/paper_files/paper/2024/file/724be4472168f31ba1c9ac630f15dec8-Paper-Conference.pdf</a></p>\n<p>[62] LMSYS Blog - Fast JSON Decoding with Compressed Finite State Machine. <a href=\"https://lmsys.org/blog/2024-02-05-compressed-fsm/\">https://lmsys.org/blog/2024-02-05-compressed-fsm/</a></p>\n<p>[63] ROCm Blogs - SGLang: Fast Serving Framework on AMD Instinct GPUs. <a href=\"https://rocm.blogs.amd.com/artificial-intelligence/sglang/README.html\">https://rocm.blogs.amd.com/artificial-intelligence/sglang/README.html</a></p>\n<p>[64] llama.cpp Wikipedia. <a href=\"https://en.wikipedia.org/wiki/Llama.cpp\">https://en.wikipedia.org/wiki/Llama.cpp</a></p>\n<p>[65] DeepWiki - ggml-org/llama.cpp. <a href=\"https://deepwiki.com/ggml-org/llama.cpp\">https://deepwiki.com/ggml-org/llama.cpp</a></p>\n<p>[66] DeepWiki - Backend Architecture and Registration. <a href=\"https://deepwiki.com/ggml-org/llama.cpp/4.1-cpu-backend\">https://deepwiki.com/ggml-org/llama.cpp/4.1-cpu-backend</a></p>\n<p>[67] DeepWiki - Metal Backend. <a href=\"https://deepwiki.com/ggml-org/llama.cpp/4.5-sycl-backend\">https://deepwiki.com/ggml-org/llama.cpp/4.5-sycl-backend</a></p>\n<p>[68] DeepWiki - Flash Attention and Optimizations. <a href=\"https://deepwiki.com/ggml-org/llama.cpp/7.4-flash-attention-and-optimizations\">https://deepwiki.com/ggml-org/llama.cpp/7.4-flash-attention-and-optimizations</a></p>\n<p>[69] Medium - llama.cpp performance breakthrough for multi-GPU setups. <a href=\"https://medium.com/@jagusztinl/llama-cpp-performance-breakthrough-for-multi-gpu-setups-04c83a66feb2\">https://medium.com/@jagusztinl/llama-cpp-performance-breakthrough-for-multi-gpu-setups-04c83a66feb2</a></p>\n<p>[70] llama.cpp GitHub - Quantize README. <a href=\"https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md\">https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md</a></p>\n<p>[71] Enclave AI - The Practical Quantization Guide for iPhone and Mac. <a href=\"https://enclaveai.app/blog/2025/11/12/practical-quantization-guide-iphone-mac-gguf/\">https://enclaveai.app/blog/2025/11/12/practical-quantization-guide-iphone-mac-gguf/</a></p>\n<p>[72] llama.cpp GitHub Discussion - Difference in quantization methods. <a href=\"https://github.com/ggml-org/llama.cpp/discussions/2094\">https://github.com/ggml-org/llama.cpp/discussions/2094</a></p>\n<p>[73] AIMultiple Research - LLM Inference Engines: vLLM vs LMDeploy vs SGLang. <a href=\"https://research.aimultiple.com/inference-engines/\">https://research.aimultiple.com/inference-engines/</a></p>",
            "url": "https://blog.ecitis.org/local-ai-inferencing-techniques/",
            "title": "Advanced Techniques for Local AI Inferencing",
            "summary": "Make local models faster and smaller with quantization, speculative decoding, attention optimization, caching, and runtime selection.",
            "image": "https://blog.ecitis.org/open-graph/local-ai-inferencing-techniques.png",
            "date_modified": "2026-01-18T00:00:00.000Z",
            "date_published": "2026-01-18T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Inference & Deployment",
                "Local AI",
                "Quantization",
                "Inference"
            ]
        },
        {
            "id": "https://blog.ecitis.org/llm-finetuning-guide/",
            "content_html": "<p>Fine-tuning large language models has become accessible to individual developers and small teams. What once required clusters of expensive GPUs can now run on consumer hardware with the right techniques and tools. This guide covers the practical aspects of fine-tuning: when to do it, how to do it efficiently, and which tools to use.</p>\n<h2 id=\"when-fine-tuning-makes-sense\">When Fine-Tuning Makes Sense</h2>\n<figure><figcaption><strong>When to Fine-Tune</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/llm-finetuning-guide-0.webp\" width=\"512\" height=\"572\" alt=\"When to Fine-Tune\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>Before investing in fine-tuning, consider whether simpler approaches will work. The decision framework looks like this:</p>\n<p><strong>Start with prompt engineering.</strong> If you can solve your problem by crafting better prompts, do that. Prompt engineering takes hours to days and requires no infrastructure changes. It works well for prototypes, MVPs, and cases where the base model’s capabilities are sufficient.</p>\n<p><strong>Use RAG when you need external knowledge.</strong> If your model hallucinates or lacks company-specific information, Retrieval-Augmented Generation is the answer. RAG connects the model to a knowledge base and retrieves relevant context before generating responses. Customer service chatbots and documentation assistants benefit from this approach.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-1\" id=\"user-content-fnref-1\">1</a></sup></p>\n<p><strong>Fine-tune when you need behavioral changes.</strong> If your system fails on reasoning, planning, or strict policies even with good prompts and retrieval, fine-tuning makes the difference. Use cases include:</p>\n<ul>\n<li>Interpreting clinical notes or legal documents</li>\n<li>Enforcing a specific output format consistently</li>\n<li>Teaching domain-specific terminology and relationships</li>\n<li>Matching a particular writing style or tone</li>\n</ul>\n<p>Fine-tuning is also appropriate when you want to distill capabilities from a larger model into a smaller one for faster inference or reduced costs.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-2\" id=\"user-content-fnref-2\">2</a></sup></p>\n<p>The tradeoff is clear: prompt engineering is a sprint, RAG is a marathon with hydration stations, and fine-tuning is building a Formula 1 car from scratch. Each serves different purposes.</p>\n<h2 id=\"types-of-fine-tuning\">Types of Fine-Tuning</h2>\n<h3 id=\"full-fine-tuning\">Full Fine-Tuning</h3>\n<p>Full fine-tuning updates every parameter in the model. For a 7B parameter model, this requires over 28GB of GPU memory just for the weights in full precision. Training consumes additional memory for gradients and optimizer states.</p>\n<p>The results can be excellent. Full fine-tuning gives the model maximum flexibility to adapt to new tasks. But the costs are significant:</p>\n<ul>\n<li>High compute requirements (multiple high-end GPUs)</li>\n<li>Risk of catastrophic forgetting (the model forgets pre-trained knowledge)</li>\n<li>One full-size checkpoint per task<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-3\" id=\"user-content-fnref-3\">3</a></sup></li>\n</ul>\n<h3 id=\"parameter-efficient-fine-tuning\">Parameter-Efficient Fine-Tuning</h3>\n<p>PEFT methods update only a fraction of parameters while keeping the base model frozen. The Hugging Face PEFT library demonstrates this: training bigscience/mt0-large with PEFT touches just 2,359,296 parameters out of 1,231,940,608 total. That is 0.19% of the model.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-4\" id=\"user-content-fnref-4\">4</a></sup></p>\n<p>Performance remains competitive. With just 200 parameters projected into a space of millions of dimensions, you can achieve 90% of full fine-tuning performance. Some studies show PEFT techniques outperforming full fine-tuning on certain code intelligence tasks.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-5\" id=\"user-content-fnref-5\">5</a></sup></p>\n<h3 id=\"lora\">LoRA</h3>\n<p>Low-Rank Adaptation, introduced by Microsoft Research in 2021, freezes the pre-trained weights and injects trainable low-rank decomposition matrices into transformer layers. Instead of updating billions of parameters, you train small adapter matrices amounting to 1-5% of the original parameters.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-6\" id=\"user-content-fnref-6\">6</a></sup></p>\n<p>The key insight: weight updates during fine-tuning have low intrinsic rank. You can represent them as the product of two smaller matrices (A and B) without losing much information.</p>\n<p>LoRA has virtually no downsides for most use cases: memory usage is minimal, training is fast, and quality is high. The adapter weights can merge into the base model for inference, eliminating latency overhead.</p>\n<h3 id=\"qlora\">QLoRA</h3>\n<p>QLoRA extends LoRA by adding 4-bit quantization of the base model. The frozen weights are stored in 4-bit precision while LoRA adapters train in higher precision. Gradients backpropagate through the quantized model.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-7\" id=\"user-content-fnref-7\">7</a></sup></p>\n<p>Three innovations make QLoRA work:</p>\n<ol>\n<li><strong>NF4 quantization</strong>: Uses the known distribution of neural network weights (zero-centered normal) to quantize effectively</li>\n<li><strong>Double quantization</strong>: Quantizes the quantization constants themselves to reduce memory overhead</li>\n<li><strong>Paged optimizers</strong>: Manages memory spikes during training</li>\n</ol>\n<p>QLoRA enables fine-tuning 70B parameter models on hardware that would struggle with 7B models using full fine-tuning. A single A100 80GB handles models that would otherwise require 4-8 GPUs.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-8\" id=\"user-content-fnref-8\">8</a></sup></p>\n<p>The tradeoff: QLoRA achieves 80-90% of full fine-tuning quality compared to LoRA’s 90-95%. For many applications, this is acceptable given the massive resource savings.</p>\n<h3 id=\"adapters\">Adapters</h3>\n<p>Adapters are small neural networks inserted into transformer layers, typically after attention or feed-forward sublayers. They have a bottleneck architecture similar to autoencoders: the input is projected down to a smaller dimension, transformed, then projected back up.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-9\" id=\"user-content-fnref-9\">9</a></sup></p>\n<p>The original adapter paper showed BERT trained with adapters reached performance comparable to full fine-tuning while training only 3.6% of parameters. More recent work demonstrates that adapter-based PEFT in 7B parameter models can match or exceed the zero-shot performance of 175B parameter models on reasoning tasks.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-10\" id=\"user-content-fnref-10\">10</a></sup></p>\n<h2 id=\"unsloth\">Unsloth</h2>\n<p><a href=\"https://unsloth.ai/\">Unsloth</a> is an open-source library that accelerates LLM fine-tuning. The claims are substantial: 2x faster training with 70% less VRAM compared to standard Hugging Face methods.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-11\" id=\"user-content-fnref-11\">11</a></sup></p>\n<h3 id=\"why-unsloth-is-fast\">Why Unsloth is Fast</h3>\n<p>The speed comes from several optimizations:</p>\n<ul>\n<li>Custom Triton kernels for RoPE embeddings and MLP layers</li>\n<li>Fused operations that reduce memory transfers</li>\n<li>“Uncontaminated sequence packing” that combines sequences efficiently</li>\n<li>Chunked cross-entropy loss computation</li>\n</ul>\n<p>A December 2025 update combined these optimizations to achieve up to 3x faster training throughput with 60% lower VRAM usage.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-12\" id=\"user-content-fnref-12\">12</a></sup></p>\n<h3 id=\"supported-models\">Supported Models</h3>\n<p>Unsloth supports the models people actually use:</p>\n<ul>\n<li>Llama 3.x and Llama 4 (including multimodal variants)</li>\n<li>Qwen 2.5 and Qwen 3</li>\n<li>Mistral and Mixtral</li>\n<li>Gemma 2 and Gemma 3</li>\n<li>DeepSeek models</li>\n<li>Phi-3 and Phi-4</li>\n<li>Vision models: Llama 3.2 Vision, Qwen 2.5 VL, Pixtral</li>\n</ul>\n<p>The library also supports NVIDIA GPUs from Tesla T4 to H100, with portability to AMD and Intel GPUs.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-13\" id=\"user-content-fnref-13\">13</a></sup></p>\n<h3 id=\"memory-requirements\">Memory Requirements</h3>\n<p>VRAM requirements depend on model size and quantization:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>QLoRA 4-bit</th><th>LoRA 16-bit</th></tr></thead><tbody><tr><td>7-9B parameters</td><td>6.5GB</td><td>24GB</td></tr><tr><td>20B parameters</td><td>14GB</td><td>-</td></tr><tr><td>70B parameters</td><td>48GB</td><td>-</td></tr><tr><td>120B parameters</td><td>65GB</td><td>-</td></tr></tbody></table></div>\n<p>These numbers make consumer GPU fine-tuning realistic. A 24GB RTX 4090 handles most 7-9B parameter models with LoRA.</p>\n<h3 id=\"installation\">Installation</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>pip</span><span> install</span><span> unsloth</span></span></code></pre>\n<p>For Conda environments:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>conda</span><span> create</span><span> --name</span><span> unsloth</span><span> python=</span><span>3.11</span><span> pytorch-cuda=</span><span>12.1</span><span> pytorch</span><span> cudatoolkit</span><span> xformers</span><span> -c</span><span> pytorch</span><span> -c</span><span> nvidia</span><span> -c</span><span> xformers</span><span> -y</span></span>\n<span class=\"line\"><span>conda</span><span> activate</span><span> unsloth</span></span>\n<span class=\"line\"><span>pip</span><span> install</span><span> unsloth</span></span></code></pre>\n<h3 id=\"code-example\">Code Example</h3>\n<p>Here is a complete example fine-tuning Llama 3.2 3B with QLoRA:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> unsloth </span><span>import</span><span> FastLanguageModel</span></span>\n<span class=\"line\"><span>from</span><span> trl </span><span>import</span><span> SFTTrainer</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> TrainingArguments</span></span>\n<span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model with 4-bit quantization</span></span>\n<span class=\"line\"><span>model</span><span>,</span><span> tokenizer </span><span>=</span><span> FastLanguageModel</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    model_name</span><span>=</span><span>\"unsloth/Llama-3.2-3B-Instruct-bnb-4bit\"</span><span>,</span></span>\n<span class=\"line\"><span>    max_seq_length</span><span>=</span><span>2048</span><span>,</span></span>\n<span class=\"line\"><span>    load_in_4bit</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Add LoRA adapters</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> FastLanguageModel</span><span>.</span><span>get_peft_model</span><span>(</span></span>\n<span class=\"line\"><span>    model,</span></span>\n<span class=\"line\"><span>    r</span><span>=</span><span>16</span><span>,  </span><span># LoRA rank</span></span>\n<span class=\"line\"><span>    target_modules</span><span>=</span><span>[</span><span>\"q_proj\"</span><span>, </span><span>\"k_proj\"</span><span>, </span><span>\"v_proj\"</span><span>, </span><span>\"o_proj\"</span><span>,</span></span>\n<span class=\"line\"><span>                    \"gate_proj\"</span><span>, </span><span>\"up_proj\"</span><span>, </span><span>\"down_proj\"</span><span>],</span></span>\n<span class=\"line\"><span>    lora_alpha</span><span>=</span><span>16</span><span>,</span></span>\n<span class=\"line\"><span>    lora_dropout</span><span>=</span><span>0</span><span>,</span></span>\n<span class=\"line\"><span>    bias</span><span>=</span><span>\"none\"</span><span>,</span></span>\n<span class=\"line\"><span>    use_gradient_checkpointing</span><span>=</span><span>\"unsloth\"</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load and format dataset</span></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> load_dataset</span><span>(</span><span>\"yahma/alpaca-cleaned\"</span><span>, split</span><span>=</span><span>\"train\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> format_prompt</span><span>(</span><span>example</span><span>):</span></span>\n<span class=\"line\"><span>    instruction </span><span>=</span><span> example</span><span>[</span><span>\"instruction\"</span><span>]</span></span>\n<span class=\"line\"><span>    input_text </span><span>=</span><span> example</span><span>[</span><span>\"input\"</span><span>]</span></span>\n<span class=\"line\"><span>    output </span><span>=</span><span> example</span><span>[</span><span>\"output\"</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    if</span><span> input_text</span><span>:</span></span>\n<span class=\"line\"><span>        text </span><span>=</span><span> f</span><span>\"\"\"### Instruction:</span></span>\n<span class=\"line\"><span>{</span><span>instruction</span><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>### Input:</span></span>\n<span class=\"line\"><span>{</span><span>input_text</span><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>### Response:</span></span>\n<span class=\"line\"><span>{</span><span>output</span><span>}</span><span>\"\"\"</span></span>\n<span class=\"line\"><span>    else</span><span>:</span></span>\n<span class=\"line\"><span>        text </span><span>=</span><span> f</span><span>\"\"\"### Instruction:</span></span>\n<span class=\"line\"><span>{</span><span>instruction</span><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>### Response:</span></span>\n<span class=\"line\"><span>{</span><span>output</span><span>}</span><span>\"\"\"</span></span>\n<span class=\"line\"><span>    return</span><span> {</span><span>\"text\"</span><span>:</span><span> text</span><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>dataset </span><span>=</span><span> dataset</span><span>.</span><span>map</span><span>(format_prompt)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Configure training</span></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> SFTTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>model,</span></span>\n<span class=\"line\"><span>    tokenizer</span><span>=</span><span>tokenizer,</span></span>\n<span class=\"line\"><span>    train_dataset</span><span>=</span><span>dataset,</span></span>\n<span class=\"line\"><span>    dataset_text_field</span><span>=</span><span>\"text\"</span><span>,</span></span>\n<span class=\"line\"><span>    max_seq_length</span><span>=</span><span>2048</span><span>,</span></span>\n<span class=\"line\"><span>    args</span><span>=</span><span>TrainingArguments</span><span>(</span></span>\n<span class=\"line\"><span>        per_device_train_batch_size</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>        gradient_accumulation_steps</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>        warmup_steps</span><span>=</span><span>10</span><span>,</span></span>\n<span class=\"line\"><span>        max_steps</span><span>=</span><span>100</span><span>,</span></span>\n<span class=\"line\"><span>        learning_rate</span><span>=</span><span>2e-4</span><span>,</span></span>\n<span class=\"line\"><span>        fp16</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>        logging_steps</span><span>=</span><span>10</span><span>,</span></span>\n<span class=\"line\"><span>        output_dir</span><span>=</span><span>\"outputs\"</span><span>,</span></span>\n<span class=\"line\"><span>    ),</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Train</span></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>train</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Save adapter weights</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>save_pretrained</span><span>(</span><span>\"lora_model\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># For inference</span></span>\n<span class=\"line\"><span>FastLanguageModel</span><span>.</span><span>for_inference</span><span>(model)</span></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> tokenizer</span><span>(</span><span>\"### Instruction:\\nWrite a haiku about coding.\\n\\n### Response:\\n\"</span><span>,</span></span>\n<span class=\"line\"><span>                   return_tensors</span><span>=</span><span>\"pt\"</span><span>).</span><span>to</span><span>(</span><span>\"cuda\"</span><span>)</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(</span><span>**</span><span>inputs, max_new_tokens</span><span>=</span><span>64</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(tokenizer.</span><span>decode</span><span>(outputs[</span><span>0</span><span>], skip_special_tokens</span><span>=</span><span>True</span><span>))</span></span></code></pre>\n<p>The <code>FastLanguageModel.for_inference()</code> call enables Unsloth’s optimized inference mode, providing 2x speedup over standard generation.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-14\" id=\"user-content-fnref-14\">14</a></sup></p>\n<h2 id=\"thinking-machines-tinker\">Thinking Machines’ Tinker</h2>\n<p><a href=\"https://thinkingmachines.ai/\">Tinker</a> is a training API from Mira Murati’s Thinking Machines Lab, released in October 2025. The premise is handling infrastructure complexity while giving you control over algorithms and data.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-15\" id=\"user-content-fnref-15\">15</a></sup></p>\n<h3 id=\"how-tinker-works\">How Tinker Works</h3>\n<p>You write simple Python scripts with four core functions, and Tinker runs them across distributed GPUs. The platform handles:</p>\n<ul>\n<li>Multi-GPU and multi-node distribution</li>\n<li>Checkpoint management</li>\n<li>Memory optimization</li>\n<li>Gradient synchronization</li>\n</ul>\n<p>Research teams at Princeton, Stanford, and Berkeley use Tinker for their work. The platform supports models up to 1 trillion parameters, including Kimi K2.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-16\" id=\"user-content-fnref-16\">16</a></sup></p>\n<h3 id=\"supported-training-approaches\">Supported Training Approaches</h3>\n<p>Tinker supports the full range of post-training methods:</p>\n<p><strong>Supervised fine-tuning</strong> for instruction following and task adaptation</p>\n<p><strong>Preference learning</strong> with a three-stage RLHF pipeline:</p>\n<ol>\n<li>Supervised fine-tuning on demonstrations</li>\n<li>Training a reward model on preferences</li>\n<li>RL optimization against the reward model</li>\n</ol>\n<p><strong>Prompt distillation</strong> for internalizing long instructions into model weights</p>\n<p><strong>Multi-agent optimization</strong> for training models to interact with other models or themselves</p>\n<p>The <a href=\"https://github.com/thinking-machines-lab/tinker-cookbook\">Tinker cookbook</a> on GitHub contains examples for these approaches.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-17\" id=\"user-content-fnref-17\">17</a></sup></p>\n<h3 id=\"lora-research\">LoRA Research</h3>\n<p>Thinking Machines published research on LoRA hyperparameters. Key finding: when picking optimal learning rates for each setting, training progresses almost identically for LoRAs with different sizes and full fine-tuning. Similar results appeared on AIME 2024 and AIME 2025 evaluations.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-18\" id=\"user-content-fnref-18\">18</a></sup></p>\n<h2 id=\"other-fine-tuning-tools\">Other Fine-Tuning Tools</h2>\n<h3 id=\"hugging-face-trl\">Hugging Face TRL</h3>\n<p>TRL (Transformer Reinforcement Learning) is Hugging Face’s library for post-training foundation models. It supports supervised fine-tuning, GRPO (Group Relative Policy Optimization), DPO (Direct Preference Optimization), and PPO.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-19\" id=\"user-content-fnref-19\">19</a></sup></p>\n<p>Key features:</p>\n<ul>\n<li>Built on the Transformers ecosystem</li>\n<li>Native support for distributed training (DDP, DeepSpeed, FSDP)</li>\n<li>Integration with PEFT for memory-efficient training</li>\n<li>OpenEnv integration for RL and agentic workflows</li>\n</ul>\n<p>The GRPOTrainer implements the algorithm used to train DeepSeek’s R1. It is more memory-efficient than PPO.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-20\" id=\"user-content-fnref-20\">20</a></sup></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> trl </span><span>import</span><span> SFTTrainer</span><span>,</span><span> DPOTrainer</span><span>,</span><span> GRPOTrainer</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoModelForCausalLM</span><span>,</span><span> AutoTokenizer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Basic SFT example</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.2-3B\"</span><span>)</span></span>\n<span class=\"line\"><span>tokenizer </span><span>=</span><span> AutoTokenizer</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"meta-llama/Llama-3.2-3B\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>trainer </span><span>=</span><span> SFTTrainer</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>model,</span></span>\n<span class=\"line\"><span>    train_dataset</span><span>=</span><span>dataset,</span></span>\n<span class=\"line\"><span>    tokenizer</span><span>=</span><span>tokenizer,</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>trainer</span><span>.</span><span>train</span><span>()</span></span></code></pre>\n<p>The combination of TRL and PEFT enables fine-tuning gpt-neo-x 20B (40GB in bfloat16) on a 24GB consumer GPU.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-21\" id=\"user-content-fnref-21\">21</a></sup></p>\n<h3 id=\"axolotl\">Axolotl</h3>\n<p><a href=\"https://github.com/axolotl-ai-cloud/axolotl\">Axolotl</a> is Modal’s recommended framework for beginners. It offers flexibility, ease of use, and rapid adoption of new models and techniques.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-22\" id=\"user-content-fnref-22\">22</a></sup></p>\n<p>2025 brought significant updates:</p>\n<ul>\n<li>February: LoRA optimizations for memory and speed, GRPO support</li>\n<li>May: Quantization Aware Training (QAT)</li>\n<li>August: NVFP4 support, GPT-OSS model support</li>\n<li>ND Parallelism combining Context Parallelism, Tensor Parallelism, and FSDP</li>\n</ul>\n<p>Axolotl supports multi-GPU training, unlike Unsloth. If you have a large GPU cluster, Axolotl is the better choice.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-23\" id=\"user-content-fnref-23\">23</a></sup></p>\n<p>Configuration uses YAML files that span the full pipeline: dataset preprocessing, training, evaluation, quantization, and inference.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>base_model</span><span>:</span><span> meta-llama/Llama-3.2-3B-Instruct</span></span>\n<span class=\"line\"><span>model_type</span><span>:</span><span> AutoModelForCausalLM</span></span>\n<span class=\"line\"><span>tokenizer_type</span><span>:</span><span> AutoTokenizer</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>load_in_4bit</span><span>:</span><span> true</span></span>\n<span class=\"line\"><span>adapter</span><span>:</span><span> qlora</span></span>\n<span class=\"line\"><span>lora_r</span><span>:</span><span> 32</span></span>\n<span class=\"line\"><span>lora_alpha</span><span>:</span><span> 16</span></span>\n<span class=\"line\"><span>lora_dropout</span><span>:</span><span> 0.05</span></span>\n<span class=\"line\"><span>lora_target_modules</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>q_proj</span></span>\n<span class=\"line\"><span>  - </span><span>v_proj</span></span>\n<span class=\"line\"><span>  - </span><span>k_proj</span></span>\n<span class=\"line\"><span>  - </span><span>o_proj</span></span>\n<span class=\"line\"><span>  - </span><span>gate_proj</span></span>\n<span class=\"line\"><span>  - </span><span>down_proj</span></span>\n<span class=\"line\"><span>  - </span><span>up_proj</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>datasets</span><span>:</span></span>\n<span class=\"line\"><span>  - </span><span>path</span><span>:</span><span> yahma/alpaca-cleaned</span></span>\n<span class=\"line\"><span>    type</span><span>:</span><span> alpaca</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>sequence_len</span><span>:</span><span> 4096</span></span>\n<span class=\"line\"><span>micro_batch_size</span><span>:</span><span> 2</span></span>\n<span class=\"line\"><span>gradient_accumulation_steps</span><span>:</span><span> 4</span></span>\n<span class=\"line\"><span>num_epochs</span><span>:</span><span> 3</span></span>\n<span class=\"line\"><span>learning_rate</span><span>:</span><span> 2e-4</span></span>\n<span class=\"line\"><span>optimizer</span><span>:</span><span> adamw_torch</span></span>\n<span class=\"line\"><span>lr_scheduler</span><span>:</span><span> cosine</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>output_dir</span><span>:</span><span> ./outputs</span></span></code></pre>\n<h3 id=\"llama-factory\">LLaMA-Factory</h3>\n<p><a href=\"https://github.com/hiyouga/LlamaFactory\">LLaMA-Factory</a> provides a WebUI for fine-tuning, making it accessible to non-technical users. The toolkit supports over 100 models and received the ACL 2024 best paper award.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-24\" id=\"user-content-fnref-24\">24</a></sup></p>\n<p>2025 additions include:</p>\n<ul>\n<li>Orthogonal Finetuning (OFT and OFTv2) in August</li>\n<li>GPT-OSS and Intern-S1-mini support</li>\n<li>GLM-4.1V, Qwen3, InternVL3, Llama 4, Qwen2.5-Omni</li>\n</ul>\n<p>Training approaches span supervised fine-tuning, continuous pre-training, and preference tuning (PPO, DPO, KTO, ORPO). The toolkit integrates FlashAttention-2, DeepSpeed, GaLore, and BAdam optimization.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-25\" id=\"user-content-fnref-25\">25</a></sup></p>\n<p>Inference uses OpenAI-style API, Gradio UI, or CLI with vLLM or SGLang workers.</p>\n<h3 id=\"torchtune\">torchtune</h3>\n<p><a href=\"https://github.com/meta-pytorch/torchtune\">torchtune</a> is PyTorch’s native post-training library. It offers composable building blocks without framework abstractions.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-26\" id=\"user-content-fnref-26\">26</a></sup></p>\n<p>If you prefer working directly with PyTorch, torchtune is the choice. The library is designed with memory efficiency in mind, with recipes tested on consumer GPUs with 24GB VRAM.</p>\n<p>Features include:</p>\n<ul>\n<li>LoRA fine-tuning on single device</li>\n<li>Knowledge distillation</li>\n<li>DPO training</li>\n<li>Multi-node training (added February 2025)</li>\n<li>Export to ExecuTorch for mobile and edge inference</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Single-device LoRA fine-tuning</span></span>\n<span class=\"line\"><span>tune</span><span> run</span><span> lora_finetune_single_device</span><span> --config</span><span> llama3_2/3B_lora_single_device</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Knowledge distillation</span></span>\n<span class=\"line\"><span>tune</span><span> run</span><span> knowledge_distillation_distributed</span><span> --config</span><span> qwen2/1.5B_to_0.5B_KD_lora_distributed</span></span></code></pre>\n<h2 id=\"data-preparation\">Data Preparation</h2>\n<h3 id=\"dataset-formats\">Dataset Formats</h3>\n<p>Two formats dominate: Alpaca and ShareGPT.</p>\n<p><strong>Alpaca format</strong> suits instruction-following tasks. Each example has three fields:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"instruction\"</span><span>:</span><span> \"Write a function to calculate factorial\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"input\"</span><span>:</span><span> \"5\"</span><span>,</span></span>\n<span class=\"line\"><span>  \"output\"</span><span>:</span><span> \"def factorial(n):\\n    if n &lt;= 1:\\n        return 1\\n    return n * factorial(n-1)\\n\\nprint(factorial(5))  # 120\"</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>The original Alpaca dataset contains 52,000 instruction-output pairs generated by GPT-4 from 175 seed instructions.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-27\" id=\"user-content-fnref-27\">27</a></sup></p>\n<p><strong>ShareGPT format</strong> handles multi-turn conversations:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>{</span></span>\n<span class=\"line\"><span>  \"conversations\"</span><span>:</span><span> [</span></span>\n<span class=\"line\"><span>    { </span><span>\"role\"</span><span>:</span><span> \"user\"</span><span>,</span><span> \"content\"</span><span>:</span><span> \"What is the capital of France?\"</span><span> }</span><span>,</span></span>\n<span class=\"line\"><span>    { </span><span>\"role\"</span><span>:</span><span> \"assistant\"</span><span>,</span><span> \"content\"</span><span>:</span><span> \"Paris is the capital of France.\"</span><span> }</span><span>,</span></span>\n<span class=\"line\"><span>    { </span><span>\"role\"</span><span>:</span><span> \"user\"</span><span>,</span><span> \"content\"</span><span>:</span><span> \"What is its population?\"</span><span> }</span><span>,</span></span>\n<span class=\"line\"><span>    {</span></span>\n<span class=\"line\"><span>      \"role\"</span><span>:</span><span> \"assistant\"</span><span>,</span></span>\n<span class=\"line\"><span>      \"content\"</span><span>:</span><span> \"Paris has a population of about 2.1 million in the city proper, and over 12 million in the metropolitan area.\"</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>  ]</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p>ShareGPT supports additional roles: human, GPT, observation, and function. Human and observation entries appear in odd positions; GPT and function in even positions.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-28\" id=\"user-content-fnref-28\">28</a></sup></p>\n<p>Use Alpaca for single-turn question-and-answer tasks. Use ShareGPT for conversational chatbots, especially those with function calling.</p>\n<h3 id=\"data-quality\">Data Quality</h3>\n<p>Quality matters more than quantity. A few thousand well-curated examples often outperform hundreds of thousands of noisy ones.</p>\n<p>Check for:</p>\n<ul>\n<li>Correct answers (verify factual claims)</li>\n<li>Consistent formatting</li>\n<li>Diverse task coverage</li>\n<li>Balanced class distribution</li>\n<li>No personally identifiable information</li>\n</ul>\n<p>Clean datasets exist for common starting points. <a href=\"https://github.com/gururise/AlpacaDataCleaned\">AlpacaDataCleaned</a> removes problematic examples from the original Alpaca dataset.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-29\" id=\"user-content-fnref-29\">29</a></sup></p>\n<h3 id=\"synthetic-data-generation\">Synthetic Data Generation</h3>\n<p>When human-labeled data is scarce, synthetic data fills the gap. Teacher-student distillation uses a larger model (like GPT-4) to generate training examples for a smaller model.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-30\" id=\"user-content-fnref-30\">30</a></sup></p>\n<p>Strategies for synthetic data:</p>\n<p><strong>Question-answer generation</strong>: Use retrieval-augmented pipelines to generate QA pairs from documents. A multi-stage framework with retriever, generator, and refinement model produces higher quality than single-pass generation.</p>\n<p><strong>Active synthetic data generation</strong>: Generate data iteratively based on the current model’s weaknesses. Simple selection criteria from active learning perform well.</p>\n<p><strong>Rephrasing</strong>: Augment existing questions with paraphrased versions. This is robust even with weaker augmentation models.</p>\n<p>Challenges exist. LLMs can hallucinate, producing incorrect labels or logically inconsistent examples. Validate synthetic data before training, either through automated checks or sampling for human review.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-31\" id=\"user-content-fnref-31\">31</a></sup></p>\n<p>Red Hat’s <a href=\"https://developers.redhat.com/articles/2025/11/25/building-domain-specific-llms-synthetic-data-and-sdg-hub\">SDG Hub</a> provides an open-source toolkit for synthetic data workflows.</p>\n<h2 id=\"training-configurations\">Training Configurations</h2>\n<h3 id=\"learning-rate\">Learning Rate</h3>\n<p>The learning rate is the most important hyperparameter. For LoRA and QLoRA fine-tuning:</p>\n<ul>\n<li>Start with 2e-4 for small models (7B)</li>\n<li>Scale down to 2e-5 for larger models (70B+)</li>\n<li>If training is unstable, reduce to 1e-4 or 3e-5</li>\n</ul>\n<p>Learning rate of 1e-4 has become standard for LoRA fine-tuning. Reducing it further helps with occasional training loss instabilities.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-32\" id=\"user-content-fnref-32\">32</a></sup></p>\n<p>Use warmup steps (5-10% of total steps) and a scheduler. Cosine annealing works well for most cases.</p>\n<h3 id=\"batch-size\">Batch Size</h3>\n<p>Maximize tokens-per-second without running out of memory. Use gradient accumulation to achieve larger effective batch sizes.</p>\n<p>A common configuration:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>per_device_train_batch_size </span><span>=</span><span> 4</span></span>\n<span class=\"line\"><span>gradient_accumulation_steps </span><span>=</span><span> 4</span></span>\n<span class=\"line\"><span># Effective batch size = 16</span></span></code></pre>\n<p>Larger batch sizes provide more stable gradients but require more memory. If you hit OOM errors, reduce per-device batch size and increase gradient accumulation.</p>\n<h3 id=\"lora-rank-and-alpha\">LoRA Rank and Alpha</h3>\n<p><strong>Rank (r)</strong>: Controls the capacity of the adapter. Common values are 8, 16, 32, or 64.</p>\n<ul>\n<li>Higher rank = more parameters = more capacity for diverse tasks</li>\n<li>Lower rank = fewer parameters = faster training, less overfitting risk</li>\n<li>Research suggests the most critical factor is applying LoRA to all linear transformer layers, not the specific rank value<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-33\" id=\"user-content-fnref-33\">33</a></sup></li>\n</ul>\n<p><strong>Alpha</strong>: Scaling factor for LoRA weights. Recommendations vary:</p>\n<ul>\n<li>Set alpha equal to rank (alpha/rank = 1)</li>\n<li>Set alpha to 2x rank (alpha/rank = 2), as Microsoft does in their examples</li>\n<li>The original LoRA paper suggests fixing alpha at 16</li>\n</ul>\n<p>If rank is 16, start with alpha of 16 or 32. Adjust based on results.</p>\n<p><strong>Target modules</strong>: Apply LoRA to both attention and MLP layers for best results. The minimal set is attention projections (q, k, v, o). Adding gate_proj, up_proj, and down_proj improves performance.</p>\n<h3 id=\"number-of-epochs\">Number of Epochs</h3>\n<p>For instruction-based datasets, 1-3 epochs work well. Training beyond 3 epochs offers diminishing returns and increases overfitting risk.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-34\" id=\"user-content-fnref-34\">34</a></sup></p>\n<p>Monitor validation loss. If it starts increasing while training loss continues decreasing, you are overfitting.</p>\n<h2 id=\"evaluation-and-deployment\">Evaluation and Deployment</h2>\n<h3 id=\"benchmarks\">Benchmarks</h3>\n<p>No single benchmark suffices. Select benchmarks relevant to your use case:</p>\n<p><strong>Instruction following</strong>: MT-Bench, AlpacaEval, LMSYS Chatbot Arena</p>\n<p><strong>Reasoning</strong>: GSM8K (math), MATH (competition math), BBH (Big Bench Hard), DROP</p>\n<p><strong>Knowledge</strong>: MMLU (15,000+ multiple-choice questions across 57 subjects)</p>\n<p><strong>Truthfulness</strong>: TruthfulQA (800+ questions across 38 subjects)</p>\n<p><strong>Domain-specific</strong>: Build custom evaluation sets matching your deployment scenario<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-35\" id=\"user-content-fnref-35\">35</a></sup></p>\n<h3 id=\"llm-as-judge\">LLM-as-Judge</h3>\n<p>MT-Bench introduced using LLMs to evaluate other LLMs. GPT-4 serves as a judge to score response quality. This scales better than human evaluation for rapid iteration.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-36\" id=\"user-content-fnref-36\">36</a></sup></p>\n<p>Open-source alternatives exist. Prometheus is a Llama-2-Chat model fine-tuned on 100K GPT-4 feedback examples. It achieves comparable evaluation capabilities when given appropriate reference materials.<sup><a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fn-37\" id=\"user-content-fnref-37\">37</a></sup></p>\n<h3 id=\"tracking-progress\">Tracking Progress</h3>\n<p>Use AlpacaEval to compute win-rate deltas before and after fine-tuning. Track metrics on held-out validation sets throughout training.</p>\n<p>Tools for experiment tracking:</p>\n<ul>\n<li>Weights &amp; Biases</li>\n<li>MLflow</li>\n<li>TensorBoard</li>\n<li>LlamaBoard (LLaMA-Factory specific)</li>\n</ul>\n<h3 id=\"deployment\">Deployment</h3>\n<p>After fine-tuning, merge LoRA weights into the base model for inference:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># With Unsloth</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>save_pretrained_merged</span><span>(</span><span>\"merged_model\"</span><span>, tokenizer, save_method</span><span>=</span><span>\"merged_16bit\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or save as GGUF for llama.cpp</span></span>\n<span class=\"line\"><span>model</span><span>.</span><span>save_pretrained_gguf</span><span>(</span><span>\"model_gguf\"</span><span>, tokenizer, quantization_method</span><span>=</span><span>\"q4_k_m\"</span><span>)</span></span></code></pre>\n<p>Quantization reduces model size for deployment. Common options:</p>\n<ul>\n<li>GGUF format for llama.cpp and Ollama</li>\n<li>AWQ for vLLM</li>\n<li>GPTQ for text-generation-inference</li>\n</ul>\n<p>For production inference, use vLLM, SGLang, or TGI rather than Hugging Face Transformers directly.</p>\n<h3 id=\"continuous-monitoring\">Continuous Monitoring</h3>\n<p>Static benchmarks alone will not catch performance drift. Monitor in production:</p>\n<ul>\n<li>Response latency</li>\n<li>User feedback signals (thumbs up/down, regeneration requests)</li>\n<li>Domain-specific quality metrics</li>\n<li>Hallucination rates on known-answer queries</li>\n</ul>\n<p>Retrain periodically as your data distribution shifts.</p>\n<h2 id=\"references\">References</h2>\n<section class=\"footnotes\"><h2 class=\"sr-only\" id=\"footnote-label\">Footnotes</h2>\n<ol>\n<li id=\"user-content-fn-1\">\n<p><a href=\"https://www.news.aakashg.com/p/rag-vs-fine-tuning-vs-prompt-engineering\" rel=\"noopener noreferrer\">RAG vs. Fine-tuning vs. Prompt Engineering: The Complete Guide to AI Optimization</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-1\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-2\">\n<p><a href=\"https://moveo.ai/blog/fine-tuning-rag-or-prompt-engineering\" rel=\"noopener noreferrer\">Fine-Tuning, RAG, or Prompt Engineering? LLM Decision Guide - Moveo.AI</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-2\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-3\">\n<p><a href=\"https://apxml.com/courses/introduction-to-llm-fine-tuning/chapter-4-parameter-efficient-fine-tuning-peft/comparing-peft-full-fine-tuning\" rel=\"noopener noreferrer\">PEFT vs Full Fine-Tuning Comparison - APXML</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-3\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-4\">\n<p><a href=\"https://huggingface.co/blog/peft\" rel=\"noopener noreferrer\">Parameter-Efficient Fine-Tuning using PEFT - Hugging Face</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-4\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-5\">\n<p><a href=\"https://dl.acm.org/doi/10.1145/3714461\" rel=\"noopener noreferrer\">Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation - ACM</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-5\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-6\">\n<p><a href=\"https://www.analyticsvidhya.com/blog/2023/08/lora-and-qlora/\" rel=\"noopener noreferrer\">Parameter-Efficient Fine-Tuning of Large Language Models with LoRA and QLoRA - Analytics Vidhya</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-6\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-7\">\n<p><a href=\"https://introl.com/blog/fine-tuning-infrastructure-lora-qlora-peft-scale-guide-2025\" rel=\"noopener noreferrer\">Fine-Tuning Infrastructure: LoRA, QLoRA, and PEFT at Scale - Introl</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-7\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-8\">\n<p><a href=\"https://www.index.dev/blog/top-ai-fine-tuning-tools-lora-vs-qlora-vs-full\" rel=\"noopener noreferrer\">Best GenAI Fine-Tuning Tools for 2025 - LoRA vs QLoRA - Index.dev</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-8\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-9\">\n<p><a href=\"https://magazine.sebastianraschka.com/p/finetuning-llms-with-adapters\" rel=\"noopener noreferrer\">Finetuning LLMs with Adapters - Sebastian Raschka</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-9\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-10\">\n<p><a href=\"https://aclanthology.org/2023.emnlp-main.319/\" rel=\"noopener noreferrer\">LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning - ACL Anthology</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-10\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-11\">\n<p><a href=\"https://github.com/unslothai/unsloth\" rel=\"noopener noreferrer\">Unsloth GitHub Repository</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-11\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-12\">\n<p><a href=\"https://medium.com/coding-nexus/unsloth-ai-makes-llm-fine-tuning-3-faster-and-90-cheaper-love-you-0f32eb3d7b98\" rel=\"noopener noreferrer\">Unsloth AI Makes LLM Fine-Tuning 3x Faster and 90% Cheaper - Medium</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-12\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-13\">\n<p><a href=\"https://unsloth.ai/\" rel=\"noopener noreferrer\">Unsloth AI - Open Source Fine-tuning &amp; RL for LLMs</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-13\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-14\">\n<p><a href=\"https://medium.com/@matteo28/qlora-fine-tuning-with-unsloth-a-complete-guide-8652c9c7edb3\" rel=\"noopener noreferrer\">QLoRA Fine-Tuning with Unsloth: A Complete Guide - Medium</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-14\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-15\">\n<p><a href=\"https://siliconangle.com/2025/12/12/thinking-machines-makes-tinker-ai-fine-tuning-service-generally-available/\" rel=\"noopener noreferrer\">Thinking Machines makes its Tinker AI fine-tuning service generally available - SiliconANGLE</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-15\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-16\">\n<p><a href=\"https://www.deeplearning.ai/the-batch/thinking-machines-new-tinker-api-makes-it-easier-to-fine-tune-models-on-many-gpus/\" rel=\"noopener noreferrer\">Thinking Machines’ New Tinker API Makes It Easier To Fine-Tune Models On Many GPUs - DeepLearning.AI</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-16\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-17\">\n<p><a href=\"https://github.com/thinking-machines-lab/tinker-cookbook\" rel=\"noopener noreferrer\">Tinker Cookbook - GitHub</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-17\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-18\">\n<p><a href=\"https://thinkingmachines.ai/blog/lora/\" rel=\"noopener noreferrer\">LoRA Without Regret - Thinking Machines Lab</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-18\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-19\">\n<p><a href=\"https://huggingface.co/docs/trl/en/index\" rel=\"noopener noreferrer\">TRL - Transformer Reinforcement Learning - Hugging Face</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-19\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-20\">\n<p><a href=\"https://github.com/huggingface/trl\" rel=\"noopener noreferrer\">TRL GitHub Repository</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-20\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-21\">\n<p><a href=\"https://huggingface.co/blog/trl-peft\" rel=\"noopener noreferrer\">Fine-tuning 20B LLMs with RLHF on a 24GB consumer GPU - Hugging Face</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-21\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-22\">\n<p><a href=\"https://modal.com/blog/fine-tuning-llms\" rel=\"noopener noreferrer\">Best frameworks for fine-tuning LLMs in 2025 - Modal</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-22\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-23\">\n<p><a href=\"https://github.com/axolotl-ai-cloud/axolotl\" rel=\"noopener noreferrer\">Axolotl GitHub Repository</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-23\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-24\">\n<p><a href=\"https://github.com/hiyouga/LlamaFactory\" rel=\"noopener noreferrer\">LLaMA-Factory GitHub Repository</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-24\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-25\">\n<p><a href=\"https://llamafactory.readthedocs.io/en/latest/\" rel=\"noopener noreferrer\">LLaMA-Factory Documentation</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-25\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-26\">\n<p><a href=\"https://pytorch.org/blog/torchtune-fine-tune-llms/\" rel=\"noopener noreferrer\">torchtune: Easily fine-tune LLMs using PyTorch - PyTorch</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-26\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-27\">\n<p><a href=\"https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/datasets-guide\" rel=\"noopener noreferrer\">Datasets Guide - Unsloth Documentation</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-27\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-28\">\n<p><a href=\"https://blog.usee.ai/fine-tuning-llms-with-llama-factory-guidance-and-insights-5cc619f56a7f\" rel=\"noopener noreferrer\">Fine-Tuning LLMs with LLaMA-Factory: Guidance and Insights - Usee.ai</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-28\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-29\">\n<p><a href=\"https://github.com/gururise/AlpacaDataCleaned\" rel=\"noopener noreferrer\">AlpacaDataCleaned - GitHub</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-29\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-30\">\n<p><a href=\"https://scale.com/blog/synthetic-data-fine-tuning-llms\" rel=\"noopener noreferrer\">Synthetic Data Generation Strategies for Fine-Tuning LLMs - Scale</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-30\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-31\">\n<p><a href=\"https://labelyourdata.com/articles/llm-fine-tuning/synthetic-data\" rel=\"noopener noreferrer\">Synthetic Data: Benefits and Techniques for LLM Fine-Tuning in 2025 - Label Your Data</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-31\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-32\">\n<p><a href=\"https://magazine.sebastianraschka.com/p/practical-tips-for-finetuning-llms\" rel=\"noopener noreferrer\">Practical Tips for Finetuning LLMs Using LoRA - Sebastian Raschka</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-32\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-33\">\n<p><a href=\"https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/lora-hyperparameters-guide\" rel=\"noopener noreferrer\">LoRA fine-tuning Hyperparameters Guide - Unsloth Documentation</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-33\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-34\">\n<p><a href=\"https://amirteymoori.com/fine-tuning-llms-with-lora-a-practical-guide-for-2025/\" rel=\"noopener noreferrer\">Fine-Tuning LLMs with LoRA: 2025 Guide - Amir Teymoori</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-34\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-35\">\n<p><a href=\"https://labelyourdata.com/articles/llm-fine-tuning/llm-evaluation\" rel=\"noopener noreferrer\">LLM Evaluation: Benchmarks to Test Model Quality in 2025 - Label Your Data</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-35\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-36\">\n<p><a href=\"https://www.evidentlyai.com/llm-guide/llm-benchmarks\" rel=\"noopener noreferrer\">30 LLM evaluation benchmarks and how they work - Evidently AI</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-36\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n<li id=\"user-content-fn-37\">\n<p><a href=\"https://alopatenko.github.io/LLMEvaluation/\" rel=\"noopener noreferrer\">Awesome LLM Evaluation - GitHub</a> <a href=\"https://blog.ecitis.org/llm-finetuning-guide/#user-content-fnref-37\" class=\"data-footnote-backref\">↩</a></p>\n</li>\n</ol>\n</section>",
            "url": "https://blog.ecitis.org/llm-finetuning-guide/",
            "title": "Fine-Tuning LLMs: A Practical Guide to Tools, Techniques, and Best Practices",
            "summary": "Decide when fine-tuning is worth it, then choose the right data, method, tooling, and evaluation path for an efficient training run.",
            "image": "https://blog.ecitis.org/open-graph/llm-finetuning-guide.png",
            "date_modified": "2026-01-09T00:00:00.000Z",
            "date_published": "2026-01-09T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "Fine-tuning",
                "LoRA",
                "Training"
            ]
        },
        {
            "id": "https://blog.ecitis.org/vision-language-models/",
            "content_html": "<p>Vision Language Models (VLMs) combine computer vision and natural language processing into unified systems capable of reasoning about images and text together. Unlike earlier approaches that treated vision and language as separate pipelines, modern VLMs learn joint representations that enable tasks like visual question answering, image captioning, document analysis, and multimodal reasoning.</p>\n<p>This article covers VLM architecture from the ground up: how vision encoders transform pixels into tokens, how projection layers bridge modalities, and why inference is more complex than text-only LLMs. We also survey popular architectures, quantization strategies, and practical deployment options.</p>\n<hr />\n<h2 id=\"what-vlms-can-do\">What VLMs Can Do</h2>\n<p>VLMs accept images (or videos) alongside text prompts and generate text responses. Concrete capabilities include:</p>\n<ul>\n<li><strong>Visual question answering</strong>: “What color is the car in the background?”</li>\n<li><strong>Document understanding</strong>: Extract text, tables, and structure from PDFs, invoices, forms</li>\n<li><strong>Chart and diagram analysis</strong>: Interpret graphs, flowcharts, architectural diagrams</li>\n<li><strong>OCR and text extraction</strong>: Read text embedded in images without separate OCR pipelines</li>\n<li><strong>Object localization</strong>: Identify bounding boxes or regions corresponding to natural language queries</li>\n<li><strong>Image captioning</strong>: Generate descriptions of visual content</li>\n<li><strong>Multi-image reasoning</strong>: Compare multiple images, summarize differences</li>\n<li><strong>Video understanding</strong>: Summarize events, answer questions about temporal sequences</li>\n</ul>\n<p>The key shift from earlier vision models: VLMs use the same interface as chat LLMs. You prompt them with natural language, optionally include images, and receive text back. This unified interface simplifies integration into existing LLM-based applications.</p>\n<hr />\n<h2 id=\"how-vlms-work-layer-by-layer\">How VLMs Work: Layer by Layer</h2>\n<p>Most VLMs follow a three-component architecture:</p>\n<figure><figcaption><strong>Vision Encoder</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/vision-language-models-0.webp\" width=\"512\" height=\"302\" alt=\"Vision Encoder\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>Let’s examine each component.</p>\n<h3 id=\"vision-encoder\">Vision Encoder</h3>\n<p>The vision encoder transforms raw pixels into a sequence of feature vectors. The dominant architecture is the Vision Transformer (ViT), introduced by Dosovitskiy et al. in 2020.</p>\n<p><strong>How ViT works:</strong></p>\n<ol>\n<li>\n<p><strong>Patch extraction</strong>: The input image is divided into fixed-size patches (typically 14x14 or 16x16 pixels). A 224x224 image with 14x14 patches yields 256 patches.</p>\n</li>\n<li>\n<p><strong>Linear embedding</strong>: Each patch is flattened and projected through a linear layer to create patch embeddings. For a 14x14x3 patch (588 values), this produces a vector of dimension D (commonly 768, 1024, or 1152).</p>\n</li>\n<li>\n<p><strong>Position encoding</strong>: Since transformers have no inherent notion of spatial order, learnable position embeddings are added to each patch embedding. These encode the patch’s location in the original image grid.</p>\n</li>\n<li>\n<p><strong>CLS token</strong>: A learnable classification token is prepended to the sequence. After processing, this token aggregates global image information.</p>\n</li>\n<li>\n<p><strong>Transformer blocks</strong>: The sequence passes through standard transformer encoder layers (self-attention + feedforward). Each patch attends to all other patches, capturing global context.</p>\n</li>\n</ol>\n<p>The output is a sequence of vectors: one per patch plus the CLS token. For a 224x224 image with 14x14 patches, you get 257 vectors (256 patches + 1 CLS).</p>\n<p><strong>Popular vision encoders:</strong></p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Encoder</th><th>Parameters</th><th>Patch Size</th><th>Output Dim</th><th>Notes</th></tr></thead><tbody><tr><td>CLIP ViT-L/14</td><td>428M</td><td>14x14</td><td>1024</td><td>Trained on 400M image-text pairs</td></tr><tr><td>SigLIP-So400m</td><td>400M</td><td>14x14</td><td>1152</td><td>Sigmoid loss, better small-batch training</td></tr><tr><td>InternViT-6B</td><td>6B</td><td>14x14</td><td>3200</td><td>Scaled up for InternVL</td></tr><tr><td>DaViT</td><td>88M</td><td>varies</td><td>varies</td><td>Dual attention (spatial + channel)</td></tr></tbody></table></div>\n<p><strong>CLIP vs SigLIP</strong></p>\n<p>CLIP (Contrastive Language-Image Pre-training) trains vision and text encoders jointly using a contrastive softmax loss. Given a batch of N image-text pairs, the model learns to match correct pairs while pushing incorrect pairs apart. The softmax normalization requires computing similarities across the entire batch.</p>\n<p>SigLIP (Sigmoid Loss for Language-Image Pre-training) replaces softmax with sigmoid loss that operates on each image-text pair independently. This removes the need for cross-batch normalization, enabling:</p>\n<ul>\n<li>Better performance with smaller batch sizes</li>\n<li>Reduced memory usage (4096 batch on 4 TPUs vs 2048 for CLIP)</li>\n<li>Comparable or better accuracy on downstream tasks</li>\n</ul>\n<p>SigLIP 2 (February 2025) adds self-distillation, masked prediction, and captioning objectives for improved localization and dense prediction.</p>\n<h3 id=\"projectionadapter-layers\">Projection/Adapter Layers</h3>\n<p>Vision encoders and language models operate in different representation spaces. The projection layer bridges this gap by transforming vision features into the language model’s embedding space.</p>\n<p><strong>Simple linear projection (LLaVA 1.0)</strong></p>\n<p>The original LLaVA used a single linear layer:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>vision_features: [num_patches, 1024]</span></span>\n<span class=\"line\"><span>      │</span></span>\n<span class=\"line\"><span>      ▼</span></span>\n<span class=\"line\"><span> Linear(1024, 4096)</span></span>\n<span class=\"line\"><span>      │</span></span>\n<span class=\"line\"><span>      ▼</span></span>\n<span class=\"line\"><span>projected_features: [num_patches, 4096]</span></span></code></pre>\n<p>This works but limits how much the vision representation can be transformed.</p>\n<p><strong>MLP projection (LLaVA 1.5+)</strong></p>\n<p>LLaVA 1.5 switched to a two-layer MLP with GELU activation:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>vision_features: [num_patches, 1024]</span></span>\n<span class=\"line\"><span>      │</span></span>\n<span class=\"line\"><span>      ▼</span></span>\n<span class=\"line\"><span> Linear(1024, 4096)</span></span>\n<span class=\"line\"><span>      │</span></span>\n<span class=\"line\"><span>      ▼</span></span>\n<span class=\"line\"><span>    GELU</span></span>\n<span class=\"line\"><span>      │</span></span>\n<span class=\"line\"><span>      ▼</span></span>\n<span class=\"line\"><span> Linear(4096, 4096)</span></span>\n<span class=\"line\"><span>      │</span></span>\n<span class=\"line\"><span>      ▼</span></span>\n<span class=\"line\"><span>projected_features: [num_patches, 4096]</span></span></code></pre>\n<p>The additional nonlinearity allows more complex transformations and improved multimodal alignment.</p>\n<p><strong>Cross-attention adapters (Flamingo-style)</strong></p>\n<p>Some architectures use cross-attention layers that let language model tokens attend to vision features rather than concatenating them directly. This approach:</p>\n<ul>\n<li>Keeps vision and language computations more separate</li>\n<li>Allows selective attention to relevant image regions</li>\n<li>Can be more parameter-efficient</li>\n</ul>\n<p><strong>Perceiver resampler</strong></p>\n<p>Perceiver-based adapters use a fixed number of learnable query tokens that cross-attend to the variable-length vision features, producing a constant number of output tokens regardless of image resolution.</p>\n<h3 id=\"language-model-backbone\">Language Model Backbone</h3>\n<p>The language model processes the projected vision tokens alongside text tokens. Any decoder-only transformer can serve this role:</p>\n<ul>\n<li>LLaMA family (7B, 13B, 70B)</li>\n<li>Vicuna (instruction-tuned LLaMA)</li>\n<li>Qwen (7B, 14B, 72B)</li>\n<li>Phi-3 (3.8B)</li>\n<li>Gemma (2B, 7B, 27B)</li>\n<li>InternLM (7B, 20B)</li>\n</ul>\n<p>The vision tokens are typically inserted at the position of special <code>&lt;image&gt;</code> tokens in the input. The language model then processes the combined sequence autoregressively, attending to both vision and text tokens.</p>\n<p><strong>Token sequence structure:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>[BOS] [system prompt tokens] [image_tok_1] [image_tok_2] ... [image_tok_N] [user prompt tokens] [assistant response]</span></span></code></pre>\n<p>The number of image tokens depends on the vision encoder’s output and any token compression applied.</p>\n<h3 id=\"position-encoding-for-images\">Position Encoding for Images</h3>\n<p>Standard language models use 1D position encoding. Images require 2D spatial information. VLMs handle this in several ways:</p>\n<p><strong>Flattened 1D positions</strong>: Simply assign positions sequentially to flattened patch tokens. Row-major order: patch (0,0) gets position 0, patch (0,1) gets position 1, etc.</p>\n<p><strong>Separate 2D position embeddings in ViT</strong>: The vision encoder learns its own 2D-aware position embeddings during pretraining. These spatial relationships are then encoded into the patch representations before projection.</p>\n<p><strong>M-RoPE (Multimodal Rotary Position Embedding)</strong>: Qwen2-VL introduced M-RoPE, which encodes position information across text, images, and videos in a unified way. For images, it captures both spatial positions and temporal positions for video frames.</p>\n<p><strong>Variable Visual Position Encoding (V2PE)</strong>: InternVL3 uses smaller, more flexible position increments for visual tokens, improving long-context understanding.</p>\n<hr />\n<h2 id=\"popular-vlm-architectures\">Popular VLM Architectures</h2>\n<h3 id=\"llava-family\">LLaVA Family</h3>\n<p>LLaVA (Large Language and Vision Assistant) established the dominant paradigm: frozen CLIP encoder + MLP projector + instruction-tuned LLM.</p>\n<p><strong>LLaVA 1.0 (2023)</strong></p>\n<ul>\n<li>Vision: CLIP ViT-L/14 (frozen)</li>\n<li>Projector: Linear layer</li>\n<li>LLM: Vicuna-13B</li>\n<li>Training: Two-stage (projector pretraining, then full fine-tuning)</li>\n</ul>\n<p><strong>LLaVA 1.5 (2023)</strong></p>\n<ul>\n<li>Projector: Two-layer MLP (significant improvement)</li>\n<li>Higher resolution support: 336x336</li>\n<li>Better instruction-following data</li>\n</ul>\n<p><strong>LLaVA-NeXT / LLaVA 1.6 (2024)</strong></p>\n<ul>\n<li>Dynamic high resolution: Up to 672x672</li>\n<li>AnyRes: Handles varying aspect ratios by tiling</li>\n<li>Multiple LLM backends (Vicuna, Mistral, Hermes)</li>\n</ul>\n<p><strong>Training process:</strong></p>\n<p>Stage 1: Projector pretraining</p>\n<ul>\n<li>Freeze vision encoder and LLM</li>\n<li>Train only the MLP projector</li>\n<li>Use image-caption pairs</li>\n<li>Goal: Align vision features with language embedding space</li>\n</ul>\n<p>Stage 2: Full fine-tuning</p>\n<ul>\n<li>Freeze vision encoder</li>\n<li>Unfreeze projector and LLM</li>\n<li>Train on visual instruction data</li>\n<li>Goal: Learn to follow multimodal instructions</li>\n</ul>\n<h3 id=\"qwen-vl-qwen2-vl-qwen25-vl\">Qwen-VL / Qwen2-VL / Qwen2.5-VL</h3>\n<p>Alibaba’s Qwen vision models introduced Naive Dynamic Resolution, processing images at their native resolution rather than forcing fixed sizes.</p>\n<p><strong>Key innovations:</strong></p>\n<p><strong>Dynamic resolution</strong>: Images are converted to a variable number of visual tokens based on their actual dimensions. A small icon might use 256 tokens; a high-resolution document might use 4096+.</p>\n<p><strong>M-RoPE (Multimodal RoPE)</strong>: Unified position encoding for text, images, and video. Enables the model to understand:</p>\n<ul>\n<li>Absolute text positions</li>\n<li>2D spatial positions within images</li>\n<li>Temporal positions across video frames</li>\n</ul>\n<p><strong>2D RoPE in ViT</strong>: The vision encoder itself uses 2D rotary position embeddings, enabling better adaptation to varying resolutions during inference.</p>\n<p><strong>Model sizes</strong>: 2B, 8B, 72B parameters.</p>\n<p><strong>Qwen2.5-VL improvements (2025)</strong>:</p>\n<ul>\n<li>ViT trained from scratch with native dynamic resolution</li>\n<li>Window attention for reduced compute</li>\n<li>Dynamic FPS sampling for video</li>\n<li>Hours-long video understanding with second-level localization</li>\n<li>Strong document/chart/table extraction</li>\n</ul>\n<h3 id=\"internvl\">InternVL</h3>\n<p>OpenGVLab’s InternVL scales both vision and language components, with InternViT reaching 6B parameters.</p>\n<p><strong>Architecture</strong>: ViT-MLP-LLM paradigm with a pixel unshuffle operation that reduces visual tokens to 1/4 of the original count.</p>\n<p><strong>InternVL 2.5 / 3.0 (2025)</strong>:</p>\n<ul>\n<li>Native Multimodal Pre-Training: Interleaves vision-language data with text corpora in a single pretraining stage (rather than adapting a text-only model)</li>\n<li>Variable Visual Position Encoding (V2PE)</li>\n<li>Dynamic High Resolution from InternVL 1.5</li>\n<li>Mixed Preference Optimization for alignment</li>\n</ul>\n<p><strong>InternVL 3.5 (August 2025)</strong>:</p>\n<ul>\n<li>Cascade Reinforcement Learning for improved reasoning</li>\n<li>Visual Resolution Router (ViR): Dynamically adjusts visual token resolution</li>\n<li>Decoupled Vision-Language Deployment: Separate vision and language servers for async pipelining</li>\n<li>4.05x inference speedup over InternVL3</li>\n<li>State-of-the-art among open-source models</li>\n</ul>\n<h3 id=\"paligemma\">PaliGemma</h3>\n<p>Google’s PaliGemma combines SigLIP with Gemma in a straightforward architecture.</p>\n<p><strong>Components</strong>:</p>\n<ul>\n<li>Vision: SigLIP-So400m (400M params)</li>\n<li>Language: Gemma-2B</li>\n<li>Projection: Linear layer (1152 to 2048 dimensions)</li>\n</ul>\n<p><strong>Resolution and tokens</strong>:</p>\n<ul>\n<li>Pretrained at 224x224, 448x448, or 896x896</li>\n<li>Patch size: 14</li>\n<li>Token counts: 256 (224px), 1024 (448px), 4096 (896px)</li>\n</ul>\n<p><strong>PaliGemma 2 (2024)</strong>:</p>\n<ul>\n<li>Upgraded to Gemma 2</li>\n<li>Sizes: 3B, 10B, 28B</li>\n<li>Multiple resolution support per model</li>\n</ul>\n<p><strong>Gemma 3 with vision (2025)</strong>:</p>\n<ul>\n<li>Custom SigLIP encoder</li>\n<li>Pan and Scan algorithm for varying aspect ratios</li>\n<li>896x896 fixed encoder input</li>\n</ul>\n<h3 id=\"phi-3-vision\">Phi-3 Vision</h3>\n<p>Microsoft’s Phi-3 Vision packs vision capabilities into a 4.2B parameter model.</p>\n<p><strong>Architecture</strong>:</p>\n<ul>\n<li>Image encoder</li>\n<li>Connector</li>\n<li>Projector</li>\n<li>Phi-3 Mini language model</li>\n</ul>\n<p><strong>Specifications</strong>:</p>\n<ul>\n<li>4.2B total parameters</li>\n<li>128K context length</li>\n<li>Inputs: Text and images</li>\n</ul>\n<p><strong>Strengths</strong>:</p>\n<ul>\n<li>Chart, graph, and table understanding</li>\n<li>Document analysis</li>\n<li>Small footprint suitable for edge deployment</li>\n</ul>\n<p><strong>Phi-3.5 Vision</strong>:</p>\n<ul>\n<li>Multi-frame capabilities (image comparison, video summarization)</li>\n<li>Improved single-image benchmarks</li>\n<li>Training: 6 days on 256 A100-80G GPUs with 500B tokens</li>\n</ul>\n<h3 id=\"cogvlm\">CogVLM</h3>\n<p>CogVLM introduces a Visual Expert module that adds trainable parameters to the language model specifically for processing vision features.</p>\n<p><strong>Architecture difference</strong>: Instead of just projecting vision features into the LLM’s input space, CogVLM adds parallel QKV matrices and MLP layers for visual tokens at each transformer layer.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>Standard layer:</span></span>\n<span class=\"line\"><span>  text_tokens ──▶ [QKV + MLP] ──▶ output</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>CogVLM layer:</span></span>\n<span class=\"line\"><span>  text_tokens  ──▶ [QKV_text + MLP_text]  ──┐</span></span>\n<span class=\"line\"><span>                                            ├──▶ output</span></span>\n<span class=\"line\"><span>  image_tokens ──▶ [QKV_vision + MLP_vision]──┘</span></span></code></pre>\n<p><strong>Benefits</strong>:</p>\n<ul>\n<li>Deep fusion of vision and language (not just early concatenation)</li>\n<li>Preserves language model performance on text-only tasks</li>\n<li>Visual expert parameters initialized from pretrained weights</li>\n</ul>\n<p><strong>Scale</strong>: CogVLM-17B has 10B vision parameters + 7B language parameters.</p>\n<p><strong>CogVLM2 (2024)</strong>:</p>\n<ul>\n<li>Improved training recipes</li>\n<li>Up to 1344x1344 input resolution</li>\n</ul>\n<h3 id=\"florence-2\">Florence-2</h3>\n<p>Microsoft’s Florence-2 takes a different approach: a unified sequence-to-sequence model for all vision tasks.</p>\n<p><strong>Architecture</strong>:</p>\n<ul>\n<li>Vision encoder: DaViT (Dual Attention Vision Transformer)</li>\n<li>Text processing: BART-style encoder-decoder</li>\n<li>Sizes: 0.2B and 0.7B parameters</li>\n</ul>\n<p><strong>Key insight</strong>: All vision tasks (detection, segmentation, captioning, grounding) can be formulated as sequence-to-sequence problems. The model outputs text, including coordinate tokens for localization tasks.</p>\n<p><strong>Training data</strong>: FLD-5B dataset with 5.4B annotations across 126M images (boxes, masks, captions, grounding).</p>\n<p><strong>Performance</strong>: Despite small size, Florence-2 outperforms much larger models on zero-shot captioning. On COCO, the 232M model (score: 133) and 771M model (score: 135.6) both beat DeepMind’s 80B Flamingo.</p>\n<hr />\n<h2 id=\"why-vlm-inference-is-more-complex\">Why VLM Inference is More Complex</h2>\n<p>VLM inference presents challenges beyond text-only LLMs.</p>\n<h3 id=\"variable-image-resolutions\">Variable Image Resolutions</h3>\n<p>Text inputs have predictable token counts (roughly 4 chars per token). Images vary wildly:</p>\n<ul>\n<li>Thumbnail: 224x224 = 256 tokens (with 14x14 patches)</li>\n<li>Document: 896x896 = 4096 tokens</li>\n<li>High-res photo with dynamic tiling: 10000+ tokens</li>\n</ul>\n<p>Dynamic resolution models like Qwen2-VL convert resolution directly to token count. Memory and compute scale accordingly.</p>\n<h3 id=\"token-count-variability\">Token Count Variability</h3>\n<p>A batch of requests might contain:</p>\n<ul>\n<li>Request 1: 100 text tokens, no images</li>\n<li>Request 2: 50 text tokens, 256 image tokens</li>\n<li>Request 3: 200 text tokens, 2048 image tokens (high-res document)</li>\n</ul>\n<p>This variability complicates batching and memory allocation.</p>\n<h3 id=\"image-preprocessing-overhead\">Image Preprocessing Overhead</h3>\n<p>Before the vision encoder runs, images require preprocessing:</p>\n<ol>\n<li><strong>Decode</strong>: Decompress JPEG/PNG to raw pixels</li>\n<li><strong>Resize</strong>: Scale to encoder’s expected resolution</li>\n<li><strong>Normalize</strong>: Convert to float, apply mean/std normalization (typically ImageNet stats: mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])</li>\n<li><strong>Patch/tile</strong>: For dynamic resolution, split into tiles</li>\n<li><strong>Tensor conversion</strong>: Move to GPU memory</li>\n</ol>\n<p>This preprocessing happens on CPU and can bottleneck throughput when image counts are high.</p>\n<h3 id=\"memory-requirements\">Memory Requirements</h3>\n<p>High-resolution images consume memory at multiple stages:</p>\n<ul>\n<li>Raw pixels: 896x896x3 = 2.4MB per image</li>\n<li>Vision encoder activations: Intermediate states during forward pass</li>\n<li>Projected tokens: 4096 tokens x 4096 dim x 2 bytes (fp16) = 33MB per high-res image</li>\n<li>KV cache: Each image token needs KV cache entries for all subsequent generation</li>\n</ul>\n<p>For a 72B model processing a high-res document, the image tokens alone can consume several GB of KV cache memory.</p>\n<h3 id=\"batching-challenges\">Batching Challenges</h3>\n<p>Text-only continuous batching works because all tokens are homogeneous. With images:</p>\n<ul>\n<li>Vision encoder runs separately from LLM</li>\n<li>Different images in a batch may have different resolutions</li>\n<li>Prefill (processing prompt + images) is much heavier than decode (generating tokens)</li>\n</ul>\n<p><strong>vLLM’s solution</strong>: Hybrid parallelism where the vision encoder uses data parallelism (each GPU processes different images) while the LLM uses tensor parallelism. This avoids synchronization overhead during vision encoding.</p>\n<p><strong>Prefix caching</strong>: vLLM V1 supports prefix caching for multimodal inputs, so repeated images don’t require re-encoding.</p>\n<hr />\n<h2 id=\"vlm-quantization\">VLM Quantization</h2>\n<p>Quantization reduces model size and speeds inference by using lower-precision weights and activations.</p>\n<h3 id=\"quantizing-the-vision-encoder\">Quantizing the Vision Encoder</h3>\n<p>Vision encoders are often left at higher precision because:</p>\n<ul>\n<li>They’re smaller than the LLM (typically &lt; 1B params)</li>\n<li>Visual features are sensitive to quantization artifacts</li>\n<li>The encoder runs once per image (not autoregressive), so speed gains are less impactful</li>\n</ul>\n<p>However, for edge deployment, encoder quantization matters. Research shows:</p>\n<ul>\n<li>INT8 quantization typically maintains quality</li>\n<li>INT4 can work with careful calibration</li>\n<li>The vision modality is generally less sensitive than language</li>\n</ul>\n<h3 id=\"quantizing-the-language-model\">Quantizing the Language Model</h3>\n<p>Standard LLM quantization techniques apply:</p>\n<ul>\n<li><strong>GPTQ</strong>: Post-training quantization using calibration data</li>\n<li><strong>AWQ</strong>: Activation-aware weight quantization</li>\n<li><strong>GGUF</strong>: llama.cpp’s format with various quantization levels (Q4_K_M, Q5_K_S, etc.)</li>\n<li><strong>bitsandbytes</strong>: 4-bit and 8-bit for training/inference</li>\n</ul>\n<p>The language model dominates parameter count (e.g., 72B LLM vs 400M vision encoder), so LLM quantization provides most of the size/speed benefits.</p>\n<h3 id=\"joint-quantization-strategies\">Joint Quantization Strategies</h3>\n<p><strong>Q-VLM</strong> (NeurIPS 2024) proposes cross-layer dependency mining for VLM quantization. Rather than quantizing layer-by-layer, it considers how quantization errors propagate through the full vision-language pipeline.</p>\n<p>Results: 2.78x memory compression, 1.44x speedup on 13B LLaVA without performance degradation.</p>\n<p><strong>MBQ (Modality-Balanced Quantization)</strong> (CVPR 2025) observes that language modules are more sensitive to quantization than vision modules. Their approach:</p>\n<ul>\n<li>4-bit quantization for vision (ViT modules)</li>\n<li>8-bit quantization for language</li>\n<li>Up to 4% improvement under W3A16 and 11% under W4A8 vs uniform quantization</li>\n</ul>\n<h3 id=\"quality-tradeoffs\">Quality Tradeoffs</h3>\n<p>Quantization affects different VLM capabilities differently:</p>\n<ul>\n<li>OCR accuracy can degrade with aggressive quantization</li>\n<li>Fine-grained visual details may be lost</li>\n<li>Reasoning over complex diagrams is sensitive</li>\n<li>General image description is more robust</li>\n</ul>\n<p>Recommendation: Benchmark your specific use case. A model that’s fine for photo captioning might fail on document analysis when heavily quantized.</p>\n<hr />\n<h2 id=\"practical-vlm-inference\">Practical VLM Inference</h2>\n<h3 id=\"hugging-face-transformers\">Hugging Face Transformers</h3>\n<p>The most straightforward approach for experimentation.</p>\n<p><strong>LLaVA example:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoProcessor</span><span>,</span><span> LlavaForConditionalGeneration</span></span>\n<span class=\"line\"><span>from</span><span> PIL </span><span>import</span><span> Image</span></span>\n<span class=\"line\"><span>import</span><span> requests</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model_id </span><span>=</span><span> \"llava-hf/llava-1.5-7b-hf\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Load model and processor</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> LlavaForConditionalGeneration</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    model_id,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>torch.float16,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>processor </span><span>=</span><span> AutoProcessor</span><span>.</span><span>from_pretrained</span><span>(model_id)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Prepare inputs</span></span>\n<span class=\"line\"><span>url </span><span>=</span><span> \"https://upload.wikimedia.org/wikipedia/commons/thumb/3/3a/Cat03.jpg/1200px-Cat03.jpg\"</span></span>\n<span class=\"line\"><span>image </span><span>=</span><span> Image</span><span>.</span><span>open</span><span>(requests.</span><span>get</span><span>(url, stream</span><span>=</span><span>True</span><span>).raw)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>conversation </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>    {</span></span>\n<span class=\"line\"><span>        \"role\"</span><span>:</span><span> \"user\"</span><span>,</span></span>\n<span class=\"line\"><span>        \"content\"</span><span>:</span><span> [</span></span>\n<span class=\"line\"><span>            {</span><span>\"type\"</span><span>:</span><span> \"image\"</span><span>},</span></span>\n<span class=\"line\"><span>            {</span><span>\"type\"</span><span>:</span><span> \"text\"</span><span>,</span><span> \"text\"</span><span>:</span><span> \"What animal is in this image? Describe it briefly.\"</span><span>}</span></span>\n<span class=\"line\"><span>        ]</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> processor</span><span>.</span><span>apply_chat_template</span><span>(conversation, add_generation_prompt</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> processor</span><span>(images</span><span>=</span><span>image, text</span><span>=</span><span>prompt, return_tensors</span><span>=</span><span>\"pt\"</span><span>).</span><span>to</span><span>(model.device)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Generate</span></span>\n<span class=\"line\"><span>output </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(</span><span>**</span><span>inputs, max_new_tokens</span><span>=</span><span>100</span><span>, do_sample</span><span>=</span><span>False</span><span>)</span></span>\n<span class=\"line\"><span>response </span><span>=</span><span> processor</span><span>.</span><span>decode</span><span>(output[</span><span>0</span><span>], skip_special_tokens</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(response)</span></span></code></pre>\n<p><strong>Qwen2-VL example:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> Qwen2VLForConditionalGeneration</span><span>,</span><span> AutoProcessor</span></span>\n<span class=\"line\"><span>from</span><span> PIL </span><span>import</span><span> Image</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> Qwen2VLForConditionalGeneration</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"Qwen/Qwen2-VL-7B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>processor </span><span>=</span><span> AutoProcessor</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"Qwen/Qwen2-VL-7B-Instruct\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Qwen2-VL supports dynamic resolution</span></span>\n<span class=\"line\"><span>image </span><span>=</span><span> Image</span><span>.</span><span>open</span><span>(</span><span>\"document.png\"</span><span>)</span><span>  # Any resolution</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>messages </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>    {</span></span>\n<span class=\"line\"><span>        \"role\"</span><span>:</span><span> \"user\"</span><span>,</span></span>\n<span class=\"line\"><span>        \"content\"</span><span>:</span><span> [</span></span>\n<span class=\"line\"><span>            {</span><span>\"type\"</span><span>:</span><span> \"image\"</span><span>,</span><span> \"image\"</span><span>:</span><span> image</span><span>},</span></span>\n<span class=\"line\"><span>            {</span><span>\"type\"</span><span>:</span><span> \"text\"</span><span>,</span><span> \"text\"</span><span>:</span><span> \"Extract all text from this document.\"</span><span>}</span></span>\n<span class=\"line\"><span>        ]</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>text </span><span>=</span><span> processor</span><span>.</span><span>apply_chat_template</span><span>(messages, tokenize</span><span>=</span><span>False</span><span>, add_generation_prompt</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> processor</span><span>(text</span><span>=</span><span>[text], images</span><span>=</span><span>[image], return_tensors</span><span>=</span><span>\"pt\"</span><span>).</span><span>to</span><span>(model.device)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>generated_ids </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(</span><span>**</span><span>inputs, max_new_tokens</span><span>=</span><span>512</span><span>)</span></span>\n<span class=\"line\"><span>output </span><span>=</span><span> processor</span><span>.</span><span>batch_decode</span><span>(generated_ids, skip_special_tokens</span><span>=</span><span>True</span><span>)</span></span></code></pre>\n<p><strong>Pipeline API (simpler):</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> pipeline</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>pipe </span><span>=</span><span> pipeline</span><span>(</span><span>\"image-text-to-text\"</span><span>, model</span><span>=</span><span>\"llava-hf/llava-1.5-7b-hf\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>output </span><span>=</span><span> pipe</span><span>(</span></span>\n<span class=\"line\"><span>    images</span><span>=</span><span>\"https://example.com/image.jpg\"</span><span>,</span></span>\n<span class=\"line\"><span>    text</span><span>=</span><span>\"Describe this image in detail.\"</span><span>,</span></span>\n<span class=\"line\"><span>    max_new_tokens</span><span>=</span><span>200</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<h3 id=\"llamacpp-multimodal\">llama.cpp Multimodal</h3>\n<p>llama.cpp supports vision models via the <code>mmproj</code> (multimodal projector) architecture.</p>\n<p><strong>Installation:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># macOS</span></span>\n<span class=\"line\"><span>brew</span><span> install</span><span> llama.cpp</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or build from source</span></span>\n<span class=\"line\"><span>git</span><span> clone</span><span> https://github.com/ggml-org/llama.cpp</span></span>\n<span class=\"line\"><span>cd</span><span> llama.cpp</span><span> &amp;&amp;</span><span> make</span></span></code></pre>\n<p><strong>Download model and projector:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Pre-quantized models available from ggml-org</span></span>\n<span class=\"line\"><span># https://huggingface.co/collections/ggml-org/multimodal-ggufs</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Example: Qwen2.5-VL-7B</span></span>\n<span class=\"line\"><span>wget</span><span> https://huggingface.co/ggml-org/Qwen2.5-VL-7B-Instruct-GGUF/resolve/main/qwen2.5-vl-7b-instruct-q4_k_m.gguf</span></span>\n<span class=\"line\"><span>wget</span><span> https://huggingface.co/ggml-org/Qwen2.5-VL-7B-Instruct-GGUF/resolve/main/qwen2.5-vl-7b-instruct-vision.gguf</span></span></code></pre>\n<p><strong>Run inference:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>./llama-mtmd-cli</span><span> \\</span></span>\n<span class=\"line\"><span>    -m</span><span> qwen2.5-vl-7b-instruct-q4_k_m.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    --mmproj</span><span> qwen2.5-vl-7b-instruct-vision.gguf</span><span> \\</span></span>\n<span class=\"line\"><span>    -p</span><span> \"Describe this image:\"</span><span> \\</span></span>\n<span class=\"line\"><span>    --image</span><span> photo.jpg</span></span></code></pre>\n<p><strong>Supported models:</strong></p>\n<ul>\n<li>Gemma 3 (4B, 12B, 27B)</li>\n<li>Qwen 2 VL (2B, 7B)</li>\n<li>Qwen 2.5 VL (3B, 7B, 32B, 72B)</li>\n<li>SmolVLM variants</li>\n<li>Pixtral 12B</li>\n<li>InternVL 2.5</li>\n<li>Mistral Small 3.1 24B</li>\n</ul>\n<p><strong>Limitations</strong>: Currently supports static images only. Video processing is not yet implemented.</p>\n<h3 id=\"vllm\">vLLM</h3>\n<p>vLLM provides high-throughput VLM serving with continuous batching.</p>\n<p><strong>Installation:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>pip</span><span> install</span><span> vllm</span></span></code></pre>\n<p><strong>Offline inference:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span><span>,</span><span> SamplingParams</span></span>\n<span class=\"line\"><span>from</span><span> vllm</span><span>.</span><span>multimodal </span><span>import</span><span> MultiModalData</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(model</span><span>=</span><span>\"llava-hf/llava-1.5-7b-hf\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> \"USER: &lt;image&gt;\\nWhat is in this image?\\nASSISTANT:\"</span></span>\n<span class=\"line\"><span>sampling_params </span><span>=</span><span> SamplingParams</span><span>(temperature</span><span>=</span><span>0.0</span><span>, max_tokens</span><span>=</span><span>256</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Single image</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> llm</span><span>.</span><span>generate</span><span>(</span></span>\n<span class=\"line\"><span>    {</span></span>\n<span class=\"line\"><span>        \"prompt\"</span><span>: prompt,</span></span>\n<span class=\"line\"><span>        \"multi_modal_data\"</span><span>: {</span><span>\"image\"</span><span>: </span><span>\"path/to/image.jpg\"</span><span>}</span></span>\n<span class=\"line\"><span>    },</span></span>\n<span class=\"line\"><span>    sampling_params</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(outputs[</span><span>0</span><span>].outputs[</span><span>0</span><span>].text)</span></span></code></pre>\n<p><strong>Server deployment:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Start server</span></span>\n<span class=\"line\"><span>python</span><span> -m</span><span> vllm.entrypoints.openai.api_server</span><span> \\</span></span>\n<span class=\"line\"><span>    --model</span><span> llava-hf/llava-1.5-7b-hf</span><span> \\</span></span>\n<span class=\"line\"><span>    --chat-template</span><span> llava</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Query via OpenAI-compatible API</span></span>\n<span class=\"line\"><span>curl</span><span> http://localhost:8000/v1/chat/completions</span><span> \\</span></span>\n<span class=\"line\"><span>    -H</span><span> \"Content-Type: application/json\"</span><span> \\</span></span>\n<span class=\"line\"><span>    -d</span><span> '{</span></span>\n<span class=\"line\"><span>        \"model\": \"llava-hf/llava-1.5-7b-hf\",</span></span>\n<span class=\"line\"><span>        \"messages\": [</span></span>\n<span class=\"line\"><span>            {</span></span>\n<span class=\"line\"><span>                \"role\": \"user\",</span></span>\n<span class=\"line\"><span>                \"content\": [</span></span>\n<span class=\"line\"><span>                    {\"type\": \"text\", \"text\": \"What is in this image?\"},</span></span>\n<span class=\"line\"><span>                    {\"type\": \"image_url\", \"image_url\": {\"url\": \"https://example.com/image.jpg\"}}</span></span>\n<span class=\"line\"><span>                ]</span></span>\n<span class=\"line\"><span>            }</span></span>\n<span class=\"line\"><span>        ]</span></span>\n<span class=\"line\"><span>    }'</span></span></code></pre>\n<p><strong>Performance tuning:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Use data parallelism for vision encoder</span></span>\n<span class=\"line\"><span>python</span><span> -m</span><span> vllm.entrypoints.openai.api_server</span><span> \\</span></span>\n<span class=\"line\"><span>    --model</span><span> Qwen/Qwen2-VL-7B-Instruct</span><span> \\</span></span>\n<span class=\"line\"><span>    --mm-encoder-tp-mode</span><span> data</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Limit image tokens to control memory</span></span>\n<span class=\"line\"><span>--limit-mm-per-prompt</span><span> '{\"image\": 2048}'</span></span></code></pre>\n<p><strong>Supported models (partial list)</strong>:</p>\n<ul>\n<li>LLaVA 1.5, LLaVA-NeXT</li>\n<li>Qwen-VL, Qwen2-VL</li>\n<li>InternVL, InternVL2</li>\n<li>PaliGemma</li>\n<li>Phi-3 Vision</li>\n<li>BLIP-2</li>\n<li>Chameleon</li>\n</ul>\n<h3 id=\"ollama\">Ollama</h3>\n<p>Ollama provides the simplest local VLM experience.</p>\n<p><strong>Installation:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># macOS/Linux</span></span>\n<span class=\"line\"><span>curl</span><span> -fsSL</span><span> https://ollama.com/install.sh</span><span> |</span><span> sh</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or download from ollama.com</span></span></code></pre>\n<p><strong>Run vision models:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Pull a vision model</span></span>\n<span class=\"line\"><span>ollama</span><span> pull</span><span> llava</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Interactive mode</span></span>\n<span class=\"line\"><span>ollama</span><span> run</span><span> llava</span></span>\n<span class=\"line\"><span>&gt;&gt;&gt; [paste image path or URL]</span></span>\n<span class=\"line\"><span>&gt;&gt;&gt; </span><span>What</span><span>'s in this image?</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Or with llava-phi3 (smaller, faster)</span></span>\n<span class=\"line\"><span>ollama pull llava-phi3</span></span>\n<span class=\"line\"><span>ollama run llava-phi3</span></span></code></pre>\n<p><strong>API usage:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> ollama</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>response </span><span>=</span><span> ollama</span><span>.</span><span>chat</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>'llava'</span><span>,</span></span>\n<span class=\"line\"><span>    messages</span><span>=</span><span>[{</span></span>\n<span class=\"line\"><span>        'role'</span><span>: </span><span>'user'</span><span>,</span></span>\n<span class=\"line\"><span>        'content'</span><span>: </span><span>'Describe this image'</span><span>,</span></span>\n<span class=\"line\"><span>        'images'</span><span>: [</span><span>'./photo.jpg'</span><span>]</span></span>\n<span class=\"line\"><span>    }]</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(response[</span><span>'message'</span><span>][</span><span>'content'</span><span>])</span></span></code></pre>\n<p><strong>Available vision models on Ollama:</strong></p>\n<ul>\n<li><code>llava</code> (7B and 13B variants)</li>\n<li><code>llava-phi3</code> (3.8B, faster)</li>\n<li><code>bakllava</code></li>\n<li><code>moondream</code> (smaller, 1.8B)</li>\n</ul>\n<hr />\n<h2 id=\"code-examples-for-common-tasks\">Code Examples for Common Tasks</h2>\n<h3 id=\"document-ocr\">Document OCR</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> Qwen2VLForConditionalGeneration</span><span>,</span><span> AutoProcessor</span></span>\n<span class=\"line\"><span>from</span><span> PIL </span><span>import</span><span> Image</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> Qwen2VLForConditionalGeneration</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"Qwen/Qwen2-VL-7B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>processor </span><span>=</span><span> AutoProcessor</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"Qwen/Qwen2-VL-7B-Instruct\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Use high resolution for documents</span></span>\n<span class=\"line\"><span>image </span><span>=</span><span> Image</span><span>.</span><span>open</span><span>(</span><span>\"invoice.png\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>messages </span><span>=</span><span> [</span><span>{</span></span>\n<span class=\"line\"><span>    \"role\"</span><span>:</span><span> \"user\"</span><span>,</span></span>\n<span class=\"line\"><span>    \"content\"</span><span>:</span><span> [</span></span>\n<span class=\"line\"><span>        {</span><span>\"type\"</span><span>:</span><span> \"image\"</span><span>,</span><span> \"image\"</span><span>:</span><span> image</span><span>},</span></span>\n<span class=\"line\"><span>        {</span><span>\"type\"</span><span>:</span><span> \"text\"</span><span>,</span><span> \"text\"</span><span>:</span><span> \"Extract all text from this invoice. Format as structured data.\"</span><span>}</span></span>\n<span class=\"line\"><span>    ]</span></span>\n<span class=\"line\"><span>}</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> processor</span><span>(</span></span>\n<span class=\"line\"><span>    text</span><span>=</span><span>processor.</span><span>apply_chat_template</span><span>(messages, tokenize</span><span>=</span><span>False</span><span>, add_generation_prompt</span><span>=</span><span>True</span><span>),</span></span>\n<span class=\"line\"><span>    images</span><span>=</span><span>[image],</span></span>\n<span class=\"line\"><span>    return_tensors</span><span>=</span><span>\"pt\"</span></span>\n<span class=\"line\"><span>).</span><span>to</span><span>(model.device)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>output </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(</span><span>**</span><span>inputs, max_new_tokens</span><span>=</span><span>1024</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(processor.</span><span>decode</span><span>(output[</span><span>0</span><span>], skip_special_tokens</span><span>=</span><span>True</span><span>))</span></span></code></pre>\n<h3 id=\"multi-image-comparison\">Multi-Image Comparison</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> LlavaNextProcessor</span><span>,</span><span> LlavaNextForConditionalGeneration</span></span>\n<span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>from</span><span> PIL </span><span>import</span><span> Image</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> LlavaNextForConditionalGeneration</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"llava-hf/llava-v1.6-mistral-7b-hf\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>torch.float16,</span></span>\n<span class=\"line\"><span>    device_map</span><span>=</span><span>\"auto\"</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>processor </span><span>=</span><span> LlavaNextProcessor</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"llava-hf/llava-v1.6-mistral-7b-hf\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>image1 </span><span>=</span><span> Image</span><span>.</span><span>open</span><span>(</span><span>\"product_v1.jpg\"</span><span>)</span></span>\n<span class=\"line\"><span>image2 </span><span>=</span><span> Image</span><span>.</span><span>open</span><span>(</span><span>\"product_v2.jpg\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> \"[INST] &lt;image&gt;\\n&lt;image&gt;\\nCompare these two product images. What are the differences? [/INST]\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> processor</span><span>(prompt, [image1, image2], return_tensors</span><span>=</span><span>\"pt\"</span><span>).</span><span>to</span><span>(model.device)</span></span>\n<span class=\"line\"><span>output </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(</span><span>**</span><span>inputs, max_new_tokens</span><span>=</span><span>300</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(processor.</span><span>decode</span><span>(output[</span><span>0</span><span>], skip_special_tokens</span><span>=</span><span>True</span><span>))</span></span></code></pre>\n<h3 id=\"chart-analysis\">Chart Analysis</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Using Florence-2 for detailed chart understanding</span></span>\n<span class=\"line\"><span>from</span><span> transformers </span><span>import</span><span> AutoProcessor</span><span>,</span><span> AutoModelForCausalLM</span></span>\n<span class=\"line\"><span>from</span><span> PIL </span><span>import</span><span> Image</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> AutoModelForCausalLM</span><span>.</span><span>from_pretrained</span><span>(</span></span>\n<span class=\"line\"><span>    \"microsoft/Florence-2-large\"</span><span>,</span></span>\n<span class=\"line\"><span>    torch_dtype</span><span>=</span><span>\"auto\"</span><span>,</span></span>\n<span class=\"line\"><span>    trust_remote_code</span><span>=</span><span>True</span></span>\n<span class=\"line\"><span>).</span><span>to</span><span>(</span><span>\"cuda\"</span><span>)</span></span>\n<span class=\"line\"><span>processor </span><span>=</span><span> AutoProcessor</span><span>.</span><span>from_pretrained</span><span>(</span><span>\"microsoft/Florence-2-large\"</span><span>, trust_remote_code</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>image </span><span>=</span><span> Image</span><span>.</span><span>open</span><span>(</span><span>\"sales_chart.png\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Florence-2 uses task-specific prompts</span></span>\n<span class=\"line\"><span>prompt </span><span>=</span><span> \"&lt;MORE_DETAILED_CAPTION&gt;\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>inputs </span><span>=</span><span> processor</span><span>(text</span><span>=</span><span>prompt, images</span><span>=</span><span>image, return_tensors</span><span>=</span><span>\"pt\"</span><span>).</span><span>to</span><span>(</span><span>\"cuda\"</span><span>)</span></span>\n<span class=\"line\"><span>generated_ids </span><span>=</span><span> model</span><span>.</span><span>generate</span><span>(</span></span>\n<span class=\"line\"><span>    input_ids</span><span>=</span><span>inputs[</span><span>\"input_ids\"</span><span>],</span></span>\n<span class=\"line\"><span>    pixel_values</span><span>=</span><span>inputs[</span><span>\"pixel_values\"</span><span>],</span></span>\n<span class=\"line\"><span>    max_new_tokens</span><span>=</span><span>512</span><span>,</span></span>\n<span class=\"line\"><span>    num_beams</span><span>=</span><span>3</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>result </span><span>=</span><span> processor</span><span>.</span><span>batch_decode</span><span>(generated_ids, skip_special_tokens</span><span>=</span><span>True</span><span>)</span><span>[</span><span>0</span><span>]</span></span>\n<span class=\"line\"><span>print</span><span>(result)</span></span></code></pre>\n<h3 id=\"batch-processing-with-vllm\">Batch Processing with vLLM</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> vllm </span><span>import</span><span> LLM</span><span>,</span><span> SamplingParams</span></span>\n<span class=\"line\"><span>import</span><span> os</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>llm </span><span>=</span><span> LLM</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"Qwen/Qwen2-VL-7B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    limit_mm_per_prompt</span><span>=</span><span>{</span><span>\"image\"</span><span>: </span><span>1</span><span>}</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>sampling_params </span><span>=</span><span> SamplingParams</span><span>(temperature</span><span>=</span><span>0.0</span><span>, max_tokens</span><span>=</span><span>256</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Prepare batch of requests</span></span>\n<span class=\"line\"><span>image_dir </span><span>=</span><span> \"./images\"</span></span>\n<span class=\"line\"><span>requests </span><span>=</span><span> []</span></span>\n<span class=\"line\"><span>for</span><span> img_file </span><span>in</span><span> os</span><span>.</span><span>listdir</span><span>(image_dir):</span></span>\n<span class=\"line\"><span>    if</span><span> img_file</span><span>.</span><span>endswith</span><span>((</span><span>'.jpg'</span><span>, </span><span>'.png'</span><span>)):</span></span>\n<span class=\"line\"><span>        requests</span><span>.</span><span>append</span><span>({</span></span>\n<span class=\"line\"><span>            \"prompt\"</span><span>: </span><span>\"USER: &lt;image&gt;\\nDescribe this image briefly.\\nASSISTANT:\"</span><span>,</span></span>\n<span class=\"line\"><span>            \"multi_modal_data\"</span><span>: {</span><span>\"image\"</span><span>: os.path.</span><span>join</span><span>(image_dir, img_file)}</span></span>\n<span class=\"line\"><span>        })</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Process batch</span></span>\n<span class=\"line\"><span>outputs </span><span>=</span><span> llm</span><span>.</span><span>generate</span><span>(requests, sampling_params)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> i</span><span>,</span><span> output </span><span>in</span><span> enumerate</span><span>(outputs):</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Image </span><span>{</span><span>i</span><span>}</span><span>: </span><span>{</span><span>output.outputs[</span><span>0</span><span>].text</span><span>}</span><span>\"</span><span>)</span></span></code></pre>\n<hr />\n<h2 id=\"references\">References</h2>\n<h3 id=\"papers\">Papers</h3>\n<ul>\n<li>Dosovitskiy et al. “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale” (ViT, 2020)</li>\n<li>Radford et al. “Learning Transferable Visual Models From Natural Language Supervision” (CLIP, 2021)</li>\n<li>Zhai et al. “Sigmoid Loss for Language Image Pre-Training” (SigLIP, ICCV 2023) - <a href=\"https://arxiv.org/abs/2303.15343\">https://arxiv.org/abs/2303.15343</a></li>\n<li>Liu et al. “Visual Instruction Tuning” (LLaVA, NeurIPS 2023)</li>\n<li>Wang et al. “Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution” (2024) - <a href=\"https://arxiv.org/abs/2409.12191\">https://arxiv.org/abs/2409.12191</a></li>\n<li>Qwen Team. “Qwen2.5-VL Technical Report” (2025) - <a href=\"https://arxiv.org/abs/2502.13923\">https://arxiv.org/abs/2502.13923</a></li>\n<li>Chen et al. “InternVL: Scaling up Vision Foundation Models” (CVPR 2024)</li>\n<li>OpenGVLab. “InternVL3” (2025) - <a href=\"https://internvl.github.io/blog/2025-04-11-InternVL-3.0/\">https://internvl.github.io/blog/2025-04-11-InternVL-3.0/</a></li>\n<li>Wang et al. “CogVLM: Visual Expert for Pretrained Language Models” (NeurIPS 2024) - <a href=\"https://arxiv.org/abs/2311.03079\">https://arxiv.org/abs/2311.03079</a></li>\n<li>Xiao et al. “Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks” (2024)</li>\n<li>Shao et al. “Q-VLM: Post-training Quantization for Large Vision-Language Models” (NeurIPS 2024) - <a href=\"https://arxiv.org/abs/2410.08119\">https://arxiv.org/abs/2410.08119</a></li>\n<li>Li et al. “MBQ: Modality-Balanced Quantization for Large Vision-Language Models” (CVPR 2025)</li>\n<li>Vasu et al. “FastVLM: Efficient Vision Encoding for Vision Language Models” (CVPR 2025)</li>\n</ul>\n<h3 id=\"model-resources\">Model Resources</h3>\n<ul>\n<li>LLaVA: <a href=\"https://llava-vl.github.io/\">https://llava-vl.github.io/</a></li>\n<li>Qwen2-VL: <a href=\"https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct\">https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct</a></li>\n<li>Qwen2.5-VL: <a href=\"https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct\">https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct</a></li>\n<li>InternVL: <a href=\"https://github.com/OpenGVLab/InternVL\">https://github.com/OpenGVLab/InternVL</a></li>\n<li>PaliGemma: <a href=\"https://ai.google.dev/gemma/docs/paligemma\">https://ai.google.dev/gemma/docs/paligemma</a></li>\n<li>Florence-2: <a href=\"https://huggingface.co/microsoft/Florence-2-large\">https://huggingface.co/microsoft/Florence-2-large</a></li>\n<li>Phi-3 Vision: <a href=\"https://huggingface.co/microsoft/Phi-3-vision-128k-instruct\">https://huggingface.co/microsoft/Phi-3-vision-128k-instruct</a></li>\n<li>CogVLM: <a href=\"https://github.com/THUDM/CogVLM\">https://github.com/THUDM/CogVLM</a></li>\n</ul>\n<h3 id=\"inference-tools\">Inference Tools</h3>\n<ul>\n<li>Hugging Face Transformers: <a href=\"https://huggingface.co/docs/transformers/en/model_doc/llava\">https://huggingface.co/docs/transformers/en/model_doc/llava</a></li>\n<li>vLLM: <a href=\"https://docs.vllm.ai/en/latest/models/supported_models/\">https://docs.vllm.ai/en/latest/models/supported_models/</a></li>\n<li>llama.cpp multimodal: <a href=\"https://github.com/ggml-org/llama.cpp/blob/master/docs/multimodal.md\">https://github.com/ggml-org/llama.cpp/blob/master/docs/multimodal.md</a></li>\n<li>Ollama vision models: <a href=\"https://ollama.com/search?c=vision\">https://ollama.com/search?c=vision</a></li>\n</ul>\n<h3 id=\"tutorials-and-guides\">Tutorials and Guides</h3>\n<ul>\n<li>Hugging Face VLM blog: <a href=\"https://huggingface.co/blog/vlms\">https://huggingface.co/blog/vlms</a></li>\n<li>Hugging Face VLMs 2025: <a href=\"https://huggingface.co/blog/vlms-2025\">https://huggingface.co/blog/vlms-2025</a></li>\n<li>SigLIP 2 announcement: <a href=\"https://huggingface.co/blog/siglip2\">https://huggingface.co/blog/siglip2</a></li>\n<li>PaliGemma architecture: <a href=\"https://developers.googleblog.com/gemma-explained-paligemma-architecture/\">https://developers.googleblog.com/gemma-explained-paligemma-architecture/</a></li>\n<li>LLaVA architecture deep dive: <a href=\"https://learnopencv.com/llava-training-a-visual-assistant/\">https://learnopencv.com/llava-training-a-visual-assistant/</a></li>\n</ul>",
            "url": "https://blog.ecitis.org/vision-language-models/",
            "title": "Vision Language Models: Architecture, Inference, and Practical Deployment",
            "summary": "Trace how visual encoders, projectors, and language models combine—and what it takes to train and deploy multimodal systems.",
            "image": "https://blog.ecitis.org/open-graph/vision-language-models.png",
            "date_modified": "2026-01-07T00:00:00.000Z",
            "date_published": "2026-01-07T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Models & Training",
                "VLM",
                "Multimodal",
                "Computer vision"
            ]
        },
        {
            "id": "https://blog.ecitis.org/understanding-gpus/",
            "content_html": "<h2 id=\"why-this-guide\">Why this guide</h2>\n<p>Modern research in machine learning and artificial intelligence requires understanding the fundamental architecture that powers computational workloads. Whether you’re training large language models, processing computer vision tasks, or running complex simulations, Graphics Processing Units (GPUs) have become the backbone of high-performance computing. This comprehensive guide explores the intricate details of GPU architecture, parallel computing paradigms, and practical implementation strategies for research-grade applications on state-of-the-art hardware.</p>\n<p>Understanding these concepts isn’t just about knowing how to run code faster. It’s about fundamentally rethinking how we approach computational problems. The shift from CPU-centric to GPU-accelerated computing enabled breakthroughs in deep learning, scientific computing, and data processing that were previously impossible.</p>\n<hr />\n<h2 id=\"what-is-a-computer-understanding-the-foundation\">What is a Computer? Understanding the Foundation</h2>\n<h3 id=\"hardware-software-the-fundamental-duality\">Hardware + Software: The Fundamental Duality</h3>\n<p>At its most basic level, a computer is a sophisticated machine that transforms electrical signals into meaningful computation through the interaction of hardware and software components. The hardware provides the physical substrate (transistors, memory cells, interconnects) while software provides the logical instructions that orchestrate these components into useful work.</p>\n<p>This duality is crucial to understand because modern GPU programming requires intimate knowledge of both layers. Unlike traditional CPU programming where hardware details are often abstracted away, GPU programming demands awareness of the underlying architecture to achieve optimal performance.</p>\n<p>The hardware layer consists of:</p>\n<ul>\n<li><strong>Processing units</strong>: CPU cores, GPU cores, specialized accelerators</li>\n<li><strong>Memory hierarchy</strong>: Registers, caches, main memory, storage</li>\n<li><strong>Interconnects</strong>: Buses, networks, on-chip communication paths</li>\n<li><strong>Control systems</strong>: Instruction decoders, schedulers, memory controllers</li>\n</ul>\n<p>The software layer encompasses:</p>\n<ul>\n<li><strong>System software</strong>: Operating systems, drivers, runtime libraries</li>\n<li><strong>Development tools</strong>: Compilers, debuggers, profilers</li>\n<li><strong>Application software</strong>: Your actual programs and algorithms</li>\n<li><strong>Middleware</strong>: Libraries, frameworks, abstraction layers</li>\n</ul>\n<h3 id=\"how-is-computation-performed\">How is Computation Performed?</h3>\n<p>Computation at the hardware level involves the coordinated execution of billions of transistor switches that can be in one of two states: conducting (1) or non-conducting (0). These binary states form the foundation of all digital computation through Boolean algebra and logical operations.</p>\n<p>The basic computation cycle follows the von Neumann architecture:</p>\n<ol>\n<li><strong>Fetch</strong>: Retrieve instruction from memory</li>\n<li><strong>Decode</strong>: Interpret what the instruction means</li>\n<li><strong>Execute</strong>: Perform the required operation</li>\n<li><strong>Store</strong>: Write results back to memory</li>\n</ol>\n<p>This cycle repeats billions of times per second, with modern processors executing multiple instructions simultaneously through techniques like:</p>\n<ul>\n<li><strong>Pipelining</strong>: Overlapping instruction execution stages</li>\n<li><strong>Superscalar execution</strong>: Multiple execution units working in parallel</li>\n<li><strong>Out-of-order execution</strong>: Reordering instructions for efficiency</li>\n<li><strong>Speculative execution</strong>: Predicting and pre-executing likely code paths</li>\n</ul>\n<hr />\n<h2 id=\"cpus-and-serial-computing-the-traditional-approach\">CPUs and Serial Computing: The Traditional Approach</h2>\n<h3 id=\"single-core-architecture-and-sequential-execution\">Single-Core Architecture and Sequential Execution</h3>\n<p>Central Processing Units (CPUs) were historically designed around the principle of sequential execution: processing one instruction at a time in a predetermined order. This design philosophy prioritizes:</p>\n<ul>\n<li><strong>Low latency</strong>: Minimize time between instruction issue and completion</li>\n<li><strong>Complex control logic</strong>: Handle arbitrary branching, exceptions, and interrupts</li>\n<li><strong>Large caches</strong>: Reduce memory access latency through prediction and prefetching</li>\n<li><strong>Sophisticated branch prediction</strong>: Anticipate program flow to maintain instruction pipeline efficiency</li>\n</ul>\n<p>A typical CPU core contains:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>┌─────────────────────────────────────────────────────┐</span></span>\n<span class=\"line\"><span>│                CPU Core Architecture                │</span></span>\n<span class=\"line\"><span>├─────────────────────────────────────────────────────┤</span></span>\n<span class=\"line\"><span>│  Instruction Fetch Unit                             │</span></span>\n<span class=\"line\"><span>│  ├─ Branch Predictor                                │</span></span>\n<span class=\"line\"><span>│  ├─ Instruction Cache (L1-I)                        │</span></span>\n<span class=\"line\"><span>│  └─ Prefetch Buffer                                 │</span></span>\n<span class=\"line\"><span>├─────────────────────────────────────────────────────┤</span></span>\n<span class=\"line\"><span>│  Decode and Dispatch                                │</span></span>\n<span class=\"line\"><span>│  ├─ Instruction Decoder                             │</span></span>\n<span class=\"line\"><span>│  ├─ Micro-op Cache                                  │</span></span>\n<span class=\"line\"><span>│  └─ Reorder Buffer                                  │</span></span>\n<span class=\"line\"><span>├─────────────────────────────────────────────────────┤</span></span>\n<span class=\"line\"><span>│  Execution Units                                    │</span></span>\n<span class=\"line\"><span>│  ├─ Integer ALU (multiple units)                    │</span></span>\n<span class=\"line\"><span>│  ├─ Floating Point Unit                             │</span></span>\n<span class=\"line\"><span>│  ├─ Vector/SIMD Unit                                │</span></span>\n<span class=\"line\"><span>│  └─ Load/Store Unit                                 │</span></span>\n<span class=\"line\"><span>├─────────────────────────────────────────────────────┤</span></span>\n<span class=\"line\"><span>│  Memory Subsystem                                   │</span></span>\n<span class=\"line\"><span>│  ├─ L1 Data Cache                                   │</span></span>\n<span class=\"line\"><span>│  ├─ L2 Cache                                        │</span></span>\n<span class=\"line\"><span>│  └─ Memory Management Unit                          │</span></span>\n<span class=\"line\"><span>└─────────────────────────────────────────────────────┘</span></span></code></pre>\n<h3 id=\"multi-core-cpus-parallel-but-still-serial\">Multi-core CPUs: Parallel But Still Serial</h3>\n<p>The introduction of multi-core CPUs represented the first major shift toward parallelism in mainstream computing. However, each core still operates on the same serial computing principles; they provide multiple independent serial processors on the same chip.</p>\n<p>Multi-core systems excel at:</p>\n<ul>\n<li><strong>Task-level parallelism</strong>: Running different programs simultaneously</li>\n<li><strong>Coarse-grained parallelism</strong>: Dividing work into large, independent chunks</li>\n<li><strong>Thread-level parallelism</strong>: Multiple execution contexts with shared memory</li>\n</ul>\n<p>However, they face fundamental limitations:</p>\n<ul>\n<li><strong>Amdahl’s Law</strong>: The sequential portions of code limit overall speedup</li>\n<li><strong>Memory bandwidth</strong>: Multiple cores competing for the same memory subsystem</li>\n<li><strong>Cache coherence</strong>: Overhead of maintaining consistent data across cores</li>\n<li><strong>Load balancing</strong>: Difficulty in evenly distributing work across cores</li>\n</ul>\n<strong>Code Example: Multi-core CPU Utilization (Python)</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> multiprocessing </span><span>as</span><span> mp</span></span>\n<span class=\"line\"><span>import</span><span> numpy </span><span>as</span><span> np</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> cpu_intensive_task</span><span>(</span><span>data_chunk</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Simulate CPU-intensive computation on a data chunk\"\"\"</span></span>\n<span class=\"line\"><span>    result </span><span>=</span><span> np</span><span>.</span><span>zeros_like</span><span>(data_chunk)</span></span>\n<span class=\"line\"><span>    for</span><span> i </span><span>in</span><span> range</span><span>(</span><span>len</span><span>(data_chunk)):</span></span>\n<span class=\"line\"><span>        # Simulate complex computation</span></span>\n<span class=\"line\"><span>        result</span><span>[</span><span>i</span><span>]</span><span> =</span><span> np</span><span>.</span><span>sin</span><span>(data_chunk[i])</span><span> *</span><span> np</span><span>.</span><span>cos</span><span>(data_chunk[i])</span><span> +</span><span> np</span><span>.</span><span>sqrt</span><span>(</span><span>abs</span><span>(data_chunk[i]))</span></span>\n<span class=\"line\"><span>    return</span><span> result</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> serial_computation</span><span>(</span><span>data</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Traditional serial computation\"\"\"</span></span>\n<span class=\"line\"><span>    start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    result </span><span>=</span><span> cpu_intensive_task</span><span>(data)</span></span>\n<span class=\"line\"><span>    end_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    return</span><span> result</span><span>,</span><span> end_time </span><span>-</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> parallel_computation</span><span>(</span><span>data</span><span>,</span><span> num_cores</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Multi-core parallel computation\"\"\"</span></span>\n<span class=\"line\"><span>    start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Split data into chunks</span></span>\n<span class=\"line\"><span>    chunk_size </span><span>=</span><span> len</span><span>(data)</span><span> //</span><span> num_cores</span></span>\n<span class=\"line\"><span>    chunks </span><span>=</span><span> [data</span><span>[</span><span>i</span><span>:</span><span>i </span><span>+</span><span> chunk_size</span><span>]</span><span> for</span><span> i </span><span>in</span><span> range</span><span>(</span><span>0</span><span>, </span><span>len</span><span>(data), chunk_size)</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Process chunks in parallel</span></span>\n<span class=\"line\"><span>    with</span><span> mp</span><span>.</span><span>Pool</span><span>(processes</span><span>=</span><span>num_cores)</span><span> as</span><span> pool</span><span>:</span></span>\n<span class=\"line\"><span>        results </span><span>=</span><span> pool</span><span>.</span><span>map</span><span>(cpu_intensive_task, chunks)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Combine results</span></span>\n<span class=\"line\"><span>    final_result </span><span>=</span><span> np</span><span>.</span><span>concatenate</span><span>(results)</span></span>\n<span class=\"line\"><span>    end_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    return</span><span> final_result</span><span>,</span><span> end_time </span><span>-</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Example usage</span></span>\n<span class=\"line\"><span>if</span><span> __name__</span><span> ==</span><span> \"__main__\"</span><span>:</span></span>\n<span class=\"line\"><span>    # Generate test data</span></span>\n<span class=\"line\"><span>    data_size </span><span>=</span><span> 1000000</span></span>\n<span class=\"line\"><span>    test_data </span><span>=</span><span> np</span><span>.</span><span>random</span><span>.</span><span>random</span><span>(data_size)</span><span> *</span><span> 100</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Compare serial vs parallel performance</span></span>\n<span class=\"line\"><span>    num_cores </span><span>=</span><span> mp</span><span>.</span><span>cpu_count</span><span>()</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Available CPU cores: </span><span>{</span><span>num_cores</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Serial execution</span></span>\n<span class=\"line\"><span>    serial_result</span><span>,</span><span> serial_time </span><span>=</span><span> serial_computation</span><span>(test_data)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Serial execution time: </span><span>{</span><span>serial_time</span><span>:.2f</span><span>}</span><span> seconds\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Parallel execution</span></span>\n<span class=\"line\"><span>    parallel_result</span><span>,</span><span> parallel_time </span><span>=</span><span> parallel_computation</span><span>(test_data, num_cores)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Parallel execution time: </span><span>{</span><span>parallel_time</span><span>:.2f</span><span>}</span><span> seconds\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Speedup: </span><span>{</span><span>serial_time </span><span>/</span><span> parallel_time</span><span>:.2f</span><span>}</span><span>x\"</span><span>)</span></span></code></pre>\n<hr />\n<figure><figcaption><strong>CPU vs GPU Architecture Comparison</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/understanding-gpus-0.webp\" width=\"512\" height=\"468\" alt=\"CPU vs GPU Architecture Comparison\" loading=\"lazy\" decoding=\"async\" /></figure>\n<h2 id=\"gpus-and-parallel-computing-the-paradigm-shift\">GPUs and Parallel Computing: The Paradigm Shift</h2>\n<h3 id=\"understanding-parallel-computing-fundamentals\">Understanding Parallel Computing Fundamentals</h3>\n<p>Graphics Processing Units represent a fundamental departure from CPU design philosophy. Instead of optimizing for serial performance with complex control logic, GPUs are designed around the principle of <strong>massive parallelism</strong>: executing thousands of simple operations simultaneously.</p>\n<p>This architectural difference stems from their original purpose: rendering graphics. Graphics operations are inherently parallel; each pixel can be processed independently, and the same operations (like texture sampling, lighting calculations, or color transformations) are applied to millions of pixels simultaneously.</p>\n<p>The key philosophical differences between CPU and GPU design:</p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Aspect</th><th>CPU</th><th>GPU</th></tr></thead><tbody><tr><td><strong>Design Goal</strong></td><td>Minimize latency for complex tasks</td><td>Maximize throughput for simple tasks</td></tr><tr><td><strong>Core Count</strong></td><td>Few (2-64) powerful cores</td><td>Thousands of simple cores</td></tr><tr><td><strong>Memory</strong></td><td>Large caches, complex hierarchy</td><td>High bandwidth, simpler hierarchy</td></tr><tr><td><strong>Control Logic</strong></td><td>Complex branch prediction, out-of-order</td><td>Simple, in-order execution</td></tr><tr><td><strong>Best For</strong></td><td>Sequential code, complex branching</td><td>Parallel code, regular patterns</td></tr></tbody></table></div>\n<h3 id=\"simd-vs-simt-two-approaches-to-parallelism\">SIMD vs SIMT: Two Approaches to Parallelism</h3>\n<p>Understanding the distinction between SIMD (Single Instruction, Multiple Data) and SIMT (Single Instruction, Multiple Threads) is crucial for effective GPU programming.</p>\n<h4 id=\"simd---single-instruction-multiple-data\">SIMD - Single Instruction Multiple Data</h4>\n<p>SIMD represents the traditional approach to data parallelism found in CPU vector units (like AVX-512) and some GPU architectures. In SIMD:</p>\n<ul>\n<li><strong>Single control unit</strong>: One instruction decoder feeds multiple execution units</li>\n<li><strong>Lockstep execution</strong>: All units execute the same instruction simultaneously</li>\n<li><strong>Data-level parallelism</strong>: Each unit operates on different data elements</li>\n<li><strong>Rigid structure</strong>: All units must execute the same operation or remain idle</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> numpy </span><span>as</span><span> np</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> demonstrate_simd_concept</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Demonstrate SIMD-style operations using NumPy\"\"\"</span></span>\n<span class=\"line\"><span>    # Create large arrays</span></span>\n<span class=\"line\"><span>    a </span><span>=</span><span> np</span><span>.</span><span>random</span><span>.</span><span>random</span><span>(</span><span>1000000</span><span>)</span></span>\n<span class=\"line\"><span>    b </span><span>=</span><span> np</span><span>.</span><span>random</span><span>.</span><span>random</span><span>(</span><span>1000000</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # SIMD-style operation: same instruction applied to all elements</span></span>\n<span class=\"line\"><span>    # This gets compiled to vectorized instructions (AVX, SSE, etc.)</span></span>\n<span class=\"line\"><span>    result </span><span>=</span><span> a </span><span>+</span><span> b  </span><span># Single instruction, multiple data elements</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Element-wise operations are also SIMD-friendly</span></span>\n<span class=\"line\"><span>    complex_result </span><span>=</span><span> np</span><span>.</span><span>sin</span><span>(a)</span><span> *</span><span> np</span><span>.</span><span>cos</span><span>(b)</span><span> +</span><span> np</span><span>.</span><span>sqrt</span><span>(np.</span><span>abs</span><span>(a </span><span>*</span><span> b))</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> result</span><span>,</span><span> complex_result</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Contrast with scalar operations</span></span>\n<span class=\"line\"><span>def</span><span> scalar_equivalent</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Show what the SIMD operation looks like in scalar form\"\"\"</span></span>\n<span class=\"line\"><span>    a </span><span>=</span><span> np</span><span>.</span><span>random</span><span>.</span><span>random</span><span>(</span><span>1000000</span><span>)</span></span>\n<span class=\"line\"><span>    b </span><span>=</span><span> np</span><span>.</span><span>random</span><span>.</span><span>random</span><span>(</span><span>1000000</span><span>)</span></span>\n<span class=\"line\"><span>    result </span><span>=</span><span> np</span><span>.</span><span>zeros</span><span>(</span><span>1000000</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # This would be much slower - one operation at a time</span></span>\n<span class=\"line\"><span>    for</span><span> i </span><span>in</span><span> range</span><span>(</span><span>len</span><span>(a)):</span></span>\n<span class=\"line\"><span>        result</span><span>[</span><span>i</span><span>]</span><span> =</span><span> a</span><span>[</span><span>i</span><span>]</span><span> +</span><span> b</span><span>[</span><span>i</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> result</span></span></code></pre>\n<h4 id=\"simt---single-instruction-multiple-threads\">SIMT - Single Instruction Multiple Threads</h4>\n<p>SIMT, pioneered by NVIDIA’s CUDA architecture, provides a more flexible approach:</p>\n<ul>\n<li><strong>Thread-level parallelism</strong>: Each thread has its own program counter and registers</li>\n<li><strong>Flexible execution</strong>: Threads can follow different execution paths (with caveats)</li>\n<li><strong>Warp-based execution</strong>: Threads are grouped into warps (typically 32 threads) for execution</li>\n<li><strong>Dynamic scheduling</strong>: Hardware scheduler manages thread execution</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// CUDA kernel demonstrating SIMT execution</span></span>\n<span class=\"line\"><span>__global__ void simt_vector_add(float* a, float* b, float* c, int n) {</span></span>\n<span class=\"line\"><span>    // Each thread computes its unique index</span></span>\n<span class=\"line\"><span>    int tid = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Boundary check - threads can have different execution paths</span></span>\n<span class=\"line\"><span>    if (tid &lt; n) {</span></span>\n<span class=\"line\"><span>        c[tid] = a[tid] + b[tid];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Conditional execution - some threads may execute this, others may not</span></span>\n<span class=\"line\"><span>        if (a[tid] &gt; 0.5f) {</span></span>\n<span class=\"line\"><span>            c[tid] *= 2.0f;</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>    // Threads that fail the boundary check essentially become idle</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"task-parallelism-vs-warp-execution\">Task Parallelism vs Warp Execution</h3>\n<p>GPUs implement parallelism through a hierarchical execution model:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>Grid (entire kernel launch)</span></span>\n<span class=\"line\"><span>├─ Block 0</span></span>\n<span class=\"line\"><span>│  ├─ Warp 0 (threads 0-31)</span></span>\n<span class=\"line\"><span>│  ├─ Warp 1 (threads 32-63)</span></span>\n<span class=\"line\"><span>│  └─ ...</span></span>\n<span class=\"line\"><span>├─ Block 1</span></span>\n<span class=\"line\"><span>│  ├─ Warp 0 (threads 0-31)</span></span>\n<span class=\"line\"><span>│  └─ ...</span></span>\n<span class=\"line\"><span>└─ ...</span></span></code></pre>\n<p><strong>Warp Execution Model</strong>:</p>\n<ul>\n<li><strong>Warp</strong>: Group of 32 threads executed in lockstep</li>\n<li><strong>SIMT execution</strong>: All threads in a warp execute the same instruction</li>\n<li><strong>Divergence</strong>: When threads take different paths, some become inactive</li>\n<li><strong>Convergence</strong>: Threads rejoin at control flow merge points</li>\n</ul>\n<p>Here’s a detailed CUDA example showing warp behavior:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>#include &lt;cuda_runtime.h&gt;</span></span>\n<span class=\"line\"><span>#include &lt;stdio.h&gt;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>__global__ void demonstrate_warp_behavior(int* data, int n) {</span></span>\n<span class=\"line\"><span>    int tid = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span>    int warp_id = tid / 32;</span></span>\n<span class=\"line\"><span>    int lane_id = tid % 32;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    if (tid &lt; n) {</span></span>\n<span class=\"line\"><span>        // All threads in warp execute this</span></span>\n<span class=\"line\"><span>        data[tid] = tid;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Warp divergence example</span></span>\n<span class=\"line\"><span>        if (lane_id &lt; 16) {</span></span>\n<span class=\"line\"><span>            // First half of warp executes this</span></span>\n<span class=\"line\"><span>            data[tid] *= 2;</span></span>\n<span class=\"line\"><span>        } else {</span></span>\n<span class=\"line\"><span>            // Second half executes this</span></span>\n<span class=\"line\"><span>            data[tid] *= 3;</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span>        // Warp reconverges here</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        data[tid] += warp_id;</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>// Host code</span></span>\n<span class=\"line\"><span>int main() {</span></span>\n<span class=\"line\"><span>    const int n = 1024;</span></span>\n<span class=\"line\"><span>    int* h_data = new int[n];</span></span>\n<span class=\"line\"><span>    int* d_data;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Allocate GPU memory</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_data, n * sizeof(int));</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Launch kernel with appropriate grid/block dimensions</span></span>\n<span class=\"line\"><span>    dim3 block_size(256);  // 8 warps per block</span></span>\n<span class=\"line\"><span>    dim3 grid_size((n + block_size.x - 1) / block_size.x);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    demonstrate_warp_behavior&lt;&lt;&lt;grid_size, block_size&gt;&gt;&gt;(d_data, n);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Copy result back</span></span>\n<span class=\"line\"><span>    cudaMemcpy(h_data, d_data, n * sizeof(int), cudaMemcpyDeviceToHost);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Print results for first few elements</span></span>\n<span class=\"line\"><span>    for (int i = 0; i &lt; 64; i++) {</span></span>\n<span class=\"line\"><span>        printf(\"Thread %d: %d\\n\", i, h_data[i]);</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    delete[] h_data;</span></span>\n<span class=\"line\"><span>    cudaFree(d_data);</span></span>\n<span class=\"line\"><span>    return 0;</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"why-parallelism-videos-matrices-and-modern-applications\">Why Parallelism? Videos, Matrices, and Modern Applications</h3>\n<p>The motivation for parallel computing extends far beyond graphics:</p>\n<p><strong>Graphics and Video Processing</strong>:</p>\n<ul>\n<li><strong>Pixel independence</strong>: Each pixel can be processed without knowledge of others</li>\n<li><strong>Regular memory patterns</strong>: Predictable access patterns enable efficient memory usage</li>\n<li><strong>High computational density</strong>: Simple operations repeated millions of times</li>\n</ul>\n<p><strong>Matrix Operations</strong>:\nModern AI workloads are fundamentally based on linear algebra operations that are highly parallelizable:</p>\n<strong>Code Example: GPU vs CPU Matrix &amp; Convolution Operations (PyTorch)</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> demonstrate_gpu_matrix_operations</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Show the performance difference between CPU and GPU matrix operations\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Create large matrices</span></span>\n<span class=\"line\"><span>    size </span><span>=</span><span> 4096</span></span>\n<span class=\"line\"><span>    cpu_a </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(size, size, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"><span>    cpu_b </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(size, size, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # GPU versions</span></span>\n<span class=\"line\"><span>    gpu_a </span><span>=</span><span> cpu_a</span><span>.</span><span>cuda</span><span>()</span></span>\n<span class=\"line\"><span>    gpu_b </span><span>=</span><span> cpu_b</span><span>.</span><span>cuda</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # CPU matrix multiplication</span></span>\n<span class=\"line\"><span>    start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    cpu_result </span><span>=</span><span> torch</span><span>.</span><span>matmul</span><span>(cpu_a, cpu_b)</span></span>\n<span class=\"line\"><span>    cpu_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # GPU matrix multiplication</span></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span><span>  # Ensure clean timing</span></span>\n<span class=\"line\"><span>    start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    gpu_result </span><span>=</span><span> torch</span><span>.</span><span>matmul</span><span>(gpu_a, gpu_b)</span></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span><span>  # Wait for GPU to complete</span></span>\n<span class=\"line\"><span>    gpu_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Matrix size: </span><span>{</span><span>size</span><span>}</span><span>x</span><span>{</span><span>size</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"CPU time: </span><span>{</span><span>cpu_time</span><span>:.4f</span><span>}</span><span> seconds\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"GPU time: </span><span>{</span><span>gpu_time</span><span>:.4f</span><span>}</span><span> seconds\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Speedup: </span><span>{</span><span>cpu_time </span><span>/</span><span> gpu_time</span><span>:.2f</span><span>}</span><span>x\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Verify results are equivalent (within floating point precision)</span></span>\n<span class=\"line\"><span>    difference </span><span>=</span><span> torch</span><span>.</span><span>max</span><span>(torch.</span><span>abs</span><span>(cpu_result </span><span>-</span><span> gpu_result.</span><span>cpu</span><span>()))</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Maximum difference: </span><span>{</span><span>difference</span><span>:.2e</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># More complex example: Convolution operations</span></span>\n<span class=\"line\"><span>def</span><span> demonstrate_convolution_parallelism</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Show how convolution operations benefit from GPU parallelism\"\"\"</span></span>\n<span class=\"line\"><span>    import</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>functional </span><span>as</span><span> F</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Create sample input (batch_size=32, channels=256, height=64, width=64)</span></span>\n<span class=\"line\"><span>    input_tensor </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>32</span><span>, </span><span>256</span><span>, </span><span>64</span><span>, </span><span>64</span><span>)</span></span>\n<span class=\"line\"><span>    # Convolution kernel (out_channels=512, in_channels=256, kernel_size=3x3)</span></span>\n<span class=\"line\"><span>    kernel </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>512</span><span>, </span><span>256</span><span>, </span><span>3</span><span>, </span><span>3</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # CPU convolution</span></span>\n<span class=\"line\"><span>    start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    cpu_output </span><span>=</span><span> F</span><span>.</span><span>conv2d</span><span>(input_tensor, kernel, padding</span><span>=</span><span>1</span><span>)</span></span>\n<span class=\"line\"><span>    cpu_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # GPU convolution</span></span>\n<span class=\"line\"><span>    input_gpu </span><span>=</span><span> input_tensor</span><span>.</span><span>cuda</span><span>()</span></span>\n<span class=\"line\"><span>    kernel_gpu </span><span>=</span><span> kernel</span><span>.</span><span>cuda</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>    start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    gpu_output </span><span>=</span><span> F</span><span>.</span><span>conv2d</span><span>(input_gpu, kernel_gpu, padding</span><span>=</span><span>1</span><span>)</span></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>    gpu_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"</span><span>\\n</span><span>Convolution operation:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Input shape: </span><span>{</span><span>input_tensor.shape</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Output shape: </span><span>{</span><span>cpu_output.shape</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"CPU time: </span><span>{</span><span>cpu_time</span><span>:.4f</span><span>}</span><span> seconds\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"GPU time: </span><span>{</span><span>gpu_time</span><span>:.4f</span><span>}</span><span> seconds\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Speedup: </span><span>{</span><span>cpu_time </span><span>/</span><span> gpu_time</span><span>:.2f</span><span>}</span><span>x\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>if</span><span> __name__</span><span> ==</span><span> \"__main__\"</span><span>:</span></span>\n<span class=\"line\"><span>    demonstrate_gpu_matrix_operations</span><span>()</span></span>\n<span class=\"line\"><span>    demonstrate_convolution_parallelism</span><span>()</span></span></code></pre>\n<hr />\n<h2 id=\"other-key-hardware-elements-the-complete-system\">Other Key Hardware Elements: The Complete System</h2>\n<h3 id=\"memory-hierarchy-understanding-the-storage-pyramid\">Memory Hierarchy: Understanding the Storage Pyramid</h3>\n<p>Modern computing systems implement a sophisticated memory hierarchy designed to balance capacity, speed, and cost. Understanding this hierarchy is crucial for writing efficient GPU code.</p>\n<h4 id=\"system-ram-random-access-memory\">System RAM (Random Access Memory)</h4>\n<p>System RAM serves as the primary storage for data and programs currently in use. Key characteristics:</p>\n<ul>\n<li><strong>Capacity</strong>: Typically 16GB to 1TB in modern systems</li>\n<li><strong>Bandwidth</strong>: 100-400 GB/s for high-end systems</li>\n<li><strong>Latency</strong>: 50-100 nanoseconds access time</li>\n<li><strong>Volatile</strong>: Data is lost when power is removed</li>\n<li><strong>Shared</strong>: Accessible by CPU and (through PCIe) GPU</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> psutil</span></span>\n<span class=\"line\"><span>import</span><span> numpy </span><span>as</span><span> np</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> analyze_system_memory</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Analyze system memory characteristics\"\"\"</span></span>\n<span class=\"line\"><span>    memory </span><span>=</span><span> psutil</span><span>.</span><span>virtual_memory</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"System Memory Analysis:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Total RAM: </span><span>{</span><span>memory.total </span><span>/</span><span> (</span><span>1024</span><span>**</span><span>3</span><span>)</span><span>:.2f</span><span>}</span><span> GB\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Available RAM: </span><span>{</span><span>memory.available </span><span>/</span><span> (</span><span>1024</span><span>**</span><span>3</span><span>)</span><span>:.2f</span><span>}</span><span> GB\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Used RAM: </span><span>{</span><span>memory.used </span><span>/</span><span> (</span><span>1024</span><span>**</span><span>3</span><span>)</span><span>:.2f</span><span>}</span><span> GB\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Memory utilization: </span><span>{</span><span>memory.percent</span><span>:.1f</span><span>}</span><span>%\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Demonstrate memory bandwidth measurement</span></span>\n<span class=\"line\"><span>    def</span><span> measure_memory_bandwidth</span><span>():</span></span>\n<span class=\"line\"><span>        \"\"\"Rough estimate of memory bandwidth\"\"\"</span></span>\n<span class=\"line\"><span>        size </span><span>=</span><span> 100_000_000</span><span>  # 100M floats = ~400MB</span></span>\n<span class=\"line\"><span>        data </span><span>=</span><span> np</span><span>.</span><span>random</span><span>.</span><span>random</span><span>(size).</span><span>astype</span><span>(np.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        import</span><span> time</span></span>\n<span class=\"line\"><span>        # Memory copy operation</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>        copy </span><span>=</span><span> np</span><span>.</span><span>copy</span><span>(data)</span><span>  # This tests memory bandwidth</span></span>\n<span class=\"line\"><span>        end_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        bytes_transferred </span><span>=</span><span> size </span><span>*</span><span> 4</span><span> *</span><span> 2</span><span>  # 4 bytes per float, read + write</span></span>\n<span class=\"line\"><span>        bandwidth </span><span>=</span><span> bytes_transferred </span><span>/</span><span> (end_time </span><span>-</span><span> start_time) </span><span>/</span><span> (</span><span>1024</span><span>**</span><span>3</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"Estimated memory bandwidth: </span><span>{</span><span>bandwidth</span><span>:.2f</span><span>}</span><span> GB/s\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    measure_memory_bandwidth</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>analyze_system_memory</span><span>()</span></span></code></pre>\n<h4 id=\"cache-memory-sram---static-random-access-memory\">Cache Memory (SRAM - Static Random Access Memory)</h4>\n<p>Cache memory provides ultra-fast access to frequently used data through a hierarchy of increasingly smaller but faster storage levels:</p>\n<p><strong>L1 Cache</strong>:</p>\n<ul>\n<li><strong>Size</strong>: 32-64 KB per core</li>\n<li><strong>Latency</strong>: 1-2 CPU cycles</li>\n<li><strong>Organization</strong>: Separate instruction and data caches</li>\n<li><strong>Bandwidth</strong>: ~1 TB/s per core</li>\n</ul>\n<p><strong>L2 Cache</strong>:</p>\n<ul>\n<li><strong>Size</strong>: 256KB - 1MB per core</li>\n<li><strong>Latency</strong>: 3-10 CPU cycles</li>\n<li><strong>Shared</strong>: Often shared between cores</li>\n<li><strong>Bandwidth</strong>: ~500 GB/s per core</li>\n</ul>\n<p><strong>L3 Cache</strong>:</p>\n<ul>\n<li><strong>Size</strong>: 8-64MB total</li>\n<li><strong>Latency</strong>: 10-20 CPU cycles</li>\n<li><strong>Shared</strong>: Across all cores</li>\n<li><strong>Bandwidth</strong>: ~200 GB/s total</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>#include</span><span> &lt;chrono&gt;</span></span>\n<span class=\"line\"><span>#include</span><span> &lt;iostream&gt;</span></span>\n<span class=\"line\"><span>#include</span><span> &lt;vector&gt;</span></span>\n<span class=\"line\"><span>#include</span><span> &lt;random&gt;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>class</span><span> CacheAnalyzer</span><span> {</span></span>\n<span class=\"line\"><span>public</span><span>:</span></span>\n<span class=\"line\"><span>    static</span><span> void</span><span> demonstrate_cache_hierarchy</span><span>() {</span></span>\n<span class=\"line\"><span>        // Test different data sizes to show cache effects</span></span>\n<span class=\"line\"><span>        std</span><span>::</span><span>vector</span><span>&lt;size_t&gt;</span><span> sizes </span><span>=</span><span> {</span></span>\n<span class=\"line\"><span>            1024</span><span>,</span><span>        // 4KB - fits in L1</span></span>\n<span class=\"line\"><span>            64</span><span> *</span><span> 1024</span><span>,</span><span>   // 256KB - fits in L2</span></span>\n<span class=\"line\"><span>            4</span><span> *</span><span> 1024</span><span> *</span><span> 1024</span><span>,</span><span>  // 16MB - fits in L3</span></span>\n<span class=\"line\"><span>            64</span><span> *</span><span> 1024</span><span> *</span><span> 1024</span><span>  // 256MB - exceeds cache</span></span>\n<span class=\"line\"><span>        };</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        for</span><span> (</span><span>size_t</span><span> size </span><span>:</span><span> sizes) {</span></span>\n<span class=\"line\"><span>            double</span><span> time </span><span>=</span><span> measure_access_time</span><span>(size);</span></span>\n<span class=\"line\"><span>            std</span><span>::</span><span>cout </span><span>&lt;&lt;</span><span> \"Size: \"</span><span> &lt;&lt;</span><span> size </span><span>*</span><span> sizeof</span><span>(</span><span>int</span><span>) </span><span>/</span><span> 1024</span><span> &lt;&lt;</span><span> \" KB, \"</span></span>\n<span class=\"line\"><span>                      &lt;&lt;</span><span> \"Time per access: \"</span><span> &lt;&lt;</span><span> time </span><span>&lt;&lt;</span><span> \" ns\"</span><span> &lt;&lt;</span><span> std</span><span>::</span><span>endl;</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>private</span><span>:</span></span>\n<span class=\"line\"><span>    static</span><span> double</span><span> measure_access_time</span><span>(</span><span>size_t</span><span> size) {</span></span>\n<span class=\"line\"><span>        std</span><span>::</span><span>vector</span><span>&lt;int&gt;</span><span> data</span><span>(size);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        // Initialize with random values</span></span>\n<span class=\"line\"><span>        std</span><span>::</span><span>random_device rd;</span></span>\n<span class=\"line\"><span>        std</span><span>::</span><span>mt19937 </span><span>gen</span><span>(</span><span>rd</span><span>());</span></span>\n<span class=\"line\"><span>        for</span><span> (</span><span>size_t</span><span> i </span><span>=</span><span> 0</span><span>; i </span><span>&lt;</span><span> size; i</span><span>++</span><span>) {</span></span>\n<span class=\"line\"><span>            data</span><span>[i] </span><span>=</span><span> gen</span><span>() </span><span>%</span><span> size;</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        // Measure random access time</span></span>\n<span class=\"line\"><span>        volatile</span><span> int</span><span> sum </span><span>=</span><span> 0</span><span>;</span><span>  // Prevent optimization</span></span>\n<span class=\"line\"><span>        auto</span><span> start </span><span>=</span><span> std</span><span>::</span><span>chrono</span><span>::</span><span>high_resolution_clock</span><span>::</span><span>now</span><span>();</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        size_t</span><span> iterations </span><span>=</span><span> 1000000</span><span>;</span></span>\n<span class=\"line\"><span>        for</span><span> (</span><span>size_t</span><span> i </span><span>=</span><span> 0</span><span>; i </span><span>&lt;</span><span> iterations; i</span><span>++</span><span>) {</span></span>\n<span class=\"line\"><span>            sum </span><span>+=</span><span> data</span><span>[i </span><span>%</span><span> size];</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        auto</span><span> end </span><span>=</span><span> std</span><span>::</span><span>chrono</span><span>::</span><span>high_resolution_clock</span><span>::</span><span>now</span><span>();</span></span>\n<span class=\"line\"><span>        auto</span><span> duration </span><span>=</span><span> std</span><span>::</span><span>chrono</span><span>::</span><span>duration_cast</span><span>&lt;std</span><span>::</span><span>chrono</span><span>::</span><span>nanoseconds</span><span>&gt;(end </span><span>-</span><span> start);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> static_cast&lt;double&gt;</span><span>(</span><span>duration</span><span>.</span><span>count</span><span>()) </span><span>/</span><span> iterations;</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>};</span></span></code></pre>\n<h4 id=\"vram-video-random-access-memory\">VRAM (Video Random Access Memory)</h4>\n<p>VRAM is specialized high-bandwidth memory attached directly to the GPU:</p>\n<ul>\n<li><strong>Capacity</strong>: 8GB to 80GB on modern GPUs</li>\n<li><strong>Bandwidth</strong>: 500GB/s to 3TB/s (much higher than system RAM)</li>\n<li><strong>Latency</strong>: Similar to system RAM but with much higher throughput</li>\n<li><strong>Architecture</strong>: Wide memory buses (384-bit, 512-bit, or wider)</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> analyze_vram_characteristics</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Analyze GPU memory characteristics\"\"\"</span></span>\n<span class=\"line\"><span>    if</span><span> not</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>is_available</span><span>():</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>\"CUDA not available\"</span><span>)</span></span>\n<span class=\"line\"><span>        return</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Get GPU memory info</span></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>current_device</span><span>()</span></span>\n<span class=\"line\"><span>    gpu_props </span><span>=</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>get_device_properties</span><span>(device)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"GPU: </span><span>{</span><span>gpu_props.name</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Total VRAM: </span><span>{</span><span>gpu_props.total_memory </span><span>/</span><span> (</span><span>1024</span><span>**</span><span>3</span><span>)</span><span>:.2f</span><span>}</span><span> GB\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Multiprocessors: </span><span>{</span><span>gpu_props.multi_processor_count</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Measure memory bandwidth</span></span>\n<span class=\"line\"><span>    def</span><span> measure_gpu_bandwidth</span><span>():</span></span>\n<span class=\"line\"><span>        \"\"\"Measure GPU memory bandwidth\"\"\"</span></span>\n<span class=\"line\"><span>        size </span><span>=</span><span> 100_000_000</span><span>  # 100M floats</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Allocate GPU memory</span></span>\n<span class=\"line\"><span>        data </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(size, device</span><span>=</span><span>'cuda'</span><span>, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Warm up</span></span>\n<span class=\"line\"><span>        for</span><span> _ </span><span>in</span><span> range</span><span>(</span><span>10</span><span>):</span></span>\n<span class=\"line\"><span>            copy </span><span>=</span><span> data</span><span>.</span><span>clone</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Measure bandwidth</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>        iterations </span><span>=</span><span> 50</span></span>\n<span class=\"line\"><span>        for</span><span> _ </span><span>in</span><span> range</span><span>(iterations):</span></span>\n<span class=\"line\"><span>            copy </span><span>=</span><span> data</span><span>.</span><span>clone</span><span>()</span></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        end_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        bytes_per_iteration </span><span>=</span><span> size </span><span>*</span><span> 4</span><span> *</span><span> 2</span><span>  # read + write</span></span>\n<span class=\"line\"><span>        total_bytes </span><span>=</span><span> bytes_per_iteration </span><span>*</span><span> iterations</span></span>\n<span class=\"line\"><span>        bandwidth </span><span>=</span><span> total_bytes </span><span>/</span><span> (end_time </span><span>-</span><span> start_time) </span><span>/</span><span> (</span><span>1024</span><span>**</span><span>3</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"Measured GPU memory bandwidth: </span><span>{</span><span>bandwidth</span><span>:.2f</span><span>}</span><span> GB/s\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Compare with theoretical peak</span></span>\n<span class=\"line\"><span>        # This is approximate and varies by GPU</span></span>\n<span class=\"line\"><span>        theoretical_peak </span><span>=</span><span> {</span></span>\n<span class=\"line\"><span>            'A100'</span><span>:</span><span> 1555</span><span>,</span><span>  # GB/s</span></span>\n<span class=\"line\"><span>            'V100'</span><span>:</span><span> 900</span><span>,</span></span>\n<span class=\"line\"><span>            'RTX 4090'</span><span>:</span><span> 1008</span><span>,</span></span>\n<span class=\"line\"><span>            'RTX 3090'</span><span>:</span><span> 936</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        for</span><span> gpu_name</span><span>,</span><span> peak </span><span>in</span><span> theoretical_peak</span><span>.</span><span>items</span><span>():</span></span>\n<span class=\"line\"><span>            if</span><span> gpu_name</span><span>.</span><span>lower</span><span>()</span><span> in</span><span> gpu_props</span><span>.</span><span>name</span><span>.</span><span>lower</span><span>():</span></span>\n<span class=\"line\"><span>                efficiency </span><span>=</span><span> (bandwidth </span><span>/</span><span> peak) </span><span>*</span><span> 100</span></span>\n<span class=\"line\"><span>                print</span><span>(</span><span>f</span><span>\"Theoretical peak: </span><span>{</span><span>peak</span><span>}</span><span> GB/s\"</span><span>)</span></span>\n<span class=\"line\"><span>                print</span><span>(</span><span>f</span><span>\"Achieved efficiency: </span><span>{</span><span>efficiency</span><span>:.1f</span><span>}</span><span>%\"</span><span>)</span></span>\n<span class=\"line\"><span>                break</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    measure_gpu_bandwidth</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>analyze_vram_characteristics</span><span>()</span></span></code></pre>\n<h3 id=\"storage-the-persistent-foundation\">Storage: The Persistent Foundation</h3>\n<h4 id=\"ssds-solid-state-drives\">SSDs (Solid State Drives)</h4>\n<p>Modern SSDs provide fast persistent storage crucial for handling large datasets:</p>\n<ul>\n<li><strong>Technology</strong>: NAND flash memory with sophisticated controllers</li>\n<li><strong>Performance</strong>: 3-7 GB/s sequential read/write (NVMe SSDs)</li>\n<li><strong>Latency</strong>: 50-100 microseconds (much lower than HDDs)</li>\n<li><strong>Durability</strong>: Limited write cycles but very reliable for reads</li>\n</ul>\n<h4 id=\"hdds-hard-disk-drives\">HDDs (Hard Disk Drives)</h4>\n<p>Traditional magnetic storage still used for large, infrequently accessed datasets:</p>\n<ul>\n<li><strong>Capacity</strong>: Up to 20TB+ per drive</li>\n<li><strong>Performance</strong>: 100-250 MB/s sequential throughput</li>\n<li><strong>Latency</strong>: 3-10 milliseconds (due to mechanical movement)</li>\n<li><strong>Cost</strong>: Much cheaper per GB than SSDs</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> os</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"><span>import</span><span> tempfile</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> benchmark_storage</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Benchmark storage performance characteristics\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> benchmark_write_read</span><span>(</span><span>file_path</span><span>,</span><span> size_mb</span><span>):</span></span>\n<span class=\"line\"><span>        \"\"\"Benchmark write and read performance\"\"\"</span></span>\n<span class=\"line\"><span>        data </span><span>=</span><span> b</span><span>'0'</span><span> *</span><span> (</span><span>1024</span><span> *</span><span> 1024</span><span>)  </span><span># 1MB chunk</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Write benchmark</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>        with</span><span> open</span><span>(file_path, </span><span>'wb'</span><span>)</span><span> as</span><span> f</span><span>:</span></span>\n<span class=\"line\"><span>            for</span><span> _ </span><span>in</span><span> range</span><span>(size_mb):</span></span>\n<span class=\"line\"><span>                f</span><span>.</span><span>write</span><span>(data)</span></span>\n<span class=\"line\"><span>        write_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Read benchmark</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>        with</span><span> open</span><span>(file_path, </span><span>'rb'</span><span>)</span><span> as</span><span> f</span><span>:</span></span>\n<span class=\"line\"><span>            while</span><span> f</span><span>.</span><span>read</span><span>(</span><span>1024</span><span> *</span><span> 1024</span><span>):</span></span>\n<span class=\"line\"><span>                pass</span></span>\n<span class=\"line\"><span>        read_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Clean up</span></span>\n<span class=\"line\"><span>        os</span><span>.</span><span>remove</span><span>(file_path)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        write_speed </span><span>=</span><span> size_mb </span><span>/</span><span> write_time</span></span>\n<span class=\"line\"><span>        read_speed </span><span>=</span><span> size_mb </span><span>/</span><span> read_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> write_speed</span><span>,</span><span> read_speed</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Test with different file sizes</span></span>\n<span class=\"line\"><span>    test_sizes </span><span>=</span><span> [</span><span>100</span><span>,</span><span> 1000</span><span>]  </span><span># 100MB, 1GB</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    for</span><span> size </span><span>in</span><span> test_sizes</span><span>:</span></span>\n<span class=\"line\"><span>        with</span><span> tempfile</span><span>.</span><span>NamedTemporaryFile</span><span>(delete</span><span>=</span><span>False</span><span>)</span><span> as</span><span> tmp</span><span>:</span></span>\n<span class=\"line\"><span>            write_speed</span><span>,</span><span> read_speed </span><span>=</span><span> benchmark_write_read</span><span>(tmp.name, size)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>            print</span><span>(</span><span>f</span><span>\"File size: </span><span>{</span><span>size</span><span>}</span><span> MB\"</span><span>)</span></span>\n<span class=\"line\"><span>            print</span><span>(</span><span>f</span><span>\"Write speed: </span><span>{</span><span>write_speed</span><span>:.2f</span><span>}</span><span> MB/s\"</span><span>)</span></span>\n<span class=\"line\"><span>            print</span><span>(</span><span>f</span><span>\"Read speed: </span><span>{</span><span>read_speed</span><span>:.2f</span><span>}</span><span> MB/s\"</span><span>)</span></span>\n<span class=\"line\"><span>            print</span><span>(</span><span>\"-\"</span><span> *</span><span> 40</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>benchmark_storage</span><span>()</span></span></code></pre>\n<hr />\n<h2 id=\"a-short-history-of-gpus-evolution-of-parallel-computing\">A Short History of GPUs: Evolution of Parallel Computing</h2>\n<p>Understanding the historical development of GPUs provides crucial context for their current capabilities and future directions.</p>\n<h3 id=\"1990-1992-the-dawn-of-3d-rendering\">1990-1992: The Dawn of 3D Rendering</h3>\n<p>The early 1990s marked the beginning of consumer 3D graphics acceleration. Before this era, all graphics rendering was performed by the CPU, limiting the complexity and realism of computer graphics.</p>\n<p><strong>Key developments</strong>:</p>\n<ul>\n<li><strong>Software rendering</strong>: Early 3D games like Wolfenstein 3D used clever CPU-based techniques</li>\n<li><strong>Fixed-function pipelines</strong>: Early 3D accelerators implemented fixed graphics operations in hardware</li>\n<li><strong>Limited programmability</strong>: Graphics cards could only perform predefined operations</li>\n</ul>\n<h3 id=\"1996-3dfx-voodoo-graphics-the-first-consumer-3d-accelerator\">1996: 3DFX Voodoo Graphics - The First Consumer 3D Accelerator</h3>\n<p>The 3DFX Voodoo Graphics card revolutionized PC gaming by providing hardware acceleration for 3D graphics operations:</p>\n<p><strong>Technical specifications</strong>:</p>\n<ul>\n<li><strong>Dedicated 3D acceleration</strong>: Separate card just for 3D rendering</li>\n<li><strong>Texture mapping</strong>: Hardware support for applying textures to 3D surfaces</li>\n<li><strong>Z-buffering</strong>: Hardware depth testing for proper 3D object ordering</li>\n<li><strong>Glide API</strong>: Proprietary programming interface for 3D applications</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// Example of early 3D graphics programming (pseudo-Glide API)</span></span>\n<span class=\"line\"><span>void</span><span> render_triangle</span><span>() {</span></span>\n<span class=\"line\"><span>    // Set rendering state</span></span>\n<span class=\"line\"><span>    grColorCombine</span><span>(GR_COMBINE_FUNCTION_SCALE_OTHER</span><span>,</span></span>\n<span class=\"line\"><span>                   GR_COMBINE_FACTOR_ONE</span><span>,</span></span>\n<span class=\"line\"><span>                   GR_COMBINE_LOCAL_NONE</span><span>,</span></span>\n<span class=\"line\"><span>                   GR_COMBINE_OTHER_TEXTURE);</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Define triangle vertices</span></span>\n<span class=\"line\"><span>    GrVertex </span><span>vertices</span><span>[</span><span>3</span><span>] </span><span>=</span><span> {</span></span>\n<span class=\"line\"><span>        {</span><span>100.0</span><span>f</span><span>,</span><span> 100.0</span><span>f</span><span>,</span><span> 0.5</span><span>f</span><span>,</span><span> 1.0</span><span>f</span><span>,</span><span> 0.0</span><span>f</span><span>,</span><span> 0.0</span><span>f</span><span>}</span><span>,</span><span>  // x, y, z, s, t, w</span></span>\n<span class=\"line\"><span>        {</span><span>200.0</span><span>f</span><span>,</span><span> 100.0</span><span>f</span><span>,</span><span> 0.5</span><span>f</span><span>,</span><span> 1.0</span><span>f</span><span>,</span><span> 1.0</span><span>f</span><span>,</span><span> 0.0</span><span>f</span><span>}</span><span>,</span></span>\n<span class=\"line\"><span>        {</span><span>150.0</span><span>f</span><span>,</span><span> 200.0</span><span>f</span><span>,</span><span> 0.5</span><span>f</span><span>,</span><span> 1.0</span><span>f</span><span>,</span><span> 0.5</span><span>f</span><span>,</span><span> 1.0</span><span>f</span><span>}</span></span>\n<span class=\"line\"><span>    };</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Render the triangle</span></span>\n<span class=\"line\"><span>    grDrawTriangle</span><span>(</span><span>&amp;</span><span>vertices</span><span>[</span><span>0</span><span>]</span><span>,</span><span> &amp;</span><span>vertices</span><span>[</span><span>1</span><span>]</span><span>,</span><span> &amp;</span><span>vertices</span><span>[</span><span>2</span><span>]);</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"1997-nvidia-riva-128-and-ati-rage-pro\">1997: NVIDIA RIVA 128 and ATI Rage Pro</h3>\n<p>This period saw the integration of 2D and 3D graphics capabilities:</p>\n<p><strong>NVIDIA RIVA 128</strong>:</p>\n<ul>\n<li><strong>Unified architecture</strong>: Combined 2D GUI acceleration with 3D rendering</li>\n<li><strong>AGP interface</strong>: Faster connection to system memory than PCI</li>\n<li><strong>Hardware transform</strong>: Basic 3D coordinate transformation in hardware</li>\n</ul>\n<p><strong>ATI Rage Pro</strong>:</p>\n<ul>\n<li><strong>Video acceleration</strong>: Hardware support for video decoding and playback</li>\n<li><strong>Dual-head support</strong>: Multiple monitor outputs</li>\n<li><strong>Improved memory bandwidth</strong>: Better utilization of available memory</li>\n</ul>\n<h3 id=\"1990s-gpgpu-research-begins\">1990s: GPGPU Research Begins</h3>\n<p>Forward-thinking researchers recognized that graphics hardware could be useful for non-graphics computations:</p>\n<p><strong>Early GPGPU techniques</strong>:</p>\n<ul>\n<li><strong>Texture-based computing</strong>: Encoding data as textures and using pixel shaders for computation</li>\n<li><strong>Fragment shader programming</strong>: Using graphics shaders for general-purpose calculations</li>\n<li><strong>Render-to-texture</strong>: Computing results and storing them in textures for further processing</li>\n</ul>\n<p><strong>1996: Stanford BrookGPU - Stream/Kernel Programming Model</strong></p>\n<p>Stanford researchers developed BrookGPU, one of the first systematic approaches to general-purpose GPU programming. This pioneering work introduced the concept of <strong>stream processing</strong> - treating data as streams that flow through computational kernels.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// BrookGPU example: vector addition using stream processing</span></span>\n<span class=\"line\"><span>kernel </span><span>void</span><span> vector_add</span><span>(</span><span>float</span><span> a</span><span>&lt;&gt;</span><span>,</span><span> float</span><span> b</span><span>&lt;&gt;</span><span>,</span><span> out </span><span>float</span><span> result</span><span>&lt;&gt;</span><span>) {</span></span>\n<span class=\"line\"><span>    result </span><span>=</span><span> a </span><span>+</span><span> b;</span><span>  // Applied to entire streams element-wise</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>void</span><span> main</span><span>() {</span></span>\n<span class=\"line\"><span>    float</span><span> input_stream_a</span><span>&lt;</span><span>1000</span><span>&gt;</span><span>;</span><span>    // Stream of 1000 floats</span></span>\n<span class=\"line\"><span>    float</span><span> input_stream_b</span><span>&lt;</span><span>1000</span><span>&gt;</span><span>;</span><span>    // Stream of 1000 floats</span></span>\n<span class=\"line\"><span>    float</span><span> output_stream</span><span>&lt;</span><span>1000</span><span>&gt;</span><span>;</span><span>     // Output stream</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Initialize streams</span></span>\n<span class=\"line\"><span>    streamRead(input_stream_a</span><span>,</span><span> host_array_a)</span><span>;</span></span>\n<span class=\"line\"><span>    streamRead(input_stream_b</span><span>,</span><span> host_array_b)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Execute kernel on GPU</span></span>\n<span class=\"line\"><span>    vector_add(input_stream_a</span><span>,</span><span> input_stream_b</span><span>,</span><span> output_stream)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Read results back</span></span>\n<span class=\"line\"><span>    streamWrite(output_stream</span><span>,</span><span> host_result_array)</span><span>;</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"2007-cuda-revolution-direct-gpu-programming\">2007: CUDA Revolution - Direct GPU Programming</h3>\n<p>NVIDIA’s introduction of CUDA (Compute Unified Device Architecture) marked the true beginning of mainstream GPGPU computing. For the first time, developers could write GPU code using familiar C/C++ syntax instead of graphics shaders.</p>\n<p><strong>CUDA’s revolutionary features</strong>:</p>\n<ul>\n<li><strong>C/C++ programming model</strong>: No need to encode computation as graphics operations</li>\n<li><strong>Explicit memory management</strong>: Direct control over GPU memory allocation and transfers</li>\n<li><strong>Thread hierarchy</strong>: Organized parallel execution through grids, blocks, and threads</li>\n<li><strong>Shared memory</strong>: Fast on-chip memory for thread cooperation</li>\n<li><strong>Synchronization primitives</strong>: Tools for coordinating parallel threads</li>\n</ul>\n<strong>Code Example: CUDA Matrix Multiplication (Naive Implementation)</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>// CUDA example: Matrix multiplication (naive implementation)</span></span>\n<span class=\"line\"><span>__global__ void matrix_multiply(float* A, float* B, float* C, int N) {</span></span>\n<span class=\"line\"><span>    // Calculate thread's position in the grid</span></span>\n<span class=\"line\"><span>    int row = blockIdx.y * blockDim.y + threadIdx.y;</span></span>\n<span class=\"line\"><span>    int col = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    if (row &lt; N &amp;&amp; col &lt; N) {</span></span>\n<span class=\"line\"><span>        float sum = 0.0f;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Compute dot product of row and column</span></span>\n<span class=\"line\"><span>        for (int k = 0; k &lt; N; k++) {</span></span>\n<span class=\"line\"><span>            sum += A[row * N + k] * B[k * N + col];</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        C[row * N + col] = sum;</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>// Host code</span></span>\n<span class=\"line\"><span>int main() {</span></span>\n<span class=\"line\"><span>    const int N = 1024;</span></span>\n<span class=\"line\"><span>    size_t size = N * N * sizeof(float);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Allocate host memory</span></span>\n<span class=\"line\"><span>    float *h_A, *h_B, *h_C;</span></span>\n<span class=\"line\"><span>    h_A = (float*)malloc(size);</span></span>\n<span class=\"line\"><span>    h_B = (float*)malloc(size);</span></span>\n<span class=\"line\"><span>    h_C = (float*)malloc(size);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Allocate device memory</span></span>\n<span class=\"line\"><span>    float *d_A, *d_B, *d_C;</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_A, size);</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_B, size);</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_C, size);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Initialize matrices (omitted for brevity)</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Copy data to GPU</span></span>\n<span class=\"line\"><span>    cudaMemcpy(d_A, h_A, size, cudaMemcpyHostToDevice);</span></span>\n<span class=\"line\"><span>    cudaMemcpy(d_B, h_B, size, cudaMemcpyHostToDevice);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Launch kernel</span></span>\n<span class=\"line\"><span>    dim3 blockSize(16, 16);</span></span>\n<span class=\"line\"><span>    dim3 gridSize((N + blockSize.x - 1) / blockSize.x,</span></span>\n<span class=\"line\"><span>                  (N + blockSize.y - 1) / blockSize.y);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    matrix_multiply&lt;&lt;&lt;gridSize, blockSize&gt;&gt;&gt;(d_A, d_B, d_C, N);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Copy result back</span></span>\n<span class=\"line\"><span>    cudaMemcpy(h_C, d_C, size, cudaMemcpyDeviceToHost);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Cleanup</span></span>\n<span class=\"line\"><span>    cudaFree(d_A); cudaFree(d_B); cudaFree(d_C);</span></span>\n<span class=\"line\"><span>    free(h_A); free(h_B); free(h_C);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    return 0;</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"2009-opencl-cross-platform-parallel-computing\">2009: OpenCL - Cross-Platform Parallel Computing</h3>\n<p>The Khronos Group introduced OpenCL (Open Computing Language) to provide vendor-neutral parallel computing across different hardware platforms.</p>\n<p><strong>OpenCL advantages</strong>:</p>\n<ul>\n<li><strong>Cross-platform</strong>: Works on NVIDIA, AMD, Intel GPUs, and CPUs</li>\n<li><strong>Heterogeneous computing</strong>: Can coordinate CPU and GPU execution</li>\n<li><strong>Fine-grained control</strong>: Explicit management of compute devices and memory</li>\n</ul>\n<strong>Code Example: OpenCL Vector Addition</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>// OpenCL example: Vector addition</span></span>\n<span class=\"line\"><span>const</span><span> char*</span><span> kernel_source </span><span>=</span><span> R</span><span>\"(</span></span>\n<span class=\"line\"><span>__kernel void vector_add(__global const float* A,</span></span>\n<span class=\"line\"><span>                        __global const float* B,</span></span>\n<span class=\"line\"><span>                        __global float* C) {</span></span>\n<span class=\"line\"><span>    int idx = get_global_id(0);</span></span>\n<span class=\"line\"><span>    C[idx] = A[idx] + B[idx];</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"><span>)\"</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>int</span><span> main</span><span>() {</span></span>\n<span class=\"line\"><span>    const</span><span> int</span><span> ARRAY_SIZE </span><span>=</span><span> 1024</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Get platform and device info</span></span>\n<span class=\"line\"><span>    cl_platform_id platform_id;</span></span>\n<span class=\"line\"><span>    cl_device_id device_id;</span></span>\n<span class=\"line\"><span>    clGetPlatformIDs(</span><span>1</span><span>,</span><span> &amp;</span><span>platform_id</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"><span>    clGetDeviceIDs(platform_id</span><span>,</span><span> CL_DEVICE_TYPE_GPU</span><span>,</span><span> 1</span><span>,</span><span> &amp;</span><span>device_id</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Create context and command queue</span></span>\n<span class=\"line\"><span>    cl_context context </span><span>=</span><span> clCreateContext(</span><span>NULL</span><span>,</span><span> 1</span><span>,</span><span> &amp;</span><span>device_id</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"><span>    cl_command_queue queue </span><span>=</span><span> clCreateCommandQueue(context</span><span>,</span><span> device_id</span><span>,</span><span> 0</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Create and build program</span></span>\n<span class=\"line\"><span>    cl_program program </span><span>=</span><span> clCreateProgramWithSource(context</span><span>,</span><span> 1</span><span>,</span><span> &amp;</span><span>kernel_source</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"><span>    clBuildProgram(program</span><span>,</span><span> 1</span><span>,</span><span> &amp;</span><span>device_id</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Create kernel</span></span>\n<span class=\"line\"><span>    cl_kernel kernel </span><span>=</span><span> clCreateKernel(program</span><span>,</span><span> \"vector_add\"</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Create buffers</span></span>\n<span class=\"line\"><span>    size_t</span><span> size </span><span>=</span><span> ARRAY_SIZE </span><span>*</span><span> sizeof</span><span>(</span><span>float</span><span>);</span></span>\n<span class=\"line\"><span>    cl_mem buffer_A </span><span>=</span><span> clCreateBuffer(context</span><span>,</span><span> CL_MEM_READ_ONLY</span><span>,</span><span> size</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"><span>    cl_mem buffer_B </span><span>=</span><span> clCreateBuffer(context</span><span>,</span><span> CL_MEM_READ_ONLY</span><span>,</span><span> size</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"><span>    cl_mem buffer_C </span><span>=</span><span> clCreateBuffer(context</span><span>,</span><span> CL_MEM_WRITE_ONLY</span><span>,</span><span> size</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Set kernel arguments</span></span>\n<span class=\"line\"><span>    clSetKernelArg(kernel</span><span>,</span><span> 0</span><span>,</span><span> sizeof</span><span>(cl_mem)</span><span>,</span><span> &amp;</span><span>buffer_A)</span><span>;</span></span>\n<span class=\"line\"><span>    clSetKernelArg(kernel</span><span>,</span><span> 1</span><span>,</span><span> sizeof</span><span>(cl_mem)</span><span>,</span><span> &amp;</span><span>buffer_B)</span><span>;</span></span>\n<span class=\"line\"><span>    clSetKernelArg(kernel</span><span>,</span><span> 2</span><span>,</span><span> sizeof</span><span>(cl_mem)</span><span>,</span><span> &amp;</span><span>buffer_C)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Execute kernel</span></span>\n<span class=\"line\"><span>    size_t</span><span> global_work_size </span><span>=</span><span> ARRAY_SIZE;</span></span>\n<span class=\"line\"><span>    clEnqueueNDRangeKernel(queue</span><span>,</span><span> kernel</span><span>,</span><span> 1</span><span>,</span><span> NULL</span><span>,</span><span> &amp;</span><span>global_work_size</span><span>,</span><span> NULL</span><span>,</span><span> 0</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Read result</span></span>\n<span class=\"line\"><span>    float*</span><span> result </span><span>=</span><span> (</span><span>float*</span><span>)</span><span>malloc(size)</span><span>;</span></span>\n<span class=\"line\"><span>    clEnqueueReadBuffer(queue</span><span>,</span><span> buffer_C</span><span>,</span><span> CL_TRUE</span><span>,</span><span> 0</span><span>,</span><span> size</span><span>,</span><span> result</span><span>,</span><span> 0</span><span>,</span><span> NULL</span><span>,</span><span> NULL</span><span>)</span><span>;</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    // Cleanup (omitted for brevity)</span></span>\n<span class=\"line\"><span>    return</span><span> 0</span><span>;</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"2014-cublascudnn-optimized-libraries-for-deep-learning\">2014: cuBLAS/cuDNN - Optimized Libraries for Deep Learning</h3>\n<p>As deep learning gained momentum, NVIDIA recognized the need for highly optimized library implementations of common operations.</p>\n<p><strong>cuBLAS (CUDA Basic Linear Algebra Subprograms)</strong>:</p>\n<ul>\n<li><strong>Optimized BLAS</strong>: Hardware-specific implementations of linear algebra operations</li>\n<li><strong>Multiple precision</strong>: Support for FP64, FP32, FP16, and mixed-precision</li>\n<li><strong>Batched operations</strong>: Efficient processing of multiple small matrices</li>\n</ul>\n<strong>Code Example: cuBLAS Optimized Matrix Multiplication</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>#include &lt;cublas_v2.h&gt;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>void optimized_matrix_multiply() {</span></span>\n<span class=\"line\"><span>    const int M = 2048, N = 2048, K = 2048;</span></span>\n<span class=\"line\"><span>    const float alpha = 1.0f, beta = 0.0f;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Allocate GPU memory</span></span>\n<span class=\"line\"><span>    float *d_A, *d_B, *d_C;</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_A, M * K * sizeof(float));</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_B, K * N * sizeof(float));</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_C, M * N * sizeof(float));</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Create cuBLAS handle</span></span>\n<span class=\"line\"><span>    cublasHandle_t handle;</span></span>\n<span class=\"line\"><span>    cublasCreate(&amp;handle);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Perform matrix multiplication: C = alpha * A * B + beta * C</span></span>\n<span class=\"line\"><span>    // Note: cuBLAS assumes column-major storage</span></span>\n<span class=\"line\"><span>    cublasSgemm(handle, CUBLAS_OP_N, CUBLAS_OP_N,</span></span>\n<span class=\"line\"><span>                N, M, K,  // Dimensions (swapped for column-major)</span></span>\n<span class=\"line\"><span>                &amp;alpha,   // Scaling factor</span></span>\n<span class=\"line\"><span>                d_B, N,   // Matrix B and leading dimension</span></span>\n<span class=\"line\"><span>                d_A, K,   // Matrix A and leading dimension</span></span>\n<span class=\"line\"><span>                &amp;beta,    // Beta scaling factor</span></span>\n<span class=\"line\"><span>                d_C, N);  // Result matrix C and leading dimension</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // This single call is typically 10-100x faster than naive implementation</span></span>\n<span class=\"line\"><span>    // due to:</span></span>\n<span class=\"line\"><span>    // - Optimized memory access patterns</span></span>\n<span class=\"line\"><span>    // - Use of Tensor Cores on supported hardware</span></span>\n<span class=\"line\"><span>    // - Advanced tiling and blocking strategies</span></span>\n<span class=\"line\"><span>    // - Assembly-level optimizations</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    cublasDestroy(handle);</span></span>\n<span class=\"line\"><span>    cudaFree(d_A); cudaFree(d_B); cudaFree(d_C);</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p><strong>cuDNN (CUDA Deep Neural Network library)</strong>:</p>\n<ul>\n<li><strong>Convolution operations</strong>: Highly optimized 2D and 3D convolutions</li>\n<li><strong>Activation functions</strong>: Hardware-accelerated ReLU, sigmoid, tanh</li>\n<li><strong>Normalization</strong>: Batch normalization, layer normalization</li>\n<li><strong>RNN support</strong>: LSTM, GRU implementations</li>\n</ul>\n<strong>Code Example: cuDNN Optimized Convolution</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>#include &lt;cudnn.h&gt;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>void optimized_convolution() {</span></span>\n<span class=\"line\"><span>    cudnnHandle_t cudnn;</span></span>\n<span class=\"line\"><span>    cudnnCreate(&amp;cudnn);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Create tensor descriptors</span></span>\n<span class=\"line\"><span>    cudnnTensorDescriptor_t input_desc, output_desc, bias_desc;</span></span>\n<span class=\"line\"><span>    cudnnFilterDescriptor_t filter_desc;</span></span>\n<span class=\"line\"><span>    cudnnConvolutionDescriptor_t conv_desc;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    cudnnCreateTensorDescriptor(&amp;input_desc);</span></span>\n<span class=\"line\"><span>    cudnnCreateTensorDescriptor(&amp;output_desc);</span></span>\n<span class=\"line\"><span>    cudnnCreateTensorDescriptor(&amp;bias_desc);</span></span>\n<span class=\"line\"><span>    cudnnCreateFilterDescriptor(&amp;filter_desc);</span></span>\n<span class=\"line\"><span>    cudnnCreateConvolutionDescriptor(&amp;conv_desc);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Set descriptors (batch=32, channels=256, height=32, width=32)</span></span>\n<span class=\"line\"><span>    cudnnSetTensorNdDescriptor(input_desc, CUDNN_DATA_FLOAT, 4,</span></span>\n<span class=\"line\"><span>                              (int[]){32, 256, 32, 32},    // dimensions</span></span>\n<span class=\"line\"><span>                              (int[]){262144, 1024, 32, 1}); // strides</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Set filter descriptor (512 output channels, 256 input channels, 3x3 kernel)</span></span>\n<span class=\"line\"><span>    cudnnSetFilterNdDescriptor(filter_desc, CUDNN_DATA_FLOAT, CUDNN_TENSOR_NCHW, 4,</span></span>\n<span class=\"line\"><span>                              (int[]){512, 256, 3, 3});</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Set convolution parameters</span></span>\n<span class=\"line\"><span>    cudnnSetConvolutionNdDescriptor(conv_desc, 2,           // 2D convolution</span></span>\n<span class=\"line\"><span>                                   (int[]){1, 1},           // padding</span></span>\n<span class=\"line\"><span>                                   (int[]){1, 1},           // stride</span></span>\n<span class=\"line\"><span>                                   (int[]){1, 1},           // dilation</span></span>\n<span class=\"line\"><span>                                   CUDNN_CROSS_CORRELATION, // correlation mode</span></span>\n<span class=\"line\"><span>                                   CUDNN_DATA_FLOAT);       // data type</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Find best algorithm</span></span>\n<span class=\"line\"><span>    cudnnConvolutionFwdAlgoPerf_t algo_perf[10];</span></span>\n<span class=\"line\"><span>    int returned_algo_count;</span></span>\n<span class=\"line\"><span>    cudnnFindConvolutionForwardAlgorithm(cudnn, input_desc, filter_desc, conv_desc, output_desc,</span></span>\n<span class=\"line\"><span>                                        10, &amp;returned_algo_count, algo_perf);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // cuDNN automatically selects optimal algorithm based on:</span></span>\n<span class=\"line\"><span>    // - Hardware capabilities (Tensor Cores, memory bandwidth)</span></span>\n<span class=\"line\"><span>    // - Input dimensions and data types</span></span>\n<span class=\"line\"><span>    // - Available memory for workspace</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    cudnnDestroy(cudnn);</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"2020-present-modern-gpu-computing-era\">2020-Present: Modern GPU Computing Era</h3>\n<p>The current era of GPU computing is characterized by:</p>\n<p><strong>Advanced Data Types and Precision</strong>:</p>\n<ul>\n<li><strong>FP8 (8-bit floating point)</strong>: Extreme compression for inference</li>\n<li><strong>BF16 (Brain Float 16)</strong>: Better numerical stability than FP16</li>\n<li><strong>TF32</strong>: Tensor Float 32 for automatic mixed precision</li>\n</ul>\n<p><strong>Sparsity Acceleration</strong>:\nModern GPUs include hardware support for sparse computations, crucial for efficient neural network inference:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>utils</span><span>.</span><span>prune </span><span>as</span><span> prune</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Example: Structured sparsity with 2:4 pattern (2 non-zeros in every 4 elements)</span></span>\n<span class=\"line\"><span>def</span><span> apply_structured_sparsity</span><span>(</span><span>model</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Apply 2:4 structured sparsity pattern\"\"\"</span></span>\n<span class=\"line\"><span>    for</span><span> module </span><span>in</span><span> model</span><span>.</span><span>modules</span><span>():</span></span>\n<span class=\"line\"><span>        if</span><span> isinstance</span><span>(module, torch.nn.Linear):</span></span>\n<span class=\"line\"><span>            # Apply structured pruning that's hardware-accelerated on A100/H100</span></span>\n<span class=\"line\"><span>            prune</span><span>.</span><span>structured</span><span>(module, </span><span>'weight'</span><span>, amount</span><span>=</span><span>0.5</span><span>, dim</span><span>=</span><span>1</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> model</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Modern sparse operations can achieve:</span></span>\n<span class=\"line\"><span># - 2x speedup in compute</span></span>\n<span class=\"line\"><span># - 2x reduction in memory bandwidth</span></span>\n<span class=\"line\"><span># - Minimal accuracy loss when properly trained</span></span></code></pre>\n<p><strong>Attention-Specific Kernels</strong>:\nThe transformer architecture has driven development of specialized kernels:</p>\n<strong>Code Example: Flash Attention Concept (Python)</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>functional </span><span>as</span><span> F</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> flash_attention_concept</span><span>(</span><span>Q</span><span>,</span><span> K</span><span>,</span><span> V</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"</span></span>\n<span class=\"line\"><span>    Conceptual implementation of Flash Attention algorithm</span></span>\n<span class=\"line\"><span>    Real implementation uses custom CUDA kernels for optimal memory usage</span></span>\n<span class=\"line\"><span>    \"\"\"</span></span>\n<span class=\"line\"><span>    B</span><span>,</span><span> H</span><span>,</span><span> N</span><span>,</span><span> D </span><span>=</span><span> Q</span><span>.</span><span>shape  </span><span># Batch, Heads, Sequence, Dimension</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Traditional attention: O(N²) memory usage</span></span>\n<span class=\"line\"><span>    # scores = Q @ K.transpose(-2, -1) / math.sqrt(D)  # N x N matrix</span></span>\n<span class=\"line\"><span>    # attn = F.softmax(scores, dim=-1)</span></span>\n<span class=\"line\"><span>    # output = attn @ V</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Flash Attention: Tiled computation with O(N) memory</span></span>\n<span class=\"line\"><span>    block_size </span><span>=</span><span> 128</span><span>  # Tile size chosen based on SRAM capacity</span></span>\n<span class=\"line\"><span>    output </span><span>=</span><span> torch</span><span>.</span><span>zeros_like</span><span>(Q)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    for</span><span> i </span><span>in</span><span> range</span><span>(</span><span>0</span><span>, N, block_size):</span></span>\n<span class=\"line\"><span>        # Load Q tile into SRAM</span></span>\n<span class=\"line\"><span>        Q_block </span><span>=</span><span> Q</span><span>[:,</span><span> :,</span><span> i</span><span>:</span><span>i</span><span>+</span><span>block_size</span><span>,</span><span> :]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        for</span><span> j </span><span>in</span><span> range</span><span>(</span><span>0</span><span>, N, block_size):</span></span>\n<span class=\"line\"><span>            # Load K,V tile into SRAM</span></span>\n<span class=\"line\"><span>            K_block </span><span>=</span><span> K</span><span>[:,</span><span> :,</span><span> j</span><span>:</span><span>j</span><span>+</span><span>block_size</span><span>,</span><span> :]</span></span>\n<span class=\"line\"><span>            V_block </span><span>=</span><span> V</span><span>[:,</span><span> :,</span><span> j</span><span>:</span><span>j</span><span>+</span><span>block_size</span><span>,</span><span> :]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>            # Compute attention block entirely in SRAM</span></span>\n<span class=\"line\"><span>            scores_block </span><span>=</span><span> Q_block </span><span>@</span><span> K_block</span><span>.</span><span>transpose</span><span>(</span><span>-</span><span>2</span><span>, </span><span>-</span><span>1</span><span>)</span></span>\n<span class=\"line\"><span>            attn_block </span><span>=</span><span> F</span><span>.</span><span>softmax</span><span>(scores_block, dim</span><span>=-</span><span>1</span><span>)</span></span>\n<span class=\"line\"><span>            output_block </span><span>=</span><span> attn_block </span><span>@</span><span> V_block</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>            # Accumulate results (with proper scaling for softmax)</span></span>\n<span class=\"line\"><span>            output</span><span>[:,</span><span> :,</span><span> i</span><span>:</span><span>i</span><span>+</span><span>block_size</span><span>,</span><span> :]</span><span> +=</span><span> output_block</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> output</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Benefits of Flash Attention:</span></span>\n<span class=\"line\"><span># - Reduces memory usage from O(N²) to O(N)</span></span>\n<span class=\"line\"><span># - Enables training on much longer sequences</span></span>\n<span class=\"line\"><span># - 2-4x faster than standard attention on A100/H100</span></span></code></pre>\n<p><strong>Advanced Compiler Stacks</strong>:</p>\n<strong>Code Example: Triton GPU Kernel Programming</strong><p><strong>Triton</strong>: Python-based GPU kernel programming</p><pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> triton</span></span>\n<span class=\"line\"><span>import</span><span> triton</span><span>.</span><span>language </span><span>as</span><span> tl</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>@triton</span><span>.</span><span>jit</span></span>\n<span class=\"line\"><span>def</span><span> vector_add_kernel</span><span>(</span><span>x_ptr</span><span>,</span><span> y_ptr</span><span>,</span><span> output_ptr</span><span>,</span><span> n_elements</span><span>,</span><span> BLOCK_SIZE</span><span>:</span><span> tl</span><span>.</span><span>constexpr):</span></span>\n<span class=\"line\"><span>    \"\"\"High-performance vector addition using Triton\"\"\"</span></span>\n<span class=\"line\"><span>    pid </span><span>=</span><span> tl</span><span>.</span><span>program_id</span><span>(axis</span><span>=</span><span>0</span><span>)</span></span>\n<span class=\"line\"><span>    block_start </span><span>=</span><span> pid </span><span>*</span><span> BLOCK_SIZE</span></span>\n<span class=\"line\"><span>    offsets </span><span>=</span><span> block_start </span><span>+</span><span> tl</span><span>.</span><span>arange</span><span>(</span><span>0</span><span>, BLOCK_SIZE)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Mask for boundary checking</span></span>\n<span class=\"line\"><span>    mask </span><span>=</span><span> offsets </span><span>&lt;</span><span> n_elements</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Load data</span></span>\n<span class=\"line\"><span>    x </span><span>=</span><span> tl</span><span>.</span><span>load</span><span>(x_ptr </span><span>+</span><span> offsets, mask</span><span>=</span><span>mask)</span></span>\n<span class=\"line\"><span>    y </span><span>=</span><span> tl</span><span>.</span><span>load</span><span>(y_ptr </span><span>+</span><span> offsets, mask</span><span>=</span><span>mask)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Compute</span></span>\n<span class=\"line\"><span>    output </span><span>=</span><span> x </span><span>+</span><span> y</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Store result</span></span>\n<span class=\"line\"><span>    tl</span><span>.</span><span>store</span><span>(output_ptr </span><span>+</span><span> offsets, output, mask</span><span>=</span><span>mask)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> vector_add</span><span>(</span><span>x</span><span>,</span><span> y</span><span>):</span></span>\n<span class=\"line\"><span>    output </span><span>=</span><span> torch</span><span>.</span><span>empty_like</span><span>(x)</span></span>\n<span class=\"line\"><span>    n_elements </span><span>=</span><span> output</span><span>.</span><span>numel</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Auto-tuning finds optimal block size</span></span>\n<span class=\"line\"><span>    grid </span><span>=</span><span> lambda</span><span> meta</span><span>: (triton</span><span>.</span><span>cdiv</span><span>(n_elements, meta[</span><span>'BLOCK_SIZE'</span><span>]),</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    vector_add_kernel</span><span>[</span><span>grid</span><span>](x, y, output, n_elements, BLOCK_SIZE</span><span>=</span><span>1024</span><span>)</span></span>\n<span class=\"line\"><span>    return</span><span> output</span></span></code></pre><p><strong>CUDA Graphs</strong>: Reduce kernel launch overhead</p><pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> cuda_graph_optimization</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Demonstrate CUDA Graphs for reducing launch overhead\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Warmup</span></span>\n<span class=\"line\"><span>    x </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>10000</span><span>, device</span><span>=</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"><span>    y </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>10000</span><span>, device</span><span>=</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Capture computation graph</span></span>\n<span class=\"line\"><span>    g </span><span>=</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>CUDAGraph</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Warmup</span></span>\n<span class=\"line\"><span>    s </span><span>=</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>Stream</span><span>()</span></span>\n<span class=\"line\"><span>    s</span><span>.</span><span>wait_stream</span><span>(torch.cuda.</span><span>current_stream</span><span>())</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    with</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>stream</span><span>(s):</span></span>\n<span class=\"line\"><span>        for</span><span> _ </span><span>in</span><span> range</span><span>(</span><span>3</span><span>):</span><span>  # Warmup runs</span></span>\n<span class=\"line\"><span>            z </span><span>=</span><span> x </span><span>+</span><span> y</span></span>\n<span class=\"line\"><span>            z </span><span>=</span><span> z </span><span>*</span><span> 2</span></span>\n<span class=\"line\"><span>            z </span><span>=</span><span> torch</span><span>.</span><span>relu</span><span>(z)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>current_stream</span><span>().</span><span>wait_stream</span><span>(s)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Capture</span></span>\n<span class=\"line\"><span>    with</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>graph</span><span>(g):</span></span>\n<span class=\"line\"><span>        z </span><span>=</span><span> x </span><span>+</span><span> y</span></span>\n<span class=\"line\"><span>        z </span><span>=</span><span> z </span><span>*</span><span> 2</span></span>\n<span class=\"line\"><span>        z </span><span>=</span><span> torch</span><span>.</span><span>relu</span><span>(z)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Replay graph (much faster than individual kernel launches)</span></span>\n<span class=\"line\"><span>    g</span><span>.</span><span>replay</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Benefits:</span></span>\n<span class=\"line\"><span>    # - Eliminates CPU-GPU synchronization overhead</span></span>\n<span class=\"line\"><span>    # - Reduces kernel launch latency from ~10µs to ~1µs</span></span>\n<span class=\"line\"><span>    # - Enables optimization across kernel boundaries</span></span></code></pre>\n<p>This evolution from graphics-centric hardware to general-purpose parallel computing platforms represents one of the most significant developments in modern computer science, enabling the current AI revolution and continuing to drive innovations in scientific computing, machine learning, and high-performance computing applications.</p>\n<hr />\n<h2 id=\"how-does-a-gpu-work-understanding-modern-gpu-architecture\">How Does a GPU Work? Understanding Modern GPU Architecture</h2>\n<h3 id=\"streaming-multiprocessors-the-core-of-gpu-parallelism\">Streaming Multiprocessors: The Core of GPU Parallelism</h3>\n<p>Modern GPUs are built around <strong>Streaming Multiprocessors (SMs)</strong>, which are the fundamental compute units that execute parallel workloads. Each SM contains hundreds of cores, shared memory, registers, and specialized compute units like Tensor Cores.</p>\n<p>The GPU execution hierarchy follows a well-defined structure:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>GPU Device</span></span>\n<span class=\"line\"><span>├─ Streaming Multiprocessor 0 (SM0)</span></span>\n<span class=\"line\"><span>│  ├─ CUDA Cores (FP32/INT32): ~64-128 cores</span></span>\n<span class=\"line\"><span>│  ├─ Tensor Cores: 4-8 units (for AI workloads)</span></span>\n<span class=\"line\"><span>│  ├─ Shared Memory: 48-164 KB (configurable)</span></span>\n<span class=\"line\"><span>│  ├─ Register File: 256 KB (65,536 32-bit registers)</span></span>\n<span class=\"line\"><span>│  ├─ L1 Cache/Texture Cache: 128 KB</span></span>\n<span class=\"line\"><span>│  └─ Warp Schedulers: 4 units</span></span>\n<span class=\"line\"><span>├─ Streaming Multiprocessor 1 (SM1)</span></span>\n<span class=\"line\"><span>│  └─ [Same structure as SM0]</span></span>\n<span class=\"line\"><span>└─ ... (up to 132 SMs in A100)</span></span></code></pre>\n<h3 id=\"thread-hierarchy-threads-warps-blocks-grid\">Thread Hierarchy: Threads → Warps → Blocks → Grid</h3>\n<p>Understanding the GPU thread hierarchy is crucial for writing efficient parallel code:</p>\n<p><strong>1. Thread</strong>: The smallest unit of execution</p>\n<ul>\n<li>Each thread has its own program counter and register space</li>\n<li>Executes the same kernel function with different data</li>\n<li>Can have unique thread ID for data indexing</li>\n</ul>\n<p><strong>2. Warp</strong>: Group of 32 threads executed in lockstep</p>\n<ul>\n<li>All threads in a warp execute the same instruction simultaneously</li>\n<li>Hardware scheduling unit - warps are scheduled for execution</li>\n<li>Divergence occurs when threads take different code paths</li>\n</ul>\n<p><strong>3. Block</strong>: Collection of threads (up to 1024) that can cooperate</p>\n<ul>\n<li>Threads within a block can synchronize using <code>__syncthreads()</code></li>\n<li>Share fast on-chip shared memory</li>\n<li>Entire block is assigned to one SM</li>\n</ul>\n<p><strong>4. Grid</strong>: Collection of blocks that compose the entire kernel launch</p>\n<ul>\n<li>Blocks are independent and can execute in any order</li>\n<li>Grid dimensions define the problem size</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>#include &lt;cuda_runtime.h&gt;</span></span>\n<span class=\"line\"><span>#include &lt;stdio.h&gt;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>__global__ void demonstrate_thread_hierarchy(int* data, int n) {</span></span>\n<span class=\"line\"><span>    // Thread identification within the hierarchy</span></span>\n<span class=\"line\"><span>    int global_thread_id = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span>    int warp_id = global_thread_id / 32;</span></span>\n<span class=\"line\"><span>    int lane_id = global_thread_id % 32;</span></span>\n<span class=\"line\"><span>    int block_id = blockIdx.x;</span></span>\n<span class=\"line\"><span>    int thread_in_block = threadIdx.x;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    if (global_thread_id &lt; n) {</span></span>\n<span class=\"line\"><span>        // Each thread processes its unique data element</span></span>\n<span class=\"line\"><span>        data[global_thread_id] = global_thread_id;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Demonstrate warp-level operations</span></span>\n<span class=\"line\"><span>        // All threads in the same warp see the same warp_id</span></span>\n<span class=\"line\"><span>        if (lane_id == 0) {</span></span>\n<span class=\"line\"><span>            printf(\"Warp %d in Block %d starting execution\\n\", warp_id, block_id);</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Warp-level primitive: ballot (which threads meet condition)</span></span>\n<span class=\"line\"><span>        int mask = __ballot_sync(0xFFFFFFFF, data[global_thread_id] % 2 == 0);</span></span>\n<span class=\"line\"><span>        if (lane_id == 0) {</span></span>\n<span class=\"line\"><span>            printf(\"Warp %d: threads with even values: 0x%08x\\n\", warp_id, mask);</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>void launch_hierarchical_kernel() {</span></span>\n<span class=\"line\"><span>    const int N = 2048;</span></span>\n<span class=\"line\"><span>    int* h_data = new int[N];</span></span>\n<span class=\"line\"><span>    int* d_data;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_data, N * sizeof(int));</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Configure kernel launch parameters</span></span>\n<span class=\"line\"><span>    int threads_per_block = 256;  // 8 warps per block</span></span>\n<span class=\"line\"><span>    int blocks_per_grid = (N + threads_per_block - 1) / threads_per_block;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    printf(\"Launching kernel with:\\n\");</span></span>\n<span class=\"line\"><span>    printf(\"  Threads per block: %d\\n\", threads_per_block);</span></span>\n<span class=\"line\"><span>    printf(\"  Blocks per grid: %d\\n\", blocks_per_grid);</span></span>\n<span class=\"line\"><span>    printf(\"  Total threads: %d\\n\", blocks_per_grid * threads_per_block);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    demonstrate_thread_hierarchy&lt;&lt;&lt;blocks_per_grid, threads_per_block&gt;&gt;&gt;(d_data, N);</span></span>\n<span class=\"line\"><span>    cudaDeviceSynchronize();</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    cudaMemcpy(h_data, d_data, N * sizeof(int), cudaMemcpyDeviceToHost);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    delete[] h_data;</span></span>\n<span class=\"line\"><span>    cudaFree(d_data);</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"kernel-execution-functions-launched-over-a-grid-of-blocks\">Kernel Execution: Functions Launched Over a Grid of Blocks</h3>\n<p>A <strong>kernel</strong> is a function that runs on the GPU and is executed by thousands of threads in parallel. Unlike CPU functions that execute once, kernels are launched over a grid of thread blocks:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>// Kernel declaration with __global__ qualifier</span></span>\n<span class=\"line\"><span>__global__ void matrix_vector_multiply(float* matrix, float* vector, float* result, int rows, int cols) {</span></span>\n<span class=\"line\"><span>    // Each thread computes one element of the result vector</span></span>\n<span class=\"line\"><span>    int row = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    if (row &lt; rows) {</span></span>\n<span class=\"line\"><span>        float sum = 0.0f;</span></span>\n<span class=\"line\"><span>        for (int col = 0; col &lt; cols; col++) {</span></span>\n<span class=\"line\"><span>            sum += matrix[row * cols + col] * vector[col];</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span>        result[row] = sum;</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>void launch_matrix_vector_kernel() {</span></span>\n<span class=\"line\"><span>    const int ROWS = 4096, COLS = 4096;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Host memory allocation</span></span>\n<span class=\"line\"><span>    float* h_matrix = new float[ROWS * COLS];</span></span>\n<span class=\"line\"><span>    float* h_vector = new float[COLS];</span></span>\n<span class=\"line\"><span>    float* h_result = new float[ROWS];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Initialize data</span></span>\n<span class=\"line\"><span>    for (int i = 0; i &lt; ROWS * COLS; i++) h_matrix[i] = (float)rand() / RAND_MAX;</span></span>\n<span class=\"line\"><span>    for (int i = 0; i &lt; COLS; i++) h_vector[i] = (float)rand() / RAND_MAX;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Device memory allocation</span></span>\n<span class=\"line\"><span>    float *d_matrix, *d_vector, *d_result;</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_matrix, ROWS * COLS * sizeof(float));</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_vector, COLS * sizeof(float));</span></span>\n<span class=\"line\"><span>    cudaMalloc(&amp;d_result, ROWS * sizeof(float));</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Data transfer to device</span></span>\n<span class=\"line\"><span>    cudaMemcpy(d_matrix, h_matrix, ROWS * COLS * sizeof(float), cudaMemcpyHostToDevice);</span></span>\n<span class=\"line\"><span>    cudaMemcpy(d_vector, h_vector, COLS * sizeof(float), cudaMemcpyHostToDevice);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Kernel launch configuration</span></span>\n<span class=\"line\"><span>    int block_size = 256;</span></span>\n<span class=\"line\"><span>    int grid_size = (ROWS + block_size - 1) / block_size;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Launch kernel over grid of blocks</span></span>\n<span class=\"line\"><span>    matrix_vector_multiply&lt;&lt;&lt;grid_size, block_size&gt;&gt;&gt;(d_matrix, d_vector, d_result, ROWS, COLS);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Copy result back</span></span>\n<span class=\"line\"><span>    cudaMemcpy(h_result, d_result, ROWS * sizeof(float), cudaMemcpyDeviceToHost);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Cleanup</span></span>\n<span class=\"line\"><span>    delete[] h_matrix; delete[] h_vector; delete[] h_result;</span></span>\n<span class=\"line\"><span>    cudaFree(d_matrix); cudaFree(d_vector); cudaFree(d_result);</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"thread-cooperation-via-shared-memory\">Thread Cooperation via Shared Memory</h3>\n<p>Threads within the same block can cooperate through <strong>shared memory</strong>, a fast on-chip memory space that enables data sharing and synchronization:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>__global__ void cooperative_reduction(float* input, float* output, int n) {</span></span>\n<span class=\"line\"><span>    // Shared memory allocation - shared by all threads in the block</span></span>\n<span class=\"line\"><span>    __shared__ float shared_data[256];  // Assuming block size of 256</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    int tid = threadIdx.x;</span></span>\n<span class=\"line\"><span>    int global_id = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Load data into shared memory</span></span>\n<span class=\"line\"><span>    if (global_id &lt; n) {</span></span>\n<span class=\"line\"><span>        shared_data[tid] = input[global_id];</span></span>\n<span class=\"line\"><span>    } else {</span></span>\n<span class=\"line\"><span>        shared_data[tid] = 0.0f;</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Synchronize to ensure all threads have loaded their data</span></span>\n<span class=\"line\"><span>    __syncthreads();</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Parallel reduction in shared memory</span></span>\n<span class=\"line\"><span>    for (int stride = blockDim.x / 2; stride &gt; 0; stride &gt;&gt;= 1) {</span></span>\n<span class=\"line\"><span>        if (tid &lt; stride) {</span></span>\n<span class=\"line\"><span>            shared_data[tid] += shared_data[tid + stride];</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span>        __syncthreads();  // Synchronize after each reduction step</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Thread 0 writes the block's result</span></span>\n<span class=\"line\"><span>    if (tid == 0) {</span></span>\n<span class=\"line\"><span>        output[blockIdx.x] = shared_data[0];</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<hr />\n<h2 id=\"gpu-memory-architecture-the-foundation-of-performance\">GPU Memory Architecture: The Foundation of Performance</h2>\n<h3 id=\"global-memory-hbmvram-the-primary-storage\">Global Memory (HBM/VRAM): The Primary Storage</h3>\n<p><strong>High Bandwidth Memory (HBM)</strong> or <strong>VRAM</strong> serves as the main memory for GPU computations:</p>\n<p><strong>Characteristics</strong>:</p>\n<ul>\n<li><strong>Capacity</strong>: 40-80GB on modern data center GPUs (A100, H100)</li>\n<li><strong>Bandwidth</strong>: 1.5-3.0 TB/s peak throughput</li>\n<li><strong>Latency</strong>: 200-800 cycles (relatively high)</li>\n<li><strong>Access Pattern Sensitivity</strong>: Coalesced access patterns crucial for performance</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> demonstrate_memory_coalescing</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Show the importance of memory access patterns\"\"\"</span></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span> if</span><span> torch.cuda.</span><span>is_available</span><span>() </span><span>else</span><span> 'cpu'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Large matrix to stress memory system</span></span>\n<span class=\"line\"><span>    N </span><span>=</span><span> 4096</span></span>\n<span class=\"line\"><span>    matrix </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(N, N, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Coalesced access: consecutive threads access consecutive memory</span></span>\n<span class=\"line\"><span>    def</span><span> coalesced_access</span><span>(</span><span>mat</span><span>):</span></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Row-major access - coalesced for C-style arrays</span></span>\n<span class=\"line\"><span>        result </span><span>=</span><span> mat</span><span>.</span><span>sum</span><span>(dim</span><span>=</span><span>1</span><span>)</span><span>  # Sum along rows</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        return</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Non-coalesced access: strided memory access</span></span>\n<span class=\"line\"><span>    def</span><span> strided_access</span><span>(</span><span>mat</span><span>):</span></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Column-major access - non-coalesced, poor cache usage</span></span>\n<span class=\"line\"><span>        result </span><span>=</span><span> mat</span><span>.</span><span>sum</span><span>(dim</span><span>=</span><span>0</span><span>)</span><span>  # Sum along columns</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        return</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Memory bandwidth test</span></span>\n<span class=\"line\"><span>    def</span><span> measure_bandwidth</span><span>():</span></span>\n<span class=\"line\"><span>        size_mb </span><span>=</span><span> 1000</span></span>\n<span class=\"line\"><span>        elements </span><span>=</span><span> (size_mb </span><span>*</span><span> 1024</span><span> *</span><span> 1024</span><span>) </span><span>//</span><span> 4</span><span>  # 4 bytes per float32</span></span>\n<span class=\"line\"><span>        data </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(elements, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Simple copy operation to measure raw bandwidth</span></span>\n<span class=\"line\"><span>        copy </span><span>=</span><span> data</span><span>.</span><span>clone</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        elapsed </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Calculate bandwidth (read + write)</span></span>\n<span class=\"line\"><span>        bytes_transferred </span><span>=</span><span> elements </span><span>*</span><span> 4</span><span> *</span><span> 2</span><span>  # read + write</span></span>\n<span class=\"line\"><span>        bandwidth_gb_s </span><span>=</span><span> bytes_transferred </span><span>/</span><span> elapsed </span><span>/</span><span> (</span><span>1024</span><span>**</span><span>3</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> bandwidth_gb_s</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    coalesced_time </span><span>=</span><span> coalesced_access</span><span>(matrix)</span></span>\n<span class=\"line\"><span>    strided_time </span><span>=</span><span> strided_access</span><span>(matrix)</span></span>\n<span class=\"line\"><span>    bandwidth </span><span>=</span><span> measure_bandwidth</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Coalesced access time: </span><span>{</span><span>coalesced_time</span><span>:.4f</span><span>}</span><span>s\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Strided access time: </span><span>{</span><span>strided_time</span><span>:.4f</span><span>}</span><span>s\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Performance ratio: </span><span>{</span><span>strided_time</span><span>/</span><span>coalesced_time</span><span>:.2f</span><span>}</span><span>x slower for strided\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Measured memory bandwidth: </span><span>{</span><span>bandwidth</span><span>:.1f</span><span>}</span><span> GB/s\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>demonstrate_memory_coalescing</span><span>()</span></span></code></pre>\n<h3 id=\"l2l1-caches-hardware-managed-performance\">L2/L1 Caches: Hardware-Managed Performance</h3>\n<p>Modern GPUs implement a sophisticated cache hierarchy to reduce memory latency:</p>\n<p><strong>L2 Cache</strong>:</p>\n<ul>\n<li><strong>Size</strong>: 40-50MB on A100/H100</li>\n<li><strong>Shared</strong>: Across all SMs on the GPU</li>\n<li><strong>Management</strong>: Hardware-managed, transparent to programmer</li>\n<li><strong>Purpose</strong>: Reduce Global Memory access latency</li>\n</ul>\n<p><strong>L1 Cache/Texture Cache</strong>:</p>\n<ul>\n<li><strong>Size</strong>: 128KB per SM</li>\n<li><strong>Management</strong>: Hardware-managed with some programmer control</li>\n<li><strong>Configurable</strong>: Can be partitioned with shared memory</li>\n<li><strong>Optimization</strong>: Benefits from spatial and temporal locality</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>__global__ void cache_friendly_stencil(float* input, float* output, int width, int height) {</span></span>\n<span class=\"line\"><span>    int x = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span>    int y = blockIdx.y * blockDim.y + threadIdx.y;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    if (x &gt; 0 &amp;&amp; x &lt; width-1 &amp;&amp; y &gt; 0 &amp;&amp; y &lt; height-1) {</span></span>\n<span class=\"line\"><span>        int idx = y * width + x;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // 5-point stencil operation - benefits from L1 cache due to spatial locality</span></span>\n<span class=\"line\"><span>        float center = input[idx];</span></span>\n<span class=\"line\"><span>        float left   = input[idx - 1];</span></span>\n<span class=\"line\"><span>        float right  = input[idx + 1];</span></span>\n<span class=\"line\"><span>        float up     = input[idx - width];</span></span>\n<span class=\"line\"><span>        float down   = input[idx + width];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Weighted average - neighboring elements likely in L1 cache</span></span>\n<span class=\"line\"><span>        output[idx] = 0.2f * (center + left + right + up + down);</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"shared-memory-sram-programmer-managed-performance\">Shared Memory (SRAM): Programmer-Managed Performance</h3>\n<p><strong>Shared Memory</strong> is the programmer’s primary tool for optimizing memory access patterns:</p>\n<p><strong>Characteristics</strong>:</p>\n<ul>\n<li><strong>Size</strong>: 48-164KB per SM (configurable)</li>\n<li><strong>Speed</strong>: ~19 TB/s per SM (much faster than global memory)</li>\n<li><strong>Scope</strong>: Shared among threads in the same block</li>\n<li><strong>Management</strong>: Explicitly managed by programmer</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>#define TILE_SIZE 16</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>__global__ void tiled_matrix_multiply(float* A, float* B, float* C, int N) {</span></span>\n<span class=\"line\"><span>    // Shared memory tiles - allocated per thread block</span></span>\n<span class=\"line\"><span>    __shared__ float tile_A[TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span>    __shared__ float tile_B[TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    int row = blockIdx.y * blockDim.y + threadIdx.y;</span></span>\n<span class=\"line\"><span>    int col = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    float sum = 0.0f;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Loop over tiles</span></span>\n<span class=\"line\"><span>    for (int tile = 0; tile &lt; (N + TILE_SIZE - 1) / TILE_SIZE; tile++) {</span></span>\n<span class=\"line\"><span>        // Collaborative loading: each thread loads one element</span></span>\n<span class=\"line\"><span>        int tile_row = tile * TILE_SIZE + threadIdx.x;</span></span>\n<span class=\"line\"><span>        int tile_col = tile * TILE_SIZE + threadIdx.y;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Load tile of A</span></span>\n<span class=\"line\"><span>        if (row &lt; N &amp;&amp; tile_row &lt; N) {</span></span>\n<span class=\"line\"><span>            tile_A[threadIdx.y][threadIdx.x] = A[row * N + tile_row];</span></span>\n<span class=\"line\"><span>        } else {</span></span>\n<span class=\"line\"><span>            tile_A[threadIdx.y][threadIdx.x] = 0.0f;</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Load tile of B</span></span>\n<span class=\"line\"><span>        if (tile_col &lt; N &amp;&amp; col &lt; N) {</span></span>\n<span class=\"line\"><span>            tile_B[threadIdx.y][threadIdx.x] = B[tile_col * N + col];</span></span>\n<span class=\"line\"><span>        } else {</span></span>\n<span class=\"line\"><span>            tile_B[threadIdx.y][threadIdx.x] = 0.0f;</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Synchronize to ensure tile loading is complete</span></span>\n<span class=\"line\"><span>        __syncthreads();</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Compute using data in shared memory (much faster than global memory)</span></span>\n<span class=\"line\"><span>        for (int k = 0; k &lt; TILE_SIZE; k++) {</span></span>\n<span class=\"line\"><span>            sum += tile_A[threadIdx.y][k] * tile_B[k][threadIdx.x];</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Synchronize before loading next tile</span></span>\n<span class=\"line\"><span>        __syncthreads();</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Write result</span></span>\n<span class=\"line\"><span>    if (row &lt; N &amp;&amp; col &lt; N) {</span></span>\n<span class=\"line\"><span>        C[row * N + col] = sum;</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<h3 id=\"registers-per-thread-fastest-access\">Registers: Per-Thread Fastest Access</h3>\n<p><strong>Registers</strong> provide the fastest memory access but are limited in quantity:</p>\n<p><strong>Characteristics</strong>:</p>\n<ul>\n<li><strong>Size</strong>: 256KB per SM (65,536 32-bit registers)</li>\n<li><strong>Distribution</strong>: Divided among active threads</li>\n<li><strong>Speed</strong>: Single-cycle access</li>\n<li><strong>Limitation</strong>: Too many registers per thread reduces occupancy</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>__global__ void register_usage_example(float* data, int n) {</span></span>\n<span class=\"line\"><span>    int tid = blockIdx.x * blockDim.x + threadIdx.x;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    if (tid &lt; n) {</span></span>\n<span class=\"line\"><span>        // These variables are stored in registers (fastest access)</span></span>\n<span class=\"line\"><span>        float reg1 = data[tid];</span></span>\n<span class=\"line\"><span>        float reg2 = reg1 * 2.0f;</span></span>\n<span class=\"line\"><span>        float reg3 = reg2 + 1.0f;</span></span>\n<span class=\"line\"><span>        float reg4 = reg3 / 3.0f;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Complex computation using registers</span></span>\n<span class=\"line\"><span>        for (int i = 0; i &lt; 10; i++) {</span></span>\n<span class=\"line\"><span>            reg1 = sinf(reg1) + cosf(reg2);</span></span>\n<span class=\"line\"><span>            reg2 = sqrtf(reg3) * reg4;</span></span>\n<span class=\"line\"><span>            reg3 = reg1 + reg2;</span></span>\n<span class=\"line\"><span>            reg4 = reg3 - reg1;</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        data[tid] = reg4;</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<hr />\n<h2 id=\"performance-analysis-understanding-gpu-bottlenecks\">Performance Analysis: Understanding GPU Bottlenecks</h2>\n<h3 id=\"memory-bound-vs-compute-bound-kernels\">Memory-Bound vs Compute-Bound Kernels</h3>\n<p>Understanding whether your kernel is <strong>memory-bound</strong> or <strong>compute-bound</strong> is crucial for optimization:</p>\n<p><strong>Memory-Bound Kernels</strong>:</p>\n<ul>\n<li>Runtime dominated by data transfer between Global Memory and SMs</li>\n<li>Low arithmetic intensity (few operations per byte transferred)</li>\n<li>Examples: Simple element-wise operations, reductions, transposes</li>\n</ul>\n<p><strong>Compute-Bound Kernels</strong>:</p>\n<ul>\n<li>Runtime dominated by actual computation (ALU or Tensor Core operations)</li>\n<li>High arithmetic intensity (many operations per byte transferred)</li>\n<li>Examples: Dense matrix multiplication, complex mathematical functions</li>\n</ul>\n<h3 id=\"arithmetic-intensity-the-key-metric\">Arithmetic Intensity: The Key Metric</h3>\n<p><strong>Arithmetic Intensity (AI) = FLOPs / Bytes transferred from HBM</strong></p>\n<p>This metric determines the performance ceiling of your kernel:</p>\n<strong>Code Example: Arithmetic Intensity Analysis</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> analyze_arithmetic_intensity</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Demonstrate the concept of arithmetic intensity\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"><span>    N </span><span>=</span><span> 4096</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Example 1: Low arithmetic intensity (memory-bound)</span></span>\n<span class=\"line\"><span>    def</span><span> vector_add_analysis</span><span>():</span></span>\n<span class=\"line\"><span>        a </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(N, N, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"><span>        b </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(N, N, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>        c </span><span>=</span><span> a </span><span>+</span><span> b  </span><span># 1 FLOP per element</span></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        elapsed </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        elements </span><span>=</span><span> N </span><span>*</span><span> N</span></span>\n<span class=\"line\"><span>        flops </span><span>=</span><span> elements  </span><span># 1 FLOP per element</span></span>\n<span class=\"line\"><span>        bytes_transferred </span><span>=</span><span> elements </span><span>*</span><span> 4</span><span> *</span><span> 3</span><span>  # 3 arrays × 4 bytes per float32</span></span>\n<span class=\"line\"><span>        arithmetic_intensity </span><span>=</span><span> flops </span><span>/</span><span> bytes_transferred</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"Vector Addition:\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Arithmetic Intensity: </span><span>{</span><span>arithmetic_intensity</span><span>:.3f</span><span>}</span><span> FLOPs/Byte\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Time: </span><span>{</span><span>elapsed</span><span>:.4f</span><span>}</span><span>s\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Performance: </span><span>{</span><span>flops </span><span>/</span><span> elapsed </span><span>/</span><span> 1e9</span><span>:.2f</span><span>}</span><span> GFLOPS\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> arithmetic_intensity</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Example 2: High arithmetic intensity (compute-bound)</span></span>\n<span class=\"line\"><span>    def</span><span> matrix_multiply_analysis</span><span>():</span></span>\n<span class=\"line\"><span>        a </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(N, N, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"><span>        b </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(N, N, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>        c </span><span>=</span><span> torch</span><span>.</span><span>matmul</span><span>(a, b)</span><span>  # N³ FLOPs for N×N matrices</span></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        elapsed </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        flops </span><span>=</span><span> 2</span><span> *</span><span> N </span><span>*</span><span> N </span><span>*</span><span> N  </span><span># N³ multiply-add operations</span></span>\n<span class=\"line\"><span>        bytes_transferred </span><span>=</span><span> N </span><span>*</span><span> N </span><span>*</span><span> 4</span><span> *</span><span> 3</span><span>  # 3 matrices × 4 bytes per float32</span></span>\n<span class=\"line\"><span>        arithmetic_intensity </span><span>=</span><span> flops </span><span>/</span><span> bytes_transferred</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"</span><span>\\n</span><span>Matrix Multiplication:\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Arithmetic Intensity: </span><span>{</span><span>arithmetic_intensity</span><span>:.3f</span><span>}</span><span> FLOPs/Byte\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Time: </span><span>{</span><span>elapsed</span><span>:.4f</span><span>}</span><span>s\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Performance: </span><span>{</span><span>flops </span><span>/</span><span> elapsed </span><span>/</span><span> 1e12</span><span>:.2f</span><span>}</span><span> TFLOPS\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> arithmetic_intensity</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Example 3: Medium arithmetic intensity</span></span>\n<span class=\"line\"><span>    def</span><span> convolution_analysis</span><span>():</span></span>\n<span class=\"line\"><span>        # 3D convolution: batch=32, channels=256, spatial=64x64, kernel=3x3</span></span>\n<span class=\"line\"><span>        input_tensor </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>32</span><span>, </span><span>256</span><span>, </span><span>64</span><span>, </span><span>64</span><span>, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"><span>        kernel </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>512</span><span>, </span><span>256</span><span>, </span><span>3</span><span>, </span><span>3</span><span>, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>        output </span><span>=</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>functional</span><span>.</span><span>conv2d</span><span>(input_tensor, kernel, padding</span><span>=</span><span>1</span><span>)</span></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        elapsed </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Approximate FLOP count for convolution</span></span>\n<span class=\"line\"><span>        batch_size</span><span>,</span><span> in_channels</span><span>,</span><span> height</span><span>,</span><span> width </span><span>=</span><span> input_tensor</span><span>.</span><span>shape</span></span>\n<span class=\"line\"><span>        out_channels</span><span>,</span><span> _</span><span>,</span><span> kernel_h</span><span>,</span><span> kernel_w </span><span>=</span><span> kernel</span><span>.</span><span>shape</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        flops </span><span>=</span><span> batch_size </span><span>*</span><span> out_channels </span><span>*</span><span> height </span><span>*</span><span> width </span><span>*</span><span> in_channels </span><span>*</span><span> kernel_h </span><span>*</span><span> kernel_w </span><span>*</span><span> 2</span></span>\n<span class=\"line\"><span>        input_bytes </span><span>=</span><span> input_tensor</span><span>.</span><span>numel</span><span>()</span><span> *</span><span> 4</span></span>\n<span class=\"line\"><span>        kernel_bytes </span><span>=</span><span> kernel</span><span>.</span><span>numel</span><span>()</span><span> *</span><span> 4</span></span>\n<span class=\"line\"><span>        output_bytes </span><span>=</span><span> output</span><span>.</span><span>numel</span><span>()</span><span> *</span><span> 4</span></span>\n<span class=\"line\"><span>        bytes_transferred </span><span>=</span><span> input_bytes </span><span>+</span><span> kernel_bytes </span><span>+</span><span> output_bytes</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        arithmetic_intensity </span><span>=</span><span> flops </span><span>/</span><span> bytes_transferred</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"</span><span>\\n</span><span>Convolution:\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Arithmetic Intensity: </span><span>{</span><span>arithmetic_intensity</span><span>:.3f</span><span>}</span><span> FLOPs/Byte\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Time: </span><span>{</span><span>elapsed</span><span>:.4f</span><span>}</span><span>s\"</span><span>)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"  Performance: </span><span>{</span><span>flops </span><span>/</span><span> elapsed </span><span>/</span><span> 1e12</span><span>:.2f</span><span>}</span><span> TFLOPS\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> arithmetic_intensity</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    ai_low </span><span>=</span><span> vector_add_analysis</span><span>()</span></span>\n<span class=\"line\"><span>    ai_high </span><span>=</span><span> matrix_multiply_analysis</span><span>()</span></span>\n<span class=\"line\"><span>    ai_medium </span><span>=</span><span> convolution_analysis</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> ai_low</span><span>,</span><span> ai_medium</span><span>,</span><span> ai_high</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>analyze_arithmetic_intensity</span><span>()</span></span></code></pre>\n<h3 id=\"a100-performance-analysis-understanding-theoretical-limits\">A100 Performance Analysis: Understanding Theoretical Limits</h3>\n<p>Let’s analyze the NVIDIA A100’s performance characteristics:</p>\n<strong>Code Example: A100 Performance Model</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> a100_performance_model</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Model A100 performance characteristics\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # A100 specifications</span></span>\n<span class=\"line\"><span>    specs </span><span>=</span><span> {</span></span>\n<span class=\"line\"><span>        'peak_fp32_compute'</span><span>:</span><span> 19.5</span><span>,</span><span>      # TFLOPS</span></span>\n<span class=\"line\"><span>        'peak_hbm_bandwidth'</span><span>:</span><span> 1.5</span><span>,</span><span>      # TB/s</span></span>\n<span class=\"line\"><span>        'hbm_capacity'</span><span>:</span><span> 40</span><span>,</span><span>             # GB (base model)</span></span>\n<span class=\"line\"><span>        'sm_count'</span><span>:</span><span> 108</span><span>,</span></span>\n<span class=\"line\"><span>        'tensor_core_fp16'</span><span>:</span><span> 312</span><span>,</span><span>        # TFLOPS with sparsity</span></span>\n<span class=\"line\"><span>        'l2_cache'</span><span>:</span><span> 40</span><span>,</span><span>                 # MB</span></span>\n<span class=\"line\"><span>        'shared_memory_per_sm'</span><span>:</span><span> 164</span><span>,</span><span>    # KB</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Calculate ridge point (arithmetic intensity threshold)</span></span>\n<span class=\"line\"><span>    ridge_point </span><span>=</span><span> specs</span><span>[</span><span>'peak_fp32_compute'</span><span>]</span><span> /</span><span> specs</span><span>[</span><span>'peak_hbm_bandwidth'</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"NVIDIA A100 Performance Analysis:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Peak FP32 Compute: </span><span>{</span><span>specs[</span><span>'peak_fp32_compute'</span><span>]</span><span>}</span><span> TFLOPS\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Peak HBM Bandwidth: </span><span>{</span><span>specs[</span><span>'peak_hbm_bandwidth'</span><span>]</span><span>}</span><span> TB/s\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Ridge Point: </span><span>{</span><span>ridge_point</span><span>:.1f</span><span>}</span><span> FLOPs/Byte\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"L2 Cache: </span><span>{</span><span>specs[</span><span>'l2_cache'</span><span>]</span><span>}</span><span> MB\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Shared Memory per SM: </span><span>{</span><span>specs[</span><span>'shared_memory_per_sm'</span><span>]</span><span>}</span><span> KB\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"\\nPerformance Predictions:\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Analyze different operation types</span></span>\n<span class=\"line\"><span>    operations </span><span>=</span><span> [</span></span>\n<span class=\"line\"><span>        (</span><span>\"Vector Add\"</span><span>,</span><span> 1</span><span>,</span><span> \"Memory-bound\"</span><span>)</span><span>,</span></span>\n<span class=\"line\"><span>        (</span><span>\"Convolution (typical)\"</span><span>,</span><span> 50</span><span>,</span><span> \"Balanced\"</span><span>)</span><span>,</span></span>\n<span class=\"line\"><span>        (</span><span>\"Dense GEMM\"</span><span>,</span><span> 100</span><span>,</span><span> \"Compute-bound\"</span><span>)</span><span>,</span></span>\n<span class=\"line\"><span>        (</span><span>\"Sparse GEMM (2:4)\"</span><span>,</span><span> 200</span><span>,</span><span> \"Highly compute-bound\"</span><span>)</span></span>\n<span class=\"line\"><span>    ]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    for</span><span> op_name</span><span>,</span><span> ai</span><span>,</span><span> classification </span><span>in</span><span> operations</span><span>:</span></span>\n<span class=\"line\"><span>        if</span><span> ai </span><span>&lt;</span><span> ridge_point</span><span>:</span></span>\n<span class=\"line\"><span>            max_performance </span><span>=</span><span> ai </span><span>*</span><span> specs</span><span>[</span><span>'peak_hbm_bandwidth'</span><span>]</span></span>\n<span class=\"line\"><span>            bottleneck </span><span>=</span><span> \"Memory bandwidth\"</span></span>\n<span class=\"line\"><span>        else</span><span>:</span></span>\n<span class=\"line\"><span>            max_performance </span><span>=</span><span> specs</span><span>[</span><span>'peak_fp32_compute'</span><span>]</span></span>\n<span class=\"line\"><span>            bottleneck </span><span>=</span><span> \"Compute throughput\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"</span><span>{</span><span>op_name</span><span>:20</span><span>}</span><span> | AI: </span><span>{</span><span>ai</span><span>:3.0f</span><span>}</span><span> | Max: </span><span>{</span><span>max_performance</span><span>:5.1f</span><span>}</span><span> TFLOPS | </span><span>{</span><span>bottleneck</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> specs</span><span>,</span><span> ridge_point</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>specs</span><span>,</span><span> ridge_point </span><span>=</span><span> a100_performance_model</span><span>()</span></span></code></pre>\n<h3 id=\"overhead-bound-operations\">Overhead-Bound Operations</h3>\n<p>Many applications suffer from <strong>overhead-bound</strong> performance, where kernel launch and dispatch overhead dominates:</p>\n<strong>Code Example: Kernel Launch Overhead Analysis</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> demonstrate_overhead_bound</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Show how small kernels can be overhead-bound\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    def</span><span> measure_kernel_overhead</span><span>(</span><span>size</span><span>,</span><span> num_launches</span><span>):</span></span>\n<span class=\"line\"><span>        \"\"\"Measure overhead for small vs large kernels\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        a </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(size, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"><span>        b </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(size, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Warm up</span></span>\n<span class=\"line\"><span>        for</span><span> _ </span><span>in</span><span> range</span><span>(</span><span>10</span><span>):</span></span>\n<span class=\"line\"><span>            c </span><span>=</span><span> a </span><span>+</span><span> b</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        for</span><span> _ </span><span>in</span><span> range</span><span>(num_launches):</span></span>\n<span class=\"line\"><span>            c </span><span>=</span><span> a </span><span>+</span><span> b</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        elapsed </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        time_per_kernel </span><span>=</span><span> elapsed </span><span>/</span><span> num_launches </span><span>*</span><span> 1000</span><span>  # milliseconds</span></span>\n<span class=\"line\"><span>        elements_per_second </span><span>=</span><span> size </span><span>*</span><span> num_launches </span><span>/</span><span> elapsed </span><span>/</span><span> 1e9</span><span>  # billion elements/sec</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> time_per_kernel</span><span>,</span><span> elements_per_second</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"Kernel Launch Overhead Analysis:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"Size\\t\\tTime/Kernel\\tBandwidth\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"-\"</span><span> *</span><span> 40</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    sizes </span><span>=</span><span> [</span><span>100</span><span>,</span><span> 1000</span><span>,</span><span> 10000</span><span>,</span><span> 100000</span><span>,</span><span> 1000000</span><span>,</span><span> 10000000</span><span>]</span></span>\n<span class=\"line\"><span>    num_launches </span><span>=</span><span> 1000</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    for</span><span> size </span><span>in</span><span> sizes</span><span>:</span></span>\n<span class=\"line\"><span>        time_per_kernel</span><span>,</span><span> throughput </span><span>=</span><span> measure_kernel_overhead</span><span>(size, num_launches)</span></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"</span><span>{</span><span>size</span><span>:8d</span><span>}</span><span>\\t</span><span>{</span><span>time_per_kernel</span><span>:8.3f</span><span>}</span><span>ms</span><span>\\t</span><span>{</span><span>throughput</span><span>:8.2f</span><span>}</span><span> GEl/s\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"\\nObservations:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"- Small kernels show high overhead (low throughput)\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"- Large kernels amortize launch overhead\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"- Async execution can help hide overhead\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>demonstrate_overhead_bound</span><span>()</span></span></code></pre>\n<hr />\n<h2 id=\"performance-optimization-strategies\">Performance Optimization Strategies</h2>\n<h3 id=\"strategy-1-operator-fusion-for-memory-bound-operations\">Strategy 1: Operator Fusion for Memory-Bound Operations</h3>\n<p><strong>Operator Fusion</strong> combines multiple operations to reduce memory traffic and kernel launch overhead:</p>\n<strong>Code Example: Operator Fusion with Triton</strong><pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> triton</span></span>\n<span class=\"line\"><span>import</span><span> triton</span><span>.</span><span>language </span><span>as</span><span> tl</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Traditional unfused operations</span></span>\n<span class=\"line\"><span>def</span><span> unfused_operations</span><span>(</span><span>x</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Separate kernels for each operation\"\"\"</span></span>\n<span class=\"line\"><span>    temp1 </span><span>=</span><span> x </span><span>+</span><span> 1.0</span><span>      # Kernel 1: load x, store temp1</span></span>\n<span class=\"line\"><span>    temp2 </span><span>=</span><span> torch</span><span>.</span><span>relu</span><span>(temp1)</span><span>  # Kernel 2: load temp1, store temp2</span></span>\n<span class=\"line\"><span>    result </span><span>=</span><span> temp2 </span><span>*</span><span> 2.0</span><span>       # Kernel 3: load temp2, store result</span></span>\n<span class=\"line\"><span>    return</span><span> result</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Fused kernel using Triton</span></span>\n<span class=\"line\"><span>@triton</span><span>.</span><span>jit</span></span>\n<span class=\"line\"><span>def</span><span> fused_kernel</span><span>(</span><span>x_ptr</span><span>,</span><span> output_ptr</span><span>,</span><span> n_elements</span><span>,</span><span> BLOCK_SIZE</span><span>:</span><span> tl</span><span>.</span><span>constexpr):</span></span>\n<span class=\"line\"><span>    \"\"\"Fused kernel: x -&gt; (x + 1) -&gt; ReLU -&gt; (* 2) in one pass\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    pid </span><span>=</span><span> tl</span><span>.</span><span>program_id</span><span>(axis</span><span>=</span><span>0</span><span>)</span></span>\n<span class=\"line\"><span>    block_start </span><span>=</span><span> pid </span><span>*</span><span> BLOCK_SIZE</span></span>\n<span class=\"line\"><span>    offsets </span><span>=</span><span> block_start </span><span>+</span><span> tl</span><span>.</span><span>arange</span><span>(</span><span>0</span><span>, BLOCK_SIZE)</span></span>\n<span class=\"line\"><span>    mask </span><span>=</span><span> offsets </span><span>&lt;</span><span> n_elements</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Load input once</span></span>\n<span class=\"line\"><span>    x </span><span>=</span><span> tl</span><span>.</span><span>load</span><span>(x_ptr </span><span>+</span><span> offsets, mask</span><span>=</span><span>mask)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Perform all operations in registers</span></span>\n<span class=\"line\"><span>    temp </span><span>=</span><span> x </span><span>+</span><span> 1.0</span></span>\n<span class=\"line\"><span>    temp </span><span>=</span><span> tl</span><span>.</span><span>where</span><span>(temp </span><span>&gt;</span><span> 0</span><span>, temp, </span><span>0.0</span><span>)</span><span>  # ReLU</span></span>\n<span class=\"line\"><span>    result </span><span>=</span><span> temp </span><span>*</span><span> 2.0</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Store result once</span></span>\n<span class=\"line\"><span>    tl</span><span>.</span><span>store</span><span>(output_ptr </span><span>+</span><span> offsets, result, mask</span><span>=</span><span>mask)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> fused_operations</span><span>(</span><span>x</span><span>):</span></span>\n<span class=\"line\"><span>    \"\"\"Fused version using Triton\"\"\"</span></span>\n<span class=\"line\"><span>    output </span><span>=</span><span> torch</span><span>.</span><span>empty_like</span><span>(x)</span></span>\n<span class=\"line\"><span>    n_elements </span><span>=</span><span> output</span><span>.</span><span>numel</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Auto-tune block size for optimal performance</span></span>\n<span class=\"line\"><span>    grid </span><span>=</span><span> lambda</span><span> meta</span><span>: (triton</span><span>.</span><span>cdiv</span><span>(n_elements, meta[</span><span>'BLOCK_SIZE'</span><span>]),</span><span>)</span></span>\n<span class=\"line\"><span>    fused_kernel</span><span>[</span><span>grid</span><span>](x, output, n_elements, BLOCK_SIZE</span><span>=</span><span>1024</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> output</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> benchmark_fusion</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Compare fused vs unfused performance\"\"\"</span></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"><span>    x </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(</span><span>10_000_000</span><span>, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Benchmark unfused</span></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>    start </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    for</span><span> _ </span><span>in</span><span> range</span><span>(</span><span>100</span><span>):</span></span>\n<span class=\"line\"><span>        result1 </span><span>=</span><span> unfused_operations</span><span>(x)</span></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>    unfused_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Benchmark fused</span></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>    start </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"><span>    for</span><span> _ </span><span>in</span><span> range</span><span>(</span><span>100</span><span>):</span></span>\n<span class=\"line\"><span>        result2 </span><span>=</span><span> fused_operations</span><span>(x)</span></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>    fused_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Verify results are equivalent</span></span>\n<span class=\"line\"><span>    assert</span><span> torch</span><span>.</span><span>allclose</span><span>(result1, result2, atol</span><span>=</span><span>1e-6</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Unfused time: </span><span>{</span><span>unfused_time</span><span>:.4f</span><span>}</span><span>s\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Fused time: </span><span>{</span><span>fused_time</span><span>:.4f</span><span>}</span><span>s\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Speedup: </span><span>{</span><span>unfused_time </span><span>/</span><span> fused_time</span><span>:.2f</span><span>}</span><span>x\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Memory traffic analysis</span></span>\n<span class=\"line\"><span>    element_count </span><span>=</span><span> x</span><span>.</span><span>numel</span><span>()</span></span>\n<span class=\"line\"><span>    unfused_traffic </span><span>=</span><span> element_count </span><span>*</span><span> 4</span><span> *</span><span> 6</span><span>  # 6 memory transactions</span></span>\n<span class=\"line\"><span>    fused_traffic </span><span>=</span><span> element_count </span><span>*</span><span> 4</span><span> *</span><span> 2</span><span>    # 2 memory transactions</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Memory traffic reduction: </span><span>{</span><span>unfused_traffic </span><span>/</span><span> fused_traffic</span><span>:.2f</span><span>}</span><span>x\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># benchmark_fusion()  # Uncomment to run</span></span></code></pre>\n<h3 id=\"strategy-2-tiling-for-compute-intensive-operations\">Strategy 2: Tiling for Compute-Intensive Operations</h3>\n<p><strong>Tiling</strong> is a fundamental optimization technique that improves the performance of compute-intensive operations by exploiting the GPU’s memory hierarchy. The core idea is to load data into fast shared memory and reuse it multiple times, thereby increasing arithmetic intensity and reducing the number of expensive global memory accesses.</p>\n<p><strong>Why Tiling Works:</strong></p>\n<p>When performing operations like matrix multiplication, the naive approach requires each thread to load data from global memory multiple times. For example, in a matrix multiplication C = A × B, each element of the result matrix C[i,j] requires loading an entire row from matrix A and an entire column from matrix B. Without tiling, this leads to:</p>\n<ol>\n<li><strong>Redundant memory accesses</strong>: Multiple threads load the same data from global memory</li>\n<li><strong>Poor cache utilization</strong>: Data is loaded from global memory and immediately discarded</li>\n<li><strong>Low arithmetic intensity</strong>: The ratio of computation to memory access is suboptimal</li>\n</ol>\n<p>Tiling addresses these issues by:</p>\n<ol>\n<li><strong>Cooperative loading</strong>: Threads within a block collaborate to load data into shared memory</li>\n<li><strong>Data reuse</strong>: Once loaded, data in shared memory is reused by multiple threads</li>\n<li><strong>Increased arithmetic intensity</strong>: More computation per byte transferred from global memory</li>\n</ol>\n<p><strong>Detailed Implementation Example:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>#define TILE_SIZE 32</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>__global__ void tiled_gemm_optimized(float* A, float* B, float* C, int M, int N, int K) {</span></span>\n<span class=\"line\"><span>    // Shared memory tiles - allocated per thread block</span></span>\n<span class=\"line\"><span>    // Each tile can hold TILE_SIZE × TILE_SIZE elements</span></span>\n<span class=\"line\"><span>    __shared__ float As[TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span>    __shared__ float Bs[TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Block and thread indices for navigation</span></span>\n<span class=\"line\"><span>    int bx = blockIdx.x, by = blockIdx.y;</span></span>\n<span class=\"line\"><span>    int tx = threadIdx.x, ty = threadIdx.y;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Global indices for this thread</span></span>\n<span class=\"line\"><span>    int row = by * TILE_SIZE + ty;  // Row in matrix C</span></span>\n<span class=\"line\"><span>    int col = bx * TILE_SIZE + tx;  // Column in matrix C</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    float sum = 0.0f;  // Accumulator for this thread's result</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Main tiling loop: iterate through tiles of A and B</span></span>\n<span class=\"line\"><span>    // Number of tiles needed to cover dimension K</span></span>\n<span class=\"line\"><span>    int num_tiles = (K + TILE_SIZE - 1) / TILE_SIZE;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    for (int tile_idx = 0; tile_idx &lt; num_tiles; tile_idx++) {</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Phase 1: Collaborative loading of tiles into shared memory</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Load tile from matrix A</span></span>\n<span class=\"line\"><span>        int a_row = row;  // Same row as output</span></span>\n<span class=\"line\"><span>        int a_col = tile_idx * TILE_SIZE + tx;  // Column shifts with tile</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        if (a_row &lt; M &amp;&amp; a_col &lt; K) {</span></span>\n<span class=\"line\"><span>            As[ty][tx] = A[a_row * K + a_col];</span></span>\n<span class=\"line\"><span>        } else {</span></span>\n<span class=\"line\"><span>            As[ty][tx] = 0.0f;  // Handle boundary cases</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Load tile from matrix B</span></span>\n<span class=\"line\"><span>        int b_row = tile_idx * TILE_SIZE + ty;  // Row shifts with tile</span></span>\n<span class=\"line\"><span>        int b_col = col;  // Same column as output</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        if (b_row &lt; K &amp;&amp; b_col &lt; N) {</span></span>\n<span class=\"line\"><span>            Bs[ty][tx] = B[b_row * N + b_col];</span></span>\n<span class=\"line\"><span>        } else {</span></span>\n<span class=\"line\"><span>            Bs[ty][tx] = 0.0f;  // Handle boundary cases</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Synchronization barrier: ensure all threads have finished loading</span></span>\n<span class=\"line\"><span>        __syncthreads();</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Phase 2: Computation using shared memory data</span></span>\n<span class=\"line\"><span>        // Each thread computes its portion using the loaded tiles</span></span>\n<span class=\"line\"><span>        for (int k = 0; k &lt; TILE_SIZE; k++) {</span></span>\n<span class=\"line\"><span>            // This access is to shared memory (very fast)</span></span>\n<span class=\"line\"><span>            sum += As[ty][k] * Bs[k][tx];</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Another synchronization barrier before loading next tile</span></span>\n<span class=\"line\"><span>        __syncthreads();</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Phase 3: Write result to global memory</span></span>\n<span class=\"line\"><span>    if (row &lt; M &amp;&amp; col &lt; N) {</span></span>\n<span class=\"line\"><span>        C[row * N + col] = sum;</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>// Advanced tiling with register blocking for even better performance</span></span>\n<span class=\"line\"><span>__global__ void tiled_gemm_register_blocked(float* A, float* B, float* C, int M, int N, int K) {</span></span>\n<span class=\"line\"><span>    // Shared memory tiles</span></span>\n<span class=\"line\"><span>    __shared__ float As[TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span>    __shared__ float Bs[TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Register blocking: each thread computes multiple output elements</span></span>\n<span class=\"line\"><span>    #define REG_TILE_M 4</span></span>\n<span class=\"line\"><span>    #define REG_TILE_N 4</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    float reg_c[REG_TILE_M][REG_TILE_N] = {{0.0f}};  // Register tile for results</span></span>\n<span class=\"line\"><span>    float reg_a[REG_TILE_M];  // Register cache for A values</span></span>\n<span class=\"line\"><span>    float reg_b[REG_TILE_N];  // Register cache for B values</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    int bx = blockIdx.x, by = blockIdx.y;</span></span>\n<span class=\"line\"><span>    int tx = threadIdx.x, ty = threadIdx.y;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Each thread block handles a larger tile due to register blocking</span></span>\n<span class=\"line\"><span>    int block_row_start = by * (TILE_SIZE * REG_TILE_M);</span></span>\n<span class=\"line\"><span>    int block_col_start = bx * (TILE_SIZE * REG_TILE_N);</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // This thread's starting position within the block's output region</span></span>\n<span class=\"line\"><span>    int thread_row_start = block_row_start + ty * REG_TILE_M;</span></span>\n<span class=\"line\"><span>    int thread_col_start = block_col_start + tx * REG_TILE_N;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Main computation loop over K dimension</span></span>\n<span class=\"line\"><span>    for (int tile_k = 0; tile_k &lt; K; tile_k += TILE_SIZE) {</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Load tiles into shared memory (similar to previous example)</span></span>\n<span class=\"line\"><span>        // ... loading code omitted for brevity ...</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        __syncthreads();</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Inner loop over the tile in shared memory</span></span>\n<span class=\"line\"><span>        for (int k = 0; k &lt; TILE_SIZE; k++) {</span></span>\n<span class=\"line\"><span>            // Load values into registers for reuse</span></span>\n<span class=\"line\"><span>            for (int i = 0; i &lt; REG_TILE_M; i++) {</span></span>\n<span class=\"line\"><span>                reg_a[i] = As[ty * REG_TILE_M + i][k];</span></span>\n<span class=\"line\"><span>            }</span></span>\n<span class=\"line\"><span>            for (int j = 0; j &lt; REG_TILE_N; j++) {</span></span>\n<span class=\"line\"><span>                reg_b[j] = Bs[k][tx * REG_TILE_N + j];</span></span>\n<span class=\"line\"><span>            }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>            // Compute outer product of register vectors</span></span>\n<span class=\"line\"><span>            for (int i = 0; i &lt; REG_TILE_M; i++) {</span></span>\n<span class=\"line\"><span>                for (int j = 0; j &lt; REG_TILE_N; j++) {</span></span>\n<span class=\"line\"><span>                    reg_c[i][j] += reg_a[i] * reg_b[j];</span></span>\n<span class=\"line\"><span>                }</span></span>\n<span class=\"line\"><span>            }</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        __syncthreads();</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Write register tile back to global memory</span></span>\n<span class=\"line\"><span>    for (int i = 0; i &lt; REG_TILE_M; i++) {</span></span>\n<span class=\"line\"><span>        for (int j = 0; j &lt; REG_TILE_N; j++) {</span></span>\n<span class=\"line\"><span>            int global_row = thread_row_start + i;</span></span>\n<span class=\"line\"><span>            int global_col = thread_col_start + j;</span></span>\n<span class=\"line\"><span>            if (global_row &lt; M &amp;&amp; global_col &lt; N) {</span></span>\n<span class=\"line\"><span>                C[global_row * N + global_col] = reg_c[i][j];</span></span>\n<span class=\"line\"><span>            }</span></span>\n<span class=\"line\"><span>        }</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p><strong>Performance Analysis of Tiling:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"><span>import</span><span> numpy </span><span>as</span><span> np</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> analyze_tiling_performance</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Comprehensive analysis of tiling performance benefits\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Test different matrix sizes</span></span>\n<span class=\"line\"><span>    sizes </span><span>=</span><span> [</span><span>512</span><span>,</span><span> 1024</span><span>,</span><span> 2048</span><span>,</span><span> 4096</span><span>]</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"Tiling Performance Analysis:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"=\"</span><span> *</span><span> 60</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"</span><span>{</span><span>'Size'</span><span>:&gt;6</span><span>}</span><span> {</span><span>'Naive (ms)'</span><span>:&gt;12</span><span>}</span><span> {</span><span>'Tiled (ms)'</span><span>:&gt;12</span><span>}</span><span> {</span><span>'Speedup'</span><span>:&gt;10</span><span>}</span><span> {</span><span>'AI Improvement'</span><span>:&gt;15</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"-\"</span><span> *</span><span> 60</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    for</span><span> size </span><span>in</span><span> sizes</span><span>:</span></span>\n<span class=\"line\"><span>        # Create matrices</span></span>\n<span class=\"line\"><span>        A </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(size, size, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"><span>        B </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(size, size, device</span><span>=</span><span>device, dtype</span><span>=</span><span>torch.float32)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Naive implementation (PyTorch automatically uses optimized kernels)</span></span>\n<span class=\"line\"><span>        # For demonstration, we'll use a simple operation that isn't optimized</span></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Simulate naive memory access pattern</span></span>\n<span class=\"line\"><span>        for</span><span> _ </span><span>in</span><span> range</span><span>(</span><span>10</span><span>):</span></span>\n<span class=\"line\"><span>            C_naive </span><span>=</span><span> torch</span><span>.</span><span>matmul</span><span>(A, B)</span><span>  # This is actually optimized</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        naive_time </span><span>=</span><span> (time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time) </span><span>/</span><span> 10</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Optimized tiled implementation (cuBLAS)</span></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        for</span><span> _ </span><span>in</span><span> range</span><span>(</span><span>10</span><span>):</span></span>\n<span class=\"line\"><span>            C_tiled </span><span>=</span><span> torch</span><span>.</span><span>matmul</span><span>(A, B)</span><span>  # Uses highly optimized tiled kernels</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>        tiled_time </span><span>=</span><span> (time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time) </span><span>/</span><span> 10</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Calculate theoretical arithmetic intensity improvement</span></span>\n<span class=\"line\"><span>        # Naive: each element loaded once per use</span></span>\n<span class=\"line\"><span>        # Tiled: each element loaded once per tile and reused TILE_SIZE times</span></span>\n<span class=\"line\"><span>        tile_size </span><span>=</span><span> 32</span><span>  # Typical tile size</span></span>\n<span class=\"line\"><span>        ai_improvement </span><span>=</span><span> tile_size  </span><span># Simplified calculation</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        speedup </span><span>=</span><span> naive_time </span><span>/</span><span> tiled_time </span><span>if</span><span> tiled_time </span><span>&gt;</span><span> 0</span><span> else</span><span> float</span><span>(</span><span>'inf'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        print</span><span>(</span><span>f</span><span>\"</span><span>{</span><span>size</span><span>:6d</span><span>}</span><span> {</span><span>naive_time</span><span>*</span><span>1000</span><span>:11.2f</span><span>}</span><span> {</span><span>tiled_time</span><span>*</span><span>1000</span><span>:11.2f</span><span>}</span><span> {</span><span>speedup</span><span>:9.1f</span><span>}</span><span>x </span><span>{</span><span>ai_improvement</span><span>:14.1f</span><span>}</span><span>x\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"\\nMemory Access Pattern Analysis:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"-\"</span><span> *</span><span> 40</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Demonstrate memory access patterns</span></span>\n<span class=\"line\"><span>    def</span><span> calculate_memory_accesses</span><span>(</span><span>matrix_size</span><span>,</span><span> tile_size</span><span>):</span></span>\n<span class=\"line\"><span>        \"\"\"Calculate memory accesses for tiled vs naive GEMM\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Naive GEMM: each output element requires loading full row and column</span></span>\n<span class=\"line\"><span>        naive_loads </span><span>=</span><span> matrix_size </span><span>**</span><span> 3</span><span>  # N³ elements loaded</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Tiled GEMM: each tile loaded once and reused</span></span>\n<span class=\"line\"><span>        tiles_per_dim </span><span>=</span><span> (matrix_size </span><span>+</span><span> tile_size </span><span>-</span><span> 1</span><span>) </span><span>//</span><span> tile_size</span></span>\n<span class=\"line\"><span>        tile_loads </span><span>=</span><span> 2</span><span> *</span><span> tiles_per_dim </span><span>**</span><span> 3</span><span> *</span><span> tile_size </span><span>**</span><span> 2</span><span>  # A and B tiles</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        return</span><span> naive_loads</span><span>,</span><span> tile_loads</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    matrix_size </span><span>=</span><span> 1024</span></span>\n<span class=\"line\"><span>    tile_size </span><span>=</span><span> 32</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    naive_loads</span><span>,</span><span> tiled_loads </span><span>=</span><span> calculate_memory_accesses</span><span>(matrix_size, tile_size)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Matrix size: </span><span>{</span><span>matrix_size</span><span>}</span><span>×</span><span>{</span><span>matrix_size</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Tile size: </span><span>{</span><span>tile_size</span><span>}</span><span>×</span><span>{</span><span>tile_size</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Naive memory loads: </span><span>{</span><span>naive_loads</span><span>:,</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Tiled memory loads: </span><span>{</span><span>tiled_loads</span><span>:,</span><span>}</span><span>\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Memory access reduction: </span><span>{</span><span>naive_loads</span><span>/</span><span>tiled_loads</span><span>:.1f</span><span>}</span><span>x\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Run the analysis</span></span>\n<span class=\"line\"><span>analyze_tiling_performance</span><span>()</span></span></code></pre>\n<p><strong>Advanced Tiling Strategies:</strong></p>\n<p><strong>1. Hierarchical Tiling:</strong>\nUses multiple levels of the memory hierarchy by combining L2 cache, shared memory, and register tiling:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>__global__ void hierarchical_tiled_gemm(float* A, float* B, float* C, int M, int N, int K) {</span></span>\n<span class=\"line\"><span>    // Level 1: L2 cache tiling (handled by block scheduling)</span></span>\n<span class=\"line\"><span>    // Level 2: Shared memory tiling</span></span>\n<span class=\"line\"><span>    __shared__ float As[TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span>    __shared__ float Bs[TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Level 3: Register tiling</span></span>\n<span class=\"line\"><span>    float reg_tile[4][4] = {{0.0f}};</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Implementation combines all three levels for maximum performance</span></span>\n<span class=\"line\"><span>    // ... (detailed implementation would be quite complex)</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p><strong>2. Double Buffering:</strong>\nOverlaps computation with memory loading to hide latency:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>__global__ void double_buffered_gemm(float* A, float* B, float* C, int M, int N, int K) {</span></span>\n<span class=\"line\"><span>    // Double buffer: two sets of shared memory</span></span>\n<span class=\"line\"><span>    __shared__ float As[2][TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span>    __shared__ float Bs[2][TILE_SIZE][TILE_SIZE];</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    int read_buf = 0, compute_buf = 1;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>    // Pipeline: load next tile while computing current tile</span></span>\n<span class=\"line\"><span>    for (int tile = 0; tile &lt; num_tiles; tile++) {</span></span>\n<span class=\"line\"><span>        // Swap buffers</span></span>\n<span class=\"line\"><span>        read_buf = 1 - read_buf;</span></span>\n<span class=\"line\"><span>        compute_buf = 1 - compute_buf;</span></span>\n<span class=\"line\"><span></span></span>\n<span class=\"line\"><span>        // Asynchronously load next tile</span></span>\n<span class=\"line\"><span>        // Compute using current tile</span></span>\n<span class=\"line\"><span>        // Overlap memory and compute operations</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"><span>}</span></span></code></pre>\n<p><strong>Real-World Application: Convolution Optimization:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>import</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>functional </span><span>as</span><span> F</span></span>\n<span class=\"line\"><span>import</span><span> time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>def</span><span> optimized_convolution_tiling</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Demonstrate tiling in convolution operations\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Large convolution problem</span></span>\n<span class=\"line\"><span>    batch_size </span><span>=</span><span> 32</span></span>\n<span class=\"line\"><span>    in_channels </span><span>=</span><span> 256</span></span>\n<span class=\"line\"><span>    out_channels </span><span>=</span><span> 512</span></span>\n<span class=\"line\"><span>    height</span><span>,</span><span> width </span><span>=</span><span> 64</span><span>,</span><span> 64</span></span>\n<span class=\"line\"><span>    kernel_size </span><span>=</span><span> 3</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Input tensor and kernel</span></span>\n<span class=\"line\"><span>    input_tensor </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(batch_size, in_channels, height, width, device</span><span>=</span><span>device)</span></span>\n<span class=\"line\"><span>    kernel </span><span>=</span><span> torch</span><span>.</span><span>randn</span><span>(out_channels, in_channels, kernel_size, kernel_size, device</span><span>=</span><span>device)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"Convolution Tiling Analysis:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"-\"</span><span> *</span><span> 40</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Standard convolution (uses optimized tiled kernels internally)</span></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>    start_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    for</span><span> _ </span><span>in</span><span> range</span><span>(</span><span>100</span><span>):</span></span>\n<span class=\"line\"><span>        output </span><span>=</span><span> F</span><span>.</span><span>conv2d</span><span>(input_tensor, kernel, padding</span><span>=</span><span>1</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    torch</span><span>.</span><span>cuda</span><span>.</span><span>synchronize</span><span>()</span></span>\n<span class=\"line\"><span>    optimized_time </span><span>=</span><span> time</span><span>.</span><span>time</span><span>()</span><span> -</span><span> start_time</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Calculate theoretical arithmetic intensity</span></span>\n<span class=\"line\"><span>    # Each output pixel requires in_channels × kernel_size² operations</span></span>\n<span class=\"line\"><span>    flops </span><span>=</span><span> (batch_size </span><span>*</span><span> out_channels </span><span>*</span><span> height </span><span>*</span><span> width </span><span>*</span></span>\n<span class=\"line\"><span>             in_channels </span><span>*</span><span> kernel_size </span><span>*</span><span> kernel_size </span><span>*</span><span> 2</span><span>)  </span><span># multiply-add</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Memory traffic (input + kernel + output)</span></span>\n<span class=\"line\"><span>    input_bytes </span><span>=</span><span> input_tensor</span><span>.</span><span>numel</span><span>()</span><span> *</span><span> 4</span></span>\n<span class=\"line\"><span>    kernel_bytes </span><span>=</span><span> kernel</span><span>.</span><span>numel</span><span>()</span><span> *</span><span> 4</span></span>\n<span class=\"line\"><span>    output_bytes </span><span>=</span><span> output</span><span>.</span><span>numel</span><span>()</span><span> *</span><span> 4</span></span>\n<span class=\"line\"><span>    total_bytes </span><span>=</span><span> input_bytes </span><span>+</span><span> kernel_bytes </span><span>+</span><span> output_bytes</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    arithmetic_intensity </span><span>=</span><span> flops </span><span>/</span><span> total_bytes</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Convolution performance: </span><span>{</span><span>optimized_time</span><span>/</span><span>100</span><span>*</span><span>1000</span><span>:.2f</span><span>}</span><span> ms per forward pass\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Arithmetic intensity: </span><span>{</span><span>arithmetic_intensity</span><span>:.2f</span><span>}</span><span> FLOPs/byte\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Throughput: </span><span>{</span><span>flops</span><span>/</span><span>optimized_time</span><span>/</span><span>1e12</span><span>*</span><span>100</span><span>:.2f</span><span>}</span><span> TFLOPS\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Tiling strategy for convolution:</span></span>\n<span class=\"line\"><span>    # 1. Tile input feature maps to fit in shared memory</span></span>\n<span class=\"line\"><span>    # 2. Load kernel weights once per tile</span></span>\n<span class=\"line\"><span>    # 3. Compute multiple output channels simultaneously</span></span>\n<span class=\"line\"><span>    # 4. Use register blocking for output accumulation</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> arithmetic_intensity</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>optimized_convolution_tiling</span><span>()</span></span></code></pre>\n<p><strong>Tiling Best Practices:</strong></p>\n<ol>\n<li>\n<p><strong>Choose Optimal Tile Size:</strong></p>\n<ul>\n<li>Consider shared memory capacity (typically 48-164KB per SM)</li>\n<li>Balance between memory efficiency and occupancy</li>\n<li>Power-of-2 sizes often work well (16, 32, 64)</li>\n</ul>\n</li>\n<li>\n<p><strong>Minimize Bank Conflicts:</strong></p>\n<ul>\n<li>Avoid simultaneous access to the same memory bank</li>\n<li>Use padding or data layout transformations when necessary</li>\n</ul>\n</li>\n<li>\n<p><strong>Maximize Data Reuse:</strong></p>\n<ul>\n<li>Each element loaded into shared memory should be used multiple times</li>\n<li>Consider register blocking for even more reuse</li>\n</ul>\n</li>\n<li>\n<p><strong>Handle Boundary Conditions:</strong></p>\n<ul>\n<li>Properly handle matrices that don’t divide evenly by tile size</li>\n<li>Use conditional statements or padding as appropriate</li>\n</ul>\n</li>\n<li>\n<p><strong>Consider Memory Coalescing:</strong></p>\n<ul>\n<li>Ensure global memory accesses are coalesced</li>\n<li>Adjacent threads should access adjacent memory locations</li>\n</ul>\n</li>\n</ol>\n<p>Tiling is one of the most powerful optimization techniques for GPU computing, often providing 5-50x performance improvements for memory-bound operations by fundamentally changing the arithmetic intensity profile of algorithms.</p>\n<hr />\n<h2 id=\"modern-gpu-computing-state-of-the-art-hardware-for-research\">Modern GPU Computing: State-of-the-Art Hardware for Research</h2>\n<h3 id=\"understanding-high-performance-computing-hpc-infrastructure\">Understanding High-Performance Computing (HPC) Infrastructure</h3>\n<p>Modern research computing relies heavily on <strong>High-Performance Computing (HPC)</strong> systems that combine powerful hardware with sophisticated software stacks. Understanding how to effectively utilize these systems is crucial for researchers working at the cutting edge of computational science, machine learning, and artificial intelligence.</p>\n<p><strong>HPC System Architecture:</strong></p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>┌─────────────────────────────────────────────────────────────────┐</span></span>\n<span class=\"line\"><span>│                       HPC Cluster Overview                      │</span></span>\n<span class=\"line\"><span>├─────────────────────────────────────────────────────────────────┤</span></span>\n<span class=\"line\"><span>│  Login Nodes        │  Compute Nodes           │  Storage       │</span></span>\n<span class=\"line\"><span>│  ┌─────────────────┐ │  ┌─────────────────────┐ │  ┌───────────┐ │</span></span>\n<span class=\"line\"><span>│  │ User Access     │ │  │ CPU: 2x64-core AMD  │ │  │ Parallel  │ │</span></span>\n<span class=\"line\"><span>│  │ Job Submission  │ │  │ RAM: 512GB-2TB      │ │  │ File      │ │</span></span>\n<span class=\"line\"><span>│  │ Code Development│ │  │ GPU: 4-8x A100/H100 │ │  │ System    │ │</span></span>\n<span class=\"line\"><span>│  │ Data Analysis   │ │  │ NVLink/Infiniband   │ │  │ (Lustre/  │ │</span></span>\n<span class=\"line\"><span>│  └─────────────────┘ │  └─────────────────────┘ │  │ GPFS)     │ │</span></span>\n<span class=\"line\"><span>│                      │                         │  │ 10-100 PB │ │</span></span>\n<span class=\"line\"><span>│  Resource Mgmt       │  Scheduler (SLURM)      │  └───────────┘ │</span></span>\n<span class=\"line\"><span>│  ┌─────────────────┐ │  ┌─────────────────────┐ │               │</span></span>\n<span class=\"line\"><span>│  │ Job Queues      │ │  │ GPU Allocation      │ │  Networking   │</span></span>\n<span class=\"line\"><span>│  │ User Limits     │ │  │ Memory Management   │ │  ┌───────────┐ │</span></span>\n<span class=\"line\"><span>│  │ Fair Sharing    │ │  │ Process Isolation   │ │  │ InfiniBand│ │</span></span>\n<span class=\"line\"><span>│  └─────────────────┘ │  └─────────────────────┘ │  │ 200-400   │ │</span></span>\n<span class=\"line\"><span>└─────────────────────────────────────────────────────────────────┤  Gbps     │ │</span></span>\n<span class=\"line\"><span>                                                  │  │ Low Latency│ │</span></span>\n<span class=\"line\"><span>                                                  │  └───────────┘ │</span></span>\n<span class=\"line\"><span>                                                  └─────────────────┘</span></span></code></pre>\n<h3 id=\"advanced-gpu-architectures-for-research\">Advanced GPU Architectures for Research</h3>\n<p>Understanding modern GPU architectures is fundamental for researchers working with high-performance computing. The evolution from graphics-focused processors to general-purpose parallel computing engines has created unprecedented opportunities for scientific discovery and computational research.</p>\n<p><strong>NVIDIA A100 (Ampere Architecture) - The Research Workhorse:</strong></p>\n<p>The A100 represents a paradigm shift in GPU design, specifically engineered for AI and HPC workloads. Unlike consumer GPUs that balance gaming and compute performance, the A100 is purpose-built for sustained computational throughput with features that directly address the needs of research computing.</p>\n<p><strong>Architectural Deep Dive:</strong></p>\n<p>The A100’s architecture is built around the concept of <strong>Streaming Multiprocessors (SMs)</strong>, each containing multiple types of processing units optimized for different computational patterns:</p>\n<ul>\n<li><strong>CUDA Cores</strong>: Traditional scalar processors optimized for single-precision floating-point operations</li>\n<li><strong>Tensor Cores</strong>: Specialized matrix processing units that can perform mixed-precision matrix operations at high throughput</li>\n<li><strong>RT Cores</strong>: Ray-tracing acceleration units (present in gaming variants but not emphasized in data center models)</li>\n</ul>\n<p><strong>Memory Subsystem Design:</strong></p>\n<p>The A100’s memory architecture reflects a deep understanding of modern workload requirements:</p>\n<ul>\n<li><strong>HBM2 Memory</strong>: High Bandwidth Memory providing massive parallel access to large datasets</li>\n<li><strong>Memory Controllers</strong>: Multiple independent controllers enabling concurrent access patterns</li>\n<li><strong>Error Correction</strong>: Full ECC support ensuring data integrity for long-running scientific computations</li>\n</ul>\n<p><strong>Multi-Instance GPU (MIG) Technology:</strong></p>\n<p>One of the A100’s most innovative features is MIG, which allows a single GPU to be partitioned into up to seven independent instances. This is particularly valuable for research environments where:</p>\n<ul>\n<li>Multiple researchers need guaranteed GPU resources</li>\n<li>Different experiments have varying resource requirements</li>\n<li>Cost efficiency requires maximizing utilization across diverse workloads</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>def</span><span> analyze_a100_capabilities</span><span>():</span></span>\n<span class=\"line\"><span>    \"\"\"Comprehensive analysis of A100 computational capabilities\"\"\"</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # A100 specifications demonstrate the scale of modern research hardware</span></span>\n<span class=\"line\"><span>    a100_specs </span><span>=</span><span> {</span></span>\n<span class=\"line\"><span>        'sm_count'</span><span>:</span><span> 108</span><span>,</span><span>                    # Massive parallelism through many SMs</span></span>\n<span class=\"line\"><span>        'cuda_cores_per_sm'</span><span>:</span><span> 64</span><span>,</span><span>           # Traditional scalar processing power</span></span>\n<span class=\"line\"><span>        'tensor_cores_per_sm'</span><span>:</span><span> 4</span><span>,</span><span>          # Matrix acceleration per SM</span></span>\n<span class=\"line\"><span>        'total_cuda_cores'</span><span>:</span><span> 6912</span><span>,</span><span>          # Total scalar processing units</span></span>\n<span class=\"line\"><span>        'hbm2_memory'</span><span>:</span><span> 40</span><span>,</span><span>                 # GB - Can be 80GB in larger variants</span></span>\n<span class=\"line\"><span>        'memory_bandwidth'</span><span>:</span><span> 1555</span><span>,</span><span>          # GB/s - Critical for data-intensive research</span></span>\n<span class=\"line\"><span>        'nvlink_bandwidth'</span><span>:</span><span> 600</span><span>,</span><span>           # GB/s bidirectional for multi-GPU scaling</span></span>\n<span class=\"line\"><span>        'pcie_bandwidth'</span><span>:</span><span> 64</span><span>,</span><span>              # GB/s (PCIe 4.0) for host communication</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Compute Performance demonstrates specialization for different precisions</span></span>\n<span class=\"line\"><span>        'fp32_performance'</span><span>:</span><span> 19.5</span><span>,</span><span>          # TFLOPS - Traditional scientific computing</span></span>\n<span class=\"line\"><span>        'tensor_fp16'</span><span>:</span><span> 312</span><span>,</span><span>                # TFLOPS - AI training acceleration</span></span>\n<span class=\"line\"><span>        'tensor_bfloat16'</span><span>:</span><span> 312</span><span>,</span><span>            # TFLOPS - Improved numerical stability</span></span>\n<span class=\"line\"><span>        'tensor_int8'</span><span>:</span><span> 624</span><span>,</span><span>                # TOPS - Inference optimization</span></span>\n<span class=\"line\"><span>        'tensor_int4'</span><span>:</span><span> 1248</span><span>,</span><span>               # TOPS - Extreme quantization with sparsity</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>        # Research-specific features</span></span>\n<span class=\"line\"><span>        'mig_support'</span><span>:</span><span> True</span><span>,</span><span>               # Multi-tenancy for shared research clusters</span></span>\n<span class=\"line\"><span>        'sparsity_support'</span><span>:</span><span> True</span><span>,</span><span>          # 2:4 structured sparsity acceleration</span></span>\n<span class=\"line\"><span>        'multi_precision'</span><span>:</span><span> True</span><span>,</span><span>           # Mixed-precision training capabilities</span></span>\n<span class=\"line\"><span>        'nvlink_generation'</span><span>:</span><span> 3</span><span>,</span><span>            # High-speed inter-GPU communication</span></span>\n<span class=\"line\"><span>        'memory_error_correction'</span><span>:</span><span> True</span><span>,</span><span>   # Critical for long-running simulations</span></span>\n<span class=\"line\"><span>    }</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"NVIDIA A100 Research Capabilities Analysis:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>\"=\"</span><span> *</span><span> 60</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # The ridge point calculation reveals the balance between compute and memory</span></span>\n<span class=\"line\"><span>    ridge_point </span><span>=</span><span> a100_specs</span><span>[</span><span>'fp32_performance'</span><span>]</span><span> /</span><span> (a100_specs</span><span>[</span><span>'memory_bandwidth'</span><span>]</span><span> /</span><span> 1000</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Peak FP32 Performance: </span><span>{</span><span>a100_specs[</span><span>'fp32_performance'</span><span>]</span><span>}</span><span> TFLOPS\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Peak Memory Bandwidth: </span><span>{</span><span>a100_specs[</span><span>'memory_bandwidth'</span><span>]</span><span>}</span><span> GB/s\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Ridge Point (AI threshold): </span><span>{</span><span>ridge_point</span><span>:.1f</span><span>}</span><span> FLOPs/Byte\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Operations with AI &lt; </span><span>{</span><span>ridge_point</span><span>:.1f</span><span>}</span><span> are memory-bound\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Operations with AI &gt; </span><span>{</span><span>ridge_point</span><span>:.1f</span><span>}</span><span> can achieve compute-bound performance\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"</span><span>\\n</span><span>Specialized Compute Units for Research:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"FP16 Tensor Performance: </span><span>{</span><span>a100_specs[</span><span>'tensor_fp16'</span><span>]</span><span>}</span><span> TFLOPS\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; </span><span>{</span><span>a100_specs[</span><span>'tensor_fp16'</span><span>] </span><span>/</span><span> a100_specs[</span><span>'fp32_performance'</span><span>]</span><span>:.1f</span><span>}</span><span>x faster than FP32 for compatible workloads\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"INT8 Performance: </span><span>{</span><span>a100_specs[</span><span>'tensor_int8'</span><span>]</span><span>}</span><span> TOPS\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; </span><span>{</span><span>a100_specs[</span><span>'tensor_int8'</span><span>] </span><span>/</span><span> a100_specs[</span><span>'fp32_performance'</span><span>]</span><span>:.1f</span><span>}</span><span>x throughput for quantized inference\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Sparse INT4 Performance: </span><span>{</span><span>a100_specs[</span><span>'tensor_int4'</span><span>]</span><span>}</span><span> TOPS\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; </span><span>{</span><span>a100_specs[</span><span>'tensor_int4'</span><span>] </span><span>/</span><span> a100_specs[</span><span>'fp32_performance'</span><span>]</span><span>:.1f</span><span>}</span><span>x theoretical peak with aggressive optimization\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Memory hierarchy analysis shows the importance of data locality</span></span>\n<span class=\"line\"><span>    l2_cache_size </span><span>=</span><span> 40</span><span>  # MB - Shared across all SMs</span></span>\n<span class=\"line\"><span>    shared_mem_per_sm </span><span>=</span><span> 164</span><span>  # KB - Programmer-controlled fast memory</span></span>\n<span class=\"line\"><span>    register_file_per_sm </span><span>=</span><span> 256</span><span>  # KB - Fastest per-thread storage</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"</span><span>\\n</span><span>Memory Hierarchy for Optimal Performance:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"L2 Cache: </span><span>{</span><span>l2_cache_size</span><span>}</span><span> MB (shared across all </span><span>{</span><span>a100_specs[</span><span>'sm_count'</span><span>]</span><span>}</span><span> SMs)\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Hardware-managed, reduces global memory pressure\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Shared Memory per SM: </span><span>{</span><span>shared_mem_per_sm</span><span>}</span><span> KB\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Programmer-controlled, enables collaborative algorithms\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Register File per SM: </span><span>{</span><span>register_file_per_sm</span><span>}</span><span> KB\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Per-thread storage, highest bandwidth but limited capacity\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Total Fast Memory: </span><span>{</span><span>shared_mem_per_sm </span><span>*</span><span> a100_specs[</span><span>'sm_count'</span><span>] </span><span>/</span><span> 1024</span><span>:.1f</span><span>}</span><span> MB\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Critical for staging data and avoiding memory bottlenecks\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    # Research implications</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"</span><span>\\n</span><span>Research Computing Implications:\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Multi-Instance GPU: Up to 7 independent partitions\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Enables secure multi-tenancy in shared research environments\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"NVLink Bandwidth: </span><span>{</span><span>a100_specs[</span><span>'nvlink_bandwidth'</span><span>]</span><span>}</span><span> GB/s\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Supports large model training across multiple GPUs\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"Error Correction: Full ECC support\"</span><span>)</span></span>\n<span class=\"line\"><span>    print</span><span>(</span><span>f</span><span>\"  -&gt; Essential for long-running simulations and reliable results\"</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    return</span><span> a100_specs</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Demonstrate A100 capabilities for research applications</span></span>\n<span class=\"line\"><span>a100_specs </span><span>=</span><span> analyze_a100_capabilities</span><span>()</span></span></code></pre>\n<p><strong>NVIDIA H100 (Hopper Architecture) - The Next Generation:</strong></p>\n<p>The H100 represents the latest evolution in GPU architecture, incorporating lessons learned from the AI revolution and pushing the boundaries of what’s possible in computational research. Built on advanced process technology and featuring revolutionary architectural improvements, the H100 addresses the scaling challenges faced by modern AI research.</p>\n<p><strong>Key Architectural Innovations:</strong></p>\n<p><strong>Transformer Engine:</strong> The H100 introduces hardware specifically designed for transformer architectures, the foundation of modern large language models. This includes:</p>\n<ul>\n<li><strong>Attention-specific acceleration:</strong> Hardware units optimized for the attention mechanism</li>\n<li><strong>Dynamic precision management:</strong> Automatic adjustment of numerical precision based on training stability</li>\n<li><strong>Sparsity acceleration:</strong> Hardware support for structured sparsity patterns common in modern neural networks</li>\n</ul>\n<p><strong>Advanced Memory Architecture:</strong></p>\n<ul>\n<li><strong>HBM3 Integration:</strong> Next-generation memory providing higher bandwidth and capacity</li>\n<li><strong>Memory Deduplication:</strong> Hardware-level optimization for memory-efficient model serving</li>\n<li><strong>Confidential Computing:</strong> Hardware-level security features for sensitive research data\n<strong>GPU Generation Evolution for Research Computing:</strong></li>\n</ul>\n<p>This comparison illustrates the exponential improvement in research computing capabilities across three major GPU generations:</p>\n<h4 id=\"nvidia-v100-volta-2017---foundation-of-modern-ai\">NVIDIA V100 (Volta, 2017) - Foundation of Modern AI</h4>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Specification</th><th>Value</th></tr></thead><tbody><tr><td><strong>Memory</strong></td><td>32 GB</td></tr><tr><td><strong>Memory Bandwidth</strong></td><td>900 GB/s</td></tr><tr><td><strong>FP32 Performance</strong></td><td>15.7 TFLOPS</td></tr><tr><td><strong>Tensor FP16 Performance</strong></td><td>125 TFLOPS</td></tr><tr><td><strong>Architecture</strong></td><td>Volta (1st gen Tensor Cores)</td></tr><tr><td><strong>Process Node</strong></td><td>12nm TSMC</td></tr><tr><td><strong>Key Innovation</strong></td><td>First Tensor Cores - mixed-precision training</td></tr><tr><td><strong>Research Impact</strong></td><td>Enabled modern deep learning research</td></tr></tbody></table></div>\n<h4 id=\"nvidia-a100-ampere-2020---large-scale-ai-training\">NVIDIA A100 (Ampere, 2020) - Large-Scale AI Training</h4>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Specification</th><th>Value</th></tr></thead><tbody><tr><td><strong>Memory</strong></td><td>80 GB (max variant)</td></tr><tr><td><strong>Memory Bandwidth</strong></td><td>2,039 GB/s</td></tr><tr><td><strong>FP32 Performance</strong></td><td>19.5 TFLOPS</td></tr><tr><td><strong>Tensor FP16 Performance</strong></td><td>312 TFLOPS</td></tr><tr><td><strong>Architecture</strong></td><td>Ampere (2nd gen with sparsity)</td></tr><tr><td><strong>Process Node</strong></td><td>7nm</td></tr><tr><td><strong>Tensor Cores</strong></td><td>Gen 3 - more data types</td></tr><tr><td><strong>Key Innovation</strong></td><td>MIG + 2<div></div> structured sparsity</td></tr><tr><td><strong>Research Impact</strong></td><td>Large-scale language model training</td></tr></tbody></table></div>\n<h4 id=\"nvidia-h100-hopper-2022---trillion-parameter-research\">NVIDIA H100 (Hopper, 2022) - Trillion-Parameter Research</h4>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Specification</th><th>Value</th></tr></thead><tbody><tr><td><strong>Memory</strong></td><td>80 GB</td></tr><tr><td><strong>Memory Bandwidth</strong></td><td>3,350 GB/s</td></tr><tr><td><strong>FP32 Performance</strong></td><td>67 TFLOPS (with sparsity)</td></tr><tr><td><strong>Tensor FP16 Performance</strong></td><td>1,979 TFLOPS (with sparsity)</td></tr><tr><td><strong>Architecture</strong></td><td>Hopper (4th gen with transformer engine)</td></tr><tr><td><strong>Process Node</strong></td><td>4nm</td></tr><tr><td><strong>Tensor Cores</strong></td><td>Gen 4 - transformer-optimized</td></tr><tr><td><strong>Key Innovation</strong></td><td>Transformer Engine, FP8 support, confidential computing</td></tr><tr><td><strong>Research Impact</strong></td><td>Trillion-parameter model research</td></tr></tbody></table></div>\n<h4 id=\"performance-evolution-baseline-v100\">Performance Evolution (baseline: V100)</h4>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Metric</th><th>A100 vs V100</th><th>H100 vs V100</th></tr></thead><tbody><tr><td><strong>Memory Capacity</strong></td><td>2.5x</td><td>2.5x</td></tr><tr><td><strong>Memory Bandwidth</strong></td><td>2.3x</td><td>3.7x</td></tr><tr><td><strong>Tensor Compute</strong></td><td>2.5x</td><td>15.8x</td></tr><tr><td><strong>Research Capability</strong></td><td>~5.7x</td><td>~9.3x</td></tr><tr><td><strong>Training Efficiency</strong></td><td>~5.7x</td><td>~58.5x</td></tr></tbody></table></div>\n<h4 id=\"research-milestones-by-generation\">Research Milestones by Generation</h4>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>V100 (2017)</th><th>A100 (2020)</th><th>H100 (2022)</th></tr></thead><tbody><tr><td>BERT</td><td>GPT-3</td><td>GPT-4 class models</td></tr><tr><td>GPT-1</td><td>PaLM</td><td>Multi-modal transformers</td></tr><tr><td>ResNet at scale</td><td>Stable Diffusion</td><td>Real-time inference</td></tr></tbody></table></div>\n<h3 id=\"framework-and-library-ecosystem\">Framework and Library Ecosystem</h3>\n<p>The software ecosystem surrounding GPU computing has evolved, creating a complex landscape of tools, frameworks, and libraries. Understanding this ecosystem is crucial for making informed decisions about research infrastructure and development approaches.</p>\n<p><strong>The Modern Deep Learning Framework Landscape:</strong></p>\n<p>The choice of deep learning framework fundamentally affects research productivity, debugging capabilities, and ultimate performance. Each framework embodies different philosophies about how computational graphs should be constructed, optimized, and executed.</p>\n<p><strong>PyTorch - The Research Standard:</strong></p>\n<p>PyTorch’s success in research environments stems from its <strong>eager execution model</strong>, which means operations are executed immediately as they are encountered in Python code. This design philosophy prioritizes developer experience and debugging ease over execution efficiency.</p>\n<p><strong>Key Research Advantages:</strong></p>\n<ul>\n<li>\n<p><strong>Dynamic Computation Graphs:</strong> The computational graph is built on-the-fly, enabling easy implementation of dynamic architectures like recursive neural networks, attention mechanisms with variable sequence lengths, and reinforcement learning algorithms where the computation depends on runtime decisions.</p>\n</li>\n<li>\n<p><strong>Pythonic Interface:</strong> PyTorch feels like native Python, making it intuitive for researchers who are primarily domain experts rather than systems programmers. This reduces the cognitive overhead of implementing complex algorithms.</p>\n</li>\n<li>\n<p><strong>Debugging Experience:</strong> Standard Python debugging tools work seamlessly with PyTorch, enabling researchers to set breakpoints, inspect tensor values, and trace execution flow using familiar tools.</p>\n</li>\n<li>\n<p><strong>Research Ecosystem:</strong> The majority of recent research papers provide PyTorch implementations, creating a network effect where researchers can easily build upon previous work.</p>\n</li>\n</ul>\n<p><strong>JAX - High-Performance Scientific Computing:</strong></p>\n<p>JAX represents a different approach, emphasizing <strong>functional programming</strong> and <strong>mathematical composability</strong>. It’s designed around the principle that research code should be mathematically clear and amenable to automatic optimization.</p>\n<p><strong>Key Research Advantages:</strong></p>\n<ul>\n<li>\n<p><strong>Automatic Differentiation:</strong> JAX can automatically compute derivatives of arbitrary Python functions, including higher-order derivatives, which is crucial for certain research areas like meta-learning and physics-informed neural networks.</p>\n</li>\n<li>\n<p><strong>Transformation System:</strong> JAX provides composable transformations (jit, grad, vmap, pmap) that can be combined to automatically parallelize, vectorize, and optimize research code without manual kernel writing.</p>\n</li>\n<li>\n<p><strong>NumPy Compatibility:</strong> JAX provides a drop-in replacement for NumPy, making it easier to accelerate existing scientific computing code.</p>\n</li>\n<li>\n<p><strong>XLA Backend:</strong> Leverages Google’s XLA compiler for aggressive optimization, often achieving better performance than PyTorch for research code that fits the functional programming paradigm.</p>\n</li>\n</ul>\n<p><strong>TensorFlow - Production and Scale:</strong></p>\n<p>While less popular in pure research settings, TensorFlow remains important for research that needs to scale to production or requires specific ecosystem features.</p>\n<p><strong>Key Capabilities:</strong></p>\n<ul>\n<li><strong>TensorBoard Integration:</strong> Sophisticated visualization and monitoring tools</li>\n<li><strong>Serving Infrastructure:</strong> Built-in support for model deployment and serving</li>\n<li><strong>Mobile/Edge Deployment:</strong> TensorFlow Lite for edge device research</li>\n<li><strong>Federated Learning:</strong> TensorFlow Federated for distributed learning research</li>\n</ul>\n<h4 id=\"deep-learning-framework-comparison\">Deep Learning Framework Comparison</h4>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Aspect</th><th>PyTorch</th><th>JAX</th><th>TensorFlow</th></tr></thead><tbody><tr><td><strong>Execution Model</strong></td><td>Eager with dynamic graphs</td><td>Functional with transformations</td><td>Graph with eager option</td></tr><tr><td><strong>Research Suitability</strong></td><td>Excellent</td><td>Excellent</td><td>Moderate</td></tr><tr><td><strong>Debugging</strong></td><td>Excellent</td><td>Good</td><td>Moderate</td></tr><tr><td><strong>Performance</strong></td><td>Good</td><td>Excellent</td><td>Excellent</td></tr><tr><td><strong>Learning Curve</strong></td><td>Moderate</td><td>Steep</td><td>Moderate-High</td></tr><tr><td><strong>Auto-optimization</strong></td><td>Limited</td><td>Excellent (XLA)</td><td>Good</td></tr><tr><td><strong>Community</strong></td><td>Academic research</td><td>Scientific computing</td><td>Industry/production</td></tr></tbody></table></div>\n<p><strong>Primary Use Cases:</strong></p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Framework</th><th>Best For</th></tr></thead><tbody><tr><td><strong>PyTorch</strong></td><td>Novel architectures, rapid prototyping, CV, NLP, RL, generative models</td></tr><tr><td><strong>JAX</strong></td><td>Scientific ML, physics-informed NNs, Bayesian DL, meta-learning, optimization</td></tr><tr><td><strong>TensorFlow</strong></td><td>Production systems, distributed training, federated learning, mobile/edge AI</td></tr></tbody></table></div>\n<p><strong>Framework Selection Quick Reference:</strong></p>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Research Area</th><th>Recommended Framework</th></tr></thead><tbody><tr><td>Novel Architecture Development</td><td>PyTorch</td></tr><tr><td>High-Performance Scientific Computing</td><td>JAX</td></tr><tr><td>Large-Scale Production Research</td><td>TensorFlow</td></tr><tr><td>Rapid Prototyping</td><td>PyTorch</td></tr><tr><td>Mathematical/Physics-Informed ML</td><td>JAX</td></tr><tr><td>Distributed Training Research</td><td>PyTorch or TensorFlow</td></tr><tr><td>Edge/Mobile AI Research</td><td>TensorFlow</td></tr><tr><td>Optimization Algorithm Development</td><td>JAX</td></tr><tr><td>Computer Vision Research</td><td>PyTorch</td></tr><tr><td>Natural Language Processing</td><td>PyTorch</td></tr><tr><td>Federated Learning</td><td>TensorFlow</td></tr><tr><td>Bayesian Deep Learning</td><td>JAX</td></tr></tbody></table></div>\n<p><strong>Performance-Optimized Libraries - The Hidden Performance Multipliers:</strong></p>\n<p>Beyond the main frameworks, the choice of supporting libraries can impact research productivity and computational efficiency. Many researchers unknowingly accept performance bottlenecks by using standard libraries when high-performance alternatives exist.</p>\n<p><strong>The Performance Library Ecosystem:</strong></p>\n<p>Modern research computing benefits enormously from libraries specifically designed for high-performance workloads. These libraries often provide 5-50x performance improvements over standard Python libraries through techniques like:</p>\n<ul>\n<li><strong>Lazy Evaluation:</strong> Deferring computation until necessary and optimizing entire operation chains</li>\n<li><strong>Parallel Processing:</strong> Automatically utilizing multiple CPU cores and SIMD instructions</li>\n<li><strong>Memory Optimization:</strong> Minimizing memory allocations and improving cache utilization</li>\n<li><strong>Native Code Generation:</strong> JIT compilation to machine code rather than Python bytecode interpretation</li>\n</ul>\n<h4 id=\"performance-optimized-library-ecosystem\">Performance-Optimized Library Ecosystem</h4>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Domain</th><th>Standard</th><th>Optimized</th><th>Speedup</th><th>Memory Improvement</th></tr></thead><tbody><tr><td><strong>Data Processing</strong></td><td>pandas</td><td>polars</td><td>5-50x</td><td>2-10x less</td></tr><tr><td><strong>Numerical Computing</strong></td><td>numpy</td><td>numba</td><td>10-100x</td><td>Near C-level</td></tr><tr><td><strong>Deep Learning</strong></td><td>PyTorch (eager)</td><td>JAX + XLA</td><td>2-10x</td><td>Better fusion</td></tr><tr><td><strong>Visualization</strong></td><td>matplotlib</td><td>bokeh + datashader</td><td>3-20x</td><td>Streaming reduces needs</td></tr><tr><td><strong>Machine Learning</strong></td><td>scikit-learn</td><td>rapids-cuml</td><td>10-50x</td><td>GPU utilization</td></tr></tbody></table></div>\n<strong>Data Processing: pandas → polars</strong><p><strong>Key Technical Innovations:</strong></p><ul>\n<li>Lazy evaluation with query optimization</li>\n<li>Parallel processing across multiple cores</li>\n<li>Apache Arrow memory format</li>\n<li>Rust-based implementation for zero-cost abstractions</li>\n</ul><p><strong>Research Applications:</strong></p><ul>\n<li>Large-scale dataset preprocessing</li>\n<li>Time-series analysis for financial research</li>\n<li>Genomics data processing</li>\n<li>Social network analysis</li>\n</ul><p><strong>When to Use:</strong> Datasets &gt; 1GB, complex transformations, memory constraints</p>\n<strong>Numerical Computing: numpy → numba</strong><p><strong>Key Technical Innovations:</strong></p><ul>\n<li>Just-In-Time (JIT) compilation to machine code</li>\n<li>Automatic parallelization of loops</li>\n<li>GPU acceleration via CUDA</li>\n<li>Type specialization for optimal performance</li>\n</ul><p><strong>Research Applications:</strong></p><ul>\n<li>Monte Carlo simulations</li>\n<li>Numerical optimization algorithms</li>\n<li>Signal processing research</li>\n<li>Computational physics simulations</li>\n</ul><p><strong>When to Use:</strong> Compute-heavy loops, numerical algorithms, custom kernels needed</p>\n<strong>Deep Learning: PyTorch → JAX + XLA</strong><p><strong>Key Technical Innovations:</strong></p><ul>\n<li>XLA compiler for graph optimization</li>\n<li>Automatic operator fusion</li>\n<li>Advanced memory layout optimization</li>\n<li>Cross-device optimization</li>\n</ul><p><strong>Research Applications:</strong></p><ul>\n<li>Large-scale transformer training</li>\n<li>Scientific machine learning</li>\n<li>High-throughput inference research</li>\n<li>Multi-device training optimization</li>\n</ul><p><strong>When to Use:</strong> Performance-critical training, large models, production research</p>\n<strong>Visualization: matplotlib → bokeh + datashader</strong><p><strong>Key Technical Innovations:</strong></p><ul>\n<li>Web-based rendering with GPU acceleration</li>\n<li>Automatic data aggregation for large datasets</li>\n<li>Interactive exploration without memory limits</li>\n<li>Streaming data visualization</li>\n</ul><p><strong>Research Applications:</strong></p><ul>\n<li>Large-scale data exploration</li>\n<li>Real-time training monitoring</li>\n<li>Interactive research presentations</li>\n<li>Collaborative data analysis</li>\n</ul><p><strong>When to Use:</strong> &gt;1M data points, interactive analysis, web deployment needed</p>\n<strong>Machine Learning: scikit-learn → rapids-cuml</strong><p><strong>Key Technical Innovations:</strong></p><ul>\n<li>GPU-accelerated implementations of ML algorithms</li>\n<li>CUDA-native data structures</li>\n<li>Automatic CPU-GPU memory management</li>\n<li>Scikit-learn compatible API</li>\n</ul><p><strong>Research Applications:</strong></p><ul>\n<li>Large-scale clustering analysis</li>\n<li>High-dimensional data preprocessing</li>\n<li>Ensemble method research</li>\n<li>Traditional ML on large datasets</li>\n</ul><p><strong>When to Use:</strong> Traditional ML on large datasets, GPU resources available</p>\n<h4 id=\"library-migration-guide\">Library Migration Guide</h4>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Migration</th><th>Difficulty</th><th>Time</th><th>API Compatibility</th><th>Tip</th></tr></thead><tbody><tr><td>pandas → polars</td><td>Easy</td><td>1-2 hrs</td><td>High</td><td>Start with <code>.lazy()</code> operations</td></tr><tr><td>numpy → numba</td><td>Moderate</td><td>4-8 hrs</td><td>Medium</td><td>Add <code>@numba.jit</code> to heavy functions</td></tr><tr><td>matplotlib → bokeh</td><td>Moderate</td><td>6-10 hrs</td><td>Low</td><td>Start simple, add interactivity</td></tr><tr><td>scikit-learn → rapids</td><td>Easy</td><td>2-4 hrs</td><td>High</td><td>Ensure CUDA compatibility first</td></tr></tbody></table></div>\n<hr />\n<h2 id=\"summary\">Summary</h2>\n<p>This comprehensive guide explores the evolution of computing from traditional CPUs to modern GPU-accelerated systems, providing researchers and engineers with essential knowledge for high-performance computing.</p>\n<h3 id=\"key-topics-covered\">Key Topics Covered</h3>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n<div class=\"table-wrap\"><table><thead><tr><th>Section</th><th>Topics</th></tr></thead><tbody><tr><td><strong>Fundamental Concepts</strong></td><td>Hardware-software duality, CPU architecture, parallel computing paradigm shift</td></tr><tr><td><strong>GPU Architecture</strong></td><td>SIMD vs SIMT, thread hierarchy, memory architecture</td></tr><tr><td><strong>Performance Analysis</strong></td><td>Arithmetic intensity, memory-bound vs compute-bound kernels, A100 modeling</td></tr><tr><td><strong>Optimization Strategies</strong></td><td>Operator fusion, tiling techniques, hierarchical optimization</td></tr><tr><td><strong>Research Hardware</strong></td><td>V100 → A100 → H100 evolution, HPC infrastructure, MIG</td></tr><tr><td><strong>Software Ecosystem</strong></td><td>PyTorch/JAX/TensorFlow comparison, performance libraries, migration strategies</td></tr></tbody></table></div>\n<h3 id=\"research-impact\">Research Impact</h3>\n<p>This guide enables researchers to make informed decisions about hardware selection, software frameworks, and optimization strategies. Understanding these concepts is crucial for maximizing computational efficiency in modern research workflows, from small-scale experimentation to large-scale distributed training.</p>\n<h3 id=\"target-audience\">Target Audience</h3>\n<p>Researchers, graduate students, and engineers in computational science, machine learning, and high-performance computing who need to understand and optimize GPU-accelerated workflows.</p>",
            "url": "https://blog.ecitis.org/understanding-gpus/",
            "title": "Understanding GPUs and Parallel Computing",
            "summary": "Build a mental model of GPU architecture, memory, kernels, and parallel execution for modern machine-learning workloads.",
            "image": "https://blog.ecitis.org/open-graph/understanding-gpus.png",
            "date_modified": "2025-09-19T00:00:00.000Z",
            "date_published": "2025-09-19T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Hardware & Systems",
                "GPU",
                "CUDA",
                "Parallel computing"
            ]
        },
        {
            "id": "https://blog.ecitis.org/generative-ai-guide/",
            "content_html": "<h2 id=\"why-this-guide\">Why this guide</h2>\n<p>Modern generative AI moves quickly. Don’t lock yourself into defaults like “BERT + FAISS + all-MiniLM” just because a quick tutorial said so. Learn the landscape, then pick the right tools for your use case. This guide covers Hugging Face (models, datasets, formats), local and server runtimes (llama.cpp, Ollama, vLLM), and multi-provider access via OpenRouter. It also includes practical patterns for building solid projects.</p>\n<hr />\n<h2 id=\"key-concepts-quick-primer\">Key concepts (quick primer)</h2>\n<ul>\n<li>Foundation vs instruction-tuned models: base models (e.g., Llama 3.1 base) vs chat/instruct versions fine-tuned for following instructions.</li>\n<li>Context window: maximum tokens per request; large contexts (e.g., 128k) change RAG and prompting strategies.</li>\n<li>Tokenization: different families (SentencePiece/BPE) affect token count and latency.</li>\n<li>Throughput vs latency: batch/continuous batching, quantization, and hardware determine cost/perf.</li>\n</ul>\n<figure><figcaption><strong>Generative AI Ecosystem</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/generative-ai-guide-0.webp\" width=\"512\" height=\"712\" alt=\"Generative AI Ecosystem\" loading=\"lazy\" decoding=\"async\" /></figure>\n<hr />\n<h2 id=\"hugging-face-the-ecosystem-youll-actually-use\">Hugging Face: the ecosystem you’ll actually use</h2>\n<p>Hugging Face is a community hub for models, datasets, and apps (Spaces). You’ll use it to:</p>\n<ul>\n<li>Discover models and their licenses, sizes, and evals</li>\n<li>Pull canonical weights/configs and tokenizers</li>\n<li>Load and stream datasets at scale</li>\n<li>Reproduce and compare baselines</li>\n</ul>\n<h3 id=\"models-on-the-hub\">Models on the Hub</h3>\n<p>Typical repository contents for Transformer LLMs:</p>\n<ul>\n<li><code>config.json</code> (architecture and hyperparameters)</li>\n<li><code>tokenizer.json</code> or <code>tokenizer.model</code> (+ vocab files)</li>\n<li>Weight files: <code>.safetensors</code> (preferred) or legacy <code>.bin</code></li>\n<li><code>generation_config.json</code> (sampling defaults)</li>\n</ul>\n<p>You’ll find families like Llama, Mistral, Qwen, Phi, Gemma, OpenMistral, etc., in multiple sizes and instruction variants. Read model cards for:</p>\n<ul>\n<li>License (commercial use?), safety notes, and evaluation metrics</li>\n<li>Context length, tokenizer details, and memory requirements</li>\n<li>Recommended inference backends (e.g., vLLM, TGI, llama.cpp, TensorRT-LLM)</li>\n</ul>\n<h3 id=\"datasets-on-the-hub\">Datasets on the Hub</h3>\n<p>Use the <code>datasets</code> library for large and streaming loads:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> datasets </span><span>import</span><span> load_dataset</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Stream JSONL shards directly from the Hub</span></span>\n<span class=\"line\"><span>ds </span><span>=</span><span> load_dataset</span><span>(</span><span>\"my-org/my-dataset\"</span><span>, split</span><span>=</span><span>\"train\"</span><span>, streaming</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>for</span><span> row </span><span>in</span><span> ds</span><span>.</span><span>take</span><span>(</span><span>3</span><span>):</span></span>\n<span class=\"line\"><span>    print</span><span>(row)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Local JSONL</span></span>\n<span class=\"line\"><span>ds_local </span><span>=</span><span> load_dataset</span><span>(</span><span>\"json\"</span><span>, data_files</span><span>=</span><span>{</span><span>\"train\"</span><span>: </span><span>\"data.jsonl\"</span><span>})</span></span></code></pre>\n<p>Datasets support map/filter/shuffle, memory-mapped Arrow, and push/pull for collaboration. Prefer canonical public datasets for benchmarking when possible.</p>\n<h3 id=\"model-formats-youll-encounter\">Model formats you’ll encounter</h3>\n<ul>\n<li>PyTorch weights: <code>.safetensors</code> (zero-copy, safe); the default for Transformers</li>\n<li>GGUF: quantized, CPU/GPU-friendly format for <code>llama.cpp</code> and <code>Ollama</code></li>\n<li>ONNX: graph format for cross-runtime acceleration (ONNX Runtime, DirectML)</li>\n<li>TensorRT-LLM/MLC/AWQ/GPTQ: specialized compilers/quantization stacks; check model card support</li>\n</ul>\n<p>Use format-native runtimes when you want portability (GGUF), maximum throughput (TensorRT-LLM), or server features (vLLM).</p>\n<hr />\n<h2 id=\"local-inference-llamacpp-and-ollama\">Local inference: llama.cpp and Ollama</h2>\n<h3 id=\"llamacpp-cc-low-dependency-runtime\">llama.cpp (C/C++ low-dependency runtime)</h3>\n<p>Best for lightweight local inference and experimentation; strong CPU quantization support and GPU offload.</p>\n<ol>\n<li>Convert HF weights to GGUF (scripts in the <code>llama.cpp</code> repo):</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Example flow (exact script names vary by model family)</span></span>\n<span class=\"line\"><span>python</span><span> llama.cpp/convert.py</span><span> --model</span><span> /path/to/hf-model</span><span> --out</span><span> model.gguf</span><span> --dtype</span><span> q4_k_m</span></span></code></pre>\n<ol>\n<li>Run inference:</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>./main</span><span> -m</span><span> model.gguf</span><span> -p</span><span> \"Explain transformers in 3 bullet points\"</span><span> -n</span><span> 256</span><span> --temp</span><span> 0.7</span></span></code></pre>\n<p>Tips:</p>\n<ul>\n<li>Choose quantization to fit RAM/VRAM (e.g., Q4_K_M vs Q5_K_S). Heavier quantization → smaller, faster, lower fidelity.</li>\n<li>Stable prompts and stop tokens matter for repeatable results.</li>\n</ul>\n<h3 id=\"ollama-developer-friendly-local-model-runner\">Ollama (developer-friendly local model runner)</h3>\n<p>Ollama wraps GGUF models with a simple CLI and local REST API.</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># Install from docs, then:</span></span>\n<span class=\"line\"><span>ollama</span><span> pull</span><span> llama3.1</span></span>\n<span class=\"line\"><span>ollama</span><span> run</span><span> llama3.1</span></span></code></pre>\n<p>Custom <code>Modelfile</code>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>FROM llama3.1</span></span>\n<span class=\"line\"><span>PARAMETER temperature 0.7</span></span>\n<span class=\"line\"><span>SYSTEM You are a helpful, terse assistant.</span></span></code></pre>\n<p>Then:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>ollama</span><span> create</span><span> my-assistant</span><span> -f</span><span> Modelfile</span></span>\n<span class=\"line\"><span>ollama</span><span> run</span><span> my-assistant</span></span></code></pre>\n<p>Use the local HTTP API for apps. Good for prototypes and offline demos.</p>\n<hr />\n<h2 id=\"server-grade-inference-vllm-openai-compatible\">Server-grade inference: vLLM (OpenAI-compatible)</h2>\n<p>vLLM is a high-throughput LLM server with PagedAttention and continuous batching. It serves HF models via an OpenAI-compatible API.</p>\n<p>Run a model:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>pip</span><span> install</span><span> vllm</span></span>\n<span class=\"line\"><span>python</span><span> -m</span><span> vllm.entrypoints.openai.api_server</span><span> \\</span></span>\n<span class=\"line\"><span>  --model</span><span> meta-llama/Meta-Llama-3.1-8B-Instruct</span><span> \\</span></span>\n<span class=\"line\"><span>  --max-model-len</span><span> 32768</span><span> \\</span></span>\n<span class=\"line\"><span>  --tensor-parallel-size</span><span> 1</span></span></code></pre>\n<p>Client (OpenAI SDK-style):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> openai </span><span>import</span><span> OpenAI</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>client </span><span>=</span><span> OpenAI</span><span>(base_url</span><span>=</span><span>\"http://localhost:8000/v1\"</span><span>, api_key</span><span>=</span><span>\"sk-no-key\"</span><span>)</span></span>\n<span class=\"line\"><span>resp </span><span>=</span><span> client</span><span>.</span><span>chat</span><span>.</span><span>completions</span><span>.</span><span>create</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-8B-Instruct\"</span><span>,</span></span>\n<span class=\"line\"><span>    messages</span><span>=</span><span>[{</span><span>\"role\"</span><span>: </span><span>\"user\"</span><span>, </span><span>\"content\"</span><span>: </span><span>\"Write a haiku about GPUs.\"</span><span>}],</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(resp.choices[</span><span>0</span><span>].message.content)</span></span></code></pre>\n<p>Notes:</p>\n<ul>\n<li>Tune <code>--gpu-memory-utilization</code>, <code>--max-model-len</code>, and tensor parallel degree to hardware.</li>\n<li>vLLM supports many model families and quantization schemes; check docs for AWQ/GPTQ integration.</li>\n</ul>\n<hr />\n<h2 id=\"multi-provider-access-openrouter\">Multi-provider access: OpenRouter</h2>\n<p>OpenRouter aggregates frontier and open models behind a single, OpenAI-compatible API. You pick the model; they route to providers (price/perf vary by model).</p>\n<p>Setup:</p>\n<ol>\n<li>Create an API key in your dashboard.</li>\n<li>Use an OpenAI-compatible client with <code>base_url</code> and your key.</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> openai </span><span>import</span><span> OpenAI</span></span>\n<span class=\"line\"><span>import</span><span> os</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>client </span><span>=</span><span> OpenAI</span><span>(</span></span>\n<span class=\"line\"><span>    base_url</span><span>=</span><span>\"https://openrouter.ai/api/v1\"</span><span>,</span></span>\n<span class=\"line\"><span>    api_key</span><span>=</span><span>os.environ[</span><span>\"OPENROUTER_API_KEY\"</span><span>],</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>resp </span><span>=</span><span> client</span><span>.</span><span>chat</span><span>.</span><span>completions</span><span>.</span><span>create</span><span>(</span></span>\n<span class=\"line\"><span>    model</span><span>=</span><span>\"anthropic/claude-3.5-sonnet\"</span><span>,</span></span>\n<span class=\"line\"><span>    messages</span><span>=</span><span>[{</span><span>\"role\"</span><span>: </span><span>\"user\"</span><span>, </span><span>\"content\"</span><span>: </span><span>\"Compare Llama 3.1 and Qwen 2.\"</span><span>}],</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(resp.choices[</span><span>0</span><span>].message.content)</span></span></code></pre>\n<p>Why use it:</p>\n<ul>\n<li>Try multiple state-of-the-art models quickly (e.g., Claude, GPT-4o, Llama, Mistral, Qwen) without new SDKs.</li>\n<li>Compare quality, latency, and price before you commit to one vendor.</li>\n</ul>\n<hr />\n<h2 id=\"dont-default-your-stack-embeddings-vector-stores-and-rag\">Don’t default your stack: embeddings, vector stores, and RAG</h2>\n<p>Instead of always picking <code>all-MiniLM</code> + FAISS:</p>\n<ul>\n<li>Evaluate multiple embedding models: <code>bge-base(-en)-v1.5</code>, <code>gte-*</code>, <code>e5-*</code>, <code>voyage-*</code> (paid), domain-specific encoders. Use public evals (MTEB) and your own retrieval tests.</li>\n<li>Vector DBs: try FAISS for in-process; or production stores like Qdrant, Weaviate, pgvector, Milvus, LanceDB. Consider recall, filtering, hybrid search, and operational fit.</li>\n<li>Chunking: semantic-aware chunking often beats naive fixed-size. Keep overlap small; store titles/headers separately.</li>\n<li>RAG pipeline: retrieval → re-ranking (e.g., cross-encoder) → structured prompts. Cache retrieved context and responses.</li>\n</ul>\n<p>Minimal retrieval sketch (HF embeddings + Qdrant):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> sentence_transformers </span><span>import</span><span> SentenceTransformer</span></span>\n<span class=\"line\"><span>from</span><span> qdrant_client </span><span>import</span><span> QdrantClient</span></span>\n<span class=\"line\"><span>from</span><span> qdrant_client</span><span>.</span><span>models </span><span>import</span><span> Distance</span><span>,</span><span> VectorParams</span><span>,</span><span> PointStruct</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> SentenceTransformer</span><span>(</span><span>\"sentence-transformers/all-MiniLM-L6-v2\"</span><span>)</span><span>  # swap for stronger embeddings</span></span>\n<span class=\"line\"><span>client </span><span>=</span><span> QdrantClient</span><span>(host</span><span>=</span><span>\"localhost\"</span><span>, port</span><span>=</span><span>6333</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>client</span><span>.</span><span>recreate_collection</span><span>(</span></span>\n<span class=\"line\"><span>    collection_name</span><span>=</span><span>\"docs\"</span><span>,</span></span>\n<span class=\"line\"><span>    vectors_config</span><span>=</span><span>VectorParams</span><span>(size</span><span>=</span><span>384</span><span>, distance</span><span>=</span><span>Distance.COSINE),</span></span>\n<span class=\"line\"><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>texts </span><span>=</span><span> [</span><span>\"Doc A...\"</span><span>,</span><span> \"Doc B...\"</span><span>]</span></span>\n<span class=\"line\"><span>vecs </span><span>=</span><span> model</span><span>.</span><span>encode</span><span>(texts, normalize_embeddings</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>client</span><span>.</span><span>upsert</span><span>(</span><span>\"docs\"</span><span>, [</span><span>PointStruct</span><span>(id</span><span>=</span><span>i, vector</span><span>=</span><span>vecs[i], payload</span><span>=</span><span>{</span><span>\"text\"</span><span>: t}) </span><span>for</span><span> i, t </span><span>in</span><span> enumerate</span><span>(texts)])</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>query </span><span>=</span><span> \"What is in Doc B?\"</span></span>\n<span class=\"line\"><span>qvec </span><span>=</span><span> model</span><span>.</span><span>encode</span><span>([query], normalize_embeddings</span><span>=</span><span>True</span><span>)</span><span>[</span><span>0</span><span>]</span></span>\n<span class=\"line\"><span>hits </span><span>=</span><span> client</span><span>.</span><span>search</span><span>(</span><span>\"docs\"</span><span>, query_vector</span><span>=</span><span>qvec, limit</span><span>=</span><span>3</span><span>)</span></span></code></pre>\n<p>Swap the encoder, add reranking, and evaluate retrieval quality. Don’t assume defaults are best.</p>\n<hr />\n<h2 id=\"building-good-projects-patterns-that-ship\">Building good projects: patterns that ship</h2>\n<ul>\n<li>Prompting as an interface: define schemas (JSON Mode, function/tool calling) and validate outputs.</li>\n<li>Evals: build small, automatic tests early. Use task-specific metrics (e.g., exact match for QA, Rouge/BLEU for summarization) and Golden Sets. Tools like Ragas can help for RAG.</li>\n<li>Guardrails: Pydantic validation, regex/JSON schema, allow lists. Consider safety filters for public apps.</li>\n<li>Observability: log prompts, latencies, tokens, costs. Redact PII. Track drift.</li>\n<li>Caching: local + hosted caches (e.g., KV stores). Deduplicate identical requests and retrieved contexts.</li>\n<li>Cost control: use smaller models for most traffic; escalate to larger ones only when needed (router policies).</li>\n</ul>\n<hr />\n<h2 id=\"choosing-the-right-runtime\">Choosing the right runtime</h2>\n<ul>\n<li>Prototyping/offline: <code>Ollama</code> (GGUF) or <code>llama.cpp</code>; low setup, great UX.</li>\n<li>Single-node serving: <code>vLLM</code>; high throughput, OpenAI-compatible API for apps.</li>\n<li>Cloud APIs: <code>OpenRouter</code>; fast model switching/comparison across providers.</li>\n<li>Hardware-optimized: <code>TensorRT-LLM</code>, <code>Hugging Face TGI</code>, or vendor-specific services for maximum perf.</li>\n</ul>\n<hr />\n<h2 id=\"quick-starts\">Quick starts</h2>\n<ol>\n<li>Local chat with Ollama</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>ollama</span><span> pull</span><span> mistral</span></span>\n<span class=\"line\"><span>ollama</span><span> run</span><span> mistral</span></span></code></pre>\n<ol>\n<li>Serve Llama 3.1 with vLLM and call it with OpenAI SDK</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>python</span><span> -m</span><span> vllm.entrypoints.openai.api_server</span><span> --model</span><span> meta-llama/Meta-Llama-3.1-8B-Instruct</span></span></code></pre>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> openai </span><span>import</span><span> OpenAI</span></span>\n<span class=\"line\"><span>client </span><span>=</span><span> OpenAI</span><span>(base_url</span><span>=</span><span>\"http://localhost:8000/v1\"</span><span>, api_key</span><span>=</span><span>\"sk-no-key\"</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(client.chat.completions.</span><span>create</span><span>(model</span><span>=</span><span>\"meta-llama/Meta-Llama-3.1-8B-Instruct\"</span><span>, messages</span><span>=</span><span>[{</span><span>\"role\"</span><span>: </span><span>\"user\"</span><span>, </span><span>\"content\"</span><span>: </span><span>\"Say hi\"</span><span>}]).choices[</span><span>0</span><span>].message.content)</span></span></code></pre>\n<ol>\n<li>Try multiple frontier models via OpenRouter</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> openai </span><span>import</span><span> OpenAI</span></span>\n<span class=\"line\"><span>import</span><span> os</span></span>\n<span class=\"line\"><span>client </span><span>=</span><span> OpenAI</span><span>(base_url</span><span>=</span><span>\"https://openrouter.ai/api/v1\"</span><span>, api_key</span><span>=</span><span>os.environ[</span><span>\"OPENROUTER_API_KEY\"</span><span>])</span></span>\n<span class=\"line\"><span>print</span><span>(client.chat.completions.</span><span>create</span><span>(model</span><span>=</span><span>\"openai/gpt-4o\"</span><span>, messages</span><span>=</span><span>[{</span><span>\"role\"</span><span>: </span><span>\"user\"</span><span>, </span><span>\"content\"</span><span>: </span><span>\"Give 3 project ideas with data sources.\"</span><span>}]).choices[</span><span>0</span><span>].message.content)</span></span></code></pre>\n<hr />\n<h2 id=\"references\">References</h2>\n<ul>\n<li>Hugging Face Hub: <a href=\"https://huggingface.co/\">huggingface.co</a></li>\n<li>Transformers: <a href=\"https://huggingface.co/docs/transformers\">huggingface.co/docs/transformers</a></li>\n<li>Datasets: <a href=\"https://huggingface.co/docs/datasets\">huggingface.co/docs/datasets</a></li>\n<li>llama.cpp: <a href=\"https://github.com/ggerganov/llama.cpp\">github.com/ggerganov/llama.cpp</a></li>\n<li>Ollama: <a href=\"https://ollama.com/\">ollama.com</a></li>\n<li>vLLM: <a href=\"https://github.com/vllm-project/vllm\">github.com/vllm-project/vllm</a></li>\n<li>OpenRouter: <a href=\"https://openrouter.ai/\">openrouter.ai</a></li>\n</ul>",
            "url": "https://blog.ecitis.org/generative-ai-guide/",
            "title": "Modern Generative AI Guide",
            "summary": "A grounded map of models, datasets, runtimes, and provider tools for building modern generative-AI systems without cargo-cult defaults.",
            "image": "https://blog.ecitis.org/open-graph/generative-ai-guide.png",
            "date_modified": "2025-09-10T00:00:00.000Z",
            "date_published": "2025-09-10T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Ecosystems & Tooling",
                "Generative AI",
                "Runtimes",
                "Model selection"
            ]
        },
        {
            "id": "https://blog.ecitis.org/hpc-guide/",
            "content_html": "<h2 id=\"overview\">Overview</h2>\n<p>This guide shows how to use Astral’s <code>uv</code> for fast, reproducible Python environments on HPC systems, install PyTorch with CUDA for NVIDIA A100 GPUs, tune performance (data loading, AMP/TF32, memory format, compile, CUDA Graphs, DDP), and monitor with <code>nvidia-smi</code>, <code>nvtop</code>, and NVIDIA Nsight tools.</p>\n<figure><figcaption><strong>HPC Development Workflow</strong></figcaption><img src=\"https://blog.ecitis.org/feeds/figures/hpc-guide-0.webp\" width=\"512\" height=\"989\" alt=\"HPC Development Workflow\" loading=\"lazy\" decoding=\"async\" /></figure>\n<p>Links you’ll use often:</p>\n<ul>\n<li>uv docs: <a href=\"https://docs.astral.sh/uv/\">docs.astral.sh/uv</a></li>\n<li>PyTorch install matrix: <a href=\"https://pytorch.org/get-started/locally/\">pytorch.org/get-started/locally</a></li>\n<li>PyTorch performance tuning: <a href=\"https://pytorch.org/tutorials/recipes/recipes/tuning_guide.html\">pytorch.org/tutorials/recipes/recipes/tuning_guide.html</a></li>\n</ul>\n<hr />\n<h2 id=\"install-uv-recommended\">Install uv (recommended)</h2>\n<p>uv is a fast package/project manager and Python launcher. It replaces <code>pip</code>, <code>virtualenv</code>, some <code>poetry</code> tasks, and even manages Python interpreters.</p>\n<p>On Linux/macOS:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>curl</span><span> -LsSf</span><span> https://astral.sh/uv/install.sh</span><span> |</span><span> sh</span></span>\n<span class=\"line\"><span>uv</span><span> --version</span></span></code></pre>\n<p>On Windows (PowerShell):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>powershell </span><span>-</span><span>ExecutionPolicy ByPass </span><span>-</span><span>c </span><span>\"irm https://astral.sh/uv/install.ps1 | iex\"</span></span>\n<span class=\"line\"><span>uv </span><span>--</span><span>version</span></span></code></pre>\n<p>If your cluster forbids curl-based installers, you can fall back to:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>pip</span><span> install</span><span> --user</span><span> uv</span></span>\n<span class=\"line\"><span>~</span><span>/.local/bin/uv --version</span></span></code></pre>\n<hr />\n<h2 id=\"manage-python-versions-with-uv\">Manage Python versions with uv</h2>\n<p>Keep project-specific Python consistent to avoid ABI mismatches with CUDA.</p>\n<ul>\n<li>List installed Pythons:</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>uv</span><span> python</span><span> list</span></span></code></pre>\n<ul>\n<li>Install specific versions (uses Python standalone builds):</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>uv</span><span> python</span><span> install</span><span> 3.11</span><span> 3.12</span></span></code></pre>\n<ul>\n<li>Pin the project to a version (writes <code>.python-version</code>):</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>uv</span><span> python</span><span> pin</span><span> 3.11</span></span></code></pre>\n<p>When running on shared clusters, ensure your job loads the same Python (via pinning) on login nodes and compute nodes.</p>\n<hr />\n<h2 id=\"create-and-use-virtual-environments-venvs\">Create and use virtual environments (venvs)</h2>\n<p>Recommended: put a <code>.venv/</code> in the project root and commit <code>.venv/</code> to <code>.gitignore</code>.</p>\n<ul>\n<li>Create venv with pinned Python:</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>uv</span><span> venv</span><span> --python</span><span> 3.11</span></span></code></pre>\n<ul>\n<li>Activate (Linux/macOS):</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>source</span><span> .venv/bin/activate</span></span></code></pre>\n<ul>\n<li>Activate (Windows):</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>.venv\\Scripts\\Activate</span></span></code></pre>\n<ul>\n<li>Upgrade pip/setuptools/wheel inside the venv (uv provides a fast shim):</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>uv</span><span> pip</span><span> install</span><span> -U</span><span> pip</span><span> setuptools</span><span> wheel</span></span></code></pre>\n<ul>\n<li>Freeze or export deps if you aren’t using <code>pyproject.toml</code> yet:</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>uv</span><span> pip</span><span> freeze</span><span> &gt;</span><span> requirements.txt</span></span></code></pre>\n<p>Tip: uv can also manage project deps via <code>pyproject.toml</code> + <code>uv.lock</code> with <code>uv add</code> and <code>uv sync</code> if you prefer lockfiles.</p>\n<hr />\n<h2 id=\"install-pytorch-with-cuda-for-a100\">Install PyTorch with CUDA for A100</h2>\n<ol>\n<li>Confirm drivers and CUDA runtime on the node:</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>nvidia-smi</span></span>\n<span class=\"line\"><span>nvidia-smi</span><span> topo</span><span> -m</span><span>   # NUMA / PCIe fabric view</span></span></code></pre>\n<p>Check the “CUDA Version” reported by <code>nvidia-smi</code> (this reflects the driver-supported CUDA runtime). On many A100 systems you’ll see CUDA 11.8 or 12.1+.</p>\n<ol>\n<li>In your venv, install the matching PyTorch wheels. Use the official index URL for CUDA builds (replace cu118/cu121 with what you need):</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># CUDA 11.8</span></span>\n<span class=\"line\"><span>uv</span><span> pip</span><span> install</span><span> torch</span><span> torchvision</span><span> torchaudio</span><span> --index-url</span><span> https://download.pytorch.org/whl/cu118</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># CUDA 12.1</span></span>\n<span class=\"line\"><span>uv</span><span> pip</span><span> install</span><span> torch</span><span> torchvision</span><span> torchaudio</span><span> --index-url</span><span> https://download.pytorch.org/whl/cu121</span></span></code></pre>\n<p>If your cluster uses modules, load them first (example):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>module</span><span> load</span><span> cuda/12.1</span><span>  # or the site-provided module that matches your drivers</span></span></code></pre>\n<ol>\n<li>Verify GPU visibility in Python:</li>\n</ol>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>print</span><span>(torch.</span><span>__version__</span><span>)</span></span>\n<span class=\"line\"><span>print</span><span>(torch.version.cuda)</span></span>\n<span class=\"line\"><span>print</span><span>(torch.cuda.</span><span>is_available</span><span>())</span></span>\n<span class=\"line\"><span>print</span><span>(torch.cuda.</span><span>get_device_name</span><span>(</span><span>0</span><span>))</span></span></code></pre>\n<p>If <code>is_available()</code> is False, check that you’re on a GPU node, the right module is loaded, and your PyTorch wheels match the available CUDA runtime.</p>\n<hr />\n<h2 id=\"data-loading-keep-the-gpu-fed\">Data loading: keep the GPU fed</h2>\n<p>The A100 is very fast; most training slowdowns are input-bound. Key <code>DataLoader</code> knobs:</p>\n<ul>\n<li><code>num_workers</code>: parallel CPU workers. Start near the number of physical cores per GPU slice, then sweep.</li>\n<li><code>pin_memory=True</code>: enables page-locked host buffers for faster H2D copies.</li>\n<li><code>persistent_workers=True</code>: avoid worker teardown cost between epochs (when <code>num_workers&gt;0</code>).</li>\n<li><code>prefetch_factor</code>: number of batches preloaded by each worker (default 2). Increase if GPU starves.</li>\n<li>Use non-blocking transfers with pinned memory: <code>.to(device, non_blocking=True)</code>.</li>\n</ul>\n<p>Example:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> os</span></span>\n<span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>utils</span><span>.</span><span>data </span><span>import</span><span> DataLoader</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>num_cpu </span><span>=</span><span> os</span><span>.</span><span>cpu_count</span><span>()</span><span> or</span><span> 8</span></span>\n<span class=\"line\"><span>loader </span><span>=</span><span> DataLoader</span><span>(</span></span>\n<span class=\"line\"><span>    dataset,</span></span>\n<span class=\"line\"><span>    batch_size</span><span>=</span><span>64</span><span>,</span></span>\n<span class=\"line\"><span>    shuffle</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    num_workers</span><span>=</span><span>min</span><span>(</span><span>8</span><span>, num_cpu),</span></span>\n<span class=\"line\"><span>    pin_memory</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    persistent_workers</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    prefetch_factor</span><span>=</span><span>4</span><span>,</span></span>\n<span class=\"line\"><span>)</span></span></code></pre>\n<p>Watch <code>GPU-Util</code> vs CPU load. If GPU &lt; 90% while CPUs are busy, increase <code>prefetch_factor</code>/<code>num_workers</code>. If CPUs are idle and GPU is low-util, you may be I/O bound (optimize storage, caching, or preprocessing).</p>\n<hr />\n<h2 id=\"math-throughput-amp-tf32-memory-formats\">Math throughput: AMP, TF32, memory formats</h2>\n<ul>\n<li>Automatic Mixed Precision (AMP): use <code>torch.cuda.amp</code> for FP16/BF16 where safe.</li>\n<li>TF32 on A100: accelerates FP32 matmuls transparently; you can explicitly allow it.</li>\n<li>Channels-last memory format (NCHW → NHWC) helps convolution throughput on Ampere.</li>\n</ul>\n<p>Minimal training loop sketch:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>amp </span><span>import</span><span> autocast</span><span>,</span><span> GradScaler</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> model</span><span>.</span><span>to</span><span>(device)</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> model</span><span>.</span><span>to</span><span>(memory_format</span><span>=</span><span>torch.channels_last)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Enable cudnn autotuner (best algos per input shapes)</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>backends</span><span>.</span><span>cudnn</span><span>.</span><span>benchmark </span><span>=</span><span> True</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># Prefer higher TF32 throughput where numerically OK</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>backends</span><span>.</span><span>cuda</span><span>.</span><span>matmul</span><span>.</span><span>allow_tf32 </span><span>=</span><span> True</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>backends</span><span>.</span><span>cudnn</span><span>.</span><span>allow_tf32 </span><span>=</span><span> True</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>scaler </span><span>=</span><span> GradScaler</span><span>()</span><span>  # use torch.amp.GradScaler for BF16 on newer PyTorch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> images</span><span>,</span><span> targets </span><span>in</span><span> loader</span><span>:</span></span>\n<span class=\"line\"><span>    images </span><span>=</span><span> images</span><span>.</span><span>to</span><span>(device, non_blocking</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>    images </span><span>=</span><span> images</span><span>.</span><span>to</span><span>(memory_format</span><span>=</span><span>torch.channels_last)</span></span>\n<span class=\"line\"><span>    targets </span><span>=</span><span> targets</span><span>.</span><span>to</span><span>(device, non_blocking</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>    optimizer</span><span>.</span><span>zero_grad</span><span>(set_to_none</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>    with</span><span> autocast</span><span>():</span></span>\n<span class=\"line\"><span>        outputs </span><span>=</span><span> model</span><span>(images)</span></span>\n<span class=\"line\"><span>        loss </span><span>=</span><span> criterion</span><span>(outputs, targets)</span></span>\n<span class=\"line\"><span>    scaler</span><span>.</span><span>scale</span><span>(loss).</span><span>backward</span><span>()</span></span>\n<span class=\"line\"><span>    scaler</span><span>.</span><span>step</span><span>(optimizer)</span></span>\n<span class=\"line\"><span>    scaler</span><span>.</span><span>update</span><span>()</span></span></code></pre>\n<p>Notes:</p>\n<ul>\n<li>Prefer BF16 on A100 when supported by your model for stability without scaling.</li>\n<li>Keep batch shapes stable to maximize <code>cudnn.benchmark</code> benefit.</li>\n</ul>\n<hr />\n<h2 id=\"compile-and-cuda-graphs\">Compile and CUDA Graphs</h2>\n<ul>\n<li><code>torch.compile</code> (PyTorch 2.x) can fuse kernels and reduce Python overhead. Try safe modes first:</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>compiled_model </span><span>=</span><span> torch</span><span>.</span><span>compile</span><span>(model, mode</span><span>=</span><span>'max-autotune'</span><span>)</span></span></code></pre>\n<ul>\n<li>CUDA Graphs reduce per-iteration launch overhead when your step is shape-stable:</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> torch</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>static_inp </span><span>=</span><span> torch</span><span>.</span><span>empty_like</span><span>(example_inp, device</span><span>=</span><span>'cuda'</span><span>)</span></span>\n<span class=\"line\"><span>g </span><span>=</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>CUDAGraph</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># warm-up allocations</span></span>\n<span class=\"line\"><span>_ </span><span>=</span><span> model</span><span>(example_inp)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>optimizer</span><span>.</span><span>zero_grad</span><span>(set_to_none</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>with</span><span> torch</span><span>.</span><span>cuda</span><span>.</span><span>graph</span><span>(g):</span></span>\n<span class=\"line\"><span>    out </span><span>=</span><span> model</span><span>(static_inp)</span></span>\n<span class=\"line\"><span>    loss </span><span>=</span><span> criterion</span><span>(out, static_target)</span></span>\n<span class=\"line\"><span>    loss</span><span>.</span><span>backward</span><span>()</span></span>\n<span class=\"line\"><span>    optimizer</span><span>.</span><span>step</span><span>()</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>for</span><span> inp</span><span>,</span><span> tgt </span><span>in</span><span> loader</span><span>:</span></span>\n<span class=\"line\"><span>    static_inp</span><span>.</span><span>copy_</span><span>(inp.</span><span>to</span><span>(</span><span>'cuda'</span><span>, non_blocking</span><span>=</span><span>True</span><span>))</span></span>\n<span class=\"line\"><span>    static_target</span><span>.</span><span>copy_</span><span>(tgt.</span><span>to</span><span>(</span><span>'cuda'</span><span>, non_blocking</span><span>=</span><span>True</span><span>))</span></span>\n<span class=\"line\"><span>    g</span><span>.</span><span>replay</span><span>()</span></span></code></pre>\n<p>Both features benefit stable shapes and control flow; avoid random control paths inside the captured region.</p>\n<hr />\n<h2 id=\"multi-gpu-on-a-single-node-ddp\">Multi-GPU on a single node (DDP)</h2>\n<p>Use Distributed Data Parallel (NCCL) to scale across A100s.</p>\n<p>Launcher (single node, 8 GPUs):</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>torchrun</span><span> --standalone</span><span> --nproc_per_node=8</span><span> train.py</span><span> --arg1</span><span> ...</span></span></code></pre>\n<p>Inside <code>train.py</code>:</p>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> os</span><span>,</span><span> torch</span><span>,</span><span> torch</span><span>.</span><span>distributed </span><span>as</span><span> dist</span></span>\n<span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>nn</span><span>.</span><span>parallel </span><span>import</span><span> DistributedDataParallel </span><span>as</span><span> DDP</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>dist</span><span>.</span><span>init_process_group</span><span>(</span><span>'nccl'</span><span>)</span></span>\n<span class=\"line\"><span>rank </span><span>=</span><span> dist</span><span>.</span><span>get_rank</span><span>()</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>cuda</span><span>.</span><span>set_device</span><span>(rank </span><span>%</span><span> torch.cuda.</span><span>device_count</span><span>())</span></span>\n<span class=\"line\"><span>device </span><span>=</span><span> torch</span><span>.</span><span>device</span><span>(</span><span>'cuda'</span><span>, rank </span><span>%</span><span> torch.cuda.</span><span>device_count</span><span>())</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>model </span><span>=</span><span> model</span><span>.</span><span>to</span><span>(device)</span></span>\n<span class=\"line\"><span>model </span><span>=</span><span> DDP</span><span>(model, device_ids</span><span>=</span><span>[device.index])</span></span></code></pre>\n<p>Use <code>DistributedSampler</code> for datasets and call <code>sampler.set_epoch(epoch)</code> each epoch.</p>\n<hr />\n<h2 id=\"cpu-affinity-threads-and-io\">CPU affinity, threads, and I/O</h2>\n<ul>\n<li>Respect scheduler allocations: in SLURM, set <code>--cpus-per-task</code> to match <code>num_workers</code> and set threads:</li>\n</ul>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>import</span><span> os</span><span>,</span><span> torch</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>set_num_threads</span><span>(</span><span>int</span><span>(os.environ.</span><span>get</span><span>(</span><span>'SLURM_CPUS_PER_TASK'</span><span>, </span><span>'4'</span><span>)))</span></span>\n<span class=\"line\"><span>torch</span><span>.</span><span>set_num_interop_threads</span><span>(</span><span>1</span><span>)</span></span></code></pre>\n<ul>\n<li>Prefer fast local storage (NVMe, node-local scratch). If reading from network storage, increase DataLoader prefetching and use binary formats (e.g., WebDataset/TFRecord) to reduce per-file overhead.</li>\n</ul>\n<hr />\n<h2 id=\"monitoring-and-profiling-on-a100\">Monitoring and profiling on A100</h2>\n<h3 id=\"quick-health-and-utilization\">Quick health and utilization</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>nvidia-smi</span><span>                 # snapshot</span></span>\n<span class=\"line\"><span>watch</span><span> -n</span><span> 1</span><span> nvidia-smi</span><span>      # live refresh</span></span>\n<span class=\"line\"><span>nvidia-smi</span><span> dmon</span><span> -s</span><span> pucm</span><span>    # per-GPU perf counters (Power/Util/Clock/Memory)</span></span>\n<span class=\"line\"><span>nvidia-smi</span><span> topo</span><span> -m</span><span>         # topology (NVLink/PCIe/NUMA)</span></span></code></pre>\n<p>How to read:</p>\n<ul>\n<li>GPU-Util: aim for ~90%+ under steady-state training.</li>\n<li>Memory-Usage: match batch size to available FB memory with headroom (fragmentation, caches).</li>\n<li>Power/Thermals: A100 40GB/80GB will throttle if constrained; file an ops ticket if sustained throttling appears.</li>\n</ul>\n<h3 id=\"nvtop-interactive-gpu-top\">nvtop (interactive GPU top)</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>sudo</span><span> apt-get</span><span> install</span><span> nvtop</span><span>  # Debian/Ubuntu</span></span>\n<span class=\"line\"><span>nvtop</span></span></code></pre>\n<p>Shows per-process SM/memory utilization and graphs. Useful to identify which PID is starving or leaking memory.</p>\n<h3 id=\"pytorch-profiler-cpucuda\">PyTorch Profiler (CPU+CUDA)</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>from</span><span> torch</span><span>.</span><span>profiler </span><span>import</span><span> profile</span><span>,</span><span> record_function</span><span>,</span><span> ProfilerActivity</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>with</span><span> profile</span><span>(</span></span>\n<span class=\"line\"><span>    activities</span><span>=</span><span>[ProfilerActivity.CPU, ProfilerActivity.CUDA],</span></span>\n<span class=\"line\"><span>    record_shapes</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    profile_memory</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>    with_stack</span><span>=</span><span>True</span><span>,</span></span>\n<span class=\"line\"><span>)</span><span> as</span><span> prof</span><span>:</span></span>\n<span class=\"line\"><span>    for</span><span> step</span><span>,</span><span> (x</span><span>,</span><span> y) </span><span>in</span><span> enumerate</span><span>(loader):</span></span>\n<span class=\"line\"><span>        if</span><span> step </span><span>==</span><span> 100</span><span>:</span><span> break</span></span>\n<span class=\"line\"><span>        with</span><span> record_function</span><span>(</span><span>'train_step'</span><span>):</span></span>\n<span class=\"line\"><span>            x </span><span>=</span><span> x</span><span>.</span><span>to</span><span>(</span><span>'cuda'</span><span>, non_blocking</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>            y </span><span>=</span><span> y</span><span>.</span><span>to</span><span>(</span><span>'cuda'</span><span>, non_blocking</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"><span>            out </span><span>=</span><span> model</span><span>(x)</span></span>\n<span class=\"line\"><span>            loss </span><span>=</span><span> criterion</span><span>(out, y)</span></span>\n<span class=\"line\"><span>            loss</span><span>.</span><span>backward</span><span>()</span></span>\n<span class=\"line\"><span>            optimizer</span><span>.</span><span>step</span><span>()</span><span>; optimizer</span><span>.</span><span>zero_grad</span><span>(set_to_none</span><span>=</span><span>True</span><span>)</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span>print</span><span>(prof.</span><span>key_averages</span><span>().</span><span>table</span><span>(sort_by</span><span>=</span><span>'cuda_time_total'</span><span>, row_limit</span><span>=</span><span>15</span><span>))</span></span>\n<span class=\"line\"><span>prof</span><span>.</span><span>export_chrome_trace</span><span>(</span><span>'trace.json'</span><span>)</span></span></code></pre>\n<p>Open <code>trace.json</code> in Chrome’s tracing viewer or TensorBoard’s profiler plugin for timelines.</p>\n<h3 id=\"nsight-systems-end-to-end-timeline\">Nsight Systems (end-to-end timeline)</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>nsys</span><span> profile</span><span> -o</span><span> profile_report</span><span> python</span><span> train.py</span></span></code></pre>\n<p>Use the GUI to inspect kernel launches, H2D/D2H copies, CPU stalls, and NCCL collectives.</p>\n<h3 id=\"nsight-compute-kernel-level\">Nsight Compute (kernel-level)</h3>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span>ncu</span><span> --set</span><span> full</span><span> --target-processes</span><span> all</span><span> python</span><span> train.py</span></span></code></pre>\n<p>Drill into specific kernels to see occupancy, memory throughput, and bottlenecks.</p>\n<hr />\n<h2 id=\"common-pitfalls-and-fixes\">Common pitfalls and fixes</h2>\n<ul>\n<li>Mismatched CUDA: wheel suffix (e.g., <code>cu118</code>/<code>cu121</code>) must be compatible with your driver’s CUDA runtime. Reinstall with the right <code>--index-url</code>.</li>\n<li>Data starvation: increase <code>num_workers</code>, <code>prefetch_factor</code>, enable <code>pin_memory</code>, move preprocessing to workers.</li>\n<li>Unstable AMP: try BF16 first on A100, or lower LR, or disable AMP for sensitive ops.</li>\n<li>Irregular shapes: pad/pack batches to stabilize shapes; enables better autotuning, compile, and graphs.</li>\n<li>Multi-GPU under-utilization: check <code>nvidia-smi topo -m</code> for NVLink/PCIe topology, ensure NCCL is the backend, and avoid CPU oversubscription.</li>\n</ul>\n<hr />\n<h2 id=\"minimal-end-to-end-recipe-a100-cuda-121\">Minimal end-to-end recipe (A100, CUDA 12.1)</h2>\n<pre class=\"copy-code-block\"><code><span class=\"line\"><span># 0) (If needed) load site modules</span></span>\n<span class=\"line\"><span>module</span><span> load</span><span> cuda/12.1</span><span>  # as required by your cluster</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># 1) Install uv and create project venv</span></span>\n<span class=\"line\"><span>curl</span><span> -LsSf</span><span> https://astral.sh/uv/install.sh</span><span> |</span><span> sh</span></span>\n<span class=\"line\"><span>uv</span><span> python</span><span> install</span><span> 3.11</span></span>\n<span class=\"line\"><span>uv</span><span> python</span><span> pin</span><span> 3.11</span></span>\n<span class=\"line\"><span>uv</span><span> venv</span><span> --python</span><span> 3.11</span></span>\n<span class=\"line\"><span>source</span><span> .venv/bin/activate</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># 2) Install PyTorch CUDA wheels</span></span>\n<span class=\"line\"><span>uv</span><span> pip</span><span> install</span><span> torch</span><span> torchvision</span><span> torchaudio</span><span> --index-url</span><span> https://download.pytorch.org/whl/cu121</span></span>\n<span class=\"line\"></span>\n<span class=\"line\"><span># 3) Sanity check</span></span>\n<span class=\"line\"><span>python</span><span> -</span><span> &lt;&lt;</span><span>'PY'</span></span>\n<span class=\"line\"><span>import torch</span></span>\n<span class=\"line\"><span>print('Torch:', torch.__version__, 'CUDA:', torch.version.cuda)</span></span>\n<span class=\"line\"><span>print('GPU available:', torch.cuda.is_available())</span></span>\n<span class=\"line\"><span>print('Device:', torch.cuda.get_device_name(0))</span></span>\n<span class=\"line\"><span>PY</span></span></code></pre>\n<p>Then apply the loader/AMP/TF32/compile practices above and watch <code>nvidia-smi</code>/<code>nvtop</code> for sustained high utilization.</p>\n<hr />\n<h2 id=\"references\">References</h2>\n<ul>\n<li>uv documentation: <a href=\"https://docs.astral.sh/uv/\">docs.astral.sh/uv</a></li>\n<li>PyTorch install (CUDA wheels): <a href=\"https://pytorch.org/get-started/locally/\">pytorch.org/get-started/locally</a></li>\n<li>PyTorch performance tuning: <a href=\"https://pytorch.org/tutorials/recipes/recipes/tuning_guide.html\">pytorch.org/tutorials/recipes/recipes/tuning_guide.html</a></li>\n<li>Nsight Systems/Compute: <a href=\"https://developer.nvidia.com/\">developer.nvidia.com</a></li>\n</ul>",
            "url": "https://blog.ecitis.org/hpc-guide/",
            "title": "HPC Guide",
            "summary": "Build fast, reproducible Python and PyTorch workflows on GPU clusters, from environment setup to profiling and distributed training.",
            "date_modified": "2025-09-08T00:00:00.000Z",
            "date_published": "2025-09-08T00:00:00.000Z",
            "author": {
                "name": "Aadit Agrawal",
                "url": "https://blog.ecitis.org/"
            },
            "tags": [
                "Hardware & Systems",
                "HPC",
                "PyTorch",
                "CUDA"
            ]
        }
    ]
}